Is there an existing issue for this?
#212 was about dates without a time zone and is fixed; this is a different date shape.
Are you using the latest version of this package?
Tested on v3.4.0 and on main (527c391); main adds one regression, see (4).
Can other PDF readers read the file?
pdfinfo (poppler) and smalot/pdfparser read all of them, dates included.
Context
We were gently encouraged in smalot/pdfparser#845 to switch to this library instead of patching smalot 😊, so we did our homework and ran it on 104 PDFs (smalot's samples/, plus files from our e-signature platform), compared with pdfinfo. The good news first: page counts are 103/103 and last-page MediaBox is 102/103, better than smalot. Nice work!
The not-so-good news: the information dictionary fails on 14 of 103 files, mostly real-world files from macOS (Quartz), pdf-lib/pyHanko and Oracle BI Publisher. For us that is a blocker, because we read Title/Producer/dates when validating uploads.
When running this snippet
This builds minimal one-page PDFs in memory (the three small attachments below are its output):
use PrinsFrank\PdfParser\PdfParser;
function pdfWithInfo(string $info): string {
$objects = [
1 => '<< /Type /Catalog /Pages 2 0 R >>',
2 => '<< /Type /Pages /Kids [3 0 R] /Count 1 >>',
3 => '<< /Type /Page /Parent 2 0 R /MediaBox [0 0 595 842] /Resources << >> >>',
4 => $info,
];
$pdf = "%PDF-1.4\n";
$offsets = [];
foreach ($objects as $number => $body) {
$offsets[$number] = strlen($pdf);
$pdf .= "{$number} 0 obj\n{$body}\nendobj\n";
}
$xref = strlen($pdf);
$pdf .= "xref\n0 5\n0000000000 65535 f \n";
foreach ($offsets as $offset) {
$pdf .= sprintf("%010d 00000 n \n", $offset);
}
return $pdf . "trailer\n<< /Size 5 /Root 1 0 R /Info 4 0 R >>\nstartxref\n{$xref}\n%%EOF\n";
}
$cases = [
'Z00\'00\' in CreationDate' => "<< /Title (Sample) /CreationDate (D:20250221091448Z00'00') >>",
'Z00\'00\' in ModDate' => "<< /Title (Sample) /ModDate (D:20250221091448Z00'00') >>",
'/Type /Info' => '<< /Type /Info /Title (Sample) >>',
];
foreach ($cases as $label => $info) {
$info = (new PdfParser())->parseString(pdfWithInfo($info))->getInformationDictionary();
var_dump($info?->getTitle(), $info?->getCreationDate(), $info?->getModificationDate());
}
I run into the following issue/exception
Attached. The three small ones are the output of the snippet above; each fails on v3.4.0 and main, while pdfinfo reads it fine:
For the indirect-object case in (1) and the regression in (4), use Issue391.pdf linked below.
1. D:YYYYMMDDHHmmSSZ00'00' is rejected. This is what macOS Quartz (and tools built on it) writes, so it shows up in a lot of files. Depending on where the value sits, the symptom differs:
| Where the date is |
Result |
ModDate, inline |
getInformationDictionary() throws ParseFailureException: Value "(D:20250221091448Z00'00')" for dictionary key ModDate could not be parsed to a valid value type, so Title, Producer and the rest become unreachable too |
CreationDate, inline |
getCreationDate() throws InvalidArgumentException: Expected value with value CreationDate to be of type …DateValue |
an indirect object (/CreationDate 261 0 R) |
getCreationDate() throws ParseFailureException: Unable to parse content "(D:20111228211201Z00'00')" of referenced object 261 |
Public files: Issue608.pdf (inline), Issue391.pdf (indirect object) from smalot's samples, and IncrementalUpdateObjectStream.pdf.
The cause seems to be in DateValue::fromValue(): str_replace("'", '', …) turns Z00'00' into Z0000, which the P format does not accept, so it returns null. Z alone (D:20250221091448Z) and +01'00' both parse fine.
2. /Type /Info in the information dictionary makes getInformationDictionary() throw Value "/Info" for dictionary key Type could not be parsed to a valid value type. The spec doesn't define a /Type for this dictionary, but some producers write one (we have it from Oracle BI Publisher), and other readers ignore it.
3. One bad field takes down the whole dictionary (case 1, first row; case 2). I'd expect a value that can't be parsed to come back as null from its getter, or at least to throw only from that getter, rather than making every other field unreachable.
4. Regression on main since #538 (527c391, not released yet): on Issue391.pdf, getTitle() and getProducer() now throw ParseFailureException: Unrecognized format (Microsoft Word - GVTW70SPAH1R0_20101210_REV_A.doc). On v3.4.0 and on ea963a8 (the commit before) they return the values. I bisected the 7 commits since v3.4.0 to find it.
Possible improvement
A small suggestion, in case it helps:
- parse dates with a lenient pattern that follows ISO 32000-1, 7.9.4 and tolerates common variants, for example
^D:(\d{4})(\d{2})?(\d{2})?(\d{2})?(\d{2})?(\d{2})?(?:([+\-Z])(\d{2})?'?(\d{2})?'?)?$, with Z followed by an optional 00'00' treated as UTC. That also covers truncated dates such as D:2025 or D:202502, which the spec allows;
- when a key's value doesn't match the expected type, return
null for that key instead of failing the whole dictionary.
Happy to open a PR for the date part if you'd like one.
Do you allow attachment files to be used in tests to prevent regressions?
Is there an existing issue for this?
#212 was about dates without a time zone and is fixed; this is a different date shape.
Are you using the latest version of this package?
Tested on
v3.4.0and onmain(527c391);mainadds one regression, see (4).Can other PDF readers read the file?
pdfinfo(poppler) and smalot/pdfparser read all of them, dates included.Context
We were gently encouraged in smalot/pdfparser#845 to switch to this library instead of patching smalot 😊, so we did our homework and ran it on 104 PDFs (smalot's
samples/, plus files from our e-signature platform), compared withpdfinfo. The good news first: page counts are 103/103 and last-pageMediaBoxis 102/103, better than smalot. Nice work!The not-so-good news: the information dictionary fails on 14 of 103 files, mostly real-world files from macOS (Quartz), pdf-lib/pyHanko and Oracle BI Publisher. For us that is a blocker, because we read Title/Producer/dates when validating uploads.
When running this snippet
This builds minimal one-page PDFs in memory (the three small attachments below are its output):
I run into the following issue/exception
Attached. The three small ones are the output of the snippet above; each fails on
v3.4.0andmain, whilepdfinforeads it fine:Z00'00'inModDate:getInformationDictionary()throws, see (1) and (3)Z00'00'inCreationDate:getCreationDate()throws, see (1)/Type /Info, see (2)For the indirect-object case in (1) and the regression in (4), use
Issue391.pdflinked below.1.
D:YYYYMMDDHHmmSSZ00'00'is rejected. This is what macOS Quartz (and tools built on it) writes, so it shows up in a lot of files. Depending on where the value sits, the symptom differs:ModDate, inlinegetInformationDictionary()throwsParseFailureException: Value "(D:20250221091448Z00'00')" for dictionary key ModDate could not be parsed to a valid value type, so Title, Producer and the rest become unreachable tooCreationDate, inlinegetCreationDate()throwsInvalidArgumentException: Expected value with value CreationDate to be of type …DateValue/CreationDate 261 0 R)getCreationDate()throwsParseFailureException: Unable to parse content "(D:20111228211201Z00'00')" of referenced object 261Public files:
Issue608.pdf(inline),Issue391.pdf(indirect object) from smalot's samples, andIncrementalUpdateObjectStream.pdf.The cause seems to be in
DateValue::fromValue():str_replace("'", '', …)turnsZ00'00'intoZ0000, which thePformat does not accept, so it returnsnull.Zalone (D:20250221091448Z) and+01'00'both parse fine.2.
/Type /Infoin the information dictionary makesgetInformationDictionary()throwValue "/Info" for dictionary key Type could not be parsed to a valid value type. The spec doesn't define a/Typefor this dictionary, but some producers write one (we have it from Oracle BI Publisher), and other readers ignore it.3. One bad field takes down the whole dictionary (case 1, first row; case 2). I'd expect a value that can't be parsed to come back as
nullfrom its getter, or at least to throw only from that getter, rather than making every other field unreachable.4. Regression on
mainsince #538 (527c391, not released yet): onIssue391.pdf,getTitle()andgetProducer()now throwParseFailureException: Unrecognized format (Microsoft Word - GVTW70SPAH1R0_20101210_REV_A.doc). Onv3.4.0and onea963a8(the commit before) they return the values. I bisected the 7 commits sincev3.4.0to find it.Possible improvement
A small suggestion, in case it helps:
^D:(\d{4})(\d{2})?(\d{2})?(\d{2})?(\d{2})?(\d{2})?(?:([+\-Z])(\d{2})?'?(\d{2})?'?)?$, withZfollowed by an optional00'00'treated as UTC. That also covers truncated dates such asD:2025orD:202502, which the spec allows;nullfor that key instead of failing the whole dictionary.Happy to open a PR for the date part if you'd like one.
Do you allow attachment files to be used in tests to prevent regressions?
Issue*.pdffiles come from smalot/pdfparser's sample set, so their terms apply.