Skip to content

[Bug]: Information dictionary fails on common dates (D:…Z00'00') and /Type /Info, plus a regression from #538 #541

Description

@mgralikowski

Is there an existing issue for this?

  • I have searched the existing issues and found no similar reports

#212 was about dates without a time zone and is fixed; this is a different date shape.

Are you using the latest version of this package?

  • The issue I'm reporting exists in the latest release

Tested on v3.4.0 and on main (527c391); main adds one regression, see (4).

Can other PDF readers read the file?

  • The PDF I'm trying to read opens correctly in at least one other PDF reader

pdfinfo (poppler) and smalot/pdfparser read all of them, dates included.

Context

We were gently encouraged in smalot/pdfparser#845 to switch to this library instead of patching smalot 😊, so we did our homework and ran it on 104 PDFs (smalot's samples/, plus files from our e-signature platform), compared with pdfinfo. The good news first: page counts are 103/103 and last-page MediaBox is 102/103, better than smalot. Nice work!

The not-so-good news: the information dictionary fails on 14 of 103 files, mostly real-world files from macOS (Quartz), pdf-lib/pyHanko and Oracle BI Publisher. For us that is a blocker, because we read Title/Producer/dates when validating uploads.

When running this snippet

This builds minimal one-page PDFs in memory (the three small attachments below are its output):

use PrinsFrank\PdfParser\PdfParser;

function pdfWithInfo(string $info): string {
    $objects = [
        1 => '<< /Type /Catalog /Pages 2 0 R >>',
        2 => '<< /Type /Pages /Kids [3 0 R] /Count 1 >>',
        3 => '<< /Type /Page /Parent 2 0 R /MediaBox [0 0 595 842] /Resources << >> >>',
        4 => $info,
    ];
    $pdf = "%PDF-1.4\n";
    $offsets = [];
    foreach ($objects as $number => $body) {
        $offsets[$number] = strlen($pdf);
        $pdf .= "{$number} 0 obj\n{$body}\nendobj\n";
    }
    $xref = strlen($pdf);
    $pdf .= "xref\n0 5\n0000000000 65535 f \n";
    foreach ($offsets as $offset) {
        $pdf .= sprintf("%010d 00000 n \n", $offset);
    }

    return $pdf . "trailer\n<< /Size 5 /Root 1 0 R /Info 4 0 R >>\nstartxref\n{$xref}\n%%EOF\n";
}

$cases = [
    'Z00\'00\' in CreationDate' => "<< /Title (Sample) /CreationDate (D:20250221091448Z00'00') >>",
    'Z00\'00\' in ModDate'      => "<< /Title (Sample) /ModDate (D:20250221091448Z00'00') >>",
    '/Type /Info'              => '<< /Type /Info /Title (Sample) >>',
];
foreach ($cases as $label => $info) {
    $info = (new PdfParser())->parseString(pdfWithInfo($info))->getInformationDictionary();
    var_dump($info?->getTitle(), $info?->getCreationDate(), $info?->getModificationDate());
}

I run into the following issue/exception

Attached. The three small ones are the output of the snippet above; each fails on v3.4.0 and main, while pdfinfo reads it fine:

File Case
info-moddate-z00.pdf Z00'00' in ModDate: getInformationDictionary() throws, see (1) and (3)
info-creationdate-z00.pdf Z00'00' in CreationDate: getCreationDate() throws, see (1)
info-type-info.pdf /Type /Info, see (2)
IncrementalUpdateObjectStream.pdf a real-world file saved on macOS, case (1)

For the indirect-object case in (1) and the regression in (4), use Issue391.pdf linked below.

1. D:YYYYMMDDHHmmSSZ00'00' is rejected. This is what macOS Quartz (and tools built on it) writes, so it shows up in a lot of files. Depending on where the value sits, the symptom differs:

Where the date is Result
ModDate, inline getInformationDictionary() throws ParseFailureException: Value "(D:20250221091448Z00'00')" for dictionary key ModDate could not be parsed to a valid value type, so Title, Producer and the rest become unreachable too
CreationDate, inline getCreationDate() throws InvalidArgumentException: Expected value with value CreationDate to be of type …DateValue
an indirect object (/CreationDate 261 0 R) getCreationDate() throws ParseFailureException: Unable to parse content "(D:20111228211201Z00'00')" of referenced object 261

Public files: Issue608.pdf (inline), Issue391.pdf (indirect object) from smalot's samples, and IncrementalUpdateObjectStream.pdf.

The cause seems to be in DateValue::fromValue(): str_replace("'", '', …) turns Z00'00' into Z0000, which the P format does not accept, so it returns null. Z alone (D:20250221091448Z) and +01'00' both parse fine.

2. /Type /Info in the information dictionary makes getInformationDictionary() throw Value "/Info" for dictionary key Type could not be parsed to a valid value type. The spec doesn't define a /Type for this dictionary, but some producers write one (we have it from Oracle BI Publisher), and other readers ignore it.

3. One bad field takes down the whole dictionary (case 1, first row; case 2). I'd expect a value that can't be parsed to come back as null from its getter, or at least to throw only from that getter, rather than making every other field unreachable.

4. Regression on main since #538 (527c391, not released yet): on Issue391.pdf, getTitle() and getProducer() now throw ParseFailureException: Unrecognized format (Microsoft Word - GVTW70SPAH1R0_20101210_REV_A.doc). On v3.4.0 and on ea963a8 (the commit before) they return the values. I bisected the 7 commits since v3.4.0 to find it.

Possible improvement

A small suggestion, in case it helps:

  • parse dates with a lenient pattern that follows ISO 32000-1, 7.9.4 and tolerates common variants, for example ^D:(\d{4})(\d{2})?(\d{2})?(\d{2})?(\d{2})?(\d{2})?(?:([+\-Z])(\d{2})?'?(\d{2})?'?)?$, with Z followed by an optional 00'00' treated as UTC. That also covers truncated dates such as D:2025 or D:202502, which the spec allows;
  • when a key's value doesn't match the expected type, return null for that key instead of failing the whole dictionary.

Happy to open a PR for the date part if you'd like one.

Do you allow attachment files to be used in tests to prevent regressions?

  • Yes, for the four attached files and the snippet above. The two Issue*.pdf files come from smalot/pdfparser's sample set, so their terms apply.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions