Skip to content

[Bug]: Raw bytes are passed to hexdec() in two places, corrupting text and emitting E_DEPRECATED per character group #534

Description

@wouterschreurs

Version: v3.3.0 (still present in v3.4.0) · PHP: 8.4

Two call sites hand a non-hex string to hexdec(). Each produces wrong output
and raises one E_DEPRECATED ("Invalid characters passed for attempted
conversion, these have been ignored") per group, so a single document can emit
thousands of identical notices in well under a second.


1. EncodingNameValue::decodeString() receives raw bytes for Identity-H/Identity-V

TextSegment::decode() has two sibling branches that disagree:

// src/Document/ContentStream/PositionedText/TextSegment/TextSegment.php
if ($toUnicodeCMap !== null) {
    return $toUnicodeCMap->textToUnicode(bin2hex($binaryString));   // line 40 — hex
}

if ($encoding !== null) {
    return $encoding->decodeString($binaryString);                  // line 44 — raw bytes
}

For Identity-H/Identity-V, decodeString() forwards to
ToUnicodeCMap::textToUnicode(), which chunks the string and calls hexdec()
on each chunk — it expects hex, exactly as line 40 supplies. Line 44 does not,
so this fires whenever a font resolves no ToUnicode CMap (no /ToUnicode and no
descendant-font /CIDSystemInfo).

Reproduce:

use PrinsFrank\PdfParser\Document\Dictionary\DictionaryValue\Name\EncodingNameValue;

set_error_handler(fn (int $n, string $s): bool => !printf("%s\n", $s));
EncodingNameValue::IdentityH->decodeString('Invoice 2026-0912 reference ABCDEF');
// 10 × Deprecated: Invalid characters passed for attempted conversion, these have been ignored

Expected: the same output line 40 produces for the same bytes.
Actual: garbage, plus one deprecation per chunk.


2. TextSegment::getCodePoints() reads the un-normalised hex string

// src/Document/ContentStream/PositionedText/TextSegment/TextSegment.php:57-59
} elseif (str_starts_with($this->textString->textStringValue, '<') && str_ends_with(..., '>')) {
    foreach (str_split(substr($this->textString->textStringValue, 1, -1), 4) as $char) {
        $codePoints[] = is_int($codePoint = hexdec($char)) ? $codePoint : throw new ParseFailureException();

TextStringValue::getBinaryString() already strips PDF whitespace
([\x00\x09\x0A\x0C\x0D\x20]), validates the hex and pads odd lengths — but
getCodePoints() splits the raw value instead. PDF producers routinely wrap long
hex strings across lines, so whitespace lands inside a 4-character group. PHP
skips leading whitespace in hexdec() silently but deprecates on whitespace
inside the group, so the symptom depends on alignment while the wrong result
does not:

use PrinsFrank\PdfParser\Document\ContentStream\PositionedText\TextSegment\TextSegment;
use PrinsFrank\PdfParser\Document\Dictionary\DictionaryValue\TextString\TextStringValue;

(new TextSegment(new TextStringValue('<00480065>'),  null))->getCodePoints(); // [72, 101]  ✅
(new TextSegment(new TextStringValue('<0048 0065>'), null))->getCodePoints(); // [72, 6, 5] ❌ no notice
(new TextSegment(new TextStringValue('<00 480065>'), null))->getCodePoints(); // [4, 32774, 5] ❌ + deprecation

Those code points feed getWidthForChars() → glyph advance → the word-break
threshold, so this also affects where spaces are inserted in extracted text.


Suggested fix

Both call sites become consistent with how the library already handles the same
data elsewhere:

--- a/src/Document/Dictionary/DictionaryValue/Name/EncodingNameValue.php
+++ b/src/Document/Dictionary/DictionaryValue/Name/EncodingNameValue.php
@@ -17,7 +17,7 @@
     public function decodeString(string $characterGroup): string {
         return match ($this) {
             self::IdentityH,
-            self::IdentityV => (new Identity0())->getToUnicodeCMap()->textToUnicode($characterGroup),
+            self::IdentityV => (new Identity0())->getToUnicodeCMap()->textToUnicode(bin2hex($characterGroup)),
             self::WinAnsiEncoding => WinAnsi::textToUnicode($characterGroup),
             self::MacRomanEncoding => MacRoman::textToUnicode($characterGroup),
             default => throw new ParseFailureException(sprintf('Unsupported encoding %s', $this->name)),
--- a/src/Document/ContentStream/PositionedText/TextSegment/TextSegment.php
+++ b/src/Document/ContentStream/PositionedText/TextSegment/TextSegment.php
@@ -55,7 +55,7 @@
                 $codePoints[] = ord($char);
             }
         } elseif (str_starts_with($this->textString->textStringValue, '<') && str_ends_with($this->textString->textStringValue, '>')) {
-            foreach (str_split(substr($this->textString->textStringValue, 1, -1), 4) as $char) {
+            foreach (str_split(bin2hex($this->textString->getBinaryString()), 4) as $char) {
                 $codePoints[] = is_int($codePoint = hexdec($char)) ? $codePoint : throw new ParseFailureException();
             }
         } else {

With both applied, all three getCodePoints() cases above return [72, 101],
the Identity-H path returns what the CMap branch returns, and the deprecations
stop.

One intentional behaviour change: an odd-length hex string previously
decoded unpadded (<048> → 72) and now goes through getBinaryString()'s
trailing-zero padding (<048> → 1152). That is what §7.3.4.3 of the PDF spec
requires and what the rest of the library already does, so this aligns the two —
flagging it in case it is load-bearing somewhere.

Happy to open a PR with this plus tests if the approach looks right.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions