Version: v3.3.0 (still present in v3.4.0) · PHP: 8.4
Two call sites hand a non-hex string to hexdec(). Each produces wrong output
and raises one E_DEPRECATED ("Invalid characters passed for attempted
conversion, these have been ignored") per group, so a single document can emit
thousands of identical notices in well under a second.
1. EncodingNameValue::decodeString() receives raw bytes for Identity-H/Identity-V
TextSegment::decode() has two sibling branches that disagree:
// src/Document/ContentStream/PositionedText/TextSegment/TextSegment.php
if ($toUnicodeCMap !== null) {
return $toUnicodeCMap->textToUnicode(bin2hex($binaryString)); // line 40 — hex
}
if ($encoding !== null) {
return $encoding->decodeString($binaryString); // line 44 — raw bytes
}
For Identity-H/Identity-V, decodeString() forwards to
ToUnicodeCMap::textToUnicode(), which chunks the string and calls hexdec()
on each chunk — it expects hex, exactly as line 40 supplies. Line 44 does not,
so this fires whenever a font resolves no ToUnicode CMap (no /ToUnicode and no
descendant-font /CIDSystemInfo).
Reproduce:
use PrinsFrank\PdfParser\Document\Dictionary\DictionaryValue\Name\EncodingNameValue;
set_error_handler(fn (int $n, string $s): bool => !printf("%s\n", $s));
EncodingNameValue::IdentityH->decodeString('Invoice 2026-0912 reference ABCDEF');
// 10 × Deprecated: Invalid characters passed for attempted conversion, these have been ignored
Expected: the same output line 40 produces for the same bytes.
Actual: garbage, plus one deprecation per chunk.
2. TextSegment::getCodePoints() reads the un-normalised hex string
// src/Document/ContentStream/PositionedText/TextSegment/TextSegment.php:57-59
} elseif (str_starts_with($this->textString->textStringValue, '<') && str_ends_with(..., '>')) {
foreach (str_split(substr($this->textString->textStringValue, 1, -1), 4) as $char) {
$codePoints[] = is_int($codePoint = hexdec($char)) ? $codePoint : throw new ParseFailureException();
TextStringValue::getBinaryString() already strips PDF whitespace
([\x00\x09\x0A\x0C\x0D\x20]), validates the hex and pads odd lengths — but
getCodePoints() splits the raw value instead. PDF producers routinely wrap long
hex strings across lines, so whitespace lands inside a 4-character group. PHP
skips leading whitespace in hexdec() silently but deprecates on whitespace
inside the group, so the symptom depends on alignment while the wrong result
does not:
use PrinsFrank\PdfParser\Document\ContentStream\PositionedText\TextSegment\TextSegment;
use PrinsFrank\PdfParser\Document\Dictionary\DictionaryValue\TextString\TextStringValue;
(new TextSegment(new TextStringValue('<00480065>'), null))->getCodePoints(); // [72, 101] ✅
(new TextSegment(new TextStringValue('<0048 0065>'), null))->getCodePoints(); // [72, 6, 5] ❌ no notice
(new TextSegment(new TextStringValue('<00 480065>'), null))->getCodePoints(); // [4, 32774, 5] ❌ + deprecation
Those code points feed getWidthForChars() → glyph advance → the word-break
threshold, so this also affects where spaces are inserted in extracted text.
Suggested fix
Both call sites become consistent with how the library already handles the same
data elsewhere:
--- a/src/Document/Dictionary/DictionaryValue/Name/EncodingNameValue.php
+++ b/src/Document/Dictionary/DictionaryValue/Name/EncodingNameValue.php
@@ -17,7 +17,7 @@
public function decodeString(string $characterGroup): string {
return match ($this) {
self::IdentityH,
- self::IdentityV => (new Identity0())->getToUnicodeCMap()->textToUnicode($characterGroup),
+ self::IdentityV => (new Identity0())->getToUnicodeCMap()->textToUnicode(bin2hex($characterGroup)),
self::WinAnsiEncoding => WinAnsi::textToUnicode($characterGroup),
self::MacRomanEncoding => MacRoman::textToUnicode($characterGroup),
default => throw new ParseFailureException(sprintf('Unsupported encoding %s', $this->name)),
--- a/src/Document/ContentStream/PositionedText/TextSegment/TextSegment.php
+++ b/src/Document/ContentStream/PositionedText/TextSegment/TextSegment.php
@@ -55,7 +55,7 @@
$codePoints[] = ord($char);
}
} elseif (str_starts_with($this->textString->textStringValue, '<') && str_ends_with($this->textString->textStringValue, '>')) {
- foreach (str_split(substr($this->textString->textStringValue, 1, -1), 4) as $char) {
+ foreach (str_split(bin2hex($this->textString->getBinaryString()), 4) as $char) {
$codePoints[] = is_int($codePoint = hexdec($char)) ? $codePoint : throw new ParseFailureException();
}
} else {
With both applied, all three getCodePoints() cases above return [72, 101],
the Identity-H path returns what the CMap branch returns, and the deprecations
stop.
One intentional behaviour change: an odd-length hex string previously
decoded unpadded (<048> → 72) and now goes through getBinaryString()'s
trailing-zero padding (<048> → 1152). That is what §7.3.4.3 of the PDF spec
requires and what the rest of the library already does, so this aligns the two —
flagging it in case it is load-bearing somewhere.
Happy to open a PR with this plus tests if the approach looks right.
Version: v3.3.0 (still present in v3.4.0) · PHP: 8.4
Two call sites hand a non-hex string to
hexdec(). Each produces wrong outputand raises one
E_DEPRECATED("Invalid characters passed for attemptedconversion, these have been ignored") per group, so a single document can emit
thousands of identical notices in well under a second.
1.
EncodingNameValue::decodeString()receives raw bytes forIdentity-H/Identity-VTextSegment::decode()has two sibling branches that disagree:For
Identity-H/Identity-V,decodeString()forwards toToUnicodeCMap::textToUnicode(), which chunks the string and callshexdec()on each chunk — it expects hex, exactly as line 40 supplies. Line 44 does not,
so this fires whenever a font resolves no ToUnicode CMap (no
/ToUnicodeand nodescendant-font
/CIDSystemInfo).Reproduce:
Expected: the same output line 40 produces for the same bytes.
Actual: garbage, plus one deprecation per chunk.
2.
TextSegment::getCodePoints()reads the un-normalised hex stringTextStringValue::getBinaryString()already strips PDF whitespace(
[\x00\x09\x0A\x0C\x0D\x20]), validates the hex and pads odd lengths — butgetCodePoints()splits the raw value instead. PDF producers routinely wrap longhex strings across lines, so whitespace lands inside a 4-character group. PHP
skips leading whitespace in
hexdec()silently but deprecates on whitespaceinside the group, so the symptom depends on alignment while the wrong result
does not:
Those code points feed
getWidthForChars()→ glyph advance → the word-breakthreshold, so this also affects where spaces are inserted in extracted text.
Suggested fix
Both call sites become consistent with how the library already handles the same
data elsewhere:
With both applied, all three
getCodePoints()cases above return[72, 101],the Identity-H path returns what the CMap branch returns, and the deprecations
stop.
One intentional behaviour change: an odd-length hex string previously
decoded unpadded (
<048>→72) and now goes throughgetBinaryString()'strailing-zero padding (
<048>→1152). That is what §7.3.4.3 of the PDF specrequires and what the rest of the library already does, so this aligns the two —
flagging it in case it is load-bearing somewhere.
Happy to open a PR with this plus tests if the approach looks right.