A native Normalizer class + normalizer_*() functions for when intl isn't available, built on utf8proc.
php-norm implements ext-intl's Normalizer class and its
normalizer_normalize() / normalizer_is_normalized() global function
aliases in C, on top of utf8proc
— a small, fast, well-tested Unicode library — instead of intl's own ICU
dependency (which pulls in a ~30MB icu.dat and isn't built by
php-wasm-compiler
today).
It exists to replace
kirigami/php-prepros's
pure-PHP Normalizer fallback (a port of
symfony/polyfill-intl-normalizer, with its Unicode decomposition tables
embedded directly in the file) with a real, compiled extension:
Normalizer::normalize("\u{00E9}"); // "é" (NFC — already composed, returned as-is)
Normalizer::normalize("e\u{0301}", Normalizer::NFC); // "é" (composes e + combining acute)
Normalizer::isNormalized("é", Normalizer::NFC); // trueThe class and its constants (NONE, FORM_D/NFD, FORM_KD/NFKD,
FORM_C/NFC, FORM_KC/NFKC) match ext-intl's real numeric values —
this is meant as a genuine drop-in, not a namespaced workalike. If a
Normalizer class (or normalizer_normalize()/normalizer_is_normalized()
function) already exists — e.g. because intl is also loaded —
php-norm detects that at startup and registers nothing on top of it,
rather than crashing or shadowing it (see docs/DECISIONS.md, decision 3, for how).
Part of the Kirigami project ecosystem.
Builds natively and under Emscripten: statically linked into
@kirigami/php-wasm
by php-wasm-compiler,
where kirigami/php-prepros uses its Normalizer. Verified against real
Unicode test vectors and alongside ext/intl; see
docs/STATUS.md.
class Normalizer
{
const NONE = 1;
const FORM_D = 2; const NFD = 2;
const FORM_KD = 3; const NFKD = 3;
const FORM_C = 4; const NFC = 4;
const FORM_KC = 5; const NFKC = 5;
public static function normalize(string $string, int $form = self::FORM_C): string|false;
public static function isNormalized(string $string, int $form = self::FORM_C): bool;
}
function normalizer_normalize(string $string, int $form = Normalizer::FORM_C): string|false;
function normalizer_is_normalized(string $string, int $form = Normalizer::FORM_C): bool;normalize() returns false for a $string that isn't valid UTF-8, and
throws ValueError for an invalid $form. isNormalized() returns
false for invalid UTF-8 rather than throwing (matching ext-intl).
Normalizer::NONE is a no-op passthrough in both.
getRawDecomposition() (ext-intl, PHP 7.3+) isn't implemented — it isn't
used anywhere in kirigami/php-prepros's current fallback.
php-norm declares intl as an optional module dependency purely to
control startup order: if intl is compiled in, its MINIT runs first,
so by the time php-norm's own MINIT runs, it can see the real
Normalizer class and normalizer_*() functions already registered — and
skips registering its own, deferring entirely. This isn't intl-specific:
the same checks (class_exists/function-table lookups, not "is intl
loaded") would defer to any extension that got there first with the same
names. phpinfo() reports which case applies.
Native build (fast iteration, no Docker/Emscripten):
cd vendor/build && bash stage.sh # downloads utf8proc
cd ../..
phpize
./configure --enable-norm
makevendor/ is gitignored and rebuilt from source every time — nothing
vendored is hand-edited. There is no Emscripten/WASM build yet (see
Status).
- PHP
>= 8.5 - A C compiler,
phpize
- utf8proc — the C library this extension wraps
- kirigami/php-prepros —
the userland
Normalizerfallback this extension replaces - php-wasm-compiler —
builds the
php.wasmruntime this extension will eventually target - php-mdhtml / php-jsonk — sibling extensions built the same way (native PECL first, WASM second)
GPL-2.0-or-later. See LICENSE for the full text.
vendor/utf8proc (downloaded, not committed) is MIT-licensed — see
utf8proc's own LICENSE.md.
Maxime Larrivée-Roy