When do JSON-constrained decoding engines get JSON Schema wrong, can we predict it, and does it matter when a real LLM generates?
LLMs are increasingly asked to return structured JSON that must follow a given JSON Schema. Constrained decoding frameworks enforce this by masking, at each step, the tokens that would break the schema. JSON Schema is very expressive, though, and these engines don't always get it right. This project measures, explains and predicts their conformance errors, test by test, for three widely used frameworks: Guidance / LLGuidance, Outlines and XGrammar. A follow-up pilot then checks the extracted rules against real LLM generation.
🎓 M1 research internship at LIP6 (Sorbonne Université, Faculté des Sciences et Ingénierie), 15 June – 31 July 2026 Supervisor: Mohammed Amine Baazizi 📄 Internship report (French):
docs/internship_report_fr.pdf· Slides:docs/internship_presentation.pptxBuilt on top of the JSONSchemaBench benchmark (Geng et al., 2025)
Each test pairs a JSON Schema with a JSON instance whose validity is known. We replay it through a framework's constrained decoder and compare the framework's decision with the expected one:
| Outcome | Meaning |
|---|---|
| ✅ Correct | Valid instance accepted, or invalid instance rejected |
| UNDER-constraint | An invalid instance is accepted: the framework is too permissive |
| OVER-constraint | A valid instance is rejected: the framework is too restrictive |
| ⏱️ Timeout / crash | Grammar compilation or validation doesn't finish, so there's no usable decision |
A single global accuracy score hides which constraints fail and in which context. The goal was to go further. We wanted to find the schema and instance configurations that trigger UNDER and OVER errors, and to check whether those errors can be predicted on schemas the model has never seen.
- Test-level execution harness. It replays every schema–instance pair through each framework and logs the decision, the expected validity and any errors. It covers ≈ 76,000 framework × test decisions over the GitHub subsets of JSONSchemaBench (trivial → ultra) and Kubernetes.
- Profiling and robust long runs. It times each stage separately (grammar compilation, tokenization, token-by-token validation, and for Outlines, regex → index → guide). A per-schema timeout, resumable checkpoints and supervision let the runs survive crashes on a shared server.
- Exploratory analysis and feature engineering. I engineered hundreds of schema and instance features: numeric bounds and boundary cases, object structure and
additionalProperties, regex andpatternProperties,allOf/anyOf/oneOf/notcombinators, and schema × instance interactions. Risk is analysed with support and lift, at both the test level and the schema level. - Predictive models. One classifier per framework × error type, comparing Logistic Regression, Random Forest, HistGradientBoosting and LightGBM. Train, validation and test splits are grouped by
schema_id, so every test of a given schema lands in exactly one split. Features are then filtered by importance and domain knowledge, and the models are retrained on the reduced feature lists. - Interpretable rules. Shallow decision trees turn the risky configurations into human-readable rules.
- Out-of-distribution evaluation. The GitHub-trained models are applied without retraining to Kubernetes schemas.
- Validation with a real LLM (follow-up pilot). A small local model generates JSON under XGrammar, to check that the cases flagged by the rules are the ones where generation actually goes wrong.
| Framework | Error | Model | Precision | Recall | F1 | ROC-AUC | PR-AUC |
|---|---|---|---|---|---|---|---|
| Outlines | UNDER | LightGBM | 0.96 | 0.75 | 0.84 | 0.99 | 0.94 |
| Outlines | OVER | Random Forest | 0.89 | 0.74 | 0.81 | 0.90 | 0.84 |
| Guidance | OVER | Random Forest | 0.71 | 0.87 | 0.78 | 0.84 | 0.86 |
| XGrammar | UNDER | Random Forest | 0.71 | 0.78 | 0.74 | 0.95 | 0.83 |
| XGrammar | OVER | Random Forest | 0.78 | 0.70 | 0.74 | 0.93 | 0.81 |
Guidance produced no UNDER errors at all in the modeling data, so there was nothing to learn for that target.
A 5-fold cross-validation grouped by schema, run with earlier feature lists, gives ROC-AUC between 0.84 and 0.96 across the five models.
✅ Reproducible: reloading the saved models and re-running predict_proba on the test split gives exactly the reported scores (see the notebook).
XGrammar: UNDER rate among invalid instances and OVER rate among valid instances, for schemas that use each keyword. Keywords such as contains, if/then/else and patternProperties mostly lead to over-restrictive rejections.
Outlines UNDER (LightGBM): instance string lengths, schema complexity and numeric boundary cases carry most of the signal.
Example of an extracted rule (Outlines UNDER). When numeric constraints are nested (depth ≥ 2) and the tested number sits on or outside a bound (other than exactly the minimum), Outlines accepts the invalid instance. On held-out schemas, this rule has 90.7% precision and covers 61.5% of Outlines UNDER errors.
Shallow decision tree used to extract the Outlines UNDER rules (click to enlarge).
- Guidance OVER transfers well: F1 0.77, recall 0.86.
- XGrammar and Outlines OVER keep a very high precision (0.99–1.00), but their recall drops to 0.23–0.37. The Kubernetes schemas trigger error patterns that are rare in the GitHub data, which is a clear case of distribution shift.
- UNDER couldn't be evaluated on Kubernetes, because neither XGrammar nor Outlines produced any UNDER error there.
Everything above replays known instances through the token mask: no model generates anything. As a follow-up, a small pilot puts a real LLM in the loop (details in extension_jsonschemabench/llm_validation/).
Goal. Check that a rule saying "risk of error" corresponds to a real failure when an LLM writes the JSON, and try a rule-guided fix.
Method. Qwen2.5-0.5B-Instruct runs locally under XGrammar constrained decoding on a copy task: the model must return a known instance as JSON. The rules are not shown to the model; they act as a router that decides, before generation, whether a case is at risk.
- OVER (valid instance, Kubernetes): it is an error if the model cannot reproduce the instance, because the grammar blocked a valid output.
- UNDER (invalid instance, held-out GitHub schemas): it is an error if the output fails
jsonschemavalidation, because the grammar let an invalid output through. - For each error type, 15 cases flagged by a rule are compared with 15 unflagged controls.
Share of cases where generation goes wrong, for cases flagged by a rule and for unflagged controls.
Results.
- The rules carry over to real generation. Flagged cases fail far more often than controls: 15/15 vs 4/15 for OVER, 8/15 vs 2/15 for UNDER. The control errors are mostly copy mistakes of the small model, not framework errors.
- Rule-guided mitigation is limited with a 0.5B model. For OVER risk, generating without the grammar gives schema-valid output in 15/15 cases, but only 1 reproduces the instance exactly (13/15 when ignoring whitespace in keys). For UNDER risk, adding the schema to the prompt rarely makes the model correct the invalid value; a
jsonschemacheck still catches every remaining invalid output.
Example (OVER). The schema does not forbid extra properties, so this instance is valid:
Asked for {"pdID": "disk-12345", "fsType ": "ext4"} <- note the key "fsType " with a trailing space
Rule fired the instance has an extra property that the schema allows
XGrammar output {"pdID": "disk-12345"} <- the extra key cannot be generated
Without the grammar {"pdID": "disk-12345", "fsType": "ext4"} <- valid, but the model "fixed" the key
This is a pilot (60 cases, one small model, one framework), meant to check the approach rather than to give statistically solid numbers.
.
├── README.md
├── requirements.txt
├── setup_benchmark.sh # downloads JSONSchemaBench/MaskBench + applies the patch
├── patches/
│ └── outlines_engine_offline_vocabulary.patch
├── docs/
│ ├── figures/ # figures used in this README
│ ├── internship_report_fr.pdf # internship report (French)
│ └── internship_presentation.pptx # internship defense slides
└── extension_jsonschemabench/ # all the internship work
├── README.md # detailed documentation (French)
├── visualisation_etudes_frameworks.ipynb # main results notebook (executed)
├── visualisation_etudes_frameworks_grouped.ipynb # EDA plots per framework (executed)
├── scripts/ # runners, profiling, feature extraction, modeling, rules
├── llm_validation/ # follow-up pilot: rules vs real LLM generation (XGrammar)
├── results/per_dataset_runs/ # per framework × dataset summaries and plots
└── coverage_prediction/
├── modeles_predictifs/ # final models, modeling tables, metrics, feature importance
├── feature_documentation/ # definition of every feature
├── feature_filter/ # feature selection audit
├── rules/ # decision trees and extracted rules
├── external_eval/Kubernetes/ # GitHub models applied to Kubernetes
└── kubernetes_eval/ # models trained and tested on Kubernetes
The benchmark itself (data/, maskbench/) isn't copied into this repository. It belongs to its authors, and setup_benchmark.sh downloads it.
git clone https://github.com/wafaaBerrais/constrained-decoding-error-prediction.git
cd constrained-decoding-error-prediction
python -m venv .venv && source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -r requirements.txt
# Browse the results: no benchmark download needed
jupyter notebook extension_jsonschemabench/visualisation_etudes_frameworks.ipynbTo re-run the frameworks or rebuild the features from scratch:
./setup_benchmark.sh # fetches data/ and maskbench/ (guidance-ai/jsonschemabench)
pip install -r maskbench/requirements.txt # Guidance/LLGuidance, Outlines, XGrammar, transformers, torchEvery script is described in extension_jsonschemabench/scripts/README.md. The saved models require scikit-learn 1.9.0. Raw per-test logs (per_test_results.jsonl) aren't included because of their size, but the run_* scripts regenerate them.
Python · scikit-learn · LightGBM · NumPy · jsonschema · Jupyter · Guidance/LLGuidance · Outlines · XGrammar · Hugging Face Transformers
This work extends JSONSchemaBench and MaskBench:
Saibo Geng et al. JSONSchemaBench: A Rigorous Benchmark of Structured Outputs for Language Models. arXiv:2501.10868, 2025. [paper] [code]
Thanks to Mohammed Amine Baazizi for supervising this internship at LIP6.
Wafaa Berrais, Master's student in Artificial Intelligence, Machine Learning & Data GitHub: @wafaaBerrais