- Deserializer is an advanced Abstract Syntax Tree (AST) static analysis engine designed to identify insecure object reconstruction and state persistence sinks across the Python ecosystem, with strong focus on AI, LLM, Robotics, Data Science, Machine Learning and Deep Learning fields; but not limited.
- Deserializer was designed to scan and analyse any king of project based on Python source code, like apps, scripts, frameworks, libraries, Python pip packages, and much more.
Far beyond a traditional scanner, it provides a generic, high-performance framework for auditing over 120 libraries and formats—including YAML, Msgpack, CBOR, and custom JSON hooks—where traditional trust in "safe" serialization hides critical logic-based RCE vectors. By resolving imports, aliases, and complex dotted attributes, the tool serves as a high-fidelity signal amplifier that prioritizes dangerous code paths in modern distributed architectures and AI/ML repositories.
The project operates through a modular, multi-phased workflow that transitions from raw detection to deep technical audit. Following the initial high-velocity SAST scan, the ecosystem leverages specialized relationship mappers to trace execution flows and result processors to generate detailed security reports. This systematic approach ensures that every finding is contextualized within the application's broader architecture, transforming high-volume telemetry into actionable research assets and structured milestones that simplify the mapping of infrastructure-level attack surfaces.
At its most advanced tier, Deserializer integrates an autonomous AI Security Agent (Phase 4), explicitly designed to navigate "self-command-injection" limitations and synthesize functional reproduction guides, based on HuggingFace inference API, Local LLM (like llama.cpp), or OpenAI API compatible.
As the project evolves, it continues to define the frontier of automated vulnerability research by bridging the gap between abstract syntax tree static analysis, relationship mapping, reporting, and functional exploit development based on documented research.
Deserializer directly supports security research by locating RCE and Insecure Deserialization paths across various large-scale AI, Robotics, and Data Science projects and environments, like Genesis World (v0.2.1), MuJoCo (v3.7.0), LeRobot (v0.5.1), Brax (v0.14.2), TensorFlow (v2.21.0), LangGraph (v1.1.6), VibeVoice (v0.0.1), Hugging Face Hub (v1.11.0), PyGlove (v0.4.5), and many others.
Its capabilities have directly powered the discovery of critical vulnerabilities in industry-leading frameworks, proving its efficacy in auditing complex MLOps and agentic AI environments.
The project is structured into four distinct phases, each designed to shift the analysis from high-volume automated telemetry to deep, functional security research:
| Phase | Title | Tooling / Engine | Objective |
|---|---|---|---|
| 1 | High-Velocity Detection | deserializer.py (Triple-Pass) |
Perform massive-scale SAST to identify potential deserialization sinks. |
| 2 | Relationship Mapping | Result Processors / Mappers | Contextualize findings by tracing execution flows and component interdependencies. |
| 3 | Technical Synthesis | Research Documentation | Formalize findings into technical writeups, mapping infrastructure-level attack surfaces. |
| 4 | Autonomous AI Agent | AI Security Agent | Automate 0-day discovery and generate functional reproduction guides/exploits using HuggingFace inference API, Local LLM (like llama.cpp), or OpenAI API compatible. |
A high-level map of the repository's organization and the technical purpose of each specialized directory:
agent/
Implementation of the Phase 4 AI Security Agent, including orchestration logic for deep analysis, reproduction guide synthesis, and LLM interaction.docs/
Central repository for project documentation, architectural diagrams, branding assets, and supporting high-level technical materials.exploit_development/
Dedicated space for engineering functional exploits, researching deserialization primitives (hooks/methods), and documenting cross-platform orchestration techniques.modules/
Core logic components, auxiliary utility functions, and scanner extensions that power the detection and relationship mapping phases of the toolchain.reports/
Storage for structured security analysis outputs, telemetry, and audit results generated by the scanner during multi-phased project evaluations.research/
A collection of high-fidelity vulnerability writeups, validated Proof-of-Concept (PoC) scripts, and deep audits conducted on modern AI/ML frameworks.templates/
Standardized reporting and research templates used to maintain technical consistency across vulnerability writeups and correlation reports.
This scanner features a high-performance parallel execution engine built on Python's concurrent.futures.ProcessPoolExecutor. It is designed to scale across all available CPU cores (controllable via the -j or --concurrency flag), making it capable of scanning tens of thousands of files in seconds.
- Non-Blocking UI: Features a stable, docked progress bar with a one-line gap for clean results presentation.
- Native Signal Handling: On Windows, it utilizes a native
SetConsoleCtrlHandlerviactypesto ensure thatCtrl+Cis 100% responsive, even during heavy processing. - Silent Tracebacks: Worker processes are silented to ensure that interrupts and internal errors don't clutter the technical output.
- Native Windows UI Support: Automatically enables Virtual Terminal Processing via
ctypesfor native ANSI color support in modern CMD and PowerShell environments.
Warning
Performance Warning: When analyzing extremely large or complex files (e.g., over 1MB, 2MB, or 3MB in size), the tool may experience significant slowdowns or appear "stuck" while parsing deep AST trees. If you encounter such bottlenecks, consider using the --timeout (to skip slow files) and --max-size (to skip huge files) flags to maintain scan velocity.
- Triple-Pass Scanning Engine: Implements an "unbreakable" multi-tiered approach:
- Pass 1 (AST): High-fidelity grammatical decomposition for precise analysis.
- Pass 2 (Token Fallback): Automated fallback to a stream-based tokenizer when encountering syntax errors or unparsable sections.
- Pass 3 (Regex Emergency): A final pattern-matching layer ensuring coverage in extremely hostile or fragmented source files.
- Native Template Neutralization: Built-in support for sanitizing Jinja2 and Mako tags, allowing the scanner to process web templates and code-generation files without choking on non-Python syntax.
- Call & Reference Detection: Identifies not just direct execution sinks (e.g.,
pickle.loads()) but also dangerous function references (e.g.,func = pickle.load), tracking assignments across the local namespace. - Tracks imports/aliases to resolve calls like:
import pickle as p→p.loads(...)import torch as t→t.load(...)- Deep attributes like
pkg.pickle.loads(...)ortorch.serialization.load(...)
- Emits findings as JSONL, one object per line, including:
- file path + location (
lineno,col_offset) module,name,qualified_namecategoryandseverity(extra fields, backwards-compatible)parser: Metadata indicating which engine pass made the discovery (ast,tokenize_fallback, orregex_fallback).
- file path + location (
- Guards against pathological inputs with:
- max file size (
MAX_FILE_BYTES) - max visited AST nodes (
MAX_AST_NODES)
- max file size (
- Optional: loads a custom ruleset from
rules.jsonand warns on malformed entries/typos.
Minimum Python version 3.9+ recommended and 3.10+ tested.
python --version
# Python 3.9+ recommendedSetup Requirements:
- Environment: Create a
.envfile in the root directory depending on your provider:# For Hugging Face provider HF_TOKEN=your_token_here # For OpenAI / Local LLM provider HA_LLM_TOKEN=your_jwt_token_here
Before running the scanner, you MUST install the dependencies:
python -m venv .venv
.venv\Scripts\activate
pip install -r requirements.txtRecommended usage:
python deserializer.py --path cloned-repo --rules-file rules.json -j 4 --out cloned-repo/cloned-repo.jsonlpython deserializer.py --path cloned-repo --rules-file rules.json -j 4 --out cloned-repo/cloned-repo.jsonl --agent --agent-provider local --llm-api-url http://127.0.0.1:8181/v1python deserializer.py --path cloned-repo --rules-file rules.json -j 4 --out cloned-repo/cloned-repo.jsonl --agent --agent-provider openai --llm-api-url http://127.0.0.1:8181/v1This phase integrates a specialized AI Security Agent to perform deep code reviews and map complex 0-day RCE vectors. The agent analyzes findings to reverse "self-command-injection" contexts and generate technical reproduction guides with a multi-platform focus (e.g., Attacker UNIX/Raspberry vs Victim Windows). Note: This phase is only executed if the --agent flag is provided.
Inference Providers:
- Hugging Face (Default): Uses cloud-based API models such as MiniMax-M2.5.
- Requires
HF_TOKENin a.envfile or environment variable. - Command:
python deserializer.py --path /path/to/repo --agent --agent-provider huggingface
- Requires
- Local LLM (
llama.cpp): Uses a local REST server (llama-server.exe).- Command example to launch server:
.\llama-server.exe --model .\models\model.gguf --host 127.0.0.1 --port 8181 --ctx-size 600000 --jinja - Command:
python deserializer.py --path /path/to/repo --agent --agent-provider local --llm-api-url http://127.0.0.1:8181/v1
- Command example to launch server:
- OpenAI Compatible:
- Requires
HA_LLM_TOKEN(orOPENAI_API_KEY) in.env. - Command:
python deserializer.py --path /path/to/repo --agent --agent-provider openai --llm-api-url http://127.0.0.1:8181/v1
- Requires
Scan the current directory and print findings to the terminal (writes JSONL to stdout by default):
python deserializer.pyScan a specific repository and save findings to a JSONL file:
python deserializer.py --path /path/to/my-repo --out audit_results.jsonlUse 8 concurrent processes and a 5-second timeout per file to keep the scan moving:
python deserializer.py -j 8 --timeout 5 --out findings.jsonlDisable the banner and stream JSONL directly to stdout for pipe processing (human logs will go to stderr):
python deserializer.py --no-banner --out - | jq .Limit processing to files under 1MB and skip specific data directories:
python deserializer.py --max-size 1048576 --skip-dirs "data,samples,tests"Use a proprietary ruleset to detect logic-specific calls:
python deserializer.py --rules-file my_custom_rules.json --out legacy_audit.jsonlTrigger the AI-Driven Deep Analysis (Phase 4) after the scan and mapping are complete:
python deserializer.py --path /path/to/repo --out findings.jsonl --agent-
--path <dir>
Root directory to scan. Default:. -
--rules-file <path>
Path to a JSON ruleset. If provided, it overrides the built-inDEFAULT_RULES. -
--out <path|->
Output destination for JSONL results. Use-to write JSONL tostdout. Default:- -
-j, --concurrency <int>
Number of concurrent processes to use. Default:(CPU cores - 2). -
-t, --timeout <float>
Timeout in seconds for each file analysis. Only effective in parallel mode. Default:None(no timeout). -
--max-size <bytes>
Maximum file size in bytes to process. Default:10,485,760(10 MiB). -
--skip-dirs <list>
Comma-separated list of directory names to ignore (e.g.,tests,.git,env). -
--no-banner
Disable the ASCII branding banner for cleaner output in scripts. -
--agent
Run AI-driven deep analysis (Phase 4). Optional; requires valid credentials (HF_TOKENorHA_LLM_TOKEN) in.envor local LLM server. -
--agent-provider <provider>
LLM inference provider for AI agent:huggingface,local, oropenai(default:huggingface). -
--llm-api-url <url>
Base URL for LLM inference API (default:http://127.0.0.1:8181/v1).
Each line is a standalone JSON object. Example finding:
{
"file": "some/path/module.py",
"kind": "call",
"module": "pickle",
"name": "loads",
"qualified_name": "pickle.loads",
"category": "deserialize",
"severity": "high",
"lineno": 34,
"col_offset": 11
}Errors (parse/read/stat/limits) are also emitted as JSONL objects:
{ "file": "bad.py", "error": "syntax_error:..." }Exit code:
0if no errors occurred during scanning1if any IO/parse/limit errors occurred (useful for CI)
Rules are a JSON object keyed by a logical module name, with:
imports: list of import roots to trackcalls: list of[module, function]pairs
Example:
{
"pickle": {
"imports": ["pickle"],
"calls": [["pickle","load"], ["pickle","loads"]]
},
"torch": {
"imports": ["torch"],
"calls": [["torch","load"]]
}
}The loader normalizes imports (keeps the root token) and includes a Heuristic Typo Detector that prints warnings to stderr if rule definitions contain near-matches to tracked modules (e.g., alerting on picle vs pickle).
This tool is a signal amplifier, not a verdict generator. A pickle.load(s) finding is often high-risk, but real exploitability depends on whether an attacker can influence the loaded artifact (local file, downloaded model, CI artifact, bucket object, etc.) and on any integrity/provenance controls in the pipeline.
Copyright (c) 2026 Joshua Provoste. All rights reserved. No license is granted to use, copy, modify, or distribute this software without explicit permission.
