Skip to content

Repository files navigation

Deserializer

Deserializer Banner

About the Project

  • Deserializer is an advanced Abstract Syntax Tree (AST) static analysis engine designed to identify insecure object reconstruction and state persistence sinks across the Python ecosystem, with strong focus on AI, LLM, Robotics, Data Science, Machine Learning and Deep Learning fields; but not limited.
  • Deserializer was designed to scan and analyse any king of project based on Python source code, like apps, scripts, frameworks, libraries, Python pip packages, and much more.

Key technical capabilities

Far beyond a traditional scanner, it provides a generic, high-performance framework for auditing over 120 libraries and formats—including YAML, Msgpack, CBOR, and custom JSON hooks—where traditional trust in "safe" serialization hides critical logic-based RCE vectors. By resolving imports, aliases, and complex dotted attributes, the tool serves as a high-fidelity signal amplifier that prioritizes dangerous code paths in modern distributed architectures and AI/ML repositories.

The project operates through a modular, multi-phased workflow that transitions from raw detection to deep technical audit. Following the initial high-velocity SAST scan, the ecosystem leverages specialized relationship mappers to trace execution flows and result processors to generate detailed security reports. This systematic approach ensures that every finding is contextualized within the application's broader architecture, transforming high-volume telemetry into actionable research assets and structured milestones that simplify the mapping of infrastructure-level attack surfaces.

Autonomous AI Security Agent

At its most advanced tier, Deserializer integrates an autonomous AI Security Agent (Phase 4), explicitly designed to navigate "self-command-injection" limitations and synthesize functional reproduction guides, based on HuggingFace inference API, Local LLM (like llama.cpp), or OpenAI API compatible.

As the project evolves, it continues to define the frontier of automated vulnerability research by bridging the gap between abstract syntax tree static analysis, relationship mapping, reporting, and functional exploit development based on documented research.

Research (AI, Robotics, Data Science, Machine Learning and Deep Learning)

Deserializer directly supports security research by locating RCE and Insecure Deserialization paths across various large-scale AI, Robotics, and Data Science projects and environments, like Genesis World (v0.2.1), MuJoCo (v3.7.0), LeRobot (v0.5.1), Brax (v0.14.2), TensorFlow (v2.21.0), LangGraph (v1.1.6), VibeVoice (v0.0.1), Hugging Face Hub (v1.11.0), PyGlove (v0.4.5), and many others.

Its capabilities have directly powered the discovery of critical vulnerabilities in industry-leading frameworks, proving its efficacy in auditing complex MLOps and agentic AI environments.

The 4 Phases of Deserializer

The project is structured into four distinct phases, each designed to shift the analysis from high-volume automated telemetry to deep, functional security research:

Phase Title Tooling / Engine Objective
1 High-Velocity Detection deserializer.py (Triple-Pass) Perform massive-scale SAST to identify potential deserialization sinks.
2 Relationship Mapping Result Processors / Mappers Contextualize findings by tracing execution flows and component interdependencies.
3 Technical Synthesis Research Documentation Formalize findings into technical writeups, mapping infrastructure-level attack surfaces.
4 Autonomous AI Agent AI Security Agent Automate 0-day discovery and generate functional reproduction guides/exploits using HuggingFace inference API, Local LLM (like llama.cpp), or OpenAI API compatible.

Project Structure

A high-level map of the repository's organization and the technical purpose of each specialized directory:

  • agent/
    Implementation of the Phase 4 AI Security Agent, including orchestration logic for deep analysis, reproduction guide synthesis, and LLM interaction.
  • docs/
    Central repository for project documentation, architectural diagrams, branding assets, and supporting high-level technical materials.
  • exploit_development/
    Dedicated space for engineering functional exploits, researching deserialization primitives (hooks/methods), and documenting cross-platform orchestration techniques.
  • modules/
    Core logic components, auxiliary utility functions, and scanner extensions that power the detection and relationship mapping phases of the toolchain.
  • reports/
    Storage for structured security analysis outputs, telemetry, and audit results generated by the scanner during multi-phased project evaluations.
  • research/
    A collection of high-fidelity vulnerability writeups, validated Proof-of-Concept (PoC) scripts, and deep audits conducted on modern AI/ML frameworks.
  • templates/
    Standardized reporting and research templates used to maintain technical consistency across vulnerability writeups and correlation reports.

Performance & Multiprocessing

This scanner features a high-performance parallel execution engine built on Python's concurrent.futures.ProcessPoolExecutor. It is designed to scale across all available CPU cores (controllable via the -j or --concurrency flag), making it capable of scanning tens of thousands of files in seconds.

  • Non-Blocking UI: Features a stable, docked progress bar with a one-line gap for clean results presentation.
  • Native Signal Handling: On Windows, it utilizes a native SetConsoleCtrlHandler via ctypes to ensure that Ctrl+C is 100% responsive, even during heavy processing.
  • Silent Tracebacks: Worker processes are silented to ensure that interrupts and internal errors don't clutter the technical output.
  • Native Windows UI Support: Automatically enables Virtual Terminal Processing via ctypes for native ANSI color support in modern CMD and PowerShell environments.

Warning

Performance Warning: When analyzing extremely large or complex files (e.g., over 1MB, 2MB, or 3MB in size), the tool may experience significant slowdowns or appear "stuck" while parsing deep AST trees. If you encounter such bottlenecks, consider using the --timeout (to skip slow files) and --max-size (to skip huge files) flags to maintain scan velocity.

What it does

  • Triple-Pass Scanning Engine: Implements an "unbreakable" multi-tiered approach:
    • Pass 1 (AST): High-fidelity grammatical decomposition for precise analysis.
    • Pass 2 (Token Fallback): Automated fallback to a stream-based tokenizer when encountering syntax errors or unparsable sections.
    • Pass 3 (Regex Emergency): A final pattern-matching layer ensuring coverage in extremely hostile or fragmented source files.
  • Native Template Neutralization: Built-in support for sanitizing Jinja2 and Mako tags, allowing the scanner to process web templates and code-generation files without choking on non-Python syntax.
  • Call & Reference Detection: Identifies not just direct execution sinks (e.g., pickle.loads()) but also dangerous function references (e.g., func = pickle.load), tracking assignments across the local namespace.
  • Tracks imports/aliases to resolve calls like:
    • import pickle as p → p.loads(...)
    • import torch as t → t.load(...)
    • Deep attributes like pkg.pickle.loads(...) or torch.serialization.load(...)
  • Emits findings as JSONL, one object per line, including:
    • file path + location (lineno, col_offset)
    • module, name, qualified_name
    • category and severity (extra fields, backwards-compatible)
    • parser: Metadata indicating which engine pass made the discovery (ast, tokenize_fallback, or regex_fallback).
  • Guards against pathological inputs with:
    • max file size (MAX_FILE_BYTES)
    • max visited AST nodes (MAX_AST_NODES)
  • Optional: loads a custom ruleset from rules.json and warns on malformed entries/typos.

Installation

Minimum Python version 3.9+ recommended and 3.10+ tested.

python --version
# Python 3.9+ recommended

Setup Requirements:

  • Environment: Create a .env file in the root directory depending on your provider:
    # For Hugging Face provider
    HF_TOKEN=your_token_here
    
    # For OpenAI / Local LLM provider
    HA_LLM_TOKEN=your_jwt_token_here

Before running the scanner, you MUST install the dependencies:

python -m venv .venv
.venv\Scripts\activate
pip install -r requirements.txt

Usage

Recommended usage:

Without AI Agent

python deserializer.py --path cloned-repo --rules-file rules.json -j 4 --out cloned-repo/cloned-repo.jsonl

With AI Agent (Local LLM)

python deserializer.py --path cloned-repo --rules-file rules.json -j 4 --out cloned-repo/cloned-repo.jsonl --agent --agent-provider local --llm-api-url http://127.0.0.1:8181/v1

With AI Agent (OpenAI / Local LLM)

python deserializer.py --path cloned-repo --rules-file rules.json -j 4 --out cloned-repo/cloned-repo.jsonl --agent --agent-provider openai --llm-api-url http://127.0.0.1:8181/v1

AI-Driven Deep Analysis (Phase 4) with an AI Security Agent

This phase integrates a specialized AI Security Agent to perform deep code reviews and map complex 0-day RCE vectors. The agent analyzes findings to reverse "self-command-injection" contexts and generate technical reproduction guides with a multi-platform focus (e.g., Attacker UNIX/Raspberry vs Victim Windows). Note: This phase is only executed if the --agent flag is provided.

Inference Providers:

  • Hugging Face (Default): Uses cloud-based API models such as MiniMax-M2.5.
    • Requires HF_TOKEN in a .env file or environment variable.
    • Command: python deserializer.py --path /path/to/repo --agent --agent-provider huggingface
  • Local LLM (llama.cpp): Uses a local REST server (llama-server.exe).
    • Command example to launch server: .\llama-server.exe --model .\models\model.gguf --host 127.0.0.1 --port 8181 --ctx-size 600000 --jinja
    • Command: python deserializer.py --path /path/to/repo --agent --agent-provider local --llm-api-url http://127.0.0.1:8181/v1
  • OpenAI Compatible:
    • Requires HA_LLM_TOKEN (or OPENAI_API_KEY) in .env.
    • Command: python deserializer.py --path /path/to/repo --agent --agent-provider openai --llm-api-url http://127.0.0.1:8181/v1

1. Basic Scan

Scan the current directory and print findings to the terminal (writes JSONL to stdout by default):

python deserializer.py

2. Targeted Audit

Scan a specific repository and save findings to a JSONL file:

python deserializer.py --path /path/to/my-repo --out audit_results.jsonl

3. Turbo Mode (Performance Tuning)

Use 8 concurrent processes and a 5-second timeout per file to keep the scan moving:

python deserializer.py -j 8 --timeout 5 --out findings.jsonl

4. CI/CD & Pipeline Integration

Disable the banner and stream JSONL directly to stdout for pipe processing (human logs will go to stderr):

python deserializer.py --no-banner --out - | jq .

5. Hardened / Safety Scan

Limit processing to files under 1MB and skip specific data directories:

python deserializer.py --max-size 1048576 --skip-dirs "data,samples,tests"

6. Custom Detection Rules

Use a proprietary ruleset to detect logic-specific calls:

python deserializer.py --rules-file my_custom_rules.json --out legacy_audit.jsonl

7. AI Deep Analysis

Trigger the AI-Driven Deep Analysis (Phase 4) after the scan and mapping are complete:

python deserializer.py --path /path/to/repo --out findings.jsonl --agent

CLI Flags

  • --path <dir>
    Root directory to scan. Default: .

  • --rules-file <path>
    Path to a JSON ruleset. If provided, it overrides the built-in DEFAULT_RULES.

  • --out <path|->
    Output destination for JSONL results. Use - to write JSONL to stdout. Default: -

  • -j, --concurrency <int>
    Number of concurrent processes to use. Default: (CPU cores - 2).

  • -t, --timeout <float>
    Timeout in seconds for each file analysis. Only effective in parallel mode. Default: None (no timeout).

  • --max-size <bytes>
    Maximum file size in bytes to process. Default: 10,485,760 (10 MiB).

  • --skip-dirs <list>
    Comma-separated list of directory names to ignore (e.g., tests,.git,env).

  • --no-banner
    Disable the ASCII branding banner for cleaner output in scripts.

  • --agent
    Run AI-driven deep analysis (Phase 4). Optional; requires valid credentials (HF_TOKEN or HA_LLM_TOKEN) in .env or local LLM server.

  • --agent-provider <provider>
    LLM inference provider for AI agent: huggingface, local, or openai (default: huggingface).

  • --llm-api-url <url>
    Base URL for LLM inference API (default: http://127.0.0.1:8181/v1).

Output format (JSONL)

Each line is a standalone JSON object. Example finding:

{
  "file": "some/path/module.py",
  "kind": "call",
  "module": "pickle",
  "name": "loads",
  "qualified_name": "pickle.loads",
  "category": "deserialize",
  "severity": "high",
  "lineno": 34,
  "col_offset": 11
}

Errors (parse/read/stat/limits) are also emitted as JSONL objects:

{ "file": "bad.py", "error": "syntax_error:..." }

Exit code:

  • 0 if no errors occurred during scanning
  • 1 if any IO/parse/limit errors occurred (useful for CI)

Rules file format (rules.json)

Rules are a JSON object keyed by a logical module name, with:

  • imports: list of import roots to track
  • calls: list of [module, function] pairs

Example:

{
  "pickle": {
    "imports": ["pickle"],
    "calls": [["pickle","load"], ["pickle","loads"]]
  },
  "torch": {
    "imports": ["torch"],
    "calls": [["torch","load"]]
  }
}

The loader normalizes imports (keeps the root token) and includes a Heuristic Typo Detector that prints warnings to stderr if rule definitions contain near-matches to tracked modules (e.g., alerting on picle vs pickle).

Notes on interpretation

This tool is a signal amplifier, not a verdict generator. A pickle.load(s) finding is often high-risk, but real exploitability depends on whether an attacker can influence the loaded artifact (local file, downloaded model, CI artifact, bucket object, etc.) and on any integrity/provenance controls in the pipeline.

License

Copyright (c) 2026 Joshua Provoste. All rights reserved. No license is granted to use, copy, modify, or distribute this software without explicit permission.

About

AST-based Static Code Analyzer with Agentic LLM-Powered Relationship Mapping to discover Python RCE paths and deep deserialization chains on AI, LLM, Robotics, Data Science, Machine Learning and Deep Learning (but not limited).

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages