Skip to content
@poisson-labs

Poisson Labs

Open-source tools for finding where AI agents and RL policies fail. Brooklyn.

Popular repositories Loading

  1. vjepa-stress vjepa-stress Public

    Where V-JEPA 2.1's dense video features hold up under noise, blur, frame drops, occlusion and low light, across all four model sizes. Code and data for the Poisson Labs post.

    Jupyter Notebook 2

  2. mhc-repro mhc-repro Public

    Reproduction of DeepSeek's mHC (Manifold-Constrained Hyper-Connections), from 10M to 2.5B parameters, logging how much each residual connection amplifies the signal.

    Python

  3. policy-robustness-sweep policy-robustness-sweep Public

    Maps where a Unitree Go1 walking policy falls over across 400 combinations of floor friction and push, then compares the map before and after retraining.

    Python

  4. babel babel Public

    Cooperative multi-agent RL benchmark for communication robustness. Shows how training metrics can hide a channel agents don't actually use, and a reward exploit.

    Python

  5. cable-insertion cable-insertion Public

    MuJoCo environments and PPO baselines for robotic cable insertion, including a UR5e arm with a Robotiq gripper.

    Python

  6. jev-replay jev-replay Public

    Measures whether TypeSafe's Jev decision model flips actions on identical replays near shipped thresholds across 12,705 calls. Replication data, harness, and independent verification.

    Python

Repositories

Showing 6 of 6 repositories
  • jev-replay Public

    Measures whether TypeSafe's Jev decision model flips actions on identical replays near shipped thresholds across 12,705 calls. Replication data, harness, and independent verification.

    poisson-labs/jev-replay's past year of commit activity
    Python 0 MIT 0 0 0 Updated Sep 23, 2026
  • cable-insertion Public

    MuJoCo environments and PPO baselines for robotic cable insertion, including a UR5e arm with a Robotiq gripper.

    poisson-labs/cable-insertion's past year of commit activity
    Python 0 MIT 0 0 0 Updated Sep 22, 2026
  • babel Public

    Cooperative multi-agent RL benchmark for communication robustness. Shows how training metrics can hide a channel agents don't actually use, and a reward exploit.

    poisson-labs/babel's past year of commit activity
    Python 0 MIT 0 0 0 Updated Sep 22, 2026
  • mhc-repro Public

    Reproduction of DeepSeek's mHC (Manifold-Constrained Hyper-Connections), from 10M to 2.5B parameters, logging how much each residual connection amplifies the signal.

    poisson-labs/mhc-repro's past year of commit activity
    Python 0 MIT 0 0 0 Updated Sep 22, 2026
  • policy-robustness-sweep Public

    Maps where a Unitree Go1 walking policy falls over across 400 combinations of floor friction and push, then compares the map before and after retraining.

    poisson-labs/policy-robustness-sweep's past year of commit activity
    Python 0 MIT 0 0 0 Updated Sep 22, 2026
  • vjepa-stress Public

    Where V-JEPA 2.1's dense video features hold up under noise, blur, frame drops, occlusion and low light, across all four model sizes. Code and data for the Poisson Labs post.

    poisson-labs/vjepa-stress's past year of commit activity
    Jupyter Notebook 2 MIT 0 1 0 Updated Sep 20, 2026

Top languages

Loading…

Most used topics

Loading…