Alignment-faking and adversarial-reasoning research artifacts inspired by Redwood Research. Behavioral probes and controlled failure-mode analysis.
python research alignment adversarial reasoning redwood ml-evaluation redwood-research behavioral-evaluation adversarial-reasoning alignment-faking ml-eval behavioral-probes controlled-failure
-
Updated
Sep 5, 2026 - Python