IFAS study: how tool/API failures and task-state loss affect the safety behaviour of LLM agents. Active research design.
-
Updated
Oct 5, 2026 - Python
IFAS study: how tool/API failures and task-state loss affect the safety behaviour of LLM agents. Active research design.
Benchmarking schema-valid false tool observations and defense baselines for tool-using LLM agents.
To associate your repository with the toolsandbox topic, visit your repo's landing page and select "manage topics."