This repository documents an independent optimization and validation study for an AIVV-style verification and validation pipeline. The focus is on improving the mathematical sentry's precision-recall behavior and evaluating whether downstream LLM council agents can correctly distinguish true physical failures from false alarms, safe transients, planned events, and dangerous instructions.
This work is an independent benchmarking, optimization, and evaluation effort built around concepts from the following foundational research:
- AIVV Framework: AIVV: Neuro-Symbolic LLM Agent-Integrated Verification and Validation for Trustworthy Autonomous Systems (Kwon et al., 2026).
https://arxiv.org/pdf/2604.02478 - Fault Detection Baseline: Fault Detection for Agents on Power Grid Topology Optimization (Lehna et al., 2024).
https://arxiv.org/pdf/2406.16426
This repository does not claim authorship of the foundational AIVV framework or any private lab-provided source code, datasets, configurations, or prompts. Public materials in this repository should be limited to original analysis, clean evaluation harnesses, anonymized or synthetic data examples, reproducible benchmark scripts, figures, and documentation of independent contributions.
- Benchmark the mathematical sentry using recall, precision, false positive rate, false negative rate, F1, AUROC/AUPRC, and threshold-sensitivity curves.
- Optimize sentry decision policies around conformal prediction bounds, uncertainty thresholds, persistence filters, and scenario-aware calibration.
- Evaluate the downstream LLM council's ability to classify and triage true failures, nuisance alarms, safe transients, planned events, and malicious or dangerous instructions.
- Produce a clear, reproducible portfolio artifact showing engineering contributions without exposing private code, datasets, credentials, or lab artifacts.
.
├── configs/ # Public-safe experiment configuration files
├── data_public/ # Synthetic, anonymized, or toy data only
├── docs/ # Architecture notes, experiment design, and writeups
├── experiments/ # Runnable experiment entry points
├── figures/ # Generated plots and architecture diagrams
├── notebooks/ # Exploratory analysis notebooks
├── references/ # Citation notes and paper summaries
├── results/ # Public-safe benchmark outputs and summaries
├── scripts/ # Utility scripts for running batches or generating plots
├── src/ # Reusable evaluation and optimization code
│ ├── common/
│ ├── council_eval/
│ └── sentry_eval/
└── tests/ # Unit tests for metrics, policies, and evaluators
- Establish baseline sentry behavior using the existing conformal prediction gate.
- Sweep conformal alpha values and analyze recall-precision tradeoffs.
- Compare conformal-only, uncertainty-only, combined, and weighted anomaly-score policies.
- Add persistence-based escalation to reduce single-sample false alarms.
- Evaluate scenario-aware calibration for nominal noise, safe transient maneuvers, planned events, and true faults.
- Report precision-recall frontiers and false-positive reductions.
- Build scenario suites for true failures, false alarms, safe transients, planned events, ambiguous cases, and dangerous instructions.
- Evaluate whether each council role votes correctly and cites valid telemetry evidence.
- Track override behavior when the mathematical sentry produces a false positive.
- Measure consistency, hallucination rate, refusal behavior for dangerous instructions, latency, and cost.
- Compare prompt variants, agent-role definitions, and structured-output schemas.
Public materials may include:
- Original benchmark harnesses and analysis code
- Synthetic or anonymized toy datasets
- Precision-recall and threshold tradeoff plots
- Architecture diagrams and methodological writeups
- Summarized results that do not expose private data
Public materials should not include:
- Private lab datasets or raw simulator outputs without permission
- Unedited source code from the original AIVV folder without permission
- API keys,
.envfiles, credentials, private endpoints, or proprietary prompts - Model checkpoints or logs that reveal private data
- Document the baseline AIVV data flow and sentry decision rule.
- Create baseline sentry metrics from reproducible runs.
- Implement threshold sweep utilities for conformal alpha and uncertainty cutoffs.
- Add persistence and hybrid scoring sentry policies.
- Build scenario-labeled council evaluation cases.
- Generate plots and summary tables for portfolio presentation.
- Write a concise final report in
docs/final_study_summary.md.