Popular repositories Loading
-
llm-jury
llm-jury PublicTraining-free committee-of-6 LLM judge for RAG hallucination detection: AUROC 0.88 / F1 0.77 on RAGTruth, reproducible at $0 from frozen verdicts — and it surfaces provable label errors in the benc…
Python
-
driftbet
driftbet PublicLabel-free drift attribution (model / noise / world / annotator) with anytime-valid guarantees and a pre-registered change-ledger — closes the zero-label misattribution gap 0.50→0.00 on the recover…
Python
-
whest-teardown
whest-teardown PublicReproducible teardown of ARC WhiteBox Estimation 2026: plain Monte Carlo beats the analytic reference ~9×, plus a 4.4× FLOP-accounting arbitrage in the scoring harness.
Python
If the problem persists, check the GitHub status page or contact support.
