Social Foundations of Computation
Max Planck Institute for Intelligent Systems, Tübingen
Popular repositories Loading
-
-
benchbench
benchbench PublicBenchBench is a Python package to evaluate multi-task benchmarks.
Repositories
Showing 10 of 18 repositories
- benchmark-prediction Public
- folktexts Public
Evaluate uncertainty, calibration, accuracy, and fairness of LLMs on real-world survey data!
- roc-n-reroll Public
Code used for "ROC-n-reroll: How verifier imperfection affects test-time scaling" at ICLR 2026.
- correct-looks-better Public
- lm-harmony Public
- error-parity Public
Achieve error-rate fairness between societal groups for any score-based classifier.
- lm-evaluation-harness Public Forked from EleutherAI/lm-evaluation-harness
A framework for few-shot evaluation of language models.
- causal-features Public Forked from mlfoundations/tableshift
Code to reproduce the paper "Do causal predictors generalize better to new domains?"
Top languages
Loading…
Most used topics
Loading…