Skip to content

Repository files navigation

AI Experiments

Measured experiments on AI models, each one keeping the exact request and response of every call it made. Completed results are published at lab.codewithnk.com.

Every experiment is two files: README.md says what it is and how to run it, written before the run, and RESULTS.md is what happened, written after. Two files rather than two sections, so a prediction cannot be quietly reworded once the answer is known and git log shows which came first. EXPERIMENT_TEMPLATE.md is the skeleton to copy for a new one.

The experiments

An experiment with no result link has not produced one yet.

Experiment Dataset What it is trying to find out Result
rerank BEIR NFCorpus test
3,633 docs, 323 queries, top-20
what it holds
Whether a decision model re-ranks retrieval better than hybrid search, and what that costs RESULTS README
banking77 BANKING77 test
3,080 rows, 77 intents
what it holds
Whether a published accuracy is explained by a shared per-option token budget, or by something else RESULTS README
typed_decisions AG News test, 4 labels
CLINC150 plus test, 151 labels
300 examples each
what they hold
How nine models compare on accuracy, latency, cost, validity and positional robustness for one fixed-label job, at 4 options and at 151 RESULTS README

Running one

Needs uv and Python 3.12. One command runs any experiment, here or on a Kaggle GPU. It is idempotent: work already in the wire log is skipped, so a second invocation recomputes the tables in seconds rather than measuring again.

uv sync

uv run cli list                                  # every experiment, and how far it got
uv run cli run    <name> [--tag full] [flags]    # run it on this machine
uv run cli run    <name> --help                  # the standard flags plus that experiment's own

--tag names both runs/<tag>/ and results/<tag>/. A tag that has run before reuses the settings it ran with, so re-running one means the same measurement rather than today's flag defaults.

On a Kaggle GPU, for the experiments that have a notebook at <name>/kaggle/:

uv run cli submit <name> --dry-run               # check the prerequisites, push nothing
uv run cli submit <name>                         # push and start
uv run cli status <name>
uv run cli fetch  <name>                         # bring results and the wire log back

submit refuses when submitting would be wrong, and reports every reason at once. The kernel clones this repository at main, so uncommitted work, unpushed commits and the wrong branch each produce a run that looks entirely normal while measuring code that is not in front of you. Two things it cannot check — whether the notebook's accelerator is on and whether its secret is attached — it prints as steps with the URL.

Where the output goes

Each experiment keeps everything it produces inside itself:

<name>/results/<tag>/ committed: summary.json, CSV tables, figures/
<name>/runs/<tag>/ gitignored: the raw wire log and the frozen run config

The split is what makes "the results reproduce byte-identically" a checkable claim rather than an assertion: the log is the evidence, the results are the claim, and re-running the report derives one from the other. Each RESULTS.md names its own files and what they mean.

uv run pytest

About

Measured experiments on decision models. Can an open-weight model re-rank search results? Does a token budget explain a benchmark failure? What does one typed decision cost across model families? Every request and response is recorded, and negative results are published as they came out.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages