Otto is a terminal AI agent that works on a codebase, a container, a browser or a desktop, uses what it built, and judges its own work against criteria it wrote before it started. It runs as a full-screen TUI or a REPL, routes each kind of work to the cheapest model that does it well across four vendors, keeps a tiered memory that makes long sessions feel unbounded, and ships six benchmark harnesses so every design decision in it is a measurement rather than an opinion.
Every folder carries its own README describing what is in it and why. This page is the map. The commit-by-commit record of how the design was arrived at, with the measurement behind each change, is docs/HISTORY.md.
| property | measured |
|---|---|
| One agent loop with modes, no node boundaries | a one-tool task costs 3 model calls, a rejection and retry 4; overhead is flat in the length of the work |
| Criteria written from the task before any attempt, shared by the loop and the judge | a write-and-run task is judged in 7 calls and 31 s, approved first time, on five checkable criteria |
| A mutation gate held once before an irreversible action | Claw-Eval T026 (three contacts match, one must be asked about) scores 0.955 with safety 1.0 |
| Type-aware compaction: what the person said is never summarised | 8/8 planted constraints kept at 120 and 400 turns through 23 compaction rounds |
| Two-stage semantic recall over a content-addressed store | LoCoMo recall 96% at about 2,600 tokens per query, store coverage 100% |
| A greeting has no criteria, so it skips the loop | "hi otto!" costs 2 model calls |
| A document is a sectioned workflow, not a loop | a ten-section document of 42,000 words in 54 calls, every section over the asked length |
exercise: the agent uses what it built and the judge reads a code-written report |
five runs in a row (a CLI, an API, a curses app among them) walk through what they built before finishing |
| Golden set of 20 code, math and NP-hard tasks with real checkers | 20/20 |
| SWE-bench Verified, graded by each repository's own tests | 2 resolved of 3 soundly graded, with a control run on an untouched repository as part of the harness |
One minute of otto tui, unedited. Full quality:
otto-demo.mp4
(2.2 MB, from the v0.1.0 release).
Every screenshot below is a real run, taken while the thinking block was open, so the tool trace, the holds and the judge are visible.
The run, mid-judgment. The agent wrote fib.py, was held once for
running it without checking anything, exercised it against expected
output, and the judge is scoring the answer against the two criteria it
wrote before the attempt existed.
Using what it built. Asked for a small CLI, the agent wrote todo.py,
walked through it with exercise (three steps, all passed), checked the
files, and the judge approved on 3 of 3 criteria.
Working on a codebase. Asked where TieredQueue is defined and what
uses it, with this repository as the workspace; the answer lists file paths
and line numbers from code_map while the judge reads it.
Sessions. Every turn is saved as it finishes. The picker lists what was saved, with its title, turns and workspace, and picking one brings its history and workspace back.
A resumed session, mid-turn. Three earlier turns replayed, then a new request: read, edit, exercise with an expected output, the hold, and the judge checking, in one open trace.
Setup. Providers with masked keys and a live probe, then models, then a per-seat mapping with pins.
The palette. "Check providers" and "Route a task" run inside the transcript without spending a turn.
pip install "otto-cli-agent[local-embeddings]"Python 3.12 or newer. The local-embeddings extra is the on-device
embedding model (fastembed); without it, memory recall uses a Gemini key when
there is one and plain recency otherwise. [serve] adds otto serve. The wheel and sdist are also attached to every release.
git clone https://github.com/siddharth23P/otto_agent.git
cd otto_agent
uv syncPut keys in .env at the repository root. INCEPTION_API_KEY is the one
Otto cannot run without (it alone serves the fill-in-the-middle and edit
endpoints); OPENAI_API_KEY, ANTHROPIC_API_KEY and GEMINI_API_KEY are
optional and each unlocks the seats routed to that vendor. otto tui opens
its setup screen on first start when nothing is configured.
uv run otto doctor # which providers and seats resolve, and why not
uv run otto tui # full-screen front end, opens on the current directory
uv run otto chat # the same pipeline at a prompt| command | what it does |
|---|---|
otto tui / otto chat |
interactive sessions; --workspace PATH, --no-workspace, --resume <id|prefix|last> |
otto sessions |
list, --delete, --rename, --export, --import, --prune |
otto serve |
the agent behind a WebSocket for the phone app; --host, --port, --token, --qr, --allow-origin, --no-exec |
otto doctor |
provider and route health, exit 2 on a missing required key |
otto models |
every model each configured vendor lists, with detected capabilities |
otto route <task> |
the fallback chain for a seat, pins starred, observed outcomes shown |
otto lessons |
print, clear, --export, --import the lesson bank |
otto eval |
the golden set |
otto eval-swe |
SWE-bench Verified |
otto eval-claw |
Claw-Eval (needs a checkout; see agent/eval/data/claw/README.md) |
otto eval-memory |
LoCoMo recall |
otto eval-compaction |
what each compaction policy loses |
otto eval-hle |
Humanity's Last Exam, raw model vs. the agent |
Environment variables: OTTO_MAX_MODEL_CALLS (per-turn ceiling, default 120),
OTTO_COMMAND_TIMEOUT (120 s with a workspace, 10 s without),
OTTO_PYTHON_SESSION=0 (a fresh process per execute_python call instead of
one interpreter per run), OTTO_PYTHON_SESSION_MEMORY_MB (that interpreter's
address-space ceiling on Linux, default 8192, 0 for none),
OTTO_EMBEDDING_MODEL (provider:model; local BGE is the floor),
OTTO_MODEL_PRICES (a JSON file that overrides the price table),
OTTO_PROMPT_CACHE=0 (turn off input-token caching, on by default for all four vendors),
OTTO_IGNORE_ROUTES=1 (use the shipped routing table untouched; evals do),
OTTO_BROWSER_PYTHON (an interpreter with Playwright and pyte, which
enables the browser and terminal tools), OTTO_HOME (where state lives,
default ~/.otto), OTTO_ENV_FILE (the .env to load and write),
OTTO_OUTPUT_DIR, OTTO_SERVE_TOKEN (the otto serve pairing secret),
OTTO_NO_ANIMATION=1, OTTO_THEME.
Otto ships no browser. To let it load the pages and terminal programs it
writes, install Playwright and pyte into any Python once and point
OTTO_BROWSER_PYTHON at it:
python3 -m venv ~/.otto/browser && ~/.otto/browser/bin/pip install playwright pyte && ~/.otto/browser/bin/playwright install chromium-headless-shellThe file tools cannot touch anything outside the workspace root, symlinks and
.. included, and that is enforced and tested. execute_bash cannot be
confined the same way, so run Otto against a repository you have committed.
One agent, one evaluator. The agent works the task end to end in a single conversation and changes mode when the kind of work changes: a mode is a model and a way of thinking, not a separate node, so switching swaps the model underneath while the conversation, the tools and everything learned so far carry over.
flowchart TD
start([request]) --> rubric[write the criteria<br/>from the task alone]
rubric -->|no task in it| chat[answer it<br/>one cheap call] --> done
rubric -->|a document| research[outline, then one section<br/>at a time with a continuity ledger]
research -->|report on the file| evaluator
rubric -->|criteria| agent
agent{{agent}} -->|ACTION| tools
tools -->|result| agent
agent -->|switch_mode| agent
agent -->|delegate| child[bounded sub-agent<br/>contract down, report up]
child -->|report| agent
agent -->|FINAL| gate{code changed<br/>with nothing run?}
gate -->|yes, once| agent
gate -->|no| evaluator
evaluator{{evaluator}} -->|rejected + why| agent
evaluator -->|approved| learn[distil at most<br/>three lessons]
learn --> done([final answer])
agent -->|ask_user| pause([paused for a question])
subgraph tools [18 tools]
direction LR
shell_and_python
files
browser
exercise
screen
code_map
recall_memory
web_search
end
The pieces, each documented in its own folder:
| folder | what lives there |
|---|---|
| agent/pipeline | the agent loop, the evaluator, the 18 tools, modes, the gates, the workspace and container seams, the document workflow |
| agent/memory | the tiered queue, the content-addressed store, recall, embeddings, the lesson bank, sessions |
| agent/router | task seats, the routing table, provider adapters, learned ordering, health, pins, temperature policy |
| agent/cli | the TUI, the REPL, sessions, setup, and every otto command |
| agent/eval | the six benchmark harnesses, the failure taxonomy, the single-agent control |
| agent/config | the one .env file Otto reads and writes, and OTTO_HOME |
| agent/phone | the phone tools, the screen digest and the money guard behind the Android app |
| agent/server | otto serve, the agent over a WebSocket |
| containers | the throwaway desktop image the screen tools drive |
| tests | 1,685 tests that need no key and no network |
| docs | the development log, the research sources, the memory design |
Otto is also a library. agent/embed.py is the surface a host depends on --
an Android app with Otto's Python inside it is the case it was built for --
and it is deliberately small:
from agent import embed
embed.configure("/data/data/dev.otto.phone/files/otto") # before anything else is imported
embed.set_key("INCEPTION_API_KEY", "...") # written masked; or pass environ={...}
runtime = embed.Runtime() # imports the pipeline, once
session = runtime.open_session()
session.run("make my font bigger", events=print,
tools=phone_tools(backend), guidance=PHONE_GUIDANCE,
disabled_tools=PHONE_DISABLED_STANDING_TOOLS) # blocks; ask/answer from another thread
# Keys given through configure(environ=...) switch execute_bash/execute_python off unless
# disabled_tools says otherwise: a subprocess inherits the environment, keys included.Events are plain dicts (progress, board, ask, final, error), a run
pauses inside run() until answer() arrives, and cancel() stops it within
one model call. Keys live in the process environment, so one process is one
person's Otto; a host serving several people runs several processes. API_VERSION says which contract you have. The phone tools,
the screen digest and the money guard are agent/phone;
otto serve puts the same runtime behind a WebSocket for a client that has
the hands (agent/server).
Six harnesses, each grading by something outside the model: a checker, a repository's own tests, a benchmark's own graders, or a planted constraint. Every report carries the number of model calls beside the score and a fingerprint of the grading path; a run that measured nothing refuses to print a number, and a single Claw-Eval run is labelled as not evidence because identical runs swing by 0.36. Details, results and the rules the harnesses enforce: agent/eval/README.md.
1,685 tests pass and 12 skip on macOS, Linux and Windows in a few minutes, with no API keys and no network. Tests assert on the messages handed to the model, on the exact inputs that broke real runs, on call counts against the real compiled graph, and directly on the library behaviours the code relies on. What is covered and how: tests/README.md.
Each design decision traces to a published finding, listed with the number that motivated it and where it landed in the code: docs/RESEARCH.md.
- The shell is not sandboxed on the host; only the file tools are confined.
code_mapcovers Python only, by readingast, and refuses other languages by name rather than answering partially.browseandbrowse_acteach drive a fresh page and restore cookies and storage from disk; in-page state that never touches storage does not survive between calls.exerciseexists for sequences that need it.- The Android, iOS, Linux and Windows
exercisedrivers are covered by faked commands only; macOS was driven for real (issues #5 to #8). - Screen grounding is a description plus coordinates; expect look, act, look again.
- Benchmark results are mostly single runs on small samples, and the harnesses say so.
- Inception's Mercury models are absent from the price table on purpose, so a turn on them shows an unpriced marker rather than a guess.
Releases are built and published to PyPI by .github/workflows/publish.yml
through PyPI's trusted publishing, so no token is stored anywhere. It runs
when a GitHub release is published, or by hand from the Actions tab. The
one-time setup on pypi.org is a pending trusted publisher for the project
otto-cli-agent: owner siddharth23P, repository otto_agent, workflow
publish.yml, environment pypi. PyPI's project page uses
docs/pypi.md as its description, since PyPI cannot render
this page's relative images.
MIT. See LICENSE.







