Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
11 changes: 8 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -48,9 +48,9 @@ For each question, the server:
4. Keeps only the logits for valid answer codes and applies softmax.
5. Maps those probabilities back to the answer names supplied by the caller.

Questions are evaluated independently. Each question gets its own prompt and
another copy of the state. This prevents one answer from affecting another,
but runtime and input-token use grow with the number and length of questions.
The default MLX and vllm-metal backends evaluate questions independently. The
DiffusionGemma backend instead places all answers in one shared canvas and
runs one denoising forward.

See [How it works](docs/how-it-works.md) for the prompt format, answer-code
registry, token limits, and confidence calculation.
Expand Down Expand Up @@ -275,6 +275,11 @@ Select a profile before starting the server:
SYSTEM_ONE_MODEL=larger uv run uvicorn system_one_lite.api:app --port 8010
```

To serve the same model through vllm-metal, see the
[vllm-metal backend guide](docs/vllm-metal.md).
For seeded, single-forward reads with the 4-bit DiffusionGemma model, see the
[DiffusionGemma backend guide](docs/diffusion-gemma.md).

The demo, benchmark, and evaluation tools also accept `--model default` or
`--model larger`. Each supported model has a checked-in registry of answer
codes that must remain single tokens with the pinned tokenizer.
Expand Down
8 changes: 4 additions & 4 deletions docs/api.md
Original file line number Diff line number Diff line change
Expand Up @@ -99,8 +99,8 @@ How likely the statement is to be true.
| --- | --- | --- |
| `model` | `string` | The engine's own model ID. The default is `mlx-community/Qwen3-1.7B-4bit`. |
| `answers` | `map<string, Answer>` | One answer per question, keyed by your IDs, in request order. |
| `usage.input_tokens` | `integer` | Total tokens evaluated across the independent question prompts. |
| `usage.output_tokens` | `integer` | Always 0. Nothing is generated. |
| `usage.input_tokens` | `integer` | Prompt tokens evaluated by the selected backend. |
| `usage.output_tokens` | `integer` | 0 for local MLX and read-only DiffusionGemma. The vllm-metal backend uses one internal output token per question for each group of up to 128 options. |

### Choice answer

Expand Down Expand Up @@ -152,5 +152,5 @@ You get an error, never a wrong answer.

Only one inference request runs at a time. If the engine is already in use,
the server returns `503` with `detail: "the inference engine is busy"`.
Each question uses a full independent prompt. `usage.input_tokens` counts
the state again for each question.
The default MLX and vllm-metal backends count one full prompt per question.
The DiffusionGemma backend counts one shared prompt for the whole request.
69 changes: 69 additions & 0 deletions docs/diffusion-gemma.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,69 @@
# Run with 4-bit DiffusionGemma

This backend uses `mlx-community/diffusiongemma-26B-A4B-it-4bit` at the pinned
revision in `engine.py`. It sends every question in one prompt. It places one
active canvas token per question. This fits System One's limit of 64 questions.
One read-only model pass returns the requested answer-code scores. The canvas
uses the same fixed seed for each request.

The backend needs the `feature/structured-reads` branch of MLX-VLM. That
branch adds `/v1/diffusion/reads` without changing normal generation.

Start the MLX-VLM server from its checkout:

```bash
cd /path/to/mlx-vlm
uv run --with-editable . python -m mlx_vlm.server --port 8080
```

Start System One in another terminal:

```bash
cd server
SYSTEM_ONE_BACKEND=mlx-vlm-diffusion \
uv run uvicorn system_one_lite.api:app --port 8010
```

Use the compact profile for short, latency-sensitive questions:

```bash
SYSTEM_ONE_BACKEND=mlx-vlm-diffusion-fast \
uv run uvicorn system_one_lite.api:app --port 8010
```

The compact profile accepts exactly one question. It removes the long prompt
instructions. Use it only when the state, question, and answer labels make the
task clear without extra guidance. It still runs the full model. Its speed
comes only from the shorter prompt.

The first System One startup downloads the pinned 4-bit model snapshot. The
weights are about 16.5 GB. Both processes use the same Hugging Face cache.

The read endpoint accepts prompt tokens, an active seed canvas, and exact token
IDs for each answer slot. It checks every token and position before the model
runs. It scores only the requested answer codes and skips the full vocabulary
head. System One then applies its usual softmax over each question's valid
answer codes.

This path differs from the vllm-metal backend in two ways:

- All questions share one prompt and one denoising forward.
- The answer slots can attend to the fixed answer template and to one another.

The public System One request and response formats do not change. Read-only
canvas work reports zero output tokens because no tokens are committed.

## Profile the backend

Start the MLX-VLM server, then run the warmed concurrency sweep from
`server/`:

```bash
uv run python -m tools.benchmark_diffusion
```

Use `--json` to save machine-readable results. Set `--requests`,
`--concurrency`, and `--questions` to change the load. Use
`--profile fast --questions 1` to measure the compact profile. The report
includes request throughput, decision throughput, client latency, and request
round-trip latency.
12 changes: 8 additions & 4 deletions docs/how-it-works.md
Original file line number Diff line number Diff line change
@@ -1,9 +1,13 @@
# How it works

The engine (`server/src/system_one_lite/engine.py`) never decodes text. It builds one full prompt
for each question. The next token must be an answer code. The engine runs a
forward pass and reads the code probabilities at that position. It generates
no tokens.
The default engine (`server/src/system_one_lite/engine.py`) never decodes text.
It builds one full prompt for each question. The next token must be an answer
code. The engine runs a forward pass and reads the code probabilities at that
position. It generates no tokens.

The optional DiffusionGemma backend keeps the same public API but uses one
shared prompt and a seeded answer canvas. See [Run with 4-bit
DiffusionGemma](diffusion-gemma.md).

This page walks through the pieces so you can read the code alongside it.

Expand Down
4 changes: 2 additions & 2 deletions docs/quickstart.md
Original file line number Diff line number Diff line change
Expand Up @@ -149,8 +149,8 @@ Things worth noticing:
never asks for money back.
- These probabilities are model scores, not calibrated odds. A value near
one does not prove that an answer is correct.
- `usage.output_tokens` is always 0. The engine never generates. Each question
uses a full independent prompt.
- `usage.output_tokens` is 0 for the default and read-only DiffusionGemma
backends. The default backend uses a full independent prompt per question.

## 4. Use the SDK (optional)

Expand Down
62 changes: 62 additions & 0 deletions docs/vllm-metal.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,62 @@
# vllm-metal backend

System One can use a separate vllm-metal server instead of loading the model
in the API process. The HTTP API and answer shapes stay the same.

Install vllm-metal in its own Homebrew environment:

```bash
brew tap vllm-project/vllm-metal https://github.com/vllm-project/vllm-metal
brew install vllm-project/vllm-metal/vllm-metal
```

Start vllm-metal with the same pinned model revision as the answer-code
registry:

```bash
vllm serve mlx-community/Qwen3-1.7B-4bit \
--revision 3b1b1768f8f8cf8351c712464f906e86c2b8269e \
--max-model-len 32769 \
--port 8000
```

The extra model-length slot is for vllM's internal output token. System One
still limits each prompt to 32,768 input tokens.

Then start System One in another terminal:

```bash
cd server
SYSTEM_ONE_BACKEND=vllm-metal \
SYSTEM_ONE_VLLM_BASE_URL=http://127.0.0.1:8000 \
uv run uvicorn system_one_lite.api:app --port 8010
```

The backend builds each prompt with System One's pinned tokenizer. It sends
the exact prompt token IDs in completion batches. It also asks vllm-metal
for the exact token IDs of every allowed answer code, then normalizes those
scores at System One's temperature. It does not use a top-k cutoff or guess a
score for a missing answer. Question IDs are still used only to route answers
and are never sent to the model.

vLLM accepts up to 128 requested token IDs in one read. A question with more
options uses more than one exact read. Each read makes vllm-metal sample one
internal token so it can return the requested log probabilities. The System
One response reports that work in `usage.output_tokens`. The local MLX backend
still reports zero output tokens.

Before inference, the backend checks both token limits. It also asks
vllm-metal to decode the full answer-code registry. A server tokenizer
mismatch stops the request instead of reading the wrong token IDs.

## Current limit

This backend does not use the seeded parallel canvas from vLLM PR 57250. Use
the [4-bit DiffusionGemma backend](diffusion-gemma.md) for that execution
shape. MLX-VLM already implements DiffusionGemma on Apple Silicon, so the
seeded read endpoint lives there instead of adapting vllm-metal's
autoregressive scheduler.

Keep vllm-metal outside the `server` uv environment. Its release pins an exact
MLX build for its native Metal extension, while System One has its own MLX
dependency.
30 changes: 27 additions & 3 deletions server/src/system_one_lite/api.py
Original file line number Diff line number Diff line change
Expand Up @@ -11,6 +11,7 @@
from fastapi import FastAPI, HTTPException

from .engine import Engine, RequestContractError
from .mlx_vlm_diffusion import MlxVlmDiffusionEngine, MlxVlmError
from .prompts import as_text, confidence, label
from .schemas import (
ChoiceAnswer,
Expand All @@ -22,11 +23,24 @@
ScoreAnswer,
Usage,
)
from .vllm_metal import VllmMetalEngine, VllmMetalError


def configured_engine():
"""Build the process-wide engine from a named profile or exact model ID."""
return Engine(os.environ.get("SYSTEM_ONE_MODEL"))
backend = os.environ.get("SYSTEM_ONE_BACKEND", "mlx")
if backend == "mlx":
return Engine(os.environ.get("SYSTEM_ONE_MODEL"))
if backend == "vllm-metal":
return VllmMetalEngine(os.environ.get("SYSTEM_ONE_MODEL"))
if backend == "mlx-vlm-diffusion":
return MlxVlmDiffusionEngine(os.environ.get("SYSTEM_ONE_MODEL"))
if backend == "mlx-vlm-diffusion-fast":
return MlxVlmDiffusionEngine(
os.environ.get("SYSTEM_ONE_MODEL"),
compact=True,
)
raise ValueError(f"unknown SYSTEM_ONE_BACKEND {backend!r}")


def labels_for(question):
Expand Down Expand Up @@ -86,13 +100,20 @@ def evaluate(req: EvaluateRequest):
raise HTTPException(503, "the inference engine is busy")

runtime_engine = app.state.engine
engine_questions = [
(as_text(question.instructions), labels_for(question)) for _, question in items
]
try:
results, input_tokens, elapsed_ms = runtime_engine.evaluate(
as_text(req.state),
[(as_text(question.instructions), labels_for(question)) for _, question in items],
engine_questions,
)
except RequestContractError as error:
raise HTTPException(422, str(error)) from error
except VllmMetalError as error:
raise HTTPException(502, str(error)) from error
except MlxVlmError as error:
raise HTTPException(502, str(error)) from error
finally:
app.state.inference_slot.release()

Expand All @@ -107,7 +128,10 @@ def evaluate(req: EvaluateRequest):
question_id: answer_for(question, probabilities)
for (question_id, question), probabilities in zip(items, results)
},
usage=Usage(input_tokens=input_tokens, output_tokens=0),
usage=Usage(
input_tokens=input_tokens,
output_tokens=runtime_engine.output_tokens_for(engine_questions),
),
)

return app
Expand Down
Loading