Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
30 changes: 30 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -79,6 +79,36 @@ uv run uvicorn system_one_lite.api:app --port 8010
The default profile uses `mlx-community/Qwen3-1.7B-4bit`. On first startup, the
server downloads the pinned model revision and compiles the Metal kernels.

### Try the OpenJEV backend

The experimental NLI backend keeps the same HTTP contract. It turns each
allowed answer into a premise-hypothesis pair, batches those pairs, and maps
their entailment probabilities back to Choice, Score, and Noul responses.

```bash
cd server
uv sync --group nli
SYSTEM_ONE_BACKEND=nli uv run uvicorn system_one_lite.api:app --port 8010
```

The first run downloads `AlexWortega/openjev` and loads its
`qwen3.5-4b-nli` checkpoint through Transformers. Override its settings with
`SYSTEM_ONE_NLI_MODEL`, `SYSTEM_ONE_NLI_SUBFOLDER`,
`SYSTEM_ONE_NLI_REVISION`, `SYSTEM_ONE_NLI_DEVICE`, `SYSTEM_ONE_NLI_BATCH_SIZE`,
`SYSTEM_ONE_NLI_MAX_LENGTH`, `SYSTEM_ONE_NLI_TEMPERATURE`, and
`SYSTEM_ONE_NLI_PREFIX_CACHE`. A custom model needs an explicit
`SYSTEM_ONE_NLI_REVISION`.

This first experiment batches every candidate pair. It does not yet reuse the
shared premise cache by default, so batching lowers latency without removing
repeated premise computation. Set `SYSTEM_ONE_NLI_PREFIX_CACHE=1` to run each
candidate batch's shared causal prefix once. It then branches the option
suffixes from that cache. Float16 cached scores can differ slightly from
full-pair scores because the model runs in different chunk shapes. Cache mode
also scores one question at a time. This can reduce batching across separate
questions. The published checkpoint was trained on 256-token pairs, so longer
inputs need their own quality checks. The MLX backend remains the default.

Send a request to `POST /evaluate`:

```bash
Expand Down
25 changes: 15 additions & 10 deletions docs/api.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,8 +14,12 @@ a typed answer for every question, keyed by the IDs you chose.
| `questions` | `map<string, Question>` | yes | At least one entry. Each key is an ID you pick; the answer comes back under it. The engine never sees your IDs. |

There is no `model` field. The engine is fixed for the life of the server
process. Set `SYSTEM_ONE_MODEL=default` or `SYSTEM_ONE_MODEL=larger` before
startup. The response's `model` field reports the exact model that answered.
process. `SYSTEM_ONE_BACKEND=mlx` is the default. Set it to `nli` to use the
experimental OpenJEV backend. `SYSTEM_ONE_MODEL` selects an MLX profile or
model. `SYSTEM_ONE_NLI_MODEL`, `SYSTEM_ONE_NLI_SUBFOLDER`, and
`SYSTEM_ONE_NLI_REVISION` select an NLI checkpoint. The response's `model`
field reports the exact checkpoint that answered. A custom NLI model needs an
explicit revision.

IDs only route answers back to your code. Two requests can use the same option
names with different IDs, and renaming an ID never changes the probabilities.
Expand All @@ -35,7 +39,7 @@ Chooses one option from your list.
| --- | --- | --- | --- |
| `type` | `"choice"` | yes | |
| `instructions` | `string \| object \| array` | yes | What to decide. |
| `criteria` | `map<string, Content \| null>` | yes | At least one option. Option name → description. `null` means the name says it all. The model sees both the name and description. Options are answer-coded `A`–`Z`, then `AA`, `AB`, …; at most 578 options. |
| `criteria` | `map<string, Content \| null>` | yes | At least one option. Option name → description. `null` means the name says it all. The model sees both the name and description. The MLX backend supports at most 578 options in one Choice. The NLI backend allows 4,096 candidates across the whole request. |

```json
"language": {
Expand Down Expand Up @@ -99,7 +103,7 @@ How likely the statement is to be true.
| --- | --- | --- |
| `model` | `string` | The engine's own model ID. The default is `mlx-community/Qwen3-1.7B-4bit`. |
| `answers` | `map<string, Answer>` | One answer per question, keyed by your IDs, in request order. |
| `usage.input_tokens` | `integer` | Total tokens evaluated across the independent question prompts. |
| `usage.input_tokens` | `integer` | Tokens processed by the selected backend. MLX counts each independent question prompt. NLI counts each candidate pair, with reused prefix tokens counted once per candidate batch when prefix caching is on. |
| `usage.output_tokens` | `integer` | Always 0. Nothing is generated. |

### Choice answer
Expand Down Expand Up @@ -141,16 +145,17 @@ use an array of error objects.
| Unknown `type` | Must be `choice`, `score`, or `noul`. |
| Unknown request field | Misspelled and unsupported fields are rejected. |
| Choice with no options | A Choice needs at least one option. |
| More options in a Choice than the engine has answer codes | Each option needs its own single-token answer code (`A`–`Z`, then two-letter codes). The model registry holds 578. |
| More options in a Choice than the MLX engine has answer codes | Each option needs its own single-token answer code (`A`–`Z`, then two-letter codes). The model registry holds 578. |
| More than 4,096 NLI candidates | The NLI limit counts every Choice option, Score level, and Noul side in the request. |
| Fewer than 2 or more than 10 Score levels | The schema enforces both bounds. |
| Content or ID over its character limit | State: 100,000; instructions and each criterion: 20,000; IDs and option names: 200. |
| Request content over 1,000,000 bytes | The combined validated request must stay within the aggregate limit. |
| Prompt or request over its token limit | Each prompt may use at most 32,768 tokens; one request may use at most 131,072 input tokens. |
| MLX prompt or request over its token limit | Each prompt may use at most 32,768 tokens; one request may use at most 131,072 input tokens. |
| NLI pair over its token limit | Each formatted premise-hypothesis pair may use at most 4,096 tokens by default. Set `SYSTEM_ONE_NLI_MAX_LENGTH` to change this limit. |

If an answer code does not land as a single token, the server returns `422`.
You get an error, never a wrong answer.
With the MLX backend, an invalid single-token answer code returns `422`.

Only one inference request runs at a time. If the engine is already in use,
the server returns `503` with `detail: "the inference engine is busy"`.
Each question uses a full independent prompt. `usage.input_tokens` counts
the state again for each question.
MLX uses a full prompt for each question. NLI uses one pair for each candidate.
Without prefix caching, both backends process the state more than once.
4 changes: 4 additions & 0 deletions server/pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -24,6 +24,10 @@ test = [
labeling = [
"httpx>=0.28.1",
]
nli = [
"torch>=2.9.0",
"transformers>=5.15.0",
]
ocean = [
"numpy>=2.5.3",
]
Expand Down
27 changes: 22 additions & 5 deletions server/src/system_one_lite/api.py
Original file line number Diff line number Diff line change
Expand Up @@ -10,7 +10,7 @@

from fastapi import FastAPI, HTTPException

from .engine import Engine, RequestContractError
from .errors import RequestContractError
from .prompts import as_text, confidence, label
from .schemas import (
ChoiceAnswer,
Expand All @@ -24,9 +24,26 @@
)


def build_mlx_engine(model_id=None):
from .engine import Engine

return Engine(model_id)


def build_nli_engine(model_id=None):
from .nli_engine import NLIEngine

return NLIEngine(model_id)


def configured_engine():
"""Build the process-wide engine from a named profile or exact model ID."""
return Engine(os.environ.get("SYSTEM_ONE_MODEL"))
"""Build the selected process-wide inference backend."""
backend = os.environ.get("SYSTEM_ONE_BACKEND", "mlx")
if backend == "mlx":
return build_mlx_engine(os.environ.get("SYSTEM_ONE_MODEL"))
if backend == "nli":
return build_nli_engine(os.environ.get("SYSTEM_ONE_NLI_MODEL"))
raise ValueError(f"unknown SYSTEM_ONE_BACKEND: {backend!r}")


def labels_for(question):
Expand Down Expand Up @@ -63,8 +80,8 @@ def answer_for(question, probabilities):


def create_app(
engine: Engine | None = None,
engine_factory: Callable[[], Engine] = Engine,
engine=None,
engine_factory: Callable[[], object] = build_mlx_engine,
) -> FastAPI:
"""Create an app whose model loads during startup, not module import."""

Expand Down
9 changes: 1 addition & 8 deletions server/src/system_one_lite/engine.py
Original file line number Diff line number Diff line change
Expand Up @@ -19,6 +19,7 @@
from huggingface_hub import snapshot_download
from mlx_lm import load

from .errors import RequestContractError, TooManyOptions
from .prompts import chat_filled

QWEN3_1_7B_MODEL = "mlx-community/Qwen3-1.7B-4bit"
Expand Down Expand Up @@ -50,14 +51,6 @@
_REGISTERED_IDS = object()


class RequestContractError(ValueError):
pass


class TooManyOptions(RequestContractError):
pass


def resolve_model_id(model_id):
"""Expand a built-in profile name, or keep an exact model ID or path."""
return MODEL_PROFILES.get(model_id, model_id)
Expand Down
9 changes: 9 additions & 0 deletions server/src/system_one_lite/errors.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,9 @@
"""Errors that an inference backend may report to the HTTP layer."""


class RequestContractError(ValueError):
pass


class TooManyOptions(RequestContractError):
pass
Loading