Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
46 changes: 26 additions & 20 deletions ARCHITECTURE-MAP.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@
# Architecture Map

**Topic:** architecture / provider / streaming / routing / usage / models / thinking
**Updated:** 2026-08-15
**Updated:** 2026-09-23
**Tags:** #architecture #provider #streaming #routing #usage #models #thinking #byok #autocomplete
**Supersedes:** -
**Related:** `docs/architecture/01-20260514-open-code-provider-architecture.md` · `docs/architecture/02-20260809-provider-adapter-architecture.md`
Expand Down Expand Up @@ -40,7 +40,7 @@ Total `src/` ≈ **16,310 lines** across ~109 files (excl. tests). Grouped by do
| **Core (registry/routing)** | `src/core/` — `routing.ts` (441), `registry.ts` (142), `transport.ts` (70) | ~653 | Data-driven model registry (`MODEL_REGISTRY`), transport resolution (`resolveModelRouting`), Responses/Google SSE normalization, shared `StreamRequestOptions` contract | **Pure** — no `vscode` import, no side effects |
| **Models (metadata)** | `src/models/` — `metadata.ts` (523), `modelTables.ts` (141), `metadataFetcher.ts` (102), `modelLimits.ts` (52), `modelCapabilities.ts` (16), `modelNames.ts` (29), `pricing.ts` (88) | ~951 | models.dev live metadata + bundled fallback snapshot (static data tables in `modelTables.ts`), limit/capability resolution, pricing | Live fetch may fail → bundled snapshot MUST exist |
| **Usage** | `src/usage/` — `tracker.ts` (585, thin class shell), `trackerTypes.ts` (96), `trackerWindows.ts` (140), `trackerSummary.ts` (238), `dashboard.ts` (19 barrel) + `dashboard/` (`webview.ts`, `webviewData.ts`, `webviewHtml.ts`, `state.ts`, `statusBar.ts`, `targetEditor.ts`, `tooltip.ts` ≈ 1,405), `history.ts` (378), `usage.ts` (146), `goUsageSync.ts` (128), `formatting.ts` (129), `usageProfile.ts` (74), `pricing.ts` (62) | ~3,571 | Go usage tracker (types/windows/summary split out of the old god file), per-profile tracking, CLI SQLite history reader, server-usage sync, status bar + usage webview + quick-pick (webview split into state/status/webview modules) | Server meters authoritative for Session/Weekly/Monthly; device-local for Today/Yesterday |
| **Thinking** | `src/thinking/` — `provider.ts` (78), `base.ts` (75), `resolve.ts` (66), `deepseek.ts` (53), `glm.ts` (53), `kimi.ts` (81), `minimax.ts` (54), `mimo.ts` (72), `openai.ts` (57), `qwen.ts` (103), `fallback.ts` (39), `schema.ts` (106), `payload.ts` (26), `types.ts` (51) | ~914 | Per-family thinking strategy classes + config resolution (single authority = per-model config) | **Pure** — no `vscode` import; family from registry |
| **Thinking** | `src/thinking/` — `provider.ts` (78), `base.ts` (75), `resolve.ts` (94), `deepseek.ts` (53), `glm.ts` (53), `kimi.ts` (81), `minimax.ts` (54), `mimo.ts` (72), `openai.ts` (57), `qwen.ts` (103), `fallback.ts` (39), `schema.ts` (106), `payload.ts` (26), `types.ts` (51) | ~942 | Per-family thinking strategy classes + config resolution (per-model config wins over workspace — but schema-default echoes are stripped first, `stripSchemaDefaultEcho` in `resolve.ts`, issue #226) | **Pure** — no `vscode` import; family from registry |
| **Request builders** | `src/request/` — `anthropic.ts` (233), `google.ts` (176), `types.ts` (117), `schema.ts` (94), `openai.ts` (95), `headers.ts` (110), `builders.ts` (17), `shared.ts` (11) | ~853 | Per-endpoint request-body builders + shared header builders (`x-opencode-session` / `x-opencode-request`) | `builders.ts` is the public barrel |
| **Commands** | `src/commands/` — `agentsWindow.ts` (130), `diagnostics.ts` (41), `providers.ts` (38), `thinkingPicker.ts` (31) | ~240 | Command handlers: diagnostics, agents-window BYOK bridge, provider enable/disable, thinking picker | Thin — delegates to provider/usage modules |
| **Autocomplete** | `src/autocomplete/` — `index.ts` (157), `engine.ts` (143), `provider.ts` (127), `context.ts` (89), `usage.ts` (88), `throttle.ts` (79), `prompt.ts` (60), `types.ts` (32) | ~775 | Inline code suggestions (opt-in) — FIM emulation over chat-completions, debounce/throttle, usage counters | Separate subsystem; not wired into Go cost tracker yet |
Expand Down Expand Up @@ -282,7 +282,7 @@ flowchart LR
flowchart TD
A[resolve apiKey: BYOK config → per-model cache → SecretStorage cold-start] --> B[convertMessage per message<br/>tool calls / tool results / images / thinking echo]
B --> C[flatten messages + source index]
C --> D[resolve thinking config<br/>per-model config = single authority]
C --> D[resolve thinking config<br/>per-model config wins, schema-default echo stripped<br/>issue #226]
D --> E[vision proxy: text-only model + images → relay to vision model]
E --> F[normalizeMessages + trimOldImagesFromHistoryInPlace<br/>keep MAX_HISTORY_IMAGES_KEPT=2]
F --> G[estimatePromptTokenCount → modelLimits]
Expand Down Expand Up @@ -411,23 +411,29 @@ Reuse these before writing new logic (all under `src/` root unless noted):

## 8. Tests, Tooling & CI

### Unit tests (`src/test/` — 22 files, 312 cases)

Pure/domain modules get co-located tests. **Harness:** Node's built-in `node --test` runner via `scripts/run-unit-tests.ts`, which collects the compiled `out/test/*.test.js` files; tests import from compiled `out/` modules with explicit `.js` extensions (Node16/ESM-style) and use `node:test` + `node:assert/strict`. Tests never need a live model — they cover the deterministic parts (message conversion, chunk parsing, token estimation, routing). Per-file case counts (312 total):

| Test file | Cases | | Test file | Cases |
| ----------------------------- | ----- | --- | --------------------------- | ----- |
| `thinking.test.ts` | 52 | | `metadata.test.ts` | 23 |
| `goUsageTracker.test.ts` | 51 | | `autocomplete.test.ts` | 23 |
| `utils.test.ts` | 20 | | `retry.test.ts` | 17 |
| `visionProxy.test.ts` | 17 | | `registry.test.ts` | 14 |
| `toolCallAccumulator.test.ts` | 14 | | `config.test.ts` | 13 |
| `goUsageSync.test.ts` | 9 | | `autocompleteUsage.test.ts` | 9 |
| `usageProfile.test.ts` | 9 | | `reasoningHistory.test.ts` | 9 |
| `responsesRequest.test.ts` | 8 | | `modelLimits.test.ts` | 5 |
| `imageNormalizer.test.ts` | 5 | | `apiKeyResolution.test.ts` | 3 |
| `chatParts.test.ts` | 3 | | `modelNames.test.ts` | 3 |
| `providerEnablement.test.ts` | 3 | | `tokenEstimate.test.ts` | 2 |
### Unit tests (`src/test/` — 34 files, 466 cases)

Pure/domain modules get co-located tests. **Harness:** Node's built-in `node --test` runner via `scripts/run-unit-tests.ts`, which collects the compiled `out/test/*.test.js` files; tests import from compiled `out/` modules with explicit `.js` extensions (Node16/ESM-style) and use `node:test` + `node:assert/strict`. Tests never need a live model — they cover the deterministic parts (message conversion, chunk parsing, token estimation, routing). Per-file case counts (466 total, audited 2026-09-23):

| Test file | Cases | | Test file | Cases |
| --------------------------------- | ----- | --- | ------------------------------- | ----- |
| `thinking.test.ts` | 71 | | `goUsageTrackerWindows.test.ts` | 17 |
| `retry.test.ts` | 35 | | `responsesRequest.test.ts` | 16 |
| `goUsageTracker.test.ts` | 35 | | `registry.test.ts` | 16 |
| `metadata.test.ts` | 28 | | `routing.test.ts` | 15 |
| `utils.test.ts` | 22 | | `messages.test.ts` | 15 |
| `autocomplete.test.ts` | 21 | | `extractors.test.ts` | 15 |
| `toolCallAccumulator.test.ts` | 19 | | `config.test.ts` | 15 |
| `visionProxy.test.ts` | 17 | | `openai-request.test.ts` | 12 |
| `goUsageSync.test.ts` | 10 | | `usageProfile.test.ts` | 9 |
| `sse.test.ts` | 9 | | `reasoningHistory.test.ts` | 9 |
| `autocompleteUsage.test.ts` | 9 | | `deprecatedFilter.test.ts` | 8 |
| `schema.test.ts` | 7 | | `modelLimits.test.ts` | 5 |
| `imageNormalizer.test.ts` | 5 | | `session-header.test.ts` | 4 |
| `issue216-217-regression.test.ts` | 4 | | `providerEnablement.test.ts` | 3 |
| `modelNames.test.ts` | 3 | | `engine.test.ts` | 3 |
| `chatParts.test.ts` | 3 | | `apiKeyResolution.test.ts` | 3 |
| `tokenEstimate.test.ts` | 2 | | `agentProvider.test.ts` | 1 |

### Scripts (`scripts/`)

Expand Down
2 changes: 2 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,8 @@ All notable changes to the **OpenCode Go BYOK Provider** extension are documente

- **`[Provider]` History-trim cuts are cache-stable — a session at the context ceiling keeps its prefix-cache hits (~99% instead of ~11%).** When the trimmed history landed just under the input budget, the minimal-fit trim moved the cut point on nearly every following turn — and the provider's prefix cache only reuses the bytes before the first changed message, so each moved cut re-billed the whole conversation at full input price (measured on a 614K-token session: hit rate collapsed from ~99% to ~11.4%, with only system + tools — 69,888 tokens — still cached; 207 trims fired across a ~12-hour span). Two changes now keep the cut still, with the same unit granularity and tool-group safety rules: a **low-water mark** (`budget − headroom`, new `HISTORY_TRIM_HEADROOM_*` constants: 3% of the budget, clamped to 8,192–32,768 tokens, never more than 10% of a small budget) and **cut-step alignment** to the next `HISTORY_TRIM_CUT_STEP_TOKENS` (32,768, capped at 10% of the budget) boundary of dropped payload. The step is what makes it robust: the crossing alone still hugs the mark within one unit, so sessions with ~2.7K-token units against ~1K of growth per request kept moving the cut every 1-3 requests (the live evening run: 12.4% misses for hours, hit rate down to 37-68%), while a smaller-unit morning session only looked stable by luck (large tool-result units). Simulation with production parameters: cut moves fall from 84/300 to 12/300 (evening regime) and 244/300 to 14/300 (smaller-unit regime). Four tests pin the low-water landing, the no-re-trim behavior, the re-supplied-history shape (constant cut → nested payload prefixes), and the cut-step stability; the first live ceiling crossing confirmed 15 trims with every landing ≤ 595,534 tokens (budget 613,952) and high-context misses down from 100/243 pre-fix to 3/55. Documented in `docs/issues/101-20260920-history-trim-cache-hysteresis.md`.

- **`[Thinking]` Global `opencodego.thinking.*` settings now take effect for models without a per-model pick (#226).** VS Code merges our picker schema defaults into the per-model `modelConfiguration` on every request, so any reasoning-capable model the user never configured arrived with `reasoningEffort: "off"` attached — and the resolver treated any delivered `modelConfiguration` as the single authority, letting the echoed `"off"` beat the global setting every time (the #214 symptom; diagnosis by @nickchomey). `resolveThinkingConfig` now strips override keys equal to the family's picker schema default before applying them — lossless, because VS Code itself strips default-equal values when persisting user picks, so such a value can never be a genuine user choice. Non-default per-model picks still win; the Agents-window default path is untouched (removing the schema default instead would have made host-side fallbacks pick `medium`/`high`). Documented in `docs/issues/100-20260923-issue226-thinking-default-echo.md`.

- **`[Models]` Bundled offline fallback data synced to models.dev (2026-09-08 snapshot).** Three drift classes fixed in the offline-only tables: (1) stale context limits — `kimi-k2.7-code` 256K → 262,144, `minimax-m3` / `qwen3.6-plus` / `qwen3.7-plus` → 1M (models.dev now differentiates Go vs Zen and the bundled Go values matched the Zen ones; the deliberately-capped `deepseek-v4-flash` 131,072 output and Go `minimax-m2.5` 65,536 output are kept); (2) wrong static pricing in the usage tracker — `mimo-v2.5-pro` and `deepseek-v4-pro` were 2–4x over-reported, `mimo-v2-omni` / `deepseek-v4-flash` under-reported, and `kimi-k3` ($3/$15) fell back to a 6x-under-estimating generic price; the table was rewritten from the current registry and 15 new models added; (3) the Go `fallbackModels` catalog still listed the retired `hy3-preview` and pre-June models — refreshed to the 27-model curated active set. Stale tests referencing `hy3-preview` updated to `hy3`. Documented in `docs/issues/100-20260908-bundled-model-data-sync.md`.

## [0.7.5] — 2026-09-08
Expand Down
Loading
Loading