Skip to content

Add structured read inference backends - #4

Draft
snellingio wants to merge 4 commits into
mainfrom
feature/structured-read-backend
Draft

snellingio wants to merge 4 commits into
mainfrom
feature/structured-read-backend

Conversation

@snellingio

Copy link
Copy Markdown
Owner

Summary

  • Add two optional structured-read backends.
  • Keep the current local backend as the default.
  • Add a pinned answer-code registry for seeded canvas reads.
  • Validate remote model identity, token counts, slots, and finite scores.
  • Document setup, limits, usage counts, and the compact profile.

Tests

  • Server: 161 passed.
  • Python SDK: 4 passed.
  • JavaScript SDK: type check, tests, and build check passed.
  • Ruff, plaincheck, and Git diff checks passed.

Draft notes

The seeded-canvas path needs an external runtime branch. Its compact profile also needs task-specific quality checks before release.

@coderabbitai

coderabbitai Bot commented Sep 17, 2026

Copy link
Copy Markdown

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Add a warmed concurrency profiler for structured reads.
Report throughput, client latency, and request round-trip latency.
Accept full-read responses without encoder metadata.
Add question batch and compact profile controls.
Keep the compact path exact on the current server.
Use fixed prompts so every timed request has the same shape.
@snellingio

Copy link
Copy Markdown
Owner Author

Full request benchmark

I ran the warmed concurrency sweep on an Apple M4 Pro. The test used the pinned 26B 4-bit model and the full encoder.

Settings:

  • Three decisions per request.
  • 64 timed requests per level after two warmups.
  • Concurrency: 1, 2, 4, 8, 16, and 32.
  • Read batch wait: 2 ms.
  • Maximum read batch: 4.
  • Old pair: System One ff7a5e3 and MLX-VLM 272791b.
  • New pair: System One fdad9ff and MLX-VLM 1cab7d25.
  • The old result is one sweep. The new result is the mean of two sweeps.

Command:

.venv/bin/python -m tools.benchmark_diffusion \
  --requests 64 \
  --concurrency 1,2,4,8,16,32 \
  --questions 3 \
  --profile full \
  --json
Clients Old req/s New req/s Gain Old median New median
1 2.869 3.330 16.1% 344 ms 299 ms
2 2.888 3.361 16.4% 686 ms 594 ms
4 2.814 3.275 16.4% 1,423 ms 1,225 ms
8 2.749 3.214 16.9% 2,909 ms 2,490 ms
16 2.785 3.215 15.4% 5,697 ms 4,970 ms
32 2.783 3.227 16.0% 10,594 ms 9,871 ms

At 32 clients, p95 fell from 15.52 seconds to 9.92 seconds.

The old request used 159 prompt tokens and 15 canvas tokens. The new request used 155 prompt tokens and 3 canvas tokens.

I also tested only the MLX-VLM changes while using the new canvas on both sides. Those average changes ranged from 0.7% slower to 2.5% faster. Most of the measured gain comes from the smaller prompt and canvas.

The graph request batch did not raise total throughput in this test. Throughput stayed near 3.2 to 3.4 requests per second as concurrency rose.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant