Add structured read inference backends - #4
snellingio wants to merge 4 commits into
Conversation
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: trueThanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Add a warmed concurrency profiler for structured reads. Report throughput, client latency, and request round-trip latency. Accept full-read responses without encoder metadata.
Add question batch and compact profile controls. Keep the compact path exact on the current server. Use fixed prompts so every timed request has the same shape.
Full request benchmarkI ran the warmed concurrency sweep on an Apple M4 Pro. The test used the pinned 26B 4-bit model and the full encoder. Settings:
Command: .venv/bin/python -m tools.benchmark_diffusion \
--requests 64 \
--concurrency 1,2,4,8,16,32 \
--questions 3 \
--profile full \
--json
At 32 clients, p95 fell from 15.52 seconds to 9.92 seconds. The old request used 159 prompt tokens and 15 canvas tokens. The new request used 155 prompt tokens and 3 canvas tokens. I also tested only the MLX-VLM changes while using the new canvas on both sides. Those average changes ranged from 0.7% slower to 2.5% faster. Most of the measured gain comes from the smaller prompt and canvas. The graph request batch did not raise total throughput in this test. Throughput stayed near 3.2 to 3.4 requests per second as concurrency rose. |
Summary
Tests
Draft notes
The seeded-canvas path needs an external runtime branch. Its compact profile also needs task-specific quality checks before release.