Fast Qwen vanilla and Huihui passed a direct non-strict tool call and result continuation, but that does not prove full agent harness compatibility. Track the rollout gate for ordinary coding and voice tool loops.
Audit actual Hermes (default/Mira, including LiveKit), OMP and pi.dev requests on the installed machines. Use the existing provider compatibility knobs (OMP/pi supportsStrictMode=false), preserve schema validation in the tool executor, and avoid silently dropping guarantees a workflow needs. Cover streaming argument assembly, multiple tools, tool-result replay, invalid arguments with repair, reasoning on/off, cancellation and long-context continuation for both fast model aliases.
The pinned runtime rejects strict=true, required/named tool_choice, parallel_tool_calls=false with callable tools, and constrained JSON. Record which clients send each option and provide minimal redacted reproductions. Use automatic tool choice where semantically acceptable. A workflow requiring enforcement must remain unsupported until the separate engine work is done.
Acceptance: per-harness results, exact versions/settings and regression tests or bounded probes, updated PRIMARY_MODEL_SWITCH.md, and a clear remaining-limit matrix. Respect the 6000 borrowing guide; no installation/default change is authorized by this tracking issue.
Related: #28 and #29. Engine source: satellitedown/cinference b74044fb0a319cd2a737cb7108012e6344b96dac, docs/serving.md.
API prerequisite: https://github.com/kortexa-ai/api.server/issues/95. Separate engine work: #32 (schemas) and #33 (tool-choice enforcement).
Fast Qwen vanilla and Huihui passed a direct non-strict tool call and result continuation, but that does not prove full agent harness compatibility. Track the rollout gate for ordinary coding and voice tool loops.
Audit actual Hermes (default/Mira, including LiveKit), OMP and pi.dev requests on the installed machines. Use the existing provider compatibility knobs (OMP/pi supportsStrictMode=false), preserve schema validation in the tool executor, and avoid silently dropping guarantees a workflow needs. Cover streaming argument assembly, multiple tools, tool-result replay, invalid arguments with repair, reasoning on/off, cancellation and long-context continuation for both fast model aliases.
The pinned runtime rejects strict=true, required/named tool_choice, parallel_tool_calls=false with callable tools, and constrained JSON. Record which clients send each option and provide minimal redacted reproductions. Use automatic tool choice where semantically acceptable. A workflow requiring enforcement must remain unsupported until the separate engine work is done.
Acceptance: per-harness results, exact versions/settings and regression tests or bounded probes, updated PRIMARY_MODEL_SWITCH.md, and a clear remaining-limit matrix. Respect the 6000 borrowing guide; no installation/default change is authorized by this tracking issue.
Related: #28 and #29. Engine source: satellitedown/cinference b74044fb0a319cd2a737cb7108012e6344b96dac, docs/serving.md.
API prerequisite: https://github.com/kortexa-ai/api.server/issues/95. Separate engine work: #32 (schemas) and #33 (tool-choice enforcement).