Skip to content

feat(local): support vLLM and SGLang OpenAI-compatible servers - #228

Open
davidliuk wants to merge 1 commit into
mainfrom
codex/issue-203-local-openai-compatible
Open

davidliuk wants to merge 1 commit into
mainfrom
codex/issue-203-local-openai-compatible

Conversation

@davidliuk

@davidliuk davidliuk commented Sep 13, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Closes #203.

  • Discover models from Ollama /api/tags or OpenAI-compatible /v1/models; support vLLM/SGLang root URLs, explicit /v1, and proxy prefixes.
  • Use the same normalized base and optional server-side LOCAL_GPU_API_KEY for streaming chat/tool calls. Keep loopback-only URL validation and reject redirects.
  • Make Settings provider-neutral, expose only nonsecret URL defaults, hide Ollama-only model pulls on compatible servers, and synchronize changed URL/model settings with already-mounted chat.
  • Remove the old migration that silently rewrote localhost:8000 into Ollama's port.
  • Document CPU/local/tunneled servers and tool-call requirements.

Validation

  • 14 focused tests covering validation, auth, discovery, prefixes, Ollama metadata, request construction, non-Ollama pull rejection, and an actual queryLocalGPU agent loop (fragmented SSE tool-call → temp-file Read → continued streamed answer).
  • Browser regression: Settings loads backend URL, tests/saves configuration and updates already-mounted chat model state; no Ollama pull controls and no page errors. Uses mock model HTTP responses, not a live GPU/model.
  • npm run typecheck and npm run build pass.
  • Full committed-branch suite passed under isolated HOME: 232 tests. The temporary verification checkout shares node_modules via symlink, so this run used NODE_OPTIONS=--preserve-symlinks for the existing bundled-CLI path assertion. [Feature]: more local llm support other than ollama #203 adds 14 Vitest tests plus one browser test.
  • Local default-HOME full suite initially hit an existing Gemini fixture/config interaction; clean HOME avoids it. No change to Gemini code in this PR.

Scope

No real vLLM/SGLang service, model download, paid inference, or user credential setup was run. The API key is configured in server .env, never browser storage. This PR does not merge or close any other pending issue/PR.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Feature]: more local llm support other than ollama

1 participant