GreenGauge recommends the lowest-cost coding model likely to finish a GitHub issue successfully. This hackathon foundation includes:
- a Next.js issue dashboard;
- a FastAPI service with a local SQLite cache;
- a GitHub issue-opened and label-change webhook flow;
- project-level Codex lifecycle hooks plus an MCP server for per-turn metrics;
- a replaceable recommendation-engine boundary, currently backed by a deterministic placeholder.
GitHub issue opened ──webhook──▶ FastAPI ──▶ recommendation engine
└──▶ SQLite cache
Codex lifecycle hooks ──every turn──▶ FastAPI ──▶ session/work-item aggregates
Codex MCP ──rich deltas──┘
Next.js dashboard ──GET /api/v1/issues──▶ FastAPI/SQLite
The dashboard reads cached recommendations; it never runs recommendation inference during page load.
Add OPENAI_API_KEY to the repository-root .env to enable semantic analysis. GreenGauge uses pinned gpt-5-nano-2025-08-07 structured output for the five-part complexity rubric and text-embedding-3-small for description embeddings. Without a key or when OpenAI is unavailable, it stores an explicitly marked local heuristic analysis and never invents an embedding.
At issue creation or whenever its labels/body change, the API:
- separates actionable labels (
accessibility,bug,documentation,enhancement) from routing and disposition labels; - excludes questions, duplicates, invalid, and wontfix issues from benchmark routing;
- stores acceptance criteria, embeddings, the five rubric scores, model/prompt/rubric versions, confidence, and evidence;
- scores historical issues as 30% actionable-label Jaccard overlap, 50% embedding cosine similarity, and 20% complexity similarity;
- recommends the lowest similarity-weighted cost-per-green
(model, reasoning effort)combination meeting the configured success threshold.
Comparable issues below GREENGAUGE_RECOMMENDATION_MIN_SIMILARITY are ignored. If no issue clears that threshold, GreenGauge uses global completed-run evidence; if no completed evidence exists, it returns an explicitly labeled complexity-only route. Recommendations are cached in SQLite, and projected per-ticket dollar cost remains intentionally absent.
Requirements: Node 20+, Python 3.9+, and uv.
cp .env.example .env
npm install
uv sync --project apps/apiStart the API:
uv run --project apps/api uvicorn greengauge_api.main:app --reload --port 8000Start the dashboard in another terminal:
npm run dev:webOpen http://localhost:3000. The API seeds a few representative issues on its first run.
For the hackathon, use one repository token instead of building GitHub App OAuth.
- Set
GITHUB_REPOSITORY=owner/repo. For a private repository, also set a fine-grainedGITHUB_TOKENwith read-only Issues and Metadata access. - Start the API and expose port 8000 with
ngrok http 8000. - In the repository's Settings → Webhooks, create a webhook pointing to
https://<ngrok-host>/api/v1/github/webhooks. - Choose
application/json, set the same secret in GitHub andGITHUB_WEBHOOK_SECRET, and subscribe only to Issues events.
When an issue is opened, the API saves it, generates a placeholder recommendation once, and caches both records in SQLite. The same Issues webhook receives labeled and unlabeled actions; those refresh the cached issue category without rerunning recommendation inference. The engine boundary in apps/api/src/greengauge_api/services/recommendation.py is where the similarity lookup and small-model call belong later.
Use Sync GitHub on the dashboard once to import the repository's existing open issues. Public repositories work without a token; a token is recommended for private repositories and to avoid GitHub's low anonymous rate limit. Syncing generates recommendations only for newly seen issues and removes demo/stale open issues from the local cache.
The project-scoped .codex/config.toml registers the STDIO MCP server directly from source, while .codex/hooks.json provides deterministic lifecycle capture. After installing dependencies and starting the API, restart Codex in this repository, run /hooks, and approve the project hooks once. Codex then runs these automatically:
SessionStartcreates or resumes a session record.UserPromptSubmitidentifies the current turn and asks the agent to classify genuine clarification episodes without transmitting prompt content.PostToolUserecords recognized test/CI command attempts and outcomes, including explicitly named acceptance and regression suites.Stopsends one idempotent delta containing active working time, test counters, changed paths/modules, change type, and logic-branch deltas.SessionEndcloses the session.
The MCP tools complement the hook data:
attach_coding_sessionlinks the Codex session to an issue and/or PR.record_turn_metricssends one entry per runtime-reported model call plus semantic deltas before every final response. Clarification messages share a stable episode ID so follow-ups count once.finish_coding_sessionrecords whether the run went green, gave up, or hit a turn, time, or cost limit;SessionEndremains the unknown-outcome fallback.
All numeric payloads are deltas for one turn, never cumulative totals. Event IDs make retries safe, and a turn reported by both the hook and MCP increments turn_count only once. No prompt, response, command, tool output, source content, or secret is sent. If the API is unavailable, the hook queues delivery under .git/greengauge-telemetry/ and does not block Codex.
SQLite maintains one work_item_metrics row per issue (or standalone PR), any number of coding_sessions, and raw idempotent telemetry_events. This lets one PR aggregate multiple Codex sessions while preserving session_count.
On Codex Desktop, the Stop hook reads exact per-call token counters from the local session transcript and transmits only those counters and the model name—never transcript content. Other runtimes must expose exact counters through the MCP tool; the collector deliberately does not estimate them. Configure placeholder prices in GREENGAUGE_MODEL_PRICING_JSON; exact model names can use an exact rate or a matching class alias such as sol or terra, while unknown models cost $0 until configured.
The dashboard intentionally does not show a projected dollar cost for an open issue. With the current evidence, that number would imply precision the system does not have. Instead it reports historical completed-run economics across all work by model:
- Cost per green issue (CPGI): all model spend from completed attempts—including failed, abandoned, and limited runs—divided by issues where both acceptance and regression tests passed.
- Autonomous CPGI: the same spend numerator divided by green issues with zero distinct clarification episodes.
- Interruptions per green issue: distinct clarification episodes across completed attempts divided by green issues.
- Green / attempted: the observed success sample size shown alongside CPGI.
For each model call, uncached input is input - cached input - cache-write tokens. Cost applies the configured uncached, cached, cache-write, and output rates to those token buckets. Reasoning tokens are retained as a diagnostic field but are not charged separately when included in output tokens. Multiple Codex sessions using the same model on one issue/PR are treated as one model attempt for CPGI; a different model on the same work item is a separate attempt. Issue type, labels, modules, files, and change types remain attached to each work item as recommendation features, but issue type is not a hard filter on the dashboard's model comparison.
You can verify the server independently with:
npm --workspace @greengauge/codex-mcp run buildGET /healthGET /api/v1/issuesGET /api/v1/issues/{issue_number}GET /api/v1/issues/{issue_number}/analysisPOST /api/v1/github/webhooksPOST /api/v1/github/syncPOST /api/v1/recommendations/{issue_number}/refreshPOST /api/v1/telemetry/sessions/startPOST /api/v1/telemetry/sessions/{session_id}/turnsPOST /api/v1/telemetry/sessions/{session_id}/finishGET /api/v1/metrics/work-itemsGET /api/v1/metrics/work-items/{issue_number}GET /api/v1/metrics/cost-effectivenessPOST /api/v1/sessions/eventsGET /api/v1/sessions
Interactive API documentation is available at http://localhost:8000/docs.
Deploy apps/web as one Vercel project and apps/api as a second Python project. SQLite is intentionally local-only for this three-hour prototype; before production deployment, swap the small repository module for Postgres, Neon, Turso, or another persistent hosted database. No dashboard code needs to change.