Skip to content
View JeremyGracey-AI's full-sized avatar
:electron:
Looking for new opportunities
:electron:
Looking for new opportunities

Highlights

  • Pro

Block or report JeremyGracey-AI

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
JeremyGracey-AI/README.md
Jeremy Gracey. Agent systems that hold up under audit.

jeremygracey.ai · LinkedIn · Hugging Face · [email protected] · [email protected]

Founder and applied AI engineer in Seattle. I build agent systems that hold up under audit: pipelines that log their decisions, memory with governance and replay, RAG that cites the exact page behind every claim.

Before software I worked emergency medicine, acute psychiatric care, and special education. Nobody in those rooms accepts "trust me" as an answer, and I never learned to accept it from software either. Everything below is built to that standard.

Built with Roboflow

prevera-guardian-lidar is a fall-detection research prototype on a Jetson Orin Nano: a floor-level 2D LIDAR that finds a person lying down from geometry alone, plus RF-DETR served on the device by Roboflow Inference for the poses the LIDAR misses.

Segment B of the field test: the floor LIDAR sees two small clusters and raises no event, while RF-DETR finds the person in 33 of 33 frames on both cameras.

  • The blind spot. Lying end-on with feet toward the sensor, a person is two 0.3 m clusters to the LIDAR, and the detector raised 0 events in 33 seconds. Stock RF-DETR found the person in 33 of 33 frames on each camera. That ran on recorded frames; it is not in the live loop yet.
  • The fine-tune. I fine-tuned RF-DETR on Roboflow to predict pose as a class, on public fall data, under a plan committed before training started. The model the plan picked passed its bars on the public test split. The device-sized model passed its bars on room frames from the counter camera; from the floor camera it read a head-first lie-down as standing in 30 of 30 frames. Both results are in the repo.
  • Upstream. The JetPack 6.2.0 Inference image returned HTTP 500 for every RF-DETR request under the documented --read-only container command. I sent the one-line fix and a unit test as roboflow/inference#3072; Roboflow merged it on 2026-10-02.

Every evaluation is declared before it runs and the failures are published beside the passes. One subject, one room; not a medical device; patent pending.

tests release

What I build

Seven more threads, each with public code behind it.

Forecasting. Prime Radiant is a CDC FluSight forecaster: it predicts weekly flu hospital admissions for all 53 FluSight locations, 23 quantiles at a time, and is registered with the CDC FluSight hub as JGracey-prime_radiant. It is built vintage-honest: the model never sees data dated after the forecast moment, and the backtest replays the same pipeline the weekly job runs, over 85 weekly origins across three seasons. See the dashboard, or pip install prime-radiant.

Agent infrastructure. The plumbing that makes autonomous agents accountable. Compass BlackBox IQ is a flight recorder for agents: git-backed memory, decision records, and a skill forge, exposed over MCP. llm-council-mcp runs multi-model deliberation as an MCP server for Claude Code and ships on PyPI as mcp-llm-council.

Agent governance. governance-drift-researcher detects drift in an AI-agent estate: every finding carries verifiable evidence, findings that can't be re-verified are dropped, and nothing publishes without human sign-off. Run it live on WeaveMind Cloud for about $0.03, or pip install governance-drift. guss is the same philosophy on 7 watts: a Jetson-hosted agent that monitors itself, heals itself, paper-trades against a benchmark, and publishes its own dashboard.

Compilers and GPU performance. triton-kernel-lab is hand-written Triton kernels with honest benchmarks on a Jetson Orin Nano, plus a working study of LLVM, MLIR, and TorchDynamo/Inductor.

Neurotech and edge hardware. nexus-neuromirror is offline-first EEG neurofeedback for the Mind Media NeXus-10, from EDF verification to a live dashboard. pip install nexus-neuromirror.

Safety evaluation. KŪPUNA-AI Bench ("Let's talk story.") scores AI replies to older adults for both overrefusal and harmful compliance, on an S0–S3 severity scale, with a judge from a model family not under test. I'm the technical lead, working with The Gerontechnology Foundation. Pre-pilot; DOI 10.5281/zenodo.22667583.

Clinical AI on open standards. clinical-ai-agent is citation-traceable decision support on SMART on FHIR: a five-agent pipeline with dual citations back to patient data and clinical sources. Hospital-Readmission-Prediction-Model covers the classical ML side, synthetic EHR data through SHAP-based clinical interpretation. dbq-qualifier-agent maps a VA knee exam to the 38 CFR Part 4 rating criteria and cites the page and field behind every finding: decision support for a VSO or attorney, never an automated rating. phi-scrub is PHI/PII redaction in Rust, on PyPI and crates.io (pip install phi-scrub).

Off GitHub: a custom-engine retrieval agent I built for a client and took through their tenant-admin approval into Microsoft 365 Copilot Chat.

Beyond the pins

The pins are the front door. These hold up past the first click too.

Six more repositories
  • provenance: RAG over textbook page images that verifies every claim against the exact page that proves it. Cohere Embed v4 retrieval, Claude vision answers.
  • calibrated-readiness: multi-agent exam-readiness scoring with a 60-second reliability-diagram check. Microsoft Foundry Agent Framework + Foundry IQ.
  • rag-healthcare-ai: fully local medical Q&A over the Merck Manual. Mistral-7B on llama.cpp, ChromaDB, no API in the loop.
  • healthcare-vjepa2-agent: V-JEPA 2 and Claude read medical procedure videos and generate teaching material (step breakdowns, narration, quizzes, safety notes).

Classical ML lives in helmnet, VGG-16 transfer learning with thresholds tuned for zero-harm deployment, and renewind-predictive-maintenance, seven Keras architectures against imbalanced turbine sensor data.

Now

Consulting through the Claude Partner Network. Heading to Roboflow's Visual Intelligence Summit in San Francisco on October 22. Digging into what neuropsychology's models of memory can teach agent memory design.

jeremygracey.ai

Pinned Loading

  1. prevera-guardian-lidar prevera-guardian-lidar Public

    Privacy-preserving LIDAR fall-detection prototype (ROS 2 Humble, Jetson Orin Nano) with on-device RF-DETR. Apache-2.0. Patent pending.

    Python 1

  2. kupuna-bench-public kupuna-bench-public Public

    KŪPUNA-AI Bench "Let's talk story.": measuring overrefusal and harmful compliance in AI conversations with older adults (public mirror of the working repository)

    Python 1

  3. nexus-neuromirror nexus-neuromirror Public

    Offline-first EEG neurofeedback prototype for the Mind Media NeXus-10: EDF/EDF+ verifier, montage configs, experiment scaffold, and a web dashboard.

    TypeScript 2

  4. llm-council-mcp llm-council-mcp Public

    Multi-model LLM Council (Karpathy-style 3-stage deliberation over OpenRouter), exposed as an MCP server for Claude Code.

    Python 1

  5. governance-drift-researcher governance-drift-researcher Public

    Deterministic AI-governance drift detection: evidence-backed findings, stated coverage gaps, human-approved publishing. Runs on WeaveMind Cloud (~$0.03/run); includes the first external evaluation …

    Python 1

  6. guss guss Public

    A governed autonomous agent on 7 watts — Jetson Orin Nano + local Qwen3-8B that monitors itself, heals itself, paper-trades against a benchmark, and publishes its own dashboard. Public case study.

    Python 1