Skip to content
View tc3oliver's full-sized avatar
🫠
Focusing
🫠
Focusing

Highlights

  • Pro

Block or report tc3oliver

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
tc3oliver/README.md

Oliver Yu

LLM Systems / Inference Engineer

Runtime · Serving · Prefill · Caching · Speculative Decoding · Heterogeneous Inference

I work on LLM inference runtimes and serving, mostly on Apple silicon with MLX, Core ML and the Neural Engine: cache behavior, concurrency, correctness and measured performance. When a measurement turns up a defect in a runtime I depend on, the fix goes upstream, so far to Apple coremltools and oMLX.

Portfolio · Research Notes · Email

Selected OSS Contributions

Apple coremltools

  • Open · #2876 — Synchronous MLModel.predict() holds the GIL for the entire native Core ML call, blocking other Python threads in the process that need it. Releases it around the native prediction only, with a threading regression test. Came out of the laya-apple GPU + ANE research.

oMLX

  • Merged · #3685 — The SDPA256 prefill route was chosen from live memory headroom, so an identical request gave different temperature-0 output in different processes. The route now depends on the call shape alone.
  • Merged · #3840 + #3842 — On hybrid attention + recurrent models the SpecPrefill draft cache never produced a usable hit: a restored cache was read as empty, and recurrent state was never saved at block boundaries. Two stacked fixes.
  • Merged · #3664 — The Responses API dropped namespace tool groups, the shape Codex uses for MCP servers, so their tools never reached the model. They now round-trip under their namespace.
  • Open · #3964 — A sparse prefill leaves no reusable prefix, so an append-heavy session keeps re-prefilling a growing suffix. Rebuilds that state in bounded background slices while the process is idle; off by default.
  • Open · #3811 — On mRoPE vision-language models, SpecPrefill wrote every selected token after the first at the wrong position. Selected tokens now keep their original positions.

Status is generated from GitHub by a weekly workflow. The full record is under More OSS Contributions.

Featured Systems Work

laya-apple — adaptive heterogeneous inference on Apple silicon

Serves Laya on the MLX GPU and the Apple Neural Engine at the same time, gated on parity with upstream Laya. Its research traced the rise in GPU tail latency beside a thread-placed ANE to the finished GPU result waiting for the GIL held by synchronous Core ML predict (GPU return P50 7.67 → 0.14 ms once released, in a 2×2 intervention), which led to apple/coremltools#2876. v1.5 does not depend on that patch: it runs eligible ANE work through Core ML's asynchronous API, detects a host-side slow state from its own request trace, and falls back to the known-safe 1.4 path. On one M4 Max:

  • GPU result return P50 4.28–8.60 → 0.035–0.043 ms against the 1.4 path.
  • 154 production validation episodes; 0 mismatches, routing failures, lost requests or crashes.
  • No slow state occurred in those runs. In a separate controlled test, the fallback recovered 12 of 12 slow episodes, back to 1.4 latency within 164–414 ms.

Research map · PyPI

llm-inference-systems — reproducible LLM inference systems research

What an inference optimization leaves behind for the next request. Three experiments with their data and figures: a sparse prefill that stops the reusable prefix from advancing, speculative decoding whose break-even is set by verify-cycle cost rather than acceptance rate, and background recovery of the reusable state. Threads on inference correctness and SpecPrefill admission economics continue from them. The recovery mechanism (#3964) and the SpecPrefill fixes above came out of this work.

Case study · Engineering

version-aware-code-mcp — version-aware code retrieval over MCP

Confines code search, call-graph queries and source reads to one repository, branch and revision, so a coding agent cannot quietly answer from the wrong version. Written in Go.

Other engineering work

SignalForge (self-hosted intelligence pipeline) · deepseek-v4-flash-mi300x (vLLM serving baseline on AMD MI300X) · Shouri · AI Coding Skills

Research & Writing

Python · Objective-C++ · MLX · Core ML · PyTorch · vLLM · ROCm · Go · TypeScript

More OSS Contributions

18 pull requests to projects I don't maintain: 9 merged · 9 open.

Apple coremltools — 1 open

  • ○ #2876 Release the GIL during native MLModel prediction

oMLX — 5 merged · 7 open

  • ✓ #3685 fix(attention): keep SDPA256 prefill on the bounded route
  • ✓ #3842 feat(specprefill): preserve draft recurrent state at cache boundaries
  • ✓ #3840 fix(specprefill): derive draft cache position from attention layers
  • ✓ #3746 fix(ane): avoid impossible sequence-length guidance below the ANE minimum
  • ✓ #3664 fix(responses): route namespace tool groups through to the model and back
  • ○ #4035 fix(scheduler): clear SpecPrefill state after failures and cache rejects
  • ○ #3964 feat(specprefill): recover reusable prefix state during idle time
  • ○ #3962 test: reset the image decode cache between tests
  • ○ #3811 fix(specprefill): keep selected tokens at their original positions on mRoPE VLMs
  • ○ #3792 docs(scheduler): fix the SpecPrefill note in the prefill-OOM requeue path
  • ○ #3762 feat(anthropic): accept per-request SpecPrefill overrides on /v1/messages
  • ○ #3756 fix(specprefill): preserve the full static system/tool prefix

Other projects — 4 merged · 1 open

Superseded, not counted

Pinned Loading

  1. deepseek-v4-flash-mi300x deepseek-v4-flash-mi300x Public

    Python 2

  2. llm-inference-systems llm-inference-systems Public

    Reproducible systems research on LLM inference under real workloads.

    Python

  3. laya-apple laya-apple Public

    Correctness-validated heterogeneous Laya runtime for Apple Silicon (MLX GPU + Apple Neural Engine)

    Python 15 6

  4. piship piship Public

    Build and ship your own Pi-based coding agent — branded, reproducible, no fork required.

    TypeScript 1