Skip to content
View YuchenHe985's full-sized avatar
  • Joined Aug 30, 2026

Block or report YuchenHe985

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
YuchenHe985/README.md

Yuchen He

GPU / AI Systems · Systems Performance · Embedded / SoC

M.S.E. Electrical Engineering @ University of Pennsylvania (2026–2028)

Building hardware-aware software from real measurements: multi-GPU inference, reliable serving, performance-oriented C++, and register-level systems.

Python · Go · C/C++ · CUDA/NCCL · Linux

Open to Summer 2027 internships in GPU/AI systems, embedded software, and hardware-aware performance engineering.

yuchenhe985.github.io · LinkedIn · [email protected]

Selected systems work

Project Engineering focus Evidence
RadixGates Failure-tolerant Go gateway for multi-GPU LLM serving: prefix affinity, health-aware failover, circuit breaking, admission control, and OpenAI-compatible routing 85.2% → 99.7% clean completions during node-crash tests; 55 tests under go test -race; real evaluation on 4× RTX 4090 and 4× A100
llm-serving-eval-kit GPU-memory sizing, topology interpretation, startup-log diagnosis, repeated benchmark matrices, and confounder-aware comparison 11 failure signatures, bootstrap intervals, cost/SLO reporting, and 39 unit/end-to-end tests
cdc-chunker C++17 Gear/Rabin content-defined chunking with streaming and bit-identical parallel output About 2.0 GB/s sequential and 7.4 GB/s with 8 threads on Apple M1; 31 tests plus 23/23 mutation checks
llm-finetune-lab Local text-to-SQL feasibility study: LoRA/DDP, execution-based evaluation, GGUF deployment, and guardrail analysis 398 MB Q4 model, 82.6% execution accuracy, 1.76× two-GPU training speedup; evidence-based human-in-the-loop release decision

What connects the projects

  • Measure before optimizing: record topology, software versions, failure modes, repetitions, and uncertainty before attributing a result.
  • Design for failure: make overload, node loss, partial streams, and invalid model output explicit instead of hiding them behind averages.
  • Follow the hardware: connect routing, cache locality, communication topology, memory limits, and parallelism choices to observed behavior.

Current direction

I am extending this systems work toward embedded and accelerator design through register-level firmware, FPGA/SoC architecture, and digital IC/VLSI coursework. My earlier engineering experience includes PySpark/Hive pipelines and data-quality checks for a regional logistics forecasting platform.

Popular repositories Loading

  1. YuchenHe985 YuchenHe985 Public

    Profile README

  2. cdc-chunker cdc-chunker Public

    Content-defined chunking in C++17: Gear and Rabin rolling hashes, streaming API, differential tests and mutation checks

    C++

  3. radixgates radixgates Public

    Go gateway for multi-GPU LLM serving on SGLang: prefix-affinity routing, failover, circuit breaking, admission control, and a failure-injection benchmark against the original.

    Go

  4. llm-serving-eval-kit llm-serving-eval-kit Public

    Sizing, failure diagnosis and reproducible benchmarking for LLM inference deployments: GPU memory estimator, startup-log diagnoser, benchmark matrices with confidence intervals, confounder-aware co…

    Python

  5. llm-finetune-lab llm-finetune-lab Public

    Feasibility study of an on-device text-to-SQL assistant: LoRA fine-tuning with PyTorch DDP, evaluation by executing the generated SQL, guardrails against wrong answers, GGUF quantization.

    Python

  6. yuchenhe985.github.io yuchenhe985.github.io Public

    Personal site of Yuchen He: systems under AI, measured.

    Python