Site Reliability / Systems Engineer — I build, measure, and operate reliable infrastructure.
Currently AI Researcher (Infrastructure & Platform) at the Korea Environment Institute.
Reliability is treated as something measured, not claimed.
- Operating a multi-server GPU compute fleet (dual NVIDIA A40) serving LLM inference at KEI
- Building KEIwi — an on-prem GPU fleet observability & incident platform (flagship project, in progress)
- Preparing for Site Reliability Engineer roles abroad (Google / Datadog / Anthropic)
Languages
Infrastructure & Observability
Cloud
| Project | What it is |
|---|---|
| KEIwi | On-prem GPU fleet observability & incident platform — flagship, in progress |
| tpuserv | TPU-aware pod scheduling on Kubernetes — ~70% faster than default scheduling, published KIISE 2024 |
| KEIAdminSuperv | On-prem RAG for internal regulations — 100% citation guarantee, Hit@1 60.0%→82.9% |
| MineSweeper | Conflict-of-interest detection for recruitment, VLM-based extraction pipeline |
| dev-booth | Autonomous multi-agent development system — 3 LLM agents, Kanban-only coordination |
More detail on each (architecture, measured results) → seanchoi.excusa.uk
- Scheduling Techniques for Improving the Efficiency of TPU Work in a Cluster Environment — KIISE, 2024
- Research on Improving Pothole Object Detection Accuracy Through Similar-Object Data — KIISE, 2024
- Proposal of a Korean History Education Application Based on Kubernetes Clustering — 2023
- A Study on Chatbot for a Safe Harbor — 2023






