Mathematics–Computer Science @ UC San Diego
Software Engineering · Applied AI/ML · Research
I build software and machine-learning systems with an emphasis on practical engineering: modular architecture, APIs, testing, data pipelines, structured LLM workflows, and reproducible experiments.
JavaScript · Jasmine · Playwright · Docker · GitHub Actions
11-person software-engineering project for interactive HTML/CSS typing practice. My contributions include the game-state engine, sandboxed iframe rendering, WPM/accuracy/error metrics, end-screen flow, themes, practice reminders, tests, and documentation.
Python · Fetch.ai uAgents · ASI-1 · REST APIs
Hackathon prototype exploring an AI-assisted referral workflow. I contributed agent/backend integration for structured referral analysis, specialty and urgency detection, missing-information checks, preparation tasks, and multilingual patient-facing explanations.
Python · LangChain · Gemini 1.5 Flash · Pydantic · PyMuPDF
LLM-powered pipeline that extracts text from PDF resumes and converts unstructured content into validated structured JSON through a modular extraction, prompting, and schema-validation workflow.
Python · NLP · TF-IDF · Random Forest · scikit-learn
5-person COGS 108 machine-learning project using 1,465 Amazon listings to predict electronics prices from product-description text. I worked with four teammates across data preparation, EDA, modeling, interpretation, and presentation; my contributions centered on dataset sourcing/cleaning, preprocessing, TF-IDF feature extraction, EDA, and model development. The tuned Random Forest achieved R² = 0.63, MAE = $28.13, and RMSE = $80.27 on the held-out test set.
Python · scikit-learn · TF-IDF · Logistic Regression · SQLite
Research-oriented experiment codebase for referral routing, including a classical NLP baseline and infrastructure for comparing accuracy, confidence, latency, question count, and future inference-cost tradeoffs.
Python · pandas · scikit-learn · Random Forest · Statistical Testing
Analyzed 12,529 professional matches from the 2022 season, tested Blue-side advantage, and built leakage-aware models using information available at 10 minutes. The final Random Forest reached 66.68% accuracy versus 62.62% for the Logistic Regression baseline.
- Software Engineering: modular JavaScript systems, automated testing, CI/CD, Docker, REST APIs
- AI / LLM Systems: LangChain, Gemini, structured outputs, agent workflows, evaluation
- Data / ML: pandas, scikit-learn, TF-IDF, feature engineering, statistical testing, model evaluation
- Research: efficient interactive systems, uncertainty-aware workflows, reproducible experiments
Languages: Python · JavaScript · Java · C++ · SQL · HTML/CSS
AI / Data: LangChain · Gemini · scikit-learn · pandas · Pydantic · TF-IDF
Engineering: Git · GitHub Actions · Jasmine · Playwright · Docker · REST APIs

