AI Software Engineer | LLM Evaluation Specialist | Generative AI | Python | JavaScript
Software Engineer building AI training data and coding challenges. I work at the intersection of software engineering and AI evaluation - writing coding problems, reviewing AI-generated code, and building the datasets that help LLMs get better at reasoning.
- π οΈ Design and review coding challenges for AI training platforms (Python, JavaScript), using Docker for containerized setups and Git for version control
- π§ Evaluate LLM outputs for correctness, bias, and reasoning quality across platforms like Turing and Handshake
- π Build bilingual (Hindi/English) datasets and Chain-of-Thought math reasoning content for LLM fine-tuning
- π Work daily in Linux/Bash environments
Python JavaScript Bash Docker Git Linux
β’ AI Agents β’ LLM Evaluation β’ Code Generation β’ Benchmark Creation β’ Prompt Engineering β’ Search Quality β’ NLP β’ Reinforcement Learning from Human Feedback
Working on AI code evaluation and technical assessment design as a contractor with Handshake, alongside LLM evaluation work with Turing.
- π AI Coding Challenges β Production-style Python & JavaScript coding challenges with automated testing
- π€ LLM Evaluation Toolkit β Rubric-based framework for evaluating AI model responses
- π Python Automation β Dataset validation and workflow automation tools
- π³ Docker Playground β Containerized FastAPI applications and Docker examples
- π§ Linux Bash Toolkit β Practical Bash scripts for automation and system utilities
- β¨ Prompt Engineering Lab β Prompt patterns, evaluation templates, and structured prompting examples
- Write clean, testable code
- Prefer reproducible environments
- Build reliable evaluation pipelines
- Document decisions clearly
- Keep learning through hands-on projects
π« [email protected] | LinkedIn