Skip to content
View hamzaqureshi5's full-sized avatar
🌏
Available
🌏
Available

Block or report hamzaqureshi5

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
hamzaqureshi5/README.md

Hi, I'm Hamza! 👋👨🏾‍💻

AI Compiler Engineer — I turn high‑level models (PyTorch · TensorFlow · JAX · ONNX) into optimized inference artifacts across CPU, GPU, FPGA, and custom accelerators, working across the TVM · MLIR · LLVM · IREE stack to cut build time, memory footprint, and on‑device latency.

🔧 Tech Stack

Languages

Coding

C++ Python

Compilers & IR

TVM MLIR LLVM XLA Triton CUTLASS

Frameworks & Serving

PyTorch TensorFlow JAX ONNX vLLM SGLang TensorRT-LLM

Hardware & Acceleration

CUDA cuDNN cuBLAS TensorRT Jetson Thor


🚀 What I Do

  • 🔭 I’m currently an AI Software Engineer at DreamBig Semiconductor, optimizing LLM inference through graph- and kernel-level transformations in TVM and IREE for GPU, FPGA, and custom accelerators.

  • 🧱 I architect compiler passes and analyses that cut build time, memory footprint, and runtime latency while preserving model fidelity—across the TVM, LLVM, MLIR, IREE stack targeting CUDA, ROCm, TensorRT, OpenVINO, Metal, Vulkan, and custom silicon.

  • 🚗 I deploy and optimize models for edge / embedded AI accelerators, including NVIDIA Jetson Thor, for low‑latency on‑device inference.

  • 🏅 I’m a Certified Artificial Intelligence Developer.

  • 💬 Ask me about Compilers, LLM inference, Transformers, GPU/CUDA, and HPC.

  • 📧 Contact me at: [email protected].

🌐 Connect with me

LinkedIn

Pinned Loading

  1. apache/tvm apache/tvm Public

    Open Machine Learning Compiler Framework

    Python 13.6k 3.9k

  2. pytorch/pytorch pytorch/pytorch Public

    Tensors and Dynamic neural networks in Python with strong GPU acceleration

    Python 102k 28.6k

  3. llvm/llvm-project llvm/llvm-project Public

    The LLVM Project is a collection of modular and reusable compiler and toolchain technologies.

    LLVM 39.6k 18k

  4. openxla/stablehlo openxla/stablehlo Public

    Backward compatible ML compute opset inspired by HLO/MHLO

    MLIR 675 211

  5. open-etsi/smart-card-programmer open-etsi/smart-card-programmer Public

    Smart Card Programmer is a cross-platform tool designed to read, write, and manage programmable SIM/USIM cards and other ISO 7816-compliant smart cards.

    Python 1 1