Skip to content
View karl-kes's full-sized avatar

Block or report karl-kes

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
karl-kes/README.md

Karl Keshavarzi

Computer Engineering, University of Waterloo

I enjoy writing low-level systems and software, especially with C++. To foster a community that surrounds my passion, I founded the UWHPC design team. My research experience focuses on distributed GPU algorithms for maximal independent sets. I'm particularly interested in HFT, GPU programming, and ML infrastructure. Open to discussing new opportunities and collaborations. Feel free to reach out.

Projects

host ≡ device

xpu
A header-only C++23 library that lets the same allocation, layout, and math code compile for either CPU or CUDA. Forget about #ifdef wrapping.

u(r) = ar / (1 + br)

Variational Monte Carlo Engine
Stochastically computes ground-state energies no analytic solution can reach. Halving the error costs four times the samples, so accuracy is a compute budget.

∂T/∂t = α∇²T

3D Heat Solver
Hits 91% of theoretical memory bandwidth on GPU, 68× faster than the multithreaded, vectorized CPU at 134M cells. Validated with analytical solutions.

∇ × E = −∂B/∂t

Finite-Difference Maxwell Solver
11× faster with CUDA than the OpenMP path on 28M-cell grids. Cache-aligned SoA and SIMD cut CPU kernel time 44% over the AoS baseline.

F = Gm₁m₂/r²

N-Body Gravity Engine
Barnes-Hut gives 53× over a direct OpenMP kernel at 131K bodies. A 4th-order Yoshida integrator holds energy error to 2e-12 across 249 simulated years.

Contact

[email protected]

Pinned Loading

  1. UWHPC/variational-monte-carlo UWHPC/variational-monte-carlo Public

    Variational Monte Carlo engine in CUDA C++ for simulating the homogeneous electron gas.

    C++ 4 1

  2. UWHPC/xpu UWHPC/xpu Public

    A header-only C++23 library that lets the same allocation, layout, and math code compile for either CPU or CUDA.

    C++ 1 1

  3. archer archer Public

    An N-Body Gravitational Simulator developed in C++ and 3D rendered in Python.

    Cuda 2

  4. kestrel kestrel Public

    Finite-Difference Time-Domain (FDTD) solver for Maxwell's equations in 3D. Implemented in C++ with parallelization and interactive visualization.

    C++ 1