This repository shows how to build a DeepSeek language model from scratch using PyTorch. It includes clean, well-structured implementations of advanced attention techniques such as key–value caching for fast decoding, multi-query attention, grouped-query attention, and multi-head latent attention.
transformers pytorch multi-query-attention grouped-query-attention multi-head-latent-attention deepseek-from-scratch
-
Updated
Jan 10, 2026 - Jupyter Notebook