Pretraining and inference code for a large-scale depth-recurrent language model
-
Updated
Dec 29, 2025 - Python
Pretraining and inference code for a large-scale depth-recurrent language model
Research and training stack for AVA — a tool-using, memory-aware virtual assistant targeting 4 GB VRAM. Spans custom transformers, verifier-RL, external memory, multi-domain benchmarks, and Gemma 4 inference optimization.
Awesome list of papers, code, models and blogs on Looped / recurrent-depth / weight-tied Transformers — depth as a third scaling axis
Awesome list + deep-dive reports on looped / recurrent-depth transformers (2018-2026): the Loop technology behind GPT-6 Astra. 122 papers, ZH/EN reports, translated PDFs, survey, slides.
Recurrent-depth transformer, fixed. Fork of kyegomez/OpenMythos with scatter-based MoE (2.94x faster), proper ACT halting, DeepSeekMoE load balancing, SDPA kernels, and a working training loop.
Loop pretrained language model layers for deeper reasoning without new parameters. Open training, evaluation, and export to standard Transformers and vLLM checkpoints.
Does the Universal/Looped-Transformer trick (reuse one block N times instead of stacking N) work on Mamba? Weight-shared depth in a state-space model, with honest per-seed results.
Recurrent-depth, P/R/C looped transformer with MoE, GQA, and stable recurrence
To associate your repository with the recurrent-depth topic, visit your repo's landing page and select "manage topics."