Skip to content
View ArtyomITA's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report ArtyomITA

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Pinned Loading

  1. llm2vec-8gb-offload-lossless-batching llm2vec-8gb-offload-lossless-batching Public

    Run LLM2Vec (Llama-3-8B) on an 8 GB GPU via RAM offload, and batch it without changing the embeddings (padding shifts them: cosine 0.76 vs 0.999956)

    Python

  2. bonsai-webgpu-tiled-vae-fix bonsai-webgpu-tiled-vae-fix Public

    Tiled VAE fix: run PrismML's Bonsai-Image 4B WebGPU at 512px+ on GPUs without shader-f16 (GTX 10xx / Pascal & older). Root cause analysis + 1.4KB patch, no quality loss on f16 devices.

    JavaScript 2

  3. hildanext hildanext Public

    AR → discrete diffusion LLM backend (LLaDA 2.1 recipe, Qwen3-0.6B, Pascal/sm_61)

    Python 1

  4. llama.cpp llama.cpp Public

    Forked from ggml-org/llama.cpp

    LLM inference in C/C++

    C++

  5. stt-whisper-pad stt-whisper-pad Public

    Whisper Pad: a small Windows desktop app for local speech-to-text with faster-whisper

    Python 1

  6. flytollm flytollm Public

    The whole male Drosophila CNS connectome (166,700 neurons) used uncut as the spiking core of a language model. Scope, numbers, honest controls, pipeline animation.

    Python