sevren-ai
Popular repositories Loading
-
FlashPrefillv2-MLX
FlashPrefillv2-MLX PublicFlashPrefill V2 block-sparse long-context prefill attention for MLX on Apple silicon.
-
mlx-heimr-sparse-down
mlx-heimr-sparse-down PublicSparse MoE down projection, one-layout prefill and MTP speculative decoding for NVIDIA Nemotron 3.5 Lightning 30B A3B on MLX, from the stock mlx-community 4-bit checkpoint
Repositories
Showing 2 of 2 repositories
- mlx-heimr-sparse-down Public
Sparse MoE down projection, one-layout prefill and MTP speculative decoding for NVIDIA Nemotron 3.5 Lightning 30B A3B on MLX, from the stock mlx-community 4-bit checkpoint
- FlashPrefillv2-MLX Public
FlashPrefill V2 block-sparse long-context prefill attention for MLX on Apple silicon.
People
This organization has no public members. You must be a member to see who’s a part of this organization.
Top languages
Loading…
Most used topics
Loading…