Harvesting forks from GitHub — first visit takes a few seconds…
Harvesting forks from GitHub — first visit takes a few seconds…
No comments yet. Be the first.


llama.cpp fork with TurboQuant WHT-rotated KV cache & weight compression + Gemma 4 MTP and Qwen 3.6 NextN speculative decoding (+30-50% throughput).
View on GitHub ↗LLM inference in C/C++
## Merge of TurboQuant & Vanilla llama.cpp's MTP recency; buffed with help from the Atomic fork.
View on GitHub ↗"TurboQuant llama.cpp fork synced with upstream master (Apr 2026) — adds Gemma 4 support"
View on GitHub ↗
Comments (0)