AI News
Hugging Face Transformers Now Runs llama.cpp’s GGUF Quants Natively
September 22, 2026
Hugging Face Transformers can now load and serve llama.cpp’s GGUF quantized models directly, reusing the original ggml kernels to close the speed gap while opening quantized checkpoints up to fine-tuning and evaluation inside standard PyTorch workflows.
Read moreRBS-Attention Claims a 20x Prefill Speedup Without Retraining a Single Model Weight
September 21, 2026
A new training-free sparse-prefill method claims up to a 20x attention speedup for long-context LLMs by fixing a subtle flaw called mean dilution, and it barely dents accuracy in the process.
Read moreNVIDIA Retires GenAI-Perf, Ships AIPerf to Fix a Benchmarking Tool That Was Lying to Engineers
September 20, 2026
NVIDIA has replaced GenAI-Perf with AIPerf, a multiprocess benchmarking tool built to stop the load generator itself from becoming the bottleneck it’s supposed to be measuring around.
Read moreAnalysis
Guides

Cross-Post 5 Platforms in 15 Minutes: The Typefully Workflow
May 22, 2026
Native posting across 5 platforms = 60 minutes. Typefully cross-posting = 15. The 6-step workflow that saves 20 hours/month, and the 4 common mistakes to skip.









