#llm inference
Llm Inference: 4 AI articles covering llm inference news, analysis, and research
Articles
DeepSeek-V4.1-Flash Launches with 1M Context, FP4 KV Cache, and Cross-Layer Attention ReuseNEWβ9
DeepSeek-V4.1-Flash launches with 1M context, FP4 KV cache, and cross-layer attention reuseβcutting KV memory to 890 bytes per token for input-heavy agent workl...
Workload-Driven HBF Substrate for Capacity-Scalable LLM Inferenceβ8
Workload-driven HBF substrate enables capacity-scalable LLM inference with 3.2x bandwidth gains and 40% lower energy.
Accelerating LLM Inference via Vector Index Based Output Embeddingsβ10
Vector index-based output embeddings accelerate LLM decoding by replacing dense projections with HNSW retrieval, boosting throughput up to 82% while preserving ...
Hybrid HBM-HBF Architecture in LLM Inferenceβ9
Hybrid HBM-HBF architecture for LLM inference cuts memory bottlenecks. Oxford research combines high-bandwidth memory and flash for scalable AI performance.
