#vllm
Vllm: 5 AI articles covering vllm news, analysis, and research
Articles
IFM Unveils K2 Horizon: Six Apache 2.0 Models from 0.9B to 375Bβ8
IFM launches K2 Horizon: six Apache 2.0 models from 0.9B to 375B-A23B with open weights, training data, and code for fully transparent AI development.
Speculative Decoding on CPUs: Nearly 4x Faster Token Generationβ10
Speculative decoding with DFlash accelerates CPU token generation up to 3.92x on Qwen3.5-9B, cutting costs by 74%. Learn how vLLM makes it possible.
Pokee AI Releases Pokee-Isaac 28B: A 10M-Token Context Agenticβ7
Pokee AI's Pokee-Isaac 28B offers a 10M-token context, on-premise deployment, and 93.3% RULER score for regulated industries.
Liquid AI Ships LFM2.5-230M with llama.cpp, MLX, vLLM, SGLang,β7
Liquid AI launches LFM2.5-230M, a 230M-parameter model for on-device inference with llama.cpp, MLX, vLLM, SGLang, and ONNX support. Achieves 213 tokens/s
Run a vLLM Server on HF Jobs in One Commandβ10
Deploy a vLLM inference server on Hugging Face Jobs in one commandβno manual setup needed. Full guide for scalable LLM serving.
