#mixture-of-experts
Mixture-Of-Experts: 13 AI articles covering mixture-of-experts news, analysis, and research
Articles
GLM-5.3-Flash vs Qwen3.8-Flash-Next: Two Chinese AI Labs Independently Converge on the Same Model Architectureโญ8
Two Chinese labs independently release similar frontier MoE models: GLM-5.3-Flash and Qwen3.8-Flash-Next, sharing a hybrid attention recipe.
From Mixtral to Kimi K3: The Evolution of Mixture-of-Experts Modelsโญ8
From Mixtral to Kimi K3, explore Mixture-of-Experts evolution from 8 to 896 experts per layer, compression, and stability for trillions of parameters.
Z.ai Unveils GLM-5.3-Flash: A 320B-A18B Multimodal MoE Model with 1M-Token Contextโญ7
GLM-5.3-Flash: Z.ai's 320B MoE model, MIT-licensed, multimodal, 1M-token context, rivals Claude at 1/10 the cost.
NVIDIA Unveils Nemotron 3.5 Lightning: A 30B Open MoE with 3Bโญ8
NVIDIA launches Nemotron 3.5 Lightning, a 30B open MoE with 3B active params, plus NeMo Switchyard router for faster, efficient AI agents.
TEXAS: Task-Expert-Aware Supervision for Downstreamโญ9
TEXAS method improves MoE LLM adaptation by identifying task-relevant experts and optimizing token supervision, boosting performance across benchmarks.
Cursor Open-Sources Mixture-of-Kittens (MoK): A Deterministicโญ7
Cursor open-sources MoK, a deterministic MoE training megakernel for NVIDIA GB300 NVL72 racks. It boosts throughput up to 2.37x for Composer models.
Alibaba Qwen Unveils Qwen3.8-Max: A 2.4-Trillion-Parameter MoEโญ6
Alibaba unveils Qwen3.8-Max, a 2.4T-parameter MoE model with open weights next week, plus Qwen3.8-27B for practical deployment.
Topology-Aware Data Movement for Disaggregated GPU Inferenceโญ10
Topology-aware KV cache transfer cuts GPU inference latency by 3-18x via NVLink, InfiniBand, and CXL optimizations.
Thinking Machines Lab Releases Inkling-Small: A 276B Total, 12Bโญ8
Thinking Machines Lab unveils Inkling-Small, an open 276B-parameter MoE with 12B active, multimodal support, 1M context, and Apache 2.0 weights.
AMD Releases Instella-MoE-16B-A3B: A Fully Openโญ7
AMD unveils Instella-MoE-16B-A3B, a fully open MoE LLM with 2.8B active parameters, trained on Instinct GPUs, offering complete transparency for research.
