#moe
Moe: 8 AI articles covering moe news, analysis, and research
Articles
GLM-5.3-Flash vs Qwen3.8-Flash-Next: Two Chinese AI Labs Independently Converge on the Same Model Architecture⭐8
Two Chinese labs independently release similar frontier MoE models: GLM-5.3-Flash and Qwen3.8-Flash-Next, sharing a hybrid attention recipe.
From Mixtral to Kimi K3: The Evolution of Mixture-of-Experts Models⭐8
From Mixtral to Kimi K3, explore Mixture-of-Experts evolution from 8 to 896 experts per layer, compression, and stability for trillions of parameters.
Meet FreeToken: An Edge-Native MoE Serving Engine that Runs 753B GLM-5.2 on a Single Workstation GPU⭐7
FreeToken enables running massive 753B-parameter MoE models on a single workstation GPU, democratizing high-scale AI inference for individual developers.
NVIDIA Unveils Nemotron 3.5 Lightning: A 30B Open MoE with 3B⭐8
NVIDIA launches Nemotron 3.5 Lightning, a 30B open MoE with 3B active params, plus NeMo Switchyard router for faster, efficient AI agents.
Cursor Open-Sources Mixture-of-Kittens (MoK): A Deterministic⭐7
Cursor open-sources MoK, a deterministic MoE training megakernel for NVIDIA GB300 NVL72 racks. It boosts throughput up to 2.37x for Composer models.
Alibaba Qwen Unveils Qwen3.8-Max: A 2.4-Trillion-Parameter MoE⭐6
Alibaba unveils Qwen3.8-Max, a 2.4T-parameter MoE model with open weights next week, plus Qwen3.8-27B for practical deployment.
Thinking Machines Lab Releases Inkling-Small: A 276B Total, 12B⭐8
Thinking Machines Lab unveils Inkling-Small, an open 276B-parameter MoE with 12B active, multimodal support, 1M context, and Apache 2.0 weights.
Knowledge Injection in MoE? Expert-Aware Contrast Decoding for⭐7
Expert-aware contrast decoding leverages distinct expert activation patterns to mitigate LLM hallucinations in MoE models, outperforming baselines across benchm...
