#moe

Moe: 8 AI articles covering moe news, analysis, and research

Articles

GLM-5.3-Flash vs Qwen3.8-Flash-Next: Two Chinese AI Labs Independently Converge on the Same Model Architecture
From Mixtral to Kimi K3: The Evolution of Mixture-of-Experts Models
Meet FreeToken: An Edge-Native MoE Serving Engine that Runs 753B GLM-5.2 on a Single Workstation GPU
NVIDIA Unveils Nemotron 3.5 Lightning: A 30B Open MoE with 3B
Cursor Open-Sources Mixture-of-Kittens (MoK): A Deterministic
Alibaba Qwen Unveils Qwen3.8-Max: A 2.4-Trillion-Parameter MoE
Thinking Machines Lab Releases Inkling-Small: A 276B Total, 12B
Knowledge Injection in MoE? Expert-Aware Contrast Decoding for