ZGCM-1: A Fully Open and Extremely Efficient 7B Foundation Model

ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search


Authors: Jiyan He, Guang Liang, Hao Liu, Haoxiang Guan, Jinbo Sun, Junyi Guo, Wenjun Feng, Yantai Xie, Yifei Shen, Bin Shao, Chuyang Wei, Kai Chen, Kexin Zhou, Minghang Zhu, Shuxin Zheng, Tie-Yan Liu, Taine Zhao, Wenhui Zhu, Xueyin Xu, Xiaoqing Zhang, Yatao Li, Yuxuan Ren


Submitted: 11 Sep 2026 | arXiv: 2609.13356 (cs.AI)




Abstract


In this work, we present ZGCM-1, a fully open 7B dense foundation model trained from scratch with extreme data, system, and algorithmic efficiency. ZGCM-1 is founded on a core premise: compact models cannot passively memorize the open web, but can overcome parametric capacity limits by coupling deliberate internal thinking with active external tool use.


To support this paradigm across a 256K context window, we develop an end-to-end, high-efficiency open training recipe.


Key Contributions


1. Architecture & System Co-design

  • Interleaved gated sliding-window and full attention for efficient long-context processing.
  • A stable FP8 Muon optimizer that reduces compute cost while maintaining training stability.

2. Progressive Curriculum & MDP Mid-Training

  • Context scaling across 16K, 64K, and 256K stages.
  • Reformulation of interaction traces into Markov Decision Processes (MDPs) to better align training with agentic behaviors.

3. AI-Native R&D Workflow

  • Agent swarms autonomously manage cluster operations, data curation, and rapid diagnostic evaluation — a workflow that in 2026 is increasingly standard among frontier labs, but rarely disclosed at full transparency for open models.

Evaluation Highlights


Extensive evaluations show that ZGCM-1-7B is competitive across the 7B model family on general benchmarks. On several challenging mathematical reasoning and agentic search suites, it remains competitive with frontier models orders of magnitude larger, such as Qwen3-235B-A22B and GLM-5.1 — underscoring that deliberate thinking plus tool use can substitute for raw parameter scale in targeted domains.




This article summarizes an arXiv preprint. For the full paper and appendices, refer to arXiv:2609.13356.

via ArXiv AI

Related