Google DeepMind Unveils Gemini 4 Argon with 1M Output Tokens for

Google DeepMind has announced Gemini 4 Argon, the first model of the Gemini 4 generation and its newest frontier model. Argon targets long-horizon software engineering, enterprise knowledge work in legal and finance, and cybersecurity defense. Its most significant technical leap is output length: it can generate up to 1 million tokens in a single response, a substantial increase from the 64K token limit on earlier Gemini models. That expanded output capacity positions Argon for tasks that require sustained, coherent generation, such as large-scale code synthesis, extended document drafting, and multi-step reasoning workflows. In the context of 2026, where agentic AI systems are increasingly expected to operate over hours or days rather than single prompts, the 1M output limit represents a deliberate push toward long-form autonomous execution.

What Google Announced

Google DeepMind described Argon as built for complex workflows across coding, enterprise knowledge work, and cybersecurity defense. The model is being positioned not just as a general-purpose chatbot but as a production-grade engine for high-stakes, multi-step tasks.

Google is taking a phased approach to release. The company is participating in the U.S. government’s voluntary process for pre-release model access, a framework that has become more formalized by 2026 as regulators and labs collaborate on safety evaluations. Google will gather feedback from early testers and iterate on guardrails before a wider release.

Pricing is already public. Argon launches at an introductory $2 per 1 million input tokens and $10 per 1 million output tokens. Cached input tokens receive a 95% discount, bringing the effective rate to $0.10 per 1 million. After the introductory period, pricing moves to $4 per 1 million input tokens and $20 per 1 million output tokens. Logan Kilpatrick confirmed the introductory $2 input and $10 output pricing.

Why the 1M Output Limit Matters

The 1M output token ceiling is not just a numerical upgrade; it changes the class of tasks the model can handle in a single pass. Earlier models with 64K output limits often required chunking, recursive prompting, or external orchestration to produce long artifacts such as full codebases, legal contracts, or security incident reports. Argon’s expanded output window reduces the need for those workarounds, enabling more coherent and contextually consistent long-form generation.

This matters especially for long-horizon software engineering, where a model may need to generate or refactor thousands of lines of code while maintaining architectural consistency. It also matters for enterprise knowledge work in legal and finance, where documents can run hundreds of pages and require precise cross-referencing. In cybersecurity defense, longer outputs allow for detailed threat analysis, playbook generation, and forensic reporting without losing thread.

Positioning in the 2026 AI Landscape

By 2026, the frontier model race has shifted from raw parameter counts to practical output length, tool use, and agentic reliability. Competitors have been pushing context windows and output limits as key differentiators. Argon’s 1M output tokens place it among the most capable models for sustained generation tasks, though the real-world utility will depend on how well it maintains coherence, factuality, and instruction adherence over such long outputs.

Google’s phased release and participation in pre-release government access also reflect the industry’s broader move toward responsible deployment. With cybersecurity defense as a stated target, the model’s guardrails will be closely watched, as will its potential for dual-use applications.

For developers, enterprises, and security teams, Argon’s introductory pricing makes it accessible for experimentation. The 95% cached input discount is particularly relevant for agentic workflows that repeatedly reference large codebases or knowledge bases. As the Gemini 4 generation unfolds, Argon sets a high bar for what the rest of the lineup will need to deliver.

via MarkTechPost

Related