Gemini 4 Argon Beats GPT-6 Astra at 40% the Cost
Google has unveiled Gemini 4 Argon, its most powerful AI model yet, meant for coding, research, enterprise work, and cybersecurity.
Google has unveiled Gemini 4 Argon, its most powerful AI model yet, meant for coding, research, enterprise work, and cybersecurity.
Article outline
- What happened
- The key numbers
- Why it matters
- Official response
- The bottom line
Key points
- Vals notes Argon's overall score exceeds Claude Sonnet 5.5 (67.04%), Claude Opus 5.5 (66.97%), and GPT-6 Astra (63.13%).
- Vals at present ranks Gemini 4 Argon first out of 41 models on its overall Vals Index with a score of 68.90%.
- Google notes Argon assisted migrate sizeable C/C++ codebases to Rust, including projects exceeding 800, 000 lines.
- The model supports up to a 1 million-token output limit, significantly higher than the 64, 000-token limit cited for previous Gemini models.
- Add ProPakistani to Preferred Sources and see more of our stories in Google Search and Top Stories.
Gemini 4 Argon delivers GPT-6 Astra-level performance and even beats it in some of the benchmarks, at only 40% of the cost.
Google is initially rolling out Argon to selected cyber defenders through its Fairwind Program, rather than making it generally available. Google notes it is gathering feedback and strengthening safeguards before expanding access to developers, enterprises, and consumers. GPT 6.1 Sol Brings Astra-Level Performance at 5x Lower Cost.
Google notes Argon will initially cost $2 per million input tokens and $10 per million output tokens when broader paid access begins. After the introductory period, pricing will rise to $4 and $20, respectively. Coding and Long-Running Tasks.
Google notes Gemini 4 Argon is built for complex, long-horizon workflows and is already being employed internally for debugging, research, algorithm design and large-scale code migrations.
Meanwhile, the firm states Argon scored 77.9% on DeepSWE v1.1, a benchmark focused on real-world software engineering tasks, and 51.3% on AutomationBench, where Google notes it ranked first. It additionally scored 91.7% on LVBench. It evaluates long-video understanding.
Google notes Argon assisted migrate sizeable C/C++ codebases to Rust, including projects exceeding 800, 000 lines. In another internal example, it supported optimize a Rust video decoder to run 2.7x faster than an earlier Rust implementation.
Notably, the model supports up to a 1 million-token output limit, significantly higher than the 64, 000-token limit cited for previous Gemini models.
Claude Sonnet 5.5 is Faster, Smarter and Up to 30% Cheaper.
Vals notes Argon's overall score exceeds Claude Sonnet 5.5 (67.04%), Claude Opus 5.5 (66.97%), and GPT-6 Astra (63.13%). Nevertheless, it does not lead in every individual test.
Google trained Argon specifically for defensive cybersecurity and notes the model can autonomously find, validate, and patch software vulnerabilities.
On CWE-bench v1, which measures vulnerability remediation, Argon scored 68% and tied for first place. Google additionally states Wiz applied the model to identify a critical vulnerability affecting healthcare software that earlier frontier models had missed.
Fairwind gives selected governments, Google Cloud customers and cybersecurity partners access to advanced defensive AI tools. Stay Connected with ProPakistani.
Obtain the latest tech news, telecom insights, and product launches wherever you prefer. Follow on Google Discover.
Add as a preferredSource on Google Follow on Google News Join WhatsApp.
Technology and Automotive Specialist covering the latest cars, smartphones, AI breakthroughs, and.
For now, gemini 4 Argon Beats GPT remains the part of the story worth watching, and further updates are likely as more details are confirmed.




