Thursday, 1 October 2026 | Updating Daily AI insight, written for builders

Gemini 4 Argon vs GPT-6.1 Sol vs Claude Sonnet 5.5: Three New AI Models Compared

  • Three big AI models arrived in three days: Claude Sonnet 5.5 (Anthropic, 28 September 2026), GPT-6.1 Sol (OpenAI, 29 September) and Gemini 4 Argon (Google, 30 September).
  • All three launch at the same developer price: $2 per million input tokens and $10 per million output tokens. Argon’s price is introductory and rises to $4 / $20 later.
  • The real difference is who can use them today: Sonnet 5.5 is open to everyone including free Claude users; GPT-6.1 Sol is for paid ChatGPT plans (in ChatGPT Work and Codex) and the API; Gemini 4 Argon is not publicly available yet.
  • Each company reports different tests, so there is no fair single winner yet. Google claims the top scores; OpenAI and Anthropic stress cost per task.

The last week of September 2026 brought an unusual pile-up of launches. Within three days, Anthropic, OpenAI and Google each released a new model — and all three landed on exactly the same developer price. If you are trying to work out which one matters to you, this side-by-side comparison sticks to what each company has actually published.

The three models at a glance

Gemini 4 Argon GPT-6.1 Sol Claude Sonnet 5.5
Company Google DeepMind OpenAI Anthropic
Announced 30 September 2026 29 September 2026 (DevDay) 28 September 2026
Position Google’s new frontier model Upgraded mid-tier model, below GPT-6 Astra Mid-tier model, below Claude Opus 5.5
API price (input / output per 1M tokens) $2 / $10 introductory, then $4 / $20 $2 / $10 $2 / $10
Cached input per 1M tokens 95% off input price $0.10 $0.20 (cache reads)
Standout claim 1M-token output limit; 77.9% on DeepSWE v1.1 Near-Astra results at a fifth of Astra’s price 30%+ faster than Sonnet 5, up to 30% cheaper per task
Free consumer access Not announced No (paid ChatGPT plans) Yes, reported on Claude’s free plan
Available today? No — trusted testers only Yes Yes

Sources: Google, OpenAI and Anthropic announcements.

Who can use each model right now

For most people this is the part that matters, because a model you can’t open is not much use.

Claude Sonnet 5.5: available to everyone

Sonnet 5.5 is live in the Claude apps and on every major cloud. Claude’s Free plan includes Sonnet models, and reports say free users get Sonnet 5.5, while Opus 5.5 stays on paid plans. Full details: Claude Sonnet 5.5 explained.

GPT-6.1 Sol: paid ChatGPT plans and the API

OpenAI says GPT-6.1 Sol is available to Plus, Pro, Business, Enterprise and Edu users in ChatGPT Work and Codex, and to developers as gpt-6.1-sol. OpenAI adds that it is “not yet available in Chat”, the regular ChatGPT conversation screen. Free ChatGPT users don’t get it. Not sure which ChatGPT plan you need? See ChatGPT Free vs Go vs Plus vs Pro. For the DevDay announcements around it, read our DevDay report.

Gemini 4 Argon: not yet

Google has released Argon only to selected cybersecurity defenders through its Fairwind Program and to its own staff. It says access will widen “as soon as possible”, starting with paid API customers and Google AI Ultra subscribers, but has not given a date. Full details: Gemini 4 Argon explained.

Price: the same $2 / $10, with a catch

It is rare for three rival flagship-class launches to match on price. For developers paying per token:

  • Claude Sonnet 5.5 keeps Sonnet 5’s $2 / $10, and Anthropic says it needs fewer tokens per task, making it up to 30% cheaper in practice.
  • GPT-6.1 Sol keeps GPT-6 Sol’s $2 / $10 and halves cached input to $0.10. OpenAI’s pitch is that it nearly matches GPT-6 Astra — which costs $10 / $50 — at a fifth of the price.
  • Gemini 4 Argon starts at $2 / $10 but only for an introductory period of unannounced length, then doubles to $4 / $20 — the same as Claude Opus 5.5.

Price per token is only half the story: a model that uses fewer tokens, or gets the job done in fewer attempts, can be cheaper even at the same rate. Both OpenAI and Anthropic now publish cost-per-task figures for exactly that reason. To estimate a bill for your own usage, try the AI API cost calculator or compare all current models in the AI models database.

Benchmarks: what each company reports

Here is the honest problem with comparing these three: they mostly report different tests, often in different versions and at different “effort” settings. The table below lists each company’s own headline numbers side by side, but a higher number in one column does not prove one model beats another.

Area Gemini 4 Argon (Google) GPT-6.1 Sol (OpenAI) Claude Sonnet 5.5 (Anthropic)
Coding DeepSWE v1.1: 77.9% (state of the art) DeepSWE v1.1: matches GPT-6 Astra at ~⅕ of the cost; 6.4 points above GPT-6 Sol’s best Terminal-Bench 4.0: 70.6%; CursorBench 4.0: 55.5%
Business workflows AutomationBench: 51.3% (#1) AutomationBench: 2.2 points above Claude Opus 5.5 at medium effort, ~⅓ of the cost GDPval-AA v2.1: 1844 (Opus 5.5: 1846)
Using a computer — OSWorld 2.0 offline: within 2.1 points of Astra at ~⅐ of the cost per task OSWorld 2.1: 80.1%
Documents, charts, video LVBench (long video): 91.7% GDP.pdf: above Opus 5.5 at under half the cost per task Chartography: 61.6%
Accuracy — Responses with a factual error: 7.7% vs 11.4% for GPT-6 Sol (hard prompts, low effort) Humanity’s Last Exam (with tools): 64.5%
Security CWE-bench v1: 68% (tied first) — Launches with Opus-level cyber safeguards

A few fair observations from these numbers:

  • Coding: DeepSWE v1.1 is the one coding test both Google and OpenAI cite. Google gives Argon a firm 77.9%; OpenAI gives Sol no absolute number, only that it matches GPT-6 Astra. Until independent results appear, Argon’s claim is the strongest.
  • Business workflows: AutomationBench is the other test both Google and OpenAI cite. Google says Argon is #1 at 51.3%. OpenAI says Sol beats Claude Opus 5.5 there by 2.2 points. Anthropic doesn’t report it for Sonnet 5.5.
  • Sonnet 5.5 vs Opus 5.5: Anthropic’s numbers show its mid-range model nearly level with its own top model on office work, at half the price.

Treat every figure here as marketing until outside evaluators like Artificial Analysis or LMArena publish their own results, which usually takes days to weeks after a launch.

Safety: an unusually cautious week

All three launches came with more safety discussion than usual:

  • OpenAI released GPT-6.1 Sol instead of the larger GPT-6.1 Astra, which it held back after internal tests raised concerns about deception and acting without permission (our report). OpenAI says Sol shows better alignment than GPT-6 Sol.
  • Google is rolling Argon out in phases, joining the U.S. government’s voluntary pre-release access process and adding monitors that can stop the model mid-task.
  • Anthropic gave Sonnet 5.5 the same cyber safeguards as its top models: higher-risk security requests fall back to the older Sonnet 5.

Which one should you use?

  • You want something free, today: Claude Sonnet 5.5 is the only one of the three on a free plan. Gemini users keep the current models in the Gemini app until Argon arrives.
  • You already pay for ChatGPT: try GPT-6.1 Sol in ChatGPT Work or Codex for coding, documents and multi-step tasks.
  • You pay for Google AI Ultra: watch for Argon in the Gemini app, but there is no date. Our guide to Google AI plans explains what Ultra includes now.
  • You build apps: Sonnet 5.5 and GPT-6.1 Sol are both available at $2 / $10 today, so test both on your own tasks. Argon is worth a trial once the API opens, but plan for the price doubling after the introductory period.
  • You are choosing a chatbot, not a model: the features around the model matter as much as the model itself — see Claude vs ChatGPT.

Frequently asked questions

Which is the best AI model right now: Gemini 4 Argon, GPT-6.1 Sol or Claude Sonnet 5.5?

There is no fair single answer yet. Google reports the strongest headline scores for Argon, but Argon isn’t publicly available. Among models you can use today, GPT-6.1 Sol and Claude Sonnet 5.5 cost the same and each company reports strengths on different tests.

Do Gemini 4 Argon, GPT-6.1 Sol and Claude Sonnet 5.5 cost the same?

At launch, yes: $2 per million input tokens and $10 per million output tokens through each company’s API. Argon’s price is introductory and rises to $4 / $20 later.

Which of the three can I use for free?

Claude Sonnet 5.5, on Claude’s free plan with usage limits. GPT-6.1 Sol needs a paid ChatGPT plan, and Gemini 4 Argon has no public release yet.

When will Gemini 4 Argon be available?

Google has not given a date. It says the rollout will start with paid API customers and Google AI Ultra subscribers.

Is GPT-6.1 Sol better than GPT-6 Astra?

No. OpenAI says it nearly matches Astra on coding, computer use and professional work at a fifth of Astra’s price, but Astra still scores higher on the hardest tasks, such as Terminal-Bench Science (68.1%).

Written by Mustafa Ihsan

Mustafa Ihsan is the founder and editor of Convly.ai. He built and maintains the site's live AI models database, its price-performance index, and its free calculators for VRAM requirements, API costs and self-hosting economics. He writes about model pricing, benchmark results and the hardware needed to run AI models locally, and consistently prefers measured numbers to vendor claims.

Scroll to Top