Todos los artículos

Gemini 3.1 Pro: Benchmark Leader at Half the Cost — But Read the Fine Print

23 de febrero de 2026

Gemini 3.1 Pro: Benchmark Leader at Half the Cost — But Read the Fine Print

Google released Gemini 3.1 Pro in preview on February 19th. It now leads the Artificial Analysis Intelligence Index at 57 points — 4 ahead of Claude Opus 4.6 (53), 6 ahead of GPT-5.2 (51). That's across 116 evaluated models, with gains in agent-based coding, knowledge, scientific reasoning, and physics. The hallucination rate dropped 38 percentage points from 3 Pro. And the ARC-AGI-2 score hit 77.1% — more than double the reasoning performance of its predecessor.

The cost story is even more dramatic. Running the full Intelligence Index benchmark suite costs 892withGemini3.1Pro.Thesameruncosts892 with Gemini 3.1 Pro. The same run costs2,486 with Claude Opus 4.6 and 2,304withGPT−5.2.At2,304 with GPT-5.2. At2 per million input tokens and 12permillionoutputtokens,GeminiundercutsOpus4.6(12 per million output tokens, Gemini undercuts Opus 4.6 (15/$75) by 7.5x on input and 6.25x on output. It did this while using just 57 million tokens — roughly half what GPT-5.2 requires for the same evaluation.

Google is also pushing 3.1 Pro into Antigravity (their agentic dev platform), Gemini CLI, and Android Studio. This is a clear signal they're going after developer workflows, not just benchmark bragging rights.

The Fine Print

Benchmarks measure capability in controlled conditions. They don't measure how well a model performs when you hand it a real codebase and say "go."

In real-world agent tasks, Gemini 3.1 Pro still falls behind Claude Sonnet 4.6, Opus 4.6, and GPT-5.2. In fact-checking tests, it verified only about 25% of statements — worse than both Opus and GPT-5.2. And Claude Opus 4.6 at full Adaptive Thinking capacity still scores roughly 64 on the Intelligence Index, the highest of any model currently available.

The pattern is consistent: Gemini 3.1 Pro leads on more general benchmarks while costing dramatically less. Claude holds the edge in expert-level tasks and agentic software engineering. GPT-5.3-Codex remains the specialist for dedicated coding workflows.

What This Means for Engineering Teams

If you're evaluating AI models for enterprise use, this release changes the cost conversation. Google is delivering benchmark-competitive intelligence at a fraction of the price. For teams running high-volume, general-purpose workloads — summarization, classification, knowledge retrieval — Gemini 3.1 Pro deserves serious consideration.

But if your use case is agentic development — autonomous coding sessions, multi-step engineering workflows, system-level reasoning — benchmark parity doesn't translate to task parity. The models that perform best in controlled evaluations aren't always the ones that perform best when you give them real autonomy on real codebases.

Watch Gemini 3.1 Pro closely. The cost pressure alone will force Anthropic and OpenAI to respond. But don't switch your agentic workflows based on benchmark scores.