PeerLM logoPeerLM
All Comparisons

Anthropic: Claude Sonnet 4.6 vs Google: Gemini 3.1 Pro Preview: Coding Performance with 10 Evaluators

This comparative analysis evaluates Anthropic: Claude Sonnet 4.6 vs Google: Gemini 3.1 Pro Preview, focusing on their respective Coding Performance with 10 Evaluators.

Anthropic: Claude Sonnet 4.6

6.6

preference score

vs

Google: Gemini 3.1 Pro Preview

3.4

preference score

Judges ranked the responses in this Run against each other; the rank is mapped onto a 0–10 scale. It shows which response was preferred, not how good either one is — and it is not a percentage, a pass rate, or a check that the output was correct.

Sample size for this comparison was not recorded. Treat it as directional.

Evidence clarification: this article predates recorded sample provenance. Treat its conclusions as claims about the displayed examples; they do not establish general model superiority, verified correctness, or production suitability.

Key Findings

Coding AccuracyAnthropic: Claude Sonnet 4.6

Claude Sonnet 4.6 achieved a significantly higher accuracy score of 6.58.

Instruction FollowingAnthropic: Claude Sonnet 4.6

Claude demonstrated superior adherence to complex technical instructions.

Cost EfficiencyAnthropic: Claude Sonnet 4.6

Claude Sonnet 4.6 proved more cost-effective per request in this evaluation subset.

Specifications

SpecAnthropic: Claude Sonnet 4.6Google: Gemini 3.1 Pro Preview
Provideranthropicgoogle
Context Length1.0M1.0M
Input Price (per 1M tokens)$3.00$2.00
Output Price (per 1M tokens)$15.00$12.00
Max Output Tokens128,00065,536
Tierfrontierpremium

Our Verdict

Anthropic: Claude Sonnet 4.6 is the clear winner for coding-heavy tasks, outperforming Google: Gemini 3.1 Pro Preview in both accuracy and instruction following. While Gemini offers greater verbosity per response, Claude's higher score and lower total cost make it the superior choice for high-stakes software engineering workflows.

Overview

In the rapidly evolving landscape of Large Language Models, selecting the right architecture for software development tasks is critical. This report provides a detailed comparative analysis of Anthropic: Claude Sonnet 4.6 vs Google: Gemini 3.1 Pro Preview. By utilizing PeerLM’s rigorous evaluation framework, we assessed both models across a series of complex coding challenges to determine which handles technical instructions and logic with higher reliability.

Benchmark Results

The evaluation focused on two primary pillars: Accuracy and Instruction Following. Using a comparative ranking methodology, we observed a distinct performance gap between the two models when subjected to 10 independent evaluators.

ModelOverall ScoreAccuracyInstruction Following
Anthropic: Claude Sonnet 4.66.586.586.58
Google: Gemini 3.1 Pro Preview3.423.423.42

Criteria Breakdown

The evaluation was centered on how well each model translates natural language prompts into executable, high-quality code. Anthropic: Claude Sonnet 4.6 emerged as the leader, demonstrating a superior ability to adhere to constraints and produce accurate logical structures. Google: Gemini 3.1 Pro Preview, while capable, showed a wider variance in its output, which resulted in lower consistency across the 10-evaluator test suite.

Cost & Latency

Efficiency is as vital as accuracy in production environments. Below is the cost breakdown based on the evaluated response set.

  • Anthropic: Claude Sonnet 4.6: Total cost of $0.014196 for 4 responses, with an average output length of 189 tokens.
  • Google: Gemini 3.1 Pro Preview: Total cost of $0.079106 for 4 responses, with an average output length of 1612 tokens.

Interestingly, while Gemini 3.1 Pro Preview produced significantly more completion tokens per request, the higher total cost suggests a different pricing profile per token compared to Claude Sonnet 4.6.

Use Cases

Anthropic: Claude Sonnet 4.6 is recommended for mission-critical coding tasks, architectural planning, and debugging where high instruction fidelity is required. Its performance consistency makes it a reliable choice for automated code generation pipelines.

Google: Gemini 3.1 Pro Preview may be better suited for exploratory tasks or scenarios where longer-form output is beneficial, provided the application can accommodate the higher cost and lower strict-instruction adherence observed in this benchmark.

Verdict

The comparison of Anthropic: Claude Sonnet 4.6 vs Google: Gemini 3.1 Pro Preview highlights a clear performance advantage for Claude in technical contexts. With a score spread of 3.16, Claude Sonnet 4.6 outperformed Gemini 3.1 Pro Preview in both accuracy and instruction adherence, establishing itself as the more capable model for rigorous coding requirements.

Backed by real data

View the Full Evaluation Report

See every response, score, and evaluator judgment behind this comparison. All data from PeerLM's blind evaluation pipeline.

View Report

Run your own Monitor

Compare Anthropic: Claude Sonnet 4.6 and Google: Gemini 3.1 Pro Preview on sampled production prompts, with frozen criteria and inspectable evidence.

Start a Monitor

Get a free managed report

We'll run a full evaluation with your real prompts and deliver a detailed recommendation. Free for qualified teams.

Request Report

Methodology

Evaluated using PeerLM's blind evaluation pipeline with 4 responses per model across 2 criteria.