PeerLM logoPeerLM
All Comparisons

Google: Gemini 3.1 Pro Preview vs Qwen: Qwen3.5 397B A17B: Coding Performance with 10 Evaluators

We evaluate the coding prowess of Google: Gemini 3.1 Pro Preview vs Qwen: Qwen3.5 397B A17B using 10 expert evaluators to determine the superior model for software development tasks.

Google: Gemini 3.1 Pro Preview

6.2

preference score

vs

Qwen: Qwen3.5 397B A17B

3.9

preference score

Judges ranked the responses in this Run against each other; the rank is mapped onto a 0–10 scale. It shows which response was preferred, not how good either one is — and it is not a percentage, a pass rate, or a check that the output was correct.

Sample size for this comparison was not recorded. Treat it as directional.

Evidence clarification: this article predates recorded sample provenance. Treat its conclusions as claims about the displayed examples; they do not establish general model superiority, verified correctness, or production suitability.

Key Findings

Coding AccuracyGoogle: Gemini 3.1 Pro Preview

Gemini secured a higher overall score, proving more reliable for complex logic.

Cost-EfficiencyQwen: Qwen3.5 397B A17B

Qwen is significantly more affordable, ideal for high-volume boilerplate tasks.

Instruction FollowingGoogle: Gemini 3.1 Pro Preview

Gemini demonstrated better adherence to strict coding constraints.

Specifications

SpecGoogle: Gemini 3.1 Pro PreviewQwen: Qwen3.5 397B A17B
Providergoogleqwen
Context Length1.0M262K
Input Price (per 1M tokens)$2.00$0.55
Output Price (per 1M tokens)$12.00$3.50
Max Output Tokens65,536235,929
Tierpremiumadvanced

Our Verdict

Google: Gemini 3.1 Pro Preview is the superior choice for high-stakes coding tasks, offering better accuracy and instruction following. However, Qwen: Qwen3.5 397B A17B remains a strong contender for cost-sensitive projects requiring high-volume code generation.

Overview

In the rapidly evolving landscape of large language models, selecting the right tool for coding tasks is critical for developer productivity. This report provides a detailed comparative analysis of Google: Gemini 3.1 Pro Preview vs Qwen: Qwen3.5 397B A17B, specifically focusing on their Coding Performance with 10 Evaluators. By utilizing a comparative ranking methodology, we highlight how these models handle complex programming challenges, instruction adherence, and overall output accuracy.

Benchmark Results

Our evaluation suite utilized 10 independent evaluators to rank the performance of these models across real-world coding scenarios. The results demonstrate a clear hierarchy in model intelligence for technical tasks.

ModelOverall ScoreAccuracyInstruction Following
Google: Gemini 3.1 Pro Preview6.156.156.15
Qwen: Qwen3.5 397B A17B3.853.853.85

Criteria Breakdown

The evaluation focused on two primary pillars of coding success: Accuracy and Instruction Following.

Accuracy

Google: Gemini 3.1 Pro Preview demonstrated superior precision in generating syntactically correct and logically sound code blocks. It showed a higher capability for debugging and identifying edge cases compared to the Qwen variant, which struggled with more complex architectural prompts.

Instruction Following

Instruction following is paramount when integrating LLMs into IDEs or automated pipelines. Gemini 3.1 Pro Preview consistently adhered to specific formatting constraints and architectural requirements provided by our evaluators, whereas Qwen 3.5 397B showed occasional drift in complex multi-step instructions.

Cost & Latency

While performance is the primary metric, operational costs are a significant factor for scaling development tools. Below is the cost breakdown based on the tokens processed during our evaluation.

ModelCost per Output TokenTotal Cost (USD)
Google: Gemini 3.1 Pro Preview$0.01227$0.079106
Qwen: Qwen3.5 397B A17B$0.002374$0.025549

As shown, Qwen: Qwen3.5 397B A17B offers a significantly lower cost per token, making it an attractive option for high-volume, lower-complexity tasks, whereas Gemini 3.1 Pro Preview justifies its higher cost through superior coding accuracy.

Use Cases

Google: Gemini 3.1 Pro Preview is best suited for complex software engineering tasks, architectural planning, and debugging legacy codebases where high accuracy requirements outweigh cost constraints. Its ability to navigate nuanced coding instructions makes it a powerful partner for senior developers.

Qwen: Qwen3.5 397B A17B excels in high-throughput environments such as automated unit test generation, boilerplate code creation, and rapid prototyping, where cost-efficiency is the primary business driver.

Verdict

When comparing Google: Gemini 3.1 Pro Preview vs Qwen: Qwen3.5 397B A17B, the Gemini model emerges as the clear leader in coding performance. While the Qwen model provides a substantial cost advantage, the consistency and accuracy delivered by Gemini 3.1 Pro Preview provide a more reliable experience for mission-critical development workflows.

Backed by real data

View the Full Evaluation Report

See every response, score, and evaluator judgment behind this comparison. All data from PeerLM's blind evaluation pipeline.

View Report

Run your own Monitor

Compare Google: Gemini 3.1 Pro Preview and Qwen: Qwen3.5 397B A17B on sampled production prompts, with frozen criteria and inspectable evidence.

Start a Monitor

Get a free managed report

We'll run a full evaluation with your real prompts and deliver a detailed recommendation. Free for qualified teams.

Request Report

Methodology

Evaluated using PeerLM's blind evaluation pipeline with 4 responses per model across 2 criteria.