PeerLM logoPeerLM
All Comparisons

Qwen: Qwen3.5 397B A17B vs Google: Gemini 3.1 Pro Preview: Coding Performance with 10 Evaluators

We evaluated Qwen: Qwen3.5 397B A17B vs Google: Gemini 3.1 Pro Preview using 10 expert evaluators to determine the superior model for coding tasks.

Qwen: Qwen3.5 397B A17B

4.3

preference score

vs

Google: Gemini 3.1 Pro Preview

5.7

preference score

Judges ranked the responses in this Run against each other; the rank is mapped onto a 0–10 scale. It shows which response was preferred, not how good either one is — and it is not a percentage, a pass rate, or a check that the output was correct.

Sample size for this comparison was not recorded. Treat it as directional.

Evidence clarification: this article predates recorded sample provenance. Treat its conclusions as claims about the displayed examples; they do not establish general model superiority, verified correctness, or production suitability.

Key Findings

Top PerformerGoogle: Gemini 3.1 Pro Preview

Ranked #1 with an overall score of 5.68 in coding tasks.

Best ValueQwen: Qwen3.5 397B A17B

Offers a significantly lower cost per output token at $0.00237.

Instructional AccuracyGoogle: Gemini 3.1 Pro Preview

Excelled in following complex coding constraints better than Qwen.

Specifications

SpecQwen: Qwen3.5 397B A17BGoogle: Gemini 3.1 Pro Preview
Providerqwengoogle
Context Length262K1.0M
Input Price (per 1M tokens)$0.55$2.00
Output Price (per 1M tokens)$3.50$12.00
Max Output Tokens235,92965,536
Tieradvancedpremium

Our Verdict

Google: Gemini 3.1 Pro Preview is the clear winner for coding performance, providing higher accuracy and better instruction following. However, Qwen: Qwen3.5 397B A17B remains a strong, cost-effective alternative for high-volume development tasks.

Overview

In the rapidly evolving landscape of large language models, selecting the right architecture for software engineering tasks is critical. This comparative analysis examines Qwen: Qwen3.5 397B A17B vs Google: Gemini 3.1 Pro Preview, specifically focusing on their Coding Performance with 10 Evaluators. By utilizing a comparative ranking methodology, we provide a clear view of how these models perform when tasked with complex programming challenges.

Benchmark Results

Our evaluation utilized 10 expert evaluators to rank the outputs of both models across a series of coding prompts. The results highlight a clear preference for one model over the other in terms of overall code quality and instruction adherence.

ModelRankOverall ScoreAccuracyInstruction Following
Google: Gemini 3.1 Pro Preview15.685.685.68
Qwen: Qwen3.5 397B A17B24.324.324.32

Criteria Breakdown

The evaluation focused on two primary pillars: Accuracy and Instruction Following. In coding tasks, these metrics are vital for ensuring that the generated code is not only syntactically correct but also aligns with specific project constraints and logic requirements.

  • Accuracy: Gemini 3.1 Pro Preview demonstrated a higher capacity for generating functional, bug-free code compared to the Qwen variant.
  • Instruction Following: When provided with complex architectural constraints, Gemini 3.1 Pro Preview proved more adept at maintaining those guidelines throughout the response.

Cost & Latency

Understanding the economic trade-offs is essential for developers integrating these models into production pipelines. While Gemini 3.1 Pro Preview leads in performance, it comes at a higher cost-per-token compared to the Qwen: Qwen3.5 397B A17B.

ModelTotal Cost (USD)Cost per Output TokenAvg Completion Tokens
Google: Gemini 3.1 Pro Preview$0.079106$0.012271612
Qwen: Qwen3.5 397B A17B$0.025549$0.002372691

Use Cases

Google: Gemini 3.1 Pro Preview is best suited for high-stakes enterprise coding tasks where precision and adherence to complex instructions are paramount. Its superior ranking in our benchmark suggests it is the more reliable choice for automated refactoring or generating complex boilerplate code.

Qwen: Qwen3.5 397B A17B, while ranking second in this specific suite, offers significant value for developers looking for a cost-effective alternative. It is well-suited for high-volume tasks where rapid prototyping or lower-cost code generation is required, provided the task complexity remains within its operational threshold.

Verdict

The comparison of Qwen: Qwen3.5 397B A17B vs Google: Gemini 3.1 Pro Preview reveals a performance gap of 1.36 points in favor of the Google model. While Qwen remains a highly efficient and cost-effective option, Gemini 3.1 Pro Preview is the current leader for coding accuracy and instruction compliance.

Backed by real data

View the Full Evaluation Report

See every response, score, and evaluator judgment behind this comparison. All data from PeerLM's blind evaluation pipeline.

View Report

Run your own Monitor

Compare Qwen: Qwen3.5 397B A17B and Google: Gemini 3.1 Pro Preview on sampled production prompts, with frozen criteria and inspectable evidence.

Start a Monitor

Get a free managed report

We'll run a full evaluation with your real prompts and deliver a detailed recommendation. Free for qualified teams.

Request Report

Methodology

Evaluated using PeerLM's blind evaluation pipeline with 4 responses per model across 2 criteria.