PeerLM logoPeerLM
All Comparisons

OpenAI: gpt-oss-120b vs Qwen: Qwen3.5 397B A17B: Coding Performance with 10 Evaluators

This analysis compares OpenAI: gpt-oss-120b and Qwen: Qwen3.5 397B A17B on their Coding Performance with 10 Evaluators, highlighting differences in efficiency and precision.

OpenAI: gpt-oss-120b

5.1

preference score

vs

Qwen: Qwen3.5 397B A17B

4.9

preference score

Judges ranked the responses in this Run against each other; the rank is mapped onto a 0–10 scale. It shows which response was preferred, not how good either one is — and it is not a percentage, a pass rate, or a check that the output was correct.

Sample size for this comparison was not recorded. Treat it as directional.

Evidence clarification: this article predates recorded sample provenance. Treat its conclusions as claims about the displayed examples; they do not establish general model superiority, verified correctness, or production suitability.

Key Findings

Overall PerformanceOpenAI: gpt-oss-120b

Achieved a higher overall score of 5.13 compared to 4.87.

Cost EfficiencyOpenAI: gpt-oss-120b

Significantly lower total cost at $0.00036 vs $0.025549.

Instruction AdherenceOpenAI: gpt-oss-120b

Scored consistently better in instruction following across the evaluation suite.

Specifications

SpecOpenAI: gpt-oss-120bQwen: Qwen3.5 397B A17B
Provideropenaiqwen
Context Length131K262K
Input Price (per 1M tokens)$0.04$0.55
Output Price (per 1M tokens)$0.17$3.50
Parameters70b+70b+
Max Output Tokens117,964235,929
Tierstandardadvanced

Our Verdict

OpenAI: gpt-oss-120b emerges as the top-performing model in this coding evaluation, combining higher accuracy scores with vastly superior cost-efficiency. While Qwen: Qwen3.5 397B A17B remains a capable engine for complex tasks, its higher resource consumption makes it less competitive for standard coding benchmarks. For most development environments, the OpenAI model provides the best balance of quality and value.

Overview

In the rapidly evolving landscape of large language models, choosing the right architecture for software development tasks is critical. This comparative analysis examines OpenAI: gpt-oss-120b vs Qwen: Qwen3.5 397B A17B through the lens of PeerLM's rigorous Coding Performance with 10 Evaluators suite. By utilizing a comparative ranking methodology, we shed light on how these models handle complex coding instructions and logical accuracy.

Benchmark Results

The evaluation was conducted using a standardized set of prompts designed to test real-world coding scenarios. Below is the summary of the performance metrics observed during the benchmarking process.

ModelOverall ScoreAccuracyInstruction Following
OpenAI: gpt-oss-120b5.135.135.13
Qwen: Qwen3.5 397B A17B4.874.874.87

Criteria Breakdown

The evaluation focused on two primary pillars: Accuracy and Instruction Following. In coding tasks, these criteria are non-negotiable; a model must not only generate syntactically correct code but also adhere strictly to the specific architectural constraints provided by the user.

  • Accuracy: OpenAI: gpt-oss-120b demonstrated a slight edge in maintaining logical consistency within complex code snippets.
  • Instruction Following: Both models were highly effective at interpreting complex requirements, though the scoring spread of 0.26 indicates a discernible preference from our 10 human-in-the-loop evaluators for the output style of the top-ranked model.

Cost & Latency Analysis

For engineering teams, cost-efficiency is as vital as code quality. The following table highlights the economic footprint of each model based on the processed token volume during the evaluation.

ModelTotal Cost (USD)Avg Completion TokensCost per Output Token
OpenAI: gpt-oss-120b$0.00036414$0.000218
Qwen: Qwen3.5 397B A17B$0.0255492691$0.002374

As shown, OpenAI: gpt-oss-120b offers a significantly lower cost profile, making it a highly attractive option for high-volume coding tasks where budget optimization is a priority.

Use Cases

OpenAI: gpt-oss-120b excels in scenarios requiring rapid iteration, bug fixing, and boilerplate generation where cost-per-token is a primary concern. Its performance in this benchmark suggests it is well-suited for integration into CI/CD pipelines or IDE extensions.

Qwen: Qwen3.5 397B A17B, while showing higher costs, is designed for heavy-duty reasoning tasks. Given its higher completion token volume, it may be better suited for complex refactoring projects or architectural design discussions where deep context maintenance is required over longer outputs.

Verdict

The comparison between OpenAI: gpt-oss-120b vs Qwen: Qwen3.5 397B A17B reveals a clear winner in terms of immediate value and evaluation score. While both models perform admirably in coding tasks, the efficiency of the OpenAI model makes it the superior choice for most standardized coding workflows.

Backed by real data

View the Full Evaluation Report

See every response, score, and evaluator judgment behind this comparison. All data from PeerLM's blind evaluation pipeline.

View Report

Run your own Monitor

Compare OpenAI: gpt-oss-120b and Qwen: Qwen3.5 397B A17B on sampled production prompts, with frozen criteria and inspectable evidence.

Start a Monitor

Get a free managed report

We'll run a full evaluation with your real prompts and deliver a detailed recommendation. Free for qualified teams.

Request Report

Methodology

Evaluated using PeerLM's blind evaluation pipeline with 4 responses per model across 2 criteria.