LLM Comparisons — Page 8
MoonshotAI: Kimi K2.5 vs xAI: Grok 4: Coding Performance with 10 Evaluators
We evaluate the coding capabilities of MoonshotAI: Kimi K2.5 and xAI: Grok 4 using 10 specialized evaluators to determine the superior model for software development tasks.
MoonshotAI: Kimi K2.5
7.0
xAI: Grok 4
3.0
MoonshotAI: Kimi K2.5 vs Anthropic: Claude Sonnet 4.6: Coding Performance with 10 Evaluators
We evaluate MoonshotAI: Kimi K2.5 vs Anthropic: Claude Sonnet 4.6 on Coding Performance with 10 Evaluators to determine the superior model for developers.
MoonshotAI: Kimi K2.5
4.5
Anthropic: Claude Sonnet 4.6
5.5
Z.ai: GLM 5 vs xAI: Grok 4: Coding Performance with 10 Evaluators
In our latest benchmark for Coding Performance with 10 Evaluators, Z.ai: GLM 5 outperforms xAI: Grok 4 in both accuracy and cost-efficiency.
Z.ai: GLM 5
6.3
xAI: Grok 4
3.7
Z.ai: GLM 5 vs Google: Gemini 3.1 Pro Preview: Coding Performance with 10 Evaluators
We compare Z.ai: GLM 5 and Google: Gemini 3.1 Pro Preview on their Coding Performance with 10 Evaluators, analyzing accuracy and instruction following.
Z.ai: GLM 5
5.1
Google: Gemini 3.1 Pro Preview
4.9
Z.ai: GLM 5 vs Anthropic: Claude Sonnet 4.6: Coding Performance with 10 Evaluators
This analysis compares Z.ai: GLM 5 vs Anthropic: Claude Sonnet 4.6, evaluating their Coding Performance with 10 Evaluators to determine the superior model for software development tasks.
Z.ai: GLM 5
2.6
Anthropic: Claude Sonnet 4.6
7.4
Qwen: Qwen3.5 397B A17B vs Google: Gemini 3.1 Pro Preview: Coding Performance with 10 Evaluators
We evaluated Qwen: Qwen3.5 397B A17B vs Google: Gemini 3.1 Pro Preview using 10 expert evaluators to determine the superior model for coding tasks.
Qwen: Qwen3.5 397B A17B
4.3
Google: Gemini 3.1 Pro Preview
5.7
Qwen: Qwen3.5 397B A17B vs xAI: Grok 4: Coding Performance with 10 Evaluators
We compare Qwen: Qwen3.5 397B A17B vs xAI: Grok 4 using our Coding Performance with 10 Evaluators benchmark to determine the superior model for software development tasks.
Qwen: Qwen3.5 397B A17B
6.0
xAI: Grok 4
4.0
Qwen: Qwen3.5 397B A17B vs Anthropic: Claude Sonnet 4.6: Coding Performance with 10 Evaluators
In our latest evaluation of Coding Performance with 10 Evaluators, we compare the Qwen: Qwen3.5 397B A17B vs Anthropic: Claude Sonnet 4.6 to see which model dominates in real-world programming tasks.
Qwen: Qwen3.5 397B A17B
3.0
Anthropic: Claude Sonnet 4.6
7.0
DeepSeek: DeepSeek V3.2 vs Google: Gemini 3.1 Pro Preview: Coding Performance with 10 Evaluators
We put DeepSeek: DeepSeek V3.2 and Google: Gemini 3.1 Pro Preview through a rigorous Coding Performance with 10 Evaluators test to see which model reigns supreme in technical tasks.
DeepSeek: DeepSeek V3.2
1.8
Google: Gemini 3.1 Pro Preview
8.2
DeepSeek: DeepSeek V3.2 vs Anthropic: Claude Sonnet 4.6: Coding Performance with 10 Evaluators
We evaluated DeepSeek: DeepSeek V3.2 vs Anthropic: Claude Sonnet 4.6 on Coding Performance with 10 Evaluators to determine the superior model for complex programming tasks.
DeepSeek: DeepSeek V3.2
3.2
Anthropic: Claude Sonnet 4.6
6.8
Meta: Llama 4 Maverick vs OpenAI: GPT-5.4 Mini: Coding Performance with 10 Evaluators
In our latest Coding Performance with 10 Evaluators benchmark, we compare the capabilities of Meta: Llama 4 Maverick against OpenAI: GPT-5.4 Mini.
Meta: Llama 4 Maverick
0.5
OpenAI: GPT-5.4 Mini
9.5
Qwen: Qwen3 32B vs Mistral: Mistral Small 3.2 24B: Coding Performance with 10 Evaluators
This analysis compares Qwen: Qwen3 32B vs Mistral: Mistral Small 3.2 24B, focusing on Coding Performance with 10 Evaluators to determine the superior model for development tasks.
Qwen: Qwen3 32B
3.4
Mistral: Mistral Small 3.2 24B
6.6