DEV Community

#benchmark

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
Can You Beat an LLM? Building Humans vs. Humanity's Last Exam

Can You Beat an LLM? Building Humans vs. Humanity's Last Exam

8
Comments
4 min read
One RTX 5090 vs a 12-GPU Cluster — Benchmarking a Decade of GPUs on the Same Go Proof

One RTX 5090 vs a 12-GPU Cluster — Benchmarking a Decade of GPUs on the Same Go Proof

Comments
4 min read
SDABench: A New Benchmark for Evaluating LLMs in Scientific Discovery

SDABench: A New Benchmark for Evaluating LLMs in Scientific Discovery

Comments
4 min read
Model Showdown Round 9: Qwen 3.6 27B vs Qwen 3.6 35B-A3B vs Qwythos-9B vs GLM-4.7-Flash vs Nemotron-3-Nano

Model Showdown Round 9: Qwen 3.6 27B vs Qwen 3.6 35B-A3B vs Qwythos-9B vs GLM-4.7-Flash vs Nemotron-3-Nano

Comments
14 min read
DeepSeek vs GLM vs Qwen: Which Free LLM API is Best for Your Project?

DeepSeek vs GLM vs Qwen: Which Free LLM API is Best for Your Project?

Comments
4 min read
AdvancedMathBench: A New Benchmark for LLM Advanced Mathematical Reasoning

AdvancedMathBench: A New Benchmark for LLM Advanced Mathematical Reasoning

Comments
3 min read
TurboQuant, Four Months Later: Chasing Google's 6x VRAM Claim Into the Wild

TurboQuant, Four Months Later: Chasing Google's 6x VRAM Claim Into the Wild

Comments
6 min read
Your agent's memory remembers what you chose. Does it remember what you rejected?

Your agent's memory remembers what you chose. Does it remember what you rejected?

3
Comments
5 min read
Which LLM should I actually code with? I built a small benchmark to find out

Which LLM should I actually code with? I built a small benchmark to find out

Comments
2 min read
I Benchmarked 42 Compression Formats Spanning Four Decades. Here's What to Actually Use.

I Benchmarked 42 Compression Formats Spanning Four Decades. Here's What to Actually Use.

Comments
5 min read
ComfyUI, Lemonade, and LocalAI: Scouting the Next Wave of Homelab AI Tools

ComfyUI, Lemonade, and LocalAI: Scouting the Next Wave of Homelab AI Tools

Comments
7 min read
AI Coding Tools Benchmark 2026: Cursor vs Copilot vs Windsurf vs Claude Code

AI Coding Tools Benchmark 2026: Cursor vs Copilot vs Windsurf vs Claude Code

1
Comments
5 min read
The Same RTX 5090, but the GPU Sat Idle — a CPU-Bound Go Solver and the Case for L2 Cache

The Same RTX 5090, but the GPU Sat Idle — a CPU-Bound Go Solver and the Case for L2 Cache

Comments
6 min read
I built a neutral benchmarking layer for quantum simulators in Rust — and it revealed a silent disagreement between two backends

I built a neutral benchmarking layer for quantum simulators in Rust — and it revealed a silent disagreement between two backends

Comments
1 min read
Debugging Deployments with Gemma 12B, TPU v6e-4, MCP, and Antigravity CLI

Debugging Deployments with Gemma 12B, TPU v6e-4, MCP, and Antigravity CLI

5
Comments
16 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.