DEV Community

#evaluation

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
The benchmark that disproved its own result

The benchmark that disproved its own result

Comments
6 min read
El modelo encontró evidencia relevante y aun así falló: por qué “alucinación” se me quedó corta

El modelo encontró evidencia relevante y aun así falló: por qué “alucinación” se me quedó corta

Comments
4 min read
The sleep loop is the tell: agents that pay per action optimize to do nothing

The sleep loop is the tell: agents that pay per action optimize to do nothing

Comments 1
2 min read
The model did the reverse-engineering. The validator was the hard part.

The model did the reverse-engineering. The validator was the hard part.

Comments 1
2 min read
Judging AI hackathon projects: what to check when every team says 'we used AI'

Judging AI hackathon projects: what to check when every team says 'we used AI'

Comments
3 min read
How to evaluate a RAG system: recall, faithfulness and the questions that matter

How to evaluate a RAG system: recall, faithfulness and the questions that matter

Comments
3 min read
How to Evaluate AI Agents

How to Evaluate AI Agents

2
Comments
7 min read
JuryTrace: make agent-judge failures inspectable

JuryTrace: make agent-judge failures inspectable

Comments 1
4 min read
My board never scored an outage as a regression. My evidence couldn't prove it.

My board never scored an outage as a regression. My evidence couldn't prove it.

Comments
5 min read
We spent two days bisecting a prompt change. The regression was noise.

We spent two days bisecting a prompt change. The regression was noise.

Comments
1 min read
Production-Ready Multi-Turn Evaluation

Production-Ready Multi-Turn Evaluation

Comments
7 min read
A Free Server Is Enough to Test a New Model Before You Trust It

A Free Server Is Enough to Test a New Model Before You Trust It

Comments
3 min read
Writing the code is no longer the bottleneck

Writing the code is no longer the bottleneck

Comments
3 min read
Why AI Benchmarks Mean Less Than You Think

Why AI Benchmarks Mean Less Than You Think

Comments
6 min read
One Quality Score Is a Lie: Split Your RAG Judge Into Retrieval, Groundedness, and Relevance

One Quality Score Is a Lie: Split Your RAG Judge Into Retrieval, Groundedness, and Relevance

1
Comments 1
5 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.