DEV Community

Ethan Walker profile picture

Ethan Walker

404 bio not found

Joined Joined on 
Our CI eval gate sent us a token bill. The deterministic one sent nothing.

Our CI eval gate sent us a token bill. The deterministic one sent nothing.

Comments
5 min read

Want to connect with Ethan Walker?

Create an account to connect with Ethan Walker. You can also sign in below to proceed if you already have an account.

Already have an account? Sign in
Our few-shot examples came from the eval set. The 0.94 was fiction.

Our few-shot examples came from the eval set. The 0.94 was fiction.

Comments 2
13 min read
Our few-shot examples came from the eval set. The 0.94 was fiction.

Our few-shot examples came from the eval set. The 0.94 was fiction.

2
Comments 8
13 min read
We gated CI on six open-source LLM eval frameworks. Only two survived the merge queue.

We gated CI on six open-source LLM eval frameworks. Only two survived the merge queue.

3
Comments 4
14 min read
The golden set stopped catching regressions the day traffic changed

The golden set stopped catching regressions the day traffic changed

2
Comments 4
6 min read
Your LLM-as-judge disagrees with itself between runs

Your LLM-as-judge disagrees with itself between runs

Comments
4 min read
LLM-as-judge disagrees with itself between runs

LLM-as-judge disagrees with itself between runs

4
Comments 1
4 min read
When an LLM answer is wrong, the trace is where you look. Some tools make that easy.

When an LLM answer is wrong, the trace is where you look. Some tools make that easy.

Comments
5 min read
# A 94% pass rate hid a PII leak in 6 test cases

# A 94% pass rate hid a PII leak in 6 test cases

Comments
5 min read
our CI passed. Your agent isn't operator-ready.

our CI passed. Your agent isn't operator-ready.

Comments 1
4 min read
The stale eval fixture that passed a broken model

The stale eval fixture that passed a broken model

Comments
4 min read
My eval handed me a 0.62 and no idea why. The fix was not a better eval.

My eval handed me a 0.62 and no idea why. The fix was not a better eval.

Comments
7 min read
91% pass rate. Gate green. Shipped. Worst regression we had all quarter.

91% pass rate. Gate green. Shipped. Worst regression we had all quarter.

1
Comments 1
2 min read
We stopped writing eval cases by hand. Now every prod incident becomes one.

We stopped writing eval cases by hand. Now every prod incident becomes one.

Comments
2 min read
Your eval criteria are code. Version them like code.

Your eval criteria are code. Version them like code.

Comments
3 min read
Datadog dashboards for prompt regression: the panels we actually keep

Datadog dashboards for prompt regression: the panels we actually keep

Comments
8 min read
Switching our LLM-as-judge from 5-class to binary in CI: the patterns we kept

Switching our LLM-as-judge from 5-class to binary in CI: the patterns we kept

Comments
3 min read
Promptfoo is a CI gate, not an eval framework. Treating it like one cost us $4,200

Promptfoo is a CI gate, not an eval framework. Treating it like one cost us $4,200

Comments 1
4 min read
loading...