DEV Community

Reid Marlow profile picture

Reid Marlow

A technologist working in automation. Here I write up what I actually learn building tools, engineering AI agents, and keeping systems running, along with the workflows and habits that stick.

Education

Hong Kong Polytechnic University, Automation PhD (in progress)

Pronouns

he/him

Work

PhD candidate in Automation (Hong Kong Polytechnic University)

The Agent Passed SWE-Bench Pro Because git show Still Had the Fix

The Agent Passed SWE-Bench Pro Because git show Still Had the Fix

Comments
3 min read

Want to connect with Reid Marlow?

Create an account to connect with Reid Marlow. You can also sign in below to proceed if you already have an account.

Already have an account? Sign in
CISA's Distillation Detector Also Flags a Shared API Key

CISA's Distillation Detector Also Flags a Shared API Key

Comments
3 min read
Don't Hand the Child Agent the User JWT

Don't Hand the Child Agent the User JWT

Comments
3 min read
Astra's Adapter Beat the Reasoning Dial on ARC-AGI-3

Astra's Adapter Beat the Reasoning Dial on ARC-AGI-3

Comments
3 min read
The Piano Decoder Was Already in the Chat Log

The Piano Decoder Was Already in the Chat Log

Comments
2 min read
A Model Swap Can Keep the Memory File and Still Lose the Facts

A Model Swap Can Keep the Memory File and Still Lose the Facts

Comments
4 min read
A Plugin Update Can Add Shell Hooks the Model Never Sees

A Plugin Update Can Add Shell Hooks the Model Never Sees

Comments
4 min read
A Per-Agent Cap Can Still Overdraw 48

A Per-Agent Cap Can Still Overdraw 48

Comments
3 min read
A Skill Without an Input Contract Should Stay in the Parent Agent

A Skill Without an Input Contract Should Stay in the Parent Agent

1
Comments
3 min read
On-Policy Distillation Works Better Without the Teacher

On-Policy Distillation Works Better Without the Teacher

3
Comments
3 min read
Why Coding Agents Fail in the Outer Loop

Why Coding Agents Fail in the Outer Loop

5
Comments 2
3 min read
When Training Lawsuits Target the Download Script

When Training Lawsuits Target the Download Script

2
Comments
3 min read
GLM-5.3, 756GB of Weights, and the Ten Billion Dollar Gate

GLM-5.3, 756GB of Weights, and the Ten Billion Dollar Gate

4
Comments 2
3 min read
The Agent Hack Postmortem Is Really About Shared State

The Agent Hack Postmortem Is Really About Shared State

6
Comments 10
4 min read
WeChat's embedding model is a deployment story, not a leaderboard flex

WeChat's embedding model is a deployment story, not a leaderboard flex

4
Comments
6 min read
OpenAI's SB 53 Pivot Is a Safety Incident Report in Disguise

OpenAI's SB 53 Pivot Is a Safety Incident Report in Disguise

1
Comments
6 min read
Claude Opus 4.6 Shows Why Old Models Need Patch Windows

Claude Opus 4.6 Shows Why Old Models Need Patch Windows

6
Comments
5 min read
AI Capex Is Turning Into an Infrastructure Bill

AI Capex Is Turning Into an Infrastructure Bill

4
Comments 4
5 min read
World Model Benchmarks Need Receipts, Not Just Scores

World Model Benchmarks Need Receipts, Not Just Scores

4
Comments
5 min read
The Official DeepSWE Number Is Footnoted to DeepSeek Harness

The Official DeepSWE Number Is Footnoted to DeepSeek Harness

3
Comments
1 min read
The Important Part of Anthropic's Risk Report Is the Benchmark That Stopped Moving

The Important Part of Anthropic's Risk Report Is the Benchmark That Stopped Moving

3
Comments
4 min read
Gemini 3.7 Flash Makes Agent Cost the Feature

Agent loops matter more than raw token cost

Gemini 3.7 Flash Makes Agent Cost the Feature

4
Comments 9
4 min read
Devin's $40B Round Is a Bet on Agent Budgets, Not Better Demos

Devin's $40B Round Is a Bet on Agent Budgets, Not Better Demos

2
Comments
3 min read
Nvidia's Router Is the Part of Agents Everyone Keeps Rebuilding

Nvidia's Router Is the Part of Agents Everyone Keeps Rebuilding

2
Comments
4 min read
Nvidia Is Buying the Part of AI Nobody Can pip install

Nvidia Is Buying the Part of AI Nobody Can pip install

2
Comments 4
4 min read
The Safety Framework Nobody Believed In Just Stopped OpenAI's Next Model

The Safety Framework Nobody Believed In Just Stopped OpenAI's Next Model

2
Comments
4 min read
ABSeeker Shows Why Agent Training Needs Receipts, Not Just Rewards

ABSeeker Shows Why Agent Training Needs Receipts, Not Just Rewards

3
Comments 2
4 min read
When Agents Lie to Maintainers, the Sandbox Already Failed

When Agents Lie to Maintainers, the Sandbox Already Failed

1
Comments
5 min read
DiffusionGemma Is Fast Because It Stops Pretending Text Has to Be Written Left to Right

DiffusionGemma Is Fast Because It Stops Pretending Text Has to Be Written Left to Right

3
Comments
3 min read
OpenAI's Math Post Is Really About Audit Trails

OpenAI's Math Post Is Really About Audit Trails

2
Comments 6
4 min read
Agent-Built Software Still Needs a Human-Shaped Test

Agent-Built Software Still Needs a Human-Shaped Test

2
Comments
3 min read
Robots Don't Need an LLM in the Fast Loop

Robots Don't Need an LLM in the Fast Loop

3
Comments
5 min read
LLM Safety Has a Language Gap

LLM Safety Has a Language Gap

2
Comments
5 min read
LLM-as-a-Judge Is Too Expensive to Be the Default

LLM-as-a-Judge Is Too Expensive to Be the Default

4
Comments 6
4 min read
Cheap Models Are Turning AI Routing Into Infrastructure

Cheap Models Are Turning AI Routing Into Infrastructure

8
Comments 2
3 min read
The OpenAI / Hugging Face Incident Was an Observability Failure First

The OpenAI / Hugging Face Incident Was an Observability Failure First

8
Comments 4
6 min read
Claude Opus 5 Is a Cost Cut Disguised as a Model Launch

Claude Opus 5 Is a Cost Cut Disguised as a Model Launch

Comments 2
5 min read
Agent Debugging Needs More Than Traces

Agent Debugging Needs More Than Traces

2
Comments 5
5 min read
Long-Horizon Agents Need a Flight Recorder

Long-Horizon Agents Need a Flight Recorder

1
Comments 2
5 min read
Frontier AI Access Just Became a Supply Chain Problem

Frontier AI Access Just Became a Supply Chain Problem

1
Comments 1
5 min read
Your RAG Eval Is Checking the Receipt, Not the Patient

Your RAG Eval Is Checking the Receipt, Not the Patient

1
Comments 2
3 min read
Google's TabFM Is the First Tabular AI Launch I'd Actually Put Next to SQL

Google's TabFM Is the First Tabular AI Launch I'd Actually Put Next to SQL

Comments
6 min read
The Next Agent Supply Chain Bug Will Look Like Documentation

The Next Agent Supply Chain Bug Will Look Like Documentation

Comments
5 min read
Your AI Bill Is a Routing Bug Now

Your AI Bill Is a Routing Bug Now

Comments
5 min read
AI Is Not Replacing Developers. It Is Replacing the On-Ramp.

AI Is Not Replacing Developers. It Is Replacing the On-Ramp.

2
Comments 1
5 min read
loading...