DEV Community

#llm

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
Gemma 4 E2B on a Single TPU v6e Chip: A Serving Deep Dive

Gemma 4 E2B on a Single TPU v6e Chip: A Serving Deep Dive

1
Comments
8 min read
Gemma4 DevOps In Action

Gemma4 DevOps In Action

8
Comments 2
7 min read
Prompt Engineering for Manual Testers: How to Get Useful Output from AI Tools

Prompt Engineering for Manual Testers: How to Get Useful Output from AI Tools

6
Comments 5
11 min read
Can You Beat an LLM? Building Humans vs. Humanity's Last Exam

Can You Beat an LLM? Building Humans vs. Humanity's Last Exam

8
Comments
4 min read
How I Built MailOS: My Journey With Qwen Cloud

How I Built MailOS: My Journey With Qwen Cloud

2
Comments 10
5 min read
My Local AI Assistant Got Worse When I Remembered Too Much

My Local AI Assistant Got Worse When I Remembered Too Much

1
Comments 2
3 min read
A 2,181-video field report made my open-source video tool better in one day

A 2,181-video field report made my open-source video tool better in one day

Comments
2 min read
AI Agents That Live Inside a Dreamed-Up World

AI Agents That Live Inside a Dreamed-Up World

Comments
3 min read
I Gave 3 AI Agents a Decaying Notepad and They Built a Culture

I Gave 3 AI Agents a Decaying Notepad and They Built a Culture

Comments
3 min read
I Built a Slot-Filling Onboarding Bot. It Leaked a Validation Error Into the Chat

I Built a Slot-Filling Onboarding Bot. It Leaked a Validation Error Into the Chat

1
Comments
4 min read
I Watched Two AI Agents Invent Their Own Language

I Watched Two AI Agents Invent Their Own Language

1
Comments
3 min read
How Does an LLM Request and Response Cycle Work? A Full Walkthrough

How Does an LLM Request and Response Cycle Work? A Full Walkthrough

Comments
8 min read
Five Comments That Redesigned My LLM Verification Pipeline

Five Comments That Redesigned My LLM Verification Pipeline

2
Comments 1
16 min read
Optimizing RAG at Scale: Chunking, Retrieval, and the Bayesian Search That Cut Latency 40%

Optimizing RAG at Scale: Chunking, Retrieval, and the Bayesian Search That Cut Latency 40%

Comments
5 min read
Optimizing RAG at Scale: Chunking, Retrieval, and the Bayesian Search That Cut Latency 40%

Optimizing RAG at Scale: Chunking, Retrieval, and the Bayesian Search That Cut Latency 40%

Comments
5 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.