DEV Community

#localllm

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
An LLM router that tells you why it chose local, and what "why" has to mean

An LLM router that tells you why it chose local, and what "why" has to mean

Comments
3 min read
I Ran DeepSeek V4 Flash Across Two DGX Sparks Over Ethernet

I Ran DeepSeek V4 Flash Across Two DGX Sparks Over Ethernet

Comments
11 min read
VRAM and RAM for local LLMs — honest planning bands, not a GPU tier list

VRAM and RAM for local LLMs — honest planning bands, not a GPU tier list

Comments
4 min read
Running Ollama on a 32 GB MacBook Air: A Practical First Setup

Running Ollama on a 32 GB MacBook Air: A Practical First Setup

Comments
6 min read
Running llama.cpp on a 32 GB MacBook Air: A Direct Comparison with Ollama

Running llama.cpp on a 32 GB MacBook Air: A Direct Comparison with Ollama

Comments 1
9 min read
Running a 35B MoE Model on an 8 GB Laptop GPU: Testing FreeToken

Running a 35B MoE Model on an 8 GB Laptop GPU: Testing FreeToken

Comments 3
7 min read
Temperature 0 is not reproducible. I measured 30 percent of my output changing between identical runs.

Temperature 0 is not reproducible. I measured 30 percent of my output changing between identical runs.

1
Comments 1
3 min read
Our 4B beat Claude Opus on a 440K-token corpus. Then it came last on the public benchmark.

Our 4B beat Claude Opus on a 440K-token corpus. Then it came last on the public benchmark.

2
Comments
4 min read
I told the model to separate fields with <TAB>. It did exactly that, and I lost 79 percent of my data.

I told the model to separate fields with <TAB>. It did exactly that, and I lost 79 percent of my data.

1
Comments
3 min read
A 4B on a 6GB laptop matched frontier-model accuracy on aggregation — except when the answer is a number

A 4B on a 6GB laptop matched frontier-model accuracy on aggregation — except when the answer is a number

1
Comments
4 min read
Your agent truncates the corpus and answers anyway. Two harnesses, and a router that picks between them.

Your agent truncates the corpus and answers anyway. Two harnesses, and a router that picks between them.

1
Comments
5 min read
How to Run a Free AI Coding Assistant Locally with VS Code, opencode, and LM Studio

How to Run a Free AI Coding Assistant Locally with VS Code, opencode, and LM Studio

Comments
4 min read
What really fits in 8GB VRAM

What really fits in 8GB VRAM

Comments
7 min read
Moving Scheduled LLM Curation from Cloud APIs to Local Models

Moving Scheduled LLM Curation from Cloud APIs to Local Models

Comments
9 min read
Nine ways to talk to a local model

Nine ways to talk to a local model

Comments
9 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.