DEV Community

#vllm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Gemma 4 E2B on a Single TPU v6e Chip: A Serving Deep Dive

Gemma 4 E2B on a Single TPU v6e Chip: A Serving Deep Dive

2
Comments
8 min read
tpu-management: a Claude Code skill for running Gemma 4 on Cloud TPUs

tpu-management: a Claude Code skill for running Gemma 4 on Cloud TPUs

2
Comments
3 min read
Does a Second GPU Increase Ollama's Context Window? (Quadro P2000 + RTX 3090 Tested)

Does a Second GPU Increase Ollama's Context Window? (Quadro P2000 + RTX 3090 Tested)

Comments
3 min read
vLLM vs llama.cpp vs Ollama: What Happens When Your Model Doesn't Fit in 24GB VRAM

vLLM vs llama.cpp vs Ollama: What Happens When Your Model Doesn't Fit in 24GB VRAM

Comments
6 min read
AI Inference Optimization in Cloud-Native Environments: GPU Orchestration, Edge Deployment, and Latency Reduction at Scale

AI Inference Optimization in Cloud-Native Environments: GPU Orchestration, Edge Deployment, and Latency Reduction at Scale

Comments
4 min read
AI Inference at the Edge: Running Real-Time LLMs in Kubernetes Without a GPU Farm

AI Inference at the Edge: Running Real-Time LLMs in Kubernetes Without a GPU Farm

Comments
3 min read
Qwen3.6-35B NVFP4 runs on one H100 — A100 owners are out

Qwen3.6-35B NVFP4 runs on one H100 — A100 owners are out

Comments
8 min read
I built an open-source alternative to Microsoft's KAITO that works on ANY Kubernetes cluster

I built an open-source alternative to Microsoft's KAITO that works on ANY Kubernetes cluster

Comments
2 min read
Prefix caching at scale: when it saves you 80% of prefill cost, and the eviction policies that quietly turn it into 5%

Prefix caching at scale: when it saves you 80% of prefill cost, and the eviction policies that quietly turn it into 5%

Comments
9 min read
KV cache quantization: what FP8/INT8 K and V actually buy you, and where they break

KV cache quantization: what FP8/INT8 K and V actually buy you, and where they break

1
Comments
8 min read
What a green GPU dashboard hides

What a green GPU dashboard hides

Comments
9 min read
The cheapest speedup is your load balancer

The cheapest speedup is your load balancer

Comments
8 min read
One box, eight GPUs, and the wires between them

One box, eight GPUs, and the wires between them

Comments
10 min read
Qwen3.6-27B + vLLM + Hermes on 24GB VRAM: May 2026 Recipe

Qwen3.6-27B + vLLM + Hermes on 24GB VRAM: May 2026 Recipe

1
Comments
4 min read
Two Qwen3 Models on One DGX Spark: The Residency Math for Local LLM Coding

Two Qwen3 Models on One DGX Spark: The Residency Math for Local LLM Coding

1
Comments
5 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.