Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
vllm
Follow
Hide
Posts
Left menu
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
Gemma 4 E2B on a Single TPU v6e Chip: A Serving Deep Dive
xbill
xbill
xbill
Follow
for
Google Developer Experts
Jul 21
Gemma 4 E2B on a Single TPU v6e Chip: A Serving Deep Dive
#
tpu
#
llm
#
vllm
#
googlecloud
2
 reactions
Comments
Add Comment
8 min read
tpu-management: a Claude Code skill for running Gemma 4 on Cloud TPUs
xbill
xbill
xbill
Follow
for
Google Developer Experts
Jul 21
tpu-management: a Claude Code skill for running Gemma 4 on Cloud TPUs
#
googlecloud
#
tpu
#
vllm
#
claudecode
2
 reactions
Comments
Add Comment
3 min read
Does a Second GPU Increase Ollama's Context Window? (Quadro P2000 + RTX 3090 Tested)
Arsen Apostolov
Arsen Apostolov
Arsen Apostolov
Follow
Jul 9
Does a Second GPU Increase Ollama's Context Window? (Quadro P2000 + RTX 3090 Tested)
#
llm
#
ollama
#
vllm
#
gpu
Comments
Add Comment
3 min read
vLLM vs llama.cpp vs Ollama: What Happens When Your Model Doesn't Fit in 24GB VRAM
Arsen Apostolov
Arsen Apostolov
Arsen Apostolov
Follow
Jul 5
vLLM vs llama.cpp vs Ollama: What Happens When Your Model Doesn't Fit in 24GB VRAM
#
llm
#
homelab
#
vllm
#
ai
Comments
Add Comment
6 min read
AI Inference Optimization in Cloud-Native Environments: GPU Orchestration, Edge Deployment, and Latency Reduction at Scale
The Cyber Sidekick
The Cyber Sidekick
The Cyber Sidekick
Follow
Jun 23
AI Inference Optimization in Cloud-Native Environments: GPU Orchestration, Edge Deployment, and Latency Reduction at Scale
#
kubernetesgpuscheduling
#
aiinferenceoptimization
#
llmdeployment
#
vllm
Comments
Add Comment
4 min read
AI Inference at the Edge: Running Real-Time LLMs in Kubernetes Without a GPU Farm
The Cyber Sidekick
The Cyber Sidekick
The Cyber Sidekick
Follow
Jun 18
AI Inference at the Edge: Running Real-Time LLMs in Kubernetes Without a GPU Farm
#
edgeai
#
kubernetes
#
llminference
#
vllm
Comments
Add Comment
3 min read
Qwen3.6-35B NVFP4 runs on one H100 — A100 owners are out
Creeta
Creeta
Creeta
Follow
Jun 18
Qwen3.6-35B NVFP4 runs on one H100 — A100 owners are out
#
qwen3
#
nvfp4
#
vllm
#
nvidia
Comments
Add Comment
8 min read
I built an open-source alternative to Microsoft's KAITO that works on ANY Kubernetes cluster
GaeaRuiW
GaeaRuiW
GaeaRuiW
Follow
Jun 9
I built an open-source alternative to Microsoft's KAITO that works on ANY Kubernetes cluster
#
kubernetes
#
vllm
#
devops
#
opensource
Comments
Add Comment
2 min read
Prefix caching at scale: when it saves you 80% of prefill cost, and the eviction policies that quietly turn it into 5%
Tech_Nuggets
Tech_Nuggets
Tech_Nuggets
Follow
Jun 7
Prefix caching at scale: when it saves you 80% of prefill cost, and the eviction policies that quietly turn it into 5%
#
llm
#
ai
#
infrastructure
#
vllm
Comments
Add Comment
9 min read
KV cache quantization: what FP8/INT8 K and V actually buy you, and where they break
Tech_Nuggets
Tech_Nuggets
Tech_Nuggets
Follow
Jun 6
KV cache quantization: what FP8/INT8 K and V actually buy you, and where they break
#
llm
#
ai
#
vllm
#
performance
1
 reaction
Comments
Add Comment
8 min read
What a green GPU dashboard hides
Harshit Luthra
Harshit Luthra
Harshit Luthra
Follow
Jul 2
What a green GPU dashboard hides
#
gpu
#
observability
#
vllm
#
prometheus
Comments
Add Comment
9 min read
The cheapest speedup is your load balancer
Harshit Luthra
Harshit Luthra
Harshit Luthra
Follow
Jul 2
The cheapest speedup is your load balancer
#
gpu
#
inference
#
routing
#
vllm
Comments
Add Comment
8 min read
One box, eight GPUs, and the wires between them
Harshit Luthra
Harshit Luthra
Harshit Luthra
Follow
Jul 2
One box, eight GPUs, and the wires between them
#
gpu
#
nvlink
#
nccl
#
vllm
Comments
Add Comment
10 min read
Qwen3.6-27B + vLLM + Hermes on 24GB VRAM: May 2026 Recipe
Xavier Rey-Robert
Xavier Rey-Robert
Xavier Rey-Robert
Follow
Jun 19
Qwen3.6-27B + vLLM + Hermes on 24GB VRAM: May 2026 Recipe
#
ai
#
llm
#
vllm
#
agents
1
 reaction
Comments
Add Comment
4 min read
Two Qwen3 Models on One DGX Spark: The Residency Math for Local LLM Coding
Devashish
Devashish
Devashish
Follow
Jun 16
Two Qwen3 Models on One DGX Spark: The Residency Math for Local LLM Coding
#
localllm
#
vllm
#
ai
#
nvidia
1
 reaction
Comments
Add Comment
5 min read
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account