All stories
6 stories found

Running Two Local Models on One GPU Without Crashing
A coding model plus an embedding model on one 16 GB card is doable — but only if you budget VRAM deliberately instead of hoping.

Ollama Not Using My NVIDIA GPU in WSL2: The Fix That Finally Worked
Ollama says "no compatible GPUs were discovered" inside WSL2 even though nvidia-smi works. Here is the exact driver, toolkit and verification sequence that fixed it on my RTX 3060.

Which GPU Should You Buy for Local AI in 2026? (I Tested Five Price Brackets)
VRAM is the only spec that decides what you can run. Here is what each budget tier actually gets you, with measured tokens per second and the models that fit.

CUDA Out of Memory: 9 Fixes That Actually Worked on My GPU
The full error, what each number in it means, and nine fixes ordered by how likely they are to solve your problem — starting with the two that fix it 80% of the time.

NVIDIA Blackwell Ultra: What It Means for AI Startups in 2026
Blackwell Ultra reshapes AI economics. B200 Ultra, GB300, and NVLink 6 for inference cost, training, and startup strategy.

Nvidia's Next Move: What the Rubin Roadmap Means for AI Buyers Right Now
Nvidia's Rubin roadmap is starting to shape 2027 capacity planning. Here is what US AI buyers should factor into contracts and cloud commitments right now.