Blog
All stories
3 stories found
Active filters:
Tag: performance

AI Tools
Why Ollama's First Response Is Slow (Cold Start Fix)
The first prompt after a pause takes forever, then everything is fast again. That is a model load, not a slow GPU — and it is fixable.
August 9, 2026 · 5 min read

AI Tools
Running Two Local Models on One GPU Without Crashing
A coding model plus an embedding model on one 16 GB card is doable — but only if you budget VRAM deliberately instead of hoping.
August 9, 2026 · 5 min read

AI
Ollama Running Slow? 7 Fixes That Actually Worked on My Machine
Same laptop, same 7B model: 4 tokens/s before, 31 after. Seven settings were wrong. Here is the exact order I check them in.
July 30, 2026 · 7 min read