All stories
24 stories found

Make Ollama Start on Boot: A systemd Service That Works
The install script sets up a service that works on a desktop and quietly fails on a headless box. Here is the unit I actually run.

Why Ollama's First Response Is Slow (Cold Start Fix)
The first prompt after a pause takes forever, then everything is fast again. That is a model load, not a slow GPU — and it is fixable.

Running Two Local Models on One GPU Without Crashing
A coding model plus an embedding model on one 16 GB card is doable — but only if you budget VRAM deliberately instead of hoping.

Ollama Answers Get Cut Off Mid-Sentence: num_ctx vs num_predict Explained
Truncated replies and models that "forget" the start of a long document come from two different settings. Here is how I tell them apart and the values I use.

Ollama Error "model requires more system memory": How I Fixed It in 10 Minutes
Ollama refuses to load a model and says it needs more system memory than you have. Here is exactly what the number means and the four fixes that worked on my 32 GB / RTX 3060 box.

DeepSeek-R1 Repeats Itself or Outputs Gibberish: The 4 Settings That Fixed It
Endless loops, repeated sentences or mojibake from DeepSeek-R1 are almost always sampling or quantization — not a broken model. The exact parameters I now use.

Ollama Filled My Disk: How I Moved the Models Directory Safely
Ollama stores every model on the system drive by default. The exact steps I used to move ~/.ollama/models to a second drive on Linux, Windows and macOS without re-downloading.

Ollama Not Using My NVIDIA GPU in WSL2: The Fix That Finally Worked
Ollama says "no compatible GPUs were discovered" inside WSL2 even though nvidia-smi works. Here is the exact driver, toolkit and verification sequence that fixed it on my RTX 3060.

Ollama Pull Fails with "max retries exceeded" or EOF: How I Get Downloads to Finish
Model downloads that stall at 97% or die with EOF. The five things that fixed it for me: resume behaviour, MTU, DNS, proxy variables and a corrupted blob cache.

Ollama "connection refused on 127.0.0.1:11434": The 5 Causes I Have Actually Hit
Your client cannot reach Ollama on port 11434. Here are the five real causes — service not running, wrong host binding, WSL networking, port conflict, and Docker — with the command that fixes each.

DeepSeek-R1 Shows Its <think> Tags in the Output — Here Is How I Strip Them
DeepSeek-R1 prints its whole reasoning chain before the answer. Three ways I remove it: prompt-level, API-level parsing, and a small wrapper I use in every script.

أخطاء Ollama الشائعة وحلولها (من واقع 6 أشهر استخدام يومي)
تسعة أخطاء تكررت معي أكثر من غيرها، ومعها الحل الذي نجح فعلاً — من نقص الذاكرة إلى مخرجات JSON غير الصالحة.