All stories
250 stories found

DeepSeek-R1 Repeats Itself or Outputs Gibberish: The 4 Settings That Fixed It
Endless loops, repeated sentences or mojibake from DeepSeek-R1 are almost always sampling or quantization — not a broken model. The exact parameters I now use.

Ollama Filled My Disk: How I Moved the Models Directory Safely
Ollama stores every model on the system drive by default. The exact steps I used to move ~/.ollama/models to a second drive on Linux, Windows and macOS without re-downloading.

Ollama Not Using My NVIDIA GPU in WSL2: The Fix That Finally Worked
Ollama says "no compatible GPUs were discovered" inside WSL2 even though nvidia-smi works. Here is the exact driver, toolkit and verification sequence that fixed it on my RTX 3060.

Ollama Pull Fails with "max retries exceeded" or EOF: How I Get Downloads to Finish
Model downloads that stall at 97% or die with EOF. The five things that fixed it for me: resume behaviour, MTU, DNS, proxy variables and a corrupted blob cache.

Ollama "connection refused on 127.0.0.1:11434": The 5 Causes I Have Actually Hit
Your client cannot reach Ollama on port 11434. Here are the five real causes — service not running, wrong host binding, WSL networking, port conflict, and Docker — with the command that fixes each.

DeepSeek-R1 Shows Its <think> Tags in the Output — Here Is How I Strip Them
DeepSeek-R1 prints its whole reasoning chain before the answer. Three ways I remove it: prompt-level, API-level parsing, and a small wrapper I use in every script.

أخطاء Ollama الشائعة وحلولها (من واقع 6 أشهر استخدام يومي)
تسعة أخطاء تكررت معي أكثر من غيرها، ومعها الحل الذي نجح فعلاً — من نقص الذاكرة إلى مخرجات JSON غير الصالحة.

كيف شغّلت نموذج ذكاء اصطناعي على لابتوب بدون كرت شاشة (تجربتي الكاملة 2026)
معالج i5 و16 جيجا رام وبدون كرت شاشة منفصل: 9 كلمات في الثانية على نموذج 7B. الإعدادات التي غيّرت النتيجة والأخطاء التي أهدرت وقتي.

How Much RAM Do You Actually Need for Local AI? I Tested 8, 16, 32 and 64GB
Four machines, same four tasks. 16GB is the real threshold, and the 64GB results were identical to 32GB on everything except one case.

I Replaced Copilot With a Fully Offline AI Assistant in VS Code
Continue + Ollama + Qwen Coder: code completion and chat inside VS Code with no account, no API key, and no code leaving the laptop.

Ollama Running Slow? 7 Fixes That Actually Worked on My Machine
Same laptop, same 7B model: 4 tokens/s before, 31 after. Seven settings were wrong. Here is the exact order I check them in.

Transcribing Audio Locally with Whisper: My Actual Workflow (No Uploads)
Two hours of interview audio transcribed on a laptop in eleven minutes, offline. The model size that is worth it, the flags that matter, and how I handle speaker names.