Ollama (Stable) 0.34.1
What's Changed * MLX safetensors ollama create no longer experimental. GGUF model creation now requires using llama.cpp tooling for safetensor conversion and quantization. * Improved MLX memory handling on Apple Silicon * Runaway repeat token detection now requires 100 repeat tokens for reduced false positives (e.g. OCR) * /api/tags is much faster on large model libraries (3.1 s → 294 ms cold in testing), and model capabilities are now reported consistently. * Deprecated typical_p: it can n…