Ollama: MLX Performance Push and Tool-Call Reliability
The dominant pattern is MLX work to cut latency and memory overhead on Mac GPUs, backed by runtime version bumps. A second cluster tightens OpenAI compatibility and model parser handling for tool calls, plus reliability fixes for downloads and timeouts.
Duration: PT2M30S
Episode overview
This episode is a short developer briefing from Ollama.
It explains recent repository work in plain language.
- Show: Ollama
- Published: 2026-10-06T13:00:57Z
- Audio duration: PT2M30S
Transcript excerpt
This excerpt keeps the crawler page concise. Listen to the episode or use the RSS feed for the full update.
Good morning, it's Tuesday, October sixth, twenty twenty-six, and this is your Ollama developer briefing.
The headline is a performance and stability push for Mac GPU models, alongside a second effort to make tool calling behave consistently.
First, MLX. Merged change eighteen eight oh seven tackles high latency after the GPU sits idle, by enabling a residency refresh every second while still honoring explicit overrides, fixing issue eighteen seven four four. Related merged work in eighteen eight oh six cuts overhead when resolving model names and…
Second, tool calling. Merged pull request eighteen seven two two keeps tool message content together in one message, preserving the tool call ID and name so renderers don't see extra unnamed results. Three open proposals cover the edges: eighteen eight oh four closes open text before streaming tool calls to keep…
Finally, reliability fixes. Open proposal eighteen eight one three would surface disk-full errors during blob download instead of reporting one hundred percent progress, and eighteen eight oh oh would clamp very large keep-alive and load timeout values before they overflow. Merged work in eighteen four seven oh…
What's next: watch…
Nearby episodes from Ollama
- Tokenizer Parity and RC-Aware Updates
- Weekly Recap - System One Decisions Take Center Stage
- MLX Tokenizer Fixes and Thinking Model Reliability
- Tool Call Ordering Fixes and Decision Models
- Decision Models Take Shape
- Pointer-Head Scoring and Reliability Fixes
- System One Moves to Explicit Capabilities
- System One Decisions Launch