Ollama: Weekly Recap - Single-Pass Structured Outputs and MLX Speedups

Ollama eliminated the double generation for structured outputs on thinking models and sped up Apple-silicon prompt processing. OpenAI compatibility fixes and desktop responsiveness work rounded out a reliability-focused week.

Duration: PT3M3S

Episode overview

This episode is a short developer briefing from Ollama.

It explains recent repository work in plain language.

  • Show: Ollama
  • Published: 2026-09-28T09:10:02Z
  • Audio duration: PT3M3S

Transcript excerpt

This excerpt keeps the crawler page concise. Listen to the episode or use the RSS feed for the full update.

Hello and welcome to the Ollama weekly recap for September 21st through September 28th, 2026.

This week brought 50 pull request activity items, 19 additional commits this week.

The lead pattern is less wasted work. Across server, runners, and desktop, the project removed second passes, second prefills, and blocked waiting that slowed developers down.

First, structured outputs on thinking models go single-pass. Previously, requesting a format on a reasoning model ran an unconstrained generation, cancelled it once thinking closed, then re-ran the prompt under a grammar. That cost a second prefill, dropped the chunk crossing the boundary, stitched metrics across…

Second, the MLX stack gets faster and more accurate. Change 18550 uses the gated delta kernel and folds scaling into the activation for Qwen 3 point 8, lifting prompt processing roughly 14 to 19 percent in reported M5 Max tests. Change 18603 replaces a fixed image budget for Gemma 4 with dynamic per-image selection,…

Third, compatibility and desktop reliability. On the OpenAI side, fixes in 18635 and related work accept reasoning content as an alias for reasoning, so replayed DeepSeek-style reasoning is no longer silently…

Nearby episodes from Ollama

  1. Tool-Call Reliability Fixes
  2. Desktop Fixes and API Correctness
  3. Mac Polish, MLX Reliability, and Compatibility Fixes
  4. GPU Efficiency Fix and MLX Hardening
  5. Single-Pass Structured Outputs and MLX Speedups
  6. Agent Reliability Fixes and Offline Model Transfer
  7. Thinking Separation Fix and GPU Visibility
  8. Weekly Recap - MLX Graduation, Reasoning Fixes, and Onboarding Polish