Ollama: Silent Sampling Overrides Fixed
Two linked OpenAI-compatibility fixes stop omitted sampling options from overriding Modelfile settings, while a Glimmer parser fix addresses dropped JSON fields. Docs and community listings round out the day.
Duration: PT54S
Episode overview
This episode is a short developer briefing from Ollama.
It explains recent repository work in plain language.
- Show: Ollama
- Published: 2026-09-28T13:01:24Z
- Audio duration: PT54S
Transcript excerpt
This excerpt keeps the crawler page concise. Listen to the episode or use the RSS feed for the full update.
Good morning, it's Monday, September 28th, 2026 — this is your Ollama briefing.
The headline: omitted API parameters should no longer silently wipe out your Modelfile settings.
That's the thread linking two OpenAI compatibility fixes. In PR 18691, the chat completions endpoint no longer forces top P to 1 when the caller leaves it out, so a custom sampling value set in the Modelfile is now preserved. PR 18694 applies the same principle to the legacy completions endpoint, covering top P plus…
Second theme is structured output reliability for Glimmer. PR 18687 addresses cases where, with thinking turned off and a JSON format requested, replies lost their first field — seen as corrupted output starting mid-object on recent testing with the 30 billion parameter Glimmer model. The cause was the parser…
Finally, lightweight docs maintenance. One change adds the missing REQUIRES entry to the Modelfile reference table of contents, and two others list new community integrations: AgentBridge as a self-hosted assistant, and since-cutoff as a terminal tool for cutoff-aware dependency checks. No action is needed for the…
What's next: review any workarounds you added for default sampling values,…
Nearby episodes from Ollama
- Weekly Recap - Single-Pass Structured Outputs and MLX Speedups
- Tool-Call Reliability Fixes
- Desktop Fixes and API Correctness
- Mac Polish, MLX Reliability, and Compatibility Fixes
- GPU Efficiency Fix and MLX Hardening
- Single-Pass Structured Outputs and MLX Speedups
- Agent Reliability Fixes and Offline Model Transfer
- Thinking Separation Fix and GPU Visibility