Timeline
What’s new in open models
Releases, launches and announcements across models, hardware and tools — each linked to the things it affects.
New versions of the programs that run models.Show everything
2026 90
- Toolruntime release GitHubAll LiteLLM Docker images are signed with cosign. Every release is signed with the same key introduced in commit 0112e53.
- Toolruntime release GitHubYou can now run Qwen-Image-2.1 locally with Unsloth! This release also includes custom Agent Skills, and easier chat/project management. It also brings 2x faster reasoning blocks (60 FPS vs 30 FPS), more reliable training, and improved Linux installs and updates. Qwen Image 2.1 G
- Toolruntime release GitHub📉 Far smaller slim image. A slim build now comes down at around 175 MB, near enough 89% smaller than the last release, the local models, the packages around them and the tools that installed them all gone from it; what that changes about the way an instance behaves is set out un
- Toolruntime release Official feedv0.30.0 Highlights This release features 762 commits from 315 contributors (104 new)! New models : DeepSeek-V4.1-Flash ( #56214 , #56228 , #56208 ) with the whole KV stored in MXFP8 through the FlashMLA V4.1 record on SM100 ( #56893 ), DeepGEMM Mega-mHC ( #56962 ), and async…
- Toolruntime release GitHub713 PRs from 237 contributors.
- Toolruntime release GitHubWe're releasing our new updated Docker image along with multi-user accounts, RDNA1+2, FP8/INT8 diffusion support, ARM64 CUDA Windows support and many training, GRPO and inference improvements. Also a Qwen3.8-Flash-Next 2x faster MTP hotfix.
- Toolruntime release Official feedRelease vllm-proto 0.3.0
- Toolruntime release GitHub
- Toolruntime release Official feedValidated by PR #56538 CI at fa2a26f .
- Toolruntime release GitHubAdded first-run setup when running ollama, with options to sign in or continue locally. Setup completion is shared with the desktop app on macOS and Windows.
- Toolruntime release GitHubAll LiteLLM Docker images are signed with cosign. Every release is signed with the same key introduced in commit 0112e53.
- Toolruntime release GitHubMLX safetensors ollama create no longer experimental. GGUF model creation now requires using llama.cpp tooling for safetensor conversion and quantization.
- Toolruntime release GitHubllama.cpp 0.4.1 adds Maple 20B-A1B, Tencent Hy 4, and Spark2.5 support, improves JSON schema handling, chat parsing, logging, and server child-process management, and updates ggml to v0.24.0.
- Toolruntime release Official feedvllm-proto 0.1.0
- Toolruntime release GitHubAll LiteLLM Docker images are signed with cosign. Every release is signed with the same key introduced in commit 0112e53.
- Toolruntime release GitHub72 commits since v0.18.0 (July 17, 2026).
- Toolruntime release GitHubThis is a large performance and reliability + bug fix release for Unsloth
- Toolruntime release GitHubThis release features 594 commits from 277 contributors (91 new)!
- Toolruntime release GitHubAll LiteLLM Docker images are signed with cosign. Every release is signed with the same key introduced in commit 0112e53.
- Toolruntime release GitHubOllama models can now be used directly in ChatGPT Desktop, so you can keep your existing workflow while running open models. Setup is available from the Ollama app on MacOS.
- Toolruntime release GitHub786 PRs from 214 contributors.
- Toolruntime release GitHubRun Qwen3.8-Flash-Next and GLM-5.3-Flash up to 2x faster with MTP. MTP is enabled by default, you can still disable it.
- Toolruntime release GitHubThis release ships container images only. There is no PyPI package for `1.99.1`.
- Toolruntime release GitHubgemma4 now supports images and audio on MLX engine
- Toolruntime release GitHubA fix-focused release. The main fixes are for few-shot leakage, a multiple-choice filter bug, and group stderr, alongside two new ONNX backends and eight new benchmark suites. Also updated most configs for datasets>=4, which accounts for much of the diff by volume.
- Toolruntime release GitHub♿ Accessibility mode reaches the menus. Accessibility mode now marks the menu entry you are pointing at and the model already chosen with a stronger background, across the dropdown menus, their submenus, and the model picker together with its filter and compare controls, so those
- Toolruntime release GitHub🖼️ Richer previews for terminal files. Word documents and slide decks produced in the terminal are now previewed as the finished document rather than an approximation, and every document preview gains a page strip down the side with numbered thumbnails you can click to jump stra
- Toolruntime release GitHubOllama's app now follows the system appearance again, restoring dark mode support
- Toolruntime release GitHubQwen3.8-Flash-Next and GLM-5.3-Flash can now run locally in Unsloth!
- Toolruntime release GitHubMLX: Qwen3.8 Flash Next support
- Toolruntime release GitHubThis release features 584 commits from 270 contributors (76 new)!
- Toolruntime release GitHub🚦 Human in the loop tool approval. Where an administrator has turned it on, you can switch a conversation from letting tools run freely to being asked first, so a model that wants to use a tool stops and waits for you to allow or deny it, one call at a time in a saved conversati
- Toolruntime release GitHubThanks for the support for Qwen3.8-27B and Unsloth Desktop! This is a bug fix release with 170+ PRs.
- Toolruntime release GitHub710 PRs from 212 contributors.
- Toolruntime release GitHubDevelopers can now easily configure Claude Desktop to seamlessly work with Ollama as a third-party gateway provider.
- Toolruntime release GitHubThanks for the support for Qwen3.8-27B and Unsloth Desktop last week! For this release, we merged 200+ PRs to introduce many new features, fixes including:
- Toolruntime release GitHubNew desktop onboarding flow on first launch
- Toolruntime release GitHubllm: transcode WebP images for llama-server
- Toolruntime release GitHubqwen3.8: support developer instructions
- Toolruntime release GitHubThis release adds the support of Qwen 3.8 27B. Qwen3.8 delivers substantial gains across coding, professional work, research, and long-horizon agentic tasks.
- Toolruntime release GitHubQwen3.8-27B and Qwen3.8-2.4T can now be run locally in Unsloth!
- Toolruntime release GitHubollama launch dsh now supports DeepSeek Harness, DeepSeek's open-source agent harness
- Toolruntime release GitHubUnsloth Desktop is here! The first desktop app to run and train AI models locally. Research, export and deploy from the same open-source app on Windows, macOS and Linux.
- Toolruntime release GitHubThis is a patch release on top of v0.27.0.
- Toolruntime release GitHubThis release features 561 commits from 242 contributors (64 new)!
- Toolruntime release GitHub582 PRs from 194 contributors.
- Toolruntime release GitHub🎨 Redesigned interface. Open WebUI has been visually rebuilt from the ground up. All aspects of the User Interface, from the chat view to the admin panel. Now with a narrower conversation column, lighter typography, tidier spacing, consistent menus and dropdowns, clearly outline
- Toolruntime release GitHubThis release features 411 commits from 212 contributors (61 new)!
- Toolruntime release GitHub574 PRs from 169 contributors.
- Toolruntime release GitHubWe've been hard at work doing low level improvements in the kernels. Over 90 commits since v0.17.0 (June 3, 2026), themed around fine-tuning very large sparse-MoE models cheaply: 4-bit expert LoRA/QLoRA (NVFP4, MXFP4, bnb) that runs fast and stays memory-flat at long context on B
- Toolruntime release GitHubThis release features 2 commits from 2 contributors (1 new)!
- Toolruntime release GitHubv0.5.15.post1 includes a few patches, mostly for GLM 5.2
- Toolruntime release GitHubThis release features 558 commits from 232 contributors (64 new)!
- Toolruntime release GitHubGLM-5.2 NVFP4, tuned for production: We took time this cycle to tune GLM-5.2 NVFP4 on Blackwell for optimized production serving. It now runs at 500+ tok/s/user on 8x B300, 450 on 4x GB300 (bs=1). Run GLM-5.2 with our cookbook.
- Toolruntime release GitHub💭 Streamed reasoning display. Models that emit thinking or reasoning now show that content as it streams, and it renders correctly in the chat overview and in exported conversations. Commit, Commit, Commit, Commit
- Toolruntime release GitHubThis release features 571 commits from 256 contributors (77 new)!
- Toolruntime release GitHub🤝 Shared folder read-only chats no longer sign users out. Opening or reading chats from shared folders now keeps the current session active when a resource-level access error is returned, instead of incorrectly showing "Session expired. Please sign in again."
- Toolruntime release GitHub🤝 Share folders with your team. You can now share a folder and the chats inside it with specific users, groups, or everyone, with read or write access; people you share with see shared folders in their sidebar and open the chats in a read-only view when they are not the owner, a
- Toolruntime release GitHubNew Model Support: GLM-5.2, LiquidAI LFM2.5, Kimi-K2.7-Code, Poolside Laguna-M.1, DiffusionGemma, Zyphra ZAYA1, MiMo-V2-ASR
- Toolruntime release GitHub
- Toolruntime release GitHubStable release of the Continue VS Code extension (final release): removes the CLI-install banner and Generate Rule feature, switches onboarding and the new-config template to explicit model definitions (no Hub slugs), and updates the deprecation banner export link. Re-cut of v1.2
- Toolruntime release GitHubStable release of the Continue VS Code extension (final release): removes the CLI-install banner and Generate Rule feature, switches onboarding and the new-config template to explicit model definitions (no Hub slugs), and updates the deprecation banner export link.
- Toolruntime release GitHubPlease note that Minimax M3 is not yet supported in this version. Please follow vLLM recipe for usage guides for M3.
- Toolruntime release GitHubNew Model Support:
- Toolruntime release GitHubThis release features 8 commits from 6 contributors (1 new)!
- Toolruntime release GitHubAnother packed release. ~84 commits since v0.16.1 last month, bringing Expert Parallelism for MoE training, BitNet 1.58-bit fine-tuning, remote training via Tinker, context parallelism for hybrid SSM models, MXFP4 ScatterMoE-LoRA, fused RMSNorm+RoPE kernels for the Qwen3 family,
- Toolruntime release GitHub📦 Official knowledge base sync tool. A new companion tool from Open WebUI, oikb, keeps a knowledge base in sync with a local directory, GitHub repo, S3 bucket, Confluence space, or any of more than 40 other sources, uploading only new and changed files using the incremental sync
- Toolruntime release GitHubv0.5.12.post1 is a stability patch on top of v0.5.12. It cherry-picks 12 fixes — primarily for DeepSeek V4 — onto the release branch.
- Toolruntime release GitHubDeepSeek V4 support: Full inference path for DeepSeek-V4 (#23882), including:
- Toolruntime release GitHubNew release with four new model backends, tensor parallel support for transformers based models (hf), new benchmarks, a TaskManager refactor, and a long tail of task correctness fixes.
- Toolruntime release GitHub🛡️ Redirect-based SSRF protection. All outbound HTTP requests now block 3xx redirects by default via a new AIOHTTP_CLIENT_ALLOW_REDIRECTS environment variable, preventing redirect-based SSRF where a public URL silently redirects to internal addresses (RFC 1918, loopback, cloud-m
- Toolruntime release GitHub📜 Chat scroll position on load. Opening a chat conversation now reliably scrolls to the bottom of the message history, fixing a regression caused by content-visibility: auto where estimated element sizes prevented the initial scroll from reaching the true bottom.
- Toolruntime release GitHubLots of bugfixes
- Toolruntime release GitHubCaching system prompt and user messages for non-trimmable caches
- Toolruntime release GitHubExample YAML: https://github.com/axolotl-ai-cloud/axolotl/blob/main/examples/gemma4/26b-a4b-moe-qlora.yaml
- Toolruntime release GitHubWe’re very excited to share this new packed release. We had ~80 new commits since v0.15.0 (March 6, 2026).
- Toolruntime release GitHubfix: ensure config.yaml exists and is populated when accessed by @RomneyDa in https://github.com/continuedev/continue/pull/11915
- Toolruntime release GitHubFix save/load of CacheList by @angeloskath in https://github.com/ml-explore/mlx-lm/pull/886
- Toolruntime release GitHubThis release brings new model support, significant MoE improvements, infrastructure updates with Torch 2.10.0 and uv builds, and a collection of quality-of-life fixes across the board.
- Toolruntime release GitHubMinor release. Stay tuned for bigger changes next release.
- Toolruntime release GitHubFix Kimi Linear by @kernelpool in https://github.com/ml-explore/mlx-lm/pull/853
- Toolruntime release GitHubTransformers v5 by @awni in https://github.com/ml-explore/mlx-lm/pull/811
- Toolruntime release GitHubThis is a major release marking our migration to Transformers v5. Along with this significant core dependency upgrade, we are introducing major performance optimizations for MoE models and new fine-tuning methods.
- Toolruntime release GitHubThe big change this release: the base package no longer installs model backends by default. We've also added new benchmarks and expanded multilingual support.
- Toolruntime release GitHubimport logging as it throws no logging error in place of actual error by @Maanas-Verma in https://github.com/ml-explore/mlx-lm/pull/778
- Toolruntime release GitHubThis is a patch release introducing GDPO support and updating core infrastructure, including newer CUDA defaults and Python versions.
- Toolruntime release GitHubThis release brings support for PyTorch 2.9.1, expands our ecosystem with new experiment trackers (SwanLab and Trackio), and introduces support for a wide range of new models including Olmo3, Ministral 3, InternVL 3.5, and Kimi. We’ve also included significant improvements to qua
- Toolruntime release GitHubAdd AWQ/GPTQ weight transformation utilities by @ericcurtin in https://github.com/ml-explore/mlx-lm/pull/730
- Toolruntime release GitHubFix mlx-lm release by @awni in https://github.com/ml-explore/mlx-lm/pull/733
- Toolruntime release GitHubcustom dsv32 chat template by @awni in https://github.com/ml-explore/mlx-lm/pull/693
2025 24
- Toolruntime release GitHubfix: server busy-waiting during idle request polling by @zenyr in https://github.com/ml-explore/mlx-lm/pull/674
- Toolruntime release GitHubThis release is packed with major new features, including Streaming SFT for massive datasets, a new Text Diffusion training plugin, and a significant upgrade to our Quantization-Aware Training (QAT) capabilities with NVFP4 support. We're also thrilled to announce support for a hu
- Toolruntime release GitHubThis release continues our steady stream of community contributions with a batch of new benchmarks, expanded model support, and important fixes. A notable change: Python 3.10 is now the minimum required version.
- Toolruntime release GitHubAdded support for all GPT-5 models.
- Toolruntime release GitHubThis v0.4.9.1 release is a quick patch to bring in some new tasks and fixes. Looking aheas, we're gearing up for some bigger updates to tackle common community pain points. We'll do our best to keep things from breaking, but we anticipate a few changes might not be fully backward
- Toolruntime release GitHub
- Toolruntime release GitHubAdded support for new Gemini models including gemini-2.5-pro, gemini-2.5-flash, and gemini-2.5-pro-preview-06-05 with thinking tokens support.
- Toolruntime release GitHubEnhanced Backend Support:
- Toolruntime release GitHubAdded support for new Claude models including the Sonnet 4 and Opus 4 series (e.g., claude-sonnet-4-20250514,
- Toolruntime release GitHub
- Toolruntime release FixtureAutomated build release.
- Toolruntime release FixtureNew multimodal engine.
- Toolruntime release GitHubAdd Qwen3 and Qwen3MoE support
- Toolruntime release GitHubAdded support for gemini-2.5-pro-preview-05-06 models.
- Toolruntime release GitHubAdd Torch 2.7.0 wheels (big thanks to @kingbri1 for unborking the build action)
- Toolruntime release GitHubSupport for GPT 4.1, mini and nano.
- Toolruntime release GitHubAdded support for the openrouter/openrouter/quasar-alpha model.
- Toolruntime release GitHubOpenRouter OAuth integration:
- Toolruntime release GitHubAdded support for SOTA Gemini 2.5 Pro.
- Toolruntime release GitHubAdded support for thinking tokens for OpenRouter Sonnet 3.7.
- Toolruntime release GitHubBig upgrade in programming languages supported by adopting tree-sitter-language-pack.
- Toolruntime release GitHubNew Backend Support:
- Toolruntime release GitHubSupport Qwen2.5-VL
- Toolruntime release FixturevLLM V1 engine enters alphaRe-architected core engine with lower CPU overhead.
2024 7
- Toolruntime release GitHubBasic video support for Qwen2-VL
- Toolruntime release GitHubThis release includes several bug fixes, minor improvements to model handling, and task additions.
- Toolruntime release GitHubSome small fixes, most notably for Qwen2-VL inference on Windows
- Toolruntime release GitHubInitial support for Qwen2-VL (images for now, no video)
- Toolruntime release GitHubThis release brings important changes to chat template handling, expands our task library with new multilingual and multimodal benchmarks, and includes various bug fixes.
- Toolruntime release GitHubSupport Pixtral
- Toolruntime release GitHubNo longer use safetensors for loading weights (fix virtual memory issues on Windows especially)