Release monitor / Verified records

AI Model Launch Tracker

Every consequential model release, structured.

A source-backed record of AI model launches, developers, availability, licenses, context windows, release types, and disclosed pricing.

Release distributionModels entering the market
22 of 67
Release date (X) · Disclosed context window (Y) · Undisclosed releases use a separate lane.Hover or focus to inspect
Verified records67
Source evidence197
Named participants108
Named participants108

Maintained dataset

Latest verified model launches signals.

Records update as sources arrive. Open any row to inspect captured facts, confidence, completeness, and supporting evidence.

Showing 41-50 of 67 verified records
Aug 20, 20262 sources
deepseek-v4-flash-vision-expDeepSeek
experimental API modelLicense not disclosed

Available through Chat Completions, Messages, and Responses requests

Details

What we captured

DeepSeek announced an experimental vision model supporting Chat Completions, Messages, and Responses requests, with a 384-token billing ceiling per image.

  • The model accepts mixed text and image requests.
  • Each image is capped at 384 billable V4-Flash tokens.
  • The production V4-Flash endpoint remains text-only.
  • Peak uncached input pricing is $0.44 per million tokens and peak output pricing is $1.32 per million tokens.
Aug 20, 20261 source
V4-Flash-Vision-ExpDeepSeek
experimental API modelLicense not disclosed

DeepSeek API

Details

What we captured

DeepSeek released V4-Flash-Vision-Exp, an experimental multimodal model available through its API and compatible with OpenAI and Anthropic API formats.

  • The model adds image processing to DeepSeek-V4-Flash.
  • DeepSeek says it nearly matches Opus 4.8 on internal multimodal agent benchmarks.
  • The model supports image inputs and visual agent workflows.
  • Harness version 0.1.1 supports the model out of the box.
Aug 20, 20262 sources
Falcon 2Murf AI
proprietary foundation voice modelLicense not disclosed

Availability not disclosed

Details

What we captured

Murf AI announced Falcon 2, a proprietary voice foundation model positioned against leading real-time voice models on naturalness, latency, and cost.

  • Murf AI announced Falcon 2 on August 21, 2026.
  • The article reports that Falcon 2 ranked ahead of ElevenLabs Turbo v2.5, ElevenLabs Flash v2.5, and OpenAI Realtime on Artificial Analysis’s Speech Arena.
  • Murf states that Falcon 2 provides sub-100-millisecond time-to-first-audio and costs $0.01 per generated minute.
  • The model supports 150+ voices across 35+ languages and up to 10,000 concurrent calls.
Aug 18, 20260 sources
Ornith-1.5Ornith AI
open weightsLicense not disclosed

Hugging Face collection; local deployment and quantized mobile version described

Details

What we captured

Ornith AI released Ornith-1.5, a three-size open model family designed to generate its own training tasks, scaffolds, and solution rollouts.

  • The release date is August 19, 2026.
  • The family contains 397B MoE, 35B MoE with 3B active parameters per token, and 9B dense variants.
  • Ornith AI says the models generate harder tasks, task-specific scaffolds, and solution rollouts within a reinforcement-learning loop.
  • Weights, serving instructions, and quantized formats were published for developer testing.

Source evidence

Confidence 100% / Completeness 87%

Aug 17, 20263 sources
GPT-5.6 SolOpenAI
APILicense not disclosed

Limited release to a small group of customers; wider access planned later

Details

What we captured

OpenAI announced a limited API release of an Ultrafast mode for GPT-5.6 Sol, powered by Cerebras and capable of up to 750 output tokens per second.

  • OpenAI is testing an Ultrafast mode for GPT-5.6 Sol.
  • The mode is available to a small group of customers through the OpenAI API.
  • The mode is powered by Cerebras and can produce up to 750 output tokens per second.
Aug 15, 20263 sources
GLM-5.3Zhipu (Z.ai)
Model releaseLicense not disclosed

Availability not disclosed

Details

What we captured

Zhipu announced GLM-5.3 and claimed it outperformed Anthropic's Mythos 5 on a cybersecurity benchmark and narrowed the coding capability gap.

  • Zhipu announced the GLM-5.3 model.
  • Zhipu said GLM-5.3 outperformed Anthropic’s Mythos 5 in a key cybersecurity benchmark.
  • Zhipu said GLM-5.3 closed the gap on coding capabilities.
  • Article date: 2026-08-16.
Aug 12, 20262 sources
GLM-5.3Z.ai
1M contextLicense not disclosed

Together AI serverless and dedicated infrastructure

Details

What we captured

Z.ai released GLM-5.3, a 1M-context coding and reasoning model, with API access through Together AI serverless and dedicated infrastructure.

  • Together AI lists the endpoint as zai-org/GLM-5.3.
  • Thinking is always enabled, with low, high, and max effort settings.
  • The model is listed with a 1M-token context length.
  • The listed prices are $1.40 per million input tokens and $440 per million output tokens.
  • Weights are described as released following a post-launch safety evaluation, but no current license is specified.
Aug 12, 20263 sources
Gemini 3.7 FlashGoogle
API and product availabilityLicense not disclosed

Available via Google AI Studio, Android Studio, Google Antigravity, Gemini Enterprise Agent Platform, the Gemini Enterprise app, and Gemini Spark for Google AI Pro and Ultra subscribers in supported countries.

Details

What we captured

Google launched Gemini 3.7 Flash for coding, agentic, software engineering, web development, and knowledge-work workflows, with availability across its API, enterprise, and consumer products.

  • Google announced the model on August 13, 2026.
  • The introductory price through the end of 2026 is $0.75 per million input tokens and $3.75 per million output tokens.
  • Google reports higher performance than Gemini 3.6 Flash across coding, web development, document reasoning, and business-workflow benchmarks.
  • The model includes updated safeguards for CBRN and cyber-offense misuse domains.
Aug 11, 20260 sources
Qwen3.8-27B-CRACK-GGUFJinho Jang (dealign.ai)
262K contextLicense not disclosed

Hugging Face repository; downloadable GGUF files and a 0.9 GB vision projector for local use

Details

What we captured

Dealign.ai published Qwen3.8-27B-CRACK-GGUF, a 27B-parameter, safety-modified derivative of Qwen3.8-27B distributed as seven GGUF quantizations and a vision projector for local llama.cpp inference.

  • Repository named Qwen3.8-27B-CRACK-GGUF published and reproduced by dealign.ai
  • Described as a 27B-parameter derivative of Qwen3.8-27B with refusal behavior removed ("abliterated")
  • Repository metadata lists creation date as 2026-08-12
  • Distributed as seven GGUF quantizations (10.5 GB to 29.0 GB) and a 0.9 GB F16 vision projector
  • Recommended local build: 17.0 GB Q4_K_M quantization
  • Packaged for local multimodal inference with llama.cpp and provides OpenAI-compatible local server commands

Source evidence

Confidence 92% / Completeness 91%

Aug 10, 20261 source
Nemotron 3.5 LightningNvidia
1M contextOpenMDW-1.1

Available through Nvidia and model repositories including Hugging Face and ModelScope; NeMo Switchyard available on GitHub

Details

What we captured

Nvidia announced Nemotron 3.5 Lightning (30B MoE, 3B active per token), released Aug. 11, with open weights under OpenMDW-1.1 and availability via Nvidia, Hugging Face, and ModelScope; paired with open-source NeMo Switchyard routing library.

  • Model: Nemotron 3.5 Lightning — 30B parameters, mixture-of-experts
  • Active parameters per token: 3B
  • Released Aug. 11 (article text)
  • Context window: up to 1,000,000 tokens
  • License: OpenMDW-1.1, open weights/training data/recipes
  • Availability: Nvidia, Hugging Face, ModelScope; NeMo Switchyard on GitHub