Release monitor / Verified records
AI Model Launch Tracker
Every consequential model release, structured.
A source-backed record of AI model launches, developers, availability, licenses, context windows, release types, and disclosed pricing.
Maintained dataset
Latest verified model launches signals.
Records update as sources arrive. Open any row to inspect captured facts, confidence, completeness, and supporting evidence.
Aug 20, 20262 sourcesdeepseek-v4-flash-vision-expDeepSeekexperimental API modelLicense not disclosedAvailable through Chat Completions, Messages, and Responses requests
Details
What we captured
DeepSeek announced an experimental vision model supporting Chat Completions, Messages, and Responses requests, with a 384-token billing ceiling per image.
- • The model accepts mixed text and image requests.
- • Each image is capped at 384 billable V4-Flash tokens.
- • The production V4-Flash endpoint remains text-only.
- • Peak uncached input pricing is $0.44 per million tokens and peak output pricing is $1.32 per million tokens.
Aug 20, 20261 sourceV4-Flash-Vision-ExpDeepSeekexperimental API modelLicense not disclosedDeepSeek API
Details
What we captured
DeepSeek released V4-Flash-Vision-Exp, an experimental multimodal model available through its API and compatible with OpenAI and Anthropic API formats.
- • The model adds image processing to DeepSeek-V4-Flash.
- • DeepSeek says it nearly matches Opus 4.8 on internal multimodal agent benchmarks.
- • The model supports image inputs and visual agent workflows.
- • Harness version 0.1.1 supports the model out of the box.
Source evidence
Confidence 99% / Completeness 87%
Aug 20, 20262 sourcesFalcon 2Murf AIproprietary foundation voice modelLicense not disclosedAvailability not disclosed
Details
What we captured
Murf AI announced Falcon 2, a proprietary voice foundation model positioned against leading real-time voice models on naturalness, latency, and cost.
- • Murf AI announced Falcon 2 on August 21, 2026.
- • The article reports that Falcon 2 ranked ahead of ElevenLabs Turbo v2.5, ElevenLabs Flash v2.5, and OpenAI Realtime on Artificial Analysis’s Speech Arena.
- • Murf states that Falcon 2 provides sub-100-millisecond time-to-first-audio and costs $0.01 per generated minute.
- • The model supports 150+ voices across 35+ languages and up to 10,000 concurrent calls.
Aug 18, 20260 sourcesOrnith-1.5Ornith AIopen weightsLicense not disclosedHugging Face collection; local deployment and quantized mobile version described
Details
What we captured
Ornith AI released Ornith-1.5, a three-size open model family designed to generate its own training tasks, scaffolds, and solution rollouts.
- • The release date is August 19, 2026.
- • The family contains 397B MoE, 35B MoE with 3B active parameters per token, and 9B dense variants.
- • Ornith AI says the models generate harder tasks, task-specific scaffolds, and solution rollouts within a reinforcement-learning loop.
- • Weights, serving instructions, and quantized formats were published for developer testing.
Source evidence
Confidence 100% / Completeness 87%
Aug 17, 20263 sourcesGPT-5.6 SolOpenAIAPILicense not disclosedLimited release to a small group of customers; wider access planned later
Details
What we captured
OpenAI announced a limited API release of an Ultrafast mode for GPT-5.6 Sol, powered by Cerebras and capable of up to 750 output tokens per second.
- • OpenAI is testing an Ultrafast mode for GPT-5.6 Sol.
- • The mode is available to a small group of customers through the OpenAI API.
- • The mode is powered by Cerebras and can produce up to 750 output tokens per second.
Source evidence
Confidence 94% / Completeness 87%
Aug 15, 20263 sourcesGLM-5.3Zhipu (Z.ai)Model releaseLicense not disclosedAvailability not disclosed
Details
What we captured
Zhipu announced GLM-5.3 and claimed it outperformed Anthropic's Mythos 5 on a cybersecurity benchmark and narrowed the coding capability gap.
- • Zhipu announced the GLM-5.3 model.
- • Zhipu said GLM-5.3 outperformed Anthropic’s Mythos 5 in a key cybersecurity benchmark.
- • Zhipu said GLM-5.3 closed the gap on coding capabilities.
- • Article date: 2026-08-16.
Aug 12, 20262 sourcesGLM-5.3Z.ai1M contextLicense not disclosedTogether AI serverless and dedicated infrastructure
Details
What we captured
Z.ai released GLM-5.3, a 1M-context coding and reasoning model, with API access through Together AI serverless and dedicated infrastructure.
- • Together AI lists the endpoint as zai-org/GLM-5.3.
- • Thinking is always enabled, with low, high, and max effort settings.
- • The model is listed with a 1M-token context length.
- • The listed prices are $1.40 per million input tokens and $440 per million output tokens.
- • Weights are described as released following a post-launch safety evaluation, but no current license is specified.
Source evidence
Confidence 99% / Completeness 96%
Aug 12, 20263 sourcesGemini 3.7 FlashGoogleAPI and product availabilityLicense not disclosedAvailable via Google AI Studio, Android Studio, Google Antigravity, Gemini Enterprise Agent Platform, the Gemini Enterprise app, and Gemini Spark for Google AI Pro and Ultra subscribers in supported countries.
Details
What we captured
Google launched Gemini 3.7 Flash for coding, agentic, software engineering, web development, and knowledge-work workflows, with availability across its API, enterprise, and consumer products.
- • Google announced the model on August 13, 2026.
- • The introductory price through the end of 2026 is $0.75 per million input tokens and $3.75 per million output tokens.
- • Google reports higher performance than Gemini 3.6 Flash across coding, web development, document reasoning, and business-workflow benchmarks.
- • The model includes updated safeguards for CBRN and cyber-offense misuse domains.
Aug 11, 20260 sourcesQwen3.8-27B-CRACK-GGUFJinho Jang (dealign.ai)262K contextLicense not disclosedHugging Face repository; downloadable GGUF files and a 0.9 GB vision projector for local use
Details
What we captured
Dealign.ai published Qwen3.8-27B-CRACK-GGUF, a 27B-parameter, safety-modified derivative of Qwen3.8-27B distributed as seven GGUF quantizations and a vision projector for local llama.cpp inference.
- • Repository named Qwen3.8-27B-CRACK-GGUF published and reproduced by dealign.ai
- • Described as a 27B-parameter derivative of Qwen3.8-27B with refusal behavior removed ("abliterated")
- • Repository metadata lists creation date as 2026-08-12
- • Distributed as seven GGUF quantizations (10.5 GB to 29.0 GB) and a 0.9 GB F16 vision projector
- • Recommended local build: 17.0 GB Q4_K_M quantization
- • Packaged for local multimodal inference with llama.cpp and provides OpenAI-compatible local server commands
Source evidence
Confidence 92% / Completeness 91%
Aug 10, 20261 sourceNemotron 3.5 LightningNvidia1M contextOpenMDW-1.1Available through Nvidia and model repositories including Hugging Face and ModelScope; NeMo Switchyard available on GitHub
Details
What we captured
Nvidia announced Nemotron 3.5 Lightning (30B MoE, 3B active per token), released Aug. 11, with open weights under OpenMDW-1.1 and availability via Nvidia, Hugging Face, and ModelScope; paired with open-source NeMo Switchyard routing library.
- • Model: Nemotron 3.5 Lightning — 30B parameters, mixture-of-experts
- • Active parameters per token: 3B
- • Released Aug. 11 (article text)
- • Context window: up to 1,000,000 tokens
- • License: OpenMDW-1.1, open weights/training data/recipes
- • Availability: Nvidia, Hugging Face, ModelScope; NeMo Switchyard on GitHub
Source evidence
Confidence 95% / Completeness 95%
