Just shipped

New launches.

What just shipped. We read the vendor's own announcement before anything lands on this page, so nothing here is a rumour.

Updated Sep 24, 2026

New in this update: Gemini 3.8 Flash TTS, GPT-6 Sol and GPT-6 Luna, Claude Opus 5.5. Listed below by launch date, so they may not be at the top.

The briefing

Get this in your inbox.

Email me the AI TIP intelligence briefing when it starts: verified launches, pricing moves and security events, twice a week. No vendor marketing. Unsubscribe any time.

Your address is used for this briefing and nothing else — we don't sell or share it. See our Privacy Policy.

Latest launches

Gemini 3.8 Flash TTS

Google DeepMindNewVoice modelsVerifiedSep 23, 2026

Promptable voice design with 2,000+ voices and 30-second consented cloning.

Gemini 3.8 Flash TTS and the cheaper Flash-Lite tier turn a text prompt into a designed voice, or clone a real one from a 30-second consented sample, across 100+ languages and hours-long two-speaker scenes. Every generated or cloned voice carries SynthID and C2PA provenance markers. Flash ranked #1 on the Hume Voice Design benchmark; Flash-Lite ranked #2 on Hume Overall Quality despite being the budget tier.

GPT-6 Sol and GPT-6 Luna

OpenAINewLLMs & AssistantsVerifiedSep 22, 2026

Two cheaper GPT-6 tiers below Astra, priced at half the outgoing 5.6 line.

Sol adds more reasoning headroom for everyday work; Luna is the lightweight, high-volume option built for fast responses at the lowest cost in the GPT-6 family. Both carry Astra's gains in factuality, coding, computer use and alignment down to a cheaper tier, and both are priced at half of what the equivalent GPT-5.6 models cost — Sol at $2/$10 per million tokens versus $4/$20 before. It's the second GPT-6 release in three weeks, after Astra's September 3 debut.

Claude Opus 5.5

AnthropicNew★ HeadlineLLMs & AssistantsVerifiedSep 22, 2026

Fable-level performance at 40% lower cost than Opus 5.

Opus 5.5 is the first release in Anthropic's 5.5 family: it matches Claude Fable 5.1 on most work while costing 40% less to run than Opus 5, with output speeds up over 30%. Pricing dropped 20% to $4/$20 per million tokens and cache reads fell 60% to $0.20/million. It's live on AWS, Google Cloud and Microsoft Azure now, with Sonnet 5.5 and Haiku 5.5 expected within weeks — and it lands the same week OpenAI cut GPT-6 Sol and Luna pricing in half, making September a price-war month across all three labs.

Gemini 3.8 Live

Google DeepMindVoice modelsVerifiedSep 15, 2026

A realtime voice-to-voice model at under half OpenAI's asking price.

Gemini 3.8 Live holds a live spoken conversation, sees what your camera sees and can call tools and APIs mid-sentence while switching among 97 languages. It runs about $0.84/hour at standard settings ($3.50/hour with Extended Thinking) against $5.83/hour for OpenAI's GPT-Live-1 Astra — Google topping its closest rival on benchmark scores while charging less than half as much for the API.

iOS 27 with the new Siri

Apple★ HeadlinePersonal AI agentsVerifiedSep 14, 2026

Apple's Gemini-powered Siri rebuild ships for US and UK users.

iOS 27 shipped a rebuilt Siri that reads context across your calendar, messages and email and acts on multi-step requests, using Apple's Private Cloud Compute to de-identify queries before routing the hard ones to a customized Gemini model Apple is paying Google roughly $1 billion a year to license. It landed for US and UK users with compatible hardware on September 14; EU iPhones don't get it yet, held up by a Digital Markets Act dispute, and older devices get a scaled-back version.

Agentforce named agents

SalesforceEnterprise AI agentsVerifiedSep 11, 2026

Seven purpose-built agents — Casey, Paige, Carter and four more — ship inside Customer 360.

Salesforce introduced seven named Agentforce agents, each scoped to one business function (support, sales coaching, workforce scheduling and more) and grounded in a company's existing Customer 360 data and business rules rather than a single general-purpose bot. Early customers report billions of agentic work units delivered, pointing at a broader shift from one do-everything agent toward a roster of narrow, auditable ones.

Fugu Max & Fugu Ultra v2.0

Sakana AIOpen-weight modelsVerifiedSep 11, 2026

A multi-agent orchestration system sold as a single model.

Rather than one large model, Fugu assembles and orchestrates a swappable pool of expert agents on the fly, learning how to route a task instead of following a hand-designed workflow. Fugu Ultra v2 posts the top or tied-top score on 5 of 8 benchmarks against frontier models it wasn't even trained alongside (its data cutoff predates Fable 5.1 and GPT-6 Astra), at $5 input / $30 output per million tokens; the smaller Fugu Max runs $2/$6. Both are available now through an OpenAI-compatible API.

Agents API

OpenAIAgent frameworksVerifiedSep 10, 2026

The Codex agent harness, rented by the API call.

The Agents API is OpenAI's own Codex infrastructure — session orchestration, context compaction over long tasks, sub-agent coordination, lazy tool loading and crash recovery — opened to any developer through four primitives: agent, environment, session, event. It's live in public beta for all developers, with sandbox compute available from OpenAI or partners including Cloudflare, Vercel and Oracle, and no fee beyond the tokens and tools an agent actually uses. It competes directly with the orchestration layer LangChain, CrewAI and similar frameworks occupy: skip building an agent runtime and rent OpenAI's instead.

DeepSeek V4.1 Flash

DeepSeek AIOpen-weight modelsVerifiedSep 10, 2026

MIT-licensed 552B MoE model at $0.15 per million input tokens off-peak.

V4.1 Flash is a 552B-parameter mixture-of-experts model on a new causal encoder-decoder architecture, with native image understanding and a 1M-token context window. Weights are MIT-licensed on Hugging Face, and off-peak pricing of $0.15 input / $0.60 output per million tokens undercuts the V4 Flash tier it replaces by half; peak-hour traffic (01:00-04:00 and 06:00-10:00 UTC weekdays) runs $0.30 / $1.20. From September 14 it also absorbs all V4 Pro traffic at Flash pricing until a dedicated V4.1 Pro ships.

Muse

Meta★ HeadlinePersonal AI agentsVerifiedSep 8, 2026

Meta's agent for booking, emailing and long-term goals, not just chat.

Muse is a personal AI agent, separate from the Meta AI chatbot, that takes a goal — book this trip, plan a year of workouts, set up a business — and works it over time: opening a browser, filling out forms, sending the email, negotiating on your behalf. It's available now in the US through its own app, inside WhatsApp, and on the web, with a free tier good for 100 million tokens a week and paid Power ($20/month) and Maximum ($100/month) tiers for heavier use. Each session runs in an isolated virtual machine; Meta says a Confidential VM mode encrypted end-to-end, unreadable even to Meta, ships later in 2026.

GPT-Live

OpenAIVoice modelsVerifiedSep 8, 2026

A native voice model behind ChatGPT Voice, no text pipeline in between.

GPT-Live replaces the speech-to-text-to-speech pipeline behind ChatGPT Voice with a model that hears and speaks directly, cutting response latency below 300ms and carrying tone and emotional nuance that a text intermediary strips out. It's rolling out now as the engine for ChatGPT's voice mode.

GPT-6 Astra

OpenAI★ HeadlineFrontier modelsVerifiedSep 3, 2026

OpenAI's new flagship, built around computer use.

GPT-6 Astra is OpenAI's most capable model yet, with the biggest gains in computer use — navigating a screen and operating software the way a person would — alongside coding, scientific reasoning, cybersecurity and long-context retrieval. It ships first to a limited set of organizations, then to ChatGPT Plus, Pro, Business and Enterprise and the API, also available on Azure and Bedrock. API pricing is $10 per million input tokens and $50 per million output tokens, with a Fast mode at double the speed and double the price. It's also the model OpenAI paused in July after it crossed the 'Critical' cybersecurity threshold in the company's own Preparedness Framework; advanced cyber capability stays restricted to vetted defenders through the Daybreak program, alongside a new $1B commitment to subsidize that access for frontline defenders.

Grok Bot for Enterprise

xAIAgentic assistantsVerifiedSep 3, 2026

Autonomous AI workers for the enterprise, with access controls.

xAI opened Grok Bot, its autonomous computer-using agent, to enterprise customers with access, network and audit controls it lacked in its August consumer beta. Each bot runs in its own isolated cloud environment, uses a browser and applications the way a person would, has no account access by default, and can repeat a workflow after a person demonstrates and corrects it once. Grok and Cursor Enterprise customers get two weeks free and can invite their whole organization to try it.

K2 Horizon

IFMOpen-weight modelsVerifiedSep 3, 2026

Six fully open models, weights to training data, from 0.9B to 375B.

The Institute of Foundation Models released K2 Horizon: six Apache 2.0 models (0.9B, 3.7B, 7B, 32B, 36B-A4B and 375B-A23B) sharing one architecture, vocabulary, training method and eval infrastructure, so results are comparable across sizes. IFM is publishing model weights and code plus training data, intermediate checkpoints, configs and logs where licenses allow — a rare level of openness among fully open releases. Day-zero support covers vLLM, SGLang and Ollama across NVIDIA, AMD and Cerebras hardware, with an intended deployment range from wearables (0.9B) up through enterprise serving (375B-A23B).

Gemini 3.8 Flash and Gemini 3.8 Flash Cyber

Google DeepMindFrontier modelsVerifiedSep 2, 2026

Google's third Flash release in six weeks, plus a dedicated cybersecurity variant.

Gemini 3.8 Flash is Google's latest workhorse model, improving on 3.7 Flash (released three weeks earlier) across software engineering, agentic tasks and multi-step reasoning in specialized domains; it posts 54.9% on HLE-Verified and leads on benchmarks like Vals Finance Agent V2 and Harvey's Legal Agent Benchmark. Gemini 3.8 Flash Cyber is a separate variant tuned specifically for vulnerability detection, which Google positions as frontier-level for cybersecurity work. Both are generally available now through the Gemini API, AI Studio, Antigravity, Android Studio and Gemini Enterprise, at the same introductory price as 3.7 Flash: $0.75 per million input tokens, $3.75 per million output tokens. The rapid Flash cadence — three releases in six weeks — is Google keeping pace with Anthropic and OpenAI's release tempo rather than a single headline jump in capability.

Claude Fable 5.1 and Claude Mythos 5.1

Anthropic★ HeadlineFrontier modelsVerifiedSep 1, 2026

Fable 5 gets a cheaper, more capable refresh; Mythos 5.1 loosens safeguards for vetted cybersecurity and life-sciences work.

Fable 5.1 and Mythos 5.1 are the same underlying model shipped with different safeguard levels, three months after Fable 5's June debut. Fable 5.1 is generally available now, built for jobs that run for hours across many tools: working through a Cowork backlog, operating a browser, or running unattended as a managed agent. Anthropic says it matches or beats Fable 5 at low and medium effort and is noticeably stronger at high effort, with Claude Code users seeing roughly 60% fewer cybersecurity false positives. The commercial story is the 75% cut to cache-read pricing, from $1.00 to $0.25 per million tokens, which Anthropic estimates lowers typical bills by 25% and up to 45% for heavily agentic workloads. Mythos 5.1 carries more permissive safeguards and stays restricted to vetted organizations in Anthropic's trusted-access program for cybersecurity and life-sciences work, with US-government coordination underway to widen access.

Solaris

Runway AIInterface generationVerifiedAug 31, 2026

Runway's first 'Interface World Model' renders app interfaces frame by frame with no code underneath.

Solaris generates a working software interface as live video rather than running code: it renders frames at interactive speeds (under 500ms) at 720p and responds to clicks, drags and voice commands, with nothing resembling a DOM behind it. Runway calls it the first entry in a new 'Interface World Models' category — a research direction, not a shipping product. It isn't public: Runway is collecting early-access requests while it works with select partners, and has disclosed no price or launch date. Text rendering is still error-prone and there's no support for screen readers, so the system is incompatible with accessibility requirements most production software has to meet. It's a signal of where Runway is pointing its video-generation research next, not something a team can use today.

GLM-5.3-Flash

Z.ai (Zhipu)Frontier modelsVerifiedAug 26, 2026

Open-weight multimodal MoE model at a tenth of GLM-5.2's price.

A 320B-parameter, 18B-active mixture-of-experts model that's natively multimodal across text, image and video, with a 1M-token context window and MIT-licensed open weights. Z.ai says it beats GLM-5.2 on its own evaluations for coding, agentic and long-context work while running at roughly one-tenth the cost — list price is $0.15/$0.50 per million tokens, discounted 50% through September 9. It's a serious open-weight option for teams building agentic coding tools who don't want to pay frontier-model rates.

Jalapeño

OpenAI★ HeadlineAI infrastructureVerifiedAug 25, 2026

OpenAI's first in-house inference chip, co-developed with Broadcom and Celestica, claims a work-per-watt win over Nvidia's Blackwell systems.

Jalapeño is an inference-only accelerator — it serves trained models, it does not train them — built under OpenAI's October 2025 deal with Broadcom to co-develop 10 gigawatts of custom silicon. Shown publicly at Hot Chips, OpenAI's own benchmarks claim 1.5x to 1.9x more work per watt at peak throughput than the best commercially available systems across three tested models, with 1.7x to 3.6x lower end-to-end latency and 2.1x to 4.1x higher throughput on interactive workloads. SemiAnalysis notes the fairer comparison is Nvidia's newer Vera Rubin platform rather than Blackwell, and that Jalapeño still comes out ahead there on tokens per megawatt even without the multi-token prediction optimization Vera Rubin uses. It deploys inside OpenAI's own infrastructure by year-end; there is no external product or API.

Gemini Enterprise for Legal

Google CloudEnterprise AIVerifiedAug 25, 2026

A law-firm-specific edition of Gemini Enterprise with pre-built agents for contract review, diligence and regulatory monitoring.

Gemini Enterprise for Legal packages Google's Gemini Enterprise Agent Platform with legal-specific skills and connectors into the document management, e-discovery and legal research systems firms already run, aimed at executing full workflows — contract review, diligence, regulatory monitoring, privacy requests — rather than just answering questions about a document. It launched in preview on August 25 with Cleary Gottlieb, Freshfields, Weil and Williams & Connolly as named firms, and is Google's first industry-specific packaging of Gemini Enterprise; a financial-services edition shipped alongside it, with healthcare versions described as coming. For firms already inside Google Workspace, this is a narrower, more opinionated product than pointing a general assistant at a contract.

ChatGPT for Teens

OpenAIFrontier modelsVerifiedAug 18, 2026

A separate ChatGPT experience for under-18 accounts, with age prediction, parental controls and a rewritten set of behavior rules.

ChatGPT for Teens auto-enrolls accounts OpenAI predicts belong to minors into a distinct experience: default-on restrictions around self-harm, romantic/sexual and relational content (the model can no longer call itself a friend or imply it has feelings), a Study Mode that walks through problems instead of answering them outright, and parent-set quiet hours. Adults wrongly flagged can confirm their age through Persona identity verification. The rollout started August 18 across Free and paid personal plans and is expected to finish within two weeks; enterprise and API products are unaffected. It follows wrongful-death lawsuits over chatbot safety and an FTC inquiry into how AI chatbots affect minors.

GLM-5.3

Z.aiFrontier modelsVerifiedAug 14, 2026

An open-weight coding model that jumps sharply on long-horizon tasks without retraining its base.

GLM-5.3 reuses the same 743B base model as GLM-5.2 — every reported gain comes from scaled-up post-training rather than a new base. The jump is steepest on long-horizon coding: Terminal-Bench 3.0 moves from 4.6 to 28.3 and DeepSWE v1.1 from 46.2 to 66.9. Cybersecurity capability grew faster than Z.ai says it expected as training scaled, reaching 84.5% on CyberGym, which is why the company is staging the release: API access and the GLM Coding Plan are live now, but open weights are held back roughly two weeks for safety hardening.

Qwen3.8-27B

Alibaba CloudFrontier modelsVerifiedAug 14, 2026

A 28B open-weight model distilled from Qwen3.8-Max, sized to run on a single high-end GPU or an AI PC.

Qwen3.8-27B is the open-weight, locally runnable sibling to Alibaba's frontier-scale Qwen3.8-Max, released under Apache 2.0 with a native 262K-token context window and text, image and video input. It distills the larger model's agentic coding, computer-use and vision-language gains into a dense 27.78B-parameter form that fits on high-end consumer GPUs or Ryzen AI Max hardware, which is the actual pitch: frontier-family capability without the API bill or the data-residency question. Weights are live on Hugging Face and ModelScope.

Gemini 3.7 Flash

GoogleFrontier modelsVerifiedAug 13, 2026

Google's fast workhorse model for coding and agents, replacing 3.6 Flash three weeks after it shipped.

Gemini 3.7 Flash is a full replacement of Gemini 3.6 Flash built from the same base via algorithmic improvements and user feedback rather than retraining from scratch, aimed at coding, web development and agent workflows. Google's own numbers put it at 65.3% on DeepSWE v1.1 versus 49.0% for 3.6 Flash. It ships at an introductory $0.75/$3.75 per million input/output tokens — half the prior launch price, good through the end of 2026 — and lands across AI Studio, Android Studio, Antigravity and the Gemini Enterprise Agent Platform, notably ahead of the still-delayed flagship Gemini 3.5 Pro.

Grok 4.6

xAIFrontier modelsVerifiedAug 12, 2026

A post-training upgrade on Grok 4.5 tuned for long-running agents, matching GPT-5.6 Sol at the same price.

Grok 4.6 keeps Grok 4.5's 500K context window and adds a new xhigh reasoning-effort level aimed at sustaining multi-step work: researching, analyzing, working across codebases and building applications without losing the thread. On the Artificial Analysis Intelligence Index it scores 61, matching GPT-5.6 Sol at max reasoning effort and trailing Fable 5 Max by one point, at $2/$6 per million tokens (doubling above 200K-token prompts). It shipped simultaneously in Cursor, Grok Build, the xAI API and GitHub Copilot, with OpenRouter, Vercel and Cloudflare also carrying it.

GPT-5.6-Cyber

OpenAIFrontier modelsVerifiedAug 10, 2026

A cybersecurity-specialized GPT-5.6 variant, released only to vetted defenders through Daybreak Red.

GPT-5.6-Cyber is built on GPT-5.6 Sol and trained specifically for finding zero-days and building exploit chains, scoring 95% on OpenAI's Advanced Cybersecurity Completion Rate versus 1.5% for standard Sol. OpenAI used it internally to find two previously unknown Chrome V8 vulnerabilities, now patched as CVE-2026-15903. It ships only through Daybreak Red, OpenAI's applicant-vetted defender tier — identity verification, account security requirements, monitoring and approved-use attestations gate access, with Accenture, IBM, CrowdStrike and Cloudflare among the first trusted partners. There is no general-availability path; this is a controlled-access release, not a product launch in the usual sense.

GPT-5.6 Sol update + free unlimited GPT-5.6 Luna

OpenAIFrontier modelsVerifiedAug 6, 2026

Plus/Pro get a more accurate Sol; free users get unlimited Luna chats and a Think button.

OpenAI updated GPT-5.6 Sol for Plus and Pro users toward more focused answers and less reflexive agreement, reporting 68% fewer factual errors than GPT-5.5 Instant in its own evaluation. Free users move to GPT-5.6 Luna as the default with unlimited text chats, plus a Think button for harder questions. The rollout covers ChatGPT's web, mobile and desktop apps; ChatGPT Work and Codex keep their current Sol version, unchanged.

Muse Code + Muse Spark 1.2

Meta★ HeadlineCoding agentsVerifiedAug 5, 2026

Meta's first coding agent, and a contributor tier that is 12x cheaper if you let them train on your work.

Muse Code is a terminal coding agent that plans, writes and validates changes across large repositories, coordinating persistent background subagents that stay alive for a whole session rather than spawning per task, with a local append-only log of every model call, tool run, approval and edit. It runs on Muse Spark 1.2, co-trained with the harness itself. Standard API pricing is $1.25/$4.25 per million input/output tokens; a separate contributor tier drops that to $0.10/$0.20 in exchange for permission to train future Meta models on your prompts and completions. Beta, macOS and Linux.

Qwen3.8-Max

Alibaba Cloud★ HeadlineFrontier modelsVerifiedAug 3, 2026

Alibaba's largest model yet: a 2.4T-parameter MoE priced to match US closed models, with open weights due within the week.

Qwen3.8-Max is a 2.4-trillion-parameter mixture-of-experts model (95B active per token) with native text, image and video input and a 1M-token context window. Alibaba priced DashScope API access at $2/$6 per million input/output tokens, in the same range as Sol and Opus 5, and confirmed full weights follow within the week — the first time a model at this parameter scale has shipped open rather than staying API-only. It's live now through Qwen Chat and the DashScope API; self-hosting waits on the weights drop.

DeepSeek-V4-Flash-0731

DeepSeekFrontier modelsVerifiedJul 31, 2026

DeepSeek's production Flash refresh — same 284B-parameter MoE, retrained for a big jump in agentic coding and tool-use.

DeepSeek-V4-Flash-0731 keeps the April preview's architecture (284B total parameters, 13B active, 1M-token context, 384K max output) and re-post-trains it for agentic workflows, landing coding and tool-use scores close to Claude Opus 4.8 at a fraction of the price. A new reasoning_effort parameter (low/high/max) lets callers trade latency for deliberation per request.

Ling-3.0-Flash

Ant GroupFrontier modelsVerifiedJul 27, 2026

Ant Group's hybrid-reasoning model matches rivals 2-3x its size, built specifically for production agent workflows.

Ling-3.0-Flash is a 124B-parameter MoE model (5.1B active) tuned for multi-agent collaboration — different agents divide labor and cross-check each other's output to cut single-model misjudgments. A cluster-level caching system cuts time-to-first-token on long inputs by 60-80%. It's free via OpenRouter and Vercel AI Gateway through 3 August, with full weights open-sourced after that window closes.

Only launches we've verified against the vendor's own announcement make this page. For the wider news tape — pricing moves, funding, research and security — see the Intelligence feed.