What happens when the good stuff stops being the expensive stuff? Wednesday answered from three directions at once. Anthropic's Claude Haiku 5.5 arrived at a tenth of the old small-model price while scoring like a much bigger model, OpenAI moved GPT-6 and a far more visual ChatGPT onto its free tier, and Microsoft made the case for running a serious coding model on your own PC with no inference bill at all. Around those headlines sat a cluster of practical releases for builders: open decision and search models from Liquid AI and Perplexity, a game maker from Google Labs, a public AI-content checker, and new homes for Haiku 5.5 on Snowflake and OpenRouter within hours of launch. Nobody claimed a new frontier, but the floor rose sharply, and that matters more to most people than another benchmark crown.
Claude Haiku 5.5 arrives as the cheapest, fastest small model Anthropic has shipped
If you run anything high-volume on Claude (summarising, classifying, routing, or acting as a sub-agent for a bigger model), your bill is about to drop sharply. Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens on prompts up to 100K tokens, a 90 per cent cut on Haiku 4.5, and Anthropic says real workloads come out around 75 per cent cheaper once its new tokenizer is accounted for. It is also a genuine step up in ability: Anthropic's own figures put it at 72.4 per cent on OSWorld computer use against 15.7 per cent for Haiku 4.5, and it is the first Haiku with an adjustable reasoning effort setting.
The catch is the 100K rule. Above that size input and output prices jump fivefold, so long-context jobs save far less, and VentureBeat points out that the lower tier exactly matches OpenAI's GPT-6 Luna, so this is a price war rather than a free lunch. It is live now in the Claude Platform, Claude Code, AWS, Google Cloud and Azure, and Anthropic still steers complex agentic coding towards Sonnet 5.5 or Opus 5.5, with Haiku as the fast helper underneath.
Claude Max and Team plans now come with monthly API credits
If you already pay for Claude Max or Team, you now have money to spend on the API every month, which removes the main excuse for not building your own tools. Max 5x plans get $100 a month, Max 20x plans get $200 and Team plans get up to $500, pooled across the workspace. The credits work on any model, Haiku 5.5 included, in your own code or in third-party tools that take an API key. It follows the Claude for Startups expansion earlier in the week, and both point the same way: Anthropic wants more people building on the platform, not just chatting in the app.
Sonnet 5.5 gets cheaper too, with cache reads cut in half
Tucked into the Haiku announcement is a change that matters to anyone running agents on Sonnet 5.5. Anthropic has halved the price of cache reads, and because agent loops re-read the same cached context again and again, it says this trims the cost of most agentic work by about 20 per cent. If you already use prompt caching, there is nothing to change: the saving simply turns up on your next invoice.
Computer use and browser use toolsets are now built into Claude's Python and TypeScript SDKs
If you have tried to build an agent that clicks around a desktop or a website, the fiddliest part was writing the loop that turns Claude's instructions into real mouse clicks and key presses. The official SDKs now run that loop for you: plug in a driver and let it go. Anthropic lists ready-made drivers from Browser Use, Browserbase, E2B and Daytona, with example drivers in its quickstarts repository if you would rather write your own. Paired with Haiku 5.5's much stronger computer-use scores, cheap browser agents just became far more practical.
GPT-6 and Intelligent UI roll out to every ChatGPT user
If you use ChatGPT every day, the answers are about to look very different. Intelligent UI builds charts, buttons, forms, maps and small working tools straight into a reply, so asking for a bill splitter, a savings calculator or an explanation of how a bicycle works gets you something to tap and adjust rather than a wall of text. Plus, Pro, Business and Enterprise users get GPT-6 Sol in the Chat tab from today, and Free and Go users get GPT-6 Luna from 8 October, which moves most of ChatGPT's enormous user base onto the new generation in a single week.
OpenAI also says GPT-6 can start answering while it is still reasoning, beginning about 44 per cent sooner than GPT-5.6 Instant on questions that need a web search. The Work and Codex models are unchanged. Nearly five weeks after GPT-6 Astra launched for paying users, this is the moment the GPT-6 family stops being a premium perk.
ChatGPT for Teens adds flashcards, quizzes and a College Planner preview
If there is a teenager in your house using ChatGPT for school, the teen version is becoming much more of a study tool. Students can turn notes or a topic into flashcards and saved decks, generate interactive quizzes from uploaded notes, and stitch several photographed pages into a single PDF on iOS. OpenAI also previewed College Planner, which gathers application requirements, deadlines and financial aid steps into one plan; it is coming soon for US students in grades 10 to 12. The launch landed on the same day Common Sense Media gave ChatGPT for Teens a poor risk rating, which OpenAI disputes, so the argument about teenagers and AI is far from settled.
Windows bets on "hybrid intelligence", with a coding model that runs on your own PC
If you code with GitHub Copilot or own a powerful Windows machine, Microsoft wants more of your AI to run locally, which means lower costs and less data leaving your desk. MAI-Code-1.1-Flash, Microsoft's own coding model, has been squeezed to 3-bit precision so it runs on the device with a 256K context window, and Microsoft AI says local model calls inside GitHub Copilot carry no inference charge. Cloud and on-device routing reaches experimental preview in the GitHub Copilot app, Copilot CLI and VS Code later in October, though you will need serious hardware such as the new Surface Laptop Ultra with up to 128GB of unified memory.
Alongside it, Copilot on Copilot+ PCs will be able to use your files, recent activity and local models (with your permission) over the coming months. Microsoft Execution Containers are now generally available on Windows 11, letting IT decide which files and networks an agent such as Codex or GitHub Copilot can touch, with Claude Code and Perplexity listed as next in line.
Google Labs launches Playground, a place to make browser games by describing them
If you have ever wanted to make a game but never learned to code, Playground lets you describe one in a chat and play it in your browser minutes later. It runs on Gemini, Nano Banana and Lyria, lets you choose 2D or 3D and single or multiplayer, and then lets you tweak the physics, rules, characters and scenery. Games can stay private, be shared by link or go into a public gallery with ratings and leaderboards tied to Play Games profiles, and Google previewed a Unity integration for more ambitious builds. It is rolling out in the US first, free to try with bigger limits for Google One members, and its launch post was the most-liked AI announcement on X all day.
SynthID Detector opens to everyone and now checks OpenAI and NVIDIA content too
If you are ever unsure whether an image, video or audio clip was AI-generated, you can now check it yourself at SynthID.com. Google has opened the detector to everyone, globally and in English, and it now recognises watermarks from partners including OpenAI, NVIDIA and Kakao as well as Google's own tools, with Apple to follow. You sign in with a Google, OpenAI or Apple account and get roughly ten checks a day. It only spots content carrying a SynthID watermark, so a clean result does not prove something is real, and Google has not published accuracy figures.
Liquid AI's Open d1 puts small decision models on the edge
If you need an app or a device to make quick calls about what it sees or hears without a round trip to the cloud, these are worth a look. Liquid AI has released d1-3B, which handles text and images, and d1-omni-600M, an experimental model for text with images or audio. Liquid says d1-3B answers a single question in under 50 milliseconds on every device it measured, and its demos include live content moderation, gesture-controlled games and a robot finding its way through a simulation on an NVIDIA Jetson. The weights are on Hugging Face with day-one llama.cpp support, making this an open, local answer to hosted services such as OpenAI's Decisions API.
Perplexity's pplx-embed-v2-late reads PDFs and slides without OCR
If you are building search over documents, decks or scanned pages, these models could save you a whole preprocessing step. Perplexity has released two open-weight late-interaction embedding models, at 0.6B and 9B parameters, that keep a small vector for each token rather than squashing a whole document into one, which helps with long or visual pages. They can embed rendered PDF pages and slides directly, skipping OCR and chunking, and the two sizes share an embedding space, so the small model can query an index built by the large one. Weights are on Hugging Face now with API access promised in the coming months; the benchmark gains are Perplexity's own, and you will need a vector database that supports multi-vector search.
Cohere opens a private beta for Compass Cloud, its hosted search and retrieval platform
If you have wanted Cohere's enterprise search and retrieval stack without running it on your own infrastructure, a hosted version is on its way. Cohere announced Compass Cloud on its own X account and is inviting teams to apply for a private beta; at the time of writing there was no blog post or press coverage, so pricing, regions and the models underneath are still to come. Worth an application if retrieval quality is the weak link in your RAG system.
Claude Haiku 5.5 lands in Snowflake Cortex AI on day one
If your company's data lives in Snowflake, you can already point Haiku 5.5 at it without moving anything. Snowflake says the model is available now in public preview on Cortex Inference, with its other AI products to follow. For high-volume jobs such as tagging, summarising or classifying rows, the new Haiku price makes this one of the cheapest ways to run Claude over warehouse data.
ElevenLabs voice models arrive on OpenRouter
If you already use OpenRouter to switch between language models, you can now add voice through the same account and key. ElevenLabs has put nine text-to-speech models covering more than 90 languages, plus two speech-to-text models with speaker labels and word-level timestamps, onto the platform, which makes it much simpler to prototype a voice agent that pairs Claude or GPT for reasoning with ElevenLabs for speech. OpenRouter also listed Claude Haiku 5.5, GPT-6 Luna Decisions and Perplexity's open-weight Decider v1.1 the same day.
Industry themes
Small models are now a price war. Haiku 5.5's lowest tier matches GPT-6 Luna to the cent, and Liquid AI's open d1 models push the same fast, cheap decision-making onto local hardware. Routing simple jobs to a small model is no longer an optimisation for specialists; it is the obvious default for anyone paying per token.
Local AI is getting serious. Microsoft running a capable coding model on the PC with no inference charge, and Liquid AI targeting edge devices, both appeal to anyone worried about privacy or tired of watching a usage meter. The hardware bar is still high, but the direction is unmistakable.
New models spread in hours, not weeks. Haiku 5.5 was on Snowflake and OpenRouter the same day it launched, while Cohere's Compass Cloud beta surfaced only on its own X account, a reminder that company channels still break news before the press does.