Step 3.7 Flash goes live on DeepInfra and Novita at budget pricing
StepFun's Step 3.7 Flash, a 198B-parameter mixture-of-experts vision-language model (approximately 11B active parameters, 256K context), is now served by both DeepInfra and Novita at roughly $0.20 per million input tokens and $1.15 per million output tokens. That undercuts most frontier coding APIs, and it matters because Step 3.7 Flash is designed for coding agents and search-heavy workflows where token volume runs high.
The model itself launched on 29 May, so what is new in this window is broad, cheap third-party serving rather than the model itself. For anyone routing agentic or vision tasks through OpenRouter or a multi-provider gateway, this adds two more competitively priced fallbacks. It is also the clearest example yet of how open-weight coding momentum is translating into access economics: the interesting competition right now is not between new capabilities but between providers racing to serve the same capable models at the lowest price.
OpenAI's fresh posts in the window were research rather than product: LifeSciBench, a new scientific benchmark, and a write-up of a near-autonomous AI chemist improving a reaction. Neither is something users can act on today, but both point to a sustained push into scientific applications. OpenAI News, 17 Jun →
Anthropic announced the opening of a Seoul office and new Korean AI partnerships. Corporate positioning rather than product news, but it signals continued expansion into regulated markets following the DXC alliance earlier this month. Anthropic News, 17 Jun →
Today's themes
The day's single qualifying story says something important about where the action is for everyday builders right now. Open-weight coding momentum continues to drive the conversation, with Z.ai's GLM-5.2 (released 13 June, MIT-licensed, and topping frontend-coding leaderboards) still pulling third-party providers to compete on price. Step 3.7 Flash's endpoint expansion is the direct result: access economics, not raw capability, is the active front. Cheap, multi-provider serving of capable open models is becoming a feature in its own right, and it narrows the gap between what smaller teams and large enterprises can afford to run in production. The major US labs, meanwhile, spent the day on research and corporate announcements rather than shipping, a reminder that not every window brings a usable launch.