Step 5 Preview: Advancing the Pareto Frontier
stepfun.com
stepfun.com
> Interesting! It turns out there's already an existing project here [...] The project is fully built [...]
I'm always astounded how little effort is put into checking the AI answers displayed in these announcements. Back when I paid more attention, I remember OpenAI's and Google's demos constantly showed their AIs giving wrong answers.
> Error: OpenAI API error (403): {"message":"model water18-new is not available for user i-yuliang [trace_id=bfcfdd6bcdc236ca18d009c65cca52e4 code=40004]", "type":"invalid_request_error", "param":null, "code":null}
Here's a still frame for posterity (apologies for the quality, the original video is tiny): https://boppreh.com/room.jpg
Demos are demos and recreating scenarios is to be expected, but oh boy, somebody should review these things before publishing.
> Built on a sparse Mixture-of-Experts architecture, Step 5 Preview has 600B total parameters, with 27B active per token, and supports a 1M-token context window and vision input.
> Step 5 Preview scores 44 on the Artificial Analysis Intelligence Index.
> The model will be released with open weights on October 15.
I guess being Chinese company they decided to skip version 4, while also giving impression to be on the similar iteration with leading companies (claude opus 5).
I wonder if other Chinese labs like Kimi/Moonshot will follow suit.It’s all marketing anyways, but that’s at least how a lot of labs have been naming things (sometimes).
I agree with you though, ChronVer all the way.
Finally FireRed is being used as a benchmark again! I believe Astra can beat it in 18 hours. Not sure how that compares.
If this performs similarly in the real world, we're approaching a level of capability where for most devs, it only makes sense to pay for Anthropic or OpenAI subscriptions if they are heavily subsidized and actually cheaper than these alternative options.
* Oddly, because I perceived Devin as being kind of a joke before trying SWE-2.
Unless you meant step-3.7-flash, the input cache hits are $0.05 per mil for step-5-preview.
> $8-12 per day only for cached reads
Pretty decent "API" rates for ~500M+ tokens on Step Fun 5, a Kimi K3 / GLM 5.3 level model?
Their "Step Plan" is ridiculous, by comparison: ~$60 usage on $6.99/mo; ~$220 on $9.99/mo. https://platform.stepfun.ai/docs/en/step-plan/overview
I have used Kimi 2.5 and GLM 5.3 (& 5.3 Flash). Do not need them for what I do outside of spec hardening (basically, a lot of chatting).
I tend to know exactly what I want and most of the weaker models are enough to get me there. I have mainly been using MiMo, DeepSeek V4 Flash and MuseSpark Contributor over the last month or so.
“the question it raises matters more than the answer” ???
“So ‘the river drifts from cool to warm’ is not a figure of speech.” Oh really?