Harmony: OpenAI's response format for its open-weight model series
github.com
github.com
My initial thought was how hacky the whole thing feels, but then the fact that it works and gives rise to complex behaviour (like coercing specific tool selection in the Manus post) is quite simple and elegant.
Also as an aside, it is good that it appears that each standard tag is a single token in the OpenAI repo.
[1] https://manus.im/blog/Context-Engineering-for-AI-Agents-Less... [2] https://github.com/NousResearch/Hermes-Function-Calling
I have a branch of llm-consortium where I was noodling with giving each member model a role. Only problem is it's expensive to evaluate these ideas so I put it on hold. But maybe now with oss models being cheap I can try and it on those.
Pardon me but are you thinking that this method is superior than mixture of experts? What are your thoughts?
MOEs are a single model. An 'expert' is a subset of layers chosen by a router model for each token. This makes them run faster. A consortium is a type of parallel reasoning that uses multiple of the same or different models to generate parallel response and find the best one.
All models have a jagged frontier with weird skill gaps. A consortium can bridge those gaps and increase performance on the frontier.
I'd love to see how they performed.
I wish someone would extract the Grok Heavy prompts to confirm, but I guess those jailbreakers don't have the $200 sub.
I assume they are using the concept of harmony to refer to the consistent response format? Or is it their intention for an open weights release?
That's pretty cool and seems like a logical next step to structure AI outputs. We started out with a stream of plaintext. In the future perhaps we'll have complex typed output.
Humans also emit many channels of information simutaneously. Our speech, tone of voice, body language, our appearance - it all has an impact on how our information is received by another.
- https://openai.com/index/introducing-gpt-oss/
- https://cdn.openai.com/pdf/419b6906-9da6-406c-a19d-1bb078ac7...
This kind of coordination failure is surprisingly common with AI releases lately. Remember when everyone was trying to access GPT-4 on launch day? Or when Anthropic's Claude had those random outages during their big announcements?
Makes you wonder if they're rushing to counter Google's Genie 3 news and got caught with their pants down during the GitHub outage. The timing seems too coincidental.
At least when it does go live, having truly open weights models will be huge for the community. Just wish they'd test their deployment pipeline before hitting 'publish' on the blog post.
https://www.bleepingcomputer.com/news/artificial-intelligenc...
... but these links aren't active yet. I presume they will be imminently, and I guess that means that OpenAI are releasing an open weights GPT model today?
Also creates a walled garden on purpose.
- https://gpt-oss.com/ Auth required?
- https://openai.com/open-models/ seems empty?
- https://cookbook.openai.com/topic/gpt-oss 404
- https://openai.com/index/gpt-oss-model-card/ empty page?
Am I holding the internet wrong?
> GPT OSS is a hugely anticipated open-weights release by OpenAI, designed for powerful reasoning, agentic tasks, and versatile developer use cases. It comprises two models: a big one with 117B parameters (gpt-oss-120b), and a smaller one with 21B parameters (gpt-oss-20b). Both are mixture-of-experts (MoEs) and use a 4-bit quantization scheme (MXFP4), enabling fast inference (thanks to fewer active parameters, see details below) while keeping resource usage low. The large model fits on a single H100 GPU, while the small one runs within 16GB of memory and is perfect for consumer hardware and on-device applications.
Not a fan of this presentation of communication.
we have a lot of new stuff for you over the next few days!
something big-but-small today.
and then a big upgrade later this week.
EDIT: nevermind, I spoke too soon! I guess this was referring to GPT 5 later this week. https://openai.com/open-models/ is liveLikely, considering every single one opens right up for me.
OpenAI might have tried coordinating the press release of their open model to counter Google Genie 3 news but got stuck in the middle of the outage.
The GitHub outage is delaying them on their release.