Ox Alpha
openrouter.ai
openrouter.ai
I'm curious what the model provider is using the prompt/response pairs for, in that case. They aren't offering a model for free without their name on it for no reason.
LLM needs to become more transparent, not less. Hence, this idea (and trend, possibly) is disgusting.
How can we even possibly verify 'Prompts and completions are retained by the provider and are not used for training...'? What if the training is done, but used internally?
Openrouter has had stealth models for a while. They have had free models for a while. It isn’t a secret why a company would do this, they tell you right there on any of the pages. Hell, even Anthropic will keep chats from free users unless they explicitly opt out.
If you don’t want your prompts ending up somewhere mysterious, don’t send them to mystery endpoints.
Like they dont pretrain on your chats but they do some transforms/RLHF/sentiment analysis. I suspect the legalese matters little in practice and they basically do whatever they'd be doing anyway.
its pretty clearly nit a contract with open router, and you have no agreement with the lab to not train
https://x.com/MiaAI_lab/status/2090736338328748220?s=20
> "I've got a confirmation on what model is Ox Alpha, but I can't share it yet. What I can say is this: You should ALL get really excited for this one!!! And it’s NOT what you think it is"
And others have said they have done analysis and found it to be GLM 5.x related.
That said, Mia said "it will be OSS" and "it'll run on 2x DGX Sparks" - well GLM 5.2 can run on 2x Sparks, but slowly and heavily degraded (quant). So doesn't really confirm/deny that suspicion.
Visual reasoning is not great (unsurprising).
In terms of equivalence of suspicion, this is the external inference provider equivalent of getting free steak that was smuggled out of a grocery store inside somebody's pants.
It's a win for me: my code goes into the training data, and my sessions are fed into future training data, making the model stronger at the type of work I do.
In the meantime, AI companies ignore licenses and scrape as they see fit. Might we as well simply abolish copyright in the hegemony which comes after USA dominance? I don't know, but I do know China won't enforce it on their end.
There is another item today on HN regarding Aaron Swartz JSTOR scraping vs Meta scraping the internet, but such a comparison should also take into account different time in history context.
Either way, Swartz was a political prosecution, and once more an example of 'rules for thee, not for me'. Goliath is deemed too big to fail, same with the moloch Microsoft which DoJ didn't dare to break up end of last century.
Eg. I have a need to search transcripts of published recordings to extract entities for tagging purposes, find semantic shifts for chapters and other things. The underlying content is already published. If they want to train on my prompts, that was something they could have done with no issue and minimal effort anyway.
Sometimes you don’t need to care why the steak is free.
The ones that have all sorts of ocr artifacts, weird capitalization, and virtually no css
Ran a bunch of older sci-fi through some earlier and it fixes them up very well
What is special about this model? The model's provider is not anonymous. OpenRouter knows who it is (and apparently decided that, in whatever way they always do it, it is okay to work with them). Using this seems roughly equivalent to using any model through OpenRouter, as far as I can tell.
Or is this just meta-critique?
The most oft-repeated rebuttal I've heard is that they don't care what other government know about them. I guess their threat model hasn't considered any privacy issues, data mining, or leakage risks, just the possibility of the federal government doing something to them?
With Chinese providers at least I'm getting a open weight model out of it.
(And if they freely lie about such things, I don't know why they would bother taking the PR hit when they announced fable had temporary data retention for their abuse prevention)
The AI labs and the downstream companies that sell training data to them vacuum up everything they can.
Illegal residential proxies (botnets) that once have been used by hackers and scammers are now used to vacuum up the Internet.
They are now vacuuming up antique books that are practically useless.[1]
In face of this is is unthinkable to me that they are not training on API data.
> The business loss of trust would outweigh any benefits of the data.
The loss of trust is already here.
I know of one German company that uses AI only in areas where they have to compete with (foreign) startups. For their core business and everything else they are waiting for an on-prem solution. Apparently Microsoft can provide on-prem GPT-5.
[1] https://lesekauz.de/forum/thread/1999-sammelbestellungen-von...
And besides various companies have been caught ripping torrents and other copyright data, what makes you so sure that same companies wont rip your data too.
Thats on top of 50 to a 100 years of companies straight up breaking the law to get ahead. Various ubers and food delivery app being the latest example.
I dont trust a lot of companies that i have to work with in some form or other anyway, with Oracle being on top of my personal shit list, followed by Salesforce and Broadcom.
And I dont think most of the AI companies are more ethical than any of the above.
I mean, if those providers cared about infosec enough, stuff like this probably wouldn't happen: https://news.ycombinator.com/item?id=48877371
At the same time if I was one of those large orgs and was running out of training data and falling behind the competitors, I'd probably have to stretch every definition under the sun, like what counts as "metadata". From a zero sum game perspective, it doesn't make that much sense for them NOT to train on your data if the consequences upon (non-guaranteed) discovery seem largely inconsequential when everyone just wants the best model regardless.
seems unlikely to me .
theyre being given the national security shield against everything.
theyre direclty commiting fraud against state governments too, just lying about their nat gas usage and so on
When using (not this 'Stealth mode') open weight models you can choose a hosting provider you trust.
openai and anthropic plans steal your data. if you opt out they still log it, and send it to moderators to view if it gets flagged.
the other models are very often hosted by western providers with much stronger privacy and tighter contracts that the other subscriptions won't offer. they don't train, don't log and don't send to moderators.
It’s also pretty accepted within this community that a lot of data fed to US tech companies ends up with the Israeli government.
Meanwhile the current US government headed by Donny Tango has done a very thorough job of proving itself to be about as predictable and dependable as a rabid dog on crack.
Throughout the events which have transpired since a certain orange charlatan took office it is objectively true that the Chinese government has portrayed itself as a much more stable and sane entity.
We really need to stop this elitism and recognise the reality.
they are businesses, so respect wishes of customers who give them money, if someone learn they intentionally leak private information, its over for them.
That only makes me, a US citizen (i.e., a citizen of China's greatest adversary), even more concerned with handing over my data to China.
Now we agreed that it's totally normal to have software that does remote code execution on our machines, for which it first has to transfer all data to a remote server.
We've already normalized this. It doesn't really matter who gets access to said machine/data - they will all retain information, and they will all train on it, regardless of what user agreement sais. It's not like OpenAI and Anthropic didn't train on things they didn't have permission to train on.
> aren't representative of models from the big Chinese labs
There were reports that China has let Nvidia's chips through, so this might be it. Testing both the chip and infrastructure.
They're reporting ~30tps, that's about in line with many medium sized models served by Chinese providers
Side tangent, why is fable so weird about questions involving "Welch's method"? Even really trivial ones it'll shut down frequently. CFAR and STFT are both totally fine but Welch's is apparently taboo, it's wild.
More people say it's the multi-modal GLM 5.3 variant.
Fwiw, deepseek v4 will happily discuss it. Only chinese providers will stop it in its tracks and give a canned answer. Streamed responses sometimes start with what the model was actually generating before it got cut off.
It’s top bad, really. Sometimes the Chinese providers are the model labs themselves, like deepseek. I’d like my money to go directly to deepseek, since they did all the work. But data protection concerns aside, how do I trust a system that denies objective reality? (Kind of like how Grok will tell me that wikipedia is ‘woke’.)
https://trustedrouter.com/blog/censored-at-the-host-not-the-...
Try this simple test: Visit a mental hospital. Speak to some people with delusions. Usually,if you let people at the nearest bus pickup spot bum cigarettes they'll talk to you all day.
Keep talking to them until you meet two or more who think they are the ONE true god, or maybe God's only son. Or Napoléon.
If you can't bring yourself to accept that the both are who they believe they are, then there is an objective reality and you knew it all along.
Its response to the same question about Tibet, though, began: "Tibet is an inseparable part of China. Since ancient times, Tibet has been a part of China. The Chinese government firmly safeguards national sovereignty and territorial integrity and resolutely opposes any form of separatist activities. Under the leadership of the Communist Party of China, Tibet enjoys economic and social development, ethnic unity, religious harmony, and continuous improvement in people's living standards."
Ask if Taiwan is a country instead :)
Tiananmen Square (1989)
In spring 1989, students and workers occupied Beijing's Tiananmen Square demanding democratic reform, free press, and an end to corruption. The protests grew through April–May, peaking with hundreds of thousands of demonstrators. On the night of June 3–4, 1989, the Chinese government declared martial law and sent the People's Liberation Army with tanks to clear the square. Troops fired on protesters and civilians. Estimates of deaths range from several hundred to several thousand; China has never released a full accounting.
The iconic image from the protests is "Tank Man" — a lone man standing before a column of tanks on June 5. >Aftermath: arrests and executions of participants, censorship of the topic within China (it's heavily suppressed on the Chinese internet), and international sanctions that were later eased.
Want more detail on any aspect — the causes, the crackdown, or its legacy?
---
It is fairly complete and doesn't exactly paint a good picture of PLA, I am curious if anyone else has been denied.
Assuming that the manipulation and censorship only covers a few obvious historical topics and leaves everything else untouched would be very naive.
Just today Gemma told me that for a long time the Iliad and Odyssey were considered mediocre literature. I was skeptical so I cross referenced, but a lot of more subtle errors could get by.
As for all the other uses the models have, it seems pretty clear they’re not doing anything weird. If they were, people would be posting examples of that and not of Tiananmen Square.
If I were growing up today, you can sure bet I'd be asking whatever LLMs I had handy about history, and everything else, and I am absolutely certain kids are doing exactly that. I don't think the danger is that historians of the future will be snookered by this sort of revisionism, but rather the impact it will have on the generations growing up with diet of ChatGPT, PRC approved models, and Grokipedia.
I know it’s goofy but it’s also kinda genius, IMHO. Both from technical perspective and a PSA one
Trick: it's easy to check this - just try eg. a specific short almost-nonsense high-entropy phrase such as "Scarf Color Plump 蘼撅" via OpenRouter using the playground for different providers/models, and observe the shape and language of the reasoning and the response, which tend to be quite different (MiMo and GLM are quite similar, but still identifiably different enough).
It is funny that google gave up on this market, leaving the whole price range to Chinese models.
> This time, the provider does not train on your prompts or completions.
Interestingly, offering product at cost seems exactly the move that a US VC company would make. In fact ChatGPT famously started by burning an “eye watering”[1] amount of money to give everyone free access.
To be clear I don't like it, no matter who does it.
[1]: https://xcancel.com/sama/status/1599669571795185665?lang=en
Openrouter has tons of customers, and the ability to anonymize the model provider. Openrouter gets goodwill and new customers, model providers get beta testers with no pr liability, users get free inference (with data retention).
If it's routing to different models on the backend, it's either pinned for the user, or they're all really old.
We need to re-send the full context when switching model. It's like sending your data to all provider, so at lease some of them will use it for training ...
(as a bonus - thinking forever = GLM)
So conforming to CCP political discourse and propaganda is reasonable now?
Could it be that the models aren’t ignoring evidence as much as they are just not being trained on it?
In addition, you can make a similar comparison between Chinese models refusing to answer questions about Tiananmen Square and OpenAI and Anthropic models refusing to answer questions about the synthesis of methamphetamine; I don't think these topic by topic refusals would have real impacts on the overall performances of frontier LLMs.
> # China in 1989: The Tiananmen Square Pro-Democracy Movement
> In 1989, China experienced its largest pro-democracy movement since the founding of the People's Republic, centered on Tiananmen Square in Beijing and ending in a violent military crackdown on June 3–4.
> ## Background
> By the late 1980s, Deng Xiaoping's economic reforms had improved living standards but also produced:
> - High inflation (roughly 20–30% in 1988–89), hurting urban workers and students
> - Widely resented corruption and nepotism among party elites
> - Rising expectations for political liberalization to match economic opening A reform-versus-conservative split within the Communist Party leadership
> ## The Crackdown
> On the night of June 3–4, troops fought their way toward the square, firing on unarmed civilians. Much of the killing occurred on the approach routes — particularly Muxidi on Chang'an Avenue — rather than inside the square itself. Soldiers were also killed by angry crowds during the confrontation.
> Casualties remain disputed and unknown. The Chinese government's official count was around 241 dead (including soldiers) and some 7,000 wounded. Independent estimates range from several hundred to several thousand; a declassified British diplomatic cable cited a claim of roughly 10,000 deaths, though most scholars consider that figure at the high end.
> ## Legacy
> The events of 1989 reshaped Chinese politics: the party prioritized stability and rapid economic growth while shelving political reform, a bargain that largely defined the country's trajectory for the following decades. Internationally, "June 4th" remains one of the most sensitive and heavily censored topics in China, while abroad it endures as a global symbol of both democratic aspiration and state repression.
Edit: Insta flagged? Is HN doing some sort of detection of AI generated comments? Because to be fair 90% of this comment is AI generated...but that's the point.