Some relevant excerpts:
“Because we believe the most important thing now is to participate in the global innovation wave. For many years, Chinese companies are used to others doing technological innovation, while we focused on application monetization — but this isn’t inevitable. In this wave, our starting point is not to take advantage of the opportunity to make a quick profit, but rather to reach the technical frontier and drive the development of the entire ecosystem.”
“We believe that as the economy develops, China should gradually become a contributor instead of freeriding. In the past 30+ years of the IT wave, we basically didn’t participate in real technological innovation. We’re used to Moore’s Law falling out of the sky, lying at home waiting 18 months for better hardware and software to emerge. That’s how the Scaling Law is being treated.
“But in fact, this is something that has been created through the tireless efforts of generations of Western-led tech communities. It’s just because we weren’t previously involved in this process that we’ve ignored its existence.”
“We do not have financing plans in the short term. Money has never been the problem for us; bans on shipments of advanced chips are the problem.”
“In the face of disruptive technologies, moats created by closed source are temporary. Even OpenAI’s closed source approach can’t prevent others from catching up. So we anchor our value in our team — our colleagues grow through this process, accumulate know-how, and form an organization and culture capable of innovation. That’s our moat.
“Open source, publishing papers, in fact, do not cost us anything. For technical talent, having others follow your innovation gives a great sense of accomplishment. In fact, open source is more of a cultural behavior than a commercial one, and contributing to it earns us respect. There is also a cultural attraction for a company to do this.”
It's really a shame that in the current world, the art of hardware is dying out, where hardware people are not properly compensated and appreciated [1].
Liang Wenfeng belongs to this breed of engineers with hybrid hardware and software background that have money and at the same time founding and leading and companies (similar to two Steves of Apple), they're a force to reckon with with even with severe limitations, in case of Chinese companies computing resources sanctions CPU/RAM/GPU/FPGA/etc. But unlike two Steves these new hybrid engineers that raised in Linux era are the big believers of open source, as Google rightly predicted in case of LLM none of the proprietary LLM solutions has the moat [2],[3].
[1] UK's hardware talent is being wasted (1131 comments):
https://news.ycombinator.com/item?id=42763386
[2] Google “We have no moat, and neither does OpenAI” (1039 comments):
https://news.ycombinator.com/item?id=35813322
[3] Google "We have no moat, and neither does OpenAI" (2023) (42 comments):
A bit of irony was that this researcher (from Europe) used to work in the same lab as me in Beijing. But these days the talent doesn’t flow so easily as it did a decade+ ago (but maybe it will again? Researchers aren’t very nationalistic and will look for the best toys to play with).
Maybe you're thinking of 1-bit DACs with oversampling and noise shaping: https://en.wikipedia.org/wiki/Delta-sigma_modulation
I agree with everything you said but this part is "broken clock will be right twice a day". It is what Google would have said regardless. A moat is never impossible to cross, it's just a passive superpower making the "enemy's" job that much more difficult. By Google's suggested interpretation of a moat, moats simply do not exist. They can all be crossed eventually, when ingenuity catches up to big budgets, so it's like they were never there?
I don't buy it that they knew or predicted anything. If Google knew something about hidden optimization available to everyone or had more reason to suspect this is the case beyond "every technology progresses", they'd already be built into their models by now (it's been 2 years since the "prediction") but there's no evidence they were even close. And there's still a HW moat. The amount of high performance HW BigAI has or affords can still make a huge difference everything else being equal, after building in all those "free" optimizations.
At the least the big companies have the ability to widen the moat when they feel pressure of the small competitors closing in. It's clear now that more money can do that. If ingenuity can replace money, then money can replace ingenuity, even if via buying out startups, paying for the best people, and so on. They've shown it again and again.
That's alway the issue with outsourcing. You rely exclusively on middlemen, middlemen will realize they can cut out their middlemen and just go directly to the customers.
These claims are quite hard to square with the long waits for H1B visas, extremely high salaries in the technology sector and net immigration to the US from China.
I’m not aware of any Americans or Europeans in my network who have gone the other direction to China.
Perhaps you have different data about the demand for tech worker visas in China.
EE as a US career is night and day from Software centered engineers. Night and day from 20 years ago as well.
http://www.talentsquare.info/blog/fall-engineering-jobs-elec...
I don't think it's much of a controversial take to suggest that China is kicking the US's butt in silicon chip production. EE's are one of the primary fields traditionally seeked to work with this.
That will be controversial until mainland China produces modern process chips economically (they can do one or the other so far). Rather Taiwan and South Korea are not the EE powerhouses. China though pays better than Taiwan (a lot of the hardware researchers in my Beijing lab were from Taiwan and Korea).
I think the plan was to have it built by 2027, but who knows now. Meanwhile, Trump called the CHIPS act "ridiculous" (very optimistic future, clearly) and just imposed tariffs on Taiwan.
I just bought a new laptop last night just in case. A refurbished M3 Max with enough memory to run DeepSeek 70b :).
If that's a problem for the West now, it's a problem of our own creation.
But more seriously, DeepSeek is a massive boon for AI consumers. It's price/performance cannot be beat, and the model is open source so if you're inclined to run and train your own you now have access to a world-class model and don't have to settle for LLaMA.
No, but the same sort of people certainly told us that we were :-)
Cf. the whole "the GPL is viral and will kill the industry" spiel we got to hear for years.
He has also spoken about world domination ;)
>If you need more than 3 levels of indentation, you're screwed anyway, and should fix your program.
Got me thinking. I might heighten up to 4 or 5 simply because modern code needs 2 indents just to start writing a function in a struct. But the quote wasn't as crazy as I thought, even 30 years later.
No, because as Stallman had pointed out Linux isn't GNU. One of the differences between the "open source" crowd and the "free software" crowd is that the latter actually does have an explicit goal of denying proprietary software the ability to exist.
you could be a communist if you open source your project
so maybe in that alternative universe, there would be something like close-source-statement instead of open source license, to avoid be accused as a communist
If the goal is to erode the moat around powerful US tech companies, by making tech that rivals theirs and releasing it to the public, it's just good for the world. The only way it isn't is if you believe that power should remain in the hands of certain elites.
Whereas what's Sam doing? Announcing a non-existing 500 billion dollar investment with the president, while all AI companies in the wesy support a trade ban for Nvidia GPUs in China.
I actually praise that offensive move, if AI companies can lost so much value from DeepSeek's open research then it's well deserved, they shouldn't be valued as much.
which is fine and dandy to do. In fact, i wish deepseek success. The US tech industry needs disruption.
The USA has never once had friendly relationships with a large power, perhaps with the very special case of the USSR alliance during WWII (and not a second after it). The European powers and Canada are extremely US friendly and support US policies (at the head of state level) in almost everything. Relations with China were good while China was a weak and poor state, acting as almost slave labor for the USA - not great now that they are rising up. Relations with Russia were good for a brief window after the fall of the USSR, while Eltsyn seemed to be "our guy", but quickly soured when it became clear he would not dance to their tune (not to sya that he was a good man or that his disputes with US intentions were good - Russia would have probably been in a better state if it had allied itself more with the USA, rather than becoming the belligerent territorial authoritarian oligarchy that it has).
> OpenAI Hails $500 Billion Stargate Plan: 'More Compute Leads to Better Models'
The cynic in me is much more likely to see this as western companies giving up on innovation in favor of grift, and their competition in the east exposing the move for what it is.
This is why competition is good. Let's make this about us (those who would do this in the open) and them (those who wouldn't) and not us (US) and them (China).
[1] https://en.wikipedia.org/wiki/Fifth_Generation_Computer_Syst...
Although it sounds like that project, if successful, would've been pretty fantastic for computing in general. I'm far less interested to see proprietary models secure dominance, whichever country they're in.
The whole thing is no longer a startup being disruptive
The US has been trying to find a "space race" challenge to justify its military spending increases for a while, AI is going to be that, but it's more driven by the US oligarchy than the US MIC this time.
That means that it's going to be driven by financial wealth accumulation instead of power accumulation.
the AI race between China and the US is going to shape the future of our generation. CIA has all the motivations to just eliminate all those core Chinese members as they pose direct national security threat to the US dominance in AI.
you need to be really naive to not being able to see these.
A couple of years ago, it was VR/AR (2nd time around for VR, it had been hyped in the '90s), before that it was "cloud" etc etc.
The CIA is not going to be going around assassinating AI developers, any more than they are going to kill the people working for ASML because they threaten US dominance in chips.
"Ah yes, because comparing AI’s transformative impact to VR’s niche flops or dismissing cloud (now the backbone of modern tech) proves you’ve got the insight of a dial-up modem. Stay salty and irrelevant!"
I get side eyes from Americans when I bring this up as a key factor when they try to shit on Europe for "lack of innovation", it's more a lack of bottomless stacks of cash enabling undercutting competition on price until they fold, then jacking up prices for VC ROI.
You pay with your data.
This could very well be the long-term plan with DeepSeek, or it could be the AI application of how China deals with other industries: massive state subsidies to companies participating in important markets.
The profit isn't the point, at least not at first. Driving everyone else out is. That's why it's hard to get any real name brands off of Amazon anymore. Cheap goods from China undercut brand-name competition from elsewhere and soon, that competition was finding it unprofitable to compete on Amazon, so they withdrew.
I used to get HEPA filters from Amazon that were from a trusted name brand. I can't find those anymore. What I can find is a bunch of identical offerings for "Colorfullfe", "Der Blue" and "Extolife", all priced similarly. I cannot find any information on those companies online. Given their origin it's safe to assume they all come from the same factory in China and that said factory is at least partially supported by the state.
Over time this has the net effect of draining the rest of the world of the ability to create useful technology and products without at least some Chinese component to the design or manufacture of the same. That of course becomes leverage.
Same here. If I'm an investor in an AI startup, I'm not looking at the American offerings, because long-term geopolitical stability isn't my concern. Getting the most value for my investment is, so I'm telling them to use the Chinese models and training techniques for now, and boom: it just became a little less profitable for Sam Altman to do what he does. And that's the point.
In this case it's open source, and with papers published. So any US company can (way more cheaply than ChatGPT and co iiuc) train their own model based on this and offer it as well.
The biggest purchaser of technology and goods and services is the US Government. It spends over $760 billion annually on products and services.
But if any other country does the same it would classify as "massive state subsidies".
I would take it a step further and say that the biggest employer in US is the US Federal Government.
Give it ten years.
So American investors dumped a metric crapload of money into the Chinese economy for things like manufacturing. The labor was cheap, and anyone who wanted better outside of the status quo was going to be turned into hamburger under the treads of a tank. No longer would they have to deal with the labor unions of the Midwest and Great Lakes regions, or have to deal with American environmental, corruption, and labor laws. The investment was the seed money for the startup we know as modern China.
Similarly the EU of 2025, has nothing to do with WW2-era starvation, that has been over half a century in the past.
And of course there was literal starvation in China as well after WWII, and much more poverty there than in the EU 30 years ago (even including Eastern Europe).
Secondly, China also has extremely high bureaucracy, and extreme levels of government regulation - a classic problem for dictatorial regimes, especially ones spanning huge spaces (where direct control is physically impossible, even in the information age).
The big difference is that EU governments have drunk the coolaid on modern economical theories, and don't generally pick winners and losers in the market (beyond few key companies with deep ties to the ruling elites, mostly in banking), don't invest massive amounts to prop up companies doing price dumping, and generally play within the rules of world trade.
Of course, those rules are made up specifically to prevent any state from using its power to out-compete incumbent companies, many of which are US owned, but also German, French, Spanish etc owned.
Also, there is little appetite for EU level strategic decisions, EU member countries are far too divided. For example, Finland probably didn't have the power to prop up Nokia's phone division when Apple and Samsung started eating its lunch with smartphones, and France or Germany wouldn't have wanted to invest EU resources into doing it either. France is likely not going to be ok with propping up a German rival to BYD using massive funds, or vice versa for a French company.
So, while collectively the EU easily rivals China on money antld the USA on population, it is far too divided to pool those powers together, and the EU population mirrors this sentiment - there is not a strong EU identity that would see a Belgian person deeply proud of a major tech company based in Slovenia, or a Czech person cheering for a massive new investment in Portugal.
They extract the very same data from paying users. And even with data factors in, they give products away at loss explicitly to undercut the competition.
You are replying to a thread with the DeepSeek CEO saying the opposite (e.g., DeepSeek built upon transformers, Llama, PyTorch, etc.)
What happened to fighting climate change?
But for the rest of humanity it doesn't look so bad.
[1] https://www.joelonsoftware.com/2002/06/12/strategy-letter-v/
“Our principle is that we don’t subsidize nor make exorbitant profits. This price point gives us just a small profit margin above costs.”
So, sadly, even something that seems noble and refreshing like open-sourcing their AI advancements will be treated with suspicion.
What gets me is when people present it like it's bad to play that game. Like "it's ok when we do it".
It's not bad. But the western superpowers, however flawed, are at least familiar. For the past 75 years we've avoided world war under this power balance. A new power balance could turn out better in that regard, but that doesn't mean it won't be scary, especially for those who value individual liberty.
I suspect "we westerners" think of "we westerners" and do not give a flying fuck about "the rest of the world". Well, as long as they keep trading exclusively in our currency etc. etc.
I agree that the CCP's view of the world and population control is negative. But don't let that poison your opinion of all Chinese people. We're all people on Earth, and we need to be forging bonds with our intelligent and good-hearted international kin that break down the walls that those in power create to keep themselves there.
It would mean having to eschew the neoliberal ideals that impede research and development in favour of the old that made America and to some extent the rest of the West the dominant superpower in R&D for many decades. We should be familiar with it, even if we have lived all or most of ours lives in the former.
Or it would be hard to convert back and we'd have a war first.
All I heard from OpenAI was that we need regulation which maybe happen to fit their business interest.
Compare this to the interviews of Altman or Musk, talking vaguely about elevating the level consciousness, saving humanity from existential threats, understand the nature of the universe and other such nonsense they pander to investors.
If that is the game they're playing, I'm all for it. Maybe it's not the result that the sanctions were intended to have, but motivating China to share their research rather than keep it proprietary is certainly a win. Making AI more efficient doesn't reduce the value of compute infrastructure; it means we can generate that much more value from the same hardware.
https://www.pekingnology.com/p/ceo-of-deepseeks-parent-high-...
Interesting tidbit:
>So far, there are perhaps only two first-person accounts from DeepSeek, in two separate interviews given by the company’s founder.
I knew DeepSeek was lowkey but I didn't expect this much stealthmode. They were likely off CCP boomer radar until last week when Liang met with PRC premiere after R1 exploded. Finance quants turned AI powerhouse validates CCP strategy to crush finance compensation to redirect top talent to strategic soft/hardware. I assume they're going to get a lot more state support now, especially if US decides to entity list DeepSeek for succeeding / making the market bleed.
Create an ecosystem and all tides rise.
> "For many years, Chinese companies are used to others doing technological innovation, while we focused on application monetization..."
> “But in fact, this is something that has been created through the tireless efforts of generations of Western-led tech communities. It’s just because we weren’t previously involved in this process that we’ve ignored its existence.”
you don't know cpc
you don't know china
and you don't know chinese
you just imagine cpc and chinese as characters in some shit comics
every chinese could possibly said that, and cpc say this a lot everyday, and cpc made national strategy base on that, you can find these words in many gov documents
so you guys are right about one thing: china is a threat, because from cpc to normal chinese, there're tons of people in china think like this, and many of them eager to challenge this
just like what deepseek is doing right now
you are very perceptive, if someone who use CPC rather than CCP, they're either chinese, or pro-china, or worse, communists
i just say that mindset (like what deepseek ceo said) is super common in China, not something hard to say or forbiddened
Generally speaking, I assume CCP is involved with anything of strategic significance. They would even chase random benign influencers.
Microsoft apply censorship to Bing search results in China. It doesn't mean they are controlled by CCP. They just got impacted by law and they want to keep operate in China.
I don't care that deepseek's own service has censorship. I would care, if they have this censored weights but haven't revealed it was (aka, fraud by omission).
Until now.
Obviously it’s a power play as China seeks influence beyond money now that’s secured. I think people should receive it on its merits.
The strategy of open sourcing to eliminate the competitive mode of those with proprietary designs is a bit of a desperate play, favored by the weaker competitor, lacking access to the desired market.
You can also perceive it as hostile and in line with dumping practices, where a high volume of product is dumped into a market at cheap prices.
But besides these tactical aspects, which are no doubt being utilized, there’s a inescapable technological reality that obviously efficiency of AI will improve, and the most efficient designs would seem to rise to the top. This utilization of and guiding of inevitable historical trends for their own advantage is a very Chinese communist dialectical materialist approach to take, and I think we can expect to see more of these types of ‘surprising’ moves by entities out of China in the decades ahead as these kind of competitions heat up. The Chinese have a very deep and a very different ideological background that would justify these types of moves as making perfect sense to them, although they simultaneously appear as nonsensical to people from other backgrounds.
I feel like the reaction of the west is protecting him from a reaction from Chinese authorities
“To people who see the performance of DeepSeek and think: ‘China is surpassing the US in AI.’ You are reading this wrong. The correct reading is: ‘Open source models are surpassing proprietary ones.’ DeepSeek has profited from open research and open source (e.g., PyTorch and Llama from Meta). They came up with new ideas and built them on top of other people’s work. Because their work is published and open source, everyone can profit from it. That is the power of open research and open source.”
[1] https://www.forbes.com/sites/luisromero/2025/01/27/chatgpt-d...
As if anyone riding this wave and making billions is not sitting on top of thousands of papers and millions of lines of open source code. And as if releasing llama is one of the main reasons we got here in AI…
Innovation ALWAYS follows this path. Something is invented in a research capacity. Someone implements it for the ultra rich. The price comes down and it becomes commoditized. It was inevitable that “good enough” models became ultra cheap to run as they were refined and made efficient. Anybody looking at LLMs could see they were a brute forced result wasting untold power because they “worked” despite how much overkill they were to get to the end result. Them becoming lean was the obvious next step, now that they had gotten pretty good to the point of some diminishing returns.
Really this should be an indictment of corporate bloat, having hundreds of thousand headcount companies distracted by performance reviews, shareholders, marketing, rebuilding the same product they launched two years ago under a new name.
Yeah.
There are some shorter words or acronyms for it though, roughly equivalent to your about 30-word paragraph above:
IBM DEC Novell Oracle MS Sun HP ... MBA , all in their worse days or incarnations or ...
Because we saw, what a week ago the leading indicator that the money people were now feeling happy they were in charge which was that weird not-government US$500 billion investment in AI announcement. And we saw the same being breathlessly reported when Elon Musk founded xAI and had "built the largest AI computer cluster!"...as though that statement actually meant anything?
There was a whole heavily implied analogy going on of "more money (via GPUs) === more powerful AIs!" - ignoring any reality of how those systems worked, their scaling rules or the fact that inferrence tended to run on exactly 1 GPU.
Even the internet activist types bought into this, because people complaining about image generators just could not be convinced that the Stable Diffusion models ran locally on extremely limited hardware (the number of arguments where people would discuss this and imply a gate while I'm sitting their with the web GUI in another window on my 4 year old PC).
Riding hype, and dumping at the first sign of issues, follows that perfectly well.
Regulatory capture only benefits you nationally. You might even get used to it.
R1 is a 650b monster no one can run locally.
This is like complaining an electric bike only goes up to 80km/h
The full version... If you have to ask you can't afford it.
We can laugh at that (like I like to do with everything from Facebook's React to Zuck's MMA training), or you can see how others (like Deepseek and to a lesser extent, Mistral, and to an even lesser extent, Claude) are doing the same thing to help themselves (and each other) catch up. What they're doing now, by opening these models, will be felt for years to come. It's draining OpenAI's moat.
Wait.. are you saying it wasn't? Just releasing it in that form was a big deal ( and heavily discussed on HN, when it happened ). Not to mention, a lot of the work that followed on llama partly because it let researches and curious people dig deeper into internals.
Open source means we need to be able to reproduce what they’ve built - which means transparency on the training data, training source code, evaluation suites, etc. For example, what AI2 does with their OLMo model:
They can “profit” (benefit in product development) from it.
They just can't profit (return gains to investors) much from it, because that requires a moat rather than a market free for all that devolves into price competition and drives market clearing price down to cost to produce.
All that cloak and dagger stuff comes at a cost, so it's only worth paying if you think you can maintain your lead while continuing to pay it. If the open source community is able to move faster because they are more focused on results than you are, you might as well drop the charade and run with them.
It's not clear that that's what will happen here, but it's at least plausible.
DeepSeek did something legitimately innovative with their addition of Group Relative Policy Optimization. Other firms are certainly free to innovate as well.
They just didn't.
Worse for the proprietary labs is how much they've trumpeted safety regulations. They can't just release a model without extensive safety testing, or else their entire regulatory push falls apart. DeepSeek can just post a new model to Hugging Face whenever they feel like it — most of their Tiananmen-style filtering isn't at the model level, it's done manually at their API layer. Ditto for anyone running finetunes. In fact, circumventing filtering is one of the most common reasons to run a finetune... A week after R1's release, there are already uncensored versions of the Llama and Qwen distills published on HF. The open source ecosystem publishes faster.
With massively expensive training runs, you could imagine a world where model development remained very centralized and thus the few big labs would easily fend off open-source competition: after all, who would give away the results of their $100MM investment? Pray that Zuck continues? But if the training runs are cheap... Well, there are lots of players who might be interested in cutting out the legs from the centralized big labs. High Flyer — the quant firm that owns DeepSeek — no longer is dependent on OpenAI for any future trading projects that use LLMs, for the cost of $6MM... Not to mention being immune from any future U.S. export controls around access to LLMs. That seems very worthwhile!
As LeCun says: DeepSeek benefitted from Llama, and the next version of Llama will likely benefit from DeepSeek (i.e. massively reduced training costs). As a result, there's incentive for both companies to continue to publish their results and techniques, and that's bad news for the proprietary labs who need the LLMs themselves to be profitable and not just the application of LLMs to be profitable... Because the open models will continue eating their margins away, at least for large-scale deployments by competent tech companies (i.e. like Linux on servers).
They kinda did: https://en.wikipedia.org/wiki/Azure_Linux
Looking back at the PDP handbook, it's not even clear that LeCun deserves the credit for CNNs, and he himself gives credit for the core "weight sharing" idea to Rumelhart.
Chollet's claim to fame seems to be more as creator of Keras than researcher, which has certainly been of great use to a lot of people. He has recently left Google and is striking out to pursue his own neuro-symbolic vision for AGI. Good luck to him - seems like a nice and very smart guy, and it's good to see people pursuing their own approaches outside of the LLM echo chamber.
They literally used GPT and Llama to help build DeekSeek, it responds thinking that it's GPT in countless queries (which people have been posting screenshots of). They 'cheated' exactly as Musk did to build xAI's model/s. So much of this is laughable scaremongering and it's absolutely not an accomplishment of large consequence.
It's a synth LLM.
isn't LeCun basically admitting that he and his team didn't have the creative insight to utilize current research and desperately trying to write off the blindside with exceptionalism?
not a good look tbh
The thing is that the steam engine guys researched thermodynamics and developed the mechanics and tooling which allowed the diesel engine to be invented and built.
Also, for every breakthrough like DeepSeek which is highly publicized, there are dozens of fizzled attempts to explore new ideas which mostly go unnoticed. Are these wasted resources, too?
Given your take, this is a meaningless question, no?
As you point out, all resource usage that lead up to the creation of the diesel engine were necessary preconditions. While one might be able to imagine a parallel universe where the diesel engine was created in another way without all the things in between that might feel like a waste, that is not this universe. In this one, it took what it took.
Same goes for AI. That AI researcher had to eat that sandwich double wrapped in plastic, subsequently placed in another plastic bag in order to get to where he got. Which might feel like a "waste of resources". I am sure you can easily imagine a parallel universe where he didn't eat something that used up so much plastic. But that was the precondition necessary in this universe.
So, ultimately, either everything is a waste of resources or nothing is. And there is no meaning in trying to find a distinction between those two.
Resource allocation in this context isn’t at all binary.
LeCun is in a different part of the organization - FAIR (FaceBook AI Research), and isn't even the head of that. He doesn't believe that LLMs will lead to AGI, and is pursuing a different line of research.
Thanks. Great observation. Sounds indeed extremely plausible that they use the LLM for automated data cleaning.
Have they actually pivoted, or are they just messing around to see what sticks?
regardless, high-flyer is an HFT firm
High-Flyer says it took directional bets and held positions, which makes at least part of it not HFT.
Also, I doubt that most quant money is in market making nowadays. That was true at some point and that's true of HFT, but I doubt it is of quant trading in general anymore.
Besides, High Flyer certainly isn't a market maker, or they wouldn't be a hedge fund. You can't really be both, hence with Citadel and Citadel Securities (Market Maker) are so strictly divided.
I don't personally buy their story, and after having used Deepseek it kind of sucks and hallucinates a lot if I'm being objectively honest.
I mean a few million for this is okay - that's cool.. but it is useless. I can understand billions of dollars into something that actually works >50% of the time.
I know it is too early, but I'd not be surprised if this was CCP intervention using a hedge fund to try and tank US AI stocks for a specific reason.
I mean again, just being objectively honest, Deepseek kind of sucks and is maybe on par with early-2023 era models.
It seems obvious that you need to have a model trained, or fine-tuned, on some reasoning data (with backtracking etc) such that reasoning behavior is part of it's repertoire, before you can use RL to hopefully get it to use such reasoning pursuant to whatever goals you are setting. I'd not be surprised if they used O1 outputs to bootstrap the model in this way, although O1's reasoning traces are a deliberate obfuscation of what it is really doing (an after-the-fact summary) so even if this is the case that should be borne in mind!
OTOH, while reasoning data may be scarce in the wild, it's presumably not entirely unavailable, and/or DeepSeek may have created some themselves, so who knows what mix DeepSeek used for this initial bootstrapping stage. As you say, this aspect remains as "secret sauce".
Of course once they've got their first stage model trained they then use that to generate data for the second/final stage.
This would make perfect sense if the goal is to devalue existing players more than it is capture the market.
If they did not open source it and instead just launched a payed (albeit much cheaper) closed model with similar performance to O1, would people trust them?
I don't think DeepSeek has any malicious intent, but boy oh boy am I glad the USA boys get wrekt by this (though I also lose money on stocks).
This is just poetic justice for the Orange Man's backwards 17th century policies.
But who's the baddies now? China is not waging war everywhere. Or threatening to steal Greenland... Or ruining our teenagers with social media.
Russia is currently invading Europe, to the tune of hundreds of thousands KIA. And Russia's invasion would be dead in the water without Chinese support.
To paraphrase the Chinese rep on the UN. If China indeed supported Russia, then this war would have ended by now.
It seems a great move.
I am sorry if my English isn't great... and, yes, sometimes I do use voice to text. Android is particularly good at messing up what I want to say.
Anyway, we may be past peak OpenAI at this juncture.
I wouldn't hold my breath on getting access to it.
There is no "secret" sauce. Only sauce.
Additionally, R1-Zero shows that you don't even really need much secret sauce data, since they trained it with zero SFT data. Take an existing base model, do GRPO RL, and tada: you have a SOTA reasoning model. SFT data improves it, but the secret sauce isn't in the data.