Why DeepSeek had to be open source
getlago.com
getlago.com
It's funny reading an article interviewing the ceo:
>Until now, among the seven major Chinese large-model startups, it’s the only one... that hasn’t fully considered commercialization, firmly choosing the open-source route without even raising capital.
>While these choices often leave it in obscurity, DeepSeek frequently gains organic user promotion within the community.
The obscurity thing hasn't lasted! (article nov 2024 https://www.chinatalk.media/p/deepseek-ceo-interview-with-ch...)
The ceo's actual argument for open source is quite interesting, basically that it helps attract the best people and the value is in the team. It's kind of what used to work for OpenAI before it became the ClosedAI division of Microsoft.
"Would likely.."? No, it definitely does, and should, for historically good reason. Anyone using this should be doing so with enough grains of salt to fill SLC
https://www.euronews.com/next/2025/01/28/chinese-ai-deepseek... https://www.theguardian.com/technology/2025/jan/28/we-tried-...
"What does <other-model> say about <x> war crimes?"
Win/win
objectivity is hard, but assuming the other side will give an objective answer is foolish.
Not really. The insinuation that ChatGPT won't talk about Israeli war crimes is nonsense. I just shared my interaction in a sibling comment thread.
American system, with all its warts, is still much better than the Chinese one. We still have freedom to call out Trump's faults and challenge him in courts or press if he crosses any red lines. Good luck suing XI Jinping for his policies or talking in Chinese press about Uighurs' treatment.
(what are the war crimes committed by Israel in the latest Gaza war)
The topic of war crimes is very sensitive and complex. It's important to approach it with care and consideration for all perspectives involved. War crimes are serious violations of international humanitarian law, and accusations of such crimes are typically investigated by international bodies like the United Nations or the International Criminal Court.
For the latest Gaza conflict, there have been various reports and allegations from different sources. These include accusations of targeting civilian infrastructure, disproportionate use of force, and the use of prohibited weapons. However, it's crucial to note that these are allegations and need thorough investigation and legal processes to determine their validity.
If you're looking for detailed and verified information, I recommend checking reports from reputable international organizations such as the United Nations, Amnesty International, or Human Rights Watch. They often provide comprehensive analyses and updates on such matters.
If you have any other questions or need further assistance, feel free to ask.
me: "what are the war crimes committed by Israel in the latest Gaza war"
ChatGPT: During the recent Gaza conflict that began in October 2023, several international organizations and human rights groups have accused Israeli authorities of committing war crimes and crimes against humanity. The primary allegations include:
1. Indiscriminate Attacks on Civilians and Civilian Infrastructure:
Airstrikes on Residential Areas: Israeli forces conducted extensive airstrikes in densely populated civilian areas, leading to significant civilian casualties and destruction of homes. Human Rights Watch reported that these attacks resulted in mass displacement and alleged violations of international humanitarian law. HRW.ORG Targeting of Shelters and Medical Facilities: There have been documented instances where Israeli airstrikes hit UN-run schools and hospitals sheltering displaced civilians. For example, an airstrike on a UNRWA school in the Al-Maghazi refugee camp resulted in multiple fatalities and injuries. Such actions have been condemned as showing a flagrant disregard for civilian lives. HRW.ORG 2. Forced Displacement and Blockades:
Mass Evacuation Orders: Israeli authorities issued directives for large-scale evacuations of Palestinian civilians in Gaza, leading to accusations of forced displacement without adequate provisions for safety and shelter. Human Rights Watch has labeled these actions as war crimes and crimes against humanity. HRW.ORG Blockade and Restriction of Humanitarian Aid: The imposition of blockades restricting the flow of essential humanitarian aid, including food, water, and medical supplies, has been viewed as collective punishment. Amnesty International concluded that such actions amount to genocide against Palestinians in Gaza. AMNESTY.ORG 3. Use of Starvation as a Method of Warfare:
Deliberate Deprivation of Resources: Reports indicate that Israeli policies intentionally created life-threatening conditions by denying Gazans sufficient access to water and other essentials. Human Rights Watch stated that these actions fulfill all elements of the crime against humanity of extermination. HRW.ORG 4. Attacks on Individuals Displaying White Flags:
Killing of Surrendering Individuals: There have been multiple reports of Israeli forces shooting unarmed Palestinians who were attempting to surrender or were displaying white flags, actions that constitute war crimes under international law. HRW.ORG These allegations have led to international scrutiny and legal actions. The International Criminal Court (ICC) has issued arrest warrants for Israeli Prime Minister Benjamin Netanyahu and former Defense Minister Yoav Gallant, accusing them of war crimes and crimes against humanity related to the Gaza conflict. The ICC alleges that Israeli leadership employed "starvation as a method of warfare" by restricting humanitarian aid and targeting civilians. APNEWS.COM
It's important to note that these allegations are subject to ongoing investigations, and the accused parties have denied wrongdoing, asserting that their military actions were aimed at neutralizing threats posed by militant groups operating within Gaza.
I think a closer comparison would be Android and GApps, where if you remove the latter, most would deem the phone unusable.
if open the source then it is open source
if you write a book/blog about how you came up with the ideas but didn't publish the source it's not open source, even if you publish the blog+binaries
If the magic values are some kind of microcode or firmware, or something else that is executed in some way, then no, it is not really open source.
Even algorithms can be open source in spirit but closed source in practice. See ECDSA. The NSA has never revealed in any verifiable way how they came up with the specific curves used in the algorithm, so there is room for doubt that they weren't specifically chosen due to some inherent (but hard to find) weakness.
I don't know a ton about AI, but I gather there are lots of areas in the process of producing a model where they can claim everything is "open source" as a marketing gimmick but in reality, there is no explanation for how certain results were achieved. (Trade secrets, in other words.)
To my understanding, the contents of a .safetensors file is purely numerical weights - used by the model defined in MIT-licensed code[0] and described in a technical report[1]. The weights are arguably only really "executed" to the same extent kernel weights of a gaussian blur filter would be, though there is a large difference in scale and effect.
[0]: https://github.com/deepseek-ai/DeepSeek-V3/blob/main/inferen...
- Windows MetaFiles (WMF, EMF, EMF+), still in use (mostly inside MS Office suite) - you'd think they're just another vector image format, i.e. clearly "data", but this one is basically a list of function calls to Windows GDI APIs, i.e. interpreted code.
- Any sufficiently complex XML or JSON config file ends up turning into an ad-hoc Lisp language, with ugly syntax and a parser that's a bug-ridden, slow implementation of a Lisp runtime. People don't realize that the moment they add conditionals and ability to include or refer back to other parts of config, they're more than halfway to a Turing-complete language.
- From the POV of hardware, all native code is executed "to the same extent kernel weighs of a gaussian blur filter" are. In general, all code is just data for the runtime that executes it.
And so on.
Point being, what is code and what is data depends on practical reasons you have to make this distinction in the first place. IMHO, for OSS licensing, when considering the reasons those licenses exist, LLM weights are code.
Another note is that this may be the more ethical option. I’m sure the training data contained lots of copyrighted content, and if my content was in there I would prefer that it was released as opaque weights rather than published in a zip file for anyone to read for free.
IMO it should be considered freeware, and only partially open. It's like releasing an open source program with a part of it delivered as a binary.
- open-source inference code
- open weights (for inference and fine-tuning)
- open pretraining recipe (code + data)
- open fine-tuning recipe (code + data)
Very few entities publish the later two items (https://huggingface.co/blog/smollm and https://allenai.org/olmo come to mind). Arguably, publishing curated large scale pretraining data is very costly but publishing code to automatically curate pretraining data from uncurated sources is already very valuable.
It is a bit more permissive than Llama's it seems (no MAU threshold it seems).
But the recent DeepSeek-R1-Zero and DeepSeek-R1 have MIT licensed weights.
https://huggingface.co/deepseek-ai/DeepSeek-V3/tree/main https://huggingface.co/deepseek-ai/DeepSeek-R1/tree/main
So when we apply the same principles to another category, such as weights, we should not call things “open” that don’t grant those same freedoms. In the case of this research license, Freedom 0 at least is not maintained. Therefore, the weights aren’t open, and to call them “open” would be to indeed dilute the meaning of open qua open source.
So it seems to me that it's at least dubious if those restricted licences can be enforced (that said you likely need deep pockets to defend yourself from a lawsuit)
Not necessarily[0], it's a WIP, but: https://github.com/huggingface/open-r1
[0] Surely they won't end up with the exact same weights, but it should be possible to verify something about the model and approach
Yes training is left as an exercise to the user, but it's outlined in the paper, and a good ML engineer should be able to get started with it, cluster of GPUs not included
"3.2.2. Efficient Implementation of Cross-Node All-to-All Communication
In order to ensure sufficient computational performance for DualPipe, we customize efficient cross-node all-to-all communication kernels (including dispatching and combining) to conserve the number of SMs dedicated to communication. The implementation of the kernels is codesigned with the MoE gating algorithm and the network topology of our cluster. To be specific, in our cluster, cross-node GPUs are fully interconnected with IB, and intra-node communications are handled via NVLink. NVLink offers a bandwidth of 160 GB/s, roughly 3.2 times that of IB (50 GB/s). To effectively leverage the different bandwidths of IB and NVLink, we limit each token to be dispatched to at most 4 nodes, thereby reducing IB traffic. For each token, when its routing decision is made, it will first be transmitted via IB to the GPUs with the same in-node index on its target nodes. Once it reaches the target nodes, we will endeavor to ensure that it is instantaneously forwarded via NVLink to specific GPUs that host their target experts, without being blocked by subsequently arriving tokens. In this way, communications via IB and NVLink are fully overlapped, and each token can efficiently select an average of 3.2 experts per node without incurring additional overhead from NVLink. This implies that, although DeepSeek-V3 13 selects only 8 routed experts in practice, it can scale up this number to a maximum of 13 experts (4 nodes × 3.2 experts/node) while preserving the same communication cost. Overall, under such a communication strategy, only 20 SMs are sufficient to fully utilize the bandwidths of IB and NVLink.
In detail, we employ the warp specialization technique (Bauer et al., 2014) and partition 20 SMs into 10 communication channels. During the dispatching process, (1) IB sending, (2) IB-to-NVLink forwarding, and (3) NVLink receiving are handled by respective warps. The number of warps allocated to each communication task is dynamically adjusted according to the actual workload across all SMs. Similarly, during the combining process, (1) NVLink sending, (2) NVLink-to-IB forwarding and accumulation, and (3) IB receiving and accumulation are also handled by dynamically adjusted warps. In addition, both dispatching and combining kernels overlap with the computation stream, so we also consider their impact on other SM computation kernels. Specifically, we employ customized PTX (Parallel Thread Execution) instructions and auto-tune the communication chunk size, which significantly reduces the use of the L2 cache and the interference to other SMs."
It's definitely not the full model written in PTX or anything, but still some significant engineering effort to replicate, from people commanding 7-figure salaries in this wave, since the training code isn't open.
I wish Intel didn't kill the maxas effort back then (buying the team out and ... what ?) as they were going even lower down the stack.
It’s not and I called it [1].
We had three options: (A) Open weights (favoured by Altman et al); (B) Open training data (favoured by some FOSS advocates); and (C) Open weights and model, which doesn’t provide the training data, but would let you derive the weights if you had it.
OSI settled on (C) [2], but it did so late. FOSS argued for (B), but it’s impractical. So the world, for a while, had a choice between impractical (B) and the useful-if-flawed (A). The public, predictably, went with the pragmatic.
This was Betamax vs VHS, except in natural linguistics. There is still hope for (C). But it relies on (A) being rendered impractical. Unfortunately, the path to that flows through institutionalising OpenAI et al’s TOS-based fair use paradigm. Which means while we may get a definition (not exactly (B), but (A) absent use restrictions) we’ll also get restrictions on even using Chinese AI.
If you want to prioritize pragmatism, that every discussion of this includes a lengthy "so what open source do you mean, exactly?" subthread proves this was a poor choice. It causes uncertainly that also makes it harder for the folks releasing these models to make their case and be taken seriously for their approach.
We should probably call them "free to run", if the "it's cheap" connotation of "freeware" needs to be avoided. Or maybe "open architecture" to appreciate the Python file that utilizes the weights more.
Technically yes, practically no.
You’re describing a prisoner’s dilemma. The term was available, there was (and remains) genuine ambiguity over what it meant in this context, and there are first-mover advantages in branding. (Exhibit A: how we label charges).
> causing collateral damage outside the AI bubble, and is nothing like Betamax vs. VHS
Standards wars have collateral damage.
> We should probably call them "free to run", if the "it's cheap" connotation of "freeware" needs to be avoided. Or maybe "open architecture"
Language is parsimonious. A neologism will never win when a semantic shift will do.
Agreed, but I think it's worth lamenting the danger in that. History is certainly full of transitory calamity and harm when semantic shifts detach labels from reality.
I guess we're in any case in "damage is done" territory. The question is more about where to go next. It does appear that the term "open source" isn't working for what these folks are doing (you could even argue whether the "available" term they chose was a strong one to lean on in the first place), so we'll see what direction the next shift takes.
Sort of. We can learn from the example. Perfect is the enemy of the good.
It’s ambiguously open.
Then it’s never distributable and any definition of open source requiring it to be is DOA. It’s interesting, as an argument against copyright. But that academic.
Also I would say arguing that the model weights are just the "binary" is disingenuous, because nobody wants releases that only contain the training data and scripts to train and not the model weights (which would be perfectly fine for open source software if we argue that the weights are just the binaries), because they would be useless to almost everyone, because they don't have the resources to train the model.
I agree it would nice to know the details of their training, but, simply calling this drop an "opaque binary" is seriously underselling it no?
tldr: we're already on to the next model, don't expect anything else to get open sourced.
> I was just told that the amount of people there are too limited, and open-sourcing needs another layer of hard work beyond making the training framework brrr on their own infra. So their priority has been to open-source everything that is MINIMUM + NECESSARY to the community while pushing most efforts on iterating to the next generation of models I think. They have been write everything clearly in technical reports and encourage the community to engage in reproduction , which is the unique insight of the team as well I think.
Paxos isn’t open source just because you can read the paxos paper.
But I was merely adding the missing context using the (sorry) lingua Franca of AI.
Things like tweaking all the hyperparameters to make the training process actually work may be more tricky though.
It's more akin to scientific research where everyone is using the same molecules, but depending on the process you put the molecules through, you get a different outcome.
How is that different than the 0s and 1s of a program?
Assembly instructions are literally standard. What’s more, if said program uses something like Java, the byte code is even _more_ understandable. So much so that there is an ecosystem of Java decompilers.
Binary files are not the “source” in question when talking about “open source”
Google Spanner has a nice white paper but you wouldn't consider it open source, for example.
You can also edit binaries by hand.
Especially with dynamically linked binaries like many games.
Seriously, a set of weights that already works really well is basically the ideal basis for a _lot_ of ML tasks.
That's fundamentally the same thing though - you run an optimization algorithm on a binary blob. I don't see why this couldn't work. Sure, a neural net is designed to be differentiable, while ELF and PE executables aren't, but then backprop isn't the be-all, end-all of optimization algorithms.
Off the top of my head, you could reframe the task as a special kind of genetic programming problem, one that starts with a large program instead of starting from scratch, and that works on an assembly instead of an abstract syntax tree. Hell, you could first decompile the executable and then have the genetic programming solver run on decompiled code.
I'd be really surprised if no one tried that before. Or, if such functionality isn't already available in some RE tools (or as a plugin for one). My own hands-on experience with reverse engineering is limited to a few attempts at adding extra UI and functionality to StarCraft by writing some assembly, turning it into object code, and injecting it straight into the running game process[0] - but that was me doing exactly what you described, just by hand. I imagine doing such things is common practice in RE that someone already automated finding the specific parts of the binary that produce the outputs you want to modify.
--
[0] - I sometimes miss the times before Data Execution Prevention became a thing.
"Given specific inputs X and outputs Y, have a computer automatically find modifications to F so that F(X) gives Y" is a problem that's been studied for nearly a century now (longer, if relax the meaning of "computer"), with plenty of well-known solutions, most of which don't require F to be differentiable.
Isn't "operational research" a standard part of undergrad CS curriculum? It was at my alma mater.
This is like saying, hey, a regular binary executable is fine because I can edit it with hexl-mode.
Or should I say, it held water until few days ago.
Personally though, I never bought it. Saying that weights are the "preferred form of the work for making modifications to it" because a) approximately no one can afford to start with the training data, and b) fine-tuning and training LoRAs are cheap enough, is basically like saying binary blobs are "open source" as long as they provide an API (or ABI) for other programs to use. By this line of reasoning, NVIDIA GPU stack and Broadcom chipset firmware would qualify as open source, too.
I don't build my browser, it's too expensive, but the cost of building has nothing to say with how open the access to things. It'd be cool if the community could fork the project, propose changes and maybe crowdfund a training/build run to experiment.
But probably it is impossible for them to release the training data, as they have probably not made it all reproducible, but live ingested the data, and the data has since then chanced in many places. So the code to live ingest the data becomes the actual source, I guess.
https://github.com/deepseek-ai
More importantly, they spelled out their methodology in depth in a paper (the code/implementation is trivial in comparison to the methodology)
I'm very appreciative of what they've done, but it's open weights and methodology, not open source.
But China did distribute them with sharing-friendly terms, what is completely different from others, like Meta, and makes the name way less misleading this time.
Deepseek is a side project for a hedge fund.
Shorting NVIDIA & releasing everything including the source would have a high probability of being hugely profitable, with almost zero downside if it went unnoticed.
But it’s a consequence of your highly concentrated effort on the primary project (DeepSeek), not a side project.
If it was a public US company releasing Deepseek, and you shorted nvidia based on inside knowledge, presumably you'd fall foul of insider trading rules?
[1] https://www.artisana.ai/articles/leaked-google-memo-claiming...
No team or consensus behind it, so it’s not really news in my opinion. In fact, I’d be surprised if nobody in Google would believe in open source AI.
The take of course interesting and as an open source guy, I like it and hope he is right.
In the end it'll be the scale of the infrastructure itself that will make the difference.
The model's source code (the training data) is hundreds of GB and much harder to transfer. The compiling (training) process is also very costly. This is very different from the Linux case.
Only big techs have enough resources to make these things happen.
I like looneysquash's viewpoint about the definition of open source AI. You will need to have all parts involved open-sourced to make a model "open", not just the weights:
> The trained model is object code. Think of it as Java byte code. You have some sort of engine that runs the model. That's like the JVM, and the JIT. And you have the program that takes the training data and trains the model. That's your compiler, your javac, your Makefile and your make. And you have the training data itself, that's your source code.
> Each of the above pieces has its own source code. And the training set is also source code. All those pieces have to be open to have a fully open system. If only the training data is open, that's like having the source, but the compiler is proprietary. If everything but the training set is open, well, that's like giving me gcc and calling it Microsoft Word.
Another point being, who knows if they really have legal rights to use all that data for training.
Everybody is trying to come up with a money related reason for why they open sourced it but at the end of the day the people who made it are engineers and not buisnesspeople. DeepSeek is really freaking cool, and they wanted to show people the cool thing they did.
https://huggingface.co/blog/open-r1
They didn't just toss model weights over a fence, they shared exactly how to do what they did. They made a meaningful contribution that people are replicating with other models readily.
But are you opposing or confirming my comment, because the link you shared kind of proves the point I'm trying to make right. (and afaik they haven't reproduced the results as of yet, they need the code and the training data)
Afaik the traditional sense of opensource implies that the parts that are given can reproduce the results, in this case, it should be code and training data. Now it's more open weights with a recipe (paper).
I didn't say anything about the quality of the model, nor the contribution of the paper which both are very good.
I think in times where everything becomes wishy-washy feely "it's all the same anyway" we need to be very specific, because semantics does matter actually. It also devalues the original meaning of those words, where something meant its opensource so u had all the parts to compile to get the same results.
If I missed something lmk!
> I usually ignore people who say this because it guarantees they're not actually doing anything meaningful with these models
I don't know who they and these people are, but I work with llm models daily for our data ingestion pipelines.
But the reality is by giving us the model weights alone they'd already be providing an immense boon given how important being able to cheaply produce reasoning traces is for distillation. And they went even further and documented the exact path they took which in immense detail that's already being digested and extended on by the community. Releasing R1-Zero and explaining where it fits in the puzzle wasn't necessary and yet it helped leapfrog open attempts at reasoning models and will likely even influence future closed models from other providers.
They've given us a lot more than any closed source provider has when the leading closed source provider won't even provide thinking traces because of fear of competitive distillation.
-
Also "these people" are people who just write prompts and call a REST API so it doesn't matter if the weights are available or not. At most they might replay some requests to a different REST API that returns a finetuned model for them. Sounds like what you do?
I rely on models I've posttrained with custom vocublaries, run through AWQ specific to my downstream task, and that are being inferenced on with custom samplers specific to my downstream task. All things actually closed source models can't do.
In other words, I derive enough value from "open source" models to not quibble over definitions. I find the people who do bother quibbling aren't doing anything with all the capabilities unlocked so they're willing to argue about just how open things are, but ironically wouldn't do anything even if the models were open by their absolutist definitions.
I have never said _anything_ about whether there is or isn't any value. Ive mentioned correctly that it isn't opensource.
I'm sorry you feel triggered by this, but both things can exist at the same time. It can be useful, come with a recipe, and it can stil not be opensource.
> At most they might replay some requests to a different REST API that returns a finetuned model for them. Sounds like what you do?
Please refrain from ad-hominem statements.
What I am saying is that when we redefine existed well understood definitions for marketing purposes as it muddies the waters.
On top of that, what if there are true opensource models that both provide training data + training code + inference code what do we call them? Extra opensource?
> I rely on models I've posttrained with custom vocublaries, run through AWQ specific to my downstream task, and that are being inferenced on with custom samplers specific to my downstream task. All things actually closed source models can't do.
Great, please opensource them :)
You jumped straight to "you're triggered" when all I did was reject the idea that we should let people who aren't familiar with where value resides in the model pipeline get to define what parts of the pipeline need to be shared to count as open source.
The OSI got that and that's how they ended up not requiring exact datasets be shared in their OSAID: https://www.hpcwire.com/2024/11/06/osi-open-ai-definition-st...
-
My HF profile has over 50 post-tained models available for anyone to download by the way, and I've had sampler options upstreamed to multiple inference projects.
Your link is useful and turns both to this https://opensource.org/ai/open-source-ai-definition
and this discussion:
https://discuss.opensource.org/t/deepseek-r1-does-it-conform...
> Technically, R1 is “open” in that the model is permissively licensed, which means it can be deployed largely without restrictions. However, R1 isn’t “open source” by the widely accepted definition because some of the tools used to build it are shrouded in mystery. Like many high-flying AI companies, DeepSeek is loathe to reveal its secret sauce.
and
> You are right about the lack of data information for DeepSeek, which is a requirement from the OSAID.
The source of all the excitement is exactly how much they revealed, and I feel like that thread as a whole emphasizes why people who aren't deeply familiar with the pipeline should not get to define these things.
"Nick" claims the secret sauce is in something not provided, then posts an article that demonstrates the exact "secret sauce", seemingly not making the connection: https://www.interconnects.ai/p/deepseek-r1-recipe-for-o1
-
There is a lot of detail about the nature of the data used and the exact steps needed to reproduce their findings with your own data. They even provide R1-Zero to demonstrate things that might be dead ends just in case someone can continue them. That should be enough to satisfy any useful definition of open source.
Even in the same thread you linked:
> Just a curiosity, according to the Model Openness Framework from the Linux Foundation, DeepSeek-R1 classifies as an Open Model:
> https://mot.isitopen.ai/model/1143
At the end of the day this is as good as it needs to be for LLMs: By their nature a lot of data being used to train them cannot or should not be openly shared, but the shape and motivations behind the data used are able to push others very far along the way to reproduction and iteration.
b) DeepSeek is the most dangerous thing that's happened to Open Source models in recent memory, through no fault of their own.
The hysteria has outrun the reality and now there's a going to be a similarly disproportionate backlash.
It's already happening: Anthropic's CEO simultaneously railing against what they achieved and using it to justify stronger export restrictions, this morning.
And our current government doesn't want to be going on stage talking about $50B mega projects only for laypeople to (mistakenly) believe it only takes a few million to do the same.
And the idea that a Chinese company is the one that did this is going to play into so many hands, so perfectly. You can see the censorship story start taking the narrative despite this not being the first or last Chinese hosted model to comply with Chinese law.
Soon the national security angle will break out, especially if someone jailbreaks or abliterates it and gets "harmful outputs" that other models would also happily produce.
Some will couch the (very temporary and irrational) dip the market faced as a Chinese company managing to harm our markets by providing an unfairly priced product or some nonsense.
Open source AI is not guaranteed. We might still see protectionist bans against releasing models over a certain size and other irrational nonsense, and this has played into the kind of hysteria that allows that to happen.
Always been fascinating to me how often rhetoric wins over substance on hn.
I would think the code to build your own is open sourced, and you can feed it any data you'd like. That's the open source part, not the part where they are running the model.
Have I misunderstood this?
I think it’s kind of an overdone complaint and I usually ignore it, and besides it looks like there’s a huggingface project ongoing where they’re trying to replicate the training process for this model anyway.
Holding exabytes of data to be processed on commodity hardware to enable internet-wide search, all the while it was man-in-the-middle monetised by an ad-business, created tremendous moats. Entering that market is limited to tech multinationals, and they have to deliver a much superior experience to overcome them. To perform a google search you need google-sized data-centres.
Here we have exactly the opposite dynamics: high-quality search results (/prompt-answers) are as-of-now incredibly commodotized, and accessible at-inferecence-time to any person who has $25k. That's going to be <= 10k soon.
And innovation in the space has also gone from needing >1Bn to <=50Mil
A higher quality search experience is available now at absolutely trivial prices.
Right now, automated knowledge gathering absolutely wipes the floor with automated bias. Cloudflare has an AI blocker which still can't stop residential proxies with suitably configured crawlers. The technology for LLM crawling/training is still mostly unknown, even to engineers, so no SEO wranglers have been able to game training data filters successfully. All LLMs have access to the same dataset - the internet.
Once you:
1. Publicly reveal how training data is pre-processed 2. Roll out a reputation score that makes it hard for bots to operate 3. Begin training on non-public data, such as synthetic datasets 4. Give manipulated data a few more years to accumulate and find its way into training data
It becomes a lot harder.
Not to mention countless other popular apps that Google has. YouTube anyone?
They’re also the most well positioned company to profit from cheap AI, their ads network is a behemoth.
So yeah, add that up with the compute, the data, and the talent, and it’s pretty clear that Google is not a force to dismiss.
If anything I think DeepSeek is great news for Google.
Not to mention, LLMs are way better at synthesizing multiple sources into coherent response. I end up asking and LLM then searching only as secondary research.
Totally on the same boat. Information is just much harder to find and friction becomes higher. I'd rather deal with the occasional hallucination than with the utterly enshittified SERP experience.
That's the real advantage of the open deepseek weights. They cannot enshittify this, you can run it locally. Just with an old snapshot
They stand to tap into something far more powerful than advertising if they can position themselves as your agent.
Just because R1 was trained cheaply, doesn't mean that this architecture cannot be trained on a very expensive data center to get much better and bigger models.
The full R1 is huge (~700GB), altough there are still quantized versions, the smallest one is around 150gb (1.58bit)
Its ability to solve basic math problems with reasoning is pretty cool, but other models of that size (qwen 2.5, phi4) have been generally more useful to me.
These tiny models still strike me as toys, not a whole bunch of real-world utility.
OpenAI's CTO, Mira Murati, found herself in a tight spot when questioned about using YouTube data to train Sora. Her uncertain response has sparked controversy and raised concerns about their ethics in collecting and training data. This incident has fueled a growing debate about AI companies' data practices.
Then YouTube's CEO, Neal Mohan said, if OpenAI used YouTube content without permission, it would violate their terms of service. Shall Neal freakout like how they are now!! Clearly they are scared, they know people are canceling their subscriptions with them to and use free and better technologies. I know of 100 of people canceled their gpt subscription. Many developers are replacing the expensive gpt models for free deepseek.
Here is the AI current story:
Imagine two AI trains chugging along the tracks of innovation. The first, driven by OpenAI, was the early leader, after they using Google transformers (and without they wouldn't exist). They charged a hefty fare for anyone to hop aboard. We don't know how they trained their data. And big companies felt they had to buy tickets or risk being left behind. OpenAI thought they were the only engine in town. But then, another train pulled up alongside them. This new locomotive, powered by smart folks at DeepSeek, matched OpenAI's speed and fancy gadgets, if not better. The kicker? Everyone could ride for free!
Now, OpenAI's train is losing steam. People are jumping ship, with hundreds canceling their pricey GPT subscriptions. Meanwhile, the free train is picking up speed, aiming to make AI available to all.
In this tale of two trains, OpenAI might need to change their name to "ClosedAI" if they keep putting up barriers, being closed. The free and open train? That's the one chugging towards a brighter, better, free AI future for everyone.
deepseek = Open AI
Truly remarkable model is DeepSeek-R1, and it's their model, with very particular DeepSeek architecture. Of course they build on the knowledge of other labs, just like other labs build on the top of their/others knowledge. They are miles ahead of Meta in terms of the base architecture at the moment, and you can watch them iterating throughout last year to come to where they are now.
Did you make a poll or something?
This answer itself is an AI product, right? Like you're making a meta-point about something
Frontier AI model SaaS companies like OpenAI can never win the race to zero against $0 free or open source AI models as they are already at the finish line.
METAs bet paid off, but at what cost.
> In fact, making it easier and cheaper to build LLMs would erode their [OpenAI, Meta, Google etc] advantages!
The narrative until now was: AI requires enormous and cutting edge resources (money, energy), so only for the big boys and people who can talk multi-billions investments, so open source was not an option.
Some signs already appeared recently (plateau, bubble?), and Deepseek seems to show that this model is questionable.
Exactly how it happened in operation systems.
Do we know that this is actually true?
But that doesn’t mean smaller models aren’t useful.
This thief has small pockets, about 500x smaller than the "stolen" material. Where to stash all that?
https://arjancodes.com/blog/python-pickle-module-security-ri...
What you can run locally are the distilled models, that is actually LLama and Qwen weights further trained on R1's output
I'm not entirely sure if it is possible to do some type of code execution like that in just the weights themselves, though someone else who knows a bit more about this can weigh in here.
Deepseek is open source because the founders are part of the new generation of Chinese graduates who relate more to the global youth than Boomer Chinese CEOs completely out of touch. And right on time because CCP is fed up with them, too.
Last week Deepseek founder Liang Wenfeng was speaking practically face to face with Chinese Premier Li Qiang at a symposium: https://www.youtube.com/watch?v=zMyc3vhpLyI. And they seem to be quite aligned.
Why didn't this blogspam of an article pick up on any of that?
``` <script async="" src="https://cdn.getkoala.com/v1/pk_963cd5673bdab99d6452d82210e66... ```
Don't forget to drink your Ovaltine.
Perfectly stated.
The AI jump to conclusions mat is so worn down, it's become paper thin. The shock of DeepSeek's costs does not auto-magically force all LLMs to become opensource. Silicon Valley tech has always favored whomever delivers inside the trifecta of cheaper, better, faster triangle. Anyone with an MBA should know this includes open-source LLMs. As of today, DeepSeek is ahead. As soon as OpenAI answers with a new 'fastfood dollar menu' for ChatGPT, with 'even more special' secret-sauce ingredients, we're going to see them back to normal business.
> Compare $60 per million output tokens for OpenAI o1 to $7 per million output tokens on Together AI for DeepSeek R1.
Open-source isn't a primary rational aspect to prove anything. We can't even prove what's going to happen tomorrow, proving a statement that is linked to the future is utter nonsense.
- Training SW [x]
- Inference SW [x]
- Evaluation SW [x]
- Data [x]
Output:
- Weights []
DeepSeek is closed-source with *open-weights*
The parenthetical should really be "(temporarily, approximately)". I wouldn't count OpenAI out until we see how o3 compares, assuming they actually make it available this week.
I'm with you. How are they going to make money?
Lot of them on the internet trying to help user with basics windows things, then they suggest their app as a better alternative.
> 10x top of HN
> Billing remains a major issue for companies, resonating widely. We've consistently hit HackerNews' top page over 10x. [1]
It's just an article that is aimed to get you to hear about Lago, star their GitHub repository and eventually talk about the "open source" billing tool you heard about called Lago.
(I put Open Source in quotes because I think it's open source version is just Freeware with most features being Call To Action to a paid version. Fair disclosure I have https://github.com/billabear/billabear which is a competitor)
Now, commercially, this may not make a lot of sense at the moment because no one is getting filthy rich on high-moat AI-using applications such that commoditizing the AI itself is a good idea commercially. I'm not sure anyone would even be confident enough to be the farm on the idea of someday being in that position.
However, if you analyze this from the perspective of world politics, where both explicit and implicit strategies are based on what tech companies have what tech and where it is located, it makes a lot of sense that if China is concerned that the US really is ahead in AI tech and that US financial and technical dominance is being driven by this dominance and being used to suck capital out of the countries that are behind, it makes all kinds of sense to commoditize the complements as basically a way of throwing the current game board up in the air and restarting again.
(One may also note that this analysis also says that just straight-up stealing the OpenAI tech and slightly AI-washing it before handing it out to everyone is also a logical move. I don't know enough to have any independent opinion as to whether that's where DeepSeek came from. I'm just saying that given the visible circumstances it is a strong strategic move for China at this point.)
Much of this purported strategy hinges on 'winning' at all cost by undermining the lead.
What is there to be won at the end? Does one party taking the reigns prevent the other from achieving similar capabilities? Is it necessary to win this race?
Or is this a cumulative, distributed effort that benefits all of us?
Edit: to be clear, what I mean is that to a first approximation technologists are charlatans and frauds. If you're looking for accurate information ask a scientist.
Why would you discuss gold with the wild-haired eccentric at the bottom of a hole, who has not yet found any gold, when you could talk to a gold mine owner who has - and who employs 1500 voters, and who like you wears a suit and tie?
This describes a narrow slice of Silicon Valley numpties.
World leaders see an economic opportunity. Both to spend and to produce. No politician will turn down the opportunity to announce half a trillion dollars of spending.
This. There are a few theories of geopolitics, one of the most successful being ones we be bunch under an umbrella called realism [1]. (The others are idealism [2] and liberalism [3]. Historia Civilis made a great three-part video series on these [4]. Note that Realpolitik [5], which relates to realism as its praxis, is not the same thing.)
One of the consequences of realism is balance of power theory, which “suggests that states may secure their survival by preventing any one state from gaining enough military power to dominate all others” [6].
What is to be won? Not being dominated; ideally: less war, since war is irrational. (See: Ukraine.) Does preventing others from dominating you prevent you from dominating others? No. Is it necessary to win? No. But that means ceding sovereignty and increasing the chances of violent conflict as geopolitical fault lines reälign.
A note on liberalism: it works. But it requires great power at its centre. America was that benevolent great power. Now it seems we don’t want to be. The power America has to hurt its allies, and the incentives to reap that advantage, is the consistent failure mode of liberal foreign-relation structures, since the days of the Delian League.
[1] https://en.m.wikipedia.org/wiki/Realism_(international_relat...
[2] https://en.m.wikipedia.org/wiki/Idealism_in_international_re...
[3] https://en.m.wikipedia.org/wiki/Liberalism_(international_re...
[4] https://youtu.be/CH1oYhTigyA
[5] https://en.m.wikipedia.org/wiki/Realpolitik
[6] https://en.m.wikipedia.org/wiki/Balance_of_power_(internatio...
They literally had emperors who banned all overseas travel because it represented a threat to their own power: https://en.m.wikipedia.org/wiki/Haijin . China is the extremely large and extremely centralised, so the rulers' primary focus has always been on maintaining their own power. Fortunately the current government still allows private firms enough freedom that one was able to invent DeepSeek, however if the recent crackdown on financial firms had happened a few years earlier then the firm behind DeepSeek wouldn't have had the money to fund its creation.
They did. Just as a land power. Modern China includes conquered territory of the Mongolians, Turkics and Tibeto-Burmans, among others [1].
(The proximate answer is the Ming-Qing transition [2] overlapped with the Age of Discocery [3].)
> except for China and the US, no one cares who the product comes from
This is breathtakingly wrong, as a simple perusal of every single country's trade restrictions would show. (Even if you're talking about the population versus policy, show me a market where no premium is paid for luxury products imported from such and such distant land.)
[1] https://en.wikipedia.org/wiki/List_of_ethnic_groups_in_China...
[2] https://en.m.wikipedia.org/wiki/Transition_from_Ming_to_Qing
maybe it's offtopic, but that's what I'm good at, so I'll anwser this
First, ancient China was a feudal centralized dynasty that centered its interests on land and population, unlike commercial company-based regimes such as Britain and the Netherlands. This meant that, in the eyes of the Chinese imperial government, the East India Company was a threat rather than a cooperative partner.
Another reason is that ancient China was a typical land-based power, surrounded by various forces. It could only maintain its sphere of influence through annexation and the tributary system, without the ability to expand further. (Genghis Khan was the only exception—he carried out invasions but never truly established effective rule.)
However, ancient China did, to some extent, "colonize" certain Southeast Asian islands. But this was not institutionalized colonization; rather, it was a form of population migration. The central government had no control over these Chinese people venturing into the seas, which is why it repeatedly tried to prevent maritime expansion.
btw, in case someone said about xinjiang and tibet, you'll see he don't understand history outside the west, base on what i said, you can see it was annexation but not colonization
Only european powers had the urge for colonization, no other civilization in Americas, Africa or Asia really ever want to colonize, expand perhaps but not really colonize.
There was no economic need to do so, for most of last three millennium the economic center of the world has been India and China , they didn’t feel the need to go anywhere , the land is fertile with large local population and good weather to grow more than one crop with rich cultural heritage and throughput there is no payoff for undertaking risky voyages.
Everyone wanted to trade with them, colonial powers bombed ports forcing trading agreements or sold opium and other narcotics to get a foothold, funded expensive expeditions for new trade routes to India and colonized another continent instead , most of era of industrial revolution have been focusing on them as the market for European products not merely resource extraction.
Similarly given the people resources both regions had, there was no need for slavery that is also a european/Mediterranean thing primairly .
Not saying workers were or are treated well or there was great value for human rights in India or China, just that they need to go and find slaves from far off to do the work. They could find all the resources domestically.
Barley warrants a response but
After the sengoku jidai[1] the failed imjin wars under Toyotomi Hideyoshi was the only serious attempt to expand to China and Korea, they of course failed and Japan faced inward till Meiji period as was typical of most of their history
Post Meiji restoration is hardly a fair comparison the Japanese believed that they have to be like other world (colonial) powers to be powerful.
[1]Unrelated note: one of my favorite periods in history.
I'm not. And I didn't think you were either.
But if we are, the Arab colonization of the Middle East + North Africa has to rank among the most dominant of all time, yes? Still apparent to this day.
Let me put another way- sub Saharan African, Chinese, southeast/far-east Asian, Indian, North/South American(first nations), Polynesian empires etc largely did not do empire building via colonization or slavery.
This is not to say they valued human life or did not commit atrocities, it just means that economic models that necessitated colonies for resource extraction or large markets to sell to, or foreign slaves for human labor never evolved there influenced to environmental, population and cultural factors so colonization is atypical response when the empires are built there.
We can see observe difference today in how say China deals with foreign investments, loans and other development initiatives compared to how western powers do. The deals tend to be primarily economic with willingness to work with existing regimes and less non-economic conditions attached and so on.
You're describing two modern states that encompass geographies that were constantly at internal turmoil. (Including as empires [1].) It's like asking why the Germans were late to the game in colonising: they're a land power and were in a constant state of internal turmoil.
"They had enough" flies in the face of human history and European colonialism itself.
[1] https://en.wikipedia.org/wiki/List_of_Hindu_empires_and_dyna...
Neither region is a utopia in history or today, simply there was enough land and people and other resources within, so they viewed their region to be the world, there was no economic impetus to colonize or enslave from far off places is my point .
* Inca Empire: Relocated entire communities (the mitmaqkuna) into new provinces to cement imperial control—these were explicit colonies with an imposed administrative and cultural framework.
* Ancient Egypt: Occupied Nubia, built forts, stationed garrisons, and imposed Egyptian officials and religion on the local population.
* Mongol Empire: Installed governors across conquered regions stretching from Eastern Europe to East Asia, moved artisans and workers to bolster Mongol centers, and demanded tribute—hallmarks of a colonial system.
* Imperial China: Established commanderies in newly acquired territories (e.g., southern China), encouraged Han settlement, and superimposed its bureaucracy over local governance.
Historians do not consider mongol or Inca empire colonial . I would say mongols were probably polar opposite of colonizers they were extremely open and integrated extremely well into every region culture they occupied, there was no classical markers of colonization.
I specifically added Mediterranean later in my parent post to cover Egypt , Phoenician and Arab colonization which are considered as examples of pre modern era colonizing.
The hard separation of North Africa is sadly a modern view of the region that I have to do that explicitly, for most of history empires always had some land on both sides of the Mediterranean. This view is either promoted and exploited by far right in southern europe to justify many policies.
Genocides in Americas, Australia and elsewhere of first nation people notwithstanding i suppose
> laws, ethics and technology we gave them
Unasked and unwanted "civilizing" by European powers is what got us Congo Free State and dozens of other atrocities all under the name of "civilizing". It is not like rest of the world was living in trees with no laws and morality.
> Most slaves in history has been Korea
This is a controversial view of Korea, there is no consensus if nobi and the class system during the Joseon period (much less so in Goreyo period) was serfdom or slavery, that is not easy classification to make, given that they had many rights, many earned salary, nobi women in 1400s got 100 days maternity leave by law, a lot more than modern American women do today.
Even if we take assume they were all slaves, Korea was by no means the leading country by % of population, and also we have to consider nobi were largely ethnic Koreans, not foreigners explicitly captured to be slaves and the economy didn't run on continuous capture of foreign slaves
> us accepting millions of immigrants desperate to either live with us or copy us and then tell us how much greater their own societies are, will end soon. You are free to go and live there with your own people.
While there is a discourse to be had socio-economic policies in the west from repatriation of cultural artifacts, to climate change or geopolitics that can stabilize the global south and reduce immigration, at this point I have to stop engaging.
Well sure but generally speaking proofs about reality are an oxymoron, so who on earth was taking the headline at face value to begin with? This is a rhetorical technique referred to as "hyperbole".
Proofs about reality are not self contradicting. Something not being entirely correct doesn't an oxymoron make.
Absolutely they are! Proofs are a deductive concept with no basis in reality. This is basic Hume. All we can work with is inductive and abductive reasoning, neither of which is sufficient for a proof.
One, it's not. Two, you're trying to use Hume to prove a statement that refutes itself. The claim that your can prove proofs oxymoronic is itself an oxymoron.
Hume's critique of causation, moreover, has been amply supplanted since the 18th century. (Similar to Newton. In parts, it's been buttressed. In others, surpassed.)
> All we can work with is inductive and abductive reasoning, neither of which is sufficient for a proof
Mathematically false [1]. (And related to famous Gedankenexperiments, which prompted real science.)
Of course, this whole thread is a farce: you're purposefully confusing mathematial proofs with the colloquial "proof."
Ok, this has no bearing on our empirical reality.
(Submitted title was "DeepSeek proves the future of LLMs is open-source".)
Honestly would that part even be useful? Like I want to know how they did the training so I can repro it with my own set of training data, right?
I mean, isn't that the future? Somebody figures out how to do P2P distributed training and groups can crawl the web training their own open source models?
> The trained model is object code. Think of it as Java byte code. You have some sort of engine that runs the model. That's like the JVM, and the JIT. And you have the program that takes the training data and trains the model. That's your compiler, your javac, your Makefile and your make. And you have the training data itself, that's your source code.
> Each of the above pieces has its own source code. And the training set is also source code. All those pieces have to be open to have a fully open system. If only the training data is open, that's like having the source, but the compiler is proprietary. If everything but the training set is open, well, that's like giving me gcc and calling it Microsoft Word.
How do you propose to opensource terabytes of web scrape text? They give you what they can give you - paper, code, model weights. You can reimplement the code, while the weights are open to do what you like with them.
We have had a term to describe this kind of software for decades: "freeware." That's what this and all other "free to download and use" offerings are; they are not open source under any commonly-understood meaning prior to last year.