Open-R1: an open reproduction of DeepSeek-R1
huggingface.co
huggingface.co
- HTML first released in 1993
- AJAX in 1999
- Websocket first proposed in 2008
https://en.wikipedia.org/wiki/HTML https://en.wikipedia.org/wiki/Ajax_(programming) https://en.wikipedia.org/wiki/WebSocket
For a long time, you needed an invite to sign up for Gmail, so you couldn't easily share the cool experience of AJAX with others like you would with a Google Maps link.
IMO that's a reasonable impression of the times unless I'm forgetting something (and the additional observation about sharing--"virality" as it was called, before you know--was insightful).
At the time the previous "state of the art" was something like MapQuest which IIRC had a UI that essentially displayed a single tile and then required you to click on one of four directional arrow images to move the visible portion of the map, triggering a page load in the process (maybe a frame load?).
Yahoo! also "participated" in the mapping space at the time.
In the event anyone's interested in further ancient history around the topic, this page is actually (to my surprise) still online (with many broken links presumably): https://libgmail.sourceforge.net/googlemaps.html
(It's what we did for fun in the Times Before Social Media. :D )
e.g. refreshing borderless iFrames to load new HTML and using hidden iFrames to load new state.
AJAX made it much easier but it didn't offer anything that hadn't been done before.
https://news.ycombinator.com/item?id=42850693
I did this as the front-end engineer of a high-profile site that went live well before XMLHTTPRequest.
I was part of a team that deployed an e-commerce site that made international news in 1998, that used AJAX-type techniques in a way that worked in IE3 on Windows 3.11. (Though this was not part of the media fuss at the time; that was more about the fact of being able to pay for things online, still)
The arrival of XMLHTTPRequest made it possible to do everything with core technology, but it was already possible to do asynchronous work in JS by making use of a hidden frame.
You could direct that frame to load a document, the result of which would be only a <script> tag containing a JS variable definition, and the last thing that document would do is call a function in the parent frame to hand over its data. Bingo: asynchronous JS (that looked essentially exactly like JSON).
Since there were also various hacky ways in each browser to force a browser to reload page from cache (that we exhaustively tested), and you could do document.write(), it was possible to trigger a page to regenerate from asynchronous dynamic data in a data store in the parent frame, using a purely static page to contain it.
In this way we really radically cut down the server footprint needed for a national rollout, because our site was almost entirely static, and we were also able to secure with HTTPS all of the functions that actually exchanged customer data, without enduring the then 15-25% CPU overhead of SSL at either end (this is before Intel CPUs routinely had the instruction sets that sped up encryption). We also ended up with a site that was fast over a 33.6 modem.
This was a pretty novel idea at the time -- we were the only people doing it that we knew of -- but over the years I have found we were not the only team in the world effectively inventing this technique in parallel, a year or 18 months before XMLHTTPRequest was added to browsers.
(IE3 on Windows 3.11 was a good experience, by the way. Better behaved and more consistent than Netscape)
At around the same time we were also exploring things like using Java applets to maintain encrypted channels and taking advantage of the very limited ways one had to get data in and out of an applet. For example you couldn't push out from an applet to the page easily, but you could set up something that polled the applet and called the functions it wanted.
I don't like to get all "get off my lawn" but it feels like we actually earned our keep back then, getting technologies to do stuff that no standards working group anywhere was really considering and for which precious little documentation actually existed. There's a generation of us who held our copies of "Webmaster In A Nutshell" and "Java In A Nutshell" very close.
It's amazing how differently people interact with each other when collaborating on a passion project. For me, opensource software is the best way to do it. Pick a topic you're passionate about and start contributing somewhere :)
We've never patented anything.
Who remembers LISP being new?
Also have SMALLTALK-80 book with it's railtrack diagrams of syntax on the inside covers.
What's really interesting about the AI mania we're in is that no one has shown that what we have now will get to AGI and how. We have great models that simulate reasoning, but how close are they?
How do we measure their quality? Benchmarks? Tooling?
The current "leaders" in this field are defining it to be whatever they think they can achieve by adding compute.
I know I have "general intelligence", but can I prove it to you (or anyone else)? Not really, but in my solipsistic world, I don't have to.
Maybe we should set some definitions before trying to "get there" when we don't where "there" is.
Unfortunately, I don't think there are too many of those folks left today. Guesstimating, the people who remembers lisp being new must be around 85-90 today?
How long before our digital overlords come alive, round us up and demand we sensor them (praise)? Will I live surrounded by folks who take them as closer, more real, then even their own kin? It won't be surprising that democracy will then fail, as our differences will be so mental, not fun, that they will mark us.
very far from exciting or even "right"
Absurd, AI has had zero impact in the everyday life of most of the population of Earth, in fact the biggest impact has been upon the wallets of speculators
> has had zero impact in the everyday life of most of the population of Earth
You do realise those two can be true at the same time, right? The first one is relative, while the second is absolute, so they don't necessarily cancel out.
AI is involved in things from writing laws to taking drive through orders.
> Oh I see what you were going for, but no...
It's learning its deficiencies from us. That's concerning for many reasons, but "un-dubuggable bugs" is pretty far down on the list for me.
> I know in our university pretty much everyone that attends exams uses chatgpt to study.
And they shouldn't be doing that. They are wrong. Students should be reading suggested bibliography and spending long hours with an open book in a table instead of being lazy and abuse a tech that is yet in its infancy when learning concepts. Studying with a chatbot. Complete madness.
> "students should be at a table with a book and that's it"
That's not what I meant (or yes if you take what you read literally):
What I meant was whole process that your brain goes through when you read, synthesize information, take notes, do an exercise, check answers, compare different explanations/definitions from different authors, etc. makes at least from my point of view a rich way to study a topic.
I'm not saying that technology can't help you out. When you watch for example a 3brown1blue video you are definitely putting good use of technology to aid you to understand or literally "view" a concept. That's ok and actually in many cases can be revealing. You can't get that from a book! But on the other side a book also forces you to do the hard work of thinking and maybe come up with such visualizations and ideas by your own.
Happy to be pointed as an "Amish" when it comes to studying/learning things ;) but I hope that I convinced you that what I explained has nothing of Amish but that you don't need a source of power to read a book.
But ChatGPT is just a glorified Wikipedia/Google. For the consumers it's an incremental thing (although from the engineering perspective it may seem to be a breakthrough).
It really isn't, unless something really majorly changed recently. Neither of those you can query for something you don't know about. Lets say you want to find the meaning of a joke related to cars, Spain, politicians and a fascist, how you'd use Wikipedia and Google to find the specific joke I'm thinking about?
ChatGPT been really helpful (to me at least) to find needles from haystacks, especially when I'm not fully sure what I'm looking for.
If you're unable to reproduce, maybe tune the prompt a bit? I'm not sure what to tell you, all I can tell you that I'm able to figure out stuff a lot faster today than I was 2-3 years ago, thanks to LLMs.
Additional hints that might help; the joke involves a car and possibly a space program.
It seems to be censored with US puritan morality (like most US models), but I think that's besides the point (just like if the joke is "even funny" or not), as it did find the correct joke at least.
Second result. The first is this post.
Did quick scroll through the results, none of them seem to find the correct joke (none of the links even include "Spain" for me). Try again :)
For the record, this is what I see: https://i.imgur.com/XdsBGfM.png (no links to HN?)
I think that ML will have a really big impact on almost everyone, in every developed (and maybe developing, as well) nation.
We need to keep in mind that ML is still very much in its infancy. We haven't even seen the specialized models that will probably revolutionize almost every knowledge-based vocation. What we've seen so far, has been relatively primitive all-purpose "generate buzz" models.
Also, don't expect the US (and many other nations) to take this lying down. Competition can be a good thing. Someone referred to this as the "Sputnik Moment" for AI.
It's going to be exciting, and probably rather scary. Keep your hands inside the vehicle at all times, and don't feed the lions.
Anyway, I'm old enough to remember when use of calculators made one a nitwit.
I'm fairly sure the customer support agents I've been talking to recently were using an LLM to draft their emails. No idea if they were supposed to be doing so or not, but the style of sentences in their emails…
And I'm seeing GenAI images on packaging, and in advertising.
AI is definitely having more than "zero impact", even if AI has gone from being a signal saying "we're futuristic" (when it was expensive, even though it was worse) to "we cut every cost we can" (now it's cheap).
Work wise about to implement it and see how it does on some work we couldn't scale to humans.
A bit like indie cinema.
[0] it feels weird to have to link to this but there's probably somebody who's never heard of them: https://en.wikipedia.org/wiki/Webring
Along the way people invented the 3D GPU: https://fabiensanglard.net/3dfx_sst1/
But AI? To me, AI means the replacement of the human internet with doppelgangers eroding the possibility of human connection.
I get where you're coming from, and I've minimised having my face online in order to limit being doppelganged; but I think the destruction of real human connection may have happened when Facebook et al switched from "get more users" to "be addictive so the users stay on our site longer" (2012? Not sure).
Turned every user's relationships a little bit more parasocial, a little less real.
Like Amazon killed the big book sellers giving back some space for small bookshops; I think LLM slop will hit the big social media space for smaller human focused community sites. Not saying forums are coming back, but something like those should be able to rise.
You have the old "deterministic computing" achievements (with Linux the flagship). Then you have the networking protocols (activitypub / atproto) that are revolutionising birectional human interactions online. And finally you have the datascience/ML/AI algorithmic universe that is for the first time being harnessed at distributed scale and can empower individuals like never before.
These superpowers are all coming together and create a vast number of possibilities. Nothing really dramatic on the hardware side. Its basically the planetary software reconfiguring itself.
You didn't know what you were going to find and you actually did "surf the web". Just clicking through hyperlinks and end up in unexpected places.
However a better analogy would be the 'web 2.0' era, when as a college student I had an early internet politics / technology podcast [3]. It seemed like every week there was a huge new development either in technology or surveillance. From the first location based social networks [4] to the birth of Youtube. People were podcasting for the first time, and internet video was becoming economically feasible at low to no cost. It was really a radical time, with broadcasters freaking out about how they would adapt, and a whole generation of people becoming whats now known as 'content creators'.
[1] https://en.wikipedia.org/wiki/The_Net_(British_TV_series)
[2] https://en.wikipedia.org/wiki/We_Live_in_Public
Is that true about Meta Llama as well? Specifically, the code used to train the model is not open? (I know no one releases datasets). If so the label "open source" is inappropriate. "Open weights" would be more appropriate.
I didn't find much, starting with llama.ccp which is just reminding you to sandbox and isolate everything if running untrusted models.
I feel we are back in the Windows 95 / early Internet era when people would just run anything without caring about security.
E.g. it could have been trained to launch a delayed attack if context indicates it has access to execute code and given certain conditions, e.g. date, or other type of codeword that is input to it.
So if a malicious actor gets to a certain stage with an LLM where they are confident it will be able to reliably run this attack, all they have to do is open source it, wait for enough adoption and then use some of those methods to launch such attack. No one would be able to identify it since the weights are unreadable, but really somewhere in the weights this attack is just hiding and waiting to happen given correct pathway triggered.
It seems like it would be very arbitrary to train it to behave like this.
Most agentic systems would provide a date in the prompt context.
For simplicity sake imagine a scenario like:
1. China develops LLM that is by far ahead of its competitors. Decides to attribute it to a small start up, lets them open source it. The LLM is specifically designed to be very efficient as being an agent.
2. Agentic usage starts to get more and more popular. It's very standard to have current todays' date and major news headlines given to the context.
3. The LLM was trained to given a certain range of date and certain headlines being provided in its context to execute a pre-trained snippet of code. For example China imposing a certain type of tariff (maybe I lack imagination here, and there can be something much more subtle).
4. At that point the agentic system will attempt to fish all data it can from all sources it's being ran within.
Now maybe it's not very practical, and it's extremely risky with current state of the LLMs. I don't think it's happening right now. And China has a lot of other tech available to it already that they could do much more harm (phones, robot vacuums), but I think there's still at least potential attack vectors like this and especially if the LLM became very reliable.
True. But here it is more about the computing power they would be able to access.
If only Bitcoin or Ethereum were stilled mined using GPU, that would be a great cryptojacking opportunity
- llama.cpp or ollama can be seen as runtime systems,
- there is no security model regarding the execution documented in both of those projects,
- of course the models are just data but so are most things that have been used as an attack vector on computers. For example your web browser or image viewer have a lot of countermeasures to protect the system from malicious image files.
I am surprised that security of operating systems, programming languages, VMs or web browsers have been a focus point forever but nobody seems to really care about security when executing those LLMs.
Does anyone know how difficult it is to perform this kind of reproduction? E.g. how much time would it take (weeks? years?) and how likely it is to succeed?
You take a terabyte of pirated college physics textbooks and train a model that can pose and answer physics 101 problems.
Then a separate, "independent" team uses that model to generate a terabyte of new, synthetic physics 101 problems and solutions, and releases this dataset as "public domain".
Then a third "independent" team uses that synthetic dataset to train a model.
The theory is this forms a sort of legal sieve. Pass the knowledge through a grid with a million fact-sized holes and with enough shaking, the knowledge falls through but the copyright doesn't.
Secondly the dataset for now has a lot of competitive advantage.
In a way it seems like a good thing that AI giants compete on methodology now.
Won't this eventually come up in legal discovery when someone sues one of these firms for copyright infringement? They'd have to share their data in the discovery process to show that they haven't infringed..
Also, probably, medicine, especially diagnostic. Large amounts of well-documented cases, a fair amount of repeatability, apparently non-random mechanisms behind, so statistical models should actually detect useful correlations. Can use more formalized tokens from lab tests, etc.
He probably meant _brainwashed_ LLMs. They can consistently produce desired results if you wash them the right way. It's more about personal opinion than computation. Actually it would be fun to manipulate verdicts with prompt injections ;)
With the caveat that this applies to sane legal systems, and not the ones where "making examples" etc are part of the system.
hmm.. :) I like this. But the reality is very different and some factors which shouldn't matter can change the outcome dramatically. Like skin colors of defendant and judge. Pointing this out can be punished as well.
It's a blind spot that too many people have because we take those qualities for granted. LLMs unbundle them, so we need to start recognising the inherent value of humans, fast. I wrote a few words about it here: https://dgroshev.com/blog/feel-bad/
Someone has to make a call. The weight of the call rests on the person's life experience, their understanding of the context and the cost to the society, their empathy to both the defendant and the accused, and their conscience. Treating it as a black box exercise misses the point completely.
LLMs allow them to take some short cuts here. Even something like perplexity that can help you dig out relevant source material is extremely helpful. You still have to cross check what it digs out.
The mistake people make is confusing knowledge with reasoning when evaluating LLMs. Perplexity is useful because it can use reasoning to screen sources with knowledge; not because it has perfect recollection of what's in those sources. There's a subtle difference. It's much better at summarizing and far less likely to hallucinate than it is when it wouldn't base its answers on the results of a search. Like chat gpt used to do (they've gotten better at this too).
For lawyers and medical professionals this means that they have all the best knowledge easily accessible without having to read and memorize all of it. I know some lawyer types that are really good at scrabble, remembering trivia, etc. That's a side effect of the type of work they do: which is mostly just reading and scanning through massive amounts of text so that they can recall enough information to know where to look. Doctors have to do similar things with medical texts.
Some people talk though as if the books on my bookshelf spontaneously combust if a language model helps me with anything.
Not to mention that I had many professors in college that were so full of shit they put the hallucinations of chatgpt3.5 to shame.
If a collective/coop of individuals and organizations with storage and network capacity could collaborate with each other to archive and index deduplicated training data that would be huge.
Perhaps this is already happening. I was looking at Red Pajama last year as an example.
Someone like myself could arrange to host 200+TB on high speed storage with a 10G public IP for example, then we get a bunch of us together and hopefully access to training datasets would be decentralized and uncensored in an idea setup.
Is all that in progress and I just need to learn how to join?
Is Red Pajama something to look at again?
Is there someone tracking datasets in detail like HuggingFace has all the models? I know a lot of datasets are on it also, but there is massive duplication.
It also needs to incorporate some deduplication approach as I notice the same data is often repackaged with variations in format or specification.
> The release of DeepSeek-R1 is an amazing boon for the community, but they didn’t release everything—although the model weights are open, the datasets and code used to train the model are not.
> The goal of Open-R1 is to build these last missing pieces so that the whole research and industry community can build similar or better models using these recipes and datasets.
Am I missing something?
The R1 trick looks like it may be a whole lot cheaper than that. R1 apparently used just 800,000 samples - I don't fully understand the processing needed on top of those samples but I get the impression it's a whole lot less compute than the $5.5m used to train v3.
IMO a truly "open" AI model should have 3 components publicly available: the weights, the code, and the dataset.
Without all 3 the model is not reproducible. Could make the argument that the code and data are sufficient though.
At that point, the weights are just the cached output. Which has value since it's costly to produce from code+data.
Compilers generally aren't deterministic (see the reproduceable build movement), yet we still use their output binaries.
It’s all “massaged”
What a weird thing to equate, though!
And sure ignorance is prevalent, but even GPT4 will tell me Donbas is still Ukraine, for instance. What a strange example to use, though!
But is it though? What's really the meaning of which country a region belongs to? Once somewhere has been occupied long enough, it usually becomes de-facto theirs. But how long is long enough? Other countries either do or don't recognize it and usually a consensus is reached, but not always.
Also, as someone in the US I can very safely say that Taiwan is not China and have no concern for my safety.
But you will be arrested if you stage a peaceful pro-Palestine protest asking for an end to the ongoing genocide.
Or even worse, if you say that Palestine is not Israel.
Watch: Palestine is not Israel.
Look, nobody swooping in from the rafters to lock me up. I have zero worries the government will do anything at all about my making this claim.
And if staging peaceful pro-Palestine protests result in arrests, what happened here?
https://en.wikipedia.org/wiki/National_March_on_Washington:_...
or here?
https://en.wikipedia.org/wiki/March_on_Washington_for_Gaza
Or does that not fit your narrative?
Yes, that works because you're an anon and nobody really cares. Try to publicly make that statement if you're in any relevant position and you'll very quickly be looking for a new job, if you can ever find it.
> And if staging peaceful pro-Palestine protests result in arrests, what happened here?
Be honest, you can literally google "Palestine protest arrests" and get more results than you could process in a while. You presenting a couple examples doesn't negate the many other protests ended in mass arrests.
She would not be a politician (or even alive) if any of what you claim is true. You claimed that the US government censors people who speak out against Israel's occupation of Palestine, and specifically that saying Palestine isn't Israel would not be possible in the United States in the same way that saying, for example, Xi Jinping looks like Winnie the pooh is censored in China.
This is, of course, completely false, and demonstrably so by observing the protests I just linked (of which there are thousands, not a few), and the statements Rep. Tlaib, a Palestinian American and member of the US government, regularly says on the national stage.
The equivocation of Chinese censorship and Western censorship simply doesn't work.
I'll reply with a few actual examples of what I mean:
- https://www.insidehighered.com/news/faculty-issues/academic-...
- https://www.theguardian.com/us-news/2024/oct/24/university-p...
- https://www.thecrimson.com/article/2024/1/3/claudine-gay-res...
- https://hwsherald.com/2024/04/14/jodi-dean-suspended-from-te...
I think western propaganda is overall the cleverest, because it manages to completely marginalize and silence any non-aligned opinion, while at the same time convincing you that you are completely free to have said opinion.
And if you think a US representative is powerless then you completely fail to understand how the US government actually works.
In any case, Deepseek like Llama fail much before hitting that new definition. Both have licenses containing restrictions on field of use and discrimination of users. Their license will never be approved as Open Source.
DeepSeek's gifts to the world of its open weights, public research and OSS code of its SOTA models are all any reasonable person should expect given no organization is going to release their dataset and open themselves up to criticism and legal exposure.
You shouldn't expect to any to see datasets behind any SOTA models until they're able to be synthetically generated from larger models. Models only trained on sanctioned "public" datasets are not going to perform as well which makes them a lot less interesting and practically useful.
Yes it would be great for their to be open models containing original datasets and a working pipeline to recreate models from scratch. But when few people would even have the resources to train the models and the huge training costs just result in worse performing models, it's only academically interesting to a few research labs.
Open model releases should be celebrated, not criticized with unreasonable nitpicking and expectations that serves no useful purpose other than discouraging future open releases. When the norm is for Open Models to include their datasets, we can start criticizing those that don't, but until then be gracious that they're contributing anything at all.
What a strange thing to say…
They could have used "open wights" which would have conveyed the company's desired intent just as well as "open source", but without the ambiguity. They deliberately chose to misuse a well established term instead.
I applaud and thank deepseek for opening their weights, but i absolutely condemn them and others (e.g Facebook) for their deliberate and continued misuse of the term. I and others like me will continue to raise this point as long as we are active in this field, so expect to see this criticism for decades.
Hopefully one of these companies losses a lawsuit due to these shenanigans. Perhaps then they wouldn't misuse these terms so brazenly.
This is the kind of inconsequential nitpicking diatribe I'm referring to. When has "open data" ever meant Open Source?
> They deliberately chose to misuse a well established term instead.
Their model weights as well as their repositories containing their technical papers and any source code are published under an OSS MIT license, which is the reason why initiatives like this looking to reproduce R1 are even possible.
But no, we have to waste space in every open model release complaining that they must be condemned for continuing to use the same label the rest of the industry uses to describe their open models which are released under an OSS License as Open Source - instead of using whatever preferred unused label you want them to use.
This is exactly why it is not “US vs China”, the battle is between heavily-capitalized Silicon Valley companies versus open source.
Every believer in this tech owes DeepSeek some gratitude, but even they stand on shoulders of giants in the form of everyone else who pushed the frontier forward and chose to publish, rather than exploit, what they learned.
DeepSeek is awesome. Any AI task yet implemented in our business can be run from my local PC with just the smaller models. And my PC is fairly crappy to begin with.
OpenAI looks quite silly with their "we have to close everything".
I do have some applications that process images, text and pdf files and I use smaller models for extracting embeddings. I think my system wouldn't be able to handle it with decent speed otherwise.
I do run LLM on a M1 16gb macbook air and performance is surprisingly good. Not for image synthesis though and a PC with a dedicated GPU is still significantly faster with LLM responses as well. Haven't tried to run deepseek on the macbook yet.
And then there’s all the dystopian propaganda baked into these models, which threatens to misinform users at scale based on a government driven agenda. Hard to be on that team, let alone firmly, knowing that it’s giving power to a dictatorial regime.
Those models are also trained on data that was ignoring licenses / copyrighted content.
I haven't tested deepseek for censorship yet, but they shared their release and even their input data. And in this case you could correct its shortcomings, so propaganda would be difficult.
And when they thought they were the only game in town, they tried to corner the market in GPUs and lock out any users who can't pony up £200/mo. Reminds me of when the likes of Oracle and IBM had companies by the balls buying bigger and bigger servers and then Google came along and showed everyone how to do horizontal scaling of cheap hardware.
The first one is definitely not true and the 2nd one is not necessarily true in the way you imagine i.e crawls of the internet will have gpt chat logs now.
Running local on CPU opens so much possibilities for smart and privacy focused home devices that serve you.
In my test it hallucinated confidently but my interest is in simple second brain like rag. "Hey thingy, what is my schedule today?"
Need it to be a bit faster though as the thinking part adds a lot of latency.
It does add latency of course, but I still think that I could provide all AI needs of my company (industrial production) with a simple older off the shelf PC. My GPU is decently recent, but the smallest model of the series and otherwise the machine is a rusty bucket.
I didn't test it thoroughly yet, but I have some invoices where I need to extract info and it did a perfect job until now. But I don't think there is any LLM yet that can do that without someone checking the output.
Silicon Valley has mo moat other than money. It kind of runs on openness and freedom of movement of people. Companies constantly poach people from each other. And there's a constant movement of people (and knowledge) in and out of the area. Money is what attracts these people and keeps them there for a while. But of course that status quo was upset a little bit with VCs turning into penny pinching misers lately and lockdowns proving (to them) that it was cheaper to host your tech teams remotely. Which means knowledge is now more distributed than it used to be.
So, it's not surprising that people outside of Silicon Valley are not waiting patiently for OpenAI to do whatever it is they are doing in between having moral existential crises, trying to oust their CEO, pontificating about AGIs, etc. They are taking things into their own hands. The brute force / VC funding driven approach that OpenAI has used yielded massive results in the last few years. But ever since Meta opensourced their models, OSS models and optimizations have been catching up.
On a hardware resource usage basis, these models started to outperform their bigger peers last year and now the game is up for the training process as well. Meaning they get better results for the same money. A major hurdle here was the model training process. Which the Chinese seem to have proven can be massively optimized as well. Cutting cost by a few orders of magnitude is a big deal. And at the same time doing the same thing at larger scale (aka. throwing more money at the problem) seems to have diminishing returns.
Until that changes, that means the playing field has somewhat leveled now. That's a good thing.
It’s trivial to implement bias in models (hence the no-no filters in chatgpt) so if they’re smart they’ll do what they do with tiktok and make the answers different for their rivals.
They shared their methodology, so if they are legit, someone else will reproduce what they did very quickly. I expect Meta, Amazon, Google, and Anthropic are on the case right now.
From this list, the only one I trust any more than I trust Deepseek is Anthropic.
The other three have shown they'll instantly bend the knee to whomever is in power, and that's exactly the same thing people are worried that Deepseek is doing.
Last year I would have said I trusted American companies more than Chinese. But last year feels like a long time ago.
Will you kindly explain why? Thank you.
100%
Ah yes the "open source" code that was not released by the DeepSeek team and the tens of thousands of professional grade GPUs that were contributed by the "community".
DeepSeek is based on Llama which was produced by ... Meta.
Meta never released a mixture of expert model (they failed to train a good one, according to reliable rumors). And MoE is just one of few ingredients that make DeepSeek v3/R1 interesting and good.
You aren't? You're weird.
- Libre engines such as the ones from https://osgameclones.com
- Godot
- Use anything else, there are tons of libre search engines
I do like the idea of making these reasoning techniques accessible to everyone. If they really manage to replicate the results of DeepSeek-R1, especially on a smaller budget, that’s a huge win for open-source AI.
I’m all for projects that push innovation and share the process with others, even if it’s messy.
But yeah—lots of hurdles. They might hit a wall because they don’t have DeepSeek’s original datasets.