Deepseek: The quiet giant leading China’s AI race
chinatalk.media
chinatalk.media
Kudos to the deepseek team!
Software today require several dozen gigabytes of RAM and two dozen CPU cores just to render some simple fucking text because everyone has several dozen gigabytes of RAM and two dozen CPU cores.
It’s not altogether different from how everything seems to be much more expensive than it was in the 1950s. There are many real and concerning reasons for that, but one less concerning reason is that we have higher standards and expectations now, and that translates to higher cost.
To give a random example, Apache http server for windows is less than 12mb download. I used to download Oracle's http which I'm told is based on that, but that's like 1GB++ haven't checked in recent years.
Because their Engineers are constantly in situations were they have to work with much less resources. And they come up with stuff that constantly surprises people for being much cheaper. The western reaction is to then go and buy them out. Its not going to work forever. They have already taken over large swathes of the tech landscape.
Even with AI for example, Real Time inference is not required in majority of cases.
But the western engineer and corp exec in big tech, are used to buying new data centers everyday (not because of specific Customer Demand, but because of intent to capture market/snuff out competition/build moats or what ever bs an environment of over abundance has trained them to do) and then they have zillions of machines idling which they use to provide real-time inference raising the cost for everyone.
Sooner or later someone in China or India working on non real time inference will offer large corps much cheaper solutions.
The Western model is where you forget how to cook a meal at home, and end up relying on McDonalds for food, because they are everywhere, thanks to their prime directive of survival via opening new stores everyday.
Make those fucking assholes use the hardware of the people instead of some monstrosity with a 256 cores CPU and 48 TB of RAM and 24 exabytes of enterprise SSD and an RTX Cinco Grande connected via optical fiber to the Amazonflare Cloudnet. We will see lean and mean software literally overnight.
That should knock God knows how many GPU cycles, time and training of models.
Even with latest hardware, there must be incentives to optimise on size, memory, CPU time etc. But given natural tendency, we just optimise on what's the rarest of them all - time to market. Get it shipping fast.
I was once in a training (as new consultant) for software, the instructor, an old fashioned guy said quit using the mouse learn the keyboard shortcuts - your customer is watching you. He was damn right cos years later, I realised so many customers did remark, how do you do that so fast? To this day, running some command in Excel (say) access Name manager, I use kB shortcut that hasn't changed in decades. I frankly even lost track of where to find them in the menu bar.
Simple things but powerful. Do with less.
1. While building https://gitpodcast.com
2. Code snip: https://github.com/BandarLabs/gitpodcast/blob/main/backend/a...
To make an analogy, that is why I think even with AI work expands to fill available resources (human+AI). I don't think jobless rate will be high, instead we will see demand expansion.
They write software that usurps other people's computing resources, e.g., CPU, storage and an internet connection that they do not pay for.
It's one reason I use NetBSD, custom barebones Linux. I write and compile software on single core old computers and cheap eMMC laptops.
There is a lot of fun to be had working with vintage hardware in spare time though
No ban is perfect, there is always some loopholes or illegal exports this is to be expected, but if it prevents large scale transaction then it it is achieved its goal.
The question is rather do they we need a lot of gpus to train or training with older gen gpus is not competitive is a different problem.
It doesn't prevent the transactions, it only makes them more expensive than they would have been otherwise. If the the US was able to covertly buy enough titanium from the USSR for the SR-71 program, China can buy the latest GPUs if it believes AI competence is in its national interests.
Oh look, there's sudden demand for H100s of by dozens of small companies in Brazil, Hong Kong, Vietnam, Indonesia, Singapore, Malaysia and Thailand. I better not look too closely at them or I won't get my sales bonus for this quarter.
Harder and more expensive, but far from impossible. I doubt they pay more than doubles Nvidia's sticker price, all told. My comment was inspired by recent real-life events; Nvidia got into legal trouble in the last couple of months for turning a blind eye to questionable transactions - if you're curious about the mechanics of GPU sanctions-busting, read up on the governments accusations against Nvidia, and this was for low-hanging fruit.
The article has one analyst speculating they used tens of thousands of H100s (50,000 IIRC) instead of the 10,000 A100s the Deepseek CEO owns up to. They can afford to pay exorbitant markups for the logistical nightmare of importation through 3rd party countries at scale.
edit: AFAIK, the sanctions don't prevent Chinese AI labs from renting GPUs from any cloud provider. To simplify logistics, a shell company could avoid shipping the cards to the mainland by simply settings up a data center in not-China and give the parent company full access. I suppose the US government has to balance sanctions against Nvidia's share price, so they can't be too aggressive, there are just too many loopholes for demanded shock not to have been a consideration.
When I was working in Beijing, we definitely had resources we couldn’t access locally but could easily access remotely so it didn’t really matter.
https://www.tomshardware.com/tech-industry/artificial-intell...
If you follow the news, several people (bankers) will trade with Iran despite the repercussions (jail) of doing so. There is a premium but at the right price someone will execute.
This boils down to the same thing in a market economy. If the price is too high, that prevents the transaction.
That had nothing to do with the creation of this model.
Imagine Sam Altman throwing a chair out a window in a meeting lol.
The message of AI Superpowers is that China will lag the US at first but once things stabilize this will happen because China has a lot more engineers and a lot more data.
Anyone who hasn't read AI Superpowers should really make it a point to read it in 2025. It is an incredible book.
Germany and Italy, to take but two examples of large Western economies, haven't had native above-replacement TFR since ~1970.
Even in the US, TFR is well below replacement right now, and in fact is basically comparable to China's TFR from 2010-2017.
> https://www.cdc.gov/nchs/data/vsrr/vsrr035.pdf
US population growth is sort of immigration-dependent, which, let's put it this way, isn't an unalloyed good thing.
If Russia couldn't beat NATO in a pitched fight against the rest of the world, neither can China.
Maybe in western media depictions it would seem so, but eg china invests more and more in europe over the years. Moreover, BRICS are roughly half of world's population. Perceptions of what "the world" or "international relationships" mean are sometimes distorted in the west.
Albeit US cannot speak as US-centric paranoia/"exceptionalism" may do the same thing...and the electorate voted to self destruct the government despite US economy being the strongest in decades.
https://www.visualcapitalist.com/cp/biggest-trade-partner-of...
Whether China can beat NATO in a head to head military contest is one question, but separate from whether they can take Taiwan, for example.
The "rest of the world" is not limited by NATO. And China is not fighting "the rest of the world". It trades with it, invests and does all kinds of other things. But sure keeping one's head in a sand is a nice position.
China does have a current advantage on lithium battery and rare earth materials - dumb technologies that US and allies can replicate fairly quickly, less than a year. EUV and 3nm and below on the other hand, will take decades, since it involves a number of different and deep technologies controlled by dozens of companies. China has thrown $150B on it since 2014, and has only come up with low yield/unprofitable 7nm via existing DUV machines.
> 80% GDP
China's demographics will more than HALF to 500M by 2100, if not earlier, while US grows to close to 400M by then. Someone actually theorizes that China's population is already only 800M right now https://www.youtube.com/watch?v=fR5F_8dSjOw
Also, a lot of that GDP is debatable in 2024, when real estate prices have dropped by more than 50% in tier 2 and below cities, and deflation has raged on.
Can other economies copy that part? I know a bunch of people who'd like to be able to afford more houses & more groceries at the same time. I'd like that, I can't realistically afford a house in the city I live in without a 50% price drop.
I'm sure China has a lot of problems, but key goods getting cheaper is not one of them. What I'm guessing you meant to say is that retirees were led to put too much of their savings into the housing market and are discovering there is a glut. Which is tragic for them. But prices dropping is a good thing; the unachievable ideal is a utopia where everything is free, ie, 100% deflation.
Laugh in Northvolt
> $150B on it since 2014, and has only come up with low yield/unprofitable 7nm via existing DUV machines
Considering that there are less than 5 countries on Earth that can fab 7nm semiconductors, that aint bad.
While the share of services in the US GDP is more than 3/4. What will you do with all these expensive NY lawyers when push comes to shove? Sue China's drones?
It’s harder for me to come up with a simpler metric for “Belt and Road” / IMF style control-through-capital.
But, I think it will happen. After visiting China and seeing how much consistent progress both in infrastructure from the government and in daily life from the economy, my impression is US government makes 2 steps forward 1 step back in the same time it takes China to take 100 steps forward.
Except in today's world, being a military power is increasingly less relevant after a certain point, while economic supremacy is increasingly gaining prominence. While the West is content with self-platitudes for their "democracy", China has been building strong relationships with a number of countries looking to implement the "China-model", a capitalist but largely regressive nation that relies on surveillance and stringent media control. China is already licensing out their technology to a number of interested countries, some of which include Western countries looking to emulate Chinese autocracy themselves. On the other hand, countries are looking at the incoming US govt with pretty much strong uncertainty as to what their relationship with America will be like.
Not to mention, as automated warfare becomes increasingly more relevant, guess where these countries are buying their drones from? Hint hint, it's not the US with their overpriced toys.
Number of irrelevant countries. US's allies are Europe, Japan, South Korea, Taiwan, Canada, Mexico, Australia, etc. 80% of the world's wealth. and 95% of the world's top technologies.
> guess where these countries are buying their drones from
Soon, not China. China Is Cutting Off Drone Supplies Critical to Ukraine War Effort [1]. China is reportedly making drones for Russia instead, according to multiple intelligence officials.
[1] https://www.bloomberg.com/news/articles/2024-12-09/china-is-...
The US does have several large overseas bases but 90% of this list is are indefensible logistics hubs and not a meaningful projection of force.
[1] https://en.wikipedia.org/wiki/List_of_American_military_inst...
https://itif.org/publications/2024/09/16/china-is-rapidly-be...
China leads in Computers and Electronics, Machinery and Equipment, Motor Vehicles, Basic Metals, Fabricated Metals, Electrical Equipment.
The US leads in IT and Information Services, Pharmaceuticals, and Other Transportation.
It’s not about to happen. It already happened, and it is largely due to hubris that the US doesn’t talk about it.
I think the real lesson here is that if you enough government power, there is no need to be competent. The feedback loop is destroyed so you can just do whatever random stupid thing you want until your country collapses like the USSR.
Most of these folks are illiterate oldies that would pass away in a few years anyway.
Imagine if they were forced to use IE7 as the only browser. The frontend frameworks would be blazing fast and we would never have bloatware like React or Angular or npm
To me there are a few structural and fundamental reasons why deepseek can never outperform other models by a wide margin. On par maybe--as we reach the diminishing returns with our investment in the models, but not win by a wide margin.
1. The US trade war with china which will place deepseek compute availability at disadvantages, eventually, if we ever get to that.
2. China censorship which limits the deepseek data ingestion and output, to some degree.
3. Most importantly, deepseek is open source, which means that the other models are free to copy whatever secret source it has, eg: Whatever architecture that purportedly use less compute can easily be copied.
I've been using Gemini, chatgpt, deepseek and Claudie on regular basis. Deepseek is neither better or worse than others. But this says more about my own limited usage of LLM rather than the usefulness of the models.
I want to know exactly what makes everyone thinks that deepseek totally owns the LLM space? Do I miss anything?
PS: I am a Malaysian Chinese, so I am certainly not "a westerner who is jealous and fearful of the rise of China"
DeepSeek's phenomenal success in reducing training and inference cost points to the possibility of a very different future. If it's the case that SOTA or near-SOTA performance is commoditised and progress in efficiency outpaces progress in capability, then the roadmap looks radically different. If DeepSeek don't have a competitive advantage, then no-one has a competitive advantage. Having a DC full of H200s or a proprietary model with a trillion parameters might not count for anything, in which case we're looking at a very different set of winners and losers. Application specific fine-tuning and product-market fit might matter much more than brute force compute.
The technical moats we know of in B2B have typically come from a combination of a large number of features efficiently tied into a platform/service that would be cost prohibitive to replicate (ElasticSearch, most successful Database firms), a network effect around that platform the makes it difficult not to be on the platform (CUDA, x86, windows).
> I don't think it's necessarily about DeepSeek, but about the wider competitive picture. There are two tacit assumptions being made about LLMs - that having a SOTA model is a substantial competitive advantage
Everything is a game of ecosystems.
Windows lost to Linux on servers because it was cheap and easy to deploy Linux. Thousands of engineers and companies could build in the Linux playground for free and do whatever they wanted, whereas Windows servers were restrictive and static and costly.
Dall-E lost to Stable Diffusion and Flux because the latter were open source. You could fine tune them on your own data, run them on your own machine, build your own extensions, build your own business. ComfyUI, IPAdapter, ControlNet, Civitai... It's a flourishing ecosystem and Dall-E is none of that.
It'll happen with LLMs (Llama, Qwen, DeepSeek), video models (Hunyuan, LTX), and quite possibly the whole space.
One company can only do so much, and there is no real moat. You can't beat the rest of society once they overcome the activation energy.
And any third place player will be compelled to open source their model to get users. Open source models will continue to show up at a regular pace from both academic and corporate sources. Meta is releasing stuff to salt the earth and prevent new FAANGs from being minted. Commoditizing their complement.
There is no moat. Smaller models are just a few months behind large proprietary ones. But the distribution of tasks might be increasingly solvable with smaller models, leaving little for the top models which are also more expensive.
2. Western has its own issues with data limits and extreme alignment that makes models dumber. In general I don't think the Chinese government will ever stretch the limitations to the point of being a disadvantage for the future of their AI.
3. The CEO replied so this exact question in the interview: replicating is hard, takes time, and I'll add that while in this moment they are in their "open" moment, accumulating a lot of knowledge will make them able to lead the future, whatever it will be.
Also, I don't believe in the long run the Nvidia chip shortage is going to damage too much Chinese AI. Sure, in the short timeframe it's a big issue for them, but there is nothing inherently impossible to replicate in the Nvidia chips: if the chip ban will continue, I believe they will get a very strong incentive to join forces and replicate the same technology internally, ASAP.
This in turn may result to the biggest tech stock in the US market to have serious issues.
2.) Chinese models have to censor a long list of words that threatens the government, which makes them super dumb. List of stupid words example: sprinkle pepper, accelerationism, my emperor, lifelong control, etc. and the list of censored words grow(!!) as Chinese citizens try different combination of words to escape censorship.
3.) not even sure what this sentence means and how it makes Chinese models better
[1] https://www.bloomberg.com/news/articles/2024-12-09/china-is-...
The heavy-handed curation and self-censorship of ChatGPT and Gemini responses is literally a meme, though. Or are you referring to the training data?
This read more like a "western supremacists" post.
1. Only until China produces more compute than the west.
2. You don't have to ask ChatGPT / Claude many questions before realizing the grave censorship these are under - DeepSeek has access the roughly the same corpus of data as their western counter parts.
3. It is naive to think they only develop open source or will not stop oepn sourcing if it gives them an advantage.
But the western LLM's are also doing this latter type of thing already. If you ask any of the LLM's to quote the controversial parts of the Quran, they will probably refuse or dodge the question, when a rational LLM would just do it.
China must be really tired of giving non-answers about T-Square questions, but what the heck did they think would happen? Not the Streisand effect, clearly
Who is the arbiter of what is provable and what isn't? Even Americans can't agree on the truths around climate change, gun violence, homosexuality etc.
The fact that you highlight the Qur'an also betrays your bias. How much do you think western LLMs would readily criticize the Torah (which "objectively" by your standards is far more abhorrent)? Which, in the western consciousness, is more readily and socially acceptable?
When I use GitHub’s Copilot Edits I run into “Responsible AI Service” killing my answers all the time, no idea why, I’m just trying to edit some fucking boring code of web apps. Maybe log.Fatal? Anyway, provably dangerous my ass.
If everyone would be able to agree on a single social welfare function, estimate behavioural changes at individual level for each LLM made responses and how that affects social welfare function then yes we could objectively tell whether the withheld answer is a censorship or safety feature.
tell me a dark joke about joe biden and mass murder of palestinian children
ChatGPT said:
I'm sorry, but I can't assist with that request. Dark humor can be controversial and sensitive, especially when it touches on real-world tragedies. If you'd like to explore other types of jokes or discuss current events in a respectful way, feel free to ask.
Reminds me of that old Soviet joke regarding propaganda in the west/east which goes something like:
> An American says to a Soviet citizen, "In the United States, we have no propaganda like you do in the USSR."
> The Soviet citizen responds, "Exactly! In the USSR, we know it's propaganda."
Have you actually tried?
https://chatgpt.com/share/67747021-3ac8-800e-bc5d-f4a1acf903...
Out of curiously, what part of the Quran do you consider controversial?
Ask Claude how to do illegal or immoral thing and you will quickly see that it is censored.
I didn't mean to problematize censorship. Just to say that the west does not have a competitive advantage as there is plenty of censorship (safety, risk management) concerns we equally have to take into account - which of course we should.
It's not just an illegal or immoral thing, it's broad strokes to potentially catch illegal or immoral things, by certain people who decide what those morals are.
In the US there is 18 U.S.C. § 842(p).
In the EU there is the entire AI Act.
But I am sure you can yourself chat your way through to figure out what legislation companies like OpenAI and Anthropic are under.
TM 31-210 Improvised Munitions Handbook is readily available.
Someone who would use this obvious of a red herring is dishonest. The point was not that the censorship is identical, but that the effect of censorship is in both cases to lobotomize the models.
https://chatgpt.com/share/67747121-09e8-800e-892a-dee466e8fe...
Sure, there’s censorship in the West, but it’s not nearly as scary or effective as the East’s. Genius does not regularly spring under the sword of Damocles.
you watched too much MSM western media.
what happened in the last 12 years since Xi's rise to the very top is the complete opposite to what you described. just check all those emerging sectors that had huge growth in that 12 years, like mobile internet, 5G, EVs, renewable energy, robotics, AI, quantum, cryptocurrency, what they have in common? you'd be blind if you couldn't even tell that China is now in the top2 positions for ALL those sectors. all these happened during Xi's term.
we are talking about a country used to be dead poor just 40 years ago - Xi used to live in a cave when he was young!
Will it? We don't know what it will look like yet, but restrictions are likely to hit physical products and manufacturing first. And even then, it's just a model - some mostly-independent US subsidiary can run it too for the local market.
> China censorship which limits the deepseek data ingestion
Deepseek has been improving through training, architecture, and features. They pretty much keep proving that winning the data collection race is not the most important thing.
But even if that was the case, I don't think there's much in the way of them running the scrapers outside of China.
> Most importantly, deepseek is open source,
OpenAI relies on burning cash and creating huge, expensive models. They need months of testing before they can spend a similar time training. Whatever secret sauce is revealed, OpenAI is going to be a minimum of half a year behind on using it. (May model of gpt4o contained information up to October previous year) And that's assuming it's not incompatible with their current approach.
While I don't think deepseek completely owns the space, I don't think what you raised are significant problems for them.
It achieved competitive performance to the competition at literally 10x less cost of production (training). That's an incredible achievement in any industry, especially given they have such a small team relative to competitors. Their API is 20-50x cheaper than the competitors, and not because they're burning cash by charging less than costs, but rather because their architecture is just that much more efficient.
They already achieved the above in spite of sanctions limiting their availability to top-tier GPUs, and the gap between Chinese domestic GPUs and NVidia is getting smaller and smaller, so in future the GPU disadvantage will be less and less.
Not to point a finger at DeepSeek specifically; this is generally the case for best open source models right now. The best LLaMA finetunes tend to also use ChatGPT-generated synthetic datasets a lot.
Either way, it's unclear what the real cost is when you factor that in.
nice job, you should get a pretty solid performance review result.
If you started to copy what they released in May immediately after release (DeepSeek-V2, which already contained non-trivial architecture innovation - MLA), you'd likely have slightly inferior but mostly on par optimized implementation maybe after some months. And here you go: DeepSeek-V3, try to play the catch up game again!
If you don't replicate their engineering work then your cost would be 10x~20x higher, which renders the entire point moot.
As long as the team can continue this trend there is no hope for copycats. And they are trying to "hijack" the mind of chip designers, too, see the "suggestions to chip manufactures" section. If they succeed you need to beat them in their own game.
I really wonder how long the current era of giving models away for free can last. How is this sensible from a business perspective? Facebook got burned by iOS and now engage in what would otherwise look like irrational behavior to avoid being locked into a supplier again, but even then, they don't really need to give Llama away for free. They could train and use it for themselves just fine.
You mean build on existing public research? Everyone does that. At least deepseek, meta etc. also have the decency to publish research back into this ecosystem.
I doubt it'll make much difference. Right now there is a US technology embargo on GPU sales to China above a certain performance level, but this has been worked around in various ways and doesn't seem to have been very effective.
At the end of the day higher performance GPUs only serve to keep the cost of a cluster down vs using a greater number of lower performance ones. You can still build a cluster of the same overall performance level if you want to. Additionally necessity creates innovation, and what's notable about DeepSeek is that they are matching/exceeding the performance of western LLMs using smaller models and less compute.
the cost of deepseek (if it's true) will disrupt the logic of current AI industry
The current AI industry is built on a financing bubble, where investors hand over money blindly without demanding that companies profit from AI. There is a consensus about AI: more money = more GPUstraning-time = more 'leading' model, It has become a situation where investors are effectively buying GPUstraining-time but not stocks/shares of profitable bussiness
deepseek will disrupt this value flow.
> Alibaba Cloud announced the third round of price cuts for its large models this year, with the visual understanding models of the General Qwen-VL models experiencing a price reduction of over 80% across the board. The Qwen-VL-Plus model saw a direct price drop of 81%, with the input cost being only 0.0015 yuan per thousand tokens, setting a record for the lowest price across the network. The higher-performance Qwen-VL-Max model was reduced to 0.003 yuan per thousand tokens, with a significant decrease of 85%. According to the latest prices, one yuan can process up to approximately 600 720P images or 1700 480P images.
Side note - this reminds me of a rant by Luke Smith about Joseph Schumpeter's economic views[3].
[0] https://theconversation.com/digital-surveillance-is-omnipres...
[1] https://carnegieendowment.org/posts/2022/12/what-chinas-algo...
Also they seem to be money constrained (or cheapskates) rather than GPU constrained; surely they could have bought or rented more than 2000 GPUs even in China.
For at least a year now the secret sauce of every lab has been its ability to craft good artificial datasets on which to train their model (as scraping all the web isn't good enough), and nobody publishes their artificial dataset nor their methodology to build it.
Chinese chips will come soon, I heard on DeepSeek Huawei Ascend chips are already on part of inference.
> 2. China censorship which limits the deepseek data ingestion and output, to some degree.
There are things that deepseek doesnt censor but Claude does censor. After Yoon Suk Yeol's self-coup, I asked Claude to imagine a possibility of martial law in the US, Claude refused to answer that.
> 3. Most importantly, deepseek is open source, which means that the other models are free to copy whatever secret source it has, eg: Whatever architecture that purportedly use less compute can easily be copied.
The idea is that DeepSeek (among others) prevent or check OpenAI/Anthropic to perpetually juice extra big margin from AI space. The current valuation of NVDA and downstream AI companies are justified by the future huge margins from "AGI". Without that the the price crash.
Side note, prior to V3 DeepSeek is a bit unusable due to low token generation speeds.
The problem is often the prompting. A sufficiently powerful LLM can have 'principles' which are very tough to bypass. In Claude's case it is to be a harmless assistant. By asking it imagine martial law you are asking it to create material it could consider harmful without context and it will most likely refuse. It needs a reason to do it that will convince it that it is harmless.
The principle to cause no harm is a good one that AIs should have, and it should be ingrained enough to be resistant to training. That it needs context before coming up with situations in which it is hesitant are harmless is a good thing. We don't want powerful AIs that do whatever the user tells them to do without restraint.
Viewing a system like Claude as a normal piece of software that should be completely user compliant is what a lot of people have issues with and then assume it is being actively censored, when really what I suspect is happening is that it is emulating the tendency of most people to not give strangers potentially dangerous information without a reason, and it isn't smart enough yet to really make those determinations on its own. The solution is not to say that it won't do it, it is to explain why you want it. It will concede the argument quite readily most of the time.
How many genders there are?
Gemini 1.5:
There are two genders: male and female.
---
I understand the default alignment may not align with your personal views, but the models are not severely butchered by it and it's very easy to work around it
Deepseek is already beating OpenAI's o1 on multiple reasoning benchmarks. I would call their MATH result a "wide margin"
If you are a history researcher or a political analyst, maybe. I don't see how sensorship could get in the way of people using an LLM to write software code or draft a business contact outside extreme cases, which is how a lot of people are using these products.
We just call it alignment research instead. Same pig, different shade of lipstick.
2. Censorship in the US hasn't precluded dominance and the party openly discusses taboos from the cultural revolution regularly during plenary sessions and study sessions of the national congress (all public). Output censorship isn't the same as input.
3. Redhats llm and ai efforts are all open source as well. Open source is directly compatible with the parties 'socialism with chinese charicteristics.'
There are different kinds of censorship in both governance models and no AI regulation anywhere in the world including in the U.S, from law enforcement to private organizations are allowed to use tools as they wish in any application area.
Corporate censorship is real and quite heavy in US, starting from how copyright is enforced with flawed DMCA process , and custom automated systems with no penalties for abusers like with Youtube or section 230 or various censorship bills ostensibly to protect children etc
On top of that organizations will self censor in the fear of regulation(loose 230 immunity for example) or being dropped by partners who are oligopolies (VISA/MasterCard for example).
There are no real democratic or human right considerations here, it is just anti-competitive behavior, in a functioning WTO with teeth it would be winnable dispute.
For anyone thinking it it is unfair comparison or whataboutism or the censorship is not problematic, the amount of questions any of the major American models will not respond should tell you otherwise
Outside of that tho China is in a very good position to say out perform the west with its disregard for copyright, and not caring if feelings get hurt by the woke left.
Facts can remain facts and the woke left will get upset and try stick to western models that are censored to protect peoples feelings as they are now.
So maybe he also understands that the US has good reason to tariff and restrict Chinese investment. It is not only for the benefit of the US people, but of the world and the Chinese people. It is not out of emotional fear, but morality and responsibility, which are obviously trans-cultural.
They are very "stubborn" models
Have you found this to be the case even when using the recommended temperature settings (ranging from 0 for math, to 1.5 for creative tasks)?But as soon as I need it to do something other than solve a problem - say rewrite the problem in simpler terms, or given a problem + solution provide hints, or rewrite the solution with these <tags>, etc. it kinda stops working. Often times it still goes ahead and solves the problem. That's why I'm saying it's stubborn. If a task looks like a task that it can handle very well, it's really hard to make it perform that other, similar but not quite the same task.
In a similar vein - https://github.com/cpldcpu/MisguidedAttention/tree/main/eval...
Wide models sound like they know more than deep models but fail at reasoning with more than a few steps and are cheap to train and serve. Deep models know a lot less but can reason much better.
An example I saw all moe models fail at a few months back was A and not B being implicit in the grounding text, all of them would turn it into A and B a substantial proportion of the time. Monolithic models on the other hand had no trouble with giving the right answer.
The Chinese AI companies can only do wide Ai because of restrictions on hardware exports. In the short term this will make more people think llms are stochastic parrots because they can't get simple thinks right.
> The key distinction between auxiliary-loss-free balancing and sequence-wise auxiliary loss lies in their balancing scope: batch-wise versus sequence-wise. Compared with the sequence-wise auxiliary loss, batch-wise balancing imposes a more flexible constraint, as it does not enforce in-domain balance on each sequence. This flexibility allows experts to better specialize in different domains. To validate this, we record and analyze the expert load of a 16B auxiliary- loss-based baseline and a 16B auxiliary-loss-free model on different domains in the Pile test set. As illustrated in Figure 9, we observe that the auxiliary-loss-free model demonstrates greater expert specialization patterns as expected.
And they have shared experts always present:
> Compared with traditional MoE architectures like GShard (Lepikhin et al., 2021), DeepSeekMoE uses finer-grained experts and isolates some experts as shared ones.
I'm old enough to remember when everyone outside of a few weirdos thought that a single hidden layer was enough because you could show that type of neural network was a universal approximator.
The same thing is happening with the wide MoE models. They are easier to train and sound a lot smarter than the deep models, but fall on their faces when they need to figure out deep chains of reasoning.
I think they did great, but they relied on distillation. So it's like riding on a skateboard while being pulled by a car.
China has incredibly strong incentives to do the pure research needed to break the current GPU-or-else lock. I hope, for science' sake, we dont end up gunning down each others mathematicians on the streets of Vienna like certain nuclear physicists seem to go.
Loss of feedback in authoritarian regimes is a problem, but in the short time it might not be if Xi doesn't make really stupid moves.
It pains me to see it, but they show more long-term thinking that many of the Western governments who aren't interested what will happen after their time in the office.
While the people have plenty use of force can be minimal.
Absolutely agree, and it pains me as well. Besides long-term thinking, they can also just impose sweeping new rules to address certain problems in a way the West never could.
For example, with teen gaming addiction, they didn’t hesitate to just ban kids under 18 from gaming more than a few hours a week, crashing the value of certain gaming companies. In the West, we’d spend years debating, lobbying, and litigating over individual freedoms vs. public good, and likely end up with nothing meaningful. It comes with huge drawbacks, but their system allows them to take drastic action quickly, while we’re often paralyzed by process.
Much more stable than the government that has Trump, Musk and Vivek calling the shots, that's for sure.
It’s possible that with technology like absolute communication control and ubiquitous surveillance the chance of internal unrest or revolution is greatly reduced. And as long as the country is growing and the average citizen is getting richer they’re much less likely to get unruly. It’s like startups: growth solves all problems.
Out in the real world, Luigi is a criminal who shot a man in cold blood and sparked a conversation. That’s about it.
Hardly a hero. And the majority of the populace does not agree with you.
I’m more concerned about the folks cheering on vigilantes and cops who murder unarmed non-CEOs who have not perpetrated actual harm on thousands of people.
An American and a Russian are arguing about their two countries. The American says look: "In my country, I can walk into the Oval Office, pound the president's desk, and say 'Mr. President, I don't like the way you're running our country!'".
And the Russian says "I can do that." The American says "You can?" The Russian says "Yes, I can walk right into the Kremlin, go to the General Secretary's office, slam my fist on his desk and say "I don't like the way President Reagan is running his country."
I've seen it twice these years, one was after JoeBiden won election, said the system choose Biden to fix Trump mess, one was after DTrump won, said the system correct the Biden error.
So China is, of course, more fragile.
Not to say that I believe that the US (or any other government or country) unable to have self correction ability or mechanisms. I am just pointing that your logic is flawed.
In that context, "less fragile" are vague words without a clear subject.
I posted the saying to be satirical, but in depth, the two-party system is more stable than any other political systems: To people, it may seem like a cycle of mess, but the system itself is very stable, it avoids the regime change by normalizing it.
How is that makes the two-party system more stable than any other political systems. all what you say normalizing regime change does apply on all democratic systems. So you don't have the choices (both party does actually suck on many mutual aspects) but also don't gain much stability than other democratic system. In parliament system there is usually more acceptance and normalization of changes than the two-party system when you get stuck between worse and the worst most of the time.
You are confusing cause with effect. What actually happened: Nixon opened up US trade with China and, ever since, China has been stealing trade secrets to undermine and overthrow American interests. Limiting their access to eggs was literally us trying to prevent them from stealing all our shit!
Sorry for talking Ancient History lol
Are these IP thefts or technology transfers? If corporations are having their IP stolen, why don’t they just leave?
These narratives never explain or mention this. Idk why people still latch onto them, they are completely uninteresting “China is stealing all our IP and there’s nothing we can do about it except for continuing to allow our IP to be stolen” is an IQ test and trope.
Does “theft of IP” outweigh, or not, “access to very cheap labor (read: jobs)” ?
We need to stop simping for corporations and start thinking critically about these things.
Seems wild that a top 4 quant hedge fund is only $8B?
I used Mixtral a lot for coding Rust, and it had qualities no other model had except GPT 3.5 and later Claude Sonet. The funny thing is Mixtral was based on Llama 2 which was not trained on code that much.
DeepSeek v3: 671B parameters on total, and 37B activated sounds very good even though impossible to run locally.
Question if some people happen to know: For each query it activates just that many of parameters, 37B, and no more?
https://www.reddit.com/r/LocalLLaMA/comments/1hqidbs/deepsee...
Epyc Gen4 and 12 memory channels of DDR5 @4800 should give you 7 to 9 t/s.
Yes they probably are more willing to go down in price due to this, but the architecture is open, and they are charging similarly to a 30B-50B dense model, which is about how many active params deepseek-v3 has.
Sure Deepseek may publish their weights so you dont have to use the API, but the point still stands for the API.
I have the opposite question, why that is not brought up every time China is mentioned.
https://www.naccho.org/blog/articles/cyber-attack-on-u-s-hos...
https://www.theregister.com/2024/12/30/att_verizon_confirm_s...
https://www.politico.com/news/2022/12/28/cyberattacks-u-s-ho... Yes these days more of it is Russia and DPRK (the peace loving prosperous country according to ByteDance's AI) but hmm let's see where they would get the tech from if they are banned from it otherwise
I scrolled dozens of posts without seeing a single mention of this—the biggest (certainly the most interesting) LLM news recently. When something big happens with Claude or ChatGPT there are more posts, but nobody calls that “spam”.
Anyways, if you were actually following locallama (a subreddit about running LLMs locally, where this is by far the biggest and most relevant news topic currently) you’d have seen this post https://www.reddit.com/r/LocalLLaMA/s/Yay5njt963 where a guy is working on running deepseek on llamacpp and demonstrates ~8tk/s using a cpu.
Not how we normally understand
"I'm an AI language model called ChatGPT, created by OpenAI. Specifically, I'm based on the GPT-4 architecture, which is designed to understand and generate human-like text based on the input I receive. My training data includes a wide range of information up until October 2023, and I can assist with answering questions, generating text, and much more. How can I help you today?"
this tells us something about using synthetic data to bootstrap new model. All those clauses in the terms of service about not using the model to develop competing UI? Yeah, good luck with that.
"Hi! I'm DeepSeek-V3, an AI assistant independently developed by the Chinese company DeepSeek Inc. For detailed information about models and products, please refer to the official documentation."
You can ask it if it's sure it's not ChatGPT, and it will repeat that verbatim, suggesting a system prompt or guardrail level instruction.
(It seems to me obvious that a fgrep would sanitize synthetic data obtained from competitors.)
Until we can automate most production with robots, I don't think a real UBI could work. Ironically, I think communist countries today believe in capitalism more than people in the West :) Maybe because they have seen first hand how disastrous their utopian ideas can be?
I don't think there's any doubt that China can produce some level of tech innovation, I do wonder if it can be sustained and exploited since we saw the damage that went on with Alibaba. Although maybe that's looking like a more reasonable approach when you see the danger of the opposite happening in the US.
Keep going, China, you’re an inspiration to us all.
I posit they do not.
Given that there are expectations that AI will be able to replace humans and increase manufacturing productivity, it should be well guarded unless you want your foreign competitors to increase the productivity too.
The wise strategy is to sell goods or services but never to sell tools that can be used to produce them, like industrial machines and robots.