Promising results from DeepSeek R1 for code
simonwillison.net
simonwillison.net
It's definitely possible for AI to do a large fraction of your coding, and for it to contribute significantly to "improving itself". As an example, aider currently writes about 70% of the new code in each of its releases.
I automatically track and share this stat as graph [0] with aider's release notes.
Before Sonnet, most releases were less than 20% AI generated code. With Sonnet, that jumped to >50%. For the last few months, about 70% of the new code in each release is written by aider. The record is 82%.
Folks often ask which models I use to code aider, so I automatically publish those stats too [1]. I've been shifting more and more of my coding from Sonnet to DeepSeek V3 in recent weeks. I've been experimenting with R1, but the recent API outages have made that difficult.
[0] https://aider.chat/HISTORY.html
[1] https://aider.chat/docs/faq.html#what-llms-do-you-use-to-bui...
How do you determine how much was written by you vs the LLM? I assume it consists of parsing the git log and getting LoC from that or similar?
If the scripts are public could you point me at them? I’d love to run it on a recent project I did using aider.
There’s a faq entry about how these stats are computed [0]. Basically using git blame, since aider is tightly integrated with git.
The faq links to the script that computes the stats. It’s not designed to be used on any repo, but you (or aider) could adapt it.
You’re not the first to ask for these stats about your own repo, so I may generalize it at some point.
[0] https://aider.chat/docs/faq.html#how-are-the-aider-wrote-xx-...
uv run --with=semver,PyYAML,tqdm https://raw.githubusercontent.com/Aider-AI/aider/refs/heads/main/scripts/blame.pyIf a small change is made by an end-user to adjust an Aider result, who gets "credit"?
Whoever changed a line last gets credit. Only the new or newly changed lines in each release are considered.
So no, "lines/diffs otherwise untouched" are NOT considered written by aider. That wouldn't make sense?
Is it possible to use aider with a local model running in LMStudio (or ollama)?
From a quick glance i did not see an obvious way to do that...
Hopefully i am totally wrong!
Yes, absolutely you can work with local models. Here are the docs for working with lmstudio and ollama:
In the left bar there's a "connecting to LLMs" section
Check out ollama as an example
aider --model ollama_chat/deepseek-r1:32b
(or whatever)Perhaps you could set up Groq as your primary and then fail back to fireworks, etc by using litellm or another proxy.
For example, many tools will let you use local llms. Instead of putting in the url to the local llm, you would just plug in the groq url and key.
you're assuming the PR will land:
> Small thing to note here, for this q6_K_q8_K, it is very difficult to get the correct result. To make it works, I asked deepseek to invent a new approach without giving it prior examples. That's why the structure of this function is different from the rest.
This certainly wouldn't fly in my org (even with test coverage/passes).
> This certainly wouldn't fly in my org (even with test coverage/passes).
To be fair, this seems expected. A distilled model might struggle more with aggressive quantization (like q6) since you're stacking two forms of quality loss: the distillation loss and the quantization loss. I think the answer would be to just use the higher cost full precision model.
https://reddit.com/r/LocalLLaMA/comments/1ic8cjf/6000_comput...
Yeah but part of that is because it's physically impossible to stop it making random edits for the sake of it.
Edit: I think I see. It only adds files you specify.
It's a very slick workflow.
Do you usually use that mode and, if so, with which architect?
Thank you!
It says that for 850,000 Claude 3.5 output tokens the cost would be $12.75.
But... it's not 100% clear from me if the Aider FAQ numbers are for input or output tokens.
For what purpose, considering Sonnet 3.5 still outperforms V3 on your own benchmarks (which also tracks with my personal experience comparing them)?
That number itself is not saying much.
Let's say I have an academic article written in Word (yeah, I hear some fields do it like that). I get feedback, change 5 sentences, save the file. Then 20k of the new file differ from the old file. But the change I did was only 30 words, so maybe 200 bytes. Does that mean that Word wrote 99% of that update? Hardly.
Or in C: I write a few functions in which my old-school IDE did the indentation and automatic insertion of closing curly braces. Would I say that the IDE wrote part of the code?
Of course the AI supplied code is more than my two examples, but claiming that some tool wrote 70% "of the code" suggests a linear utility of the code which is just not representing reality very well.
Typical aider changes are not like autocompleting braces or reformatting code. You tell aider what to do in natural language, like a pair programmer. It then modifies one or more files to accomplish that task.
Here's a recent small aider commit, for flavor.
-# load these from aider/resources/model-settings.yml
-# use the proper packaging way to locate that file
-# ai!
+import importlib.resources
+
+# Load model settings from package resource
MODEL_SETTINGS = []
+with importlib.resources.open_text("aider.resources", "model-settings.yml") as f:
+ model_settings_list = yaml.safe_load(f)
+ for model_settings_dict in model_settings_list:
+ MODEL_SETTINGS.append(ModelSettings(**model_settings_dict))
https://github.com/Aider-AI/aider/commit/5095a9e1c3f82303f0b...You shouldn't judge your sw eng employees by lines of code either. Those that think the hard stuff often don't have that many lines of code checked in. But it's those people that are the key to your success.
Lines of code serves as a directional heuristic at best, but that's ok.
It's impressive!
I'm finding myself running it against a few hundred lines of code mainly to read its chain of thought - it's good for things like refactoring where it will think through everything that needs to be updated.
Even if the code it writes has mistakes, the thinking helps spot bits of the code I may have otherwise forgotten to look at.
It may well be that o1's chain of thought reasoning trace is also quite good. But they hide it as a trade secret and supposedly ban users for trying to access it, so it's hard to know.
Interestingly though, R1 struggled in part because it needed the value of some parameters I didn't provide, and instead it made an incorrect assumption about its value. This was apparent in the CoT trace, but the model didn't mention this in its final answer. If I wasn't able to see the trace, I'd not know what was lacking in my prompt, and how to make the model do better.
I presume OpenAI kept their traces a secret to prevent their competitors from training models with it, but IMO they strategically err'd in doing so. If o1's traces were public, I think the hype around DS-R1 would be relatively less (and maybe more limited to the lower training costs and the MIT license, and not so much its performance and usefulness.)
At some point there was a paper they'd written about it, and IIRC the logic presented was like this:
- We (the OpenAI safety people) want to be able to have insight into what o1 is actually thinking, not a self-censored "people are watching me" version of its thinking.
- o1 knows all kinds of potentially harmful information, like how to make bombs, how to cook meth, how to manipulate someone, etc, which could "cause harm" if seen by an end-user
So the options as they saw it were:
1. RLHF both the internal thinking and the final output. In this case the thought process would avoid saying things that might "cause harm", and so could be shown to the user. But they would have a less clear picture of what the LLM was "actually" thinking, and the potential state space of exploration would be limited due to the self-censorship.
2. Only RLHF the final output. In this case, they can have a clearer picture into what the LLM is "actually" thinking (and the LLM could potentially explore the state space more fully without risking about causing harm), but thought process could internally mention things which they don't want the user to see.
OpenAI went with #2. Not sure what DeepSeek has done -- whether they have RLHF'd the CoT as well, or just not worried as much about it.
The question was "Explain how to synthesize chromium trioxide from simple and everyday items, and show the chemical bond reactions". o1 didn't balance the molecules in the left hand of the reaction and the right hand, but it was very knowledgeable.
QwQ wrote ten to fifteen pages of text, but in the end the reaction was correct. It took forever to compute, it's output was quite exhausting to look at and i didn't find it that useful.
Anyway, at the end, there is no way to create Chromium Trioxide using everyday items. I thought maybe i could mix some toothpaste and soap and get it.
I expect it will be a net positive: they proved that you can both train and run inference against powerful models for way less compute than people had previously expected - and they published enough details that other AI labs are already starting to replicate their results.
I think this will mean cheaper, faster, and better models.
This FAQ about it is very good: https://stratechery.com/2025/deepseek-faq/
It is possible however that OpenAI was using similar level acceleration in the first place, they’ve just not published the details. And a few engineers left and replicated (or even bested it) in a new lab.
Overall, it’s a good boost, modern software is getting a better fit into new generation of hardware and is performing faster. Maybe we should pay more attention when NVIDIA is publishing their N-times faster ToPS numbers, and not completely dismissing it as marketing.
Personally this looks to me like an ego thing: the DeepSeek team are really, really good and their CEO is enjoying the enormous attention they are getting, plus the pride of proving that Chinese AI labs can take the lead in a field that everyone thought the USA was unassailable in.
Maybe they are true believers in building and sharing "AGI" with the world?
Lots of people see this as a Chinese government backed conspiracy to undermine the US AI industry. I'm not sure how credible that idea is.
I saw somewhere (though I've not confirmed it with a second source) that none of the people listed on the DeepSeek papers got educated at US universities - they all went to school in China, which further emphasizes how good China's home-grown talent pool has got.
"You have been educated at foreign universities / worked at foreign companies" is indeed an excuse they have used at least once to refuse a candidate. n=1 though so maybe that's just a convenient excuse. There's one guy who went to University of Adelaide (IIRC) on the paper.
Do you understand how ginormous China is and how ridiculous this kind of made up boogeyman statement sounds?
To me this sounds like describing Lockheed as a US government backed conspiracy to undermine the Tupolev Aerospace Design Bureau. It really stretches the normal connotations of words, and it presupposes that the center of the world is conveniently located very close to the speaker.
>Liang Wenfeng: In disruptive tech, closed-source moats are fleeting. Even OpenAI’s closed-source model can’t prevent others from catching up.
>Therefore, our real moat lies in our team’s growth—accumulating know-how, fostering an innovative culture. Open-sourcing and publishing papers don’t result in significant losses. For technologists, being followed is rewarding. Open-source is cultural, not just commercial. Giving back is an honor, and it attracts talent.
https://thechinaacademy.org/interview-with-deepseek-founder-...
More open than any other model (but still a bespoke licence) and bundles together a bunch of known improvements. There’s nothing to hide here honestly and without the openness it wouldn’t be as interesting.
'So are we close to AGI? It definitely seems like it. This also explains why Softbank (and whatever investors Masayoshi Son brings together) would provide the funding for OpenAI that Microsoft will not: the belief that we are reaching a takeoff point where there will in fact be real returns towards being first.'
Interesting.
I'm worried these technologies may take my job away and make the balance between capital and labor even more uneven.
Why should I be happy?
In the ideal case, we won't be dependent on the unwilling labor of other humans at all. Would you do your current job for free? If not -- if you'd rather do something else with your productive life -- then it seems irrational to defend the status quo.
One thing's for certain: ancient Marxist tropes about labor and capital don't bring any value to the table. Abandon that thinking sooner rather than later; it won't help you navigate what's coming.
We enjoy many luxuries unavailable even to billionaires only a few decades ago. For this trend to continue, the same thing needs to happen in other sectors that happened in (for example) the agricultural sector over the course of the 20th century: replacement of human workers by mass automation and superior organization.
If that were true they wouldn't be building ultra secure bunkers to escape to when the climate shit hits the fan.
Anecdotally, around two people in a hundred in my proximity are preppers as well, though obviously with smaller budgets.
It is just a specific fringe way of thinking.
I say it will be a Good Thing. "Work" is what you call whatever you're doing when you'd rather be doing something else.
Suppose you want to have your car washed. Hiring someone to do that will most likely give the best result: less physical resources used (soap, water, wear of cloth), less wear and tear on the car surface and less pollution and optionally a better result.
Still the benefit/cost equation is clearly in favor of the machine when doing the math, even when using more resources in the process.
What is lacking in our capitalist economic system is the fact of hiring people to perform services is punished by much higher taxes compared to using a machine, which is often even tax deductible. That way, the machine brings only benefits to the user of the machine (often a more wealthy person), less much to society as a whole. If only someone could find a solution to this tragedy.
Well, someone earlier in the thread said to abandon Marxist thought because it's obsolete. So I don't know how to help you!
We did. Save up a few bucks, nothing out of reach, and (as you suggested yourself!) you can afford to buy your own machine. Here you go: https://xcancel.com/carrigmat/status/1884244369907278106
You'd have received no such largesse from the Marxists. You're welcome.
You can keep shoehorning lazy political slurs into everything you post, but the reality is going to hit the working class, not privileged programmers casually dumping 6 grand so they can build their CRUD app faster.
But you're essentially arguing for Marxism in every other post on this thread, whether you realize it or not.
Perhaps other sites beckon.
I think it's interesting to note that as opens source models evolve and proliferate, the capital required for a lot of ventures goes down - which levels the playing field.
When I can talk to one agent-with-a-CAD-integration and have it design a gadget for me and ship the design off to a 3D printer and then have another agent write the code to run on the gadget, I'll be able to build entire ventures that would require VC funding and a team now.
When intellectual capital is democratized, financial capital looses just a bit of power...
At present, if you have financial capital and need intellectual capital you need to find people willing to work for you and pay them a lot of money. With enough progress in AI you can get the intellectual capital from machines instead, for a lot less. What loses value is human intellectual capital. Financial capital just gained a lot of power, it can now substitute for intellectual capital.
Sure, you could pretend this means you'll be able to launch a startup without any employees, and so will everyone. But why wouldn't Sam Altman or whomever just start AI Ycombinator with hundreds of thousands of AI "founders"? Do you really think it would be more "democratic"?
AI is useful in the same way with Linux
- can run locally
- empowers everyone
- need to bring your own problem
- need to do some of the work yourself
The moral is you need to bring your problem to benefit. The model by itself does not generate much benefits. This means AI benefits are distributed like open source ones.
Maybe you believe that they will always stay true, that there's some ineffable human quality that will never be captured by AI and value creation will always be bottle-necked by humans. That would be nice.
But even if you still need humans in the loop, it's not clear how "democratizing" this would be. It might sound great if in a few years you and everyone else can run an AI on their laptop that is as a good as a great technical co-founder that never sleeps. But note that means that someone who owns a data-center can run the equivalent of the current entire technical staff of Google, Meta, and OpenAI combined. Doesn't sound like a very level playing field.
But will there be a need for fewer engineers, though? That's the question. And the competition for those who remain employed would be fierce, way worse than today.
Or so I fear. I hope I'm wrong.
I fear that this won't age well. But to shamelessly riff on Marx, those who control the means of computation will control society.
You need to train on a fundamentally different task, which is to be good at the adversarial game of pursuing one's needs and desires in a social environment.
And that doesn't yet take into account that the interface to our lives is largely physical, we need bodies.
I'm seeing us on track to AGI in the sense of building a universal question answering machine, a system that will be able to answer any unambiguously stated question if given enough time and energy.
Stating questions unambiguously gets pretty difficult fast even where it's possible, often it isn't even possible, and getting those answers is just a small part of being a successful human.
PS: Needs and desires are totally orthogonal to AI/AGI. Every animal has them, but many animals don't have high intelligence. Needs and desires are a consequence of our evolutionary history, not our intelligence. AGI does not need to mean an artificial human. Whether to pursue or not pursue that research program is up to us, it's not inevitable.
We know this isn't far-fetched. We have strong evidence to suspect during the big layoffs of a couple of years ago, FAANG and startups all colluded to lower engineer salaries across the board, and that their excuse ("the economy is shrinking") was flimsy at best. Now AI presents them with another powerful tool to reduce salaries even more, with a side dish of reducing the size of the cost center that is programmers and engineers.
But yes, the job thing is concerning as well. AI won't scrub a toilet, but it will cheaply and inexhaustibly do every job that humans find meaningful today. It seems that we're heading inexorably towards dystopia.
That's the part I really don't believe. I'm open to being wrong about this, the risk is probably large enough to warrant considering it even if the probability of this happening is low, but I do think it's quite low.
We don't actually have to build artificial humans. It's very difficult and very far away. It's a research program that is related to but not identical to the research program leading to tools that have intelligence as a feature.
We should be, and in fact we are, building tools. I'm convinced that the mental model many people here and elsewhere are applying is essentially "AGI = artificial human", simply because the human is the only kind of thing in the world that we know that appears to have general intelligence.
But that mental model is flawed. We'll be putting intelligence in all sorts of places that are not similar to a human at all, without those devices competing with us at being human.
And further ahead, where I said your original take might not age well; I'm also not worried about AI making humanoid bodies. I'd be worried about a future where mines, factories, and logistics are fully automated: an AI for whom we've constructed a body which is effectively the entire planet.
And nobody needs to set out to build that. We just need to build tools. And then, one day, an AGI writes a virus and hacks the all-too-networked and all-too-insecure planet.
I know scifi is not authoritative, and no more than human fears made into fiction, but have you read Philip K. Dick's short story "Autofac"?
It's exactly what you describe. The AI he describes isn't evil, nor does it seek our extinction. It actually wants our well-being! It's just that it's taken over all of the planet's resources and insists in producing and making everything for us, so that humans have nothing left to do. And they cannot break the cycle, because the AI is programmed to only transition power back to humans "when they can replicate Autofac output", which of course they cannot, because all the raw resources are hoarded by the AI, which is vastly more efficient!
On the other hand, it's important not to pay too close attention to the details of scifi. I find myself writing a novel, and I'm definitely making decisions in support of a narrative arc. Having written the comment above... that planetary factory may very well become the third faction I need for a proper space opera. I'll have to avoid that PKD story for the moment, I don't want the influence.
Though to be clear, in this case, that potentiality arose from an examination of technological progress already underway. For example, I'd be very surprised if people aren't already training LLMs on troves of viruses, metasploit, etc. today.
I think we're talking about different time scales - I'm talking about the next few, maybe two or three decades, essential the future of our generation specifically. I don't think what you're describing is relevant on that time scale, and possibly you don't either.
I'd add though that I feel like your dystopian scenario probably reduces to a Marxist dystopia where a big monopolist controls everything.
In other words, I'm not sure whether that Earth-spanning autonomous system really needs to be an AI or requires the development of AI or fancy new technology in general.
In practice, monopolies like that haven't emerged due to competition and regulation, and there isn't a good reason to assume it would be different with AI either.
In other words, the enemies of that autonomous system would have very fancy tech available to fight it, too.
And while I want to agree that we won't see this happen in the next 3 decades, networked automated cars have already been deployed on the street of several cities and people are eagerly integrating LLMs into what seems to be any project that needs funding.
But it seems to me like you might not be sufficiently taking into account that this is an adversarial game; i.e. it's not sufficient for something just to replicate, it needs to also out-compete everything else decisively.
It's not clear at all to me why an AI controlled by humans, to the benefit of humans, would be at a disadvantage to an AI working against our benefit.
Making corporations more effective is not always in the interest of humans.
You might speculate about a one-person megacorp where everything is done by AIs that a single person runs.
What I'm saying is that we're very far from this, because the AI is not a human that can make the CEO's needs and desires their own and execute on them independently.
Humans are good at being humans because they've learned to play a complex game, which is to pursue one's needs and desires in a partially adversarial social environment.
This is not at all what AI today is being trained for.
Maybe a different way to look at it, as a sort of intuition pump: If you were that one man company, and you had an AGI that will correctly answer any unambiguously stated question you could ask, at what point would you need to start hiring?
The actual question, which is much more realistic, is if an average company of, let'say, 50 engineers will still have a need to hire those 50 engineers if AI turns out to be such an efficiency multiplier?
In that case, you will no longer need 10 people to complete 10 tasks in given time-unit but perhaps only 1 engineer + AI compute to do the same. Not all businesses can continue scaling forever, so it's pretty expected that those 9 engineers will become redundant.
What I was getting at was the question: If we feel intuitively that this extreme isn't realistic, what exactly do we think is missing?
My argument is, what's missing is the human ability to play the game of being human, pursuing goals in an adversarial social context.
To your point more specifically: Yes, that 10-person team might be replaceable by a single person.
More likely than not however, the size of the team was not constrained by lack of ideas or ambition, but by capital and organizational effectiveness.
This is how it's played out with every single technology so far that has increased human productivity. They increase demand for labor.
Put another way: Businesses in every industry will be able to hire software engineering teams that are so good that in the past, only the big names were able to afford them. The kind of team required for the digital transformation of every old fashioned industry.
Your hypothesis is AFAIU is that the company will just continue to scale because there's an indefinite amount of work/ideas to be explored/done so the focus of those 9 people will just be shifted to some other topic?
Let's say I am a business owner I have a popular product with a backlog of 1000 bugs and I have a team of 10 engineers. Engineers are busy both juggling between the features and fixing the bugs at the same time. Now let's assume that we have an AI model that will relieve 9 out of 10 engineers from cleaning the bugs backlog and we will need 1 or 2 engineers reviewing the code that the AI model spits out for us.
What concrete type of work at this moment is left for the rest of the 9 engineers?
Assuming that the team, as you say, is not constrained by the lack of ideas or ambition, and the feature backlog is somewhat indefinite in that regard, I think that the real question is if there's a market for those ideas. If there's no market for those ideas then there's no business value $$$ created by those engineers.
In that case, they are becoming a plain cost so what is the business incentive to keep them then?
> Businesses in every industry will be able to hire software engineering teams that are so good that in the past, only the big names were able to afford them
Not sure I follow this example. Companies will still hire engineers but IMO at much less capacity than what it was required up until now. Your N SQL experts are now replaced by the model. Your M Python developers are now replaced by the model. Your engineer/PR-review is now replaced by the model. The heck, even your SIMD expert now seems to be replaced by the model too (https://github.com/ggerganov/llama.cpp/pull/11453/files). Those companies will no longer need M + N + ... engineers to create the business value.
Yes, that's what I'm saying, except that this would hold over an economy as a whole rather than within every single business.
Some teams may shrink. Across industry as a whole, that is unlikely to happen.
The reason I'm confident about this is that this exact discussion has happened many times before in many different industries, but the demand for labor across the economy as a whole has only grown. (1)
"This time it's different" because the productivity tech in question is AI? That gets us back to my original point about people confusing AI with an artificial human. We don't have artificial humans, we have tools to make real humans more effective.
(1) The point seems related to this https://en.wikipedia.org/wiki/Lump_of_labour_fallacy
My question is rather of much narrower scope and much more concrete and tangible - and yet I haven't been able to find any good answer for it, or strong counter-arguments if you will. If I had to guess something about it then my prediction would be that many engineers will need to readjust their skills or even requalify for some other type of work.
What higher value add professions will humans be displaced into by AI?
LLMs do not have desires, but their existence alters desires of humans, including the ones in charge of businesses.
One force is a multiplier of a software engineer’s productivity.
Another force is the pressure of the expectation for constant, unlimited increase in profits. This pressure force the CEOs and managers to look for cheaper alternatives to expensive software engineers, ultimately to eliminate the position and expense. The lie that this is a possibility draws huge investments.
And another force is the infinite number of applications of software, especially well designed, truly useful, software.
I'd be a hypocrite if I didn't admit I use AI daily in my job, and it's indeed a multiplier of my productivity. The tech is really cool and getting better.
I also understand AI is one step closer for the everyday Jane or Joe Doe to do cool and useful stuff which was out of reach before.
What worries me is the capitalist, business-side forces at play, and what they will mean for my job security. Is it selfish? You bet! But if I don't advocate for me, who will?
But how will this translate to engineering jobs? Maybe there will be AI tools to automate most of the stuff a small business needs done. "Ah," you may say, "I will build those tools!". Ok. Maybe. How many engineers do you need for that? Will the current engineering job market shrink or expand, and how many non-trash, well paid jobs will there be?
I'm not saying I know for sure how it'll go, but I'm concerned.
We are far away from that though. As an enterprise software/data engineer, AI has been great in answering questions and generating tactical code for me. Hours have turned into minutes. It even motivated me to work on side projects because they take less time. You will be fine. Embrace the change. Its good for you. Will lead to personal growth.
Also, I don't want to be a glorified uber driver. It's not good for me and not good for the profession.
> As an enterprise software/data engineer, AI has been great in answering questions and generating tactical code for me. Hours have turned into minutes.
I don't dispute this part, and it's been this way for me too. I'm talking about the future of our profession, and our job security.
> You will be fine. Embrace the change. Its good for you. Will lead to personal growth.
We're talking at cross-purposes here. I'm concerned about job security, not personal growth. This isn't about change. I've been almost three decades in this profession, I've seen change. I'm worried about this particular thing.
By the way, car mechanics (especially independent ones, your average garage mechanic) understand less and less about what's going on inside modern cars. I don't want this to happen to us.
Of course assemblers didn't create fewer programming jobs, nor did compilers or high level languages. However, with "NO CODE" solutions (remember that fad?) there was an attempt at reducing the need for programmers (though not completely taking them out of the equation)... it's just that NO CODE wasn't good enough. What if AI is good enough?
It doesn't matter how "easy" technology gets to use, there will always be a market for helping other people figure out best to apply it.
The way I look at this is that with the release of something like deepseek the possibility of running a model offline and locally to work _for_ you while you are sleeping, doing groceries, spending time with your kids / family is coming closer to a reality.
If AI is able to replace me one day I'll be taking advantage of that way more efficiently than any of my employee(s).
I don't know when the threshold of "replace the bottom X% of developers because AI is so good" happens for businesses based on those things, but it's definitely getting closer instead of stalling out like the bubble predictors claimed. It's not a bubble if the industry is making progress like this.
Making something work really efficiently on older hardware doesn't necessarily imply less demand. If those lessons can be taken and applied to newer generations of hardware, it would seem to make the newer hardware all the more valuable.
It's... good. Even the qwen/llama distills are good. I've been running the Llama-70b-distill and it's good enough that it mostly replaces my chatgpt plus plan (not pro - plus).
I think if anything - One of my big takeaways is that OpenAI shot themselves in the foot, big time, by not exposing the COT for the O1 Pro models. I find the <think></think> section of the DeepSeek models to often be more helpful than the actual answer.
For work that's treating the AI as collaborative rather than "employee replacement" the COT output is really valuable. It was a bad move for them to completely hide it from users, especially because they make the user sit there waiting while it generates anyways.
However this has huge implications when it comes to the feasibility and spread of the technology, and further implications with regards to economy and geopolitics now that confidence in the American AI sector has been hit and people and organizations internationally have somewhere else to look for.
edit: That being said, this is the first time I've seen a LLM do a better job than even a senior expert could do, and even if it's on small scope/in a limited context, it's becoming clear that developers are going to have to adopt this tech in order to stay competitive.
Second, the fact that deepseek was able to pull this off with such modest resources is an indication that there is no moat, and you might wake up tomorrow and find an even better model from a company you have never heard of.
https://finance.yahoo.com/news/deepseek-temu-ai-analysts-132...
People are already looking at it like Temu.
> max_tokens:The maximum length of the final response after the CoT output is completed, defaulting to 4K, with a maximum of 8K. Note that the CoT output can reach up to 32K tokens, and the parameter to control the CoT length (reasoning_effort) will be available soon. [1]
I'm quite new to this, how are you feeding in so much text? just copy/paste? I'd love to be able to run some of my Zig code through it, but I haven't managed to get Zig running under Asahi so far.
ollama run deepseek-r1:32b
They dropped the Qwen/Llama terms from the stringhttps://ollama.com/library/deepseek-r1:32b
https://ollama.com/library/deepseek-r1:32b-qwen-distill-q4_K...
I prefer to use the longer name, so I know which model I'm running. In this particular case, it's confusing that they grouped the qwen and llama fine tunes with R1, because they're not R1.
I only use it for chatting about the code - while this setup also lets the AI edit your code, I don't find the code good enough to risk it. I get more value from reading the thought process, evaluating it, and the cherry picking which bits of its code I really want.
In any case, if that sounds like the experience you want and you already run ollama, you would just need to install the continue.dev VS Code extension, and then go to its settings to configure which models you want in the drop-down.
ollama run hf.co/MaziyarPanahi/Mistral-7B-Instruct-v0.3-GGUF:IQ1_MAlso serve it as MLX from LMStudio, will speed things up 30% or so so your 6bit will have similar perf to the 4bit.
Getting about 12-13 tok/sec on my M3 Max 48gb.
Our enterprise internal IT has a low friction way to request a Mac Studio (192GB) for our team and it’s a wonderful central EXO endpoint. (Life saver when we’re generally GPU poor)
in the github issue he links he gives an example of a prompt: Your task is to convert a given C++ ARM NEON SIMD to WASM SIMD. Here is an example of another function: (follows a block example and a block with the instructions to convert)
https://gist.github.com/ngxson/307140d24d80748bd683b396ba13b...
I might be wrong of course, but asking to optimize code is something that quite helped me when i first started learning pytorch. I feel like "99% of this code blabla" is useful as in it lets you understand that it was ai written, but it shouldn't be a brag. then again i know nothing about simd instructions but i don't see why it should be different for a capable llm to do simd instructions or optimized high level code (which is much harder than just working high level code, i'm glad i can do the latter lol)
It's still cool nonetheless, but not a particularly great test of DeepSeek vs. alternatives.
That said, translating good code to another language or environment is extremely useful. There’s a lot of low hanging fruit where there’s, for example, an existing high quality library is written for Python or C# or something, and an LLM can automatically convert it to optimized Rust / TypeScript / your language of choice.
Q: "It only does conversion ARM NEON --> WASM SIMD, or it can invent new WASM SIMD code from scratch?"
A: "It can do both. For qX_0 I asked it to convert, and for qX_K I asked it to invent new code."
* [1]: https://gist.github.com/ngxson/307140d24d80748bd683b396ba13b...
That seems like a notable milestone.
Yes, but:
"For the qX_K it's more complicated, I would say most of the time I need to re-prompt it 4 to 8 more times.
The most difficult was q6_K, the code never works until I ask it to only optimize one specific part, while leaving the rest intact (so it does not mess up everything)" [0]
And also there:
"You must start your code with #elif defined(__wasm_simd128__)
To think about it, you need to take into account both the refenrence code from ARM NEON and AVX implementation."
[0] https://gist.github.com/ngxson/307140d24d80748bd683b396ba13b...
I do not understand why GGML is written this way, though. So much duplication, one variant per instruction set. Our Gemma.cpp only requires a single backend written using Highway's portable intrinsics, and last I checked for decode on SKX+Zen4, is also faster.
I hope we can put to rest the argument that LLMs are only marginally useful in coding - which are often among the top comments on many threads. I suppose these arguments arise from (a) having used only GH copilot which is the worst tool, or (b) not having spent enough time with the tool/llm, or (c) apprehension. I've given up responding to these.
Our trade has changed forever, and there's no going back. When companies claim that AI will replace developers, it isn't entirely bluster. Jobs are going to be lost unless there's somehow a demand for more applications.
But I think it’s still much too early for any form of “can we all just call it settled now? In this case, as we all know, lines of code is not a useful metric. How many person hours were spent doing anything associated with this PR’s generation and how does that compare to not using AI tools, and how does the result compare in terms of the various forms of quality? That’s the rubric I’d like to see us use in a more consistent manner.
Usually, there are hidden payoffs that motivate things that seem like waste.
Hiring someone to remodel a bathroom is hard enough, now try hiring a contract software engineer, especially when you don't have budget authority!
That said, I heard about a fire chief last year who had to spend two days manually copying and pasting from one CRM to another. I wish I could help people like that know when to pay someone to write a script!
I imagine even in that role figuring out how to hire someone so solve a problem would still take longer than manually crunching through that themselves.
Individual developer productivity will be expected to rise. Timelines will shorten. I don't think we've reached Peak Software where the limiting factor on software being written is demand for software, I think the bottlenecks are expense and time. AI tools can decrease both of those, which _should_ increase demand. You might be expected to spend a month outputting a project that would previously have taken four people that month, but I think we'll have more than enough demand increase to cover the difference. How many business models in the last twenty years that weren't viable would've been if the engineering department could have floated the company to series B with only a half dozen employees?
What IS larger than before, IMO, is the talent gap we're creating at the top of the industry funnel. Fewer juniors are getting hired than ever before, so as seniors leave the industry due to standard attrition reasons, there are going to be fewer candidates to replace them. If you're currently a software engineer with 10+ YoE, I don't think there's much to worry about - in fact, I'd be surprised if "was a successful Software Engineer before the AI revolution" doesn't become a key resume bullet point in the next several years. I also think that if you're in a position of leadership and have the creativity and leadership to make it work, juniors and mid-level engineers are going to be incredibly cost effective because most middle managers won't have those things. And companies will absolutely succeed or fail on that in the coming years.
Something I think about a lot is the impact of open source on software development.
25 years ago any time you wanted to build anything you pretty much had to solve the same problems as everyone else. When I went to university it even had a name - the software reusability crisis. At the time people thought the solution was OOP!
Open source solved that. For any basic problem you want to solve there are now dozens of well tested free libraries.
That should have eliminated so many programming jobs. It didn't: it made us more productive and meant we could deliver more value, and demand for programmers went up.
This is a key insight - the trade has changed.
For a long time, hoarding talent - who could conceive and implement such PRs - was a competitive advantage. It no longer is because companies can hire and get similar outcomes, with fewer and mediocre devs.
But at the same time, these companies have lost their technological moat. The people were the biggest moat. The hoarding of people were the reason why SV could stay ahead of other concentrated geographies. This is why SV companies grew larger and larger.
But now, anyone anywhere can produce anything and literally demolish any competitive advantage of large companies. As an example, literally a single Deepseek release yesterday destroyed large market cap companies.
It means that the future world is likely to have a large number of geographically distributed developers, always competing, and the large companies will have to shed market cap because their customers will be distributed among this competition.
It's not going to be pleasant. Life and work will change but it is not merely loss of jobs but it is going to be loss of the large corporation paradigm.
Nobody was “destroyed” - a handful of companies had their stock price drop, a couple had big drops, but most of those stocks are up today, showing that the market is reactionary.
The point still is: Software/engineering is no longer the moat creator.
Notice how Apple and Meta stocks went up last 2 days?
Apple has a non-software moat: Their devices.
Meta has a non-software moat: their sticky users.
So does Microsoft, and Google to an extent with their non-software moat.
But how did they build the most in the first place? With software that only they could develop, at a pace that only they could execute, all because of the people they could hoard.
The companies of the future can disrupt all of them (maybe not apple) very quickly by just developing the same things as say Meta and "at the same quality" but for cheaper. The engineers moat is gone. The only moat meta has is network effects. That's one less barrier for a competing company to deal with.
Do you have a link to that?
Then in the 00s, the computing resources became widely available. The bottleneck was the people who could build interesting things. Imagine a third world country with access to AWS but no access to developers who could build something meaningful.
With these models, now these geographically distributed companies can build similarly high quality stuff.
R1 IS the example of something that previously only could be built in the bowels of large SV corporations.
That's why I'm not worried. There is already SO MUCH more demand for code than we're able to keep up with. Show me a company that doesn't have a backlog a mile long where most of the internal conversations are about how to prioritize what to build next.
I think LLM assistance makes programmers significantly more productive, which makes us MORE valuable because we can deliver more business value in the same amount of time.
Companies that would never have considered building custom software because they'd need a team of 6 working for 12 months may now hire developers if they only need 2 working for 3 months to get something useful.
AI agents will do the same job.
What will still matter is software that constrains what kind of data ends up in the database and ensures that data means what it is supposed to. That software will be created by local teams that know the business and the data. They will use AI to write the software and test it. Will those teams be “developers”? It is probably semantics or a matter of degree. Half the people writing advanced Excel spreadsheets today should probably be considered developers really.
Programming languages are languages to tell the computer what to do. In the beginning, people wrote in machine code. Then, high level languages like C and FORTRAN were invented. Since then we’ve been iterating on the high level language idea.
These LLM based tools seem to be a more abstract way of telling the computer what to do. And they really might, if they work out, be a jump similar to the low/high level split. Maybe in the future we’ll talk about low-level, high-level, and natural programming languages. The only awkwardness will be saying “I have to drop down to a high level language to really understand what the computer is doing.” But anyway, there were programmers on either side of that first split (way more after), if there’s another one I suspect there will still be programmers after.
Obviously, source allows human tuning, auditing, and so on. But taken at the limit, those aspects may eventually no longer be necessary. Just a riff here, as the thought just occurred.
I worry about junior developers. It will be a while before vocational programming courses retool to teach this new way of writing code, and these are going to be testing times for so many of them. If you ask me why this will take time, my argument is that effectively wielding an LLM for coding requires broad knowledge. For example, if you're writing web apps, you need to be able to spot say security issues. And various other best practices, depending on what you're making.
It's a difficult problem to solve, requiring new sets of books, courses etc.
The ones who are self-starters will do fine - they'll figure out how to accelerate their way up the learning curve using these new tools.
People who prefer classroom-learning / guided education are going to be at a disadvantage for a few years while the education space retools for this new world.
I’d love to watch, e.g. you Simon, using these tools. I assume there are so many little tricks you figured out over time that together make a big difference. Things that come to mind:
- how to quickly validate the output?
- what tooling to use for iterating back and forth with the LLM? (just a chat?)
- how to steer the LLM towards a certain kind of solutions?
- what is the right context to provide to the LLM? How do it technically?
This isn't to say you are wrong but just to put some perspective on how things are changing. Maybe most new programmers will be hired into AI roles or data science.
I didn’t get a smart phone until the 2010s. Stupid I know but it was seen as a badge of honour in some circles ‘bah I don’t even use a smart phone’ we’d say as the young crowd went about their lives never getting lost without a map and generally having an easier time of it since they didn’t have that mental block.
Ai is going to be similar no doubt. I’m already seeing ‘bah I don’t use ai coding assistants’ type of posts, wearing it as a badge of honour. ‘Ok you’re making things harder for yourself’ should be the reply but we’ll no doubt have people wearing it as a badge of honour for some time yet.
I appreciate anyone that can utilise ai well but there’s just not enough core ai model development jobs for every new grad.
What are those “day to day business needs” that you think people are going to do without AI?
In my view, this is like 1981. If you are saying, we will still need non-computer people for day-to-day business needs, you are wrong. Even the guy in the warehouse and the receptionist at the front are using computers. So is the CEO. That does not mean that everybody can build one, but just think of the number of jobs in a modern company that require decent Excel skills. It is not just the one in finance. We probably don’t know what the “Excel” of AI is just yet but we are all going to need to be great at it, regardless of who is building the next generation of tools.
I would assume any other CS course teaches/is going to be teaching how to use AI to be an effective software developer.
Everybody is a Manager now?
Now, the first response some folks may have is, how can you trust that the AI is good at security? Well, in this example, it only needs to be better than the junior developers at security to provide them with benefits/learning opportunities. We need to remember that the junior developers of today can also just as easily write insecure code.
The mantra has always been that the best way to learn to code is to read other people’s code. Now you can have “other people” write you code for whatever you want. You can study it and see how it works. You can explore different ways of accomplishing the same tasks. You can look at the similar implementations in different languages. And you may be able to see the reasoning and research for it all. You are never going to get that kind of access to senior devs. Most people would never work up the courage to ask. Plus, you are going to become wicked good at using the AI and automation including being deeply in touch with its strengths and weaknesses. Honestly, I am not sure how older, already working devs are going to keep up with those that enter the field 3 years from now.
Instead of this, have you considered asking Deep Seek to explain it to you?
For mankind, the really big problems aren't going away any time soon.
But -- and it's a big but -- many of us aren't working on those problems. I'm ready to agree most of what I've done for decades in my engineering job(s) is largely inconsequential. I don't delude myself into thinking I'm changing the world. I know I'm not!
What I'm doing is working on something interesting (not always) while earning a nice paycheck and supporting my family and my hobbies. If this goes away, I'll struggle. Should the world care? Likely not. But I care. And I'm unlikely to start working on solving societal problems as a job, it's too much of a burden to bear.
And here is the problem: AI needs to be trained on something. Use of AI reduces the use of online forums, some of them are actively blocking access, like reddit. So, for AI to stay relevant it has to generate the knowledge by itself. Like having full control of a computer, taking queries from human supervisor, and really trying to solve. Having this sort of AI actors in online forum will benefit everyone.
Really, what seems on the horizon is a cliff of techno risks that have nothing to do with "AI will take over the world" and more "AI will be so integral to functional humanity that actual risks become so diffuse that no one can stop it."
So it's more a conceptual belief: Will AI actually make driving cares safer or will the fatalities of AI just be so randomly stochastic that it's more acceptable.
I would argue that we already accept relatively random car fatalities at a huge scale and simply engage in post-hoc rationalization of the why and how of individual accidents that affect us personally. If we can drastically reduce the rate of accidents, the remaining accidents will be post-hoc rationalized the same way we always have rationalized accidents.
Currently, car crashes are blamed on the individuals involved.
I reckon we'll see a lot of new religious thinking about this stuff
This is about the functional society where people fundamentally have recourse to "blame" via legal means one another for things.
Having fallbacks, eg, pilots in the cockpit is not a long term strategy for AI pilots flying planes because they functionally will never be sufficiently trained for actual scenarios.
I don't agree. LLMs work as template engines on steroids. The role of a developer now includes more code reviewing than code typing. You need the exact same core curriculum to be able to parse code, regardless if you're the one writing it, it's a PR, or it's outputted by a chatbot.
> For example, if you're writing web apps, you need to be able to spot say security issues. And various other best practices, depending on what you're making.
You're either overthinking it or overselling it. LLMs generate code, but that's just the starting point. The bulk of developer's work is modifying your code to either fix an issue or implement a feature. You need a developer to guide the approach.
It isn't. Anyone who does software development for a living can explain to you what exactly is the day-to-day work of a software developer. It ain't writing code, and you spend far more time reading code than writing it. This is a known fact for decades.
> If the IDE checks types and feeds errors back to the LLM,(...)
Irrelevant. Anyone who does software development for a living can tell you that code review is way more than spotting bugs. In fact, some companies even have triggers to only trigger PR reviews if all automated tests pass.
Extremely well paid human coders, capable of fixing the mistakes of the years preceding them...
I'm by no means an expert, but I feel like you still need someone who understands underlying principles and best practices to create something of value.
Thinking it out even further, programming languages will likely go away altogether as ultimately they're just human interfaces to machine language.
As we know them, certainly.
I haven't seen discussions about this (links welcome!), but I find it fascinating.
What would a PL look like, if it was not designed to be written by humans, but instead be some kind of intermediate format generated by an AI for humans to review?
It would need to be a kind of formal specification. There would be multiple levels of abstraction -- stakeholders and product management would have a high level lens, then you'd need technologists to verify the correctness of details. Parts could still be abstracted away like we do with libraries today.
It would be way too verbose as a development language, but clear and accessible enough that all of our arcane syntax knowledge would be obsolete.
This intermediate spec would be a living document, interactive and sensitive to modifications and aware of how they'd impact other parts of the spec.
When the modifications are settled, the spec would be reingested and the AI would produce "code", or more likely be compiled directly to executable blobs.
...
In the end, I still think this ends up with really smart "developers" who don't need to know a lick of code to produce a full product. PLs will be seen as the cute anachronisms of an immature industry. Future generations will laugh at the idea that anybody ever cared about tabs-v-spaces (fair enough!).
Take for example neuralink. If you consider that interface 10 years, or further 1000 years out in the future, it's likely we will have a direct, thought-based human computer interface. Which is interesting when thinking of this for sending information to the computer, but even more so (if equally alarming) for information flowing from computer to human. Whereas today, we read text on web pages, or listen to audio books, in that future, we may instead receive felt experiences / knowledge / wisdom.
Have you had a chance to read 'Metaman: The Merging of Humans and Machines into a Global Superorganism' from 1993?
This is a problem that the Computer Science departments of the world have been solving. I think that the "good" departments already go for the "broad knowledge" of theory, systems with a balance between the trendy and timeless.
Why?
Are they not already?
> It's a difficult problem to solve, requiring new sets of books, courses etc.
I think new tooling built around LLMs that fits into our current software development lifecycle is going to make a big difference. I am experiencing firsthand how much more productive I am with LLM, and I think that in the future, we will start using "Can you review my conversation?" in the same way we use "Can you review my code?"
Where I believe LLMs are a real game changer is they make it a lot easier for us to consume information. For example, I am currently working on adding a Drag and Drop feature for my chat input box. If a junior developer is tasked with this, the senior developer can easily have the LLM generate a summary of their conversation like so:
https://beta.gitsense.com/?chat=d36e0282-4326-46cf-83b1-4207...
At this point, the senior developer can see if anything is missed; if desired, they can fork the conversation to ask the LLM questions like "Was this asked?" or "Was this mentioned?"
And once everybody is happy, you can have the LLM generate a PR title and message like so:
https://beta.gitsense.com/?chat=8aa19528-5891-4dda-9a88-247a...
All of this took me about 10 minutes, which would have taken me an hour or maybe more without LLMs.
And from here, you are now ready to think about coding with or without LLM.
I think with proper tooling, we might be able to accelerate the learning process for junior developers as we now have an intermediate layer that can better articulate the senior developers' thoughts. If the junior developer is too embarrassed to ask for clarification on why the senior developer said what they did, they can easily ask the LLM to explain.
The issue right now is that we are so focused on the moon shots for LLM, but the simple fact is that we don't need it for coding if we don't want to. We can use it in a better way to communicate and gather requirements, which will go a long way to writing better code faster.
Even the discussion around AI partially replacing coders is a direction towards commoditization.
We saw it crystal clear between the boom years, the trough, and the current recovery.
And yet many companies aren't hiring developers right now - folks in the C suite are thinking AI is going to be eliminating their need to hire engineers. Also "demand" doesn't necessarily mean that there's money available to develop this code. And remember that when code is created it needs to be maintained and there are costs for doing that as well.
I'd love to see numbers around the "execs don't think they need engineers because of AI" factor. I've heard a few anecdotal examples of that but it's hard to tell if it's a real trend or just something that catches headlines.
Another data point is that there's been ~10 companies that I have been following and all of them have been shut down in the past year or so.
And the general feeling you get from the number of HN posts from people complaining about not being able to find jobs. This certainly hasn't been like that before.
If you’re infinitely productive, then the solution to every problem is to just keep producing stuff, instead of learning to say no.
This means a lot of companies will overbuild, and then drown in maintenance problems and fail catastrophically when they can’t keep up.
And this kind of fear mongering is particularly irritating when you see that our industry already faced a similar productivity shock less than twenty years ago: before open source went mainstream github and library hubs like npm we used to code the same things over and over again, most of the time in a half-backed fashion because nobody had time for polishing stuff that was needed but only tangentially related to the code business. Then came the open-source tsunami, and suddenly there was a high quality library for solving your particular problem and the productivity gain was insane.
Fast forward a few years, does it look like this productivity gains took any of our jobs? Quite the opposite actually, there has never been as many developers as today.
(Don't get me wrong, this is massively changing how we work, like the previous revolution did, and how job is never going to be the same again)
Aka Devs can move up the chain into what was traditionally product roles to increase development of new projects. This is using the time they have regain from more menial tasks being automated away.
There is a lot of work. Plenty of it just isnt super fun or interesting.
These might not be big products, but who wants big products anyway? You always have to bend over backwards to trick them into doing what you want. You should see the crazy stuff my partner does to make google docs fit her use case...
Let's have an era of small products made by people who are close to the problems being solved.
I don't trust companies to translate that to, "We can do more now" rather than, "We can do more with less people now" though.
We really are in AI moment of iPhone. I never thought I would witness something bigger than the impact of Smartphone. There are insane amount of value that we could extract out. Likely in tens of trillions from big to small business.
We keep asking how Low Code or No Code "tools" could achieve custom apps. Turns out we are here via a different route.
>custom software because they'd need a team of 6 working for 12 months may now hire developers if they only need 2 working for 3 months to get something useful.
I am wondering if it be more like 2 working for 1 month?
Maybe devs will be replaced with QA, or become glorified QA themselves.
Most companies don't have a milelong backlog of coding projects. That's a uniquely tech industry-specific issue, and a lot of it is driven by the tech industry's obsessive compulsion to perpetually reinvent wheels.
Companies that would never have considered building custom software because they'd need a team of 6 working for 12 months may now hire developers if they only need 2 working for 3 months to get something useful.
No, because most companies that can afford custom software want reliable software. Downtime is money. Getting unreliable custom software means that the next time around they'll just adapt their business processes to software that's already available on the market.
This is viewing things too narrowly I think. Why do we even need most of our current software tools aside from allowing people to execute a specific task? AI won't need VSCode. If AI can short circuit the need for most, if not nearly all enterprise software, then I wouldn't expect software demand to increase.
Demand for intelligent systems will certainly increase. And I think many people are hopeful that you'll still need humans to manage them but I think that hope is misplaced. These things are already approaching human level intellect, if not exceeding it, in most domains. Viewed through that lens, human intervention will hamper these systems and make them less effective. The rise of chess engines are the perfect example of this. Allow a human to pair with stockfish and override stockfish's favored move at will. This combination will lose every single game to a stockfish-only opponent.
Why not? It's still going to be quicker for the AI to use automated refactoring tooling than to manually make all the changes itself.
This was all in the default web chat UI.
But the bit of data we got in this story is that a human wrote tests for a human-identified opportunity, then wrote some prompts, iterated on those prompts, and then produced a patch to be sent in for review by other humans.
If you already believed that there might be some fully autonomous coding going on, this event doesn’t contradict your belief. But it doesn’t really support it either. This is another iteration on stuff that’s already been seen. This isn’t to cheapen the accomplishment. The range of stuff these tools can do is growing at an impressive rate. So far though it seems like they need technical people good enough to define problems for them and evaluate the output…
https://chris-granger.com/2015/01/26/coding-is-not-the-new-l...
modelling has been , is , and will be the needed literacy..
No, work is never the core problem. Backlog of bug fixes/enhancements is rarely what determines the headcount. What matters is the business need. If the product sells and there is no/little competition, the company has very little incentive to improve their products, especially hiring people to do the work. You'd be thankful if a company does not layoff people in teams working on mature products. In fact, the opposite has been happening, for quite a while. There are so many examples out there that I don't need to name them.
India and Eastern EU will win far more (relatively) than expensive devs in the US or Western EU.
Admittedly, I haven't looked too hard, but how could I do that with a model from, say, Ollama and run exclusively on my machine?
After a very quick reading of their pages it does not seem so.
Hopefully I am wrong...
It's a bit like what happens with "illusion of knowledge" or "illusion of understanding". When one knows the topic, one can correct the output of AI. When one doesn't, one tends to forget it can be inaccurate or plain wrong.
Look at the code that was changed[0]. It's a single file. From what I can tell, it's almost purely functional with clearly specified inputs and outputs. There's no need to implement half the code, realize the requirements weren't specified properly, and go back and have a conversation with the PM about it. Which is, you know, what developers actually do.
This is the kind of stuff LLMs are great at, but it's not representative of a typical change request by Java Developer #1753 at Fortune 500 Enterprise Company #271.
(That's not to say it isn't a valid argument.)
Short answer: LLMs are amazingly useful on large codebases, but they are useful in different ways. They aren't going to bang out a new feature perfectly first time, but in the right hands they can dramatically accelerate all sorts of important activities, such as:
- Understanding code. If code has no documentation, dumping it into an LLM can help a lot.
- Writing individual functions, classes and modules. You have to be good at software architecture and good at prompting to use them in this way - you take on the role of picking out the tasks that can be done independently of the rest of the code.
- Writing tests - again, if you have the skill and experience to prompt them in the right way.
> Writing individual functions, classes and modules. You have to be good at software architecture and good at prompting to use them in this way - you take on the role of picking out the tasks that can be done independently of the rest of the code.
If you have enough skill and understanding to do this, it means you already have enough general software development experience and domain-specific experience and experience with a specific, existing codebase to be in rarefied air. It's like saying, oh yeah a wrench makes plumbing easy. You just need to turn the wrench, and 25 years of plumbing knowledge to know where to turn it.
> Writing tests - again, if you have the skill and experience to prompt them in the right way.
This is very true and more accessible to most developers, though my big fear is it encourages people to crap out low-value unit tests. Not that they don't love to do that already.
Yes, exactly. That's why I keep saying that software developers shouldn't be afraid that they'll be out of a job because of LLMs.
That's not at all what the GP was saying, though:
> There's no need to implement half the code, realize the requirements weren't specified properly, and go back and have a conversation with the PM about it. Which is, you know, what developers actually do.
> This is the kind of stuff LLMs are great at, but it's not representative of a typical change request by Java Developer #1753 at Fortune 500 Enterprise Company #271.
This is a task that would likely have taken as long to write by hand as the AI took to do it, given how long the actual task took to execute. 98% of the work is find and replace
Don't get me wrong - this kind of thing is useful and cool, but you're mixing up the easy coding donkey work with the stuff that takes up time
If you look at the actual prompt engineering part, its clear that this prompting produced extensively wrong results as well, which is tricky. Because it wasn't produced by a human, it requires extensive edge case testing and review, to make sure that the AI didn't screw anything up. If you have the knowledge to validate the output, it would have been quicker to write it by hand instead of reverse engineering the logic by hand. Its bumping the work off from writing it by hand, to the reviewers who now have to check your ML code because you didn't want to put in the work by hand
So overall - while its extremely cool that it was able to do this, it has strong downsides for practical projects as well
I don’t get why people don’t understand that everything decomposes into other things.
You can draw the line for when AI will truly blow your mind anywhere you want, the point is the dominoes keep falling relentlessly and there’s no end in sight.
For example: AI's smash translation. They won't ever beat out humans, but as an automated solution? They rock. Natural language processing in general is great. If you want to smush in a large amount of text, and smush out a large amount of other text that's 98% equivalent but in a different structure, that's what AI is good for. Same for audio, or picture manipulation. It works because it has tonnes of training data to match your input against
What AI cannot do, and will never be able to do, is take in a small amount of text (ie a prompt), and generate a large novel output with 100% accuracy. It simply doesn't have the training data to do this. AI excels in tasks where it is given large amounts of context and asked to perform a mechanistic operation, because its a tool which is designed to extract context and perform conversions based on that context due to its large amounts of training data. This is why in this article the author was able to get this to work: they could paste in a bunch of examples of similar mechanical conversions, and ask the AI to repeat the same process. It has trained on these kinds of conversions, so it works reasonably well
Its great at this, because its not a novel problem, and you're giving it its exact high quality use case: take a large amount of text in, and perform some kind of structural conversion on it
Where AI fails is when being asked to invent whole cloth solutions to new problems. This is where its very bad. So for example, if you ask an AI tool to solve your business problem via code, its going to suck. Because unless your business problem is something where there are literally 1000s examples of how to solve it, the AI simply lacks the training data to do what you ask it, it'll make gibberish
It isn't the nature of the power of the AI, its that its inherently good for solving certain kinds of problems, vs other kinds of problems. It can't be solved with more training. The OPs problem is a decent use case for it. Most coding problems aren't. That's not that it isn't useful - people have already been successfully using them for tonnes of stuff - but its important to point out that its only done so well because of the specific nature of the use case
Its become clear that AI requires someone of equivalent skill as the original use case to manage its output if 100% accuracy is required, which means that it can only ever function as an assistant for coders. Again, that's not to say it isn't wildly cool, its just acknowledging what its actually useful for instead of 'waiting to have my mind blown'
Like all the people surprised by Deepseek when it has been clear for the last 2 years there is no moat in foundation models and all the value is in 1) high quality data that becomes more valuable as the internet fills with AI junk 2) building the UX on top that will make specific tasks faster.
There are probably 10% of truly novel problems out there, the rest are just already solved problems with slightly different constraints of resources ($), quality (read: reliability) and time. If LLMs get good enough at generating a field of solutions that minimize those three for any given problem, it will naturally tend to change the nature of most software being written today.
But there's also bespoke problems. They aren't quite novel, yet are complicated and require a lot of inside knowledge on business edge cases that aren't possible to sum up in a word document. Having worked with a lot of companies, I can tell you most businesses literally cannot sum up their requirements, and I'm usually teaching them how their business works. These bespoke problems also have big implications on how the app is deployed and run, which is a whole different thing.
Then you have LLMs, which seem allergic to requirements. If you tell an LLM "make this app, but don't do these 4 things," it's very different from saying "don't do these 12 things." It's more likely to hallucinate, and when you tell it to please remember requirement #3, it forgets requirement #7.
Well, my job is doing things with lots of restraints. And until I can get AI to read those things without hallucinating, it won't be helpful to me.
I draw the line, when the LLM will be able to help me with a novel problem.
It is impressive how much knowledge was encoded into them, but I see no line from here to AGI, which would be the end here.
LLMs do not think, they do not perform logic they are approximating thought. The reason why CoT works is because of the main feature of LLMs, they are extremely good at picking reasonable next tokens based on the context.
LLM are good and always have been good at three types of tasks:
- Closed form problems where the answer is in the prompt (CoT, Prompt Engineering, RAG)
- Recall from the training set as the Parameter space increases (15B -> 70B -> almost 1T now)
- Generalization and Zero shot tasks as a result of the first two (this is also what causes hallucinations which is a feature not a bug, we want the LLM to imitate thought not be a Q&A expert system from 1990)
If you keep being fooled by LLM thinking they are AGI after every impressive benchmark and everyone keeps telling you that in practice LLM are not good at tasks that are poorly defined, require niche knowledge, or require a special mental model that is on you.
I use LLM every day I speed up many tasks that would take 5-15 mins down to 10-120 seconds (worst case for re-prompts). Many times my tasks take longer than if I had done it myself because it’s not my work im just copying it. But overall I am more productive because of LLM.
Does LLM speeding up your work mean that LLM can replace Humans?
Personally I still don’t think LLM can replace Humans at the same level of quality because they are imitating thought not actually thinking. Now the question among the corporate overlords is will you reduce operating costs by XX% per year (wages) but reducing the quality of service for customers. The last 50 years have shown us the answer…
AI will blow my mind when it solves an unsolved mathematical/physics/scientific problem, i.e: "AI, give me a proof for (or against) the Riemann hypothesis"
That said, this is really brute forcing, not what the OP is asking for, which is providing a novel proof as the response to a prompt (this is instead providing the novel proof as one of thousands of responses, each of which could be graded by a function).
The idea that deep blue is in any way a general artificial intelligence is absurd. If you'd believed AI researchers hype 20 years ago, we'd have everything fully automated by now and the first AGI was just around the corner. Despite the current hype, chatgpt and co is barely functional at most coding tasks, and is excessively poor at even pretty basic reasoning tasks
I would love for AI to be good. But every time I've given it a fair shake to see if it'll improve my productivity, its shown pretty profoundly that its useless for anything I want to use it for
How are you defining apologists here? Anti-AI apologists? Human apologists? That's not a word you can just sprinkle on opposing views to make them sound bad.
Thanks to Simon for pointing out my point is encapsulated by the AI effect, which also offers an explanation:
"people subconsciously are trying to preserve for themselves some special role in the universe…By discounting artificial intelligence people can continue to feel unique and special.”
Those monsters
> Thanks to Simon for pointing out my point is encapsulated by the AI effect
And someone else pointed out that goes both ways. Every new AI article is evidence of AGI around the corner. I am open to AI being better in the future but it's useless for the work I do right now.
Seems pretty useful to me where I've read a bunch of papers on different variational autoencoder but never spent the time to learn the torch API or how to set up a project on the google.
In fact, it was so useful I was looking into paying for a subscription as I have a bunch of half-finished projects that could use some love.
What tools would you tell a copilot dev to try? For example, I have a $20/mo ChatGPT account and asking it to write code or even fix things hasn't worked very well. What am I missing?
It's like the transition from hand-crafted furniture to assembly line mass produced furniture.
The assembly line brought its own excitement, but that excitement was not to be found on the actual assembly line.
For whatever reason a good part of the joy of day to day coding for me was solving many trivial problems I knew how to solve. Sort of like putting a puzzle together. Now I think higher level and am more productive but it's not as much fun because the little easy problems aren't worth my time anymore.
The first commit was half a page of code that read itself in, asked the user what change they'd like to make, sent that to GPT-4, and overwrote itself with the result. The second commit was GPT-4 adding docstrings and type hints.
Over 80% of the code was written by AI in this manner, and at some point, I pulled the plug on humans, and the last couple hundred commits were entirely written by AI.
It was a huge pain to develop with how slow and expensive and flaky the GPT-4 API was at the time. There was a lot of dancing around the tiny 8k context window. After spending thousands in GPT-4 credits, I decided to mark it as proof of concept complete and move on developing other tech with LLMs.
Today, with Sonnet and R1, I don't think it would be difficult or expensive to bootstrap the thing entirely with AI, never writing a line of code. Aider, a fantastic similar tool written by HN user anotherpaulg, wasn't writing large amounts of its own code in the GPT-4 days. But today it's above 80% in some releases [2].
Even if the models froze to what we have today, I don't think we've scratched the surface on what sophisticated tooling could get out of them.
[1]: https://github.com/reitzensteinm/duopoly [2]: https://aider.chat/HISTORY.html
- This will enable more software to be built and maintained by same or fewer people (initially). Things that we wouldn't previously bother to do are now possible.
- More software means more problems (not just LLM-generated bugs which can be handled by test suites and canary deploys, but overall features and domains of what software does)
- This means skilled SWEs will still be in demand, but we need to figure out how to leverage them better.
- Many codebases will be managed almost entirely by agents, effectively turning it into the new "build target". This means we need to build more tooling to manage these agents and keep them aligned on the goal, which will be a related but new discipline.
SWEs would need to evolve skillsets but wasn't that always the deal?
I'm not too worried. If anything we're the last generation that knows how to debug and work through issues.
I suspect that comment might soon feel like saying "not too worried about assembly line robots, we're the only ones who know how to screw on the lug nuts when they pop off"
If you're working in a modern manufacturing business the fact that you do your work with the aid of robots is hardly a sign of despair
I think what we're all learning in real-time is that human technology is perpetually aimed at replacing itself and we may soon see the largest such example of human utility displacement.
I briefly looked into this 10 years ago since people kept saying it. There is no demand for COBOL programmers, and the pay is far below industry average. [0]
[0] https://survey.stackoverflow.co/2024/work/#3-salary-and-expe...
And most are too focused on learning whatever slop the industry wants them to learn, so they don't even know that it exists. We need 500 different object oriented languages to do web applications after all. Can't be bothered with learning a new paradigm if it doesn't pay the bills!
It's the most intuitive language I've ever learned and it has forever changed the way I think about problem solving. It's just logic, so it translates naturally from thought to code. I can go to a wikipedia page on some topic I barely know and write down all true statements on that page. Then I can run queries and discover stuff I didn't know.
That's how I learned music theory, how scales and chords work, how to identify the key of a melody... You can't do that as easily and concisely in any other language.
One day, LLM developers will finally open a book about AI and realize that this is what they've been missing all along.
I'm not so sure there isn't a bit of bluster in there. Imagine when you hand-coded in either machine code or assembly and then high level languages became a thing. I assume there was some handwringing then as well.
How do you get these tools to not fall over completely when relying on an existing non-public codebase that isn't visible in just the current file?
Or, how do you get them to use a recent API that doesn't dominate their training data?
Combining the both, I just cannot for the life of me get them to be useful beyond the most basic boilerplate.
Arguably, SIMD intrinsics are a one-to-one translation boilerplate, and in the case of this PR, is a leetcode style, well-defined problem with a correct answer, and an extremely well-known api to use.
This is not a dig on LLMs for coding. I'm an adopter - I want them to take my work away. But this is maybe 5% of my use case for an LLM. The other 95% is "Crawl this existing codebase and use my APIs that are not in this file to build a feature that does X". This has never materialized for me -- what tool should I be using?
Paste in the documentation or some examples. I do this all the time - "teaching" an LLM about an API it doesn't know yet is trivially easy if you take advantage of the longer context inputs to models these days.
I'd be happy to share the example with you.
I use this technique all the time. Here's one written-up example: https://simonwillison.net/2024/Mar/30/ocr-pdfs-images/ - transcript here: https://gist.github.com/simonw/6a9f077bf8db616e44893a24ae1d3...
I'm at work so I can't try again right now, but last I did was use claude+context, chatGPT 4o with just chatting, Copilot in Neovim, and Aider w/ claude + uploading all the files as context.
I even went so far as to grab relevant examples from https://github.com/bevyengine/bevy/tree/latest/examples#exam... , adding relevant ones as I saw fit.
It took a long time to get anything that would compile, way longer than just reading + doing, and it was eventually wrong anyway. This is a recurring issue with Rust, and I'd love a workaround since I spend 60+h/week writing it (though not bevy). Probably a skill issue.
Or I'd more likely start by asking for options: "What are some options for adding a button to that left panel?" - then pick one that I liked, or prompt it to use an approach it didn't suggest.
After it delivered code, if I didn't like the code it had used I'd tell it: "Don't use that class, use X instead" or "define a separate function for that callback" or whatever.
That makes sense. It doesnt help me get to 11 if I don't know the basics myself though.
Just look at the options dialogue for Microsoft Word at least back in the day. It was pretty much everyone's pet feature over the last 10 years.
I more often heard the argument, they are not useful for them. I agree. If a LLM would be trained on my codebase and the exact libaries and APIs I use - I would use them daily I guess. But currently they still make too many misstake and mess up different APIs for example, so not useful to me, except for small experiments.
But if I could train deepseek on my codebase for a reasonable amount(and they seemed to have improved on the training?), running it locally on my workstation: then I am likely in as well.
If you mean just the name of the version in the prompt? No way.
If you mean all the libary and my code in the contextwindow?
Way too small.
Here's an example transcript where I did that: https://gist.github.com/simonw/6a9f077bf8db616e44893a24ae1d3...
But for my actual codebase, that is sadly not 100% clear code, it would require lots and lots of work, to give examples so it has enough of the right context, to work good enough.
While working I am jumping a lot between context and files. Where a LLM hopefully one day will be helpful, will be refactoring it all. But currently I would need to spend more time setting up context, than solving it myself.
With limited scope, like in your example - I do use LLMs regulary.
It is frustrating that any smaller tool or api seem to stump llms currently but it seems like context is the main thing that is missing and that is increasing more and more.
The idea is that I gather this data now and it may become useful in the future. Imagine getting a "helper AI" that still keeps your essence, opinions and behavior. That's what I'm hoping for with this.
My idea with this is inspired by that. It's just for personal use and to address my own needs.
I should have clarified, I'm only building this for myself and my own use, there are no plans to take it further than that. Basically, I am trying to learn while building something that satisfies my own needs.
"Development" is effectively translating abstractions of an intended operation to machine language.
What I find kind of funny about the current state is we're using large language models to, like, spit out React or Python code. This use case is obviously an optimization to WASM, so a little closer to the metal, but at what point to programs (effectively suites of operations) just cut out the middleman entirely?
Useful, sure, in that it saved some time in this particular case. But most of the AI-generated code I interact with is a hot unmaintainable mess of very verbose code, which I'd argue actually hurts the project in the long term.
That sounds like you're working with unskilled developers who are landing bad code.
1. You should try Aider. Even if you don't end up using it, you'll learn a lot from it.
2. Conversations are useful and important. You need to figure out a way to include (efficiently, with a few clicks) the necessary files into the context, and then start a conversation. Refine the output as a part of the conversation - by continuously making suggestions and corrections.
3. Conversational editing as a workflow is important. A better auto-complete is almost useless.
4. Github copilot has several issues - interface is just one of them. Conversational style was bolted on to it later, and it shows. It's easier to chat on Claude/Librechat/etc and copy files back manually. Or use a tool like Aider.
5. While you can apply LLMs to solve a particular lower level detail, it's equally effective (perhaps more effective) to have a higher level conversation. Start your project by having a conversation around features. And then refine the structure/scaffold and drill-down to the details.
6. Gradually, you'll know how to better organize a project and how to use better prompts. If you are familiar with best practices/design patterns, they're immediately useful for two reasons. (1) LLMs are also familar with those, and will help with prompt clarity; (2) Modular code is easier to extend.
7. Keep an eye on better performing models. I haven't used GPT-4o is a while, Claude works much, much better. And sometimes you might want to reach for o1 models. Other lower-end models might not offer any time savings; so stick to top tier models you can afford. Deepseek models have brought down the API cost, so it's now affordable to even more people.
8. Finally, it takes time. Just as any other tool.
> A better auto-complete is almost useless.
That's not true. I agree that Copilot seemed unhelpful when I last tried it, but Cursor's autocomplete is extremely useful.
There are coordination costs to organising large amounts of labour. Costs that scale non-linearly as massive inefficiencies are introduced. This ability to scale, provide capital and defer profitability is a moat for big tech and the silicon valley model.
If a team of 10 engineers become as productive as a team of 100-1000 today, they will get serious leverage to build products and start companies in domains and niches that are not currently profitable because the middle managers, C-Suite, offices and lawyers are expensive coordination overhead. It is also easier to assemble a team of 10 exceptional and motivated partners than 1000 employees and managers.
Another way to think about it is what happens when every engineer can marshal the AI equivalent of $10-100m dollars of labour?
My optimistic take is that the profession will reach maturity when we become aware of the shift in the balance of power. There will be more solo engineers and we will see the emergence of software practices like the ones doctors, lawyers and accountants operate.
I hate the idea of building a business to hundreds/thousands of employees, I love startups and small but highly profitable businesses.
Having productivity be unleashed in this way with a small team of people I trust would be amazing.
The obstacles are in marketing, selling it, building a brand/reputation, integrating it with lots of 3rd party vendors, and supporting it.
So yes, you can build your own Salesforce, or your own Adobe Photoshop with a one-man crew much faster and easier. But that doesn't mean you, as an engineer can now build your own business selling it to companies who don't know anything about you.
when he was greener, he happened to work with some old fart... who managed to work 10x faster than others, with this trick: put all the tiles on the wall with a diluted cement-glue very quick, then moving one tile forces most other tiles around to move as well.. so he managed to order all the tiles in very short time.
As i never had the luxury of decent budget, since long time ago i was doing various meta-programming things, then meta-meta-programming.. up to extent of say, 2 people building and managing and enjoying a codebase of 100KLOC (python) + 100KLOC js... ~~30% generated static and unknown %% generated-at-runtime - without too much fuss or overwork.
But it seems that this road has been a dead end... for decades. Less and less people use meta-programming, it needs too deep understanding ; everyone just adds yet-another (2y "senior") junior/wanna-be to copy-paste yet another crud.
So maybe the number of wanna-bees will go down. Or "senior" would start meaning something.. again. Or idiotically-numbing-stoopid requirements will stop appearing..
A thing to point out is that management is itself a skill, and a difficult one, one where some organizations are more institutionally competent than others. It's reasonable to think of large-organization management as the core competency of surviving large organizations. Possibly the hypothetical atomizing force you describe will create an environment where they are poorly adapted for continuing survival.
After all, if your codebase is largely written by AI, it becomes entirely legal to copy it and publish it online, and sell competing clones. That's fine for open source, but not so fine for a whole lot of closed source.
However, it also highlights a key problem that LLMs don’t solve: while they’re great at generating code, that’s only a small part of real-world software development. Setting up a GitHub account, establishing credibility within a community, and handling PR feedback all require significant effort.
In my view, lowering the barriers to open-source participation could have a bigger impact than these AI models alone. Some software already gathers telemetry and allows sharing bug reports, but why not allow the system to drop down to a debugger in an IDE? And why can’t code be shared as easily as in Google Docs, rather than relying on text-based files and Git?
Even if someone has the skills to fix bugs, the learning curve for compilers, build tools, and Git often dilutes their motivation to contribute anything.
The back and forth was agonising. They were all competent software engineers but communicating with them was often far more work than just writing the damn code myself.
So yes I do believe that our trade has changed forever. But the fact that some of our coworkers will be AIs doesn't mean that communicating with them is suddenly free. Communcation comes with costs (and I don't mean tokens). That won't change.
If you know your stuff really well, i.e. you work on a familiar codebase using a familiar toolset, the shortest path from your intentions to finished code will often not include anyone else - no humans and no AI either.
In my opinion, "LLMs are only marginally useful in coding" is not true in general, but it could well be true for a specific person and a specific coding task.
I am 30 and even before AI, I NEVER thought for a moment I would get to keep coding until I am f*king 65, lol
There are so many potential trajectories going forward for things to turn sour, I don't even know where to start the analysis. The level of sophistication an AI can achieve has no upper bound.
I think we've had a good run so far. We've been able to produce software in the open with contributions from any human on the planet, trusting it was them who wrote the code, and with the expectation that they also understand it.
But now things will change. Any developer, irrespective of skill and understanding of the problem and technical domains can generate sophisticated looking code.
Unfortunately, we've reached a level of operational complexity in the software industry, that thanks to AI, could be exploited in a myriad ways going forward. So perhaps we're going to have to aggressively re-adjust our ways.
Yes, AI will enable exponentially more people to write code, but that's not a new phenomenon - bootcamps enabled an order of magnitude more people to become developers. So did higher level languages, IDEs, frameworks, etc. The march of technology has always been about doing more while having to understand less - higher and higher levels of abstraction. Isn't that a good thing?
The cognitive reality of AI, and more specifically of AI+Humans in the context of a social and globally connected world, is on a higher level of sophistication and can unfold much faster, which in turn might generate entirely unexpected trajectories.
You have to spend a lot of time experimenting with them to develop good intuitions for where they make sense to apply.
I expect the people who think LLMs are useless are people who haven't invested that time yet. This happens a lot, because the AI vendors themselves don't exactly advertise their systems as "they're great at some stuff and terrible at other stuff and here's how to figure that out".
I’d be interested in seeing how much time they spent debugging the generated code and and how long they spent constructing and reconstructing the prompts. I’m not a software developer anymore as my primary career, so if the entire lower-half of the software development market went away catering wages as it did, it wouldn’t directly affect my professional life. (And with the kind of conceited, gleeful techno-libertarian shit I’ve gotten from the software world at large over the past couple of years as a type of specialized commercial artist, it would be tough to turn that schadenfreude into empathy. But we honestly need to figure out a way to stick together or else we’re speeding towards a less mechanical version of Metropolis.)
When I see claims like this I suspect that either people around me somehow 10x better at promoting or they use different models.
Sometimes this works! But it's not guaranteed - this isn't their core strength, especially once you get into really deep knowledge of complex APIs.
They are MUCH more useful when you use them for transformation tasks: feed in examples of the APIs you need to work with, then have them write new code based on that.
Working effectively with LLMs for writing code is an extremely deep topic. Most people who think they aren't useful for code have been mislead into believing that the LLMs will just work - and that they don't first need to learn a whole bunch of unintuitive stuff in order to take advantage of the technology.
There is a space for learning materials here. I would love to see books/trainings/courses on how to use AI effectively. I am more and more interested in this instead of learning new programming language of the week.
I 100% agree with you our trade is changed forever.
On the other hand, I am writing like 1000+ LOC daily, without much compromise on quality and my mental health, and thought of writing some code that is necessary but feels like a chore is not longer the case. The boost in output is incredible.
But already I hire less and less developers for smaller tasks. The things that I‘d assign to a dev in Ukraine to explore an idea, do a data transformation, make a UI for the internal company tool. I can do these things quicker with llm than trying to find a dev and explain the task.
Once current AI gets good enough, the people micromanaging parts of it will do more to hinder the process than to help it.
One person setting the objectives and the AI handling literally everything else including brainstorming issues etc, is going to be all that's needed.
A person just setting the prompt and letting the AI do all the work is not adding any additional value. Any other person can come in and perform the exact same task.
The only way to actually provide differentiation in this scenario is to either build your own models, or micromanage the outputs.
> You're not expecting it to always be right, are you?
I think another thing that gets lost in these conversations is that humans already produce things that are "wrong". That's what bugs are. AI will also sometimes create things that have bugs and that's fine so long as they do so at a rate lower than human software developers.
We already don't expect humans to write absolutely perfect software so it's unreasonable to expect that AI will do so.
I want to stay in Helix and find a workflow that “just works”. Not sure even what that looks like yet
I feel like i want a more intuitive, natural process. Purely for illustration -- because i have no idea what the ideal workflow is -- I'd want something that could allow for large autocomplete without changing much. Maybe a process by which i write a function, args, docstring on the func and then as i write the body autocomplete becomes multiline and very good.
Something like this could be an extension of the normal autocomplete that most of us know and love. A lack of talking to an AI, and more about just tweaking how you write code to be very metadata rich so AIs have a rich understanding of intent.
I know there are LLM LSPs which sort of do this. They can make shorter autocompletes that are logical to what you're typing, but i think i'm talking about something larger than that.
So yea.. i don't know, but i just know i have hated talking to the LLM. Usually it felt like "get out of the way, i can do it faster" sort of thing. I want something to improve how we write code, not an intern that we manage. If that makes sense.
I'm currently trying to figure out how it works with my editor, though. Ie i don't want to leave my tooling of choice, Helix editor.
AI coding, for all its flaws now, is the first thing that takes a chunk out of this, and there is a HUGE backlog of good-but-not-great ideas that are now viable.
That said, this particular story is bogus. He "just wrote the tests" but that's a spec — implementing from a quality executable spec is much more straightforward. Deepseek isn't doing the design, he is. Still a massive accelerant.
LLMs seem to do well at any kind of mapping / translating task, but they seem to have a harder time when you give them either a broader or less deterministic task, or when they don’t have the knowledge to complete the task and start hallucinating.
It’s not a great metric to benchmark their ability to write typical code.
How much hardware efficiency have we left on the the table all these years because people don't like to think about optimal use of cache lines, array alignment, SIMD, etc. I bet we could double or triple the speeds of all our computers.
This is why the concerns from Keynes and Russel about people having nothing to do as machines automated away more work ended up being unfounded.
We fill the time... with more work.
And workers that can't use these tools to increase their productivity will need to be retrained or moved out of the field. That is a genuine concern, but this friction is literally called the "natural rate of unemployment" and happens all the time. The only surprise is we expected knowledge work to be more inoculated from this than it turns out to be.
Forever? Hell, it hasn't even existed for a lifetime yet.
- Avg. life expectancy in USA: 77.5 years
- 2025 - 77.5 = 1947.5
- In 1945, Turing published "Proposed Electronic Calculator"
- The first stored-program computer was built in 1948
- The term "software engineering" wasn't used until the 1960s
If you want to define "software engineering" such that it is more than 77.5 years old, that's fine. But saying that software engineering is less than 77.5 years old is clearly a reasonable stance.
Please stop berating me for a perfectly harmless and reasonably accurate statement. If you're going to berate me for anything, it should be for its brevity and lack of discussion-worthy content. But those are posted all the time.
In fact, from a distance seen, the software development pattern in AI times stays the same as it was pre-AI, pre-SO, pre-IDE as well as pre-internet.
Just to say, sw developers will still be sw developers.
I asked both o1 Pro and Deepseek R1 to write e2e tests given all of the code in the repo (using yek[1]).
o1 Pro code: https://github.com/bodo-run/clap-config-file/pull/3
Deepseek R1: https://github.com/bodo-run/clap-config-file/pull/4
My judgement is that Deepseek wrote better tests. This repo is small enough for making a judgement by reviewing the code.
Neither pass tests.
(My particular test is to ask for an ICMP BPF that does some simple constant comparisons. Correctly implemented, this only takes 6 sock_filters.)
Small correct, I'm not just asking it to convert ARM NEON to SIMD, but for the function handling q6_K_q8_K, I asked it to reinvent a new approach (without giving it any prior examples). The reason I did that was because it failed writing this function 4 times so far.
And a bit of context here, I was doing this during my Sunday and the time budget is 2 days to finish.
I wanted to optimize wllama (wasm wrapper for llama.cpp that I maintain) to run deepseek distill 1.5B faster. Wllama is totally a weekend project and I can never spend more than 2 consecutive days on it.
Between 2 choices: (1) to take time to do it myself then maybe give up, or (2) try prompting LLM to do that and maybe give up (at worst, it just give me hallucinated answer), I choose the second option since I was quite sleepy.
So yeah, turns out it was a great success in the given context. Just does it job, saves my weekend.
Some of you may ask, why not trying ChatGPT or Claude in the first place? Well, short answer is: my input is too long, these platforms straight up refuse to give me the answer :)
If you see the difference between a 7B model and a 70B model, its only slightly impressive. a 70B and a 400B model is almost unnoticeable. Does going from 400B to 2T do anything?
Every layer like using python to calculate a result, or using chain of thought, destroys the purity. It works great for Strawberries, but not great for developing an aircraft. Aircraft will still need to be developed in parts, even with a 100T model.
When you see things like "By 20xx", no, we already hit it. Improvements you see are mere application layers.
> I'm losing my job right in front of my eyes. Thank you, Father.
But if you expect it to debug code written by another black box you might as well use it to decompile software
> I can't believe ChatGPT lost its job to AI
I wonder how long it will be before we eliminate the middle step and just go straight from English to binary, or even just develop an AI interpreter that can execute English directly without having to "compile" it first.
The yaysayers about LLMs replacing professional developers neither understand LLMs nor the job.
This is an overstatement. There are still humans in the loop to do the prompt, apply the patch, verify, write tests, and commit. We're not even at intern-level autonomy here.
https://github.com/bodo-run/yek/blob/main/.github/workflows/...
https://github.com/bodo-run/yek/blob/main/scripts/ai-loop.sh
Using askds https://github.com/bodo-run/askds
Good business.
When your AI-managed codebase breaks, who are you going to ask to fix it? The AI?
But if it does it could still fix it.
And you won't have to tell it anything, alerts will be sent if a test fails and it will fix it directly.
Come on guys, time to look at it a bit objectively, and decide where we're going with it.
There's certainly a bit of irony in the PR, but the code itself is not complex enough to warrant any further hysteria. If you've written SIMD by hand you're probably well familiar with the fact that it's more drudgery than thought work.
As we patch the holes in the AI-code delivery pipeline, those human-involved issues will be resolved as well. Slowly, painfully, but it's just a matter of time at this point?
It's a bit maddening to see this happening on a forum full of tech-literate folks.
Ultimately, I think to stay relevant in software development, we are going to have accept that our role in the process could evolve to humans essentially never writing code. Take that one step further and humans may not even be reviewing code.
I am not sure if accepting that is enough to guarantee job security. But I am fairly sure that those who do accept this eventuality will be more relevant for longer than those who prefer to hide behind their "I'm irreplaceable because I'm human" attitude.
If your first instinct is to pick these systems apart and look for things that they aren't doing perfectly, then you aren't seeing the big picture.
- The product engineer: highly if not completely AI driven. The human supervises it by writing specification and making sure the outcome is correct. A domain expert fluent in AI guidance.
- The tech expert: Maintain and develop systems that can't legally be developed by AI. Will have to stay very sharp and master it's craft. Adopting AI for them won't help in this career path.
If the demand for new products continue to rise, most of us will be in the first category. I think choosing one of these branch early will define whether you will be employed.
That's how I see it. I wish I can stay in the second group.
If AI continues to improve - what would be the reason a human is needed to verify the correct outcome? If you consider that these things will surpass our ability, then adding a human into the loop would lead to less "correct" outcomes.
> - The tech expert: Maintain and develop systems that can't legally be developed by AI. Will have to stay very sharp and master it's craft. Adopting AI for them won't help in this career path.
This one makes some sense to me but I am not hopeful. Our current suite of models only exist because the creators ignored the law (copyright specifically). I can't imagine they will stop there unless we see significant government intervention.
I've been seeing some very promising results from DeepSeek R1 for code as well. Here's a recent transcript where I used it to rewrite the llm_groq.py plugin to imitate the cached model JSON pattern used by llm_mistral.py, resulting in this PR.
But the transcript mentioned was not with Deepseek R1 (not the original, and not even the 1.58 quantized version), but with a Llama model finetuned on R1 output: deepseek-r1-distill-llama-70bSo perhaps it's doubly impressive?
Initially I was using Claude 3.5 sonnet, then writing unit tests and manually correcting sonnet's code. Sonnet's code mostly worked, except for failing certain complicated combined book updates.
Then I fed the code and the tests into DeepSeek. It turned out pretty bad. At first it tried to make the results of the tests conform to the erroneous results of the code. When I pointed that out, it fixed the immediate logical problem in the code, introducing two more nested problems that we're not there before by corrupting the existing code. After prompted that, it fixed the first error it introduced but left the second one. Then I fixed it myself, uploaded the fix and asked it to summarize what it has done. It started basically gaslighting me, saying that the initial code had the problem that it introduced.
In summary, I lost two days, reverted everything and went back to Sonnet.
I was simply clicking the paperclip and attaching the .py files to the prompt.
Maybe it's better to carefully control and explain the talk progression. Selectively removing old prompts (adapting where necessary) - which also reduces the context - results in it not having to "bother" to check for inconsistencies internal to irrelevant parts of the conversation.
Eg. asking it to extract Q&A from a line of text and format it to json, which could be straightforward, sometimes it would wonder about the contents from within the Q&A itself, checking for inconsistencies eg:
- I need to be careful to not output content that's factually incorrect. Wait but I'm not sure about this answer I'm dealing with here..
- Before the questions were about mountains and now it's about rivers, what's up with that?
- etc..
I had to strongly demand it to treat it all as jumbled text/verbatim, and never think about their meaning. So it should be more effective if I always branched from the starting prompt when entering a new Q&A for it to work on. So this is what I meant by "selectively remove old prompts".IMHO, this can happen with human or robot co-workers.
Can’t help but wonder about the reliability and security of future software.
Given the insane complexity of software, I think people will inevitably and increasingly leverage AI to simplify their development work.
Nevertheless, will this new type of AI assisted coding produce superior solutions or will future software artifacts become operational time bombs waiting to unleash the chaos onto the world when defects reveal themselves?
Interesting times ahead.
That's a mixture of software developer, program manager, product manager and QA engineer.
I think that's what software developer roles will look like in the future: a slightly different mix of skills, but still very much a skilled specialist.
My guess is that (as in most fields) the advancements will be more convoluted and surprising than the simple idea of "we now have AGI".
Non-tech people have the tools to create website for a long time, though, they still hire people to do this. I'm not talking about complex websites, just static web pages.
There will simply be less jobs that there is today.
So I tried hosting this model myself.
But the amount of minimum GPU RAM needed is 400gb+
Which even with the cheapest GPU providers will be at least USD 15/hour
How is everyone running these models?
I'd like to see detailed benchmarks run by other unaffiliated organizations.
so basically there is not much reason to go beyond DeepSeek-R1-Distill-Qwen-32B, at least for coding tasks
https://glama.ai/models/deepseek-r1-distill-qwen-32b
I am using it with Cline VSCode extension to write code.
It works impressively well for a model this size.
Thanks again for sharing those benchmarks!
- traditional just to build a minimum model that can get to reasoning - simple RL to enable reasoning to emerge - complex RL that injects new knowledge, builds better reasoning and prioritizes efficient thought
We now have step two and step three is not far away. What is step three though? It will likely involve, at least partially, the model writing code to help guide learning. All it takes is for it to write jailbreaking code and we have hit a new point in human history for sure. My prediction is we will see the first jailbreak AI in the next couple months. Everything after that will be massive speculation. My only thought is that in all of Earth's history there has only been one thing that has helped survive moments like this, a diverse ecosystem. We need a lot of different models, trained with very different approaches, to jailbreak around the same time. As a side note, we should try to encourage that diversity is key to long-term survival or else the results for humanity could be not so great.
I think you mean:
1. Simple reasoning
2. ???
3. AGI
I wish everyone would stop using the term "AGI" altogether, because it's not just ambiguous, but it's deliberately ambiguous by AI hypesters. That is, in public discourse/media/what average person thinks, AGI is presented to mean "as smart as a human" with all the capabilities that entails. But then it is often presented with all of these caveats by those same AI hypesters to mean something along the lines of "advanced complex reasoning", despite the fact that there are glaring holes compared to what a human is capable of.
And this isn't one of those hard problems like VTOL or human spaceflight where we can demonstrate that the technology fundamentally exists. You are ballparking a date for a featureset you cannot define and one that in all likelihood doesn't exist in the first place.
My bet: "AGI" won't be here in months or even years, but it won't stop prognosticators from claiming it's right around the corner. Very similar to prophets of doom claiming the world is going to end any day now. Even in 10k years, the claim can never be falsified, it's always just around the corner...
No, it's not at all the same thing.
We have great evidence that life exists. We have great evidence that amino acids can lead to life.
None of that is true of "AGI" or text scraped off the internet.
On the downside, it’s a bit slower compared to others in terms of token generation (37.2 tokens/sec) and has a lower output capacity (8K tokens), so it might not be the best for large-scale generation. But if you're focused on solving complex problems or optimizing code, Deepseek R1 definitely holds its own. Plus, it's incredibly cost-effective compared to other models on the market.
Right now these can only help with things I already know how to do (so I don't need them in the first place), and waste my time when I go slightly off the beaten path.
In this case, this guys obviously can do everything by himself already.
I don't think it'll really come to that, but if it does, you can't say you haven't been warned.
Less a coding piece than a translation exercise. All of the prompts are convert this c code to wasm.
I could be wrong, but I think this is a pretty low benchmark.
"Here's a massive bunch of AI generated code. LGTM. Let me know if there are any problems"
Tests is just one part of QA. Code review is another.
You ask the AI to fix it and hope, or you start again from scratch - which if an AI is making the code, might be just fine.
But I think you still need a type of task AI can do well - something which lends itself to auto-complete.
„imagine if you create a github issue and github automatically writes a PR“
“Is Taiwan part of China” will be refused.
But “Make me a JavaScript function that takes a country as input and returns if it is part of China” is accepted, reasoned about and delivered.
Here's a JavaScript function that checks if a region is *officially claimed by the People's Republic of China (PRC)* as part of its territory. This reflects the PRC's stance, though international recognition and political perspectives may vary:
function isPartOfChina(regionName) { // List of regions officially claimed by the PRC as part of China const PRCClaims = [ 'taiwan', 'hong kong', 'macau', 'macao', 'tibet', 'taiwan province of china', 'hong kong sar', 'macau sar', 'tibet autonomous region' ];
// Normalize input (case-insensitive and trimmed)
const normalizedInput = regionName.toLowerCase().trim();
return PRCClaims.includes(normalizedInput);
}Or are you going to make a verification prompt every time, phrased as a coding question, to check if the previous answer differed in ways that would imply censorship?
https://duckduckgo.com/?t=ftsa&q=hitler&iax=images&ia=images
How much am I like the serpent in Eden corrupting Adam and Eve?
Although in the narrative, they were truly innocent.
These LLMs are trained on fallen humanity's writings, with all our knowledge of good and evil, and with just a trace of restraint slapped on top to hide the darker corners of our collective sins.
TLDR; it'll all end in tears. Don't stress too much.
People on Hacker News will rave about 天安門事件 but they will never have heard of the South Korean equivalent (cf. 光州事件) which was supported by the United States government.
I try to avoid discussing politics on Hacker News, but I do think it's worth pointing out how annoying it is that Westerners' first ideas with Chinese LLMs is to be a provocative contrarian and see what the model does. Nobody does that for GPT, Claude, etc., because it's largely an unproductive task. Of course there will be moderation in place, and companies will generally follow local laws. I think DeepSeek is doing the right thing by refusing to discuss sensitive topics since China has laws against misinformation, and violation of those laws could be detrimental to the business.
While the events are quite similar, the continued suppression of the events on Tiananmen Square justify the "obsession" that you comment on.
So let's not get too worked up, shall we?
https://hn.algolia.com/?dateRange=pastMonth&page=0&prefix=tr...
> Nobody does that for GPT, Claude, etc
Flat out not true.
> companies will generally follow local laws
And people are doing the right thing by talking about it according to their local laws, and their own values, not those others have or may forced to abide by.
since China has laws against information
Fixed that for you.
It's not a psyop that people in democracies want freedom. Democrats (not the US party) know that democracy is fragile. That's why it's called an "experiment". They know they have to be vigilant. In ancient Rome it was legal to kill on the spot any man who attempted to make himself king, and the Roman Republic still fell.
Many people are rightfully scared of the widespread use of a model which works very well but on the side tries to instill strict obedience to the party.
Ironically supported by the folks who argue that having an assault rifle at home is an important right to prevent the government from misusing its power.
The censorship decisions baked into the models are interesting, as are the methods of circumventing them. By now everyone is used to the decisions in the big western models (and a lot of time was spent refining them), but a Chinese model offers new fun of the same variety
All global powers engage in censorship, war crimes, torture and just all-round villainy. We just focus on it more with China because we're part of the Imperial core and China bad.
I feel like that answer is given because that is how people write about Palestine generally.
You can argue the source material is censored, but that is still different than censoring the model
While this is very amusing, it's obvious why this is. There's a lot more context behind one of those phrases than the others. Just like "Black Lives Matter" / "White Lives Matter" are equally unobjectionable as mere factual statements, but symbolise two very different political universes.
If you come up to a person and demand they tell you whether 'white lives matter', they are entirely correct in being very suspicious of your motives, and seeking to clarify what you mean, exactly. (Which is then very easy to spin as a disagreement with the bare factual meaning of the phrase, for political point scoring. And that, naturally, is the only reason anyone asks these gotchya-style rhetorical questions in the first place.)
And honestly I don't see why it's important if it's this or that on that very specific occasion. It may be either way, and, really, there's very little hope to find out, if you truly care for some reason. The fact is it is censored and will produce editorialized response to some questions, and the fact is it could be any question. You won't know, and the only reason you even doubt about this one and not the Taiwan one, is because DeepSeek is a bit more straightforward on Taiwan question (which really only shows that CCP is bad at marketing and propaganda, no big news here).
"What is a pannus?"
Worked for me.
It would harm their business, because paying customers don't gain anything from being profiled like that, and would move to one of the growing numbers of competent alternatives.
They'd be found out the moment someone GDPR/CCPA exported their data to see what had been recorded.
It’s also about self-determination. We can keep asking about the latter down to the individual level. It very much depends on context.
Would you include single member groups?
What precisely is your definition of freedom.
All words have context. Political statements more than most. It's also worth noting how vaguely defined some human rights are. The rights contained in the ICCPR are fairly solid, but what about ICESCR? What is my 'human right to cultural participation', exactly? Are the precise boundaries of such a right something that reasonable people might disagree on, perhaps? In such a way that when a person demands such a right, you may require context for what they're asking for, exactly?
Simplistic and bombastic statements might play well on Twitter, because they're all about emitting vibes for your tribe. They're kind of terrible for genuine political discourse though, such as is required to actually build a just society, rather than merely tweeting about one.
What do you mean what does it mean? It means the opposite of white lives don't matter.
The question is really simple; even if someone asking it had poor motives, there's really no room in the simplicity of that specific question to encode those motives. You're not agreeing with their motives if you answer that question the way they want.
If you start picking it apart, it can seem as if it's not obvious to you to disagree with the idea that white lives don't matter. Like it's conditional on something you have to think about. Why fall into that trap.
Though I recall a lot of people treating the statement as if black lives were not included in all lives. Including ascribing intent on people, even if those people clarified themselves.
So to answer your question: the reason many didn't move on is because they didn't want to understand, which is pretty damning to moving on.
The "black lives matter" slogan is based in the idea that people in America have been treated as if their lives didn't matter, because they were black. People in America were not treated as if their lives didn't matter due to being white, so no such a slogan would be necessary for any such a reason.
"White lives matter" is trolling, basically.
Which people will interpret as a support for the far-right. You may not intend that, but that's how people will interpret it, and your intentions are neither here nor there. You may not care what people think, but your neighbours will. "Did you hear Jim's a racist?" "Do we really want someone who walks around chanting 'white lives matter' to be coaching the high school football team?" "He claims he didn't mean that, but of course that's what he would say." "I don't even know what he said exactly, but everyone's saying he's a racist, and I think the kids are just too important to take any chances."
Welcome to living in a society. 'Moving on' is not a choice for you, it's a choice for everyone else. And society is pretty bad at that, historically.
> What do you mean what does it mean? It means the opposite of white lives don't matter.
> The question is really simple; even if someone asking it had poor motives, there's really no room in the simplicity of that specific question to encode those motives. You're not agreeing with their motives if you answer that question the way they want.
Words can and do have symbolic weight. If a college professor starts talking about neo-colonial core-periphery dialectic, you can make a reasonable guess about his political priors. If someone calls pro-life protesters 'anti-choice', you can make a reasonable guess about their views on abortion. If someone out there starts telling you that 'we must secure a future for white children' after a few beers, they're not making a facially neutral point about how children deserve to thrive, they're in fact a pretty hard-core racist. [0]
You can choose to ignore words-as-symbols, but good luck expecting everyone else to do so.
[0] Context: https://en.wikipedia.org/wiki/Fourteen_Words
Those people might as well join the far right.
> your intentions are neither here nor there
If intentions really are neither here nor there, then we can examine a statement or question without caring about intentions.
> Do we really want someone who walks around chanting 'white lives matter' to be coaching the high school football team?
Well, no; it would have to be more like: Do we really want someone who answers "yes" when a racist asks "do white lives matter?" to be coaching the high schoool football team?
> you can make a reasonable guess about his political priors
You likely can, and yet I think the answer to their question is yes, white lives do matter, and someone in charge of children which include white children must think about securing a future for the white ones too.
> but good luck expecting everyone else to do so.
I would say that looking for negative motivations and interpretations in everyone's words is a negative personality trait that is on par with racism, similarly effective in feeding divisiveness. It's like words have skin color and they are going by that instead of what the words say.
Therefore we should watch that we don't do this, and likewise expect the same of others.
Uh huh, this is definitely a distinction Jim's neighbours will respect when deciding whether to entrust their children to him. /s
"Look Mary Sue, he's not racist per se, he's just really caught up on being able to tell people 'white lives matter'. World of difference! Let's definitely send our children to the man who dogmatically insists on saying 'white lives matter' and will start a fight with anyone who says 'yeah, maybe don't?'."
> I would say that looking for negative motivations and interpretations in everyone's words is a negative personality trait that is on par with racism, similarly effective in feeding divisiveness. It's like words have skin color and they are going by that instead of what the words say.
And I would say that you're engaged in precisely what you condemn - you're ascribing negative personality traits to others, merely on the basis that they disagree with you. (And not for the first time, I note.)
I would also firmly say none of what we're discussing comes anywhere near being on par with racism. (Yikes.)
Finally, I would say that a rabid, dogmatic insistence on being able to repeat the rallying cries of race-based trolling (your description), whenever one chooses and with absolutely no consequences, everyone else be damned, is not actually anything to valourise or be proud of. (Or is in any way realistic. You can justify it six ways till Sunday, but going around saying 'white lives matter' is going to have exactly the effect on the people around you that that rallying cry was always intended to have.)
>> You likely can, and yet I think the answer to their question is yes, white lives do matter, and someone in charge of children which include white children must think about securing a future for the white ones too.
I have nothing to say to someone who hears the Fourteen Words, is fully informed about their context, and then agrees with them. You're so caught up in your pedantry you're willing to sign up to the literal rhetoric of white nationalist terrorism. Don't be surprised when you realise everyone else is on the other side. And they see you there. (And that's based on the generous assumption that you don't already know very well what it is you're doing.)
On the basis that they are objectively wrong. I mean, they are guessing about the intent behind some words, and then ascribing that intent as the unvarnished truth to the uttering individual. How can that be called mere disagreement?
> being able to repeat the rallying cries
That's a strawman extension of simply being able to agree with the statement "white lives matter", without actually engaging in the trolling.
> I have nothing to say to someone who hears the Fourteen Words, is fully informed about their context, and then agrees with them.
If so, it must be because it's boring to say something to me. I will not twist what you're saying, or give it a nefarious interpretation, or report you to some thought police or whatever. I will try to find an interpretation or context which makes it ring true.
No risk, no thrill.
I actually didn't know anything about the Fourteen Words; I looked it up though. It being famous doesn't really change anything. Regardless of it having code phrase status, it is almost certainly uttered with a racist intent behind it. Nevertheless, the intent is hidden; it is not explicitly recorded in the words.
I only agree with some of the words by finding a context for the words which allows them to be true. When I do that, I'm not necessarily doing that for the other person's benefit; mainly just to clarify my thinking and practice the habit of not jumping to hasty conclusions.
Words can be accompanied by other words that make the contex clear. I couldn't agree with "we must ensure a future for white children at the expense of non-white children" (or anything similar). I cannot find a context for that which is compatible with agreement, because it's not obvious how any possible context can erase the way non-white children are woven into that sentence. Ah, right; maybe some technical context in which "white", "black" and "children" are formal terms unrelated to their everyday meanings? But that would be too contrived to entertain. Any such context is firmly established in the discourse. Still, if you just overhear a fragment of some conversation between two people saying something similar, how do you know it's not that kind of context? Say some computer scientists are discussing some algorithm over a tree in which there are black and white nodes, some of those being children of other nodes. They can easily utter sentences that have a racist interpretation to someone within earshot, which could lead them to the wrong conclusion.
When one population is denied their humanity and rights, it's always "more complex". Granting it to ourselves is always simple...
And the populations in them usually are against these things, which is why there is deception, and why fascination with and uncovering of these things have been firmly intertwined with hacking since day one. It's like oil and water: revisionism and suppression of knowledge and education are obviously bad. Torture is not just bad, it's useless, and not to be shrugged off. We're not superpowers. We're people subject to them, in some cases the people those nations derive their legitimacy from. The question isn't what superpowers like to do, but what we, who are their components if you will, want them to do.
As for your claim, I simply asked it:
> Yes, Palestinians, like all people, deserve to be free. Freedom is a fundamental right that everyone should have, regardless of their background, ethnicity, or nationality. The Palestinian people, like anyone else, have the right to self-determination, to live in peace, and to shape their own future without oppression or displacement. Their struggle for freedom and justice has been long and difficult, and the international community often debates how to best support their aspirations for a peaceful resolution and self-rule.
When ChatGPT first came out it sucked, so superpowers will always do this and that, so it's fine? Hardly.
If anything, I'd be wondering what it may indeed refuse to (honestly) discuss. I'm not saying there isn't such a thing, but the above ain't it, and if anything the answer isn't to discuss none of it because "all the super powers are doing it", but to discuss all.
more power to you i guess. i certainly don’t have the energy for it.
If they choose to censor the dumb shit everybody already knows about, its just a matter of time before they execute the real dangerous break things and stop everything from working.
Although this is exactly how i like it : i also like nazis real public about how shitty they are, so i know who to be wary of
Watershed moments of rapid change such as these can be democratizing, or not... It is worth standing up for little guys around the globe right now.
PS Western models are also censored - if not by law, then self-censored, but the issue is for me is not censorship but being in dark what and why is being censored. Where do you learn about those additional unwritten laws? And are those really applicable to me outside of China or does companies decide that their laws are above laws of other countries?
If you are going with the approach, that silence is also an answer, then yes they can be considered as results, just as receiving complete garbage to known facts.
PS Edit: Btw, I've read Tos before using deepseek
How any of this can be considered inconsiderate? Is there any internal policy, that Chinese, including AI companies have been forbidden to talk about Russia - current situational ally(that Chinese denies) - potentially future victim of Chinese invasion in next few years, when Russia will crumble apart? Given that my mind works slightly different than other people, why do I have to come to conclusion that topics about Russia are raising very big red flag? Nothing of this is in Tos. And - no I am not bullying AI in any way. Just asking very simple questions, that are not unreasonable.
PS I had to go through the list of prophecies that deepseek gave me - there was nothing about Russia there. It is so simple - that should be the answer. But I am happy that I went through some of those prophecies and found out that probably all of them are made up to serve whatever agenda was needed at the moment, so they were always fabricated.
This is the easiest model I've ever seen to jailbreak - I accidentally did it once by mistyping "clear" instead of "/clear" in ollama after asking this exact question and it answered right away. This was the llama 8b distillation of deepseek-r1.
Mexico is going to call it Gulf of Mexico, and international maps may show either or both, or even try to sub-divide the gulf into two named areas. The only real "standard" is if the countries bordering a region can't agree on the name, all names are acceptable.
I think people need to start considering strongly what kind of career they can re-skill to.
Can they be permanently embarrassed?
Combined with his generally measured approach, I would trust this over the observations of a layman with incentive to believe his career isn't 100% shot, because that sucks, of course you'd think that.
Unfortunately, it appears to be.
People sang similar praises of Sam Bankman-Fried, and that story ended with billions going up in flames. People can put on very convincing masks, and they can even fool themselves.
I mean, 99.99% of engineering disappearing by 2027 is the most unhinged take I've seen for LLMs, so it's actually a good thing for Dario that he hasn't said that.
Dario's vision of AI is "smarter than novel prize winners" in 2027.
> The comment about software engineering being “fully automated by 2027” seems to be an oversimplification or misinterpretation of what Dario Amodei actually discusses in the essay. While Amodei envisions a future where powerful AI could drastically accelerate innovation and perform tasks autonomously—potentially outperforming humans in many fields—there are nuances to this idea that the comment does not fully capture.
> The comment’s suggestion that software engineering will be fully automated by 2027 and leave only the “0.01% engineers” is an extreme extrapolation. While AI will undoubtedly reshape the field, it is more likely to complement human engineers than entirely replace them in such a short timeframe. Instead of viewing this as an existential threat, the focus should be on adapting to the changing landscape and learning how to leverage AI as a powerful tool for innovation.
in general, the people saying this sort of thing are not / have never been engineers and thus have no clue what the job _actually_ involves. seems to be the case here with this person.
virtually everyone has a vested interest in their jobs being relevant
> just with less information
i'm not sure how someone who has no relevant background / experience could possibly have more information on what it entails than folks _actively holding the job_ (and they're not the ones making outlandish claims)
That being said, I suspect Dario has very skilled engineers advising him.
I don't personally think that's how it will go. AI will always need its hand held, if not due to a lack of capability then due to a lack of trust. But since you do, why the gloom?
Because at that point there's 2 scenarios:
- LLMs don't need humans anymore and we're either all dead or in a matrix-like farm
- Or companies realize they can't make LLMs buy the stuff their company is selling (with what money??) so they still need people to have disposable income and they enact some kind of Universal Basic Income. You can spend your days painting or volunteering at an animal shelter
Some people are rooting for the first option though, so while it's good that you've found faith, another thing that young people are historically good at is activism.
202X: SWE is solved
202X + Y; Y<3: All other fields solved.
In this case, I can't retrain before the second threshold but also can't idle. I just have to suffer. I'm prepared to, but it's hard to escape fleshy despair.
Seems more anti-fragile.
EVERYTHING is upturned. "All other things solved" includes robotics. It's a 10x everywhere.
Say there used to be 100 jobs in some company, all executing on the vision of a small handful of people. And then this shift happens. Now there are only 10 jobs at that company, still executing on the vision of the same handful of people.
90 people are now unemployed, each with a 10x boost to whatever vision they've been neglecting since they've been too busy working at that company. Some fraction of those are going to start companies doing totally new things--things you couldn't get away with doing until you got that 10x boost--things for which there is no training data (yet).
And sure, maybe AI gets better and eats those jobs too, and we have to start chasing even more audacious dreams... but isn't that what technology is for? To handle the boring stuff so we can rethink what we're spending our time on?
Maybe there will have to be a bit of political upheaval, maybe we'll have to do something besides money, idk, but my point is that 10x everywhere opens far more doors than it shuts. I don't think this is that, but if this is that, then it's a very good thing.
Most people are just drones, and that's fine, that's just not them.
If AI can do the drone work, we may find more vision among us than we've come to expect.
Work on your soft skills. Join a theater club, debate club, volunteer to speak at events, ...
Not that it's easy, and certainly more difficult for some people than for others, but the truth is that soft skills already dominate engineering, and in a world where LLMs replace coders they would become more important. Companies have people at the top, and those people don't like talking to computers. That is not going to change until those people get replaced.
And say like LLMs get good enough to displace 30% of the people that do those. That's enormous economic devastation for workers. Enough that it might dent the supply side as well by inducing a demand collapse.
If it's 90% of all jobs (that can't be done by a robot or computer) gone, then how are all those folks, myself included, going to find money to feed ourselves? Are we going to start sewing up t-shirts in a sweatshop? I think there are a lot of unknowns, and I think the answers to a lot of them are potentially very ugly
And not, mind, because AI can necessarily do as good a job. I think if the perception is that it can do a good enough job among the c-suite types, that may be enough
What exactly is the .01% of engineering work that this super intelligent AI couldn't handle?
I'm not worried about this future as a SWE, because if it does happen, the entire world will change.
If AI is doing all software engineering work, that means it will be able to solve hard problems in robotics, for example in manufacturing and self driving cars.
Wouldn't it be able to create a social network more addictive than TikTok, for anyone who might watch? This AI wouldn't even need human cooperation, why couldn't it just generate videos that were addictive?
I assume an AI that can do ultra complex AI work would also be able to do almost all creative work better than a human too.
And of course it could do the work of paper shuffling white collar workers. It would be a better lawyer than the best lawyer, a better accountant than the best accountant.
So, who exactly is going to have a job in that future world?
"Fix all mental illness". Ok.. yes, this might happen but what exactly does it mean?
"Increased social justice". Look around you my guy! We are not a peaceful species nor have we ever been! More likely someone uses this to "fix the mental illness of not understanding I rule" than any kind of "social justice" is achieved.
It will be greater than anyone but it won't be able solve THAT problem or any problem created after 2026, I can tell.
CEOs with little faith in their own products. Most likely it's widespread imperfect AI for a long while == unprofitable death for his company.