Then there's chain-of-thought being positioned as the next big step forwards, which works by throwing more inferencing at the problem, so that cost can't be amortized over time like training can...
2) Absolutely. Thats like one hour of an engineer salary for a whole month.
Isn't each new model bigger and heavier and thus requries more compute to train?
Until the competition outcompetes you with their new model and you have to train a new superior one, because you have no moat. Which happens what, around every month or two?
> the hardware is getting much better and prices will fall drastically once there is a bit of a saturation of the market or another company starts putting out hardware that can compete with NVIDIA
Where is the hardware that can compete with NVIDIA going to come from? And if they don't have competition, which they don't, why would they bring down prices?
Huge margins lead to a lot of competition trying to catch up, which is what makes market economies so successful.
Eventually one of you runs out of money, but your customers keep getting better models until then; and if the loser in this race releases the weights on a suitable gratis license then your businesses can both lose.
But that still leaves your customers with access to a model that's much cheaper to run than it was to create.
Google however does not sell these, you can only lease time on them via GCP.
Raising the max intelligence of the models tends to raise the intelligence of all the models via distillation.
Short term pricing inefficiency is not relevant to long term impact.
If there is a common shift to using additional runtime compute to improve quality of output, such as OpenAI's GPT-o1, then FLOPs required goes up massively (OpenAI has said it takes exponential increase in FLOPS/cost to generate linear gains in quality).
So, while costs will of course decrease, those $20-30K NVIDEA chips are going to be kept burring, and are not going to pay for themselves ...
This may end up like the shift to cloud computing that sounds good in theory (save the cost of running your own data center), but where corporate America balks when the bill comes in. It may well be that the endgame for corporate AI is to run free tools from the likes of Meta (or open source) in their own datacenter, or maybe even locally on "AI PCs".
Drop the S, I think. There’s no time dimension.
And FLOP is a generalized capability meaning you can do any operation. Hardware optimizations for ML can deliver the same 100B computations faster and cheaper by not being completely generalized. Same way ray tracing acceleration works: it does not use the same amount of compute as ray tracing in general CPU’s.
Still, even with modern accelerators it's lot of computation, and is what drives the price per token of larger models vs smaller ones.
Llama, Mixtral, Stable diffusion and Flux are a lot of fun and free to run locally, you should try them out.
Let me use CG rendering as an example. Back in the day only the big companies could afford to do photoreal 3D rendering because only they had access to the compute and even then it would take days to render a frame.
Eventually people could do these renders at home with consumer hardware but it still took forever to render.
Now we can render photoreal with path tracing at near realtime speeds.
If you could go back twenty years and show CG artists the Unreal Engine 5 and show them it’s all realtime they would lose their minds.
I see the same for A.I., now it’s only the big companies that can do it, then we will be able to do it at home but it will be slow and finally we will be able to train it at home for quick and cheap.
It was an entire web app, with search filters, tree based drag and drop GUIs, the backend api server, database migrations, auth and everything else.
Not once did he need to ask me a question. When I asked him "how long did this take" and expected him to say "a few weeks" (it would have taken me - a far more experienced engineer - 2 months minimum).
His answer was "a few days".
What I'm not saying is "AGI is close" but I've seen tangible evidence (only in the last 2 months), that my 20 year software engineering career is about to change and massively for the upside. Everyone is going to be so much more productive using these tools is how I see this.
Step outside of building basic web/CRUD apps and its accuracy drops off substantially.
Also almost every library it uses is old and insecure.
I have also been developing for 20+ years.
And have heard the exact same thing about IDEs, Search Engines, Stack Overflow, Github etc.
But in my experience at least how fast I code has never been the limiting factor in my project's success. So LLMs are nice and all but isn't going to change the industry all that much.
If nothing really changes in 3-5 years, then I'd call it a flop. But the writing is on the wall that "scale = smarts", and what we have today still looks like a foundational stage for LLM's.
> If nothing really changes in 3-5 years, then I'd call it a flop
Transformers have been used for what 6 years now? Will you in 6 years say "I'll decide if they don't change the world in another 6 years?"
Working with AI-generated code to add new features feels like working with Dreamweaver-generated code, which was also unpleasant. It's not written the same way a human would write it, isn't written with ease of modification in mind, etc.
I've tried using LLMs for some libraries I'm working on, and they failed miserably. Trying to make an LLM implement a trait with a generic type in Rust is a game of luck with very poor chances.
I'm sure LLMs can massively speed up tasks like front-end JavaScript development, simple Python scripts, or writing SQL queries (which have been written a million times before).
But for anything even mildly complex, LLMs are still not suited.
Succeeding on the most common tasks (which isn't exactly what you said) is identical to "they're useful".
But it's also utterly failed to handle mundane tasks, like porting legacy code from one language and ecosystem to another, which is frankly surprising to me because I'd have assumed it would be perfectly suited for that task.
But the worst LLMs? One of my personal tests is "write Tetris as a web app", and the worst local LLM I've tried, started bad and then half way through switched to "write a toy ML project in python".
It’s a very useful tool, not magic.
front-end JS can easily also become very complex
I think a better metric is how close you are to reinventing a wheel for the thousands time. Because that is what LLMs are good at: Helping you write code which nearly the same way has already been written thousands of times.
But that is also something you find in backend code, too.
But that is also something where we as a industry kinda failed to produce good tooling. And worse if you are in the industry it's kinda hard to spot without very carefully taking a hounded (mental) steps back from what you are used to and what biases you might have.
We should not have an entire industry of 10,000,000 devs reinventing the JS/React/Spring/FastCGi wheel. Im sure those humans can contribute in much better ways to society and progress.
I'd have said the opposite. I think LLMs facilitate disposable code. It might use the same paradigms and patterns, but my bet is that most LLM written code is written specifically for the app under development. Are there LLM written libraries that are eating the world?
The code itself is not clean and reusable across implementations, but you don't even need that clean packaged library. You just have an LLM regenerate the same code for every project you need it in.
The LLM itself, combined with your prompts, is effectively the reusable code.
Now, this generates a lot of slop, so we also need better AI tools to help humans interpret the code, and better tools to autotest the code to make sure it's working.
I've definitely replaced instances where I'd reach for a utility library, instead just generating the code with AI.
I think we also have an opportunity to merge the old and the new. We can have AI that can find and integrate existing packages, or it could generate code, and after it's tested enough, help extract and package it up as a battle tested library.
E.g. let's say I'm working on a production thing and features/bugfixes accumulate and some file in the codebase starts to resemble spaghetti. The LLM can help me unfuck that way faster and get to a state of very clean code, across many files at once.
Could potentially mean just a change in time allocation/priority. As it's easier and faster to locate and potentially resolve issues later, it is less important for code to be consistent and perfectly documented.
Not fool proof and who knows how that could evolve, but just an alternative view. One of these big names in the industry said we'll have AGI when it speaks it's own language. :P.
LLMs are definitely suited for tasks of varying complexity, but like any tool, their effectiveness depends on knowing when and how to use them.
SQL is a good target language because the translation from ideas (or written description) is more or less linear, the SQL engine uses entirely different techniques to turn that query into a set of relational operators which can be rewritten for efficiency and compiled or interpreted. The LLM and the SQL engine make a good team.
1. Aasked ChatGPT to write a simple echo server in C but with this twist: use io_uring rather than the classic sendmsg/recvmsg. The code it spat out wouldn't compile, let alone work. It was wrong on many points. It was clearly pieces of who-knows-what cut and pasted together. However after having banged my head on the docs for a while I could clearly determine from which sources the code io_uring code segments were coming. The code barely made any sense and it was completely incorrect both syntactically and semantically.
2. Asked another LLM to write an AWS IAM policy according to some specifications. It hallucinated and used predicates that do not exist at all. I mean, I could have done it myself if I just could have made predicates up.
> But for anything even mildly complex, LLMs are still not suited.
Agreed, and I'm not sure we are any close to them being.
This is probably why there’s such a divide when you try to talk about software dev online. One camp believes that it boils down to duct taping as many ready made components together all in pursuit of impact and business value. Another wants to really understand all the moving parts to ensure it doesn’t fall apart.
Personally, I just use the hell out of Django for that. And since tools like that are already ridiculously productive, I don't see much upside from coding assistants. But by and large, so many of our tools are so surprisingly _bad_ at this, that I expect the LLM hype to have a lasting impact here. Even _if_ the solutions aren't actually LLMs, but just better tools, since we reconfigured how long something _should_ take.
Bad tools often falls in three categories. Too simple, too complex, or unsuitable. For the last two, you'd better switch but there's the human element of sunken costs.
> Current LLMs fail if what you're coding is not the most common of tasks. And a simple web app is about as basic as it gets.
These two complexity estimates don’t seem to line up.
I've used LLMs to generate quite a lot of Rust code. It can definitely run into issues sometimes. But it's not really about complexity determining whether it will succeed or not. It's the stability of features or lack thereof and the number of examples in the training dataset.
What I meant by complexity is not "a task that's difficult for a human to solve" but rather "a task for which the output can't be 90% copied from the training data".
Since frontend development, small scripts and SQL queries tend to be very repetitive, LLMs are useful in these environments.
As other comments in this thread suggested: If you're reinventing the wheel (but this time the wheel is yellow instead of blue), the LLM can help you get there much faster.
But if you're working with something which hasn't been done many times before, LLMs start struggling. A lot.
This doesn't mean LLMs aren't useful. (And I never suggested that.) The most common tasks are, per definition, the most common tasks. Therefore LLMs can help in many areas, and are helpful to a lot of people.
But LLMs are very specialized in that regard, and once you work on a task that doesn't fit this specialization, their usefulness drops, down to being useless.
And for me that is the best case scenario, it takes away the part we have to code / solve already solved problems again and again so we can focus more on the other parts of software engineering beyond writing code.
Personally I much prefer Chatgpt. I give it specific small problems to resolve and some context. At most 100 lines of code. If it gets more the quality goes to shit. In fact copilot feels like chatgpt that was given too much context.
All of my experiences with LLMs have been that for anything that isn't a braindead-simple for loop is just unworkable garbage that takes more effort to fix than if you just wrote it from scratch to begin with. And then you're immediately met with "You're using it wrong!", "You're using the wrong model!", "You're prompting it wrong!" and my favorite, "Well, it boosts my productivity a ton!".
I sat down with the "AI Guru" as he calls himself at work to see how he works with it and... He doesn't. He'll ask it something, write an insanely comprehensive prompt, and it spits out... Generic trash that looks the same as the output I ask of it when I provide it 2 sentences total, and it doesn't even work properly. But he still stands by it, even though I'm actively watching him just dump everything he just wrote up for the AI and start implementing things himself. I don't know what to call this phenomenon, but it's shocking to me.
Even something that should be in its wheelhouse like producing simple test cases, it often just isn't able to do it to a satisfactory level. I've tried every one of these shitty things available in the market because my employer pays for it (I would never in my life spend money on this crap), and it just never works. I feel like I'm going crazy reading all the hype, but I'm slowly starting to suspect that most of it is just covert shilling by vested persons.
In theory I'm learning from the LLM during this process (much like a real code review). In practice, it's very rare that it teaches me something, it's just more careful than I am. I don't think I'm ever going to be less slap-dash, unfortunately, so it's a useful adjunct for me.
I use it to write test systems for physical products. We used to contract the work out or just pay someone to manually do the tests. So far it has worked exceptionally well for this.
I think the core issue of the "do LLMs actually suck" is people place different (and often moving) goalposts for whether or not it sucks.
It works well as a smart documentation search where you can ask follow-up questions or when you know what the output should look like if you see it but can't type it directly from the memory.
For code assistants (aka copilot / cursor), it works if you don't care about the code at all and ok with any solution if it's barely working (I'm ok with such code for my emacs configuration).
When they say treat it like an intern, I'm so confused. An intern is there to grow and hopefully replace you as you get promoted or leave. The tasks you assign to him are purposely kept simple for him to learn the craft. The monotonous ones should be done by the computer.
[0]: https://gist.github.com/simonw/97e29b86540fcc627da4984daf5b7...
After 20 years of being held accountable for the quality of my code in production, I cannot help but feel a bit gaslit that decision-makers are so elated with these tools despite their flaws that they threaten to take away jobs.
Lots of people are terrible at going from 0 to 1 in any project. Me included. LLMs helped me a lot solving this issue. It is so much easier to iterate over something.
And that’s fine if the dev realizes what’s going on but when they attribute their own quirks to AI magic, that’s a problem.
It's almost as if the horde of former kleptocurrency bros have found a promising new seam of fool's gold to mine
I did it in little pieces and started over with fresh context each time the LLM started to get off in the weeds. I'm very happy with the result. The code is clean and well commented, the tests are comprehensive and the app looks nice and performs well.
I could have done all this manually too but it would have taken longer and I probably would have skimped out on some tests and gave up and hacked a few things in out of expedience.
Did the LLM get things wrong on occasion? Yes. Make up api methods that don't exist? Yes. Skip over obvious standard straightforward and simple solutions in favor of some rat's nest convoluted way to achieve the same goal? Yes.
But that is why I'm here. It's a different style of programming (and one that I don't enjoy nearly as much as pounding the keyboard). It's more high level thinking and code review involved and less worrying about implementation detail.
It might not work as well in domains which training data doesn't exist in. Also certainly if someone expects to come in with no knowledge and just paste code without understanding, reading and pushing back, they will have a non working mess pretty shortly. But overall these tools dramatically increase productivity in some domains is my opinion.
If you aren’t on board then it looks impressive but flawed and not even close to living up to the hype.
I've been impressed with the ability to generate "throw away" code for testing out an idea or rapidly prototyping something.
I have a good mental map of the projects I work on because I wrote them myself. When new business problems emerge, I can picture how to solve them using the different components of those applications. If I hadn't actually written the application myself, that expertise would not exist.
Your colleague may have a working application, but I seriously doubt he understands it in the way that is usually needed for maintaining it long term. I am not trying to be pessimistic, but I _really_ worry about these tools crippling an entire generation of programmers.
That sounds quite useful. Does Cursor feed your entire project code (traversing all folders and files) into the context?
My constant suspicion is that most results people are so impressed with were just never validated.
> Analogously, I imagine future software dev to consist mostly of writing specs in natural language.
https://www.commitstrip.com/en/2016/08/25/a-very-comprehensi...?
I like this take. I feel like a significant portion of building out a web app (to give an example) is boilerplate. One benefit of (e.g., younger) developers using AI to mock out web apps might be to figure out how to get past that boilerplate to something more concise and productive, which is not necessarily an easy thing to get right.
In other words, perhaps the new AI tools will facilitate an understanding of what can safely be generalized from 30 years of actual code.
I’d argue web frameworks don’t even help a lot in this regard still. They pile on more concepts to the leaky abstractions of the web. They’re written by people that love the web, and this is a problem because they’re reluctant to hide any of the details just in case you need to get to them.
Coworker argued that webdev fundamentally opposes abstraction, which I think is correct. It certainly explains the mountains of code involved.
It does seem inevitable that some large change will happen to our profession in the years to come. I find it challenging to predict exactly how things will play out.
Isn’t that the point? Degrade the user long enough that the competing user is on-par or below the competence of the tool so that you now have an indispensable product and justification of its cost and existence.
P.S. This is what I understood from a lot of AI saints in news who are too busy parroting productivity gains without citing other consequences, such as loss of understanding of the task or expertise to fact-check.
I'm not a mathematician, hell i did general maths at school. Currently I've been talking through scripting a method to mix dsd audio files natively without converting to tradional pcm. I'm about to use gpt to craft these scripts. There is no way I could have done this myself without years of learning. Now all I have to do is wait half a day so I can use my free gpt o credits to code it for me (I'm broke af so can't afford subs). The productivity gains are insane. I'd pay for this in a heartbeat if I could afford it.
In niche situations it's not helpful at all in writing code that works (or even close). It is helpful as a quick lookup for docs for libs or functions you don't use much, or for gotchas that you might otherwise search StackOverflow for answers to.
It's good for quick-and-dirty code that I need for one-off scripts, testing, and stuff like that which won't make it into production.
Yeah, if you want tic-tac-toe or snake, you can simply ask ChatGPT and it will spit out something reasonable.
But this is not much better than a search engine/framework to be honest.
Asking it to be "creative" or to tweak existing code however ...
The sad part is beginners using the boilerplate code won't get any practice building apps and will completely fail at the complex parts of an app OR try to use AI to build it and it will be terrible code.
10% more productive. What does that mean? If you mean lines of code, then it's an incredibly poor metric. They write more code, faster. Then what? What are the long-term consequences? Is it ultimately a wash, or even a detriment?
https://stackoverflow.blog/2024/03/22/is-ai-making-your-code...
Personally I am having a lot of fun, as an iOS developer, creating web games. No market in that, not really, but it's fun and I wouldn't have time to update my CSS and JS knowledge that was last up-to-date in 1998.
However, for a more typical software engineer, where every project is different, you have full lifecycle responsibility from design through coding, occasional production support, future enhancements, refactorings, updates for 3rd party library/OD updates, etc/etc, then how much of your time is actually spent pure coding (non-stop typing) ?! Probably closer to 10-25%, and certainly no-where near 100%. The potential overall time saving from a tool that saves, let's say, 10-25% of your code typing is going to be 1-5%, which is probably far less than gets wasted in meetings, chatting with your work buddies, or watching bullshit corporate training videos. IOW the savings is really just inconsequential noise.
In many companies the work load is cyclic from one major project to the next, with intense periods of development interspersed with quieter periods in-between. Your productivity here certainly isn't limited by how fast you can type.
If you pay Silicon Valley salaries this seems like a no-brainer. There are bigger time wasters elsewhere, but this is an easy win with minimal resistance or required culture change
Ok models already run locally; that aside, as the hosted ones are kinda similar quality to interns (though varying by field), the answer is "what you'd pay an intern". Could easily be £1500/month, depending on domain.
- ok it works, but it won't be useful.
- ok it's useful, but it won't scale.
- ok it scales, but it won't make any money.
- ok it makes money, but it's not going to last.
etc etc
After all, you could have used the exact same response in defense of web3 tech. That doesn't mean LLMs are fated to be like web3, but similarly the outcome that the current expenditure can be recouped is far from a certainty just because there are doubters.
I'm probably one of those "heavy users", though I've only been using it for a month to see how well it does. Here's my review:
Large completions (10-15 lines): It will generally spit out near-working code for any codemonkey-level framework-user frontend code, but for anything more it'll be at best amusing and a waste of time.
Small completions (complete current line): Usually nails it and saves me a few keystrokes.
The downside is that it competes for my attention/screen space against good old auto-completion, which costs me productivity every time it fucks up. Having to go back and fix identifiers in which it messed up the capitalization/had typos, where basic auto-complete wouldn't have failed is also annoying.
I'd pay about about $40 right now because at least it has some entertainment value, being technologically interesting.
If what I give it is too open ended, doesn't have enough info, etc, I'll still get a low quality output. Though I find I can steer it by asking it to ask clarifying questions. Asking it to build unit tests can help a lot too in bolstering, a few iterations getting the unit tests created and passing can really push the quality up.
100,000 ? 500,000 ?
If Copilot came for free and Azure cost a tiny bit more, nobody would even blink.
For ChatGPT in its current state, probably $1K/month.
When GPT4 was launched last year, the API cost was about $36/M blended tokens, but you can now get GPT4o tokens for about $4.4/M tokens, Gemini 1.5 Pro for $2.2/M or DeepSeek-V2 (as 21B A/236B W model that matches GPT4 on coding) for as low as $0.28/M tokens (over 100X cheaper for the same quality output over the course of about 1.5 years).
The just released Qwen2.5-Coder-7B-Instruct (Apache 2.0 licensed) also basically matches/beats GPT4 on coding benchmarks and quantized can not only can run at a decent speed on just about any consumer gaming GPU, but on most new CPUs/NPUs as well. This is about a 250X smaller model than GPT4.
There are now a huge array of open weight (and open source) models that are very capable and that can be run locally/on the edge.
> I'm more productive than ever before.
You realize that another way to read that sentence is "I am a really bad coder".
There are people who are getting “pretty rich” by trafficking humans, or selling drugs. Would you want to live in a society where such activities are encouraged? In the end, we need to look at technological progress (or any progress for that matter) as where it will bring us to in the future, rather than what it allows you to do now.
It also pisses me off that software engineering has such a bad reputation that everyone, from common folks to the CEO of nvidia, is shitting on it. You don’t hear phrases like “AI is going to change medicine/structural engineering”, because you would shit your pants if you had to sit in a dentist chair, while the dentist would ask ChatGPT how to perform a root canal; or if you had to live in a house designed by a structural engineer whose buddy was Claude. And yet, somehow, everyone is ready to throw software engineers under the bus and label them as "useless"/easily replaceable by AI.
But we're getting paid right? And the Prime van will keep delivering.
Op is addressing the hype that there is some linear path of improvement here and chatgpt 8.5 will be AGI.
To which people always seem to jump in with but it’s useful for me and makes me code faster. Which is fine and valid, just beside the point
I use Scheme a lot, but the 1970s MIT AI folks' contention that LISPs encapsulate the core of human symbolic reasoning is clearly ridiculous to 2020s readers: LISP is an excellent tool for symbolic manipulation and it has no intelligence whatsoever even compared to a jellyfish[1], since it cannot learn.
GPTs are a bit more complicated: they do learn, and transformer ANNs seem meaningfully more intelligent than jellyfish or C. elegans, which apparently lack "attention mechanisms" and, like word2vec, cannot form bidirectional associations. Yet Claude-3.5 and GPT-4o are still unable to form plans, have no notions of causality, cannot form consistent world models[2] and plainly don't understand what numbers actually mean, despite their (misleading) successes in symbolic mathematics. Mice and pigeons do have these cognitive abilities, and I don't think it's because God seeded their brains with millions of synthetic math problems.
It seems to me that transformer ANNs are, at any reasonable energy scale, much dumber than any bird or mammal, and maybe dumber than all vertebrates. There's a huge chunk we are missing. And I believe what fuels AI boom/bust cycles are claims that certain AI is almost as intelligent as a human and we just need a bit more compute and elbow grease to push us over the edge. If AI investors, researchers, and executives had a better grasp of reality - "LISP is as intelligent as a sponge", "GPT is as intelligent as a web-spinning spider, but dumber than a jumping spider" - then there would be no winter, just a realization that spring might take 100 years. Instead we see CS PhDs deluding themselves with Asimov fairy tales.
[1] Jellyfish don't have brains but their nerve nets are capable of Pavlovian conditioning - i.e., learning.
[2] I know about that Othello study. It is dishonest. Unlike those authors, when I say "world model" I mean "world."
I need to be convinced that an LLM is smarter than a honeybee before I am willing to even consider that it might be as smart as a human child. Honeybees are smart enough to understand what numbers are. Transformer LLMs are not. In general GPT and Claude are both dramatically dumber than honeybees when it comes to deep and mysterious cognitive abilities like planning and quantitative reasoning, even if they are better than honeybees at human subject knowledge and symbolic mathematics. It is sensible to evaluate Claude compared to other human knowledge tools, like an encyclopedia or Mathematica, based on the LLM benchmarks or "demonstrated LLM abilities." But those do not measure intelligence. To measure intelligence we need make the LLM as ignorant as possible so it relies on its own wits, like cognitive scientists do with bees and rats. (There is a general sickness in computer science where one poorly-reasoned thought experiment from Alan Turing somehow outweighs decades of real experiments from modern scientists.)
[1] People dishonestly claim LLMs fail at counting because of minor tokenization issues, but
a) they can count just fine if your prompt tells them how, so tokenization is obviously not a problem
b) they are even worse at counting if you ask them to count things in images, so I think tokenization is irrelevant!
But at the same time there is a lot of value to capture here by building solid applications around the capabilities that already exist. It might be a winter more like the "winter" image recognition went through before multimodal LLMs than the previous AI winter
a) childish motivated reasoning led people to think a fairly simple technology could solve profoundly difficult business problems in the real world
b) a culture of "number goes up, that's just science"
c) uncritical tech journalists who weren't even corrupt, just bedazzled
In particular I don't think generative AI is like cryptocurrency, which was always stupid in theory, and in practice it has become the rat's nest of gangsters and fraudsters which 2009-era theory predicted. After the dust settles people will still be using LLMs and art generators.
but if you remove implicit subventions from the AI/AGI hype then for many such tools the cost to benefit calculation of creating and operating will become ... questionable
furthermore the places where such tools tend to shine the most often places where the IT industry has somewhat failed, like unnecessary verbose and bothersome to use tools, missing tooling and troublesome code reuse (so you write the same code again and again). And this LLM based tools are not fixing the problem they just kinda hiding it. And that has me worried a bit because it makes it much much less likely for the problem to ever be fixed. Like I think there is a serious chance for this tooling causing the industry to be stuck on a quite sub-par plato for many many years.
So while they clearly help, especially if you have to reinvent the wheel for a thousands time, it's hard to look at them favorably.
E.g. I feel like it should be possible to first blast out a lot of repetitive code and then for LLM to go over all of it and abstract it reasonably, while tests are still passing.
How will that ever get solved, in this universe? Look at what C++ does to C, what TypeScript does to JavaScript, what every standard does to the one before. It builds on top, without fixing the bottom, paving over the holes.
If AI helps generate sane low level code, maybe it will help you make less buffer overflow mistakes. If AI can help test and design your firewall and network rules, maybe it will help you avoid exposing some holes in your CUPS service. Why not, if we're never getting rid of IP printing or C? Seems like part of the technological progress.
Similar this law isn't really a law for a good reason, it doesn't always work. Not everything gets cheaper (in a relevant amount) at scale.
No offense, but I have only seen people who barely coded before describe being "very productive" with AI. And, sure, if you dabble, these systems will spit out scripts and simpler code for you, making you feel empowered, but they are not anywhere near being helpful with a semi-complex codebase.
Let’s see in some years… long winter ahead.
- Generate boilerplate
- Generate extremely simple code patterns. You need a simple CRUD API? Yeah, it can do it.
- Generate solutions for established algorithms. Think of solutions for leetcode exercises.
So yeah, if that's your job as a developer, that was a massive productivity boost.
Playing with anything beyond that and I got varying degrees of failure. Some of which are productivity killers.
The worst is when I am trying to do something in a language/framework I am not familiar with, and AI generates plausibly sounding but horribly wrong bullshit. It sends me in some deadends that take me a while to figure out, and I would have been better just looking it up by myself.
- Generate boilerplate : Snippets, templates, and code generators
- Generate extremely simple code patterns : Frameworks
- Generate solutions for established algorithms : Libraries.
My point is that I don't think AI can meaningfully output code that would be useful beyond that, because that code is not available in its training data.
Whenever I see people going on about how AI made then super productive, the only thing I ask myself is "My brother in Christ, what the fuck are you even coding?"
Like it will generate code like `x && Array.isarray(x)` because `x && x is something` is a common pattern I guess - but it's completely pointless in this context.
It will often do roundabout shit solutions when there's trivial stuff built into the tool/library when you ask it to solve some problem. If you're not a domain expert or search for better solutions to check it you'll often end up with slop.
And the "reasoning" feels like the most generic answers while staying on topic, like "review this code" will focus on bullshit rather than prioritizing the logic errors or clearing up underlying assumptions, etc.
That said it's pretty good at bulk editing - like when I need to refactor crufty test cases it saves a bunch of typing.
It's like looking at the first version of an IDE that got intellisense/autocomplete and deciding that we'll be able to write entire programs by just pressing tab and enter 10,000 times.
I do not claim to know what the future holds, but I do feel the clock is ticking on the AI hype. OpenAI blew people's minds with GPTs, and people extrapolated that mind-blowing experience into a future with omniscient AI agents, but those are nowhere to be seen. If investors have AGI in mind, and it doesn't happen soon enough, I can see another winter.
Remember, the other AI winters were due to a disconnect between expectations and reality of the current tech. They also started with unbelievable optimism that ended when it became clear the expectations were not reality. The tech wasn't bad back then either, it just wasn't The General Solution people were hoping for.
They seem very good at writing SQL for example. All the commas are in the right place and exactly the right amount of brackets square curly and round. But when they get it wrong, it really shows up the lack of intelligence. I hope the froth and bubble in the marketing of these tools matures into something with a little less hyperbole because they really are great just not intelligent.
Coding AI assistants have done some impressive things, I’ve been amazed at how they sniffed out some repetitive tasks I was hacking on and I just tab completed pages of code that was pretty much correct. There is use. I pay for the feature. I don’t know if it’s worth 35% of the world’s energy consumption and all new fabrication resources over the next handful of years being dedicated to ‘ai chips.’ We arent looking for a better 2.0, we are expecting an exponentially better “2.0” and those are very rare.
I don't think the bulk of this VC money is predicated on AGI being around the corner.
But the general trend hopping nature of big VC money is real. Still, VCs manage to continue to make a profit despite this, otherwise the industry would have died off or shrunk the 10 other years HN critiqued this behaviour, so on the whole they must be doing something right.
There are so many things that can be automated out there that currently aren't. Other industries are extremely manual and process driven still. Many here tend to underestimate this.
Some programmers here will argue it's error prone or creating technical debt but most people don't care, if it works it works, one can worry about it breaking in 5 years time after its saved you considerably time and money.
Uhh, that's a pretty big change.
Who are you and what are you being so productive in?
These code assistants are wholly unable to help with the day to day work I do.
Sometimes I use them to remind me what flags to use with a tarball[0], so replaced SO, but anything of consequence or creativity and they flounder.
What are you getting out of this excess productivity? A pay raise? More time with your loved ones?
[0] https://xkcd.com/1168/ (addressing the tool tip, but hilariously, in regards the comics content that would be a circumstance where I would absolutely avoid trusting one of these ‘assistants’)
Outside of very very short isolated template creation for some kind of basic script or poorly translating code from one language to another, they have wasted more time for me than they saved.
The area they seem to help people, including me, the most in is giving me code for something I don't have any familiarity with that seems plausible. If it's an area I've never worked in before, it could maybe be useful. Hence why the less breadth of knowledge in programming you have, the more useful it is. The problem is that you don't understand the code it produces so you have to entirely be reliant on it, and that doesn't work long term.
LLMs are not and will not be ready to replace programmers within the next few years, I guarantee it. I would bet $10k on it.
Either you pay more and more to keep your job as it gets better, or the company pays any amount for it so they can replace you over and over as a barely useful cog.
The current state of it being cheap only exist as it is in beta and they need more info from you, the expert, until it no longer needs you
But they lived in grass huts and the highest they had ever been off the ground was when they climbed a tree.
One day a genius was born on the island. She built a structure taller than the tallest tree. "I call it a stepladder," she said. The people were amazed. They climbed the stepladder and looked down upon the treetops.
The people proclaimed "All we have to do now is make this a little higher and we can reach the moon!"
A mere 5 years ago I was firmly on the camp (alongside linguists mainly) that believed intelligence requires more than just lots of data and a next token predictor. I was clearly wrong and would have lost a $1000 bet it I had put my money where my mouth was back then. Anyone not noticing how far things have come I think are mostly moving goalposts and falling to hindsight bias.
A better analogy is that the genius person in the village built a step ladder made of carbon nanotubules. Some people proclaimed "All with have to do is keep going and we can reach the moon with a very tall ladder!' Other people - many quite smart - proclaimed reasonably: "This is impossible. You are not realizing the unique challenges and materials we have not yet researched we need to build something like that."
Some in society kept building. The ladder kept getting higher. They run into issues like oxygen and balance so the ladder is redesigned into an elevator. Challenges came and were thought insurmountable until redesigns were still found to work with the miraculous carbon nanotubule material which seemed like a panacea for every construction ill.
Regardless of how high the now elevator gets and regardless of how many times the elevator gets higher than what the naysayers firmly believed is impossible the same naysayers keep saying they will never get much higher.
And higher the elevator grows.
Eventually a limit is reached, but that limit ends up being far higher than any naysayers ever thought possible. And when the limit is reached the naysayers all gathered and said "told you this would be the limit and that it was impossible!'
The naysayers failed to see the carbon nanotubules for the revolutionary potential that it had, even if they were correct that it wasn't enough.
And little did everyone know that their society was mere months ago from another genius being born that would give them another catalyst on the order of carbon nano tubules that would again lead to dramatic unexpected and long-term gains to how high the elevator can grow.
> "Reminds me of autonomous vehicles a couple of years back".
I don't think any reasonable interpretation of "autonomous vehicle" includes the ability to change a tyre. My point is that sometimes hype becomes reality. It might just take a little longer than expected.
It has taken 10+ years to get to present day, from the start of the "deep learning revolution" around 2010. I vaguely recall Uber promising self-driving pickups somewhere around 8-10 years ago. A main difference between current AI systems and the systems behind the cyclical hype cycles ongoing since the 1950s is that these systems are actually delivering impressive and useful results, increasingly so, to a much larger amount of people. Waymo alone services tens of thousands of autonomous rides per month (edit: see sibling comment, I was out of date, it's currently hundreds of thousands of rides per month -- but see, increasingly), and LLMs are waaaaay beyond the grandparent's flippant characterization of "plausible-looking but incorrect sentences". That's markov chains territory.
But they aren't particularly autonomous, there's a fleet of humans watching the Waymos carefully and frequently intervening for the case where every 10-20 miles or so the system makes a stupid decision that needs human intervention: https://www.nytimes.com/interactive/2024/09/03/technology/zo...
I think Waymo only releases the "critical" intervention rate, which is quite low. But for Cruise the non-critical interventions was every 5 miles and I suspect Waymos are similar. It appears that Waymos are way too easily confused and left to their own devices make awful decisions about passing emergency vehicles, etc.
Which is in fact consistent with what self-driving skeptics were saying all the way back in 2010: deep learning could get you 95% of the way there but it will take many decades - probably centuries! - before we actually have real self-driving cars. The remote human operators will work for robotaxis and buses but not for Teslas.
(Not to mention the problems that will start when robotaxis get old and in need of automotive maintenance, but the system didn't have any transmission problem scenarios in its training data. At no time in my life has my human intelligence been more taxed than when I had a tire blowout on the interstate while driving an overloaded truck.)
If this is the end result, this is already a substantial business savings.
What "critical" intervention rate are you talking about? What network magically supports the required low latencies to remotely respond to an imminent accident?
How does your theory square with events like https://www.sfchronicle.com/sf/article/s-f-waymo-robotaxis-f... that required a service team to physically go and deal with the stuck cars, rather than just dealing with them via some giant remotely intervening team that's managed to scale to 10x rides in a year? (Hundreds of thousands per month absolutely.)
Sure, there's no doubt a lot of human oversight going on still, probably "remote interventions" of all sorts (but not tele-operating) that include things like humans marking off areas of a map to avoid and pushing out the update for the fleet, the company is run by humans... But to say they aren't particularly autonomous is deeply wrong.
I would be interested if you can dig up some old skeptics, plural, saying probably centuries. May take centuries, sure, I've seen such takes, they were usually backed by an assumption that getting all the way there requires full AGI and that'll take who knows how long. It's worth noticing that a lot of such tasks assumed to be "AGI-complete" have been falling lately. It's helpful to be focused on capabilities, not vague "what even is intelligence" philosophizing.
Your parenthetical seems pretty irrelevant. First, models work outside their training sets. Second, these companies test such scenarios all the time. You'll even note in the link I shared that Waymo cars were at the time programmed to not enter the freeway without a human behind the wheel, because they were still doing testing. And it's not like "live test on the freeway with a human backup" is the first step in testing strategy, either.
I was being vague - Waymo tests the autonomous algorithms with human drivers before they are deployed in remote-only mode. Those human drivers rarely but occasionally have to yank control from the vehicle. This is a critical intervention, and it seems like the rates are so low that riders almost never encounter a problem (though it does happen). Waymo releases this data, but doesn't release data on "non-critical interventions" where remote operators help with basic problem solving during normal operations. This is the distinction I was making and didn't phrase it very clearly. I think those people are intervening at least every 10-20 miles. And since those interventions always involve common-sense reasoning about some simple edge case, my claim is that the cars need that common-sense reasoning in order to get rid of the humans in the loop. I am not convinced that there's even enough drivers in the world to generate the data current AI needs to solve those edge cases - things like "the fire department ordered brand new trucks and the system can't recognize them because the data literally doesn't exist."
> First, models work outside their training sets.
This is incredibly ignorant, pure "number go up" magical thinking. Models work for simple interpolations outside their training data, but a mechanical failure is not an interpolation, it's a radically different change which current systems must be specifically trained on. AI does not have the ability to causally extrapolate based on physical reasoning like humans. I had never experienced a tire blowout but I knew immediately what went wrong, relying on tactile sensations to determine something was wrong in the rear right + basic conceptual knowledge of what a car is to determine the tire must have exploded. Even deep learning's strongest (reality-based) advocates acknowledge this sort of thinking is far beyond current ANNs. Transformers would need to be trained on the scenario data. There are mitigations that might work: simply coming to a slow stop when a separate tire diagnostic redlines, etc. But these might prove bitter and unreliable.
> Second, these companies test such scenarios all the time.
No they don't! The only company I am aware of which has tested tire blowouts is Kodiak Robotics, and that seemed to be a slick product demo rather than a scientific demonstration. I am not aware of any public Waymo results.
I'd also like to challenge people to actually consider how often humans are correct. In my experience, it's actually very rare to find a human that speaks factually correctly. Many professionals, including doctors (!), will happily and confidently deliver factually incorrect lies that sound correct. Even after obvious correction they will continue to spout them. Think how long it takes to correct basic myths that have established themselves in the culture. And we expect these models, which are just getting off the ground, to do better? The claim is they process information more similarly to how humans do. If that's true, then the fact they hallucinate is honestly a point in their favor. Because... in my experience, they hallucinate exactly the way I expect humans to.
Please try it, ask a few experts something and I guarantee you that further investigation into the topic will reveal that one or more of them are flat out incorrect.
Humans often simply ignore this and go based on what we believe to be correct. A lot of people do it silently. Those who don't are often labeled know-it-alls.
Implementation detail that will be solved as the price of AI training decreases. Right now only inference is feasible at scale. Transformers are excellent here since they show great promise at 'one shot' learning meaning they can be 'trained' for the same cost as inference. Hence the sudden boom in AI . We finally have a taste of what could be should we be able to not only inference but also train models at scale.
The winter is going to be warm because of all the heat generated by GPUs ;)
global warming is killing us all
Is it going to scale to "superintelligence?" Is it going to be "the last invention?" I doubt it, but it's going to be a big deal. At the very least, comparable to google search, which changed how people interact with computers/the internet.
LLMs, irrespective of how powerful, are all subject to the fundamental limitation that they don't know anything. The stochastic parrot analogy remains applicable and will never be solved because of the underlying principles inherent to LLMs.
LLMs are not the pathway to AGI.
So what exactly is the usefulness of this discussion? You think "I'll trust my gut" is a useful argument in a debate?
FWIW my gut happens to agree with yours.
It hurts the pride of technical people that there's a revolution going on that they aren't involved in. Easier to just deny it or act like it's unimpressive.
How can I account for the cynicism that's so common on HN? It's got to be a psychological mechanism.
I thought a forum of engineers would be more interested in the practical applications and possible future capabilities of LLMs, than in all these semantic arguments about whether something really is knowledge or really is art or really is perfect
Hard disagree. LLMs merely present the illusion of knowledge to the casual observer. A trivial cross examination usually is sufficient to pull back the curtain.
Repeatedly, we’ve thought that humans and animals were different in kind, only to find that we’re actually just different in degree: elephants mourn their dead, dolphins have sex for pleasure, crows make tools (even tools out of multiple non-useful parts! [1]). That could be true here.
LLMs are impressive. Nobody knows whether they will or won’t lead to AGI (if we could even agree on a definition – there’s a lot of No True Scotsman in that conversation). My uneducated guess is that that you’re probably right: just continuing to scale LLMs without other advancements won’t get us there.
But I wish we were all more humble about this. There’s been a lot of interesting emergent behavior with these systems, and we just don’t know what will happen.
[1]: https://www.ox.ac.uk/news/2018-10-24-new-caledonian-crows-ca...
LLM proponents seem unwilling to accept that we comprehend the words we speak/write in a way that LLMs are not capable of doing.
Maybe their salary depends on them not understanding it.
That effective theory is knowledge, literally.
People harping about “stochastic parrot” are just people repeating a shallow meme — ironically, like a stochastic parrot.
LLM models are very far off from humans in reasoning ability, but acting like most of the things humans do aren't just riffing on or repeating previous data is wrong, imo. As I've said before, humans been the stochastic parrots all along.
No it isn't. The previous state of the art was markov chain level random gibberish generation. What OP described is an enormous step up from that.
Why? Text training data is already exhausted.
Next focus will hopefully be on reasoning abilities. Probably gonna take another decade and a similar paper to attention is all you need before we see any major improvements...but then again all eyes are on these models atm so perhaps it'll be sooner than that.
Literally today I used Bing and it was making up API parameters.
Code example looked fine, but didnt reflect reality.
That'll be investor types who bring this stupid "winter" on, because they run their lives on hype and baseless predictions.
Technology types on the other hand don't give a shit about predictions, and just keep working on interesting stuff until it happens, whether it takes 1 year or 20 years or 500 years. We don't throw a tantrum and brew up a winter storm just because shit didn't happen in the first year.
In early 2022 there was none of this ChatGPT stuff. Now, we're only 2 years later. That's not a lot of time for something already very successful. Humans have been around for several tens of thousand years. Just be patient.
If investors ran the show in the 1960s expecting to reach the moon with an 18 month runway, we'd never have reached the moon.
The difference is current ML already has real use cases right now in its current form. Some examples are OCR, text to speech, speech to text, translation, recommendations (for eg. Facebook Tiktok etc.) and simple NLP tasks ("was [topic] mentioned in the following paragraph"). Even if AGI is proved impossible, these are real use cases that hold billions in value. And ML research is also considered a prestigious and interesting field academically and that will likely not change even if investors give up on funding AGI.
You missed the point of the parent comment's post. He's talking about the current post chatbot GenAI hype (i.e., the massive amounts of funding being poured into companies specifically after this turning point).
1. You don't need massive amounts of funding to work on ML. A good deal of important work in ML was done in universities (eg. GANs, DPM, DDPM, DDIM) or were published before the hype (Attention). The only qualifier here is that training cost a lot right now. Even so, you don't need billions to train and costs may go down as memory costs come down and hardware competition increases.
2. You don't need VC type investors to fund ML research. Large tech companies like Facebook, Google, Microsoft, ByteDance and Huawei will continue investing in ML no matter what, even if the total amount they invest goes down (which I personally don't think it will). Even if they shift away from chatbots and only focus on simpler NLP tasks as described above, related research will still continue as all these tasks are related. For example, Attention was originally developed for translation and Llama 3.2 isn't just a chatbot and can also do general image description, which is clearly important to Facebook and ByteDance for recommendations and to Google for image search and ads. Understating what people like and what they are looking at is a difficult NLP problem and one that many tech companies would like to solve. And better image descriptions could then improve existing image datasets by allowing better text-image pairs, which could then improve image generation. So hard NLP, image generation and translation are all related and are increasingly converging into single multimodal LLMS. That is, the best OCR, image generation translation etc. models may be ones that also understand language in general (ie. broad and difficult NLP tasks). The issue is that OP assumes it must be AGI or bust.
The fault lies with humans using AI for something sensitive, without having the AI pass through certification etc. Part of the problem is glacial pace of laws around things, but that's nothing new isn't it; us humans being whiny, argumentative, inefficient, emotional meat bags about every little thing. I wonder, once we do make AGI, if it will wonder why it took us so damn long to tax the disgustingly wealthy, implement ww public healthcare, UBI, etc and solve the housing crisis by gasp building more houses...
We evolved, so our deep, deep underlying motivations pretty much always circulate around self-preservation and reproduction (resource contention).
If a machine spits out some plausible looking text (or some cookie-cutter code copy-pasted from Stack Overflow) the human brain is basically hardwired to go "wow this is a human friend!". The current LLM trend seems designed to capitalize on this tendency towards sympathizing.
This is the same thing that made chatbots seem amazing 30 years ago. There's a minimum amount of "humanness" you have to put in the text and then the recipient fills in the blanks.
This is not a reasonable take on the current capabilities of LLMs.
Definitely in the coming decade, we can prepare for a lot of the simpler tasks in an office to be taken over by AI. There are plenty of scenarios in which someone is managing a spreadsheet because an SME doesn't have the money to hire developers to automate & maintain that process - with advanced LLMs they can get it done by asking it to.
It really feels like a substantive step forward in terms of computer utility kind of like spreadsheets, databases, apps. We'll see how far it takes us down the line of human replacement though.
IMO it will be done vertical by vertical, with no standard interface coming for a while.
Like, are we using entirely different products? How are we getting such different results?
A lot of the rest of the world are using it for other things. And at these other things, the results are less impressive. If you've had to correct a family member who got the wrong idea from whatever chat bot they asked, if you've ever had to point out the trash writing in an email someone just trusted AI to write on their behalf before it got sent to someone that mattered, or if you've ever just spent any amount of time on twitter with grok users, you should be exceptionally and profoundly aware of how unimpressive AI is for the rest of the world.
I feel we need less people complaining about the skepticism on HN and more people who understand these skeptics that hang out here already know how wonderful a productivity boost you're getting from the thing they're rightly skeptical about. Countering with "But my code productivity is up!" is next to useless information on this site.
I appreciate your anecdotes on failures/embarrassment for people outside of tech- there's pretty clearly a gap in experience, understanding, and marketing hype.
I don't think it's useless to ask what that gap is, and why GP got such poor results.
Our everyday lives should make it evident how much the working of our brain doesn't resemble that of our computers. Our experiences change our brains somehow but exactly how we don't have the faintest idea about and we can re-live these experiences somewhat which creates a memory but the mechanism is by no means perfect. There's the Mandela Effect https://pubmed.ncbi.nlm.nih.gov/36219739/ and of course "tip of my tongue" where we almost remember a word and then perhaps minutes or hours later it just bursts into our consciousness. If it's a computer why is learning so hard? Read something and bam, it's written in your memory, right? Right? Instead, there's something incredibly complex going on, in 2016 an fMRI study was made among the survivors of a plane crash and large swaths of the brain lit up upon recall. https://pubmed.ncbi.nlm.nih.gov/27158567/ Our current best guess is somehow its the connections among neurons which change and some of these connections together form a memory. There are 100 trillion connections in there so we certainly have our task cut.
And so we are here where people believe they can copy human intelligence when they do not even know what they are trying to copy falling for the latest metaphor of the workings of the human brain believing it to be more than a metaphor.
At the end of the day, does it matter? If humans can be fooled by artificial intelligence in pretty much all areas, and that intelligence surpasses ours by every possible measurement, does it really matter that it's not powered by biological brains? We haven't quite reached that stage yet, but I don't think this will matter when we do.
This is just preposterous. You can be fooled if you have no knowledge in the area but that's about it. With current tech there is, there can not be anything novel. Guernica was novel. No matter how you train any probabilistic model on every piece of art produced before Guernica it'll never ever create it.
There are novel novels (sorry for the pun) every few years. They delight us with genuinely new turns of prose, unexpected plot twists etc.
Also harken to https://garymarcus.substack.com/p/this-one-important-fact-ab... which also happens to include a verb made up on spot.
And yes we have cars which move faster than a human can but they don't compete in high jumps or climb rock walls. Despite we have a fairly good idea about the mechanical workings of the human body, muscles and joints and all that we can't make a "tin man", not by far. As impressive as Boston Dynamics demos are they are still very very far from this.
I wasn't talking about current tech, which is obviously not at human levels of intelligence yet. I would still say that our progress in the last 100 years, and the last 50 in particular, has been astonishing. What's preposterous is expecting that we can crack a problem we've been thinking about for millennia in just 100 years.
Do you honestly think that once we're able to build AI that _fully_ mimics humans by every measurement we have, that we'll care whether or not it's biological? That was my question, and "no" was my answer. Whether we can do this without understanding how biological intelligence works is another matter.
Also, AI doesn't even need to fully mimic our intelligence to be useful, as we've seen with the current tech. Dismissing it because of this is throwing the baby out with the bath water.
What made you think that is measurable and if it is then we can build something like that ever?
I already linked https://garymarcus.substack.com/p/this-one-important-fact-ab... did you read it?
What makes you think it isn't, and that we can't? The Turing test was proposed 75 years ago, and we have many cognitive tests today which current gen AI also passes. So we clearly have ways of measuring intelligence by whatever criteria we deem important. Even if those measurements are flawed, and we can agree that current AI systems don't truly understand anything but are just regurgitation machines, this doesn't matter for practical purposes. The appearance of intelligence can be as useful as actual intelligence in many situations. Humans know this well.
Yes, I read the article. There's nothing novel about saying that current ML tech is bad at outliers, and showcasing hallucinations. We can argue about whether the current approaches will lead to AGI or not, but that is beside the point I was making originally, which you keep ignoring.
Again, the point is: if we can build AI that mimics biological intelligence it won't matter that it's not biological. And a sidenote of: even if we're not 100% there, it can still be very useful.
Hydraulics, gear systems, and computers are all Turing complete. If you're not a dualist, you have to believe that each of these would be capable of building a brain.
The history described here is one where humans invent a superior information processor, notice that it and humans both process information, and conclude that they must be the same physically. The last step is obviously flawed, but they were hardly going to conclude that the brain processes information with electricity and neurotransmitters when the height of technology was the gear.
Nowadays, we know the physical substrate that the brain uses. We compare brains to computers even though we know there are no silicon microchips or motherboards with RAM slots involved. We do that because we figured out that it doesn't matter what a machine uses to compute; if it is Turing complete, it can compute exactly as much as any other computer, no more, no less.
LLMs don’t have any ability to choose to update their policies and goals and decide on their own data acquisition tasks. That’s one of the key needs for an AGI. LLM systems just don’t do that / they are still primarily offline inference systems with mostly hand crafted data pipelines offline rlhf shaping etc…
There’s only a few companies working on on-policy RL in physical robotics. That’s the path to AGI
OpenAI is just another ad company with a really powerful platform and first mover advantage.
They are over leveraged and don’t have anywhere to go or a unique dataset.
For instance, a number of my conversations with ChatGPT contain messages it attempted to steer its own future training with (were those conversations to be included in future training).
And.its not just incorrect sentences it's weird questions which are getting answered a lot better than ever before.
Why are you so dismissive? Have you ever talked or wrote with a computer which felt anything like a modern LLM? I have not
I have no idea about AGI but honestly how can you use claude or chatgpt and come away unimpressed? It's like looking at spaceX and saying golly the space winter is going to be harsh because they haven't gotten to Mars yet.
Mars is hard but there are paths forward. More efficient engines, higher energy density fuels, lighter materials, better shielding, etc, etc. It's hard but there are paths forward to make it possible with enough time and money. We have an understanding of how to get from what we have now to what makes Mars possible.
With LLMs, there is no path from LLM -> gAI. No amount of time, money or compute will make that happen. They are fundamentally a very 'simple' tool that is only really capable of one thing - predicting text. There is no intelligence. There is no understanding. There is no creativity or problem solving or thought of any kind. They just spit out text based on weighted probabilities. If you want gAI you have to go in a completely different direction that has no relationship with LLM tools.
Don't get me wrong, the work that's been done so far took a long time and is incredibly impressive, but it's a lot more smoke and mirrors than most people realize.
And "LLM's just make plausible looking but incorrect text" is silly when that text is more correct than the average adult a large percentage of the time.