I remember back in like 2011 or 2012 I wanted to use an SSD for a project in order to spend less time dealing with disk seeks. My internet research suggested that there were a number of potential problems with most brands, but that the Intel Extreme was reliable.
So I specified that it must be only that SSD model. And it was very fast and completely reliable. Pretty expensive also, but not much compared to the total cost of the project.
Then months later a "hardware expert" was brought on and they insisted that the SSD be replaced by a mechanical disk because supposedly SSDs were entirely unreliable. I tried to explain about the particular model being an exception. They didn't buy it.
If you just lump all of these together as LLMs, you might come to the conclusion that LLMs are useless for code generation. But you will notice if you look hard that OpenAIs models are mostly nailing the questions.
That's why right now I only use OpenAI for code generation. But I suspect that Falcon 180B may be something to consider. Except for the operational cost.
I think OpenAI's LLMs are not the same as most LLMs. I think they have a better model architecture and much, much more reinforcement tuning than any open source model. But I expect other LLMs to catch up eventually.
Except this isn't new. This is after throwing massive amounts of resources at it multiple decades after arrival.
The transformer architecture on which (I think) all recent LLMs are based dates from 2017. That's only "multiple decades after" if you count x0.6 as "multiple".
Neural networks are a lot older than that, of course, but to me "these things are made out of neural networks, and neural networks have been around for ages" feels like "these things are made out of steel, and steel has been around for ages".
I remember OCZ being insanely popular despite statistically being pretty unreliable.
Me, Kubernetes Haikus, time taken 84 seconds:
----------
Kubernetes rules
With its smooth orchestration
You can reach web scale
----------
Kubernetes sucks
Lost in endless YAML hell
Why is it broken?
Imagine you have a lot more computing resources in a multimodal LLM. It sees your request of count the syllables and realizes it can't do them from text alone (hell I can't and have to vocalize it). It then sends your request to a audio module and 'says' the sentence, then another listening module that understand syllables 'hears' the sentence.
This is how it works in most humans, now if you do this every day you'll likely make some kind of mental shortcut to reduce the effort needed, but at the end of the day there is no unsolvable problem on the AI side.
No.
I doubt you would fully trust a LLM to replace high risk jobs such as lawyers, doctors or pilots such that when something goes wrong as it is used unattended, there is no-one held to account for it to transparently explain its own mistakes and errors.
It is just nonsense to suggest that such systems are capable of ‘reasoning’ when it pretends to do so and repeats itself without understanding their own errors.
Thus, LLMs and other black-box AIs cannot be trusted for those high risk situations over a consensus of human professionals.
And, as an obligate customer of many large companies, you should be in favor of that as well. Most companies already automate, poorly, a great deal of customer service work; let us hope they do not force us to interact with these deeply useless things as well.
If the primary complaint is the blues that GPT-4 wrote is not that great, I think it is definitely worth the hype, given that a year before people argued that AI can never pass turing test.
And this is my biggest issue with the AI mania right now -- the models don't actually understand the difference between correct or incorrect. They don't actually have a conceptual model of the world in which we live, just a model of word patterns. They're auto complete on steroids which will happily spit out endless amounts of garbage. Once we let these monsters lose with full trust in their output, we're going to start seeing some really catastrophic results. Imagine your insurance company replaces thier claims adjuster with this, or chain stores put them in charge of hiring and firing. We're driving a speeding train right towards a cliff and so many of us are chanting "go faster!"
No they won't.
>they can go find someone else with more expertise when they are lacking.
They can but they often don't.
>the models don't actually understand the difference between correct or incorrect.
They certainly do
I see similar comments everywhere where AI is praised, and I don't get why you need to comment this. Literally no one ever said LLM surpassed experts in their field, so basically you aren't arguing against anyone.
To use the canonical example of "internet service support call," most issues are because the rep either can't do what you're asking (e.g. process a disconnect without asking for a reason) or because they have no visibility into the thing you're asking about (e.g. technician rolls).
I honestly think we'd be in a better place if companies freed up funding (from contact center worker salary) to work on those problems (enhancing empowerment and systems integration).
https://www.cnn.com/2023/08/30/tech/gannett-ai-experiment-pa...
If the AI is a lot cheaper than a human, then it can make business sense to replace the human even if the AI is not nearly as good.
We are updating our expectations very fast. We are fighting over a growing pie. Maybe the cost reduction from not having to pay human wages is much smaller than the productivity increase created by human assisted AI. Maybe it's not an issue to pay the humans. AI works better with human help for now, in fact it only works with humans, never capable of serious autonomy.
Capitalism baby! You must continually earn more to enrich the investor class regardless of the cost to society as a whole. Just because the pie grows in size doesn't mean those with the capitol have to share it with anyone else. Greed, unfortunately, is limitless.
If it takes a whole business day to "spin up" a human for a task, and takes literally 5 seconds to call an OpenAI API, then guess what? The API wins.
That's impossible, LLMs are not that good. They might be firing people and crashing service quality.
You'd also want to look at models that are well-suited to what you're doing -- some of these are geared to specific purposes. Folks are pursuing the possibility that the best model would fully-internally access various skills, but it isn't known whether that is going to be the best approach yet. If it isn't, selecting among 90 (or 9 or 900) specialized models is going to be a very feasible engineering task.
> The 12-bar blues progressions seem mostly clueless.
I mean, it's pretty amazing that they many look coherent compared to the last 60 years of work at making a computer talk to you.
That being said, I played GPT4's chords and they didn't sound terrible. I don't know if they were super bluesy, but they weren't _not_ bluesy. If the goal was to build a music composition assistant tool, we can certainly do a lot better than any of these general models can do today.
> The question is will any of these ever get significantly better with time, or are they mostly going to stagnate?
No one knows yet. Some people think that GPT4 and Bard have reached the limits of what our datasets can get us, some people think we'll keep going on the current basic paradigm to AGI superintelligence. The nature of doing something beyond the limits of human knowledge, creating new things, is that no one can tell you for sure the result.
If they do stagnate, there are less sexy ways to make models perform well for the tasks we want them for. Even if the models fundamentally stagnate, we aren't stuck with the quality of answers we can get today.
I expect additional advances at some point in the future.