Generative AI is overrated, long live old-school AI
encord.com
encord.com
Now, "AI" in people's minds means generative models.
That's it, it doesn't mean generative models are replacing CNNs, just like CNNs don't replace SVMs or regression or whatever. It's just that pop culture has fallen in love with something else.
But neither the traditional nor generative models are "AI" in the sense that normal people think when they hear "AI".
In my grad school, we were working on something similar - using computer vision to analyze reactor flows to then change process variables. The results would be fed back into the system for RL. Too bad the project sorta froze after I graduated.
We actually use more than one neural network for this. The software is designed so the NN component is a plugin. The reason we do this is because some types of neural nets work better for some tasks than others.
Most (but not all) of our nets are convolutional.
Since you've already done some work with this sort of thing, I'm unsure about what level of overview would be of value to you, but this looks reasonable for a technically competent person who is new to the topic:
https://towardsdatascience.com/a-comprehensive-guide-to-conv...
As with all neural nets, the "secret sauce" isn't the code, it's the training.
Imagine asking an AI assistant to perform a certain industrial control task. The assistant, instead of executing the task “itself”, could figure out which model/system should perform the task and have it do it. Then even monitor the task and check it’s completion.
ChatSpot can understand your commands and then perform actions in the system for you, for example add a lead, change their contact info, write a blog post, publish it, add an image…
Edit: but if you connected it with physical actions, it could control your house, maybe check your smart refrigerator, order food on Instacart, send you recipe, schedule the time to cook in your calendar, request an Uber to pick you up from work, invite someone over, play music…
There’s a discussion about this on another homepage thread here: https://news.ycombinator.com/item?id=35172362
Also, even if a LLM could do that, so could a shell script, without the risks involved in using "AI" for it, or for now the ridiculous external dependence that would involve.
I wonder if in 10 years people will be stuck debugging Rube-Goldberg machines composed of LLM api calls doing stuff that if-statements can do, probably cobbled together with actual if-statements
In other words, old-school expert systems.
What your are saying is: “why use the washing machine, if I my clothes are even cleaner when I wash them myself - I also spend less detergent and less water”.
You are free to keep doing your laundry by hand.
But I bet most people prefer the washing machine.
Like it or not, an AI’s behavior is a black box and can’t be “proven” to execute exactly the same every time for the scenarios you are targeting.
A shell script will do exactly what it has been written to do every time, unless tampered with. And if changes need to be made, it can be done quickly without need for retraining, god knows how long that would take for an AI to learn something new. God help you if you need to maintain “versions” of your AI, trained for different things.
Face it, AI are pointless and slow for certain classes of problems.
Or unless some magic environment variable changes, or one of the runtime dependencies changes, or it is run on a different operating system, or permissions aren't setup right, or one of its tasks errors out.
Shell scripts are digital duct tape, the vast majority of shell scripts do not come close to being reliable software.
> god knows how long that would take for an AI to learn something new
Did you watch OpenAI's demo yesterday? They pasted in new versions of API docs and GPT4 updated its output code. When GPT forgot a parameter, the presenter fed back the error message and GPT added the parameter to the request.
You don’t have to feed a developer code or docs, you can give them a high level idea and they’ll figure it out on their own if you want.
The big thing everyone in this single thread is missing is that AI is a metaheuristic.
I wouldn't expect to use AI to run_script.py. That's easy. I'd expect it to look at the business signals and do the work of an intern. To look at metrics and adjust some parameters or notify some people. To quickly come up with and prototype novel ways to glue new things together. To solve brand new problems.
It’s not there yet.
Sooner or later it may (or may not) be a true statement, but it's awfully hard for me to say that it's any different right now than it has been before.
You could solve 5% of cases now and over five years drive it up beyond 100% (where you're getting new customers and startups that had never even tried the previous methods).
More shitty CRUD apps? Which get easier and easier to pump out everyday with the growing numbers of frameworks, libraries, copying and pasting snippets from stack overflow?
Or will AI really write all the code for all our critical systems, operating systems, financial markets, planes, factories, vehicles, spacecraft? And will they do it so confidently and accurately that humans can safely just forget how to code?
Sure, perhaps. But by then AI will also be so advanced and independent in its problem solving that it will have no need to listen to human prompts.
I don’t really see the point in that.
E.g. I posted a while back how I had it write the guts of a DNS server. It produced a rough outline after the first request, and would fill out bit by bit as I asked it to elaborate or adjust specific points. The typical intern would not know where to start and I'd need to point them to the RFC, and they'd go off and read them and produce something overwrought and complex (I've seen what even quite experienced software devs produce when given that task; and I know how much work it took me the first time I did it).
So it may not exactly replace an intern, in that there are classes of problems that require low-level reasoning and a willingness and ability to go off and research that it's just not set up for yet and that will be harder to replace. But the problem set will change. Both in that what gets to the intern will be things where LLMs don't produce good result fast enough (I wouldn't ask an intern to do something what ChatGPT can do well with little prompting), and that interns will be more likely to go off and learn a bit and then spend more time prompting LLMs and in that sense produce more value than they could before.
The problem is Jira workflows are designed for reporting results to upper management.
If the cost savings / productivity benefits[1] are there, new workflows will appear.
[1] Then again there are huge productivity benefits to be gained by simplifying existing Jira workflows, yet such steps are not taken.
I wish I lived in that universe.
"Write me a shell script that runs run_control.py with the speed argument set to one hundred."
to get,
"./run_control.py --speed 100"
As an AI language model, I do not have direct access to image editing software, but I can provide you with a shell script that could achieve the described image manipulation using ImageMagick, assuming it is installed on your system.
Here's an example shell script that you could use as a starting point:
bash Copy code #!/bin/bash
# Set the input image filename input_file="input.jpg"
# Set the output image filename output_file="output.jpg"
# Increase the redness of the image convert "$input_file" -modulate 100,150,100 -channel R -evaluate multiply 1.2 +channel temp1.jpg
# Flip the image vertically convert temp1.jpg -flip temp2.jpg
# Crop the bottom 25% of the image convert temp2.jpg -gravity South -chop 0x25% temp3.jpg
# Compress the image to 80% convert temp3.jpg -quality 80 "$output_file"
# Clean up temporary files rm temp*.jpg To use this script, save it to a file (e.g., "image_manipulation.sh"), make it executable (e.g., "chmod +x image_manipulation.sh"), and run it in the same directory as your input image. The output image will be saved in the same directory with the name "output.jpg".
Please note that this is just an example script and may need to be modified depending on your specific image manipulation needs.
Shell scripts can't create intermediate bitmap images in memory and operate on them with normal drawing and filtering apis like opencv, without lossily serializing and deserializing them as jpg files, or even operate on json structures directly.
It would be a much stronger argument if your example used python and opencv or any other normal library, instead of incredibly inefficient and hard to write and maintain invocations to full blown unix commands run in separate processes instead of direct library calls to manipulate in-memory images.
There's no reason "AI" code generation has to use the worst possible language and technique to generate code to solve a problem.
It's like having a food replicator from Star Trek: TNG, and asking it to make you rancid dog food instead of lobster.
write me a function in python that ...
I've always forgot the syntax for alot of functions/libraries etc.
Also, I haven't really written lot of python until recently.
> I have a hard time imagining a future where someone tells an LLM, "Write me a shell script that runs run_control.py with the speed argument set to one hundred."
I'll point out that we already live in a world where single lines of pure function code are distributed as an NPM packages or API calls.
Run run control with speed argument 100.
AI: "Scheduling you for a speech therapist session to work on your stutter"
I've been a developer for a long-ass time, though I don't have super frequent occasion where I find it worthwhile to write a shell script. It comes up occasionally.
In the past 2 weeks I've "written" 4 of them via ChatGPT for 1-off cases I'd have definitely found easier to just perform manually. It's been incredible how much easier it was to just get a working script from a description of the workflow I want.
Usually I'd need to double check some basic things just for the scaffolding, and then, maybe double check some sed parameters too, and in one of these cases look up a whole bunch of stuff for ImageMagick parameters.
Instead I just had a working thing almost instantly. I'm not always on the same type of system either, on my mac I asked for a zsh script but on my windows machine I asked for a powershell script (with which I'd had almost no familiarity). Actually I asked for a batch file first, which worked but I realized I might want to use the script again and I found it rather ugly to read, so I had it do it again as a powershell script which I now have saved.
Sure though, someone won't tell an LLM to write a shell script that just calls a python script. They'd have it make the python script.
Whose time and effort do they think they're saving by bringing more shell scripts into existence, anyway?
Shell scripts are objectively and vastly worse than equivalent Python scripts along all dimensions.
It's like having a food replicator from Star Trek: TNG, and asking it to make you rancid dog food instead of lobster.
Sounds like an extension of https://en.wikipedia.org/wiki/Wirth%27s_law. How many times have I done some simple arithmetic by typing it into my browser's bar and checking out the google calculator results? When a generation ago I would have plugged it into a calculator on my desk (or done it in my head, for that matter...). I would be entirely unsurprised to hear that in another generation we're using monstrously complicated "AI" systems to perform tasks that could be done way more simply/efficiently just because it's convenient.
There are lots of systems where you're taking some information about a user and making a best guess at what action the system should take. Even without a need for super high accuracy these rule systems can get surprisingly complex and adding in new possible decisions can be tricky to maintain. In LLM world you just maintain a collection of possible actions and let the LLM map user inputs to those.
With a neural network you have a black box and for example with ChatGPT it doesn't even have a specification. It turns the verification process upside down.
We are going to be able to automate everything and anything with the proper feedback loops.
For example, you could have an app that writes itself, deploys itself, tests itself, receives feedback, updates itself based on the feedback, writes additional tests, does CI/CD.
At that point you will be just creating and directing. Or you can choose whatever you actually want to execute.
And then if those same kind of processes are given access to physical tools, they could do all of our manufacturing, design and build their own machines and infrastructure.
We could essentially collaborate with our systems in the most amazingly seamless way.
Once an "AI" system becomes reliable, we quickly take it for granted and it no longer seems impressive or interesting. It's just a database. Or an image classifier. Or a chatbot.
Regression, both linear and logistic are from the mid 1800s to early 1900s. Neural networks, at least the basics are from around 1950.
What has really changed is the engineering, the data volume and the number of fields we can apply the mathematics to. The math itself (or what is the basis of AI) is really old.
and it was only in the last decade that the vanishing gradients problem was tamed.
my impression is that ML researchers were stumbling along in the mathematical dark, until they hit a combination (deep neural nets trained via stochastic gradient descent with ReLU activation) that worked like magic and ended the AI winter.
One of the big pieces was Schmidhuber's lab's highway nets, done ~30 years ago, but just didn't land until a more limited version was rediscovered.
like, regression, sure - because it's a tool to measure how well a hypothesis (polynomial function) matches the data (points.) and CNNs are still foundational in computer vision. but the first and last time I heard of SVMs was in college, by professors who were weirdly dismissive of these newfangled deep neural networks, and enamored by the "kernel trick."
but aren't SVMs basically souped up regression models? are they used in anything ML-esque, i.e. besides validating a hypothesis about the behavior of a system?
Yes they are. They allow for non-linear decision boundaries and more dimensions than rows of data, which for many other ML methods is a problem.
Linear regression, logistic regression, SVM and CART decision trees are all still very popular in the real world where data is hard to come by.
LOL. Exact same experience in my college courses. Glad to know it's universal.
I'm pretty cynical on LLMs(i.e. they're not intelligent and won't take all our jobs soon), but am coming around on their importance and capabilities.
We love the model because it speaks our language as if it's "one of us", but this may be deceiving, and the complete lack of model for truth is disturbing. Making silly poems is fun but the real uses are in medicine and biology, fields that are so complex that they are probably impenetrable to the human mind. Can Reinforcement learning alone create a model for the truth? The Transformer does not seem to have one, it only works with syntax and referencing. How much % of truthfulness can we achieve, and is it good enough for scientific applications? If a blocker is found in the interface between the model and reality, it will be a huge disappointment
there seems to be accumulating evidence that "finding the optimal solutions" means (requires) building a world model. Whether it's consistent with ground truth probably depends on what you mean by ground truth.
Given the hypothesis that the optimal solution for deep learning presented with a given training set, is to represent (simulate) the formal systemic relationships that generated that set, by "modeling" such relationships (or discovering non-lossy optimized simplifications),
I believe an implicit corollary, that the fidelity of simulation is only bounded by the information in the original data.
Prediction: a big enough network, well enough trained, is capable of simulating with arbitrary fidelity, an arbitrarily complex system, to the point that lack of fidelity hits a noise floor.
The testable bit of interest being whether such simulations predict novel states and outcomes (real world behavior) well enough.
I don't see why they shouldn't, but the X-factor would seem to be the resolution and comprehensiveness of our training data.
I can imagine toy domains like SHRDLU which are simple enough that we should be able to build large models well enough already to "model" them and tease this sort of speculation experimentally.
I hope (assume) this is already being done...
Was this ever in doubt? This has been the case forever (even before "AI"), and I thought it was well-established. The fidelity of the model is the core problem. What "AI" is really providing is a shortcut that allows the creation of better models.
But no model can ever be perfect, because the value of them is that they're an abstraction. As the old truism goes, a perfect map of a terrain would necessarily be indistinguishable from the actual terrain.
Not sure why but I find this incredibly insightful…
That is a pretty good description of human brains/bodies. You could also say that quantum physics is where our noise floor might be.
Without sensing/experiencing the world, there is no truth.
The only truth we can ever truly know, is the present moment.
Even our memories of things that we “know” that happened, we perceive them in the now.
Language doesn’t have a truth. You can make up anything you want with language.
So the only “truth” you could teach an LLM, is your own description of it. But these LLMs are trained on thousands or even million different versions of “truth”. Which is the correct one?
There are a whole range of tasks that can’t be done today with an LLM because of the hallucination issues. You can’t rely on the information it gives you when writing a research paper, for example.
The rules don’t determine the interpretation.
An LLM will pretty much always respect the rules of language, but it can use them to tell you completely fake stuff.
https://arxiv.org/abs/2212.03827
Another approach - a model can learn the distribution - is this fact known or not in the training set, how many times does it appear, is the distribution unimodal (agreement) or multi-modal (disagreement or just high variance). Knowing this a model can adjust its responses accordingly, for example by presenting multiple possibilities or avoiding to hallucinate when there is no information.
In all seriousness though, what you are asking is whether an objective reality exists which is not a settled debate. There is also the whole solipsism thing though many disregard as a valid view of the world because it can be used to justify anything and is not a particularly interesting position.
Of course there is also the whole local realism thing with QM and of course the whole relativity thing and time flowing at different speeds destroying a universal “now”.
Then there is the whole issue with our senses being fallible and our brains hallucinating reality in a manner that is as confident as GPT3.5 is when making up facts.
In fact, it’s all just information and information doesn’t need a medium.
To my mind generative AI is great at finding needles in the haystack of stuff we already know. Of course it just as often gives you a fake needle right now, just to see if you notice.
On the other hand "traditional"/predictive AI is often better at the things we don't already know or understand.
I mean, the only thing GPT does is predict the next word, which makes it not so different from a compression algorithm. And diffusion models (the image generating stuff) are essentially fancy denoisers.
Depending on how you assemble the big building blocks, you get generation or you get prediction.
If that is the definition of old school AI, I wonder how symbolic AI should be named.
It is not. Symbolic, deductive reasoning engines have the same claim to being old-school AI as predictive statistic models.
Although we do have a litmus test in asking it "What is the meaning of life the universe and everything?"
“Unsupervised generative AI” is useless IMO.
For example: "Multiple-choice questions in 57 subjects (professional & academic)" - https://openai.com/research/gpt-4
Meanwhile, there are off-shelf models that you can train very efficiently, on relevant data, privately, and you can run these on your own infrastructure.
Yes, GPT4 is probably great at all the benchmark tasks, but models have been great at all the open benchmark tasks for a long time. That's why they have to keep making harder tasks.
Depending on what you actually want to do with LMs, GPT4 might lose to a BERTish model in a cost-benefit analysis--especially given that (in my experience), the hard part of ML is still getting data/QA/infrastructure aligned with whatever it is you want to do with the ML. (At least at larger companies, maybe it's different at startups.)
What happens with completely new questions from totally different subject. The generative model will produce nonsense.
Basically a form of adversarial training/generation.
I believe that we'll see the most success/accuracy once you have generative AI compare itself to itself, monitored by a GAN, which then spits out it's answer while retaining some knowledge as to how it came to the conclusion. A tricameral mind.
Their power does not only lie in their ability to _generate_ new data, but to _model_ existing data.
Interpretable models with transparent loss functions are easy to grok.
How LLMs might fail on a classic task is (afaict right now) difficult to predict.
If I train a classic deep net as a classifier and there are 5 possible classes, it will only ever output those 5 classes (unless there's a bug).
With ChatGPT, for example, it could theoretically decide to introduce a 6th class - what I would call an alien failure mode, even if you explicitly told it not to.
I think formally / provably constraining the output of LLM APIs will help mitigate these issues, rather than needing to use an embedding API / use the LLM as a featurizer and train another model on top of it.
(You give it the image and prompt it with the 1000 classes and ask it which one the image belongs to).
I'm surprised ClosedAI didn't include this kind of benchmark. I guess it doesn't do too well?
https://www.pinecone.io/learn/zero-shot-image-classification...
Good answer but I feel that most users/people do not understand the difference between generative and predictive machine learning and that will probably cause unpredictable failures and false flags. So yes it has been overhyped in my opinion
Basically I think it's overhyped by the use of the term "AI" and how easy we are to accept it generally. Some aspect of them being generative models could have been the term used to market/describe them, but instead a much broader term is used.
We're just years into generative approaches. And I think we'll more combinations of methods used in the future.
The goal of AI has never been to build an all knowing perfect system. It has also never been to replicate the way the human brain works. But its been to build an artificial system that can learn -- and AGI specifically to be able to give the appearance of human learning.
I feel like we've turned this corner where the question now is, "Can we build something that knows everything that has been documented and can also synthesize and infer all of that data at a level of a very smart human". The fact that this has become the new bar is IMO one of the biggest tech changes in history. Not the biggest, but up there.
The word "know" is doing some heavy lifting there, as is "synthesize" and "infer".
Now "infer" and "synthesize" I meant the standard human definition of "synthesize" and "infer". In my interactions with relatively bright people, they really expect ChatGPT to be able to synthesize text at the level of a very sharp HS/college student. They don't want simple regurgitation of a text or a middel school analysis -- they want/expect ChatGPT to analyze nuance, and pull in its vast database to make connections to things that maybe aren't apparent at first glance.
The bar has raised so high so quickly -- it's crazy.
Generative models are about characterising probability distributions. If you ever predict more than just the average of something using data, then you are doing generative modelling.
The difference between generative modelling and predictive modelling is similar to the difference between stochastic modelling and deterministic modelling in the traditional applied mathematical sciences. Both have their place. Neither is overrated.
Grab the best tool for the job.
"Don't be dazzled by AI computer vision's creative charm! Classical computer vision, though less flashy, remains crucial for solving real-world challenges and unleashing computer vision's true potential."
Meant for those in classical computer vision before ML ate the field.
It depends how you define "generate." For example, is software that controls a robot arm generating anything? I guess it's generating the movements of the arm. But when people use the term "generative" with regards to machine learning models right now, they generally mean content—e.g. text or images for consumption.
Generative AI is essentially the opposite of a classifier. You give it a prompt that could mean many different things, and it gives you one of those things. A robotic arm could use generative AI, because there are many different sets of electrical signals that would result in success for, say, catching a ball.
Classification is an example of a non-generative AI in that there is only 1 correct answer, but it still requires machine learning to acquire the classification function.
You may twist the language to say that they are generating a list of validations and errors, but even then it's definitely a different use case than merely creating new items.
TLDR; Don't be dazzled by generative AI's creative charm! Predictive AI, though less flashy, remains crucial for solving real-world challenges and unleashing AI's true potential. By merging the powers of both AI types and closing the prototype-to-production gap, we'll accelerate the AI revolution and transform our world. Keep an eye on both these AI stars to witness the future unfold.