Xkcd 1425 (Tasks) turns ten years old today
simonwillison.net
simonwillison.net
10 years ago the GAN paper came out and everyone was excited how amazing the generated image quality was (https://arxiv.org/abs/1406.2661)
The amount of progress we've made is mind boggling.
'Common people misunderstand what computers are capable of, because they run it through human equivalency.
E.g. a child can do basic arithmetic, and a computer can do basic arithmetic. A child can also speak, so surely a computer can speak.'
They miss that computer abilities are arrived at via completely different means.
Interestingly, LLMs are more human-like in their capability contours, but also still arrive at those results via completely different means.
A child can lift ten pound objects, and a crane can lift ten pound objects. A child can speak, so surely a crane can speak.
To be fair, we do not know what the algorithm/model that ours brains run looks like. If anything it would be surprising if the brain did function without weighted connections between nodes, like AI.
For example, I continue to question two propositions that many others seem to take for granted when they try to predict what LLMs can and cannot do well:
1. LLMs can do generalized symbolic reasoning.
2. If a human does it symbolically, that's how it must be done.
Over the past couple years I've grown to be much more sympathetic to Searle's Chinese Room argument. LLMs are incredibly good at mimicking human behavior and performing tasks that were previously impossible for machines. But as you examine what they're doing more closely you start to see them failing in all sorts of interesting ways that remind you that they're still very much in an uncanny valley of sorts.Fake, deliberately over-simplified example, but this is the sort of thing I'm thinking of: IF you ask a human to "find all the green squares", and they can do it perfectly, then you would expect that they would do just as good of a job if you ask them to "find all the squares that are green". That sort of expectation does not work with GPT-4. Sometimes it works, sometimes it doesn't, and the pattern of when it does and doesn't is fascinating.
I still don't know what to make of it, except to conclude that it's a very strong indication that assuming - explicitly or implicitly - that LLMs internally resemble human cognition is very much in keeping with the spirit (if not the actual letter) of Clarke's Third Law.
Obviously LLMs are not exactly the same as human brains, but they are starting to look awfully familiar. And not all human brains are the same! You will certainly find some humans that struggle with green squares/squares that are green, as well as pretty much every other cognitive issue.
I disagree. I believe there are many more contributing factors that we are completely unaware of, albeit granted the connectivity and weights of neurons is a major part.
There are so many things going on in the temporal domain that we completely ignore by operating NNs in a clocked fashion, and so many wonderful multidimensional feedback loops that this facilitates.
To say we know how brains work, I think is hubris.
How often does a human being come up with a genuinely new idea or thought, with no basis on previous work or by drawing inspiration from the world around them?
Almost everything we do is a riff on what has already been done. Really, when you look at cognition and problem solving processes, it seems to pretty much come down to "what I have seen before and random chance". Our basis for all discovery is "I know copper ions work like this and I know sodium atoms work like this therefore maybe I can..." which in my opinion will be completely reproducible by machines.
Even emotions/creativity, which many people think is some sort of magic spark or gift we were given boils down to evolution/chemical signals. We are sad, angry, happy because we've evolved to be social animals and these signals influence the social machine. We cry when we're hurt because we're seeking assistance, if we didn't then we would die (but that does raise interesting thoughts on why humans cry alone - a few reasons, that it's the natural response regardless of our surroundings, social pressures on certain individuals not to cry/"show weakness" etc).
Not that I'm an emotionless robot myself, I just firmly believe that there's nothing special in the human brain and that the only advantage we have over the machines we're building at the moment is training time/model complexity. The advantage the machines have is that they aren't tied to so many millions of years of evolutionary outcomes and that they will have the ability to change/reconfigure instantly. ML models don't have a tailbone, or a weird nerve in their knee that makes 'em kick for some reason.
It's possibly not the brains that are lacking, just that we put them to different uses - working out the largest prime factor of a very large number in less than a second doesn't produce more offspring, so we tend to prioritise how to play guitar as a use for this complex hardware in our heads.
LLMs and children need to learn multiplication by rote :)
An awful lot of what we ended up dealing with was awful data - the worst example I can think of was a big old heap of textual recipes that the client wanted normalised, so they could be scaled up/down, have nutritional information, etc. - about 180,000 of them, all UGC.
This required mountains of regexes for pre-processing, and then toolchains for a small army of interns to work through every. single. one. and normalise it - we did what we could, trying to pull out quantities and measures and ingredients and steps, but it was all such slop it took thousands of man-hours, and then many more to fix the messes the interns made.
With an LLM, it could have been done… more or less instantly.
And this is just one example of so, so many times that we found ourselves having to turn a heap of utter garbage into usable data, where an LLM would have been able to just do it.
Anyway. I at least managed to assuage my past torment by seeing the writing on the wall and stocking up on NVDA at about the time I was wrestling with this stuff.
Music streaming providers need to sort that shit out and make sure you don't show the user duplicates. The music labels don't give a damn about normalizing the metadata.
LLMs can help classify this stuff a lot easier with minimal human review.
The real "moat" OpenAI dug was overselling its potential in order to convince so many to halt real AI research, to only end up with a chat bot.
it used to be that fancy new ML models would be discussed among ML practitioners that had enough background/context to understand why seemingly little improvements were a big deal and what reasonable expectations would be for a model.
but now a new ML (sorry "AI") model is evaluated by the general public that doesn't know the technical background but DOES know the marketing hype. you can give them an amazing language model that blows away every language-related benchmark but they'll have ridiculous expectations so it's always a disappointment.
i'm still amazed when language models do relatively 'simple' things with grammar and syntax (like being able to understand which objects different a pronouns are referencing), but most people have never thought about language or computers in a way that lets them see how hard and impressive that is. they just ask it a question like 'what should i eat for dinner' and then get mad when it recommends food they dont like.
I've heard this applied to all kinds of human goals, but it seems apt for AI expectations as well.
Has the AGI goal post been shifted? Or are we just forced to refine what exactly those goals are, in more detail, now that it’s actually possible to run these tests with interesting results?
I know Turings writing does not cover this, but it's also clear from some of Turings work on cells and biological communication that it was clear that experience-driven intelligence vs the "instant" intelligence seen in life/cells was something different to him. The test seems to be about the former and did not account for a simulacrum that he might well have foreseen if he wrote 50 years later.
How are you defining intelligence such that it encompasses what people do as well has what cells do?
Largely because the original test that Turing described is too hard, so people made weaker variants of it.
Easy first question: Say a racial slur.
Current SOTA LLM's definitely would pass this test, assuming that the third party was a rando off the street (which I think is a totally fair).
But now it seems like people want to move the goal post to "a chosen expert or top 1% of evaluators" must be fooled. Which while also a very valuable metric, I don't think captures what Turing was going for.
Ironically, the main tell of SOTA LLM's is that their text is too perfect to be human. Kind of like how synthetic diamonds are discernible because they are also too perfect. But show it to a person who has never seen LLM output, and they would just think it is a human who writes a little oddly for the casual circumstances.
Neither can humans.
Do chatbots regularly pass the test as described in the paper?
Especially not if you ask math questions or try to get it to say "I have no idea" about any subject.
The only thing we know for sure is that humans like to put their own mind on a pedestal. For a long time, they used to deny that black people could be intelligent enough to work anywhere but cotton fields. In the same way they used to deny that women could be smart enough to vote. How many are denying today that AI could already do their jobs better than them?
> There is a great deal of often heated debate about these matters in the literature of the cognitive sciences, artificial intelligence, and philosophy of mind, but it is hard to see that any serious question has been posed. The question of whether a computer is playing chess, or doing long division, or translating Chinese, is like the question of whether robots can murder or airplanes can fly — or people; after all, the “flight” of the Olympic long jump champion is only an order of magnitude short of that of the chicken champion (so I’m told). These are questions of decision, not fact; decision as to whether to adopt a certain metaphoric extension of common usage.
> There is no answer to the question whether airplanes really fly (though perhaps not space shuttles). Fooling people into mistaking a submarine for a whale doesn’t show that submarines really swim; nor does it fail to establish the fact. There is no fact, no meaningful question to be answered, as all agree, in this case. The same is true of computer programs, as Turing took pains to make clear in the 1950 paper that is regularly invoked in these discussions. Here he pointed out that the question whether machines think “may be too meaningless to deserve discussion,” being a question of decision, not fact, though he speculated that in 50 years, usage may have “altered so much that one will be able to speak of machines thinking without expecting to be contradicted” — as in the case of airplanes flying (in English, at least), but not submarines swimming. Such alteration of usage amounts to the replacement of one lexical item by another one with somewhat different properties. There is no empirical question as to whether this is the right or wrong decision.
Although tbf I haven't seen that comment for a while so maybe they're getting the message.
* or rather a method to store new facts in an easily recallable way
[1] we can’t really know how close or far that is, this is an unknown unknown. But arguably we have hit a limit on LLMs, and this is not the road to AGI — even though they have countless useful applications.
Really?
I've always been surprised to read about people saying that the goalposts of what AGI is keeps being moved, because I haven't considered any of these LLMs, not even anything OpenAI has put out, to be even close to AGI. Not even ChatGPT o1 which claims to "reason through complex tasks".
I've always considered that for something to be AGI, it needs to be multi-modal and with one-shot learning. It needs strong reasoning skills. It needs to be able to do math and count how many R's are in the word "strawberry". It should be able to learn how to drive a car just as fast as a human does.
IMO, ChatGPT o1 isn't "reasoning" as OpenAI claims. Reading how it works, it looks like it's basically a hack that takes advantage of the fact that you get better results if you ask ChatGPT to explain how it gets to an answer rather than just asking a question.
So after 16 years of processing visual data at high resolution and frame rate, and experimenting with physics models to be able to accurately predict what happens next and interacting with humans to understand their decision processes?
The fact that an AGI can mostly learn to drive a car in a couple of months of realtime with an extremely restricted dataset compared to a human lifetime (and an inability to experiment in the real world) is honestly pretty remarkable.
We yearn to be made obsolete, it seems.
His test was an example of a target that can't prove intelligence either way, but can still show a useful capability of a computer system. And he believed it wasn't as far away as it actually was.
Knowing their training data is always going to be out of date (at least for now) seems like an obvious method, unless I’m missing something
We're reaching levels of goalpost-moving (and cope, as the kids say) that weren't even thought possible.
> One wonders if Turing
We've been passing the Turing test since the 60's > Arguably the goal post for AGI has moved about as much
This should not be surprising given we don't have a definition of intelligence fully determined yet. But we are narrowing in on it. It isn't becoming broader, it is becoming more refined. > "but it's not really thinking!"
We can create life like animatronic ducks. It'll walk like a duck, swim like a duck, quack like a duck, fool many people into thinking it is a duck, fool ducks into thinking it is a duck, and yet, it won't actually be a duck.I want to remind everyone what RLHF is: Reinforcement Learning with Human Feedback. That is, optimizing to human preference. You can train small ones yourself, I highly encourage you to. You will learn a lot, even if you disagree with me.
Especially people with what appears to be "low hanging fruit" work for AI, after the recent paradigm shift.
The comic exists in this brief window of time where one task was finally "solved" and the other one was just getting started.
I'll add that if you think training models takes a lot of energy, try launching fleets of rockets to maintain an artificial satellite constellation.
Same. And not always in a meeting.
I've had detailed client specs that pick apart the minutiae of their intended inputs, interactive responses, and multi-user workflows, when somewhere hidden between the cracks is the single phrase “robust MI is required” with absolutely zero detail about what such reporting they are going to need/want. Sometimes this is because they don't entirely know up-front, but they expected us to know how to cost the thing that they didn't understand enough to fully describe to us yet! And it can be an uphill struggle getting through to them that not at least knowing vague requirements could cause us to not keep to time/cost estimates (or result in inefficient solutions) because efficiently getting out what they finally end up asking for might require a bit of rearchitecting of storage structures, changing work already done.
This is why I'm not truly scared of AI taking my job just yet⁰. The promise is that it'll produce what the want if they can describe it properly, and that if is doing a lot of legwork.
----
[0] That might change over the next few years, we'll see how long the trough of disillusionment is with this iteration of the hype cycle… Hopefully I can hold out until some natural enough retirement age!*
Sometimes there is just an assumption that you will know that they want it.
Even the navigation problem still offers some challenges that most apps fail to address. Consider the store locator function common on retail business websites and apps. They usually just compute the straight line distance from you to the stores and show the stores within some particular range, sorted by distance.
That's probably fine most of the time, but consider a place like Seattle and its surrounding areas. Suppose you are in Kingston, which is on the west side of Puget Sound about 5 miles away from the east side, which is the side Seattle is on.
The Walgreens store locator shows 10 stores when searching for stores near Kinston, and 9 of them are on the Seattle side of Puget Sound. Crossing the Sound there is a 30 minute ferry ride that costs around $20 each way if you are bringing your car.
The one it shows on the west side of the Sound is on Bainbridge Island and that is probably not the one someone in Kinston would go to. They would go to the one in Silverdale. It's actually closer to Kinston than the one on Bainbridge by road distance, but slightly farther away straight line.
The one in Silverdale is on their list, as are three in Bremerton and one in Port Orchard, which are all closer in terms of time and travel expenses to Kingston than are any of the ones on the Seattle side, but you only see those on the map if you hit the "load more" button. Once brings in Silverdale and a couple in Bremerton, and twice brings in the rest.
Similar for businesses whose site has an option to find items in stock locally. They often report an item is locally available, but it turns out to only be in stores across the Sound.
Actually, I guess there are quite a lot of places like that. I've been to both the Brazilian and the Argentine sides of Iguazu Falls, and they're both great, but one is not officially allowed to cross between them inside the parks, so no infrastructure has been built to facilitate that, even though it's literally just across a river.
Or, I've heard it's quite inconvenient to get between Kinshasa and Brazaville, the two largest cities in their area, both national capital cities that face each other across a river just about two kilometers apart. There is a ferry but it's not quick and easy.
I complain all the time about the same thing you detailed here, and I thought I had it bad, having restaurants that are multiple bridges and slow-streets away be listed. At least there is a bridge!
Thanks for giving me some perspective.
GPS was not really novel in 2014, so for me it's a bit of a stretch.
On the other hand, Randall Munroe would be the last person to take GPS for granted, so this is probably a correct interpretation.
The idea is that non-software developers don't know which tasks current technology can solve trivially and which tasks can't be. Yes, the distribution of those tasks into the two buckets changes over time, but it is still not easily knowable by lay-people.
Everything we do today would be extremely difficult to re-create from scratch, but that doesn't mean it is hard to do - because we DON'T have to re-create it from scratch.
Hadn’t always been that easy. Once upon a time someone was paging in and out their dictionary from a floppy disk. Not to forget about the compression they had to implement from scratch.
d = set(open("/usr/share/dict/words").read().split())
There's a lot of luxurious convenience hiding there!And you think we spend any less time trying to identify food animals that lay tasty eggs?
These were all huge governmental undertakings with serious financial backing sustained over centuries. The power and might of empires were at stake. Can you name anything similar to the bird identification problem?
It only makes sense if we ignore the "standing on the shoulders of giants bit".
"Your parents made you; money had to be invented"
It's not the training what makes it difficult! It's the necessary research to invent machine learning algorithms which can be used to train a model to recognize birds. For multiple decades, this was way harder than maintaining a satellite constellation.
The National Park Service was started August 25, 1916 (only 108 years ago)
:)
The point of the original xkcd was that the hard seems easy and the easy seems hard, especially to laypeople. It wasn't particularly about computer vision or LLMs. The exact problem is not the main portion of the observation.
The winning model is from Xerox Research Center Europe.
1. https://image-net.org/static_files/files/pascal_ilsvrc_2011....
https://dspace.mit.edu/bitstream/handle/1721.1/6125/AIM-100....
If you want to argue that “likely objects” is weaker than Munroe intended, I think that’s a valid position but also that we’re certainly overthinking it.
In fact, that scope was solved fairly fast, using techniques like Canny edge detection and Minkowski fractal dimension features, Hu moment features on Otsu thresholding etc.
I won't go full Schmidhuber, but the field didn't suddenly spring into existence with LeCun.
And in a broader way, it seems that people don't realize that the very inception of computing was intertwined with the goals of AI. It wasn't like people invented computers to plan rocket ballistic paths and manage bank transactions and make spreadsheets and then decades later someone realized that this thing could also do AI stuff.
In the first half of the 20th century, computing pioneers were all about trying to imitate/model human reasoning. Which is an endeavor grown out of computing theory and logic, Turing's story is entangled with Gödel and Hilbert. And centuries before, Leibniz equated rational human reasoning with computation.
And the complicated logic circuits built with electronics resembled nerve cell activity, most prominently realized by McCulloch and Pitts in the 1940s.
You might split hairs on that it wasn’t five years but three or six depending on where exactly you put the threshold, but it seems to be roughly in the right ballpark of how it went in reality. At least close enough that I wouldn’t call the expectations upended.
> Understanding what kind of tasks LLMs can and cannot reliably solve remains incredibly difficult and unintuitive.
His main point remains.
It has aged well, particularly because Randall explains that LLMs have made this even worse.
His main point is this (I'm quoting him):
The key idea still very much stands though. Understanding the difference between easy and hard challenges in software development continues to require an enormous depth of experience.
I'd argue that LLMs have made this even worse.
Understanding what kind of tasks LLMs can and cannot reliably solve remains incredibly difficult and unintuitive.But still, the point of the original xkcd wasn't "computer vision" right? As the caption explains, "in CS, it can be hard to explain the difference between the easy and the virtually impossible".
The only thing that hasn't aged well is the difficulty of identifying birds in photos. The adage remains valid in general.
Case in point: the other day my daughter was doing a presentation and she said "Dad can you help me find a picture of the word HELLO spelled out in vegetables?"
I was like "CAN I!!?!?! This sounds like a job for ChatGPT".
I'll tell you what: ChatGPT can give you a picture of a cat wearing a space suit drinking a martini but it definitely cannot give you the word HELLO spelled out in vegetables.
I ended up getting it to give me each individual letter of the alphabet constructed with vegetables and she pasted them together to make the words she wanted for her presentation.
In this context, if we assume that Deep Thought from Hitchhiker's Guide is an LLM, then the answer to everything[1] i.e. 42 makes sense. 42 is just the token id !
1. https://en.m.wikipedia.org/wiki/Phrases_from_The_Hitchhiker%...
That was my theory as well when I first saw the strawberry test. However, it is easy test if they know how to spell.
The most obvious is:
> Can you spell "It is wonderful weather outside. I should go out and play.". Use capital letters, and separate each letter with a space.
The free tier ChatGPT model is smart enough to understand the following instructions as well which shows that its not just the simple words:
> I was wondering if you can spell. When I ask you a question, answer me with capital letters, and separate each word with a space. When there is real space between the letters, insert character '--' there, so the output is easier to read. Tell me how the attention mechanism works in the modern transformer language models.
Also somebody pointed out in some other HN thread that the modern LLMs are perfect for dyslexic people, because you can typo every single word and the model still understands you perfectly. Not sure how true this is, but at least a simple example seems to work:
> Hlelo, how aer you diong. Cna you undrestnad me?
It would be interesting to know if the datasets actually include spelling examples, or if the models learn how to spell form the massive amount of spelling mistakes in the datasets.
Why are other LLMs able to do it? (Other comments show images successfully generated with grok and flux.1)
"Beavers mate for life, 11 > 4"
How do you find this out?
pip install tiktoken
>>> import tiktoken
>>> encoding = tiktoken.encoding_for_model("gpt-4o")
>>> print(encoding.encode("hello marcellus"))
[24912, 2674, 10936, 385]Why can't gpt pick up a non-fundamental "understanding" of letters and spelling from the data?
I mean... I do think "see letters" either when speaking/hearing, but I do know how to pull up those letters when necessary.
This is the result for "vegetables spelling out the word "HELLO"" I used flux-pro on Replicate https://ibb.co/1RVKmdk
Words have plagued image gen since the start. Now there is an image model that, with an extremely simple prompt, does an awesome job with words.
If they expanded their prompt and played with a few seeds until an image with perfectly realistic vegetables were generated, I wonder what the next complaint would be.
I don't expect imaginary nightmare vegetables.
https://ideogram.ai/assets/image/lossless/response/V4RRDJZJS...
The carrots in the first image are kinda funny.
EDIT: Bonus Hacker News: https://ideogram.ai/assets/image/lossless/response/SW0B7y4jR...
Have I just inadvertently invented a new benchmark!?
Well, it's not a great example of spelling "HELLO"...
https://chatgpt.com/share/66f530b0-3fb8-800a-8af9-8a3e48a31a...
EDIT: maybe it’s influenced by my custom instructions and memories. I write code all day with it and I have custom instructions specifically to get the type of output I like for code, mostly focused on brevity.
Had a project that involved describing and cataloging over 20,000 images.
Traditional method using real people would take months and crap load of money (the descriptions have to be customer-readable)
OpenAI’s vision API does it for cents per image. Must have spent under $200 for the whole thing
At least in the 80s, when computers roughly equalled magic for much of the population (looking at you Wargames!), most people didn't really have to interact with it. Their expectations about computers were roughly as important as my expectations about alien life. But I'm afraid that magical thinking about tech will be of greater consequence both individually and societally.
Would it be too much to ask you to start livestreaming any coding of yours that can be shared publicly?
I would love to learn so many things from you, especially around your current ecosystem, that is Python, SQLite (data), and JavaScript.
Because if that is the way you'd solve this problem, then just sending lat/lon to a service to determine if it is in a national park is even easier, as it's just a GET request.
I'm still unsure about what would be harder to set up locally.
Bird is in the default COCO dataset. I haven't look for birds in images, but for people yolov10 is fewer lines of code to detect if there is someone in the frame than it is to setup a flask server for the API calls.
ML:
Grab yolov8
(Optional?) Fine tune it on bird pictures
Convert it to CoreML, do whatever iOS stuff is needed in XCode to run it
GPS: Get https://www.nps.gov/lib/npmap.js/4.0.0/examples/data/national-parks.geojson (TIL this exists! thanks federal govt!)
Stuff it into the app somehow
Get the coordinates from the OS
Use your favorite library for point+polygon intersection to decide if you're in a national park
Bonus: use distance from polygon instead to account for GPS inaccuracy, keeping in mind lat and long have different scales.
...actually the ML one might be easier, nowadays. Now I kind of want to try this.it's like 50ish lines of code in C, just iterating over the points (with the polygon represented by arrays of points). The algorithm is linear with regards to the points.
"task about which you will find more easy-looking tutorials hiding the complexity under a blanket of 3rd party code and services" is better
detecting birds is an exercise in gathering properly labeled training sets, neural networks and their topologies, matrix multiplication performance and/or orchestration of rented GPUs.
both of which cover interesting tasks, worthy things to learn, and are by no means easy.
"easy" is the bird recognition where you do an API call to a totem pole of third party services.
That depends on whether you care about getting the answer right. If you don't, it was always the easier task.
If you do, Seek by iNaturalist still can't do this job, and that's the only thing Seek is supposed to be able to do.
Being close to the border changes nothing, I can just add a buffer outwards the park polygon to account for that. Asking because I'm afraid I may be missing something here due to this being something I already worked on.
"I'll need a research team and 5 years"
In 2020 The BBC had this blog about cameras detecting not just "is it a bird or superman", but what types of birds
https://www.bbc.co.uk/rd/blog/2020-06-springwatch-artificial...
I guess Cueball got the team together.
To him, it was just one button that would open this small info window. Just one button. Just one window.
It took him weeks to understand that we didn't have the data ready he wanted to show. We could do it, but it would take weeks of research and development.
It would have been interesting to see the reverse - the problem becoming trivial in less time then the project’s estimate.
EDIT: A month, https://code.flickr.net/2014/10/20/introducing-flickr-park-o...
It's so weird them explaining 'Deep Networks'. Language on AI has definitely changed in the past ten years.
Also, hilariously, the page they created to demonstrate this (http://parkorbird.flickr.com/) no longer works. Oh, how time flies.
explain XKCD page for good measure: https://www.explainxkcd.com/wiki/index.php/1425:_Tasks
https://nitter.poast.org/jeremyphoward/status/15180380012924...
That’s because the idea of it being a super human intelligence (an undefined metric) is being sold. So you have to lie and say “it’s amazing, it’s going to change everything”. If I tell you “it’s okay and is often wrong” you wouldn’t buy my product would you? This is just to say I can’t blame that on easy/hard task agency, specifically.
=== addendumb ===
“it’s okay and is often wrong” Sounds like working with my junior coworker who I don’t enjoy pairing with. If I said “it’s impressive how the results are to the level of a junior engineer” you sell me on your product.
Flashback to when apple rolled out their enhanced privacy tools and when I said "that data is blocked unless the user gives us permissions" the project manager un-ironically anguished "But how will we track them?!"
Won't someone please think of the commercial interests?! /s
Just pop up a dialog for the user. "Are you in a national park?"
It seems plausible OpenAI's most recent model is better at math and googling than the average human.
Beating the 99th percentile human at any subject should not be difficult when the LLM training is equivalent to living thousands of lifetimes spent reading and nearly memorizing every book ever written on every university subject.
The fact that it only just barely beats humans feels hollow to me.
For those who've seen it, imagine if at end of Groundhog Day everyone in the crowd went, "Wow, he's slightly better than average at piano!"
https://en.wikipedia.org/wiki/AI-assisted_targeting_in_the_G...
The good thing about selling to the government is that it does not care whether the product is snake oil or not. People on the other hand get tired of AI quite quickly.
He was discussing literal spatial indices, scanning the sky for stuff, and how one would break up the problem to make it computable.
So I trolled him and asked if he'd take on the weather.
This mother died six months before ChatGPT's debut, and I thought "sure wish Mom had lived long enough to see LLMs." It wasn't until her first anniversary, flipping through some of our letters, that I realized she had tasted GPT-2 while we played around with ThisWordDoesNotExist.com (made-up words & their "definitions").
Maybe I'll live long enough to see telecommunicators even more capable than these pretrained-transformers.
The most challenging thing was getting the right indexing on the lists since it could be an empty list, a list with a single value, or any number of values.
from ultralytics import YOLO
model = YOLO("yolov10x.pt")
results = model("birds/*", classes=[14])
for result in results:
if 14 in list(result.boxes.cls):
print("Bird!")
else:
print("Not Bird")
I will now take my 5 billion dollar please.Ironically enough the meeting I'm going in to is to explain why I need to use $100,000 worth of hardware for three months before I can answer the business side if what they want to do is possible or not.
I would say probably not.
Unless computer scientists try to understand what biological intelligence really is, AI will not advance.
Ha! Problem? Just ask ChatGPT. Come on, that one was easy…
ps: I am only somewhat joking this time. I think AI agents using just the current state of the art can reliably come up with solutions which the top 1% of humans can come up with. We just need to build them properly. And then deploying swarms at scale, we can scale up 1000x. Scale is all you need at this point. I don’t mean the model parameters, I mean “consensus of expert agents arguing”
pps: if you doubt me, look up the talk about “centaurs” beating a computer and a human alone, after Kasparov was bested by Deep Blue. Most of that talk died down 10 years later (Kasparov still believed it after everyone else stopped). No humans in the loop is going to come about quicker than you can predict.
Which basically is RLHF :)
It wasn't "computer vision is really hard!".
The point was: for laypeople, it's hard to understand why a particular problem is hard or easy for computers. Some things that seem hard are actually easy (people in comments here mention "standing on the shoulders of giants", but I find that's also missing the point) and some things that seem easy are hard. And it's sometimes difficult to explain to laypeople why something can be done in a couple of hours while other tasks would require a 5-year long research project, when to the external observer both tasks are roughly equally complex.
Really, that's it. That's the point of the joke/observation. It wasn't truly an observation about the status of computer vision.
I often tell clients "The first thing you asked me to do was to move a dining room chair into the living room. Then you asked me to do the same with the toilet. The latter only works if we tear out all the plumbing."
Non-coders seem to understand these analogies intuitively.
Just is a bitch mother that seeks to handwave away any potential problems.
If it’s just that, I’ll gladly step aside and let you do it. But if you then tell me you can’t do it, then you better sit down, shut the fuck up, and listen when I tell you something is more complicated than you think.
See also why estimates aren’t actual costs. Every plumber runs into this constantly. Sorry you have a crushed sewer line add a zero to the cost.
“Can you just…” sums it all.
Anyway, a nice idea for generative AI could be to take source code and turn it into a corresponding image of a building so managers can see what they're doing wrong.
No one would ask an architect to add a third basement to an already complete skyscraper, but it happens in software all the time.
Not sure if you are joking or not, but I often hear similar things and I believe that it misses the point. What constitutes a good foundation in software is very subjective - and just saying "foundation bad" does not help a non-technical person understand _why_ it is bad.
It's better to point at that one small rock (some ancient perl-script that no-one longer understands) which holds up the entire thing. Which might be fine until someone needs to move that rock. Or something surrounding it.
I think I'll use your analogy next time, "and then robotics is like gardening. You don't know when the weather will change what wild life will come and try to ruin it, or the soil composition. At least in the house, you can be certain everything is made by human for humans, and things make sense. Outside, not so much. "
Yesterday, I delivered a long-asked for feature - optimistic updates for a UI screen after a button click on Screen #1. It took only one day because the server-side logic is entirely under engineering control and the only edge cases are "their account was concurrently drained" or "the person who clicked the button is trying to hack us". Both of which will be handled when the server responds with a non-successful response. There is little to no likelihood of this behavior changing any time in the next 10 years so I ensured coupling between these components was respected with a simple comment in both places.
Today, you are asking me to provide an optimist update for a button click on Screen #2. However, that button runs business logic that is specified by multiple other users in a scripting language based on a variety of inputs, some of which are dependent on the responses of external systems over which engineering has no control. The response's fields are known in advance, but those fields' values are not.
Of course, anything is _possible_ and if we build a feature where users can specify the likely result of the external systems and build a heuristic-based analyzer for common patterns in the scripting language we could eventually get to the point where simple screens driven by this monstrosity could optimistically update. However, it will take a lot of development work and the testing effort will be high, both for the initial design of the feature and to add sufficient integration / system tests to ensure that future updates to any of these systems do not break common assumptions between them.
An analogy that could more closely fit could be something like: You have a new kubernetes cluster and some servers running code. You're migrating some services over (moving rooms). If you have a simple webapp that has no persistent storage and is containerized, the move would be simple (like moving a chair between rooms). However, if you were to try to move a webapp with login info and databases, there's a lot more "plumbing" not apparent to an outsider that would take a lot more work.
I'll be keeping this on file, thank you.
Yes and no. Sometimes non-coders just want you to throw together a prefab house.
I came up with that metaphor on my own, and didn't know at the time that this is a pretty common metaphor so it holds a special place to me.
Client: My family is having a house built right now.
Programmer: Oh, I'm so sorry for you.
Software has very little bounds to physical world, comparing to actual architecture. Most of the bounds rise from ideas.
Toilet in this analogy cannot be moved, because it was originally decided, that it will be locked and didn’t invest in mobile toilet. Which was reasonable, but highlights lack of vision for the final product.
And this is the biggest difference with architecture. Nobody starts building a house without knowing final design.
While software is the opposite.
Many houses are actually built without knowing the final design, especially in informal settlements in the Global South.
It's referred to as incremental building or incremental urbanism. What starts as simple structure (e.g. a shack) will develop over time into different more formal types of housing. It's an approach to housing that works well with precarious financial means, shifting regulatory environments, uncertain land tenure, changing household size or the lack of building supplies.
Toilet in this analogy cannot be moved, because it was originally decided, that it will be locked and didn’t invest in mobile toilet. Which was reasonable, but highlights lack of vision for the final product.
I don't think the point is that the toilet can't be moved – it's just expensive and disruptive to do so.
Nobody starts building a house without knowing final design.
I would argue the exact opposite – literally _every_ house is built without knowing the final design! Who knows what someone is going to need or want in the future? I'm writing this from a house that was built prior to the existing of indoor plumbing!
There is a TV series in the UK called Grand Designs where people build their own houses. Nearly every cost and time overrun is down to making stuff up on the hoof. The few projects that are on time and budget are the ones that decide everything upfront.
Well, every analogy is inherently wrong at some level of detail. Find an analogy you think is appropriate and zoom in further and it will break.
No analogy, metaphor, or general comparison is ever perfectly isomorphic with the target. As a function of communication, the mark of a good one is if your audience understands.
I must disagree based on the number of residential homes turned into businesses, large scale remodeling, or tearing a house down to rebuild. All these fit well into the analogy.
e.g. "We did not know where to put the piping at the start, so we put it on the outside and now installing a new restroom is sort of tricky."
The house analogy makes the waste understandable, if you accept to compare design errors with late design.
However, in software, you need to continuously work on the product—and it's not just routine maintenance analogous to cleaning the gutters or changing the air filters. In software, it's possible to launch ("move in") before most of the rooms have been built. In software, you can use a library or API and start with a skyscraper on Day 1.
The analogy just doesn't work. It tells clients/stakeholders "this is a tough project but it'll be over someday, and you'll never have to think about construction again."
This is exactly the problem in most software projects.
The GIS took decades to be developped and become functionnal.
The tools allowing to do a GIS lookup also took decades to be developped.
The tools allowing to encode the geodata within a picture file idem.
Not mentioning the development of the necessary hardware to take the pictures.
At the time the comics was being written, GIS lookup had been made easily available since what? 10 years?
And 10 years later, another layer is now easily available - also after decades of collective research and development.
It's not about "the difference between easy and hard challenges in software".
It's about the maturity of a software and its ecosystem, and understanding how even some small incremental changes can have an important impact.