Unpredictable black boxes are terrible interfaces
magrawala.substack.com
magrawala.substack.com
We’ve had this in the engineering world a long time in the form of Google and StackOverflow searches.
The great engineers i’ve worked with know most stuff they need day-to-day, but when they run into issues it’s their ability to search for and find the right answer that gives them advantage over others who don’t know where to look (or even start).
Getting excellent, exact performance out of deterministic systems is an impressive feat, but autonomy means variability, and getting excellent performance out of variable systems (especially illegible ones) is a different game.
> Simply prompting a model to “think step by step” can lead it to perform well on entire categories of math and reasoning problems that it would otherwise fail on (Kojima et al., 2022). Similarly, even observing that an LLM consistently fails at some task is far from sufficient evidence that no other LLM can do that task (Bowman, 2022).
The eerie thing to me is that this is coaching. I have never coached anything that isn’t alive by any definition. My feelings about this realization are ambivalent.
I’ve never seen code look so correct, sort of work but be so strangely crap.
Complication meaning something slightly beyond changing variable names: you need a different if condition, the snippet doesn't contain the required import, the version has advanced slightly.
Google can do that too. (Well, it could if it wasn't permanently broken from SEO spam and monetization efforts.)
Which is the key takeaway: the "chat" part of ChatGPT is a stupid smoke and mirrors gimmick. The salient useful parts are the large and well-curated training database.
Which is technically quite possible to do without "AI" or text generators; though possibly not realistic from a business sense. There's no business case for "information search without spam and ads".
This is my experience of GPT generated code too. It needs refactoring as soon as it's generated for me to be happy with it.
The real question is whether that actually matters though. If an AI tool is writing ugly code that works, and the 'developer' is only interacting with that code through prompts, then what the code looks like is irrevelant. No one will see it. It'd be like worrying about what the inside of a Photoshop file or a Word doc looks like.
I think we're some way off using AI to write and modify code like it's an opaque file, but once we're there it won't matter if that code is ugly and terrible. No one will look.
Another nice thing with language models is that they condense the information, it doesn't matter how many webpages there are for a product, if you ask ChatGPT about it, it will list it exactly once. It knows it is the same product each time and the list it produces will be created specifically for your query. Meanwhile Google gets cluttered with duplicate information, since your query is implicitly about webpages mentioning a product, not products themselves.
That said, there is still plenty of work to do here. BingChat is completely unusable when it comes to product search, much worse than Google, as it will just pick the first three search results, which tend to be SEO spam, and summarize their content. ChatGPT is much better here, but its lack of direct Web access also handicaps it rather heavily.
In fact one trick is to go to LLM to get the "closed-book" answer, and then use that answer to formulate a proper search query for Google.
Q: What is the height of Everest.
A: The height of Everest is 8800m
Search: "The height of Everest is 8800m"
Result: 8849m
Using the generated answer as search phrase, even if it is factually wrong, works because it has the right structure and good keywords.
I wonder if this skill can become more useful in the age of "prompt engineering"? Will my googling skills be useless and I need to start over? Or will my ability to google naturally convert into a decent prompt-engineering skill given a reasonable amount of practice?
there's two parts, good google query, but also figuring out which of the results returned is the most likely to be useful.
I've often found results for people that were on the first page of the most obvious query in the world to start with, but they were not the first result but down a bit and maybe looked a bit weird.
That's the basic — sometimes only — input to a search query. I suspect the ability to comprehend and describe a problem is a skill that translates quite smoothly to a world of prompt-based interaction.
Yesterday I wanted to write a bash script to do a particular task. "
I never properly learnt bash and it takes me a very long time to do something simple, furthermore, in this case, I didn't actually know the particular tools which were needed for what I wanted to write.
After inputting, it correctly solves the exact problem for me which was perfect.
Working backwards, I then asked it to explain what each part of the script was doing -- and learned the syntax for sed etc, along with gaining a better understanding of what what each command was doing.
The point i'm trying to make here, is that it's amazing to tackle some interactions more naturally. Top down, instead of bottom up.
This can help reduce the paralysis that occurs when working with the unknown. Someone very skilled with photoshop might understand the tool very well and work with particular tools to come up with a result. but Natural language can have you provide the desired end result and work down to fine details.
Did you actually learn it though, enough so you understand it and can recall it? If you did then next time you need a bash script you'll be able to quickly figure out the basics of it, or at least Google for something and modify it with your new knlwledge. I don't think it's much of a stretch to suggest that you won't do that and you'll just use GPT again instead. And next time you might not bother asking it to explain everything...
That's not a criticism of you or of GPT. You have an awesome new tool that magically writes things for you. The most efficient way to use it is to let it do its magic and move on with more important things. All I'm suggesting is that most people, given such a tool, will use it and not learn from it.
If their GPT generated code doesn't work they'll use GPT to refine it until it does. People will learn GPT rather than bash.
An image program I use has sliders for things called "High", "Low ", and "Gamma". I don't know what these mean, but I can immediately see the effects they have when I use them. I can also reverse the effects back to the original image by putting them back.
I don't know if AI could easily reverse back changes, at least over a couple of iterations. But I wonder if things like sliders or knobs could be added to AI interfaces to allow more gradual iteration of a text or image, or idea of a text or image, than having the AI generate an iteration based on a new plain text prompt, or even "conversational interface".
You could have a ton of sliders with names that the user doesn't understand. You could link sliders together so that moving one moves them all, and then unlink individual ones as desired to see what they do. This would provide more immediate and tailorable feedback than trying to get your words just right.
but yes, I'd still welcome more dials. more is a bigger pain to get started but useful after you figure them out
Just going from GPT-3.5 to GPT-4 I'd say intent inference has improved by an order of magnitude. It almost feels mind-reader-ish.
I've been waffling because I'm trying to cope with the absurdity of this product by highlighting the flaws (it can fail badly on tough problems), but we've gone from CleverBot to a genie that can instantly analyze your written intent and produce a highly logically coherent, well-worded analysis in a very short period of time.
I have a similar experience but when I actually look at what’s been returned I find strange things in there.
Like it often works, but it’s rarely as correct as it seems. At least for a given use case.
I ended up having to point it out directly the next day. It was a spurious -A NUM argument that would, in a production scenario, cause it to recognize valid data as invalid because the full context of the data would be missing.
All this to say, it is a very useful tool, but we should be highly suspicious of its output, and really make sure we understand what it's giving us. The ability to read and mentally parse and execute code will become extremely important, and give people with that ability a leg up.
EDIT: Well I'll be damned.
"However, there is a bug in the script as it only checks for the tags attribute in the first 10 lines of the resource block using the grep -A10 command. "
Sometimes, it can one-shot the problem.
Gonna go with Carmack and say that "hand-coding" isn't going to last.
Another thing it can do is reliably debug a file, or point me in the right direction, if I explain the bug and provide the code.
It's scary as hell because I can reliably assume that it will always understand the nuance to the content I provide.
Scarily, it'll probably become more scarce at the same time. It's a lot easier to parse code in your head when you're writing it every day from the ground up. Even a couple weeks away from your own code makes it harder to follow, no matter how well commented it is. Working with other people's code is always an order of magnitude harder, but there's an unspoken trust that they at least put their mind to it and tested it against the tasks it was supposed to perform.
Reading mostly-AI-generated code to spot bugs could be a specialty indeed, but it's one perfectly adapted to black hat hackers.
: No.
//damn you, explain why I wrote this code.
: Who knows, meat dummy?
Signed, yourself from the future.
I go back and look at shit I wrote ten or fifteen years ago and feel like it was written by a super intelligent alien. The again, I feel a weak version of that when I look at code I wrote last month.
So yeah, commenting
Extracting intent from one piece of text when a number of interpretations exist that I cannot count is not.
The instant feedback is better than with humans. Like humans, these models can take one thing to mean another, but because we are still using them for tasks that require instant feedback, we can catch the errors early and correct. It becomes worse if the are to be used on longer lasting tasks. The entire project could be derailed by one misunderstood concept. Just like a human project.
Natural language is a fair interface. But with lots of ambiguity. When I say 'clear', do you mean understandable, easily recognisable or clear like water. I faced that dilemma when I used that word as a design prompt because it can mean things which are opposites of each other. Some languages are so rich they avoid such ambiguities.
A common language can be good. It should not be designed by a committee, it should evolve with these tools naturally.
Stable diffusion (and to a lesser extent midjourney) exposes a LOT more knobs to tune. Layer on controlnet and other extensions, and you can really iteratively refine an image to get closer and closer to what you want.
The key is to be able to 'lock in' a depth map, rough composition, seed, etc, and then tune the other parameters to refine what you're looking for.
It's definitely not a one-shot solution, but its still miles faster than learning how to do digital painting and manually rendering things. I don't feel like it matters if you have to generate 16 images to find one to start building off of when each image can be rendered in 1/1000th the time a human could.
The challenging part of this is inverting everything and figuring out what all of your classes are and how to arrange training data to fit them. Set theory is another new dragon here. Smaller models like Ada and Babbage are ideal for this, but we are looking at 90's tech too (SVMs). If we need 10k separate models because we have so many classes, then I'd prefer to be able to train them in a few seconds/minutes on a CPU.
Fine tuning gigantic models like Davinci is a dead-end at this point IMO. A fleet of microscopic models working together to feed 1 big powerful (untunable) model feels a lot more manageable.
To answer the hypothetical - I do think handcrafted solutions by professionals are still the best option in a lot of places. Less weird edges and things to worry about. But, a lot of hard things are quickly reduced to routine crap once you have effectively 'done it' a few times.
Existence is wonderfully observable & learnable. It took human kind to create this particularly infernal information void, to have manufactured such a degrading & omnipresent local system of black holes.
These are woefully against all spirit & should be torn down. More positively, there's so many positive pro-human pro-Augment-Intellect things tech could be doing; we need new powerful positive symbols that establish visible legible expansion of understandability.
But can it do it without baby steps?
I fear people will start taking the beautiful and profound for granted, it'll just be content.
But it just feels like humanity secedes more and more ground & understanding to humanity each month, and there's so few visible notable frontiers where humanity is staking down real wins. Rather than have advances in general purpose computing, we keep getting more and more apps, more and more far-off cloud systems.
> but there's nothing a machine puts into this
What makes you so sure?
What makes you think a machine has any wants and desires?
That depends on your prompt. You can make it longer if you have more to say. It's a conditional language model, you can select the place it starts from and how to go. But you can't blame it for bad prompts.
Instead of seeing LLM as "nothing there, just token probabilities", I'd rather see it as a distillation of human culture. It's like a mirror house with infinite reflexions and complexities. A place to contemplate, a microscope where you can study something in detail, a simulator where you can deploy experiments.
We know that users are increasingly likely to leave a page as load time increases. Now what happens when you have an unpredictable black box that sometimes doesn't do what the user expects, or flat out refuses to cooperate? This is poor UX for many tasks.
ChatGPT gave me "ffmpeg -i input.mp4 -vf "vflip" output.mp4" and then explained what each part of the command does. I searched on google and the results were all for random tools when I specified ffmpeg, and the ones that were correct had complex walls of text, adverts everywhere, and other bloat.
Knowing more precisely that rotation isn't what I want, googling for "ffmpeg vertical flip video" sends me to [1] in the very first result, which is the same as what you got. Except it also specifies copying the audio without reencoding it, I'm not sure what your version would do without including that.
[0] https://stackoverflow.com/questions/3937387/rotating-videos-...
[1] https://filme.imyfone.com/video-edit-tutorials/ffmpeg-flip-v...
The ChatGPT one just leads with the answer and explains it after. It's a non destructive command so if I later decide I also want audio, I can just continue the conversation and ask for audio to be included.
Also reminds me of Monty Python's Stay Here and Make Sure He Doesn't Leave https://youtube.com/watch?v=2f5MvVx8RM8
ChatGPT almost always understands what I want. Even if it doesn't, it's easy add another prompt to put it on the right track. It's shockingly good at that. The answers it provides might not always compile, but neither would most human written code when written without access to a compiler for testing. And getting facts wrong is quite understandable as well without access to outside references. I also find that all the stuff ChatGPT fails at the most, are simply things that don't have any easy answers. It's not magic, it knows what it was feed and little beyond that.
Image generation meanwhile uses a far more primitive prompting style. There are no follow questions you can ask to correct the result. And that one prompt you get is not natural language to begin with, it's just keywords used for photobashing, one that is dramatically underspecified no less. If you tell it to make a picture of New York, well, New York is big, you didn't tell where to put the camera. So of course the results will be pretty random. Worse yet, they'll change dramatically with minimum prompt changes, the moment you add a person to a photo, that person will become the primary subject and you'll have a very hard time making it just a background character by prompt. There are just a lot of image features that you can't accurately describe with a few keyword, or even if you could come up with a fitting keyword, as long as that keyword isn't in training set it's not going to help.
But that's really nothing more than a current day limitation. The language understanding will get better and so will the integration ways to manipulate the image, we already have inpainting to change specific parts of the image, outpainting to change the framing, people have been playing around with projecting the images onto depthmaps or 3D geometry.
In the end what matters is not that the AI is a blackbox, the inner workings are largely irrelevant, but that the output of the AI fits into my personal predictions of what it should be. And as far as ChatGPT is concerned, it's already better than any other user interface I ever touched.
There’s so much “hey look at this from AI” and I look at what I get and it’s… not good/ frustrating.
And it’s true that humans are not necessarily better.
But that’s not entirely true. My closest friends know me. If I say “put up some turnstile,” they know I’m probably talking about the band because I like the band.
Taking to these AIs and voice assistants is like talking to an anonymous person I’ve just met for the first time, every time, so you always need extra qualifiers.
Like, you know, managing humans?
So I'd rather buy wood and nails to build the hut I like than to buy some hut and hammer it into the hut I wanted.
The first is a really fulfilling task, the last annoying. And so I rather drive my car to my destination than let it drive itself and have to control if it doesn't make a mistake.
And this, ladies and gentlemen, is the reasoning for working on your prompt engineering.
The prompt is the interface. Prompting is a communication skill, a management skill, and a technical skill.
I doubt many will be hired as a "prompt engineers" in coming years (some already have been, by idiotic corps) - but being a bad prompt engineer will be a noticeable lack.
Start practicing.
If not, what is the specific thing people should be practicing right now?
Sounds like you're looking for a function "I : a -> I". Precision really seems important. Programming in English is like being granted wishes from a djinn. You'll get what you asked for, but beware. But currently you don't even get that.
If it's useful now it's not idiotic for a company to hire for people good at solving their problems.
It's not :) - they're hype-hiring.