GPT-4 generates simple app from Whiteboard photo
twitter.com
twitter.com
The average person could be using computing to solve all kinds of problems but it requires expertise and a time investment that is just not a good value proposition for a lot of people. If GPT or some other thing comes around and makes it so my uncle could say to his computer "hey build me a site so I can show off all my fishing trophies" and it just does it? the world would have a lot more niche software, more weird stuff and overall everyone would be more productive and happy.(or so I hope)
Isn't that itself the province of an LLM? Say I have a bunch of text. How do I save the text search "by similarity"? Sphinx and semantic search was hard, I remember. Facebook had Faiss. And here we are supposed to just save vectors on commodity hardware BEFORE using an LLM?
It is! The steps are:
1. Take a bunch of text, run it through an LLM in embedding mode. The LLM turns the text into a vector. If the text is longer than the LLM context window, chunk it.
2. Store the vector in a vector DB.
3. Use the LLM to generate a vector of your question.
4. Query the vectordb for all similar vectors (that fit in the context window)
5. Get the text from all those vectors. Concatenate the text with the question from step 3.
6. Step 5 is your prompt. The LLM can now answer your question with a collection of similar/relevant text already provided to the LLM in the context window along with your question.
So no change from today?
Can the same be said for facts? Hard to generalize. Like lets say I asked GPTV the year a painted was painted. Way back when, I was right, I will never need to know the exact year of a painting - whatever that means - after I finished my modern art history exam. Does GPTV? No. I am happy with ranges in my answers about general knowledge, because at the end of the day, there isn't some error-intolerant process using those outputs anyway. Even order of magnitude errors may be fine, getting the zero wrong in the token sometimes just doesn't matter. It always matters when programming though.
Writers increasingly have the same opinion. As more people use GPT4 and learn how stilted and uncreative it is, and that a one token error can make a plot stop making sense, it isn't clear how it's going to deliver on bulk writing. Spam and content farming sure, but even short prose is not in the realm of reality for actual, real breathing human consumption.
How do you categorize this? I don't know, I'm not in the field. There's some huge difference between one token error means 0% correct, versus one token error means 99% correct. Maybe transformers can solve this. Maybe they never will.
Even "logic" errors can be generated in a go and detected by another instance. Like a "this is what i wanted but you did this" works sometimes too.
...yet.
But besides that, the LLM is less likely to demand unscheduled time off - especially as a fraction of the hours it can put in. If I have a family emergency once per year, and I need my eight hour day off, I've just removed 1/200 of my yearly output potential. The LLM would need to be down for over 400 hours per year to get to that type of output reduction. Realistically, that is unlikely to happen.
So does my code though. You can take this code, put it into xcode and then ask GPT about any errors. Reiterate and repeat. I did the same while building some Chrome plugins.
Look, seriously, "it hasn't crashed lately" is the standard by which 99.9% of human-written software is judged.
Potentially you could do the same for fact-related questions, by fact checking the result against the (properly indexed) training set and feeding the results back.
I wonder if it would work.
This was also true in the early days of cars, planes and—for that matter—computers. Most coding, even today, doesn’t require a full-fledged engineer.
It’s my understanding that this feature is still rolling out to everyone, so if there are some use cases beyond “simple” - I am sure we’ll hear about it in a more detailed article / blog post.
Not taking anything away, just pointing out the facts.
https://twitter.com/skirano/status/1706823089487491469
https://twitter.com/GabGarrett/status/1706872805214593173
are much more impressive demonstrations. It shows the model understanding ui and figma elements to a non trivial degree.
There is a submission on the front page,
https://news.ycombinator.com/item?id=37673409
But it isn’t about code, so I am sure something like that is coming.
ChatGPT will be next year’s biggest excuse for bugs in products.
And it undoubtedly reduces the number of developers needed.
It doesn't replace all human coders yet. But a senior, super skilled engineer can now effectively do the task of a team of 10. This will be true for many teams that I have personally seen.
There will be exceptions, too. Not all teams work like that.
I can confidently say that LLMs will cause a reduction in number of developers required at least in the short term.
Factors that cause an increase in demand might propagate simultaneously in a greater rate and undo this but that is unlikely in my opinion.
Everyone loves to say this and yet I’ve never seen a single concrete example aside from toy demos.
Here’s the thing: output is pretty easy to see. If one developer can 10x their output, that means they’re doing in a month what would take them nearly a year without AI.
So, show me one example of an individual or team who has done a year’s worth of work in a month.
Is it the act of writing code or is it creating things? I often wonder this when the threads discussing raw typing speed come up, or the "extra typing" involved in optional curly braces. Even today, without LLMs, only 20% or less of my work is writing actual code. Requirements/research, designing, optimizing, testing, etc make up the lion's share.
So you take out that 20%, I'll be more productive, but my work doesn't change.
So yeah, bleak.
Seems inefficient. Shouldn't the AI just be generating machine code for the computer to run?
Its much easier in the short term to generate larger snippets of Python code based on prompts and let the programmer fine tune it
Computer: I understand. OpenAI requires you to provide a $10,000 advance on this task, and $10,000 more when it's ready. Do you agree?
User: Compared to the money I will make, this is peanuts. I agree.
Computer: Alright, your account was charged. I am also legally required to notify Amazon Hive of every group of users that try to clone their store each day. Don't worry, their legal subAI only sues 0.3% of the time. Have a good afternoon!
Tons of indications how ai will/is able to learn flows and architecture and code etc.