Malleable software in the age of LLMs (2023)
geoffreylitt.com
geoffreylitt.com
I wish, I have tried. LLM's don't understand the DOM. Unless it is as simple as the address being in element id=address, an LLM is useless for generating scraping code. You are better off with an element picker and some heuristics to generate a selector / xpath query. Now there are some specialized models, I found a paper that takes the DOM tree and fits a vector to each node, but I think they are too much effort for little gain, unless someone integrates them into an open source scraping library so they are easy to use.
Bummer. I wanted to try my hand at this. There has to be some trick where you can combine LLM and some element picker to get a really robust solution right?
I even got it usefully updated when the format changed!
Even if/when LLMs improve enough to bring their code quality up to par, humans should still be comfortable reading that code to ensure it aligns with the goal.
This is the one area of AI I’m actually pretty positive about. Computing largely passed people by, and computers have not lived up to their potential. I could see this approach making computers more useful for a large number of people.
Especially semi-technical people who’s specialty isn’t programming.
That would be worse kind . Imaging fixing bug of those people with LLM Hallucinated code that runs but having wrong logic all over the place.
I reduced use of LLM for coding these days. I only ask it to generate templates or write some repetitive codes and sampledata.
What i love most is to draft out project specs from user's requirements and to generate user stories.
Or let it write leet-code like problems that is too boring to do manually.
The actual layer of coding is still best done by experienced engineer's biological brains.
And of course with iteration and feedback loops people can definitely learn how to specify what they want in a fairly precise way.
There was a while there where it seemed like conversational (either text or by voice-to-text) interfaces were the only way people could imagine using LLMs. Everything being an empty text box staring back at you.
It seems like that may have just been due to ease of implementation to an API. Now that we’ve all had some time to work a bit, some interesting UI experiments are starting to peek out.
Malleable Software in the Age of LLMs (2023) - https://news.ycombinator.com/item?id=40188435 - April 2024 (38 comments)
Tools will no longer need operators as adding multi modality to tools will ensure we can actually have the operator interface on top of the tool.
Tuning the operator is where the playing field shifts to as operators which can translate fuzzy inputs to creative specs or workflows which can in turn be executed by the downstream engine.
> Jensen Huang CEO of NVIDIA said: “Every single pixel will be generated soon. Not rendered: generated” (https://x.com/icreatelife/status/1639363377255309328?lang=en)
If we take this to the extreme, user interfaces are going to be generated on the fly, specific for the user. They will adapt to the user based on the device they are using, time of day, what data is being displayed, what the user prefers, etc.
These large models generating views might be streamed from servers (ala stadia), it might pass off some of the work to edge devices.
The models will be able to store things and communicate with other models as needed. Models might spin up that perform certain things well and have access to specific resources.
Seriously, think harder about what you're suggesting here. It's ridiculous.
If it does come to that, whether or not there is code will be the least of our worries.
This concept obviously flies well over the heads of the LLM/AI/AGI bros who think that our creative lives is going to be rendered obsolete by vegetable silicon. They lack creativity and imagination.