Not sure why no discussion at all - maybe the design is underwhelming.
Not sure why no discussion at all - maybe the design is underwhelming.
And while tool use is being worked on, the results I've seen are are the "that's an interesting tech demo" level rather than the mind-blowing change when InstructGPT demonstrated the ability for a language model to generate any meaningful code at all from natural language instruction.
that's what rag is supposed to solve: they chunk your 60M loc, and then retrieve and process only relevant depending on your inquery.
I use copilot every day, but it's only so good.
The LLM hype feels like it's been driven by FOMO.
Ever saw someone really bad at googling? It's the exact same thing with LLMs (for now). They're not magic crystal balls, and they certainly can't read everyone's minds at the same time. But give them a bunch of context, and they'll surprise you.
Other than the automation aspect, it is a pretty good alternative to in-depth googling.
BTW, why the disparaging reference to "little toy apps"?
It's an unmaintainable, single use piece of software (that doesn't even implement the features, it just glues together already existing code) that any CS student could write in a week. Congrats on getting a really fast CS student I guess ? Not to mention the fact that perfectly viable, better alternatives are available in many places.
It's like me nailing two 2x4s together to make a shelf. Yeah, sure, I made it myself and I didn't need any woodworking knowledge, but let's just hope I don't put grandma's heavy china on it.
The upside is that I can produce a shit-ton of one-shot code in record time, so I've got time to face the downside.
search_documents "search term" --bm25 --table-prefix some_project | fulltext -hl -v | less
search_documents "search term" --bm25 --table-prefix some_project | metadata
It inserts documents by piping paths into another script, eg, find /some/path/*.pdf | insert_documents --table-prefix some_project
The documents end up in a Postgres database with pg_search bm25, tsvector, and semantic embeddings (from a local model).I would estimate that I only wrote 5% of the code in the project with the rest coming from the LLM.
Sure, it's just a few hundred lines of code but it's been stable and helpful to get through some very large tranches of discovery material.
As a programmer (which is the requisite to build such tools even with LLMs), I have a plethora of tools to do the tasks, what I choose and how much time I invested in in that depends on something similar to this chart, but with an added dimension: interest.
Take for example the URL extraction. For one single occasion, I'd probably use VIM and macros to quickly do it. If it were many pages, I'd write a script. If it were infrequent, but recurrent, I'd take the time to write a better script and would only write a web page if the use case was shared with other people or if I wanted a cross platform solution.
I believe the first question one should ask before building is why. That leads you to find a better UX than shoehorning everything inside a web app.
In that aspect, I am hopeful. Maybe if "waste of time" activities are commoditized, "professionals" can instead focus on "what is important," whatever that might be.
What LLMs promise is endless drag. I try to structure my work to ensure that the final velocity is high.
This app for example - which runs OCR against PDF files entirely in the browser - was assembled by pasting in an example of PDF.js usage and an example of Tesseract.js usage and having it figure out the rest: https://simonwillison.net/2024/Mar/30/ocr-pdfs-images/
The kind of project I work on is more like this: Build an Android app for a quiz game. The quiz takes a list of random question from a set. Each set is a package that can be installed and upgraded when online. While the app is free, there is an activation code to be able to download the main packages. The app should work offline except for the activation and downloading packages. It also should notify when a new version is ready for a package. etc...
I don't know if LLMs could have helped me at the time (pre 2020), but I doubt it. Not because the code was complex, but mostly how cohesive the whole thing should be while taking care they're not tightly coupled and be maintainable by a single person. The IDE was a great helper once I got the design and the architecture outlined, mostly because it was deterministic and I already know what the end result should be.
The best current LLMs GPT-4o, Gemini 1.5 Pro, Claude 3.5 Sonnet - are just about at the point now where I'd expect them to be able to get a useful chunk of your Android spec there done. Which is pretty wild!
Perhaps how quickly we become jaded should be taken as evidence of how quickly the world is changing right now.
When I looked at the examples they seemed like the kind of one off scripts, of limited complexity, that we’ve seen many times in the last year or so.