Bonus points: I had never written a browser plugin, but GPT4 helped me do it in under half an hour.
I’m only vaguely familiar with the API.. if I had to guess I would say you send
- a system instruction that its job is to filter unwanted content
- examples of unwanted content
- an instruction like “filter the following html:”
For every web request you want to filter, you would re-send all of those messages followed by the page HTML as the final message. Is that close?
- Examples of unwanted content
- Then I give a large numbered list of comments and ask which numbers should be filtered
- The plugin then just deletes those comment nodes from the DOM. If HN ever updates their HTML I will have to tweak this code.
The reason to send a large list of comments is just to save on costs. It's cheaper to do it this way than one comment at a time.
So the main difference from what you've proposed is GPT never sees the HTML. My code enumerates the comments in the HTML and splices them in to the prompt in a nice numbered list, then does the reverse translation from list number to DOM element in the other direction.
How sure are you that all that code has been reviewed by a 3rd party? How many CVEs a year impact your laptop/desktop?
Do you have any reason to think that increased productivity with LLM assistance will result in lower quality code? Personally I find LLM assistance increases productivity, decreases the penalty of using a more difficult language like rust, and makes it more palatable to spend more (LLM assisted) time writing tests.
I just don't see it chatbot assisted programming any worse than what we have today.
You are hereby removed from the discourse.
/s
Also, unless you are reading every comment in every thread, you are going to miss a few interesting ideas anyway. That's ok too.
I've already used LLMs in my work as a data scientist but it requires a ton of work to just make the results tractable (and I have been using GPT4, which behaves pretty well). These smaller language models ain't so regular. Ok, like consider a basic thing you want to do with a classifier: understand its behavior on a held out data set. Since no one knows what is really in the training data (since its so large), its quite hard to understand what the model can generalize about and what it has just accidentally memorized. CF reports that GPT4 doesn't perform nearly as well on even simple programming exercises that are chosen in such a way as to be sure they weren't in the training data.
There is enormous potential for statistical fuck ups here. Prompt engineering, for instance, is an easy place for over-fitting to happen as a prompt is fine tuned on data the prompt engineer has and thus fails to generalize to new data.
I do think there is a lot of value here, but I'm also sure that sloppy use of large language models is going to cause a bunch of trouble in the short to medium term, generate a lot of garbage, pollute a lot of databases, etc, while we figure all this stuff out.
But I posit that most text classification tasks don't have such strict accuracy requirements. For one, no text classifier is 100% accurate. For instance, I have genuine mail in my spam folder frequently. I see spam on social networks, etc. I struggle to think of cases that aren't at least somewhat tolerant to some amount of incorrect classification.
it will be hard to trust anything at all
as always happens