For the past 8-10 years it has all felt like a bunch of apps that just aim to be mediocre middlemen/gig economy brokers with bad customer service.
For the past 8-10 years it has all felt like a bunch of apps that just aim to be mediocre middlemen/gig economy brokers with bad customer service.
And I'd wager that there are silent revolutions happening all across colossus that's the tech industry that will become apparent in the next decade.
Jeff Bezos put it best during his recent interview at the 2024 NYTimes Dealbook Summit, "We're living in multiple golden ages at the same time." There's never been a better time to be alive.
It can sometimes be useful to input a more "human" search and have something get spit out but 60% of the time it completely lies to you. I'm talking about questions related to web specifications which are public documents. Section numbers, standards names, etc.. will be completely made up.
This is such an exhausting conversation
I had the same impression about the hallucinations 2 years ago. The reality is in at the end of 2024, you can get incredible value from LLMs.
I've used copilot to code almost exclusively now for the past few months. Anyone still comparing it to text completion I feel is operating on completely out of date information either intentionally or unintentionally.
My general rubric is: “would I trust someone on Reddit to correctly guide me on this”. If the answer is “yes” then ChatGPT is likely going to do well. If the volume on a particular subject is low / susceptible to false information then it’ll lie.
Recently it lied hard about how to configure MikroTik routers. I lost many hours. But for a large construction project recently it completely balled out.
Are you doing cutting edge / complicated stuff? Have you examples of where it lies?
ChatGPT lies a lot about RouterOS, I don't know why. Claude helped me a lot on the other hand with all things MikroTik.
No specific prompts, but most were related to the XHR/Fetch specs and behaviors within. It would say "X.Y.Z sections defines this" but that section didn't exist at all and the answer provided was not accurate.
> My general rubric is: “would I trust someone on Reddit to correctly guide me on this”. If the answer is “yes” then ChatGPT is likely going to do well
I see. Well, I don't know if I find that very valuable but if others do, then so be it.
- completely made up books
- real books that were only marginally related
- real books with really bad reviews
I'd estimate that only 30-40% of the time did I find the results at all useful.it's just that every thread about LLMs in AI invariably has someone complaining about best results from a query best described as `SELECT * FROM...`
They can make stuff up, but saying "60% of the time they lie to you" hasn't been true for years.
If you're using them to fill knowledge gaps, what scaffolding have you set up to ensure that those gaps aren't being filled with incorrect-but-plausible-sounding information?
So much this. So many times I've argued with hired experts saying "can't be done" just to see yes, it can be done.
I'm glad ChatGPT didn't lead you astray, but I'm not seeing what it's added here besides shuffling up the user interface in a way that you presently and subjectively prefer?
Time is my most precious thing, I already don't have enough time to do all the things that I want to do, I don't want to waste that trying to find and test solutions when ChatGPT gives me instant answers. I'd rather spend time playing with my cats or riding a bike instead. It's not a matter of UI, it's a matter of preventing waste of time, energy and money, and less frustration. For that alone, €20/month is a very good value. And that's just for my personal life.
This. But in the same sense the past 50 years merely changed interface from dusty textbooks in libraries to Google Search, and the past 100 years gave us dusty textbooks over writing to Royal Society, and that just replaced the option of asking a local whisperer or hoping you'll find answers on the Sunday mass.
Do not underestimate the power of being able to get an answer to your problem described, visualized, and perhaps complete with interactive demo to explore it further, in time it would previously take you to formulate the right search query that finally gives you relevant information.
EDIT:
And that's on top of all the arbitrary data transformations prior tools couldn't do. E.g. I'm increasingly often using GPT and Claude models to turn photos of (possibly hand-written) notes or posters into iCAL files I can immediately import into our family shared calendar.
Another frequent use case, data normalization. Paste a whole dump of inconsistently structured data multiple people collected (say, addresses of various local businesses that helped a local NGO and now are supposed to get a thank-you card for Christmas). Like, you get 200 rows of addresses in a single column, with spelling mistakes, repetitions, junk at the end, arbitrary capitalization, wrong order of address segments, and such; you need to separate it out into 5+ columns (name line 1, name line 2, street address, zip code, city, etc.) and have it all normalized.
The fastest and most robust way to do it as a one-off job, today, is to paste the whole thing to GPT-4o or Claude 3.5 Sonnet, tell it how the output should look (give one-two examples, mention some mistakes you saw), then send the message and wait 30 seconds for the job to be done for you.
(Yes, it may make mistakes - it didn't for me in recent memory, but it can. But for that, I quickly add an extra verification column for each one in LLM output, and do a simple case-insensitive substring match with original, and eyeball any data row that shows an error. And guess what, the formulas don't take much time either, since LLMs are good at writing them for you, too!)
I wouldn't discount this effect. As someone with sensory issues, one thing I like about ChatGPT as opposed to the "raw" internet is that I can see the answer to my questions in a nice and calm textual format without some website who created the article specifically to catch my search terms, but is trying to get me to deceptively click on ads or pull me into buying something through their affiliate links. That's absolutely increased my own enjoyment and productivity.
In the past week I have used it for helping write a script in a framework I'm not super familiar with (OpenSCAD), I was able to finish a project in 5 minutes that otherwise would have taken me hours. I have used it to help make movie recommendations (none of them were hallucinated). I have used it to translate a conversation with a non-english speaker, etc. There are other tools that can help me do all of these things, but none quite as fast or painlessly.
It might not be useful for your use case of asking questions related to specific web specs, but that doesn't mean that the technology has no value. Horses for courses...
Imaging being graded on your ability to quote exact line numbers of particular parts of your codebase as a senior software engineer without being able to look at it!
LLMs are not, in isolation, a search product.
The hard part is, despite actually having some "real" value delivered, you still have to sort through the 99% of bullshit that comes along with it anyways.
I'm also going to stand up for AR/VR here. I'm in a long-distance relationship and me and my partner spend an hour or so in VRChat around two to three times a week. The power that has to reduce the badness of an LDR is well well well well worth the three hundred bucks I paid for a Quest. That and some of the golf games on it are fun.
I've had an HTC Vive and an Oculus Rift 3 (Walkabout Mini Golf is one I tried!) and while I wouldn't try to argue NOBODY has found a use for it (somebody somewhere found uses for all of the things I mentioned, just not me and just not the majority of people like big new things are promised to) it never really ticked the "new value" box before they ended up in the closet for me.
That and the ergonomics do still suck, even if I've mostly gotten used to them.
I do think VR will make it, though - starting with the kids. Apparently Gorilla Tag broke 1.5 million players recently, and those are mostly under-15s. The next generation is going to have a strange relationship with computers.
Lots of engineering involved
Isn't this the new LLM playbook?
I pay Claude/ChatGPT trivial amounts of money for metered API access to their models, and they in turn provide it to me.
Middlemen/marketplace models like "Uber for x" or "Etsy for x" or "Betterhelp for x" is a totally different business model.
Yes.
> and adding value.
No. The only breakthrough innovation LLMs gave us is the ability to speedrun the making of racist pictures. Not sure the world really benefited.