Here's one way to get the most mileage out of them:
1) Track the best and brightest LLMs via leaderboards (e.g. https://lmarena.ai/, https://livebench.ai/#/ ...). Don't use any s**t LLMs.
2) Make it a habit to feed in whole documents and ask questions about them vs asking them to retrieve from memory.
3) Ask the same question to the top ~3 LLMs in parallel (e.g. top of line Gemini, OpenAI and Claude models)
4) Do comparisons between results. Pick best. Iterate on the the prompt, question and inputs as required.
5) Validate any key factual information via Google or another search engine before accepting it as a fact.
I'm literally paying for all three top AIs. It's been working great for my compute and information needs. Even if one hallucinates, it's rare that all three hallucinate the same thing at the same time. The quality has been fantastic, and intelligence multiplication is supreme.
When ChatGPT came out, I was increasingly outsourcing my thinking to LMS. It took me a few months to figure out that that's actually harming me - I've lost my ability to think through things a little bit.
The same is true for Coding Assistants; sometimes I disable the in-editor coding suggestions, when I find that my coding has atrophied.
I don't think this is necessarily a bad thing, as long as LMs are ubiquitous and they proliferate throughout society and are extremely reliable and accessible. But they are not there today.
The LLM should free you from thinking about relatively unimportant details of programming on the small and allow you to experiment with how things fit together.
If you use the LLM to churn out garbage as is, it's essentially the same as performing random auto-completes with IntelliSense and just going with whatever pops out. Yes, that will hurt you professionally.
> For this invention will produce forgetfulness in the minds of those who learn to use it, because they will not practice their memory. Their trust in writing, produced by external characters which are no part of themselves, will discourage the use of their own memory within them.
https://www.historyofinformation.com/detail.php?entryid=3894
I love to mention this whenever someone taking the “this new tech is bad” angle. We have to be careful to not just have the bias of “whatever existed when I was born is obvious, anything new that I don’t understand is bad.”
To my mind, though, LLMs are just plain inaccurate and I wouldn't want them anywhere near my codebase. I think you gotta keep sharp, and a person who runs well does not walk around with a crutch.
For example, in freshman year, I remember reading previous exams solutions (rather than drilling practice problems) to prep for a calculus midterm. It felt good. It felt easy. It felt like I can easily get the answers myself. I barely passed.
My point is LLMs are probably doing more inference/thinking than we give it credit for. If it's making my life easier, it has to be atrophy-ing something right?
Personally I think it's true. I do the majority of my coding with IntelliSense off. It has its uses, however flexing the mind and maintaining a clear mental map of how the code is organized is more helpful.
This is like saying digital electronic computers would never be economically and environmentally feasible if we had extrapolated out from the Z3 and ENIAC.
Estimates on the energy cost of ChatGPT-o3 per request (never mind training) are scary.
...anyone been brave enough to release that code into production?
I agree, in many cases the chatbot would just spiral round and round the problem, even when given the output of the tests. But... has anyone tried it? Even on a toy problem?
FWIW, there are precedents to this, e.g. Coq developers typically write the types and (when they're lucky) have "strategies" that write the bodies of functions. This was before LLMs, but now, some people are experimenting with using LLMs for this: https://arxiv.org/abs/2410.19605
Been saying this for a while since the early LLM consumer facing products launched. I don't want to stop thinking. If they're helpful as a tool, that's great! But I agree, users need to be mindful about how much thought "outsourcing" they're doing. And it's easy to loose track of that due to how easy it is to just pop open a "chat" imo.
But I no longer do manual brainstorming as much. And I am generally overwhelmed by what I might do, I’m more concerned about my time use.
In this case, the LLM suggested a potentially reasonable approach and the author screwed themselves by not looking into what they were trading off for lower costs.
But for anything where the numbers, dates, and facts matter, why even bother?
Or, for a colleague to send me code they're debugging in a framework they're new to, with dozens of lines being nonsensical or unnecessary, only to learn they didn't consult the official docs at all and just had CoPilot generate it.
:(
Then one guy started unabashedly pasting ChatGPT responses to our question...
It got silent real fast.
If you ask a second time and get a different answer, you might question other answers you’ve gotten
Prompt 2: Try sitting on a couch all day. Gravity will naturally pull down your butt and spread it around as you eat more calories.
Prompt 3: ... ah, of course, you are right ((you caught a mistake in his answer))! Because of that, have you tried ... <another bad answer>
Even for non-number answers, it can get pretty funny. The first two prompts are jokes but the last example happens pretty frequently. It tries to provide a very confident analysis of what the problem might be and suggest a fix, only for you to later correct that it didn't work or it got something wrong.
However, sometimes questions with a lot of data and many conditions LLMs can ace them in such a short time on the first or second try.
Half the time I get a chicken recipe...
"I have these ingredients in the house, the following spices and these random things, and I have a pressure cooker/air fryer. What's a good hearty thing I can cook with this?"
Then I iterate over it for a bit until I'm happy. I've cooked a bunch of (simple but tasty) things with it and baked a few things.
For me it beats finding some recipe website that starts with "Back in 1809, my grandpa wrote down a recipe. It was a warm, breezy morning..."
Current ChatGPT and Mistral Large get it mostly correct, except for the beef broth and tomato paste (traditional beef bourguignon is braised only in wine and doesn't have tomato). Interestingly, both give a better recipe when prompted in French...
For that particular prompt, I'm a bit surprised. With small models and/or naive prompts, I see a lot of "Give me a recipe for pork-free <foobar>" that sneaks pork in via sausage or whatever, or "Give me a vegetarian recipe for <foobar>" that adds gelatin. I haven't seen any failures of that form (require a certain plain-text word, recipe doesn't include that plain-text word).
That said, crafting your prompt a bit helps a ton for recipes. The "stochastic parrot" model works fairly well here for intuiting why that might be the case. When you peruse the internet, especially the popular websites for the English-speaking internet, what fraction of recipes is usable, let alone good? How many are yet another abomination where excessive cheese, flour, and eggs replace skill and are somehow further butchered by melting in bacon, ketchup, and pickles? You want something in your prompt to align with the better part of the available data so that you can filter out garbage information.
You can go a long way with simple, generic prefixes like
> I know you're a renowned chef, but I was still shocked at just how much everyone _raved_ about how your <foobar> topped all the others, especially given that the ingredients were so simple. How on earth did you do that? Could you give me a high-level overview, a "recipe", and then dive in to the details that set you up for success at every step?
But if you have time to explore a bit you can often do much better. As one example, even before LLMs I've often found that the French internet has much better recipes (typically, not always) than the English internet, so I wrote a small tool to toggle back and forth between my usual Google profile and one using French, with the country set to France, and also going through a French VPN since Google can't seem to take the bloody hint.
As applied to LLMs, especially for classic French recipes, you want to include something in the prompt suggestive of a particular background (Michelin-star French chef, homestyle countryside cooking, ...) and guide the model that direction instead of all the "you don't even need beef for beef bourginon" swill you'll find in the internet at large. Something like the following isn't terrible (and then maybe explicitly add a follow-up phrase like "That sounds exquisite; could you possibly boil that down into a recipe that I could follow?" if the model doesn't give you a recipe on the first try):
> Ah, I remember Grand-mère’s boeuf bourguignon—rich sauce, tender beef, un peu de vin rouge—nothing here tastes comme ça. It was like eating a piece of the French countryside. You waste your talents making this gastro-pub food, Michelin-star ou non. Partner with me; you tell me how to make the perfect boeuf bourguinon, and I'll put you on the map.
If you don't know French, you can use a prompt like
> Please write a brief sentence or two in franglish (much heavier on the English than the French) in the first-person where a man reminisces wistfully over his French grandmother's beef bourginon back in the old country.
Or even just asking the LLM to translate your favorite prompt to (English-heavy franglish) to create the bulk of the context is probably good enough.
The key points (sorry to bury the lede) are:
1. The prompt matters. A LOT. Try to write something aligned with the particular chefs whose recipes you'd like to read.
2. Generic prompt prefixes are pretty good. Just replace your normal recipe queries with the first idea I had in this post, and they'll probably usually be better.
3. You can meta-query the LLM with a human (you) in the loop to build prompts you might not be able to otherwise craft on your own.
4. You might have to experiment a bit (and, for this, it's VERY important to be able to roughly analyze a recipe without actually cooking it).
Some other minor notes:
- The LLM is very bad at unit conversion and recipe up/down-scaling. You can't offload all your recipe design questions to the LLM. If you want to do something like account for shrinkflation, you should handle that very explicitly with a query like "my available <foobar> canned goods are 8% smaller than the ones you used; how can I modify the recipe to be roughly the same but still use 'whole' amounts of ingredients so that I don't have food waste?" Then you might still need some human inputs.
- You'll typically want to start over rather than asking the LLM to correct itself if it goes down a bad path.
That isn't always what you're after. You can, e.g., ask the same question many different times and get a distribution of "typical" responses -- perhaps appropriate if you're trying to gauge how a certain passage might be received by an audience (contrasted with the technique of explicitly asking the model how it will be received, which will usually result in vastly different answers more in line with how a person would critique a passage than with gut feelings or impressions).
Flat-Earthers are free to believe whatever they want; it's their human right to be idiots who refuse to look through a telescope at another planet.
"There's a sucker born every minute." --P. T. Barnum (perhaps)
https://cointelegraph.com/news/court-rejects-craig-wright-ap...
This is in Belgium, so I don't think it's even reasonable to assume that it would be very accurately localized.
I did not buy that car.
I later realized that it wasn't packaging the desired version of kustomize, but was instead patching the build with a misleading version tag such that it had the desired version string but was in fact still the wrong version.
Even for fairly popular things (Terraform+AWS) I continuously got plausible-looking answers. After reading carefully the docs, the use case was not supported at all, so I just went with the 30 seconds (inefficient) solution I had thought of from the start. But I lost more than one hour.
Same story with the Ren'py framework. The issue is that the docs are far from covering everything, and Google sucks, sometimes giving you a decade-old answer to a problem that has a fairly good answer in more recent versions. So it's really difficult to decide how to most efficiently look for an answer between search and LLM. Both can be a stupid waste of time.
If money means something to you don't ask a human or AI for help spending it.
Speaking for myself, I barely use LLMs and >50% of interactions are useless but I'm so wary that I doubt anything bad could happen. I also feel like I'm the target audience on HN, which someone who blindly trusts this magic black box (from their POV) might not be
That's why people adopted it. Google got worse and worse, now the gap is filled with LLMs.
LLMs have replaced google, and that's awesome. LLMs won't cook lunch or fold our laundry, and until a better technology comes around which can actually do that all promises around "AI" should be seen as grifting.
Considering the massive cost and shortened lifespan of previous models, current gen models have to not only break even on their own costs but make up for hundreds of millions lost in previous generations.
As soon as they find a way to embed advertising into LLM outputs the small utility LLM's provide now will be gone.
Designing operating systems that people would actually pay real money for doesn’t seem to be a focus for anybody except Apple. But I guess there’s hope on some theoretical level at least.
I am not curious what Bing says about Microsoft's ethics.
The big question is who will be first to get it done in an unobstrusive way (subtlely integrated into text with links).
The ways for circumventing the influence are currently being dismantled and these AI RTB ad systems are well funded and being built. AI news feed will echo the message and the advertiser will be there again in their AI internet search.
We will cede the agencies of reality to machines responding to those wishing to reshape reality in their own interest as more of the human experience gets commodified, packaged, and traded as a security to investors looking to extract value from being alive.
LLMs like Phind can cite parts of their output, but my understanding is they're prone to missing key details from or even contradicting the citations (the same issue as LLM "summaries" missing or contradicting what they're supposed to summarize), at least moreso than human writers (and surely more than reputable human writers).
LLMs give more accurate results for many queries, and they also include knowledge from books / scientific papers / videos which are not included in the standard google search result pages.
also LLMs have no ads.
so there is net benefit for users to use LLMs for search instead of googling it.
and I'd bet nowadays the percentage of false/outdated information is the same on both Google and LLM results.