HNHacker News
TopNewBestAskShowJobs

sysmax

179 karma · joined September 8, 2022

submissionscomments
sysmax··on Don’t let an LLM make decisions or execute business logic
LLMs are a glorified regex engine with fuzzy input. They are brilliant at doing boring repetitive tasks with known outcome.

- Add a 'flags' argument to constructors of classes inherited from Record.

- BOOM! Here are 25 edits for you to review.

- Now add "IsCaseSensitive" flag and update callers based on the string comparison they use.

- BOOM! Another batch of mind-numbing work done in seconds.

If you get the hang of it and start giving your LLMs small, sizable chunks of work, and validating the results, it's just less mentally draining than to do it by hand. You start thinking in much higher-level terms, like interfaces, abstraction layers, and mini-tests, and the AI breeze through the boring work of whether it should be a "for", "while" or "foreach".

But no, don't treat it as another human capable of making decisions. It cannot. It's a fancy machinery for applying known patterns of human knowledge to the locations where you point based on a vague hint, but not a replacement for your judgement.

sysmax··on Googling Is for Old People. That's a Problem for Google
I think, they are trying to push back against generated pages. I faced this exact problem myself. We recently published an interactive source code navigation tool [0] where you can find examples for commonly used functions from some embedded SDKs. Google indexed it immediately and almost immediately it got a spike of views.

Then, an interesting thing happened. Most pages simply disappeared from the results. Search console shows them as indexed, no problems, no manual actions, but if you google up those functions, there results are not there.

It took some statistical analysis to figure out that they appear to be capping the number of pages. Out of all the pages Google crawled, it picked some percentage of the "most important" ones and it's showing those. The importance, by the looks of it, was computed from the number of incoming links, prioritizing pages for common stuff like int32_t that nobody googles.

It's not ideal, but it kinda makes sense. It's 2024. You can use AI to generate plausible content for any search query you can think of. And unless they put some kind of limits, we'll get overrun with completely useless LLM-churned stuff.

[0] https://sourcevu.sysprogs.com/

sysmax··on QwQ: Alibaba's O1-like reasoning LLM
Well, to be perfectly honest, it's hard question for an LLM that reasons in tokens and not letters. Reminds me of that classic test that kids easily pass and grownups utterly fail. The test looks like this: continue a sequence:

  0 - 1
  5 - 0
  6 - 1
  7 - 0
  8 - 2
  9 - ?
Grownups try to find a pattern in the numbers, different types of series, progressions, etc. The correct answer is 1 because it's the number of circles in the graphical image of the number "9".
sysmax··on The industry structure of LLM makers
LLMs are a very good tool for a particular class of problems. They can sift through endless amounts of data and follow reasonably ambiguous instructions to extract relevant parts without getting bored. So, if you use them well, you can dramatically cut down the routine part of your work, and focus on more creative part.

So if you had that great idea that takes a full day to prototype, hence you never bothered, an LLM can whip out something reasonably usable under an hour. So, it will make idea-driven people more productive. The problem is, you don't become a high-level thinking without doing some monkey work first, and if we delegate it all to LLMs, where will the next generation of big thinkers come from?

sysmax··on OpenCoder: Open Cookbook for Top-Tier Code Large Language Models
I was just messing around with LLMs all day, so had a few test cases open. Asked it to change a few things in a ~6KB C# snippet in a somewhat ambiguous, but reasonable way.

GPT-4 did this job perfectly. Qwen:72b did half of the job, completely missed the other one, and renamed 1 variable that had nothing to do with the question. Llama3.1:70b behaved very similar to Qwen, which is interesting.

OpenCoder:8b started reasonably well, then randomly replaced "Split('\n')" with "Split(n)" in unrelated code, and then went completely berserk, hallucinating non-existent StackOverflow pages and answers.

For posterity, I saved it here: https://pastebin.com/VRXYFpzr

My best guess is that you shouldn't train it on mostly code. Natural language conversations used to train other models let them "figure out" human-like reasoning. If your training set is mostly code, it can produce output that looks like code, but it will have little value to humans.

Edit: to be fair, llama3.2:3b also botched the code. But it did not hallucinate complete nonsense at least.

← PreviousPage 2 of 2