Don't classify, hallucinate
softwaredoug.com
softwaredoug.com
It's the same for so many things:
- reading documentation (what do I expect this function to be called?)
- finding clothes in a shop (something long-sleeved and light)
- picking the fridge for dinner
- finding a book in the library...
so many analogues where I'm not coming cold to a choice.
He clearly didn't know enough about vector embeddings.
What I meant is something more like explicitly programmed vs. learned. Intelligence can result from learned behavior, but not from explicit programming of rules by humans.
An aspect of this is that “learning” is unpredictable - we can’t predict in advance exactly how the resulting model will behave, except broadly. It seems non-deterministic if only by virtue of its complexity, which is beyond anything we’re able to predictively model.
This is incorrect.
Simulated intelligence can and has been encoded explicitly by humans defining rules programmatically in the form of expert systems[0].
I also have worked on/with expert systems.
> I don't agree that they achieve "simulated intelligence". They're preprogrammed with a set of domain-specific rules that are trivially simple by comparison to even relatively simple and small neural networks.
This position does not account for fuzzy logic[0] nor an expert system's ability to produce an answer of "I do not know and here is why", which neural networks are incapable of doing.
I am not saying expert systems are "better" than ANNs as both are algorithms having significant value for what they provide. What I am saying is neural networks are pattern-matching algorithms, quite useful in their own right, and do not possess the ability to identify the lack of existence.
Comparing the two in 2026 seems like a bit of a joke to me. I'm not saying there's no role in future for traditional expert systems or fuzzy logic (or hand-written code, for that matter), but to claim they're "intelligence" or even "simulated intelligence" implies such a trivial definition of "intelligence" as to make it a useless term.
> an expert system's ability to produce an answer of "I do not know and here is why", which neural networks are incapable of doing.
Why do you believe that? Here's an excerpt from a response I received from Claude tonight:
> "I want to be honest about a limitation: I can't reliably confirm fine construction details — like exactly which sub-assembly is bolted to the spoke flange versus the fixed axle — from a marketing cutaway graphic at typical web resolution. Those images tend to be stylized/exploded-view illustrations meant to show 'there's a battery and a motor in here,' not engineering-accurate cross-sections with clear rotating/stationary boundaries marked."
This is after it examined two images I provided it with, and related it to the discussion we'd been having.
This demonstrates that it can indeed answer "I do not know and here is why", so your idea about what neural networks "are incapable of doing" is clearly incorrect.
And even if I grant your trivial threshold for intelligence, an interaction like that one clearly demonstrates a far superior degree of multi-modal intelligence, reasoning, and understanding that no expert system or fuzzy logic has ever even come close to achieving.
Because LLMs are artificial neural networks[0] (ANN), which are statistical in nature, and thus intrinsically non-deterministic. Pretty much every AI algorithm has randomness involved in its definition and many (most?) incorporate probabilities.
0 - https://en.wikipedia.org/wiki/Neural_network_(machine_learni...
Fixing the seed is still intentional bias. Or you could force it to always take the one token with the highest probability, but that is still biased sampling. Deterministic, sure, but intentionally wrong just to avoid a technically
This assertion is "oddly" similar to the GPT answer "neural network inference determinism" produced:
Neural network inference is often non-deterministic due to
factors like floating-point arithmetic and concurrent
execution, which can lead to variations in output even with
the same input.
Surely this is but a coincidence.Regarding your previous statement:
> However, on a technical level, neural network inference truly is inherently deterministic.
This holds for a vanishingly small set of conditions, none of which include randomness, nor when context and transformers are involved, let alone underlying model evolution (thus making model use over time non-deterministic).
My statement and GPT's statement are both correct answers to the same question, so I think it makes sense that they would be similar. Are you accusing me of having paraphrased an LLM in writing my answer? I did not, I just remembered having read Thinky's post on the subject [0], which GPT has probably read also.
> This holds for a vanishingly small set of conditions, none of which include randomness, nor when context and transformers are involved, let alone underlying model evolution (thus making model use over time non-deterministic).
There are plenty of ways to introduce nondeterminism into any system. By your standards, I doubt you could point to a single deterministic system in the world. print("hello, world") is only deterministic if your CPU is properly shielded from cosmic rays and your OS isn't out of memory etc. There are some inherently nondeterministic processes, like the stochastic methods used to train models or the random sampling used at inference time if you have temperature!=0, but inference under greedy decoding is conceptually deterministic.
[0]: https://thinkingmachines.ai/blog/defeating-nondeterminism-in...
Often it's a difference between repeatable versus predictable, or whether a system has chaotic aspects like the configurations of a double-pendulum or weather-forecasting.
Sometimes it's the difference between determinism in-theory versus in-practice, especially when various optimizations are being applied to save money.
The LLM inference process can be 100% deterministic but the weights can still make the end result quite chaotic. Just because temperature>0 improves results doesn't mean its an innate part of the mechanism. Just because scale-out architectures introduce jitter in communication doesn't mean that's an innate part of the mechanism.
Except not as much as I'd like... they often also don't know what the hell I'm talking about, and it still takes them twenty minutes of Googling to find the right page!
It nearly always works.
The entire problem of search is that the user has the wrong data and wants to use it to receive the correct data. That was the start, not the state we’ve ended up at - it is unironically how we got to LLMs.
Worse is when you don't know whether the answers you have are totally wrong.
“Hallucinate” is misleading here. In the given example, a classification is being done very successfully - it’s just that it requires an extra step to map it to an arbitrary predefined list of classifications.
If you can articulate why you think this isn’t a good approach, I’d be interested to hear it.
This method is sensitive to the thresholds (what is the maximum distance between embeddings for them to be still considered part of the same semantic group), so I run it all in an agentic loop where an agent tries different thresholds and clustering algorithms until it's satisfied with the result, plus it may deduplicate some groups.
I run it all on self-hosted hardware, so it costs nothing to leave it running for, like, a night, and as a bonus, none of the corporate data leaves the office. I think a rigid set of manually created classifications may not capture all the possible classifications that can exist. Needs a review by a human, though.
We had a similar problem where you can literally millions of email that we were pretty sure came from only a limited set of bad actors.
We first started classifying emails into buckets by From, mailserver relay chains etc as that's all we had to to go on.
Over time, those buckets got linked to spammer signatures and then we narrowed down from there.
Fascinating to see this happening nowadays with LLMs.
You can slice and dice it a ton of different ways, but the significance of groups is incidental.
It's a good starting point, but having done this a few times for a few companies it always seems like it needs substantial human review.
Even better is to search the corpus first with like naive BM25 / embedding search, aggregate over top N to get most representative categories, then have the LLM categorize in that set.
But honestly, it only works for common knowledge that's already in the LLM. If the target document contains very niche or private information, then the hallucinated answer's embedding can be even farther away than the query's.
Isn't this begging the question that the hallucinated classification will be more selective with respect to the real schema than the query itself? What would the dot product of <E(search query), E(schema)> have given?
Even if that is too vague, smaller LLMs are capable rerankers; return the top N matching true categories and ask for a contextual ordering.
I've found, though, getting it in the language of the vocabulary has generally improved performance.
Further, when searching for "blue shoes" you want to separate the color from the item type. So its useful to have a dumb LLM do this for you. And with the LLM in the loop, its further useful to get it into the language of the taxonomy to improve embedding retrieval accuracy.
There are of course many ways to skin the cat here :)
But no classification is perfect. In search in particular, you will also want to have places for manual intervention for high priority queries.
Which accomplishes the same thing as HyDE, but in the model instead of in text space. If the encoder has a query mode you skip the rewrite.
Voyage has an example of this in `input_type` https://docs.voyageai.com/reference/embeddings-api
Additionally you could experiment with a reranker instead of an LLM or after reranking take top-3 results and then feed to LLM as input in order to reduce input token costs.
Putting all the labels into the LLM is super expensive per call when you have millions of items to classify.
You can't reduce the number of labels becasue they are correctly organizes/structured. This class of problem exists in many different domains.
This problem has also been solved for 3 decades now.
Orginally I started writing that as sarcasm, and now I'm not quite so sure.
{ rationale, categories }
Where you don’t really care about the rationale but you’re using it as a pseudo thinking for models that don’t support it.Luna is surprising capable and cheap, and I haven’t done this type of thing since before GPT 5 so might not be such a useful trick now
It didn't end up being very useful - I ran a comparison where I just had a bigger agent do the organization in a more straightforward way, and that had better results.
I did find that Flash 3.6 High was >9x faster than Luna xhigh for this task, and got very similar results, though.
But if accuracy matters, you can't rely on embedding sort to get a closet match. With a real test set they usually don't hold up under scrutiny.
Everything in AI is like this. You get an idea, try it once or twice, "LGTM" and you ship. Then it never survives contact reality.
Embedding sort gives you a better shortlist than the whole list, but you will probably want a heavier model to vet candidates.
```
Request 1: "brown coffee table: " + {Root Schema} => "Furniture"
Request 2: "brown coffee table: Furniture / " + {Furniture Schema} => "Living Room Furniture"
Request 3: "brown coffee table: Furniture / Living Room Furniture / " + {Living Room Furniture Schema} => "Coffee Tables"
```
Many more round trips, but classifying products is not a latency sensitive task.
https://github.com/aurelio-labs/semantic-router
I guess it is based on the same fundamentals as well.
[1] https://www.postgresql.org/docs/current/textsearch-intro.htm...
[2] https://www.elastic.co/docs/explore-analyze/query-filter
You'd think they would have solved it by now.
(Lets ignore for now that no one seems to agree to what should be the spec sheets)
Instead: use an LLM to build a large (old-school) database of products with all their specifications. The LLM can also build the schema for that database as it finds more data.
Then use an LLM to query that database based on the user's specifications (+ add some intelligence to find nice suggestions for a birthday if wanted, but I'd consider that an extra).
The idea is to make structured queries using these protocols which can be used to fetch top products matching the user needs instead of just relying on semantic search.
https://developers.openai.com/commerce/specs/file-upload/pro...
But you'll be amazed by the abundance.
I noticed yesterday when browsing on mobile that there used to be a box where I could search reviews and it got swapped with a Rufus box. I guess somebody needs to juice their engagement numbers for an investor briefing.
honestly, Amazon doesnt even need AI it just needs a better UI, more metadata for its products and to make reviews less scammy.
Asking an LLM for a list of every county in the USA for example, or every county with a population of more than 100,000 people.
Even if those county names and their populations are mixed up in their weights, the nature of next-token-prediction does not lend them to effectively answering comprehensive, detailed questions like that.
An agent system build on top of an LLM can do it, if it has access to tools which can help access eg a table of counties and then filter them with SQL or Pandas or similar.
Considering that agents are not a new concept, why isn't this a solved problem by now?
We've got the early LLM-based AI agents in 2023, and it only became a popular, mainstream thing in 2025 - with Claude Code.
Pretty cool technique honestly. You could do it the other way as well right?
If you had a list of categories you have the model to generate a sample query and then do embedding on that?