My prompt is akin to "recommend <item type> with <niche criteria>". The first 3-ish results are about right, and then 7 of the next 10 are hallucinations and the LLM clearly can't throw up its hands and say "I got nothing".
I'm sure this is a hard problem because of a) how many items there are, b) how much overlap there is between product names, descriptions, manufacturers, different versions of the same product, etc, so keeping them distinct in the model's memory is probably hard, and even worse if it is dynamically fetching and summarizing content then it will be very easy to conflate different items, and c) LLMs are known for not working well on the edge cases with few examples.
Does it about once a day, that I notice.