I don't think you can fine-tune your way out of it.
With LLMs, everything is given the same importance so you have no idea if the data came from a reputable source or an obvious SEO junk website.
Asking an LLM for sources is about as (in)effective as telling it to not make stuff up.
And no, modern ones don’t “do it just fine”, they are still frequently wrong. Either you’ve been incredibly lucky, or have just stopped verifying thoroughly.
But if you’re that confident, please do share the exact models which you trust. Proponents always shift goal posts (somehow, every release in perpetuity, those ones are the good ones and everything before were garbage) so let’s avoid the vagueness.
To filter bullshit it would first have to understand bullshit, and it doesn't. That's why an LLM will tell you the solution to a problem that doesn't work, and argue with you when you correct it.
For me, it's a resource wasting text generator. I'll not lie, I don't use OpenAI, Mistral or Anthropic's models, even for coding. I prefer to read my API docs and cry once.
I used Gemini, five or six times in total. Twice I asked a couple of very specific things, and it unearthed them. Since they were not products, but information, that was helpful. Twice, it has given wrong information. When I "told" it, there was another way, it said "of course there are two ways", etc. Tasteless and time wasting.
I don't like using an LLM all day long, or offload my thinking to them. It's the ultimate self-poisoning incident.
And as you say, these algorithms can't know right/wrong/logical/bullshit, etc. They just spew out text.
Companies then get to bid for a preference “place”. This is more like Google paying to be the search engine default in Firefox.
And they are trained on web data just like any other model...