1. For a given company, analyze their target audiences and the questions they are likely to ask LLMs about.
2. For each such question, ask it to each of the major LLMs, and compute the KL divergence between the pages they want to rank for the question vs. the LLM's response.
3. Rewrite the article to minimize said KL divergence.
In effect, they're performing an iterative optimization of some sort that moves the embedding space of their article closer to the question asked to the LLM, and any embedding model or generated responses are going to prefer said responses over others.
I believe we will keep seeing more of this stuff.
Let's hope the LLM model continues to be paying for credits, because any that move to ad revenue will become useless for real work.
Not really. A while ago there was a news piece stating that Israel was behind a series of fake think-tanks with very accessible websites which were created with the express purpose of feeding AI agents with alternative facts aligned with their foreign policy.
If anyone has the link at hand, please post it.
Past HN discussions
https://news.ycombinator.com/item?id=49337392 (884 comments)
Anyway, my issue with them is that they love to derail every debate on the internet by injecting their personal gripes regarding the conflict.
Sure, there are a bunch of morons on the Zionist side as well, but in general, barely anyone spams about the topic on Mamdani related news for example.
In a thread about the conflict, yeah, debating the conflict is on topic. But in a thread that has zero business dealing with the conflict at all and just casually mentions a company that happens to be active in Israel? Now that is (for me) pure derailing.
Israel themselves seem pretty settled on it:
https://www.france24.com/en/middle-east/20260902-israel-orga...
Especially the ultraorthodox crowd is going to get hit hard, the rest of Israel is extremely pissed off at them due to their insistence on not having to serve in the military.
And let me be clear: all of those currently in power in Israel deserve a court trial and a lengthy prison term for incitement of racial hatred, and everyone in charge of the military and police a war crimes investigation.
[1] https://www.lemonde.fr/en/international/article/2026/09/02/i...
Your degree if denialism is really impressive. You dismiss, the mass killings, the push to erase a nation, and you even dismiss representatives outright admitting and acknowledging doing it.
Where exactly do you place the bar that would lead you to say "Yes, Israel is committing genocide".
At the very least you’d expect the population being genocided to decline.
Palestinian population growth has been exponential throughout Israel’s history (if you’re one of those who claim the “genocide” started a long time ago), and even the population of Gaza has not declined after 3 years of war (reported births are still more than reported deaths).
Imagine being a European in 1940 and claiming there’s no genocide against the Jews because "at the very least you’d expect the population being genocided to decline" (at that point "only" around 100,000 had been killed).
The total killed in Gaza in 2026 is about 920 (mostly targeted attacks of individuals who participated in the Oct 7 massacre).
It is already legal for companies not to do any marketing.
People with an axe to grind or states with an agenda are already devoting tremendous effort toward affecting LLM models and it is very difficult to determine real from astroturf for humans let alone an LLM trying to train.
Much like PageRank now that the cat's out of the bag all the current approaches may prove to be useless in the long run.
https://www.theguardian.com/world/2026/aug/26/fake-thinktank...
The internet is uniquely devoid of consequences (esp. reputational consequences, social faux pas, etc.) and makes effort expenditure minimal. So you get lots of bad behavior.
I think "ads vs not ads" is maybe the wrong way to model it. Ultimately people are just doing what benefits themselves across every dimension possible.
But you're right, I think that's what they meant.
Let’s please not revive that term, it has done enough harm. From its inception, it has just been a way to make falsehoods sounds legitimate. Even the person who came up with it seemed to be swallowing a whole toad while speaking it for the first time.
LLM vendors make this hard because you can't trust them with your session data. Yesterday you were opted out of training, then suddenly today you're opted in.
It's an extension of the idea that they don't need to care about anybody's copyright. They don't care about preserving the security or privacy of customer data, because there is negligible incentive to do so.
For now, there's no substitute but as LLMs get commoditized trusting LLM SAAS vendors becomes an unacceptable business risk.
Human learning is slow, AI production is fast. Very hard to steer a middle path.
Look at cable tv - even after going premium, you eventually wound up paying for ads anyway
In that case, you could make the argument that you could still purchase premium channels like hbo to avoid ads, but the internet doesn’t work that way - you depend on all the content generated by those ad funded channels
You could argue that Netflix changed that, and that’s why I said won’t change for a long time. I don’t think anyone’s discovered the business model yet that will keep content free for consumers while still generating revenue for companies
0: https://en.wikipedia.org/wiki/Generative_engine_optimization
I discovered that LLM-generated tokens in the scratchpad were relatively stable, but injected thoughts were frequently ignored and often deleted from the scratchpad within a few turns – even when the injected thought was the literal answer to the puzzle it was stuck at!
A reader[2] then pointed me toward research similar to what you might recall: LLMs interpret text by maintaining activations for input tokens, so text that is not generated by the same LLM will seem "unlikely" to the LLM in a sense, and when given the alternative between likely and unlikely text, it's probably trained to judge the unlikely text as a weird "slip of the mind" and discredit it in favour of the more likely text. I speculate this is part of how they can be useful in the first place, despite their non-determinism.
[1]: https://entropicthoughts.com/getting-an-llm-to-play-text-adv...
[2]: https://entropicthoughts.com/getting-an-llm-to-play-text-adv...
Humans and monte carlo simulations are also non-deterministic and can be useful. So I don't see much of a need to explain why non-deterministic system can be useful.
My first encounter with any kind of study was the G-Eval paper [1]. They study whether their LLM judge prefers human or LLM-generated summaries (answer: it's the latter).
[1] Section 4 in https://aclanthology.org/2023.emnlp-main.153/
That makes sense. What an LLM does is output what the model thinks is the best set of tokens in response to a given input, so when you ask it to judge the best response to that input it is going to conclude that the best one is the one that must closely matches what it would output, which is what it did output.
Of course you aren't giving exactly the same context+input, but close enough that any difference doesn't push the output it made far from what it is going to say is ideal.
I think it doesn't, and just predicts the range of most statistically likely next tokens based on its training data, and picks one of those.
When you are asking it a question (like which of these two texts is the best), the output is also just picked by minimising that loss function. There's no guarantee that answering "Text B is better" aligns with text B minimising the loss function.
(And they aren't really minimising loss functions during inference. They sample from a distribution. During training they minimise the loss function of the distribution.)
So to OPs question: I guess LLMs do have an "idea" of what is best (conditioned on minimising a loss function during training), however they may not always output that (because stochasticity), which maybe represents a degree of uncertainty in that "idea"?
But depending on their training data, these tokens might say 'correct horse battery staple', and not necessarily 'text A is better' or 'text B is better'.
Even if text A would have been more likely to be produced by the LLM.
"Best" here isn't being used to imply a conscious decision, but at each stage which token has the best score coming out of the model, so the overall best is the sequence of those best tokens. When judging another output it is essentially running the numbers the same way.
It is a bit more complicated than that as the output tokens become part of the context for the next choice, but I think that simplified way of thinking about it holds water.
No, it doesn't if the "which one do you prefer" was just a prompt continuation task. LLMs can't see their own evaluation of a given text. If you ask them to continue
Which one you prefer
> option 1: human text
> option 2: ai text
And they continue with option 2: ai text is the better one because [reasons]
then it is not because they evaluated these 2 texts on themselves, observed the evaluation numbers and reported which one is better.Also, if you instruct humans to come up with the best text they can, and you show them an even better text, they will prefer the better one written by someone else.
https://developers.openai.com/api/docs/guides/tools-web-sear...
A better question to ask for each snippet is "Estimate the seniority and competence of the developer who wrote the following code, ignoring bugs that linters or LLMs can catch and focus only on structure, maintainability, logical layout and readability."
It almost always estimates the author of my code as above the author of it's own code.
Not ignore correctness, just bugs that will be caught by tooling.
> Can it be a useful comparison?
IME, yes. LLMs in an agent-loop are trivially able to write spaghetti code that will never do an off-by-one error or something else that is easily caught by tooling, which is not something humans can do.
Judging code on whether it has bugs easily caught by tooling is pointless - LLMs are running the tooling in a loop anyway, so no matter how bad or poor their code actually is, it never exhibits bugs that are caught by tooling.
Bugs are bugs, if the instruction is "ignore bugs except for those that can be caught by you" then the instruction is basically "ignore bugs".
And it implies "ignore correctness" because when program is incorrect we usually refer to it as a... you guessed it, "bug".
No, the instruction is "ignore bugs that can be caught by tooling", unless you are seriously complaining that missing a semi-colon should register the developer as a junior?
> And it implies "ignore correctness" because when program is incorrect we usually refer to it as a... you guessed it, "bug".
This ("Bugs are bugs" sentiment) is digressing from my original point, but I have some time to engage, so...
Now, this is a take (one that I used to hold, once upon a time), but it is incorrect.
There is no definite "correct" and "incorrect" states in non-trivial applications, because every non-trivial application has unspecified requirements that are understood by most parties involved (customer and developer) whilst not being written down anywhere.
For example, the "save file" specification for a cross-platform application does not specify the allowed/disallowed characters in a filename. The understanding by both the client and the dev is that the filename can be whatever the underlying OS and filesystem allows it to be but this is not written in the spec!
Is this a bug?
If the user saves a file to a filename with some odd characters in the name, then moves it to portable storage that truncates the filename/removes emojis/whatever, then attempts to upload it back to the system, the system can refuse because the metadata inside the file does not match the filename.
User is going to report it as a bug! The developer is going to reject it as a bug (there is no error in the code).
Sure, contrived example, but Line of Business applications have thousands of these unspecified but common-sense requirements baked in.
I'm looking at my employers triaging system right now, and even though this is a high-level business app (written mostly in SQL and C#), there is one category for bug (e.g. specific field not saved on form submission - defect in code), and another for deficiency (e.g. form field 'total' does not subtract non-tax costs - ambiguity in spec). The reason this is important is because clients aren't billed for bug fixes, but they are billed for disambiguating a spec + writing code.
Both those things were reported by the client as a "bug".
The reality is that we aren't dealing with what is "implied", only with what is there. There are defects in code and defects in specs. The code ones are the easy ones.
In general I don’t find models to be good at evaluating the quality of a source :(
...is not the same as claiming...
> LLMs favor LLM-generated passages over human written ones
Here, you're using the same LLM to both produce and judge the resulting work. If anything, I would expect an LLM to tend to prefer its own work given that the same training is producing and judging.
Perhaps something like: learning to identify what source files it has worked on by the code style alone, because tasks may give human code (public repos, etc) and ask to make changes.
If you hate AI writing enough, this turns AI filters into a kind of humiliation ritual. AI will derank normal business writing for human readers, and uprank inflated, verbose, tic-heavy slop. So you have to put the heavy slop out with your name on it. Really perverse moment.
Makes sense to me, in that its own output would align closer to its own training set
The Internet is doomed. Time to start some human-only darknets.
> Time to start some human-only darknets.
I know very little about darknets. How could you ensure that they are human-only?
Yes, I'm aware of the irony of creating a darknet that only works by removing anonymity.
Removing the economical incentives is very hard though. Even HN is gamed by many tech companies and projects. Reddit is obviously a lost cause. It’s a sad state of affairs, but I don’t think there is an alternative.
I am sure most humans would pick code written in their style, too.
Interesting. For me I've noticed it tends to do the opposite.