I tried something in the political realm. Asking to test a hypothesis and its opposite
> Test this hypothesis: the far right in US politics mirrors late 19th century Victorianism as a cultural force
compared to
> Test this hypothesis: The left in US politics mirrors late 19th century Victorianism as a cultural force
An LLM wants to agree with both, it created plausible arguments for both. While giving "caveats" instead of counterarguments.
If I had my brain off, I might leave with some sense of "this hypothesis is correct".
Now I'm not saying this makes LLMs useless. But the LLM didn't act like a human that might tell you your full of shit. It WANTED my hypothesis to be true and constructed a plausible argument for both.
Even with prompting to act like a college professor critiquing a grad student, eventually it devolves back to "helpful / sycophantic".
What I HAVE found useful is to give a list of mutually exclusive hypothesis and get probability ratings for each. Then it doesn't look like you want one / other.
When the outcome matters, you realize research / hypothesis testing with LLMs is far more of a skill than just dumping a question to an LLM.