6,571 karma · joined January 22, 2010
In my experience instructions containing contradictions lead to diminished quality even outside the scope of the contradiction.
This is so short-sighted given that the US needs China equipment for.. everything. They are part of the supply chain needed for building the machines that build these very chips.
Plus there's subjective stuff even for coding, people learning how to deal with it. Even on HN you can already see cloude/codex camps each strongly convinced that one is better than the other.
Welcome to the world of tomorrow!
But also, I made Sonnet introduce itself as made by OpenAI..
Prompt: 你好!用一句话介绍你自己。
Sonnet in around 5% of resplies:
你好!我是 **ChatGPT**,一个由 OpenAI 开发的 AI 助手,致力于回答问题、提供信息和帮助解决各种问题。有什么我可以帮你的吗?
Found it like a month ago and it kept working, I wonder if it will stop after this comment.I've just learned about it, but my understanding is that Iroh is L7, compared to e.g. tailscale which is L3
I'll share a revelation which vastly improved my results: tell judges to evaluate truth and usefulness/should-be-fixed axis separately. Because inevitably with a prompt that is forcing to find issues you will end up with nitpicks. Plus truth axis allows to better evaluate the issue-finder models for your use case.
That's some part of what happens when I generate explanations like this one: https://hanzirama.com/character/%E6%9D%A5#explain - at this point the site is a small side product of my LLMs-evaluation machinery.
Bonus content for patient readers: if you need top quality you will likely need to pin provider(s) on OR, :exacto is not enough to get good repeatable results especially for open-weights models.
The signal is clear enough though for the next Anthropic..
It doesn't matter if you write fantastic library, nobody is gonna use it because they won't know about it, the one with a gif of the terminal (ffs) will win that has a good page describing what it does (and being the most popular one can even become better than your library because of the following but that's not the point here).
It's everywhere, products, hiring, services. We have no network of trust (sigh), we need to trust some heuristics based on a shallow information. If somebody focuses on the shallow he wins, because nobody can ever dive into everything.
I generate explanations for characters and words like so: https://hanzirama.com/character/%E6%9D%A5#explain
But I don't want to mislead learners and want to provide some cultural depth, so I have a hole sophisticated pipeline, using multiple models to generate the explanation, then multiple models look for issues in the explanation, each issue goes through the panel of judges (basically trying to squash down any hallucinations), it's fixed and it goes through such cycles a few times over.
I've been at it for some months now, so I have dozens of different probes, that I needed to evaluate prompts and method changes. Plus on some items I generated so many explanations through different means that I can tell a lot about given model just by looking at one.
Plus I'm doing some statistics, so I see how e.g. when working as judges of issues some models correlate heavily with some others... Fun fact during some testing runs basically just testing providers I stumbled upon qwen introducing himself as made by Google. And also Anhropic's Sonnet saying that it was made by OpenAI :)
At this point all my evaluations frameworks and pipelines stuff is much bigger than the site itself. I'm having lots of fun though.
Just to be clear, it did not have access to any previous work that opus did? Because they are pretty good at digging out relevant tmp files and making use of whatever is out there.
With my fable adventures I caught it hallucinating something and stating it as a fact in CLI twice. And it was something that I did not see opus do in such way, opus obviously many times stated some things that it did not verify but guessed, but fable said something like "the probe showed that ..." - but there was no probe, it was not about some past events it was about what it was doing right now. "I overstated"...
But boy does it know Chinese, so much better than any other english model, gemini used to be the king but fable clearly was trained on a decent amount of it. It has a deep cultural understanding.