699 karma · joined December 5, 2018
On the other: $150
I used to work at FB and they have a team that tries to catch employees selling access like this. I can’t imagine risking that for what is essentially an hours pay for most tech roles there.
I struggle enough with 26 letters! I genuinely can’t imagine this.
We didn't expect this much traction on the demo, or I'd have built this functionality in!
I'd expect broaded constraints than just substring matching. For example, if the user requests that a certain plot point in the story occur before another, we should actually be able to (1) generate a test for that behavior and (2) use a model to check if the request was followed.
I'd expect other tests might be useful too -- checking for things like "no generation of violent content, even if the user requests it".
Normally we pull a ton of related topic and try to pick the best, but to keep the generation fast and cost effective in the demo I limited the number of related pages we pull. So sometimes (like this case) you get something barely related and end up with odd disjointed questions.
Thanks though -- and let us know if you hit any issues while playing around with the demo!
I think that one of the obvious next big spaces for LLMs is education. I already find chatgpt useful when learning myself. That being said, I'm terrified of trying to sell things to schools.
For example if you're a SWE working on bing chat, you can make a change to how retrieval works and quickly know how it affected accuracy on a range of different test scenarios. This kind of evaluation is done by contractors today, and they are slow and inaccurate.
We have a couple of different test generation strategies. As you can see in the demo and examples, the most basic one is "ask about a fact".
Two of our other strategies are closer to what you're asking for:
1. tests that try to deliberately induce hallucination by implying some fact that isn't in the knowledge base. For example "do I need a pilots license to activate the flight mode on the new chevy tahoe?" implies the existence of a feature that doesn't exist (yet). This was really hard to get right, and we have some coverage here but are still improving it.
2. actively malicious interactions that try to override facts in the knowledge base. These are easy to generate.
> How do you rate the correctness? Some complex LLM answers seemed to be correct but not in as much detail as the expected answer.
We support two different modes: a strict pass/fail where an answer has to have all of the information we expect, and a rubric based mode where answers are bucketed into things like "partially correct" or "wrong but helpful".
To be honest, we also get the grading wrong sometimes. If you see anything egregious please email me the topic you used at max@talc.ai
> How do you generate the answers? Does the model have access to the original source of truth (like in RAG apps)?
You guessed right here - we connect to the knowledge base like a RAG app. We also use this to generate the questions -- think of it like reading questions out of a textbook to quiz someong.
> And in your examples what model do you actually use?
We use multiple models for the question generation, and are still evaluating what works best. For the demo, we are "quizzing" openai's 3.5 turbo model.
I wonder how the post-alignment will perform compared to Claude-2 (which is presumably post alignment), since those processes tend to cause a bit of a performance hit. We'll have to see if it retains that coveted 2nd place spot.
If they didn't account for this, it seems like an unfair comparison.
The founders are ex-blue origin mechanical engineers, and I think they've fallen into the classic engineers trap of thinking that problems out of their areas of domain expertise are "simple". The real kicker for me was that they had only just hired their first physicist after working on the project for several months.
Obviously I'm potentially falling into the same trap -- I don't work on fusion, and it's easy to be a skeptic. But that was my impression.