help make jev!
typesafe.ai/careers
262 karma · joined October 19, 2014
help make jev!
typesafe.ai/careers
we actually use astra (and fable) in this way for our evals: evals.typesafe.ai
someone on the team cooked hard on that and it shows example traces comparing our model to opus/sol
(shilling my blog post on that jaggedness: https://www.completeskeptic.com/p/lies-damned-lies-and-bench...)
the one nuance I'd get into is I'd call it "zero-shot" over "instruction-tuned" (the latter often implies a particular distribution), but very safe for sharing
I very much agree with this and want to hone in on where do actually disagree. Would you say a linear classifier hallucinates?
(though ideally everyone gets off the waitlist and can try it out for themselves )
we do think there's a lot of potential though and do want coding themed releases soon
constrained decoding (OpenAI-style structured outputs) make models dumber unfortunately - the short+dense version is that simply masking logits is insufficient because if ever a model was assigning probability to an invalid token, the model is by definition confused. you'd be better off erroring IMO
I don't want to shill my blog too much, but I will say data is probably far most interesting than architecture: https://www.completeskeptic.com/p/the-bitterest-lesson
> Also "can't hallucinate" seems wrong? Sure, it can't emit an invalid type, but it can still emit a completely wrong valid value.
that is likely true of all ML! perhaps we could debate semantics, but I don't think it's fair to say a random forest "hallucinates" in the way LLMs do
we also thought the voice at the end was AI-ish, but apparently that's a real voice actor but slightly sped up
it is a structured data model, but technically not a language model (it doesn't generate language)
this is complex, but generating text is highly complicated and requires mode dropping to make long cohesive text
I do agree that the comparison to LLM tokens is hard to understand (also because output tokens are not comparable).
But yes, text or structured state (like a JSON with multiple pieces of text in) -> decisions out (e.g. choice maps to "match" statement, "score" maps to sorting, "noul" short for bernoulli maps to if-statements)
An interesting follow up would be when did you realize GenAI wasn't as good as you thought in that "oh shit" moment
[1] co-author of InstructGPT/RLHF/ChatGPT
TypeSafe AI (https://typesafe.ai/) is a well-funded stealth mode company making a new class of AI models which enable a reliable way to embed composable intelligence into traditional software, founded by a co-author of ChatGPT/RLHF/GPT4 (answering the questions "If AI is so smart, why are we barely automating any work? Why does today's AI always need a human-in-the-loop?").
We're especially looking for people to join our Synthetic Data research team - no previous data experience needed (synthetic or otherwise). We need strong engineers who care deeply about the end use cases and whose instincts are to NOT trust LLMs blindly and look at the reality of their generations (much more investigator in the spectrum of investigator vs optimizer).
Short 5 min video (of me) describing the company: https://www.youtube.com/watch?v=LE3bGTaAgOE
See open roles here: https://typesafe.ai/careers#opportunities
Seems so to me - see GPT-5.4[1] and 5.2[2] announcements.
Might be an tacit admission of being behind.
[1] https://openai.com/index/introducing-gpt-5-4/ [2] https://openai.com/index/introducing-gpt-5-2/