HNHacker News
TopNewBestAskShowJobs

CompleteSkeptic

262 karma · joined October 19, 2014

submissionscomments
CompleteSkeptic··on Ask HN: Who is hiring? (October 2026)
typesafe ai | ONSITE in SF | basically all forms of engineering (front-end / back-end / infra)

help make jev!

typesafe.ai/careers

CompleteSkeptic··on Introducing System One Models and Jev
yes and can do many of those in parallel
CompleteSkeptic··on Introducing System One Models and Jev
unfortunately all hand-written :( my chief-of-staff does unironically handwrite em dashes though
CompleteSkeptic··on Introducing System One Models and Jev
the edit is right - jev would be cheaper, faster, and more self-consistent (in general)

we actually use astra (and fable) in this way for our evals: evals.typesafe.ai

someone on the team cooked hard on that and it shows example traces comparing our model to opus/sol

CompleteSkeptic··on Introducing System One Models and Jev
1. yes a general model 2. no training at all 3. but it is focused on "System 1" tasks (more human judgment, less math reasoning)
CompleteSkeptic··on Introducing System One Models and Jev
strings (and all sequential data structures) are not allowed at all - this is how we make sure all outputs can be computed in parallel (thus no output token cost)
CompleteSkeptic··on Introducing System One Models and Jev
They don't like adding stealth startups :(
CompleteSkeptic··on Introducing System One Models and Jev
1. I am extremely on the same page 2. I do think that subconscious is not only much smarter than we give it credit for, but also much more robust than the "jagged frontier" of current LLMs

(shilling my blog post on that jaggedness: https://www.completeskeptic.com/p/lies-damned-lies-and-bench...)

CompleteSkeptic··on Introducing System One Models and Jev
very accurate!

the one nuance I'd get into is I'd call it "zero-shot" over "instruction-tuned" (the latter often implies a particular distribution), but very safe for sharing

CompleteSkeptic··on Introducing System One Models and Jev
> Type safety is not factual correctness.

I very much agree with this and want to hone in on where do actually disagree. Would you say a linear classifier hallucinates?

CompleteSkeptic··on Introducing System One Models and Jev
inputs are structured program state. there is an example at around second 30 of the doom demo

(though ideally everyone gets off the waitlist and can try it out for themselves )

CompleteSkeptic··on Introducing System One Models and Jev
we hope so! the bigger hope is to not just eat LLM market share, but to allow for people to use AI much more in the inner loop of software
CompleteSkeptic··on Introducing System One Models and Jev
the hard part for coding is actually state engineering (e.g. getting your dependencies in context) - we haven't even tried it yet (because my philosophy is we should automate the easy tasks before the hard and we've been working on getting the model smart on the former)

we do think there's a lot of potential though and do want coding themed releases soon

CompleteSkeptic··on Introducing System One Models and Jev
thanks a ton!

constrained decoding (OpenAI-style structured outputs) make models dumber unfortunately - the short+dense version is that simply masking logits is insufficient because if ever a model was assigning probability to an invalid token, the model is by definition confused. you'd be better off erroring IMO

CompleteSkeptic··on Introducing System One Models and Jev
architecture is close to the chest for now, but we have talked about writing a paper

I don't want to shill my blog too much, but I will say data is probably far most interesting than architecture: https://www.completeskeptic.com/p/the-bitterest-lesson

CompleteSkeptic··on Introducing System One Models and Jev
love that you love the manifesto! letting the first batches off the waitlist now, but we do have some early users describing their experience (https://x.com/danshipper/status/2099947471518474522)
CompleteSkeptic··on Introducing System One Models and Jev
that's right, but because these models are probabilistic, it's also possible to be confidently wrong (and all future models will be smarter still and still have that possibility)
CompleteSkeptic··on Introducing System One Models and Jev
I'm biased but I wouldn't call it misleading - generating text is super awesome and flexible, (we describe that in the blog post - and I personally use string models all the time) but it's true you pay a high tax for autoregressive generation

> Also "can't hallucinate" seems wrong? Sure, it can't emit an invalid type, but it can still emit a completely wrong valid value.

that is likely true of all ML! perhaps we could debate semantics, but I don't think it's fair to say a random forest "hallucinates" in the way LLMs do

CompleteSkeptic··on Introducing System One Models and Jev
exactly right!
CompleteSkeptic··on Introducing System One Models and Jev
definitely not AI-generated - this is my real wardrobe

we also thought the voice at the end was AI-ish, but apparently that's a real voice actor but slightly sped up

CompleteSkeptic··on Introducing System One Models and Jev
we have played with this! the fascinating thing we've found so far is that adversarial examples for our model are quite different from that of LLMs so that they work even better together
CompleteSkeptic··on Introducing System One Models and Jev
it is just a model, no harness yet ;)

it is a structured data model, but technically not a language model (it doesn't generate language)

CompleteSkeptic··on Introducing System One Models and Jev
it's our output tokens that are free (under the system one / jev column)
CompleteSkeptic··on Introducing System One Models and Jev
you could, but it the model is not optimized for text

this is complex, but generating text is highly complicated and requires mode dropping to make long cohesive text

CompleteSkeptic··on Introducing System One Models and Jev
CEO here - that is right!

I do agree that the comparison to LLM tokens is hard to understand (also because output tokens are not comparable).

But yes, text or structured state (like a JSON with multiple pieces of text in) -> decisions out (e.g. choice maps to "match" statement, "score" maps to sorting, "noul" short for bernoulli maps to if-statements)

CompleteSkeptic··on Better Models: Worse Tools
constrained decoding tends to make models dumber - this is why it's rarely used
CompleteSkeptic··on Ask HN: What was your "oh shit" moment with GenAI?
I helped train some of the first "magic" models at OpenAI[1] and it was a wild ride. We were a pretty sane + skeptical team and we weren't totally convinced the models were as general as they seemed, but the query that convinced me (and later got included in the paper[2]) was "Why is it important to eat socks after meditating?" (something that almost certainly did not appear on the internet before).

An interesting follow up would be when did you realize GenAI wasn't as good as you thought in that "oh shit" moment

[1] co-author of InstructGPT/RLHF/ChatGPT

[2] https://arxiv.org/pdf/2203.02155

CompleteSkeptic··on Ask HN: Who is hiring? (June 2026)
TypeSafe AI | Software Engineer | ONSITE (San Francisco) | Full-time

TypeSafe AI (https://typesafe.ai/) is a well-funded stealth mode company making a new class of AI models which enable a reliable way to embed composable intelligence into traditional software, founded by a co-author of ChatGPT/RLHF/GPT4 (answering the questions "If AI is so smart, why are we barely automating any work? Why does today's AI always need a human-in-the-loop?").

We're especially looking for people to join our Synthetic Data research team - no previous data experience needed (synthetic or otherwise). We need strong engineers who care deeply about the end use cases and whose instincts are to NOT trust LLMs blindly and look at the reality of their generations (much more investigator in the spectrum of investigator vs optimizer).

Short 5 min video (of me) describing the company: https://www.youtube.com/watch?v=LE3bGTaAgOE

See open roles here: https://typesafe.ai/careers#opportunities

CompleteSkeptic··on GPT-5.5
Is this the first time OpenAI has published comparisons to other labs?

Seems so to me - see GPT-5.4[1] and 5.2[2] announcements.

Might be an tacit admission of being behind.

[1] https://openai.com/index/introducing-gpt-5-4/ [2] https://openai.com/index/introducing-gpt-5-2/

CompleteSkeptic··on PP-YOLO Surpasses YOLOv4 – State-of-the-art object detection techniques
There actually is some work (https://arxiv.org/abs/2003.13630) claiming that FLOPS are a poor measure of real-world performance - with some of the more recent FLOP-efficient models actually running slower than older models.
Page 1 of 2Next →