4 karma · joined February 11, 2026
Jev answers yes/no, picks from a few choices or gives a score, with probabilities, and it's fast and cheap. See https://docs.typesafe.ai/introduction for more details on Jev. It performs on par with small LLMs while being faster and cheaper (independent study: https://www.ayautomate.com/blog/jev-vs-llm-benchmark), and in my benchmark it was better calibrated than DeepSeek V4.1 Flash (see below).
I wanted agents to be able to use such a System 1 in the middle of a task without writing a script each time. So jevpipe is a small CLI that can be used like grep or other Unix tools: pipe lines or files in and get classifications out.
Example:
cat abstracts.jsonl | jevpipe filter "Does this paper report a randomized trial?"
cat abstracts.jsonl | jevpipe map -q '{"design": {"type": "choicsign?", "criteria": {"rct": "randomized trial", "observational":"observational study", "review": "review or meta-analysis", "other": "anything else"}}}' | jq -r .answers.design.choice | sort | uniq -c
filter prints the lines that pass, map returns typed answers as JS
basically program their own System 1 for anything and apply it to a large chunk of data for
"gut feeling" decisions.I also added an agent skill for jevpipe that teaches the agent when and how to use it.
If you are interested in technical details, I also did a benchmark because that seemed to be missing from the recent evaluations. But it is by no means coding only or just a search tool. https://github.com/fabianboth/jevpipe/blob/main/bench/README...
It needs an OpenRouter key. I'm not affiliated with TypeSafe.
I think what they want to achieve here is less "kill openclaw" or similar and more "keep our losses under control in general". And now they have a clear criteria to refer when they take action and a good bisection on whom to act on.
In case your usage is high they would block / take action. Because if you have your max subscription and not really losing them money, why should they push you (the monopoly incentive sounds wrong with the current market).
Curious how the capability tradeoff plays out in practice though. SWE-Bench Pro scores are noticeably lower than full 5.3-Codex. For quick edits and rapid prototyping that's probably fine, but I wonder where the line is where you'd rather wait 10x longer for a correct answer than get a wrong one instantly.
Also "the model was instrumental in creating itself" is doing a lot of heavy lifting as a sentence. Would love to see more details on what that actually looked like in practice beyond marketing copy.
The window to embed privacy protections into the IEEE 802.11bf standard is closing. Once this is ratified without safeguards, retrofitting privacy will be much harder.
I already had built a hook with desktop notification and window highlighting myself. But I have to admit, making it fun like this beats it by a lot.
The real insight the article misses is that AI coding tools actually widen the gap. The more senior you are, the more you accelerate. You're a better reviewer, you scope tasks correctly, you catch the nonsense faster. The friends in this post didn't fail because the tools are bad. They failed because reviewing code is harder than writing it, and you can't review what you don't understand.
The thing that surprised me most was how unreliable even basic guardrails were once you gave agents real tools. The gap between "works in a demo" and "works in production with adversarial input" is massive.
Curious how you handle the evaluation side. When someone claims a successful jailbreak, is that verified automatically or manually? Seems like auto-verification could itself be exploitable.