HNHacker News
TopNewBestAskShowJobs

dom96

7,203 karma · joined January 2, 2011

Currently working as a Software Engineer at Cloudflare on Workers (and particularly on Python Workers). Previously at Meta.

Since 2010 I have been involved with building the Nim programming language, as well as many tools and libraries for it. I have also published a book about Nim called Nim in Action.

These days I'm working on various personal projects- some big, some small. Most recently https://onlyhumanhub.com, https://listifications.app and https://mousetrack.co.uk

Github: http://github.com/dom96

Bluesky: https://bsky.app/profile/dom96.picheta.me

Twitter: http://twitter.com/d0m96

Mastodon: https://mastodon.social/@dom96

Website: http://picheta.me, my contact info is here, always happy to receive emails.

submissionscomments
dom96··on Gemini 4 Argon
I built a brand new language to test this[1]. Not only is the language different to basically any other language but it also tries to be adversarial against LLM understanding.

The best models can still make sense of it[2], though the tasks so far have been pretty basic. But I do think it gives some evidence that languages which aren’t well represented in an LLM’s training can still be reasoned about and written well by LLMs.

1 - https://killswitch-lang.org

2 - https://bench.killswitch-lang.org

dom96··on Gemini 4 Argon (High): Intelligence, Performance and Price Analysis
I'd love to run it on my benchmark but alas, Google not making it public prevents this.
dom96··on Gemini 4 Argon
Why announce this if it’s not available yet? Why not at least announce when it will be released to the public?

None of the other AI labs do this. Really frustrating.

dom96··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
Well, it is genuine.

Opus 5.5 fails the "understanding" tasks which Opus 5 passes. I feed it a script which takes two numbers and prints the max of the two numbers. Opus 5.5 thinks it prints 1/0 instead of the max numbers. Opus 5 gets it right.

Here are the outputs from both: https://gist.github.com/dom96/b5bce82b6e6c1ebd5271ed70ad941b....

Looking at that Opus 5.5 fails to deduce that the "hack statement" is actually an if statement in disguise, but Opus 5 gets this right. I feel like this is a pretty good test and shows Opus 5's greater intelligence.

dom96··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
Based on my benchmark[1] it is the same price as Opus 5.5 and just as capable.

1 - https://bench.killswitch-lang.org

dom96··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
Surprisingly (or maybe not) it matches the performance of Astra on my benchmark[1], but is much cheaper. It is also head to head with Opus 5.5 on both the price and pass rate, but edges it out slightly.

1 - https://bench.killswitch-lang.org/

dom96··on Sonnet 5.5
Funnily enough the Fable safeguards are the worst and testing Sonnet 5.5 didn't trigger them as much as it even did for Opus on my benchmarks[1].

1 - https://bench.killswitch-lang.org

dom96··on Sonnet 5.5
I built an adversarial esoteric programming language to benchmark LLM models and just ran it on Sonnet 5.5 It does worse than Sonnet 5. Mainly because it is more reluctant to keep going to get an answer, instead it returns to ask the user questions whether to keep going.

https://bench.killswitch-lang.org/

    Claude Sonnet 5    17.8%
    Claude Sonnet 5.5  7.4%
dom96··on Evolving programming languages in the AI era
There are multiple phrases that can be used[1]

1 - https://github.com/dom96/KillSwitch/blob/main/SPEC.md#abrupt...

dom96··on Evolving programming languages in the AI era
Agents are fairly good even at languages designed to trick them. I built one[1] and it does make for a good benchmark[2] to see which LLMs are actually good. I think that a language which is largely similar to others will be a piece of cake for most and any advantage that an existing language will have will be minor enough to not matter.

Btw if folks have ideas of how to make Killswitch even harder for LLMs I’d appreciate them.

1 - https://killswitch-lang.org

2 - https://bench.killswitch-lang.org

dom96··on Grok 4.7
The Democrats are not a left-wing party, I'm not sure why you think "the Left" would be in love with him after all the other horrible things he's done if he just gave money to a political party that is at best centrist.

If you want him to get the benefit of the doubt about his "hand gesture" then it would help if he wasn't promoting far-right parties all over Europe.

dom96··on GPT-6 Sol and Luna
It's surprising but MiMo V2.6 Pro performs better and is cheaper than GPT 6 Sol on my benchmark[1]. Open weight models are really snapping at the heels of the major western models.

1 - https://bench.killswitch-lang.org

dom96··on Claude Opus 5.5
Just updated KillSwitch-Bench with this new model: https://bench.killswitch-lang.org/

It does perform slightly worse than Opus 5, but it is significantly cheaper and faster.

dom96··on MiMo-v2.6-Pro: Intelligence, Performance and Price Analysis
Not yet. But can do so if there is interest.
dom96··on MiMo-v2.6-Pro: Intelligence, Performance and Price Analysis
I’m using openrouter which I think is a fair representation of what the typical user will experience.
dom96··on MiMo-v2.6-Pro: Intelligence, Performance and Price Analysis
It is an impressive model. Agreed on most that is written on this page, with the exception of it being fast. I ran it on my own LLM benchmark suite[1] and it is faster than DeepSeek but still much slower than leading models. But it's pricing is where it really shines.

KillSwitch-Bench 1.0

  Claude Opus 5           66.9
  GPT-6 Astra             57.9
  Claude Fable 5.1        46.7
  MiMo-V2.6-Pro           38.8
  Muse Spark 1.3          36.5
1 - https://bench.killswitch-lang.org/
dom96··on MiMo v2.6
How do you think other benchmarks work? Every single one is going to have a budget and time limit. There has to be a cap on those, you can't just let it spin forever and use unlimited funds.

I would argue that this benchmark is uniquely suited to how most people use LLMs because it actually tests common harnesses and it is a true coding benchmark for a language that is unseen, thus testing the LLMs actual ability to understand nuance and learn.

Internet access is restricted to ensure that over time models cannot just look up the source code of KillSwitch, which would allow them to cheat.

dom96··on MiMo v2.6
Very capable model. I just ran it on my own LLM benchmark suite[1] and it matches Muse Spark 1.3 in pass rate but is significantly cheaper.

KillSwitch-Bench 1.0

  Claude Opus 5           66.9
  GPT-6 Astra             57.9
  Claude Fable 5.1        46.7
  MiMo-V2.6-Pro           38.8
  Muse Spark 1.3          36.5
1 - https://bench.killswitch-lang.org/
dom96··on MiMo v2.6
Why not? In my own benchmark Opus 5 does in fact come out on top[1]

1 - https://bench.killswitch-lang.org/

dom96··on Python Workers are now generally available
(I'm one of the authors of this post)

> Being tied to use only one version of Python/Pyodide (the one that Workerd embeds)

This isn't quite the case, you can choose between different versions using compatibility flags. For example, `python_workers_314` is the compat flag for Python 3.14[1]. You've also got compat flags for 3.13 and 3.12. Though it is worth noting that by using those older versions you will also be using older Pyodide versions too, which have fewer features (for example they lack JSPI support).

> Architecturally tied to the JS/v8 world, which may show some challenges as they aim to reduce cold start times

That is indeed a challenge. But our memory snapshot implementation has improved the cold starts significantly already and we will be working to reduce these even further. We also have sharding these days which reduces cold start frequency a lot. We wrote about cold starts (and sharding) in a previous blog post[2] which includes some numbers.

1 - https://developers.cloudflare.com/workers/configuration/comp...

2 - https://blog.cloudflare.com/python-workers-advancements/

dom96··on Python Workers are now generally available
Gyeongjae, who is one of the authors too, is also a Pyodide core developer. Definitely couldn't have got Python Workers this far without both of their help.
dom96··on Python Workers are now generally available
It looks like that link is incorrect, it should be pointing at https://cold.picheta.me/#packages.
dom96··on Grok 4.7
It’s a shame this model has such negative political baggage associated with it. It’s the only one I decided not to run in my LLM benchmarks[1].

1 - https://bench.killswitch-lang.org

dom96··on Python Workers are now generally available
We’ve got some cold start numbers in our previous blog post: https://blog.cloudflare.com/python-workers-advancements/.

There is still more to do here and it’s what we will be focusing on next.

dom96··on Show HN: KillSwitch – a programming language designed to be difficult for LLMs
Glad you like it :)
dom96··on Show HN: KillSwitch – a programming language designed to be difficult for LLMs
In case you missed the link on the page, the benchmarks themselves are pretty interesting too and give a good indication of the performance of different models: https://bench.killswitch-lang.org/.
dom96··on Apple Reference Image: A New Approach for Verified Photography
How is this different to Content Credential and why isn’t Apple implementing that?
dom96··on Steam Frame starts at $1059
It's pretty good, but it would have been good enough as a non-VR title. The thought of digging out my VR headset to play it makes me shudder. I never even managed to finish it, played through most of it and then couldn't bring myself to get my VR headset out again.
dom96··on Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires
This is great. I've been building my own model benchmark lately and it has indeed been so easy to mess up the scoring. It's simply much harder to come up with an algorithm that combines all your individual scores into something that isn't broken in some special circumstances. That's why I think many just start capping the results.
dom96··on Anthropic boss Dario Amodei calls for AI development to slow down
This feels like all the labs have reached the peak of what is possible with the LLM architecture and just need an excuse to spend the next decade finding the next big jump in intelligence.
Page 1 of 34Next →