75 karma · joined February 3, 2026
It synthesizes comments on “RL Environments” (https://ankitmaloo.com/rl-env/), “World Models” (https://ankitmaloo.com/world-models/) and the real reason that the “Google Game Arena” (https://blog.google/innovation-and-ai/models-and-research/go...) is so important to powering LLMs. In a sense it also relates to the notion of “taste” (https://wangcong.org/2026-01-13-personal-taste-is-the-moat.h...) and how / if it’s moat-worthiness can be eliminated by models.
I remember over hearing some normal people on the bus talking about essentially orchestrating some agent scraper to pull and summarise news from 40 different sites he identified as important which put him quite ahead of his peers. These were non-technical people orchestrating an agent workflow to make them better at work.
Though there’s not much that tickles my software brain here. But the agents are coming for us all.
This reminded me of Kairos which came up a few days ago (https://www.kairos.computer/) however I actually feel much better and more inspired at the angle OpenAI took than the angle kairos took. OpenAI’s genuinely feels like a platform for a coworker while Kairos is yet another cool landing page, yet another agent platform with X amount of data integrations. The use cases in OpenAI’s article also felt more concrete and impressive to be honest.
The fact that “as agents have gotten more capable, the opportunity gap between what models can do and what teams can actually deploy has grown.” is definitely true. I think the analogy whose source I have forgot commented that we have F1 cars driving at 60/kmh so for a lot of enterprises they are not even at the deployment limit where improving benchmarks matter. They are still at the level of not being able to provide the right info, not having the right evaluation and improvement frameworks .etc.
Using “Opening the AI Frontier” as a heading would be in really poor taste before OpenAI released their OSS models (earning their ClosedAI moniker) but I guess it’s a bit less offensive now. I think this product combined with OpenAI FDEs is going to make a lot of large industries inaccessible to startups but there may still be value in companies like Kairos watching what OpenAI does in this space and copying them.
I read an article a while ago about how “taste is a moat” (https://wangcong.org/2026-01-13-personal-taste-is-the-moat.h...) and it kind of applies here. In that article a technically correct kernel patch was rejected since it actually just re-implemented functionality htat was available elsewhere. In the tldraw repo, users seem to clone the repo, spin up claude and then make a PR without any kind of “taste” involved.
What confuses me is the fact that tldraw is actually very good for trying to get the best out of models, and indeed, internal to tldraw, models are expected to be used and the author gets value out of them. And yet, people leave sloppy unvetted PRs. This is a social issue that we didn’t really have before since it was producing code was the difficult part. Now producing code and PRs is easy the signal v.s. Noise ratio has collapsed completely and it’s just not worth it for people to actually review this stuff.
It would be better for people to leave one line issues with video demonstrations and allow the internal team to /fix them: “In a world of AI coding assistants, is code from external contributors actually valuable at all? If writing the code is the easy part, why would I want someone else to write it?”. Is code really needed to convey problems with open source repos or is it something unnecessary that we are now unshackled from? In the case of tldraw a lot of the PRs are just the result of people running claude on issues and therefore they add absolutely zero value.
The article has some really odd low level descriptions of bash orchestration which I suppose are important to illustrate how barebones it was. However I always feel it odd when we’re talking about agents that are lauded as borderline super intelligence and there is still low level bash being slung around – feels like we’re talking about things at the wrong level.
The point about writing extremely high quality tests reminds me a bit of the “hot mess theory of AI” (https://alignment.anthropic.com/2026/hot-mess-of-ai/) also made by anthropic where they essentially say that long horizon tasks are more likely to fall to incoherency than for a model to purposefully pursue incorrect results. This is phrased in the article as “Claude will work autonomously to solve whatever problem I give it. So it’s important that the task verifier is nearly perfect, otherwise Claude will solve the wrong problem”.
The author also observes something that I’ve realised after the initial joy of seeing an agent one shot a task wore off – for a 30 minute agent task, 25 minutes may be spent doing exploration of the environment. While it would be an offence to give a human unvetted model generated documentation and runbooks (I’m looking at you emoji ridden README.md files becoming more common across Show HN), models should commit things like this to memory for themselves to avoid repeatedly paying the “discovery tax” on every new action. Errors, hallucinations or changes cause the generated docs to fail create more busywork for the agent but agent time is less valuable than finite human life.
Some arguments are made about retaining focus and single-mindedness while working on AI. I think these points are important. It’s related to the article on cutting out over-eager orchestration and focusing on validation work (https://sibylline.dev/articles/2026-01-27-stop-orchestrating...). There are a few sides to this covered in the article. You should always have high value task to switch to when the agent is working (instead of scrolling tiktok, instagram,X, youtube, facebook, hackernews .etc). In my case I might try start to read some books that I have on the backburner like Ghost in the Wires. You should disable agent notifications and take control of when you return to check the model context to be less ADHD ridden when programming with agents and actually make meaningful progress on the side task since you only context switch when you are satisfied. The final one is to always have at least one agent and preferably only one agent running in the background. The idea is that always having an agent results in a slow burn of productivity improvements and a process where you can slowly improve the background agent performance. Generally, always having some agent running is a good way to stay on top of what current model capabilities are.
I also really liked the idea of overnight agents for library research, redevelopment of projects to test out new skills, tests and AGENTS.md modifications.
I was shocked to see that in the prompt for one of the landing pages the text “lavender to blue gradient” was included as if that’s something that anybody actually wants. It’s like going to the barber and saying “just make me look awful”.
This was my first time actually seeing what the GDPval benchmark looked like. Essentially they benchmark for all the artifacts that HR/finance might make or work on (onboarding documents, accounting spreadsheets, powerpoint presentations .etc). I think it’s good that models are trained to generate things like this well since people are going to use AI to do such anyway. If the middlemen passing AI ouputs around are going to be lazy I’m grateful that at least OpenAI researchers are cooking something behind the scenes.
Some of the takes in this article relate to the "Agent Native Architecture" (https://every.to/guides/agent-native), an article that I critiqued quite heavily for being AI generated. This article presents many of the concepts explored there in a real-world, pragmatic lens. In this case, the author brings up how initially they wanted their agent to invoke specific pre-made scripts but ultimately found out that letting go of the process is where the inner model intelligence was able to really shine. In this case, parity, the property whereby anything a human can do an agent can do was achieved most powerfully buy simply giving the agent a browser-use agent which cracked open the whole web for the agent to navigate through.
The gradual improvement property of agent native architectures was also directly mentioned by the article, where the author commented on giving the model more and more context allowed him to “feel the AGI”.
ClawdBot is often reduced to “just AI and cron” but that might be overly reductive in the same way that one could call it a “GPT wrapper” in the same way that one could call a laptop an “electricity wrapper”. It seems like the scheduler is a significant aspect of what makes ClawdBot so powerful. For example the author, instead of looking for sophisticated scraper apps online to monitor prices of certain items will simply ask ClawdBot something like: “Hey, monitor hotel prices” and ClawdBot will handle the rest asynchronously and communicate back with the author over slack. Any performance issues due to repeated agent invocations are ameliorated by problem context and runbooks that are automatically generated and probably cost less time than maintaining pipelines written in plain code for a single individual who wants a hands-off agent solution.
Also, the article actually explains the obsessions with Mac Mini’s which I thought was some kind of convoluted scam (though apple doesn’t need scams to sell Macs…). Essentially you need it to run a browser or multiple browsers for your agents. Unfortunately that’s the state of the modern web.
I actually have my own note taking system and a pipeline to give me an overview of all of the concepts, blogs and daily events that have happened over the past week for me to look at. But it is much more rigid than ClawdBot: 1) I can only access it from my laptop, 2) it only supports text at the moment, 3) the actions that I can take are hard coded as opposed to agent-refined and naturally occuring (e.g. tweet pipeline, lessons pipeline, youtube video pipeline), 4) there’s no intelligent scheduler logic or agent at all so I manually run the script every evening. Something like ClawdBot could replace this whole pipeline.
Long story short, I need to try this out at some point.
The point about filtering signal vs. noise in search engines can’t really be stated enough. At this point using a search engine and the conventional internet in general is an exercise in frustration. It’s simply a user hostile place – infinite cookie banners for sites that shouldn’t collect data at all, auto play advertisements, engagement farming, sites generated by AI to shill and produce a word count. You could argue that AI exacerbates this situation but you also have to agree that it is much more pleasant to ask perplexity, ChatGPT or Claude a question than to put yourself through the torture of conventional search. Introducing ads into this would completely deprive the user of a way of navigating the web in a way that actually respects their dignity.
I also agree in the sense that the current crop of AIs do feel like a space to think as opposed to a place where I am being manipulated, controlled or treated like some sheep in flock to be sheared for cash.
Possibly the most salient point in the article is the following: “for the love of god, put [...] whatever tool du jour you're using to blow up your codebase, and make sure every claim in your README, every claim in your docs (you have docs, right?), every claim on your website is 100% tested and validated. Run actual rigorous benchmarks. Set up E2E tests driven by behavioral specs. Take your users seriously enough to deliver a good experience out of the box rather than trying to use hype to drive uptake then hoping they'll provide you with free QA”.
Personally this really resonated with the absolute fatigue I feel inside when I see a new “Show HN” to a GitHub repository in the year of our lord 2026. I’ve been burned by “slop” repos so much that my I already feel the Claude emoji drivel coming and sure enough a lot of the time that’s all a repo is, the abandoned and uncared for orphan child born of a passionate one night stand with Claude Code. Not a single screenshot or demo video in sight, just plausible promises dumped into a file for end users to figure out.
As a recovering emacs user I had a pretty visceral reaction to seeing a “cons” cell in code. In all my ramblings about a post-syntax world and chasing higher and higher abstractions seeing the level of abstractions of “cons” cells, a linked list implementation detail, mingled in with the orchestration of a whole operating system just feels like a despicable and cruel joke. It’s like studying system design and hearing some peanut brain in the other room say something like “let’s use a for loop” - not the level that one should be thinking at.
Anyway, more on the actual article what he’s done is really cool and features a lot of stuff that has proven to work at the forefront of automatic programming – he has a massive test suite against all major model providers, he runs his agent against known eval suites as well.
The optimisation hierarchy in the article is also cool going from:
Agent composes atomic tools to achieve actions
Agent invokes domain specific skills
Agent flow is translated into code
Code is optimised in a lower level language
…
Assembly?
I like how it extends our normal notion of optimisation with new “agentic programs” which are nothing but model intelligence, simple but powerful tools and user intent.
Ultimately I couldn’t finish the article. “This isn't about a one-to-one mapping of UI buttons to tools—it's about achieving the same outcome”. Everytime I read contrast framing I take psychological damage and it really really hurts. Sometimes I feel like contrast framing is left in as a conspiracy by AI companies to easily make known people who are so lazy that they can’t even be bothered asking their agents to omit it. Like a watermark hidden in plain sight….
The article in fairness says that it was authored by Claude. And you can tell. When I first saw that the article was co-written by Claude I actually thought that it would be better thought out and structured. Surely, if somebody is brave enough to admit that they used AI for a blog, where the author’s voice is valued so highly, it must mean that they have cracked the code and produced something truly magnificent and they want to show it off! How wrong I was.