HNHacker News
TopNewBestAskShowJobs

ttul

10,427 karma · joined March 19, 2009

Founder of MailChannels. Defender of open communications on the internet.
submissionscomments
ttul··on Clef: Open-source decision models, and new RL fine-tuning platform
That's probably driven by their own internal need to show the model images of emails and webpages to detect phishing, despite obfuscation of the underlying HTML.
ttul··on Gemini 4 Argon
Indeed. They have this amazing model and you can’t access it in their own branded app. It’s insane.
ttul··on Dots: Always-on agents
Scary thought: AI is already directing humanity. Even when you think you’re overseeing its output, by making use of the output, it is in some material way directing you.
ttul··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
Was it a last-minute panic, or just OpenAI releasing an update when the had a bit more training under their belt to make 6.1-sol a whole lot better? Either way, I'm extremely pleased and will be giving this model a shot.
ttul··on Sonnet 5.5
Project Glasswing is by invitation only. https://www.anthropic.com/glasswing
ttul··on Sonnet 5.5
Yeah, I did say the bar is low :)

Daybreak Blue is the not the same thing as Daybreak Red, which has a more significant hurdle. I don't know anyone who has gotten access to Red.

ttul··on Sonnet 5.5
Daybreak Blue is not bad and the bar to get into OpenAI's program is reasonable.
ttul··on The problem is not AI code, but not knowing about system architecture or intent
You could say the same thing about the transition from assembler to C in the 1970s and early-1980s, albeit that was at vastly smaller scale of impact. It’s not that nobody knows _anything_. We still direct the machines, just in a different way. And if what comes out the other side satisfies our needs, does it matter what lies beneath?

Tail risks have always existed in software development. The tail risk of a bug introduced by some dev who quit five years ago is similar to the tail risk of a bug introduced by Claude six months ago. Deal with it by building better visibility into how your systems work. Demand that your agents write good documentation to accompany their code-writing.

If you’re doing it right these days, it means you’re thinking of a much bigger picture and containing downside risks as boldly as you’re expanding the frontier of upside opportunities.

ttul··on Jevmem – automatic project memory for Claude Code, built on Jev
The AI language really stands out, too. "What is automatic and what depends on the agent". A human might write, "jevmen watches your coding session with Claude Code or Codex and picks just the right moment to remember important things that you decided along the way. There are some differences in how jevmem works, depending on which coding harness you are using. The table below summarizes these differences:"

I don't know why the models were generally trained to be so brief, but it's definitely not the way anyone I know actually writes. A second pass is always a good idea to clean this stuff up. And, thankfully, the models are all pretty good at that.

ttul··on Gemini 3.8 text-to-speech
I gave 3.8 a whirl today, replacing Eleven v3 TTS in an internal application that uses TTS to provide a listening function. The Google model produces extremely expressive output. To my ears, it’s on par with Eleven v3, which was already amazing.

Nice work. It’s awesome to have these capabilities so close at hand and so trivially easy to integrate with.

ttul··on Claude discovers a novel enzyme system with CRISPR-like repeats
Here's my grok of it: Deep learning models progressively abstract a concept presented at the input by passing the input through many sequential layers () until an output layer transforms the output of the final layer into something interpretable, such as an indication of what token to predict next, or a classification, or whatever. The transformer architecture futhermore offers layers that allow different parts of the previous layer's output to sort of mix with each other in complex ways. As you get into greater levels of abstraction, the attention process is mixing very abstract concepts with each other in a nonetheless highly structured manner. I believe this is where the intelligence lives.

sometimes with residual connections, but we can ignore that for sake of simplicity.

ttul··on AI Has No Wisdom and Neither Will You
I think this kind of blog post reflects a fear of infantilization by AI. By asserting that only humans know what good code looks like, the author is attempting to mend his own ego.

Build the systems around the code and let the agents do their work.

ttul··on AI Has No Wisdom and Neither Will You
I’m taking this view now as well. If you’re reading code, you’re probably doing it wrong. You should absolutely be setting criteria that can be objectively measured and rejecting code that doesn’t meet those criteria or perform as specified. We are all senior software engineering managers now, with a fleet of cheap and ambitious young engineers doing all the authoring.

But reading code? What does that accomplish, other than to slow your dev process down enormously? Serious question.

ttul··on Cloudflare Quick Tunnels
I’m not sure your comment is load-bearing enough…
ttul··on Xiaomi Mimo 2.6 live post-training dashboard
$5 per second if my eyes don’t fool me. That’s ~$432K per day. Enough to rent 3,000 B300 nodes on Modal.
ttul··on EU chief opens door for Canada to become 'associate member'
Every "yank" I know is also pretty great. The problem is a political system that has entrenched an advantage for a party that would otherwise be consistently in the minority, reducing the ability of the center to moderate policy.

Canadians are right to be skeptical that the next presidency will be significantly better than the current one. A serious amount of goodwill has been burned and I think the average view up north is that it's time to build a very solid backup plan.

ttul··on Introducing System One Models and Jev
Well, I'm stoked to try it out. We have about a billion reasons a day to call a model like this to rid the world of spam and phishing.
ttul··on Introducing System One Models and Jev
For many day-to-day computing use cases, Jev seems far better suited than an autoregressive language model, if for no other reason than it is not wasting compute thinking about anything other than how to spit out a decision.

Do you have an architectural explainer yet for Jev or are you holding that close to your chest and letting the magic rip for now?

ttul··on Cartesian – AI 3D Modeling for Design
It's just a matter of time. Astra was trained on a new 3D modelling dataset. They will surely train it on CAD apps (if they have not already done so). I don't hold out much hope for competitors hoping to get in front of that train.
ttul··on Cloudflare AKE cuts origin HelloRetryRequests from 52% to 3.7%
If it's so easy, why hasn't the next Cloudflare-destroyer shown up yet? Anyone?
ttul··on Cloudflare AKE cuts origin HelloRetryRequests from 52% to 3.7%
That's not what I'm saying at all. I guarantee Cloudflare engineers are using coding agents every day for everything they do. The point is that discovering what you need to build comes only with the experience of operating at tremendous scale. I do not believe that someone could replicate what Cloudflare offers today just because they, say, have invested $20K/mo for 100x Claude Max plans, all of which are cranking 24x7.
ttul··on Cloudflare AKE cuts origin HelloRetryRequests from 52% to 3.7%
Cloudflare is the canonical example of why you can't vibe-code infrastructure. Knowing that this optimization was even necessary, let-alone having the ability to build it, is something that doesn't become apparent until you're operating at considerable scale. Once there, of course if you're Cloudflare, you use coding agents to build it. But outside of these temples of scale, good luck even knowing it was needed.

If you work at a SaaS of any kind, I think it's worthwhile considering what things will look like when scale is the only thing that is really defensible anymore.

ttul··on Pion, an agent designed to run any company autonomously
Most of our use of agents for automation is inside of internal processes. I think most real companies have tons of internal processes that could safely be automated today. The reason they aren't yet automated likely falls to a) lack of awareness that this is possible, and b) lack of resources to conduct the automation work.

OpenAI and Anthropic have recently hired legions of "forward-deployed engineers" specifically to help companies do this automation work. It's a solid move. And, if you look at some recent product announcements, they are also hard at work building the necessary plumbing. For instance, the OpenAI Agents API lets you, "Build and run cloud agents with the Codex harness, fully managed by OpenAI."

This kind of enterprise-ready, cloud-hosted stuff really accelerates implementation of AI workflows within large organizations. Not every company is in the tech space (not by a long shot). Slop isn't the primary concern. Accuracy and reliability is the primary concern, and beyond that, just the capacity to actually make the changes happen.

ttul··on Pion, an agent designed to run any company autonomously
A company is a collection of processes, capabilities, and resources. Many of the processes are currently run by humans, but over time, more processes will be automated - something that has been progressing for decades, but which LLMs greatly accelerated. One way to view things as they currently are (at least from my view as a tech CEO): We now use agentic LLMs every day to inform us on strategy and process implementation. And agents now run several processes, with more on the way every week.

But here's the thing: As we use agents to automate previously manual processes, we are elevating the humans to do work that is less amenable to automation. And the surface area of that work keeps expanding because the competitive market we exist in demands it of us.

To stretch an analogy, businesses are like organisms in a pond. A new nutrient (agentic LLMs) was recently added to the pond that makes business organisms more efficient and able to eat new kinds of food and explore new areas. As a result, those organisms that do the extra exploring and consuming grow much faster than their peers who do not. At the end of the day, the new nutrient will just be part of the pond and the old kind of organism will be a fossil.

ttul··on Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases
I think it’s more than what Aider gives you, because this is a semantic atlas as well as just a symbol lookup.
ttul··on Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases
A bunch of things. First off, I got Astra to build a semantic map itself. So, not using embeddings. Just Astra looking at our Helm charts and then the underlying repositories to see how the different parts of the system talk to each other and rely on each other.

It also had access to our internal docs (Confluence), JIRAs, Slack conversation history… All via MCPs. So it could dig around to its heart’s content as would a human developer trying to figure out the same problem.

I did also add a Vectorize database (the whole thing is Cloudflare hosted behind zero trust OAuth) as a second step and that can be helpful in surfacing concepts via the atlas’ MCP interface.

ttul··on Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases
We built a “code atlas” that provides the LLM with a semantically queryable map of how things connect and relate in a very large and sprawling codebase that evolved over 15 years. It tends to dramatically reduce the length of time models have to spend reading code while also making sure they are aware (within their context window) of nuances that are important that might be missed were they forced to just rely on reading the code in hundreds of repositories.

I strongly recommend trying this approach out yourself. The recipe is not rocket science. Get your coding agent to take a first cut at building the atlas itself, and then manually correct it. Once you’re happy that it got things right, put an MCP on it or a CLI or whatever. And your LLMs will know what to do from there.

ttul··on Show HN: ResolveHQ – A Helpdesk Built on Cloudflare Workers, D1, R2 and Queues
Just have Codex or whatever take the screenshots and post them to the repo. Easy peasy.
ttul··on Matt Mullenweg tells Automattic staff in Slack he's back in control after ouster
Most surprising in this article is the vibrant HDR magenta circle on Matt’s Twitter.
ttul··on Tell HN: OpenAI brings back 5 hour limit for plus and business standard users
I think Sam would be saying something pithier. And the 20x Pro plan absolutely runs at a loss, so I think he would be the last one to promote it.
Page 1 of 34Next →