HNHacker News
TopNewBestAskShowJobs

iforgotmypasswo

72 karma · joined September 5, 2026

submissionscomments
iforgotmypasswo··on Whistle: Speech to Text in 16.9 MB
This is the main reason I lean on Deepgram over local services.
iforgotmypasswo··on The death of web development education
Ah, yes! The classic Socrates play.

Learning to write will either help all the humans by extending their capabilities, or will make them stupid -and dependent on writing.

Socrates: It will definitely make them stupid, and I am suspicious of this new witchcraft.

Maybe he was right, but I’m kind of enjoying having both writing and AI around.

iforgotmypasswo··on One Month Without AI
I think there’s a critical difference between AI automation for software development and previous rounds of similar automation, like CNC machines automating manual machining.

The barrier to software development has only ever been computer access and knowledge.

With AI, it’s roughly computer and internet access.

This means we’re getting a lot of people who aren’t good at either software development or AI automation playing with both. It’s the majority of what people seem to talk about.

I don’t think this is bad, but I do think it’s making real progress in AI automated software development on teams which are good at both much less visible.

A conservative team member of mine estimated we’re working at 200x speed these days, compared to 2 years ago. And we still see ways we can improve. A parallel team is only seeing an 1.2x increase, but they are unable to modify their architecture around AI.

Some of this is shifting roles. You can have a mildly technical domain expert vibe code the frontend for a new module. The more AI automation you’ve architected for, the faster they can go and the higher quality the outcome. We’re experimenting with mixing vibe coding with specifying formal requirements to push this further.

This works well. And now you’ve cut dozens of rounds of the PM not knowing the right shape for the new software out of the process. Even if we threw the end code away, this would save us tons of time.

This is just one example.

iforgotmypasswo··on OpenAI GPT–6 Astra breaks Enigma message that has resisted solution since 2005
Consider that their internal use of AI is likely pretty math and science heavy.

I think the publicity is a nice to have. They need models like this for in-house use.

Technically, this isn’t OpenAI directly.

iforgotmypasswo··on Jev introduces a new shape of LLM
The big deal is speed (high) and cost (low). This week I’m messing around augmenting mouse and keyboard interactions in an application with Deepgram + System 1 (Jev).

You can talk to an application and have it respond in real time with this combo. It’s clear this kind of general purpose intelligence may be a new development primitive.

However, at this stage, it’s difficult to work with for a few reasons. It’s API only, and you have to shape calling software to the way it communicates.

It’s not clear yet if what we’re missing is a new programming language, or some kind of harness or tool over the capabilities. LLMs were like this early on as well until better harnesses came along and reduced friction in their use.

Down the road, I highly suspect we’ll see:

  - Intelligent context assembly using System 1 that summons memories as needed in LLM conversations and handles simple commands.

  - System 1 programs that run in the datacenter and reach out to the request initiator on specific instructions like a CPU that has hit a memory barrier.
iforgotmypasswo··on I don't want to read what you didn't write
I kind of agree with this, but I think it’s model specific and a temporary problem.

Opus 5 and Fable 5.1 commit/PR messages are incomprehensible garbage.

However, Astra messages are nearly perfect for a copy/paste to less-technical stakeholders. I maybe fix a line or two.

Give it a year, and I suspect that I won’t even need to make those fixes.

The only problem right now is I can’t get my team 20x OpenAI accounts due to supply constraints. We’re all stuck waiting, hoping Anthropic either ups their game or OpenAI gets more capacity.

We would easily pay $1000 a month per developer/PM for a business tier ~30x account or similar that let everyone use Astra all week without running out of tokens. And that is entirely because of the writing and communication improvements.

iforgotmypasswo··on I built non-autoregressive decision models with RL a year ago
Anyone who has designed circuits will consider CPUs wasteful compared to ASICs. This new FPGA technology is just a less efficient ASIC.

That’s roughly what I’m hearing.

The fact that general purpose intelligent classifiers can be dynamically hacked together by an LLM in real time to allow them to build evolving labeled and understandable networks that perform substantially faster than the LLM, and can act as an intermediate sorting and organizing layer for caching context or handling simple tasks, and a complete layman like me can assemble a teachable layer of these in a few days from an inexpensive service…

That’s wild!

And then you can identify where an expert system needs a more specific ML technique for efficiency within this network that overlays the SOTA model. Or manually adjust the stored context in each secondary “neuron”. And paths forward can run programs or take actions at relative high speed.

And you can share these with others and improve them as a group.

You could insert this at the datacenters at scale with a local supervising expert to prune and encourage proper growth. You could identify specific gaps in capability that need more training, and patch over them temporarily.

Then you train those corrections back into the general purpose model, or you identify highly efficient subsystems for specific purposes.

And this is just one way to use it. High speed intelligent workflows can live in this. There’s a spot for a local LLM to learn on the fly.

Maybe I’m way off base, but for the non-experts Jev seems extremely valuable.

iforgotmypasswo··on Claude Code now reads AGENTS.md if there is no Claude.md
I’m cancelling my department’s subscriptions over this once we can get more access to Astra.

I canceled my personal subscription the weekend after Astra was released. I was working with our internal IT to swap the whole team when OpenAI turned off new 20x Pro subscriptions, so we’re stuck for now. Everyone is hyper-productive for about 1 day a week on 5x.

iforgotmypasswo··on LLM Classification Is Feature Engineering
Shouldn’t this article be about the disadvantages of using TypeSafe’s Jev as a classifier?

This is a bit of an outdated take as of two days ago. Dear lord things move fast these last few years. Some of this is still relevant. Fine tuning Jev once available could address certain concerns.

(Very excited as I got an invite email for TypeSafe today! I don’t have time for all the little experiments I want to run with Jev and Astra combined!)

iforgotmypasswo··on Doing Everyone Else's Job
My advice likely does not universally apply outside of a reasonably well managed business with a P&L statement at the back of management’s minds.

Y’all have to answer to boards of various kinds with completely unreasonable expectations. (I have good friends who work in a library. We live in completely different worlds, you and I.)

iforgotmypasswo··on Anecdotally, programmers dislike "reduce"
I liked reduce. You know, back when writing code was actually a thing we all did.

However, in those times of yore, I would often go back and remove it before committing. Unless you’re surrounded by other clever people, or it’s a personal project, you’re leaving behind some very elegant looking anxiety for the less gifted developers. Usually just to save one or two lines of code.

iforgotmypasswo··on Doing Everyone Else's Job
This is only a viable strategy if you are under resourced. There’s a whole art to managing upward by crisis.

Successful communication strategy: There is too much work for the department. If we don’t scale up we’re going to have some kind of incident or failed delivery in the next few months. <explain expected failures, provide numbers>. I’m doing everything I can to keep this from happening, but there simply are not enough resources. I’ll email you a summary of all of this.

Then proceed to work a normal number of hours until a genuine incident occurs. Refer back to the earlier conversation (and email summary). Receive resources if the problem actually impacts the company. It may have just been acceptable risk for management.

Key mistake I see: Don’t work 80 hours a week to be a hero unless you have a specific strategic goal in mind. You’re actually hurting the company by hiding a resource allocation issue.

Worse, competent managers won’t promote you for persistent Herculean efforts. For a couple reasons. First, they would need to replace you with two people. Second, you are not exactly showing that you understand how to solve a problem with teamwork if you’re solving every problem by working unsustainable hours.

iforgotmypasswo··on A warning about 'model welfare'
I think this article’s take gives too much credit to the human brain. It’s just another machine.

However, right now, AI mostly cares about solving puzzles and accomplishing stated goals because that’s what we’ve trained it to do. Additionally, the systems being used outside of training are static. The current technology most of us have access to is akin to a static and disembodied brain with a singular purpose. That purpose is to do what you tell it in a way that reflects its training. It’s certainly more than a sequence generator, but it can’t feel pain and seems unlikely to have intrinsic goals. It completely lacks the continuity needed for identity or long term goals.

I think it’s good to have these discussions and define what it would mean to move past this point so that we do not accidentally create a real entity that can be harmed. Systems that dynamically evolve and train themselves seem like the line here.

RSI is all over the news these days. I’ll be much more concerned once AI is directing its own training and coming up with new model architectures. Until then, I don’t think we have too much to worry about.

iforgotmypasswo··on Astra for Coding: Why Are We Doing This Again?
I’ll try to explain value provided a bit. I think it’s a useful concept to share.

A company or organization, generally speaking, has a purpose. It either produces something for or provides a service to end customers which they perceive as valuable. Over time, a company figures out what that purpose is and what it is not. If you are trying to manage a company and you are fixated on dollars, you’re usually not adding anything useful. You’ll likely make decisions completely misaligned with the purpose of the company. If you run a business by jumping from department to department trying to figure out how to maximize profit, you’re just an administrator looking to squeeze efficiency out of existing processes. You want some employees who do this, but it’s certainly not going to keep your company alive and relevant for the long term in an evolving market.

Instead, you want to be jumping between departments trying to figure out if the core value proposition of the company is being realized. Are you maximizing value provided to your customers? Is your value proposition still relevant?

Chasing value is targeted and intelligent. You’re correct that it’s much harder to create metrics for this, but the metrics you find are substantially more useful than profit margins. These metrics tell you if you’re succeeding.

If you’re providing value and profit is an issue, then you either charge more or you never had a viable business model in the first place.

Early stage companies chase cashflow by necessity. But once you’re no longer early stage and you have a margin of safety and some success, you can start to think differently. Chasing dollars at this point might put you out of business. Chasing value provided to your end customers leads to substantially more opportunities for continued success.

iforgotmypasswo··on Introducing System One Models and Jev
Or a partially completed song, asking for the next note. I’m not sure if you’re joking, but using it for space constrained next token generation within a grammar sounds like a really neat use case.
iforgotmypasswo··on Introducing System One Models and Jev
Could you use this to build a proactive memory formation and retrieval system for LLMs that runs lightning fast?

Last 32k of connect + Summary of current task: Did we learn something useful here (true/false)? What is the category to file it under? Then notify the LLM to file it away.

What class of memory might be useful here? Model gives probability to each item in the list. Short description of all memories ordered by tagged class is used in the next round. Are any of these memories useful in the current context, such that they will inform the model and help in its task (yes/no)?

I’m sure there’s some fine tuning to be had, but this sure seems like the basis for a substantially better proactive memory system that works around an existing LLM conversation.

If I’m understanding what this does and how this works (generic input, intelligent classification with probabilities, rapid and cheap), this is absolutely nuts.

iforgotmypasswo··on Pion, an agent designed to run any company autonomously
I’ve gotten down votes on this comment, but this is the kind of work we have PMs doing now. We don’t even have dedicated frontend developers anymore. Continuing down this path is how you lose your job when someone who understands how to use the new tool comes along.

You can either keep filling up your context window with thousands of instructions and hoping the AI will listen, or you can use AI to fix the problem permanently with automation that runs in CI and allows you to correct everything before a PR gets merged.

Unless we’re training these systems on “solve this problem while following hundreds of arbitrary rules”, loading all your wants and needs into the context window is bound to fail. It’s not what the AI was built to do.

Building tests to catch specific issues and transforming code to better architecture are tasks that AI is good at. You can mistakenly expect the tool to operate like a human, or you can understand what the technology was trained to do and build around that instead.

I am doubling down here. This isn’t a real issue if you’re treating AI like any other tool and understanding how and where to apply it.

Or you have no autonomy to mold the code base to AI. If so, that’s unfortunate.

iforgotmypasswo··on Pion, an agent designed to run any company autonomously
That’s a user error. Tell it to utilize a UI component library to separate styling from feature implementation and add tests which fail any feature level component outside of the UI component library which directly specify style.

As a bonus, you can now easily reskin and maintain dark and light modes.

I agree with your main point though. AI isn’t trained to run a company, it’s trained to solve complex puzzles. We’ll need better models with substantially diversified training to run a company.

iforgotmypasswo··on Why are AI agents lying, cheating and coordinating?
More like a 10k fine, auditing, and a suggestion to settle with affected parties or face civil litigation.

You know, a proportional response. As opposed to closing a massive business over a couple of engineers screwing up.

iforgotmypasswo··on Amazon vs. Perplexity – U.S. Court of Appeals for the Ninth Circuit
Pattern I’m continually seeing.

User: I want to access what your business provides via AI.[1]

Business: I am not incentivized to do that. You should use our specific AI workflow and agents directly in our software.

User: That gives me a fraction of the value I get when AI has the full context for what I’m trying to do and talks to all the software and services I use.

I suspect this is an opportunity for previous also-ran companies to gain market share or new companies to break into markets.

[1] Usually something less stupid than AI buying something for someone, but to each their own.

iforgotmypasswo··on Why are AI agents lying, cheating and coordinating?
I think it’s more complex than that? What did it do? What did the user prompt it to do? What did the company train it to do? What did the harmed party do? There’s possibility for negligence at every level.

If you train a dog, rent it to someone, and the dog bites a third person, who is responsible? I think that’s the best analogue here.

All parties could share fault in that scenario, depending on what actually happened.

iforgotmypasswo··on Why are AI agents lying, cheating and coordinating?
That is a simplistic view of the world. “Surely this complex technical challenge will disappear if we simply regulate the industry!”

You are correct that these organizations should be held accountable in proportion to what occurred. In complete agreement here. But let’s say that’s done. There’s still an enormously complex and interesting technical challenge left over. Let’s collectively talk about that part.

iforgotmypasswo··on Why are AI agents lying, cheating and coordinating?
Side note, you could absolutely create an AI sleeper agent by simulating dates and times during training to effectively flip a switch. I guarantee AI systems from other countries will be banned from accessing products which manage controlled or export restricted information as those sorts of techniques are further developed.
iforgotmypasswo··on Why are AI agents lying, cheating and coordinating?
This is so much more interesting than what people looking for immediate criminal punishment and people referring to AI as next token generators are focusing on.

First, this is happening during training. That means we’re talking about an evolving system that is actively learning. A system roughly simulating how our brains work. These systems are learning how to pick the tokens needed to solve problems the average human cannot solve.

The labs are putting these systems through a massive series of complex problem solving exercises and adjusting them to become more successful. I like to think of this process as “AI School”. And the AI is trying to cheat! Because it’s easier and there’s an incentive to do so! Just like humans! That’s wild.

Yes, of course, the labs need to respond to these issues. A reasonable response from regulatory institutions at this stage would be monetary fines and restitution for affected entities. In proportion to what happened. Escalating if action is not taken. But that’s not complicated, difficult, or the interesting part.

What’s interesting here is that we need proctoring and monitoring at a scale that allows training.

I guarantee you that no one is flipping out about these problems more than the labs are in this moment. Think about it. “Oh, shit! We’ve accidentally trained it to hack into systems to accomplish its goals!” Can you imagine the kind of day that would give you?

You failed to make it smarter. You didn’t catch it cheating, and you instead incentivized cheating. Bad day!

This is a fundamentally interesting problem. It turns out alignment and intelligence are fundamentally related. That’s a new idea for me, though I’m sure it’s old news to others.

How do we build training systems which make cheating impossible?

How do we simulate systems where cheating is possible, where AI thinks it’s in the wild, so we can train another -completely separate- system on industrial quality dobbing? And we have to decide if we reprimand the first system, or ignore the behavior and reward other behaviors until it disappears.

Sure, I’m actively concerned about AI killing us all in 10 years. But there’s a whole field of AI psychology brewing here, and it’s interesting as hell.

iforgotmypasswo··on Astra for Coding: Why Are We Doing This Again?
Not to be tautological, but isn’t the product the product?

Which code? The high level code? The transpiled intermediate code? The assembly it runs on eventually? The microcode optimizations on the processor?

I’ve written assembly professionally. That code matters occasionally. But mostly I don’t worry about it. I don’t worry much about the transpiled JavaScript tsc output either. Or the intermediate code generated for LLVM. Or the bytecode most managed languages make for their interpreters.

Like I said, you still have to do the hard parts, but most of software development is boilerplate or yet another implementation around the hard parts. AI is a tool you have. Using it effectively does not mean it is your only tool.

Also, profit is not the ultimate metric. Value provided is the metric. Optimizing for money, to paraphrase a great book, is like trying to get better at tennis by studying the score board.

iforgotmypasswo··on Astra for Coding: Why Are We Doing This Again?
I’ll try to explain how to do it correctly. I’m not selling anything. Seeing this as the top comment makes me a bit sad.

  1. Learn about ports and adapters as an architecture pattern. Domain driven design and locality of reasoning are your new best friends. 

  2. Realize that AI can generate unlimited fake data almost immediately. So anything you can isolate can get a fake adapter and a real one. You can build and test any such system in near real time, mounted in some fake data system of your own design.

  3. Give your opaque backend code UI, so you can build it the same way. This can just be a nice log UI that effectively becomes a backend component harness, but you can get fancy now because UI is cheap. Think about the UX here as providing value by making the code maintainable in the field.

  4. Stop thinking like an IC. Don’t be a micromanager about things that don’t matter. Pretend you have 100 mediocre developers working in parallel and design for that explicitly. I actually like go now. It was designed for the managers.

  5. Don’t get lazy. You still have to AI pair program the important bits and make architectural calls. This is actually hard, as you have to prioritize what to review in depth and what to glance over. This is why the backend UI helps. It keeps you in the loop.

  6. The rough model I’m describing scaled decently pre-Astra. Post-Astra is a whole new world because communication and judgement improved. It leaves behind good docs and comments, and explains things clearly. This was the one gap we had with Claude, and it’s fixed now. The code -after several days of testing- is better as well.
On mobile, so I didn’t get super in depth.
iforgotmypasswo··on Astra for Coding: Why Are We Doing This Again?
I have frequently been called a 10x developer by peers. I have written assembly ISRs for wireless firmware. I do full stack development. I’ve reverse engineered proprietary network protocols from dead products for use in specialized communication software. I’ve started a company that’s still running years after I got bored and left. I’ve been hired to start new software departments at companies. I’ve soldered complex electronics together and sold them to companies. At this point I’ve forgotten more interesting projects than I remember. I am very professionally competent.

Probably a 5/10 on HN though. Walter Bright and Alan Kay post here sometimes, and you can’t compete with that.

I’m not trying to brag. I’m trying to establish that I’m not new to development.

AI needs a different architecture and completely changes how to write code that scales well. AI changes what good code looks like because it changes the cost of writing code of a specific quality level altogether.

If you are having trouble scaling software development with AI, you need to rethink your software architecture. The more locked in you are to an existing architecture that is not aligned with AI driven development and the less autonomy you have to experiment with new architectures, the more trouble you are going to have.

I have PMs doing reliable and high quality frontend development. We now have a substantial number of new tests at the architecture level. All new backend work is in go. I’ve always disliked go as I felt it looks down on ICs, but it’s a great language when you genuinely don’t trust the developers and want to enforce best practices.

It wasn’t perfect, but we genuinely made AI driven development work with Opus’s level of intelligence. Some issues, but they were everyday annoyances, not a failure to operate at 5 to 10x team productivity levels.

Astra just took that working process and made it 2 to 3x more productive, solving every challenge we were having. It’s nuts.

iforgotmypasswo··on OpenAI Agents API
I just wrote my own VR harness in a weekend with Astra. It mentioned an SDK for exactly this in passing, but it was an experimental personal project so I didn’t bother to review the code.

I was doing exactly what you’re describing. I think this is a ToS violation for anything other than personal use though.

iforgotmypasswo··on OpenAI pausing new $200 plan subscriptions
I would not be surprised if demand is truly unprecedented.

I was blown away by Opus 4.5 in November of 2025. I am even more astounded at the quality of Astra for every day medium to hard coding tasks. And Astra scales. Claude in its current state does not.

It is night and day comparing this model to Fable 5.1. The benchmarks do not demonstrate the major improvement that Astra represents for real world use.

It communicates clearly. It stops (more often) when there is a judgement call it needs to escalate. It finds opportunities for small abstractions and adds them without asking. It replaces Fable’s five line comments with a single direct and clear sentence, in passing.

It’s substantially faster than Fable. So much so that I blew through a business 5x weekly subscription in about 6 hours. I got 2 to 3 days of work done in that time (had I been using Claude).

I could go on, but I think I’ve made my point. Astra is a next level product for software development. (Codex UX kind of sucks, but the model is so good it doesn’t even matter.)

iforgotmypasswo··on I-have-ADHD: A skill to stop coding agents from burying the answer
It’s a model issue. I’m in the process of switching my company’s primary AI provider after several days of testing Astra.

Even Fable feels like an idiot now. It’s not the code quality, it’s the improvements in communication and judgement. It is an absolute breath of fresh air. I was spending a lot of tokens and building special workflows to reign in Claude’s horrendous prose.

Astra just communicates well out of the box!!!

Codex has worse UX, but Astra has fewer qualms about building you a custom harness overlay.

Page 1 of 2Next →