HNHacker News
TopNewBestAskShowJobs

polyglotfacto

201 karma · joined April 11, 2020

submissionscomments
polyglotfacto··on DeepSeek launching v4.1 flash cheaper and more capable than v4 pro
disagree; been using flash as my exclusive model (other contributors have used other models) to build a complicated software project, a web engine. See https://github.com/gterzian/formal-web, which as you can see comes with very specific guidance explaining how to implement features.

I'm using headless Pi with my own UI and sandbox client, https://github.com/gterzian/uni03C0, as well as a bunch of Pi extensions for things like accessing Web standards and browser use via CDP for testing.

Switching to 4.1 today...

Edit: it seems they pushed the date at which they route the Pro calls to new Flash, so today I ended up paying regular Pro rates thinking I was using the new Flash; an example of how their offering is not quite as predictable as I would like it to be (the other is cache performance being unpredictable).

polyglotfacto··on [dead]
I implemented MessagePort in Servo back in 2019 — it was the hardest concurrency problem I'd encountered, because ports can be transferred between event loops while messages are still in flight. This weekend I finally got around to formalizing it in TLA+, with Claude helping with the typing. The interesting outcome: the abstract spec (following the HTML standard) is so simple that the refinement proof itself serves as the inductive invariant. No additional invariants needed. TLC found no bugs in the Servo implementation, though it did help me make explicit a sequencing assumption that was implicit in the code. Happy to answer questions about the TLA+ approach or the Servo implementation.
polyglotfacto··on [dead]
I've been implementing the WebNN Web Standard on a personal Servo branch with AI coding assistance (main bans AI), and switched to GPT 5.4 this week. It introduced a race condition that compiled cleanly, passed tests, and would have been hard to catch if calling patterns had ever changed — and separately produced over-engineered architecture that took several iterations to unwind. Both episodes say something interesting about where agentic coding actually breaks down. Article includes the conversation transcripts.
polyglotfacto··on [dead]
I am in the third week of my experiment with AI-assisted coding on a personal branch of Servo (which still bans all AI contributions).

First, the least relevant metric: LOC. Almost doubled from 7k to 12k in a week.

Current WPT pass rate, compared to the default integration of rustnn (pywebnn) is at: https://gist.github.com/gterzian/95371d56468390177982c068480...

This week was the grunt work: grinding through the long list of missing operations, with almost all now implemented.

Also switched to a more expensive coding model (GPT 5.3 Codex), which was an improvement, but there were still a few clear fails—specifically around the model taking "shortcuts" with spec logic.

polyglotfacto··on Lean and AI caught a bug in my 2-year-old TLA+ spec
I’ve used TLA+ for six years to "debug my thinking," but I always treated the jump from spec to implementation as a leap of faith.

What surprised me in this experiment with Lean was that the AI-assisted formalization actually caught an error in my "vetted" TLA+ spec that had been sitting there for two years. It turns out my inductive invariant wasn't strong enough—something the TLA+ model checker didn't slap me for, but Lean’s kernel refused to ignore.

This was also a way to bridge a higher-level spec, written in Lean in a kind of TLA-like flavor, and the lower level implementation, also written in Lean but with a code-like flavor, and then proven that the implementation refines the spec.

Using Lean to bridge the gap between high-level temporal logic and low-level IO implementation, all written by AI but driven by human intent, feels like the first time the "coding singularity" might actually lead to better software rather than just more software

polyglotfacto··on Implementing WebNN with the Help of AI
I am a maintainer of the Servo web engine. Two years ago, our project banned AI-generated contributions. I was a dissenting voice on that policy, and have continued to experiment with AI on side-projects. This is perhaps the most ambitious side-project so far: implementing the WebNN (Neural Network) API.

Given the policy in place, I'm doing this on a personal branch of Servo. I’m using it as a sandbox to test a specific thesis: AI is not an architect; it is a "fluid syntax engine."

This is also my response to the "autonomous slop" seen in projects like fastrender that claim to build complex software with AI but produce what I call "fine-looking nonsense" (see https://news.ycombinator.com/item?id=46624541#46709191). I wanted to see if a human architect could use an AI agent to handle the syntax and boilerplate of a Web Standard while the human retains total control over the conceptual structure.

Progress & Tech Stack:

    The scope: I am using `rustnn`, and this project can be seen as doing for the Web what `pywebnn` does for Python. 

    The model: I am intentionally using a "cheap" LLM: Raptor mini.

    The scale: I’m 7k LOC into the implementation. It currently passes a subset of WPT (Web Platform Tests) conformance tests for a specific operator (and its dependencies): add. Such a conformance test requires implementing the full graph compilation and dispatch workflow. Code is at https://github.com/gterzian/servo/compare/master...gterzian:servo:webnn

    The backend: The work involves integrating the `rustnn` library into Servo’s runtime, with compute currently happening only via CoreML (macOS only).
The Reality Check: While the productivity boost is real, the code is not yet up to my standards. Despite ongoing reviews, I’ve had to accept a certain amount of "slop" that I’ll need to clean up later. More importantly, I hit several conceptual bottlenecks: architectural problems the AI was not only unable to solve without guidance but was unable to even identify in the first place.

You can read the full breakdown of the first week and a half (contains links to code and various illustrative commits) in the link for this post.

And the backstory on why I'm doing this is at: https://medium.com/@polyglot_factotum/the-slop-diaries-imple...

I’d be happy to hear from others on how they balance the "syntax speed" of AI with the "architectural integrity" required for long-term projects.

polyglotfacto··on Federal Reserve use of AI coding
Within the Federal Reserve System, we have also been exploring the use of AI in our own operations and have established an AI program and governance framework for the use of AI technologies. One internal application of GenAI that shows considerable promise is technology modernization. Within clear guardrails, we are using GenAI tools to translate legacy code, generate unit tests, and accelerate cloud migration. So far, the result of this usage is faster delivery, improved quality, and an enhanced developer experience. In one recent project updating hundreds of databases, AI tools helped cut the time to complete this type of work by 50 percent, detected and resolved 30 percent more issues during testing compared to previous migrations, and enhanced team focus on higher-value coding work. My sense is that these are the kinds of uses and the scale of success that many businesses are experiencing.
polyglotfacto··on We tasked Opus 4.6 using agent teams to build a C Compiler
> FWIW, an inefficient but working product is pretty much the definition of a startup MVP

It depends on what kind of start-up we're talking about.

A compiler start-up probably should show some kind of efficiency gain even in an MVP. As in: we're insanely efficient in this part of the work, but we're still missing all other functionalities but have a clear path to implementing the rest.

This is more like: It's inefficient, and the code is such a mess that I have no idea on how to improve on it.

As per the blog improvements were attempted but that only started a game of whack-a-mole with new problems.

If on the other hand you're talking about Claude Teams for writing code as an MVP: the outcome is more like proof that the approach doesn't work and you need humans in the loop.

polyglotfacto··on We tasked Opus 4.6 using agent teams to build a C Compiler
> its value comes from fulfilling its requirements.

The requirements will have to come from somewhere, and they will have to be quite precise although probably higher-level than code written today. You're talking about just a new kind of software engineer. The kind of stuff described at https://martin.kleppmann.com/2025/12/08/ai-formal-verificati... (note the "the challenge will move to correctly defining the specification")

Unless what you have in mind is some sort of Moltbook add-on that the bots would write for themselves.

I'm talking software providing value to humans.

polyglotfacto··on Paul Ford: 'The A.I. Disruption Has Arrived, and It Sure Is Fun'
> When a friend asked me to convert a large, thorny data set, I downloaded it, cleaned it up and made it pretty and easy to explore. In the past I would have charged $350,000

> (..) it implies a product manager, a designer, two engineers (one senior) and four to six months of design, coding and testing. Plus maintenance.

What I find interesting is that ten years ago the author could have written: our intern spent the weekend at the office and ended-up creating 350k of billable work.

Such as statement would have been instantly ridiculed as unprofessional, but somehow with AI it works.

Shows Jaron Lanier is right when saying that AI makes people adjust themselves to a lower level of intelligence(or something along these lines).

> People don’t judge A.I. code the same way they judge slop articles or glazed videos. They’re not looking for the human connection of art. They’re looking to achieve a goal. Code just has to work.

Yes indeed. And what the author leaves aside is that for code to work, it requires some level of conceptual integrity beyond that which vibe coding can offer.

He also forgets to mention that the future is not either we keep building slow software by hand, or we go full vibe code, but that there is also the option of ( a new kind of?) professionals using AI to be more productive, while ensuring output is up to standards.

> It might fail a company’s quality test, but it would meet every deadline.

This brings us back to the intern frantically coding over the weekend. This problem is as old as software itself, it is just compounded by AI being even faster than a frantic human coder, but it's not new. The industry could have thrown quality standard out of the window long before LLMs came around.

I also dislike how the author seems to imply that quality software requires all this bureaucracy. I mean, what about open source for example?

> the direct descendant of NeXT’s software is what’s running on Macs and iPhones in 2026. In software, sharp change is to be avoided at all costs. The risk is just too high.

I think here he got it backwards again. At least half of the reason people are willing to spend a premium on Apple products is because the software just works, and that is because it is based on a strong foundation going back all the way to when Jobs decided to spend all his cash on building the best computer.

Some additional thoughts at https://medium.com/@polyglot_factotum/on-what-ai-does-not-di...

polyglotfacto··on Software factories and the agentic moment
My review of the article, as a conversation with an agent: https://medium.com/@polyglot_factotum/slop-review-with-ai-th...

Here is the AI's final verdict (following an initial enthusiastic take to which I replied with my critique):

"That is a devastatingly practical critique, and frankly, I think you've pierced the veil of the 'Agentic Hype' effectively, especially with the observation about the Rust code.

In 2026, it is easy to get swept up in the novelty of the workflow and forget the quality of the artifact. Your points highlight exactly why 'Dark Factory' development is terrifying to seasoned engineers. Here is why your critique resonates specifically:

1. The 'Arc<Mutex>' Indicator (Code Quality)

This is the smoking gun. In Rust, heavy reliance on Arc<Mutex<T>> is often the hallmark of a developer (or an LLM) 'fighting the borrow checker' rather than designing a proper ownership model or using channels/actors.

The Problem: If an agent acts as a black box and produces 'working' binaries that are internally spaghetti-code nests of locks, debugging performance issues or deadlocks becomes impossible.

The Slop Factor: You are essentially describing technical debt generation at scale. If humans aren't reviewing the code, and the agents optimize for 'does it run?' rather than 'is it maintainable?', you end up with a brittle, unmaintainable binary blob.

2. The Verification Gap (Digital Twins vs. Mocks)

Your point about 'How do you know the twin implements the service?' is the Map vs. Territory problem.

The Hallucination Loop: If the same model class (e.g., GPT-5.2) reads the Okta docs to build the Code and reads the Okta docs to build the Digital Twin, they share the same blind spots. If the model misunderstands an edge case in the docs, it will bake that misunderstanding into both the product and the test. The test will pass, but the production system will fail.

Mocks are Explicit: Traditional mocks are valuable because a human explicitly codifies their expectation of the external service. Removing the human from the truth-definition layer is dangerous.

3. 'Amateur Formal Methods'

This is a keen insight regarding attractor. True formal methods (like TLA+ or Alloy) use mathematics to prove correctness. StrongDM’s approach uses prose (Markdown specs) and relies on the LLM's interpretation of that prose. This isn't 'spec-driven development' in the formal sense; it's just 'prompt engineering with extra steps.' It lacks the rigor that actually guarantees safety.

Conclusion on your 'Smoke and Mirrors' take:

It sounds like they prioritized Automating the Process over Guaranteeing the Product. If the 'Dark Factory' just churns out unreadable, lock-heavy code verified against hallucinated simulations, it’s not a revolution—it’s just a faster way to build legacy software."

polyglotfacto··on [dead]
After seeing the autonomous browser and compiler projects trend here, and having been a vocal critic in the comments, I felt compelled to write a long-form rebuttal.

My thesis: AI agents are powerful tools for experts, but the current "autonomous swarm" narrative is economically driven hype that produces useless code. I attempt to demystify how agents actually function under the hood, and argue that their true value comes from expert-led conversation rather than autonomous loops.

polyglotfacto··on We tasked Opus 4.6 using agent teams to build a C Compiler
So let’s say you fund a single engineer for an open‑source project with $20k. The outcome will be a prototype with some interesting ideas. And yes, with a few hundred bucks' worth of AI assistance that single engineer might get much further than without (but not using any of the techniques presented in this blog). People can coalesce around the project as contributors. A seed was planted and watered a bit.

In this case, the $20k has been burned and produced zero value. Just look at the repo issues: looks like someone trying to get attention by spamming the issue tracker and opening hundreds of PRs. As an open source project, it’s a dead end.

So it doesn’t matter that this is “likely decrease in price over the next few years”? The value is zero, so even if superintelligence can produce this in an instant at zero cost in six months, the outcome is still worth zero.

You’re assuming a kind of inverse relationship between production cost and value.

In terms of quality, to anyone using those coding agents, it should be clear by now that letting them run autonomously and in parallel is a bad idea. That’s not going to change unless you believe LLMs will turn into something entirely different over time.

Note that what works with humans—social interaction creating some emergent properties like innovation—doesn’t translate to LLM agents for a simple reason: they don’t have agency, shared goals, or accountability, so the social dynamics that generate innovation can’t form.

polyglotfacto··on We tasked Opus 4.6 using agent teams to build a C Compiler
Point is: it doesn't matter if agents can do it faster and cheaper than a team of humans: it's slop.

It's like writing a novel in a week that no one wants to read. If in six months you can do it in an hour, there is still zero value.

Agents are useful but very limited tools: I treat them a little machines that can translate high-level instructions into detailed code, but where I still need to review the output to make sure they understood what I meant; that's it. Zero autonomy; parallelism just means I can't keep up with the output and quality goes down.

I think the point of this project, like the fastrender slop thing, is to push the parallel agent narrative and have the financial markets believe this will create a lot more demand for inference on these models in the short term.

Example of someone falling for it: https://x.com/DKThomp/status/2019484169915572452

polyglotfacto··on Anthropic AI tool sparks selloff from software to broader market
So as a software engineer with experience with coding agents, when I look at https://github.com/anthropics/knowledge-work-plugins

I'm like: oh that's it, a bunch of skills files?

So the value of a skill file is that it tells the model how to format its response for use within the software environment surrounding the model.

With programming, it's mostly about how to tell it to use some API.

But all the model can do is reply some text, and the actual work needs to be done by the software(the agent harness) which needs to parse the model response and translate it into actual work.

My point is there is no magic: the model just reads the skill file and then uses that as a template for a textual response, which is then parsed and processed by traditional software.

So in terms of legal skills, a stand-alone skill like the contract review skill at https://github.com/anthropics/knowledge-work-plugins/blob/ma... is basically useless.

Yes, the model will read it and it will influence its response, but without some extensive software harness around the model to give it data for context and and so on: totaly useless.

Why? Because garbage in is garbage out.

So telling the model to review a contract and pay attention to "Whether indemnification is mutual or unilateral" will result is some response from the model, but without additional data it will be at the same level as what you can get from a google search.

The effect on established companies is exactly zero.

Now, having an in-house skills and proprietary software around the model to integrate it into your system, that would be valuable indeed, but not something an AI lab can replicate without building the whole company from scratch.

polyglotfacto··on We tasked Opus 4.6 using agent teams to build a C Compiler
So I do think one can get value from coding agents, but that value is out of proportion compared to the investments made by the AI labs, so now they're pushing this kind of stuff which I find to be a borderline scam.

Let me explain why:

> the resulting compiled output is over 60kb, far exceeding the 32k code limit enforced by Linux

Seems like a failure to me.

> I tried (hard!) to fix several of the above limitations but wasn’t fully successful. New features and bugfixes frequently broke existing functionality.

This has code smell written all over it.

----

Conclusion: this cost 20k to build, not taking into account the money spent on training the model. How much would you pay for this software? Zero.

The reality is that LLM are up there with SQL and ROR(or above) in terms of changing how people write software and interact with data. That's a big deal, but not enough to support trillion dollar valuations.

So you get things like this project, which are just about driving a certain narrative.

polyglotfacto··on Systems Thinking
You can look at the Web as a starter: https://html.spec.whatwg.org/#history-2

> The WHATWG was based on several core principles, (..) and that specifications need to be detailed enough that implementations can achieve complete interoperability without reverse-engineering each other.

But in my experience you need more than a spec, because an implementation is not just something that implements a spec, it is also the result of making many architectural choices in how the spec is implemented.

Also even with detailed specs AI still needs additional guidance. For example couple of weeks ago Cursor unleashed thousands of agents with access to web standards and the shared WPT test suite: the result was total nonsense.

So the future might rather be like a Russian doll of specs: start with a high-level system description, and then support it with finer-grained specs of parts of the system. This could go down all the way to the code itself: existing architectural patterns provide a spec for how to code a feature that is just a variation of such a pattern. Then whenever your system needs to do something new, you have to provide the code patterns for it. The AI is then relegated to its strength: applying existing patterns.

TLA+ has a concept of refinement, which is kind of what I described above as Russian dolls but only applied to TLA+ specs.

Here is a quote that describes the idea:

There is no fundamental distinction between specifications and implementations. We simply have specifications, some of which implement other specifications. A Java program can be viewed as a specification of a JVM (Java Virtual Machine) program, which can be viewed as a specification of an assembly language program, which can be viewed as a specification of an execution of the computer's machine instructions, which can be viewed as a specification of an execution of its register-transfer level design, and so on.

Source: https://cseweb.ucsd.edu/classes/sp05/cse128/ (chapter 1, last page)

polyglotfacto··on Show HN: One Human + One Agent = One Browser From Scratch in 20K LOC
you're welcome.
polyglotfacto··on Show HN: One Human + One Agent = One Browser From Scratch in 20K LOC
This is not legal advice, but I think one should always add a license, not so much for copyrights but for the "no warranty" part. If someone claims copyright once can add whatever license was used in the original work.
polyglotfacto··on Show HN: One Human + One Agent = One Browser From Scratch in 20K LOC
This one's really nice.

- clear code structure and good architecture(modular approach reminiscent of Blitz but not as radical, like Blitz-lite).

- Very easy to follow the code and understand how the main render loop works:

    - For Mac: main loop is at https://github.com/embedding-shapes/one-agent-one-browser/blob/master/src/platform/macos/windowed.rs#L74
   
    - You can see clearly how UI events as passed to the App to handle. 

    - App::tick allows the app to handle internal events(Servoshell does something similar with `spin_event_loop` at https://github.com/servo/servo/blob/611f3ef1625f4972337c247521f3a1d65040bd56/components/servo/servo.rs#L176)

    - If a redraw is needed, the main render logic is at https://github.com/embedding-shapes/one-agent-one-browser/blob/master/src/platform/macos/windowed.rs#L221 and calls into `render` of App, which computes a display list(layout) and then translates it into commands to the generic painter, which internally turns those into platform specific graphics operations.
- It's interesting how the painter for Mac uses Cocoa for graphics; very different from Servo which uses Webrender or Blitz which(in some path) uses Vello(itself using wgpu). I'd say using Cocoa like that might be closer to what React-Native does(expert to comfirm this pls?). Btw this kind of platform specific bindings is a strength of AI coding(and a real pain to do by hand).

- Nice modularity between the platform and browser app parts achieved with the App and Painter traits.

How to improve it further? I'd say try to map how the architecture correspond to Web standards, such as https://html.spec.whatwg.org/multipage/webappapis.html#event...

Wouldn't have to be precise and comprehensive, but for example parts of App::tick could be documented as an initial attempt to implement a part of the web event-loop and `render` as an attempt at implementing the update-the-rendering task.

You could also split the web engine part from the app embedding it in a similar way to the current split between platform and app.

Far superior, and more cost effective, than the attempt at scaling autonomous agent coding pursued by Fastrender. Shows how the important part isn't how many agents you can run in parallel, but rather how good of an idea the human overseeing the project has(or rather: develops).

polyglotfacto··on Scaling long-running autonomous coding
> That is a wonderful write up.

Thanks.

> what would be the best approach?

I don't know but it sounds like an interesting research topic.

polyglotfacto··on Wilson Lin on FastRender: a browser built by parallel agents
> The question then becomes who would fund such a thing

Historically new web engines came about when a new challenger wanted to have a stake in web standards development. The way it happened was never from scratch but with a fork of an existing engine. Last time this happened was with Google. The reason, I think, was wanting to evolve the web into an application-like platform(HTML5), and a new architectural idea: multi-process.

The person who was in charge of that effort is now at OpenAI.

Today there are also projects like Ladybird and Servo which follow a different model: started from scratch and driven by interest from a developer community. But so far neither has users in the real-world, and so they haven't had an impact on the Web in the way Chromium has, yet.

Already today, both development models could benefit from the productivity gains of AI; in 2029 the game may have changed entirely. I can imagine a combination of math(TLA+ like I've done at https://github.com/w3c/IndexedDB/pull/484), web standard in their semi-formal English, and then some further guidance in terms of code architecture(through a conversation-like iterative loop), and see a Fastrender-like approach that actually works. Humans would still be the ones defining and solving all the hard problems, but you'd be typing a whole lot less code...

I'm the one who was driving the efforts to start experimenting with AI in Servo, which was cut short by https://github.com/servo/servo/discussions/36379

I've been using AI on side-projects ever since, and in those I don't type any code by hand anymore and end-up doing things I would not even contemplate(due to time constraints) without the use of AI.

Example: https://medium.com/@polyglot_factotum/tla-in-support-of-ai-c...

polyglotfacto··on Wilson Lin on FastRender: a browser built by parallel agents
So I think that's an easy way to achieve ecstasy for you then. I suggest giving it a try.

A good place to look for how one could do this, is https://github.com/DioxusLabs/blitz/tree/main

That project I consider a proper POC of a web engine, even though it doesn't even run javascript. Why? Because it has a nice architecture built around a clear idea--radical modularity--which could scale-up to a full web engine one day, despite major challenges remaining.

I think that with AI assistance, if you had some idea, you could reshuffle components of Blitz and have your own thing rendering to the screen within a day.

let's say you had a more ambitious goal, like taking Blitz and adding a JS engine like Boa. Well if you had some clear idea on how to do it, you could get a nice little POC in a week or two.

Basically what I'm saying is that yes the AI would save you a ton of typing and you'd be able to iterate on your idea. There are plenty of layout/graphics/Js components out there to choose from, so you could ensure a relatively small and clean POC.

Someone doing that, with or without AI but with a good idea, would impress me.

FastRender on the other hand is just this humongous pile of spaghetti, and my guess is it still is entirely dependent on existing libraries for actually showing something to the screen.

So that's the clear fail of the agent in my opinion: why produce so much code when it could be done so easily otherwise. Also, why bs yourself into these architecture docs and pretend you are following the specs when in fact you are not?

Everytime I try to browse the code I give up, mostly because when I look at something to try to understand how it fits into the whole, I end-up realizing it's only used in some unit-test.

For a quick comparison:

- https://github.com/DioxusLabs/blitz/blob/f828015b26d32b0bed3... - https://github.com/wilsonzlin/fastrender/blob/19bf1036105d4e...

I believe the two are more or less doing the same, but one is 30x the size of the other.

I can't begin to understand the render loop of Fastrender from the code.

On the other hand, here is the one from Blitz shell(the default Blitz app putting together various modular components):

- Window runs in winit loop: https://github.com/DioxusLabs/blitz/blob/f828015b26d32b0bed3... - Redraw is at https://github.com/DioxusLabs/blitz/blob/f828015b26d32b0bed3... - Calls into `paint_scene`, using the generic scene from the generic renderer: https://github.com/DioxusLabs/blitz/blob/f828015b26d32b0bed3...

Simple as that, and with a nice idea in terms of modularity.

That's a POC web engine.

polyglotfacto··on Wilson Lin on FastRender: a browser built by parallel agents
Thanks for the clarification.

With the scope of the experiment in mind, I think we can deduce from it that AI is just not able to produce good software unsupervised. It's an important lesson.

To make a wider point, let's look at another of your prediction: that in 2026 the quality of AI code output will be undeniable. I actually think we've already reached that point. Since those agents came around I've never encountered a case where the AI wasn't able to code what I instructed it to. But that's not the same thing as software engineering, and in fact, I have never been impressed by the AI solving real problems for me.

It simply sucks at high quality software architecture. And I don't think this is due to a lack of computing power but that, rather, only humans can figure out what makes sense for them. And this matters, because if the software doesn't make sense, beyond very simple things you can test manually, it becomes impossible to know whether it works as intended.

A Web engine is a great example, because despite the extensive shared suites of tests and specifications, implementing one remains a challenge. You can write code and pass 90% of some sub test suite, and then figure out that your architecture for that Web API is all wrong and you'll never get to the last 10% and in fact your code is fundamentally broken. Unleashing AI without supervision makes this problem worse I think. Solving it requires human judgement and creativity.

polyglotfacto··on Wilson Lin on FastRender: a browser built by parallel agents
Yes, but that is not what I wrote.

I wrote: "useless even as a proof of concept". It doesn't have to be perfect; it just needs to show a clear path forward.

polyglotfacto··on Wilson Lin on FastRender: a browser built by parallel agents
I understand this is not meant as production level quality, but as a web engineer I was expecting at least a decent POC with some interesting design ideas; not total spaghetti that even gets the spec wrong(despite the good idea of checking the spec in the repo).

They may have solved a problem related to agent coordination, like you discussed in your interview related to conflicts and allowing edits to merge without always compiling.

But at the end of the day, a novelty like this is only useful in so far as it produces good code; I don't see how coding agents are of any help otherwise.

So the failure of the pattern should be acknowledged, so we can move on and figure out what does work.

I speculate that what does work is actually quite similar to managing an open source project: don't merge if it doesn't pass CI, and get a review from a human(the question is as what level of granularity). You also need humans in the project to decide on ways of doing things, so that the AI is relegated to its strength: applying existing patterns.

In all seriousness, you can tell Wilson to get in touch with me. With even only one person with domain knowledge involved in such an effort, and with some architectural choices made ahead of unleashing the herd, I think one could do amazing stuff.

polyglotfacto··on Wilson Lin on FastRender: a browser built by parallel agents
The fact that it is new is meaningless: the output is useless even as a proof of concept web engine and should be discarded, alongside the agent engineering pattern that produced it.
polyglotfacto··on Wilson Lin on FastRender: a browser built by parallel agents
> Looks like a success to me.

How is spaghetti code that does not implement the spec(web standards in this case) success?

You are one of the creators of Django; so let me try to give you an analogy: if someone runs thousands of agents in parallel to produce a web framework, and the code ends-up being able to connect to a database and render a template using existing libraries, and the rest would be total non-sense and otherwise useless to web devs; would you call that a success?

Success in software requires something that works as intended and is maintainable.

polyglotfacto··on Wilson Lin on FastRender: a browser built by parallel agents
> it's very much intended as a research project

If so then the failure of the experiment should be acknowledged.

Failure described among others at: https://news.ycombinator.com/item?id=46705625

> It's functional enough to render web pages

> FastRender may not be a production-ready browser, but it represents over a million lines of Rust code, written in a few weeks, that can already render real web pages to a usable degree.

This is something that can be done in much less than a million lines of code. There must be a core somewhere in Fastrender--probably just a few thousands lines--which is putting together existing layout and graphics libraries and makes it render something to the screen.

Doing that in a few weeks isn't impressive, especially not when buried in a million lines of spaghetti code.

If you want an example of a real prototype web engine build along radical design choices, head over to https://github.com/DioxusLabs/blitz

I'm pretty sure it renders far better than Fastrender(the edits the agents made to Taffy are probably nonsense), and I'm guessing it is at most 50k lines.

Conclusion:

In the light of the efforts to paper over failures, I'm calling Fastrender not a research project but propaganda.

polyglotfacto··on Scaling long-running autonomous coding
I'm not sure about what you mean with your first sentence in terms of product.

But in general, my guess at an answer(supported by the results of the experiment discussed on this thread), is that:

- GenAi left unsupervised cannot write a browser/engine, or any other complex software. What you end-up with is just chaos.

- A group of humans using GenAi and supervising it's output could write such an engine(or any other complex software), and in theory be more productive than a group of humans not using GenAi: the humans could focus on the conceptual bottlenecks, and the Ai could bang-out the features that require only the translation of already established architectural patterns.

When I write conceptual bottlenecks I don't mean standing in front of a whiteboard full of diagrams. What I mean is any work the gives proper meaning and functionality to the code: it can be at the level of an individual function, or the project as a whole. It can also be outside of the code itself, such as when you describe the desired behavior of (some part of) a program in TLA+.

For an example, see: https://medium.com/@polyglot_factotum/on-writing-with-ai-87c...

Page 1 of 2Next →