HNHacker News
TopNewBestAskShowJobs

flail

694 karma · joined June 16, 2016

I lead Lunar Logic, a company with no managers.
submissionscomments
flail··on They Shouldn't Know It's a "User" Interview
Out of interest: why?
flail··on We 3.5x'd Our Pull Requests with AI: Now We Catch Fewer Bugs
Opus was pretty decent at that before the whole Fable drama.

In fact, for our work, which is absolutely not rocket science most of the time, we see little upside in using Fable. The output is still "good enough," but the token burnout rate is through the roof. For many tasks, we stick to Opus, or even Sonnet, and the output is just fine.

Fun fact, we recently prepared an AI recruitment task. The basic idea was that there are conflicting goals in the context. We assumed that AI would lead candidates to a dead end, and they'd have to figure out what was happening. We abandoned the idea as even Sonnet was handling it fine enough. And it burned way fewer tokens than Fable would.

flail··on We 3.5x'd Our Pull Requests with AI: Now We Catch Fewer Bugs
We also find that the agents tend to find false positives. Especially when it comes to security. In theory, you can get an automated pen test every week. In practice, it drowns you in triaging reported vulnerabilities, as you find a significant part of them aren't something that you'd like to address, like ever.
flail··on Check Your Fucking Sources, People
A quick smoke test, then. Gemini 3, Thinking Mode. The article: https://techtrenches.dev/p/the-human-cost-of-10x-how-ai-is-p... The prompt: literally what you suggested.

Gemini: The article focuses on the environmental and human labor costs of scaling Artificial Intelligence, specifically focusing on water usage, electricity, and "ghost work."

Which is hilarious, since the article doesn't even mention the words "water" or "electricity." Gemini remains unfazed, reporting the links that are not in the article (some don't exist at all) to make the final ruling: "The Tech Trenches document is highly accurate in its citations."

Now, I know. Had I used Claude Code with relevant skills, it would have done better. But would it be good?

flail··on Check Your Fucking Sources, People
That thought crossed my mind. However, for such a product to work, there would have to be a human in the loop. With data-starved edge cases, which are many in fact-checking landscape, it would be relatively easy for an LLM to make stuff up or mislabel the context (which it inherently does not understand).

Also, thorough validation would cost a ton in tokens. So it would be both expensive from the tech perspective (AI bills) and labor. Now, whose interest would be to fund such a product? I don't see too many takers...

flail··on Check your fucking sources, people
Um, is Snopes wrong about city cleaning crows, though? As that was the context of the original post. Which, by the way, doesn't say "Go, trust Snopes with everything; they can't be wrong!"
flail··on Check Your Fucking Sources, People
Fundamentally, yes, it is a different "search engine."

BTW, as critical as I can be to AI, using an argument that something didn't work 3 years ago, so it must be crap, doesn't work in this context. 3 years ago, AI could barely generate several lines of consistent code. Now, it generates working apps with a prompt (it's another discussion how good the code is, but still).

I guess 3 years ago, Gemini couldn't tell how many r's are in the word refrigerator.

Same for research. At some point, I switched from ChatGPT and Gemini to Perplexity as it promised AI-powered search. It worked visibly better. Until it didn't, as GPT and Gemini models made a leap.

Back to the point, as long as we understand that, for now, it's all just a probabilistic machine generating the most likely output, no one should expect bulletproof answers. Search was/is way more deterministic than LLMs.

flail··on Check Your Fucking Sources, People
Ultimate credibility? Sure, they never did. Yet the whole thing Google was built upon was using links as tokens of credibility.

You'd assume an outgoing link from a CNN website has more credibility than one from an anonymous blog. That is, I reckon, still true. Although the credibility either link conveys is degrading. Again, it has been so since we started playing the game of SEO, yet AI-generated content in this context is basically a weapon of mass destruction. The deterioration has sped up dramatically.

flail··on Check your fucking sources, people
There's nuance to that. An LLM is quite capable of suggesting relevant reading, given the context. Especially when the context is broad enough that there's enough training data.

"Find me research on code reviews, their size, and quality" would give you more than enough reading. Yet, if you start with a claim, like "Longer PRs mean worse defect detection," the relevant data points fall to few enough for AI to start hallucinating.

You get "something, something, PR length, defect detection, IDK, I don't read research papers." Such output is fine as long as the author cares to validate it.

Skip the second step, and you might be good if you ask about something generic, like "What's the Slack story?" or "How did Blockbuster go bust?" Ask about some specific details, though, and you're bound to end up with made-up stuff that sounds just about right, while it's actually wrong.

flail··on The VC-Funded Company Is an Obsolete Organizational Form
Everything else being the same, more resources are better than less. Yet, VC money comes with strings attached.

VC doesn't want a startup to become just a healthy business. It needs to grow at a breakneck speed. In fact, for a VC, it's better to put pressure on somewhat successful startups to take a moonshot at becoming unicorns, even at the grave risk of going bust instead.

The expectation to spend the funding round in 12-18 months is a well-established pattern. So you get millions, but you have to spend it fast.

Running a product development consultancy, I routinely see products/businesses that could have been built for a fraction of what they cost. You don't need to hire hundreds of developers (pre-2026) and instantly have huge misalignment and coordination issues. You don't need to tokenmaxx the crap of everything (2026), ballooning your AI spend and generating a ton of bloat. That is, unless someone pressures you to spend fast because it's their shot at you becoming a unicorn.

flail··on The Ultimate Question: What Does the Endgame Look Like?
"Even without AI-generated code, [code review] is already a major failing."

By all means, yes. Yet, it feels like we were playing a catch-up game (to a degree), and no one intentionally shipped unreviewed code. Now, reviewed without comprehension becomes standard, unreviewed & unread increasingly happens. That's a different kind of reckless.

"Open source adds thousands of (unpaid) eyes to code review."

True. And the open source community sees a massive inflow of AI-generated pull requests, which floods their capabilities to review. Leaving the ecosystem as it is means it will be dead. Thus, I assume resistance or evolution. And we definitely see some of the former, with some open source codebases being closed for AI contributions.

"Hiring is now 'My AI versus Your AI; and the former real need of 'A qualified person for a suitable job' is lost in the fallout."

Yes, that's where hiring has headed. Which, coincidentally, has made everyone worse off (save for AI-for-hiring apps providers). Candidates have it harder to land a decent job. Companies talk to people who play the AI hiring game better, not the most suitable candidates. All while having the same number of candidates and the same number of jobs, but 100x as many resumes exchanged: https://brodzinski.com/2025/08/broken-ai-hiring.html

Which basically means that a resume has lost its value as a token of information exchange. And since we base the whole process on this very assumption (resume as a token of information), the system is due to be rewired eventually. And sooner rather than later. One random idea: how about creating limited traffic where people actually care at least enough to pay some token money: https://brodzinski.com/2025/12/pay-for-resume-read.html

"My humble suggestion is that our ultimate question be phrased as 'How much is enough?'"

Perfect question if we start from the grand scheme of things. I am afraid, though, that there is never enough. At some point, another billion means increased status. You could buy everything with the billions you had previously, so right now it's a virtual leaderboard between you and other billionaires. And the status game is, indeed, infinite. If you aren't winning now, you can chase the leader. If you are the leader, you try to escape the chase.

The "enough" question doesn't work just as well in a finer-grained context. If we want to figure out things like the evolution of a specific profession. Or consider how digital products will be built in the future. Or how well outsourcing your content generation to an AI agent would work in the long run.

flail··on AI-Generated Products Won't Trigger a SaaSpocalypse
Yes. And the more autonomously we create code, the more of these (and not only these) vulnerabilities we'll be adding. Combine that with the AI-automation in attacks, and you have an all-out security mess.

It's like a Petri dish for inventing new angles of security attacks.

Oh, and let's not forget that coding agents are non-deterministic. The same prompt will yield a different result each time. Especially for more complex tasks. So it's probably enough to wait till the vibe-coded product "slips." Ultimately, as a black hat hacker, I don't need all products to be vulnerable. I can work with those few that are.

flail··on AI-Generated Products Won't Trigger a SaaSpocalypse
Security is even a bigger issue than it looks at first glance. While security risk by omission was always a thing (AI or not), now we face a whole new level of risks, from prompt injection to creating malicious libraries to be used by coding agents: https://garymarcus.substack.com/p/llms-coding-agents-securit...

The most shallow security, however, seems easier. Now, you can get through an automated AI security audit every day for (basically) free. You don't have to hire specialists to run pen tests.

Which makes the whole thing even more challenging. Safe on the surface while vulnerable in the details creates the false sense of safety.

Yet, all these would be a concern only once a product is any successful. Once it is, hypothetically, the company behind should have money to fix the vulnerabilities (I know, "hypothetically"). The maintenance cost hits way earlier than that. It will kick in even for a pet personal project, which is isolated from the broader internet. So I treat it as an early filter, which will reduce the enthusiasm of wannabe founders.

flail··on "SaaSpocalypse" Is Merely a Regression to Normal
The question is not whether we like or want subscriptions, but rather whether we're used to them. And the answer is yes.

Given the choice, we'd be using Spotifys and Netflixes for free, and have ad-free Google. I don't expect that choice to be given to us.

AI tools won't change anything on that account. At best, we'll switch one subscription for another one, except that the latter will add a bill for the tokens we use.

flail··on [dead]
There's a huge difference between nurses or teachers and Ivy League students. Namely, the former are not remotely as prestigious roles. I highly doubt there are 20 candidates for each nurse or teacher job.

Affirmative action happens when we discuss privileged positions. Spots at Ivy League colleges definitely are positions of privilege.

So if the situation under consideration were nursing, there wouldn't be such a discussion because there wouldn't be affirmative action in place.

flail··on A trillion dollars (potentially) wasted on gen-AI
> do Altman and Andreesen really believe that, or is it just a marketing and investment pitch?

As for Andreessen, I don't think he even cares. As the author writes:

"for the venture capitalists that have driven so much of field, scaling, even if it fails, has been a great run: it’s been a way to take their 2% management fee investing someone else’s money on plausible-ish sounding bets that were truly massive, which makes them rich no matter how things turn out"

VCs win every time. Even if it's a bubble and it bursts, they still win. In fact, they are the only party that wins.

Heck, the bigger the bubble, the more money is poured into it, and the bigger the commissions. So VCs have an interest in pumping it up.

flail··on A trillion dollars (potentially) wasted on gen-AI
> Have LLMs learned to say "I don't know" yet?

Can they, fundamentally, do that? That is, given the current technology.

Architecturally, they don't have a concept of "not knowing." They can say "I don't know," but it simply means that it was the most likely answer based on the training data.

A perfect example: an LLM citing chess rules and still making an illegal move: https://garymarcus.substack.com/p/generative-ais-crippling-a...

Heck, it can even say the move would have been illegal. And it would still make it.

flail··on A trillion dollars (potentially) wasted on gen-AI
> We've got something that seems to be general and seems to be more intelligent than an average human.

We've got something that occasionally sounds as if it were more intelligent than an average human. However, if we stick to areas of interest of that average human, they'll beat the machine in reasoning, critical assessment, etc.

And in just about any area, an average human will beat the machine wherever a world model is required, i.e., a generalized understanding of how the world works.

It's not to criticize the usefulness of LLMs. Yet broad statements that an LLM is more intelligent than an average Joe are necessarily misleading.

I like how Simon Wardley assesses how good the most recent models are. He asks them to summarize an article or a book which he's deeply familiar with (his own or someone else's). It's like a test of trust. If he can't trust the summary of the stuff he knows, he can't trust the summary that's foreign to him either.

flail··on A trillion dollars (potentially) wasted on gen-AI
What's the lifecycle length of GPUs? 2-4 years? By the time OpenAIs and Anthropics pivot, many GPUs will be beyond their half-life. I doubt there would be many takers for that infrastructure.

Especially given the humungous scale of infrastructure that the current approach requires. Is there another line of technology that would require remotely as much?

Note, I'm not saying there can't be. It's just that I don't think there are obvious shots at that target.

flail··on A Non-Obvious Answer to Why the AI Bubble Will Burst
> I stopped reading here, which is at the very start of the article (...) > (...) this article is low quality and honestly full of basic errors.

Just curious: How do you know it's full of errors, given that you stopped reading at the very start?

flail··on A Non-Obvious Answer to Why the AI Bubble Will Burst
One more interesting aspect: the infrastructure doesn't age that well. We basically need to renew all that infrastructure every, like, 2-4 years or so? (And I think I'm being optimistic here.)
flail··on A Non-Obvious Answer to Why the AI Bubble Will Burst
I don't think FB was an outlier. I can't be sure, but I don't think there were many (any?) companies that took more than 10 years to profitability pre-2015.

I think Twitter took 11 years, and it was 2017.

Uber is actually a good counterexample for more reasons than just how long it took to reach profitability. It also raised a lot of money $13B+ (compared to Facebook's ~$2B and Twitter's ~$3.5B), and ~$8B from IPO (that's another interesting fact; IPO when bleeding money).

However, it would rather make Uber an outlier, not vice versa. I guess Tesla and SpaceX fall into the "Uber" bucket, too (SpaceX would actually be profitable pre-2015, right?). How many others can you list?

So yes, we have extending timelines, but pouring money into a leaky bucket for 10 years is still predominantly a losing bet. For each that eventually made it you would have Foursquare, We Work, Better Place, Jawbone, Theranos (!), Fisker Automotive, etc.

And for each of those, you would have dozens that are even more forgotten because investors pulled the plug after just a few years (anyone remember fab.com perchance?). I would put Groupons of this world in the same bucket.

But even if we treated Uber and Tesla as the norm, OpenAI has already beaten them all in terms of how much funding it raised (and Anthropic is on its way there, too). Both with no signs of profitability round the corner and an absurd burn rate that can't be carried by any single customer group (and I already think about their geography as global).

That's why corporate results are so important, as they can afford to pay a premium. ChatGPT users will not.

So even among the wildest outliers, AI companies are extreme outliers.

flail··on Poor leadership slows down game development
There's Peter's Principle that says that everyone will be promoted till they eventually become incompetent at their job: https://en.wikipedia.org/wiki/Peter_principle

And then, gamedev isn't known for their progressive approach to management (to say the least). A couple of years back, it made the major news in Poland that CD Projekt RED adopted Agile. They actually pumped PR efforts in that.

In 2023.

Give them two more decades, and they might as well adopt modern management approaches or even Lean Startup.

I would speculate that a relatively high degree of incompetence of leadership in gamedev is a combination of Peter's Principle and the fact that it's an industry romanticized by many. Thus, they can afford not to fix many issues that would be fatal for an average boring corporation. There will always be new blood coming.

flail··on The price of mandatory code reviews
I think these two are two dimensions. You can have any combination of: a) single branch vs feature branches b) code review as a norm vs not required

(I'd rather draw a line with code review being/not being a norm, rather than whether it's mandatory. It can be mandatory and still shit.)

And as you suggest, I would expect that trunk-based development leads to greater care for quality. Add to that code reviews that seem to improve quality even further. I don't see a contradiction here.

Also, what the data suggests is that, for good productivity, it may be more important to have short lead times (from development to production) rather than just "no mandatory code reviews."

If you can expect code review to be done just-in-time, you retain the context, limit the tax of context switching, avoid Zeigarnik effect (https://en.wikipedia.org/wiki/Zeigarnik_effect), etc. So I guess this may be a sweet spot reconciling two sources.

flail··on Development speed is not a bottleneck
YC couldn't care less how well these companies fare post-IPO. Post-IPO, their job is essentially done, and they cash their investment.

They're probably busting champagne if the peak valuation is instantly after IPO. That means they maximized the potential payoff.

You're right that one perspective to consider how well these companies are doing is to look at their continuous growth, and that includes post-IPO valuation.

At the same time, YC backed a couple of dozen unicorns. Including those that, despite having a valuation lower than at the time of their IPO, are still in the 10-digit range. Well, I'd be the last one to complain if I got to seed fund a company that ended up at $50B+ valuation.

Having said that, it's just a financial aspect of the argument, and it's only focused on post-IPO, whereas most YC-backed unicorns have not yet done so.

If you look at product innovation, you'll see Dropbox, Airbnb, Gitlab (save for Airbnb, all failures by your standards), Stripe, Deel, Zapier, all household names in my world. Why weren't these products developed by big techs?

You could easily point to one that would be especially interested in taking over that part of the pie.

flail··on Development speed is not a bottleneck
Sure, it was way easier to move at Google in the early 2000s than it is now. Yet, one has to admit they still keep trying. The list of products that they tried and killed doesn't show signs of stagnation: https://killedbygoogle.com/

And that's only the things that they have released. I'd bet that there are lots more that never make it to the public.

And I expect no less from Microsoft, by the way. Microsoft is, in fact, a great case in point of how failed releases don't hurt the company's PR long-term. How many failures have they scored trying to catch up with the missed opportunities of the 2000s? Smartphones & tablets, search, music players, social media.

They were late to move the Office to the cloud, and kept pumping dollars into the Explorer/Edge lost cause, too.

I don't know enough details, but Xbox seems more like an outlier than a norm.

Yet they rebounded with Azure and made some good bets with AI, and are doing better than ever. However, we don't see a stream of new product bets coming from them.

Oh, and on Apple: I wouldn't discount the role of cult-like following in repeated product success. Neither of the other big techs has such a relationship with its user base. You don't see many raving fans of Facebook or Google. And you definitely have millions of people who would buy any new Apple product simply because it is a new Apple product.

It's like Joel Spolsky but on a global scale. In the 2000s, whatever Joel Spolsky touched turned into gold. Stack Overflow? Check. Trello? Check. Was there something unique about these products? Details, sure. But the biggest thing was Joel's leverage.

Having run a highly popular blog for developers, he could instantly reach out to his early adopters. Given that many of the readers were actual fans, they'd jump on the opportunity, whatever it was. So the early traction was not a problem (which was especially crucial for the developers' forum).

Scale that up to the big tech context, and you get Steve Jobs.

A side note: I wonder how long it will take Tim Cook to dismantle that. You can already see cracks.

flail··on Development speed is not a bottleneck
The last YC batch was like ~170 companies, correct? Each year, there are like 150 million startups. So let's not take YC stable for the whole startup ecosystem.

And I'm with you with a critical view on their all-in move toward AI. It's just what all the VCs do, and it's hard to say who's parroting who in this setup (I think that others are parroting YC, but feel free to challenge me on that).

Having said all that, I wouldn't be surprised if a couple of companies from this year's cohort made it big. If you look at YC's biggest successes year by year, you will often (but not always) find a household name.

Was there anyone who predicted these would be the greatest hits? Of course not! That's the whole point of having an investment portfolio. You can be wrong a lot of times if you secure an early investment in a unicorn every other year or so.

Also, "one recent example" of poor investment decision doesn't invalidate 2 decades of rather successful investment portfolios (as a whole, not individually).

In no way is it a YC defense. I'm very critical of the whole startup funding ecosystem, and they are a prominent player. Yet, if they were consistently stupid with their decisions, they wouldn't exist, let alone be the most desired accelerator out there.

Also, if it's that simple to copy what they do and what the companies in their portfolio do, why wouldn't Google et al. take their almost infinite funds and get the competing offers for non-BS ideas up and running in no time?

I bet that if you had an idea that could pay off thousandfold, you'd get enough eager ears to hear you out in any big tech. And still, it's the makeshift mass of startups that come through with new products.

One has to wonder why things like Shopify, Stripe, Zapier, or Figma did not come from the big tech. Each would have an ideal match. Even if you look at the AI landscape, how come Lovable made such a career? After all, they repackage the AI capabilities rented elsewhere. Somehow, with all the ingenuity of building ChatGPT, OpenAI and the rest didn't get it.

flail··on Development speed is not a bottleneck
OK, let's assume you can get food for free (or close enough). Like if you were super rich, and the cost was absolutely marginal for you.

How many dinners a day can you have?

You would still rely on alternative proxies, like recommendations or reviews.

flail··on Development speed is not a bottleneck
Big Techs do have ways of rolling out new services step by step.

Paul Buchheit's stories about Gmail and AdSense are good examples. I was an early Gmail user when it was invitation-only and invitations were scarcely distributed (only as fast as the infrastructure could handle).

So, while I understand the difference in PR costs, it's not like they don't have tools to run smaller experiments.

I agree with the huge bureaucracy cost. On the other hand, they really have (relatively) infinite resources if they care to deploy them. And sometimes they do. And they still fail.

They often fail even when they try a Skunk Works-like approach. Google Wave was famously developed as a corporate Lean Startup (before there was Lean Startup). It was a disaster. Precisely because they did close to zero validation pre-release.

A side note, a huge flop it was (although Buzz and Google+ were bigger), it didn't hurt them long term in PR or reputation.

flail··on Development speed is not a bottleneck
Of course, Big Techs have leverage of their bottomless coffers. What they can't develop, they buy. What was the last successful product idea coming from, say, Facebook?

Or on a smaller scale, what's the last genuine Attlassian success?

Yet, when it comes to product innovation, the momentum is always on the side of the new players. Always has been.

Project management/work organization software? Linear. Async communication? Slack. Social Media? TikTok. One has to be curious how Zoom is doing so well, given that all the big competition actually controls the channels for setting up meetings. Self-publishing? Substack. Even with AI, everyone plays catch-up with Sam Altman, and many of the most prominent companies are newcomers.

We could go on and on.

Yes, Big Techs will survive because they have enough momentum to survive events such as the Balmer-era MS. But that doesn't mean they lead product innovation.

And it's expected. Conflicting priorities, growing bureaucracies, shareholders' expectations, old business lines (and more), all make them less flexible.

Page 1 of 3Next →