HNHacker News
TopNewBestAskShowJobs

noodletheworld

431 karma · joined December 14, 2024

submissionscomments
noodletheworld··on JetBrains reports revenue growth, net financial loss for 2025
> I certainly don't need my subscription anymore

Sure. I still use mine though; privately and at work.

> Nobody is going to use their tools in five years.

> they're dead. Just like half the people expecting to still be writing code by hand.

Hyperbole much. Leave it on LinkedIn man.

Who knows?

5 years is a long time.

They're trying to navigate the AI era just like everyone else.

You know what we got told at work? “Scrappy AI startups are snapping at our heels, we need to move fast so the competition doesn't over take us”

…but if you ask: what startups? Silence. If you ask, why are we scared of some teenager vibe coding a platform and stealing our customers? Silence.

Dont question the narrative.

Of course AI is an existential threat. Of course there are vibed startups trying to eat our market.

Have you not seen our share price?

Such hyperbolic BS. Ffs. Calm down.

noodletheworld··on Getting the most out of Opus 5.5 in Claude and Claude Code
Hm… I don’t think you should let claude run for 5 hours with a ten line prompt.

You’ll get “something” but what you get is certainly not going to be what you wanted.

Look, its not complicated:

1) be precise in what you want

2) have a feedback loop to verify it.

3) check in often, not once a day.

> Migrate the payment endpoints from the old client to the new one. Done means: every endpoint uses the new client, the old client is deleted, and the test suite passes.

What end point? What test suite? What is a client?

Maaaaybe the model can infer it from your code, but look, if a human taking that jira ticket would go: what does this mean? …then your model (yes, even 5.5) is going to make a bunch of assumptions.

Correct assumptions? Maybe. Maybe not. …but do you really want to find that out hours later?

Just say:

> Plan out the following as a set high level tasks: …

Then, review the plan and tell it to execute the plan, maybe something like; after each step, ensure the code base compiles and the tests pass.

Then, come back after one hour and make sure it’s on the right track.

Poorly specified long running tasks is a recipe for “rollback all those changes…”

noodletheworld··on Make Tmux the OS
> Filesystems as a weak abstraction.

Is a concept the author needs to get over.

> The ideal would be "across my entire computer and applications, what are the recent files I have interacted with" and expand that out to include emails and Teams/Slacks and everything. I should be able to see everything going on with my machine, search through it and not care if its stored inside of Slack or iCloud or whatever. Single pane of glass.

Nope. This would be great if it worked.

This idea of getting rid of the file system is very old, tired and has failed. And failed. And failed.

…and you know apps hate? Letting you out of their walled gardens.

> Microsoft tried it, people didn't like it,

Yup.

> but I think the concept makes sense.

Nope.

You’re going to have to do better than trying this again.

Something something local LLM? Mm.

It’s a Hard Problem. The kind of problem that random engineers look at and go “doesn’t seem that hard, I could do that in a weekend”.

…

> In the interest of full disclosure I think this is gonna take many weekends of work to even get a functional demo running. So I'll do my best, but be patient.

Okay, to be fair, the author did acknowledge that it’s a lot of work and it’s hard; and I look forward to seeing the result, because experiments in this space are interesting, the idea of a visual tmux maybe has some legs…

…but I think this project is too ambitious to succeed.

Not being the default isn't what stops these concept desktops from taking off; it’s that people don’t like them.

No file system, you just get “tasks” and an LLM to pray to and hope you get your data from?

:/ I don’t think it’ll fly.

noodletheworld··on Ember-1
Seems irrelevant? Of course we don’t use data for training.

…trust me bro.

It’s obviously easier to believe when they’re not training models.

Eh, anyway this whole thing is just an ad:

> Looking to take Ember-1 one step further, and optimize it for your use case? We are also launching training support for Ember-1, enabling enterprises to build customized, token-efficient models tailored to their needs with their own data. The future of open models is specialized models trained on your specific workload.

Probably, I guess, fancy serverless infrastructure actually makes virtually no difference to hosting really large models that people want to use, and “just” being an inference provider for open weight models turns out to have no moat.

So this is a bit of a pivot to “use our training infrastructure too…!” imo.

Pivot? Sure. Go them. Not what I signed up for though. /shrug

noodletheworld··on You can defeat the Dream Devourer from Chrono Trigger using an int overflow
> It does imply that you’re blind to all the other aspects of games that bring people enjoyment.

Like what?

I’m not making a sweeping generalisation here about how people play games. We’re talking about chrono cross.

For a story driven JRPG like CT / CC specifically, if you don’t care about the characters or story, there is no reason to achieve any goal in the game.

I’m not so proud I can’t admit to being wrong; but in a JRPG there is no game without the character driven story. If you imagine replacing Kid and Serge with cardboard cutouts labeled A-san and B-san, what are you enjoying about the game?

noodletheworld··on You can defeat the Dream Devourer from Chrono Trigger using an int overflow
Bluntly? Because it’s a meaningless opinion.

You like coding, but not the algorithms.

You like books but not the story.

This isn't “I will play it my way”; its heres a BS reason why legitimate critique can be dismissed out of hand.

“Oh, but for me its not about of the food tastes good”

Ok. Sure. You do You.

…but if someone says: that meal was rubbish and tasteless, is it so outrageous to call out “but I don't care about the taste” as incomprehensible nonsense to the majority of people?

Am I being outrageous here?

What does the the “rp” stand for in rpg?

You tell me.

noodletheworld··on You can defeat the Dream Devourer from Chrono Trigger using an int overflow
You can downvote the parent post all you like, but…

> Personally I don't care much about stories in games

I cant relate to this.

Chrono trigger is a game I love because of the great story.

If you’re playing rpgs and you don't care about the story or characters…

I have no idea. Its like reading a book and saying you don't enjoy the story; you're what, just there for the physical motion of turning the pages?

No idea. Chrono cross left no impression on me. I agree with the parent post.

noodletheworld··on Bend 2 and the Vibe-Coding Trap
I feel like this is the same black hole as small local models.

Things people want to be awesome and true, and things that are actually awesome and true don't intersect the way people want them to.

…so if there was an easy way to do provably correct AI code, it would be nice.

…but I’d also like a frontier that runs on my raspberry pi and a cheap fully autonomous self driving car that just uses a single cell phone camera.

Unfortunately wanting those doesn't make them exist; and people telling you they do exist usually are either a) uninformed, or b) selling something.

noodletheworld··on Navier-Stokes Announcement
My opinion is that it is pretty clear that they’re not going to do that.

> The rules governing the prizes describe the process for evaluating what has been achieved and for assigning credit. The process is deliberately unhurried, but we will provide updates.

I think “you don’t get anything straight away for rushing your AI into the maths problems, not even credit” aligns pretty fairly with what the fields medalists are concerned with.

noodletheworld··on Our decision on Cursor following its acquisition by SpaceX
Days until cursor announces open models… 3, 2, 1…

Seriously though, when the api providers start blacklisting cursor and they’re forced to swap to open weight models…

If that happens, what even is the difference between openrouter and cursor?

noodletheworld··on Small Models Have Arrived
So can anyone using sol and luna.

Don't complain to me if someone responds using a stupid metaphor that proves the opposite of the point they were trying to make.

noodletheworld··on Small Models Have Arrived
Anyone who can’t tell the difference between driving those two cars isn't actually driving.

You cant just go “oh hey, I guess they're both cars so I’m taking your lexus away, catch a cab its cheaper” and expect people to just hug you be be like “yay, thanks! I still have a job I guess! :party:”

:P

noodletheworld··on Small Models Have Arrived
Is sol better?

Yes. Categorically. Anyone who tells you otherwise and that luna is “just as good” does not know what they are talking about.

Going from sol to luna is a downgrade.

It is not a question, it is a fact.

> Is sol actually worth the extra cost?

Is a question only you can answer, because it has no generic answer.

Right now, for me, being able to use sol is worth the cost, but using it all the time is not.

I’m sure going from using it to using luna feels rubbish; but there are realities about costs you have to face sooner or later.

Maybe like… give your team credits and make them pick the right tool for the job; and if they burn their credits on sol in 20 minutes, well, tough luck buddy, looks like you're coding by hand for the rest of the month.

Team will quickly shift. People hate losing access to ai.

noodletheworld··on Apple introduces M6 and M5 Ultra for a big leap in performance and AI compute
How much memory does that come with?
noodletheworld··on Slack Code
What even is this feature?

You can’t just “spin Claude up in slack”; it can’t magically just connect to GitHub and your code.

I presume this feature basically is like… if you’re already using your agents in cloud hosted runtimes, then you can connect to those agents like you already are, but via slack.

…

Did I misunderstand?

This seems like a non feature; if you’re bring to want it, you can’t have it because you have to setup Complicated External System, and if you’re already using cloud hosted agents, you’re already using it (eg. Claude tag or whatever).

So it’s what… a diff viewer in slack? An agents tab so slack becomes like cursor?

The videos seem enormously light on examples of how it actually works, did someone find a more detailed example of usage to show why you want this?

noodletheworld··on My friends all hate AI; I just joined an AI startup
Thats because it’s flame bait; there no actual discussion going on, just people shouting.
noodletheworld··on My friends all hate AI; I just joined an AI startup
Can we just flag obvious flame bait comments like this?

Theres clearly no meaningful thought behind a comment like this, and all it does is stir up engagement for laughs.

The comment thread clearly shows this isn't a meaningful discussion, it’s people shouting past each other.

How can AI be bad? I only see good things! :troll face: it must be a woke left thing! :troll face:

Come on.

noodletheworld··on Prevent cognitive debt by manually retyping LLM-generated code
I work on mobile native applications.

Without an active harness (eg. Appium) that can end-to-end deterministically verify the changes you make continue to work correctly it is almost impossible to continue to keep the same pace on the app.

Unsupervised LLMs (even fabel) are categorically incapable of running parallel unsupervised mobile app feature development.

That is my personal, first hand experience working in a team in this space.

What you are (I guess?) experiencing is user-in-the-loop light touch LLM development where you can 80% most tasks quite quickly (much faster than without assistance!) with a small number of human developers working on largely unrelated features and manually verifying they are correct and manually fixing the platform specific issues you encounter.

Maintaining a strong appium end-to-end test suite is still extremely challenging with notifications and maps.

Honestly, it blows my mind you could even being to claim that of all things, native apps using obscure languages like swift are suitable for this, compared to the much much easier path of web + react.

You might say “yeah yeah, but one month? Come on!”

…but have you actually seen how much code fabel can write in a month?

Its a lot.

So sure, you say, work at a slower pace. Don't just endlessly run a frontier model in unsupervised feature development mode.

Yes… you see, thats the point. Thats what the op is saying.

Move more slowly, and you can avoid building a spaghetti castle (ok sure! If you dont wanna, maybe don't retype every character by hand, but the point of that practice is not upping your wpm typing speed. :p It is to take the time to think, design and collaborate, not rush rush rush)

noodletheworld··on Qwen3.8-Max: A New Bar for Coding and Cowork
> You need to setup a harness that works against your model - You might need to setup additional websearch tools, image tools, etc since harnesses like pi don't come with the model

Pi has a nice guide on it (https://pi.dev/docs/latest/llama-cpp) and it is really not that hard.

How is that hypocrisy? Self hosting is somehow anti AI? Its not anti AI. Its literally using AI!

…and honestly, at a higher technical level than slapping your wallet against a token provider and running prompts in a hosted sandbox you can't even see the prompts in.

noodletheworld··on I regret migrating to Codeberg
That thread is from 2020.

Times change.

Free services owe no one, anything. If you are joining a community, be respectful.

If you're not, and you get banned, that sucks. If it causes a whole category of your peers to be banned that sucks too.

…but its a bit rich to say, oh yeah, the promises you made before LLMs were a thing didn't include a note saying “except if you spam AI generated code”.

Come on.

They dont owe anyone anything.

Certainly not for some comment from ten years ago.

This could entirely legitimately be a case of “was this decision made the right way according to their own decision making process?” … but “is it ok for then to have made this decision?”

Yes. Its fine. It’s their platform. They can do whatever they want.

noodletheworld··on AI Mania Is Eviscerating Global Decision-Making
Look, if they have data and it says 0%, and you have vibes that say that can’t be true, who should we believe?

Do you work with lots of companies and see large AI success stories?

Or do you just vibe that you personally find AI useful so it must also be a business success?

Look, I honestly don’t care, but I think “it must be false” is also unsubstantiated hyperbole. If an agency says they see no AI success, I see no particular reason to believe they’re lying.

They’re not saying AI can’t be a success. They’re saying they haven’t seen it. That matches my experience too. Proven AI success stories seem… vague, when you dig into the details, in my personal experience.

It doesn’t seem surprising to me.

noodletheworld··on Fable turned reMarkable into Tom Riddle's diary from Harry Potter
Whatever man.

If you slap chat GPT on a tablet or a website or a watch or a smart fridge you can make a song and dance about how great it is.

…but bluntly, it would have been impressive 20 years ago.

It is not impressive today; at least, not to me. I’m not 6.

noodletheworld··on Fable turned reMarkable into Tom Riddle's diary from Harry Potter
o_O am I missing something?

Did you show them the github repo, or the disappointing 5 second video of chat gpt from twitter?

I was excited too until I saw it actually running and was like… huh. ChatGPT for tablet. Righto.

What was your son so impressed by?

noodletheworld··on Zuckerberg says AI agent development going slower than expected
> plenty of evidence … although not qualitative or obviously causal

Those two things are the opposite of each other (evidence, but only anecdotally; you cant be both).

Anyway.

More tangible to your argument; what is your argument that this will be more effective than just prompt engineering?

Ive long believed that prompt engineering is a losers game; if there is a trivial set of tricks that improve the output, they will simply be automatically applied.

We see this playing out with the system prompts in coding agents and image gen.

The value of learning “photo realistic studio lighting…” was non existent. The nano banana api is capable of taking a naive prompt and expanding it with these tricks.

People who devoted themselves to learning these “magical incantations” wasted their time and effort; and it was obvious, from the beginning this would be true.

Now.

With managing agents; if a trivial set of management tricks can drastically improve the results, why are you better off learning them now, rather than waiting for them to be baked into cursor/codex/claude in easy mode?

What makes you believe this is a valuable investment in time and effort?

Even if we accept that right now assigning personas to agents and managing them as a manager yields good results, the horizon for change right now is so short, it seems extraordinary to suggest mass management and leadership training for engineers.

We should just wait and see.

All in investments like this would just be tokenmaxing in a funny hat.

noodletheworld··on Jamesob's guide to running SOTA LLMs locally
I sat in a meeting 7 weeks ago where senior leaders said they expect token prices to drop significantly over the next 6 months, and we should all be using as much AI as possible; our team goal was set to use more tokens.

This week, we are banned from using anything more expensive than opus 4.6 and encouraged to use sonnet (but not sonnet 5! Thats expensive!) or lower for daily tasks to help manage costs.

Weeks ago, they gave exactly the same justification as you just gave; and it makes sense!

…but maybe not over the next 6 months.

> Maybe not over the next 18-24 months

Maybe not. Probably not, I guess.

A lot of money has been invested on the expectation that the current gen of hardware is going to reap a colossal profit, and the capex to replace it, is vanishing into investor skepticism as we speak.

It seems like most people have a very very low ability to forecast long horizon change in the current environment, but, in general… it seems like until demand drops, the chances of prices dropping is dubious; at best we get a price war with chinese models or a bubble pop; and even then, there are plenty of startups lurking to snap up cheap hardware.

For individuals, the horizon for buying cheap AI capable compute doesn’t seem close, at all, to me.

noodletheworld··on Vite+ Beta
I appreciate the effort to bring things together in this but…

> Vite+ will manage your global Node.js runtime and package manager.

What? Why?

You’re really going all-in if you adopt this; and… for what? A bit of cozy tooling around existing standard ways of doing things?

Ok, sure; I like tools, like vite.

…but even for an opinionated tool, this is extraordinarily opinionated. Like next.js

Im skeptical.

The pitch of bringing things together seems strong, but did we go too far here?

Reading reviews of people using this didn't really convince me.

It seems to be running on the coat tails of the vite name, rather than its own merit.

noodletheworld··on Tokenmaxxing is dead, long live tokenmaxxing
? What is your point? That the OP is obviously finically motived to encourage tokenmaxing?

Here’s what they said, $$$ aside:

> That’s no longer true. We’ve entered a different regime, where spending more tokens generally results in better results. We call this “compounding correctness” — the more tokens you spend on getting a task correct

> Compounding correctness flips the calculus. If more token spend leads to better outcomes, then you’re going to want to spend a lot of time running tokens. Which sure as hell sounds like tokenmaxxing to me! The original incentives to tokenmax are gone, but eventually folks will realize that a new and more powerful incentive has take its place.

> There were ways to get loops to work, but it was hard. You had to think a lot about how to prompt the agent, which in turn required a pretty deep familiarity with how these things work.

> Now, though, it’s easy. Compounding correctness makes it easy

Go on, tell me I’m quoting the OP out of context.

It’s pretty clear this person believes in compounding correctness, while other, more serious people (1) are perhaps more skeptical.

..and Armins company owns pi. You can’t get much more all in on AI.

Compounding correctness sounds cool, but the real examples of people spending lots of tokens are not compounding correctness; they are wide parallel exploration; like Mythos. The OP is confused, and wrong; they’ve made some basic (flawed) assumptions, and based their entire reasoning on them.

…and are selling AI things. How surprising.

[1] - https://lucumr.pocoo.org/2026/6/23/the-coming-loop/

noodletheworld··on The Coming Loop
Theres a deep insight in this post about the value of looping for throw away code to explore a problem space, rather than brute force a problem by just applying more tokens and hoping.

The more I play in this space, the more I’m drawn to the idea that some kind of back tracking constraint solver is a better solution than then the current naive while loop / brute force approach here.

The results I see are similar to what you get from a greedy brute force constraint solver; solves trivial problems, sometimes solves harder problems after a long time, takes too long to solve really hard problems; solutions are increasingly non optimal on average as complexity goes up.

We have so much existing knowledge about building good constraint solvers, if we could just figure out how to apply it here somehow.

noodletheworld··on The Coming Loop
Use appium or XCTest or swift testing; generate the tests first (failing) from the spec.

The loop is basically then a while loop:

While (tests fail) { trigger agent: spec, failures list }

for bugs, write failing tests.

Its basically TDD.

Loops do nothing useful beyond making the “spec -> code” step more “hands off” and let you be confident that the code you write does what is intended.

Obviously you see the issue: writing the loop harness is > effort than not having it…

…but the idea is that you run “spec first” and are totally hands off on the code, just updating the validation step and then waiting while the agent iterates over and over to solve for some solution that passes the loop harness.

People suggest that it is possible to go, eg. directly figma/jira to harness via (random tool here), saving even more time and invoking even fewer humans, but thats currently, as far as I can tell, actually just hype.

No one is actually doing that effectively.

Loops are currently carefully hand crafted, which makes them tedious and of questionable value imo.

noodletheworld··on You're probably using Agent Skills wrong
Agree; posts like this frustrate me.

Tldr: you're doing it wrong but I will not show you how to do it right. I also did not run the bench using my approach but it definitely “vibes better” to me, and I reject your actual research paper.

Come on, show us some actual skills.

That one you use all the time looks a hell of a lot like “I wont a deterministic shell script for something a skill saying ‘run the shell script’”

Is that what you do? How much time do you spend on them? How do you stop the agent from making a bunch of very similar skills? How do you deal with the explosion of the total number of skills impacting your token use? Do you use skills from github, or is that bad practice? Why?

So many unanswered questions; so little content. :/

Page 1 of 8Next →