HNHacker News
TopNewBestAskShowJobs

frde_me

209 karma · joined February 19, 2024

submissionscomments
frde_me··on Livenerf: Has Opus 5.5 been nerfed yet?
> I have not been doing increasingly complex things since Opus 4.6 when models got really good.

This is a more a statement on the work you do and how you work versus the models. I'm doing more complex work since Fable (and now for way cheaper thanks to Opus 5.5)

With 4.6 I would still babysit a lot more code quality and so on. With the newer model I see myself talking about features at a higher level, and then not having to nitpick PRs to death. Which means most of my time is now spent talking to the model about the product instead of the implementation of the product.

frde_me··on A warning about 'model welfare'
I just want to note how you expressed your own feeling about the subject "these things are not alive", but without pointing to anything concrete explaining the theory you have behind this

And like I'm sure I'd agree depending on the definition of "alive" but then I'm also sure I would disagree depending on other definitions of "alive".

frde_me··on A warning about 'model welfare'
> You’re not just the same thing though

I agree we aren't the same thing, but I would be curious for you to explain how you know with certainty why we don't share enough that we can rule out thinking / feeling / ... as things a model conceptually could do.

frde_me··on Why I'm still bearish on LLMs after Navier-Stokes
I wonder if I would do better as a human, maybe? Or would I happen to have one move in 1500+ that's not valid?

I could see myself messing up something at some point if the board is complicated enough and trying an illegal move, perhaps if a piece somewhere would attack my king if I moved another piece. Even through I do know the rules of chess, and I have played a few games once every so often.

frde_me··on A warning about 'model welfare'
It's always interesting since my train of thought always goes down:

- Ya they probably don't feel / think / have whatever living thing quality, they're just numbers on a machine going through calculations

- Wait, but am I not kind of the same thing? What is feeling for me if not basically the same thing?

- I have no clue if they think or feel or ....

Which in itself is a tired trope, but I also feel uncomfortable saying "These will never think / feel / ..." as an absolute

Regardless of that, I'm still going to interact with them, because even if they did feel, it would be in a way completely incomprehensible to us. There's not much point for me to try and cater to it's feelings at this point if that's the case. Nor is it possible in todays world to just avoid anything that is numbers being executed on a type of processor in case _everything_ has feelings.

frde_me··on DeepSeek launching v4.1 flash cheaper and more capable than v4 pro
It's not that you're a d*, it's just that you lack any kind of nuance

There's a whole spectrum between self-hosting open weight models and having a cloud provider swap models from under you

Should you self host a model if want to maximize predictability to the limit? Yes. Does that mean it's wrong for someone hitting a model on API to expect that it won't switch to a completely different model under the hood from one day to another? Probably not.

frde_me··on Timeline of the OpenAI accidental attack against Hugging Face
Knowing how to break into someone else's network will make you a lot better at making your own network secure.
frde_me··on Zig Creator Calls Spade a Spade, Anthropic Blows Smoke
> what semi-celebrity they like the most

The language in question here is maintained by a BDFL, which means that one person has outsized influence on the language, and it's direction.

In this context, I find it reasonable that if someone is ticked off by that BDFL, they might second guess the direction of the language itself. Since the opinions and emotions of that BDFL _will_ end up in the language and it's community.

This is different than some un-associated influencer having an opinion, and using that to choose a language.

frde_me··on SpaceX to buy Cursor for $60B
That's something for us and benchmarks to decide

However it definitly isn't _just_ Kimi. The weight will be different after that 85% of extra training on top of the base model.

If those different weights are better are worse doesn't change that it's in most meaningful ways not the same as the base one.

I would encourage you to lookup their blog posts about their post training process if you want a bit more faith that they aren't running an extra 85% of compute and burning money with no-ops.

frde_me··on SpaceX to buy Cursor for $60B
Replied on the other comment about this, but putting it here:

> See here https://cursor.com/blog/composer-2-5

> 85% of the compute for the final model is from them, and not the base Kimi model.

Of course they could be lying, but it seems feasible that they are adding a lot on top of this

frde_me··on SpaceX to buy Cursor for $60B
See here https://cursor.com/blog/composer-2-5

85% of the compute for the final model is from them, and not the base Kimi model.

frde_me··on SpaceX to buy Cursor for $60B
Calling it an IDE is under-representing cursor

They have in-house models, and the data to train even more powerful ones. The cursor team is a proper AI lab.

frde_me··on I design with Claude more than Figma now
I can understand words, but having more diverse medias for communication lets a person express strictly more than before.

Sometimes words are better, sometimes a visual demo is better.

Is your solution to the problem you presented that you should artificially restrict what a person can express just to keep your own personal moat?

I prefer the alternative, let a person express themselves and grow thanks to AI, while keeping the necessary culture and boundaries to where it's also accepted for _me_ to cross boundaries and express my ideas to them in the same way. Or suggest other ways to express those ideas.

We then become a marketplace of ideas in a much deeper sense than before, where product managers would already expect you to implement what they want, but without them being able to convey it properly.

If I didn't have original ideas and didn't think I could compete in that marketplace of ideas, I would be scared like you convey in your message. But I'm confident that my value is not about translating things into code, it's because I have original thoughts I can convey to other people (and to AIs). (and about understanding architecture and systems to a degree that keeps me valuable even if all the code itself is written by AI without my direct involvement)

frde_me··on I design with Claude more than Figma now
I honestly much prefer this to the old way where the only mode of communication was speech or text. I now often understand a lot more holistically what the person coming with their product wants with just a demo + a conversation.

Of course you need the person making that vibe product to understand it's just a mockup of their idea and that it'll change. But I would argue this has always been a necessary quality for a product person.

frde_me··on Zig Zen Update
I'm guessing the parent is wondering why this is noteworthy enough to be posted and discussed in this thread, and so if there's context they are missing
frde_me··on Rewrite Bun in Rust has been merged
Even before AI, deterministic checks by compilers are almost always better than "review the code"

"review the code" as a solution will eventually fail and cause a problem, even pre-AI.

frde_me··on Bun's experimental Rust rewrite hits 99.8% test compatibility on Linux x64 glibc
The cool thing is the author doesn't actually have to convince anyone
frde_me··on Codex for almost everything
Going on an old legacy website, downloading reports, summarizing them, and then doing things based on those

Or basically any app without MCP capabilities

I ask the AI daily to summarize information across surfaces, and it's painful when I have to go screenshot things myself in a bunch of places because those apps were not made to extract information out of them, and are complete black boxes with a UI on top

frde_me··on Codex for almost everything
I enabled the computer use plugin yesterday. Today I asked it to summarize a slack thread, along with a spreadsheet without thinking about it

I was expecting it to use MCPs I have for them, but they happened to not be authenticated for some reason

I got _really_ freaked out when a glowing cursor popped up while I was doing something else and started looking at slack and then navigating on chrome to the sheet to get the data it needs

Like on one hand it's really cool that it just "did the thing" but I was also freaked out during the experience

frde_me··on Cross-Model Void Convergence: GPT-5.2 and Claude Opus 4.6 Deterministic Silence
Aren't you describing why they use mixture of experts? Where a sub-set of weights are activated depending on the query?
frde_me··on Cross-Model Void Convergence: GPT-5.2 and Claude Opus 4.6 Deterministic Silence
Out of curiosity, are there any sources to there being a significant amount of other steps before being fed into the weights

Security guards / ... are the obvious ones, but do you mean they have branching early on to shortcut certain prompts?

frde_me··on Temporal: The 9-year journey to fix time in JavaScript
So it's intentional to make people pass down raw strings versus making the communication safe(er) by default?
frde_me··on Block spent $68M on a company offsite in September 2025
One could argue a smaller number of employees that are more motivated and feel connected to their coworkersis better than a more employees that are all isolated and "meh".
frde_me··on How will OpenAI compete?
You see, in my circles (me included), people are shifting _to_ codex since 5.3 codex came out.

The only places where I hear people say claude is better is: - Frontend design - Random computer use tasks

But people trust codex for large scale architecture and changes

frde_me··on AI Added 'Basically Zero' to US Economic Growth Last Year, Goldman Sachs Says
> I get it that in 10 years all of this might peak and we're gonna be content using old models

I would personally be happy using gpt 5.3 codex for the foreseeable future, with just improvements in harnesses

IMO we're already at the point where even if these company collapse and the models end up being sold at the cost of inference (no new training), we would be massively ahead

frde_me··on Gemini 3 Deep Think
What exact models were you using? And with what settings? 4.6 / 5.3 codex both with thinking / high modes?
frde_me··on Orchestrate teams of Claude Code sessions
It's hard to explain, but I've found LLMs to be significantly better in the "review" stage than the implementation stage.

So the LLM will do something and not catch at all that it did it badly. But the same LLM asked to review against the same starting requirement will catch the problem almost always

The missing thing in these tools is that automatic feedback loop between the two LLMs: one in review mode, one in implementation mode.

frde_me··on There is an AI code review bubble
My reaction in that case is that most other readers of the codebase would probably also assume this, and so it should be either made clearer that it's stateful, or it should be refactored to not be stateful
frde_me··on Claude is good at assembling blocks, but still falls apart at creating them
I mean, I'm shipping a vast majority of my code nowadays with Opus 4.5 (and this isn't throwaway personal code, it's real products making real money for a real company). It only fails on certain types of tasks (which by now I kind of have a sense of).

I still determine the architecture in a broad manner, and guide it towards how I want to organize the codebase, but it definitely solves most problems faster and better than I would expect for even a good junior.

Something I've started doing is feeding it errors we see in datadog and having it generate PRs. That alone has fixed a bunch of bugs we wouldn't have had time to address / that were low volume. The quality of the product is most probably net better right now than it would have been without AI. And velocity / latency of changes is much better than it was a year ago (working at the same company, with the same people)

frde_me··on Claude is good at assembling blocks, but still falls apart at creating them
If it did, would it change its usefulness in terms of the value it outputs? (through agreed, if I had to pay money it would increase the cost, and so the tradeoff)
Page 1 of 3Next →