HNHacker News
TopNewBestAskShowJobs

sothatsit

914 karma · joined February 22, 2021

I really like The Royal Game of Ur for some reason.
submissionscomments
sothatsit··on The short leash AI coding method for beating Fable
There are always concepts that some people think are a basic, that others haven't heard of. The entire benefit here is that AI can point out what we miss. There are certainly techniques you don't know about, or just didn't think to apply to a problem, that others would find to be pretty standard.
sothatsit··on The short leash AI coding method for beating Fable
You can have a nuanced discussion with an LLM. But LLMs also have failure modes where they start making up justifications. The two are not mutually exclusive.
sothatsit··on The short leash AI coding method for beating Fable
I disagree with keeping an eye on the model as it is working, approving every command, and denying and stopping the model when you think it has gone wrong. It is not that it is actively harmful to do this, but rather that it is a waste of time and you can avoid the need for it through better design discussions and review.

Micro-managing and keeping the AI on a "short leash" also lends itself better to telling models to do smaller units of work at a time instead of discussing broader design concerns. That is why I think someone doing this would miss the MILP solution, because they might never discuss the overall design with the model but rather just tell it what to implement next.

sothatsit··on The short leash AI coding method for beating Fable
"Nuanced discussions" is more about describing a design to a model, asking the model to critique your design and ask you for clarifications, and then you providing those clarifications and the model "getting it" and proceeding to additional levels of detail before implementation. In particular the models being able to highlight concerns you have not yet thought about is a pretty good sign of this. Fable is noticeably better at this compared to Opus.

I was not talking about models making mistakes. Mistakes, and then models making up justifications for those mistakes, is a failure mode of any LLM, and Fable is no different in that regard. Newer models might make less mistakes, or at least make less egregious mistakes, but they still make mistakes.

sothatsit··on The short leash AI coding method for beating Fable
This “short leash” seems like more of a crutch to me, and a sign of not giving the AI enough detail on the problem to begin with, or not reviewing and iterating on its output.

Hand-holding great models like Fable through implementation is a waste of time, and a waste of Fable. You can have increasingly nuanced discussions with stronger models, and they write a lot better code than they used to. The process of discussing designs and their implementations, questioning things that look weird to you, and actually reading the AI’s responses also helps to find better solutions.

For example, one time I wanted to write a greedy solver for a problem, and in my discussion with Opus on the idea it suggested using an existing MILP library to solve the problem exactly. I’d never even heard of MILP, but my final implementation ended up being better and simpler than what I’d have done alone.

sothatsit··on Claude Fable 5 Promotional Access
You can get away with a lot when you have the best models… I’m looking forward to OpenAI or open-source catching up so we have some competition again.
sothatsit··on Previewing GPT‑5.6 Sol: a next-generation model
Refusals, presumably.
sothatsit··on Bun has an open PR adding shared-memory threads to JavaScriptCore
It’s pretty incredible to me that a mammoth change like this is possible to prototype now using LLMs.

It makes me wonder how much of our software stack will become more malleable to big ideas and experiments in the future, like Filip’s idea here. Even if you don’t want to merge the code, it’s still an incredible existence proof that something like this could work.

sothatsit··on Did Anthropic ask for this?
Would the US government have slapped Anthropic with this export control if Anthropic never fearmonger'ed about Mythos? I think the answer is very likely no.

But is this the type of regulation Anthropic has been asking for? Not at all.

This is a failure of Anthropic's politicking, and a warning that they need to be more careful with their communication in the future. If they truly want constructive regulations because of their fears about AI, they will need to repair their relationship with the administration, and it is still unclear to me how they plan to do that.

sothatsit··on Open source AI must win
The big AI labs are also accumulating huge datasets of expert work in a wide range of fields, which is very expensive to re-create. It seems pretty plausible that this this gives them a big advantage that is compounded by their larger training runs and larger models.
sothatsit··on Statement on US government directive to suspend access to Fable 5 and Mythos 5
It is gone for me now.

> There's an issue with the selected model (claude-fable-5). It may not exist or you may not have access to it.

sothatsit··on Claude Fable 5
Seems to just be a bigger model.
sothatsit··on Claude Fable 5
The Team plan is ~125 USD / month / user. Big enterprises like Uber are paying upwards of $1500 USD / month / user. Anthropic can raise their revenue a lot more by selling to big enterprises than they can by selling more team plan seats.
sothatsit··on When AI Builds Itself: Our progress toward recursive self-improvement
Definitely, it is quite an extreme change. But the upsides of better access to support and advice are huge, even if the potential downsides are scary as well. This feels like one area where we need better transparency and regulation due to how much ChatGPT and others can affect people who listen to them.
sothatsit··on When AI Builds Itself: Our progress toward recursive self-improvement
It is not solely or even primarily the big AI labs that would need to prepare. They have a better idea of what’s coming, and they’re positioned to benefit from it.

It is governments, big companies, and individuals who could all experience fundamental changes if any of these predictions come true. If people within the labs believe these possibilities are around the corner, it would be responsible to try to let people know so they can be more ready if RSI suddenly hits and in a couple years time all our work is fundamentally changed.

That’s not to say I agree with their predictions, but rather I’m just saying that there are good reasons for Anthropic to publish stuff like this that are not just PR.

sothatsit··on Ask HN: What was your "oh shit" moment with GenAI?
I gave GPT-4 some source code and my existing tests, and asked it to write a new test, and it did it! It didn’t even run straight away, I had to fix it, but it still blew my mind.

Later, I wrote a ~5k line proxy for work in C, and gave the whole thing to ChatGPT o1 and asked it to review it. It found several real memory bugs, and now that service has been running since with no problems.

Just this week, I was trying to write a greedy solver to pick the best subset of block sizes to keep from a larger sweep for shorter testing. Opus 4.8 suggested that this could actually be solved as a MILP problem, and found the perfect solution in 5 mins. I’d never even heard of MILP before.

sothatsit··on When AI Builds Itself: Our progress toward recursive self-improvement
I would be very surprised if this is an actual thought-out PR strategy. I am far more inclined to believe that their employees are just bought-in to the future where AI is genuinely transformative.

Whether they are right of wrong is another matter, but their claims also don’t seem too far out of the realm of possibility to me.

Coding agents have fundamentally changed my day-to-day job. In the last year, my work has shifted from me writing all of my code, to me writing very little code and spending most of my time on understanding problems better and setting direction, and reviewing, verifying, and polishing the output of coding agents. It has been quite a drastic change.

It is not that outlandish to suggest that coding agents could continue to improve at such a drastic rate over the next year. And the implications of that could be quite large! Even just the implications of more white-collar workers adopting tools like Cowork seems potentially very large, with tools that already exist today. It seems sensible to at least consider this as a possibility.

sothatsit··on When AI Builds Itself: Our progress toward recursive self-improvement
Or: Anthropic genuinely believes the future scenarios they outline are realistic possibilities, and they want more people to take them seriously.
sothatsit··on When AI Builds Itself: Our progress toward recursive self-improvement
Maybe my bar for what constitutes a breakthrough is lower than other people's, but all of these seem like breakthroughs to me:

NLP as a field saw huge shifts. NLP tasks that used to be complex and inaccurate can now be setup very easily and quickly using structured outputs from LLMs, often with greater accuracy.

A small charity I help with has now been able to build their own website to manage their day-to-day operations. It saves them a lot of time, and it was vibe-coded using Manus. I don't think people appreciate how much room there is left for bespoke software to have big impacts on small organisations that can't afford to hire developers. The cost for software like the one they made has gone from 10s of thousands of dollars to $10/month and volunteer hours.

My brother has recently been setting up Cowork to do an automatic review of contracts before human review, and he said it is far more diligent than people when it comes to routine things to check. This is another huge breakthrough for not just efficiency, but the quality of work.

I really don't think we can discount AI finding bugs and vulnerabilities. If you care about code quality and keep up review standard, LLMs can help you write more robust software. AI has found a huge number of bugs for me before they hit production, including potential out-of-bounds memory accesses and segfaults.

ChatGPT has 1 billion MAU. People are now getting life advice, financial advice, and mental health help from chatbots at a scale and cost that no human support network could match.

sothatsit··on They’re made out of weights
You can claim the use of AI is unethical, or the work as derivative, but AI being used as a tool in no way precludes something from being art. It is thought provoking and challenging, it seems like textbook art to me, and it’s clearly struck a chord here. There is no “minimum effort” required for something to be art.

I personally found the contrast with the original “They’re made out of meat” to be really interesting. I don’t care that AI was used during its creation at all.

sothatsit··on Uber's $1,500/month AI limit is a useful signal for AI tool pricing
This completely ignores all the other huge costs the AI labs are paying in data center builds, researcher salaries, experiments, and training models.

The fact that Anthropic is rumoured to have a profitable quarter indicates that their margins on API priced inference are very strong.

sothatsit··on Uber's $1,500/month AI limit is a useful signal for AI tool pricing
You must not be using coding agents. You can sneeze and spend $1 on Opus in Claude Code.
sothatsit··on Can the stockmarket swallow Anthropic, SpaceX and OpenAI?
Enterprises are paying API prices, which are ~9x the price of the plan for the same usage. A lot of people on the plans are not maxing them out either.
sothatsit··on I think Anthropic and OpenAI have found product-market fit
I don’t remember ever hearing Dario or Sam recommend replacing people. Rather they say that smaller groups of people can do more work, so hiring will slow because small teams can do more.

The only times when people talk about actual full replacement of people is always when they are talking about some “future AGI” that is far more capable than the tools we have today.

sothatsit··on Is AI Profitable Yet?
Or, tokens are more like energy and prices will drop over time until they reach some equilibrium.

The big labs are actively moving into the application layer, where they’ll have more pricing power. Maybe that layer will end up with a Mac (Anthropic) vs Windows (OpenAI) vs Linux (open-source) dynamic as well if they can create a moat. But so far it’s pretty easy to move between providers.

sothatsit··on Vibe coding and agentic engineering are getting closer than I'd like
I think the distinction is that for experiments and prototypes the behaviour of the final system is what we are trying to design. We can experiment and see the tradeoffs and explore the design space before committing to a direction. And then we can sit down and produce the final code to a quality we are happy with. If you are serious about this process, there is no way you are producing 1000s of lines of code a day, unless it is trivial boilerplate.

In terms of higher-level abstractions, I agree this is one particularly treacherous rung on the ladder of abstractions. Previous abstractions like compilers or garbage collectors have at least had more structure/rules to rely upon. I don't know exactly how that will look but I don't think we will solely be relying on banging on the output, we will also be spot-checking the source code, using profilers or other tools to inspect the behaviour of systems, and asking the agent to explain the architectural decisions made. I'm not sure exactly how this will look, but I do believe that people who care will still find ways to do good work.

sothatsit··on Vibe coding and agentic engineering are getting closer than I'd like
Well you have obviously already made up your mind, so have fun with your confirmation bias. We'll all be over here having a good time, getting more work done. Feel free to come over when you put down your grudge.
sothatsit··on Vibe coding and agentic engineering are getting closer than I'd like
The entire mistake you are making is comparing using AI to skimming textbooks, or taking shortcuts. Your entire premise is wrong.

People who care about craft will care about the quality of what they produce whether they use AI or not.

The code I ship now is better tested and better thought through now than before I used AI because I can do a lot more. That extra time goes into additional experiments, jumping down more rabbit holes, and trying out ideas I previously couldn’t due to time constraints. It’s freeing to be able to spend more time to improve quality because the ROI on time spent experimenting has gone up dramatically.

sothatsit··on Vibe coding and agentic engineering are getting closer than I'd like
You don’t need to write code by hand to learn from iterations and experiments. I run more experiments and try out more different solutions than I ever could before, and that leads to better decisions. I still read all the code that gets shipped, and don’t want to give that up, but the idea that all craft and learning is lost when you don’t is a bit silly. The craft/learning just moves.
sothatsit··on Agents for financial services and insurance
It still surprises me how effective the /simplify skill is.

I’ve also had some great results with a /reflect skill that asks the agent to look at the work in the broader context of the project. But those are the only two skills I use regularly that aren’t specific to our company, codebase, or tools.

← PreviousPage 2 of 13Next →