HNHacker News
TopNewBestAskShowJobs

Snuggly73

167 karma · joined February 19, 2025

submissionscomments
Snuggly73··on GDPVal: Measuring the performance of our models on real-world tasks
https://huggingface.co/datasets/openai/gdpval/viewer/default...
Snuggly73··on GDPVal: Measuring the performance of our models on real-world tasks
Apparently producing a react component that returns a piece of html with aria tags set up. Long horizon my ass.
Snuggly73··on RustGPT: A pure-Rust transformer LLM built from scratch
well, hopefully the author did learn something or at least enjoyed the process :)

(the code looks like a very junior or a non-dev wrote it tbh).

Snuggly73··on RustGPT: A pure-Rust transformer LLM built from scratch
To my untrained eye, this looks more like an instruct dataset.

For just plain text, I really like this one - https://huggingface.co/datasets/roneneldan/TinyStories

Snuggly73··on RustGPT: A pure-Rust transformer LLM built from scratch
Congrats - there is a very small problem with the LLM - its reusing transformer blocks and you want to use different instances of them.

Its a very cool excercise, I did the same with Zig and MLX a while back, so I can get a nice foundation, but since then as I got hooked and kept adding stuff to it, switched to Pytorch/Transformers.

Snuggly73··on Base44 sells to Wix for $80M cash
I can tell you a variation from Bulgaria (Soviet era joke as well).

In the newspapers there were news that Bulgaria sent 8000 personal computers to Japan. On the next day, there was a slight correction published:

1. It wasn’t 8000. But 8000000 2. It wasn’t computers, but jars of marmalade 3. They weren’t sent to Japan, but returned back by Japan

Snuggly73··on Generative AI coding tools and agents do not work for me
New benchmark for competitive coding dropped yesterday - https://livecodebenchpro.com/

Apparently models are not doing great for problems out of distribution.

Snuggly73··on OpenAI to buy AI startup from Jony Ive
The means of production - yes. But to have a working economy I suspect you have to have a matching consumption. Not sure how this works without involving people.
Snuggly73··on OpenAI to buy AI startup from Jony Ive
And Honest Zuck talking how his AGI is going to specialize in ads and entertainment. To whom are you going to sell ads…

I dunno - I am rather thinking that they are hedging.

Snuggly73··on OpenAI to buy AI startup from Jony Ive
Interesting - how do you reconcile mass unemployment with working economy? (And this is honest question from one that is invested and hopeful that their life savings won’t evaporate overnight)
Snuggly73··on OpenAI to buy AI startup from Jony Ive
I am utterly confused. If AGI is around the corner, this means that the economy is going to be destroyed and money are going to lose their meaning. Who is going to buy your AI gadget? Why spend money on that and Windsurf?
Snuggly73··on A Research Preview of Codex
I mean that there is the possibility that swe bench is being specifically targeted for training and the results may not reflect real world performance.
Snuggly73··on A Research Preview of Codex
I can be completely off base, but it feels to me like benchmaxxing is going on with swe-bench.

Look at the results from multi swe bench - https://multi-swe-bench.github.io/#/

swe polybench - https://amazon-science.github.io/SWE-PolyBench/

Kotlin bench - https://firebender.com/leaderboard

Snuggly73··on As an experienced LLM user, I don't use generative LLMs often
Emmm... why has Claude 'improved' the code by setting SQLite to be threadsafe and then adding locks on every db operation? (You can argue that maybe the callbacks are invoked from multiple threads, but they are not thread safe themselves).
Snuggly73··on LLM-powered tools amplify developer capabilities rather than replacing them
Just to continue my train of thought, because I keep coming back to this.

I think I was expecting that it will turn me into a FE developer and it will feel as natural and smooth as usual when I am in my element.

It didn’t. And the results weren’t what you would get from a real FE dev. And it felt unsatisfactory, stressful and ultimately hollow.

I guess _for me_ it would be fine for a throw away MVP - something that I don’t want to put my heart into.

Snuggly73··on LLM-powered tools amplify developer capabilities rather than replacing them
Well... In my case the scaffolding was done by the Expo template, used Expo libraries for social login and I wrote the Expo API backend functions.

Was it productivity boost for me - yeah, cause I know mostly shit about React. But as an end result it just felt very underwhelming. Discussing it today with my brother (who lives and breathes FE) it apparently was.

I guess I was just expecting... I dunno... more - people are claiming nX productivity boosts, and considering how the UI is mostly boilerplate...

Snuggly73··on LLM-powered tools amplify developer capabilities rather than replacing them
I spend most of my time talking to stakeholders and BAs to help them better understand what their actual problem is instead of coming to me with "we want this" :/

Writing the code is the trivial part.

Snuggly73··on LLM-powered tools amplify developer capabilities rather than replacing them
2. is mostly what works for me.

Usually when I am in the flow of writing code, I can think, write, tab away and review without breaking it. If I need a smallish (up to 100-ish lines) piece of code that I know the shape of - I would use the chat to generate it and merge it back after review.

Letting the agent rip always has led to more pain and suffering down the line :(

Snuggly73··on LLM-powered tools amplify developer capabilities rather than replacing them
I had the weirdest experience the other day. I wanted to write an Expo React Native application - something I have zero experience with (I’ve been writing code non-stop since I was a kid, starting with 6502 assembly). I’ve leaned heavily on Sonnet 3.7 and off we went.

By the end of the day (10-ish hours) all I got to show was about 3 screens with few buttons each… Something a normal React developer probably would’ve spat out in about a hour. On top of that, I can’t remember shit about the application itself - and I can practically recite most of the codebases that I’ve spent time on.

And here I read about people casually generating and erasing 20k lines of code. I dunno, I guess I am either holding it wrong, or most of the time developing software isn’t spent vomiting code.

Snuggly73··on Shopify CEO tells teams to consider using AI before growing headcount
Consider is too mild. Its basically justify why adding more AI isnt the solution. (#5 here https://x.com/tobi/status/1909251946235437514 )
Snuggly73··on OpenAI's o1-pro now available via API
Probably not great (or even unnatural) example. There are tons of examples of PCM players as Flutter plugins on the net and Gemini from the free AI Studio spits an implementation out in about 20 seconds and 0$.

YMMV

Snuggly73··on Ask HN: What less-popular systems programming language are you using?
"Microsoft Macro Assembler, a product I grew to hate with a passion."

Turbo Assembler FTW :)

Snuggly73··on Hallucinations in code are the least dangerous form of LLM mistakes
Maybe be, after all - I dont write web servers (btw, the PQ and JQ libraries doesnt seem to use the arena allocator, which makes the whole proposition a bit dubious, but lets say that its me being picky).

What I meant was, that IMO the code is not very robust when dealing with memory allocations:

1. The "string builder" for example silently ignores allocation failures and just happily returns - https://github.com/williamcotton/webdsl/blob/92762fb724a9035...

2. In what seems most of the places, the code simply doesnt check for allocation failures, which leads to overruns (just couple of examples):

https://github.com/williamcotton/webdsl/blob/92762fb724a9035...

https://github.com/williamcotton/webdsl/blob/92762fb724a9035...

Snuggly73··on Hallucinations in code are the least dangerous form of LLM mistakes
Agree - case in point - dealing with race conditions. You have to reason thru the code.
Snuggly73··on Hallucinations in code are the least dangerous form of LLM mistakes
..."request-based memory arena"...

there are some very questionable things going on with the memory handing in this code. just saying.

Snuggly73··on Yes, Claude Code can decompile itself. Here's the source code
“Using the above technique you can clean-room any software in existence in hours or less.”

Having spent my misguided youth doing horrible things to Sentinel Rainbow and its cousins - I can only chuckle.

Snuggly73··on Yes, Claude Code can decompile itself. Here's the source code
The other linked "oh fuck" article https://ghuntley.com/oh-fuck/ where it supposedly converts https://github.com/RustAudio/cpal at least has a linked video. As far as I can see, based on that video's file tree/Haskell files - its not exactly doing what's described in the article. I'll extrapolate from there...
Snuggly73··on Yes, Claude Code can decompile itself. Here's the source code
Indeed, if you read the attached Claude transcript to the original post, you will see that it finds some strings and infers the functionality (i.e. it has .wav files, thus it probably plays them).

https://claude.ai/share/3eecebc5-ff9a-4363-a1e6-e5c245b81a16

Snuggly73··on Claude 3.7 Sonnet and Claude Code
I've played with it the whole day (so take it with a grain of salt). My gut feeling is that it can produce a bigger ... "thing". I am calling it a "thing", because it looks very much as what you want, but the bigger it is - the more the chances of it being subtly (or not) wrong.

I usually ask the models to extend a small parser/tree-walking interpreter with a compiler/VM.

Up until Claude 3.7 the models would propose something lazy and obviously incomplete. 3.7 generated something that looks almost right, mostly works, but is so overcomplicated and broken in such a way, that I rather delete it and write it from scratch. Trying to get the model to fix it resulted in running in circles, spitting out pieces of code that didn't fit the existing ones etc.

Not sure if I prefer the former or the latter tbh.

Snuggly73··on Some critical issues with the SWE-bench dataset
I only trust benchmarks that I’ve faked myself :)
← PreviousPage 2 of 3Next →