167 karma · joined February 19, 2025
(the code looks like a very junior or a non-dev wrote it tbh).
For just plain text, I really like this one - https://huggingface.co/datasets/roneneldan/TinyStories
Its a very cool excercise, I did the same with Zig and MLX a while back, so I can get a nice foundation, but since then as I got hooked and kept adding stuff to it, switched to Pytorch/Transformers.
In the newspapers there were news that Bulgaria sent 8000 personal computers to Japan. On the next day, there was a slight correction published:
1. It wasn’t 8000. But 8000000 2. It wasn’t computers, but jars of marmalade 3. They weren’t sent to Japan, but returned back by Japan
Apparently models are not doing great for problems out of distribution.
I dunno - I am rather thinking that they are hedging.
Look at the results from multi swe bench - https://multi-swe-bench.github.io/#/
swe polybench - https://amazon-science.github.io/SWE-PolyBench/
Kotlin bench - https://firebender.com/leaderboard
I think I was expecting that it will turn me into a FE developer and it will feel as natural and smooth as usual when I am in my element.
It didn’t. And the results weren’t what you would get from a real FE dev. And it felt unsatisfactory, stressful and ultimately hollow.
I guess _for me_ it would be fine for a throw away MVP - something that I don’t want to put my heart into.
Was it productivity boost for me - yeah, cause I know mostly shit about React. But as an end result it just felt very underwhelming. Discussing it today with my brother (who lives and breathes FE) it apparently was.
I guess I was just expecting... I dunno... more - people are claiming nX productivity boosts, and considering how the UI is mostly boilerplate...
Writing the code is the trivial part.
Usually when I am in the flow of writing code, I can think, write, tab away and review without breaking it. If I need a smallish (up to 100-ish lines) piece of code that I know the shape of - I would use the chat to generate it and merge it back after review.
Letting the agent rip always has led to more pain and suffering down the line :(
By the end of the day (10-ish hours) all I got to show was about 3 screens with few buttons each… Something a normal React developer probably would’ve spat out in about a hour. On top of that, I can’t remember shit about the application itself - and I can practically recite most of the codebases that I’ve spent time on.
And here I read about people casually generating and erasing 20k lines of code. I dunno, I guess I am either holding it wrong, or most of the time developing software isn’t spent vomiting code.
YMMV
Turbo Assembler FTW :)
What I meant was, that IMO the code is not very robust when dealing with memory allocations:
1. The "string builder" for example silently ignores allocation failures and just happily returns - https://github.com/williamcotton/webdsl/blob/92762fb724a9035...
2. In what seems most of the places, the code simply doesnt check for allocation failures, which leads to overruns (just couple of examples):
https://github.com/williamcotton/webdsl/blob/92762fb724a9035...
https://github.com/williamcotton/webdsl/blob/92762fb724a9035...
there are some very questionable things going on with the memory handing in this code. just saying.
Having spent my misguided youth doing horrible things to Sentinel Rainbow and its cousins - I can only chuckle.
https://claude.ai/share/3eecebc5-ff9a-4363-a1e6-e5c245b81a16
I usually ask the models to extend a small parser/tree-walking interpreter with a compiler/VM.
Up until Claude 3.7 the models would propose something lazy and obviously incomplete. 3.7 generated something that looks almost right, mostly works, but is so overcomplicated and broken in such a way, that I rather delete it and write it from scratch. Trying to get the model to fix it resulted in running in circles, spitting out pieces of code that didn't fit the existing ones etc.
Not sure if I prefer the former or the latter tbh.