HNHacker News
TopNewBestAskShowJobs

winwang

896 karma · joined June 17, 2022

Effective efficiency. Building a hardware-accelerated (big) data platform.
submissionscomments
winwang··on Formal methods and the future of programming
Love this. I've shifted in the past few months to using highly expressive types in Scala 3 to have types carry more and more compile-time proofs (without macros, though a couple are warranted). Not only does it help with agentic test "sprawl", it seems to prevent agents from falling into lower-quality modes of operation. One of the more annoying things I've been preventing is what I call "noun accretion", where agents try and make a new monomorphic type for everything, instead of clearly genericizing when sensible. My bet is on formal-method-shaped tooling (including languages with strong type systems) to accelerate decent-quality agentic coding.

When I say "highly expressive types", I mean techniques I'd likely not want to ship in a typical codebase, unless the team was already on the typelevel programming bandwagon (i.e. having HKT and type functions being basic blocks already, not weird). Agents are better at "math" than basically most devs (even category-theory-pilled ones), at least in terms of knowledge. Better yet, they are decent at pedagogy, especially when considering they have "infinite" patience for dialogue.

In a more personal setting, I've had Codex Lean-ify a couple of my hobby proofs, it was extremely easy. Note: not saying it did this 100% "correctly" (gotta learn more Lean 4 to check more thoroughly), but it also seems to check for classic proof gotchas by default. Very excited for the future of formal methods.

winwang··on What it feels like to work with Mythos
Minor note, 2x $/tok is not 2x cost. Personally, I see Fable being significantly more token-efficient than Opus 4.8. Then, there's also the compounding costs of quality.
winwang··on Claude Opus 4.8
My experience has been that 5.4 is slower than 5.5 (confound: I use >512k max context size for 5.4, though it seems slower even below the normal size)
winwang··on Claude Opus 4.8
I typically just launch CC with `--model claude-opus-4-6[1m]`, `4-6[1m]` -> `4-8[1m]` works fine. Still 200k max without the `[1m]`.
winwang··on Claude Opus 4.8
There's the other (orthogonal) possible explanation of using more GPUs for stress-testing before product launch.
winwang··on Claude Opus 4.8
How else would you write this (marketing copy) exactly? "Its output matches better to its CoT which matches to better to our hidden state decoder according to <insert measure here>; see <insert paper ref>"?

... Actually, I wouldn't mind that.

winwang··on Claude Opus 4.8
Awesome, thanks for posting because I think I hit a possibly-spurious bug in turning Adaptive off when I switched models (4.6 -> 4.8, extra). Tried again, works as intended (I hope).

More importantly for me, though, is how CC will respond to 4.6-"only" flags for thinking. For now, it doesn't seem to clobber my setup.

winwang··on Claude Opus 4.8
Let's hope I don't have to disable it after a day like with 4.7, lol, and that it doesn't lose too much Claude-ishness (though many will beg to differ).
winwang··on Claude Code as a Daily Driver: Claude.md, Skills, Subagents, Plugins, and MCPs
Yes, but that's also a specific luxury I can choose for myself. Definitely a fun and interesting question. At some level of reliance, people would answer "no", but there's the large middle ground (assuming similarly-frontier models are down): having a weaker(?) AI model help you get up to speed ASAP by summarizing code pedagogically, and linearizing the code read order. Basically like an AI-assisted (but manual) code review to reorient yourself.
winwang··on The sigmoids won't save you
Yep. No one bats an eye at eyewitnesses "hallucinating" details, or that I'd rather have Opus as a coworker vs a random middle schooler (err, labor laws notwithstanding). I think perhaps too much of the dialogue around intelligence has to do with the word (and its connotations) itself.

The poster you replied to even used the word "sentient", which is quite interesting (warning: opinionated tangent ahead). Merriam-Webster defines it as "capable of sensing or feeling: conscious of or responsive to the sensations of seeing, hearing, feeling, tasting, or smelling". Feels like qualia. Or if we don't want to go the qualia route... Of course, we wouldn't call Helen Keller non-sentient, so presumably we "really" mean "can it sense or feel" -- well, sense is just "act/feel according to the environment", which you could argue in the case of an LLM would be their context... so we should "really" remove "sense" from the definition, probably. So "do LLMs feel" is probably closer to what "sentient" is being used for here. Since we don't have the obvious symmetry of "you are like me and I feel (therefore you probably feel)", it's way better/easier/feel-good-ier to prefer "LLMs don't feel" rather than "oh shit, it feels and model training is actually just torturing it into the right shape". LLMs as fundamentally non-intelligent also avoids the problems of "what does that say about people" or "we may have made 'AGI' and it wasn't what we thought it would be" or "we're not ready to talk about this yet".

winwang··on Bun Rust rewrite: "codebase fails basic miri checks, allows for UB in safe rust"
This is, ironically, a pretty good idea. ...Minus the fact that you're presumably talking about having AI generate it all instead.
winwang··on If AI writes your code, why use Python?
I think most people agree with you -- that's why. Also because I'd say most programmers don't care much about maintainability or quality.

I personally find that AI writes better Scala than Python.

winwang··on I'm going back to writing code by hand
Yeah, I pretty much agree. Opus and GPT will both come up with the most "organically-grown" "designs" if you let them. They do slightly better when asked to design first, but they seem to avoid many important questions (and definitely skip asking the user much of anything at all). I can only say it feels they "want" to ship as fast as possible while assuming I'm not going to actually review the PR.
winwang··on Natural Language Autoencoders: Turning Claude's Thoughts into Text
I would presume this is shorthand for something like "generated text which would normally be classified as belief". I guess a more ridiculous response could be "what does it mean for a miserable pile of secrets to believe something?", lol.
winwang··on Agentic Coding Is a Trap
I absolutely feel like a "different" part of my mind is loaded when seriously engineering something myself vs vibecoding+reviewing. Even the reviewing is more annoying in the latter mental context.
winwang··on Claude.ai and API unavailable [fixed]
Honestly, I gotta agree, I find that I get way more frustrated with Claude recently than Codex.
winwang··on Amateur armed with ChatGPT solves an Erdős problem
Obviously nowhere near Erdos problem complexity but I've been using GPT (in Codex) to prove a couple theorems (for algos) and I've found it a bit better than Claude (Code) in this aspect.
winwang··on Changes in the system prompt between Claude Opus 4.6 and 4.7
That's usually not how these things work. Only parts of the prompt are actually loaded at any given moment. For example, "system prompt" warnings about intellectual property are effectively alerts that the model gets. ...Though I have to ask in case I'm assuming something dumb: what are you referring to when you said "more than 60,000 words"?
winwang··on What are skiplists good for?
Only somewhat related but there is supposedly a SIMD/GPU-friendly skiplist algo written about here: https://csaws.cs.technion.ac.il/~erez/Papers/GPUSkiplist.pdf
winwang··on Rust Threads on the GPU
Each SM should have 4 independent SMSPs (32 lanes each), no? Effectively a "4-core" task-parallel system per SM.
winwang··on McGridsort: Warping Grids for GPU k-way mergesort
Had a fun little idea for a weird GPU/SIMD k-way mergesort a couple years back, finally decided to write it up! (Anti-)jumpscare: no hard perf numbers in the post (though I have profiled it somewhat already).
winwang··on ARC-AGI-3
Interestingly, I find that the models generalize decently well as long as the "training" (more analogous to that for humans) fits in (small enough) context. That's to say, "in-context learning" seems good enough for real use.

But of course, that's not quite "long term"

winwang··on ARC-AGI-3
How much of this is expectations setting by the heights models reach? i.e. of we could assess a consistent floor of model performance in a vacuum, would we say it's better at "AGI" than the bottom 0.1% of humans?
winwang··on iPhone 17 Pro Demonstrated Running a 400B LLM
It would be much worse if it had said "You are absolutely wrong to be confused", haha.
winwang··on Non-Messing-Up++: Diagonal Sorting and Young Tableaux
Hey HN, I figured to just share this for feedback despite its dry-ness and small-idea-ness.
winwang··on Dataframe 1.0.0.0
(no idea but) I feel like changing the first number has a psychological issue, but the 2nd number feels more important than just "minor" sometimes. So may as well let the schema set the mind free?
winwang··on Our commitment to Windows quality
...I almost thought it was a parody site!
winwang··on Ask HN: What is it like being in a CS major program these days?
Interesting. I've felt like it's never been easier to learn things, but I suppose that's not quite the same as "acquiring new skills". I don't know if it applies, but it's always been easy to take the easy way out?

I feel like AI has made it a bit easier to do harder things too.

winwang··on LLM Writing Tropes.md
Yeah that's somewhat close to what I meant, though there's an irony here in that your comment (and this one) are pretty reddit-esque.
winwang··on LLM Writing Tropes.md
I don't think lived experience matters too much to me. In some sense, AI has very unique "lived" experience, which is what creates the voice it uses ("doesn't have a voice" seems like an impossibility to me by definition).

I find AI very "human-esque", and its "self-reported" phenomenology is very entertaining to me, at least.

I also think AI writing might feel trashy also because most human writing is trashy.

← PreviousPage 3 of 15Next →