HNHacker News
TopNewBestAskShowJobs

Mathnerd314

3,005 karma · joined November 1, 2009

Making the ultimate programming language https://mathnerd314.github.io/stroscot/

In the past I developed SuperTux (http://supertux.lethargik.org) when I was bored.

submissionscomments
Mathnerd314··on The Big Oops: Anatomy of a Thirty-Five-Year Mistake [video]
So, this is pretty difficult to test in a real-world environment, but I did a little LLM experiment. Two prompts, (A) "Implement a consensus algorithm for 3 nodes with 1 failure allowed." vs. (B) "Write a provably optimal distributed algorithm for Byzantine agreement in asynchronous networks with at least 1/3 malicious nodes". Prompt A generates a simple majority-vote approach and says "This code does not handle 'Byzantine' failures where nodes can act maliciously or send contradictory information." Prompt B generates "This is the simplified core consensus logic of the Practical Byzantine Fault Tolerance (PBFT) algorithm".

I would say, if you have to design a good consensus algorithm, PBFT is a much better starting point, and can indeed be scaled down. If you have to run something tomorrow, the majority-vote code probably runs as-is, but doesn't help you with the literature at all. It's essentially the iron triangle - good vs. cheap. In the talk the speaker was clearly aiming for quality above all else.

Mathnerd314··on The Big Oops: Anatomy of a Thirty-Five-Year Mistake [video]
The dates are the dates of the sources, he says in the talk he wasn't going to try to infer the dates these ideas were invented. Also he barely talked about Alan Kay.
Mathnerd314··on Jank is C++
Ok so jank is Clojure but with C++/LLVM runtime rather than JVM. So already all of its types are C++ types, that presumably makes things a lot easier. Basically it just uses libclang / CppInterOp to get the corresponding LLVM types and then emits a function call. https://github.com/jank-lang/jank/blob/interop/compiler%2Bru...
Mathnerd314··on Zig Community Mirrors
3 mirrors? Arch has 827.
Mathnerd314··on LLMs pose an interesting problem for DSL designers
Python is just a beautiful, well-designed language - in an era where LLM's generate code, it is kind of reassuring that they mostly generate beautiful code and Python has risen to the top. If you look at the graph, Julia and Lua also do incredibly well, despite being a minuscule fraction of the training data.

But Python/Julia/Lua are by no means the most natural languages - what is natural is what people write before the LLM, the stuff that the LLM translates into Python. And it is hard to get a good look at these "raw prompts" as the LLM companies are keeping these datasets closely guarded, but from HumanEval and MBPP+ and YouTube videos of people vibe coding and such, it is clear that it is mostly English prose, with occasional formulas and code snippets thrown in, and also it is not "ugly" text but generally pre-processed through an LLM. So from my perspective the next step is to switch from Python as the source language to prompts as the source language - integrating LLM's into the compilation pipeline is a logical step. But, currently, they are too expensive to use consistently, so this is blocked by hardware development economics.

Mathnerd314··on Fine-tuning LLMs is a waste of time
OK, so this intuition is actually a bit hard to unpack, I got it from bits and pieces. So this is this post https://www.fast.ai/posts/2023-09-04-learning-jumps/. Essentially, a single pass over the training data is enough for the LLM to significantly "learn" the material. In fact if you read the LLM training papers, for the large-large models, they generally explicitly say that they only did 1 pass over the training corpus, and sometimes not even the full corpus, only like 80% of it or whatever. The other relevant information is the loss curves - models like Llama 3 are not trained until the loss on the training data is minimized, like typical ML models. Rather they use these approximate estimates of FLOPS / tokens vs. performance on benchmarks. But it is pretty much guaranteed that if you continued to train on the training data it would continue to improve its fit - 1 pass over the training data is by no means enough to adequately learn all of the patterns. So from a compression standpoint, the paper I linked previously says that an LLM is a great compressor - but it's not even fully tuned, hence "not trained to saturation".

Now as far as how fine-tuning affects model performance, it is pretty simple: improves fit on the fine-tuning data, decreases fit on original training corpus. Beyond that, yeah, it is hard to say if fine-tuning will help you solve your problem. My experience has been that it always hurts generalization, so if you aren't getting reasonable results with a base or chat-tuned model, then fine-tuning further will not help, but if you are getting results then fine-tuning will make it more consistent.

Mathnerd314··on Fine-Tuning LLMs Is a Waste of Time
Wasn't there that thing about how large LLM's are essentially compression algorithms (https://arxiv.org/pdf/2309.10668)? Maybe that's where this article is coming from, is the idea that finetuning "adds" data to the set of data that compresses well. But that indeed doesn't work unless you mix in the finetuning data with the original training corpus of the base model. I think the article is wrong though in saying it "replaces" the data - it's true that finetuning without keeping in the original training corpus increases loss on the original data, but "large" in LLM really is large and current models are not trained to saturation so there is plenty of room to fit in finetuning if you do it right.
Mathnerd314··on Surprisingly fast AI-generated kernels we didn't mean to publish yet
> we didn't mean to publish yet

I was thinking this was about leaking the kernels or something, but no, they are "publishing" them in the sense of putting out the blog post - they just mean they are skipping the peer review process and not doing a formal paper.

Mathnerd314··on You do not need NixOS on the desktop
I'd say about 90% of the pain in NixOS comes from its non-FHS (Filesystem Hierarchy Standard) layout. But that's also a fundamental part of Nix/NixOS's design-it was built that way from the start.

For complex packages like Steam, it's both possible and recommended to use FHS-compatible containers on NixOS. Still, I've seen people say things like, "All I do is set up containers-why not just use Docker instead of NixOS?" The thing is, if you dig deeper, tools like Docker or Flatpak are actually less powerful than Nix when it comes to container management.

I've been toying with an idea: using filesystem access tracing to replace the current approach of using random hashes for isolation. This could allow an FHS-style layout while preserving many of the guarantees of the Nix model. It would dramatically improve compatibility out-of-the-box, enable capabilities that aren't possible today, and reduce network and disk usage-since files could be modified in-place instead of being remade or redownloaded.

It's on my backlog, though. Starting a new distro doesn't seem particularly rewarding at the moment.

Mathnerd314··on Heart disease deaths worldwide linked to chemical widely used in plastics
1,1,1-trichloroethane doesn't seem particularly toxic - "probable carcinogen", some neurological and liver effects but I'd say it's probably still safer than e.g. isopropyl alcohol which definitely leads to neurological issues long-term. The reason it's banned is because of the ozone layer, not because it's unsafe to individual humans.

I feel like it's probably the wrong chemical though, far too many similar names. Maybe you meant https://en.wikipedia.org/wiki/Trichloroethylene

Mathnerd314··on Australian who ordered radioactive materials walks away from court
there was a much more "interesting" incident circa 1995, look up "The Radioactive Boy Scout"
Mathnerd314··on You Can Be a Great Designer and Be Completely Unknown
Hot take but you can be a terrible designer and be completely unknown too. I've been getting into music and there are a lot of wannabes and very few "gems hidden in the dirt" or whatever - if your music is good you'll at least be able to get some decent bookings.
Mathnerd314··on Yes, Claude Code can decompile itself. Here's the source code
There is a question of originality. If the variable names, comments, etc. are preserved, then yes, it is probably a derivative work. But here, where you are starting from the obfuscated code, there is an argument that the code is solely functional, hence doesn't have copyright protection. It's like how if I take a news article and write a new article with the same facts, there's no copyright protection (witness: news gets re-reported all the time). There is a fine line between "this is just a prompt, not substantial enough to be copyrightable" and "this is a derivative work" which is still being worked out in the legal system.
Mathnerd314··on Everyone at NSF overseeing the Platforms for Wireless Experimentation is gone
That is how a large portion of the internet works, e.g. in most subreddits certain viewpoints will be instantly banned without any discussion. HN is kind of strange in that respect.
Mathnerd314··on Everyone at NSF overseeing the Platforms for Wireless Experimentation is gone
There's certainly an argument that anything the government can do, the private sector can do better. That argument would conclude that the government should indeed not exist, and consequently have no programs. The reality is more complicated, something like the microkernel vs. monolithic kernel debate, but it is hard to say that the current distribution between private and public sectors is optimal.
Mathnerd314··on Everyone at NSF overseeing the Platforms for Wireless Experimentation is gone
All of the important programs have temporary restraining orders. That's actually the standard the judge applies, "is there a possibility of irreparable harm?" (e.g. lives lost). It's not perfect but no system is.
Mathnerd314··on Everyone at NSF overseeing the Platforms for Wireless Experimentation is gone
Sometimes the only way to know something is important is to shut it off and see if anyone complains. For example, lots of stories in https://news.ycombinator.com/item?id=9629714. Now certainly the Trump administration could have been more careful, but they only have 4 years so the Facebook motto of "move fast and break things" applies.
Mathnerd314··on Carbon capture more costly than switching to renewables, researchers find
> If scenarios with different mixes of CC/DAC and WWS were performed, it would not be possible to conclude whether one is an opportunity cost. Instead, using a mixture requires assuming that both CC/DAC and WWS should be used before determining whether one has any benefit relative to the other.

Seems stupid - they are both being used, so even the business-as-usual scenario is a mixture. If indeed the 100% WWS + 0% CC/DAC scenario is better than than the 95% WWS + 5% CC/DAC scenario, then it is logical to conclude that CC/DAC is useless, but according to the tables and figures, they didn't even look at whether a 50/50 WWS + CC/DAC split would be better or worse. Yet their conclusion is still "policies promoting CC and SDACC should be abandoned". They have these really complex models but at the end of the day it is garbage in, garbage out.

Mathnerd314··on No one is disrupting banks – at least not the big ones
Well, it's half true and half false. There are a lot of "new" fintech-ish banks competing on fees, transaction speed, overdrafts, etc. - bank-type things that matter to consumers. But it's true, there are no fintech banks competing to be "too big to fail" and getting that government bailout money. You have to look at crypto for equivalents of the Federal reserve, and people don't recognize those as banks. Although I would say, Coinbase is getting pretty close to a consumer-level "crypto bank".
Mathnerd314··on People are bad at reporting what they eat. That's a problem for dietary research
There are apps, but they are incredibly inaccurate. For starters, they don't recognize the food right. Usually you have to pick from a menu of 10 items. Then they have to estimate a 3D quantity (volume) from a 2D image, then they have to estimate the density... the amazing thing though is despite all this, they are still more accurate than recall diaries.
Mathnerd314··on Is the world becoming uninsurable?
So let me try to put the author's argument in order:

(1) The author tried to get homeowner's insurance, but was denied because their home was a significant hurricane risk

(2) The author (maybe?) got insurance through a state-run FAIR program, but then cites news reports that these programs are close to insolvency (As are a significant amount of non-state-run homeowner insurance programs).

(3) The author is like, "well, if it's so hard to insure my house, maybe I should think about living somewhere else." And then generalizes to "a lot of places should be uninsurable and uninhabited - apocalypse here we come"

Mathnerd314··on How can a top scientist be so confidently wrong? R. A. Fisher and smoking (2022)
To quote: "Ironically, Fisher was proven right, albeit in a very limited way: such genes [that increase both the tendency to smoke and the tendency for lung cancer] do exist."

The actual issue was not that (Cornfield wrote a paper in 1959 showing the effect was too small). It was that Fisher continued to repeat one finding in one study, despite that said finding had not been replicated in new studies (namely, lung cancer patients described themselves as inhalers less often than the controls), and continued to obstinately ignore all the other research coming out. But it was only 3 years between Cornfield and Fisher's death in 1962, so perhaps Fisher simply did not have time to change his mind.

Mathnerd314··on How can a top scientist be so confidently wrong? R. A. Fisher and smoking (2022)
It's kind of the same situation with alcohol now: a lot of people denouncing it, the alcohol industry throwing a lot of shade, and top scientists making confident pronouncements (on both sides).
Mathnerd314··on Using coding skills to make passive income
> you just go to SE asia

I could see justifying a trip like that on a cost-of-living basis. If you go to a place like Thailand, you are going to be spending pennies on the dollar vs. the EU, even after paying incredible amounts (in local currency) for first-world conveniences like clean drinking water and internet. So in that sense if you are going to be coding and living life online, you might as well live someplace cheap IRL. But that's different from a tourist crawl where you are just spending money like water, which maybe was more your idea.

Mathnerd314··on Year old startup overloaded GitHub – Incident report
Well, the app store, for example. Sure, it is of course a good idea to comply with the app store policies to the extent possible, but ultimately there not much you can do to prevent Google or Apple saying "we don't like this app" and pulling it. For example with the UTM emulator. So how then can Google or Apple making such a decision be "your responsibility"?

As another example, let's say you build a house in a hurricane-prone area. It's your responsibility to ensure the owner buys hurricane insurance, as mandated by law. It's not your responsibility to build a nuclear-bunker-grade house that is impervious to hurricanes. It is easy to point the finger and say "you should have thought of that", but in practice it is easier to deal with such catastrophes as they happen.

Mathnerd314··on Year old startup overloaded GitHub – Incident report
Well, Github explicitly took responsibility. The first action Github did once Lovable reached out for support was "reinstate our app and apologize for the issues it caused us and our users."

And no, you are not responsible for every 3rd party service you use. Some services are unavoidable, some services are just nice-to-have, but if you can't trust a service to perform its advertised function, it is the service's fault.

Mathnerd314··on De-smarting the Marshall Uxbridge
I'm curious about price - sure, the speakers were free ($240 value), but I don't think printing up a PCB is cheap, and those are some pretty big capacitors.
Mathnerd314··on Apple's new AI feature rewords scam messages to make them look more legit
IIRC there was something about how the scammers don't bother making their scams that legitimate looking, because if they are too legit, they get a lot of people who waste the scammer's time during the next phases.

But yeah, you'd think spam filtering would be more important. I use GMail and I haven't seen a spam message in years, besides when I check my spam folder. Even there, most of the "spam" is false positives.

Mathnerd314··on Into CPS, Never to Return
Well, with CPS it just works. You start with

  let
    x k = if x then k 1 else k 2
    E k h = if h == 1 then k 2 else k 3
  in
    \k -> x (E k)
and then when you inline E into x it becomes

  let
    E k h = if h == 1 then k 2 else k 3
  in
    \k -> if x then E k 1 else E k 2
which can again be locally simplified to

  \k -> if x then k 2 else k 3
Unlike ANF, no blind duplication is needed, it is obvious from inspection that E can be reduced. That's why ANF doesn't handle this transformation well - determining which expressions can be duplicated is not straightforward.
Mathnerd314··on Into CPS, Never to Return
If there was a clear distinction between let-bindings and join points, I would be happy. But there is not - there is this contification process and, rather than proving that their algorithm finds all join points, they just say "it covers most of the ground [in practice]", and then they cite (Section 4) two other papers that suggest contification is undecidable - whether a function returns to a specific continuation or function is a behavioral property, not syntactic. Even if you accept that the dominator-based analysis is optimal, it is a global property. not a local property like in a proper IR. So what I see is a muddy mess.
← PreviousPage 3 of 34Next →