HNHacker News
TopNewBestAskShowJobs

germanjoey

225 karma · joined March 27, 2018

submissionscomments
germanjoey··on Did Claude increase bugs in rsync?
IMO "bugs per commit" is even worse than that, because, in addition to what you say, it also hides the extraordinary spike of commit activity of a project that had previously been stable. [0]

It is the exact metric you'd choose if you wanted to make the current situation of rsync look like not a big deal.

[0] https://github.com/RsyncProject/rsync/graphs/commit-activity

germanjoey··on We Made CUDA Optimization Suck Less
TBH, the 2x-4x improvement over a naive implementation that they're bragging about sounded kinda pathetic to me! I mean, it depends greatly on the kernel itself and the target arch, but I'm also assuming that the 2x-4x number is their best case scenario. Whereas the best case for hand-optimized could be in the tens or even hundreds of X.
germanjoey··on The Missing Nvidia GPU Glossary
This is really incredible, thank you!
germanjoey··on Trillium TPU Is GA
Sambanova's RDU is a dataflow processor being used for ML/AI workloads! It's amazing and actually works.
germanjoey··on Llama 3.1 405B now runs at 969 tokens/s on Cerebras Inference
Pretty amazing speed, especially considering this is bf16. But how many racks is this using? The used 4 racks for 70B, so this, what, at least 24? A whole data center for one model?!
germanjoey··on Cerebras Trains Llama Models to Leap over GPUs
the title says "Cerebras Trains Llama Models"...
germanjoey··on Cerebras Inference now 3x faster: Llama3.1-70B breaks 2,100 tokens/s
They said in the announcement that they've implemented speculative decoding, so that might have a lot to do with it.

A big question is what they're using as their draft model; there's ways to do it losslessly, but they could also choose to trade off accuracy for a bigger increase in speed.

It seems they also support only a very short sequence length. (1k tokens)

germanjoey··on Civilization VII recommends 16 cores and 32GB RAM for 4K gameplay
Simply increasing processing power for the AI isn't enough. Gameplay mechanics are intimately related to the capabilities of the AI.

For example, when they redesigned combat around the 1-Unit-Per-Tile (1UPT) mechanic for CIV 5, this crippled the ability of the AI to wage war. That's because even if a high-difficulty AI could out-produce the player in terms of military, they were logistics-limited in their ability to get those units to the front because of 1UPT. That means that the AI can't threaten a player militarily, and thus loses it's main lever in terms of it's ability to be "difficult."

Contrast this to Civ 4, where high-difficulty AIs were capable of completely overwhelming a player that didn't take them seriously. You couldn't just sit there and tech-up and use a small number of advanced units to fend off an invasion from a much larger and more aggressive neighbor. This was especially the case if you played against advanced fan-created AIs.

I'm hoping they get rid of 1UPT completely for Civ 7, but I have a feeling that it is unlikely because casual players (the majority purchaser for Civ) actually like that 1UPT effectively removes tactical combat from the game.

germanjoey··on We fine-tuned Llama 405B on AMD GPUs
How are you verifying accuracy for your JAX port of Llama 3.1?

IMHO, the main reason to use pytorch is actually that the original model used pytorch. What can seem to be identical logic between different model versions may actually cause model drift when infinitesimal floating point errors accumulate due to the huge scale of the data. My experience is that debugging an accuracy mismatches like this in a big model is a torturous ordeal beyond the 10th circle of hell.

germanjoey··on A post by Guido van Rossum removed for violating Python community guidelines
Looks like some kind of power play...

Originally discussed here: https://news.ycombinator.com/item?id=41234180

germanjoey··on Model Explorer: intuitive and hierarchical visualization of model graphs
Is there a demo of a model visualized using this somewhere? Even if it's just a short video... it's hard to tell what it's like from screenshots.
germanjoey··on Groq CEO: 'We No Longer Sell Hardware'
cost effective in what sense? groq doesn't achieve high efficiency, only low latency. but that's not done in a cost-effective way. compare sambanova achieving the same performance with 8 chips instead of 568, and with higher precision.
germanjoey··on Try SambaNova chat: 1T param LLM, 500 tokens/SEC
We're showing off our 1.05T param Composition of Experts LLM! It's 150 experts running on 1 node consisting of 8 SN40L RDU chips.

Each of our nodes has a huge amount of DDR attached, in addition to copious amounts of on-chip HBM and SRAM. This allows the system to switch between a variety of different models of different sizes and architectures at lightning speed. A highlight is one based on Llama2 7b, similar to the Groq demo, but executing with bf16/fp32 instead of int8. (And using only 8 chips instead of 568!)

germanjoey··on A hacker's guide to language models [video]
Sambanova just launched something similar to what you're describing. It's a demo of their new chip running a 1T param MoE model 150 7B llama2s, each retrained to be an expert in a different topic. So one of them is a "law" expert, another on "physics", etc.

They've got a video here [1] (scroll down slightly) that compares it against a 180B Falcon model that's running on GPUs on HuggingFace. The MoE results are not only just as good quality-wise, but also ridiculously fast. Like, nearly instant. A big benefit is that the experts can be swapped-out and retrained with new data, which is obviously not as easy with the more monolithic 180B model.

[1] https://sambanova.ai/launch2023

germanjoey··on GPT-4
welp,

This report focuses on the capabilities, limitations, and safety properties of GPT-4. GPT-4 is a Transformer-style model [33 ] pre-trained to predict the next token in a document, using both publicly available data (such as internet data) and data licensed from third-party providers. The model was then fine-tuned using Reinforcement Learning from Human Feedback (RLHF) [34 ]. Given both the competitive landscape and the safety implications of large-scale models like GPT-4, this report contains no further details about the architecture (including model size), hardware, training compute, dataset construction, training method, or similar.

germanjoey··on GPT-4
How big is this model? (i.e., how many parameters?) I can't find this anywhere.
germanjoey··on The maze is in the mouse: what ails Google
I worked with the author for a couple of years, pre- and post- acquisition, and I have to admit that he drove me somewhat crazy sometimes too. Leaving that aside, I also had an immense amount of personal respect for him as I could see how much he very genuinely cares about what he is doing. And, that's actively doing his best to do right by his customers. I think the author is 110% spot-on with his critique of Google here.
germanjoey··on Amazon to Lay Off over 17,000 Workers, More Than First Planned
What's the new performance process?
germanjoey··on The type system is a programmer's best friend
> You don't introduce more coupling, you don't the coupling that already exists.

This is true at the code level. But at the system-design level, this documentation is the extra coupling.

I feel like it's important to understand this. I agree with the original commenter; in engineering, nothing is truly free. In many cases, this extra coupling helps keep a system strong and stable, like extra nails holding planks of wood together. In other cases, you may find that part of a system's spec actually missed the mark and now needs to be ripped up and redone. That extra coupling might now work against you!

Again, that doesn't mean that it wasn't worth having it. It is just important to understand tradeoffs in engineering.

germanjoey··on Steve Yegge Joins as Head of Engineering of Sourcegraph
It is interesting reading that second paragraph many years later. Most of the things that Steve Yegge brags about that Google "does right" (e.g. how they do recruiting, their engineering "standards", SREs running products rather than the engineers, their cushy offices and benefits packages, etc) now read to me the opposite of the intended way. As in, they read more like a list of reasons why Google slowly declined from being the shining city on the hill to the dysfunctional embarrassment it is today.
germanjoey··on In defence of garlic in a jar
Great post; this is how I felt about it too.
germanjoey··on The Edited Latecomer’s Guide to Crypto
The article (or, rather, the commentary in the link above on the article) talks about the fallacious notion of "market cap" in regards to cryptocurrencies. That is to say, e.g., multiplying the number of bitcoins in existence times the current market price is a silly metric because the entire market would never be able to cash-out at that maximum price.

What I was wondering was: is there a better number? e.g., is there a way to calculate the amount of USD put into a cryptocurrency across a timeframe? What I'm imagining is a metric like (sum of all bitcoins bought by USD purchase price) - (sum of all bitcoins sold by USD sale price) = amount of USD that has been put "into" bitcoin. That first glance, one might expect this number to equal zero, but it should be greater than zero because of the new coins created by mining.

germanjoey··on The Competitive Duality of Slay the Spire
It means you play through the four characters in order in four successive runs, e.g. ironclad -> silent -> defect -> watcher, rather than e.g. just playing the watcher (the character generally considered the easiest to win with at high skill levels) 4 times in a row.
germanjoey··on Nature Neuroscience offers open access publishing for $11k per article
...and what service is that?
germanjoey··on Ask HN: Whatever happened to Wolfram Alpha?
You're severely underselling Google's incapability, e.g. https://i.imgur.com/UoIZSU2.png
germanjoey··on A Tale of Two Optimisations
"whitespacer" reminds me of Damian Conway's classic "Acme::Bleach" perl module...

https://metacpan.org/pod/Acme::Bleach

germanjoey··on MangaDex infrastructure overview
Many manga fans have a love/hate relationship with mangadex. On one hand, it's provided hosting for countless hours of entertainment over the years. Their "v3" version of the site was basically perfect from a usability point of view, to the point that the entire community chose to unite itself under its flag.

On the other hand, directly because of the above, their hasty self-inflicted take down earlier this year nearly killed the entire hobby. Many series essentially stopped updating for the ~5 months the site was down, and many more are likely never coming back again.

The decision to suddenly take the site down for a full site rewrite feels completely inexplicable from the outside. (A writeup the above or the previous one[1], both of which read like they were written by a Google Product Manager, especially don't help as they conspicuously avoid any comment to the one question on everyone's mind: "leaving aside the supposed security issues with the backend, why on earth also rewrite and redesign the entire front end from scratch at the same time?")

[1] https://mangadex.dev/why-rebuild/

germanjoey··on Publish and Perish
The actual impact of gain-of-function research papers isn't actually what is being critiqued here. That's putting the cart before the horse. What is being criticized is that the authors of such papers are pushing their research efforts into this area because they see it as their highest expected value avenue for generating the "impact" they need to secure funding, tenure, etc.
germanjoey··on Dungeon Crawl Stone Soup
The state of DCSS is interesting because it's like a software Ship of Theseus. It's been actively hacked (and slashed) upon for so long that it is hard to claim that its the same game, even if the git commit history shows a direct lineage otherwise. Or, at least, the aspects of the game that I once so greatly admired (which I wrote about here [1]) are longer present. Only a superficial similarity remains. A very opinionated group of core developers with a minimalist design philosophy has driven the project for years now, and the game now is not the game I knew.

[1] https://rhizzone.net/articles/crawl/

germanjoey··on Unsettling capital letters
My five year old son does this!! He draws each extra horizontal line with ever increasing gusto...
Page 1 of 2Next →