HNHacker News
TopNewBestAskShowJobs

bufo

1,594 karma · joined August 20, 2010

submissionscomments
bufo··on C++20 Modules: Practical Insights, Status and TODOs
Were you using pre-compiled headers before?
bufo··on Gemini CLI
Grateful that this one supports Windows out of the box.
bufo··on Microsoft begins turning off uBlock Origin and other extensions in Edge
Brave.
bufo··on Starship Flight 5: Launch and booster catch [video]
The plan is to have many, many Mechazillas.
bufo··on Russ Cox is stepping down as the Go tech lead
Because Go has massive traction both inside and outside of Google, whereas Dart/Flutter never got big traction.
bufo··on Reproducing GPT-2 in llm.c
The RTX 4090 has about the same BF16 Tensor Core TOPs than the A100, assuming 50% MFU (like the A100 40 GB PCIe) it would take 8x longer on 1 RTX 4090 vs 8x A100 80GB SXM, so 12 hours. Datasheet here for the TOPs https://images.nvidia.com/aem-dam/Solutions/geforce/ada/nvid... 50% MFU should be achievable on the 4090.
bufo··on Perplexica: Open-source Perplexity alternative
I recommend Phind.com, it’s been much better and faster for me than Perplexity Pro. I typically use their custom 70B model but you can also use GPT4 o or Turbo, or Claude 3 Opus.
bufo··on Vice website is shutting down
It takes a while to take down job posts. Everyone likely learned the decision recently. I don’t think the employees who are going to be laid off care about updating the job posts at the moment…
bufo··on Vice website is shutting down
“Vice website is shutting down” is 100% accurate. Note that they will also lay off most of their staff.
bufo··on Sam Altman Seeks Trillions of Dollars to Reshape Business of Chips and AI
Was Google a chip designer before the first TPU?
bufo··on Seniors spend the equivalent of 3 weeks a year on health care, study says
That doesn’t seem like a bad number.
bufo··on --libcurl
Seriously!!
bufo··on How to do OCR on a Mac using the CLI or just Python
Way, way better than Tesseract!
bufo··on Show HN: A pure C89 implementation of Go channels, with blocking selects
Kqueue! Not the same design or as flexible as io_uring though.
bufo··on Show HN: A pure C89 implementation of Go channels, with blocking selects
Build 22000 is Windows 10 21H2.
bufo··on Intel launches Core Ultra processors
Actual benchmarks and useful info here https://youtu.be/WH-qtuVRS2c
bufo··on Show HN: A pure C89 implementation of Go channels, with blocking selects
Oh yeah I meant io_uring too. Plus Windows copied it so you can implement things very similarly for Windows.
bufo··on Show HN: A pure C89 implementation of Go channels, with blocking selects
Great! I was looking into something like this. I assume ending up with epoll will be better?
bufo··on How many lines of C it takes to execute a + b in Python
It’s about 100 for x86_64 https://www.computerenhance.com/p/waste
bufo··on Apple unveils M3, M3 Pro, and M3 Max
It was pretty hard to saturate the memory bandwidth on the M2 on the CPU side (not sure about the GPU).
bufo··on FlashAttention-2, 2x faster than FlashAttention
Tri Dao and Tim Dettmers ftw
bufo··on More than 75% of Steam games tested are playable or verified on the Steam Deck
I don’t mind the slow burn at all! I however did not like being forced to do frustrating platforming / movements while having to start from scratch every time I run out of time or die.
bufo··on More than 75% of Steam games tested are playable or verified on the Steam Deck
I had the same experience after the jellyfish, at which point I gave up and just watch YouTube to know what happens.
bufo··on More than 75% of Steam games tested are playable or verified on the Steam Deck
I found the world and exploration very fun, but the “platforming” challenges were extremely frustrating for me, and I didn’t enjoy the random messages and the miscellaneous details that you translated. Basically the gameplay loop was filled with things that didn’t quite ring with me, even though the overarching design and story were compelling.
bufo··on More than 75% of Steam games tested are playable or verified on the Steam Deck
I also love that feature for the exact same reason!

Somewhat disappointed with Outer Wilds though ;)

bufo··on Apple Introduces M2 Ultra
There is a difference. We train with large batch sizes these days. The ANE silicon size is tiny and can't do the large matrix multiplications for big LLMs with or without a batch size higher than 1. Meaning that it cannot saturate the RAM bandwidth and that you're better using off the much bigger GPU on the Apple die.
bufo··on Apple Introduces M2 Ultra
Yes, you are correct in that the ANE does have the equivalent of tensor cores and that I didn’t mention that. I just don’t expect it to be usable beyond inference because the number of compute units will not work for batches in medium/large/huge networks. That’s obviously by design! The ANE silicon size is tiny compared to the GPU area. I wouldn’t be actually surprised if Apple strategically only invests in using their GPU for LLM (1B+ params) work.

Note that if you are currently using CoreML for LLMs all the work is done in the GPU.

bufo··on Apple Vision Pro: Apple’s first spatial computer
This is completely wrong. Learn about lenses.
bufo··on Apple Introduces M2 Ultra
The neural engine has severe limitations at the moment. I tried using it for BERT about a year ago and kept crashing its API because of "out of memory" issues. The theoretical TOPs you mention also don't necessarily translate into usable TOPs because of memory bandwidth and caches. This is why for example the comparison of the M1 Max with a RTX 3090 was completely off.
bufo··on Apple Introduces M2 Ultra
The memory bandwidth is still a bit lower than Nvidia's best cards, and it doesn't have the equivalent of Tensor Cores. If they wanted they could compete, but it's clearly not their desire. They build consumer end products.
Page 1 of 3Next →