HNHacker News
TopNewBestAskShowJobs

boulos

13,551 karma · joined May 15, 2013

If you're trying to reach me about GCE (or Google Cloud generally) my username at Google is the same as this one.
submissionscomments
boulos··on Humans have caused 1.5 °C of long-term global warming according to new estimates
While people are excited about AI and datacenter use, it's still tiny in comparison to global energy consumption (1-2%) though that excludes all the crypto folks who are another 100 TWh per year or so:

> Estimated global data centre electricity consumption in 2022 was 240-340 TWh1, or around 1-1.3% of global final electricity demand. This excludes energy used for cryptocurrency mining, which was estimated to be around 110 TWh in 2022, accounting for 0.4% of annual global electricity demand.

You're hearing about the potential for a Gigwatt site, but a Gigwatt full out is less than 10 TWh per year (8960 hours/year). These things make the news, but they're pretty efficient electrically. The question is whether they have utility.

https://www.iea.org/energy-system/buildings/data-centres-and...

boulos··on SVDQuant: 4-Bit Quantization Powers 12B Flux on a 16GB 4090 GPU with 3x Speedup
The next sentence there:

> To achieve measured speedups, both weights and activations must be quantized to the same bit width; otherwise, the lower precision is upcast during computation, negating any performance benefits.

tries to explain that.

What it means though is that if you only store the inputs in lower precision, but still upcast to say bf16 or fp32 to perform the operation, you're not getting any computational speedup. In fact, you're paying for upconverting and then downconverting afterwards.

boulos··on SVDQuant: 4-Bit Quantization Powers 12B Flux on a 16GB 4090 GPU with 3x Speedup
As others have replied, this is reasonable general feedback, but in this specific case the work was done carefully. Table 1 from the linked paper (https://arxiv.org/pdf/2411.05007) includes a variety of metrics, while an entire appendix is dedicated to quality comparisons.

By showing their work side-by-side with other quantization schemes, you can also see a great example of the flavor of different results you can get with these slight tweaks (e.g., ViDiT INT8) and that their quantization does a much better job in reproducing the "original" (Figure 15).

In this application, it's not strictly true that you care to have the same results, but this work does a pretty good job of it.

boulos··on How a Mumbai Drugmaker Is Helping Putin Get Nvidia AI Chips
https://archive.is/VvBcB
boulos··on An Update on Apple M1/M2 GPU Drivers
You'll have to see what the instruction set / features are capable of, but most likely the "hardware ray tracing" support means it can do ray-BVH and ray-triangle intersection in hardware. You can reuse ray-box and ray-triangle intersection for collision detection.

The other parts of ray tracing like shading and so on, are usually just done on the general compute.

boulos··on Crossing the USA by Train
Wow, that's really unfortunate. As I just wrote in a top-level comment, Denver to Salt Lake is definitely the highlight. You get to be up high above some of those creeks and passes where the road doesn't go. I hope you're able to try again in the future and sit in the panorama car.
boulos··on Crossing the USA by Train
You can also do any segment. Having done Chicago to Emeryville, I'd suggest folks focus on the Denver to Salt Lake portion.

Chicago to Denver is mostly flat, and at least when I did it, mostly in the dark. Salt Lake to Reno is also boring, but surprisingly the section of rail around Tahoe doesn't have good views either. If you've driven on 80, or 50 or 89, you've seen better terrain.

As other comments have said though, the timeline is unpredictable, since freight takes priority. So you could easily end up in the dark around Grand Junction, CO and miss out on the views.

boulos··on C++ proposal: There are exactly 8 bits in a byte
The current proposal says:

> A byte is 8 bits, which is at least large enough to contain the ordinary literal encoding of any element of the basic character set literal character set and the eight-bit code units of the Unicode UTF-8 encoding form and is composed of a contiguous sequence of bits, the number of which is bits in a byte.

But instead of the "and is composed" ending, it feels like you'd change the intro to say that "A byte is 8 contiguous bits, which is".

We can also remove the "at least", since that was there to imply a requirement on the number of bits being large enough for UTF-8.

Personally, I'd make a "A byte is 8 contiguous bits." a standalone sentence. Then explain as follow up that "A byte is large enough to contain...".

boulos··on AMD Unveils Its First Small Language Model AMD-135M
I assume by "open inference" you mostly mean "weights available"?
boulos··on Committing to Rust in the Kernel
Rustc lowers to LLVM IR, so it can target anything LLVM can. For a long while, that meant there were obscure architectures that were excluded, but IIRC coverage has improved and the Linux kernel has decided to drop support for them.

The main push has actually been from the BSD family pushing clang (and thus LLVM) to support a broader swath of less common architectures.

boulos··on g1: Using Llama-3.1 70B on Groq to create o1-like reasoning chains
Reminder: you need to escape the * otherwise you end up with emphasis (italics here).
boulos··on Show HN: Wealthfolio: Private, open-source investment tracker
I've used ofxtools to download things locally. It basically supports the old Quicken/QuickBooks interfaces which many US banks, brokerages, and credit cards support. It's pretty clumsy though.
boulos··on Oxford commercializes its 20% more powerful solar panels in the US
The direct announcement (https://www.oxfordpv.com/news/20-more-powerful-tandem-solar-...) seems to have more nuanced wording:

> The 72-cell panels, comprised of Oxford PV’s proprietary perovskite-on-silicon solar cells, can produce up to 20% more energy than a standard silicon panel.

I think "standard silicon panel" and "up to" is doing a lot of heavy lifting. They might also be using the lab number (26.9%) rather than the 24.5% number. But then it says:

> The first Oxford PV panels available on the market have a 24.5% module efficiency, offering performance significantly above traditional silicon technology.

AFAICT, the real news isn't a massive efficiency win, but that these are actually going into production.

boulos··on Creating invariant floating-point accumulators
This seems to keep coming up, and I see confusion in the comments. There is a standard: IEEE 754-2008. There are additional things people add like approximate reciprocals and approximate sqrt. But if you don't use those, and you don't make an association error, you get consistent results.

The question here with association for summation is what you want to match. OP chose to match the scalar for-loop equivalent. You can just as easily make an 8-wide or 16-wide "virtual vector" and use that instead.

I suspect that an 8-wide virtual vector is the right default for people currently, since systems since Haswell support it, all recent AMD, and if you're using vectorization, you can afford to pay some overhead on Arm with a double-width virtual vector. You don't often gain enough from AVX512 to make the default 16-wide, but if you wanted to focus on Skylake+ (really Cascadelake+) or Genoa+ systems, it would be a fine choice.

boulos··on America’s Transit Exceptionalism
Note that "this Fall", Caltrain's electrified service is finally happening:

  https://www.caltrain.com/projects/electrification/project-benefits/caltrain-electrified-service-plan
Most of the stations will get faster with more frequent service. I think the electric EMUs also means the ancient clunker ones will be gone.
boulos··on Run CUDA, unmodified, on AMD GPUs
Note: gold uses troy ounces, so adjust by ~10%. It's easier to just use grams or kilograms :).
boulos··on SF's AI boom can't stop real estate slide, as office vacancies reach new record
In case it's not obvious, that's an annual number.
boulos··on OpenAI was hacked year-old breach wasn't reported to the public
Huh? I just went to OpenAI.com and there is a little "Security" link in the bottom pile that points to https://openai.com/security-and-privacy/ .

Under "Reporting security issues" it points you to a bug bounty page: https://bugcrowd.com/openai with a bunch of explanations.

I'm guessing if you also just send an email to security@openai.com it'll go to someone. Using Bugcrowd just seems like a nice way to also run a bug bounty as part of their normal intake.

boulos··on How Waymo outlasted the competition and made robo-taxis a real business
For folks that are interested in the business specifically, we have some open roles listed at https://waymo.com/careers/ with the word Commercialization in the titles.
boulos··on How Waymo outlasted the competition and made robo-taxis a real business
Disclosure: I work for Waymo.

No, there isn't. We do have a team of folks to support situations where we aren't confident and choose to "phone a friend". This recent blog post covers some of it in more detail:

https://waymo.com/blog/2024/05/fleet-response/

Most importantly: at no point does someone remotely "drive" the vehicle. They can direct it to say "hey make a u-turn and go to this new point", but they aren't remotely driving.

boulos··on How Waymo outlasted the competition and made robo-taxis a real business
Disclosure: I work for Waymo.

We handle both dense fog and heavy rain on the latest vehicles. The best blog post is probably https://waymo.com/blog/2021/11/a-fog-blog/ but you can find a lot of videos in the rain.

Snow and very cold weather is a challenge for sensor cleaning. We've done some testing in both NYC and Buffalo (https://waymo.com/blog/2023/11/road-trip-how-our-cross-count...) to collect data.

boulos··on Codestral: Mistral's Code Model
This is why I prefer the term "weights available" just like "source available". It makes it clear that you can get your hands on the copy, you could run this exact thing locally if they go out of business, etc. but it is definitely not open in the OSS sense.
boulos··on Companies with return-to-office mandates face losing their most valuable workers
How have you handled synchronization? I agree with async-primary, but it's also important to get some overlap depending on the situation. Did folks in Manila work off hours to align a bit with San Francisco?
boulos··on New exponent functions that make SiLU and SoftMax 2x faster, at full accuracy
Yes, but a direct `exp` implementation is only like 10-20 FMAs depending on how much accuracy you want. No gathering or permuting will really compete with straight math.
boulos··on Add sysctl to disable Nagle's algorithm (RFC 896 – Congestion Control)
I kind of love that this continues to haunt you :).
boulos··on Why the CORDIC algorithm lives rent-free in my head
Depends on what you're doing. The main issue here is reductions / accumulations.

That is, if you have a bunch of floats like:

  float sum = 0.f;
  for (int i = 0; i < N; i++) {
    sum += x[i];
  }
and that you vectorize that to something like (typing in this comment, errors are likely):

  __mm256 sum_8wide = _mm256_setzero_ps();
  for (int i = 0; i < N/8; i++) {
    sum_8wide = _mm256_add_ps(sum_8wide, _mm256_load_ps(x[8*i]);
  }
  // Now sum up the 8 values to get the final sum
  float sum = _mm256_hadd_ps(...);
that will result in a different accumulation than if you did those are 4-wide and then a reduction. The usual solution is to either use the lowest common denominator (e.g., use AVX instead of AVX2) or more performance oriented, use the 4-wide SIMD units on ARM to "emulate" an 8-wide virtual vector (~15 years since I wrote NEON... and again, this is in a comment):

  float32x4_t sum_lo = ; // zero;
  float32x4_t sum_hi = ; // zero;
  for (int i = 0; i < N/8; i++) {
    sum_lo = vaddq_f32(sum_lo, vload1q_f32(x[8*i));
    sum_hi = vaddq_f32(sum_hi, vload1q_f32(x[8*i + 4]));
  }
  // Reduce the sum in the same order
You would want to write a "virtual SIMD" wrapper library, so you don't do this manually in lots of places.
boulos··on Why the CORDIC algorithm lives rent-free in my head
The same reasoning applies though. The compiler is just another program. Outside of doing constant folding on things that are unspec'ed or not required (like mmsqrtps and most transcendentals), you should get consistent results even between architectures.

Of course the specific line linked to in that GH issue is showing that LLVM will attempt constant folding of various trig functions:

https://github.com/llvm/llvm-project/blob/faa43a80f70c639f40...

but the IEEE-754 spec does also recommend correctly rounded results for those (https://en.wikipedia.org/wiki/IEEE_754#Recommended_operation...).

The majority of code I'm talking about though uses constants that are some long, explicit number, and doesn't do any math on them that would then be amenable to constant folding itself.

That said, lines like:

https://github.com/llvm/llvm-project/blob/faa43a80f70c639f40...

are more worrying, since that may differ from what people expect dynamically (though the underlying stuff supports different denormal rules).

Either way, thanks for highlighting this! Clearly the answer is to just use LLVM/clang regardless of backend :).

boulos··on Why the CORDIC algorithm lives rent-free in my head
That blog post is now a decade old, but includes an important quote:

> The IEEE standard does guarantee some things. It guarantees more than the floating-point-math-is-mystical crowd realizes, but less than some programmers might think.

Summarizing the blog post, it highlights a few things though some less clearly than I would like:

  * x87 was wonky

  * You need to ensure rounding modes, flush-to-zero, etc. are consistently set

  * Some older processors don't have FMA

  * Approximate instructions (mmsqrtps et al.) don't have a consistent spec

  * Compilers may reassociate expressions

For small routines and self-written libraries, it's straightforward, if painful to ensure you avoid all of that.

Briefly mentioned in the blog post is IEEE-754 (2008) made the spec more explicit, and effectively assumed the death of x87. It's 2024 now, so you can definitely avoid x87. Similarly, FMA is part of the IEEE-754 2008 spec, and has been built into all modern processors since (Haswell and later on Intel).

There are still cross-architecture differences like 8-wide AVX2 vs 4-wide NEON that can trip you up, but if you are writing assembly or intrinsics or just C that inspect with Compiler Explorer or objdump, you can look at the output and say "Yep, that'll be consistent".

boulos··on Cold brew coffee in 3 minutes using acoustic cavitation
Maybe ask them to make you "an Americano, but with cold brew coffee"? Trick them into it being an on-the-menu item.
boulos··on Jolie, the service-oriented programming language
Agreed. I should have said so, too.
← PreviousPage 5 of 34Next →