HNHacker News
TopNewBestAskShowJobs

smueller1234

1,121 karma · joined November 25, 2016

submissionscomments
smueller1234··on A beginning for mathematics
Here's another example of a similar but slightly different failure mode of the live defense: The department where I got my degree had a practice around PhD defense where there would be a prof each representing the four research pillars + the dean. It was well known that certain professors disliked each other enough that they would snipe each other's candidates. You'd just hope not to work for one whose enemies got picked for the panel. If you were a superstar, you'd pass, but if you were an okay candidate or someone strong but with stage fright, you be could toast even if you did good research.

The defense was only briefly about the actual thesis, then switched over to whatever research interest the committee members had, and they'd drill into their pet subjects. This was in physics, hardly politics forward normally, and at a very well respected university.

smueller1234··on Cherenkov Radiation - traveling faster than light
Cherenkov radiation is also a key mechanism behind ultra high energy cosmic ray (UHECR) detectors. UHECR are understood to be nuclei, not gamma rays. For ground based detectors, light tight water tanks are fitted with photo multipliers to capture the flash of cherenkov light even secondary particles pass through the water.

Additionally, they can use telescopes pointed at the atmosphere, using the atmosphere as a calorimeter, basically. The primary signal those telescopes look for is fluorescence, but when the direction of travel points at the telescope, cherenkov light far outstrips it in brightness, so it has to be included in the event reconstruction.

Most prominent contemporary example is the Pierre Auger Observatory https://auger.org

smueller1234··on Biggest dark matter detector spots a single weird particle
As another comment points out: that's nothing compared to high energy/particle physics!

The Pierre Auger Observatory certainly is a large collaboration for the astroparticle physics domain though. It's a big international collaboration.

Quick anecdote: my name (S Mueller) is not on that paper's author list because we had a rule that you had to be in the collaboration for a year before getting authorship. You stayed on for a year after leaving. Very reasonable! At the time I was nonetheless a bit bummed about missing out on the big Science paper. I guess I'm on the retraction though ;)

smueller1234··on Biggest dark matter detector spots a single weird particle
Or this[2] 2007 Science paper on ultra high energy cosmic ray source candidates ("anisotropy") that we had to retract because significance started dropping almost the day the paper was approved.

It was a fascinating experience as a junior member to follow the collaboration internal conversation and investigation on this, because a lot of extremely principled scientists were clearly deeply worried about losing their hard earned reputation. In the end, I am convinced that we were simply unlucky.

[2] https://arxiv.org/pdf/0712.2843

smueller1234··on ASML became the chokepoint for cutting-edge chips
The timelines matter as well: They were working on EUV at Zeiss (who make the lensing/mirroring systems) already in 2005. That's about 20 years of development.
smueller1234··on The $100B megadeal between OpenAI and Nvidia is on ice
Even if it's success rather than money, you still have survivorship bias to contend with, so it's not really much of a helpful distinction.
smueller1234··on Internet Archive's Storage
IIRC, the most recent and most technical public content we (Google) have published on Colossus are these:

https://cloud.google.com/blog/products/storage-data-transfer...

https://cloud.google.com/blog/products/storage-data-transfer...

Facebook's published content on Tectonic is quite good and I think it's well more recent than 2010-14.

(Current Google employee, just pointing to public content, hope that's helpful.)

smueller1234··on Kids Rarely Read Whole Books Anymore. Even in English Class
Slight problem with that if you would like to live in a functioning, thriving democracy: democracy in the sense of "one person, one vote" requires or at least greatly benefits from a broadly educated population. It's not sufficient, but very likely necessary.
smueller1234··on Wolfram Compute Services
You're right -- the theoretical particle physicists at my faculty were using Mathematica very heavily when I was still in academia and maintained a dedicated compute cluster for it.

They really did not appreciate the debugging experience, but maybe that's improved in 15 years. :)

smueller1234··on How AWS S3 serves 1 petabyte per second on top of slow HDDs
I realize you're making a general point about space/IO ratios and the below is orthogonal, no contradiction.

It's actually a lot less user-facing per disk IO capacity that you will be able to "sell" in a large distributed storage system. There's constant maintenance churn to keep data available: - local hardware failure - planned larger scale maintenance - transient, unplanned larger scale failures (etc)

In general, you can fall back to using reconstruction from the erasure codes for serving during degradation. But that's a) enormously expensive in IO and CPU and b) you carry higher availability and/or durability risk because you lost redundancy.

Additionally, it may make sense to rebalance where data lives for optimal read throughput (and other performance reasons).

So in practice, there's constant rebalancing going on in a sophisticated distributed storage system that takes a good chunk of your HDD IOPS.

This + garbage collection also makes tape really unattractive for all but very static archives.

smueller1234··on Anandtech.com now redirects to its forums
I think Chips and Cheese is more like a fine replacement for realworldtech.com sans the toxic and highly educational and entertaining forums. Anandtech was much more accessible to the general tech public, but also more commercial and thus hit and miss on the content (no judgement intended, gotta eat).
smueller1234··on Colossus for Rapid Storage
Google's internal systems have been written against the Colossus semantics for many, many years and thus benefit from it's upsides (performance, cost efficiency, reliability, strong isolation for a multi tenant system, ability to scale byte and IO usage fairly independently, tremendously good abstraction against and automation of underlying physical maintenance, etc) while not really having too much of an issue with any of the conscious trade-offs (like no random writes).

On the other hand, if you've been building your applications against expectations of different semantics (like POSIX), retrofitting this into your existing application is really hard, and potentially awkward. This is (IMO) why there hasn't been an overtly Colossus based Google Cloud offering previously. (Though it's well publicized that both Persistent Disk and GCS use Colossus in their implementation.)

One of the reasons why it would be extremely hard to just set up or build CFS elsewhere or on a different abstraction level is that while it may look quite achievable to implement the high level architecture, there is vast complexity in the practical implementation side. The tremendous user isolation it affords for an MT system, the resilience it has against various types of failures and high throughput planned maintenance, the specialization it and its dependencies have to use specific hardware optimally.

(I work on Google storage part time, I am not a Colossus developer.)

smueller1234··on Colossus for Rapid Storage
Concur, Colossus is one of the examples where Google built what almost feels like magic technology. I work on Google Storage (among other things), and I've wished for a Cloud offering that exposes Colossus for years.

I don't know that it took "AI branding" to convince anybody. I think these workloads potentially enabled additional demand/market for such a product that may not have been there before.

One of the challenges with exposing native Colossus was always that it's just different enough from how people elsewhere are used to use Storage that there was a lot of uncertainty about the addressable market of a "native" Colossus offering. It's not a POSIX file system. Some of the specific differences (eg. no random writes) are part of what makes Colossus powerful and performant on HDDs, but it means you have to write your application to work well within its constraints. Google has been doing that for a long time. If you haven't, even if it's an amazing product, is it worth rewriting your applications or middleware?

Rapid Storage basically addresses this by adding the object store API on top if it (TIL from this thread that there's a lower abstraction client in the works as well).

Anyway, the team behind this is awesome. Awesome tech, awesome people. Seeing this launched at Next and seeing some appreciation on HN makes me very grateful.

smueller1234··on Oracle customers confirm data stolen in alleged cloud breach is valid
4% of revenue is terrifying for large corporations.
smueller1234··on It is no longer safe to move our governments and societies to US clouds
https://www.s3ns.io/en

This is Google + Thales doing the 3rd party operator model, with the operator being a subsidiary of Thales and not Google.

(NB: I work for Google in the EU.)

smueller1234··on MySQL at Uber
Was about to say that we managed upwards of 4k servers worth of MySQL databases (as in 4k baremetal servers worth, not 4k small VMs each having a small MySQL) using "Orchestrator" at Booking.com ten years ago. I checked if it was still kicking before writing this and found that the GitHub repo was archived last year. The next Google hit I find is this article from three days ago:

https://www.percona.com/blog/orchestrator-for-managing-mysql...

Please do some additional research into the state of maintenance of that piece of tech before jumping on it. But it certainly did a lot of powerful things for us back in the day. The automatic promotion of followers was key to our deployment.

smueller1234··on The perils of transition to 64-bit time_t
They make it easier, but just at a source code level. They're not a real (and certainly not full) abstraction. An example that'll be making it obvious: if you replace the underlying type with a floating point type, the semantics would change dramatically, fully visible to the user code.

With larger types that otherwise have similar semantics, you can still have breakage. A straightforward one would be padding in structs. Another one is that a lot of use cases convert pointers to integers and back, so if you change the underlying representation, that's guaranteed to break. Whether that's a good or not is another question, but it's certainly not uncommon.

(Edit: sibling comments make the same point much more succinctly: ABI compatibility!)

smueller1234··on Cisco slashes thousands of workers as it announces yearly profit of $10.3B
It's actually quite likely something else (unless it's just an excuse to reap short term savings): in a company large enough, deciding to shift staffing from one area to another is hard. As an executive with thousands of staff, you can tell your management team to each cough up a certain number or (somewhat) suitably qualified people. But again, if large enough, incentives diverge, so you don't necessarily end up with the top talent you thought you needed for your big new thing.

An "easy" solution is to do a layoff, then open roles elsewhere, allowing for selection.

It's commonly practiced across the large companies in the industry.

(Not speaking for my employer)

smueller1234··on How Meta trains large language models at scale
Multiple types of TPUs.

(I work for Google, but the above is public information.)

smueller1234··on 93% of paint splatters are valid Perl programs (2019)
Former Perl language contributor here. A sibling comment to this already pointed out that you must use strict mode with Perl to retain your well being.

The two languages certainly both have their terrible warts. I think in the implicit conversion gotchas, JS is actually markedly worse. Perl has polymorphic values, but somewhat typed (for its basic types) operators (eg "eq" for strings, "==" for integers). JS has both implicit value type conversions and overloaded operators. That leads to an unholy level of indeterministic mess.

smueller1234··on 93% of paint splatters are valid Perl programs (2019)
Many moons ago, I made a case for making strict mode the default in Perl. We settled on the current backwards compatibility compromise, which is that breaking changes are hidden behind a minimum version toggle:

Eg. putting "use v5.14.0;" or similar on top of your file (or compilation unit/scope) will indeed turn on strict mode for you, along with adding a number of features as well.

At the time, also auto-toggling warnings was considered unacceptable because technically, using the warnings pragma anywhere had some edge case action at a distance. This has been remedied in some later release after I wasn't involved in the language development anymore, and from some more recent version, warnings are also part of the standard import.

I imagine you (TheDauthi) already know that, though.

smueller1234··on DwarFS – Deduplicating Warp-Speed Advanced Read-Only File System
It's not an idle use case by the way. mhx wrote and maintains a library that provides important backwards compatibility for native (typically C based) extensions for Perl across decades of language releases.
smueller1234··on I discovered a critical exploit in ZeroMQ with mostly pure luck
See my response to a sibling of the comment you're responding to. The library had shocking code quality issues. It's unlikely that they're all peachy now.

The other side of this is that while Pieter's writing was marketing genius, it was also woefully understating the complexity of any practical use case. The way I tried to summarize that to folks who were keen to try zeromq then was that they should start at the back of the book with the most complex example, and that's by far the simplest setup that they could hope to end up with once they start thinking about putting something into production. And everything leading up to that - a book no less - was exclusively educational/toy use cases.

smueller1234··on I discovered a critical exploit in ZeroMQ with mostly pure luck
Zeromq will have changed a lot since then, but some time in the 2010s, I prototyped a system using it (which was going to be a major production system in a large tech company) and had weird unexpected blocking issues with it. To debug, I sat down to read a bunch of the zeromq code, just to realize that it was using assert() to handle wire protocol errors (unrelated to the blocking bug).

I've never dropped a piece of software as quickly as that.

smueller1234··on America's Great Poet of Darkness: A Reconsideration of Robert Frost at 150
"Before I built a wall I’d ask to know What I was walling in or walling out, And to whom I was like to give offense."

Only poem I can cite by heart a quarter of a century after spending time with it in school.

smueller1234··on The life and death of open source companies
I haven't ever used a 3d printer. But your comment made me realize that if PrusaSlicer is based slic3r, it's actually also using software that I wrote many, many years ago.

That's another side of open source: if you don't rely on it to make a living (though it did help in getting my first job as a developer!), there's that pure joy in seeing your software get picked up and used by others. This little discovery made my day.

smueller1234··on P vs NP: The most important unsolved problem in computer science
Apologies for nitpicking, but n^m doesn't mean m loops (ie n^1000 dient mean 1000 passes): that would be mn (1000n in your example). I think your intuition argument kind of breaks down there.
smueller1234··on Five-Disk Floppy RAID: 4MB of Blistering Fast Storage (2009)
It's easy enough to repeat, but I'd be leaking sensitive info that way. Sorry :/
smueller1234··on Five-Disk Floppy RAID: 4MB of Blistering Fast Storage (2009)
Sorry for a second response to the same comment. While on the topic of entertainment, consider the opposite of putting erasure codes on many weird devices!

Years ago, some colleagues and I did the napkin math over lunch to estimate what a single hard drive would look like that could store all of Google's data. This was for a humorous internal talk. The linear speeds of the outer sectors were enough to leave the solar system (not just higher than Earth's escape velocity). Good laughs were had, I think.

smueller1234··on Five-Disk Floppy RAID: 4MB of Blistering Fast Storage (2009)
I would say for sure LTO tape drives.
Page 1 of 13Next →