HNHacker News
TopNewBestAskShowJobs

aetimmes

576 karma · joined April 17, 2013

@NVIDIA. Former employers: Princeton University, Bloomberg, Twitter, SeatGeek.
submissionscomments
aetimmes··on Meta disbanded its Responsible AI team
People do what they are incentivized to do.

Engineers are incentivized to increase profits for the company because impact is how they get promoted. They will often pursue this to the detriment of other people (see: prioritizing anger in algorithmic feeds).

Doing Bad Things with AI is an unbounded liability problem for a company, and it's not the sort of problem that Karen from HR can reason about. It is in the best interest of the company to have people who can 1) reason about the effects of AI and 2) are empowered to make changes that limit the company's liability.

aetimmes··on Val, a high-level systems programming language
Communities aren't monoliths.

The groups of people who criticize real tangible languages and people who have the skills to critique the design of abstract language design don't necessarily overlap.

aetimmes··on We Don’t Do Legacy (2012)
Maybe not MIT, but there are many universities whose bottom lines are significantly impacted by the success of their sports programs and the television deals that come with them.
aetimmes··on Twitter Is DDOSing Itself
Anecdotally, you are wildly incorrect.
aetimmes··on My ranking of every Shakespeare play
What an amazing series. My favorite bit has to be the Love's Labor's Lost bit at the top of ep. 3, followed by an excerpt from A Bit of Fry and Laurie that makes a joke at the expense of folks like John Barton that he's able to laugh along with.

There's so much about the language of Shakespeare that needs to be vocalized and heard to understand. It took me way too long to realize that not all of his iambic pentameter was "correct' rhythmically, and that those variances _meant_ something. Watching and hearing these actors work through their scenes (and sometimes having their work adjusted live!) was truly eye-opening as to how the language of Shakespeare, arcane though it is, was deliberate, purposeful and useful.

aetimmes··on My ranking of every Shakespeare play
But what is most important is that he was correct about vanishingly few of them.
aetimmes··on Everything happening on Bluesky
Social media is a service.
aetimmes··on Twitter's Recommendation Algorithm
With forced RTO and "hardcore" mandates, it's difficult to find 6+ hours of time to interview at other companies (assuming there are any open visa sponsorships available).
aetimmes··on The Age of Advertising Must Come to an End
Anecdotally, I know several people (outside the tech field) who have told me that they get value specifically out of product-based FB/Instagram targeted ads, to the point that they don't _want_ to install ad-blockers for those sites. I don't think the experience of the HN crowd resembles the experience of the average user in any way.
aetimmes··on An update on two-factor authentication using SMS on Twitter
It's very straightforward. If you have money, you're allowed to be stupid. This is clearly the policy of Twitter 2.0.
aetimmes··on Ask HN: Why is everybody copying layoffs?
It's amazing how many wrong answers there are in this thread. This is the only correct one.
aetimmes··on Server BMCs can need to be rebooted every so often
Anecdote, but: I've seen a previous employer blackball a hardware vendor because of terrible BMC support.
aetimmes··on Production Twitter on one machine? 100Gbps NICs and NVMe are fast
I think there are several TCO issues you'd run into here:

- vendor lock-in: anyone who has worked at a shop running Sun SPARC machines when they got purchased by Oracle can speak to the pain involved with negotiating software licenses or hardware support contracts with the Only Game In Town.

- the price/scarcity of mainframe talent: you're going to have to pry IBM z-series experts away from banks who are paying 50-100% over market rate, oftentimes in straight cash, to maintain systems that are propping up the United States economy in its entirety. Not to mention - my first job out of college >10 years ago had a mainframe, and I was incredulous that _anyone_ still had or needed one in the 21st century. Now I can appreciate the specific tradeoffs being made that caused the business to choose a mainframe, but attracting top-tier cost-effective junior dev talent out of college becomes several orders of magnitude more difficult once the word "mainframe" leaves your recruiters' lips.

- scalability: in the event that you ever decide to add features or functionality (or, say, increase your tweet character limit by an order of magnitude), you have now committed yourself to scaling your systems in units of mainframes costing millions per unit, as opposed to servers costing five figures per unit (not to mention, you probably need a dev environment that's airgapped from your prod environment, which means yet _another_ mainframe...)

- build vs. buy: using the same commodity x86_64/ARM hardware and Linux kernel that everyone else is using allows you to take advantage of all of the open-source datacenter software being built for that happy-path profile. The minute you stray from that path, the engineering-hour cost of everything you do has the potential to skyrocket, because you can't use anything off-the-shelf and need to recompile everything for z/Architecture. In fact, based on some cursory web searches, it doesn't appear that you can compile the Rust toolchain to even _run_ on z/OS as of today, so at minimum, OP would be committing to implementing that.

But at the end of the day, the constraining resource in every software organization I've encountered has been engineering hours, and by choosing a mainframe you're drastically limiting the potential number of engineering hours available to you in the employee market.

aetimmes··on Production Twitter on one machine? 100Gbps NICs and NVMe are fast
It sounds like those engineers did good work.
aetimmes··on Twitter Layoffs Continue into 2023
The $4M/day figure was based on the aforementioned $1B/yr in interest payments.
aetimmes··on Production Twitter on one machine? 100Gbps NICs and NVMe are fast
It depends on what you (or OP) mean by "one machine".

There was a HPC cluster at Princeton when I worked there (which, looking at their website, has since been retired) that was assembled by SGI and outfitted with a customized Linux unikernel that presented itself as a single OS image, despite being comprised several disparate racks of individual 2-4u servers. You might be able to metaphorically duct-tape enough machines together with a similar technique to be able to run the author's pared-down scope within a single OS image.

With respect to the IBM z-series specifically - if the goal of the exercise is to save money on hardware costs, I'm imagining purchasing an IBM mainframe is in direct opposition to that goal. :) I'm not familiar enough with its capabilities to say one way or the other.

aetimmes··on Production Twitter on one machine? 100Gbps NICs and NVMe are fast
(Disclaimer: ex-Twitter SRE)

> There’s a bunch of other basic features of Twitter like user timelines, DMs, likes and replies to a tweet, which I’m not investigating because I’m guessing they won’t be the bottlenecks.

Each of these can, in fact, become their own bottlenecks. Likes in particular are tricky because they change the nature of the tweet struct (at least in the manner OP has implemented it) from WORM to write-many, read-many, and once you do that, locking (even with futexes or fast atomics) becomes the constraining performance factor. Even with atomic increment instructions and a multi-threaded process model, many concurrent requests for the same piece of mutable data will begin to resemble serial accesses - and while your threads are waiting for their turn to increment the like counter by 1, traffic is piling up behind them in your network queues, which causes your throughput to plummet and your latency to skyrocket.

OP also overly focuses on throughput in his benchmarks, IMO. I'd be interested to see the p50/p99 latency of the requests graphed against throughput - as you approach the throughput limit of an RPC system, average and tail latency begin to increase sharply. Clients are going to have timeout thresholds, and if you can't serve the vast majority of traffic in under that threshold consistently (while accounting for the traffic patterns of viral tweets I mentioned above) then you're going to create your own thundering herd - except you won't have other machines to offload the traffic to.

aetimmes··on Twitter Layoffs Continue into 2023
It's directly tied to the $1B/yr in additional interest payments being incurred by the debt used to finance the acquisition. Twitter pre-Elon was profitable. Twitter post-Elon is bleeding cash.
aetimmes··on The Twitter Files
> Infamously, Twitter did not even remove reported CSAM until after Musk purchased the company and fired the director of Trust and Safety

Source on this?

Anecdotally, the reaction of everyone I know working in law enforcement to the Musk purchase and subsequent layoffs was "this is going to make it infinitely more difficult to get CSAM taken down from Twitter".

aetimmes··on The Gervais Principle, or the Office According to “The Office” (2009)
I had been looking for this specific quote for months; thank you for surfacing it!
aetimmes··on We reduced 502 errors by caring about PID 1 in Kubernetes
"Services you didn't build yourself" include Kubernetes and Docker, and "system dependencies that you might not have touched before" include glibc and and the Linux kernel. The sentence you quoted is true of damn near every SRE.
aetimmes··on 100 People with rare cancers who attended same NJ high school demand answers
Both of my parents have had cancer, but neither were the rare types seemingly caused by CHS (melanoma and kidney cancer). Certainly worth keeping an eye on, though.
aetimmes··on 100 People with rare cancers who attended same NJ high school demand answers
Another story on the topic: https://www.nj.com/news/2022/04/a-mystery-in-colonia.html

Including an anecdote from finding a radioactive rock in the school in 1999:

"But when the teacher moved to an unremarkable, slate-gray, grapefruit-sized rock, the Geiger counter erupted like an alarm clock, sending the tiny gauge on the device to whiz to the highest reading levels, Gallo said."

(My hometown, but did not attend the high school, AMA)

aetimmes··on Celery in production: Three more years of fixing bugs
Threading is local to a machine; task queues are generally a solution to distributing a large number of tasks across many distinct compute nodes by using some remote network service (redis, rmq, etc) to track the progress of jobs.
aetimmes··on The Container Throttling Problem
Because then you have a snowflake service with a non-standard environment and still haven't solved the problem for all the other services that are still on Mesos.
aetimmes··on Fastly Outage
> People must be held accountable to have good incentives to reduce such outtages in the future.

Holding specific people "accountable" for outages doesn't incentivize reducing outages; it incentivizes not getting caught for having caused the outage.

As a result, post-mortems turn into finger-pointing games instead of finding and resolving the root cause of the issue, which costs the company more money in the long run when a political scapegoat is found but the actual bug in the code is not.

aetimmes··on Bare-Metal Kubernetes with K3s
Hmm - what's the overlap between your definition of "bare metal" and the current definition of "embedded"?

I will say, this comment section is the first time I'm hearing about "bare-metal" meaning "without an OS", but the above question is genuine curiosity.

aetimmes··on Bare-Metal Kubernetes with K3s
I recently set up a homelab of ~10 k3os+k3s nodes on NUCs. Setting up MetalLB on top of the base k3s installation made exposing services on their own IP addresses pretty dead simple.
aetimmes··on Kentucky wants to make it a felony to post law enforcement names online
IANAL, but this seems like a pretty clear violation of the First Amendment. Even if this bill passes (which, ridiculous bills get proposed all the time and never go anywhere), I can't imagine it standing up to challenges in the courts.
aetimmes··on Rocky Linux: A CentOS replacement by the CentOS founder
Assuming based on GP that this is in a HPC environment, there is often a delineation between the people writing HPC software and the people maintaining the clusters and the software installed on them. Telling a brand-new graduate student with zero software development experience to just throw everything into a container results in running code that is not optimized for the hardware it's running on, which in turn negatively impacts the other users competing for compute time on HPC clusters.

There is a movement to incorporate technologies like Singularity into the HPC workflow but for established projects, it often looks like a lot of bikeshedding for negative results compared to just running the code on bare metal.

← PreviousPage 2 of 5Next →