HNHacker News
TopNewBestAskShowJobs

wickberg

93 karma · joined June 18, 2011

submissionscomments
wickberg··on UPS plane crashes near Louisville airport
Newer airports usually try to have space, that's the only thing helping with the physics involved here.

Older airports might have EMAS [1] retrofitted at the ends to help stop planes, but that's designed more for a landing plane not stopping quickly enough (like [2]) - not a plane trying to get airborne as in this case.

[1] https://en.wikipedia.org/wiki/Engineered_materials_arrestor_... [2] https://en.wikipedia.org/wiki/Southwest_Airlines_Flight_1248

wickberg··on Ham radio enthusiasts who help keep the NYC Marathon running smoothly
The US technician exam has almost zero relevance to operating a radio, outside of some legal reminders to broadcast your call sign every 10-minutes.

I'll second the recommendation for hamstudy.org. 90 minutes cramming that the night before the exam and I aced the test in 7 minutes.

wickberg··on A day in the life of the fastest supercomputer
That's 200Gbps from that card to any other point in the other 9,408 nodes in the system. Including file storage.

Within the node, bandwidth between the GPUs is considerably higher. There's an architecture diagram at <https://docs.olcf.ornl.gov/systems/frontier_user_guide.html> that helps show the topology.

wickberg··on A day in the life of the fastest supercomputer
Frontier is 4x 200Gbps links per node into the interconnect. The interconnect is designed for 540TB/s of bisection bandwidth. <https://icl.utk.edu/files/publications/2022/icl-utk-1570-202...>

Bisection bandwidth is the metric these systems will cite, and impacts how the largest simulations will behave. Inter-node bandwidth isn't a direct comparison, and can be higher at modest node counts as long as you're within a single switch. I haven't seen a network diagram for LambdaLabs, but it looks like they're building off 200Gbps Infiniband once you get outside of NVLink. So they'll have higher bandwidth within each NVLink island, but the performance will drop once you need to cross islands.

wickberg··on A day in the life of the fastest supercomputer
"Aurora" at Argonne National Labs is intended to be a bit bigger, but has suffered through a long series of delays. It's expected to surpass Frontier on the TOP500 list this fall once they some issues resolved. El Capitan at LLNL is also expected to be online soon, although I'm not sure if it'll be on the list this fall or next spring.

As others note, these systems are measured by running a specific benchmark - Linpack - and require the machine to be formally submitted. There are systems in China that are on a similar scale, but, for political reasons, have not formally submitted results. There are also always rumors around the scale of classified systems owned by various countries that are also not publicized.

Alongside that, the hyperscale cloud industry has added some wrinkles to how these are tracked and managed. Microsoft occupies the third position with "Eagle", which I believe is one of their newer datacenter deployments briefly repurposed to run Linpack. And they're rolling similar scale systems out on a frequent basis.

wickberg··on A day in the life of the fastest supercomputer
On a node-level, usually these are aiming for around 90-95% allocated. Note that, compared to most "cloud" applications, that usually involves a number of tricks at the system scheduling level to achieve.

At some point, in order to concurrently allocate a 1000-node job, all 1000 nodes will need to be briefly unoccupied ahead of that, and that can introduce some unavoidable gaps in system usage. Tuning in the "backfill" scheduling part of the workload manager can help reduce that, and a healthy mix of smaller single-node short-duration work alongside bigger multi-day multi-thousand-node jobs helps keep the machine busy.

wickberg··on A day in the life of the fastest supercomputer
Almost no modern systems are running Torus these days - at least not at the node level. The backbone links are still occasionally designed that way, although Dragonfly+ or similar is much more common and maps better onto modern switch silicon.

You're spot on that the bandwidth available in these machines hugely outstrips that in common cloud cluster rack-scale designs. Although full bisection bandwidth hasn't been a design goal for larger systems for a number of years.

wickberg··on A day in the life of the fastest supercomputer
Frontier runs unclassified workloads. Other Department of Energy systems, such as the upcoming "El Capitan" at LLNL (a sibling to Frontier, procured under the same contract) are used for classified work.
wickberg··on Backdoor in upstream xz/liblzma leading to SSH server compromise
I know Golang has their own implementation of sd_notify().

For Slurm, I looked at what a PITA pulling libsystemd into our autoconf tooling would be, stumbled on the Golang implementation, and realized it's trivial to implement directly.

wickberg··on The art of high performance computing
Cluster authentication actually remains the sore point for all interactions - for most systems the least-common denominator remains whether one can SSH into a login node. At which point the main job commands - sbatch/squeue - are pretty stable.

There have been some attempts to standardize basic job management APIs in the past - DRMAA being one noteworthy example. Although DRMAA v2 was only ever implemented by Grid Engine, and is effectively an lightly-abstracted version of their internal APIs, that has never really seen first-class adoption by Slurm/PBS/LSF.

For Slurm, the REST API is meant to be the way forward. It punts the authentication problem to, potentially, anything the admins may care to wire up through an Apache / NGINX proxy. And the basic job submission and status APIs have stablizied to the point that a client application should be able to consume nearly any version going forward.

wickberg··on Lawrence Livermore National Lab's powerful new supercomputer
Slurm is endian and word-size agnostic; there's nothing about the POWER9 platform that would be insurmountable for it to run on. There are occasionally some teething issues with NUMA layout and other Linux kernel differences on newer platforms, but this tends to affect everyone, and get resolved quickly.

My understanding is that Spectrum (formerly Platform) LSF was included as part of their proposal.

wickberg··on Intel Begins EOL Plan for Xeon Phi 7200-Series ‘Knights Landing’ Host Processors
In terms of KNL deployments, LANL's Trinity system (#9 on June 2018) is slightly larger. Trinity has 9,984 KNL nodes, vs 9,688 in Cori.
wickberg··on Designing Google Maps for Motorbikes
I use it on my motorcycle in the US, but there are some major pain points.

I have a Bluetooth headset in my helmet, so I can hear the directions... which are usually close enough. But you still do get some occasionally confusing voice prompts that - without the screen to glance at - may lead you to a wrong turn.

I do have a handlebar mount I've resorted to on a few occasions, but as others have noted (a) the phone can overheat, (b) subjecting a $500 device to the vibration and chance of falling is not ideal, and (c) I've had the acceleratorometer actions kick in unexpectedly, leading to the phone battery dying faster since it turned the LED flashlight mode on.

That said, I have only one serious complain - the opt-out-of-route-change that they've made the default is a huge issue, and I really wish there was a setting to disable that. If I've gone in and picked between a few potential routes, having it decide a different one is faster and that I'll need to press a button on screen to stay on my preferred heading ends up leaving me off course and/or on the side of the road digging my phone back out.

(And before someone suggests voice commands to fix it... that's not a workable solution. While the helmet cuts down noise, it doesn't eliminate it, and I have enough problems with voice recognition that I leave it permanently disabled.)

wickberg··on Accelerated Computing Powering World’s Fastest Supercomputer
I don't have any customer data on hand at the moment, but I'd roughly describe it as basic scheduling on strict priority order getting the system to 80+% usage, and then backfill boosting that up to 90%+. Careful backfill tuning, and care in defining the priority structure gets you to 95+%.

97% is the highest specific value I can recall for any of the larger sites with a heavily mixed workload, absent having a nearly infinite supply of short+small jobs at hand to use to fill those gaps. And a lot of users aren't trying to chase that - for these large scale "capability" systems the goal is to scale out as large as you can anyways, the smaller stuff is usually relegated to "capacity" systems elsewhere with a less expensive architecture.

One thing that at least some schedulers can manage is the idea of a min+max runtime for a job, combined with a min+max node/cpu count. If you have users willing to 'scavenge' otherwise wasted time by running under such a regime that can put you closer to full usage.

wickberg··on Accelerated Computing Powering World’s Fastest Supercomputer
If you're the one that came up with the idea of the "quiesce", I feel like you owe quite a few of us in the systems software community a beer or three...
wickberg··on Accelerated Computing Powering World’s Fastest Supercomputer
99% is a bit higher than most systems expect to run at - any mix of job sizes in the queue will tend to leave small gaps when nodes need to stay idle until jobs "fit" perfectly again. 95% is a pretty common target for these larger systems.

There isn't necessarily a lot of research into packing these better - the basic algorithms have been unchanged for quite some time, and a lot more effort goes into deciding how to prioritize different groups that are sharing access into the same system.

wickberg··on The Disappearing American Grad Student
Part time PhD's are rare - most professors need their grad students to be working on their grant-related research projects as RAs. That work is then used as the basis for your dissertation. As a part time PhD student you'd need to either have your own research planned out (and separate from the professor's grant-related work) which also implies you're likely to be paying out of pocket for tuition. Or find a professor who was willing to accept a slower pace of work on something that's presumably backed by a grant, which is pretty unlikely as the funding agency is usually holding them to strict deadlines.

Part-time Masters degrees are more common (and what I did myself), but come with similar trade-offs and usually need to be paid for by yourself or by your employer. And as you're not committing to being a part of a professor's research group for the next 5+ years, it's harder to find a good match for an advisor and a good thesis.

Co-terminal degrees (B.S. + M.S. all at once, with an extra year after undergrad to do your thesis) mitigate this, but you're going to be out of pocket for that extra year of tuition on top of the bachelor's.

wickberg··on Massachusetts is considering leaving the Eastern Time Zone
It's even more subtle than that - about half of Connecticut is heavily involved with business in NYC, and would probably be quite unwilling to jump an hour over.

To a lesser extent, south-western Vermont is economically linked to the Albany metro area, and would also likely vote against a shift.

Were you to poll on support for moving New England over to Atlantic time, I'd bet you'd wind up with a map closely matching that of local support for the Boston Red Sox vs New York Yankees: https://www.nytimes.com/interactive/2014/04/23/upshot/24-ups...

wickberg··on IBM Preps Power9 for AI and HPC Launch, Forges Big NUMA Iron
The Netflix Open Connect appliances are custom-built hardware they send out to ISPs. [1]

As evidenced by their heroic efforts to tunes performance[2], and their careful choice of hardware[3], they could certainly deploy POWER if they chose to.

[1] https://openconnect.netflix.com/en/

[2] https://medium.com/netflix-techblog/cdb51dda3b99

[3] https://openconnect.netflix.com/en/hardware/

wickberg··on Adafruit acquires RadioShack?
Indeed, it does look like the set she's holding is one of three that were sold in the auction. If Adafruit has actually acquired the brand, this is certainly an odd way to go about announcing it. I could see this having been a joke that's now gotten misinterpreted.

But, if so, the @adafruit twitter account is definitely further confusing things.

wickberg··on AMD ThreadRipper 1950X can compile the entire Linux kernel in 36 seconds
The Xeon... maybe.

The Phi... definitely slower. The single-core performance is very low, and there are usually at least a few linking stages in any source code compilation that collapse the process down to a single thread - you can't beat Amdahl's Law here.

wickberg··on Ask HN: Who is hiring? (December 2016)
SchedMD | Software Engineer | Lehi, Utah | ONSITE https://www.schedmd.com

SchedMD are the developers of the open-source Slurm resource manager (aka job scheduler) used by roughly half of the top500 systems.

We're looking for experienced Linux systems programmers and software support personnel to develop additional capabilities, build out a further CI/regression stack, and help our customers keep their clusters running at peak capacity.

Apply through the website (https://www.schedmd.com/careers.php), or email me directly at tim@schedmd.com .

wickberg··on OpenCAPI Unveiled: AMD, IBM, Google, Xilinx, Micron and Mellanox Join Forces
This is meant as a competitor, at least in the high-performance computing market, to Intel's Xeon Phi (aka Knights Landing) based systems which will start including their OmniPath network fabric on-die (based on QLogic's Infiniband tech).

If those gain widespread adoption, that leave little room for Mellanox's IB platform, hurts NVIDIA's sales of accelerator cards, and cuts AMD's and IBM's processors out of the picture entirely.

wickberg··on Photos from inside the Baikonur Cosmodrome
STS was never fully automated, they would have needed to install a temporary cable that allowed the landing gear to be automatically deployed.

It looks like https://en.wikipedia.org/wiki/STS-3xx has details on this cable, referred to as the Remote Control Orbiter (RCO) in-flight maintenance (IFM) cable.

wickberg··on Multiple vulnerabilities released in NTP
Mix in the just-announced CVE-2014-9322, among others, and you have a fairly obvious path to root.

Generally, at any time, it's safer to assume there's at least one active local root exploit in any system.

wickberg··on Google's POWER8 server motherboard
I wouldn't rely on the Watson connection to influence their chip division strategy.

It's not something they spread around, but from what I recall from a Q+A with IBM engineers the Watson prototype was developed on AMD xSeries servers, and only moved to pSeries machines late in the development before the Jeopardy! matches.

wickberg··on Fizz Buzz codegolf challenge in 15 languages
I was working on the same lines, and managed to bump it up to a score of 106:

  #define P printf(
  main(n){(n%3?0:P"Fizz"))+(n%5?0:P"Buzz"))?:P"%d",n);P"\n");n>99?:main(n+1);}
wickberg··on Inside the Titan Supercomputer: 299K AMD x86 Cores and 18.6K Nvidia GPU Cores
A few different reasons we keep them split:

- Jitter. Others have mentioned it, but I'll repeat it. These systems run carefully tuned microkernels to try to avoid any spurious interrupts, placing disk storage directly in the system would cause a lot of problems. Most applications require a careful lock-step progression to keep the calculations relevant; having one node take even 1% longer on a step wastes tremendous amounts of compute power for the system as a whole.

- System lifecycle management. As mentioned in the article, supercomputers are only cost effective to run for ~ 4 years; but data storage can outlast that. De-coupling storage and compute makes the transitions a bit easier; you can still get at the old filesystem even when the last generation machine has been scrapped. Also, this helps with access from related-but-disparate systems - you do need to get the data in and out of the machine to do anything useful with it, and potentially interrupting or degrading compute performance for external file access would be a problem.

- Power, and power stability - filesystems, especially large distributed systems such as Lustre/GPFS/PVFS2, do NOT handle power loss well. Best current practice for HPC centers is to keep the storage subsystems, file servers and disk arrays on backup power, but the compute side is directly run from the grid. Embedding disk in the compute would either require UPS'ing the compute platform, or anticipating filesystem corruption.

As with pretty much everything in the industry, people are looking at approaches to solve these. There is some inertia as you speculated, but it's becoming obvious that I/O is the new bottleneck as systems scale up any further, and you'll likely see some storage start moving in closer to the compute system.

wickberg··on Building a supercomputer from 64 Raspberry Pis and Lego
There are a few different companies that have ARM + custom interconnect systems out there or in development. They're not necessarily cost-competitive yet, but they're an interesting start.

Dell's "Project Copper" - http://content.dell.com/us/en/enterprise/d/campaigns/project...

Boston Viridis - http://www.boston.co.uk/solutions/viridis/default.aspx

wickberg··on Building a supercomputer from 64 Raspberry Pis and Lego
Yes, from a performance standpoint there's no reason to go after this type of system. No one is arguing this is a good system for production work.

The application of this is a teaching model. It's a lot easier to demonstrate parallelism gains on this type of platform. Scaling beyond a single ARM core is going to give you immediate performance benefits. Scaling further out to the entire cluster will continue to show returns.

With a single desktop, once you go beyond ~4 cores the gains will drop off too quickly. You just won't be able to see gains out to 64 threads on a single CPU, where on this you should.

It also doesn't hurt to have a quirky architecture to get students excited by. And yes, you could also spend some time discussing the architectural trade-offs and why this is not a cost-effective system for production use.

Page 1 of 2Next →