HNHacker News
TopNewBestAskShowJobs

jleahy

924 karma · joined October 20, 2013

submissionscomments
jleahy··on The Latest Linux File-System: TernFS
Yes hard links aren't supported in TernFS. They would actually be really difficult to make work in this kind of sharded metadata design as they would need to be reference counted and all the operations would need to go via the CDC. It wouldn't really have matched with the design philosphy of simple and predictable performance.
jleahy··on The Latest Linux File-System: TernFS
(disclaimer: CTO of XTX)

It was a long long time ago that we were only using NFS, it ran on top of a Solaris machine running ZFS. It did its job at the very beginning, but you don't build up hundreds of petabytes of data on an NFS server.

We did try various solutions in between NFS and developing TernFS, both open source and properietary. However we didn't name these specifically in the blog post because there's little point in bad mouthing what didn't work out for us.

jleahy··on The Latest Linux File-System: TernFS
(disclaimer: CTO of XTX)

These limits aren't quite as strict as they first seem.

Our median file size is 2MB, which means 50% of our files are <2MB. Realistically if you've got an exabyte of data with an average file size of a few kilobytes then this is the wrong tool for the job (you need something more like a database), but otherwise it should be just fine. We actually have a nice little optimisation where very small files are stored inline in the metadata.

It works out of the box with "normal" tools like rsync, python, etc despite the immutability. The reality is that most things don't actually modify files, even text editors tend to save a new version and rename over the top. We had to update relatively little of our massive code base when switching over to this. For us that was a big win, moving to an S3-like interface would have required updating a lot of code.

Directory creation/deletion is "slow", currenly limited to about 10,000 operations per second. We don't current need to create more than 10,000 directories per second so we just haven't prioritised improving that. There is an issue open, #28, which would get this up to 100,000 per second. This is the sort of thing that, like access control, I would love to have had in an initial open source release, but we prioritised open sourcing what we have over getting it perfect.

jleahy··on TernFS – An exabyte scale, multi-region distributed filesystem
> So no RDMA?

We can saturate the network interfaces of our flash boxes with our very simple Go block server, because it uses sendfile under the hood. It would be easy to switch to RDMA (it’s just a transport layer change) but right now we didn’t need to. We’ve had to make some difficult prioritisation decisions here.

PRs welcome!

> Implementing distributed consensus correctly from scratch is very hard - why not use some battle-tested implementations?

We’re used to building things like this, trading systems are giant distributed systems with shared state operating at millions of updates per second. We also cheated, right now there is no automatic failover enabled. Failures are rare and we will only enable that post-Jepsen.

If we used somebody else’s implementation we would never be able to do the multi-master stuff that we need to equalise latency for non-primary regions.

> This is not true for NFSv3 and older, it tends to be stateless (no notion of open file).

Even NFSv3 needs a duplicate request cache because requests are not idempotent. Idempotency of all requests is hard to achieve but rewarding.

jleahy··on TernFS – An exabyte scale, multi-region distributed filesystem
2MB median to be fair, so half of our files are under 2MB.
jleahy··on TernFS – An exabyte scale, multi-region distributed filesystem
Trading using ML.
jleahy··on TernFS – An exabyte scale, multi-region distributed filesystem
could be a tectonic shift in the open source filesystem landscape?
jleahy··on TernFS – An exabyte scale, multi-region distributed filesystem
The seamless realtime intercontinental replication is a key feature for us, maybe the most important single feature, and AFAIK you can’t do that with Ceph (even if Ceph could scale to our original 10 exabyte target in one instance).
jleahy··on TernFS – An exabyte scale, multi-region distributed filesystem
It's just not optimised for tiny files. It absolutely would work with no problems at all, and you could definitely use it to store 100 billion 1kB files with zero problems (and that is 100 terabytes of data, probably on flash, so no joke). However you can't use it to store 1 exabyte of 1 kilobyte files (at least not yet).
jleahy··on The Telechron master station clock was used to maintain power grid frequency
The point is that they need to be synchronised in order to respond to changes in load.
jleahy··on The Telechron master station clock was used to maintain power grid frequency
That’s not correct. The frequency provides an indication of how much power is being drawn. If the frequency is low then generators supply more power to bring it back to 60 Hz.

If there was not a well known fixed frequency it would be impossible to evenly distribute load over power stations. All generators have a %load vs frequency delta curve built into them which is precisely calibrated.

jleahy··on Georgia needs to produce more electric power for data centers
ICE is the name of the derivatives exchange, which is in Chicago. The cash market is called NYSE.

The parent post said ‘ICE AND NYSE’.

jleahy··on Georgia needs to produce more electric power for data centers
ICE is actually in Chicago.
jleahy··on Ask HN: How to get into quantitative trading?
In my experience a large % of the daily trade volume comes from me.
jleahy··on Ask HN: How to get into quantitative trading?
The spread between SPX and ES is purely mechanical. It’s a function of expected future interest rates and dividends over the remaining life of the future.

There is no such thing as support/resistance in reality.

jleahy··on Ask HN: How to get into quantitative trading?
ES trumps the others.
jleahy··on Do I need to get out the soldering iron again? (2018)
The Topping MX3 is also pretty good in my opinion and about half the price. That’s probably next down on the pareto frontier.
jleahy··on Quint: A specification language based on the temporal logic of actions (TLA)
At very least it looks like it can do what TLA can do whilst being dramatically less painful to learn / work with. That is very much enough to be interesting.
jleahy··on Rest in Peace, Optane
Even hitting RAM takes the best part of 100ns. They probably mean 5-10us given the ‘6x faster’ thing.
jleahy··on PCIe vs. CXL for Memory and Storage
Well it says "Limited adoption so far; CXL expected to accelerate through 2022 and 2023" (which didn't happen), so presumably 2021.
jleahy··on What happens when you shift a register by more than the register size?
Well it does work in hardware, but not with the normal shift instruction. Just use a ‘double shift’ which you can get with __int128 (or with inline asm, or just full asm).
jleahy··on Coincidentally-identical waypoint names foxed UK air traffic control system
Yes, it happens to people every day for all kinds of boring reasons. Normally because they made a mistake in their flight plan.

Remember anyone flying IFR needs a flight plan, lots of private pilots etc. Plenty of people make mistakes.

jleahy··on Coincidentally-identical waypoint names foxed UK air traffic control system
Well they can reject the bad flight plan and the plane will not be allowed to follow it.
jleahy··on Requiring ink to scan a document–yet another insult from the printer industry
I used to have a Brother laser but now have a Kyocera color laser MFP, it works beautifully so that’s another good option.

Also a networked printer is a lot less faff if you have a lot of Linux machines at home, from a drivers perspective.

jleahy··on Every time you click this link, it will send you to a random Web 1.0 website
The ‘how do I exit vim’ problem.
jleahy··on Ada Outperforms Assembly: A Case Study
There are still jobs where you end up writing large amounts of asm regularly and this is absolutely table-stakes for that. Otherwise there would be no point using asm at all, you’d just use C or another high-level language.

There was once a case where someone dropped from asm to machine code to shave a little more performance off. Sometimes asm is too high-level.

jleahy··on All IP addresses are equal? “Dot-zero” addresses are less equal (2013)
ah, onlink, indeed that is another way.
jleahy··on All IP addresses are equal? “Dot-zero” addresses are less equal (2013)
It’s very nasty, but there are a lot of options. For example you can give the gateway an address in the link local range and add a route pointing at that. I think that’s the best way.

Another option would be to set the subnet mask to /0 and enable ARP proxy on the gateway (that is truly diabolical).

Another way is to have a private /30 or /31 as the linknet and then add the /32 public ip as an additional one with a /0 route to the routers ip in the private /30 (and the router can have a /32 route to your ip in the private /30).

1:1 NAT is another option (but that’s not quite the question).

jleahy··on In New Paradox, Black Holes Appear to Evade Heat Death
Yes, and you can calculate exactly how much time it takes (it was an exam question in my final exams for undergrad iirc). The amount of time depends on the mass of the black hole but it's definitely less than a minute.
jleahy··on Google Launches AI Supercomputer Powered by Nvidia H100 GPUs
250
Page 1 of 12Next →