HNHacker News
TopNewBestAskShowJobs

ttfvjktesd

49 karma · joined September 18, 2025

submissionscomments
ttfvjktesd··on Rubygems.org AWS Root Access Event – September 2025
> failed to rotate the AWS root account credentials ... stored in a shared enterprise password manager

Unfortunately, many enterprises follow the poor practice of storing shared credentials in a shared password manager without rotating them when an employee with prior access leaves the company.

ttfvjktesd··on Building the heap: racking 30 petabytes of hard drives for pretraining
How is it going to work when the GPU is in the cloud and the storage is miles away in a local colo in SF down the street? I was under the impression that the GPUs has to go multiple times over the training dataset, which means transfer 30 PB multiple times in and out of the clouds. Is the data link even fast enough? How much are you charged for data transfer fees.
ttfvjktesd··on Building the heap: racking 30 petabytes of hard drives for pretraining
> do you think managing IAM and Terraform is free?

No, but I would argue that a SaaS offering, where the whole maintenance of the storage system is maintained for you actually requires less maintenance hours than hosting 30 PB in a colo.

In terraform you define the S3 bucket and run terraform apply. Afterwards the company's credit card is the limit. Setting up and operating 30 PB yourself is an entirely different story.

ttfvjktesd··on Building the heap: racking 30 petabytes of hard drives for pretraining
How about all the other infrastructure. Since you are obviously not using the cloud, you must have massive amounts of GPUs and operating systems. All of that has been working together, it's not just keep watching for the physical disks and all is set.

Don't get me wrong, I buy the actual numbers regarding hardware costs, but in addition to that presenting the rest as basically a one man show in terms of maintenance hours is the point where I'm very sceptical.

ttfvjktesd··on Building the heap: racking 30 petabytes of hard drives for pretraining
You are under the assumption that only Ceph (and similar complex software) requires staff, whereas plain 30 PB can be operated basically just by rebooting from time to time.

I think that anyone with actual experience of operating thousands of physical disks in datacenters would challenge this assumption.

ttfvjktesd··on Building the heap: racking 30 petabytes of hard drives for pretraining
The biggest part that is always missing in such comparisons is the employee salaries. In the calculation they give $354k/year of total cost per year. But now add the cost of staff in SF to operate that thing.
ttfvjktesd··on Typst: A Possible LaTeX Replacement
> even though I'm going to have to make a pixel-perfect clone of my university's LaTeX template

I'm not sure if you really mean pixel perfect or if it's just an exaggeration. There are packages in latex which are almost impossible to replicate in a pixel perfect way, one widely used example is microtype, which is especially useful in scientific works.

ttfvjktesd··on Titanic's sister, Britannic, sank in 1916. Divers have recovered artifacts
> Britannic has long been one of the major bucket-list dive that many deep wreck sport divers pursue.

There were less people non-commercial diving below 100m, than people reaching the top of the everest. If someone has Britannic on his list, then he's ether an extremely talented very serious technical diver with 500+ logged dives or it's just a pipe dream.

ttfvjktesd··on How AWS S3 serves 1 petabyte per second on top of slow HDDs
This piece is interesting background, but worth noting that the actual numbers are highly speculative. The NSA has never disclosed hard data on capacity, and most of what's out there is inference from blueprints, water/power usage, or second-hand claims. No verifiable figures exist.
ttfvjktesd··on Unlocking a Million Times More Data for AI
I think one important point is missing here: more data does not automatically lead to better LLMs. If you increase the amount of data tenfold, you might only achieve a slight improvement. We already see that simply adding more and more parameters for instance does not currently make models better. Instead, progress is coming from techniques like reasoning, grounding, post-training, and reinforcement learning, which are the main focus of improvement for state-of-the-art models in 2025.
ttfvjktesd··on Top Programming Languages 2025
Different compiler versions, target architectures, or optimization levels can generate substantially different assembly from the same high-level program. Determinism is thus very scoped, not absolute.

Also almost every software has know unknowns in terms of dependencies that gets permanently updated. No one can read all of its code. Hence, in real life if you compile on different systems (works on my machine) or again but after some time has passed (updates to compiler, os libs, packages) you will get a different checksum for your build with unchanged high level code that you have written. So in theory given perfect conditions you are right, but in practice it is not the case.

There are established benchmarks for code generation (such as HumanEval, MBPP, and CodeXGLUE). On these, LLMs demonstrate that given the same prompt, the vast majority of completions are consistent and pass unit tests. For many tasks, the same prompt will produce a passing solution over 99% of the time.

I would say yes there is a gap in determinism, but it's not as huge as one might think and it's getting closer as time progresses.

ttfvjktesd··on Top Programming Languages 2025
On the other hand, one can see it as another layer of abstraction. Most programmers are not aware of how the assembly code generated from their programming language actually plays out, so they rely on the high-level language as an abstraction of machine code.

Now we have an additional layer of abstraction, where we can instruct an LLM in natural language to write the high-level code for us.

natural language -> high level programming language -> assembly

I'm not arguing whether this is good or bad, but I can see the bigger picture here.

ttfvjktesd··on How AWS S3 serves 1 petabyte per second on top of slow HDDs
> tens of millions of disks

If we assume enterprise HDDs in the double digit TB range then one can estimate that the total S3 storage volume of AWS is in the triple digit Exabyte range. That's propably the biggest storage system on planet earth.

ttfvjktesd··on TernFS – An exabyte scale, multi-region distributed filesystem
Digital Ocean is also using Ceph[1]. I think these cloud providers could easily have 100s of PBs Clusters at their size, but it's not public information.

Even smaller company's (< 500 employees) in today's big data collection age often have more than 1 PB of total data in their enterprise pool. Hosters like Digital Ocean hosts thousands of these companies.

I do think that Ceph will hit performance issues at that size and going into the EB range will likely require code changes.

My best guess would be that Hetzner, Digital Ocean and similar, maintain their own internal fork of Ceph and have customizations that tightly addresses their particular needs.

[1]: https://www.digitalocean.com/blog/why-we-chose-ceph-to-build...

ttfvjktesd··on KDE is now my favorite desktop
> By the way, the crop and blur from that screenshot above ....

I just want to mention that blurring secret information is not secure. Use black bars instead.

ttfvjktesd··on TernFS – An exabyte scale, multi-region distributed filesystem
How does TernFS compare to CephFS and why not CephFS, since it is also tested for the multiple Petabyte range?