On the one hand, storing positions on JSON is quick to implement, easy to understand, easy to read, easy to hack on, and junior engineers and the people who have to deal with your code later will be able to pick it up and run with it easily.
On the other hand, when just about everyone is making this same ease-of-use/performance tradeoff, software bloat happens. Our computers are so much faster and beefier, but we never actually seem to be able to enjoy the benefits of that, in part because software engineers keep optimizing for quick and easy.
Maybe we shouldn't dogmatically reach for the easy no-nonsense solution every time, and instead consider whether maybe a little nonsense might, over time, save people a lot of time.
Most likely unrelated, I know! Just a wondering in passing.
I doubt you could fit that JSON file in its NVRAM. Sometimes I wonder how much sooner we could have had what smartphones offer us today if some technical choices had been made more... wisely.
But that's the endless conundrum Worse is better [1].
I've only seen 90s (nineties), never 90ies (ninety-ies). I am imagining the second one pronounced differently.
JSON is for storing data as text. Not work with that text all the time.
[edit] shorter version of the above: it stores the values, but doesn’t store what they mean.
There are different applications for different things: If you want to host a website with real-world tournament results involving only humans, you probably can get away with using more bytes. But if you're writing an engine that uses pre-computed positions, you want to be as compact as possible.
https://en.wikipedia.org/wiki/Endgame_tablebase#Computer_che...
I did laugh a bit at this bit because "conventional server" and "64 TB RAM" is hilarious to think about in 2023, but will probably be the base config in a Raspberry Pi in 2035 or so:
> In 2020, Ronald de Man estimated that 8-man tablebases would be economically feasible within 5–10 years, as just 2 PB of disk space would store them in Syzygy format, and they could be generated using existing code on a conventional server with 64 TB of RAM
Also I'd add the sizes involved here are kind of insane. I wrote a database system that was using a substantially better compression that averaged out to ~19 bytes per position IIRC. And I was still getting on the order of 15 gigabytes of data per million games. Ideally you want to support at least 10 million games for a modern chess database, and 150 gigabytes is already getting kind of insane - especially considering you probably want it on an SSD. But if that was JSON, you'd be looking at terrabytes of data, which is just completely unacceptable.
People like to use bots to run thousands or millions of simulated games to test how good their bot is at chess and have it ranked.
Other people like to use the bots that were created to play chess as practice toward a certain skill level. Beginners can pick bots that are proven to be beginner level, through thousands or millions of simulated games.
The smaller the data footprint for the games, the faster and more efficiently the bots can play, which reduces cost and time. In a more practical sense, AI/ML algorithms can be more efficient with tiny data sizes for a bunch of complicated reasons.
So, overall, this is "nerd sniping" to develop better chess players, both human and automated. It's not the most extensible presentation, I'll grant you, but I'm sure it's as fun as Regex Golf, or any other data-packing stuff.
P.S. I'm sure the comp-sci dev in you already knew all this; this is just a bill in case anyone read your comment and truly didn't already know all of this.
You'd write a wrapper to extract our the dense form to something with a nice interface.