Vesuvius Challenge 2023 Grand Prize awarded: we can read the first scroll
scrollprize.org
scrollprize.org
Today I sit with the same amazement, taken aback again, appreciating how ridiculously awesome this is. Congratulations to the winners and everyone involved!
The fact I have a computer writing flowery alt text descriptions of my photos with unnerving accuracy is something I would not have predicted for another 20 years. But, here we are...
State-of-the-art machine learning architectures aren't actually that complex. Diffusion models and transformers can be explained to a bright high schooler. I'm sure Archimedes and Euclid would have no problem understanding them.
What they might have a problem understanding (or even imagining) is the mind-boggling amount of computation required to make those systems do anything useful. Getting Llama to produce a single token of text takes more calculations than all of humanity did by hand during all of Classical Antiquity.
Imagine all the stuff... transistors, Turing/Von Neumann machines, lithography, theoretical computer science, OS and compilers, the Internet... and lastly there's modern day machine learning that builds on top of all the above.
The base level stuff isn't exactly protons and electrons, but given the nanometer scale of our chips, it's not that far away from the truth, and we (humanity) has somehow built amazing stuff on top of that.
Smart "intellectual" people would certainly be willing to challenge basically everything they assume about nature, but I don't think your run of the mill farmer would be able to do that.
It would be like explaining a baking recipe just talking about wheat and flour and heat. The point is that from first principles, ML is a huge jump. From first principles, baking is not.
https://github.com/mlang/tracktales
The fact I have a computer generate spoken narration for my MPD playlist with descriptions of album art included just blows my mind. 2023 was indeed a fucking milestone.
My grandma was blind, and I just spent 6 months looking after a guy who was blind (but has now had surgery and beat me in the eye test at the doctor's!), so I think about blindness a lot when I design.
And this isn't only about the monetary value itself, but also the fact that a large cash prize attached to a challenge boosts the prestige of finding a solution. Nobel Prizes come with about a million bucks on top of them, after all.
I'm quite confident that if someone offered $100 million for deciphering the Voynich manuscript or Linear A, we'd have a solution within 3 years.
Make it $8m or $12m and FAANG employees can actually start to justify working on it seriously from a money perspective
And yes, $1 million is very substantial for an individual. And the cool thing about offering it as a prize (from the point of view of the organizers, that is) is they only have to pay one person or team, although potentially thousands ultimately contribute to the solution, directly or indirectly.
But it's not clear what current tech could help with. Machine learning can't be applied to something you don't have training data for.
```
xxxxxxxxxx <- The surface of the scroll
xxxxxxxxxx
...
xxxxxxxxxx <- The bottom of the scroll
```
So, by tiling the image on the surface you get data that is size_x * size_y * n_layers. So, it can be seen as a video stream with size_x * size_y * 1 channel * n_layers where the layers replace the temporal dimension.
It's a cut through the scroll, with the time dimension in this video representing the location of the cut along the scroll lengthwise.
As you can see from the mess it's far from trivial to find the surface of any of the sheets in the scroll, often they layers are blended together messes.
You may be thinking of the scans they used of an unwrapped sheet, those were as you describe and were used to help figure out methods for the real challenge.
Absolutely insane the level of wizardry being applied here to turn a lump of blackened, charred scrolls into readable text.
Having only cursory knowledge with machine learning are some of the techniques used in the article only recently discovered or have they been around for a while?
Is it due to us having reached an inflection point with these types of algorithms that they have become more popular and thus we are seeing new ways to apply them to old problems?
Imagine what we'll be able to do to brains, dead or alive, in 100 years.
And in 10,000, maybe we'll be reconstructing the light cone. Maybe that's what we are right now. (Not serious, but it's a fun thought experiment.)
("No real humans harmed.")
But if the future can reverse the light cone, nobody is immune to that fate.
Who knows what the future holds. These are just sci-fi flights of fancy.
(I agree with you).
Basically reverse time.
Not to make a low effort comment, but this.
But that'll depend on things going smoothly and non-evil, non-dictators winning in the end. It'd be horrific if evil and malicious entities won and decided they just wanted to fuck with everyone.
You raise your parents from the dead, they raise their parents, and so on and so forth.
Doesn't apply to raising random people from a thousand years ago, but it's a reason people today might be resurrected in a couple of centuries.
It does put into perspective though, how distracted and selfish our species can be; oh, let's fight each other because of skin colour, sexuality, ethnicity, religion, fighting wars over resources, etc. Meanwhile people are getting cancer, have disabilities like blindness and paraplegia and generally just...dying, especially when it comes early, after a hard life. It's just so sad and disappointing that we have the resources to give everyone a pretty decent life while we work on solving these bigger problems...but we just don't.
It looks like there are two main bottlenecks to reading more: the need for manual intervention in segmenting the scanned scrolls, and the cost in scanning new scrolls.
The 2019 scans were done at 7.91 micrometer resolution with 88KeV monochromatic x-rays. Synchrotrons like DLS can also use their uniquely coherent x-rays to do diffraction imaging but I see no evidence they've tried that with Vesuvius scrolls. (Because the object is too thick?)
That scan is 5.5 terabytes of .tif files. The 2023 scan apparently is at 3.24 micrometer resolution and four separate scans of 53/70/88/105KeV each, which should produce a vast, data processing-chokingly wad of data. If all x thousand scrolls need to be scanned at that detail then the Vesuvius Challenge people are going to be juggling a lot of hard drives.
As for segmentation: get some sort of collective solution going, like the Seti@Home did, but for people who are bored as hell, instead of them scrolling Reddit or Twitter all day. Maybe do it like a CAPTCHA so you get it done for free? I'd segment for a few hours a month if I had the ability to do so.
This is a cool project that has taken a community to build to this point, why not try and open and expand the collective of humans working to understand the scrolls? Get millions of people involved and you don't need to rely on technological crutches and development, though that is not the worst way to go either.
If you know anyone that would help chip in for the Phase 2 of the project (scaling up, please let Nat know! (not directly affiliated with the project management team, just pointing to him as a great contact for that.... <3 :')))) ) )
I highly recommend spending a few hours wandering the site, it is an absolute wonder.
1: https://www.icloud.com/photos/#08dJAA5eM9jpbhlEa3fzkl5ng 2: https://www.icloud.com/photos/#076Pof4FziA7WgcI8hZrGZmzg
I hope to visit Herculaneum some day.
It looks like, from what we can gather, the author decides that should something be hard to get, that doesn’t lead to greater enjoyment. But, it seems that the archaeologists have found an awful lot of joy in how “rare” access to these scrolls is!
So much has been lost to well-meaning archaeologists who dug up and threw away things that they didn't think were important. They tried cleaning and preservation techniques on artifacts without testing, sometimes ruining them in the process. They ripped things out of context, and "restored" them based on guesses that were sometimes flagrantly wrong.
Of course they couldn't be expected to know everything that would come in the future, so blame can sometimes perhaps be muted. But it's especially positive that they extricated these particular objects very carefully and just waited for a way to extract information that they could hardly have hoped for.
Rainbows End by Vernor Vinge (ChatGPT helped with the search)
> The book you're describing sounds like "Rainbows End" by Vernor Vinge. In this near-future sci-fi novel, set in 2025, one of the subplots involves a project called the "Library Project," where the UCSD (University of California, San Diego) library decides to digitize its entire collection. The process is somewhat as you described: books are destructively scanned by being shredded into tiny pieces, which are then scanned and digitized, with the text being reconstructed from the scans. This process is a part of the broader themes of the book, which include the effects of technology on society and the concept of "wearable computing" and augmented reality. Vernor Vinge, a retired San Diego State University professor of mathematics, computer scientist, and Hugo Award-winning author, is well-known for his works in the science fiction genre, especially for exploring the concept of the technological singularity.
Similarly, a friend recently read Ghost Fleet and I decided to pick it up and read it. The first chapter seemed familiar, and every once in a while there were "scenes" that I absolutely remembered having read years prior, but I had no memory of the overall plot.
Thanks for reminding me about Rainbows End.
That's not a sci-fi novel, that's OpenAI's business model!
Instead of relying upon machinery, some zillionaire has their body dry frozen and stashed in a lunar south pole crater, with a foundation funding interstellar propulsion research to move the body to the coldest stable points discovered along the way towards the Boomerang Nebula (1° Kelvin) and research to revive back from burnt-crisp state.
The foundation incites all sorts of advancements along the way like working out practical fusion and ever more exotic energy generation, AGI, gravity manipulation, Drexlerian nanotech, Dyson swarm, star wisps, self-modifying bodies and so on, in its quixotic quest to fulfill its mandate.
https://www.smithsonianmag.com/smart-news/archaeologists-reb...
He of the terracotta army. Not excavated yet for fear of damage, but I would so love to know...
One of the suppositions is that the main chamber contained a model of his entire kingdom, replete with rivers of mercury.
So yes. Archaeology is a bit destructive, and sometimes the destruction can go both ways. Proceed with caution.
In the early days they wouldn't have accomplished anything by pushing forward, so it doesn't take all that much restraint.
I'm more impressed by people in, say, the 1990s or early 2000s. They might've had a shot but there was still too much risk, so they restrained themselves until it was a safer bet.
It is a bit of miracle that they were preserved, and not just discarded.
https://www.nationalgeographic.com/history/article/mummy-eat...
Amazing achievement, let's hope the Italian government allows for additional excavation of the villa.
But we have only read 5% of this scroll and there are a ton more already excavated, it will probably take years before we manage to process what we already have.
In the direction things are going ... maybe a few months :)
Maybe that's another AI application.
I found some conversation on the difficulty here https://www.reddit.com/r/AcademicBiblical/comments/16dj68i/h...
"ChatGPT, give me the highlights of these ancient Greek scrolls ..."
This sounds like the money IS a huge issue. How expensive can it be to buy out the locals? We're talking about priceless cultural artifacts
I wonder how far Italy will go once - if - they get rid of the mafia, it is like trying to drive a car with the handbrake on.
First word discovered in unopened Herculaneum scroll by CS student - https://news.ycombinator.com/item?id=37857417 - Oct 2023 (207 comments)
The Vesuvius Challenge - https://news.ycombinator.com/item?id=35322809 - March 2023 (32 comments)
Vesuvius Challenge - https://news.ycombinator.com/item?id=35169869 - March 2023 (32 comments)
Can AI Unlock the Secrets of the Ancient World? - https://news.ycombinator.com/item?id=39261465 - Feb 2024 (1 comment)
and this tweet which presumably covers the same ground as OP:
The $700k Vesuvius Challenge prize has been won - https://news.ycombinator.com/item?id=39261933 - Feb 2024 (2 comments)
But the main problem is probably not reading thrown out drives - it's that stuff is just too transient nowadays. People put something up on the net and decides to withdraw it next year. archive.org can't store the walled gardens, and even if a social network for example wants to archive stuff they might not be allowed to for legal issues (not that it has prevented them before, but anyway...)
2000 years later we scan the carbonised scrolls with (basically) magic rays and use thinking machines to reconstruct what Philodemus wrote.
I wish we could tell him. Sounds like he was a thinker, he would really appreciate it.
"You recovered... uh... everything?"
And of course, rm just unlinks, doesn’t actually delete, so even going a step further and recovering deleted content is hardly magic.
This is more like if, sometime in the future, they somehow successfully reconstructed a snapshot of our computers’ volatile memory by examining the power supply, or something ridiculous like that.
We’ll lose a lot of digital data simply because we won’t have the means to read it. CD-readers aren’t manufactured anymore in volume. It’s easy to imagine society in 40 years not having any CD readers handy but having a bunch of CDs they want to read. Now multiply that by all the funny storage formats we’ve created over the years.
On HDDs. On SSDs it'll lead to now-unusued space getting TRIMed which actually erases the blocks. Back to scraping the papyrus.
https://en.wikipedia.org/wiki/Erotic_art_in_Pompeii_and_Herc...
Priapus had it goin' on! Reading the Priapeia for the first time is a treat...
Maybe something along these lines?
This is the most exciting thing in the world to me right now, these scrolls, along with the thought that there might be literally thousands more still in the ground.
So many there's a lengthy list on Wikipedia about it. It's fascinating reading ancients casually referencing works that we otherwise know nothing else about. Without the careful, laborious copying (often imperfect) over the centuries most things would've been lost completely. There's also other works such as maps that did not survive, the Tabula Peutingeriana for example is thought to be a derivative work of one commissioned by Augustus of the known world at the time (to Romans) and of which there's a few mentions in some works by historians at the time.
The worse case would be that it was 800 copies of the same scroll waiting to be sold off to other libraries.
Everyone interested in this story should read Stephen Greenblatt's The Swerve (https://www.pulitzer.org/winners/stephen-greenblatt).
It traces the story of a Renaissance humanist who tracked down and translated the Epicurean philosopher/poet Lucretius' De Rerem Natura, which Greenblatt describes as portraying a strikingly modern way of seeing the world.
In particular Lucretius and the Epicureans denied the existence of supernatural causes, were opposed to religious fear, and posited the ideas of atomism and biological evolution. Of course they're better known for their approach to living life, which Greenblatt shows is more sophisticated than sometimes caricatured, and which he portrays as a breath of fresh air compared to the oppressive moralism and hypocrisy of the Church at the time. (Jefferson and many of the American Founders described themselves as Epicureans.)
He goes on to imply that Epicureanism was influential and widespread in the ancient world but suppressed by the early Church, so that we now know little of it.
Anyone, one of the tantalizing parts of the book is where he describes the carbonized and unreadable Herculaneum scrolls, since they were the private library of a wealthy patron of the Epicureans. I think he thinks being able to read the scrolls will really change our understanding of the ancient world.
And remember: if they hadn't been carbonized, they would have crumbled to dust. That's why we only have the texts that managed to get copied. (Anthony Doerr's Cloud Cuckoo Land is a novel about the survival and 21st century rediscovery of an imaginary Greek play, and ... I'll let you read it yourself - https://www.anthonydoerr.com/books/cloud-cuckoo-land)
(Apologies for any errors above, as basically all I know about this subject is what I read in the book!)
It's interesting reading for a layperson, but as with any other pop-history book, one should read this with a heaping plate of salt at hand. (I'm... not sure what that metaphor actually means or if this is an appropriate way to extend it.)
Things are always more nuanced than can be laid out in a sweeping narrative format and the compression required can lose some critical information, even with the best of intentions. There's also just getting things wrong, which most non-historians do and many historians will do on topics that aren't their expertise.
I'd read this criticism from AskHistorians (not infallible, I know)
https://old.reddit.com/r/AskHistorians/comments/ejfxe5/comme...
So the extended metaphor makes no literal sense according to the Pliny text, but it makes sense according to our interpretation of it, which is what matters.
In all seriousness though, I find this such an amazing project to follow regardless of the outcome(s)
The developers of the transformer, as a group, should win some sort of significant prize; it has had more impact in a short time than anything I've seen before. Will we find better architectures in the near future?
Utterly brilliant. I'm so glad it appears to be a bit of scholarly writing too, which is what I know most classicists secretly love!
I wonder what this means for the maya codices, many of which are in similar shape: https://en.wikipedia.org/wiki/Maya_codices#Other_Maya_codice...
This reminded me, since they're scarce but also abundant... Has anyone actually eaten these giant waterbugs at Nue in Seattle? Is that like, a reasonable thing to subject a date to?
Big ups for the winners - this is so cool and hopefully can be replicated for deciphering many other lost manuscripts.
I think the biggest success isn't the recovered text but organizing such an endeavor with such success.
Totally inspiring!
It is wild to me, though, that if I have an SSD fail it's essentially unrecoverable, but a 2,000-year old, rolled-up, lava-burnt scroll of Papyrus can be read using Technology™! I love to see it!
https://en.wikipedia.org/wiki/Complaint_tablet_to_Ea-n%C4%81...
Most of what we have left from the ancient world is material that people felt worth copying for _centuries_ after they were written. That's a fairly amazing quality filter and gives people a skewed perspective on the overall quality of material written at the time.
A selection of work from a random library has also likely to be filtered for quality to an extent, and for the most part, the "good stuff" is going to be works that we _already_ have copies of or fragments from. Anything we don't already have a copy is is most likely going to be something that there weren't many copies made of, and usually there's a reason for that.
Which isn't to say that anything we find isn't going to be interesting for other reasons -- even bad writing is going to be incredibly useful for historical research.
The general subject of the text is pleasure, which, properly understood, is the highest good in Epicurean philosophy. In these two snippets from two consecutive columns of the scroll, the author is concerned with whether and how the availability of goods, such as food, can affect the pleasure which they provide. Do things that are available in lesser quantities afford more pleasure than those available in abundance? Our author thinks not: “as too in the case of food, we do not right away believe things that are scarce to be absolutely more pleasant than those which are abundant.” However, is it easier for us naturally to do without things that are plentiful? “Such questions will be considered frequently.”
oh good grief, even back then, we were limited to this value
Technical reproduction. The Vesuvius Challenge Technical Review Team reproduced the winning submissions manually. We made sure to clearly understand every part of the code, and that when we run it independently we get similar output images. Since all code and training data is now open source, you can do the same!
Multiple submissions of the same area. You might have noticed that all submission images above show the same area of the scroll. This is because we released 3d-mapped papyrus sheets within the CT-scan (“segments”) created by our segmentation team, which were then used by all contestants. The resulting output images — created by different ML models and training labels — have produced extremely similar results. This holds not just for the winners and runner ups, but also for the other submissions that we received.
Small input/output windows. The ink detection models are not based on Greek letters, optical character recognition (OCR), or language models. Instead, they independently detect tiny spots of ink in the CT scan, the writing appearing later when these are aggregated. As a result, the text appearing in the images is not the imagined output of a machine learning model, but is instead directly tied to the underlying data in the CT scan.But, without that process, you can have no comfort than what is occurring is valid. The software might reliably see a letter in some noise, but so? It doesn't mean the letter is actually there... One can't verify the scroll, and one hasn't verified the process.
One would always want to test stuff in software development, especially if it was fraught and can easily be tested.
Mostly there are no tests to be undertaken in history -hence it it's so much hearsay. But here is an opportunity to gain some genuine certainty, in a way that is normally unavailable! The implementors of this method should absolutely test their process!
However, you are right - if you go to Tutorials and Scanning there is reference to the creation of a 'campfire scroll'. And now we have some detail..... and the detail is problematic.
In tutorial 3, halfway down this page (https://scrollprize.org/tutorial3) there is a before and after comparison of the scroll. They ask a question "If you look back to the last page of the campfire scroll (before carbonization), can you see which area of the scroll this segment came from?".
This is mean to be obvious to answer - and one does have the 2 images to compare.
My thoughts on the comparison is - yes at a glance the there is a section that appears to match up - the 'angular squiggle' next to the @ symbol. However, if I look more closely at the 'angular squiggle' I see features and spaces in the generated image that do not correlate with the photo before carbonisation. The troughs are too deep, the spaces are too big. It seems to be a superficial similarity only.
I wish I could show what I mean by referencing the images.. But I will provide links for others to see what I mean when I say there is a superficial correlation only.
https://scrollprize.org/img/tutorials/vc-segment.png
https://scrollprize.org/img/tutorials/campfire-last-page.jpg
Final thought - why not map the 2 images one over the other on the site? Why ask a leading question rather than provide a proof? I hate that kind of presentation - it smacks of providing enough information for someone to make an incorrect snap judgement.
The before pic, needs to be flipped to compare to the after pic.
The before pic is not flat - it is taken on a curve. This makes it hard to use for comparisons.
The after pics are terrible quality.
As we are dealing with a long squiggle I can't describe it easily - people will be confused at what I am talking about.
However, once you get the images more or less in a comparable state, I can see things do no line up as you would expect. The first 'peak' points a different way. The lines do not line up. Stuff like that.