Using generative AI as part of historical research: three case studies
resobscura.substack.com
resobscura.substack.com
> had unusually legible handwriting, but even “easy” early modern paleography like this is still the sort of thing that requires days or weeks of training to get the hang of.
Why would you need weeks of training to use some OCR tool? No comparison to any used alternatives in the article. And only using "unusually legible" isn't that relevant for the… usual cases
> This is basically perfect,
I’ve counted at least 5 errors on the first line, how is this anywhere close to perfection???
Same with translation: first, is this an obscure text that has no existing translation to compare the accuracy to instead of relying on your own poor knowledge? Second, what about existing tools?
> which I hadn’t considered as being relevant to understanding a specific early modern map, but which, on reflection, actually are (the Peter Burke book on the Renaissance sense of the past).
How?
> Does this replace the actual reading required? Not at all.
With seemingly irrelevant books like the previous one, yes, it does, the poor student has a rather limited time budget
To your point about OCR: I think you'll find that the existing OCR tools will not know where to begin with the 18th century Mexican medical text in the second case study. If you can find one that is able to transcribe that lettering, please do let me know because it would be incredibly useful.
Speaking entirely for myself here, a pretty significant part of what professional historians do is to take a ton of photos of hard-to-read archival documents, then slowly puzzle them out after the fact - not by using any OCR tool (because none of them that I'm aware of are good enough to deal with difficult paleography) but the old fashioned way, by printing them out, finding individual letters or words that are readable, and then going from there. It's tedious work and it requires at least a few days of training to get the hang of.
If anyone wants to get a sense of what this paleography actually looks like, this is something I wrote about back in 2013 when I was in grad school - https://resobscura.blogspot.com/2013/07/why-does-s-look-like...
For those looking for a specific example of an intermediate-difficulty level manuscript in English, that post shows a manuscript of the John Donne poem "A Triple Fool" which gives a sense of a typical 17th century paleography challenge that GPT-4o is able to transcribe (and which, as far as I know, OCR tools can't handle - though please correct me if I'm wrong). The "Sea surgeon" manuscript below it is what I would consider advanced-intermediate and is around the point where GPT-4o, and probably most PhD students in history, gets completely lost.
re: basically perfect, the errors I see are entirely typos which don't change the meaning (descritto instead of descritta, and the like). So yes, not perfect, but not anything which would impact a historical researcher. In terms of existing tools for translation, the state of the art that I was aware of before LLMs is Google Translate, and I think anyone who tries both on the same text can see which works better there.
re: "irrelevant books," there's really no way to make an objective statement about what's relevant and what's not until you actually read something rather than an AI summary. For that reason, in my own work, this is very much about augmenting rather than replacing human labor. The main work begins after this sort of LLM-augmented research. It isn't replaced by it in any way.
My point about OCR is you haven't done any comparison and is now making the same mistake of claiming without any evidence. The most basic one from google translate does "know where to begin", it even doesn't make the "physical" mistake, though makes others. It also does know where to begin with the image from your second post, although it seems worse. And it's not the state of the art, and I don't know what that is for spanish either, but again, that wasn't my point. You do not have a care-free option, to be able to understand that "physical" mistake you'd still need to read the source, which means you still need those days/weeks of training
> none of them that I'm aware of are good enough to deal with difficult paleography
And you haven't demonstrated anything re. difficult paleography for the LLMs in your article either!
> entirely typos which don't change the meaning
First, you'd need to actually demonstrate that, and that would require the full accounting which you haven't done (and no, I don't plan to do that either) This could be a typo in a name or a year, which is bound to have some impact on a historical researcher? He'd try searching for a misspelled name and find nothing while there could've been an interesting connection in some other text?
>translation, the state of the art that I was aware of before LLMs is Google Translate, and I think anyone who tries both on the same text can see which works better there.
Yes, do try it, for example, in Deepl, to see that it's not any worse
> no way to make an objective statement about what's relevant and what's not until you actually read something rather than an AI summary
Sure, but presumably you've done that before making the claim of relevance "on reflection"? So how is it relevant to demand this "human labor" of the students?
Likewise with using Google translate on both case studies #1 and #2 - the results are self-evidently far worse. In both cases there were multiple errors in each line and in case study #2 it was entirely unable to transcribe or translate the title line. If you see this, please email me at bebreen [at] ucsc dot edu to share the better results you are seeing because I genuinely am interested and open to using alternative tools - I just am not seeing what you are seeing, apparently.
In terms of typos not changing the meaning, yes naturally a real human needs to double check absolutely everything if it's being used in research. We agree on that - the point is simply that this significantly speeds up the initial research process, not that it replaces the expertise necessary to, for instance, double check that a name or year is transcribed correctly. A huge amount of historical research is simply about skimming through documents looking for relevent info to zero in on - this is where LLMs can really help.
I do understand that a mere user of e.g. OCR tooling does not perform a systematic evaluation with the available tools, although it would be the scientific way to decide for one. For a researcher, however, the lack of knowledge about the tooling ecosystem seems concerning.
> Granted, Monte had unusually legible handwriting, but even “easy” early modern paleography like this is still the sort of thing that requires days or weeks of training to get the hang of.
He isn't talking about weeks of training to learn to use OCR software, he means weeks of training to learn to read that handwriting without any assistance from software at all.
Or, to get back to my original comment, if it's ok to be illiterate, why would you need weeks to learn using an alternative OCR tool?
Experts in the field might know more specialized tools, or how to train an actually better Transkribus model without deep technical knowledge required.
Most of the library is untranslated Latin. I have a book that was recently professionally translated but it has not yet been published. I’d like to benchmark LLMs against this work by having experts rate preference for human translation vs LLM, at a paragraph level.
I’m also interested in a workflow that can enable much more rapid LLM transcriptions and translations — whereby experts might only need to evaluate randomized pages to create a known error rate that can be improved over time. This can be contrasted to a perfect critical edition.
And, on this topic, just yesterday I tried and failed to find English translations of key works by Gustav Fechner, an early German psychologist. This isn’t obscure—he invented the median and created the field of “empirical aesthetics.” A quick translation of some of his work with Claude immediately revealed concept I was looking for. Luckily, I had a German around to validate the translation…
LLMs will have a huge impact on humanities scholarship; we need methods and evals.
The tendency to reaffirm popular beliefs would make current LLMs almost useless for actual historical work, which often involves sifting fact from fiction.
Instead, you should find the primary sources through other means and then paste them into the LLMs to help translate/evaluate/etc, which is what this author is doing.
Handwriting recognition, a classic neural network application, and surfacing information and ideas, however flawed, that one may not have had themselves.
This is really cool. This is AI augmenting human capabilities.
Then again, maybe not; OCR is one of the most worked on problems, so the quality of parsing characters into text maybe shouldn't be as surprising.
Off topic: it's wild to me that in 2025 sites like substack don't apply `prefers-color-scheme` logic to all their blogs.
I’m not a historian. I don’t speak old spanish. I am not a domain expert at all. I can’t do what the author of this post can do: expertly review the work of an LLM in his field.
My expertise is in software testing, and I can report that LLMs sometimes have reasonable testing ideas— but that doesn’t mean they are safe and effective when used for that purpose by an amateur.
Despite what the author writes, I cannot use an LLM to get good information about history.
I get a ton of value out of LLMs as a programmer partly because I have 20+ years of programming experience, so it's trivial for me to spot when they are doing "good" work as opposed to making dumb mistakes.
I can't credibly evaluate their higher level output in other disciplines at all.
The problem is, how do you know? I've seen developers go completely off-course just from bad search engine results and one did admit he felt something wasn't right but kept going because he didn't know better; now imagine he's being told by a very confident but incorrect LLM, and you can see how hazardous that'll be.
"You don't know what you don't know."
> There are, again, a couple errors here: it should be “explicación phisica” [physical explanation] not “poetic explanation” in the first line, for instance.
The image seems to say "phicica" (with a "c"), but that's not Spanish. "ph" is not even a thing in Spanish. "Physical" is "física", at least today, IDK about the 1700's. So, if you try to make sense of it in such a way that you assume a nonsense word is you misreading rather than the writer "miswriting", I can see why it assumes it might say "poética", even though that makes less sense semantically.
> I agree that my read may not be correct either
Just in case, by "you", I meant from the POV of the AI, not you the author.
That's interesting to know about "ph". I didn't know it was present in Latin, and I wonder if that's also the case with Spanish.
https://corpus.rae.es/cordenet.html
and it found 33 hits for "phisica" and 99 for "phisico", mostly from the 1490s. Now some of these can be deceptive, like a few are from a bilingual Spanish-Latin book and occur in the Latin portions rather than the Spanish portions, but it seems like some authors in the 1400s wrote "ph" in some Spanish words, at least when they knew the Latin or Greek etymologies.
I don't know when the Iberian languages first got their more phonetic orthographies, especially suppressing that h (that was originally in Latin digraphs used to transliterate Greek letters θ, φ, χ).
Edit: There are also about two dozen hits for physico/physica, interestingly more from the 1700s rather than 1400s.
You know, that might be analogous to Spanish speakers familiar with English writing "tweet" in Spanish text, while being ignorant that RAE added "tuit"[1] to the language, which is more in-line with general language rules. IDK if any Spanish speaker has ever written "tuit" in real life.
I love this line and the “flattening of human complexity into numbers” quote above it. It sums up perfectly how I feel about the whole LLM to AGI hype/debate (even though he’s talking about consciousness).
Everyone who develops a model has to jump through the benchmark hoop which we all use to measure progress but we don’t even have anything approaching a rigorous definition of intelligence. Researchers are chasing benchmarks but it doesn’t feel like we’re getting any closer to true intelligence, just flattening its expression into next token prediction (aka everything is a vector).
I agree a narrowing has happened. But the narrowing is to move us closer to saying "if it's not implemented in a brain, located inside a skull, in a body that was developed by DNA-coded cells replicating in a controlled manner over a period of years, it's not really AI."
There's an emotional attachment to intelligence being what makes us human that causes people to lose their minds when machines approach our intelligence. Machines aren't humans. If we value humanity, we should recognize that distinction—even as machines become intelligent and even sentient.
And we should definitely think twice, or, you know, many many many many more times, before building intelligent machines. But I don't think pretending we're not doing that right now is helpful.
I think what both viewpoints show is that, at the end of the day, intelligence is a broad, fuzzily defined thing, and attempting to claim that a single capability is evidence of intelligence always seems to be insufficient (from either direction).
I also think your points about our own emotional attachment and thinking carefully about intelligent machines are superb. I see a lot of people chasing certain tech right now and I see a far smaller number asking whether or not this tech is something we need or want. I personally don't need to live in a world in which robots are 1:1 emulations of humans (or better). I'd be just as content to live in a world of highly specific and highly optimized collections of robots or "intelligences" only capable of doing one thing really well (a unix theory of "agents", as it were)
https://zwischenzugs.com/2023/12/27/what-i-learned-using-pri...
This isn't a limitation, this is critically dangerous. Commercial AI is a centralized, controlled, biased LLM. At what point will someone train it to say something they want people to believe? How can it be trusted?
Consensus based information is still best, and I don't feel LLMs will give us that.
That's exactly not how that happened. That happened because Google's summaries are based on their search results and one of the search results contained that.
What I learned using private LLMs to write an undergraduate history essay - https://news.ycombinator.com/item?id=38813297 - Dec 2023 (81 comments)
That's an excellent way to put it. It's the default mode of an LLM. You can ask an LLM for biases, and get them, of course.
An easy way to make it not be true would be to emphasize some sources in pretraining by putting them in the corpus multiple times.
As far as I know, modern LLMs try to strike a balance between being somewhat neutral, while not being too neutral on topics outside of the overton window. They'll give you a "both sides have their good points" argument on abortion, religion, guns or immigration, but won't do that for obvious racism or nazi viewpoints.
Early LLMs had a problem with getting this balance right, I feel like many of them were a lot more left-leaning. I don't know how much of the change is caused by us understanding the technology better and how much is just the political winds shifting, though.
I felt like we had a moment there when some models were a bit too "well it depends", even on very uncontroversial subjects.
"Reality has a liberal bias"
[0]: https://epoch.ai/blog/will-we-run-out-of-data-limits-of-llm-...
E.g., the models won't report that unicorns are real because the majority of the internet doesn't report that unicorns are real. Of course, there may be issues (like ghosts?) where the majority of the internet isn't accurate?
But the gist of its argument just seems to be that they don't know fine details of history, and make the same generalized assumptions that humans would make with only a cursory knowledge of a particular topic. This seems unavoidable for a model that compresses a broad swath of human knowledge down to a couple hundred gigabytes.
Using AI as a research tool instead of a fact database is of course a whole different thing.
E.g. I have this recollection of a quote, slightly pithy, from around the 19 hundreds about hobby clubs controlling social life, maybe from Mark twain, maybe not.
I just cannot come up with the prompt that gets me the answer, instead I just get hallucination after hallucination, just confirming whatever I put in, like a student who didn't study for the test and is just going along with what the professor is asking at the oral exam.
Lorentz developed the Lorentz contraction independently of Einstein already, he was just hampered by the fact he adhered to the idea of the luminiferous ether as a medium for light propagation. I fully believe Hilbert+knowledge of tensors (Einstein didn't know the concept of tensors in 1905! [1]) would have developed general relativity as well.
[1] Einstein actually had an idea for general relativity much before he actually figured it out, he literally just lacked the mathematical knowledge to formalize it. He had to learn tensors in order to develop the Einstein Field Equations https://www.quora.com/How-did-Einstein-get-the-idea-that-he-...
https://www.metafilter.com/201537/O-brave-new-world-that-has...
> "In 1945, the French public said the Soviets did the most to defeat Nazi Germany - but in 2024 they're most likely to say it was the Americans"[0]
[0] https://yougov.co.uk/politics/articles/49613-d-day-anniversa...
Stalin thought he would’ve lost if it wasn’t for Lend Lease.
And this is extremely remarkable if you think about it. Germany basically declared war on the world, and very nearly won.
Eventually the fact that the war was happening in their land would grind them down. Plus, the US had nukes and aircraft carriers by the end, which would have presented a challenging situation.
So what would have happened in this scenario is difficult to even imagine, because Germany would have been under far less pressure. They were already working on the development of 'Projekt Amerika' [3]. It went nowhere, but without the pressures of the Red Army they would have had vastly more resources to expend on such ventures.
[1] - https://en.wikipedia.org/wiki/Red_Army
[2] - https://www.nationalww2museum.org/students-teachers/student-...
Getting a plane across the Atlantic with the limited bomb and fuel load that entails might not have accomplished much. If only a couple bombs could have done the job, I guess allies wouldn’t have been building all those bomber fleets, right?
And they didn’t build a serious surface navy before the war, or in the first year-or-so when they were still trading partners with the Soviet Union. It seems there was something beyond manpower pressure holding them back.
Now I think your argument is basically the same, but for America. In that the Nazis were in no position to invade and occupy America - which I fully agree with. But victory in a war doesn't require you occupy every single enemy nation. The most common way wars end is in settlement. And I see no possible path where the Nazis would not have been able to achieve a favorable settlement.
Americans were in position to fund and aid it. Both Stalin and Khrushchev acknowledged that American material help was of great importance.
Sorry but I’ve read tons of AskHistorians answers about this. I recommend you go there and search this same question. The Soviets needed that 11%. My understanding is that a lot of it helped mechanise the Soviet Army, ie trucks were delivered in large quantities.
From Claude I got this:
“In Khrushchev's memoirs, he recalled Stalin saying in a private conversation that without American aid through Lend-Lease, the Soviet Union "would not have been able to cope because we lost so much of our industry." Khrushchev himself wrote that Lend-Lease was "of utmost importance" and that "we would have been in a difficult position without it." “
“ The significance of Lend-Lease goes far beyond the raw percentage of industrial output. Here's why it was so crucial:
1. Timing and Critical Shortages: - The aid arrived during the most critical period (1941-1942) when Soviet industry was being relocated east of the Urals - During this vulnerable period, American trucks, food, and materials helped keep the Soviet army mobile and fed - Without this bridge of support during the industrial relocation, the USSR would have faced severe shortages at its most vulnerable moment
2. Strategic Materials and Bottlenecks: - The Soviets received specific materials that were severe bottlenecks in their production: - Aviation fuel and high-octane gas - Aluminum for aircraft production - Radio equipment and communications gear - Special grades of steel and industrial equipment - These materials were critical multipliers that enabled Soviet production
3. Transportation and Logistics: - Nearly 450,000 trucks were provided, which revolutionized Soviet logistics - Before Lend-Lease, the Red Army relied heavily on horse transport - American Studebaker trucks allowed for rapid troop movements and superior logistics - This mobility was crucial for later Soviet offensive operations”
Too many people forget about that. They were allies at first.
As far as I can tell, the Americans and Brits took too much credit. Then the Soviets and Russians insisted on more credit -- arguably too much. Of late I'm hearing historians say "Yeah, the Germans overextended themselves at the start and likely would have lost even if Hitler hadn't betrayed Stalin". I'm sure that analysis too will change.
The US industrial contribution is easier to understand looking back. It makes a lot of sense to us nowadays, looking at it in a table (not to cheapen it, it was an astonishing amount of stuff that was produced).
It seems entirely possible that the Soviets gave up more for the victory, while the US contributed more to victory.
Granted, you’d need to spend about a year on this and for a lot of that time your graphics card (and possibly whole computer) would be unusable, but then if the results were compelling you’d get a cool 15 minutes of internet fame when you posted your results.
Reduce the dataset to “knowledge as of year 1880” - and it’s not certain you’d even be able to “interact” with the LLM in any meaningful way…
I know this is possible, but the further away I get from my core domains, the harder it is for me to use these tools in a way that doesn’t feel like too much blind faith (even if it works!)
Like if you have a friend who's very well-read and talkative but is also extremely confident and loves the sound of their own voice. You quickly learn to treat them as a source of probably-correct information, but only part of they way you learn any given topic.
I do this with LLMs all the time: I'm constantly asking them clarifying questions about things, but I always assume that they might be making mistakes or feeding me convincing sounding half-truths or even full hallucinations.
Being good at mixing together information from a variety of sources - of different levels of accuracy - is key to learning anything well.
Let's say you're trying to get a university degree, but having a professor who makes up 20% of what they say. Is that helping you "learn well"?
And 20% is way overstated, esp for a SOTA model when it comes to verifiable facts.
It's easy to laugh and say, well I'm smart enough to defeat this. I know the trick. I'll just mentally discount this information so that I'm not unduly influenced by it. But I suspect you are over-indexing on fields where you are legitimately an expert—where your expertise gives you a good defense against this effect. Your expertise works as a filter because you can quickly discard bad information. In contrast, in any area where you're not an expert, you have to first hold the information in your head before you can evaluate it. The longer you do that, the higher the risk you integrate whatever information you're given before you can evaluate it for truthfulness.
But this all assumes a high degree of motivation and effort. Like the opening to this article says, all empirical evidence clearly points in the direction of people simply not trying when they don't need to.
Personally, I solve the problem in my friend circle by avoiding overconfident people and cultivating friendships among people who have a good understanding of their own certainty and degree of expertise, and the humility to admit when they don't know something. I think we need the same with these AIs, though as far as I understand getting the AI to correctly estimate its own certainty is still an open problem.
I can't speak to everyone's experience - but whenever I'm having a conversation around relatively complex topics with MY friends - the deeper they dive, the more they're constantly referring back to their dive computer. They'll also try to make arguments that are principally anchored to the pegs that they're convinced will hold. I'm aware I'm mixing metaphors here but the point stands.
As far as "mixing information" - yes there are commonly known tricks to trying to get a more accurate answer:
- Query several LLMs
- Query the same LLM multiple times with different context histories
- Socratically force it to re-assess itself
- Provide RAG / documents / access to search engines
- Force quantitative tests in the form of virtualized envs though this is more for Compsci/Tech/Math
etc.
LLMs don't currently have a good sense of their boundaries - they can't provide realistic confidence scores and weight their outputs accordingly - the human equivalent of saying, "I only have a passing familiarity with the original greek of the Septuagint, but I think...."
It's a poor use of an LLM as a glorified fact checker - it's far better as a tool for free form exploration.
> Being good at mixing together information from a variety of sources - of different levels of accuracy - is key to learning anything well.
I have a pretty extensive background in teaching/education and I would heavily disagree with this assertion - at least when starting as a complete novice. The key to learning well is to establish a strong foundation by learning from the most accurate resources as possible. When you pick up a musical instrument, you don't want a teacher who's just one page ahead of you in the lesson book.
I tend to ask multiple models and if they all give me roughly the same answer, then it's probably right.
You can see this effect in the ARC-AGI evals, too much context impacts even o3(high).
... or they had a lot of overlapping training data in that area.
https://beta.gitsense.com/?chat=ed907b02-4f03-477f-a5e4-ce9a...
If you click on the Evaluation links, you can see how you can use multiple LLMs to validate LLM response. The evaluation of the accurate response is interesting since Llama 3.3 was the most critical.
https://beta.gitsense.com/?chat=fdfb053d-f0e2-4346-bdfc-7305...
At this point, you would ask Llama to explain why the response was not 100% which you can use to cross reference other LLMs or to do your own research.
You can't. Because LLM's are statistical generative text algorithms, dependent upon their training data set and subsequent reinforcement. Think Bayesian statistics.
What you are asking for is "to reliably interrogate topics of interest", which is not what LLM's do. Concepts such as reliability are orthogonal to their purpose.
I have always felt that LLMs would fall apart beyond the summarization. Maybe they would be able to regurgitate someone else's analysis. The author seems to think there's some level of intelligent creativity at play
I'm hopeful that the author is right. That truly creative thinking may be beyond the abilities of LLMs and be decades away.
I think the author doesn't consider the implications of broad use of LLM societally. Will people be willing to fund human historian grad students when they can get a LLM for a fraction of the price? Will prospective historians have gained the training necessary if they've used an LLM through all of school?
I believe the education system could figure it out over time. I'm more worried that LLMs like this will be used as further justification to defund or halt humanities research. Who needs a history department when I can get 80% for the cost of a few chatGPT queries?
Good historians? Ehhhhhhh.
The problem is one of trust, and it's very difficult to trust the output of LLMs to be correct/true vs "truthy" without extensive verification that may be either as laborious as doing the original research or that may be difficult or impossible without knowledge and understanding of the internals and sources that may not be available.
A hobby of mine is editing Wikipedia articles about Australian motorsport (yes, I have an odd hobby, sue me).
The vehicles in the premier domestic auto racing category in Australia, the Supercars Championship, are unique to the category. Like NASCAR, they're built on a dedicated space frame chassis with body panels that look like either a Mustang or a Camaro draped over the top.
I'd seen occasional claims on forums that when the organising body was deciding on the design of the current generation of cars, they considered using the "Group GT3" rules that are used for a bunch of racing series around the world (including the German DTM championship, the GT World Challenge events raced across Europe, Asia, and Australia, and the IMSA GTD and GTD Pro categories). If true, it might be an interesting side note to the article about the Supercars Championship.
So I asked Copilot (the paid model) to find articles in motor sport media about this (there are a number of professional online publications that cover the series extensively). It confidently claimed that yes, indeed, there was some interest in using GT3 cars in the Supercars championship, and pointed me to three articles making this case.
The first was an article featuring quotes from the promoter of the DTM series saying what a good idea it was to have a common car across different national series. So the first article was relevant, but didn't actually show that anyone involved in the administration of the Supercars Championship was interested in the idea.
The second and third references were articles about drivers and teams whose core business is the Supercars championship also running cars in the local GT3 championship (while not explicitly mentioned in the article, they do this for a large wad of cash from the rich hobbyists who co-drive and fund most GT3 racing). Copilot's interpretation of the articles was just flat-out wrong.
Yes, this was a sample size of one historical query, but its response was very poor.
If those 5-10 results aren't great the LLM's response won't be great either.
Was this using o1? The author of the article was quite clear that his opinion doesn't apply to previous models.
Clearly, I’ll need to give o1 a try at some point to see if it does better.
The idea that there will be one model to rule them seems very unlikely.
While I welcome the rise of parallel shadow institutions as civilization grows spiritlessly utilitarian, the future for common sense looks bleak.
"Please don't post shallow dismissals, especially of other people's work. A good critical comment teaches us something."
"I’m told that OpenAI’s newish o1 model is genuinely helpful and creative when it comes to thinking through open problems in the sciences..."
"Likewise, although my knowledge of Italian is not great, I can read it well enough to confirm that the translation it offers is good enough to use for research:"
Using a translation for research that you couldn't perform yourself seems extraordinarily substandard for a historian.
The only part I agree with is a simple search to identify sources that may be relevant that you had not considered. i.e. Primary sources to be examined directly -- the way historians have done it for millennia.
I don't think history should be filtered through model hallucinations. It seems an invitation to mistakes.
The reference to a "medallion, or seat of humors, or badge of office" is obviously something being held. The historian specifically says it is not any of these, but a "urine flask." A historian obviously should not take an LLM model as factual.
Later, the author writes, "After all: when you get down to it, o1 talking about a panopticon and Foucault in the above snippet is very, very similar to what a first year history PhD student might produce."
This is exactly the point. A mediocre average of writing that a first year student would produce. Sure it could be used for "I hadn't considered that," but it surely should not be used for any factual interpretation.
Are you saying historians should only ever consider sources in languages they are personally fluent in?
Robert Nozick (in Examined Life) asked how we feel if we found out, say, Beethoven seriously composed music based on a secret formula, which is entire mechanical and required no effort for him at all.
Would we still appreciate the music in the same way? If not, does our appreciation really stem from the fact that we feel he has also struggled like we do, and nevertheless produced something incredible.
I remember as a very small child watching figure skaters on TV and thinking "that's no big deal". And before I started programming: "it's just logic, all very straightforward". But that was before I first entered an ice rink or centre-d a div
Maybe we don't really appreciate something unless we appreciate it is hard in a visceral way.
If he discovered the formula, then yes I imagine most people would appreciate the music just as much, if not more so.
If he copied the formula from somebody else, then he was just turning the crank - a far more sterile and mechanical affair.
Using Suno to "create" music is just turning the crank.
Related but much of Bach's music is just as incredible for its incredible mathematically relational structures as it is for its pure virtuosity and brilliance.
"Yet our experience of Beethoven's string quartets would be diminished if we discovered he had stumbled upon someone else's rules for musical composition, which he applied mechanically". (p. 38 of https://archive.org/details/examinedlife00robe/page/n15/mode...)
I guess another way to put the question is this. Suppose there is an alien civilisation where their brains are hard-wired to make Beethoven level music automatically. Most of us can hum a tune without effort: these aliens can hum music that would strike us as original and compelling without much effort. How would we react then?
Plus: nice pointer on the math scholar link. I remember loving the musical parts of Godel Escher Bach. Wish there is a good interactive website where I can revisit all the content (and listen to all the music) there in the browser.
Whatever music that follows, however hard fought for - even if 1:1 Bach output that he kept in a drawer, isn't beautiful? Just string plucking? That's not music appreciation, that's love of reputation you can easily grasp and associate with.
In Anarchy, State and Utopia, he tackles some utopian theorists' claim that, if equality prevails, everyone will rise up to the level of the greatest writers and artists. Would people be content then? Or will they still want to vie for "eyeballs"? If the latter, should we just admit that there is just a deep-seated human desire to compete for dominance?
For what it is worth, I've written up my reflections on skimming Anarchy, State and Utopia here: https://books-blog.3willows.xyz/posts/2024-10-26-anarchy-sta...
Count me out of that "we" -- I appreciate the artist who put in the work because they put in the work to make the thing I like, but I don't appreciate the thing because of the effort. I can marvel at the effort required to produce art in a certain way, but I'm (largely) indifferent to the effort in my actual appreciation of the thing (or lack of it).
I look forward to the time when I can have as much high-quality (to me) fiction to read as I like, because it's all generated by LLM. Some time after that, I'd love to see the main Star Wars sequence done properly. I won't care that it isn't created by a vast team of humans.
Since it is mirroring human culture, why do you see it in such a negative light? Instead see it like what it is, an interactive reconstruction, or maybe like a microscope to zoom into any idea.
It’s just in the context of poetry, and literary writing in general, that I feel differently about them. There’s also the fact that I haven’t read all that human writers and poets have already written (and will never be able to in this short life) so there’s no need to turn to synthetic output. No supply problem exists. Poetry in particular is something to ponder over and over. You can’t really run out.
You can't know what any poet felt or didn't feel while writing a poem. Perhaps it was a commission piece, or an experiment or an emulation of something the poet had heard elsewhere.
And more generally, whether the specific emotion another man feels is similar or even comparable to your own is also unknowable. He might use the same word to describe it, but the subjective experience associated with it might be completely different, and completely impossible to share.
Also poems are not really puzzles to be solved. If it produces an effect and is solid craft-wise, that is enough. There’s a lot to the craft side btw in the Urdu and Persian ghazal form which is what I had in mind while writing my original comment. LLMs can easily master the latter but have nothing to do with the former. Their output is pure form without substance.
Edit: I want to add that ambiguity (ابہام) is even a desirable property in Urdu ghazal, specifically. The more interpretations a couplet can have, the greater is the accomplishment in terms of craft.
Are your tastes so hyper-specific that we aren't already in this world? Fiction is (even pre-LLM) easier to find in whatever genre you want than ever.
I gave ChatGPT a list of my favorite SF novels, and a brief description of why, and asked for similar works. It recommended 10 novels, three of which I've read and weren't in the sweet spot. Also, everything it recommended was 30+ years old -- to be fair, the same is true of the list I gave it, but I think it goes against your point that there's an unlimited supply.
So I told it about the three and asked it to adjust and to give more recent works, and it obliged. One of the new recommendations was in the Culture series, which I've read one of and it wasn't my jam. Another was Project Hail Mary by Andy Weir, which I've read and enjoyed the Martian, but I'm betting that's the only Andy Weir I'll like. The others I'll have to check out.
It's an interesting exercise.
I think the problem here is analogous to the "500 channels and nothing to watch" issue in the heydey of cable.
Ok let's say you have an LLM in your hand that can generate any story you want. High quality. So you say: "tell me a story" and it tells you as story. But what story? Who is in it? What characters? Why are they there?
The only novelty that's going into this the prompt. Everything else is regurgitated weights and probability associations. The question is: does the full infinite closure of recombination over some finite learning set (no matter how large) encompass enough of the essence of creativity to produce something "new"?
This is a hard question to answer because it forces us to try to define creativity, or lacking that - at least try to identify where it comes from.
I don't have a clear answer to this but I'll suggest a line of thinking that seems plausible.
When a person writes a story, it's not derived as an amalgam of everything they have read. It's not some probabilistic weighted average of all those associations. The story they write is also derived from their lived experience. Their personal interactions, their observations, their musings, their passions, their fears.. and how all of those things interact with their circumstance, influencing their reactions, those reactions influencing their environment, and that feeding back into the above process.
There are two components that seem important here: the first is the existence of a rich, dynamic, and active _dialogue_ between the mind and its environment. It's not static, and it involves a feedback loop between the mind and the environment it models.
The second is a motive force. For humans the origin is biological. Fear, hunger, satiation, arousal, etc. - those core primitive emotional drives that originally developed to help us survive, but then were layered over with an intellect that elaborated on them. What originated as a motive force to drive the mating instinct evolves into a sonnet about an unattainable maiden. The fear of the dark that keeps us away from the places where we would be eaten.. evolves into a stories about unfathomable creatures and impossible colours that drive men insane.
And I think there's a third one that's unelaborated and implied but should be made explicit: introspection & reflection. The ability to consider your choices and consequences with respect to your motivations, and adjust any number of things - from the motivations themselves, to expectations/understanding.
This creature would have a lived experience, some underlying motivations, and a feedback loop established between the two using introspection. I have no idea how you'd build any of that.. but it feels like that's what you'd need before you got yourself a good storyteller.
But by that point, you'd also be compelled to question whether or not it's even ethical to force it to tell you a story anymore.
I don't think it's impossible that some broader AI system eventually is capable of genuinely creating creative output. LLMs are not that, though.
They seem more like a substrate.
Mountains get meaning as aspects of our environment, but try and name the your top 10 most aesthetically pleasing mountains. At least for me, I may appreciate a scenic view but I just don’t think of them in that kind of context.
But so does a user. Users don't prompt "draw a dog" but give 3 lines of intricate details and iterate a dozen times until it looks right. It's not like these models work all on their own.
Our brains are drawn to some things visually for instinctive reasons, and I don't need a big message when I'm decorating or wanting to please the eye.
Unless there are ghosts in the shell, MidJourney gives an aproximation of how a painting of a rose looks like. Its like an aggrregate function that averages a million artists. Its a weird concept.
I would say yes, this dash cam technique can be an artistic method. Reminds me of Jon Rafman's wonderful Nine Eyes project - he captures screenshots from Google Streetview, see https://9-eyes.com
Its a fools errand - it’s an infinitely small set of people who can accurately describe their reasoning - even fewer have consistent reasoning - fewer still have coherence between beliefs
The ones that do we call either monks or crazy
I’d argue people aren’t even coherent enough to know how or what to appreciate
The classic internet philosopher’s lament: everyone else is irrational, inconsistent, and incapable of coherent thought—except, of course, the enlightened commentator making the claim. The irony is that this kind of sweeping generalization is itself an incoherent mess, built on vague cynicism rather than any serious engagement with human reasoning. If you actually believe that consistency and coherence are so rare, what exactly do you think you’re demonstrating here? Because from where I’m sitting, it looks less like deep insight and more like self-important nihilism masquerading as wisdom.
But that wasn't the claim. @AndrewKemendo said "it’s an infinitely small set of people who can accurately describe their reasoning - even fewer have consistent reasoning - fewer still have coherence between beliefs" So he didn't say that everyone else is irrational, he said that very few can accurately describe their reasoning. And I think this is true. Very few take the time to introspect. Fewer still will do so to the point that they are consistent in their thinking. And fewer still will will analyze their values and beliefs and get them to square up with each other. Their is nothing controversial here. It's demonstratively true, all one has to do is listen to people carefully and probe them to motivate their reasoning every now and again.
> The irony is that this kind of sweeping generalization is itself an incoherent mess, built on vague cynicism rather than any serious engagement with human reasoning.
I reject that it's a "sweeping generalization" – I assert that most if not all people who spend enough time carefully introspecting and observing others necessarily must come to this conclusion. What about the claim is an "incoherent mess"? Clearly this is a personal peeve of yours because your response is emotional and doesn't refute the claim in any decent way.
> If you actually believe that consistency and coherence are so rare, what exactly do you think you’re demonstrating here?
That's a logical fallacy.
> Because from where I’m sitting, it looks less like deep insight and more like self-important nihilism masquerading as wisdom.
Twaddle.
Emotional twaddle.
No thanks. Not really interested in the views of a murderer.
In my mind, art has always been a technological endeavor. Language, writing, and grammar are all tools. Brushes, stroke technique, and paint composition are all tools. I heard a story about Tony Hawk pioneering some skate board move, being the first in the world to get it right. And then seeing some teenagers doing the same thing years later in a park.
Real artists learn what is possible and then develop tools to break those limits.
And to your point, knowing how much effort really goes into something often requires a bit of experience to really appreciate it.
fantasic mr fox, coraline, jack stauber's opal, etc. are also very beautiful
At this point Aardman is also doing a lot of non clay stop motion, but it's still the core of their work.
Appreciating a piece of music is purely on its own merits of its content not if it was easy to create or not.
The background, ethics, skill or even creative process of the people behind it have no bearing on whether the music itself is good and how we much we like it, even if Hitler wrote the 9th symphony it would be still be a just as good a masterpiece. To consider anything but the merits of the output is a slippery slope of what biases are acceptable and not, that inevitably ends up being racial or at least exclusionary.
Even it was not as difficult as you imagined it to be, he still was the first to find it, or even just the first to popularize it and that is all that matters.
Sure they may functionally have the same effect or enjoyment, but appreciation of a fine work goes deeper than its function.
(eg: imagine pretraining where some of the documents are prepended with a "this is a bad example" token.)
A historian works with (and may even seek out in musty rooms) primary and secondary sources to produce novel research and interpretation.
An AI is at best limited to ~reading sources that human historians/archivists/librarians have already identified and digitized.
Certainly value to be had here wrt to finding needles in and making sense of already-digitized historical records, but that's more like a research assistant.
I do know some people working in classical literature that have been testing LLMs against untranslated sources and finding them perform reasonably well. It's completely within the scope of possibility to imagine them becoming more useful for academic work over time.
I don't agree. You won't cite an LLM in an academic paper as a source (since it's unverifiable and not reproducible), and claiming than an LLM's result is your own original would be fraud. So unless you never plan on publishing anything ever, what's the point?