- can’t hear the differences in spelling between American and British English
- are listening to someone read. That person probably is able to read either American or British English.
That said, some words really are different and could cause confusion/disconnect so probably are potential candidates for being changed even in audio books.
- the one he gave “apartment” vs “flat” is a good example. I think British people would all be fine with “apartment” whereas I think “flat” might confuse some Americans who hadn’t heard that usage.
- “pavement” means the road surface in the US vs it means the pedestrian path on the side of the road in the UK (that in the US you would call a “sidewalk”). “Sidewalk” wouldn’t confuse a British person but it’s not a word a British person would use. “Pavement” meaning the road surface does confuse British people.
- “jumper” in the UK vs “sweater” in the US
- Don’t even ask about “pants”
- etc
On the flipside, I was delighted to hear stories of US kids confusing their parents with British English because they'd watched so much Peppa Pig during lockdown. Very good.
Imagine learning something about the world after reading (or hearing) a book!
But here's an example that's worth thinking through. This movie https://www.imdb.com/title/tt0110428/?ref_=fn_al_tt_1 was called "The Madness of King George". It's based on a play called "The Madness of George III". They changed the name when producing the movie because of the difference between Britain and the US.
In Britain pretty much everyone would know that George III refers to King George III. However when they first started discussing this project in the US, a common reaction was "I haven't seen "The Madness of George 1 and 2". Now the differences between the titles wasn't that important and clearly this was just an impediment to understanding with no benefit.
Now you could say "Well the people could go on wikipedia and find out that there was a King George and a King George II and then this guy, George III" and that's true but the problem is negative self-selection. People who don't know that don't realise that's what's at play here and so won't do that search so nobody learns anything.
Don't know why someone would "rent a flat?" You go look the phrase up and discover that it's British usage. Confusion over... because you learned. Reading Moby Dick and don't know what "scraggy scoria" is? You look the words up and learn they're the perfect fit for the landscape Melville is painting.
But there are difficult examples, even in Rowling. She uses "revision" in a way that American readers don't know ("studying") and might have a hard time piecing together from context. It's not like the reader will be prompted to open a dictionary, since "revision" is a common word in AE. It just makes for a poor reader experience.
> Righto, mate, gimme a shout as Ishmael. Few donkey's years ago—don't get your knickers in a twist 'bout when exactly—findin' me wallet as dry as a dead dingo's donger, and not a bloody thing worth a squiz on the land, reckon I'd chuck a U-ey on life, put a bit of water under me bridge. Wanted a stickybeak at the wet half of this great wide world, didn't I?
(With ChatGPT assist!)
It's never stopped me from reading a book before, but it does diminish my enjoyment quite a bit. For nonfiction, I don't care, I'm reading actively regardless. But when reading passively for enjoyment, it's noticeable and irritating.
(For those who can't read it - the text is English transliterated into Hebrew script, not actually Hebrew. I'm impressed that Google managed to make sense of it)
> И I struggled for a bit with this one, but I don't know of a better way to represent it other then "ai" in Cyrillic I guess.
The "ѵ" is interesting too, never seen that before since you can make do with the sorta B-looking letter instead.
- instead þ could be used ѳ;
- с in тектс is missed, it should be текстс or better теѯтс;
- ѵ is basically Greek upsilon which appeared in English mostly as u [auto] or y [system], therefore шѵд would be better written as шоѵд where /oѵ/ in old Cyrillic pronounsed as /u/ as in Greek, or just шꙋд;
- дж in лангѵадж could be writtend as џ;
And so on.
No, never was a thing. Early Cyrillic had a bunch of unique letters that are long gone, but it never had "þ" in there.
> The "ѵ" is interesting too
It's from Greek "Y" (upsilon) and it - depending on the place - could've meant either /i/ or /v/ sound (and /u/ when in "оѵ" digraph, so parent comment has it wrong): https://en.wikipedia.org/wiki/Izhitsa
If you want to see how it should look in context, here's a typical 17th century example: https://www.raptisrarebooks.com/images/86066/paradise-lost-a...
The prints from the 18th century are even nicer with better quality and consistency in general. The Caslon typefaces are pretty exemplary here: https://upload.wikimedia.org/wikipedia/commons/4/45/A_Specim...
Roman typefaces with a long-s are a very 17th-18th century thing. A blackletter (gothic) typeface gets you closer, but all of that is still not very much medieval. What you really want is a meticulously manuscript in Carolingian Miniscule[1] or Uncial script[2], complete with killer rabbits[3] and knights fighting snails[4].
[1] https://en.wikipedia.org/wiki/Carolingian_minuscule
[2] https://en.wikipedia.org/wiki/Uncial_script
[3] https://blogs.bl.uk/digitisedmanuscripts/2021/06/killer-rabb...
[4] https://blogs.bl.uk/digitisedmanuscripts/2013/09/knight-v-sn...
Upvoting because this is seriously interesting. I can see old or Tolkien English being a hurdle, so I relate to what you’re saying. But small changes in diction tend to draw me into the world, by othering it from the familiar.
“Sophomore”, maybe?
and then run this prompt on each pair of words that you want it to do. When it's done, run a quick diffcheck on it & check its work to learn of any gotchas.
Are we really going to let big corp automate away their role in the social contract? Ostensibly Random House needs to provide real output not just be a brand that farms out all the work to AI, and thus captures value on the fiat ledger.
The number of rent seeker non-contributors is too damn high. We cannot base society off dying peoples hallucinations about how the world works anymore. This is bonkers.
I never asked to exist. Why am I constrained to beliefs like Bill Gates and the like are divine in contemporary ways?
That said, I found myself doing little text mangling tasks with GPT-4 instead of appropriate command line tools, because fuck if I remember the flags to sort and unique, and it's much faster to just describe what I want and have the LLM take a crack at it.
This is why, I think, you'll see a lot of simple(ish) tasks being done by LLMs, at couple orders of magnitude more compute cost - it's just that much more convenient.
EDIT: nevermind. That was about just doing regex replace.
But if you want to do it correctly in an automated way, LLMs are indeed the best tool for the job in the general case, because they understand how language work, in a way that you can't really formalize in code. No, feeding the complete definition of English grammar to the computer won't help, because a) it probably doesn't exist, and b) even if it does, it's merely a suggestion - natural language is a living thing, and is not bound by fixed rulesets.
"Oh, I live near there too! House or flat?"
Stanley hesitated. These questions were getting more and more personal. Was this her idea of casual conversation? Or was she trying to get to know him personally? Well, he thought, what could go wrong if I treat this like a conversation. "Flat," he said. And then asked a question of his own. "What kind of pop do you like?"
How are you going to decide if "flat" is talking about non-fizzy pop/soda, or a dwelling space with a regex? It's a lot closer to the word "pop."
Of course, GPT handles this perfectly https://chat.openai.com/share/fb6a56a2-749c-4e71-9506-541bb8...
But I'll admit I agree that this is well beyond the scope of regex.