Ars Astronomica – English translations of rare Hebrew and Latin astronomy texts
arsastronomica.com
arsastronomica.com
Sharing some artifact of your own added value, or things you learned in the process, would be more interesting to me than the output.
You are welcome to try. It is not rocket science, but what I developed would take much effort to reproduce.
I wrote an origin story about how I got started: https://aicentral.substack.com/p/pipeline-to-the-past
It describes some of the added value and lessons learned.
https://github.com/sweisman/translation-pipeline
What I haven't found in either the OP or Github repo is what your background is - software development? astrophysics? paleography? two or three of those? [edit: history of science?] - and what human review is involved. (My apologies if I overlooked anything; I don't have time to read it all!)
Your suggestion makes sense, and you are not the first to offer it.
Maybe someday, although probably not CC licensed.
Does copyright actually belong to Scott Weisman if all the words are outputs of LLMs? Is there any relevant case law here? I'm very curious about whether LLM outputs are copyrightable by the LLM user (I'm guessing potentially it varies throughout the world).
An original human-authored translation is a derivative work that can hold copyright, but it is the human-authored parts that give it protection.
No amount of human authored pipelines that are automated (no human input) would give it this status as that is not human authorship (the pipeline itself can be however). The prompts for the LLM can also be copyright (again…human authorship), however they would be difficult to enforce since the direct output can’t be copyrighted.
In what respect are they being compared, from your perspective? How do you think they are not comparable? What is the harm of comparing them? How would Copernicus and Kepler help and what is the drawback of multiple Brahe entries?
> I can't really know what your intention was
My intention? I have nothing to do with the OP. Many assumptions jumping up here.
That being said, Ralbag, as we refer to him, or Gersonides, as you do, was both a highly regarded (and controversial) rabbi and esteemed astronomer who pioneered empirical observation and measurement.
My comment is indeed badly written. But if you go over the further comments the submitter made, it should become clear what bothered me.
So, for clarity (and not trying to be politically correct, but just say what I think is true): there is the unfortunate tendency among today's jews, seeing their success to bring US to fight on their behalf for their interests, to see in that a proof of them to be the Chosen People - in a very concrete and materialistic way.
One of the conflicts that such point of view generates, is conflict with history: in no way scientific thinking came from Judeah, it came of course from Greece ("but how can this be? It must have been the Chosen People who invented science! Look at all these Nobel laureates!").
And the revival of the Greek thought happened with the Italian Rinascimento, and from this came Tyho Brahe, Copernicus, Kepler and so many others.
While the jews were always closed within themselves (it is not true that it was only goyim who wanted them to be in ghettos, they themselves did not want very much to go out..). And did not start to occupy themselves seriously with science at least until the Enlightenment.. and these were not rabbis.
It is not the first time I see my conationals trying to prove that it is false. One could have just laughed at this, were it not that these people are so determined in their efforts. Really, just read the submitter comments..
You have not provided any clarity. It's a mishmash of bitterness and anger.
There is no "of course". You are ignorant of history.
In any case, there is no connection between "invented science" centuries or millennia ago and "all these Nobel laureates."
The documentary historical record is clear. Some of the Jews in my corpus pioneered things during the Renaissance period. One is a very early documented use of decimal numbers, as one example (https://arsastronomica.com/works/immanuel-bonfils-shesh-kena...). They pioneered empirical measurement centuries before Tycho. They invented instruments relied up for astronomy and navigation.
Further, Jews who are not ignorant know that the claim of Chosenness does not rest on intelligence or accomplishments.
Gans met and wrote about his encounters with Tycho. It was these documented encounters that drove my interest in translating the work.
In chapter 25 of the same book I mentioned, on the famous Gemara where Chaza"l concede to the chachmei goyim in Pesachim, Tycho says the Jews were right and the goyim wrong, and the Jews were wrong to concede.
Gans also met Kepler.
In any case, my email is on the site and in every book. Feel free to contact me privately to continue the discussion.
The mere fact that the translation was AI-assisted shouldn't be a reason to flag a comment. AI translations, especially supervised translations, can actually achieve a reasonable level of quality.
Disagreement on AI philosophy shouldn't be a reason to flag a contribution.
Others have responded equally positively to the Latin translations.
AI as a panacea is overwrought, but there are genuine positive uses of it. Would critics rather have NO translations of these pivotal works?
You don't, yet I still stand by my work.
I translated some books up to four times as I refined the process. The default behavior is repeatedly to not attempt to translate text it can't resolve for whatever reason. AIs do have a tendency to hallucinate, and I expended a lot of effort on minimizing this problem to the point it can be considered mitigated.
I sent a Classicist the link because it's adjacent to his area of interest, and by chance he addressed this question. I don't think he'll mind me quoting him:
"There's a lot of this shite popping up at the moment. [...] There seems to be some idea that publishing a shit translation of an untranslated work is better than no translation, but that's obviously wrong."
Speaking for myself: Some of the anomaly rules are obviously patches to deal with a specific situation, and might be better off in code. But relying on the chatbot to flag wider anomalies is odd. How does it know what it doesn't know?
At a minimum, I'd use a spread of models to gain binocular vision, and I wouldn't publish until a human was prepared to sign their reputation to it.
In other words, I think what you've got there is a first draft of a translation, not a translation. Given that, if you are going to publish, I think the NoDerivatives restriction is a mistake. But that's a minor issue.
Translations of human language have no analogous mechanism. So yes, the translation might be perfect, but until a human puts their reputation on the line and says "I certify this translation is accurate", it's still shit. It's shit because of the way it was created, not its absolute accuracy (or otherwise). He doesn't need to read it.
I'm sorry, I'm not trying to upset you, and I know I won't change your mind. We just have different philosophical positions on this, I think.
Numerous people have commented to me about Hebrew and Latin works, who know what they are talking about, and none has criticized the fidelity of the translation to the source. The criticisms proffered are minor, while the praise extensive.
A couple years ago, I shared your opinion. Now I don't. You are the one who won't change your mind. Like you said, it's "philosophical" for you, while for me, it isn't. If you insist on calling it shit regardless, you are retarded.
I've done it myself. Transcription of 19th century newspaper articles to markdown, mostly. Some earlier wills (which were an absolute pig - secretary hand). Oh, and categorisation of postcards. That's why I was poking around your pipeline - to see if I could learn anything. You're right, tabular data is hard. Also columns, and proper nouns.
Feeding the LLM a context-aware cheat sheet helped with the nouns. BTW, what I said about using multiple models for parallax was good advice.
Do you know you're very spiky?
Hebrew is also much worse than Latin, which uses Arabic numerals. Hebrew conventionally uses the letters for numeric representation too.
I am getting good (and improving results), but the token cost is heavy.
Have these people gone on the record saying that? What are their credentials?
I got to be honest, claims that some unnamed person privately told you they thought your product was good is pretty meaningless. I feel like that makes me trust you less not more.
At least my name is public and on the record, along with my work. You are irrelevant. So is what you say.
You haven't read anything. Are you retarded or disingenuous?
C'mon man, I know criticisms of your baby feel like attacks on you, but take a step back. You must see that "the name of the author of the pipeline is public knowledge" is not a sensible response to "nobody who can read both versions has attested to its accuracy in public".
Right now, everybody believes that LLMs can't self-correct (see https://arxiv.org/abs/2310.01798) and that they hallucinate when asked to do OCR tasks. It doesn't matter if that's true or not, that's the prevailing belief you're working against. If you want to convince people your pipeline can self-correct, you're going to need to supply evidence. I see two possibilities:
(1) Pay someone to audit one of the translations publicly.
Crucially they need to answer the question: does the pipeline introduce inaccuracy? If it does then it really is worse than useless, because it tells lies. Clunky, inconsistent, partial translations would still be better than no translation - it's the risk of hallucination that's the killer.
(2) Are you doing test runs on similar documents that have also been translated by humans? (Preferably very recent translations). Publishing those test runs would allow poor uneducated slobs like me to line up a human translation and your machine translation and see for ourselves that your pipeline works.
This is meant as helpful advice - I'm trying suggest paths that respond to valid critique with something other than bluster. I wish you and your project well.
These are the kind of comments I also typically get regarding regressions. Problems with mostly mechanically repairable aspects of the translation, but not with the quality of the translation itself. From people who know of what they speak. That I do not put them on the record is not relevant, except to people like you. They are real, and their attestations are real.
The problem is not the AI, but rather a weakness in the design. In other words, a bug. Or a regression and we call them today. It is fixable.
I can't imagine that metivta and mesivta are not both transliterations of the word מתיבתא.
You need to identify the target audience and stick to their language. Are you translating for a general lay audience? An academic audience? Contemporary religious Jews? If the latter, is it a yeshivish audience or a wider group? How much familiarity with Hebrew is expected from the reader?
If you'd like, you can contact me privately and I will let you know when the updates are published. Or just check the site occasionally for updates.
Amen!
That may be true, but also not really what i'm worried about. With AI on an ancient book written from a world view very different from the contemporary one, i'd worry more about misleading translations.
Check out a work or two.
My queue of Hebrew and Latin works is already lengthy, and this is my focus for now.
The site currently includes English translations of rare Hebrew and Latin works, including texts that have never previously appeared in English and others that survive only in manuscript form. All editions are released under a Creative Commons license.
I'm interested in feedback from historians of science, classicists, medievalists, digital-humanities researchers, and anyone interested in AI-assisted scholarly publishing.