Others have responded equally positively to the Latin translations.
AI as a panacea is overwrought, but there are genuine positive uses of it. Would critics rather have NO translations of these pivotal works?
Others have responded equally positively to the Latin translations.
AI as a panacea is overwrought, but there are genuine positive uses of it. Would critics rather have NO translations of these pivotal works?
The problem is not the AI, but rather a weakness in the design. In other words, a bug. Or a regression and we call them today. It is fixable.
I can't imagine that metivta and mesivta are not both transliterations of the word מתיבתא.
You need to identify the target audience and stick to their language. Are you translating for a general lay audience? An academic audience? Contemporary religious Jews? If the latter, is it a yeshivish audience or a wider group? How much familiarity with Hebrew is expected from the reader?
If you'd like, you can contact me privately and I will let you know when the updates are published. Or just check the site occasionally for updates.
I sent a Classicist the link because it's adjacent to his area of interest, and by chance he addressed this question. I don't think he'll mind me quoting him:
"There's a lot of this shite popping up at the moment. [...] There seems to be some idea that publishing a shit translation of an untranslated work is better than no translation, but that's obviously wrong."
Speaking for myself: Some of the anomaly rules are obviously patches to deal with a specific situation, and might be better off in code. But relying on the chatbot to flag wider anomalies is odd. How does it know what it doesn't know?
At a minimum, I'd use a spread of models to gain binocular vision, and I wouldn't publish until a human was prepared to sign their reputation to it.
In other words, I think what you've got there is a first draft of a translation, not a translation. Given that, if you are going to publish, I think the NoDerivatives restriction is a mistake. But that's a minor issue.
Translations of human language have no analogous mechanism. So yes, the translation might be perfect, but until a human puts their reputation on the line and says "I certify this translation is accurate", it's still shit. It's shit because of the way it was created, not its absolute accuracy (or otherwise). He doesn't need to read it.
I'm sorry, I'm not trying to upset you, and I know I won't change your mind. We just have different philosophical positions on this, I think.
Numerous people have commented to me about Hebrew and Latin works, who know what they are talking about, and none has criticized the fidelity of the translation to the source. The criticisms proffered are minor, while the praise extensive.
A couple years ago, I shared your opinion. Now I don't. You are the one who won't change your mind. Like you said, it's "philosophical" for you, while for me, it isn't. If you insist on calling it shit regardless, you are retarded.
I've done it myself. Transcription of 19th century newspaper articles to markdown, mostly. Some earlier wills (which were an absolute pig - secretary hand). Oh, and categorisation of postcards. That's why I was poking around your pipeline - to see if I could learn anything. You're right, tabular data is hard. Also columns, and proper nouns.
Feeding the LLM a context-aware cheat sheet helped with the nouns. BTW, what I said about using multiple models for parallax was good advice.
Do you know you're very spiky?
Hebrew is also much worse than Latin, which uses Arabic numerals. Hebrew conventionally uses the letters for numeric representation too.
I am getting good (and improving results), but the token cost is heavy.
Have these people gone on the record saying that? What are their credentials?
I got to be honest, claims that some unnamed person privately told you they thought your product was good is pretty meaningless. I feel like that makes me trust you less not more.
At least my name is public and on the record, along with my work. You are irrelevant. So is what you say.
You haven't read anything. Are you retarded or disingenuous?
C'mon man, I know criticisms of your baby feel like attacks on you, but take a step back. You must see that "the name of the author of the pipeline is public knowledge" is not a sensible response to "nobody who can read both versions has attested to its accuracy in public".
Right now, everybody believes that LLMs can't self-correct (see https://arxiv.org/abs/2310.01798) and that they hallucinate when asked to do OCR tasks. It doesn't matter if that's true or not, that's the prevailing belief you're working against. If you want to convince people your pipeline can self-correct, you're going to need to supply evidence. I see two possibilities:
(1) Pay someone to audit one of the translations publicly.
Crucially they need to answer the question: does the pipeline introduce inaccuracy? If it does then it really is worse than useless, because it tells lies. Clunky, inconsistent, partial translations would still be better than no translation - it's the risk of hallucination that's the killer.
(2) Are you doing test runs on similar documents that have also been translated by humans? (Preferably very recent translations). Publishing those test runs would allow poor uneducated slobs like me to line up a human translation and your machine translation and see for ourselves that your pipeline works.
This is meant as helpful advice - I'm trying suggest paths that respond to valid critique with something other than bluster. I wish you and your project well.
These are the kind of comments I also typically get regarding regressions. Problems with mostly mechanically repairable aspects of the translation, but not with the quality of the translation itself. From people who know of what they speak. That I do not put them on the record is not relevant, except to people like you. They are real, and their attestations are real.
You don't, yet I still stand by my work.
I translated some books up to four times as I refined the process. The default behavior is repeatedly to not attempt to translate text it can't resolve for whatever reason. AIs do have a tendency to hallucinate, and I expended a lot of effort on minimizing this problem to the point it can be considered mitigated.