Vāgdhenu: A Sanskrit Chanting TTS System
prathosh.in
prathosh.in
From the introduction;
Classical Sanskrit recitation, parāyaṇa, is a chanted rather than a read register. A faithful synthesizer must hold long vowels, sustain a terminal visarga, articulate retroflex and aspirated consonants, render dense consonant conjuncts cleanly, and respect the metrical structure of the verse.
None of these is well served by general-purpose text-to-speech, and there is essentially no chant-domain training data available off the shelf. The problem is therefore doubly hard: it is low-resource, and the target prosody is a specialized melodic contour rather than ordinary read speech.
This report describes a system, Vāgdhenu, that solves the practical version of this problem well enough to ship two large deployments, and it documents the design decisions, the dead ends, and the one negative result that turned out to be the most useful finding. We do not claim a new model. We claim an honest account of what it takes to build a faithful Sanskrit chant pipeline on top of current open backbones, what works, and what is architecturally out of reach.
Our framing is that of an experience report. The evidence we offer is the comparative lineage across architecture families, a reproducible production system, two shipped artifacts at real scale, and a public release of code, weights, data, and a live demonstration. Formal listening studies are limited to expert evaluation, which we state plainly and treat as a limitation rather than a result.
I have lots of questions: what's the use case for this? Primarily religious / liturgical? It also seems that the fact that this tool works means that any arbitrary Sanskrit sentence can be translated into a chant by some sort of procedure (dare I say algorithm)? I'm terribly curious and fascinated by this!
Hindu Vedic pandits and priests who come in and run priestly events and functions at Hindu homes are usually well known within their communities are not treated as strangers and neither is there an element of inconvenience. If anything their absence adds an element of inconvenience. For one most rituals in sanskrit are available on youtube for those who want to diy. But most discerning Hindus prefer the real thing. Same as Christians.
Now you can't automate him either. His presence is there because religious rituals and practice and tradition mandate his presence, but if you could just play the chants over a speaker and have people understand and follow the ritual instructions some other way, nothing would really change.
For what it's worth, people can and do use AI visuals and sermons in Christianity as well
If you have been to many hindu events you would know Vak (speech) and Prana (breath) are both required to deliver vedic chants. Why do people go to musical concerts when they can turn on spotify? Use common sense.
Priestless hindu functions at your home will get laughed at so save yourself some embarrassment.
Were their presence not required, absolutely people would turn to recorded voice as 90% of people don't know what is being said anyways
It does not. The majority of the world is Muslim or Christian and this does not apply to those religious rituals
Sankrit is the mother language of a lot of indian languages that Hindu’s speak, and can be reproduced without loss of fidelity in many languages. Sanskrit is also easily understandable and readable if you know the language you speak enough to be able to write it. I dont think you can reproduce arabic version of the quran in bengali or urdu or indonesian. And again, majority of muslims do not understand Arabic to be able to use it for more than daily prayers.
Having a spoken reference that too as a chant makes it easy to recite the verses.
Besides, it’s a fun exercise, so why not?
From what I can tell, it used 5.3 hours of single voice fine-tune data.
I've noticed a lot of people seem quick to dismiss something as vibe-coded these days.
This is most likely done by claude as per the aesthetics. Every AI has its own aesthetics as well.
Sanskrit does not do schwa deletion. So engines must map the phonemes to Kannada/Telugu (which support the full complement of Sanskrit sounds) if they want to get somewhere. A lot of this can be solved if the backend supports ipa.
But this is for classical Sanskrit. Vedic recitation is too complex (and I know too little about it) to handle without a LOT of work.
It is neither.
The fundamental issue is that there is no way to represent the schwa sound in English[1]. All of a,e,i,o,u have been used to communicate the schwa (or schwa-like) sound in English.
- The `a` in about
- The `e` in taken
- The `i` in cousin
- The `o` in button
- The `u` in upon
1 - https://ashishb.net/linguistics/schwa/Oddly enough amongst the Indic languages, Tamil has precisely the same problem: க, ப, ட may be voiced or voiceless plosives depending on context. Sanskrit loanwords in Tamil are typically pronounced very differently compared to their original in Sanskrit, or even in Telugu.
Personally when transliterating any Indic language I tend to use ISO-15919, and in that scheme, it is Rāma, and 'a' represents the schwa. Or Rāmaḥ, Rāmam, Rāmē, or Rāmasya, whichever grammatical attachment is appropriate.
The level of phoneticism between English language and its latin script is not evenly spread.
For example, the letter "c" might mean /s/ or /k/ sound. However, the letter "k" almost certainly means /k/ sound.
In some ways, the schwa sound in English the worst as there is no symbol which is committed even partially towards it.
> it is Rāma, and 'a' represents the schwa.
Sure, but if my name is Rāma, I have to choose either "Ram" or "Rama" or "Raama" for my passport name. Or legal name in most places, non-alphabetic symbols do not work with all modern systems.
> Many other languages using the Latin script (German, Italian, Finnish) don't have this problem, what you see is what get.
English imports spellings from other cultures and adds its own layer of pronunciations. Other languages probably don't do it as often.
> English imports spellings from other cultures and adds its own layer of pronunciations. Other languages probably don't do it as often.
English is almost unique in how inconsistent its orthography is especially with respect to vowel sounds, and loanwords are only the latest manifestation of this problem. Even consonants aren't spared. Consider the digraph 'th'. This can be a voiced or voiceless dental fricatives. English used to have letters for each: ð and þ. 'ough' has nine pronunciations—plough, though, through, thorough, cough, rough, bought, lough, hiccough.
English has experienced contact with such a wide variety of unrelated foreign languages to an extent few others ever have.
nb: Bhatta is Sanskrit for priest.
I just tried this
यं ब्रह्मा वरुणेन्द्ररुद्रमरुत: स्तुन्वन्ति दिव्यै: स्तवै-
र्वेदै: साङ्गपदक्रमोपनिषदैर्गायन्ति यं सामगा: ।
ध्यानावस्थिततद्गतेन मनसा पश्यन्ति यं योगिनो
यस्यान्तं न विदु: सुरासुरगणा देवाय तस्मै नम: ॥
Interestingly, the word ब्रह्मा which should say Brahmaaaa says baa. Please take a look.
What's the point anyway? If you really love your traditions, slokas and chanting, please keep it as natural as possible, instead of plasticizing it to the core.
This is third level of mechanizing the sacred rituals. First - we lost the live chanting due to recorded recitals with human voice. Then we lost the experience of presence due to online streaming. And now even human voice is lost - the broken chanting is generated. So, combined all together, you have a streaming of the rituals which was performed by recorded chanting that never really involved a human reading the sloka.
Just because we can pollute the world, we need not do it.
Most people don't. Slokas and chanting are just traditions we do begrudgingly to enable the actual social objectives (family bonding, etc) of the rituals; you'd be hard-pressed to find anyone who actually understands them, even if they have it memorized by heart. I'd say automating the recitation would be quite a good sell for most families, it gives a more convenient and reliable option in place of having to invite strangers over (and reinforcing caste roles) just for that.
> Vāgbodhinī A Sanskrit chant tutor. Paste any śloka (or prose) in any script, hear its metre-aware reference chant rendered by Vāgdhenu, then chant along — a Sanskrit speech model scores every syllable and shows you what to fix. For classical (laukika) Sanskrit.
It doesn't have to be a replacement. Anyone who has ever been in the physical presence of an expert reciting sanskrit verses knows that this can never replace that experience.
It will only make the language and the chants more accessible.
I have typed the word in English literals, and it made a horrible pronunciation of the word.
> It will only make the language and the chants more accessible.
No, it will turn sacred chanting into an AI slop. Entirely due to desire of people wishing to make some quick bucks or cheap reputation.
I tried out the Gayathri Manthra (a simple test) using Sanskrit, Tamil and English scripts and as expected, the first two render fine but not the last.
This is another tool to help correct that balance.
The author seems a pretty knowledgeable person - https://prathosh.in/
His talk "The Non-negotiability of Shastra-s in Modern Science" - https://www.youtube.com/watch?v=i5Tp4gdQiTM
His goals need not be what you want them to be; his project; his rules.
Nobody is claiming this is something earth-shattering; just that it is an interesting project on a difficult subject matter. The author has linked to the paper for the project on the site itself which you can read for your edification.
From the introduction;
Classical Sanskrit recitation, parāyaṇa, is a chanted rather than a read register. A faithful synthesizer must hold long vowels, sustain a terminal visarga, articulate retroflex and aspirated consonants, render dense consonant conjuncts cleanly, and respect the metrical structure of the verse.
None of these is well served by general-purpose text-to-speech, and there is essentially no chant-domain training data available off the shelf. The problem is therefore doubly hard: it is low-resource, and the target prosody is a specialized melodic contour rather than ordinary read speech.
This report describes a system, Vāgdhenu, that solves the practical version of this problem well enough to ship two large deployments, and it documents the design decisions, the dead ends, and the one negative result that turned out to be the most useful finding. We do not claim a new model. We claim an honest account of what it takes to build a faithful Sanskrit chant pipeline on top of current open backbones, what works, and what is architecturally out of reach.
Our framing is that of an experience report. The evidence we offer is the comparative lineage across architecture families, a reproducible production system, two shipped artifacts at real scale, and a public release of code, weights, data, and a live demonstration. Formal listening studies are limited to expert evaluation, which we state plainly and treat as a limitation rather than a result.
As a practical example; if you have ever heard even the common Gayathri Mantra being chanted by a Tamil Thanjavur Brahmin priest vs. a UP Allahabad Brahmin priest you would know the wide difference in utterances. Having something like this gives you a reference (albeit with a South Indian Kannada bent) to compare against.
Do not post on HN if the goals/rules/project are not be questioned. This is not a bhajan forum.
> You do need to know about a person to understand why he did a project that he did. What were his motivations and goals? That is the point here which can be inferred from his bio/talks.
If your approach to assess something is only by learning about the person, and not by what they did or said then you need to know about me as well. But you came to conclusions without knowing about me, which violates your own principle.
Also, if you think it is blasphemous to question the work done by your great person, please do not post it here. I'm not into "person-worshipping" or andh-bhakt cult. And this is not a thread to discuss about the person. The post talks about the work and my comments are in that context. I'm not required to get into bio/talks/profile or the perceived greatness of a person.
Who said they cannot be questioned? But the author/submitter is not obligated to anticipate/satisfy/answer any of your questions/assumptions/etc. Something is posted on its own terms and you take it or leave it with/without criticism. Everything is volitional.
In this context this recent cnn article "‘Whataboutism’ makes the internet exhausting. Why people think this way" (Bean Soup Theory) is very relevant - https://edition.cnn.com/2026/07/16/health/bean-soup-theory-s...
> If your approach to assess something is only by learning about the person, and not by what they did or said then you need to know about me as well. But you came to conclusions without knowing about me, which violates your own principle.
Learning about a person behind some work is one factor but not the only factor. In this case it was relevant since the author is both a AI researcher and a Sanskrit scholar and so the motivations/goals are easily inferred.
If you want people to know about you (presumably w.r.t. to your scholarship in this domain) then you should have something in your profile instead of being anonymous and/or made it explicit in your comment. In its absence, people can't imagine things.
Using talking points like "bhajan forum", "person-worshipping", "andh-bhakt cult" etc. is simply you projecting your assumptions/biases and nothing more. Pointing to a person's background has to do with establishing domain expertise and technical credentials.
You can post any comment you want but so can others on your comment whether you like it or not. People can dismiss you the way you dismiss others.
This is not true at all. Live chanting, live pUjAs, etc. are still very much practiced. It's just that recorded versions also became available, making our lives richer.
The same applies to other art forms. E.g., bharatanATyam is still performed live, but we also have recorded performances. That is a good thing!