Ask HN: How to Build an LLM of the Buddha?
Some ideas that pull me to thinking about this:
- LLMs build a world model and are not 'just' parrots
- The Buddha spoke and people wrote it down, maybe the 'original' world model can be extracted from the writing?
- The Pali canon has 15-17k pages, sounds like a lot but it's a small dataset compared to how much data is used to train models these days
- There's orders of magnitude more commentary that could be used, but does it detract/dilute the 'original' world model?
- Say we dump everything written about buddhism, there's probably redundancy but maybe enough material to get 'enough' data to train