You might want to highlight chunking and how embeddings can/should represent subsections of your document as well. It seems relevant to me for cases like similarity or semantics search, getting the reader to the relevant portion of the document or page.
Theres probably some interesting ideas around tokenization and metadata as well. For example, if you’re processing the raw file I expect you want to strip out a lot of markup before tokenization of the content. Conversely, some markup like code blocks or examples would be meaningful for tokenization and embedding anyways.
I wonder if both of those ideas can be combined for something like automated footnotes and annotations. Linking or mouseover relevant content from elsewhere in the documentation.