At most, your project is a simple language model, but definitely not a large language model.
After looking at your code, you also seem to have a wrong understanding of embeddings. Other people in this thread have already shared great resources that offer good explanations on these topics.
You seem to be very dismissive of Markov chains (or HMMs) while they've been used in NLP for years and produce really great results. Your project seems to use some concepts of Markov chains but more simplified. Here [1] is a random article that does next-token prediction using Markov chains. It's basically what your project does but in a more scalable and probabilistic way.
>Use cases include: Auto-completion, auto-correct, spell checking, search/lookup, conversation simulation (chatbot), and more.
These use cases are not possible using the concepts used in your project.
[1] https://bespoyasov.me/blog/text-generation-with-markov-chain...