I crawled very deep into this particular rabbit hole. I built text compressors that exploited linguistic structure. I was able to quantify improvements in my understanding of text by observing improvements in the compression rate. The beautiful aspect of this project is its rigor: lossless compression is such a demanding challenge that you cannot possibly deceive yourself. If you have a bug in your code, you will either decode the wrong result or fail to get a codelength reduction. Using this methodology I was able to build a statistical parser without using any labeled training data.
Unfortunately for me, the advent of GPT and other DNN-based LLMs made this research obsolete. It may still be interesting for humans to build theories of syntax and grammar, but GPT has this information encoded implicitly in its neural weights in a much more sophisticated way than linguistic theories can achieve.