I saw this on HN before, but I thought it was another from-scratch llama implementation... Which is fine, but much less interesting to me, as a from-scratch implementation probably not as fast/feature packed as llama.cpp or the TVM implementation.
Keeping up with llama.cpp's rapid evolution is very difficult, and there's a need for projects like this.