Hey! This is really cool! We did the same but for Presto and Spark by following a similar approach (mostly focusing on Partitioning and Bucketing, but the intermediate metadata datasets ended up being leveraged by multiple teams in the company) Looking forward for the VLDB presentation.