How will they get the data to train the model. They have everyone’s documents, and could potentially produce amazing results. But how will they protect against content leaks?
This seems to me like the "grounding" (getting documents/data) won't be fed back into the system. It will just be used to verify that output has an actual artifact that is been based on.
It's a three part model. The LLM does not know the docs, it gets a modified prompt that includes the context