You don't necessarily need to provide everything. Arthur (our AI) is smart enough to see exactly which information it needs to answer a given question. but, yes, the more information you provide the easier of a time the AI will have in answering your question. Arthur doesn't guess. if there is crucial information it needs he will ask for it. it doesn't have to be a Plaid hook up, a csv or even a simple user response is a start.
On your third point — you'd need an outcomes dataset — that's true for traditional ML, but it's not how this works. The normative layer is finance itself (life-cycle theory, tax rules, amortization) implemented as deterministic calculators, with the LLM doing explanation and elicitation. The paper under discussion is sort of the proof: the models already give theory-aligned advice with zero outcome training. The gap it found is input quality and statelessness, not a missing training set.