One of the harder parts of the problem is that CIMs aren't standardized. The same metric can appear multiple times across a document in different formats (tables, narrative text, adjusted figures, projections, etc.).
Our pipeline extracts candidate values, normalizes them into structured metric families, and links them back to document locations so analysts can verify where each number came from.
We're still improving accuracy around projections and adjusted metrics since those are often labeled inconsistently across documents.