What about the non giant llm families? are they worth it when it comes to direct data extraction?
Veryfi and Taggun seem like good data extraction options and also have on prem versions
becomes a pain when dealing with docs in bulk, or with docs that have multiple pages. Token length limits with GPT with a single request and I see a loss of context accuracy when extracting insights in multiple requests.