For example, I copy-pasted some text from a HTML page where each table cell ended up on its own line, like so:
Enable feature?
Yes
Expand things?
No
This could be fixed with some multi-line regex, except that it was interleaved with headings that messed up the key-value pairings.I fed this into GPT 4 and it correctly surmised which rows were headings, keys, and values. This is an easy task for a human, but a shocking thing to see a computer solve without years of programming effort put into solving this specific problem!
You've got to wonder how many of these single-purpose AI models like Table Transformer are going to be subsumed into LLMs. For example, Table Transformer comes with a bunch of labelled training data. Just point the variant of GPT 4 that has vision at the training data to make a tuned version. That should outperform a "small" model because it both understands what it's seeing and has been tuned on the special-purpose data set.
What I'm saying is that if you have a terabyte of specific training data and train a model on it, that's fine, but then you'll have a model with "1 TB of knowledge" at most. If you start with a pre-trained LLM with petabytes of knowledge crammed into it, then adding that 1 TB would give you the benefits of both, but the benefit of the petabyte vastly outstrips the extra terabyte!
With LLMs being quantized down to just a few gigabytes and able to run on mobile devices, I wonder if this is what the future of AI will look like. No more training models from scratch...