2023 office software already uses 1000x more ressources than 1990s'. I bet we are ready to do that again.
2023 office software already uses 1000x more ressources than 1990s'. I bet we are ready to do that again.
The "Information Extraction from semistructured and unstructured documents" task is seeing a huge leap, just 3 years ago it was very tedious to train a model to solve a single use case. Now they all work.
But if you do make the effort to train a specialised model for a single document type, the narrow model surpasses GPT3.5 and 4.
But I wonder how much more productive our economies could be if everyone was taught programming the same way we teach reading & writing, and open standards were ubiquitous.
Prompt engineering is turning coding problems into language problems. It’s conceivable that humans writing code becomes artisanal in a century.
Pedantically, sure. The field ChatGPT is most impactfully commoditizing is low-level coding. Instead of someone giving natural language instructions to a team of humans, they're increasingly able to give them to an LLM. It's an open question how far this can scale. But we may be near the zenith of the practicality of large-scale coding expertise.
Pedantic, maybe, but “coding expertise” isn’t going anywhere.
At the pace we’re moving at now we’re talking a few decades away at the most, well within most peoples’ career span. I feel sorry for any junior coder just entering the industry.
I see it as incredible. Most PDFs that i see are basically just thin wrappers around image scans of documents that don’t exist anywhere anymore. Archives from estates, manuals, etc.
These techniques of using LLMs to clean ocr output is game changing because best in class before was human-in-the-loop systems that required huge amounts of rewriting to get useable output.
Now LLMs are unlocking for significantly cheaper previously difficult data sources for relatively cheap.
If LLMs are deployed in large enough scale, the convenience really could justify the cost.