A transformer-based method for zero and few-shot biomedical NER
arxiv.org
arxiv.org
This paper is extremely similar in domain (same corpus, etc - but that's not surprising since everyone uses these) but they're leaning heavily on the pretraining allowing capabilities towards few and zero shot, which is already well understood. Ultimately I think it's a good resource to use for the code, and if the API ends up being easier to change some of the internals on compared to the many options out there such as scispacy, or any of the pipelines used to achieve Pubtator, then it's a welcome addition.
My assessment is that this is a useful alternative where there are many solutions, but mostly an engineering product, and quite far away from any scientific contribution.
So clearly you need to have some very good hardware to process all of it.
However, compare that with some of the simpler encoder models that have far fewer params that are targeted for specific tasks. These systems can plow through 10^5 or 10^6 tokens per second. So now that 16 years is a week.
This is why small models that are task specific are so important. They make much more possible in reasonable time frames and reduce CO2 emissions by orders of magnitude. Along the same lines of "why use an LLM to extract everywhere the string 'Starbucks4{:digit:}' appears in text when you can use a regex?". You can get the output in a few seconds on billions of articles with a DB, whereas an LLM would take more than a decade.
:-P
There's always tradeoffs. Regex is fast, linear CRFs are quite fast too. Simple LSTMs are fast-ish, as well as BERT-like systems, can offer a decent tradeoff in speed/performance. LLMs are much slower, and need some type of distill step by step to get anything very useful out (for task specific model).
Ultimately, you're right in that regexes along with other older techniques should be understood and weighed for what is the optimal solution for the task.