You give it a pattern for each token type, and a function to be called on each match, and you get back a list of processed tokens.
Importantly, it processes the list in one pass and ensures the matches are contiguous, where a naive `re.findall` with capture groups will ignore unmatched characters. You also get a reference to the running scanner, so you can record the location of the match for reporting errors.
import re
scanner = re.Scanner([
(r"[0-9]+", lambda scanner, token:("INTEGER", int(token))),
(r"[a-z_]+", lambda scanner, token:("IDENTIFIER", token)),
(r"[,.]+", lambda scanner, token:("PUNCTUATION", token)),
(r"\s+", None), # None == skip token.
])
results, remainder = scanner.scan("45 pigeons, 23 cows, 11 spiders.")
assert not remainder
print(results)
[('INTEGER', 45),
('IDENTIFIER', 'pigeons'),
('PUNCTUATION', ','),
('INTEGER', 23),
('IDENTIFIER', 'cows'),
('PUNCTUATION', ','),
('INTEGER', 11),
('IDENTIFIER', 'spiders'),
('PUNCTUATION', '.')]
[0]: https://stackoverflow.com/a/693818/252218[1]: https://en.wikipedia.org/wiki/Lexical_analysis#Tokenization