uh actually, _we_ did (generates a Docker-style manifest on the fly)
1,982 karma · joined September 10, 2010
uh actually, _we_ did (generates a Docker-style manifest on the fly)
BTW the reporting in this article is sloppy/incorrect
Might think about it:)
GitHub repo: https://github.com/bigscience-workshop/promptsource
SQuAD is the prototypical example of a dataset for this task, see https://rajpurkar.github.io/SQuAD-explorer/
I've done it for one of my team members, it's pretty easy.
(nor on the same scale of corpora, as far as I can tell)
Main features: - Encode 1GB in 20sec - Provide BPE/Byte-Level-BPE/WordPiece/SentencePiece... - Compute exhaustive set of outputs (offset mappings, attention masks, special token masks...) - Written in Rust with bindings for Python and node.js
Github repository and doc: https://github.com/huggingface/tokenizers/tree/master/tokeni...
To install: - Rust: https://crates.io/crates/tokenizers - Python: pip install tokenizers - Node: npm install tokenizers
Disclaimer: built it.
Thanks for the contributions everyone, and glad to have bet on PyTorch a few years back.
[disclaimer: made it]