HNHacker News
TopNewBestAskShowJobs

ImperiumOfMan

25 karma · joined April 11, 2024

submissionscomments
ImperiumOfMan··on New AI Books5 Dataset for training LLMs further (July 2024 update)
- More than 600,000 fiction and non-fiction book full-texts

- More than 8,000,000 scholarly publications, magazines, and manuals full-texts

- More than 5,000,000 US patents

- 164,000,000 metadata records

Format is described at https://t.me/nexus_search/226

ImperiumOfMan··on AI Libstc3 (July 2024) Dataset for training LLMs further
What? - More than 600,000 fiction and non-fiction book full-texts - More than 8,000,000 scholarly publications, magazines, and manuals full-texts - More than 5,000,000 US patents - 164,000,000 metadata records

Format Zstd compressed file, JSON lines, one per book/publication, 356GB - abstract, content - description and content in markdown format - issued_at - time of issuing of the object (not of the record itself) - metadata - ISBNs, publishers, series etc - id - identifier in external systems, if applicable (i.e. DOI) other fields should be self-descriptive

Download magnet:?xt=urn:btih:37504c50e6f318e8fee6e3f82a8150a888a4cdc8&dn=http://libstc3-libstc.cc.jsonl.zst&tr=udp%3A%2F%http://2Ftra...

ImperiumOfMan··on AI Dataset Libstc2 a.k.a. Books4
More than 400,000 fiction and non-fiction book full-texts. Multiple languages, curated, deduplicated.

More than 6,000,000 scholarly publications, magazines, and manuals full-texts. Multiple languages, curated, deduplicated.

150,000,000 metadata records

Download:

Torrent file below or magnet:?xt=urn:btih:a904e660355c49006b2e7d43893d31bf3c2be9cc&dn=libstc2.jsonl.zst&tr=udp://tracker.opentrackr.org:1337/announce&tr=https://tracker1.ctix.cn:443/announce&tr=udp://open.demonii....

~240GB

ImperiumOfMan··on AI Books4 Dataset for training LLMs further
Origin of the data:

- LibGen collections gathered over decades

- Sci-Hub

- Additional fresh scholar papers and books collected by Library STC