Show HN: Mapping AI research, from 496k indexed papers
foundinghires.tech
foundinghires.tech
It is one index of the AI research literature, assembled mainly from arXiv, OpenAlex, ORCID, DBLP and the pages researchers publish about themselves. 496,295 papers, 790,725 people, 553,497 of them with a country. Every figure is a count on that index and carries the number of profiles it stands on. It is updated at least weekly.
The page is a single HTML file. The data ships inside it as JSON, about 100 KB of the 165 KB, so there is no API call, no framework, no chart library and no tracker. Nothing on it is typed by hand: a scheduled job recounts everything and regenerates the page and the published aggregates in the same pass. The aggregates are CC BY.
Author disambiguation has been an important part of building this. Merging records that describe the same person across four sources, on shared identifiers where they exist, and on the concordance of names, coauthors and institutions where they do not. That is where the errors that remain are hiding.
What it does not measure, said now rather than defended later: publication, not all of research, so people who ship models without publishing are missing. A country is the country of the institution, as I found it way more interesting to understand the strength of an ecosystem. Subjects come from the OpenAlex taxonomy, whose buckets are uneven, which is why a label as wide as Topic Modeling sits at the top.
I work on recruiting tools, so the index behind the map is also what I work with in my day to day life. The map publishes counts only, never an individual record. Happy to go into any part of the pipeline.