HNHacker News
TopNewBestAskShowJobs

bnewbold

764 karma · joined April 21, 2011

protocol engineer at bluesky (atproto.com). formerly built scholar.archive.org

bluesky handle: @bnewbold.bsky.team

https://bnewbold.net/

[ my public key: https://keybase.io/bnewbold; my proof: https://keybase.io/bnewbold/sigs/FoU08TLKiIIyXeAuGV71MNeYz_XKz0iojLiM0NMlBqA ]

submissionscomments
bnewbold··on Internet Archive Scholar: Search Millions of Research Papers
Great question!

BASE, SHARE (https://share.osf.io/), and CORE (https://core.ac.uk) all primarily pull metadata via OAI-PMH, though they may also incorporate other sources these days. We have worked with CORE to check content overlap. We have also done our own OAI-PMH bulk scraping and broad preservation crawling, but most of this content has not ended up indexed in fatcat yet.

The main reason that we haven't done bulk imports from OAI-PMH directly or from any of these sources is that we haven't gotten a handle on the metadata quality yet. There are many, many duplicate records out there, and we think it is important to merge these correctly (under "work" entities in fatcat). Until recently we didn't have a mechanism to fuzzy-match new records to prevent duplicate creation. Once we get a policy figured out and polish the de-dupe code, we expect to significantly increase the amount of content in the catalog from these sources.

Another thing we haven't figured out yet is accurately tagging OAI-PMH feeds as journals (eg, OJS instances) vs. institutional repositories or subject repositories. That distinction will change how we classify imported records (eg, "published" versions vs. "pre-print" or "manuscript").

bnewbold··on Internet Archive Scholar: Search Millions of Research Papers
This is great feedback, thank you.

For future follow-up, my work email is my handle here (bnewbold) at archive.org

bnewbold··on Internet Archive Scholar: Search Millions of Research Papers
I think it is in a good place for simple bibliometric queries. The fatcat elasticsearch API is open at https://api.fatcat.wiki/fatcat_release/ (behind a proxy to filter "unsafe" requests). That works pretty well for jupyter notebook style experimentation if you are willing to learn the elasticsearch query DSL for aggregations and things.

I don't think the catalog has high enough metadata quality today for use in published research. There are some glaring errors and omissions when you actually starting digging in. On the other hand, almost all bibliographic catalogs seem to have such problems. Fatcat, by being open and having an API, does have the potential to aggregate corrections, fixes, and contributions directly from researchers over time.

A particular missing piece today is that there is no categorization or "discipline" metadata of almost any type. This sort of metadata is more subjective, and the catalog currently carefully only includes factual information. We will likely start collecting metadata at the journal ("container") level and can trickle that down to papers. Aggregating, editing, and curating that metadata in Wikidata first, then importing to Fatcat, might be the best and most sustainable path forward.

bnewbold··on Internet Archive Scholar: Search Millions of Research Papers
Ah, sorry to hear. We in particular want to include content from outside the US/Europe publishing world.

For Japanese publishing, we have done metadata imports from JaLC (Japanese DOI registrar), and crawled a lot of open content from J-Stage (https://www.jstage.jst.go.jp/) and I hoped that coverage was pretty good. If you get a chance, could you try searching for metadata records on https://fatcat.wiki, with both Japanese and English titles and names (if applicable)?

For Korean publishing, the regional DOI registrar (https://www.kisti.re.kr/eng/) does not provide open metadata, which is a known hole in our coverage. IIRC it looked like there might be a way to scrape at least DOIs, titles, and author names, but haven't had time to take a crack at it.

Mainland Chinese publishing is probably the biggest single hole in coverage by absolute numbers. There are two DOI registrars and neither have open metadata.

Regarding the u-tokyo.ac.jp, it looks like we are able to consume metadata and do crawls via the OAI-PMH protocol. We crawled over 112k URLs from that domain via that protocol about a year ago, and they should be preserved/mirrored in web.archive.org but they haven't ended up in fatcat or scholar yet. We want to go slow with pulling in OAI-PMH content, and ensure we de-duplicate records and add filters to ensure we are getting clean metadata and content. Also preserving repository content hasn't been as urgent as getting to small OA publishers which might lack a preservation scheme and vanish off the web.

bnewbold··on Internet Archive Scholar: Search Millions of Research Papers
Fixed, thanks!
bnewbold··on Internet Archive Scholar: Search Millions of Research Papers
We are mostly not indexing on a journal-by-journal basis, but try to import from large, broad sources. For example, DOI registrars (Crossref, Datacite, J-Stage), DOAJ article and journal metadata (for OA publications), etc. Some field-specific indexes we have imported from include JSTOR early journals subset, PubMed, and dblp.

Some fields/disciplines are probably still systemically under-represented. For example, I bet we are missing a bunch of scholarship on art and history published before 1980. We have a couple ideas up our sleeves which we hope will help with "completeness" across more disciplines.

To answer your question directly, you can search journal names here: https://fatcat.wiki/container/search

And click through to see how many articles we know about, and what we think the preservation status is. Click through again to the "coverage" tab for a more detailed breakdown. (improving the usability and ranking on the journal search results is on our short list)

bnewbold··on Internet Archive Scholar: Search Millions of Research Papers
Thank you for the kind words!

We are friendly with Semantic Scholar, and have used their "open corpus" dumps as one of several URL seed lists for crawling in the past. Their search and discovery tech is more sophisticated than ours is likely to be any time soon (https://medium.com/ai2-blog/building-a-better-search-engine-...). We would love to get to the place where groups like AI2, which are primarily research-oriented, could build on an existing open catalog and corpus, and not need to duplicate time crawling, merging catalogs, cleaning metadata, etc. As of today Microsoft Academic (used by Semantic Scholar) might be a better option.

Want to be thoughtful about ranking signals, and are deeply skeptical of journal impact factor, h-index, and most bibliometrics. "Has this been cited more than a handful of times" seems like a reasonable coarse boost. Hope to include more curated signals, like "won a paper prize", "journal in DOAJ and other reviewed indices", etc.

Have been working on a citation graph, keep an eye out for something about that in coming months. One cool thing we hope to do with the citation graph is find "missing works" not yet in the catalog (eg, don't have a DOI, especially for pre-1990 era).

bnewbold··on Internet Archive Scholar: Search Millions of Research Papers
This service was hinted at back in September, but is now formally announced and live at https://scholar.archive.org

Related previous post: https://news.ycombinator.com/item?id=24485444

Much of the catalog functionality can be accessed from the fatcat.wiki API (https://api.fatcat.wiki/redoc). Scholar adds a search index over the body content of papers, and we are still thinking through how to make this available through a public API without slowing down query latency even more.

Folks here might also be interested in this CLI for interfacing with the catalog and making edits: https://gitlab.com/bnewbold/fatcat-cli

bnewbold··on UC’s termination of Elsevier contract has had limited negative impact (2020)
At the Internet Archive, we are working on one aspect of this problem: https://fatcat.wiki/

Other notable efforts, mostly envisioned and led by librarians, are "dark" digital archives (LOCKSS, CLOCKSS, Portico, etc) which usually have a mechanism to flip and "trigger" public access if the publisher vanishes; microfilming and other microform throughout the 20th century; Hathitrust / Google Books; and as others have mentioned, scalable and sustainable (cost-wise) off-campus depositories.

bnewbold··on Internet Archive Infrastructure
There are APIs and it would be great if more people and organizations built on top of them, and specifically build content or collection-specific interfaces.

Here is the entry point for API documentation: https://archive.org/services/docs/api/

Hot linking, CORS, and other things to support third-party integration are usually supported, though there are a lot of special cases for security or to prevent malicious use. If you run in to technical problems we are usually responsive to the main contact routes on the archive.org site.

The system is not designed to allow multi-party curation and editing of metadata, but there is nothing stopping folks from building third-party catalogs on top of the content stored (and served) from IA. That is sort of what openlibrary.org is for books. The same thing could be done for music, video, specific documents, etc.

Note: I work at IA but not on the APIs or archive.org collections

bnewbold··on Internet Archive Infrastructure
Performance is fun!

One aspect is that our data centers are in California, with no CDN. If you are on the other side of the world, you will have higher round-trip latency on every request, for all services.

Another is layers of caching. Popular or recently requested Wayback content is more likely to be in either an explicit cache (eg, redis), or implicitly in kernel page caches across all layers of the request.

Every wayback replay request hits several layers of index indexes (sorted by domain, path, and timestamp), which are huge and thus actually served from spinning disk over HTTP (!). This includes a timeline summary for the primary document, to display the banner. Then the actual raw records are fetched from another spinning disk over HTTP. This may result in one or more layers of internal redirect (additional fetches) if there was a "revisit" (identical HTTP body content, same URL, different timestamp). Then finally the record is re-written for replay (for HTML, CSS, Javascript, etc, unless the raw record was requested). Some pages will have many sub-resources, so this process is repeated many times, but that is the same as page load and you can see which resources are slow or not.

As mentioned in the video, depending on where we are in the network hardware upgrade lifecycle, sometimes outbound bandwidth is tight also, which slows down transfer.

And of course most of these services operate without a ton of overhead, so if there is a spike in traffic everything will slow down a bit. There is a lot of multi-tenancy-like situations also, so if there is a very popular zip file or Flash game getting served from the same storage disk as the WARC file holding a wayback resource, the replay for that specific resource will be slow due to disk I/O contention.

If you are curious about why a specific HTML wayback replay was slow, you can look in the source code of the re-written document and see some timing numbers.

Several organizations run large web archives that operate similarly to web.archive.org, and have described cost/benefit trade offs for different components. Eg, National Library of Australia has an alternative CDX index called OutbackCDX, which uses RocksDB on SSDs. I believe other folks store WARC files in S3 or S3-like object storage systems. The Wayback Machine is somewhat unique in the amount of (read) traffic it gets, the heterogeneity of archived content (from several crawlers, in older ARC as well as WARC), volume of live crawling ("save paper now" results show up pretty fast in the main site, which is black magic), running on "boring" general purpose hardware, and deep integration with our general purpose storage cluster.

Note: I work at IA but not on the Wayback system

bnewbold··on Internet Archive Infrastructure
The disks are spinning all the time, and most disks are seeing fairly frequent reads to some content or another. A lot of content is very rarely accesses, but almost every disk has some content which gets accessed. If spinning disks had only frequently-accessed content, they would be unable to keep up with the read rate or read throughput, things balance out reasonably on average.

Wayback content is on the same disks as most other content, in the form of WARC files, with individual records fetched out of the middle of WARC files via HTTP range request.

Note: I work at IA but am not on core infrastructure team

bnewbold··on Internet Archive Infrastructure
If you are interested in cost modeling for long-term digital preservation, check out this blog series: https://blog.dshr.org/2019/02/economic-models-of-long-term-s...
bnewbold··on How Internet Archive Is Ensuring Permanent Access to Open Access Journals
Link back to recent discussion on this topic: https://news.ycombinator.com/item?id=24422593
bnewbold··on Dozens of scientific journals have vanished from the internet
From what I have seen, the least technically resourced journals often use hosted platforms or free software like OJS (basically wordpress for journals), which comes with features like HTML meta tags and OAI-PMH by default.

The trickier cases are when folks write their own platforms, or even write their own raw HTML with no templating, in which case adding tags to all landing pages or supporting an API would be a relatively large amount of work.

bnewbold··on Dozens of scientific journals have vanished from the internet
Unpaywall is very helpful! However, even for direct PDF links, publishing platforms will often do things like check for a session cookie; if you don't have the correct cookie you get bounced back to the landing page, where you need find and follow another link. This isn't super complicated to work around (persist a cookie jar, use a headless browser, etc), but it doesn't work out-of-the-box with our crawlers, the same way crawling youtube doesn't work out-of-the-box so we rely on the youtube-dl community.
bnewbold··on Dozens of scientific journals have vanished from the internet
In Latin America, the SciELO network has been very successful at providing shared, low-cost, stable infrastructure for digital journal hosting using state funding: https://en.wikipedia.org/wiki/SciELO
bnewbold··on Dozens of scientific journals have vanished from the internet
At the Internet Archive, we are working on this exact problem, and have been in communication with the pre-print's authors. We have built open infrastructure (open source, open data) tracking "preservation coverage", for example:

https://fatcat.wiki/coverage/search?q=is_oa%3Atrue+year%3A%3...

and are working to improve crawling. There is a "save paper now" feature, as well as an API for bots. Organizations like DOAJ, ISSN, DOI registrars (Crossref, Datacite, others) are crucial for this. In the broader ecosystem, we hope this can complement existing efforts that partner with large publishers (like LOCKSS, Portico, JSTOR) and institutional repositories. A natural niche for us is web-native (HTML) content, which we have crawled a lot of but are just getting started to index. For example, publications like d-lib, first monday, and distill.pub.

If folks want to help, it would be great to have a "youtube-dl for open access papers". There is a lot of content on large platforms and publishers which have anti-crawling measures (even for gold OA and hybrid content!), as well as a long tail of small publishers that don't use simple/common mechanisms like OAI-PMH and the `citation_pdf_url` HTML meta tag to identify fulltext content. The OAI-PMH ecosystem sadly is not very complete or helpful for the use case of mirroring.

bnewbold··on ArchiveBox: Open-source self-hosted web archive
Have you looked at the WARC format? It's ridiculously simple, basically concatenated raw HTTP requests and responses, with some extra HTTP metadata headers mixed in (a la extra JSON metadata keys). You can open it with a text editor. Very simple and efficient to manipulate, and very efficient to iterate over or generate.

https://iipc.github.io/warc-specifications/specifications/wa...

Arguably the biggest problem is that it isn't complex enough: there is no index of contents built-in (the standard .csv-like index format of URL/timestamp/hash/offset is called CDX).

There aren't a ton of tools in the web archiving space in general, but almost all of the ones that do exist work with WARC. Existing tools (for interchange) include bulk indexing (for search, graph analysis, etc) and "replay" via web interface or browser-like application. Apart from specific centralized web archiving services that use WARC, there are several large public datasets, like Common Crawl, that are released in WARC format.

bnewbold··on More than 9M broken links on Wikipedia are now rescued
The Archive currently has about 46 Petabytes of content ("bytes archived"), and over 120 PB of raw disk capacity; the difference is due to data replication, "currently filling" storage, non-storage infrastructure, etc.

We save a lot on web content storage by de-duplicating "revists" when the page hasn't changed. This works out to save a whole lot for content like jQuery served from a common CDN URL; it doesn't work well when there is a page counter or any trivial changing content on a page.

If you are interested in the storage back-end, it's actually pretty simple: HTTP requests/responses are concatenated and compressed in WARC files (sort of like .tar.gz) that get stored on regular old ext4 filesystems. An index of "what URL captures are in what WARC files on what servers" is continuously generated in the form of, basically, a giant sorted (and shareded) .tsv file; replay requests on web.archive.org look up the URL and timestamp and get a reference to a machine, file, and file offset, and make an HTTP 1.1 range request for the content in question. There are a bunch of other details, like checking robots.txt status, but the core design is super simple, cheap, and (relatively) easy to operate at scale.

Apart from web crawl content (including, these days, "heavy" video content which is difficult to de-dupe), we have a large amount of live recorded TV, scanned books (raw photos), etc.

(I currently work at IA)

bnewbold··on Sci-Hub Proves That Piracy Can Be Dangerously Useful
They are very different beasts. Arxiv (and most pre-print repos) accept submissions (with filters, like requiring academic affiliation or vouching, and requiring reasonably-formatted metadata), have a moderator do a skim-level review of the work, and then post it. CiteseerX is an automated crawler, like Google Scholar, which finds PDFs on author homepages and extracts metadata from the PDF. There is no human review or cleanup process, and there can be a long delay before content gets discovered.
bnewbold··on Sci-Hub Proves That Piracy Can Be Dangerously Useful
I could be misinterpreting your comment, but it reads like a broad misunderstanding of the role, economics, and value-add of contemporary publishing.

Here is a (now somewhat old) breakdown of per-article costs, contrasting open access and subscription journals: https://www.nature.com/news/open-access-the-true-cost-of-sci...

From my perspective the main values that an modern open access publication provide are primarily emotional labor: 1) herding cats: all those reviewers and editors that are volunteering are incredibly flaky and need to be hounded by phone, in person, etc. This can't be automated because everybody ignores bots. 2) resolving disputes. hopefully the median paper goes through the process smoothly, but any dispute involving prominent researchers will require dozens of hours of careful mediation, fact-checking, and follow-up. 3) generally being a responsible institution (not suffering fraud or embezzlement, planning for change over time, etc), which is valuable because humans don't have the time or capacity to judge every piece of content from scratch and fall back on reputation. All of these tasks require competent, savy, and highly trained folks, or the whole thing falls apart. Many in the old (subscription) world would say that open access publishers have already cut corners at $3k/article, and if you read around on the internet everybody complains about how long it takes to get responses or resolve disputes, so maybe they are right.

If these numbers sound totally bogus to you, maybe you should jump in to the publishing business and undercut everybody with superior service? Try writing a business model. As to whether we need publishing at all, i'd also try convincing researchers and authors to just post their results on blogs or wikis, which have zero barriers, require no new development, etc. I think authors/researchers value the services publishers provide, even apart from the whole branding/reputation/incentives game. Even pre-print repositories involve a degree of labor to review, moderate, and resolve disputes.

"automatically typeset using LaTeX" is only a norm in a few fields, which are all relatively well resourced and tech-savy. The fact that extra typesetting labor is involved in fields that don't use LaTeX might be one reason that pre-prints have been slower to take of in fields not covered by, eg, arxiv.org.

Another class of costly labor is copy editing and getting details right. This means things like tracking down every single citation and making sure they exist and are formatted correctly, spelling corrections, making graphs color-blind accessible, etc. This is hard to automate with current tooling; a lot of time and resources are currently being spent on next-generation open source publishing/manuscript pipelines to make this sort of thing more machine-verifiable, but it's going to be a long struggle to get authors (legitimately busy and distracted and over please-do-everything-this-new-disruptive-way-please-ed) to change their established field-specific workflows.

bnewbold··on Archive.org and California to start a data sharing and preservation project
I can't speak formally for the Internet Archive, but the existing content and services are not going to disappear overnight: funding comes from several sources, thought has been put in to organizational structure, and things have been designed to keep core access and preservation infrastructure running with minimal cost and effort (eg, if the economy tanks).

Getting the content coverage people sometimes assume we already have is another matter. Additional funding (thanks for you donation!) go towards additional crawling and keeping up with the endless treadmill of media types and protocols. Eg, headless browser crawling development and deployment to capture javascript-heavy sites (https://github.com/internetarchive/brozzler); this is much more expensive than "classic" crawling.

For more on increasing storage costs and the under-funded state of web archiving in general, I recommend David Rosenthal's blog, eg:

https://blog.dshr.org/2018/05/longer-talk-at-msst2018.html

https://blog.dshr.org/2014/03/the-half-empty-archive.html

Far more effective and robust than hoping the archive is "suck it up for us" is to upload snapshots/dumps/exports yourself! Anybody can create an archive.org account and upload content (recommend https://github.com/jjjake/internetarchive over the HTML form), within reasonable limits. Obviously, care needs to be taken to remove sensitive (and personal) information first.

bnewbold··on Designing the Wayback Machine Loading Animation
Well, I oversimplified a bit. The "tarball" (WARC or ARC file) is a single large file with individually compressed HTTP response objects concatenated together. The index stores the byte offset and length (compressed) into this file, so an HTTP/1.1 range request can pull only the data needed. So in theory it's efficient. I don't work on that team directly and haven't looked in to exactly what the performance latency bottleneck is, sorry for the misleading response.
bnewbold··on Designing the Wayback Machine Loading Animation
The Wayback backend is much simpler than one might imagine. The index lookup ("where is the data for this URL at this timestamp?") is called the CDX API and is pretty fast, given that it's basically looking up a line in a sorted many-terabyte text file. Slower are the processes that extract the original HTTP response from what is effectively a giant multi-GB tarball sitting on an archival-grade (not IOPS-optimized) spinning disk, and the code that parses HTML and/or javascript and re-writes "embed" URLs for playback. Another speed limit is that we serve everything from our datacenters in the bay area with no CDN for most content, which you'll really notice when connecting from outside North America.

The priority is to have as much data as accessible as possible, and for the same cost we can crawl and store far more web content on slow spinning disks than with SSDs or large RAM caches. Most RAM on storage nodes, which would otherwise be used as disk cache, gets used for derive tasks (like OCR or video conversion) or crawling/crunching tasks. That being said, if the service is so slow as to be unusable then there's no point operating it in the first place; hopefully we can get the latency a it lower.

bnewbold··on Net Neutrality Day of Action: Help Preserve the Open Internet
To state the obvious, this proposal is antithetical to the concept of network neutrality.

Also, the belief that there would be "competitors in the user's area that don't play stupid games" that offer comparable services (eg, within a factor of 5x bandwidth-per-dollar value) seems to be a misunderstanding of the utility/monopoly/duopoly economics in play in many regions of the USA.

bnewbold··on Ask HN: Who is hiring? (July 2017)
Internet Archive | San Francisco, CA | PM, SRE, Book Curator | Full Time, ONSITE

The Internet Archive is a US-based non-profit which has been backing up the web and pursuing "Universal Access to All Knowledge" for more than 20 years. Some of our larger projects include the Wayback Machine, a national news TV archive, openlibrary.org, and archive-it.org. We own an operate all of our digital compute and redundant storage hardware, including our data center real estate, and run a predominantly free/open source software stack. We have some remote employees, but the majority of engineers work full time from our beautiful headquarters in the Richmond district of San Francisco (the roles below are on-site, though exceptional and travel-flexible remotes are welcome to apply anyways).

As of July 2017 we are actively hiring for:

- Manager Site Reliability & Infrastructure

- Hardware Design Engineer, for our custom book scanning pipeline

- Curator of Books & Linked Data, who will improve our book cataloging, metadata management, and scan prioritization process

- Project Manager for an existing 5-engineer team developing and operating a successful production Python/Django web application (no listing for this yet, contact me)

- ... and more!

The Internet Archive is an equal opportunity employer and additionally encourages applicants from all educational and career backgrounds. As a mission-driven non-profit we do not compete with Silicon Valley market rates, but we do pay a livable Bay Area salaries, provide full benefits including paternal leave, and our employees have the freedom to discuss their work publicly. Our funding is stable, we have weathered several tech bubbles, and intend to be operating well through the 21st century.

Applications and more information at http://archive.org/about/jobs.php, or contact me directly and I will route your message internally.

bnewbold··on Leibniz – A Digital Scientific Notation
"Scientific models are much simpler than most software"

I've been noodling with some similar ideas around collaborative editing of interconnected declarative models, and I don't think the above is necessarily true. A diagram of the Linux Kernel might be even more complex than a diagram of some wild systems biology[0], but surely there is ever more complexity to discover. The tools for massively distributed digital collaboration seem much more polished in software than scientific fields today. I hope the author's work is a step in the right direction!

For similar Scheme-y computer algebra stuff, see Sussman and Wisdom's work in "Structure and Interpretation of Classic Mechanics" and follow-ons.

[0]: https://s-media-cache-ak0.pinimg.com/originals/35/98/92/3598...

bnewbold··on AlphaGo Beats Lee Sedol in Final Game
Nothing! This is already very much a "thing" in the Go community. Bots like GNU Go, CrazyStone, and Zen are all running under multiple accounts on most of the popular online Go servers (KGS, etc). There are enough bots of differing age and ability that one can either hop up a chain of different bots or try to configure a strong bot into a weak configuration. The GNU Go bot is also downloadable free software and is frequently integrated into, eg, mobile apps (it is, however, old and not as strong as other bots, I think around 6k level). The game of Go also has a wonderful and essential handicap system to allow players (human or bots) of differing abilities (within a reasonable range; eg not possible for a novice to play an even game against Lee Sedol even with 9 stones).

As far as I can tell the vast majority of amateur players play against bots online and review games to improve their skills. It would be nice if it was easier to select a bot with a given skill rating, but you can figure this out pretty easily by playing some games or reading up on bots. Playing against a skilled human who cares about your advancement is still the best way to advance though, in my opinion. Getting good feedback on your mistakes and style of play is extremely helpful.

bnewbold··on Cleanflight Flight Control Software
Hmmm, looks like some license cleanup is needed:

> modified version of StdPeriph function is located here. > TODO - what license does apply here? > original file was lincesed under MCD-ST Liberty SW License Agreement V2 > http://www.st.com/software_license_agreement_liberty_v2

From "ultimate-liberty-v2.txt":

> 4. This software, including modifications and/or derivative works of this software, must execute solely and exclusively on microcontroller or microprocessor devices manufactured by or for STMicroelectronics.

Maybe switch to libopencm3 for basic hardware support?

← PreviousPage 2 of 3Next →