To clarify I'm not asking about HN itself but articles linked from HN.
As you said the HN api is great and there are at least 2 existing published crawls of it that help a lot.
As you said the HN api is great and there are at least 2 existing published crawls of it that help a lot.
I might not have a clear picture of what you're looking for, but items of type "story" returned by the HN API do have a URL field, which I believe correspond to submitted links.
You can scrape the text field of comment items, but that takes a bit more work.
I don't understand your question. If you have the URL, you just GET it, like any regular URL? Is there something that I'm missing?