Here’s another Christmassy alternative: https://display.archive.org/xmas
I’m one of the makers of OneMillionScreenshots.com and I’m currently working on an update to it.
2,564 karma · joined November 22, 2009
product at: https://urlbox.io
twitter: http://twitter.com/jot
podcast on customer interview examples: https://empathydeployed.com/
coworking space in Brighton, UK: http://theskiff.org
event: http://SussexFounders.com
Here’s another Christmassy alternative: https://display.archive.org/xmas
I’m one of the makers of OneMillionScreenshots.com and I’m currently working on an update to it.
I’m one of the makers of this - thank you for posting!
We built it in early 2024 and it’s due an update.
While the main visualisation is currently out of date we’ve been grabbing screenshots monthly for the last 18 months. There’s a a free (no key required) API to all the data at https://ScreenshotOf.com
I’d love to hear any ideas you have for improving it.
It’s one of the top reasons larger organisations prefer to use hosted services rather than doing it themselves.
It’s rare that anyone makes that kind of mistake. It probably helps that our rate limits are relatively low compared to other APIs and we email you when you get close to stepping up a tier. If you did make such a mistake we would, like all good dev tools, work with you to resolve. If it happened a lot we might introduce some additional controls.
We’ve been in this business for over 12 years and currently have over 700 customers so we’re fairly confident we have the balance right.
In my experience our customers are more worried about having the service stop when they hit the limit of a tier than they are about being charged a few more dollars.
- https://browserless.io - low level browser control
- https://scrapingbee.com - scraping specialists
- https://urlbox.com - screenshot specialists*
They’re all profitable and have been around for years so you can depend on the businesses and the tech.
* Disclosure: I work on this one and was a customer before I joined the team.
If you also want to grab an accurate screenshot with the markdown of a webpage you can get both with Urlbox.
We have a couple of free tools that use this feature:
His book “Hourly Billing is Nuts” is particularly good: https://jonathanstark.com/hbin
I had so much fun with this with my 7 year old. Was super easy to take their art and make games with it. You can start by editing one of the many example games it comes with.
With some configuration you can get most of the way there.
The challenge there is that the content is in an iframe.
If you get the URL used for the iframe you can get the content: https://url2text.com/u/kJWaZY
But that's frustrating as it requires two steps.
We might be able to help you get the content from URLs like these in one step. We have quite a bit of power in the Urlbox API that url2text isn't using.
Drop us an email: support@urlbox.com and we'll see what we can do.
Sorry it's not clearer but you can skip the screenshot in the Urlbox API if you want to with:
curl -X POST \
https://api.urlbox.io/v1/render/sync \
-H 'Authorization: Bearer YOUR_URLBOX_SECRET' \
-H 'Content-Type: application/json' \
-d '
{
"url": "example.com",
"format": "md"
}
'
Here's the result of that:
https://renders.urlbox.io/urlbox1/renders/5799274d37a8b4e604...Sorry the pricing isn't a good fit for you. Urlbox has been running for over 11 years. We're bootstrapped and profitable with a team of 3 (plus a few contractors). We're priced to be sustainable so our customers can depend on us in the long term. We automatically give volume discounts as your usage grows.
We were lucky to build this on a mature API that already solves loads of the edge cases around rendering different kinds of pages.
I built a similar tool last year that doesn't have those features: https://url2text.com/
Apologies if the UI is slow - you can see some example output on the homepage.
The API it's built on is Urlbox's website screenshot API which performs far better when used directly. You can request markdown along with JS rendered HTML, metadata and screenshot all in one go: https://urlbox.com/extracting-text
You can even have it all saved directly to your S3-compatible storage: https://urlbox.com/s3
And/or delivered by webhook: https://urlbox.com/webhooks
I've been running over 1 million renders per month using Urlbox's markdown feature for a side project. It's so much better using markdown like this for embeddings and in prompts.
If you want to scrape whole websites like this you might also want to checkout this new tool by dctanner: https://usescraper.com/
If you'd rather do it for free weasyprint[2] is the best open source alternative.
Another more affordable option you might want to consider is Urlbox[3]. (Disclosure: I work on this)
Urlbox's rendering engine is based on Chrome. It's been refined over the last 11 years to render pages as images or PDFs[4] that look great. I was a customer for 5 years before I joined the team. Everything we'd tried before Urlbox was a disappointment.
Urlbox probably can't match the power of either Onedoc or DocRaptor, but pricing starts at less than $0.01 per document and drops significantly with scale. If your PDF looks great when saving as PDF in Chrome it should look identically brilliant with Urlbox.
[1]: https://docraptor.com [2]: https://weasyprint.org [3]: https://urlbox.com [4]: https://urlbox.com/html-to-pdf
It's primarily purpose is to render screenshots full-page or limited to viewport or an element. To do that well as it does the HTML has to be rendered perfectly first.
It's not as cheap as other solutions but we have customers who render millions of pages per month with us. They value the accuracy and reliability that's come from over a decade of refinements to the service.
Larger projects can request preferential pricing based on the specifics of the kinds of pages they are rendering.
I send the URLs I want scraped to Urlbox[0] it renders the pages saves HTML (and screenshot and metadata) to my S3 bucket[1]. I get a webhook[2] when it's ready for me to process.
I prefer to use Ruby so Nokogiri[3] is the tool I use for scraping step.
This has been particularly useful when I've want to scrape some pages live from a web app and don't want to manage running Puppeteer or Playwright in production.
Disclosure: I work on Urlbox now but I also did this in the five years I was a customer before joining the team.
[0]: https://urlbox.com [1]: https://urlbox.com/s3 [2]: https://urlbox.com/webhooks [3]: https://nokogiri.org
His daily email list gives me regular reminders of how to improve in sales and pricing.
His books are brilliant too. Start with Hourly Billing is Nuts: https://jonathanstark.com/hbin
Urlbox helps web developers render the web with precision. We've been focused on generating screenshots, images and PDFs from HTML or URLs for over a decade. Our customers include over 500 design or compliance led organisations. They depend on us to get the intricacies of browser rendering right so they can focus on their core products and services.
We're bootstrapped, profitable and ready to add a third full-time engineer to our team. Our stack is primarily TypeScript. It's a bonus if you're also interested in learning how to orchestrate and scale headless browsers on our Kubernetes clusters. There's also opportunities to create/maintain libraries and SDK's in a range of other languages.
We're excited to hear from people early in their tech career as well as more experienced folk.
Read more: https://urlbox.com/jobs/typescript-developer
A large part of it is dedicated to old tech donated by locals over the years. Highly recommended if you’re in the area.
He also wanted 3D but once we added some great looking dinosaur sprites (generated with DALL E) he was fully engaged. I'm a ruby developer and it's been a joy learning the differences between web and game dev.
Knowing that we can easily distribute on mobile platforms, web, Steam and Switch once we're ready has kept us coming back.
One situation where you should always have screenshots in your documentation is when the documentation is for a screenshot API.
We recently updated our docs to do a better job of this: https://www.urlbox.io/docs/options#url-examples