HNHacker News
TopNewBestAskShowJobs

jot

2,564 karma · joined November 22, 2009

Spent my first 1000 days on Hacker News as: http://news.ycombinator.com/user?id=madmotive

product at: https://urlbox.io

twitter: http://twitter.com/jot

podcast on customer interview examples: https://empathydeployed.com/

coworking space in Brighton, UK: http://theskiff.org

event: http://SussexFounders.com

submissionscomments
jot··on One million (small web) screenshots
So good to see this different approach! The clustering looks really cool and love that the focus is not on the most popular websites.

Here’s another Christmassy alternative: https://display.archive.org/xmas

I’m one of the makers of OneMillionScreenshots.com and I’m currently working on an update to it.

jot··on One Million Screenshots
Thank you! This is exactly what I want to do with this. I have embeddings for all the images but hadn’t figured out the last step to getting them onto a grid like that. Can’t wait to try this.
jot··on One Million Screenshots
Hi HN,

I’m one of the makers of this - thank you for posting!

We built it in early 2024 and it’s due an update.

While the main visualisation is currently out of date we’ve been grabbing screenshots monthly for the last 18 months. There’s a a free (no key required) API to all the data at https://ScreenshotOf.com

I’d love to hear any ideas you have for improving it.

jot··on Gridfinity: The modular, open-source grid storage system
There’s a filament saving variant where you can use toilet rolls or other waste cardboard for the walls: https://www.printables.com/model/880256-cardboard-gridfinity...
jot··on Postgres Just Cracked the Top Fastest Databases for Analytics
They list "Managed Iceberg tables" top of list of features on that page.
jot··on Postgres Just Cracked the Top Fastest Databases for Analytics
How is this different from Crunchy Warehouse which is also built on Postgres and DuckDB?

https://www.crunchydata.com/products/warehouse

jot··on Show HN: An API that takes a URL and returns a file with browser screenshots
Too many developers learn this the hard way.

It’s one of the top reasons larger organisations prefer to use hosted services rather than doing it themselves.

jot··on Show HN: An API that takes a URL and returns a file with browser screenshots
That’s right. On our standard self-service plans we automatically charge a better rate as volume increases. You only pay the difference between tiers as you move through them.

It’s rare that anyone makes that kind of mistake. It probably helps that our rate limits are relatively low compared to other APIs and we email you when you get close to stepping up a tier. If you did make such a mistake we would, like all good dev tools, work with you to resolve. If it happened a lot we might introduce some additional controls.

We’ve been in this business for over 12 years and currently have over 700 customers so we’re fairly confident we have the balance right.

jot··on Show HN: An API that takes a URL and returns a file with browser screenshots
I’m sure we can do better here.

In my experience our customers are more worried about having the service stop when they hit the limit of a tier than they are about being charged a few more dollars.

jot··on Show HN: An API that takes a URL and returns a file with browser screenshots
If you’re worried about the security risks, edge cases, maintenance pain and scaling challenges of self hosting there are various solid hosted alternatives:

- https://browserless.io - low level browser control

- https://scrapingbee.com - scraping specialists

- https://urlbox.com - screenshot specialists*

They’re all profitable and have been around for years so you can depend on the businesses and the tech.

* Disclosure: I work on this one and was a customer before I joined the team.

jot··on Show HN: HTML-to-Markdown – convert entire websites to Markdown with Golang/CLI
We do that with Urlbox’s markdown feature: https://urlbox.com/extracting-text
jot··on Show HN: HTML-to-Markdown – convert entire websites to Markdown with Golang/CLI
This is great!

If you also want to grab an accurate screenshot with the markdown of a webpage you can get both with Urlbox.

We have a couple of free tools that use this feature:

https://screenshotof.com https://url2text.com

jot··on Day Rates (2023)
I highly recommend reading Jonathan Stark’s material on this topic. It changed the way I think about billing for software projects and advisory work.

https://jonathanstark.com/

His book “Hourly Billing is Nuts” is particularly good: https://jonathanstark.com/hbin

jot··on Ask HN: Platform for 11 year old to create video games?
DragonRuby https://dragonruby.org/

I had so much fun with this with my 7 year old. Was super easy to take their art and make games with it. You can start by editing one of the many example games it comes with.

jot··on Show HN: I made a tool to clean and convert any webpage to Markdown
This is worth having a look at: https://mixmark-io.github.io/turndown/

With some configuration you can get most of the way there.

jot··on Show HN: I made a tool to clean and convert any webpage to Markdown
Our tool sadly also fails on this: https://url2text.com/u/KYkpBj

The challenge there is that the content is in an iframe.

If you get the URL used for the iframe you can get the content: https://url2text.com/u/kJWaZY

But that's frustrating as it requires two steps.

We might be able to help you get the content from URLs like these in one step. We have quite a bit of power in the Urlbox API that url2text isn't using.

Drop us an email: support@urlbox.com and we'll see what we can do.

jot··on Show HN: I made a tool to clean and convert any webpage to Markdown
Thanks!

Sorry it's not clearer but you can skip the screenshot in the Urlbox API if you want to with:

  curl -X POST \
    https://api.urlbox.io/v1/render/sync \
    -H 'Authorization: Bearer YOUR_URLBOX_SECRET' \
    -H 'Content-Type: application/json' \
    -d '
  {
    "url": "example.com",
    "format": "md"
  }
  '
Here's the result of that: https://renders.urlbox.io/urlbox1/renders/5799274d37a8b4e604...

Sorry the pricing isn't a good fit for you. Urlbox has been running for over 11 years. We're bootstrapped and profitable with a team of 3 (plus a few contractors). We're priced to be sustainable so our customers can depend on us in the long term. We automatically give volume discounts as your usage grows.

jot··on Show HN: I made a tool to clean and convert any webpage to Markdown
Last time I tried readability it worked well with articles but struggled with other kinds of pages. Took away far more content than I wanted it to.
jot··on Show HN: I made a tool to clean and convert any webpage to Markdown
It's not easy working around things like that. But here's how it could work: https://url2text.com/u/wYVake

We were lucky to build this on a mature API that already solves loads of the edge cases around rendering different kinds of pages.

jot··on Show HN: I made a tool to clean and convert any webpage to Markdown
Great idea to offer image downloads and filtering with GPT!

I built a similar tool last year that doesn't have those features: https://url2text.com/

Apologies if the UI is slow - you can see some example output on the homepage.

The API it's built on is Urlbox's website screenshot API which performs far better when used directly. You can request markdown along with JS rendered HTML, metadata and screenshot all in one go: https://urlbox.com/extracting-text

You can even have it all saved directly to your S3-compatible storage: https://urlbox.com/s3

And/or delivered by webhook: https://urlbox.com/webhooks

I've been running over 1 million renders per month using Urlbox's markdown feature for a side project. It's so much better using markdown like this for embeddings and in prompts.

If you want to scrape whole websites like this you might also want to checkout this new tool by dctanner: https://usescraper.com/

jot··on Launch HN: Onedoc (YC W24) – A better way to create PDFs
It sounds like this is as advanced as DocRaptor[1]. They have what I consider to be the best PDF generation API, giving complete control over the documents you need to create. The pricing is similar.

If you'd rather do it for free weasyprint[2] is the best open source alternative.

Another more affordable option you might want to consider is Urlbox[3]. (Disclosure: I work on this)

Urlbox's rendering engine is based on Chrome. It's been refined over the last 11 years to render pages as images or PDFs[4] that look great. I was a customer for 5 years before I joined the team. Everything we'd tried before Urlbox was a disappointment.

Urlbox probably can't match the power of either Onedoc or DocRaptor, but pricing starts at less than $0.01 per document and drops significantly with scale. If your PDF looks great when saving as PDF in Chrome it should look identically brilliant with Urlbox.

[1]: https://docraptor.com [2]: https://weasyprint.org [3]: https://urlbox.com [4]: https://urlbox.com/html-to-pdf

jot··on Web Scraping in Python – The Complete Guide
Urlbox will save the whole page.

It's primarily purpose is to render screenshots full-page or limited to viewport or an element. To do that well as it does the HTML has to be rendered perfectly first.

It's not as cheap as other solutions but we have customers who render millions of pages per month with us. They value the accuracy and reliability that's come from over a decade of refinements to the service.

Larger projects can request preferential pricing based on the specifics of the kinds of pages they are rendering.

jot··on Web Scraping in Python – The Complete Guide
This is how I do it.

I send the URLs I want scraped to Urlbox[0] it renders the pages saves HTML (and screenshot and metadata) to my S3 bucket[1]. I get a webhook[2] when it's ready for me to process.

I prefer to use Ruby so Nokogiri[3] is the tool I use for scraping step.

This has been particularly useful when I've want to scrape some pages live from a web app and don't want to manage running Puppeteer or Playwright in production.

Disclosure: I work on Urlbox now but I also did this in the five years I was a customer before joining the team.

[0]: https://urlbox.com [1]: https://urlbox.com/s3 [2]: https://urlbox.com/webhooks [3]: https://nokogiri.org

jot··on I'm an engineer that needs to sell my services. Any good books on sales?
I recommend Jonathan Stark's writing on this: https://jonathanstark.com/

His daily email list gives me regular reminders of how to improve in sales and pricing.

His books are brilliant too. Start with Hourly Billing is Nuts: https://jonathanstark.com/hbin

jot··on Ask HN: Who is hiring? (February 2024)
Urlbox | TypeScript / Next.js / DevOps (k8s) | REMOTE (UK) | Full-time | £30K to £60K | https://urlbox.com

Urlbox helps web developers render the web with precision. We've been focused on generating screenshots, images and PDFs from HTML or URLs for over a decade. Our customers include over 500 design or compliance led organisations. They depend on us to get the intricacies of browser rendering right so they can focus on their core products and services.

We're bootstrapped, profitable and ready to add a third full-time engineer to our team. Our stack is primarily TypeScript. It's a bonus if you're also interested in learning how to orchestrate and scale headless browsers on our Kubernetes clusters. There's also opportunities to create/maintain libraries and SDK's in a range of other languages.

We're excited to hear from people early in their tech career as well as more experienced folk.

Read more: https://urlbox.com/jobs/typescript-developer

jot··on Martello Tower
The one nearest me has a fantastic museum [0] inside.

A large part of it is dedicated to old tech donated by locals over the years. Highly recommended if you’re in the area.

[0]: https://seafordmuseum.co.uk/

jot··on Ask HN: 9-yo son wants to build a game, I'm lost. What can I do?
Not 3D but my 7 year old and I have been having loads of fun with DragonRuby[0]

He also wanted 3D but once we added some great looking dinosaur sprites (generated with DALL E) he was fully engaged. I'm a ruby developer and it's been a joy learning the differences between web and game dev.

Knowing that we can easily distribute on mobile platforms, web, Steam and Switch once we're ready has kept us coming back.

[0]: https://dragonruby.org/

jot··on The Invisible Screen – An E-Paper Smart Display
I purchased one of these after seeing it posted to hn a few weeks ago. Very happy with it. Super easy to setup and just works. Maker has clearly put loads of work into refining the experience.
jot··on Should you add screenshots to documentation?
It's not easy keeping screenshots up to date.

One situation where you should always have screenshots in your documentation is when the documentation is for a screenshot API.

We recently updated our docs to do a better job of this: https://www.urlbox.io/docs/options#url-examples

jot··on I spent 3 years working on a coat hanger [video]
Cheers. That was thanks to Lukas: https://lw.works/en
Page 1 of 7Next →