HNHacker News
TopNewBestAskShowJobs

portInit

96 karma · joined October 12, 2015

nic (@) crul.com
submissionscomments
portInit··on Show HN: Crul – Query Any Webpage or API
Appreciate you taking the time to play around with crul and share your thoughts! They're incredibly valuable.

1. Although it has limitations, were you able to try the normalize command? https://www.crul.com/docs/queryconcepts/api-normalization

2. Will need to think about this some more.

3. Although we don't yet handle pdfs, the rest of the flow is one we aiming to accomplish with crul. Other than pdf, the pieces should be there and would love to understand this further

portInit··on Show HN: Crul – Query Any Webpage or API
We certainly have, it is admittedly quite new to both of us, so we have been exploring the best way to introduce something like this, as well as tying up technical bits. The open core model with commercial features is certainly appealing.

Open to any perspectives on this.

portInit··on Show HN: Crul – Query Any Webpage or API
Thank you! Will make this more clear.
portInit··on Show HN: Crul – Query Any Webpage or API
Thank you! The query below would get you the full page text, although it likely wont be too legible. I'll read up on readability JS some, it looks quite magic and could possibly be added in.

open https://news.ycombinator.com/item?id=34970917 || filter "(nodeName == 'BODY')" || table innerText

Often you can use crul to discover the html structure in the results table. With a `find "text string"` to filter rows and then a filter on the column values that identify the desired elements.

portInit··on Show HN: Crul – Query Any Webpage or API
I'm having trouble replicating. Although I did get a prompt from Safari to allow downloads from crul.com. The screenshot looks like it is redirecting from the button, is that what you are seeing - rather than the button triggering a download?
portInit··on Show HN: Crul – Query Any Webpage or API
Yeah exactly. Living up to your username :) nice find! Note this is a global default, unlike domain policies which are associated with a domain.
portInit··on Show HN: Crul – Query Any Webpage or API
Thanks for the heads up. Could you share what browser you are using? And are you downloading from https://www.crul.com/account?
portInit··on Show HN: Crul – Query Any Webpage or API
You would need to do a bit filtering to get the exact text you need. For example, to get the text for this post you could run:

open https://news.ycombinator.com/item?id=34970917 || filter "(attributes.class == 'toptext')" || table innerText

portInit··on Show HN: Crul – Query Any Webpage or API
We hand wrote a number of integrations, sometimes it was a simple as reusing a schema with slightly different values, we are also using the awesome https://www.benthos.dev/!
portInit··on Show HN: Crul – Query Any Webpage or API
That's unfortunate. A quick look suggests ISP but maybe not with the hotspot. Will do a little more digging.
portInit··on Show HN: Crul – Query Any Webpage or API
Thank you! Docker image has everything for functionality, there's just an update check that pings our end to check for updates.
portInit··on Show HN: Crul – Query Any Webpage or API
Sorry to hear that - we do need to think about this. It's our first pass at product tiers and features and we may need to adjust.

Scheduling and Domain Policies were the main features we chose to gate initially as they don't affect core functionality other than performance and deployment.

portInit··on Show HN: Crul – Query Any Webpage or API
We'll have to see how close our current postgres integration is! Would like to understand more and will reach out.
portInit··on Show HN: Crul – Query Any Webpage or API
Thanks for this, we're still trying to figure out these details ourselves.

Related to the quote, we've seen interest for API data wrangling, where prepackaged data feeds can be cumbersome to edit, or other implementation details become challenging like credential management, domain throttling, scheduling, checkpointing, export, etc.

It's also interesting for webpage data when you need to use a particular page as an index of links to filter and crawl. We've tried to build an abstraction layer around that.

Initially, we were mainly focused on webpages, and wanted to bypass the visuals of the browser and use a headless browser to fulfill network requests, render js, etc. then convert the page to flat table of enriched elements. With the APIs as another data set there is some work for us to do around language.

We're now trying to figure out what workflows are most relevant for crul to optimize around, as well, honestly - we just built what we thought was cool. Some features/workflows will certainly be more straightforward with existing tools and software - especially for a technically savvy user.

portInit··on Show HN: Crul – Query Any Webpage or API
Ah! We didn't quite get an html table command in this release but it will be in the next one.

Here's a query that shows an option, but the table command will be far more straightforward.

open https://www.w3schools.com/html/html_tables.asp --dimension || filter "(nodeName == 'TD')" || groupBy boundingClientRect.top || table _group.0.innerText _group.1.innerText _group.2.innerText

portInit··on Show HN: Crul – Query Any Webpage or API
Thank you! Really means a lot to us
portInit··on Show HN: Crul – Query Any Webpage or API
Agreed, caching does come with its own set of quirks and mind-numbing bugs, crul does have a caching override flag at the command/stage level which alleviates some of this: https://www.crul.com/docs/queryconcepts/common-flags#--cache

Your provided links are interesting and something for us think about some more. Honestly, I would be quite interested in hearing more about your experiences.

portInit··on Show HN: Crul – Query Any Webpage or API
Yes that is possible, although we currently have the scheduler set as an enterprise feature. We should look into a free trial, I enjoyed the flow of the EasyDataTransform installation with the free trial option.
portInit··on Show HN: Crul – Query Any Webpage or API
Would love to chat and try a few use cases together! At first glance I see you can drag in csv files, which could be generated by crul and either manually downloaded or scheduled and written to the filesystem.

Crul is a really easy way of populating tools that need data, whether it's just a one time thing for a demo/static data set or a scheduled data feed, so this kind of usage makes sense to us.

portInit··on Show HN: Crul – Query Any Webpage or API
Yeah the default domain throttle policy is 1 req per second per domain. Configurable through domain policies https://www.crul.com/docs/features/domain-policies - although currently an enterprise feature.

We found that it becomes too easy to break API request limits or spam a website otherwise.

However if you rerun that query it should load pretty instantly due to the caching layers, so the actual querying/filtering of the data part is smoother/faster.

portInit··on Show HN: Crul – Query Any Webpage or API
Thank you and your comment really warms our hearts. It's been fun building in the "cave" but comes with self doubt.

We've built using a microservice architecture to allow us to scale out the parts that need to scale, mainly the workers, which interact with a queue, although we'll need to move from the parts that are currently nodejs for some perf gains and a smaller footprint. All those microservices are consolidated for the desktop variants.

Network concurrency is mostly throttled by our domain policy manager (named "gonogo" - lol) at 1 req/per outbound domain a sec. It's a little slow for a default, but also configurable and provides a nice guardrail for api request limits, etc. Overall async networking has been quite tricky, esp with retries, etc., and we're still iterating on it. Agreed on the profiling difficulties.

portInit··on Show HN: Crul – Query Any Webpage or API
This makes sense and I am curious about this. Was there consistency between those 1k client sites or were they all rather different? Mind if I reach out?
portInit··on Show HN: Crul – Query Any Webpage or API
lol - yeah. https://www.imdb.com/title/tt0085811/?ref_=tt_urv

We've been thinking of ways to make the pronunciation a bit clearer, maybe a mascot or something. Open to ideas!

portInit··on Show HN: Crul – Query Any Webpage or API
It's a tricky question. Part of it is looking at APIs as the main source of data for scheduled queries/data feeds.

Crul sort of operates as a text only browser when interacting with a single page at a time, but when you expand and open up multiple tabs it becomes a little more challenging. We have the concept of domain policies which allow you to control how quickly/slowly you access something. There are also some puppeteer level options that could be relevant, even a headful toggle.

We have not invested too much time into this yet as we focused on getting the core functionality working. We think there are use cases (particularly with APIs) that don't run into this problem, but if it comes up more often we'll come up with some options.

portInit··on Show HN: Crul – Query Any Webpage or API
At this point we're considering it a foundational concept to build around - web content changes, so our best option currently is to make the query as easy as possible to change, and alert when things break.

We have done some preliminary work in some AI or other intelligence for pattern recognition to be able to handle structural changes better, but still have lots of work.

But the expanding and querying concepts also make a lot of sense with APIs, which tend to be a little more stable.

portInit··on Show HN: Crul – Query Any Webpage or API
Will absolutely reach out! Our experience has been that just getting data is often really challenging, so we've really focused on that piece, and being able to easily share with destinations that are purpose built for analytics, viz, etc.

Thanks for checking it out!

portInit··on Show HN: Crul – Query Any Webpage or API
Thanks! The find (https://www.crul.com/docs/commands/find) command works really well if you are trying to construct a filter expression and just want to quickly look for results containing a particular string so you can see the defining attributes/column+row values.

There's a short writeup of this pattern here: https://www.crul.com/docs/examples/how-to-find-filters

portInit··on Starting a Tiny Indoor Garden
Much appreciated! I'm only scratching the surface of what is possible but it was so surprisingly easy for someone with minimal starting knowledge to get something going.

I can definitely start to see that the nutrient, ph, light details are going to be important for more consistency and better yields but having had this simple setup work ok has certainly been motivating to expand my knowledge.

Will checkout that channel. Looks like a fun rabbit hole :)