Show HN: Puppeteer scripts in the Browser, DevTools on remote pages
pptrconsole.com
pptrconsole.com
I dislike when apps/services drop you into the experience without any background, or some scenarios/use-cases/reasons why I should consider using the app/service...or even something as basic as an intro paragraph describing what it is that I'm seeing. Clearly, I'm missing the point; so would be very appreciative if someone would kindly provide a little description/context. Thank you!
- DevTools doesn't display the viewport. I'm not sure if this is due to a change in the latest Chrome to which I just updated (~90) or because I broke my serving of it by updating it. A workaround will be serving a static snapshot of the devtools front-end rather than just (simply, as I'm doing right now) pulling it out of Chrome's RDP endpoint each time. This may take some time to do.
- DevTools doesn't seem to work on iOS (as I've tested it, Safari or Chrome).
- There are many more issues, and a lot, but not all, of them are edge cases but they'll be fixed eventually.
More bug reports, UI/UX tips and advice, and other feedback are very welcome! Unfortunately the whole app is not open source but some parts are open source, namely, the virtualized browser[0], and the devtools-front-end[1].
When I posted a remote browser project similar to this Show HN, some time ago, Joel mentioned on the thread he might like to work together and to email him. I mailed him, then followed up, but he never replied. I guess he didn't mean what he said. Maybe he got told he wasn't allowed to work with me!
But looking at this link, this look fantastic! I love the design of browserless overall, and how Joel has built it into a thrumping business. Totally thrumping!
Architecturally, it looks to be a custom reworking of chrome-devtools-frontend (I guess the viewport is adapted or spliced directly out of the devtools frontend viewport, but maybe nto) and integrating with the puppeteer window, at least on the front-end.
And it's a high coincidence that we both released a similar thing within 1 month. I did not see his thing lately, and did not directly reference it consciously as inspiration for what I built, but I believe I've seen this thing before, maybe about a year ago (if it existed in some form then? or maybe that was another site) -- so maybe there was some inspiration about putting all these parts together from that, operating for me.
I love the look of his puppeteer debugger console, his is much more polished that mine :)
Some things I noticed about https://chrome.browserless.io
- after using it for a bit the actual tab in my browser crashed with error "STATUS_BREAKPOINT"
- Does not work on Mobile, it loaded, and ran the script, but I couldn't type into the editor on my Android phone
- I tried to bork it with while(true) page.browser().newPage() it seemed OK
- I tried to bork it with a bandwidth speedtest, it seemed OK
- No multiple tabs
- No back buttons, so I needed to go to console to history.back()
- No paste into page, no right-click context menu
- Did not raise a pseudo-modal for remote page modal dialogs, I assume it silently closes them. I.e. https://infosimples.github.io/detect-headless/
- Fails (or passes, depending on your perspective) the "Are you Headless?" test at https://arh.antoinevastel.com/bots/areyouheadless
- Did not raise a pseudo-modal for remote page file chooser dialogs, I assume it silently closes them. I.e. https://blueimp.github.io/jQuery-File-Upload/
- Appears to close the browser every time you press the |> play button. I think that makes debugging fluid.
- Does not raise a pseudo-modal for remote page basic auth modal dialogs, it seems to silently close them. I.e https://jigsaw.w3.org/HTTP/Basic/ (at https://jigsaw.w3.org/HTTP/)
All in all I love how the DevTools window is integrated with the page, even tho it does eat up screen real-estate, it's ok. I also love how product-focused Joel is and how the provided scripts are really laser focused on the use cases of his customers. That's what I think is Joel's biggest strength, how much of a good businessman he is, how focussed on his customers he is, and how he delivers that value for them. That's also what I think is my biggest weakness. Compared to how he's building products and features, I'm like wandering around in the dark shining my flashlight on whatever looks interesting to me. I'm not so honed and focused toward satisfying a specific customer goal, I guess, and that's to my detriment, I think. Design is another weakness I have. But we shall see how things go!
Maybe I'll catch up to him. Maybe I'll overtake him! But I'm not sure that side of the business, "puppeteer automation at scale" is necessarily what I want to go into. We'll see what happens tho! :)
Meanwhile a product like this is good competitor to https://mightyapp.com/
But is mighty intended for use on a single computer? Or for running headless browser tests / reducing computing consumption from mass browser testing?
Seems the former, which seems an odd value prop and somewhat niche, considering the rent a performance increase vs build/upgrade equation.
- This is webp images (where available, else jpeg) demand-streamed to the client over WebSocket binary channel. For webp, they are encoded server-side using cwebp from the original JPEGs.
- The servers in this demo are hosted in GCP us-west3 (Salt Lake City, Utah)
Also, some possible improvements to streaming I intend to look at are:
- stream frames as they are available, rather than when the client requests them (which it does at a small regular interval, or whenever the client performs some action)
- encode the raw PNG frames to h264 with ffmpeg
- use Chrome Extension desktopCapture API (similar to navigator.mediaDevices.getDisplayMedia) with xvfb (no headless) and send the resulting stream either through the server, or p2p using WebRTC
I initially didn't develop it to handle high framerates or high quality (it's more of a debug tool, and delivery system for a web scraping app), but people are requesting this.
MightApp has been in beta for a long time. Handling massive streaming and combining it with interactivity, and making the economics work, is not trivial.
Despite people's requests for better video quality, and my willingness to be responsive to those, I still have my heart firmly set on providing the best experience with the minimum amount of bandwidth and the lowest framerate and lowest quality (as in, resolution) possible. I just think this is more efficient, and will end up being more scalable, and it fits well for my initial web-scraping tool use-case. My biggest fear about this feature is lag to the point of unusability, which happens whenever the bandwidth is larger than the capacity. You get a backlog of inflight frames, and the usability goes to shit because everything takes x seconds to occur, and then you're behind anyway. I'm reminded of that nightmare scape video of people inside an Oculus delay chamber trying to pick up a feather from the floor. Glitch in the matrix.
I know about the abuse. The first time I launched something like this I kept getting hit with massive CPU spikes from innocuous looking pages. When I looked into it, I found it must have been some sort of crypto mining attempt.
Now I have scripts monitoring usage, and use cgroups, cpulimit, and process killing to prevent such resource abuse.
I mean except obvious Testing purpose usage, as many other tools give (Puppeteer, Selenium, WebDriver ...).
The first is a malicious website that uses ph/fishing or some other social engineering type attack. Unless that attack also included the need to download some sort of payload file to create a back door or whatever this additional barrier would not provide any protection against social engineering type attacks.
Another case is where there's a link in email to download a payload file that's like a word document or PDF that contains some sort of exploit that will you know initiate it take over on the user's computer. I guess a lot of them will be protected again by using a barrier like this because those files are downloaded to a separate server and then converted into images and then displayed back over the web so the client only receives an image of each page of that document and the document itself is never executed except that it's converted to images.
Another case is a malicious website that uses a browser zero day or other type of exploit that enables remote code execution or sandbox escape in a browser. In that case you probably be protected by using this barrier because the only thing you're getting from the remote website is pixels of screenshots and any exploit that occurs will execute on the remote server.
Here are the weaknesses is in the current setup. In order to successfully take over the remote server the exploit needs to escape the browser sandbox achieve remote code execution and achieve privilege escalation. I'm sure that's possible so it's not totally secure with respect to the server itself being pawned and then of course everybody who's using the server becomes vulnerable. So by no means is it perfect and from that point of view it's quite weak but from the point of view of the other cases it's quite a strong barrier.
On the server every browser instance runs in its own temporary user. That user is no login user whose processes and home directory only exists for the duration of the browser session. And that user has limits on the amount of memory and CPU and disk space it can use. So from that point of view this system is relying on existing Unix process isolation and user privilege isolation to provide a layer of security on the server.
In addition to that the controller server of every browser instance is a separate process for each client, on that separate process is run under that temporary user assigned to that client. So each controller server only how's the privileges of that temporary user, it only talks to that client and to that browser and every public API endpoint is authenticated. In addition the internal chrome remote debugging ports are prevented from being accessed from the outside web and other iptables rules drop packets on Google compute engine internal network endpoints.
The reason I didn't wrap each browser instance inside a docker container to add yet another barrier and require the chaining together of yet another exploit is because it was a trade-off and I don't think it was worth it between the extra effort and overhead of running inside docker versus running inside a lightweight operating system sandbox using cgroups and temporary users.
With that being said I have a couple of ideas on improving security if that's needed. Firstly rather than using the latest Linux server I could use SELinux and add additional hardening. Secondly I could like I mentioned before wrap each browser instance in its own docker container to require a container escape exploit to be chained together as well. Thirdly I could fall back on the security of Google cloud to use instance level isolation and require a hypervisor escape to compromise anything but the temporary virtual private server. In other words each browser instance could run on its own tiny VPS.
The con side of the trade-off equation in each of these has been: no demand for it, and high cost in terms of application performance, implementation time, additional complexity and or money.
For instance running each browser on its own tiny VPS would incur both an additional complexity and implementation cost as well as cut the performance because on average the performance of a tiny VPS is less good than the amortized same amount of processes on a much larger machine. In addition to that VPS bandwidth scales with number of processors, and bandwidth is a key component of the performance equation in this application. Even at the same raw number of processes and amount of memory than a larger machine I guess the money cost would be higher to run multiple tiny machines, and it would certainly be higher when each instance was enlarged to match the amortized performance available using a larger machine.
hopefully that answers your question I'm sorry if it wasn't totally related to what you're wanting to know. Tho at the same time I'm pretty sure someone else'll find that information useful. Thanks for asking.
How much time did you spend developing all this? And resource wise how hungry is it? (Servers, Memory, CPU...)
Edit: Wait a sec, I just saw this is part of open source demo of ViewFinderJS, bit confused now, is it commercial or opensource? How do you prevent someone of other 8000ish will not create a bit same and steal your thunder so to say?
2 years development in total so far. Not totally full time since roughly end of 2018
It's very resource lightweight for what it is. This is mainly due to how node and chrome work. When they don't do anything, they literally don't do anything. And chrome headless is normally pretty low CPU for everything. The node servers are very simple, there's not much that's thread blocking.
It's a commercial service, but a couple of components are open-source. The browser you use in this demo and in the paid service is based on the ViewFinderJS open-source project, but it's more advanced, and is like the "pro" version of that.
That's on my to-do list tho, I think I can make a good wrapper for the existing extension apis using only devtools API. Some things will be unsupported of course, but it's probably not a problem because many extensions use a diverse set of APIs but not all