HNHacker News
TopNewBestAskShowJobs

unlog

377 karma · joined February 26, 2020

submissionscomments
unlog··on Solid 2.0 RC: The Big <Reveal>
Yeah, like solid-universal (lets you target non-dom environments, state being derived and computed as little as possible, never again a glitch or disagreement because theres nothing to sync)
unlog··on Solid 2.0 RC: The Big <Reveal>
I keep wondering if the next step of Ryan should be to raise this a step further and bring it to anything-ui, not just the web. It's obviously a step forward, but that step could be mean a lot more than it really seems to be.
unlog··on Our commitment to Windows quality
Used Windows since forever because "it just worked for me". Last year switched to Fedora + Plasma, as I started to consider staying on windows was risky.

The feedback/forum tool, has been a thing for years. Submited many bugs that I wanted fixed, and always been ignored.

Thanks, but Im not looking back.

unlog··on Web Components: The Framework-Free Renaissance
Getting tired of their framework-free narrative.

What they are doing is backing in the browser, via specifications and proposals to the platform, their ideas of a framework. They are using their influence in browser makers to get away in implementing all of this experiments.

Web Components are presented as a solution, when a solution for glitch-free-UI is a collaboration of the mechanics of state and presentation.

Web Components have too many mechanics and assumptions backed in, rendering them unusable for anything slightly complex. These are incredible hard to use and full of edge cases. such ElementInternals (forms), accessibility, half-style-encapsulation, state sharing, and so on.

Frameworks collaborate, research and discover solutions together to push the technology forward. Is not uncommon to see SolidJS (paving the way with signals) having healthy discussions with Svelte, React, Preact developers.

On the other hand, you have the Web Component Group, and they wont listen, they claim you are free to participate only to be shushed away by they agreeing to disagree with you and basically dictating their view on how things should be by implementing it in the browser. Its a conflict of interest.

This has the downside that affects everyone, even their non-users. Because articles like this sell it as a panacea, when in reality it so complex and makes so many assumptions that WC barely work with libraries and frameworks.

unlog··on The time is right for a DOM templating API
Yep, `lit` is contaminating the browser API with their ideas just because their group of people writes the code for the browsers. They should be competing from the outside. Instead of pushing this kind of apis that only fit their mental models.
unlog··on Solidjs: Simple and performant reactivity for building user interfaces
SolidJS and dom-expressions are the best things that have happened in the front-end since React, it is influencing the whole ecosystem, from templating to Signals. It will be very, very hard to come up with better ideas, it may not be that popular, but it's leading the way.
unlog··on The Leningrad botanists who saved the first seed bank
> The uploader has not made this video available in your country

We really need a civilization changing event to rethink some stuff.

unlog··on Namespace JSX – Elements Table (various frameworks)
Made this table listing the typings used by SolidJS, Voby, Vue, Preact, React, Pota, VSCode-LSP and Chrome, for easily comparing type definitions between frameworks and the browser.

Note: There are a few inaccuracies, like the `<audio>` tag not including typings, that's because extended interfaces aren't resolved, but _most_ of the stuff is there.

unlog··on HTML Form Validation is underused
In an all honest reply, is that the people that writes these specifications, live disconnected from the reality, they don't use the stuff they specify. That stuff works for very simple things, but then when your forms evolve you realise you will be better off just writing the whole thing yourself.
unlog··on Show HN: Crawl a modern website to a zip, serve the website from the zip
If the URL is public you may post it here or in a GitHub issue, so I can take a look to what's wrong with it.
unlog··on Show HN: Crawl a modern website to a zip, serve the website from the zip
Thanks for sharing!
unlog··on Show HN: Crawl a modern website to a zip, serve the website from the zip
I think you are misunderstanding, your application is expected to give mostly 200s codes, if you get a 404, then a link is broken or a page misbehaving which is exactly why that page url is displayed on the console with a warning.
unlog··on Show HN: Crawl a modern website to a zip, serve the website from the zip
It tries to fetch a sitemap for in case there's some missing link. But it starts from the root and crawls internal links. There's a new mode added this morning for spa with the option `--spa` that will write the original HTML instead of the generated/rendered one. That way some apps _will_ work better.
unlog··on Show HN: Crawl a modern website to a zip, serve the website from the zip
It saves the generated/rendered html, but I have just added a `spa` mode, that will save the original HTML without modifications. This makes most simple web app work.

I have also updated the local server for fetching from origin missing resources. For example, a webapp may load some JS modules only when you click buttons or links, when that happens and the requested file is not on the zip, it will fetch it from origin and update the zip. So mostly you can back up an SPA by crawling it first and then using it for a bit for fetching the missing resources/modules.

unlog··on Show HN: Crawl a modern website to a zip, serve the website from the zip
let me know how that goes I am interested!
unlog··on Show HN: Crawl a modern website to a zip, serve the website from the zip
Modern websites execute JavaScript that render DOM nodes that are displayed on the browser.

For example if you look at this site on the browser https://pota.quack.uy/ and do `curl https://pota.quack.uy/` do you see any of the text that is rendered in the browser as output of the curl command?

You don't, because curl doesn't execute JavaScript, and that text comes from JavaScript. One way to fix this problem, is by having a Node.js instance running that does SSR, so when your curl command connects to the server, a node instance executes JavaScript that is streamed/served to curl. (node is running a web server)

Another way, without having to execute JavaScript in the server is to crawl yourself, let's say in localhost, (you do not even need to deploy) then upload the result to a web server that could serve the files.

unlog··on Show HN: Crawl a modern website to a zip, serve the website from the zip
> Why would someone crawl their own website?

My main use case is that the docs site https://pota.quack.uy/ , Google cannot index it properly. On here https://www.google.com/search?q=site%3Apota.quack.uy you will see some tiles/descriptions won't match what the content of the page is about. As the full site is rendered client side, via JavaScript, I can just crawl myself and save the html output to actual files. Then, I can serve that content with nginx or any other web server without having to do the expensive thing of SSR via nodejs. Not to mention, that being able to do SSR with modern JavaScript frameworks is not trivial, and requires engineering time.

unlog··on Show HN: Crawl a modern website to a zip, serve the website from the zip
Yes! You know, I was considering this the previous couple of days, was looking around on how to construct a `mhtml` file for serving all the files at the same time. Unrelated to this project, I had the use case of a client wanting to keep an offline version of one of my projects.

> Although UNIX philosophy posits that it's good to have many small files, I like your idea for its contribution to reduceing clutter (imagine running 'tree' in both scenarios) and also avoiding running out of inodes in some file systems (maybe less of a problem nowadays in general, not sure as I haven't generated millions of tiny files recently).

Pretty rare for any website to have many files, as they optimize to have as few files as possible(less network requests, which could be slower than just shipping a big file). I have crawled react docs as a test, and it's a zip file of 147mb with 3.803 files (including external resources).

https://docs.solidjs.com/ is 12mb (including external resources) with 646 files

unlog··on Show HN: Crawl a modern website to a zip, serve the website from the zip
Big fan of HTTrack! reminds me of the old days and makes me sad of the current state of the web.

I am not sure if HTTTrack progressed from fetching resources, long time since I used it for last time, but what my project does, is spin a real web-browser(chrome in headless mode which means it's hidden) and then it lets the JavaScript on that website execute, which means it will display/generate some fancy HTML that you can then save it as is into an index.html. It saves all kind of files, it doesn't care the extension or mime types of files, it tries to save them all.

unlog··on Show HN: Crawl a modern website to a zip, serve the website from the zip
Status codes, I am displaying the list because mostly on a JavaScript driven application you don't want other codes than 200 (besides media).

I thought about robots.txt but as this is a software that you are supposed to run against your own website I didn't consider it worthy. You have a point on speed requirements and prohibited resources (but is not like skipping over them will add any security).

I haven't put much time/effort into an update step. Currently, it resumes if the process exited via checkpoints(it saves current state every 250 URLs, if any is missing then it can continue, else it will be done)

Thanks, btw what's your project!? Share!

unlog··on Show HN: Crawl a modern website to a zip, serve the website from the zip
added
unlog··on Show HN: Crawl a modern website to a zip, serve the website from the zip
That's something I haven't explored, sounds interesting. Right now, the zip file contains a mirror of the files found on the website when loaded in a browser. I've ended with a zip file by luck, as mirroring to the file system gives predictable problems with file/folder names.
unlog··on Show HN: Crawl a modern website to a zip, serve the website from the zip
Sure, I forgot about that detail, what license do you suggest?
unlog··on Show HN: Crawl a modern website to a zip, serve the website from the zip
That the page HTML is indexable by search engines without having to render in the server. Such unzipping to a directory served by nginx. You may also use it for archiving purposes, or for having backups.
unlog··on Show HN: Crawl a modern website to a zip, serve the website from the zip
I'm a big fan of modern JavaScript frameworks, but I don't fancy SSR, so have been experimenting with crawling myself for uploading to hosts without having to do SSR. This is the result
unlog··on SeaMonkey All-in-One Internet Application Suite
I have ported multiple tab handler from piro to seamonkey back in the day, I miss xul so much, the browser used to be a very powerful tool
unlog··on Demystifying the Shadow DOM
The first link is broken, can you please fix it, thanks!
unlog··on WinRAR 7.0
A hidden gem of WinRar is that the internal file viewer can open multi gigabit text files faster than many editors.
unlog··on Show me the prompt
imho, just to let you know, I think the original title, more than being baity, it clearly describes a growing sentiment when using the never ending list of AI tools. The prompt is important and not being able to see it, we cant tell if could be improved, and can't test the same prompt on different models. My perception is that the original title is more acurrate. Anyway, thanks!
unlog··on Not all TLDs are Created Equal
My ccTLD gives me much more confidence than some random company that rents TLDs governed by who knows. It's provided by my ISP which is owned by the state, which serves the population.
Page 1 of 3Next →