NetSurf, a multi-platform web browser
netsurf-browser.org
netsurf-browser.org
I really recommand to have a look at the code.
https://source.netsurf-browser.org/libdom.git/tree/src/core/...
https://source.netsurf-browser.org/libdom.git/tree/src/core/...
https://source.netsurf-browser.org/libcss.git/tree/src/parse...
Looks like the authors have a severe case of goto-phobia that causes a quadratic explosion of copy-pasted code in the error-return paths. Some of the files also feel like they have been translated into C from C++ or some other OO language by some automated tool, resulting in some very long "namespaced"-looking identifiers.
Then again, I don't think Firefox or WebKit code is that much better either, so my impression of this codebase is neither great nor horrible. As "beautiful" is subjective, to give a reference for what I'd consider beautiful, look at BSD or early UNIX.
The third one is true goto-phobia of course. Crazy how people can't see the mess they make by avoiding goto for the sake of avoiding goto.
My comment may seem exagerated but that retranscribes very well how I felt when I discovered this code. Often, I'm overwhelmed by the code tree of medium to big-sized projects, with many abstractions and complicated folder structures. This is a browser and yet it was easy to figure out where to find the parts I was interested in and to understand what the code did.
Now, one can always find problems and discuss about the lack of gotos or arrays, but if as a complete outsider I can navigate and understand the code, and even feel that I could hack it quite easily with no documentation, something must be right with it.
I'm not in any way connected to the project, I don't even use this browser (sadly, it is too impractical).
Asking because: DNS-level adblocking is stupid.
It's great, but DNS over HTTPS will end the party soon enough (if I was a smart TV manufacturer, I would be prioritizing adding dns over https to the device firmware to subvert network blocks).
(It'll need to be a constantly-updating blocklist, but the DNS-adblock lists are also that already.)
[1] https://support.mozilla.org/en-US/kb/canary-domain-use-appli...
Much easier I would think for application developers to just make an ad blocking extension, e.g., uMatrix, stop working. For example, they could say this is because the application now has its own built-in ad blocker. Nevermind that the developers are paid from the sale of web advertising services.
I also use a local proxy in addition to DNS which allows me to serve alternative resources or block/redirect certain URLs based on prefix/suffix/regex.
It has made me lose my focus on work repeatedly, and Stack Overflow really is a site related to work for me. So I block that column.
I am more of a non-interactive user and do not use a graphical, javascript-enabled browser much.
Here is a snippet I used to remove the annoying "hot network questions" from the page:
sed '/./{/div id=\"hot-network-questions/,/<\/ul>/d;}' page.html
Out of curiosity I wanted to see if I could access all these networking questions non-interactively. That is, download all the questions, then download all the answers.Some years ago, like 10 years or more, I was making some incremental page requests on SO, e.g., something like /q/1, /q/2, ... and I got blocked by their firewall. What amazed me at the time was the block was for many months, it may even have been a year. This is one of the harshest responses to crawling I ever encountered. One of the very few times I have ever been blocked by any site and the only time I ever got blocked for more than a few hours.
Things have definitely changed since then. To get all the networking questions, I pipelined 277 HTTP requests in a single TCP connection. No problems.
Here is how I got the number of pages of networking questions:
y=https://networkengineering.stackexchange.com/questions
x=$(curl $y|sed -n 's/.*page=//;s/\".*//;N;/rel=\"next\"/P;N;')
echo no. of pages: $x
To generate the URLs: n=1;while true;do test $n -le $x||exit;
echo $y?page=$n;n=$((n+1));done
I have simple C programs I wrote for HTTP/1.1 pipelining that generate HTTP, filter URLs from HTML and process chunked encoding in the responses.Fastly is very pipelining friendly. No max-requests=100. Seems to be no limits at all.
There were 13,834 networking questions in total.
Wondering just how many requests Fastly would allow in one shot, I tried pipelining all 13,834 in a single TCP connection. It just kept going, no problems. Eventually I lost the connection but I think the issue was on my end, not theirs. At that point I had received 6,837 first pages of answers. 211MB of gzipped HTML.
So, it is quite easy these days to get SO content non-interactively.
It was also easy to split the incoming HTML into separate files, e.g., into a directory that I could then browse with a web browser.
x=$(zgrep -c ^HTTP answers.gz)
mkdir newdir; cd newdir;
zcat ../answers.gz|csplit -k - '/^HTTP/' '{'$x'}'As an aside, I've seen mirrors of Stack Overflow pop up when I use DuckDuckGo to search. Google seems to filter these out.
Curious if those SO mirrors did not show cruft like "Hot Network Questions" on every page, would you use them instead of relying on an ad blocker.
I certainly block plenty of Google-controlled domains. I normally do not use Google search and even when I do I never seems to trigger any ads. Maybe I am just not searching for things people want to sell. In the rare event I do trigger an ad, because I am not using a "modern" browser to do searches, these ads are not distracting and I can easily edit them out of the text stream if I want to.
Literally any search that uses the "expensive" keywords[1]. "car insurance quotes" would do nicely, for instance.
[1] https://www.wordstream.com/articles/most-expensive-keywords
Interestingly, the /aclk? Ad URLs do not use HTTPS.
Seeing that these Ad URLs are still unobtrusive, I am wondering why anyone would want to remove them from the search results page. For cosmetic reasons?
I prefer searching from the command line. To remove the /aclk? Ad URL's I used sed and tr.
#!/bin/sh
# usage: $0 query > 1.htm
#
x=$(echo y|tr y '\004');
z=$(echo https://www.google.com/search?q=$@\&num=100|sed 's/ /%20/g');
curl --resolve www.google.com:443:172.217.17.100 -Huser-agent: "$z"|sed "s/<a href=\"\/url?q=/"$x"&/g"|tr '\004' '\012'|sed -n '/url?q=/{s/.url?q=//;s/&sa=.*\"><h3/\"><h3/;s/&sa=.*\"><span/\"><span/;s/$/<br>/;/aclk?/d;p;}'133,215 network filters + 155,733 cosmetic filters
In the stats. Network filters being URL based not just domain based. The lists are easy to view from the uBlock settings page if you want an endless supply of examples. They are used in pretty much any style list: ad, privacy, annoyances, cookie banners, tracking
Pivots to what?
Is there a well-maintained fork that fixes these issues? The main repo[1] hasn't seen updates in over a year now. I read that development slowed down after the main maintainer left suckless, but I'm hoping the community will pick it up. It's an excellent minimal browser.
There's zsurf based on QtWebEngine/Chromium too, FWIW: https://github.com/SteveDeFacto/zsurf
You can use it with QtWebKit instead, but given that's based on a 2016 WebKit with no process isolation or sandboxing, I wouldn't recommend it.
But the former Opera devs started Vivaldi some years ago. And it's really great.
In Settings, search for "Allow Text Selection in Links" and deactivate it. Now you can drag+drop links.
Edit: just a warning, the Windows version has problems with the <select> element, it won't display the options when you click it.
2018 https://news.ycombinator.com/item?id=18692837
It's still relatively big, because web standards are incredibly complicated, but it's almost tractable.
If you care about browser diversity, it's important.
https://www.ekioh.com/flow-browser/ is fully independent and can run Gmail, but it's commercial / non-FOSS. It's funded by contracts in the set-top-box industry.
https://sciter.com/ is non-FOSS with a free edition but you shouldn't run untrusted javascript on it
And there are some lightly-maintained forks of the leaked Presto 12.x engine from before Opera became a Chromium derivative, but of course it's copyright infringement.
Flow browser passes ACID3 [1], closed source but independent (afaik, dont't really know much about it)
Previous HN discussion: https://news.ycombinator.com/item?id=23508979
http://www.netsurf-browser.org/documentation/progress.html
The general impression I get is that it is not bad when it comes to HTML and CSS but not even close for Javascript (disabled by default because of minimal support apparently), so comparing this with Chrome is like apples to oranges, given the way the modern Web is going.
As a counterpoint, when you come across a site which should but doesn't work in one of these minimal browsers, maybe the blame should be put on the site for using needless complexity instead of the browser for not implementing it.
I'm ancient enough to remember when mainstream browsers came in packages of a few megabytes and got us through the day just fine with a featureset less than (or at the very most equal to) the current state of NetSurf. In a somewhat more rational world there would be massive popular pressure for this to remain the case.
Yes, sites break in NetSurf. This is squarely their own fault for not providing civilized degradation. Although it certainly wouldn't go amiss if the NetSurf engine encorporated the worthwhile elements of CSS 3, basically meaning Grid and possibly Flexbox.
I think its a very cool and admirable project
Reading an article: https://l.sr.ht/ZSwt.png
e.g, Redox, maybe Haiku pre-WebKit port? Haiku I could easily be misremembering as there was probably a BeOS browser that ran, albeit outdated...
tl;dr I appreciate the simplicity of it in contrast to modern browsers.
"Copyright 2003 - 2009 The NetSurf Developers". (eg; Downloads)
Also, the duktape JS capabilities are behind even edbrowse.
I notice there's a framebuffer version, so building a nano-linux distro that boots instantly into that shouldn't be that hard.
So for the framebuffer code, maybe an MSDOS UNIVBE wrapper could be used.
For netcode, I have no idea, but a wrapper for POSIX calls must surely exist there also.
Anyyways. I guess the userbase would be pretty much nil, for this as alternatives do exist on DOS :)
The biggest problem however will be compiling the code. The only viable compiler for msdos is openwatcom so the question is if the code uses any fancy extensions or gnuisms. Also the only version of openwatcom that didn't produce broken binaries is the Windows version (or at least was about two years ago).
Happy porting :-)
Its last CVE was 10 years ago[1].
It's not actually necessary for software to change if it already does what you want it to.
1: https://www.cvedetails.com/vulnerability-list/vendor_id-1005...
https://nvd.nist.gov/vuln/search/results?form_type=Basic&res...
The utter waste of resources that is the mindless trendchasing of the "modern web" is in desperate need of some strong opposition.