HNHacker News
TopNewBestAskShowJobs

textmode

234 karma · joined June 15, 2016

   int main(){
   char *b[16];
   b[0]="curvecpclient";
   b[1]="server2"; // server name
   b[2]="174b288f3edb7e930040d9f0dd417577b210eb4ef6b9103e6a197210df27bc33";
   b[3]="127.0.0.1"; // server address
   b[4]="80";
   b[5]="90ce4d782b0211e8811d77379f4da525"; // server extension
   b[6]="/usr/bin/curvecpmessage";
   b[7]="-c";
   b[7]="/usr/bin/tcpclient";
   b[8]="-RHDl0";
   b[9]="127.0.0.1";
   b[10]="80";
   b[11]="/usr/bin/http-get";
   b[12]="http://127.0.0.1/1.htm";
   b[13]="1.htm"; 
   b[14]="1.tmp"; 
   b[15]=(void *)0;
   execve("/usr/bin/curvecpclient",b,(void *)0);
   }
submissionscomments
textmode··on Freeing the Web from the Browser
"Different people have different perspectives on how information should be connected, so why do we not allow these range of perspectives to be represented and shared digitally? Why limit ourselves to just one point of view?

...

Why re-create code editors, simulators, spreadsheets, and more in the browser when we already have native programs much better suited to these tasks?"

The title is something I contemplated and began to address long ago, only on a personal level.

With respect to the first question, perhaps this goes to the poor mechanism promoted by Google, to rank the www's contents by "popularity".

This mechanism obviously succeeds for purposes of measuring www user opinion and selling advertising (the later not anticipated by the founders in the early years). However it falls short in the non-commercial context, e.g., the academic setting out of which the company grew. Anyone remember "Knol"?

Today Google search (and probably others seeking to emulate its commercial success) intentionally promote a pattern of usage of their cache/database where its users never reach "page 2" of search results. The company has built their ad sales business on the idea that one perspective ("the top search result") should not only prevail but also that, optimally, other results need not even be considered. It should be obvious that in a non-commercial research context, this is not optimal.

If the www is 100% commercial then of course this is not an issue. But "the www" is difficult to define. All httpd's on any accessible network? All httpd's listening on accessible addresses with corresponding ICANN-registered domainnames? All pages crawled by a commercial bot, deposited in a commercial www cache and made accessible to the public? And so on. In any event, if users only view the www's supposed contents through the lense of a commercial entity, the perception of what the www actually comprises may be manipulated in a way that suits commercial interests, e.g. the sale of advertising.

As to the second question, when given the choice I do not use a popular web browser. The author mentions the utility of "native programs". I would prefer the term "dedicated programs". Programs that perform essentially one task, or "do one thing". Whether such programs can perform their dedicated tasks better than an omnibus-styled program that performs many, varied tasks is a question for the user to decide. For example, the author answers that native programs are "better suited" than the web browser.

The "web browser" has become a conglomeration of once dedicated programs.

There are such dedicated programs for making TCP connections over which HTTP commands can be sent and www content retrieved. This is a task that web browsers can perform, although some users may prefer a dedicated program. In this way content retrieval can be separated from content consumption, alleviating many of the www annoyances such as user tracking, manipulation and advertising.

textmode··on Internet Archive, decentralized
Is there a delay between the time the robots.txt changes and the time when the content becomes inaccessible via Wayback Machine? How often does the archive.org_bot crawl robots.txt?

Can a script check robots.txt periodically for changes and if changes are detected, then download the content from Wayback Machine before it becomes inaccessible?

Additionally, can a script check the domain registration for an anticipated expiration date, or perhaps monitor domainname "drop lists"?

textmode··on The default OpenSSH key encryption is worse than plaintext
I have been using tinysshd for a number of years and I am hooked. Keen to experiment, I have also been using ed25519 keys instead of rsa since this option was added to openssh. No one told me to use tinysshd or ed25519 keys. As someone else pointed out, it seems like most "guides" on ssh, even ones written after ed25519 was added, still advocate rsa keys.
textmode··on Facebook’s New Message to WhatsApp: Make Money
In an interview a number of years ago, one of the founders of WhatsApp said the $1 fee was actually not a business model. I believe the interview is on YouTube.

He said he instituted the $1 fee to try to slow down a relentless increase in new users because he was afraid of potential outages.

textmode··on The Bullshit Web
"I don't think there is anything wrong with user agents downloading resources (like images and stylesheets) linked to by an html document."

Neither do I. For some websites, this is both necessary and appropriate.

However, in cases where the user does not want/need these resources, or where she does not trust the provider, I do not think there is anything wrong with not downloading images, stylesheets, unnecessary scripts, fonts, spyware, advertisements, etc.

textmode··on The Bullshit Web
"I just loaded the New York Times front page. It was 6.6mb."

   ftp -4o 1.htm https://www.nytimes.com

   du -h 1.htm

   206K
For the author, 206K somehow grew to 6.6M.

Could it have anything to do with the browser he is using?

Does it automatically load resources specified by someone other than the user, without any user input?

Above I specified www.nytimes.com. I did not specify any other sources. I got what I wanted: text/html. It came from the domain I specified. (I can use a client that does not do redirects.)

But what if I used a popular web browser to download the front page?

What would I get then? Maybe I would get more than just text, more than 206K and perhaps more from sources I did not specify.

If the user wants application/json instead of text/html, NYTimes has feeds for each section:

    curl  https://static01.nyt.com/services/json/sectionfronts/$1/index.jsonp
where $1 is the section name, e.g., "world".

The user can use the json to create the html page she wants, client-side. Or she can let a browser javascript engine use it to construct the page that someone else wants, probably constructing it in a way that benefits advertisers.

textmode··on The Bullshit Web
Salesforce uses a large quantity of DNS indirection, more than even the large CDNs. I measure the amount of lookups that sites and apps require, and the delay it causes. Most sites on the www only require two lookups.

This is perhaps an example of "... the benefits primarily accrue to the developers of the app and not to the customer."

textmode··on How F5Bot Slurps All of Reddit
I apologise if I confused you. I was simply wondering why he is not using pipelining, which IME can be ideal for the sort of text retrieval he is performing.
textmode··on How F5Bot Slurps All of Reddit
You might be right.

With the "&limit" parameter he can change how many items he receives per HTTP request. This has nothing to do with a limit on how many HTTP requests he can make per TCP connection (pipelining). Maybe that is the "100" he is complaining about, i.e., 100 items per HTTP request.

However you failed to answer my question: Is he making 100 TCP connections to make 100 HTTP requests?

Does the Reddit server set a limit on how many HTTP requests he can make per connection? (100 is a common limit for web servers)

Sometimes the server admins may set a limit of 1 HTTP request per TCP connection. This prevents users from pipelining outside the browser, e.g., with libcurl or some other method.

textmode··on How F5Bot Slurps All of Reddit
"Turns out that Reddit [API] has a limit. It'll only show you 100 posts at a time."

100 sounds like a typical "max-requests" pipelining limit.

He does not mention CURLMOPT_PIPELINING.

Does this mean he makes 100 TCP connections in order to make 100 HTTP requests?

textmode··on An XMPP/Jabber echo bot written in sed
In processing HTML, XML or JSON with sed I have often used tr (e.g., delete newlines, add non-printable delimiter, then replace delimiter with newline) to reformat into sed-friendly input. However, an easy alternative to using tr for this is flex.

As an example, below is a one-off/reusable HTML/XML reformatter in flex. This makes HTML/XML easier for me to read. It also makes it very easy to process with sed and other line-based utilities.

    ftp -4o 1.xml http://web.archive.org/web/20130814000845/http://zombofant.net/blog/ |a.out |less

    flex -8iCrfa 038.l 
    cc -static lex.yy.c

    cat 038.l

    #define echo ECHO
    #define jmp BEGIN
    #define nl putchar(10)
    #define ind fputs("\40\40\40",stdout)
   %s xa xb 
   xa \11|\40 
   %%
   ^\x0d\x0a jmp xb;
   \<{xa}*script nl;ind;echo;jmp xa;
   <xa>\<{xa}*\/script{xa}*\> echo;jmp xb;
   <xa>{xa}{xa}* putchar(32);
   <xa>. echo;
   <xb>\< nl;ind;echo;
   <xb>\> echo;
   <xb>{xa}{xa}* putchar(32);
   <xb>. echo;
   .|\n
   %%
   int main(){ yylex();}
   int yywrap(){ nl;}
textmode··on An XMPP/Jabber echo bot written in sed
STDBUF=U is one of the benefits of using NetBSD libc.

This solution is however limited to applications compiled with the libc C "standard" stdio.

textmode··on Freezing Python’s Dependency Hell
Naive question: Why does this url 302 redirect to medium.com and then medium.com forwards back to the same original url?

Is there some commercial advantage?

Why not just post the medium url

https://medium.com/p/f1076d625241

This 302 redirects to tech.instacart.com

textmode··on An XMPP/Jabber echo bot written in sed
"I would be interested in how you would do especially the tr part in sed..."

There is more than one tr part. :) I tried to be careful to refer only to where the task is replacing characters.

As for joining lines, I have used the same tr -d '\12' technique quite often in the past if the input is not contained in a file. For files I would sometimes use ed scripts.

IME, sed is one of the most ubiquitous programs in UNIX-like OS. I use it daily. However when sed cannot do the job or cannot do it fast enough, I use flex to write relatively small, single-purpose programs quickly. There is a good chance they will compile cleanly on a variety of OS. Perhaps flex is as available to the user as tr, mkfifo, openssl, etc., i.e., when the user has access to all those programs she also probably has access to flex, cpp, as, ld and cc.

I would be happy to provide an example directed at the xmpp-echo-bot solution if you can provide some sample input and the desired output, ICPC-style.

textmode··on An XMPP/Jabber echo bot written in sed
"All you have is openssl, bash, dig, stdbuf and sed?"

s/sed/GNU sed/

echoz.sh also uses cut, base64, stdbuf, tr, mkfifo, tee and sort.

echoz.sh does not seem to have any "bashisms" so perhaps one does not actually need bash and other sh will suffice.

stdbuf is absent from some BSD UNIX. Below is a small script to try to unbuffer output from any program without using LD_PRELOAD. base64 is also absent from some BSD UNIX and can be replaced with

   openssl enc -a
tee can be replaced with

   sed -u w/dev/stdout 
The uses of cut to isolate substrings and tr to replace characters can also be acomplished using sed.

   #!/bin/sh
 
   test $# -ge 1||exec echo usage: $0 program arguments ;
   maxarg=5;
   test $# -le $maxarg||exec echo max $maxarg arguments ;
   {  
   a=$#;
   b=$(command -v $1);
   echo -n '   #include <stdio.h>
   int main(){
   setvbuf(stdin,NULL,_IONBF,0);
   setvbuf(stdout,NULL,_IONBF,0);
   char *b['$((a+1))'];'
   n=0;while true;do
   test $n -le $maxarg||exit;
   a=$((a-1));
   test $a -ge 0||if export n;then echo;break;fi ;
   echo -n '
   b['$n']="'$1'";'
   shift 1;
   n=$((n+1));   
   done;
   echo '   b['$n']=(void *)0;
   execve("'$b'",b,(void *)0);
   }'; 
   }|exec cc -xc - -static && exec a.out
textmode··on We need a new model for tech journalism
Heres an MP3 of the Kara Swisher interview that is mentioned.

https://content.production.cdn.art19.com/episodes/1454f1b6-4...

textmode··on How SSH port became 22
Does it have rules such as "Bitcoin only"? If not, will it open up to other uses in the future?
textmode··on Why Gov.uk content should be published in HTML and not PDF
"[PDFs: ] They're quick and easy to create

... they can be easily created from popular applications that people are already using to author and share documents."

This appears under the heading "Why do people use PDFs?"

However I would have listed this as the sole reason that documents should be distributed as HTML. The reasoning is simple.

Imagine a hypothetical where one has a choice of distributing documents in two formats, A and B, and there are particular advantages to each format. As such, some users prefer format A, while others prefer format B. Not to mention those users who would like to have both formats available.

In the hypothetical, users can easily convert from format A to B however converting from format B to A is difficult.

Assuming one can distribute the documents in format A, it makes no sense to distribute in format B. Users who prefer format A will be unhappy.

Distributing in format A keeps users who like format B happy because they can easily convert from A to B.

textmode··on APL at Its Core

   cat 1.txt

   http://example.net/?p=
   https://www.example.com/?page=
   ...
For a given URL prefix in the file 1.txt, print URLs for pages 25-35. Desired URL prefix is on line 100 of the file.

   cat 1.sh

   #!/bin/sh
   l=$(exec head -$2|exec tail -1);
   n=$3;while true;do
   test $n -le $4||break;
   echo $l$n;
   n=$((n+1));
   done;
1.sh 1.txt 100 25 35

   cat 1.k
  
   g:{_getenv x};
   u:0:g "u";
   l:0$g "l";
   b:0$g "b";
   e:0$g "e";
   `0:{,/$u[l],x}'b _!e;
u=1.txt l=99 b=25 e=36 k 1
textmode··on Djbsort: A new software library for sorting arrays of integers
"The page teaches users to use su to lower privileges ..."

In the example, he could have used his own utilities for dropping privileges (setuidgid, envuidgid from daemontools).

If I am not mistaken, busybox includes their own copies of setuidgid and envuidgid, meaning it is found in myriad Linux distributions. I believe OpenBSD has their own program for dropping privileges. Maybe there are others on other OS.

Instead he picked a ubiquitous choice for the example, su.

It is interesting to see someone express disdain for the version.txt idea. I had the opposite reaction. To me, it is beautiful in its simplicity.

As a user I like the idea of accessing a tiny text file, version.txt, similar to robots.txt, etc., that contains only a version number and letting the user insert the number into an otherwise stable URL.

This is currently how it works for libpqcrypto.

https://libpqcrypto.org/install.html

I would actually be pleased to see this become a "standard" way of keeping audiences up to date on what software versions exist.

By simplifying "updates" in this way, any user can visit the version.txt page or write scripts that retrieve version.txt to check for updates, in the same way any user can visit/retrieve robots.txt to check for crawl delay times, etc.

It is not necessary to "copy and paste" from web pages. Save the "installation" page containing the stable URL as text, open it in an editor, insert the desired version number into the stable URL.

Save the file. Repeat when version number changes, appending to the file.

I like to keep a small text file containing URLs to all versions so I can easily retrieve them again at any time.

textmode··on Djbsort: A new software library for sorting arrays of integers
https://twitter.com/hashbreaker/status/1016951373005455360
textmode··on Show HN: Browsh – A modern, text-based browser
The latest version of links does support SNI.

I use several clients that do not support SNI and one workaround is to connect through a program that does support it, e.g. haproxy, socat, etc.

Privacy/censorship conscious users may dislike SNI, SSL/TLS implementors are now trying to "fix" it, and in fact most SSL/TLS-enabled websites do not require SNI to be sent (popular browsers send it anyway). If requested, I can post stats on whether SNI is required for any list of websites. I have already done this a couple of times with the list all sites currently posted on HN: only a minority require SNI.

textmode··on Show HN: Browsh – A modern, text-based browser
"As of writing in 2018, the average website requires downloading around 3MB and making over 100 individual HTTP requests. Browsh will turn this into around 15kb and 2 HTTP requests - 1 for the HTML/text and the other for the favicon."

Does "100 individual HTTP requests" mean 100 TCP connections?

As far as I know, according to RFC 2616, connection keep-alive was intended to promote making numerous HTTP requests. In fact, IME, most web servers default to setting max-requests at 100. Some are higher.

Following the guidance of the RFCs, for decades I have been using this HTTP feature to make 100 requests to a site in a single connection. That is 100 pages of HTML in one quick TCP connection. If I retrieve 3MB, it is 3MB of HTML from the website and zero from third parties. (Further, I have written filters to remove junk from the HTML and print only the content I want, e.g. for reading or to import into database. If I need to split into separate files, which is rare, csplit works nicely.)

In order to achieve this efficiency I could not and do not use a popular browser authored by ad-supported entities. I am an avid text-based web user who has no need for ads, graphics, and other external resources, e.g. Javascript. As such, I can use clients that can support http pipelining according to the RFCs. It works very well; no complaints.

Best of luck with this project.

textmode··on Where grep came from [video]
tl;dr The idea for grep, like sed, began in the mind of Lee McMahon. The program grep was implemented by Thompson. I recall reading that the program sed was first implemented by Ritchie, and later McMahon himself. Is this correct?
textmode··on The History of Ice Cream
No. You do not have to Javascript to read it either.

The text of the article can be retrieved without any use of Javascript.

textmode··on My home lab setup for highly-available Internet
"Past failures

I used to use a Soekris net6501 as my home gateway, but its CPU maxes out NAT'ing about 300 Mbps, sadly, so I started looking at alternatives when I got Centurylink fiber.

I used to use a UniFi Security Gateway Pro but it failed one day and wouldn't power on any more. Dave had a backup for me handy, but the Unifi controller software wedged itself and wouldn't let me remove the old (dead) one ..."

There is much adoration of Ubiquiti hardware on forums and message boards. I do not doubt for a moment it has been well-deserved.

However, I have a question about the software. I would like to use own kernel and custom utilities.

If I understand correctly, installing one's own choice of OS on Ubiquiti hardware is not always possible and even if successful it carries a penalty in terms of performance versus retaining the Ubiquiti pre-installed proprietary OS.

Soekris made it easy for the user to install the OS of her choice. Tradeoff: More user "control", but a slower router.

The question is: Are there other alternatives to Soekris that can exceed 300mbps and allow for user-chosen OS?

This is another line of (faster) routers where the vendor has allowed for easy installation of user-chosen OS.

https://protectli.com/product-comparison/

There are comments in some other forums and message boards about these computers but I have not seen this company discussed on HN before.

Note the website claims models FW1, 2 and 4 have no Intel ME, SPS or TXE.

https://protectli.com/kb/intel-management-engine-vulnerabili...

textmode··on Evaluating the privacy implications of a canvas fingerprinting countermeasure
How does canvas fingerprinting work when Javascript is disabled in the browser?

What if the HTTP client used to fetch the page does not run third party Javascript?

textmode··on Privacy risks with Facebook’s PII-based targeting: auditing a data broker
tl;dr Facebook tracking pixels placed on websites around the web are automatically loaded by popular browsers' default settings. These benefit Facebook and their customers (advertisers, data brokers) but have created privacy risks for users.

Solution for users: Block loading of tracking pixels, e.g., via browser extensions, DNS, filtering proxy, etc.

textmode··on Against privacy defeatism: why browsers can still stop fingerprinting
"There's no point in securing DNS because your ISP can see what IP addresses you connect to."

Perhaps the ISP has fingerprinted the pages of every website on every shared IP address so they can easily determine which website the user is visiting? Why rely on unencrypted DNS or unencrypted SNI when it is so trivial to map IP addresses to domains.

A poor example, but consider this single IP address 216.239.36.21 with a reverse lookup returning a 23M list of 1,283,151 domains.

https://api.hackertarget.com/reverseiplookup/?q=216.239.36.2...

textmode··on Against privacy defeatism: why browsers can still stop fingerprinting
"Another lesson is that privacy defenses don't need to be perfect. Many researchers and engineers think about privacy in all-or-nothing terms: a single mistake can be devastating, and if a defense won't be perfect, we shouldn't deploy it at all."

This "all-or-nothing" perspective is rampant on www forums discussing computer topics and certainly HN is no exception. It is particularly acute in any discussions of "privacy" or "security".

There are countless examples.

Earlier this week the topic of SNI rose again to HN's front page.

A minor percentage1 of TLS-enabled websites require SNI. An unfortunate side effect of SNI is that it makes it easier for third parties to observe which websites users are accessing via TLS because it sends domainnames unencrypted in the first packet.

Forum commenters will thus argue because there are other, more difficult means for some third parties to observe these domainnames, e.g., through traffic analysis, that the unencrypted SNI is therefore not an issue worth addressing.

All-or-nothing. If the privacy achieved by some proactive measure is not "perfect" then to these commenters it is worthless.

But the HN front page reference suggested otherwise: It was an RFC describing how the IETF is taking a proactive measure, trying to "fix" SNI, encrypting it to prevent third parties from using it in ways detrimental to users.

There is an easier proactive measure. The popular browsers send SNI by default, even if the website does not require it. The default behaviour is to accomodate a minority of TLS-enabled websites at the expense of all users, including those who may not be using this minority of websites.

To make an analogy to fingerprinting, imagine sending 17 unique identifiers with every HTTP transaction when, say, only 5 are actually needed. The all-or-nothing perspective adopted by forum commenters would dictate that it makes no sense to reduce the number unless the number can be reduced to zero.

Amongst the security folks there is a concept sometimes called "defense in depth". Commenters in discussions about security often agree there is no such thing as "perfect" security and they cannot rely on a single, "silver bullet". They must use multiple tactics.

Is privacy somehow different? There are many tactics users can take that, cumulatively, can make things more difficult for the data collectors.

1 Survey of websites currently appearing on HN

Number of unique urls: 367

Number of http urls: 43

Number of https urls: 324

Number of https urls requiring SNI: 38

Number of https urls requiring correct SNI: 26

"Requiring correct SNI" means SNI must match Host header.

Summary

One can fetch 286 of the 324 https urls currently posted on HN with a HTTP client that does not send SNI.

An additional 12 can be retrieved by sending a decoy SNI name that does not match the Host header.

← PreviousPage 2 of 8Next →