HNHacker News
TopNewBestAskShowJobs

cben

68 karma · joined August 31, 2008

Long-time pythoneer. Created mathdown.net (collaborative markdown with math). Live in Israel. github.com/cben, gitlab.com/cben, beni.cherniavsky@gmail.com

[ my public key: https://keybase.io/cben; my proof: https://keybase.io/cben/sigs/JGI5fI4zkb-2MZpVugiAm-7WyT2s9AWqqNtSYT5QTIg ]

submissionscomments
cben··on Diagram.codes
Mermaid recently added this, still "experimental": https://mermaid-js.github.io/mermaid/#/README?id=git-graph-e...

But I want more control — name the commits "A", "M" instead of random hash, free-form messages near commits, etc...

cben··on Slack’s new WYSIWYG input box is terrible
FWIW, the editor they use, at least on the web, seems to be https://github.com/quilljs/quill.

I'm guessing Slack implemented something a-la https://github.com/patleeman/quill-markdown-shortcuts, but not very well...

In general the model of allowing separate positions for cursor inside vs outside a span is a huge improvements over WYSIWYG assigning "sticky formatting" to characters. (Hat tip to TeXmacs which was maybe first time I've met that UI pattern.) But you have to do it consistently. E.g. end of literal block has separate inside/outside but start doesn't.

(Of course, in source editing, it comes for free – formatting is delimited by characters like * or ` and cursor is clearly before or after the delimeter.)

----

I believe there is a saner middle ground — formatting as syntax highlighting. You're still editing the source, no hidden state, just immediate in-place feedback. (Nowdays some call this "WYSIWYM" though it's not exactly same) Examples: https://stackedit.io/, https://codemirror.net/demo/variableheight.html, https://simplemde.com/, https://laobubu.net/HyperMD (the latter is borderline, hides too many formatting chars IMHO, but does reveal them when you move cursor to edit.)

cben··on Ask HN: What are the most fundamental books on computer science?
You don't really need much math for SICP. You can learn a few random pieces of math from SICP, and without background that will be harder and you might feel like you're missing out, but these are just examples; you can just trust the math works and focus on how they implement it in code, it's structure of the code they want to teach you about.

And, it gets better, IIRC from chapter 2 there is much less math.

What happens in the first chapter, is that SICP is exploring various control structures (and abstractions over them) before teaching any compound data structures at all. So all it has to work with is numbers. But as number theory teaches us, numbers actually have a rich internal structure, and SICP uses that to make "interesting" examples!

(This aspect of SICP is one of the big critiques in the paper http://cs.brown.edu/~sk/Publications/Papers/Published/fffk-h... which explains the motivations beyond HtDP textbook. In particular, HtDP shows compound data structures first, and uses those to start with explicit methodology for data-directed design of code. These are smart programmers and smart educators, and I trust they made a better intro course. I won't recommend SICP for people who don't know any programming at all. But I feel SICP has a magical voice about the art of programming that HtDP didn't replicate — though I haven't read HtDP, only leafed through it a bit.)

So I don't recommend to to cram algebra/calculus before you begin SICP journey! Just start on SICP, and look up math stuff if/as you want.

At the beginning, IIRC you'll encounter: prime vs composite integers; base-2 representation of integers; some calculus but it's enough to understand derivative = slope. There is some symbolic algebra stuff later that's maybe trickier, but I think (hope?) SICP is actually a great book to demystify that.

---

If you do want to learn math relevant to CS, but mainly for the purpose of learning math:

- If you're comfortable with code more than math: A Programmer’s Introduction to Mathematics by Jeremy Kun.

- "Concrete Math" by Knuth/Oren/Patashnik is wonderful. Don't feel you have to finish it! I stopped ~60%.

cben··on Ask HN: What are the most fundamental books on computer science?
My recommendations are not exactly what you ask but what I suspect you'll enjoy if you're asking this.

(1) "Understanding Computation" by Tom Stuart. Not "fundamental" as a deep textbook, but very approachable for programmers intro into a big chunk of CS, explaining deep ideas about languages using rigorous working clean code (in Ruby, no prior knowledge needed). I especially loved the first few chapters about what it means to define a programming languange and various kinds of formal semantics.

(2) Designing Data-Intensive Applications, Martin Kleppmann. This gives you a phenomenally good survey of concepts and practice of distributed systems. This is more software engineering than pure CS, but in my view you can't approach the field of distributed systems without blending both anyway.

(3) POODR — Practical Object-Oriented Design, in Ruby, by Sandi Metz. This is 100% software engineering, where there is no single definition of "foundational", but many people who read this swear by it. It's remarkably thin but lucid distillation of ideas that were "in the air" but Sandi nailed them down. An important thesis is that good code is not an aesthetic judgement of how it _now_ looks, but objective question how easy it will be to _change in the future_. Not Ruby-specific at all, but it teaches the original Smalltalk "message-passing" view of OOP, that for people that only learnt statically-typed Java, C++ etc view of OOP is a fundamental idea they're missing on.

Finally, not a book, but "the morning paper" https://blog.acolyer.org/ is excellent "return on your time" if you want to sample academic papers, both classic foundational ones, as well as cutting edge.

cben··on Ask HN: What are the most fundamental books on computer science?
Same here about volume 1 — I only have it because my father got it somewhere but neither of us read it much, it's just collecting dust.

However I'm now buying & progressing through volume 4 which Knuth is gradually releasing, and it's _way_ more readable, more novel for me (I already learnt my quicksort from more approachable sources, but this contains recent advanced stuff) as well as fun, if you're mathematically inclined.

Knuth still formats algorithms pretty badly IMHO, with little indentation, one-liner loops, "return to step 2" jumps, single-letter names etc. But at least they're not in assembly any longer :-)

This puzzles me from a man who pays tremendous attention to presentation and readability, literally wrote the book on typesetting algorithms & system he devised _for writing this book_, and introduced the idea of literate programming. Surely this is not oversight but very deliberate style, I just don't enjoy it much, but I'd love to learn why he chose it...

cben··on Larry Wall has approved renaming Perl 6 to Raku
That would be Rakuv רקוב, just Raku doesn't have any meaning in Hebrew. It does sound similar when pronounced though.

(Hebrew word formation is from 3-consonant roots, not concatenation, so removing a letter doesn't give a related word. It's either an invalid word or a different root with 3rd voiceless consonant רקוא which is a not a root / רקוע Raku'a "flattened" but that doesn't sounds quite differently.)

cben··on SQL queries don't start with SELECT
Similar tip: don't start a command from `rm -rf ...`, type `rm /foo/bar/` first only when finished append ` -rf` at the end.

Similarly `git push ... -f` etc.

cben··on TSV Utilities: Command line tools for large, tabular data files
`keep-header` is a superb gem of do-one-thing composability.
cben··on Introducing nushell
Null-terminated mode: turns out tons of tools have these: https://github.com/fish-shell/fish-shell/issues/3164#issueco... But yeah, 2nd "standard" hurts compositionality.

Default posix/bash splitting of output by spaces is very big-prone and having to type "$array[@]" all over the place is annoying. But splitting by lines is much saner. Now newlines in file names, that's just evil and should be outlawed at OS level. I think OSX did this (?)

In bash life can be easier by setting IFS (Google "bash strict mode" though I have reservations about the IFS part).

But really, its easy to fix in a non-posix shell. Fish splits by lines and interpolates variables as arrays by default.

cben··on Introducing nushell
Comments are great for human-edited config files, but not a clear win for an stdout/in format.

The trouble is, comments are normally defined as having no effect, not part of the data model. But iff the comments contain useful information, how do you pick it out using the next tool? Imagine having to extend `jq` with operators to select "comment on line before key Foo" etc... And if it is extractable, are they still comments?

How do you even preserve comments from input to output, in tools like grep, sort, etc.

XML did make comments part of its "dataset". That is, conforming parsers expose them as part of the data. Similarly for whitespace (though it's only meaningful in rare cases like pre tag but that's up to the tool interpreting data), and some other things that "should not matter" like abbreviated prefixes used for XML namespaces). This does allow round-tripping, but complicates all tools, most importantly by disallowing assumptions that some aspect never matters, e.g. that's it's safe re-indent.

I'd argue the depest reason JSON won over XML is that XML's dataset was so damn complicated.

cben··on Show HN: I wrote a book on Python regular expressions
No experience with them in python, but look also for PEG grammars. They are significantly simpler than the traditional tower of lexer + ambiguous limited lookahead grammar.

Plus the memoized Packrat algorithm allows throwing in functions with custom conditions. (Somewhat like parser combinators, they also support custom logic.)

cben··on LightSail 2 Spacecraft Successfully Demonstrates Flight by Light
However, one side reflective/white one absorbing black is possible. I guess the weight of extra paint is not attractive, compared to attitude control which you need anyway.

EDIT: this matters because full 180° reflection is like "bouncing" the photon to opposite velocity, can up to double exchanged momentum vs absorption. (This is why reflective mylar is nice)

cben··on Svgbob: Convert your ASCII diagram scribbles into happy little SVG
Found another similar tool, also preserves the grid & uses monospace text — aafigure: https://github.com/aafigure/aafigure/blob/master/documentati... -> https://aafigure.readthedocs.io/en/latest/examples.html
cben··on Svgbob: Convert your ASCII diagram scribbles into happy little SVG
Very nice! Reminds me of Markdeep diagrams [1]. The examples are almost identical, so I'm curious if one of them should give credit to the other? Or was there a shared ancestor?

Either way, the core insight that it's OK to map the the monospace grid 1:1 to diagram coordinates was IMHO a breakthrough in text->diagram conversion that previous tools like ditaa missed by trying too hard to parse "semantic structure". (After that it's 99% perspiration of course.)

The editor is pretty great.

[1] https://casual-effects.com/markdeep/features.md.html#basicfo...

cben··on Phantom OS, a Russian OS where “everything is an object”
Exhibit A: MS Word "fast save" format: https://web.archive.org/web/20160308183811/http://1017.songt... https://www.joelonsoftware.com/2008/02/19/why-are-the-micros...

From CS perspective, it's an impressive feat of engineering. They used persistent data structures to optimize in-memory operations, and essentially a write-ahead log of these same structures to optimize commits to disk. They also created "OLE" technology in the OS+libraries to support embedding live documents one in another - remarkably not unlike Phantom's goals. The trouble is, they did these in document interchange formats! Office succeeded, and the insane complexity of reverse-engineering its formats actually became a great business "moat".

But decades later, after switching the surface structure from binary insanity to XML, it still contains attributes like `autoSpaceLikeWord95` and `footnoteLayoutLikeWW8` :-( http://www.robweir.com/blog/2007/01/how-to-hire-guillaume-po...

A lot has been written assuming malice in OOXML, but you can also argue it'd require unbelievable competence to evolve their codebase to anything other than this! If you never needed pervasive dump/load abstractions, your formats map to 1:1 to your semantics, and the only way to evolve is keeping old semantics around, forever...

EDIT: I can't help copy pasting more from the last link:

- lineWrapLikeWord6 (Emulate Word 6.0 Line Wrapping for East Asian Text)

- mwSmallCaps (Emulate Word 5.x for Macintosh Small Caps Formatting)

- shapeLayoutLikeWW8 (Emulate Word 97 Text Wrapping Around Floating Objects)

- truncateFontHeightsLikeWP6 (Emulate WordPerfect 6.x Font Height Calculation)

- useWord2002TableStyleRules (Emulate Word 2002 Table Style Rules)

- useWord97LineBreakRules (Emulate Word 97 East Asian Line Breaking)

- wpJustification (Emulate WordPerfect 6.x Paragraph Justification)

- shapeLayoutLikeWW8 (Emulate Word 97 Text Wrapping Around Floating Objects)

cben··on Name It, and They Will Come
I think "Name it and they will come" is now a pretty good quotable name for the Naming / storytelling Solution.
cben··on How I'm able to take notes in mathematics lectures using LaTeX and Vim
eqn (part of troff) from 1974, predating TeX, had a very decent syntax! It's not as complete as TeX, and the layout quality is worse, and I wouldn't want to touch the rest of troff syntax. So not really practical, but good inspiration.

Open/LibreOffice math formulas use a syntax that's clearly inspired by it. Unfortunately it's very lacking in documentation.

Another important question is are you looking for presentational or semantic math? For presentational fine tuning I'm afraid nothing can beat the flexibility of TeX combined with a huge body of Q&A (on stack exchange and elsewhere)...

cben··on HTTP/3 explained
Indeed. But (not sure if that's what GP meant just guessing) it's sad that tons of middleboxes that peeked too much inside packets effectively broke "end-to-end" goals, so now it's impractical to deploy any new protocols (such as SCTP) alongside existing UDP and TCP. Most practical progress is made by layering more and more on top of existing protocols :-(

Actually I'm kinda relieved QUIC succeeded at all with much less layering on top existing stuff than usual. (Compared to say Websockets-over-HTTPS-over-TLS-over-TCP-over-someIPv6-over-IPV4-tunnel...). If it's feasible to deploy a major new protocol over just UDP, that's practically as good as directly over IP!

P.S. I think encryption is the main force that held back the (economically almost inevitable) desire of middleboxes to "add value" by manipulating inner layers.

cben··on For the Love of Pipes
TAOUP had a chapter "a tale of 5 editors" discussing emacs, vi, and more, and does point out emacs is an outlier (and outsider) to many unix principles. It does quote Doug McIlroy speaking against it (but also against vi?). It attempts to generalize from discussing "The Right Size for an Editor" question to discussing how to think about "The Right Size of Software".

I don't know if it's possible to have impartially "fair" discussion of editors. Skimming now, I can see how vi lovers would hate some characterizations there. But it does try to learn interesting lessons from them.

It does NOT simply equate "Emacs has UNIX nature" so you can't just prove something like "TAOUP mentions Emacs, Emacs is GNU, Gnu is Not Unix => TAOUP is not UNIX, QED" ;-)

http://www.catb.org/esr/writings/taoup/html/ch13s02.html

bias disclaimers: I learnt most of what I know of unix from within Emacs, which I still use ~20 years later. I learnt more from Info pages than man pages (AIX had pretty bad man pages). I suspect you have a different picture of unix than I. And I now know better than arguing which editor is better ;)

But I found TAOUP articulated ideas I only learnt through osmosis. I'm looking forward to reading a better articulation if you know one.

cben··on How to teach Git
I only have a sample of 1 (my wife) but I've found Ungit: https://github.com/FredrikNoren/ungit unparalleled for explaining the graph model behind git — what's a merge, what it means to fetch vs pull, what's the difference between committing locally vs push etc...

The specific points that make it great for such teaching:

1. pretty graph

2. hovering over actions such as Commit / Merge / Rebase / Push shows what would happen to the graph if you do it.

3. you can manually "Move" local & remote branches anywhere you want! This is mildly risky as a habit, but much clearer to explain than fast-forward and push, especially with multiple remotes.

4. automatic fetch that works pretty well (though explicit fetch UI with multiple remotes is clunky). For people scared of merging and conflicts, it's liberating to teach "fetch is always safe" and "local commit is always safe", and that you can fast-forward or merge/rebase separate step.

cben··on Input: Fonts for Code
Note that Red Hat had this designed for styling large headings (similar to the font used on highway signs), not for reading code.

https://github.com/RedHatBrand/Overpass https://overpassfont.org

Hmm, nice selection of math operators! If I ever need to put math formulas on highway signs... But I don't want to get people killed, and anyway TeX would have better spacing ;-)

cben··on I’m Chinese, Google’s DragonFly Must Go On
Why did I care when I discovered Infoseek, Excite, Open Directory Project, Dogpile (I'm fuzzy on the exact order) and eventually Google? Do you remember your feeling when playing with a new search engine? Differences in access to information are real and quickly felt.

(disclaimer: ex-googler. In words Google itself likes to use, because long-tail search matters. Also I should mention I now default to DuckDuckGo, on privacy grounds, but it's not a step up just "surprisingly not bad", I frequently repeat a search in Google and get more results)

cben··on Ask HN: What's your favorite elegant/beautiful algorithm?
In the hashing department:

- Zobrist hashing. It's completely incredible that such a trivial scheme is a near-perfect hash for sets. That's before you even consider the useful ability to compute hashes of related sets with added/removed-elements by simple XORing. (Caveat: I'm repeatedly tempted to use this magic for large sets; but it stops working well when number of elements > number of bits, because the matrix is over-determined, so there are easy-to-find subsets that contribute 0)

- Cuckoo hashing. Before going into the specific way it pushes around elements to usually fit exactly 1 per slot, there is a basic idea to grok that makes 2 possible places for an element work much better than 1: "the power of 2 choices". This is also the reason even a little load-balancing is effective. http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.25.8... is a great overview.

- Merkle DAGs. Immutable content hashes are about the only non-painful form of distributed pointers humanity has found. OK, that's an idea not exactly an algorithm, but for example the easy ability to recursively diff 2 git trees is a neat algorithm.

cben··on Ask HN: What's your favorite elegant/beautiful algorithm?
rsync was indeed a dramatic "this should not be possible" feat. But the way it uses rolling hashes is not the simplest possible. One side sends hashes of fixed-size chunks, the other scans input for rolling hash matches, which only works for 2 sides and requires non-cachable CPU work on at least one side.

There is an even simpler idea at the core of many modern backup programs (eg. bup, attic, borg)! Possibly pioneered in `gzip --rsyncable`, not sure. https://en.wikipedia.org/wiki/Rolling_hash#Content-based_sli... Each side slices files into chunks at points where rolling hash % N == 0, giving chunks of variable size averaging N. These points usually self-synchronize after insertions and deletions. This allows not just comparing 2 files for common parts, but also deduplicated storage of ANY number of files as lists of chunk hashes, pointing to just one copy of each chunk!

cben··on Ask HN: What's your favorite elegant/beautiful algorithm?
Yep; it's one-dimentional variant of Error Diffusion (which itself is a notably beautiful dithering algorithm), where "input" = slope of line and output is quantized dy=0 or dy=1.

Dithering also appears in audio. The old PC Speaker had merely 1/0 states. It's supposed to be limited to square waves, limited and unpleasant, right? But at some point people figured out if you flip 1/0 fast enough you can approximate continuous values: http://bespin.org/~qz/pc-gpe/speaker.txt

I once played with using "error diffusion" like dithering to play wav files. In theory that should produce better audio than the simple fixed-table and PWM dithering suggesting in above article, but I did this on a much newer computer (Pentium?) by which time PC Speaker was irrelevant (only for hack value) and only reached 30-40 bits per input sample at 100% CPU. Plus as that article explains, simple PWM might(?) actually produce less noise from timing irregularity.

But the technique is not merely an obsolete hack! Every CD player that visibly brags it has "1-bit DAC" uses a closely related technique, though in fast hardware. See: https://en.wikipedia.org/wiki/Delta_modulation https://en.wikipedia.org/wiki/Delta-sigma_modulation Take me with a grain of salt, I always get confused trying to grok these... Full explanations of the noise shaping properties that make Delta-Sigma highly useful involve quite a bit of signal-processing, but the core accumulate-compare loop is a very simple algorithm...

cben··on Show HN: Markdown New Tab – A new tab replacement to jot down notes in Markdown
Also the venerable https://stackedit.io/. Also the new GitBook https://docs.gitbook.com/v2-changes/feature-highlights#unifi...

I'm sure there are a few more markdown "WYSIWYM" I'm forgetting...

To build your own see also https://simplemde.com/. Basically CodeMirror is friendly to mixed-font syntax highlighting.

Also yours truly https://mathdown.net

cben··on “I'm basically giving myself a permanent vacation from being BDFL”
I had the same gut reaction when I heard the PEP even exists. But now I actually read the PEP, it's very well written and makes a reasonably good case that in specific scenarios it will be an improvement.

(My second gut reaction was it should have used `as NAME` postfix syntax. The PEP debunks that too. Turns out it previously proposed that and the switch to := was a major improvement :-)

The bigger question is whether the gains justify making the language larger... That's subjective, impossible to settle by debate, and the kind of thing a BDFL can help decide one way or the other :-)

cben··on Mermaid: Markdown-like generation of diagrams and flowcharts from text
Client-side is not really feasible (emscripten TeX has been done, but the packages to load are just too huge...).

Server-side, https://tex.s2cms.com/ does a nice job and outputs svg; https://upmath.me/ by same author is markdown editor integrating it - see tikz examples there.

cben··on Mermaid: Markdown-like generation of diagrams and flowcharts from text
GitLab renders math (via KaTeX iirc): https://docs.gitlab.com/ee/user/markdown.html#math
cben··on Fish: A user-friendly command line shell for macOS, Linux, etc
Alt+. is a bash (libreadline really) key.

In fish Alt+Up does something similar, but repeated presses iterates over all words of all commands rather than last words of bash commands. A big improvement is you can type some chars first and then Alt+Up only gets words matching this substring — like Up but word granularity! (Closest bash key is Ctrl+Alt+I)

If you don't mind the differences, and want Alt+. muscle memory to work in fish too, do:

    bind \e. history-token-search-backward
(plus Alt+Up doesn't work for me in linux console, Alt+. does)
← PreviousPage 3 of 8Next →