Ask HN: What's a small example of some of the best code ever written?
Does anyone have any suggestions for such examples? Any language is fine. As long as there's some way to say "look, here is where people say this is great code".
Does anyone have any suggestions for such examples? Any language is fine. As long as there's some way to say "look, here is where people say this is great code".
https://www.gnu.org/software/mes/manual/html_node/LISP-as-Ma...
It's terrible code from a software engineering perspective though. The readability is roughly that of x86 Assembly. The density obfuscates what is actually happening. If someone opened a pull request containing one of these function implementations, I would probably assume they're trying to insert a backdoor somehow. Either way, I'd never accept code like this.
Great code is simple and clear in a human sense, not just minimal in a conceptual or computational sense.
Old time programmers can't get past the low resources of old systems so they continue to prefer clever and tight code over simple and clear code. Last I was in school many of my teachers preferred clever code so that's the coding style we strived to copy. I suspect that's still the gold standard. But I agree, simple and clear in a human sense, is much better. Specially in a team environment where we need to read other's code.
Having to explain your code is just wasted effort and time.
The university system does sometimes result in students having strange ideas that are a bit divorced from industry best practices (and in some cases reality), due to the attitudes of their teachers rubbing off on them. I had one professor who continually insisted we were on the verge of a major resurgence of semantic web technologies and strict XHTML, and a friend who took a dev tools course (mostly focusing on git, makefiles, etc.) and was told tools like Maven would never take off and you're better off managing dependencies by hand (this was around ~2017).
I don't want to be harsh on academics / the tertiary education system in general, but I think it's kind of an unavoidable consequence. Maybe "boot camp" students fare better?
well, it's an eval/apply loop, which by definition, is going to be able to execute code!
I dont think it's obfuscated at all. Of course, reading it requires knowledge of how the read/eval loop works in lisp - something the viewer might not have studied, but that doesn't make it obfuscated.
Not deliberately obfuscated, true. But any code that contains tokens like "CAR", "caar", "CDR", "cdar", "CONS", "caddr", etc doesn't belong in a modern production environment where auditability and maintainability are paramount. What's the goal here, shortest possible text? Why not just replace each of those with a Unicode Greek letter and save two bytes and three characters each?
If that means Lisp itself is the problem, so be it. The assumption that "software engineers" in the late 1950s somehow designed the holy grail of programming languages is laughable anyway.
Do you code in Lisp? I do and I don't find it obfuscated at all. On the other hand I find it absolutely clear and beautiful.
car, caar, cdr, etc. are no more obfuscated than chmod, chown, cp, ls, etc. in shell. If you are unfamiliar with Lisp, sure they would seem obfuscated but so would those commands in shell to someone who has never done shell. But after using the shell for a while those commands become clear, concise and convenient. So do car, cdr, caar, etc. "cons" just constructs a cell. (cons "foo" "bar") constructs a cell with "foo" as its car (first element) and "bar" as its cdr (second element). It can't get simpler than that!
> It can't get simpler than that!
It can, because in your description you introduce multiple levels of indirection.
What would improve things would be to have the mirror image functions:
rac
rdc
raac
rdac
radc
rddc
and so on. We might as well drop the final c, too.These would read out the traversal from left to right. (radad obj) would take the car, then cdr, then car, then cdr, left to right.
This would be better in a left-to-right threading macro:
(--> (whatever) radad (+ 3) ...)
because it would be equivalent to (--> (whatever) car cdr car cdr (+ 3) ...)
the expansion to car/cdr is in the same order.Why not simply (nth 2 obj) which tells clearly that this is a function call that picks the 2nd element in obj and returns it?
Seriously though, arguing about syntax to criticize a language is such a trivial and banal activity that it detracts from any useful discussion about the true merits of a language, about its expressiveness, flexibility and paradigms it supports.
Syntax is superficial (and Lisp has very little of it anyway). What syntax someone likes depends on what syntax they are familiar with. I am familiar with Lisp syntax (the very little there is of it) and C, C++, Python, Java, JavaScript, Go, Rust, APL, etc. and I find Lisp syntax the most intuitive and clear. You might find the C/Java syntax clearer and more intuitive. And that is why a discussion about syntax is superficial and pointless.
Semantics and paradigms make much better topics of discussion when talking about programming languages.
Even if we set nesting aside, our z could be in these positions:
z ;; z is the object itself
(z)
(y . z)
(y)
(x y . z)
(x y z)
and so on.Numeric navigation of the tree structure is possible in various forms. TXR Lisp has the functions cxr and cyr which use a bitmask "address" argument to navigate the car/cdr links. The address 0 is invalid, 1 is a noop. Other than that, bit 1 means car and 0 means cdr. cxr and cyr differ in the order of the bits.
1> (cxr #b10101 '((x (y . z))))
z
I.e. 1> (cxr 21 '((x (y . z))))
z
To interpret the bitmask, we ignore the upper framing bit and follow the 0101 right to left: 1=car, 0=cdr, 1=car, 0=cdr. Of course, this just being a function argument, it can be computed. Any positive integer is a potentially valid address in some cons structure.Indexing is supported, but doens't reach dotted positions; it is geared toward sequences, not cons structure.
2> [[['(a b (c d e (f))) 2] 3] 0]
f
That could be compressed into some kind of path mechanism where you just have the indices as 2 3 0.The numeric address idea is inspired by the tree addressing system in Urbit's Nock. This is explained as a layer-by-layer enumeration using successive natural numbers:
1
2 3
4 5 6 7
Address is 1 is "this object". Then 2 is car, and 3 is cdr. Note that these have a binary structure 10 and 11. Similarly 4, 5, 6, 7 are 100 101 110 and 111. You can see that the bits mean 0=car 1=cdr, left to right, after skipping the framing 1 bit.I swapped their meaning in my implementation and provided a right-to-left function.
I don't remember the exact reason I swapped the roles, but likely it was because linear lists are more common than nested lists. You do more cdr than car, loosely speaking. For instance, the simple positions in a linear list are all cdr operations followed by car: car, cadr, caddr, cadddr, ... if cdr is 0, then these look like 11, 110, 1100, 11000:
40> (cxr #b11 '(1 2 3 4 5 6))
1
41> (cxr #b110 '(1 2 3 4 5 6))
2
42> (cxr #b1100 '(1 2 3 4 5 6))
3
43> (cxr #b11000 '(1 2 3 4 5 6))
4
44> (cxr #b110000 '(1 2 3 4 5 6))
5
In other words, 3 is car, and then we just arithmetically left shift the 3 to get the other positions. If cdr is 1, then we have to left shift and mask in a 1 to replace the 0 that was shifted in.The accessor functions like caddr should probably be avoided in code which is using cons cells as linear lists. If someone is writing a server that pops requests from a list, and I see caddr in that code, I will probably flag that in the review.
The functions are useful in code that manipulates nested cons cell structure, and it's important to be familiar with them.
I'd prefer destructuring or pattern matching to be used.
Some Lisp dialects don't have a pattern matching library included; it requires third party code. Destructuring can be cumbersome because it has a variable in every position of the pattern.
Say we have a variable called obj which holds something like (1 ((2 3)) ...) and we are interested in extracting the 3, and don't care about the other values. Destructuring can be cumbersome; e.g.
(destructuring-bind (a ((b c)) &rest r) obj
c) ;; a, b and r are ignored
Here, there may be excellent justification in just doing (cadaadr obj), if you happen to have the function to that depth. You likely don't, so a break-up into two is required such as (cadar (cadr obj)) and that may even be more readable, since "cadar" and "caddr" are recognizable idioms.In some lower level parts of a Lisp implementation, you may not have any pattern matching or destructuring available, due to the bootstrapping chicken-and-egg problem: the code is helping to implement that.
The meta-circular interpreter is like that. It could use destructuring, but then the goal is for the meta-circular interpreter to be able to handle all the constructs that are used in its own code; thus a definition of destructuring would be required, making it a lot larger.
w for w in words if w in WORDS
There's just two concepts there, w and words, doing a whole lot. I mean, that's literally a filter, but Python refuses to embrace these things because they're supposedly confusing or hard. In F#, that would be: List.filter (fun word -> List.contains word WORDS) words
I would of course rename WORDS and words to be something distinct rather than just differ by capitalization. The Python is confusing because the way to trace this comprehension is by starting with the w in the middle, then the w at the end, and then the w at the beginning (i.e., 2-3-1 sequence).Also, I highly dislike using lines of code as some sort of ultimate measurement. The program is what it is, and then requires a page 5 times the length of the program to explain the program. The program should be clearer or more comments added to it or both.
The program also uses Python's "feature" of evaluating default arguments once when the function is defined as a sort of global variable, because the P function is never used by passing in a value for N. Seems clever, which is the problem with it. Lastly, function names should be more clear. "known", "P", and honestly the others are not good function names.
filter(lambda x: x in WORDS, words)
would I think do pretty much the same as what he wrote, and then depending on how you use the output, you may want to wrap that in list() or something. If you are going to iterate over it, it's gtg as-is.There are much more idiomatic and clear (imo) ways of definining a function-static block (which gets evaluated just the first-time a function is called) than that default arg trick.
Although, F# has list comprehensions as well.
[for word in words do if List.contains word WORDS then word]
That's verbose because I spelled out word, but I feel it reads well while the Python version is weird to me. Mathematical set definitions follow the form of {e <- set | <condition on e>} which is read as "e in set such that e satisfies condition" or "for e in set and e satisfies condition". The Python version kind of reads weird to me as it's more like {e | e <- set | <condition on e>}. (I note that Counter is some iterable object and not a list, I suppose.) I really like comprehensions, but I always find Python's to be a bit strange. I think it's because the "in" in the posted comprehension has two semantic meanings, if I understand correctly. It's a nitpick for sure, but slippery is the best way I can describe Python.> then depending on how you use the output, you may want to wrap that in list() or something
Shudder. Haha.
Yea, the default argument trick threw me off, which is why I don't think it's a good example of the "best" code. The best code doesn't have tricks.
It's a neat domain program, but not the "best" code.
words.filter { WORDS.contains($0) }.forEach { w in
...
} array_map(fn($match) => print($match), array_filter($words, fn($_) => in_array($_, WORDS)));
or less verbosely: array_map(fn($match) => print($match), array_intersect($words, WORDS));https://betterexplained.com/articles/understanding-quakes-fa...
Also a lot of 4k demo scene have some great looking videos rendered purely by a small amount of code.
In the year 2033 Earth was discovered by Survey Fleet MCXII of the Galactic Empire. The Emperor ordered Earth to send a representative to the court, with strict instructions to “bring something beautiful”.
Emperor: What beautiful things does Earth have?
Earth Representative: Your excellency, our mathematician Euler proved in our year 1737 that
+/(1+⍳∞)*-s ←→ ×/÷1-(⍭⍳∞)*-s
Emperor: What is the ⍭ symbol?Earth Rep.: ⍭i is the i-th prime.
Emperor: And what is ∞? Does it do anything useful?
Earth Rep.: It denotes infinity, your excellency.
Emperor: (ponders equation for a minute or two) Respect!
Emperor: Neat notation you have there. Tell me more.
Earth Rep.: Your excellency, it’s called APL. It was invented by the Canadian Kenneth E. Iverson…
⍳ is range 0…n–1
÷ is reciprocal of everything to its right
×/ is product of everything to its right
+/ is sum of everything to its right
At least, that's how my CS prof 20+y ago described it to us :)
The C 'Hello, world!' program documented the style of a whole programming language in 5 lines of code. Every computer engineer who sees it can now recognize C code. Even non -programmers understand vaguely what's going on.
Just about every programming language today starts with the same example. That's how wildly succesfull it became as demo.
This isn't just a small implementation of the C standard library, it's an absolute joy to read and usually every easy to understand. It takes real talent to write code which is both this legible and efficient.
Often when I'm trying to understand what glibc is doing I'll read the musl version of a function and that will give me an indication of what the glibc source is trying to communicate.
He was one of the Gang of Four who wrote Design Patterns, but he wrote this library years before the book came out. It always seemed to me that he invented all the main patterns in the Interviews library before they were subsequently (nicely!) documented in the book (which went on to become a smash hit).
I've since realised that I much prefer functional programming (Clojure!) over OO, BUT, oh my goodness, what a truly groundbreaking and beautiful bit of work it was for its time. I read every line of code in that library, and I became a much better programmer for it.
So much wasted human-centuries of software labor.
I think of things like the PowerBuilder datawindow widget, something that HTML+JS couldn't touch for 15 years until I started to see things like ExtJS and jquery. And it ran on 486s and pentiums.
As for the original question, practically all the superuseful UNIX utils that had to be performant on ancient MHz-class hardware is amazing stuff.
You could charge a credit card with just 7 lines of code.
But I've heard good things about the Fast Inverse Square Root Calculation algorithm in the Doom 3 source code. https://www.youtube.com/watch?v=p8u_k2LIZyo