Web application from scratch, part I
defn.io
defn.io
We should change the definition of "full-stack developer" to "given a computer, an opcode reference sheet, a way to enter bytes into a computer, and enough time, is capable of making an operating system and a bunch of applications" :)
Yeah, as long as “enough time” is measured in years I could still be a fullstack developer under that definition.
There's still a good reason to treat layers as magic though. Mental working space is limited and abstractions with minimal leakage are the only way to get things done at ever increasing speeds.
New version comes out, something breaks in a non-obvious way, the industry moves on? They're left with a head full of nothing.
Abstraction is a tool, and like all others it can be abused.
> All problems in computer science can be solved by another level of indirection - except for the problem of too many layers of indirection.
I was a developer for a decade or two before I finally realized something critical to my career: other developers are actively selling me on things that may or may not be a net benefit to me as the years pass.
Developers are a market just like any other, so people make abstraction layers and then become "evangelists" that go out and sell people on stuff. What's to sell? The idea that the framework takes away all the hard work. (I was a big sucker for frameworks with slick UIs back in the day.) We want to believe with a minimum amount of thought, something super cool that will come out that users and other developers will fawn over.
Once I started teaching developers, I ran into people who called themselves professionals but were unable to think through simple activities. Yes, these people were lazy, but they were also products of an environment where they were just moving from shiny object to shiny object. They knew a little bit about various abstraction layers here or there, but nothing about how anything worked. Most of the time, once they went off "happy path", they just banged away at the compiler until things looked like they were working. And these were people being paid to program!
I completely agree this is because of laziness. I don't think it's the entire picture, however.
I personally have a hard time working on something when I can't see through the abstractions to how it works, at least to the CPU and network level. Without knowing how something is working, I don't have a mental model of failure modes or performance characteristics, and I don't know how to compose a solution from the bottom up that will scale well and work reliably.
Thus when I come into something new, I'm frequently less productive than I've observed other people to be, who just take the abstraction at face value, and start coding in terms of it. I catch up in the medium term and overtake in the long term though, because I can debug implementation issues and don't fall into traps where the abstraction is a poor fit.
I consider these two different styles of working as top-down (abstraction first) vs bottom-up (needing to understand the mechanism).
The two styles use very different methods for evaluating third-party libraries and frameworks. Top-down tends to evaluate for popularity and social proof above all else. Even if it only covers 5% of the solution and needs workarounds and add-ons for core solution use cases, it gets used because it's popular. Whereas bottom-up tends to evaluate for conceptual and implementation simplicity, possibly at the cost of under-abstracting the solution, and running the risk of NIH causing redundant work.
Top-down tends gets something up and running and demoable / sellable faster. If you're ambitious in a big organization, or early in a startup, it's the way to go, because it looks like you're moving really quickly, and for MVP purposes you are. Bottom-up tends to build something that will last and can be maintained in the long term, with a smaller semantic gap between the problem domain and the chosen abstractions for the solution, usually because there's an extra abstraction layer that's missing in the top-down solution.
That layer that bottom-up tends to build, and top-down doesn't, is very similar to what Paul Graham talks about in programming bottom up in Lisp: http://www.paulgraham.com/progbot.html - changing the language to suit the problem. You build a library that lets you compose solution primitives at a higher level with just the right chunking for your problem.
“Furthermore, the whole idea of an approximate method was beyond him, even though a cubic root often cannot be computed exactly by any method. So I never could teach him how I did cube roots or explain how lucky I was that he happened to choose 1729.03.”
As I replied to another comment below, I wrote "is capable of doing", not "does".
> ... we all know that we human beings are composed of an enormous number of cells (around wenty-five trillion), and therefore that everything we do could in principle be described in terms of cells. Or it could even be described on the level of molecules. Most of us accept this in a rather matter-of-fact way; we go to the doctor, who looks at us on lower levels than we think of ourselves. We read about DNA and "genetic engineering" and sip our coffee. We seem to have reconciled these two inconceivably different pictures of ourselves simply by disconnecting them from each other. We have almost no way to relate a microscopic description of ourselves to that which we feel ourselves to be, and hence it is possible to store separate representations of ourselves in quite separate "compartments" of our minds. Seldom do we have to flip back and forth between these two concepts of ourselves, wondering "How can these two totally different things be the same me?"
-Douglas Hofstadter, "Gödel, Escher, Bach", Chapter 10
I understand the distinction you're making but in my experience the words are interchangeable. Just because I treat a system as magic one moment does not mean that I can't later dive under the hood and fix it the next. It's just a matter of context switching.
For example i always thought that Windows was developed by people who are much more qualified than myself but when I saw some parts of the source code I quickly realized that it was written by human beings without any magic.
Treat things as a black box but never as magic.
If magic wasn't a black box it wouldn't be magic.
My high-school maths teacher referred to it as "just plumbing".
Those who do not build upon the achievements of others are far less likely to reach any new heights.
It's good to remember what we take for granted, and it's good to understand what is happening in layers below our concern. But pragmatically speaking, it's more valuable to know how to study and understand what's beneath us when we need rather than spending time proving ourselves just for the sake of proving.
Do you still hunt, kill, and butcher your own meat?
Um, yes? it is not that hard. You do need land for it.
I think reaching new heights requires a full understanding of how things work in general. Here is the thing with computers, it isn't that hard to get that knowledge. You just have to apply yourself and build a great deal of things over a career.
A human being should be able to change a diaper, plan an invasion, butcher a hog, conn a ship, design a building, write a sonnet, balance accounts, build a wall, set a bone, comfort the dying, take orders, give orders, cooperate, act alone, solve equations, analyze a new problem, pitch manure, program a computer, cook a tasty meal, fight efficiently, die gallantly. Specialization is for insects.
- Robert A. Heinlein
The socket API is nontrivial and is hardly explained at all. A lot of implementation magic is just coded right in without a word of explanation. The shiniest Python wizardry is used throughout without much of a benefit to clarity - you've got a generator, f-literals, type annotations, typed name tuple syntax, etc. Even if you can read and like type annotations as a documentation aid, the code as written says weird things like 'an http request is a a named tuple and depends on sockets'.
First project where I spent any appreciable amount of time with programmers was eye opening. I was perplexed that not a single one of them had ever run a network sniffer or strace or had the faintest clue about HTTP or how browsers maintained state with their application. I squatted in a room with them for 4-5 weeks and acted a bit like an xray machine for any problems they were encountering. I didn't know anything about programming, so they would tell me what the code was supposed to do and I would show them what was actually happening on the wire or filesystem or whatever. It was a lot of fun and we all left the exercise having learned a lot more than expected. This was nearly 20 years ago, tooling has improved a good bit since then but it still helps to be able to get into the weeds.
I do wish that the skills of sysadmins weren't so downplayed though. (Companies hiring SWE for Ops positions and the notion of 'NoOps'[0] being examples)
[0] https://go.forrester.com/blogs/11-02-07-i_dont_want_devops_i...
Had to do a fair bit of work with Wireshark to troubleshoot protocol issues. It was eye-opening too, especially the amount of traffic generated in the background by Facebook.
Sadly, developers rarely go any further than the network tab in a browser debug tools.
Then we have a true full stack developer. One that shouldn’t feel bad about himself.
Nobody is suggesting that people who don't know all this stuff should feel bad about themselves, just that striving to understand it is admirable.
Also, the idea that you can continue in the same thought process all the way to mining is a little bit weird. Understanding your CPU, its ISA, your OS, etc. are arguably all potentially relevant for software developers. Knowing how to mine is not.
Knowing about the abstraction layers under your javascript helps a lot with efficiency. If more people knew that maybe I wouldn't need 16 Gb of ram just for the browser tabs...
Edit: that doesn't mean you should write your own web server every time you do a new web site, but it would help if you do that once to get an idea of what's going on.
“If you wish to make an apple pie from scratch, you must first invent the universe.” -Carl Sagan
This is the most important part. Is there a tutorial on sockets, perferably not written from a C perspective and in a language agnostic way? How OS specific is it? Is there a common basis over most OSes?
[0] https://www.freebsd.org/doc/en_US.ISO8859-1/books/developers...
UDP is pretty similar but you have to change the code to be connectionless.
There was a book I borrowed that was useful, UNIX Network Programming, Volume 1, Networking APIs: Sockets and XTI
I found beej's guide useful as well. http://beej.us/guide/bgnet/html/multi/syscalls.html
Some languages provide their own useful abstractions (e.g. Python's Twisted/Tornado).
Berkeley sockets is a weak API that became a standard. And you need to know what is happening at the C layer in order to reason about what is going on.
People who know a lot about a thing get Stockholm syndrome. So, a couple of quick notes for why I think it is a bad settlement.
It is riddled with special-cases that you need to know about. For example, there is a key function called select(2). When you call it, it arranges lists that tell you sockets that are currently readable, writeable, in-error. But this knowledge is contextual. If select(2) tells you that a socket is writeable, you need to know what that socket was doing to understand what that means. (Is it an outbound client socket, which is currently making a connection? Or is it a server socket, that is listening? Or is it an established client socket, for which you are the server?)
The API does not give you a convenient place to do this tracking. So you, the developer, need to build structures to one side to track it yourself. This is a form of engineered coincidence that the API foists on your system.
Berkeley sockets has notions of Server and Client. In practice, this causes problems because the handling you do for an inbound Client when you are the server is completely different to the handling you do when making an an outbound Client connection. (It would have been better if they had created separate ideas of Server, Client and Accept.)
Windows has some high-performance APIs that take a quite different approach to Berkeley sockets. They have their own problems, but they demonstrate that the design choices of Berkeley sockets are arbitrary rather than inevitable.
The Berkeley sockets API is good at one thing: once you have learnt all of its dumb tricks, you can reason about what the OS is doing. For this reason, there is a danger to putting convenience APIs on top of it. Stuff will go wrong, and your layer now acts as a barrier between the developer and their problem. (This does not stop people from trying, my own convenience effort is at github.com/solent-eng/solent.)
Step back. I was discussing why you need to know C. When you do socket programming in python, you might be tempted to do this,
lst = [my_socket]
(rlst, wlst, xlst) = select.select(lst, lst, lst, 0)
Underneath, the OS is modifying lst. Hence, the code above does not work as most would expect. rlst, wlst and xlst will be pointers to the same lst, and this code is broken. Approaching the domain from C, this would be more obvious."How OS specific is it?"
The low-performing stuff has been standard in unix for 30 years. Windows NT has a version of select that is similar to unix but has quirks: you can't select on a file in NT, you can't supply empty lists to select.
But here is a critical thing. You set out on a journey, thinking you want to learn about sockets. But over time that changes. You realise: in order to be effective with sockets, you must to construct your system in a way that is different to your familiar procedural or functional approaches.
There are several approaches to chose from, each with their own quirks and substantial learning curves: pure async, threading with blocking sockets, event-sourced threading, forking child processes. The decisions you make here will affect how you do other types of IO. The platforms vary a great deal on these other APIs.
I like the elegance of the pure-async approach, but it is only effective for general-purpose systems if you are on BSD, and you need coroutines to do non-trivial work. Event-sourced threading is the best multi-platform approach, but makes it harder to reason about what your hardware is doing.
This is the real issue. Socket programming is a nasty subtopic of a larger issue: how does your program coordinate IO? You will be operating against many systems APIs for that study, all built by people who think in C.
https://docs.python.org/3/howto/sockets.html
Which should go well with:
http://aosabook.org/en/500L/a-simple-web-server.html
And a follow-up to that might be:
http://aosabook.org/en/twisted.html
Now, this is all python - but lessons from the first two should be helpful in understanding most socket client/servers.
Next I wanted to replace the browser with my own client. Soon realized how huge of a obstacle I was heading towards. Got bogged down by the complexity of a simple browser and still havn't got started on this. Maybe someday :D
http://www.kylheku.com/cgit/tamarind/tree/
Tamarind allows users to manage throw-away e-mail aliases.
I've been using this almost daily for a few years.
It's self-contained; no framework or external libraries are required other than what is in that directory. It just requires an installation of the TXR language in which it is written.
Though it doesn't include its own HTTPD server, it's in a language that I made myself. The arbitrary boundary defining "from scratch" could easily be moved to encompass "own TCP/IP stack and ethernet driver" or "own filesystem", and so on down to the hardware.
https://web.archive.org/web/20180226040031/https://defn.io/2...
Edit: I've read the article after this comment, now I see what the parent comment meant. I guess in this case sth like "\r\n".join(["line1", "line2", "line3"]) is best.