The beauty of CGI and simple design
rubenerd.com
rubenerd.com
[1]: https://spin.fermyon.dev/
[2]: WASI is a POSIX-ish standard for WASM that gives you all the low level stuff like standard input and output. It includes all the bits and pieces needed for CGI to work.
That makes sense if you want to compile existing software to WASM but that isn't the case here is it?
I don't have a strong opinion on the AssemblyScript controversy[1]. I agree that WASI is imperfect, especially in browsers, but I'm also glad that _something_ exists with not-horrible runtime support. WASI was pretty nice to use as someone who just wanted to get a C program running on WASM.
[1] https://www.assemblyscript.org/standards-objections.html
The minute you want to do url routing and have middleware, a persistent process makes everything much simpler and keeps everything in one place. I much prefer routing inside a framework than using Mod_Rewrite.
Edit: This article covers what I was thinking of:
"Myths About CGI Scalability"
http://z505.com/cgi-bin/qkcont/qkcont.cgi?p=Myths%20About%20...
Previously discussed:
2023 - Use single binary caddy. It's like 2 lines of config total to remove the extension from your CGI url's.
My point really was when having url routing, such as ids/slugs in the url, it is much easier to implement in your application code than in the web server configuration. It keeps application code in one place.
I started with PHP and never trusted persistent processes for anything.
What if your persistent process hangs or crashes for whatever reason?
Is your whole site down then, not just the part with the bug?
The only site I ever made in persistent process paradigm was in ASP.NET and it didn't cause too many problems but it was more of an app, not website, with limited number of logged in users and pretty much no anonymous/guest part.
You've got the same problem with your web server, and your operating system kernel. Simply write a program that can't hang or crash – at least, not due to an application bug. It takes attention to detail, but this is a lot easier than writing a program that always behaves correctly; if you limit the syscalls you make, use timeouts appropriately, and don't segfault, it's often achievable.
Rails? Django? Express? .NET? Do they kill long running requests so they don't affect whole application?
Can the application still mostly work despite syntax error in one of the files?
Another thing is level of trust. I trust Linux kernel, I trust apache webserver. But some webserver written in scripting language few years ago? Not so much.
Not sure about the built-in .net server, but IIS does exactly that. It supports timeout and error/time based restarts for the persistent process.
I trust more the code that I wrote than the code written from others... not because I'm better than others, but because I know what it does, and I know how to fix it if it breaks.
Well it's (maybe was, it did get a bit nicer over years but old shit code is still there) terrible language so no wonder you have no trust.
> What if your persistent process hangs or crashes for whatever reason?
depend on language. For example in Go (in most frameworks at least) panic() in goroutine handling request will just... nuke that request and nothing else. Sure, memory leaks can be issue if the code is terrible but essentially restarting app after every request is just shitty workaround for shoddy code.
That's how most other handle code. Other, that do multi-core processing badly like Ruby, just spawn multiples of server so it isn't really that different from PHP model, you get "supervisor" that spawns few processes of app and feed them requests, if one dies it just gets restarted.
The reason someone who was shaped by working in PHP might not trust long-running processes probably has less to do with any specific shortcomings of it as a language and more to do with the fact that overwhelming execution model is single-request shared-nothing.
This conceptual model works well enough (especially for the web) there's entire cloud services built around recreating it for other languages called "serverless," and they do well because it's a useful simplification which means the issues associated with persistent processes become somebody else's problem.
Devs being what we are, of course there's always those of us who prefer certain things to remain our problem. This is fortunate as someone needs to assume problems like long-running processes in order for others to make it somebody else's problem. Software development has some things in common with comparative advantage.
It hits that sweetspot of speed of development and responsiveness. You also have the benefit of staying in one language.
Django has so much out of the box that you don't have to decide for yourself after the fact. So much accumulated knowledge of what works too.
Sure go ahead rewrite it if it starts making money and you can afford development time/salary. But these older technologies are easier to fit inside one head and let you get enough speed to take off.
Stealing this for work, couldn't agree more.
CGI executables also caused problems when I switched to an Alpha AXP RISC server which emulated x86, bringing the performance to a crawl. That made me switch to classic ASP, and about 10 years later, to ASP.NET MVC with routing, unit tests, abstractions, jQuery, all the shiny things at the time.
Web didn't stop there of course; there came SPA's, React, Vue and whatnot.
Now, seeing the yearning for CGI in 2023 feels funny. Have we come full circle? :)
___
This is only true if you've done it before, and know what you're doing. In reality, it looks like a mess of `mod_cgi` configuration, trying different combinations of file permissions, finding the magic `cgi_bin` directory, finding the right obscure log files when there are inevitably errors, wrestling with CORS and other subtleties of HTTP headers, and other complexities that are only easy to navigate if you're already an experienced CGI user.
That being said, I love the philosophy of using CGI for scripts. Instead of using CGI itself, though, I wrote a (single-file, statically-linked) web server called "QuickServ" to bring this philosophy into the twenty-first century. It has all of the upside of CGI, but is much easier to set up and run, especially for beginners.
One of its benefits is that it automatically parses submitted HTTP forms, and converts the form fields to command line arguments. That means it's extremely easy to put existing CLIs on the web with minimal changes other than writing an HTML form front-end.
If you like CGI, I will (shamelessly) ask that you check it out!
Then in your config file
AddHandler cgi-script .cgi
Options +ExecCGII was talking to a friend the other day who wants to learn some web dev for fun - he had no clue where to start because there's now SO much, so many 'best practice' posts, so much 'you must learn like this'. I know the world changes (and often for the better), but it was a magic time when you could get so far so quickly.
As an aside, it's why I have huge amounts of respect for people like Pieter Levels and his 'just get it done' approach to building things.
By the way, the best HTML tutorial I found so far is Scrimba. The idea of having interactive videos where you can pause and edit the code is amazing.
Equally, half the magic is ability to see it online quickly. Cheap, basically set up free web hosting and FTP made that very easy.
And it's improved a lot PHP, nowadays it's not that bad language...
The thing is that PHP is unbeatable for doing simple stuff, for example a simple contact form in a website, or a page where you have to book slots, and similar stuff like that, since you add a few PHP instruction to the site, upload it to a webserver, enable PHP that is super easy to install, and done. Need to do a modification? You can connect through SSH to the server and edit the PHP files in /var/www with vim, no need to compile, have a build system, restart services, nothing like that!
I got so jaded dealing with this JS nonsense just to get a simple front-end going for my pet project that I just said "screw it" and started using straight up ES6 in the browser. No libraries, no transpilers, no minifiers.
mod_perl teaches you quickly about state.
On that note, I recently revived some nearly 20 year old perl CGI stuff. Had some minor syntax issues (language ambiguities which got fixed), but it was running just fine after having fixed these. There was also some C code with it (some tooling), needed header include fixing + function renaming, otherwise still worked. Small project all in all, not reliant on a lot of external things.
IIRC, the issue was that they didn't have a maintainer for the cgi library anymore, and that is why they removed it.
It's still there, you can still use it. Works great.
The result is here: https://www.thran.uk/cgi-win/fleg.pl
It is a flag generator affectionately titled flegger.
It even does event-based and websockets
What replaced CGI scripts is slightly more complex, but also more naturally grows to other use-cases.
In short: I personally don't miss CGI. :)
- Take untrusted input and pass it directly to shell scripts on your filesystem. What could go wrong?
- Executables in cgi-bin were nearly always included in your document root on shared hosting servers.
- File drops and web shells were rampant
- A trivial misconfiguration (not setting +x?) could result in the web server just serving up the source code of your scripts directly to the client. There's your database credentials hard-coded right there at the top of your file now exposed for all to see.
- You had to pay very close attention to resource limits on the server itself to prevent malicious things like fork-bombs
- Apache was a nightmare to configure securely for CGI
The list goes on. I don't miss it at all.
Who passes non-sanitized or otherwise untrusted input to executables, any more than is necessary? Why do you (seem to) believe modern approaches to dynamism are immune when the bulk of user input still comes from HTML forms? How do modern web applications avoid input from untrusted sources?
Same applies to the directory layout of servers they were hosted on. That’s more a lack of server admin competence than a fault of CGI. Was far easier than configuring Tomcat to safely host Java web apps. I recall it being far easier to configure than WebObjects, PHP, Rails, etc.
CGI programs were far from the only form of web application effected by local security implementations. In my experience common gateway interface applications consist of fewer files than something written against Spring, WebObjects, or 99% of server side web app frameworks. Fewer files means less system configuration and less attack surface, no?
As to forking leading to DDOS, that was something that could happen to any forking server, which was pretty much all servers when CGI first came into existence. FastCGI helped address that issue.
I think what people really want when they pine for the days of CGI is simple file-based routing for dynamic content. Look around and you can see new frameworks are coming back around to this... or consider plain old PHP.
CGI is _computationally_ slow, but in practice it's faster than it needs to be for most sites. Some examples of CGI-powered sites i'm aware of include:
- https://sqlite.org/wasm (same domain, different sub-site)
- https://fossil-scm.org/forum (same domain, separate sub-site)
- https://fossil.wanderinghorse.net
All CGI, all the time, and _plenty_ performant for the purpose.
But if that process is Python or Java it's going to be tens to hundreds of milliseconds before it's ready to serve the request.
It tested in a loop fork/exec and vfork/exec for a dynamic binary and for a static binary with musl.
On an Intel Core i5 6600k:
regular:
fork 13179 ms / 100000 = 131 µs, 7587 per sec
vfork 14423 ms / 100000 = 144 µs, 6933 per sec
static:
fork 3173 ms / 100000 = 31 µs, 31515 per sec
vfork 3833 ms / 100000 = 38 µs, 26089 per sec
However, perhaps 15+ years ago I converted an application to an Apache module from a cgi-bin binary and the performance benefit was immense in an embedded Linux device. We certainly were not talking even about tens requests per second in the CGI version.
You can share state between processes using temporary files/shared memory. Like CGI itself, it doesn’t scale as well, but it’s not even a blip on resource usage unless you’re dealing with thousands of requests per second.
> In particular, it can maintain a database connection all the time (with a retry logic for the case when it gets disconnected).
You can run a daemon that performs the connection pooling for you, and have each request connect to the daemon instead.
This is true, CGI is sensitive to startup latency. It can be addressed with pre-forking, but doing so bears the consequence of increased memory usage at idle workloads.
> As for the connection pooling daemon, you still need to connect to it. It might be faster than talking to the database directly on a different server, but no new connection still beats new local connection.
It’s a nothingburger. A connection pool’s bottleneck is in the TCP connection to the database. A UNIX socket has orders of magnitude less latency (a couple μs) and supports orders of magnitude more throughput (a few GBs/sec). Compared to spawning new processes running interpreters, it’s noise.
Doesn't necessarily extrapolate back another 20 years, but perhaps sheds some light on how "slow" CGI would've been.
Not so uncommon these days tbh
On a related note, "modern" cloud-based application development strategies (like AWS Lambda) seem no faster than CGI. We've come full circle. Though now things are more distributed, of course.
When i first discovered fossil, in December 2007, its ability to run as a CGI was one of its two "killer features" for me, and it's still right at the top of its list of killer SCM features for me.
(Disclaimer: i'm a long-time fossil contributor, but its CGI support pre-dates my joining the project.)
Now, CGI and FastCGI are cool, but who else here built CGIs for running on Pre-OSX Macs? I remember AppleScript/AppleEvents based CGI. The Web Server would be a GUI program running on the Mac and the CGI would also be one. They'd communicate to one another sending AppleEvents. I once made good money with a system for running school exams with it.
I miss those days a bit.
Yes, I’m old.
projectile vomits profusely
> No dependencies or libraries to install, upgrade, maintain, test, and secure.
Except entire webserver. Just install <your language's preferred web server lib> if you want simple.
> Works with any programming language ever that can read text and print text.
You won't use it with those that can't anyway unless for mental excercise
> Zero external configuration, other than telling your webserver to enable CGI on your file.
Except entire webserver. Just install <your language's preferred web server lib> if you want simple.
Seriously, what kind of nostalgia someone needs to have to have any positive thoughts over cgi-bin?
FCGI is an orchestration system. It starts up more worker processes if there are more requests, up to some limit it figures out by itself. If there are few requests, some of the worker processes get an end of file notification and exit. If a worker process crashes, it loads a fresh copy. Little or no human attention required. One of my systems has been up for 1788 days now.
FCGI works fine with Go programs, Python programs, etc. If you use Go, you can get performance reasonably close to what the hardware can do. Until you get to the point that you're renting additional servers as load increases, you don't need more than that for most applications.
You can run FCGI on some cheap shared hosting systems. Why pay more?
Yes.
> Do you have any example I can have a look at?
https://github.com/John-Nagle/vehiclelogserver/blob/master/v...
There's not much machinery required.
package main
import (
"database/sql"
"net/http"
"net/http/fcgi"
)
//
// Called by FCGI for each request
//
func (sv FastCGIServer) ServeHTTP(w http.ResponseWriter, req *http.Request) {
body := make([]byte, 5000) // buffer for body, which should not be too big
if req.Body != nil {
len, _ := req.Body.Read(body) // body of HTTP request
bodycontent := body[0:len] // take correct part of buffer
Handlerequest(sv, w, bodycontent, req) // handle request
}
}
// Run FCGI server
func main() {
sv := new(FastCGIServer)
fcgi.Serve(nil, sv)
}
// Handlerequest -- handle one request from a client
func Handlerequest(sv FastCGIServer, w http.ResponseWriter, bodycontent []byte,
req *http.Request) {
// GENERATE REPLY CONTENT HERE
w.WriteHeader(statuscode) // internal server error
w.Write(content) // report error as text ***TEMP***
}
That will run on low-end hosting at Dreamhost.So there you are, in a modern language, with hard-compiled code with good performance, and reasonably good fault tolerance. Next step up would be a load balancer and redundant databases.
Look at
https://github.com/John-Nagle/vehiclelogserver, which has database connections.
Each FCGI program has a database connection. That database connection persists as long as the program instance is running. So it's not establishing a new database connection for each transaction.
""" Haserl is a small cgi wrapper that allows "PHP" style cgi programming, but uses a UNIX bash-like shell or Lua as the programming language. It is very small, so it can be used in embedded environments, or where something like PHP is too big.
It combines three features into a small cgi engine:
It parses POST and GET requests, placing form-elements as name=value pairs into the environment for the CGI script to use. This is somewhat like the uncgi wrapper.
It opens a shell, and translates all text into printable statements. All text within <% ... %> constructs are passed verbatim to the shell. This is somewhat like writing PHP scripts.
It can optionally be installed to drop its permissions to the owner of the script, giving it some of the security features of suexec or cgiwrapper. """
So you don't actually want this program to do nearly anything useful unless you write it all yourself.
> Zero external configuration, other than telling your webserver to enable CGI on your file.
So now you have to have a webserver external dependency and also configure it. Configuring Apache for CGI securely was no fun unless you were an expert webmaster.
People really put their rose-colored glasses on when they reminisce about this stuff. But there's a reason why we don't use it anymore. We don't need to make up nonsense about why it was so great.
I think the idea is that there's no required dependencies. Unless I misunderstand, you can take on whatever dependencies you want from your chosen language.
> > No dependencies or libraries to install, upgrade, maintain, test, and secure.
This is mentioned as opposed to what we have today with Node.js, Pip and so on. Old CGI programs surely had dependencies. But dependency management was dramatically simpler. Of course there are pros and cons in the comparison. But simplicity is the key reminiscence here.
> But there's a reason why we don't use it anymore.
However I wonder: are all those reasons still valid? To what extent were those reasons motivated by rational decisions instead of pure trend-following? We shall not fool ourselves: not every technical decision we make have grounds on reason. I suspect most of them are not.
https://docs.oracle.com/cd/A97335_02/apps.102/a90099/feature...
I kinda wonder if anyone actually ever used that in production.
I am more interested to know in which situations it's a bad idea to use CGI ?
As one of the previous comments mentions, operating systems do a better job of caching now, so there's less overhead there. Fast CGI also gets rid of that problem by running your CGI script as a single persistent process.
CGI may not be the best solution for really huge applications that have hundreds of routes or very complex business logic.
That said, it's still incredibly useful. I use it on some personal sites, and it's nice to know that I can drop it into any new site without having to worry about dependencies and without having to do much configuration. It's simple and it works. Those are two big pluses in a world where things have gotten so complex.
There is something elegant about PHP and CGI. Just put file somewhere and query it by browser.
How would you fix the design of PHP and CGI to be scalable?
I wrote a userspace scheduler that multiplexes lightweight threads onto kernel threads. It is a 1:M:N scheduler since there is a scheduler thread that preempts other thread's loops. I think I can write a runtime that uses files. I also have an epollserver. If I merge them together and write a HTTP router, then in theory you could have a modern server runner that can execute in-process with threads.
I would need to think how to execute Ruby, Lua, Javascript code in a thread. Perhaps that's similar to NodeJs but it's a file runner rather than a server application.
https://github.com/samsquire/preemptible-thread https://github.com/samsquire/epoll-server
A lot of sites get pretty far with just FastCGI. If that's not enough, then a user-space/mmap cache like Apcu. Then, if that's not enough, horizontal scaling and Redis or similar.
I don't want to make a Google survey or a Doodle or any of that. It's not for me, not off the clock anyway.
I built a simple HTML form to start with. Just typed it out like it was 1997 again (and then added a little Javascript so it remembers that you've submitted the form to soft-block people from submitting more than once; and some CSS to make it look a little nicer and work on phones and such).
Then I thought I had to build a small backend with a database... I'd originally planned to do a simple CGI thing that stuck the form in SQLite or whatever, but then, when I was doing the frontend, I saw that the web server that I use to serve the file actually puts the entire query string in the access log on disk.
So that's it. That's the backend. I can just grep the log file when I want the answers.
Someone should make a framework like Electron or Tauri or NeutralinoJS that works like CGI. Neutralino on is probably the closest but... could be simpler. Just push everything over a pipe.
I’d argue that already exists in the form of your $SHELL.
Though I'm not sure requiring end users to install such a niche thing (assuming it exists) is a great distribution strategy.
Personally though, I don’t think HTML and JavaScript is a particularly nice abstraction for writing UIs. But each to their own
[FastCGI] keeps most of the CGI's simplicity of the interface, while keeping the worker processes running, instead of starting a fresh copy every time. It's very widely used to this day.
[FastCGI]: https://en.wikipedia.org/wiki/FastCGI
FastCGI needs multiplexed transport of HTTP traffic in to and out of a single process, changing code complexity drastically and taxing memory management. It's far from keeping CGI's simplicity.
But if someone else set it up, or you're using one of the cloud ones, FaaS sort of reinvents the cgi-bin mindset. Here is a talk by one of the Fission.io devs showing the similarities.
https://www.youtube.com/watch?v=X-XV6vvwhuo
Not a plug for fission, not affiliated. I mainly like how they integrate with Keda for event/message queue setups.
It makes perfect sense but always felt out of place to me for some reason.
It's not that you can't do what you need to do using the ISAPI, but the ecosystem is just to small for things like the Delphi stuff and keeping up with expectations of customers becomes to much work.
But I assume ISAPI didn't have anything like CPAN or pip. It seems likely you'd have to write everything yourself - perhaps worthwhile if you were a very large corporation with the manpower. But something as simple as PHP would swamp it for the rest of us.
https://www.mit.edu/~yandros/doc/specs/fcgi-spec.html
Easier than integrating a 3rd party tool.
http://python.ca/scgi/protocol.txt
I started implementing a FastCGI client library for Racket, and (though I have experience getting into the bytes and bits of protocol implementation) it seemed harder than it needed to be. So I decided to try SCGI first, and SCGI ended up working great in production:
https://web.archive.org/web/20210822121843if_/https://halest...
https://webcache.googleusercontent.com/search?q=cache:https:...
"If I'm ever in the situation again of helping new people learn web technology then I'm going to get or convert them to use CGI right off the bat. It's easier to teach, it's easier to understand, easier to get working on most webservers and isn't locked in to any particular language or framework.
The only downside of CGI that I know about is the fact it starts a new process to handle each user request. Yes that's a problem in big sites handling hundreds or thousands of visitors per second. But by the time a new student gets to running a big site they will have already encountered many, many other scalability issues in their code and backend/storage. Let alone teaching them database and security concepts. There's a reason we have quotes like "premature optimisation is the root of all evil".
I don't think students new to webdev should be started on anything other than CGI. They can use any language they want. They can actually understand what they're doing. And they're not hitting any artificial barriers or limits set by frameworks or libraries."
OpenWRT still uses CGI
Combined with single binary caddy and sqlite (+wal) gets you down to copy + paste production-ready deployments.
* Perl was very memory inefficient so you could often only serve 4 users at the same time, given the limited RAM available (eg. 512M - 1G). Note that Perl's speed was not a problem. I suspect this one has probably gone away with modern servers with massive amounts of RAM, plus the cloud allows you to quickly scale up and down with demand. Nevertheless using a compiled language would be better. (I later wrote a vastly more efficient framework in C).
* Every action on the page had to round-trip to the server, and always involved a database check (at minimum to map the user's cookie to a user ID). This is sort of solved now that Javascript is widely supported, although that brings its own issues along. Also memcached neatly solves the database access problem. Our website was developed a little bit before memcached.
* It's very clumsy and time-consuming to write any non-trivial CRUD-style action using CGI. eg. Just having a table with update/delete buttons is going to involve writing paging code, update form, update submit form (repeat for every possible action). This is the kind of thing that Ruby on Rails solved quite nicely.
* Organizational problems interfacing with the database. We had to negotiate with the DBA for every schema change, which is a problem when the DBA is a sociopath supported by management. Devops sort of solves this (but also I appreciated not being on call).
* The general hassle of dealing with the web, like setting all the non-obvious HTTP headers to make it secure. I think modern frameworks just deal with this, although I've not used them very much.
However if I was to go back and change anything, it wouldn't have been to change the technology. It would have been to extend the service first to university students and later to the whole world :-)
I cannot even imagine the strange complexities of this circa year 2000 web application if it could only fit 4 simultaneous page loads on 1 GB of RAM. It's baffling, in fact. Perl isn't really memory hungry these days. Was it really that bad 20 years ago? I would be surprised if the Perl 5 interpreter wasn't in fact much leaner back then. At least my immediate impression of this bullet point is that it wasn't a programming language problem, but a software design problem.
In modern terms Perl isn't a problem, especially when you have servers with hundreds of gigabytes.
But, certainly, I'll concede that if the platform is completely CGI-based and for some reason is written such that there's a minimum footprint of 50-100 MB for any single invocation, then servers will be needing a bit of memory in lieu of a software rethink.
Also Perl FCGI existed much earlier than 2000 (looking at CPAN ~'96-97 appear to be first versions), althought obviously the code would be bit more complex than just "dump html on stdout"
It's a petty that in Python, long running threads are the norm.
I hope the Python ecosystem can slowly convert to a CGI style approach.
it’s a nice way to orchestrate long running workers actually: the main process exists exclusively to keep the app up/ talking to gunicorn/uvicorn and register PIDs of a worker process in a queue.
but we’re all in on wsgi, i don’t see that changing.
I prefer it when it just runs for a single request.
There's a big difference between get shit done and enjoy it.
And I started around PHP 4
I found this useful when I create a HTTP service to check a remote service's status over SSH.
Writing CGI scripts with node.js is theoretically possible but you are in for a lot of pain.
And, well, simplicity can be relative: if you already know language X, using it will likely be simpler for you than some other better language that you would have to learn first. The beauty of CGI is that its simplicity makes it a language-agnostic API, letting you pick the simplest language for yourself.
So true it hurts.
The lesson from survivorship bias is to optimize for survival :).
Completely forgot how fast Rails makes everything. Based on the experience I’m probably going to force myself to just use Rails for any personal stuff now and figure if any of them actually take off I’ll rebuild if I need to.
Try not to be influenced by the "omg, that tech is SOOO old!!!" crowd.
Doing what I'm trying to do with CGI would be a huge headache by comparison. When I said "how fast rails makes things" I meant from a standpoint of productivity.
I need to print this out:
"JUST USE RAILS FOR NOW."
If it is just me, for me, I will write it simply and from scratch or use a familiar general purpose framework like rails or laravel. One I have lots of experience with so the project actually gets done, instead of spending all my energy learning another new tool.
That could also be reduced down to "JUST USE RAILS".
Why not get the best of both worlds where you finish projects AND be happy about it? The "for now" makes you think it's a bad decision and you hacked something together to get it done.
But you can finish, stick with Rails and be insanely successful based on what's important to you (start, finish, maintain, profit, IPO, etc.).
A quote from Linus Torvalds someone posted on HN and I saved almost a year ago
- John Gall
However, I wonder how this works for physical things, how does one create a simple version of a train or bridge? Some things have a complexity floor.
(That's a half serous question, of course.)
"The right final architecture" is never achieved by just immediately going out and building that architecture, i.e. by hooking up all the tools required to support that architecture. That's cargo-culting the architecture.
Facebook's use of Cassandra and CI lint-checks and blue-green deployments is just like military cargo planes' use of radio towers — they didn't build those first; they scaled the thing they were doing to the point that these things became necessary support structures, and then they built them.
The "right way" — the right process for engineering a solution — has very little to do with up-front architectural design. The "right way" — the way that'll be most likely to get you to that "right final architecture" eventually — is really the tenable way: the iterative approach that allows you to build up your solution while keeping only one change or consideration in your head at a time. Which means that engineering "the right way" involves not doing all those cargo-cult practices unless/until they become necessary, and even then, only adopting them one at a time. Just like you wouldn't try to make ten different refactorings in a codebase in one patch-set.
Or, to put that another way: YAGNI applies to processes and tools just as much as it does to code. Some projects never exceed 1000 lines. Do those projects need CI cyclomatic-complexity checkers? No.
Only introduce support structures to a project, as the pain of not having them starts to outweigh the pain of adding them.
This is so true. In fact, a lot of the technology choices companies at the scale of Facebook use are only necessary because they're so large/popular, and were added later to stop everything falling over.
Because it's so unlikely your project will ever need to contend with a billion users or whatever, you really shouldn't be designing under the assumption that's likely to happen.
"Tell HN: I was tired of being a perfectionist so I built an app within 24 hours"
More modern languages like Go or Rust provides a lot of safety and is pretty easy to statically compile, so it's easily chrooted. It would however be weird to do a CGI program in Go, because that language already have a really good build in webserver and more powerful/easily accessible features for interacting with requests compared to what CGI provides.
Noting that not everyone has the ability to run standalone servers on their web hoster. Shared hosters generally enable CGIs but cannot support per-user standalone servers, in particular not on common ports like 80 or 443.
https://www.ionos.co.uk/digitalguide/server/configuration/ht...
You could write this off as a library implementation problem, but it comes about because of unfortunate mapping of HTTP parameters to well-known environment variables in CGI, and has tripped up multiple library authors.
* Code runs as web server user * Code is deployed as web server user, because web monkey doesn't understand unix permission (or alternatively chmod 777 everything for same effect) * Code gets hacked, attacker can modify any file so they just leave backdoor in code
If you deploy code as different user attacker can still of course steal all the data but at the very least they can't modify the code that is running.
On top of that, a lot of poor security practices when it comes to qouting stuff. Perl had kinda interesting feature regarding that where in special mode (recommended for apps dealing with untrusted inputs) the variable was considered tainted [1] till something "cleaned" it by passing it thru regexp. So say taking argument directly from environment (so headers for web) or stdin and passing it with no processing to system() would trigger it.
There’s other ways to get that model, but lambda makes it pretty easy.
Personally I think the enthusiasm for lambdas stems from several things:
- A (legitimate) interest in outsourcing non-essential things like server maintenance to somewhere else with more ops experience. This frees up the devs to work on business logic instead of worrying about backups and log rotations.
- Some people are REALLY enthusiastic about the whole scale-down-to-zero thing. I personally think that that is somewhat shortsighted, since that means trading (expensive!) engineering time to save a few 100 USD per month on your AWS bill. Sometimes it is definitely worth it to cut costs, but usually it's engineers overoptimizing the one aspect of profit they have significant influence over (ie hosting costs).
- Some people are very enthusiastic about how you can instantly get to "web scale" with lambdas by just putting a very large max concurrency number. That is somewhat true, but massive overkill for the vast majority of companies. It's "engineer it like Google" syndrome but with Amazon instead.