OpenBSD, C, httpd and SQLite – Web App stack
learnbchs.org
learnbchs.org
I generally think websites should do their processing in the backend, whenever possible. Running code to generate the website should not be the user's problem. This also comes with the advantage of not being constrained to *Script languages. C is alright, but if you're of the defeatist camp who thinks writing C safely is impossible, you can adapt the BCHS philosophy just as well to C++, Rust, Go, D, Python, AWK, Common Lisp…
From memory management to buffer overflows to understanding the encoding types.
It has essentially the predictability and simplicity of C, just with a minimal amount of dynamic binding to make it comfortable. And NSString variants tend to take care of string handling and encodings.
Swift is an entirely new language, more like Rust or Kotlin. It's also not based on the clang compiler, clang is a C/Objective-C/C++ front end for LLVM.
As to "nicer/modern", well...it does clean up some of the effects of having a hybrid (syntax duplication). Other than that, it is in many ways a step back towards C++ style languages, static/brittle, incredibly complex and incredibly slow to compile.
But anyway for a larger codebase that is about GUI interaction I figure you want more control than what these models provide. You don't want to write event handlers that "push" actions. You want to pull them and process them in a known context.
String handling in C is in general much better than its reputation. You just need to write a couple lines that format to allocated buffers. And you want to error out in case of OOM. Can't check each allocation on the spot. And you need some sort of pool that can release all allocations of the same lifetime at once. (I don't know that this approach is practical for typical web applications, which are extremely string-heavy. It is indeed quite practical for games or compilers.)
Maintaining a C webapp is terrible.
You don't have any primitives, you need to manage your own memory which is extremely dangerous, and the amount of best practices you need to learn so you won't shoot yourself in n the foot and create an exploitable/unstable application is immense.
If you want good performance, then, for non-trivial sites, you cannot afford to spawn the interpeter on each request, so you have to use some asynchronous tooling. Similarly, spawning connections is much heavier weight than it should be, so you end up with backend connection pools.
At work, we used fabric for this, and it blew up in our face at scale, and we had to rewrite.
In order to get off the ground, you also need to use pip, which causes an inordinate number of problems vs. the bsd ports tree or .deb files. Managing dependencies and versions between the host os and pip space still seems to be an open problem. You can use virtual environments for each script, but then patching zero days becomes infeasible.
If you want basic type safety, you need to rely on external tooling. In my experience, third party python libraries evolve types and api’s without warning. This shows up at runtime, so it introduces security holes.
I could go on for hours.
Which Python framework is spawning one interpreter on each request? Most frameworks rely on WSGI where requests are served from a fixed number of interpreters. This approach achieves some honest performance. See for example https://klen.github.io/py-frameworks-bench/.
> At work, we used fabric for this, and it blew up in our face at scale, and we had to rewrite.
What is fabric? In Python, this is an automation tool for SSH. Unlikely to solve the problem of serving web requests.
You mention this as a case against using Python, but this is entirely unrelated to Python as a framework. In other words, had they used PHP or Node or %s, they'd still have to face the issue of optimizing the request-to-response path.
This relates much more to the issues that often come with the common practice of separating your webserver, from your business logic (Flask, in our case).
> "so you have to use some asynchronous tooling"
Solutions such as mod_cgi and wsgi have been the standard for years at this point. That's, in many ways, a solved problem.
> "In order to get off the ground, you also need to use pip, which causes an inordinate number of problems vs. the bsd ports tree or .deb files."
Any specific issues you can expand on? Because this is a rather vague form of criticism.
Using pip is ridiculously easy. I've been using Python extensively in various production environments and rarely had to point a finger at pip. Package and dependency management in Python are solid (that isn't to say things can't improve).
> "You can use virtual environments for each script, but then patching zero days becomes infeasible."
Again, this is an entirely subjective form of criticism. I haven't had any issues with keeping my virtualenvs up to date.
> "If you want basic type safety, you need to rely on external tooling."
Again, you just seem to favor strongly typed languages and that's fine, but that's hardly a point against using Python here. Type safety is simply not a Python thing, much like weak typing is not a C thing.
> "In my experience, third party python libraries evolve types and api’s without warning. This shows up at runtime, so it introduces security holes."
Developers will sometimes break APIs. How is this a Python-specific issue? These things are literally happening everywhere.
What tools do you use keep all of your virtualenvs up to date?
That's actually not happening literally everywhere, that's part of the development culture surrounding (not only) Python. In the C world, maintaining a stable API is the norm, not the exception.
Yes, PHP and Node.JS and Ruby all are just as bad or worse than Python. That's the whole point!
The original comment I was replying to was taking a stab at Python in particular. All I'm saying, is that in the open-source community, in that specific regard, Python doesn't stand out. Many developers care about backwards compatibility, many (dare I say, most) others simply don't.
I've had a large Java repo that broke recently, after we've updated some of our Maven deps to slightly newer versions (that was a very large, well known and used package). So those things happen in the "Java world" too.
(Rewrote big chunks in python multiple times...)
You've mentioned many issues you've had with your Python code base. Would you kindly expand on how/why C++ or Rust solve those in your particular case?
https://learnbchs.org/tools.html
This set of tooling is much more powerful than what I’ve seen deployed in large python environments.
BHCS has static and dynamic analysis out of the box, automatic penetration testing, hardware-assited, os level sandboxing and privilege whitelisting.
The framework provides performant and hardened process management, and hardened json and sql processing out of the box.
Edit: Reading through depends on what you need. I suggest you take some time to look at the tooling provided and what you need (if you have prior experience on the field great). Don't let the hype carry you. Yet, good idea to listen, read or chat with people that had a similar problem to solve (with a grain of salt).
And also OpenBSD pledge / "sandboxing" on other OSs
https://man.openbsd.org/pledge
Bringing sandboxing to all applications and requiring listing of entitlements and digital signing.
Microsoft has decided if the apps don't come to the store, the store goes to the apps, and are merging Win32 with UWP sandboxing concepts via Windows containers and the new MSIX package format.
Android sandboxing is still the most complicated one.
And the answer to this is writing text-processing functions that you expose to the world, in C... skeptical face
Rust is a good solution, but it isn't trivial to learn, and might not suite all situations. I think it's great to have a C solution like this.
> When a TCP packet or UDP packet arrives with a particular destination port number, inetd launches the appropriate server program to handle the connection. For services that are not expected to run with high loads, this method uses memory more efficiently, since the specific servers run only when needed. Furthermore, no network code is required in the service-specific programs, as inetd hooks the sockets directly to stdin, stdout and stderr of the spawned process.
Just ask init or getty.
Init and getty are small, self-contained system demons that are part of the O/S rather than application servers inviting long-running, single address-space processes for business logic.
If you're serving thousands of requests per second, you probably need a pretty beefy server anyway, regardless of the stack you're running. Forking makes good use of multiprocessor capabilities if nothing else.
Shell scripting is a handy tool but it is also slow so you wouldn't write hot paths in shell scripts. And for the same reasons you wouldn't write busy web servers in CGI. Both are fork() heavy.
CGI is also the way the web used to work. Spawning processes has only gotten faster since then. The entire "process spawning doesn't scale to the modern web" argument is completely and totally bogus. Today, spawning a process in Linux is only 10-20 microseconds slower than creating a thread: http://www.bitsnbites.eu/benchmarking-os-primitives/
The performance problems are elsewhere.
The folks who started at a small scale in serverless but see traffic growing and kind of are in the middle space before jumping to dedicated servers have an entire art form of keeping their functions "hot" since both Lambda and AppEngine actually keep your functions hot loaded when once it's spun up for some time.
https://www.google.com.sg/search?q=serverless+warm+up&oq=ser...
Isn't the learning curve of C pretty similar or even higher than rust's if you concider all the best practices you'll have to learn and understand and use? Imho in that regard rust pays off very well (and fast), because it's compiler _enforces_ most of these "best practices" to everyone.
Let SQLite deal with it! The JSON1 extension is quite capable of parsing, extracting, sorting and filtering JSON data and by default it's enabled.
C still has some unique portability characteristics that are very hard to match in other languages. If your requirements lead you down this path, it can be done.
You rightly point out however that this isn’t a path one travels lightly. Maybe the flip side of all this is, “if you’re prepared to invest the kind of rigor into testing that SQLite has, then rock on.”
Maybe another opportunity to mention Zig again as a possible good alternative in the learning curve front? (I learned about it when it appeared recently in the HN front page)
It isn't hard to build tools like Valgrind and AFL into your C workflow. I would take that over simply taking for granted that a higher-level language was secure just because it is interpreted. After all it's probably C under the hood anyway.
For example you could implement a string class in C++ using strlen, strdup etc. if you were feeling insane and it would probably be fine since you only use strdup once and then all the users of you class don't have to worry about getting it wrong.
If you write in C you have to use strdup every time you copy a string and there's no way you get it right 1000 times.
You'd be surprised how far large static buffers and snprintf will get you. Easy to deal with, no leaks, valgrind is happy...
You got me thinking about ways one could get strdup wrong:
- input is not a string -> possible UB
- input is a string, but the character encoding wasn't what you thought -> possible UB
- input is a string, but it was the pointer-plus-length kind -> possible UB
- input is modified by another thread -> possible UB
- strdup called from within a signal handler -> possible UB
- failure to handle error return values -> possible UB
- failure to free the memory when it's no longer needed -> memory leak
- freed the memory more than once -> possible UB
- used the memory after freeing it -> possible UB
I've personally seen several of these in real-world code.
- forget to #include string.h, so strdup is implicitly declared, so its return type is int, which implicitly converts to char*, so everything still compiles -> possible UB
Why not just build the ecosystem to do that and still have it all be “C”?
Why not some fixes to the standard and update then old but useful code, and work on better compilers?
This is what’s been confusing me for a while
If you can write safe enough C for the core of an interpreted language, why not abstract that into tools and patterns that generate better, safer C and learn how to do that over the last 30 years
Instead of JS and dozens of flavors, Python, ruby, lua...
DRY right?
My suspicion is “vanity projects generate a sense of novelty that’s easier to sell.”
But if the OS and bulk of the stack are “C inside anyway” why the extra nonsense?
Also what is the point of calling out on projects that are not using "cool" languages? You don't agree with the developer choice that is OK, probably that developer does not agree with your choice and has other priorities and if is an open source project probably he wants also to have fun while coding it.
Most people who write code in C are not that good at avoiding memory safety issues. The process is mostly to write some code, run it, then fix it until it doesn't crash immediately anymore.
Have you ever been the first person to run valgrind on a codebase? Lots of uninitialized reads and use-after-free issues will be flagged, because they don't normally crash the program and therefore go undetected without using an additional checker.
It can be fun to watch the error messages flow by, until you realize that now you have to file tickets for all of them. And some people won't even see the problem with those errors that don't crash the program, because potentially exploitable vulnerabilities are not as obviously bad as crashes.
Btw, I am not a C fan, in fact I think I have no favorite language, I use what I have to for the project
The time and effort to secure a C program against malicious input is absolutely huge. It would be quicker to write a secure dsl that using raw C. Which brings us back to scripting languages that have that implemented already as a library you load to run your requests through.
What could possibly go wrong
Uh... am I missing the joke? The C programs don't "call binary files", the C programs become the binary files through compilation.
Yes, yes you are.
Seasoned developers would use a tool more fit for the job, like Java, Python or even PHP. Seasoned hipsters would use Node, Clojure or Elixir.
OpenBSD as a base operating system, running relayd and httpd to face the hostile web proxying for efficient services written in Rust would be my ideal choice. SQLite for most simple data store needs, and Postgres if the data management gets big enough that it might spill to another machine.
Second point. We do a lot of RESTful microservices. Which might be good for this in the sense that with C you would want to keep your code small (cause no classes and manual mem-management). But what people do not always appreciate at first about microservices is that you need to rely heavily on libraries (either custom or external like Spring Security) to handle cross cutting concerns, else you are writing the same code multiple times. Essentially if you write a peice of code that you expect will go on more than one microservice, then it should go into library. So my concern is how easily would it be to create custom libraries to deal with cross cutting concerns like logging, and security in this stack, AND does C ecosystem offer external libraries for dealing with these common concerns that a modern RESTful microservice would encounter?
I like the idea of a web framework having a C API. We can make bindings to it in different languages and have a kind of semantic standard across the board, like is done with numerous other things: numeric libraries, crypto, GUI, ...
One issue that seems to fly over the head of “safe” high level language proponents are that you need to know what data types you are manipulating in order to write safe code.
Also, language-level isolation is much harder to implement correctly than process-level isolation (and OpenBSD has spent tons of time hardening process-level isolation).
I’d love to see a security bake off of a few mature applications built with this and random-language-de-jour.
As many mentioned this might be true, but only if you don't have to deal with user input, strings, unicode and anything but extremely simple manual memory management.
The essence is, C and similar languages like Rust are great for writing an OS or a browser, but you really, really, really do not want to write complex server side web apps in those languages.
Sure, you can force your way through and if you have 30+ years of C in your head, it might even be safe. But it is not reasonable nor advisable for the general case. IMHO.
Not sure why you're lumping Rust in with C here. Rust has automatic memory management as part of the language, without a runtime, and it has excellent libraries for dealing with Strings, unicode, and concurrency and asynchronous functions. It's a perfect language for writing a complex server side web app.
The docs seem to suggest otherwise:
> if a column is of type INTEGER and you try to insert a string into that column, SQLite will attempt to convert the string into an integer. If it can, it inserts the integer instead.
>> if a column is of type INTEGER and you try to insert a string into that column, SQLite will attempt to convert the string into an integer. If it can, it inserts the integer instead.
There's nuance to your quote. From my recollection, this means that "321a" will be inserted as "321", but "foo" will be inserted as "foo" (into an INTEGER column). Definitely a wart, on an otherwise fantastic system.
The expression "CAST('321a' AS INTEGER)" will do as you suggest and ignore the trailing 'a' character, yielding an integer 123 result. But that only happens for an explicit CAST. Automatic type conversions must be reversible. That means that '321a' is inserted as a string in an INTEGER column, but '321' (without the trailing 'a') will be converted into an integer 123.
PostgreSQL, MySQL, and SQL Server do exactly the same thing for the '321' case. For the '321a' case, the other three throw an error whereas SQLite just cancels the type conversion and inserts the original string.
Either way, the fact that you can end up with strings in an Integer column is certainly surprising...
sqlite> create table test (foo INTEGER);
sqlite> insert into test (foo) values (123);
sqlite> insert into test (foo) values ("blah");
sqlite> insert into test (foo) values ("123a");
sqlite> select * from test;
123
blah
123aThere is some explanation here: https://learnbchs.org/ksql.html
I assume from importing stdlib etc, which probably give you additional variables to use, that you would need to know about explicitly. One thing I like about zeit.co’s micro server (js) was the argument about not using bodyparser and magically having res.body available, but using async/await and assigning it to a variable. This is popping up more with things like es module imports, render props in react, etc. They make it clear where variables come from and allow you to avoid stepping on existing variables. What other 10-100 variables can I step on potentially with global patterns as a new user?
- pledge: unistd.h[1]
- puts: stdio.h[2]
- EXIT_SUCCESS and EXIT_FAILURE: stdlib.h[3]
You can probably safely assume a C programmer is familiar with puts and EXIT_SUCCESS/EXIT_FAILURE and that an OpenBSD programmer has heard of pledge(2).
[1] https://man.openbsd.org/pledge.2
[2] ISO C11, § 7.21.7.9 no. 1; https://man.openbsd.org/puts.3
[3] ISO C11, § 7.22 no. 3
Simply having all the bells, whistles and free candy, doesn't determine 'toy' status
"Why would anyone need this" was also the typical dismissive response I got when looking around for help.
I appreciate the fact that httpd doesn't attempt to be the end-all server for everybody - its primarily a lightweight way to run things like the BGPd looking glass, other tools, a simple website, or something via fastcgi. The term they use often when denying pull requests is 'Featuritis', which Apache & Nginx suffer from.
With relayd you can add custom headers, as this random example shows:
https://github.com/reyk/httpd/wiki/Using-relayd-to-add-Cache...
However that's probably not great still, because it's a global solution to a per-file problem.
Can someone briefly summarize the security issues in C? If you manage memory properly and take a conservative approach to handling input, where is the risk?
Like I said, I'm young and have only be programming in C for ~3 years.
If you're serious about writing safer C code, certainly check up the following two resources from Robert Seacord:
- SEI CERT C Coding Standard: https://wiki.sei.cmu.edu/confluence/display/c/SEI+CERT+C+Cod...
- Secure Coding in C and C++: https://resources.sei.cmu.edu/library/asset-view.cfm?assetid...
- Decays of arrays into pointers.
- Decays of enumeration into numeric types
- Default signess is implementation defined
- Implicit conversions
- Currently ISO C11 lists about 200 UB cases, C2X plans to list even more
- Strings are a pointer to somewhere in memory that you hope the caller actually terminated them.
- While the pre-processor seems basic, compared with real macros, you can still be very creative with it
- No way to validate security issues in binary libraries
The LLVM and PVS Studio blogs have quite a few examples of C gotchas.
That sounds like a benefit, not a downside. See, for instance, http://libcello.org/
Doing code fixes on a server code which was using clever tricks to convert between memory handles and the real memory addresses teached me that.
Not only does this survey some interesting C behaviors in practice but asks "Do you know of real code that relies on it?" and there's always a number of "yes" / "yes, but it shouldn't" and also "no, that would be crazy" answers to each :D
Put it this way; Would you invest in s startup intending to use this stack ? Can this stack ever be the result of pragmatic decision making?
If you can genuinely answer yes to both in your use case , then why not ...
That said OkCupid famously used a custom HTTP-server to power their site:
Maybe non-mustachioed out of the box but I've done the mustachio thing to generate code for my ASDL parser backend.
Well...I guess it was a C++ library since it was "header only" but I'm sure I could have found a C library if part of the project wasn't to learn boost::spirit.
> C is a straightforward
Indeed, what can be more straightforward than manual memory management and pointer arithmetic...
My favourite flag in clang is -weverything. Everything I write in C I compile with it and fix whatever I reasonably can.
I strangely really like the css formatting used in the website's source code (e.g. https://github.com/kristapsdz/bchs/blob/master/json.css)
You don't /have/ to use it, but it's safer.
I'm aware what FastCGI is ;) Btw. if you're interested in native HTTP I'd be looking into nghttp2 rather than FastCGI and OBSD's httpd.
Edit: see also https://ef.gy/fastcgi-is-pointless and https://news.ycombinator.com/item?id=9202039 (previous discussion of OBSD's httpd)
CGI always spawns off an new process for each request, and always was slow compared to FastCGI, so slowcgi(8) seems like a good name.
Everyone here is completely glossing over this as if it were suggested that this is intended for everybody, for every solution. If you're not extremely confident in your C skills, or have the experience to back it up, This is NOT for you! If you think you can only do it with another language, then This is NOT for you! If you think you're gonna write the next Amazon, This is NOT for you!
And lets not forgot just how many things are written in C and still running the internet. It's just another tool in your toolbox.
Sick of the C Apologism Task Force's idea that "if only rockstar programmers wrote code we wouldn't have any problems", completely ignoring the reality of software development and the evidence of 50 years of exploits.
As for Linus, the kernel was written when he was a student, and I doubt he personally still makes the same low hanging fruit vulnerabilities that are the big issue.
That’s not to say that other languages are perfect, but there are much safer options that are still very fast, such as Rust, or even C# or Java.
I hate to break this to you, but that describes almost every TCP/IP stack and network adapter driver in deployment.
(Not to mention httpd servers, and programming language interpreters written in C.)