Kore: a fast web server for writing web apps in C
kore.io
kore.io
> Only HTTPS connections allowed
Seems a bit painful. HTTPS on the webserver itself is a fair bit more painful to setup and administer than HTTPS in your reverse proxy / loadbalancer (I prefer nginx). Web servers should support plain HTTP.
If you have many web servers behind a reverse proxy that takes care of TLS it's often pointless to waste processing power on encryption in between.
$ make BENCHMARK=1
It is not a run time option by design, but it is there.
I want Kore to have sane defaults for getting up and running. That means TLS (1.2 default by only), no RSA based key exchanges, AEAD ciphers preferred and the likes.
edit: spelling
The HTTPS only choice would annoy me a lot because I run most HTTPS services in behind a reverse proxy in a FreeBSD jail on the same host. HA proxy and nginx are still superior to most applications in regard to reliable TLS termination.
Using HTTPS by default a the right choice for a new project but offering no HTTP support (outside of a benchmark) patronizes the user.
All in all this looks like a nice way to export C APIs through HTTPS.
I agree the BENCHMARK build option is a bit confusing. I might end up renaming it altogether.
The complaints about C++ seem to mostly be around the ability to abuse the language, not specific issues that C solves. Something like https://github.com/facebook/proxygen seems like a better API.
And I don't quite buy portability- if it's not a modern compiler with decent security checks then I'm note sure it should be building web-facing code.
This is a bit more difficult for kernel-level code running on an RTOS, as there might be a lot more unique APIs to stub out, but it can still be done.
Although you can add that better test that than nothing, assuming that the respective OS vendor doesn't offer tooling similar to Valgrind, which I fully agree.
EDIT: added some clarifications.
The long story. There are three dominant Turing-complete axiomatizations: numbers (Dedekind-Peano), sets (Zermelo-Fraenkel) and functions (Church). The Curry-Howard correspondence shows that all Turing-complete axiomatizations are mirrors of each other.
If "everything is an object", it means that there must exist a Turing-complete axiomatization based on objects.
Well, such axiomatization does not exist at all. Nobody has phrased one. Therefore, object orientation and languages like C++ are snake oil. They fail after simple mathematical scrutiny. C++ is simply a false belief primarily inspired by ignorance.
Using a threaded model with tiny stacks, and std::lock_guard for atomic operations.
The biggest downside is you have to run the same OS your server uses on your dev box (which is what I do); or you have to upload the source and compile the binaries on your server directly. (or have fun with cross-compilation, I guess.)
To answer the inevitable "why?" -- for fun and learning. Kind of cool to have a fully LAMPless website+forum in 50KB of code. Not planning to displace nginx and vBulletin at a Fortune 500 company or anything.
Still wishing I could do HTTPS without requiring a complex third-party library.
You should check out Docker.
You can't run it on Windows, Mac or the various BSDs without a VM, at which point you're running "the same OS your server uses".
I agree that the reality of the situation is that it's abstracted away enough that the distinction isn't really all that meaningful.
You can't run any containers on osx, but docker itself runs fine as a client binary to docker running on a linux server. I use this configuration daily using a native osx docker binary on my workstation speaking TLS to a docker service running on a CoreOS server.
I use OS X for my main machine, and can happily deploy something to a server running 'not OS X' and be confident that all libraries and dependencies are exactly the same in production as what I was using in development.
Sure Docker is running through a stripped down VM on the Mac, and so technically it's "the same OS my server uses", except it's abstracted away so that for all intents and purposes I'm using OS X for development (and email, browsing, and other things) and deploying hassle free to Linux.
http://www.openbsd.org/cgi-bin/man.cgi/OpenBSD-current/man3/...
But I doubt such tiny processors are being used in boards with network capabilities.
So, although technically the existence of C doesn't make sense, as it is superseded by C++ (except couple of things), C is winning in the branding department.
In the 90s, C++ was much more popular than it is now. It was used as the go-to general purpose language for all kinds of "serious" software (not addressed by VB or Delphi). Early on, almost everyone was very impressed with the power C++ brought, but after a few years, as codebases aged, it became immediately clear that maintaining C++ codebases is a nightmare. The software industry lost a lot of money, developers cried for a simpler, less clever language, and C++ (at least on the server side) was abandoned en-masse -- almost overnight -- in favor of Java, and on the desktop when MS started pushing .NET and C#. So while today C++ is servicing its smallish niche quite well, as a general-purpose programming language for the masses it actually proved to be an unmitigated disaster. It is most certainly not the case that C++'s infamous complexity is "intimidating"; C++'s complexity is infamous because of the heavy toll it took on the software industry. Which is why, to this day, a whole generation of developers tries to avoid another "C++ disaster", and you see debates on whether or not complex, clever languages like Scala are "the new C++" (meant as a pejorative) or not.
https://geeki.wordpress.com/2010/11/21/ken-thompson-on-c/
(when this topic comes up, I always find it funny that url's rarely contain ++ or # when talking about the descendants of C)
and apple of course, with their objective c.
Now both Java and .NET eco-systems are getting there AOT optimizing compilers. .NET with MDIL and .NET Native, Java targeted for Java 10 (if it still goes to plan).
Also, thanks to Oracle's disregard by mobile platforms by not providing neither JIT nor AOT compilers for Java, C++ became the language to go for portable code across mobile OSes when performance matters.
But that's because Java is meant to cover a very wide "middle" -- plus maybe a few corners where possible -- rather than every possible niche.
> Also, thanks to Oracle's disregard by mobile platforms by not providing neither JIT nor AOT compilers for Java
Well, it didn't start out that way, did it? But mobile platforms -- because they're rather tightly controlled -- are, and have always been, much more driven by politics than technical merit.
C++ adds masses of complexity and implicit behaviour. While development in C++ can be quicker and might be 'safer' it can also produce all sorts of unexpected problems.
It also encourages all sorts of nested template types that can make existing codebases incredibly hard to read.
Further, in embedded situations, you may not have space for its standard library.
I have developed and enjoyed developing in both. C has an elegant simplicity about it and you can do literally anything. C++ can be quicker, and it has a bunch of useful standard stuff, but it does have some downsides and quirks. There's room in the world for both.
For anything else where the option was between C and C++, I always picked C++ when given liberty of choice.
Programming Languages are in the domain of UX. Type systems, syntax, RAII -- they all just a serve a means to an end, which is useability.
In my experience, a lot of the developers who prefer C to C++ are developers who wrote a lot of C++, found that it only improved productivity in solving problems that it created, and went back to C and realized how much easier it is to write software in C.
C++ gets you caught up thinking about problems that don't even matter.
which i find curious, given the truism about debugging and understanding code taking longer than writing code.
personally i would prefer 1980s problems people in general know the answer to, rather than up to the minute esoterica.
all that said, the scott meyers modern c++ book is out, so i should really start and finish that before forming an worthy opinion in 2015.
>"says that 'don't use malloc and free, unless you want to debug 1980s' problems."
The problem with malloc and free is that they left to the developer to keep track of the references and memory corruptions bug are not pleasant to debug. RAII idiom actually helps a lot in that context, and that is what Strouptrup was referring to.
i think malloc and free are a lesser evil than this stuff:
http://thbecker.net/articles/rvalue_references/section_08.ht...
feel free to disagree!
http://thbecker.net/articles/rvalue_references/section_07.ht...
Also that example uses a Factory pattern which is actually verbosing the example, but it's like comparing apple and oranges since that's OOP which you actually wouldn't be doing in C.
What you would be comparing is something like:
mystruct * stcVar = malloc(sizeof(mystruct)); free (stcVar);
Vs.
std:shared_ptr<mystruct> stcVar;
Of course that is a stupid example; However things gets more interesting when you have pieces of code sharing the same structure (threads maybe?) and you don't know exactly which code should be the one in charge of releasing the pointer.
clearly, it is C++, not C.
instead of a simple malloc and free, you will invariably end down the rabbit hole learning about rvalues. perfect forwarding, move semantics, lvalues - the list goes on - these topics arise from the simple concept of RAII.
which i find worse than malloc and free.
I got that, what I said is that you wouldn't have that problem in C because it arises when you are doing OOP, and most C devs would not use OOP. Also, keep in mind that it is perfectly fine not using OOP in C++.
> instead of a simple malloc and free, you will invariably end down the rabbit hole learning about rvalues
That's not true. Actually you can be a very decent C++ developer without knowing the notion of what an rvalue is. Move semantics is just an optimization to avoid extra copy of objects, so it is completely optional.
At the moment of writting this, there is an entry on the front page with an example of a modern C++ piece of code. You would notice that there is not a single explicit heap allocation nor any other crazy stuff.
> instead of a simple malloc and free, you will invariably end down the rabbit hole learning about rvalues
truth is subjective to me, in my experience, when you program C++ you will end up exploring a myriad of vast expanses of language features. you may feel that someone can be a very decent C++ developer without knowing that, plenty others would be aghast that someone ignorant of rvalues would describe themselves as 'very decent'.
remember the ostensible creator of the language rates himself at 8/10.
i think that front page example is rather interesting, given the number of #using directives - it gives an indication of the time that developer has spent learning the language. most of them are not trivial to understand to the degree that this guy has. also reading his background, i'm quite sure he knows what an rvalue is.
as you can imagine, i don't really mind explicit allocations - as an awful lot of knowledge is required in C++ to deal with implicit allocations.
in my opinion, there's a lot of 'crazy' stuff there, the short linecount is a product of the author having done his homework.
i get the impression that you and i would use very different subsets of C++. mine would be far smaller! i usually confine myself to whatever idoms the libraries i pull in use and go no futher.
Can we move away from this horrible language already?
http://harmful.cat-v.org/software/c++/ http://harmful.cat-v.org/software/c++/coders-at-work http://harmful.cat-v.org/software/c++/linus http://harmful.cat-v.org/software/c++/rms
I think Torvalds and Stallman fall in that category actually. Stallman even mention generics which C++ doesn't have and it is very different to the templates mechanism that C++ does have.
As for the people who actually uses the language (specially after C++11), most of them would tell you that there is a great subset of the language that works for them. A valid argument I have heard before is that subset is different with each people and the problem arises when maintaining somebody else code. Well that happens with every language (Have you ever had to debug a memory corruption bug in an old C code with void* all over the place? Hint: it's not pleasent )
C++ is a language that allow you to use abstractions with a very reasonable performance and very reasonable resources, and in that niche there is not a real alternative.
Yes, you can have a fine tuned VM running Java or .Net code that might be comparable with C++ but with a cost in memory. I don't think Rust or D are ready yet to be a real competitor.
So... no... we can't move away from this "horrible" language yet.
Now the people who defend C++ or push it to every project (like in this thread) is almost always someone who sit exclusively in C++. You know, the people who think they are experts in C just because they know C++ (like in this thread)? There is the real incompetence.
I also fail to see how hunting down a memory corruption bug in C will make C++ look better, and since you mention it is legacy code it is the same in C++.
It does not matter what language I suggest, the real fact is that C++ user will use C++ for ANY project regardless. So no, we can't move away from this horrible language yet, but not for the right reasons.
It doesn't, I meant it as just an example that every laguange has its own caveats.
>It does not matter what language I suggest, the real fact is that C++ user will use C++ for ANY project regardless.
That just not true in my experience. Actually none of the C++ developers that I know uses just one programming language for everything. They pick the language depending of the requeriments of the project (and that's how it should be isn't it?)
The trick to it is not to over-use or abuse the language's many features. I tend to write C++ that is a fairly thin layer on top of C, using the STL for data structures and algorithms but using things like my own templates very sparingly.
I like Go, but it's really only well supported server side. Rust is promising but not ready for prime time.
Server side due to its memory safety and easy concurrency.
Commandline utilities due to its easy to distribute staticaly linked binaries.
Of course, I have wondered for quite a while if it might not be interesting to do away with DLLs in favor of static linking and then use both disk and memory deduplication to handle the efficiency issues.
Of course the problem here is: what happens when a major bug like Heartbleed is found in a library used by 500 things?
You run away screaming. Static compilation used to be the only option 30 years ago, and there are multiple reasons why mainstream moved away from it.
Unless you write on an embedded system, a game, or a high performance number crunching application, C is premature optimisation.
And even in the above we see drastic changes today, embedded systems have become so powerful that they can run scripting languages (http://www.eluaproject.net), game engines are written in C and scripted with other things (http://docs.unity3d.com/ScriptReference/), and inmemory-bigdata systems like spark offer significant advantages over classical HPC frameworks like MPI (http://www.dursi.ca/hpc-is-dying-and-mpi-is-killing-it/).
While JS is horrid, it at least doesn't have manual memory management.
Many interpreters also make assumptions (e.g. expecting a POSIX'y system) about the host system that many embedded platforms doesn't necessarily meet.
I've more than once looked at interpreters for embedding and found most of the alternatives sorely lacking. Very few interpreters are well suited for embedding on constrained platforms at all (Lua, admittedly is probably one of the more solid exceptions). And once you start having to write lots of support code in C to port or sandbox your interpreter of choice, the reason for considering an interpreter quickly becomes less compelling.
On the other hand, there is a fair bit of negativity in this thread, just because it is C. That might not be in the hacker spirit, so to speak.
We've advanced the state of the art quite a bit with dramatically more expressive languages than C that are sufficiently efficient in terms of memory and CPU. This is especially true when communications are occurring over HTTP and not direct socket-to-socket comms.
Why use C instead of D, Rust, Go, C#, Java, Perl, Python, Ruby, Scala, Clojure, Erlang, Elixir, Haskell, Swift, OCaml, Objective-C...?
I didn't miss C++, it just seems a worse alternative than C.
C is just awesome.
So... you guys added something random and silly like <complex.h> to the standard, but still couldn't get around to a working string implementation? OK. Well, good luck with all that.
Not solved yet.
What does Kore use? libsrt? Let's have a look at how you're supposed to program a web app in its examples:
int
serve_file_upload(struct http_request *req)
{
int r;
u_int8_t *d;
struct kore_buf *b;
u_int32_t len;
char *name, buf[BUFSIZ];
b = kore_buf_create(asset_len_upload_html);
kore_buf_append(b, asset_upload_html, asset_len_upload_html);
if (req->method == HTTP_METHOD_POST) {
http_populate_multipart_form(req, &r);
if (http_argument_get_string("firstname", &name, &len)) {
kore_buf_replace_string(b, "$firstname$", name, len);
} else {
kore_buf_replace_string(b, "$firstname$", NULL, 0);
}
if (http_file_lookup(req, "file", &name, &d, &len)) {
(void)snprintf(buf, sizeof(buf),
"%s is %d bytes", name, len);
kore_buf_replace_string(b,
"$upload$", buf, strlen(buf));
} else {
kore_buf_replace_string(b, "$upload$", NULL, 0);
}
} else {
kore_buf_replace_string(b, "$upload$", NULL, 0);
kore_buf_replace_string(b, "$firstname$", NULL, 0);
}
d = kore_buf_release(b, &len);
http_response_header(req, "content-type", "text/html");
http_response(req, 200, d, len);
kore_mem_free(d);
return (KORE_RESULT_OK);
}
Uh oh. So Kore invented yet another safe string/buf type? Why didn't they use libsrt? What's all this kore_buf stuff?Strings in C are NOT a solved problem. They're a hot mess.
P.S. Before implementing libsrt strings I did a wide study of many string and generic C libraries, implementing the best from all, and adding things that were not still covered (I'll investigate the Kore string/buffer implementation, too):
https://github.com/faragon/libsrt/blob/master/doc/references...
As for ANSI C, maybe someday this will get folded into the standard and we can pass around ss_t* 's rather than char* 's whenever we use third-party libraries.
P.S. I don't expect any standard committee adopting that, not even wide usage (I'm glad just having some feedback! :-D).
This is such a retarded opinion. A person expressing it probably doesn't know shit about C++ or just plainly an idiot.
Because C runs pretty much anywhere? There are plenty of platforms where C is available where I doubt you'd find any of the others above (e.g. C64; yes there are C compilers for them; yes, I'm mentioning it tongue in cheek)
Because you can generate small, compact static executables? E.g. I used to write network monitoring software and an accompanying SNMP server for a system with 4MB RAM and 4MB flash, the latter of which had to include the Linux kernel and a shell on top of the application in question. The system was so limited we did not run a normal init, and couldn't fit bash - instead we ended up running ash as the init...
There are plenty of use-cases where "web application" == "user interface for a tiny embedded platform".
I always hear this argument, but as time has progressed, for better or worse 'anywhere' has become a much smaller target. If your language runs on Intel and ARM then it's good enough. There are a lot of reasons I might choose C for a project, but 'run anywhere' is not one of them.
* List traversal order matters a lot when it's something like a list of monsters getting struck by a spell and the spell has complicated side effects. Brushing it under the rug with abstract iterators or functional Array.map's is a recipe for not knowing how your own game works.
* Realtime is an illusion, it really means "fast turn-based", you don't want players with fast connections to get an advantage by spamming commands and having them executed the instant they're received. You want to queue commands and execute them fairly at regular pulses. So much for all your abstract events infrastructure!!
* Certain object-oriented idioms become eye-rollingly silly when your application actually involves _objects_ (in the in-game sense). Suppose it's a game where players build in-game factories, suddenly the old "FactoryFactory" joke just got a million times worse.
I'm not saying C is the best for those sorts of applications, but it's certainly not bad, and a lot of modern language features just aren't appropriate.
C and Rust are not in the same play field of Ruby, Python or PHP. These languages are typed, compiled and MUCH faster.
You'll obviously build 99% of your application in Ruby, but you might need C or Rust for high-volume calculations.
An example that happened to me a few weeks ago. Scaling a financial application to make millions of calculations. The core App is made with PHP, and the difference between 0.1sec and 0.000764sec gets important here.
"Its main goals are security.."
Is it actually?
I also don't really see an advantage of using something like this over something like the Go net/http package.
Web-type API stuff is usually high enough level that something like C doesn't make sense. Go has nice enough standard packages for system things that even if I was doing a lot of system-y stuff I would be alright. I don't really see the type of work I would be doing where I want to use this.
Sure high level language can help with memory management... but plenty of CVE are because of sloppy coding, not because of low level language.
PHP bugs tend to be more exploitable, because you're doing something supported by the language.
I'm not putting any kind of weight behind that, though; I just feel that it's a bit odd for people (not specifically meaning you, just the whole thread) to criticise this purely on language choice and not put any substance behind their criticisms that actually relate to the software in question.
A high level language should reduce boilerplate and 'force' you to write concise and predictable code.
PHP does none of that especially in the context of error/exception handling.
For example being strict on the network input path and doing proper validation of incoming data is a strong part of the design.
Or was the question more related to, it is C therefor security cannot be part of the process?
C runs "everywhere". Your old C64 from the 80's? Has C compilers.
Go does not, even with gccgo.
I'd bet there are also still likely at least two orders of magnitude more programmers that know C well, and still more programmers with in depth experience of embedded development in C than have even tried Go.
I've still yet to meet anyone outside of the startup devops bubble that have written any Go, and often what language the developers have experience with matters more.
Yeah, but the idea to use them back then was like trying to use Ruby for real time applications in modern days, given how shitty they were.
Home computers might have had implementations of C and Pascal dialects, but we all used Assembly when stepping out of the built-in Basic and Forth enviroments.
Architecture looks pretty interesting too. Wonder why was there a need for an accept lock? Ordinary accept() socket call already allows for simultaneous threads/process wait on a single socket.
Maybe if it is using epoll/select/etc. it would exhibit the thundering herd issue.
The accepting socket is shared between multiple workers which each have its own fd for epoll or kqueue. Because of this a form of serialising the accepts between said workers is needed to avoid unnecessary wakeups.
http://lwn.net/Articles/633422/
See part about EPOLLEXCLUSIVE
Some fantastically quick points from a very cursory glance at the code. Feel free to ignore this.
- The code uses the convention to put the argument of return inside parentheses, making it look like a function call. This is very strange, to me.
- It treats sizeof as a function too (i.e. always parentheses the argument).
- It is not C99, which always seems so fantastically defensive these days.
- It's not (in my opinion) sufficiently const-happy.
- I saw at least one instance (in cli.c) of a long string not being written as a auto-concatenated literal but instead leading to multiple fprintf() calls. Very obviously not in a performance-critical place, so perhaps it's not indicative of anything. It just made me take notice.
I see you picked out the few things that I consistently hear on the coding style I adopted which is based on my time hacking on openbsd. I have no real points to argue against those as it is based on preference in my opinion.
I am curious why you arrived on it not being sufficiently constified however. I'll gladly make sensible changes.
As for the multiple fprintf() calls ... to me it just reads better and the place it occurs in is as you stated pretty obvious non performance critical.
I still don't see the point, or why any sane guide would prefer to treat return as a function. It just never seems helpful to me, and always wasteful/more complicated. I realize it's just two tokens, so it's probably not "important" in any real sense of the word, but it irks me. I like to point it out since it can help others cargo-culting this.
It's not sufficiently const if there are places where a variable could be const but still isn't. :) To be super-specific, the variable 'r' here: https://github.com/jorisvink/kore/blob/master/src/cli.c#L542 is one such case. It should be declared inside the loop, i.e. as "const ssize_t r = write(...);" since once assigned the return value from write(), it's read-only.
Of course, many ancient-smelling style guides seem to outlaw declaring variables as close to their point of usage, too. Note that declaring variables inside scopes other than the "root" one in a function isn't even C99, but many people seem to think you can't do that.
I strongly dislike declaring variables anywhere else but the function root, but I agree with you on the example you provided that those kind of variables could be constified to be sane.
Lastly, you can't just have an event loop without also creating an entirely async platform. For an event loop to work well, all operations from file reading to network requests need to be completely async.
I completely agree with need to async. The hard part is that many operations are async without an async interface. For example memory allocation, or even memory usage if the memory was not truly allocated by malloc.
1) Client libraries you might need to use in your web service might not be available in asynchronous versions.
2) Writing blocking code is much easier to write than asynchronous code.
3) Your server code is CPU bound, so there's no benefit to an asynchronous model.
4) If your web app runs in an asynchronous server and your app crashes, it'll crash the whole server. On the other hand, in a forking model, only the client that the child is serving will be impacted; the other workers will be unaffected.
5) Memory leaks are easier to contain in a forking model, assuming the child can exit or be killed after N requests.
I actually can't think of a case where a multi-threaded/forking-only web server would be faster than that. Again, assuming complete support for async libraries used throughout the web application.
Are there any web servers that have this architecture? NodeJS obviously doesn't. *
* Actually, for maximum absurdity, it looks like Kore, the web server we are currently discussing, has this architecture
| Event driven architecture with per CPU core worker processes
So each process should be able to handle a lot of concurrent connections, just like nginx.
And I tried the websocket example, and saw only the first worker process responding whenever a websocket is created.
It uses an event driven architecture with per CPU worker processes. The number of workers you have can be controlled by the config.
I am doing a lot of that and will keep a look at Kore. Unfortunately, HTTPS only and non-evented core is a no-go for me.
I am currently relying on the web server embedded in libevent, as well as wslay for websockets and some additional code for SSE. To easily start a project, I am using a cookiecutter template: https://github.com/vincentbernat/bootstrap.c-web
https://github.com/jorisvink/kore/tree/master/examples/sse https://github.com/jorisvink/kore/tree/master/examples/webso...
I use uwsgi extensively (not just with python), and I think it sets the bar these days.
The only valid argument to avoid a single event based I/O is some sort of hard blocking I/O such as disk or non-queuing chardev.
However I'm still not biting, this is solved...and as usual the answer is somewhere in between. For example, in RIBS2 there are two models for connection handling, event loops for connections and "ribbons" for the non-queuing bits [1]. RIBS2 is also written in C for C.
[1] https://github.com/Adaptv/ribs2/blob/master/README
Edit - mention RIBS2 is also written in C
It uses per cpu worker processes which multiplex I/O over either epoll or kqueue.
Workers are spawned when the server is started. Each of them deals with tens of thousands of connections on its own via the listening socket they share.
This is a common technique and scales incredible well.
By whom, exactly? There are still plenty of reasons to use a forking web server (see my other comment in this discussion). Saying it "does not scale" is misleading; even with a event-driven model there are only so many CPU resources that can be used to serve responses to clients.
Event-driven webservers are fantastic compared to forking ones, for keeping open many thousands of relatively idle connections (if that is your definition of "scale"). But many web services simply don't do that.
Preforking webservers, like event-driven ones, still have a rightful place in this world. As with all things technology, you have to pick the right tool for the job.
var kore = require("kore");
kore.on("request", http_request);
function http_request(req, resp) {
var statusCode = 200;
resp.write("Hello world", statusCode);
}Can anyone who works in web (I don't, but I'm curious) explain what kind of services is this good and bad for?
It is evented I/O with multiple worker processes.
It is literally in the documentation and easily spottable in the code.