Althttpd: Simple webserver in a single C file
sqlite.org
sqlite.org
Something I absolutely love about text based protocols such as HTTP/1 is how easy you can implement it in any virtually programming language. Sure, the implementation is not top-of-the-notch, but it just damned works, it is portable, it is understandable by humans. That's something what's got lost with HTTP/2 and HTTP/3, respectively.
Text protocols have difficult problems like escaping or detecting the end of particular field that are frequent source of mistakes.
The issue is that many (especially scripting) languages treat binary data as second class.
The only real issue is that inspecting the binary message visually is little bit more difficult. You can usually easily tell if your data structure is broken if it is in text form. What I do is I usually have some simple converter that converts the binary message to text form and this helps me inspect the message easily.
I will agree that exhaustively defining text protocols is extremely hard, starting from character set / encoding and getting worse from there.
Typically you create message description which is then compiled to code that can serialize/deserialize messages in BER-TLV or PER-TLV.
I know because I wrote a complete parser/serializer for BER-TLV. It is simple protocol and any security issue is in the parser/serializer and not the protocol itself. That for simple reason that the protocol is nothing more than a format to serialize/deserialize the data.
BER-TLV is really nice protocol. I have worked with it for couple of years when I worked on an EMV application. It uses BER-TLV to communicate with the credit card but it is also very convenient format for all sorts of other uses and I would use it wherever I could. Think of it as Json but in binary form. It is not complicated and I would not even bother parsing the messages -- I could interpret hex dumps of them by sight very easily.
>>* a legal notice, here is a blessing:
>>*
>>* May you do good and not evil.
>>* May you find forgiveness for yourself and forgive others.
>>* May you share freely, never taking more than you give.
Above all, this note at the beginning of the source code impressed me.
I don't disagree in principle but I've come across a handful of very poorly documentated binary protocols in my years and that is an extremely painful thing to deal with compared to text-based protocols.
GET /asdfasdfasdfdsaf
X-Asdf-Asdf: 83e7234
HTTP/1.0 202
X-83e7233: 1
X-83f730b: 4
It's text, but you still have no idea what's going on.Overall, I think it's kind of a wash. It's basically equally easy to take a documented text or binary protocol and write a working parser. Neither format solves any intrinsic problems -- malicious input has to be accounted for, you have to write fuzz tests, you have to deal with broken producers. It's a wash.
People like text because they can type it into telnet and see something work, which is kind of cool, but probably not a productive use of time. I can type HTTP messages, but use curl anyway. (SMTP was always my favorite. "HELO there / MAIL FROM foo / RCPT TO bar / email body / ." Felt like a real conversation going on. Still not sure how to send an email that consists of a single dot on its own line though.)
It seems you prefix any line of data starting with a period with an additional period, to therefore distinguish it from the end of mail period.
A line starting with a period is escaped by adding an extra period; the receiving side removes the first character of a line if it is a period.
However, HTTP/2 and HTTP/3 certainly are, though the reasons why they are complicated have nothing to do with choosing to use a binary based format. (They are complicated for good reason, though, and I hope that browsers and servers can continue to support HTTP/1 as the baseline till the sun burns out, just to make life easier.)
Nobody does that anymore and debugging is easily solved by converting your binary protocol to a textual form.
* https://en.wikipedia.org/wiki/History_of_email
Perusing document-based information was also done via clients, first with Gopher and then with the WWW: either GUIs like Mosaic, or on the CLI via (e.g.) Lynx.
well, other side of HTTP was telnet in the terminal - this is why it is ASCII to begin with
https://en.wikipedia.org/wiki/Line_Mode_Browser
But that didn't just spit out the HTTP response - it parsed and rendered the HTML.
Sure enough, the function StrAppend potentially overflows a size_t size (without checking), and then writes into memory could be past the end of the allocated buffer. Given 5 minutes, I didn't look thoroughly if this is actually exploitable, but it's definitely a red-flag for the code. Be careful out there! Hopefully I am missing something, or this is just a simple oversight, but I would carefully audit this code before using it.
Submitted a ticket through the Althttpd website.
static char StrAppend(char zPrior, const char zSep, const char zSrc){ char zDest; size_t size; size_t n0, n1, n2;
if( zSrc==0 ) return 0;
if( zPrior==0 ) return StrDup(zSrc);
n0 = strlen(zPrior);
n1 = strlen(zSep);
n2 = strlen(zSrc);
size = n0+n1+n2+1;
zDest = (char*)SafeMalloc( size );
memcpy(zDest, zPrior, n0);
free(zPrior);
memcpy(&zDest[n0],zSep,n1);
memcpy(&zDest[n0+n1],zSrc,n2+1);
return zDest;
}How should this happen in practice? The three strings would have to be larger than the available address space...
However, the presence of one piece of code that is not integer-overflow safe definitely makes me nervous. This is just the one I found in 5 minutes, what else is in there?
At core, it's just request-reply + key-value metadata. Whenever it's text, or binary, it does not matter much. But writing HTTP/2 frame types in letters would not make them any easier to understand.
Now, HTTP/2 isn't even conceptually simple, I agree about that... it seems ugly.
well because the hard part is the client-side. caching is a client-side only thing with http, keep-alive is a thing that a server pushes to a client, the same with chunked transfer, which is not as easy to implement for a client like it was with content length.
basically a server does only need to implement certain headers, but a client needs to know all. also most clients even accept bad servers, like content for head requests, etc.. most stateless protocols put a lot of burden into clients.
h2 on the other hand is stateful and keeps the same hard semantics onto the client side and also makes servers more complex, because it's a state machine.
While '408 Request Timeout' is somewhat dubious fig leaf over TCP RST
200 FINE 200 KTHX 404 MISS 403 NO 500 FUCK
Another similar server in one file is busybox httpd command if you are interested https://git.busybox.net/busybox/tree/networking/httpd.c
flex -8iCrfa <<eof
int fileno (FILE *);
xa "\15"|"\12"
xb "\15\12"
%option noyywrap nounput noinput
%%
^[A-Fa-f0-9]+{xa}
{xa}+[A-Fa-f0-9]+{xa}
{xb}[A-Fa-f0-9]+{xb}
%%
int main(){ yylex();exit(0);}
eof
cc -std=c89 -Wall -pipe lex.yy.c -static -o yy045
ExampleYahoo! serves chunked pages
printf 'GET / HTTP/1.1\r\nHost: us.yahoo.com\r\nConnection: close\r\n\r\n'|openssl s_client -connect us.yahoo.com:443 -ign_eof|./yy045This is what it produces for me when I run `lexit.sh us.yahoo.com` - https://stuff-storage.sfo3.digitaloceanspaces.com/ee.txt
The "gibberish" is GZIP compressed data. "yy054" is a simple filter I wrote to extract a GZIP file from stdin, i.e., discard leading and trailing garabage. As far as I can tell, the compressed file "ee.txt" is not chunked transfer encoded. If it was chunked we would first extract the GZIP, then decompress and finally process the chunks (e.g., filter out the chunk sizes with the filter submitted in the OP).
In this case all we need to do is extract the GZIP file "ee.txt" from stdin, then decompress it:
printf "GET /ee.txt\r\nHost: stuff-storage.sfo3.digitaloceanspaces.com\r\nConnection: close\r\n\r\n"|openssl s_client -connect 138.68.34.161:443 -quiet|yy054|gzip -dc > 1.htm
firefox ./1.htm
Hope this helps. Apologies I initially guessed wrong on here doc. I was not sure what was meant by "gibberish". Looks like the here doc is working fine.The compiled program is only useful for filtering chunked transfer encoding on stdin. Most "HTTP clients" like wget or curl already take care of processing chunked transfer encoding. It is when working with something like netcat that chunked tranfser encoding becomes "DIY". This is a simple program that attempts to solve that problem. It could be written by hand without using flex.
Similarly, if you write a server that only speaks HTTP/1.1 (or HTTP/1.0 + Host header), you can put it behind a reverse proxy or load balancer that handles higher versions, does connection management, and terminates TLS. It will work perfectly fine, only without some of the latest performance optimizations that you might or might not even need.
This is even standard practice in many production deployments. Typically you want the proxy or load balancer anyway and there's often little benefit (if any) to using HTTP/2 or HTTP/3 over a very low-latency, high reliability local network.
However even if it didn't, js based client-side apps are probably still attackable in the right set of circumstances.
The Lagrange browser seems quite polished.
I think the solution is to start with a simple protocol and upgrade to more complex protocols after that. While technically you don't need to support HTTP/1 to support 2 and 3, the upgrade over TCP happens mostly in that way.
This is not actually true of HTTP/2 (the upgrade doesn’t use HTTP/1), and slightly misleading of HTTP/3 (there’s nothing to upgrade because it sits beside TCP HTTP; but advertising HTTP/3 support is currently done over TCP HTTP).
Now for the details:
HTTP/2 upgrade is done at the TLS level, via ALPN. Essentially the client TLS handshake says “hello there! BTW do you do HTTP/2?” and the server either responds “hi!” and starts talking HTTP/1, or “hi! Let’s talk h2!” and starts talking HTTP/2.
So it’s perfectly possible (though not a good idea) to have a fully-functioning HTTP/2 server that doesn’t speak a lick of HTTP/1.
(HTTP/2 over cleartext, h2c, does use the old HTTP/1 Upgrade header mechanism, but h2c is more or less just not used by anyone.)
HTTP/3 upgrade, well, “upgrade” is the wrong word. It’s operating over UDP rather than TCP, so you’re not upgrading an existing HTTP thing to HTTP/3, you’re starting a new connection because you’ve learned that the server supports HTTP/3. This bootstraping is currently done by advertising h3 support via the Alt-Svc header (HTTP/1+) or by an ALTSVC frame (HTTP/2+), which the client can then remember so it uses the best protocol next time.
For best results, they’re working on allowing you to advertise HTTP/3 support on DNS, so that after a while it should be genuinely possible (though a very bad idea) to have a fully-functioning HTTP/3 server that doesn’t speak any TCP at all, yet works with sufficiently recent browsers in network environments that don’t break HTTP/3. https://blog.cloudflare.com/speeding-up-https-and-http-3-neg... is good info on this part.
No, it is (some) ASCII plus "opaque octets". ("A recipient SHOULD treat other octets in field content (obs-text) as opaque data.") If you want to say that historically it was such, that's not correct either; it was that plus "supporting other charsets only through use of [RFC2047] encoding" which is a nightmare of an encoding scheme.
> HTTP itself is mostly very simple, text "name: value" pairs separated by newlines
Except when it isn't: headers can be folded across multiple lines. (Which is also obsolete, and discouraged.)
Add to that 1xx responses, transfer encodings, chunks, chunk extensions, trailers. HTTP is far from "simple"…
I really hate when people do that. Static should only ever be const, init once, or something intrinsically singleton. There are very few exceptions.
But if this only ever wants to be an app binary, I guess it's sort of okay.
This program isn't meant to be multithreaded, it's meant to be multi-/process/. The difference is that each instance of the application has its own independent address space. So that's one address space per instance. And because of that there are no threading issues. Also as a result any communication between different processes has to be explicitly defined.
Making this program multi-threaded would be a big mistake because then you couldn't use any of the HTTP proxies out there to monitor for connections and hand them off to this program.
This is arguably a much better architecture from a safety/security standpoint then trying to spin multiple threads each sharing an address space. It forces you to use OS-provided mechanisms for shared state rather than simply manipulating the memory to share information.
From a practical perspective, with a binary protocol, it can be difficult to use across different languages or add support for a new language. If you use the simplest possible encoding, you’d send raw struct data. But this doesn’t always work across different OS/arch/versions/etc. if the server is in C, but the client is in Python, reading the binary protocol would require a far more complicated parser.
Obviously a more formal encoding (protobuf, etc) would be preferred, but if you already need to use an encoding mechanism, why not wrap it in a text format? It’s easier to write clients that can read/write text protocols in any language. The reason why text protocols are so popular aren’t because they are necessarily “better” but easier to adopt. This is why the most popular protocols are text based.
> with a binary protocol, it can be difficult to use across different languages
This is also true of text protocols that aren't well-designed. I don't think it's necessarily the case that binary protocols are more difficult to deal with. You just have a different set of concerns to address.
> If you use the simplest possible encoding, you’d send raw struct data.
This is the "simplest" in the sense that it's definitely easy to just copy this data on the wire, but I think this is a straw man. I don't think it's really any more difficult to write a simple protocol that uses binary data compared to text.
I don’t really think so either… I mean, I’ve done both and it’s really not terrible to use binary. I think text is marginally easier to parse, but once to have the routines to read the right endian-ness, the advantage is minor. As you said, the biggest concern (as always) should be the design. A good design can be implemented easily with either mode.
However, it is significantly easier to debug a text protocol. Attaching a monitor or capturing packets is easier with text as the parsers are much easier and more generic.
All virtual users finished Summary report @ 09:39:57(-0400) 2021-06-08 Scenarios launched: 33645 Scenarios completed: 2573 Requests completed: 2573 Mean response/sec: 42.57 Response time (msec): min: 0 max: 9029 median: 2 p95: 6027.7 p99: 8778.8 Scenario counts: Get index.html: 33645 (100%) Codes: 200: 2573 Errors: ETIMEDOUT: 31008 EPIPE: 48 ECONNRESET: 16
Choose the right tool for the job.
Changing the https://sqlite.org/ website to run off of Nginx or Apache instead of alhttpd would just increase the time I spend on administration and configuration auditing.
But it seems to be "good enough", no? As stated on the page, it serves 500k requests a day.
Were you running your tests using xinetd or stunnel?
Most people don’t need FANG tools.
It absolutely will fail under a DDoS-like punishing load which, say, nginx would have a chance to fend off.
It's still plenty adequate for many real-world configurations and load patterns, much like Apache 1.x has been. Only this is like 2% the size of the Apache 1.x.
It’s not optimized for high ‘performance’. It’s optimized for low resource usage, and the ability to reliably serve large amounts of requests on a small budget, right?
They state that the website is currently serving 500K requests & 50GB of bandwidth per day. Respectfully, this is quite the opposite of your ‘only good for small embedded devices’ claim.
I think this is very interesting, and I’m glad I know this exists now! Worth considering if you have the right type of use case.
My hobby website serves more traffic for a 1/4 of the cost and is easy to configure.
This?
/*
** Test procedure for ParseRfc822Date
*/
void TestParseRfc822Date(void){
time_t t1, t2;
for(t1=0; t1<0x7fffffff; t1 += 127){
t2 = ParseRfc822Date(Rfc822Date(t1));
assert( t1==t2 );
}
}
There's only two billion integers, guess we can test them all. Well, substantially fewer than two billion with that skip. I wonder if that completes in a few seconds. $ time ./althttpd-time-parse
Test completed in 10518961 us
real 0m10,521s
user 0m10,486s
sys 0m0,004s
The "Test completed" line is from my main() "driver", I wanted to measure time inline too and the measurements seem to agree.This is on a Dell Latitude featuring a Core i5-7300U at 2.6 GHz, running Ubuntu 20.10.
IMO most tests are fundamentally flawed. The way most testing is done it would be easier and better to just write everything twice, have a method to compare results, and hope you got it right at least once.
I was curious to see how my M1 compares to my intel 2019 macbook pro:
M1:
/tmp/tt 8.42s user 0.01s system 99% cpu 8.435 total
2,6 GHz 6-Core Intel Core i7
/tmp/tt 15.69s user 0.03s system 98% cpu 15.888 total
/tmp/tt 15.69s user 0.03s system 98% cpu 15.888 total
It fails an assert if the parse doesn't work I guess?
That the sqlite website is able to run this way is more a testament to Linux's work on a lightweight/fast fork() than anything else. This would perform terribly on a more traditional Unix.
Forking for every request is slow, sure.
But if your code is written with it in mind it's faster than most people might expect, and most people never get to a scale where it matters.
It's not the right choice for everything, but people have ironically gotten obsessed with things we introduced a long time ago as workarounds for slow hardware (and fork used to be slow on Linux too) decades after the original problems were largely solved.
I do agree there are times this won't be useful, though.
Yes, but that met expectations of that time period, and expectations for a webmail service. I'm curious if you also forked for every static asset...that's what this setup appears to do.
I just don't see the benefit of mysql choosing to use this today. It works, but there are other minimal http servers that would be just as simple, but would be faster and use fewer resources. I suppose they don't need to change it, but it's not really a great example of anything other than "fork is cheap on linux" to me.
We didn't fork for for every static asset, but the vast majority of overall requests were dynamic past the initial pageload, so the vast majority of requests resulted in fork.
In terms of benefits, the simplicity is attractive. It's an approach that is in general quite resilient to errors.
As peer poster said, expectations years ago were higher than today. It is frustrating how common it is for sites to think that downloading multiple MB of code just to show a simple page and having it take seconds to render is somehow ok.
The expectations used to be ~100ms for a page load and render back when I was working on high performance web servers (~15 years ago).
- Uses sendfile() on FreeBSD, Solaris and Linux
- Event loop, single threaded - no fork() or pthreads
- Supports If-Modified-Since, Keep-Alive, IPV4, 301 redirects
And appears to be just a little larger than Althttpd. Sounds good to me.
static const char *azDisallow[] = {
"skidrowcrack.com",
"hoshiyuugi.tistory.com",
"skidrowgames.net",
};
Anyone know why? }else if( strcasecmp(zFieldName,"Referer:")==0 ){
zReferer = StrDup(zVal);
if( strstr(zVal, "devids.net/")!=0 ){ zReferer = "devids.net.smut";
Forbidden(230); /* LOG: Referrer is devids.net */
}
Which I can appreciate why it's there, it's still odd.Not sure why they're blocked
Caddy is just really awesome as a reverse proxy (2 line config!!) and I am in the processes of moving all my projects to it. It is fast enough as well since other things will be the bottle neck way before that.
I am not affiliated with Caddy in any way, just blown away by the quality of it.
This might solve my problem with older servers that no longer support the latest SSL.
I really need to upgrade those rickety old machines.
Runtime dependencies create a nuisance as you have to update several things together. On the other hand, they can allow components with separate update cycles and responsibilities to be update separately.
Build dependencies create maintainability and security problems. They can also solve maintainability and security problems. It depends on what your consideration is. But as a matter of practice, many developers seem too concerned with possible behavioral/API breakage, that they like to pin to specific versions of their dependencies, which now means that you aren't getting any security fixes.
(Technically, Althttpd doesn't achieve zero runtime dependencies in comparison to a modern http server that does HTTPS, because it requires a separate program to terminate TLS. But these connect through general mechanisms that are much easier to combine and update separately.)
Everyone has to make a judgement about how they maintain their own systems, but being excited about "zero (runtime) dependencies!" isn't the way the judgement concludes.
It feels like then I'd probably need either shared storage for the certificate files (which goes against the idea of decentralization somewhat) or to use a DNS challenge type.
Anyone have experience with something like that?
I'm doing this exact thing, with the Redis plugin behind DNSRR and it works seamlessly.
10 minutes of caddy, I had everything running exactly as I wanted and the job was done.
I don't think I'd choose to use it again. Instead, I'll try Caddy, or HAProxy if I need massive performance.
For very tiny hooks, you might be able to get away with using request matchers[1] and respond[2].
[0]: https://caddy.community/t/missing-starlark-documentation/958...
[1]: https://caddyserver.com/docs/caddyfile/matchers
[2]: https://caddyserver.com/docs/caddyfile/directives/respond
But yeah, it's still something at the back of our minds, and we were considering Starlark for this, but that hasn't really materialized because it's usually easier to just go with the plugin route.
It makes you wonder just how "heavy" operating system processes actually are. We may not need to worry about the complexity of trying to multiple run async requests in a single process/thread in all cases.
If you only have around 100 concurrent confections, a separate thread per connection is entirely feasible. A whole new process is probably fine on Linux, but e.g. Windows takes pretty long to spawn a process
Isn't it an appealing model to not even have to talk about threads because every process is 1 thread by definition?
It's also more secure because you should not mix different users' requests in the same process if you can avoid it.
nginx runs on a one process per core model and more or less does everything correctly.
Granted, 10k concurrent requests is a problem for the 1% of websites, so processes were (and still are) good enough for the long tail of personal or small-scale websites.
Not a problem for most folks, but when you want the greatest possible performance, you want to avoid these kinds of transitions. Basically, the same reason some folks use user-space networking stacks.
a lot of people just did not get the difference between concurrency vs. parallelism. threads and processes are basically parallelism, while concurrency is async programming. good talk about that stuff from rob pike (go): https://www.youtube.com/watch?v=oV9rvDllKEg
[0] https://gist.github.com/nicowilliams/a8a07b0fc75df05f684c23c18d7db234I was not familiar with darkhttpd. Both of these are similar to the sqlite server in security design (chroot capability), but unlike in that a single process serves all requests and does not fork.
I have used stunnel in front of thttpd, and chrome has no complaints.
Justine Tunney is a treasure!
I was curious about the number of lines and calculating characters was a simple select all from there. There are 2,592 lines.
Also, I'd like to complain about the hn hug of death, because it isn't happening.
** May you do good and not evil.
** May you find forgiveness for yourself and forgive others.
** May you share freely, never taking more than you give.wow. Thanks also for elaborating the xinetd & stunnel4 configs.
I mostly use C#, and a while back I settled on a middle ground, where closely-related classes and interfaces are grouped together in a single file.
When I'm working on web apps/APIs, I usually follow the "feature folder" concept too, where all the most central parts are together in the same file.
althttpd -exec some_executable {}.method {}.body
So you could quickly call executable from a browser and redirect the output in the response.
What have I missed?
I don't know what the landscape was like in 2004 really, but probably at least an order of magnitude less than today's bazillion (whatever that would be!).
On a more personal note, wow! I had no idea I started using the Internet for realz before the release of Apache, in 1994. This young made me feel, not.
It also isn't brand new: it's been around since 2004. So that probably narrows the range of possible competitors even more.
If you can find a webserver that meets all of those constraints, please let us know.
I wrote it because no other web server could serve files fast enough on my system (not lighttpd, not nginx, not Apache httpd, not thttpd) to keep movies from buffering.
Could you expand on that? What type of files, how many clients? I seem to recall plain apache2 from spinning rust streaming fine to vlc over lan - but last time I did that was before HD was much of a thing... Now I seem to stream 4k hdr over ZeroTierOne over the Internet to my Nvidia Shield via just DLNA/UpNP (still to vlc) just fine. But I'm considering moving to caddy and/or http/webdav - as a reasonable web server with support for range request seem to handle skipping in the stream much better.
This was for serving MPEG4-TS files with, IIRC, H.264 video and MPEG-III audio streams -- nothing fancy -- from a server running a container living on a disk attached via USB/1.1.
While USB/1.1 has enough bandwidth to stream the video, the other HTTP servers were too slow with Range requests, because they would do things like wait for logs to complete and open the file to serve (which is synchronous and requires walking the slow disk tree).
Ah, ok. That makes sense. USB 1.1 can certainly challenge cache layers and software assumptions.
I do wonder how far apache2 might have been pushed, dropping logs and adjusting proxy/cache settings.
Back then, there weren't bazillion web servers out there. A patchy server was still.. patchy. And engine that solves problem X (c10k) was not created yet :)
(for whoever reads my comment, I am referring to Apache and nginx)
But hang on is dependency anxiety really the reason or did you just make that up?
In 2001?
One file is easy to add into a project, and the compiler optimizes translation units better, so you get a bit of a performance increase in some cases.
Having "Find symbol in file" is nice too if you know you are looking for it just in this one file related to the code. Most editors aren't as ergonomic for finding "symbol in current directory" as they are for "symbol in current file".
But now I want to see how badly you could Uncle Bob this thing. My screen should be wide enough for the resulting function names.
Comparing to purely event based web servers forking can still be better as no request should fully block another, or usually crash another (which is more likely with threads), and thread or fork based servers can make better use of concurrency which is significant for CPU heavy jobs.
So swings & roundabouts. Each type (event, thread, process, some hybrid of the above) has strengths, and of course weaknesses.
That's because internally it's nearly the same thing. Both forking and starting a new thread on Linux is a variant of the clone() system call, the only difference being which things are shared between parent and child.
It's the distributed version control (and more) used by SQLite. Most people have no idea about how cool the SQLite ecosystem is, and how it's used even on avionics!
No, having a commit "fix typo" in the main branch's history is not at all useful and won't ever be. It's noise.
In a work setting it's much better to reduce noise.
The difference is that Fossil does not promote the use of commit-squashing. While it can be done, it takes a little work and knowledge of the system. Consider the premature-merge problem in which a feature branch is merged into trunk before it is ready, and subsequent typo fixes need to be added. To do this in Fossil you first move the errant merge onto a new branch (accomplished by adding a tag to the merge check-in) then fix the typo on the original feature branch, then merge again. So in Fossil it is a multi-step process. Fossil does not have a "rebase" command to do all that in one convenient step. Also, Fossil preserves the original errant check-in on the error branch, rather than just "disappearing" the check-in as Git tends to do.
The difference here is a question of priorities. What is more important to you, an accurate history or a clean history that tells a story? Fossil prioritizes truth over beauty. If you prefer a retouched or "photoshopped" history over an auditable record of what really happened, Fossil might not be the right choice for you.
To put it another way, Fossil can squash commits, but another system might work better for you if commit-squashing is your go-to method of dealing with configuration management problems.
That isn't the normal use case for commit squashing though. Generally, trunk/master/main isn't ever rewritten. Squashing is usually done on feature branches _before_ merging. What does that look like in fossil?
It seems like part of the problem is that fossil is designed for a very different workflow. See https://www.fossil-scm.org/home/doc/trunk/www/fossil-v-git.w.... The autosync, don't commit until it is ready to be merged workflow might work well for a small flat organization, but I'm not sure how that scales to large organizations that have different privilege levels, formal review requirements, and hundreds or thousands of contributors.
In terms of Git and the usual commercial SCM practices, I'm speaking empirically. Everywhere I worked in a team, leaders and managers wanted main branch's history to be a bird's-eye view, to have every commit fully build in CI/CD, and be able to find who introduced a problem (admittedly this requires a little more digging compared to Fossil, though). Squashed commits help when auditing for production breakages, and apparently also helps managers do release management (as well as billing customers sometimes).
Do I have all sorts of minor commits in my own projects? Sure! I actually pondered using fossil for them but alas, learning new tools just never gets enough priority due to busy life. I'd love to learn and use it one day. I'm sick of Git.
But I don't think your analogy with a photoshopped / retouched picture is fair. Squashing PRs into a single commit is not done for aesthetic reasons or for deliberately disappearing information -- a link to the original PR with its branch and all commits in it remain and can be fully audited after all. No information actually disappeared.
I believe a better analogy would be with someone who prefers to have one big photo album that contains smaller albums which in turn contain actual photos of separate life events that are mostly (but not exactly) in chronological order -- as opposed to Fossil's approach which can be likened to a classic big photo album with all semantically unrelated photos put in strict chronological order.
I'll reiterate that my observations and opinions are mostly empirical. And let me say that I don't like Git at all. But the practice I described does help in a classic team of programmers and managers.
I concede that Git and its quirks represent a local maxima that absolutely can be improved upon, but at least to me the jury is still out on what's the better approach -- and I'm not sure a flat history is it.
I frequently use git rebase in interactive mode to rearrange and curate my commits to form whatever narrative I'm aiming for. Commits are semi-independent stories which can be merged, in order, at any rate and still make sense. Each commit makes sense with respect to history, but doesn't care about the future.
I squash and rearrange and fixup commits until they look they way I want, and would want to see if I was looking at a history, and then send them for review.
Whether you merge my branches, or fast-forward and rebase the individual patches, makes little difference to me. But please don't squash my hard work.
Squashing buys you next to nothing, and costs you the ability to dive into the history in greater detail.
I suppose if your project is truly huge, it becomes worth it to reduce load on your VCS, but beyond that...
The "greater detail" part can cost me a lot of time.
Sure it does, but sometimes that level of detail in history is not helpful. Individual keystrokes are an even finer/"more accurate" representation of history; but who wants that? At some point, having more granular detail becomes noise - the root of the disconnect is that people have a difference in opinion on which level that is: for some (like you), it's at individual commit-level. For others (like me), it's at merge-level: inspecting individual commits is like trying to parse someone's stream-of-consciousness garbage from 2 years ago. I really don't care to know you were "fixing a typo" in a0d353 on 2019-07-15 17:43:32, but your commit is just tripping-up my git-bisect for no good reason.
I would say what you're describing is a break down in CI/CD and code review. How is code that is that broken getting into your default branch in the first place?
As to rebases to clean up history (and not just the PR itself)... personally, I don't think that's worth it. My experience with history like this is that it's relevant during review, and then around 95% of it is irrelevant - you may not know which 5 % are relevant beforehand, but it's always some small minority. It's worth cleaning up a PR for review, but not for posterity. And when it comes to review, I like commits like "pay respect to the linter gods" and the like, because they're easy to ignore, whereas if you touch code and reformat even slightly in one commit, it's often harder to skip the boring bits; to the point that I'll even intentionally commit poorly formatted code such that the diff is easy to read and then do the reformat later. Removing clear noise (as in code that changes back and forth and for no good reason) is of course nice, but it's easy to overdo; a few typo commits barely impact review-ability (imho), and rebases can and do introduce bugs - you must have encountered semantic merge conflicts before, and those are 10 times as bad with rebases, because they're generally silent (assuming you don't test each commit post-rebase), but leave the code in a really confusing situation, especially when people fix the final commit in the PR, but not the one were the semantic merge conflict was introduced, and laziness certainly encourages that.
It also depends on how proficient you are with merge conflicts and git chicanery. If you are; then the history is yours to reshape; but not everybody is, and then I'd rather review an honest history with some cruft, rather than a frankenstein history with odd seams and mismatched stuff in a commit that basically exists because "I kept on prodding git till it worked".
All due diligence is done there, not in the main branch.
The main branch only needs to have one big commit saying "merging PR #2169". If you need more details you'll go that PR/branch and get your info.
The "fix typo" commit being in the main branch buys you nothing. It's only useful in its separate branch.
Why not merge?
It gives you better high-level observability and a good bird's-eye view. And again -- if you need more details you can go and check all separate commits in the PR/branch anyway.
And finally, squashed commits are kind of atomic commits. Imagine history of three separate PRs happening at roughly the same time. And now all those commits are interspersed in the history of the main branch.
How is that useful or informative? It's chaos.
EDIT: my bad, I conflated merging with rebasing. Still, I prefer a single squashed commit for most of the reasons above, plus those of the other two commenters (useful git blame output and buildable history).
"Encrypt project's sensitive fields (#1234)"
With the number being a PR or an issue # (which does contain a link to the PR).
I do care about history in branches though. And many others do. I agree that it varies from team to team.
Also, in case it helps you in the future, `blame -wC` is what I use when doing blame; it ignores whitespace changes and tracks changes across files (changes happened before a rename, for example.)
I've come across "fix indentation" or "fix typo" commits where a bug was introduced, like someone accidentally comitted a change (maybe they were debugging something, or just accidentally modified it).
For example: I'm tracing a bug where a value isn't staying cached. I find a line of code DefaultCacheAge=10 (which looks way too short) and git blame shows the last change was modifying that value from 86400. What I do next will be very different if the commit message says "fix indentation" vs "added new foobar feature" or "reduced default cache time for (reason)".
As I replied to @SQLite above, I am not saying this is the optimal state of affairs -- not at all. But it's what is required everywhere I ever worked for 19.5 years (and similar output was desired when we worked with CVS and Subversion and it was harder to achieve there).
But I'll disagree this is lack of discipline. It's not that at all. It's a compromise between programmers, managers/supervisors, CTOs / directors of engineering, and release engineers. They want a bird's-eye view of the project in the main branch.
Sometimes those insights are just as useful as the packaged whole.
You surely aren't going to have 50 contributors all simultaneously working on the same mainline or feature (or have 50 working together at the same time if ever, if you do, I'd say that's poor project management). The reality is a portion of the developers work on this feature in this part of the code base, a few over here, they'll be on their own branches or lines, and everything will be fine.
This scenario where we have 50+ devs all crashing and bumping into each other rarely happens, if ever. I have personally only seen one instance of it happen in kernel development. And even then it was relatively straightforward to sort out.
To go further, in a hypothetical scenario where there is one feature and 50+ open source developers are all vying to push their patches in, there is still going to be one reference point to work off of, and reviewers are going to base everything off that. It's a sequential process, not concurrent.
And since that day, not one looked at Fossil the same way, at least not the same way they looked at git
https://www.mail-archive.com/fossil-users@lists.fossil-scm.o...
I can't tell for sure how much of impact this had on fossil's adoption, its hard to beat git no matter how good you are, but I think it was a bit hit
1) No, it didn't. Please read the part of the thread after Zed's initial panic attack.
2) To the best of our[1] knowledge, fossil itself has never caused a single byte of data loss. When fossil checks in new data, it reads that data back (in the same SQL transaction the data was written in) to ensure than it can read what it wrote with 100% fidelity, so it's nearly impossible to get corrupted data into fossil without going into the db and massaging it by hand. People have lost data by storing their only copy of a repository on failing/failed storage or on a network drive, but no software can protect against hardware failure and nobody in their right mind tries to maintain an active sqlite db over a network drive (plenty of people do it, despite the repeated warnings of anyone who knows anything about sqlite, and they have only themselves to blame when it goes pear shaped). Fossil makes syncing to/from a remote copy absolutely trivial, so any failure to regularly sync copies to a backup is end-user error.
[1] = the fossil developers.
> Let me put it this way. Suppose you’ve got Zed Shaw. No, wait, say you’ve got “a person.” (We’ll call this person “Hannah Montana” for the sake of this exercise.) And you look outside and this young teen sensation is yelling, throwing darts at your house and peeing in your mailbox. For reals. You can see it all. Your mailbox is soaked. Defiled. The flag is up.
> Now, stop and think about this. This is a very tough situation. This young lady has written one of THE premiere web servers in the whole wide world. Totally, insanely RFC complaint. They give it away on the street, but everyone knows its secretly worth like a thousand dollars. And there was nothing in that web server that hinted to these postal urinations.
Still an impressive bug – but the first rule of “my CVS has broken” is “stop running commands, and copy the CVS directory somewhere else”. (I've needed to do this to a .git twice.) Had he done this, he wouldn't've lost work. While he shouldn't've had to, “stop doing everything and take a read-only copy” is the first step when any database containing important data has corrupted.
I finally gave Fossil a serious try last year for a few months, just on my own but I don't think my opinions would change if I tried in a team setting. I still love the idea, but the execution is... ghetto. Serviceable, certainly, but ghetto is the best overarching description I have for it. You have to be willing to look past a lot of things, sure many of them are petty, in order to take it over dedicated polished services for the things it combines. And git+[choice of issue tracker]+[choice of forum]+[choice of wiki]+[choice of project website]+etc. is the real competition against Fossil, not git+nothing, so even if it was better on the pure version control bits it would still be a tough battle. (Not to mention the elephant git+github is a pretty good kitchen sink on its own if you don't want to think about choices and have something serviceable that's also a lot less ghetto.)
I also realized how much I love git's staging area concept once it was gone -- even when I had to use Perforce a lot, at least Perforce has the concept of pending changelists so you have something similar. I've never been a big fan of git's rebase, but it's also brought up a lot as a feature people are unwilling to give up, and I see the appeal. In summary, I think the adoption issue is just that people who do eventually give it a shot find usability issues/missing functionality they aren't willing to put up with.
I think that this is only true for some projects. For some people, the ease of self hosted setup (one executable) and the fact that you can change the documentation, edit code and close bugs offline is a big win that no centralized service can compete with.
For project management I'm less familiar with the options out there but I'd be surprised if there was nothing that gives a really stellar offline experience. I'd give a preemptive win to Fossil on the narrow aspect that your issue tracking changes can be synced and merged automatically with a collaborative server when you come back online, whereas if you stood up your own instance of Trac for instance I'm not sure if they have any support for syncing. If you're working by yourself, though, then there's no problem, Trac and many others work just like Fossil and stand up a local server (or are dedicated standalone programs) and work the same whether you're offline or online. But when I'm working solo I prefer low-tech over anything that resembles Jira (and I don't even really dislike Jira) -- I've played with https://github.com/dspinellis/git-issue as another offline/off-platform option but in my most recent ongoing solo project I'm quite happy with a super low-tech issues text file that has entries like (easy to make with https://github.com/dhruvasagar/vim-table-mode)
+--------------+
| Add thing |
+==============+
| Done whens / |
| other info |
+--------------+
and when I'm closing one I just move it to the issues-closed file as part of the closing commit. I might give it an identifier if I need to reference it in the code/over multiple commits.> I just cloned your repo. Everything is working fine. Breath, Zed.
seems like he panicked and made it worse, then rage-quit
I understand both that he is and why he is such a polarizing figure. Just wanted to put my positive anecdote on the pile, since they seem to be a less common when he comes up in online comments.
My personal website, https://dbohdan.com/, is powered by Fossil. A year ago I was shopping for a wiki engine, didn't love any I looked at, and realized I could try something I was already familiar with: Fossil. It did take a few hacks to make it work how I wanted. The wiki lacks category and transclusion features and, at least for now [1], can't generate tables of contents. I've invented a simple notation for tags and generate a "tag page" [2] using a Tcl script [3]. The script runs every time I synchronize my local repository with dbohdan.com. The TOC is generated in IE11-compatible JavaScript in the reader's browser [4]. The redirects are in the Caddyfile (not in the repo). Maybe I'll migrate to a more full-featured wiki later [5], but I am enjoying this setup right now. I am happy I gave Fossil a try.
Fossil also has a built-in forum engine [6]. I am thinking of migrating a forum running on deprecated software to it.
Edit: My favorite music page and sitemap are generated on sync, too. [7] The sitemap uses Fossil's "unversioned content" feature to avoid polluting the timeline (commit history). [8]
-----
[1] In the forum thread https://fossil-scm.org/forum/forumpost/b635dc56cb?t=h DRH talks about implementing a server-side TOC.
[2] The page lists the tags and what pages are tagged with each. Tags on other pages link to their section of the tag page. https://dbohdan.com/wiki/special:tags.
[3] https://dbohdan.com/artifact/8297b54f5d
[4] https://dbohdan.com/artifact/d81bb60a0e
[5] PmWiki seems like a nice lightweight option—an order of magnitude less code than its closest competitor DokuWiki, very stable, and has a better page history view. Caveat: it is written in old school PHP. https://pmwiki.org/.
[6] https://fossil-scm.org/home/doc/trunk/www/forum.wiki
[7] https://dbohdan.com/wiki/music-links with https://dbohdan.com/artifact/053d0ff993, https://dbohdan.com/uv/sitemap.xml with https://dbohdan.com/artifact/c21444f7c9.
Just FYI: we recently improved the internals to be able to add propagating tags to wiki pages[1], so it will eventually be possible to use those to categorize/group your wiki pages. What's missing now is UIs which can make use of that feature. The CLI tag command can make use of them, but that doesn't help your UI much.
> ... and transclusion features
For the wiki it seems unlikely to me that transclusion will ever be a thing. It can hypothetically be done with the embedded docs feature if the fossil binary is built with "th1-docs" support, but, alas, we can't currently support propagating tags on file-level content. (i have an idea how it might be integrated, but figuring out whether or not it internally makes sense requires trying it out (and that doesn't have a high priority).)
As for transclusions, I don't expect Fossil to implement them. While something like https://www.pmwiki.org/wiki/PmWiki/IncludeOtherPages would be cool, it seems probably out of scope for Fossil.
Side note: I've been exploring the ecosystem around Fossil/SQlite as well. I've been working with Pikchr a lot recently as a way to create diagrams that can be version controlled. Because Pikchr is implemented as a single C file, I was able to compile using Emscripten as a WASM file, and embed that file in a single HTML page that gives me a "live editing" experience (basically I call a render method when the text area gets updated). The way these pieces of software are written minimizes dependencies and allows for them to be used in a huge variety of environments. I've really enjoyed working with them.
Interesting. If the load avg is consistently low, it could mean they're over-paying for CPU. If this was a non-dedicated AWS instance you might want low load so you don't chew up CPU credits, but you'd also want to use an instance type that creates some load so you're utilizing what you're paying for. Linode VPCs don't use cpu credits so the calculation is a bit simpler. I'm also curious how much of that bandwidth couldn't be offset by a CDN or mirrors.
If you were using a serverless platform, you'd ideally want to use something like static site hosting feature where you're mostly just paying for storage and egress. Or a serverless application platform to auto-scale traffic as needed. The main problem with doing this, of course, is the cost of egress: cloud providers with fancy serverless platforms often have redonkulous egress costs, so using a plain old VM on a VPC provider can be cheaper if you have more bandwidth demands than compute.
(I am aware none of this is a concern if you'd rather just spend $40 and forget about it. I am a nerd.)
As you say, at $40/m it's academic for a lot of people, but AFAIK, the whole site is static, so presumably if you put it behind Cloudflare’s free tier it would serve all but the file downloads from the edge. A pure guess, but I'd imagine that would mean serving 75% of requests from the edge.
The core SQLite website has a lot of static content, but there are dynamic elements, such as Search (https://www.sqlite.org/search?s=d&q=sqlite) and the source code repository (https://www.sqlite.org/src/timeline?n=100&y=ci).
So far today, 23.48% of HTTP requests to the sqlite.org domain are for dynamic content, according to server logs.
If it was written in Haskell or Rust for example you could be more sure about correctness. I believe for something like this, correctness is fairly important. Not to mention you probably wont even lose speed. As for if the code is understandable, it is a 2600 lines of terse-ish C [0]. Do you really think about the entire blob at the same time?