Old box, dumb code, few thousand connections, no big deal
rachelbythebay.com
rachelbythebay.com
It might sound like I'm trying to steal her thunder, but mostly what I'm trying to say is she is right. Listen to her. Here is further evidence that she is right.
As I wrote in https://gitlab.com/kragen/derctuo/blob/master/vector-vm.md, single-threaded nonvectorized C wastes on the order of 97% of your computer's computational power, and typical interpreted languages like Python waste about 99.9% of it. There's a huge amount of potential that's going untapped.
I feel like with modern technologies like LuaJIT, LevelDB, ØMQ, FlatBuffers, ISPC, seL4, and of course modern Linux, we ought to be able to do a lot of things that we couldn't even imagine doing in 2005, because they would have been far too inefficient. But our imaginations are still too limited, and industry is not doing a very good job of imagining things.
—
⁰ http://canonical.org/~kragen/sw/dev3/server.s
¹ http://canonical.org/~kragen/sw/dev3/httpdito-readme
² https://news.ycombinator.com/item?id=6908064
† It's actually bloated up to 2060 bytes now because I added PDF and CSS content-types to it, but you can git clone the .git subdirectory and check out the older versions that were under 2000 bytes.
As a self-taught programmer I would say that what all these less efficient bit easier to learn technologies have done is enable people like me who evidently are not geniuses like yourself to write software. Should programming always be an ivory tower thing?
It’s fair to point out how far you can get just programming one computer using traditional and well understood concepts like sockets and threads. And how weird it is that we live in a world where Kubernetes is mainstream and fun but threads are esoteric.
Adding unnecessary complexity doesn't always make things easier to learn.
Can you elaborate on what this means exactly? For example, is there some reasonable C code that runs 33 times slower than some other ideal code? In what sense are we wasting 97% of our computer's computational power?
Includes a 9-level nested for loop, which is always great to see.
Roughly that 3000× is 18× from multithreading, 3× from SIMD instructions, 15× from tuning access patterns for locality of reference, and 3× for turning on compiler optimization options. This is a really great slide deck!
I was assuming "single-threaded nonvectorized C" already had compiler optimization turned on and locality of reference taken into account. As the slide deck notes, you can get some vectorization out of your compiler — but usually it requires thinking like a FORTRAN programmer.
So I think in this case reasonable C code runs about 54× slower than Leiserson's final code. However, you could probably get a bigger speedup in this particular case with GPGPU. Other cases may be more difficult to get a GPU speedup, but get a bigger SIMD speedup. So I think my 97% is generally in the ballpark.
A big problem is that we can't apply this level of human effort to optimizing every subroutine. We need better languages.
- http://spcl.inf.ethz.ch/Research/DAPP/
- http://tiramisu-compiler.org/
This way you can have a researcher implementing the algorithm (say bilinear filtering) and a HPC expert who tunes it with parallelism, SIMD, tiling.
I wrote an overview of most DSL for high performance or image processing in this issue: https://github.com/mratsim/Arraymancer/issues/347#issuecomme...
I'm still kind of a newb myself but from what I understand these are special CPU instructions that allow you execute the same instruction in parallel against multiple data points. This allows you to eke out a lot more performance. It's how simdjson[1] is able to outperform all other C++ json parsers.
It's pretty variable: some things we haven't figured out how to speed up with SIMD, sometimes we have a GPU, sometimes we can get 8 or 16 SIMD lanes out of SSE3 or AVX128 or 32 of them out of AVX256, sometimes you only have four cores, sometimes make -j is enough parallelism to win you back the 8× factor from the cores (though not SIMD and GPGPU). But I think 97% is a good ballpark estimate in general.
It's also not reasonable to expect to vectorize everything; for example a web server is unlikely to benefit from vectorization.
Fifteen years ago we thought regular expression matching and compilers were unlikely to benefit from vectorization, but now we have Hyperscan and Co-dfns, so they did. So I think it's likely that we will figure out ways to do a wider variety of computations in vectorized ways, now that the rewards are so great.
As an example, if you are checking a boolean flag (1 bit) on an object, and it ends up being a cache miss (and x86_64 cache line size is 64 bytes), then your computer just went through all the expense of pulling in 512 bits from RAM yet it only used 1 of them. You are achieving 0.2% of the machine's possible throughput.
It's totally reasonable for a company to choose the Python/Gunicorn option if they already have a bunch of people who know Python and they don't need to serve tons of requests per second.
Even if they do need to serve tons of requests per second, it's totally reasonable for them to still choose Python/Gunicorn if the cost of the additional servers is less than the cost of having to support multiple languages. Or if they get a lot of value from libraries that are unique to the Python ecosystem. Or if they care more about quickly iterating on features than driving down server costs.
I agree that there's a point where it stops making sense, and there are plenty of engineers who don't recognize when they're past that point because they keep doubling down on sunk costs and things they're familiar with. But let's not be too quick to assume people are in that camp when we don't know all the tradeoffs they're facing.
Is what the author did supposed to be impressive? Is it supposed to make python look bad?
I don't get it. Seems like run of the mill stuff. Python might struggle at the same level of concurrency (was it like 15k?) but you can still do 10k connections easy enough iirc.
Just having a lot of connections that do trivial stuff is not very difficult. It becomes way more interesting when all of those connections need to access shared data structures and whatnot.
Async is typically less prone to error and complexity than thread code, and also typically faster/lighter for io (in python). I can't see a reason for this not to be the case in other languages too.
In terms of error-proneness, I would say the hierarchy is this:
Actors < CSP ~ Async << Threads
In terms of "getting started" difficulty, I would put the order like this:
Async < Actors ~ CSP < Threads
In terms of first-order maintainability in large projects I would put the order like this:
Actors >> CSP ~ Threads > Async
Async is on its face easy to grok, and saves you a ton of problems with locking and the like, and you can run it on a single threaded system with the right type of dispatcher. However, small amounts of complexity rapidly devolve into towers of async calls that cannot be untangled from a spaghetti pile, and there's not really satisfying ways of figuring out how to unwind error-handling in async, except for the basic case of "I truly don't care if this async fails".
I will say though, it's spectacularly easy to write messy and garbage code in all four of these concurrency models (I certainly have). My preference for actors comes from four opinions:
- some systems (if you want to be flippant, microservices architecture) use actors to encapsulate failure domains, which is really fantasic, and truly the #1 reason to use actors
- 99.8% of the time no need to write mutexes, and good actor systems basically won't deadlock unless you really try hard. (also true for async system btw)
- gives you an organizational framework to write well-designed and well-engineered systems.
- with a small amount of discipline "not having a spaghetti ball" scales with complexity (I find it takes a lot of discipline to not have a spaghetti ball with async in the more complex cases)
I can't even say the same for synchronous Django. Sometimes it's just the quality of the tool you're using, not the higher-level concept it implements.
Okay. Get over it.
Python is fine. Threads are fine. Go is fine. Clojure is fine. Java is fine. PHP is ... okay I won't go that far. :)
You fix the problem when the cost required to fix the problem is finally less than the opportunity cost of fixing something else.
I've mainly worked at companies that would be bankrupt if they used Python the way Instagram does.
You need a market. You need paying customers. You need cashflow. You need features. You need a business plan.
You don't need scaling. Ever. To first, second and third order approximations.
Your company has a higher probability of bankruptcy than needing scaling.
It's a resounding no from me. If you're stuck with PHP, sure, this beats the WordPress style of...well...I can't really find words to describe how bad it is..., but it's still miles behind anything else. It full of strings and other things that you just have to memorize with IDE integration even with specialised extensions still worse than other languages with standard IDEs.
'gender' => 'in:male,female,other'
You can now do: 'gender' => Rule::in(['male', 'female', 'other'])
Which, coming from Go, I very much prefer.Modern PHP is pushing for type-safety (via type-hinting) and Laravel is following this direction as well.
PHP has grown up a lot over the last ten years.
It's too stringly-typed for my tastes, seems like a lot of errors you can make will only be caught at runtime.
How hard is it to get up to speed on any other tech stack? ASP.NET Core is extremely fast and the learning curve is close to none, for example.
If someone was able to wrap his head around backend development with Python I'm pretty sure they have the mental fortitude to onboard a tech stack that doesn't suffer from major performance problems.
Now you have to figure out how to run a CI on a build server. Deployment on your production systems, be it containers or even not. Monitoring, alerting, profiling, tuning under load. You have to support a new database connection library with new shenanigans. In general, you need to integrate the new language into the existing ecosystem. The latter may even be impossible depending on the stack and solution chosen. Just look up those weird Java only caching servers.
All of that is possible, sure. Due to company acquisitions and specialized teams in some areas, we're kinda running the full bingo card of languages.
But there's little denying: We overall spent less time handling language runtimes and language-specific monitoring back when we were java+mysql and that's it.
Most if not all of those items are either trivial or non-issues.
In fact, I would argue that deploying a Python app is a far more convoluted process than getting an ASP.NET Core app up and running.
With Docker, the problem simply disappears.
Getting it to build and test on a CICD pipeline is as hard as typing $ dotnet build, or $ dotnet test.
> All of that is possible, sure.
Not only it is possible, it's laughably easy.
We are supposed to avoid our problems, not perpetuate and aggravate them as eternum because we are too lazy to look for ways to make our life easier.
(On the plus side, seems my Jetbrain kit includes PyCharm, so I've even got an IDE!)
Or a pool of asynchronous "green threads" using any of a number of libraries (gevent is the one I'm most familiar with). The thing to avoid is mixing the two (multiple processes and green threads).
> A python webserver is more optimal when used in multiprocess configuration
For CPU bound worker tasks, yes, this is true. For I/O bound applications, not so much; asynchronous I/O can handle the same I/O load with much fewer resources (particularly as forking Python processes uses a lot more memory because so much of the memory in the Python interpreter is dynamic, so you don't get a lot of benefit from what in a compiled language would be shared read-only code that doesn't need to be copied for every fork).
> (as opposed to multithreaded configuration, where python sorely suck at)
Yes, the limitation of the GIL is one of Python's worst warts. (It looks like there are finally efforts to remove it, but it's taken a long, long time.)
I actually don't mind the heavy framework thing because .net core mvc is the same deal. The developer experience in python was just way worse for me. Autogenerating swagger docs didn't seem possible without manually adding annotations to all my code (for flask at least), the battle-tested python libraries for web dev aren't async, mypy is obviously way worse than an actual type system, model validation (the thing forms/serializers handle) was more tedious, vscode was worse than visuals tudio.
Flask is lighter and it is more DIY but easier if you have something that is more generic.
"The developer experience in python was just way worse for me. Autogenerating swagger docs didn't seem possible without manually adding annotations to all my code (for flask at least), the battle-tested python libraries like django and sqlalchemy aren't built for async, mypy is obviously way worse than an actual type system and having static types really helps with understanding a large new code base, model validation (the thing forms/serializers handle) was more tedious, vscode was worse than visual studio (code navigation was very hard to do in python. I could jump 1 layer into library code, but when I tried to jump deeper vscode couldn't find anything.)"
I've seen people say you should only choose rails or django for a startup, but .net core mvc provides the same batteries included approach and is built with the modern web ecosystem in mind. It also runs on linux and easily integrates with postgres so its just as cheap now as well. I think people that write off C# haven't really worked with it in its present form. I didn't find python significantly more terse than C# (other than having to define properties on types, which I already stated I prefer), just less feature rich.
Back then we already had Zope and it wasn't blazing fast.
I followed the tutorial using the aspnet CLI tool and got nothing but errors. I Google’d for a while until I gave up.
Switching to a programmig language based on an entirely different programming paradigm is not comparable to switching to a language based on the exact same programming paradigm to develop the exact same application using the exact same design patterns.
This was from the official site and many of the commands threw errors.
I think it was a case of incomplete Mac support, which should have been called out.
IIRC I posted on HN about it and got an apology for the state of tooling
Otherwise the doc is quite good for ASP.net Core, and it’s rare I get stuck on a problem for too long with it.
2) C# is a very verbose language, that requires a lot of typing.
3) F#, the best language in .NET, is largely ignored by the .NET community.
And this was at the time when MS would announce a new blessed way to do things on a reasonably frequent basis. When I started, the blessed way to access data was DAOs, then it was ADO.NET (note, those could be around the wrong way, I have trouble figuring it out now in hindsight), then it was Linq2Sql, then that was deprecated for Entity Framework (which, I'll give credit, they seem to have stuck with, even if it does feel like a half-cribbed NHibernate).
I was frantically trying to learn the blessed thing, because the MS shops I was familiar with only used the blessed thing - the server was IIS, the database was SQL Server, the language was C# (I was the only member of the DUG who coded in F#, and few others knew of it), and you used the blessed patterns and the blessed frameworks.
Incidentally, this is why FOSS has had such a hard time in .NET, as soon as MS releases something that is reasonably feature complete, a lot of single-vendor minded companies switch to it.
And I met a lot of developers in their early 40s who were quietly terrified of getting left behind on the MS technology treadmill, and trying just as frantically to learn the new blessed thing as I was.
And then I got hired by a Java shop, and faced a paradigm where shit code cough java.util.Calendar, java.util.Date cough stuck around for yonks because it was good enough and replacing it had to be done very thoughtfully and gently.
Point #2 isn't super relevant in the age of IDEs, but I agree wholeheartedly that F# deserves a lot more love than it gets.
ASP.NET Core 2.1 was released in 2018 and will be supported until late 2021.
ASP.NET Core 3.1 was released a few months ago and there is no end of support in sight. Moreover, the changes between 2.1 and 3.1 were not that many. I've migrated a whole ASP.NET Core 2.1 web service to 3.1 in less than 1 hour.
> 2) C# is a very verbose language, that requires a lot of typing.
Nonsense. The only added verbosity to C# when compared with Python are the type declarations, which arguably are a problem plaguing Python. The first class support for events and async programming and properties in C# more than make up for it.
> 3) F#, the best language in .NET, is largely ignored by the .NET community.
I fail to see what point you were trying to make.
Which is way too unstable, especially for the kind of corporate environment c# has typically been used in. Getting those places to upgrade to stable supported versions of the framework has always been a battle even when backwards compatibility was great, if they have to deal with breaking changes every few years they will never upgrade.
This is why so many companies stick with their ancient COBOL systems, most modern alternatives don't offer the stability they need.
ASP.NET Core 2.1 is the LTS release of ASP.NET Core 2, which was released in 2017. I fail to see how a first class framework with a LTS that was released years ago can be described with a straight face as "way too unstable".
> Getting those places to upgrade to stable supported versions of the framework has always been a battle
ASP.NET Core 2 is stable since at least 2 or 3 years ago, depending on how you decide to count.
> This is why so many companies stick with their ancient COBOL systems, most modern alternatives don't offer the stability they need.
This assertion is simply wrong at so many levels. Don't confuse "why waste money maintaining working software" linesof reasoning as a sign of respect for stability.
More importantly, it's disingenuous to even think of the technical debt that keeps cobol on the map as relevant to the world of web services.
I fail to see how you can call 2 years of support an LTS with a straight face, it's taking the piss out of the term. The LTS of the OS I'm likely to run it on is supported for 8 years. 2 years isn't even enough time to finish many projects on the same LTS it started on.
At work we've got 30 year old c/c++ code bases that still run, they'll probably run for another 20 at least, we've got 20 year old python code that still runs (for now) and we've got 20 year old c# projects that still run. That last one will never get rewritten in .net core, in part because they've pissed away the stability the framework had. It would be crazy to use tools with 2 years of support for any of those projects.
> Don't confuse "why waste money maintaining working software" linesof reasoning as a sign of respect for stability.
Why should they waste money maintaining working software when there are stable options available? What does upgrading to asp.net core get them? Why should tens of thousands of companies waste money modifying working software just because someone on the core team thought the existing API was inelegant or too hard to maintain compatibility?
Notice a language can't be everything. They're either too verbose, too terse/cryptic, or trade ease of development for lack of control/performance/efficiency.
IMO, C# strikes a good middle ground on such matters...
You will eventually cut it down, but I see better ways to use my time.
If I find myself debugging python tools, I usually just add debug statements to figure out WTF it is trying to do, and reimplement it in bash. It invariably is less than 10% as many lines of code, and also more debuggable / readable than the original.
Granted, most of the python scripts I see these days are build processes or cluster coordinators.
> more debuggable / readable
More debuggable / readable for you, perhaps.
Yes, if done properly. The issues with the particular Python/Gunicorn setup the author described in an earlier article (linked to in this one) were not so much Python/Gunicorn issues as "not understanding how to properly use Python/Gunicorn" issues, or more generally "not understanding how the tool you are trying to use actually works" issues. (I actually shudder to think what such a group would have done trying to program the same application in C.)
They abandoned their message queue (Starling) written in Ruby early on. That abandonment was often misattributed to them abandoning Rails, I guess because of the shared usage of Ruby. The confusion was compounded to back when someone at Twitter (may have been Odeo at time?) posted a message board rant about how Rails does not scale because ActiveRecord did not allow connecting to multiple databases out of the box. That event seems to be the origins of the "Rails does not scale" mantra that swept the internet for a time. But, humorously, someone replied with a solution a few minutes later.
Twitters problem was a too centralized architecture not Rails.
And I say that as someone who at the time hated Rails and who still hates Rails. I find it bloated and over-complicated. It may even have led them to make bad architectural choices because of how it was structured.
But they still did make bad architectural choices, and they fixed those choices at the same time they moved off Rails.
Had they started with them and there wouldn't exist blue wales.
My first commercial use of Ruby was in 2005. Not web facing, but messaging middleware. As in a pub-sub type passing of messages between various endpoints.
We had a C version. It was about 7k lines to support the bare minimum we needed. As an experiment to teach myself Ruby I wrote a Ruby implementation. With the usual caveats (it's often easy to make a rewrite better in all kinds of ways, including size), it was ~700 lines, far easier to read, and supported far more functionality, so I put it in production.
Was it slower? As usual that depends what you mean by "slower". It consumed 10x more CPU, but it also did much more work (e.g. supporting more flexible routing of messages etc.). The throughput, however was the same, and 10x more CPU means that maxing out the network connection took 10% of a single core instead of 1% of a single core.
For some types of tasks CPU is the most important thing, but for a lot of tasks you'll be IO limited. And for a lot of tasks that people think are CPU limited are really down to poor IO handling (causing excessive context switches e.g. through lots of small reads is a common one)
You don't need to give up on JIT and AOT compilation to use high level languages, and this is where current tooling for Ruby and Python ends up losing.
side-note when starting with D: make sure to install dub. it's the package manager and basically eliminates makefiles from the compilation process. Just "dub init" and "dub run" and you're off to the races.
With the proliferation of microservices, I find this increasingly not true. Sure, python definitely is plenty fast when you need to send something to a user many miles away. But with microservices, writing to the network might mean writing to a machine in the same data center, or even the same host in a different container. That's as fast as a few memory copies and a few context switches.
Give us some numbers.
Inside the data centre you have 10G, 40G, 100G ethernet connections. I know for a fact that you will struggle to soak a 10G connection using a single thread so I know you can't do this in Python without multiple processes using SO_REUSEPORT.
That said by far most programs don't need to worry about saturating a 10G connection. I'm not writing a file server in python, I'll leave that to nginx or S3. I'm writing business logic in python which tends to be bottlenecked by a database in any case.
Python is great for plumbing together other functionality, which it turns out means most backends you'd be writing anyways. Python is less great at handling a large quantity of data, though most of the time you can get away with handing the data handling to some library (e.g. numpy or libuv or any one of thousands of libraries).
Worst case you can easily plumb in some C-calling-convention code into python. With FFI it's a matter of copying the header definition and you're off to the races. That way you can still write the bulk of the program in python, delegating the bulk data wrangling to C or D or rust or go or whatever you prefer.
There simply isn't an excuse for using Python for any infrastructure. It does nothing particularly well - or even right - other than very purpose-specific scripting. It can tie things together well enough. And your codebase becomes a liability rather than an asset.
I always say this opinion when it comes to Python discussion in HN and I always get downvoted but hey, "all it takes for evil to triumph...".
This coming from someone who has literally never written a line of python in his life.
People expressing unpopular opinions and defending them with evidence is how we find out when the popular opinion is wrong, and how, when our discourse is functioning properly, we can gradually change the popular opinion to be less wrong. You, and the people downvoting the comment, are throwing a monkey wrench in those works.
— ⁂ —
In 2000 or 2005 Python was a simple, consistent, practical language with a policy of strict error handling that was very useful for producing reliable code: "Errors should never pass silently. Unless explicitly silenced. In the face of ambiguity, refuse the temptation to guess. There should be one-- and preferably only one --obvious way to do it." Its runtime cost was significant but bearable, about a factor of 20–40: if you wrote your code in C instead of Python, it would run 20 to 40 times faster, and that was all the machine could do.
In 2020 Python is an overcomplicated, inconsistent, slow, unreliable language with a persistent schism resulting from the core developers' poor choice to make backwards-incompatible changes to simplify the language. It has metaclasses, superclass method resolution order linearization to enable mixins (two different ones, in Python 2), a lazy sequence construct that has been gradually Frankensteined into a general coroutine construct (with an additional lazy sequence construct added on top), two different incompatible language constructs to compensate for the lack of block arguments or full-fledged lambdas (decorators and context managers — I'm excluding generators here since they're more powerful than block arguments), and on and on. The reference documentation for "import" alone is 20 pages, and that's in Python 3, the simplified version of Python.
Python's performance cost has not increased in absolute terms — in fact, it's even improved a bit — but it's increasingly painful. In 2000 we could rest easy knowing that whatever we wrote in Python would be sped up by Moore's Law and Dennard scaling, roughly a doubling in speed every 18 months, so in three years it would be four times as fast, and in three more years it would be 16 times as fast. That, together with a little judicious implementation of inner loops in C, was a small price to pay for getting things done sooner and not having to open core files in a debugger.
But then Dennard scaling slammed into a wall around 2006 and Moore's Law sank into a swamp around 2016. Meanwhile, manycore meant that without multithreading, or at least multiprocessing, your program suffered an additional order of magnitude slowdown. Even US$40 hand computers now feature quad-core CPUs. Today, the gap between what the machine can do in absolute terms and what it can do when saddled with Python is a gap of 1000 or 10,000, not 20. If you can cope with the limitations of PyPy (it supports Numpy now! Since 2017) then you can get up to the speed of single-threaded C, which is about 3% of what your computer is capable of. But it's not going to get faster just because hardware progressed: computers will maybe be twice as fast in five years, at best, and maybe not. If it's too slow today, it'll probably be too slow then too.
But that's not the worst part. Python's completely botched Unicode handling introduces bugs into most Python programs that handle strings from the outside world, latent bugs that only surface once those strings contain non-ASCII characters — similar to the situation with bash scripts and filenames containing spaces, although that can be detected by purely local analysis (missing doublequotes around a $var, red alert!). Plan 9 had already demonstrated one correct way to handle the situation (the one used in Golang and Rust) and Markus Kuhn's UTF-8B proposed another, one which was eventually partially implemented in Python as PEP 383 ("surrogateescape") but turned off by default. I've had bugs in on-orbit satellite control software that I couldn't track down because Python generated a UnicodeDecodeError when it tried to log the stack trace.
— ⁂ —
At the same time, other alternatives got a lot better. Java grew into a mildly reasonable language, and Kotlin and Clojure are outstanding ones. Haskell, defying everyone's expectations, became practical. Microsoft started trying to embrace and extend free software, so now we have F# on Mono, which is almost OCaml — almost as convenient and concise as Python, but enormously less bug-prone. Mike Pall, a superhuman intelligence from the future, wrote LuaJIT, which gives you performance on par with C in a language as friendly as Python — not modern overcomplicated Python, old Computer Programming For Everybody Python. 100 million people, including little kids, program computer games in Roblox using Lua. (It's bug-prone as hell, though. Lua has a footgun for each toe, as Sean Palmer says.)
Even C++ has been tamed somewhat. And of course we have Rust and Golang. Golang is only a little bit uglier to program in than Python, and both of these new systems-programming languages make it a lot easier to take advantage of manycore, though not SIMD.
Switching from Python to Golang is as easy as switching from Perl to Python, but with a lot more benefits. And that's a big reason why a lot of the important infrastructure software written over the last decade has been written in Golang.
On the horizon, we have things like arcfide's Co-dfns APL compiler, Matt Pharr's ISPC, and GLSL to show us how massively parallel programming, including SIMD, can become accessible to mere mortals. They aren't yet practical options as alternatives to Python (except that GLSL is practical for its original purpose, of making nice graphics), but they might be pointing the way to something that is.
— ⁂ —
So I think it's eminently defensible that Python should now be consigned to "a few hundred[] lines worth of utility". Python is great for scripting TensorFlow, and it's a far superior substitute for MATLAB. But writing infrastructure in Python in 2020 is like writing infrastructure in Perl in 2005.
...which I was also doing. I think I probably owe an apology to a lot of folks at Aruba Networks who are maintaining that code today.
The difference in productivity between C# and Python is such that they might as well have been developed by different civilizations.
(I have occasionally used Python since about 2003 and C# since 2018)
I think they were. Is that a bad thing?
One important reason for using Python is that it almost always forces people and companies to release the source code.
I will take a shitty Python script over shitty C/Rust/Go/Java code any day for that reason ALONE.
With source code, I can fix your shitty program (and all programs are shitty--even mine). If it's compiled, that path is blocked.
First of all, it does not take "that much machine" to serve a fair number of clients. Ever since I wrote about the whole Python/Gunicorn/Gevent mess a couple of months back, people have been asking me "if not that, then what". It got me thinking about alternatives, and finally I just started writing code.
```
Another day, another questionable premise for a blog post. No, @rachelbythebay, the question is not "if not that, then what", and the answer is not reinventing the wheel. The question is just "what ___" -- what do you plan on doing, what does your software need to do, what does it need to support? Use the right tool for the right job. If you have a language you're proficient in and with an ecosystem that supports you developing something rapidly, it's borderline malpractice not to start there. When you need to optimize, optimize then. Maybe that means you carve out a subcomponent into a new service, and you choose a language purpose built for speedily doing what you need. Maybe it means a lot of things, but it doesn't mean you throwing out the baby with the bathwater and setting out to recreate the baby, the bathtub, and the bathwater from scratch to answer the question of why your tub is overflowing.
I wish this blog post was about solving real engineering problems instead of writing code to provide mediocre answers to poor questions.
I think the point of this post is that most new programmers do not know things can be faster than their monstrous JS blob. The solution to slow requests is more servers instead of fixing the code.
Related to this: programmers who know an ORM or two but never learn SQL proper.
At my job I refactored a giant, slow, memory hungry reporting task into a single sql query. It used to take 10s of minutes to collate 1000s of datapoints. Now it takes 10s of miliseconds to collate a factor 10 more data. Never mind after I added some indexes to speed the thing up.
Knowing about the layer below the abstraction you're working at can be rather useful at times.
Notably, you did NOT decide to write your own dialect of SQL to "scratch an itch". Knowing the layer below the abstraction you're working at can be useful at times, but only in relation to understanding the context of how all the layers fit together and making effective, pragmatic decisions. Pointless rewrites are anything but. Good on you for avoiding that impulse and doing the right thing.
Besides, I'm not suggesting (or "demanding") that all articles explicitly state their target audience. Just that this one does a poor job of indicating it.
So I think it's a poor argument when the parent comment (to which I was responding) assumes an exclusive target audience: "not introductory programmers... more experienced programmers".
You could argue that one was supposed to assume that from context. However, I didn't make that assumption whilst reading, so I especially wouldn't expect "introductory programmers" to pick up on that either.
I can DoS it with a single java client running 50 threads. If I use 100 the p95 shoots up to 30 - 40 seconds.
But the kicker is that no one (other than me) really cares. 50 concurrent threads is probably around the peak load it will get in prod, and various people involved think why bother trying to fix it?
It's driving me nuts.
Thanks for listening
All of them sit at 1% cpu as well.
- Web-anything is fetched off of a DB nowadays. That's another crapton of latency, because a) it's just easier to understand its characteristics if it runs on a separate system and b) most companies IME have either no DBAs at all or the DBAs have no time to look in depth at every system being built. So the DB resources are vastly underutilized, and the DB itself is badly understood.
- Cache invalidation is still the hardest problem, and every caching framework I've looked at seems to gloss over that part. Just cache all the things and hope people retry enough times to get the latest update. I would love someday to work on a system where things are aggressively cached at every level and invalidated at every level and with perfect granularity.
- Building for web scale from the beginning is premature optimization for the vast majority of companies. In the vanishingly unlikely scenario that the company actually grows 100-fold or more it makes sense to start investing heavily in performance. Of course, this also has the knock-on effect that the vast majority of software developers never get anywhere near a web scale system. OTOH it creates jobs for millions of developers, some of whom might end up building at scale someday.
Another elephant in the room is that building anything at web scale is just not something anybody straight into the workforce is anywhere near qualified for. We desperately need more focused learning (mentoring, pairing, etc.) across the board to bring everybody up to speed faster.
IMO she's right, I wish we didn't mess wrap around gunicorn and gevent for some of our services. Certainly would've made my life easier and the services faster.
[0] - https://www.appdynamics.com/blog/engineering/a-performance-a...
bjorn is fast because it has a minimal feature set. No threads, no multiprocessing, no nothing. If it works for you, great, but it's never satisfied my requirements.
I'd avoid uWSGI, it's performance is good but it's so complicated with so many features that I never felt confident using it.
Never used waitress except in development, but people seem to have had success in production.
It seems like the new trend for Python servers is ASGI (https://asgi.readthedocs.io/en/latest/), e.g. as in uvicorn (https://www.uvicorn.org/).
Some fun extra features that uWSGI provides:
- Daemon management: It can manage other daemons, for example Celery, for you, so you can put your whole environment into a single uWSGI.conf.
- A cron-like interface for generating events on a schedule.
- Emperor - hosting of multiple apps. This is extremely configurable, you can have the server look at the filesystem for config files for the simplest setup but it also supports AQMP, a Postgres database, MongoDB, a shell command or a bunch of other things. It has a bunch of interesting isolation options, like running each app in its own Linux namespace. It also has (optional) socket activation, so apps won't be started until the first request.
- Auto-scaling
- Soon, multiple event-loop/IO subsystems to choose from, including asyncio
I've been meaning to have a look at nginx-unit though.
Its primary use case is that it is pure python, doesn't rely on any specific libraries or compilers to run/build, and is a threaded WSGI implementation so it uses Python threads to run a WSGI app.
It works well for what it needs to do, and hopefully it is fairly robust. I've personally ran waitress directly facing the internet, but will readily admit that in most cases running it behind a load balancer is a good idea, especially since it doesn't support SSL out of the box (yet, I should say, it's on my roadmap).
It won't win any speed contests and it won't win performance contests, but it holds its own.
If you have any issues, please drop by https://github.com/pylons/waitress/issues and I'll see if I can help you out :-)
File systems too. A cron job writing a csv makes a surprisingly simple cache the OS can keep in memory.
Here, to me, is the key item:
The "listener" thread owns all of the file descriptors (listeners and clients both), and manages a single epoll set to watch over them.
This is exactly what any async server does: it centralizes all the file descriptor management and handling in one place, and only uses workers (whether they are threads or "green threads" or whatever) to read from/write to fd's that are marked as ready in the epoll set.
For your case, unless I'm misreading something, what the workers are doing in between the read/write is CPU intensive (or at least it's CPU work and not I/O work, even though it's not very "intensive" CPU work), so actual OS threads are a better choice for the workers since you can't rely on cooperative scheduling.
If what the workers were doing was I/O work (for example, sending a request to a remote database and waiting for a response), "green threads" would work fine (since their only real purpose would be to organize the I/O--the actual fd's are going to be managed by the central server that manages all the fd's and checks which ones are ready for read/write). And one definitely should not try to run "green threads" for the same server in multiple O/S threads (or worse still, multiple OS processes). For an I/O bound server, one shouldn't need to anyway.
Totally possible with another thread acting as watchdog timer and sending a signal which causes the read to return with EINTR which can then check a flag whether it should retry or abort. And that's for file IO. For socket IO you can just set it to non-blocking.
An alternative to putting the watchdog timer in another thread is to use alarm(2) and use the kernel's built-in watchdog timer, and the default behavior for SIGALRM is probably adequate. This might be easier than non-blocking I/O.
Are you thinking of, like, CICS systems from the 1970s connected to SNA or something? I mean they didn't have select(2) but they did serve many clients in a single process.
If there was a time when all network servers were written using select(2) or similar event-driven APIs, it wasn't 1988 or later.
No, it isn't, it's doing the same thing as a select(2) server would do, except it's using epoll to avoid scaling issues when you have a lot of file descriptors in the polling set. The only difference is that the workers are doing something that requires CPU, not I/O, so OS threads are being used for them (a single threaded server would be fine if the workers were just doing more I/O, like sending a request to a remote database and waiting for a response). But the worker threads are not doing any I/O management at all; they read from or write to an fd only when the central server that is calling epoll tells them to. In the listen/accept/fork model, the central server forgets about an fd once it has passed it to a handler process, and the handler process using blocking I/O.
I actually want to know: then what? As a web developer who usually reaches for Django or Flask with Gunicorn because I just don't know any better, is there a better stack that doesn't face these problems? Or is this a 'call to action' for somebody to build a better web server that follows this advice?
Apparently, it's C, or ...assembly? According to the comment above yours.
Because developer time is apparently cheaper than CPU cycles. Weird.
And what's the opportunity cost of the new features they can't create because they're rewriting existing apps?
If you're arguing that Python has more runtime overhead than Rust, I don't disagree.
But there's a reason people invented higher level languages than C. Rust is a far nicer systems language than C, but is it faster to develop in than Python or Kotlin? It really depends.
That said, with Rust or C++ you can – depending on the situation – be clearly as effective, or even more effective than with Python, and you gain an enormous performance benefit (which in turn saves you money, which in turn means you can hire even more developers)
Put in the same effort, and get 10x-1000x faster code or lower resource needs. Why not?
C and CGI are a great simple combo. When you get down to it 99% of web development is shuffling data from sql to html and vice versa. It's so stupidly simple that the main code doesn't even hit the rough edges of c, there are almost no allocations for isntance so no memory management to worry about, everything that could possible leak memory or cause security issues is in a handful of utility functions.
No MVC, no DTO's, no view models, no service layers, no templating, no client side javascript, just reading from and sql connection and printf'ing html to stdout. Without all the useless complications I ended up with far less code than the typical equivalent in a high level language with a framework. I'm now fairly convinced that most of the projects I've worked on would have been better off this way.
I'll just stick to python thanks.
I'm a bit mystified by Rachel's use case—A green thread hogging the process for so long that another request times out on the client? That means that requests are doing large amounts of processing and latency requirements are super tight, which sounds like a very specialized use case.
Go, any JVM language, any .NET language, OCaml, Haskel, D, JavaScript/TypeScript
As alternative you can try to use PyPy instead of CPython.
`it kicks off a "serviceworker" thread to handle it.`
It doesn't even tell me how I should do this. I don't know what 'serviceworker' code looks like. I don't know what 'kicking off' means.
This post reads like it was made for maybe 3 or 4 people in the world who truly 'get it' and if you aren't in that elite club you're a terrible engineer, apparently. There's not even a lick of example code to get an idea of what is going on. Engineering is a big word that shouldn't be used in this blog post.
C10K was solved by switching to an event system like epoll or kqueue. This decreased the big O complexity of kernel<->user information flow so that you don't pay more for listening on more sockets.
C10M seems to be solved by colocating the network stack and app stack data structures by running the driver and the app in the same context. This can be achieved via DPDK-like schemes or pushing more into the kernel like Netflix's work to push TLS into kernel sockets that you can sendfile to.
> By the early 2010s millions of connections on a single commodity 1U server became possible: over 2 million connections (WhatsApp, 24 cores, using Erlang on FreeBSD), 10–12 million connections (MigratoryData, 12 cores, using Java on Linux).
Code, changelog and a bit of history: https://lugdunum.shortypower.org/kiten.html
I wonder what app that could be :-)
- Slack
- Keybase
- Mattermost
- Discord
- Microsoft Teams
- Facebook Messenger (incl. Work)
All of these fine folks apparently use Electron or a similar technology...
I think of Electron as Gresham's law in action.
Slack on the other hand can be very slow at times.
Now in order to make it slow you must know enough stuff to introduce complexity to the program. To make a slow program you probably have to learn about frameworks, php, micro services, cloud databases, etc.
I remember when C10K was the big challenge, but even the naive approach of spawning a thread per connection now can handle that.
Modern fast network code in Rust is based on async/cooperative multitasking distributed over num cpu worker threads which works very well.
Old box, complicated code, few thousand connections, good fucking luck.
Does she even see the size of the batteries included in a framework like Django? Here is some brand new information. Everyone knows Python/Django is slow and inefficient (same for rails) but they still use it for those batteries.
It is easy to write some web service in a low level language that just crunches some numbers for your benchmark. It quickly gets complicated when you start thinking about accounts, security, passwords, databases, orders, carts etc.
What is useful is to spend 2 days researching massively concurrent network applications, and find the ones that were already written, and see how they evolved over time. For the most part, it's just understanding your operating system and network protocols. Once you learn how they all work, the answers come quickly. This problem has been solved many times.
But for the most part, none of you need to know how this stuff works. You can write the crappiest network server in the world, and there are still so many other components you can wrap around it to make it scale that you never even need to get close to network optimization. So you may have to run 10 instances of your app; who cares? We have virtually unlimited everything these days. The answer to a poorly performing app is "just throw more cloud at it".
Also, I feel the need to remind everyone that single app instance handling 100K+ concurrent connections is a terrible fucking idea. What happens when the app crashes? What happens when the hardware dies? What happens when you need to, like, upgrade/restart the app? Several million SYN packets in 100ms, and default kernel TIME_WAIT settings, do not make great bedfellows.
How about two app instances?
said it has no real purpose yet. that's key. no doubt she knows what she is writing about and also picked a favorable language. but the gunicorn folks have also been doing this server thing forever now. they probably have their share of stories.
hope some code gets released soon.
I know amount of users doesn't equal users/second. But we shouldn't pretend like they weren't able to handle high traffic before Facebook aquired them.
Each time football Cup is active Twitter goes down. Same for plenty of big names when they launch a hyped service (Blizzard for example is another one). Scaling up from lab to real life dealing 24/7 with those millions/second is an entire different beast.
It's about 30% slower than the fastest golang equivalent.
For multiple-queries benchmark, its about 40% slower. https://www.techempower.com/benchmarks/#section=data-r18&hw=...
Once I start bringing Numba for accelerating computational usecases... I guess the difference will be even smaller.
Is it? I haven't used threads directly in a while, but I remember dealing with sinchronization issues. Problems that just don't exist in single threaded node with async await.
I find the async Promise or Task to be a more useful abstraction than the thread. Although, you need threads or a task dispatcher with a pool of threads if you need to run cpu intensive stuff.
No, but if your single listener/server thread is managing epoll for all your file descriptors, you do have to have a way of synchronizing the worker threads with it, so they know when and when not to read from/write to their fd's. I assume Rachel is using some kind of semaphore or other threading synchronization mechanism for this.
Try it today with NGINX.
This is basically the erlang design. And yes it works really well.
And this is to the detriment of all of us. That’s why you’ll never hear me call myself an engineer.
We don’t. We built wind tunnel models, we taxi them around at higher and higher speeds, we put them in machines that wiggle the wings at high loads. The first flight is a little hop and then right back down. Months later there might be a big ceremony with VIPs where the new plane takes a lap around the airport as it’s “first flight”.
And there are mistakes song the way, giant ones that add years to the schedule and tiny ones that engineers argue over even telling their boss about.
Sure there’s planning and experience, just like she knew to use epoll and not select, and that Linux can handle thousands of threads per process. But there’s no magic to it, just lots and lots of human attention and testing along the way.
Didn’t Boeing just get caught trying that?.
Whipping something up and then seeing what it's capable of doesn't qualify.
The gist of it was: Other engineering disciplines use the techniques you've mentioned because of the the costs, both time and money, associated with getting it wrong.
Software engineering lends itself to different methods of development and construction, as the costs associated with getting it wrong or making changes after the fact are much lower. (For most applications, anyways).
As such, (this definition would be another sticking point) these less rigid methods should still be considered engineering, with engineering being a balancing of resources with outcomes, not fixation on mathematical models.
What is actually happening we don't know. IT is dramatically changing everything. It's not all bad but it isn't all good either. I hate to give examples since it is really vague what goes in and what comes out and people tend to mistake an example for full coverage but... for example, we have no idea what drives suicide rates.
This is a good point but I wanted to make it a bit more clear.
Across a population we have a good idea of the risk factors and of the things that increase rates of suicide: deprivation, abuse, substance misuse[1], previous self harm.
The bit we have no idea about is how to apply these to an individual person to see if they're high or low risk of suicide.
There are a load of different tools that input lots of different information and put out a risk rating, and none of them are as good as just asking the person "what do you think your risk is?"
Most large companies do some stress load testing first for any new system. Then they release to few percent of the users (less than 10%) and gradually increase to the rest of the userbase.
I worked at Spotify, when Tidal was launched. They had failed to do proper capacity testing, and the service failed under the load the first week and it showed. But most mature large company tend to be really good on this.
It is remarkable how many tech companies have managed to stay mostly up with very few outages, given this whole pandemic situation, where everybody is online.
I'd say working in large scale deployment is proper engineering .... creating a simple webpage, maybe not. Deploying that webpage to millions of users, it is.
Also, don't forget that bridges have been made since the dawn of time, while tech and the internet are very very young in human terms.
I must admit that it was very gratifying watching all our engineering decisions pay off when our system started seeing record high traffic every day and scaled effortlessly and with no outages - the traffic that came with lockdown is, even on our slower days, twice our planned for and tested for high water mark.
But it took us about 4 years of consistent effort in changing our organisational mindset all the way through, devs, testers, product managers, c-suite members, to get here. 4 years ago, we would've been waking everyone up and going without sleep for a couple of days trying to get something back online, then fixing the next system down the line that failed because of the load, then the next one, and then writing long post-mortems for our business team.
And that's not engineering?
I was a mechanical engineer prior to switching to software. As a general rule, the things we do in software are very distant from engineering.
A seasoned engineer notices that everyone is only building suspension bridges all of a sudden.
They point out that you can span a stream with some bricks or rocks and a bit of mortar, and are ridiculed.
Next, they build a highway overpass out of concrete pylons, and stress test it to 10x the necessary engineering load, and point out it cost 10% as much as a typical contemporary suspension bridge. That’s this article.
I think this is akin to NASA working on the Apollo program vs. someone in their garage attempting to build a go-cart for the first time.
When you just slap things together and see if they work - are you really engineering? Can you exactly repeat the process and achieve exactly the same result every time?
I think we often cross "research and development" with "engineering". Exploring a problem space and tinkering with concepts isn't engineering. Taking what you've learned, planning out and executing a solution to a precise set of requirements, and being able to repeat your steps and achieve those results again and again - is engineering.
Why not? That's basically what testing is. Which was one of the attributes you attributed to "Engineer" >formal testing
>I think we often cross "research and development" with "engineering"
My general take:
Scientists primarily focus on learning and proving new knowledge & ideas. (i.e. they research)
Engineers focus in using proven knowledge and applying it to design and create things or solve problems. When things are not perfectly certain, they can prototype and do tests similar to how scientists do experiments (e.g. aerodynamics in wind tunnels). (i.e. development)
We don't run test suites on our software to see what it does. We run test suites to validate it operates as it is supposed to.
I think the way you described testing is more in line with tinkering and research rather than engineering. It's experimentation, not testing.
When the outcome is unknown and unreliably unpredictable, it's research (tinkering). When it's predictable and has a known, repeatable outcome, it's engineering.
Yeah that's fair.
I had originally skimmed the article, but after re-reading it the author apparently admits they didn't put any careful thought into what they made. No real goal. Just slapped stuff together so to speak. Which I'd agree doesn't quite sit as engineering to me...
As the title says, this is dumb code. But it makes use of years of engineering effort to deliver a result which can get you very far without thinking about the low level.