Apache releases first major new version of popular Web server in six years
zdnet.com
zdnet.com
It frustrates me when people use ASCII instead of packed bitmaps for things like this (packet transmitted once a second from potentially hundreds or thousands of nodes, that each frontend proxy has to parse into a binary form anyway before using it). Maybe it's a really small amount of CPU but it's just one of many things which could easily be more efficient.
Some of this stuff is so simple and useful it's a wonder they weren't there before:
http://httpd.apache.org/docs/2.4/mod/mod_auth_form.html http://httpd.apache.org/docs/2.4/mod/mod_session_dbd.html http://httpd.apache.org/docs/2.4/mod/mod_buffer.html http://httpd.apache.org/docs/2.4/mod/mod_data.html http://httpd.apache.org/docs/2.4/mod/mod_ratelimit.html http://httpd.apache.org/docs/2.4/mod/mod_lua.html
"mod_ssl can now be configured to share SSL Session data between servers through memcached"
"mod_cache is now capable of serving stale cached data when a backend is unavailable (error 5xx)."
"Translation of headers to environment variables is more strict than before to mitigate some possible cross-site-scripting attacks via header injection."
"mod_rewrite Allows to use SQL queries as RewriteMap functions."
"mod_ldap adds LDAPConnectionPoolTTL, LDAPTimeout, and other improvements in the handling of timeouts. This is especially useful for setups where a stateful firewall drops idle connections to the LDAP server."
"rotatelogs May now create a link to the current log file."
Very nice that they rewrote the mod_rewrite and caching guides with more examples and ease of use in mind. Here are the API changes: http://httpd.apache.org/docs/2.4/developer/new_api_2_4.html
I think that's a premature optimization. If it becomes a performance problem, optimize it then. Otherwise, I doubt it's worth the cost of humans not being able to read the information on the wire, and deciding on and implementing a binary format to represent the information.
If you're writing a piece of some software with one specific function and it doesn't affect anything else and it doesn't take any more time to make it efficient, just do it efficiently the first time. This way you don't have to come back later and rewrite it (which by the time that's required, somebody already wrote something dependent on your crappy design, which now has to be fixed too, etc)
100% agreement on "if it's not showing up in your profiler, don't optimize it".
Yeah it's probably insignificant, but they could have just done it that way the first time. There'd be no optimizing needed thereafter, and no side-effect as it's only used in two places for a limited function. It's actually smaller and easier and faster to write AND to run. You make the right choices at design and implementation time and you come out with better code which doesn't need to be optimized.
It parses that string and reads it into a struct at a rate of about 3 million per second on a single core on my Macbook Pro. That means for your example there of 850 machines in a cluster, we're talking about roughly a quarter of a millisecond of one core of CPU time. Even if that's off by an order of magnitude, it doesn't matter.
THERE IS ABSOLUTELY NO REASON TO OPTIMIZE THIS CASE
The enemy of optimization is hard to maintain code, not small inefficiencies.
Using text in this case makes the wire protocol easier to debug and generally easier to maintain for forward compatibility. Need to add an extra argument? It'll be both self-documenting and backwards compatible, instead of some versioned mess of bitflags.
Could I be wrong? Sure. There's an easy way to prove such: show me the profiler output. Your eagerness for micro-optimizations significantly hints at inexperience.
Then why add string serialization to something that doesn't remotely need it?
Using text in this case makes the wire protocol easier to debug and generally easier to maintain for forward compatibility.
If anything the use an unspecified string format makes debugging potentially harder here (UDP truncation...).
The wire protocol you're talking about is a heartbeat protocol. This doesn't need to be human-readable because there is no input from a human, nor is it intended for a human to read. Debugging it would be as complicated as a "printf". Oh the horror. We wouldn't want to add debugging to our application - better make people read the raw data with a packet sniffer (which you can't run unless you have root, so good luck getting your application debugged quickly, developers).
Adding an extra argument would require appending an extra field and incrementing the version. Oh noes the mess! Since both solutions are versioned, forward compatibility is just updating code in the server, and if you're going to change things you have to update code anyway, so this doesn't sound like a good reason to oppose it.
Realistically, will more than a handful of shops have a big enough cluster for any of this to matter? No. But even your example of using sscanf is faster than Apache's code (http://pastie.org/3430085) while being 100% compatible with the rest of the code, and still goes to show that doing it right the first time is better than just slapping something together and waiting until you have to optimize.
So do we need to use an unreadable, simpler solution? No. But it would work just as well as anything else and take just as much time to write - if not less.
If you find yourself saying that, then your work is done. There's nothing to optimize. Optimized code is non-idiomatic, usually clever, often obtuse, and as a result, more buggy. Sure, copying the data off the network is easier, but you still have to figure out an encoding on the sender side, document it, implement it, and probably in the future extend it. There is a very real cost to having around "optimized" code, which is why you had better be sure that the performance gains outweigh the software engineering costs.
Parsing strings is most certainly not the default in C.
And since the data being communicated between the heartbeat client and server is computer-generated and automatic, "text" is not even not required, it is an unnecessary step in communication. Apache is implemented in C. C is good at accessing, copying, transmitting, and inspecting raw memory. It is not good at parsing and comparing strings. Adding text to a protocol nobody will ever interface with is adding an unnecessary layer which is not only more complex than my suggestion but is itself more prone to security flaws and bugs unless proper care is taken when parsing the strings.
Really I think this is about people's comfort levels. You feel more comfortable looking at a string and knowing what's in it. Either way will work; it doesn't matter how you skin this cat. But I think it's disingenuous to suggest that comparing raw bits is somehow more likely to blow up than comparing the raw bits after you parsed them from a string.
Further, I think you are underestimating the benefits of people and other programs being able to read the messages without having to read up on, and reimplement such a format.
A protocol which transmits data that dynamically changes according to the use of the application servers, which is implemented by services using a language which makes it easy to communicate raw memory and not need to parse it, passing a tiny low-latency message over an unreliable transport layer without even a concept of whether messages are ordered properly or not.
Forget the common developer. Forget compatibility. Forget usability. Forget reliability. This thing only has one function: tell the frontend proxy i'm alive and how many slots are active or busy. The only thing that matters to the backend server is how fast you can spit out a packet, and to the frontend how many machines are alive. Shit, there might not even be a valid checksum on the udp packet. And you're worried about how much effort it takes to implement this? It probably takes less time to write the code than it does to add the source into your makefiles and make a unit test. The existing function that parses these packets is one small function, which without parsing a string would be one memcpy (or the slightly-less-flexible sscanf as shown by wheels earlier).
There is no crime in writing apps in a way that fits their use. Not all apps are the same, and not every "default" design decision is appropriate. Sometimes you just do what you're familiar with. In this case the Apache guys used an HTTP query string for a heartbeat packet. I would have used something more akin to a network protocol packet. It doesn't really matter as long as it meets the application's requirements and works.
But I still like my way better. :)
I can buy more hardware, but operating and maintaining systems that have been optimized to hell and back for no compelling reason is a massive ongoing human resources cost, not even counting the initial unnecessary development time.
My original point was not to encourage over-optimization, but to design better. But deploy whatever you want, I don't care.
This is a textbook premature optimization. You're designing for academic purity instead of the real world.
v=1&ready=75&busy=0
Consumers should handle new variables besides busy and ready, separated by '&', being added in the future.
I think that last line is why they avoid a binary format. Everyone who's consuming this packet already has a function available which can parse url query parameters with arbitrary name/value pairs, and that's the format being used here. I'm also not sure that the user needs to turn this into a binary format before using it; a string comparison against "75" and "0" is just as capable and nearly as fast as a numeric comparison, especially if you're using a scripting language for your heartbeat monitor rather than C.
It's very easy to glance over and see what's been set up compared to Apache's verbose pseudo-XML syntax which is about the worst syntax you can come up with: the verbosity of XML but without the benefit of being able to generate or parse it using standard XML tools!
The other thing that fucks me off is that the server will reboot happily and fail to start without having first checked that it can start, and configtest doesn't pick up things like directories being missing.
Uh, last time I checked (Apache 2.2.17 (Ubuntu)) it did.
# apache2ctl configtest
Warning: DocumentRoot [/dev/null/doesnt/even/exist] does not existJust read http://wiki.nginx.org/IfIsEvil for a start. Lots of things can only be done with ifs (unless you write a custom module) The "What to do instead" doesn't work for things that aren't simple rewrite rules.
There are seemingly arbitrary rules for what contexts you can have conditionals in, what commands can be in conditionals. There is no support for ands, ors, or nested ifs.
Personally, i'm very glad to see performance considerations being taken seriously, and even if nginx or node.js don't take over the world, its nice to see that they're forcing others to sit up and think.
Apache is "good enough" for most people.
NGINX or Node.js really don't bring much new to the table of "good enough".
It's why plan 9 isn't as popular as it probably should be. UNIX was good enough.
Just reading an nginx.conf and comparing it to an httpd.conf should be enough to convince you what nginx brings to the table.
The mandatory configuration is actually pretty minimal, especially if you are using it as an app server. Take, for example, http://kasparov.skife.org/blog/src/wombat/httpd-conf-cool.ht... which sets up a bunch of mod_wombat (precurser to now-bundled mod_lua) stuff.
On the high end, provisioning is fully automated.
Not to mention the fact that after > 15 years(!), amateur sysadmins like me can pretty much dream the config settings.
Which leaves a relatively small target audience for whom a simpler config is even remotely relevant.
And why should these five lines be buried in a 5000 line conf file?
The average user only ever has to access one small file with 10-15 lines.
There has never been a concerted effort to move Plan 9 beyond its roots as a research platform. The real world doesn't run on research platforms, and those of us trying to get real work done aren't going to invest our time in a research platform that needs enormous work to bring it up to a usable state.
You certainly didn't use nginx before it had any english documentation then.
The Mapnik GIS software (and the stack around it) used to be similar: a total pain in the ass to compile and configure, but people used it anyway because it was the best thing out there.
Mind you, I think these are real exceptions, and don't agree with the implied idea that good software will thrive regardless of how hard it is to work with.
Nobody has put effort into Plan 9's usability (which, by the way, is a much bigger problem than documentation issues), which was my whole point.
I look forward to testing it out down the road.
In this context, "major new version" means a non-patch release, with new features.
"major" new version is just English - it's a new version with significant enhancements or impact.
An example of which would be http://semver.org/
7
and
3
Oh and ... Node.js is probably going to be mentioned.
This is good for everyone ranging from small site to large because it means reduces costs etc.
Much less is a new release indicator that Zdnet will be "eating their words in a year or two" regarding Ngix/Apache market share.
On the contrary, if this release DOES improve performance a lot and reduces memory usage, on top of all the other savings, it would make it even LESS possible for NGINX to win over Apache.
Besides raw performance, there are lots of reasons to use Apache still, from the fact that it's a battle tested server with tons of documentation, know best practices, tools and modules support, knowledgable admins etc available for it. So, if performance is improved, many people won't bother switched that otherwise might have.
http://news.netcraft.com/archives/2012/02/07/february-2012-w...
The main reason is: I don't care for that much performance in the raw one standalone server case, and I can always put Varnish on top. But I do care for easy of installation, breadth of documentation, etc, and with Apache you got that in spades. This can be in even very simple things, that I can solve in Nginx in 15 minutes, like, say, Wordpress having specific rewrite rules and support for Apache built-in and not for Nginx. It's trivial to solve, but if I used Apache, I wouldn't even have to.
Given two platforms, one of which is better than the other but less popular, I usually stick to the most popular one (within reason. Like, I won't go for PHP, but I would go with Rails, not Padrino or some even less known thing).
Fewer problems down the road, and if you get into those, people have already encountered them.
Is it just me or the more you pay for software, the more it hurts you these days?
I am using and administering it from 6 years, seems pretty much okay to me if you use the .NET stack. If 18% of the top million sites are using it over completely free alternatives, they must be doing something right.
edit: forgot the "burn, karma, burn" line...
VBScript must be the hottest programming language, right?
Netcraft confirms it.
Do you really think most MC* people are eager to move on to tools that make their certifications worthless? When they move, they do along Microsoft's designated path and rarely stray from it.
Better tools like what?
Maybe there are many people out there that think that ASP.NET/C#/MVC and Visual Studio are actually better tools for them?
I know that could be an alien concept around these parts and for you but that doesn't make it any less true.
Saying that they should discard their knowledge and move to Ruby/PHP/Node is as idiotic as saying Ruby developers should ditch Ruby and move to Visual Studio/C#/.NET since it might be one of the best platforms around.
>Do you really think most MC* people are eager to move on to tools that make their certifications worthless? When they move, they do along Microsoft's designated path and rarely stray from it.
Microsoft has been building up support for Python, PHP, Node and github. Anyway, the reality is nothing close to the dystopian light that you paint them in. The jobs are no more dead-end than Java or PHP or Ruby jobs.
And no, sorry, VBScript is barely around in maybe in around 5% of companies running IIS, people have moved on to new technologies, unlike constant MS bashers who seem to be stuck atleast a decade back in their criticisms.
Except that VS/C#/.NET isn't one of the best platforms around unless you code for Windows. And that's one more reason not to move to other platforms - because the tools they use don't support the alien technology as well as what they've been using.
> The jobs are no more dead-end than Java or PHP or Ruby jobs.
You obviously have a different idea of what constitutes a dead-end job. I imagine it's a job at a company that thinks of IT as a cost of doing business, something competitive advantages are not to be derived from. Those companies will not invest in new things until everybody else is doing it and hire the same kinds of professional other companies hire. They use certifications instead of interviews because then the whole hiring process can be done within HR. I've seen a lot of them.
> VBScript is barely around in maybe in around 5% of companies running IIS
I know it's not a thorough review, but I see plenty of .asp URLs around within corporate confines.
>Except that VS/C#/.NET isn't one of the best platforms around unless you code for Windows. And that's one more reason not to move to other platforms - because the tools they use don't support the alien technology as well as what they've been using.
Code for Windows? As oppposed to what? Code for the Mac or Linux? These days most of the effort is in coding for the web.
And it doesn't really really matter to the user if the website is running on Windows Server, Linux or BSD.
>You obviously have a different idea of what constitutes a dead-end job. I imagine it's a job at a company that thinks of IT as a cost of doing business, something competitive advantages are not to be derived from. Those companies will not invest in new things until everybody eles is doing it. I've seen a lot of them.
Sure there are, but if they ran on Ruby or PHP, they would be doing the exact same thing, I don't see how IIS is relevant here. The companies want a well supported product, with an available developer base and some of them choose the .NET stack based on MS' really long support cycles.
>I know it's not a thorough review, but I see plenty of .asp URLs around within corporate confines.
If you're seeing more ASP URLs than ASPX URLs, I would say your corporate selection is skewed. There have been and continues to be a massive number of migrations away from asp over the past ten years.
Go compare the number of job listings on Dice for classic ASP developers vs. ASP.NET, it's not even a contest.
Since when exploring and using different, possibly better, technology is "on a whim"?
> Code for Windows? As oppposed to what? Code for the Mac or Linux?
I don't think Mac is a popular platform for running server applications, but I am sure Linux is a very relevant one, if not so popular in environments that reject "new" technology.
> And it doesn't really really matter to the user if the website is running on Windows Server, Linux or BSD.
Actually, it does. If your choice of technology implies higher prices, longer time for bug-fixes or added features or lower reliability, it directly impacts user experience. If your choice of technology fails to attract the best developers, software quality will suffer. That will impact user experience.
> The companies want a well supported product
If they are exploring competitive differentials through the adoption of newer technologies, they'd better realize it's not possible.
> with an available developer base
This certainly impacts the price of your labor. If you chose a technology that has lots of developers readily available, you'll be able to offer them lower compensation.
> Go compare the number of job listings on Dice for classic ASP developers vs. ASP.NET, it's not even a contest.
Job listings reflect open positions, not the number of people using a given technology.