How to set up a safe and secure Web server
arstechnica.com
arstechnica.com
Also a good idea to have your log files backed up somewhere else where your server does not have sufficient access to delete (or modify) them.
Also if you have multiple web apps running, chroot them if at all possible so that if something does break out it can't (so easily) wreak havok over your entire filesystem.
If you are using PHP also bare in mind that a common default is for all sessions to be written to /tmp which is world read and writeable. So if others have access to your server they can steal or destroy sessions easily.
I also didn't see mention of an update strategy for security updates. You can use apticron to email you with which updates are available and which are important for security.
You can set updates to go automatically (I recommend security only) but if you are more cautious you might want to test on a VM first. But keep an eye on them! This is very important, especially if you are managing wordpress etc through apt.
And so many other things that I have probably forgotten.
Having some form of audit (that tripwire can provide) is vital in those "oh fuck" moments where something doesn't seem quite right and you start wondering if you have been pwned but have no real way of actually knowing.
> If you are using PHP also bare in mind that a common default is for all sessions to be written to /tmp which is world read and writeable. So if others have access to your server they can steal or destroy sessions easily.
I'm slightly confused by this, within the context of this article. Yes, /tmp is readable and writable by all, but that doesn't mean that everything in it is readable or writable to other people. The sessions that PHP creates will be owned by the webserver user (nobody/www-data/something else), and shouldn't be readable or writable by other users.This is still a problem with shared hosting, where you might have multiple websites running on the same server. One shared host would be able to read another's sessions, because they are all running under the webserver's user. This[1] suggests overriding the session_set_save_handler and writing to a resource that only you control, such as your DB.
http://securityreactions.tumblr.com/post/36736148501/rkhunte...
I know it's a meme-grade blog but it really fits.
I sense a little bit of bias.
As a multiplatform developer I can think of a number of reasons why someone might opt to go the Windows Server route. ASP.NET MVC 4 is a first class framework and many prefer it over other popular alternatives on other platforms such as Django, Rails, and Cake. In addition, Visual Studio is arguably the best IDE available and publishing to an IIS server is dead simple.
As for cost, full versions of Visual Studio and Windows Server can both be obtained for free through the DreamSpark program for college students and through the similar BizSpark program for startups and small businesses.
1. You have to apply. 2. You have to be approved, which takes a minimum of 5 business days, and you may not get approved. 3. At some point you "graduate" and lose your licenses and have to pay for software. The faq does not say when this graduation happens, but I seem to remember hearing something about 6 months.
So, no it is not really free. And it is definitely harder to get it than simply typing:
sudo apt-get install apache2
1. It does take a few days.
2. You do need to be approved. I've heard of plenty of people signing up and getting approved, not about anyone who hasen't. I don't think it's a problem to get approved assuming you're really a startup. And for the record, you don't need to give them almost any information, save a website and filling out a few forms.
3. It lasts for 3 years. I think there's also a limit on how profitable you can be and still be eligible for the program, but it's something high. E.g., if you're making 1 million dollars in profit already, you're not eligible.
All in all, I sense from your tone that you're really looking down on this. But Microsoft is basically giving away thousands of dollars of software to any startup that wants it. Not just servers, but also Windows, the Office Suite, etc. This is a great opportunity and startups that aren't taking advantage of it are missing.
Also people discussing the Spark programs never mention one basic fact - in many countries starting a company is a bureaucratic nightmare and it's not free of charge either. You also need to pay an accountant too, as bookkeeping can get complicated even if you don't have any revenues, depending on the local legislation.
The Spark programs are basically incompatible with most startups, considering how most startups are started - an individual or a bunch of people starting something and experimenting on the side.
> Microsoft is basically giving away thousands of dollars of software to any startup that wants it
Price does not necessarily correlate with provided value, but rather with the scarcity of a resource in the context of supply/demand and the scarcity of many Microsoft licenses is rather artificial.
I use "thousands of dollars of software" every day for free, without having to be approved, without bureaucracy and without a 3-years ticking time bomb.
You also get to keep any software downloaded during the time you were a member. This includes licensing for servers in production.
There is a fairly clear explanation at http://www.microsoft.com/bizspark/about/Graduation.aspx
When one pick between WAMP and LAMP, the question one should ask is if a) what framework matches best with the developers skill/knowledge and what features the site is going to have, b) what performance requirements are there, and c) what legacy code will have to run along side. Everything else is just buzz.
The A in WAMP/LAMP is for Apache, so that does not apply here. I'm not sure what the acronym is for Windows/IIS/<database>/.NET.
The article's target audience is not professional sysadmins, but people who would like to set-up a server and tweak it for their needs.
>> I can think of a number of reasons why someone might opt to go the Windows Server route.
Sure, but practically speaking this happens most often in the corporate env. See above re target audience.
>> can both be obtained for free through the DreamSpark program for college students and through the similar BizSpark program for startups and small businesses.
Why would anyone do this? Just type in terminal three commands and you are done. I am honestly curious - why would anyone use IIS except it mandatory per corporate P&P...?
Whether or not a particular software can be used in a commercial context has more to do with licensing than how "full version" it is. In fact, even Trial versions of softwares can often be used in a commercial context during the trial period.
Versus, say, the full version of VMware Player, which while full featured and having no trial period, is prohibited via license to be used commercially.
Instead, go with what your distribution gives you. The people who put your favourite distribution together work on making the system safe and secure as a whole. People who don't think it is safe and secure file bugs and they get fixed. And you have one place to get all your updates in case fixes are needed.
If you start adding third party sources, you're on your own as to managing any implications of the way you've put it together. Just because each individual component is safe and secure doesn't mean that it is as a whole. For example, Ubuntu add hardening (AppArmor) for various server daemons which you won't get if you just download apache from the project website.
If you need a guide on putting a system together yourself, then you aren't someone who can manage these implications yourself, and you're trusting the guide author in not having made any mistakes. Are you really in a position to judge his competence?
Just use your distribution's standard web server and you'll get your safe and secure Web server in one command.
1-> auto indexing enabled (should be disabled)
2-> user directories enabled (should be disabled)
3-> server signatures 'on' (should be 'off')
4-> server tokens set to 'full' (should be 'prod')
5-> hidden (dot prefixed) files not always blacklisted as unauthorised files in Apache config
6-> same as above for editor specific back up files (eg file~). But this is only an issue for people who bulk upload or edit files on the live server (very naughty).
7-> having less optimal SSL configurations (Apache's default SSL set up isn't PCI compliant).
(IIRC there's a couple of other Apache tweaks, but that's just off the top of my head)
Then you have issues with PHP (eg logging to STDOUT so web users can view PHP errors).
And that's just covering the webserver configuration. There's still holes that need plugging if you actually want to run a secure web server; one of my favourite tools for that is using fail2ban which will auto-blacklist IPs in iptables (Linux firewall) based on if certain attacks are detected. For example it can prevent brute force attacks on SSH brute force attacks (which negates the need to run denyhosts as recommended in the article), HTTP auth, FTP, mail servers, etc.
And touching on SSH, many data centres will enable root log ins on their OS pre-installs by default. Your first job on any such box will be creating a user account and then disabling root log ins (PermitRootLogin no -> /etc/ssh/sshd_config).
Then there's a whole stack of optimisations you can run on iptables to adaptively prevent port scanning, forged TCP/IP packets and such like.
Also many distributions ship FTP, which is insecure by default. Anyone sys admin requiring FTP hosting would be better off using a chrooted SFTP environment (so the security benefits of SSH plus the sandboxed (chroot) environment that FTP often benefits).
Let's also not forget the plethora of internal services (databases, MTAs (mail servers), etc). If you're reading this article then you're probably not a trained systems administrator or not working for a large enough company where each server is running in it's own environment. Thus your MySQL / Postgre / whatever databases are running on your webservers. So you need to ensure that your database is only listening on localhost (127.0.0.1). Same applied for your MTA, unless you are intentionally providing a mail services (it is advisable to have an MTA installed even if you're not providing mail services because you can then set up daemons like fail2ban to automatically e-mail you when attempted break ins happen; which will in turn allow you to spot any determined hackers and permanently ban their IP from your box).
And lastly, if you're really paranoid, run your webserver inside a Linux/UNIX container (FreeBSD Jail, Solaris Zone, Linux OpenVZ) instead of inside a virtual machine as there have been hacks where attackers can break out of the VM and gain access to the host system (as of yet, I've not heard such attacks from containers, but if anyone has any evidence of that then I'd love to know). Running inside a container will give your webserver the sandboxed environment to work with plus an easy route for backups / snapshotting meaning minimal downtime and data loss should the worst happen and your server does become compromised.
(source: http://www.youtube.com/watch?v=hCPFlwSCmvU - it's a plugged vulnerability, but it will give you an idea about the kind of issues VMs face when compared to containers)
In short; distribution's don't ship secure defaults. They ship a happy medium between usability and security. However if you have an internet facing box (particularly one not hidden behind a hardware firewall, as many budget set ups are not), then the happy medium isn't secure enough.
1) Changes that give you relatively little gain from a security point of view at the cost of massive usability. I'd be happy to run production servers without these changes, and so I claim that these are debatable.
2) Changes that stop inexperienced admins accidentally compromising their servers, at the cost of usability/complexity. Debatable because they're less about security and more about not shooting yourself in the foot. Eg. fail2ban. I use ssh keys only. Brute forcing my ssh won't work, and I'd prefer to not allow outsiders to cause firewall rule changes on my servers based on a script that parses log strings with regular expressions. I trust ssh's security more than I trust fail2ban in not having a vulnerability.
> And lastly, if you're really paranoid, run your webserver inside a Linux/UNIX container
This is pretty much what AppArmor hardening does. LXC (Linux containers) are implmented via AppArmor. A compromised httpd daemon restricted in this way won't be able to much anyway (eg. open outbound ports or go to other areas of the filesystem). And this is what Ubuntu ships by default for many daemons.
> Also many distributions ship FTP, which is insecure by default.
Nowadays this stuff is only installed if you requested it, which is different from "ship"ing it. If you use distribution defaults, you won't have an FTP server installed. If you choose to install one, you'll get the weakness whether you use the distribution, tune it yourself or install from a third party source.
> In short; distribution's don't ship secure defaults. They ship a happy medium between usability and security. However if you have an internet facing box (particularly one not hidden behind a hardware firewall, as many budget set ups are not), then the happy medium isn't secure enough.
Secure enough for production. I have plenty of servers on distribution defaults that haven't been compromised (and I have experience in dealing with others' servers when they have been compromised, so I don't think I'm oblivious). You may prefer adding additional hardening, but it is unnecessary and relies on you knowing what you are doing.
Finally, you've missed my biggest point. For someone new to running servers on the Internet, who can he trust for guidance? You, me, a random guide on the Internet or a distribution vendor?
If you don't mind, I will address your points individually:
"Changes that give you relatively little gain from a security point of view at the cost of massive usability. I'd be happy to run production servers without these changes, and so I claim that these are debatable."
Which items specifically prevents usability? None of them do aside the FTP (but SFTP is supported by nearly every FTP client already) and turning off PHP error logging to STDOUT (really not recommended for production servers!). Everything else doesn't reduce usability for authorised users. And I resent the claim that there's little gain as even on my servers which have a low public profile, I get reports of multiple brute force attacks a week (on average, about separate 5 attacks a week, all of which I manually add to a permanent firewall blacklist when it becomes obvious that they're just retrying attacks everytime fail2ban's autoban expires).
"Changes that stop inexperienced admins accidentally compromising their servers, at the cost of usability/complexity. Debatable because they're less about security and more about not shooting yourself in the foot. Eg. fail2ban. I use ssh keys only. Brute forcing my ssh won't work, and I'd prefer to not allow outsiders to cause firewall rule changes on my servers based on a script that parses log strings with regular expressions. I trust ssh's security more than I trust fail2ban in not having a vulnerability."
You don't understand how fail2ban works if that's your primary concern. Fail2ban isn't public facing, it just monitors logs and then adds an iptables (or any firewall you chose) rules based on the results of the log files. It's iptables that is public facing and thus needs to be protected against vunrabilities; and iptables is proven technology already. What's more, preventing brute force attacks doesn't reduce usability nor prevent you from shooting yourself in the foot; preventing brute force attacks is the !!bare minimum!! you need to do if you have an internet facing server. This is security 101! I do agree with you that SSH keys are preferable to password log ins, but that also adds complexity and can reduce usability if users need access from multiple locations and platforms (there are workarounds, eg memory stick with the key on), and then you still have the issue of brute force attacks on every other log in service.
"This is pretty much what AppArmor hardening does. LXC (Linux containers) are implmented via AppArmor. A compromised httpd daemon restricted in this way won't be able to much anyway (eg. open outbound ports or go to other areas of the filesystem). And this is what Ubuntu ships by default for many daemons."
You missed the point about easy snapshotting though, which is why I raised the point about containers in the 1st place. I will grant you that you can have back up solutions on bare metal systems, but snapshotting is way more usable (which is ironic given that's been your key argument). Plus I did say that advice was for the paranoid (ie not really all that necessary unless you fancy tinkering).
"Nowadays this stuff is only installed if you requested it, which is different from "ship"ing it. If you use distribution defaults, you won't have an FTP server installed. If you choose to install one, you'll get the weakness whether you use the distribution, tune it yourself or install from a third party source."
You're just reiterating what I said though as I didn't say distributions install it by default. However you missed my point there as well, as I was describing how FTP isn't ideal for production systems (FTP sends clear text passwords and doesn't behave itself behind firewalls and/or NATing without adaptive routing (which means that FTPS often doesn't work behind some firewalls).). Thus chrooted SFTP is a much saner solution for production systems.
"Secure enough for production. I have plenty of servers on distribution defaults that haven't been compromised (and I have experience in dealing with others' servers when they have been compromised, so I don't think I'm oblivious). You may prefer adding additional hardening, but it is unnecessary and relies on you knowing what you are doing."
You've been lucky. But what your advocating is little more than 'security by obscurity' rather than pro-actively hardening your box. Having worked on a number of high profile web infrastructures (including farms that take web payments), I wouldn't be doing my job if I ignored my steps; in fact our business wouldn't exist as we would break UK laws.
The Apache configuration I described is ridiculously easy to set up, as is fail2ban. They are the bare minimum any web server should implement. Which is why I placed emphasis on them by placing those suggestions at the top of my post. However if we're going to advocate complacency in order to make our lives easy, then lets also not bother keeping our software up to date, because that's a real pain in the arse at times :p
Preventing brute force attacks is not the bare minimum. The bare minimum is to not be vulnerable to brute force attacks in the first place. ssh already has built in protection for this, and it slows brute force attacks to the point where such an attack is impossible in practice, unless you have weak passwords. Mandate large enough keys only, ban passwords, and you're done. Allow password authentication is your real security vulnerability here. Patching over it with fail2ban just hides the issue.
I do understand how fail2ban works. "Just" monitoring logs isn't good enough. The data that appears in logs is not generally considered to be a security sensitive channel. It's string data with poorly defined delineation. It should not be trusted for automatic use, since every channel that dumps data into log files is not vetted for security.
> You've been lucky.
I disagree. You claim I'm advocating security by obscurity, but you're the one who seems to think that hiding version strings gains in security. That's security by obscurity, since you can determine the version of software by observing its behaviour (or just not caring and trying your attack anyway).
I'm not advocating complacency. We just disagree on what complacency is. I claim that if you keep your software up to date (easiest if you do follow distribution defaults), then you are sufficiently secure. The overwhelming majority of security compromises happen because people fail to run updates. The second largest cause is because of vulnerable misconfigurations that people have introduced. Only a tiny fraction of compromises come through a default distribution installation that your hardening would catch, and these holes are rapidly patched by vendors and a simple update will close them. I think that it is much more likely that you'll open the second cause (an introduced misconfiguration) if you try hardening and you don't know what you're doing. Which is why I recommend sticking to distribution defaults.
I didn't say reports increase security, and Fail2ban does more than just reporting, it actively blocks brute force attacks. It's a bit difficult us discussing the merits of certain security measures when you keep focusing on the irrelevant as if those were my security suggestions - it's almost as if you're trying to 'death by a thousand paper cuts' my whole post simply to win an internet argument <_<
"Preventing brute force attacks is not the bare minimum. The bare minimum is to not be vulnerable to brute force attacks in the first place. ssh already has built in protection for this, and it slows brute force attacks to the point where such an attack is impossible in practice, unless you have weak passwords. Mandate large enough keys only, ban passwords, and you're done. Allow password authentication is your real security vulnerability here. Patching over it with fail2ban just hides the issue."
For SSH, you'd be right. But as I've repeatedly said, not all services offer key based log ins and sometimes there's a business requirement for password log ins on systems that could be managed with keys. You're original arguments were about usability yet the arguments you're making now are the lest flexible suggestions raised thus far!
"I do understand how fail2ban works. "Just" monitoring logs isn't good enough. The data that appears in logs is not generally considered to be a security sensitive channel. It's string data with poorly defined delineation. It should not be trusted for automatic use, since every channel that dumps data into log files is not vetted for security."
Logs are fine for parsing as you have to be compromise before the logs are comprimised. But which point, it's already too late.
"I disagree. You claim I'm advocating security by obscurity, but you're the one who seems to think that hiding version strings gains in security. That's security by obscurity, since you can determine the version of software by observing its behaviour (or just not caring and trying your attack anyway)."
In theory I'd agree with you, however a great number of compromised systems were attacked by opportunists scanning version numbers looking for boxes to target with known vulnerabilities. Plus, and once again I'm having to repeat myself, those specific changes I advised are actually required to comply with many compliance laws (eg PCI compliance, required if you make finance transactions in the UK).
"I'm not advocating complacency. We just disagree on what complacency is. I claim that if you keep your software up to date (easiest if you do follow distribution defaults), then you are sufficiently secure. The overwhelming majority of security compromises happen because people fail to run updates."
There's no such thing as 'sufficiently secure' as that only depends on the attackers targeting your system. Today you might be 'sufficiently secure' because your box has not been spotted by any keen attackers, tomorrow might be different.
Also, I'd be more inclined to agree with you if all of the examples you've given weren't off topic from the points I raised or just incorrect (eg changing Apache config being dangerous and/or hard, fail2ban reducing usability, etc).
"The second largest cause is because of vulnerable misconfigurations that people have introduced."
I'd go along with that. I've often said 'users are the biggest security risks' :)
"Only a tiny fraction of compromises come through a default distribution installation that your hardening would catch, and these holes are rapidly patched by vendors and a simple update will close them."
None of the configurations I mentioned (bar the list of paranoid ones) fall into that category; non-optimal configuration isn't a hole that gets patched. Plus even if it was, it wouldn't be fixed with software updates as package managers tend to avoid over-writing live config files else they'd risk doing more damage than good.
"I think that it is much more likely that you'll open the second cause (an introduced misconfiguration) if you try hardening and you don't know what you're doing. Which is why I recommend sticking to distribution defaults."
You can't make a system less secure by changing the settings I recommended as the distro defaults are already on the most open defaults. Plus, and once again I'm repeating myself, the configurations I'm recommending are incredibly easy to implement.
I find it odd that we're actually arguing about whether it's worth making the most basic of changes based on the assumption that those people in question are stupid and their box will probably be ok. Surely a better approach would be to suggest optimisations; guiding them through the process if needs be? After all, it's too late to regret using the defaults if and when you get hacked (and in my line of work, I've had to fix quite a number of boxes where the sys admins have been content just running with the default settings).
Not unless we're people with excellent reputations that "those people" can recognise, or it is possible to determine that people with these excellent reputations endorse our advice. Otherwise we're just adding to the muddle of information on the Internet, some of which is bad, some of which is good, and it is impossible for non-experts to tell the difference.
My argument is that the distribution is such a reputable source, and that a random article upvoted on HN isn't. If a distribution ships insecure defaults, then you should petition them to fix the defaults rather than publishing "fixes" elsewhere. You'll have to fight your corner, of course, against a bunch of people who might differ from you in your opinion about what is and isn't secure :)
https://gist.github.com/c6fd22f73468b26e01b0
I built it from scratch (ish) so I know what all the parameters do.
Do you have more info on the SSL PCI compliance?
For SSL PCI compliance, have the following config as part of your SSL settings (which you've got commented out currently):
SSLHonorCipherOrder On
SSLCipherSuite ECDHE-RSA-AES128-SHA256:AES128-GCM-SHA256:RC4:HIGH:!MD5:!aNULL:!EDH
SSLProtocol -ALL +SSLv3 +TLSv1 +TLSv1.1 +TLSv1.2
This should force Apache not to default to older insecure SSL protocols and disable SSL compression (HTTP compression via mod_deflate still works here) which leaves HTTPS open to attacks like BEAST.Bare in mind I'm still testing the above code myself (funny enough, that's actually what I'm doing this very minute) as the BEAST vulnerability is still relatively new (or rather, new enough where it wasn't part of PCI compliance until the last month or so). I'll update this thread in the next few hours if that code doesn't work, but I can't see there being a problem as it follows the standards defined in Apache's manual.
Also make sure you have OpenSSL version 1.0.1 installed (required for TLS1.1 & 1.2). You can check this by running: openssl version from the command line. However if your system is built from a package manager and has been kept relatively up to day, then you shouldn't have a problem there.
SSLProtocol ALL -SSLv2Next up they just HAVE to show us how to setup our own quakeworld or UnrealTournament'99 or Quake3 server! ;-)
It was an amazingly short sighted move, but such things were typical back then. And it's only from learning the hard way that we've managed to get to the stage we're at now.
However I think it's often forgotten that servers need a different set of security profiles depending on the server's role and where it is sat. For example, a webserver sat behind a hardware load balancer wouldn't necessarily need much SSH protection as the webfarm HTTP traffic should be on a different VLAN to the internal systems administration traffic (which in turn, would be another different VLAN to the company's staff VLAN). So it would be almost impossible to get access to an OpenSSH log in, let alone attack it. Where as most consumer VPS solutions put all their customer servers in the DMZ, which means it's up to the customer to provide software preventions to harden against access that would normally be protected with a complex hardware solution in more professional / clustered set ups.
And this is why you can't fully trust default configs; there simply is no "one size fits all" solution so package maintainers instead opt for the best compromises.
My point is that there is more to security than installing an OS and running regular updates.
This goes on with the choice for Ubuntu Server. Why? Is it an article about "safe and secure web server" or about "how does my grandma set up a server"? There are much more choices in terms of reliability and proven track record like FreeBSD, OpenBSD, Debian, RHEL/CentOS. The choice was made because it's easier to set up and apparently the author is too lazy to _really_ do his homework.
In the end, i'd say if the articles title would be "beginners guide how to setup a server" i wouldn't comlain..
In my case that's Node.js, Redis, nginx and occasionally python 2.7, and of those I'd be installing Redis (I often run beta releases) and nginx (want to compile in my own modules) by hand on Ubuntu as well. Sure this is slightly more work, but it gives me more control and I feel a more stable server environment.
My personal view:
- my personal server runs Debian but i would place ubuntu a second for a one-person thing with no personal data on it. I know it runs a stable and updated OS and it just runs and runs and runs. You need a newer software version? You can selectively take some from testing or even unstable if you have to. You still get timely updates and live in an eco-system that is largely seen as proven, stable, excellent. Ubuntu (afaik) only get's Debian testing packages anyway, but still needs to patch them or tweak them. Or you get non-upstream packages. Also my experience is that Debians configuration is quite secure by default. Whereas Ubuntu tries to make it easy for you but also more open to security flaws.
- for an enterprise i can totally see why many choose RHEL (because of support, fast security updates and great in-house know-how. Do you know how many kernel developers are employed by Ubuntu? You don't want to know.. I give a lot of Kudos to RedHat for being such an open and contributing company.)
- for a start-up i can totally see how FreeBSD might fit as an internet-facing frontend. It's a stable and fast workhorse, and you'll end up compiling, configuring, installing the core components of your system all the time anyway.
Talking about it i feel a bit sad of not mentioning Solaris anymore. Solaris "was" awesome.
http://h10010.www1.hp.com/wwpc/us/en/sm/WF05a/15351-15351-42...
Has ECC RAM support. Takes 4 3.5" hard disks, and runs very quiet and cool.
curl http://www1.hp.com
curl: (6) Could not resolve host: www1.hp.com; nodename nor servname provided, or not knownThat combined with the costs of getting a static IP and a few other things I'd want for hosting @ home, and I decided to use a VPS instead.
Do you have one of these? Can you answer some questions for me, as I'm just about ready to impulse buy!
Is it passively cooled? Is the "embedded raid" an actual RAID or some junk software emulation (I'll probably use ZFS anyway, and that should be used w/out hardware RAID)?
EDIT
Not passively cooled :( Full review here: http://www.silentpcreview.com/HP_Proliant_MicroServer
I feel like doing this stuff by hand should be considered insecure and outdated..
Virtualization was build for server providers to make easy money, not for server owners to gain performance advantages.
Vistualization is not for production. Production servers need less code, not more.
It is the same kind of mistake as JVM - we need less code, integrated with OS, not more "isolated" crapware which needs networking, AIO and really quick access to the code from shared libraries.
And, of course, a setup without middle-ware (python-wsgi, etc) and several storage back-ends (redis, postgres) is meaningless.
Update:
Well, production is not about having a big server which is almost always 100% idle, and can be partitioned (with KVM, not a third-party product) to make a few semi-independent virtual servers 99% idle. This is virtual, imaginary advantage.
On the other side, your network card and your storage system cannot be partitioned efficiently, despite all they say in commercials. And that VM migration is also nonsense. You are running, say, a MySQL instance. Can you migrate it without a shutdown and then taking a snapshot of an FS? No. So, what migration you're talking about? It is all about your data, not about having a copy of a disk-image.
It is OK to partition development, or 100% idle machines - like almost all those Linode instances, which have a couple of page request in a day - this is what it was made for, same as old plain Apache virtual-hosting. But as long as one needs performance and low latency, all the middle-men must go away.
Most server operators don't care about performance. They have performance coming out of their ears. They care about redundancy and maintenance, or to put another way cost centres.
Your post is on the wrong side of history. Virtualisation is being rolled out in a massive scale right now. Essentially you can abstract your entire physical infrastructure away from your logical infrastructure.
You have a physical server die? The HV has already moved the image to a new node and started it. Before you even receive the e-mail notification the new server is already booting.
So now a hardware failure goes from being a massive panic, to being a small annoyance. You pull the dead hardware from the rack, and plug a new generic node in and that now becomes available for the HV to use.
You want to back up a server? Take a copy of the ENTIRE image in one go. You want to deploy a template? Well that's trivial with images. You want to do change management with the servers? Just put the images in GIT. Boom done.
What you're suggesting is essentially taking the cheapest bits of server management (i.e. the physical hardware) and acting like they're the most expensive bits (i.e. people, time, and flexibility).
Automated server management has been done long before virtualization stacks emerge, and it about utilizing monitoring and network boot.
What virtualization stack Google uses on its servers? None.
Most businesses don't have hundreds of identical servers and services. They have a few dozen very specific or niche ones which need high up-time. This is one area where virtualisation can play a great role.
Another example is someone like a host where they want to distribute resources without any human intervention (e.g. Virtual Private Servers, shared hosting, etc).
All in all you're now starting to see data centres turn into "dumb" hardware farms, with the logical design and deployment being handled up-stream. This even extends to things like networking (routes, switches, etc - all centrally controlled).
Is it possible that the inefficiency and added unnecessary complexity makes a virtualized servers unusable for the most common server's tasks, due to I/O interference and cache/memory access complications?)
It has to do with organisation and about being able to abstract logical servers away from physical hardware.
There is very little inefficiency (see hardware assisted virtualisation point above) and very little overt complexity (go play with any modern HyerVisor solution).
As I said above you don't understand how virtualisation works. Your points about "I/O interference" and "cache/memory access complications" just make you sound ignorant.
On the other hand, I'd been involved in a few projects, which includes optimization of a big centralized databases, so, I think, I know a bit about flows of data, access patterns and where the bottlenecks are (hint: around serializing and scheduling low-level I/O operations).
Try to look under the surface structure, which plain words are.)
They even developed a cluster management tool for Xen/KVM: http://code.google.com/p/ganeti/
To further on this excellent and valid point, very cheap virtualisation was available on a lot of other very solid and very productive architectures and has been for decades; it was just x86 finally catching-up late. There are a whole lot of reasons speaking FOR virtualisation and only a few very specific applications where it might be a bad idea. I have no idea what OP up there is all about and against it, they make no sense.
Also...Is Amazon EC2 considered virtualization for you?
Running closer to the metal often means higher development costs, which vastly outweigh the performance overhead abstraction layers they incur.
You're right about some things getting messed up by virtualization, but you're over-estimating the impact this has on running an online service.
Chances are your physical boxes will have one resource almost always 100% idle and having your workloads virtualized allows you to better use them.
Also, virtualization allows you to more easily fix problems - you just delete the box and rebuild it from a base image and the packages you need. Reimaging a physical box is a pain.
Obviously, if you are into HPC and very low-latency, you need physical hardware. You probably need some specialty networking hardware too and maybe less OS than most of us would like to have. But then we are no longer in the general purpose amd64 box world anymore.
I remember one instance where we were having some trouble with one batch-processing workload. The amd64 box we had for that was not enough, despite having much more memory and processing cores than its peers. I suggested we moved the workload to an almost abandoned Itanium box we had played with for some time. The HP machine crunched the workload more than 3 times as quickly. I noticed the working datasets could fit completely inside its (then) huge L3 cache (16MB, IIRC) but not in the 4MB of the amd64 machine, more or less eliminating the external memory bus penalties.
The last one is more interesting. Have you noticed that your solution was in a splitting workload, not in sharing it with other processes?
In general, splitting (dividing/partitioning/isolating) is a good idea, sharing (sharing resources) is a bad one.
Sharing, unless your data is read-only, like a txt segments of code, is a source of problems, not a solution. Partitioning, on the other hand, is a universal, natural way.
The first two paragraphs are about one benefit: using otherwise idle resources. At my lab, we have five powerful servers for data-processing experiments and non-time-critical analytics. Most of the time they sit completely unused. By installing Xen and running things in virtual machines, we have been able to also put tens of differently configured web servers for various small projects on each.
The third paragraph is also very valuable from experience. It has been very useful to keep documentation around how different servers are set up, and verifying that the actual setup still matches the documentation. This has been easiest to ensure by having scripts that create a new VM, install packages and make configuration changes - if in doubt, completely delete the VM and re-create it from scratch with a single command-line.
I don't buy your argument about latency one bit, do you have ANY data to back up your statement? You know - Google and many other big players run huge virtual machine clusters, you'd think they wouldn't if that almost unmeasurable difference in latency had an impact.
I would run KVM on top of a machine, even if there was only to be one VM on it, just because moving, backing up or scaling up or down can easily be done. You can also easily migrate the host to another physical machine if you are experiencing hardware issues, or just upgrading your rig. And yes, you can migrate a running server, I'm doing it all the time.
And yes, you can snapshot and take a backup of a running server, but what goes on inside of MySQL, or other applications for that matter, has nothing to do with the consistency of the disk image. When you snapshot a disk image, you naturally don't know whats in RAM.
All you can do is instruct MySQL to pause, flush all its write buffers, and then take a snapshot of FS, then move it on. But this procedure has nothing to do with whether or not it runs under, say, VmWare or not. It has anything to do with does this particular disk volume supports FS snapshoting.
The next question is - what use of that snapshot when your system crashes and you got lost all the changes made since the last snapshot? How a VMWare helps you here?
Now consider what a bottleneck naively virtualized (represented as a file in a host system) disk volumes become, when I/O operations on, say, your DB's physical (transaction) log got interfered by I/O operations of your syslog daemon, or whatever other activity is going on.
In a database world the solution is about decoupling, partitioning and avoiding any I/O sharing possible. So, virtualization is just another layer of complexity which makes everything less predictable and controllable.
For total system failures you still need a proper backup solution on the machine. This is true even if it's dedicated hardware or a virtual machine.
You seem to always resort to talking about databases, and this might be true for huge centralized databases, but that is a pretty damn specific task. Also, I thought we were past the "put everything on one box"-model.
The complexity you talk about is just not there. It acts and feels just like a normal machine, and you have yet to provide any data that would support your claim, even for edge cases.
Something worth reading: http://en.wikipedia.org/wiki/X86_virtualization#Hardware_ass...
What virtualization allows you to do is optimize the organizational responsibility. If I am developing internally on my VM, when it gets certified by ops management, and rolled out onto the live VM stack, I can be very sure what I am being responsible for; the ops also can be sure (because they maintain a huge VM filesystem, mostly) that they don't have to deal with things at a very thin slice.
For web operations, look, its simple: VM is good enough because there is so much overhead all over the typical web stack that a few frames of difference are, largely, irrelevant to the use case. Okay, don't put your multiplayer live realtime universe hashes in a VM; those belong bare metal. But the web front ends that are serving the same old content, over and again .. these are well worth putting in a package which can be massively deployed at whim.
A typically well-packed VM, consisting of only the hand-optimized built image of a development team with this in mind, can be a very, very tight package. I've seen reflectors and mongodb proxies and so on, packed into 256meg boot image that can be replicated simply by giving it a new name .. deploy 2,000 copies of these VM's, and you have a massive solution to the front-door problem..
Rubbish. I do this regularly.
I agree entirely with this comment. Between lun size limits, license management (which is hell), the general additional work and the fact that the benefits are suspect, I agree.
Processes and multiprogramming were invented to virtualise batch machines. That's as far as it needs to go.
I'm not impressed with the industry trend.
Many if not most of the production servers on the web run some sort of virtualization; Xen alone powers EC2, Rackspace, Linode, Rimuhosting and many others. In my experience it only adds about 3% of CPU overhead as a disadvantage.
If you have a small server, I'd really recommend checking out these scripts that assist with configuring and setting up a server very quickly: http://lowendscripts.com/wiki/shell_scripts
I personally used a fork of lowendscript last year to set up some servers, but if I had to set up a new server today, I'd check out some of the other other options at that link, like Minstall: https://github.com/maxexcloo/Minstall But this Xeoncross lowendscript fork is still very active: https://github.com/Xeoncross/lowendscript
Say I worked out your home IP (not hard), then sent a large number of failed SSH attempts with the IP address forged as yours. You are now locked out if your home server.
Fail2ban is actually a vulnerability in itself.
That's a bit harsh. It's true that you may have to tweak some settings to prevent or minimize DoS attacks, but even that risk is a far cry from an attacker gaining a login or rooting the box. Fail2ban has proven to be safe and reliable in the years I've used it. Nonetheless, the old maxim holds true: Know your tools.
If you are already able to whitelist your (valid) login points, why would you need fail2ban? Just whilteliste them in your firewall and/or /etc/hosts.allow.
Personally I've yet had anyone bruteforce my ssh-key (although, as I run Debian, that is just luck as it turned out...). Still, fail2ban wouldn't really have helped against an attacker that knows/can figure out my access token off line...
Fail2ban isn't a firewall. It monitors logs for suspicious activity and responds with an action (not limited to banning an IP). When you expose services publicly, it's one of many tools you can use to limit bad behaviour without penalizing or inconveniencing legitimate users. I also use iptables (including the recent and string modules), RBLs, and a host of other access controls. Security is a layered approach and redundancy isn't a bad thing.
BAM.
I always make sure I have some access to a network KVM or remote console for my servers so I would be able to unblock myself.
If you want to host a server from home, you have a few choices:
a) Obtain a static IP address from your ISP.
b) Obtain a domain name from a dynamic DNS provider and map that to your current IP address. Update the mapping as necessary.
c) Just use your existing IP address -- this is how P2P or Skype works, for example.
Of course there is the downside that no one except those very few people that has ipv6 will be able to use that service, but it is free and it can be a very useful way to access a computer behind nat.
Otherwise, you just use your IP and make sure to edit the DNS yourself when it changes.
[0]: http://dyn.com/dns/
No big surprise here, since the router is really just a nice little Linux box.
can't believe that hasn't been on HN before, so added it: http://news.ycombinator.com/item?id=4841329
Locked: "This question exists because it has historical significance, but it is not considered a good, on-topic question for this site, so please do not use it as evidence that you can ask similar questions here. This question and its answers are frozen and cannot be changed. More info: FAQ."
At least the list of references there is useful.
http://howto.biapy.com/en/debian-gnu-linux/system/security/h...
also, chrooted sftp-only accounts
http://en.wikibooks.org/wiki/OpenSSH/Cookbook/SFTP#Chrooted_...
Another solution that is much easier to configure compared to iptables/fail2ban.
It's a bit different if you're expecting said site to get 100K+ views per day or is going to host some big database, but even then I'd probably run it in the cloud to save on bandwidth costs.
I do wish the article gave more attention to VPS providers like AWS or Rackspace, but I don't see how anything in the article could discourage someone from administering their own server (that is, of the set of people who aren't already put off by scary command lines). 90% of the article is relevant to someone with a VPS.
FWIW, my current costs for a single dev server at Rackspace (initially Slicehost) are about $12-$20/month ($240 max for a year), which probably costs less than the electricity my current computer is using.
Even were I to want one on a webserver on the localhost I'd probably just install webmin or something, which you can do equally well on a VM somewhere else.
That said, I wouldn't use a GUI for such things, but I would still prefer my own box. The big power button is a lot easier than EC2s endless menus and options. I say this from the perspective of someone who's used EC2 for a couple of project - imagine someone who's never seen it before.
A harddisk seek + reading 1 MB sequentially is something like 30 times slower than SSD and 120 times slower than reading from RAM. Disk seeks are what really kills you, as a disk seek is 20 times slower than a roundtrip within the same datacenter.
Sending 1 MB of data over a 1Gbps network using a disk seek and a sequential read is 3.6 times slower than doing the same with SSD. For big files with a server that has to serve many concurrent requests you'll end up doing several disk seeks for the same file. And we've been talking about the ideal too, because data on disks can get fragmented and so reading from the same file does not guarantee a sequential read.
This is why Varnish does caching, even when serving static files. If you want high performance with low latency, then Varnish works better than something like Nginx.
As an example, right now I'm working on integrating with an OpenRTB ads exchange platform. My server has to respond within 100ms total roundtrip (with an upper bound of 200ms, but they complain if you consistently respond in over 100ms). It's enough to say that everything to be served has to be already in memory, as doing anything else means that window cannot be met (consider how a single disk seek can cost something like 10ms).
So yeah, SSDs can be useful.
That's not really the point though, since the big motivator for SSDs isn't bandwidth, it's latency.
> Temporary files usually start with a dot or a dollar-sign.. to make sure that Nginx never serves any files starting with either of those characters...
> location ~ ~$ { access_log off; log_not_found off; deny all; }
Wouldn't that regex match temporary files ending with ~ (as it should)?