GitHub's Metal Cloud
githubengineering.com
githubengineering.com
That's pretty funny yet sounds a lot familiar to many of us as every now and then we all do these sort of nasty hacks.
Actually, on second thoughts, I'm not surprised so much. I've read many threads on Supermicro IPMI and people's frustration with it (reliant on outdated Java, and hacked together wrappers over VNC) that make it seem like a deliberate choice to obfuscate things -just- enough to make other tools difficult.
* boot a custom live cd (a la Knoppix) over PXE
* Live CD places node into database if it doesn't exist yet,
using dmidecode to find serialnumbers and such
* Live Cd keeps querying database for instructions
* Engineer adds a profile to the node in the database
* Live Cd slices up disk to match the profile
* Live Cd fetches a tarball of the base OS from an URL
and throws it on the metal. Runs grub setup. Reboots.
Tumblr did something like this with http://tumblr.github.io/collins/.When you want to automate your complete infra, including rolling out hardware, you need hardware to develop and test on. Entropy of life will ensure that exactly that moment you need to reinstall that PostgreSQL slave from scratch, the PXE server is unreachable, or the server has a different diskcontroller, or an iLO certificate expired. Or something stupid.
Test your code. And for Ops this means: machines that are solely for the testing pleasure of Ops. No other function.
I've got systems PXE booting to CoreOS where everything else runs in Docker containers (even the odd KVM VM).
https://github.com/cobbler/cobbler
What is interesting in this article in particular are the auto-firmware upgrade state transitions, which seem pretty neat.
The chat bot is I guess neat, but if a lot of people are using the cloud that could be a hard way to get status.
PS - Hi Marty, unknown nick here but I'm sure you can figure out who I am. :D
Collins - asset management
Phil - ipxe booting based on Collins state
Genesis - base hardware configuration (firmware, bios, raid, etc), burnin, kickstart
Configrr - state / config management
0: https://github.com/xcat2/xcat-core/blob/master/docs/source/i...
http://theforeman.org/plugins/foreman_discovery/4.1/index.ht...
Maybe you were not aware because it's a plugin, we kind of have that problem in the Foreman community, plugins are not as visible as they should and they can contain key features.
But did anybody else see Hubot with a Santa hat and think that was adorable? Because I did.
I'm honestly interested in what the valid reasons to prefer Ubuntu over Debian (particularly on the server side) are.
I'm pretty dedicated to Debian on the server, a good part of my business is providing infrastructure support and setup, and I work from a MacBook Pro.
They're both pretty awful due to their automagic(al tendency to break down in mysterious ways), but if I have to choose, I'll go with Ubuntu.
Disclaimer: My main box runs Gentoo and I own no Mac machines, if that changes anything in your vision of Ubuntu users.
Those two are in opposition: Debian (generally) has new-enough packages, but it's stable, which is what one wants on a server system. Meanwhile, Debian's LTS story is better than Ubuntu's: just upgrade, and know it will work.
> easier to get non-free firmware/drivers
But how often is that needed for server systems? And of course, there're the ethical & engineering issues with using proprietary software in the first place.
> seemingly more support from third parties
There is that, but if we all wanted more support from third parties, we'd have stuck with Windows, no?
> They're both pretty awful due to their automagic(al tendency to break down in mysterious ways)
I've not experienced that with Debian in a long time. I used to have issues with Ubuntu, but I don't think that they were generally all that bad. Better than what I used to experience with Macs and Windows back in the 90s, anyway.
This is a leading word. As is a lot of that paragraph. For many tools some companies use, Debian certainly is NOT "new-enough" with many package choices. Nor is Ubuntu inherently NOT "stable" - and still trails a little behind the leading edge. As for upgrades, I've watched many a server upgrade seamlessly from 10.04 LTS to 12.04 to 14.04. I'm sure there can be and has been many a person, many a thread who've not had seamless experiences. But the same applies for Debian - heck, even the release manual has a section entitled "How To Recover A Broken System" with reference to system upgrades.
"non-free firmware/drivers"
How often needed? In this article alone, IPMI, BMC, RAID, BIOS.
"And of course, there're the ethical & engineering issues"
This is a derailment. What exactly are the ethical issues for a closed source company in using other proprietary software?
I'm by no means an Ubuntu fanatic. It has its share of issues, absolutely. I have everything from FreeBSD to Debian to RHEL to OmniOS to administer, and they all have strengths and weaknesses.
Also Dell's maintenance tooling barely works on Ubuntu at all; they don't even officially support their OpenManage stack on it. And forget about online firmware upgrades.
Maybe they'd actually used them before.
CentOS is a turd of a distribution, I'm certain the only reason it has any marketshare is because it's the only supported "free" OS for cPanel/WHM which a lot of web hosts provide for non-technical customers.
What's the matter with it?
Has someone first hand experience with such Hubot usage? Do you prefer such commands or would you want to write more informal short sentences?
To me, a wiki would be even better, because you could retroactively include expository comments along with the command history.
Does anyone knows what app does GitHub use for chats? Looks like a simple and elegant UI over Basecamp.
> [gPanel] Deploying DNS via Heaven... > hubot is deploying dns/master (deadbeef) to production. > hubot's production deployment of dns/master (deadbeef) is done! (6s)
Is this just an IMO odd use of the word "deploying" or does a DNS change really mean building and deploying a new package/image?
>Once we've gathered all the information we need about the machine, we enter configuring, where we assign a static IP address to the IPMI interface and tweak our BIOS settings. From there we move to firmware_upgrade where we update FCB, BMC, BIOS, RAID, and any other firmware we'd like to manage on the system.
EDIT - looks like it is straightforward if you control the IPMI locally. So the software would send commands to do it locally.
It's saved us a lot of engineering time and let us offer the same interface as our VMs for provisioning baremetal machines with OpenStack Nova.
There are other costs, like cooling, power, peering, networking gear, colocation/building costs, having spare parts on hand, paying sysadmins, et cetera, that are going to vary based on your requirements and region.
IMHO, requiring hardware, or the entire machine should be exception, not the rule.
And for important fileservers and databases you're going to run on specific hardware for a long time.
Now there's probably some room for debate on whether these guys job should just be outsourced to Amazon, but github has some pretty good uptime and they seem to know what they're doing, thus they've probably already won that debate.
It's easy to forget just how ridiculously powerful bare iron is these days. Go to Dell.com and see how much RAM you can cram into a U or three or four today in 2015. Or see how many IOPS a modern NetApp or Symmetrix (EMC) can push with 'flashcache' or million-dollar SSDs. It is ridiculous, and while a lot of those platforms are meant for 'building your own private cloud', etc, there's a non-trivial amount of workloads/projects where bare-iron is the best tool for the job.
Yeah, any serious I/O load is unsuited for virtualization.
* boot a live image into memory
* point LXD/RKT/Docker to /containers
* ...
* profit!Also, 1 time out of a 4-5000 or so network boot fails. Not sure why.
But it does mean no IPMI. However I built a small circuit that sits on a power cable that can interrupt said power cable with a relais that sits on a bus plugged into our server, so we can do the reboot thing.
I've been meaning to redo that power cable circuit using wifi as the linking technology, now that we have esp8266 available.
When we didn't have PXE, we had a script that told iLO to boot from CD, and that the CD was located at http://something/bootme.iso. iLO would always have network, and would pass the .iso magically to the server as device to boot from.
[1] https://github.com/joyent/sdc
[2] https://github.com/joyent/sdc/blob/master/docs/developer-gui...
[3] https://www.youtube.com/watch?v=ieGWbo94geE
[4] https://www.joyent.com/blog/triton-docker-and-the-best-of-al...
[5] https://www.joyent.com/blog/spin-up-a-docker-dev-test-enviro...
At some point someone needs to manage the actual hardware, whether that's you, or a middleman, and when you're handling hundreds or even thousands of devices its just not practical without automation.
We may eventually offer higher order infrastructure or platform services internally, but it's not our current focus.