Sysadmin war story: the network ate my font
verticalsysadmin.com
verticalsysadmin.com
And then the printer you are migrating the MICR device to has decided it doesn't understand the PCL in the font file anymore.
This is 2017 and I just dealt with this today.
They also pride themselves on the quality of PDF output (and have full (iirc) support for PDF-A).
The font check alone can be done in a few lines of Python: http://beza1e1.tuxen.de/articles/preflight.html
And it turns out that the answer will probably be, "well, we asked the lawyers, and you need rights from the font owner to do so, and it makes files bigger, and most stuff is in whatever Microsoft Office uses as their default font this year anyway."
And so we're torn between "yes, of course the designer who built the font should get paid" and "a sane default became a checkbox because of US copyright law".
And most of the time, the letter spacing is the default given by the font, so you just have to encode the position of the beginning of a run of text. So it is pretty space efficient, too.
To make this work effectively, printers would cache fonts. That saved on overall file size, which was important for storage and transmission. But the real driver was rendering speed. Most documents are pages of text at a small number of sizes and there are a small number of letters.
If you're going to print 300 lower-case "e" characters, all in 10.5 point Times New Roman, it would have been ridiculous to do the hard work of rendering the bitmap from vector each time. You render it once, cache the bitmap, and then just plop the bitmap in the right spot.
I know this because circa 1993 one client had me build a custom font that varied letterforms slightly to mimic a hand-lettered effect. They ran these weekly newspaper ads for their big wine store, and the were paying a guy to hand-draw the whole thing. They wanted to keep the casual look, but save on the cost. (And I presume the guy was kinda tired of writing the same things over and over, but I never met him.)
I learned enough PostScript to make it happen, decomposed each letter into strokes, and then drew the strokes with slightly different alignment each time. It worked fine in simple tests, but the first time I rendered a page, I thought I had broke the printer. Instead of the printer's top speed of 8 pages/sec, a full page took over 15 minutes.
So as usual with "why don't they just" questions, the answer is, "because 'just' is sweeping some things under the rug". It was harder than it looked at first glance.
Graphic designer designs cheque. For the design to be signed off he/she includes 'lorem ipsum' placeholder text for the special numbers at the bottom.
Design gets signed off, a template is made for the programmers to use.
In the code a third party library is used to make it 'easy' to create a PDF. This process consists of opening the template, adding a line of text to it and printing it as a PDF to another file, ready for the printer.
A little while later the graphic designer edits the template to make a few amends. The file is re-saved, this fresh copy no longer contains the fonts not in the document. The placeholder text having long gone, the font for it is not saved. The other fonts for the cheque are, they moved across and were catered for by the software in the updated template.
The software runs exactly as before, just the template file has been updated. However now the font is not found unless installed or cached on the computer or printer.
The programmer never had to embed the font, his/her third party library abstracted that requirement away. The programmer had worked with the library before and knew that it was best to use Helvetica because PDF knows that is a built in and therefore does not need to bloat the document with the default fonts. Any other font would add megabytes to the document. So there was probably no oversight by the programmer.
However there may have been a micro-manager that was 'responsible' for micro-managing the update to the template. This probably involved meetings and conference calls and deadlines for 'the project' on the whiteboard. Not wishing to overstretch programmer resources, the micro-manager took it on himself to make sure the programmer was not 'interrupted' so he got another lackey to upload the new cheque template. This all worked fine initially.
Had there not been a micro-manager then the graphic designer would have had to have worked with the programmer, without the micro-manager or his lackey. The programmer would have picked up on the smaller file size as this would be a noticeable change. Instinctively the programmer would have made a test run and, not having fonts on his/her dev box, would have detected the problem right away. Meanwhile, the lackey with no knowledge of things like version control just uploaded the template as told, blindly unaware of the requirements to 'check your work'.
Why didn't the micro-manager check that the font wasn't embedded? That is what I want to know.
Historically the size of hard-disks have been very limited, so it was important to save a file with as few bytes as possible.
Also the serial transfer speed to the printer was very low, so it was important to limit how much data that was sent for each print job.
1. That can't happen.
2. That doesn't happen on my machine.
3. That shouldn't happen.
4. Why does that happen?
5. Oh, I see.
6. How did that ever work?To be fair, fonts really do pave the way for malware into your kernel - https://googleprojectzero.blogspot.no/2015/07/one-font-vulne...
I got pretty sick of opening this file after half a day, because it had tons and tons of CR/LF's and my editor at the time rendered these with strange characters on the screen .. and I didn't like that, as a junior guy, so I just replaced the LF's with 00's. For some reason, this just worked fine in my local environment, and I was able to validate the data in the file before sending it off to the spool for printing, later in the afternoon.
About 3 minutes after I closed the job, the building got a fire alarm, and we all had to exit. Apparently there had been smoke detected in the operations room, where the printer was located, so the Halon systems went off, and we went into full-blown "Ops Reset" mode.
After an hour of hanging around the parking lot, I was called in by the head honcho's in the Ops Room, sat down in front of the printer, and told to explain myself.
Well, turns out, I was responsible. The lack of LF's in the text file meant that the printer was printing - as fast as it could - every single line on the very first character position .. and after a few minutes, the printer simply caught fire.
Oh man, since that day (mid-80's), I've eschewed any job that requires me to deal with printers, and I've been anti-printer ever since. ;)
The network ate my webpage...
Dang cache.
My own personal strangest experience was having two autonegotiating 10/100baseT devices that wouldn't speak to each other but would via any other switch.
One or two years ago, I was tasked to solve the following issue:
Installing an obsolete RHEL 4 on a brand new computer, with all the drivers issues that entails. Virtualization was not an option.
The solution I chose was to "backport" the latest RHEL 5 kernel in the RHEL 4. A few weeks of works later repackaging the kernel, adapting the mkinitrd script and a lot of headaches around the install iso (it was actually the hardest part, with a lot of hacks around anaconda), I finally managed to have a working server.
Then a few days latter, I was notified of a regression.
There was a somewhat crazy application that was managing various configuration files on various devices of this particular infrastructure which stopped working.
Digging in this application, I discovered that it was "pushing" the updated conf files through NFS, more exactly, it was notifying a service on the device targeted, and then this device mounted an NFS share on the RHEL 4 server, recovered the conf file, and unmounted the NFS share (I told you, "crazy").
The targeted devices were using RH 7 (RH, not RHEL, I'm talking about the one with a 2.4 kernel).
Strangely enough, mounting the share by hand and doing an ls showed that the files were indeed present, and there was no permission issues.
Reading the source code a little further and looking at quite old QT versions (the service on the device was QT based). I finally managed to find which line of code was not working: it was a simple readdir().
So I created a simple C program that just did the readdir and I finally managed to reproduce the bug, indeed, readdir was not listing the files present in the directory.
But it was not helping me much... So I decided to do some network captures, everything looked OK. I did the same network captures with the old RHEL 4 kernel, and it looked exactly the same.
After hours at steering at my screen with the 2 captures side by side, I finally spotted a subtle difference. The fileids (64 bits) was padded with zeros on the first 32 bits with the old kernel (ex: 0x0000000000000000A489097654456F97), and it was not the case with the new kernel (ex: 0xFC871902B9086456A489097654456F97).
Basically, the old RH 7 was violating the NFS RFC by only handling 32 bits fileids (by the way, the RFC predates the RH 7 by 5 years).
As it was impossible to upgrade the old RH 7, I finally backported the "bug" in my "RHEL 5 kernel on an RHEL 4" by adding a small two lines patch in the RHEL 5 kernel that ensured 32 bits fileids.
It was a system with a lot of cruft accumulated overtime, and a lot of domain specific applications that needed to be ported to a newer environment (new OSes, migrating from QT3 to QT4/QT5 and other newer libraries millions of line of codes). Actually they were planning to move to a newer base system when I finally left the company. We actually did it for other part of the system, and it was a several years process.
I've also another horror story about NFS and RHEL 4 (at nearly the same time as the first one).
I was tasked to update a small internal infrastructure used by another project (replacing an old active directory, an old exchange server, adding a few other service, renewing the hardware, reworking the backups, etc).
For the most part, I managed to completely replace the old stuff by newer things (postfix+dovecote, openldap, samba4, CentOS7, bind, bacula, everything hosted in VMs).
There were a few things that were too hard to migrate without blocking the users too much.
So, to at least migrate away from a +10 years server, I made the choice to just transform it as a VM.
This server was hosting various things (300 subversion repositories, a custom tracker using php 4, a viewvc crazily deployed (apache listened on 10 different ports with some static pages to point to the different ports), probably some other stuff I didn't know, and, you guessed it, NFS.
There was a lot of stuff relying on this particular NFS server (lot of scripts with the server name and paths hardcoded).
Also, the company had decided that it was completely unacceptable to risk losing one line of code on any developer desktop. So they put their homes on this NFS server (And if you are wondering, compiling code on an old NFS server, over a 100MB/s network, with 10 other devs doing the same thing, it is just miserable...). The amount of data (~1TB) was quite large for a server that old. And yeah, it was also acting as a NIS server.
So I started preparing the migration:
* boot on a livecd
* assign a temporary IP to the new server
* creating the partitions by hand
* rsyncing the system from the live, old system
* and few other things like installing grub
(basically a Gentoo installation minus the compilation and packages choosing parts)
Then I checked the services, and everything was working correctly.
So I scheduled a final downtime in the late afternoon for the final synchronization and at the scheduled hour I shutdown all the services except ssh and did the final rsync. I finally shutdowned the old server and switch the new server to the old IP from its temporary one.
A final check showed me that most services seemed to basically work.
I arrived the next morning, and everyone was panicking, the NFS server was behaving badly. I logon on the server and saw that the partition used by NFS was mounted read only. I rebooted the server, and again the partition became read only.
I urgently shutdowned the new VM and restarted the old server.
Then I investigate the issue and I finally found that there was a likely candidate : https://access.redhat.com/security/cve/CVE-2006-3468
The good thing was that given it was a CVE, I managed to find an exploit and it did reproduced the issue on my system.
The issue was triggered because some desktops were not shutdown during the migration, they saw the old server disappear and a new one reappear under the same IP. Consequently they were sending old file handles from the old server to the new server. As described in the CVE the NFS server didn't handled bogus file handles correctly, remounting the partition Read Only.
I updated the kernel, replaned another maintenance window and everything went fine.
A final note: the kernel version installed was 2.6.9-41, the one that fixed the issue was 2.6.9-42, if this machine was updated ONCE in its life, the migration would have gone smoothly...
I left this company a few months later, with somewhat of a deep hatred for NFS ^^.
We found this out with an oscilloscope.
But, yes, still my favorite one ^^.