The problem was the warehouse owner wanted partial kitting. So if the warehouse only had (in our example) A and B, the code would send the picker to put A and B into a box, then direct him to drop it off in a special partial kit area. When C was back in stock, the system would have the workers fill out the partial kits and ship them. This way if a kit required a dozen items and you were just waiting for one to arrive, you could get most of the work done beforehand.
The problem was now A and B are in boxes and not in "inventory". So when someone orders a kit that contains A, B, and D the A and B bins are empty (as all items A and B are already part of a kit and thus not available) and the code would direct him to put D in a box and put it in the partial kit area. Eventually the D bin is empty, so when an order comes for a kit that requires D and E, we get another flood of partial kits, all going to the same location (which was just a square painted on the warehouse floor).
Anyway, long story short, if the right few items were out of stock and the right orders came in the right sequence, nearly the entire inventory of the warehouse ended up in a giant pile of boxes that was too large for the workers to sort through even when the needed items arrived.
Everything was humming along just peachy for weeks and then BAM! Red faces all around. It took days for them to put all the inventory back into the proper bins and fix all the data, and that probably cost into seven figures, all told.
In my defense, I wasn't the last one to touch that module.
If there is a fault the safest thing for them to do is bleed out energy slowly unfortunately in this case sounds like this crushed the 'obstruction' in the process.
My Industrial plant code screwup story was not caused by me but was pretty impressive, what was supposed to be a "simple firewall change" knocked out communication between two interlinked parts of our plant which caused the line to stop and a big delay with a few million dollars lost. I believe the root cause was someone fat fingered the addition of a new firewall rule and we ended up dropping every incoming packet.
Freezing = you hope that whatever the problem is, maintaining power doesn't make it worse.
In general, make the code easy to interpret, especially when debugging - which means organizing your code and using simple abstractions. State machines can be useful because they're easy to interpret if you comment your states, leading to high-level understandings like "the machine is waiting for an item" rather than "it will do something when I0.5 goes high".
He sent the machine an update, rebooted it, yet the staging machine was acting unchanged from the previous version. Moments later a supervisor ran into the room, yelled to put his hands up from the keyboard, reached over him, ran some commands, and disappeared into another room. A few minutes later he announced that my friend logged into the main assembly line machines and rebooted them with code that used a sensor that didn't exist yet, which stopped the entire chain for ~20 minutes. The company suffered $250k production losses during that time.
One day I was told that my software had a bug. The tool wasn't being retracted (pulled out of the material being cut) before being rapidly moved to a new location. As a result, the cutting tool was being broken off.
I asked [I think it was our application engineer] if we sold the replacement tools to the customer and I was told "yes". Then I asked him: "Then isn't breaking off tools kind of a feature"?
"Just fix it Chris. Just fix it."
Thank God, the patient was a pig. We hadn't made it into clinical use yet.
My first worry was that your measurement would be wrong, not that the power wouldn't be killed! Any redundancy on that side? Or was it not necessary?
If we measured wrong, we could either be high or low. If we measured high (that is, the reading is higher than reality), we would either turn down the power until it read right, or else kill power completely. If we read low, though, IIRC we would limit how high we'd turn up the gain to try to get the power we were asking for.
There was also a feedback loop based on temperature. If the power was double what we asked for, the temperature would quickly climb, and we'd reduce power. It would have worked, even with inaccurate power readings, though not as smoothly as it should with accurate power readings. But when we got 20 times the power we asked for (due to the power control failure), it was too much too fast.
Huh, I'd expect sensitive systems like these to have some sort of hardware redundancy/voting system.
Unfortunately, I cannot share much details except that I wrote code that was meant to manage the amount of risk that a certain really big financial institution was supposed to take. My code may or may not have shipped after I left that institution. If it did ship, maybe it did not do what it was supposed to do. If it did not ship, maybe it failed to replace the broken system that it was supposed to replace. Either way, months after I left, the head of the institution acknowledged on TV that they were taking on more risk that they intended to.
Well, the engineer writing the test code knew these devices were odd, and that sometimes they'd just fail. So, s/he put in an if block to the effect that, "if this fails once, run the test 30 times and, if it passes 25/30 times, call it a pass." So, every now and again, the entire automated testing line comes to a halt and sits there for 31x the amount of time it should take, and it's not a short test (maybe sat there 30 minutes each iteration).
The inner loop was something like:
while message:
converted_event = new Event()
for event in message.events():
converted_event.set_fields(event)
write_to_hdfs(converted_event)
Can you spot the bug? Led to a month of corrupted data before I noticed..The `set_fields` method does not clear all fields, so every event had more and more junk data than the one before it. All because i thought i would be clever and get some performance gains by initializing `converted_event` outside the inner loop.
Maybe there should be a separate database of historic students who used to go to that school, and currently enrolled.
It's not just "isDeceased", but "goesToThisSchool". Nobody want to get some notification from a school about something, when their kid doesn't go there any more for any reason.
Or, rather than duplicating data, just use a view with appropriate criteria to limit to currently active, living students for most queries. But a developer that's called into build a query generally isn't going to get a lot of mileage out of suggesting rearchitecting the database, in either of those ways.
It's like you were there :)
It was a third party product so changing the structure was out of the question. We had some views but they pulled in the entire database and ended up with so much duplicate and irrelevant information they were unusable.
I tried creating a clean set of views like "v_currentStudents" that could then be joined on for information relevant to the current report. I even built a small test suit for them, but getting the support devs (who I was covering for when this happened) to change their cowboy ways was too hard. Management didn't like they idea either, cut into the billable hours.
One hour before the end of my internship, I was ready to leave, my work done, ready to be used for the next person taking over the project. I want cleanup my files and documentation so it is all tidy and I try to commit my work. Of course SVN cannot commit because the repo and my work have nothing left in common. So I type (on a Linux system): svn delete to cleanup the repo so that I can push my files... I lost months of work and I was not able to recover my lost files from the file system... I had to leave for my country of origin since this internship was part of an exchange program. I felt so bad about it, it still haunts me.
Someone else added it to a group policy for all corporate servers, including all our Exchange servers, where the active database transaction logs are named .log.
At some point, I found out that inserting 'use strict' at the beginning of each Node.js module, enabled the experimental ES6 (harmony) features in Node v4. Needless to say, I was super excited and immediately started using classes and other ES6 goodies everywhere, even refactoring already existing modules.
Shorty after that, we noticed that our servers were leaking memory and started crashing almost every day. At the time, I had no idea what the problem was - and believe me I tried everything to find a solution - until a couple of months later we switched to Node v6, and everything miraculously returned back to normal. In the meantime though, during those 2 dreadful months between v4 and v6, we had to setup cron to restart our servers every single day at 04:00...
Never use experimental features.
I (thought I) set an actuator running an auger (screw) that offloaded the feed from the mixer into a leg (12 meter tall vertical screw) to "run always." That should be safe right? The auger runs all the time, carrying away anything that dumps into it. What I had actually done was set a hay belt to "run always" it was stuffing the mixer with more and more hay until it was a solid mass inside the box.
Everything seemed fine when we started that next batch of feed... then the mixer started. The lights dimmed and there was this shriek of metal and a bang from the mill floor. We shut down and went out to see this huge mixer hanging off a drive chain at 45 degree angle from the floor. Bolt heads the size of manhole covers had sheared off and were lying nearby. Fortunately no one had been standing nearby. I don't know if my memory matches reality, but I recall a light from one of skylights shining down on it in the grain dust like a spotlight.
I was pretty sure this was going to be my last day on the job.
I walked over to stand next to the Mill Manager, a salty fellow named Marvin with three fingers on his right hand. Marvin looked up the chain and back down to the bolts on the floor and said "Yep, it'll do that."
Two workers lowered it down and welded the bolt heads back in place like they did it every day.
I was with the company for five years. I don't recall every having a support call from that mill after we finished the installation.
tons and tonnes/'metric tons' are roughly interchangeable, like yards and meters. :)
I was so mortified, I guess it stuck well enough that that's the worst thing off the top of my head.
But it seems like I'm an underachiever based on this thread.
Is that a jab at Hughes Aircraft? :) Looks like "Firefinder" is some kind of radar system developed by them.
I'm sorry for how you must feel.
It isn't a pleasant thing to live with.
Imagine someone who works as a knife grinder. If he do his job right the knifes will be much more dangerous, they may cause accidents or even some will be used intentionally as a deadly weapon.
Then considering these possibilities: an ethical knife grinder should do a shitty job, should quit, or should live in self-reproach?
It may be more complicated if you are a gunsmith. But those guns are used by your customers - so in what extent your ethical evaluation will depend on their actions in this case?
For example if your guns are used in an arming race, and eventually they help avoiding war then you are a saint? If your guns are used in a victorious war then you are a hero? And if they are used for killing innocents then you are an evil person? Or you should be judged by the average probabilities of the global gun usage? Or what?
Poorly sharpened knives are far more likely to cause accidental injuries, and serious ones, than well-sharpened knives. At least in kitchen use, though I'd expect the “a dull knife is more likely to fail to cut what you meant, slideshow off, and strike something else” effect would apply in most uses of knives.
The Ohio Bureau of Worker's Compensation[2] recommends the same.
The Bureau of Industrial and Labor Statistics[3] cites dull knives as a common cause for injury, and recommends keeping knives 'sharp and in good trim' to prevent accidents.
In short, "a sharp knife is a safe knife" isn't hokum. When you're pushing a knife into something, you're storing and releasing kinetic energy. A sharper knife requires less kinetic energy to begin cutting the object, which is ostensibly dangerous, but not as dangerous as a failure to cut, which releases all that kinetic energy in uncontrollable fashion.
Past that, in the event that you do get cut by a knife, a sharper knife makes a cleaner cut, which means easier healing, easier care, and (if dire enough) easier reattachment. Oh, and less scarring to boot.
[1] - https://www.osha.gov/SLTC/youth/restaurant/knives_foodprep.h...
[2] - https://www.bwc.ohio.gov/downloads/blankpdf/SafetyTalk-Preve...
[3] - https://books.google.com/books?id=W0M4AQAAMAAJ&pg=PA190&lpg=...
Suppose you work on the grep program GNU Coreutils. Harmless, right?
Some regime could use that to grep out a list of innocent people to put on a hit list.
If you had no idea what the purpose was, that means the program had conceivable purposes of various kinds, not related to killing, just like grep.
If you develop something that is pretty much only for killing, obviously so, then you know, right? Or else are capable of incredible denial.
One of my first implementations at a bank many years ago... bunch of 'C' levels are in the main branch for my first big launch demo...
Tape a few keys...
**ERR ** HELL FROZE OVER!
LPT: never use this in an else case. default: /* unreachable case */
assert(0 && "hell musta just froze over");
that type of thing? Impossible case throwing funny error message? "rm -rf /var/scratchdir /"
Yeah the space was a typo. Wasn't running as root but was able to make a pretty big mess regardless. rm -rf $MISPELED_ROOT_DIR/lib/
Oops, the script didn't have "set -u", and I happened to run that as root. So, /lib directory gone.I managed to recover that machine by copying libs from another one running the same distro.
Came to work the next day, nightly build still had not finished on slave servers, had errors about non-existent home folder when tried to log in.
What made it worst was that every server mounted a NFS share that contained fingerprints and binaries of different versions of software modules built on different platforms.
Killed all slaves, restored the NFS share from week old backups on tapes, tens of developers could not create new versions of software and send previous versions/patches to customers for a while.
I was working on Cloud Management software for a Private Cloud at a major tech company in SV. We had software which would reserve Prod IP space for hypervisors, e.g. this hardware SKU can support up to 5 VMs, therefore it needs to reserve 5 IP addresses in the corresponding subnet.
Turned out the API call to reserve the IP space from the IP Manager wasn't asynchronous and because the manager tried to get consecutive space, the runtime increased exponentially with the requested # of IP addresses.
In preparation for Holiday traffic, we were onboarding a new SKU of Hardware. This hardware supported more tenants and so instead of requesting 7 IP addresses per HV, now we're asking for 15. This took the latency of a call to the IP Manager from 3-5 seconds to 5-10 minutes. To round off the perfect storm, the code was retrying requests which failed, without propagating the failure to the Cloud Admins using the software.
One day in October, I received a panicky call from our Capacity manager, customers are trying to spin-up VMs but are being told there's no IP space left. He knows we've onboarded all the racks, and he's done the math on the subnets (which are showing as fully reserved), and there still isn't IP space...WTF!!
Turned out the IP manager's VIP was cutting off requests after a few minutes, (never a possibility when reserving only 7 spaces) but the reservation process wasn't stopping, the IP was being reserved, marked as in-use, but never actually making it to the networking service to be used by VMs.
Solution: At 2am on a Friday night I ran a script to manually mark tens of thousands of production IP records as not-in-use in the IP manager, purely based on grepping through logs from my service, and nslookups. But don't worry, we pinged each IP just to be safe :)
I had some code which handled a temporary loss of carrier. It would poll for the carrier to come back for a few moments, otherwise indicate to layers higher up that carrier is lost, so the user can be logged out.
Problem is, in that piece of code, I forgot to pop something off the stack that I pushed onto the stack. I had a user who was a bit of a cracker. I got a note from the guy, "I got into your operating system by dialing touch tones while connected".
Dialing a touch tone interrupted the carrier sense in the modem, triggering that code with the bad stack handling that would crash the BBS program, leaving the I/O hooks still connected to the modem driver, giving the caller full access to the system.
This didn't reproduce during the usual case when the carrier was lost permanently, only when it recovered.
- function initMultiowned(address[] _owners, uint _required) {
+ function initMultiowned(address[] _owners, uint _required) internal {
This bug led directly to over $30 million dollars being stolen yesterday. Not my code, but impressive nonetheless.Hackers have stolen $32 million in Ethereum in the second heist this week http://www.businessinsider.com/report-hackers-stole-32-milli...
Fix initialisation bug. https://github.com/paritytech/parity/commit/e06a1e8dd9cfd8bf...
He told me to: -write out a script that waits 1 second and then runs the application -run the script in a separate process -kill the application
I bet that code is still there. It works, but damn is that cringy.
Edit: OK, this seems like a very common issue!
I forgot to check my inputs. Ran in production for a backup system.
Essentially production was not acceptable for a little bit....
Nobody got fired, because we had a QA team, and their testing procedure didn't test for something like that.