I shipped a word processor that formatted the hard drive every 1024 saves
twitter.com
twitter.com
I have written variants of "GIANT BUG... causing /usr to be deleted... so sorry...." into commit messages over the years. Such a classic part of history.
[1] Commit thread: https://github.com/MrMEEE/bumblebee-Old-and-abbandoned/commi...
[2] Issue thread (not as good): https://github.com/MrMEEE/bumblebee-Old-and-abbandoned/issue...
EDIT: Github is timing out trying to load it. May want to use archive.org: https://web.archive.org/web/20130613012555/https://github.co...
At least it's hilarious now, for some people it must've been horrifying.
The grandparent bug is as much a failing of the grandparent's script as the OS, which exposes itself as one global mutable namespace.
At any rate, lots of people at Apple back then brand new to Unix, NeXT was still integrating, everything was still coming together. And they made one of the absolute most classic Unix newbie whoops moments: they wanted to clean up old versions of iTunes, so the used rm -rf... without quoting the path. It had this IIRC:
rm -rf $2Applications/iTunes.app 2
with $2 as the path. But of course classic Mac users were used to having spaces in drive names and folders and so on. If you only had the startup drive no problem. But if you'd partitioned or had an external drive and it had a space in the name, ie "Disk 1", then that'd become rm -rf Disk 1/Applications/iTunes.app 2 and you were off to the races. There were some fun discussion threads about it, although unfortunately the only Apple Discussions bookmarks I have saved from back then all seem to be dead. Not sure if they're still archived somewhere and the links just no longer redirect or it was cleared out sometime.They got that pulled real quick too, but I always secretly wondered if the genesis of Time Machine was somewhere around there... backups were pretty rough at that stage of the game. Well, everything about Mac OS X was pretty rough, though exciting too.
Edit 2: Did find an old /. discussion about it that still works. Bit of a blast from the past reading through some of those, both in what has changed and what hasn't:
https://apple.slashdot.org/story/01/11/04/0412209/itunes-20-...
It didn't succeed because I wasn't 'root'. Maybe it was a local misconfiguration in this company. Not sure. But I've never looked at Go the same way again.
1. Create the directory "C:\APPNAME-tmp"
2. Change the current directory to "C:\APPNAME-tmp"
3. Recursively delete everything under the current directory.
This worked fine for thousands of users until I came along -- running Windows 2000, with the "C:\" directory read-only for unprivileged users, running as an unprivileged user. At that point it became:
1. Try, and fail, to create the directory "C:\APPNAME-tmp"
2. Try, and fail, to change the current directory to "C:\APPNAME-tmp"
3. Recursively delete everything under the current directory, which is my home directory.
The authors were very apologetic when I pointed out the bug.
https://github.com/valvesoftware/steam-for-linux/issues/3671
Featuring probably the most unflappable human in the history of time as a bug reporter. One day I hope to have the fortitude this man does.
One of my professors used to say that he should be able to destroy your laptop, buy you an equivalent new one, and you should be up and running again within a few hours. Hard drives fail all the time and computers get lost/damaged/stolen. Losing your home directory on a computer should be expected, and definitely not the end of the world.
I've been using actual raid5 (via adaptec raid card) for years and very recently had one of my trusty 5TB HGST drives fail (after 3+ yrs of uptime).
Fortunately the rebuild worked, but there are so many horror stories of raid5 rebuilds NOT working it has me contemplating going back to simple mirroring.
As an IT Exec (retired)... this was a lesson learned long ago. For my own/home systems I use rclone [i] (in conj w raid5) for the most critical files.
If you are going to use hardware raid, please do make sure you have a spare raid controller, same firmware, same model. If not, you will be SOL when your controller dies.
RAID5 these days (with our very large disks) is basically asking for trouble - the odds of a second disk failing during the reconstruction are very high. But I guess you already know that!
Now I use a more sophisticated RAID10 setup and really appreciate the way failed and replaced drives automatically rebuild themselves without my dumb interaction.
I would also mention it's probably a good idea to use something like ZFS rather than a hardware RAID card. In general, free software RAID is better audited and you don't need specific hardware for the recovery process. But also, ZFS has proper checksumming, so you won't have to worry about all sorts of silent corruption (which RAID blissfully ignores -- most implementations will just recompute the parity in case of a mismatch, which is the wrong thing to do (1-1/n) of the time).
Was hoping to utilize ZFS in 20.04... but the whole SNAP issue has me sticking with my current version.
Perhaps it's time to find a dummies guide to setting up ZFS on 18.04 LTS.
You can say 'robust' but there are limits of what I think one can reasonably expect from a user. That they had backups at all is not necessarily the norm.
If you want to guard against rogue software (or clumsy fingers in the terminal), you'll probably need to have remote backups. It sounds like copies on the cloud saved this user, and it's not unrealistic to suggest users backup to the cloud (I think many already do with OneDrive/iCloud/Dropbox). If you're a Linux user who likes to tinker, you can set up a Raspberry Pi with a hard drive attached and use restic over SFTP (or any of the other numerous choices).
That is fair. I was commenting from the perspective of what is rather than what should be. Alongside making software as safe as possible, we should also be encouraging and expecting people to do this.
Not sure if rsync can do local encryption these days? So I guess 'for the paranoid' (as Tarsnap's tagline is), Tarsnap with write-only keys might be better.
By old files, that can mean any snapshot schedule you choose to set up (daily, weekly, monthly, quarterly ...).
No one can delete the files afterwards, not even root, without removing the attribute first
It worked well with Mercurial, which is also append only. So I could commit as usually to my Mercurial repository, but it could not be deleted.
Ironically, I stopped doing that, because it messed my backups up. When running the backups as root, some tools would add the flag to the backuped files, but then later backups as non-root could not replace the file with new versions.
What do you use for that?
[1] https://syncthing.net/ [2] https://restic.net/ [3] https://nixos.org/
Secondly, treat all machines as cattle. No customization, unless done programmatically or repeated easily, and absolutely essential.
In practice, I have a tiny Dir in a Resilio share from which I bootstrap. It contains some cfg files/dies, a bashrc, passwords and share keys, some written notes for customizations that are not possible or unreliably so to automate. In my notes you will find for instance a package list for fresh installs, notes on which Firefox extensions to install, how to configure certain software that have only a GUI to reliably do so (and I test such instructions before I consider the tool ready for use), a zip of my thunderbird profile, and so on.
I started this way of thinking when my dad accidentally formatted my drive in 2000, and it has been bullet proof since. When I use new software, I do not consider it usable before I made a note or cfg backup in my 'bootstrap' share. I do not rely on my memory and I do not rely on a particular machine for anything, and it costs barely any effort other than being disciplined in never customizing a machine without recording and testing how to repeat that someday. If that is too much work? Then I will not use that software, apparently its not worth it.
rm -rf "$STEAMROOT/"*
Yes. I have bookmarked this issue as a reminder to always read every single shell script before execution. set -Eeuo pipefailAs for the other conclusion, I have a story of my own. I once shut down all of my company's asynchronous task processing for about 20 minutes, completely by mistake. Our web servers kept serving, and nobody external would have noticed a difference, but, for 20 minutes, our scrapers stopped scraping, our NLP classifiers stopped classifying, and a bunch of stuff that made us money wasn't happening. Had I not noticed it and fixed it ASAP, we probably would have lost money. Instead, I immediately announced what I had done, got help, and we fixed it.
The real moral of the story is that an honest mistake is nothing to be ashamed of. I make mistakes all the time. Some of them make it into production. Big deal. Everybody does.
This didn't happen on Windows, just in Wine. It also went away in the 1.10 patch IIRC, so whatever mysterious behavior resulted in that got fixed.
edit: Found it! [2] It was even more recent than I remembered.
[1] - https://forums.gentoo.org/viewtopic-p-2723361.html?sid=3c70a...
The records were orders. Millions of orders.
mkfs.vfat /dev/sda1
That's not that much different from a plausible error on Linux, with a USB drive and a hard drive.(At least, the man page for my mkfs.vfat doesn't include any options suggesting it protects against formatting already-formatted or mounted filesystems.)
Killed a windows install (which had the first partition on the disk) only a year ago with that one.
"Yes by all means, use a nonexistent file as input to overwrite an existing file That doesn't even have a .tar extension. That's exactly what I wanted"
And it shipped actual, physical CDs across the world for free. It was 2004!
Then it shipped a free LTS that may not have been as stable as RHEL, but it was definitely fresher and with guaranteed updates (for main) for 5 years (3 on desktops early on). It was 2006.
“I shipped a hard disk formatting utility that doubled as a word processor.”
There's a reporting & data extraction language called Focus (nowadays WebFocus). I was working in a very old version, and had a few nested GOTO statements to form a loop where the final output was a customized email sent for each row of a query returned. I had to make a simple update, but accidentally deleted a line first. I reinserted it... one line off from where it had been. My careful GOTO structure went from a loop, to recursive. Instead of sending an email to each row returned, I sent an email to the first person. Then I sent an email to the second person and the first. Then to the third, and the second, and the first.
The list was about 1,000 recipients... but luckily it usually finished very quick, and I would monitor as it went, and noticed the long run time after about 100 iterations and killed it. I then checked the logs to see why it ran so long, and traced back my error and fixed it... and sent an apology email to the 100 recipients I accidentally spammed.
Pro-tip: turn-on the alerts after deploying the site.
I've been writing code for about as long as him, but because I started with Asm, which makes one become really careful with buffer sizes, doing something like making one thing larger would not be done without carefully looking "down the line" to see if any further changes were required; and they almost certainly were. That's not to say I haven't corrupted files before, but fortunately nothing quite as catastrophic as wiping the disk.
Might be some memory corruption caused by a wrong pointer in any other part besides the saving?
I don't think it's realistic to say that all incompetent engineers think they're competent.
I'm not aware of any way to make a modern x86 processor self-destruct. Apparently you can make an FPGA fry itself but even then not with good heat conduction away from it.
Not entirely sure (and I have no intention of given Bethesda enough money to find out)
In '87 my first 256 computer had 640k of memory with a hd that was less than 1mb IIRC. It cost nearly 5 thousand dollars including the $700 dot matrix printer.
Let the good times roll!
Even though it was a Tandy, I appear to be mistaken about the HD capacity.
Excuse me?! So if I'm writing a prng, should I write the numerical constants like
int modulus = 1 + (1+1+1)*(0 + (1+1+1)*(-1+... )...);const int MAGIC_THING = 0xdeadbeef
& then say MAGIC_THING throughout the code, rather than using 0xdeadbeef in all your expressions
Protip: If a smart person says something that seems obviously dumb to you, it's worth trying to find an interpretation that isn't dumb. Doubly so when, as here, it's from a piece that others are clearly finding smart and useful.
Also, you are referring to an orthogonal problem—you would presumably want to put magic constants into a static constant to reduce the chances of disagreement across references. Hell even if there’s a single usage I would expect magic constants to be clearly demarcated and documented.
Edit: didn’t see sibling constant, didn’t mean to dupe the reply itself. This certainly refers to in-line use of numeric literals vs those used in constant definitions.
Of course, if you are repurposing the code to do something new, you need to be extra careful not to use the constant just because "it's the right number". And sure, it sometimes gets tricky, but using literals will not make it any less so.
If you want to go full spectrum the other way.
There are other numbers like that (12, 24, 60 come to mind; 365 has a similar familiarity but has lots of gotchas; I am sure many developers will immediately understand 1024 too).
But that's also mostly due to how they will be used. Contrast and compare:
— a = b * 24
— a = b * HOURS_PER_DAY
— uptime_in_hours = 24 * uptime_in_days
I like the last one best, but if given a choice between the first or the second, I'd prefer to be reading the second.
FWIW, I am not advocating for use of units in variable names at all times — but if you are converting between units, put them in either your variable or constant names.
Basically, I interpret the rule of no-literals as a reminder to think of the readibility of any statement involving them.
Now that is an excellent rule, if slightly too specific.
What he is suggesting is that you generally not include literals in your code, and instead, use a constant/variable that can be traced back to a single place in order to make changes more visible and easier to deal with. That he makes exceptions for 1, 0, and -1 is explained by conventions in various languages where those values in context end up as generic and meaningful as any other language keyword and thus would not benefit from being referenced from a constant. That doesn't mean you wouldn't ever end up with constants that are 0, 1 or -1 though, just that you wouldn't assume every random place you might use those values (such as sorting algorithms) justifies a constant placeholder.
For further review on the topic, may as well start here: https://en.wikipedia.org/wiki/Magic_number_%28programming%29
I'll admit, my example was overwrought with comedic intent, but I stand firm on this issue. Use a magic number twice? Yeah, go ahead and put that numeric literal somewhere convenient (but you still have a numeric literal) to the consumers of it and easy to find by your readers (not a top-level magic_numbers.h)
I'm a mathematician. The fear of numbers in code is an affront to the domain that I work in. YMMV. I put a lot of work into documenting my code, but for the love of pete, 2 is not a magic number: my example was no more overwrought than OP's rule.
In some cases, I'd say that 1 can be a magic number. For example, it would better to write
fprintf(STDOUT, output_string)
than fprintf(1, output_string)
However, 0 and ±1 have a special role in picking out the first/next/previous elements of a sequence (and related things), which is so common that it'd be silly to insist on defining POSSIBLE_NEXT_INDEX and POSSIBLE_PREDECESSOR. That said, people do seem to love Python's itertools, so....maybe that's where we're headed. I had somebody aggressively complain about the readability of the all-pairs for-loop, which I thought was basically standard. for(auto i=0; i<N-1; i++)
for(auto j=i+1; j<N; j++)
process(X[i], X[j])
The flip side is that I don't think you always need to do this. For example, this is silly: double triangle::area() const {
const double NORMALIZATION = 1/2.0;
return this->base * this->height * NORMALIZATION;
}
The point is just to make it clear why that particular number is being used and where it came from. for i < j <= N
as this is super clear.However, Python has itertools.combinations(X,2) and Julia has IterTools.subsets(X,2) if that's what you want.
# see __magic_numbers.py
itertools.combinations(X, number_of_things_in_a_pair)Some things only/obviously make sense over pairs of items. For those, go with a literal 2 for x,y in itertools.combinations(obj, 2): if is_overlapping(x,y): raise OverlapError(x,y)
However, sometimes the subset’s size just happens to be two (but you might change it), maybe a constant would be good. Ditto if you’re doing some math where there are “real” 2s that are part of a formula and incidental ones that are due to the subset size. For example:
subset_size = 2
for subset in itertools.combinations(X, subset_size):
mse.append(sum((subset - target)**2)/subset_size)
is a bit more flexible and more clear than using all 2s, IMO, and at very little cost (YAGNI blah blah, but I think that’s an argument for not making a whole configuration system that lets you set the size at runtime).Only in my choice of variable name. OP drew a line at 0, 1 and -1. What I did there was highlight that the implications of that rule are absurd. See how your sum of squares also contains a 2? VERBOTEN!!!! And don't you dare re-use "subset_size" ;)
This is somewhat akin to Dijkstra's opinion on goto. Which is actually great advice when you're doing apps in javascript, but doesn't get you very far if you're writing or generating assembly. When such advice is promoted to a taboo, I side with Churchill: this is the type of arrant pedantry up with which I will not put. Or Emerson: foolish consistency is the hobgoblin of little minds.
Yes, too many numeric literals can make code hard to read. But math is hard to read[1] because you need to really think about it -- you can't avoid the complexity; you can only rearrange it. Where you put it is a matter of taste, and absolutes have absurd consequences.
As for the original notion of indexing a triangular array, I have a greater concern:
for i in range(n):
for j in range(i):
foo(array[i][j])
this is self-documenting in the sense that I immediately know the relationship between i and j -- without reading documentation, I can't recall if itertools.combinations will give me the upper or lower triangle. In this case, I'd avoid the 2 for entirely different reasons :)[1] and I don't mean "let's go shopping!" -- I mean that reading a math paper, even for experts, can take days per page.
Yes, of course if you have a formula that involves dividing by 2 or something you just include the number directly.
But anything that's a magic number, even if only used once, is better to have a named constant defined. The name of the constant becomes the documentation itself.
The other reason is that if someone needs to come in and maintain the code and change a magic number for whatever reason... if it's a defined constant, they know it should only need to be changed in that one place, assuming all magic numbers are defined in a sufficiently wide scope. If it's just a number, they have to search the whole codebose for all instances of that number, and investigate each and every one to see if it needs to be replaced as well or not.
Or replace it with >> 1
I have some code where I needed to calculate what the 1st percentile of a list of numbers was. Since it doesn't make a lot of statistical sense to calculate a 1st percentile of fewer than 100 numbers, I inserted a condition like
if len(numbers) < 100:
# skip the calculation
and included a comment stating that 100 is explicitly not a magic number here. By that, I meant that you'll never want to change this 100, so, why bother obfuscating it behind a name?It's perfectly readable, if you understand why the calculation is skipped in the first place. If you don't understand why the calculation is skipped, giving it a name like TOO_FEW_NUMBERS_TO_CALCULATE_1ST_PERCENTILE isn't really going to give you much insight into why the calculation is skipped, anyway.
To me "if len(numbers) < MINIMUM_NUMBERS_FOR_1ST_PERCENTILE" still reads better and requires me not to need to refer to comments, which I fall back on when I do not immediately get the code.
So maybe the risk is not lowered, but readability is still improved.
(Oh, and I wonder what do you return for a list of all equal numbers when there are more than a 100? If it's that number, why would suddenly going from 100 to 99 change that? ;)
/* From Park and Miller (1988, pg 1195), using their notation */
const int A = 16807;
const int M = 2147483647;
const int q = 127773;
const int r = 2836;
...
instead of "inlining" them into the code like
test = 16807 * (seed % 127773) - 2836 * (seed/127773)The "ban" on literals obviously exempts their definition. No one, outside a number theory textbook, is defining A as successor(successor(...(successor(1)))
That's pretty cool. Did you at least have to look it up?
I recommend the citation-in-the comments idea though. We do it for lab stuff and it's very helpful for everything from debugging to writing up results.