62 karma · joined July 6, 2012
If they showed individuals with already high acceptance rates received lower acceptance rates on the same code when submitted from a gender identifiable account then I'd find it to be more compelling.
Gain more experience in IT? I mean i've been at it professionally for 18 years soooo I guess it depends on what you consider "a lot of experience"
I get what you are saying. "A disgruntled amazon employee could just delete the world's data" right? Except for thats not the case. Its all stored in triplicate(at a minimum) across a vast number of independently available data centers. The durability guarantees are hard to fathom . Honestly I bet they'd have a hard time deleting data permanently even if they wanted to.
But once against I'm not advocating that you only store your data in one place. I'm just saying dollars to donuts you stored at least one copy of your data on s3, catastrophe has struck and only one copy of your data is left where do you think it is?
I know where I am placing my bets.
I didn't call you an "alarmist" I said worrying about whether amazon will stay in business, "get hacked", or if S3 is going to disappear without warning is "alarmism" They were responsible for 39% of all commercial internet transactions last year and store 2,000,000,000,000 objects. If they go under we are are all in a heap of trouble. A NAS device in some remote part of the netherlands isn't going to save us.
>I'm a piece of shit because you disagree with me, but it just so happens that you've been all over this thread with a continuous stream of purposeful mis-understandings and/or downright trolling.
This is just not true. I called you that because you tried to do just that to me. A quick re-read of all my comments and I'm sure you'll see I never made any personal reference to you. Now if you consider me telling you that you are incorrect an attack then I'm sorry this is something you'll need to get used to. In this case because you are wrong.
I've offered you ample opportunity to proffer a statistical argument, but all you can give me is "if I didn't build its not safe" Sorry friend I just don't trust you. I've been doing this a long time and I don't trust myself.
I never told someone to do something "cause I said so" but then again you never address the meat of my position. "Cloud providers have much lower failure rates than you historically I trust them more than I do you when building my understanding of the risks". Its trivial to undermine my position. Show you have lower failure rates than a prime time storage provider(bonus points if its amazon).
Lets just a get a few non sequitars that you can't seem to get your head wrapped around. I never said don't have a back up. I said the back up in the cloud is the most durable one.
I never said amazon couldn't fail. I just said the chances they fail vs. you are orders of magnitude lower.
I never said you were bad at your job. I said you seem to have a shaky grasp on statistics and possible a fair amount of NIH syndrome.
You are correct my anecdote doesn't prove my point, but it provides anecdotal evidence to corroborate my statistical position.
I didn't say "don't be open or transparent" I said pulling a persons personal details into an internet argument crosses a very clear line.
Get over yourself and remember not everyone that disagrees with you is trying to undermine your career. But you step over a line with you get personal and start pulling personal details into an internet conversation.
So getting back to risk management. So we about to make a QB selection for the big game. Now that we understand that the implementation of folks like jaque here are like the 12 year old pee wee standout vs. the cloud provider's seasoned NFL quarterback. Which one do you pick to safe guard your business.
Insured by who? Not governments. Governments collapse all the time. The only safe place to keep it is in your own custom bank at home that you built yourself.
This is from you and this is from the wikipedia article on NIH(not invented here) syndrome ...
>Not invented here (NIH) is the philosophical principle of not using third party solutions to a problem because of their external origins. False pride often drives an enterprise to use less-than-perfect invention in order to save face by ignoring, boycotting, or otherwise refusing to use or incorporate obviously superior solutions by others.
I am not saying don't have a copy, I am saying the if you have several copies the safe one is in the cloud.
I can make this even easier. Its almost certainly going to boil down to this question. Have you ever lost even a single byte of data? Cause if you have, you aren't even in the same ballpark.
The failure probabilities are relative. And everyone here is equivocating. As if the likelihood of failure is evenly dispersed. You indicate that storing customer data in the cloud is "playing fast and loose" but I'd argue the opposite. Not having a cloud backup is what is "fast and loose"
Imagine you are an independent bank. You need to move your customers deposits. Do you load sacks of cash into your corporate minivan or hire an armored car service? Well lets look closer. With the latter you are giving up control, right? You have no control over the quality measures an armored car service might take. Yet somehow not contracting them seems like foolishness. The reason is obvious. Its because you know in this case that transporting money isn't your expertise. Its not something you can focus adequate time and resources on perfecting. You also can't spread the risk of failure across a lot of customers, absorb that failure, and make your service better for the next go. The exact same logic applies to long term storage of data. If that isn't your only function its extremely hard to get it right.
>Do you control Amazon? No? Then you're not the captain
This is textbook NIH syndrome. Rather than looking at the relative probabilities you look at something from an ideological perspective. Your argument isn't "Its safer with me" its "I want to be in control of it" This attitude is generally harmful. Think about your money. Are you in control of the bank? no right, but you still keep your money there instead of in a shoe box under your bed. Why? Because the bank is better at protecting your money than you are. They have big heavy metal doors and men with guns to move it from place to place. Amazon is like a bank for your data.
But the fact remains that where ever you choose to put it "offline and offsite" the odds of it being lost are orders of magnitude greater then its persistent and redundant storage on a reputable cloud provider. Even if you put it on the most stable storage you can find and lock it in a underground safe. You can't guarantee its integrity a handful of years from now let alone centuries.
To put it in other words you are advocating storing your money in a mattress because you don't "trust the banks"
If Amazon went out of business there is no chance at this point that the loss would be total and catastrophic. Its essentially a bank at this point. Its shutdown would be orderly.
This is just alarmism. If you really wanted to demonstrate your point you would show me some data. The data would demonstrate that over the millions and millions of users on a number of cloud platforms that their rates of data loss are significantly higher than your "home spun" storage. Then you would take out the outliers and show an honest distribution on the most stable vendors(since I'd expect the majority of data loss over all vendors to have some amount of locality within a specific vendor). The result of this will be the rate of data loss on the most reliable cloud platforms. You can then compare this against your own success rates.
If you have ever(even once) lost even a single byte of data(for any reason), I do believe you will find yourself hopelessly outmatched.
I don't understand this logic. Amazon's S3 offers service level agreements with failure rates that at one point implied the statistical likelihood of losing an object to be once in "thousands of years". When dealing with any sort of stable storage this is simply something I cannot offer. I couldn't produce a set up locally with the resources I have for making guarantees on the decade level let alone millennia.
With that said I keep personal copies, but the authoritative copy is what's in the cloud, because its a hell of a lot more stable.
TL;DR I hear this argument all the time. The cloud isn't perfect but its a hell of a lot closer to anything I could achieve. "not invented here" syndrome won't save your data.
Look at something like java. Java has a core set of design elements. It was built by a person with some philosophical leanings("Everything is a class", "Checked/Unchecked exceptions being different things", "no first class functions") Then it has standard libraries(i/o, collections, threading/synchronization) each of these was built by a person with there own set of biases and understandings. The original Collections implementation has lots of mutable data structures then later Josh Bloch decided that he didn't like that anymore and stopped adopting "immutability" as core design philosophy. Immutability was not previously considered important when evaluating a java implementation. What you end up with is a mish mash of different opinions that only gets more different as you go.
Some 3rd party libraries like Guava didn't jibe with the Philosophical leanings of the language itself and looked for work arounds. They went as far as to create their own implementation of functions as a first class citizen in a language that was expressly designed to omit them. Some commonly used Android libraries do this as well.
My point here is that "language" can mean lots of things. It can refer to the language itself, its run time, the syntax+runtime+community. What is idiomatic and on, on, on. People and groups of people have the philosophies. I'd like to see the author point to something a little more critical about the differences between these two concepts. This premise is a little muddled.
The reason for this isn't clear to me. PRs are nothing but a `git merge` wrapped in a web ui that shows you a preview of the diff. Since those concepts are equivalent you have to then say "merging is pretty clearly outside the realm of continuous integration", but I'd say thats the concept that makes CI possible. By virtue of their equivalence PRs(if you choose to use them) can make CI possible too.
I want to be super clear about this using "Forking" and "Pull Requests" have zero limitations when it comes to CI/CD. In fact compared to working out of a shared repo there is only exactly one difference. Since your clone's master can diverge from authoritative master you have to periodically synchronize masters(thats that "sync your fork" link I put in the original post) I'll admit this is almost a justification for not using forks. Its annoying and tedious to explain to new git users.
>It's not continuous at all -- it places the onus of "merges" on people looking at stuff instead of trusting the tests
This sounds a bit off to me. Continuous integration is about getting good code into the delivery pipeline. Once that code is there you want to get it public as quick as possible. There is a relationship that describes the cost of bugs and bad design decisions as exponential given their proximity to getting into the customers hands. The tests will find regressions but won't stop a bad design or a design with a new bug, or a piece code without any tests at all. There are two things to note here catching things early is hugely cost effective and that code reviews are a vehicle for probing a completely different class of problems. The GH workflow is built around code reviews. This is because in the "Social Coding" model you want to make sure whatever rando is delivering code into your repo is respecting your style/development guide. Most organizations want this benefit as well. Do code reviews slow things down? ... yes. But I'd say thats a feature not a bug :) Are "pull requests" not "continuous". I don't understand exactly what you mean by that or what its value is. But it doesn't seem terribly useful by itself.
Branching is a separate concept and its certainly accounted for in the forking model. But we are talking about the process driven meta layer that github has stacked on top of git. Their original proposed workflow did not have a bunch of folks working out of the same shared repo instead(and this even extends to their enterprise model) every user that start working on a project would fork it. You'd make your changes and contribute them back to the authoritative repo as a "Pull Request". Branches model an evolutionary line of code, but a "Pull Request" models a change that you want to introduce typically into another repository. The implication here is that forks and branches are separate concepts that share some overlap. I often branch in my fork and contribute my change from my forked branch to be introduced into `master` in the authoritative repo without ever delivering the branch itself. Rather simply merging the changes into `master` as a discrete unit of work. Github's polite suggestion that you work this way is further evidenced by encouraging you to delete merged branches. I find this to be satisfactory but many people have personal hangups with deleting branches for reasons that are not entirely clear to me. This model allows me to sync my fork's version of "master" periodically with master in the repository of record(typically referred to as "upstream") and bring those changes into my branched work.
"Forks" and "Pull Requests" are not git concepts they are github concepts. Forking provided a nice little metaphor for locking down a repo because you couldn't create a disaster for anyone but yourself. Many git saavy organizations do not allow direct access to the authoritative repo instead only allowing "Pull only" access. This allows the authoritative repo to pull the changes that add value and reject those that do not without adding lots of cruft. This harkens the "Social Coding" aspect that github wanted to develop in their earlier days. With everyone contributing to everything forking left and right. Businesses wanted to piggy back of the toolchain they created but most folks weren't familiar with the idea of social coding and/or had needs that weren't addressed well by the social coding model. Github said no worries I think I can tailor this to a business. Which is why you have lots of fork based tools for private repos( organizations can take ownership of private forks, can force forks to be private, can set organizational ACLs on forks and private repos, etc, etc)
Some not so git/github saavy organizations work out of a shared repo for reasons that aren't very clear to me. Its usually a misinterpretation of who owns forks and/or the visibility of private code. I get a sense that github fought these ideas for a long time just saying "c'mon friends just use forks" and that this is their aim at a compromise.
>branches are the "git way," no?
This takes us to this. The answer is sort of. Really cloning is the git way. You clone a repo and synchronize it with other people's repos. Most organizations realize pretty quickly that some repo has to be the repository of record, but git doesn't care. To it a repo is a repo and you know what you are doing. Github just layers a little process on top of the "git way" turning a "clone" into a "fork" and add a little ceremony to the contribution process.
oof talk about over-engineering. What is the point? Just don't allow force pushes to protected branches is a much simpler model than. "Keep a bunch of bookkeeping meta data in the event of an unanticipated force push". All this complexity would only save you during the span of time those objects were collectible but not yet collected.