There's "need" and there's "need". The person without the solid computer science background may very well be able to solve all the problems at hand, but the person with the solid CS background is far more likely to come up with the fast elegant solution that doesn't fall over in obscure corner cases.
To use a concrete example for early in my career; I spent days and days trying to solve a problem by building a bigger and bigger pile regexps and if-else statements. Then my project manager (who had a PhD in CS) came along a just wrote a custom parser that solved the whole thing.
This can go both ways, though. Sometimes the fast elegant solution is too elegant for the nasty corner cases and a much bigger, uglier, but ultimately straightforward procedural block of crap is better.
The examples I've seen are from established businesses that occasionally slipped up in refactoring out hacks that were introduced to support deadline-driven requirements. A year or two later, that hack is powering the reporting for a huge portion of the traffic, and you have to be able to balance between (a) writing maintainable code that supports the hack for the short term and (b) working with the rest of the business to clean up the requirements so you can improve the system for the long term.
(b) is a skill that most interview processes I've seen completely ignore. And (a) often is downplayed (as "simple" or "easy") compared to more clever tricks - but knowing how to do the clever tricks might turn into a temptation to use them more than you should.
I know the plural of anecdote is not data, but I'd say that the typical PhD that I have interviewed has generally poorer useful coding skills than the person from a similar background who stopped at a masters or even BS and then wrote a lot of high quality code during the time the PhD was doing a whole bunch of theoretical work.
>To use a concrete example for early in my career; I spent days and days trying to solve a problem by building a bigger and bigger pile regexps and if-else statements. Then my project manager (who had a PhD in CS) came along a just wrote a custom parser that solved the whole thing.
It's MUCH easier to come along after someone has already spent days exploring the problem and can demonstrate the dead ends, then say "oh, you need a custom parser" than to start at the beginning and do the same thing. USUALLY writing a custom parser is overkill and it takes a decent amount of digging into a system to know when it's not. If you start out writing one every time something looks like it can be solved by an "if" statement, you'd never get anything useful built.
Maybe you aren't near water or writing performance-sensitive code 95% of the time, but the 5% of the time you're in over your head, being able to implement (or identify when to use) basic algorithms will save you a whole lot of time, money, and headache.
However, I strongly disagree with your statement!
I work on distributed systems day in and day out, and more and more I find that I'm using or building distributed data structures analogous to BST, HashTable, etc.
Not knowing the foundations of my field in and out would be a mistake!
Disclaimer: Never worked in financial tech, I have thought about this problem for 5 mins. This is unlikely to be an optimal solution to keeping a sorted set of distributed data. I'm just trying to show how basic knowledge of data structures helps in the field of distributed systems.
Lets say you need to keep a sorted set of sooooo much data that you cannot keep it all on one machine. I would imagine all financial services firms have some custom datastore which has "potential buy orders" and need to keep them all sorted by profitability to efficiently search/insert/remove from that datastore.
In such a situation, my instincts would be to create a distributed, redundant Heap. Now, how do you build a distributed, redundant Heap, without understanding how a Heap works? How do you even know what to look for if you don't know what a Heap is?
Now, lets say you've built it such that each node in your distributed network acted like a node in a Heap, and the left "pointer"(i.e. url to other computer) meant "less than" and the right "pointer" meant "greater than". Now you have this data store in production and all of a sudden you realize performance worsens over time... "WHY is this happening? Is there some memory leak?". You investigate and realize that all of your data isn't being distributed evenly! "Do I need to like.. shuffle the data around? How do I get it so that the data is evenly distributes?"
At this point if you do not know what a red-black tree is, what do you even google look for? Lets say you dig around for a while and find out about re-balancing trees. Now you have to implement it, and eventually you figure out that one of the best ways is to color your nodes. When a future colleague/manager asks, "What is this system, please describe it to me?" wouldn't it be nice to just say "Oh visualize it kinda like a distributed red-black tree. It holds XXX data, and ensures that it is always sorted and re-balanced for optimal performance."
With dozens of database stacks to choose from, many companies won't have anybody actively and regularly working in the database's code, nor in the OS kernel code either.
If, then, an individual can successfully contribute to the company without ever hearing of a red-black tree, what's the value in testing for it, past "this is how we've always done it"?
(How long it would take someone who doesn't know CS lingo to string together "tree" and "balance" and plug that into Google, I don't know, but I suspect it's not impossible to find.)
You've misunderstood me. I'm literally talking about building a distributed B-Tree or Heap. As in, a Heap which cannot fit in memory or storage of any one node.
Think: S3/DynamoDB is a distributed, highly available HashTable; _____ is a distributed, highly available Heap.
Just because those topics show up infrequently doesn't mean that no one's using them. Plenty of companies outside of Google, etc. are doing complicated things that require a solid CS background.