30 karma · joined August 1, 2016
Don't spend more than a sentence or two discussing your previous role. Describe their business model and leave it at that.
90% of your interview should be finding out what skills and tech your new employer needs and explaining to them why you're perfect for delivering those skills.
If candidates could see the other applicant's CVs it would kick off an arms race.
If you apply for a job, the experience you list should match the tech stack in the job spec. Don't list roles that aren't relevant (i.e. no .net if it's a Python role).
I helped hire a DS this week. Their CV was among 20 excellent candidates all in London with their latest role matching most of the tech stack we're looking for. All the candidates are in London. This is what you're competing against.
Where is the transactional requirement? This person is working with a copy of the real data.
ETLs only need to be written once and if he decided on a PSQL approach he'd be writing ETLs to send the data there too. He's probably going to find a number of consistency problems so trying to normalise all this data again will just result in more work that won't make his team of DS' more productive.
If he's at ~1 TB of data today, where will he be in a few years time? What's the point of putting infrastructure in place that won't last for the next 10+ years?
Storing that data on S3 is probably 50% the price of storing it on EBS and you won't have the durability guarantees of S3 when you're using PostgreSQL on EBS volumes.
If you're both exploring data and building models then Spark is fine. Its APIs are no more complicated that anything else out there for these tasks.
Hive is doing nothing more than offering schema on read and shouldn't be something you're thinking much about.
PostgreSQL is row-oriented and won't be able to offer features like row-group statistics that allow queries to get minimum and maximum values for every 10-15K rows of data for the columns their interested in. This gives queries a huge speed up over needing to scan over rows rather than just the statistics for the columns their interested in.
Remember that you can have a single engineer run a single query on Spark and distribute it across several servers. This allows you to scale CPU and memory bandwidth in a way you won't be able to with PostgreSQL.
It sounds like your data isn't well organised. If you moved it around and put some consistent naming conventions in place that could help. You could also look to build an atlas of the data for newcomers to get an overall picture of what data you're storing and where it lives.
> Linuxbrew does not currently support 32-bit x86 platforms. It would be possible for Linuxbrew to work on 32-bit x86 platforms with some effort. An interested and dedicated person could maintain a fork of Homebrew to develop support for 32-bit x86.