2,392 karma · joined September 6, 2012
The blog itself is great, providing detailed accounts about iconic products and business initiatives in the tech industry.
Having a data lake - which I understand as a repository of raw data of diverse types, regardless of the tools - structured in a tool like S3 is very useful when you have multiple use cases over data of different kinds.
For example, you could store audio files from customer calls and have them processed automatically by Spark jobs (e.g. for transcript and stats generation), structure and store call stats on a database for analytics, and do further analysis via notebooks on data science initiatives (e.g. sentiment analysis). This is akin to having a staging area for complex and diverse data types, and S3 is useful for this because of its speed, scale and management features.
Teradata or Snowflake aren't a great fit for use cases like these, but they are great if the use case is to get answers to questions like "top 3 operators per team in volume of calls, by department and region, in last quarter" if the volume of calls is big.
If I understood correctly, I think your comment was more focused on why use new tools when the existing are mature, but I think big data tools have had to become more specialized and targeted for specific use cases. But if the question is "why build more than one data lake", the only reason I can see is organizational: teams or different areas of an organization either need their own data lake because they have specific needs (which is rare) or won't/can't collaborate with others to have a shared asset.
What is the benefit of using Kubernetes to deploy Spark jobs then? Is that approach meant to achieve independence from the hardware?
I'm asking because that is fairly trivial to achieve using, at least, a provider like AWS: you can build a CloudFormation template (or use the AWS API or the web UI) to launch AWS EMR clusters with specific hardware and run any spark jars, and you can use services like DataPipeline or Glue to schedule and/or automate the whole process. So you can use AWS services to set up a schedule that will periodically spin up a cluster with whatever machines you need to run a Spark app and decommission it as soon as its done.
In this case, the EMR cluster comes with the myriad of Hadoop tools and services (and Spark, and other relevant software) preinstalled and ready to use. And most relevant Spark settings are already optimized for the cluster's hardware; but not for the Spark app itself, which is what this solutions seems to address.
Although I've experienced very rare crashes, like I have with Chrome/Chromium, my experience is that Firefox is stable and does everything I need, including some basic web dev/debug.
That's exactly how that problem has been solved successfully for the past 20 years.
Filtering ssh connections at firewall level helps, and certainly reduces log entries for port scans and can halt less sophisticated attackers, but it doesn't mitigate attack vectors like a DDoS or a well funded actor.
It stems from automated port scans sweeping IP ranges looking for vulnerable/exploitable systems.
A simple way to address it is to use something like fail2ban to ban repeat offenders. Blacklist services like abuseipdb.com may help further, but ban with caution.
While I don't have hard data on this (and it would be hard to establish causality anyway), most of the ERP implementation failures I've witnessed and know about stem from line of business people who firmly believe their company's processes - e.g. collection, billing, etc - are very special and unique, when it is very rarely the case; and when it is, it shouldn't be, the process is likely overly complex due to inertia or lack of will/skills to improve or just legacy. Then, the customizations that need to be implemented to support the "special and unique processes" are so complex that the project's schedule and budget inevitably explode...
ERP implementations could actually be seen as a great opportunity to revisit, optimize and document current processes before the ERP is implemented...
What's actually mind boggling for me, and I wonder if it's a bit of over-engineering, is people going for complex setups (oh, it's just Airflow scripts with k8s and a little bit of SystemD services and configs plus some shell scripts) when there are COTS tools that do more for less engineering cost. Yes, these carry a price tag, but it's usually quite less than paying for engineers to babysit a tool with a ton of moving parts...
These activities are usually managed by cron and more often by advanced scheduler tools (depending on the vendor), so it's quite a core part of any architecture that needs to e.g. load/reload/refresh data periodically.
If the requirement is simply to connect notebooks to a data lake, then the only scheduling required is to load the data lake, and something like Airflow may be overkill for this, depending on what/how the data is processed and loaded.
Spark would have a been a simple option to do this kind of processing, with less lines of code, and could also run on "spare compute". Same goes for the "How does one process 10 million JSON files taking up just over 1 TB of disk space in an S3 bucket?": there are appropriate file formats for storing and querying big datasets, text/json is simply the least efficient option and likely the cause of the "$2.50 USD per query" number...
But I agree that this level of privacy invasion should be illegal.
So, everything that is currently delivered in the real world by technologies like OutSystems, which have an enterprise customer base, 1-click deploys, and that you can actually try for free.
" Let the officers and the directors and the high-powered executives of our armament factories and our steel companies and our munitions makers and our ship-builders and our airplane builders and the manufacturers of all other things that provide profit in war time as well as the bankers and the speculators, be conscripted — to get $30 a month, the same wage as the lads in the trenches get."
in https://en.wikipedia.org/wiki/War_Is_a_Racket
Edit: corrected "general's number of stars"