Outside of legacy systems, Hadoop isn't widely used anymore.
Outside of legacy systems, Hadoop isn't widely used anymore.
Parquet is great, but it’s simply nowhere near as ubiquitous as CSV.
What’s the Parquet equivalent of going to the store to buy Mentos and Diet Coke now?
It's easy to forget just how much analytic "stuff" Excel still powers.
If you don't mind being old-school, the data is ASCII text, and you're tired of some of CSV's little issue, then ASCII has the FS, GS, RS, and US control characters - specifically intended for such uses. Micro conceptual overhead, none of the CSV issues which screw up many *nix text-handling programs and little script files, and a decent modern filesystem can handle the compression separately.
There are many reasons why CSV is flawed for the purposes of storing tabular data (e.g. loss of column type information) but the alternatives are just so unergonomic that CSV remains a viable choice in many situations.
> If you don't mind being old-school, the data is ASCII text, and you're tired of some of CSV's little issue, then ASCII has the FS, GS, RS, and US control characters
I have used those before, and yet I still had those characters appear in data. The only places I'd ever seen them were in the wiki page and in customer delivered data. Absolute pain to dig through and remove.
On top of that "if your data is ASCII" is something I'd be nervous about for many use cases even if it is right now.
Beyond that, then you need everyone to swap out their parsing to use those characters.
CSV is fine until it totally blows up in your face. All it takes is one "oh it's fine we'll use awk" stage somewhere or a CSV parser that isn't good enough and one person to put a newline where nobody had expected it before.
Oh, yes - which is why I emphasized "is". But ASCII text is easy to test for, which lets you fast-track into exception handling - "Tell Sales that Customer data is not as represented", "Trouble-shoot internal data source", etc.
(My experience is that substantial Customer data is never, ever as initially represented. Nor as represented after you point out the first set of issues with it. Nor as represented after you point out the second set of issues. Nor as...)
Oh with that I mean the ASCII control characters appearing in inputs. So some columns would have record end markers in for example.
If I'm able to make everyone dealing with the reading and writing add specific characters to be used for start/end/etc I'd rather just tell them to swap to a parquet reader unless they've got a really good reason.
Feather is a layer on top of arrow and was a proof of concept (so I'm not sure how heavily it's used now), and arrow is fast becoming the interchange format. It's exactly laid out as things will be in memory - which means zero copy for shuttling it around from one place to another. I _think_ there is less support for feather but that is likely changing as everything converges.
Parquet should be
* Faster to write * Faster to read (even if you're reading the whole file, which actually isn't required, the format helps you read just sections of the columns you need) * Smaller * Better at handling actual floating points
than CSV, while having actual standards alongside it. Be a little wary of pandas guessing the right column types for you if you're creating partitioned files btw.
When you're working with pandas, etc (check out Dask) you can pretty much just swap out some reading and writing functions. You can also use pyarrow directly if you need to be very careful about column types.
For your use case you may want to explicitly use a single column for the features that is a list, I'm not sure if that's better/worse than having so many columns. If a reader may want to find just some images where a small subset of features are > X, you might benefit from multiple columns so that the reader only processes the data it needs.
Worth testing out, but I expect you should be able to try it out in an afternoon if you're already working with pandas/similar. Just install pyarrow and use a to_parquet. Things like dask (or straight pyarrow) give you partitioned files as output if you want too, if there's a useful column or columns to split on https://arrow.apache.org/docs/python/parquet.html#partitione...
Structure packing[1] and consideration of locality of reference[2] would need to be applied for high performance applications where a computer scientist has considered the algorithm needing to be implemented and the most efficient data format that the source data would need to be provided in.
This is not true at all. Almost all the cloud providers have their own Hadoop distributions that is used a lot in many companies.
> However, a very common setup is to use Flink to analyze data stored in the Hadoop Distributed File System (HDFS). -- https://wints.github.io/flink-web//faq.html
I just write SQL in Snowflake and it replaces 95% of what I would otherwise have done in custom MapReduce or Spark code.
1) Why would you want to maintain your own Spark infrastructure? Spark on Kube is a huge improvement over YARN but you still have to deal with OOMEs, filled disks, Kube upgrades, pushing custom images to container registries, etc etc etc.
2) Snowflake is probably 10-50x as performant as Spark for data manipulation. I don't know what kind of unholy demonic incantations Snowflake is doing on the backend to support their SQL performance, but it's really freaking fast. There's just no other way to cut it.
I've spent 5-10 years eking every ounce of performance I can get out of a Hadoop/Spark cluster. I'm not trying to be unreasonable about this. I would love for OSS to be competitive; it's great for the world, and it would be great for my skill set and earning potential.
But it's not a contest, and if you think standalone Spark is going to be a viable competitor in a couple years, you are deluding yourself. Make informed choices about your career and investment.
My only concern is that they offer just a managed cloud product. That's cool for startups, but large enterprises sometimes need more governance and ownership than that.
https://en.wikipedia.org/wiki/Data_vault_modeling
Edit: As the Wikipedia article has no Criticism section I will add some references:
http://kejser.org/the-data-vault-vs-kimball-round-2/
https://timi.eu/blog/data-vaulting-from-a-bad-idea-to-ineffi...
Wow is this for a fact? I haven't used either in a while but I saw the blog post from databricks and Spark was more performant than snowflake.
I assumed that's what I'll also get when i run spark on kube
In the early 2000s, columnar relational data warehouses were not sophisticated and scalable enough to handle the scale of data encountered at Yahoo, Google and other internet companies. MapReduce (and the many evolutions of Hadoop ecosystem) was created to scale processing through low-level instructions and algorithms.
Eventually columnar data warehouses caught up and are now capable of handling petabyte scale, regardless of whatever language you use to query them. The fundamental storage and compute primitives haven't really changed that much, just offered in a much more user-friendly way now.
It's all using the same principles underneath.
This is by design. The less skill required to use a tool, the less some CTO has to pay the person using it.
- skip streaming entirely and have near real time solutions using just storage + a query engine
- have streaming using message queues and lambda architecture
In both cases the goal is that your freshest data shows up on a dashboard.
Or, as the original article says, some companies just use some command line tools, shell scripts.
It's been a couple of years since I was interested in Data Engineering, so my knowledge on this topic is some years behind.
I think we're seeing a big shift with Hadoop-like workloads being moved onto cloud providers, so BigQuery, Amazon EMR etc.
My general rule of thumb is whether it is too big to put on my laptop. So greater than a couple of Tb's.
Big data is a moving target, but I’m comfortable defining it as data too large to fit in memory. Obviously, you can always get a bigger node, my rule is thumb is that if you need generators, you are working with big data.
GCP: Dataproc
Those are just the obvious ones though.
Quote from AWS: "EMRFS is an implementation of the Hadoop file system ..."
https://docs.aws.amazon.com/emr/latest/ManagementGuide/emr-p...
Then we probably need the big tools.
> Big data does exist in the wild.
So does little data.
The problem is that a "one-size-fits-all" approach has become common, not just in data analysis; think of all the low-medium traffic webpages that use giant, complex frameworks and huge distributed systems just to display essentially a small CRUD app that would have been ALOT easier to cobble together in plain JS on a simple LAMP server.
What the article shows is the importance on deciding for the right tool for the job: When I want to plant a little tree in my backyard, bringing one of these https://upload.wikimedia.org/wikipedia/commons/0/01/Bucket_w... to dig the hole is proooobably overengineering it a tiny little bit, and will likely take longer than getting a shovel.