Apache Arrow and MinIO
blog.min.io
blog.min.io
- Arrow is an efficient object format that helps skip serialization
- Spark works in memory
- Minio makes storage of data scalable and cheaper than hard disks
- (from the article) Spark jobs don't (didn't??) saturate the network
- (from the article) Spark jobs needed to write the same data to many buffers to get anywhere (exporting?)
- memory mapped IO is here to help (this seems like they're mixing memory mapped device IO
Is the gist that because arrow does not require much serialization/deserialization, memory mapped network devices can shoot it out to another server (w/ S3 as an intermediary) at network speed, skip the disk, and get used at the destination faster? I guess the other method was program-1 -> disk -> network -> disk -> program-2 and now it's program-1 -> s3 -> program-2 ?
To me, the article says "columnar formats that don't require deserialization are good, oh, and you can use Minio to store data!"
And yes it has nothing to do with minio any more than it has with s3 or any cloud FS. There are easy to use writers and readers that also handle streaming.
Last time I checked,
> The most common usage is parquet handling.
> python to JAVA IPC was in place.
> The communication across systems uses protobuf and that is optimised for web-level transfers and not large scale data transfers yet. So transfers still have work to be done
There was also an in-memory object sharing module. I think its no longer part of the repo now.
Most of the time I don't necessarily want disparate languages in one system but it seems inevitable and knowing that Arrow exists seems like it will help me one day.
So if your company is ready to invest in MinIO without S3 compatibility it's a nice software, and my kudos to the team who took the efforts to build it. It's just that it's not fully S3 compatible and MinIO buckets do not behave the same as S3 buckets.
The key thing to note is that it’s probably not fully compatible, but it’s definitely compatible enough for many use cases.
I can also see in the issue you reported that you changed the title of the issue from "Minio is not fully S3 compatible" to "Minio is not S3 compatible", which strikes me as you having some personal beef with them.
In my view it's misleading as minIO is not a drop-in replacement for S3 API based applications. My team spend significant efforts to diagnose and find the error, as I believed in this line and asked them to retry different ways by changing code again and again, even though code was working fine with Amazon S3.
I changed the title so that it can warn others, to not go through the same shooting the foot believing MinIO is a drop in replacement for Amazon S3 based applications. If you read the thread you can notice there isn't any plan to support posix like folder API in minIO which is supported in Amazon S3.
If you look at S3 implementations from the various clouds and NAS appliances and whatnot, you'll see that none of them is a 1:1 compatibilite with each other.
If you wanted to warn others that, instead of Minio not providing "full" S3 compatibility, but Minio does not provide S3 compatibility at all (which is what the title edit suggests), then I believe you are mistaken. Minio definitely provides S3 compatibility, all the main tools and SDKs can talk perfectly with Minio, it's just that there are different implementation details around the edges.
I didn't believe it blindly and that's the reason there is a ticket, after thorough investigation of the error in code as well as self-hosted instance configuration and log file analysis.
Also to make sure others do not need to go through the same long process of debugging, try to request the project maintainers to change this marketing claim as evident in the ticket. Please check the ticket and see for yourself. I did it to make sure others facing the same problem can see it before spending too much efforts to debug.
I can see the amount of work minIO team has put in this project and again kudos to the team. Open source is hard and requires a lot of commitment.
This is how open source works, you give back to community in the form of feedback, documentation, issues and if capable in code.
However, in this case, MinIO makes the deliberate choice to deviate from AWS S3 behavior, on a very common call (listing objects). I do agree that this has the risk to break applications in non-obvious ways, e.g. the call succeeds, bu t your application behavior differs, compared to running against AWS S3, due to the missing directory entry.
I also find their argumentation in the ticket a bit worrisome, calling the AWS S3 behavior a "blunder". I saw the same hubris when looking at their data path, where they choose to ignore battle-tested approaches like Paxos / Raft for distributed coordination, and instead build their own distributed lock algorithm and implementation: it does not seem to persist state, so resilience to power failure and server/process crashes might not be what you expect.
From what I could find, Arrow supports reading (writing?) data from (to?) memory-mapped files (i.e., memory regions created through mmap and friends). However, this has no relation to how the IO is being done, hence not related to access to IO devices using either IO ports, or memory mappings (DMA and such).
This section seems to be mixing up two fairly distinct concepts, i.e., talking about ways to access IO devices and transfer data to/from them (among which memory mapping is an option), where the memory mapping (mmap of files) as used by Arrow is something a little different.
> Thanks to Wes McKinney for this brilliant innovation, its not a surprise that such an idea came from him and team, as he is well known as the creator of Pandas in Python. He calls Arrow as the future of data transfer.
I assume the confusion is with the author of the blog post and not Wes mcKinney, so this callout in that context is a real disservice.
> The output which displays the time, shows the power of this approach. The performance boost is tremendous, almost 50%.
Keep in mind that this is reading a 2 MB file with 100k entries, which somehow manages to consume half a second of CPU time. The author compares wall time and not CPU time; both runs consume somewhere between 600 ms and over a second of wall time (again, handling 2 MB of data). I wouldn't be surprised if the first call simply takes so long because it is lazily loading a bunch of code.
Later on memory consumption is measured, and one of the file format readers manages to consume -1 MB.
This article has a very bad smell.
This is not true. On a Linux system run "cat /proc/ioports" to see the Port mapped I/O space. There are no network card or storage controller registers in Port mapped I/O space on my machine so network and disk I/O does not involve Port mapped I/O. Modern devices tend to use Memory Mapped I/O registers (see "cat /proc/iomem").
Is the author confusing MMIO vs PIO with mmap vs syscalls? They are two different things. Using mmap can be more efficient than using syscalls because the application has direct access to the page cache, no syscall overhead, and no memory copies.
Having a data lake - which I understand as a repository of raw data of diverse types, regardless of the tools - structured in a tool like S3 is very useful when you have multiple use cases over data of different kinds.
For example, you could store audio files from customer calls and have them processed automatically by Spark jobs (e.g. for transcript and stats generation), structure and store call stats on a database for analytics, and do further analysis via notebooks on data science initiatives (e.g. sentiment analysis). This is akin to having a staging area for complex and diverse data types, and S3 is useful for this because of its speed, scale and management features.
Teradata or Snowflake aren't a great fit for use cases like these, but they are great if the use case is to get answers to questions like "top 3 operators per team in volume of calls, by department and region, in last quarter" if the volume of calls is big.
If I understood correctly, I think your comment was more focused on why use new tools when the existing are mature, but I think big data tools have had to become more specialized and targeted for specific use cases. But if the question is "why build more than one data lake", the only reason I can see is organizational: teams or different areas of an organization either need their own data lake because they have specific needs (which is rare) or won't/can't collaborate with others to have a shared asset.
> they are great if the use case is to get answers to questions like "top 3 operators per team in volume of calls..."
You are straw-manning Snowflake/BQ. Just because they are SQL database systems doesn't mean you have to do 100% of your analysis in SQL. You can use other systems, like Spark, PyTorch, Tensorflow, to work with data that you manage inside a RDBMS. There are some issues with bandwidth getting data between systems, but these issues are getting solved (by Arrow!) and in the meantime unload-to/load-from S3 is a good workaround.
I've heard a lot of people make these same arguments and I've tentatively concluded it's mostly motivated reasoning. Engineers like to engineer things. They start by trying to make the obvious, boring system work, but when they run into an obstacle they immediately jump to "I need to build a new system using $TECHNOLOGY."
Databricks Delta Lake has some use cases, but there are some aspects that are rough around the edges. vacuuming is very slow, the design decision to store partitioned data on disk in folders has certain pros / cons, etc.
There are a lot of great products in the data lake space, but lots more innovation is needed going forward.
This is just flatly untrue---they are nearly the same cost/TB as object storage, and they store everything in compressed columnarized format, so they're about as efficient as you can get.
I have heard many people make the same claim. I can't figure it out. Is there something wrong with my calculator???
We switched to it to Parquet on HDFS and did some (ok, a lot...)performance tuning and exactly the same processing ended up at 6 seconds.
TL;DR: You really really need to understand how Spark does data procesisng.
[1]https://medium.com/pinterest-engineering/open-sourcing-terra...
(Any index scan is extremely fast really. You can build your own indexes as separate Parquet files if you want to avoid HBase for some reason)
Of course, you can also get these savings with an online service like Wasabi or Backblaze B2. I just feel that using S3 as your only storage is fine, but once you move to the budget options you need an external backup. So we are using Minio and backing up to Wasabi, and still saving a lot.
S3 is considerably simpler than SFTP or rsync over SSH on non-Unix-like platforms, which makes it a very interoperable API. Even Wordpress or a small Go service can implement S3 support without much effort or any extra dependencies. Being HTTPS-based, it can also be reverse proxied more easily because HTTP presents a Host header, unlike SSH.