Building an Open, Multi-Engine Data Lakehouse with S3 and Python
tower.dev
tower.dev
provision spark on emr or duckdb on beefy ec2 -> run sqlmesh -> wipe resources.
i'm still in MVP phase of revamping my company's current data platform, so maybe there are better alternatives -- which i'd love to hear about.
For building single engine AWS based data lake house you can refer to this article [1], or just use Amazon Sagemaker that also support Iceberg.
Fun Amazon AWS data storage dictionary:
S3: Data Lake
Glacier: Archival Storage
DocumentDB: NoSQL Document Database ala MongoDB
DynamoDB: NoSQL KV and WC Database
RDS: SQL Database
Timestream: Time-Series Database
Neptune: Graph Database
Redshift: Data Warehouse
SageMaker: Data Lakehouse
Islander: Data Mesh (okay kidding, just made this up)
[1] Build a Lake House Architecture on AWS:
https://aws.amazon.com/blogs/big-data/build-a-lake-house-arc...
An error occurred (AccessDenied) when calling the ListObjectsV2 operation: Access Denied
error when trying to run
aws s3 ls s3://mango-public-data/lakehouse-snapshots/peach-lake --recursive