Meltano: ELT for the DataOps era
meltano.com
meltano.com
Meltano and pipelinewise are both ways how to orchestrate Singer.io taps and targets while Airbyte is its own thing.
EDIT: I also don't know how just released Airbyte can be ahead of something that's surely in production for a while.
Since Stitch, Meltano and Pipelinewise all use Singer.io taps and targets under the hood, I wonder if there's any reason to choose one over the other?
Meltano and Pipelinewise are open source projects that someone built for themselves but are sharing. You can just start playing with it and change the code or whatever, but there's no support to pay for.
For example where I'm standing the best one would be "Stitch I can self-host for free for a PoC and then eventually engage vendor about a support contract while still self-hosting it for security reasons" but there doesn't seem to be anything like it.
[0] https://gitlab.com/meltano/meltano/-/issues/2616 [1] https://github.com/singer-io/getting-started/blob/master/doc...
the good news is the implicitly typed json examples look arrow friendly, so users can to/from_json if they don't care about data speed/quality like when prototyping and not think about it. there may be other data-engineering-friendly formats that'd work too.
prefect, dask, and friends solve it by abstracting over it. you can send whatever you want.. and it happens to be friendly to dataframes (pydata) / compact & typed data. but there projects seem to be more about source/sink, so encouraging structure by default would be helpful...
It's up to the target what it does with the JSON messages it receives, so you can for example have a target-avro that takes JSON records and outputs them as an Avro file and translates the JSON schema to the corresponding Avro schema.
Then the holy grail would be to have bunch of taps-targets running in parallel for a single pipeline, each working on a subset of streams.
[0] https://about.gitlab.com/handbook/business-ops/data-team/pla... [1] https://gitlab.com/gitlab-com/www-gitlab-com/-/merge_request...
PS. Like Taylor, I'm on the Meltano team at GitLab.
[0] https://blog.getcensus.com/dbt-the-etl-elt-disrupter/ [1] https://www.getdbt.com/
VC valuations in Silicon Valley for traditional "ETL" tools isn't great. The term "ETL" has baggage of Talend and other big enterprise tools that have an aura of "not being user friendly".
"ELT" is intended to signal that a tool is trying really hard to not be like those sad, difficult to use, dirty "ETL" tools.
That's the real reason. There really isn't an actual difference in how the tools work. Traditional ETL tools are perfectly capable of loading the data and initiating the transform in the database it's being stored in. This has been done for decades.
I have been repeatedly lectured by senior strategy team executives about using the proper vernacular to ensure perception of value is maximized in the trendy, buzzword driven VC world. So fucking stupid.
Doesn't make much sense to me, because anyone who's been in the data space quickly sees the rampant marketing of different tools that all do the same thing but simply change the order of the pipeline.
In other words just get the data in house into a queryable form and worry about the T part later.
How would Meltano or the other mention tools handle this?
Example is EDIFACT or FHIR or BDT
[0] https://www.dataengineeringpodcast.com/meltano-data-integrat...