Data Version Control (DVC) for Machine Learning Projects
medium.com
medium.com
What I don't understand about them is they have a built-in pipeline. it is strange to me, because I would use airflow to handle CI/CD and keeps the version control simple. Mixing things together seems to be against "Do One Thing and Do It Well".
Are those pipelines function like git's pre/post commit hooks?
DVC provides a command `dvc repro` to actually execute the pipeline definition and doesn't dictate where it's run; so you can decide to run it in your CI environment or pipeline orchestration tool or on GPU-powered hardware.
Btw, what do you mean by "handle CI/CD"?
In my opinion, what's versioned in DVC should have been cleaned up already. So I would offload ETL to airflow and only version the result of it.
Thank you for your explanation. that makes sense to me. I think it's like the project.json of a nodejs project, where you can define a set of commands, such as npm build, npm run, npm serve ...