Something that I believe we do that I don't see with Plynx is our support for "debugging" pipeline in realtime through extensive Jupyter integration.
When a pipeline step is a Jupyter notebook, you can execute previous steps in the pipeline (partially execute the pipeline graph). It passes data to the running notebook kernel and which allows you to explore it interactively. This makes it a lot easier to guarantee your incoming data matches your expectations.
We also put each step in a pipeline in a separate container (with its own image), which will simplify a lot of things when taking the effort to productionize pipelines to make them suitable for horizontal scale-out.
Edit: running locally can be useful if you want to use your own resources (i.e. your own GPU-rack) to run compute heavy data science pipelines. But we envision for most real-world team usage they'll want to run in the cloud. We're working on making it run directly on top of cloud provided Kubernetes engines.
I would find a comparison table vs. existing tools useful, to help me consider Orchest by placing it in my existing workflow.
I personally like something like GitLab's https://about.gitlab.com/devops-tools/. We'll try to put up something similar on our website at some point.
I'm not deeply familiar with MLFlow, but from what I have seen/read it is more of a tracking framework that you can integrate into an existing codebase.
While Orchest allows you to take your existing codebase and structure it into a pipeline to get a visual and containerized way of interacting with the codebase (allow a mix of notebooks and .py/.sh/.R scripts), running the pipeline, and visually inspecting success/failure of pipeline runs/steps.
Another key point of difference is how we are more concerned with managing the flow of data. Since we let you build pipelines we can give you abstractions to separate data flow from the pipeline code. I.e. letting you define generic pipelines that take any data source (in some standard form, like a schema'd database) and produce reports. Because we control the data source in relation to the containerized pipelines we can also make sure the whole thing performs well when it's being executed in parallel (i.e. same version of the pipeline running grid search over paramaterized pipelines). In other words, we also control more of the underlying infrastructure when executing the pipelines.