180 karma · joined October 9, 2015
However, both the Megatron-LM[3] and the DeepSpeed[4] don't use pipedream-2bw scheduling. Could anyone share me some insights or ideas about why such an efficient scheduling scheme doesn't get popular in the LLM pretraining community? Does it suffer convergence/accuracy issue in practice? Or are there any other concerns that blocking it become the default / most popular pipeline parallelism scheduling?
[1]: https://arxiv.org/abs/2006.09503
[2]: https://arxiv.org/abs/2101.06840
Will GLT be part of graphscope[1] and replacing the current graphscope-for-learning implementation?
Closing the "zero changes" merge request have confused in my previous experience when I continue to push re-added commits to the original pull request branch.
It would cause some confusion for project management, I guess.
It is not a real problem that allows anyone to push into any repo, but is a real problem that shows a incorrect/unexpected "Merged" pull request status on any repository.
> 2) Is it just a very confusing message?
From the links in paste in the comments you could see that github shows the unauthorized user "merged" the pull request into main, and, the repo's owner received a email says:
FROM: XXXXX Content: Merged #xxx into main.
It is exactly the SAME as email notification of a normal authorized merge event.
After further inspection, we found that the merge event is triggered by the creator of the pull request pushing current main commits to the PR's "from" branch.
Moreover, when pushing current main to a pull request,
- The pull request is displayed as "Merged"
- A "PR merged into main" email is sent to all subscribers (mainly the repo owners)
- A "PR merged" contribution is displayed on the creator's Github profile
Closing dangling pull requests is a quite resonable design, but mark it as "Merged" rather than "Closed" would confuse people to let them think they are hacked at the first glance (note that there even an email notification "Merged #xxxx into main" to repo's owners).
If such a feature is misused, it may lead to chaos to more open source repos in the futuer, especially those famous ones.
See some example links in
GraphScope is a one-stop graph computing systems from Alibaba aimed to address challenges in large-scale graph computation in real production environments. GraphScope releases v0.9, enabling data scientists to develop graph computing workflows for analytical, interactive query and GNN workloads on small graphs in jupyter notebooks in a interactive manner. Once finishing the development and debugging, users can easily deployed their workflows to Kubernetes with one-line change!
To try GraphScope, you could find it on Colab[1], Jupyter Hub[2], or install GraphScope to your environment using pip by:
pip3 install graphscope
For more details of our v0.9 release, please refer to https://github.com/alibaba/GraphScope/releases/tag/v0.9.0
[1]: https://colab.research.google.com/github/alibaba/GraphScope/...
See how vineyard helps for ETL jobs that large volume of complex data need to be shared between tasks: https://registry.astronomer.io/dags/v6d-etl-pandas
There's also a tutorial to get started: https://v6d.io/notes/airflow.html
After several months of development we are glad to announce our improvements that briging vineyard with the bigdata computing ecosystem, more specifically,
Vineyard v0.2.7 brings many new features into experimental stage, including
+ Airflow Integration with Vineyard: vineyard serves as a XCom backend for airflow to support the tasks in user-defined workflow shares large volume of intermediate data using vineyard, and enables more possibilities for orchestrating bigdata analytical workloads with Airflow. To try Airflow with vineyard, please refer to https://v6d.io/notes/airflow.html
+ Vineyard brings dask integration in this release, data in vineyard can be consumed and produced by dask procedures, and can be further shared by other compute engines that works on Vineyard. For a end-to-end demo, please refer to https://v6d.io/notes/dask.html
+ Vineyard now can serve as the data source for ML frameworks, e.g., tensorflow, pytorch and xgboost. (thanks @chaitravi-ce )
+ Vineyard Go SDK (thanks @linlih) and Java SDK (thanks @zhanglei1949) comes into a very preliminary stage.
Vineyard v0.2.7 will the last release of v0.2.x series and v0.3.0 will come next week!
1. The implementation of lazy evaluation has been improved in this releases, providing users better one-stop experience for graph computation on large-scale data.
2. The GAIA[2] execution engine has been integrated into graphscope, bringing better performance for interactive queries[3].
3. This release brings better support for running GraphScope on MacOS, Ubuntu and CentOS to ease the environment setup for developers.
[1]: https://news.ycombinator.com/item?id=26000243
[2]: https://www.usenix.org/conference/nsdi21/presentation/qian-z...
[3]: https://graphscope.io/blog/tech/2021/08/05/GAIA-Deep-Dive-Bo...
Presentation: https://www.youtube.com/watch?v=p-falphSJq8&list=PLj6h78yzYM...
Live demo: https://www.youtube.com/watch?v=vPbF1l5nwwQ&list=PLj6h78yzYM...
Features and improvements included in v0.2.0 are highlighted as follows:
+ Vineyard has been integration with Helm for easier deployment on kubernetes clusters by simply typing `helm install vineyard vineyard/vineyard`
+ Vineyard objects are now represented using CRDs on kubernetes and are observable to tools built upon kubernetes.
+ A vineyard operator has been included in release v0.2.0, where a scheduler plugin could works in a way where pods are aligned with vineyard server where its required data is placed when possible, which reduce the data migration overhead for big-data analytical pipelines a lot.
+ A malloc-like allocator enables applications to allocate memory chunks directly on vineyard's shared memory arenas, implemented using jemalloc, avoiding the cost of copy and the cost of refactoring the source code to ease the integration with vineyard.
[1]: https://github.com/alibaba/libgrape-lite
[2]: https://github.com/alibaba/libgrape-lite/blob/master/Perform...
We do have a concrete plan for k8s-less deployment and we already have an issue [1] to track that. That will be available before the end of March 2021.
To simplify the environment setup process we will release a docker image for end-users, but without docker will be ok as well (requires building from sources).
GraphScope use vineyard [2] as the storage layer for im-memory graph data structures. And current the graph type (aka. ArrowPropertyFragment in GraphScope) uses a set of arrow tables and arrays under the hood.
GraphScope supports a `to_vineyard_dataframe` method on the computation context [3]. We also has a plan for integration between vineyard and dask (may could be delivered in March as well). At that time the interop between dask would be straightforward.
[1]: https://github.com/alibaba/GraphScope/discussions/113
[2]: https://github.com/alibaba/libvineyard
[3]: https://graphscope.io/docs/reference/context.html#graphscope...
1. The programming model is different. Euler provides a message-passing style API to define new graph models, while GIE provides a sampling API first, and abstracts a batch of seed nodes or edges(named ‘ego’) and their receptive fields (multi-hops neighbors) as a "EgoGraph", which can be turned into a "EgoTensor" as features.
2. Euler implements operation on graphs as tensorflow ops, while GIE takes a more flexible design and doesn't couple with a specific machine learning framework. Upon the sampling interface, developer could build their GNN models using either tensorflow, pytorch or any other machine learning framework, and even for non-GNN tasks.
3. Most importantly, euler is just for graph learning. In graphscope, the learning engine could co-work with the analytical engine and with the interactive query. In real world cases usually so complicated and the problem is not just a GNN training task, and different dedicated systems may involve for different kinds of workload, that means users need to move and transform data back and forth between systems. In GraphScope, engines for different proposals share the same graph, and, live in the same jupyter notebook to delivery the ability of one-stop large-scale graph computation. That is more user-friendly for data scientists.
Typically the graph may have billions of nodes and 10x billions of edges. Obviously the graph data cannot be fit into a single machine to run alogrithms like SSSP or pagerank. And a single machine usually doesn't have enough cores for the computation, e.g., an interactive query couldn't return within milliseconds. That why we need distributed graph processing system for such big graphs.
We just released the version 0.2.0. And along with the release, we launched a public JupyterLab service where you can have a try in your browser: https://try.graphscope.app
Github: https://github.com/alibaba/graphscope. (stars are welcome :)
Website: https://graphscope.io
Documentation: https://graphscope.io/docs
Any comments and contributions from the community are welcomed!
+ feather: as well as arrow, are focused on interoperable data stroage format. feather and arrow define a IPC serailziation schema for common data structures, e.g., tensor and DataFrame. Rather than do serailization for zero-copy (when possible) deserialization, vineyard organized data as a "metadata" and a set of "blobs", the metadata decides how to interpret those blobs and blobs are shared in a zero-copy fashion. Moreover, vineyard could manages large that cannot be fit into a single machine as a "GlobalObject" and global objects can be shared efficiently as well.
+ fst: fst is also a serailization/stroage format, but vineyard is not such a thing.
+ diskframe: diskframe is similar to fst, and works for data that cannot be fit into memory. Vineyard shares in-memory data that cannot be fit into a single machine between different engines (might be implemented in different languages).
+ duckdb: vineyard is not a SQL execution engine.
No, in vineyard "distributed" doesn't mean data replicated across machines, rather, big data partitioned across machines. We address cases where the data cannot be fit into a single machine.
> Can this be run outside kubernetes or is it meant to be part of a big set of infra?
Yes, vineyard can run outside of kubernetes, but it would be a bit complicated to setup on many machines as a cluster.
Devops your vineyard cluster on kubernetes could leverage the ability of Kubernetes. We already have integrated with helm and there's a vineyard-operator in our roadmap.
It is easy to do it on a single machine in a single python runtime because data structures like tensors can be easily shared within a single python runtime.
It will become a bit harder to share the data across processes/Python runtime efficiently without expensive IO or serialization (doable with the help of things like plasma in Apache Arrow).
However, if we want to deal with "big data" that cannot be handled on a single machine, it will become very hard to share data without IO/serialization. Vineyard can be used for such cases. It makes sharing big distributed data structures easy for different runtime.
There will be some added benefits with vineyard: 1. vineyard can handle the the IO, sharding/partitioning, failover and data migrations for the applications.
2. vineyard supports many out-of-the-box highly efficient data formats (such as Apache Arrow) and high-level abstractions which ease the difficulties of developing big data applications.
3. With stream data support, there is a possibility of enabling cross-process optimizations.
4. Fits K8s very well, where common data structures can live in a separated pod/container with its own resource limit.
1. distributed in-memory immutable data storage;
2. zero-copy data access through shared-memory;
3. out-of-the-box high-level data abstractions for commonly used data structures;
4. pre-built drivers for I/O, migration, checkpoint, etc;
5. C++ and Python API; and
6. Kubernetes-integration for large-scale big data applications
Github: https://github.com/alibaba/libvineyard (stars are welcomed!)
Documentation: https://v6d.io
Helm charts: https://artifacthub.io/packages/helm/vineyard/vineyard
Any comments and contributions from the community are welcomed!