HNHacker News
TopNewBestAskShowJobs

sighingnow

180 karma · joined October 9, 2015

sighingnow@gmail.com
submissionscomments
sighingnow··on GraphScope Journey: Visualizing Progress&Lessons in Graph Computing at Alibaba
We have unveiled a new webpage that tells the story of GraphScope, charting its course from inception to its future prospects in the field of graph computing. The page explores how Alibaba Cloud has developed GraphScope from its early stages, overcoming challenges and achieving significant milestones to become a crucial tool in graph data processing. This journey highlights not only technological advancements but also practical lessons in scaling, performance, and adaptability.
sighingnow··on Why async gradient update doesn't get popular in LLM community?
The pipedream-2bw paper[1] and the zero-offload paper[2] both show that 1-step delayed asynchronous gradient update doesn't affect the convergence (and perplexity) while improve the training efficiency (by fully utilize the bubbles in pipeline parallelism) at a large margin.

However, both the Megatron-LM[3] and the DeepSpeed[4] don't use pipedream-2bw scheduling. Could anyone share me some insights or ideas about why such an efficient scheduling scheme doesn't get popular in the LLM pretraining community? Does it suffer convergence/accuracy issue in practice? Or are there any other concerns that blocking it become the default / most popular pipeline parallelism scheduling?

[1]: https://arxiv.org/abs/2006.09503

[2]: https://arxiv.org/abs/2101.06840

[3]: https://github.com/nvidia/Megatron-LM

[4]: https://github.com/microsoft/DeepSpeed/issues/1110

sighingnow··on Show HN: Graphlearn-for-PyTorch, distributed graph learning on PyTorch
Optimizing distributed sampling and feature lookup looks really attractive. It's really challenging to deploy GNN training at an industrial-scale for a large graph.

Will GLT be part of graphscope[1] and replacing the current graphscope-for-learning implementation?

[1]: https://github.com/alibaba/GraphScope

sighingnow··on Tencent TDSQL breaks TPC-C record
On March 30th, the Transaction Processing Performance Council (TPC) released its latest rankings: Tencent Cloud's TDSQL set a new record with a performance of 814 million tpmC and a cost of 1.27 RMB per tpmC. This achievement placed TDSQL at the top spot for both tpmC performance and cost-effectiveness in the world.
sighingnow··on Show HN: Libobjectfs – zero-dependency, single-artifact library for many storage
More language bindings, e.g., Python and Rust, will be added later. Don't hesitate to open issues for feature request if you are struggling with handle various different I/O and its dependency hell trouble and find this library might be helpful in your application!
sighingnow··on GitHub “allows” unauthorized users “merging” PRs, bypass write permission check?
Allowing "zero changes" state for a merge request does make sense, as during the development workflow developers may doing some rebase & force push work and they may forget commit changes after reset to main and before push to the pull requests.

Closing the "zero changes" merge request have confused in my previous experience when I continue to push re-added commits to the original pull request branch.

sighingnow··on GitHub “allows” unauthorized users “merging” PRs, bypass write permission check?
The "zero-commit-merges" should be treated as a "Closed" event better than being treated as a "Merged" event.
sighingnow··on GitHub “allows” unauthorized users “merging” PRs, bypass write permission check?
+1 for the feature of marking as merged, as in many projects, e.g., apache arrow, the pull requests are merged in a different way and all pull requests are showed as "closed" rather than "merged".

It would cause some confusion for project management, I guess.

sighingnow··on GitHub “allows” unauthorized users “merging” PRs, bypass write permission check?
You don't need to get accept to "merge" a PR. When pushing the main branch to your PR, the PR would get closed as "Merged" and there will be a "Merged" contribution badage showed in your github profile.
sighingnow··on GitHub “allows” unauthorized users “merging” PRs, bypass write permission check?
The concern here is not about who "commit" this, the concern is who "click the merge button" in this case.
sighingnow··on GitHub “allows” unauthorized users “merging” PRs, bypass write permission check?
> 1) Is it a real problem that allows anyone to push into any repository?

It is not a real problem that allows anyone to push into any repo, but is a real problem that shows a incorrect/unexpected "Merged" pull request status on any repository.

> 2) Is it just a very confusing message?

From the links in paste in the comments you could see that github shows the unauthorized user "merged" the pull request into main, and, the repo's owner received a email says:

FROM: XXXXX Content: Merged #xxx into main.

It is exactly the SAME as email notification of a normal authorized merge event.

sighingnow··on GitHub “allows” unauthorized users “merging” PRs, bypass write permission check?
We recently noticed some strange email notifications from Github with content "Merged #948 into main.", which told that pull requests on our repo has been merged by stranger who actually doesn't have the write permission to our repo!

After further inspection, we found that the merge event is triggered by the creator of the pull request pushing current main commits to the PR's "from" branch.

Moreover, when pushing current main to a pull request,

- The pull request is displayed as "Merged"

- A "PR merged into main" email is sent to all subscribers (mainly the repo owners)

- A "PR merged" contribution is displayed on the creator's Github profile

Closing dangling pull requests is a quite resonable design, but mark it as "Merged" rather than "Closed" would confuse people to let them think they are hacked at the first glance (note that there even an email notification "Merged #xxxx into main" to repo's owners).

If such a feature is misused, it may lead to chaos to more open source repos in the futuer, especially those famous ones.

See some example links in

- https://github.com/v6d-io/v6d/pull/948

- https://github.com/alibaba/GraphScope/pull/1931

sighingnow··on GraphScope on Colab: Large-Scale Graph Computing from Notebooks to Kubernetes
We are glad to announce the landing of GraphScope on Colab: https://colab.research.google.com/github/alibaba/GraphScope.

GraphScope is a one-stop graph computing systems from Alibaba aimed to address challenges in large-scale graph computation in real production environments. GraphScope releases v0.9, enabling data scientists to develop graph computing workflows for analytical, interactive query and GNN workloads on small graphs in jupyter notebooks in a interactive manner. Once finishing the development and debugging, users can easily deployed their workflows to Kubernetes with one-line change!

To try GraphScope, you could find it on Colab[1], Jupyter Hub[2], or install GraphScope to your environment using pip by:

pip3 install graphscope

For more details of our v0.9 release, please refer to https://github.com/alibaba/GraphScope/releases/tag/v0.9.0

[1]: https://colab.research.google.com/github/alibaba/GraphScope/...

[2]: https://try.graphscope.app/

sighingnow··on Vineyard serves as the data back end for airflow
Vineyard, an open-source immutable in-memory data-manager that is designed for optimize the data exchange between tasks inside a bigdata analytical workflow, now supports airflow and can serve as the data backend for efficient data sharing between tasks.

See how vineyard helps for ETL jobs that large volume of complex data need to be shared between tasks: https://registry.astronomer.io/dags/v6d-etl-pandas

There's also a tutorial to get started: https://v6d.io/notes/airflow.html

sighingnow··on Vineyard 0.2.7: Airflow, Dask, and better ML experience
We previously have seen many possible scenarios that vineyard could be helpful in https://news.ycombinator.com/item?id=25832831

After several months of development we are glad to announce our improvements that briging vineyard with the bigdata computing ecosystem, more specifically,

Vineyard v0.2.7 brings many new features into experimental stage, including

+ Airflow Integration with Vineyard: vineyard serves as a XCom backend for airflow to support the tasks in user-defined workflow shares large volume of intermediate data using vineyard, and enables more possibilities for orchestrating bigdata analytical workloads with Airflow. To try Airflow with vineyard, please refer to https://v6d.io/notes/airflow.html

+ Vineyard brings dask integration in this release, data in vineyard can be consumed and produced by dask procedures, and can be further shared by other compute engines that works on Vineyard. For a end-to-end demo, please refer to https://v6d.io/notes/dask.html

+ Vineyard now can serve as the data source for ML frameworks, e.g., tensorflow, pytorch and xgboost. (thanks @chaitravi-ce )

+ Vineyard Go SDK (thanks @linlih) and Java SDK (thanks @zhanglei1949) comes into a very preliminary stage.

Vineyard v0.2.7 will the last release of v0.2.x series and v0.3.0 will come next week!

sighingnow··on Show HN: GraphScope v0.6.0: lazy evaluation and GAIA
The GraphScope[1] team is pleased to announce the v0.6.0 release after 2 months of development work and 98 commits in the main repository. We would like to highlight the following major added features:

1. The implementation of lazy evaluation has been improved in this releases, providing users better one-stop experience for graph computation on large-scale data.

2. The GAIA[2] execution engine has been integrated into graphscope, bringing better performance for interactive queries[3].

3. This release brings better support for running GraphScope on MacOS, Ubuntu and CentOS to ease the environment setup for developers.

[1]: https://news.ycombinator.com/item?id=26000243

[2]: https://www.usenix.org/conference/nsdi21/presentation/qian-z...

[3]: https://graphscope.io/blog/tech/2021/08/05/GAIA-Deep-Dive-Bo...

sighingnow··on Show HN: Vineyard v0.2.0: big-data applications optimization on Kubernetes
Vineyard, an open-source in-memory data manager, was previously announced at three months ago[1]. After three-months iteration vineyard release v2.0.2 finally comes out, with a more tight integration with kubernetes and enables orchestration of both "data" and "tasks" for data-intensive applications in cloud-native environment. Refer to our presentation and live demo for more details:

Presentation: https://www.youtube.com/watch?v=p-falphSJq8&list=PLj6h78yzYM...

Live demo: https://www.youtube.com/watch?v=vPbF1l5nwwQ&list=PLj6h78yzYM...

Features and improvements included in v0.2.0 are highlighted as follows:

+ Vineyard has been integration with Helm for easier deployment on kubernetes clusters by simply typing `helm install vineyard vineyard/vineyard`

+ Vineyard objects are now represented using CRDs on kubernetes and are observable to tools built upon kubernetes.

+ A vineyard operator has been included in release v0.2.0, where a scheduler plugin could works in a way where pods are aligned with vineyard server where its required data is placed when possible, which reduce the data migration overhead for big-data analytical pipelines a lot.

+ A malloc-like allocator enables applications to allocate memory chunks directly on vineyard's shared memory arenas, implemented using jemalloc, avoiding the cost of copy and the cost of refactoring the source code to ease the integration with vineyard.

[1]: https://news.ycombinator.com/item?id=25832831

sighingnow··on GraphScope: A One-Stop Large-Scale Graph Computing System
We don't have a benchmark between the analytical engine in GraphScope (aka. GAE) with GraphX/Giraph. But we do have evaluated the performance of the underlying engine of GAE (libgrape-lite) with LDBC Graph Analytics Benchmark and it achieves higher performance comparably to the state-of-the-art systems [2].

[1]: https://github.com/alibaba/libgrape-lite

[2]: https://github.com/alibaba/libgrape-lite/blob/master/Perform...

sighingnow··on GraphScope: A One-Stop Large-Scale Graph Computing System
Thanks for you interests on GraphScope!

We do have a concrete plan for k8s-less deployment and we already have an issue [1] to track that. That will be available before the end of March 2021.

To simplify the environment setup process we will release a docker image for end-users, but without docker will be ok as well (requires building from sources).

GraphScope use vineyard [2] as the storage layer for im-memory graph data structures. And current the graph type (aka. ArrowPropertyFragment in GraphScope) uses a set of arrow tables and arrays under the hood.

GraphScope supports a `to_vineyard_dataframe` method on the computation context [3]. We also has a plan for integration between vineyard and dask (may could be delivered in March as well). At that time the interop between dask would be straightforward.

[1]: https://github.com/alibaba/GraphScope/discussions/113

[2]: https://github.com/alibaba/libvineyard

[3]: https://graphscope.io/docs/reference/context.html#graphscope...

sighingnow··on GraphScope: A One-Stop Large-Scale Graph Computing System
The learning engine in GraphScope (aka. GIE) has a simiary programming interface and functionlity with euler. They both support graph neural network training, e.g., GraphSAGE, GCN, GAT, etc. However there are many differences under the hood.

1. The programming model is different. Euler provides a message-passing style API to define new graph models, while GIE provides a sampling API first, and abstracts a batch of seed nodes or edges(named ‘ego’) and their receptive fields (multi-hops neighbors) as a "EgoGraph", which can be turned into a "EgoTensor" as features.

2. Euler implements operation on graphs as tensorflow ops, while GIE takes a more flexible design and doesn't couple with a specific machine learning framework. Upon the sampling interface, developer could build their GNN models using either tensorflow, pytorch or any other machine learning framework, and even for non-GNN tasks.

3. Most importantly, euler is just for graph learning. In graphscope, the learning engine could co-work with the analytical engine and with the interactive query. In real world cases usually so complicated and the problem is not just a GNN training task, and different dedicated systems may involve for different kinds of workload, that means users need to move and transform data back and forth between systems. In GraphScope, engines for different proposals share the same graph, and, live in the same jupyter notebook to delivery the ability of one-stop large-scale graph computation. That is more user-friendly for data scientists.

sighingnow··on GraphScope: A One-Stop Large-Scale Graph Computing System
GraphScope addresses on computations on extermely large graphs, e.g., the friendship networks of all facebook users, the hyperlink relations between all webpages around the world, and so on.

Typically the graph may have billions of nodes and 10x billions of edges. Obviously the graph data cannot be fit into a single machine to run alogrithms like SSSP or pagerank. And a single machine usually doesn't have enough cores for the computation, e.g., an interactive query couldn't return within milliseconds. That why we need distributed graph processing system for such big graphs.

sighingnow··on GraphScope: A One-Stop Large-Scale Graph Computing System
GraphScope is a unified distributed graph computing platform that provides a one-stop environment for performing diverse graph operations on a cluster of computers through a user-friendly Python interface. GraphScope makes multi-staged processing of large-scale graph data on compute clusters simple by combining several important pieces of Alibaba technology for analytics, interactive, and graph neural networks (GNN) computation, respectively, and the vineyard store that offers efficient in-memory data transfers.

We just released the version 0.2.0. And along with the release, we launched a public JupyterLab service where you can have a try in your browser: https://try.graphscope.app

Github: https://github.com/alibaba/graphscope. (stars are welcome :)

Website: https://graphscope.io

Documentation: https://graphscope.io/docs

Any comments and contributions from the community are welcomed!

sighingnow··on Vineyard: An open-source in-memory data manager
Thanks for liking this project! Your understanding is correct. If you want to give vineyard a try, you are welcome to drop me an email (sighingnow@gmail.com), I can help you set it up.
sighingnow··on Vineyard: An open-source in-memory data manager
Thanks for your interest on this project. Training multiple models using the same data is truely a typical use case for vineyard. We will release more integeration with machine learning engines (tensorflow and pytorch) in the next major release.
sighingnow··on Vineyard: An open-source in-memory data manager
Vineyard addresses on sharing "distributed abstracted immutable data" in memory, e.g., dask[1] generates a distributed dataframe (since the size of data cannot be fit into a single machine), and feed it to tensorflow as features for distributed training. Vineyard aims at sharing the "DataFrame" between these two engines, by memory sharing.

+ feather: as well as arrow, are focused on interoperable data stroage format. feather and arrow define a IPC serailziation schema for common data structures, e.g., tensor and DataFrame. Rather than do serailization for zero-copy (when possible) deserialization, vineyard organized data as a "metadata" and a set of "blobs", the metadata decides how to interpret those blobs and blobs are shared in a zero-copy fashion. Moreover, vineyard could manages large that cannot be fit into a single machine as a "GlobalObject" and global objects can be shared efficiently as well.

+ fst: fst is also a serailization/stroage format, but vineyard is not such a thing.

+ diskframe: diskframe is similar to fst, and works for data that cannot be fit into memory. Vineyard shares in-memory data that cannot be fit into a single machine between different engines (might be implemented in different languages).

+ duckdb: vineyard is not a SQL execution engine.

sighingnow··on Vineyard: An open-source in-memory data manager
> Does distributed mean it’s replicated across machines?

No, in vineyard "distributed" doesn't mean data replicated across machines, rather, big data partitioned across machines. We address cases where the data cannot be fit into a single machine.

> Can this be run outside kubernetes or is it meant to be part of a big set of infra?

Yes, vineyard can run outside of kubernetes, but it would be a bit complicated to setup on many machines as a cluster.

Devops your vineyard cluster on kubernetes could leverage the ability of Kubernetes. We already have integrated with helm and there's a vineyard-operator in our roadmap.

sighingnow··on Vineyard: An open-source in-memory data manager
Let's say someone want to use python first to create a few tensors with numpy, and then use another python package networkx to create a graph from the tensors just created and next to do some graph analysis with networkx.

It is easy to do it on a single machine in a single python runtime because data structures like tensors can be easily shared within a single python runtime.

It will become a bit harder to share the data across processes/Python runtime efficiently without expensive IO or serialization (doable with the help of things like plasma in Apache Arrow).

However, if we want to deal with "big data" that cannot be handled on a single machine, it will become very hard to share data without IO/serialization. Vineyard can be used for such cases. It makes sharing big distributed data structures easy for different runtime.

There will be some added benefits with vineyard: 1. vineyard can handle the the IO, sharding/partitioning, failover and data migrations for the applications.

2. vineyard supports many out-of-the-box highly efficient data formats (such as Apache Arrow) and high-level abstractions which ease the difficulties of developing big data applications.

3. With stream data support, there is a possibility of enabling cross-process optimizations.

4. Fits K8s very well, where common data structures can live in a separated pod/container with its own resource limit.

sighingnow··on Vineyard: An open-source in-memory data manager
Vineyard is a distributed in-memory data manager. Features:

1. distributed in-memory immutable data storage;

2. zero-copy data access through shared-memory;

3. out-of-the-box high-level data abstractions for commonly used data structures;

4. pre-built drivers for I/O, migration, checkpoint, etc;

5. C++ and Python API; and

6. Kubernetes-integration for large-scale big data applications

Github: https://github.com/alibaba/libvineyard (stars are welcomed!)

Documentation: https://v6d.io

Helm charts: https://artifacthub.io/packages/helm/vineyard/vineyard

Any comments and contributions from the community are welcomed!