HNHacker News
TopNewBestAskShowJobs

rajatarya

36 karma · joined January 19, 2020

submissionscomments
rajatarya··on The future of large files in Git is Git
xet is open source, check out https://github.com/huggingface/xet-core.

We (I'm on the xet team at HF) are open sourcing our spec + protocol in the coming months. So, with the spec and protocol open-sourced, anyone can create xet clients and implement the protocol to build a xet backend.

The specific implementation of our xet backend is deeply integrated into HF backend so open-sourcing it directly wouldn't be very helpful. Once we get the spec + protocol released it should be easy to generate a compatible backend.

rajatarya··on Benchmarking Versioning Tools: S3, DVC, Git LFS, and XetHub
re: "modern development experience" - more and more technical projects can be seen as "just" software projects with huge data dependencies. For gaming, the data dependencies are binary texture files, sound files, and more. For biotech these are binary formats from equipment and for analysis. And for ML these are images, video, text (and binary formats for text like Parquet).

One of core motivations behind XetHub is to enable teams across industries to benefit from the workflow we've used in software for 15+ years. We've used this workflow for so long it is easy to overlook its benefits.

Software teams have a clear picture of who is working on what, what is in flight, what is in review, and what is remaining. Anyone on the team can easily pick up work in progress from someone else or start a new derivation of work without concern about interference. Teams can be distributed across timezones and yet everyone feels connected to the project and is able to contribute without disruption.

The power of a GitHub-style workflow for team collaboration comes from being able to experiment freely (branches or forks), review easily (pull requests), and observe (passively learn) best practices from the team (issues, code review feedback).

rajatarya··on Benchmarking Versioning Tools: S3, DVC, Git LFS, and XetHub
(disclaimer: XetHub co-founder here) What other tools should we add to this benchmarking set?

Last year we benchmarked this set along with LakeFS, should we add LakeFS back to this set?

rajatarya··on Show HN: I git commit my home directory every night
Can you share any dedupe results as you've been running this? How long does it take nightly?
rajatarya··on GitHub Monorepo with Meta's Seamless Models and Code
Why do this? What is the benefit of having a monorepo for these models?
rajatarya··on Show HN: Version code, models, & datasets together in GitHub
When you try this out, I'd love to know how much dedupe was possible for your Git LFS files (I'm a co-founder at XetHub).
rajatarya··on NFS > FUSE: Why We Built Our Own NFS Server in Rust
Repo link: https://github.com/xetdata/nfsserve
rajatarya··on Show HN: Gitopia: Decentralized GitHub Alternative for Open Source Collaboration
1st off - like the design touches on the website. And kudos for getting something off the ground: 0->1 is hard!

Is this just another Gogz/Gitea derivative (the Hub repo looked Golang so I am guessing one of these projects)?

Is there something decentralized about the Hub part? From my quick 2m glance I couldn't see anything.

rajatarya··on Show HN: GitEase: Python3 CLI to simplify Git usage with LLM-assisted commits
In my using this over the last couple days (I work with xdss and he asked me to look over the code a couple days ago) gitease works best if you find calling 'git add/commit/push' repeatedly to break your flow with git.

The main benefit is you just call 'ge save' and the tool takes care of calling 'git add/commit (uses LLM for summarizing the diff)'. If you call 'ge share' then also pushes the changes to remote.

Using gitease doesn't limit your usage of git in any way, I think of it as a utility to make it a little easier when first starting out on a project - when you just want to commit and keep going - without thinking too deeply about exactly what changed.

Hope this helps!

rajatarya··on Oxen.ai: Fast Unstructured Data Version Control
Since launching in December we’ve had a few teams use XetHub specifically as a drop-in replacement for Git LFS and their main reasons for switching are: easier to use (no .gitattributes file management), faster performance, and less storage.

Love to have you try XetHub and give us your feedback!

rajatarya··on Oxen.ai: Fast Unstructured Data Version Control
Great to see more people in this space! We are the authors of XetHub (posted in Dec ‘22, ShowHN: https://news.ycombinator.com/item?id=33969908) and also think a git-like workflow is perfect for ML dataset management, except that we actually integrate with git (like LFS). <A quick benchmark suggests we are 2x your published performance!>
rajatarya··on Show HN: We scaled Git to support 1 TB repos
Not yet. Would be happy to try - can you point me to a project to use?

Do you have a repo you could try us out with?

We have tried a couple Unity projects (41% smaller due to republication) but not much from Unreal projects yet.

rajatarya··on Show HN: We scaled Git to support 1 TB repos
Yes, see this for more details of how XetHub deduplication: https://xethub.com/assets/docs/xet-specifics/how-xet-dedupli...
rajatarya··on Show HN: We scaled Git to support 1 TB repos
XetHub Co-founder here. Yes, we use the same Git extension mechanism as Git LFS (clean/smudge filters) and we store pointer files in the git repository. Unlike Git LFS we do block-level deduplication (Git LFS does file-level deduplication) and this can result in a significant savings in storage and bandwidth.

As an example, a Unity game repo reduced in size by 41% using our block-level deduplication vs Git LFS. Raw repo was 48.9GB, Git LFS was 48.2GB, and with XetHub was 28.7GB.

Why do you think using a Git-based solution is a hack compared to p4? What part of the p4 workflow feels more natural to you?

rajatarya··on Show HN: We scaled Git to support 1 TB repos
XetHub Co-founder here. Yes, one illustrative example of the difference is:

Imagine you have a 500MB file (lastmonth.csv) where every day 1MB is changed.

With file-based deduplication every day 500MB will be uploaded, and all clones of the repo will need to download 500MB.

With block-based deduplication, only around the 1MB that changed is uploaded and downloaded.

rajatarya··on Show HN: We scaled Git to support 1 TB repos
XetHub Co-founder here. We are still trying to figure out pricing and would love to understand what sort of pricing tier would work for you.

In general, we are thinking about usage-based pricing (which would include bandwidth and storage) - what are your thoughts for that?

Also, where would you be mounting your repos from? We have local caching options that can greatly reduce the overall bandwidth needed to support data center workloads.