HNHacker News
TopNewBestAskShowJobs

sluongng

163 karma · joined May 20, 2022

submissionscomments
sluongng··on AI coding has made CI a bottleneck, so we reworked ours to keep up
Yes. Bazel is half of the equation. The other half which other enterprises rely on is the server side of Bazel's Remote Build protocol. There, the scheduler can be implemented with logic that accounts for sandbox sizing as well as other requirements/constraints.

I gave a talk about how we do it at BuildBuddy at a recent BazelCon here: https://youtu.be/iQqLtuBzkKE?t=848

sluongng··on AI coding has made CI a bottleneck, so we reworked ours to keep up
The Bazel ecosystem builds container images not by using Dockerfile, which contains non-reproducible primitives such as RUN and others. We do it by actually constructing the file trees and tarballs manually, then using them to compose the JSON manifest and indices. This is done via smaller hermetic and reproducible Bazel actions and thus enables the ecosystem to scale way beyond what alternative BuildKit-based solutions can.

https://www.youtube.com/watch?v=biYXmAv4Ppk&t=314s should be a good talk to study up on the matter. The speaker is now working at Apple.

sluongng··on AI coding has made CI a bottleneck, so we reworked ours to keep up
What i have seen on my end is that it’s pretty easy to setup a cloud coding agent with a warm Bazel cache. “Warm” here can means multiple layers: same disk to CoW/hardlink, different disk, network disk/block devices, same rack/datacenter, same AZ, etc…

And yes, there are a ton of investments going toward Bazel recently to unlock these newer use cases.

sluongng··on AI coding has made CI a bottleneck, so we reworked ours to keep up
We solve it 2 ways in the Bazel ecosystem: for the intermediate artifacts, we only fetch the digest (hash + size) of the blobs to calculate the merkle tree forward. The blob itself can stay on the remote cache server.

For the bigger final artifacts, we support using Content Defined Chunking (rolling gear hashing) to only fetch the missing chunks between incremental builds. Binaries executable with stable layout benefits from this quite a lot.

We are definitely not done with all of the improvements here. But since all the major AI labs are using Bazel, we know that the tools can support “Agent Scale”. https://webazel.dev/

sluongng··on Should you wash your solar panels?
isn't that more of a storage issue, which a home battery would solve?
sluongng··on Demis Hassabis has a plan to harness AI safely
how would this help smaller labs? would it put more burdens on them when trying to compete with trillion-dollar companies or would it help?
sluongng··on Google copybara: moving code between repositories
Yeah I vibe coded https://github.com/sluongng/capyfun during a hackathon recently to add a generative transformation layer on top of the traditional imperative transformations.
sluongng··on Cutting inference cold starts by 40x with LP, FUSE, C/R, and CUDA-checkpoint
There are plenty of cool advancements in reducing inference cold start when I was meeting with folks in person at FOSDEM this year. However, I still struggle to understand: why would folks care about this?

Major AI Labs all have secured their own compute in the form of hardware, data center, and power generation. That means their resource pool is fixed, and they can do all sorts of tricks to pre-load, pre-allocate, etc... to improve on inference latency.

Cold start is usually a solution for "cloud" environment when your pool is flexible, and you only pay for what you use. Its effectiveness lowered in bare-metal settings as folks do not care about scaling up and down as much.

So my question is: who is this for? AWS and GCP running Anthropic models?

sluongng··on Show HN: Git bayesect – Bayesian Git bisection for non-deterministic bugs
You can run bisect with first-parent
sluongng··on Ninja is a small build system with a focus on speed
My teammate has a great time reimplementing Ninja (slop-free) in Go here https://github.com/buildbuddy-io/reninja to make it even faster with Remote Build Execution.
sluongng··on Parallel coding agents with tmux and Markdown specs
Yeah, I don't disagree with your assessment at all. I think the H2A ratio is still a good metric for the AI adoption rate of an organization. At a higher H2A ratio, you will also start to hear people measuring things using token volumes, which I think is also a similar metric (because most models nowadays run on a relatively fixed Tokens/second speed).

All of this is not a direct signal to a productivity boost. I think at higher volumes, you will need to start to account for the "yield" rate of the token volumes above: what are the volumes of tokens that get to the final production deployment? At which stage is it a constraint on the yield? Is it the models, or is it the harness, or something else (i.e. Code Review, CI/CD, Security Scans etc...)? And then it becomes an optimization problem to reduce the Cost of Goods Sold while improving/maintaining Revenues. The "productivity" will then be dissolved into multiple separate but more tangible metrics.

sluongng··on Parallel coding agents with tmux and Markdown specs
I do. The reason why the current generation of agents are good at coding is because the labs have sufficient time and computes to generate synthetic chain-of-thoughts data, feed those data through RL before use them to train the LLMs. These distillation takes time, time which starts from the release of the previous generation of models.

So we are just now getting agents which can reliably loop themselves for medium size tasks. This generation opens a new door towards agent-managing-agents chain of thoughts data. I think we would only get multi-agents with high reliability sometimes by the mid to end of 2026, assuming no major geopolitical disruption.

sluongng··on Parallel coding agents with tmux and Markdown specs
Yeah the 8 agents limit aligns well with my conversations with folks in the leading labs

https://open.substack.com/pub/sluongng/p/stages-of-coding-ag...

I think we need much different toolings to go beyond 1 human - 10 agents ratio. And much much different tooling to achieve a higher ratio than that

sluongng··on Cracking the Python Monorepo
Most of the time, the CI resources in a python monorepo is not spent on packaging. It’s spent on running the tests.

I would love to read more about how the author is tackling the testing problem in their setup.

sluongng··on Move tests to closed source repo
https://sluongng.substack.com/i/186718212/test-is-king I wrote about this less than a month ago. Things are moving pretty fast in this direction.
sluongng··on Show HN: I ported Tree-sitter to Go
Oh this is really neat for the Bazel community, as depending on tree-sitter to build a gazelle language extension, with Gazelle written in Go, requires you to use CGO.

Now perhaps we can get rid of the CGO dependency and make it pure Go instead. I have pinged some folks to take a look at it.

sluongng··on Putting Gemini to Work in Chrome
Not yet in Linux?
sluongng··on Rust at Scale: An Added Layer of Security for WhatsApp
I suspect they just use no_std whenever its applicable

https://github.com/facebook/buck2/commit/4a1ccdd36e0de0b69ee...

https://github.com/facebook/buck2/commit/bee72b29bc9b67b59ba...

Turn out if you have strong control over the compiler and linker instrumentations, there are a lot of ways to optimize binary size

sluongng··on I made my own Git
Zstd dictionary compression is essentially how Meta's Mercurial fork (Sapling VCS) stores blobs https://sapling-scm.com/docs/dev/internals/zstdelta. The source code is available in GitHub if folks want to study the tradeoffs vs git delta-compressed packfiles.

I think theoratically, Git delta-compression is still a lot more optimized for smaller repos. But for bigger repos where sharding storaged is required, path-based delta dictionary compression does much better. Git recently (in the last 1 year) got something called "path-walk" which is fairly similar though.

sluongng··on Transfering Files with gRPC
The evolving schema is much more attractive than a bunch of plain text HTTP headers when you want to communicate additional metadata with the file download/upload.

For example, there are common metadata such as the digest (hash) of the blob, the compression algorithm, the base compression dictionary, whether Reed-Solomon is applicable or not, etc...

And like others have pointed out, having existing grpc infrastructure in place definitely helps using it a lot easier.

But yeah, it's a tradeoff.

sluongng··on Transfering Files with gRPC
https://github.com/googleapis/googleapis/blob/master/google/... is a more complete version of this. It supports resumable uploads, and the download can start from an offset within a file, allowing you to download only part of the file instead of the whole.

Another version of this is to use grpc to communicate the "metadata" of a download file, and then "side" load the file using a side channel with http (or some other light-weight copy methods). Gitlab uses this to transfer Git packfiles and serve git fetch requests iirc https://gitlab.com/gitlab-org/gitaly/-/blob/master/doc/sidec...

sluongng··on A faster path to container images in Bazel
The underlying problem is that most container images are not cache efficient. Compressed tarballs arent and that’s what most of container images are. And Bazel relies heavily on caching to stay fast.

Most of the hyper scaler actually do not store container images as tarballs at scale. They usually flatten the layers and either cache the entire file system merkle tree, or breaking it down to even smaller blocks to cache them efficiently. See Alibaba Firefly Nydus, AWS Firecracker, etc… There is also various different forms of snapshotters that can lazily materialize the layers like estargz, soci, nix, etc… but none of them are widely adopted.

sluongng··on Fast trigram based code search
> They use Google's web indexing technology adapted for trigrams, which was mostly developed to support their massive internal monorepo

Do you have a source for this? I would love to read more about it.

In the doc of the Zoekt repo, it says

> What does cs.bazel.build run on?

> Currently, it runs on a single Google Cloud VM with 16 vCPUs, 60G RAM and an attached physical SSD.

https://github.com/sourcegraph/zoekt/blob/main/doc/faq.md#wh...

so at least they were using Zoekt up until a certain point in the past.

sluongng··on Introducing architecture variants
Nice. This is one of the main reasons why I picked CachyOS recently. Now I can fallback to Ubuntu if CachyOS gets me stuck somewhere.
sluongng··on Modern CI is too complex and misdirected (2021)
The most concerning part about modern CI to me is how most of it is running on GitHub Actions, and how GitHub itself has been deprioritizing GitHub Actions maintenance and improvements over AI features.

Seriously, take a look at their pinned repo: https://github.com/actions/starter-workflows

> Thank you for your interest in this GitHub repo, however, right now we are not taking contributions.

> We continue to focus our resources on strategic areas that help our customers be successful while making developers' lives easier. While GitHub Actions remains a key part of this vision, we are allocating resources towards other areas of Actions and are not taking contributions to this repository at this time.

sluongng··on Claude says “You're absolutely right!” about everything
I don't view it as a bug. It's a personality trait of the model that made "user steering" much easier, thus helping the model to handle a wider range of tasks.

I also think that there will be no "perfect" personality out there. There will always be folks who view some traits as annoying icks. So, some level of RL-based personality customization down the line will be a must.

sluongng··on Show HN: Linux CLI tool to provide mutex locks for long running bash ops
Why not https://man7.org/linux/man-pages/man2/flock.2.html?
sluongng··on Netflix’s Media Production Suite
More on Netflix's Remote Workstation setup for artists https://aws.amazon.com/solutions/case-studies/netflix-workst...
sluongng··on Why Apple's Severance gets edited over remote desktop software
I think this is a common practice by now.

Here is a talk from Netflix about cloud workspace for their artists https://aws.amazon.com/solutions/case-studies/netflix-workst...

sluongng··on There isn't much point to HTTP/2 past the load balancer
Hmm it’s weird that this submission and comments are being shown to me as “hours ago” while they are all 2 days old
Page 1 of 3Next →