HNHacker News
TopNewBestAskShowJobs

zxt_tzx

186 karma · joined November 25, 2023

submissionscomments
zxt_tzx··on Open sourcing SemHub, a semantic search tool for GitHub (by popular demand)
My previous HN post (https://news.ycombinator.com/item?id=43299659) got some traction, and several folks reached out asking for the source code, so here it is!

Rumor has it a certain dedicated vector database provider (whose name may or may not start with "Q" and end with "rant") might be forking this repo to showcase how you'd build a similar semantic search feature using their offering. Keep an eye out for that! (I'm not affiliated with that effort, so no guarantees.)

Feel free to ask any questions about the code (am particularly pleased with the Easter Egg).

Have a great weekend!

zxt_tzx··on Long Read: Lessons from Building Semantic Search for GitHub and Why I Failed
Thank you for the comment, compared to you I have only touched the bare surface of this quite complex domain, would love to get more of your input!

> building HNSW indices in Postgres is still extremely slow (even with parallel index builds), so it is difficult to experiment with index hyperparameters at scale.

Yes, I experienced this too. I from 1536 to 256 and did not try more values than I'd have liked because spinning up a new database and recreating the embeddings simply took too long. I’m glad it worked well enough for me, but without a quick way to experiment with these hyperparameters, who knows whether I’ve struck the tradeoff at the right place.

Someone on Twitter reached out and pointed out one could quantizing the embeddings to bit vectors and search with hamming distance — supposedly the performance hit is actually very negligible, especially if you add a quick rescore step: https://huggingface.co/blog/embedding-quantization

> But (as mentioned in other comments) keeping your data in sync is a huge issue.

Curious if you have any good solutions in this respect.

> The other challenge I’ve found is that filtering is often the “special sauce” that vector store providers bring to the table, so it’s pretty difficult to reason about the performance and recall of various types of filters.

I realize they market heavily on this, but for open source databases, wouldn't the fact that you can see the source code make it easier to reason about this? or is your point that their implementation here are all custom and require much more specialized knowledge to evaluate effectively?

zxt_tzx··on Long Read: Lessons from Building Semantic Search for GitHub and Why I Failed
I'm glad you found it helpful :)
zxt_tzx··on Long Read: Lessons from Building Semantic Search for GitHub and Why I Failed
> And overall, the fewer people use CF (or another provider of their size) the better.

I understand your sentiment, but I vehemently disagree.

The cloud provider space has rapidly become an oligopoly and CloudFlare is one of the few new entrants that (1) has sufficient scale to compete with the incumbents; (2) has new ideas that the incumbents cannot easily match (region earth, durable objects etc.).

For most production workloads, I would not even consider the newer cloud providers, but I sincerely meant it when I said I hope Cloudflare will succeed. They've also been very responsive to the feedback raised in the blogpost when I DM-ed them.

(On a side note re: difficulty for newcomers in this market, I used to be part of a team that would run e.g. staging and testing environments on a new serverless db provider, but would run prod on AWS Aurora. In retrospect, this did not make much sense either as you want your environments to be as similar as possible, which means new cloud providers have an even tougher time getting started.)

zxt_tzx··on Long Read: Lessons from Building Semantic Search for GitHub and Why I Failed
After this failed experience with SemHub, I am actually thinking of building something like this, for open source maintainers like you are definitely the ICP! (nuqs seems really cool btw, storing state in the URL param is definitely the way to go)

To elaborate, I was thinking of:

- running a cron that checks repos every X minutes

- for every new issue someone has opened, I will run an agent that (1) checks e.g. SemHub to look for similar issues; (2) checks the project's Discord server or Slack channel to see if anyone has raised something similar; (3) run a general search

- use LLMs to compose a helpful reply pointing the OP to that other issue/Discord discussion etc.

From other OSS maintainers, I've heard that being able to reliably identify duplicates would be a huge plus. Does this sound like something you'd be interested to try? Let me know how I can reach you if/when I have built something like this!

I am personally quite annoyed by all the AI slop being created on social media and even GitHub PRs and would love to use the same technology to do something pro-social.

zxt_tzx··on Long Read: Lessons from Building Semantic Search for GitHub and Why I Failed
Thanks for sharing! Do you have more details to share, e.g. did you just have a vector db, or did you have a main db as well?

In my research, Qdrant was also the top contender and I even created an account with them, but the need to sync two dbs put me off

zxt_tzx··on Long Read: Lessons from Building Semantic Search for GitHub and Why I Failed
Ah I was doing semantic search of GitHub _issues_, not the actual code on GitHub.

For code search, I have used grep.app, which works reasonably well

zxt_tzx··on Long Read: Lessons from Building Semantic Search for GitHub and Why I Failed
> There’s so much complexity that comes with keeping your vector db in sync with you main db (especially once you start filtering with metadata)

Ohh do you speak from experience? I know I will likely never do this, but curious how did you do it? When I looked into this, I found that Airbyte has something to connect the vector db with the main db, but I never bit that bullet (thankfully)

zxt_tzx··on Long Read: Lessons from Building Semantic Search for GitHub and Why I Failed
oh wow that's super cool, I tried it and it's very fast indeed. thanks for sharing! will spend more time to understand how it's implemented
zxt_tzx··on Long Read: Lessons from Building Semantic Search for GitHub and Why I Failed
Thanks for the feedback, to be honest, my own experience is actually very similar to yours.

The original pain point probably only exists for small minority of open source maintainers who manage multiple repos and actually search across them regularly. Most devs are probably like you and I, and the mediocre GitHub search experience is more than compensated by using Google.

In its current iteration, it's quite hard to get regular devs to change their searching behaviour, and, even for those who experience this pain point, it probably isn't large enough for them to change their behavior.

If I continue to work on this, I would want to (1) solve a bigger + more frequent pain point; (2) build something that requires a smaller change in user behavior.

zxt_tzx··on Long Read: Lessons from Building Semantic Search for GitHub and Why I Failed
> Have you looked into chunking (breaking input into smaller chunks and doing vector search on the chunks)?

Ohh I had not seriously considered this until reading this. I could have multiple embeddings per issue and search across those embeddings and if the same issue is matched multiple times, I would probably take the strongest match and dedupe it.

I could create embeddings for comments too and search across those.

Thanks for the suggestion, would be a good think to try!

> Choosing a chunking strategy seems to be a deep rabbit hole of its own.

Yes this is true. In my case, I think the metadata fields like Title and Labels are probably doing a lot of the work (which would be duplicated across chunks?) and, within an issue body, off the top of my head, I can't see any intuitive ways to chunk it.

I have heard that for standard RAG, chunking goes a surprisingly long way!

zxt_tzx··on Long Read: Lessons from Building Semantic Search for GitHub and Why I Failed
> it's weird you consider this a failure. you spent a few months and learned how to work with embedding models to build an efficient search. the fact that your search works well is a successful outcome.

Thank you for your encouragement! I take your point that it was not a technical failure, but I think it's still a product failure in the sense that SemHub was not solving a big enough pain point for sufficiently many people.

> if you want to turn your search into a business now that's a new and different effort, mostly marketing and stuff that most self respecting engineers gives zero shits about, but if that's your real goal don't call it a failure yet because you haven't even tried.

Haha to be honest, my goal was even more modest, SemHub is intended to be a free tool for people to use, we don't intend to monetize it. I also did try to market it (DMing people, Show HN), but the initial users who tried it did not stick around.

Sure, I could've marketed SemHub more, but I think the best ideas carry within themselves a certain virality and I don't think this is it.

zxt_tzx··on Long Read: Lessons from Building Semantic Search for GitHub and Why I Failed
Ohh apologies, I think there was a bug that led to the Internal Server Error, please try again, I _think_ it should be working now!

> I think a project like yours is going to be helpful to OSS library maintainers to see which features are used in downstream projects and which have issues.

That was indeed the original motivation! Will see if I can convince Ammar to reconsider shutting down the project, but no promises

> For this use case, I deployed my own instance to index all OSS repos implementing OSLC REST or using our Lyo SDK

Ohh, in case it's not clear from the UI, you could create an account and index your own "collection" of repos and search from within that interface. I had originally wanted to build out this "collection" concept a lot more (e.g. mixing private and public repos), but I thought it was more important to see if there's traction for the public search idea at all

zxt_tzx··on Long Read: Lessons from Building Semantic Search for GitHub and Why I Failed
Totally fair point. Thanks for taking the time to read through it! I guess I didn't want to use a VPS and then have to switch to something else if the product really worked, but I guess that rhymes with premature optimization.

Some other clarifications:

- I was also surprised with how expensive Supabase turned out to be and only got there because I was trying to sync very big repos ahead of time. I could see an alternative product where the cost here would be minimal too

- I did see this project as an opportunity to try out Cloudflare. as mentioned in the post, as a full stack TypeScript developer, I thought Cloudflare could be a good fit and I still really want it to succeed as a cloud platform

- deploying two separate API server and auth server is actually simpler than it sounds, since each is a Cloudflare Worker! will try to open source this project so this is clearer

- the durable objects rate limiter was wholly experimental and didn't make it into production

> All that being said though, maybe all it would've done is prolong the inevitable death due to the product gap the author concludes with.

Very true :(

zxt_tzx··on Long Read: Lessons from Building Semantic Search for GitHub and Why I Failed
Author here. Over the last few months, I have built and launched a free semantic search tool for GitHub called SemHub (https://semhub.dev/). In this blog post, I share what I’ve learned and why I’ve failed, so that other builders can learn from my experience. This blog post runs long and I have sign-posted each section. I have marked the sections that I consider the particularly insightful with an asterisk (*).

I have also summarized my key lessons here:

1. Default to pgvector, avoid premature optimization.

2. You probably can get away with shorter embeddings if you’re using Matryoshka embedding models.

3. Filtering with vector search may be harder than you expect.

4. If you love full stack TypeScript and use AWS, you’ll love SST. One day, I wish I can recommend Cloudflare in equally strong terms too.

5. Building is only half the battle. You have to solve a big enough problem and meet your users where they’re at.

zxt_tzx··on Managing Oneself (2005)
It's always a little dubious when modern people pretend to have high confidence about the behaviors of long-dead people to serve their modern purposes. (Another example: oh you're an INFJ, just like Moses from the Bible!)
zxt_tzx··on We improved the performance of a userspace TCP stack in Go
I met one of the founders of Coder.com, he's a really cool dude. It's a pity that it is a product aimed more at enterprises than individual developers, else it would have far more developer mindshare.

Unlike, say, GitHub Codespaces, running something like this on your own infra means your incentives and Coder.com's are aligned, i.e. both of you want to reduce your cloud costs (as opposed to, say, GitHub running on Azure gives them an opportunity and incentive to mark up on Azure cloud costs).

zxt_tzx··on In Response to Google
Mildly interesting that the author runs a media relations company in SF.

On the one hand, I am sympathetic to the general perspective of the original article.

On the other hand, that the same person is writing “hit pieces” and running a media relations company bears a certain resemblance to protection rackets of yore

zxt_tzx··on MemoryDB: A fast and durable memory-first cloud database
Interesting stuff. We use MemoryDB as the underlying service for BullMQ, a NodeJS queue that’s built on top of Redis. We trade off a bit of speed and cost (MemoryDB costs more than Elasticache) for persistence and BullMQ’s many features, which is a good tradeoff for most apps.
zxt_tzx··on Ask HN: What apps have you created for your own use?
I was inspired by Hey.com’s screener feature but I didn’t want to move my existing Gmail accounts, so I created one for myself: https://app.inboxhero.org/

The idea is first-time senders will be moved out of your inbox and you will get a screener to manually whitelist/blacklist them.

Am now trying to productize it, my thesis is emails will becoming even spammier now that LLMs can pass the Turing test, so a feature like this will be key to restore sanity to your inbox.

zxt_tzx··on Ask HN: Can we do better than Git for version control?
Hey, I have only just started to use it. My colleagues swear by it, which is how I've found out about it in the first place.

I think there's no magic and we'll still have to resolve merge conflicts on our own, but my sense is it does simplify repetitive operations.

Hope this helps!

zxt_tzx··on Ask HN: Can we do better than Git for version control?
I think Git itself is probably too entrenched to be displaced by now, but I recently came across Graphite (https://graphite.dev/) and, while it’s all still Git under the hood, it abstracts away many of the common pain points (stacking PRs, rebasing) and has nice integrations with GitHub and VS Code.
zxt_tzx··on If buying isn't owning, piracy isn't stealing
Instead of getting into a shouting match over whether piracy is or isn’t stealing, here’s an interesting story.

Bill Gates made an interesting comment regarding software piracy in the 1990s. At that time, software piracy was rampant in China, and Microsoft’s Windows operating system was one of the most pirated products. In response to this situation, Gates remarked that if people were going to pirate software, he would prefer they pirate Microsoft’s software.

His rationale was strategic: if people in China became accustomed to using Microsoft’s Windows and Office software, they would be more likely to continue using these products in the future, including in business environments where legitimate software licenses are more commonly purchased.

(Whether things worked out as he envisioned is a slightly different matter.)

zxt_tzx··on Apple cuts off Beeper Mini's access
”It is difficult to get a man to understand something when his salary depends upon his not understanding it.” Upton Sinclair
zxt_tzx··on 'Greedflation' study finds many companies were lying to you about inflation
> While this obviously contributed to rising prices, the report finds that company profits increased at a much faster rate than costs did, in a process often dubbed “greedflation.”

I think there is this expectation that companies "should" only raise prices necessary to cover their cost increases. But companies charge what the market can bear and, depending on the specific circumstances, this could be equal to, less than, or more than the cost increases.

To use a HN-friendly example, let's say AWS cuts their EC2 costs. If you're providing an undifferentiated computation service using EC2, you'll likely have to cut your costs or you will lose all your business to your competitors. If you're providing a highly differentiated SaaS product, you can probably keep the cost savings to yourself.

You can reason in the opposite direction for price hikes. Depending on the exact company in question, they could be in a privileged position of the supply chain or have marketing power that allows them to increase prices even more.

One main difference here is price stickiness [1], i.e. prices tend to be fixed for a period of time, even if the underlying economics has changed. I think this is underlying reason for perceptions of "greedflation" because, during period of inflation, prices become less sticky and companies use this to adjust prices to better match the underlying economics.

[1] https://en.wikipedia.org/wiki/Nominal_rigidity

zxt_tzx··on 'Greedflation' study finds many companies were lying to you about inflation
I would frame it slightly differently: people (and by extension, these companies) have always been greedy, so what is the new variable that actually enabled them to convert this greed into higher prices?
zxt_tzx··on Apple cuts off Beeper Mini's access
Are you referring to the step where Beeper's servers make a persistent connection with Apple's APN service to listen to new messages ?

So your point is Apple can presumably distinguish between an actual iOS connection and Beeper's connection by looking at "how many connections per IP"? Still seems prone to false positives to me, unless there is something else I missed.

(Upon re-reading the post, I realized that the phone number registration is actually done by Apple. Wonder if this might provide another basis to block Beeper, i.e. all this SMS infrastructure is not cheap to maintain and Beeper's integration is arguably using it in an "unauthorized" way.)

zxt_tzx··on IA Writer in Paper
I am not much of a pen-and-paper kinda person, but I recently came across CPG Grey's Sidekick Notepad that made me think twice.

I like how it's designed to be integrated into a digital workflow and the "signaling" aspect of opening and closing the notepad is a nice touch too.

YouTube: https://www.youtube.com/watch?v=z_7N8MFRJkc

Website: https://cottonbureau.com/p/XT9MRF/journal/sidekick-notepad

zxt_tzx··on Apple cuts off Beeper Mini's access
> just the IPs Beeper Mini was using to connect to the APN service.

Hmm, wouldn't blocking IPs be overly broad and risked affecting regular users? Considering that IPs are scarce and constantly recycled by ISPs etc. Blocking device identifiers sounds more targeted and, for that reason, realistic.

zxt_tzx··on Apple cuts off Beeper Mini's access
> We've really done one over on ourselves by adopting the mental model that only a vertically integrated corp can deliver privacy and security to users. This rigid tendency towards homogeneity is bound to suffer a tragic systemic failure before too long.

Look no further than the other news that came out this week re: government spying via push notifications. (https://www.reuters.com/technology/cybersecurity/governments...) Consumers rationally trust the few big companies which are incentive-aligned to protect their data and government then goes after those few big companies. I thought this was particularly galling:

> In a statement, Apple said that Wyden's letter gave them the opening they needed to share more details with the public about how governments monitored push notifications.

> "In this case, the federal government prohibited us from sharing any information," the company said in a statement. "Now that this method has become public we are updating our transparency reporting to detail these kinds of requests."

Page 1 of 3Next →