3,239 karma · joined October 20, 2010
Always happy to meet a stranger, please reach out any time: hn @ peterdowns.com
[ my public key: https://keybase.io/peterldowns; my proof: https://keybase.io/peterldowns/sigs/9N-85LOZH1eJMXHLR70WroivxJB1is_s3ye5IT5xzxs ]
https://github.com/tegonhq/tegon/blob/158b54af8d6f7cf4195c61...
Separately, there seems to be a ton of unused or broken or dead code sprinkled throughout — for instance, in the auth code, I can't tell if you're doing basic email/password auth or using Supertokens and a third-party login via Google. You have code for both and some routes seem dead or missing.
https://github.com/tegonhq/tegon/blob/158b54af8d6f7cf4195c61...
Also, I mentioned the lack of documentation for how to run Tegon locally because your docs are entirely insufficient. The main docs page is just a README template. https://github.com/tegonhq/tegon/tree/main/docs
The quickstart guide has a broken link to instructions on how to self-host https://github.com/tegonhq/tegon/blob/main/docs/quickstart.m...
The oss/local-setup guide is entirely empty https://github.com/tegonhq/tegon/blob/main/docs/oss/local-se...
The oss/deploy-tegon guide does not explain anything and the script it references seems out of date https://github.com/tegonhq/tegon/blob/main/docs/oss/deploy-t...
I'm done looking at this project. I strongly recommend hiring the best engineer you can find as quickly as you can.
Aside — the architecture diagrams on https://xeiaso.net/blog/xesite-v4/ are not displaying in-browser for me, but do transfer correctly and are viewable on my computer. Maybe another mime issue, as the network inspector shows them being transferred with type "octet-stream"?
Repro by visiting the URL in latest Firefox, Safari, or Chrome.
For instance, I have a migrations library that I descibre as "modern" because it is designed for a continuous deployment environment where you're automatically running migrations on container startup — there are tons of existing popular migrations libraries, but none of them work this way because they were written in the era that you'd manually run sql commands in prod. I say "modern" so that if anyone finds my library, they realize that it was created recently based on more recent dev/ops trends.
Maybe I should drop the "modern"? I do see a lot of people describe their code as "minimal" or "clean", which is pretty meaningless to me, so I get that "modern" could come across that way as well.
My condolences to the OP and I was happy to read such a well-written article describing a fantastic outcome for his father's company.
- commenting under a pseudonymous profile
- asking for emails by saying "please email me. contact at cyberscarecrow.com"
- describing yourself in your FAQ entry for "Who are you?" by writing "We are cyber security researchers, living in the UK. We built cyber scarecrow to run on our own computers and decided to share it for others to use it too."
I frequently use pseudonymous profiles for various things but they are NOT a good way to establish trust.
https://pigweed.dev/docs/overview.html
(If you're reading this, hello Keir!)
It's interesting to note that JVector accomplishes this differently than how DiskANN described doing it. My understanding (based on the links below, but I didn't read the full diff in #244) is that JVector will incrementally compress the vectors it is using to construct the index; whereas DiskANN described partitioning the vectors into subsets small enough that indexes can be built in-memory using uncompressed vectors, building those indexes independently, and then merging the results into one larger index.
OP, have you done any quality comparisons between an index built with JVector using the PQ approach (small RAM machine) vs. an index built with JVector using the raw vectors during construction (big RAM machine)? I'd be curious to understand what this technique's impact is on the final search results.
I'd also be interested to know if any other vector stores support building indexes in limited memory using the partition-then-merge approach described by DiskANN.
Finally, it's been a while since I looked at this stuff, so if I mis-wrote or mis-understood please correct me!
- DiskANN: https://dl.acm.org/doi/10.5555/3454287.3455520
- Anisotropic Vector Quantization (PQ Compression): https://arxiv.org/abs/1908.10396
- JVector/#168: How to support building larger-than-memory indexes https://github.com/jbellis/jvector/issues/168
- JVector/#244: Build indexes using compressed vectors https://github.com/jbellis/jvector/pull/244
See https://www.postgresql.org/docs/17/sql-set-constraints.html for more information regarding constraint checking.
> Would love to hear what you think about it!
Also hi joe :^)
Embed the demo video (the tweet you linked) too!
Looks great and if I still used records I would try it out! Congratulations on the launch.
EDIT: keep the rest of the writing if you'd like, I have no objection to it and I thought it was funny. But lead with what the app is and what it does!
- Why? Strongly suggest updating the README to explain why this is useful, or what kind of workflow makes this useful. Just a list of features isn't that compelling to me because I'm not sure I want it.
- Because this is a gh cli extension, and not a standalone program, searching is unfortunately going to be fairly slow. There doesn't seem to be any local caching or syncing of the Github data.
- Not really a user-impacting issue, just a fun fact: because the pagination is handled via Graphql's PageInfo support in their Graphql API, you'll never be able to page through more than 1000 results in a given response.
- Authors might consider linking to https://docs.github.com/en/search-github/searching-on-github in the README, it makes figuring out how to write a custom search a lot easier.
- I'm happy to see XDG_CONFIG_HOME support, but the authors should consider also checking for a config file in $RepoRoot/.config/gh-dash/ and $RepoRoot/.gh-dash/, so that a single config could be more easily shared by members of a team who all check out the same repository.
I'm working on something like this (customizable filter views for Github; quickly drill down into PRs by different authors, affecting different files, across multiple repos, filtered by regex search on the changed files / pr body / pr title / pr comments) but based in the browser, not a TUI. The primary goal of my project is to help engineering leaders understand what is actually getting shipped, and communicate that to other non-technical leaders. If anyone is interested in being a beta tester let me know, I'm hoping to have a release publicly available at the end of this week!
That said, the JS ecosystem is so weird that I totally understand the urge to bail entirely and do things in Go. But JS/TS and all the related frameworks really are decent if you pick a reasonable subset of them.