HNHacker News
TopNewBestAskShowJobs

moab

838 karma · joined June 14, 2014

fun with algorithms and data structures
submissionscomments
moab··on Diets high in advanced glycation end products promote insulin resistance
Coffee is very low in AGEs as per Table 2 of https://www.jandonline.org/article/S0002-8223(10)00238-5/ful...
moab··on Databases in 2024: A Year in Review
FWIW, as someone who had multiple friends who worked closely with Andy over 5+ years and at all stages of their career (both BS / PhD) those comments reek of someone with an axe to grind. All of the many anecdotes I have about Andy paint a picture of a great advisor and mentor. I suppose I should say "shenanigans aside", but if you can't separate his jokes from his academic side you need to develop a sense of humor.
moab··on Immutable Data Structures in Qdrant
> bog standard FP community lies > constrain mutable structures to the same memory layouts and ideas as immutable ones

What are you even talking about, "dude"? I don't think you have the background or knowledge you seem to think you have to argue about this space. It's OK, as you blithely pointed out in your earlier post, there is a place called medium where you'll find likeminded folks that will eat up your drivel.

moab··on Immutable Data Structures in Qdrant
Your analysis sounds reasonable to the non-expert, but recent work on purely-functional trees suggests that the gap is smaller than you suggest ("orders of magnitude slower").

E.g., see the nice work on the PAM library (https://arxiv.org/abs/1612.05665). Ideas from this work were used to build lots of cool things (immutable graph data structures, segment trees, databases) that are very fast, and all immutable.

moab··on Eating Foods with Xylitol Can Be a Risk to Your Heart
Unfortunately none of these news articles or papers discuss whether consuming ~2--3g of xylitol/day in the form of gum (e.g., Pur) is OK. I'm guessing it's fine, given that the warning is for products that are artificially sweetned with xylitol with a lot more xylitol, and they say that toothpaste/mouthwash with xylitol should be OK.
moab··on GPT-4o
I misunderstood; my apologies.
moab··on GPT-4o
Pretty sure the snark is unnecessary.
moab··on Build a simple LSM-Tree storage engine in a week
Big kudos to Alex. I hope he continues towards his PhD; these tutorials he's been putting out are fantastic.
moab··on Benjamín Labatut Will Not Be Profiled
Thanks for sharing your thoughts---I'll give the Labatut book a try but disappointing to hear that it pales in comparison to the real history masterfully told by Rhodes...
moab··on Benjamín Labatut Will Not Be Profiled
Thanks for sharing this. On a related theme I highly recommend "the making of the atomic bomb" (this is a narrative history, not an Ikiru-like summoning of a life by different narrators).
moab··on A new quantum algorithm for classical mechanics with an exponential speedup
It would be a good idea to read the blog post before making comments.

The result looks very interesting, and the blog post is well written (e.g., I did not know about the prior work re. Grover's algorithm and pendulum systems).

The blog post is also based on a recent FOCS paper, and the authors are reputable people in CS theory, if that convinces anyone to take a closer look.

moab··on NYT interviews former OpenAI CEO Sam Altman (11/20/23)
I experienced this recently as well on archive.is / archive.ph.
moab··on Bandcamp has been sold to Songtradr. What does this mean for the musicians?
I'm glad you found it useful :-)

The other useful resource I will mention is "soulseek" (https://www.slsknet.org/news/) which is basically old-school direct downloads, and includes a lot of obscure music that you won't find on AppleMusic/Spotify.

Happy trails!

moab··on Bandcamp has been sold to Songtradr. What does this mean for the musicians?
Check out https://rateyourmusic.com/ which is good for discovery, e.g., https://rateyourmusic.com/genre/atmospheric-black-metal/ and the all-time lists: https://rateyourmusic.com/charts/top/album/all-time/g:atmosp...

Users also make wonderful charts on this website. also, hello from a fellow what refugee!

moab··on I’m a climate scientist and I think society will collapse by 2050. I’m preparing
What a ridiculous article. One of the temperature predictions hinges on us burning all of the fossil fuels. Poor article quality to see on HN.
moab··on An early look at HNSW performance with pgvector
Thanks for the response. I wonder whether HNSW will still perform well if it needs to page neighbor-lists to/from disk. Do you plan to benchmark the setting where the dataset is too large to fit in-memory?
moab··on An early look at HNSW performance with pgvector
See https://big-ann-benchmarks.com/neurips21.html

They're not OpenAI embeddings, but they are realistic, and much larger (number of vectors).

I think many production embeddings at non-OpenAI companies will use lower-dimensional vectors than 1536, so it makes sense to focus on non-OpenAI embeddings as well in your benchmarking.

moab··on An early look at HNSW performance with pgvector
Glad to see more work on pgvector but why test on such small datasets on a large memory machine? The big ann datasets have 1B points and are much more interesting/representative of current embedding use cases (eg from dual encoder models).

I’m also curious if there is a way to not store everything in memory for pgvector. Is that possible?

Lastly, what is the parallelism story? Is it just using a thread pool under the hood? OpenMP?

Understanding if pgvector plans to support point insertions and deletions is also important in practice.

moab··on Investors are happy to pay premium for tech, but not for AI
> LLM-enabled senior engineers are 100X more efficient and safe than brand new junior devs.

Come on man. Having seen the inside of big-tech-TM and the senior engineers there, yes they are fast and good, but they are not 100x better than the new guy. Maybe 3--5x at best.

Anyway how do you train good senior engineers? They don't just pop up out of thin air.

moab··on USearch: Smaller and faster single-file vector search engine
Do you have plans to support metadata filtering?
moab··on Nvidia DGX GH200 Whitepaper
Unfortunate that they don't mention the running times for any of the applications they benchmark (e.g., PageRank). Does anyone in the know have some idea how long this takes?
moab··on Death of the physical library
A bit pessimistic no? The interested always find a way, and it has never been easier with the online databases that the author acknowledges as being extremely useful (sci-hub, libgen, etc). What he is lamenting is the loss of serendipity from browsing a physical library shelf; but it is not at all clear that this experience is incompatible with virtual libraries.
moab··on Stanford president resigns over manipulated research, will retract 3 papers
Well not really, right? Let's suppose some well known, well respected author that has a history of correct results puts up a new paper. I (and I think most people) will assume that the result is correct. We start to apply more doubt once the claimed result is a solution to a longstanding open problem, or importantly, if the researcher has a spotty track record for correctness (in math/TCS) or falsifying results (in experimental fields).

But really we shouldn't be talking about math errors and falsification in the same category.

moab··on Stanford president resigns over manipulated research, will retract 3 papers
I completely agree. It's a pity that this isn't becoming standard in fields affected by the replication crisis. I would be happy to be corrected if someone has heard / experienced otherwise.
moab··on Stanford president resigns over manipulated research, will retract 3 papers
This is an issue at the department politics level. For the scientific field, once someone starts retracting papers (and arguably, even before this), everybody knows that you should take person X's papers with a huge grain of salt.

E.g., in math / theory, if someone has a history of making big blunders that invalidate their results, you will be very hesitant to accept results from a new paper they put on arXiv until your community has vetted the result.

So yes, I do trumpet science as a model of a rational, self-correcting social enterprise, at least in CS.

Other sciences like biology and psychology have some way to go.

moab··on Stanford president resigns over manipulated research, will retract 3 papers
So just because one person is cheating, it means all academics are cheating?

FWIW, most top-ranked CS conferences have an artifact evaluation track, and it doesn't look good if you submit an experimental paper and don't go through the artifact evaluation process. Things are certainly changing in CS, at least on the experimental side.

It's also possible that theorems are incorrect, but subsequent work that figures this out will comment on it and fix it.

The scientific record is self-correcting, and fraud / bullshit does get caught out.

moab··on PhD Simulator
You had a terrible advisor. I'm sorry.
moab··on GPT-4’s secret
https://ai.googleblog.com/2022/05/language-models-perform-re...

You can google for more results / papers.

moab··on Vector support in PostgreSQL services to power AI-enabled applications
Not sure why you're being downvoted. IIUC, pgvector uses a clustering implementation similar to ones implemented by FAISS. These are pretty simple / straightforward to implement but do not give the best performance. For more on the current SOTA, which are primarily graph-based algorithms like HNSW and Vamana I would check out https://big-ann-benchmarks.com/
moab··on AMD EPYC 97x4 “Bergamo” CPUs: 128 Zen 4c CPU Cores for Servers, Shipping Now
Work stealing makes it arguably much easier to program than CUDA or OpenCL.
← PreviousPage 2 of 8Next →