HNHacker News
TopNewBestAskShowJobs

lsb

4,183 karma · joined August 22, 2007

classicist, computer scientist, founder of PoetaExMachina and NoDictionaries, biker around San Francisco

leebutterman@gmail.com

submissionscomments
lsb··on Fast B-Trees
Clojure, for example, uses Hash Array Mapped Tries as its associative data structure, and those work well
lsb··on 48% of NYC riders do not pay the bus fare
All New Yorkers pay 80% of the cost to run the bus before they even think about boarding. (“Farebox recovery ratio”)

This is an excuse to fund more cops. Transit should be free, like sidewalks and parks.

lsb··on Ethically Sourced Lena Picture
The history is still there, in archives, and modern usage is what we’re talking about
lsb··on Ethically Sourced Lena Picture
The model rescinded her consent to be used as a technical benchmark photo, a while ago
lsb··on TPU transformation: A look back at 10 years of our AI-specialized chips
Doesn’t Google Cloud’s AI infrastructure include Colab? That’s useful for so many things
lsb··on The six dumbest ideas in computer security (2005)
Default permit, enumerating badness, penetrate and patch, hacking is cool, educating users, action is better than inaction
lsb··on Don’t try to sanitize input, escape output (2020)
Of the “six famous bad ideas in computer security”, the first and second are “default permit” and “enumerating badness”.

http://www.ranum.com/security/computer_security/editorials/d...

lsb··on Daylight eInk Computer
Dasung has eink HDMI monitors that have quite a high refresh rate. For a 25” monitor it runs about $1700. https://shop.dasung.com/
lsb··on Modos Paper Monitor Pre-Launch on Crowd Supply
Looks similar to the Dasung monitor latency. Would be interested to see a comparison, both for features and performance and price
lsb··on $64B Gamble: SoftBank Arm Plan to Launch AI Chip in 2025
Exactly, like how a “bitcoin mining” chip would implement the SHA in hardware.

And, CPUs prioritize hiding latency with all sorts of caches, and GPUs prioritize cores and bandwidth to hide latency, so there’s different tradeoffs about memory bandwidth versus latency.

lsb··on $64B Gamble: SoftBank Arm Plan to Launch AI Chip in 2025
The earliest Intel chips (8086) had an extra add-on to do floating point math in hardware, instead of software (8087). There’s a lot of “AI” workloads that include a lot of matrix multiplication, which the Google TPU and the Nvidia tensor cores implement in hardware instead of software.
lsb··on I’m writing a new vector search SQLite Extension
Would this implement indexing strategies like HNSW? Linear scan is obviously great to start out with, and can definitely be performant, especially if your data is in a reasonable order and under (say) 10MB, so this shouldn't block a beta release.

Also do you build with sqlite-httpvfs? This would go together great https://github.com/phiresky/sql.js-httpvfs

lsb··on I’m writing a new vector search SQLite Extension
It depends on your data and your embedding model. For example, I was able to quantize embeddings of English Wikipedia from 384-dimensions down to 48 7-bit dimensions, and the search works great: https://www.leebutterman.com/2023/06/01/offline-realtime-emb...
lsb··on I’m writing a new vector search SQLite Extension
Binary, quantize each dimension to +1 or -1

You can try out binary vectors, in comparison to quantize every pair of vectors to one of four values, and a lot more, by using a FAISS index on your data, and using Product Quantization (like PQ768x1 for binary features in this case) https://github.com/facebookresearch/faiss/wiki/The-index-fac...

lsb··on The Elements of Differentiable Programming
Yes, these are “adversarial patches” in image classification, like https://arxiv.org/abs/1712.09665 . Similarly you can take these adversaries and add them to your own larger model, in an arms race.
lsb··on Compressing Images with Neural Networks
You’re looking for what’s called upscaling, like with Stable Diffusion: https://huggingface.co/stabilityai/stable-diffusion-x4-upsca...
lsb··on The Vicuñas and the $9k Sweater
It’s an absolute scandal, as if France only sold wine grapes, and no one in France had ever made or tasted wine. Most of the value add is in the processing.
lsb··on FCC rules AI-generated voices in robocalls illegal
This is an example of the Trust On First Use policy, like when you SSH to a machine whose cert you don't have and you are invited to trust it.

https://en.wikipedia.org/wiki/Trust_on_first_use

lsb··on Mistral looks set to challenge AI frontrunners Google and OpenAI
The iOS app “MLCChat”
lsb··on Mistral looks set to challenge AI frontrunners Google and OpenAI
I can run Mistral 7B, in 3-bit quantization, on my phone, and I get helpful answers for (just today) seeing environment variables in Python.
lsb··on Japan's earthquake alters coastline, extends it by two football fields
That's not _nearly_ the most unusual unit of measurement one could find https://en.wikipedia.org/wiki/List_of_unusual_units_of_measu...
lsb··on Show HN: Stable Diffusion Clock
That’s why I made the only-urban clock face, or the only-train clock face :)
lsb··on Show HN: Stable Diffusion Clock
I can easily generate one every few minutes with a few CPU cores on my desktop, could be a cute dynamic background :) it’s a delicate balance to run on stable enough hardware that has low latency. The Ampere card that’s running this is taking 280W running this in PyTorch and cuda so it’s running a little warm for residential computing
lsb··on Show HN: Stable Diffusion Clock
All the models are available on Huggingface, happy to answer any questions
lsb··on Wikipedia search-by-vibes through millions of pages offline
Yes! The quality probably isn’t as good as Similar Website Finder https://explore2.marginalia.nu/ ;) and I bet using a more recent sentence embedding would lead to better results, I gotta collect more data
lsb··on Wikipedia search-by-vibes through millions of pages offline
I average the embeddings of every 512 bytes of the page text
lsb··on Wikipedia search-by-vibes through millions of pages offline
Thanks! The goal was to demo the database engine and show off how everything can work airgapped (after the browser downloads everything). I think there’s a lot of parameters to tune (Use just the first paragraph of the article, or everything? Search within some short distance of a particular article?) and I haven’t yet.

Wikipedia is a great demo dataset, and I’m definitely up for adding more datasets. Specifically, just like iPhoto lets you search “mountain” and you can get pictures with mountains, might be cool to search with some multi modal models like CLIP on various datasets

lsb··on Wikipedia search-by-vibes through millions of pages offline
Lemme try a few other embedding models after the weekend :)
lsb··on Wikipedia search-by-vibes through millions of pages offline
“Gist” to me implies accuracy, and has a meaning in GitHub, whereas the averaged embedding of 512-character chunks of text is more, uh, impressionistic.
lsb··on Wikipedia search-by-vibes through millions of pages offline
Yeah it’s an off the shelf sentence-transformer model from over a year ago. The demo was more to show off the embedding database, but the embeddings themselves are slightly useful too.

I don’t keep any analytics on the page about what people find and don’t find, so I haven’t set myself up to improve the search results :/

← PreviousPage 2 of 33Next →