HNHacker News
TopNewBestAskShowJobs

kernelsanderz

226 karma · joined August 8, 2016

submissionscomments
kernelsanderz··on An AI agent published a hit piece on me
Theo’s snitch bench is a good data driven benchmark on this type of behavior. But in fairness the models are prompted to be bold to take actions. And doesn’t necessarily represent out of the box or models deployed in a user facing platform.

https://snitchbench.t3.gg/

kernelsanderz··on Amazon just funded a streamer that lets you use AI to make your own TV shows
https://archive.is/vXfbb
kernelsanderz··on Embedding user-defined indexes in Apache Parquet
I’ve been excited about lancedb and its ability to support vector indexes and efficient row level lookups. I wonder if this approach would work for their design goals and still allow broader backwards compatibility with the parquet ecosystem. Have been intrigued by Ducklake, and they’ve leaned into parquet. Perhaps this approach will allow more flexible indexing approaches with support for the broader parquet ecosystem which is significant.
kernelsanderz··on A Python-first data lakehouse
Marimo is really special and solves most of the problems that you have with Jupyter. For those Marimo curious I strongly recommend checking out their YouTube channel. So much effort gone into making these videos really great. https://youtube.com/@marimo-team?si=ZGaf8Zgq5WN3LKRg
kernelsanderz··on Getting Past Procrastination
I’ll read this tomorrow
kernelsanderz··on The best way to use text embeddings portably is with Parquet and Polars
For another library that has great performance and features like full text indexing and the ability to version changes I’d recommend lancedb https://lancedb.github.io/lancedb/

Yes, it’s a vector database and has more complexity. But you can use it without creating indexes and it has excellent polars and pandas zero copy arrow support also.

kernelsanderz··on S3 as a Git remote and LFS server
Also worth checking out https://github.com/jasonwhite/rudolfs

Been using it to store datasets via lfs. Written in rust and has been very reliable.

kernelsanderz··on S3 as a Git remote and LFS server
I’ve been using https://github.com/jasonwhite/rudolfs - which is written in rust. It’s high performance but doesn’t have all the features (auth) that you might need.
kernelsanderz··on James Earl Jones, Actor Whose Voice Could Menace or Melt, Dies at 93
https://archive.is/S7UEJ
kernelsanderz··on OpenAI in Talks for Funding Round Valuing It Above $100B
https://archive.is/do6tG
kernelsanderz··on Did you lose your AirPods?
Some heroes don’t wear capes. They wield scripts, API calls, and a bit of luck.
kernelsanderz··on Former OpenAI employee quit to avoid 'working for the Titanic of AI'
The actual podcast and interview is here https://youtu.be/dzQlRt3y5mU?si=qfS0DlPVcjBn-0zd
kernelsanderz··on Show HN: A modern Jupyter client for macOS
I use jupyterlab via an SSH tunnel
kernelsanderz··on Anthropic's Claude 3.5 Sonnet model now available in Amazon Bedrock
Is anyone else not seeing this in their console. I still only see the old sonnet.
kernelsanderz··on OpenAI expands lobbying team to influence regulation
Non paywall link https://archive.is/lCdK3
kernelsanderz··on iTerm 3.5.1 removes automatic OpenAI integration, requires opt-in
No, there's a configurable system prompt which is:

Return commands suitable for copy/pasting into \(shell) on \(uname). Do NOT include commentary NOR Markdown triple-backtick code blocks as your whole response will be copied into my terminal automatically.

The script should do this: \(ai.prompt)

And then you type your prompt in, and it returns the answer. And then you can choose to edit the command that gets returned or execute it directly.

So essentially what you'd do if you were using the API directly, just more convenient.

kernelsanderz··on iTerm 3.5.1 removes automatic OpenAI integration, requires opt-in
For those who have the funds and means, you can support the maintainer through their donation page here: https://iterm2.com/donate.html
kernelsanderz··on Tantivy – full-text search engine library inspired by Apache Lucene
In practice, a combination of full text and vector databases often gives superior performance than just one of the types. It's called hybrid search. Here's an article that talks a bit about this: https://opster.com/guides/opensearch/opensearch-machine-lear...

Often you take the results from both vector search and lexical search and merge them through algorithms like Reciprocal Rank Fusion.

kernelsanderz··on Tantivy – full-text search engine library inspired by Apache Lucene
Tantivy is also used in an interesting Vector Database product called LanceDb - https://lancedb.github.io/lancedb/fts/ to provide full text search capabilities. Last time I looked it was only through the python bindings, though I know they're looking to implement the rust bindings natively to support other platforms.
kernelsanderz··on DuckDuckGo was down
down for me. doesn't get search results
kernelsanderz··on Energy-Efficient Llama 2 Inference on FPGAs via High Level Synthesis
What's really cool, is that this was built on Karpathy's llama2.c repo https://github.com/karpathy/llama2.c

A great example of the unexpected things that happen when put great code into the commons.

I bet Andrej never expected anything like this when he released it.

kernelsanderz··on Show HN: a Rust based CLI tool 'imgcatr' for displaying images
I just noticed that tmux has some basic sixel support now - https://raw.githubusercontent.com/tmux/tmux/3.4/CHANGES

discovered at https://www.arewesixelyet.com/

kernelsanderz··on Show HN: A fast HNSW implementation in Rust
If people are interested in fast, scalable HNSW implementations which have solid rust bindings then I’d highly recommend checking out usearch https://github.com/unum-cloud/usearch
kernelsanderz··on Nvtop: Linux Task Monitor for Nvidia, AMD and Intel GPUs
There’s also asitop https://github.com/tlkh/asitop
kernelsanderz··on S3 Express Is All You Need
thank so much for sharing. Amazing product BTW!
kernelsanderz··on Powering cost-efficient AI inference at scale with Cloud TPU v5e on GKE
a post about cost-efficient AI without any costs or comparisons?
kernelsanderz··on S3 Express Is All You Need
I'd be fascinated if you could share your insights from using this. Where does the pricing fall down? And is the latency/throughput a big improvement for this use case? (ie. externalizing a search index).
kernelsanderz··on S3 Express Is All You Need
Very excited about being able to build scalable vector databases on DiskANN like turbopuffer or lancedb. These changes in latency are game changing. The best server is no server. The capability a low latency vector database application that runs in lambda and S3 and is dirt cheap is pretty amazing.
kernelsanderz··on Hey, computer, make me a font
Has the code and models been release for this? Looks like amazing work!
kernelsanderz··on LlamaIndex: Unleash the power of LLMs over your data
It's not really. It was unfortunate timing. This used to be called GPT-Index, and then they changed names before Meta released their LLM. So their use predated it.

I feel sorry for the amazing team behind this great library. Changing names is hard.

Page 1 of 3Next →