HNHacker News
TopNewBestAskShowJobs

fancy_pantser

1,001 karma · joined September 20, 2013

works at startups

mail me at github at hella dot rip

submissionscomments
fancy_pantser··on Ask HN: What is the international distribution/statistics of HN visitors?
https://news.ycombinator.com/item?id=35567986

https://news.ycombinator.com/item?id=44572750

https://news.ycombinator.com/item?id=30210378

fancy_pantser··on Tell HN: Merry Christmas
my relatives' is always on a sticker under the AP
fancy_pantser··on I rebuilt FlashAttention in Triton to understand the performance archaeology
When OpenAI announced the Triton language, I was worried I'd be confused one day while reading something because of Nvidia's open-source Triton inference server. I made it quite a long time, but it finally happened today! I was so intrigued for the first few pages and then deeply confused.
fancy_pantser··on Google Titans architecture, helping AI have long-term memory
Mods usually apply [Dupe] to later submissions if a recent (last year or so) one had a fair amount of discussion.
fancy_pantser··on Google Titans architecture, helping AI have long-term memory
Student: Look, there’s hundred dollar bill on the ground! Economist: No there isn’t. If there were, someone would have picked it up already.

To wit, it's dangerous to assume the value of this idea based on the lack of public implementations.

fancy_pantser··on Chips for the Rest of Us
> The team then created VeriGen, the first specialized AI model trained solely to generate Verilog code.

Perhaps it's the first open one. I was an eng manager at a hyperscaler helping one of our clients, a large semiconductor design company, build models to use internally. It was trained on their extensive Verilog repos, tooling, and strict style guides. I see this being repeated across industries, at least since 2023 there are quite a few deep-pocketed S&P 500 orgs creating models from scratch or extensively finetuning to give unique advantages they require. They're rarely announced specifically, but you can often infer from the initial investment or partnership announcements that they're working on it.

fancy_pantser··on The healthcare market is taxing reproduction out of existence
> Prices are entirely hidden

Recent legal changes have made pricing more transparent. In 2020, the federal government issued the "transparency in coverage" final rule under the Federal No Surprises Act. This limited the expenses for emergency care when out-of-network and a few other things, but even more exciting is that hospitals and insurers are now required to publish a comprehensive machine-readable file with ALL items and services. They have to provide all negotiated rates and cash prices for the services and include a display of "shoppable" services in a consumer-friendly format. The machine-readable files are impractical to process yourself for comparison shopping (picture: different formats, horribly de-normalized DB dumps), but many sites and APIs have emerged to scrape them and expose interfaces to do so.

fancy_pantser··on Building more with GPT-5.1-Codex-Max
> Attention is quadratic

Exactly. Standard Multi-Head Attention uses a matrix that grows to 4B parameters for a 64K sequence as a starting place. FlashAttention v2 helps slightly, but as you grow to 128K context length, you still need over 1TB/s memory bandwidth to stay compute-bound in practice even with this optimization.

So there has been a lot of research in this area and model architectures released this year are showing some promising improvements. Sliding windows lose context fidelity and if you go fully linear, you sacrifice math, logic, and long multi-turn (agentic) capabilities, so everyone is searching for a good alternative compromise.

MiniMax-M1 had lightning attention to scale up to 1M context lengths. It's "I/O aware" via tiling and calculates attention two ways block-wise (intra-block traditional attention and inter-block linear attention), thereby avoiding the speed-inhibiting cumulative summation.

DeepSeek V3.2 uses DeepSeek Sparse Attention (DSA), which is sub-linear by only computing "interesting" pairs. For example, in 128K context lengths this requires only 10-20% of attention pairs to be materialized.

Both Qwen3-Next and Kimi Linear adopt a Gated DeltaNet, which is borrowed from Mamba2. In Qwen3-Next it alternates three Gated DeltaNet (linear attention) layers for every one gated [full] attention. The speedup is from a delta rule, which basically amounts to caching in a hand-wavy way.

There's no universally-adopted solution yet, as these are all pretty heavy-duty compromises, but the search is going strong right now for linear or better attention mechanisms that still perform well.

fancy_pantser··on A graph explorer of the Epstein emails
Software like i2 Analyst's Notebook.
fancy_pantser··on AI World Clocks
Have you given using MCPs to provide documentation and examples a shot? I always have to bring in docs since I don't work in Python and TS+React (which it seems more capable at) and force it to review those in addition to any specification. e.g. Context7
fancy_pantser··on Show HN: AI toy I worked on is in stores
Would love to see this connect to a smartphone running a matching app, doing the inferencing on the phone so you wouldn't have to bill for increments. It will be a few more years until that can be done in a low-enough latency way (improved models, more compute and memory available).
fancy_pantser··on Show HN: OWhisper – Ollama for realtime speech-to-text
I scratched a similar itch and found local LLMs plus Whisper worked really well to listen in and "DJ" a soundtrack while playing tabletop RPGs with a group. If you want to check it out: https://github.com/sean-public/conductor
fancy_pantser··on The Math Is Haunted
Claimify from MS research aims in this direction. There's a paper and video explainer from a few months ago.

https://www.microsoft.com/en-us/research/blog/claimify-extra...

fancy_pantser··on Internet Archive is now a federal depository library
There are utilities to help, waybackpack comes to mind, but I haven't looked in a while. https://github.com/jsvine/waybackpack
fancy_pantser··on Intel CEO Letter to Employees
As OP, I think I'm supposed to arrange for the card to be collected... oh hey you're Kirk! It's Sean formerly of Amino as well :)
fancy_pantser··on Pgactive: Postgres active-active replication extension
Are you looking for a tool like Barman?
fancy_pantser··on The Zed Debugger Is Here
I've been using the unofficial builds via scoop for the last two months. It's working great so far. I use it on a Macbook as well and I haven't found any features that are missing or buggier on Win11. Really enjoying the new agent version of the AI assistant, which I use with both Claude API and devstral locally via Ollama.

https://github.com/deevus/zed-windows-builds

fancy_pantser··on Kagi Assistant is now available to all users
Maybe their privacy pass is useful then?

https://help.kagi.com/kagi/privacy/privacy-pass.html

fancy_pantser··on LLM Workflows then Agents: Getting Started with Apache Airflow
Been having a great time with postgresml for this exact kind of thing. If you don't need a complex DAG but have a simple pipeline or work queue that can be easily represented in postgres anyway, it's very straightforward to work with and nicely encapsulates all of your processing (traditional data munging and LLM calls) together with a modest extension of a familiar system.
fancy_pantser··on Prime numbers so memorable that people hunt for them
You can use the number with your local area code just about anywhere at the pump to get a gas discount as well (a common loyalty reward program benefit).
fancy_pantser··on Kelly Can't Fail
A very similar card game played by deciding when to stop flipping cards from a deck where red is $1 and black is −$1 as described in Timothy Falcon’s quantitative-finance interview book (problem #14). Gwern describes it and also writes code to prove out an optimal stopping strategy: https://gwern.net/problem-14
fancy_pantser··on 1-800-ChatGPT
have you checked daily.co pricing for audio-only? they do SIP
fancy_pantser··on Show HN: I built an open-source data pipeline tool in Go
> specifying and visualizing DAGs

Do you mean like Airflow or Pachyderm? I am also very interested in new tooling in this space that has these features.

fancy_pantser··on Ask HN: To those with successful browser extension(s), how did you grow it?
Recipe Filter: https://chromewebstore.google.com/detail/recipe-filter/ahlcd...
fancy_pantser··on 'I grew up with it': readers on the enduring appeal of Microsoft Excel
here ya go ;) https://news.ycombinator.com/item?id=24665282
fancy_pantser··on Goldman and Apple 'illegally sidestepped' obligations to credit-card customers
apple glazing:apple::navel gazing:orange
fancy_pantser··on Amazon reveals first color Kindle, new Kindle Scribe, and more
Look into PocketBook readers. They have a full lineup with color and grayscale in various sizes. They run linux, you can ssh into it and install koreader, etc. They have a good privacy story and little vendor lock-in. The experience is a lot like my Kobo, but the matching mobile app is much better. I continue where I leave off with my iPad sometimes, it's a nice feature.

https://pocketbook.ch/en-ch/catalog

fancy_pantser··on Ask HN: What are you working on (August 2024)?
I wondered this myself and found some. I've played: When Rivers Were Trails, Thunderbird Strike, Dialect, Zulu Dawn, This Land Is My Land. There are a lot of options in boardgames (maybe more than with video games?), such as Burn the Fort.
fancy_pantser··on Ollama now supports tool calling with popular models in local LLM
I see Command-R+ but not Command-R marked for tool use. The model is geared for it, much easier to fit on commodity hardware like 4090s, and Ollama's own description for it even includes tool use. I think it's just not labeled for some reason. It works really well with the provided ollama-python package and other tools that already brought function calling capabilities via Ollama's API.

https://ollama.com/library/command-r

fancy_pantser··on Txtai: Open-source vector search and RAG for minimalists
No fine-tuning is necessary. You can use something reasonably good at RAG that's small enough to run locally like the Command-R model run by Ollama and a small embedding model like Nomic. There are dozens of simple interfaces that will let you import files to create a RAG knowledgebase to interact with as you describe, AnythingLLM is a popular one. Just point it at your locally-running LLM or tell them to download one using the interface. Behind the scenes they store everything in LanceDB or similar and perform the searching for you when you submit a prompt in the simple chat interface.
← PreviousPage 2 of 9Next →