HNHacker News
TopNewBestAskShowJobs

tmostak

642 karma · joined March 29, 2012

Twitter: @toddmostak
submissionscomments
tmostak··on Use your Nvidia GPU's VRAM as swap space on Linux
GPU-accelerated databases have a long history. I founded HeavyAI (previously MapD/OmniSci) in 2013, but there are or have been many other startups in this space, such as Voltron Data, Kinetica, Sqream, etc. And now you have major players like IBM, Starburst, and Microsoft (which just announced Fabric SQL on GPU today) working on their own GPU-accelerated systems. GPUs have a huge advantage in terms of compute, memory, and interconnect bandwidth over CPU, as long as you can keep them fed with data.

I believe within 2-3 years databases and data warehouses on GPU will be common. The widespread use of agents to query data will be a part of this, as there will be a need to run far more queries at lower latency than needed for the ETL and BI workloads of the past.

tmostak··on Waymo robotaxi hits a child near an elementary school in Santa Monica
Evidence (preferably with recent Teslas/HW4)?
tmostak··on Waymo robotaxi hits a child near an elementary school in Santa Monica
Evidence of this? I own a Tesla (HW4, latest FSD) as well as have taken many Waymo rides, and have found both to react well to unpredictable situations (i.e. a car unexpectedly turning in front of you), far more quickly than I would expect most human drivers to react.

This certainly may have been true of older Teslas with HW3 and older FSD builds (I had one, and yes you couldn't trust it).

tmostak··on Waymo robotaxi hits a child near an elementary school in Santa Monica
Do you have data to back this claim up, specifically with HW4 (most recent hardware) and FSD software releases?
tmostak··on Prefix sum: 20 GB/s (2.6x baseline)
Even without NVLink C2C, on a GPU with 16XPCIe 5.0 lanes to host, you have 128GB/sec in theory and 100+ GB/sec in practice bidirectional bandwidth (half that in each direction), so still come out ahead with pipelining.

Of course prefix sums are often used within a series of other operators, so if these are already computed on GPU, you come out further ahead still.

tmostak··on Modern Minimal Perfect Hashing: A Survey
We've made extensive use of perfect hashing in HeavyDB (formerly MapD/OmniSciDB), and it has definitely been a core part of achieving strong group by and join performance.

You can use perfect hashes not only the usual suspects of contiguous integer and dictionary-encoded string ranges, but also use cases like binned numeric and date ranges (epoch seconds binned per year can use a perfect hash range of one bin per year for a very wide range of timestamps), and can even handle arbitrary expressions if you propagate the ranges correctly.

Obviously you need a good "baseline" hash path to fall back to you, but it's surprising how many real-world use cases you can profitably cover with perfect hashing.

tmostak··on Show HN: TabPFN v2 – A SOTA foundation model for small tabular data
This looks amazing!

Just looking through the code a bit, it seems that the model both supports a (custom) attention mechanism between features and between rows (code uses the term items)? If so, does the attention between rows help improve accuracy significantly?

Generally, for standard regression and classification use cases, rows (observations) are seen to be independent, but I'm guessing cross-row attention might help the model see the gestalt of the data in some way that improves accuracy even when the independence assumption holds?

tmostak··on All You Need Is 4x 4090 GPUs to Train Your Own Model
You should be able to train/full-fine-tune (i.e. full weight updates, not LoRA) a much larger model with 96GB of VRAM. I generally have been able to do a full fine-tune (which is equivalent to training a model from scratch) of 34B parameter models at full bf16 using 8XA100 servers (640GB of VRAM) if I enable gradient checkpointing, meaning a 96GB VRAM box should be able to handle models of up to 5B parameters. Of course if you use LoRA, you should be able to go much larger than this, depending on your rank.
tmostak··on How Meta trains large language models at scale
This assumes that you can linearly scale up the number of TPUs to get equal performance to Nvidia cards for less cost. Like most things distributed, this is unlikely to be the case.
tmostak··on GPT-4.5 or GPT-5 being tested on LMSYS?
Are you measuring tokens/sec or words per second?

The difference matters as generally in my experience, Llama 3, by virtue of its giant vocabulary, generally tokenizes text with 20-25% less tokens than something like Mistral. So even if its 18% slower in terms of tokens/second, it may, depending on the text content, actually output a given body of text faster.

tmostak··on Ollama v0.1.33 with Llama 3, Phi 3, and Qwen 110B
But it's likely to be much slower than what you'd get with a backend like llama.cpp on CPU (particularly if you're running on a Mac, but I think on Linux as well), as well as not supporting features like CPU offloading.
tmostak··on Show HN: Use natural language to query and visualize 400M tweets
Thank you, it's been a major team effort!
tmostak··on Explore 400M tweets with LLM-powered conversational analytics
More info can be found here: https://www.heavy.ai/heavyiq/overview
tmostak··on Ask HN: Who is hiring? (February 2024)
HEAVY.AI | SQL Analyst/Wrangler | Part-time or Full-time | Remote

HEAVY.AI builds a GPU-accelerated analytics platform that allows users to interactively query and visualize billions of records of data in milliseconds.

We’re looking for someone who really knows SQL. If you can decipher schemas, figure out what’s wrong with SQL statements and correct them, as well as generate queries in response to user questions, we'd love to talk to you.

The work would initially be on contract, but could lead to full-time employment. Geospatial analytics, data science background, and Python programming skills would be very useful to have as well, but are not absolute requirements.

If interested please reach out to pey.silvester@heavy[dot]ai.

tmostak··on Drawing.garden
These are awesome!
tmostak··on OpenAI investors keep pushing for Sam Altman’s return
I assume if MSFT/Satya are supportive it won't be an issue.
tmostak··on Fine-tuning GPT-3.5-turbo for natural language to SQL
It wasn't clear to me what evaluation method was being used, the chart in the blog says Execution Accuracy, but the numbers that seem to be used appear to correlate with "Exact Set Match" (comparing on SQL) instead of the "Execution With Values" (comparing on result set values). For example, DIN-SQL + GPT-4 achieves an 85.3% "Execution With Values" score. Is that what is being used here?

See the following for more info:

https://yale-lily.github.io/spider https://github.com/taoyds/spider/tree/master/evaluation_exam...

tmostak··on Fine-tuning GPT-3.5-turbo for natural language to SQL
I agree that Spider queries are not necessarily representative of the SQL you might see in the wild from real users, but looking at some analysis I did of the dataset around 43% of the queries had joins, and a number had 3, 4, or 5-way joins.
tmostak··on Getting to the bottom of web map performance
You could try OmniSci, it’s a database, rendering engine, and interactive analytics frontend (or any combination of the above) and can easily query and render millions to tens of billions of points interactively while allowing for things like tooltips on the data. See omnisci.com/demos for some live examples.
tmostak··on OmniSci launches free edition of platform for interactive visual analytics
OmniSci can run on any Nvidia GPU with sufficient RAM (we'd generally recommend >= 8GB), including a 3080. (I have two 3090s myself!) It also can run purely (and performantly) on CPU, and with the Intel's help we're further optimizing our capabilities on X86. Note however that currently you can't use our rendering engine without a GPU, however there is some initial support to run on CPU and render on an Nvidia GPU if you're interested, and soon enough we hope to support AMD and Intel integrated/discrete GPUs for rendering as well.
tmostak··on International Space Station 437.800 MHz cross band FM repeater activated
Definitely can pick it up with a HT, just caught it for ~3 minutes in the Bay Area on a Baofeng HT with whip antenna, and even picked up the first part while I was still indoors. There was a lot of static although I could make out some of the sentences.
tmostak··on 1.1B Taxi Rides Using OmniSciDB and a MacBook Pro
Unfortunately it's probably not in the cards in the near term just do to other priorities and insufficient demand (plus alternatives like HIP for AMD). I will say a lot of us here at OmniSci would kill to leverage the latent GPUs in our Macs and other places, so we'd welcome any community help towards this end (it's not a trivial thing to add, but also not particularly difficult either, just work).

We'll plan to update that issue with the above.

tmostak··on 1.1B Taxi Rides Using OmniSciDB and a MacBook Pro
Hi @nikita, good to reconnect.

When you say an array and not a hash table, do you just mean a simple perfect hash table indexed by the offset of the dictionary id? We use this fairly extensively for inputs of bounded domain (i.e. dictionary-encoded strings, moderately-sized integer ranges, even binned values, numeric or timestamp), but call it a perfect hashing. Assume we're talking about the same thing but wanted to clarify.

tmostak··on 1.1B Taxi Rides Using OmniSciDB and a MacBook Pro
For those wanting to try it for themselves, we recently released a preview of our full stack for Mac (containing both OmniSciDB as well as our Immerse frontend for interactive visual analytics), available for free here: https://www.omnisci.com/mac-preview. This is a bit of an experiment for us, so we'd love your feedback! Note that the Mac preview doesn't yet have the scalable rendering capabilities our platform is known for, but stay tuned.

You can also install the open source version of OmniSciDB, either via tar/deb/rpm/Docker for Linux (https://www.omnisci.com/platform/downloads/open-source) or by following the build instructions for Mac in our git repo: https://github.com/omnisci/omniscidb (hopefully will have standalone builds for Mac up soon). You can also run a Dockerized version on your Mac, but as a disclaimer the performance, particularly around storage access, lags a bare metal install.

tmostak··on 1.1B Taxi Rides Using OmniSciDB and a MacBook Pro
Just to clarify, most of the query engine is built around LLVM-based JIT compilation, and CUDA is not really used per say except for GPU-specific operators like atomic aggregates and thread synchronization, and of course we use the driver API to manage the GPUs, allocate memory, etc. Supporting AMD GPUs or the upcoming Intel Xe GPUs (or frankly anything that has an LLVM-backend) would not be particularly hard, it would just require adding similar supporting infra.
tmostak··on 1.1B Taxi Rides Using OmniSciDB and a MacBook Pro
To be fair, the c5d.9xlarge instances are $1.728 each per hour, or $5.18 for the 3-server cluster (looks to be about $3.06/hr for reserved 1-year pricing). Even with reserved pricing, that's $26,806 a year, or 6.5X more than a $4K laptop that likely will last for years and would be bought anyway (or at least a cheaper variant, which would also run these queries nearly as quickly). Of course that's very apples-to-oranges, so another way to look at this is that OmniSci would probably see significantly better performance on a single c5d.9xlarge than what we saw on this Mac (would need to benchmark, but informally I can say that OmniSci was 2-3X faster running on CPU on my Linux workstation compared to my Mac).

Disclaimer: No disrespect to ClickHouse here, it's an amazing system that I'm sure beats out OmniSci for certain workflows.

tmostak··on 1.1B Taxi Rides Using OmniSciDB and a MacBook Pro
Mark did a benchmark of SQLite using its internal file format a few years ago (https://tech.marksblogg.com/billion-nyc-taxi-rides-sqlite-pa...), clocking the import at 5.5 hours. It looks like this was done though on a spinning disk, so given a proper SSD, and a newer version of SQLite, it might be much faster.
tmostak··on Tools I Recommend for Building Geospatial Web Applications
If you sign up for a cloud trial, we can provide you with an ODBC interface. Other options are using our JDBC or our Python/JS interfaces, which are in our open source: https://github.com/omnisci .
tmostak··on Tools I Recommend for Building Geospatial Web Applications
Hi John, thanks!

One option is our cloud (pricing is on the page). https://www.omnisci.com/cloud . You can also spin us up from the AWS, Azure, or GCP marketplaces. If you need an on-prem option, feel free to reach out: info@omnisci.com.

tmostak··on Tools I Recommend for Building Geospatial Web Applications
Also worth checking out OmniSci, (https://www.omnisci.com), formerly known as MapD, which can query and visualize tens of billions of geospatial records in sub-second timeframes. See here for an interactive demo of 11.6 billion ship AIS records: https://www.omnisci.com/demos/ships. Note that the backend database is open source, but the rendering engine and web front end are part of the paid offering. [Disclaimer: I work there]
Page 1 of 5Next →