HNHacker News
TopNewBestAskShowJobs

platypii

342 karma · joined September 4, 2012

submissionscomments
platypii··on A visual explainer of how to scroll through billions of rows in the browser
Sylvain Lesage’s cool interactive explainer on visualizing extreme row counts—think billions—inside the browser. His technical deep dive explains how the open-source library HighTable works around scrollbar limits by:

- Lazy loading - Virtual scrolling (allows millions of rows) - "Infinite Pixel Technique" (allows billions of rows)

Hyperparam sponsored Sylvain’s work as part of our broader effort to invest in open-source infrastructure and get ahead of the data-scale problems that are emerging with LLMs. With a regular table, you can view thousands of rows, but the browser breaks pretty quickly. We created HighTable with virtual scroll so you can see millions of rows, but that still wasn’t enough for massive unstructured datasets. What Sylvain has built virtualizes the virtual scroll so you can literally view billions of rows—all inside the browser. His write-up goes deep into the mechanics of building a ridiculously large-scale table component in react.

platypii··on Ask HN: Where are you keeping your LLM logs?
We're willing to spend money, but I've had the "datadog billing problem" before where it starts reasonable and then grows to a non-trivial percent of saas budget, and then theres a scramble to refactor. Trying to get ahead of that as the LLM logs are MUCH larger that my APM logs.
platypii··on What UI do you use on top of data engineering tools to look at data?
Makes sense. I'm not currently in snowflake because I'm mostly working with local parquet files. Would prefer not to have to pay for snowflake just to explore my data. I'm interested in better data UIs though so I might need to check it out.
platypii··on Show HN: We built an AI tool for working with massive LLM chat log datasets
I started Hyperparam one year ago because I knew that the world of data was changing, and existing tools like Python and Jupyter Notebooks were not built for the scale of LLM data. The weights of LLMs may be tensors, but the input and output of LLMs are massive piles of text.

No human has the patience to sift through all that text, so we need better tools to help us understand and analyze it. That's why I built Hyperparam to be the first tool specifically designed for working with LLM data at scale. No one else seemed to be solving this problem.

platypii··on Lessons from Hyperparam's year of open source data transformation
This is a Q&A I did on what I learned from a year of open source data transformation. Most of all, it reinforced my belief that browser-native tools aren’t “toys” that don’t work for real systems. When Hugging Face integrated my libraries, it confirmed that the browser can handle serious data work, and maybe there's an opportunity for more browser-based data tools.
platypii··on Ask HN: How far can we push the browser for large-scale data parsing?
As with anything, there are engineering tradeoffs.

What I've found is that moving data processing toward the browser has been for one, a refreshing developer experience because I don't need to build a pair of backend+frontend. From a user experience point of view, I think you can build MORE interactive data applications by pushing it toward the frontend.

platypii··on From GPT-4 to GPT-5: Measuring progress through MedHELM [pdf]
Why not? We are trying to evaluate AI's capabilities. It's OBVIOUS that we should compare it to our only prior example of intelligence -- humans. Saying we shouldn't compare or anthropomorphize machine is a ridiculous hill to die on.
platypii··on The Quest for Instant Data
This is the story of how I spent a year making the world's fastest Parquet loader in JavaScript. The goal:

- Make a faster, more interactive viewer for AI datasets (which are mostly parquet format)

- Simplify the stack by doing everything from the browser (no backend)

TLDR: My open-source library Hyparquet can load data in 155ms, which would take 3466ms in duckdb-wasm for the same file.

platypii··on Show HN: Hyperparam: OSS Tools for Exploring Datasets Locally in the Browser
I don’t have benchmarks specifically against duckdb. I’m sure native C++ will run faster than JavaScript.

But whats important is that with Hyperparam you can do it in the browser, where the bottleneck will always be network-bound not cpu-bound.

platypii··on Show HN: Hyperparam: OSS Tools for Exploring Datasets Locally in the Browser
Funny you say that, because I built these tools because I wanted to build something very much like what you're describing!

I was trying to look at, filter, and transform large AI datasets, and I was frustrated with how bad the existing tool was for working with datasets with huge amounts of text (web scrapes, github dumps, reasoning tokens, agent chat logs). Jupyter notebook is woefully bad at helping you to look at your data.

So I wanted to build better browser tools for working with AI datasets. But to do that I had to build these tools (there was no working parquet implementation in JS when I started).

Anyway I'm still working on building an app for data processing using LLM chat assistant to help a single user curate entire datasets singlehandedly. But for now I'm releasing these components to the community as open source. And having them "do a single task each" was very much intentional. Thanks for the comment!

platypii··on Show HN: Hyperparam: OSS tools for exploring datasets locally in the browser
Yea except with parquet you don't need to load the entire file, the parquet metadata let's you do http range requests for just the data you need.

For example this parquet is the entire english wikipedia (400mb) but loads less than 4mb including html and all js to display the first rows:

https://hyperparam.app/files?key=https%3A%2F%2Fs3.hyperparam...

This way you can have huge AI datasets in cloud storage, and still have a nice interface for looking at your data.

In particular, a lot of modern AI datasets are huge walls of text (web scrapes, chains of thought, or agentic conversation histories), and most datasets on huggingface are in parquet. So you can much more quickly look at your data this way versus say jupyter notebooks.

Here's the glaive reasoning dataset on the Hyperparam hugging face space:

https://huggingface.co/spaces/hyperparam/hyperparam?url=http...

platypii··on Show HN: Hyperparam: OSS Tools for Exploring Datasets Locally in the Browser
That's fair criticism... to be honest when I started the project it was more focused on hyperparameters, and it evolved into this javascript-for-ai mission. But now I just kind of liked the name.
platypii··on Show HN: Hyperparam: OSS tools for exploring datasets locally in the browser
It does support using S3 presigned requests, but it's admittedly a little awkward to ask a server for a presigned request before every fetch. But does still have the benefit that you can have a small and light server just handing out signed requests, and then the user and their browser does the heavy lifting. This can save a lot on scaling out server costs.

That being said, I wish there was a better auth story. Open to suggestions if anyone has ideas!

platypii··on Show HN: Hyperparam: OSS Tools for Exploring Datasets Locally in the Browser
Duckdb and datafusion are super cool! But they are VERY large wasm blobs (30-40mb each). This is often larger than the data you’re trying to load. And they add complexity with serving and deploying wasm files.

Hyparquet is 10kb of pure js, and so its trivial to deploy on a modern webapp, and wins hands down on time-to-first-data metric.

platypii··on Show HN: Hyperparam: OSS tools for exploring datasets locally in the browser
Zero telemetry, fully local. It spawns `http-server` on port 2048 and opens your browser at `localhost`. Similar pattern as Jupyter Notebooks. Feel free to audit the code... the server is <200 LOC.
platypii··on Tesla sales plummet in the UK, France, and Germany
FSD makes Tesla superior to any car out there. No other car comes even close.

Although I heard that FSD was already crippled in eu so maybe they aren't missing out as much.

platypii··on Ask HN: Teams using AI – how do you prevent it from breaking your codebase?
AI is even better at writing tests than writing code. So have it write the tests first and then write the code.
platypii··on Hyparquet.js: World's Smallest and Most Conformant Parquet File Parser
My goal is to build tools which enable working with large-scale ML datasets in the browser. The browser is critical for building compelling UIs, but previous parquet js libraries had gone abandoned.

Apache Parquet is a very complicated format. It has 22 data types, 9 encodings, 8 compression codecs. However, I can confidently say that Hyparquet is now the most conformant parquet parser in existence. It can open all the parquet files: more than PyArrow and DuckDB. I dare you to find a file that Hyparquet can’t open!

Hyparquet is MIT licensed, and there is a demo github page which can open parquet files in the browser with no backend server.

platypii··on Show HN: Hyparquet 1.0 – Apache Parquet Parser for the Browser
I definitely think that UX is an underappreciated area for machine learning data. I want to make a set of libraries and tools that make it easier for people to work with ML data in the browser. The first step of good data science is to become one with your data.

I started with parquet because most datasets for modern LLMs are in parquet format. But there are other formats like JSONL which are common too.

platypii··on Guardrails AI wants to crowdsource fixes for GenAI model problems
Pretty interesting to go from a world of deterministic code, to LLMs which can do incredible things, but unreliably. In a world of LLMs, I could imagine guardrails being a table-stakes part of engineering an ML system, just like unit tests, and CI/CD would be for traditional software.
platypii··on On the dangers of stochastic parrots: Can language models be too big? (2021)
This paper is embarrassingly bad. It's really just an opinion piece where the authors rant about why they don't like large language models.

There is no falsifiable hypothesis to be found in it.

I think this paper will age very poorly, as LLMs continue to improve and our ability to guide them (such as with RLHF) improves.

platypii··on Code CAD – Use code to create CAD models
Shout out to JSCAD -- Javascript solid object CAD software. I really like it. Some things I've been able to do with JSCAD that would be hard to imagine with any other CAD software:

- From the same source files, I can generate either an "assembled" model or a "print version" arranged for 3D printing.

- Integrated into github CI/CD so that I know immediately if I broke my designs. Easy to write tests for things like bounding box size.

- Use eslint to enforce style rules. Mocha for tests. Typescript if you want.

- Browserify to bundle the design into a website. Directly from the source CAD files, without rendering to a huge mesh file format.

It's cool because you get access to the whole javascript ecosystem, and it's native to the browser.

https://openjscad.xyz/

platypii··on Android's new Bluetooth stack rewrite (Gabeldorsh) is written with Rust
I went through the same thing recently. The best reference I've ever seen on Android Bluetooth is this series of posts:

https://medium.com/@martijn.van.welie/making-android-ble-wor...

platypii··on A Book about Aircraft Scale Drawings Creating with Inkscape and Gimp
I work in inkscape a lot, and aesthetically like to keep things aligned on integers where possible. Makes it easier to edit by hand if needed, keeps things aligned, and keeps files small.

You have to be careful about when you move, scale, group, ungroup, etc. But I have found it pretty reasonable to stay on nice round numbers with inkscape. I feel like inkspace is BETTER for that than say illustrator.

Worst case you can export as "Optimized SVG" and reduce the number of decimals. Make sure you check it after though since it can change the design when rounding.

platypii··on Andean condor can fly for 100 miles without flapping wings
No you're wrong. Galileo's experiment doesn't apply here. Heavier objects DO fall faster at terminal velocity. A glider in sustained flight is an example of that.

Galileo just showed that objects accelerate at the same rate due to gravity. But they do NOT fall at the same rate unless in a vacuum.

platypii··on Wildlife is reclaiming Yosemite National Park
https://www.nps.gov/yose/planyourvisit/visitation.htm

April is when visits ramp up.

platypii··on Mate Desktop 1.24
Really grateful that Mate exists. To this day, gnome3 has not reached parity with gnome2 on flexibility of configuration. I don't even make a lot of customizations to my desktop, but having the option when I need it is really key.
platypii··on Twitter for Mac is incapable of accepting certain letters in the password field
Obligatory:

"I included emoji in my password and now I can't log in to my Account"

https://apple.stackexchange.com/questions/202143/i-included-...

platypii··on Issue 914451: Autofill does not respect autocomplete="off"
Thank you chrome, for recognizing that the USER is what matters, not the website developer. Websites started blocking legitimate uses of autocomplete, and that makes the web less secure. As a user I sincerely hope that "autocomplete=off" dies just like the blink tag, as it is user hostile.
platypii··on Equifax Faces Multibillion-Dollar Lawsuit Over Hack
Submit it to debt collectors and put it on Equifax's credit report
Page 1 of 3Next →