HNHacker News
TopNewBestAskShowJobs

AaronFriel

9,073 karma · joined July 23, 2012

Opinions are my own and not those of my employer, etc.

https://bsky.app/profile/aaronfriel.bsky.social

mayreply at my name dot com

submissionscomments
AaronFriel··on Apple must pay 13B euros in back taxes, EU's top court rules
Sure, but I think it's the sort of "self-dealing" that's the problem.

Suppose it were possible for a wholesaler in Ireland to purchase a product in bulk at around 1/3 of MSRP. Market equilibrium would drive the price of that product down, right? If any other company could do that, price competition would prevail and eventually the delta between the import cost (in Ireland) and the export price (to an Italian phone shop) would shrink. Likewise, the retailers that wholesaler sells to would want to have some margin as well. This would put pressure on the wholesaler - likely competing with other wholesalers - to have a small margin as their "value added" is insubstantial.

But, crucially, this is not a case of three independent entities: a manufacturer, a wholesaler, and a retail business. This is one entity, with three subsidiaries and setting prices between them to minimize tax burden, and setting prices in ways that are simply nonsensical, like selling products from one subsidiary to another at or below cost, and then to another at full retail price. If they were three separate companies, the manufacturer and the retailer would go under. In this scheme, the wholesaler is somehow adding all of the value to the product, despite doing nothing more than acting as a shipping hub.

AaronFriel··on Llama 3.1
Is there pricing available on any of these vendors?

Open source models are very exciting for self hosting, but the per-token hosted inference pricing hasn't been competitive with OpenAI and Anthropic, at least for a given tier of quality. (E.g.: Llama 3 70B costing between $1 and $10 per million tokens on various platforms, but Claude Sonnet 3.5 is $3 per million.)

AaronFriel··on Llama 3.1
If previous quantization results hold up, fp8 will have nearly identical performance while using 405GiB for weights, but the KV cache size will still be significant.

Too bad, too, I don't think my PC will fit 20 4090s (480GiB).

AaronFriel··on 10% of Cubans left Cuba between 2022 and 2023
Blanket accusations like this are racist. Genital mutilation is an abhorrent practice. While the impact is uniquely horrible for women, it is definitely "throwing stones in glass houses" to act as though it's a strange cultural practice to damage the nerves of a child's genitals, a practice done to millions of boys in Europe by those of multi-generational European descent.

Now, you could be opposed to both practices of course, but you are choosing to make a blanket accusation to suit a point of view.

AaronFriel··on Never Update Anything
A feature I've wanted for ages, for every OS package manager (Windows, apt, yum, apk, etc.), every language's package manager (npm, pypi, etc.), and so on is to update but filter out anything less than one day, one week, or one month old. And it applies here, too.

Now, some software, they effectively do this risk mitigation for you. Windows, macOS, browsers all do this very effectively. Maybe only the most cautious enterprises delay these updates by a day.

But even billion dollar corporations don't do a great job of rolling out updates incrementally. This especially applies as tools exist to automatically scan for dependency updates, the list of these is too long to name - don't tell me about an update only a day old, that's too risky for my taste.

So for OS and libraries for my production software? I'm OK sitting a week or a month behind, let the hobbyists and the rest of the world test that for me. Just give me that option, please.

AaronFriel··on Google presents method to circumvent automatic blocking of tag manager
This is incorrect, the documentation in the article involves configuring an L7 load balancer to route a path on the same domain as the origin to Google Tag Manager. This means even `SameSite=strict`, `Secure`, `HttpOnly` cookies will be sent to GTM, if the instruction I quoted is followed to pass all cookies and query strings.

It's weird that the document specifically says "all cookies" - that gives GTM access to every cookie sent to your application.

AaronFriel··on Google presents method to circumvent automatic blocking of tag manager
> Override the Host header to be equal to GTM-123456.fps.goog. Allow all cookies and query strings to be forwarded.

Did a security team review this? This leaks session cookies for your domain to Google in a way GTM did not previously capture.

AaronFriel··on I'm not a fan of strlcpy(3)
> Imagine you are writing performance sensitive code. You want to get a substring from a string, one that is not going to live outside your hot loop. In standard C you can just reference a part of a string with a pointer offset.

If you want your substring to terminate in the same place as the original, at a null terminator. But that sadly is almost never the case, and as many C practitioners know, references like this are often unsafe and so APIs that substring tend to copy. That's just what they have to do to pass address sanitizer and static analysis checks.

If you want arbitrary views on a null terminated string, well, it's no longer null terminated and that's just the start of your problems in C.

In languages like Rust and Go, taking a view of a string or array is safe and doesn't copy the underlying data or require an allocation. So if you are writing performance sensitive code where substrings are a major contributor to CPU cycles, best go with those language (or C++) rather than C.

AaronFriel··on Show HN: A fast OSS voice assistant
In autoregressive models we can "feed forward" the model by injecting additional tokens. Computing the KV cache entries for those tokens (called"prefill"), then resuming decoding. If we can do this quickly, and on the same node that has a hot KV cache (or otherwise low latency access to shared KV cache), we are quite a ways closer to offering a full duplex, or at least near zero latency, language model API. This does require a full duplex connection (i.e.: Websocket).

For true full duplex communication, including interruption, it will be more challenging but should be possible with current model architectures. The model may need to be able to emit no-op or "pause" tokens or be used as the VAD, and positional encoding of tokens might need to be replaced or augmented with time and participant.

I imagine the first language model which has "awkward pauses" is only a year or so away.

AaronFriel··on Show HN: A fast OSS voice assistant
Because the model has been trained to do what you tell it to do? That's what instruction pretraining/fine-tuning is.
AaronFriel··on Show HN: A fast OSS voice assistant
I'm impressed by the latency using a request response. It looks this uses speech detection locally using Silero voice activity detector model using the ONNX web runtime, collects audio, then performs a POST. It doesn't look like the POST is submitted though until I'm done speaking. The response depends on chaining together several AI APIs that themselves are very, very fast to provide a seamless experience.

This is very good. But this is, unfortunately, still bound by the dominant paradigm of web APIs. The speech to text model doesn't get its first byte until I'm done talking, the LLM doesn't get its first byte until the speech to text model is done transcribing, and the speech to text model doesn't get its first byte until the LLM call is complete.

When all of these things are very fast, it can be very seamless, but each of these contributes to a floor of latency that makes it hard to get to lifelike conversation. Most of these models should be capable of streaming prefill - if not decode (for the transformer like models) - but inference servers are targeting the lowest common denominator on the web: a synchronous POST.

When only 3 very fast models are involved, that's great. But this only compounds when trying to combine these with agentic systems, tool calling.

The sooner we adopt end-to-end, bidirectional streaming for AI, the sooner we'll reach more lifelike, friendly, low latency experiences. After all, inter-speaker gaps in person to person conversations are often in the sub-100ms range and between friends, can even be negative! We won't have real "agents" until models can interrupt one another and talk over each other. Otherwise these latencies compound to a pretty miserable experience.

Relatedly, Guillermo - I've contributed PRs to reduce the latency of tool calling APIs to the AI SDK and Websockets to Next.js. Let's break free of request-response and remove the floor on latency.

AaronFriel··on Run the strongest open-source LLM model: Llama3 70B with just a single 4GB GPU
A good rule of thumb is that models can be quantized to 6 to 8 bits per weight without significantly degrading quality. This is convenient for the math: 70GB plus some overhead for the attention matrices (ongoing requests). This overhead depends on workload and context lengths, but you should expect about 30% more. So, around 100GB for a server under load.
AaronFriel··on Electricity prices in France turn negative as renewable energy floods the grid
Oil prices went negative due to an exogenous, rare event shocking demand. No one expected that would recur. (And indeed, many people did try to figure out how to store oil in one-off vessels, short term leases for storage, etc., though I don't know if anyone succeeded.)

Solar panel production on the other hand is exponential and growing much faster than overall power consumption. It can be very, very favorable to build batteries and that's why grid scale battery production is taking off. There is in fact a storage business dependent on energy arbitrage over time, it's lucrative, and all indications suggest it will continue to be lucrative for many years to come.

AaronFriel··on Cost of self hosting Llama-3 8B-Instruct
These costs don't line up with my own experiments using vLLM on EKS for hosting small to medium sized models. For small (under 10B parameters) models on g5 instances, with prefix caching and an agent style workload with only 1 or a small number of turns per request, I saw on the order of tens of thousands of tokens/second of prefill (due to my common system prompts) and around 900 tokens/second of output.

I think this worked out to around $1/million tokens of output and orders of magnitude less for input tokens, and before reserved instances or other providers were considered.

AaronFriel··on After 6 years, I'm over GraphQL
This is a vague recollection, but I seem to recall Meta/Facebook engineers on HN having said they have a tool that allows engineers to author SQL or ORM-like queries on the frontend and close to where the data is used, but a compiler or post-processor turns that into an endpoint. The bundled frontend code is never given an open-ended SQL or GraphQL interface.

And perhaps not coincidentally, React introduced "server actions" as a mechanism that is very similar to that. Engineers can author what looks, ostensibly, like frontend code, merely splitting the "client" side and "server" side into separate annotated functions, and the React bundler splits those into client code, a server API handler, and transforms the client function call into the annotated server function into an HTTP API call.

Having used it for a bit it's really nice, and it doesn't result in yielding so much control to a very complex technology stack (GraphQL batchers, resolvers, etc. etc.)

AaronFriel··on Gemini Flash
The PagedAttention paper is a good starting point as it represents the first major open source inference engine that had "pretty good" batch performance for transformers.

https://arxiv.org/pdf/2309.06180

AaronFriel··on Gemini Flash
The attention mechanism is vastly more efficient to train when it can attend to larger, more meaningful tokens. For inference servers, a significant amount of memory goes into the KV cache, and as you note, to build up the embedding through attention would then require correlating far more tokens, each of which is "less meaningful".

I think we may get to this point eventually, in the limit we will want multimodal LLMs that understand images and sounds down to the pixel and frequency, and it seems like for text, too, we will eventually want that as well.

AaronFriel··on RFC 9562: Universally Unique IDentifiers (May 2024)
"Only" 10k machines producing a combined 100 billion transactions per second is pretty hard to imagine, least of all that would all be producing transactions that are part of the same namespace. Virtually all UUIDs are meaningless outside of a particular system in which they were created.

There is a solution that doesn't require extending UUIDs (which has a storage cost everyone pays), which is to use a URI/URN instead of a UUID to provide a namespace. In practice this already occurs, except the namespace (scheme, path) containing the UUID is implicit, as it hasn't been named.

AaronFriel··on RFC 9562: Universally Unique IDentifiers (May 2024)
True, as universally unique identifiers, 128 (less a few) bits is not enough. You're talking about humanity generating 505 exabytes per year of just UUIDs. That won't happen any time soon.
AaronFriel··on GPT-4o
It has only been a little over one year since GPT-4 was announced, and it was at the time the largest and most expensive model ever trained. It might still be.

Perhaps it's worth taking a beat and looking at the incredible progress in that year, and acknowledge that whatever's next is probably "still cooking".

Even Meta is still baking their 400B parameter model.

AaronFriel··on RFC 9562: Universally Unique IDentifiers (May 2024)
Care to share your math? My understanding of the birthday paradox is that it is astoundingly unlikely.
AaronFriel··on Conical Slicing: A different angle of 3D printing
Very cool project, breaking the assumption that the nozzle must lay down material while moving on the same plane as the bed is a great innovation.

One nit though about the website: I found it very distracting that the images kept cycling between slides, and by the time I was halfway down the article, I realized that they were cycling even when off-screen.

AaronFriel··on IBM nearing a buyout deal for HashiCorp, source says
Pulumi is open source, and committed to Apache 2.0!

We liken our model to git and GitHub, building the best service for the Pulumi engine and related tools.

https://www.pulumi.com/blog/pulumi-hearts-opensource/

AaronFriel··on Show HN: I made a multiple runtime version manager that can be used on Windows
For runtimes with tooling like .NET has, it's a niche use case.

At $dayjob, we have to support non-current compilers/SDKs for five different languages and runtimes where our tool will invoke (shell out, often) the CLIs for those tools. When triaging a bug report, it's great to have a version manager to use exactly the customer's version of the CLI.

Likewise, we need to make sure all of our examples and templates build even if the user has an old version, and the surest way to validate that is to hide newer CLIs and tools, and to test on the range of binaries that a customer could have installed.

AaronFriel··on German state ditches Microsoft for Linux and LibreOffice
I've seen this story every few years, perhaps this is just how they Munich negotiates its contract with Microsoft.
AaronFriel··on What even is a JSON number?
I'd strongly recommend against this default - it's a major blocker for using the Haskell library with web APIs as it transforms JSON RPC into into readily available denial of service attacks.

8 billion digits (~100 bits?) is far more than should be used.

Would it possible to use const generics to expose a `BigDecimal<N>` or `BigDecimal<MinExp, MaxExp, Precision>` type with bounded precision for serde, and disallow this unsafe `BigDecimal` entirely?

If not, I expect BigDecimal will be flagged in a CVE in the near future for causing a denial of service.

AaronFriel··on LLaMA now goes faster on CPUs
30gb+, yeah. You can't get by streaming the model's parameters: NVMe isn't fast enough. Consumer GPUs and Apple Silicon processors boast memory bandwidths in the hundreds of gigabytes per second.

To a first order approximation, LLMs are bandwidth constrained. We can estimate single batch throughput as Memory Bandwidth / (Active Parameters * Parameter Size).

An 8-bit quantized Llama 2 70B conveniently uses 70GiB of VRAM (and then some, let's ignore that.) The M3 Max with 96GiB of VRAM and 300GiB/s bandwidth would have a peak throughput around 4.2 tokens per second.

Quantized models trade reduced quality for lower VRAM requirements and may also offer higher throughput with optimized kernels, largely as a consequence of transfering less data from VRAM into the GPU die for each parameter.

Mixture of Expert models reduce active parameters for higher throughput, but disk is still far too slow to page in layers.

AaronFriel··on LLMs use a surprisingly simple mechanism to retrieve some stored knowledge
It indeed is. An attention mechanism's key and value matrices grow linearly with context length. With PagedAttention[1], we could imagine an external service providing context. The hard part is the how, of course. We can't load our entire database in every conversation, and I suspect there are also challenges around training (perhaps addressed via LandmarkAttention[2]) and building a service efficiently retrieve additional key-value matrices.

The external service vector database may require tight timings necessary to avoid stalling LLMs. To manage 20-50 tokens/sec, answers must arrive within 50-20ms.

And we cannot do this in real-time, pausing the transformer when a layer produces a query vector stalls the batch, so we need a way to predict queries (or embeddings) several tokens ahead of where they'd be useful and inject the context in when it's needed, and to know when to page it out.

[1] https://arxiv.org/abs/2309.06180

[2] https://arxiv.org/abs/2305.16300

AaronFriel··on Show HN: Dropflow, a CSS layout engine for node or <canvas>
Unfortunately not for nested inline nodes, like spans of text with formatting. For a lot of uses, that will be OK - but for rendering say, markdown text, Satori won't work.

The upstream layout engine handles flexbox layout, and it's unclear if Facebook needs inline layout or if Vercel would pick it up and close the gap: https://github.com/facebook/yoga

Then again, for the main purpose Satori is advertised for - generating URL unfurl previews - Dropflow looks like it might be the answer.

AaronFriel··on Show HN: Dropflow, a CSS layout engine for node or <canvas>
Does this depend on a browser? It looks like it doesn't - which is pretty impressive!
← PreviousPage 2 of 34Next →