HNHacker News
TopNewBestAskShowJobs

AaronFriel

9,073 karma · joined July 23, 2012

Opinions are my own and not those of my employer, etc.

https://bsky.app/profile/aaronfriel.bsky.social

mayreply at my name dot com

submissionscomments
AaronFriel··on Pg_hint_plan: Force PostgreSQL to execute query plans the way you want
Postgres' query planner is likely considering random vs sequential page costs, and preferring an index scan on created_at.

What are the current values of `random_page_cost` and `seq_page_cost`?

    SHOW seq_page_cost;
    SHOW random_page_cost;
The default is typically 4, and in practice with modern disks you should use a lower value closer to 1.
AaronFriel··on Claude 3 model family
Oh, I see. That must be frustrating to folks at OpenAI. Their product rests on the quality of their models, and making users unable to see which results came from their best doesn't help.

FWIW, GPT-4 and GPT-4 Turbo via developer API call both seem to produce the result you expect.

AaronFriel··on Claude 3 model family
Are you aware you're using GPT-3 or weaker in those chats? The green icon indicates that you're using the first generation of ChatGPT models, and it is likely to be GPT-3.5 Turbo. I'm unsure but it's possible that it's an even further distilled or quantized optimization than is available via API.

Using GPT-4, I get the result I think you'd expect: https://chat.openai.com/share/da15f295-9c65-4aaf-9523-601bf4...

This is a good PSA that a lot of content out on the internet showing ChatGPT getting things wrong is the weaker model.

Green background OpenAI icon: GPT 3.5

Black or purple icon: GPT 4

GPT-4 Turbo, via API, did slightly better though perhaps just because it has more Drizzle knowledge in the training set, and skips the SQL command and instead suggests modifying only db.ts and page.tsx.

AaronFriel··on The Era of 1-bit LLMs: ternary parameters for cost-effective computing
In undergrad, some of us math majors would joke that there's really only three quantities: 0, 1, infinity.

So, do we need the -1, and/or would a 2.32 bit (5 state, or 6 with +/-0) LLM perform better than a 1.58 bit LLM?

AaronFriel··on Our next-generation model: Gemini 1.5
If you can dig in further - prompt engineer out - the prompt for minors that would be fascinating to report out.
AaronFriel··on Our next-generation model: Gemini 1.5
Of course, if you get exactly the answer you want in the first reply.
AaronFriel··on License University of Michigan's Databases of Academic Speech and Student Papers
Most university athletic programs are unprofitable, if examined solely through the lens of outlays vs direct revenue. It's of course harder to measure the consequences on recruitment, alumni attachment to the school, etc.

https://www.insidehighered.com/blogs/just-explain-it-me/shou...

> The 2020 report found only 25 Division I programs had revenues exceeding expenses. No Division II or III program had revenues exceeding expenses. There are 1,102 Division I, II and III schools.

I was active in budget committees as a student, and while my alma mater was in Division I, it was always in the red. It's a really hard decision for the student services committee every year to raise fees or cut funding, and the pressure from the athletics program to cut their budget last was intense. We tried to keep fee increases at or under the "higher education price index", but that itself is a flawed measure as it consistently is greater than CPI.

AaronFriel··on Our next-generation model: Gemini 1.5
There will always be more data that could be relevant than fits in a context window, and especially for multi-turn conversations, huge contexts incur huge costs.

GPT-4 Turbo, using its full 128k context, costs around $1.28 per API call.

At that pricing, 1m tokens is $10, and 10m tokens is an eye-watering $100 per API call.

Of course prices will go down, but the price advantage of working with less will remain.

AaronFriel··on Show HN: Integer Map Data Structure
Interesting! Reminds me a great deal of Judy Arrays: https://en.m.wikipedia.org/wiki/Judy_array

Judy Arrays are a radix trie with branching and a few node types designed to be cache line width optimized.

AaronFriel··on US developers can offer non-app store purchasing, Apple still collect commission
Those businesses operating at a loss are a tiny fraction of the "real economy", even though they may be giants in the future.
AaronFriel··on Benchmarks and comparison of LLM AI models and API hosting providers
Variance would be good, and I've also seen significant variance on "cold" request patterns, which may correspond to resources scaling up on the backend of providers.

Would be interesting to see request latency and throughput when API calls occur cold (first data point), and once per hour, minute, and per second with the first N samples dropped.

Also, at least with Azure OpenAI, the AI safety features (filtering & annotations) make a significant difference in time to first token.

AaronFriel··on Gvisor: Application Kernel for Containers
Substantially different.

Talos is a Linux operating system distribution tailored for running Kubernetes and container workloads. It runs runc, containerd, and other binaries, which spawn containers which themselves run at the same level of virtualization as the kernel. It's hardened for security. Containers that run in privileged mode can make syscalls that affect the host kernel.

gVisor is an OCI runtime - a tool for running containers, typically on Linux - implementation that presents a virtualized Linux kernel surface area to the applications and containers it runs. The applications have any syscalls they make intercepted by gVisor. It's not quite the same as hardware virtualization, but it reduces attack surface area by disallowing containers to make syscalls to the OS kernel.

---

Addenda:

gVisor is closer in spirit to Firecracker when used with Kata containers. Although the way the two work is very different, both effectively prevent containers from manipulating the Linux OS they run on.

Talos is closer, I think, to Bottlerocket OS, which is a Linux distribution created by Amazon for their container workloads. Bottlerocket is also a hardened, minimal operating system designed for running containers. Both of these are very similar to CoreOS aka Container Linux.

AaronFriel··on Fly Kubernetes
If there were a virtual kubelet per unit of granularity (datacenter, in their case?) then you would be able to use affinity rules just fine.
AaronFriel··on Paying Netflix $0.53/H, etc.
Important to remember that the distributions here are multi-modal, so averages don't represent the data well. For AAA games I think you're way off. Looking at the 5 top played games in 2023, "main story" to "completionist":

    Title                  Main Story (hrs)      Completionist (hrs)
    Baldur's Gate 3              57                    146
    Super Mario Wonder           10                     19
    Starfield                    22                    145
    Zelda: ToTK                  59                    235
    Hogwarts Legacy              27                     68
Of these, the first two are co-op but not the same kind of multiplayer I think you allude to. Super Mario Wonder is not, I think, what many people would call "AAA".

If we add Spider-Man 2, that's still a 17 hour main story, and has been criticized for being "too short". Many people who purchase AAA titles do expect 20-30 hours of entertainment, for a ~$2/hour return. And the Game of the Year is the title with the second best ROI!

AaronFriel··on What's Going on with Language Rankings?
RedMonk and many other language rankings have methodological issues when languages are very similar. Many people tag TypeScript questions as JavaScript, repositories with large amounts of JS checked in are considered JavaScript even if they have complete TypeScript typings (see: Webpack), and so-on.

I think this is one of the strongest signals that TypeScript is the dominant way JavaScript is written:

https://npmtrends.com/@types/react-vs-react

The type definitions for React are more popular than React itself. Why is that? Even users who aren't choosing to use using TypeScript are benefiting from their IDE installing typings on their behalf, at a rate large enough to exceed the CI/CD systems and users running "npm/yarn/pnpm install" and only installing "react".

It's like this for other packages as well; but the sheer popularity of react makes the point well. Many, many developers are using TypeScript even if indirectly, it's what makes their IDE light up.

AaronFriel··on The Linux Scheduler: A Decade of Wasted Cores (2016) [pdf]
The Linux 6.6 kernel ships with a new default scheduler[1]. Is the testing methodology in this paper relevant and used to assess the current (EEVDF) scheduler, or is it irrelevant?

[1] - https://lwn.net/Articles/925371/

AaronFriel··on Beeper Mini is back
If vendors were required to conform to standards, I suspect we would have had something like RCS a lot sooner.
AaronFriel··on Mistral: Our first AI endpoints are available in early access
Imagine a search engine that tries to persuade you a sponsored product is the answer to your problems, and it can use its knowledge about you to do so.

I think search products will be OK - it will just be stranger than we expect.

AaronFriel··on Apple Patent Shows GPU Dynamic Caching Has Been in Development for Years
Is macOS Dynamic Caching similar to Windows hardware-accelerated GPU Scheduling[1]? That feature purports improve latency and GPU scheduling efficiency. Is the feature here that the OS is delegating more scheduling to the M3's ASC coprocessor[2] - assuming it's similar to the M1?

It seems to me that every explanation of dynamic caching in terms of memory is "wrong" - as seen here and in several articles written by folks more familiar with PC hardware.

I think where some folks might get it wrong is thinking of Apple silicon as being like PC hardware, where VRAM and RAM are distinct pools of memory and using the CPU to move data between them (or other devices) is very inefficient. PCI bus attached devices having separate pools of memory has given rise to a plethora of technologies to allow directly read from RAM (DMA), or to enable a GPU to read from NVMe (DirectStorage on Windows), and so on.

The patent from the article seems to describe unified memory, part of the M1's architecture, not whatever "Dynamic Caching" is; but I'll admit Apple makes it a bit hard to understand what exactly is the case.

There's only one person I'd trust to describe what this feature actually is - I'll wait for Asahi Lina to break down what dynamic caching is and whether this is a hardware or OS feature.

[1] https://devblogs.microsoft.com/directx/hardware-accelerated-...

[2] https://asahilinux.org/2022/11/tales-of-the-m1-gpu/

AaronFriel··on X rolls out new ad format that can't be reported, blocked
On prominent accounts I follow, the top replies are universally all paid accounts posting clickbait, totally unrelated (or possibly AI generated?) to the poster. Here's a great example. I follow Jake Sherman, founder of Punchbowl News, to stay up to date on happenings in Congress. This is a reply:

https://twitter.com/3MTS_/status/1710360772612706416

All of that account's replies are spam like that. They're not even the worst offender, there are a handful of other accounts that just post emojis, complete nonsense, to every single popular account.

Musk had already destroyed the value of looking at a tweet's replies, which I consider the "accuracy" of the platform. But that was never the main draw anyway.

So far, my feed remains pretty good - the "recall". I've opted in to follow certain people or topics and truly, there's no better place to get the latest news. But if they destroy the value of the feed, there really will be no reason for anyone to stay.

AaronFriel··on Shaving 40% Off Google’s B-Tree Implementation with Go Generics (2022)
The article is about performance or CPU time, not code size. Per the first paragraph, "In this blog post I’m going to show how, using the generics, we got a 40% performance gain in an already well optimized package, the Google B-Tree implementation."
AaronFriel··on Fixing for loops in Go 1.22
I think that's an ahistorical reading of events. They did have the opportunity, but there were very few languages doing what Go was at the time it was designed. My recollection of the C# 3 to 5 and .NET 3 to 4.5 is a bit muddled, but it looks like the spec supports a different reading:

C# 3.0 in 2007 introduced arrow syntax. I believe this was primarily to support LINQ, and so users were typically creating closures as arguments to IEnumerable methods, not in a loop.

C# 4.0 in 2010 introduced Task<T> (by virtue of .NET 4), and with this it became much more likely users would create a closure in a loop. That's how users would add tasks to the task pool, after all, from a for loop.

C# 5.0 in 2012 fixes loop variable behavior.

I think the thesis I have is sound: language designers did not predict how loops and lightweight closures would interact to create error-prone code until (by and large) users encountered these issues.

AaronFriel··on Fixing for loops in Go 1.22
Ah, thanks for the correction! That's right. But:

1. Did JS introduce this change after Go's creation? Yes. (And also after C#.)

2. Did arrow functions support precede let and const support? Yes.

Answering the second question and finding the versions with support and their release dates answers the first question.

             Arrow functions    let and const
    Firefox    22 (2013)          44 (2016)
    Chrome     45 (2015)          49 (2016)
    Node.js     4 (2015)           6 (2016)
    Safari     10 (2016)          10 (2016)
This places it nearly 10 years after the creation of Go. And with the exception of Safari, arrow functions were available for months to years prior to let and const.

This is somewhat weak evidence for the thesis though; these features were part of the same specification (ES6/ES2015), but to understand the origin of "let" we also need to look at the proliferation of alternative languages such as Coffeescript. A fuller history of the JavaScript feature, and maybe some of the TC39 meeting minutes, might help us understand the order of operations here.

(I'd be remiss not to observe that this is almost an accident of "let" as well, there's no intrinsic reason it must behave like this in a loop, and some browsers chose to make "var" behave like "let". Let and const were originally introduced, I believe, to implement lexical scoping, not to change loop variable hoisting.)

AaronFriel··on Fixing for loops in Go 1.22
Oh, today I learned. I think this was an issue in Scala (with `var`), but this seems like a great compromise for Java core.

I suppose Java had many years after C#'s introduction of closures to reflect on what went well and what did not. Go, created in 2007, predates both languages having lightweight closures. Not surprising that they made the decision they did.

Your comment inspired me to ask what Rust does in this situation, but of course, they've opted for both a different "for" loop construct, but even if they hadn't, the borrow checker enforces a similar requirement as Java's effectively final lambda limitation.

AaronFriel··on Fixing for loops in Go 1.22
Many languages have made this mistake, despite having engineers and teams with many decades or centuries of total experience working on programming languages. Almost all languages have the loop variable semantics Go chose: C/C++, Java, C# (until 5.0), JavaScript (when using `var`), Python. Honestly: are there any C-like, imperative languages with for loops, that _don't_ behave like this?

That decision only becomes painful when capturing variables by reference becomes cheap and common; that is, when languages introduce lightweight closures (aka lambdas, anonymous functions, ...). Then the semantics of a for loop subtly change. Language designers have frequently implemented lightweight closures before realizing the risk, and then must make a difficult choice of whether to take a painful breaking change.

The Go team can be persuaded, it's just a tall order. And give them credit where credit is due: this is genuinely a significant, breaking change. It's the right change, but it's not an easy decision to make more than a decade into a language's usage.

That said, there may be a kernel of truth to what you're alluding to: that the Go team can be hard to persuade and has taken some principled (I would argue, wrong) positions. I'm tracking several Go bugs myself where I believe the Go standard library behaves incorrectly. But I don't think this situation is the right one to make this argument.

AaronFriel··on Fixing for loops in Go 1.22
The C# language team encountered this as well, after introducing lightweight closures in C# 4.0 it quickly became apparent that this was a footgun. Users almost always used loop variables incorrectly, and C# 5.0 made the breaking change.

Eric Lippert has a wonderful blog on the "why" from their perspective: https://ericlippert.com/2009/11/12/closing-over-the-loop-var...

I had a bit of trouble finding the original C# 5 announcement; that's hopefully not been lost in the (several?) blog migrations on the Microsoft domain since 2012.

AaronFriel··on Run LLMs at home, BitTorrent‑style
That's true for conventional fine-tuning, but is it the case for parameter efficient fine tuning and qLORA? My understanding is that for a N billion parameter model, fine tuning can occur with a slightly-less-than-N gigabyte of VRAM GPU.

For that 70B parameter model: an A100?

AaronFriel··on Bun v1.0.0
Do you have an empirical basis for that? I'm curious because on one hand, Bun is advertising they are solving a major challenge for library authors, really betting quite heavily on that being persuasive for adoption.

The risk for Bun's adoption is that they're wrong, of course.

The risk for Node's usage is that they're right, and library authors begin urging users to use Bun because it's 100x easier than making a useful library with Node. (This will happen very slowly, but things that happen very slowly can start happening all at once very quickly.)

AaronFriel··on Bun v1.0.0
How common are packages with no imports, exports, requires, or "module.exports="?

I imagine that's quite rare, because such a package would only be imported or required for its side effects; and because that package cannot import or require any others, could you safely just assume that if it was `require()`ed use CJS, and if it was `import`ed, use ESM?

AaronFriel··on Persimmon-8B
I hope this is only a slight tangent; since the authors talk about their model serving throughput and I hope I can get a gut-check on my understanding of the state-of-the-art of model serving.

The success of ChatGPT and my current work has had me thinking a lot about the "product" applications of large language models. I work at Pulumi on www.pulumi.com/ai; it's a GPT-3.5 and GPT-4 interface using retrieval augmented generation to generate Pulumi programs, and user experience is top of mind for me.

(Fingers crossed this doesn't hug our site to death here for the reasons I'm about to explain.)

To be blunt: I have found it surprisingly difficult to find the right tools to host models without dramatically worsening the UX. In theory we should be able to fine-tune a model against our own SDKs and synthetically generated code to improve the model's output and to guard against hallucination when retrieval fails. In practice, self-hosted model serving APIs have really poor time-to-first-token or even completely lack streaming behavior. It's a non-starter to build a product on something where a user has to sit and watch a spinner for a minute or more. I've been looking at the vLLM project with great interest, but haven't found much else.

---

For folks in MLops, deploying models with streaming APIs:

1. Is it mostly accurate that none of the model serving tools created prior to ChatGPT are great for streaming, interactive use cases?

2. How are you currently serving these models as an API and what upcoming tools are you exploring?

For the authors: How does your inference optimization compare to vLLM, or other tools using techniques such as continuous batching and paged attention?

← PreviousPage 3 of 34Next →