1,253 karma · joined October 26, 2024
The marketing value of this to Anthropic (if Anthropic even cares, this might just be the Bun team selling past the close) is to show that such a rewrite is possible and delivers engineering value. If exactly this project is $800K today, it'll be $200K and then $80K soon, so it's not so important to the story that it's cheap, just that a big "cool" rewrite is possible, and delivers velocity to the buisness.
They are back to pre 2022 levels now! https://fred.stlouisfed.org/series/APU0000708111
I don’t need it to be optimal, just … handy as an option!
The last time I brought this up here, folks offered a bunch of options that don’t quite do this, and the best candidate was this 15 year old compiler project that is Intel specific!
Could some programming language nerd build this?
(While you are at it give me a clear idiomatic way to pay the cost to switch from array of structs to struct of arrays)
Be fun to ask Fable to write a search program to find more counter examples using only early grad theory to guide the search.
Every 6-12 months, give out $200K to the first model to hit a min threshold on a set of ~5-10 hard benchmarks (+ perhaps one secret benchmark) using a total of 16GB / 32GB / 64GB / 128GB of VRAM (at a min context length of 200K), then move the threshold up. Quantization etc. is dealers choice, it just needs to nail the benchmark on a reference machine by using exactly that much VRAM (no mapping to RAM / disk etc.)
You could crowdsource the funding, and cross subsidize by adding targeted prizes focused on corporate needs (the classic one is PDF processing benchmarks), and say that 25% of each corporate prize funding also flows into the general prize pool.
For a lot of these open-source model companies, it's less about the $s (though $200K is nothing to sneeze at), it's the clear recognition that helps their model efforts stand out, gain usage etc.
You need to generate a QR code then scan it from the signing mobile device, which opens a secret menu option to sign (fine just brings up a confirmation dialog).
Incredibly annoying but perhaps more secure vs the threat of randomly tapping at prompts
Sometimes it feels like what people want is to only serve websites and content to good normal users but not evil bad “scrapers” (because maybe maybe your content will be monetized in some nebulous way) but … you put your content up publicly on the web! That should be part of reasonable use!
EDIT: Lwn.net is perhaps not a fair target of my ire.
“There is also a desire to not impede the operation of legitimate search engines, the Internet Archive, and other such groups. Some sites may add explicit allowlists to, for example, give the dominant search engine access to the site. Such measures have the effect of further entrenching a monopoly that already serves us poorly and should be avoided. We have, thus far, succeeded in that.”
Is reasonable! Many others are not
I just wish it had a way for me to downlimit tool access (I love midsized local models, and I'd love to enable only 30% of the power for some usecases).
Wasn’t this the point of the web?
I can imagine Anthropic wanting to acquire Bun without the gimmicks.
If you let a modern LLM do even the first, they’d crush this specific benchmark.
What is interesting is understanding how LLMs are able to beat 70+% on this benchmark or getting some of the poorly framed questions right? Are they implicitly learning the test writers style? Are the solutions leaking into their training set?
Perhaps reassuring is that even Fable stalls out at ~72% (on the hidden set which OpenAI did not run this analysis on), so perhaps training on the bench is not happening in anything but the most indirect ways.
I care a lot because small open models can never learn idiosyncrasies like this, so I really want good ways to judge models fairly.
EDIT: Humm OpenAI is muddying the water a bit. Only 20%ish of problems are broken in ways that are unfair to the agent, 4-10% are broken in favorable ways, so the benchmark ceiling is probably closer to 80-85%
Has extreme Curtains for Zoosha (https://amp.knowyourmeme.com/memes/curtains-for-zoosha) energy