Konwinski Prize
andykonwinski.com
andykonwinski.com
The prizes scale with the model’s score; the total prize pool is between $100,000 and $1,225,000, depending on the top scores.
If your AI can do this, it's worth several orders of magnitude more. Just FYI.
If you are serious, you should put the funds in an escrow contract and announce a bounty.
There are many brilliant people who would work on this for you.
Also, you cut off the "from the benchmark" part; this doesn't expect it to solve any random Github issue, just the ones from the (presumably manually vetted and cleaned up) bench dataset.
Also, the prize doesn't require you to train a new foundational model, just that whatever you use is open weights or open source.
Theoretically, might be get away with a Llama3.3 (or any other model which you think makes sense) with a cleverly designed agentic system and a fresh codebase-understanding approach, with minimal compute cost.
(ok, probably not that easy, but just saying there's much more to AI coding that the underlying model)
I followed your link, but it doesn't seem to bear out upur assertion. The two numbers mentioned in the article are 176 mil and 612 mil. Mind you those weren't an estimate of cost, but rather an estimate to replace. Article is dated 2004, with an update in 2011.
Using the lines-of-code estimation it crossed a billion in 2010 - again to replace. That has no relation to what it did actually cost.
Getting from there to "tens of billions" seems a stretch. Assuming a bottom value in your estimate of 20 billion, and assuming a developer costs a million a year, that's 20 000 man-years of effort. Which implies something like 2000 people (very well paid people) working continuously for the last decade.
Which seems, well, unlikely.
So doesn't seem that unlikely based on your estimates.
Let's ignore the years before dotcom boom since the dev community was probably much smaller, and assume an average of 3500 contributors since.
That's 25 years * 3500 contributors on average * 200k salary (total employee cost, not take home) = $17.5b
Napkin math, but order of magnitude checks out.
Those two numbers are from the intro. The postscript and the updates at the end mention $1.4b and $3b respectively.
The real cost is probably impossible to calculate, but that order of magnitude is a reasonable estimate IMHO, and absolutely comparable, or even larger, than compute costs for SOTA LLMs
one of my goals is to inspire and honor those that work on open source AI. Those people tend to be motivated by things like impact and the excitement of being part of something big. i know that's how i always feel when i'm around Berkeley and get to meet or work with OG BSD hackers or the people who helped invent core internet protocols.
those people are doing this kind of OSS work and sharing it with the world anyway, without any cash prize. i think of this as a sort of thank you gift for them. and also a way to maybe convince a few people to explore that path who might not have otherwise.
In a perfect world this wouldn't be necessary, but in the current research environment where benchmarks are the primary currency and are usually taken at face value, more unbiased evals with known methodology but hidden tests are exactly what we need.
Also one reason why, for instance, I trust small but well-curated benchmarks such as Aider (https://aider.chat/docs/leaderboards/) or Wolfram (https://www.wolfram.com/llm-benchmarking-project/index.php.e...) over large, widely targeted, and increasingly saturated or gamed benchmarks such as LMSYS Arena or HumanEval.
Goodhart's law is thriving and it's our duty to fight it.
Also, I answered a bunch of questions yesterday on LocalLLaMA that people here might find interesting https://www.reddit.com/r/LocalLLaMA/comments/1hdfng5/ill_giv...
In reponse to my comment of "Realistically, an AI that can perform that well is worth a lot, lot more than $1M.", he said:
> yeah i agree. one of my goals is to inspire and honor those that work on open source AI.
> people who work on open source tend to be motivated by things like impact and the excitement of being part of something bigger than themselves - at least that's how i always feel when i'm around Berkeley and get to meet or work with OG BSD hackers and people who helped invent core internet protocols or the guys who invented RISC or more recently RISC-V
> those people are going to do this kind of OSS work and share it with the world anyway, without any cash prize. i think of this as a sort of thank you gift for them. and also a way to maybe convince a few people to explore that path who might not have otherwise.
But what I appreciate even more is that we keep pushing the bar for what an AI can/should be able to do. Excited to track this benchmark over time.
Likely worth at least $500M, even if his stake was smaller than some of the other co-founders.
Isn’t that the core issue in the first place?
Clearly the world’s nations aren’t guaranteed to share your views…
Who gets to define “protect the people”…?
Dodging the core issue over and over again isn’t going to lead anywhere…
1) since we are creating a contamination-free version of SWE-bench (i.e. scraping a new test set after submissions are frozen) it is guaranteed that agents in this contest can't "cheat", i.e., models can't have trained on the benchmark / agents cant memorize answers.
2) as a general rule in life, don't cheat on things (not that there aren't exceptions)