“Everything that has an ICAO code is on Earth”, and “flights are at most a few days long”. The whole list makes you despair of structuring any data, ever.
13,501 karma · joined October 23, 2009
I’ve lost more than $1bn twice - failure is the best teacher! He says to himself..
peter@capital6.com
“Everything that has an ICAO code is on Earth”, and “flights are at most a few days long”. The whole list makes you despair of structuring any data, ever.
If you lift comment threads that have participation from the karma cutoff you chose you still get some discoverability.
I was going to publish it but .. it took like three hours and that was nine months ago. But I recommend making one yourself if modern hn is getting you down. Ultimately I gave myself a time out which also worked :)
ALSO I believe I have found one of the sources of claude’s “load bearing” tic — the author Elizabeth Van Nostrand’s blog https://acesounderglass.com/2019/12/11/hows-that-epistemic-s... uses the phrase “load bearing facts” in a comprehensible way, and was written by someone rationalist adjacent writing in their own voice.
Seriously, this is big. I’m going to pester claude as to whether or not it’s copying Elizabeth.
I didn't see any details on this on the announce page. And I don't believe it. Astra x-high pass@1 on DeepSWE is 74% +/- 3%. (https://deepswe.datacurve.ai).
That said, love seeing some of these new architectures get people exploring. But, surely somebody is incorrect here inre: those numbers.
I just went ahead and built some stuff with Jev to get a feel for it, including a small chat harness — in this case the harness sends out like 40 parallel API calls to get a probability distribution against the 800 or so tokens in that call, and then combines up the likely ones and runs it through another narrowing process. With that in place, Jev can talk. Although it’s not very talkative, but it definitely can respond to queries.
I tried it out for some computer use usecases, and it has potential to be very fast there — it had enough comprehension to do the selecting and tool calling and pass back control to the harness at the right times.
So, upshot - useful and interesting tool. I didn’t benchmark it against any of the open jev clones because a) it’s cheap, b) I’m not using it for anything major right now and c) like I said above, I’d be surprised if that team just spent two years wasting time on a weekend project.
I’ll double down and say that if this arch turns out to be genuinely useful, (and I think it could be), then when we get good broad benchmarks, this release of Jev will benchmark higher against the weekend clones than it does now.
Think of it this way: there’s some time needed to optimize / design an architecture, and the world gets that for free when it’s described. As to the rest of the last two years spent, is it more likely a former oAI lead spent them fucking around, or adding as many RL environments as possible to its model that is supposed to be a generalized classifier?
Right now my prior is that jev is probably better than these rando weekend models, whether or not we know how to test and demonstrate that in a benchmark. It’s also super cheap, so I don’t think there’s a strong reason not to try out building with it first, then walk down the ladder to an open model if you need to for some reason.
We haven’t seen any of these copy cats play doom or street fighter for instance; just categorize email.
I imagine once the author cools down and evaluates on a broad harness of tasks he may find that his new thing has a lot of engineering work ahead.
Upshot, if there are real errors in published work, I think most mathematicians want to know about them.
I think there are few moats in the engineering use case, and a single new model can absolutely drive compute demand.
Here's my answer: If you think we're getting to AGI in the next 9 months, then you believe, with all your heart, that these problems will fall soon. However, there's an ocean to boil in terms of what you could point your limited clusters at. In the meantime, the market is desperate for any sign your company might be first to AGI. Therefore, news of tractability with current models might focus an organization intensely - internally they have a huge leg up on the public, and therefore it's minimal compute to check - and if they are successful, they get approximately $50 million of free PR, likely adding 10-20% to their valuation.
Likewise someone like Tristan is fighting for his (metaphorical) life right now, hoping to preserve his claims of primacy and have a shot at some of that prize money, despite being only partway to a full solution for N-S.
I don't think we see any behavior at all that isn't simple to understand and well described by the setup here, but tell me what you see differently.
oAI has made clear they did not specifically pull in any user data to context for this run.
There really aren't good messaging libraries that provide what I wanted, so I built it. Basically I started with "signal but no need to have a phone number." I've used it to build a group messaging iOS app for friends, and just pass it to an agent all the time if I want them to be able to direct message.
The group API approval is modeled off of multi signature approval mechanics from Ethereum - so, you could use it like you describe: "Only allow this call if it passes safety checks from n reviewers", or you could have human in the loop, or a program that checks business rules + an agent and a human, etc. etc. I just wanted something that let us control API calls properly.
It's end to end encrypted, and has group messaging support, so if you wanted to read what the agents are saying you'd just add them to groups you're in. It also has a web ui. Version 0.6.0 should get pushed this evening pacific time, with some additional agent specific features.
Bug reports welcome! If I did a good job on architecture, you should be able to have your blackboard up tonight.
Should it be competitive? Is it more useful to be collaborative? How collaborative can it be when it's fundamentally competitive? Is it only fundamentally competitive because of some common 'quirks' of math types, or are there deeper forces pressuring it to be competitive?
These are all questions that I think are worth discussing, as is the note that the pendulum seems to be swinging away from cooperation in the face of competing for $trillion+ valuations (and a real enthusiasm for proving cool math stuff). The alternative, tweeting complaints on twitter without some context, is mostly a waste of space. I mentioned the history in hopes we could get informed complaints on twitter.
Mathematics has always been highly competitive.
All that was just kicked in the teeth by a group with a lot of compute that was like “bro I heard on twitter that Navier stokes could be solved. Let’s try it.” That’s an existential level of engagement that almost no mathematician in history would like.
Similarly there have been few positive externalities from nuclear industry, making it easier to make the case to wind down research. This same set of concerns in biotech is much harder to get compliance with, precisely for this reason.
Anyway I’m especially wary of over analogizing to nuclear era concepts: I think they’re a trap.
I’m not a doomer, although I don’t think doomers are dumb, just wrong. I think you should design your systems around the possibility that people who disagree with you are correct , hence my nod to negative sum. If you have more than 30 years to live, I’d personally rep to the most likely outcomes being very positive. With a lot of disruption in the middle.