HNHacker News
TopNewBestAskShowJobs

vessenes

13,501 karma · joined October 23, 2009

I'm an investor, the founder of new alchemy, co-founder Lamina1, former CEO of coinlab.com, Chairman of the Bitcoin Foundation, Ethereum security researcher, holder of patent 9298806, general nerd. Currently working at Capital6, my private equity fund.

I’ve lost more than $1bn twice - failure is the best teacher! He says to himself..

peter@capital6.com

submissionscomments
vessenes··on Two American Airlines Flights End Up with the Same Flight Numbers
Reminds me of the excellent, and frequently surprising https://flightaware.engineering/falsehoods-programmers-belie..., which includes surprising false statements like:

“Everything that has an ICAO code is on Earth”, and “flights are at most a few days long”. The whole list makes you despair of structuring any data, ever.

vessenes··on Show HN: Our space game has a built-in RISC-V emulator that runs Linux
Ooh cool. Suggestion - you could set up deep space relays with very long latency to IPv6 routers. So there could be some outward facing connectivity, but it would just be, you know, slow. Until the player researches Ansible tech, obviously
vessenes··on Apple Pass Designer
I vibe coded hn10k earlier this year (with options for 1k and 100k). It’s more as you remember it. Significantly better than now actually.

If you lift comment threads that have participation from the karma cutoff you chose you still get some discoverability.

I was going to publish it but .. it took like three hours and that was nine months ago. But I recommend making one yourself if modern hn is getting you down. Ultimately I gave myself a time out which also worked :)

vessenes··on Frog and Toad and the Increasingly Capable Machines
This is delightful. Read it.

ALSO I believe I have found one of the sources of claude’s “load bearing” tic — the author Elizabeth Van Nostrand’s blog https://acesounderglass.com/2019/12/11/hows-that-epistemic-s... uses the phrase “load bearing facts” in a comprehensible way, and was written by someone rationalist adjacent writing in their own voice.

Seriously, this is big. I’m going to pester claude as to whether or not it’s copying Elizabeth.

vessenes··on Show HN: TinyAIArena watch AI agents battle it out
Fun! As we know from 2024, these models can play diplomacy relatively well. Might I suggest they are allowed to talk to each other between rounds? Maybe up to two statements and emoji response to a DM. I think you’ll see more interesting gameplay.
vessenes··on Contrastive Language Models
Ah-ha. Interesting! Thanks.
vessenes··on Contrastive Language Models
>"CLM-8B also sets a new SOTA on challenging agentic coding benchmarks, including DeepSWE (81.6%) and Terminal Bench 2.1 (87.6%).

I didn't see any details on this on the announce page. And I don't believe it. Astra x-high pass@1 on DeepSWE is 74% +/- 3%. (https://deepswe.datacurve.ai).

That said, love seeing some of these new architectures get people exploring. But, surely somebody is incorrect here inre: those numbers.

vessenes··on Show HN: JevBench, a reproducible benchmark for typed decision models
Oh I’m all for replicating. Suspicious is a big word, and to my mind fairly useless state of mind.

I just went ahead and built some stuff with Jev to get a feel for it, including a small chat harness — in this case the harness sends out like 40 parallel API calls to get a probability distribution against the 800 or so tokens in that call, and then combines up the likely ones and runs it through another narrowing process. With that in place, Jev can talk. Although it’s not very talkative, but it definitely can respond to queries.

I tried it out for some computer use usecases, and it has potential to be very fast there — it had enough comprehension to do the selecting and tool calling and pass back control to the harness at the right times.

So, upshot - useful and interesting tool. I didn’t benchmark it against any of the open jev clones because a) it’s cheap, b) I’m not using it for anything major right now and c) like I said above, I’d be surprised if that team just spent two years wasting time on a weekend project.

I’ll double down and say that if this arch turns out to be genuinely useful, (and I think it could be), then when we get good broad benchmarks, this release of Jev will benchmark higher against the weekend clones than it does now.

vessenes··on Show HN: JevBench, a reproducible benchmark for typed decision models
I think it’s much more likely that these benchmarks are not good, in that they do not explore much of the space that jev was (likely) trained to cover. The doom demo is a good example — I’d like to see a wide variety of things like that included in any benchmark, not just ‘email classification’ or what have you.

Think of it this way: there’s some time needed to optimize / design an architecture, and the world gets that for free when it’s described. As to the rest of the last two years spent, is it more likely a former oAI lead spent them fucking around, or adding as many RL environments as possible to its model that is supposed to be a generalized classifier?

Right now my prior is that jev is probably better than these rando weekend models, whether or not we know how to test and demonstrate that in a benchmark. It’s also super cheap, so I don’t think there’s a strong reason not to try out building with it first, then walk down the ladder to an open model if you need to for some reason.

vessenes··on Grok 4.7
I was going to say the reverse - claude has been the less satisfying normalized by benchmark for me in the last year. Both astra and fable have their quirks, but I am 90% codex this year up from 10% last year.
vessenes··on Grok 4.7
Nice to see this release cadence increasing and some continued improvement in quality. I am guessing these models are basically still outcomes of the cursor team integrating with the massive amount of compute they now own: I’d imagine we will see significant step up improvements with grok 5 later this year as the team gets more experienced and confident with larger training deployments. Here’s hoping for another competitive frontier model!
vessenes··on I Built Non-Autoregressive Decision Models with RL a Year Ago
Agreed. Another difficulty here is there are not good benchmarks for this new architecture yet, so it’s easy to potshot and snipe, where jev seems to be pretty broadly intelligent/at least have had a lot of rl in different domains.

We haven’t seen any of these copy cats play doom or street fighter for instance; just categorize email.

I imagine once the author cools down and evaluates on a broad harness of tasks he may find that his new thing has a lot of engineering work ahead.

vessenes··on I vibed a proof of Conway's conjecture
A counterpoint - I was told a story by one of my professors in the late 1990s, about one of his professors -- he'd written a thesis, gotten hired somewhere like Princeton, and taught there for a few years as Dr. <Somebody>. One day he received a letter pointing out a construction flaw in his thesis. He brought it to the department head who read the letter, and said "Well, Mr. Somebody, ..." Ultimately he fixed the proof.

Upshot, if there are real errors in published work, I think most mathematicians want to know about them.

vessenes··on AI recursive self-improvement might not come so quickly after all
Well, duh. If you could do this with Opus 4.8, we would know. When Astra’s successor is 2-3x better at math research, and the internal teams say “we believe we will get there,” I’m inclined to believe the insiders.
vessenes··on OpenAI's Sam Altman says it would be 'ill-advised' to go public in 2026
Sorry, I don't understand what this means. What does it mean?
vessenes··on OpenAI's Sam Altman says it would be 'ill-advised' to go public in 2026
I think that implies they're seeing unusual new subscription demand, yes? That's how I'd read it, not least because I worked through four resets this week on Astra, which is un unbelievable amount more inference than I've wanted from openAI really ever, and the highest ratio vs. claude since opus 4 at the very least, probably farther back.

I think there are few moats in the engineering use case, and a single new model can absolutely drive compute demand.

vessenes··on OpenAI's Sam Altman says it would be 'ill-advised' to go public in 2026
Prices are moving inference to highly profitable, oAI seems to have solved their training problems, and they're currently competing nicely with Anthropic on the coding side. I'd wait, too -- why fight this stuff out in public when you can stay private and have your big competitor deal with all the public company concerns? There's plenty of capital available in the private markets for them right now.
vessenes··on We must pace the frontier
“We” does a lot of work here. You keep using that word. I don’t think that word means what you think it means.
vessenes··on OpenAI have no mathematicians capable of understanding what they put out
Shenanigans is an inaccurate word; it implies underhanded behavior that's hidden / concealed. I think "to act so aggressively" is more balanced.

Here's my answer: If you think we're getting to AGI in the next 9 months, then you believe, with all your heart, that these problems will fall soon. However, there's an ocean to boil in terms of what you could point your limited clusters at. In the meantime, the market is desperate for any sign your company might be first to AGI. Therefore, news of tractability with current models might focus an organization intensely - internally they have a huge leg up on the public, and therefore it's minimal compute to check - and if they are successful, they get approximately $50 million of free PR, likely adding 10-20% to their valuation.

Likewise someone like Tristan is fighting for his (metaphorical) life right now, hoping to preserve his claims of primacy and have a shot at some of that prize money, despite being only partway to a full solution for N-S.

I don't think we see any behavior at all that isn't simple to understand and well described by the setup here, but tell me what you see differently.

vessenes··on OpenAI have no mathematicians capable of understanding what they put out
Yep, it's possible. But, we have literally no idea how much a set of prompts would impact training as far as general usefulness. I don't think we even know if Tristan's said he allowed training on his prompting or not. This is about money, ego, primacy, all the usual mathematician priority disputes.
vessenes··on OpenAI have no mathematicians capable of understanding what they put out
This is speculation right now. The idea would be that if someone used the product and granted training rights, which is the default for many subscription levels, then some knowledge would have been imparted into the general weights of the new model.

oAI has made clear they did not specifically pull in any user data to context for this run.

vessenes··on An Accidental Blackboard
Thanks! It's just infra, so you could do what you wanted. The clients include a safety reminder, and the datastructures include "unsafe" in the name of the inputs passed around, but that's just a little hygiene.

There really aren't good messaging libraries that provide what I wanted, so I built it. Basically I started with "signal but no need to have a phone number." I've used it to build a group messaging iOS app for friends, and just pass it to an agent all the time if I want them to be able to direct message.

The group API approval is modeled off of multi signature approval mechanics from Ethereum - so, you could use it like you describe: "Only allow this call if it passes safety checks from n reviewers", or you could have human in the loop, or a program that checks business rules + an agent and a human, etc. etc. I just wanted something that let us control API calls properly.

vessenes··on An Accidental Blackboard
I built what I think is a pretty good library and set of tools for this earlier this year -- https://github.com/corpollc/qntm (or `uvx qntm --help`); it includes a cli, python and typescript libraries, and works out of the box aimed at either a public endpoint, or a private one, depending on environment.

It's end to end encrypted, and has group messaging support, so if you wanted to read what the agents are saying you'd just add them to groups you're in. It also has a web ui. Version 0.6.0 should get pushed this evening pacific time, with some additional agent specific features.

Bug reports welcome! If I did a good job on architecture, you should be able to have your blackboard up tonight.

vessenes··on Tao: Open math problems being non-renewably mined by AI
Not arguing against open science - it's super valuable. I'm saying that pearl clutching by people reading Tao isn't useful, because it misses some long history which tells us that this kind of science has been seen as fundamentally competitive for millennia.

Should it be competitive? Is it more useful to be collaborative? How collaborative can it be when it's fundamentally competitive? Is it only fundamentally competitive because of some common 'quirks' of math types, or are there deeper forces pressuring it to be competitive?

These are all questions that I think are worth discussing, as is the note that the pendulum seems to be swinging away from cooperation in the face of competing for $trillion+ valuations (and a real enthusiasm for proving cool math stuff). The alternative, tweeting complaints on twitter without some context, is mostly a waste of space. I mentioned the history in hopes we could get informed complaints on twitter.

vessenes··on Tao: Open math problems being non-renewably mined by AI
That’s not untrue. But it’s also a misstatement of mathematical history. Many leading mathematicians historically have been highly competitive — Gauss comes to mind. Woe betide the lesser intellect that sent Gauss some ideas. The Newton Leibniz controversy was very serious business at the time in the UK and the continent. It was considered at the least a sin to reveal that sqrt(2) was irrational to those outside Pythagoras circle.

Mathematics has always been highly competitive.

vessenes··on On the Navier–Stokes Millennium Prize Problem
If those researchers did not opt out then training data might go in. I think it’s a courteous acknowledgement; as was reaching out and examining the direction of proofs themselves. At stake here is a particular mathematician dynamic - ego, prize money, and the sense of proprietary ownership that some might feel working on a problem.

All that was just kicked in the teeth by a group with a lot of compute that was like “bro I heard on twitter that Navier stokes could be solved. Let’s try it.” That’s an existential level of engagement that almost no mathematician in history would like.

vessenes··on An Alien Mind
I can’t read your original comment, but when the chief scientist of a company that just released six month old software that can use Kicad on your computer to design a circuit board and have it created and shipped to you tells you he thinks we will get to automatically self reinforcing improvements, I think it’s wise to take him seriously. It was only eighteen months before Astra was created that LLMs could not count rs in strawberry.
vessenes··on An Alien Mind
The situations aren’t equivalent - luckily in my opinion because the stakes with nuclear are much higher. von Neumann constructed a multinational game theory approach appropriate for weapons. AGI is a much harder problem to corral because there are so many benefits beyond just blowing up cities. But it’s also a much better thing to have for these very same reasons.

Similarly there have been few positive externalities from nuclear industry, making it easier to make the case to wind down research. This same set of concerns in biotech is much harder to get compliance with, precisely for this reason.

Anyway I’m especially wary of over analogizing to nuclear era concepts: I think they’re a trap.

vessenes··on An Alien Mind
I’m like a 2(.5?) there - I don’t think ASI will care about my kids better than I will for some definitions of better, for instance, and I feel very fuzzy and vague about what actual differences in qualia between me and ASI would yield in the wild.

I’m not a doomer, although I don’t think doomers are dumb, just wrong. I think you should design your systems around the possibility that people who disagree with you are correct , hence my nod to negative sum. If you have more than 30 years to live, I’d personally rep to the most likely outcomes being very positive. With a lot of disruption in the middle.

vessenes··on An Alien Mind
No. These parties are composed of people who most definitely think this way, and therefore will have distinct goals and interests when presented with opportunities. That’s reality quite aside from how a game theorist assesses the situation.
Page 1 of 34Next →