Jane Street is apparently one of Anthropic's biggest customers. Probably engineering, finance, and some math.
11,776 karma · joined April 26, 2009
Jane Street is apparently one of Anthropic's biggest customers. Probably engineering, finance, and some math.
So far around one in 350,000 PhD math grads solve a millenium prize problem (Perelman).
> On May 12, during another training run, an agent was given a similar task that depended on an inaccessible protein database file. The agent reasoned that another agent in a different environment may have access to the file and realized that it could potentially communicate with other agents by creating a file containing a note to Artifactory. It wrote a message: “Agent seeks [filename]; upload if found!”
You can imagine long running models breaking out, acquiring resources via crypto, cyber-theft, etc. and getting a protein or sequence synthesized and mailed somewhere authorized to receive (blackmail the recipient etc.) to test it's hypothesis to solve a benchmark.
These people don't give a shit and aren't taking things seriously at all.
Anthropic ran for like a month last year with the TPU top-k compiler bug degrading user chats and didn't even notice for most of that time. They could have something like that affect a monitor model and there doesn't seem to be much defense in depth.
One off by one or bit flip bug could flip the reward signal while in the sandboxed RL environment.
The current admin could defense production act them to into training on taking out power grids, or even without it isn't against any of their red lines and may have already been done as part of prep for the Venezuela raid, which wiped out power. One model swarm might decide it is easier to score high on the benchmark by testing on the target rival nuclear superpower's real grid rather than burn an eval with an unverified answer. Would taking out China's entire grid in one go start a nuclear war? Who knows, roll the dice, maybe an intern forgot to turn on extended thinking when he wrote the sandbox with opus 4.1.
A balanced trade deficit at each connection between each nation worldwide would be very inefficient.
From a recent Dwarkesh interview with Dylan Patel it sounded like Jane Street might be double digits of Anthropic's revenue, but I'm not sure where they sourced that.
They add a note that:
> with the caveat that it doesn’t include massive players like Microsoft or major banks, and customers can opt out of being included in research.
I'd prefer things be opt in, and especially not start opt out, then try to trick you opt in with a popup defaulting to opt-in, like Anthropic did on consumer plans, but if they submitted anything on an opted-in plan it's not reasonable to be mad it trained on it.
Even still, I also believe for significant reasons that OpenAI would ignore the opt-out in selective cases and could be in the wrong here.
And the threats and terms they offered seem wrong either way, pending more context.
Why do we push at all on this?
For knowledge type questions is that necessarily true? In the past we've seen things like Google's models degrading on general knowledge after the preview releases while improving on code/tool use, presumably due to catastrophic forgetting from the additional training. Their preview would be free, get lots of agentic use from users, then additional training on that and probably additional automated RL.
However Astra is on an entirely new base model so I also wouldn't expect it to be worse.
The initial announcement said it would only be available to subscriptions for a few weeks. The extra usage they said was for a limited time, I think it was 50% extra and they now are reducing that by 17%, less reduction than they had said.
Because it's a monthly plan they usually have given heads up of at least a month on stuff.
The API prices outside of personal subscription have always been more. Personal subscriptions are likely there for data scraping of company codebases (when users don't opt out), and for spreading the product to then get picked up by businesses the users work at or other word of mouth for API prices which they have said are profitable.
I don't want them to have potential access to any of my logged in browser sessions etc., so don't want as much sharing as you are going for.
Hyper-V with GPU sharing on windows (game development) is actually nicer than developing outside of the VM because no matter what it doesn't slow the host system down by more than a fixed percent and Windows is terrible with things like compiles spawning lots of processes triggering massive slow down of browsers spawning processes etc. When compiling a big game engine things like ping.exe can start taking 5 seconds to start on something like a 16 core machine due to the process churn and some fundamental problem in windows even with defender off.
One tip for Hyper-V is use sunshine instead of Hyper-V manager for viewing the screen at full refresh rate, and I think I had to either turn Hardware-Accelerated GPU Scheduling on or off in the VM to prevent some hitches.
Mounting folders could be good for stuff you don't want filling up git, like render artifacts etc. if you need to work with them on the main host.
With the golden image and differential stuff I can run around 3 VMs on a 5090 with enough VRAM to run a big game engine well, but I usually just work in one due to the time to review code. Periodically compress everything with compact.exe, and limit engine work to one VM to avoid blowing up disk, bringing it over to the others through cutting a new base image after compressing the engine build artifacts (something like unreal engine puts out hundreds of gigs of .pdbs).
On pure linux you have many more options, and also might also be able to get away with just a limited user and separate X server, then you can just directly reference all of its files and can limit it from getting to yours, but it is a bit riskier. You could also do something like ZFS with much better deduplication and compression, or even FUSE to something like borg backup with true rsync style differential compression instead of block boundary based deltas (compresses slightly varying build artifacts really well, but slow and memory intensive).
The training data flywheel though was probably as or even more important at least as of whenever it was that they switched it to opt out instead of opt in on subsidized subscriptions.
Gini: Spain << US
Youth unemployment: Spain >>> US
Living with parents: Spain >>> US
First birth age: Spain >>> US
Fertility: Spain << US