I'm happy
I'm happy
If you're using a customised harness you should make sure you don't have something that's e.g. changing your system prompt on some requests or rewriting history - it can be tempting to do stuff like strip old thinking tokens or compact tool call results to reduce context size but it's a trap - you want to never change history because of how cheap cache is, even more so with deepseek because their cache hit pricing is so low compared to most other models.
I assume it's done as load balancing/latency mitigation, but it's put me off of OpenRouter for my use cases (limited use, limited need for changing models).
This prompted me to dig back through the logs, it seems to be closer to 96% at its absolute worst and 99.3 at best.
I only sell to openrouter, would much rather support Deepseek directly.
The absurd goal was to be able to simulate all of the active traffic in NYC, so 3-500,000 cars, without using any of the macro flow corner cutting that you see commercially or academically.
Initially thought that one machine was not going to be enough to do this in realtime so I spun off a very big fork and built out a webtransport stack to split the effort over a local network. It was promising until I also wanted to cover the highway in thousands of giant beach balls [1] [2].
In the process of leveraging codegen to shrink the data that needed to be relayed by >10000x (packing and dynamically updating sparse continous arrays of floats), threw in a fuckton of LOD work (both spatial & temporal), statistical aggregation, and a lot of differential equation bullshit to derive LUTs. It's at the point where a base model M2 mini can handle much more than that by itself. All of the networking effort paid large dividends in cross thread coordination and lock-free data passing. Went from struggling to fit each of the sim kernels for just 32 cars @ 60hz (~11ms) to 0.03, 0.0003ms p99s with the corresponding jumps in car count (north of 32k/thread). Multiplayer works well enough, it falls apart where it should (~128 people or LLMs driving around in the same square quarter mile, and a great deal more NPCs, latency permitting)
The remaining work is making it look and sound cool as fuck. Building out a physically-based audio synthesis engine that simulates the pulses of exhaust gas starting from the cylinder count/size/firing order, intake and exhaust count and size, header and exhaust configuration, and another one for the tire sounds, and another one for the collisions, and then rendering those out to wavetables so it scales and frees up the cycles for occlusion. And then getting it to visually render and control performantly in a browser tab, lots of instancing and shader work.
There's also some bullshit cooking that uses the motion estimation built into GPU video compression engines for ... other purposes at stupid low latency. I figured that's what Waymo had to be doing so I let it rip
It's fucking nuts, I'm having so much fun :) It will be done when it's finished
[1]: https://i.imgur.com/BDQSuLv.png [2]: https://i.imgur.com/CpreOWT.png
You can rack up quite a lot of tokens if you ask it to try out a lot of things, eg for performance investigations and trying out optimisation ideas.
Two premier sources: (1) look at good contributions someone already tried to make, but that got stuck in review or were otherwise abandoned. (2) look at user reported bugs and see if we can find a user reported bugs, and see if we can reproduce and fix cheaply.
Both are explicitly scoped as best-effort affairs: move on, if you can't quickly make progress.
Most of the work I have to do as a human is review and navigating the submission process: tokens are cheap these days, so you really need to make sure the contribution is actually worth someone's time to review.
I'm not sure if you think that's a lot, but that was barely even 8 hours. I've had 200B+ months lol