We Built a Video Rendering Engine by Lying to the Browser About What Time It Is
blog.replit.com
blog.replit.com
here is their Breakpoint 2007 demo, a 177 Kb executable including 3d assets and textures. https://www.youtube.com/watch?v=wqu_IpkOYBg
Some enterprise software also have it, mainly for testing and they have lint tools that check that you never use Date.now()
What you should do is put everything that was scheduled on a timeline (every setTimeout, setInterval, requestAnimationFrame), then "play" through it until you arrive at the next frame, rather than calling each setTimeout/setInterval callback only for each frame.
Also their main loop will let async code "escape" their control. You want to make sure the microtask queue is drained before actually capturing anything. If you don't care about performance, you can use something like await new Promise(resolve => setTimeout(resolve, 0)) for this (using the real setTimeout) before you capture your frame. Use the MessageChannel trick if you want to avoid the delay this causes.
For correctness you should also make sure to drain the queue before calling each of the setTimeout/setInterval callbacks.
I'm leaning towards that code being simplified, since they'd probably have noticed the breakage this causes. Or maybe, given that this is their business, their whole solution is vibe-coded and they have no idea why it's sometimes acting strange. Anyone taking bets?
And then you'd need to maintain the code so it works with future Chrome versions.
Just as I got it working AI came along which made it pointless then I realised when I got to the bottom that you cannot do it perfectly because it’s not a deterministic renderer. So I called the project to an end.
What? What does AI have to do with anything here?
https://source.chromium.org/chromium/chromium/src/+/main:com...
You give the agent a URL it records itself going through UX flows, give that video to a coding agent and you have quite a feature.
It works but only in a limited way there's lots of problems and caveats that come up.
I dropped it in the end partly because of all the problems and edge cases, partly because its a solution looking for a problem an AI essentially wipes out any demand for generating video in browsers.
I ended up writing code that modified chromium and grabbed the frames directly from deep in the heartof the rendering system.
It was a big technical challenge and a lot of fun but as I say, fairly pointless.
And there are other solutions that are arguably better - like recording video with OBS / the GPU nvenc engine / with a hardware video capture dongle and there's other ways too that are purely software in Linux that work extremely well.
You can see some of the results I got from my work here:
https://www.youtube.com/watch?v=1Tac2EvogjE
https://www.youtube.com/watch?v=ZwqMdi-oMoo
https://www.youtube.com/watch?v=6GXts_yNl6s
https://www.youtube.com/watch?v=KzFngReJ4ZI
https://www.youtube.com/watch?v=LA6VWZcDANk
In the end if you want to capture browser video - use OBS or ffmpeg with nvenc or something - all the fancy footwork isn’t needed.
https://www.amazon.com.au/AVerMedia-Streaming-Passthrough-Re...
Or use ffmpeg with nvenc it allows simultaneous capture of 12 sessions.
Toss away all the hard work futzing with the browser just put in one ffmpeg command.
On top you could use that technique to record at frame-rates higher than native. There's no reason why you shouldn't be able to redraw a basic page with some animations at a few hundred fps.
That is only because your view omits some other problems this solves/products this enables.
There is an incredible ecosystem of tools out the browser land, to create animation.
If you can capture frames from the browser you can render these animations as videos, with motion blur (render 2500 frame for a second of video, blend 100 frames each with a shutter function) to get 25fps with 100 motion blur samples (a number AfterEffects can't do, e.g).
Also you must understand that chrome is not a deterministic renderer. You cannot get the per frame control because it is fundamentally designed to get frames in front of the user fast.
They did some work around the concept of virtual time a few years ago with this sort of thing in mind and eventually dropped it.
Not sure what market you are talking about.
What I was talking about: people pay for motion graphics. LLMs are excellent at creating motion graphics from/around browser technology ...
Advertising is a huge market and motion graphics is everywhere in video/film-based advertising.
> Also you must understand that chrome is not a deterministic renderer. You cannot get the per frame control because it is fundamentally designed to get frames in front of the user fast.
It absoluetly deterministic if you control the input. There is no "add random number to X" in Chrome. The non-determinism is user inputs and time.
I know this because the company I work for did extensive tests around this last year. I was one of the people working on that part.
We looked into the same approach as replit. The only reason we gave up on it was product-related which changed our needs. Not because it is impossible (which, I guess, their blog post prooves).
An exact color from a customer's brand book? Complex animated typography? Forget it.
Check out my GH profile and who I work for. We're ex blockbuster VFX professionals. We use Ai everywhere. We know what is possible and what isn't.
The market we serve is huge. And every ad we create is bespoke. And the product the ads are for is unique.
A scratch on a rim of the left front wheel? When we do a turntable that scratch needs to be in every frame and the rim can not change number of spokes or the like (that's what latest gen models like Nano Banana v2 still do).
Show me a model that can do this level of detail. They barely manage now for a few seconds for special things like humans. Even there you may get subtle changes of eye or hair color. Anything that is not a human and needs to stay exacty the same each frame: good luck.
But let's assume the models were there today.
Still a dead end because: cost.
You know how much it costs to create a 10-15 second clip with a state-of-art/somewhat useful model on a high VRAM GPU instance vs such a clip that is rendered by a headless Chrome browser on a cheap, low RAM, CPU spot-instance?
We don't use Chrome (see previous post). We use a bespoke 2D/3D renderer with custom Ai/human-fed pipeline. But cricually, there are no final frames ever coming from Ai for now because of the reasons I mentioned at the top.
We're talking multiple orders of magnitude difference in cost. As of this writing, a 15sec clip w. Veo 2 costs about 7.50 USD. This needs to be two orders of magnitude cheaper to become viable.
tl;dr if we relied on Ai for this, our business would not exist.
It does not have deterministic frame control.
Just use nvenc or intel or AMD hardware video capture.
- no special framework. No library buy-in. Just a URL
- Advance clock. Fire callbacks. Capture. Repeat. Every frame is deterministic, every time.
- We render dozens of frames that nobody will ever see, just to keep Chrome's compositor from going stale.
- The fundamental insight that you could monkey-patch browser time APIs ... is genuinely clever
- Where we diverged
The whole post is like this, but these examples stand out immediately. We haven't quite collectively put a name on this style of writing yet, but anyone who uses these tools daily knows how to spot it immediately.
I'm okay with using LLMs as editors and even drafters, but it's a sign of laziness and carelessness when your entire post feels written by an LLM and the voice isn't your own.
It feels inauthentic and companies like replit should consider the impact on their brand before just letting people write these kind of phoned-in blog posts. Especially after the catastrophe that was the Cloudflare Matrix incident (which they later "edited" and never owned up to).
And the lede is buried at the very end: This is just a vibe-coded modification of https://github.com/Vinlic/WebVideoCreator, and instead of making their changes open source since they're "standing on the shoulders of giants", the modifications are now proprietary.
In the end, being an AI company is no excuse for bad writing.
Also yikes for the proprietary modifications. AI companies: "what's yours is mine, and what's mine is mine only"
I'm not even against using AI per se, but when something is obviously written in ChatGPTese I'm not going to read it if I don't have to.
what's the issue with this one? it sounds like something I might write, tbh.
"We X, just to keep Y from Z" and its variations are a pattern I've seen come up a lot.
See --virtual-time-budget
https://peter.sh/experiments/chromium-command-line-switches/
Lesson: if your going to do LLM assisted writing, say to it “make sure this has a distinct tone that consistent and clearly quite different”.
Short sentences. Plenty of newlines. Enumerate everything. Always.
> The page behind that URL might use framer-motion, plain CSS animations [...]
And the code example does something with css:
`await seekCSSAnimations(currentTime); // sync CSS`
But by faking the performance of your webpage, maybe you are lying to your potential users too?
I think you're missing the point of it a little. The "user" is someone who wants to watch a rendered video of the brower's display, but if it takes longer than one frame (where you read the word frame in this comment, think of a frame of video or film, not a browser "frame" like people used to make broken menus with) to actually draw the visual the browser will skip it.
Instead this appears to just tell the browser it's got plenty of time, keep drawing, and then capture the output when it's done.
It's not too different to how you'd do for example stop motion animation - you'd take a few minutes to pose each figure and set up the scene, trip the shutter, take a few more minutes to pose each figure for the next part of each movement, trip the shutter again, and so on. Say it took five minutes to set up and shoot each frame then one second of film would take an hour of solid work (assuming 12 frames per second, or "shooting on twos").
It's just saying "take all the time you want, show me it when it's done" and then worrying about making it into smooth video after the work is done.
While such a person might indeed exist, I think the more common situation is a vendor showing a demo of how a website might work. In that situation the consumer wants a realistic depiction of someone interacting with the site. Though of course for the user of the video service it might be very useful if the video hides all manner of performance issues.