1,131 karma · joined March 27, 2019
Another thing to keep in mind is that rideshare revenue in the US is extremely geographically concentrated in urban cores. This is why every AV company was targeting SF as their first city (excepting Waymo, which did some stuff in PHX). 'Hyperfocused expansion' probably looks a lot closer to tackling new, novel areas in different metro areas rather than, say, expanding down in to San Jose and the central valley.
These things, they take time.
They've clearly hit (or projections confidently show they'll hit) a point where each car is profitable. I worked in the space for a while - platform upgrades (new cars, sensors, etc) are planned out years in advance and are pretty complex processes. But generally, each upgrade was a massive decrease in cost per car. (usually 50% cheaper or more). So also possible they want to wait for the next platform transition.
F1 on the other hand was maybe the worst offender as far as literalism is concerned.
I'll admit that the way we were using the image tag was a little unusual, but still something that was imminently supported by a plain HTML image tag.
My point is more that Next is such a bizarre black box that things like this were a regular occurrence.
To clarify: yes, it was the next Image tag. The moment we switched to using a plain image tag it resolved itself.
Isn’t the entire aim of world models (at least, in this particular case) to learn a very high quality 3D representation from 2D video data? My point is if that you manage to train a navigable world model for a particular location, that model has managed to fit a very high quality 3D representation of that location. There’s lots of research dealing with NERFs that demonstrate how you can extract these 3D scenes as meshes once a model has managed to fit it. (NERFs are another great example of learning a high quality 3D representation from sparse 2D data.)
>That said, our belief is that model-imagined experiences are going to become a totally new form of storytelling, and that these experiences might not be free to be as weird and whacky as they could because of heuristics or limitations in existing 3D engines. This is our focus, and why the model is video-in and video-out.
There’s a lot of focus in the material on your site about the models learning physics by training on real world video - wouldn’t that imply that you’re trying to converge on a physically accurate world model? I imagine that would make weirdness and wackiness rather difficult
> To be clear, we don't yet know what shape these new experiences will take. I'm hoping we can avoid an awkward initial phase where these experiences resemble traditional game mechanics too much (although we have much to learn from them), and just fast-forward to enabling totally new experiences that just aren't feasible with existing technologies and budgets. Let's see!
I see! Do you have any ideas about the kinds of experiences that you would want to see or experience personally? For me it’s hard to imagine anything that substantially deviates from navigating and interacting with a 3D engine, especially given it seems like you want your world models to converge to be physically realistic. Maybe you could prompt it to warp to another scene?
Additionally, curious about what exactly the difference between the new mode of storytelling you’re describing and something like a crpg or visual novel is - is your hope that you can just bake absolutely everything into the world model instead of having to implement systems for dialogue/camera controls/rendering/everything else that’s difficult about working with a 3D engine?
In my experience, the best immersive theater experiences find very clever ways to make the atmosphere work. In Sleep No More and other Punchdrunk shows, all of the guests are given masquerade masks to wear, the venue is fogged, and the lighting is dim. The dim and foggy atmosphere hides stuff that would otherwise take you out of the dreamlike 1920s noir setting of the show - that the other guest walking next to you is wearing a graphic tee, for example. The masks cast the audience as a shuffling horde of vengeful spirits haunting the characters for the sins they commit throughout the show - so when you see a big crowd following a character, it doesn't instantly feel at odds with the setting the way I imagine seeing a 5 year old in a Pokemon t shirt on the Starcruiser would.
The way to make something that’s fun is to try to make something that’s fun over and over again until you’ve got it down. It’s not by obsessively reading what investors or people who are essentially glorified influencers say.
Also, I’m sorry dude, but your blog sounds exactly like every metaverse pitch I’ve ever seen. Down to the “ROBLOX, MINECRAFT, ETC” highlight at the end. This doesn’t make me feel like I’m reading something written by someone that cares about games - it makes me think the author reads a lot of VC substacks and played League of Legends for a while.
Even mentioning “GTM strategy” and “ideal customer profile” tells me you’re drinking a lot of Kool aid. Stop watching YC videos aimed at B2B SaaS founders and reading a16z blogposts and work a lot harder on this stuff if you ever want to show it publicly.
Also keep in mind, there’s probably literally 100 other companies with the same pitch, vibe, and idea that you have.
If I were you, I’d think deeply about whether this is actually something you’re passionate about working on or if you just want to start a tech startup and chase trends. If the former is true, I’d close this down and work on it for a lot longer before ever showing it publicly. If it’s the latter I’d pivot to an LLM B2B SaaS company like everyone else like you and try your luck at that.
https://en.m.wikipedia.org/wiki/Ancient_Semitic-speaking_peo...
There’s an entire subgenre of YouTube channels that consist solely of creators updating videos promising that they have inside information on the creative conflicts at Lucasfilm/Amazon/etc, all of which happen to align perfectly with whatever the fandom is outraged with that week.
https://youtube.com/@mikezeroh
This guys channel is a great example - most of the channels discussing “female custodes Henry cavil warhammer 40k tv series” Amazon follow a similar format.
Edit:
I’d also add that I don’t think Rings of Power is bad because they cast minorities - most of the actors are fine, really. The plotting and pacing is just horrendous. In a show that has 5+ active plot lines and threads scattered all over the world, they’ve spent a quarter of their screen time on a plot line that’s completely disconnected not just from the lore but from the wider story being told and doesn’t look like it’s going to connect anytime soon. Which is funny, because I’d imagine Amazon execs felt that they were obligated to include that plot line (the Hobbit one) to appease viewers.
In one case, greentheonly realized some fraction of Tesla’s cars are shipped out of the factory still in dev mode, with debug mode enabled and increased privileges. He found someone with a car like this who was down to helped and swapped part of his cars hardware with their car, and from then on was able to get a much better view of what was running on his car.
Unfortunately twitter is awful to search and a lot of his info is buried deep in old threads, but a few (old) examples to illustrate that he regularly does this.
Visualizing the outputs of the models running in the car: https://x.com/greentheonly/status/1404164587927314435
Tesla’s dev debug menus circa 2020: https://x.com/greentheonly/status/1336467014727110656
It’s been a while since I’ve worked in the space so I haven’t followed green as closely.
If you want a grounded explanation of how Tesla’s stack works, follow @greentheonly on twitter. He’s a Tesla reverse engineer who regularly posts about the software that’s actually running on the car.
If you want an explanation about how real AV companies stacks work, I’d read Sebastian Thrun’s robotics textbook - then imagine what’s outlined in that book but with ML plugged in to a ton of spaces throughout the stack. This is also similar to how Tesla’s stack works, btw - greens just good to follow because a lot of people refuse to believe Tesla isn’t running some kind of “LLM but for driving” fully end to end black box model.
Tesla won’t launch a robotaxi anytime soon because they can’t use remote support or HD maps - although I think they’ve been stepping up their mapping efforts. Even the demo at Universal studios a few weeks ago was HD mapped - per @greentheonlys twitter.
I worked in the space for years and have seen the internal of both a traditional robotaxi company’s stack and Tesla’s.
Contrary to what other people in this thread are saying, the remote support isn’t remote direct driving of the car - essentially what will happen is that if the car finds itself in a situation where it’s unsure of how to proceed and it’s safe to stop, it will pause for a few seconds and wait for a remote operator to clarify a situation for it.
A good example might be road construction - if the car detects new road construction work that doesn’t match its map of the area, and its onboard systems determine that it’s not sure how to proceed through the construction with confidence, it will send what it thinks the top five likeliest ways to proceed to a remote operator. The operator then selects the proper path (or says that none of them are proper). The car will then follow the path presented by the operator, but actual driving behavior /collision detection / pathfinding is still determined locally. Think of it like ordering a unit around in StarCraft.
It’s also possible they shifted direction because the long term vision of robotaxis is much more lucrative.