1,630 karma · joined October 13, 2013
It seems like a dumbed down reskin of Codex/ChatGPT Work but with the power-user features e.g. visibility/mentions removed. As a serious engineer why would I want that? Then the word agent becomes dot.
It also seems to be running a VM so the agent has its own computer.
I guess the intent was to pull together the Codex, claw and ChatGPT Work paradigms and simplify them?
I used to really like Windsurf. (Now Devin. Kind of? But also now Antigravity.) I still use it as my editor but haven't touched the agent for a while simply due to the rise of Codex.
They allow 3rd party sellers in their platform and in their warehouses
There's many reasons for this but they include:
Somebody will compete in every segment: building the taxis, operating them, and vertically integrating them. Put another way - some people will buy taxis in this model either way, so Tesla is incentivized to participate in this also. That helps them get economies of scale.
It also of course minimize risk. But not just the obvious kind. It also minimises risk of a niche competitor taking the buy to own robotaxi market and from that wedge becoming a substantial competitor.
But also like Amazon - operators should worry about Tesla taking the data they have about most successful routes and using that to compete directly in the must lucrative identified markets.
You need to clean, inspect, repair, insure, secure and charge the cars. To do so efficiently you will need to custom develop premises full of chargers and efficient charging and cleaning infrastructure.
That absolutely is operationally intense. Premises, permits, construction and then significant operations.
Also if Elon Musk has a demonstrable superpower it's the ability to raise capital. So they really could order 100,000 of these themselves.
If you are transforming anyway you're looking at virtually zero dev time cost and runtime cost when no probe is active of less than 1 ms.
This approach survives any environment I know of and has almost zero runtime cost.
How does the probe function technically. Inspector API I believe is unavailable on CloudFlare etc.
I once wrote something like this which could work on serverless platforms without the Inspector API. It used Typescript AST transforms to insert no-op listeners at every line, so they would dynamically eval or dump breakpoint style if a listenToLine parameter equalled their line, otherwise no-op. So trivial but not technically zero runtime cost.
How does it work? Using the NodeJs inspector API or other language equivalent to drop breakpoints? Those APIs are unavailable in many serverless environments and are challenging to use alongside bundlers.
Welcome to the club
How is it easier to sign up and manage a different service, implement a different API, etc.
And from the company side the fatal flaw is that these types of tools rely upon 1% of their users having huge spend. Nobody is going to be a huge spender here because it's easier to hand roll than navigate procurement on this (not to mention impossible to justify the spend, additional security/privacy risk, etc.)
It feels approximately impossible for this company to have large accounts.
In the tradition of boring software, even before LLMs it was much simpler to just use your existing tools and hand-roll. With LLMs I cannot fathom reaching for a product for something small like this.
Why would I not just use Codex directly?
The we wrote a bunch of prompts argument is kind of meh. That sort of thing has not only diminishing value with subsequent model releases but I actually believe will turn negative. The model will know better by default.
For example, initially giving the model some advice on code best practice was helpful. But now it's unhelpful because the model already knows best.
Although tbh the article makes a lot of this obvious and trivial change in syntax.
Example: for large Eloqua/Marketo/HubSpot emails we would previously make a planner which delegated the sections to their own call.
GPT5.6 can do the whole thing. The planner is unhelpful.
My suggestion: feature flag your complex implementations so you can rapidly contrast with and without it. (Or a formal eval suite if you have one).
If you prefer the simpler path, delete the old path.
Note: the challenge is making things compatible with these tools. Obviously generating html directly has been simple for ages.
(Source: mopsy.ai)
I actually started typing the same point that the chances are actually high because of train/eval overlap then realised you answered your own question with that same observation.
It is interesting though!
Perhaps in some way this means we should decide which eval set aligns best with our taste?
Back to the blog post. This is an excellent write up of an excellent technical achievement.
I have a lot of respect for the Cognition/Devin (always "Windsurf" to me) and Cursor teams.
I found it interesting - but justified - that they referred to themselves as a foundation lab rather than a dev tools company.
Source: Gary Tan, who can write more code than Jeff Dean and John Carmack.
If you want to go really fast, you put a Ralph Wiggum dungeon. It's where you orchestrate a team of Ralph Wiggum loops together using subagents to win the economic Darwin Award which companies are handing out for who can burn the most tokens.
Haiku = essentially phased out Sonnet = the Haiku use cases Opus = the new Sonnet class Fable = the new Opus class
If I am right, the other "5.0" models will be conspicuously absent, possibly even for a couple of months. (If Opus 5 follows soon and is even modestly better than 4.8 then I was wrong.)
Certainly if you compare it to another likely scenario where Vercel buys them and fast forward 2 years, it's plausible that a huge number of projects went one way or the other because of what the AIs defaulted to.
The agents already reach for Vite. When they reach for Vite it's very logical they will default to CloudFlare after. (Much like they will guide users to setup Vercel for NextJS).
This could be a $20m acquisition which will generate $billions from the increase in the agent equivalent of SEO.
But it's also possible they haven't spent much of that money.
The investors don't need to be happy. They just need to be made whole (assuming they have a minority control).
It could literally be that only $2m ever got spent and that's been paid back.
It could also be that when literally nobody said they would pay for Vite+ the investors and team in general lost confidence and were actually very happy just to get their money back and pivot into this acquisition.
The article didn't mention what happens to paying Vite+ users. Is that because there basically aren't any?
I know a lot of people are/will build this. I would be specifically interested in Black Magic doing it first party.
Having said that, for all the AI features, the big one would be setting key frames etc. with an agent, driving the general editing workflow with text,etc. I realize this is non trivial but it's certainly viable for a team of this calibre.
I think if BM added a paid for agent which helped execute their traditional video editing tools (even if it "only" supported a subset) then that's a subscription a lot of people would be willing to pay for, especially as their core tool is so generous.
Anthropic capitalized upon a brief window of being more code-focused, which turned into enterprize contracts.
Then on renewal rug-pulled those same enterprises - going from your seat includes all the usage a user would reasonably need, to being you pay for the seat + all tokens at API pricing. (Which they raised by how many times in a year? I don't know the actual number.)
Revenue spikes like crazy through basically hostage taking made possible by Sonnet 3.5 era sentiment + enterprise purchasing lag.
Parlay the revenue spike into the valuation.
Crazy. Those same enterprises will get sticker shock and leave. Absurd short-term thinking.
OpenAI is the better company (transparency, open sourcing things, how they handle things in general e.g. OpenClaw, how they compete, etc.) and they have the vastly better brand, the better consumer presence, and (for me and many others) they have the better coding app + models.
Anthropic doing deeply customer hostile stuff - again and again - to produce a short term revenue spike does NOT make for a long-term sustainable business.
For such a young business to have such a long history of bait-and-switch is absolutely crazy. (Raising prices repeatedly, lowering rate-limits repeatedly, changing the terms, banning calls which contain "OpenClaw", turning on their IDE partners, turning on their enterprise partners.)
AFAICT anyone who's ever shown faith in Anthropic has been immediately exploited by them to some degree. They will quickly get the reputation of being "the Oracle of AI companies".
I wouldn't even value them at half of OpenAI.