180 karma · joined June 20, 2019
One bad cycle and the next Opus 6 or GPT 7 might flop (think what happend to LLama4 or Gemini), and the user quickly switch to the next best thing. So it make sense to build your tooling to be model-agnostic.
alternative is to collect/sell usage data, which is icky, or provide something as a service (token usage, infra, plugins, etc)
The terminal multiplexer / multi-agent coding space is getting very crowded. YC alone has funded many competing startups in this space: herdr, Superset, cmux, Emdash, Orca, Bullet, Conductor (Conductor was in YC before pivoting from chat->coding). I’m probably missing a few, and that’s before counting companies outside YC such as Superlogical and Agentastic.dev (by yours truly).
One thing I find particularly interesting is the role of open source in this market.
A lot of these products started open source (i assume as a tactical way to gain traction and build a community). But after the early traction phase or raise funding, the incentives seem to change.
Being open source by itself is not a product differentiator in this market. The features required for string positioning (orchestration, cloud, custom agents, enterprise features, hosted infrastructure, etc.) often end up either closed source or available only through the hosted product.
That makes me wonder how much developers actually value OSS for this category. How much does it matter to you whether your IDE or coding environment is open source?
Is OSS important because you want to inspect the code, fork it, self-host it, avoid vendor lock-in, or simply because you trust open-source developer tools more? Or, for desktop software like IDEs is open source mostly a nice-to-have rather than a requirement? So I’m curious where HN lands on this
Well, i'm trying to do the later, and it is very hard to position the product and differentiate without vertical integration. When Warp owns both agent and the terminal, they can a) give terminal for free and charge on the agent, and b) lock the user in, so it is a good business but poor user experience.
on the other hand, i have to make sure every agent is working which more surface area to maintain, and i'm also limited to the common subset of features. Also it is much harder to convince folks to switch.
Lets try: I support Wezterm (and ghostty, and others) with 40+ different agents (Claude, Codex, etc) in agentastic.dev, would you give it a try?
OpenAI did not launch any hardware (yet), and I don't think their IPO pricing hinges on them launching an iPhone competitor. To me, it looks like Apple wants to launch a similar AI hardware and is slowing-down the competition to be the first-to-market.
Funny thing: they are barking at the wrong tree here.
Personally i think it's fair. You get to use this model as long as you don't sue Meta because of the model's weights or outputs.
Thank you for keeping the token furnace burning!
IIRC Doom Dark Age sold 3m? copies, so it would cost $10m in Unreal Engine fees. Doom Engine team likely cost more.
Microsoft problem is it is giving away games for free on gamepass, so it sees its engineers as cost center.
Apple M3 Ultra chip 819GB/s memory bandwidth
They have 96, 128, 256, and 512GB variants my friend.
And for a long time (pre-AI) youtube was the biggest load on Google's entire infra. The number i recall was ~30% of all Google's cpu utilization was for youtube, and google spent a lot of effort optimizing it.
For example, it writes the whole front-end twice, once is claude design and then later it has to read it again and re-implement it in code. Also a lot of stuff (e.g. Claude.md, skill files, etc) are not supported, and they have their own set of ui-design and design systems, which claude code doesn't support.
I think Claude Design is a wonderful product, i'm just pointing that it is an independent product to Claude Code, heck even today it works better with Codex than Claude Code (that is how i use it with my own browser-use agent, i tell it to browse the design in Claude.ai/design and re-implement the design, works much better than downloading the zip file and asking the model to implement in that way.
He also saw LLM would replace search before anyone else, and that is something to look at the Lamda or GPT-1's output and think: yeah this will answer all of our questions one day.
- The price is actually competitive. They want to compete with Meta's Orion. However the product is ... lacking. None of the demos actually show-cased the dual wavelength display. The current usecases are available today at Meta Rayban Display for more cheaper and more polish.
- Dual snapdragon, and i assume one is dedicated to CV alg / scene-understanding/slam, without an external puck, i wonder how the thermal performance would look like.
- Looks very ugly, C'mon. This is the same team that desigend the original spectacle? where is Evan?
- I liked the Los Angeles text on the side. well-done.
Standard approach for training MM-LLMs is we train the encoder first, there are O(2-10B) good images on the internet, so encoder needs to see each image O(10-100) times, that is O(100T) tokens, which is more than the entire pre-training budget for most runs. That is the reason we train the encoder separately (smaller model, 2B active vs 30B or 200B active LLM); there is nothing magical about training the encoder and LLM together, it is just more token-efficient to train the image modality first.
* in our experience, in our evals and codebase, 4.6 was a bad model. This is over 60k developers, so statistically significant.
In my experience, Opus 4.0 was fantastic, major jump from 3.7. it was creative, super slow and expensive, and would sometime forget what it was doing, but it was getting the job done.
4.1 they made it much faster, so a lot of infra improvements.
4.5 was the time it could work on longer task, didn't make a lot of obvious mistakes of 4.0, and i think this was about the time the opus went mainstream, and all of the anthropic's compute crisis began, so instead of making the model better they tried to optimize it to reduce cost instead.
4.6 was such a bad model, they switched to adaptive thinking and it had so many bugs. poor api design, benchmaxxed and poor real-world results. i switched back to 4.5.
4.7 they just fixed the bugs they added in 4.6. Better than 4.5.
haven't fully tested 4.8 yet.