That was about my conclusion as well: Slightly less than GLM 5.3 performance but made in Europe. So, maybe it answers Tiananmen Square questions correctly, and in French. All in all, a reasonable model, but not frontier.
I sometimes switch mid-session. Depends what I’m doing. Sometimes a model is slow or it’s not giving me what I want and I don’t want to dump the context. But if there’s a natural break, I’ll start a new session with a new model.
Cache cost is the primary metric, imo. Most people look at output cost, but you only pay that once. Cache cost you pay every turn and it grows each turn as the context length grows.
Well, even ignoring Claude’s capabilities, you would expect human testers (whether Apple employees or external beta testers) to be creating bug reports for all these things. It’s embarrassing if your self-proclaimed bug-fix/stability release is buggy and unstable. If the bugs are filed, then Claude should be able to help fix them.
All the major streaming providers are also adding live content. Once they have the infrastructure built out and the customers, it’s relatively easy to add the live content.
Spot on. They discounted to encourage the transition away from cable and DVR (remember TiVo?) to streaming and then they raise prices to the magical “fraction of monthly budget for entertainment” level that everyone (cable, too) works up to. Every company is going to try to get you to give them all of what you’re willing to pay. When cancellations are greater than new subs, they’ll stop and wait for inflation and rising wages to catch up again. That’s the game. That’s always going to be the game.
Right, and I’d add that in some games there is actually a “win” state that serves as a binary indicator. Given this, you can actually view the model as the thing being classified into two states: (1) consistently wins the game and (2) doesn’t consistently win the game.
One problem we’re going to have with AI hardware is coming up with a standard set of specifications that are comparable. I don’t really care about the CPU GHz and the memory bandwidth, at least not directly. What I really want to know is how many tokens per second this will deliver, but that also depends on the model. We need a standard metric for that. Perhaps we agree on a specific open weight model (e.g. GLM 5.3 Flash or Qwen vWhatever) and then measure TPS on the hardware of interest.
Because a car killing somebody means there is someone (the owner, “driver,” manufacturer, etc.) for the victim’s family to sue. And while anyone can bring a lawsuit against anyone at any time, regardless of how frivolous, you can’t sue God for an act of God.
Yes, but those situations are irrelevant to this issue. The question is, who pays when your self-driving car injures or kills somebody? I assure you that the answer to that question is not “nobody.”
> If we want adoption, we would need it be placed in individual vehicles and it should look at removing driver liability.
Liability is the biggest issue holding back widespread adoption, IMO. The tech is pretty good already. Yet accidents are still going to happen. The question is m who pays when they do? The “driver”? The taxi operator? The car manufacturer? It’s all unclear at best right now.
The speed of GLM 5.3 Flash on OpenRouter seems to vary considerably by provider. Some are fast and some are slow. OpenRouter does provide some tuning knobs, but not enough for my taste. It’s also token-heavy with reasoning, though I found it better than Deepseek V4 Flash previously.
I think you’re going to be swimming against the tide on this as computer interfaces are increasingly infused with AI at all levels (OS and apps). IMO, it’s more natural to anthropomorphize. Photoshop in 2010 was just an app, a tool that you wielded. But Photoshop in 2030 is going to be a “digital artistic assistant” that collaborates with you to design whatever you need. In that new world, “I” is more natural. Whether you like that or want it, it seems like that’s where we’re headed.
This is one of the reasons I use Pi. Pi’s minimal system prompt avoids contradiction between what the harness writer thinks is best and what the user thinks is best. The user specifies what the user wants and that’s pretty much the end of it.