For my Figma needs, having Codex do computer use seems just as good as their mcp. I can tell it, “go download the assets for what I need and take a few screenshots for reference”.
6,579 karma · joined March 13, 2013
For my Figma needs, having Codex do computer use seems just as good as their mcp. I can tell it, “go download the assets for what I need and take a few screenshots for reference”.
- Maintaining a significant new moving part in the dev process. KMP updates, management, integration, futures awareness, this all takes non-zero time.
- If an iOS dev wants to work on a KMP project what's the ramp up time? How much less efficiently can they contribute compared to a pure native app?
- How much time is spent on problems, bugs, issues, related to KMP. This is guaranteed non-zero for any tool.
- Is there less benefit to KMP with coding agents? Previously with two native apps, a bug fix had to be done twice. Currently, say a fix in done on Android, it's quite easy to say "propose an equivalent fix for the iOS repo". I do not claim this removes all benefits of a single code base, just that some things are not as bad now.
So I’m sure they won’t be fixing it then.
They are often too verbose, or miss capturing an important concept or purpose, or add bullets for parts of the changes that no one cares about.
This comparison doesn’t even make sense.
Lots of papers have great results that don’t depend on the latest models.
However in this case it’s problematic:
- They specifically make claims about the state of “current LLMs”. o3 is not representative of this.
- They ask are LLMs capable of X and arrive at a negative result.
If their claim was LLM’s can write coherent sentences, and their conclusion was positive, then there would be no issue using old models because the end result would be factual.
However, when you have a negative result that makes a claim about the current state of all LLMs and the ones you were using are not current, by definition it draws the whole conclusion into question.
2. Chess seems to be a poor benchmark for generalized strategic reasoning. People who are good at it rely more on experience and deep domain expertise than on skills that generalize to make them experts at unrelated tasks.
3. The study sounds like proving humans will never fly because they don’t have wings. In reality, humans do fly, and Claude Fable would destroy any human at chess by coding a strong enough engine on the fly.
- Give Jev and LLM the same input
- Lock down both to approved/rejected/unknown (LLM restricts on decoding)
- Both can be wrong, but neither can hallucinate (invent an another option).
"Jev: New frontier model 40-400x cheaper and 20-200x faster"
I'm not the gatekeeper of who gets to call themselves a frontier model, but I don't think most people would count Jev in that group. It sounds false.
If their specific claims hold up, then it would make more sense to say something like:
"Advanced the speed/cost frontier for structured decisions"
The original title before it changed less than an hour ago was:
"Jev: New frontier model 40-400x cheaper and 20-200x faster"
I'm going to agree that was misleading.
And on the second point:
>>Also "can't hallucinate" seems wrong? Sure, it can't emit an invalid type, but it can still emit a completely wrong valid value.
>that is likely true of all ML! perhaps we could debate semantics, but I don't think it's fair to say a random forest "hallucinates" in the way LLMs do"
Also going to disagree here, and I don't think it's semantics.
Type safety is not factual correctness.
Any attributes or keywords relating to Objective-C or interop are not everyday baggage for most people.
For SwiftUI, attributes seem like as good a choice as anything else.
As a whole concurrency is a mishmash syntactic mess, but I agree with the direction and the result, they're trying to manage the improvements they keep adding over years of investment.
To buy presents for a family Christmas list Mom drives to Store A and Dad drives to Store B.
As more items get added to the list, they must decide who should drive to a new store location to buy the present. How can they minimize total driving distance while kids are randomly adding new items to their list?
The proof above guarantees its possible to never drive more than twice the mileage you would knowing all the items in advance.
The big news is this guarantee works for any number of drivers with any arrangement of gifts.
The algorithm to do this was already known, what we’ve learned is it’s not possible to do any better.
One way to think about competitive analysis is bulk discounts. In life we’re constantly having to choose between quantity and discount. We could buy 1 item for a higher price, or say quantity 5 or 10 to get better discounts. The problem comes when we don’t know in advance exactly how many we’re going to need.
What should be our strategy for choosing how many to buy, and whatever the strategy is how well does it compare with having perfect knowledge upfront?
I mean, I otherwise agree with the sentiment of your post but I feel like most people are rating him pretty well.
In your generalized example I think the concern is when the additional evaluation effectively becomes a replacement for CoT, where something like the coconut research could replace it completely.
However, I don’t think we’re anywhere close to that with Astra.
Looping transformers uses additional calculations (repeating layers) to generate a token.
Reasoning (in this context) is test time generation of multiple tokens that allow a model to have a scratch pad to refine its thoughts, chain of thought reasoning in other words.
Doing the former in no way means that you have to hide the latter.
Raschka is right in this post, The Information article was wrong. The Astra system card does concede reasoning traces are sometimes smaller, but this could be for a lot of reasons, including simple efficiency. And it absolutely doesn’t mean they are going away or completely obscured.
The Last Week in AI podcast from Sept 8 seems to have gotten this wrong as well. Jeremie Harris rages that OpenAI implemented latent reasoning, ala the coconut paper, which could potentially actually obscure reasoning traces. But for the life of me, I do not know how he arrived at this conclusion and see no evidence that this has happened in Astra.
For example, we don’t even know whether perfect White play can possibly overcome two correctly placed Black stones against perfect play.
Basically it was a bug in the wiki code. They transferred the POST form parameters to GET URL parameters, and wiki internally doesn’t distinguish between the two.
For example, it's not even close to Opus 5 on Terminal-bench 4.0, 19.1% vs. 51.8%.
“get inflated in anecdotes”
Just because one study has the number 34% in it, does not prove the numbers I gave were wrong and does not mean they are anecdotal.
This is something that’s been studied from many angles and there’s more than one number that’s defensible based on the context and dataset.
For example, even the study you cited says: “In a study of 344 patients by Montejo-Gonzalez et al, 58% of patients reported sexual dysfunction when physicians directly inquired”, and “14% with spontaneous reporting”.
In another example, a 1,000 patient study found 57.7–72.7% for common SSRIs. https://pubmed.ncbi.nlm.nih.gov/11229449/
The “scientific numbers” you allude to are not one value. It’s a body of research not one definitive number.
Taken a whole I don’t think anyone is being that mislead when it’s suggested more than half of people may experience these side effects.
Sexual effects alone hit about 50–70% people. Then you could face nausea, insomnia, profuse sweating when you’re still, emotional blunting, and weight gain.
I don’t want to discourage anyone, they can save lives. It is a net benefit for some people, but even then most people will experience side effects and wish they didn’t have to take them.
People like to reduce papers to a simple hot take, but the paper is more than that, offers viewpoints that are non-standard and speculation about new possibilities.
Fraud is not the reason it is happening. Even the people who are making the cuts in this case have clarified the purported reason and it has nothing to do with fraud.
Imagine if they had GPU resources of western labs.
They left Opus in and got beat in all but one benchmark.
Nothing wrong with trying to improve, but why the marketing games?
Instead of trying to say in the post you’re “closer” to frontier, first set a clear goal to beat the Chinese labs on price or performance and demonstrate it convincingly.
Then when your ready, come back and talk frontier without playing hide the model.
You’ve got to be kidding me. That one variable could make a huge difference in the results. I can’t understand why they would leave that out.
Dumbed down quantization?
No. Full intended inference weights preserved, so far so good.
Slow performance?
No again. Looks like you could get over 150 tokens/second.
Give up context window size?
Yes. Original model is trained for and served at 1M, this is 256k. A very practical tradeoff though. Codex is in this range, and quality does start to drop off toward the full size.