My understanding is that code mode is supposed to be implemented by the harness, not the MCP provider. You chain multiple MCP providers as well as other harness provided tools inside the sandbox.
Can we talk about the point the article makes that the MCP servers should implement JSON apis instead of text APIs? It makes sense if your harness supports code mode, but an embedded code sandbox is a pretty big piece of machinery and not everyone implements it. It seems like a massive change to the way one should build an MCP. Also why again are we using MCPs instead of OpenAPI when it's now supposed to be just stateless JSON APIs?
You cannot introspect model training by prompting the model. Whatever answer it hallucinated on your query "where do you know that from" has almost certainly very low bearing on reality.
Personally I've found the opposite useful - I mostly use OAI models with OMP harness and I have a rule instructing it to summarize if any work was omitted or if it made any surprising changes. Sometimes the agent will forget to do some piece of work or invent a new bizarre way of doing things, but quite often it is able to self reflect on this at the end of the session.
The new GUI is... not great. But the UX old UI was downright awful. There were like 5 different paradigms how to use individual screens, not just layouts, completely different ways to navigate and interact with controls.
After couple tens of hours you memorized the arcane keyboard invocations to do what you wanted, but it never felt intuitive.
My understanding is that most OpenAI users are on a free tier. Secondary effect of this is that OpenAI free tier model capability (assuming Luna) is what what most users associate with frontier AI capability giving somewhat warped view to many people.
If all of these meshes use the same shader (which they do, its just PBR) they can be all drawn with a single multidraw indirect call.
Edit: thinking more about it, since the objects are small they could be rendered with mesh shaders achieving further perf improvements. Also fine grained decomposition helps with culling. IMO the real problem is that rendering in game engines does not map well to modern hardware. Unreal engine went full software rendering with nanite replacing the whole pipeline (which brings significant drawbacks) and everyone else is using graphics with techniques 15 years out of date which leads to this mismatch between what is possible with hardware and what is good practice with game engines.
There are several tiers to these services, some are selling real us phone number verifications at about 0.5usd/text while others are selling virtual phone number verifications at much cheaper. From some limited experience with the former, there is rarely if ever any problems with rejections.
You can set different compaction strategy, currently "Summarize in place and keep the current session", "Generate handoff and continue in a new session", "Drop heavy content in place, recover via artifact", Snapcompact as mentioned
I think that we are starting to see that API inference prices for US labs are excessive and subscription prices are closer to real costs so 200/m can be realistic longer term.
I've been measuring waiting for tool calls/waiting for model response in my OMP with Sol 5.6 and usually it's 85%-95% of time spent waiting for model to respond, so 14x speedup in model perf would still be very significant. YMMV but speeding up tests and improving DX is somewhat well understood.
The ethics group isn't at OpenAI to make it ethical, it is there so that it can be pointed to outsiders such as press with words "we take ethics very seriously". It is very similar in function to security group (which should be responsibility of everyone) and DEI group. Ethics head quiting has about the same importance as head of PR quiting.
I'm very torn on these laws. On one hand I understand the purpose, but on the other recoding phone calls and meetings is immensely useful. On Android I had an option to automatically record phone calls and it has saved my ass many times, while on IOS it's not possible at all.
Chromium is actually fairly efficient when shared across multiple applications. If the webview2, which is probably what the weather app uses, would not create a whole browser per application but rather just the renderer process and the rest shared system wide it would be a couple hundred megs at most.
Using OMP through ACP inside Zed is drastically worse experience than just running OMP inside Zed terminal. The ACP UI is dreadful and I had bunch of crashes even for first party Codex support.
In my experience personal projects are the greatest indicator of IC competence, especially for young people. You may not like it, but turns out that when you do a thing in your free time because you like it, you get better at the thing than the people that only do it because they have to.
Personally I really dislike when the agents generate super long composed shell commands because they are really hard to audit. ffmpeg I'd whitelist, but if it makes a mistake in some super long chained git command it can have pretty scary consequences.
That's really nice. It would be really good for game GUIs too where the situation is quite poor and would work well with underlays/overlays/worldspace UIs. That said while binary size may be around 10mb, it still baloons to 500mb at runtime for your TODO list example which is more than some electron apps.
That's really nice. Have you tested if it works well with longer and more detailed prompts? For example adding more whole product specs and so on. It would be nice to generate a design system from generated UI you like instead of recreating that UI directly.
If the frontier models will take as much money to train as they do now, there is no way the wealthy are able to afford their training just for their own consumption. Financing of this whole thing rests on the models being available to companies and consumers who are willing to pay astronomical (compared to other software) sums for it.