There is also a 2.5D paradigm. I made a small demo of a 2.5D where I generated pure 2D pixelart sprites and embedded them in a Blender+Unity 3D level. Personally I find this mix very exciting.
Looking at the logprobs on tokens works for the local models, but not on the frontier ones. It's been more or less broken since GPT-4o for example. I wrote about it two years ago: https://medium.com/data-science/9-11-or-9-9-which-one-is-hig.... Also, I've done some work in estimating confidence and on rubric evals using the same method, and you actually get better correlation to "real confidence" by just getting the LLM to say it.
I'm seeing lot of negative sentiment around this on X. I think the key is to establish independence. If Accenture stems to gain (financially, in some way) from Anthropic success - for example if they rely on Mythos class models to reduce costs or generate revenue - then this will obviously not fly. Then there is a more philosophical question - can independent evaluation be established at all? Since we are all exposed to these models, and some people REALLY like Claude models.
At this point, you are more likely to be stabbed to death, gunned down, or mowed down with a car, by any random psycho. For the nasty or disgruntled actors, they are today able to build bombs or chemical weapons without the use of AI, they have proven this time and time again. Japanese PM Shinzo Abe was assassinated with a home made gun.
(1) Are guns open source? (2) Far more people - order of magnitude higher - have access to guns than they do to open source software - I count access as in actually being able to do something with it.
You can always open source and follow the example from Linux and all the amazing things that came out of the open source community. This is the only way to reach true equilibrium globally, where for every misalignment you have equal and opposite effort working on the alignment.
I think we can soon include "recursive depth" strategy that Astra is employing, which (I suspect) is using recursive internal state changes in the transformer as opposed to full forward-pass + sampling which has traditionally been the case with thinking/CoT. Similar method was used here (but different context - encoding tools inside the transformer weights for fast execution): https://www.percepta.ai/blog/can-llms-be-computers
It’s been part of the strategy from both OpenAI and Anthropic to split users into “devs” and “knowledge workers”. Hence Codex and Work (or Claude Code vs Cowork), and Chat is stuck in between. Codex can do everything Work can do and most non devs I know use Codex - from sales ppl doing weekly prioritisation of pipelines and customised email reach outs, to project managers using it as a living LLMWiki of all the projects and teams. In fact the biggest shift in business I’ve seen is the embrace of coding agents as defacto AI tool across knowledge workers.
Codex cli is open source rust project (Claude Code is not). Whisper is an open weight ASR/STT model. GPT-OSS is an open weight LLM that can run locally and is still very capable. CLIP is an open source model used for image encoding. Just some of the open contributions from OpenAI
What's very promising here is the number of tokens-vs-accuracy ratio. I am assuming their "output tokens" means tokens generated as part of thinking and any tool calls (what are referred to as "input tokens" from billing PoV by service providers). The Chinese models like Qwen3.8 and GLM 5.2 are insanely overthinking in our workloads (which are highly complex data analysis tasks). It's a factor of 3-4x over Opus and GPT models. Even with cheaper prices per 1M tokens, the cost ends up being higher, in some cases 2x. So this is very promising from GLM 5.3. Looking forward to trying it.
Don't see this happening. Both Nemo and TensorRT from Nvidia are heavily invested in low precision formats and higher inference throughput as well as low memory use. It's more about what happens in relation to GGUF, MLX and non CUDA-native frameworks/formats.
HF has been a close part of my ML/AI career, coinciding exactly when I moved into this space 10 years ago. There are lot of nuances here (if the deal goes through). Some people say it's a loss for EU sovereign AI but HF is technically an American corporation. On the positive note, the founders (Julien, Thomas and Clem - all French) stem to make significant amount of money, which they are likely to pour into a new frontier AI lab in Europe. So potentially it's a big win. For Nvidia this is a great strategic play as it potentially gains control of the "AI app store" and can influence the direction of Transformers, Diffusers, PEFT and various other HF libs inits favour. At the same time Nvidia has been (somewhat paradoxically) a dominant force in actual "open source" AI and has contributed significantly via Nemotron and various optimisers close to the metal. This would likely accelerate further. Nvidia has everything to gain from having a massive open model ecosystem, instead of a consolidated market consisting of 2-3 players. Will that have impact on MLX and other contributions? We'll see!
How is input token efficiency/verbosity on this model? Has anyone tried? GLM 5.2 was doing lot of turns and thinking piling up input tokens in the context (compared to Claude and GPT models). Then Qwen3.8-27B was 2x of that. Both delivered good output results but those cumulative input token costs were not cheap. Note this is on our specific business workloads. Genuinely interested in other people's experience (if you are able to try it out).
Basic statistics and basic economics/finance are two subjects that are completely underdeveloped within the general population. Coincidentally those two effectively rule our entire lives. As an aside, STAT 200 looks like a solid stats course.
Same here. The absolute star in my team is the youngest. But! I think what's happening is that the distribution in this age group has shifted. Essentially the absolute best (high IQ, super ambitious, great communicator, cracked builder types) are now even better positioned than ever before. But the majority of mass in this age category now suffers, i.e. everyone that's "good" or "very good" but not exceptional/brilliant.
It's interesting because a study here in Sweden [1] produces identical results. To quote from the study:
"An event study documents an accelerating decline in employment
of 22–25-year-olds in high-AI-exposure occupations, reaching 5.5 per cent by
early 2025 relative to less exposed occupations within the same employers"
This is a great read and I can extend this argument really to any small business. Tax filing, VAT, bookkeeping, privacy, consumer, workplace and sectoral requirements frequently impose a disproportionately large cost on microbusinesses. EU is more and more split along the enterprise and VC backed line, and everyone in between has disproportionate costs. My barber in Stockholm often complains about this - the amount of money they would need to earn to overcome all these costs is staggering. So they often do lots of gigs on the side.
It's best not to take advice on direction of careers from anyone. It's better to find and work on things that interest you, and then take advice from people that are amazing in that specific field. In mid 2000s in Australia all the "top people" were telling me not to get into a software engineering career because it was dead. It's certainly challenged right now, but it took off during those 15+ years.