This takes me back, my first adventure in programming was as a 12 year old creating a Pong clone using Turbo Pascal on DOS. It came in handy later as a 20 year old when I was doing freelance as a Delphi dev. Amazing to see Pascal on HN.
I got into computers in the 90s and back then hackers like Kevin Mitnick and Kevin Poulsen were all the rage. They all faced the law and prison sentences. What's weird now is that we have something between gross negligence and malice, and nothing is being done, except maybe coordinated consolidation of AI power under the guise of "safety".
Haven't tried recent frontier ones, Astra for example doesn't support logprobs emission on the API, and Sol and Luna supposedly support it with reasoning disabled. Haven't tried local models like Qwen 3.8 27b (I'm actually exploring their thinking trace, it's a lot of fun)
There is also a 2.5D paradigm. I made a small demo of a 2.5D where I generated pure 2D pixelart sprites and embedded them in a Blender+Unity 3D level. Personally I find this mix very exciting.
Looking at the logprobs on tokens works for the local models, but not on the frontier ones. It's been more or less broken since GPT-4o for example. I wrote about it two years ago: https://medium.com/data-science/9-11-or-9-9-which-one-is-hig.... Also, I've done some work in estimating confidence and on rubric evals using the same method, and you actually get better correlation to "real confidence" by just getting the LLM to say it.
I'm seeing lot of negative sentiment around this on X. I think the key is to establish independence. If Accenture stems to gain (financially, in some way) from Anthropic success - for example if they rely on Mythos class models to reduce costs or generate revenue - then this will obviously not fly. Then there is a more philosophical question - can independent evaluation be established at all? Since we are all exposed to these models, and some people REALLY like Claude models.
At this point, you are more likely to be stabbed to death, gunned down, or mowed down with a car, by any random psycho. For the nasty or disgruntled actors, they are today able to build bombs or chemical weapons without the use of AI, they have proven this time and time again. Japanese PM Shinzo Abe was assassinated with a home made gun.
(1) Are guns open source? (2) Far more people - order of magnitude higher - have access to guns than they do to open source software - I count access as in actually being able to do something with it.
You can always open source and follow the example from Linux and all the amazing things that came out of the open source community. This is the only way to reach true equilibrium globally, where for every misalignment you have equal and opposite effort working on the alignment.
I think we can soon include "recursive depth" strategy that Astra is employing, which (I suspect) is using recursive internal state changes in the transformer as opposed to full forward-pass + sampling which has traditionally been the case with thinking/CoT. Similar method was used here (but different context - encoding tools inside the transformer weights for fast execution): https://www.percepta.ai/blog/can-llms-be-computers
It’s been part of the strategy from both OpenAI and Anthropic to split users into “devs” and “knowledge workers”. Hence Codex and Work (or Claude Code vs Cowork), and Chat is stuck in between. Codex can do everything Work can do and most non devs I know use Codex - from sales ppl doing weekly prioritisation of pipelines and customised email reach outs, to project managers using it as a living LLMWiki of all the projects and teams. In fact the biggest shift in business I’ve seen is the embrace of coding agents as defacto AI tool across knowledge workers.
Codex cli is open source rust project (Claude Code is not). Whisper is an open weight ASR/STT model. GPT-OSS is an open weight LLM that can run locally and is still very capable. CLIP is an open source model used for image encoding. Just some of the open contributions from OpenAI
What's very promising here is the number of tokens-vs-accuracy ratio. I am assuming their "output tokens" means tokens generated as part of thinking and any tool calls (what are referred to as "input tokens" from billing PoV by service providers). The Chinese models like Qwen3.8 and GLM 5.2 are insanely overthinking in our workloads (which are highly complex data analysis tasks). It's a factor of 3-4x over Opus and GPT models. Even with cheaper prices per 1M tokens, the cost ends up being higher, in some cases 2x. So this is very promising from GLM 5.3. Looking forward to trying it.
Don't see this happening. Both Nemo and TensorRT from Nvidia are heavily invested in low precision formats and higher inference throughput as well as low memory use. It's more about what happens in relation to GGUF, MLX and non CUDA-native frameworks/formats.
HF has been a close part of my ML/AI career, coinciding exactly when I moved into this space 10 years ago. There are lot of nuances here (if the deal goes through). Some people say it's a loss for EU sovereign AI but HF is technically an American corporation. On the positive note, the founders (Julien, Thomas and Clem - all French) stem to make significant amount of money, which they are likely to pour into a new frontier AI lab in Europe. So potentially it's a big win. For Nvidia this is a great strategic play as it potentially gains control of the "AI app store" and can influence the direction of Transformers, Diffusers, PEFT and various other HF libs inits favour. At the same time Nvidia has been (somewhat paradoxically) a dominant force in actual "open source" AI and has contributed significantly via Nemotron and various optimisers close to the metal. This would likely accelerate further. Nvidia has everything to gain from having a massive open model ecosystem, instead of a consolidated market consisting of 2-3 players. Will that have impact on MLX and other contributions? We'll see!
How is input token efficiency/verbosity on this model? Has anyone tried? GLM 5.2 was doing lot of turns and thinking piling up input tokens in the context (compared to Claude and GPT models). Then Qwen3.8-27B was 2x of that. Both delivered good output results but those cumulative input token costs were not cheap. Note this is on our specific business workloads. Genuinely interested in other people's experience (if you are able to try it out).
Basic statistics and basic economics/finance are two subjects that are completely underdeveloped within the general population. Coincidentally those two effectively rule our entire lives. As an aside, STAT 200 looks like a solid stats course.