HNHacker News
TopNewBestAskShowJobs

armcat

1,131 karma · joined September 3, 2021

I build sh*t.
submissionscomments
armcat··on Modern Object Pascal Introduction for Programmers
This takes me back, my first adventure in programming was as a 12 year old creating a Pong clone using Turbo Pascal on DOS. It came in handy later as a 20 year old when I was doing freelance as a Delphi dev. Amazing to see Pascal on HN.
armcat··on Revealing the details of how OpenAI agents hacked Hugging Face
I got into computers in the 90s and back then hackers like Kevin Mitnick and Kevin Poulsen were all the rage. They all faced the law and prison sentences. What's weird now is that we have something between gross negligence and malice, and nothing is being done, except maybe coordinated consolidation of AI power under the guise of "safety".
armcat··on 2DWillNeverDie
That's a great game. Another I played recently is REPLACED: https://store.steampowered.com/app/1663850/REPLACED/
armcat··on Jev in 25 Lines of Python
Haven't tried recent frontier ones, Astra for example doesn't support logprobs emission on the API, and Sol and Luna supposedly support it with reasoning disabled. Haven't tried local models like Qwen 3.8 27b (I'm actually exploring their thinking trace, it's a lot of fun)
armcat··on 2DWillNeverDie
There is also a 2.5D paradigm. I made a small demo of a 2.5D where I generated pure 2D pixelart sprites and embedded them in a Blender+Unity 3D level. Personally I find this mix very exciting.

Demo: https://acatovic.github.io/afterlight-play/ Tool I used to create 2D sprites: https://github.com/acatovic/ai-game-studio

armcat··on Dutch governments builds alternative for Microsoft based on NixOS
The link doesn't work for me.
armcat··on Jev in 25 Lines of Python
Looking at the logprobs on tokens works for the local models, but not on the frontier ones. It's been more or less broken since GPT-4o for example. I wrote about it two years ago: https://medium.com/data-science/9-11-or-9-9-which-one-is-hig.... Also, I've done some work in estimating confidence and on rubric evals using the same method, and you actually get better correlation to "real confidence" by just getting the LLM to say it.
armcat··on Grim Fandango Puzzle Document (1996) [pdf]
I would sometimes just let the game idle at the casino level and just listen to the mellow jazz. That game was a piece of art.
armcat··on Anthropic partnering with Accenture on embedded evaluation
I'm seeing lot of negative sentiment around this on X. I think the key is to establish independence. If Accenture stems to gain (financially, in some way) from Anthropic success - for example if they rely on Mythos class models to reduce costs or generate revenue - then this will obviously not fly. Then there is a more philosophical question - can independent evaluation be established at all? Since we are all exposed to these models, and some people REALLY like Claude models.
armcat··on Claude Code now reads AGENTS.md if there is no Claude.md
About time!
armcat··on Astra for Law
How do these use cases stack up (for real) in legal AI tools like Legora and Harvey?

Disclaimer: I used to work in legaltech, but not those two companies.

armcat··on Hacking OpenAI
Related: Mistral seems to also have been hacked, https://frenchbreaches.com/blog/mistral-ai-de-nouveau-pirate... (NOTE: in French).
armcat··on We must pace the frontier
At this point, you are more likely to be stabbed to death, gunned down, or mowed down with a car, by any random psycho. For the nasty or disgruntled actors, they are today able to build bombs or chemical weapons without the use of AI, they have proven this time and time again. Japanese PM Shinzo Abe was assassinated with a home made gun.
armcat··on We must pace the frontier
(1) Are guns open source? (2) Far more people - order of magnitude higher - have access to guns than they do to open source software - I count access as in actually being able to do something with it.
armcat··on We must pace the frontier
It's interesting that the default thinking is that no one on the planet can be trusted except a privileged few. Event Karpathy has gone this way: https://x.com/karpathy/status/2098811935114551617

You can always open source and follow the example from Linux and all the amazing things that came out of the open source community. This is the only way to reach true equilibrium globally, where for every misalignment you have equal and opposite effort working on the alignment.

armcat··on ChatGPT Is Throwing 404
It’s Muse Spark 1.3 taking out the competition
armcat··on The efficient frontier of LLM inference
I think we can soon include "recursive depth" strategy that Astra is employing, which (I suspect) is using recursive internal state changes in the transformer as opposed to full forward-pass + sampling which has traditionally been the case with thinking/CoT. Similar method was used here (but different context - encoding tools inside the transformer weights for fast execution): https://www.percepta.ai/blog/can-llms-be-computers
armcat··on Fastpotify
You had me at Winamp
armcat··on Understanding ChatGPT Work
It’s been part of the strategy from both OpenAI and Anthropic to split users into “devs” and “knowledge workers”. Hence Codex and Work (or Claude Code vs Cowork), and Chat is stuck in between. Codex can do everything Work can do and most non devs I know use Codex - from sales ppl doing weekly prioritisation of pipelines and customised email reach outs, to project managers using it as a living LLMWiki of all the projects and teams. In fact the biggest shift in business I’ve seen is the embrace of coding agents as defacto AI tool across knowledge workers.
armcat··on Hy4 preview
Less tokens. These models already overthink like crazy especially for complex tasks.
armcat··on Hy4 preview
That would be very interesting but only the open models allow you to see the reasoning trace
armcat··on Our decision on Cursor following its acquisition by SpaceX
You can still use it via API keys (answering my own question from earlier), https://x.com/thsottiaux/status/2093515916076343774?s=46
armcat··on Our decision on Cursor following its acquisition by SpaceX
Does this also apply to API endpoints, let’s say if you want to provide API keys in Cursor to Azure or Bedrock deployment of a GPT model?
armcat··on Our decision on Cursor following its acquisition by SpaceX
Codex cli is open source rust project (Claude Code is not). Whisper is an open weight ASR/STT model. GPT-OSS is an open weight LLM that can run locally and is still very capable. CLIP is an open source model used for image encoding. Just some of the open contributions from OpenAI
armcat··on GLM-5.3 is now open-weight
I am with you 100% but Microsoft (partially) open sourced MS-DOS (v1.25 and v2.0) in 2018, 37 years after its initial release.
armcat··on GLM-5.3 is now open-weight
What's very promising here is the number of tokens-vs-accuracy ratio. I am assuming their "output tokens" means tokens generated as part of thinking and any tool calls (what are referred to as "input tokens" from billing PoV by service providers). The Chinese models like Qwen3.8 and GLM 5.2 are insanely overthinking in our workloads (which are highly complex data analysis tasks). It's a factor of 3-4x over Opus and GPT models. Even with cheaper prices per 1M tokens, the cost ends up being higher, in some cases 2x. So this is very promising from GLM 5.3. Looking forward to trying it.
armcat··on Nvidia agrees to acquire Hugging Face for $13B
Don't see this happening. Both Nemo and TensorRT from Nvidia are heavily invested in low precision formats and higher inference throughput as well as low memory use. It's more about what happens in relation to GGUF, MLX and non CUDA-native frameworks/formats.
armcat··on Nvidia agrees to acquire Hugging Face for $13B
HF has been a close part of my ML/AI career, coinciding exactly when I moved into this space 10 years ago. There are lot of nuances here (if the deal goes through). Some people say it's a loss for EU sovereign AI but HF is technically an American corporation. On the positive note, the founders (Julien, Thomas and Clem - all French) stem to make significant amount of money, which they are likely to pour into a new frontier AI lab in Europe. So potentially it's a big win. For Nvidia this is a great strategic play as it potentially gains control of the "AI app store" and can influence the direction of Transformers, Diffusers, PEFT and various other HF libs inits favour. At the same time Nvidia has been (somewhat paradoxically) a dominant force in actual "open source" AI and has contributed significantly via Nemotron and various optimisers close to the metal. This would likely accelerate further. Nvidia has everything to gain from having a massive open model ecosystem, instead of a consolidated market consisting of 2-3 players. Will that have impact on MLX and other contributions? We'll see!
armcat··on Qwen3.8-Flash-Next
How is input token efficiency/verbosity on this model? Has anyone tried? GLM 5.2 was doing lot of turns and thinking piling up input tokens in the context (compared to Claude and GPT models). Then Qwen3.8-27B was 2x of that. Both delivered good output results but those cumulative input token costs were not cheap. Note this is on our specific business workloads. Genuinely interested in other people's experience (if you are able to try it out).
armcat··on More than half of adults in U.S. say they lack basic statistical understanding
Basic statistics and basic economics/finance are two subjects that are completely underdeveloped within the general population. Coincidentally those two effectively rule our entire lives. As an aside, STAT 200 looks like a solid stats course.
Page 1 of 7Next →