HNHacker News
TopNewBestAskShowJobs

armcat

1,374 karma · joined September 3, 2021

I build sh*t.
submissionscomments
armcat··on 2DWillNeverDie
There is also a 2.5D paradigm. I made a small demo of a 2.5D where I generated pure 2D pixelart sprites and embedded them in a Blender+Unity 3D level. Personally I find this mix very exciting.

Demo: https://acatovic.github.io/afterlight-play/ Tool I used to create 2D sprites: https://github.com/acatovic/ai-game-studio

armcat··on Dutch governments builds alternative for Microsoft based on NixOS
The link doesn't work for me.
armcat··on Jev in 25 Lines of Python
Looking at the logprobs on tokens works for the local models, but not on the frontier ones. It's been more or less broken since GPT-4o for example. I wrote about it two years ago: https://medium.com/data-science/9-11-or-9-9-which-one-is-hig.... Also, I've done some work in estimating confidence and on rubric evals using the same method, and you actually get better correlation to "real confidence" by just getting the LLM to say it.
armcat··on Grim Fandango Puzzle Document (1996) [pdf]
I would sometimes just let the game idle at the casino level and just listen to the mellow jazz. That game was a piece of art.
armcat··on Anthropic partnering with Accenture on embedded evaluation
I'm seeing lot of negative sentiment around this on X. I think the key is to establish independence. If Accenture stems to gain (financially, in some way) from Anthropic success - for example if they rely on Mythos class models to reduce costs or generate revenue - then this will obviously not fly. Then there is a more philosophical question - can independent evaluation be established at all? Since we are all exposed to these models, and some people REALLY like Claude models.
armcat··on Claude Code now reads AGENTS.md if there is no Claude.md
About time!
armcat··on Astra for Law
How do these use cases stack up (for real) in legal AI tools like Legora and Harvey?

Disclaimer: I used to work in legaltech, but not those two companies.

armcat··on Hacking OpenAI
Related: Mistral seems to also have been hacked, https://frenchbreaches.com/blog/mistral-ai-de-nouveau-pirate... (NOTE: in French).
armcat··on We must pace the frontier
At this point, you are more likely to be stabbed to death, gunned down, or mowed down with a car, by any random psycho. For the nasty or disgruntled actors, they are today able to build bombs or chemical weapons without the use of AI, they have proven this time and time again. Japanese PM Shinzo Abe was assassinated with a home made gun.
armcat··on We must pace the frontier
(1) Are guns open source? (2) Far more people - order of magnitude higher - have access to guns than they do to open source software - I count access as in actually being able to do something with it.
armcat··on We must pace the frontier
It's interesting that the default thinking is that no one on the planet can be trusted except a privileged few. Event Karpathy has gone this way: https://x.com/karpathy/status/2098811935114551617

You can always open source and follow the example from Linux and all the amazing things that came out of the open source community. This is the only way to reach true equilibrium globally, where for every misalignment you have equal and opposite effort working on the alignment.

armcat··on ChatGPT Is Throwing 404
It’s Muse Spark 1.3 taking out the competition
armcat··on The efficient frontier of LLM inference
I think we can soon include "recursive depth" strategy that Astra is employing, which (I suspect) is using recursive internal state changes in the transformer as opposed to full forward-pass + sampling which has traditionally been the case with thinking/CoT. Similar method was used here (but different context - encoding tools inside the transformer weights for fast execution): https://www.percepta.ai/blog/can-llms-be-computers
armcat··on Fastpotify
You had me at Winamp
armcat··on Understanding ChatGPT Work
It’s been part of the strategy from both OpenAI and Anthropic to split users into “devs” and “knowledge workers”. Hence Codex and Work (or Claude Code vs Cowork), and Chat is stuck in between. Codex can do everything Work can do and most non devs I know use Codex - from sales ppl doing weekly prioritisation of pipelines and customised email reach outs, to project managers using it as a living LLMWiki of all the projects and teams. In fact the biggest shift in business I’ve seen is the embrace of coding agents as defacto AI tool across knowledge workers.
armcat··on Hy4 preview
Less tokens. These models already overthink like crazy especially for complex tasks.
armcat··on Hy4 preview
That would be very interesting but only the open models allow you to see the reasoning trace
armcat··on Our decision on Cursor following its acquisition by SpaceX
You can still use it via API keys (answering my own question from earlier), https://x.com/thsottiaux/status/2093515916076343774?s=46
armcat··on Our decision on Cursor following its acquisition by SpaceX
Does this also apply to API endpoints, let’s say if you want to provide API keys in Cursor to Azure or Bedrock deployment of a GPT model?
armcat··on Our decision on Cursor following its acquisition by SpaceX
Codex cli is open source rust project (Claude Code is not). Whisper is an open weight ASR/STT model. GPT-OSS is an open weight LLM that can run locally and is still very capable. CLIP is an open source model used for image encoding. Just some of the open contributions from OpenAI
armcat··on GLM-5.3 is now open-weight
I am with you 100% but Microsoft (partially) open sourced MS-DOS (v1.25 and v2.0) in 2018, 37 years after its initial release.
armcat··on GLM-5.3 is now open-weight
What's very promising here is the number of tokens-vs-accuracy ratio. I am assuming their "output tokens" means tokens generated as part of thinking and any tool calls (what are referred to as "input tokens" from billing PoV by service providers). The Chinese models like Qwen3.8 and GLM 5.2 are insanely overthinking in our workloads (which are highly complex data analysis tasks). It's a factor of 3-4x over Opus and GPT models. Even with cheaper prices per 1M tokens, the cost ends up being higher, in some cases 2x. So this is very promising from GLM 5.3. Looking forward to trying it.
armcat··on Nvidia agrees to acquire Hugging Face for $13B
Don't see this happening. Both Nemo and TensorRT from Nvidia are heavily invested in low precision formats and higher inference throughput as well as low memory use. It's more about what happens in relation to GGUF, MLX and non CUDA-native frameworks/formats.
armcat··on Nvidia agrees to acquire Hugging Face for $13B
HF has been a close part of my ML/AI career, coinciding exactly when I moved into this space 10 years ago. There are lot of nuances here (if the deal goes through). Some people say it's a loss for EU sovereign AI but HF is technically an American corporation. On the positive note, the founders (Julien, Thomas and Clem - all French) stem to make significant amount of money, which they are likely to pour into a new frontier AI lab in Europe. So potentially it's a big win. For Nvidia this is a great strategic play as it potentially gains control of the "AI app store" and can influence the direction of Transformers, Diffusers, PEFT and various other HF libs inits favour. At the same time Nvidia has been (somewhat paradoxically) a dominant force in actual "open source" AI and has contributed significantly via Nemotron and various optimisers close to the metal. This would likely accelerate further. Nvidia has everything to gain from having a massive open model ecosystem, instead of a consolidated market consisting of 2-3 players. Will that have impact on MLX and other contributions? We'll see!
armcat··on Qwen3.8-Flash-Next
How is input token efficiency/verbosity on this model? Has anyone tried? GLM 5.2 was doing lot of turns and thinking piling up input tokens in the context (compared to Claude and GPT models). Then Qwen3.8-27B was 2x of that. Both delivered good output results but those cumulative input token costs were not cheap. Note this is on our specific business workloads. Genuinely interested in other people's experience (if you are able to try it out).
armcat··on More than half of adults in U.S. say they lack basic statistical understanding
Basic statistics and basic economics/finance are two subjects that are completely underdeveloped within the general population. Coincidentally those two effectively rule our entire lives. As an aside, STAT 200 looks like a solid stats course.
armcat··on AI is hitting entry-level jobs hardest, Stanford study finds
Same here. The absolute star in my team is the youngest. But! I think what's happening is that the distribution in this age group has shifted. Essentially the absolute best (high IQ, super ambitious, great communicator, cracked builder types) are now even better positioned than ever before. But the majority of mass in this age category now suffers, i.e. everyone that's "good" or "very good" but not exceptional/brilliant.
armcat··on AI is hitting entry-level jobs hardest, Stanford study finds
It's interesting because a study here in Sweden [1] produces identical results. To quote from the study:

"An event study documents an accelerating decline in employment of 22–25-year-olds in high-AI-exposure occupations, reaching 5.5 per cent by early 2025 relative to less exposed occupations within the same employers"

[1] Same Storm, Different Boats: Generative AI and the Age Gradient in Hiring, Lodefalk et al, https://www.oru.se/globalassets/oru-sv/institutioner/hh/work...

armcat··on How Europe is killing makers and micro-entrepreneurs
This is a great read and I can extend this argument really to any small business. Tax filing, VAT, bookkeeping, privacy, consumer, workplace and sectoral requirements frequently impose a disproportionately large cost on microbusinesses. EU is more and more split along the enterprise and VC backed line, and everyone in between has disproportionate costs. My barber in Stockholm often complains about this - the amount of money they would need to earn to overcome all these costs is staggering. So they often do lots of gigs on the side.
armcat··on I were 17, I'd learn how to build LLMs from scratch
It's best not to take advice on direction of careers from anyone. It's better to find and work on things that interest you, and then take advice from people that are amazing in that specific field. In mid 2000s in Australia all the "top people" were telling me not to get into a software engineering career because it was dead. It's certainly challenged right now, but it took off during those 15+ years.
← PreviousPage 2 of 8Next →