HNHacker News
TopNewBestAskShowJobs

AdamConwayIE

64 karma · joined August 22, 2022

adam@xda-developers.com
submissionscomments
AdamConwayIE··on The newest ESP32 can run Linux and it's getting close to a Raspberry Pi
Hey, article author here. No, it's not an AI article. I've been sick the past couple of days and overlooked that heading. What I meant by it was that peripheral support changes how the device positions itself compared to others on the market.

I'll fix it, and thanks for mentioning it, but I also wanted to add that the "AI article" accusation was unnecessary. I spent a lot of time going through Espressif's documentation and the datasheet, and working out how Sv32 works and differs from the "MMU" implementation Espressif typically touts. It just feels like a strange accusation to tack on to an otherwise fair point regarding the article's readability.

AdamConwayIE··on Hy4 preview
Likely something that was first made especially obvious by Chinese models and then became something worth optimizing for in English too.

Chinese can be extremely information-dense in token terms, though it depends on the tokenizer. Roughly speaking, you can pack more "meaning" into a short sequence than English often allows for. That's why "caveman" reasoning is a pretty good fit.

There's a difference between bolting caveman speak onto an existing model and training a model to reason that way, though. If you just force an existing model to be concise in outputs, you're artificially reducing its available reasoning steps and can possibly prevent useful exploration or verification. If it's trained specifically to use compressed reasoning, it can learn to represent the same intermediate ideas in fewer generated tokens, cutting the number of sequential inference steps without necessarily sacrificing the useful reasoning itself.

It's not so much inherently a Chinese-model trait, but Chinese models could definitely have helped demonstrate how effective very compressed reasoning traces can be.

There are few tests of this, but one example I thought was interesting was here: https://github.com/PastaPastaPasta/llm-chinese-english

I wouldn't say it was Chinese specifically that was emulated, but it got people thinking about tokenizers and representation efficiency, and how natural English is rather inefficient.

AdamConwayIE··on I gave Qwen 3.8 27B a reverse-engineering job and it finished in 30 minutes
Maybe so, but there were other elements that I've seen frontier models struggle with in the past, which was the perspective I had coming into this. It's the type of test I run frequently and this is the first small local model I've seen pull it off.

It had a very non-standard RSA key implementation that was obfuscated heavily. As well, it has an online license check at first run, and that part typically trips up most of the local models I've tried. I've been running this test for about a year now with different models, and it was the first I've seen not only figure out the RSA key implementation, but the first that didn't just give up once it saw the online license check. Even though it's only a first-time launch check.

That's why I call it one of the hardest, because in my experience, it has been. It's the first local model I've seen pull it off end-to-end. For some of the reverse engineering work that I've done with LLMs, none have been as consistent as this particular test at highlighting a model's failure in this domain.

I have access to Daybreak Blue and I'm approved for Anthropic's Cybersecurity program, so I might run the same test with both of those just to see, because it's been a while since I used a frontier model on this test. I imagine they'll make relatively light work of it, though, assuming it doesn't trip the relaxed guardrails.

AdamConwayIE··on I gave Qwen 3.8 27B a reverse-engineering job and it finished in 30 minutes
I added a line to address this, sorry it wasn't there before! It was Pi and only used Bash-based tools.
AdamConwayIE··on I gave Qwen 3.8 27B a reverse-engineering job and it finished in 30 minutes
Nope, it wasn't.
AdamConwayIE··on I gave Qwen 3.8 27B a reverse-engineering job and it finished in 30 minutes
Ah, my bad! This image came from our backend, used for an unrelated article. I selected it by mistake rather than inserting the actual image that I'd uploaded. I'm updating it, thanks for the heads up!

For what it's worth, that image couldn't have been related. The other screenshots all showed thinking traces, and Claude doesn't share those.

AdamConwayIE··on Japan tried to build an operating system for the world, the US intervened
Oh wow! Article writer here, can't believe I didn't come across this when I was researching. Kicking myself haha
AdamConwayIE··on Be skeptical of OpenAI's rogue hacker agent story
Yeah, I'm reminded of container escapes, VM escapes etc. There have been plenty in the past; VirtualBox E1000 (I think?) comes to mind from a few years ago. If we are to believe this model can find 0days, I'm on board with the idea it could do so in sandbox.

That's not to say I believe it outright, but people are being oddly dismissive and acting as if it's impossible to break out of a sandbox. Which we've seen time and time again that it absolutely can be.

AdamConwayIE··on Phoenix LiveView 1.2
I've been loving LiveView. Been using it for a project for a client recently and it's so... chill. I like it a lot.
AdamConwayIE··on Web Browsers on Video Game Consoles
I remember using Orb on the Wii way back then for media streaming. Good times!
AdamConwayIE··on Running local models on an M4 with 24GB memory
Had a very similar experience recently.

Built a basic authentication handler for this test just so it wouldn't be in the training data of either model. It had deliberately planted bugs. One was a hardcoded secret, another was a wrap-on-0xFFFFFFFF bug as a result of a malloc(length+1).

Qwen 3.6 found both, alongside two other issues I hadn't even considered, and the location of the magic value. GPT-5.4, though, missed the malloc issue (flagging memory exhaustion as the only risk), it missed a separate timing bug (it explicitly said the function was safe), and it hallucinated the location of the magic value. Qwen correctly identified the integer overflow. GPT-5.4 did not.

I then compared basic research between them using SearXNG for web search. For example, the current status of MTP in llama.cpp. Qwen 3.6 27B found the current PR, but flagged a related issue that shows the current implementation can be slower than just using a draft model right now. GPT-5.5 Thinking found the same PR, but didn't flag the downsides.

In a similar comparison, I asked both models how I should get started with ESPHome as a total beginner. ChatGPT suggested an ESP32-S3 and a BME280, which is... just not a good idea. It also talked about the ESP32-P4 not having Wi-Fi, and installing with HA or Docker. Meanwhile, Qwen3.6 27B said regular ESP32, DHT22, and mentioned HA, Docker, and pip as installation methods. While GPT was good, it was just throwing out jargon for a prompt that explicitly requested it for a beginner.

It kind of blew my mind that in all three of these, Qwen landed it better.

AdamConwayIE··on Wine 11 rewrites how Linux runs Windows games at kernel with massive speed gains
Hey, article author here!

I've been writing for nearly a decade, and I can assure you, all of this is human written. I've long been writing about the Linux kernel where it's been relevant to my coverage, and there are articles under my name talking about low-level technical aspects in drivers and kernels from as far back as 2017.

I get that it's hard to know what to trust out there given that Dead Internet Theory is beginning to feel like a reality, but comments like this can be quite upsetting after spending days researching and writing an article like this. I totally get criticism of the article itself, and I'm fine with that, but it feels as if people are too quick to jump on the "must be written by AI" bandwagon. I receive it, my colleagues receive it, and for the people who I know put in so much effort into their work, it can be upsetting to them as well.

As was mentioned in another thread, there were actually a couple of typos in this article when it went live. I cleaned those up once they were pointed out, but AI doesn't make typos. I get it to an extent; hostility and accusations of all kinds have been levied at writers for the years and years I've been in this industry writing long-form content and analysis. But with the proliferation of AI, that hostility has really ramped up over the last couple of years.

AdamConwayIE··on Claude Sonnet 4.6
There aren't really any of the typical benchmark suites targeting Codex 5.3 because it's still not in the API.

SWE bench for example creates a predictions file and evaluates the results in the harness. Without Codex 5.3 being in the API, it can't.

AdamConwayIE··on Ireland rolls out basic income scheme for artists
You don't have to wonder whether or not it returns value to the tax payer. The Irish government already monitored the pilot program for two years, publishing all of the details and findings. [1]

"The headline finding from this social CBA is that for every €1 of public money invested in the pilot, society received €1.39 in return"

This came about as a mixture of greater economic activity from participants, cultural impacts that saw public-facing artist activities increase, and improvements to wellbeing of participants that reduced their requirement for psychological interventions by the state. The state also predicts that the further roll-out of this program will benefit consumers with lower prices for artistic works, as there will be more supply overall.

The scheme has been quite popular here in Ireland. Given the history of Ireland when it comes to art (both in the sense of spoken and written word, and in other mediums), it makes sense to introduce a scheme like this to safeguard and uplift those who produce art.

[1] https://www.gov.ie/en/department-of-culture-communications-a...

AdamConwayIE··on Please stop using OpenClaw, formerly known as Moltbot
Article author here: you'd be surprised! XDA these days has quite a bit of mainstream outreach, and this article has been getting shared on some socials. Even saw it getting passed around on LinkedIn.
AdamConwayIE··on A history of ARM, part 1: Building the first chip (2022)
There's even an interview with Steve Furber, who co-designed it, where he talks about it. https://www.youtube.com/watch?v=1jOJl8gRPyQ&t=508s
AdamConwayIE··on Distillation makes AI models smaller and cheaper
People always forget that back when OpenAI accused DeepSeek of distillation, o1's reasoning process was locked down, with only short sentences shared with the user as it "thought." There was a paper published in November 2024 from Shanghai Jiao Tong University that outlined how one would distill information from o1[1], and it even says that they used "tens of thousands" of o1 distilled chains. Given that the primary evidence given for distillation, according to Bloomberg[2], was that a lot of data was sent from OpenAI developer accounts in China in late 2024, it's not impossible that this (and other projects like it) could also have been the cause of that.

The thing is, given the other advances that were outlined in the DeepSeek R1 paper, it's not as if DeepSeek needed to coast on OpenAI's work. The use of GRPO RL, not to mention the training time and resources that were required, is still incredibly impressive, no matter the source of the data. There's a lot that DeepSeek R1 can be credited with in the LLM space today, and it really did signify a number of breakthroughs all at once. Even their identification of naturally emergent CoT through RL was incredibly impressive, and led to it becoming commonplace across LLMs these days.[3]

It's clear that there are many talented researchers on their team (their approach to MoE with its expert segmentation and expert isolation is quite interesting), so it would seem strange that with all of that talent, they'd resort to distillation for knowledge gathering. I'm not saying that it didn't happen, it absolutely could have, but a lot of the accusations that came from OpenAI/Microsoft at the time seemed more like panic given the stock market's reaction rather than genuine accusations with evidence behind them... especially given we've not heard anything since then.

https://github.com/GAIR-NLP/O1-Journey https://www.bloomberg.com/news/articles/2025-01-29/microsoft... https://github.com/hkust-nlp/simpleRL-reason

AdamConwayIE··on The Google Pixel 6a highlights the problem with the U.S. phone market
Does it? The Pixel 6a's primary sensor has really been showing its age for a while now, and often struggles to contrast really dark spots with really bright spots. We talked a lot about that in our comparison of the Pixel 5 to the iPhone 13 Pro. It's part of why Google needed to upgrade the Pixel 6 Pro camera, and I wouldn't really say the 6a is a "top-tier" anymore. It's still really good, but there are a ton of phones that do way better nowadays.
AdamConwayIE··on The Google Pixel 6a highlights the problem with the U.S. phone market
Hey there! Author here.

This isn't a criticism of the Pixel 6a, per se. This article is not meant to be a critique of the 6a, but rather, is using it as a tool to illustrate an overall greater point. The reason you don't understand the criticism of the 6a is because it's not supposed to be criticism.

The phone costs a lot, and a lot more than most other devices that are in a similar boat of "mid-range". It has a lot of bells and whistles that normal consumers won't necessarily care for, because it doesn't matter how great the camera is when a lot of people are just using their phones for the likes of Instagram, Twitter, and Facebook. A Nothing Phone, Nord 2T etc will get 80% of the way there in the camera department after the photo is on social media, and most people won't care for the difference.

However, the primary argument of the article is not to critique the Pixel 6a. Far from it. The problem is how the reason it's considered good value is because of the US carrier market. This article is primarily taking aim at the US carrier market and not the Pixel 6a. It's just a tool being used to illustrate a point.