HNHacker News
TopNewBestAskShowJobs

arkmm

250 karma · joined April 15, 2020

submissionscomments
arkmm··on Show HN: I trained a 125M model to autocomplete piano on-device
Very cool! Can you say a little bit about the size of the DPO training examples and how long training took?
arkmm··on Claude: System Prompts
"Claude keeps responses focused, brief, and concise to avoid overwhelming the person."

Claude and I must have a different idea of what brief and concise mean.

arkmm··on Why does Opus 5 feel worse to work with?
Since reasoning tokens are just text, I think the models have learned to squeeze in some computation in their output writing as well. So they're incentivized to be correct but long-winded, as it gives them more time to think. It's kind of the equivalent of filler words for humans, except LLMs can actually word-vomit something intelligible.
arkmm··on Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
I'd also be really curious about the cost to run something like this, and what things you think it's particularly helpful for?
arkmm··on Ask HN: What are you working on? (August 2026)
Building a simple sandboxed way to run open weight models over a copy of all my personal data (email, docs, messaging, photos, etc). Think it'd be cool to have a self hosted AI "chief of staff" that has full read access to my personal info.
arkmm··on The Tower Keeps Rising
And the human has to explain it at 2 tokens/sec (average speaking speed ~130-150 wpm). That's another constraint against slop - humans need to be able to explain their code succinctly.
arkmm··on The Tower Keeps Rising
I think part of the problem is the context windows for humans are actually much smaller than what an LLM can keep track of today. The small context window of humans is a feature that forces modularity and abstraction in software engineering so that you can decompose what you're working on into something that can fit into your head. But since LLMs can fit so much more in their head, so to speak, they don't have this same incentive, and you get the unorganized mess of spaghetti code that current agents often produce.
arkmm··on Building an HTML-first site doubled our users overnight
Seems like LLMs embrace that last point as well.
arkmm··on Agent Memory: An Anatomy
everything is computer
arkmm··on Faster asin() was hiding in plain sight
Didn't know this technique had a name, but I would think a modern compiler could make this optimization on its own, no?
arkmm··on Qwen3.5 Fine-Tuning Guide
You can fine tune a small LLM with a few thousand examples in just a few hours for a few dollars. It can be a bit tricky to host, but if you share a rough idea of the volume and whether this needs to be real-time or batched, I could list some of the tradeoffs you'd think about.

Source: Consulted for a few companies to help them finetune a bunch of LLMs. Typical categorical / data extraction use cases would have ~10x fewer errors at 100x lower inference cost than using the OpenAI models at the time.

arkmm··on Qwen3.5 Fine-Tuning Guide
Can you share more details about your use case? The good applications of fine tuning are usually pretty niche, which tends to make people feel like others might not be interested in hearing the details.

As a result it's really hard to read about real-world use cases online. I think a lot of people would love to hear more details - at least I know I would!

arkmm··on Payment fees matter more than you think
Payment fees are crazy when you think about them from the perspective of a merchant in a low margin business. E.g. in retail or restaurants, margins aren't much better than ~10%. If they didn't have to pay ~3% credit card fees, they'd have 30% more profit!
arkmm··on Nano Banana 2: Google's latest AI image generation model
I used to also have this optimistic take, but over time I think the reality is that most people will instead just distrust unknown online sources and fall into the mental shortcuts of confirmation bias and social proof. Net effect will be even more polarization and groupthink.
arkmm··on The First Fully General Computer Action Model
Get ready for the acquisition offers.
arkmm··on Do not apologize for replying late to my email
Sorry Ploum, just getting a chance to read this now and comment. Great insights!
arkmm··on Claude Code is your customer
this is a really cool insight, going to use this on my team from now on!
arkmm··on Raspberry Pi's New AI Hat Adds 8GB of RAM for Local LLMs
They're still very good for finetuned classification, often 10-100x cheaper to run at similar or higher accuracy as a large model - but I think most people just prompt the large model unless they have high volume needs or need to self host.
arkmm··on Sergey Brin's Unretirement
Maybe a bit off-topic, but how'd you meet your partner while on your adventures?
arkmm··on Space Elevator
As a follow-up to this, even though water makes up 70% of the Earth's surface, it's only 0.02% of the Earth's mass.
arkmm··on DeepSeek OCR
Wow, this deserves its own submission.
arkmm··on A stateful browser agent using self-healing DOM maps
Neat approach, but seems like the eventual goal of caching DOM maps for all users would be a privacy nightmare?
arkmm··on NanoChat – The best ChatGPT that $100 can buy
What's misleading about that? You rent $100 of time on an H100 to train the model.
arkmm··on Gemini 2.5 Computer Use model
What sorts of automations were you able to get working with the Chrome dev tools MCP?
arkmm··on Circular Financing: Does Nvidia's $110B Bet Echo the Telecom Bubble?
The irony of this is so much of Reddit comments these days are AI generated.
arkmm··on After 50 years, The Magic Circle finally inducts Penn and Teller
Unfortunately I think they have stopped doing this since COVID.
arkmm··on Nvidia buys $5B in Intel
This misses the forest from the trees IMO:

- The datacenter GPU market is 10x larger than the consumer GPU market for Nvidia (and it's still growing). Winning an extra few percentage points in consumer is not a priority anymore.

- Nvidia doesn't have a CPU offering for the datacenter market and they were blocked from acquiring ARM. It's in their interest to have a friend on the CPU side.

- Nvidia is fabless and has concentrated supplier and geopolitical risk with TSMC. Intel is one of the only other leading fabs onshoring, which significantly improves Nvidia's supplier negotiation position and hedges geopolitical risk.

arkmm··on We’re Not So Special: A new book challenges human exceptionalism
Looking forward to reading corroborating essays from other non-human species.
arkmm··on Show HN: Building a web search engine from scratch with 3B neural embeddings
"There was one surprise when I revisited costs: OpenAI charges an unusually low $0.0001 / 1M tokens for batch inference on their latest embedding model. Even conservatively assuming I had 1 billion crawled pages, each with 1K tokens (abnormally long), it would only cost $100 to generate embeddings for all of them. By comparison, running my own inference, even with cheap Runpod spot GPUs, would cost on the order of 100× more expensive, to say nothing of other APIs."

I wonder if OpenAI uses this as a honeypot to get domain-specific source data into its training corpus that it might otherwise not have access to.

arkmm··on How to sell if your user is not the buyer
How are you guys reaching users with such a technical value proposition? Cold emailing engineers first and then expanding the conversation from there?
Page 1 of 2Next →