HNHacker News
TopNewBestAskShowJobs

c7b

1,733 karma · joined September 3, 2022

submissionscomments
c7b··on DSpark: Speculative decoding accelerates LLM inference [pdf]
The question is also what game they're playing. Deepseek came out of a hedge fund. I think it's no coincidence that their publications tend to have a large impact on AI stock prices.

Destroying the growth story of overvalued stocks is an interesting investment strategy. It's not even new. Shortsellers understandably get terrible rep from execs, but their actions are more often in the public interest than you'd think. Normally it's exposing fraud, but here we get the really fortunate side benefit of what could eventually amount to the most significant contribution to the general software community since Linux.

c7b··on Early adversity leaves lasting molecular imprint across the body: primate study
Almost anything can be made sense of from an evolutionary perspective. Often even the opposite of what's being observed. Can be a fun game to play. The corollary is it's not useful for vetting theories for plausibility.
c7b··on GLM-5.2 – How to Run Locally
Can someone explain the math to me? Why is 1-bit only ten percent less memory than 2-bit?
c7b··on GLM-5.2 – How to Run Locally
You can get a 128GB Strix Halo for under $3k. Used to be under $2k. Even if you believe it'll be completely obsolete for AI in two years, it'll still be good for many other things. Games for at least several more years, a great home server and/or desktop almost indefinitely. Plus, we might actually reach good enough levels for some AI use cases, if we're not already there.

And never underestimate the potential for enshittification. Your local rig will only deliver better performance over time as more and more tweaks come out. With cloud services expect the opposite to happen as subsidies run out. It's entirely possible that they will intersect on a bang per buck basis within two years.

c7b··on The frontier is open-source today
What's your hardware stack?
c7b··on Vacation With An Artist – Mini-Apprenticeships with Artists in Their Studios
Why? Afaik, apprentices are employees, with all the benefits that that entails.
c7b··on The AirPods Effect
Sounds like the opposite of what the OP and GP were advocating for.
c7b··on Leaked financial docs show OpenAI is losing billions of dollars a year
I make an important distinction between cloud services and local AI. My lifetime spending on cloud AI is probably less than $500, and I don't intend to spend any more. But I've already dropped $2.5k on new hardware for local inference, and could easily see myself spending more in the future. In fact, I'm regularly browsing for deals. I would also be open to paying for local models, if there was a way to make that compatible with fully open models.

AI is so important, I want to have it under my control. Even if I have to pay a penalty in terms of capabilities.

c7b··on Launch HN: Adam (YC W25) – Open-Source AI CAD
Congrats! I like the code-based CAD paradigm, but one question about that: why did you choose OpenSCAD instead of more powerful alternatives like CadQuery?
c7b··on Apple is about to make Hide My Email useless
If you don't mind trusting another company with forwarding your emails, it's definitely less hassle to set up an equivalent service for yourself.
c7b··on GPT‑NL: a sovereign language model for the Netherlands
And if you ask a bit more, you'll find that those same people are very likely to daily-drive AI models, desktop and phone operating systems and various other software critical to their professional and personal lives from US companies. And buy tons of Chinese products over Chinese or US e-commerce platforms.

What people say matters much less than what they do.

c7b··on Show HN: Kage – Shadow any website to a single binary for offline viewing
Probably a stupid question, but could this archive embedded videos as well?
c7b··on Open source AI must win
Since it's not mentioned in the article, the distinction between open source and open weights is important. Open weights models are almost like a 'first shot is free' entry drug. Without at least the original training data your ability to meaningfully upgrade it is so limited that its utility will quickly fall behind the latest versions of continuously developed models. So much that it'll leave you craving for another release, or have you going back to the provider's API. Even simple things like moving the knowledge cutoff forward will noticeably improve the UX, and that's not to speak of more fundamental improvements like reasoning, quantization-aware training and all the goodness that's yet to come.

Sure, we can do research to bring improvements to open weights models, but it's the same thing: it's either open source or it won't benefit the general public nearly as much.

c7b··on Open source AI must win
Could you put some numbers and examples behind the efficiency gap between data center and consumer-grade AI hardware? Did you include examples like the RTX Spark on the consumer side? I was always amazed at the low power consumption of unified memory style architectures. In absolute terms and even more so compared to consumer-grade GPUs. I'd be genuinely interested in a comparison with data-center-grade hardware.
c7b··on Show HN: Script to bulk delete Claude chats from the web UI
Wouldn't downloading conversations be more useful? If you've input something into Claude that you don't want a future trillion-dollar US company (a bicorn?) to know, this script won't help I'm afraid. Free reminder that local AI exists and works well if you're willing to tinker a bit.
c7b··on Show HN: FablePool – pool money behind a prompt, and Fable builds it in public
Interesting that this doesn't seem to use blockchains. Arguably it would have been a good use case. OP, could you elaborate on the reasons for the choice (if it was a conscious one at all)?
c7b··on Surprise, pay $1000
I'm old enough to remember that capitalism used to always feel like this. Even worse. Predatory subscriptions, targeting minors in some cases (remember phone ringtones?), toll numbers, unclear fees, surprise bills (my favorite - roaming fees in your home country because the phone stayed logged into the foreign network after an abroad trip). Silicon Valley/YC style startups felt like a breath of fresh air. Generous return policies, subscriptions can be canceled monthly and even do so automatically in a lot of cases (eg disabled credit card). I guess the fact that a lot of what they offered has near zero marginal cost helped. AI changed that part, so we're back to what I would consider the natural state.

The least-resistance path out of such a bad equilibrium is regulation. And they did add a lot of protections in the last decades, which probably helped too. Some would say places like the EU even added too many. But I'm pretty sure that "if you're using our product, you need to pay" would fly even in the most customer-friendly jurisdictions today.

c7b··on Uber's $1,500/month AI limit is a useful signal for AI tool pricing
1,5k. For two months of that spend you could buy a machine that can self-host decent models, plus a year's worth of electricity. It's not up there in terms of quality, but with a bit more effort it works pretty decently. I'm completely baffled that that's not way more common, is it really just the quality?
c7b··on Nvidia RTX Spark
It's not even anything new, it's basically the mobile version of the DGX Spark. The two chips (N1X/GB10) are pretty similar in terms of architecture and specs. I don't get why this seems to be getting so much attention now.

But I like it. It's a copy of Apple's SoC design philosophy, same as AMD's Strix Halo, which I always thought was really cool both for laptops and home PCs. NVidia's traditional consumer cards pull way too much power and are too noisy to comfortably put them in a living or office environment.

c7b··on I made a million dollar product from my dorm room (2025)
I admit, I barely understand what the product does, much less how there's 50k people wanting this. This is a component you can use if you're building a DIY keyboard and want to make it wireless? Seems profoundly niche to me. Am I missing something?

Anyway, congrats on finding and reaching your market! The Internet at its best (although part of me wishes this nerd community had found a more self-hosted way of connecting online than Discord).

c7b··on DuckDuckGo search saw 28% more visits after Google said people love AI mode
Then we should take your word over mine. My assumption was that those A/B tests will lead to products that do increase the numbers they were measuring (retention, conversion,...) at the expense of enshittified UX (up to the point of things feeling objectively broken, like notification badges re-appearing for the same items, settings that reset after user changes, search results missing,...). At least that was my explanation for how products by major tech giants like LinkedIn, Facebook, Outlook,... could end up being shipped with such flaws. What would you say?
c7b··on DuckDuckGo search saw 28% more visits after Google said people love AI mode
I would assume that they've A/B-tested any such important change extensively and basically know that it won't affect their numbers for the worse.
c7b··on The bootstrapper's EU stack for under €10 per month
That's highly misleading to outright misinformation.

> Passkeys don't have to be remembered

Because you need an app for the login flow. You also don't have to remember passwords if you use a password manager app.

> don't need 2FA

Not true, a second factor in the form of eg a biometric ID or PIN is mandatory.

Phishing resistance exists, but only truly so if you completely surrender control over your device and access to your credentials. Something that the same organizations who you'll depend on for Passkeys are actively pushing for through various initiatives.

c7b··on Omarchy Is Not A Distro
I agree with the author's general sentiment, but I would give Omarchy credit for some design features that I find either really innovative or well executed.

I find the launch menu the most interesting I've seen out of any desktop environment / OS. It puts easy access to your apps front and center and above esthetics, and yet it still looks great. It's not a sane request, but I'd love to be able to swap the launcher/start menu in another DE for a fully customizable version of that.

Omarchy also has a very smooth way to 'install' web apps. There exist packages that do that for you, but I've tested several and wouldn't recommend any of them. They're bloated (this should really only be a couple lines of bash, like Omarchy's [0]), some use web views (I really want an actual browser under the hood so I can have my extensions), and all that I've tested leave litter on your machine after removing apps or the software. I find web apps as .desktop files so useful that I'm actually using a hacky DIY script now (which I'm considering releasing under GPL if I ever find time to clean it up).

Also, whether you're a beginner or a pro, having a sane starting config for Hyprland is just convenient. Which tells you something about it imho. My conclusion would be similar to the OP's, if you're Omarchy-curious, try Cosmic on your distro of choice. Or at least a cleaned up version with the most egregious personal preferences (like a global keybind for opening Twitter/X) removed, if anybody cares to maintain that.

[0] https://github.com/basecamp/omarchy/blob/dev/bin/omarchy-web...

c7b··on Where are all the UK red telephone kiosks?
Some suggestions:

- Not sure what they're called, but I've seen a lot of fully automated outdoor "locker stations" for packet deliveries

- Power bank "banks" or charging stations for smartphones in indoor spaces like malls

- QR codes on stickers/ads in public spaces are a sort of bridge between the physical and digital worlds

c7b··on How fast is N tokens per second really?
Regarding the first, parallel requests to the same loaded model seem to work pretty well, I'm trying to find time to look more into it myself, but this may be something that might already be within reach for local models.
c7b··on An OpenAI model has disproved a central conjecture in discrete geometry
One could hardly ask for a task better suited for LLMs than producing math in Lean. Running a restaurant is so much fuzzier, from the definition of what it even means to the relation of inputs to outputs and evaluating success.
c7b··on How fast is N tokens per second really?
Do you have ideas/suggestions for agentic workflows that only start making sense at such speeds?
c7b··on Show HN: Forge – Guardrails take an 8B model from 53% to 99% on agentic tasks
So as an oversimplified PoC, I get:

    llama-parallel -m ~/models/Qwen3.5-4B-Q8_0.gguf -ns 4 -p "Fix this Python code, answer with code only: prnt('Hello World)" -pps

    llama_perf_context_print:        load time =    1181.90 ms
    llama_perf_context_print: prompt eval time =     190.57 ms /   374 tokens (    0.51 ms per token,  1962.49 tokens per second)
    llama_perf_context_print:        eval time =    3612.25 ms /   159 runs   (   22.72 ms per token,    44.02 tokens per second)
    llama_perf_context_print:       total time =    4302.84 ms /   533 tokens
    llama_perf_context_print:    graphs reused =        155
and four answers (3 of which are immediately usable), with -ns 1 I get :

    llama_perf_context_print:        load time =    1185.61 ms
    llama_perf_context_print: prompt eval time =     187.55 ms /   305 tokens (    0.61 ms per token,  1626.27 tokens per second)
    llama_perf_context_print:        eval time =     158.92 ms /     7 runs   (   22.70 ms per token,    44.05 tokens per second)
    llama_perf_context_print:       total time =     468.85 ms /   312 tokens
    llama_perf_context_print:    graphs reused =          6
Now this is probably not the right way to use it, you should probably also use vLLM instead and it's also not a good model to use for this. But there is a real effect here that others have demonstrated, that the GPU is apparently not always maxed out while handling a single request, so sending concurrent requests can yield substantial parallelization benefits. The idea with this application would be something like this: send off the same query in parallel requests, triggering parallel tool calls, and then filter the results (filter out all failing ones, rank the rest by some simple metric of code complexity). There are probably better applications as well, I'm basically just thinking what kinds of tasks could benefit from parallelization.
c7b··on Copy Fail, Dirty Frag, and Fragnesia kernel vulnerabilities
If we were to start from security first, we would be asking questions like 'how can we make sure that new code is safe?'. Manual review is great, but we can likely think of some desirable invariants for program behavior that could be tested automatically, or even formally verified. Those would come at the very start. The entire mindset right now is that the existing code is probably unsafe and we'll ship fixes as we discover its vulnerabilities. Not immediately applying updates is seen as a kind of moral failure. All major OS and most software projects were developed with this mindset of crossing your fingers at launch and then changing the tires while driving. So much so that we think of it as the natural state of software. If you start from a base of verified code, the mindset shifts. Not that there are zero vulnerabilities guaranteed in the existing code, but you become a lot more suspicious of new code.
← PreviousPage 5 of 20Next →