HNHacker News
TopNewBestAskShowJobs

lcampbell

324 karma · joined November 16, 2011

submissionscomments
lcampbell··on Sonnet 5.5
I was under the impression that the cache write fee was added to both the input and output costs (except in cases where the cache write is explicitly disabled via e.g. DISABLE_PROMPT_CACHING). The output becomes part of the context, after all; if they don't (for some reason, due to disaggregated inference perhaps) then I'd expect output tokens get charged both output then input+cache_write on the subsequent completion request.

The pricing model confuses me though (I presume by design, Hanlon be damned).

lcampbell··on DeepSeek API Pricing Update
> Do all other providers somehow overcharge by that much?

This, I think. Cached inputs have an opportunity cost (keeping the KV cache until use) but a hit is basically free. “Basically” - if the cache is offloaded to system RAM or NVMe there’s some scheduling overhead.

From a consumer viewpoint a more interesting metric than the raw costs is

    cached cost * hitrate + input cost * (1 - hitrate)
from a personal standpoint rather than a per-provider one (e.g. if OpenRouter is blindly dispatching your requests you might have a bad time).
lcampbell··on DeepSeek V4 Flash 0731
You're gonna want a custom vLLM build.

Here's a runbook: https://github.com/local-inference-lab/rtx6kpro/blob/master/...

If the newer builds aren't working, you might try running the old v6 build (based on the eldritch-enlightenment image). gilded-gnosis gave me some problems that I haven't bothered to track down, the old builds are still gonna blow away llama-server performance. And that's before you get hooked on vLLM's PagedAttention and can run multiple sequences without a ton of extra overhead.

lcampbell··on Self-hosting Kimi K3: 20% more hardware cost, 20% better task resolution
> Aggregate token throughput is about 30% lower (122 vs 170 tok/s at 16 users)

Is that actually the aggregate throughput? 8xB300 (with 4TB/s/GPU bandwidth) is only pushing 8 tg/s/session? That seems… incredibly low, even for an A100B model.

Is it actually 122 tg/s per session? (1952 tg/s aggregate throughout)?

lcampbell··on Engineering management after the cost of code collapsed
More than that, the AI disclaimer itself is a hallucination:

> Gemini 4 helped with the editing.

There is no Gemini 4, unless the author is writing from the future.

lcampbell··on Kimi K3 exploited the latest Redis server
FWIW, you can get a 16x RTX 6000 Pro setup running for significantly less than that. Even considering the electrical hookup fees (it’s a lot of power and cooling). Home ownership might push you into the original quote though.
lcampbell··on Migrating a production AI agent to GPT-5.6: 2.2x faster, 27% cheaper
> The fix that worked is a schema transform at the provider boundary. For OpenAI-family models only, we rewrite every optional property to be required but nullable, using anyOf: [T, null], which gives the model an explicit way to say “not using this.”

I admit, I've only used a bastardized form of MCP, but this smells... wrong? It's not clear to me why the Typescript type definitions would have any influence on (what I presume is) JSONSchema being sent from the agent to the inference backend as part of the completion request. The MCP specification (which the OpenAI backend might not use, I don't know) has an explicit field to signify "optional" parameters in the JSONSchema; my read on this is there's a bug somewhere between the Typescript layer(??) and the generated tool description which is actually sent to the inference backend.

It's possible the inference backend has changed from "generate valid tool responses" to "generate valid tool responses according to the JSON schema [where no parameters are optional]" but it's impossible to tell without seeing the actual requests sent to the inference backend (which I didn't see in TFA).

lcampbell··on LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active
I don't think llama.cpp supports any of the LongCat models, actually.

They haven't posted weights/inference solutions for LongCat-2.0 [1], but LongCat-Next had transformers support, which I assume means it works with vLLM/SGLang.

Given it's 1.6T, "common hardware" is probably out of the question; even 2bpw is going to measure out at 400GB, even before considering the bandwidth requirements for 48B active. I haven't read the LongCat-2.0 architecture docs, but if you're not running GLM-5.2, you're probably not running this either.

[1] https://huggingface.co/meituan-longcat/LongCat-2.0: "Model weights coming soon — stay tuned!"

lcampbell··on Professor denounces mass AI fraud on an exam at Brown
At UVA many years ago, one of my roommates was one of the unfortunate 20 or so annually expelled -- the only outcome of being convicted of breaking the "no cheating, stealing, or lying" honor code. It didn't take repeat offenses, expulsion was a first offense consequence.

Interestingly, it seems like you weren't joking about the decline:

> Finally in the spring of 2022, a sanction reform referendum succeeded with more than 80% of the vote, changing the penalty for an Honor violation from expulsion to a two semester suspension. [1]

[1] https://en.wikipedia.org/wiki/Honor_system_at_the_University...

lcampbell··on A robot is sprinting towards you. Do you want it running on Claude or Grok?
> I want to be careful here.

was the giveaway for me

lcampbell··on AI coding is gambling
The reasoning generally isn't kept in the context, so after choosing the secret word in the first reasoning block, the LLM will have completely forgotten it in the second and subsequent requests.

So, it technically didn't change the secret word so much as it was trying to infer what its own secret word might have been, based on your guesses.

lcampbell··on Linux gamers on Steam cross over the 3% mark
Epic Anti-Cheat fully supports Linux[1]. I believe what the GP comment means is that the Fortnite publishers opted not to tick the “allow Linux” checkbox on the developer portal website.

There is probably more nuance behind that decision than I’m giving them credit for, but from a technical standpoint it’s just a checkbox.

[1] https://dev.epicgames.com/docs/game-services/anti-cheat/usin...

lcampbell··on Upcoming coordinated security fix for all Matrix server implementations
Without knowing anything about Tor, I'd guess you've got it backwards. I imagine Tor leaks your OS through TCP/IP fingerprinting, and whether that fingerprint matches your `navigator.platform` is probably a factor into whether e.g. Cloudflare hellbans you.

Then again, I'd also assume Cloudflare just de facto hellbans all Tor exit node IPs, so...

lcampbell··on AI is killing some companies, yet others are thriving – let's look at the data
If you're given a button to click, your browser has successfully passed the environment integrity checks and you have not been flagged as a bot.

You'll be flagged as a bot if your browser configuration has something "weird" (e.g. webrtc is disabled to reduce your attack surface) and you will be completely unable to access any site behind cloudflare with the anti-bot options turned on. You'll get an infinite redirect loop, not a button to click.

lcampbell··on Test if a number is even
Interestingly, not only do both versions emit the same assembly, but clang both autovectorizes and unrolls the loop:

https://godbolt.org/z/qxsKWfz9s

lcampbell··on How is my Browser blocking RWX execution?
This looks like it, I think:

https://searchfox.org/mozilla-central/source/toolkit/xre/dll...

lcampbell··on OpenSSH introduces options to penalize undesirable behavior
> what if someone turns [password authentication back] on

sshd_config requires root to modify, so you've got bigger problems than weak passwords at this point.

lcampbell··on Psychological tricks rich people use to look generous without spending more
Less than skips over, utility based shopping is explicitly derided:

> The narrative that you just told me [about utility shopping] is “I am a very analytical person who only has book smarts and no emotions”. And that narrative is boring!

lcampbell··on Curl HTTP/3 security audit
This is briefly mentioned in the article, but from the report[1]:

> It should be noted that the scope of the code reviewed within this audit is relatively narrow. In particular, while we audited cURL’s use of the third-party libraries ngtcp2, nghttp3, quiche, and msh3 to implement HTTP/3 functionality, we did not investigate the internals of those libraries—which is where the majority of the low-level parsing and data transformation necessitated by the HTTP/3 protocol occurs.

the report goes on to concede

> [we] did not observe any coverage of the nghttp3 library code. We suspect that, as the HTTP/3 protocol itself is significantly intertwined with TLS, the encryption makes it hard for a fuzzer to progress to the point where data can be decoded and parsed meaningfully.

[1] https://curl.se/docs/audit/trail-of-bits-http3-report.pdf

lcampbell··on “Clickless” iOS exploits infect Kaspersky iPhones with never-before-seen malware
Given the exploit vector looks like yet another iMessage attachment bug,

> The target iOS device receives a message via the iMessage service, with an attachment containing an exploit.

and that one of the effects of Lockdown Mode is

> Messages - Most message attachment types are blocked, other than certain images, video, and audio. Some features, such as links and link previews, are unavailable.

It might be prevented. Pretty sure disabling iMessage altogether sidesteps this class of bugs too. I've lost track of how many times iMessage has been the root cause of "unattended iOS RCE," at this point it's almost user negligence to have left on.

lcampbell··on My daughter's school took over my personal Microsoft account
> banned by Microsoft for breach of TOS, whatever that might have been

FWIW, I had the same thing happen and found out the ban reason was "fraud (please insert phone number)".

lcampbell··on Chat GPT is the birth of the real Web 3.0, and it's not going to be fun
I believe GP is referring to the multiple AI-driven vtubers on Twitch (vedal987 and motherv3 I think?). They're not yet 24/7 because they still require human supervision for reasons -- vedal987 was recently banned for holocaust denial, IIRC.
lcampbell··on Stop the proposal on mass surveillance of the EU
Probably just an extension to existing ETSI legal interception interfaces[1] that I believe are required to be implemented by all service providers of a certain size in the EU. Your personal email server and private IRC network are probably out of scope.

[1] e.g. for email: https://www.etsi.org/deliver/etsi_ts/102200_102299/10223202/... - the specs are really boring to read and there are lots and lots of them, if you want to deep dive. From what I can tell it's basically just the service provider implementing an API client that feeds everything they're required to a centralized endpoint.

lcampbell··on Regulators of Facebook, Google and Amazon also invest in the companies’ stocks
> but prohibitions on Congressmen remain in force [i.e. 2020 congressional insider trading]

Is it really "in force" when, despite much ado, no charges were brought in the linked scandal [1][2]? I don't really have a horse in this race, I just took issue with the particular example you referenced.

[1] https://www.nytimes.com/2020/05/26/us/politics/senators-stoc...

[2] https://www.cnbc.com/2021/01/19/doj-will-not-charge-sen-rich...

lcampbell··on Google pays ‘enormous’ sums to maintain search-engine dominance, DOJ says
Are you using the no-javascript version (“HTML” version), by any chance? If so, you might find the “fully-featured” version has better^W results, at the cost of requiring Javascript and all that entails.
lcampbell··on Debian's Chromium changes default search engine to DDG
As a regular user of the Javascript-less page, several months ago it started returning wildly different results than the “fully featured” version for the same queries. My uneducated guess is that it’s using a different index. There also appears to be some sort of rate-limiting wherein the results will frequently just be empty (using the JS version and same query resolves the issue).

I’m guessing they’re intentionally degrading the non-Javascript page as an anti-bot measure, but it’s so bad that I find it disingenuous to suggest that the non-Javascript page even a valid alternative at this point.

lcampbell··on tcpcp - passing TCP connections between hosts (2005)
This might make more sense in the situation where multiple hosts share the same IP, as in a CARP[1] setup. There’s probably other useful use-cases but that’s the one that comes to my mind.

[1] https://en.wikipedia.org/wiki/Common_Address_Redundancy_Prot...

lcampbell··on Tell HN: I got 10x Hetzner storage at the same price
Direct link: https://www.youtube.com/watch?v=5eo8nz_niiM

Really enjoyed the video, thanks for the suggestion!

lcampbell··on How to send a real number using a single bit (and some shared randomness)
Anything over a network is going to have layer 2 framing overhead. Something like a serial port would do the trick.
lcampbell··on Sod – An Embedded Computer Vision and Machine Learning Tiny C Library
Doesn't seem like it. The GPU module required for training CNNs is sold separately[1]:

> A GPU capable special SOD release which is not available to the public is required in order to train your own CNN model.

There are some bits included for training "RealNet" models[2] but I'm not sure what those are or how they work. The documentation suggests that RealNet models can be trained in reasonable timeframes with a CPU.

[1] https://sod.pixlab.io/cnn_train.html

[2] https://sod.pixlab.io/c_api/sod_realnet_train_start.html

Page 1 of 4Next →