The pricing model confuses me though (I presume by design, Hanlon be damned).
324 karma · joined November 16, 2011
The pricing model confuses me though (I presume by design, Hanlon be damned).
This, I think. Cached inputs have an opportunity cost (keeping the KV cache until use) but a hit is basically free. “Basically” - if the cache is offloaded to system RAM or NVMe there’s some scheduling overhead.
From a consumer viewpoint a more interesting metric than the raw costs is
cached cost * hitrate + input cost * (1 - hitrate)
from a personal standpoint rather than a per-provider one (e.g. if OpenRouter is blindly dispatching your requests you might have a bad time).Here's a runbook: https://github.com/local-inference-lab/rtx6kpro/blob/master/...
If the newer builds aren't working, you might try running the old v6 build (based on the eldritch-enlightenment image). gilded-gnosis gave me some problems that I haven't bothered to track down, the old builds are still gonna blow away llama-server performance. And that's before you get hooked on vLLM's PagedAttention and can run multiple sequences without a ton of extra overhead.
Is that actually the aggregate throughput? 8xB300 (with 4TB/s/GPU bandwidth) is only pushing 8 tg/s/session? That seems… incredibly low, even for an A100B model.
Is it actually 122 tg/s per session? (1952 tg/s aggregate throughout)?
> Gemini 4 helped with the editing.
There is no Gemini 4, unless the author is writing from the future.
I admit, I've only used a bastardized form of MCP, but this smells... wrong? It's not clear to me why the Typescript type definitions would have any influence on (what I presume is) JSONSchema being sent from the agent to the inference backend as part of the completion request. The MCP specification (which the OpenAI backend might not use, I don't know) has an explicit field to signify "optional" parameters in the JSONSchema; my read on this is there's a bug somewhere between the Typescript layer(??) and the generated tool description which is actually sent to the inference backend.
It's possible the inference backend has changed from "generate valid tool responses" to "generate valid tool responses according to the JSON schema [where no parameters are optional]" but it's impossible to tell without seeing the actual requests sent to the inference backend (which I didn't see in TFA).
They haven't posted weights/inference solutions for LongCat-2.0 [1], but LongCat-Next had transformers support, which I assume means it works with vLLM/SGLang.
Given it's 1.6T, "common hardware" is probably out of the question; even 2bpw is going to measure out at 400GB, even before considering the bandwidth requirements for 48B active. I haven't read the LongCat-2.0 architecture docs, but if you're not running GLM-5.2, you're probably not running this either.
[1] https://huggingface.co/meituan-longcat/LongCat-2.0: "Model weights coming soon — stay tuned!"
Interestingly, it seems like you weren't joking about the decline:
> Finally in the spring of 2022, a sanction reform referendum succeeded with more than 80% of the vote, changing the penalty for an Honor violation from expulsion to a two semester suspension. [1]
[1] https://en.wikipedia.org/wiki/Honor_system_at_the_University...
was the giveaway for me
So, it technically didn't change the secret word so much as it was trying to infer what its own secret word might have been, based on your guesses.
There is probably more nuance behind that decision than I’m giving them credit for, but from a technical standpoint it’s just a checkbox.
[1] https://dev.epicgames.com/docs/game-services/anti-cheat/usin...
Then again, I'd also assume Cloudflare just de facto hellbans all Tor exit node IPs, so...
You'll be flagged as a bot if your browser configuration has something "weird" (e.g. webrtc is disabled to reduce your attack surface) and you will be completely unable to access any site behind cloudflare with the anti-bot options turned on. You'll get an infinite redirect loop, not a button to click.
https://searchfox.org/mozilla-central/source/toolkit/xre/dll...
sshd_config requires root to modify, so you've got bigger problems than weak passwords at this point.
> The narrative that you just told me [about utility shopping] is “I am a very analytical person who only has book smarts and no emotions”. And that narrative is boring!
> It should be noted that the scope of the code reviewed within this audit is relatively narrow. In particular, while we audited cURL’s use of the third-party libraries ngtcp2, nghttp3, quiche, and msh3 to implement HTTP/3 functionality, we did not investigate the internals of those libraries—which is where the majority of the low-level parsing and data transformation necessitated by the HTTP/3 protocol occurs.
the report goes on to concede
> [we] did not observe any coverage of the nghttp3 library code. We suspect that, as the HTTP/3 protocol itself is significantly intertwined with TLS, the encryption makes it hard for a fuzzer to progress to the point where data can be decoded and parsed meaningfully.
[1] https://curl.se/docs/audit/trail-of-bits-http3-report.pdf
> The target iOS device receives a message via the iMessage service, with an attachment containing an exploit.
and that one of the effects of Lockdown Mode is
> Messages - Most message attachment types are blocked, other than certain images, video, and audio. Some features, such as links and link previews, are unavailable.
It might be prevented. Pretty sure disabling iMessage altogether sidesteps this class of bugs too. I've lost track of how many times iMessage has been the root cause of "unattended iOS RCE," at this point it's almost user negligence to have left on.
FWIW, I had the same thing happen and found out the ban reason was "fraud (please insert phone number)".
[1] e.g. for email: https://www.etsi.org/deliver/etsi_ts/102200_102299/10223202/... - the specs are really boring to read and there are lots and lots of them, if you want to deep dive. From what I can tell it's basically just the service provider implementing an API client that feeds everything they're required to a centralized endpoint.
Is it really "in force" when, despite much ado, no charges were brought in the linked scandal [1][2]? I don't really have a horse in this race, I just took issue with the particular example you referenced.
[1] https://www.nytimes.com/2020/05/26/us/politics/senators-stoc...
[2] https://www.cnbc.com/2021/01/19/doj-will-not-charge-sen-rich...
I’m guessing they’re intentionally degrading the non-Javascript page as an anti-bot measure, but it’s so bad that I find it disingenuous to suggest that the non-Javascript page even a valid alternative at this point.
[1] https://en.wikipedia.org/wiki/Common_Address_Redundancy_Prot...
Really enjoyed the video, thanks for the suggestion!
> A GPU capable special SOD release which is not available to the public is required in order to train your own CNN model.
There are some bits included for training "RealNet" models[2] but I'm not sure what those are or how they work. The documentation suggests that RealNet models can be trained in reasonable timeframes with a CPU.
[1] https://sod.pixlab.io/cnn_train.html
[2] https://sod.pixlab.io/c_api/sod_realnet_train_start.html