HNHacker News
TopNewBestAskShowJobs

nicce

5,733 karma · joined February 27, 2021

submissionscomments
nicce··on Speculative Decoding in vLLM on AMD GPUs
2x RTX 3090 is not enough for proper use. ideal is 64GB+ so that you get proper cache and 256k context with good enough models. (e.g. Qwen 3.8 27B with MXFP4). And that is tight already.

You would need 3x RTX 3090 - but then the PCIe bandwidth comes an issue if you really want tensor parallelism with three cards. 3x PCIe5 x16 is not cheap with direct CPU access.

So surprisingly, 2x r9700 starts be a nice deal.

> Also these 20-30t/s jumping to 150-200... Watch out for the massaged numbers coming from vendors.

Well, luckily these are not vendor numbers. Prefill also scales almost linearly with the amount of GPUs.

nicce··on LG smart TVs caught logging audio with screen off and snooping on local devices
I bought Apple TV many years ago and haven’t plugged tv into internet since.
nicce··on It's time for Mark Zuckerberg to resign from Meta
Duolingo has joined to the addiction club
nicce··on How we monitor internal coding agents for misalignment
Both can be true. I was saying marketing trick about the response and selected words ”how great it was”.
nicce··on How we monitor internal coding agents for misalignment
Sounds like marketing trick once again, to be honest.
nicce··on QBittorrent breaks out of sandbox to commit crimes
Not far away if it could be completely removed with all the other competitive technologies:

https://www.uschamber.com/technology/data-privacy/impacts-of...

nicce··on There's No Limit to How Bad Code Can Get
People already look me like crazy when I still review all the code changes AI does, and not even using the agent directly. And make it do things differently if it is unmaintainable slop. Their argument is that are you really going to modify the code by hand in the future?
nicce··on GPT-6 Astra might be too powerful to understand or control
Meanwhile it still can’t write a sentence without adding three semicolons.
nicce··on Qwen 3.8 27B available on Cerebras at 1500 tokens/s
You can run it with 2x r9700 with 150-200 tokens per second. It is intelligent enough if you just point the docs / whatever for it.
nicce··on Qwen 3.8 27B available on Cerebras at 1500 tokens/s
They seem to have good enough general intelligence that missing knowledge is not that big thing. If you are able to have a proper [free search engine], they can do almost anything. Having own local search index about relevant stuff can help a lof if you don’t want to pay for search API.
nicce··on GPT-6 Astra
Yeah. Unless human can verify it, not sure if it is certain or useful.
nicce··on Qwen 3.8 27B available on Cerebras at 1500 tokens/s
Many. Too lucrative for certain companies and even governments to allow that to happen
nicce··on Pre-Release of Polars 2.0
> Version bumps should really be about removing deprecated cruft rather than shiny new features.

Can there be deprecated cruft without new features? :-D

nicce··on Quasar 438B: Europe's Leading AI Model
Issue is that it is extremely hard to notice if there is backdoor or some training-related hallucinations that are completely random.

E.g. I had random, completely unrelated and irrelevant fetch by Qwen3.8 27B to " https://routify-file-proxy-sg.oss-ap-southeast-1.aliyuncs.co..." and I only noticed it because I have allowlist rules for what they can do.

nicce··on Claude Fable 5.1 and Claude Mythos 5.1
I used Fable once. Used through API and asked it to review one 2k word plan. It costed me 15 dollars and haven't used it since.
nicce··on Apple caught off guard by AI demand for Mac Mini and Mac Studio
> Apple doesn't design GPUs on-par with Nvidia's efficiency yet

How much it matters in inference? Most GPUs have enough computing for that and the bottleneck is the RAM speed and size. And M5 Ultra is becoming to challenge this.

nicce··on A CVE Dispute
First thing that came to my mind... why people are so self-centric... and yet, somehow, you feel bad because the same people flex with these CVEs while you can't tell a single huge one you had found under NDAs, while they keep downplaying or ignoring you. I think its better to treat these people as less professional, and try to influence the common reception what is actually professional and what is not.
nicce··on EVE Online moves to Python 3
Depends on what you do. For prototyping GUIs without well-defined framerwork, Rust is terrible. TypeScript wins completely. For back-ends, not so clear, since it is easier to write the intent there and even prototype must be somewhat correct, and TS is not as explicit as Rust is.
nicce··on Asahi Linux Progress Report: Linux 7.2
> I’ve got some news to share about your new Macbook.

As long as you pay for Apple Care, not a problem. But maybe not the point…

nicce··on Coding expertise is going to collapse from AI reliance
Thank you! Have to look into it.
nicce··on Nvidia agrees to acquire Hugging Face for $13B
Money always wins, sadly.
nicce··on Coding expertise is going to collapse from AI reliance
What is your config for getting 80tps decode?
nicce··on Why your local LLM feels dumber than it is
Having a server in the basement helps a lot :-D Then tailscale from everywhere.
nicce··on Bun 1.4
> That total assumes API billing rates, which isn't what Anthropic pays for the inference. So I'm confused why people keep quoting that. It's not representative of what it'd cost for you either because you could use cheaper models.

It is, because this rewrite was marketing for Claude. So it reflects the price for what other companies might need to pay if they want to rewrite something else with similar size.

nicce··on Anthropic IPO valuation hinges on $190-200B 2028 revenue forecast
Look information about the parent company High-Flyer which has has been breaking profit records with AI based stock market prediction. They already had the computing before LLM era.
nicce··on Anthropic IPO valuation hinges on $190-200B 2028 revenue forecast
> The Chinese companies will need to pay their bills eventually too.

What bills? Deepseek has been profitable for long time.

nicce··on Firefox for iOS now has a native adblocker
How is Firefox on iOS so fast compared to Safari or Orion? It is easily noticeable
nicce··on Firefox for iOS now has a native adblocker
> without it they won't produce quality content.

If the content would be really high quality and exceptional, people would pay for it. But if it close to something people publish even without ads, then they don’t pay for it.

nicce··on What Happened to HackerOne?
> You can report a self-XSS sev:hi (and bounty hunters do) and get many orgs to take them seriously, because they don't have serious security practices.

Which can be definitely high, if it can be triggered by giving specific URL, for example.

I think there is too much generalization happening here.

nicce··on What Happened to HackerOne?
> Every application has those bugs; on a software pentest, we'd sev:lo them.

Every application has a bug that can bring the whole application down for every user without owning a botnet? That comes often with a significant business cost, if someone exploits it. Many companies take them seriously. I have reported many as high and business has agreed. Not with HackerOne thought. If there is a bug where someone can make your whole product down with a single laptop isn't really something you can just ignore.

← PreviousPage 5 of 34Next →