HNHacker News
TopNewBestAskShowJobs

_ache_

818 karma · joined October 30, 2021

submissionscomments
_ache_··on OpenDLSS: A Vulkan Reimplementation of Nvidia's DLSS 5 Neural Rendering Network
From the LICENSE

Copyright (c) 2026 maan

So... Either alooshdenny stole the commits, or it's an alias for maan.

_ache_··on Qwen 3.8 Omni Flash
Yes, that was my point. Good but too niche.
_ache_··on Qwen 3.8 Omni Flash
You point is valid for textual LLM, with large CoT, not Omni which will respond quick with a voice. In this case, token price is a good enough proxy.

What matter the most and isn't told by token price is the latency. You expect a voice LLM to respond very quick. If it takes 5s to response to a simple "Hello, what the weather today?", them not much people will use it.

_ache_··on Qwen 3.8 Omni Flash
Very capable yes but very slow. 27B is relatively easy to run, but the 125b one need around 128Gb of RAM (DDR4 isn't enough, you need DDR5 to be quick enough, that's $3000 alone, you also need a graphic card). DDR4 is caped @20tps.

So, the cost of a setup to run Qwen-Flash-Next at +40tks is around $3000. Too much for most people.

With only a RTX 4090, you will reach 30tps (with DDR5...), not +40tks, and it's about the limit to be usable. Oh ! I forget Apple device too, it's a good option to run this model I guess, but still slow.

Yet, as you said, it's still a wip implementation, it may improve soon (MTP support is about to be merged in llama.cpp soon).

_ache_··on Shapelearn Qwen 3.8 27B (13.1 GB VRAM)
Ahahah, thank you.

Yes, I'm not a native English speaker, and I'm definitely not an AI. ;)

I can't say the same thing because I just don't notice typos (mine or others).j I just assume my English is bad.

Isn't the definition of "being human" is "not to be perfect"? In French, it kinda is. We say "He/she humain after all" to mean that someone made a mistake.

_ache_··on Shapelearn Qwen 3.8 27B (13.1 GB VRAM)
Ahahah :'D

Yes... I'm not english native. And definitively not an AI.

_ache_··on Qwen 3.8 Omni Flash
I know, but it's a good enough proxy.
_ache_··on Shapelearn Qwen 3.8 27B (13.1 GB VRAM)
From my own test. It's not faster than the unsloth model.

Disclarer: I'm unsing Vulkan on an AMD GC.

_ache_··on Qwen 3.8 Omni Flash
If the performances are comparable, and there is no evidence it's not.

in/out ($) Gemini : 1.5 / 9.0 | Qwen 3.8: 0.15 / 0.47

That is a massive cost reduction.

Refs: https://www.alibabacloud.com/help/en/model-studio/model-pric... https://runware.ai/gemini-omni

_ache_··on Qwen 3.8 Omni Flash
I don't think Qwen3.8-Omni-X will ever be released.

The last one was: Qwen3-Omni-30B-A3B https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct

And maybe Qwen4 won't be released, they only release Qwen3.8 27B (and a mostly unusable 125B). There are definitively slowing down open weight release.

_ache_··on Mistral X Mozilla: Private, Multilingual AI Browsing
The strategies of Google and Apple, regarding how to provide a LLM, seam to disagree with you. Gemini run on a potato and Apple is local first.

So, you may actually have very good performance with local model. Just not yet on *every* device. So the Mozilla strategy here feel very reasonable. A Cloud provider specialised in local models, to be able to switch once local models will be quick enough on most devices.

_ache_··on AI models don't kill people – people kill people
Ok, so OpenAI is going to compensate RubyGem for the cyberattack on its servers?
_ache_··on GPT-6 Astra
https://ache.one/gpt6_now_down.png

Big claims, expensive and not release to the public yet.

_ache_··on OpenAI begins rolling out GPT-6 Astra
It's up then down again. https://openai.com/index/gpt-6-astra/

What a bunch of amateurs. Here is it anyway :

https://ache.one/gpt6_now_down.png

The claims: https://share-md.com/view?id=870ba228-a25c-4169-bbc9-12d7f25...

And some others like this bugged Karts Game:

https://tidal-rush-paradise-gp.skirano.chatgpt.site/

This impressive spaceship construction game:

https://voidexplorer-shipyard.openai.chatgpt.site/?fleetSeed...

And a lot of graphs, some without even Astra on it. Oh and the logo is a Galaxy.

_ache_··on I trained a small transformer in 1.5hrs and it beats many LLMs
In computer science, that is technically a language. A formal language if you want to look it up on Wikipedia.
_ache_··on Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)
I will rephrase it. Will it run on any consumer hardware?
_ache_··on Claude Code is going reduce limits by 25% from September 14
Reduce limits or usage? Twitter is blocked.

Can you do more or less?

_ache_··on Show HN: See fiber breaks linked to a map
Is it linked to the last month COLT problem?
_ache_··on Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)
I'm hearing Tencent, Zhipu and Baidu shaking from here. It's fair to assume BATX / 6 Tigers don't sleep very well either.
_ache_··on Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)
I think a reasonable expectation of MAX requirement to claim "runable on consumer hardware" is to 32G VRAM and 128GB RAM and it run at +10tps.
_ache_··on Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)
What will be the requirement, like 128G of RAM and 12G of VRAM ?
_ache_··on Anthropic Risk August 2026 [pdf]
They need to train a new model every month to keep at the top of most benchmarks. They don't own any DC, the price is insane. Most of people are aiming at smaller models because Claude one's are too expansive. Evolution of intelligence of bigger models start to stagnate, smaller models are catching up.

27B local model just dropped, it's 6/8-month old SOTA. General ROI of AI investment is expected on a baseline of >10y.

_ache_··on Qwen 3.8 27B
From your benchmark, Qwen3.8 is nearer than Opus 4.8 than Qwen3.6. 0.1pp but still.

Also, a lot of people don't really care about german language capacity, maybe people programming in DDP idk.

PS: You benchmark seems saturated. Most values sit @>75% in a benchmark generally indicate that it's no longer as useful as a <70% one. I mean, Qwen3.8 is 77.5% and Fable5 80%, the poll of values is from 65% to 90%.

_ache_··on Anthropic Risk August 2026 [pdf]
I actually expect them to explain me how they will manage to not go bankrupt soon.
_ache_··on Anthropic Risk August 2026 [pdf]
It's crazy how Anthropic talks so much about their "AGI risk" and not enough about the risk of bankruptcy.
_ache_··on France's tax authority had data stolen on 680k taxpayers
It is already. You can buy it online. There is not a lot of places where hackers sell that kind of stuffs.

1k lines are already shared, seems legit. Most of them is just <30k€ people. Only a handful of millionaires (8 >10M if I remember correctly), 0 billionaires. 2/3 people, 1/3 pro, there is no information about tax of professionals.

Seems to be an API endpoint about people who ask questions to the DGFiP.

_ache_··on GLM-5.3: Frontier coding with emergent cyber capabilities
No yet finished! Still waiting for tonight Qwen3.8-27B and the unsloth Q5_K_M/S quantification.

Hopping for an AgentWorld variant from Qwen but I guess, I have too high expectations.

_ache_··on AI At Home Part 1: A Box Of Scraps
What can you do with "only" 64G of VRAM that a 32G can't? Also, the R9700 are so loud!
_ache_··on Docker Sandboxes – Disposable, isolated sandboxes for AI agents
I planed to do exactly this, with podman instead of docker, volume support.

Like:

$ podman run -it --rm -v .:/workspace local-dev-ia /usr/bin/oc

Configured with a .env file. Hope to do it hopefully before the end of the week.

_ache_··on Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
It is interesting but it does look like a careful distillation of (Spark and) biggers open-weight models.

The progress compared to Qwen3.6 27B is good, not that impressive, it's a 4 months old model. (kuto to them to compare to 27B dense and not 35B MoE, it's more fair to do so). It is very probable that Qwen3.8 27B will crush Glimmer-30B on most benchmarks.

Page 1 of 13Next →