HNHacker News
TopNewBestAskShowJobs

trouve_search

179 karma · joined May 27, 2026

submissionscomments
trouve_search··on Trump administration is suspending Microsoft from a green card program
> Foreign labour depreciates the wages companies have to pay for employees. It has been shown again and again that concerning quantities of foreign labour is used exactly for that purpose

That's a (citation needed) right there.

The literature on the matter finds little to no effects of immigration on local wages. See the literature review by Kerr & Kerr:

https://www.nber.org/system/files/working_papers/w16736/w167...

Key sentence in the conclusion: "The likelihood and magnitude of adverse labor market effects for natives from immigration are substantially weaker than often perceived. Within the large empirical literature looking at the effects of immigration on native employment and wages, most studies only minor displacement effects even after very large immigrant flows."

What you find in practice is both that immigrants accept lower wages, AND that local wages for natives are unaffected. Which happens because the labor market can generally take in the new labor comfortably.

trouve_search··on ADHD, autism or complex trauma? [pdf]
It's rarely reminded that when dealing with professionals (doctors, lawyers, accountants, etc.), it's important to get second opinions if you have any doubts.
trouve_search··on ADHD, autism or complex trauma? [pdf]
It's an article with no clear conclusion it's normal to feel confused.

My read of the article is balancing the fact that there's a lot of overlap between CPTSD/ADHD/ASD in the symptoms (emotional dysregulation, hyperarousal, etc) and in hereditary factors (undiagnosed parents causing trauma more often on average). Also that traumatic childhood experiences are more likely to stick around as CPTSD in adulthood if there's also neurodivergence.

The author says there can be incredible relief to be correctly diagnosed with ADHD/ASD, so obviously she says it's helpful.

But she also warns that wrong treatment can often happen (eg. Giving stimulants to perpetual fight or flight PTSD brains), or inneffective therapy.

trouve_search··on I can't stop thinking about Papua New Guinea
It's important to note that languages evolve much faster without a writing tradition to anchor the language down over time.
trouve_search··on Your car is selling your data
Couldn't you put some sort of faraday cage around the antenna?
trouve_search··on ChatGPT Is Throwing 404
pi.dev + anything else
trouve_search··on DiffusionGemma Technical Report
That's AMD's fault.

RDNA4 is pretty similar to CDNA4, yet over a year after the release of "pro AI" cards like the r9700, they had basic kernels lacking in vllm (like w4a16 int4 kernels) while they were implemented in the datacenter CDNA4 cards.

AMD hardware runs well on llama.cpp because basically anything runs on llama.cpp, especially with vulkan. It's not high praise of AMD's software team to say llama.cpp runs well on their hardware

trouve_search··on DiffusionGemma Technical Report
Yes, it runs in vllm happily. It gets >900TPS output reliably on a single 5090 with the nvfp4 model.

It's clearly worse than vanilla 26B-A4B, and lacks some things like structured outputs, and gets some tool calls wrong.

So you have to find a usecase or a hand rolled harness that leverages the cerebras-level TPS while not going off track during (even short) tasks.

trouve_search··on Qwen3.8 27B scores 52 on Artificial Analysis
thanks for posting your setup! I think it's smart to set the reasoning effort default to something saner in the base config.

Here's a VLLM command for 3.6 (I'll update to 3.8 today) to test out:

```

PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True \

vllm serve Qwen/Qwen3.6-27B-FP8 \

--dtype auto \

--kv-cache-dtype fp8 \

--enable-chunked-prefill \

--enable-prefix-caching \

--trust-remote-code \

--enable-auto-tool-choice \

--reasoning-parser qwen3 \

--tool-call-parser qwen3_coder \

--speculative-config '{"method":"qwen3_next_mtp","num_speculative_tokens":3}' \

--default-chat-template-kwargs '{

    "enable_thinking": true, 

    "reasoning_effort":"medium"

 }' \

 --tensor-parallel-size 2 \

 --max-model-len 250000 \

 --gpu-memory-utilization 0.9 \

 --max-num-batched 12000 \

 --max-num-seqs 24
```

I took the liberty of adding your reasoning effort chat template to my setup. You can play around with the last few parameters. In generall VLLM will be better in higher concurrency scenarios, so if you only use it for a personal vibe coding assistant and less as a general home model for task execution llama.cpp may be better.

trouve_search··on Qwen3.8 27B scores 52 on Artificial Analysis
Yes, I mentioned the setup, but on vllm you can only use TP with speculative decoding or pipeline parallelism without, so there's tradeoff to both.

I gave general numbers of what I'm getting above, the performance ratios seemed similar regardless of setup (eg. getting a AWQ-in4 quant on a single GPU vs PP without speculative decoding vs TP with speculative decoding).

Overall single GPU is fastest, and TP+speculative decoding is still faster than PP, but for fp8 models you need dual GPUs whether you want it or not.

trouve_search··on Qwen3.8 27B scores 52 on Artificial Analysis
I think it's a vllm vs llama_cpp performance thing, will pay more into it.

One note I had between the two is that gemma has a much higher prefix cache hit rate in general.

trouve_search··on Qwen3.8 27B scores 52 on Artificial Analysis
What configuration are you using? On both vllm and llama-cpp, I get significantly higher speeds from gemma4 than qwen3.6 (with their respective speculative decoding methods).

Output TPS in vllm for instance:

- Gemma4 26B-A4B: 200-300TPS

- Qwen3.6 35B-A3B: 120-180TPS

- Gemma4 31B: 80-120TPS

- Qwen3.6 27B: 60-80TPS

This is for a first request on a dual 5090 setup, with their respective speculative decoding methods.

trouve_search··on llama.cpp
Using 98.css would still leave you with the AI slop text wording.

The core problem is that some people don't even seem to notice / care.

trouve_search··on Nvidia Nemotron 3.5 Lightning and NeMo Switchyard
Laguna XS is MoE, however.
trouve_search··on Manus will return to operating as an independent company
Interesting, what does your tasks & workflow look like with them?

I generally found the quality decent (say, similar to other competitors), but the speed of task completion was very slow. I think because they would depend too much on Sonnet as a core backend, and relied on big/expensive models more than other harnesses.

trouve_search··on Manus will return to operating as an independent company
Does anyone here use Manus actively?

I found it worse than alternatives in the similar space (Claude, genspark, kagi research, etc.) and slightly baffling they were being acquired at a $2B valuation in the first place.

trouve_search··on DeepSeek costs OpenCode Go user $1.14/day; dual DGX breaks even in 24 years
The value prop really depends on what you're doing.

If you're just vibe coding with giant frontier models, yes, the value will be worse. Especially now, where GPU prices have spiked another 20% last month.

For some tasks where owning the setup and full kv cache matters, the payoff calculation is ridiculously in favor of running your own deployment.

For instance for some batch classifications jobs where the prefix cache hit rate will be >95%.

The calculus also changes if you just use AI as a light tool while coding and don't need the giant models; qwen3 27B runs at 80TPS on a 5090 properly deployed.

trouve_search··on The Absurdity of Albert Camus
I love these, thanks!
trouve_search··on The Absurdity of Albert Camus
Everyone has to find an answer to the meaning if life. For the non religious/escapist, you have to stare into the void and find an answer at some point.

Existentialism (Sartre) says you have to find your own meaning. Camus says there cannot be meaning; life is inherently "absurd".

Any of the existentialist philosophies will be adjacent to angsty teenager stuff; they dance closely with nihilism and cynicism.

trouve_search··on Read this before you buy that TV streaming stick
The nvidia shield is pretty damn good as well, even if old at this point.
trouve_search··on Be skeptical of OpenAI's rogue hacker agent story
From my reading, the sandbox escape came from the JS packages in the harness still having an internet connection (somehow!), the agent having access to the source of those packages, reading it and executing code from them to access the internet.
trouve_search··on I co-founded Wikipedia, but an anonymous mob runs the show – and now I'm banned
The fact that he could only get this story published in the Washington examiner of all places should be a signal that more reputable places don't want to attach their name to this
trouve_search··on OpenAI unveils its first custom chip, built by Broadcom
Cerebras is a whole lot of SRAM, basically a ton more L1/L2 cache, hence increasing throughput.

They're pretty supply constrained right now though and their production costs seem prohibitive.

The interesting players at the moment are from Toronto: taalas (print the model onto the silicon) and tenstorrent (dataflow programming based hardware)

trouve_search··on GLM 5.2 Performance Benchmarks
A lot of benchmarks are setup to not punish false positives (irrelevant answers or extra text) and punish false negatives (missing the snippet being looked for).

This leads to answer bloat and/or hallucination if you benchmaxx on those

trouve_search··on Running local models is good now
gemma 12B 4bit quant; try something with MTP and an AWQ quant
trouve_search··on Running local models is good now
On a 5090, gemma4 26B runs at 350TPS with the command below [1] and gemma4 31B is around 150TPS with a similar command.

I'm really surprised how much slower a DGX spark is for the same price.

1. Here's my command.

PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True \ vllm serve cyankiwi/gemma-4-26B-A4B-it-AWQ-4bit \ --dtype auto \ --gpu-memory-utilization 0.95 \ --kv-cache-dtype fp8 \ --enable-chunked-prefill \ --enable-prefix-caching \ --trust-remote-code \ --enable-auto-tool-choice \ --tool-call-parser gemma4 \ --reasoning-parser gemma4 \ --max-num-batched 16000 \ --max-model-len 64000 \ --max-num-seqs 12 --speculative-config '{"model": "./gemma-4-26B-A4B-it-assistant", "num_speculative_tokens": 4}'

trouve_search··on How is Groq raising more money?
Cerebras are only serving kimi for dedicated endpoint customers; for that you need a >$5m annual deal with them

Cerebras also seems to be killing off their regular APIs, they're deprecating models and GLM is still stuck on GLM 4.7, a whole 2 versions behind.

trouve_search··on It's hard to justify buying a Framework 12
Not sure if the M5 is that massively different but I have a M2 max laptop and the screen is noticeably brighter on the Asus.
trouve_search··on It's hard to justify buying a Framework 12
It hasn't been so bad for me to notice. Compiling rust you'll hear the fan, but you'll also hear it in a MBP.

The MBP will compile faster however.

trouve_search··on Domain expertise has always been the real moat
The moat is the difference between knowledge and know-how.

You can read all the plumbing books, but you need to get your hands dirty a few times, mess it up and fix it, to get mentally comfortable and efficient with the work

Page 1 of 2Next →