HNHacker News
TopNewBestAskShowJobs

redrove

883 karma · joined December 18, 2022

submissionscomments
redrove··on Show HN: Open-source model routing for coding agents at Astra-level performance
Is the model you trained available as open weights?
redrove··on Beware Management Consultants
I'm sorry but that's like saying communism is 'user error'.

If the ideal scenario never happens in practice, it's a flawed concept to begin with.

redrove··on Launch HN: machine0 (YC S26) – Persistent CPU and GPU VMs from the CLI
YC gives DO credits for startups.
redrove··on Degraded performance for multiple models
Could’ve just used fewer tokens and redirected to the Codex signup page.

ba dum tsss

(sorry couldn’t help myself)

redrove··on Qwen 3.8 27B is excellent, but it defaults to overthinking things
I’m not so sure I agree, given the overwhelming concentration of capital and regulatory capture they have, I just don’t see them going anywhere; becoming more niche rather than even more of a standard is “going away” to a certain extent as far as I can see.
redrove··on Qwen 3.8 27B is excellent, but it defaults to overthinking things
Thanks for the very detailed answer!

Would you say ChatGPT does everything you need today or do you still see some gaps? i.e. things you think it should be able to do but currently doesn’t, or things that still take too much effort on your part to setup chatgpt to do it.

redrove··on Qwen 3.8 27B is excellent, but it defaults to overthinking things
I would also be interested in what you use AI for. Is it marketing? bureaucracy?
redrove··on Qwen 3.8 27B is excellent, but it defaults to overthinking things
I don’t disagree but saying “cloud will change in the next 8-20y-ish” is a bit of a non-argument, you’re not really stating any thesis to speak of; Change is a given over that time frame.
redrove··on How Organizations Use AI: Evidence from ChatGPT [pdf]
I think you're supposed to use AI to digest it
redrove··on Delta
So..pseudo code?
redrove··on Grok Bot
Oh this looks very interesting, both for personal and work; I’ve been looking for something similar for quite a while.

However, no OpenAI API support (just Anthropic + openai.com) means I can’t use it for either.

redrove··on Grok Bot
I’ve tried using this as a self hosted instance and it’s been a little rough around the edges with Hermes.
redrove··on Melatonin impairs morning cognition in healthy young adults (2023)
Could it just be dosage? Commonly found dosage is 10-20x what you'd actually secrete and IIRC absorption is actually quite good.
redrove··on Show HN: SIEMatic, a fair-sourced observability and security platform
Is there a screenshot and a feature matrix of some sort? I couldn’t find on the docs site and I don’t understand what this tool is capable of, even though I’m familiar with SIEM as a general concept.
redrove··on LLMs reward expertise
Exactly, sounds to me like OP is the one with the LLM skill issue.
redrove··on Show HN: NixOS-DGX-Spark – Nix and NixOS on the DGX Spark
No worries! Would love to run K3 but lack the hardware for a 2.8T model, like most people.
redrove··on Show HN: NixOS-DGX-Spark – Nix and NixOS on the DGX Spark
I’d be happy to test this out if you want to publish an alpha branch or something.
redrove··on Show HN: NixOS-DGX-Spark – Nix and NixOS on the DGX Spark
The FP8 version from DeepSeek themselves [0], around 1800 tps prefill and 45 tokens per second decode.

I’ve been running a custom VLLM image with b12x as well as nvfp4_ds_mla.

I would say it’s quite fantastic in day to day, I use it mostly in Hermes and sometimes for coding.

I have qwen 3.6 27b on an rtx 6000 pro as well so I use that as a workhorse in pi with DS as a reviewer/planner.

[0] https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731

Edit: I think you may have misread my post. k3s is NOT kimi k3, and I did mention I was running deepseek.

redrove··on Show HN: NixOS-DGX-Spark – Nix and NixOS on the DGX Spark
Been running this on a few Asus GX10 machines with k3s on top, it’s been great. I’m running the new deepseek.

Thank you for your work!

redrove··on MicroVMs: Run isolated sandboxes with full lifecycle control
No, it doesn’t seem like it.
redrove··on U.S. government will decide who gets to use GPT-5.6
I’m sorry but I can’t take a European Commission link seriously about training SOTA LLMs.
redrove··on The worthlessness of Vitamin D is mildly exaggerated
Do you have any sources for this? Genuinely interested in reading more.
redrove··on The European Social Stack
No it’s not.

You’re generalizing, DACH != the entire EU.

redrove··on Local Qwen isn't a worse Opus, it's a different tool
> My dream would be a local model that can do, say, 80% of the day to day tasks I need; "how does X Handler connect to Y storage?", "commit that feature, but leave out the bits that relate to billing" etc.

Qwen 3.6 27B can do that today, but setup properly and in a good quant, I run an autoround [0] with weights in int8 and attention heads in f16 on a single RTX 6000 Pro Blackwell Max-Q via vllm with mtp=2 and full context, --max-num-seqs 3, KV in f16, mamba f32.

>It would have 99% reliable tool calling

I managed to score 93/100 in tool-eval-bench [1]. For me this is very good already, at least in the pi coding harness I've never had an issue that wasn't auto-fixed in the next turn(s).

>the ability to go "this task is beyond my skills" and refer to a Big Boy Online Model in a gigantic datacenter somewhere

This is heavy on the harness engineering side I think, but also quite contrary to the nature of LLMs today. If you figure this out I'd love to know.

[0] https://huggingface.co/Minachist/Qwen3.6-27B-INT8-AutoRound/...

[1] https://github.com/SeraphimSerapis/tool-eval-bench

redrove··on GPT‑NL: a sovereign language model for the Netherlands
They’re still debating that, they’ll get back to us soon I’m sure.
redrove··on Home alone: Remote work, isolation, and mental health
I read the entire thing and it seems to me like they started with the conclusion and tried to find proof, like a lot of psychology papers.

Thoroughly unreliable.

redrove··on MAI-Code-1-Flash
It’s about bang for buck. That high a score for 5B params is pretty good, nigh unbelievable a short while ago.

It is my belief that smaller models will get better and better, and even cloud SOTA models will shrink.

Yet another reason the current buildout will feel like the railroads.

redrove··on Odysseus – self-hosted AI workspace
Wow, if you you could've just done docker compose up. I guess we'll never know.
redrove··on Can we have the day off?
So the theory goes. Can you point out one place where this is working out?
redrove··on I bypassed AWS API Gateway auth with a trailing slash. Got $12K bounty
Don’t vibe code your auth path folks.
Page 1 of 10Next →