HNHacker News
TopNewBestAskShowJobs

Philpax

9,356 karma · joined October 1, 2013

submissionscomments
Philpax··on Kusama Yayoi has died
No.
Philpax··on Omarchy development practices lead to predictable security issues
"political histrionics" is an interesting way to refer to calling for ethnic cleansing
Philpax··on Omarchy development practices lead to predictable security issues
https://jakelazaroff.com/words/dhh-is-way-worse-than-i-thoug...
Philpax··on Behaviorally fingerprinting Ox Alpha's provenance
Stealth launch: builds hype, allows them to collect user preference data and see where the model fails. Why people care: it's free, decent, and people love a good mystery.
Philpax··on Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)
Strongly recommend https://github.com/Neroued/ninfer, which can pull ~180 TPS on 5090 with 3.8, and 500 (!) with 3.6 35B-A3B.
Philpax··on Building an (almost) fully self-hosted, sandboxed, agentic software factory
Using a RTX 5090 ($3~4k), I can run Qwen 3.8 27B at ~180 TPS with ninfer [0]. With its thinking maxed out, I can confirm that the quality of output is roughly on par with Opus 4.5~4.6 - that is, this 20GB file really can write software by itself, but the amount of thinking required makes it strictly slower than larger models, even at 180 TPS.

Still, it's largely replaced the cheap tier of the frontiers that I would otherwise be using. It can be run with older GPUs (a 3090 is ~1k), but the time spent thinking will become a fairly noticeable impediment for staying in the flow.

The next step up would be to run DeepSeek V4 Flash 0731 on two DGX Sparks (~$10k), which serve at 60 TPS and sit somewhere around Opus 4.7 level without 3.8-tier thinking.

However, it is worth noting that, if you are buying this hardware just to serve LLMs, it is not cost-effective. It would take over a decade of continuous use to make back the cost of the DGX Spark setup in 0731 tokens. I'm running this setup because I happen to have a 5090, and the two 3090s in my server were cheap enough when amortised over several years.

[0]: https://github.com/Neroued/ninfer

Philpax··on DiffusionGemma Technical Report
DFlash 2 is a diffusion-based speculative decoding head for autoregressive models; there's nothing to accelerate here, because this is already wholly diffusion.
Philpax··on Maximizing the value of your Claude Code sessions
> Are they literally adding a hidden system prompt that says "effort level: $level" ?

Yes. https://magazine.sebastianraschka.com/p/controlling-reasonin...

Philpax··on Accelerating GPT-5.6 Sol Ultrafast
I think it's pretty obvious that, in that world, the AIs will simply be tasked with making the compilers faster. It's already happening with their own stack, after all.
Philpax··on Grok 4.6
I think you might be underestimating how many people genuinely despise Musk.
Philpax··on Qwen/Qwen3.8-2.4T-A95B
I was under the impression that you could fit the full 1M context within the 192GB VRAM as a result of DeepSeek's various architectural advancements, but I'll grant that DSpark + a larger pool for concurrency may necessitate more VRAM, yes.
Philpax··on Qwen3.8-2.4T
What do you need the extra 2 for? Tensor parallelism?
Philpax··on Learning more about Claude's mathematical capabilities
For me, personally, it's that the Bun guy - specifically him, not a mathematician - indirectly progressed the Riemann hypothesis by repeatedly telling a model to ganbatte!

It's a ridiculous position we find ourselves in.

Philpax··on Learning more about Claude's mathematical capabilities
> Jarred Sumner, an Anthropic staff member (and non-mathematician) prompted Claude to “take a real stab” at the hypothesis itself, leaving the mathematical choices from there up to the model. Initially, Claude generated and tried 650 ideas, none of which worked. Jarred prompted Claude to try again, and it spent a day and a half coordinating about 60 Claude subagents, which this time went much deeper: between them, they ran 2,400 shell commands and wrote hundreds of Python scripts.1 The subagents ran thousands of numerical checks against known zeta zeros and refereed one another’s work. Throughout this process, Jarred's input was mostly limited to sending Claude messages of encouragement (mostly variants of “keep going” or “believe in yourself”).2 This seems to have helped Claude overcome some initial skepticism that it could make meaningful progress.

The world we live in is beyond parody.

Philpax··on Mea Culpa – Dark Hours
When we have evidence that that's what actually happened and it's not the known-deceitful author trying to cover their own arse again?
Philpax··on Pi's Minimalism Is Its Advantage
In my experience, the Rust community very much cares - it is a struggle to find a modern Rust application of any popularity that does not abide by XDG.

.cargo's placement is a historical mistake that can't be undone now, but ecosystem participants are generally good participants.

Philpax··on Waymo in Dallas
Are you running your replies through a LLM, or is the LLM producing them entirely?
Philpax··on Show HN: Fuse – statically typed functional programming language
I've only used it for small scripts and can't say how well the type system actually functions/how it scales, but what I've seen so far has pleased me. (They have a _comptime_-equivalent! https://luau.org/types/type-functions/)

You may be interested in https://lute.luau.org/, which is a node.js-style runtime for the language.

Philpax··on Ten advances in mathematics and theoretical computer science
4. I think they can, especially if the problem statement is well-specified and, importantly, autonomously testable. Of course, specifying a problem that meets these requirements is non-trivial, but the claim requests _a_ counterexample :P
Philpax··on Show HN: Fuse – statically typed functional programming language
Luau is statically typed and interpreted/JIT'd: https://luau.org/
Philpax··on Anime User Interfaces
Copeland! I was quite surprised to see that myself.
Philpax··on Get ready to flee, Americans in ten countries warned
You may not want to engage with politics, but politics will engage with you.
Philpax··on DeepSeek-V4-Flash Update
That will cease to be a problem in the next 24 hours, now that the weights are out: https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731
Philpax··on Our position on open-weights models
This is where I'm at, or rather, will be. I like open-weight models and I've done my part in facilitating them, but if you believe in the increasing capability of these systems - and I do, to some measured extent - it seems plausible to me that an incident will happen at some point in the future.

I don't agree with his argument as a whole, especially not on some of the specifics (it is not great that this technology is being developed under the current US government), but I am sympathetic to the idea that some bells can't be unrung, and thus we should proceed with caution.

Philpax··on OpenAI and Hugging Face address security incident during model evaluation
You... haven't shown any evidence that we're near collapse. That's what I'm asking you for. Show me some evidence that we are losing capabilities with we have today.
Philpax··on OpenAI and Hugging Face address security incident during model evaluation
Assumes facts not in evidence. Please show your working.
Philpax··on OpenAI and Hugging Face address security incident during model evaluation
Over the last two years, this weight class has doubled its scores and/or saturated several benchmarks in the Qwen lineup alone without loss of generality: https://claude.ai/public/artifacts/9f249169-3623-417e-86cd-7...

There is undoubtedly a limit somewhere (there is only so much you can pack into a given size) but it's really not particularly clear where that limit is. I don't think it's superintelligence - that much I agree with you - but I think "We already have a 1gb model that is as capable as it will ever be" is strictly false.

Philpax··on OpenAI and Hugging Face address security incident during model evaluation
Please source this claim. What 1GB models are capable of has increased generation-on-generation.

> For example: you can't make a mice-sized brain as smart as a human brain no matter how hard you try.

Sure. We don't know where the ceiling is for our digital minds, though.

Philpax··on OpenAI and Hugging Face address security incident during model evaluation
> This is science fiction, these models don't have access to their own weights.

The models are being used to train, and improve the infrastructure for training, other models [0][1]. Several RL techniques rely on using the currently-being-trained weights as part of their process. I really would not take "don't have access" as a given, especially during the training phase.

> What would be a lot more scary is a model as capable as sol that's able to run on consumer hardware without taking up several terabytes of storage, but of course that is simply not possible as we need 4t parameters to even begin emulating a small fraction of what a human brain can do.

The Poolside Laguna S 2.1 model [2] purports to compete with models several times its size, and inference compute is becoming increasingly plentiful. Again, would not hold anything here as a given.

[0]: https://openai.com/index/gpt-5-6/ ("GPT-5.6 accelerates OpenAI")

[1]: https://www.kimi.com/blog/kimi-k3#coding

[2]: https://poolside.ai/blog/introducing-laguna-s-2-1

Philpax··on OpenAI and Hugging Face address security incident during model evaluation
I've been rewatching Person of Interest for related reasons, and it hits uncomfortably close to things that are playing out today (e.g. https://youtu.be/zRL2sRkUvYk)

We live in interesting times.

← PreviousPage 2 of 34Next →