HNHacker News
TopNewBestAskShowJobs

PhilippGille

1,311 karma · joined August 4, 2017

https://github.com/philippgille

meet.hn/city/de-Leipzig

submissionscomments
PhilippGille··on GitLab Outage
It would be nice to disclose that you're the creator of the tool which you are promoting here.
PhilippGille··on Opus 5.5 is good at explainer videos
Websites can just present information.

When I visit a website, I'm usually looking for information and not for a message.

PhilippGille··on Android 17 is the first since 3.x to add new APIs without releasing to the AOSP
https://grapheneos.org/faq#upstream

https://android-review.googlesource.com/q/status:merged+auth...

PhilippGille··on Corporate America is getting hooked on open-source AI
> Open source does not apply to AI

Isn't that a bit overgeneralized?

There's more than weights for the Olmo models for example: https://allenai.org/olmo

Similar for Nvidia's Nemotron models IIRC.

Artificial Analysis has an "openness" ranking: [1]

[1] https://artificialanalysis.ai/models?model-filters=open-sour...

PhilippGille··on Show HN: My Claude quota ran out in 10 minutes, so I made a tool to find out why
Token usage is generally lower in greenfield projects.
PhilippGille··on GLM-5.3-Flash
NovitaAI is already hosting it: https://openrouter.ai/z-ai/glm-5.3-flash#providers
PhilippGille··on GeForce Now on Linux Native App Exits Beta
Official blog post: https://blogs.nvidia.com/blog/geforce-now-thursday-linux-nat...
PhilippGille··on DeepSeek costs OpenCode Go user $1.14/day; dual DGX breaks even in 24 years
According to OpenRouter, DeepSeek trains on input. At least when disabling all providers that train on your data, DeepSeek gets disabled.

Other providers don't.

You can also choose to route to ZDR only.

https://openrouter.ai/deepseek/deepseek-v4-flash-0731#provid...

PhilippGille··on OpenAI Daybreak Blue
Existing recent discussion with 127 points, 68 comments: https://news.ycombinator.com/item?id=49246704
PhilippGille··on GitHub Actions and Pages are experiencing degraded availability
98.33 according to https://mrshu.github.io/github-statuses/
PhilippGille··on DeepSeek-V4-Flash Update
Yes that's my point. The old and the new version are different in capabilities, but now when someone talks about DeepSeek V4 Flash (in benchmarks, on inference providers), you don't know which exact version it's about.

Some providers like OpenRouter now call it `deepseek-v4-flash-0731`, but even in places like here on HackerNews people say things like "Sonnet is better than DeepSeek" without specifying a version or a reasoning effort, certainly no one will mention that `-0731` suffix when talking about DeepSeek V4 Flash.

PhilippGille··on DeepSeek-V4-Flash Update
That's what I mean. On DeepSeek it's now just `deepseek-v4-flash`, while OpenRouter calls it `deepseek/deepseek-v4-flash-0731`, so now when someone talks about DeepSeek V4 Flash, like in benchmarks, or other inference providers, which version do they actually mean?

The `-0731` style suffix is worse compared to a proper version bump like V4.1.

PhilippGille··on DeepSeek-V4-Flash Update
The previous V4 version wasn't called “Preview” by most inference providers. For example, the OpenRouter model slug was `deepseek/deepseek-v4-flash`. So now there will be confusion when someone talks about V4 Flash or when someone offers V4 Flash inference.

Why not call it V4.1?

PhilippGille··on Advancing the price-performance frontier with GPT‑5.6
Depends on the reasoning effort, see https://deepswe.datacurve.ai (add Luna via model selection drop down, if it's not shown by default)
PhilippGille··on Show HN: Bento - An entire PowerPoint in one HTML file (edit+view+data+collab)
The project looks very interesting, thanks for sharing!

You seem to have created a new GitHub account just for this project a week ago. Do you have any other GitHub accounts that enable us to see a track record of your work (maintenance, security)?

PhilippGille··on Transcribe.cpp
Handy already supports streaming transcription models, and you can see the words in the small Handy pop-up while you are talking.

So in general this definitely works. Handy is just missing the feature to insert these streamed words into the app where the cursor is.

PhilippGille··on GPT-5.6
That was the case for early models (Llama etc), but they got much better since then. Not perfect, but good enough.

This is from Ministral 3 14B, a 2025 model without reasoning, that you can run on your PC:

> Write a Haiku involving HackerNews, and the capability of large language models like you to reply in an exact number of words or syllables.

    Silicon whispers,
    exact words in code’s embrace—
    Haiku blooms anew.
Across multiple tries it got it wrong a couple times (by ~2 syllables). But syllables are extra tricky (because of how LLMs use tokens) and the point is that for things like "summarize in 5 bullet points" you will mostly get 5 bullet points, maybe 6, but not 10 or 20, and no need for a tool that count bullet points.
PhilippGille··on Monetization Gateway: Charge for any resource behind Cloudflare via x402
Currently this is for payments with stablecoins.

For Bitcoin / Lightning these kind of pay-per-request API paywalls have existed for many years already (e.g. my own from 8 years ago [1], but others as well).

Flattr [2] existed for non-crypto micropayments.

None became mainstream. I think the friction is always the extra setup on the client side. In all 3 cases the user (API consumer) has to set up a special wallet (browser extension or something for the agent) and deposit some money/crypto on the client side first. This part needs to become simpler.

[1] https://github.com/philippgille/ln-paywall

[2] https://en.wikipedia.org/wiki/Flattr

PhilippGille··on GLM-5.2 is a step change for open agents
> Kimi and GLM models have coined a new term: Thinkslop. > [...] > So for now I'm happy with just two models: GPT and DeepSeek.

1. DeepSeek V3.2, V4 Flash, V4 Pro, at high or max thinking, ... when recommending a model it should always be a precise model, not just an AI lab

2. DeepSeek V4 Flash at max thinking is the most verbose model (among top models) in the AA benchmarks. See the "Intelligence Index Token Use" chart: [1]

[1]: https://artificialanalysis.ai/models?models=gpt-5-5-high%2Cg...

PhilippGille··on MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
The interesting bits on how they achieved it:

> On the model side, we applied FP4 quantization

> introduced DFlash, an efficient speculative decoding method based on block-level masked parallel prediction

> On the system side, TileRT perfectly adapts to the dynamic characteristics of these algorithms

> 1000+ tokens/s output [...] using just a single standard 8-GPU commodity node

PhilippGille··on MiniMax M3: The First Open-Weights Model to Combine Three Frontier Capabilities
The blog post has more info: https://www.minimax.io/blog/minimax-m3
PhilippGille··on Step 3.7 Flash
Do you mean MiMo V2 Flash? V2.5 doesn't have a Flash version.
PhilippGille··on Quack: The DuckDB Client-Server Protocol
It's in the article:

> HTTP also allows the DuckDB-Wasm distribution to speak Quack natively! So DuckDB running in a browser can e.g., directly connect to a DuckDB instance running in an EC2 server using Quack.

PhilippGille··on Using Claude Code: The unreasonable effectiveness of HTML
Both the original Markdown spec [1] as well as CommonMark [2] clearly specify support for inline HTML. With that you can kind of get the best of both words depending on your use case.

For the most parts you just write the regular Markdown headers and paragraphs, embed images, insert tables etc without the need for any HTML tags, making it readable in source form. And if you want to embed an SVG file for example, which the author of the article mentions as one use case, you just embed the SVG directly, and people can render the Markdown in their favorite viewer.

Let's say you're viewing a raw Markdown file in VS Code. You come onto an HTML tag, so you hit Cmd+Shift+V to open the preview and that's it.

Of course for full-fledged web pages with interactive buttons and fully customized styling and all of that, which the author shows in some examples, this is not feasible. But you can get very far when you have mostly text/images/tables and just want to add some extras here and there.

[1] https://daringfireball.net/projects/markdown/syntax#html

[2] https://spec.commonmark.org/0.31.2/#html-blocks

PhilippGille··on DeepSeek 4 Flash local inference engine for Metal
On max it uses more than twice as many tokens as on high when running the ArtificialAnalysis benchmark suite, and then it's indeed the model with the highest token usage (among the current top tier models). See the "Intelligence vs. Token Use" chart here:

https://artificialanalysis.ai/models?models=gpt-5-5%2Cgpt-5-...

PhilippGille··on Ask HN: Best Embedding Models?
Benchmarks only paint part of the picture, but it's still a decent place to start looking into recent models:

https://huggingface.co/spaces/mteb/leaderboard

PhilippGille··on DeepSeek v4
When you say "Gemini", which exact model do you mean? You know there are several and they vary a lot in how capable they are? Pro 3.1 Preview, 2.5 Pro (their latest non-preview pro model), Flash 3 Preview, ...

Same with GPT-5: Latest 5.5, prior 5.4, or actually the original 5 (.0)?

You can't talk about model performance without specifying the exact model.

PhilippGille··on High-Level Rust: Getting 80% of the Benefits with 20% of the Pain
> C# [...] only really works properly in Windows

What do you mean with this? Maybe you are thinking of the old ".NET Framework" runtime, which only runs on Windows? Nowadays there is ".NET Core" which runs on macOS and Linux as well.

PhilippGille··on I run multiple $10K MRR companies on a $20/month tech stack
He specifically mentions that he is using GitHub Copilot because of how Microsoft bills per request instead of token.
PhilippGille··on Old laptops in a colo as low cost servers
> it is possible with some software to have everything massively cached, with the cloud doing that, with the origin server in my basement, only accessible from the allowed cache arrangement

Do you mean a setup like:

    client -> cloud(HAProxy+Varnish) -WireGuard-> basement(backend)
Or something else?
Page 1 of 14Next →