1,311 karma · joined August 4, 2017
meet.hn/city/de-Leipzig
When I visit a website, I'm usually looking for information and not for a message.
Isn't that a bit overgeneralized?
There's more than weights for the Olmo models for example: https://allenai.org/olmo
Similar for Nvidia's Nemotron models IIRC.
Artificial Analysis has an "openness" ranking: [1]
[1] https://artificialanalysis.ai/models?model-filters=open-sour...
Other providers don't.
You can also choose to route to ZDR only.
https://openrouter.ai/deepseek/deepseek-v4-flash-0731#provid...
Some providers like OpenRouter now call it `deepseek-v4-flash-0731`, but even in places like here on HackerNews people say things like "Sonnet is better than DeepSeek" without specifying a version or a reasoning effort, certainly no one will mention that `-0731` suffix when talking about DeepSeek V4 Flash.
The `-0731` style suffix is worse compared to a proper version bump like V4.1.
Why not call it V4.1?
You seem to have created a new GitHub account just for this project a week ago. Do you have any other GitHub accounts that enable us to see a track record of your work (maintenance, security)?
So in general this definitely works. Handy is just missing the feature to insert these streamed words into the app where the cursor is.
This is from Ministral 3 14B, a 2025 model without reasoning, that you can run on your PC:
> Write a Haiku involving HackerNews, and the capability of large language models like you to reply in an exact number of words or syllables.
Silicon whispers,
exact words in code’s embrace—
Haiku blooms anew.
Across multiple tries it got it wrong a couple times (by ~2 syllables). But syllables are extra tricky (because of how LLMs use tokens) and the point is that for things like "summarize in 5 bullet points" you will mostly get 5 bullet points, maybe 6, but not 10 or 20, and no need for a tool that count bullet points.For Bitcoin / Lightning these kind of pay-per-request API paywalls have existed for many years already (e.g. my own from 8 years ago [1], but others as well).
Flattr [2] existed for non-crypto micropayments.
None became mainstream. I think the friction is always the extra setup on the client side. In all 3 cases the user (API consumer) has to set up a special wallet (browser extension or something for the agent) and deposit some money/crypto on the client side first. This part needs to become simpler.
1. DeepSeek V3.2, V4 Flash, V4 Pro, at high or max thinking, ... when recommending a model it should always be a precise model, not just an AI lab
2. DeepSeek V4 Flash at max thinking is the most verbose model (among top models) in the AA benchmarks. See the "Intelligence Index Token Use" chart: [1]
[1]: https://artificialanalysis.ai/models?models=gpt-5-5-high%2Cg...
> On the model side, we applied FP4 quantization
> introduced DFlash, an efficient speculative decoding method based on block-level masked parallel prediction
> On the system side, TileRT perfectly adapts to the dynamic characteristics of these algorithms
> 1000+ tokens/s output [...] using just a single standard 8-GPU commodity node
> HTTP also allows the DuckDB-Wasm distribution to speak Quack natively! So DuckDB running in a browser can e.g., directly connect to a DuckDB instance running in an EC2 server using Quack.
For the most parts you just write the regular Markdown headers and paragraphs, embed images, insert tables etc without the need for any HTML tags, making it readable in source form. And if you want to embed an SVG file for example, which the author of the article mentions as one use case, you just embed the SVG directly, and people can render the Markdown in their favorite viewer.
Let's say you're viewing a raw Markdown file in VS Code. You come onto an HTML tag, so you hit Cmd+Shift+V to open the preview and that's it.
Of course for full-fledged web pages with interactive buttons and fully customized styling and all of that, which the author shows in some examples, this is not feasible. But you can get very far when you have mostly text/images/tables and just want to add some extras here and there.
[1] https://daringfireball.net/projects/markdown/syntax#html
https://artificialanalysis.ai/models?models=gpt-5-5%2Cgpt-5-...
Same with GPT-5: Latest 5.5, prior 5.4, or actually the original 5 (.0)?
You can't talk about model performance without specifying the exact model.
What do you mean with this? Maybe you are thinking of the old ".NET Framework" runtime, which only runs on Windows? Nowadays there is ".NET Core" which runs on macOS and Linux as well.
Do you mean a setup like:
client -> cloud(HAProxy+Varnish) -WireGuard-> basement(backend)
Or something else?