HNHacker News
TopNewBestAskShowJobs

npn

232 karma · joined June 7, 2021

submissionscomments
npn··on DeepSeek Harness Desktop for macOS and Windows
He thinks his ideas are still valuable and wants to sell them somehow.
npn··on AI Has No Wisdom and Neither Will You
> Sad to say, but this is no different from human written code.

I don't think so. It's true that human also write shitty code but the key difference is we actually remember what is the intention behind those crappy implementations so someone can fix it later. aka it is the matter of long term memory that currently LLM architecture is not capable of.

You can argue that claude can read the whole linux codebase and report bugs, but they can only report local bugs, not systematic one. 1M context windows seems like huge, but the effective range is actually pretty limited, and it still does not equal to human insight.

npn··on Kev: Tiny Jev-like family of decision models built on top of Qwen3.5
> And even if BERT wasn't woefully underintelligent for the task... have 100+ instances of BERT running locally faster than Jev API response times? Sweet rig you must have...

why the heck do you need 100+ instances of bert. do you even attempt to research about this before?

the laya paper show that you can do the similar stuff with jev using modern bert only: https://laya.convaiinnovations.com/

and even without the newer wave of applying llm techniques to the older bert models, even flan-t5 was trained for handling 1800+ tasks.

npn··on Kev: Tiny Jev-like family of decision models built on top of Qwen3.5
Show me a single example how is this jev thing better than a modern Bert solution?

Or even llm if you claim about versatility. You can easily modify the llm inference code to make it predict a single token represent the classification choice and extract the probability that way.

Sure jev will still be faster, but a local deployed Bert model is way faster than both.

And to get the most out of it you still need to fine tune the models anyway, unless your classification task is just one of those mainstream ones.

npn··on Kev: Tiny Jev-like family of decision models built on top of Qwen3.5
You got it reversed. Bert is encoder only and gpt is decoder only.
npn··on Bend – A language that blocks AI mistakes via proof, on CPU and GPU
I read the readme and the guide file. There is just one thing I can comment: might as well solve the NP hard problems. I think you can do it easily, author. As you can already solved harder problems than those with your language.
npn··on Ask HN: How can DigitalOcean not afford $50, but can afford $3M to Omarchy?
People do love auditing everyone else’s donations.
npn··on So you want to use OpenRouter?
As expected vibecoding bros cannot even read the manual properly.

It is pretty trivial to pin a single provider for a model. Better yet, instead of calling the model directly, use presets instead. You can easily change the setting on openrouter without having to update your app every time.

npn··on DeepSeek v4.1 Flash
it is a way bigger model with extra 200B engram so of course the score improves.

can't wait for deepseek v4.1 pro

npn··on DeepSeek launching v4.1 flash cheaper and more capable than v4 pro
Crazy that they still keep the price -- or actually decrease it, even -- despite it is a big improvement. I hope it retains some of the tps speed of the preview release though, 300 tps means gemini flash is no longer "the fastest option" any more.
npn··on Kimi K3 (2.8T) at 1 token/s on a MacBook Pro, streamed from four SSDs
yeah raid or not you still get the hard limitation by the pcie lanes

it is even worse with 40 macbooks.

if 40 macbooks is all that take to serve a 1TB model with decent speed then you would see everyone selling the models for very cheap right now.

npn··on Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out
luckily I have purchased a used amd mi50 32gb card for pretty cheap back then. while I haven't used it extensively, it feels pretty great having a backup plan that does not depend on any external 3rd parties.
npn··on Muse Spark 1.3
openai did human crafted chain of thought dataset training. deepseek didn't have the resources so they attempted RL. doing RL correctly is hard because of the risk of model collapsing.
npn··on Gemini 3.8 Flash and 3.8 Flash Cyber
very important actually. just try to generate code for fresher frameworks/libraries. gemini sucks so bad in real work usage, everything it suggests are outdated and mostly useless.
npn··on Gemini 3.8 Flash and 3.8 Flash Cyber
Still refuse to search internet for stuff it thinks does not exist lol.

And even when searching for internet, it still cannot suggest a up-to-date approach to the problem.

For example I'm using crystal, it recently revamped the concurrency/parallel model. Even using web search, gemini still does not aware of the new feature and still give the outdated code.

I'm sure my crystal usage is not the unique case here.

npn··on Our decision on Cursor following its acquisition by SpaceX
Because you can also use grok and a dozen other models. In fact grok is preferred choice for cursor right now, so obviously the other models get sidelined
npn··on GLM-5.3-Flash Intelligence, Performance and Price Analysis
1. they still have revenue though. it might not enough to cover all the r&d but it is surely enough to cover the hardware cost.

2. people tend to ignore this, but the salary budget of a US frontier lab and chinese frontier lab is nowhere comparable, the first can easily outdone the later by 100x.

3. us labs, like other US style startups, always throw ton of money to capture the market. I don't see the chinese company doing the same scheme at all.

so, surely chinese AI providers also lost money making new models, but they are not spending nearly as much as US ones.

npn··on Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)
I'm confused? Can you just define some presets and call them instead? With preset you can pinpoint a lot of things, especially the providers
npn··on Death to px, long live ch
? Why are you assuming that I don't know about that.

Actually have you really design practical webapps before? Because your post only have empty pretty words with no substance.

Semantic html tags are just semantic, it is irrelevant for the discussion.

Responsive design is hard, have you ever try to reize your windows and see how your pretty website behaves? Leave aside stupid people thinking you only need to care about phone device (in portrait mode even) and laptop, most people also think 3 or 4 device widths is enough. Sure some smart one already promoted to break the screen by content, not fixed device size, but even that comes with some downsides.

To make a website works decent on all the possible screen sizes is a lot of work, I spent like 4 months updating and fixing my webapp when I first developed it.

npn··on Death to px, long live ch
what hard is making a fully featured website, with panels no wider than 60ch.

typically 60ch equal to 480px (font size 16px), so you need sidebars to the left and the right. which is fine, the holy grail was like that. but if you want to design a layout that look good in all screen resolutions, then a fixed layout does not work. you would have a dozen compromises and layout switches to make it work. at that point, forcing a block to have at max 60ch is painful.

and, if you do the main + sidebar layout, with maximum width for the whole site is around 960px (or 1200px if you have multiple sidebars) then suddenly you do not need to think about 60ch width anymore, because it naturally just fits (except for some width points)

npn··on Death to px, long live ch
I also did some experiment with ch many years ago. I found that 60ch is ideal width for block text for easy reading. too bad it is pretty hard to make websites with only 60ch wide.
npn··on Ox Alpha
didn't zai already do that with their coding plan? I mean they surely had to pay users to use claude models (paying the differences). they also funded some newapi token resale websites.
npn··on Ox Alpha
I hope it is glm air. We need more "small" models. Big models are more capable and useful, but for majority of tasks some smaller models can work just fine.

It is funny that google gave up on this market, leaving the whole price range to Chinese models.

npn··on Ox Alpha
Ok that antise... We all know who you really want to criticize here.
npn··on AI companies destroy physical books – let's scan rare books before it's too late
No but with 100% clean data you can easily train a model to filter ai generated content.
npn··on Bun 1.4
I wonder how many posts in this thread are AI generated or shill posted.

nobody ever reports that they installed it and ran it on production or something.

personally I only use bun to replace yarn as script executor now. Used to follow it and tried to replace node server sveltkit runtime with bun, but it didn't work reliably so I stopped the experiment.

it's too bad. when you are spoiled by crystal/go/rust written servers that consume 20MB-50MB ram only, seeing a nodejs server consume 400-500MB ram looks ridiculous bun was such an attractive option.

npn··on The AI Credit Resale Economy
There are like thousands sites with similar features all using newapi core.

You can easily find them in Chinese tech forum linux.do

npn··on GLM-5.3: Frontier coding with emergent cyber capabilities
> used up internet-scale data

yet but it is still contain a lot of trash. you need better models to process those trash and create a curate dataset. this will happen again and again until there is no more juice to squeeze. and I'm sure we are still not done with it.

> post training

yeah this will be crucial. the big models are already too capable, they are just not that aligned with current agent tasks.

> parameter count doesn’t seem to be a direct correlation anymore

I don't think so, remember that chinese labs do not have as much compute power compare to US frontier labs. that's why deepseek v4 flash had that huge jump and deepseek v4 pro is kinda a disappointment, they just do not have the compute power to proper posttrain the pro model like they wanted. glm is also a relative small model so you also can see the huge jump with just post training. so it does not mean the size does not matter, it is just mean that the chinese labs currently only capable of training smaller models effectively.

npn··on Gemini 3.7 Flash
it is partly true, but like I said it is not 2025 anymore. models now get released more often, and still have notable progress so they can safely replace the old models while being faster/cheaper. and thank to chinese models the pricing is pretty much stable and affordable now.

and now we have ai agents to automatic migrate the system with new models. in the past we would need to spend hours to design the prompts, then test the output, then write codes to babysitting it. nowadays any ai agent can do it effortlessly.

npn··on Gemini 3.7 Flash
> * For 3.6 and 3.7 Flash, introductory price expires on December 31, 2026. Starting January 1, 2027, $1.50/1M input tokens and $7.50/1M output tokens will apply.

this is hilarious. it is not 2025 any more, by Jan 2027 there will be at least 3 newer generation of models (from other provider) released already. nobody would use flash 3.7 at that time.

sure we used to cling to gemini models in the past, demanding 2.5 models to continue to serve, but since google betrayed us with those price hike, people already spent their time making their production pipeline less dependent on google since then.

heck, even now I'm not sure I even care if they cut the pricing even lower. there are too many models with cheaper price and similar performance now.

Page 1 of 8Next →