232 karma · joined June 7, 2021
I don't think so. It's true that human also write shitty code but the key difference is we actually remember what is the intention behind those crappy implementations so someone can fix it later. aka it is the matter of long term memory that currently LLM architecture is not capable of.
You can argue that claude can read the whole linux codebase and report bugs, but they can only report local bugs, not systematic one. 1M context windows seems like huge, but the effective range is actually pretty limited, and it still does not equal to human insight.
why the heck do you need 100+ instances of bert. do you even attempt to research about this before?
the laya paper show that you can do the similar stuff with jev using modern bert only: https://laya.convaiinnovations.com/
and even without the newer wave of applying llm techniques to the older bert models, even flan-t5 was trained for handling 1800+ tasks.
Or even llm if you claim about versatility. You can easily modify the llm inference code to make it predict a single token represent the classification choice and extract the probability that way.
Sure jev will still be faster, but a local deployed Bert model is way faster than both.
And to get the most out of it you still need to fine tune the models anyway, unless your classification task is just one of those mainstream ones.
It is pretty trivial to pin a single provider for a model. Better yet, instead of calling the model directly, use presets instead. You can easily change the setting on openrouter without having to update your app every time.
can't wait for deepseek v4.1 pro
it is even worse with 40 macbooks.
if 40 macbooks is all that take to serve a 1TB model with decent speed then you would see everyone selling the models for very cheap right now.
And even when searching for internet, it still cannot suggest a up-to-date approach to the problem.
For example I'm using crystal, it recently revamped the concurrency/parallel model. Even using web search, gemini still does not aware of the new feature and still give the outdated code.
I'm sure my crystal usage is not the unique case here.
2. people tend to ignore this, but the salary budget of a US frontier lab and chinese frontier lab is nowhere comparable, the first can easily outdone the later by 100x.
3. us labs, like other US style startups, always throw ton of money to capture the market. I don't see the chinese company doing the same scheme at all.
so, surely chinese AI providers also lost money making new models, but they are not spending nearly as much as US ones.
Actually have you really design practical webapps before? Because your post only have empty pretty words with no substance.
Semantic html tags are just semantic, it is irrelevant for the discussion.
Responsive design is hard, have you ever try to reize your windows and see how your pretty website behaves? Leave aside stupid people thinking you only need to care about phone device (in portrait mode even) and laptop, most people also think 3 or 4 device widths is enough. Sure some smart one already promoted to break the screen by content, not fixed device size, but even that comes with some downsides.
To make a website works decent on all the possible screen sizes is a lot of work, I spent like 4 months updating and fixing my webapp when I first developed it.
typically 60ch equal to 480px (font size 16px), so you need sidebars to the left and the right. which is fine, the holy grail was like that. but if you want to design a layout that look good in all screen resolutions, then a fixed layout does not work. you would have a dozen compromises and layout switches to make it work. at that point, forcing a block to have at max 60ch is painful.
and, if you do the main + sidebar layout, with maximum width for the whole site is around 960px (or 1200px if you have multiple sidebars) then suddenly you do not need to think about 60ch width anymore, because it naturally just fits (except for some width points)
It is funny that google gave up on this market, leaving the whole price range to Chinese models.
nobody ever reports that they installed it and ran it on production or something.
personally I only use bun to replace yarn as script executor now. Used to follow it and tried to replace node server sveltkit runtime with bun, but it didn't work reliably so I stopped the experiment.
it's too bad. when you are spoiled by crystal/go/rust written servers that consume 20MB-50MB ram only, seeing a nodejs server consume 400-500MB ram looks ridiculous bun was such an attractive option.
You can easily find them in Chinese tech forum linux.do
yet but it is still contain a lot of trash. you need better models to process those trash and create a curate dataset. this will happen again and again until there is no more juice to squeeze. and I'm sure we are still not done with it.
> post training
yeah this will be crucial. the big models are already too capable, they are just not that aligned with current agent tasks.
> parameter count doesn’t seem to be a direct correlation anymore
I don't think so, remember that chinese labs do not have as much compute power compare to US frontier labs. that's why deepseek v4 flash had that huge jump and deepseek v4 pro is kinda a disappointment, they just do not have the compute power to proper posttrain the pro model like they wanted. glm is also a relative small model so you also can see the huge jump with just post training. so it does not mean the size does not matter, it is just mean that the chinese labs currently only capable of training smaller models effectively.
and now we have ai agents to automatic migrate the system with new models. in the past we would need to spend hours to design the prompts, then test the output, then write codes to babysitting it. nowadays any ai agent can do it effortlessly.
this is hilarious. it is not 2025 any more, by Jan 2027 there will be at least 3 newer generation of models (from other provider) released already. nobody would use flash 3.7 at that time.
sure we used to cling to gemini models in the past, demanding 2.5 models to continue to serve, but since google betrayed us with those price hike, people already spent their time making their production pipeline less dependent on google since then.
heck, even now I'm not sure I even care if they cut the pricing even lower. there are too many models with cheaper price and similar performance now.