HNHacker News
TopNewBestAskShowJobs

osti

584 karma · joined January 5, 2016

submissionscomments
osti··on DeepSeek V4 Peak Valley Pricing Change
This announcement also mentioned that they will release the next version (official non preview version) of v4 in mid July.
osti··on Why can't India's government build a decent website?
Is it only the Indian government? I don't think that's in any way unique to India, I've seen many poor government websites.
osti··on Bipartite Matching Is in NC
Ah got it. Reread previous comment and that makes sense.
osti··on Previewing GPT‑5.6 Sol: a next-generation model
Sol? Looks like openai is jealous of anthropics good model naming ability and wants to emulate it.
osti··on Bipartite Matching Is in NC
Hmm your last sentence seems to exactly agree that it's a class of algos that parallelize well? What does sped up arbitrarily mean? It's still polynomial speed up right?
osti··on Bipartite Matching Is in NC
So is it a class of problems that can be parallelized well?
osti··on Windows 10 quietly gets one more year of support and updates
Does that support modern gaming?
osti··on Anthropic says Alibaba illicitly extracted Claude AI model capabilities
I read quite a few posts about this on RedNote.
osti··on Anthropic says Alibaba illicitly extracted Claude AI model capabilities
That's not true, some of them are indeed fake, but a lot of them are actually providing real opus at low cost doing what op said.
osti··on GLM-5.2 is a step change for open agents
Even as a GLM z.ai fan, I wouldn't pay for their plans. They are just way worse values than gpt or anthropic plans, in terms of both usage and capabilities.
osti··on Steam Machine launches today
If AI is so good, then why don't we have perfect compatibility layers between OS's for games yet?
osti··on GLM-5.2 is the new leading open weights model on Artificial Analysis
The official API is FP8, which should imply that it's lossless.
osti··on GLM-5.2 is the new leading open weights model on Artificial Analysis
There are definitely some Chinese bots + actual people (imagine that!) who like to talk up Chinese models, I'm one of them but I like to find out how good these models really are before saying anything.

GLM definitely isn't opus level yet but it's for sure good. I think it lacks some knowledge (when coding) that the frontier models possess, which is expected given that the model is probably quite small when compared to the frontier.

But people don't say much about Mistral, probably because they are nowhere as good.. And they don't have large population behind them to actually use them.

osti··on GLM-5.2 is the new leading open weights model on Artificial Analysis
Fun fact: Zhipu aka Z.ai, Knowledge Atlas etc., the company that made GLM, is listed on Hong Kong stock exchange, is up over 10x since the IPO at the beginning of this year.
osti··on GLM-5.2 is the new leading open weights model on Artificial Analysis
I indeed got a few timeouts yesterday using the official API, I imagine for the coding plan users it'll be even worse.
osti··on GLM-5.2 is the new leading open weights model on Artificial Analysis
Many other open source models have vision but they don't compare to GLM in terms of coding quality. So I don't think it's because of vision that the frontier models are better, it's more that they are probably just much bigger models.
osti··on GLM-5.2: Frontier Intelligence, Open Weights
I don't know what people mean when they say design lol, is it for frontends?
osti··on GLM-5.2: Frontier Intelligence, Open Weights
Given that DeepSwe is one of the very few coding benchmarks worth taking a look at, this achieves rather excellent result at it (not far from opus 4.8).

From looking at the results and my own impression of 5.1 and other models, I think this is the best Chinese coding model by some non-insignificant margin.

osti··on Z.ai GLM 5.2
Given that DeepSwe is one of the very few coding benchmarks worth taking a look at, this achieves rather excellent result at it (not far from opus 4.8).

From looking at the results and my own impression of 5.1 and other models, I think this is the best Chinese coding model by some non-insignificant margin.

osti··on Cohere's First Model for Developers
Well, the level of tech is at least on a whole different level at those companies than whatever cohere is doing.
osti··on Anthropic requires 30 day data retention for Fable and Mythos
Hmmm no? The only way is to deploy your own local model, using anyone else's you are at their whim on what happens to your data.
osti··on Claude Fable 5
And notable absence of DeepSWE benchmark where they do badly, but somehow a benchmark that was published yesterday is in this announcement.
osti··on MiMo-v2.5-Pro-UltraSpeed: 1T model with 1000 tokens per second
But I think the eventual goal is that documentations won't even be needed. LLM should just itself understand the nuances of frameworks by analyzing their codebase.
osti··on I design with Claude more than Figma now
I think Jane Street is an Anthropic investor, so take it fwiw.
osti··on Claude Opus 4.8
Yeah I guess it being vague is more what I meant. But even if you told AI you need to wash the car, then why are you asking AI in the first place whether you should walk there or drive there. The question just doesn't make too much sense to me, doesn't look like it makes sense to the AI's either.
osti··on Claude Opus 4.8
Meh, I feel that the car wash test is probably the worst question of all of those LLM test questions. The question is basically logically inconsistent and expect the model to work around the inconsistency.
osti··on I think Anthropic and OpenAI have found product-market fit
For coding I wouldn't say a year, last year this time claude or gpt definitely weren't able to do what GLM is able to do today, but easily 6 months I'd say.

Not sure about other domains though.

osti··on Fuck You, Bambu Lab. Go Ahead, Sue Us
Reality has both negatives and positives. Gamers Nexus clearly lie on the other side of the spectrum where they overwhelmingly choose the negative stuffs, hence rage baits.
osti··on Fuck You, Bambu Lab. Go Ahead, Sue Us
I understand their points. But to me they do too much rage baits these days that I can't bring myself to watch their stuff. Why would I let my self get angry lol.
osti··on Bun's experimental Rust rewrite hits 99.8% test compatibility on Linux x64 glibc
Even last year at this time people wouldn't believe it.
← PreviousPage 2 of 8Next →