HNHacker News
TopNewBestAskShowJobs

benxh

57 karma · joined May 2, 2023

North Macedonia based.

@nipple_nip on twitter

submissionscomments
benxh··on Big AI sets out its terms for regulatory capture
At the current pace of iteration from Deepseek I think this is a moot point, they'll keep building their own RL environments, and just keep RL on top of whatever flavor of model they can train/host on their Huawei SuperPods
benxh··on Saab has unveiled its A3 collaborative combat aircraft concept
NATO v Yugoslavia 1999 is not like the others
benxh··on Qwen 3.8 27B
Depends on what software/hardware you'll run it. GGUFs from Unsloth can run on pretty much every single potato; full weights need beefy gpus
benxh··on GLM 5.2 Performance Benchmarks
benchmark where gemini flash is better than fable btw.
benxh··on US to suspend immigrant visa processing for 75 nations, State Department says
Crazy calling sovereign states "US Puppets".
benxh··on Chinese AI models have lagged the US frontier by 7 months on average since 2023
Minimax has been great for super high speed web/js/ts related work. It compares in my experience to Claude Sonnet, and at times gets stuff similar to Opus. Design wise it produces some of the most beautiful AI generated page I've seen.

GLM-4.7 like a mix of Sonnet 4.5 and GPT-5 (the first version not the later ones). It has deep deep knowledge, but it's often just not as good in execution.

They're very cheap to try out, so you should see how your mileage varies.

Ofcourse for the hardest possible tasks that GPT 5.2 only approaches, they're not up to scratch. And for the hard-ish tasks in C++ for example that Opus 4.5 tackles Minimax feels closer, but just doesn't "grok" the problem space good enough.

benxh··on Chinese AI models have lagged the US frontier by 7 months on average since 2023
It is arguable that the new Minimax M2.1 and GLM4.7 are drastically above Sonnet 3.7 in capabilities.
benxh··on First Self-Propagating Worm Using Invisible Code Hits OpenVSX and VS Code
cline is used by a lot of devs
benxh··on OpenAI o3-pro
The longer "it" reasons, the more attention sinks are used to come to a "better" final output.
benxh··on DoppelBot: Replace Your CEO with an LLM
To prove you right, you can read up on the incredible giga-brained countrywide experiments by Kardelj in Socialist Yugoslavia [0]. The result being a country where no-one wanted to work, and everyone had a great standard of living (while the IMF didn't call in its loans). And then the entire country collapsed all at once under the accumulated mismanagement.

[0] - https://en.wikipedia.org/wiki/Workers%27_self-management#Yug...

benxh··on Llama.cpp supports Vulkan. why doesn't Ollama?
My biggest gripe with Ollama is the badly named models, e.g. under deepseek-r1, it defaults to the distill models.
benxh··on In the land of LLMs, can we do better mock data generation?
I'm pretty sure that Neosync[0] does this to a pretty good degree, it is open source and YC funded too.

[0] https://www.neosync.dev/

benxh··on Anna’s Archive approaching 1 petabyte
I am assuming this will be solved this year.
benxh··on Phind-70B: Closing the code quality gap with GPT-4 Turbo while running 4x faster
If GPT4 is 220B/8 experts, that would be in-line with 3.5 Turbo being a 20B model, and GPT4 being a 55B activation out of a total 220B parameters.

It is ultimately all speculation, until Deepseek releases their own 145B MoE model, and then we can compare the activations/results

benxh··on A million ways to die on the web
I personally was affected by this fire, although I've always kept 3 month backups of production data, encrypted, on-site, just in case of emergencies like this. Haven't touched their services for anything production related ever since
benxh··on AlphaCode2 powered by Gemini performs better than 85% of programmers
It's buried deep in the Gemini report, but goddamn are these incredible stats.
benxh··on OpenAI's board has fired Sam Altman
The Albanian takeover of AI continues. It's incredibly exciting!
benxh··on Phind Model beats GPT-4 at coding, with GPT-3.5 speed and 16k context
I can't wait to see this open sourced, there's a lot of sampling strategies that help coding.

And I also can't wait to see how much Phind will improve further if the Glaive dataset is added onto it.

Edit: Contrastive search, dynamic temperatures.

benxh··on China is becoming a data black hole, says short seller Aandahl
It was known by Polynesians for at least 1000 years before Columbus. See sweet potatoes.
benxh··on Mistral 7B
It's missing a lot of crucial details. Nothing on the dataset used, nothing on the data mix, nothing on their data cleaning procedures, nothing on the tokens trained.
benxh··on Tiny Language Models Come of Age
Not all of them per se, take a look at something like Mistral. It's a 7B model displaying incredible performance. IMO, we still haven't even scratched the surface of what is possible with small LLMs. Especially not with pre-filtered/classified pre-training data. (Interesting LLMs based on their data approach and relatively small size: Qwen, InternLM, Mistral, Phi)
benxh··on OpenAI's justification for why training data is fair use, not infringement [pdf]
Added, and reached out on Twitter.
benxh··on OpenAI's justification for why training data is fair use, not infringement [pdf]
I would like to get in touch with you related to books4. Do you happen to have discord? or would twitter be ok?

There's currently multiple attempts at creating what you describe as books4.

benxh··on Llama 2 on togetherAI is as bad of a privacy nightmare as OpenAI
I've had some success using vast.ai[0] with the Oobabooga LLM WebUI (LLaMA2) instances. One click to start up, minimal editing in the interface settings to enable OpenAI compatible interface.

[0] https://cloud.vast.ai/

benxh··on A GPT-4 capability forecasting challenge
Yeah wildly inaccurate.
benxh··on Why Host in Kosovo?
This reads like a Serbian owned business from the North of Kosovo. But anybody following the local politics would know that: 1) Corruption as an issue is disappearing in Kosovo, especially compared to Albania, Macedonia, Montenegro and Serbia where it flourishes. 2) Corruption of police is practically impossible in Kosovo. 3) The IX is literally under USAID/EU control, good luck not getting blocked. 4) The North of Kosovo until this year was lawless. What they wrote about the east makes 0 sense.

All in all, this is either a honeypot, or somebody trying to discredit Kosovo.

benxh··on Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?
Yes, but the acquisition of that data itself is illegal in almost all jurisdictions, since libgen is treated as a piracy website. Now if there were a pipeline to access books from Amazon or the Google Books project for training it would be a different story.

Still, for certain languages, only libgen and public piracy websites contain any scientific or fiction material in digital formats. E.g. my native language doesn't have easily accessible e-books at all, unless you go through illegal means.

I hope somebody undertakes the steps necessary to train on the entirety of libgen. The amount of high quality tokens in libgen should be substantial.

benxh··on Meta announces its Quest 3 VR headset
How does one go about becoming a distributor of Quest products in countries which arent served at all by Meta? Considering I can leverage existing infrastructure and network connections to bring it to market?
benxh··on Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?
To be honest, I've been asking myself the same thing, technically the amount of "good quality" data in libgen is huge, way larger than the books3 dataset. However it would probably run afoul of copyright. Then again, a huge amount of data that LLMs go through is copyrighted.
benxh··on Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?
So a model fine-tuned on libgen?
Page 1 of 2Next →