HNHacker News
TopNewBestAskShowJobs

kamranjon

2,676 karma · joined February 8, 2017

kamranjon.com
submissionscomments
kamranjon··on Show HN: Maple-Preview – Ternary 20B MoE running at 120 tok/s on a iPhone
“Current approaches to low precision primarily focus on converting models trained in full precision to lower bitwidths. We view this as fundamentally the wrong approach…”

Very excited to see how it performs, I’ve been a bit skeptical of the efficacy of converting existing models - really cool to see one trained from scratch in the ternary format.

kamranjon··on An Honest Review of AI Programming
“While technically true the hallucination rates on modern models is low…”

Isn’t this entirely context dependent? Where did you get the information that modern models have low hallucination rates? I’d love to see the benchmark if there is one, it seems like it would be useful to track.

kamranjon··on Show HN: Fine-tune an 8B model on a 4 GB laptop GPU
This seems really interesting - I was curious about this line from the website.

“The whole post-training stack in one CLI. Soup doctors your data pre-flight, picks the method, writes the config, derives evals from your own data, gates every save, and self-corrects reward hacking mid-run instead of just halting.”

How does soup auto tune the hyper parameters and make some of these more complex training decisions?

kamranjon··on Explorative modeling: Train on the best of K guesses
Do you have an example? Would love to read one.
kamranjon··on Google News is just Forrest Gump's shrimp boat now
I love this analogy and think it possibly also applies to search which just doesn’t work anymore and is full of AI generated garbage.

What the hell happened to stack overflow? I don’t think I’ve gotten a google result for stack overflow in nearly 2 years. How have they just forgot how to make competent search?

I didn’t think it’d happen so quickly but I honestly get better results from DuckDuckGo at this point.

kamranjon··on Explorative modeling: Train on the best of K guesses
In what way is it similar?
kamranjon··on Explorative modeling: Train on the best of K guesses
This is amazing and I think will probably end up being a pretty important development.

I was just reading this great breakdown of how diffusion Gemma works: https://newsletter.maartengrootendorst.com/p/a-visual-guide-...

In reference to the difficulties with applying this to autoregressive LLMs - I wonder if these type of hybrids might be a good candidate for this approach.

kamranjon··on DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
I actually run it as a server - so most of the time I don't have to listen to it right next to me - it's just sitting in another room in my house - but I often am traveling with it and will have it sitting right next to my coding laptop and yea the fan runs non-stop - it's not obnoxious so i can pretty easily tune it out - also airpods/noise canceling headphones help!
kamranjon··on DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
The really interesting thing about this is how big of a jump was achieved with just extra fine-tuning here. No structural changes to the model, just more data, compute and time. It makes me pretty excited for the future of small models - DS v4 flash is a relatively small model when compared to the class it's competing with, so likely similar gains can be made applying quality data/training pipeline to other smaller models.
kamranjon··on DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
Generally get 20-25tps - prefill is pretty good around 400-450tps. I have been using compaction at around 100k tokens but mostly just cause it was the default in pi coding agent - might see if i can expand it a bit.
kamranjon··on DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
Can't wait for the DwarfStar quants - I have been using DeepSeek v4 flash (preview) as my main coding agent for months now (running on my 128gb mbp) - it seems this model outperforms GLM 5.2 on nearly every metric. Thanks for sharing the news, I was refreshing huggingface but gave up thinking it likely would take some more time.
kamranjon··on The AI trade now runs on borrowed money, and the lenders are repricing it
I know many companies are spending quite a bit of money, I don’t know if it bears out that the increased spend has resulted in increased profits, even if there has been some increase in productivity. I think this is the tough situation many orgs are facing right now, drastic adoption without material economic gains.
kamranjon··on Neutrino-1 8B
I haven't spun up thinkingcap yet but I'm aware of it and am intending to try it out soon. How did you find it?
kamranjon··on Kimi K3 Architecture Overview and Notes
Would you happen to have a link to that interview? Sounds like an interesting read.
kamranjon··on Neutrino-1 8B
Sorry I should have clarified - I meant that a ternary 27b model would outperform a non-quantized or 8 bit quantized 9 or 12b model - which it is generally close to (or much smaller than) in size. So yeah the comparison I was trying to make was between models of equivalent size or models that could run on similarly sized hardware.
kamranjon··on Neutrino-1 8B
I’ve actually been really impressed with the 27b model they recently released - amazing performance approaching 40 tok/s on m4 max and I didn’t run into any quality issues in the small set of tasks I tried. Haven’t gone full coding with it yet but suspect it’s better than say a 9b or 12b model.
kamranjon··on Neutrino-1 8B
Unfortunately Fermion Research appears to entirely AI generate all of their content here, even for the research section: https://www.fermionresearch.com/research/neutrino-8b/

"Neutrino-1 8B was trained natively in its shipping format. There is no full-precision product model that was rounded afterward: the ternary representation is the medium the weights learned in, and the training methods that hold this quality at this depth are the lab’s unpublished work. The findings below are the part that travels."

This statement seems misleading at best.

Both the model page and the release page are basically unintelligible - I don't have a ton of faith in the work here, at least PrismML write coherent releases for their models.

Edit: Another beautiful piece of prose here, I almost wonder if they used the 8b model to generate the content for this release...

"Across the 6.95B coded weights, 62.63% sit at zero and the remainder splits 18.68% plus to 18.69% minus: sign-balanced to a hundredth of a point with no constraint asking for it."

kamranjon··on Neutrino-1 8B
There's a really interesting trend of labs using "proprietary" methods to convert existing models to a compressed ternary format.

PrismML actually targeted the same Qwen 8b model and got it down to 1.75gb here: https://prismml.com/news/ternary-bonsai

I wonder how proprietary it all is though, since the BitNet b1.58 paper has been out for a couple years now: https://arxiv.org/abs/2402.17764

From the wikipedia on 1.58 bit llms: "BitNet derives its performance from being trained natively in 1.58 bit instead of being quantized from a full-precision model after training. Still, training is an expensive process, and it would be desirable to be able to somehow convert an existing model to 1.58 bits. In 2024, HuggingFace reported a way to gradually ramp up the 1.58-bit quantization in fine-tuning an existing model down to 1.58 bits."

The section from huggingface is here: https://huggingface.co/blog/1_58_llm_extreme_quantization#fi...

I just wonder how many of these labs are basically following the huggingface recipe here and possibly tweaking it and releasing models without huge training costs.

kamranjon··on Kimi-K3 on HuggingFace
There is a really interesting startup in Prague that is doing just that. They fine-tuned Qwen 3.6 27b to have 46% fewer reasoning tokens while maintaining most of the performance characteristics. I'm interested to see if they continue down this path of optimizing reasoning for other models.

https://bottlecapai.com/post/thinkingcap-qwen3-6-27b/

kamranjon··on Sand battery: Finland's answer to a renewable energy headache
This does seem like a prett niche solution because the electricity is converted into heat and then just used directly as heating - which removes a lot of the utility of electricity.

It is a great solution for places where heating is the primary use of electricity and I hope it finds broader applications.

kamranjon··on Running a 28.9M parameter LLM on an $8 microcontroller
So while SSD streaming is interesting I'm not sure it's exactly the same thing as the per-layer embedding that is being utilized in tandem with streaming here. To utilize per-layer embedding, it would have had to be trained that way, which GLM 5.2 was not.
kamranjon··on Running a 28.9M parameter LLM on an $8 microcontroller
Pretty incredible performance for the footprint - really interested to see what could be done on slightly more powerful SBCs like some that have been mentioned in this thread.
kamranjon··on ARC-AGI Leaderboard
Interesting to place that level of trust in the providers, but I guess that’s the best you can do with closed models. Makes me wonder if Opus 5 could have been trained on data they promised they weren’t training on? One of the interesting things about LLMs is how opaque they are from the outside, even with open weights, it’s very difficult to know if a model incorporated benchmark data in their training.
kamranjon··on Show HN: Echo – Fable-level results at 1/3 the cost using open-weight models
You keep saying open weights - but you aren't sharing any information on which models you are using. What benefit does using open weights models provide to the end user if there is zero transparency?
kamranjon··on Writing by hand is good for your brain
Wow I never knew this - what a useful bit of information. Makes me want to try a fountain pen, I do like writing in cursive but I often get hand cramps doing it for extended periods.
kamranjon··on I regret migrating to Codeberg
I’ve recently moved to sourcehut and really enjoy it so far! While I think the creator has strong personal opinions about LLMs - I don’t think they have any intentions of banning LLM generated projects. Anyhow just wanted to share because I think it’s a really nice product that isn’t overbloated with JS and just gets the thing done.
kamranjon··on OpenAI’s accidental attack against Hugging Face is science fiction that happened
Is it a sandbox if the machines hosting your mock package servers have access to the internet? This just seems like a huge failure on the part of OpenAI to secure their environment during testing.
kamranjon··on Show HN: Echo – Fable-level results at 1/3 the cost using open-weight models
No benchmarks, no info on which models are used, ai generated video, just a signup page with nothing else.

Anyhow, this kinda reminds me of that quote about architecture: "We replaced our monolith with micro services so that every outage could be more like a murder mystery."

kamranjon··on “We have information that Moonshot distilled Fable for the development of K3”
So here is an important question I think.

If LLM outputs aren't copywriteable and you create your own synthetic training set using Fable and share it publicly on huggingface, and someone else uses that training set to fine-tune a model, would this be considered illegal?

I ask because this happens all the time, synthetic datasets have basically become a key aspect of training a model at this point. I even generated a synthetic set from DeepSeek v4 to aid in fine-tuning a classifier just a few weeks ago.

So I just wonder on what grounds any of this makes sense, I wouldn't be surprised if some of these American labs were using open models on their own self hosted infrastructure to generate training data, but by nature of them being open nobody has to know.

I'll make a prediction: I don't think we will ever see any of the evidence of this "distillation" before they end up implementing some type of ban.

kamranjon··on Laguna S 2.1
"It went from the start of training to launch in under nine weeks..."

This is pretty impressive.

← PreviousPage 3 of 22Next →