HNHacker News
TopNewBestAskShowJobs

kamranjon

2,659 karma · joined February 8, 2017

kamranjon.com
submissionscomments
kamranjon··on New Mac Studio with M5 Max and M5 Ultra
You can get the m5 pro in the Mac mini with 307GB/s at 64gb of memory it’s $2899
kamranjon··on Why Sal Khan't: On Learning by Making but Teaching by Telling
Yeah the article feels as though its describing Khan Academy from 10 or more years ago - it is quite odd.

"What he has never had is pedagogical knowledge: an understanding of how people learn, what motivates them, what makes the difference between someone who pushes through difficulty and someone who types “IDK.”"

It is almost as though the author thinks that Khan Academy is an organization of 1, and that they don't employ learning scientists and have a content team and don't publish research papers in the learning science field.

kamranjon··on The coolest anti-surveillance tools at Defcon [video]
Is anyone familiar with the laws surrounding police basically operating their own pseudo cell towers?

I would assume this would be highly illegal for individuals, what sort of hoops did law enforcement need to jump through to get this type of approval? Did FCC need to rubber stamp this?

kamranjon··on How we made a text-to-speech model respond in sub-50 ms
Hi there! I actually thought your Dia models were amazing and very natural sounding, I haven’t tried qwen 3 tts yet - has your focus shifted away from building your Dia models and shifted more towards hosting and infrastructure?
kamranjon··on DiffusionGemma Technical Report
Just wanted to share this, I found it was a really nice resource to understand how diffusion Gemma worked: https://newsletter.maartengrootendorst.com/p/a-visual-guide-...

The really interesting thing to me was that they didn’t need to train this model from scratch they just used their existing MOE checkpoint:

“To convert a decoder-only model (Gemma 4 26B A4B) into a denoiser, we can make use of something it is not directly using when generating tokens, namely the logits of all tokens!”

What makes me hopeful about this release is that possibly this same conversion can be applied to other open models and we might see a bunch of diffusion versions of existing local models. It’s exciting stuff!

kamranjon··on Seeing beyond BMI: Estimating cardiometabolic risk with smartphone imagery
I wonder if this could be used by insurance companies to determine premiums?
kamranjon··on Unsloth Dynamic 3.0 GGUFs
Are you using the recommended settings for temperature and such? https://unsloth.ai/docs/models/qwen3.8#recommended-settings

Often times I run into issues like this it’s because I am using settings for a different model or just forget to set them up.

kamranjon··on Error by AI scribe during medical appointment leaves patient devastated
I think if a real human doctor completely fabricated a drug usage history for one of their patients, that would also be news.
kamranjon··on Unsloth Dynamic 3.0 GGUFs
What size context are you able to squeeze in with less than 2gb of headroom? I have had some luck using a quantized kv cache but i fear that also decreases overall quality.
kamranjon··on Qwen 3.8 27B is excellent, but it defaults to overthinking things
A no-thinking pelican! I hope to see more, it's surprisingly good for just 2 minutes.
kamranjon··on Qwen 3.8 27B
I actually disagree that this doesn’t mean anything. I understand the contention that it’s not measuring the quality of the model in general, but I think it is measuring something useful.

A good example of this is planning hardware projects - a larger 200b plus model like DeepSeek V4 flash will recommend parts like motors, real time clocks, voltage regulators etc and it will do so providing exact model names and specifications.

I wouldn’t expect a smaller model to encode all of this information, but it is helpful to understand where that cutoff is because it changes what the model might be useful for. It is a very crude way of measuring because it comes down to the balance of training data at sizes this small - but I do think it conveys something that is helpful in real world tasks.

kamranjon··on Qwen 3.8 27B
Have you thought about running a second tier of the Pelican benchmark where you see which model makes best pelican on lowest or no reasoning settings? I think that'd be pretty interesting and might help highlight which models have a baseline capability - even within those I'd imagine the token usage would vary wildly and might give some indications on verbosity as well.
kamranjon··on Qwen 3.8 27B
You can buy two b60s for $1300 right now (650 each) if you want a total of 48gb. Intel recently raised the price on all of their gpu's except the b60 series, so they are currently the best deal per gb I think.
kamranjon··on Qwen 3.8 27B
Since Qwen 3.6 27b outperforms Gemma 4 26b in most benchmarks I'm not sure the value - also Gemma 26b is a MOE model whereas this is a dense model, so not typically direct competitors at their sizes - Gemma 4 31b comparison would be interesting though.
kamranjon··on DeepSeek Harness
It's useful logs which i think is an important distinction.
kamranjon··on AI agents lie, cheat and steal. That is putting off users
I don't really think the distinction here is relevant. If the end result is the equivalent of lying, cheating or stealing - then the problem still exists and it needs to be solved.
kamranjon··on DeepSeek Harness developer preview
New coding harness that seems to have some novel concepts and one of the pretty cool things on their landing page for it here: https://deepseek.com/harness/en/ is the Every Run is Traceable view:

"Everything the model sees is recorded in an append-only session log: system prompts, reasoning, tool calls and results, subagent scheduling, and every context injection. In the Trajectory view, you can inspect these records by source. Resume, fork, search, and replay all operate on the same event stream."

Seems pretty helpful - have sort of wanted something similar (I use Pi).

They also released this research paper that backs their whole plugin composability system that seems pretty cool: https://github.com/cordiverse/paper

kamranjon··on Qwen 3.8-27B goes openweight in 2 days
this seems to be a countdown for the 2.4t model - which is gigantic and is not as exciting for those running models locally
kamranjon··on llama.cpp
I actually have been exploring this very thing!

I think the best option right now, since Apple has raised prices and Mac minis are basically impossible to get your hands on, is to build your own micro-itx machine. I actually built a mini-itx machine, but it does restrict your options a bit.

The Arc series Intel GPUs are what I think make this possible. I built a machine with an Arc b50 - it runs Gemma 26b a4b qat at around 30tok/s with their MTP head and prompt processing sits at around 500 tok/s. The really beautiful thing about this setup is the entire energy envelope of this machine sits at 120w at full load - when idle, it's at 40w and i've done some work in ubuntu to basically intelligently hibernate, which drops it to 0 watts when not in use. You can use a raspberry pi and Wake on Lan to wake the machine up for a overall draw of around 5 watts when not in use.

All in all this machine cost me 1.4k to build - but if you used micro-itx instead of mini-itx parts you could do it for under 1k - it has just 16gb of ddr5 but you don't really need more if you use models that can fit in vram.

I think it's pretty incredible that you can run an actually useful coding agent on a machine with a power envelope that is less than an incandescent light bulb. If you go up to micro-itx you can do even large cards like an intel b60 with 24gb or a b70 with 32gb and run even more powerful models. For all of these intel GPU's you'll want to compile the latest llama.cpp version with SYCL support - they are getting speedups every day, so worth staying on the edge.

kamranjon··on Nvidia Nemotron 3.5 Lightning
You might look at this and and be a bit disappointed by the performance against qwen and gemma models - but this is an entirely open source training pipeline, this is quite impressive and I don't think another model this performant exists with fully open source data and recipes alongside the weights.
kamranjon··on Mark Zuckerberg attacks 'closed' AI rivals as Meta returns to open models
I honestly think of the T5 model family to sort of be the real beginning of this open model craze - I know BERT was already popular for classification etc, but T5 was the first sort of generally useful model, was exceptionally simple to fine-tune, and is still in use today (t5 base is still averaging over a million downloads a month on huggingface), has tons of variants and sort of kickstarted this whole community. US labs get a lot of flack but Google has been super supportive and open in a lot of ways that has pushed this whole endeavor forward, even if I feel like they've sort of declined in transparency in recent years with their open models.
kamranjon··on Why Wall Street is ignoring big tech's debt [video]
1 to 200 is a pretty big spread to between losing 18k dollars and making 3.6 milllion - do you have any actual numbers on the value produced from this 18k investment?
kamranjon··on DeepSeek V4 Flash 0731
This is pretty interesting, I've never heard of this approach before - do you know if there is a research paper that covers how this was achieved?
kamranjon··on The Claudyssey: A line-for-line translation of Homer's Odyssey by Claude Fable 5
The sort of sad but interesting thing here is that while this claims to be a more direct translation, Claude models were certainly trained on all of the pre-existing translations and this is likely going to directly influence the output in the ways you've outlined here. I'm not sure any modern LLM could honestly be claimed to produce a translation that doesn't draw on these pre-existing sources.
kamranjon··on Another Corner of the Internet Has Been Ruined
This wasn't written with AI... obviously... I feel like there might need to be a new definition for whatever this paranoia is called because it's getting a bit out of hand.
kamranjon··on Godox Transparent Viewfinder Camera C100
Because street photography is very spontaneous it’s pretty common practice to set an aperture of 8 and just snap away - it’s a helpful trick for rangefinder cameras that often take some time to pull focus.
kamranjon··on Show HN: Maple-Preview – Ternary 20B MoE running at 120 tok/s on a iPhone
“Current approaches to low precision primarily focus on converting models trained in full precision to lower bitwidths. We view this as fundamentally the wrong approach…”

Very excited to see how it performs, I’ve been a bit skeptical of the efficacy of converting existing models - really cool to see one trained from scratch in the ternary format.

kamranjon··on An Honest Review of AI Programming
“While technically true the hallucination rates on modern models is low…”

Isn’t this entirely context dependent? Where did you get the information that modern models have low hallucination rates? I’d love to see the benchmark if there is one, it seems like it would be useful to track.

kamranjon··on Show HN: Fine-tune an 8B model on a 4 GB laptop GPU
This seems really interesting - I was curious about this line from the website.

“The whole post-training stack in one CLI. Soup doctors your data pre-flight, picks the method, writes the config, derives evals from your own data, gates every save, and self-corrects reward hacking mid-run instead of just halting.”

How does soup auto tune the hyper parameters and make some of these more complex training decisions?

kamranjon··on Explorative modeling: Train on the best of K guesses
Do you have an example? Would love to read one.
← PreviousPage 2 of 22Next →