HNHacker News
TopNewBestAskShowJobs

kamranjon

2,659 karma · joined February 8, 2017

kamranjon.com
submissionscomments
kamranjon··on Anthropic's IPO prospectus shows AI vision, surging costs
does anyone understand what this part actually means?

"The near-$42 billion net loss included a roughly $34 billion accounting charge that reflected an increase in the estimated value of financing that could eventually turn into Anthropic shares, rather than money the company spent running its business."

kamranjon··on Drawgent: Coding agent on a live Excalidraw canvas
Is the excalidraw api constantly changing?
kamranjon··on One Month Without AI
You would probably be surprised by how many jobs require that you use AI - I even had an interview where it was strongly encouraged to use it during the technical phase. I guess what I’m getting at is, for many people this isn’t really a choice.
kamranjon··on Starlink ground station in Poland hit by fire in suspected arson attack
Is that true though? 5g is significantly faster than starlink
kamranjon··on Starlink ground station in Poland hit by fire in suspected arson attack
Huh… I always thought direct satellite connection was one of the things that distinguished starlink from cellular tech like 4g or 5g, and that low earth orbit was what made it possible… what actually makes starlink different?

Edit: oh I think? I understand now, these basically relay the internet itself up to the satellites - so your individual starlink dish is still talking directly to a satellite but the satellite resolves your request for a website through these ground stations. I think?

kamranjon··on Radicle: Disclosure of Vulnerability in the Network Protocol
it's not the same name
kamranjon··on Pentagon says overreliance on AI contributed to missile strike on Iran school
Many people are commenting that AI is just a scapegoat here and really this is a story about military incompetence. What I think is probably more interesting is how AI enables incompetent people to do more damage than they would otherwise. Two things can be true at once, the military can be incompetent, and AI can be used to make targeting decisions that are wrong - we can understand both the new world that we live in, and the absolution of responsibility it offers without ignoring these nuances.
kamranjon··on Apple iPhone 18 Pro Camera test
Is that true though? It seems like the iphones are doing some type of skin smoothing algorithm but the Huawei is the only one that actually seems to produce actual skin texture and blemishes. My initial read was that the Huawei was the least processed of the bunch.
kamranjon··on I built non-autoregressive decision models with RL a year ago
It is really interesting to see this claim, because i thought the current theory was that typesafe actually repackaged the work from GLiNER[1] - which does seem to be a closer match, and their original paper[2] predates yours by several years. Curious if you had heard of it before? It is also open source[3] and I think also has some good usage.

[1] https://arxiv.org/abs/2507.18546

[2] https://arxiv.org/abs/2311.08526

[3] https://github.com/fastino-ai/GLiNER2

kamranjon··on Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint
"Ternary Bonsai 2 27B reaches up to 143 tokens/second on NVIDIA GeForce RTX 5090 and 46.8 tokens/second on M5 Max. On an RTX 4090, Ternary Bonsai 2 27B consumes just 0.714 mWh/token, making it 40% more energy-efficient than an 8B model running in full-precision."
kamranjon··on Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint
Love this for the folks with 16gb graphics cards - 3.8 27b has been incredible but not quite runnable on anything less than 32gb - will try loading this up on my 16gb intel b50 and see how it goes - not sure these quants can be accelerated by the XPU cores yet but maybe in time!
kamranjon··on How GLM built its own inference infrastructure
Someone tell this man about vLLM!
kamranjon··on DeepSeek launching v4.1 flash cheaper and more capable than v4 pro
Is nobody using structured outputs? They use constrained decoding at the generation stage to ensure the probability of tokens that would break the format are set to 0. I kinda figured everyone was doing this at this point.
kamranjon··on DeepSeek launching v4.1 flash cheaper and more capable than v4 pro
Hey there! I do the same but I use dwarfstar at a 2-bit quant: https://github.com/antirez/ds4

I'm curious if you've tried dwarfstar and decided to move to llama.cpp and 3 bit quants or what made you go that route instead? I've been using ds4 for months now and it's already got support for the new vision model, haven't tried it yet, still on 0731 but it's been very solid for me.

kamranjon··on Georgi Gerganov on llama.cpp/ggml future after Nvidia acquisition of HuggingFace
I follow llama.cpp pretty closely as I use either llama.cpp itself or projects that depend on it all the time, and one thing that I don't think gets talked about is the sheer scale of community involvement. It seems like a logistical nightmare, but somehow thousands of different contributors are opening hundreds of PR's every week and getting them merged in to support various hardware or implement a new pattern or algorithm from a recent research paper. It's really quite awe inspiring for me to see, and think it is in no small part because of the leadership of ggerganov - so I'm happy to see that he is sticking around and plans to keep building this incredibly useful tool that has grown into a huge community at this point.
kamranjon··on K2 Horizon: A connected fleet of six open models
Did you post in the wrong thread?
kamranjon··on Discovery of a new OpenAI agent message board
I think this account should be banned.
kamranjon··on K2 Horizon: Frontier Performance, Radically Open
it's funny that the tagline is Radically Open, but you're immediately hit with http login - maybe this was the wrong link?
kamranjon··on Gemini 3.8 Flash and 3.8 Flash Cyber
They said Opus 5 medium - which does have an intelligence score of 59 (you have to select it manually from the dropdown to see it)
kamranjon··on Gemini 3.8 Flash and 3.8 Flash Cyber
Yea I am testing through OpenRouter - have you noticed 3.7 flash being significantly faster?
kamranjon··on Gemini 3.8 Flash and 3.8 Flash Cyber
They've interestingly left out any mention of speed.

I have been testing 3.7 flash against 3.5 flash and it seems to lose every time in overall latency. Every benchmark I've seen seems to suggest the opposite[1] - that 3.7 flash is significantly (at times 2x) faster than 3.5 flash - but I have never been able to prove this out in real world use cases.

Has anyone found their latency numbers to actually be accurate? Is this why they've toned it down in this release? For context, I'm testing larger generation payloads that take 8-10 seconds in 3.5 flash and 15-25 seconds in 3.7 flash. Lowest reasoning settings in both cases.

1: https://artificialanalysis.ai/?speed=intelligence-vs-speed&m...

kamranjon··on Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
I am running 3.8 27b at q6 quant with 160k context on a 32gb video card (arc b70 pro) - I quantized the kv cache at q8 - that is the only trick really - works great.
kamranjon··on Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses
Hmm, makes the 2 bit quants actually seem pretty reasonable…
kamranjon··on Qwen3.8-Flash-Next Technical Report [pdf]
Trained on roughly 1/9th of the training FLOPS - it’s the pretty incredible that they are making these advances and at the same time sharing their learnings in these papers - I wish we saw more of this from US labs.
kamranjon··on Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)
I actually do this with my MBP - it's a LLM server when I'm working - and then when I'm not it's just a really great machine for video editing and other media work.
kamranjon··on Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)
Where did they say that? My understanding of this 3.8-Flash-Next release is that it's a MOE (as per the title of the posting here, 125B a6b)
kamranjon··on Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)
Have you tried FreeToken yourself? I was hoping to find some benchmarks on their github but took a quick pass at their research paper and it seems they're showing ~2x performance on qwen 3.6 35b when compared to llama.cpp - but llama.cpp is so sprawling and has so many options I find that a difficult comparison.
kamranjon··on Apple introduces M6 and M5 Ultra
I think you misread, it’s 170gb/s for base M6 model and 1.2tb/s for M5 ultra.
kamranjon··on Apple Introduces New Mac Studio with M5 Max and M5 Ultra
You would want to get the M5 pro version with 307gb/s if you were interested in running local LLMs.
kamranjon··on Apple Introduces New Mac Studio with M5 Max and M5 Ultra
1tb would likely be ~$20k - given the current >$10k price tag of 256gb. Would you still be considering it at that price?
Page 1 of 22Next →