HNHacker News
TopNewBestAskShowJobs

kiratp

964 karma · joined September 19, 2017

Co-Founder @ osmos.io
submissionscomments
kiratp··on Anthropic's best AI model struggles to attract users as cheaper tools thrive
The actual issue, is suspect, is that Anthropic won’t provide ZDR for Fable. Makes it a non started for a large percentage of businesses.
kiratp··on Kimi-K3 on HuggingFace
AISI is capped at 100M tokens and K3 is less token efficient than Anthropic/OpenAI models. There is an argument to be made, looking at AISI results, that with uncapped tokens it would be just slightly behind the closed weight players.
kiratp··on An update on recent Claude Code quality reports
Agents making forward progress hours apart is an expected pattern and inference engines are being adapted to serve that purpose well.

It’s hard to do it without killing performance and requires engineering in the DC to have fast access to SSDs etc.

Disclosure: work on ai@msft. Opinions my own.

kiratp··on An update on recent Claude Code quality reports
OpenAI does this for all API calls

> Our systems will smartly ignore any reasoning items that aren’t relevant to your functions, and only retain those in context that are relevant. You can pass reasoning items from previous responses either using the previous_response_id parameter, or by manually passing in all the output items from a past response into the input of a new one.

https://developers.openai.com/api/docs/guides/reasoning

Disclosure - work on AI@msft

kiratp··on An update on recent Claude Code quality reports
By caching they mean “cached in GPU memory”. That’s a very very scarce resource.

Caching to RAM and disk is a thing but it’s hard to keep performance up with that and it’s early days of that tech being deployed anywhere.

Disclosure: work on AI at Microsoft. Above is just common industry info (see work happening in vLLM for example)

kiratp··on Anthropic takes $5B from Amazon and pledges $100B in cloud spending in return
At the full current retail API price.

Business buyers are paying API prices, not subscription

Disclosure: Work at Microsoft on AI

kiratp··on 1M context is now generally available for Opus 4.6 and Sonnet 4.6
GitHub Copilot CLI lets you use all these models (unless your employer disables them.

https://github.com/features/copilot/cli

Disclosure: work at Msft

kiratp··on NASA announces overhaul of Artemis program amid safety concerns, delays
Same contractors (Beoing) who built Starliner...

Explaining Why NASA's Starliner Report Is So Bad > https://www.youtube.com/watch?v=L96asfTvJ_A

kiratp··on Sam Altman’s DRAM Deal
This is missing a key part of the picture - Nvidia just announced that partners will need to source RAM themselves.

OpenAI is basically ensuring that they can actually get the chips they need for the DCs they are building.

I can’t guess as to what move came first (Nvidia policy change or these DRAM deals) but I would bet this is a large if not larger factor here than “bloc my competitors.

kiratp··on Fizz Buzz without conditionals or booleans
A loop either never halts or has a conditional. I guess a compiler could elide a “while True:” to a branch-less jump instruction.

One hack would be to use recursion and let stack exhaustion stop you.

kiratp··on Fizz Buzz without conditionals or booleans
A for loop has a conditional in it.

Unless by conditionals we mean “no if/else” and not “no branch instructions”.

kiratp··on Fizz Buzz without conditionals or booleans
A for loop has an implicit conditional in its stop condition check.
kiratp··on UnitedHealth pays its own physician groups 17% more than outside ones
This only applies to large employers. Smaller ones are just presentef a limited list of plans to pick from, and the plans change every year. Most of the time, as a startup, you can’t buy a Mag7 equivalent health plan for any amount of money off the marketplace
kiratp··on FSF announces Librephone project
Should the app builder’s ability to “trust” that the hardware will protect them from the user supersede the user’s ability to be able to trust that the hardware will protect them from the app?

In other words, should the device be responsible to enforcing DRM (and more) against its owner?

kiratp··on The Tiny Teams Playbook
The kind of people in these small teams are not ones to think "work is just work".
kiratp··on Meta’s live demo fails; “AI” recording plays before the actor takes the steps
You can put the AI on rails by just prompting by it. The latest models are very steerable.

System prompt: “stick to steps 1-n. Step 1 is…”

I can say confidently because our company does this. And we have F500 customers in production.

kiratp··on Meta’s live demo fails; “AI” recording plays before the actor takes the steps
I see no evidence of that. It seems like they tried to put the AI “on rails” with predefined steps and things went wrong.
kiratp··on Meta’s live demo fails; “AI” recording plays before the actor takes the steps
So much negativity.

I’m just excited that our industry is lead by optimists and our culture enables our corporations to invest huge sums into taking us forward technologically.

Meta could have just done a stock buyback but instead they made a computer that can talk, see, solve problems and paint virtual things into the real world in front of your eyes!

I commend them on attempting a live demo.

kiratp··on A postmortem of three recent issues
This is due to RoPE scaling.

> All the notable open-source frameworks implement static YaRN, which means the scaling factor remains constant regardless of input length, potentially impacting performance on shorter texts. We advise adding the rope_scaling configuration only when processing long contexts is required. It is also recommended to modify the factor as needed. For example, if the typical context length for your application is 524,288 tokens, it would be better to set factor as 2.0.

https://huggingface.co/Qwen/Qwen3-Next-80B-A3B-Thinking

kiratp··on Claude now has access to a server-side container environment
Hardware can be the same but scheduling is a whole different beast.

Also, if you pull too manny resources from training your next model to make inference revenue today, you’ll fall behind in the larger race.

kiratp··on Claude now has access to a server-side container environment
> Importantly, we never intentionally degrade model quality as a result of demand or other factors, and the issues mentioned above stem from unrelated bugs.

Things they could do that would not technically contradict that:

- Quantize KV cache

- Data aware model quantization where their own evals will show "equivalent perf" but the overall model quality suffers.

Simple fact is that it takes longer to deploy physical compute but somehow they are able to serve more and more inference from a slowly growing pool of hardware. Something has to give...

kiratp··on US economy added just 22,000 jobs in August, unemployment highest in 4 yrs
Source?
kiratp··on Gemini 2.5 Flash Image
It's an arms race.

https://removemysynthid.com/tools/images

kiratp··on How well does the money laundering control system work?
lol look up Civil Asset Forfeiture.
kiratp··on GPT-5: "How many times does the letter b appear in blueberry?"
> Edit: Letter frequency apparently has just become another scripted output, like doing arithmetic. LLMs don't have the ability to do this sort of work inherently, so they're trained to offload the task.

Mechanistic research at the leading labs has shown that LLMs actually do math in token form up to certain scale of difficulty.

> This is a real-time, unedited research walkthrough investigating how GPT-J (a 6 billion parameter LLM) can do addition.

https://youtu.be/OI1we2bUseI

kiratp··on Perplexity is using stealth, undeclared crawlers to evade no-crawl directives
The web browsers that the AI companies are about to ship will make requests that are indistinguishable from user requests. The ship on trying to save minimization has sailed.
kiratp··on O3 and Grok 4 Accidentally Vindicated Neurosymbolic AI
So a sequence of characters that is a python program is “neurosymbolic” but a sequence (of the same domain) in English (a different ruleset) that says “reverse this string” is not?
kiratp··on X changes its terms to bar training of AI models using its content
That will play out exactly like the "Do not track" bit did.
kiratp··on Mistral Code
How do you launch a dev tool with a “contact us” call to action?

It’s like Mistral is choosing to fail here.

Edit: I can't even tell if its a CLI tool, an IDE plugin or a standalone IDE!

Edit 2: oh man! it's at the bottom of the page

Edit 3: "Mistral Code Enterprise is currently only available with an enterprise license." :D

kiratp··on Claude Code: Best practices for agentic coding
The productivity boost can be so massive that this amount of fiddling to control costs is counterproductive.

Developers tend to seriously underestimate the opportunity cost of their own time.

Hint - it’s many multiples of your total compensation broken down to 40 hour work weeks.

Page 1 of 8Next →