HNHacker News
TopNewBestAskShowJobs

sonink

590 karma · joined November 23, 2007

https://x.com/sonink

https://nishantsoni.com

https://nonbios.ai

submissionscomments
sonink··on Introducing System One Models and Jev
Spent a lot of time - but this makes zero sense to me. It can, maybe, return type safe outputs faster than larger llms - but there is little reason to believe that it will be more accurate. It does absolutely hallucinate - and seems to me that the claim is largely misleading.

You architect your systems with typesafe - because it is marginally faster, but inaccurate - to do what ? You can just wait for the next version of LLM's to get more accuracy at the same cost - or just use a faster model right now from a different provider.

sonink··on Lovable raises $400M Series C
Not at all - how do you build a classified site on Wix/Framer ? Afaik they are for only static/marketing sites - plus a cms to host a blog.

These are all apps with multi-tenancy - database backed - payment integrations - the works. The classified site (marketcity.org) has features like user chats/dms, social media sharing, paid upgrades etc.

Actually now the capabilities are dramatically increasing - you can now vibecode like a python based image/video processing pipeline directly - all through a webapp. And very soon - something more ambitious, like finetune an LLM - again all through a simple chat webapp.

sonink··on Lovable raises $400M Series C
I think software engineers might be in a echo chamber around Lovable and apps like this. It would have been a surprise to me too, as I also build software for a living, but it just so happens that I also run a company - in the same space as lovable - so I can give you some real examples:

1. Dance school management SaaS

2. Classified sites for Netherlands

3. An old school website hosting service

All of these were built by entrepreneurs who do not have a tech background. And by no tech background, I mean most of them dont know github, some of them dont know reddit and some of them dont know linkedin. The way they 'backup' their work is by taking successive dumps of the entire folder on local.

They already know that the business of software is one of the better businesses out there. Till now the option they had was to outsource to a dev shop - which will set them back to $10k-$50k atleast - but they intutively understand that the business may or may not work. So they were on the fence.

Now they build themselves - each of the three above have spent around $2k-$5k on our platform over the last quarter. Some of the businesses are working, some not so much - some are trodding along. So these guys do some updates here and there - are figuring out marketing over time - but if it seems like its not taking off, they start building the next thing on their mind. And once they 'understand' a platform they stick to it. Couple of grands dropped to build their next idea is the cheapest they have ever paid.

sonink··on The Human Is the Loop
Not too sure - I never read that many books - and it wasnt an issue at all. I think writing more, without AI, is helping me get back to my pre-AI-stink self.
sonink··on The Human Is the Loop
"Direction" is what I mean by strategic thought - but even if the high level direction came from you - the tactical approach might fall prey to the same over confidence which destroys direction setting in the first place.

I use it very sparingly for stuff like this - and even when I do - I discount its outputs heavily and maybe just use it let me see some blind spots and sanity testing.

sonink··on The Human Is the Loop
I am coming around to the view that for somethings it is better to not use AI at all. Writing and Strategic thought are those things for me.

The problem with writing, and something I am noticing lately, is that AI robs your ability to create prose in your style. Even if you use AI just a little bit to touch up your writing - you start being influenced - And over time it adds up - you start writing like AI just a little bit every now and then. It seems to be robbing me of my ability to create original prose.

Strategic thought is the other thing - like if I am deciding what we, as a company should focus on long term, it is easy for LLM's to give you a wrong direction. I think this is more because Strategic directions require a ton of context - and it is almost impossible for you to feed the ai EVERYTHING. There might be some things which are easy to miss out, which might make all the difference, and the LLM would assume that the 'requirements' are complete and would go to town with it. I suspect, this is largely because the current SOTA LLM's are being optimised for code and there isnt enough data in the training set for them to learn long term thought.

Coding is I think fine for me. But even there, I dont let the agent use memory at all. I create the context - and it runs from that only. The biggest unlock here is that I can now program in any language. And it frees up sometime to think in systems and come up with more robust architectures.

sonink··on I tricked Claude into leaking your deepest, darkest secrets
Its a bit wild to me that there hasnt been a pushback against enabling memories by frontier AI companies. This data is something advertisers could only dream off. Before AI, most of this data was approximated by whatever little information could be gleaned from the websites we visit. But now people are handing over their deepest darkest secrets and pretty much EVERYTHING to AI on a platter.

Maybe its just me who is paranoid because I happen to spend a fair bit of time in the advertising world, but the first thing I did when memory was launched on Claude/Chatgpt - was to switch them off. And it helps that they are not even useful, and would actually downgrade your experience by polluting the context of irrelevant details. I go one step ahead - if there is a personal discussion you want to have - maybe use another account like provided by the likes of companies like openrouter etc.

I would argue that we should have regulation that should prohibit the storage of user profile information by AI companies, and any such memories feature should exclusively reside on the users servers. Infact, maybe go one step ahead, that 'memory' firms cannot be owned by AI firms and vice versa.

sonink··on MicroVMs: Run isolated sandboxes with full lifecycle control
I dont get it either - I was going to ask the same question but found this.

We have been doing the exact opposite - instead of micro VM's we are giving agents larger VMs.

Previously we were giving them 1GB RAM VM's - now we have upped to 4 GB RAM VM's. When the agent is working - the real cost is in the inference. There is no reason to keep the agent waiting because your VM is too damn slow. So we moved to larger and faster VMs.

The agent might install a package, or run a script - and now it moves along just faster. Not to mention that if the agent is installing a 'fat' SDK, like maybe android sdk, a thicker RAM just moves along everything smoothly without breakages. The incremental amount we pay for the bigger VM is more than justified by the increase in agent performance.

And all the tooling that has already been built up for standard human operated VM's just works pretty well out of the box. We are able to spin up VM's pretty much on demand and purge them clean once the work is done.

We are moving to 8 GB RAMs/4CPUs sometime this year, and GPU's hopefully sometime next.

sonink··on Om Malik has died
Om had a very honest voice - I never met him, but have read his takes for a very long time now. Om Shanti.
sonink··on U.S. government will decide who gets to use GPT-5.6
There is an assumption that everyone is making here - that China will not do the same. It is entirely possible, that China restricts their frontier models - as and when they are developed - to only Chinese citizens. And India follows along.

IMO AI is different from everything else. It is a weapon as potent as nuclear. It is only natural that it be treated as one.

sonink··on Policy on the AI Exponential
I see a lot of skepticism in Dario's position in this forum. But allow me to argue the opposite.

I think the key argument that this skepticism lies on is that he himself gained from AI - specifically building Frontier AI models - and this is basically regulatory capture disguised as doomerism.

Fair points - but I think this is a more charitable version of this. Dario is building Anthropic because that is the most valuable thing he can build, or at least that is what his conviction has been. The success of Anthropic and the impending IPO is proof that this conviction has not only been correct but has largely played out very successfully. Dario understands the true nature of AI and he has welded that power to immense personal benefit.

But maybe he also sees the potential danger to AI which he is trying to address through these posts and regulatory initiatives. There are three reasons why I would support the charitable version:

Firstly, personal gain and societal benefit can coexist in the same individual. And both of them might drive towards opposite agendas. But that doesn't necessarily have to mean that the impulse driving the societal benefit is not earnest. In fact if you would look at Dario’s proposal - like closing the data broker loophole - several of them could constrain Anthropic instead of benefitting them.

Secondly, he expects that his concerns on the negative potential of AI will be taken seriously, if he is actually running the Frontier AI company. And there is some truth to this argument. The only reason we are discussing this is because he is the CEO of Anthropic. He is probably the most influential figure outside of the government who has to be taken seriously when he claims something like this.

Thirdly, and most importantly, Dario has previously demonstrated that he is willing to sacrifice personal/corporate gains for societal benefits. The proof is the classification by the US DoD of Anthropic as a supply chain risk when Anthropic refused to completely cooperate with the military to develop fully Autonomous AI weapons and enable mass surveillance. It would have been only too easy for Dario to accept if personal benefit was his only concern - and OpenAI was more than ready to step in their place.

Even with Mythos - Anthropic could have released the model to the public broadly. But they took their time to reduce the potential danger - as best as they could. Despite the fact that GPT-5.5 was nipping in the buds in what is becoming a very competitive market.

That being said, just because Dario is acting in good faith, does not mean that this will all result in good outcomes. The FAA-styled regulation could still end up favoring incumbents - some of whom might choose to not act in good faith. A more diversified capability can potentially limit that power ending up in a small number of wrong hands. Just because Anthropic is the leader right now, doesn't mean that they will always be. Maybe someone else tomorrow benefits from this regulatory capture at the expense of everyone else - and Dario might have a hand in driving it.

sonink··on When AI Builds Itself: Our progress toward recursive self-improvement
Broadly agree to this position - I think there are some people skeptical that Anthropic is doing this for regulatory capture - but I think there are being honest about they are seeing and how regulation should catch up.

I for one, believe that we should pause all work on AI for the forseeable future. This is almost impossible to orchestrate - but we should still try nevertheless. Maybe we are not able to pause, but we are able to slow down. That might give us more room, to maybe able to pause in the future. But going ahead is too dangerous.

And its not just Anthropic which is saying this. Even Geoffry Hinton has said the same thing. If there is a non-zero chance that AI can kill all of humanity, and both Geoffry and Anthropic have the same position, then it makes sense for us to be hundred percent sure before we move ahead. Dario/Anthropic have already made their money from AI, maybe they are just being honest about what they think lies ahead.

sonink··on Claude Opus 4.8
Same here - we never bumped to 4.7 in our agentic app. Continue to use 4.6.
sonink··on Language Models Need Sleep
I think you are missing the most important part - forgetting. The missing "memory" layers is consolidation, prioritization AND forgetting (what is not important).

Also not too sure about provenance and inspectability - it is part of memory. If the source is deemed 'important' it will survive forgetting. If not, then maybe not. And its ok. I am sure you dont know the exact source who told you that the capital of France is Paris. You forgot, and its no big deal.

sonink··on Multi-Agent is a snake oil
Agree. We came to the same conclusion sometime last year. Now we solely focus on a single threaded agent - and make them run longer in a focussed way.

You dont need a team of humans to write great code - one engineer with solitude and focus is all it takes.

sonink··on Sleep research led to a new sleep apnea drug
Keep at it - there are a ton of things you can try which will help. Try to use reddit to see what other people are looking at.

If there is one thing I can recommend you - get the Wellue O2 Monitor. This wil montior your O2 saturation through the night - if the saturation is dropping, you know that there is an issue, irrespective of what the resmed cpap claims. The resmed cpap figures cant be trusted imo.

But once you know what the saturation levels are, you can try a LOT of things which might help. Different cpap masks, head of bed elevation, mouth tape, neck straps - what you end up using will depend on your condition. But imo - dont rely on doctors to diagnose and fix you - you will have to do it yourself. Fortunatately there is a ton of people who are trying things out in reddit and extremely unlikely that you have a conditoin which is not fixable.

Just keep at it and track it using the O2 ring and try out what works for you.

sonink··on Sleep research led to a new sleep apnea drug
You should consider getting an Wellue O2 ring. This is something you can use to monitor your oxygen saturation throughout the night. Use it with the CPAP and also otherwise. If your oxygen saturation is better with CPAP - you know that it is working. You will eventually feel better.

The main thing about CPAP is that, and imo almost everyone gets wrong, is that you need to titrate it. CPAP is sold as an Automatic Pressure device, but in practice it doesnt work like that. You almost always need to set it just 1 number below and 1 above your required pressure - more like a fixed pressure device. And getting it working correctly - with all the mask combinations, leaking issues, pressure calliberation, supporting gear like mouth tapes and neck bands - can take months. It is incredibly hard - BUT - it is worth it. The best resource for me has been the reddit to get this right.

The key is to track your saturation everyday with all the small tweaks you make and the only way to do it is using something like the O2 ring.

sonink··on Show HN: Semble – Code search for agents that uses 98% fewer tokens than grep
The bigger problem with solutions like these is that most AI already know how to use grep and search really well because of their training. Any such new tool that you handle to the AI, takes away from the cognitive capability of the AI. Humans would normally 'learn' how to operate tools like this - but the learning in LLM's is frozen and they already with a very strong depth in existing tools like grep.

For example, an AI would already use linux commands like tree to traverse the code base. And again it already has good training in this.

The other problem is that it is easy to cook up examples which demonstrate the efficacy of tools like these - but actually proving that the cognitive deficit that such tools result it, is surmounted by their efficacy in long horizon runs. My first contact instinct is that this will result in a net negative 'deployable intelligence' over long horizon runs - make the agent perform worse than using existing tools.

Proving the opposite is a non-trivial problem - but maybe it might be something you want to take up.

sonink··on All your agents are going async
I was of the same view - but then there is this other trend which is putting sync back in favor. And that is that agents are becoming faster. If they are faster - it makes sense to stick around and maintain your 'context' about the task and supervise in real time. The other thing which might keep sync in fashion is that LLM providers are cutting back on cheap tokens. So you have a bigger incentive to stick around and make sure that your agent is not going astray.

The only place I use async now is when I am stepping away and there are a bunch of longer tasks on my plate. So i kick them off and then get to review them when ever I login next. However I dont use this pattern all that much and even then I am not sure if the context switching whenever I get back is really worth it.

Unless the agents get more reliable on long horizon tasks, it seems that async will have limited utility. But can easily see this going into videos feeding the twitter ai launch hype train.

sonink··on OpenClaw isn't fooling me. I remember MS-DOS
Not sure why you think Agentic coding is solved - it isnt imo, and exactly because of the same memory issues.
sonink··on Show HN: Saunas lower nighttime heart rate more than exercise (n=59,000)
I looked into Saunas in detail sometime back as a replacement/complement to exercise. There is a lot of research out there which says Saunas are as beneficial - but at the end of it I reached a similar conclusion - exercise is just better understood, so no point experimenting when something can go wrong.
sonink··on OpenClaw’s memory is unreliable, and you don’t know when it will break
To clarify - we don't read user conversations - what I'm referring to is deploy volume (how many times OpenClaw gets spun up on our infra) and direct conversations I've had with people in my network who deployed OpenClaw independently
sonink··on Launch HN: Freestyle – Sandboxes for Coding Agents
Congratulations on the launch !

We run upwards of a thousand sandboxes for coding agents - but these are all standard VM's that we buy off the shelf from Azure, GCP, Akamai and AWS. I am not sure why we should use this instead of the standard VM's? Pricing could be one part, but not sure if the other features resonate.

Forking is interesting, but I would need to know how it works and if it is in the blast radius of the agent execution. If we need to modify the agent to be cognizant of forking, then that is a complexity which could be very expensive to handle in terms of context. If not, then I am not sure what is the use for it.

Sandbox start time at 500ms is definitely interesting. But its something we already are on track to reproduce with a pooled batch of VM's. So not sure if that in itself is worth paying for the premium.

My two cents on the space is that agents are rapidly becoming more capable to just use the tooling developed for humans. All clouds provide a CLI which agents can already use to orchestrate - they should just use the VM's designed for humans through the CLI. Our agent can already 'login' to any VM on the cloud and use the shell exactly like a human would. No software harness is required for this capability. The agent working on a VM is indistinguishable from humans.

sonink··on The Shell Is the Most Underrated Interface in AI
Paper: https://huggingface.co/papers/2604.00073
sonink··on Inside Nepal's Fake Rescue Racket
> The second method is more troubling. At altitudes above 3,000 metres, mild symptoms of altitude sickness are common. Blood oxygen saturation can drop, hands and feet tingle, headaches develop. In most cases, rest, hydration or a gradual descent is all that is needed. ...investigators found that Diamox (Acetazolamide) tablets, used to prevent altitude sickness, were administered alongside excessive water intake to induce the very symptoms that would justify a rescue call.

This doesnt sound accurate. I have trekked the Himalayas for over a decade - the risks of AMS are very real. Two people I have trekked with have died due to AMS on separate himalayan treks - both had trekked multiple times before, and were well aware of the risks. Both the fatalities were around 12000-14000 feet - much below the Everest Base Camp trek. When AMS hits, you need to descend - as fast as possible, with whatever means you have at your disposal. Otherwise you have unknowingly entered a Russian Roulette.

And Diamox is used as a preventative course for AMS - alongside excessive water intake - this is standard guidelines in all high altitude himalayan treks.

sonink··on OpenAI Has New Focus (on the IPO)
Suggestions are absolutely fine. But this is baiting. Chatgpt could have easily given me that information without the bait. And I would have happily consumed it. And maybe if it did it once, it was fine - but it kept on doing it - bait after bait after bait.

The objective was to increase the engagement "metrics" clearly. The seems to me as if the leadership will take all 'shortcuts' required for growth.

sonink··on OpenAI Has New Focus (on the IPO)
From the article: "You can see that in the recent iterations of ChatGPT. It has become such a sycophant, and creates answers and options, that you end up engaging with it. That’s juicing growth. Facebook style."

This is something I relalized lately. ChatGPT is juicing growth Facebook style. The last time, I asked it a medical question, it answered the question, but ended the answer with something like "Can I tell you one more thing from your X,Y,Z results which is most doctors miss ? " And I replied "yes" to it, and not just once.

I was curious what was going on. And Om nails it in this article - they have imported the Facebook rank and file and they are playing 'Farmville' now.

I was already not positive of what OpenAI is being seen as a corporate, but a "Facebook" version of OpenAI, scares the beejus out of me.

sonink··on Launch HN: Kita (YC W26) – Automate credit review in emerging markets
If your VLM based pipeline is really that good for OCR - and no reason to believe it cant be - why dont you just launch that as a product. The way I see it is that these are two separate products - VLM based OCR for messy documents, and automated credit review for developing markets.

I have some experience setting up automated OCR systems for one of the largest fintechs in Australia - and VLM based pipelines can definitely give an extra edge and this is easily a very large TAM market. However, existing players might also be upgrading their systems so might not be too easy to disrupt. That being said, credit analysis is also a very hard problem, but I am not sure how much quality OCR would help here.

Given what I know, I would focus on the VLM/OCR problem rather than the Credit scoring one.

sonink··on When does MCP make sense vs CLI?
Agree. MCP isn't really required. Skills/CLI/API is good enough.

At the AI startup I work on, we never bothered building MCP's - it just never made sense.

And we were using skills before Claude started calling them skills, so they are kind of supported by default. Skills, CLI, Curl API requests - thats pretty much all you need.

sonink··on A Model of a Mind
> I suppose it would be interesting to have an online space to discuss where things are headed on a logical level, without emotion and ideals and the ridiculous idea that humanity must persevere.

Absolutely. Happy to be part of it if you are able to set it up.

Page 1 of 6Next →