HNHacker News
TopNewBestAskShowJobs

ssivark

6,379 karma · joined January 3, 2012

http://sivark.me/

email: siva dot rk dot sw at Google's email service

(Opinions my own; not representing any entity)

submissionscomments
ssivark··on LeCun has "zero concerns" about AI wiping out humanity, recent "rogue" incidents
This is a motte & bailey moment. Parent comment stated something much sharper, that I responded to:

> I do not understand how you can see what has happened in the last 3 years and say "it will not get us there" no matter what "there" is.

--

Your statement is something much weaker, and I would still question what exactly "general" means when AI capabilities are commonly accepted to be so "jagged".

ssivark··on What Meta got right with Muse
> requires you to be way more flexible in your assumptions.

This is what happens if you try to read the data like gospel without having a mental model to interpret it... in other words, critical thinking.

Chatbots are a failure because chat capabilities have sucked for most of the past decade and chatbots weren't really capable enough to do anything meaningful. Not because chat is fundamentally bad. Similarly with auto-router.

Also, if you want to "cross the chasm" you cannot measure reactions from current AI users (early adopters) and extrapolate to the rest of the populace. Early adopters are paying you for tokens and try to wieid them like a surgical tool; across the chasm the rest really don't want to know anything about "tokens" or models tiers or "thinking level".

ssivark··on LeCun has "zero concerns" about AI wiping out humanity, recent "rogue" incidents
> 4 years ago, a program that could [...] we would have called it AGI

If you had told someone in the 1800s that a machine could instantly multiply 100 digit numbers, that would have been considered dazzlingly intelligent. And yet we are not that dazzled by our calculators today (despite how useful they might be!).

ssivark··on Make Tmux the OS
Interesting and expansive post about what the modern desktop experience should be like, and most of it has nothing to do with Tmux except if you squint in some very particular way.
ssivark··on DeepSeek Harness Desktop for macOS and Windows
Exciting -- looking forward to trying it out!
ssivark··on DeepSeek Harness Desktop for macOS and Windows
This page doesn't emphasize the cordis architecture [1], but that's the most exciting thing about this -- not just yet another harness. It has the potential to make this into something like the emacs of harnesses! I think this might be particularly potent for long-running agents.

[1] https://arxiv.org/abs/2608.25512

ssivark··on Functional Ultrasound Imaging (fUSI) from scratch
Merge labs (previously Forest Neurotech) is starting with people who have a certain surgery that needs excising a window in the skull -- and using that as an acoustic window. And in parallel hoping to figure out sufficiently good imaging across the skull for the general user.

There was also an exciting fUS imaging demo from Aleph neuro a few months ago.

ssivark··on Dots: Always-on agents
Perhaps highlight the AI assistant aspect more prominently? If it weren't for seeing this link in this HN discussion, I would have interpreted the offering as some collaborative workspace, a bit like nextcloud.
ssivark··on Dots: Always-on agents
Windows had linux, IE had firefox and iOS had Android. I'm sure we'll see a relatively open assistant harness that token providers will rally around (at some point once the burgeoning field of assistants settles a bit, with clear use cases and value propositions)
ssivark··on Dots: Always-on agents
> Other inference providers should counter this by providing their own version of standardized managed agents

The whole personal agents field is still at a nascent stage. As the field matures and we learn winning use cases and robust operating patterns, we'll also see a few open-source agent harnesses doing well. That would be the signal for commodity infra providers to start providing hosted assistants. Too much froth to keep up with, before that point.

ssivark··on Ember-1
LLM inference has two very different regimes of work: prefill & decode. You can think of the former roughly as processing a pre-specified prompt, and the latter as sequential processing (auto-regressive token generation) eg. "chain of thought". The latter is very important for LLMs and cannot be ignored; it deeply influences infra design, even necessitates copious amounts of high-bandwidth memory. Jev-like models can ignore the latter and therefore optimize much better for the former, consequently operating at both better cost and latency.
ssivark··on Ember-1
This was exactly the crux of Nathan Lambert's recent testimony to a group of US Congressional members/staff: https://www.interconnects.ai/p/the-current-balance-of-power-...
ssivark··on There are no "rogue" AI agents
> legal culpability and misalignment are two separate topics that should not be mixed

Legal culpability for AI labs is exactly the thing that would incentivize -- and hence ensure -- aligned behavior from models.

The last time there was a claim about GPT-4 exhibiting misaligned behavior [1] it turns out it was prompted and pushed to behave so by humans at OpenAI, and OpenAI clearly lied in the GPT-4 system card.

[1]: https://aiguide.substack.com/p/did-gpt-4-hire-and-then-lie-t...

ssivark··on Fable 5 – Median thinking declined in August
The actual tokens might be non-deterministic, but you could look for proxy measures that are supposed to be invariant. Eg. correctness/performance on benchmarks, "thinking level" on complex problems, etc
ssivark··on I don't like passkeys
WDYM -- I thought "try another way" is supposed to list all possible options instead of cycling through them!
ssivark··on OpenAI Discloses Six New Incidents of ‘Concerning’ A.I. Behavior
> OpenAI cautioned that the reports were individual snapshots and “shouldn’t be considered reflective of how often misalignment occurs.”

Sounds like a proper infestation of roaches!

ssivark··on OpenAI buys smartphone camera maker Glass Imaging for $300M
The time is ripening to disrupt Android/iOS. The place to start would be to rethink the device around AI. No need for apps or messy integrations with third-party crap. Just a simple device with a great new UX (not typing on a touchpad!) that is slick and does a few things very well (personal agent, coordinating with your personal knowledge base, handling external communications and your interface with the world). This could very easily supplant a phone.

AI is rapidly getting to such a place where the you don't really need a developer ecosystem to get your device off the ground. As "plugins" into your agent at some later point -- sure, why not.

ssivark··on OpenAI bots knew about the RubyGems caching vulnerability
Perhaps don't deploy random weights of unknown origin?

Also not every model provider might be capable of babysitting all your uncontrolled agent deployments. If you want SLOs, get into a contractual relationship with entities whose weights you deploy, and also monitor your agents so they don't go off the rails.

All this is just like deploying any other tech in the world eg. if you buy a car, or a chainsaw, or a book.

ssivark··on OpenAI bots knew about the RubyGems caching vulnerability
The problem with third party audits is that it allows OAI/Ant to shrug off any further responsibility and claim that they are following best practices (basically, reward hacking). The only real solution is to make them absorb liability for the actions of their agents -- because they are the ones giving agency to their models and allowing them to run amok.
ssivark··on OpenAI bots knew about the RubyGems caching vulnerability
> Source? How do you know they were "prompted to hack to get answers"? How do you guarantee they will always listen to you when you say "do not hack outside systems". They are not classical deterministic programs doing exactly what you say. They are trained to follow orders by RL, but it's not a perfect process.

Who gives a shit? Not my circus; not my monkeys! It's the responsibility of whoever deploys the agents that they are instructed / sandboxed well enough that they can't cause collateral damage. That is the only way this doesn't get out of hand with everybody deploying their agents / robots for a world of utter chaos.

It is impossible (and asinine) to audit every model and deployment; far better to impose liability and the the socio-legal system figure it out.

ssivark··on Garry Tan wants US open-weight AI labs to 'distill' frontier models, too
> to separate ZDR, attestation, and confidential routes

Could you please clarify what that means? Given what I've been searching for, I might in principle be part of your intended customer profile, but I can't figure out whether you are merely doing routing (alternative to OpenRouter) or also inference (alternative to the names I've mentioned above). If it's merely routing, then how do you protect me from any potential misbehavior on the part of the inference provider?

Just feedback for what you're building, so please take this in a positive spirit... I'm an AI researcher and not quite an infra guy, and I'm making recommendations on token APIs for several less knowledgeable around me (I've gotten a few people set up with Baseten recently), and I couldn't figure out whether/why I would be interested in TrustedRouter. You should communicate the story better :-)

EDIT: Here's what I now understand after some digging; please correct if wrong.

There are some M token providers (not the names I listed above?) who provide cryptographic guarantees about inference services. But somebody still needs to verify what they do on each request. For an individual running a single harness, that harness would be a logical place to perform this verification if possible. For an org with N users each running their own harness, TrustedRouter solves the N*M problem and becomes the single gateway for trusted inference -- provided one somehow trusts/verifies TrustedRouter.

ssivark··on Garry Tan wants US open-weight AI labs to 'distill' frontier models, too
What about inference providers like Baseten, Modal, Fireworks, Together, etc? I thought one of their value propositions was inference (using open weights models) that guarantees with crisp terms that they will not use your data.
ssivark··on More questions about whether researchers can trust OpenAI with unpublished math
There is potentially a world of difference between how you interpret what is fair and what the terms of service contractually guarantee.
ssivark··on Ask HN: How do you manage skills files?
Don't know why the below comment by killix got flagged; it's a legitimate point.

In the current version of my setup, I've decided to accept that tradeoff.

But it would also be interesting to check whether agent behavior can be controlled well enough by a skill-management skill telling them to synchronously commit any changes with their signature; that would get the best of both worlds.

ssivark··on Ask HN: Would you read a statistics textbook?
> Someone who has no idea what a standard deviation is can't intuit about distributions.

I disagree vehemently with this claim. I could cite my experience in teaching this topic to liberal arts / humanities college students in the US, but it is really more obvious than that. Anyone can understand a histogram easily and far more intuitively than they can understand the formula for a standard deviation and whether it must divide by N or N-1. Statistics courses and textbooks get stuck on that kind of pedantry, and most students end up missing the forest for the trees.

ssivark··on Ask HN: Would you read a statistics textbook?
Look at the histogram and think about what distribution one could reasonably impute from samples. And what you would set as bounds for "outliers", per your needs. While we're at it, let me also say that it might be useful to specify outlier bounds not just based on the spread in sample values, but the costs/payoffs they imply for your application.

If you are not doing something crazy, most reasonable people would agree with your judgement. Conversely, if you are making non-obvious inferences where reasonable people disagree, you are in murky water and no sophisticated statistical method will save you. Math is not magic; theorems merely recycle (launder) modeling assumptions into results.

ssivark··on Ask HN: Would you read a statistics textbook?
My personal opinion is that statistics textbooks usually come from a prescriptive perspective, and that makes it challenging for the reader to get visceral intuition for what is actually going on. Any reader would be far better off just visualizing the damn distribution / samples and using reasonable judgement, instead of implicitly assuming a Gaussians distribution and blindly memorizing tests / formulae. Making the distributions explicit allows us to model them and get an intuition for what the samples are telling us. I would whole-heartedly recommend the Model based machine learning book to anyone (online version is free) https://mbmlbook.com/
ssivark··on Ask HN: How do you manage skills files?
To the extent that skills are contextual guidance (for this author, this project, etc) and not just (raw) capabilities they are unlikely to be eaten by models.

I maintain all my skill files in a central location (like dotfile management) and have guix home sync it to the skill folders of various harnesses that I'm playing with (codex, pi, antigravity, Claude Code, Deepseek harness, etc). They're set up to be bidirectional links rather than read-only like the default configuration, so I can keep editing them / adding to the corpus from any harness.

This works well for skills since all harnesses expect the same format, but is more annoying for other features.

EDIT: This is actually an example of a potentially useful skill. You might choose to manage your skills slightly differently. All you need to do is write a skill-management skill for your agents to be able to wire things up correctly / access them for edits.

Some other nifty skills/plugins in my experience: render latex equations, cetz diagrams inline, jujutsu, guix, code reviewer, writing feedback.

ssivark··on GPT-6 Astra on robot arms
It's one of several downstream problems, including poorer quality of life by various metrics. By emphasizing a PR problem for businesses more than other human problems downstream or the root problem, you are demonstrating priorities that prosocial humans viscerally disagree with.
ssivark··on GPT-6 Astra on robot arms
That's like saying a law and order problem is a PR problem because word got out that crime is spiking.

The PR problem is downstream consequence of what can be measured, but the root cause is upstream. Focusing directly on solving the downstream problem is basically Goodharting the metric -- thereby ensuring the upstream problem never really gets solved, and the downstream metric becomes fake and decoupled from the true situation.

Once you "solve the PR problem" what is the incentive to ensure that the root problem is actually solved now that you have successfully blinded yourself (society) from measuring the true problem?

Page 1 of 34Next →