HNHacker News
TopNewBestAskShowJobs

shostack

6,016 karma · joined October 1, 2013

I like digital media and have been in the online advertising and analytics space for a long time. If you're interested in reaching out, my email is "michael.myHNusername@gmail.com".
submissionscomments
shostack··on Plan mode is dead
I am curious what your take is on skill packages like Obra/Superpowers and grillme and the like. On one hand I find them extremely useful for asking me questions. I wouldn't have thought of myself or that I don't feel I would have necessarily gotten in a structured a format just by having the AI interview me. On the other hand, I am not sure how to tell when they have become obsolete because I'm not sure where to even start with generating evals for such a thing.
shostack··on Ember-1
Last I checked they still offer no training ZDR US based hosting. It is one of my three pinned providers for deepseek v4 flash along with Parasail and Deepinfra.
shostack··on Microsoft abandons personal AI chatbot race with Copilot reboot
How timely. This is a great write-up on what I meant: https://news.ycombinator.com/item?id=49850305
shostack··on Microsoft abandons personal AI chatbot race with Copilot reboot
These are effectively the next OS. I wonder if Microsoft realizes that and if so what their strategic play is.
shostack··on GPT-6 Sol and Luna
TY, appreciate the thoughtful reply. Transparently, I have been so on the fence with moving more of my workflows and personal usage over from Hermes+VPS+ZDR model provider, or implementing the "Hermes shell over Codex Subscription" pattern because I worry about:

1. Legal loopholes given OpenAI's advertising aspirations and model training needs

2. Data retention and rising threats of fascism that historically have not served the persecuted very well when fascist regimes get access to said data

3. Risk from centralized collection of that data with a company whose software I do not control in a world where enshitification and lock-in is the norm.

I really wish OpenAI did more to espouse exactly this: "When I say no training, I mean no training. No gimmicks around data vs derived data, synthetic data, preference data, etc." and ideally provide technical reasurrances that this is impossible (eg: certain technical ZDR approaches, etc.).

Do you happen to have a favorite reference to point me at that would document some of those official assurances to the nuanced detail we've discussed here?

shostack··on GPT-6 Sol and Luna
Ted can you confirm your choice of words here to be precise for an audience who is familiar with the nuances, when you say "no training" or "training (opt out)" for personal... Is that inclusive of "sanitized" (or pseudonymized) data?

Your response to the original question is using generalized terminology when there is a very important distinction the OP made by the use of "sanitized."

People want to know to that extent derivatives of their data are being used. Synthetic data has been proven to be effective at generating training data and AI is very good at shuffling context such that you have something where you don't have to say it is "user data."

But there are many shades of gray there for people versed in how the sausage is made. I'm sure you'll appreciate then why your response leaves additional questions in light of that "sanitized" distinction.

shostack··on How GLM built its own inference infrastructure
So effectively, trying to undercut their competitor to reduce its wealth and power?
shostack··on Red and Blue America Have Found Something to Agree On: Flock Cameras Must Go
Where are all the articles stating how switching to a Flock competitor is just as bad and isn't making the issue go away?
shostack··on Don't be the out of touch Kung Fu master
Tech priests reciting blessings to the Omnisiah. Time to do the blessing of checking how much 5h usage I have left.
shostack··on No Man's Sky Cosmos
The negative comments, including the one I'll make here, are from fans who want the game to be more than it is and see the potential but continually get disappointed with its biggest flaw.

That flaw is that for all the features they bolt on for free, the actual depth never feels that much deeper. It feels like the shallow pond got wider.

The worlds and universe still feel repetitive and empty. There is very little emergent gameplay. Mechanics are not complex compared to simulator sandboxes like E:D or more in depth survival crafter games.

And so you end up with what often feels like a kids toy with the edges filed off. Its very colorful, and they keep giving us me shinier versions for free, but it still feels like playing with Duplo when I want it to be LEGO or K'Nex.

And I say the above with nothing but respect for Sean and the team and what they have accomplished. I will be a launch day LNF player as well.

shostack··on Muse – Meta’s personal AI agent
Your comment is slightly reminiscent of the response to dropbox when it was first announced on HN.
shostack··on Muse – Meta’s personal AI agent
How will that be possible with the Confidential VM version?
shostack··on GPT-6 Astra
You can. The point of having your agent keep track of it is that it will likely notice things you won't, and it can automate cataloging it with relevant metadata (prompts, environment, examples, etc.) that make it trivial to automate rerunning those tests when new models launch.
shostack··on Invisible Companies
I wish there were a clear and easy way to identify these small services owned by PE, who often go out of their way to avoid people finding out.

Usually I find out by noticing symptoms of worsening quality, costs increasing more than I might expect, ramp up in aggressive cross selling of services and subscriptions, shifting to call services that are clearly not local and know nothing of our area, etc.

In some cases I'll get an employee that knows the situation and let's spill the PE sale and then I need to find a new service provider.

shostack··on The asteroid currently hitting front end web development
Being a programmer has always meant dealing with abstraction layers and climbing up the abstraction ladder.

You are simply learning a new language, whose syntax happens to look like English (or whatever language you speak), but with new, undiscovered, and constantly changing design patterns and best practices.

shostack··on GPT-6 Astra
True, but they're still friction to be reduced here.

What I desperately want is for 1password or stripe or even Google who already has much of my data, to o come up with a secure solution for online purchases with agentic credit cards where I can effectively get a phone prompt to authorize a purchase while the agent can fully own the checkout flow.

I have seen various things coming on the market for this, but none of them appear aimed at a consumer audience. And I am a firm believer at this point in keeping my payment authorization and history and credentials harness agnostic.

shostack··on GPT-6 Astra
One suggestion is to make a list or make a skill to have your agent keep a list of things you do not feel work well with today's models. And then, when new models come out, periodically, revisit items on that list to see if you get better results.
shostack··on Understanding ChatGPT Work
I consider myself fairly AI native. I was absolutely blown away by how frictionless it felt the other day interfacing with codex using the experimental headless app server on my VPS and using voice mode on my phone connected to it while I had my browser open having a conversation about making edits to my website.

The site uses Astro to hot load edits and so the exceptional Live voice model would use some filler words in response to me asking for an edit and before I knew it, the page had refreshed with the fix.

When people talk about things like OpenClaw and Hermes being a new operating system paradigm this is the sort of UX that comes to mind.

And simply conversing with it with my phone in my pocket and air pods on its the closest I've felt to a live conversation with AI ever.

Kudos to the voice mode and Live voice model teams.

shostack··on Understanding ChatGPT Work
You can have OpenRouter filter for ZDR providers. There are nuances like contractual ZDR vs technical ZDR but definitely worth investigating
shostack··on ChatGPT Work Tool and Skill Reference
As someone who uses ChatGPT and Codex across desktop and phone and VPS, the current topology they have for their different "apps" is very confusing and disjointed. I would really like to see improvements in unifying this and reducing the confusion between what works in which part on which device and enabling seamless context and memory sharing across them.
shostack··on The internet is kind of a predatory cesspit now
It also was because it became a target to influence and manipulate people both by every day scammers and marketers as well as nation state level actors.
shostack··on Small Models Have Arrived
Yeah, I think if we reframed it as "why is there no consumer AWS?" It would make more sense.

AI is a utility that can abstract code to such a high level it is indiscernible from natural language.

shostack··on U.S. State Department pauses immigrant visa applications
The cruelty is the point.
shostack··on A week of using Codex more than Claude
For individual use without specially negotiated Enterprise stuff I can't afford both offer contractual assurances of not using your data for training or ads and measurement and selling your data.

But these are not the same level of technical assurance you get from say, a zero data retention provider on OpenRouter.

Right now I am finding I have to tolerate substantial friction to use Hermes for personal stuff with a ZDR provider and ChatGPT and Codex for less personal stuff because the products and models are simply so much better.

shostack··on Why your local LLM feels dumber than it is
Debugging any LLM output when you also have done substantial harness engineering is a total pita and I wish there were better tools for it to isolate issues.

I spent ages tracking down start appears to be an issue with the current Deepseek v4 flash 0731 version that would cause it to output giant walls of gibberish in Hermes with reasoning turned on.

shostack··on Do Chatbot LLMs Talk Too Much?
I needed this. I had no idea it existed.

I have bent over backwards trying to enforce brevity with deepseek v4 flash to the point where I think I broke some things trying to do prompt injection in my Hermes setup and was still unsuccessful.

Meanwhile Sol blows me away and I want that to be my default for everything now.

In general though I seem to have the most success with a "<=10w" requirement in my prompts.

What I don't see listed and would be a good comparison is the STS models. OpenAI's live model is an absolute joy to talk with.

shostack··on Marketers are Addicted to Bad Data (2020)
Having been at smaller companies without the data, tooling, discipline, and resourcing to conduct viable experiments, and then being at a company that is actually one of the best in the world at it and building solutions for these problems at scale for marketing teams, so many problems come back to the human element in how data is tracked, how teams collaborate or fail to which can lead to pollution and dirty data, and how decisions are made as the moment trade-offs need to be considered it becomes personal and enters messy human relationship territory.

But ultimately from a pure "did this work or not" standpoint you are right. Incrementality experiments are the gold standard.

shostack··on Mushroom behind 'tiny people' hallucinations identified
So like, some spren from Roshar?
shostack··on DeepSeek Harness developer preview
Could memories and blocks of context be made pluggable and unpluggable in this manner to "hot swap" context?
shostack··on OpenAI launches ChatGPT desktop app for Linux
Specifically, I want to use the ChatGPT Android app to use it like I would a desktop remote control session. It is a nice mobile interface. I do not want to have to use a mobile terminal app like Termius. The UX is awful.
Page 1 of 34Next →