HNHacker News
TopNewBestAskShowJobs

_bobm

12 karma · joined October 21, 2020

submissionscomments
_bobm··on One Month Without AI
Well written.

I think there might be (dare I say) a middle ground to get the productivity of the llm, esp as we evolve them, while still maintain a global and even fine-grain comprehension of a code base.

It is not a simple change, however, but a fundamental one.

Overall I think we are still living in the past and try to apply ourselves to the future. But if the ai craze is to be taken clear-headedly for what it is, it is a complete break from the von Neumann computer and all its resulting artifacts. So why should we use the same tools?

_bobm··on Plan mode is dead
So planning is tied to the spec. This much is clear.

What the author I think is hinting at is not planning alone but "the development and evolution of any program and the state of this program throughout the planning, elaboration, and eventual runtime".

I chuckle at the thought that throwing an md file or a prompt at this problem is sufficient.

So, I posit that if we want any agentic code to evolve meaningfully in the short future and over the long run, we have to have a system which holds and presents this information, the state of a program, in a coherent manner to a human operator. No other way. No other way. And this I say to both nay and yay sayers.

You can argue also that this is part of an even bigger thing. But it is not part of the current discussion on planning and speccing in agentic systems.

OpenAI and anthropic can throw all the billions they don't have at this and adjacent issues, but if this is not solved then they don't have anything.

_bobm··on Another terminal/shell AI assistant (shell native)
I also have made similar agents. cool stuff happens when you compose them and a composition is merely different prompts. this strips down the agent to the minimal.

It is not clear what is the best position for the agent or in other words how colesly one wants to integrate with the terminal, shell, or multiplexer.

But over all it does show that maybe the future of agents is not huge blobs of semi-reingeneered multiplexers or see above, but a lean agent with a session stream for context.

great job, i will follow your progress.

_bobm··on 25 Years of Mass Surveillance Is Enough
Hear, hear!
_bobm··on The Coxon psyop was a decade and more than a billion dollars in the making
https://x.com/kevinnbass/status/2098579876194263467#m
_bobm··on The Navier–Stokes Millennium Prize Problem
I also find it hilarious. I wonder what will happen if they don't solve this.
_bobm··on The Navier–Stokes Millennium Prize Problem
I not only think it is not trivial, I know it is not solved.
_bobm··on The Navier–Stokes Millennium Prize Problem
let's say that they ask a single question for each session they get. they are immediately doubling the compute they need in processing and then post-processing the same session twice.

nothing trivial about it. not saying they cannot feed "their own LLM" saying it isn't trivial especially at scale.

if you do not trust me try it without the "at scale" part.

_bobm··on The Navier–Stokes Millennium Prize Problem
People are focused on the drama but the problem showing is the data. This is the elephant in the room and I am surprised that openai can be that stupid with it.

How can openai do this, what is being claimed, at the scale of their entire userbase? If they do this only for particular sessions then how do they sieve through sessions for the good stuff?

How are sessions stored, how are they processed, how much storage and how much compute is used in these pipelines, how economical is it, how fast are the requirements on the storage on the compute growing as userbase grows and generated data grows.

All these questions are far more pertinent than the navier-stokes, but i can only imagine all at openai doubling down on this "very important" mathematical milestone.

_bobm··on The Navier–Stokes Millennium Prize Problem
what is behind "process" it and "further selection" and "refinement" and how big are these "datasets"? These companies ship the encrypted session to you not because they want to.

I agree that they have pipelines for what you are describing but how effective they are at scale and at focusing is the question.

_bobm··on The Navier–Stokes Millennium Prize Problem
Hah, what is the infrastructure which takes user sessions (chats with API keys, directions, navier-stokes math/progress) and regurgitates this into pre-training, RL, fine-tuning data? Or better, in-context data?

People talk about the "compute" but what about the "storage"? Is storage exponentially greater, or soon to be, than the compute? Is the storage going to slow down growing to some constant rate, i.e. all people on earth using chatgpt, or no, on the contrary, it will keep growing?

If there were any shady business, I do not condone it, but technologically we are not there yet for said shady business to happen.

_bobm··on An Alien Mind
All this AGI/ alien nonesense trumpeting from these guys (OpenAI, Anthropic, nVidia) makes me think we are drifting even further from stated goal and money is running dry.

The interesting question is who is left standing after the party is over.

_bobm··on An Alien Mind
I think it is simpler, but bannin open weight is a collateral worth scoring nonetheless.

AI fails to deliver on the multi trilion usd promises:

Gov/venture says: oh, wow, what will we do with all these datacenters, we need money back!

OpenAI/ Anthropic: let's monitor the citizenry. They may be plotting nefarious schemes using AI models.

Gov: great idea. Whew.

Investors: whew!

Tax payers: paying to be in prison.

So not so much as defense against foreign actors, but failure of the self-tooted AGI goal + gov being gov.

_bobm··on Agent memory as a file format
I think this is the wrong approach because everything is external to the model. You end up creating an ad-hoc externalized model scaffolded out of coarser systems, RAGs, files, and so on.

This leads to, if useful at all, to this process of ad-hoc recall inference which has to happen in time. By itself this is not a problem.

The problem is that the model has to be constantly injected in context with the newest version of the "memory state" at each turn or relevant turn.

The newest coherent memory state is also a problem. More or less 50 years of not failures, but not success either. This may be even deeper problem than the externalization problem.

I think we are just in the very beginning and we are slapping database stuff to the transformer hoping it will work, but these deep neural-net architectures categorically show that they are not databases.

There will be a synergistic middle ground, but its shape is still not clear.

_bobm··on Beyond Recall and the Illusion of Competence
In your closing thoughts, you don't mention what is this staying firmly in control? how do you stay firmly in control while reaping the benefit of the model while being more productive?

because if you start a curriculum over what the model wrote then, one can argue, you are better off writing it yourself in the first place.

_bobm··on Zellij 0.45.0: nested sessions, Kitty graphics, a fresh UI
What is the purpose of the "zellij" in the top left corner and why isn't there an option to remove it?
_bobm··on GPT-5.6
are people getting the `<!-- -->` sentinel'd reasoning summaries?
_bobm··on Identity verification on Claude
I find it funny that some comments are arguing why "the innocent users have to fall victims, becoming/ being collateral damage" to this american governmental whim and thus being deprived of access to these models.

But hold on, collateral to what? Is this "our own personal jesus" access that we cannot live without or what?

People don't and cannot learn how to code or what? We don't know how to think?

I am calling their bluff. Fill your own gddam datacenters with "meaning".

Panta rei.

_bobm··on Ask HN: Has anyone replaced Claude/GPT with a local model for daily coding?
Yes, but it isn't.
_bobm··on Ask HN: Has anyone replaced Claude/GPT with a local model for daily coding?
Yes, local models have already all that is needed, they have all the prerequisites.

But what they do not have is the correct shape, the correct approach. This is missing and it shows on multiple scales: it shows in the COT, it shows in the output itself, it shows in the infra to serve the models, it shows in the model orchestration.

This is what anthropic said one year ago:

> Finally, we've introduced thinking summaries for Claude 4 models that use a smaller model to condense lengthy thought processes. This summarization is only needed about 5% of the time—most thought processes are short enough to display in full. Users requiring raw chains of thought for advanced prompt engineering can contact sales about our new Developer Mode to retain full access.

_bobm··on Ask HN: Has anyone replaced Claude/GPT with a local model for daily coding?
idk what is "minifying outputs" in the context of what we are talking about. Opencode is opensource, you can find out what it is doing.

Last time I checked, OpenAI even send (in the response) the summary of the thinking part alreafy in markdown, so opencode has to remove the formatting to format it to their liking.

> Many models now no longer return the entire chain-of-thought (to avoid distillation attacks).

This is what they say: to avoid distillation attacks. And to some large extent this is true. I am saying there is a side- effect and this side- effect (depending on how tin-foilly you want to go) may be either a nice thing to have or it may be the "main reason" for all of this.

The side effect is splicing the inference, brokering requests, and what not, which brings huge benefits at scale.

This was my original point: openweights model to a sota model may be apples to oranges. So when will a local model catchup with its single cot run which is not even shaped properly: well never.

It is apples to oranges.

_bobm··on Ask HN: Has anyone replaced Claude/GPT with a local model for daily coding?
> Can you give an example?

Sure, connect opencode to an openai/chatgpt endpoint and use it. You will notice multiple "thinking" parts per "turn".

I put all of these in quotation because... they are part of the orchestration game. For example, it is not known if the thinking parts of a particular turn are chain of thought thinking summaries or just plain response which is masquaraded and thus orchestrated into appearing as thinking.

Further notice the cadence, word choice and sentence formation. Notice sentence construction. Notice "thinking part" construction and sequencing.

There is pretty heavy orchestration.

> I don't understand, why does it make you think this is the case?

Because not all tokens are equal. And if you waste expensive tokens on mundane tasks you will go out of business. This is the reason.

As I said, if you observe the output from these api endpoints you will notice it.

_bobm··on Ask HN: Has anyone replaced Claude/GPT with a local model for daily coding?
But, guys, when you say Claude/ GPT models, do you stop to think what are these "models"?

One day I thought about how can GPT send thinking parts one after another with a markdown header summary of the thinking block itself. Just think about it.

As a matter of fact, think about these operations, api endpoints, observe their output.

These so called SOTA models are not what meets the eye, and are not at all comparable in the infra department to local models. There is crazy orchestration going on due to the scale of these operations. But also these hard constraints lead to innovation. Innovation nobody speaks about.

I wouldn't say we cannot catchup, but serving our local models through llama, vllm is just the A, B, C of it all. In reality I think what is needed is a replication of said orchestration which I hinted at above.

The SOTA models are a deep orchestration of multiple models operating together it isn't a single model. As such no single model ever will catchup to them until it replicates through training first and then maybe through model architecture this orchestration.

Finally, I would wager that the SOTA "models", as one of these models in this orchestration setup, as served for general consumption, are not so much more capable than qwen 3.6.

I am sure that if you change your perspective you will start noticing the scale of the "magic".

_bobm··on AWS Bedrock to require sharing data with Anthropic for Mythos and future models
Very confident. But will it stick? And if it doesn't -- what then? Back to scheming?
_bobm··on Framework Laptop 13 Pro
What are the news recently?
_bobm··on The local LLM ecosystem doesn’t need Ollama
amen
_bobm··on MinIO repository is no longer maintained
hah, good on them.

nice catch.

_bobm··on Discord will require a face scan or ID for full access next month
How do you see Zulip comparing to anytype, https://anytype.io/ ?
_bobm··on Ask HN: What is the best way to provide continuous context to models?
This is how I view it as well.

And... and...

This results in a _very_ deep implication, which big companies may not be eager to let you see:

they are context processors

Take it for what it is.

_bobm··on Anthropic blocks third-party use of Claude Code subscriptions
But you are not having a free meal lunch are you? You _are paying_ for your meal.

Worse: you are the meal as well.

Do you see this?

Page 1 of 2Next →