HNHacker News
TopNewBestAskShowJobs

diwank

2,120 karma · joined September 11, 2011

Dropped out of Columbia (physics, philosophy) for a Thiel Fellowship. Built Julep AI — open-source AI agent framework, 7k stars on GitHub — because I wanted agents that could actually do things. Now building Memory Store (YC X26): persistent memory for AI agents. Turns out the hard problem isn’t making agents smart, it’s making them remember.

Still fascinated by what foundational models and cognitive architectures teach us about our own minds. Building beats theorizing, but sometimes they’re the same thing.

https://memory.store

https://diwank.space

hi@diwank.space

submissionscomments
diwank··on Tells of a Slop UI
watch this turn into a skill and become a popular meme on X...
diwank··on Show HN: DeepSeek-V4 Latent Reasoning – moving "thinking" into latent space
you lose CoT monitorability which is a big issue since models have become quite powerful and also often deceptive but i do think that efficiency pressure will keep nudging us toward latent reasoning. looped language models are an active research area. but imagine not being able to monitor mythos' thoughts as it's working through a national security need...
diwank··on The real prices of frontier models. Tokens * Price, right?
this is surprisingly high delta. to make matters worse, reasoning tokens account for the majority of tokens and they are completely opaque so it's hard to tell how much of that is prose or code
diwank··on GPT-5.6
i'm not happy with how openai is trying to pit 5.6 sol as a cheaper equivalent to fable here

for one thing, they said that on AA, sol is "within one point of fable" at 58.9 vs 59.9 but don't clarify that the latter is with safeguards where ~8% of the tasks got routed to opus

i'm not rooting for either and genuinely think that the token efficiency and cheaper price are important but this sort of thing just feels disingenuous :-/

diwank··on Department of Commerce has lifted export controls on Claude Fable 5 and Mythos 5
where did you find that? weird coz their post announcing this also mentioned Claude Code:

> Fable 5 will be available starting tomorrow, Wednesday, July 1, to users globally on the Claude Platform, Claude.ai, Claude Code, and Claude Cowork. For Pro, Max, Team, and select Enterprise plans,1 Fable 5 will be included for up to 50% of weekly usage limits through July 7, after which it will be available via usage credits. We will re-enable access on AWS, Google Cloud, and Microsoft Foundry as quickly as possible.

https://www.anthropic.com/news/redeploying-fable-5

diwank··on NSA using Anthropic's Mythos for cyber attacks
""" It remains unclear whether Anthropic’s engineers are assisting the NSA in active operations. However, one person close to the situation said Mythos would be useful for infiltrating the networks of nations such as China or Iran. """
diwank··on Local AI needs to be the norm
in order for us to get there, i think we need a standardized api at the os layer for local models so that the os could optimize, batch and safely allocate resources. something like an analog of chrome's local model "prompt" api but provided and managed by the os itself. the user can choose which model they want to primarily use and so on but all of the heavy lifting and continuous batching is done automatically by the os
diwank··on Is my blue your blue?
and device dependent. this is a very tricky thing to get rendered consistently
diwank··on Is my blue your blue? (2024)
this is actually a surprisingly rich area of debate in philosophy of mind. see: https://plato.stanford.edu/entries/qualia-inverted/
diwank··on Show HN: Postgres extension for BM25 relevance-ranked full-text search
had a bad experience with pg_search (paradedb) in the past
diwank··on Show HN: Postgres extension for BM25 relevance-ranked full-text search
we have been using pg_textsearch in production for a few weeks now, and it's been fairly stable and super speedy. we used to use paradedb (aka pg_search -- it's quite annoying that the two or so similarly named), but paradedb was extremely unstable, led to serious data corruption a bunch of times. in fact, before switching to pg_textsearch, we just switched over to plain trigram search coz paradedb was tanking our db so often...

also shoutout to tj for being super responsive on github issues!

diwank··on Day 1 of ARC-AGI-3
this is so disingenuous on symbolica's part. these insincere announcements just make it harder for genuine attempts and novel ideas
diwank··on Antimatter has been transported for the first time
Angels & Demons anyone?
diwank··on I want to wash my car. The car wash is 50 meters away. Should I walk or drive?
opus 4.6 gets it right more than half the times
diwank··on Ask HN: What are you working on? (February 2026)
Working on Memory Store: persistent, shared memory for all your AI agents.

https://memory.store

The problem: if you use multiple AI tools (Claude, ChatGPT, Cursor, etc.), none of them know what the others know. You end up maintaining .md files, pasting context between chats, and re-explaining your project every time you start a new conversation. Power users spend more time briefing their agents than doing actual work.

Memory Store is an MCP server that ingests context from your workplace tools (Slack, email, calendar) and makes it available to any MCP-compatible agent. Make a decision in one tool, the others know. Project status changes, every agent is up to date.

We ran 35 in-depth user interviews and surveyed 90 people before writing a line of product code — 95% had already built workarounds for this problem (custom GPTs, claude.md templates, copy-paste workflows). The pain is real and people are already investing effort to solve it badly.

Early users are telling us things like one founder who tracked investor conversations through Memory Store and estimated talking to 4-5x more people because his agents could draft contextual replies without manual briefing. It helped close his round.

Live in beta now. Would love feedback from anyone who's felt this pain! :)

diwank··on GPT-5.2 and GPT-5.2-Codex are now 40% faster
I dont think this is Cerebras. Running on cerebras would change model behavior a bit and it could potentially get a ~10x speedup and it'd be more expensive. So most likely this is them writing new more optimized kernels for Blackwell series maybe?
diwank··on A battle over Canada’s mystery brain disease
yeah I agree. this is really unfortunate because it seems that there is something systemic here at play which has become twisted up in a cult of personality and that's made a rigorous scientific investigation very difficult
diwank··on OpenAI are quietly adopting skills, now available in ChatGPT and Codex CLI
GitHub Copilot now supports Agent Skills

https://github.blog/changelog/2025-12-18-github-copilot-now-...

diwank··on Show HN: Why write code if the LLM can just do the thing? (web app experiment)
Just in time UI is incredibly promising direction. I don't expect (in the near term) that entire apps would do this but many small parts of them would really benefit. For instance, website/app tours could be just generated atop the existing ui.
diwank··on Claude Haiku 4.5
I am a bit mind boggled by the pricing lately, especially since the cost increased even further. Is this driven by choices in model deployment (unquantized etc) or simply by perceived quality (as in 'hey our model is crazy good and we are going to charge for it)?
diwank··on Ask HN: E-ink devices with real AI/LLM integration?
Remarkable is really lagging behind on this. I was thinking of finally biting the bullet and writing an app for the Paper Pro. Any ideas/takers?
diwank··on Google can keep its Chrome browser but will be barred from exclusive contracts
Google's response:

"Read our statement on today’s decision in the case involving Google Search."

https://blog.google/outreach-initiatives/public-policy/doj-s...

diwank··on Visualizing GPT-OSS-20B embeddings
Agreed. The fact that it has any structure at all is fascinating (and super pretty). Could signal at interesting internal structures. I would love to see a version for Qwen-3 and Mistral too!

I wonder if being trained on significant amounts of synthetic data gave it any unique characteristics.

diwank··on Gemma 3 270M: Compact model for hyper-efficient AI
also ettin is a new favorite and a solid alternative: https://huggingface.co/jhu-clsp/ettin-encoder-1b

I'd encourage you to give setfit a try, along with aggressively deduplicating your training set, finding top ~2500 clusters per label, and using setfit to train multilabel classifier on that.

Either way- would love to know what worked for you! :)

diwank··on Why LLMs can't really build software
It's coming soon! I think this experiment has really taught me a lot about the limits of agentic code assistants, stuff that they're good at, they're insanely good at, and stuff that they're horrible at and cannot seem to overcome. I did write a little bit about how I use Claude Code [1] before I started this project a while back, and I'm planning to finish a sequel pretty soon.

^[1]: https://diwank.space/field-notes-from-shipping-real-code-wit...

diwank··on Why LLMs can't really build software
yup. I started a fully autonomous, 100% vibe coded side project called steadytext, mostly expecting it to hit a wall, with LLMs eventually struggling to maintain or fix any non-trivial bug in it. turns out I was wrong, not only has claude opus been able to write up a pretty complex 7k LoC project with a python library, a CLI, _and_ a postgres extension. It actively maintains it and is able to fix filed issues and feature requests entirely on its own. It is completely vibe coded, I have never even looked at 90% of the code in that repo. it has full test coverage, passes CI, and we use it in production!

granted- it needs careful planning for CLAUDE.md and all issues and feature requests need a lot of in-depth specifics but it all works. so I am not 100% convinced by this piece. I'd say it's def not easy to get coding agents to be able to manage and write software effectively and specially hard to do so in existing projects but my experience has been across that entire spectrum. I have been sorely disappointed in coding agents and even abandoned a bunch or projects and dozens of pull requests but I have also seen them work.

you can check out that project here: https://github.com/julep-ai/steadytext/

diwank··on Genie 3: A new frontier for world models
> Future robots may learn in their dreams...

So prescient. I definitely think this will be a thing in the near future ~12-18 months time horizon

diwank··on Hierarchical Reasoning Model
Same! Guan’s work on sample packing during finetuning has become a staple. His openchat code is also super simple and easy to understand.
diwank··on Hierarchical Reasoning Model
I think that’s too harsh a position solely for not being peer reviewed yet. Neither of yhe original mamba1 and mamba2 papers were peer reviewed. That said, strong claims warrant strong proofs, and I’m also trying to reproduce the results locally.
diwank··on Hierarchical Reasoning Model
Exactly!

> It uses two interdependent recurrent modules: a *high-level module* for abstract, slow planning and a *low-level module* for rapid, detailed computations. This structure enables HRM to achieve significant computational depth while maintaining training stability and efficiency, even with minimal parameters (27 million) and small datasets (~1,000 examples).

> HRM outperforms state-of-the-art CoT models on challenging benchmarks like Sudoku-Extreme, Maze-Hard, and the Abstraction and Reasoning Corpus (ARC-AGI), where CoT methods fail entirely. For instance, it solves 96% of Sudoku puzzles and achieves 40.3% accuracy on ARC-AGI-2, surpassing larger models like Claude 3.7 and DeepSeek R1.

Erm what? How? Needs a computer and sitting down.

Page 1 of 8Next →