2,120 karma · joined September 11, 2011
Still fascinated by what foundational models and cognitive architectures teach us about our own minds. Building beats theorizing, but sometimes they’re the same thing.
https://memory.store
https://diwank.space
hi@diwank.space
for one thing, they said that on AA, sol is "within one point of fable" at 58.9 vs 59.9 but don't clarify that the latter is with safeguards where ~8% of the tasks got routed to opus
i'm not rooting for either and genuinely think that the token efficiency and cheaper price are important but this sort of thing just feels disingenuous :-/
> Fable 5 will be available starting tomorrow, Wednesday, July 1, to users globally on the Claude Platform, Claude.ai, Claude Code, and Claude Cowork. For Pro, Max, Team, and select Enterprise plans,1 Fable 5 will be included for up to 50% of weekly usage limits through July 7, after which it will be available via usage credits. We will re-enable access on AWS, Google Cloud, and Microsoft Foundry as quickly as possible.
also shoutout to tj for being super responsive on github issues!
The problem: if you use multiple AI tools (Claude, ChatGPT, Cursor, etc.), none of them know what the others know. You end up maintaining .md files, pasting context between chats, and re-explaining your project every time you start a new conversation. Power users spend more time briefing their agents than doing actual work.
Memory Store is an MCP server that ingests context from your workplace tools (Slack, email, calendar) and makes it available to any MCP-compatible agent. Make a decision in one tool, the others know. Project status changes, every agent is up to date.
We ran 35 in-depth user interviews and surveyed 90 people before writing a line of product code — 95% had already built workarounds for this problem (custom GPTs, claude.md templates, copy-paste workflows). The pain is real and people are already investing effort to solve it badly.
Early users are telling us things like one founder who tracked investor conversations through Memory Store and estimated talking to 4-5x more people because his agents could draft contextual replies without manual briefing. It helped close his round.
Live in beta now. Would love feedback from anyone who's felt this pain! :)
https://github.blog/changelog/2025-12-18-github-copilot-now-...
"Read our statement on today’s decision in the case involving Google Search."
https://blog.google/outreach-initiatives/public-policy/doj-s...
I wonder if being trained on significant amounts of synthetic data gave it any unique characteristics.
I'd encourage you to give setfit a try, along with aggressively deduplicating your training set, finding top ~2500 clusters per label, and using setfit to train multilabel classifier on that.
Either way- would love to know what worked for you! :)
^[1]: https://diwank.space/field-notes-from-shipping-real-code-wit...
granted- it needs careful planning for CLAUDE.md and all issues and feature requests need a lot of in-depth specifics but it all works. so I am not 100% convinced by this piece. I'd say it's def not easy to get coding agents to be able to manage and write software effectively and specially hard to do so in existing projects but my experience has been across that entire spectrum. I have been sorely disappointed in coding agents and even abandoned a bunch or projects and dozens of pull requests but I have also seen them work.
you can check out that project here: https://github.com/julep-ai/steadytext/
So prescient. I definitely think this will be a thing in the near future ~12-18 months time horizon
> It uses two interdependent recurrent modules: a *high-level module* for abstract, slow planning and a *low-level module* for rapid, detailed computations. This structure enables HRM to achieve significant computational depth while maintaining training stability and efficiency, even with minimal parameters (27 million) and small datasets (~1,000 examples).
> HRM outperforms state-of-the-art CoT models on challenging benchmarks like Sudoku-Extreme, Maze-Hard, and the Abstraction and Reasoning Corpus (ARC-AGI), where CoT methods fail entirely. For instance, it solves 96% of Sudoku puzzles and achieves 40.3% accuracy on ARC-AGI-2, surpassing larger models like Claude 3.7 and DeepSeek R1.
Erm what? How? Needs a computer and sitting down.