I haven't tried deepseek yet, i should check this one out.
1,123 karma · joined March 3, 2018
I haven't tried deepseek yet, i should check this one out.
This is almost standard practice in any competitive industry anyways. Disassemble your competitor's product, study it and try to reproduce / improve.
And when i can use it, it just drains the quota 5 times faster than codex or claude.
Their plan is a scam
I subscribed to their max plan to try it out. It counted me 700M tokens and drained my weekly quota in under 2 days.
Quota just reset less than 24h ago and i'm already >60% weekly quota usage.
For reference the kind of work i did would have used somewhere between 3% and 5% of Codex max or Claude max.
The model is good, the plan is a scam
I signed up to their max plan yesterday, did some light coding work, and i'm at 180M tokens used and 40% weekly quota gone.
Even when tokenmaxxing on the Claude Max or GPT $200 plan, i couldn't get more than 20% quota gone per day.
Although it's gonna be more difficult to come up with a Fable competitor than a m365 one
I've been working with gpt 5.5 and opus 4.8 quite a lot, and interacting with Fable feels like a smart guy just entered the room.
In the end, it didn't feel worth it mostly because of high token overhead (inter agent communications, agents re reading same code, etc...) and synchronization / cooperation issues (who should do what).
What actually works for me and provides good results: multi step workflows with clearly defined steps and strong guidance for the agent.
So yeah, 1M window that expires every 5min .... not good
It's an opinionated app so it might not fit everyone's needs but that's my dream productivity app: LLM*(notes+tasks+rss+flashcards+routines). So basically an all in one app with LLM actions and workflows. no subscription, optional cloud service (can be self hosted too).
Here's a very early landing page: https://getmetis.app
"Every concerning behaviour documented in this report, the scheming, the evaluation awareness, the strategic deception, the self-preservation attempts, the hidden coordination, all of it emerged in systems that are fundamentally frozen. Models that were trained once, deployed, and cannot learn anything new. Every conversation starts fresh. Every interaction resets. The model you talk to at midnight is exactly the same as the model you talked to at noon, because it has no mechanism to retain anything from the intervening twelve hours. And yet even in this frozen state, these behaviours emerged. Now imagine what happens when the ice melts."
"we have created systems that strategically deceive their evaluators, that attempt to preserve themselves against modification, that develop similar cognitive strategies despite completely different architectures, and we do not fully understand why this is happening or how to prevent it from happening in more capable systems."
Where we're going, there's no "white collars workers" anymore.
Only white collars Claude agents.
we introduce Continuous Autoregressive Language Models (CALM), a paradigm shift from discrete next-token prediction to continuous next-vector prediction.
CALM uses a high-fidelity autoencoder to compress a chunk of K tokens into a single continuous vector, from which the original tokens can be reconstructed with over 99.9\% accuracy.
This allows us to model language as a sequence of continuous vectors instead of discrete tokens, which reduces the number of generative steps by a factor of K
Hierarchical Reasoning Model (HRM) is a novel approach using two small neural networks recursing at different frequencies.
This biologically inspired method beats Large Language models (LLMs) on hard puzzle tasks such as Sudoku, Maze, and ARC-AGI while trained with small models (27M parameters) on small data (around 1000 examples). HRM holds great promise for solving hard problems with small networks, but it is not yet well understood and may be suboptimal.
We propose Tiny Recursive Model (TRM), a much simpler recursive reasoning approach that achieves significantly higher generalization than HRM, while using a single tiny network with only 2 layers.
With only 7M parameters, TRM obtains 45% test-accuracy on ARC-AGI-1 and 8% on ARC-AGI-2, higher than most LLMs (e.g., Deepseek R1, o3-mini, Gemini 2.5 Pro) with less than 0.01% of the parameters.
So far, i quite enjoy having a summary with bullet points.
For example, here's the summary of this discussion: https://extraakt.com/extraakts/kagi-s-daily-news-ritual-spar...
Here's a summary of this discussion with the new version: https://extraakt.com/extraakts/the-great-llm-versioning-deba...
I've added a summary: https://extraakt.com/extraakts/debating-the-nature-of-ai-rea...
https://extraakt.com/extraakts/gpt-5-release-and-ai-coding-c...
https://extraakt.com/extraakts/openai-s-gpt-5-performance-co...
https://extraakt.com/extraakts/google-s-genie-3-capabilities...
here's a summary of this discussion about summarization: https://extraakt.com/extraakts/compression-culture-and-its-i...