3,600 karma · joined May 28, 2020
"showClearContextOnPlanAccept": true
Boris rationale was "with 1m context window, most users don't need it anymore."What I'd love to see is a harness that deeply optimises for the best results obtainable out of non-frontier models. Many of these have 1M context windows, and most of it remains unused and under utilised in these harnesses, in my opinion.
Haha, mine too! I guess it was to calculate the fuel efficiency, which is often done automatically now.
Inference time is going to be dominated by the low memory bandwidth on these Macs, so a dense model will suffer most. It’s more of an opportunity for large MoE models with a low number of active experts since you can keep all experts in VRAM but not pay the bandwidth cost until they are used.
> you can leverage two different bit rates depending on the actual need for constant access to common layers versus sparse access to infrequent layers
This is an interesting direction that I expect to see more of. But for most models currently you need basically all experts loaded since they are chosen per token.
Apple seems to be researching longer horizon expert caching, where they keep experts swapped in for longer runs of tokens [1]. Other labs are offloading ngram caches but not sure if they’re pursuing anything like this?
1. https://machinelearning.apple.com/research/introducing-third...
Not if crafty claude finds a way to overflow vllm or something. “Hmm. Maybe i’ll return an unterminated thinking block with these special tokens and fill my cache up in exactly this pattern and…”
https://news.ycombinator.com/item?id=49424387&utm_source=cha...
I would expect the agent loop and system prompt to be basically the same. Is it the precise semantics of the tools (and how closely they match what a particular agent was trained on) or something else?
I also wish the discussion around Pi did not always use cost/token count as the metric. It's amazingly token efficient, but how does it stack up again opencode and others if you don't care about token count?
My experience is that the harness is mainly polish preventing failed tool calls, bad edits, stuff like that, but doesn't make much difference to the overall "intelligence". But that opencode seems slightly more robust against stupid errors than out of the box Pi due to the additional context it forces through every thread.
> Instead of forcing the entire model into DRAM, the full model is stored in flash memory (NAND). Because NAND-to-DRAM bandwidth is too slow to swap weights token by token, as standard MoE models require, AFM 3 Core Advanced makes routing decisions per prompt. A lightweight, dense block selects a fixed set of experts during initial processing, periodically reselecting them during generation. To minimize data movement, the model relies on a high percentage of always-active “shared experts” alongside input-dependent “routed experts” swapped into DRAM only when needed.
This is an interesting hybrid between MoE and managing entirely separate domain-specific models. Select the experts once, bring them into memory, and run inference for some period of time before re-evaluating. Saves having all experts in memory, but it's better than just selecting a whole model per query since you have a high number of small opaque experts that overlap and combine in interesting ways.
There is a probably a massive quality hit to doing this but it's interesting because it allows infinite scaling of model size.
I wonder what tricks one could use to ensure the model is actually performing on par with the reference api, like matching seeds or running exact benchmarks.
> I think I’m suspecting something is going “wrong” in the training process. The model is greatly rewarded for succeeding on long-horizon tasks, but presumably there is very little punishing going on for “shitty code.”
My suspicion is that both OpenAI and Anthropic moved their RL agendas from "being rated as useful according to human feedback" to "succeeds at long horizon tasks" in the last few months, resulting in agents that are closer to AGI in an autonomous task-completing sense, but strangely bad at communicating.
The result is that they are amazingly good at long horizon tasks, computer use, solving difficult math/ARC-AGI type problems, but becoming weirder and weirder to work with.
How are you running jobs unattended 24/7 without hitting your token limits?
In general, our economy relies a lot on these natural filters for effort. A company with a nice website doesn't necessarily mean the company is good, except historically it kind of does, because it's a proxy for effort (and budget), both of which correlate roughly with reputability and a good product.
The same applies to well-written blog post: typically, a very good writer generally is also someone with something interesting to say, even though in theory the two things don't have to be connected.
Similarly, the entire moat of many companies is that switching to a competitor is a lot of effort.
AI is effectively sending the effort required to do any of these things to zero, which means all our existing implicit filters are going to stop working.