I think SSD offload will become more feasible an the engram approach matures.
For now though, I would question if the difference in smarts is large enough to justify tolerating a generation speed measured in seconds per token compared to spending more turns refining a plan with a flash model.