I tried to explain but I fear were probably going to be adding artificial sleeps to these things to convince the masses it's doing something clever.
I tried to explain but I fear were probably going to be adding artificial sleeps to these things to convince the masses it's doing something clever.
Because it's not human and not "thinking", it's a mathematical algorithm
It’s going to be really crazy when the bottle neck for agents is the speed of the tool calls rather than the speed of inference. Imagine an agent interacting with the terminal near instantly…
Developing software becomes 95% about intent and requirements. Can’t wait for the next iteration of that.
The characters in the 3-act Shakespearean play had very little depth, many of the names were similar, and they were not very smart, but the simple plot was cohesive.
But it’s proven they can automate this (they didn’t etch eight billion weights by hand after all, obviously), so now the interesting question is whether they can scale it to more recent aka bigger models.
After all, there’s already very useful models even for productivity at 27 or 35B.
Not sure if modern models "think" only by outputting <thinking> blocks, or there is a more complex mechanism at play.
That's pretty much it - a small refinement to "Chain of Thought" prompting, where you tell the model explicitly in the prompt to "Think step by step" or similar, so it writes out more steps before giving a final answer, potentially catching some errors. The "thinking" models are tuned to do that without being prompted to, and to output the "thinking" markers around it, so they can be hidden from the user.
I am curious what the drop in thoughput is for multi-turn answers, instead of one-shot. More in line with the current "agentic" use-cases.
Probably, the usual initial suspects for “what makes computation slow” will become a focus point that needs to be optimized again: file access, network, etc.