>On a side note it feels like each command takes longer to process than the previous - almost like it is re-doing everything for each command (and that is how it keeps state).
That's because it's probably redoing everything.
But that's probably to keep the implementation simple. They are probably just appending the new input and re-running the whole network.
The typical data dependency structure in a transformer architecture is the following :
outputt0 outputt1 outputt2 outputt3 | outputt4
featL4t0 featL4t1 featL4t2 featL4t3 | featL4t4
featL3t0 featL3t1 featL3t2 featL3t3 | featL3t4
featL2t0 featL2t1 featL2t2 featL2t3 | featL2t4
featL1t0 featL1t1 featL1t2 featL1t3 | featL1t4
input_t0 input_t1 input_t2 input_t3 | input_t4
The features at layer Li at time tj only depends on the features of the layer L(i-1) at times t<=tj.
If you append some new input at the next time t4 and recompute everything from scratch it doesn't change any feature values for time < t4.
To compute the features and output at time t4 you need all the values of the previous times for all layers.
The alternative to recomputing would be preserving the previously generated features, and incrementally building the last chunk by stitching it to the previous features. If you have your AI assistant running locally that something you can do, but when you are serving plenty of different sessions, you will quickly run out of memory.
With simple transformers, the time horizon of the transformer used to be limited because the attention of the transformer was scaling quadratically (in compute), but they are probably using an attention that scale in O(n*log(n)) something like the Reformer, which allows them to handle very long sequence for cheap, and probably explain the boost in performance compared to previous GPTs.