I have written Flink pipelines that solved similar problems to this post: sessionization of time-skew data points from multiple sources with variable and possibly large delays; simple and economical is not how I would describe it :) .
If one of your datasource get lag behind, Flink would buffer a huge amount of data waiting for it to catch up due to how watermark work when joining 2 streams, and you would still encounter out-of-memory error even with RocksDB from time to time if your session window get too large. In addition, with our state size frequently reached hundred of GBs, recovering from failure was not exactly fast either.