[1] https://en.wikipedia.org/wiki/Exploration-exploitation_dilem...
[1] https://en.wikipedia.org/wiki/Exploration-exploitation_dilem...
I remember learning something akin to this in social psychology in the context of a single risk-taker fish breaking from the school of fish to explore and take risks which ultimately benefitted the whole group.
The general idea is that your agent, however it is training, has to balance trying new things to possibly find the global maxima instead of getting hooked on a rewarding local maxima.
https://en.wikipedia.org/wiki/Reinforcement_learning#Explora...
here an n-gram of the first mentions: https://books.google.com/ngrams/graph?content=Exploration-ex...
I didn’t research machine learning, but my friends did, so it could have been an idea I picked up from someone else.
Anyway, I think it’s a valuable framework because we need to make time for exploration—it’s a great way to let go, have fun, and dissolve the fear of failure.
Also, the typical framing of the problem is the same "kind" of choice being repeatedly executed (e.g., betting on a coin-flip of unknown bias, or balancing the gain of consumer purchasing information vs exploiting known information when setting items into aisle end-caps in a grocery store). That has a lot more structure than arbitrary graphs, enough so to make it worthy of its own dedicated study (especially given the real-world applicability).