This seems to not fit the criteria anymore (not tabular, not Markov). Is it related to structured bandits?
This seems to not fit the criteria anymore (not tabular, not Markov). Is it related to structured bandits?
They'll explore the arm with the highest potential payoff i.e. the highest upper-confidence-bound, which often is the arm you know the least about since they're roughly calculated as (average + confidence interval) and early on the confidence bound is large. This style of algorithm means you can add arms as you go through the experiment and they'll be explored/exploited in a reasonable way.
What you're describing sounds more like you're exploring a frontier and "discovering" new options along the way..?
Variations of this scheme include exploit vs. copy vs. innovate used in computational biology. Learning agent can copy what others do or innovate and try something new.