Explore or Exploit, explained.
The multi-armed bandit problem describes the tradeoff between exploration, which discovers better options, and exploitation, which uses the best option known so far.
Why it happens
Early results are noisy. An option that wins once may be weak, while a promising option may initially lose. Exploring buys information at the cost of occasionally choosing a lower-paying option. The value of that information depends on how many decisions remain.
A bandit problem asks how to balance immediate reward and learning from experiments.
Read the result
Change exploration and compare cumulative reward over the same horizon. The oracle benchmark knows the best machine in advance; it shows the cost of learning, not an achievable strategy for someone without that knowledge.
A worked example
Choosing a newsletter subject line
Three subject lines have unknown response rates. You can test them over repeated sends.
Send some messages using alternatives while giving most traffic to the current leader.
The tests can discover a better line, but their cost matters more when only a few sends remain.
OPTIONAL DEEPER DETAILGo deeper: inside the model
Inside this model
Three seeded machines have fixed hidden win chances. The comparison agent uses epsilon-greedy choice over sample means, beginning with one trial of each. The oracle always chooses the best hidden chance. The plot shows cumulative reward over 60 rounds.
Where this idea is useful
A practical use
Trying a new page design while still showing the strongest current design.
A common misconception
“The first winner deserves all future choices.”
A small sample can misidentify the best option. Continued exploration helps correct that mistake, especially over a long horizon.
What this explanation leaves out
- Independent stationary rewards and a single epsilon rule omit changing audiences, costs and delayed outcomes.
Does this experiment use Thompson sampling?
The automated comparison uses an exploration-based policy, not Thompson sampling. Several algorithms solve bandit problems using different ways to represent uncertainty.
How much future opportunity remains to benefit from what you learn today?
Associated thinkers
Further reading
Explore the original research or the teaching reference behind this experiment.