Notation

General

SymbolDescription
N\mathbb{N}Natural numbers, not including zero
Z\mathbb{Z}Integers
R\mathbb{R}Real numbers
C\mathbb{C}Complex numbers
Rd\mathbb{R}^dEuclidean space of dimension dd
RX\mathbb{R}^XSpace of functions from XX to R\mathbb{R}
C(X;R)C(X;\mathbb{R})Space of continuous real-valued functions
[N][N]Finite set of all integers from 11 to NN
⊕\oplusDisjoint union of sets, concatenation of sequences
𝟙(⋅)\text{𝟙}_{(\cdot )}Indicator function
O(⋅)\mathcal{O}(\cdot )Asymptotic upper bound
Ω(⋅)\Omega(\cdot )Asymptotic lower bound
Θ(⋅)\Theta(\cdot )Asymptotic upper and lower bound

Probability

SymbolDescriptionLink
(Ω,F,P⁡)(\Omega,\mathcal{F},\operatorname{\mathbb{P}})Probability spaceA.1
∼\sim Distribution of a random variableA.1
E⁡(⋅)\operatorname*{\mathbb{E}}(\cdot )Expectation of a random variableA.1
Var⁡(⋅)\operatorname{Var}(\cdot )Variance of a random variableA.1
Cov⁡(⋅,⋅)\operatorname{Cov}(\cdot ,\cdot )Covariance between two random variablesA.1
E⁡(⋅∣⋅)\operatorname*{\mathbb{E}}(\cdot \mid\cdot )Conditional expectationA.1
M1(X)\mathcal{M}_1(X)Space of probability measures over XXA.1
M1(Y∣X)\mathcal{M}_1(Y\mid X)Space of probability kernels over YY given XXA.1
δx\delta_xDirac measure centered at xxA.1
N⁡(μ,σ2)\operatorname{N}(\mu,\sigma^2)Normal distributionA.1
U⁡(X)\operatorname{U}(X)Uniform distribution over XXA.1
Ber⁡(p)\operatorname{Ber}(p)Bernoulli distributionA.1
Bin⁡(n,p)\operatorname{Bin}(n,p)Binomial distributionA.1
DKL⁡(⋅∣∣⋅)D_{\operatorname{KL}}(\cdot \mid\mid\cdot )Kullback–Leibler divergenceA.2
I(⋅;⋅)I(\cdot ;\cdot )Mutual informationA.2

Markov Decision Processes

SymbolDescriptionLink
S\mathcal{S}State spaceA.3
A\mathcal{A}Action spaceA.3
rrReward functionA.3
ppTransition kernelA.3
γ\gammaDiscount factorA.3
s0s_0Initial stateA.3
π\piPolicyA.3
Π\PiSpace of policiesA.3
V(π)(⋅)V^{(\pi)}(\cdot )Value function under policy π\piA.3
V∗(⋅)V^*(\cdot )Optimal value functionA.3
π∗\pi^*Optimal policyA.3

Episodic Decision Problems

SymbolDescriptionLink
AAAction space2.1
R\mathcal{R}Reward function class2.1
rrTrue reward function2.1
qqReward distribution in the Bayesian variant2.1
rtr_tSequence of true rewards in the adversarial variant2.1
Σ\SigmaObservation space2.1
σ\sigmaFeedback function2.1
at,σta_t, \sigma_tAction played and feedback observed in episode tt2.1
x1:tx_{1:t}Sequence from x1x_1 to xtx_t2.1
Seq⁡(X)\operatorname{Seq}(X)Space of finite-length sequences over XX2.2
ppAlgorithm for an episodic decision problem2.2
P\mathcal{P}Space of algorithms2.2
RT(p,r)R_T(p,r)Cumulative regret2.3
rT(p,r)r_T(p,r)Simple regret2.3
rc(p,r)r_c(p,r)Cost-adjusted simple regret2.3
RT(p,q)R_T(p,q)Bayesian regret2.3
VMDP⁡∗V^*_{\operatorname{MDP}}Optimal value of the underlying MDP of a Bayesian variant2.3
⪯\preceqBlackwell order over feedback functions2.3

Bayesian Models and Algorithms

SymbolDescriptionLink
pθp_\thetaPrior distribution for θ\theta2.2
py∣θp_{y\mid\theta}Likelihood for yy given θ\theta2.2
pθ∣yp_{\theta\mid y}Posterior distribution for θ\theta given yy2.2
δ\deltaDecision rule2.2
α\alphaAcquisition function2.2