Sheet 1

Prof. Leif Döring, Felix Benning
††course: Wahrscheinlichkeitstheorie 1††semester: FSS 2022††tutorialDate: 21.02.2022††dueDate: 10:15 in the exercise on Monday 21.02.2022
Exercise 1 (Properties of Conditional Expectation).

Let X and Y be two integrable random variables, 𝒢⊆ℱ σ-Algebras.

  1. (i)

    Prove 𝔼⁢[𝔼⁢[X|ℱ]]=𝔼⁢[X] and 𝔼⁢[1|ℱ]=1 a.s.

    Solution.

    Using the second defining property of conditional expectation with the set Ω we have

    𝔼⁢[𝔼⁢[X|ℱ]]=𝔼⁢[𝟙Ω⁢𝔼⁢[X|ℱ]]=𝔼⁢[𝟙Ω⁢X]=𝔼⁢[X].

    For the second claim we simply check the definition of conditional expectation for the constant function 1. Since constants are measureable with regard to any σ-algebra, the first requirement is a given, the second one is trivial. ∎

  2. (ii)

    Prove 𝔼⁢[𝔼⁢[X|ℱ]|𝒢]=𝔼⁢[𝔼⁢[X|𝒢]|ℱ]=𝔼⁢[X|𝒢] a.s.

    Solution.

    Since 𝔼⁢[X|𝒢] is by definition 𝒢 measurable and we have 𝒢⊆ℱ, it is also ℱ measurable. We therefore have

    𝔼⁢[𝔼⁢[X|𝒢]|ℱ]=linearity𝔼⁢[X|𝒢]⁢𝔼⁢[1|ℱ]=𝔼⁢[X|𝒢].

    To show that 𝔼⁢[𝔼⁢[X|ℱ]|𝒢]=𝔼⁢[X|𝒢], we will show that 𝔼⁢[X|𝒢] fulfills the defining requirements. 𝒢 measurability is obvious, what is left to show is that for any A∈𝒢⊆ℱ we have

    𝔼⁢[𝟙A⁢𝔼⁢[X|ℱ]] =A∈ℱ𝔼⁢[𝟙A⁢X]=A∈𝒢𝔼⁢[𝟙A⁢𝔼⁢[X|𝒢]].∎
  3. (iii)

    Assume that X=Y almost surely. Prove that, almost surely, 𝔼⁢[X|Y]=Y and 𝔼⁢[Y|X]=X.

    Solution.

    Due to symmetry it is sufficient to prove 𝔼⁢[X|Y]=Y. To show this we simply check the defining properties of 𝔼⁢[X|Y]. First, Y is σ⁢(Y) measurable by definition, and second, we have for all A∈σ⁢(Y)

    𝔼⁢[X⁢𝟙A] =𝔼⁢[Y⁢𝟙A].∎
Exercise 2.

The following questions are independent.

  1. (i)

    Let X and Y be i.i.d. Bernoulli variables: ℙ(X=1)=1-ℙ(X=0)=p for some p∈(0,1). We set Z≔𝟏{X+Y=0}. Compute 𝔼⁢[X|Z] and 𝔼⁢[Y|Z]. Are these random variables independent?

    Solution.

    First, due to Z=𝟙{X+Y=0}∈{0,1} we can write

    𝔼⁢[X|Z] =𝔼[X|Z=0]⋅𝟙Z=0+𝔼[X|Z=1]⋅𝟙Z=1
    =𝔼⁢[X⋅𝟙Z=0]ℙ(Z=0)⋅𝟙Z=0+𝔼⁢[X⋅𝟙Z=1]ℙ(Z=1)⏟=0⋅𝟙Z=1

    as X=0 for Z=1. So we only need to think about the first term. First, we have

    P(Z=0) =1-ℙ(Z=1)=1-ℙ(X=0)ℙ(Z=0)⏟(1-p)2
    =2⁢p-p2=p⁢(2-p),

    and due to X⋅𝟙Z=0=X we get

    𝔼⁢[X⋅𝟙Z=0]=𝔼⁢[X]=𝔼⁢[𝟙X=1]=p.

    Collecting all three results together, results in

    𝔼⁢[X|Z] =12-p⋅𝟙Z=0.

    Due to symmetry we get the same result for 𝔼⁢[Y|Z] which implies 𝔼⁢[X|Z]=𝔼⁢[Y|Z] a.s. In particular they are not independent! ∎

  2. (ii)

    Let X be a square-integrable random variable and 𝒢⊂ℱ a sub-σ-algebra. We define the conditional variance

    Var⁢(X|𝒢)≔𝔼⁢[(X-𝔼⁢[X|𝒢])2|𝒢].

    Prove the following identity:

    Var⁢(X)=𝔼⁢[Var⁢(X|𝒢)]+Var⁢(𝔼⁢[X|𝒢]).
    Solution.

    Recall that we have Var⁢(X)=𝔼⁢[X2]-𝔼⁢[X]2, the same holds true for the conditional variance:

    Var⁢(X|𝒢) =𝔼⁢[X2-2⁢X⁢𝔼⁢[X|𝒢]+𝔼⁢[X|𝒢]2|𝒢]
    =𝔼⁢[X2|𝒢]-2⁢𝔼⁢[X⁢𝔼⁢[X|𝒢]|𝒢]+𝔼⁢[𝔼⁢[X|𝒢]2|𝒢]
    =𝔼⁢[X2|𝒢]-𝔼⁢[X|𝒢]2. (1)

    With those identities we can directly prove our claim by calculation

    𝔼⁢[Var⁢(X|𝒢)]+Var⁢(𝔼⁢[X|𝒢]) =𝔼⁢[Var⁢(X|𝒢)]+𝔼⁢[𝔼⁢[X|𝒢]2]-𝔼⁢[𝔼⁢[X|𝒢]]2
    =𝔼⁢[Var⁢(X|𝒢)+𝔼⁢[X|𝒢]2]-𝔼⁢[𝔼⁢[X|𝒢]]2
    =(1)𝔼⁢[𝔼⁢[X2|𝒢]]-𝔼⁢[𝔼⁢[X|𝒢]]2
    =𝔼⁢[X2]-𝔼⁢[X]2
    =Var⁢(X)∎
Exercise 3 (Factorization lemma).

(Important!) Let X,Y:(Ω,𝒜)→(ℝ,ℬ⁢(ℝ)) be random variables. Show that, if X is (σ⁢(Y),ℬ⁢(ℝ))-measurable, then there exists a measurable function g:(ℝ,ℬ⁢(ℝ))→(ℝ,ℬ⁢(ℝ)), such that X=g⁢(Y).

Solution.

First we consider simple functions, i.e.

X=∑k=1nak⁢𝟙Ak

for ak≥0 and Ak∈𝒜 for all k=1,…,n. Without loss of generality we can assume the Ak to be disjoint and ak≠an for k≠n. This, together with the fact that X is supposed to be σ⁢(Y) measurable, implies there exists some Bk∈ℬ⁢(ℝ) s.t. Ak∈Y-1⁢(Bk) for all k, because we have

Ak=disjointX-1⁢({ak})∈σ⁢(Y)={Y-1⁢(B):B∈ℬ⁢(ℝ)}.

Hence we have

X=∑k=1nak⁢(𝟙Bk∘Y)⏟𝟙Ak=(∑k=1nak⋅𝟙Bk⏟≔g)∘Y,

where g is measurable.

Now we just assume X≥0 is (σ⁢(Y),ℬ⁢(ℝ))-measurable, which implies there exists a sequence (Xn)n∈ℕ of (σ⁢(Y),ℬ⁢(ℝ))-measurable simple functions with Xn↑X. Taking intersections of the indicator sets of Xn+1 and Xn implies that Δ⁢Xn=Xn-Xn-1 is a simple function too, which can be written as Δ⁢Xn=gn⁢(Y). Since Xn converges we have

X=limn→∞⁡Xn=limn→∞⁡∑k=1nΔ⁢Xn=∑k=1∞gk⁢(Y)=∑k=1∞gk⏟=⁣:g∘Y

using X0=0 without loss of generality. And g is measurable as a limit of measurable functions. Note that g is well defined as ∑k=1ngk with gk≥0 is monotonously increasing although it might be infinite. This is why we used the telescoping trick. Splitting X into X=X+-X- yields the general case. ∎

Exercise 4 (Best Estimators).

Let X be a real random variable.

  1. (i)

    Find the estimator m minimizing 𝔼⁢[𝟙X≠m] when X is discrete. What is the best estimator for a dice roll using this loss?

    Solution.

    The best estimator is the maximum likelihood estimator maxmP(X=m), as

    minm𝔼[𝟙X≠m]=minmℙ(X≠m)=minm[1-P(X=m)]=1-maxmP(X=m).

    So for a fair dice, any of its faces {1,…,6} is a minimizer. ∎

We define the median of X to be 𝕄⁢[X]:=inf⁡{m∈ℝ:f⁢(m)>0}, where

f:{ℝ→[-1,1]m↦ℙ(X≤m)-ℙ(X>m)
  1. (ii)

    (Unimportant) Prove that ℙ(X≤𝕄[X])≥12 and ℙ(X≥𝕄[X])≥12, and show that for all m>𝕄⁢[X] we do not have ℙ(X≤m)≥12.

    Hint.

    Calculate the limits limm↑𝕄⁢[X]⁡f⁢(m) and limm↓𝕄⁢[X]⁡f⁢(m) using the continuity of measures. Also note that f⁢(m)=2⁢FX⁢(m)-1, where FX is the cumulative distribution function of X.

    Solution.

    As f is right-continuous, since FX(m)=ℙ(X≤m) is right-continuous, we have

    2ℙ(X≤𝕄[X])-1=f(𝕄[X])=limm↓𝕄⁢[X]f(m)≥0

    and therefore ℙ(X≤𝕄[X])≥12. For the second claim we need to use continuity of measures directly

    2ℙ(X<𝕄[X])-1 =ℙ(X<𝕄[X])-ℙ(X≥𝕄[X])
    =ℙ(⋃m<𝕄⁢[X]{X≤m})-ℙ(⋂m<𝕄⁢[X]{X>m})
    =limm↑𝕄⁢[X]ℙ(X≤m)-ℙ(X>m)
    =limm↑𝕄⁢[X]⁡f⁢(m)≤0

    Therefore we have ℙ(X<𝕄[X])≤12, which implies that ℙ(X≥𝕄[X])≥12.

    Lastly, for any m>𝕄⁢[X], we can find a small ϵ>0 with m-ϵ>𝕄⁢[X]. Therefore

    2ℙ(X<m)-1≥2ℙ(X≤m-ϵ)-1=f(m-ϵ)>0,

    which implies ℙ(X<m)>12 and therefore P(X≤m)<12. ∎

    Remark.

    This property is usually the defining property of the median. If this property does not define the median uniquely, 𝕄⁢[X] is the largest median. A lower bound can be found similarly.

  2. (iii)

    Prove that the median minimizes the L1 error, i.e. 𝕄⁢[X]=arg⁡minm⁡𝔼⁢[|X-m|]. Whenever useful, you can assume continuity.

    Hint.

    Recall that for X≥0 a.s., we have 𝔼[X]=∫0∞ℙ(X≥t)dt=∫0∞ℙ(X>t)dt. Use this fact to prove

    𝔼[(X-m)+]=∫m∞ℙ(X>m)dt.
    Solution.

    Recall that for X≥0 a.s., we have 𝔼[X]=∫0∞ℙ(X≥t)dt=∫0∞ℙ(X>t)dt. Applying this fact to (X-m)+ we get

    𝔼⁢[(X-m)+] =∫0∞ℙ((X-m)+>t)⏟=ℙ(X-m>t)=ℙ(X>t-m)⁢𝑑t
    =∫m∞ℙ(X>s)ds

    And similarly

    𝔼⁢[(X-m)-] =∫0∞ℙ((X-m)-≥t)⏟=ℙ(m-X≥t) (∀t>0)=ℙ(m-t≥x)⁢𝑑t
    =∫-∞mℙ(s≥X)ds

    Put together we have

    𝔼⁢[|X-m|] =∫m∞ℙ(X>s)ds+∫-∞mℙ(X≤s)ds.

    Assuming ℙ(X>s) and ℙ(X≤x) are both continuous in s, we can use the fundamental theorem of calculus to obtain

    dd⁢m𝔼[|X-m|]=ℙ(X≤m)-ℙ(X>m)=f(m).

    Therefore we have

    𝔼⁢[|X-m|]-𝔼⁢[|X-𝕄⁢[X]|]=∫𝕄⁢[X]mf⁢(t)⁢𝑑t≥0

    as f is greater than zero for any m≥𝕄⁢[X] and less than zero for any m≤𝕄⁢[X] and sorting the integral borders results in another sign flip. Notice that f is strictly greater than zero for m≥𝕄⁢[X] which implies that 𝕄⁢[X] is the largest median, i.e. all numbers greater than 𝕄⁢[X] do not minimize 𝔼⁢[|X-m|].

    As mentioned one can similarly define a lower bound, and all numbers in-between fulfill the defining property of medians of ℙ(X≥m)≥12 and ℙ(X≤m)≥12. ∎

    Remark.

    In the non-continuous case, one still has right-continuity of f and can prove right-differentiability of 𝔼⁢[|X-m|]. This might be sufficient to prove it is monotonously increasing after 𝕄⁢[X] and similarly monotonously decreasing before. But the proof is likely going to be complex or require a strong analysis foundation. Alternatively, a short and basic (but unintuitive) proof of the general case can be found here: https://math.stackexchange.com/a/2790390/445105.