Skip to content
Sarthak Bagaria
All notes

Chapter 6 Numeraires

We change the unit of account, which is the most useful single manoeuvre in derivative pricing and the least intuitive. Girsanov’s theorem says a measure change moves drifts and cannot touch volatilities, and that asymmetry is what makes the technique work: choosing a numeraire chooses which quantity is a martingale, and the right choice makes a hard expectation trivial. We derive the forward and annuity measures that chapter 7 and chapter 15 need, and find that no-arbitrage assigns one price to each source of risk — the same projection that reappears as a hedge in chapter 23.

6.1 Change of Measure

Definition 6.1 (Radon-Nikodym Derivative).

Consider two equivalent probability measures ℙ and ℙ^ on a measurable space (Ω,Σ). The Radon-Nikodym derivative d⁢ℙ^/d⁢ℙ:Ω→ℝ is defined such that for any subset A,Ω⫆A∈Σ,

∫A𝑑ℙ^=∫A(d⁢ℙ^/d⁢ℙ)⁢𝑑ℙ. (6.1)

Suppose (Ω,ℱ,ℱt) is a filtered probability space, then note from the above definition we have

(d⁢ℙ^/d⁢ℙ)|ℱt=𝔼ℙ⁢[(d⁢ℙ^/d⁢ℙ)|ℱt] (6.2)

(d⁢ℙ^/d⁢ℙ)|ℱt is also written as (d⁢ℙ^/d⁢ℙ)t and is thus a martingale stochastic process (by iterated conditioning) in ℙ.

Example 6.1.

A die roll makes the Radon-Nikodym derivative concrete: multiple probability distributions can be assigned to the same six outcomes.

ω ℙ ℙ^ d⁢ℙ^/d⁢ℙ
1 1/6 1/2 3
2 1/6 1/4 3/2
3 1/6 1/8 3/4
4 1/6 1/16 3/8
5 1/6 1/32 3/16
6 1/6 1/32 3/16

The probability in ℙ^ of getting an odd number is ∫ω∈{1,3,5}𝑑ℙ^=1/2+1/8+1/32=21/32=3∗1/6+3/4∗1/6+3/16∗1/6=∫ω∈{1,3,5}(d⁢ℙ^/d⁢ℙ)⁢𝑑ℙ

Theorem 6.2 (Abstract Bayes’ Theorem).

Let ℙ and ℙ^ be two measures on measurable space (Ω,ℱ). Let 𝒢⊂ℱ be another sigma algebra on Ω. Then for any A∈𝒢 and random variable X

𝔼ℙ^⁢[X|𝒢]=𝔼ℙ⁢[(d⁢ℙ^/d⁢ℙ)⁢X|𝒢]𝔼ℙ⁢[(d⁢ℙ^/d⁢ℙ)|𝒢]. (6.3)
Proof.

We show that for any A∈𝒢

𝔼ℙ^⁢[X|𝒢]⁢𝔼ℙ⁢[(d⁢ℙ^/d⁢ℙ)|𝒢]=𝔼ℙ⁢[(d⁢ℙ^/d⁢ℙ)⁢X|𝒢]

Since the random variables involved are constant over A, we can check equality on integrals over A.

∫A𝔼ℙ^⁢[X|𝒢]⁢𝔼ℙ⁢[(d⁢ℙ^/d⁢ℙ)|𝒢]⁢𝑑ℙ =∫A𝔼ℙ⁢[(d⁢ℙ^/d⁢ℙ)⁢𝔼ℙ^⁢[X|𝒢]|𝒢]⁢𝑑ℙ (𝔼ℙ^⁢[X|𝒢] is 𝒢 measurable)
=∫A(d⁢ℙ^/d⁢ℙ)⁢𝔼ℙ^⁢[X|𝒢]⁢𝑑ℙ (conditional expectation)
=∫A𝔼ℙ^⁢[X|𝒢]⁢𝑑ℙ^ (Radon-Nikodym)
=∫AX⁢𝑑ℙ^ (conditional expectation)
=∫A(d⁢ℙ^/d⁢ℙ)⁢X⁢𝑑ℙ (Radon-Nikodym)
=∫A𝔼ℙ⁢[(d⁢ℙ^/d⁢ℙ)⁢X|𝒢]⁢𝑑ℙ.

The last step is the definition of conditional expectation. ∎

Remark.

Taking X=VT where Vs is ℱs adapted and G=ℱt for t<T in the above theorem, we get the very useful formula for measure change for conditional expectations on filtered spaces,

𝔼ℙ^⁢[VT|ℱt]=𝔼ℙ⁢[(d⁢ℙ^/d⁢ℙ)T(d⁢ℙ^/d⁢ℙ)t⁢VT|ℱt] (6.4)
Proof.
𝔼ℙ^⁢[VT|ℱt] =𝔼ℙ⁢[(d⁢ℙ^/d⁢ℙ)⁢VT|ℱt]𝔼ℙ⁢[(d⁢ℙ^/d⁢ℙ)|ℱt] (abstract Bayes’ theorem)
=𝔼ℙ⁢[𝔼ℙ⁢[(d⁢ℙ^/d⁢ℙ)⁢VT|ℱT]|ℱt]𝔼ℙ⁢[(d⁢ℙ^/d⁢ℙ)|ℱt] (iterated conditioning)
=𝔼ℙ⁢[𝔼ℙ⁢[(d⁢ℙ^/d⁢ℙ)|ℱT]⁢VT|ℱt]𝔼ℙ⁢[(d⁢ℙ^/d⁢ℙ)|ℱt] (VT is ℱT measurable)
=𝔼ℙ⁢[(d⁢ℙ^/d⁢ℙ)T⁢VT|ℱt](d⁢ℙ^/d⁢ℙ)t (martingale property of Radon Nikodym derivative)
=𝔼ℙ⁢[(d⁢ℙ^/d⁢ℙ)T(d⁢ℙ^/d⁢ℙ)t⁢VT|ℱt] ((d⁢ℙ^/d⁢ℙ)t is ℱt measurable).

∎

We consider processes until terminal time S i.e. ℱ=ℱS and (d⁢ℙ^/ℙ)=(d⁢ℙ^/ℙ)S. We take a strictly positive martingale process to be the Radon-Nikodym derivative:

d⁢ft=ft⁢σ⁢(t)⁢d⁢Wt;ft=e−∫0t12⁢σ2⁢(s)⁢𝑑s+∫0tσ⁢(s)⁢𝑑Ws. (6.5)
Theorem 6.3 (Girsanov Theorem).

If W1,t and W2,t are (possibly correlated) Brownian motions in ℙ and the Radon-Nikodym derivative (d⁢ℙ^/d⁢ℙ)t=ft is given by ft=e−∫0t12⁢σ2⁢(s)⁢𝑑s+∫0tσ⁢(s)⁢𝑑W2,sℙ then W1,t−∫0tσ⁢(s)⁢𝑑W1,s⁢𝑑W2,s is a Brownian motion in ℙ^, where d⁢W1,s⁢d⁢W2,s is the covariation rate of W1 and W2 (zero if they are independent, ρ⁢d⁢s if correlated at ρ, and d⁢s in the case W1=W2).

Proof.

We show that Xt=W1,t−∫0tσ⁢(s)⁢𝑑W1,s⁢𝑑W2,s follows normal distribution with mean 0 and variance t in ℙ^. Other required properties can be verified easily. We show that the moment generating function of Xt is same as that of normal distribution with mean 0 and variance t.

𝔼ℙ^⁢[e−y⁢Xt] =𝔼ℙ⁢[(d⁢ℙ^/d⁢ℙ)t⁢e−y⁢Xt]
=𝔼ℙ⁢[e−∫0t12⁢σ2⁢(s)⁢𝑑s+∫0tσ⁢(s)⁢𝑑W2,s⁢e−y⁢Xt]
=e−∫0t12⁢σ2⁢(s)⁢𝑑s⁢𝔼ℙ⁢[e∫0tσ⁢(s)⁢𝑑W2,s−y⁢Xt]
=e−∫0t12⁢σ2⁢(s)⁢𝑑s⁢𝔼ℙ⁢[e∫0tσ⁢(s)⁢𝑑W2,s−y⁢(W1,t−∫0tσ⁢(s)⁢𝑑W1,s⁢𝑑W2,s)]
=e−∫0t12⁢σ2⁢(s)⁢𝑑s⁢𝔼ℙ⁢[e∫0t(σ⁢(s)⁢d⁢W2,s−y⁢d⁢W1,s+y⁢σ⁢(s)⁢d⁢W1,s⁢d⁢W2,s)].

Define the martingale Zy,t=∫0t(σ⁢(s)⁢d⁢W2,s−y⁢d⁢W1,s). Then eZy,t−12⁢∫0t𝑑Zy,s⁢𝑑Zy,s is a martingale as well, with

d⁢Zy,s⁢d⁢Zy,s=(σ2⁢(s)+y2)⁢d⁢s−2⁢y⁢σ⁢d⁢W1,s⁢d⁢W2,s.

We then have

𝔼ℙ^⁢[e−y⁢Xt] =e−∫0t12⁢σ2⁢(s)⁢𝑑s⁢𝔼ℙ⁢[eZy,t−12⁢∫0t𝑑Zy,s⁢𝑑Zy,s+12⁢∫0t(σ2⁢(s)+y2)⁢𝑑s]
=e12⁢y2⁢t⁢𝔼ℙ⁢[eZy,t−12⁢∫0t𝑑Zy,s⁢𝑑Zy,s]
=e12⁢y2⁢t⁢(eZy,t−12⁢∫0t𝑑Zy,s⁢𝑑Zy,s)|t=0
=e12⁢y2⁢t.

∎

Structure (A measure change moves the drift and cannot touch the volatility).

The theorem’s content is an asymmetry: it changes the drift of a process and leaves its diffusion coefficient exactly where it was. That asymmetry explains several things at once, and stating it precisely matters, because the loose version is misleading in a way that bites later.

Quadratic variation is computed pathwise: ⟨M⟩t is a limit of sums of squared increments of the trajectory that actually occurred, so it is a random variable and not an expectation. Different paths have different quadratic variations — in a stochastic volatility model ⟨S⟩t=∫0tvs⁢Ss2⁢𝑑s, and a path that spent its life in a high-variance regime has a larger one than a path that did not. What is measure-invariant is therefore not a number but a function of the path: the map from trajectory to accumulated variance mentions no measure anywhere, so two equivalent measures assign the same quadratic variation to the same path. They disagree about which paths are likely, not about what each path’s variance is. Equivalence is doing exactly one job in that sentence: the limit converges in probability, so ⟨M⟩ is defined only up to null sets, and equivalent measures are precisely those that agree on which sets those are.

That is why volatility can be estimated from one realisation and drift cannot. An observer sees a single trajectory; the sums of squared increments of that trajectory converge without reference to any measure, so no averaging over paths that did not happen is required. A statement about drift is the opposite kind of statement — it is a claim about the average over the paths one did not see — so it needs either many independent histories, of which there is one, or a long span of time. Chapter 21 measures the consequence: over a fixed window, sampling faster drives the volatility error to zero and leaves the drift error exactly where it was. It is also why chapter 22 can measure realised volatility and must argue about expected returns, and why every model in these notes is calibrated on volatilities and never on drifts.

It is also why the risk neutral measure exists at all. Pricing requires the drift to be one particular thing, the measure change is free to set it, and nothing about the observable roughness of the path has to be disturbed to do so.

One consequence deserves drawing out, because collapsing it is the commonest way to misread the paragraph above. The quadratic variation of a given path is measure-free; the distribution of quadratic variation across paths is not. So

𝔼ℙ⁢[⟨S⟩T]≠𝔼ℚ⁢[⟨S⟩T]

is perfectly consistent with everything just said, and it is the variance risk premium that chapter 21 measures — the reason implied volatility exceeds realised on average. Girsanov leaves σ alone as a coefficient and does not leave the law of accumulated variance alone at all, because in a stochastic volatility model it changes the drift of v itself. The two statements to keep apart are that the realised variance of the path one saw is a fact, and that the variance one expects is a choice of measure.

Remark (What a change of measure does, and what it does not do).

The algebra above is dense enough to lose the idea in, and the idea is simple. A change of measure does not move anything. Every path the world could take is still available afterwards, ending where it always would have. What changes is how much each one counts.

The die in the example at the start of this chapter is the honest picture of it: the same six faces, reweighted. Girsanov is that operation performed on a continuum of paths rather than six outcomes, and the Radon-Nikodym derivative is the column of weights.

For the geometric Brownian motion of chapter 5 the weights can be written down. If the asset drifts at μ in the real world and at r under the risk neutral measure, the change of measure is governed by

θ=μ−rσ, (6.6)

the excess return per unit of risk, called the market price of risk. A path ending at ST carries the weight

d⁢ℚd⁢ℙ=e−θ⁢WT−12⁢θ2⁢T, (6.7)

which decreases in WT and therefore in ST: the risk neutral measure counts the good outcomes for less.

Reweighting the paths this way is the same as shifting the noise. Girsanov applied to (6.7) makes Wtℚ=Wt+θ⁢t a Brownian motion under ℚ, and substituting d⁢W=d⁢Wℚ−θ⁢d⁢t into the real-world dynamics,

d⁢St=μ⁢St⁢d⁢t+σ⁢St⁢d⁢Wt=(μ−σ⁢θ)⁢St⁢d⁢t+σ⁢St⁢d⁢Wtℚ=r⁢St⁢d⁢t+σ⁢St⁢d⁢Wtℚ,

using σ⁢θ=μ−r. The drift falls from μ to r; the diffusion coefficient σ⁢St is the same on every line. That is the whole content of the phrase “removing the risk premium”, and (6.6) says exactly how much removing costs.

Figure 6.1: A change of measure, performed rather than described. The dotted sample is drawn under the real world measure at μ=12% and never redrawn; each draw is then multiplied by the weight (6.7). It lands on the risk neutral density at r=3%, which the sampling never saw. The second panel is the weight itself, tilting down across the outcomes. Drag μ onto the interest rate and the weight flattens onto one, the two densities merge, and the sample follows them — with no risk premium there is nothing to reweight, and the two measures coincide.
Show the model behind this figure (2 functions)
/// The Radon-Nikodym derivative `dQ/dP`, as a function of the terminal price.
///
/// Every path ending at `x` carries this weight. It is a decreasing function of
/// `x` whenever the asset earns a risk premium: the risk-neutral measure counts
/// the good outcomes for less, which is the whole of what "removing the risk
/// premium" means once it is written down.
pub fn radon_nikodym(x: f64, spot: f64, mu: f64, r: f64, sigma: f64, t: f64) -> f64 {
    if x <= 0.0 || sigma <= 0.0 || t <= 0.0 {
        return 1.0;
    }
    let theta = market_price_of_risk(mu, r, sigma);
    // The standard normal draw that produced this terminal price under P.
    let z = ((x / spot).ln() - (mu - 0.5 * sigma * sigma) * t) / (sigma * t.sqrt());
    (-theta * t.sqrt() * z - 0.5 * theta * theta * t).exp()
}
reweighted_histogramquant/src/measure.rs
/// The reweighted histogram of a sample drawn under the real-world measure.
///
/// `grid` supplies the bin centres. Returns a density: the total weight landing
/// in each bin, divided by the bin width and by the total weight of the sample.
///
/// Nothing here knows the risk-neutral density. The sample is drawn once under
/// `mu`, and the only risk-neutral quantity used is the weight. The numeraires
/// chapter's claim is that this reproduces the risk-neutral density anyway, and
/// [`tests::reweighting_a_real_world_sample_gives_the_risk_neutral_density`]
/// checks it.
pub fn reweighted_histogram(
    grid: &[f64],
    spot: f64,
    mu: f64,
    r: f64,
    sigma: f64,
    t: f64,
    paths: usize,
    seed: u64,
) -> Vec<f64> {
    let n = grid.len();
    if n < 2 {
        return vec![0.0; n];
    }
    let width = (grid[n - 1] - grid[0]) / (n - 1) as f64;
    let (lo, hi) = (grid[0] - 0.5 * width, grid[n - 1] + 0.5 * width);

    let mut bins = vec![0.0; n];
    let mut total = 0.0;
    let mut rng = Rng::new(seed);

    for _ in 0..paths {
        // One draw under the real-world measure. This is the sample, and it is
        // never redrawn: the risk-neutral column below comes from reweighting
        // these same numbers.
        let z = rng.next_normal();
        let x = spot * ((mu - 0.5 * sigma * sigma) * t + sigma * t.sqrt() * z).exp();
        let weight = radon_nikodym(x, spot, mu, r, sigma, t);
        total += weight;
        if x >= lo && x < hi {
            let bin = (((x - lo) / width) as usize).min(n - 1);
            bins[bin] += weight;
        }
    }

    if total <= 0.0 {
        return bins;
    }
    for b in &mut bins {
        *b /= total * width;
    }
    bins
}
Structure (One price per risk, and no-arbitrage is what sets it).

The scalar θ=(μ−r)/σ above is a ratio that is easy to read as bookkeeping. With more than one asset it stops being bookkeeping and becomes the fundamental theorem in coordinates.

Take n assets driven by d Brownian motions, with σ the matrix of loadings and μ−r⁢𝟏 the vector of excess returns. Girsanov shifts each Brownian motion by some θj, and asset i’s drift falls by its own exposure to those shifts. Demanding that every asset end up drifting at r under one measure is therefore the linear system

μ−r⁢𝟏=σ⁢θ. (6.8)

Read the indices. There is one θj per Brownian motion and not one per asset: a price of risk belongs to the risk, and each asset’s excess return is forced to be its own loadings against those common prices. Two assets exposed to the same factor must reward that exposure at the same rate, whatever else differs between them.

Which is a statement no-arbitrage can enforce, and does. If two assets on one Brownian motion offered different Sharpe ratios, sell the one paying less per unit of risk, buy the one paying more, in the ratio that cancels the exposure — the position has no randomness left and a positive drift. Solvability of (6.8) is exactly the absence of that trade, which is chapter 4’s fundamental theorem written with matrices instead of measures:

Statement about (6.8) Statement about measures
solvable a risk neutral measure exists, so no arbitrage
uniquely solvable it is unique, so the market is complete
unsolvable an arbitrage, and the residual is the portfolio

The last row is constructive rather than merely a diagnosis. Let θ^ minimise ‖(μ−r⁢𝟏)−σ⁢θ‖2 and write

ρ=(μ−r⁢𝟏)−σ⁢θ^

for the residual — the part of the excess-return vector orthogonal to the column space of σ, so σ⊤⁢ρ=0. Hold ρ as portfolio weights. Its loading on Brownian motion j is ∑iρi⁢σi⁢j=(σ⊤⁢ρ)j=0, so the position carries no factor risk; its excess return is ∑iρi⁢(μi−r)=ρ⊤⁢(μ−r⁢𝟏), and since μ−r⁢𝟏=σ⁢θ^+ρ with σ⊤⁢ρ=0,

ρ⊤⁢(μ−r⁢𝟏)=ρ⊤⁢σ⁢θ^+ρ⊤⁢ρ=‖ρ‖2.

So ρ is a riskless portfolio earning ‖ρ‖2 over the funding rate — an arbitrage exactly when ρ≠0, which is exactly when (6.8) has no solution. The failure hands over the trade. factor_risk_prices builds it and checks both halves: the factor exposure is zero and the profit is ‖ρ‖2.

In a real market ρ is not riskless — it still carries whatever the factors do not span — so a non-zero residual is a statistical arbitrage rather than an actual one: a factor-neutral portfolio with a positive expected return. That is what a relative value book trades, and chapter 22 builds one explicitly, then measures whether a year of data can tell its return apart from zero.

This is also where chapter 8’s term structure models get their θ. Every bond on the curve is driven by the same short rate, so every bond must reward that one risk at one rate, and forming a riskless combination of two maturities is the argument above with n=2. The market price of interest rate risk is not an extra modelling assumption. It is (6.8) with one column.

Structure (The middle row: an incomplete market and a family of measures).

The table’s middle row deserves more than a name. Solvable but not unique means (6.8) is consistent and under-determined: the solutions are θ∗+ker⁡σ for any one particular solution θ∗, an affine subspace of dimension d−rank⁡(σ), positive exactly when there are more Brownian motions than the assets can pin down.

Every θ in that subspace gives a valid measure ℚθ for the n traded assets, and none is preferred over another by anything said so far. Shifting θ by a vector in ker⁡σ changes no asset’s drift at all, because only σ⁢θ enters (6.8); it does change the drift of anything driven by a Brownian combination outside the row space of σ, which is to say anything exposed to a risk the n assets do not see. So the traded market is silent between these measures, and a payoff exposed only to the row space of σ is priced the same under all of them, while a payoff exposed to ker⁡σ is not — the family 𝒬={ℚθ:θ∈θ∗+ker⁡σ} gives it a whole range of prices, 𝔼ℚθ⁢[payoff] moving as θ moves.

Definition 6.4 (Super- and sub-replication).

For a payoff X paid at T, a super-replicating strategy is a self-financing portfolio in the traded assets whose time-T value dominates X almost surely; the super-replication price is the cheapest such strategy’s cost today. A sub-replicating strategy is the mirror image, with time-T value at most X, and the sub-replication price is the most expensive such strategy’s cost.

Theorem 6.5 (Every price in 𝒬 sits between the hedge prices).

For every ℚθ∈𝒬,

sub-replication price≤𝔼ℚθ⁢[DT⁢X]≤super-replication price. (6.9)
Proof.

Let ϕ super-replicate X, with time-T value VTϕ≥X always. Being built only from the traded assets, ϕ’s discounted wealth is a ℚθ-martingale for every θ solving (6.8) — that is what solving it means. So for every such θ,

D0⁢V0ϕ=𝔼ℚθ⁢[DT⁢VTϕ]≥𝔼ℚθ⁢[DT⁢X],

using the martingale property for the equality and domination for the inequality. This holds for every super-replicating ϕ, so the cheapest one’s cost, the super-replication price, is at least 𝔼ℚθ⁢[DT⁢X] for every θ. The mirror argument, run on a sub-replicating strategy, gives the other half. ∎

Remark (What the theorem does and does not say).

It bounds every consistent price by the two hedge costs; it does not say those costs are finite, and it does not say the bound is tight. Finiteness is the subject of the example below. Tightness — that the super-replication price equals supθ𝔼ℚθ⁢[DT⁢X] exactly, not merely bounds it, so some strategy actually achieves the supremum — is a genuinely harder fact, proved in general by the optional decomposition theorem.111Kramkov (1996) in the semimartingale case, after El Karoui and Quenez (1995) for a Brownian filtration.

Example 6.2 (A payoff the market cannot hedge at any price).

Take the simplest incomplete market: one asset driven by W1 alone, and a second, independent Brownian motion W2 that drives nothing tradable. Then σ=(σ1,0), θ1=(μ−r)/σ1 is pinned, and ker⁡σ={(0,θ2):θ2∈ℝ} is the free direction of the structure above. Under ℚθ, W2 picks up drift −θ2, so W2,T∼N⁢(−θ2⁢T,T).

For the bounded payoff X=𝟏⁢{W2,T>0},

𝔼ℚθ⁢[X]=Φ⁢(−θ2⁢T)∈[0,1]

for every θ2 — bounded for free, because an expectation can never exceed the supremum of what it averages, whatever the measure. For the unbounded payoff X=eW2,T,

𝔼ℚθ⁢[X]=exp⁡(−θ2⁢T+12⁢T),

the moment generating function of a normal, and this diverges as θ2→−∞.222both closed forms checked against a direct simulation under ℚθ; the digital, swept over four orders of magnitude in θ2; the exponential, shown increasing without a ceiling. By theorem 6.5 that means the super-replication price of eW2,T is infinite: no finite position in the one traded asset can ever dominate a payoff that grows without limit in a direction nothing traded is exposed to. That is the theorem correctly reporting that this claim cannot be hedged at any finite cost, not a failure of it. It is the same question chapter 14 asks of a moment strip before trusting a transform: an expectation written down formally is not a price until its finiteness has been checked, and the check is not automatic once the payoff is unbounded.

Remark (Two different reasons a market can be incomplete).

Frictions are often filed under the same heading as ker⁡σ≠{0}, and one of them belongs there and two do not.

Trading only at finitely many dates is the same mechanism, discretised in time rather than in the number of assets: a position held between two dates produces a payoff linear in the state at the next one, which cannot replicate a general nonlinear function of it, for exactly the reason n assets cannot span d>n Brownian motions. Chapter 5 measures the residual this leaves behind directly, as a hedging error rather than as a price interval.

Transaction costs and position limits are not this. Neither removes any randomness or reduces how many independent risk factors the traded assets see; a market that would be complete without them can become impossible to hedge exactly once they are added, because the frictionless replicating strategy still exists and is merely not admissible — proportional costs make continuous rebalancing cost infinitely much, since the Black-Scholes delta has infinite variation, and a position limit forbids whatever hedge would need more than it allows. The reachable-payoff set still shrinks, and a super/sub-replication interval still results, but ker⁡σ has nothing to do with it: the relevant duality is a separate theory, for transaction costs Jouini and Kallal (1995) and Cvitanić and Karatzas (1996), for portfolio constraints its own constrained-hedging literature, none of it reducible to a solvability question about (6.8).

6.2 Numeraire Measures

Definition 6.6 (Numeraire).

A Numeraire is a strictly positive price process of a tradable relative to which prices of all other tradables are expressed.

A savings account which earns interest at the instantaneous interest rate can be taken as a numeraire. The value of the savings account at any point is given by

At=e∫0tr⁢(t)⁢𝑑t=1/D⁢(t) (6.10)

where D⁢(t) is the discount factor.

The fact that discounted trade prices are martingales in risk-neutral measure can then also be stated as: tradable prices in savings account (or money market) numeraire are martingales.

Consider another numeraire Nt. Since the numeraire is itself a price process, Dt⁢Nt is a martingale and we can take it as a Radon-Nikodym derivative, with a suitable normalization so that 𝔼ℙ⁢[(d⁢ℕ/d⁢ℙ)]=1, giving

𝔼ℕ⁢[XTNT|ℱt]=𝔼ℙ⁢[(d⁢ℕ/d⁢ℙ)T(d⁢ℕ/d⁢ℙ)t⁢XTNT|ℱt]=𝔼ℙ⁢[DT⁢NTDt⁢Nt⁢XTNT|ℱt]=1Dt⁢Nt⁢𝔼ℙ⁢[DT⁢XT|ℱt]=XtNt (6.11)

where ℙ is the risk neutral measure and ℕ is the measure corresponding to the numeraire Nt.

Remark (Prices in a numeraire are martingales in its measure).

The above equation shows that tradable prices with respect to a numeraire are martingales in the measure corresponding to the numeraire.

Risk neutral measure corresponds to the savings account (or money market) numeraire.

Definition 6.7 (T-forward measure).

If we take Nt to be the price of a riskless bond maturing at time T, the corresponding measure is the T-forward measure.

Xt/B⁢(t,T)=𝔼T⁢[XT/B⁢(T,T)]=𝔼T⁢[XT] (6.12)

We see that the expectation of XT in the T-forward measure gives the forward price Xt/B⁢(t,T) of the trade with expiration date T, hence the name T-forward measure. From the above equation we also notice that expiration T forward price process on an asset is martingale in T-forward measure.

Everything so far says which quantities are martingales under which measure, which is a statement about drifts being zero. The working question is usually the neighbouring one: given a process’s drift under one numeraire, what is it under another. The answer follows from Girsanov and the density identified above, and it is short enough to record once and use thereafter — chapter 13 in particular does almost nothing else.

Theorem 6.8 (Change of numeraire, in drifts).

Let A and B be numeraires, with ℚA and ℚB the corresponding measures, and let X be an Itô process. Then

driftB⁡(X)=driftA⁡(X)+dd⁢t⁢⟨X,ln⁡BA⟩t. (6.13)

In particular, for a strictly positive S the drift of ln⁡S picks up d⁢⟨ln⁡S,ln⁡(B/A)⟩, and so does the relative drift in d⁢St/St.

Proof.

B is the price of a tradable and A is the numeraire, so B/A is a price expressed in units of A and is a ℚA martingale by the remark above. Normalising it,

Zt=Bt/AtB0/A0,

gives a strictly positive ℚA martingale starting at one, and it is exactly the density (d⁢ℚB/d⁢ℚA)t identified earlier in this section.

Being positive it can be written d⁢Zt=Zt⁢νt⁢d⁢WtA, and Girsanov then says WB=WA−∫ν⁢𝑑s is a Brownian motion under ℚB. So for any d⁢Xt=μtA⁢d⁢t+σt⁢d⁢WtA, substituting d⁢WA=d⁢WB+ν⁢d⁢t gives

d⁢Xt=(μtA+σt⁢νt)⁢d⁢t+σt⁢d⁢WtB.

It remains to recognise σ⁢ν. By Itô, d⁢ln⁡Z=ν⁢d⁢WA−12⁢ν2⁢d⁢t, so d⁢⟨X,ln⁡Z⟩=σ⁢ν⁢d⁢t; and ln⁡Z differs from ln⁡(B/A) by a constant, which no covariation sees. That is (6.13).

For the last claim take X=ln⁡S. The relative drift follows because d⁢S/S=d⁢ln⁡S+12⁢d⁢⟨ln⁡S⟩, and the quadratic variation term is pathwise and so the same under both measures. ∎

Remark (Reading the formula).

Three things stand out, since the formula gets used more often than it gets derived.

It depends on B/A alone. Neither numeraire matters on its own, only their ratio, which is why changing from A to B and back leaves every drift where it started.

A process uncorrelated with that ratio does not move at all. The adjustment is a covariation, so it vanishes exactly when the process and the numeraire ratio share no randomness, however volatile either happens to be.

And what it adjusts is a drift, never a volatility — the structure block at the head of this chapter, arriving as a formula.

Example 6.3 (Option Pricing).

A measure change earns its place in pricing by re-deriving Black-Scholes directly. For an option on a stock with expiry T and strike K, the valuation formula is

Vt=1Dt⁢𝔼ℙ⁢[DT⁢max⁡(ST−K,0)|ℱt]

where ℙ is the risk neutral probability measure, Dt is the discount factor and St is the stock price which follows the Black Scholes diffusion

d⁢StSt=r⁢d⁢t+σ⁢d⁢Wt

where r is the riskfree interest rate and Wt is a Brownian motion in ℙ.

Denote by 𝕀ST>K the indicator variable which takes value 1 is ST>K and 0 otherwise. We then have

Vt =1Dt⁢𝔼ℙ⁢[DT⁢𝕀ST>K⁢(ST−K)|ℱt]
=1Dt⁢𝔼ℙ⁢[DT⁢𝕀ST>K⁢ST|ℱt]−1Dt⁢𝔼ℙ⁢[DT⁢𝕀ST>K⁢K|ℱt]

Taking 𝕊 to be the measure corresponding to stock price as numeraire, and taking 𝕋 to be the measure corresponding to T-expiry bond as numeraire, we have

Vt =1Dt⁢𝔼𝕊⁢[DT⁢(d⁢ℙ/d⁢𝕊)T(d⁢ℙ/d⁢𝕊)t⁢𝕀ST>K⁢ST|ℱt]−1Dt⁢𝔼𝕋⁢[DT⁢(d⁢ℙ/d⁢𝕋)T(d⁢ℙ/d⁢𝕋)t⁢𝕀ST>K⁢K|ℱt]
=1Dt⁢𝔼𝕊⁢[DT⁢(d⁢𝕊/d⁢ℙ)t(d⁢𝕊/d⁢ℙ)T⁢𝕀ST>K⁢ST|ℱt]−1Dt⁢𝔼𝕋⁢[DT⁢(d⁢𝕋/d⁢ℙ)t(d⁢𝕋/d⁢ℙ)T⁢𝕀ST>K⁢K|ℱt]
=1Dt⁢𝔼𝕊⁢[DT⁢Dt⁢StDT⁢ST⁢𝕀ST>K⁢ST|ℱt]−1Dt⁢𝔼𝕋⁢[DT⁢Dt⁢(B⁢(t,T))DT⁢B⁢(T,T)⁢𝕀ST>K⁢K|ℱt]
=St⁢𝔼𝕊⁢[𝕀ST>K|ℱt]−K⁢B⁢(t,T)⁢𝔼𝕋⁢[𝕀ST>K|ℱt]

where B⁢(t,T) is the bond price with the diffusion

d⁢B⁢(t,T)/B⁢(t,T)=r⁢d⁢t+σ2⁢d⁢W2,t

where W2,t is another Brownian motion in ℙ such that d⁢Wt⁢d⁢W2,t=ρ⁢d⁢t. Giving the bond its own volatility goes beyond chapter 5’s market, where the rate is constant and B⁢(t,T) has none; the derivation below specialises to the ordinary Black-Scholes d1,d2 when σ2=0, and otherwise prices the option under a genuinely stochastic discount factor.

Using Girsanov’s theorem, we have

d⁢Wt𝕊 =d⁢Wtℙ−σ⁢d⁢Wtℙ⁢d⁢Wtℙ =d⁢Wtℙ−σ⁢d⁢t
d⁢Wt𝕋 =d⁢Wtℙ−σ2⁢d⁢Wtℙ⁢d⁢W2,tℙ =d⁢Wtℙ−σ2⁢ρ⁢d⁢t

as Brownian motions in the measures corresponding to S, and B(t,T) numeraires respectively.

d⁢StSt=r⁢d⁢t+σ⁢(d⁢Wt𝕊+σ⁢d⁢t)=r⁢d⁢t+σ⁢(d⁢Wt𝕋+σ2⁢ρ⁢d⁢t) (6.14)

where d⁢Wt𝕊=d⁢Wt−σ⁢d⁢t is a Brownian motion in 𝕊 and d⁢Wt𝕋=d⁢Wt−σ2⁢ρ⁢d⁢t is a Brownian motion in 𝕋. Rearranging,

d⁢StSt=(r+σ2)⁢d⁢t+σ⁢d⁢Wt𝕊=(r+σ⁢σ2⁢ρ)⁢d⁢t+σ⁢d⁢Wt𝕋 (6.15)
ST=St⁢e(r+12⁢σ2)⁢τ+σ⁢Wτ𝕊=St⁢e(r+σ⁢(σ2⁢ρ−12⁢σ))⁢τ+σ⁢Wτ𝕋 (6.16)

where τ=T−t.

𝔼𝕊⁢[𝕀ST>K|ℱt] =𝔼𝕊⁢[𝕀St⁢e(r+12⁢σ2)⁢τ+σ⁢Wτ𝕊>K|ℱt]
=𝔼𝕊⁢[𝕀(r+12⁢σ2)⁢τ+σ⁢Wτ𝕊>log⁡(K/St)|ℱt]
=𝔼𝕊⁢[𝕀σ⁢Wτ𝕊>log⁡(K/St)−(r+12⁢σ2)⁢τ|ℱt]
=𝔼𝕊⁢[𝕀1τ⁢Wτ𝕊>1σ⁢τ⁢(log⁡(K/St)−(r+12⁢σ2)⁢τ)|ℱt]
=N⁢(−1σ⁢τ⁢(log⁡(K/St)−(r+12⁢σ2)⁢τ))
=N⁢(d1)

where N is the cumulative normal distribution, and d1=1σ⁢τ⁢(log⁡(St/K)+(r+12⁢σ2)⁢τ).

Similarly we have,

𝔼𝕋⁢[𝕀ST>K|ℱt] =𝔼𝕋⁢[𝕀St⁢e(r+σ⁢(σ2⁢ρ−12⁢σ))⁢τ+σ⁢Wτ𝕋>K|ℱt]
=𝔼𝕋⁢[𝕀(r+σ⁢(σ2⁢ρ−12⁢σ))⁢τ+σ⁢Wτ𝕋>log⁡(K/St)|ℱt]
=𝔼𝕋⁢[𝕀σ⁢Wτ𝕋>log⁡(K/St)−(r+σ⁢(σ2⁢ρ−12⁢σ))⁢τ|ℱt]
=𝔼𝕋⁢[𝕀1τ⁢Wτ𝕋>1σ⁢τ⁢(log⁡(K/St)−(r+σ⁢(σ2⁢ρ−12⁢σ))⁢τ)|ℱt]
=N⁢(−1σ⁢τ⁢(log⁡(K/St)−(r+σ⁢(σ2⁢ρ−12⁢σ))⁢τ))
=N⁢(d2)

where N is the cumulative normal distribution, and d2=1σ⁢τ⁢(log⁡(St/K)+(r+σ⁢σ2⁢ρ−12⁢σ2)⁢τ).

And finally we have,

Vt =St⁢N⁢(d1)−K⁢B⁢(t,T)⁢N⁢(d2)
=St⁢N⁢(d+σ⁢τ)−K⁢B⁢(t,T)⁢N⁢(d+σ2⁢ρ⁢τ)

where d=1σ⁢τ⁢(log⁡(St/K)+(r−12⁢σ2)⁢τ).

Note that σ⁢d⁢t and σ2⁢ρ⁢d⁢t are the drift correction terms we get to stock price Brownian motion using Girsanov’s theorem to change measure to stock and bond numeraires respectively.

Remark (The two terms are one event counted twice).

The derivation above deserves one more sentence: it explains something about the Black formula that is usually left as a curiosity.

N⁢(d1) and N⁢(d2) are both the probability of {ST>K} — the same event, the option finishing in the money. They are different numbers because they are computed under different measures: N⁢(d1) under the measure in which the stock is the numeraire, N⁢(d2) under the T-forward measure. Neither is “the” probability of exercise, and asking which one is misses that the question has no answer until a numeraire is named.

The gap between them is the measure change, and it is σ⁢τ worth of shift in the argument of N. It is therefore largest for long dated and volatile options and vanishes as σ⁢τ→0, where there is no randomness for the reweighting to act on and the two measures agree.

40608010012014016018020000.20.40.60.81StrikeProbability of finishing above the strike
  • Share measure, N(d₁)
  • T-forward measure, N(d₂)
Figure 6.2: The probability that a one year option finishes in the money, at a 25% volatility, under the two measures the Black formula is assembled from. The share measure weights each path by the terminal price itself, so it counts the paths finishing high for more — and those are exactly the paths on which the option is exercised. It therefore assigns the higher probability at every strike. Both curves describe the same event; the market has not been consulted twice.
Show the model behind this figure (1 function)
exercise_probabilitiesquant/src/measure.rs
/// The probability of finishing above `strike`, under two different measures.
///
/// Returns `(N(d1), N(d2))`. These are the two terms of the Black formula, and
/// the numeraires chapter derives them as the probability of the *same event*
/// under the share measure and under the `T`-forward measure. They are
/// different numbers because the measures are different, not because the event
/// is.
pub fn exercise_probabilities(forward: f64, strike: f64, sigma: f64, t: f64) -> (f64, f64) {
    if strike <= 0.0 {
        return (1.0, 1.0);
    }
    if sigma <= 0.0 || t <= 0.0 {
        let exercised = if forward > strike { 1.0 } else { 0.0 };
        return (exercised, exercised);
    }
    let vol = sigma * t.sqrt();
    let d1 = ((forward / strike).ln() + 0.5 * vol * vol) / vol;
    (norm_cdf(d1), norm_cdf(d1 - vol))
}
Structure (The pricing kernel).

A change to the risk neutral measure does two things at once, reweights the paths and discounts along the way, and the split between them depends on which numeraire was chosen. Their product does not. Write Dt for the discount factor of the savings account and Lt=(d⁢ℚ/d⁢ℙ)t for the density of the risk neutral measure against the real world one, and set

Mt=Dt⁢Lt. (6.17)

This is the pricing kernel, or stochastic discount factor. It is the single random variable that turns a real-world expectation into a price.

Every price is a real-world expectation weighted by M. For a payoff XT at T, the abstract Bayes formula of this chapter gives

Vt=𝔼ℚ⁢[DTDt⁢XT|ℱt]=𝔼ℙ⁢[LT⁢DTLt⁢Dt⁢XT|ℱt]=1Mt⁢𝔼ℙ⁢[MT⁢XT|ℱt].

Equivalently Mt⁢Vt is a ℙ-martingale for every tradable V. There is no discounting and no reweighting left over, and nothing in the formula mentions a numeraire.

Every numeraire measure is the kernel weighted by the numeraire. If N is a tradable, M⁢N is a strictly positive ℙ-martingale, so (d⁢ℚN/d⁢ℙ)t=Mt⁢Nt/(M0⁢N0) is a density, and

𝔼N⁢[XTNT|ℱt]=𝔼ℙ⁢[MT⁢NTMt⁢Nt⁢XTNT|ℱt]=VtNt,

which is the statement that prices in a numeraire are martingales in its measure. The savings account gives the risk neutral measure, the bond maturing at T gives the T-forward measure, and the stock gives the share measure of the Black formula. Choosing a numeraire only chooses how to divide the same kernel into a discount and a density.

The price of risk is the kernel’s loading. In one factor, d⁢L/L=−θ⁢d⁢W and D decays at the short rate, so d⁢M/M=−r⁢d⁢t−θ⁢d⁢W. For an asset d⁢V/V=μ⁢d⁢t+σ⁢d⁢W, requiring M⁢V to have no drift under ℙ gives μ−r−σ⁢θ=0, which is (6.6) read off the kernel: the excess return is minus the covariance of the asset with the kernel,

μ−r=−1d⁢t⁢Cov⁡(d⁢VV,d⁢MM).

An asset that is high where M is high pays in the states the kernel weights most, so it is a hedge: it is priced above its expected value and earns less than r. An asset that is high where M is low, like the stock, earns a premium.

A forward price is an expectation plus a covariance. At time zero, with M0=1 and P⁢(0,T)=𝔼ℙ⁢[MT],

F=𝔼ℙ⁢[MT⁢XT]𝔼ℙ⁢[MT]=𝔼ℙ⁢[XT]+Covℙ⁡(MT,XT)𝔼ℙ⁢[MT]. (6.18)

The forward is the real-world expectation shifted by the covariance of the payoff with the kernel. For the stock in chapter 5, which is high where M is low, the shift is negative, and the forward sits below the real-world mean by exactly S0⁢(eμ⁢T−er⁢T). Nothing about the size or sign of the shift is a fact about the payoff alone; it is a fact about how the payoff lines up with the kernel.

The kernel is the same object economics calls the marginal utility of wealth: a representative investor’s M is high in bad states, which is why an asset that pays there is worth more than its mean. That reading is an interpretation and the notes use nothing from it. Everything above follows from the existence of ℚ.333the kernel prices the stock, the bond and the call from the real world, the call against the Black-Scholes formula; the numeraire measures’ densities; the covariance identity for the stock.

References

  • -

    Girsanov, I. V. (1960). On transforming a certain class of stochastic processes by absolutely continuous substitution of measures. Theory of Probability and its Applications, 5(3), 285–301.

  • -

    Geman, H., El Karoui, N., & Rochet, J.-C. (1995). Changes of numeraire, changes of probability measure and option pricing. Journal of Applied Probability, 32(2), 443–458.

  • -

    Jamshidian, F. (1997). LIBOR and swap market models and measures. Finance and Stochastics, 1(4), 293–330.

  • -

    El Karoui, N., & Quenez, M.-C. (1995). Dynamic programming and pricing of contingent claims in an incomplete market. SIAM Journal on Control and Optimization, 33(1), 29–66.

  • -

    Kramkov, D. O. (1996). Optional decomposition of supermartingales and hedging contingent claims in incomplete security markets. Probability Theory and Related Fields, 105(4), 459–479.

  • -

    Jouini, E., & Kallal, H. (1995). Martingales and arbitrage in securities markets with transaction costs. Journal of Economic Theory, 66(1), 178–197.

  • -

    Cvitanić, J., & Karatzas, I. (1996). Hedging and portfolio optimization under transaction costs: a martingale approach. Mathematical Finance, 6(2), 133–165.