Skip to content
Sarthak Bagaria
All notes

Chapter 21 Estimation in Practice

Chapter 20 asked whether a fit means anything. This chapter asks the same question where the parameter is never observed at all — a variance, a regime, a factor driving a curve — and where the last decade of machine learning has made the boldest claims. Filtering recovers a latent state in principle and a network makes the recovery fast; neither changes what the data can identify, and telling the two apart is most of what follows. We work through what is genuinely hidden and what only looks that way, how long a regime takes to notice before it has already changed again, whether prices and a history can be fitted at once, the two parameters estimated constantly and understood least — a volatility and a correlation — and the oldest asymmetry in the subject: why implied volatility sits above realised, and what it costs to collect the difference.

21.1 What the Newer Methods Change

Everything in chapter 20 was classical. What the last decade of machine learning does and does not alter is worth setting out plainly, because the claims made for it are uneven.

21.1.1 Latent states: filtering, and the likelihood it hands over

Chapter 20 assumed throughout that the thing being estimated is observed. In fixed income it never is. Nobody sees the short rate, the variance, or the factors driving a curve; what is seen is a set of yields or option prices, each a known function of the state plus an error. That is a state-space model, and it changes what estimation means: there are now two problems, recovering the state and estimating the parameters, and the second cannot be attempted until the first is solved.

Definition 21.1 (State-space form).
xt+1 =f⁢(xt)+noise the state, unobserved
yt =h⁢(xt)+error what is quoted

A filter computes p⁢(xt∣y1:t), the state given everything seen so far.

Remark (When it is exact: the Kalman filter).

If f and h are affine and both noises Gaussian, the filtering distribution is Gaussian and the recursion is exact and cheap.

Linearity is what makes the recursion closed: an affine map of a Gaussian is Gaussian, so a distribution described by a mean and a variance stays described by a mean and a variance, and the filter never has to represent anything else.

Gaussianity is what makes it optimal among all estimators. Drop it but keep linearity and the same recursion is still the minimum-variance estimator among linear ones, for any noise with finite variance — so the filter is not useless off its assumptions, it is merely no longer using everything in the data. What does break is the likelihood: the number the recursion emits is then a Gaussian quasi-likelihood rather than the true one, which is still usable for estimation but gives standard errors that are wrong unless corrected.

In the affine Gaussian case both hold. That case is not a toy here — it is precisely a Gaussian affine term structure model, where yields are affine in the state by construction, so chapter 8’s models are estimated this way as a matter of course.

The recursion is two steps. Predict moves the state distribution forward by the dynamics, which widens it. Update conditions on the new observations, which narrows it, by a precision-weighted average of the prediction and the data — Bayes’ rule for Gaussians and nothing more.

The by-product is what makes it valuable. Each step produces a predictive distribution for the next observation, and summing its log density over the history gives p⁢(y1:T∣θ) exactly. So the filter does not merely estimate the state, it delivers the likelihood of the parameters — which is what makes maximum likelihood, or the posterior of the previous subsection, possible at all for a model whose state is never seen.111estimation::AffineStateSpace, with the likelihood peaking at the true mean reversion and the filtered state beating the best single yield by a third.

Structure (The measurement error is not a nuisance term).

Five yields and a one-factor model is an over-determined system.

The difficulty is that the curve is not free. Here yt∈ℝ5 is the vector of the five yields quoted at date t, one per maturity, and h is affine: yt,i=ai+bi⁢xt, with loading bi=(1−e−κ⁢τi)/(κ⁢τi) and intercept ai fixed by the parameters. The parameters are held still while the filter runs; the only thing that moves from one date to the next is the scalar state xt. So the observation is not five points to be joined but a five-vector,

yt=a+xt⁢b,a,b∈ℝ5⁢ fixed,

with a=(a1,…,a5) and b=(b1,…,b5) the fixed intercept and loading vectors, required to lie on a line in ℝ5 — the set {a+s⁢b:s∈ℝ} that the scalar xt sweeps out as it varies, one dimension inside five. Five numbers, one unknown, and generically no xt reproduces them. The measurement error is what makes the problem well posed.

That also settles an apparent contradiction with chapter 8, which fits today’s curve exactly with a one-factor model. It does, and it buys the fit with a time-dependent drift: a free function, in effect one number per maturity, spent entirely on the first date. It does nothing on any date after, where the same fixed loadings apply and only x has moved. A one-factor model can match any single curve and cannot match a panel of them, and it is the panel the filter sees.

Two consequences. Assuming the yields are cleaner than they are makes the state estimate worse, because the filter then trusts each observation more than it deserves and chases the noise — which is measurable and is measured. And the fitted size of the error is a diagnostic worth reading: a model whose estimated measurement error comes out far above the bid-offer is telling you it cannot fit the cross-section, whatever its time series behaviour looks like. That is a specification test that costs nothing, since the number is estimated anyway.

Example 21.1 (Estimating the short-rate dynamics and the noise together, from yields alone).

The same five-yield model as above, now with nothing handed to the filter but 3,000 daily observations: the mean reversion κ, the state’s own volatility σ, and the measurement error common to every yield, all three estimated at once by maximising the Kalman likelihood.222estimation::fit_affine_state_space.

It comes back with κ^≈0.315 against a truth of 0.3, σ^≈0.0092 against 0.01, and a measurement error of about 5.5 basis points against a true 5 — every one of the three within ten percent. The recovery is tighter than Heston’s below, and for a specific reason: this likelihood is exact and smooth in every parameter, with no resampling to make it noisy, so a coordinate search has an easy target rather than a rough one.333the recovery, measured against each of the three truths at once.

Remark (When it is not exact).

Affine and Gaussian is a strong assumption and the alternatives are ranked by how much of it they give up.

Extended and unscented filters keep the Gaussian approximation and handle a nonlinear h — the first by linearising, the second by propagating a small set of points through the nonlinearity and refitting a Gaussian. Both are cheap and both fail in the same circumstance, which is a nonlinearity strong enough that the filtered distribution is not Gaussian in shape. A square-root variance near zero is exactly that circumstance, so the models where one most wants a filter are the ones where these approximations are least safe.

Particle filters give up the Gaussian entirely: represent the distribution by a weighted sample, push each particle through the dynamics, reweight by how well it explains the new observation, and resample. Nothing is assumed about shape, which is why this is the method for a stochastic volatility model, where the state is a variance that is positive and skewed.

The cost is the usual one and it has a specific form. Weights degenerate — after a few steps one particle carries almost all of them — so resampling is required, and resampling makes the likelihood estimate noisy and discontinuous in the parameters, which breaks any optimiser that wants a gradient. The repair is to embed the particle filter’s unbiased likelihood estimate inside a Markov chain Monte Carlo scheme rather than to optimise it, which is exact despite the noise and is why particle methods and the Bayesian treatment above are usually met together.

21.1.2 Particle filters, in more detail

Stochastic volatility has a problem the sections above sidestepped: vt is not observed. Prices are, volatility is not, so the likelihood of a Heston model given a price history involves integrating over every possible volatility path — an integral of enormous dimension, which is why stochastic volatility models are so often calibrated to option prices instead of estimated from returns.

Sequential Monte Carlo solves it. Carry a population of particles, each a hypothesis about the current vt; propagate them through the model’s dynamics; reweight by how well each explains the newly observed return; resample. What comes out is the filtered distribution of the latent state, and — the part that matters here — the likelihood of the data as a by-product, which makes maximum likelihood estimation possible for models where it previously was not.

Two things about it matter before using it, because both are places it fails quietly.

Weights degenerate. After a few steps one particle carries almost all the weight and the rest carry none, so the effective sample size collapses and the filter is representing a distribution by one point. Resampling is the standard repair — discard the low-weight particles and duplicate the high-weight ones — and it trades the degeneracy for impoverishment, since the surviving particles share ancestors and the population loses diversity in the distant past. Neither problem is visible in the output unless the effective sample size is monitored, and it should be.

The proposal matters more than the particle count. The naive scheme propagates particles through the model’s own dynamics and then reweights, which is wasteful when the observation is much sharper than the dynamics: nearly every particle lands somewhere the data rules out, and the few that survive determine the answer. Proposing from a distribution that already accounts for the new observation — the auxiliary and guided variants — fixes this, and the improvement is far larger than the one bought by multiplying the number of particles.

This is the principled answer to a real problem, it is genuinely used, and it has no fashionable name.

Calculation 21.2 (A Heston variance, filtered from returns alone).

Simulate a Heston path — κ=2, θ=0.04, η=0.5, ρ=−0.6, daily steps over six years — and forget the variance was ever known. The bootstrap filter above needs one adjustment for the leverage correlation ρ: draw each particle’s variance shock before scoring the day’s return against it, so the return’s likelihood and the particle’s proposed next variance share the randomness the model says they must, rather than being drawn as though they came from two independent sources.444estimation::heston_particle_filter.

Handed the true parameters, the filter’s own estimate of vt has a root mean squared error a bit over half the size of simply reporting the long-run average θ every day: real information, and far from all of it — which is exactly why these models are usually calibrated to option prices instead of estimated from returns, a fact stated earlier in this section and measured here.555measured, against the constant-θ baseline.

Example 21.2 (Estimating without being handed the answer).

Now the harder version of calculation 21.2: nothing but the six years of returns, and all four parameters — κ, θ, η, ρ — estimated by maximising the filter’s own likelihood.666estimation::fit_heston_particle_filter. A coordinate search rather than a gradient method, deliberately: resampling makes the likelihood surface noisy and discontinuous in the parameters, for the reason given above, and a textbook implementation would go further and embed the filter inside a particle Markov chain Monte Carlo scheme, which tolerates that noise rather than being defeated by it. This is the cheaper version — common random numbers held fixed across the search to tame the discontinuity rather than remove it — enough to show the mechanism recovers something real.

It comes back with κ^≈1.76 against a truth of 2, θ^≈0.037 against 0.04, η^≈0.58 against 0.5, and ρ^≈−0.50 against −0.6 — every parameter within thirty percent of the truth, the leverage correlation within a tenth. Not exact, and not nothing: four parameters of a model whose entire state is unobserved, recovered from the one thing that was ever going to be observed.777the recovery, measured against each of the four truths at once.

Remark (Heston is affine too — so why doesn’t the Kalman filter above apply to it?).

Chapter 19’s transform methods rest on Heston’s variance being affine: its drift and its diffusion coefficient squared, η2⁢vt, are both affine functions of vt. That is a real fact about the model, and it is not the fact a Kalman filter needs.

A Kalman filter needs the map from state to observation, and the map from one state to the next, to be affine with additive, state-independent Gaussian noise — so that an affine map of a Gaussian stays Gaussian, and the filtering distribution never leaves the family it started in. Vasicek and Hull-White satisfy this: their diffusion coefficient is a constant σ, the corner of the affine family where the state-dependent part of the diffusion vanishes entirely. Heston’s diffusion coefficient is η⁢vt, genuinely state-dependent, so the noise entering the transition is neither additive nor of fixed scale, and the state’s own stationary law is non-central chi-squared rather than Gaussian. The affine family that makes chapter 19’s transforms work is strictly larger than the Gaussian corner of it that makes this chapter’s filter exact, and the two conditions are easy to conflate precisely because both go by the same name.

Example 21.3 (What linearising it costs).

Apply an extended Kalman filter to the same simulated path as calculation 21.2, given the same true parameters. Two simplifications are needed to get there at all: the leverage correlation is dropped from the observation equation, and the CIR transition is linearised around the current mean rather than carried through exactly.888estimation::ekf_heston_variance.

The particle filter’s RMSE against the true variance sits 42% below the constant-θ baseline; the extended Kalman filter’s sits only 5% below it — barely better than not filtering at all, and about 63% higher than the particle filter’s own error on the identical path. The gap is not an implementation shortfall on the linearised filter’s part; it is the measured cost of the two simplifications the remark above made necessary, on a calibration where the Feller condition is violated and vt genuinely visits the region — near zero — where the linearisation is worst.999the two filters, measured on the same path.

21.1.3 Reporting the valley: Bayesian calibration

Chapter 20 ended with a complaint: a calibration reports ρ=−0.30 when the evidence supports anything from −0.28 to −1. The Bayesian answer is direct — do not report a point. Put a prior on the parameters, condition on the quotes with a likelihood reflecting the bid-offer, and report the posterior.

The output is then the flat region itself rather than an arbitrary point inside it, and every downstream number inherits the uncertainty honestly. A risk figure computed across the posterior is a risk figure that knows the model was not fully determined, which is exactly the information a point estimate throws away.

The cost is computational and it is real, since the posterior must be sampled rather than optimised. But note that nothing has been added to the model. The information was always this thin; the Bayesian version merely declines to hide it.

Structure (A posterior over parameters is a posterior over prices).

The consequence is larger than better reporting of a calibration.

A price is a function of the parameters, V=V⁢(θ). A posterior p⁢(θ∣quotes) therefore pushes forward to a distribution of prices, obtained by pricing under each draw. That distribution is the object several later chapters have been approximating by hand.

It is the model reserve. Chapter 24 computes model risk by fixing a plausible range for an undetermined parameter, revaluing at each end, and taking the difference. That is a two point approximation to a push-forward, and it is crude in a specific way: it ignores the shape of the valley, ignores that two parameters may be jointly undetermined along a diagonal while each looks fine alone, and produces a number whose size depends on how wide somebody drew the range. A posterior quantile is the same quantity computed rather than sketched.

And it nets correctly across several jointly undetermined parameters at once, which the two-point version cannot. Chapter 24 already computes its reserve at the level of the book rather than trade by trade, for exactly this reason: two trades whose exposure to a single parameter offsets have no joint model risk even though each has plenty on its own. But that repair works only because there is one parameter to revalue against. Once several are jointly undetermined, as the diagonal case above shows, a separate two-point range in each has no notion of the joint region they occupy together and nothing to revalue against it. The push-forward has this for free: one draw of the whole parameter vector θ prices every trade in the book, so the distribution of the portfolio value already reflects every cancellation — across trades sharing one parameter, and across parameters that are jointly rather than separately undetermined — inside the same computation that produced the trade-level number.

It gives relative value a decision rule. Chapter 22 insists that a model saying a spread is wide is not a signal. A posterior says how wide the spread would have to be to mean anything: if the observed spread lies inside the band the model’s own uncertainty generates, the model has not identified a mispricing, it has failed to determine a price. Only a spread outside the band is a candidate — and even then the structural reason chapter 22 insists on is still required. This is a sharper filter than the usual one and it is available at no extra modelling cost.

Remark (What the band is really a statement about).

Along a direction the quotes constrain, the posterior is narrow and is driven by the data. Along a flat direction — and chapter 20’s whole point is that these exist — the posterior is the prior, because conditioning on uninformative data returns what was assumed. So the width of the band in those directions measures the prior rather than the market.

That is not a defect, provided it is said. It converts an assumption that was previously invisible into a number somebody has to write down and defend, which is precisely what chapter 24 says the reserve process is for. A band that is wide because the prior was wide is an honest statement that nobody knows; a point estimate in the same situation is a dishonest statement that somebody does.

The practical obstacle is cost, since sampling a posterior requires thousands of pricings of a model that may take minutes each. Which is where the next subsection’s neural approximation earns its place: a fast surrogate for the pricing map does not resolve the identification problem, but it makes the calculation that reports the problem affordable. Speed does not buy knowledge; it buys the ability to say how little one has.

21.1.4 Speed, not knowledge: neural approximation

The honest and successful use of neural networks in derivatives is unglamorous: learn the map from parameters to prices offline, once, and then evaluate it in microseconds.

That pays off because the map is expensive for models like rough volatility where each price is a simulation, and calibration needs thousands of evaluations. Learning it turns an overnight calibration into an interactive one. The model is unchanged, the prices are unchanged, and the network is a lookup table with good interpolation — which is why it works and why it is safe: the approximation error can be measured against the true pricer on as many test points as one likes.

Remark (The limit, which is the point of this chapter).

A neural network approximating the pricing map does not, and cannot, resolve chapter 20’s problem. The flat valley is a property of the map itself — of the fact that quotes near the money contain almost nothing about ρ. Learning that map faster reproduces the flatness faithfully, at speed. No architecture recovers a parameter the data does not contain.

Machine learning changes what is computationally feasible. It does not change what is identifiable, and identifiability is a property of the instruments quoted.

21.1.5 Learned models, and the two questions they do not answer

The more ambitious programme replaces the model rather than approximating it: neural stochastic differential equations, market generators, deep hedging, and network solutions of high-dimensional pricing equations. Those belong to chapter 19, since they are about computing a price rather than fitting one.

Two cautions belong here rather than there, because they are about fitting rather than about what the model can be. The first needs stating carefully, because the usual version of it is wrong.

Whether a learned model is arbitrage-free depends entirely on what is learned. A network fitted directly to the price surface has no structural guarantee: chapter 4 established that a price system without a martingale measure admits a money pump, and nothing in a training loss enforces one — a constraint can be added as a penalty, but a penalty is a preference and not a guarantee. A network placed inside the coefficients of a stochastic differential equation is in a completely different position, and chapter 19 works out why in full: chapter 2’s representation theorem fixes the drift and leaves only the diffusion coefficient free, so writing

d⁢StSt=r⁢d⁢t+σθ⁢(t,St,vt)⁢d⁢Wt (21.1)

with σθ a network produces an arbitrage-free model for every value of θ — the model class itself rather than a restriction on it, with the constraint carried by the equation rather than by a penalty. The remaining problem is the one this chapter is about, which no amount of structure removes: a finite set of quotes does not determine a function.

And chapter 9’s warning transfers exactly. A generative model that reproduces the historical distribution of surfaces perfectly is in the position local volatility was in: excellent fit, unconstrained dynamics, and no guarantee that the hedge ratios it implies are the ones that work. Fitting more flexibly does not turn an in-sample fit into evidence.

21.2 What Is Actually Latent

The previous section filtered a state on the grounds that it is not observed, and how much of that is true deserves asking.

Remark (Adapted to what?).

“Adapted” is not a property of a process but a relation between a process and a filtration, and three of them are in play whenever a latent state is posited: 𝒢t, generated by what is observed; ℱt, the model’s own, containing whatever latent factors it carries; and the one generated by the driving Brownian motions. A process can be ℱ-adapted and not 𝒢-adapted, and that gap is what the word “latent” is pointing at.

Two separate things force adaptedness, and they are usually conflated. The Itô integral needs a predictable integrand to exist at all — chapter 1’s construction fails outright if the integrand sees the increment it multiplies — and since σ is that integrand, it must be predictable in whatever filtration the integral is taken in. That is well-posedness, not economics. Separately, no-arbitrage requires the price to be a semimartingale in the market’s filtration, which is chapter 4’s use of the Bichteler-Dellacherie theorem, and requires trading strategies to be predictable.

Put those together and positing a latent state does not escape adapted volatility, it only relabels it. If S is a semimartingale in 𝒢, decompose it there; the continuous martingale part is ∫σ^⁢𝑑W^ with σ^ 𝒢-adapted, by chapter 2’s representation theorem. Whatever hidden machinery is imagined upstream, the observable description always has an adapted volatility. What changes is which one — v becomes something like 𝔼⁢[v∣𝒢].

Structure (Incompleteness is not unobservability).

There is a sharper statement, and it cuts against the usual picture of stochastic volatility as a model with something hidden in it.

Quadratic variation is a limit of sums of squared increments of the observed price, so ⟨S⟩ is measurable with respect to the price’s own filtration. Where it has a density, vt=d⁢⟨S⟩t/(St2⁢d⁢t) is therefore a function of the path being watched. And once v is known, its own innovations d⁢Z=(d⁢v−κ⁢(θ−v)⁢d⁢t)/(η⁢v) are known too. So continuous observation of a single price reveals the entire state of a standard stochastic volatility model. Nothing in it is hidden.

What remains true is that the market is incomplete: one traded asset cannot span two Brownian motions, so Z-risk cannot be hedged with S alone. That is a statement about spanning, and it is routinely confused with a statement about observation.

Which relocates the question. Genuine latency does not come from the model. It comes from the observation being discrete, from the microstructure noise that Calculation 21.6 shows drives realised variance away from the truth rather than towards it, and from jumps confounding the estimate. The filtering literature is about imperfect observation, not about a metaphysically hidden state — and that is a more tractable problem, because the imperfection can be measured.

Remark (Two projections, and only one of them is a local volatility model).

Chapter 9 projects a stochastic volatility model onto σloc2⁢(t,x)=𝔼⁢[vt∣St=x], conditioning on the current level and nothing else, and matches one-dimensional marginals. The projection onto what an observer knows conditions on the whole observed path,

σ^t2=𝔼⁢[vt|𝒢t],

and these are not the same object. The second is not a local volatility model at all: σ^t depends on the entire history, which makes it a path-dependent volatility model.

That distinction explains a fact both chapter 10 and chapter 18 arrive at from other directions. Matching marginals leaves the joint law free, so two models with the same σloc agree on every European price and disagree completely about a forward smile. The path projection keeps what the level projection throws away, which is exactly the dependence across time.

It also names the family that modern work has moved towards. If the honest observable model is a functional of the path, then write one. Guyon and Lekeufack’s volatility, for instance, is a function of two exponentially weighted moments of past returns — a trend and an activity level — fitted directly to an index’s own history rather than to a hidden factor.101010Guyon, J., & Lekeufack, J. (2023). Volatility is (mostly) path-dependent. Quantitative Finance, 23(9), 1221–1258. In full generality, volatility can be written as a functional of the path signature, which is dense in continuous path functionals, so which features of the history matter is learned rather than hand-picked — the construction Perez Arribas, Salvi and Szpruch use to build an entire pricing model this way.111111Perez Arribas, I., Salvi, C., & Szpruch, L. (2020). Sig-SDEs model for quantitative finance. Proceedings of the First ACM International Conference on AI in Finance. Such a model is fully observable by construction, and it has no latent state to filter.

Remark (What a neural stochastic differential equation does and does not escape).

A network in the coefficients of

d⁢St=μt⁢d⁢t+σθ⁢(t,St,vt)⁢St⁢d⁢Wt

is still an Itô diffusion with adapted volatility. It does not go beyond the form — chapter 2 showed the form is not a restriction — it goes beyond the parametric families within the form, which is a real gain and a different one.

The genuine extensions are three, and none is a change of equation. Richer path dependence, as above. Rough volatility, where v is driven by a fractional Brownian motion with H near 0.1: the volatility is then neither Markov nor a semimartingale, while the price remains a perfectly good semimartingale and no-arbitrage is untroubled. And jumps, which is the only one that changes the shape, and chapter 2 says what is obtained.

There is also a practical asymmetry to weigh, and it is this chapter’s argument in a new setting. A latent-state neural model needs filtering inside the training loop, since the likelihood integrates over unobserved paths — which is why such models are fitted by variational approximation and why their parameters are so poorly determined. A path-dependent model has no latent state, so every conditioning variable is observed, the likelihood is direct, and the parameters are identified by data one actually has. If identification rather than compute is the binding constraint, flexibility is better spent on the path dependence than on a hidden state.

21.3 Regimes, and Whether They Can Be Caught

One observation motivates latent states more than any other: past prices are frequently a poor guide to future ones, in a way that looks less like noise than like something changing. A policy stance, a liquidity condition, information held by somebody else — none of it visible in the price history until it acts. The standard response is a regime: volatility takes one of two values, the state switches at random times, and the state is not observed.

It is tempting to describe this as a parameter changing rather than a state being hidden, and the distinction does not survive inspection. A regime indicator is a state process; calling σ a parameter that changes is declining to model it, and modelling it turns the parameter into a state.

Remark (What actually separates a regime from a diffusive variance).

Something does separate them, and it is not what the word “latent” suggests.

Neither is hidden under continuous observation. If the regimes differ in volatility then vt=σ2⁢(Rt) is recoverable from the bracket by the argument above, so the regime is recoverable too — a two-state model is no more opaque than a square-root variance. The difference is in the predictability of the change rather than the observability of the level. A diffusive variance has continuous paths, so its next value is close to its current one and the near future is forecastable from the present. A regime switches at a totally inaccessible stopping time in the sense of chapter 2: no announcing sequence exists, so nothing in the observed past gives warning. It is fully observable and completely unforecastable, and those two properties are usually assumed to be incompatible.

Which leaves the practical difficulty exactly where the next section puts it. With continuous noiseless observation neither object is hidden, so everything that makes a regime hard is the discreteness and the noise of real observation, and (21.2) is the measure of what that costs.

There is a third case. Suppose the transition law is unknown, or the states cannot be enumerated — which is what people usually mean when they say the world changed. Then no filter applies, because a filter needs a likelihood for each state and a prior over the set of them, and neither exists. That is not a latent state but an unspecified model, and the honest response is chapter 24’s reserve rather than an estimate. Distinguishing the two is worth more than distinguishing parameters from states: one is a computation and the other is an admission.

Remark (The regress, and where to stop).

Modelling a parameter as a process does not remove an arbitrary choice, it relocates one. A constant σ becomes a variance process with its own κ, θ and η; making those stochastic in turn introduces theirs. Each level pushes the arbitrariness up rather than eliminating it, and there is no level at which the regress terminates on its own.

So the stopping rule has to come from outside the modelling, and this chapter supplies it. Add a level only where the data can see it move faster than it moves — which is (21.2) again, applied to the new state rather than to a regime. A parameter promoted to a process whose variation the observations cannot resolve has not been modelled; it has been given somewhere to hide, and the fit will improve while nothing has been learned. That is the whole of this chapter in one criterion.

Remark (The Bayesian version is the natural one).

A regime is a discrete state, so the posterior over it is a probability rather than a distribution over a continuum, and the filter is correspondingly cheap. It needs a prior over the two states to start the recursion — the stationary distribution implied by the switching probability is the natural default, so the filter does not open biased toward either regime. Predict by the switching probability, update by Bayes against the two observation densities, renormalise: two multiplications per observation, exactly.

What that buys is the object the Bayesian calibration above asked for. The output is not a regime but a probability of a regime, so every price computed from it is a mixture, and the width of that mixture is a confidence band that updates as each new price arrives. A desk holding a position can therefore report not only its value but how much of the value is a bet on the state estimate, which is the number that ought to size the position.

The question that decides whether any of this is worth building is not accuracy but speed measured in data. A regime that is detected reliably after two hundred observations, and lasts fifty, has been correctly identified and is of no use whatever. That comparison has an answer.

Calculation 21.3 (How long it takes to notice).

Sequential detection theory gives the expected delay to declare a change, subject to a false alarm rate α, as

delay≈ln⁡(1/α)Dobservations,D=12⁢(r−1−ln⁡r),r=σhigh2/σlow2, (21.2)

with D the Kullback-Leibler divergence per observation between the two regimes’ return densities. The numerator is the evidence that must be accumulated; the denominator is the rate at which each observation supplies it.

Running the filter on simulated switches, against the decision rule that declares a switch once the new regime’s posterior crosses 0.9 — a threshold, not itself the false-alarm rate, which the simulation measures rather than assumes:121212estimation::RegimeFilter, which measures the delay against (21.2) rather than assuming it.

Regime shift D per observation Measured delay Predicted
15%→30% 0.807 9.4 8.5
15%→25% 0.378 16.3 17.3
15%→18% 0.038 97.9 152.2

The measured delay tracks (21.2) for the two wide shifts and beats it for the narrow one, which is expected — the bound is asymptotic in small α and the filter is running at a false alarm rate of a third of a per cent.

Read the last row. A volatility regime moving from fifteen to eighteen takes about a hundred trading days to establish at ninety per cent confidence, which is five months. If such a regime lasts weeks, it is real, it is undetectable in time to act on, and a model containing it is describing something nobody can trade.

Structure (The screening criterion comes before the model).

Equation (21.2) turns a modelling question into an arithmetic one, and it should be asked before anything is built.

Is the expected life of the regime longer than ln⁡(1/α)/D? If not, stop. No filter, Bayesian or otherwise, recovers information the observations do not contain, and the delay above is not a property of the algorithm — it is a property of the two distributions being told apart. A better estimator cannot beat it, because it is the rate at which the data distinguishes the hypotheses.

Three consequences follow.

The gap matters quadratically, not linearly. Expanding D for a small shift gives D≈(r−1)2/4, so halving the difference between the regimes quarters the information and quadruples the delay. Regimes that are subtle are not slightly harder to detect; they are out of reach.

Sampling faster helps, in calendar time, and only inside the model. The interval cancels out of D entirely — a return carries the same information about a volatility ratio whatever span it covers — so the delay in observations is fixed by the regimes alone and the delay in calendar time is that divided by the observation frequency.

The cancellation, however, needs the returns to be independent draws of variance σ2⁢Δ, so that halving Δ doubles the count of observations and leaves each one’s information intact. Three things go wrong with that as Δ shrinks.

Microstructure noise floors the estimate, by (21.4), so past some frequency the extra observations are measuring the spread.

The returns also stop being independent, since the same noise induces autocorrelation in them — and n dependent observations do not carry n times one observation’s information, so the delay in observations stops falling as fast as the count rises.

The third is the one that is easiest to miss and hardest to repair. The quantity being estimated changes with the scale. Intraday volatility carries a pronounced daily pattern, opening and closing far above the middle of the session, and it responds to liquidity events that leave no mark on a daily series. So a regime in minute-by-minute volatility is not a finer measurement of the daily regime; it is a different object with its own dynamics. Sampling faster to detect a shift in daily volatility can converge quickly and confidently on a shift in something else.

Which gives the honest version of the criterion. The frequency that optimises the volatility measurement is a ceiling and not a target: sample as fast as (21.4) allows, but no faster than the scale on which the regime you care about is defined. If the regime is a monthly change in daily volatility, minute data does not shorten the delay for it, however many observations it supplies.

And the criterion generalises past regimes. Any latent structure — a switch, a jump in a parameter, a change in correlation — is worth carrying only if the data separates its states faster than the states change. That is a sharper version of this chapter’s recurring complaint. Elsewhere the trouble was that the calibration set did not determine a parameter; here it is that the data cannot determine it in time.

Chapter 20’s remark Sampling faster does not help reaches the same verdict for mean reversion, and the two are the same criterion at the level of the general statement — but not at the level of mechanism, which is worth keeping straight. There, sampling faster never helps, even under noiseless continuous observation: what a history identifies is bounded by elapsed half-lives alone, and slicing the same span more finely adds none. Here, sampling faster genuinely does help, without limit, in the idealisation — D is interval-invariant, so more observations in the same span is strictly more evidence — and it is only the three practical obstructions above that cap the benefit. Mean reversion fails to be identifiable in time as a theorem; a regime fails to be identifiable in time as a fact about noise, dependence, and scale that a cleaner market would not share.

Remark (What a subtle regime excuses, and what it does not).

A discrete regime that sits close to its neighbour is the coarse shadow of a continuous state: a chain with many closely spaced levels converges, under the same kind of scaling that turns a random walk into Brownian motion, to a diffusion. Recasting the problem that way trades one difficulty for another rather than removing it. There is no longer a threshold to cross or a false alarm to control — filtering vt in a stochastic volatility model reports a posterior mean and variance continuously, with no decision forced on any date — but the posterior is left just as imprecise as the discrete version already said it must be. A vol-of-vol too small, or a mean reversion too slow, is filtered with the same poverty of information that here reads as an undetectable regime; calculation 21.4 measures the continuous version of the same fact, that only a longer history pins down a slow coefficient, no faster sampling.

The other question — does an undetectable regime matter — has two different answers depending on which quantity is asked about, and they diverge by more than one might expect. A price computed today, Bayes-weighted across both regimes, is close to either regime’s own price whenever the regimes are close and the pricing map is smooth in the parameter — true from the first day, with nothing to wait for, since it is a statement about the mixture of two close numbers rather than about how much evidence has accumulated. A position already held and hedged on the wrong assumed volatility the whole time is a different question, and here subtlety excuses nothing: the gap in realised variance after N observations grows in proportion to N⁢(r−1), linear in the very quantity whose square sets the detection delay, so the realised divergence crosses any fixed materiality threshold after about 1/(r−1) observations while confident detection needs about 1/(r−1)2. The ratio of the two diverges as the regimes converge, so for a subtle enough pair the cost is already material long before anyone could have told the two regimes apart with any confidence at all. Not knowing which regime is in force is not the same as its not mattering; it only fails to matter for the number that gets recomputed fresh each morning.

What can break that first claim is a payoff whose price is not smooth in the parameter where the two regimes sit — a barrier or a digital positioned exactly where they diverge — and there the Bayes-mixed price can be sensitive to a regime probability the data barely moves. That is a statement about the curvature of the payoff, not about the difficulty of the estimation, and it is a third and different thing from the other two: one says the model risk is invisible in the number recomputed daily, the other says it accumulates silently in the number nobody recomputes until the position is closed.

21.4 Calibrating to Prices and to History at Once

The programme suggested by (21.1) is more ambitious than fitting a surface, and it is the natural one to attempt: learn σθ so that the model reprices the liquid instruments and reproduces the dynamics the price history actually exhibits. One model, two data sources, two terms in the loss.

The question is whether that is coherent, since the two datasets live under different measures. It is, and the reason is exact rather than approximate — which makes this one of the few places in the chapter where the news is good.

Remark (Girsanov decides what is shared).

A change of measure alters the drift of a diffusion and leaves its diffusion coefficient untouched. That is chapter 6’s theorem, and read as a statement about estimation it says: σ is the same object under the historical measure and the pricing measure, while μ is not.

The reason runs deeper than Girsanov’s formula happening to spare σ: a change of measure is a reweighting of path probabilities and nothing else, so it never touches a path itself. Quadratic variation is a pathwise limit of sums of squared increments of the observed price, computed without ever taking an expectation, so it cannot see that reweighting at all: ⟨S⟩t, and with it vt=d⁢⟨S⟩t/(St2⁢d⁢t) and σ=vt, are the same function of the same path under ℙ and ℚ, not two numbers that merely happen to agree. That is what licenses reading σ off a single history and carrying it into a pricing model with no translation required, and the same argument runs one level down in a stochastic volatility model: vt and vt are themselves fixed once the path is, so they too are the same pathwise object under both measures.

What a change of measure moves is vt’s own drift, and through it which of vt’s paths are the likely ones — never the value vt takes on a path already fixed. The confusion, if there is one, is forgetting that the Brownian motion driving vt has changed too: Girsanov relates the two by d⁢Wtℚ=d⁢Wtℙ−λt⁢d⁢t, and substituting one for the other inside vt’s equation exactly cancels the drift difference, leaving vt⁢(ω) untouched on every path. ℙ and ℚ do not disagree about where vt goes; they disagree about which process counts as the driving Brownian motion. §21.7 works out exactly which drift parameters move, and by how much.

Under the pricing measure the drift is not a free parameter at all — it is r for a tradeable, or the drift condition of chapter 8 for a curve. So of the two coefficients in (21.1), one is determined by no-arbitrage and the other is shared between the two measures. The only genuinely free object linking them is the market price of risk λ=(μℙ−μℚ)/σ.

That much is standard. What makes the combination pay off is a fact about what a price history can and cannot measure.

Calculation 21.4 (What a history identifies, and what no amount of it does).

Observe a diffusion with drift 5% and volatility 20% over a fixed window of one year, sampled n times, and estimate both coefficients. Averaged over four hundred independent histories:131313estimation::identification_by_frequency.

Samples in the year Volatility error Drift error
250 3.64% 0.158
2 500 1.12% 0.155
25 000 0.36% 0.159

The two columns do not merely differ in rate. The volatility error falls like 2/n and can be driven as low as the data allows. The drift error does not move at all — not slowly, but exactly not at all — because the estimator (XT−X0)/T is a function of the endpoints and the intermediate samples are not merely uninformative about the drift, they do not appear in it. Its standard error is σ/T regardless of n.

Only a longer window helps: sixteen times the history quarters the drift error, the same 1/T law applied to T instead of n.

One caveat, which the next section is about: the volatility column says what a diffusion gives up under sampling, and a price is not a diffusion when looked at closely enough. The improvement stops, and then reverses.

Structure (The division of labour is exact).

Put the two facts together and they interlock neatly.

The parameter a price history estimates well — arbitrarily well, given enough sampling frequency — is the diffusion coefficient. The parameter it estimates not at all over a fixed window is the drift. And by Girsanov the diffusion coefficient is precisely the one shared with the pricing measure, while the drift is precisely the one no-arbitrage has already determined there.

The payoff is a constructive answer to this chapter’s own complaint. Every under-identified parameter these notes have found is a parameter of the diffusion structure — β in chapter 10, the mixing weight in chapter 11, the mean reversion in chapter 12, the correlation between forwards in chapter 13. All of them are measure-invariant. So all of them are in principle visible in the time series, which is a different dataset from the one that failed to determine them, and Calculation 21.4 says the time series is the good direction to look. When this chapter says the fix for an unidentified parameter is more information, this is where the information is.

Remark (Three reasons it is harder than that).

The argument above is correct and it is not a recipe. Three things stand between it and a working joint calibration.

The processes are latent. Quadratic variation identifies the volatility of something observed. The parameters that are unidentified belong mostly to processes that are not: the variance process itself in a stochastic volatility model, which regime is in force, the state of a multi-factor curve. Estimating those from prices is a filtering problem rather than a sum of squared increments, which is why the particle filter appears earlier in this chapter, and the rates degrade accordingly.

Some of them are drift parameters of a latent process. A mean reversion is the drift of the variance or of the curve factor, so it inherits Calculation 21.4’s bad column: it needs a long history rather than a finely sampled one, and this chapter has already measured that a mean reversion fitted to a finite sample comes out biased fast. Sampling faster does not touch it.

The over-identifying restriction can fail, and its failure is information. If both sources speak about σ, they can disagree — and they routinely do, because implied volatility exceeds realised volatility on average. That gap is not an error in either estimate; it is §21.7’s subject. But it means a joint calibration that assumes the two agree will corrupt a σ the option prices had identified correctly, and one that does not assume it must model the discrepancy.

So the honest summary is narrower than the structure above suggests. Joint calibration adds genuine information about the diffusion structure, which is where these notes keep finding the gaps. It buys that at the price of requiring a model for the risk premium, and it should be run as a test before it is run as a fit: fit to prices alone, then ask what the historical dynamics of the fitted model would look like and whether they resemble the ones observed.

21.5 Estimating a Volatility

The previous section leaned on volatility being the easy parameter. Four decisions have to be made before a number comes out, and each of them moves it by more than the precision anybody quotes.

Over what window

Volatility is not constant, which is the premise of five chapters of these notes, so a sample average over a window is an average of something that moved. A long window has small sampling error and estimates a blend of regimes; a short one tracks the current regime and is noisy. There is no correct answer, only a stated one, and the honest practice is to report the window alongside the number and to check that the conclusion does not depend on it.

The choice interacts with the use. A volatility feeding a risk model over a ten day horizon and a volatility feeding a five year swaption are not the same quantity even if both are called annualised volatility, and using one for the other is the commonest way a risk number and a pricing number come to disagree for reasons nobody can locate.

At what resolution, and the answer is the hedging frequency

The sampling interval looks like a purely statistical choice and it is not. Chapter 5 showed that the profit and loss of a delta hedged position over an interval is

12⁢Γ⁢S2⁢((Δ⁢SS)2−σ2⁢Δ⁢t),

so what a hedger is paid, interval by interval, is the difference between the squared move over their rebalancing interval and the variance they sold. Accumulated over the life of the option, the quantity that determines whether the hedge made or lost money is the realised variance sampled at the rebalancing frequency.

That answers the question. A desk rebalancing daily is exposed to daily realised variance and should estimate it from daily data, and one rebalancing hourly to hourly variance. The two are not the same number — if they were, the signature plot below would be flat — and the difference is not an estimation error to be minimised but a genuine difference in what is being hedged.

The square root is not free

Almost every estimator computes a variance and roots it. The sum of squared increments is unbiased for the variance; the square root is concave, so by Jensen the volatility that comes out is biased low.

Calculation 21.5 (The correction, in closed form).

For n increments of a driftless Gaussian, n⁢σ^2/σ2 is chi-squared with n degrees of freedom, so

𝔼⁢[σ^]=σ⁢2n⁢Γ⁢(n+12)Γ⁢(n2)≈σ⁢(1−14⁢n). (21.3)

At twenty increments — a month of daily data — the factor is 0.98758, so the volatility comes out 1.24% too low, which on a 20% volatility is a quarter of a point.141414estimation::sqrt_bias_correction, with a simulation confirming that dividing by it removes the bias rather than merely that the formula is what it is.

The bias is one-signed, so averaging across months does not remove it, and it is available in closed form, so there is no reason to carry it. It is the smallest of the four difficulties here and the only one with an exact answer.

And a price is not a diffusion when looked at closely

The serious difficulty is the one that contradicts the previous section’s arithmetic. What is observed is not the efficient price but the price at which a trade happened, which alternates between the bid and the offer. Write the observation at each sampled time as Yi=Xi+ϵi, with ϵi independent noise of scale η, roughly the half spread. An observed increment is then Δ⁢Yi=Δ⁢Xi+(ϵi−ϵi−1), a difference of two independent noise draws rather than one, so its noise variance is 2⁢η2 regardless of how short the sampling interval is:

𝔼⁢[realised variance over ⁢n⁢ samples]=σ2⁢T+2⁢n⁢η2. (21.4)

The bias is linear in the sampling frequency. Sampling faster does not converge on the truth, it diverges from it. Consecutive increments also share a noise draw with opposite sign — ϵi enters Δ⁢Yi positively and Δ⁢Yi+1 negatively — which induces a negative autocorrelation between adjacent returns, the same dependence that later stops sampling faster from helping a regime be detected any sooner.

Calculation 21.6 (The volatility signature plot).

A 20% volatility over one trading day, observed with a one basis point half spread.151515estimation::realised_variance_with_noise, measured against (21.4) rather than against intuition.

Samples in the day Measured volatility Error
6 (hourly) 19.09% −4.5%
26 (quarter hourly) 19.86% −0.7%
78 (five minutes) 20.05% +0.3%
390 (one minute) 20.48% +2.4%
1 560 (fifteen seconds) 21.87% +9.3%
7 800 (two seconds) 28.16% +40.8%

At two second sampling the measured volatility is half as large again as the truth, and it is not converging — it is measuring the spread.

Structure (Two biases pointing opposite ways).

The table is not monotone, and the reason is the useful part.

At coarse sampling (21.3) dominates: few increments, so the square root pulls the estimate down, and hourly sampling is four and a half per cent low. At fine sampling (21.4) dominates and the estimate runs away upwards. The two biases point in opposite directions and there is a frequency where they cancel — here around a quarter of an hour, and the five minute sampling that the literature recommends is within a fraction of a per cent of the truth for the same reason.

So the standard advice to sample every five minutes is not a rule of thumb about market conventions. It is where two known biases of known size happen to cross for typical parameters, and (21.3) and (21.4) say how it moves: a wider spread pushes the optimum coarser, a shorter horizon pushes it finer. A desk trading an instrument with a ten basis point spread and using five minute sampling because that is what the papers say is making an error of a hundred times the size the papers were addressing.

It also settles the previous section’s overreach. The claim that a fixed history identifies volatility arbitrarily well is a statement about a diffusion. Against real observations the error falls, reaches a floor set by the spread, and then rises, so the identification argument holds down to the noise and no further — which is still a great deal better than the drift, whose error never falls at all.

Remark (What is done about it).

Three responses, in increasing order of effort. Sample at the frequency where the biases cross and accept the residual, which is what most desks do and which Calculation 21.6 shows is defensible. Subsample and average — compute the realised variance on several offset grids and average them, which reduces the variance without changing the bias. Or use an estimator built for the problem, and there are two standard ones, both exploiting the same fact: the bias in (21.4) is exactly proportional to the increment count, so two realised variances computed on the same path at two different frequencies carry the same noise process in a known, fixed ratio.

Two-scale realised volatility takes that ratio literally. Average the realised variance over k coarse, offset subsamples of the path to estimate the low-frequency side; take the ordinary realised variance over the whole path for the high-frequency side; and subtract the two in the proportion (21.4) predicts. The noise term cancels algebraically rather than merely being averaged over.161616Zhang, L., Mykland, P. A., & Aït-Sahalia, Y. (2005). A tale of two time scales. Journal of the American Statistical Association, 100(472), 1394–1411.

A realised kernel works on the same increments directly. Ordinary realised variance is the increments’ own sample autocovariance at lag zero; the noise contributes a matching autocovariance at lag one, of the opposite sign, because adjacent increments share a noise draw with opposite signs, as noted above. Adding the higher lags back in, at a weight that shrinks smoothly to zero by some bandwidth, cancels the bias instead of leaving it in — the negative autocorrelation used constructively rather than merely diagnosed.171717Barndorff-Nielsen, O. E., Hansen, P. R., Lunde, A., & Shephard, N. (2008). Designing realised kernels to measure the ex post variation of equity prices in the presence of noise. Econometrica, 76(6), 1481–1536.

The third route is the correct one: the noise is a known functional form with an unknown scale, so it can be estimated and removed rather than avoided. Equation (21.4) has η in it, and anything with a parameter in it can be fitted.

Calculation 21.7 (What the two estimators buy, on the same path).

The same simulated day as Calculation 21.6, now measured three ways at once: naive realised variance, two-scale realised volatility, and a realised kernel.181818estimation::noise_robust_estimators, with both staying within a few percent of the truth where the naive estimator does not.

Samples in the day Naive Two-scale Kernel
26 (quarter hourly) 20.06% 19.31% 19.95%
78 (five minutes) 20.12% 19.82% 20.08%
390 (one minute) 20.50% 19.93% 19.99%
7 800 (two seconds) 28.16% 20.00% 20.07%

At the finest sampling, where the naive estimator is measuring the spread rather than the volatility, both alternatives are within a tenth of a percentage point of the truth — the whole point of building an estimator around the bias rather than merely choosing a frequency to tolerate it at.

21.6 Estimating a Correlation

Chapter 13 produced an uncomfortable result: the correlation between rates moves a swaption volatility by percentage points, while the choice between two whole model families moves it by hundredths. If that is right — and it is measured — then the correlation estimate is the most consequential number in the whole apparatus, and it deserves more care than it usually gets.

It is also the hardest thing in this chapter to estimate, for a reason that is dimensional rather than statistical.

Calculation 21.8 (A correlation matrix of pure noise).

Take forty independent series — the forwards of a ten year quarterly structure, if they had no correlation at all — and estimate their correlation matrix from one year of daily observations. The truth is the identity, so every eigenvalue is one. The estimate is not close.191919estimation::sample_correlation_spectrum, against the Marchenko-Pastur edges rather than against an impression.

Series Observations Largest Smallest Ratio
40 250 1.84 0.39 4.7×
40 1 000 1.40 0.66 2.1×
40 4 000 1.19 0.83 1.4×

At a year of data the largest apparent factor looks nearly five times the smallest, in a matrix with no factors in it whatever. Anyone reading a scree plot of those eigenvalues would report a rich factor structure and would be reporting nothing.

Remark (The spread is predictable, which is what makes it usable).

The numbers above are not accidents of the simulation. Estimating p series from n observations means extracting p⁢(p−1)/2 numbers from p⁢n data points, and the ratio q=p/n controls everything. The Marchenko-Pastur law says the sample eigenvalues of independent series fill the interval

[(1−q)2,(1+q)2], (21.5)

and the measured columns above sit against those edges to within a few per cent.

That is more useful than a warning, because it supplies a null hypothesis. An eigenvalue inside (21.5) is consistent with noise and carries no information; one above the upper edge is a factor that is really there. So a scree plot becomes a test rather than an impression: draw the upper edge on it, and count what pokes out. On the US Treasury curve — the only one this book’s committed data can put through the test — exactly two eigenvalues clear the edge: level and slope, on every panel tested, including four decades of daily observations spanning several rate cycles rather than one calendar year.202020risk::only_level_and_slope_clear_the_noise_edge_on_the_full_panel and risk::only_level_and_slope_clear_the_noise_edge_across_four_decades. That is chapter 24’s level and slope, now established rather than asserted. The third shape a curve’s principal component analysis always returns, curvature, sits inside the noise bulk on every panel tested; it is a real direction to trade, as chapter 22’s butterfly shows, but not one this test can distinguish from what forty-four years of pure noise would also produce. Whether a curve with more cross-sectional structure — a swap curve segmented by credit and liquidity tier, say — would put a third eigenvalue above the edge is a real question this book’s data cannot answer; the count of two is a property of the government curve, not asserted as a property of curves in general.

The remedy follows the diagnosis. Keep the eigenvalues above the edge, replace the bulk with its average, and rebuild the matrix; or shrink the sample matrix towards a structured target, which is the same operation performed continuously rather than by a cutoff. Both are ways of refusing to estimate what the data cannot support.

Structure (Three routes, and the market model takes the third).

There are only three things one can do about Calculation 21.8.

More data. It works and it works slowly: the excess of the top eigenvalue over one falls like q, so sixteen times the history roughly quarters it. Sixteen years of daily data also spans several regimes, so what is gained in sampling error is lost in the assumption that the correlation was constant — which returns to the window problem of the previous section, now with sharper consequences.

Cleaning. Estimate the matrix, then remove the part attributable to noise, by eigenvalue clipping or by shrinkage towards a target. This is the honest general-purpose answer and it makes no structural assumption beyond the choice of target.

Parametrise. Assume a form. Chapter 13 uses ρi⁢j=e−β⁢|Ti−Tj|, one parameter for the whole matrix, which cannot overfit because there is nothing to overfit with. The cost is that it is wrong in a specific way — real curve correlations do not decay to zero at long separations, they flatten out at a positive level, so the exponential understates the coupling between the long end and the very long end. A two parameter form ρi⁢j=ρ∞+(1−ρ∞)⁢e−β⁢|Ti−Tj| fixes exactly that and is what is generally used.

The third route is the right one here and the reason is Calculation 21.8 rather than parsimony for its own sake. Forty forwards from a year of data cannot support a free matrix; the estimate would be noise wearing the shape of structure, and chapter 13 showed that this is the parameter the price is most sensitive to. When the data cannot identify a parameter and the answer depends on it, imposing a form is not laziness — it is the only alternative to letting the noise choose.

Calculation 21.9 (Checking the form against the same test that counted the factors).

A two parameter form is chosen for parsimony, but Calculation 21.8’s own machinery can ask a sharper question: does it match the shape a principal component analysis finds, or only survive by having too few parameters to fail? An equicorrelation floor ρ∞ is, by itself, a rank-one matrix with a flat eigenvector — a level factor. An exponential decay on top of it contributes a second, maturity-ordered direction — a slope. So the two parameter form’s own structure predicts exactly two real dimensions, which is either a coincidence or a genuine match to the remark above’s finding of exactly two, on four decades of Treasury data.

Fit ρ∞ and β to the sample correlation matrix of daily changes across the committed eight tenors, by least squares on the off-diagonal entries, then compare the fitted matrix’s own eigenvalues to the real ones.212121risk::fit_maturity_correlation, with the comparison measured against the same panel Calculation 21.8 and the noise-edge remark use.

Real Fitted, ρ∞=0.51, β=0.19
Level 6.28 6.40
Slope 1.12 0.80
Third (noise) 0.27 0.34

The level eigenvalue is matched to within two per cent, and neither matrix puts a third eigenvalue above the noise edge of 1.05 — the rank is right, not merely small enough to be safe. The slope eigenvalue is not matched: the fit understates it by about a quarter. A single exponential has one parameter left after ρ∞ pins the floor, and that parameter cannot simultaneously set the decay’s shape and match how much independent variance the real slope carries — so the form gets the number of real factors right and gets one of their sizes wrong, which is a more precise criticism than “it is only an approximation” and a more precise defence than “it captures the two real factors.”

Remark (And it is not constant, in the direction that hurts).

One further problem, which no amount of cleaning addresses. Correlations move, and they move up in stress: rates that normally decouple stop decoupling when everything is being sold at once.

So an estimate from a calm sample understates the correlation in exactly the scenario a risk calculation is about, and chapter 18 has already made the general form of this point — the coupling in the tail is not the coupling in the body, and a single correlation number cannot carry both. For pricing, the calm number is arguably the right one, since it is the average over the option’s life that matters. For risk it is not, and chapter 24’s stress work has to use something else. Reporting one number for both uses is the error, rather than any particular estimate of it.

21.7 Why Implied Exceeds Realised

The gap is large, persistent and one-signed, so it is not a misfit to be calibrated away. Two explanations are usually offered and they are frequently presented as alternatives. They are not, and separating them matters for §21.4, because they imply different things about whether a joint calibration is well posed.

It is a risk premium, and it sits where Girsanov says it must

The first explanation is that variance is a priced risk factor. A position that is long volatility pays off when volatility spikes, and volatility spikes when everything else is going wrong, so it is a hedge — and a hedge earns a negative expected return, which is to say it costs a premium to hold. Options are that hedge, so they are expensive relative to what the underlying goes on to do.

This is formally the same object as the equity risk premium and belongs in the same place. Under the historical measure a stock drifts faster than a bond because holding it is uncomfortable; under the pricing measure it does not. Here the variance process drifts differently under the two measures for the same reason.

Calculation 21.10 (The premium, written down).

Take the Heston variance of chapter 10 under the historical measure, and let the market price of variance risk be proportional to volatility, so that changing measure adds −λ⁢v⁢d⁢t to the variance’s drift. Then

κℚ=κ+λ⁢η,θℚ=κ⁢θκℚ, (21.6)

and η is unchanged. Take a 20% long-run volatility, mean reversion κ=2, vol-of-vol η=0.3, and the variance starting at its own long-run level so that nothing below is confused with a transient:222222heston::variance_premium, with the expected integrated variance in closed form rather than simulated.

λ κℚ One year implied Premium
−1 1.70 20.90% 0.90 pts
−2 1.40 21.89% 1.89 pts
−3 1.10 23.00% 3.00 pts

against a realised expectation of exactly 20%. The observed gap on an equity index is one to three volatility points, so a market price of variance risk of a few units reproduces it — the explanation is the right size, which is the first thing to check about any explanation.

Structure (The premium cannot be in the vol-of-vol).

If implied volatility carries a risk premium, it is tempting to suppose that vol-of-vol carries one too — that η is higher under the pricing measure the way θ is. It is not, and it cannot be. η is a diffusion coefficient, and a change of measure moves drifts and leaves diffusion coefficients alone. The premium has nowhere to go except into κ and θ.

That is §21.4’s division of labour applied one level down, and it has a consequence that is directly useful. Vol-of-vol is measure-invariant, so it is jointly identified: an η estimated from the time series of realised volatility and an η implied by the curvature of the smile are estimates of the same number, and they can be compared. If they disagree, something is wrong with the model rather than with the premium, because no premium can explain a discrepancy in a diffusion coefficient. This is one of the few genuinely testable cross-measure restrictions in derivatives pricing, and it is almost never tested.

It is one-sided flow, and dealers cannot arbitrage it away

The second explanation makes no reference to risk aversion. Banks are structurally short options, because their clients — corporates hedging liabilities, insurers hedging guarantees, funds buying protection — want to own them. A dealer who cannot lay the position off must hold it, and chapter 23 derives what happens next: a market maker holding inventory it did not choose quotes to be paid for holding it. The premium is then a price for absorbing one-sided demand, in exactly the sense chapter 22 requires of a trade — somebody is on the other side for a reason that is not an opinion.

This explanation is not a rival to the first. It is a second premium, paid for a different service, and both are collected by the same position. A single calibrated λ cannot tell them apart, though: matching the whole observed gap makes λ absorb whatever share of it is flow, not only the risk premium it is meant to represent, and the remark on telling them apart below says what to check before trusting it.

And liquidity, which is why neither gets arbitraged away

A premium of one to three volatility points invites the obvious question: why does nobody sell it until it is gone? The answer is the third component, and it is not a separate explanation so much as the reason the first two survive.

Harvesting the premium means selling options and delta hedging to expiry. The round trip costs the option’s bid-offer, which on a liquid index swaption is a fraction of a volatility point and on anything else is more, plus the accumulated cost of rebalancing the delta, which chapter 23 prices and which grows with how often one rebalances — and Calculation 21.6 has just shown that rebalancing frequency is not a free choice either. A premium smaller than the cost of collecting it is not an arbitrage; it is a fee for a service that is genuinely expensive to provide.

That produces a floor. A model error smaller than the bid-offer spread is not a model error anyone can trade, so it cannot be arbitraged away and there is no force driving it to zero. Chapter 13 used exactly this to settle the choice of coordinates: an approximation two orders of magnitude below the spread is free. The same reasoning says the variance premium, being larger than the spread but not by much, is exactly in the range where it persists — big enough to be real, small enough that collecting it is a business rather than an arbitrage.

Liquidity also enters as a premium in its own right, and in the same direction. An instrument that cannot be exited cheaply trades at a discount, which is to say a higher expected return, and chapter 22’s on-the-run spread is that discount measured directly. For options the effect concentrates where it is least convenient: far strikes and long tenors, which have the widest spreads, the thinnest volume and the least reliable marks — and which are also where a model is least constrained by quotes. So the part of the surface where calibration has the least to say is the part where the price is least trustworthy, and the two failures compound rather than offset.

Remark (Volume is a poor proxy and it is the one everybody uses).

Since liquidity is not directly observable, it is proxied — usually by volume, sometimes by quoted spread, occasionally by the price impact measures of chapter 23. Volume is the worst of the three and the most used, because it is high in two opposite circumstances: when an instrument is easy to trade, and when everybody is trying to get out of it at once. A liquidity signal built on volume therefore reads its highest exactly when liquidity is about to be worst, which is the failure mode chapter 22’s six trades all share.

Quoted spread is better and still incomplete, since a quoted spread is for a quoted size and the size is what matters in a stress. The honest measure is the one chapter 23 builds, price impact per unit of quantity, and it requires data most participants do not have.

Remark (How to tell them apart, and why it matters here).

The two have different signatures, and the distinction is testable rather than philosophical.

A risk premium is compensation for an exposure, so it should be roughly proportional to the quantity of variance risk borne and should have a term structure fixed by the model — Calculation 21.10 shows it emerging over the mean reversion time, so it grows with horizon in a shape κ determines. It should be present in every market where variance is risky, in proportion to how risky.

A flow premium is compensation for absorbing inventory, so it should track dealer capacity rather than risk: largest where client demand concentrates, which is downside strikes and specific tenors rather than uniformly across the surface; widening when balance sheet is constrained, at quarter ends and after dealer losses; and varying across markets with who happens to be buying rather than with how volatile they are.

Both signatures are visible in the data, which is the honest answer to which explanation is right: there is a persistent baseline consistent with a risk premium, and a strike- and tenor-dependent component that moves with demand.

For §21.4 that distinction is decisive rather than academic. If the gap is a risk premium, λ is a smooth function of the state with few parameters, it can be modelled, and a joint calibration to prices and history is well posed — one estimates the shared diffusion structure and the premium together. If the gap is flow, then the quantity separating the two measures depends on dealer positioning, which appears in neither dataset. A joint fit will then attribute a flow effect to dynamics, and the dynamics is what it was trying to learn. So the programme’s viability rests on how much of the gap is which, and that question has to be answered before the fit rather than by it.

Some of that positioning is not entirely unobservable, only absent from these two datasets specifically. Exchange-reported open interest and volume by strike and tenor proxies for where client demand concentrates; dealer CDS spreads and balance-sheet disclosures proxy for the capacity constraint directly, since a dealer’s own cost of capital is what widens at quarter ends and after losses; and CFTC positioning reports do the same on the futures side of the same trade. None of it substitutes for prices and history — open interest in particular inherits the remark above on volume, reading highest exactly when it is least informative — but each is a genuinely third source, carrying the one thing prices and history structurally lack, and answering the risk-versus-flow question before the fit has to be built from data like this rather than from more of what the fit already has.

The split is worth making because the two carry different consequences, not only different causes. A risk premium should persist wherever variance is priced, so a book built to harvest it can be sized and stressed like any other risk exposure; a flow premium can vanish with the client demand or dealer capacity that created it, so the same one to three points of income calls for a liquidity stress rather than a volatility one, and for sizing that tracks the business rather than the risk. The genuine way to do the split is to regress the observed gap, by strike and tenor, on the risk premium’s own term structure alongside the capacity and demand proxies above, and let the two explain what they actually explain.

21.8 What To Take Away

Three things carry over from chapter 20, sharpened by what a latent state adds.

First, speed is not knowledge. A neural surrogate for a pricing map, a particle filter, a posterior sampled instead of optimised — each turns a slow calculation into a fast one, and none of them extracts information the quotes or the history did not contain. The flat valley in a SABR fit is exactly as flat once the fit is instant.

Second, not everything travels between the two measures, but more does than is usually assumed. A diffusion coefficient is the same object under ℙ and ℚ, so it can be estimated from history and checked against what is implied; a drift, a mean reversion, or the risk premium in the gap between implied and realised volatility cannot make that trip, because a change of measure is precisely what moves them.

Third, a latent state is usually a smaller problem than it is made out to be, and a real one is a race. Continuous, noiseless observation of a price would reveal its own quadratic variation and, with it, the variance a model called hidden — but no real market grants that for free. Microstructure noise floors the estimate long before the sampling interval reaches zero; a jump adds its own squared size to quadratic variation and must be separated from the diffusive part rather than counted as more of it; and what quadratic variation converges to is itself a different object at every scale, so sampling faster answers a question about a finer regime rather than the coarser one more precisely (Calculation 21.3’s third consequence, and the hedging-frequency argument earlier in this chapter). What remains genuinely unknown is what discreteness, noise, jumps, and scale destroy — and a change worth modelling is one the data can tell apart from noise faster than the change itself unfolds, at the resolution the change is actually defined at.

The tools for computing have changed enormously. The questions have not.

References

  • -

    Doucet, A., de Freitas, N., & Gordon, N. (2001). Sequential Monte Carlo Methods in Practice. Springer.

  • -

    Horvath, B., Muguruza, A., & Tomas, M. (2021). Deep learning volatility. Quantitative Finance, 21(1), 11–27.

  • -

    Buehler, H., Gonon, L., Teichmann, J., & Wood, B. (2019). Deep hedging. Quantitative Finance, 19(8), 1271–1291.

  • -

    Han, J., Jentzen, A., & E, W. (2018). Solving high-dimensional partial differential equations using deep learning. PNAS, 115(34), 8505–8510.

  • -

    Guyon, J., & Lekeufack, J. (2023). Volatility is (mostly) path-dependent. Quantitative Finance, 23(9), 1221–1258.

  • -

    Perez Arribas, I., Salvi, C., & Szpruch, L. (2020). Sig-SDEs model for quantitative finance. Proceedings of the First ACM International Conference on AI in Finance.

  • -

    Zhang, L., Mykland, P. A., & Aït-Sahalia, Y. (2005). A tale of two time scales. Journal of the American Statistical Association, 100(472), 1394–1411.

  • -

    Barndorff-Nielsen, O. E., Hansen, P. R., Lunde, A., & Shephard, N. (2008). Designing realised kernels to measure the ex post variation of equity prices in the presence of noise. Econometrica, 76(6), 1481–1536.