Skip to content
Sarthak Bagaria
All notes

Chapter 18 Dependence

Chapter 16 left a gap: single-name curves pin every marginal default distribution and say nothing at all about the joint one. This chapter is about what to put in that gap. It builds the tools that describe dependence — Sklar’s theorem, rank correlation, and the tail dependence coefficient that turns out to be the number most portfolio instruments are actually a bet on — and then gives four things to use, in increasing order of how much structure they demand. The Gaussian copula appears last, as a diagnosis rather than a target: it is a coherent model whose one defining property is that it sets to zero the exact quantity a senior tranche is priced off, and we finish by measuring what that costs.

18.1 The Gap, Stated Precisely

A basket of n names, each with a default time τi whose distribution Fi chapter 16 bootstrapped from that name’s own CDS quotes. Anything depending on more than one name — a first-to-default swap, a tranche, the counterparty exposure of chapter 24 — depends on the joint distribution F⁢(t1,…,tn), and the marginals do not determine it.

The following theorem says exactly how much freedom that leaves.

Theorem 18.1 (Sklar).

For any joint distribution F with marginals F1,…,Fn there is a function C:[0,1]n→[0,1] with

F⁢(t1,…,tn)=C⁢(F1⁢(t1),…,Fn⁢(tn)), (18.1)

and C is itself a distribution function on the unit cube with uniform marginals. It is unique wherever the marginals are continuous. Conversely, any such C combined with any marginals gives a valid joint distribution.

Proof.

With continuous marginals the copula is not merely shown to exist, it is written down.

The probability integral transform is the whole idea: if Fi is continuous then Ui=Fi⁢(τi) is uniform on [0,1], since ℙ⁢(Fi⁢(τi)≤u)=ℙ⁢(τi≤Fi−1⁢(u))=u. So define

C⁢(u1,…,un)=F⁢(F1−1⁢(u1),…,Fn−1⁢(un)), (18.2)

which is the joint distribution of (U1,…,Un). It is a distribution function because F is; its marginals are uniform by the transform; and substituting ui=Fi⁢(ti) into (18.2) recovers (18.1). Uniqueness is immediate from the same substitution: (18.1) determines C at every point of the form (F1⁢(t1),…,Fn⁢(tn)), and continuity of the marginals makes those points the whole cube.

The converse needs no work at all. Given any copula C and any marginals, the composition C⁢(F1⁢(t1),…,Fn⁢(tn)) is a composition of a distribution function with non-decreasing right-continuous maps, so it is again a distribution function, and setting all but one argument to their upper limits returns Fi. This half is what does the damage in this chapter: it says the construction never fails, so no choice of C can ever be ruled out by the marginals.

When a marginal has an atom the argument breaks at exactly one point — Fi is no longer invertible, and (18.1) constrains C only on the range of Fi, which now has gaps. Existence survives, by interpolating across the gaps, and uniqueness does not. That is the whole content of the continuity hypothesis, and it is why a default-time distribution with a lump of probability at a coupon date is a case where “the” copula is not well defined. ∎

Remark (What it means).

A joint distribution factors cleanly into two independent pieces: the marginals, which say how each name behaves on its own, and the copula C, which says how they are coupled. The two can be chosen separately, and every choice is legitimate.

The intuition is a change of coordinates. Fi⁢(τi) is uniform on [0,1] whatever τi was — this is the probability integral transform — so applying Fi to each name strips out everything specific to that name and leaves only its rank. The copula is the joint distribution of the ranks. It is what remains of the dependence after every marginal has been standardised away.

Sklar’s theorem is usually presented as a technical result. For our purposes it is the precise statement of chapter 16’s complaint: the market gives us the Fi and says nothing whatever about C, and by the converse half of the theorem, every C is consistent with the quotes. Choosing one is not calibration. It is an assumption, and it should be argued for rather than defaulted into.

Example 18.1 (The two extremes).

Two names, each defaulting within five years with probability 10%. What is the probability both do?

If they are independent, C⁢(u,v)=u⁢v and the answer is 1%. If they are comonotone — one defaults exactly when the other does — then C⁢(u,v)=min⁡(u,v) and the answer is 10%. Both are consistent with identical CDS quotes on both names. A tenfold range in the joint probability, and nothing in the single-name market narrows it by a basis point.

18.2 Correlation Is the Wrong Word

The industry names this problem “correlation”. That is a poor choice.

Remark (Correlation is not invariant).

Linear correlation is a property of the variables, not of their copula. Apply a strictly increasing function to one of them — take a log, or a square — and the copula is unchanged by construction, since ranks are unchanged, while the correlation moves. So correlation mixes up the dependence with the marginals, which is exactly the separation Theorem 18.1 was useful for making.

Rank statistics do not have this problem. Kendall’s τ and Spearman’s ρ depend only on the copula, and are the right things to quote if a single number must be quoted.

Theorem 18.2 (Attainable correlations).

Let X=eσ1⁢Z1 and Y=eσ2⁢Z2, with Z1 and Z2 standard normal. Then whatever the joint law of (Z1,Z2),

e−σ1⁢σ2−1(eσ12−1)⁢(eσ22−1)≤Corr⁢(X,Y)≤eσ1⁢σ2−1(eσ12−1)⁢(eσ22−1). (18.3)
Proof.

The marginals fix the means and variances, so only 𝔼⁢[X⁢Y] varies with the coupling, and the correlation is increasing in it. The Hoeffding-Frechet bounds say the extremes of 𝔼⁢[X⁢Y] over joint laws with given marginals are attained by the comonotone coupling, C=min⁡(u,v), and the countermonotone one, C=max⁡(u+v−1,0). For standard normals these are Z2=Z1 and Z2=−Z1: the first sends both to the same quantile, and Φ⁢(−z)=1−Φ⁢(z) makes the second send them to opposite quantiles.

Now compute. For a standard normal Z, 𝔼⁢[es⁢Z]=es2/2, so

𝔼⁢[X]=eσ12/2,Var⁡(X)=𝔼⁢[X2]−𝔼⁢[X]2=e2⁢σ12−eσ12=eσ12⁢(eσ12−1),

and likewise for Y. With Z2=ε⁢Z1 for ε=±1, the product is X⁢Y=e(σ1+ε⁢σ2)⁢Z1, so

𝔼⁢[X⁢Y]=e(σ1+ε⁢σ2)2/2=e(σ12+σ22)/2⁢eε⁢σ1⁢σ2⟹Cov⁡(X,Y)=e(σ12+σ22)/2⁢(eε⁢σ1⁢σ2−1),

using 𝔼⁢[X]⁢𝔼⁢[Y]=e(σ12+σ22)/2. Dividing by Var⁡(X)⁢Var⁡(Y)=e(σ12+σ22)/2⁢(eσ12−1)⁢(eσ22−1), the exponential prefactor cancels and

Corr⁡(X,Y)=eε⁢σ1⁢σ2−1(eσ12−1)⁢(eσ22−1),

which is the upper end of (18.3) for ε=+1 and the lower end for ε=−1. ∎

Remark (The Gaussian copula’s correlation is not the Pearson correlation).

The same computation gives the whole range and not just its ends. If Z2=ρ⁢Z1+1−ρ2⁢W with W an independent standard normal, which is the Gaussian copula with parameter ρ, then 𝔼⁢[X⁢Y]=e(σ12+σ22+2⁢ρ⁢σ1⁢σ2)/2, and

Corr⁢(X,Y)=eρ⁢σ1⁢σ2−1(eσ12−1)⁢(eσ22−1). (18.4)

This is increasing in ρ, and ρ=±1 are the two ends of (18.3). The parameter ρ of the Gaussian copula is always attainable and the joint law it defines is always valid. The Pearson correlation of the two lognormals it produces is (18.4), which is not ρ, and for volatile positions is far closer to zero. Reading one number as the other is the mistake, and no joint law is impossible until that is done.

Inverting (18.4) turns a target Pearson correlation into a copula parameter, ρ=ln⁡(1+c⁢(eσ12−1)⁢(eσ22−1))/(σ1⁢σ2), and this fails exactly when the target c is outside (18.3): the logarithm has no argument below the lower bound, and ρ leaves [−1,1] above the upper one. It is the same bounds theorem, now stated for the copula parameter.111simulating the coupling and measuring the Pearson correlation; the two ends and the monotonicity; the inversion, and where it fails.

Example 18.2 (The bounds are not decorative).

Evaluating (18.3) numerically:222quant/src/dependence.rs.

Volatilities Lowest attainable Highest attainable
25% and 25% −0.94 +1.00
100% and 100% −0.37 +1.00
200% and 200% −0.02 +1.00
50% and 200% −0.16 +0.44

Read the last two rows. Two lognormals at 200% volatility cannot be more than two percent negatively correlated no matter how they are coupled, because the closest they can come to opposed is e2⁢Z against e−2⁢Z, whose product is identically one. They are opposed in rank, and yet their correlation is −e−σ2≈−0.02. Each has a mean of about 7.4 and a standard deviation of about 54, and that variance is carried by rare huge values, which occur on opposite sides: one is enormous only when the other is tiny, so the covariance 1−e4 is small next to the variance e4⁢(e4−1) that the tails create.333the closed form, and the mean and standard deviation quoted. And a 50% name and a 200% name cannot exceed 0.44 however tightly they are bound, because the fatter-tailed one has a variance the thinner one cannot track.

A system asked for a Pearson correlation of −0.5 between two 200% volatility positions has been asked for something no joint distribution can produce, and a careful one fails to calibrate. A system that takes the same −0.5 as the parameter of a Gaussian copula builds a valid joint law without complaint, and (18.4) says the Pearson correlation of the positions it implies is −0.016.444the number, pinned.

Remark (And zero correlation is not independence).

The familiar one, included because it is the reason the next section exists. Correlation measures one particular linear summary of the joint law, and there are many ways to be strongly dependent with none of it. In particular, two variables can be uncorrelated and still almost always crash together, which is the case a portfolio of credits actually lives in.

18.3 The Number That Actually Matters

The failures above are diagnostic. This is the one that decides prices.

Definition 18.3 (Tail dependence).

The coefficient of upper tail dependence of a copula is

λ=limu→1−ℙ⁢(U>u⁢∣V>⁢u), (18.5)

where U,V are the two uniform coordinates. It is the probability that one variable is extreme, given that the other is, in the limit of extreme.

Remark (Why this and not correlation).

Consider what a senior tranche is. It absorbs nothing until the portfolio has already lost, say, fifteen percent, which requires a large fraction of the names to have failed together. Its entire value is the probability of a joint extreme. It is a bet on λ and on almost nothing else.

Correlation, by contrast, is an average over the whole distribution, dominated by the middle where most of the probability sits. Two copulas can agree on correlation to three decimal places, agree on every marginal exactly, and disagree about λ completely — and a senior tranche will price off the disagreement rather than the agreement.

Theorem 18.4 (Tail dependence of the two standard copulas).

The Gaussian copula with correlation ρ<1 has λ=0. The Student-t copula with correlation ρ and ν degrees of freedom has

λ=2⁢tν+1⁢(−(ν+1)⁢(1−ρ)1+ρ)>0 (18.6)

for every ρ>−1 and every finite ν.

Proof.

Both halves come from the same question — conditional on one variable being far out in the tail, how far out is the other — and the two answers differ for one structural reason, which the calculation makes visible.

The reduction. Both copulas are unchanged by reflecting both coordinates, so the upper tail coefficient equals the lower one, and in terms of the standardised variables X,Y (normal or t, both symmetric),

λ=limx→−∞ℙ⁢(Y⁢<x∣⁢X<x)=limx→−∞ℙ⁢(X<x,Y<x)ℙ⁢(X<x).

Numerator and denominator both vanish, so by L’Hôpital the limit is the ratio of their derivatives in x. The denominator’s is the density f⁢(x). The numerator’s corner (x,x) moves in both coordinates,

dd⁢x⁢ℙ⁢(X<x,Y<x)=f⁢(x)⁢ℙ⁢(Y⁢<x∣⁢X=x)+f⁢(x)⁢ℙ⁢(X⁢<x∣⁢Y=x)=2⁢f⁢(x)⁢ℙ⁢(Y⁢<x∣⁢X=x),

the two terms being equal because the pair is exchangeable. So λ=2⁢limx→−∞ℙ⁢(Y⁢<x∣⁢X=x), a statement about the conditional law of Y given X=x.555the joint tail by quadrature at x=−1000, against twice the conditional probability.

The Gaussian case. Write (X,Y) as standard bivariate normal with correlation ρ and condition on X=x. Then

Y=ρ⁢x+1−ρ2⁢ϵ,ϵ∼N⁢(0,1)⁢ independent of ⁢X,

so λ is twice the limit of ℙ⁢(Y⁢<x∣⁢X=x) as x→−∞, which is

ℙ⁢(ϵ<x⁢1−ρ1−ρ2)=Φ⁢(−|x|⁢1−ρ1+ρ)⟶ 0

for any ρ<1, since the argument goes to −∞. The conditioning moved Y’s mean out by ρ⁢x, but the residual ϵ has a fixed scale, and a fixed-scale Gaussian cannot keep up with a barrier receding at rate x. So the second name is dragged part of the way into the tail and left there. Only at ρ=1 does the drag become complete.

The Student case. The t copula is the Gaussian construction with one change: the pair is divided by an independent random scale,

(X,Y)=νW⁢(Z1,Z2),W∼χν2,

with (Z1,Z2) Gaussian of correlation ρ. Now conditioning on X being extreme is informative in a way it was not before: a large |X| is evidence not only about Z1 but about W being small, and a small W inflates Y as well. The tail event is explained by the common factor rather than by the individual one, and the common factor moves both names together.

Making that quantitative needs the conditional law of Y given X=x, which for the bivariate t is again a t. Three steps show it.

What X=x says about W. Given W, X=ν/W⁢Z1 is N⁢(0,ν/W), so the joint density of (X,W) is proportional to

wν/2−1⁢e−w/2⋅w⁢e−w⁢x2/(2⁢ν)=w(ν+1)/2−1⁢exp⁡(−w2⁢(1+x2ν)).

As a function of w this is a gamma density, and it says that given X=x the variable W′=W⁢(ν+x2)/ν is χν+12. So W has one more degree of freedom than before, and it is scaled by ν/(ν+x2): a large |x| makes a small W likely, 𝔼⁢[W∣X=x]=ν⁢(ν+1)/(ν+x2).

What is left of Y. Given W and X=x, Z1=x⁢W/ν is fixed, so

Y=νW⁢(ρ⁢Z1+1−ρ2⁢ϵ)=ρ⁢x+1−ρ2⁢νW⁢ϵ,

where ϵ is standard normal and independent of (X,W).

Recognising a t. Substituting ν/W=(ν+x2)/W′,

Y=ρ⁢x+1−ρ2⁢ν+x2ν+1⋅ϵ⁢ν+1W′,

and ϵ⁢(ν+1)/W′ is by definition a t variable with ν+1 degrees of freedom, ϵ being independent of W′∼χν+12. So given X=x, Y is a tν+1 variable with mean ρ⁢x and scale 1−ρ2 times the factor (ν+x2)/(ν+1). That factor is the information about W: it equals one at x2=1 and grows like |x|, and as ν→∞ it tends to one, which returns the Gaussian case’s fixed scale 1−ρ2.666a simulation of the construction, keeping the draws whose X lies near −2.

Then

ℙ⁢(Y⁢<x∣⁢X=x)=tν+1⁢((x−ρ⁢x)⁢ν+1(1−ρ2)⁢(ν+x2)),

and as x→−∞ the ν+x2 in the denominator grows like |x| and cancels the x in the numerator, leaving the finite limit

tν+1⁢(−(1−ρ)⁢ν+11−ρ2)=tν+1⁢(−(ν+1)⁢(1−ρ)1+ρ).

Twice this limit is (18.6), the factor of two being the one from the reduction above. Note where the cancellation came from: the scale of the conditional distribution grew with |x|, and it grew because the mixing variable W is shared. That is the entire difference between the two copulas. ∎

Structure (The cancellation is the point).

The two calculations are the same calculation with one term changed, and the term is the scale of the conditional law. In the Gaussian case it is constant, so a barrier receding at rate |x| eventually outruns it and the limit is zero. In the Student case it grows like |x|, the two rates match, and a finite limit survives.

The mechanism carries beyond copulas, and it is general: a fixed-scale residual cannot produce tail dependence, and a shared random scale always can. Any model in which the extreme behaviour of several quantities is driven by a common multiplicative factor — a stochastic volatility shared across names, a funding cost that hits every position, a liquidity parameter — has tail dependence for this reason, and any model whose only coupling is through the mean of a fixed-variance residual does not, whatever correlation is fed into it. Chapter 24’s stress work is the practical form of the same observation: the scenario that matters is the one that moves the shared scale.

Remark (Read that again).

The Gaussian statement is not that tail dependence is small, or that it has been approximated away. It is exactly zero, for every correlation short of perfect. Condition on one name being in a one-in-a-million event and the probability that the other is too converges to zero. Joint extremes still occur; they become negligible next to single extremes.

The reason is regression to the mean. Given X=x, the most likely value of Y is ρ⁢x, closer to the centre, and reaching x takes a deviation of (1−ρ)⁢x from a residual whose scale 1−ρ2 is fixed. The cost of that deviation is a probability of order e−x2⁢(1−ρ)/(2⁢(1+ρ)), which vanishes as x grows for every ρ<1. The correlation controls how far along Y is expected to come, and never how far it can go in the tail.

The Student-t fixes it with one extra parameter and no extra conceptual machinery. A t is a normal divided by an independent random scale, and it is that shared scale that produces joint extremes: occasionally the whole system is drawn from a wide distribution, and then everything moves at once. This is not a mathematical trick. It is a common volatility factor, and it is what the world does.

00.20.40.60.800.20.40.60.8CorrelationCoefficient of tail dependence
  • Student-t, ν = 3
  • Student-t, ν = 6
  • Student-t, ν = 15
  • Gaussian, any ν
Figure 18.1: The tail dependence coefficient against correlation. The Gaussian curve is the horizontal axis — not close to it, on it, for every correlation up to one, where it jumps discontinuously to one. The Student-t curves sit well above and rise as the degrees of freedom fall. At a correlation of 0.3 a t with four degrees of freedom gives λ=0.16: condition on one name being in an extreme event and there is a one in six chance the other is as well. The Gaussian model, calibrated to exactly the same correlation, says the chance is zero.
Show the model behind this figure (1 function)
t_tail_dependencequant/src/dependence.rs
/// The coefficient of upper tail dependence of a Student-t copula.
///
/// Positive for every `rho > -1` and every finite `nu`, and it approaches the
/// Gaussian zero only as `nu` grows. Two names with the same correlation and the
/// same marginals can therefore have wildly different probabilities of failing
/// together, which is precisely the freedom the credit chapter said was left open.
pub fn t_tail_dependence(rho: f64, nu: f64) -> f64 {
    if rho >= 1.0 {
        return 1.0;
    }
    let argument = -((nu + 1.0) * (1.0 - rho) / (1.0 + rho)).sqrt();
    2.0 * t_cdf(argument, nu + 1.0)
}

18.4 Four Things To Use Instead

Now the constructive half. These are ordered by how much structure they ask for, and the right choice depends on what is being priced and on whether it has to be hedged.

18.4.1 Give the copula tails

The cheapest fix, and often enough. Replace the Gaussian copula by a Student-t with the same correlation matrix and one degrees-of-freedom parameter. Everything about the implementation is unchanged — the same factor structure, the same simulation, the same calibration to marginals — and λ goes from zero to (18.6).

The ν parameter can be calibrated to the tranche market, or set from history, and history gives two different answers depending on what is fitted. A t fitted by maximum likelihood to a single series measures the tail thickness of that series: EM at each ν for the location and scale, then a search over ν on the profile likelihood. That is not the parameter here. The copula’s ν governs how two names move together in the tail, and it is fitted to the pair after replacing each series by its ranks, so that only the joint behaviour is scored — maximum pseudo-likelihood, with the copula density in place of the t density. On 11,241 daily changes in Treasury yields since 1981, the marginal fits give ν=2.5 for the two year and 4.1 for the ten year, and the copula of the two has ν=5.1 at a correlation of 0.82.777the three fits on the committed panel; recovery of a known ν; recovery of the copula’s ν, with the marginals distorted by a monotone map. The marginals and the copula disagree, which is the reason to keep them apart. For equity indices, fits of a t are commonly reported with ν between three and six, and that range is the usual starting point when a t copula is set by convention; it is a rule of thumb and not something measured here, and a credit book’s own data, or the tranche market, is the better source for its own ν.

Remark (Its limitation, stated honestly).

The t copula is symmetric: it produces joint booms as readily as joint crashes. Credit is not symmetric — names default together far more often than they are jointly upgraded — so a t that is fitted to the downside will overstate the upside. Whether that matters depends on whether anything in the book pays off on the upside. For a tranche, largely not.

18.4.2 Give it asymmetric tails

If the asymmetry matters, the Archimedean family provides it. A Clayton copula has lower tail dependence and none in the upper tail; a Gumbel copula the reverse. Each is a one-parameter family generated by a single convex function, which makes them easy to simulate and easy to reason about.

The cost is that Archimedean copulas in dimension n are exchangeable: they have one parameter for the whole basket and cannot express that two names in the same sector are more tightly coupled than two names in different ones. That is a serious limitation for a real portfolio, and the usual repair — nesting them hierarchically by sector — works but restores much of the complexity that made them attractive.

18.4.3 Model the cause, not the coupling

The first two are still copulas, and every copula shares a defect that the previous section did not mention because it is not about tails.

Remark (A copula has no dynamics).

A copula on default times is a static object. It gives the joint distribution of the τi as seen from today, and it says nothing about how single-name spreads move. The only thing it can update on is which names have defaulted, so between defaults its own prices are deterministic: there is jump-to-default risk in it, and no spread risk. There is no diffusion for a replication argument to cancel, and the dynamic hedging of chapter 5 has nothing to act on. The static content of chapter 4 still applies, since a copula can be arbitrage-free at a point in time, but the hedge ratio of a tranche against spread moves has no source in the model.

This is a deeper problem than the tails. A trader running a tranche book must hedge it with single-name CDS, and the hedge ratio has to come from somewhere. In a copula model it comes from bumping a static parameter and re-running, which is a finite difference of a quantity that was never claimed to be a price process.

The alternative is to model dependence where it comes from. Chapter 2 built default as the first jump of a Cox process with intensity λt; make the intensities share a factor,

λti=ai⁢Yt+Zti, (18.7)

with Y a common process and Zi idiosyncratic ones. Now everything is a process. Names default together because their intensities rise together; the joint distribution is an output rather than an assumption; and the model has a filtration, so a tranche has a genuine delta with respect to the single-name spreads that hedge it.

Remark (What this buys, and what it costs).

It buys dynamic consistency, which is the thing a copula cannot provide at any price. Spread moves and defaults are the same mechanism seen at different scales, which matches what a credit book actually experiences: correlated spread widening first, defaults later.

It costs calibration difficulty. Getting enough tail dependence out of (18.7) requires Y to be able to jump — a diffusive common intensity produces correlated spreads and puts far less loss in the far tail than a jump factor of the same mean and variance, which is the Gaussian problem in a new costume. So Y is usually taken with jumps, and the model is harder to fit than a copula. That is the trade, and for a book that has to be hedged rather than merely priced, it is usually worth making. The calculation below shows what fitting it involves.

Calculation 18.5 (Fitting the common-factor model).

Take d⁢Y=−κ⁢Y⁢d⁢t+d⁢J with J a compound Poisson process of rate ℓ and exponential jumps of mean μ, and write B⁢(t)=(1−e−κ⁢t)/κ.

Only one integral matters. Take the idiosyncratic intensities Zi deterministic. Given Y, name i survives to T with probability exp⁡(−ai⁢ΛT−∫0TZi), where ΛT=∫0TY⁢𝑑t, and the names are independent given ΛT. So every joint quantity is an integral over the law of the single variable ΛT. Integrating Yt=y0⁢e−κ⁢t+∫0te−κ⁢(t−s)⁢𝑑Js and exchanging the order of integration,

ΛT=y0⁢B⁢(T)+∫0TB⁢(T−s)⁢𝑑Js,

a sum over the jumps of B⁢(T−s) times the jump size.

Its transform is in closed form. For a Poisson process the exponential formula gives 𝔼⁢[e−a⁢ΛT]=exp⁡(−a⁢y0⁢B⁢(T)+ℓ⁢∫0T(𝔼⁢[e−a⁢B⁢(s)⁢X]−1)⁢𝑑s), and for an exponential jump X of mean μ, 𝔼⁢[e−u⁢X]=1/(1+μ⁢u). With c=μ⁢a/κ the integrand is 1/((1+c)−c⁢e−κ⁢s)−1, so

𝔼⁢[e−a⁢ΛT]=exp⁡(−a⁢y0⁢B⁢(T)+ℓ⁢[ln⁡((1+c)⁢eκ⁢T−c)κ⁢(1+c)−T]). (18.8)

This is the affine structure of chapter 14 in the simplest case, and it is checked against an exact simulation of ΛT, which is a Poisson number of jumps at uniform times and needs no time discretisation.888the transform at three loadings, and the mean and variance.

What the jump is. It is in the common factor, so every intensity rises at once: a system-wide shock and not a default. Given the path of Y the names default independently, so a default in one name does not by itself move the others, and there is no direct contagion; all dependence is through common shocks. That is what reduces the model to the single variable ΛT. Contagion in the direct sense, where a default makes the other names’ intensities jump, needs a self-exciting term and breaks that reduction, and a default can also move the other spreads without any intensity jumping if the factor is hidden and the market learns about it from the default. This model has neither.

The first layer of the fit is the single-name curves. Because Y and Zi are independent, Qi⁢(T)=𝔼⁢[e−ai⁢ΛT]⋅e−∫0TZi. The first factor is (18.8) at ai, so a deterministic function Zi⁢(t), bootstrapped maturity by maturity exactly as the hazard rate is bootstrapped in chapter 16, reproduces each name’s CDS curve for any values of κ, ℓ, μ and ai. The single-name market is fitted exactly, as a Hull-White shift fits the rates curve, and it constrains nothing about the dependence.

The second layer is the tranches. Given ΛT, a large homogeneous portfolio loses the deterministic fraction (1−R)⁢(1−e−a⁢ΛT−z⁢T), so the expected loss of a tranche is 𝔼⁢[f⁢(ΛT)] for a call spread f, an integral over the law of ΛT. There is no closed form for it, which is the first way the model is harder than a copula: the law comes from inverting (18.8) or from sampling ΛT as above, and the calibration is a least squares over that inner computation, with common random numbers so that the objective is smooth in the parameters. The parameters κ, ℓ, μ and a are chosen to minimise the squared error against the quoted tranches.

What identifies the parameters. The n-th cumulant of ΛT is ℓ⁢n!⁢μn⁢∫0TB⁢(s)n⁢𝑑s. The mean and variance fix ℓ⁢μ and ℓ⁢μ2 and so, given κ, both ℓ and μ, and every higher cumulant is then predicted. The equity and mezzanine tranches, which price the middle of the law, therefore pin the parameters, and the senior tranche, which prices the far tail, is a prediction of the model and not a free fit. A miss there is information about the model. A copula reaches each tranche with its own correlation and so fits every quote exactly, which is why it is easier to fit and why the fit says less.

Take κ=0.5, ℓ=0.4 per year, μ=0.05, y0=0.01 and T=5, and compare with a diffusive factor, whose integral is Gaussian, of the same mean and variance of ΛT. The mezzanine tranches, 3–10% and 10–20%, lose less under the jump factor, and the 20–60% tranche loses more than twice as much: jumps take loss out of the middle of the law and put it in the tail, with the first two moments unchanged.999the tranche loss ratios.

Remark (Contagion, modelled directly).

The common-factor model has no contagion: given the path of Y the names default independently. To put it in, add to each intensity a term that jumps when another name defaults and then fades, λti=ai⁢Yt+Zti+∑j≠ibi⁢j⁢e−δ⁢(t−τj)⁢𝟏{τj≤t}. This is the infectious-defaults model of Davis and Lo, and the primary-secondary counterparty structure of Jarrow and Yu. A default now raises the other intensities at once, so defaults cluster and spreads jump when a default is observed, which the common factor can only imitate by jumping on its own.

The price is conditional independence: the intensities depend on each other’s default histories, and the reduction to ΛT is lost. What survives is the affine structure. For an exchangeable portfolio with an exponential kernel, the pair of the default count and the intensity is Markov, and the generator applied to eu⁢N+v⁢λ has a multiplier affine in λ, so the invariance test of chapter 14 passes and the transform of the loss solves a Riccati pair. Errais, Giesecke and Goldberg develop this for affine point processes. The two mechanisms are complements, shared shocks and feedback between defaults, and a book that wants both has to fit both from the same few tranche quotes.

18.4.4 Do not model dependence at all

The fourth option is the most interesting, and it is available whenever a liquid tranche market exists.

Notice what a tranche actually is. It absorbs portfolio loss L between a and d, so its payoff is

(L−a)+−(L−d)+d−a, (18.9)

which is a call spread on a single scalar variable. The whole capital structure is a strip of call spreads on L — and by chapter 9’s theorem Breeden-Litzenberger, a strip of call spreads on a variable determines that variable’s distribution.

Remark (The same move as chapter 9).

This is the local volatility argument again. There we stopped trying to guess the dynamics of the underlying and read the risk-neutral density straight out of the option prices. Here we stop trying to guess the copula and read the loss distribution straight out of the tranche prices. In both cases the object that was being modelled turns out to be observable, and modelling it was never necessary.

The advantages transfer too, and so do the limitations. What is recovered is the loss distribution at each horizon and nothing more — the same partial information chapter 11 had to work around with local-stochastic volatility, and for the same reason. A bespoke tranche with different attachment points can be priced by interpolating the recovered distribution. A product depending on which names defaulted, or on the order they defaulted in, cannot: that is information about the copula and not about L, and it is not in the tranche prices.

Exercise.

Show from (18.9) that the notional-weighted expected losses of any partition of [0,1−R] into tranches sum to the expected loss of the portfolio itself, whatever the dependence. (Hint: the call spreads telescope.) This is the credit analogue of put-call parity: a model-free constraint, which makes it the first thing to check in any implementation. quant/src/dependence.rs tests exactly this, for both copulas and three correlations.

18.5 The Gaussian Copula, and What It Cost

We can now say what went wrong precisely.

The Gaussian copula is not a broken model. It is a coherent, tractable, easily simulated dependence structure that reproduces any correlation matrix exactly, and for a great many purposes it is perfectly adequate. Its one structural property is Theorem 18.4: it sets tail dependence to zero.

It was then adopted as the market standard for pricing tranches, whose senior end is a pure bet on tail dependence. The single quantity those instruments were sensitive to was the single quantity the model asserted to be zero.

Example 18.3 (What the assumption is worth).

A large homogeneous portfolio, five percent of names defaulting by the horizon, forty percent recovery, thirty percent correlation. The same marginals, the same correlation, under a Gaussian copula and a Student-t with four degrees of freedom:101010every tranche in the table.

Tranche Gaussian Student-t Ratio
0−3% (equity) 54.11% 39.40% 0.73
3−7% 19.58% 18.25% 0.93
7−15% 5.83% 8.35% 1.43
15−30% 0.81% 2.41% 2.98
30−60% (super senior) 0.019% 0.195% 10.16

Every number a risk report records is identical between the two columns: the same default probabilities, the same recovery, the same correlation. The super senior tranche is worth ten times more under one than the other.

And note the direction. The fat-tailed copula makes the equity tranche safer — it puts more probability at low losses as well as at high ones. So the effect is not a uniform increase in risk that a conservative haircut would have caught. It is a transfer of risk up the capital structure, and it is invisible at the bottom, where the market’s attention was, and largest at the top, where the notional was.

Figure 18.2: The same portfolio sliced into three point wide tranches at every attachment point. The upper panel shows the levels, which are dominated by the equity end and reveal nothing about the top. The lower panel shows the ratio of the two models, which reveals everything: below one at the bottom, crossing at around six percent attachment, and running away above it. A model chosen for its behaviour on the tranches that trade most was applied to the tranches that carried the most notional.
Show the model behind this figure (2 functions)
Portfolio::tranche_expected_lossquant/src/dependence.rs
/// The expected loss on the tranche covering `[attach, detach]`, as a
/// fraction of the tranche's own notional.
///
/// A tranche is a call spread on the portfolio loss, so its expected loss is
///
/// ```text
///     ( E[(L - a)+] - E[(L - d)+] ) / (d - a)
/// ```
///
/// and each call is the integral of the exceedance probability above its
/// strike. This is Breeden-Litzenberger from the local volatility chapter
/// read backwards, and the dependence chapter makes something of that: the
/// capital structure of a portfolio is a strip of call spreads on one
/// variable, so a complete set of tranche quotes implies a loss
/// distribution the same way a complete set of option quotes implies a
/// density.
pub fn tranche_expected_loss(&self, attach: f64, detach: f64) -> f64 {
    if detach <= attach {
        return 0.0;
    }
    (self.call_on_loss(attach) - self.call_on_loss(detach)) / (detach - attach)
}
Portfolio::exceedancequant/src/dependence.rs
/// The probability that the loss fraction exceeds `level`.
///
/// For the Gaussian copula this is Vasicek's closed form, inverted. For the
/// Student-t the mixing variable has to be integrated out, which is done on a
/// grid over the chi-square density; the integrand is smooth and one
/// dimensional, so a few hundred points is far more accuracy than the model
/// deserves.
pub fn exceedance(&self, level: f64) -> f64 {
    let loss_given_default = 1.0 - self.recovery;
    if loss_given_default <= 0.0 {
        return 0.0;
    }
    // Convert a loss level into the default rate that produces it.
    let rate = level / loss_given_default;
    if rate <= 0.0 {
        return 1.0;
    }
    if rate >= 1.0 {
        return 0.0;
    }

    let rho = self.correlation.clamp(1e-9, 0.999_999);
    let sqrt_rho = rho.sqrt();
    let sqrt_one_minus = (1.0 - rho).sqrt();

    match self.copula {
        Copula::Gaussian => {
            // Loss exceeds the level exactly when the factor is low enough.
            let threshold = norm_inv(self.default_probability);
            let critical = (threshold - sqrt_one_minus * norm_inv(rate)) / sqrt_rho;
            norm_cdf(critical)
        }
        Copula::StudentT { nu } => {
            let threshold = t_inv(self.default_probability, nu);
            chi_square_average(nu, |w| {
                // Given w, the same inversion as the Gaussian case.
                let critical = (threshold * (w / nu).sqrt()
                    - sqrt_one_minus * norm_inv(rate))
                    / sqrt_rho;
                norm_cdf(critical)
            })
        }
    }
}
Remark (Base correlation was the tell).

The market did notice, in the way markets do. Fitting one Gaussian correlation to all tranches simultaneously turned out to be impossible: each tranche implied a different one, and the market began quoting a “base correlation” per attachment point, rising steeply with seniority.

This is precisely chapter 9’s volatility smile, and it means precisely the same thing. A parameter that the model says is a single number, quoted instead as a curve, is the market saying the model is wrong in a structured way and telling you the shape of the error. In the volatility case, the profession read the message and built chapters 10 to 12 in response. In the correlation case, the curve was largely treated as a quoting convention.

The lesson generalises past credit. When a model requires a different parameter value for each instrument, the parameter is not a parameter. It is a record of the model’s error.

Remark (The fair verdict).

It is fashionable to blame the formula, and that is too easy — and it lets the actual mistake escape unexamined. The Gaussian copula did what it says. The failure was in the layer above it: an instrument was priced with a model that had been chosen for tractability and for its fit to the part of the capital structure that traded, and nobody asked what the senior end was sensitive to and whether the model had anything to say about it.

That question — what is this instrument’s payoff a bet on, and does my model have a view on that quantity or has it assumed one — is not specific to credit, and it is the one thing to take away from this chapter.

18.6 The Same Gap in Rates

A CMS spread option pays on the difference between two swap rates of different tenors,

(S1⁢(T)−S2⁢(T)−K)+, (18.10)

and the callable range accruals and steepener notes that make up much of the structured rates market are strips of them. Chapter 12 met this payoff already, as the instrument that exposes what a one-factor model cannot do.

Now notice the shape of the pricing problem. Each rate’s marginal distribution is known: its swaption smile gives the law of the rate under its own annuity measure by Breeden-Litzenberger, without any model of the curve, and chapter 15 shows how to carry that law to the measure of a payment date at the price of one assumption about the annuity mapping, not a model of the curve. What is not known is the joint distribution, and (18.10) depends on nothing else. That is exactly the position §18.1 opened this chapter with, in a different market: marginals supplied by the market, dependence supplied by the modeller.

So the market does exactly what credit did. It takes the two marginals from the smiles, joins them with a Gaussian copula carrying a single correlation, and integrates (18.10) against the result. And it has arrived at the same symptom: no single correlation fits the quoted spread options across strikes, so a correlation is quoted per strike — the base correlation of the previous section, in rates, for the same reason.

Calculation 18.6 (What the copula is worth on a spread option).

Set up as the tranche comparison was, so that nothing but the dependence differs. Two rates at 4% and 2.5%, both with an absolute volatility of 100 basis points, five years, correlation 0.8 — and both keeping exactly the same normal marginal under either copula.111111measured. The Gaussian column is checked against Bachelier, which it must reproduce, since normal marginals joined by a Gaussian copula give a normal spread.

Against a Student-t with four degrees of freedom, the option struck at a spread of 1.5% is worth 0.95 times its Gaussian value. Struck at 4.5% it is worth 2.6 times.

Tail dependence makes extremes arrive together, and one expects two rates that move together to produce a narrower spread. What actually happens is that a Student-t copula is a bivariate normal divided by a common random scale, and a small draw of that scale makes both rates extreme without making them equal. The spread inherits the heavy tail. So the Gaussian copula understates a far strike spread option for precisely the reason it understated a senior tranche, and the mass it is missing there is the mass it has put in the middle.

Remark (Two assumptions, not one).

Rates adds a layer that credit does not have.

The marginals in credit come from CDS quotes with no modelling in between. The marginals here come from chapter 15, which reaches them only through the linear annuity mapping — an assumption in its own right, and one calibrated to different instruments than the copula is. A CMS spread option therefore rests on two dependence-free approximations stacked on each other: a mapping to get each rate’s law under the payment measure, and a copula to join them. Neither is implied by the other and a reserve against one says nothing about the other.

A quanto CMS spread adds a third, since chapter 17’s adjustment needs each rate’s correlation with the exchange rate. The object required is then a three-dimensional joint law of which the market quotes only pairwise pieces, and the standard practice — assemble it from those pieces and hope the assembly is consistent — is not guaranteed to produce a valid joint distribution at all.

Calculation 18.7 (What “valid” means, and how it fails).

Standardise the three variables and read each as a unit vector in the L2⁢(Ω) of chapter 1. A correlation is the cosine of the angle between two of them, ρ⁢(X,Y)=⟨X,Y⟩/(‖X‖⁢‖Y‖), and any three vectors have a positive semi-definite Gram matrix of their pairwise inner products — v⊤⁢G⁢v=‖∑ivi⁢ui‖2≥0 for every v. So three numbers can be the pairwise correlations of three real random variables only when

(1ρ12ρ13ρ121ρ23ρ13ρ231)

is positive semi-definite, which for three variables is one inequality beyond |ρi⁢j|≤1,

1−ρ122−ρ132−ρ232+2⁢ρ12⁢ρ13⁢ρ23≥ 0. (18.11)

Nothing in three separate calibrations enforces it. Take ρ12=0.8 for the two rates, as calibrated to the spread option, and ρ13=0.7, ρ23=−0.1 for each rate’s correlation with the exchange rate, as calibrated to a quanto adjustment on that rate alone — both plausible on their own, since the two rates are not the same variable and can disagree about their exchange rate exposure. The left side of (18.11) is −0.252. No real (rate1,rate2,FX) has this triple of correlations: the matrix has no square root, so a Gaussian copula built from it cannot be sampled, and nothing downstream of “assemble the pieces” was ever a joint distribution.121212this triple fails a Cholesky factorisation and a consistent one succeeds.

Which leaves the choice this book keeps returning to. A term structure model of chapter 12 or chapter 13 supplies the joint law by construction and gives up the exact fit to the smiles; a copula keeps the exact fit to each marginal and supplies the joint law by assumption. Chapter 12’s calculation How much a factor costs measured what the first costs on this very product. Calculation 18.6 measures what the second costs.

References

  • -

    Sklar, A. (1959). Fonctions de repartition a n dimensions et leurs marges. Publications de l’Institut de Statistique de l’Universite de Paris, 8, 229–231.

  • -

    Embrechts, P., McNeil, A., & Straumann, D. (2002). Correlation and dependence in risk management: properties and pitfalls. In Risk Management: Value at Risk and Beyond, Cambridge University Press, 176–223.

  • -

    Li, D. X. (2000). On default correlation: a copula function approach. Journal of Fixed Income, 9(4), 43–54.

  • -

    Duffie, D., & Garleanu, N. (2001). Risk and valuation of collateralized debt obligations. Financial Analysts Journal, 57(1), 41–59.

  • -

    Davis, M., & Lo, V. (2001). Infectious defaults. Quantitative Finance, 1(4), 382–387.

  • -

    Jarrow, R. A., & Yu, F. (2001). Counterparty risk and the pricing of defaultable securities. Journal of Finance, 56(5), 1765–1799.

  • -

    Errais, E., Giesecke, K., & Goldberg, L. R. (2010). Affine point processes and portfolio credit risk. SIAM Journal on Financial Mathematics, 1, 642–665.

  • -

    Frechet, M. (1951). Sur les tableaux de correlation dont les marges sont donnees. Annales de l’Universite de Lyon, Section A, 14, 53–77.

  • -

    Hoeffding, W. (1940). Massstabinvariante Korrelationstheorie. Schriften des Mathematischen Instituts der Universitat Berlin, 5, 179–233.

  • -

    Nelsen, R. B. (2006). An Introduction to Copulas, 2nd ed. Springer.

  • -

    McNeil, A. J., Frey, R., & Embrechts, P. (2015). Quantitative Risk Management, revised ed. Princeton University Press.