Chapter 18 Dependence
Chapter 16 left a gap: single-name curves pin every marginal default distribution and say nothing at all about the joint one. This chapter is about what to put in that gap. It builds the tools that describe dependence — Sklar’s theorem, rank correlation, and the tail dependence coefficient that turns out to be the number most portfolio instruments are actually a bet on — and then gives four things to use, in increasing order of how much structure they demand. The Gaussian copula appears last, as a diagnosis rather than a target: it is a coherent model whose one defining property is that it sets to zero the exact quantity a senior tranche is priced off, and we finish by measuring what that costs.
18.1 The Gap, Stated Precisely
A basket of names, each with a default time whose distribution chapter 16 bootstrapped from that name’s own CDS quotes. Anything depending on more than one name — a first-to-default swap, a tranche, the counterparty exposure of chapter 24 — depends on the joint distribution , and the marginals do not determine it.
The following theorem says exactly how much freedom that leaves.
Theorem 18.1 (Sklar).
For any joint distribution with marginals there is a function with
| (18.1) |
and is itself a distribution function on the unit cube with uniform marginals. It is unique wherever the marginals are continuous. Conversely, any such combined with any marginals gives a valid joint distribution.
Proof.
With continuous marginals the copula is not merely shown to exist, it is written down.
The probability integral transform is the whole idea: if is continuous then is uniform on , since . So define
| (18.2) |
which is the joint distribution of . It is a distribution function because is; its marginals are uniform by the transform; and substituting into (18.2) recovers (18.1). Uniqueness is immediate from the same substitution: (18.1) determines at every point of the form , and continuity of the marginals makes those points the whole cube.
The converse needs no work at all. Given any copula and any marginals, the composition is a composition of a distribution function with non-decreasing right-continuous maps, so it is again a distribution function, and setting all but one argument to their upper limits returns . This half is what does the damage in this chapter: it says the construction never fails, so no choice of can ever be ruled out by the marginals.
When a marginal has an atom the argument breaks at exactly one point — is no longer invertible, and (18.1) constrains only on the range of , which now has gaps. Existence survives, by interpolating across the gaps, and uniqueness does not. That is the whole content of the continuity hypothesis, and it is why a default-time distribution with a lump of probability at a coupon date is a case where “the” copula is not well defined. ∎
Remark (What it means).
A joint distribution factors cleanly into two independent pieces: the marginals, which say how each name behaves on its own, and the copula , which says how they are coupled. The two can be chosen separately, and every choice is legitimate.
The intuition is a change of coordinates. is uniform on whatever was — this is the probability integral transform — so applying to each name strips out everything specific to that name and leaves only its rank. The copula is the joint distribution of the ranks. It is what remains of the dependence after every marginal has been standardised away.
Sklar’s theorem is usually presented as a technical result. For our purposes it is the precise statement of chapter 16’s complaint: the market gives us the and says nothing whatever about , and by the converse half of the theorem, every is consistent with the quotes. Choosing one is not calibration. It is an assumption, and it should be argued for rather than defaulted into.
Example 18.1 (The two extremes).
Two names, each defaulting within five years with probability . What is the probability both do?
If they are independent, and the answer is . If they are comonotone — one defaults exactly when the other does — then and the answer is . Both are consistent with identical CDS quotes on both names. A tenfold range in the joint probability, and nothing in the single-name market narrows it by a basis point.
18.2 Correlation Is the Wrong Word
The industry names this problem “correlation”. That is a poor choice.
Remark (Correlation is not invariant).
Linear correlation is a property of the variables, not of their copula. Apply a strictly increasing function to one of them — take a log, or a square — and the copula is unchanged by construction, since ranks are unchanged, while the correlation moves. So correlation mixes up the dependence with the marginals, which is exactly the separation Theorem 18.1 was useful for making.
Rank statistics do not have this problem. Kendall’s and Spearman’s depend only on the copula, and are the right things to quote if a single number must be quoted.
Theorem 18.2 (Attainable correlations).
Let and be lognormal. Then whatever the joint law of ,
| (18.3) |
Proof.
The Hoeffding-Frechet bounds say the extreme joint distributions with given marginals are the comonotone one, , and the countermonotone one, . For lognormals these are and , and substituting each into with the standard lognormal moments gives the two ends of (18.3). ∎
Example 18.2 (The bounds are not decorative).
Evaluating (18.3) numerically:111quant/src/dependence.rs.
| Volatilities | Lowest attainable | Highest attainable |
|---|---|---|
| and | ||
| and | ||
| and | ||
| and |
Read the last two rows. Two lognormals at volatility cannot be more than two percent negatively correlated no matter how they are coupled, because the closest they can come to opposed is against , and both of those are mostly near zero with an occasional enormous value — which makes them nearly independent, not opposed. And a name and a name cannot exceed however tightly they are bound, because the fatter-tailed one has a variance the thinner one cannot track.
A risk system that accepts a correlation of between two volatile positions has accepted a number that no joint distribution in the world can produce. It will not complain. It will return an answer.
Remark (And zero correlation is not independence).
The familiar one, included because it is the reason the next section exists. Correlation measures one particular linear summary of the joint law, and there are many ways to be strongly dependent with none of it. In particular, two variables can be uncorrelated and still almost always crash together, which is the case a portfolio of credits actually lives in.
18.3 The Number That Actually Matters
The failures above are diagnostic. This is the one that decides prices.
Definition 18.3 (Tail dependence).
The coefficient of upper tail dependence of a copula is
| (18.4) |
where are the two uniform coordinates. It is the probability that one variable is extreme, given that the other is, in the limit of extreme.
Remark (Why this and not correlation).
Consider what a senior tranche is. It absorbs nothing until the portfolio has already lost, say, fifteen percent, which requires a large fraction of the names to have failed together. Its entire value is the probability of a joint extreme. It is a bet on and on almost nothing else.
Correlation, by contrast, is an average over the whole distribution, dominated by the middle where most of the probability sits. Two copulas can agree on correlation to three decimal places, agree on every marginal exactly, and disagree about completely — and a senior tranche will price off the disagreement rather than the agreement.
Theorem 18.4 (Tail dependence of the two standard copulas).
The Gaussian copula with correlation has . The Student- copula with correlation and degrees of freedom has
| (18.5) |
for every and every finite .
Proof.
Both halves come from the same question — conditional on one variable being far out in the tail, how far out is the other — and the two answers differ for one structural reason, which the calculation makes visible.
The Gaussian case. Write as standard bivariate normal with correlation and condition on . Then
so is the limit of as , which is
for any , since the argument goes to . The conditioning moved ’s mean out by , but the residual has a fixed scale, and a fixed-scale Gaussian cannot keep up with a barrier receding at rate . So the second name is dragged part of the way into the tail and left there. Only at does the drag become complete.
The Student case. The copula is the Gaussian construction with one change: the pair is divided by an independent random scale,
with Gaussian of correlation . Now conditioning on being extreme is informative in a way it was not before: a large is evidence not only about but about being small, and a small inflates as well. The tail event is explained by the common factor rather than by the individual one, and the common factor moves both names together.
Making that quantitative is a computation with the conditional law of given , which for the bivariate is again a — with degrees of freedom, a mean , and a scale inflated by the factor that carries the information about . Then
and as the in the denominator grows like and cancels the in the numerator, leaving the finite limit
The factor of two in (18.5) is the standard convention of adding the two tails. Note where the cancellation came from: the scale of the conditional distribution grew with , and it grew because the mixing variable is shared. That is the entire difference between the two copulas. ∎
Structure (The cancellation is the point).
The two calculations are the same calculation with one term changed, and the term is the scale of the conditional law. In the Gaussian case it is constant, so a barrier receding at rate eventually outruns it and the limit is zero. In the Student case it grows like , the two rates match, and a finite limit survives.
The mechanism carries beyond copulas, and it is general: a fixed-scale residual cannot produce tail dependence, and a shared random scale always can. Any model in which the extreme behaviour of several quantities is driven by a common multiplicative factor — a stochastic volatility shared across names, a funding cost that hits every position, a liquidity parameter — has tail dependence for this reason, and any model whose only coupling is through the mean of a fixed-variance residual does not, whatever correlation is fed into it. Chapter 24’s stress work is the practical form of the same observation: the scenario that matters is the one that moves the shared scale.
Remark (Read that again).
The Gaussian statement is not that tail dependence is small, or that it has been approximated away. It is exactly zero, for every correlation short of perfect. Condition on one name being in a one-in-a-million event and the probability that the other is too converges to zero. Joint extremes, in a Gaussian copula, do not happen.
The reason is the same one that makes a bivariate normal’s contours ellipses: as you move out along the diagonal the density falls off faster than along either axis, so the far tail of a Gaussian is dominated by one variable being extreme, never both. The correlation controls how elliptical the contours are.
The Student- fixes it with one extra parameter and no extra conceptual machinery. A is a normal divided by an independent random scale, and it is that shared scale that produces joint extremes: occasionally the whole system is drawn from a wide distribution, and then everything moves at once. This is not a mathematical trick. It is a common volatility factor, and it is what the world does.