Chapter 22 Relative Value and the Basis Trade
We look at what the other side of the industry does with the same machinery. A relative value desk is not trying to price a derivative correctly; it is trying to find two things that should be worth the same and are not. We derive the decomposition of a fixed income return into carry, roll-down and yield change, construct the curve trades that isolate one factor from another, and work through the cash-futures basis in enough detail to see why the trade exists and why it periodically destroys the people doing it. Who is on the other side, and why they stay there, is the question the rest of the chapter returns to: the same ideas cover selling volatility and what that position is short, the retail versions of these trades, and how much a signal about a future release is worth, which is a mutual information measured in nats and sized by a log-optimal rule.
22.1 What This Chapter Is
Everything so far has been derivation. A model was posited, a measure found, a price computed, and the result was true given the assumptions. This chapter describes trades that people put on to make money.
The subject here is not whether these trades work. It is what they are: what a position consists of, how its profit and loss decomposes, and — the part that is genuinely mathematical and genuinely useful — where the risk actually sits, which is frequently not where it appears to sit. A trade whose P&L looks like a small steady income and whose risk is a rare enormous loss is a recognisable object, and recognising it is worth more than any view about whether it is currently cheap.
The reader who wants to know whether to put a trade on will not find it here. The reader who wants to know what they are holding will.
Structure (What makes a spread a trade).
A model tells you a spread is wide. That is not a signal, and treating it as one is the characteristic error of the subject. Two further things are required before a wide spread is a trade, and neither is a modelling output.
A structural reason. Somebody has to be on the other side for a reason that is not an opinion — a regulatory constraint that forces a pension fund to hold a particular instrument, an index rule that requires a bond to be sold on a date, a balance sheet cost that makes a dealer unwilling to warehouse a position, a settlement convention that nobody can arbitrage away. When such a reason exists, the spread is a price being paid for a service and can be collected. When it does not, the wide spread is more likely to be the model’s error than the market’s, and §22.9’s negative number is the case where telling the two apart is the whole problem.
A horizon. Convergence is not a date. A spread that is two standard deviations wide and genuinely mean reverting still takes longer to close than the intuition suggests, and goes further against the position first. That is a question about first passage times rather than about forecasting, it is answerable, and the next section answers it — because the usual way of getting it wrong is not misjudging direction but misjudging how long being right takes.
The model’s job in all of this is narrower than it looks. It supplies the spread and the decomposition of the P&L; it has nothing to say about either of the two conditions above. Which is why the discipline in relative value work lives in chapter 20’s estimation rather than in the pricing.
22.2 Who Is On the Other Side
The structural reason the structure above demands is not an abstraction. It belongs to a specific institution, with a specific mandate and a specific regulator, and knowing who that is — and what they are required, rather than persuaded, to do — is most of what separates a genuine relative value trade from a model’s own noise.
Repo, which is how balance sheet actually gets priced
Almost everything below is financed the same way, so the mechanism comes first, before the players.
Definition 22.1 (Repurchase agreement).
A repo is a sale of a security today, combined with an agreement to repurchase the same security at a fixed later date for a fixed, slightly higher price. Economically it is a collateralised loan — the seller borrows cash and posts the security as collateral — structured as two trades rather than one so that, in a bankruptcy, the lender already owns the collateral rather than standing in a queue for it. The difference between the two prices, annualised, is the repo rate: the interest paid on the cash. The lender advances less than the collateral’s market value, the haircut, to stay protected if the collateral’s price falls before it can be sold.
Most repo is general collateral: the cash lender does not care which bond is posted, any similar one will do, and the rate sits close to a short-term benchmark. A bond becomes special when enough people need that specific security — to deliver into a short, to meet a settlement obligation, because it is the benchmark issue hedges are executed in — that they will accept a lower return on their cash purely to secure it. The repo rate on a special bond falls below the general collateral rate, and in a genuine squeeze it can fall to zero or below: the cash lender is then paying to lend, because what they are actually buying is the bond, not the yield. That is the mechanism behind §22.11’s on-the-run richness and what happened to Salomon’s target note in 1991: a repo rate is not a property of the borrower’s credit, it is a price for a specific piece of collateral, and it moves with demand for that collateral exactly like any other price.
The players
Pension funds owe payments decades into the future, frequently linked to inflation or wages, and are marked against those liabilities’ own discounted value rather than against a market index. That creates a structural demand for very long duration and for inflation exposure, in quantities a fund’s own assets rarely supply directly — which is why liability-driven investing uses swaps and long gilts or Treasuries, often levered through repo, to close the gap cheaply rather than buying enough of the underlying bond outright. The lever is exactly what turns a slow, structural trade into a fast one under stress: in September 2022 a rise in UK gilt yields fell on levered liability-driven investing positions as a fall in collateral value, triggering margin calls that could only be met by selling gilts, which pushed yields higher and triggered the next round of calls — the same feedback loop the mortgage convexity hedging entry of §22.11 describes, with a margin call standing in for a rebalancing trigger and the Bank of England’s emergency purchases standing in for the intervention that finally broke it.
Insurers face a related but differently shaped constraint. Solvency II in Europe and risk-based capital rules in the US charge capital against a mismatch between the timing of assets and the timing of liabilities, so an insurer’s own capital cost falls when it buys assets that match liability cashflows closely — long-dated credit, structured or illiquid assets that qualify for favourable treatment under the rules, specifically because they match rather than because they are attractively priced. The constraint creates demand for an asset shape, and a security that fits that shape can trade rich to an otherwise identical one that does not, for the same reason the on-the-run bond does.
Banks and dealers are constrained by capital and leverage-ratio rules, most acutely at quarter ends when those ratios are reported, and by the balance sheet a repo book requires to warehouse anything. Chapter 23 works out the consequence for quoting; here the consequence is that a dealer is structurally reluctant to hold inventory, structurally short the optionality clients want to buy, and the natural supplier of exactly the balance sheet covered interest parity’s cross-currency basis prices.
Hedge funds and relative value desks are the balance sheet that steps into the gap dealers leave — unconstrained by the same regulatory capital rules, but constrained instead by financing that can be pulled and by investors who can redeem, both on a timescale shorter than the trade’s own horizon. That asymmetry is the horizon condition named above, in institutional form: the trade is right, and the capital available to hold it is not guaranteed to survive as long as being right takes.
Corporates issue the debt the rest of this list trades around, and hedge the interest rate and currency exposure that issuing it creates — paying floating and receiving fixed on newly issued debt, or the reverse, and buying the caps that bound a floating liability. That one-directional demand — borrowers wanting caps, issuers wanting the option to call — is exactly the flow “Selling Volatility, and What That Is Short” identifies as what leaves the dealer community structurally short optionality and paid to be.
Asset managers run mandates benchmarked to an index rather than to a liability, and are marked to market continuously rather than against a discounted cashflow, which pushes them towards the more capital-efficient instrument for a given exposure — futures over cash bonds, a swap over a bond where the mandate allows it — exactly the preference behind the cash-futures basis below.
None of this is a claim that any of these institutions is behaving irrationally. Each is optimising something — a funding ratio, a capital ratio, a tracking error — that is not the price of the instrument it trades, and §22.11 is a catalogue of what is left over when several of them optimise different things against the same market.
22.3 How Long Being Right Takes
Take the simplest possible model of a converging spread: an Ornstein-Uhlenbeck process, as in chapter 20,
| (22.1) |
entered when sits two stationary standard deviations from zero, held until it returns. The mean reversion is certain — this is not a case where the trade might be wrong about direction. Everything below is what happens when it is right.
Calculation 22.2 (First passage is not the half-life).
The half-life is the number everyone quotes and it answers a different question. It says how fast an expectation decays: halves in that time. It does not say how long a path takes to reach zero, and the two are not close.
Simulating (22.1) from two standard deviations, with a one-year half-life:111estimation::ConvergenceTrade.
| half-life | years |
|---|---|
| mean time to first reach the mean | years |
| mean worst level reached first, in deviations |
So the trade takes more than twice the half-life, and before converging it goes about half a deviation further against the position than the level it was entered at. Neither number is available from the half-life, and both are what size the trade.
Remark (Where a stop turns a winning trade into a loss).
The mean worst excursion is an average, so a stop placed at it is hit about half the time. Placing it closer is worse:
| Stop, in deviations | Fraction stopped out before converging |
|---|---|
| under |
Entered at two deviations, a stop half a deviation away loses a trade that was right about direction two times in five. At twice the entry width the stop stops binding altogether, and the trade’s risk becomes the holding period rather than the loss.
That is the regime to be in, and reaching it is a sizing decision rather than a stop-placement decision: the position has to be small enough that sitting through twice the entry width is tolerable. A desk that sizes to its stop rather than to its horizon has built a trade that loses when it is right.
Structure (The same first passage problem as the doubling strategy).
This is chapter 4’s doubling strategy in respectable clothing.
There, a strategy on a martingale was certain to reach its target and the drawdown before it had a tail so heavy that its mean was infinite — which is why admissibility has to bound the loss rather than its expectation. Here the process is mean reverting rather than a martingale, so the tail is far lighter and the mean excursion is finite. But the structure is the same: a bet that is certain to win, whose risk lives entirely in what happens before it does, and whose sizing is therefore governed by a first passage distribution rather than by an expected return.
The difference is what makes relative value a business and doubling a fallacy. Mean reversion makes the excursion’s distribution thin enough to survive with finite capital, and (22.1)’s is exactly what controls how thin. Which puts an uncomfortable amount of weight on knowing .
Remark (And is the parameter that is estimated badly).
Chapter 20 measures that mean reversion estimated from a finite sample comes out too fast, and chapter 14 explains why from the spectrum: a deviation decays more quickly than the slowest mode at first, so fitting one exponential to a whole sample returns a rate above the true gap. The bias is one-signed.
Follow that through to the trade. An overstated is an understated half-life, which is an understated horizon. With the true half-life at one year and the estimate thirty per cent fast — comfortably inside what a decade of data delivers — the planned holding period is years against a realised .222measured. The desk expects to be out in twenty months and is still in the position at twenty-five.
Every part of that error points the same way. The horizon is longer than planned, so the funding cost is larger than budgeted; the position is held through more of the excursion distribution, so the drawdown is deeper than modelled; and the capital is committed longer, so the return on it is lower than advertised. There is no compensating error in the other direction.
22.4 Whether the Spread Converges at All
The section above began by granting that the mean reversion is certain, and everything in it followed from that. What happens when it is not granted?
Chapter 20 has already ruled out the obvious approach. Fitting (22.1) to a history returns a positive whether or not there is any mean reversion to find — on a random walk the estimate is positive with probability approaching one — so “the regression found convergence” establishes nothing. What is needed is not a better estimate but a test, with the random walk as the null hypothesis rather than as an alternative nobody considered.
Chapter 20’s Dickey-Fuller test is the natural tool here rather than an imported one, and its critical value of about rather than is exactly the correction that a non-stationary null requires.333estimation::unit_root_test.
Calculation 22.3 (What the test can actually see).
The test is correctly sized — a genuine random walk passes for mean reverting about five per cent of the time, as asked. The question is the other error: how often does the test find mean reversion that is really there?
| Half-life | 1 year | 3 years | 5 years | 10 years |
|---|---|---|---|---|
| months | ||||
| year | ||||
| years |
Read the middle row. A spread that genuinely reverts with a one-year half-life, watched for five years, is identified as reverting eight times in a hundred.444measured. The test is not broken; it is being asked to separate two hypotheses that a sample of that length barely distinguishes.
And the pattern across the table is the same shape chapter 20 found for the estimation bias. Power depends on the number of half-lives the sample spans and not on its length or its resolution: a one-year half-life over five years and a two-year half-life over ten give the same answer, though one history is twice the other.555measured. Sampling the spread hourly would add observations and no power at all.
Structure (Which is why the structural condition does the work).
Put Calculation 22.3 beside the two conditions this chapter opened with and the two halves of its argument fit together.
A test that finds real convergence eight times in a hundred cannot be what a desk relies on. Worse, it is not merely uninformative but dangerous under search: given instruments there are many combinations to weight, and looking through enough of them will produce one that passes at five per cent whether or not anything converges. A screen that tests a thousand spreads finds fifty by construction. The statistic that was weak as evidence becomes actively misleading as a filter.
So the structural reason demanded earlier is not a piece of good practice sitting alongside the statistics. It is what the statistics cannot supply. A butterfly is a curvature by construction, a cash-futures basis is tied to financing by an arbitrage that must close at delivery, an on-the-run spread is a liquidity premium with a named mechanism — and each of those is a reason to expect convergence that does not come from the sample, so it is not consumed by having searched the sample. The test is then a check on a prior rather than a way of forming one, which is the only role its power supports.
That is also the honest difference between this and the equity pairs trading the same mathematics is usually taught with. There the combination is typically discovered by search over a universe, and the discovery is exactly the procedure Calculation 22.3 says will manufacture false positives. Here the combination is usually written down first, for a reason, and the data is asked only whether it disagrees.
22.5 Carry, Roll-Down, and the Rest
Start with the simplest possible question. I buy a bond, fund it, and hold it for three months. Where does the money come from?
Write for the yield of a -maturity bond, for its duration and for the horizon. Over the horizon two things happen: the bond gets older, and the curve moves. Expanding the price change,
| (22.2) |
The first two terms are the carry: what the bond pays less what the funding costs. The third is the roll-down: even if the curve does not move at all, a five year bond becomes a four-and-three-quarter year bond, and on an upward sloping curve that means its yield falls and its price rises. The fourth is the only term involving an actual change in the market, and the fifth is the convexity of chapter 7.
Calculation 22.4 (How the known part compares with the unknown).
Equation (22.2) separates what is known from what is not and says nothing about their sizes, which is the comparison that decides whether a carry trade is a harvest or a bet. Both are available: the known part from today’s curve, and the risk from the realised volatility of the same maturity’s yield over the same horizon.
Taking the Treasury curve and a quarter’s holding period, with the one-month bill as funding: the known part is exact discount-factor arithmetic against today’s curve, and the risk is realised yield volatility over four decades of history where a tenor has been quoted that long, and one year where it has not.666risk::what_the_curve_currently_pays_to_hold_by_maturity, against the panel in public/marketdata.
| Maturity | yield | carry roll | risk over the quarter | ratio |
|---|---|---|---|---|
| y | bp | bp | ||
| y | bp | bp | ||
| y | bp | bp | ||
| y | bp | bp | ||
| y | bp | bp |
Remark (One year is an exception, and the long end has inverted).
From two years out the pattern is the one the two conditions above lead you to expect: the known part is a small fraction of the risk, a directional bet with a small tilt in its favour rather than an income stream with noise attached.
One year does not fit it. A ratio of is not a small tilt, and the reason is visible in the curve itself: the one month bill sits close to half a point below the one year bill, a genuinely steep front end that even four decades of yield volatility does not dominate. Whether that persists is a question about tomorrow’s curve, not a property asserted of curves in general — the same calculation answers it again whenever it is asked.
And the long end has inverted rather than merely underperformed. The thirty year known return is comfortably negative, not the largest number on the curve: the twenty year point sits above both its neighbours, so a bond ageing from thirty towards twenty years ages into a higher yield and a lower price, roll-down running backwards. The worst risk-adjusted place to hold duration on this curve is not “the longest maturity” as a rule; it is wherever the curve’s own local slope happens to be working against you, and on this curve that is squarely the long end.
The important structural feature of (22.2) is that the first three terms are known today. They are properties of today’s curve and the passage of time, not forecasts. Only the fourth is uncertain.
That is what makes carry-and-roll trades attractive and what makes them dangerous. A position with positive carry and roll makes money on every day the market does not move. It loses money when the market moves against it, and the losses are proportional to duration, which is to say much larger than the daily income. A carry trade is therefore short a large, rare loss and long a small, steady gain — which is the payoff diagram of a sold option, assembled without buying or selling one.