Chapter 12 Markovian Term-Structure Models
In these notes we return to interest rates carrying what the volatility chapters taught us. The Heath-Jarrow-Morton framework of chapter 8 is completely general and, for that reason, computationally useless: its state is the whole curve. We derive the one restriction on the volatility that collapses it to a finite state, obtain the quasi-Gaussian models, and take a mathematical excursion into why models of this shape can be solved at all.
12.1 Why HJM Is Not Enough
Chapter 8 derived the Heath-Jarrow-Morton drift condition: under the risk neutral measure, an arbitrage-free evolution of the instantaneous forward curve must take the form
| (12.1) |
This is a complete answer to the question chapter 8 asked. Choose any volatility structure whatsoever and the drift is determined; there is no further freedom and no further constraint. As a piece of theory it could hardly be better.
As a thing to compute with it is close to unusable. The state of the model at time is the entire function . Nothing smaller will do: to take the next step, (12.1) needs integrated over the whole remaining curve, and in general may depend on the curve’s current shape. So the process is Markov in an infinite dimensional state and in nothing less.
For pricing a European payoff this is survivable, though not because simulation somehow escapes the dimension. It does not. Discretise the curve onto a few hundred maturities and every one of them is a state variable that has to be stored and stepped, because (12.1) needs integrated over the whole remaining curve before it can take a step, and if depends on the curve’s shape then the previous step’s answer is what supplies it. A simulated HJM path is a simulated surface.
What makes that bearable is the shape of the cost rather than its size. Carrying three hundred tenors costs three hundred times a one-factor model per step, which is linear in the dimension; and chapter 19’s measurement is that a Monte Carlo error is set by the payoff’s variance and the path count and does not grow with the number of state variables at all. So a large state multiplies the cost of each path and does not multiply the number of paths needed. A European payoff also never asks what anything is worth at an intermediate date: one sweep from today to expiry, average at the end, and the state’s dimension is bookkeeping.
Remark (When even the stepping is unnecessary).
If is deterministic, (12.1) can be integrated in closed form:
in which the drift term is a number and the stochastic integral is Gaussian with a deterministic variance. The whole curve at is then a jointly Gaussian vector that can be drawn in one shot, with no time stepping at all. It is the dependence of on the curve that forces the step-by-step simulation, and with it the need to carry the surface.
Backward induction is where the dimension stops being bookkeeping. A Bermudan swaption must be valued by comparing, at each exercise date, the value of exercising against the value of continuing, and that comparison has to be made state by state — so it needs the value at a state rather than an average over states. And regression needs a small set of variables to regress against, which is exactly what an infinite dimensional state does not offer, and nothing in the model tells us which handful of curve summaries would do.
So the question of this chapter is: what must we give up for the curve to be a function of a handful of numbers?
12.2 What a Finite State Requires
Two separate things have to be true before the curve collapses onto a handful of numbers. The first is that the model be Markov in finitely many variables. The second is that its volatility depend on how long a rate has left to run rather than on the date it matures. The exponential comes from the two together.
Take deterministic for this section. The state-dependent case is what the model is eventually for, and it is returned to at the end.
12.2.1 Markov Is Exactly Finite Rank
Integrating (12.1), the curve at time is
| (12.2) |
for every maturity . With deterministic the drift is deterministic too, so the whole question is the second term: a continuum of random variables, one per maturity, each an integral over the whole past.
Lemma 12.1 (Finite state, finite rank, and a sum of products are the same thing).
Let be deterministic and square integrable. The following are equivalent.
-
(i)
For every , the curve is recoverable from and random variables.
-
(ii)
The functions , as varies, span a space of dimension .
-
(iii)
is a sum of products and no fewer,
(12.3)
In that case the state variables are , and
| (12.4) |
Proof.
(i) (ii) is the Itô isometry. The map is linear, and
so it is an isometry and in particular injective. A family of such integrals therefore spans a space of exactly the dimension spanned by the family of integrands. Recovering (12.2) for every from numbers is thus precisely the statement that the integrands span dimensions.
(ii) (iii) is linear algebra with no probability in it. If the sections span dimensions, choose a basis of that span. Each section is in the span, so
where the coefficients depend on which section we took, that is on alone; and that is (12.3). Conversely (12.3) exhibits every section as a combination of the same functions. Substituting into (12.2) and pulling out of the integral, since it does not involve , gives (12.4). ∎
12.2.2 Finite Rank Alone Forces Nothing
It is tempting to stop here and conclude that a Markov model needs an exponential decay. It does not.
Example 12.1 (A one-state model with an arbitrary maturity shape).
Let for any square-integrable whatsoever — the volatility of the forward rate maturing at is , and does not change as advances. This is rank one, with , so lemma 12.1 gives a single state,
The drift condition determines from as always, so the model is arbitrage free; it is Markov in one variable; and is unrestricted. Take if a concrete one is wanted. Measuring the rank of the kernel confirms it is one, for that and for anything else put in its place.111the test that pins this.
So a Markov model can carry any maturity shape at all. What is wrong with example 12.1 is not its mathematics but its economics: the volatility attached to the rate maturing in is the same number in as it is today. A thirty year forward rate and a rate one day from fixing are given whatever volatility their maturity dates happen to have been assigned, with no relation between them. Real forward rates do not behave that way — what governs a rate’s volatility is how long it has left to run.
Structure.
This is chapter 10’s complaint arriving in the rates market. There the trouble was a local volatility surface whose skew decayed along the calendar, so a forward starting option saw a flatter smile than a spot one; here it is a volatility whose shape is pinned to dates rather than to tenors, so a forward curve does not look like a spot curve. Both are the same failure: a parameter carrying calendar time where the market carries only elapsed time.
12.2.3 The Second Requirement
So we impose it, as a modelling demand and not as a theorem.
Definition 12.2 (Time-homogeneous volatility).
The forward rate volatility is time homogeneous if it depends on the maturity only through the time left to it,
| (12.5) |
for a maturity profile and a level carrying whatever randomness there is.
Now put the two requirements together. Taking deterministic and non-vanishing, the sections of (12.5) are on ; the factor is common to all of them and so cannot change the dimension of their span, which leaves the sections of . Writing
| (12.6) |
turns a section into , indexed by . This is not a trick to make the algebra work. It is what time homogeneity does to the kernel: the sections of a kernel that depends only on the difference of its arguments are the shifts of one function, and the substitution merely names them.
So lemma 12.1 becomes, under definition 12.2: the model is Markov in states if and only if the shifts of span a space of dimension . Write for the shifted profile,
and for the linear span of the over . Expanding in a basis of gives the form (12.3) again, now with the shift structure visible:
| (12.7) |
Taking shows : the profile is one of the shapes its own shifts span.
12.2.4 Rank One Is the Exponential
Proposition 12.3.
Under definition 12.2, a model Markov in one state per factor has , and no other profile will do.
Proof.
means is spanned by itself, so (12.7) reads . Setting gives , and normalising — the constant can be absorbed into — leaves
| (12.8) |
This is Cauchy’s exponential equation, whose measurable solutions are and nothing else.222Measurability is not a formality one can drop: without it the equation has pathological solutions built from a Hamel basis. It is not a real restriction on a volatility. ∎
That is the model this chapter uses.
Definition 12.4 (Separable volatility).
The forward rate volatility is separable if it factorises as
| (12.9) |
with the first factor allowed to depend on time and on the state of the world, and the second a purely deterministic decay in the maturity.
Take constant for the rest of this chapter, so the decay is . Proposition 12.3 is what earns it: given time homogeneity, the exponential is not one convenient choice among many but the only one a single state permits.
Remark (Whether a varying mean reversion is allowed).
(12.9) writes and the proposition below concludes . Let’s see how these two sit together.
Write , so that the decay in (12.9) is and
That is a function of times a function of , whatever does — an outer product, of rank one as a kernel in .333computed. So a time-dependent mean reversion is permitted and costs nothing: the state is still and everything derived below goes through with read as .
What it costs is time homogeneity, which is the proposition’s hypothesis rather than its conclusion. The maturity profile at date is , and that depends on as well as on : the shape a rate’s volatility follows as it ages is a different shape depending on when one starts watching. Read as a fixed profile in time to maturity it is not an exponential at all, and its shifts span a space of dimension well above one.444computed.
Which makes this example 12.1 again, in the notation the chapter actually uses. One state, because the volatility is indexed by the maturity date; no fixed profile, because it is not indexed by time remaining. The economic objection carries across unchanged, and so does the practical reason desks do it anyway: a varying is one of the cheapest ways to fit a term structure of volatility exactly, and what is bought with the fit is paid for in forward volatility, which is then whatever the calibration happened to leave behind.
12.2.5 Rank Is a Differential Equation
Theorem 12.5 (Quasi-exponentials, and nothing else).
Suppose is smooth and is finite. Then satisfies a linear differential equation of order with constant coefficients,
| (12.10) |
and consequently
| (12.11) |
a finite sum of polynomials times exponentials, with the the roots of the characteristic polynomial of (12.10) and one less than the multiplicity of . Complex roots appear in conjugate pairs and contribute and . Conversely every such has equal to the order of the equation it satisfies.
Proof.
is closed under shifting: a typical element is , and is again a combination of shifts of , so again in .
That makes closed under differentiation. For ,
and every difference quotient on the right lies in , because both and do and is a linear space. A finite dimensional subspace is closed, so the limit lies in too.
So is a linear operator on an dimensional space. By Cayley-Hamilton it satisfies its own characteristic polynomial , of degree , on all of ; and . That is (12.10), with , and (12.11) is its classical solution. For the converse, the shifts of a solution of (12.10) are again solutions, so sits inside a solution space of dimension .555Smoothness is assumed rather than derived. It can be got from being finite with more work, by solving (12.7) for the at well-chosen values of and bootstrapping; but a volatility profile that is not smooth is not a modelling choice anyone makes. ∎
12.2.6 What Rank Two and Three Look Like
The cheapest way to see the general case is to expand by hand. For the binomial theorem gives
| (12.12) |
which is (12.7) exactly, with basis for . So needs states, and (12.10) agrees: , a root of multiplicity .
The case is the one that matters in practice, since a volatility peaking at some tenor rather than decaying from the front is closer to what is quoted. Written out, (12.12) is
so the hump is Markov in the two states of (12.7) built from and . The first is Hull-White’s own state; the hump keeps it and adds one. This is an identity per increment rather than in a limit, so it holds pathwise and exactly.666the pathwise test.
Everything else follows straight from the characteristic polynomial:
| maturity profile | its equation | states |
| none | no finite number | |
| none | no finite number |
Computing the rank of numerically returns exactly this column,777quasigaussian::realisation_dimension..
The last two rows deserve a word, since a theorem excluding nothing would hardly need proving. Neither function is pathological: both are smooth, positive and decaying, and either would be a reasonable thing to fit to a volatility term structure.
Example 12.2 (A profile with no finite realisation).
Take and suppose some finite combination of its shifts vanishes,
with the distinct. The left side is a rational function of with simple poles at , all distinct. Multiplying by and letting kills every term but the first and leaves ; repeating kills them all. So no finite set of shifts is dependent, is infinite, and there is no Markov realisation in any number of states.
Set that beside example 12.1, which used the same and got a model with one state. Nothing about the function changed. What changed is whether it was read as a shape in calendar time or a shape in time to maturity, and that alone moves the model from one state to infinitely many. It is as sharp a statement as this section has of where the exponential actually comes from: not from Markovianity, which tolerates any shape, but from insisting that the shape ride along with the maturity.
So being Markov in finitely many variables does not force separability. It forces finite rank, by lemma 12.1; time homogeneity then turns finite rank into a differential equation, by theorem 12.5; and separability is the case of the two together. Multi-factor Cheyette models are the case , at times the state and with nothing else changed.
Structure.
This is chapter 14’s subject arriving early. There the question is which models are solvable, and the answer is that a family of functions preserved by the dynamics turns a partial differential equation into a few ordinary ones. Here the dynamics is an infinite dimensional one on the curve, the family is , and preservation is closure under shifting; the payoff is the same, an infinite dimensional object collapsing onto finitely many coordinates. Tractability keeps turning out to be an invariance, and the invariant object keeps turning out to be a small linear space — here small enough that its dimension is literally the number of state variables.
Remark (What survives when the level is stochastic).
Everything above assumed deterministic, and (12.9) explicitly permits a depending on the state — which is what makes the model useful, since it is where the smile comes from.
Sufficiency is untouched. Nothing in theorem 12.6 uses determinism: the integrals defining , and are pathwise, and hold for any adapted . So a rank- profile with an exponential shape and an arbitrary state-dependent level still gives a finite Markov state, which is the direction the chapter actually needs. The construction is safe.
Necessity is what the Itô isometry was buying, and that is where the argument breaks — though not because the isometry fails. It does not: holds for every adapted square integrable integrand, random ones included. What stops working is the use the proof makes of it.
For deterministic the family is a family of fixed vectors in , and the isometry embeds that space linearly and injectively into . So the span of the random variables has exactly the dimension of the span of the functions, and both are spans over — constant coefficients. That is the notion (12.4) needs, because there the curve is rebuilt by deterministic loadings on random variables. Linear algebra is the right tool because, with deterministic, the curve is Gaussian and everything in sight is linear.
Let depend on the state and neither remains true. The integrands are themselves random, so is no longer a deterministic family in a fixed space; and a curve recoverable from state variables is recoverable by a measurable function of them, which for a non-Gaussian curve has no reason to be linear. The span that matters is then over the state-measurable coefficients rather than over — a module rather than a finite dimensional vector space — and a rank computed over is not counting it. The isometry survives intact; what it no longer does is convert “ state variables” into “ linearly independent integrands”. Lemma 12.1 loses its proof there.
The replacement is Lie-algebraic. Reparametrising the curve by time to maturity, — the Musiela form — turns (12.1) into a stochastic partial differential equation whose drift contains the transport term , because holding a tenor fixed means sliding along the curve as time passes. The model has a finite dimensional realisation exactly when the Lie algebra generated by its drift and volatility vector fields is finite dimensional.
Applied to a deterministic volatility, that criterion reproduces theorem 12.5. Bracketing the volatility field against the transport term differentiates the maturity profile; bracketing again differentiates it again. Finite dimensionality of the algebra therefore says that the derivatives of span a finite dimensional space — which is (12.10), and so still the quasi-exponentials. The elementary proof and the geometric one arrive at the same object, closed under , by different roads.
Once depends on the curve, those brackets acquire further terms from the volatility field’s own dependence on the state, and the criterion stops reducing to a condition on alone.
-
-
What the chapter uses is the if direction, and it holds regardless. An exponential profile with any adapted level whatever gives the two-state model of theorem 12.6.
-
-
What is given up is the only if. For deterministic , lemma 12.1 is an equivalence, so we know that nothing outside the quasi-exponentials could have worked. With a stochastic level we no longer have that proof, and whether some cleverly state-dependent level could rescue a profile that is not quasi-exponential is a question nothing here settles. It has to be settled model by model, from the Lie algebra.
The deterministic characterisation is due to Bjork and Christensen, the state-dependent analysis to Bjork and Svensson, and the general geometry of invariant manifolds to Filipovic and Teichmann.
With that, we can find the state.
Theorem 12.6 (Cheyette, quasi-Gaussian state).
Under (12.9), define
Then is Markov, with
| (12.13) |
both starting from zero, and the entire curve is recovered from them by
| (12.14) |
Proof.
Integrate (12.1) from to with the separable volatility. The stochastic part is
which is the factorisation the previous remark promised: the maturity has come outside, leaving a single history-dependent quantity. Call that quantity .
For the drift, the inner integral in (12.1) is
so the accumulated drift is . Since , the integrand is
and splitting both exponentials at as before gives
| (12.15) |
Note that two accumulators have appeared, at rate and at rate , so the state looks three dimensional at this point.
It is not. Setting in (12.15) and adding the stochastic part,
That is a relation among the three, and using it to eliminate from (12.15),
in which has cancelled. The bracket on the right is , so
| (12.16) |
Two states, not three, because the short rate itself supplies the third relation.
Integrating (12.16) in and exponentiating, using from chapter 7, gives (12.14): the produces , and
produces the quadratic term, after cancelling against the initial curve.
Finally the dynamics. Differentiating the three pieces separately,
the last of which is already the second equation of (12.13). For the first, , so
and the bracket is . Note that has cancelled again: it is needed to write the curve down and not to step it forward. ∎
Three things about this deserve emphasis, because they are what make the model usable rather than merely finite.
First, has no in it. It is not a risk factor and cannot be shocked; it is a running accumulator of variance, and its only job is to make the pair Markov. It is exactly the memory that (12.1)’s drift needs and that alone cannot supply.
Second, (12.14) reconstructs every discount factor as an explicit function of two numbers. So every forward rate, every swap rate and every annuity is a closed-form function of . That is what makes backward induction possible: a two-dimensional grid, or a regression on in a Monte Carlo, and the exercise decision can be made state by state.
Third, today’s curve enters only through the ratio , which is read straight off chapter 7’s bootstrap. The initial curve is matched exactly and automatically, with nothing to calibrate.
Example 12.3 (Climbing down the ladder to Hull-White).
The quickest way to believe Theorem 12.6 is to take the one case where we already know the answer.
Set , a constant. The equation for in (12.13) has no randomness in it at all, so it is an ordinary linear differential equation, with . Solving,
which is a known function of time and not a state variable at all. The model has collapsed to one dimension.
What is left is
a Gaussian Ornstein-Uhlenbeck process with a time-dependent drift, and since , the short rate is
This is exactly the Hull-White short rate equation from chapter 8, with assembled from the initial curve and the model’s own accumulated variance — which is what chapter 8 found there too, by a completely different route.
Remark (Where Hull-White sits).
Set deterministic. Then is a deterministic function of time and drops out as a state variable, leaving a one-factor Gaussian model with an affine bond formula — which is the Hull-White model of chapter 8, met there through the short rate and arrived at here from the other direction. The ladder is:
HJM (general, infinite state) impose exponential separability quasi-Gaussian (Markov in ) make deterministic Hull-White.
Each arrow is a restriction bought for a computational gain, and knowing which one you are standing on is most of knowing what your model can and cannot do.
12.3 Excursion: Affine Models and Why They Solve
Formula (12.14) has a particular shape — an exponential of something linear in the state, plus a correction — and it is not a coincidence. It is an instance of a general structure that accounts for essentially every interest rate model with a closed form.
Definition 12.7 (Affine term structure model).
A model with state has an affine term structure if
for deterministic functions and .
The question is which state processes produce this, and the answer is clean.
Theorem 12.8 (Duffie and Kan).
Suppose the short rate is affine in the state, , and the state is a diffusion whose drift and covariance are both affine in the state:
Then the model has an affine term structure, and and solve a system of ordinary differential equations — Riccati equations — in the maturity.
Proof.
The bond price satisfies the pricing equation of chapter 5 with this model’s generator — discounted, it is a martingale — which written out for a multi-factor state is
Substitute the guess . Every derivative brings down a factor of the exponential, which cancels throughout, leaving
where the dots are derivatives in . Now the key step: this must hold for every value of , and every term is either constant in or linear in it. A polynomial vanishing identically has vanishing coefficients, so the constant and linear parts must separately be zero:
with and . These are ordinary differential equations in one variable, quadratic in the unknown — Riccati equations. ∎
Remark (What “solvable” actually means here).
The pricing problem for a general model is a partial differential equation in space dimensions plus time, and solving it costs work that grows exponentially in . What the affine structure achieves is to reduce that PDE to a system of ordinary differential equations, which cost essentially nothing and are often available in closed form.
The mechanism is the separation in the proof: because the coefficients are affine and the guess is exponential-affine, the state variable appears only linearly and can be matched off, leaving equations for functions of maturity alone. The state has been eliminated from the problem entirely.
The condition is genuinely restrictive. The drift being affine is mild. The requirement that the covariance be affine in the state is what excludes almost everything: a volatility proportional to makes the covariance proportional to , and the match fails. This is why the tractable models are the Gaussian ones ( constant) and the square-root ones (, so the covariance is proportional to ) and essentially nothing else. Heston’s variance process in chapter 10 is a square-root process for exactly this reason, and its closed-form characteristic function is Theorem 12.8 applied to the log-price rather than to a bond.
Structure (Why affine gives Riccati, in one line).
The system above is usually presented as what falls out of substituting an ansatz. See why the ansatz works, because the reason generalises and the algebra does not.
Apply the generator of chapter 3 to an exponential, , for a diffusion whose drift is and whose variance is :
The generator has mapped an exponential-affine function to an affine multiple of itself. It does not say the family is mapped into itself: is not of the form , so is not preserved by . What has happened is better suited to the purpose.
Regard the family as a two-dimensional surface in function space, with coordinates . Differentiating along each coordinate gives the tangent directions,
so the tangent space at any point of the surface is precisely the set of affine multiples of that point. The display above therefore says that carries each point of the surface into its own tangent space: the vector field is tangent to the surface.
That is exactly the condition for the surface to be invariant under the flow , which is the object we actually care about — invariance under the semigroup, not under the generator. A flow tangent to a two-dimensional surface is a pair of ordinary differential equations in its coordinates, whatever the underlying problem was. And because both sides of now lie in the same tangent space, they can be compared in its basis , which is all the next line does. Chapter 14 takes this distinction as its starting point and makes it general.
Reading the bracket gives them directly. The constant term is and the coefficient of is , so
and the second is quadratic in — a Riccati equation — for the single reason that the variance was allowed to depend on the state. Vasicek has , the quadratic term disappears, and the equation is linear; the square-root model has and it is not.
So “affine” is not a description of the coefficients so much as a statement about the generator, and the Riccati equation is not a coincidence but the shadow of a two-dimensional invariant family. Chapter 14 asks which other families work.
Example 12.4 (Solving the Riccati equations for Vasicek).
The abstract statement is more convincing once the equations have actually been solved once, and in one dimension they can be.
Take , , and : the Vasicek model. In the notation of Theorem 12.8, , , the covariance is the constant so and , and , . Writing so the equations run forward in from zero, they become
Because the first equation is linear — the quadratic term of the Riccati equation has vanished — and it integrates immediately:
which is from (12.14), arrived at yet again. The second equation is then a plain integral of known functions, giving in closed form.
So a bond price in the Vasicek model costs two evaluations of an exponential. This is what Theorem 12.8 is worth: the same problem posed as a partial differential equation would have to be solved numerically on a grid, for every maturity, every time the parameters moved.
Notice also exactly where the linearity came from. It came from , that is, from the covariance not depending on the state — the Gaussian case. In the square-root case , the quadratic term survives, and solves a genuine Riccati equation. It still has a closed form, which is why the square-root process is the other tractable one, but the algebra is a page rather than a line.
Remark (Quadratic Gaussian models).
There is one more family that solves, and it looks at first like a counterexample. Take to be an ordinary Gaussian process — constant — but make the short rate quadratic in it:
Guessing and repeating the argument of Theorem 12.8, the exponential again cancels and one matches the constant, linear and quadratic coefficients in , obtaining ordinary differential equations for , and a matrix Riccati equation for . The model solves.
It is not really an exception. Adjoin the products to the state; then is affine in the enlarged state, and the enlarged state’s drift and covariance are affine in it too, because is Gaussian. Quadratic Gaussian models are affine models wearing a smaller state.
What one buys is a smile. With non-zero, the rate is a quadratic function of a Gaussian variable, so it is not Gaussian, and it is bounded below if is positive definite — which gives a distribution with genuine skew and curvature while keeping closed-form bonds.
12.4 Where the Smile Lives
Return now to the quasi-Gaussian state equations (12.13). Nothing so far has said what is, and this is where the volatility chapters come back.
The model is called quasi-Gaussian because it is Gaussian whenever is deterministic, and departs from Gaussian only through the state dependence of .
| name | what it gives | |
|---|---|---|
| Hull-White | Gaussian, closed forms, no smile | |
| linear quasi-Gaussian | a skew; is the skew dial | |
| quadratic quasi-Gaussian | skew and curvature: a genuine smile | |
| stochastic-vol quasi-Gaussian | vol-of-vol, forward smile, real vega |
The state is a one-dimensional diffusion with volatility , so it is a local volatility model — of the state rather than of the underlying, but the same object chapter 9 studied. Two results carry straight over.
Gyongi’s theorem says the distribution of depends on the volatility only through its value at each level, so the shape of as a function of is precisely what shapes the distribution, and therefore the smile. Nothing else about the model can affect it.
And chapter 10’s midpoint rule says how: the implied volatility at a strike is averaged over the journey from the forward to that strike. Averaging preserves shape and halves slopes. So a that is
-
-
constant averages to a constant, and the smile is flat;
-
-
linear in averages to something linear, and the smile is a straight tilt at half the slope;
-
-
quadratic in averages to something quadratic, and the smile acquires curvature.
The quadratic term is the first shape that survives averaging as a bend, which is why it is the first one that produces a smile rather than a skew.