Chapter 10 Smile Dynamics and Stochastic Volatility
In these notes we show that fitting every option in the market is not enough. We derive, rather than assert, the relationship between the slope of a smile and the way that smile moves, and find that a local volatility model gets it wrong by a factor of two in the wrong direction. We then introduce models in which volatility is genuinely random, and show that the market’s quotes do not determine which one to use.
10.1 What Is Left to Get Wrong
Chapter 9 ended in an uncomfortable place. Dupire’s formula fits every European option in the market exactly, with no fitting and no error, using a model in which volatility carries no randomness at all. If a model reproduces every price we can see, in what sense can it be wrong?
Gyongi’s theorem already told us. A local volatility model matches the marginal distribution of the underlying at each date. It says nothing about the joint distribution across dates, and infinitely many joint distributions share a set of marginals. So the model is guaranteed to be right about anything determined by one date at a time, and is unconstrained about everything else.
-
-
Anything path dependent. A barrier option cares whether the underlying visited a level, which is a statement about the whole path.
-
-
Anything with an early exercise decision. A Bermudan swaption’s value depends on what the smile will look like on each exercise date.
-
-
Anything forward starting. A cliquet struck at the money in a year’s time is a bet on volatility a year from now.
-
-
The hedge. This one is easy to miss and is the most important, because it applies to the plain vanilla options themselves.
The last point deserves spelling out. Suppose we and the market agree exactly on the price of every option today. We then hold one and hedge it. The hedge ratio is
the total derivative of the option’s value with respect to the underlying — and depends on both directly and through the volatility we will use to price it after has moved. Writing and differentiating,
| (10.1) |
The first term is what chapter 5 gave us. The second is vega multiplied by the rate at which the implied volatility of this option changes when the underlying moves. Vega is large — for a one year at-the-money option it is roughly of the underlying per unit of volatility — so the second term is not a refinement.
And is not observable today. It is a statement about the future, and therefore a statement the model has to make. Two models agreeing on every price today will disagree on the hedge unless they also agree about this. That is the subject of the chapter.
10.2 Two Rules of Thumb, and What They Would Mean
Before deriving what a model says, it helps to have the two things a trader might say, because they bracket the answer.
Definition 10.1 (Sticky strike).
The smile is sticky strike if the implied volatility of a fixed strike does not change when the underlying moves: is a function of alone.
Definition 10.2 (Sticky delta, or sticky moneyness).
The smile is sticky delta if the implied volatility depends only on the strike relative to the forward: . The smile rides along with the market.
To compare them we need two quantities that are easy to confuse, so let us name them carefully. Both are measured at the money and in log-moneyness , which is the coordinate that makes them comparable.
Definition 10.3 (Skew and backbone).
The skew is the slope of today’s smile in the strike direction, with the forward held still:
The backbone is the rate at which the at-the-money volatility itself moves when the forward moves:
The skew is a photograph. The backbone is a motion. They are different numbers, they are measured on different days. The market lets us observe the first and not the second.
Before the rules, one identity, because these two numbers are constantly confused and the identity is what keeps them apart. There is a third derivative in play — how the volatility of one fixed strike moves — and all three are related by the chain rule.
Lemma 10.4 (The three derivatives).
For any smile,
| (10.2) |
both derivatives on the right evaluated at the money.
Proof.
, so moving the forward moves both arguments. Differentiating,
and the first term is the skew by definition. ∎
Now the two rules.
Under sticky delta, is a constant, so : the at-the-money volatility never moves, however far the market travels. From (10.1), with fixed and , the smile’s contribution to delta is at the money.
Under sticky strike, does not move, so as the forward slides along the fixed smile the at-the-money volatility takes the value of the smile at the new level: , and . The smile’s contribution to delta is zero, since the volatility of the option we hold does not change at all.
So the two conventions bracket the answer: at one end and at the other. In an equity market with a downward skew, , so the two say the at-the-money volatility either stays put or falls a little when the market rises.
We are now in a position to ask what a local volatility model says — and the answer is outside the bracket.
10.3 The Backbone of a Local Volatility Model
The one ingredient we need from outside is a short-maturity result about how a local volatility function turns into an implied volatility. It has a clean statement and a clean intuition.
Theorem 10.5 (Berestycki, Busca and Florent).
In a local volatility model, as the expiry shrinks to zero the implied volatility converges to the harmonic mean of the local volatility along the path in log-space from the forward to the strike:
| (10.3) |
Why it is true, before why it is a theorem.
Think of the underlying as diffusing through a medium whose diffusivity varies from place to place. The option struck at is asking how hard it is to get from to . In a medium of constant diffusivity , covering a log-distance takes a characteristic time , so the natural measure of separation is , and in a varying medium the separation of and is the accumulated . Implied volatility is by definition the constant volatility reproducing the same option price, hence the same separation, so
which is (10.3) rearranged. Implied volatility is a harmonic average because it is an average of a rate, and rates average harmonically — for the same reason that driving one mile at 30 and one mile at 60 does not average to 45. ∎
Remark (What turning that into a proof requires).
The argument above is not a proof and the gap is not a technicality, so it pays to see where the work is. The rigorous chain has four links.
One. The short-time behaviour of the transition density is governed by a distance. Varadhan’s theorem says
where is the Riemannian distance in the metric — precisely the separation the heuristic wrote down. So the heuristic names the right object; what it does not do is establish the limit. The geometry that this metric carries is taken up later in the chapter.
Two. The same asymptotics transfer from the density to the option price, since an out-of-the-money call is an integral of the density over a region whose nearest point to the forward is the strike, and a Laplace-type estimate says an integral of is governed by the smallest in the region.
Three. Black-Scholes with a constant volatility is the special case , where the distance is . Equating the two distances is the theorem.
Four. Step one is classically available for a uniformly elliptic operator with smooth coefficients on a compact space, and none of those three conditions holds here: the operator degenerates as , the domain is not compact, and coming out of a calibration is at best continuous. The route that survives this substitutes directly into the pricing equation, which turns a linear parabolic equation into a nonlinear one, and the limit as is a first-order Hamilton-Jacobi equation — an eikonal equation , whose solutions are exactly distance functions. Two facts about that equation then have to be supplied: that the limit satisfies it, and that it has only one solution with the given boundary behaviour. Neither is available in the classical sense, because a distance function is not differentiable where two geodesics meet. Viscosity solutions are the notion of solution designed for exactly that, and the comparison principle for them is what delivers uniqueness.
A degenerate limit of a linear problem became a nonlinear problem, and a nonlinear first-order equation with non-smooth solutions needs a weak notion of solution to be well posed at all. The same substitution and the same difficulty appear wherever a small-noise limit is taken; large deviations theory is this argument in general form, with the exponential rate playing the part of the distance here.
From this we get the fact we actually use.
Lemma 10.6 (The midpoint rule).
To first order in ,
that is, the implied volatility of a strike is the local volatility evaluated halfway between the forward and the strike, measured in log-space.
Proof.
Write , and let
so the integral in (10.3) runs from to and the mean is taken over a width . Expanding about the midpoint,
the term linear in integrating to zero by symmetry. So the harmonic mean is up to a correction of order , and . ∎
The midpoint rule is the whole story in one line: implied volatility is local volatility, averaged over the journey. Averaging halves a slope, and that is where the factor of two lives.
Remark (The same two as chapter 9).
This factor is not a new one. Chapter 9 established, statically, that the local volatility curve has twice the slope of the implied volatility curve — the rule of two — and for the same reason: implied volatility is an average of local volatility over the interval between the forward and the strike, and averaging a function over an interval halves its slope.
The market’s verdict is shared too. Chapter 9 measured the ratio and found it near two only close to the money, decaying further out, because both statements are leading order. And the backbone measured in the market is not either — which is the complaint this chapter is built around, arriving in a second form.
Theorem 10.7 (The local volatility backbone).
In a local volatility model, the at-the-money volatility responds to a move in the forward at twice the rate of the skew:
| (10.4) |
Proof.
The essential point is that is a fixed function of the level of the underlying. It was calibrated once, to today’s surface, and it does not move when the market does. Write it in log-space as . By Lemma 10.6,
Both quantities we want are derivatives of this single expression, taken in different directions.
For the skew, hold fixed and differentiate in , then set :
For the backbone, set first — the at-the-money option is a different contract at each level of the forward — giving , and then differentiate:
Dividing, . ∎
The factor of two comes from a genuinely simple place. The skew moves only the strike, so it moves the midpoint by half as much; the backbone moves the forward and the strike together, so it moves the midpoint by the full amount. Chapter 9 met the same factor from the other side — its rule of two, that the local volatility curve extracted from a smile is twice as steep as the smile itself, is this statement read backwards. The two derivations are independent, and local_volatility_has_twice_the_slope_of_the_implied_smile in quant/src/localvol.rs checks the result numerically against a Dupire formula that shares none of the algebra.
Remark (Why this is bad, in numbers).
Put the three predictions side by side. Take a one year at-the-money option with and a skew of volatility points per unit of log-moneyness, which is an ordinary equity index figure. Suppose the market rises .
| backbone | change in | |
|---|---|---|
| sticky delta | points | |
| sticky strike | points | |
| local volatility | points |
Equity index markets behave much closer to the first row than the third. The local volatility model does not merely get the magnitude wrong; it predicts the at-the-money volatility falls twice as fast as even the more aggressive of the two rules of thumb, and it does so because it was fitted to today’s skew. The steepness was not chosen. It was forced on the model by the requirement to match the smile.
Now feed that into (10.1). With vega around per unit volatility, an error of volatility points per move is a delta error of roughly , about five percent of the underlying’s notional, on every option in the book, in the same direction. That is not a rounding error; it is a systematic position nobody chose to take.
Remark (No vega).
A second failure follows from the same source and can be stated quickly. Because there is no independent volatility factor, the model is complete in the sense of chapter 4: the underlying and the money market replicate everything, so there is no vega to hedge and nothing to hedge it with. A desk whose business is trading volatility needs volatility to be a risk factor, and here it is not one.
10.4 The Backbone the Market Actually Has
The backbone above is a modelling choice presented as a menu. A rate’s volatility is taken proportional to some power of its level,
| (10.5) |
and three values of have names. At the model is normal, so a yield at one per cent moves in basis points exactly as violently as a yield at ten. At it is lognormal and moves ten times as violently. At it is the square-root middle that keeps the rate positive without letting the volatility scale fully. SABR carries the same exponent under the same letter, and it is usually not calibrated at all — it is set to one of the three by convention and the remaining parameters absorb the difference.
These notes can do better than convention here, because the underlying is observable. The curve is quoted every business day and has been for sixty-four years.
Remark (Why a physical measure estimate is admissible).
The objection to measuring from history is the usual one: hedging happens under the pricing measure and a time series is drawn under the physical one, so the estimate appears to answer a different question.
It does not, and chapter 6 says why. Girsanov changes the drift of a process and leaves its diffusion coefficient exactly where it was, so (10.5) — which is a statement about the diffusion coefficient and nothing else — is the same equation under both measures. The premium lives in ’s drift and in the level, not in the exponent. So is one of the few quantities a pricing model needs that can be read off history without a change of measure.
Calculation 10.8 (The exponent, from sixty-four years of the ten-year yield).
Cut the history into non-overlapping months, take the realised volatility of daily changes within each and the average level across it, and regress one log on the other. Overlapping windows would quadruple the apparent sample without adding information and make the standard error a fiction.
Over the whole series the level ranges from to , a factor of twenty-six, and it is that range which identifies the exponent at all — a few months of data cannot see it, however many days they contain. The result is
on monthly windows.111measured. Every one of the three conventional choices is rejected: normal by standard errors, square root by , and lognormal by .