Strip away all the technical details—at every moment, a market maker only makes three decisions, and there are always just these three:

1. How much is this worth today, exactly? (fair value)

2. How far do I place orders on either side of this number? (bid-ask spread)

3. When you already have inventory on hand, which direction does the entire quote tilt? (tilt)

50 years of market-making theory is a single, long attempt to answer: what should these three numbers be, and why? This piece walks the whole path: from a question in 1968 to two formulas that the trading floor still teaches today. The ending includes no strategy promotion. Not investment advice.

I. Why spreads exist: two roads to losing money

First ask a naive question: if both sides are quoting orders, how could anyone possibly make money?

The first decent answer comes from Demsetz, 1968: the spread is about immediacy pricing. If someone wants to get filled right now, they have to pay a convenience fee to someone willing to wait. Market making is professional waiting; the spread is his rent.

If the story ended here, market making would be free money. It isn’t. Because the person placing the order has exactly two systemic ways to lose money—and the spread must cover both at the same time.

Loss channel one: the information tax. The people eating your orders come in two types. One type has something to do—rebalancing, panic, a slip of the hand. Their fills are your profits. The other type knows something you don’t; their fills are targeted losses: they take your sell orders because the price is about to rise. You can’t tell in advance who is who, so each fill comes with an expected loss containing the belief: “my counterparty might be smarter than me.” That’s why quoting too narrow is guaranteed to bleed: a narrow spread draws in informed order flow, but the rent collected can’t cover that tax.

Loss channel two: the inventory tax. Even if every counterparty is uninformed, once your buy is hit, you become long. Then the price continues to wander while you sit there waiting for the next reversing trade. Someone has to pay for that variance you’re stuck holding. The bigger the position, the bigger the volatility, and the longer you expect to hold it—the larger the bill.

So: spread = convenience fee − information tax − inventory tax. Every model in the literature chooses which of these items gets to be the star.

II. Fifty years, one paragraph

1968, Demsetz: the spread gives immediacy pricing, and market making becomes a business that can be analyzed. 1976 to 1981, Garman, Stoll, Ho and Stoll: the inventory school. A market maker is risk-averse and manages inventory via skewed quotes; Ho and Stoll in 1981 already wrote down the modern form of the formulas, but the math was so heavy back then that nobody could actually use it. 1985, Kyle, and Glosten & Milgrom, in the same year: the information school. Even if a risk-neutral, zero-cost market maker mixes with informed traders in the order flow, he must quote a spread; the spread is adverse-selection pricing, and the width scales with the fraction of informed order flow. 2008, Avellaneda and Stoikov: rewriting the inventory school using stochastic control for the era of electronic limit order books, providing closed-form solutions that you can directly copy into code. 2013, Gueant, Lehalle, and Fernandez-Tapia add a hard inventory limit and cleaner asymptotic solutions. After that—queue position, microprice, and rebate games—everything is tinkering on top of this skeleton.

One thing must be singled out: Avellaneda–Stoikov is a pure inventory-school model; it assumes the information tax equals zero. That’s its biggest lie, and it’s also the key to judging what kind of order books it can survive in.

III. Two formulas

The model world has four assumptions: the fair price follows a random walk, with volatility σ that is unaffected by your trades (equivalently, there is no informed counterparty); the probability your quote gets hit decays exponentially with distance δ, with decay rate κ; you are risk-averse, with risk aversion coefficient γ; and there is an end time—at the deadline you liquidate and settle accounts.

Formula one: the reservation price. Ask: you already hold q units—how much does one more unit do for you?

r = s − q·γ·σ²·τ

s is the market fair price, and τ is your remaining holding period. Read it aloud: every unit of long inventory drags your private anchor below the market anchor by a distance equal to the risk aversion coefficient times the variance you’re about to endure. Then you place both sides of your quotes around r, not around s. Long inventory pushes both bid and ask down: sell orders move closer to the market (making unloading easier), and buy orders retreat (making it harder to keep getting filled). Skew isn’t an after-the-fact risk-control patch; it drops out directly from the utility function.

Formula two: the optimal half-spread. Put risk aside first, and just maximize expected revenue per unit time—distance times the fill rate, i.e., δ times e^(−κδ). Take the derivative and you get something quite pretty:

δ* = 1 / κ

The optimal spread is the inverse of the order-flow’s price sensitivity. If the people who take orders are picky (large κ), you must quote narrow; if they don’t care about price (small κ), you can quote wide. Market making is a monopolist of immediacy, and κ is the demand curve it faces. The full solution adds the risk term back in:

δ* = γσ²τ / 2 + (1/γ)·ln(1 + γ/κ)

The first term is the inventory tax, spread across both sides of the quote. The second term is the monopoly markup; as γ approaches zero it returns to 1/κ. The entire spread is just one insurance premium plus one monopoly rent—nothing else.

IV. What are these symbols actually saying

In every formula, σ is squared—never to the first power. Variance accumulates linearly over time, so carrying a position through chaos hurts quadratically: volatility doubles, and the model quadruples the effects of both your skew and your risk premium. The trader instinct to “pull back when the tape gets messy” isn’t mysticism; it’s a second-order term.

τ, in the original paper, is the time to settlement—the time until the close. A 24/7 market has no closing bell, so the correct reading becomes: given the current execution rate, how long until a reversing trade comes along to take this batch of inventory off my hands? On a thin order book, that answer is measured in hours, not seconds—and the inventory tax becomes immediately more expensive. This is the biggest quantitative difference between “quote a mainstream asset” and “quote a niche asset.”

γ is the only parameter you choose rather than estimate, and it has units: one divided by money. It is the exchange rate that converts variance into “how much it hurts in dollars.” That means there is no standard value—at the same psychological tolerance, if your account is ten times larger, the corresponding γ is ten times smaller. In practice, decide first how many basis points you want your quote to skew when inventory is fully loaded, then back out γ. What it encodes is your discipline, not the structure of the market. And precisely for that reason, it should be written on paper before you go live, not adjusted while you’re running.

κ is the most information-rich parameter in the model. Its reciprocal has a physical translation: on average, how deep into the order book a market order will “plow.” κ is also a profile of your counterparty: large κ means the order flow is picky and professional, staying close to buy at one and sell at one; small κ means the incoming orders basically don’t care about price—panic, forced liquidation, or mistyped orders. There’s also an empirical fact that holds across every order book people have measured: decay near the top of book is exponential, but the tails are fatter than exponential. Orders resting deeper are hit more frequently than the formula predicts, and almost always by order flow that ignores price.

A, the base arrival rate, is the venue’s traffic—the flow of cars. The elegant part is that in optimization, A is completely absorbed—your optimal distance has nothing to do with it. Trading volume determines whether this business is worth doing; κ determines how it should be run. A whole bunch of venues died on the first question, but their structure on the second question was actually quite beautiful.

V. What the model refuses to tell you

Five silences—each one ended a real professional career.

Information. The whole Glosten–Milgrom world is absent here. If your trades can predict prices, the model’s core assumption breaks and the machine bleeds exactly as designed. The diagnostic tool is markout: after each trade, re-mark at fixed time points. If markout is negative, the information tax is real, and your bid-ask spread didn’t cover it.

Queue position. At the same price, the orders in front of you execute first. The hit rate the model “sees” is the hit rate at the very front of the queue. If you ignore this in backtesting, you’ll log a pile of fills you could never actually get.

Jumps. In the model, the price path is continuous. A real order book will continuously breach your quotes on the same side within seconds. The inventory limit is not an optional decoration; it’s the only thing still standing between you and the tail.

Holding costs. On perpetual contracts, even if the price goes nowhere, inventory isn’t free to hold: funding is settled by the hour. Depending on the direction, it’s either a tailwind subsidizing one side of your book, or a leak that slowly eats away at the spread.

Delay and fees. The formulas assume quotes teleport instantly and cost nothing. Every real venue charges you back in milliseconds and order-book fees—and these two corrections point in the same direction: downward.

VI. An honest summary

Avellaneda–Stoikov won’t hand you an edge. It assumes away exactly the thing that kills most quote strategies—informed order flow. What it gives you is the correct skeleton: an anchor, a width, a skew, an upper bound—everything derived, not revered. The experience parameters (σ, κ) can be measured from data; the discipline parameters (γ, inventory limit) can’t be measured—you can only choose them—and they belong on the page before your very first order. Running with limited samples and tweaking as you go has another name: fitting noise.

This math is fifty years old, and it’s public. Even today it still distinguishes professional players from tourists. Not because it’s secret, but because the model tells you accurately: after the textbook is over, which remaining questions are left—who is my informed order flow, where is my queue, and what happens in the tail? Those who answer these three questions with measurement survive; those who answer them with hope are the counterparty’s meal.

References: Demsetz 1968 (QJE); Ho and Stoll 1981 (JFE); Kyle 1985 (Econometrica); Glosten and Milgrom 1985 (JFE); Avellaneda and Stoikov 2008 (Quantitative Finance); Gueant, Lehalle, Fernandez-Tapia 2013; Cartea, Jaimungal and Penalva 2015 Chapter 10 has the full derivations.

#做市 # Perpetual contract