Stochastic Optimal Control
1. Making a Market: the HJB a Naive Quoter Is Ignoring
Inventory Risk, the Equation That Prices It, and the Quotes That Follow
Preprint · Stochastic Optimal Control
20-SEP-2026 · 21 pages · PR-2026-29694967
A market maker who posts a fixed spread around the mid earns that spread on every fill and carries whatever inventory the fills leave behind. We treat the two halves of that sentence as two problems. The naive symmetric quoter is solved in closed form: its expected profit is $2AT\delta e^{-k\delta}$, maximised at a half-spread of $1/k$, and its inventory is a symmetric random walk whose variance grows linearly in the horizon, so that inventory risk outweighs spread-capture risk by a factor $\sigma^2 T / 2\delta^2$ that grows without bound. The optimal quoter is the solution of a Hamilton--Jacobi--Bellman equation which, under an exponential-utility ansatz, collapses to a linear system of ordinary differential equations. We solve that system exactly rather than asymptotically, and two things follow. The optimal quotes are stationary: they settle within a fraction of a unit of time and do not depend on the horizon thereafter. And the closed form in general use is the short-horizon expansion of this solution, wrong by twenty-one per cent in the spread and eighty-seven per cent in the inventory skew at parameters where it is routinely applied. Simulation of all three quoters on the same order flow, scored by certainty equivalent, confirms the ordering and measures the cost: the approximation gives up fifty-six units of certainty equivalent at a horizon of four, and the naive quoter turns negative.
2. A Leverage Cap on an Inconsistent Investor
Equilibrium Mean-Variance Control under a Stochastic Opportunity Set
Preprint · Stochastic Optimal Control
22-AUG-2026 · 35 pages · PR-2026-51037481
The mean–variance criterion does not satisfy Bellman's principle, so "optimal" splits into three incompatible strategies: a pre-commitment plan that requires a commitment device, a naive agent who re-solves continually and follows none of his plans, and an equilibrium from which no future self wishes to deviate. We compute all three for a market whose Sharpe ratio is an Ornstein–Uhlenbeck factor, under a leverage constraint of the kind every mandate carries.
Unconstrained, the equilibrium value is affine in wealth and the policy splits into a myopic term and a hedging term that exists only because the problem is inconsistent. The correlation between the asset and the factor enters through that term and nowhere else: an uncorrelated factor leaves the policy unchanged to within solver noise, however volatile it is. The term's share of the policy is governed by $\kappa T$, the number of times the opportunity set turns over inside the horizon, falling from thirty-six per cent to six across a factor of twenty-four — time-inconsistency costs something only when there is something to hedge over one's horizon.
The leverage cap destroys that structure. The binding set has two components and occupies a third of the state space throughout the horizon, while the value it destroys falls by an order of magnitude; a mandate is therefore expensive in proportion to the runway it removes rather than to how often it binds. It also changes the policy in regions it never touches, by nearly four per cent at wealths well clear of the boundary. Simulated on common noise under the cap, the three strategies rank naive, pre-commitment, equilibrium — the reverse of what unconstrained optimality suggests, separated by tens of standard errors and by under one per cent of objective. A constraint charges a strategy in proportion to the leverage it wanted, and the most modest rule survives it best.
3. Accountability Collapse under Wealth Growth: A Stochastic Control Problem with Integral Equations
We formulate the erosion of social accountability $\Omega(t)$ under growing private wealth $R(t)$ as a stochastic optimal control problem on $\mathbb{R}_+$. The wealth process follows a pure Kou double-exponential jump model with no diffusion and no drift: two independent compound Poisson processes with rates $\lambda_1, \lambda_2$ and exponential jump sizes $\mathrm{Exp}(\eta_1)$, $\mathrm{Exp}(\eta_2)$, capturing large upward and small downward jumps respectively. We derive the Hamilton-Jacobi-Bellman integro-differential equation for the value function $V(R,\Omega)$, convert it to a Fredholm integral equation of the second kind with bilateral exponential kernel $K(R,x)$, and identify the Wiener-Hopf structure on the half-line via the rational symbol $\Phi(\xi) = 1 - \hat{k}(\xi)$. The Kou model's quadratic numerator yields an analytic factorisation $\Phi = \Phi^+ \Phi^-$, reducing the problem to a second-order ODE whose characteristic roots $\beta_1, \beta_2$ are the Cramér-Lundberg exponents, giving the explicit solution $V(R) = A\,e^{-\zeta R} + V_p(R)$ with $\zeta = |\beta_2|$. The optimal control is bang-bang: a justice curve $\mathcal{J} = \{R^{\gamma}\eta = \Omega P\}$ separates prosocial from antisocial behaviour, and above $\mathcal{J}$ accountability collapse $\Omega \to 0$ is the rational optimum as $R \to \infty$.
4. Optimal Stopping and the Snell Envelope
We study the optimal stopping problem in continuous time, where an agent chooses a stopping time to maximise the expected value of a payoff process, following the classical framework of Snell (1952) and its continuous-time extension via the Doob–Meyer decomposition. The value function is characterised as the Snell envelope — the smallest supermartingale dominating the payoff — whose generator satisfies a Hamilton– Jacobi–Bellman variational inequality of obstacle type. The optimal stopping time is the first entry into the stopping region, where the value function equals the payoff, and the continuation region is determined by the strict inequality V > g. As a canonical application, we solve the perpetual American put option, obtaining the closed-form exercise boundary and value function, and illustrate how the exercise boundary moves with time-to-expiry under the finite-horizon formulation.
5. Russian Options: HJB and Reflected BSDE — Two Methods, One Price Surface
The Russian option, introduced by Shepp and Shiryaev (1993), is a perpetual American lookback contract whose payoff is the running maximum of the underlying asset. We price it via two independent frameworks: a Hamilton-Jacobi-Bellman (HJB) variational inequality and a singly reflected backward stochastic differential equation (BSDE) with barrier $M_t$. The HJB approach yields a closed-form solution through a scalar Euler ODE, while the BSDE Bellman iteration — solved by Gauss-Hermite quadrature — converges to the same value function with mean relative error $0.10\%$. A third result ties the two together: the expected stopping time $u(x) = \mathbb{E}[\tau^* \mid X_0 = x]$ satisfies a companion Poisson equation $\mathcal{L}u = -1$ under the same operator and boundary conditions, and its closed form reveals that the option premium equals, to leading order, the discount rate times the expected wait times the current value.