Publication: Optimal Market Making in Prediction Markets: Empirical Distribution Learning, Analytical Policy Characterization, and Reinforcement Learning
Open/View Files
Date
Authors
Published Version
Published Version
Journal Title
Journal ISSN
Volume Title
Publisher
Citation
Abstract
When should a market maker quote given fill probabilities, queue-clearing dynamics, and adverse-selection risk estimated from real market data? Understanding this question better allows us to understand optimal market making in prediction markets. In this thesis, we attempt to answer this question in the setting of NBA in-game markets on Kalshi. We approach this question from three different angles. We learn an empirical distribution of market events, analytically characterize optimal policy under this empirical distribution, and learn optimal policy using reinforcement learning. To model our environment, we first fit Hawkes timing models and log-normal size distributions to describe the underlying distributions of events, trades, placements, and cancellations. We make these distributions conditional on the quarter, odds bucket, and a discrete market regime to capture the evolving game state. Then, we use these fitted distributions to analytically derive a benchmark which has closed form quoting thresholds as a function of these distributions based on the hazard of queue clearing versus adverse-selection risk. Finally, we simulate and train PPO agents that quote either on the BBO or both the BBO and inside spreads in our market environment. We analyze these agents and their learned policy relative to our analytical benchmark. Our analysis reveals that optimal quotes vary significantly based on game state. We see that event size jumps in the latter quarters and specific market regimes which generally correspond to closer games. We also see that buffer regimes with faster queue-clearing present the highest risk of staleness. We find that our analytical benchmark is able to capture a lot of the structure of the optimal policy. It teaches us that quoting is only beneficial when the queue-clearing happens with high frequency in comparison to adverse-selection risk and is significantly less desirable as the game progresses and you enter riskier regimes. Finally, our RL agents are then able to recover most of these insights while also taking advantage of the rich environment to balance the benchmark intuition and create advanced policies that account for inventory management, sequential decision-making, and depth choice. At riskier times, instead of removing quotes like the single-level agent learns, the multi-level agent learns to quote deeper to compensate for risk. Overall, this thesis showcases how one can use analytical modeling and reinforcement learning together to understand optimal market making in prediction markets. The benchmark allows us to see the intuition behind optimal policy, but the learned RL policies show us how the sequential nature of the problem shifts what actions are optimal.