A
Andrew Rodn
Guest
On 22 August 2026, BTCUSDT fell 2.3% in 15 minutes on Binance. A Polymarket "BTC Up or Down" market repriced within seconds: "Up" traded at 0.49–0.63 just before Binance crossed the window's opening price at 05:07:12 and near 0.10 forty seconds later. Over a 60-day sample, VPIN gave no warning of this sweep and only a weak statistical edge overall (a 1.4x lift). One-sided taker flow reacted first, and the decision window on both venues was seconds, which argues for a circuit breaker that reacts to live microstructure rather than one that tries to predict.
Correction, 2 October 2026: An earlier version said Polymarket lagged Binance by about a minute. That came from lining up Polymarket's one-minute price history, whose points are stamped a few seconds after each minute, with the wrong Binance bar. Polymarket's trade-by-trade record shows it repriced within seconds.
Why Market Makers Need a Circuit Breaker
A market maker earns the bid-ask spread on uninformed order flow and loses on informed flow. In high-frequency trading (HFT), these losses aren't spread evenly across the day; they cluster in short bursts:
- News shocks. The fastest participants lift or hit every resting quote before it can be repriced.
- Liquidation cascades. Individual forced sellers are uninformed, but once a cascade starts, it's predictable, and anyone who sees it early trades against stale bids.
- Prediction-market sweeps. Polymarket's short-dated crypto markets trade on a central limit order book (CLOB) and price off spot. When Binance moves, the first trader to react sweeps every quote still priced off the old spot. On a 15-minute binary contract, a stale quote can be off by 60 cents on a $1 payout.
The practical defence isn't forecasting the move. It's pulling or widening quotes within milliseconds of order flow turning toxic, then restoring them once conditions normalise.
Measuring Toxicity
VPIN: Volume-Synchronized Probability of Informed Trading
The PIN model (Easley, Kiefer, O'Hara & Paperman, 1996) splits order flow into uninformed traders, arriving on both sides at rate $\epsilon$, and informed traders, who appear with probability $\alpha$ and trade one-sidedly at rate $\mu$. The probability that a trade is informed is:
$$\text{PIN} = \frac{\alpha\mu}{\alpha\mu + 2\epsilon}$$
Maximum-likelihood estimation of PIN is too slow for intraday use. VPIN (Easley, López de Prado & O'Hara, 2012) approximates it on a volume clock:
- Partition trades into buckets of equal traded volume $V$.
- Split each bucket $\tau$ into buyer-initiated volume $VB[\tau]$ and seller-initiated volume $VS[\tau]$.
- Average the normalised imbalance over the last $n$ buckets:
$$\text{VPIN} = \frac{1}{n} \sum_{\tau=1}^{n} \frac{\vert{}VB[\tau] - VS[\tau]\vert{}}{V}$$
Because the expected imbalance per bucket is about $\alpha\mu$ and the expected bucket volume is $\alpha\mu + 2\epsilon$, VPIN converges to PIN.
On crypto venues, trade direction is observed rather than inferred: each Binance trade carries an isBuyerMaker flag, and each kline reports exact taker-buy notional.
The critical parameter is $V$. The original paper used $V = \text{ADV} / 50$, where ADV is average daily volume, with $n = 50$.
If $V$ is too small relative to typical trade size, each bucket holds one or two trades, and VPIN drifts toward 1 regardless of information content. If $V$ is too large, VPIN barely moves inside an event.
Calibrate $V$ per symbol. Even then, raw VPIN levels differ between pairs because trade-size distributions differ. The comparable quantity is VPIN's rank within its own history (the paper reports it as a CDF). FollowSM publishes this rank as vpin_percentile.
L1 and L2 Orderbook Imbalance
VPIN describes executed flow; the order book describes resting intent. With best bid price and size $p_b, q_b$ and best ask $p_a, q_a$, the top-of-book (L1) imbalance is:
$$I_{\text{L1}} = \frac{p_b \cdot q_b}{p_b \cdot q_b + p_a \cdot q_a} \in [0, 1]$$
L1 is noisy and easy to spoof, so aggregate notional within a band of $\pm k$ around the mid price $m$ (the L2 view):
$$B_{\text{bid}}(k) = \sum p_i \cdot q_i \quad \text{over bids with } p_i \ge m \cdot (1 - k)$$
$$B_{\text{ask}}(k) = \sum p_j \cdot q_j \quad \text{over asks with } p_j \le m \cdot (1 + k)$$
$$\text{imbalance}(k) = \frac{B_{\text{bid}}(k)}{B_{\text{bid}}(k) + B_{\text{ask}}(k)}$$
$$\text{ob\toxicity\1pct} = \frac{B{\text{ask}}(1%)}{B{\text{bid}}(1%)}$$
An ob_toxicity_1pct above 2 means twice as much resting sell notional as buy notional within 1% of mid. The bid is thin, and a modest market sell will walk through it.
Many books are persistently skewed (in a sample of FollowSM's live full-book stream on 2 October 2026, SOLUSDT's ratio sat above 2 in about 80% of frames), so react to changes: flag when the 1% imbalance reaches an extreme tail of its own recent history (ob_imbalance_percentile), or deviates from its exponentially weighted moving average (EWMA) by more than a threshold $\delta$.
Robust Volume Z-Score
Volume is heavy-tailed, so use the median and the median absolute deviation (MAD) rather than mean and standard deviation:
$$z_V = 0.6745 \cdot \frac{V_t - \text{median}(V_{t-k..t-1})}{\text{MAD}(V_{t-k..t-1})}$$
The 0.6745 factor scales MAD to be comparable to a standard deviation under normality. Pro-rate the volume of a still-open candle by elapsed time before comparing it.
Volatility is tracked separately by the normalized average true range (NATR): the average true range divided by price, which makes it comparable across assets.
Cross-Venue Divergence
Crypto prediction markets carry an implied probability $p$ for an outcome that depends directly on spot price $S$. Assign each market a direction $d$: $+1$ if YES is bullish for spot, $-1$ if bearish, $0$ if ambiguous. The venues diverge when:
$$\text{sign}(\Delta S_{15m}) \neq \text{sign}(d \cdot \Delta p_{15m}) \quad \text{and} \quad \vert{}\Delta p_{15m}\vert{} > 0.05$$
Implementation
The examples use the followsm-sdk Python package (Python 3.10+):
Code:
pip install followsm-sdk
The SDK exposes typed pydantic models for Binance microstructure (VPIN, orderbook imbalance, volume Z-score, NATR) and Binance $\times$ Polymarket confluence snapshots.
Reading a Snapshot
Code:
from followsm_sdk import FollowSMClient, RateLimitExceededException
client = FollowSMClient() # unauthenticated: free tier, rate-limited per IP
try:
snap = client.get_confluence_snapshot("BTCUSDT")
except RateLimitExceededException as exc:
raise SystemExit(str(exc))
micro = snap.binance_microstructure
print(
f"vpin={micro.vpin:.3f} "
f"ob_toxicity_1pct={micro.ob_toxicity_1pct:.2f} "
f"spot_15m={micro.price_delta_15m_pct:+.3%}"
)
for event in snap.polymarket_confluence.active_events:
print(
f"{event.market_slug}: p={event.implied_probability:.3f} "
f"prob_delta_15m={event.prob_delta_15m:+.3f} "
f"direction={event.direction}"
)
print("divergence:", snap.composite_signals.cross_market_divergence_flag)
print("server action:", snap.composite_signals.recommended_action)
A Configurable Risk Ladder
evaluate_risk re-derives the recommended action from the snapshot's raw metrics using your own thresholds:
Code:
from followsm_sdk import FollowSMClient, RiskConfig
risk = RiskConfig(
vpin_percentile_widen_threshold=0.85, # -> WIDEN_SPREAD_2X
vpin_percentile_halt_threshold=0.97, # with divergence -> HALT_MAKER_QUOTES
ob_imbalance_percentile_low=0.005, # 1% book in either extreme tail
ob_imbalance_percentile_high=0.995, # of its own history counts as toxic
ob_toxicity_threshold=2.5, # ask/bid ratio fallback while rank warms up
min_semantic_confidence=0.70, # never halt on ambiguously mapped market
)
client = FollowSMClient(risk_config=risk)
snap = client.get_confluence_snapshot("BTCUSDT")
print(client.evaluate_risk(snap))
In order of severity:
- HALT_MAKER_QUOTES: Toxic flow (VPIN percentile at or above the halt threshold, or a toxic 1% book) and cross-venue divergence. Downgraded to WIDEN_SPREAD_1_5X if the Polymarket market's direction confidence is below min_semantic_confidence.
- WIDEN_SPREAD_2X: VPIN percentile at or above the widen threshold, or a toxic 1% book.
- WIDEN_SPREAD_1_5X: Divergence alone.
- NONE: Otherwise.
A Streaming Circuit Breaker
Polling suits a bot that re-quotes every few seconds; a market maker needs every tick. The following consumes an authenticated WebSocket stream of per-symbol metrics:
Code:
import asyncio
import os
from followsm_sdk import FollowSMClient
PCTL_TRIP, PCTL_REARM = 0.90, 0.80
BOOK_TAIL, OB_TOX_FALLBACK = 0.01, 2.0
def lopsided_book(m) -> bool:
p = m.ob_imbalance_percentile # 1% imbalance ranked vs symbol history
if p is None:
# warming up: fall back to a fixed ask/bid ratio
return m.ob_toxicity_1pct > OB_TOX_FALLBACK
return p <= BOOK_TAIL or p >= 1 - BOOK_TAIL
async def cancel_quotes(symbol: str) -> None:
print(f"KILL SWITCH {symbol}: cancelling resting maker quotes")
# call your OMS here
async def main() -> None:
client = FollowSMClient(api_key=os.environ["FOLLOWSM_API_KEY"])
halted: set[str] = set()
async for m in client.stream_toxicity():
pctl = m.vpin_percentile # None while warming up
toxic = (pctl is not None and pctl > PCTL_TRIP) or lopsided_book(m)
if toxic and m.symbol not in halted:
halted.add(m.symbol)
await cancel_quotes(m.symbol)
elif (
m.symbol in halted
and (pctl is None or pctl < PCTL_REARM)
and not lopsided_book(m)
):
halted.discard(m.symbol)
print(f"RE-ARM {m.symbol}")
asyncio.run(main())
The gap between PCTL_TRIP and PCTL_REARM is hysteresis. Without it, a reading oscillating near the trip level would cancel and re-post quotes on every tick.
A production implementation is available as open source (hft-toxicity-circuit-breaker). It adds:
- Depth-imbalance spike detection
- A second evaluation leg on cross-venue snapshots
- A re-arm cooldown
- Fail-closed reconnects, where a dropped stream halts every symbol it guards
- Non-blocking callbacks for cancellations and HMAC-signed webhooks
Its in-process evaluation measured $p50 = 4.7 , \mu\text{s}$ and $p99 = 24.9 , \mu\text{s}$ per frame on live data, so the latency budget is dominated by the network, not the decision.
Case Study: The 22 August 2026 BTC Sweep
Method
86,400 one-minute BTCUSDT klines from Binance (28 July – 26 September 2026), including exact taker-buy notional per bar.
VPIN computed with $n = 50$ in two calibrations:
- Slow: $V = \text{ADV}/50 \approx $23.3\text{M}$, the original paper's setting
- Fast: $V = $1\text{M}$, a short desk horizon
The largest 15-minute drop after a 7-day warm-up, matched with Polymarket's public CLOB price history and trades for the corresponding 15-minute "Up or Down" market. The Polymarket column below is the first price-history point after each Binance minute closed (points are stamped a few seconds past the minute).
Observations
| UTC | Close | Notional | Taker-buy | VPIN fast (7d pctl) | VPIN slow | Polymarket "Up" |
|---|---|---|---|---|---|---|
| 04:55 | 78,541 | $0.5M | 0.72 | 0.236 (23%) | 0.130 | 0.505 |
| 05:00 | 78,532 | $1.0M | 0.69 | 0.228 (21%) | 0.130 | 0.545 |
| 05:03 | 78,544 | $1.2M | 0.48 | 0.211 (16%) | 0.130 | 0.655 |
| 05:06 | 78,528 | $0.6M | 0.66 | 0.200 (14%) | 0.130 | 0.635 |
| 05:07 | 78,286 | $2.9M | 0.17 | 0.210 (16%) | 0.130 | 0.085 |
| 05:08 | 78,138 | $3.1M | 0.37 | 0.207 (16%) | 0.130 | 0.035 |
| 05:09 | 77,802 | $10.9M | 0.40 | 0.200 (14%) | 0.130 | 0.005 |
| 05:10 | 76,742 | $56.3M | 0.34 | 0.326 (63%) | 0.140 | 0.001 |
| 05:11 | 77,087 | $43.4M | 0.49 | 0.062 (0%) | 0.142 | 0.001 |
Over the 15 minutes, $84M traded, a robust volume Z-score of 22.8 against the preceding five hours. The "Up" market resolves Up if BTC ends the 05:00–05:15 window above its opening price.
Findings
VPIN did not anticipate the sweep. Slow VPIN was at the 4th percentile of its trailing week at the start of the window, and fast VPIN at the 23rd. Neither moved until the crash minute, after which fast VPIN fell as two-sided rebound volume arrived. This is consistent with Andersen & Bondarenko (2014), who showed that VPIN's reported lead before the 2010 Flash Crash depended heavily on the trade-classification method.
The conditional edge is weak. Across 5,087 non-overlapping 15-minute windows, top-decile slow VPIN raised the probability of a top-1% move ($\ge 0.70%$) from 0.95% to 1.37%. That's a 1.4x lift resting on roughly five events. With fast buckets, the relationship inverted (0.29% vs 1.09%).
One-sided flow moved first. At 05:07, 83% of taker notional was sell-side, on roughly three times the preceding minutes' notional. Tick-level buckets and 1% book toxicity capture this directly; one-minute bars blur it, and klines carry no book-depth information at all.
Polymarket repriced within seconds. Binance's first one-second close below the window's opening price (78,470.15) came at 05:07:12. Polymarket's trades show "Up" at 0.49–0.63 in the ten seconds before, 0.31–0.48 in the ten after, 0.23–0.30 in the next ten, and 0.08–0.13 by 05:07:42–52. The market settles on Chainlink's BTC/USD feed, so Binance is a proxy here.
Implications
Treat VPIN as a regime input, not a trigger. Use it to widen spreads, and let faster signals (taker imbalance, book toxicity, depth-imbalance spikes, volume bursts) trigger halts.
Calibrate per symbol, and set thresholds from each symbol's own distribution rather than a single raw level.
Optimise reaction latency rather than prediction. Here Polymarket did most of its repricing within about 20 seconds of Binance crossing the opening price, and the Binance book moved faster still.
Monitor the other venue. Divergence between two prices for the same risk is an inexpensive and direct adverse-selection signal.
Reproducibility
The replay script uses only public Binance and Polymarket endpoints and requires no API keys:
Code:
git clone https://github.com/Follow-SM/hft-toxicity-circuit-breaker.git
cd hft-toxicity-circuit-breaker
pip install httpx && python research/vpin_case_study.py
Resources
Python SDK: pip install followsm-sdk (PyPI)
TypeScript SDK: npm install @followsm/sdk (npm)
REST Polling Bot for Polymarket: polymarket-arbitrage-starter-kit
WebSocket Circuit Breaker: hft-toxicity-circuit-breaker
API Access: Free tier at 30 requests/min per IP; DEVELOPER ($199/mo) at 300 requests/min; ENTERPRISE ($499/mo) at 1,000 requests/min with the WebSocket streams used above. Details at follow-sm.com/pricing.
References
- Easley, D., Kiefer, N. M., O'Hara, M., & Paperman, J. B. (1996). Liquidity, Information, and Infrequently Traded Stocks. Journal of Finance.
- Easley, D., López de Prado, M. M., & O'Hara, M. (2012). Flow Toxicity and Liquidity in a High-Frequency World. Review of Financial Studies.
- Andersen, T. G., & Bondarenko, O. (2014). VPIN and the Flash Crash. Journal of Financial Markets.
Disclaimer: This article is educational and not financial advice. Trading crypto assets and prediction markets carries substantial risk of loss.