<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
  <author>
    <name>Farhan Sadeek</name>
    <email>farhan@farhansadeek.com</email>
  </author>
  <generator uri="https://hexo.io/">Hexo</generator>
  <icon>https://www.gravatar.com/avatar/d1b6802ae0c9e31e1433f970bdb5b449</icon>
  <id>https://projects.farhansadeek.com/</id>
  <link href="https://projects.farhansadeek.com/" rel="alternate"/>
  <link href="https://projects.farhansadeek.com/atom.xml" rel="self"/>
  <rights>All rights reserved 2026, Farhan Sadeek</rights>
  <subtitle>Systems, trading and ML projects, built from scratch and measured</subtitle>
  <title>Farhan Sadeek · Projects</title>
  <updated>2026-09-29T06:49:26.552Z</updated>
  <entry>
    <author>
      <name>Farhan Sadeek</name>
      <email>farhan@farhansadeek.com</email>
    </author>
    <category term="python" scheme="https://projects.farhansadeek.com/tags/python/"/>
    <category term="statistics" scheme="https://projects.farhansadeek.com/tags/statistics/"/>
    <category term="experimentation" scheme="https://projects.farhansadeek.com/tags/experimentation/"/>
    <category term="causal-inference" scheme="https://projects.farhansadeek.com/tags/causal-inference/"/>
    <content>
      <![CDATA[<p>I wrote expope, a small engine for the two questions a product data science team answers every week. The first is whether an A/B test moved a metric, and whether the p-value can be trusted. The second is what a new recommendation policy would have scored on traffic the old policy served, using only the old policy's logs. The first half is a DuckDB metrics layer over a raw event log plus from-scratch Welch, delta method, CUPED, sample ratio mismatch, Holm, Benjamini-Hochberg and a mixture sequential probability ratio test (mSPRT). The second half is from-scratch off-policy estimators (IPS, SNIPS, direct method, doubly robust and switch-DR) with bootstrap intervals, run on a synthetic bandit with known truth and on the Open Bandit Dataset from ZOZOTOWN.</p><p>The headline is a simulation. I ran 2,000 A/A experiments, looked at each one 50 times as users arrived, and stopped at the first look where Welch's t-test gave p below 0.05. <strong>32.8 percent</strong> of those experiments with no effect at all were declared significant. The same data tested once at the final look gave 5.25 percent, as it should. The always-valid mSPRT, looking at the same 50 points, rejected <strong>1.95 percent</strong>.</p><p>The second finding is on the off-policy side. A cross-fitted gradient-boosted reward model that looks accurate on average made the direct method <strong>11.0 percent low</strong> on a synthetic bandit, and its bootstrap intervals covered the truth in 0 of 200 replicates. Doubly robust with the same model was 0.02 percent low and covered in 192 of 200. The bias is not a bug. The model shrinks toward the mean, and the new policy concentrates on exactly the context and action pairs where shrinkage pulls predictions down the most.</p><p>On correctness, every off-policy estimator agrees with Open Bandit Pipeline (obp) to within 1e-9 on 25 randomly generated slate problems, and on the real Open Bandit sample the largest gap was 1.7e-18. On 2026-09-27 an independent rerun of the full test suite passed all <strong>38 tests</strong>, including the obp cross-checks at 1e-9. The simulation experiments were not rerun independently, so every simulated rate below is from my own single run.</p><p>One limit belongs up front. The Open Bandit experiment uses the 10,000-row sample that ships with obp, and the random-policy log in it holds <strong>38 clicks</strong>. All five estimators land inside the on-policy confidence interval, which is a sanity check, but the intervals are so wide that this data cannot rank the estimators, and I do not try to.</p><p>Code is in <code>projects/25-cuped-msprt-ope-engine</code>.</p><h2 id="table-of-contents">table of contents</h2><ul><li><a href="#what-i-wanted-to-build">what I wanted to build</a></li><li><a href="#theory">theory</a></li><li><a href="#architecture">architecture</a></li><li><a href="#implementation">implementation</a></li><li><a href="#problems">problems</a></li><li><a href="#experiments">experiments</a></li><li><a href="#results">results</a></li><li><a href="#what-i-would-change">what I would change</a></li><li><a href="#reproducibility">reproducibility</a></li></ul><h2 id="what-i-wanted-to-build">what I wanted to build</h2><p>Most statistics bugs in experimentation do not crash. A variance formula that treats sessions as independent, a dashboard that recomputes a t-test every morning and a reward model with a small average error all still return a number. The only way I know to catch them is to run the code on data where the answer is known and count how often it is wrong.</p><p>So the goal was an engine where every method has a simulation that measures it against truth. False positive rate on A/A tests, power against the noncentral t, interval coverage with a nonzero true effect, false positives under continuous monitoring, variance reduction against the CUPED formula, and bias and coverage of every off-policy estimator on a bandit whose true policy value is computed from 2 million contexts. For the off-policy estimators I also wanted an external reference, so every one is cross-checked against obp on identical arrays.</p><p>I wanted the A/B side to start from a raw event log, not a tidy table of per-user outcomes, because who counts as exposed, which events fall in the window and what a user with no events contributes all change the answer before any statistics run.</p><h2 id="theory">theory</h2><h3 id="welch-and-power">Welch and power</h3><p>For per-user means in two arms, the difference has variance s_c²/n_c + s_t²/n_t. Welch's test uses that directly and gets its degrees of freedom from the Welch-Satterthwaite approximation, so it does not assume equal variances. With equal n and a common σ, the power at effect d comes from a noncentral t distribution with noncentrality d / (σ √(2/n)). That formula is the reference the power simulation is checked against.</p><h3 id="the-delta-method-for-ratio-metrics">the delta method for ratio metrics</h3><p>Clicks per session is not a mean of per-user values. It is the sum of clicks over the sum of sessions, R = x̄ / ȳ, where x and y are per-user totals. Users are the randomisation unit, so the sessions inside one user are correlated, and treating every session as an independent observation understates the variance. A first-order Taylor expansion of R around (μ_x, μ_y) gives</p><figure class="highlight plaintext"><table><tr><td class="code"><pre><span class="line">Var(R) ≈ (1/n) [ Var(x)/μ_y²  -  2 μ_x Cov(x,y)/μ_y³  +  μ_x² Var(y)/μ_y⁴ ]</span><br></pre></td></tr></table></figure><p>which needs only five sums per arm. That matters for the architecture below.</p><h3 id="cuped">CUPED</h3><p>CUPED replaces the outcome Y with Y minus θ(X minus X̄), where X is the same metric measured before the experiment started and θ = Cov(Y, X) / Var(X). Because X is fixed before treatment, the adjustment has mean zero in both arms and the difference in means stays unbiased. The variance of the adjusted outcome is Var(Y)(1 minus ρ²), where ρ is the correlation between X and Y. A pre-period covariate with ρ = 0.7 roughly halves the variance, which is the same as doubling the sample. A covariate with ρ near 0 does nothing, and the experiments below have one of each.</p><h3 id="sample-ratio-mismatch-and-multiple-testing">sample ratio mismatch and multiple testing</h3><p>SRM is a Pearson chi-square goodness-of-fit test of arm counts against the designed split. When it fails, the randomiser or the exposure trigger is broken and no metric result should be trusted. Holm is a step-down Bonferroni that controls the family-wise error rate. Benjamini-Hochberg is a step-up rule that controls the false discovery rate.</p><h3 id="why-peeking-breaks-the-t-test-and-what-msprt-does-instead">why peeking breaks the t-test, and what mSPRT does instead</h3><p>A fixed-horizon test promises that if you test once, at a sample size chosen in advance, you reject a true null 5 percent of the time. It promises nothing about testing repeatedly. Under the null the z-statistic, viewed as a function of sample size, behaves like a scaled random walk, and a random walk eventually crosses any fixed boundary. Every extra look is another chance to cross ±1.96. With 50 looks the chances add up to about a third.</p><p>The mixture SPRT (Johari, Pekelis and Walsh, 2017) replaces the p-value with a likelihood ratio that is safe to monitor. Take the current estimate Z of the effect with variance s², and put a normal N(0, τ²) prior over the effect under the alternative. Integrating the normal likelihood against that prior gives a closed form</p><figure class="highlight plaintext"><table><tr><td class="code"><pre><span class="line">Λ = sqrt( s² / (s² + τ²) ) · exp( τ² Z² / (2 s² (s² + τ²)) )</span><br></pre></td></tr></table></figure><p>Under the null, Λ as a function of sample size is a nonnegative martingale with expectation 1. Ville's inequality says that such a martingale exceeds 1/α at any point, ever, with probability at most α. So the rule &quot;stop the first time Λ reaches 1/α&quot; has type I error at most α no matter how often you look. The always-valid p-value is the running minimum of 1/Λ, and inverting the same inequality in θ gives a confidence sequence that only shrinks.</p><p>The price is power. The test has to budget for all future looks, so at any single sample size it is more conservative than a fixed-horizon test there. The parameter τ sets where along the sample size axis the test is most sensitive. I used τ = 0.1 in outcome units, with outcomes of unit standard deviation and a real effect of 0.05 in the power scenario.</p><h3 id="off-policy-evaluation">off-policy evaluation</h3><p>The logs are rounds (x, a, r) where a context x arrived, a behaviour policy π_b chose action a with known probability (the propensity score), and reward r was observed. The question is the expected reward of a different policy π_e.</p><ul><li><strong>IPS</strong> reweights each logged reward by w = π_e(a|x) / π_b(a|x). It is unbiased when the propensities are right, and its variance grows with the weights.</li><li><strong>SNIPS</strong> divides by the mean weight instead of n, trading a small bias for lower variance.</li><li><strong>The direct method (DM)</strong> fits a reward model q̂(x, a) and averages it over π_e. It has low variance and is exactly as biased as the model.</li><li><strong>Doubly robust (DR)</strong> is DM plus the weighted residual on the logged action, DM + w(r minus q̂(x, a)). If the propensities are right, the correction term has expectation equal to the model's error under π_e, so it cancels that error. If instead the model is right, the correction has mean zero. Either suffices.</li><li><strong>Switch-DR</strong> keeps the correction only on rounds where w ≤ τ, and falls back to the model where weights are large.</li></ul><p>The Open Bandit data is a slate, three items shown per impression, so everything is computed per position following obp's conventions. Each logged round carries a position, and the evaluation policy is an (n, A, L) array of probabilities over A actions at each of L positions.</p><h2 id="architecture">architecture</h2><figure data-figure="diagram:peeking-pipeline"></figure><p>The A/B side is two SQL stages and one Python stage. The contract between SQL and Python is a set of frozen dataclasses of per-arm sums (<code>MeanStats</code>, <code>RatioStats</code>, <code>CovariateStats</code>). Every test takes those, and each dataclass has a <code>from_array</code> constructor, so the same Welch code runs on SQL output and on numpy arrays in the simulations. That is what lets the SQL layer be tested against a polars reference and the statistics be tested against scipy independently.</p><p>The off-policy side reduces every estimator to a mean of per-round terms, with SNIPS as a ratio of two means. One function turns the arrays into a <code>RowTerms</code> of per-round weight, reward, DM term and factual q̂. The bootstrap then only resamples those vectors and never touches the (n, A, L) tensors again.</p><h2 id="implementation">implementation</h2><p>Everything statistical is from scratch. scipy supplies only t, normal and chi-square distribution functions and the noncentral t for power. The package is about 1,370 lines of Python across metrics, stats, ope and sim, plus 585 lines of tests.</p><h3 id="the-metrics-layer">the metrics layer</h3><p>Metrics are SQL expressions over per-user aggregates, so revenue per user, conversion and clicks per session are one line each. The first stage builds one row per exposed user. The part that matters is who counts.</p><figure class="highlight sql"><table><tr><td class="code"><pre><span class="line">first_exposure <span class="keyword">AS</span> (</span><br><span class="line">    <span class="keyword">SELECT</span> experiment_id, user_id, <span class="built_in">MIN</span>(exposed_at) <span class="keyword">AS</span> exposed_at</span><br><span class="line">    <span class="keyword">FROM</span> exposures <span class="keyword">GROUP</span> <span class="keyword">BY</span> experiment_id, user_id</span><br><span class="line">),</span><br><span class="line">units <span class="keyword">AS</span> (</span><br><span class="line">    <span class="keyword">SELECT</span> a.experiment_id, a.user_id, a.variant, f.exposed_at, x.start_at, x.end_at,</span><br><span class="line">           x.start_at <span class="operator">-</span> to_days(x.pre_period_days) <span class="keyword">AS</span> pre_start</span><br><span class="line">    <span class="keyword">FROM</span> assignments a</span><br><span class="line">    <span class="keyword">JOIN</span> first_exposure f <span class="keyword">USING</span> (experiment_id, user_id)</span><br><span class="line">    <span class="keyword">JOIN</span> experiments x <span class="keyword">USING</span> (experiment_id)</span><br><span class="line">    <span class="keyword">WHERE</span> f.exposed_at <span class="operator">&gt;=</span> x.start_at <span class="keyword">AND</span> f.exposed_at <span class="operator">&lt;</span> x.end_at</span><br><span class="line">),</span><br></pre></td></tr></table></figure><p>Only users whose first exposure falls inside the experiment window count. Post-period events count from that first exposure to the window end, and pre-period events from <code>start_at</code> minus the pre-period length up to <code>start_at</code>. The final select wraps every aggregate in <code>COALESCE(..., 0)</code>, so an exposed user with no events contributes a zero rather than vanishing from n. That one function call is the difference between measuring the treatment effect and measuring the effect among users who happened to do something. The tests cover exactly these cases (a user first exposed before the window, a user never exposed, an exposed user with no events) against a polars reference implementation of the same metrics.</p><p>The second stage reduces <code>user_metrics</code> to sums per arm, <code>SUM(y)</code>, <code>SUM(y*y)</code>, and for ratio and CUPED metrics <code>SUM(x)</code>, <code>SUM(x*x)</code> and <code>SUM(x*y)</code>. SRM is checked twice, once on assignments (is the randomiser healthy) and once on exposed users (is the trigger healthy).</p><h3 id="cuped-from-sums-alone">CUPED from sums alone</h3><p>Most CUPED code adjusts per-user arrays. Mine had to run on the five sums the SQL returns, so the adjusted outcome's sum of squares is expanded algebraically.</p><figure class="highlight python"><table><tr><td class="code"><pre><span class="line"><span class="keyword">def</span> <span class="title function_">adjusted</span>(<span class="params">s: CovariateStats</span>) -&gt; MeanStats:</span><br><span class="line">    n = s.n</span><br><span class="line">    sum_adj = s.sum_y - theta * (s.sum_x - n * xbar)</span><br><span class="line">    c = theta * xbar</span><br><span class="line">    sum_adj2 = (s.sum_y2 + theta * theta * s.sum_x2 + n * c * c</span><br><span class="line">                - <span class="number">2</span> * theta * s.sum_xy + <span class="number">2</span> * c * s.sum_y - <span class="number">2</span> * theta * c * s.sum_x)</span><br><span class="line">    <span class="keyword">return</span> MeanStats(n=n, sum_y=sum_adj, sum_y2=sum_adj2)</span><br></pre></td></tr></table></figure><p>θ and x̄ are pooled across both arms, so the adjustment cannot create a difference between arms on its own. A hypothesis test checks that CUPED from sums equals CUPED on explicitly adjusted arrays, and another that θ = 0 reduces to plain Welch.</p><h3 id="the-msprt-in-a-few-lines">the mSPRT in a few lines</h3><p>The log likelihood ratio is vectorised so the peeking simulation can compute it for 200 experiments by 50 looks in one call, and the always-valid p-value is an accumulated minimum along the look axis.</p><figure class="highlight python"><table><tr><td class="code"><pre><span class="line"><span class="keyword">def</span> <span class="title function_">msprt_log_lr</span>(<span class="params">est, var, tau2, theta0=<span class="number">0.0</span></span>):</span><br><span class="line">    d = est - theta0</span><br><span class="line">    <span class="keyword">return</span> <span class="number">0.5</span> * np.log(var / (var + tau2)) + tau2 * d * d / (<span class="number">2.0</span> * var * (var + tau2))</span><br><span class="line"></span><br><span class="line"><span class="keyword">def</span> <span class="title function_">msprt_always_valid_p</span>(<span class="params">est, var, tau2, theta0=<span class="number">0.0</span>, axis=-<span class="number">1</span></span>):</span><br><span class="line">    p = np.minimum(<span class="number">1.0</span>, np.exp(-msprt_log_lr(est, var, tau2, theta0)))</span><br><span class="line">    <span class="keyword">return</span> np.minimum.accumulate(p, axis=axis)</span><br></pre></td></tr></table></figure><p>A test checks the closed form against numerical integration of the normal likelihood over the N(0, τ²) mixture. A streaming <code>MSPRTMonitor</code> keeps the running minimum p-value and intersects each new confidence interval with the previous one, so the sequence only shrinks. It uses plug-in variances, the standard practical choice, not a known-variance version.</p><h3 id="every-estimator-is-a-mean-of-per-round-terms">every estimator is a mean of per-round terms</h3><p>This is the core of the off-policy side.</p><figure class="highlight python"><table><tr><td class="code"><pre><span class="line">w = action_dist[idx, action, pos] / pscore</span><br><span class="line">pi_at_pos = action_dist[idx, :, pos]          <span class="comment"># (n, A)</span></span><br><span class="line">q_at_pos  = q_hat[idx, :, pos]                <span class="comment"># (n, A)</span></span><br><span class="line">dm = (q_at_pos * pi_at_pos).<span class="built_in">sum</span>(axis=<span class="number">1</span>) / pi_at_pos.<span class="built_in">sum</span>(axis=<span class="number">1</span>)</span><br><span class="line">...</span><br><span class="line"><span class="keyword">def</span> <span class="title function_">dr_terms</span>(<span class="params">t</span>):</span><br><span class="line">    <span class="keyword">return</span> t.dm + t.w * (t.reward - t.q_factual), <span class="literal">None</span></span><br><span class="line"></span><br><span class="line"><span class="keyword">def</span> <span class="title function_">switch_dr_terms</span>(<span class="params">t, tau</span>):</span><br><span class="line">    keep = (t.w &lt;= tau).astype(<span class="built_in">float</span>)</span><br><span class="line">    <span class="keyword">return</span> t.dm + keep * t.w * (t.reward - t.q_factual), <span class="literal">None</span></span><br></pre></td></tr></table></figure><p>IPS is <code>w * reward</code>, SNIPS is the pair (<code>w * reward</code>, <code>w</code>), DM is <code>dm</code>. With that shape, switch-DR at τ = infinity is DR and at τ = 0 is DM, and both limits are tested.</p><h3 id="a-bootstrap-that-is-one-matrix-product">a bootstrap that is one matrix product</h3><p>Resampling n rounds with replacement is the same as drawing a multinomial count vector over rounds. One replicate is then a dot product of counts with the per-round terms, and a chunk of replicates is one matrix product.</p><figure class="highlight python"><table><tr><td class="code"><pre><span class="line">counts = rng.multinomial(n, p, size=b).astype(<span class="built_in">float</span>)</span><br><span class="line">top = counts @ num</span><br><span class="line">reps[s : s + b] = top / (counts @ den) <span class="keyword">if</span> den <span class="keyword">is</span> <span class="keyword">not</span> <span class="literal">None</span> <span class="keyword">else</span> top / n</span><br></pre></td></tr></table></figure><p>For SNIPS this recomputes the ratio inside every replicate, so the interval includes the variability of the normaliser. I also wrote <code>bootstrap_ci_obp_compatible</code>, which reproduces obp's random draws exactly, but it resamples already-normalised SNIPS terms and so ignores that variability. It exists for the cross-check, not for the experiments.</p><h3 id="the-reward-model-is-cross-fitted">the reward model is cross-fitted</h3><p>The reward model is one LightGBM classifier for all actions, with the action index as a categorical feature, so items share strength. That matters when Open Bandit has 80 items and a few dozen clicks. Rows are split into three folds, and each fold is predicted by a model trained on the other two, so no round is ever scored by a model that saw its reward. Without cross-fitting, the model's memorisation of the logged rewards leaks straight into DR's residual term. The default is 200 trees at learning rate 0.05, 15 leaves and at least 50 rows per leaf. Those regularisation settings matter for the DM bias story below.</p><h3 id="thompson-sampling-slates-vectorised">Thompson sampling slates, vectorised</h3><p>The ZOZOTOWN evaluation policy is Bernoulli Thompson sampling with a production Beta prior per item. It does not depend on context, so its slate distribution is computed once. Each simulation draws θ_a from Beta(α_a, β_a) for all 80 items, sorts, and records which items land in the top three slots. obp does this in a Python loop. I do 20,000 simulations at a time with one <code>argsort</code>, and then broadcast the resulting (80, 3) distribution to every round.</p><h2 id="problems">problems</h2><h3 id="1-the-direct-method-was-11-percent-low-with-a-good-model">1. the direct method was 11 percent low with a good model</h3><p>This was the surprise of the project and it is now the main off-policy result. The cross-fitted LightGBM direct method came out 11.0 percent below the true policy value, averaged over 200 replicates, with bootstrap intervals that never covered the truth. My first suspicion was a feature bug, such as the action column being misaligned when scoring counterfactual actions. I checked one replicate by hand. The per-action averages of q̂ were close to the per-action averages of the true q, and yet the direct method was well below truth. (That hand check is in the devlog, and its numbers were not saved to <code>results/</code>, so I do not quote them here.)</p><p>The explanation is selection on shrinkage. The synthetic bandit's evaluation policy is a softmax of 3 times the log of the true click probability, mixed with 10 percent uniform, so π_e(a|x) is roughly proportional to q(x, a)³. It piles its mass onto the specific context and action pairs where the true reward is highest. A regularised tree model does the opposite at exactly those points. With at least 50 rows per leaf and 200 shallow trees, the model's predictions are pulled toward the mean, so where q is unusually high, q̂ is too low, and where q is unusually low, q̂ is too high. Those errors roughly cancel in a per-action average. They do not cancel in the direct method, which is</p><figure class="highlight plaintext"><table><tr><td class="code"><pre><span class="line">DM - V = E_x [ sum_a  π_e(a|x) · (q̂(x,a) - q(x,a)) ]</span><br></pre></td></tr></table></figure><p>and weights the errors by π_e, which is large where the error is negative. Increasing capacity to 600 trees, 31 leaves and 20 rows per leaf barely moved it, according to the devlog, because the model is still regularised and still sees very few logged rounds in the extreme regions π_e cares about.</p><p>DR fixes this without a better model. Its correction term, w(r minus q̂(x, a)) on the logged action, has expectation E_x[Σ_a π_e(a|x)(q(x, a) minus q̂(x, a))] when the propensities are correct, which is exactly the negative of the error above. I kept the default model and reported the bias, because this is the textbook reason DR exists, and a concrete demonstration of it is worth more than a tuned model that hides it.</p><h3 id="2-switch-dr-looked-broken">2. switch-DR looked broken</h3><p>Switch-DR at τ = 2 came out 8.5 percent low, which looked like a bug in the threshold logic. It is not. In the first replicate, 13.8 percent of rounds had importance weights above 2, and on those rounds switch-DR drops the correction and uses the model alone. Those rounds are not a random 13.8 percent. A large weight means π_e likes the logged action much more than π_b did, which means a high-q action, which is where DM's error is concentrated. So switch-DR drops the correction on a seventh of the rounds and keeps roughly three quarters of DM's bias. On Open Bandit, where only 1.04 percent of weights exceed τ = 10, it behaves well. The lesson is that τ must be chosen with the weight distribution in view, and ideally tuned from data.</p><h3 id="3-matching-obp-s-bootstrap-bit-for-bit">3. matching obp's bootstrap bit for bit</h3><p>To cross-check intervals as well as point estimates I needed obp's exact random stream. obp calls legacy <code>RandomState(seed).choice(samples, size=n)</code>, and for uniform sampling with replacement that draws <code>randint(0, n, size=n)</code> indices internally. Reproducing that call gives bootstrap bounds equal to obp's within 1e-12 in the test. It also made me look at what obp resamples, which for SNIPS is the already-normalised terms, and that is why the experiments use my multinomial version instead.</p><h3 id="4-the-thompson-sampling-slate-cannot-match-to-1e-9">4. the Thompson sampling slate cannot match to 1e-9</h3><p>Both libraries compute the BTS slate distribution by Monte Carlo with different generators, so they will never agree to machine precision. I compare them in total variation instead, which came out 0.0118, 0.0093 and 0.0120 for the three slots at 100,000 simulations. For the estimator cross-check both libraries receive the identical <code>action_dist</code> array. As a check that the Monte Carlo difference does not matter, I reran every estimator with obp's distribution in place of mine, and no estimate moved by more than 1e-5.</p><h3 id="5-small-ones">5. small ones</h3><p>pytest tried to collect the <code>TestResult</code> dataclass as a test class because of its name, fixed with <code>__test__ = False</code>. The first A/A figure had ten long metric names overlapping on the x axis, so it became a horizontal layout, and a <code>--plot-only</code> flag redraws it from saved CSVs without regenerating 30 million events.</p><h3 id="6-a-shared-machine">6. a shared machine</h3><p>The session was paused partway through to reduce load on a shared machine that was running other jobs. On resume I capped BLAS and OpenMP at 2 threads, cut the synthetic OPE default to 2 worker processes and LightGBM to 2 threads, and ran one experiment at a time. The synthetic OPE run in <code>results/</code> predates the pause and used 6 workers, which changes wall time but not the numbers, since each replicate is seeded by its index.</p><h3 id="7-the-open-bandit-sample-is-tiny-in-click-terms">7. the Open Bandit sample is tiny in click terms</h3><p>The sample shipped with obp has 10,000 rounds per logging policy, with 38 clicks in the random log and 42 in the BTS log. The estimators are correct, and the cross-check shows they compute exactly what obp computes, but the data cannot tell them apart. I say so in the results rather than reading a ranking into noise.</p><h2 id="experiments">experiments</h2><p>All runs were on an Apple M4 Pro shared with other jobs, with BLAS, OpenMP and LightGBM threads capped at 2 (the synthetic OPE run used 6 worker processes, as noted above). Timings are indicative only. None of these experiments was rerun independently.</p><ol><li><strong>A/A through the full DuckDB pipeline.</strong> 1,000 experiments of 2,000 users, 30,568,177 events and 1,800,259 exposed users, five metrics with Welch, CUPED or the delta method, plus both SRM checks, each with a Wilson interval and a Kolmogorov-Smirnov test of p-value uniformity.</li><li><strong>Power.</strong> Eleven effect sizes from 0 to 0.2 standard deviations, 1,000 users per arm, 2,000 simulations each, normal and log-normal outcomes.</li><li><strong>Coverage.</strong> 1,000 simulations of 2,000 users per arm with nonzero true effects, for Welch on zero-inflated log-normal revenue, the delta method on clicks per session and CUPED on a correlated outcome.</li><li><strong>Peeking.</strong> 2,000 experiments per scenario, up to 10,000 unit-variance users per arm, a look every 200 users for 50 looks. The naive analyst stops at the first Welch p below 0.05. The mSPRT uses τ = 0.1 and stops when the always-valid p falls below 0.05. Run under the null and under a true effect of 0.05.</li><li><strong>CUPED.</strong> The variance ratio for ρ from 0 to 0.9, 1,000 simulations of 1,000 users per arm each, plus 50 event-log experiments of 4,000 users.</li><li><strong>Synthetic OPE.</strong> A bandit with 10 actions, 5-dimensional normal contexts and a nonlinear true click probability. The behaviour policy is a softmax at inverse temperature 0.5 and the evaluation policy at 3.0. The true policy value, 0.50492, comes from 2 million fresh contexts with no reward noise. Each of 200 replicates logs 5,000 rounds, fits the cross-fitted LightGBM and a deliberately weak context-free model (the mean reward of each action), runs every estimator, and builds a 500-replicate bootstrap interval.</li><li><strong>Open Bandit.</strong> The all-campaign random-policy log (propensity 1/80 per slot, 3 slots) as the logged data, BTS with the production prior as the evaluation policy, 100,000 Monte Carlo slates, LightGBM with item features, 2,000 bootstrap replicates, and the obp cross-check. The BTS policy's own log gives the on-policy truth to compare against.</li></ol><h2 id="results">results</h2><h3 id="a-a-false-positive-rates-are-where-they-should-be">A/A false positive rates are where they should be</h3><p>Every one of the ten metric and method pairs landed in the 4 to 6 percent band over 1,000 A/A experiments.</p><table><thead><tr><th>Metric</th><th>Method</th><th>False positive rate</th></tr></thead><tbody><tr><td>Revenue per user</td><td>Welch</td><td>5.5%</td></tr><tr><td>Revenue per user</td><td>CUPED</td><td>5.4%</td></tr><tr><td>Conversion</td><td>Welch</td><td>4.2%</td></tr><tr><td>Conversion</td><td>CUPED</td><td>4.3%</td></tr><tr><td>Sessions per user</td><td>Welch</td><td>4.0%</td></tr><tr><td>Sessions per user</td><td>CUPED</td><td>4.0%</td></tr><tr><td>Clicks per session</td><td>delta method</td><td>5.1%</td></tr><tr><td>Revenue per session</td><td>delta method</td><td>5.8%</td></tr><tr><td>SRM on assignments</td><td>chi-square</td><td>4.7%</td></tr><tr><td>SRM on exposed users</td><td>chi-square</td><td>4.8%</td></tr></tbody></table><p>No KS test of p-value uniformity rejected, and the smallest KS p-value was 0.104. Across the family of five primary tests, the chance of at least one false positive in an experiment was 16.8 percent uncorrected, 3.8 percent with Holm and 4.2 percent with Benjamini-Hochberg. That 16.8 is the everyday version of the peeking problem. Five metrics tested at 5 percent each is not a 5 percent test.</p><p>The user-metrics SQL processed the 30.6 million events in 12.8 s, about 2.4 million events per second, and the sufficient-statistics query took 1.8 s. Generating and loading the events took 55 s and 49 s, so the SQL is not the bottleneck. These timings are from the shared, thread-capped machine.</p><h3 id="power-matches-the-noncentral-t">power matches the noncentral t</h3><p>The largest gap between simulated and analytic power was 0.026 on normal data and 0.031 on log-normal data, at effect 0.1 where the simulation gave 0.640 against an analytic 0.608. 19 of the 22 analytic values fell inside the simulated Wilson interval. At zero effect the simulated rejection rates were 5.45 and 4.65 percent.</p><h3 id="intervals-cover-and-cuped-more-than-halves-their-width">intervals cover, and CUPED more than halves their width</h3><p>Nominal 95 percent intervals covered the true effect in 95.2 percent of simulations for Welch on revenue, 95.1 for the delta method and 94.0 for CUPED (Wilson 92.4 to 95.3), over 1,000 simulations each. Plain Welch on the same data as CUPED covered 94.4 percent, so CUPED's slightly low number is shared with the unadjusted test on that dataset rather than introduced by the adjustment. On that data the mean CUPED interval width was 0.124 against 0.277 for Welch.</p><h3 id="peeking">peeking</h3><figure data-figure="chart:projects/cuped-msprt-ope-engine/cuped-msprt-ope-engine-peeking"></figure><p>Under the null, stopping at the first significant naive test over 50 looks gave a <strong>32.8 percent</strong> false positive rate (Wilson 30.8 to 34.9). The curve rises fastest early, 15.0 percent by the fifth look at 1,000 users per arm, because early estimates are noisy and consecutive looks early on share less data. The fixed-horizon test at the final look alone rejected 5.25 percent. The mSPRT rejected <strong>1.95 percent</strong> (Wilson 1.4 to 2.7), well under its 5 percent guarantee, because Ville's bound covers infinitely many looks and this experiment stops at 50.</p><p>The cost shows up under a real effect of 0.05. The mSPRT detected it in 72.75 percent of experiments, stopping on average at 5,112 users per arm among those that stopped. A fixed-horizon test at 10,000 users per arm detected it in 93.75 percent. So the always-valid test buys the right to look whenever you like, and pays about 21 percentage points of power at this horizon and this τ. The naive peeker &quot;detected&quot; it 97.35 percent of the time, but that number means nothing when the same procedure flags a third of null experiments.</p><h3 id="cuped-follows-1-minus-rho-squared">CUPED follows 1 minus rho squared</h3><figure data-figure="chart:projects/cuped-msprt-ope-engine/cuped-msprt-ope-engine-cuped"></figure><p>The simulated variance ratio tracks 1 minus ρ² across the grid, for example 0.831 against 0.84 at ρ = 0.4 and 0.180 against 0.19 at ρ = 0.9. The largest gap, 0.039 at ρ = 0.6, is sampling noise in a ratio of two variances each estimated from 1,000 simulations. The mean ratio of squared standard errors, which uses the analytic formulas rather than the spread of estimates, sits closer still, 0.641 against 0.64 at the same point.</p><p>On the event-log generator, where users have a persistent Gamma-distributed activity rate, sessions per user has a pooled pre-post correlation of 0.723. Theory predicts a variance ratio of 0.4778 and the observed ratio of squared standard errors was 0.4778. Revenue per user has ρ = 0.036, because purchases are rare and revenue is heavy tailed, so CUPED does essentially nothing there (0.998). That is a useful thing to tell a product team. CUPED is not a free halving of sample size, it is exactly as good as the covariate.</p><h3 id="synthetic-off-policy-evaluation">synthetic off-policy evaluation</h3><figure data-figure="chart:projects/cuped-msprt-ope-engine/cuped-msprt-ope-engine-synthetic-ope"></figure><p>Relative bias, relative RMSE and interval coverage over 200 replicates, against a true value of 0.50492.</p><table><thead><tr><th>Estimator</th><th>Reward model</th><th>Bias</th><th>RMSE, percent of truth</th><th>CI coverage</th></tr></thead><tbody><tr><td>IPS</td><td>none</td><td>minus 0.25%</td><td>2.3%</td><td>95.0%</td></tr><tr><td>SNIPS</td><td>none</td><td>minus 0.10%</td><td>1.7%</td><td>97.0%</td></tr><tr><td>DM</td><td>LightGBM</td><td>minus 11.0%</td><td>11.1%</td><td>0%</td></tr><tr><td>DR</td><td>LightGBM</td><td>minus 0.02%</td><td>1.6%</td><td>96.0%</td></tr><tr><td>Switch-DR, τ = 2</td><td>LightGBM</td><td>minus 8.5%</td><td>8.7%</td><td>0%</td></tr><tr><td>DM</td><td>weak</td><td>minus 24.4%</td><td>24.4%</td><td>0%</td></tr><tr><td>DR</td><td>weak</td><td>minus 0.17%</td><td>1.7%</td><td>96.5%</td></tr><tr><td>Switch-DR, τ = 2</td><td>weak</td><td>minus 18.1%</td><td>18.1%</td><td>0%</td></tr></tbody></table><p>DR with LightGBM is the best estimator here, with the lowest RMSE and nominal coverage, and it stays unbiased with the weak model that makes DM 24.4 percent low. That is double robustness working as advertised. The weak model ignores context entirely, so its error is large and strongly negative exactly where π_e concentrates, and DR's correction still removes it, at the cost of a slightly wider spread.</p><p>The coverage column carries the other lesson. DM's bootstrap intervals are tight and wrong. The bootstrap resamples rounds with q̂ held fixed, so it measures the variance of averaging a fixed model over contexts, about 0.008 standard deviation, and nothing about the model's bias, which is about 0.056. An interval like that covered the truth in 0 of 200 replicates. A tight interval around a biased estimate is the failure mode to warn people about, because it looks like confidence.</p><p>In this environment the importance weights are well behaved (in the first replicate the mean weight was 0.976, the 99th percentile 3.53 and the maximum 4.18), which is why plain IPS and SNIPS also do well. DR's advantage over SNIPS here is small. With heavier weights it would matter more, and with a worse model DM's disadvantage would too.</p><h3 id="open-bandit-a-sanity-check-and-not-a-ranking">Open Bandit, a sanity check and not a ranking</h3><figure data-figure="chart:projects/cuped-msprt-ope-engine/cuped-msprt-ope-engine-obd"></figure><p>The BTS policy's own log gives an on-policy click rate of 0.0042, with a bootstrap 95 percent interval of 0.0030 to 0.0055. The random policy's own click rate is 0.0038. From the random logs alone, the estimates of the BTS value were</p><table><thead><tr><th>Estimator</th><th>Estimate</th><th>95% bootstrap CI</th><th>Relative error</th></tr></thead><tbody><tr><td>IPS</td><td>0.00455</td><td>0.00149 to 0.00950</td><td>8.4%</td></tr><tr><td>SNIPS</td><td>0.00478</td><td>0.00157 to 0.00987</td><td>13.7%</td></tr><tr><td>DM</td><td>0.00475</td><td>0.00470 to 0.00479</td><td>13.0%</td></tr><tr><td>DR</td><td>0.00486</td><td>0.00168 to 0.00984</td><td>15.7%</td></tr><tr><td>Switch-DR, τ = 10</td><td>0.00461</td><td>0.00330 to 0.00621</td><td>9.8%</td></tr></tbody></table><p>All five fall inside the on-policy interval. That is the claim, and it is a weak one. IPS, SNIPS and DR have intervals roughly 0.0015 to 0.0099 wide, more than three times the width of the on-policy interval they are being compared with, because the random log has 38 clicks. With that little signal, IPS having the smallest relative error here is luck, not evidence that IPS is the best estimator for this problem. DM's narrow interval is the same illusion as in the synthetic study, a fixed model resampled, and it says nothing about the model's error.</p><p>The switch-DR threshold sweep moves the estimate between 0.00454 and 0.00486 as τ goes from 1 to 20, and at τ = 20 it equals DR exactly, because the largest weight in the log is 19.66. Only 1.04 percent of weights exceed τ = 10.</p><p>The cross-check is the strong claim on this data. With identical inputs, IPS, DM, DR and switch-DR agreed with obp exactly, and SNIPS differed by 1.7e-18, one unit in the last place. My vectorised BTS simulation took 0.42 s against obp's 1.92 s at 100,000 simulations, on the shared machine.</p><h2 id="what-i-would-change">what I would change</h2><h3 id="run-the-full-open-bandit-dataset">run the full Open Bandit Dataset</h3><p>The full release has about 26 million rounds. With thousands of clicks, the Open Bandit comparison would separate the estimators and let me report relative error across campaigns and bootstrap seeds rather than one point per estimator. That is the single most informative next run.</p><h3 id="tune-the-switch-dr-threshold-from-data">tune the switch-DR threshold from data</h3><p>A fixed τ is a guess. obp's tuning variants choose τ by minimising an estimated MSE, and the synthetic results show why that matters, since τ = 2 kept most of DM's bias.</p><h3 id="close-the-dm-gap-with-a-better-reward-model">close the DM gap with a better reward model</h3><p>A per-action linear head on context, or calibrated predictions, may reduce the shrinkage bias. The experiment is to add each and report which one closes the gap, keeping DR as the reference.</p><h3 id="put-the-msprt-on-the-sql-layer">put the mSPRT on the SQL layer</h3><p>Continuous monitoring should read cumulative per-day sums straight from DuckDB rather than from numpy simulations. A group sequential design with alpha spending would be the natural comparison, since it buys back power when the number of looks is known in advance.</p><h3 id="speed-up-generation-and-loading">speed up generation and loading</h3><p>Generating and loading the A/A events took 55 s and 49 s against 13 s for the SQL itself. Writing Parquet once and letting DuckDB read it directly would remove most of that.</p><h2 id="reproducibility">reproducibility</h2><p>The project uses uv and pins Python 3.12.</p><figure class="highlight sh"><table><tr><td class="code"><pre><span class="line"><span class="built_in">cd</span> projects/25-cuped-msprt-ope-engine</span><br><span class="line">uv <span class="built_in">sync</span> --group dev --group crosscheck      <span class="comment"># crosscheck installs obp, used only for cross-checks</span></span><br><span class="line">uv run python scripts/fetch_obd.py          <span class="comment"># copies the OBD sample shipped with obp into data/obd</span></span><br><span class="line">uv run --group crosscheck pytest -q         <span class="comment"># 38 tests, obp tests are skipped without the group</span></span><br></pre></td></tr></table></figure><p>The experiments, with threads capped as they were for the reported runs.</p><figure class="highlight sh"><table><tr><td class="code"><pre><span class="line"><span class="built_in">export</span> OMP_NUM_THREADS=2 OPENBLAS_NUM_THREADS=2 VECLIB_MAXIMUM_THREADS=2</span><br><span class="line"><span class="built_in">cd</span> experiments</span><br><span class="line">uv run python 01_aa_fpr.py --n-experiments 1000 --n-users 2000</span><br><span class="line">uv run python 02_power.py --sims 2000 --n 1000</span><br><span class="line">uv run python 03_coverage.py --sims 1000 --n 2000</span><br><span class="line">uv run python 04_peeking.py --sims 2000 --max-n 10000 --batch 200</span><br><span class="line">uv run python 05_cuped.py --sims 1000 --n 1000</span><br><span class="line">uv run python 06_synthetic_ope.py --reps 200 --n 5000 --workers 2</span><br><span class="line">uv run --group crosscheck python 07_obd_ope.py</span><br></pre></td></tr></table></figure><p><code>01_aa_fpr.py --plot-only</code> redraws its figures from saved results. The full Open Bandit run, when there is disk for it, is</p><figure class="highlight sh"><table><tr><td class="code"><pre><span class="line">curl -LO https://research.zozo.com/data_release/open_bandit_dataset.zip</span><br><span class="line">unzip open_bandit_dataset.zip -d data/obd_full</span><br><span class="line">uv run --group crosscheck python experiments/07_obd_ope.py --data-dir data/obd_full/open_bandit_dataset</span><br></pre></td></tr></table></figure><p>Code is in <code>projects/25-cuped-msprt-ope-engine</code>, with the metrics layer in <code>src/expope/metrics</code>, the A/B statistics in <code>src/expope/stats</code>, the estimators, reward model and Open Bandit loader in <code>src/expope/ope</code>, the simulators in <code>src/expope/sim</code>, the numbered experiment scripts in <code>experiments/</code>, every CSV and JSON behind this post in <code>results/</code>, the invariants and trade-offs in <code>DESIGN.md</code>, and the build log in <code>DEVLOG.md</code>.</p>]]>
    </content>
    <id>https://projects.farhansadeek.com/posts/why-peeking-breaks-ab-tests/</id>
    <link href="https://projects.farhansadeek.com/posts/why-peeking-breaks-ab-tests/"/>
    <published>2026-09-29T06:17:12.000Z</published>
    <summary>An A/B testing and off-policy evaluation engine built from scratch, where checking a test fifty times turns a 5 percent false positive rate into 33.</summary>
    <title>Why Peeking Breaks A/B Tests, and What to Use Instead</title>
    <updated>2026-09-29T06:49:26.552Z</updated>
  </entry>
  <entry>
    <author>
      <name>Farhan Sadeek</name>
      <email>farhan@farhansadeek.com</email>
    </author>
    <category term="research" scheme="https://projects.farhansadeek.com/tags/research/"/>
    <category term="transformers" scheme="https://projects.farhansadeek.com/tags/transformers/"/>
    <category term="interpretability" scheme="https://projects.farhansadeek.com/tags/interpretability/"/>
    <content>
      <![CDATA[<p>The <a href="/posts/reverse-engineering-gpt-2s-induction-circuit/">first post</a> scored every attention head in GPT-2 small and found five induction heads, <strong>L5H5, L6H9, L7H10, L5H1 and L7H2</strong>, with a gap of 0.290 to the sixth. It closed by saying an induction score is a correlation, not a cause, and that four experiments would decide whether the five heads do the copying or only look in the right place. This post reports those experiments, numbered 02 to 06 in the project.</p><p>The short version has four parts. The five heads carry the copy together but not alone. Patching all five recovers about 0.81 of clean performance, while no single head recovers more than 0.128. Their keys are fed by one upstream head, <strong>L4H11</strong> in layer 4, and not by a layer 0 head as the textbook two-layer story and our own outline assumed. Their weights implement the two halves of induction, attend back and copy, but the head with the highest induction score, L5H5, is the weakest direct copier of the five. And removing all five costs only <strong>33%</strong> of the model's in-context learning score, because GPT-2 small keeps many other induction heads in reserve.</p><p>That last result is the one the first post warned about. It does not undo the circuit. It changes what &quot;the induction circuit&quot; can mean in this model, from a pair of heads to a redundant population with a few members that matter more than the rest.</p><p>Everything runs on an M4 Pro laptop through MPS, with the same inputs as experiment 01. Every number below is read from a file in <code>projects/tiny-circuits/results/</code> unless the text says otherwise.</p><p><em>Reading note.</em> The argument runs in the order of the experiments, from what the heads attend to, through what causes the output, to what the weights say and what removal costs. The surprises are in <a href="#the-upstream-head">the upstream head</a>, the L5H5 result in <a href="#the-weights">the weights</a>, and <a href="#necessity">necessity</a>. Skip to <a href="#conclusions">conclusions</a> for the claims and their limits.</p><h2 id="table-of-contents">table of contents</h2><ul><li><a href="#what-the-first-post-left-open">what the first post left open</a></li><li><a href="#inputs-and-metrics">inputs and metrics</a></li><li><a href="#attention-patterns">attention patterns</a></li><li><a href="#activation-patching">activation patching</a></li><li><a href="#the-upstream-head">the upstream head</a></li><li><a href="#the-weights">the weights</a></li><li><a href="#necessity">necessity</a></li><li><a href="#against-the-literature">against the literature</a></li><li><a href="#what-this-does-not-establish">what this does not establish</a></li><li><a href="#instrument-notes">instrument notes</a></li><li><a href="#reproducibility">reproducibility</a></li><li><a href="#conclusions">conclusions</a></li></ul><h2 id="what-the-first-post-left-open">what the first post left open</h2><p>The claimed mechanism has two heads and one edge. A previous-token head writes &quot;the token before me was A&quot; into each position. An induction head in a later layer uses that signature as its key, so a query for token A finds the position right after the last A, and its OV circuit copies the token there into the output. Experiment 01 tested only whether heads exist whose attention lands on the induction target. Experiment 03 tests causation by patching, experiments 03 and 04 test the upstream edge, experiment 05 reads the mechanism from the weights, and experiment 06 tests necessity by ablation. The first post's plan put path patching under experiment 04; in the code it landed in 03, next to the other patching runs.</p><h2 id="inputs-and-metrics">inputs and metrics</h2><p>Every experiment uses the input from experiment 01, a BOS token followed by 50 random tokens and then the same 50 tokens again, with seed 1337. Experiments 02 to 05 use a batch of 8. Experiment 06 uses a batch of 32 whose first 8 rows are the same sequences, because loss differences between ablations are small and a larger batch steadies them.</p><p>Patching needs a corrupted input. We corrupt the <strong>first</strong> half with fresh random tokens, giving <code>[BOS][A'][A]</code>. The second half, and so every prediction target, is identical in both runs, and the only thing removed is the earlier occurrence induction needs to look back at. Corrupting the second half instead would change the labels at the scored positions, and a residual patch there would mostly re-insert the clean input token.</p><p>The patching metric is the mean log probability of the correct next token over second-half positions, normalized so the clean run scores 1 and the corrupted run scores 0. A patch that scores 0.4 recovered 40% of the gap between the two runs. In the clean run the second-half mean log probability is about -0.22 nats, and in the corrupted run it is about -12.36, so the gap is large and the normalization is well conditioned.</p><p>The ablation metric in experiment 06 is an in-context learning score, the mean loss on the second half minus the mean loss on the first half, in the spirit of the ICL score in <a href="https://transformer-circuits.pub/2022/in-context-learning-and-induction-heads/index.html">Olsson et al. (2022)</a>. The first half is random, so its loss sits near chance. The second half is predictable only by copying. For the intact model the first-half loss is 13.17 nats, the second-half loss is 0.24, and the ICL score is -12.94. A score of 0 would mean the model gains nothing from having seen the sequence before.</p><h2 id="attention-patterns">attention patterns</h2><p>Experiment 02 plots the full attention pattern of each of the five heads on one sequence, next to L4H11, the head the literature names as GPT-2 small's main previous-token head. For every head we also split the attention of second-half queries into four buckets, the induction target, the previous token, BOS, and everything else, averaged over the batch.</p><table><thead><tr><th>Head</th><th>Induction target</th><th>Previous token</th><th>BOS</th><th>Other</th></tr></thead><tbody><tr><td>L5H5</td><td>0.925</td><td>0.001</td><td>0.029</td><td>0.045</td></tr><tr><td>L6H9</td><td>0.915</td><td>0.000</td><td>0.046</td><td>0.039</td></tr><tr><td>L7H10</td><td>0.905</td><td>0.000</td><td>0.017</td><td>0.078</td></tr><tr><td>L5H1</td><td>0.902</td><td>0.000</td><td>0.089</td><td>0.009</td></tr><tr><td>L7H2</td><td>0.820</td><td>0.000</td><td>0.037</td><td>0.143</td></tr><tr><td>L4H11</td><td>0.000</td><td>0.988</td><td>0.000</td><td>0.012</td></tr></tbody></table><p><em>Source, results/02_attention_viz.csv.</em></p><p>The induction column repeats the experiment 01 scores exactly, because it is the same quantity. What 02 adds is where the rest of the attention goes. None of the five heads puts more than 0.001 on the previous token, so they are not previous-token heads that happen to score well. Their leftover mass is split between BOS, where GPT-2 heads park attention when there is nothing to attend to, and a diffuse remainder that is largest for L7H2 at 0.143.</p><p>The plotted patterns (<code>figures/02_attention_viz.png</code>) show the geometry the score implies. In the first half, where no earlier copy exists, the five heads attend almost entirely to BOS. In the second half they draw one clean stripe offset 49 positions below the diagonal. L4H11 draws the off-by-one diagonal across the whole sequence and no stripe. This is still attention, and so still correlation. It rules out only that the scores come from some broader pattern overlapping the diagonal.</p><h2 id="activation-patching">activation patching</h2><p>Experiment 03 intervenes on activations. Each run takes one activation from the clean run and puts it into the corrupted run (denoising, which tests whether the component is sufficient to restore the answer), or takes one activation from the corrupted run and puts it into the clean run (noising, which tests whether the component is necessary).</p><h3 id="the-residual-stream-handoff">the residual stream handoff</h3><p>The first intervention patches the residual stream entering a layer, <code>resid_pre</code>, at one half of the sequence at a time. If induction works as described, the information the second half needs starts out at first-half positions, since that is where the earlier copy lives, and is moved to second-half positions by the induction heads.</p><figure data-figure="chart:tiny-circuits/residual-handoff"></figure><p>That is what happens. Patching the clean first half into the corrupted run recovers 0.92 of performance at layers 4 and 5, then 0.35 at layer 6, 0.18 at layer 7 and 0.06 at layer 8. Patching the clean second half recovers nothing up to layer 5 (between -0.017 and 0.000), then 0.15 at layer 6, 0.46 at layer 7 and 0.87 at layer 8. The curves cross between layers 6 and 7. The handoff happens across exactly the layers where the five heads sit, 5 through 7, and it is complete by layer 8.</p><p>At layer 0 the two patches recover exactly 1.0 and 0.0, as they must, since <code>resid_pre</code> there is the token embedding. After layer 8 the second-half curve keeps rising to 0.998 at layer 11, so later layers add to the answer without moving it again. Patching a single position is much weaker, at most about 0.022, because the task is spread over 50 positions.</p><h3 id="heads-one-at-a-time-and-together">heads one at a time and together</h3><p>The second intervention patches one head's output, <code>hook_z</code>, at all positions, in both directions.</p><figure data-figure="chart:tiny-circuits/head-patching"></figure><p>Single heads matter little. In the denoising direction the best head, L7H2, recovers 0.128 and L6H9 recovers 0.120. L5H5, the top head by induction score, recovers only 0.031. In the noising direction the picture is flatter still. Corrupting one induction head costs almost nothing, with one exception, L5H1, whose corruption loses 0.160. The other four each lose less than 0.003.</p><p>Patched together, the five heads are a different object. Denoising all five at once recovers about <strong>0.81</strong> of clean performance, and noising all five loses about <strong>0.77</strong>. The sum of the five single-head denoising effects is about 0.46, well short of 0.81, so the heads are more than additive when restored together. The reading we favor is redundancy. With four clean induction heads still running, corrupting a fifth changes little, because the others still deliver the copy. Only when all of them are corrupted at once does the redundancy run out.</p><p>The five are also not the only heads that recover performance. L9H6 recovers 0.104 and L9H9 recovers 0.103, level with the canonical heads. These rank seventh and eighth by induction score in experiment 01 (0.509 and 0.500). L10H0 recovers 0.082 and L10H1 recovers 0.067. The patching result is already hinting at the necessity result below. Induction in GPT-2 small is done by more heads than the five with the sharpest attention.</p><h3 id="heads-that-push-against-the-answer">heads that push against the answer</h3><p>Two heads have clearly negative denoising effects. Restoring the clean output of <strong>L10H7</strong> lowers the metric by 0.092, and restoring <strong>L11H10</strong> lowers it by 0.055. Giving these heads their clean input makes the model worse at predicting the repeated token.</p><p>These are the heads <a href="https://arxiv.org/abs/2211.00593">Wang et al. (2022)</a> called negative name movers in the indirect-object circuit, and that <a href="https://arxiv.org/abs/2310.04625">McDougall et al. (2023)</a> reinterpreted as copy suppression heads. A copy suppression head attends to tokens that earlier heads are predicting from context and pushes their logits down, which calibrates overconfident copying. On a task where copying is always right, that calibration shows up as a cost. Experiment 05 confirms the mechanism from the weights.</p><h2 id="the-upstream-head">the upstream head</h2><p>The outline for this project put the previous-token head in layer 0, following the two-layer attention-only model of <a href="https://transformer-circuits.pub/2021/framework/index.html">Elhage et al. (2021)</a>. In GPT-2 small it is not there.</p><h3 id="path-patching-into-the-keys">path patching into the keys</h3><p>Path patching isolates one edge. For each of the 60 heads in layers 0 to 4, we take the difference between its corrupted and clean outputs and add it only to the key input of the five induction heads, leaving every other path, including the head's effects on MLPs and later heads, clean. We repeat this for the query and value inputs as a control. If a previous-token head feeds induction through K-composition, its effect should appear through keys and nowhere else.</p><figure data-figure="chart:tiny-circuits/path-patching"></figure><p>One sender dominates. Corrupting <strong>L4H11</strong>'s direct path into the induction keys loses <strong>0.242</strong> of clean performance. The next sender, L3H7, loses 0.013, and L2H2 loses 0.007. L4H11's query and value paths lose about 0.0001 each, and no sender's query or value path reaches 0.002. The edge is real, it runs through keys as K-composition predicts, and one head carries almost all of it.</p><h3 id="why-the-total-effect-is-smaller">why the total effect is smaller</h3><p>Here is a surprise. L4H11's total effect is smaller than its direct path. Noising its output everywhere, rather than only on the path into the induction keys, loses 0.093 (from the head patching table), less than half of the 0.242 its key path alone loses.</p><p>The difference is MLP4, the MLP in the same layer. When L4H11's output is corrupted everywhere, MLP4 sees the corrupted output too and its response partly offsets the damage. Holding MLP4 at its clean output while noising L4H11 raises the loss to about 0.26, back to the size of the direct path. Zero-ablating L4H11 costs only about 0.02.</p><p>We do not know what MLP4 computes that lets it compensate. The simplest reading is that previous-token information is available to MLP4 from other sources, so it can partly rebuild the signal L4H11 would have written. Downstream components compensating for a damaged one is documented in larger models by <a href="https://arxiv.org/abs/2307.15771">McGrath et al. (2023)</a>, who call it the hydra effect, and backup heads in GPT-2 small by Wang et al. We have not seen it reported for this MLP and this head. The practical point is about method. Had we measured L4H11 only by its total effect or by zero-ablation, we would have ranked it as a minor head. Path patching was needed to see that it is the main source of the keys.</p><h3 id="previous-token-scores">previous-token scores</h3><p>Experiment 04 scores every head on attention from each position to the one before it, on the repeated random sequences and, as a check that the score does not depend on random tokens, on a short passage of English.</p><table><thead><tr><th>Head</th><th>On random tokens</th><th>On English</th></tr></thead><tbody><tr><td>L4H11</td><td>0.987</td><td>0.999</td></tr><tr><td>L3H7</td><td>0.544</td><td>0.358</td></tr><tr><td>L2H2</td><td>0.494</td><td>0.532</td></tr><tr><td>L6H8</td><td>0.480</td><td>0.188</td></tr><tr><td>L5H6</td><td>0.415</td><td>0.140</td></tr><tr><td>L3H2</td><td>0.402</td><td>0.388</td></tr></tbody></table><p><em>Source, results/04_prev_token_scores.csv, top six of 144 heads by the random-token score.</em></p><p>L4H11 is a near-perfect previous-token head on both kinds of input. Everything else is partial. The strongest layer 0 head is L0H7, at 0.185 on random tokens and 0.264 on English, so no layer 0 head plays the role the two-layer story assigns. L2H2 is the head Wang et al. name alongside L4H11 as a previous-token head, and it shows up here as the second strongest on English.</p><h3 id="composition-scores">composition scores</h3><p>Composition scores ask the same question from the weights alone, with no input at all. Following Elhage et al., the K-composition score between an upstream head A and an induction head B is the normalized Frobenius norm of <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msubsup><mi>W</mi><mrow><mi>O</mi><mi>V</mi></mrow><mi>A</mi></msubsup><mo stretchy="false">(</mo><msubsup><mi>W</mi><mrow><mi>Q</mi><mi>K</mi></mrow><mi>B</mi></msubsup><msup><mo stretchy="false">)</mo><mi mathvariant="normal">⊤</mi></msup></mrow><annotation encoding="application/x-tex">W_{OV}^{A} (W_{QK}^{B})^{\top}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1.2605em;vertical-align:-0.4114em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.1389em;">W</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.8413em;"><span style="top:-2.4247em;margin-left:-0.1389em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight" style="margin-right:0.0278em;">O</span><span class="mord mathnormal mtight" style="margin-right:0.2222em;">V</span></span></span></span><span style="top:-3.063em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight">A</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.2753em;"><span></span></span></span></span></span></span><span class="mopen">(</span><span class="mord"><span class="mord mathnormal" style="margin-right:0.1389em;">W</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.8413em;"><span style="top:-2.4247em;margin-left:-0.1389em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight">Q</span><span class="mord mathnormal mtight" style="margin-right:0.0715em;">K</span></span></span></span><span style="top:-3.063em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight" style="margin-right:0.0502em;">B</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.4114em;"><span></span></span></span></span></span></span><span class="mclose"><span class="mclose">)</span><span class="msupsub"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.8491em;"><span style="top:-3.063em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mtight">⊤</span></span></span></span></span></span></span></span></span></span></span></span>, which measures how much of what A writes lands in the subspace B's keys read. Q- and V-composition are the analogous products with B's query side and B's OV circuit. The baseline is the same score for random matrices of the same shapes, 0.036.</p><p>Averaged over the five induction heads, L4H11 has the highest K-composition of any head in layers 0 to 4, at 0.096, against 0.052 for its Q-composition and 0.038 for its V-composition. Its K-composition is between 0.087 and 0.104 with each individual induction head. Across the 60 heads in layers 0 to 4, previous-token score and mean K-composition have a Spearman correlation of 0.78.</p><p>The scores are noisy in the way <a href="https://github.com/callummcdougall/ARENA_3.0">ARENA's tutorial</a> and Elhage et al. warn about. L4H7 has a K-composition of 0.091, almost level with L4H11, yet its previous-token score is 0.177 and its path-patching effect through the induction keys is -0.002. The weights say L4H7 could compose with the induction heads, and the activations say it does not in any way that matters here. That is why we treat path patching as the primary evidence for the edge and composition scores as support.</p><h2 id="the-weights">the weights</h2><p>Experiment 05 asks whether the weights implement induction independent of any input, which is the kind of evidence the first post said a circuit claim needs alongside patching.</p><h3 id="ov-circuits-copy">OV circuits copy</h3><p>A head's full OV circuit, embedding through value, output and unembedding, is a map from &quot;the token this head attends to&quot; to &quot;which output logits go up.&quot; A copying head raises the logit of the token it attends to. We form this map on a fixed random subset of 2,000 vocabulary tokens and count how often a token's own logit is the largest in its row (top-1) or among the five largest (top-5). Chance is 1 in 2,000 for top-1. We also report the positive-eigenvalue share of Elhage et al., the sum of the eigenvalues divided by the sum of their magnitudes, which is 1 for a pure copying map and -1 for a pure anti-copying one.</p><p>The embedding matters. GPT-2 small is known to use its first MLP as an extension of the token embedding, so we compute the circuit twice, once with the raw embedding <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>W</mi><mi>E</mi></msub></mrow><annotation encoding="application/x-tex">W_E</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.8333em;vertical-align:-0.15em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.1389em;">W</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3283em;"><span style="top:-2.55em;margin-left:-0.1389em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight" style="margin-right:0.0576em;">E</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span></span></span></span> and once with the effective embedding <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>W</mi><mi>E</mi></msub><mo>+</mo><msub><mtext>MLP</mtext><mn>0</mn></msub><mo stretchy="false">(</mo><msub><mi>W</mi><mi>E</mi></msub><mo stretchy="false">)</mo></mrow><annotation encoding="application/x-tex">W_E + \text{MLP}_0(W_E)</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.8333em;vertical-align:-0.15em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.1389em;">W</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3283em;"><span style="top:-2.55em;margin-left:-0.1389em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight" style="margin-right:0.0576em;">E</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">+</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mord"><span class="mord text"><span class="mord">MLP</span></span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3011em;"><span style="top:-2.55em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight">0</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span><span class="mopen">(</span><span class="mord"><span class="mord mathnormal" style="margin-right:0.1389em;">W</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3283em;"><span style="top:-2.55em;margin-left:-0.1389em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight" style="margin-right:0.0576em;">E</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span><span class="mclose">)</span></span></span></span>, as McDougall et al. do.</p><table><thead><tr><th>Head</th><th>Top-1, effective</th><th>Top-5, effective</th><th>Eigenvalue share</th><th>Top-1, raw</th></tr></thead><tbody><tr><td>L7H2</td><td>0.924</td><td>0.990</td><td>0.996</td><td>0.008</td></tr><tr><td>L7H10</td><td>0.892</td><td>0.965</td><td>0.995</td><td>0.021</td></tr><tr><td>L6H9</td><td>0.872</td><td>0.970</td><td>0.997</td><td>0.047</td></tr><tr><td>L5H1</td><td>0.558</td><td>0.791</td><td>0.987</td><td>0.004</td></tr><tr><td>L5H5</td><td>0.155</td><td>0.353</td><td>0.953</td><td>0.058</td></tr></tbody></table><p><em>Source, results/05_summary.json, 2,000 tokens.</em></p><p>With the effective embedding, L7H2, L7H10 and L6H9 map a token to itself for 87% to 92% of tokens, and to its top five for 96% to 99%. Their eigenvalue shares are 0.995 to 0.997. A Gaussian random head matched in norm scores a top-1 rate of 0.0006 and an eigenvalue share of -0.005. With the raw embedding, top-1 rates fall to between 0.004 and 0.058, which is the clearest single sign in this project that MLP0 acts as part of the embedding.</p><p>The eigenvalue test does not single out induction heads, and it should not be read as if it did. Among the other 139 heads the 90th percentile of the eigenvalue share is 0.988, and 33 of all 144 heads exceed 0.95. Copying OV circuits are common in the later layers of GPT-2 small. What distinguishes an induction head is copying combined with the right attention, not copying alone.</p><h3 id="qk-circuits-attend-back">QK circuits attend back</h3><p>The QK side is tested with L4H11 in the loop. For each query token t we score every key position whose previous token is one of the 2,000 subset tokens, with the key built from L4H11's output on that previous token, and ask whether the highest-scoring key is the one whose previous token is t. For all five heads it is, for 98.9% to 99.8% of tokens (L5H5 0.994, L6H9 0.998, L5H1 0.998, L7H10 0.989, L7H2 0.989). As a control, the same computation with plain token keys, no previous-token head, picks the matching key for exactly 0.0% of tokens in every head. The attend-back rule is in the weights, and it only exists once the previous-token head writes into the keys.</p><h3 id="direct-logit-attribution">direct logit attribution</h3><p>Direct logit attribution asks how much each head's output, passed straight through the final LayerNorm scale and the unembedding, raises the logit of the correct next token.</p><figure data-figure="chart:tiny-circuits/direct-logit"></figure><p>On second-half positions the five heads add L7H2 +1.51, L7H10 +1.16, L6H9 +1.08, L5H1 +0.69 and L5H5 +0.30. On first-half positions, where there is nothing to copy, all five contribute between -0.001 and +0.003. Together the five write 4.74 of the 13.37 that all 144 heads write directly, against a correct-token logit of 21.49 and a mean logit of about 0. L9H9 (+1.42), L9H6 (+1.41) and L10H0 (+1.17) each write as much as a canonical head.</p><p>The copy suppression heads show up from the other side. L10H7 writes -1.21 and L11H10 writes -0.77, the two most negative heads in the model, and their OV eigenvalue shares are -0.999 and -0.996, which is almost perfect anti-copying. Their weights do exactly what the patching result implied.</p><h3 id="the-l5h5-surprise">the L5H5 surprise</h3><p>The next surprise is L5H5. It has the highest induction score in the model (0.925) and attend-back weights as sharp as the others (0.994), and yet it is the weakest copier of the five (top-1 0.155), the smallest direct writer (+0.30), and the smallest single-head patching effect (0.031 recovered). Its eigenvalue share of 0.953 says its OV circuit is broadly positive, just not sharply diagonal.</p><p>One reading is that L5H5 matters mostly indirectly, by writing something later heads read, rather than by writing the answer itself. An early induction head whose output is consumed by later induction or copying heads would look exactly like this. We have not tested that reading. Path patching from L5H5 into later heads would, and it is the obvious next experiment. The broader lesson is that rank by induction score, which measures attention placement, does not predict rank by causal effect, copying or direct write.</p><h2 id="necessity">necessity</h2><p>Experiment 06 removes heads and measures what is lost. Zero ablation sets a head's output to 0. Mean ablation replaces it with its mean over a separate batch of 64 repeated random sequences, position by position, which removes the token-specific information while keeping the average activation the rest of the model expects. Mean ablation is the cleaner test, and we report it first.</p><h3 id="the-five-heads-together">the five heads together</h3><p>Mean-ablating all five induction heads moves the ICL score from -12.94 to <strong>-8.68</strong>, and the second-half loss from 0.24 to 4.46 nats. That removes 33% of the ICL score. Zero ablation gives a similar -8.94. Adding L4H11 to the five takes the score to -6.44 under mean ablation, removing half of it.</p><p>Twenty random sets of five heads, drawn from the other 139, are the control. Under mean ablation they average -12.86 with a standard deviation of 0.055, and the worst set reaches -12.76. The five induction heads matter far more than five random heads. They are not, on their own, the in-context learning of this model.</p><h3 id="single-heads">single heads</h3><p>No induction head matters much alone under mean ablation. Each moves the ICL score by at most 0.047 nats (L6H9). L4H11 alone moves it by 0.294, the largest single-head effect in the model, which fits its role as the one head all five induction heads read their keys from.</p><p>Zero ablation tells a partly different story, and it is the one number here that disagrees with the patching picture in an informative way. Zero-ablating L5H1 costs 0.63 nats of ICL score, more than any other single head, while mean-ablating it costs 0.011. L5H1 also stood out in patching as the only induction head whose corruption alone loses real performance (0.160). One explanation is that L5H1's output carries a large constant component the rest of the model relies on, which the mean keeps and zero removes. We have not tested it.</p><h3 id="backup-induction-heads">backup induction heads</h3><p>If the five are a subset of a larger population, ablating further heads in order of induction score should keep eroding the ICL score, and ablating random heads should not.</p><figure data-figure="chart:tiny-circuits/backup-heads"></figure><p>That is what the curve shows. With the top k heads mean-ablated, the ICL score is -12.03 at k = 4, -8.68 at k = 5, -5.34 at k = 10 and -1.13 at k = 20. By k = 20 the second-half loss is 11.80 nats against a first-half loss of 12.93, so in-context copying is nearly gone. Random sets of k heads, averaged over ten draws per k, stay near -12.7 all the way to k = 20.</p><p>The steps are informative. Most of the drop at k = 5 comes when L7H2 is added, after the first four heads together cost only 0.9 nats, so the four canonical heads above it are well covered by the rest. The curve jumps again at k = 8 when L9H9 joins L9H6, and at k = 13 when L8H1 is added. It bumps slightly the wrong way at k = 9 and k = 11, where the added heads are L10H7 and L11H10, the copy suppression heads. Removing a head that works against copying helps copying, which is the ablation result agreeing with patching and with the weights.</p><h3 id="the-knockout-sweep">the knockout sweep</h3><p>The last run mean-ablates every head alone. The largest effect by far is L4H11, at +0.294 nats. The next largest are L6H10 and L8H6, at +0.077 each, and 16 heads in total exceed +0.05. The heads whose removal most improves the ICL score are L6H6 (-0.086), L8H10 (-0.070) and L10H7 (-0.063), and all three have strongly negative OV eigenvalue shares in experiment 05 (-0.969, -0.992 and -0.999). The anti-copying heads are the heads whose loss helps.</p><p>No induction head appears near the top of the single-head sweep. That is the backup picture in its plainest form. The model has one bottleneck upstream, L4H11, and a redundant population downstream.</p><h2 id="against-the-literature">against the literature</h2><p>The table sorts each result by how directly it can be compared with published work.</p><table><thead><tr><th>Result</th><th>Reference</th><th>Comparison</th></tr></thead><tbody><tr><td>L5H5, L6H9, L5H1, L7H10 and L7H2 are the induction heads</td><td>ARENA tutorial names all five, Wang et al. 2022 name L5H5 and L6H9</td><td>Match on the head identities</td></tr><tr><td>L4H11 is the main previous-token head, L2H2 a weaker one</td><td>Wang et al. 2022, ARENA</td><td>Match on the head identities</td></tr><tr><td>The previous-token edge runs through keys, not queries or values</td><td>Elhage et al. 2021</td><td>Qualitative match, their model is a two-layer attention-only transformer</td></tr><tr><td>The previous-token head sits in layer 0</td><td>Elhage et al. 2021, our outline</td><td>Does not carry over to GPT-2 small</td></tr><tr><td>L10H7 and L11H10 push against the copied token</td><td>Wang et al. 2022, McDougall et al. 2023</td><td>Match on identities and sign</td></tr><tr><td>OV circuits of induction heads copy, with positive eigenvalues</td><td>Elhage et al. 2021, Olsson et al. 2022</td><td>Qualitative match</td></tr><tr><td>MLP0 acts as part of the embedding</td><td>McDougall et al. 2023</td><td>Match in direction</td></tr><tr><td>Ablating induction heads removes in-context learning, far more than other heads</td><td>Olsson et al. 2022</td><td>Direction matches, magnitude only qualitative</td></tr><tr><td>Backup heads and self-repair</td><td>Wang et al. 2022 in GPT-2 small, McGrath et al. 2023 in larger models</td><td>Qualitative match</td></tr></tbody></table><p>What we can match is head identities and signs. Magnitudes mostly cannot be compared, because the published work used other models (Olsson et al. and Elhage et al. study their own small models), other tasks (Wang et al. study indirect-object identification), or other metrics. We know of no published patching map, K-path size or top-k ablation curve for this setup. Every qualitative prediction we could check held, and the magnitudes here are new measurements rather than replications.</p><h2 id="what-this-does-not-establish">what this does not establish</h2><p>Four limits apply to everything above.</p><p>First, the input is still repeated random tokens. That was the right choice for isolating induction, as the first post argued, and it means nothing here shows what these heads do on English. The one exception is the previous-token score of L4H11, which is 0.999 on a paragraph of English as well.</p><p>Second, the patching and ablation runs use one seed and small batches, 8 for patching and 32 for ablation. The random-head control gives a sense of the noise for ablation (a standard deviation of 0.055 nats under mean ablation), and there is no equivalent for patching. The ordering of the large effects is safe. Small differences, such as L9H6 against L9H9, are not.</p><p>Third, the mechanism is described for one edge and one layer of heads. We have not traced what reads from L5H5, what lets MLP4 compensate for L4H11, or why zero-ablating L5H1 costs so much more than mean-ablating it. Each is a named gap, not a detail.</p><p>Fourth, &quot;the induction circuit&quot; now needs a looser definition than the one the first post started with. A strict circuit, two heads and one edge that are together sufficient and necessary, does not describe GPT-2 small. A redundant population of induction and copying heads, fed by one previous-token head and opposed by a few copy suppression heads, does. That description fits every result here, and it is harder to test to the same standard.</p><h2 id="instrument-notes">instrument notes</h2><details class="collapsible-section"><summary><strong>Numbers printed by the scripts but not written to results/</strong></summary><p>Most numbers above are read from CSV or JSON files in <code>results/</code>. A few are only printed to the console by the scripts, and so appear in the project README rather than in a results file. They are the clean and corrupted log probabilities (-0.22 and -12.36), the joint patching of all five induction heads (0.81 recovered, 0.77 lost), L4H11's effect with MLP4 held clean (0.26) and under zero ablation (0.02), and the random-k ablation curve (-12.92 at k = 1, -12.67 at k = 10, -12.66 at k = 20). To check them we reran experiments 03, 04 and 06 in a separate copy of the project. The reruns reproduced every committed results file byte for byte and printed these same values. Writing them to the results files is a small change to experiments 03 and 06.</p></details><details class="collapsible-section"><summary><strong>Why mean ablation, and why it differs from zero ablation</strong></summary><p>Zero ablation removes a head's output entirely, including any constant component that downstream layers treat as a bias. That can push the residual stream off the distribution the model was trained on, and the resulting damage says as much about the distribution shift as about the head's function. Mean ablation keeps the average and removes only the input-dependent part, which is the part a copying head is supposed to contribute.</p><p>The sweep shows why this matters. Under zero ablation, several layer 0 heads change the ICL score by large amounts in the <em>helpful</em> direction, L0H8 by -2.87 nats. That is not a head that works against copying. The ICL score is a difference of two losses, and zeroing an early head raises the first-half loss (the random set that contains L0H6 reaches a first-half loss of 14.80 under zero ablation, against 13.17 intact, while its second-half loss stays at 0.23) more than the second-half loss. We read the zero-ablation sweep for layer 0 as an artifact of the metric and rely on mean ablation throughout.</p></details><details class="collapsible-section"><summary><strong>The MPS warning</strong></summary><p>TransformerLens warns that the MPS backend may produce silently incorrect results with the installed PyTorch version. The consistency checks inside these runs argue against a problem here. Layer 0 residual patching returns exactly 1.0 and 0.0, the attention buckets in experiment 02 reproduce the experiment 01 scores to every printed digit, and the top-1 and top-5 mean-ablation rows reproduce the single-head and five-head rows exactly. The reruns above were also on MPS, so they show determinism, not correctness. A CPU rerun would settle it and has not been done.</p></details><h2 id="reproducibility">reproducibility</h2><p>Each experiment is a numbered script that writes its figures to <code>figures/</code> and its tables to <code>results/</code>.</p><figure class="highlight bash"><table><tr><td class="code"><pre><span class="line">uv run python experiments/02_attention_viz.py</span><br><span class="line">uv run python experiments/03_activation_patching.py</span><br><span class="line">uv run python experiments/04_prev_token_heads.py</span><br><span class="line">uv run python experiments/05_qk_ov_analysis.py</span><br><span class="line">uv run python experiments/06_ablation.py</span><br></pre></td></tr></table></figure><p>All five run on the same laptop as experiment 01. Experiment 03 is the slowest, since it runs more than 1,500 patched forward passes for the residual map, the head patching and the path patching. The charts in this post are drawn from JSON files copied from <code>results/</code>, and each names its source file.</p><h2 id="conclusions">conclusions</h2><p>The first post found five heads that look in the right place. This one asked whether they do the work, and the answer has more parts than the two-head story allows.</p><ul><li>The five induction heads are causally responsible for in-context copying as a group. Patching all five recovers about 0.81 of clean performance and noising all five loses about 0.77, while the best single head recovers 0.128.</li><li>The information moves from first-half to second-half positions between layers 5 and 8, across exactly the layers the five heads occupy.</li><li>One upstream head, L4H11 in layer 4, feeds the induction keys. Its direct key path carries 0.242 of performance, the next sender carries 0.013, and its query and value paths carry nothing. No layer 0 head plays this role.</li><li>MLP4 partly compensates when L4H11 is corrupted, so L4H11's total effect (0.093) understates its direct role. Path patching was needed to see it.</li><li>The weights implement both halves of induction. The QK circuit with L4H11 keys picks the right key for about 99% of tokens and plain token keys for none. The OV circuits of three heads copy for 87% to 92% of tokens, once MLP0 is counted as part of the embedding.</li><li>Rank by induction score does not predict rank by effect. L5H5 has the top score and is the weakest copier, direct writer and single-head patch of the five.</li><li>Removing the five heads costs 33% of the ICL score, against about 1% for random heads. Removing the top 20 by induction score costs over 90%. GPT-2 small has many backup induction heads, and two copy suppression heads, L10H7 and L11H10, work against all of them.</li></ul>]]>
    </content>
    <id>https://projects.farhansadeek.com/posts/testing-gpt-2s-induction-heads/</id>
    <link href="https://projects.farhansadeek.com/posts/testing-gpt-2s-induction-heads/"/>
    <published>2026-09-29T06:17:11.000Z</published>
    <summary>Activation patching, path patching, weight analysis and ablation on GPT-2 small's five induction heads. They carry the copy together, one layer 4 head feeds their keys, and removing them costs only a third of in-context learning.</summary>
    <title>Testing GPT-2's Induction Heads and the Backups Behind Them</title>
    <updated>2026-09-29T06:49:26.552Z</updated>
  </entry>
  <entry>
    <author>
      <name>Farhan Sadeek</name>
      <email>farhan@farhansadeek.com</email>
    </author>
    <category term="storage" scheme="https://projects.farhansadeek.com/tags/storage/"/>
    <category term="cpp" scheme="https://projects.farhansadeek.com/tags/cpp/"/>
    <category term="distributed-systems" scheme="https://projects.farhansadeek.com/tags/distributed-systems/"/>
    <content>
      <![CDATA[<p>I wrote minikafka, a single-broker event log in C++20 that follows Kafka's storage design closely. Topics are split into partitions, and each partition is an append-only log of segment files with a sparse offset index. Producers send batches over a small binary TCP protocol, consumers fetch from an offset with long polling, and consumer groups get partitions assigned by the broker and commit their positions into an internal log, so a group resumes where it left off after a restart. Recovery after a crash truncates a torn tail, and size-based retention deletes old segments.</p><p>The code is about 2,700 lines of C++ across the library, the broker, a CLI and a benchmark harness, plus about 830 lines of GoogleTest. There are <strong>31 tests</strong>, and on 2026-09-26 an independent clean release rebuild passed all 31. The ASan plus UBSan and ThreadSanitizer builds also passed in my own runs, but those, and the benchmarks, were not rerun independently.</p><p>The headline measurement is how much batching matters. With 100 B messages, going from one record per produce request to 1,000 moved throughput from <strong>32,886 to 12,744,103 messages per second</strong>, a factor of about 390. End to end, without an fsync, a message took <strong>41.5 µs</strong> at the median from send to receive at 10,000 messages per second. Pushing data all the way through the drive's write cache with <code>F_FULLFSYNC</code> raised that to <strong>4.0 ms</strong>. All benchmark numbers were measured on a shared Apple M4 Pro that was also running other people's builds and benchmarks, with the 1 minute load average recorded in every row (8 to 11 for the reported runs). They are indicative, not a clean benchmark.</p><p>Code is in <code>projects/04-minikafka-log-broker</code>. v0 is one broker, with no replication.</p><p>The argument of this post is that Kafka's speed comes from a few simple decisions about bytes, and that the hard part of reproducing it was not the protocol but keeping the index, the locks and the benchmarks honest. Skip to <a href="#problems">problems</a> for the bugs.</p><h2 id="table-of-contents">table of contents</h2><ul><li><a href="#what-i-wanted-to-build">what I wanted to build</a></li><li><a href="#theory">theory</a></li><li><a href="#architecture">architecture</a></li><li><a href="#implementation">implementation</a></li><li><a href="#problems">problems</a></li><li><a href="#experiments">experiments</a></li><li><a href="#results">results</a></li><li><a href="#what-i-would-change">what I would change</a></li><li><a href="#reproducibility">reproducibility</a></li></ul><h2 id="what-i-wanted-to-build">what I wanted to build</h2><p>I wanted to understand why Kafka is fast and what it actually promises, by building the core of it myself. The v0 scope was chosen to be small enough to finish and verify but real enough to measure.</p><ul><li>Topics with partitions, each an append-only log of segment files on disk.</li><li>A sparse offset index per segment, so that fetching from offset N does not scan the log.</li><li>Crash recovery that finds and removes a torn write at the end of the log.</li><li>A binary TCP protocol with batched produce and long-poll fetch.</li><li>Consumer groups with broker-side partition assignment, and committed offsets stored in an internal log so they survive a restart.</li><li>Size-based retention.</li><li>Producer and consumer libraries, a CLI, tests for ordering, durability and resume, and throughput and latency benchmarks.</li></ul><p>Replication, a controller, zero-copy <code>sendfile</code>, compaction and exactly-once delivery are out of scope for v0. They are on the roadmap, and a few of the measurements below show exactly where their absence costs something.</p><h2 id="theory">theory</h2><h3 id="a-partition-is-a-log">a partition is a log</h3><p>A Kafka partition is an ordered sequence of records where each record gets the next integer offset. Almost everything else follows from that one decision.</p><p>Appends only touch the end of a file, so writes are sequential, and the operating system's page cache absorbs them. A consumer that is caught up reads data that was written moments ago, which is still in the same cache, so a well-behaved consumer rarely touches the disk at all.</p><p>A consumer's entire state is one number per partition, the next offset to read. That makes committing progress cheap and moves the question of which messages have been seen out of the broker.</p><h3 id="ordering-and-keys">ordering and keys</h3><p>Ordering is promised only within a partition. A producer that wants all events for one customer to stay in order gives them the same key, and the key is hashed to pick the partition. Partitions are also the unit of parallelism, since each partition in a group is read by exactly one consumer at a time.</p><h3 id="finding-offset-n">finding offset N</h3><p>A consumer asks for &quot;offset 1,048,576 onwards&quot;. Records have variable length, so offset N is not at a computable byte position. The broker needs an index.</p><p>Kafka's index is sparse. Every few KiB of log, it writes an entry mapping a relative offset to a file position. A lookup is a binary search for the last entry at or below the target, then a short forward scan through record headers. Because offsets are dense, the floor entry is never more than one interval of bytes behind the target, so the scan is bounded.</p><figure data-figure="diagram:kafka-sparse-index"></figure><p>A dense index would cost 8 bytes per record for no real gain, since the fetch reads a chunk of the file anyway. At a 4 KiB interval, a 64 MiB segment needs at most 16,384 entries, or 128 KiB, small enough to keep in memory.</p><h3 id="batching">batching</h3><p>Every produce request pays fixed costs, a system call on each side, a network round trip, a lock acquisition and, if configured, an fsync. Batching amortises them over many records, and most of the experiments below ask how much that buys.</p><h3 id="durability-is-a-policy">durability is a policy</h3><p>Acknowledging a write after <code>write(2)</code> means the data is in the kernel. It survives the broker process crashing or being killed, but not a kernel panic or power loss. <code>fsync(2)</code> asks the kernel to push the data to the drive. On macOS, <code>fsync</code> hands data to the drive but does not flush the drive's own write cache, and <code>fcntl(F_FULLFSYNC)</code> does. Kafka itself acknowledges after the write by default and relies on replication for durability. A single broker has no replicas, so it has to choose, and v0 exposes the choice as a flag and measures it.</p><h2 id="architecture">architecture</h2><p>Clients hold one TCP connection each, with one request in flight. The broker accepts connections on one thread and serves each connection on its own thread. A request frame is decoded and dispatched to one of ten APIs, which are CreateTopic, Metadata, Produce, Fetch, ListOffsets, OffsetCommit, OffsetFetch, JoinGroup, Heartbeat and LeaveGroup.</p><figure data-figure="diagram:kafka-broker"></figure><h3 id="one-byte-format-for-the-wire-and-the-disk">one byte format for the wire and the disk</h3><p>The most important decision is that a record looks the same on the wire and on disk.</p><figure class="highlight plaintext"><table><tr><td class="code"><pre><span class="line">record = header (16 B) + payload</span><br><span class="line">  u64 offset        0 from producers, assigned by the broker</span><br><span class="line">  u32 payload_len   bytes after the crc field</span><br><span class="line">  u32 crc           CRC-32C of the payload</span><br><span class="line">  payload:</span><br><span class="line">  i64 timestamp_ms</span><br><span class="line">  u32 key_len</span><br><span class="line">  key bytes</span><br><span class="line">  value bytes       payload_len - 12 - key_len</span><br></pre></td></tr></table></figure><p>A record costs 28 bytes plus its key and value. Because the formats match, the broker never re-encodes anything. On produce it validates the batch, writes the assigned offsets into the headers in place and issues a single <code>pwrite</code>. On fetch it copies a byte range out of the file into the response. The consumer does the parsing and CRC checking. This is Kafka's main trick, and it is also what would make zero-copy <code>sendfile</code> a small change later.</p><p>Integers are little-endian, unlike Kafka's big-endian. Every machine this runs on is little-endian, so encoding is a <code>memcpy</code>, and the code asserts the host byte order at compile time.</p><h3 id="the-produce-path">the produce path</h3><ol><li>The producer appends encoded records to a per-partition buffer and sends it when it reaches a record count or byte limit, or on <code>flush()</code>.</li><li>The broker reads the whole frame, and <code>validate_batch</code> walks it, bounds-checking every length and verifying every CRC. Nothing touches the log until the whole batch is known good.</li><li><code>PartitionLog::append</code> takes the partition mutex, rolls to a new segment if needed, patches offsets into the headers, and writes the batch. With a flush policy other than none, the segment is flushed before the lock is released.</li><li>If any fetch is waiting, the broker bumps a long-poll epoch and wakes it. It replies with the base offset.</li></ol><h3 id="the-fetch-path">the fetch path</h3><ol><li><code>PartitionLog::read</code> takes the mutex only long enough to check the offset range, pick the segment by binary search over base offsets, look up the floor index entry, and snapshot the segment's current size.</li><li>Outside the lock, it <code>pread</code>s from the index position, skips records below the target, and cuts at the last whole record that fits in <code>max_bytes</code>. The first record is always returned whole, so an oversized record cannot wedge a consumer.</li><li>If no partition had data and the request allows waiting, the fetch sleeps on a condition variable until a produce bumps the epoch or the deadline passes, and then reads again.</li></ol><h3 id="invariants-the-code-relies-on">invariants the code relies on</h3><p>DESIGN.md lists eight. The two that carry the most weight are that offsets in a partition are dense and assigned under the partition mutex, and that bytes of a segment below its current size never change, which is what lets fetches read without the lock. A third, that every record on disk passed <code>validate_batch</code>, is what lets recovery treat any bad record as the end of the log.</p><h2 id="implementation">implementation</h2><h3 id="the-segment">the segment</h3><p>A segment is one <code>.log</code> file and one <code>.index</code> file, named by the base offset padded to 20 digits so that a directory listing sorts in offset order. The index lives in memory as a vector of <code>{u32 relative offset, u32 file position}</code> pairs, and new entries are appended to the index file after the data they point at has been written.</p><p>The index loop is the part I got wrong first (see <a href="#problems">problems</a>). The current version walks the batch record by record while appending and drops an entry at every record boundary where 4 KiB or more has been written since the previous entry.</p><figure class="highlight cpp"><table><tr><td class="code"><pre><span class="line"><span class="function"><span class="type">void</span> <span class="title">Segment::append</span><span class="params">(<span class="type">const</span> <span class="type">uint8_t</span>* data, <span class="type">size_t</span> n, <span class="type">uint64_t</span> first_offset, <span class="type">uint32_t</span> count)</span> </span>&#123;</span><br><span class="line">  <span class="type">const</span> <span class="type">size_t</span> old_entries = index_.<span class="built_in">size</span>();</span><br><span class="line">  <span class="type">size_t</span> pos = <span class="number">0</span>;</span><br><span class="line">  <span class="keyword">for</span> (<span class="type">uint32_t</span> i = <span class="number">0</span>; i &lt; count; ++i) &#123;</span><br><span class="line">    <span class="keyword">if</span> (bytes_since_index_ &gt;= index_interval_) &#123;</span><br><span class="line">      index_.<span class="built_in">push_back</span>(&#123;<span class="keyword">static_cast</span>&lt;<span class="type">uint32_t</span>&gt;(first_offset + i - base_),</span><br><span class="line">                        <span class="keyword">static_cast</span>&lt;<span class="type">uint32_t</span>&gt;(size_ + pos)&#125;);</span><br><span class="line">      bytes_since_index_ = <span class="number">0</span>;</span><br><span class="line">    &#125;</span><br><span class="line">    <span class="type">const</span> <span class="type">size_t</span> rec = kRecordHeaderSize + <span class="built_in">load_u32</span>(data + pos + <span class="number">8</span>);</span><br><span class="line">    bytes_since_index_ += rec;</span><br><span class="line">    pos += rec;</span><br><span class="line">  &#125;</span><br><span class="line">  <span class="built_in">pwrite_all</span>(log_fd_, data, n, size_);</span><br><span class="line">  <span class="comment">// new index entries go out in one write, after the data they point at</span></span><br><span class="line">  ...</span><br><span class="line">&#125;</span><br></pre></td></tr></table></figure><p>Lookups use <code>std::upper_bound</code> to find the first entry past the target and step back one. Reads go through a small buffered reader that fetches a window of the file with one <code>pread</code> and serves header lookups from it, so a fetch is usually one or two system calls rather than one per record.</p><h3 id="recovery">recovery</h3><p>Only the last segment of a partition can have a torn tail, because before rolling to a new segment the outgoing one is flushed with the configured policy. So on startup only the active segment is scanned. Each record must have the expected next offset, a sane length and a matching CRC, and the file is truncated at the first record that fails.</p><figure class="highlight cpp"><table><tr><td class="code"><pre><span class="line"><span class="keyword">while</span> (limit - pos &gt;= kRecordHeaderSize) &#123;</span><br><span class="line">  <span class="type">const</span> <span class="type">uint8_t</span>* h = r.<span class="built_in">at</span>(pos, kRecordHeaderSize);</span><br><span class="line">  <span class="type">const</span> <span class="type">uint64_t</span> off = <span class="built_in">load_u64</span>(h);</span><br><span class="line">  <span class="type">const</span> <span class="type">uint32_t</span> len = <span class="built_in">load_u32</span>(h + <span class="number">8</span>);</span><br><span class="line">  <span class="type">const</span> <span class="type">uint32_t</span> crc = <span class="built_in">load_u32</span>(h + <span class="number">12</span>);</span><br><span class="line">  <span class="keyword">if</span> (off != expect || len &lt; kPayloadFixedSize || len &gt; kMaxPayloadSize ||</span><br><span class="line">      len &gt; limit - pos - kRecordHeaderSize)</span><br><span class="line">    <span class="keyword">break</span>;</span><br><span class="line">  <span class="type">const</span> <span class="type">uint8_t</span>* body = r.<span class="built_in">at</span>(pos + kRecordHeaderSize, len);</span><br><span class="line">  <span class="keyword">if</span> (<span class="built_in">crc32c</span>(body, len) != crc) <span class="keyword">break</span>;</span><br><span class="line">  ...  <span class="comment">// rebuild index entries as we go</span></span><br><span class="line">  pos += kRecordHeaderSize + len;</span><br><span class="line">  ++expect;</span><br><span class="line">&#125;</span><br><span class="line"><span class="keyword">if</span> (pos != size_) <span class="built_in">ftruncate</span>(log_fd_, pos);</span><br></pre></td></tr></table></figure><p>The expected-offset check matters as much as the CRC. A torn write can leave a header whose length field happens to be plausible, and the offset is a second, independent check that the bytes belong where they are. Sealed segments trust their index, but if an index file is missing, has a partial entry, or is not strictly increasing in both offset and position, the segment rebuilds it by the same scan. The cost of scanning only the active segment is that corruption inside a sealed segment is not detected at startup. A consumer would see a CRC error when it reached it.</p><p>CRC-32C uses the ARMv8 CRC instructions, eight bytes per <code>__crc32cd</code>, with a table-driven software fallback. A test checks the two against each other and against known vectors.</p><h3 id="reads-without-the-lock">reads without the lock</h3><p>The partition mutex serialises appends. I did not want a slow consumer's disk read to block a producer, so a read holds the lock only to decide what to read.</p><figure class="highlight cpp"><table><tr><td class="code"><pre><span class="line">&#123;</span><br><span class="line">  <span class="function">std::lock_guard <span class="title">lk</span><span class="params">(mu_)</span></span>;</span><br><span class="line">  ...  <span class="comment">// range checks, pick segment by binary search</span></span><br><span class="line">  seg = *std::<span class="built_in">prev</span>(it);                   <span class="comment">// shared_ptr&lt;Segment&gt;</span></span><br><span class="line">  start_pos = seg-&gt;<span class="built_in">floor_position</span>(offset);</span><br><span class="line">  limit = seg-&gt;<span class="built_in">size</span>();                    <span class="comment">// snapshot, bytes below this are immutable</span></span><br><span class="line">&#125;</span><br><span class="line">res.data = seg-&gt;<span class="built_in">read_from</span>(start_pos, offset, max_bytes, limit);</span><br></pre></td></tr></table></figure><p>Two things make this safe. Bytes below the snapshotted size never change, so a concurrent append cannot tear what the read sees. And the read holds a <code>shared_ptr</code> to the segment, so if retention deletes the segment in the middle of the read, the file is unlinked but the descriptor stays open until the last reference drops.</p><h3 id="retention">retention</h3><p>Size retention deletes whole segments from the front, but only if what remains is still at least the retention target, and it never deletes the active segment. After retention runs, the log's size is between the target and the target plus one segment. The internal offsets log has retention forced off, because without compaction retention would silently forget committed offsets.</p><h3 id="long-polling-without-a-lost-wakeup">long polling without a lost wakeup</h3><p>A consumer that is caught up sends a fetch with a maximum wait. The broker should park that request and wake it when data arrives, but it must not miss a produce that lands between the fetch's read and its wait. The usual bug is a wakeup sent before the waiter is waiting.</p><p>The fetch registers itself as a waiter before its first read, snapshots an epoch counter, reads, and then waits for the epoch to change. A produce increments the epoch after its append, and only if the waiter count is above zero, so producers pay nothing when no one is polling.</p><figure class="highlight cpp"><table><tr><td class="code"><pre><span class="line"><span class="keyword">if</span> (guard.c) guard.c-&gt;<span class="built_in">fetch_add</span>(<span class="number">1</span>, std::memory_order_acq_rel);   <span class="comment">// register first</span></span><br><span class="line"><span class="keyword">for</span> (;;) &#123;</span><br><span class="line">  <span class="type">uint64_t</span> epoch;</span><br><span class="line">  &#123; <span class="function">std::lock_guard <span class="title">lk</span><span class="params">(data_mu_)</span></span>; epoch = data_epoch_; &#125;         <span class="comment">// snapshot before reading</span></span><br><span class="line">  ...  <span class="comment">// read every requested partition</span></span><br><span class="line">  <span class="keyword">if</span> (any || max_wait_ms == <span class="number">0</span> || stopping_ || now &gt;= deadline) <span class="keyword">break</span>;</span><br><span class="line">  <span class="function">std::unique_lock <span class="title">lk</span><span class="params">(data_mu_)</span></span>;</span><br><span class="line">  data_cv_.<span class="built_in">wait_until</span>(lk, deadline, [&amp;] &#123; <span class="keyword">return</span> data_epoch_ != epoch || stopping_.<span class="built_in">load</span>(); &#125;);</span><br><span class="line">&#125;</span><br></pre></td></tr></table></figure><p>The partition mutex orders the fetch's read against the producer's append. Either the read sees the new data, or the append happened after it, in which case the producer sees the registered waiter and bumps the epoch, which the wait predicate then observes.</p><h3 id="consumer-groups-and-fencing">consumer groups and fencing</h3><p>A consumer joins a group with the topics it wants. The coordinator keeps, per group, a generation number, an ordered map of members, and the current assignment. Any membership change, whether a join, a leave or a session expiry, bumps the generation and recomputes a range assignment on the broker. Each subscribed member gets a contiguous block of partitions, and the first few get one extra when the count does not divide evenly.</p><p>Kafka splits this into JoinGroup and SyncGroup, with a barrier in between so that every member has given up its old partitions before anyone gets new ones. I left SyncGroup out. The cost is that during a rebalance two consumers can briefly fetch the same partition, until the old owner's next heartbeat tells it the generation changed. What keeps that from corrupting progress is fencing on commit.</p><figure class="highlight cpp"><table><tr><td class="code"><pre><span class="line"><span class="function">ErrorCode <span class="title">GroupCoordinator::check_commit</span><span class="params">(<span class="type">const</span> std::string&amp; group, <span class="type">const</span> std::string&amp; member_id,</span></span></span><br><span class="line"><span class="params"><span class="function">                                         <span class="type">uint32_t</span> generation)</span> </span>&#123;</span><br><span class="line">  <span class="keyword">if</span> (member_id.<span class="built_in">empty</span>()) <span class="keyword">return</span> ErrorCode::None;  <span class="comment">// standalone consumer, no fencing</span></span><br><span class="line">  ...</span><br><span class="line">  <span class="keyword">if</span> (!g.members.<span class="built_in">count</span>(member_id)) <span class="keyword">return</span> ErrorCode::UnknownMember;</span><br><span class="line">  <span class="keyword">if</span> (generation != g.generation) <span class="keyword">return</span> ErrorCode::IllegalGeneration;</span><br><span class="line">  <span class="keyword">return</span> ErrorCode::None;</span><br><span class="line">&#125;</span><br></pre></td></tr></table></figure><p>A consumer that was rebalanced away cannot commit, so it cannot move the new owner's offset backwards or forwards. The overlap turns into duplicate delivery, never into lost or rewound progress, which is at-least-once.</p><p>Committed offsets are records in an internal partition, <code>__consumer_offsets-0</code>, keyed by group, topic and partition, with the 8-byte offset as the value. The offset store holds one mutex across the append and the map update, so the log order and the in-memory order agree, and startup replays the log into the map. That gives commits exactly the same durability path as data with no second storage engine.</p><h3 id="the-clients">the clients</h3><p>The producer keeps one buffer per partition and hashes keys with FNV-1a to choose one, or goes round robin for records without a key. The consumer fetches all of its assigned partitions in one long-poll request. Both sit on a <code>Client</code> with one blocking connection and one request in flight, which keeps ordering trivial and caps unbatched throughput.</p><h3 id="testing">testing</h3><p>31 GoogleTest cases cover the codec, the log (including torn and corrupt tails, a broken sealed index and retention across a restart), the group coordinator, and the broker over real TCP, where they check ordering, key stickiness, durability across a restart, long-poll wakeup, commit and resume, and fenced stale commits.</p><p>The concurrent ordering test is the one I trust most. Six producers, each with a different batch size, write 3,000 messages each to one partition. The test then reads the partition back and asserts that offsets are exactly 0, 1, 2 and so on with no gaps, and that each producer's messages appear in the order it sent them.</p><p>In my own runs the suite passed in a release build in 0.85 s, under ASan plus UBSan in 4.08 s, and under TSan (<code>results/test_summary.txt</code>). A smoke test runs the real broker binary and CLI, consumes with a group, restarts the broker, produces one more message, and checks that the group sees only the new one while a new group sees everything (<code>results/smoke_test.txt</code>). On 2026-09-26 an independent clean release rebuild passed all 31 tests. The sanitizer and TSan builds were not rerun independently.</p><h2 id="problems">problems</h2><h3 id="1-the-sparse-index-was-only-sparse-between-batches">1. the sparse index was only sparse between batches</h3><p>My first <code>Segment::append</code> added at most one index entry per call, at the start of the batch. Tests passed, because every lookup was still correct. It was only slow.</p><p>It showed up in the first consumer benchmark. Reading 10,000 B messages with 4 KiB fetches ran at <strong>14.5 MB/s</strong>, and no fetch size got the median above <strong>82.5 MB/s</strong> (<code>results/run1_batch_start_index/consume_throughput.csv</code>). A consumer reading 10 KB records should be the easy case, so the shape was the clue rather than the level. Larger messages were doing worse than small ones.</p><p>The benchmark fills the partition with batches of about 1 MiB, which for 10 KB messages is 104 records per produce. With one index entry per batch, a fetch for a record in the middle of a batch had to start from the batch's first record and skip forward through up to 1 MiB of headers. With 4 KiB fetches, each fetch returns one 10 KB record, so reading one batch meant 104 fetches each skipping on average half a megabyte, about 50 times more bytes read than delivered by my arithmetic. The index was sparse between batches but not within them, and with big batches that is not sparse at all. (The devlog describes this as one entry per 10 MB, but the benchmark code fills 10 KB messages in batches of 104, so about 1 MiB is the right figure.)</p><p>The fix is the loop shown above, which indexes every record boundary past the 4 KiB interval, including inside a batch. <code>Segment.IndexHasEntriesInsideLargeBatches</code> appends a 1 MiB batch in one call, checks there are at least <code>size / 4096 - 1</code> entries, and checks that for offsets 1, 500 and 999 the floor position is within one interval plus one record of the target and that reading from it returns exactly that record.</p><p>After the fix the same point measured <strong>225.3 MB/s</strong>, and 4 MiB fetches reached <strong>2,336.5 MB/s</strong> (<code>results/consume_throughput.csv</code>). I only claim the direction from that comparison, not the ratio. The first run was at a load average of 275 to 330 on 14 cores and the second at about 9, so machine load and the fix are confounded.</p><h3 id="2-apple-clang-s-sanitizer-runtime-hangs">2. Apple clang's sanitizer runtime hangs</h3><p>On macOS 26.5 with Apple clang 17, the ASan build deadlocked during runtime initialisation, before <code>main</code>, even for an empty program. That rules out my code. I moved the sanitizer and TSan builds to Homebrew LLVM, and <code>CMakeLists.txt</code> now prints a warning if you ask for a sanitizer with the Apple toolchain. Upstream clang emits DWARF 5, which Apple's linker warns about, so non-Apple builds pass <code>-gdwarf-4</code>.</p><h3 id="3-the-machine-was-not-mine">3. the machine was not mine</h3><p>The first two full benchmark runs happened at load averages between 138 and 330 on 14 cores, with other agents' builds and benchmarks running. The numbers were low and unstable. For example, 1000 B messages at batch 1000 measured <strong>294.5 MB/s</strong> in the loaded run (<code>results/run2_loaded/produce_throughput.csv</code>) and <strong>2,528.0 MB/s</strong> once load dropped to about 8 (<code>results/produce_throughput.csv</code>), more than eight times faster for the same code.</p><p>I could not get an idle machine, so I changed the harness instead. Every CSV row now records the 1 minute load average at the time it was measured, every throughput point is the median of 3 repetitions with the min and max kept, and I reran everything once load was around 8 to 11. The loaded runs are kept in <code>results/run2_loaded/</code> and <code>results/run1_batch_start_index/</code> rather than deleted. One detail to watch in the files is that <code>results/bench_env.txt</code> still has a &quot;load average before run&quot; of 204.79 in its header, which was recorded before the loaded run. The per-row <code>load1</code> column is what describes the reported numbers.</p><p>Even at low load, nominally identical configurations disagree. 1000 B messages at batch 100 to one partition with no flush measured <strong>1,001.6 MB/s</strong> in the producer sweep, <strong>1,187.7 MB/s</strong> in the flush sweep, and <strong>1,340.3 MB/s</strong> as the single-producer point of the scaling sweep. That spread of about a third is a fair estimate of how much to trust any single number in this post.</p><h3 id="4-one-noisy-point-even-at-low-load">4. one noisy point even at low load</h3><p>In the first low-load producer run, 100 B messages at batch 100 measured <strong>14.6 MB/s</strong>, below batch 10, with a min to max spread of 11.7 to 38.6 (<code>results/produce_throughput_noisy.csv</code>). A curve that goes down when batching goes up usually means Nagle's algorithm holding small writes, but <code>TCP_NODELAY</code> was already set. Running the same mode again gave <strong>233.4 MB/s</strong> with a spread of 229.6 to 240.2 (<code>results/produce_throughput.csv</code>). A tight spread on the rerun against a threefold spread in the original points to interference from other processes, not to my code. I kept both files and report the second.</p><h3 id="5-slower-send-rates-had-worse-tails">5. slower send rates had worse tails</h3><p>Without a flush, p99 latency at 1,000 messages per second was <strong>3.02 ms</strong>, while at 10,000 per second it was <strong>142 µs</strong> (<code>results/latency.csv</code>). Sending less made the tail about twenty times worse.</p><p>My explanation is that at the low rate the producer sleeps between sends, and the broker's connection thread and the consumer's long-poll thread park, so every message pays for waking threads up, and probably for cores leaving low-power states. At 10,000 per second the threads stay warm. The same pattern appears in the loaded and unloaded runs. I have not confirmed it by profiling, so it is a hypothesis.</p><h2 id="experiments">experiments</h2><p>All experiments use <code>mk-bench</code>, which starts a broker in the same process with a fresh data directory on the internal SSD and talks to it over loopback TCP through the real client code. Throughput is counted as payload bytes (the message value, not the 28 bytes of record overhead) in decimal megabytes. Throughput points are the median of 3 repetitions.</p><p>The hardware is an Apple M4 Pro with 14 cores and 48 GiB, running macOS 26.5, built with Apple clang 17 at <code>-O2</code> (<code>results/bench_env.txt</code>). The machine was running other jobs throughout. The 1 minute load average was 8 to 11 during the reported runs and is recorded per row. None of the benchmarks were rerun independently, so these are single-session, loaded-machine numbers.</p><ol><li>Producer throughput against batch size (1, 10, 100 and 1,000 records) and message size (100 B, 1 KB, 10 KB), one partition, no fsync. The batch is pre-encoded, so this measures transport plus broker (<code>results/produce_throughput.csv</code>).</li><li>Consumer throughput against fetch <code>max_bytes</code> (4 KiB to 4 MiB) and message size, reading a pre-filled partition from offset 0 and verifying every record's CRC (<code>results/consume_throughput.csv</code>).</li><li>End-to-end latency. A producer paced to a fixed schedule stamps each message with a monotonic clock, and a consumer long-polls and records receive time minus send time. Five configurations cover send rate, batch size and flush policy (<code>results/latency.csv</code>, with every sample in <code>results/latency_raw.csv</code>).</li><li>Cost of durability, with 1 KB messages at each batch size under the three flush policies (<code>results/flush_policy.csv</code>).</li><li>Producer scaling with 1, 2, 4 and 8 concurrent producers, each to its own partition, 1 KB messages, batch 100 (<code>results/producer_scaling.csv</code>).</li></ol><h2 id="results">results</h2><h3 id="batching-dominates-producer-throughput">batching dominates producer throughput</h3><figure data-figure="chart:projects/kafka-from-scratch/kafka-from-scratch-batching"></figure><p>With 100 B messages, going from batch 1 to batch 1000 raised throughput from <strong>32,886 to 12,744,103 messages per second</strong> (3.3 to 1,274.4 MB/s), a factor of about 390. At batch 1, every message size lands between 25,457 and 32,886 messages per second. That is the rate of one blocking round trip per connection, so at batch 1 the protocol is the limit, not the disk or the broker.</p><p>Large messages saturate earlier. 10 KB messages reach <strong>1,311.4 MB/s</strong> at batch 10 but only <strong>2,020.6 MB/s</strong> at batch 1000, because a 10-record batch of 10 KB is already 100 KB per request and the fixed costs are amortised. The best point was 1 KB messages at batch 1000, <strong>2,528.0 MB/s</strong>, or about 2.5 million messages per second into one partition.</p><h3 id="fetch-size-does-the-same-for-consumers">fetch size does the same for consumers</h3><figure data-figure="chart:projects/kafka-from-scratch/kafka-from-scratch-fetch-size"></figure><p>100 B messages went from <strong>92.8 MB/s</strong> with 4 KiB fetches to <strong>2,150.1 MB/s</strong> (21.5 million messages per second) with 1 MiB fetches, and 4 MiB did not help further (2,134.9 MB/s). These numbers include a CRC check of every record on the consumer.</p><p>10 KB messages are flat at <strong>225.3 and 219.7 MB/s</strong> for 4 KiB and 16 KiB fetches, because each fetch returns exactly one record. The first record is always returned whole even when it exceeds <code>max_bytes</code>, so the fetch size stops mattering below the record size. They reach <strong>2,336.5 MB/s</strong> at 4 MiB. 1 KB messages are the odd ones out, topping out at <strong>1,645.8 MB/s</strong>, with a wide min to max spread at 1 MiB (725.7 to 1,785.0), which I read as noise rather than a real knee.</p><h3 id="latency">latency</h3><figure data-figure="chart:projects/kafka-from-scratch/kafka-from-scratch-latency"></figure><p>Without fsync, end-to-end latency is tens of microseconds at the median. At 10,000 messages per second, p50 was <strong>41.5 µs</strong>, p99 <strong>142 µs</strong> and p99.9 <strong>1.21 ms</strong>. The produce acknowledgement alone had a p50 of 32.9 µs in the same run, which suggests that most of the median is the cost of one request crossing loopback TCP and the broker, not the long poll. At 50,000 per second in batches of 10, p50 was <strong>45.5 µs</strong> and p99 <strong>185 µs</strong>.</p><p>With <code>fsync</code> at 1,000 per second, p50 was <strong>115 µs</strong> and p99 <strong>586 µs</strong>. With <code>F_FULLFSYNC</code> at 200 per second, p50 was <strong>4.03 ms</strong> and p99 <strong>7.79 ms</strong>. That median is about 35 times the plain <code>fsync</code> median, and it is the cost of actually getting data past the drive's write cache, and it is the honest price of single-node durability on this machine.</p><h3 id="durability-costs-depend-on-batching">durability costs depend on batching</h3><figure data-figure="chart:projects/kafka-from-scratch/kafka-from-scratch-flush-policy"></figure><p>At batch 100 with 1 KB messages, throughput was <strong>1,187.7 MB/s</strong> with no flush, <strong>594.9 MB/s</strong> with <code>fsync</code> and <strong>23.2 MB/s</strong> with <code>F_FULLFSYNC</code>. At batch 1000 the gap narrows to 1,519.5, 1,324.1 and 197.6 MB/s, because each flush is amortised over about 1 MB. At batch 1, <code>F_FULLFSYNC</code> managed <strong>231 messages per second</strong>, about 4.3 ms per message.</p><p>Each flush happens per append while the partition lock is held, so concurrent producers to one partition queue behind each other's flushes. Group commit would fix that but needs asynchronous acknowledgements, which v0 does not have.</p><h3 id="producer-scaling-is-sublinear">producer scaling is sublinear</h3><p>Aggregate throughput was <strong>1,340.3 MB/s</strong> with one producer, 1,490.7 with two, 1,812.9 with four and <strong>2,061.2 with eight</strong> (<code>results/producer_scaling.csv</code>), 1.54 times the single-producer rate with eight times the producers. Each producer has its own partition and its own lock, so it is not contention on the log. Plausible limits are memory bandwidth for copying through loopback TCP and the page cache, and the other load on the machine. I have not profiled it, and given the one-third spread between nominally identical runs noted in <a href="#problems">problems</a>, I would not read much into the shape beyond &quot;sublinear&quot;.</p><h3 id="what-load-does-to-all-of-this">what load does to all of this</h3><p>For contrast, the second loaded run, at per-row load averages of 138 to 231, gave <strong>294.5 MB/s</strong> for 1 KB messages at batch 1000 and <strong>382.0 MB/s</strong> at best for consumers (<code>results/run2_loaded/</code>). The first run at 275 to 330 had a p99 of <strong>172 ms</strong> at 1,000 messages per second without fsync (<code>results/run1_batch_start_index/latency.csv</code>), against 3.02 ms at low load. Throughput from a shared machine is a lower bound, and tail latency from a shared machine mostly measures the other jobs.</p><h2 id="what-i-would-change">what I would change</h2><h3 id="pipeline-producer-requests">pipeline producer requests</h3><p>One request in flight caps unbatched throughput at about 25,000 to 33,000 messages per second per connection. Allowing several in flight is easy for throughput and hard for ordering, since a retry of an earlier batch could land after a later one. Kafka solves that with per-partition sequence numbers checked by the broker, which is also the first step toward idempotent producers and exactly-once. I would build them together.</p><h3 id="add-a-syncgroup-barrier">add a SyncGroup barrier</h3><p>Today two consumers can briefly read the same partition during a rebalance. Generation fencing turns that into duplicate delivery rather than lost progress, but it is still a weaker guarantee than Kafka's, where the old owner revokes before the new one starts.</p><h3 id="an-event-loop-sendfile-and-group-commit">an event loop, sendfile and group commit</h3><p>Thread per connection kept every handler straight-line code and was fine for a handful of benchmark connections, but it will not survive thousands of clients. A kqueue event loop is the right design there, and since the on-disk format already equals the wire format, fetches could then use <code>sendfile</code> and skip the copy through user space. The same asynchronous acknowledgement path would allow group-commit fsyncs across concurrent producers.</p><h3 id="compact-the-offsets-log">compact the offsets log</h3><p><code>__consumer_offsets</code> only grows, and startup replays all of it. Compaction, keeping only the latest record per key, is what Kafka uses to bound it, and it is on the roadmap along with time-based retention.</p><h3 id="profile-before-explaining">profile before explaining</h3><p>Producer scaling and the low-rate latency tail both have plausible explanations in this post that I did not verify. I would rerun them on an idle machine with sampling (Instruments on macOS) before trusting either, and I would rerun every benchmark on an idle machine before quoting any number here as more than indicative.</p><h2 id="reproducibility">reproducibility</h2><p>Requirements are CMake 3.24 or newer, Ninja and a C++20 compiler. The project was developed with Apple clang 17 on macOS 26.5 on an Apple M4 Pro, and GoogleTest is fetched by CMake. From <code>projects/04-minikafka-log-broker</code>, the following commands build everything.</p><figure class="highlight sh"><table><tr><td class="code"><pre><span class="line"><span class="comment"># Release (-O2)</span></span><br><span class="line">cmake -S . -B build/release -G Ninja -DCMAKE_BUILD_TYPE=Release</span><br><span class="line">cmake --build build/release</span><br><span class="line"></span><br><span class="line"><span class="comment"># ASan + UBSan, with Homebrew LLVM because Apple clang 17&#x27;s sanitizer runtime hangs on macOS 26.5</span></span><br><span class="line">cmake -S . -B build/debug -G Ninja -DCMAKE_BUILD_TYPE=Debug -DMK_SANITIZE=ON \</span><br><span class="line">      -DCMAKE_CXX_COMPILER=<span class="string">&quot;<span class="subst">$(brew --prefix llvm)</span>/bin/clang++&quot;</span></span><br><span class="line">cmake --build build/debug</span><br><span class="line"></span><br><span class="line"><span class="comment"># ThreadSanitizer</span></span><br><span class="line">cmake -S . -B build/tsan -G Ninja -DCMAKE_BUILD_TYPE=Debug -DMK_TSAN=ON -DMK_BUILD_BENCH=OFF \</span><br><span class="line">      -DCMAKE_CXX_COMPILER=<span class="string">&quot;<span class="subst">$(brew --prefix llvm)</span>/bin/clang++&quot;</span></span><br><span class="line">cmake --build build/tsan</span><br></pre></td></tr></table></figure><p>Tests and the smoke test.</p><figure class="highlight sh"><table><tr><td class="code"><pre><span class="line">ctest --test-dir build/release --output-on-failure</span><br><span class="line">ctest --test-dir build/debug --output-on-failure     <span class="comment"># ASan + UBSan</span></span><br><span class="line">./build/tsan/mk-tests                                 <span class="comment"># TSan</span></span><br><span class="line">./scripts/smoke_test.sh                               <span class="comment"># real broker + CLI, including a restart</span></span><br></pre></td></tr></table></figure><p>Benchmarks and plots. The script runs one mode at a time and writes every CSV in <code>results/</code>, with the load average in each row. <code>results/latency_raw.csv</code> is gitignored and regenerated.</p><figure class="highlight sh"><table><tr><td class="code"><pre><span class="line">./scripts/run_benchmarks.sh 3</span><br><span class="line">python3 -m venv .venv &amp;&amp; .venv/bin/pip install matplotlib</span><br><span class="line">.venv/bin/python scripts/plot.py</span><br></pre></td></tr></table></figure><p>Code is in <code>projects/04-minikafka-log-broker</code>, with the library in <code>src/</code> and <code>include/minikafka/</code>, the broker and CLI in <code>tools/</code>, the benchmark harness in <code>bench/</code>, the tests in <code>tests/</code>, the formats, invariants and trade-offs in <code>DESIGN.md</code>, and the build log in <code>DEVLOG.md</code>.</p>]]>
    </content>
    <id>https://projects.farhansadeek.com/posts/what-durability-costs-a-kafka-style-broker/</id>
    <link href="https://projects.farhansadeek.com/posts/what-durability-costs-a-kafka-style-broker/"/>
    <published>2026-09-29T06:17:11.000Z</published>
    <summary>A single-broker event log in C++ with segmented partitions, crash recovery and consumer groups, and a measured look at what each flush policy costs.</summary>
    <title>What Durability Costs a Kafka-Style Log Broker</title>
    <updated>2026-09-29T06:49:26.552Z</updated>
  </entry>
  <entry>
    <author>
      <name>Farhan Sadeek</name>
      <email>farhan@farhansadeek.com</email>
    </author>
    <category term="trading" scheme="https://projects.farhansadeek.com/tags/trading/"/>
    <category term="market-microstructure" scheme="https://projects.farhansadeek.com/tags/market-microstructure/"/>
    <category term="simulation" scheme="https://projects.farhansadeek.com/tags/simulation/"/>
    <category term="python" scheme="https://projects.farhansadeek.com/tags/python/"/>
    <content>
      <![CDATA[<p>I built a small agent-based market to watch one textbook claim happen in a real order book. A dealer earns the spread from uninformed traders and pays some of it back to informed ones, and as the informed share of order flow grows, a dealer who does not widen should make less money. In the simulator it does. As the informed fraction of taker arrivals goes from 0 to 0.6, adverse selection per fill for a fixed-formula market maker rises from <strong>0.605 to 1.266 ticks</strong>, and its PnL per fill falls from <strong>2.210 to 0.881 ticks</strong>, a drop of 60%. A second market maker that adds its own measured markout loss to its half-spread widens from <strong>6.85 to 8.16 ticks</strong> and keeps PnL per fill at 2.0 or above everywhere.</p><p>The market has four kinds of agents on one event heap. Noise traders send market orders as a Poisson process. Informed traders see a noisy signal of a hidden fundamental value that follows a random walk with jumps. Noise liquidity providers post passive limit orders. One Avellaneda-Stoikov market maker quotes around a learned fair value and skews by inventory. It never sees the fundamental, but in the default mode it is told the true informed fraction and scales its belief updates by it, so the headline numbers describe a dealer who knows the informed share exactly. Everything is pure Python and seeded, and an independent rerun reproduced the sweep, the ablation and the stylized facts byte for byte.</p><p>The market-making result was the easy part. Getting there took three bugs in the market maker's model of the world, one of which crashed the price by 98 ticks with no news, and one performance bug in CPython's <code>dict</code> that I would not have guessed. The stylized facts are more modest than the market-making story. Fat tails are present but mostly come from the jumps I put in, and volatility clustering dies within about five lags.</p><p>Code is in <code>projects/10-adverse-selection-lob-sim</code>. Every number below is read from a file in its <code>results/</code> directory, and the file is named next to the number. The exceptions are in the verification note under <a href="#reproducibility">reproducibility</a>, which come from the project's devlog.</p><p><em>Reading note.</em> The problems section is the most useful part. If you only want the claim and its limits, read <a href="#results">results</a> and <a href="#what-i-would-change">what i would change</a>.</p><h2 id="table-of-contents">table of contents</h2><ul><li><a href="#what-i-wanted-to-build">what i wanted to build</a></li><li><a href="#theory">theory</a></li><li><a href="#architecture">architecture</a></li><li><a href="#implementation">implementation</a></li><li><a href="#problems">problems</a></li><li><a href="#experiments">experiments</a></li><li><a href="#results">results</a></li><li><a href="#what-i-would-change">what i would change</a></li><li><a href="#reproducibility">reproducibility</a></li></ul><h2 id="what-i-wanted-to-build">what i wanted to build</h2><p>I wanted a market small enough to reason about completely. That meant a real limit order book with price-time priority, a clock that decides who acts when, and agents whose behavior comes from the standard microstructure models rather than from tuning. The target was the core story of market making. The dealer is paid by noise, pays informed traders, and should respond to more informed flow by quoting wider.</p><p>The second goal was honesty about stylized facts. Real returns have fat tails, almost no linear autocorrelation, and long-lived autocorrelation in absolute returns. Agent-based models are often shown reproducing these, and I wanted to know which ones a model this simple produces on its own and which ones it only produces because I built them in.</p><p>The v0 scope was an event scheduler, a Python order book, four agent types, stylized-fact measurements, and a sweep over the informed fraction. The roadmap after that is an RL market maker, latency, multiple venues, calibration to real data, and a C++ matching core shared with the exchange project. None of that is in this post.</p><h2 id="theory">theory</h2><p>Three models do the work here, plus one accounting identity.</p><h3 id="glosten-milgrom">Glosten-Milgrom</h3><p>A competitive dealer faces a stream of traders. A fraction <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>α</mi></mrow><annotation encoding="application/x-tex">\alpha</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal" style="margin-right:0.0037em;">α</span></span></span></span> of them know the true value <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>V</mi></mrow><annotation encoding="application/x-tex">V</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal" style="margin-right:0.2222em;">V</span></span></span></span>, and the rest buy or sell at random. The dealer cannot tell them apart, so it sets the ask to the expected value conditional on the next trade being a buy, and the bid to the expected value conditional on a sell.</p><p class="katex-block"><span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mtext>ask</mtext><mo>=</mo><mi mathvariant="double-struck">E</mi><mo stretchy="false">[</mo><mi>V</mi><mo>∣</mo><mtext>buy</mtext><mo stretchy="false">]</mo><mo separator="true">,</mo><mspace width="2em"/><mtext>bid</mtext><mo>=</mo><mi mathvariant="double-struck">E</mi><mo stretchy="false">[</mo><mi>V</mi><mo>∣</mo><mtext>sell</mtext><mo stretchy="false">]</mo></mrow><annotation encoding="application/x-tex">\text{ask} = \mathbb{E}[V \mid \text{buy}], \qquad \text{bid} = \mathbb{E}[V \mid \text{sell}]</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6944em;"></span><span class="mord text"><span class="mord">ask</span></span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mord mathbb">E</span><span class="mopen">[</span><span class="mord mathnormal" style="margin-right:0.2222em;">V</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">∣</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mord text"><span class="mord">buy</span></span><span class="mclose">]</span><span class="mpunct">,</span><span class="mspace" style="margin-right:2em;"></span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mord text"><span class="mord">bid</span></span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mord mathbb">E</span><span class="mopen">[</span><span class="mord mathnormal" style="margin-right:0.2222em;">V</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">∣</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mord text"><span class="mord">sell</span></span><span class="mclose">]</span></span></span></span></span></p><p>A buy is more likely to come from an informed trader when <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>V</mi></mrow><annotation encoding="application/x-tex">V</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal" style="margin-right:0.2222em;">V</span></span></span></span> is high, so the ask sits above the prior mean. To first order the half-spread is about <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>α</mi><mi>D</mi></mrow><annotation encoding="application/x-tex">\alpha D</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal" style="margin-right:0.0037em;">α</span><span class="mord mathnormal" style="margin-right:0.0278em;">D</span></span></span></span>, where <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>D</mi></mrow><annotation encoding="application/x-tex">D</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal" style="margin-right:0.0278em;">D</span></span></span></span> is the informed trader's information advantage. Three consequences matter later. The spread exists purely because of adverse selection. It widens with <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>α</mi></mrow><annotation encoding="application/x-tex">\alpha</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal" style="margin-right:0.0037em;">α</span></span></span></span>. And each trade moves the dealer's belief by an amount proportional to <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>α</mi></mrow><annotation encoding="application/x-tex">\alpha</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal" style="margin-right:0.0037em;">α</span></span></span></span>, because a trade from a population that is mostly noise carries little information. That last point turned out to be the fix for a bug.</p><h3 id="avellaneda-stoikov">Avellaneda-Stoikov</h3><p>Glosten-Milgrom has no inventory. Avellaneda-Stoikov (2008) is the opposite, a risk-averse market maker with CARA utility and risk aversion <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>γ</mi></mrow><annotation encoding="application/x-tex">\gamma</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.625em;vertical-align:-0.1944em;"></span><span class="mord mathnormal" style="margin-right:0.0556em;">γ</span></span></span></span>, facing mid-price volatility <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>σ</mi></mrow><annotation encoding="application/x-tex">\sigma</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal" style="margin-right:0.0359em;">σ</span></span></span></span> and fills that arrive with intensity <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>A</mi><msup><mi>e</mi><mrow><mo>−</mo><mi>k</mi><mi>δ</mi></mrow></msup></mrow><annotation encoding="application/x-tex">A e^{-k\delta}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.8491em;"></span><span class="mord mathnormal">A</span><span class="mord"><span class="mord mathnormal">e</span><span class="msupsub"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.8491em;"><span style="top:-3.063em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mtight">−</span><span class="mord mathnormal mtight" style="margin-right:0.0315em;">k</span><span class="mord mathnormal mtight" style="margin-right:0.0379em;">δ</span></span></span></span></span></span></span></span></span></span></span></span> at distance <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>δ</mi></mrow><annotation encoding="application/x-tex">\delta</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6944em;"></span><span class="mord mathnormal" style="margin-right:0.0379em;">δ</span></span></span></span> from mid. The optimal quotes sit around a reservation price that is shaded by inventory <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>q</mi></mrow><annotation encoding="application/x-tex">q</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.625em;vertical-align:-0.1944em;"></span><span class="mord mathnormal" style="margin-right:0.0359em;">q</span></span></span></span>.</p><p class="katex-block"><span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mi>r</mi><mo>=</mo><mi>s</mi><mo>−</mo><mi>q</mi><mtext> </mtext><mi>γ</mi><msup><mi>σ</mi><mn>2</mn></msup><mo stretchy="false">(</mo><mi>T</mi><mo>−</mo><mi>t</mi><mo stretchy="false">)</mo><mo separator="true">,</mo><mspace width="2em"/><msup><mi>δ</mi><mi>a</mi></msup><mo>+</mo><msup><mi>δ</mi><mi>b</mi></msup><mo>=</mo><mi>γ</mi><msup><mi>σ</mi><mn>2</mn></msup><mo stretchy="false">(</mo><mi>T</mi><mo>−</mo><mi>t</mi><mo stretchy="false">)</mo><mo>+</mo><mfrac><mn>2</mn><mi>γ</mi></mfrac><mi>ln</mi><mo>⁡</mo><mtext> ⁣</mtext><mrow><mo fence="true">(</mo><mn>1</mn><mo>+</mo><mfrac><mi>γ</mi><mi>k</mi></mfrac><mo fence="true">)</mo></mrow></mrow><annotation encoding="application/x-tex">r = s - q\,\gamma\sigma^2 (T - t), \qquad \delta^a + \delta^b = \gamma\sigma^2 (T - t) + \frac{2}{\gamma}\ln\!\left(1 + \frac{\gamma}{k}\right)</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal" style="margin-right:0.0278em;">r</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord mathnormal">s</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">−</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:1.1141em;vertical-align:-0.25em;"></span><span class="mord mathnormal" style="margin-right:0.0359em;">q</span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mord mathnormal" style="margin-right:0.0556em;">γ</span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0359em;">σ</span><span class="msupsub"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.8641em;"><span style="top:-3.113em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight">2</span></span></span></span></span></span></span></span><span class="mopen">(</span><span class="mord mathnormal" style="margin-right:0.1389em;">T</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">−</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mord mathnormal">t</span><span class="mclose">)</span><span class="mpunct">,</span><span class="mspace" style="margin-right:2em;"></span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0379em;">δ</span><span class="msupsub"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.7144em;"><span style="top:-3.113em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight">a</span></span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">+</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:0.8991em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0379em;">δ</span><span class="msupsub"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.8991em;"><span style="top:-3.113em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight">b</span></span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:1.1141em;vertical-align:-0.25em;"></span><span class="mord mathnormal" style="margin-right:0.0556em;">γ</span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0359em;">σ</span><span class="msupsub"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.8641em;"><span style="top:-3.113em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight">2</span></span></span></span></span></span></span></span><span class="mopen">(</span><span class="mord mathnormal" style="margin-right:0.1389em;">T</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">−</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mord mathnormal">t</span><span class="mclose">)</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">+</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:2.2019em;vertical-align:-0.8804em;"></span><span class="mord"><span class="mopen nulldelimiter"></span><span class="mfrac"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:1.3214em;"><span style="top:-2.314em;"><span class="pstrut" style="height:3em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0556em;">γ</span></span></span><span style="top:-3.23em;"><span class="pstrut" style="height:3em;"></span><span class="frac-line" style="border-bottom-width:0.04em;"></span></span><span style="top:-3.677em;"><span class="pstrut" style="height:3em;"></span><span class="mord"><span class="mord">2</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.8804em;"><span></span></span></span></span></span><span class="mclose nulldelimiter"></span></span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mop">ln</span><span class="mspace" style="margin-right:-0.1667em;"></span><span class="mspace" style="margin-right:0.1667em;"></span><span class="minner"><span class="mopen delimcenter" style="top:0em;"><span class="delimsizing size2">(</span></span><span class="mord">1</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">+</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mord"><span class="mopen nulldelimiter"></span><span class="mfrac"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:1.1076em;"><span style="top:-2.314em;"><span class="pstrut" style="height:3em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0315em;">k</span></span></span><span style="top:-3.23em;"><span class="pstrut" style="height:3em;"></span><span class="frac-line" style="border-bottom-width:0.04em;"></span></span><span style="top:-3.677em;"><span class="pstrut" style="height:3em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0556em;">γ</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.686em;"><span></span></span></span></span></span><span class="mclose nulldelimiter"></span></span><span class="mclose delimcenter" style="top:0em;"><span class="delimsizing size2">)</span></span></span></span></span></span></span></p><p>A long dealer shades both quotes down so that it is more likely to sell. I hold <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>T</mi><mo>−</mo><mi>t</mi></mrow><annotation encoding="application/x-tex">T - t</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.7667em;vertical-align:-0.0833em;"></span><span class="mord mathnormal" style="margin-right:0.1389em;">T</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">−</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:0.6151em;"></span><span class="mord mathnormal">t</span></span></span></span> constant at a rolling horizon <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>τ</mi><mo>=</mo><mn>10</mn></mrow><annotation encoding="application/x-tex">\tau = 10</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal" style="margin-right:0.1132em;">τ</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.6444em;"></span><span class="mord">10</span></span></span></span> time units, which makes the quotes stationary instead of collapsing toward a terminal time. With the defaults (<span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>γ</mi><mo>=</mo><mn>0.1</mn></mrow><annotation encoding="application/x-tex">\gamma = 0.1</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.625em;vertical-align:-0.1944em;"></span><span class="mord mathnormal" style="margin-right:0.0556em;">γ</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.6444em;"></span><span class="mord">0.1</span></span></span></span>, <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>σ</mi><mo>=</mo><mn>1</mn></mrow><annotation encoding="application/x-tex">\sigma = 1</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal" style="margin-right:0.0359em;">σ</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.6444em;"></span><span class="mord">1</span></span></span></span>, <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>k</mi><mo>=</mo><mn>0.5</mn></mrow><annotation encoding="application/x-tex">k = 0.5</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6944em;"></span><span class="mord mathnormal" style="margin-right:0.0315em;">k</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.6444em;"></span><span class="mord">0.5</span></span></span></span>, <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>τ</mi><mo>=</mo><mn>10</mn></mrow><annotation encoding="application/x-tex">\tau = 10</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal" style="margin-right:0.1132em;">τ</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.6444em;"></span><span class="mord">10</span></span></span></span>) the half-spread is</p><p class="katex-block"><span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mi>h</mi><mo>=</mo><mfrac><mrow><mi>γ</mi><msup><mi>σ</mi><mn>2</mn></msup><mi>τ</mi></mrow><mn>2</mn></mfrac><mo>+</mo><mfrac><mn>1</mn><mi>γ</mi></mfrac><mi>ln</mi><mo>⁡</mo><mtext> ⁣</mtext><mrow><mo fence="true">(</mo><mn>1</mn><mo>+</mo><mfrac><mi>γ</mi><mi>k</mi></mfrac><mo fence="true">)</mo></mrow><mo>=</mo><mn>0.5</mn><mo>+</mo><mn>10</mn><mi>ln</mi><mo>⁡</mo><mn>1.2</mn><mo>≈</mo><mn>2.32</mn><mtext> ticks</mtext><mo separator="true">,</mo></mrow><annotation encoding="application/x-tex">h = \frac{\gamma\sigma^2\tau}{2} + \frac{1}{\gamma}\ln\!\left(1 + \frac{\gamma}{k}\right) = 0.5 + 10\ln 1.2 \approx 2.32 \text{ ticks},</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6944em;"></span><span class="mord mathnormal">h</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:2.1771em;vertical-align:-0.686em;"></span><span class="mord"><span class="mopen nulldelimiter"></span><span class="mfrac"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:1.4911em;"><span style="top:-2.314em;"><span class="pstrut" style="height:3em;"></span><span class="mord"><span class="mord">2</span></span></span><span style="top:-3.23em;"><span class="pstrut" style="height:3em;"></span><span class="frac-line" style="border-bottom-width:0.04em;"></span></span><span style="top:-3.677em;"><span class="pstrut" style="height:3em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0556em;">γ</span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0359em;">σ</span><span class="msupsub"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.8141em;"><span style="top:-3.063em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight">2</span></span></span></span></span></span></span></span><span class="mord mathnormal" style="margin-right:0.1132em;">τ</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.686em;"><span></span></span></span></span></span><span class="mclose nulldelimiter"></span></span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">+</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:2.2019em;vertical-align:-0.8804em;"></span><span class="mord"><span class="mopen nulldelimiter"></span><span class="mfrac"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:1.3214em;"><span style="top:-2.314em;"><span class="pstrut" style="height:3em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0556em;">γ</span></span></span><span style="top:-3.23em;"><span class="pstrut" style="height:3em;"></span><span class="frac-line" style="border-bottom-width:0.04em;"></span></span><span style="top:-3.677em;"><span class="pstrut" style="height:3em;"></span><span class="mord"><span class="mord">1</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.8804em;"><span></span></span></span></span></span><span class="mclose nulldelimiter"></span></span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mop">ln</span><span class="mspace" style="margin-right:-0.1667em;"></span><span class="mspace" style="margin-right:0.1667em;"></span><span class="minner"><span class="mopen delimcenter" style="top:0em;"><span class="delimsizing size2">(</span></span><span class="mord">1</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">+</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mord"><span class="mopen nulldelimiter"></span><span class="mfrac"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:1.1076em;"><span style="top:-2.314em;"><span class="pstrut" style="height:3em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0315em;">k</span></span></span><span style="top:-3.23em;"><span class="pstrut" style="height:3em;"></span><span class="frac-line" style="border-bottom-width:0.04em;"></span></span><span style="top:-3.677em;"><span class="pstrut" style="height:3em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0556em;">γ</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.686em;"><span></span></span></span></span></span><span class="mclose nulldelimiter"></span></span><span class="mclose delimcenter" style="top:0em;"><span class="delimsizing size2">)</span></span></span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.7278em;vertical-align:-0.0833em;"></span><span class="mord">0.5</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">+</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:0.6944em;"></span><span class="mord">10</span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mop">ln</span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mord">1.2</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">≈</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.8889em;vertical-align:-0.1944em;"></span><span class="mord">2.32</span><span class="mord text"><span class="mord"> ticks</span></span><span class="mpunct">,</span></span></span></span></span></p><p>and the skew is <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>γ</mi><msup><mi>σ</mi><mn>2</mn></msup><mi>τ</mi><mo>=</mo><mn>1</mn></mrow><annotation encoding="application/x-tex">\gamma\sigma^2\tau = 1</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1.0085em;vertical-align:-0.1944em;"></span><span class="mord mathnormal" style="margin-right:0.0556em;">γ</span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0359em;">σ</span><span class="msupsub"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.8141em;"><span style="top:-3.063em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight">2</span></span></span></span></span></span></span></span><span class="mord mathnormal" style="margin-right:0.1132em;">τ</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.6444em;"></span><span class="mord">1</span></span></span></span> tick per unit of inventory. Quotes are rounded outward to whole ticks, bid down and ask up, so a half-spread of 2.32 gives a quoted spread of 5 or 6 ticks depending on where the reservation price falls between ticks. The measured mean for the fixed market maker is 5.64 to 5.69 across the whole sweep, which is the first sanity check the numbers pass.</p><p>The model assumes fill intensity falls with distance from the mid. That assumption turned out to matter more than anything else in it, because it is what gives the inventory skew a restoring force.</p><h3 id="kyle">Kyle</h3><p>In Kyle (1985) price impact is linear in signed order flow. My market maker's fair-value update ended up in this form. Each trade moves fair value by a fixed step in the direction of the aggressor, and in the Glosten-Milgrom mode that step is scaled by <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>α</mi></mrow><annotation encoding="application/x-tex">\alpha</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal" style="margin-right:0.0037em;">α</span></span></span></span>.</p><h3 id="pnl-decomposition">PnL decomposition</h3><p>For one market-maker fill with sign <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>s</mi></mrow><annotation encoding="application/x-tex">s</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal">s</span></span></span></span> (+1 for a buy), price <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>p</mi></mrow><annotation encoding="application/x-tex">p</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.625em;vertical-align:-0.1944em;"></span><span class="mord mathnormal">p</span></span></span></span>, size <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>q</mi></mrow><annotation encoding="application/x-tex">q</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.625em;vertical-align:-0.1944em;"></span><span class="mord mathnormal" style="margin-right:0.0359em;">q</span></span></span></span>, mid <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>m</mi><mn>0</mn></msub></mrow><annotation encoding="application/x-tex">m_0</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.5806em;vertical-align:-0.15em;"></span><span class="mord"><span class="mord mathnormal">m</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3011em;"><span style="top:-2.55em;margin-left:0em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight">0</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span></span></span></span> just before the aggressing order, mid <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>m</mi><mi>h</mi></msub></mrow><annotation encoding="application/x-tex">m_h</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.5806em;vertical-align:-0.15em;"></span><span class="mord"><span class="mord mathnormal">m</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3361em;"><span style="top:-2.55em;margin-left:0em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight">h</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span></span></span></span> one markout horizon later and final mid <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>m</mi><mi>T</mi></msub></mrow><annotation encoding="application/x-tex">m_T</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.5806em;vertical-align:-0.15em;"></span><span class="mord"><span class="mord mathnormal">m</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3283em;"><span style="top:-2.55em;margin-left:0em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight" style="margin-right:0.1389em;">T</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span></span></span></span>,</p><p class="katex-block"><span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><munder><munder><mrow><mi>s</mi><mi>q</mi><mo stretchy="false">(</mo><msub><mi>m</mi><mi>T</mi></msub><mo>−</mo><mi>p</mi><mo stretchy="false">)</mo></mrow><mo stretchy="true">⏟</mo></munder><mtext>fill PnL</mtext></munder><mo>=</mo><munder><munder><mrow><mi>s</mi><mi>q</mi><mo stretchy="false">(</mo><msub><mi>m</mi><mn>0</mn></msub><mo>−</mo><mi>p</mi><mo stretchy="false">)</mo></mrow><mo stretchy="true">⏟</mo></munder><mtext>spread capture</mtext></munder><mo>+</mo><munder><munder><mrow><mi>s</mi><mi>q</mi><mo stretchy="false">(</mo><msub><mi>m</mi><mi>h</mi></msub><mo>−</mo><msub><mi>m</mi><mn>0</mn></msub><mo stretchy="false">)</mo></mrow><mo stretchy="true">⏟</mo></munder><mtext>adverse selection</mtext></munder><mo>+</mo><munder><munder><mrow><mi>s</mi><mi>q</mi><mo stretchy="false">(</mo><msub><mi>m</mi><mi>T</mi></msub><mo>−</mo><msub><mi>m</mi><mi>h</mi></msub><mo stretchy="false">)</mo></mrow><mo stretchy="true">⏟</mo></munder><mtext>inventory carry</mtext></munder><mi mathvariant="normal">.</mi></mrow><annotation encoding="application/x-tex">\underbrace{s q (m_T - p)}_{\text{fill PnL}} = \underbrace{s q (m_0 - p)}_{\text{spread capture}} + \underbrace{s q (m_h - m_0)}_{\text{adverse selection}} + \underbrace{s q (m_T - m_h)}_{\text{inventory carry}}.</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:2.3341em;vertical-align:-1.5841em;"></span><span class="minner munder"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.75em;"><span style="top:-1.4159em;"><span class="pstrut" style="height:3em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord text mtight"><span class="mord mtight">fill PnL</span></span></span></span></span><span style="top:-3em;"><span class="pstrut" style="height:3em;"></span><span class="minner munder"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.75em;"><span class="svg-align" style="top:-2.102em;"><span class="pstrut" style="height:3em;"></span><span class="stretchy" style="height:0.548em;min-width:1.6em;"><span class="brace-left" style="height:0.548em;"><svg xmlns="http://www.w3.org/2000/svg" width="400em" height="0.548em" viewBox="0 0 400000 548" preserveAspectRatio="xMinYMin slice"><path d="M0 6l6-6h17c12.688 0 19.313.3 20 1 4 4 7.313 8.3 10 13 35.313 51.3 80.813 93.8 136.5 127.5 55.688 33.7 117.188 55.8 184.5 66.5.688 0 2 .3 4 1 18.688 2.7 76 4.3 172 5h399450v120H429l-6-1c-124.688-8-235-61.7-331-161C60.687 138.7 32.312 99.3 7 54L0 41V6z"/></svg></span><span class="brace-center" style="height:0.548em;"><svg xmlns="http://www.w3.org/2000/svg" width="400em" height="0.548em" viewBox="0 0 400000 548" preserveAspectRatio="xMidYMin slice"><path d="M199572 214c100.7 8.3 195.3 44 280 108 55.3 42 101.7 93 139 153l9 14c2.7-4 5.7-8.7 9-14 53.3-86.7 123.7-153 211-199 66.7-36 137.3-56.3 212-62h199568v120H200432c-178.3 11.7-311.7 78.3-403 201-6 8-9.7 12-11 12-.7.7-6.7 1-18 1s-17.3-.3-18-1c-1.3 0-5-4-11-12-44.7-59.3-101.3-106.3-170-141s-145.3-54.3-229-60H0V214z"/></svg></span><span class="brace-right" style="height:0.548em;"><svg xmlns="http://www.w3.org/2000/svg" width="400em" height="0.548em" viewBox="0 0 400000 548" preserveAspectRatio="xMaxYMin slice"><path d="M399994 0l6 6v35l-6 11c-56 104-135.3 181.3-238 232-57.3 28.7-117 45-179 50H-300V214h399897c43.3-7 81-15 113-26 100.7-33 179.7-91 237-174 2.7-5 6-9 10-13 .7-1 7.3-1 20-1h17z"/></svg></span></span></span><span style="top:-3em;"><span class="pstrut" style="height:3em;"></span><span class="mord"><span class="mord mathnormal">s</span><span class="mord mathnormal" style="margin-right:0.0359em;">q</span><span class="mopen">(</span><span class="mord"><span class="mord mathnormal">m</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3283em;"><span style="top:-2.55em;margin-left:0em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight" style="margin-right:0.1389em;">T</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">−</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mord mathnormal">p</span><span class="mclose">)</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.898em;"><span></span></span></span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:1.5841em;"><span></span></span></span></span></span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:2.4702em;vertical-align:-1.7202em;"></span><span class="minner munder"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.75em;"><span style="top:-1.4159em;"><span class="pstrut" style="height:3em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord text mtight"><span class="mord mtight">spread capture</span></span></span></span></span><span style="top:-3em;"><span class="pstrut" style="height:3em;"></span><span class="minner munder"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.75em;"><span class="svg-align" style="top:-2.102em;"><span class="pstrut" style="height:3em;"></span><span class="stretchy" style="height:0.548em;min-width:1.6em;"><span class="brace-left" style="height:0.548em;"><svg xmlns="http://www.w3.org/2000/svg" width="400em" height="0.548em" viewBox="0 0 400000 548" preserveAspectRatio="xMinYMin slice"><path d="M0 6l6-6h17c12.688 0 19.313.3 20 1 4 4 7.313 8.3 10 13 35.313 51.3 80.813 93.8 136.5 127.5 55.688 33.7 117.188 55.8 184.5 66.5.688 0 2 .3 4 1 18.688 2.7 76 4.3 172 5h399450v120H429l-6-1c-124.688-8-235-61.7-331-161C60.687 138.7 32.312 99.3 7 54L0 41V6z"/></svg></span><span class="brace-center" style="height:0.548em;"><svg xmlns="http://www.w3.org/2000/svg" width="400em" height="0.548em" viewBox="0 0 400000 548" preserveAspectRatio="xMidYMin slice"><path d="M199572 214c100.7 8.3 195.3 44 280 108 55.3 42 101.7 93 139 153l9 14c2.7-4 5.7-8.7 9-14 53.3-86.7 123.7-153 211-199 66.7-36 137.3-56.3 212-62h199568v120H200432c-178.3 11.7-311.7 78.3-403 201-6 8-9.7 12-11 12-.7.7-6.7 1-18 1s-17.3-.3-18-1c-1.3 0-5-4-11-12-44.7-59.3-101.3-106.3-170-141s-145.3-54.3-229-60H0V214z"/></svg></span><span class="brace-right" style="height:0.548em;"><svg xmlns="http://www.w3.org/2000/svg" width="400em" height="0.548em" viewBox="0 0 400000 548" preserveAspectRatio="xMaxYMin slice"><path d="M399994 0l6 6v35l-6 11c-56 104-135.3 181.3-238 232-57.3 28.7-117 45-179 50H-300V214h399897c43.3-7 81-15 113-26 100.7-33 179.7-91 237-174 2.7-5 6-9 10-13 .7-1 7.3-1 20-1h17z"/></svg></span></span></span><span style="top:-3em;"><span class="pstrut" style="height:3em;"></span><span class="mord"><span class="mord mathnormal">s</span><span class="mord mathnormal" style="margin-right:0.0359em;">q</span><span class="mopen">(</span><span class="mord"><span class="mord mathnormal">m</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3011em;"><span style="top:-2.55em;margin-left:0em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight">0</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">−</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mord mathnormal">p</span><span class="mclose">)</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.898em;"><span></span></span></span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:1.7202em;"><span></span></span></span></span></span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">+</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:2.3341em;vertical-align:-1.5841em;"></span><span class="minner munder"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.75em;"><span style="top:-1.4159em;"><span class="pstrut" style="height:3em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord text mtight"><span class="mord mtight">adverse selection</span></span></span></span></span><span style="top:-3em;"><span class="pstrut" style="height:3em;"></span><span class="minner munder"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.75em;"><span class="svg-align" style="top:-2.102em;"><span class="pstrut" style="height:3em;"></span><span class="stretchy" style="height:0.548em;min-width:1.6em;"><span class="brace-left" style="height:0.548em;"><svg xmlns="http://www.w3.org/2000/svg" width="400em" height="0.548em" viewBox="0 0 400000 548" preserveAspectRatio="xMinYMin slice"><path d="M0 6l6-6h17c12.688 0 19.313.3 20 1 4 4 7.313 8.3 10 13 35.313 51.3 80.813 93.8 136.5 127.5 55.688 33.7 117.188 55.8 184.5 66.5.688 0 2 .3 4 1 18.688 2.7 76 4.3 172 5h399450v120H429l-6-1c-124.688-8-235-61.7-331-161C60.687 138.7 32.312 99.3 7 54L0 41V6z"/></svg></span><span class="brace-center" style="height:0.548em;"><svg xmlns="http://www.w3.org/2000/svg" width="400em" height="0.548em" viewBox="0 0 400000 548" preserveAspectRatio="xMidYMin slice"><path d="M199572 214c100.7 8.3 195.3 44 280 108 55.3 42 101.7 93 139 153l9 14c2.7-4 5.7-8.7 9-14 53.3-86.7 123.7-153 211-199 66.7-36 137.3-56.3 212-62h199568v120H200432c-178.3 11.7-311.7 78.3-403 201-6 8-9.7 12-11 12-.7.7-6.7 1-18 1s-17.3-.3-18-1c-1.3 0-5-4-11-12-44.7-59.3-101.3-106.3-170-141s-145.3-54.3-229-60H0V214z"/></svg></span><span class="brace-right" style="height:0.548em;"><svg xmlns="http://www.w3.org/2000/svg" width="400em" height="0.548em" viewBox="0 0 400000 548" preserveAspectRatio="xMaxYMin slice"><path d="M399994 0l6 6v35l-6 11c-56 104-135.3 181.3-238 232-57.3 28.7-117 45-179 50H-300V214h399897c43.3-7 81-15 113-26 100.7-33 179.7-91 237-174 2.7-5 6-9 10-13 .7-1 7.3-1 20-1h17z"/></svg></span></span></span><span style="top:-3em;"><span class="pstrut" style="height:3em;"></span><span class="mord"><span class="mord mathnormal">s</span><span class="mord mathnormal" style="margin-right:0.0359em;">q</span><span class="mopen">(</span><span class="mord"><span class="mord mathnormal">m</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3361em;"><span style="top:-2.55em;margin-left:0em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight">h</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">−</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mord"><span class="mord mathnormal">m</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3011em;"><span style="top:-2.55em;margin-left:0em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight">0</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span><span class="mclose">)</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.898em;"><span></span></span></span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:1.5841em;"><span></span></span></span></span></span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">+</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:2.4516em;vertical-align:-1.7016em;"></span><span class="minner munder"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.75em;"><span style="top:-1.4345em;"><span class="pstrut" style="height:3em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord text mtight"><span class="mord mtight">inventory carry</span></span></span></span></span><span style="top:-3em;"><span class="pstrut" style="height:3em;"></span><span class="minner munder"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.75em;"><span class="svg-align" style="top:-2.102em;"><span class="pstrut" style="height:3em;"></span><span class="stretchy" style="height:0.548em;min-width:1.6em;"><span class="brace-left" style="height:0.548em;"><svg xmlns="http://www.w3.org/2000/svg" width="400em" height="0.548em" viewBox="0 0 400000 548" preserveAspectRatio="xMinYMin slice"><path d="M0 6l6-6h17c12.688 0 19.313.3 20 1 4 4 7.313 8.3 10 13 35.313 51.3 80.813 93.8 136.5 127.5 55.688 33.7 117.188 55.8 184.5 66.5.688 0 2 .3 4 1 18.688 2.7 76 4.3 172 5h399450v120H429l-6-1c-124.688-8-235-61.7-331-161C60.687 138.7 32.312 99.3 7 54L0 41V6z"/></svg></span><span class="brace-center" style="height:0.548em;"><svg xmlns="http://www.w3.org/2000/svg" width="400em" height="0.548em" viewBox="0 0 400000 548" preserveAspectRatio="xMidYMin slice"><path d="M199572 214c100.7 8.3 195.3 44 280 108 55.3 42 101.7 93 139 153l9 14c2.7-4 5.7-8.7 9-14 53.3-86.7 123.7-153 211-199 66.7-36 137.3-56.3 212-62h199568v120H200432c-178.3 11.7-311.7 78.3-403 201-6 8-9.7 12-11 12-.7.7-6.7 1-18 1s-17.3-.3-18-1c-1.3 0-5-4-11-12-44.7-59.3-101.3-106.3-170-141s-145.3-54.3-229-60H0V214z"/></svg></span><span class="brace-right" style="height:0.548em;"><svg xmlns="http://www.w3.org/2000/svg" width="400em" height="0.548em" viewBox="0 0 400000 548" preserveAspectRatio="xMaxYMin slice"><path d="M399994 0l6 6v35l-6 11c-56 104-135.3 181.3-238 232-57.3 28.7-117 45-179 50H-300V214h399897c43.3-7 81-15 113-26 100.7-33 179.7-91 237-174 2.7-5 6-9 10-13 .7-1 7.3-1 20-1h17z"/></svg></span></span></span><span style="top:-3em;"><span class="pstrut" style="height:3em;"></span><span class="mord"><span class="mord mathnormal">s</span><span class="mord mathnormal" style="margin-right:0.0359em;">q</span><span class="mopen">(</span><span class="mord"><span class="mord mathnormal">m</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3283em;"><span style="top:-2.55em;margin-left:0em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight" style="margin-right:0.1389em;">T</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">−</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mord"><span class="mord mathnormal">m</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3361em;"><span style="top:-2.55em;margin-left:0em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight">h</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span><span class="mclose">)</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.898em;"><span></span></span></span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:1.7016em;"><span></span></span></span></span></span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mord">.</span></span></span></span></span></p><p>It telescopes, so it is an identity. Summed over all fills, the left side is exactly the dealer's cash plus inventory marked at <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>m</mi><mi>T</mi></msub></mrow><annotation encoding="application/x-tex">m_T</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.5806em;vertical-align:-0.15em;"></span><span class="mord"><span class="mord mathnormal">m</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3283em;"><span style="top:-2.55em;margin-left:0em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight" style="margin-right:0.1389em;">T</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span></span></span></span>, since it starts flat with no cash. I use a 10-unit horizon for <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>h</mi></mrow><annotation encoding="application/x-tex">h</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6944em;"></span><span class="mord mathnormal">h</span></span></span></span>. The split is only as meaningful as the choice of <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>h</mi></mrow><annotation encoding="application/x-tex">h</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6944em;"></span><span class="mord mathnormal">h</span></span></span></span>, but the sum is exact, and I test it as an identity rather than a statistical claim.</p><h3 id="stylized-facts">stylized facts</h3><p>Real asset returns have positive excess kurtosis with a tail index near 3, almost no linear autocorrelation beyond very short lags, and autocorrelation in <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="normal">∣</mi><mi>r</mi><mi mathvariant="normal">∣</mi></mrow><annotation encoding="application/x-tex">|r|</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mord">∣</span><span class="mord mathnormal" style="margin-right:0.0278em;">r</span><span class="mord">∣</span></span></span></span> that stays positive for a long time. Kurtosis also shrinks as the return horizon grows. These are the targets I measure against, not ones I tuned toward.</p><h2 id="architecture">architecture</h2><p>Everything runs on one event heap keyed by <code>(time, seq)</code>. Agents act only through a <code>Market</code> object and hear about the world through callbacks.</p><figure data-figure="diagram:market-sim"></figure><p>Data flows one way. The hidden fundamental <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>V</mi></mrow><annotation encoding="application/x-tex">V</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal" style="margin-right:0.2222em;">V</span></span></span></span> is read by exactly two things, the informed traders and the analysis stamps on each trade. The market maker never reads it to make a decision. That separation is what makes &quot;edge against <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>V</mi></mrow><annotation encoding="application/x-tex">V</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal" style="margin-right:0.2222em;">V</span></span></span></span>&quot; a ground-truth check that only a simulator can offer and a real desk cannot.</p><h3 id="the-scheduler">the scheduler</h3><p>The scheduler is twenty lines, and its one design choice is the <code>seq</code> counter.</p><figure class="highlight python"><table><tr><td class="code"><pre><span class="line"><span class="keyword">def</span> <span class="title function_">at</span>(<span class="params">self, t, fn, *args</span>):</span><br><span class="line">    <span class="keyword">if</span> t &lt; <span class="variable language_">self</span>.t:</span><br><span class="line">        <span class="keyword">raise</span> ValueError(<span class="string">f&quot;cannot schedule in the past: <span class="subst">&#123;t&#125;</span> &lt; <span class="subst">&#123;self.t&#125;</span>&quot;</span>)</span><br><span class="line">    heapq.heappush(<span class="variable language_">self</span>._heap, (t, <span class="variable language_">self</span>._seq, fn, args))</span><br><span class="line">    <span class="variable language_">self</span>._seq += <span class="number">1</span></span><br><span class="line"></span><br><span class="line"><span class="keyword">def</span> <span class="title function_">run</span>(<span class="params">self, t_end</span>):</span><br><span class="line">    heap = <span class="variable language_">self</span>._heap</span><br><span class="line">    <span class="keyword">while</span> heap <span class="keyword">and</span> heap[<span class="number">0</span>][<span class="number">0</span>] &lt;= t_end:</span><br><span class="line">        t, _, fn, args = heapq.heappop(heap)</span><br><span class="line">        <span class="variable language_">self</span>.t = t</span><br><span class="line">        <span class="variable language_">self</span>.n_events += <span class="number">1</span></span><br><span class="line">        fn(*args)</span><br><span class="line">    <span class="variable language_">self</span>.t = t_end</span><br></pre></td></tr></table></figure><p><code>seq</code> breaks ties by insertion order, so two events at the same timestamp always run in the order they were scheduled. That gives determinism, and it means the heap never has to compare two functions, which would raise in Python.</p><h3 id="zero-delay-reactions-are-scheduled-not-called">zero-delay reactions are scheduled, not called</h3><p>When the market maker is filled, it does not requote inside the fill callback. It schedules a requote at the current time. The matching loop is still dispatching fills at that moment, and a synchronous requote would re-enter the book halfway through a match. Scheduling at <code>t + 0</code> puts the requote after everything already queued at that time, so the ordering stays well defined, and a <code>_requote_pending</code> flag collapses several fills from one aggressive order into a single requote.</p><h3 id="one-aggregated-agent-per-noise-type">one aggregated agent per noise type</h3><p>Only the total arrival rate matters for the statistics. One noise-taker agent with rate <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mo stretchy="false">(</mo><mn>1</mn><mo>−</mo><mi>α</mi><mo stretchy="false">)</mo><mi>λ</mi></mrow><annotation encoding="application/x-tex">(1 - \alpha)\lambda</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mopen">(</span><span class="mord">1</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">−</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mord mathnormal" style="margin-right:0.0037em;">α</span><span class="mclose">)</span><span class="mord mathnormal">λ</span></span></span></span> is equivalent to many small ones and much cheaper, so there is one of each agent type rather than thousands of objects.</p><h3 id="independent-random-streams">independent random streams</h3><p>Each component gets its own generator, spawned from one <code>SeedSequence(seed)</code>. Adding a random draw inside one agent does not shift any other agent's randomness, so an ablation that changes the market maker leaves the order flow it faces unchanged up to the point where the book differs.</p><h2 id="implementation">implementation</h2><h3 id="the-order-book">the order book</h3><p>Each side of the book is a dict from price to an <code>OrderedDict</code> FIFO queue, a heap of prices, and a cached quantity per level. A global <code>orders</code> index makes cancel O(1).</p><figure class="highlight python"><table><tr><td class="code"><pre><span class="line"><span class="keyword">def</span> <span class="title function_">best</span>(<span class="params">self</span>) -&gt; <span class="built_in">int</span> | <span class="literal">None</span>:</span><br><span class="line">    heap, levels = <span class="variable language_">self</span>.heap, <span class="variable language_">self</span>.levels</span><br><span class="line">    <span class="keyword">while</span> heap:</span><br><span class="line">        price = -<span class="variable language_">self</span>.sign * heap[<span class="number">0</span>]</span><br><span class="line">        <span class="keyword">if</span> price <span class="keyword">in</span> levels:</span><br><span class="line">            <span class="keyword">return</span> price</span><br><span class="line">        heapq.heappop(heap)  <span class="comment"># stale entry for an emptied level</span></span><br><span class="line">    <span class="keyword">return</span> <span class="literal">None</span></span><br></pre></td></tr></table></figure><p>Bids are stored negated so <code>heap[0]</code> is always the best price on either side. Deletion is lazy. A level that empties is removed from <code>levels</code> immediately, and its heap entry is popped the next time <code>best()</code> finds it at the top. Matching always takes the head of the best level and trades at the resting order's price.</p><figure class="highlight python"><table><tr><td class="code"><pre><span class="line">level = opp.levels[best]</span><br><span class="line"><span class="keyword">while</span> qty &gt; <span class="number">0</span> <span class="keyword">and</span> level:</span><br><span class="line">    oid = <span class="built_in">next</span>(<span class="built_in">iter</span>(level))</span><br><span class="line">    maker = level[oid]</span><br><span class="line">    traded = <span class="built_in">min</span>(qty, maker.qty)</span><br><span class="line">    fills.append(Fill(oid, maker.agent_id, agent_id, side, best, traded))</span><br><span class="line">    ...</span><br></pre></td></tr></table></figure><p>That <code>next(iter(level))</code> line is the site of Problem 5 below. I rejected a sorted price list with <code>bisect</code>, because inserting a new level is O(n), and per-price arrays over a fixed band, because <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>V</mi></mrow><annotation encoding="application/x-tex">V</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal" style="margin-right:0.2222em;">V</span></span></span></span> wanders without bound and the band would have to move with it.</p><p><code>check_invariants()</code> asserts that the book is never crossed, no level is empty, every cached level quantity equals the sum of its orders, the <code>orders</code> index and the levels hold the same set, and every live level has a heap entry.</p><h3 id="the-agents">the agents</h3><p>The fundamental is a Gaussian random walk with <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>σ</mi><mo>=</mo><mn>0.3</mn></mrow><annotation encoding="application/x-tex">\sigma = 0.3</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal" style="margin-right:0.0359em;">σ</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.6444em;"></span><span class="mord">0.3</span></span></span></span> ticks per square-root time unit, plus Poisson jumps at rate 0.004 per unit with normal sizes of standard deviation 12 ticks, stepped once per time unit.</p><p>Informed traders arrive at rate <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>α</mi><mi>λ</mi></mrow><annotation encoding="application/x-tex">\alpha\lambda</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6944em;"></span><span class="mord mathnormal" style="margin-right:0.0037em;">α</span><span class="mord mathnormal">λ</span></span></span></span>, see <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>V</mi></mrow><annotation encoding="application/x-tex">V</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal" style="margin-right:0.2222em;">V</span></span></span></span> plus <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>N</mi><mo stretchy="false">(</mo><mn>0</mn><mo separator="true">,</mo><mn>0.5</mn><mo stretchy="false">)</mo></mrow><annotation encoding="application/x-tex">N(0, 0.5)</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mord mathnormal" style="margin-right:0.109em;">N</span><span class="mopen">(</span><span class="mord">0</span><span class="mpunct">,</span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mord">0.5</span><span class="mclose">)</span></span></span></span> noise, and buy or sell one unit only if the signal is outside the quote. Noise liquidity providers post one unit at or behind the best quote and cancel after an exponential lifetime with mean 30. Noise takers arrive at rate <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mo stretchy="false">(</mo><mn>1</mn><mo>−</mo><mi>α</mi><mo stretchy="false">)</mo><mi>λ</mi></mrow><annotation encoding="application/x-tex">(1 - \alpha)\lambda</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mopen">(</span><span class="mord">1</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">−</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mord mathnormal" style="margin-right:0.0037em;">α</span><span class="mclose">)</span><span class="mord mathnormal">λ</span></span></span></span> and are mildly price-elastic, for reasons that are Problem 3.</p><p>The market maker quotes one unit a side from the formulas above, is post-only, and keeps a quote whose price has not changed so it keeps queue priority. Its quoting is short.</p><figure class="highlight python"><table><tr><td class="code"><pre><span class="line"><span class="keyword">def</span> <span class="title function_">quotes</span>(<span class="params">self</span>):</span><br><span class="line">    r = <span class="variable language_">self</span>.fair - <span class="variable language_">self</span>.inventory * <span class="variable language_">self</span>.skew_per_unit</span><br><span class="line">    h = <span class="variable language_">self</span>.half_spread()</span><br><span class="line">    bid, ask = math.floor(r - h), math.ceil(r + h)</span><br><span class="line">    <span class="keyword">if</span> ask &lt;= bid:</span><br><span class="line">        ask = bid + <span class="number">1</span></span><br><span class="line">    <span class="keyword">return</span> bid, ask</span><br></pre></td></tr></table></figure><p>It also schedules its own markout check 10 units after each fill. That check is what the adaptive variant learns from.</p><figure class="highlight python"><table><tr><td class="code"><pre><span class="line"><span class="keyword">def</span> <span class="title function_">_markout</span>(<span class="params">self, idx</span>):</span><br><span class="line">    row = <span class="variable language_">self</span>.fills[idx]</span><br><span class="line">    row[<span class="number">5</span>] = <span class="variable language_">self</span>.market.mid()</span><br><span class="line">    as_per_unit = -row[<span class="number">1</span>] * (row[<span class="number">5</span>] - row[<span class="number">4</span>])   <span class="comment"># how far the mid moved against us</span></span><br><span class="line">    <span class="variable language_">self</span>.as_cost += <span class="variable language_">self</span>.cfg.mm_as_ewma * (as_per_unit - <span class="variable language_">self</span>.as_cost)</span><br></pre></td></tr></table></figure><p>With <code>mm_adaptive</code> set, <code>half_spread()</code> adds <code>max(0, as_cost)</code> to the Avellaneda-Stoikov half-spread. This is the Glosten-Milgrom response done empirically. The maker never learns <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>α</mi></mrow><annotation encoding="application/x-tex">\alpha</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal" style="margin-right:0.0037em;">α</span></span></span></span> for its spread. It measures what trading costs it and charges that.</p><h3 id="three-market-maker-variants">three market-maker variants</h3><p>The sweep runs three variants.</p><ul><li><code>fixed</code> uses the Avellaneda-Stoikov quotes, with the fair-value step scaled by <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>α</mi></mrow><annotation encoding="application/x-tex">\alpha</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal" style="margin-right:0.0037em;">α</span></span></span></span> as in Glosten-Milgrom.</li><li><code>adaptive</code> is the same, plus its half-spread grows by the EWMA of its own measured 10-unit markout loss.</li><li><code>naive</code> uses a fixed learning rate of 0.3 whatever <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>α</mi></mrow><annotation encoding="application/x-tex">\alpha</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal" style="margin-right:0.0037em;">α</span></span></span></span> is.</li></ul><p>In the Glosten-Milgrom mode, which is the default, the maker is told <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>α</mi></mrow><annotation encoding="application/x-tex">\alpha</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal" style="margin-right:0.0037em;">α</span></span></span></span>, as the dealer is in the model. It never reads <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>V</mi></mrow><annotation encoding="application/x-tex">V</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal" style="margin-right:0.2222em;">V</span></span></span></span>, but it knows the informed share exactly, and the <code>fixed</code> and <code>adaptive</code> results below both assume a dealer with that knowledge. That is a reasonable stand-in for Glosten-Milgrom and a real simplification, and it is one reason the adaptive variant, whose spread comes from measured markouts rather than from <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>α</mi></mrow><annotation encoding="application/x-tex">\alpha</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal" style="margin-right:0.0037em;">α</span></span></span></span>, is the more interesting one.</p><h3 id="tests">tests</h3><p>There are 32 tests. The one I trust most is a fuzz test that runs 5 seeds of 3,000 random operations, 15,000 in total, against a naive reference book that keeps a flat list and scans it on every call, and compares every fill by maker id, price and quantity. The others check invariants after every simulated time unit, bit-identical determinism for a fixed seed, conservation of cash, inventory and edge across all agents, the exact PnL identity, the Avellaneda-Stoikov closed form, the variance of the fundamental's increments, and the estimators against known distributions.</p><h2 id="problems">problems</h2><p>Most of the time on this project went into the market maker's beliefs, not the book. Three of the eight problems below were the market maker being wrong about the world in a way that looked like a market phenomenon.</p><h3 id="1-the-market-maker-crashed-the-price-by-learning-from-its-own-quotes">1. the market maker crashed the price by learning from its own quotes</h3><p>My first fair-value update was the obvious one, move fair value toward each trade price.</p><figure class="highlight python"><table><tr><td class="code"><pre><span class="line"><span class="variable language_">self</span>.fair += <span class="variable language_">self</span>.learning * (tr.price - <span class="variable language_">self</span>.fair)</span><br></pre></td></tr></table></figure><p>With informed traders on and jumps off, the mid still made moves of up to 98 ticks over 100 time units while <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>V</mi></mrow><annotation encoding="application/x-tex">V</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal" style="margin-right:0.2222em;">V</span></span></span></span> barely moved. It showed up first as a kurtosis number that made no sense for a model with no jumps, and then as a price path with cliffs in it. I traced the trades around one crash. The market maker was long, so the inventory skew had shaded its bid down. Noise sellers hit that low bid. The update read the low trade price as news and lowered fair value by several ticks per trade. That lowered the bid further, the maker bought more, and the loop ran away.</p><p>The bug is that the market maker was learning from its own quotes. Its trade prices contain its inventory skew, and feeding them back into fair value turns the skew into a belief. The fix was to learn from trade direction only.</p><figure class="highlight python"><table><tr><td class="code"><pre><span class="line"><span class="keyword">if</span> <span class="variable language_">self</span>.cfg.mm_learning_signal == <span class="string">&quot;direction&quot;</span>:</span><br><span class="line">    <span class="variable language_">self</span>.fair += <span class="variable language_">self</span>.learning * tr.taker_side * <span class="variable language_">self</span>.base_half_spread</span><br><span class="line"><span class="keyword">else</span>:  <span class="comment"># &quot;price&quot;: the original, buggy update</span></span><br><span class="line">    <span class="variable language_">self</span>.fair += <span class="variable language_">self</span>.learning * (tr.price - <span class="variable language_">self</span>.fair)</span><br></pre></td></tr></table></figure><p>I kept the old rule behind <code>mm_learning_signal=&quot;price&quot;</code> so the bug stays reproducible, and the stylized-facts script runs it as an ablation. In <code>results/stylized/summary.json</code> the <code>price_learning_bug</code> run has a largest 100-unit move of 98.0 ticks and a 100-unit excess kurtosis of 19.21. The fixed model with no jumps gives 11.5 ticks and 0.39. One caveat on that comparison. The bug run also has inelastic noise and no inventory learning, the configuration the bug was found in, so it differs from the no-jumps run in three settings, not one.</p><figure data-figure="chart:projects/market-simulator/market-simulator-kurtosis"></figure><h3 id="2-adverse-selection-fell-as-informed-flow-rose">2. adverse selection fell as informed flow rose</h3><p>In my first sweep the market maker used a fixed learning rate and lost the most money at <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>α</mi><mo>=</mo><mn>0</mn></mrow><annotation encoding="application/x-tex">\alpha = 0</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal" style="margin-right:0.0037em;">α</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.6444em;"></span><span class="mord">0</span></span></span></span>, which is backwards. With no informed traders there should be nothing to be adversely selected by.</p><p>It was chasing noise. With nothing pulling the price back, every uninformed trade permanently moved its quotes toward the side that had just traded against it, so the mid moved against its fills and the markouts registered that as adverse selection. More informed traders helped, because they pushed the price back toward <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>V</mi></mrow><annotation encoding="application/x-tex">V</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal" style="margin-right:0.2222em;">V</span></span></span></span>.</p><p>This is exactly the Glosten-Milgrom point that belief updates should scale with <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>α</mi></mrow><annotation encoding="application/x-tex">\alpha</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal" style="margin-right:0.0037em;">α</span></span></span></span>. A trade from a population that is 95% noise should barely move the dealer. The <code>gm</code> learning mode sets the step to <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>w</mi><mo>=</mo><mi>α</mi></mrow><annotation encoding="application/x-tex">w = \alpha</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal" style="margin-right:0.0269em;">w</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal" style="margin-right:0.0037em;">α</span></span></span></span>. The naive variant is still in the sweep to show the effect. At <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>α</mi><mo>=</mo><mn>0</mn></mrow><annotation encoding="application/x-tex">\alpha = 0</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal" style="margin-right:0.0037em;">α</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.6444em;"></span><span class="mord">0</span></span></span></span> its adverse selection per fill is 1.005 ticks, against 0.605 for the Glosten-Milgrom learner (<code>results/sweep/aggregate.csv</code>). The two variants meet exactly at <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>α</mi><mo>=</mo><mn>0.3</mn></mrow><annotation encoding="application/x-tex">\alpha = 0.3</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal" style="margin-right:0.0037em;">α</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.6444em;"></span><span class="mord">0.3</span></span></span></span>, where both step sizes equal 0.3, and every column of the two rows in <code>aggregate.csv</code> is identical there. That coincidence doubles as a determinism check.</p><h3 id="3-inventory-drifted-to-the-limit">3. inventory drifted to the limit</h3><p>With price-insensitive noise traders, inventory drifted to the ±50 limit and stayed near it. At <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>α</mi><mo>=</mo><mn>0</mn></mrow><annotation encoding="application/x-tex">\alpha = 0</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal" style="margin-right:0.0037em;">α</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.6444em;"></span><span class="mord">0</span></span></span></span> mean absolute inventory was 29.4, and the maker sat at the limit 2.9% of the time (<code>results/inventory_ablation.csv</code>, the row with elastic noise and inventory learning both off).</p><p>The reason is that the Avellaneda-Stoikov skew only works if fills respond to price. Its fill intensity <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>A</mi><msup><mi>e</mi><mrow><mo>−</mo><mi>k</mi><mi>δ</mi></mrow></msup></mrow><annotation encoding="application/x-tex">A e^{-k\delta}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.8491em;"></span><span class="mord mathnormal">A</span><span class="mord"><span class="mord mathnormal">e</span><span class="msupsub"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.8491em;"><span style="top:-3.063em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mtight">−</span><span class="mord mathnormal mtight" style="margin-right:0.0315em;">k</span><span class="mord mathnormal mtight" style="margin-right:0.0379em;">δ</span></span></span></span></span></span></span></span></span></span></span></span> is exactly that assumption. If noise flow ignores where the quotes are, shading them changes nothing, and inventory is a random walk with a wall at 50.</p><p>The first fix was to make noise takers mildly price-elastic, so they trade with probability <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msup><mi>e</mi><mrow><mo>−</mo><mn>0.1</mn><mi>d</mi></mrow></msup></mrow><annotation encoding="application/x-tex">e^{-0.1 d}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.8491em;"></span><span class="mord"><span class="mord mathnormal">e</span><span class="msupsub"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.8491em;"><span style="top:-3.063em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mtight">−</span><span class="mord mtight">0.1</span><span class="mord mathnormal mtight">d</span></span></span></span></span></span></span></span></span></span></span></span>, where <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>d</mi></mrow><annotation encoding="application/x-tex">d</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6944em;"></span><span class="mord mathnormal">d</span></span></span></span> is how far the quote they would hit is beyond a slow EWMA of the mid.</p><figure class="highlight python"><table><tr><td class="code"><pre><span class="line"><span class="keyword">if</span> <span class="variable language_">self</span>.k &gt; <span class="number">0</span>:</span><br><span class="line">    quote = book.best_ask() <span class="keyword">if</span> side == BUY <span class="keyword">else</span> book.best_bid()</span><br><span class="line">    <span class="keyword">if</span> quote <span class="keyword">is</span> <span class="keyword">not</span> <span class="literal">None</span>:</span><br><span class="line">        d = (quote - ref) <span class="keyword">if</span> side == BUY <span class="keyword">else</span> (ref - quote)</span><br><span class="line">        go = <span class="variable language_">self</span>.rng.random() &lt; math.exp(-<span class="variable language_">self</span>.k * <span class="built_in">max</span>(d, <span class="number">0.0</span>))</span><br></pre></td></tr></table></figure><p>That fixed <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>α</mi><mo>=</mo><mn>0</mn></mrow><annotation encoding="application/x-tex">\alpha = 0</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal" style="margin-right:0.0037em;">α</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.6444em;"></span><span class="mord">0</span></span></span></span>, where mean absolute inventory fell to 7.92. With informed flow it did not. Inventory was still 25.5 at <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>α</mi><mo>=</mo><mn>0.2</mn></mrow><annotation encoding="application/x-tex">\alpha = 0.2</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal" style="margin-right:0.0037em;">α</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.6444em;"></span><span class="mord">0.2</span></span></span></span> and 27.3 at <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>α</mi><mo>=</mo><mn>0.4</mn></mrow><annotation encoding="application/x-tex">\alpha = 0.4</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal" style="margin-right:0.0037em;">α</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.6444em;"></span><span class="mord">0.4</span></span></span></span>. The reason took a while to see. Informed traders trade whenever <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>V</mi></mrow><annotation encoding="application/x-tex">V</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal" style="margin-right:0.2222em;">V</span></span></span></span> is outside the quote, so they effectively pin the reservation price <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>r</mi><mo>=</mo><mtext>fair</mtext><mo>−</mo><mi>q</mi><mo>⋅</mo><mtext>skew</mtext></mrow><annotation encoding="application/x-tex">r = \text{fair} - q \cdot \text{skew}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal" style="margin-right:0.0278em;">r</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.7778em;vertical-align:-0.0833em;"></span><span class="mord text"><span class="mord">fair</span></span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">−</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:0.6389em;vertical-align:-0.1944em;"></span><span class="mord mathnormal" style="margin-right:0.0359em;">q</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">⋅</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:0.6944em;"></span><span class="mord text"><span class="mord">skew</span></span></span></span></span> to <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>V</mi></mrow><annotation encoding="application/x-tex">V</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal" style="margin-right:0.2222em;">V</span></span></span></span>. Whatever gap exists between the maker's fair value and <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>V</mi></mrow><annotation encoding="application/x-tex">V</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal" style="margin-right:0.2222em;">V</span></span></span></span> then has to be absorbed by <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>q</mi></mrow><annotation encoding="application/x-tex">q</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.625em;vertical-align:-0.1944em;"></span><span class="mord mathnormal" style="margin-right:0.0359em;">q</span></span></span></span>. The maker's belief is wrong, and the market expresses that as inventory.</p><p>The second fix treats persistent inventory as information. Every requote tick, fair value moves by <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mo>−</mo><mn>0.02</mn><mtext> </mtext><mi>q</mi><mo>⋅</mo><mtext>skew</mtext></mrow><annotation encoding="application/x-tex">-0.02\,q \cdot \text{skew}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.8389em;vertical-align:-0.1944em;"></span><span class="mord">−</span><span class="mord">0.02</span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mord mathnormal" style="margin-right:0.0359em;">q</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">⋅</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:0.6944em;"></span><span class="mord text"><span class="mord">skew</span></span></span></span></span>. A position that will not go away means net order flow has been one-sided, and that is a signal. With both fixes, mean absolute inventory is 1.16, 1.44 and 1.62 at <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>α</mi></mrow><annotation encoding="application/x-tex">\alpha</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal" style="margin-right:0.0037em;">α</span></span></span></span> = 0, 0.2 and 0.4, and PnL per fill barely moves (2.21, 1.81 and 1.39, against 2.50, 1.74 and 1.29 with elastic noise alone).</p><figure data-figure="chart:projects/market-simulator/market-simulator-inventory"></figure><p>I do not love this fix. It works, but it is a patch on a belief model that is too simple, and I come back to it under what I would change.</p><h3 id="4-informed-traders-sometimes-lost-money">4. informed traders sometimes lost money</h3><p>A test asserting that informed traders make money failed on some seeds. I had marked informed PnL at the final <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>V</mi></mrow><annotation encoding="application/x-tex">V</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal" style="margin-right:0.2222em;">V</span></span></span></span>. Informed traders never unwind, so by the end of a run they hold hundreds of units, and the mark is dominated by how <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>V</mi></mrow><annotation encoding="application/x-tex">V</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal" style="margin-right:0.2222em;">V</span></span></span></span> drifted after they traded. In the baseline stylized run, informed PnL marked at the final <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>V</mi></mrow><annotation encoding="application/x-tex">V</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal" style="margin-right:0.2222em;">V</span></span></span></span> is negative, while their edge at trade time is positive (<code>results/stylized/summary.json</code>).</p><p>The fix was to record each agent's edge at trade time, <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mo>∑</mo><mi>s</mi><mi>q</mi><mo stretchy="false">(</mo><msub><mi>V</mi><mi>t</mi></msub><mo>−</mo><mi>p</mi><mo stretchy="false">)</mo></mrow><annotation encoding="application/x-tex">\sum s q (V_t - p)</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mop op-symbol small-op" style="position:relative;top:0em;">∑</span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mord mathnormal">s</span><span class="mord mathnormal" style="margin-right:0.0359em;">q</span><span class="mopen">(</span><span class="mord"><span class="mord mathnormal" style="margin-right:0.2222em;">V</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.2806em;"><span style="top:-2.55em;margin-left:-0.2222em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight">t</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">−</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mord mathnormal">p</span><span class="mclose">)</span></span></span></span>. Informed edge is then positive and noise edge negative on every run, and across all agents edge sums to zero, which is now a test. This is the only place a simulator can offer something a desk cannot, a markout against the true value.</p><h3 id="5-the-book-was-quadratic-in-level-depth">5. the book was quadratic in level depth</h3><p>The benchmark workload builds deep levels, with 31,603 resting orders left at the end, clustered near one price (<code>results/bench.json</code>), and throughput was poor. There were two culprits.</p><p>The first was ordinary. Matching iterated <code>for oid in list(level)</code>, which copies the whole level on every aggressive order.</p><p>The second was subtler. After I switched to reading the head with <code>next(iter(level))</code> on a plain <code>dict</code>, draining a level from the front was still slow. CPython dicts keep insertion order in a dense entries array, and deleting an entry leaves a dummy slot behind rather than compacting. Iteration from the start has to skip every dummy slot, and a FIFO queue deletes from exactly the start. So each head lookup costs time proportional to the number of orders already removed, and draining a level costs time quadratic in its depth. <code>OrderedDict</code> keeps a linked list and does not have this problem.</p><p>The microbenchmark in <code>results/level_container_microbench.txt</code> shows it plainly.</p><table><thead><tr><th>Container</th><th>Drain 10,000 head-first</th><th>Drain 40,000 head-first</th></tr></thead><tbody><tr><td><code>dict</code></td><td>0.112 s</td><td>3.554 s</td></tr><tr><td><code>OrderedDict</code></td><td>0.002 s</td><td>0.006 s</td></tr></tbody></table><p>Four times the size costs the plain dict about 32 times the time. Levels are OrderedDicts now. The devlog records that the full sweep was bit-identical before and after the change, zero differing cells in <code>aggregate.csv</code>, which is a good regression check that FIFO semantics did not move. That comparison was not saved to <code>results/</code>, so it rests on the log.</p><h3 id="6-there-are-two-spreads">6. there are two spreads</h3><p>The book spread narrowed as <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>α</mi></mrow><annotation encoding="application/x-tex">\alpha</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal" style="margin-right:0.0037em;">α</span></span></span></span> grew, from 5.48 to 4.31 ticks in the fixed market maker's runs, even though the Avellaneda-Stoikov formula does not depend on <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>α</mi></mrow><annotation encoding="application/x-tex">\alpha</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal" style="margin-right:0.0037em;">α</span></span></span></span> at all. For a while I read that as a result.</p><p>It is stale orders. When informed flow moves the market maker's fair value, liquidity-provider orders placed against the old quotes can end up inside the new ones, and they set the inside spread until they are hit or cancel. I added the market maker's own quoted spread to the recorder, and the sweep reports both. The fixed maker's own spread stays flat at 5.64 to 5.69.</p><h3 id="7-tick-discreteness-inflates-short-horizon-kurtosis">7. tick discreteness inflates short-horizon kurtosis</h3><p>On a one-tick grid with small per-unit moves, many one-unit mid changes are exactly zero, and a distribution with a spike at zero has high kurtosis for a trivial reason. That is why I read fat tails at the 10-unit horizon and report 1, 10 and 100 units side by side.</p><h3 id="8-the-benchmark-machine-was-shared">8. the benchmark machine was shared</h3><p>Fourteen other projects were building on the same 14-core machine, and <code>results/bench.json</code> records a one-minute load average of 291 at the start of the benchmark and 320 at the end. Wall time for the same 216-run sweep ranged from 61 to 226 s across reruns. The run that wrote <code>results/sweep/</code> recorded 61.2 s in <code>meta.json</code>, and <code>results/sweep_stdout.txt</code> holds the 225.7 s output of an earlier run with identical numbers. Under that load no timing in the committed files is a measurement of the code, which is why the throughput section below does not quote them as performance.</p><h2 id="experiments">experiments</h2><p>Unless stated otherwise, every run uses the <code>SimConfig()</code> defaults. Taker arrivals total 1 per time unit, liquidity providers arrive at 0.5, the book is sampled every time unit, and returns are taken over 10 units.</p><h3 id="informed-fraction-sweep">informed-fraction sweep</h3><p><code>experiments/informed_sweep.py</code> crosses 9 values of <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>α</mi></mrow><annotation encoding="application/x-tex">\alpha</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal" style="margin-right:0.0037em;">α</span></span></span></span> from 0 to 0.6 with the 3 market-maker variants and 8 seeds, each run 20,000 time units long, for 216 runs. It writes per-run rows to <code>results/sweep/runs.csv</code>, means and standard errors to <code>results/sweep/aggregate.csv</code>, and the configuration to <code>results/sweep/meta.json</code>.</p><h3 id="stylized-facts-2">stylized facts</h3><p><code>experiments/stylized_facts.py</code> runs one 200,000-unit baseline at <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>α</mi><mo>=</mo><mn>0.4</mn></mrow><annotation encoding="application/x-tex">\alpha = 0.4</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal" style="margin-right:0.0037em;">α</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.6444em;"></span><span class="mord">0.4</span></span></span></span> with the fixed market maker, plus three ablations on the same seed. They are no jumps, no informed traders, and the original buggy market maker. Each gives 20,000 ten-unit returns. Outputs are <code>results/stylized/summary.json</code>, <code>acf.csv</code> and <code>jump_event_study.csv</code>.</p><h3 id="inventory-ablation">inventory ablation</h3><p><code>experiments/inventory_ablation.py</code> crosses noise elasticity in {0, 0.1} with inventory learning in {0, 0.02} and <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>α</mi></mrow><annotation encoding="application/x-tex">\alpha</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal" style="margin-right:0.0037em;">α</span></span></span></span> in {0, 0.2, 0.4}, one seed per cell, and writes <code>results/inventory_ablation.csv</code>.</p><h3 id="benchmark">benchmark</h3><p><code>experiments/bench.py</code> times five repeats of 500,000 pre-generated book operations (60% limit, 15% market, 25% cancel, sizes 1 to 4) and three 20,000-unit simulations at <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>α</mi><mo>=</mo><mn>0.2</mn></mrow><annotation encoding="application/x-tex">\alpha = 0.2</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal" style="margin-right:0.0037em;">α</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.6444em;"></span><span class="mord">0.2</span></span></span></span>, single-threaded. It writes <code>results/bench.json</code>.</p><h2 id="results">results</h2><h3 id="spreads-and-adverse-selection-against-informed-flow">spreads and adverse selection against informed flow</h3><p>All numbers in this subsection are from <code>results/sweep/aggregate.csv</code>, means over 8 seeds, in ticks per unit filled.</p><table><thead><tr><th><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>α</mi></mrow><annotation encoding="application/x-tex">\alpha</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal" style="margin-right:0.0037em;">α</span></span></span></span></th><th>fixed PnL per fill</th><th>fixed capture</th><th>fixed adverse sel.</th><th>adaptive own spread</th><th>adaptive PnL per fill</th><th>informed edge per informed trade</th></tr></thead><tbody><tr><td>0.00</td><td>2.210</td><td>2.758</td><td>-0.605</td><td>6.85</td><td>2.803</td><td>n/a</td></tr><tr><td>0.05</td><td>2.121</td><td>2.730</td><td>-0.680</td><td>7.00</td><td>2.789</td><td>25.05</td></tr><tr><td>0.20</td><td>1.814</td><td>2.649</td><td>-0.873</td><td>7.39</td><td>2.676</td><td>6.47</td></tr><tr><td>0.40</td><td>1.399</td><td>2.478</td><td>-1.089</td><td>7.80</td><td>2.449</td><td>2.93</td></tr><tr><td>0.60</td><td>0.881</td><td>2.157</td><td>-1.266</td><td>8.16</td><td>2.003</td><td>1.85</td></tr></tbody></table><p>The Glosten-Milgrom direction shows up cleanly. Adverse selection per fill for the fixed maker roughly doubles, from 0.605 to 1.266. Its PnL per fill falls by 60%, from 2.210 to 0.881. Spread capture falls too, from 2.758 to 2.157, probably because more of its fills come when the price is moving through its quotes. Standard errors are 0.025 or less on PnL per fill and 0.023 or less on the adaptive spread, so none of these differences is noise.</p><figure data-figure="chart:projects/market-simulator/market-simulator-spreads"></figure><p>The adaptive maker, which widens by its measured markout loss, quotes a spread that rises monotonically from 6.85 to 8.16. Note that it is already about 1.2 ticks wider than the fixed maker at <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>α</mi><mo>=</mo><mn>0</mn></mrow><annotation encoding="application/x-tex">\alpha = 0</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal" style="margin-right:0.0037em;">α</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.6444em;"></span><span class="mord">0</span></span></span></span>, where there is nobody informed to protect against. The 10-unit markout against the mid also captures the maker's own footprint. After a fill, its inventory skew and inventory learning move its own quotes, the mid moves with them, and that registers as adverse selection of about 0.6 ticks per fill even with no informed traders. The adaptive rule charges for that too. A markout against the mid is what a real desk can measure, which is why I use it, but it is not a pure measure of information.</p><figure data-figure="chart:projects/market-simulator/market-simulator-pnl-per-fill"></figure><h3 id="where-the-naive-reading-of-glosten-milgrom-breaks">where the naive reading of Glosten-Milgrom breaks</h3><p>One result goes against the static model. Each informed trade earns much less as <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>α</mi></mrow><annotation encoding="application/x-tex">\alpha</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal" style="margin-right:0.0037em;">α</span></span></span></span> rises, 25.05 ticks at <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>α</mi><mo>=</mo><mn>0.05</mn></mrow><annotation encoding="application/x-tex">\alpha = 0.05</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal" style="margin-right:0.0037em;">α</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.6444em;"></span><span class="mord">0.05</span></span></span></span> down to 1.85 at 0.6. More informed flow keeps the price closer to <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>V</mi></mrow><annotation encoding="application/x-tex">V</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal" style="margin-right:0.2222em;">V</span></span></span></span>. Mean <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="normal">∣</mi><mtext>mid</mtext><mo>−</mo><mi>V</mi><mi mathvariant="normal">∣</mi></mrow><annotation encoding="application/x-tex">|\text{mid} - V|</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mord">∣</span><span class="mord text"><span class="mord">mid</span></span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">−</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mord mathnormal" style="margin-right:0.2222em;">V</span><span class="mord">∣</span></span></span></span> is 24.85 ticks at 0.05 and 1.63 at 0.6. So the loss per taker trade to informed traders, the break-even half-spread a zero-profit dealer would need, falls from 1.45 to 0.57.</p><p>The static formula <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>h</mi><mo>≈</mo><mi>α</mi><mi>D</mi></mrow><annotation encoding="application/x-tex">h \approx \alpha D</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6944em;"></span><span class="mord mathnormal">h</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">≈</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal" style="margin-right:0.0037em;">α</span><span class="mord mathnormal" style="margin-right:0.0278em;">D</span></span></span></span> treats the information gap <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>D</mi></mrow><annotation encoding="application/x-tex">D</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal" style="margin-right:0.0278em;">D</span></span></span></span> as fixed. In a dynamic market <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>D</mi></mrow><annotation encoding="application/x-tex">D</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal" style="margin-right:0.0278em;">D</span></span></span></span> shrinks as <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>α</mi></mrow><annotation encoding="application/x-tex">\alpha</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal" style="margin-right:0.0037em;">α</span></span></span></span> grows, because informed trading is itself what closes the gap. The dealer's measured adverse selection still rises, because the informed share of trades rises faster than the gap shrinks, from 0.058 to 0.310 for the fixed maker. I did not expect this going in, and it is the result I find most interesting.</p><h3 id="stylized-facts-3">stylized facts</h3><p>From <code>results/stylized/summary.json</code> and <code>results/stylized/acf.csv</code>.</p><table><thead><tr><th>run</th><th>excess kurtosis at 1 / 10 / 100 units</th><th>ACF of r, lag 1</th><th>ACF of abs r, lags 1 / 5 / 20</th></tr></thead><tbody><tr><td>baseline (<span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>α</mi></mrow><annotation encoding="application/x-tex">\alpha</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal" style="margin-right:0.0037em;">α</span></span></span></span> 0.4, jumps)</td><td>3.88 / 2.61 / 2.78</td><td>-0.135</td><td>0.192 / 0.023 / -0.006</td></tr><tr><td>no jumps</td><td>3.51 / 1.09 / 0.39</td><td>-0.321</td><td>0.171 / -0.003 / 0.003</td></tr><tr><td>no informed traders</td><td>2.28 / 0.78 / 0.38</td><td>-0.211</td><td>0.096 / -0.007 / 0.003</td></tr></tbody></table><h3 id="fat-tails-mostly-imported">fat tails, mostly imported</h3><p>At the 10-unit horizon the baseline has excess kurtosis 2.61 and a Hill tail index of 4.63, against roughly 3 for real equities. So the tails are fat but thinner than real ones. At 100 units kurtosis stays at 2.78 with jumps and drops to 0.39 without them. The long-horizon tails come from the jumps I put into the fundamental. They are not emergent.</p><p>Informed flow does thicken tails at the 10-unit horizon. Across the sweep, mean 10-unit kurtosis for the fixed maker is 1.06 at <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>α</mi><mo>=</mo><mn>0.2</mn></mrow><annotation encoding="application/x-tex">\alpha = 0.2</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal" style="margin-right:0.0037em;">α</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.6444em;"></span><span class="mord">0.2</span></span></span></span> and 7.65 at <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>α</mi><mo>=</mo><mn>0.6</mn></mrow><annotation encoding="application/x-tex">\alpha = 0.6</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal" style="margin-right:0.0037em;">α</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.6444em;"></span><span class="mord">0.6</span></span></span></span> (<code>results/sweep/aggregate.csv</code>). With many informed traders the price follows <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>V</mi></mrow><annotation encoding="application/x-tex">V</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal" style="margin-right:0.2222em;">V</span></span></span></span> closely, jumps included, so the price inherits the jumps almost undiluted.</p><h3 id="volatility-clustering-short-lived">volatility clustering, short-lived</h3><p>The ACF of <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="normal">∣</mi><mi>r</mi><mi mathvariant="normal">∣</mi></mrow><annotation encoding="application/x-tex">|r|</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mord">∣</span><span class="mord mathnormal" style="margin-right:0.0278em;">r</span><span class="mord">∣</span></span></span></span> is 0.192 at lag 1 and falls inside the iid 95% band of ±0.014 by about lag 5. Real data stays positive for hundreds of lags. The short burst comes from price discovery after a jump and from the market maker's quote dynamics. Nothing in the model produces long memory, and I would rather say so than present the lag-1 value as a match.</p><h3 id="negative-lag-1-autocorrelation">negative lag-1 autocorrelation</h3><p>Linear autocorrelation at lag 1 is -0.135 in the baseline. It comes from the temporary impact of noise trades reverting. Fair value and inventory skew both move after a trade and then decay. Real data shows the same microstructure effect at short horizons.</p><h3 id="spread-around-jumps">spread around jumps</h3><p>From <code>results/stylized/jump_event_study.csv</code>, averaging over the 836 jumps with a full window around them, mean <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="normal">∣</mi><mtext>mid</mtext><mo>−</mo><mi>V</mi><mi mathvariant="normal">∣</mi></mrow><annotation encoding="application/x-tex">|\text{mid} - V|</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mord">∣</span><span class="mord text"><span class="mord">mid</span></span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">−</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mord mathnormal" style="margin-right:0.2222em;">V</span><span class="mord">∣</span></span></span></span> goes from an average of 2.68 ticks over the 20 samples before a jump to 10.65 at the jump sample and 10.28 one unit after, then 4.25 fifty units later, and is back to 2.69 after 150 units. The book spread barely moves. It averages about 4.84 before a jump and peaks at 4.99 about 27 units after.</p><p>The fixed Avellaneda-Stoikov maker has no channel to widen on news. Its spread depends only on inventory, which inventory learning keeps small. A maker that widened on order-flow imbalance would show a real response, and that belongs to the RL milestone.</p><h3 id="market-maker-pnl-decomposition">market-maker PnL decomposition</h3><p>For the baseline run, 71,520 fills over 200,000 units, total PnL was 97,116 ticks. Spread capture contributed 177,634, adverse selection -80,589 and inventory carry 71 (<code>results/stylized/summary.json</code>). The three terms sum to the total exactly. Adverse selection took 45% of gross spread capture. Carry is negligible because mean absolute inventory is 1.7 units. In the run with no informed traders, adverse selection per fill falls from 1.127 to 0.611, and what remains is the maker's own footprint on the mid described above.</p><h3 id="throughput-not-yet-measured">throughput, not yet measured</h3><p>I do not have a throughput number I would stand behind. <code>results/bench.json</code> records 33,200 book operations per second and 13,854 simulation events per second, but it was taken single-threaded at a one-minute load average near 300 on a 14-core machine, and those figures say more about the machine than about the code. The same file set shows how unstable they are. The four 200,000-unit stylized runs in <code>results/stylized/summary.json</code>, at 1.18 to 1.26 million events each, recorded 31,271 to 33,389 events/s, more than twice the benchmark rate for the same code.</p><p>The independent verification gives a sense of how far off the loaded numbers are. At a load average near 6, a small smoke run of the benchmark put one 20,000-unit simulation at about 376,000 events/s, and each 200,000-unit stylized run took about 3 s against 35 to 64 s in the committed files. That smoke run was a single short simulation and its tiny book workload does not build the deep levels of the real one, so it is not a benchmark either. The honest status is that throughput has not been measured on a quiet machine, and the committed figures are very conservative lower bounds.</p><h2 id="what-i-would-change">what i would change</h2><h3 id="a-real-belief-model-for-the-market-maker">a real belief model for the market maker</h3><p>The maker should hold a Bayesian belief over <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>V</mi></mrow><annotation encoding="application/x-tex">V</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal" style="margin-right:0.2222em;">V</span></span></span></span> instead of a Kyle-style step plus an inventory-decay term. Both fixes in Problem 3 work, but they are patches. A posterior needs a model of the informed traders' signal, which is heavier than v0 needed and is the right answer. It would also let the maker estimate <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>α</mi></mrow><annotation encoding="application/x-tex">\alpha</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal" style="margin-right:0.0037em;">α</span></span></span></span> from the flow instead of being told it, which is the assumption the fixed and adaptive results currently rest on.</p><h3 id="smarter-liquidity-providers">smarter liquidity providers</h3><p>Stale liquidity-provider orders inside the quote muddy every spread measurement. Either they should requote when the maker moves, or they should go and the maker should provide depth at several levels, so that &quot;the spread&quot; means one thing.</p><h3 id="a-mechanism-for-long-memory">a mechanism for long memory</h3><p>Hawkes-process order arrivals or a regime-switching fundamental would give volatility clustering a cause. Then clustering would be a result I could switch on and off rather than an artifact of price discovery.</p><h3 id="a-markout-that-separates-information-from-footprint">a markout that separates information from footprint</h3><p>The adaptive maker charges itself about 0.6 ticks per fill at <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>α</mi><mo>=</mo><mn>0</mn></mrow><annotation encoding="application/x-tex">\alpha = 0</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal" style="margin-right:0.0037em;">α</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.6444em;"></span><span class="mord">0</span></span></span></span> for its own impact on the mid. A markout against a reference that excludes the maker's own quotes would separate the two.</p><h3 id="measure-performance-properly">measure performance properly</h3><p>Rerun <code>experiments/bench.py</code> in full on an idle machine before quoting any throughput, and profile the event loop before starting the C++ core. At about 120,000 events per run, the Python book is not obviously the bottleneck. The scheduler, the recorder and per-event allocation matter about as much.</p><h3 id="variable-order-sizes">variable order sizes</h3><p>Unit sizes keep the Glosten-Milgrom comparison clean, but they hide the depth and market-impact effects that the latency and multi-venue milestones will need.</p><h2 id="reproducibility">reproducibility</h2><p>You need <a href="https://docs.astral.sh/uv/">uv</a>. The project is pinned to Python 3.12.</p><figure class="highlight bash"><table><tr><td class="code"><pre><span class="line"><span class="built_in">cd</span> projects/10-adverse-selection-lob-sim</span><br><span class="line">uv <span class="built_in">sync</span></span><br><span class="line">uv run pytest -q                                   <span class="comment"># 32 tests</span></span><br><span class="line"></span><br><span class="line">uv run python experiments/informed_sweep.py        <span class="comment"># 216 runs, results/sweep/</span></span><br><span class="line">uv run python experiments/stylized_facts.py        <span class="comment"># 4 runs of 200k units, results/stylized/</span></span><br><span class="line">uv run python experiments/inventory_ablation.py    <span class="comment"># results/inventory_ablation.csv</span></span><br><span class="line">uv run python experiments/bench.py                 <span class="comment"># results/bench.json</span></span><br></pre></td></tr></table></figure><p>The sweep uses every core by default. Pass <code>--workers N</code> to limit it. Each script prints a summary table, and the saved output is in <code>results/*_stdout.txt</code>. A single run from Python looks like this.</p><figure class="highlight python"><table><tr><td class="code"><pre><span class="line"><span class="keyword">from</span> marketsim <span class="keyword">import</span> SimConfig, run</span><br><span class="line"><span class="keyword">from</span> marketsim.metrics <span class="keyword">import</span> summarize</span><br><span class="line"></span><br><span class="line">res = run(SimConfig(seed=<span class="number">1</span>, informed_frac=<span class="number">0.2</span>, mm_adaptive=<span class="literal">True</span>, t_end=<span class="number">20_000</span>))</span><br><span class="line"><span class="built_in">print</span>(summarize(res))</span><br></pre></td></tr></table></figure><p>Runs are bit-identical for the same <code>SimConfig</code>. The code is in <code>projects/10-adverse-selection-lob-sim</code>, with the book in <code>src/marketsim/orderbook.py</code>, the scheduler and market in <code>src/marketsim/engine.py</code>, the agents in <code>src/marketsim/agents.py</code>, and the metrics and PnL decomposition in <code>src/marketsim/metrics.py</code>.</p><h3 id="verification">verification</h3><p>An independent reviewer who did not build the project reran it on 2026-09-26, with threads capped at 2 and every process pool limited to 2 workers because the machine was shared. The 32 tests passed. The full informed-fraction sweep (216 runs) wrote an <code>aggregate.csv</code> byte identical to the committed one, and <code>runs.csv</code> matched in every column except the timing field <code>events_per_s</code>. The inventory ablation was byte identical, and the stylized-facts runs reproduced <code>acf.csv</code> and <code>jump_event_study.csv</code> byte for byte, with <code>summary.json</code> differing only in wall time and events per second. So every headline number in this post, from the 0.605 to 1.266 ticks of adverse selection to the 97,116 = 177,634 - 80,589 + 71 PnL split, holds exactly. The reviewer also corrected the sweep wall time range to 61 to 226 s. Timing was deferred. The full benchmark was not rerun, only a small smoke run, and throughput still needs to be measured on a quiet machine.</p>]]>
    </content>
    <id>https://projects.farhansadeek.com/posts/pricing-adverse-selection/</id>
    <link href="https://projects.farhansadeek.com/posts/pricing-adverse-selection/"/>
    <published>2026-09-29T06:17:10.000Z</published>
    <summary>An agent-based market where a market maker quotes against noise and informed traders, and informed flow eats more than half of its edge per fill.</summary>
    <title>Measuring What Informed Traders Cost a Market Maker</title>
    <updated>2026-09-29T06:49:26.552Z</updated>
  </entry>
  <entry>
    <author>
      <name>Farhan Sadeek</name>
      <email>farhan@farhansadeek.com</email>
    </author>
    <category term="pytorch" scheme="https://projects.farhansadeek.com/tags/pytorch/"/>
    <category term="autograd" scheme="https://projects.farhansadeek.com/tags/autograd/"/>
    <category term="numpy" scheme="https://projects.farhansadeek.com/tags/numpy/"/>
    <content>
      <![CDATA[<p>I wrote smolgrad, a reverse-mode automatic differentiation engine with a small PyTorch-shaped neural network library on top, in 891 lines of Python over numpy. It builds a dynamic graph on every forward pass, walks it in reverse topological order, and ships the pieces needed to train real models, which are broadcasting-aware ops, a module system, SGD with momentum and AdamW.</p><p>The number I trust most is not an accuracy. Trained side by side with a PyTorch twin from identical weights on identical batches, in float64, the two models finish three MNIST epochs with a maximum weight difference of <strong>5.3e-14</strong>, and an independent rerun landed at 3.9e-14. That is float64 rounding territory, and it leaves no room for a bug in any gradient or optimizer update on that path. In float32 the same experiment ends ten epochs <strong>0.313</strong> apart, and most of this post's middle is about why that is not a bug.</p><p>With that established, the models are almost a formality. A 784-512-256-10 MLP reaches <strong>97.89%</strong> MNIST test accuracy after ten float32 epochs (the PyTorch twin reaches 97.81%), which is one draw from a spread of about 0.2 points that depends on the BLAS thread count, and a character-level model reaches <strong>1.794 nats per character</strong> on held-out Tiny Shakespeare against 2.482 for a bigram baseline. A full training step is <strong>1.2x to 2.5x slower than PyTorch</strong> on one CPU thread, and a hand-written numpy baseline splits that gap into autograd overhead and kernel speed. Every timing was taken on a machine loaded by other jobs, so the absolute times are upper bounds and the single-threaded ratios are the useful part.</p><p>Code is in <code>projects/11-smolgrad</code>.</p><p><em>Reading note.</em> The argument is that an autograd engine is short, and that almost all of the difficulty is in proving it right and in the handful of places where numpy's semantics and the chain rule disagree about who owns an array. Skip to <a href="#problems">problems</a> for the bugs.</p><h2 id="table-of-contents">table of contents</h2><ul><li><a href="#what-i-wanted-to-build">what I wanted to build</a></li><li><a href="#theory">theory</a></li><li><a href="#architecture">architecture</a></li><li><a href="#implementation">implementation</a></li><li><a href="#problems">problems</a></li><li><a href="#experiments">experiments</a></li><li><a href="#results">results</a></li><li><a href="#what-i-would-change">what I would change</a></li><li><a href="#reproducibility">reproducibility</a></li></ul><h2 id="what-i-wanted-to-build">what I wanted to build</h2><p>I wanted to know what actually happens inside <code>loss.backward()</code> by writing it, and to hold the result to a standard that a toy would not survive. A scalar autograd in the style of micrograd is an afternoon. A tensor autograd that trains the same networks PyTorch does, to the same numbers, is a different object, because broadcasting, reductions over several axes, repeated indices and dtype promotion all have gradient rules that are easy to get almost right.</p><p>The goals for v0 were these.</p><ul><li>A <code>Tensor</code> with reverse-mode autodiff over a graph rebuilt every forward pass, and enough ops to train real models, from broadcasting arithmetic and batched matmul through reductions over arbitrary axes, indexing, softmax and cross-entropy.</li><li>A module system (<code>Module</code>, <code>Linear</code>, <code>Embedding</code>, <code>LayerNorm</code>, <code>Sequential</code>) with SGD with momentum and AdamW.</li><li>Proof that it is right, meaning a finite-difference check for every op and parity tests against PyTorch down to whole optimizer trajectories.</li><li>Proof that it is useful, meaning a trained MNIST MLP and a trained character model.</li><li>A benchmark against PyTorch that explains where the gap comes from, not just how big it is.</li></ul><p>The only runtime dependency is numpy. PyTorch appears only as a test oracle and as the benchmark baseline.</p><h2 id="theory">theory</h2><h3 id="reverse-mode-in-one-paragraph">reverse mode in one paragraph</h3><p>A training loss is one scalar <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>L</mi></mrow><annotation encoding="application/x-tex">L</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal">L</span></span></span></span> computed from a few million inputs. Forward-mode differentiation would need one pass per input to get <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="normal">∂</mi><mi>L</mi><mi mathvariant="normal">/</mi><mi mathvariant="normal">∂</mi><mi>θ</mi></mrow><annotation encoding="application/x-tex">\partial L / \partial \theta</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mord" style="margin-right:0.0556em;">∂</span><span class="mord mathnormal">L</span><span class="mord">/</span><span class="mord" style="margin-right:0.0556em;">∂</span><span class="mord mathnormal" style="margin-right:0.0278em;">θ</span></span></span></span>. Reverse mode gets all of them in a single backward sweep whose cost is a small constant multiple of the forward pass. The trick is that each op <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>y</mi><mo>=</mo><mi>f</mi><mo stretchy="false">(</mo><msub><mi>x</mi><mn>1</mn></msub><mo separator="true">,</mo><mo>…</mo><mo separator="true">,</mo><msub><mi>x</mi><mi>k</mi></msub><mo stretchy="false">)</mo></mrow><annotation encoding="application/x-tex">y = f(x_1, \dots, x_k)</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.625em;vertical-align:-0.1944em;"></span><span class="mord mathnormal" style="margin-right:0.0359em;">y</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mord mathnormal" style="margin-right:0.1076em;">f</span><span class="mopen">(</span><span class="mord"><span class="mord mathnormal">x</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3011em;"><span style="top:-2.55em;margin-left:0em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight">1</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span><span class="mpunct">,</span><span class="mspace" style="margin-right:0.1667em;"></span><span class="minner">…</span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mpunct">,</span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mord"><span class="mord mathnormal">x</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3361em;"><span style="top:-2.55em;margin-left:0em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight" style="margin-right:0.0315em;">k</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span><span class="mclose">)</span></span></span></span> never needs its full Jacobian. It needs only its vector-Jacobian product (VJP), a function that takes <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mover accent="true"><mi>y</mi><mo>ˉ</mo></mover><mo>=</mo><mi mathvariant="normal">∂</mi><mi>L</mi><mi mathvariant="normal">/</mi><mi mathvariant="normal">∂</mi><mi>y</mi></mrow><annotation encoding="application/x-tex">\bar{y} = \partial L / \partial y</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.7622em;vertical-align:-0.1944em;"></span><span class="mord accent"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.5678em;"><span style="top:-3em;"><span class="pstrut" style="height:3em;"></span><span class="mord mathnormal" style="margin-right:0.0359em;">y</span></span><span style="top:-3em;"><span class="pstrut" style="height:3em;"></span><span class="accent-body" style="left:-0.1944em;"><span class="mord">ˉ</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.1944em;"><span></span></span></span></span></span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mord" style="margin-right:0.0556em;">∂</span><span class="mord mathnormal">L</span><span class="mord">/</span><span class="mord" style="margin-right:0.0556em;">∂</span><span class="mord mathnormal" style="margin-right:0.0359em;">y</span></span></span></span> and returns <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mover accent="true"><mi>x</mi><mo>ˉ</mo></mover><mi>i</mi></msub><mo>=</mo><mi mathvariant="normal">∂</mi><mi>L</mi><mi mathvariant="normal">/</mi><mi mathvariant="normal">∂</mi><msub><mi>x</mi><mi>i</mi></msub></mrow><annotation encoding="application/x-tex">\bar{x}_i = \partial L / \partial x_i</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.7178em;vertical-align:-0.15em;"></span><span class="mord"><span class="mord accent"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.5678em;"><span style="top:-3em;"><span class="pstrut" style="height:3em;"></span><span class="mord mathnormal">x</span></span><span style="top:-3em;"><span class="pstrut" style="height:3em;"></span><span class="accent-body" style="left:-0.2222em;"><span class="mord">ˉ</span></span></span></span></span></span></span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3117em;"><span style="top:-2.55em;margin-left:0em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight">i</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mord" style="margin-right:0.0556em;">∂</span><span class="mord mathnormal">L</span><span class="mord">/</span><span class="mord" style="margin-right:0.0556em;">∂</span><span class="mord"><span class="mord mathnormal">x</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3117em;"><span style="top:-2.55em;margin-left:0em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight">i</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span></span></span></span> for each input.</p><p>The engine does three things.</p><ol><li>During the forward pass, record each op's inputs and a closure that computes its VJP.</li><li>Topologically sort the graph from the loss.</li><li>Walk that order in reverse, seed <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mover accent="true"><mi>L</mi><mo>ˉ</mo></mover><mo>=</mo><mn>1</mn></mrow><annotation encoding="application/x-tex">\bar{L} = 1</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.8201em;"></span><span class="mord accent"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.8201em;"><span style="top:-3em;"><span class="pstrut" style="height:3em;"></span><span class="mord mathnormal">L</span></span><span style="top:-3.2523em;"><span class="pstrut" style="height:3em;"></span><span class="accent-body" style="left:-0.2222em;"><span class="mord">ˉ</span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.6444em;"></span><span class="mord">1</span></span></span></span>, call each closure, and sum every gradient that arrives at a node from its consumers.</li></ol><p>The sum in step 3 is the multivariate chain rule. A node used twice (a residual connection, or <code>x * x</code>) receives two contributions, and forgetting to add them is the first bug everyone writes.</p><h3 id="the-vjps-that-matter">the VJPs that matter</h3><p>For a gradient <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>G</mi></mrow><annotation encoding="application/x-tex">G</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal">G</span></span></span></span> flowing into the output, the rules I needed are short.</p><table><thead><tr><th>op</th><th>backward</th></tr></thead><tbody><tr><td><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>C</mi><mo>=</mo><mi>A</mi><mi>B</mi></mrow><annotation encoding="application/x-tex">C = AB</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal" style="margin-right:0.0715em;">C</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal">A</span><span class="mord mathnormal" style="margin-right:0.0502em;">B</span></span></span></span></td><td><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mover accent="true"><mi>A</mi><mo>ˉ</mo></mover><mo>=</mo><mi>G</mi><msup><mi>B</mi><mi mathvariant="normal">⊤</mi></msup></mrow><annotation encoding="application/x-tex">\bar{A} = G B^{\top}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.8201em;"></span><span class="mord accent"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.8201em;"><span style="top:-3em;"><span class="pstrut" style="height:3em;"></span><span class="mord mathnormal">A</span></span><span style="top:-3.2523em;"><span class="pstrut" style="height:3em;"></span><span class="accent-body" style="left:-0.1111em;"><span class="mord">ˉ</span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.8491em;"></span><span class="mord mathnormal">G</span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0502em;">B</span><span class="msupsub"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.8491em;"><span style="top:-3.063em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mtight">⊤</span></span></span></span></span></span></span></span></span></span></span></span>, <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mover accent="true"><mi>B</mi><mo>ˉ</mo></mover><mo>=</mo><msup><mi>A</mi><mi mathvariant="normal">⊤</mi></msup><mi>G</mi></mrow><annotation encoding="application/x-tex">\bar{B} = A^{\top} G</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.8201em;"></span><span class="mord accent"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.8201em;"><span style="top:-3em;"><span class="pstrut" style="height:3em;"></span><span class="mord mathnormal" style="margin-right:0.0502em;">B</span></span><span style="top:-3.2523em;"><span class="pstrut" style="height:3em;"></span><span class="accent-body" style="left:-0.1667em;"><span class="mord">ˉ</span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.8491em;"></span><span class="mord"><span class="mord mathnormal">A</span><span class="msupsub"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.8491em;"><span style="top:-3.063em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mtight">⊤</span></span></span></span></span></span></span></span></span><span class="mord mathnormal">G</span></span></span></span></td></tr><tr><td>sum over an axis</td><td>broadcast <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>G</mi></mrow><annotation encoding="application/x-tex">G</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal">G</span></span></span></span> back along that axis</td></tr><tr><td><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>s</mi><mo>=</mo><mrow><mi mathvariant="normal">s</mi><mi mathvariant="normal">o</mi><mi mathvariant="normal">f</mi><mi mathvariant="normal">t</mi><mi mathvariant="normal">m</mi><mi mathvariant="normal">a</mi><mi mathvariant="normal">x</mi></mrow><mo stretchy="false">(</mo><mi>x</mi><mo stretchy="false">)</mo></mrow><annotation encoding="application/x-tex">s = \mathrm{softmax}(x)</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal">s</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mord"><span class="mord mathrm">softmax</span></span><span class="mopen">(</span><span class="mord mathnormal">x</span><span class="mclose">)</span></span></span></span></td><td><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mover accent="true"><mi>x</mi><mo>ˉ</mo></mover><mo>=</mo><mi>s</mi><mo>⊙</mo><mo stretchy="false">(</mo><mi>G</mi><mo>−</mo><mo stretchy="false">⟨</mo><mi>G</mi><mo separator="true">,</mo><mi>s</mi><mo stretchy="false">⟩</mo><mo stretchy="false">)</mo></mrow><annotation encoding="application/x-tex">\bar{x} = s \odot (G - \langle G, s \rangle)</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.5678em;"></span><span class="mord accent"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.5678em;"><span style="top:-3em;"><span class="pstrut" style="height:3em;"></span><span class="mord mathnormal">x</span></span><span style="top:-3em;"><span class="pstrut" style="height:3em;"></span><span class="accent-body" style="left:-0.2222em;"><span class="mord">ˉ</span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord mathnormal">s</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">⊙</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mopen">(</span><span class="mord mathnormal">G</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">−</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mopen">⟨</span><span class="mord mathnormal">G</span><span class="mpunct">,</span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mord mathnormal">s</span><span class="mclose">⟩)</span></span></span></span></td></tr><tr><td><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>y</mi><mo>=</mo><mi>log</mi><mo>⁡</mo><mrow><mi mathvariant="normal">s</mi><mi mathvariant="normal">o</mi><mi mathvariant="normal">f</mi><mi mathvariant="normal">t</mi><mi mathvariant="normal">m</mi><mi mathvariant="normal">a</mi><mi mathvariant="normal">x</mi></mrow><mo stretchy="false">(</mo><mi>x</mi><mo stretchy="false">)</mo></mrow><annotation encoding="application/x-tex">y = \log\mathrm{softmax}(x)</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.625em;vertical-align:-0.1944em;"></span><span class="mord mathnormal" style="margin-right:0.0359em;">y</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mop">lo<span style="margin-right:0.0139em;">g</span></span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mord"><span class="mord mathrm">softmax</span></span><span class="mopen">(</span><span class="mord mathnormal">x</span><span class="mclose">)</span></span></span></span></td><td><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mover accent="true"><mi>x</mi><mo>ˉ</mo></mover><mo>=</mo><mi>G</mi><mo>−</mo><mrow><mi mathvariant="normal">s</mi><mi mathvariant="normal">o</mi><mi mathvariant="normal">f</mi><mi mathvariant="normal">t</mi><mi mathvariant="normal">m</mi><mi mathvariant="normal">a</mi><mi mathvariant="normal">x</mi></mrow><mo stretchy="false">(</mo><mi>x</mi><mo stretchy="false">)</mo><mo>∑</mo><mi>G</mi></mrow><annotation encoding="application/x-tex">\bar{x} = G - \mathrm{softmax}(x) \sum G</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.5678em;"></span><span class="mord accent"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.5678em;"><span style="top:-3em;"><span class="pstrut" style="height:3em;"></span><span class="mord mathnormal">x</span></span><span style="top:-3em;"><span class="pstrut" style="height:3em;"></span><span class="accent-body" style="left:-0.2222em;"><span class="mord">ˉ</span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.7667em;vertical-align:-0.0833em;"></span><span class="mord mathnormal">G</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">−</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mord"><span class="mord mathrm">softmax</span></span><span class="mopen">(</span><span class="mord mathnormal">x</span><span class="mclose">)</span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mop op-symbol small-op" style="position:relative;top:0em;">∑</span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mord mathnormal">G</span></span></span></span></td></tr><tr><td>mean cross-entropy over <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>N</mi></mrow><annotation encoding="application/x-tex">N</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal" style="margin-right:0.109em;">N</span></span></span></span> rows</td><td><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mover accent="true"><mi>z</mi><mo>ˉ</mo></mover><mo>=</mo><mo stretchy="false">(</mo><mrow><mi mathvariant="normal">s</mi><mi mathvariant="normal">o</mi><mi mathvariant="normal">f</mi><mi mathvariant="normal">t</mi><mi mathvariant="normal">m</mi><mi mathvariant="normal">a</mi><mi mathvariant="normal">x</mi></mrow><mo stretchy="false">(</mo><mi>z</mi><mo stretchy="false">)</mo><mo>−</mo><mrow><mi mathvariant="normal">o</mi><mi mathvariant="normal">n</mi><mi mathvariant="normal">e</mi><mi mathvariant="normal">h</mi><mi mathvariant="normal">o</mi><mi mathvariant="normal">t</mi></mrow><mo stretchy="false">)</mo><mi mathvariant="normal">/</mi><mi>N</mi></mrow><annotation encoding="application/x-tex">\bar{z} = (\mathrm{softmax}(z) - \mathrm{onehot}) / N</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.5678em;"></span><span class="mord accent"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.5678em;"><span style="top:-3em;"><span class="pstrut" style="height:3em;"></span><span class="mord mathnormal" style="margin-right:0.044em;">z</span></span><span style="top:-3em;"><span class="pstrut" style="height:3em;"></span><span class="accent-body" style="left:-0.1944em;"><span class="mord">ˉ</span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mopen">(</span><span class="mord"><span class="mord mathrm">softmax</span></span><span class="mopen">(</span><span class="mord mathnormal" style="margin-right:0.044em;">z</span><span class="mclose">)</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">−</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mord"><span class="mord mathrm">onehot</span></span><span class="mclose">)</span><span class="mord">/</span><span class="mord mathnormal" style="margin-right:0.109em;">N</span></span></span></span></td></tr><tr><td>gather <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>W</mi><mo stretchy="false">[</mo><mrow><mi mathvariant="normal">i</mi><mi mathvariant="normal">d</mi><mi mathvariant="normal">x</mi></mrow><mo stretchy="false">]</mo></mrow><annotation encoding="application/x-tex">W[\mathrm{idx}]</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mord mathnormal" style="margin-right:0.1389em;">W</span><span class="mopen">[</span><span class="mord"><span class="mord mathrm">idx</span></span><span class="mclose">]</span></span></span></span></td><td>scatter-add <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>G</mi></mrow><annotation encoding="application/x-tex">G</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal">G</span></span></span></span> back into the rows of <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>W</mi></mrow><annotation encoding="application/x-tex">W</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal" style="margin-right:0.1389em;">W</span></span></span></span></td></tr></tbody></table><p>The softmax rule matters because it never forms the <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>C</mi><mo>×</mo><mi>C</mi></mrow><annotation encoding="application/x-tex">C \times C</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.7667em;vertical-align:-0.0833em;"></span><span class="mord mathnormal" style="margin-right:0.0715em;">C</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">×</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal" style="margin-right:0.0715em;">C</span></span></span></span> Jacobian <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mrow><mi mathvariant="normal">d</mi><mi mathvariant="normal">i</mi><mi mathvariant="normal">a</mi><mi mathvariant="normal">g</mi></mrow><mo stretchy="false">(</mo><mi>s</mi><mo stretchy="false">)</mo><mo>−</mo><mi>s</mi><msup><mi>s</mi><mi mathvariant="normal">⊤</mi></msup></mrow><annotation encoding="application/x-tex">\mathrm{diag}(s) - s s^{\top}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mord"><span class="mord mathrm" style="margin-right:0.0139em;">diag</span></span><span class="mopen">(</span><span class="mord mathnormal">s</span><span class="mclose">)</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">−</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:0.8491em;"></span><span class="mord mathnormal">s</span><span class="mord"><span class="mord mathnormal">s</span><span class="msupsub"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.8491em;"><span style="top:-3.063em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mtight">⊤</span></span></span></span></span></span></span></span></span></span></span></span>. It contracts the Jacobian with <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>G</mi></mrow><annotation encoding="application/x-tex">G</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal">G</span></span></span></span> analytically, which is the whole point of VJPs.</p><p>The gather rule is where the real difficulty hides. If index 7 appears three times in a batch, row 7 of the embedding table must receive the sum of three gradient rows. numpy's <code>W[idx] += G</code> keeps only one of the three writes, silently. More on that in <a href="#problems">problems</a>.</p><h3 id="broadcasting-has-an-adjoint">broadcasting has an adjoint</h3><p>Broadcasting copies a value along an axis. The adjoint of copying is summing. So every binary op has to take the gradient it received, which has the broadcast output's shape, and sum it back down to each input's shape. There are two cases. Broadcasting can prepend axes (a bias of shape <code>(512,)</code> added to activations of shape <code>(128, 512)</code>), and it can stretch a size-1 axis (shape <code>(128, 1)</code> against <code>(128, 512)</code>). Here is the function that undoes both, in full.</p><figure class="highlight python"><table><tr><td class="code"><pre><span class="line"><span class="keyword">def</span> <span class="title function_">_unbroadcast</span>(<span class="params">grad, shape</span>):</span><br><span class="line">    <span class="keyword">if</span> grad.shape == shape:</span><br><span class="line">        <span class="keyword">return</span> grad</span><br><span class="line">    extra = grad.ndim - <span class="built_in">len</span>(shape)</span><br><span class="line">    <span class="keyword">if</span> extra &gt; <span class="number">0</span>:</span><br><span class="line">        grad = grad.<span class="built_in">sum</span>(axis=<span class="built_in">tuple</span>(<span class="built_in">range</span>(extra)))</span><br><span class="line">    stretched = <span class="built_in">tuple</span>(i <span class="keyword">for</span> i, s <span class="keyword">in</span> <span class="built_in">enumerate</span>(shape) <span class="keyword">if</span> s == <span class="number">1</span> <span class="keyword">and</span> grad.shape[i] != <span class="number">1</span>)</span><br><span class="line">    <span class="keyword">if</span> stretched:</span><br><span class="line">        grad = grad.<span class="built_in">sum</span>(axis=stretched, keepdims=<span class="literal">True</span>)</span><br><span class="line">    <span class="keyword">return</span> grad.reshape(shape)</span><br></pre></td></tr></table></figure><p>Every binary op routes both parent gradients through it, and the engine asserts after every VJP that the gradient's shape equals its tensor's shape. That assertion catches an entire class of bug at the op that caused it rather than three layers later.</p><h3 id="checking-gradients-numerically">checking gradients numerically</h3><p>The oracle for every op is a central difference.</p><p class="katex-block"><span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mfrac><mrow><mi mathvariant="normal">∂</mi><mi>L</mi></mrow><mrow><mi mathvariant="normal">∂</mi><msub><mi>x</mi><mi>i</mi></msub></mrow></mfrac><mo>≈</mo><mfrac><mrow><mi>L</mi><mo stretchy="false">(</mo><mi>x</mi><mo>+</mo><mi>h</mi><msub><mi>e</mi><mi>i</mi></msub><mo stretchy="false">)</mo><mo>−</mo><mi>L</mi><mo stretchy="false">(</mo><mi>x</mi><mo>−</mo><mi>h</mi><msub><mi>e</mi><mi>i</mi></msub><mo stretchy="false">)</mo></mrow><mrow><mn>2</mn><mi>h</mi></mrow></mfrac><mo separator="true">,</mo></mrow><annotation encoding="application/x-tex">\frac{\partial L}{\partial x_i} \approx \frac{L(x + h e_i) - L(x - h e_i)}{2h},</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:2.2074em;vertical-align:-0.836em;"></span><span class="mord"><span class="mopen nulldelimiter"></span><span class="mfrac"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:1.3714em;"><span style="top:-2.314em;"><span class="pstrut" style="height:3em;"></span><span class="mord"><span class="mord" style="margin-right:0.0556em;">∂</span><span class="mord"><span class="mord mathnormal">x</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3117em;"><span style="top:-2.55em;margin-left:0em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight">i</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span></span></span><span style="top:-3.23em;"><span class="pstrut" style="height:3em;"></span><span class="frac-line" style="border-bottom-width:0.04em;"></span></span><span style="top:-3.677em;"><span class="pstrut" style="height:3em;"></span><span class="mord"><span class="mord" style="margin-right:0.0556em;">∂</span><span class="mord mathnormal">L</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.836em;"><span></span></span></span></span></span><span class="mclose nulldelimiter"></span></span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">≈</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:2.113em;vertical-align:-0.686em;"></span><span class="mord"><span class="mopen nulldelimiter"></span><span class="mfrac"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:1.427em;"><span style="top:-2.314em;"><span class="pstrut" style="height:3em;"></span><span class="mord"><span class="mord">2</span><span class="mord mathnormal">h</span></span></span><span style="top:-3.23em;"><span class="pstrut" style="height:3em;"></span><span class="frac-line" style="border-bottom-width:0.04em;"></span></span><span style="top:-3.677em;"><span class="pstrut" style="height:3em;"></span><span class="mord"><span class="mord mathnormal">L</span><span class="mopen">(</span><span class="mord mathnormal">x</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">+</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mord mathnormal">h</span><span class="mord"><span class="mord mathnormal">e</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3117em;"><span style="top:-2.55em;margin-left:0em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight">i</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span><span class="mclose">)</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">−</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mord mathnormal">L</span><span class="mopen">(</span><span class="mord mathnormal">x</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">−</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mord mathnormal">h</span><span class="mord"><span class="mord mathnormal">e</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3117em;"><span style="top:-2.55em;margin-left:0em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight">i</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span><span class="mclose">)</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.686em;"><span></span></span></span></span></span><span class="mclose nulldelimiter"></span></span><span class="mpunct">,</span></span></span></span></span></p><p>whose truncation error is <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>O</mi><mo stretchy="false">(</mo><msup><mi>h</mi><mn>2</mn></msup><mo stretchy="false">)</mo></mrow><annotation encoding="application/x-tex">O(h^2)</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1.0641em;vertical-align:-0.25em;"></span><span class="mord mathnormal" style="margin-right:0.0278em;">O</span><span class="mopen">(</span><span class="mord"><span class="mord mathnormal">h</span><span class="msupsub"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.8141em;"><span style="top:-3.063em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight">2</span></span></span></span></span></span></span></span><span class="mclose">)</span></span></span></span>. I run it in float64 with <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>h</mi><mo>=</mo><msup><mn>10</mn><mrow><mo>−</mo><mn>6</mn></mrow></msup></mrow><annotation encoding="application/x-tex">h = 10^{-6}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6944em;"></span><span class="mord mathnormal">h</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.8141em;"></span><span class="mord">1</span><span class="mord"><span class="mord">0</span><span class="msupsub"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.8141em;"><span style="top:-3.063em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mtight">−</span><span class="mord mtight">6</span></span></span></span></span></span></span></span></span></span></span></span>. For an op with a tensor output, the checker contracts the output with a fixed random tensor <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>w</mi></mrow><annotation encoding="application/x-tex">w</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal" style="margin-right:0.0269em;">w</span></span></span></span> and checks <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>L</mi><mo>=</mo><mo>∑</mo><mi>w</mi><mo>⊙</mo><mi>f</mi><mo stretchy="false">(</mo><mi>x</mi><mo stretchy="false">)</mo></mrow><annotation encoding="application/x-tex">L = \sum w \odot f(x)</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal">L</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mop op-symbol small-op" style="position:relative;top:0em;">∑</span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mord mathnormal" style="margin-right:0.0269em;">w</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">⊙</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mord mathnormal" style="margin-right:0.1076em;">f</span><span class="mopen">(</span><span class="mord mathnormal">x</span><span class="mclose">)</span></span></span></span> instead of <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mo>∑</mo><mi>f</mi><mo stretchy="false">(</mo><mi>x</mi><mo stretchy="false">)</mo></mrow><annotation encoding="application/x-tex">\sum f(x)</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mop op-symbol small-op" style="position:relative;top:0em;">∑</span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mord mathnormal" style="margin-right:0.1076em;">f</span><span class="mopen">(</span><span class="mord mathnormal">x</span><span class="mclose">)</span></span></span></span>. With plain <code>sum()</code>, every output element gets weight 1, so a backward that permutes or transposes output gradients could pass. A random projection gives each element its own weight and closes that hole.</p><h2 id="architecture">architecture</h2><p>The value and the graph node are the same object. A <code>Tensor</code> wraps one <code>np.ndarray</code> (float32 by default, float64 when given float64) and uses <code>__slots__</code>, and there is no separate <code>Node</code> or <code>Function</code> class.</p><figure data-figure="diagram:smolgrad-layers"></figure><p>One training step flows like this.</p><figure data-figure="diagram:smolgrad-step"></figure><h3 id="closures-not-function-objects">closures, not Function objects</h3><p>Each op defines its backward as a closure that captures what it needs from the forward pass. <code>exp</code> captures its output, softmax captures the probabilities. The closure <em>returns</em> parent gradients rather than writing them anywhere, which keeps all accumulation in one place (the engine) and makes each op's backward a pure function that finite differences can test in isolation.</p><p>The alternatives were a <code>Function</code> class per op, as in PyTorch's <code>autograd.Function</code> and tinygrad, or a tape of <code>(op, inputs, output)</code> records. Closures are the least code per op. Their cost is that they are opaque, so you cannot inspect, rewrite or fuse the graph. That trade-off comes back in <a href="#what-i-would-change">what I would change</a>.</p><h3 id="gradients-live-in-a-local-dict">gradients live in a local dict</h3><p>During <code>backward()</code> the gradients of intermediate nodes live in a dict keyed by <code>id(node)</code>, local to that call, and each entry is popped as soon as its node is processed. Only leaves get a <code>.grad</code>, and it accumulates across calls until zeroed, like PyTorch.</p><p>The micrograd approach of storing <code>.grad</code> on every node is simpler and has two bugs waiting. A second backward call double counts intermediate gradients unless something resets them, and every intermediate keeps a gradient array alive as long as the graph lives. The local dict avoids both.</p><h3 id="pytorch-conventions-on-purpose">PyTorch conventions, on purpose</h3><p><code>Linear.weight</code> is stored as <code>(out_features, in_features)</code>, init is <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>U</mi><mo stretchy="false">(</mo><mo>−</mo><mn>1</mn><mi mathvariant="normal">/</mi><msqrt><mtext>fan_in</mtext></msqrt><mo separator="true">,</mo><mn>1</mn><mi mathvariant="normal">/</mi><msqrt><mtext>fan_in</mtext></msqrt><mo stretchy="false">)</mo></mrow><annotation encoding="application/x-tex">U(-1/\sqrt{\text{fan\_in}}, 1/\sqrt{\text{fan\_in}})</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1.24em;vertical-align:-0.3628em;"></span><span class="mord mathnormal" style="margin-right:0.109em;">U</span><span class="mopen">(</span><span class="mord">−</span><span class="mord">1/</span><span class="mord sqrt"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.8772em;"><span class="svg-align" style="top:-3.2em;"><span class="pstrut" style="height:3.2em;"></span><span class="mord" style="padding-left:1em;"><span class="mord text"><span class="mord">fan_in</span></span></span></span><span style="top:-2.8372em;"><span class="pstrut" style="height:3.2em;"></span><span class="hide-tail" style="min-width:1.02em;height:1.28em;"><svg xmlns="http://www.w3.org/2000/svg" width="400em" height="1.28em" viewBox="0 0 400000 1296" preserveAspectRatio="xMinYMin slice"><path d="M263,681c0.7,0,18,39.7,52,119c34,79.3,68.167,158.7,102.5,238c34.3,79.3,51.8,119.3,52.5,120c340,-704.7,510.7,-1060.3,512,-1067l0 -0c4.7,-7.3,11,-11,19,-11H40000v40H1012.3s-271.3,567,-271.3,567c-38.7,80.7,-84,175,-136,283c-52,108,-89.167,185.3,-111.5,232c-22.3,46.7,-33.8,70.3,-34.5,71c-4.7,4.7,-12.3,7,-23,7s-12,-1,-12,-1s-109,-253,-109,-253c-72.7,-168,-109.3,-252,-110,-252c-10.7,8,-22,16.7,-34,26c-22,17.3,-33.3,26,-34,26s-26,-26,-26,-26s76,-59,76,-59s76,-60,76,-60zM1001 80h400000v40h-400000z"/></svg></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.3628em;"><span></span></span></span></span></span><span class="mpunct">,</span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mord">1/</span><span class="mord sqrt"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.8772em;"><span class="svg-align" style="top:-3.2em;"><span class="pstrut" style="height:3.2em;"></span><span class="mord" style="padding-left:1em;"><span class="mord text"><span class="mord">fan_in</span></span></span></span><span style="top:-2.8372em;"><span class="pstrut" style="height:3.2em;"></span><span class="hide-tail" style="min-width:1.02em;height:1.28em;"><svg xmlns="http://www.w3.org/2000/svg" width="400em" height="1.28em" viewBox="0 0 400000 1296" preserveAspectRatio="xMinYMin slice"><path d="M263,681c0.7,0,18,39.7,52,119c34,79.3,68.167,158.7,102.5,238c34.3,79.3,51.8,119.3,52.5,120c340,-704.7,510.7,-1060.3,512,-1067l0 -0c4.7,-7.3,11,-11,19,-11H40000v40H1012.3s-271.3,567,-271.3,567c-38.7,80.7,-84,175,-136,283c-52,108,-89.167,185.3,-111.5,232c-22.3,46.7,-33.8,70.3,-34.5,71c-4.7,4.7,-12.3,7,-23,7s-12,-1,-12,-1s-109,-253,-109,-253c-72.7,-168,-109.3,-252,-110,-252c-10.7,8,-22,16.7,-34,26c-22,17.3,-33.3,26,-34,26s-26,-26,-26,-26s76,-59,76,-59s76,-60,76,-60zM1001 80h400000v40h-400000z"/></svg></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.3628em;"><span></span></span></span></span></span><span class="mclose">)</span></span></span></span>, the SGD momentum buffer starts at <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>v</mi><mo>=</mo><mi>g</mi></mrow><annotation encoding="application/x-tex">v = g</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal" style="margin-right:0.0359em;">v</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.625em;vertical-align:-0.1944em;"></span><span class="mord mathnormal" style="margin-right:0.0359em;">g</span></span></span></span> on the first step, the AdamW update order matches, relu's gradient at exactly 0 is 0, and LayerNorm uses the biased variance. None of these are the only reasonable choice. They are PyTorch's choices, and matching them is what makes step-for-step parity tests possible, which is what makes the 5.3e-14 result possible.</p><h2 id="implementation">implementation</h2><p>The library is 891 lines across seven files. <code>tensor.py</code> (439 lines) holds the Tensor, the engine and every primitive op. <code>functional.py</code> holds the composed and fused ops, <code>nn.py</code> the module system, <code>optim.py</code> the optimizers and <code>gradcheck.py</code> the checker.</p><h3 id="recording-an-edge">recording an edge</h3><p>Every op builds its output through one function, and that function is the only place the graph grows.</p><figure class="highlight python"><table><tr><td class="code"><pre><span class="line"><span class="meta">@staticmethod</span></span><br><span class="line"><span class="keyword">def</span> <span class="title function_">_make</span>(<span class="params">data, parents, backward, op</span>):</span><br><span class="line">    out = Tensor(data, dtype=data.dtype)</span><br><span class="line">    <span class="keyword">if</span> _GRAD_ENABLED <span class="keyword">and</span> <span class="built_in">any</span>(p.requires_grad <span class="keyword">for</span> p <span class="keyword">in</span> parents):</span><br><span class="line">        out.requires_grad = <span class="literal">True</span></span><br><span class="line">        out._parents = parents</span><br><span class="line">        out._backward = backward</span><br><span class="line">        out._op = op</span><br><span class="line">    <span class="keyword">return</span> out</span><br></pre></td></tr></table></figure><p>Under <code>no_grad()</code>, or when no parent needs a gradient, nothing is stored, so evaluation builds no graph at all. A typical op then reads like its math.</p><figure class="highlight python"><table><tr><td class="code"><pre><span class="line"><span class="keyword">def</span> <span class="title function_">__mul__</span>(<span class="params">self, other</span>):</span><br><span class="line">    other = <span class="variable language_">self</span>._const(other)</span><br><span class="line">    a, b = <span class="variable language_">self</span>, other</span><br><span class="line"></span><br><span class="line">    <span class="keyword">def</span> <span class="title function_">backward</span>(<span class="params">g</span>):</span><br><span class="line">        ga = _unbroadcast(g * b.data, a.shape) <span class="keyword">if</span> a.requires_grad <span class="keyword">else</span> <span class="literal">None</span></span><br><span class="line">        gb = _unbroadcast(g * a.data, b.shape) <span class="keyword">if</span> b.requires_grad <span class="keyword">else</span> <span class="literal">None</span></span><br><span class="line">        <span class="keyword">return</span> ga, gb</span><br><span class="line"></span><br><span class="line">    <span class="keyword">return</span> Tensor._make(a.data * b.data, (a, b), backward, <span class="string">&quot;mul&quot;</span>)</span><br></pre></td></tr></table></figure><h3 id="the-engine">the engine</h3><p>The topological sort is an iterative depth-first search with an explicit stack of <code>(node, expanded)</code> pairs, emitting a node only after all its parents. The reverse sweep is the part worth reading.</p><figure class="highlight python"><table><tr><td class="code"><pre><span class="line">grads = &#123;<span class="built_in">id</span>(<span class="variable language_">self</span>): grad&#125;</span><br><span class="line">handed_out = <span class="built_in">set</span>()</span><br><span class="line"><span class="keyword">for</span> node <span class="keyword">in</span> <span class="built_in">reversed</span>(<span class="variable language_">self</span>._toposort()):</span><br><span class="line">    g = grads.pop(<span class="built_in">id</span>(node), <span class="literal">None</span>)</span><br><span class="line">    <span class="keyword">if</span> g <span class="keyword">is</span> <span class="literal">None</span>:</span><br><span class="line">        <span class="keyword">continue</span></span><br><span class="line">    <span class="keyword">if</span> node._backward <span class="keyword">is</span> <span class="literal">None</span>:  <span class="comment"># leaf</span></span><br><span class="line">        <span class="keyword">if</span> node.grad <span class="keyword">is</span> <span class="keyword">not</span> <span class="literal">None</span>:</span><br><span class="line">            node.grad = node.grad + g</span><br><span class="line">            <span class="keyword">continue</span></span><br><span class="line">        <span class="keyword">if</span> <span class="built_in">id</span>(g) <span class="keyword">in</span> handed_out <span class="keyword">or</span> <span class="keyword">not</span> g.flags.owndata <span class="keyword">or</span> <span class="keyword">not</span> g.flags.writeable:</span><br><span class="line">            g = g.copy()</span><br><span class="line">        handed_out.add(<span class="built_in">id</span>(g))</span><br><span class="line">        node.grad = g</span><br><span class="line">        <span class="keyword">continue</span></span><br><span class="line">    <span class="keyword">for</span> p, pg <span class="keyword">in</span> <span class="built_in">zip</span>(node._parents, node._backward(g)):</span><br><span class="line">        <span class="keyword">if</span> pg <span class="keyword">is</span> <span class="literal">None</span> <span class="keyword">or</span> <span class="keyword">not</span> p.requires_grad:</span><br><span class="line">            <span class="keyword">continue</span></span><br><span class="line">        <span class="keyword">assert</span> pg.shape == p.shape, <span class="string">f&quot;<span class="subst">&#123;node._op&#125;</span>: grad <span class="subst">&#123;pg.shape&#125;</span> vs param <span class="subst">&#123;p.shape&#125;</span>&quot;</span></span><br><span class="line">        k = <span class="built_in">id</span>(p)</span><br><span class="line">        grads[k] = pg <span class="keyword">if</span> k <span class="keyword">not</span> <span class="keyword">in</span> grads <span class="keyword">else</span> grads[k] + pg</span><br></pre></td></tr></table></figure><p>The last line is the chain rule's sum. Note that it allocates a new array (<code>grads[k] + pg</code>) rather than adding in place, because <code>pg</code> may be the same object another parent also received. The leaf branch with its three-way copy condition looks paranoid. It is the fix for the first bug in <a href="#problems">problems</a>.</p><h3 id="fused-where-it-pays-composed-where-it-does-not">fused where it pays, composed where it does not</h3><p><code>linear</code> and <code>layer_norm</code> are compositions of primitives. LayerNorm is about ten graph nodes, slower than a fused kernel, but its gradient comes free from primitives that are already checked, and it matches <code>torch.nn.LayerNorm</code> gradients to a relative 1e-8.</p><p><code>cross_entropy</code> is fused. Composing <code>log(softmax(x))</code> produces three graph nodes and computes a log of something that can underflow to zero. The fused version computes a max-shifted log-sum-exp once and has the well-known backward.</p><figure class="highlight python"><table><tr><td class="code"><pre><span class="line"><span class="keyword">def</span> <span class="title function_">backward</span>(<span class="params">g</span>):</span><br><span class="line">    grad = np.exp(logp)</span><br><span class="line">    grad[np.arange(N), t] -= <span class="number">1.0</span></span><br><span class="line">    grad *= g * scale  <span class="comment"># g is the upstream 0-d gradient, scale is 1/N for mean</span></span><br><span class="line">    <span class="keyword">return</span> (grad.reshape(logits.shape),)</span><br></pre></td></tr></table></figure><p>A test checks it against the composed version, and another feeds it logits large enough that the naive form overflows.</p><p><code>embedding</code> is fused because its fast backward cannot be expressed with the other ops. That backward turned out to be the most expensive line in the character model, which is the fourth entry in <a href="#problems">problems</a>.</p><h3 id="modules-without-registration">modules without registration</h3><p>A <code>Module</code> finds its parameters by walking <code>vars(self)</code> in insertion order, recursing into sub-modules and lists of them, with no <code>register_parameter</code>. The deterministic order matters because optimizer state is a list aligned with the parameters, and the parity tests copy weights to PyTorch by zipping the two parameter lists.</p><h3 id="adamw-with-no-allocations">AdamW with no allocations</h3><p>AdamW is the decoupled-weight-decay version, where decay multiplies the weights directly instead of being added to the gradient and then divided by the adaptive denominator. Written as ordinary numpy expressions, one step created about six parameter-sized temporaries per tensor. The version that shipped uses one preallocated scratch buffer per parameter and <code>out=</code> arguments throughout.</p><figure class="highlight python"><table><tr><td class="code"><pre><span class="line">np.multiply(g, <span class="number">1.0</span> - <span class="variable language_">self</span>.b1, out=tmp)   <span class="comment"># m = b1*m + (1-b1)*g</span></span><br><span class="line">m *= <span class="variable language_">self</span>.b1</span><br><span class="line">m += tmp</span><br><span class="line">np.multiply(g, g, out=tmp)               <span class="comment"># v = b2*v + (1-b2)*g^2</span></span><br><span class="line">tmp *= <span class="number">1.0</span> - <span class="variable language_">self</span>.b2</span><br><span class="line">v *= <span class="variable language_">self</span>.b2</span><br><span class="line">v += tmp</span><br><span class="line">np.sqrt(v, out=tmp)                      <span class="comment"># denom = sqrt(v)/sqrt(bc2) + eps</span></span><br><span class="line">tmp *= inv_sqrt_bc2</span><br><span class="line">tmp += <span class="variable language_">self</span>.eps</span><br><span class="line">np.divide(m, tmp, out=tmp)               <span class="comment"># p -= lr/bc1 * m/denom</span></span><br><span class="line">tmp *= step_size</span><br><span class="line">p.data -= tmp</span><br></pre></td></tr></table></figure><p>It still makes about ten passes over each parameter, which is where most of the remaining optimizer gap to PyTorch comes from.</p><h3 id="the-tests">the tests</h3><p><code>pytest</code> collects 101 tests in three files.</p><ul><li><code>test_gradcheck.py</code> has 58 finite-difference cases covering every op, including broadcasting variants, batched and broadcast-batch matmul, reductions over several axes and indexing with repeated indices, plus a self-test that the checker rejects a deliberately wrong backward. A checker that cannot fail proves nothing.</li><li><code>test_engine.py</code> has 19 tests of engine semantics. Diamonds, accumulation across backward calls, aliasing, <code>no_grad</code>, dtype preservation, a 20,000-iteration chain of about 40,000 nodes, and both paths of the embedding scatter against <code>np.add.at</code>.</li><li><code>test_parity_torch.py</code> has 23 tests comparing op values and gradients, modules, and 20-step SGD and AdamW trajectories against PyTorch. The trajectory tests require losses equal to a relative 1e-9 at every step and final weights equal to 1e-8 relative in float64.</li></ul><h2 id="problems">problems</h2><p>This section lists only the issues that left evidence in the code, the tests or the results files.</p><h3 id="1-leaf-gradients-aliasing-each-other">1. leaf gradients aliasing each other</h3><p><code>add</code> returns the <em>same</em> array object as the gradient for both parents, since <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="normal">∂</mi><mo stretchy="false">(</mo><mi>a</mi><mo>+</mo><mi>b</mi><mo stretchy="false">)</mo><mi mathvariant="normal">/</mi><mi mathvariant="normal">∂</mi><mi>a</mi></mrow><annotation encoding="application/x-tex">\partial(a + b)/\partial a</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mord" style="margin-right:0.0556em;">∂</span><span class="mopen">(</span><span class="mord mathnormal">a</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">+</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mord mathnormal">b</span><span class="mclose">)</span><span class="mord">/</span><span class="mord" style="margin-right:0.0556em;">∂</span><span class="mord mathnormal">a</span></span></span></span> and <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="normal">∂</mi><mo stretchy="false">(</mo><mi>a</mi><mo>+</mo><mi>b</mi><mo stretchy="false">)</mo><mi mathvariant="normal">/</mi><mi mathvariant="normal">∂</mi><mi>b</mi></mrow><annotation encoding="application/x-tex">\partial(a + b)/\partial b</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mord" style="margin-right:0.0556em;">∂</span><span class="mopen">(</span><span class="mord mathnormal">a</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">+</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mord mathnormal">b</span><span class="mclose">)</span><span class="mord">/</span><span class="mord" style="margin-right:0.0556em;">∂</span><span class="mord mathnormal">b</span></span></span></span> are both the identity. And <code>sum</code>'s backward returns <code>np.broadcast_to(g, shape)</code>, which is a read-only view with zero strides, not a real array. The first version of the engine stored whatever arrived at a leaf directly as its <code>.grad</code>.</p><p>That caused two failures. Two parameters could end up sharing one gradient array, so an in-place change to one silently changed the other. And the first time a leaf's gradient came from <code>sum</code> and was later accumulated into, the in-place add failed outright because the array was read-only.</p><p>The regression test is four lines and describes the bug exactly.</p><figure class="highlight python"><table><tr><td class="code"><pre><span class="line">a = Tensor(np.zeros(<span class="number">3</span>), requires_grad=<span class="literal">True</span>)</span><br><span class="line">b = Tensor(np.zeros(<span class="number">3</span>), requires_grad=<span class="literal">True</span>)</span><br><span class="line">(a + b).<span class="built_in">sum</span>().backward()</span><br><span class="line">a.grad += <span class="number">100.0</span></span><br><span class="line">np.testing.assert_allclose(b.grad, [<span class="number">1.0</span>, <span class="number">1.0</span>, <span class="number">1.0</span>])</span><br></pre></td></tr></table></figure><p>The obvious fix is to copy every incoming gradient at every leaf. That works and costs a full copy of every parameter's gradient every step. The shipped fix copies only when it must, which is when the array is a view (<code>not g.flags.owndata</code>), is read-only, or has already been handed to another leaf in this same call (tracked in <code>handed_out</code>). Otherwise the leaf adopts the fresh array the VJP produced. The caller's seed gradient is always copied, since it might be the caller's own array and could otherwise become a leaf's <code>.grad</code>, and a second test covers that.</p><p>The lesson is that the engine's contract with every VJP needs an ownership rule. Mine is that closures never mutate their input gradient, and the engine makes sure a leaf's gradient is its own.</p><h3 id="2-silent-float64-promotion">2. silent float64 promotion</h3><p><code>Tensor(float32) * 0.5</code> must stay float32. Wrapping a Python scalar in a plain numpy array makes it float64, and numpy's promotion rules then make the whole product float64. Nothing crashes. Activations would quietly double in size and every downstream op would run at float64 speed.</p><p>The fix is <code>Tensor._const</code>, which wraps a constant in the <em>calling tensor's</em> dtype. The optimizer had the same bug in a different place. AdamW's bias correction <code>1 / sqrt(1 - b2**t)</code> computed with numpy is a float64 scalar, and multiplying a float32 buffer by it promotes. It is now converted to a Python float first. <code>test_float32_stays_float32_with_python_scalars</code> guards the op side, and the float32 parity test asserts the loss dtype.</p><h3 id="3-the-recursion-limit">3. the recursion limit</h3><p>A recursive topological sort is five lines and breaks at Python's default recursion limit of 1,000 frames. An MLP never gets near that. A long unrolled recurrence, or any test that chains ops in a loop, does. I replaced it with the iterative DFS above, and <code>test_deep_graph_does_not_recurse</code> runs <code>y = y * 1.0 + 0.0</code> twenty thousand times and checks the gradient is exactly 1.</p><h3 id="4-np-add-at-was-the-most-expensive-call-in-a-training-step">4. np.add.at was the most expensive call in a training step</h3><p>The embedding backward has to scatter-add gradient rows into the table with repeated indices accumulating. The textbook numpy answer is <code>np.add.at(full, idx, g)</code>, and that is what the generic indexing op still uses. It is correct. It is also an unbuffered per-element loop.</p><p>Profiling 200 steps of the character model with cProfile showed it. The top of the profile, from <code>results/profile_char_step.txt</code>, was taken at a load average of 252.</p><table><thead><tr><th>function</th><th>total time over 200 steps</th></tr></thead><tbody><tr><td><code>ufunc.at</code> (the scatter)</td><td>1.117 s</td></tr><tr><td><code>optim.py</code> AdamW step</td><td>1.030 s</td></tr><tr><td><code>tensor.py</code> matmul backward</td><td>0.401 s</td></tr><tr><td><code>tensor.py</code> matmul forward</td><td>0.392 s</td></tr></tbody></table><p>The scatter alone cost more than the whole optimizer. The char model scatters 4,096 rows (a batch of 256 times a 16-character context) into a 65-row table. I benchmarked two replacements in <code>scripts/bench_scatter.py</code>.</p><ul><li>The first is a one-hot matmul. Build an <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>N</mi><mo>×</mo><mi>V</mi></mrow><annotation encoding="application/x-tex">N \times V</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.7667em;vertical-align:-0.0833em;"></span><span class="mord mathnormal" style="margin-right:0.109em;">N</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">×</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal" style="margin-right:0.2222em;">V</span></span></span></span> one-hot matrix and compute <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mrow><mi mathvariant="normal">o</mi><mi mathvariant="normal">n</mi><mi mathvariant="normal">e</mi><mi mathvariant="normal">h</mi><mi mathvariant="normal">o</mi><mi mathvariant="normal">t</mi></mrow><mo stretchy="false">(</mo><mrow><mi mathvariant="normal">i</mi><mi mathvariant="normal">d</mi><mi mathvariant="normal">x</mi></mrow><msup><mo stretchy="false">)</mo><mi mathvariant="normal">⊤</mi></msup><mi>G</mi></mrow><annotation encoding="application/x-tex">\mathrm{onehot}(\mathrm{idx})^{\top} G</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1.0991em;vertical-align:-0.25em;"></span><span class="mord"><span class="mord mathrm">onehot</span></span><span class="mopen">(</span><span class="mord"><span class="mord mathrm">idx</span></span><span class="mclose"><span class="mclose">)</span><span class="msupsub"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.8491em;"><span style="top:-3.063em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mtight">⊤</span></span></span></span></span></span></span></span></span><span class="mord mathnormal">G</span></span></span></span>. It is a single BLAS call, and wasteful in FLOPs, but at <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>V</mi><mo>=</mo><mn>65</mn></mrow><annotation encoding="application/x-tex">V = 65</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal" style="margin-right:0.2222em;">V</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.6444em;"></span><span class="mord">65</span></span></span></span> the waste is irrelevant.</li><li>The second is sort and reduceat. Sort the indices, find the start of each run of equal indices, and sum each run with <code>np.add.reduceat</code>. It does no wasted arithmetic, but pays for a sort and a gather.</li></ul><figure data-figure="chart:projects/autograd-from-scratch/autograd-from-scratch-scatter-add"></figure><p>At the char model's shape the one-hot product took 72 µs against 1,272 µs for <code>np.add.at</code>, a 17.6x speedup, with sort and reduceat at 484 µs. The one-hot matrix grows as <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>N</mi><mo>×</mo><mi>V</mi></mrow><annotation encoding="application/x-tex">N \times V</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.7667em;vertical-align:-0.0833em;"></span><span class="mord mathnormal" style="margin-right:0.109em;">N</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">×</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal" style="margin-right:0.2222em;">V</span></span></span></span>, though, and at a GPT-2-sized vocabulary of 50,257 with 8,192 rows it would be over 400 million entries, so it is not even attempted there. At that shape sort and reduceat wins, 9,399 µs against 12,530 µs.</p><p>The shipped function picks by size.</p><figure class="highlight python"><table><tr><td class="code"><pre><span class="line"><span class="keyword">def</span> <span class="title function_">_scatter_add_rows</span>(<span class="params">idx, g, num_rows</span>):</span><br><span class="line">    n = idx.shape[<span class="number">0</span>]</span><br><span class="line">    <span class="keyword">if</span> n * num_rows &lt;= (<span class="number">1</span> &lt;&lt; <span class="number">20</span>):</span><br><span class="line">        onehot = np.zeros((n, num_rows), dtype=g.dtype)</span><br><span class="line">        onehot[np.arange(n), idx] = <span class="number">1</span></span><br><span class="line">        <span class="keyword">return</span> onehot.T @ g</span><br><span class="line">    order = np.argsort(idx, kind=<span class="string">&quot;stable&quot;</span>)</span><br><span class="line">    s = idx[order]</span><br><span class="line">    starts = np.flatnonzero(np.concatenate(([<span class="literal">True</span>], s[<span class="number">1</span>:] != s[:-<span class="number">1</span>])))</span><br><span class="line">    out = np.zeros((num_rows, g.shape[<span class="number">1</span>]), dtype=g.dtype)</span><br><span class="line">    out[s[starts]] = np.add.reduceat(g[order], starts, axis=<span class="number">0</span>)</span><br><span class="line">    <span class="keyword">return</span> out</span><br></pre></td></tr></table></figure><p>Two honest caveats. At the middle shape (<span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>V</mi><mo>=</mo><mn>1,000</mn></mrow><annotation encoding="application/x-tex">V = 1{,}000</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal" style="margin-right:0.2222em;">V</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.8389em;vertical-align:-0.1944em;"></span><span class="mord">1</span><span class="mord"><span class="mpunct">,</span></span><span class="mord">000</span></span></span></span>, <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>N</mi><mo>=</mo><mn>8,192</mn></mrow><annotation encoding="application/x-tex">N = 8{,}192</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal" style="margin-right:0.109em;">N</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.8389em;vertical-align:-0.1944em;"></span><span class="mord">8</span><span class="mord"><span class="mpunct">,</span></span><span class="mord">192</span></span></span></span>), the threshold sends the call to sort and reduceat even though one-hot had the lower minimum (1,794 µs against 1,981 µs). The medians rank them the other way (4,461 µs against 3,226 µs), so under this much load I cannot call it either way, and the threshold is a memory rule as much as a speed rule. Second, sort and reduceat is not bit-identical to <code>np.add.at</code>. Its maximum absolute error in float32 was 7.6e-6 at the small shape, because it sums each run in a different order. That is rounding, not a bug, and it foreshadows problem 6.</p><p>Rerunning the same profile after the change, the step went from 23.78 ms to 11.51 ms under the profiler and the scatter dropped out of the top ten. Not all of that halving is the scatter, though. The scatter accounted for 5.6 ms per step, and the rest of the 12.3 ms drop came from lines whose code did not change, such as the AdamW step falling from 1.030 s to 0.395 s over 200 steps. That part is load, and the honest claim is a 5.6 ms per step saving.</p><h3 id="5-adamw-s-temporaries">5. AdamW's temporaries</h3><p>In the second profile the optimizer is the top line. For a 535,818-parameter MLP at small batch it is a large, batch-independent share of every step, and the expression form of AdamW allocated about six parameter-sized arrays per tensor per step. The <code>out=</code> rewrite shown in <a href="#implementation">implementation</a> removed them. It did not close the gap to PyTorch's multi-tensor AdamW, which I return to in <a href="#results">results</a>.</p><h3 id="6-float32-drifts-away-from-pytorch-and-it-is-not-a-bug">6. float32 drifts away from PyTorch, and it is not a bug</h3><p>This is the one that looked like a real bug and took the most care to rule out.</p><p>The MNIST script trains a PyTorch twin from the same initial weights on the same batches and logs the maximum absolute weight difference after each epoch. In float32 it was 0.098 after one epoch and <strong>0.313</strong> after ten. That is not rounding error by any normal standard, yet every parity test passed. Either there was a bug that only shows up over thousands of steps, or training amplifies rounding differences into large weight differences.</p><p>Those two explanations make different predictions about precision. A real bug in a gradient or an update rule does not care about dtype. It should make the float64 twins diverge just as fast. Amplified rounding should shrink by roughly the ratio of the two machine epsilons, about nine orders of magnitude. So I reran the twin experiment in float64 for three epochs.</p><table><thead><tr><th>epoch</th><th>float64 max weight difference</th><th>smolgrad test accuracy</th><th>PyTorch test accuracy</th></tr></thead><tbody><tr><td>1</td><td>2.5e-14</td><td>96.81%</td><td>96.81%</td></tr><tr><td>2</td><td>2.3e-14</td><td>97.68%</td><td>97.68%</td></tr><tr><td>3</td><td>5.3e-14</td><td>97.75%</td><td>97.75%</td></tr></tbody></table><p>After 1,407 AdamW steps on 535,818 parameters, the two implementations agree to 5.3e-14. The epoch train losses agree to 15 significant figures. Test accuracy is identical at every epoch. There is no room for a bug in the forward pass, any gradient, or the optimizer on this path.</p><p>To see how the float32 gap grows, <code>scripts/divergence.py</code> runs the twins for 300 steps in both dtypes with both optimizers and logs the weight difference every step.</p><table><thead><tr><th>run</th><th>step 1</th><th>step 30</th><th>step 100</th><th>step 300</th></tr></thead><tbody><tr><td>float32, SGD with momentum</td><td>3.7e-9</td><td>7.5e-8</td><td>4.4e-3</td><td>2.6e-2</td></tr><tr><td>float32, AdamW</td><td>8.3e-6</td><td>8.4e-6</td><td>6.8e-4</td><td>6.7e-2</td></tr><tr><td>float64, SGD with momentum</td><td>6.9e-18</td><td>1.7e-16</td><td>4.4e-16</td><td>5.6e-16</td></tr><tr><td>float64, AdamW</td><td>9.7e-15</td><td>9.8e-15</td><td>9.6e-15</td><td>1.5e-14</td></tr></tbody></table><p>In float64 the difference stays at or below 1.5e-14 for 300 steps with either optimizer. In float32 it starts at the rounding level and does not grow smoothly. It jumps. The SGD run sits at 7.5e-8 through step 30 and is at 1.8e-6 one step later, 24 times larger. The AdamW run is flat at about 8.3e-6 for 77 steps, then climbs, and jumps again between steps 91 and 93. Discrete jumps like that are consistent with a ReLU unit whose pre-activation is near zero landing on different sides of zero in the two runs, after which the two networks are computing slightly different functions. I have not traced a specific unit, so that is an interpretation, not a finding.</p><p>The implementations round differently because they sum in different orders. numpy's reductions and Accelerate's GEMM do not accumulate in the same order as PyTorch's kernels, float32 addition is not associative, and training is a chaotic map that amplifies the difference.</p><p>AdamW starts noisier than SGD, 8.3e-6 after one step against 3.7e-9, and that has a clean explanation. Adam's update is <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>m</mi><mi mathvariant="normal">/</mi><mo stretchy="false">(</mo><msqrt><mi>v</mi></msqrt><mo>+</mo><mi>ϵ</mi><mo stretchy="false">)</mo></mrow><annotation encoding="application/x-tex">m / (\sqrt{v} + \epsilon)</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1.0503em;vertical-align:-0.25em;"></span><span class="mord mathnormal">m</span><span class="mord">/</span><span class="mopen">(</span><span class="mord sqrt"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.8003em;"><span class="svg-align" style="top:-3em;"><span class="pstrut" style="height:3em;"></span><span class="mord" style="padding-left:0.833em;"><span class="mord mathnormal" style="margin-right:0.0359em;">v</span></span></span><span style="top:-2.7603em;"><span class="pstrut" style="height:3em;"></span><span class="hide-tail" style="min-width:0.853em;height:1.08em;"><svg xmlns="http://www.w3.org/2000/svg" width="400em" height="1.08em" viewBox="0 0 400000 1080" preserveAspectRatio="xMinYMin slice"><path d="M95,702c-2.7,0,-7.17,-2.7,-13.5,-8c-5.8,-5.3,-9.5,-10,-9.5,-14c0,-2,0.3,-3.3,1,-4c1.3,-2.7,23.83,-20.7,67.5,-54c44.2,-33.3,65.8,-50.3,66.5,-51c1.3,-1.3,3,-2,5,-2c4.7,0,8.7,3.3,12,10s173,378,173,378c0.7,0,35.3,-71,104,-213c68.7,-142,137.5,-285,206.5,-429c69,-144,104.5,-217.7,106.5,-221l0 -0c5.3,-9.3,12,-14,20,-14H400000v40H845.2724s-225.272,467,-225.272,467s-235,486,-235,486c-2.7,4.7,-9,7,-19,7c-6,0,-10,-1,-12,-3s-194,-422,-194,-422s-65,47,-65,47zM834 80h400000v40h-400000z"/></svg></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.2397em;"><span></span></span></span></span></span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">+</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mord mathnormal">ϵ</span><span class="mclose">)</span></span></span></span>, which normalizes each coordinate's step toward the full learning rate regardless of the gradient's magnitude. For a weight whose true gradient is near zero, a rounding-level difference in that gradient becomes a difference of order the learning rate in the update. SGD scales the update by the gradient itself, so tiny differences stay tiny.</p><p>The practical takeaway is that float32 parity with a reference implementation is only meaningful over a handful of steps, and long-run parity has to be tested in float64. The ten-epoch float32 twins still end at 97.89% and 97.81%, which is the level at which they should agree. The same sensitivity shows up between thread configurations, not just between implementations, as the MNIST results below describe.</p><h3 id="7-the-machine-was-shared">7. the machine was shared</h3><p>Every timing in this project was taken while other jobs were running on the same 14-core M4 Pro. The results files record the one-minute load average at each measurement, and it was 338 to 358 for the single-threaded benchmark, 358 to 360 for the default-threads benchmark, 194 to 202 for the scatter benchmark and 252 for the profile. A load of 300 on 14 cores means each runnable process waits most of the time.</p><p>The benchmark interleaves the implementations round-robin, so any burst of load hits all three alike, reports the minimum over all runs as the best estimate of intrinsic cost, since load can only add time, and records the load average in every results file. The medians in the CSVs measure the other jobs more than smolgrad, and I do not quote them. The absolute times below are upper bounds.</p><p>How much load distorted things shows in the training runs. The MNIST script's first epoch took 19.2 s and its tenth took 4.5 s with no code change, and the character model averaged 71 ms per step during training against a benchmark minimum of 3.4 ms.</p><h2 id="experiments">experiments</h2><ol><li>Correctness is checked by <code>pytest -q</code>, 101 tests.</li><li>MNIST uses a 784-512-256-10 ReLU MLP, 535,818 parameters, AdamW at learning rate 1e-3 and weight decay 0.01, batch 128, 10 epochs in float32, with a PyTorch twin from identical weights on identical batches. Then the same twin setup in float64 for 3 epochs.</li><li>The drift experiment runs 300 steps for each of float32 and float64 crossed with SGD with momentum and AdamW, logging the weight difference to the PyTorch twin every step.</li><li>The character model is a Bengio-style character MLP on Tiny Shakespeare. An <code>Embedding(65, 32)</code> over a 16-character context is flattened to 512, then <code>Linear(512, 512)</code>, <code>LayerNorm</code>, tanh, <code>Linear(512, 65)</code>. That is 299,105 parameters, trained with AdamW at 3e-3 with a linear decay over the last quarter, batch 256, 6,000 steps, and compared against uniform, unigram and bigram baselines on the same validation split.</li><li>The step-time benchmark times a full training step (forward, backward, AdamW) of the MNIST MLP at batch 1 to 2048 and the character model at batch 256, single-threaded and with default threading, in three implementations. They are smolgrad, PyTorch eager on CPU, and the same MLP hand-written in numpy with manual gradients, no graph, and smolgrad's own AdamW. A microbenchmark also times 2,000 chained multiplies on one-element tensors, forward plus backward.</li></ol><p>The hand-written numpy MLP is the design choice that makes the benchmark worth reading. Without it, &quot;1.8x slower than PyTorch&quot; could mean a slow graph engine or slow kernels. With it, smolgrad minus numpy is the price of autograd, and numpy minus PyTorch is the price of numpy.</p><h2 id="results">results</h2><h3 id="correctness">correctness</h3><p>All 101 tests pass. The float64 twin result above is the strongest end-to-end evidence, since it exercises every op in the MLP, the fused cross-entropy, the engine's accumulation and AdamW over 1,407 steps at once.</p><h3 id="mnist">MNIST</h3><table><thead><tr><th></th><th>smolgrad</th><th>PyTorch twin</th></tr></thead><tbody><tr><td>final test accuracy, epoch 10</td><td>97.89%</td><td>97.81%</td></tr><tr><td>final test loss</td><td>0.0863</td><td>0.0946</td></tr><tr><td>final train loss</td><td>0.0184</td><td>0.0187</td></tr></tbody></table><p>I report only the final epoch. <code>results/mnist_summary.json</code> also records a best test accuracy of 98.11% at epoch 5, but there is no validation split, so picking the best epoch means picking it on the test set, and that number is optimistic by construction. A separate validation split would fix this.</p><p>The float32 run is deterministic for a fixed thread count but not across thread counts. An independent rerun with BLAS capped at 2 threads matched itself exactly, and reached 98.07% final accuracy with a final test loss of 0.0778, against 97.89% and 0.0863 here. The runs already differ after the first epoch (96.96% against 96.77%), which points to summation order in the matrix multiplies rather than a bug, and the float64 twin rules out a gradient error. So read 97.89% as one draw from a spread of about 0.2 points.</p><p>Test loss bottoms out at epoch 5 at 0.0623 and rises afterwards while train loss keeps falling, which is ordinary overfitting with no regularization beyond weight decay. The 0.08-point accuracy difference between the twins is float32 drift, not a property of either implementation. At epoch 8 the ranking was the other way around, 97.51% against 97.88%.</p><h3 id="tiny-shakespeare">Tiny Shakespeare</h3><figure data-figure="chart:projects/autograd-from-scratch/autograd-from-scratch-char-loss"></figure><p>The character model reaches 1.794 nats per character on the validation split after 6,000 steps, 0.688 below the bigram baseline. Train loss at the end was 1.576, so the model is starting to overfit its 16-character windows, and validation loss was still falling at the last evaluation, helped by the learning-rate decay. Samples at temperature 0.8 get the play format right, with speaker names and line breaks, mostly real short words and invented longer ones.</p><figure class="highlight plaintext"><table><tr><td class="code"><pre><span class="line">KING RICHARD III:</span><br><span class="line">Ay, I woonesh thee, I say though it east hear good.</span><br></pre></td></tr></table></figure><p>That is about what a 300,000-parameter model with a 16-character window should manage.</p><h3 id="step-time-against-pytorch">step time against PyTorch</h3><figure data-figure="chart:projects/autograd-from-scratch/autograd-from-scratch-step-ratio"></figure><p>Single-threaded, taking the minimum over 320 steps for batches up to 128 and 120 steps above, at a load average of 338 to 358.</p><table><thead><tr><th>MLP batch</th><th>smolgrad</th><th>numpy by hand</th><th>PyTorch</th><th>smolgrad / PyTorch</th></tr></thead><tbody><tr><td>1</td><td>3.21 ms</td><td>2.32 ms</td><td>1.29 ms</td><td>2.5x</td></tr><tr><td>32</td><td>2.43 ms</td><td>1.62 ms</td><td>1.32 ms</td><td>1.8x</td></tr><tr><td>128</td><td>2.94 ms</td><td>2.09 ms</td><td>1.66 ms</td><td>1.8x</td></tr><tr><td>512</td><td>5.01 ms</td><td>4.97 ms</td><td>3.00 ms</td><td>1.7x</td></tr><tr><td>2048</td><td>21.2 ms</td><td>38.1 ms</td><td>11.6 ms</td><td>1.8x</td></tr></tbody></table><p>The character model step at batch 256 was 3.40 ms for smolgrad and 2.77 ms for PyTorch, a ratio of 1.2x.</p><p>The batch 2048 row, where hand-written numpy is 17 ms slower than smolgrad, is load noise. The two run the same matmuls. I keep the row because deleting inconvenient measurements is worse than labelling them, and its other numbers deserve the same suspicion.</p><h3 id="where-the-gap-comes-from">where the gap comes from</h3><figure data-figure="chart:projects/autograd-from-scratch/autograd-from-scratch-phases"></figure><p>At batch 128 the phases split like this.</p><table><thead><tr><th>phase</th><th>smolgrad</th><th>numpy by hand</th><th>PyTorch</th></tr></thead><tbody><tr><td>forward</td><td>0.42 ms</td><td>0.36 ms</td><td>0.31 ms</td></tr><tr><td>backward</td><td>1.15 ms</td><td>0.38 ms</td><td>0.35 ms</td></tr><tr><td>AdamW step</td><td>1.33 ms</td><td>1.32 ms</td><td>1.00 ms</td></tr></tbody></table><p>My reading of it, with the usual hedge that these are minima under heavy load and that phase minima do not add exactly to step minima.</p><p>The autograd machinery costs about 0.85 ms per step (2.94 minus 2.09), and almost all of it is in backward, 1.15 ms against 0.38 ms. The forward passes are within 0.06 ms of each other, so recording the graph is cheap. The backward cost is extra memory traffic. The hand-written version computes the relu mask and applies it in one expression, sums the bias gradient directly, and never builds a dict or checks an ownership flag. smolgrad runs <code>_unbroadcast</code> on every bias add, allocates a new array for every relu mask multiply, allocates again when summing gradients that arrive at the same node, and copies wherever a leaf cannot adopt its gradient.</p><p>The kernel gap is smaller, about 0.43 ms (2.09 minus 1.66), and most of it is the optimizer. PyTorch's <code>foreach</code> AdamW updates all parameters with multi-tensor kernels, while mine makes about ten numpy passes over each parameter. At batch 128, AdamW is 1.33 of smolgrad's 2.94 ms, or 45% of the step, and because its cost depends only on the parameter count it takes a larger share the smaller the batch.</p><p>Per-op Python overhead is not the problem. On one-element tensors, a forward plus backward multiply costs 6.2 µs in smolgrad and 12.6 µs in PyTorch single-threaded, and 6.2 µs against 10.0 µs with default threading. PyTorch's dispatcher does more work per op than a Python closure does. This surprised me, and it means the gap on real workloads is about how many bytes move, not how many Python calls happen.</p><p>The default-threads benchmark (40 or 15 steps per point, at a load average near 360) is less reliable. In it PyTorch at batch 128 is slower than hand-written numpy, 5.24 ms against 2.25 ms, which is what ten PyTorch threads on an oversubscribed machine look like. I use it only for the per-op overhead numbers.</p><h2 id="what-i-would-change">what I would change</h2><h3 id="record-ops-not-closures">record ops, not closures</h3><p>Closures made v0 short, and they make the graph a black box. There is no way to look at a chain of <code>add</code>, <code>relu</code>, <code>mul</code> and fuse it. The planned lazy-graph milestone needs an explicit op record (type, inputs, attributes), and I would switch to that first.</p><h3 id="fuse-the-obvious-hot-spots-before-writing-a-backend">fuse the obvious hot spots before writing a backend</h3><p>The phase table says where the 1.8x lives. A fused bias-plus-relu with an in-place mask in backward attacks the 0.77 ms of backward overhead directly. A multi-tensor AdamW over one flat parameter buffer would replace about ten numpy passes over each of the MLP's six parameter tensors with about ten passes over one buffer, and should attack most of the 0.33 ms optimizer gap. A fused LayerNorm forward and backward would replace ten graph nodes with one in the character model. All three are possible in numpy without writing a kernel.</p><h3 id="add-a-version-counter">add a version counter</h3><p>Closures read their parents' <code>.data</code> lazily at backward time, so updating a parameter in place between forward and backward silently corrupts the gradients. PyTorch catches this with a version integer on every tensor. v0 relies on the convention that only the optimizer mutates in place, and it runs after backward.</p><h3 id="write-backward-rules-in-tensor-ops-for-higher-order-gradients">write backward rules in Tensor ops for higher-order gradients</h3><p>Right now each VJP is raw numpy, so the backward pass builds no graph and grad-of-grad is impossible. Supporting it means a second set of backward rules written in Tensor ops, behind a <code>create_graph=True</code> flag so the first-order path does not pay for it.</p><h3 id="rerun-every-timing-on-an-idle-machine">rerun every timing on an idle machine</h3><p><code>scripts/run_bench_when_quiet.sh</code> already waits for the load average to drop below a threshold. The numbers here were taken at a one-minute load average of 194 to 360 and are upper bounds, with the single-threaded ratios roughly right. The float64 parity result does not depend on the machine at all, which is why it leads this post.</p><h2 id="reproducibility">reproducibility</h2><p>Needs <a href="https://docs.astral.sh/uv/">uv</a>. Python is pinned to 3.12. From <code>projects/11-smolgrad</code>, the following commands set up the environment, run the tests, download the data and reproduce every experiment.</p><figure class="highlight bash"><table><tr><td class="code"><pre><span class="line">uv <span class="built_in">sync</span></span><br><span class="line">uv run pytest -q</span><br><span class="line"></span><br><span class="line">uv run python scripts/download_data.py              <span class="comment"># MNIST (SHA-256 checked) and Tiny Shakespeare, about 13 MB</span></span><br><span class="line">uv run python scripts/train_mnist.py --torch-twin   <span class="comment"># 10 epochs, float32, with the PyTorch twin</span></span><br><span class="line">uv run python scripts/train_char.py                 <span class="comment"># 6,000 steps</span></span><br><span class="line">uv run python scripts/divergence.py                 <span class="comment"># float32 vs float64 drift, 300 steps each</span></span><br><span class="line"></span><br><span class="line">uv run python scripts/bench.py --threads 1          <span class="comment"># single-threaded step-time benchmark</span></span><br><span class="line">uv run python scripts/bench.py                      <span class="comment"># default threading</span></span><br><span class="line">uv run python scripts/plot_bench.py</span><br><span class="line">uv run python scripts/bench_scatter.py              <span class="comment"># embedding backward strategies</span></span><br><span class="line">uv run python scripts/profile_step.py               <span class="comment"># cProfile of a char-model step</span></span><br><span class="line"></span><br><span class="line"><span class="comment"># on a shared machine, wait for load average &lt; LOAD_MAX (default 20) first</span></span><br><span class="line">scripts/run_bench_when_quiet.sh</span><br></pre></td></tr></table></figure><p>The float64 MNIST twin run is the command below, and it writes <code>results/mnist_train_f64_twin.csv</code> and <code>results/mnist_summary_f64_twin.json</code>.</p><figure class="highlight bash"><table><tr><td class="code"><pre><span class="line">uv run python scripts/train_mnist.py --torch-twin --dtype float64 --epochs 3 --tag _f64_twin</span><br></pre></td></tr></table></figure><p>Every number in this post comes from a file in <code>results/</code>, which holds the raw CSV, JSON and profile output of each run, each timing file stamped with the load average at the time. The exceptions are the rerun figures from the independent verification, which are recorded in <code>DEVLOG.md</code>.</p><h3 id="verification">verification</h3><p>An independent reviewer who did not build the project reran it on 2026-09-26, with BLAS threads capped at 2 on a shared machine and every output written to a scratch folder. All 101 tests passed. The float64 MNIST twin reproduced, with 97.75% test accuracy for both implementations, every loss matching the committed CSV to about 1e-16, and a final weight difference of 3.9e-14 against the committed 5.3e-14, both at rounding level. The full character run matched <code>results/char_summary.json</code> to 1e-7, with validation loss 1.794 and the same three baselines. The drift experiment matched in float64 and for float32 SGD, while float32 AdamW agreed at step 1 and then drifted to a different end value, as expected for the chaotic float32 behavior described in problem 6. The float32 MNIST run reproduced only up to the thread-count spread described in the MNIST results. The reviewer also checked that the hand-written numpy baseline computes the same loss and gradients as smolgrad, and corrected the README, which had claimed checksums for both datasets when only MNIST is checked. Timing was deferred. The benchmark was only checked to run, and at a load average near 9 an MNIST epoch took about 1.0 s and the character run 17 s, against 4.5 to 19.2 s and 478 s here, so every timing in this post still needs a quiet machine.</p><p>Code is in <code>projects/11-smolgrad</code>, with the library in <code>src/smolgrad/</code>, the tests in <code>tests/</code>, and the design notes and full build log in <code>DESIGN.md</code> and <code>DEVLOG.md</code>.</p>]]>
    </content>
    <id>https://projects.farhansadeek.com/posts/smolgrad/</id>
    <link href="https://projects.farhansadeek.com/posts/smolgrad/"/>
    <published>2026-09-29T06:17:10.000Z</published>
    <summary>smolgrad is a small reverse-mode autodiff engine and neural network library. In float64 it tracks PyTorch to 13 decimal places across three epochs of MNIST.</summary>
    <title>smolgrad, a NumPy Autograd Engine Checked Against PyTorch</title>
    <updated>2026-09-29T06:49:26.552Z</updated>
  </entry>
  <entry>
    <author>
      <name>Farhan Sadeek</name>
      <email>farhan@farhansadeek.com</email>
    </author>
    <category term="cpp" scheme="https://projects.farhansadeek.com/tags/cpp/"/>
    <category term="trading" scheme="https://projects.farhansadeek.com/tags/trading/"/>
    <category term="market-microstructure" scheme="https://projects.farhansadeek.com/tags/market-microstructure/"/>
    <category term="testing" scheme="https://projects.farhansadeek.com/tags/testing/"/>
    <content>
      <![CDATA[<p>I wrote the core of a small exchange in C++20, a limit order book and matching engine for one symbol on one thread, with price-time priority, five order types, cancel, and modify with the usual queue position rules. Every input and output is a plain event, so the engine is a deterministic state machine that can be replayed from its log.</p><p>The part I trust most is not the engine but the test that watches it. A second matcher, written to be obviously correct rather than fast, processes the same random order flow, and the two are compared on every output event and periodically on the full priority-ordered contents of the book. Across <strong>8.3 million random events</strong> the two never disagreed. On 2026-09-27 an independent rerun with a seed I had not used (seed 11, 200,000 events) also matched with <strong>zero divergence</strong>, and a clean Release rebuild passed all <strong>39 tests</strong>.</p><p>A fuzzer that passes proves little unless it can fail. So I planted six known bugs, one at a time, in copies of the engine and ran the fuzzer against each. It caught <strong>6 of 6</strong>, the slowest within 3,487 events. The most instructive one, a modify that grows an order without sending it to the back of the queue, produced exactly the right output at the moment it happened and was only visible when the book's internal order was compared. That result is the reason the test compares state as well as outputs.</p><p>The latency numbers are less solid. They were measured on a shared Apple M4 Pro at 1 minute load averages of 172 to 352 on 14 cores, not rerun independently, and quantized by a 41.7 ns timer tick. With 10,000 resting orders the median add took 167 ns, cancel 209 ns and a single-fill match 84 ns, which I treat as indicative upper bounds.</p><p>Code is in <code>projects/06-price-time-matching-engine</code>. Skip to <a href="#problems">problems</a> for the bugs.</p><h2 id="table-of-contents">table of contents</h2><ul><li><a href="#what-i-wanted-to-build">what I wanted to build</a></li><li><a href="#theory">theory</a></li><li><a href="#architecture">architecture</a></li><li><a href="#implementation">implementation</a></li><li><a href="#problems">problems</a></li><li><a href="#experiments">experiments</a></li><li><a href="#results">results</a></li><li><a href="#what-i-would-change">what I would change</a></li><li><a href="#reproducibility">reproducibility</a></li></ul><h2 id="what-i-wanted-to-build">what I wanted to build</h2><p>A matching engine holds resting buy and sell orders, decides which ones trade when a new order arrives, and tells the world what happened. I wanted one built the way a real venue would build it, with a v0 scope fixed up front.</p><ul><li>One symbol, one thread, a limit order book with price-time priority.</li><li>Limit, market, IOC, FOK and post-only orders, plus cancel and modify, where a modify that increases size or changes price loses its place in the queue.</li><li>Integer prices and quantities, with no floating point anywhere in the engine.</li><li>An event-sourced interface, so every input and output can be logged and replayed.</li><li>An L2 market-data feed and a trade tape, and a replay tool that proves the engine is deterministic.</li><li>Randomized differential tests against a naive reference matcher on millions of orders.</li><li>Latency percentiles (p50, p99, p99.9) for the add, cancel and match paths, and throughput.</li></ul><p>Everything on that list got built. Multiple symbols, opening and closing auctions, self-trade prevention and a network gateway are the next milestones. Later projects in this series are meant to consume this engine's event and market-data streams.</p><h2 id="theory">theory</h2><h3 id="price-time-priority">price-time priority</h3><p>A limit order says &quot;buy up to Q at price P or better&quot;. An order that cannot trade immediately rests in the book. The book has two sides, bids sorted from the highest price down and asks from the lowest price up. The best bid and best ask are the touch, and the gap between them is the spread.</p><p>When an aggressive order arrives, price-time priority decides who trades with it. The best price always goes first. Among orders at the same price, the one that arrived first goes first. The trade prints at the resting order's price, so the aggressor keeps any price improvement.</p><h3 id="what-each-order-type-does-with-its-leftovers">what each order type does with its leftovers</h3><p>The five order types differ only in what happens to the part that does not fill.</p><table><thead><tr><th>Type</th><th>Price bound</th><th>Unfilled remainder</th></tr></thead><tbody><tr><td>limit</td><td>yes</td><td>rests until cancelled</td></tr><tr><td>market</td><td>none</td><td>cancelled</td></tr><tr><td>IOC</td><td>yes</td><td>cancelled</td></tr><tr><td>FOK</td><td>yes</td><td>the whole order is cancelled unless it can fill completely</td></tr><tr><td>post-only</td><td>yes</td><td>rests, and the order is rejected if it would trade on arrival</td></tr></tbody></table><p>FOK must check available quantity within its limit before it prints a single trade, because trades are irrevocable. Post-only lets a market maker be sure it pays the maker fee, and I reject it even when it would only lock the touch, since at that price it would trade.</p><h3 id="why-modify-has-to-cost-queue-position">why modify has to cost queue position</h3><p>At a given price the front of the queue fills first. If an order could grow without losing its place, a trader could sit at the front with one lot and increase it only when they liked the market. So reducing size at the same price keeps priority, and increasing size or changing price sends the order to the back, as if it had been cancelled and re-entered. Breaking this rule produces a bug that is invisible in the output, as problem 1 shows.</p><h3 id="event-sourcing-and-determinism">event sourcing and determinism</h3><p>If the engine has no hidden inputs, no clock, no randomness and no dependence on thread timing, then its state is a pure function of its input log. Replaying the log into a fresh engine reproduces every output, including trade ids, bit for bit. That is how exchanges do failover and audit, and any failure becomes a log file that reproduces it. Time priority does not need a wall clock either, because arrival order is exactly the sequence number the engine assigns.</p><h3 id="fixed-point">fixed point</h3><p>A price level is looked up by price, so prices must compare exactly, which rules out <code>double</code>. A price is an <code>int64</code> count of ticks, a quantity is an <code>int64</code> count of lots, and decimal strings are converted only at the edges. The parser refuses input that would need rounding, so <code>101.255</code> at a scale of two decimals is an error, not 101.26. Quantities are capped at 2^40, which keeps a level's aggregate quantity far from <code>int64</code> overflow.</p><h3 id="differential-testing">differential testing</h3><p>Unit tests check the cases I thought of, and the bugs that hurt live in combinations I did not, such as a modify of a partly filled order at a level about to empty. Differential testing writes the specification twice, once fast and once so simply it is hard to get wrong, feeds both the same random inputs and compares everything they produce. The two share no data structures, so a bug in the engine's incremental bookkeeping has nothing in the oracle to hide behind. The weak point is that the comparison only sees what it compares, and a bug that changes internal state without changing any output yet can hide for a long time.</p><h2 id="architecture">architecture</h2><p>The engine is a single-threaded state machine. Inputs arrive as <code>InputEvent</code>s (new, cancel, modify), each gets the next sequence number, and the engine appends <code>OutputEvent</code>s to a vector the caller owns. There are no callbacks or virtual calls in the hot path.</p><figure data-figure="diagram:matching-engine"></figure><p>For one input the outputs always come in the same order. First exactly one of <code>ACCEPTED</code>, <code>REJECTED</code> or <code>MODIFIED</code>. Then trades, in match order. Then a <code>CANCELED</code> for any unfilled market, IOC or FOK remainder. Then <code>L2</code> book updates, one per price level whose aggregate changed, bids first and then asks, each by ascending price. Here is part of the hand-written golden session in <code>examples/</code>, with each input shown as a comment above the outputs it caused.</p><figure class="highlight plaintext"><table><tr><td class="code"><pre><span class="line"># N 8 B LMT 10020 12   aggressive buy, takes 10010 and part of 10020</span><br><span class="line">7 ACCEPTED 8 B 10020 12</span><br><span class="line">7 TRADE 1 taker=8 maker=4 B 10010 10</span><br><span class="line">7 TRADE 2 taker=8 maker=5 B 10020 2</span><br><span class="line">7 L2 S 10010 0</span><br><span class="line">7 L2 S 10020 13</span><br><span class="line"># N 10 S FOK 9990 100  cannot fill completely, so no trades at all</span><br><span class="line">9 ACCEPTED 10 S 9990 100</span><br><span class="line">9 CANCELED 10 S 9990 100 fok_unfilled</span><br></pre></td></tr></table></figure><p>L2 updates carry the new absolute quantity at a level rather than a delta, so a late joiner only needs a snapshot and a dropped message cannot corrupt every later view. The <code>MarketDataFeed</code> rebuilds the public book purely from these events, and tests check that it always equals the engine's real book.</p><p>The <code>exch</code> tool wraps this in four commands, <code>gen</code>, <code>replay</code>, <code>verify</code> and <code>fuzz</code>.</p><h2 id="implementation">implementation</h2><h3 id="the-book">the book</h3><p>Each side is a <code>std::vector&lt;Level&gt;</code> sorted from the worst price to the best, so the best level is always <code>back()</code>. Most activity happens near the touch, so matching reads <code>back()</code> and removes an emptied level with <code>pop_back()</code>, both O(1). Inserting a level deep in the book costs a <code>memmove</code>, which I accepted. Finding a level scans the last 8 entries and falls back to binary search.</p><figure data-figure="diagram:order-book-levels"></figure><p>Each level owns a FIFO queue, an intrusive doubly linked list through a pool, which is a <code>std::vector&lt;Order&gt;</code> plus a free list. Orders link by 32-bit slot index rather than pointer, which keeps an order at 40 bytes and keeps links valid when the vector grows. Once the engine knows an order's slot, cancelling it from the middle of a queue is O(1).</p><h3 id="the-matching-loop">the matching loop</h3><p>This is the whole matching path, trimmed, with the construction of the trade event folded into a helper. It walks the opposite side from the best level inward and each level's queue from the head.</p><figure class="highlight cpp"><table><tr><td class="code"><pre><span class="line"><span class="keyword">while</span> (qty &gt; <span class="number">0</span> &amp;&amp; !book.<span class="built_in">empty</span>()) &#123;</span><br><span class="line">    Level&amp; L = book.<span class="built_in">back</span>();                       <span class="comment">// best opposite level</span></span><br><span class="line">    <span class="keyword">if</span> (has_limit &amp;&amp; !<span class="built_in">crosses</span>(side, limit, L.price)) <span class="keyword">break</span>;</span><br><span class="line">    <span class="built_in">touch</span>(opp, L.price, L.total);                 <span class="comment">// remember it for L2</span></span><br><span class="line">    <span class="keyword">while</span> (qty &gt; <span class="number">0</span> &amp;&amp; L.head != kNil) &#123;</span><br><span class="line">        Order&amp; m = pool_[L.head];                 <span class="comment">// oldest order at this price</span></span><br><span class="line">        <span class="type">const</span> Qty fill = std::<span class="built_in">min</span>(qty, m.qty);</span><br><span class="line">        out.<span class="built_in">push_back</span>(<span class="built_in">trade</span>(taker, m.id, next_trade_++, side, L.price, fill));</span><br><span class="line">        m.qty -= fill; L.total -= fill; qty -= fill;</span><br><span class="line">        <span class="keyword">if</span> (m.qty == <span class="number">0</span>) &#123;                         <span class="comment">// maker done: pop the queue head</span></span><br><span class="line">            <span class="type">const</span> std::<span class="type">uint32_t</span> mi = L.head;</span><br><span class="line">            L.head = m.next;</span><br><span class="line">            <span class="keyword">if</span> (L.head != kNil) pool_[L.head].prev = kNil; <span class="keyword">else</span> L.tail = kNil;</span><br><span class="line">            --L.count;</span><br><span class="line">            ids_.<span class="built_in">erase</span>(m.id);</span><br><span class="line">            free_.<span class="built_in">push_back</span>(mi);</span><br><span class="line">        &#125;</span><br><span class="line">    &#125;</span><br><span class="line">    <span class="keyword">if</span> (L.head == kNil) book.<span class="built_in">pop_back</span>();</span><br><span class="line">&#125;</span><br><span class="line"><span class="keyword">return</span> qty;                                       <span class="comment">// the caller decides: rest or cancel</span></span><br></pre></td></tr></table></figure><p>The trade prints at <code>L.price</code>, the maker's price. One of the planted bugs later changed exactly that line to print at the taker's limit, and the fuzzer caught it on the ninth event.</p><p>FOK runs a read-only pass first. <code>fillable</code> sums <code>total</code> over crossing levels from the best inward, and matching starts only if the sum reaches the order size, which is why the failed FOK above produced no trades.</p><h3 id="modify">modify</h3><p>Modify is where the priority rule lives, in one condition.</p><figure class="highlight cpp"><table><tr><td class="code"><pre><span class="line"><span class="keyword">if</span> (in.price == o.price &amp;&amp; in.qty &lt;= o.qty) &#123;</span><br><span class="line">    <span class="comment">// Same price, size not increased: amend in place, keep queue position.</span></span><br><span class="line">    <span class="built_in">touch</span>(side, L.price, L.total);</span><br><span class="line">    L.total -= o.qty - in.qty;</span><br><span class="line">    o.qty = in.qty;</span><br><span class="line">    out.<span class="built_in">push_back</span>(m);</span><br><span class="line">    <span class="keyword">return</span>;</span><br><span class="line">&#125;</span><br><span class="line"><span class="comment">// Price change or size increase: loses priority. Leave the book and</span></span><br><span class="line"><span class="comment">// re-enter as a fresh limit order under the same id, which can trade.</span></span><br><span class="line"><span class="built_in">remove</span>(oi);</span><br><span class="line">out.<span class="built_in">push_back</span>(m);</span><br><span class="line"><span class="type">const</span> Qty left = <span class="built_in">match</span>(id, side, <span class="literal">true</span>, in.price, in.qty, out);</span><br><span class="line"><span class="keyword">if</span> (left &gt; <span class="number">0</span>) <span class="built_in">rest</span>(id, side, in.price, left);</span><br></pre></td></tr></table></figure><p>A modify across the spread therefore trades immediately, with the modified order as taker. A modified order always becomes a plain limit order, so a post-only order modified across the spread trades instead of being rejected, a simplification rather than a feature.</p><h3 id="l2-coalescing">L2 coalescing</h3><p>Before the engine changes any level it calls <code>touch(side, price, quantity_before)</code>, which records the level once per input and ignores later calls for the same level. After the input is handled, it sorts the touched list into canonical order and emits an update for each level whose quantity now differs from the recorded value.</p><figure class="highlight cpp"><table><tr><td class="code"><pre><span class="line"><span class="keyword">for</span> (<span class="type">const</span> Touched&amp; t : touched_) &#123;</span><br><span class="line">    <span class="type">const</span> Qty now = <span class="built_in">level_qty</span>(t.side, t.price);</span><br><span class="line">    <span class="keyword">if</span> (now == t.before) <span class="keyword">continue</span>;   <span class="comment">// net no-op, publish nothing</span></span><br><span class="line">    out.<span class="built_in">push_back</span>(<span class="built_in">book_update</span>(seq_, t.side, t.price, now));</span><br><span class="line">&#125;</span><br><span class="line">touched_.<span class="built_in">clear</span>();</span><br></pre></td></tr></table></figure><p>A market order that eats ten orders at one price produces one L2 message, and an input that leaves a level's total unchanged produces none. The cost is that every code path that changes a level must remember to call <code>touch</code> first. Forgetting it in one branch is another planted bug, and it was the hardest of the six for the fuzzer to hit.</p><h3 id="the-id-index">the id index</h3><p>Cancel and modify arrive with an order id, so the engine needs a map from id to pool slot. I wrote <code>IdMap</code>, an open addressing table with linear probing, a power-of-two capacity, a load factor of at most one half, a splitmix64 hash so that sequential ids spread out, and backward shift deletion so that there are no tombstones. The deletion is the interesting part.</p><figure class="highlight cpp"><table><tr><td class="code"><pre><span class="line"><span class="comment">// After finding the key at slot i, pull later entries of the probe run</span></span><br><span class="line"><span class="comment">// back into the hole whenever their home slot allows it.</span></span><br><span class="line">std::<span class="type">size_t</span> hole = i, j = i;</span><br><span class="line"><span class="keyword">while</span> (<span class="literal">true</span>) &#123;</span><br><span class="line">    j = (j + <span class="number">1</span>) &amp; mask_;</span><br><span class="line">    <span class="keyword">if</span> (slots_[j].key == <span class="number">0</span>) <span class="keyword">break</span>;               <span class="comment">// end of the probe run</span></span><br><span class="line">    std::<span class="type">size_t</span> h = <span class="built_in">home</span>(slots_[j].key);</span><br><span class="line">    <span class="type">bool</span> movable = (hole &lt;= j) ? (h &lt;= hole || h &gt; j) : (h &lt;= hole &amp;&amp; h &gt; j);</span><br><span class="line">    <span class="keyword">if</span> (movable) &#123; slots_[hole] = slots_[j]; hole = j; &#125;</span><br><span class="line">&#125;</span><br><span class="line">slots_[hole] = Slot&#123;&#125;;</span><br></pre></td></tr></table></figure><p>Deletion costs time proportional to the rest of the probe run. With a good hash, runs are short. Problem 3 found a way to make one run as long as the table. The table has its own randomized test against <code>std::unordered_map</code>, 400,000 operations on a small key space to force wraparound and deletions from the middle of probe chains.</p><h3 id="the-reference-matcher">the reference matcher</h3><p>The oracle keeps every resting order in one flat vector with an arrival stamp. Finding the next maker is a linear scan.</p><figure class="highlight cpp"><table><tr><td class="code"><pre><span class="line"><span class="function"><span class="type">int</span> <span class="title">ReferenceMatcher::best_opposite</span><span class="params">(Side taker_side)</span> <span class="type">const</span> </span>&#123;</span><br><span class="line">    <span class="type">const</span> Side want = <span class="built_in">opposite</span>(taker_side);</span><br><span class="line">    <span class="type">int</span> best = <span class="number">-1</span>;</span><br><span class="line">    <span class="keyword">for</span> (std::<span class="type">size_t</span> i = <span class="number">0</span>; i &lt; orders_.<span class="built_in">size</span>(); ++i) &#123;</span><br><span class="line">        <span class="type">const</span> Resting&amp; o = orders_[i];</span><br><span class="line">        <span class="keyword">if</span> (o.side != want) <span class="keyword">continue</span>;</span><br><span class="line">        <span class="keyword">if</span> (best &lt; <span class="number">0</span>) &#123; best = <span class="built_in">int</span>(i); <span class="keyword">continue</span>; &#125;</span><br><span class="line">        <span class="type">const</span> Resting&amp; b = orders_[best];</span><br><span class="line">        <span class="type">const</span> <span class="type">bool</span> better_price = want == Side::Buy ? o.price &gt; b.price : o.price &lt; b.price;</span><br><span class="line">        <span class="keyword">if</span> (better_price || (o.price == b.price &amp;&amp; o.time &lt; b.time)) best = <span class="built_in">int</span>(i);</span><br><span class="line">    &#125;</span><br><span class="line">    <span class="keyword">return</span> best;</span><br><span class="line">&#125;</span><br></pre></td></tr></table></figure><p>That function is the definition of price-time priority with no queues or levels to get wrong. The oracle's L2 updates come from diffing full <code>std::map</code> snapshots of the book taken before and after each input, sharing none of the engine's touched-level logic. The price is speed, since those two snapshots per event dominate every fuzz run.</p><h3 id="the-order-flow-generator">the order flow generator</h3><p><code>OrderFlow</code> produces a deterministic stream from a seed with its own splitmix64 generator, because <code>std::uniform_int_distribution</code> is not specified bit for bit across standard libraries. By default about 30 percent of events are cancels and 10 percent modifies. New orders are 3 percent market, 7 percent IOC, 5 percent FOK and 10 percent post-only, the rest plain limits, priced in a band around a mid that drifts one tick at a time, and half a percent of events are deliberately invalid. The generator never looks at the book, so it naturally cancels filled orders, modifies dead ones and reuses live ids, which exercises every reject path for free.</p><h3 id="the-fuzz-loop">the fuzz loop</h3><p>The comparison itself is short. Every output event must match, and every <code>check_every</code> events the full state must match too.</p><figure class="highlight cpp"><table><tr><td class="code"><pre><span class="line">eng.<span class="built_in">process</span>(e, a_out);</span><br><span class="line">ref.<span class="built_in">process</span>(e, b_out);</span><br><span class="line"><span class="keyword">if</span> (a_out != b_out) &#123; <span class="built_in">print_divergence</span>(seed, i, e, a_out, b_out); <span class="keyword">return</span> <span class="number">1</span>; &#125;</span><br><span class="line"><span class="keyword">for</span> (<span class="keyword">auto</span>&amp; o : a_out) md.<span class="built_in">on_event</span>(o);              <span class="comment">// feed the L2 mirror</span></span><br><span class="line"><span class="keyword">if</span> (i % check_every == <span class="number">0</span>) &#123;</span><br><span class="line">    eng.<span class="built_in">check_invariants</span>();                         <span class="comment">// aborts on failure</span></span><br><span class="line">    <span class="keyword">if</span> (eng.orders_in_priority(Side::Buy)  != ref.orders_in_priority(Side::Buy)  ||</span><br><span class="line">        eng.orders_in_priority(Side::Sell) != ref.orders_in_priority(Side::Sell) ||</span><br><span class="line">        md.<span class="built_in">depth</span>(Side::Buy)  != eng.<span class="built_in">depth</span>(Side::Buy) ||</span><br><span class="line">        md.<span class="built_in">depth</span>(Side::Sell) != eng.<span class="built_in">depth</span>(Side::Sell)) &#123;</span><br><span class="line">        std::<span class="built_in">printf</span>(<span class="string">&quot;STATE DIVERGENCE seed %llu event %llu\n&quot;</span>, seed, i);</span><br><span class="line">        <span class="keyword">return</span> <span class="number">1</span>;</span><br><span class="line">    &#125;</span><br><span class="line">&#125;</span><br></pre></td></tr></table></figure><p><code>orders_in_priority</code> lists every resting order on a side in the exact order it would fill. <code>check_invariants</code> aborts if levels are unsorted or empty, if a level's total or count disagrees with its queue, if back links do not mirror forward links, if the id map holds anything but the resting orders, if a pool slot is neither resting nor free, or if the book is crossed.</p><h3 id="tests">tests</h3><p>The suite is 37 GoogleTest cases plus two command line tests. Scenario tests cover each rule. Differential tests run five flow shapes (default, narrow and heavily crossing, a wide book, an aggressive mix, and 50 short seeds) and compare state every 257 events. Replay tests check that the same inputs give the same digest and that dropping one input changes it, and the golden session replays against its reviewed output.</p><h2 id="problems">problems</h2><h3 id="1-a-priority-bug-with-no-wrong-output">1. a priority bug with no wrong output</h3><p>To check that the fuzzer could find bugs at all, <code>scripts/mutation_test.sh</code> copies the engine, applies one <code>sed</code> edit that plants a known bug, rebuilds, and runs <code>exch fuzz</code> for two seeds of 200,000 events with a state check every 100 events. The most realistic mutant is a one-token change to the modify condition.</p><figure class="highlight diff"><table><tr><td class="code"><pre><span class="line"><span class="deletion">- if (in.price == o.price &amp;&amp; in.qty &lt;= o.qty) &#123;</span></span><br><span class="line"><span class="addition">+ if (in.price == o.price) &#123;</span></span><br></pre></td></tr></table></figure><p>With this change, an order that grows at the same price keeps its queue position. I expected the fuzzer to catch it immediately through the output diff. It did not. At the moment of the modify, the correct engine and the mutant emit identical events. Both emit <code>MODIFIED</code> and then one <code>L2</code> update with the level's new total, because the correct path removes the order and re-rests it at the same price with no trade in between. The only difference is where the order sits in its queue, and queue position is invisible until a later aggressive order reaches that level and fills the wrong maker first.</p><p>The mutant was caught by the periodic comparison of <code>orders_in_priority</code>, which reported a state divergence at event 3000 of seed 1. The other five mutants were caught by the output diff, most of them quickly. Printing at the taker's price failed on event 9, LIFO queues on event 86, a post-only that may lock the touch on event 112, and an off-by-one in the FOK check (<code>&lt;=</code> instead of <code>&lt;</code>) on event 593.</p><figure data-figure="chart:projects/exchange-from-scratch/exchange-from-scratch-mutation"></figure><p>The slowest catch was the missing <code>touch</code> in the in-place amend branch, at event 3,487. My reading of why it took so long is that the bug needs a modify whose new price lands exactly on the order's current price with a size that does not increase, and the generator draws modify prices uniformly from a band about 41 ticks wide, so such modifies are rare. I have not measured that rate directly. Once one happens, the L2 stream is visibly short an update and the diff fails on that event.</p><p>The lesson I took is that an output diff tests the interface and a state diff tests the model, and a bug can live in the model for a long time before it reaches the interface. Both the tests and the fuzzer compare state for this reason.</p><h3 id="2-the-standard-library-map-was-faster-at-inserting">2. the standard library map was faster at inserting</h3><p>I wrote <code>IdMap</code> expecting it to beat <code>std::unordered_map</code> everywhere. It did not. On a burst of 1M adds into an empty book, the standard map ran at 9.8M orders per CPU second against 4.4M for mine. The reason is the hash. libc++ hashes integers with the identity function, so sequential order ids go into sequential buckets and inserts walk memory in order. My splitmix64 hash scatters them over a table of 2^21 slots of 16 bytes each, which is 32 MB, so each insert is likely a cache miss.</p><h3 id="3-an-identity-hash-made-matching-14-times-slower">3. an identity hash made matching 14 times slower</h3><p>The obvious fix was to use the identity hash in <code>IdMap</code> too, so I added it as a third benchmark variant. Adds became as fast as the standard map, 9.8M per CPU second. Matching fell to 0.27M per CPU second, 14 times slower than the 3.8M of splitmix64, with per-run values between 0.23M and 0.32M.</p><p>The cause is the backward shift loop shown above. Sequential ids with an identity hash fill one long contiguous run of occupied slots. Deletion walks forward from the hole to the end of the run. Matching removes makers in FIFO order, which is roughly ascending id order, so every deletion near the start of the run walks almost the whole remaining run. The total cost is quadratic. Cancels in random order break the run into pieces quickly, which is why the cancel burst did not show the problem and why a benchmark with only random cancels would have hidden it.</p><p>The identity hash also collapses on ids that share their low bits. Inserting 20,000 ids that are multiples of 2^16 ran at 0.06M per CPU second against 6.5M for splitmix64, because they all probe from the same home slot. A client that picks its own ids could use that to slow the engine down on purpose. I kept splitmix64, and the real fix is in what I would change.</p><h3 id="4-the-timer-is-coarser-than-the-thing-being-timed">4. the timer is coarser than the thing being timed</h3><p><code>steady_clock</code> on Apple Silicon advances in steps of 41 to 42 ns. It is a 24 MHz counter. The smallest nonzero step I measured was 41 ns and back-to-back reads had a p50 of 41 ns. An add or a match takes a few of those ticks, so every latency percentile is a multiple of about 41.7 ns (42, 83, 125, 167, 208 and so on), and a match p50 of 84 ns means &quot;two ticks&quot;, not 84 ns to three significant figures. The CDF of any path is a staircase.</p><p>User space cannot read the cycle counter on Apple Silicon without help from the kernel. I kept per-order timing for the shape of the distribution and added burst runs with no timers inside the loop, which give an averaged cost per order below the timer's resolution.</p><h3 id="5-the-machine-was-not-mine">5. the machine was not mine</h3><p>Other heavy jobs shared the machine the whole session, and during the benchmark the 1 minute load average stayed between 172 and 352 on 14 cores. In my first attempt the benchmark process got about 5 percent of a core and the same mixed-flow run varied by roughly a factor of ten in wall-clock throughput.</p><p>I changed three things. Throughput is measured in thread CPU time (<code>CLOCK_THREAD_CPUTIME_ID</code>) as well as wall time, since CPU time excludes time spent descheduled. Every configuration runs three times with the variants interleaved, so a burst of load hits all of them, and the reported number is the median. The benchmark also requests the user-interactive QoS class, since macOS does not allow pinning. None of this makes the numbers clean. CPU-time throughput for the default engine on the mixed flow still ranged from 9.2M to 17.4M events per second across runs, which I attribute to the scheduler moving the thread between performance and efficiency cores and to cache pollution from neighbours. Maximum latencies in the raw files reach tens to hundreds of milliseconds. Those are preemptions, and I do not report them.</p><h3 id="6-a-throughput-column-that-measured-nothing">6. a throughput column that measured nothing</h3><p>My first benchmark printed an operations per second figure for each latency scenario, dividing timed operations by the wall time of a loop that also contained the untimed steps keeping the book at a steady size. The number was neither the add rate nor the cancel rate. I deleted it and wrote separate burst runs, 1M adds into an empty book, 1M cancels in random order, and 1M aggressive orders that each fill exactly one maker.</p><h3 id="7-apple-clang-s-addresssanitizer-hangs">7. Apple clang's AddressSanitizer hangs</h3><p>The first sanitizer build compiled and then every binary hung, including the test discovery step. I reduced it to <code>int main() { return 0; }</code> built with <code>-fsanitize=address</code>, which still hung. UBSan alone was fine. <code>ASAN_OPTIONS=verbosity=2</code> showed the runtime stuck in <code>FindDynamicShadowStart</code>, where ASan searches for free address space for its shadow memory. So the problem was Apple clang 17's ASan runtime on macOS 26.5, not my code. Homebrew LLVM builds working ASan binaries, so the <code>asan</code> preset uses that compiler, plus <code>-gdwarf-4</code> to silence the linker warning its DWARF 5 debug info produced on every object.</p><h2 id="experiments">experiments</h2><h3 id="differential-fuzzing">differential fuzzing</h3><p><code>exch fuzz</code> in Release, with a state check every 1,000 events. Eight seeds of 1M events, run as four processes of two seeds, with price bands from 11 to 101 ticks wide depending on the seed, plus one run of 300,000 events with a 401-tick band and up to 5,000 live ids for a deeper book. Each 2M-event process took 157.5 to 163.9 s of wall time on the loaded machine and the deep book run 325.4 s, almost all of it in the reference matcher.</p><h3 id="mutation-check">mutation check</h3><p>The six mutants from problem 1, each run for two seeds of 200,000 events with a state check every 100 events.</p><h3 id="latency-by-path-and-book-size">latency by path and book size</h3><p>For each book size (1k, 10k, 100k and 1M resting orders over 100 price levels per side) and each path, 300,000 timed operations after 30,000 of warmup. Each timed add is followed by an untimed cancel, each timed cancel by an untimed add, and each timed single-fill match by an untimed replacement maker, so the book stays the same size. Three repetitions, median of each percentile.</p><h3 id="throughput">throughput</h3><p>Burst runs as described in problem 6, and a mixed flow of 3M pre-generated events from <code>OrderFlow</code>, about half of them new orders across all five types (1,495,188 of 3M in the first run) and the rest cancels, modifies and invalid inputs, with three seeds per run and three runs per variant.</p><h3 id="id-index-ablation">id index ablation</h3><p>The same engine and benchmark compiled with the default <code>IdMap</code>, with <code>std::unordered_map</code>, and with <code>IdMap</code> using an identity hash, interleaved within each repetition.</p><h3 id="replay-determinism">replay determinism</h3><p><code>exch gen --n 1000000 --seed 2026</code>, two independent <code>exch replay</code> runs, and an <code>exch verify</code> of a replay against the first run's output log and digest.</p><h2 id="results">results</h2><p>All benchmark numbers come from an Apple M4 Pro (14 cores, 48 GB) running macOS 26.5, Apple clang 17 at <code>-O2</code>, on the loaded machine described in problem 5. They were not rerun independently. The correctness results were.</p><h3 id="correctness">correctness</h3><table><thead><tr><th>Check</th><th>Result</th><th>Source</th></tr></thead><tbody><tr><td>Release test suite</td><td>39 of 39 passed, and 39 of 39 again in an independent clean rebuild</td><td><code>results/tests_release.txt</code></td></tr><tr><td>ASan and UBSan suite</td><td>39 of 39 passed, 132 s at <code>ctest -j2</code>, in the builder's rerun</td><td><code>results/tests_asan_ubsan.txt</code></td></tr><tr><td>differential fuzz, 8 seeds of 1M</td><td>8M events, 16.4M output events, 2.83M trades, 0 divergences</td><td><code>results/fuzz_seed*.txt</code></td></tr><tr><td>differential fuzz, deep book</td><td>300,000 events, 101,890 trades, 0 divergences</td><td><code>results/fuzz_deep_book.txt</code></td></tr><tr><td>independent fuzz rerun, seed 11</td><td>200,000 events, 0 divergences</td><td>verification rerun</td></tr><tr><td>planted bugs caught</td><td>6 of 6</td><td><code>results/mutation.txt</code></td></tr><tr><td>replay of 1M inputs</td><td>same digest <code>fa2ff178540c7510</code> on both runs, verify matched all 2,072,205 output lines</td><td><code>results/replay/</code></td></tr></tbody></table><p>The 1M-event replay produced 2,072,205 outputs, 358,590 trades and 675,925 L2 updates, and left 95 orders resting on 17 bid and 14 ask levels, identically on both runs. A sanitizer run earlier in the session, at a load average near 300, needed over 8 minutes for each of the two longest differential tests. The 132 s figure is from a rerun at a load average of about 11.</p><p>The fuzzer shows the engine agrees with the reference on the flows the generator produces, and the mutation check shows the comparison is sensitive to six specific kinds of error. It does not show the reference itself is right. The reference could share a misreading of the rules with the engine, since I wrote both. The golden session in <code>examples/</code>, whose output I reviewed by hand, is the check against that, and it is small.</p><h3 id="latency">latency</h3><table><thead><tr><th>Path</th><th>Resting orders</th><th>p50</th><th>p99</th><th>p99.9</th></tr></thead><tbody><tr><td>add</td><td>10,000</td><td>167 ns</td><td>541 ns</td><td>3,208 ns</td></tr><tr><td>cancel</td><td>10,000</td><td>209 ns</td><td>916 ns</td><td>7,333 ns</td></tr><tr><td>match, one fill</td><td>10,000</td><td>84 ns</td><td>542 ns</td><td>1,583 ns</td></tr><tr><td>add</td><td>1,000,000</td><td>84 ns</td><td>625 ns</td><td>1,500 ns</td></tr><tr><td>cancel</td><td>1,000,000</td><td>459 ns</td><td>1,417 ns</td><td>23,834 ns</td></tr><tr><td>match, one fill</td><td>1,000,000</td><td>209 ns</td><td>958 ns</td><td>6,083 ns</td></tr></tbody></table><figure data-figure="chart:projects/exchange-from-scratch/exchange-from-scratch-latency"></figure><p>Matching at the touch is the cheapest path, because the best level is <code>back()</code> and the maker is the head of its queue. Cancel is the most sensitive to book size. Its p50 goes from 167 ns at 1k orders to 459 ns at 1M, because it does a hash lookup and then touches a random order slot, and at 1M orders both the 40 MB pool and the id table are far larger than cache.</p><p>Two cells should not be over-read. The add p50 at 1M (84 ns) is lower than at 100k (208 ns), but that cell ranged from 83 to 166 ns across the three runs, so I read it as noise plus tick quantization rather than an effect. The p99.9 column mostly describes the machine. The same cell differs by up to ten times between runs, with cancel at 10k ranging from 5,209 to 73,417 ns.</p><h3 id="throughput-2">throughput</h3><p>On the mixed random flow the default engine's median was 13.1M events per CPU second, which is 6.5M new orders per CPU second, with a best run of 17.4M. The wall-clock median was 4.8M events per second, which says more about the load than about the engine. On the bursts of 1M orders the medians were 4.4M adds, 2.2M cancels and 3.8M single-fill matches per CPU second.</p><h3 id="the-id-index-ablation">the id index ablation</h3><figure data-figure="chart:projects/exchange-from-scratch/exchange-from-scratch-id-index-burst"></figure><table><thead><tr><th>Variant</th><th>add burst</th><th>cancel burst</th><th>match burst</th><th>strided ids</th><th>cancel p50 at 1M</th></tr></thead><tbody><tr><td>IdMap, splitmix64 (default)</td><td>4.4M/s</td><td>2.2M/s</td><td>3.8M/s</td><td>6.5M/s</td><td>459 ns</td></tr><tr><td><code>std::unordered_map</code></td><td>9.8M/s</td><td>1.4M/s</td><td>2.2M/s</td><td>10.4M/s</td><td>750 ns</td></tr><tr><td>IdMap, identity hash</td><td>9.8M/s</td><td>1.5M/s</td><td>0.27M/s</td><td>0.06M/s</td><td>459 ns</td></tr></tbody></table><p>Burst figures are orders per CPU second. The custom table wins where the id index is on the critical path of a large book. At 1M resting orders, all three <code>idmap</code> runs had a cancel p50 of 416 to 500 ns and all three <code>stdmap</code> runs 750 to 792 ns, so the ranges do not overlap even on this noisy machine, and the match burst was 1.7 times faster. It loses on inserting sequential ids, where the standard library's identity hash gets perfect locality.</p><figure data-figure="chart:projects/exchange-from-scratch/exchange-from-scratch-cancel-book-size"></figure><p>On the mixed flow, where the book stays small (about 1,500 orders resting at the end of each run), the three variants were within run-to-run noise of each other, with medians of 13.1M, 16.0M and 15.9M events per CPU second for idmap, stdmap and identity and individual runs from 9.2M to 22.1M. I would not claim the default is slower on that flow, and I would not claim it is faster either.</p><h2 id="what-i-would-change">what I would change</h2><h3 id="exchange-assigned-order-ids">exchange-assigned order ids</h3><p>Most of the id index trouble comes from accepting arbitrary client ids. If a gateway mapped each client id to an exchange id that increases by one per order, the index could be a flat array indexed by <code>id - base</code>, with no hashing at all. Inserts would be sequential, lookups O(1), and there would be no probe runs to degrade and no ids for a client to choose adversarially. That is the design I would build with the TCP gateway.</p><h3 id="a-faster-reference-matcher">a faster reference matcher</h3><p>The oracle's per-event snapshots limit the fuzzer to about 12,000 events per second on this machine. Keeping the reference's L2 aggregate in a map updated alongside its flat vector would stay independent of the engine's touched-level logic and should be much faster, and more events per second means more seeds for rare bugs like the missing amend update.</p><h3 id="measure-how-often-the-generator-hits-each-rule">measure how often the generator hits each rule</h3><p>The amend bug took 3,487 events because same-price modifies are rare in the generated flow. I would count how often each branch of the engine fires during a fuzz run, and bias the generator toward the rare ones, rather than guess at their frequency as I did above.</p><h3 id="quieter-benchmarking">quieter benchmarking</h3><p>I would rerun the suite on an idle machine, report the per-run spread next to every median, and try the kperf counters for cycle-level timing instead of 41.7 ns ticks.</p><h3 id="a-price-ladder-for-liquid-instruments">a price ladder for liquid instruments</h3><p>The sorted vector of levels is simple and fast at the touch, but inserting a new level deep in the book moves every level behind it. For an instrument with a known tick band, a dense array indexed by price with a bitmap for finding the next non-empty level would make every level operation O(1).</p><h3 id="a-binary-log">a binary log</h3><p>Replaying the 1M-event text log took 3.4 s of wall time without writing outputs and 6.9 s with the output log, while the engine itself processes 3M events of the mixed flow in 0.17 to 0.33 CPU seconds. A fixed-size binary record per event would make replay limited by the engine rather than by <code>snprintf</code>.</p><h3 id="post-only-semantics-on-modify">post-only semantics on modify</h3><p>A post-only order modified across the spread currently trades as a plain limit order. A real venue would reject the modify or reprice it, so the event model should carry the original order type and let the engine apply the right rule.</p><h2 id="reproducibility">reproducibility</h2><p>Build. The <code>asan</code> preset needs Homebrew LLVM (<code>brew install llvm</code>) because of problem 7. GoogleTest is fetched by CMake at configure time.</p><figure class="highlight sh"><table><tr><td class="code"><pre><span class="line"><span class="built_in">cd</span> projects/06-price-time-matching-engine</span><br><span class="line">cmake --preset release &amp;&amp; cmake --build --preset release</span><br><span class="line">cmake --preset asan &amp;&amp; cmake --build --preset asan</span><br></pre></td></tr></table></figure><p>Tests.</p><figure class="highlight sh"><table><tr><td class="code"><pre><span class="line">ctest --preset release          <span class="comment"># about 1 minute</span></span><br><span class="line">ctest --preset asan -j2         <span class="comment"># about 2 minutes on an idle machine</span></span><br></pre></td></tr></table></figure><p>Differential fuzzing and the mutation check.</p><figure class="highlight sh"><table><tr><td class="code"><pre><span class="line">./build/release/exch fuzz --n 1000000 --seed 1 --seeds 2</span><br><span class="line">./build/release/exch fuzz --n 300000 --seed 21 --width 200 --max-live 5000</span><br><span class="line">./build/release/exch fuzz --n 200000 --seed 11</span><br><span class="line">scripts/mutation_test.sh 200000</span><br></pre></td></tr></table></figure><p>Replay determinism.</p><figure class="highlight sh"><table><tr><td class="code"><pre><span class="line">./build/release/exch gen --n 1000000 --seed 2026 --out /tmp/in.log</span><br><span class="line">./build/release/exch replay --<span class="keyword">in</span> /tmp/in.log --out /tmp/out.log</span><br><span class="line">./build/release/exch verify --<span class="keyword">in</span> /tmp/in.log --expect /tmp/out.log</span><br></pre></td></tr></table></figure><p>Benchmarks and plots. Raw CSVs, per-run output, the load average log and the plots land in <code>results/bench/</code>.</p><figure class="highlight sh"><table><tr><td class="code"><pre><span class="line">scripts/run_bench.sh 3 300000 3000000</span><br><span class="line">python3 -m venv .venv &amp;&amp; .venv/bin/pip install matplotlib</span><br><span class="line">.venv/bin/python scripts/plot_bench.py results/bench</span><br></pre></td></tr></table></figure><p>Code is in <code>projects/06-price-time-matching-engine</code>, with the engine, reference matcher and market-data feed in <code>src/</code> and <code>include/exchange/</code>, the <code>exch</code> tool in <code>tools/</code>, the benchmark in <code>bench/</code>, the tests in <code>tests/</code>, the data structures, invariants and rejected alternatives in <code>DESIGN.md</code>, and the build log with its verification note in <code>DEVLOG.md</code>.</p>]]>
    </content>
    <id>https://projects.farhansadeek.com/posts/fuzzing-a-matching-engine/</id>
    <link href="https://projects.farhansadeek.com/posts/fuzzing-a-matching-engine/"/>
    <published>2026-09-29T06:17:09.000Z</published>
    <summary>A price-time priority matching engine in C++, checked event by event against a deliberately simple reference on 8.3 million random events.</summary>
    <title>Fuzzing a Matching Engine Against a Brute-Force Twin</title>
    <updated>2026-09-29T06:49:26.552Z</updated>
  </entry>
  <entry>
    <author>
      <name>Farhan Sadeek</name>
      <email>farhan@farhansadeek.com</email>
    </author>
    <category term="databases" scheme="https://projects.farhansadeek.com/tags/databases/"/>
    <category term="storage" scheme="https://projects.farhansadeek.com/tags/storage/"/>
    <category term="cpp" scheme="https://projects.farhansadeek.com/tags/cpp/"/>
    <content>
      <![CDATA[<p>I wrote kvd, a single-threaded TCP server that speaks a subset of RESP2, the Redis wire protocol. It stores strings with optional TTLs, executes pipelined requests, expires keys lazily and actively the way Redis does, and logs every write to an append-only file (AOF) that it replays on startup, including after a <code>kill -9</code> or a torn final record. A pipelined C++ load generator, <code>kvbench</code>, ships with it.</p><p>The server is 1,922 lines of C++20 in <code>src/</code>, headers and comments included. There are <strong>49 GoogleTest cases</strong> plus an end-to-end test written in stdlib-only Python against the real binary, and on 2026-09-27 an independent clean release rebuild passed all <strong>50</strong>. The AddressSanitizer plus UBSan build also passed in my own runs, but it was not rerun independently, and neither were the benchmarks.</p><p>The headline measurement is how little the server's own work matters. At 50 connections, SET throughput went from <strong>107,160 ops/s</strong> with one request in flight per connection to <strong>2,942,811 ops/s</strong> with 128 in flight, a factor of 27 with no change to the server. Server CPU per command fell from about 9 µs to about 0.34 µs over that range, which says nearly all the cost at depth 1 is system calls and wakeups, not the hash table. All benchmark numbers were measured on a shared Apple M4 Pro that was also running other jobs, with the 1 minute load average recorded in every row (7 to 13 for the reported runs). They are indicative, not a clean benchmark.</p><p>Two gaps should be stated up front. The epoll backend exists as source but has <strong>never been compiled</strong>, because there was no Linux machine in the session, so the Linux half of the poller abstraction is a claim rather than a result. And there was <strong>no comparison against real Redis</strong>, because neither <code>redis-server</code> nor <code>redis-benchmark</code> was installed. Every number below is internally consistent, and none of them says whether kvd is fast relative to the thing it imitates.</p><p>Code is in <code>projects/02-kvd-resp-server</code>. The argument of this post is that a single-threaded server's speed is decided by how many system calls each command shares, and that the hardest thing to keep honest was the measurement rather than the code. Skip to <a href="#problems">problems</a> for the bugs.</p><h2 id="table-of-contents">table of contents</h2><ul><li><a href="#what-i-wanted-to-build">what I wanted to build</a></li><li><a href="#theory">theory</a></li><li><a href="#architecture">architecture</a></li><li><a href="#implementation">implementation</a></li><li><a href="#problems">problems</a></li><li><a href="#experiments">experiments</a></li><li><a href="#results">results</a></li><li><a href="#what-i-would-change">what I would change</a></li><li><a href="#reproducibility">reproducibility</a></li></ul><h2 id="what-i-wanted-to-build">what I wanted to build</h2><p>Redis runs every command on one thread and is still fast enough that most people never need anything else. I wanted to understand why by building the core of it and measuring where it breaks.</p><p>The v0 scope was fixed before writing code. One event loop on kqueue, with an epoll path behind the same interface. RESP2 with pipelining and the inline form that makes <code>nc</code> usable. <code>GET</code>, <code>SET</code> with every expiry and condition flag, <code>DEL</code>, <code>EXISTS</code>, the <code>EXPIRE</code> family, <code>TTL</code>, <code>INCR</code> and friends, and <code>KEYS</code> with glob patterns. Lazy and active expiry modelled on Redis. An AOF with three fsync policies, replay and torn-tail recovery. And a load generator, so the design choices could be measured rather than asserted.</p><p>Multithreaded I/O, RDB snapshots, AOF rewrite, lists and hashes and sorted sets, pub/sub and replication were left for later milestones. <code>COMMAND</code> and <code>CONFIG GET</code> are stubs that return empty arrays, so a <code>redis-cli</code> handshake should succeed, though I could not check that here.</p><h2 id="theory">theory</h2><h3 id="why-one-thread-can-be-enough">why one thread can be enough</h3><p>A request/response server spends most of its life waiting on the network. With a thread per connection the waiting is done by blocked threads, and the cost is context switches and stack memory. An event loop inverts that. One thread asks the kernel which sockets are ready (kqueue on macOS, epoll on Linux) and touches only those. As long as every command is short, one core can serve thousands of connections, and because only one thread ever touches the data, every command is atomic without a lock.</p><p>The catch is the word short. One slow command, a <code>KEYS</code> over a million keys or an fsync on a slow disk, stalls every client at once.</p><h3 id="where-the-time-goes">where the time goes</h3><p>A <code>GET</code> that finds its key in a hash table costs well under a microsecond. The <code>read()</code>, <code>send()</code> and <code>kevent()</code> calls around it cost several. So the throughput of a server like this is set mostly by how many system calls each command pays for.</p><p>Pipelining attacks exactly that. If a client sends 16 commands before reading any replies, the server picks them up with one <code>read()</code>, executes them back to back, and answers with one <code>send()</code>. The syscall cost is shared 16 ways.</p><h3 id="little-s-law-as-a-sanity-check">Little's law as a sanity check</h3><p>In a closed-loop benchmark, where each connection keeps a fixed number of requests outstanding, Little's law ties the three quantities together.</p><p class="katex-block"><span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mtext>requests in flight</mtext><mo>=</mo><mtext>throughput</mtext><mo>×</mo><mtext>mean latency</mtext></mrow><annotation encoding="application/x-tex">\text{requests in flight} = \text{throughput} \times \text{mean latency}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.8889em;vertical-align:-0.1944em;"></span><span class="mord text"><span class="mord">requests in flight</span></span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.8889em;vertical-align:-0.1944em;"></span><span class="mord text"><span class="mord">throughput</span></span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">×</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:0.8889em;vertical-align:-0.1944em;"></span><span class="mord text"><span class="mord">mean latency</span></span></span></span></span></span></p><p>With 50 connections and one request in flight each, a server doing 100,000 requests per second must show a mean latency of <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mn>50</mn><mi mathvariant="normal">/</mi><mn>100,000</mn></mrow><annotation encoding="application/x-tex">50 / 100{,}000</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mord">50/100</span><span class="mord"><span class="mpunct">,</span></span><span class="mord">000</span></span></span></span> s <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mo>=</mo><mn>500</mn></mrow><annotation encoding="application/x-tex">= 500</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.3669em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.6444em;"></span><span class="mord">500</span></span></span></span> µs, however fast each individual command is. Latency in a saturated closed-loop benchmark is set by the queue, not by the work. I used this to check every result below.</p><h3 id="expiry-without-a-timer-per-key">expiry without a timer per key</h3><p>A key with a TTL must disappear on time, but nobody wants a timer per key. Redis combines two cheap mechanisms. Lazy expiry checks the deadline whenever a key is touched, so an expired key is never visible to a client. Active expiry runs about ten times a second, samples 20 random keys that have a TTL, deletes the expired ones, and repeats while more than a quarter of the sample was expired. That bounds how much memory dead but untouched keys can hold, at a cost proportional to the garbage rather than to the keyspace.</p><h3 id="durability-is-a-policy">durability is a policy</h3><p>The AOF is a log of write commands in the same RESP format clients send, and replaying it rebuilds the dataset. The durability knob is when to call <code>fsync</code>. <code>always</code> means before acknowledging a write, <code>everysec</code> means once a second, and <code>no</code> leaves it to the operating system.</p><h2 id="architecture">architecture</h2><p>Everything runs on one thread.</p><figure data-figure="diagram:redis-event-loop"></figure><p>One loop iteration does four things in order.</p><ol><li><code>before_sleep()</code> writes the AOF buffer to the file, fsyncs it under <code>always</code>, and only then tries to <code>send()</code> every client's pending replies.</li><li><code>poller.wait()</code> blocks until a socket is ready or the next cron tick is due. If a client still has unsent output that the kernel did not refuse, the timeout is zero.</li><li>Events are dispatched. The listening socket accepts until <code>EAGAIN</code>. A readable client gets one <code>read()</code> of up to 16 KB, and every complete command in its buffer is parsed and executed.</li><li>If the cron is due, it runs one active expiry cycle and the <code>everysec</code> fsync check.</li></ol><p>Each component is one file under <code>src/</code>. The poller is a four-method interface with kqueue and epoll implementations behind it, so the server never includes a platform readiness header.</p><h2 id="implementation">implementation</h2><h3 id="the-loop-that-makes-pipelining-work">the loop that makes pipelining work</h3><p>Pipelining is not a feature the server implements so much as a property of this loop. Every complete frame in the input buffer is executed in one pass, and all the replies land in one output buffer that goes out in one <code>send()</code>.</p><figure class="highlight cpp"><table><tr><td class="code"><pre><span class="line"><span class="function"><span class="type">void</span> <span class="title">Server::process_input</span><span class="params">(Client&amp; c)</span> </span>&#123;</span><br><span class="line">    <span class="type">int64_t</span> now = <span class="built_in">wall_ms</span>();  <span class="comment">// one clock read per batch, like Redis&#x27;s cached mstime</span></span><br><span class="line">    <span class="keyword">while</span> (!c.close_after_write &amp;&amp; c.in_pos &lt; c.in.<span class="built_in">size</span>()) &#123;</span><br><span class="line">        <span class="keyword">if</span> (c.out.<span class="built_in">size</span>() - c.out_pos &gt; cfg_.out_soft_limit) &#123;</span><br><span class="line">            c.read_paused = <span class="literal">true</span>;               <span class="comment">// backpressure, see below</span></span><br><span class="line">            poller_-&gt;<span class="built_in">set_read</span>(c.fd, <span class="literal">false</span>);</span><br><span class="line">            ++stats_.read_pauses;</span><br><span class="line">            <span class="keyword">break</span>;</span><br><span class="line">        &#125;</span><br><span class="line">        ParseResult r = <span class="built_in">parse_request</span>(std::<span class="built_in">string_view</span>(c.in).<span class="built_in">substr</span>(c.in_pos), args_);</span><br><span class="line">        <span class="keyword">if</span> (r.status == ParseStatus::Incomplete) <span class="keyword">break</span>;</span><br><span class="line">        <span class="comment">// ... protocol errors reply and close ...</span></span><br><span class="line">        c.in_pos += r.consumed;</span><br><span class="line">        <span class="keyword">if</span> (args_.<span class="built_in">empty</span>()) <span class="keyword">continue</span>;</span><br><span class="line">        <span class="keyword">if</span> (proc_.<span class="built_in">execute</span>(args_, now, c.out) == CommandProcessor::Action::Close)</span><br><span class="line">            c.close_after_write = <span class="literal">true</span>;</span><br><span class="line">    &#125;</span><br><span class="line">    <span class="comment">// compact `in`, then queue the client for before_sleep() if it has output</span></span><br><span class="line">&#125;</span><br></pre></td></tr></table></figure><p>Each client carries an input buffer with a consumed offset and an output buffer with a sent offset. Advancing an offset is O(1). Erasing the front of a <code>std::string</code> after every command would make a deep pipeline quadratic in its length, so the input buffer is cleared when fully consumed (the common case) and compacted only when the dead prefix is over 64 KB and more than half the buffer.</p><h3 id="a-parser-that-copies-once">a parser that copies once</h3><p>The RESP parser is stateless. It is handed the unconsumed tail of the input buffer and returns complete, incomplete or error. For a multibulk frame it makes a first pass that records only offsets, and copies arguments into <code>std::string</code>s only once the whole frame is present.</p><figure class="highlight cpp"><table><tr><td class="code"><pre><span class="line"><span class="keyword">struct</span> <span class="title class_">Span</span> &#123; <span class="type">size_t</span> off, len; &#125;;</span><br><span class="line">std::vector&lt;Span&gt; spans;</span><br><span class="line"><span class="keyword">for</span> (<span class="type">int64_t</span> k = <span class="number">0</span>; k &lt; n; ++k) &#123;</span><br><span class="line">    <span class="comment">// ... find &quot;$&lt;len&gt;\r\n&quot;, validate len ...</span></span><br><span class="line">    <span class="type">size_t</span> need = data + <span class="keyword">static_cast</span>&lt;<span class="type">size_t</span>&gt;(len) + <span class="number">2</span>;</span><br><span class="line">    <span class="keyword">if</span> (buf.<span class="built_in">size</span>() &lt; need) <span class="keyword">return</span> &#123;&#125;;          <span class="comment">// incomplete, nothing copied</span></span><br><span class="line">    spans.<span class="built_in">push_back</span>(&#123;data, <span class="keyword">static_cast</span>&lt;<span class="type">size_t</span>&gt;(len)&#125;);</span><br><span class="line">    pos = need;</span><br><span class="line">&#125;</span><br><span class="line">args.<span class="built_in">resize</span>(spans.<span class="built_in">size</span>());</span><br><span class="line"><span class="keyword">for</span> (<span class="type">size_t</span> k = <span class="number">0</span>; k &lt; spans.<span class="built_in">size</span>(); ++k) args[k].<span class="built_in">assign</span>(buf.<span class="built_in">data</span>() + spans[k].off, spans[k].len);</span><br></pre></td></tr></table></figure><p>An 8 MB value that trickles in across hundreds of reads is scanned for headers many times but copied once. There is a test that sends exactly that.</p><h3 id="a-dense-vector-of-ttl-keys">a dense vector of TTL keys</h3><p>Active expiry needs a uniformly random key among those with a TTL, in O(1). Redis gets that from a second hash table of expiring keys and a random bucket walk. I kept a <code>std::vector</code> of raw pointers to the <code>unordered_map</code> nodes that have a TTL, with each entry storing its own index into that vector.</p><figure class="highlight cpp"><table><tr><td class="code"><pre><span class="line"><span class="function"><span class="type">void</span> <span class="title">Store::remove_ttl</span><span class="params">(Node&amp; node)</span> </span>&#123;</span><br><span class="line">    Entry&amp; e = node.second;</span><br><span class="line">    <span class="keyword">if</span> (e.expire_at_ms == kNoExpiry) <span class="keyword">return</span>;</span><br><span class="line">    <span class="type">size_t</span> slot = e.ttl_slot;</span><br><span class="line">    Node* last = ttl_nodes_.<span class="built_in">back</span>();</span><br><span class="line">    ttl_nodes_[slot] = last;          <span class="comment">// swap the last pointer into our slot</span></span><br><span class="line">    last-&gt;second.ttl_slot = slot;</span><br><span class="line">    ttl_nodes_.<span class="built_in">pop_back</span>();</span><br><span class="line">    e.expire_at_ms = kNoExpiry;</span><br><span class="line">&#125;</span><br><span class="line"></span><br><span class="line"><span class="function"><span class="type">void</span> <span class="title">Store::erase</span><span class="params">(Map::iterator it)</span> </span>&#123;</span><br><span class="line">    <span class="built_in">remove_ttl</span>(*it);  <span class="comment">// must happen before erase, ttl_nodes_ points into the node</span></span><br><span class="line">    map_.<span class="built_in">erase</span>(it);</span><br><span class="line">&#125;</span><br></pre></td></tr></table></figure><p>This is legal only because <code>std::unordered_map</code> never moves its nodes. Rehashing invalidates iterators but not pointers or references to elements. The price of that guarantee is that the map is node-based, one allocation per key, which comes back in the replay results. <code>Store::check_invariants</code> walks the map and checks that every TTL node is in the vector at the slot it records and that the counts match, and a randomized test runs it after each of 20,000 mixed operations.</p><p>The map also uses a transparent hash, so a lookup by <code>string_view</code> does not allocate a temporary <code>std::string</code>.</p><h3 id="active-expiry-with-a-wall-clock-budget">active expiry with a wall-clock budget</h3><figure class="highlight cpp"><table><tr><td class="code"><pre><span class="line"><span class="keyword">while</span> (!ttl_nodes_.<span class="built_in">empty</span>()) &#123;</span><br><span class="line">    <span class="type">size_t</span> n = std::<span class="built_in">min</span>(kSampleSize, ttl_nodes_.<span class="built_in">size</span>());   <span class="comment">// 20</span></span><br><span class="line">    <span class="type">size_t</span> expired_this_round = <span class="number">0</span>;</span><br><span class="line">    <span class="keyword">for</span> (<span class="type">size_t</span> k = <span class="number">0</span>; k &lt; n &amp;&amp; !ttl_nodes_.<span class="built_in">empty</span>(); ++k) &#123;</span><br><span class="line">        Node* node = ttl_nodes_[<span class="built_in">pick</span>(rng_)];</span><br><span class="line">        <span class="keyword">if</span> (node-&gt;second.expire_at_ms &lt;= now_ms) &#123;</span><br><span class="line">            <span class="built_in">erase</span>(map_.<span class="built_in">find</span>(node-&gt;first));</span><br><span class="line">            ++expired_this_round;</span><br><span class="line">        &#125;</span><br><span class="line">    &#125;</span><br><span class="line">    <span class="keyword">if</span> (expired_this_round * <span class="number">4</span> &lt;= n) <span class="keyword">break</span>;                 <span class="comment">// at most 25% expired</span></span><br><span class="line">    <span class="keyword">if</span> (<span class="built_in">elapsed_us</span>() &gt;= budget_us) <span class="keyword">break</span>;                   <span class="comment">// 25% of a cron tick</span></span><br><span class="line">&#125;</span><br></pre></td></tr></table></figure><p>At the default 10 Hz the budget is 25 ms per cycle, so expiry cannot stall the loop for longer than that. The budget is measured in wall-clock time, which matters later.</p><h3 id="logs-that-do-not-depend-on-when-they-are-replayed">logs that do not depend on when they are replayed</h3><p>A relative TTL replayed a day later would extend every key's life by the downtime. So each write is rewritten before it is logged.</p><table><thead><tr><th>Client sends</th><th>Logged as</th></tr></thead><tbody><tr><td><code>SET k v EX s</code>, <code>PX ms</code> or <code>EXAT s</code></td><td><code>SET k v PXAT &lt;absolute ms&gt;</code></td></tr><tr><td><code>SET k v NX</code> that succeeded</td><td><code>SET k v</code></td></tr><tr><td><code>EXPIRE k s</code>, <code>PEXPIRE k ms</code></td><td><code>PEXPIREAT k &lt;absolute ms&gt;</code></td></tr><tr><td><code>EXPIRE k</code> with a deadline already past</td><td><code>DEL k</code></td></tr><tr><td><code>DEL missing</code>, <code>SET NX</code> on an existing key</td><td>nothing</td></tr></tbody></table><p>Keys whose deadline passed while the server was down are purged in one sweep right after replay. Replay runs through a <code>CommandProcessor</code> with no propagation hook, so replayed commands are not appended to the file again.</p><h3 id="reply-after-log-and-group-commit">reply after log, and group commit</h3><p>The ordering in <code>before_sleep</code> is the durability guarantee.</p><figure class="highlight cpp"><table><tr><td class="code"><pre><span class="line"><span class="function"><span class="type">void</span> <span class="title">Server::before_sleep</span><span class="params">()</span> </span>&#123;</span><br><span class="line">    <span class="keyword">if</span> (aof_ &amp;&amp; !aof_-&gt;<span class="built_in">flush</span>() &amp;&amp; !aof_error_logged_) &#123; <span class="comment">/* log once, retry next time */</span> &#125;</span><br><span class="line">    std::vector&lt;<span class="type">int</span>&gt; batch;</span><br><span class="line">    batch.<span class="built_in">swap</span>(pending_);</span><br><span class="line">    <span class="keyword">for</span> (<span class="type">int</span> fd : batch) &#123; <span class="comment">/* ... */</span> try_write(c); &#125;</span><br><span class="line">&#125;</span><br></pre></td></tr></table></figure><p><code>Aof::flush</code> does one <code>write()</code> of everything this iteration produced, then one <code>fsync()</code> if the policy is <code>always</code>, and only then does any reply leave the process. So under <code>always</code> an acknowledged write is on disk (in the sense macOS <code>fsync</code> means, discussed below). The end-to-end test checks it directly by writing 200 keys with <code>appendfsync always</code>, killing the server with <code>SIGKILL</code>, restarting it and finding all 200.</p><p>Because the fsync is per loop iteration and not per command, it is also a group commit. With 50 busy connections, one fsync covers every write any of them made in that iteration. That turns out to be the most important performance property of the AOF.</p><p>A torn final record, from a crash in the middle of a <code>write()</code>, is truncated on startup, which is what Redis does with <code>aof-load-truncated yes</code>. Garbage anywhere else in the file stops the server from starting, on the view that corruption in the middle means something else went wrong and silently dropping everything after it is worse than refusing to run.</p><h3 id="backpressure-by-pausing-reads">backpressure by pausing reads</h3><p>A client that pipelines faster than it reads replies makes its output buffer grow without bound. Redis disconnects such a client at a hard limit. I stop reading from it once its unsent output passes a soft limit, 1 MB by default, and resume when it drains. Because the check runs before every command, the output buffer is bounded by the soft limit plus one reply, and a fast pipelining client still makes progress at the speed it drains. The resume path has a subtlety that is in <a href="#problems">problems</a>.</p><p>Two smaller choices keep the common path cheap. Write interest is registered with the poller only when <code>send()</code> hits <code>EAGAIN</code>, so a normal reply costs no extra <code>kevent</code> call. And readiness is level-triggered with one 16 KB read per wakeup, which is simpler than draining every socket under edge triggering and fairer, since one busy client cannot monopolize an iteration.</p><h3 id="testing">testing</h3><p>Unit tests cover the parser, the glob matcher, the store and command semantics. In-process TCP tests run the real server on an ephemeral port and cover a 20,000 command pipeline, an 8 MB value, 16 clients doing 2,000 <code>INCR</code>s each, backpressure and AOF restart. <code>scripts/client_test.py</code>, an independent stdlib RESP client, drives the real binary through every command, real-time expiry, <code>SIGTERM</code>, <code>SIGKILL</code> recovery and a torn tail, standing in for <code>redis-cli</code>.</p><h2 id="problems">problems</h2><h3 id="1-apple-clang-s-sanitizer-runtime-hangs">1. Apple clang's sanitizer runtime hangs</h3><p>The Debug build with <code>-fsanitize=address,undefined</code> hung at startup before reaching <code>main</code>. It is not my code. An empty <code>int main(){return 0;}</code> built with Apple clang 17 and <code>-fsanitize=address</code> hangs the same way on this macOS 26.5 machine, and it was still hanging when I rechecked at the end, killed by a 15 second timeout.</p><p>The fix was to point the <code>asan</code> CMake preset at Homebrew LLVM's <code>clang++</code> and keep Apple clang for the release build. Both builds passed the same 50 tests in my runs. The independent rebuild covered only the release build.</p><h3 id="2-a-backpressure-test-that-passed-in-one-build-and-failed-in-the-other">2. a backpressure test that passed in one build and failed in the other</h3><p>The first version of <code>Server.BackpressurePausesAndResumesReads</code> relied on the default 1 MB soft limit being crossed. Whether it was crossed depended on how fast the kernel drained socket buffers relative to how fast the server produced replies, which differed between the fast <code>-O2</code> build and the slow ASan build. The assertion that reads had been paused was not deterministic.</p><p>The test now sets a 64 KB soft limit and pipelines 30,000 <code>GET</code>s of a 1 KB value, about 30 MB of replies, before reading anything. With the limit that far below the reply volume, the pause always happens regardless of build speed. It then reads all 30,000 replies, checks each one, and checks the pause counter is nonzero.</p><h3 id="3-a-paused-client-can-stall-forever">3. a paused client can stall forever</h3><p>This one is a hazard the design had to close rather than a bug I watched happen, but it is the kind that would only show up as a hung client in production. When a client is paused, its input buffer may already hold complete commands that were read but not executed. When its output drains, re-enabling read interest is not enough. The socket may have nothing new to read, so no readable event ever arrives, and the buffered commands sit there forever.</p><p>So the resume path runs them itself.</p><figure class="highlight cpp"><table><tr><td class="code"><pre><span class="line"><span class="keyword">if</span> (!c.read_paused) <span class="keyword">return</span> <span class="literal">true</span>;</span><br><span class="line"><span class="comment">// Output drained. Its input buffer may already hold complete commands that</span></span><br><span class="line"><span class="comment">// no new read event will ever announce, so process them now.</span></span><br><span class="line">c.read_paused = <span class="literal">false</span>;</span><br><span class="line">poller_-&gt;<span class="built_in">set_read</span>(c.fd, <span class="literal">true</span>);</span><br><span class="line"><span class="built_in">process_input</span>(c);</span><br><span class="line"><span class="keyword">if</span> (c.out.<span class="built_in">empty</span>()) <span class="keyword">return</span> <span class="literal">true</span>;</span><br><span class="line"><span class="comment">// otherwise loop and try to send what that produced</span></span><br></pre></td></tr></table></figure><p>The backpressure test exercises it, since the client in that test sends its whole pipeline before reading and the server necessarily pauses with commands still buffered.</p><h3 id="4-benchmarking-on-a-machine-twenty-times-oversubscribed">4. benchmarking on a machine twenty times oversubscribed</h3><p>The first full benchmark run happened while other heavy jobs were running. <code>results/run1_loaded/bench_pipeline.jsonl</code> records a 1 minute load average of 364.7 at the start, on 14 cores. The numbers looked like results and were not.</p><figure data-figure="chart:projects/redis-from-scratch/redis-from-scratch-loaded"></figure><p>SET at 50 connections without pipelining gave a median of 17,325 ops/s with a p99 of 51.9 ms, against 107,160 ops/s and 935.5 µs once the machine was quieter. The instability was as telling as the level. The three <code>workload-mix</code> repetitions in the loaded run gave 27,832, 40,969 and 480,780 ops/s (<code>results/run1_loaded/bench_workload.jsonl</code>), a factor of 17 between runs of the same thing.</p><p>I noticed because the p99 numbers were absurd. Two changes made the data usable. First, every row now records the load average before and after the run and the CPU time the server process used, read from <code>ps</code>. Throughput per server CPU second is a number that load degrades much less, because it measures what the server did while it had a core rather than how often it got one. At depth 1 it was 92,593 in the loaded run against 112,360 in the quiet one. When the server is saturated, wall-clock throughput and throughput per CPU second should agree, which is a quick check that the server, not the load generator, was the bottleneck. Second, when the load dropped I reran everything one experiment at a time, with <code>kvbench</code> capped at 4 client threads so the benchmark itself added little load. The loaded run is kept in <code>results/run1_loaded/</code> rather than deleted.</p><p>The CPU time comes from <code>ps</code>, which has 10 ms resolution. The shortest runs used 0.33 s of server CPU, so rounding alone can move a per CPU second figure by about 3 percent. It is a coarse instrument.</p><h3 id="5-the-expiry-experiment-raced-itself">5. the expiry experiment raced itself</h3><p>The expiry experiment loads 200,000 keys that all expire at one absolute instant, via <code>PXAT</code>, and then watches <code>DBSIZE</code>. If the deadline is chosen before encoding and sending the keys, a slow machine can reach the deadline while keys are still loading, and the drop smears across the load.</p><p>The script now encodes every command first, measures how long loading 200,000 persistent keys took, and sets the deadline to twice that plus 2 seconds. It records the margin it actually achieved, 1,779 ms before the deadline in the reported run (<code>results/bench_expiry.jsonl</code>, the <code>expiry-meta</code> row).</p><h3 id="6-a-run-too-short-to-measure">6. a run too short to measure</h3><p>My first single-connection fsync experiment used 20,000 requests per run. It showed the effect, but the AOF-off baseline swung so much between repetitions that the noise was larger than the thing being measured. I raised it to 100,000 requests per run and replaced the file, so the short run is not in <code>results/</code>. Even at 100,000 requests the AOF-off runs ranged from 27,562 to 44,434 ops/s (<code>results/bench_fsync1.jsonl</code>), which is worth keeping in mind below.</p><h3 id="7-the-session-was-interrupted">7. the session was interrupted</h3><p>Work stopped partway through to reduce CPU load on the shared machine. The code, tests and loaded run survived, and on resuming I rebuilt both configurations with <code>-j2</code>, ran the two test suites one at a time, and redid the benchmarks as described above.</p><h2 id="experiments">experiments</h2><p>All experiments ran on an Apple M4 Pro (14 cores, 48 GB) over loopback TCP, with the release build (<code>-O2</code>, Apple clang 17), <code>kvbench</code> using 4 client threads, and 3 repetitions per configuration. I report medians of each metric over the three runs. <code>scripts/run_benchmarks.py</code> writes the raw <code>results/bench_*.jsonl</code> files, one JSON line per run, and <code>scripts/plot_results.py</code> builds <code>results/summary.csv</code> and the charts. The 1 minute load average during the reported runs was between 7.2 and 12.9, recorded per row.</p><p><code>kvbench</code> is a closed-loop generator. Each connection keeps up to P requests in flight, topping the window back up as replies arrive. Latency is measured from the moment a request's batch was written to the moment its reply was parsed, so at P above 1 it includes time queued behind earlier requests in the same pipeline. That is the latency a pipelining client actually sees.</p><ol><li><strong>Pipeline depth.</strong> 50 connections, SET with 16 B values, depth 1 to 128, AOF off.</li><li><strong>Connection count.</strong> Depth 1, SET, 1 to 256 connections, AOF off.</li><li><strong>Workload mix.</strong> 50 connections, depth 16, 100,000 keys, pure SET against pure GET against a 50/50 mix.</li><li><strong>Persistence with many clients.</strong> 50 connections, SET at depth 1 and 16, with AOF off and with <code>appendfsync</code> set to <code>no</code>, <code>everysec</code> and <code>always</code>.</li><li><strong>Persistence with one client.</strong> 1 connection, depth 1, 100,000 SETs, the same four modes.</li><li><strong>Active expiry.</strong> 200,000 persistent keys plus 200,000 keys sharing one deadline, no client touching them afterwards, <code>DBSIZE</code> sampled about every 50 ms.</li><li><strong>AOF replay.</strong> Fill an AOF with 100,000, 400,000 and 1,600,000 SETs, then time replay at startup.</li></ol><p>There is no side-by-side run against real Redis. The Python client in <code>scripts/client_test.py</code> plays the interoperability role, which checks correctness against the protocol as I understood it, not against Redis's behavior.</p><h2 id="results">results</h2><h3 id="pipelining-is-the-whole-story-for-throughput">pipelining is the whole story for throughput</h3><figure data-figure="chart:projects/redis-from-scratch/redis-from-scratch-pipeline"></figure><p>At 50 connections SET throughput went from 107,160 ops/s at depth 1 to 999,599 at depth 16 and 2,942,811 at depth 128, with no change to the server (<code>results/bench_pipeline.jsonl</code>). Throughput per server CPU second tracks wall-clock throughput within about 5 percent at every depth, 112,360 against 107,160 at depth 1 and 2,941,176 against 2,942,811 at depth 128. The single server thread was the bottleneck throughout, not the load generator.</p><p>That makes the per-command cost easy to read off. At depth 1 one SET costs the server about 1 / 112,360 s, or 8.9 µs of CPU. At depth 128 it costs about 0.34 µs. The hash table insert did not get 26 times cheaper. What changed is how many commands share each <code>read()</code>, <code>send()</code> and <code>kevent()</code>.</p><p>The spread between runs grows with depth. At depth 64 the three runs gave 1,187,745, 2,276,344 and 2,631,208 ops/s, so the curve's exact shape above depth 16 should not be read closely. The p99 at depth 16 (1.74 ms) is also higher than at depth 32 (1.48 ms), which is the kind of wobble to expect from three runs on a shared machine.</p><h3 id="little-s-law-checks-out">Little's law checks out</h3><p>With 50 connections at depth 1, the law predicts a mean latency of 50 / 107,160 s = 467 µs, and the measured p50 is 438.0 µs. At depth 128 there are 6,400 requests in flight, predicting 2.17 ms against a measured p50 of 2.05 ms. The law is about the mean and I am comparing against the median, so close agreement is what I would expect rather than a proof, but a load generator that miscounted in-flight requests or timestamps would fail this check.</p><p>The p99 stayed between 1.7 and 2.4 times the p50 at every depth, 4.39 ms against 2.05 ms at depth 128. No connection was being starved, which is what level-triggered, one-read-per-event dispatch is supposed to buy.</p><h3 id="more-connections-do-not-help-without-pipelining">more connections do not help without pipelining</h3><p>One connection did 40,193 ops/s at a p50 of 21.7 µs (<code>results/bench_clients.jsonl</code>). Its throughput per server CPU second was 111,111, almost three times the wall-clock rate, which says the server was idle most of the time, waiting for the client's round trip. Throughput rose to 115,697 at 8 connections and 120,252 at 16, and then stayed between 104,740 and 114,065 all the way to 256. Past that point each extra connection only adds queueing. The p50 at 256 connections was 2.18 ms, again what Little's law predicts (256 / 112,758 s = 2.27 ms).</p><p>So at depth 1 the server saturates at around 115,000 to 120,000 SETs per second on this machine, and the only way past that is to put more commands in each system call.</p><h3 id="reads-and-writes-cost-about-the-same">reads and writes cost about the same</h3><p>At 50 connections and depth 16, GET did 1,198,850 ops/s, SET 1,072,646 and the 50/50 mix 1,046,355 (<code>results/bench_workload.jsonl</code>). The spread between repetitions is larger than the gap between workloads. SET alone ranged from 1,032,764 to 1,240,625 and GET from 940,801 to 1,269,237.</p><h3 id="with-many-clients-appendfsync-always-was-almost-free">with many clients, appendfsync always was almost free</h3><p>At 50 connections and depth 16, AOF off gave 1,217,265 ops/s and <code>always</code> gave 1,042,970, with <code>no</code> at 987,323 and <code>everysec</code> at 1,077,584 (<code>results/bench_aof.jsonl</code>). <code>no</code> coming out slowest, when it does strictly less work than <code>always</code>, shows that the differences between modes here are mostly noise. At depth 1 the four modes landed between 112,334 and 118,052.</p><p>The reason <code>always</code> is cheap is group commit. With 50 busy connections one fsync per loop iteration covers every write made in that iteration, so the fsync cost is divided among dozens of commands. Each AOF was 70.8 MB for 1.2 million SETs across the depth 1 and depth 16 runs, 59 bytes per logged command.</p><h3 id="with-one-client-always-nearly-triples-latency">with one client, always nearly triples latency</h3><figure data-figure="chart:projects/redis-from-scratch/redis-from-scratch-fsync"></figure><p>On a single connection at depth 1 there is no one to share the fsync with. Median latency went from 22.2 µs with AOF off to 59.0 µs with <code>always</code>, and throughput from 40,120 to 14,165 ops/s (<code>results/bench_fsync1.jsonl</code>). Against the <code>no</code> mode, which also writes the file but never fsyncs on the reply path, the gap is 26.5 µs, so a plain <code>fsync</code> on this machine costs a few tens of microseconds.</p><p>That is cheap because macOS <code>fsync</code> pushes data to the drive but does not flush the drive's own write cache. <code>F_FULLFSYNC</code> would, at a much higher cost, and I did not measure it. So <code>always</code> here protects against a process crash, which the <code>SIGKILL</code> test demonstrates, and not against power loss.</p><p>Writing the AOF without an fsync on the reply path (<code>no</code> and <code>everysec</code>) put the p50 at 30.7 to 32.5 µs, 8 to 10 µs above AOF off, which I would attribute to the extra <code>write()</code> per iteration. I hold that loosely. In the first repetition AOF off measured 28.0 µs and <code>no</code> 29.5 µs, almost no gap, and the per-run medians of the AOF-off mode ranged from 21.5 to 28.0 µs.</p><h3 id="active-expiry-is-fast-when-the-machine-is-quiet">active expiry is fast when the machine is quiet</h3><figure data-figure="chart:projects/redis-from-scratch/redis-from-scratch-expiry"></figure><p>All 200,000 TTL keys shared one deadline and no client touched them, so only active expiry could remove them. <code>DBSIZE</code> was still 400,000 at 63 ms after the deadline, 294,740 at 116 ms and 200,000 at 171 ms, and stayed there (<code>results/bench_expiry.jsonl</code>). The server's shutdown log confirms 200,000 keys were expired actively.</p><p>That is two 10 Hz cron cycles, the first removing about 105,000 keys and the second the remaining 95,000. Each cycle is allowed at most 25 ms. If each one ran to its budget, which is my inference and was not measured, deleting a sampled expired key costs about 0.24 µs.</p><p>Under heavy load the same experiment took until 791 ms after the deadline to finish, in eight uneven steps, one of which removed only 20 keys (<code>results/run1_loaded/bench_expiry.jsonl</code>, load average 222). The cycle budget is wall-clock time, so a server that gets descheduled in the middle of a cycle burns its budget without doing work. That bounds how long clients wait, but it means expiry slows down exactly when the machine is busiest.</p><h3 id="replay-grows-faster-than-linearly">replay grows faster than linearly</h3><p>Replaying 100,000 SETs (5.9 MB) took 24 ms, 400,000 (23.6 MB) took 121 ms and 1,600,000 (94.4 MB) took 854 ms (<code>results/bench_replay.jsonl</code>). Per command that is 240 ns, 302 ns and 534 ns. Wall time and CPU time agree to within a few milliseconds (850 ms of CPU for the largest), so this is not waiting on disk.</p><p>My guess, which I have not tested, is some mix of two things. The hash table and its node allocations fall out of cache as the keyspace grows, and <code>std::unordered_map</code>'s all-at-once rehashes get more expensive as the table grows. Calling <code>reserve</code> before replay would test the second half of that.</p><p>The loaded run makes a separate point about wall time. At load average 251 to 257, the three 1.6 million command replays took 1,685, 4,785 and 3,601 ms of wall time but 1,248, 1,241 and 1,233 ms of CPU (<code>results/run1_loaded/bench_replay.jsonl</code>). The server's own work barely changed. The scheduler added the rest. That is why the startup log prints both.</p><h3 id="measurement-conditions-mattered-more-than-any-tuning-i-did">measurement conditions mattered more than any tuning I did</h3><p>Load cost the pipeline experiment 5.5 to 10 times its throughput and 37 to 55 times its p99 (the loaded chart above). Nothing I changed in the code moved any number by that much.</p><h2 id="what-i-would-change">what I would change</h2><h3 id="replace-the-standard-hash-map-with-an-incremental-rehash-table">replace the standard hash map with an incremental-rehash table</h3><p>Redis's <code>dict</code> keeps two tables during a resize and moves a few buckets per operation, so no single command pays for the whole rehash. <code>std::unordered_map</code> rehashes everything at once, which is a latency spike for every client each time the table doubles. The replay numbers suggest something that scales badly is already visible at a million keys, and an open-addressing table would also drop the per-key node allocation. The dense TTL vector relies on node stability, so it would need to store keys or indices instead of pointers.</p><h3 id="move-the-everysec-fsync-off-the-loop">move the everysec fsync off the loop</h3><p>It runs inline in the cron now, so a slow disk shows up as client latency. Redis does it on a background thread.</p><h3 id="add-aof-rewrite">add AOF rewrite</h3><p>The file only grows, including records for keys that were later deleted or expired. A rewrite that snapshots the live keyspace into a fresh file, with a <code>fork()</code> for copy-on-write, is the v2 milestone.</p><h3 id="measure-against-real-redis">measure against real Redis</h3><p>Installing <code>redis-server</code> and running the same <code>kvbench</code> sweep against both would say whether these numbers are good, not just consistent with each other. Running <code>redis-benchmark</code> and <code>redis-cli</code> against kvd would be the real interoperability test that the Python client only stands in for.</p><h3 id="build-and-run-the-epoll-backend">build and run the epoll backend</h3><p>It is written against the same interface but has never been compiled. Until it runs on Linux, the portable poller is a claim, not a result.</p><h3 id="benchmark-on-a-quiet-machine-from-the-start">benchmark on a quiet machine from the start</h3><p>I lost the first full benchmark run to other jobs. Recording load and server CPU per row from the first run, and checking them before trusting a number, would have saved a rerun.</p><h2 id="reproducibility">reproducibility</h2><p>Build. This needs CMake 3.24 or newer, Ninja and a C++20 compiler, and GoogleTest is fetched at configure time. The <code>asan</code> preset uses Homebrew LLVM at <code>/opt/homebrew/opt/llvm/bin/clang++</code> because Apple clang 17's ASan runtime hangs on macOS 26.5.</p><figure class="highlight sh"><table><tr><td class="code"><pre><span class="line"><span class="built_in">cd</span> projects/02-kvd-resp-server</span><br><span class="line">cmake --preset release &amp;&amp; cmake --build --preset release   <span class="comment"># -O2, build/release/</span></span><br><span class="line">cmake --preset asan    &amp;&amp; cmake --build --preset asan      <span class="comment"># Debug + ASan/UBSan, build/asan/</span></span><br></pre></td></tr></table></figure><p>Tests. Each preset runs the 49 GoogleTest cases and the end-to-end Python client test.</p><figure class="highlight sh"><table><tr><td class="code"><pre><span class="line">ctest --preset release</span><br><span class="line">ctest --preset asan</span><br><span class="line">python3 scripts/client_test.py --server build/release/kvd   <span class="comment"># e2e test alone</span></span><br></pre></td></tr></table></figure><p>Run the server and poke it.</p><figure class="highlight sh"><table><tr><td class="code"><pre><span class="line">build/release/kvd --port 6379 --aof appendonly.aof --appendfsync everysec</span><br><span class="line"><span class="built_in">printf</span> <span class="string">&#x27;SET a 1\r\nGET a\r\n&#x27;</span> | nc 127.0.0.1 6379</span><br></pre></td></tr></table></figure><p>Benchmarks and plots. The runner writes one JSON line per run to <code>results/bench_*.jsonl</code>, with the load average and server CPU seconds in each row. The reported numbers used <code>--threads 4</code>.</p><figure class="highlight sh"><table><tr><td class="code"><pre><span class="line">build/release/kvbench -p 6379 -c 50 -P 16 -n 1000000 -w <span class="built_in">set</span>   <span class="comment"># one run, human readable</span></span><br><span class="line">python3 scripts/run_benchmarks.py --reps 3 --threads 4</span><br><span class="line">python3 -m venv .venv &amp;&amp; .venv/bin/pip install matplotlib</span><br><span class="line">.venv/bin/python scripts/plot_results.py                      <span class="comment"># results/summary.csv and PNGs</span></span><br></pre></td></tr></table></figure><p>Code is in <code>projects/02-kvd-resp-server</code>, with the server in <code>src/</code>, the load generator in <code>bench/loadgen.cpp</code>, the tests in <code>tests/</code>, the benchmark runner, plotter and end-to-end client in <code>scripts/</code>, the raw results in <code>results/</code> (and the loaded run in <code>results/run1_loaded/</code>), the invariants and trade-offs in <code>DESIGN.md</code>, and the build log in <code>DEVLOG.md</code>.</p>]]>
    </content>
    <id>https://projects.farhansadeek.com/posts/group-commit-in-a-redis-clone/</id>
    <link href="https://projects.farhansadeek.com/posts/group-commit-in-a-redis-clone/"/>
    <published>2026-09-29T06:17:09.000Z</published>
    <summary>A single-threaded Redis-compatible server in C++ with TTLs, pipelining and an append-only file. One fsync per event-loop pass makes durability nearly free with many clients.</summary>
    <title>Group Commit in a Single-Threaded Redis Clone</title>
    <updated>2026-09-29T06:49:26.552Z</updated>
  </entry>
  <entry>
    <author>
      <name>Farhan Sadeek</name>
      <email>farhan@farhansadeek.com</email>
    </author>
    <category term="databases" scheme="https://projects.farhansadeek.com/tags/databases/"/>
    <category term="storage" scheme="https://projects.farhansadeek.com/tags/storage/"/>
    <category term="cpp" scheme="https://projects.farhansadeek.com/tags/cpp/"/>
    <content>
      <![CDATA[<p>I wrote kvdb, a persistent single-node key-value store in C++20 with no runtime dependencies. It keeps a B+ tree in 4 KiB pages in one file, caches pages in a buffer pool with LRU eviction and pin counts, logs every commit to a write-ahead log (WAL) with a configurable sync policy, makes checkpoints crash-atomic with a page journal, and recovers after a crash by replaying the log. A small REPL sits on top.</p><p>The code is about 1,700 lines of C++ in <code>src/</code> and <code>include/</code>, headers and comments included, plus about 930 lines of GoogleTest. There are <strong>36 tests</strong>, and on 2026-09-27 an independent clean release rebuild passed all 36, including the tests that fork a writer process and kill it. The AddressSanitizer plus UBSan build also passed all 36 in my own run, but it was not rerun independently, and neither were the benchmarks.</p><p>The headline result is the crash testing. A child process writes and acknowledges each commit over a pipe, and the parent kills it with SIGKILL. In the verified release run, <strong>12 random kills covering 9,721 acknowledged operations lost none of them</strong>. A second test aims kills into the middle of a large checkpoint, and <strong>6 of its 10 killed rounds</strong> landed in the window where the journal was complete but the pages were only partly written in place, and were repaired from the journal. A third test crashes the process at each of <strong>5 named steps</strong> of the checkpoint protocol and then crashes it again at the same step inside recovery. Every one of these reopens the database, walks the whole tree checking its invariants, and compares it to a model. These are process crashes. The OS page cache survives them, so they say nothing about power loss, which I did not test.</p><p>The second story is what durability costs on a Mac. With no sync a commit took <strong>1.2 µs</strong> at the median. With <code>fsync</code> it took <strong>24 µs</strong>, and with <code>F_FULLFSYNC</code>, the only call on macOS that makes the drive flush its own cache, it took <strong>4.01 ms</strong>, about 3,300 times the unsynced cost. Batching 256 puts into one <code>F_FULLFSYNC</code> commit brought throughput from 243 to <strong>59,450 puts per second</strong>. All benchmark numbers come from a shared Apple M4 Pro that was running other jobs, with the 1 minute load average recorded in every row (3.9 to 6.3 for the reported runs). They are indicative, not a clean benchmark. The 1M key database also fits in memory, so the throughput numbers measure CPU and system call cost, not the disk.</p><p>Code is in <code>projects/01-kvdb-btree-wal-engine</code>. The argument of this post is that a small storage engine becomes trustworthy through its crash tests rather than its design document, and that on macOS &quot;durable&quot; is a choice between three very different prices. Skip to <a href="#problems">problems</a> for what went wrong.</p><h2 id="table-of-contents">table of contents</h2><ul><li><a href="#what-i-wanted-to-build">what I wanted to build</a></li><li><a href="#theory">theory</a></li><li><a href="#architecture">architecture</a></li><li><a href="#implementation">implementation</a></li><li><a href="#problems">problems</a></li><li><a href="#experiments">experiments</a></li><li><a href="#results">results</a></li><li><a href="#what-i-would-change">what I would change</a></li><li><a href="#reproducibility">reproducibility</a></li></ul><h2 id="what-i-wanted-to-build">what I wanted to build</h2><p>I wanted a storage engine I understand all the way down, from bytes in a 4 KiB page, to a page in a buffer frame, to a frame pinned by a tree operation, to a commit in a log record, to a crash that loses nothing it promised to keep.</p><p>The v0 scope was fixed up front. A pager over one file. A buffer pool with LRU eviction and pin counts. A B+ tree with insert, point lookup, range scan and lazy delete. A WAL with fsync and redo recovery on startup. A REPL. A crash test that really kills a process. And benchmarks against <code>std::map</code> and SQLite at 1M keys.</p><p>The bar was correctness first. Every claim about durability had to be backed by a test that kills a process with SIGKILL and checks what survives against a model, and every performance number had to come from a run saved under <code>results/</code>. Transactions, concurrency, page checksums, real delete and an LSM variant are later milestones and are not in this post.</p><h2 id="theory">theory</h2><h3 id="b-trees-on-disk">B+ trees on disk</h3><p>A disk-based B+ tree stores sorted keys in fixed-size pages. Internal pages hold separator keys and child pointers. Leaves hold the data and are linked left to right, so a range scan is one descent followed by a walk along the leaf chain. The point of the shape is fanout. In kvdb an internal cell for a 16-byte key is 24 bytes plus a 2-byte slot, so a 4 KiB page holds about 150 children, and 1M keys fit in a tree of height 4. A lookup touches four pages and does a binary search inside each.</p><p>Splits decide how full the leaves are. A textbook split cuts a full node in half, and under random inserts leaves settle at about 69% full (ln 2). Under ascending inserts the left half is never touched again and stays half empty, so the append pattern deserves a special case.</p><h3 id="the-write-ahead-rule">the write-ahead rule</h3><p>Durability rests on one rule. A change is committed once a description of it is on stable storage, and that description must get there before the data pages change on disk. After a crash, the committed state is rebuilt from the last consistent page file plus the log.</p><p>The general design is ARIES, which is STEAL (a dirty page may be written back before its transaction commits) and NO-FORCE (pages need not be written at commit). It is flexible and it needs both undo and redo logging, page LSNs to know which changes a page already contains, and a three-pass recovery.</p><p>I took a simpler corner of the design space. kvdb's buffer pool is NO-STEAL, meaning dirty pages never reach the page file except at a checkpoint. The page file on disk is therefore always exactly the state as of the last checkpoint, and the log only needs logical redo records (&quot;put k v&quot;). There is no undo, no page LSN, and recovery is &quot;load the snapshot, redo the operations after it&quot;.</p><h3 id="making-the-checkpoint-atomic">making the checkpoint atomic</h3><p>The remaining problem is the checkpoint itself. It overwrites many pages in place, and a crash halfway leaves a parent from the new generation pointing at a child from the old one, which is a corrupt tree. The fix is a double write. Copy every page the checkpoint will overwrite to a side file, sync it, and only then write in place. A crash before the journal is complete leaves the page file untouched. A crash after it leaves a complete journal that recovery replays. SQLite's rollback journal and InnoDB's doublewrite buffer do versions of the same thing.</p><p>The subtle part is deciding whether a journal is complete. Writing the header last is not enough, because without a barrier between them the header can reach the disk before the body. So kvdb accepts a journal only if the magic matches, the file size equals exactly the size the header's page count implies, and a CRC over the entire body matches. A journal cut short at any byte fails one of those checks.</p><h3 id="three-durability-levels-on-macos">three durability levels on macOS</h3><p>On macOS, <code>fsync()</code> pushes data to the drive but does not make the drive flush its volatile write cache. Only <code>fcntl(fd, F_FULLFSYNC)</code> does that. So there are three levels, and kvdb exposes all of them as <code>SyncMode</code>.</p><table><thead><tr><th>Mode</th><th>What it survives</th><th>Call after each WAL append</th></tr></thead><tbody><tr><td><code>kNone</code></td><td>a process crash</td><td>none</td></tr><tr><td><code>kFsync</code> (default)</td><td>a process or OS crash</td><td><code>fsync</code></td></tr><tr><td><code>kFullFsync</code></td><td>power loss, if the drive honors the flush</td><td><code>fcntl(F_FULLFSYNC)</code></td></tr></tbody></table><p><code>kNone</code> survives a process crash because a <code>pwrite</code> that has returned is in the kernel's page cache, and the kernel writes it out whether or not the process is still alive. That matters for reading the crash tests correctly, as I explain below.</p><h2 id="architecture">architecture</h2><p>A database at path <code>P</code> is three files. <code>P</code> is the page file, where page 0 is a meta page and every other page is a B+ tree node. <code>P-wal</code> holds one CRC-framed record per committed batch since the last checkpoint. <code>P-journal</code> is empty except while a checkpoint runs.</p><figure data-figure="diagram:kvdb-write-and-recovery"></figure><p>The meta page holds the magic number, format version, root page id, page count, tree height and checkpoint LSN. It lives in the buffer pool like any node, so a root split or a page allocation is checkpointed atomically with the nodes it refers to. There is no separate &quot;superblock write&quot; to get wrong.</p><p>Nodes are slotted pages. A 16-byte header is followed by an array of 2-byte slot offsets sorted by key, growing toward the end of the page, while cells (<code>klen, vlen, key, value</code>) grow backward from the end.</p><figure data-figure="diagram:kvdb-slotted-page"></figure><p>A leaf's <code>link</code> is its right sibling. An internal node's <code>link</code> is its leftmost child, and each cell's value is a 4-byte child id. Keys are capped at 128 bytes and values at 512, which makes a worst-case cell plus slot 646 bytes, so any overflowing node holds at least 6 cells and a byte-balanced split always produces two halves that fit. That cap is what lets v0 skip overflow pages.</p><p>A WAL record is <code>u32 crc32 | u32 len | u64 lsn | payload</code>, and the payload is the <code>WriteBatch</code> encoding itself, so logging a batch is one copy and one <code>pwrite</code>. The journal is <code>u64 magic | u32 page count | u32 crc32(body)</code> followed by <code>(u32 page id, 4096 bytes)</code> per page.</p><p>Three invariants carry the design. Between checkpoints the page file is byte-identical to the last completed checkpoint. A batch is committed if and only if its WAL record is intact on disk. And replaying a batch over a page file that already contains some or all of it gives the same result, because putting or deleting a whole key is idempotent.</p><h2 id="implementation">implementation</h2><h3 id="the-write-path">the write path</h3><p>A write is four lines, and the order is the whole durability argument.</p><figure class="highlight cpp"><table><tr><td class="code"><pre><span class="line"><span class="function"><span class="type">void</span> <span class="title">DB::write</span><span class="params">(<span class="type">const</span> WriteBatch&amp; batch)</span> </span>&#123;</span><br><span class="line">  <span class="keyword">if</span> (batch.<span class="built_in">count</span>() == <span class="number">0</span>) <span class="keyword">return</span>;</span><br><span class="line">  <span class="built_in">maybe_checkpoint</span>(<span class="literal">true</span>);</span><br><span class="line">  <span class="type">const</span> <span class="type">uint64_t</span> lsn = next_lsn_++;</span><br><span class="line">  wal_-&gt;<span class="built_in">append</span>(lsn, batch.<span class="built_in">rep</span>(), opts_.sync);  <span class="comment">// commit point</span></span><br><span class="line">  <span class="built_in">apply</span>(batch.<span class="built_in">rep</span>());</span><br><span class="line">  applied_lsn_ = lsn;</span><br><span class="line">  ++stats_.commits;</span><br><span class="line">&#125;</span><br></pre></td></tr></table></figure><p><code>put</code> and <code>del</code> are one-op batches, and deleting a missing key is not logged at all. <code>Wal::append</code> frames the record, computes the CRC with the ARMv8 CRC32 instructions (8 bytes per instruction, with a table fallback that a test checks agrees), issues a single <code>pwrite</code>, and syncs according to the mode. The WAL file is grown in 4 MiB <code>ftruncate</code> steps, so most appends land inside the file and the file has a zero-filled tail that replay must treat like garbage.</p><h3 id="replay-stops-at-the-first-thing-it-does-not-trust">replay stops at the first thing it does not trust</h3><figure class="highlight cpp"><table><tr><td class="code"><pre><span class="line"><span class="keyword">while</span> (off + kHeaderSize &lt;= data.<span class="built_in">size</span>()) &#123;</span><br><span class="line">  <span class="type">const</span> <span class="type">uint32_t</span> crc = <span class="built_in">load</span>&lt;<span class="type">uint32_t</span>&gt;(data.<span class="built_in">data</span>() + off);</span><br><span class="line">  <span class="type">const</span> <span class="type">uint32_t</span> len = <span class="built_in">load</span>&lt;<span class="type">uint32_t</span>&gt;(data.<span class="built_in">data</span>() + off + <span class="number">4</span>);</span><br><span class="line">  <span class="keyword">if</span> (len &gt; data.<span class="built_in">size</span>() - off - kHeaderSize) <span class="keyword">break</span>;          <span class="comment">// torn tail</span></span><br><span class="line">  <span class="keyword">if</span> (<span class="built_in">crc32</span>(data.<span class="built_in">data</span>() + off + <span class="number">4</span>, <span class="number">12</span> + len) != crc) <span class="keyword">break</span>;  <span class="comment">// corrupt</span></span><br><span class="line">  <span class="type">const</span> <span class="type">uint64_t</span> lsn = <span class="built_in">load</span>&lt;<span class="type">uint64_t</span>&gt;(data.<span class="built_in">data</span>() + off + <span class="number">8</span>);</span><br><span class="line">  <span class="keyword">if</span> (lsn &lt;= last_lsn) <span class="keyword">break</span>;                                <span class="comment">// stale bytes</span></span><br><span class="line">  <span class="built_in">fn</span>(lsn, std::<span class="built_in">string_view</span>(data.<span class="built_in">data</span>() + off + kHeaderSize, len));</span><br><span class="line">  last_lsn = lsn;</span><br><span class="line">  off += kHeaderSize + len;</span><br><span class="line">&#125;</span><br></pre></td></tr></table></figure><p>A torn write, a flipped byte and the preallocated zero tail all end replay the same way, and the file is then truncated at the last good record so new records are never appended after garbage. The WAL tests cut the log at every possible byte offset of the last record and check that replay keeps exactly the records before it.</p><h3 id="the-buffer-pool">the buffer pool</h3><p>The pool is one page-aligned allocation split into frames, an <code>unordered_map</code> page table, a free list and a <code>std::list</code> LRU. The key design choice is that the LRU list contains only frames that are both unpinned and clean. Eviction is then <code>lru_.back()</code>, and it structurally cannot pick a pinned or dirty page. If every frame is pinned or dirty, <code>fetch</code> throws rather than corrupting anything. <code>PageGuard</code> is an RAII pin that carries the dirty bit to <code>unpin</code>, so a tree operation cannot forget to release a page.</p><p>Because the pool is NO-STEAL, dirty pages accumulate until a checkpoint. The DB layer checkpoints when the WAL passes 64 MiB, at clean close, at the end of recovery, and whenever the dirty count comes within 64 frames of capacity. That last trigger is where NO-STEAL shows its cost, as the small-pool benchmark shows.</p><h3 id="the-append-split">the append split</h3><p>On overflow the tree materializes the node's entries plus the new one and splits by bytes rather than by count, since cells vary in size. One special case does a lot of work.</p><figure class="highlight cpp"><table><tr><td class="code"><pre><span class="line"><span class="keyword">if</span> (idx == n - <span class="number">1</span> &amp;&amp; node::<span class="built_in">link</span>(leaf.<span class="built_in">data</span>()) == kNoPage) &#123;</span><br><span class="line">  <span class="comment">// Appending past the end of the rightmost leaf (ascending inserts): leave</span></span><br><span class="line">  <span class="comment">// the old leaf full and start a new one. Gives ~100% fill instead of 50%.</span></span><br><span class="line">  m = n - <span class="number">1</span>;</span><br><span class="line">&#125; <span class="keyword">else</span> &#123;</span><br><span class="line">  m = std::<span class="built_in">clamp</span>(<span class="built_in">byte_midpoint</span>(ents), <span class="number">1</span>, n - <span class="number">1</span>);</span><br><span class="line">&#125;</span><br></pre></td></tr></table></figure><p>The first key of the right half is pushed into the parent, and internal nodes split the same way, except the middle key moves up rather than being copied. <code>BTree::verify()</code> walks the whole tree checking each node's format, every key against the range its parent's separators allow, that all leaves are at the same depth, and that the sibling chain visits them in order. Every crash test calls it after reopening.</p><h3 id="the-checkpoint">the checkpoint</h3><figure class="highlight cpp"><table><tr><td class="code"><pre><span class="line">tree_-&gt;<span class="built_in">set_checkpoint_lsn</span>(applied_lsn_);</span><br><span class="line"><span class="type">const</span> <span class="keyword">auto</span> pages = pool_-&gt;<span class="built_in">dirty_pages</span>();        <span class="comment">// sorted by page id, meta included</span></span><br><span class="line"><span class="comment">// stream (id, page) entries to the journal body in 1 MiB chunks, CRC as we go</span></span><br><span class="line">failpoint::<span class="built_in">hit</span>(<span class="string">&quot;ckpt.before_journal_header&quot;</span>);</span><br><span class="line"><span class="built_in">pwrite_all</span>(journal_fd_, hdr, kJournalHeader, <span class="number">0</span>);</span><br><span class="line"><span class="built_in">sync_fd</span>(journal_fd_, mode);</span><br><span class="line">failpoint::<span class="built_in">hit</span>(<span class="string">&quot;ckpt.journal_synced&quot;</span>);</span><br><span class="line"><span class="keyword">for</span> (<span class="type">size_t</span> i = <span class="number">0</span>; i &lt; pages.<span class="built_in">size</span>(); ++i) &#123;</span><br><span class="line">  pager_-&gt;<span class="built_in">write_page</span>(pages[i].first, pages[i].second);</span><br><span class="line">  <span class="keyword">if</span> (i == pages.<span class="built_in">size</span>() / <span class="number">2</span>) failpoint::<span class="built_in">hit</span>(<span class="string">&quot;ckpt.mid_data&quot;</span>);</span><br><span class="line">&#125;</span><br><span class="line">pager_-&gt;<span class="built_in">sync</span>(mode);</span><br><span class="line">failpoint::<span class="built_in">hit</span>(<span class="string">&quot;ckpt.data_synced&quot;</span>);</span><br><span class="line"><span class="comment">// then truncate the WAL (failpoint ckpt.wal_reset) and the journal</span></span><br></pre></td></tr></table></figure><p>The <code>failpoint::hit</code> calls are no-ops unless a test has armed that name, in which case the process calls <code>_exit(77)</code> on the spot, with no destructors and no flushing.</p><p>Recovery is the mirror image. If the journal is valid, every page in it is written to its home location and synced, then the journal is truncated. If it is not valid, the crash happened before any in-place write, so the page file is still the previous checkpoint and the journal is simply discarded.</p><h3 id="batches-bigger-than-the-pool">batches bigger than the pool</h3><p>One case needed care. A single batch can dirty more pages than the pool holds. <code>apply</code> then checkpoints in the middle of the batch, but without truncating the WAL and without moving the checkpoint LSN past the previous batch. After a crash, recovery redoes the whole batch on top of a page file that already contains part of it, which is safe because whole-key puts and deletes are idempotent.</p><h3 id="crash-tests-with-real-processes">crash tests with real processes</h3><p>The random-kill test forks a child that opens the database with a 256-page pool and a 64 KiB WAL threshold, so checkpoints happen constantly, and runs a deterministic stream of operations. Most are single puts, about one in five is a delete, and every 25th is a 20-op batch, to check that batches are atomic. After each <code>write()</code> returns, the child writes the op index to a pipe.</p><figure class="highlight cpp"><table><tr><td class="code"><pre><span class="line"><span class="keyword">while</span> (seen &lt; kill_after &amp;&amp; ::<span class="built_in">read</span>(fds[<span class="number">0</span>], &amp;ack, <span class="keyword">sizeof</span> ack) == <span class="keyword">sizeof</span> ack) &#123;</span><br><span class="line">  last = ack;</span><br><span class="line">  ++seen;</span><br><span class="line">&#125;</span><br><span class="line">::<span class="built_in">kill</span>(pid, SIGKILL);</span><br><span class="line"><span class="comment">// Drain acks the child managed to send before dying: those were committed.</span></span><br><span class="line"><span class="keyword">while</span> (::<span class="built_in">read</span>(fds[<span class="number">0</span>], &amp;ack, <span class="keyword">sizeof</span> ack) == <span class="keyword">sizeof</span> ack) last = ack;</span><br></pre></td></tr></table></figure><p>The parent then reopens the database, runs <code>verify()</code>, and requires the contents to equal a <code>std::map</code> model after every acknowledged op, or after exactly one more. The extra one is the op that may have reached the WAL between <code>write()</code> returning and the ack being sent. Anything else fails, including a lost acknowledged op, a resurrected delete and a half-applied batch. The test runs 12 rounds on the same file, so each round also recovers on top of the previous recovery.</p><p>The <code>flock</code> that stops two processes opening the same database dies with the process, which is what lets the parent reopen immediately after a kill.</p><h2 id="problems">problems</h2><h3 id="1-the-random-kills-never-hit-a-checkpoint">1. the random kills never hit a checkpoint</h3><p>The random-kill test was supposed to be the checkpoint test too. It was not. Checkpoints are short compared with the time between them, so a kill at a random moment almost always lands in the write path. In the final release run, <strong>0 of 12 rounds</strong> found a journal to replay. The test was passing without ever exercising the journal.</p><p>I found this because the test prints how many rounds recovered from a journal, and the number was zero. I added two tests to cover the gap. <code>KillNineDuringLargeCheckpoint</code> builds a checkpoint big enough to aim at (6,000 mixed ops plus 50,000 new keys, with an 8,192-page pool), runs one uncounted calibration round that reports how long the journal and data phases take, and then kills 10 rounds at random delays between halfway through the journal write and the end of the checkpoint. In the release run the calibration checkpoint took 6.7 ms, 1.4 ms of it writing the journal and 5.4 ms writing pages in place and truncating. <strong>6 of the 10 kills</strong> landed after the journal was complete and were repaired from it. The rest landed before it was complete and were correctly discarded.</p><p>That hit rate depends on timing. Under ASan the journal phase took 39.4 ms instead of 1.4 ms, and only 2 of 10 kills landed in the repair window. So the aimed test is probabilistic, and I did not want the protocol's correctness to rest on luck. The deterministic <code>CheckpointCrash</code> tests arm a failpoint at each of five steps (before the journal header, after the journal sync, halfway through the in-place writes, after the data sync, after the WAL reset), crash there, then arm the same failpoint and open the database, so recovery's own checkpoint crashes at the same step. A third open must produce exactly the model. All five pass in both builds.</p><h3 id="2-sqlite-s-full-sync-was-faster-than-a-full-sync">2. SQLite's full sync was faster than a full sync</h3><p>With <code>synchronous=FULL</code> and <code>fullfsync=ON</code>, SQLite committed <strong>1,006 times per second</strong> against kvdb's <strong>247</strong> with <code>F_FULLFSYNC</code>. That looked like kvdb doing something wrong, until I compared latencies. A real <code>F_FULLFSYNC</code> on this machine takes about 4 ms, in both kvdb and a bare <code>pwrite</code> loop, while SQLite's median commit took <strong>0.69 ms</strong>. SQLite cannot be issuing a full flush on every commit.</p><p>My hypothesis is that Apple's system SQLite, which is the copy the benchmark links, uses the lighter <code>F_BARRIERFSYNC</code> for WAL commits, which orders writes without waiting for the cache flush. I have not verified it, with <code>fs_usage</code> or by building SQLite from source. Until I do, the <code>F_FULLFSYNC</code> column of the durability comparison is not like for like, and I mark it that way in the figure below.</p><h3 id="3-a-comment-claimed-a-14x-speedup-nobody-had-measured">3. a comment claimed a 14x speedup nobody had measured</h3><p>A comment in <code>wal.h</code> said that on APFS an append which extends the file costs about 14 times more than a write inside a preallocated file. I had not measured that, so I wrote the <code>walgrow</code> experiment. It did not hold up. Preallocation made a 128-byte append <strong>1.70 times</strong> faster with no sync, <strong>1.09 times</strong> faster with <code>fsync</code>, and made no difference with <code>F_FULLFSYNC</code> (247 versus 248 appends per second). I kept preallocation, since it costs nothing and replay already handles a garbage tail, and corrected the comment to cite the file.</p><h3 id="4-the-machine-was-busy">4. the machine was busy</h3><p>The machine was shared with a dozen other builds. An earlier benchmark run, later interrupted, happened at a 1 minute load average of about 256 on 14 cores, and its wall-clock numbers were roughly an order of magnitude below what I measured later at a load of 4 to 6. I threw it away. Every benchmark row now records ops per CPU second and the load average, the build script uses <code>-j2</code>, and I ran one benchmark or test binary at a time. Even so, <code>results/machine.txt</code> shows a 15 minute load average of 42 when the reported run started, so the machine had been very busy shortly before.</p><h3 id="5-apple-clang-s-sanitizer-runtime-hung">5. Apple clang's sanitizer runtime hung</h3><p>An early UBSan-only run under Apple clang 17 stalled inside the large-checkpoint crash test and never finished, and the sanitizer runtime hung at startup on this macOS version more generally. I moved the sanitizer preset to Homebrew LLVM 23, which runs ASan and UBSan together without trouble. Switching compilers then surfaced two smaller problems. Three files used <code>errno</code> without including <code>&lt;cerrno&gt;</code> and had only compiled because an Apple header pulled it in transitively. And Homebrew LLVM's DWARF 5 debug info made Apple's linker print a harmless warning for every object file, burying real warnings, so the sanitizer build now passes <code>-gdwarf-4</code> for upstream Clang.</p><h2 id="experiments">experiments</h2><p>All runs were on an Apple M4 Pro (14 cores, 48 GB) running macOS 26 (Darwin 25.5.0), release build at <code>-O2</code>, with the 1 minute load average recorded next to each row. Keys are 16 bytes (<code>user%012d</code>, so lexicographic order is numeric order) and values are 100 bytes.</p><ol><li><strong>Correctness.</strong> The 36 tests in 8 suites, in a release build and in an ASan plus UBSan build. The release suite was rerun independently on 2026-09-27 and passed. The sanitizer suite was not rerun.</li><li><strong>Throughput and latency at 1M keys</strong>, 3 repetitions, medians reported. For each engine, insert 1M keys in sequential or shuffled order, then read all 1M in sequential and shuffled order, then scan everything. Every operation is timed individually with <code>steady_clock</code>, so p50, p99 and p99.9 are exact. The engines are <code>std::map&lt;std::string, std::string&gt;</code>, kvdb with a 65,536-frame (256 MiB) pool, kvdb with a 4,096-frame (16 MiB) pool, and SQLite 3.51.0 with <code>journal_mode=WAL</code>, <code>synchronous=OFF</code>, a 256 MiB cache, prepared statements and a <code>WITHOUT ROWID</code> table. kvdb ran with <code>SyncMode::kNone</code>, so both disk engines write their log on every commit and neither syncs.</li><li><strong>Durability cost.</strong> 2,000 random puts, one per commit, under each sync mode for kvdb and SQLite, then kvdb with <code>F_FULLFSYNC</code> and 1, 4, 16, 64 or 256 puts per <code>WriteBatch</code>. One run each.</li><li><strong>WAL preallocation.</strong> 3,000 raw 128-byte <code>pwrite</code> appends, each followed by no sync, <code>fsync</code> or <code>F_FULLFSYNC</code>, into a file that either grows with each append or was sized in advance.</li></ol><p>What the benchmarks do not measure matters as much. With the 256 MiB pool, the whole database is resident. The pool hit rate was 1.000 for both fills, and the page file was 122.1 MiB after the sequential fill and 170.9 MiB after the random one. Even the 16 MiB pool misses only into the OS page cache, since the machine has 48 GB and the file was just written. <strong>No benchmark here reads from the SSD.</strong> The throughput numbers are the cost of CPU work and system calls, and the only experiment where the storage device is on the critical path is the durability one.</p><h2 id="results">results</h2><h3 id="crash-safety">crash safety</h3><table><thead><tr><th>Test</th><th>What it does</th><th>Release result</th></tr></thead><tbody><tr><td><code>KillNineDuringWritesLosesNothingAcknowledged</code></td><td>12 random SIGKILLs during writes, <code>kNone</code></td><td>9,721 acknowledged ops, none lost, 2,254 WAL records redone</td></tr><tr><td><code>KillNineWithFsyncCommits</code></td><td>3 SIGKILLs after 300, 500 and 700 acks, <code>kFsync</code></td><td>passed</td></tr><tr><td><code>KillNineDuringLargeCheckpoint</code></td><td>10 SIGKILLs aimed into a 6.7 ms checkpoint</td><td>all 10 exact, 6 repaired from a complete journal</td></tr><tr><td><code>CheckpointCrash</code>, 5 steps</td><td>crash at a step, then again at that step in recovery</td><td>all 5 exact</td></tr><tr><td><code>CrashWithoutCloseThenTornWalTail</code></td><td>exit without close, chop 7 bytes off the WAL, append junk</td><td>1,999 of 2,000 records recovered, the torn one dropped</td></tr></tbody></table><p>All 36 tests passed in 2.5 s in the release build (<code>results/tests_release.txt</code>) and in 34.7 s under ASan plus UBSan (<code>results/tests_asan.txt</code>). The sanitizer run's random-kill test covered 9,693 acknowledged ops with 2,212 records redone, the different count coming from different kill timing, and it also lost none.</p><p>It is worth being exact about the scope. The random-kill test runs with <code>SyncMode::kNone</code>, and with <code>kNone</code> the checkpoint's journal &quot;sync&quot; is a no-op too. That is fine for a process crash, where every returned <code>pwrite</code> is already in the kernel's page cache and ordering between files is preserved by that cache. It would not be fine for power loss, where unsynced data can vanish and writes can reach the disk out of order. None of these tests pulls the plug, so the claim is that kvdb loses no acknowledged write to a process crash at any point I could aim at, including inside checkpoints and inside recovery. The protocol is designed for power loss when run with <code>kFullFsync</code>, but that is a design argument, not a test result.</p><h3 id="what-durability-costs">what durability costs</h3><figure data-figure="chart:projects/kvdb-btree-wal-engine/kvdb-btree-wal-engine-durability"></figure><table><thead><tr><th>Mode</th><th>kvdb commits/s</th><th>kvdb p50</th><th>SQLite commits/s</th><th>SQLite p50</th></tr></thead><tbody><tr><td>no sync</td><td>693,922</td><td>1.2 µs</td><td>70,517</td><td>10.3 µs</td></tr><tr><td>fsync</td><td>38,162</td><td>24.0 µs</td><td>4,900</td><td>67.5 µs</td></tr><tr><td>F_FULLFSYNC</td><td>247</td><td>4.01 ms</td><td>1,006 (see problem 2)</td><td>0.69 ms</td></tr></tbody></table><p>Each step costs one to two orders of magnitude. Going from no sync to <code>fsync</code> cut kvdb's commit rate by a factor of 18, and going from <code>fsync</code> to <code>F_FULLFSYNC</code> cut it by another 154. At the median, a truly durable single-put commit on this Mac costs about 4 ms, which is about 3,300 times the unsynced commit.</p><p>The <code>fsync</code> row is the uncomfortable one. It is the default in kvdb, as it is in SQLite on macOS, and it costs 24 µs, which feels like durability. It is not durability against power loss, because the data may still be sitting in the drive's cache. The honest price of that guarantee is the 4 ms row.</p><h3 id="batching-is-the-only-way-to-pay-less">batching is the only way to pay less</h3><figure data-figure="chart:projects/kvdb-btree-wal-engine/kvdb-btree-wal-engine-batching"></figure><p>With <code>F_FULLFSYNC</code> on every commit, putting more operations in each <code>WriteBatch</code> scales almost perfectly. Throughput was 243, 980, 3,960, 15,485 and 59,450 puts per second at 1, 4, 16, 64 and 256 puts per batch, while the median commit stayed at 4.00 to 4.05 ms throughout. The flush is the cost and it is paid once per batch, so the 256-put batch is about 245 times faster than single puts. A <code>WriteBatch</code> is one WAL record, one <code>pwrite</code> and one sync, and it is atomic, which the random-kill test checks with its 20-op batches.</p><p>This is group commit done by hand, because v0 is single-threaded. Letting concurrent committers share one flush automatically is on the v2 list.</p><h3 id="throughput-when-everything-is-in-memory">throughput when everything is in memory</h3><figure data-figure="chart:projects/kvdb-btree-wal-engine/kvdb-btree-wal-engine-throughput"></figure><table><thead><tr><th>Workload</th><th><code>std::map</code></th><th>kvdb 256 MiB</th><th>kvdb 16 MiB</th><th>SQLite</th></tr></thead><tbody><tr><td>sequential put, ops/s</td><td>3,852,003</td><td>728,263</td><td>719,354</td><td>94,221</td></tr><tr><td>random put, ops/s</td><td>1,098,039</td><td>439,183</td><td>206,867</td><td>49,851</td></tr><tr><td>random get after random put, ops/s</td><td>840,719</td><td>1,251,868</td><td>639,082</td><td>170,354</td></tr><tr><td>sequential get after sequential put, ops/s</td><td>6,986,963</td><td>3,668,542</td><td>3,364,736</td><td>253,240</td></tr><tr><td>random put p50 / p99, ns</td><td>875 / 1,791</td><td>1,833 / 5,083</td><td>2,583 / 5,708</td><td>10,667 / 74,041</td></tr><tr><td>random get p50 / p99, ns</td><td>1,042 / 2,125</td><td>750 / 1,250</td><td>1,542 / 2,209</td><td>5,792 / 7,875</td></tr><tr><td>full scan after random put, keys/s</td><td>11,357,624</td><td>60,934,431</td><td>20,973,686</td><td>13,935,243</td></tr></tbody></table><p>These are medians of 3 runs from <code>results/summary_main.csv</code>. Given that none of it touches the SSD, I read them as follows.</p><h3 id="kvdb-beats-a-red-black-tree-on-reads">kvdb beats a red-black tree on reads</h3><p>Random gets ran at 1.25M per second in kvdb against 0.84M for <code>std::map</code>, with a p50 of 750 ns against 1,042 ns. A red-black tree filled in random order scatters 1M nodes across the heap and follows about 20 pointers per lookup, many of them likely cache misses. The B+ tree touches 4 pages and binary-searches inside each. I did not profile this, so the explanation is a plausible one rather than a measured one. The full scan shows the same effect more strongly, 61M keys per second for kvdb against 11M for <code>std::map</code> after a random fill, because a leaf chain is sequential memory and a red-black tree walk is not.</p><h3 id="kvdb-loses-to-a-red-black-tree-on-writes">kvdb loses to a red-black tree on writes</h3><p>Sequential puts were 5.3 times slower than <code>std::map</code> and random puts 2.5 times slower. Each put encodes a batch, computes a CRC, issues one <code>pwrite</code> system call to the WAL and may split nodes. The system call per put is the obvious target, and it is the first thing on the change list below.</p><h3 id="the-gap-to-sqlite-is-mostly-generality">the gap to SQLite is mostly generality</h3><p>kvdb's random puts were 8.8 times faster than SQLite's and its random gets 7.3 times faster. I would not read that as kvdb being a better storage engine. SQLite here runs prepared statements through its bytecode VM, encodes records, and maintains a far more general B-tree, while kvdb has a specialized key-value path. It is a comparison of a narrow tool against a general one, both in memory.</p><h3 id="the-append-split-works">the append split works</h3><p>Sequential fill produced 30,304 leaves at 99% fill. Random fill produced 43,324 leaves at 69% fill, which is the textbook ln 2 value for random inserts. Both trees have height 4, and the random-fill file is 40% larger for the same data.</p><h3 id="no-steal-turns-memory-pressure-into-checkpoints">NO-STEAL turns memory pressure into checkpoints</h3><p>With a 16 MiB pool, random fill needed <strong>213 checkpoints</strong> against 3 for the 256 MiB pool, and ran 2.1 times slower (207K versus 439K puts per second, pool hit rate 0.832). Random inserts dirty leaves all over the tree, and a NO-STEAL pool cannot write any of them back without a full checkpoint. Sequential fill barely noticed (9 checkpoints, 719K versus 728K per second, hit rate 0.988), because it only ever dirties the rightmost path.</p><h3 id="wal-preallocation-is-a-small-win">WAL preallocation is a small win</h3><table><thead><tr><th>Sync</th><th>growing file, appends/s</th><th>preallocated, appends/s</th><th>ratio</th></tr></thead><tbody><tr><td>none</td><td>601,253</td><td>1,021,581</td><td>1.70</td></tr><tr><td>fsync</td><td>34,394</td><td>37,525</td><td>1.09</td></tr><tr><td>F_FULLFSYNC</td><td>247</td><td>248</td><td>1.00</td></tr></tbody></table><p>Whatever extending the file costs on APFS, it shows up when nothing else is expensive and disappears under a 4 ms flush.</p><h2 id="what-i-would-change">what I would change</h2><h3 id="make-the-pool-steal-and-the-log-physiological">make the pool STEAL and the log physiological</h3><p>NO-STEAL made recovery easy to reason about, and the crash tests are simple partly because of it. But the small-pool run shows that it converts memory pressure directly into checkpoints, 213 of them for 1M random puts, and it rules out transactions larger than memory. With concurrent transactions coming in v1 and v2, the right design is page LSNs and physiological log records, which let the pool evict dirty pages and let checkpoints be fuzzy.</p><h3 id="checksum-every-page">checksum every page</h3><p>Today a torn or bit-flipped page in the data file would be read silently. The WAL and the journal carry CRCs, and pages should too, with torn-page detection at read time.</p><h3 id="buffer-the-log-instead-of-a-pwrite-per-put">buffer the log instead of a pwrite per put</h3><p>The write path is where kvdb loses most to <code>std::map</code>, and one system call per put is the obvious cause. An in-memory log buffer flushed at commit boundaries would keep the same durability semantics for synced modes and remove most of the system calls for <code>kNone</code>.</p><h3 id="real-delete">real delete</h3><p>Lazy delete was the right cut for v0, since lookups and scans stay correct and deleting everything then reinserting works. But space is never reclaimed, so a delete-heavy workload wastes space without bound. Merge, redistribution and a free-page list are in v3.</p><h3 id="settle-the-sqlite-question-and-test-power-loss">settle the SQLite question and test power loss</h3><p>I would verify SQLite's sync behavior with <code>fs_usage</code> or a source build, so the durability comparison is like for like. And the crash tests should gain a mode that simulates power loss, for example by interposing on writes and dropping everything after the last completed sync, since the process-crash tests cannot say anything about that case.</p><h3 id="benchmark-larger-than-memory-on-a-quiet-machine">benchmark larger than memory, on a quiet machine</h3><p>Every throughput number here is an in-memory number. A database several times larger than RAM, with the page cache dropped, would make the buffer pool go to the SSD and make the random-read results reflect I/O. It should run on a machine with nothing else on it.</p><h2 id="reproducibility">reproducibility</h2><p>Build. This needs CMake 3.24 or newer, Ninja and a C++20 compiler, and GoogleTest is fetched at configure time. SQLite is optional, and the macOS SDK's copy is found automatically and enables the SQLite baseline. The <code>asan</code> preset uses Homebrew LLVM at <code>/opt/homebrew/opt/llvm/bin/clang++</code> because Apple clang 17's sanitizer runtime hangs on this macOS version.</p><figure class="highlight sh"><table><tr><td class="code"><pre><span class="line"><span class="built_in">cd</span> projects/01-kvdb-btree-wal-engine</span><br><span class="line">cmake --preset release            <span class="comment"># -O2, build/release</span></span><br><span class="line">cmake --build --preset release -j2</span><br><span class="line">cmake --preset asan               <span class="comment"># ASan + UBSan, build/asan</span></span><br><span class="line">cmake --build --preset asan -j2</span><br></pre></td></tr></table></figure><p>Tests.</p><figure class="highlight sh"><table><tr><td class="code"><pre><span class="line">./build/release/kvdb_tests        <span class="comment"># 36 tests, a few seconds</span></span><br><span class="line">./build/asan/kvdb_tests           <span class="comment"># the same 36 under ASan + UBSan</span></span><br><span class="line">ctest --preset release            <span class="comment"># or through CTest</span></span><br></pre></td></tr></table></figure><p>Benchmarks and plots. The script writes every raw CSV and log to <code>results/</code>, with the load average in each row, and takes about 5 minutes. The plotting script needs matplotlib.</p><figure class="highlight sh"><table><tr><td class="code"><pre><span class="line">scripts/run_bench.sh 1000000 3</span><br><span class="line">python3 -m venv .venv &amp;&amp; .venv/bin/pip install matplotlib</span><br><span class="line">.venv/bin/python scripts/plot_results.py      <span class="comment"># results/summary_main.csv and PNGs</span></span><br><span class="line">./build/release/kvdb_bench --exp=<span class="built_in">sync</span> --n=2000 --out=/tmp/sync.csv   <span class="comment"># one experiment</span></span><br></pre></td></tr></table></figure><p>Code is in <code>projects/01-kvdb-btree-wal-engine</code>, with the public API in <code>include/kvdb/db.h</code>, the engine in <code>src/</code>, the crash tests in <code>tests/recovery_test.cc</code>, the benchmark harness in <code>bench/bench.cc</code>, the formats, invariants and trade-offs in <code>DESIGN.md</code>, and the build log in <code>DEVLOG.md</code>.</p>]]>
    </content>
    <id>https://projects.farhansadeek.com/posts/crash-safe-b-plus-tree-store/</id>
    <link href="https://projects.farhansadeek.com/posts/crash-safe-b-plus-tree-store/"/>
    <published>2026-09-29T06:17:08.000Z</published>
    <summary>A B+ tree key-value store in C++ with a write-ahead log and journaled checkpoints. It survived kills aimed at every checkpoint step without losing an acknowledged write.</summary>
    <title>A Crash-Safe B+ Tree That Survives kill -9</title>
    <updated>2026-09-29T06:49:26.552Z</updated>
  </entry>
  <entry>
    <author>
      <name>Farhan Sadeek</name>
      <email>farhan@farhansadeek.com</email>
    </author>
    <category term="quant" scheme="https://projects.farhansadeek.com/tags/quant/"/>
    <category term="research" scheme="https://projects.farhansadeek.com/tags/research/"/>
    <category term="crypto" scheme="https://projects.farhansadeek.com/tags/crypto/"/>
    <content>
      <![CDATA[<p>I wrote down one hypothesis, a universe, a period split, a list of 18 trials, a cost model and a decision rule, and froze that file before downloading a single price. Then I did all model selection on 2020 to mid 2024 data, and opened a two year holdout (July 2024 to August 2026) exactly once.</p><p>The result splits in two. Among the 50 most liquid USDT pairs on Binance, yesterday's losers do rank above yesterday's winners the next day. The holdout mean rank information coefficient of one day reversal is 0.019 with a Newey-West t of 2.24, which clears the pre-registered bar of 2 by a small margin. That hypothesis is supported. The same signal traded as a dollar-neutral portfolio has a holdout gross Sharpe of -0.78 and a net Sharpe of -4.06. It lost money before paying a single basis point of cost. The second hypothesis, that the portfolio picked in development makes money after 15 bps per trade, is not supported. That portfolio earned a net Sharpe of 0.08 with a probabilistic Sharpe ratio of 0.55, against a required 0.95.</p><p>So the honest answer to &quot;can we predict markets&quot; from this study is yes, a little, at the level of daily cross-sectional ranks, and no, not in a way that paid after costs in the pre-registered test. There is also a post-hoc observation that the ridge and gradient boosting models did much better in the holdout than the model my selection rule picked. I explain below why I cannot claim that as a result.</p><p>Code, the pre-registration, the lock file and every result file are in <code>projects/09-crypto-reversal-preregistered</code>. Every number in this post comes from a file in its <code>results/</code> directory, named where it is used.</p><p><em>Reading note.</em> If you only want the verdicts, read <a href="#results">results</a>. The most useful sections for anyone building a backtest are <a href="#problems">problems</a> and the selection failure at the end of <a href="#experiments">experiments</a>.</p><h2 id="table-of-contents">table of contents</h2><ul><li><a href="#what-i-wanted-to-build">what I wanted to build</a></li><li><a href="#theory">theory</a></li><li><a href="#architecture">architecture</a></li><li><a href="#implementation">implementation</a></li><li><a href="#problems">problems</a></li><li><a href="#experiments">experiments</a></li><li><a href="#results">results</a></li><li><a href="#what-i-would-change">what I would change</a></li><li><a href="#reproducibility">reproducibility</a></li></ul><h2 id="what-i-wanted-to-build">what I wanted to build</h2><p>I wanted a research pipeline I could defend line by line, applied to one falsifiable question. Does yesterday's loser beat yesterday's winner tomorrow among liquid crypto pairs (short-term reversal), and if so, is any of that left after paying to trade it?</p><p>The question was chosen to be boring. Reversal is one of the oldest documented cross-sectional effects in equities, it needs nothing but prices, and it has an obvious enemy in trading costs because it turns the book over almost completely every day. A study where the answer could plausibly go either way, on free data, with a known failure mode, is a good test of the process rather than of my ability to find a clever signal.</p><p>The process was the point. Most backtests I have read fail in the same four ways. They peek at the future inside a feature or a universe definition. They quietly drop the coins or stocks that died. They re-use the test period while iterating, so it stops being a test. And they report the best of many variants without saying how many variants there were. I wanted each of those to be either impossible by construction or caught by a test.</p><h3 id="the-pre-registration">the pre-registration</h3><p><code>PREREGISTRATION.md</code> was written before any data was downloaded and fixes everything a result could be tuned against. Nothing in it may change after the holdout except an amendments section, which is still empty.</p><p>The two hypotheses and their decision rules, in full.</p><table><thead><tr><th>Hypothesis</th><th>Statistic</th><th>Supported if</th><th>Fails if</th></tr></thead><tbody><tr><td>H1, signal level. The cross-sectional rank of a coin's past 1 day return negatively predicts the rank of its next 1 day return</td><td>Holdout mean daily Spearman IC of <code>rev_1d = -ret_1d</code>, Newey-West t with 5 lags</td><td>mean IC above 0 and t above 2.0</td><td>mean IC at or below 0 (anything between is inconclusive)</td></tr><tr><td>H2, tradability. A dollar-neutral portfolio built from the dev-selected model earns a positive net Sharpe</td><td>Holdout annualized net Sharpe at 15 bps, and PSR that the true Sharpe exceeds 0</td><td>net Sharpe above 0 and PSR above 0.95</td><td>otherwise</td></tr></tbody></table><p>The file also fixes the data source, the point-in-time universe rule, the 9 features, the target and execution timing, all 9 models and both portfolio constructions (18 trials), the cost of 15 bps per unit of one-way turnover, the sensitivity grid of 0, 5, 10 and 25 bps, and the selection rule. The rule is to take the trial with the highest development net Sharpe at 15 bps, breaking ties by mean IC.</p><p>The periods are fixed too.</p><figure class="highlight plaintext"><table><tr><td class="code"><pre><span class="line">2018-01-01        2020-01-01                       2024-07-01              2026-08-31</span><br><span class="line">    |  warm-up         |  development, 9 half-year folds  |  holdout, run once       |</span><br><span class="line">    |  features and    |  walk-forward, all selection     |  refit once on data      |</span><br><span class="line">    |  training only   |  5 day embargo per fold          |  up to 2024-06-25        |</span><br></pre></td></tr></table></figure><p>The holdout script writes a lock file containing the SHA-256 of the pre-registration before it computes anything, and refuses to run if the lock exists. While writing this post I recomputed the hash of the current <code>PREREGISTRATION.md</code>, and it matches the one in <code>results/holdout/LOCK.json</code> (<code>0f38b396...79f6dc</code>), so the file has not been edited since the holdout ran.</p><p>An independent reviewer who did not build the project later checked the same thing from the other side. The hash of <code>PREREGISTRATION.md</code> still equals the one recorded in the lock when the holdout started, at 03:07 UTC on 2026-09-27. The file was last modified at 22:30:43 US Eastern the evening before, 35 seconds before the download log was created and about 37 minutes before the holdout ran, so it was fixed before any price was seen. The project was not yet in git at that point, so this rests on file timestamps plus the lock hash rather than on commit history. The same reviewer reran the development study and recomputed the holdout from the frozen selection, and both matched the committed files exactly (see <a href="#reproducibility">reproducibility</a>).</p><p>A negative answer was an acceptable outcome, and half of one is what I got.</p><h2 id="theory">theory</h2><h3 id="cross-sectional-prediction-and-the-information-coefficient">cross-sectional prediction and the information coefficient</h3><p>The study never predicts whether the market goes up. Each day it ranks the coins in the universe by a forecast and asks whether that ranking lines up with the ranking of next-day returns. The daily Spearman correlation between the two is the information coefficient.</p><p class="katex-block"><span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><msub><mrow><mi mathvariant="normal">I</mi><mi mathvariant="normal">C</mi></mrow><mi>t</mi></msub><mo>=</mo><mi mathvariant="normal">corr</mi><mo>⁡</mo><mo fence="false" stretchy="true" minsize="1.2em" maxsize="1.2em">(</mo><mi mathvariant="normal">rank</mi><mo>⁡</mo><mo stretchy="false">(</mo><msub><mi>f</mi><mrow><mi>i</mi><mo separator="true">,</mo><mi>t</mi></mrow></msub><mo stretchy="false">)</mo><mo separator="true">,</mo><mtext> </mtext><mi mathvariant="normal">rank</mi><mo>⁡</mo><mo stretchy="false">(</mo><msub><mi>r</mi><mrow><mi>i</mi><mo separator="true">,</mo><mi>t</mi><mo>+</mo><mn>1</mn></mrow></msub><mo stretchy="false">)</mo><mo fence="false" stretchy="true" minsize="1.2em" maxsize="1.2em">)</mo></mrow><annotation encoding="application/x-tex">\mathrm{IC}_t = \operatorname{corr}\big(\operatorname{rank}(f_{i,t}),\ \operatorname{rank}(r_{i,t+1})\big)</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.8333em;vertical-align:-0.15em;"></span><span class="mord"><span class="mord"><span class="mord mathrm">IC</span></span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.2806em;"><span style="top:-2.55em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight">t</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:1.2em;vertical-align:-0.35em;"></span><span class="mop"><span class="mord mathrm">corr</span></span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mord"><span class="delimsizing size1">(</span></span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mop"><span class="mord mathrm">rank</span></span><span class="mopen">(</span><span class="mord"><span class="mord mathnormal" style="margin-right:0.1076em;">f</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3117em;"><span style="top:-2.55em;margin-left:-0.1076em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight">i</span><span class="mpunct mtight">,</span><span class="mord mathnormal mtight">t</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.2861em;"><span></span></span></span></span></span></span><span class="mclose">)</span><span class="mpunct">,</span><span class="mspace"> </span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mop"><span class="mord mathrm">rank</span></span><span class="mopen">(</span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0278em;">r</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3117em;"><span style="top:-2.55em;margin-left:-0.0278em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight">i</span><span class="mpunct mtight">,</span><span class="mord mathnormal mtight">t</span><span class="mbin mtight">+</span><span class="mord mtight">1</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.2861em;"><span></span></span></span></span></span></span><span class="mclose">)</span><span class="mord"><span class="delimsizing size1">)</span></span></span></span></span></span></p><p>where <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>f</mi><mrow><mi>i</mi><mo separator="true">,</mo><mi>t</mi></mrow></msub></mrow><annotation encoding="application/x-tex">f_{i,t}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.9805em;vertical-align:-0.2861em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.1076em;">f</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3117em;"><span style="top:-2.55em;margin-left:-0.1076em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight">i</span><span class="mpunct mtight">,</span><span class="mord mathnormal mtight">t</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.2861em;"><span></span></span></span></span></span></span></span></span></span> is the forecast for coin <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>i</mi></mrow><annotation encoding="application/x-tex">i</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6595em;"></span><span class="mord mathnormal">i</span></span></span></span> at the close of day <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>t</mi></mrow><annotation encoding="application/x-tex">t</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6151em;"></span><span class="mord mathnormal">t</span></span></span></span> and <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>r</mi><mrow><mi>i</mi><mo separator="true">,</mo><mi>t</mi><mo>+</mo><mn>1</mn></mrow></msub></mrow><annotation encoding="application/x-tex">r_{i,t+1}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.7167em;vertical-align:-0.2861em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0278em;">r</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3117em;"><span style="top:-2.55em;margin-left:-0.0278em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight">i</span><span class="mpunct mtight">,</span><span class="mord mathnormal mtight">t</span><span class="mbin mtight">+</span><span class="mord mtight">1</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.2861em;"><span></span></span></span></span></span></span></span></span></span> is its return over the next day. A mean IC of 0.02 sounds like nothing. Grinold's fundamental law of active management says the information ratio of a strategy scales like <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mrow><mi mathvariant="normal">I</mi><mi mathvariant="normal">C</mi></mrow><msqrt><mi>B</mi></msqrt></mrow><annotation encoding="application/x-tex">\mathrm{IC}\sqrt{B}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1.04em;vertical-align:-0.1133em;"></span><span class="mord"><span class="mord mathrm">IC</span></span><span class="mord sqrt"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.9267em;"><span class="svg-align" style="top:-3em;"><span class="pstrut" style="height:3em;"></span><span class="mord" style="padding-left:0.833em;"><span class="mord mathnormal" style="margin-right:0.0502em;">B</span></span></span><span style="top:-2.8867em;"><span class="pstrut" style="height:3em;"></span><span class="hide-tail" style="min-width:0.853em;height:1.08em;"><svg xmlns="http://www.w3.org/2000/svg" width="400em" height="1.08em" viewBox="0 0 400000 1080" preserveAspectRatio="xMinYMin slice"><path d="M95,702c-2.7,0,-7.17,-2.7,-13.5,-8c-5.8,-5.3,-9.5,-10,-9.5,-14c0,-2,0.3,-3.3,1,-4c1.3,-2.7,23.83,-20.7,67.5,-54c44.2,-33.3,65.8,-50.3,66.5,-51c1.3,-1.3,3,-2,5,-2c4.7,0,8.7,3.3,12,10s173,378,173,378c0.7,0,35.3,-71,104,-213c68.7,-142,137.5,-285,206.5,-429c69,-144,104.5,-217.7,106.5,-221l0 -0c5.3,-9.3,12,-14,20,-14H400000v40H845.2724s-225.272,467,-225.272,467s-235,486,-235,486c-2.7,4.7,-9,7,-19,7c-6,0,-10,-1,-12,-3s-194,-422,-194,-422s-65,47,-65,47zM834 80h400000v40h-400000z"/></svg></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.1133em;"><span></span></span></span></span></span></span></span></span>, where <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>B</mi></mrow><annotation encoding="application/x-tex">B</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal" style="margin-right:0.0502em;">B</span></span></span></span> is the number of independent bets per year, and 50 names a day for 365 days a year is a lot of bets. The catch is that the law ignores costs, and it treats every bet as equally sized. Both assumptions turn out to matter here.</p><h3 id="why-reversal-might-exist">why reversal might exist</h3><p>In equities, short-term reversal is usually explained as payment for liquidity. Someone who has to trade pushes the price, and whoever absorbs that flow is paid by the rebound. Part of it is also bid-ask bounce, where a close that prints at the bid looks like a loss and the next close at the ask looks like a gain. Bounce should be small here, because Binance daily closes on liquid pairs are last trades on books with spreads of a few basis points.</p><h3 id="what-costs-do-to-a-daily-book">what costs do to a daily book</h3><p>A daily-rebalanced long-short book with average turnover <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>T</mi></mrow><annotation encoding="application/x-tex">T</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal" style="margin-right:0.1389em;">T</span></span></span></span> per day (the sum of absolute weight changes) and a one-way cost <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>c</mi></mrow><annotation encoding="application/x-tex">c</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal">c</span></span></span></span> loses <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>T</mi><mo>⋅</mo><mi>c</mi><mo>⋅</mo><mn>365</mn></mrow><annotation encoding="application/x-tex">T \cdot c \cdot 365</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal" style="margin-right:0.1389em;">T</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">⋅</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:0.4445em;"></span><span class="mord mathnormal">c</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">⋅</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:0.6444em;"></span><span class="mord">365</span></span></span></span> per year. At 15 bps and a turnover of 1.3, which is typical for a 1 day reversal signal, that is about 71% per year. The holdout confirms the arithmetic to the decimal. The reversal portfolio's annualized gross return was -17.2% and its net return -89.0%, a gap of 71.8 points, and <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mn>1.311</mn><mo>×</mo><mn>0.0015</mn><mo>×</mo><mn>365</mn><mo>=</mo><mn>0.718</mn></mrow><annotation encoding="application/x-tex">1.311 \times 0.0015 \times 365 = 0.718</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.7278em;vertical-align:-0.0833em;"></span><span class="mord">1.311</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">×</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:0.7278em;vertical-align:-0.0833em;"></span><span class="mord">0.0015</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">×</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:0.6444em;"></span><span class="mord">365</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.6444em;"></span><span class="mord">0.718</span></span></span></span> (<code>results/holdout/holdout_report.json</code>). Any gross edge has to beat that number before it is worth anything.</p><h3 id="statistics-that-account-for-how-the-numbers-were-made">statistics that account for how the numbers were made</h3><p>Daily ICs are autocorrelated when a signal changes slowly. A 20 day momentum ranking barely moves day to day, so consecutive ICs are not independent draws, and the naive t-statistic overstates confidence. I use a Newey-West t-statistic, which replaces the variance of the mean with a Bartlett-weighted sum of autocovariances up to lag <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>L</mi><mo>=</mo><mn>5</mn></mrow><annotation encoding="application/x-tex">L = 5</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal">L</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.6444em;"></span><span class="mord">5</span></span></span></span>.</p><p class="katex-block"><span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><msubsup><mover accent="true"><mi>σ</mi><mo>^</mo></mover><mrow><mi>N</mi><mi>W</mi></mrow><mn>2</mn></msubsup><mo>=</mo><msub><mover accent="true"><mi>γ</mi><mo>^</mo></mover><mn>0</mn></msub><mo>+</mo><mn>2</mn><munderover><mo>∑</mo><mrow><mi>l</mi><mo>=</mo><mn>1</mn></mrow><mi>L</mi></munderover><mo fence="false" stretchy="true" minsize="1.8em" maxsize="1.8em">(</mo><mn>1</mn><mo>−</mo><mfrac><mi>l</mi><mrow><mi>L</mi><mo>+</mo><mn>1</mn></mrow></mfrac><mo fence="false" stretchy="true" minsize="1.8em" maxsize="1.8em">)</mo><msub><mover accent="true"><mi>γ</mi><mo>^</mo></mover><mi>l</mi></msub><mo separator="true">,</mo><mspace width="2em"/><mi>t</mi><mo>=</mo><mfrac><mover accent="true"><mi>x</mi><mo>ˉ</mo></mover><msqrt><mrow><msubsup><mover accent="true"><mi>σ</mi><mo>^</mo></mover><mrow><mi>N</mi><mi>W</mi></mrow><mn>2</mn></msubsup><mi mathvariant="normal">/</mi><mi>n</mi></mrow></msqrt></mfrac></mrow><annotation encoding="application/x-tex">\hat\sigma^2_{NW} = \hat\gamma_0 + 2\sum_{l=1}^{L}\Big(1 - \frac{l}{L+1}\Big)\hat\gamma_l, \qquad t = \frac{\bar x}{\sqrt{\hat\sigma^2_{NW}/n}}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1.1111em;vertical-align:-0.247em;"></span><span class="mord"><span class="mord accent"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.6944em;"><span style="top:-3em;"><span class="pstrut" style="height:3em;"></span><span class="mord mathnormal" style="margin-right:0.0359em;">σ</span></span><span style="top:-3em;"><span class="pstrut" style="height:3em;"></span><span class="accent-body" style="left:-0.25em;"><span class="mord">^</span></span></span></span></span></span></span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.8641em;"><span style="top:-2.453em;margin-left:-0.0359em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight" style="margin-right:0.109em;">N</span><span class="mord mathnormal mtight" style="margin-right:0.1389em;">W</span></span></span></span><span style="top:-3.113em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight">2</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.247em;"><span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.8889em;vertical-align:-0.1944em;"></span><span class="mord"><span class="mord accent"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.6944em;"><span style="top:-3em;"><span class="pstrut" style="height:3em;"></span><span class="mord mathnormal" style="margin-right:0.0556em;">γ</span></span><span style="top:-3em;"><span class="pstrut" style="height:3em;"></span><span class="accent-body" style="left:-0.25em;"><span class="mord">^</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.1944em;"><span></span></span></span></span></span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3011em;"><span style="top:-2.55em;margin-left:-0.0556em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight">0</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">+</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:3.1304em;vertical-align:-1.3021em;"></span><span class="mord">2</span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mop op-limits"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:1.8283em;"><span style="top:-1.8479em;margin-left:0em;"><span class="pstrut" style="height:3.05em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight" style="margin-right:0.0197em;">l</span><span class="mrel mtight">=</span><span class="mord mtight">1</span></span></span></span><span style="top:-3.05em;"><span class="pstrut" style="height:3.05em;"></span><span><span class="mop op-symbol large-op">∑</span></span></span><span style="top:-4.3em;margin-left:0em;"><span class="pstrut" style="height:3.05em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight">L</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:1.3021em;"><span></span></span></span></span></span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mord"><span class="delimsizing size2">(</span></span><span class="mord">1</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">−</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:2.1408em;vertical-align:-0.7693em;"></span><span class="mord"><span class="mopen nulldelimiter"></span><span class="mfrac"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:1.3714em;"><span style="top:-2.314em;"><span class="pstrut" style="height:3em;"></span><span class="mord"><span class="mord mathnormal">L</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">+</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mord">1</span></span></span><span style="top:-3.23em;"><span class="pstrut" style="height:3em;"></span><span class="frac-line" style="border-bottom-width:0.04em;"></span></span><span style="top:-3.677em;"><span class="pstrut" style="height:3em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0197em;">l</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.7693em;"><span></span></span></span></span></span><span class="mclose nulldelimiter"></span></span><span class="mord"><span class="delimsizing size2">)</span></span><span class="mord"><span class="mord accent"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.6944em;"><span style="top:-3em;"><span class="pstrut" style="height:3em;"></span><span class="mord mathnormal" style="margin-right:0.0556em;">γ</span></span><span style="top:-3em;"><span class="pstrut" style="height:3em;"></span><span class="accent-body" style="left:-0.25em;"><span class="mord">^</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.1944em;"><span></span></span></span></span></span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3361em;"><span style="top:-2.55em;margin-left:-0.0556em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight" style="margin-right:0.0197em;">l</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span><span class="mpunct">,</span><span class="mspace" style="margin-right:2em;"></span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mord mathnormal">t</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:2.3748em;vertical-align:-1.13em;"></span><span class="mord"><span class="mopen nulldelimiter"></span><span class="mfrac"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:1.2448em;"><span style="top:-2.1738em;"><span class="pstrut" style="height:3em;"></span><span class="mord"><span class="mord sqrt"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.9362em;"><span class="svg-align" style="top:-3.2em;"><span class="pstrut" style="height:3.2em;"></span><span class="mord" style="padding-left:1em;"><span class="mord"><span class="mord accent"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.6944em;"><span style="top:-3em;"><span class="pstrut" style="height:3em;"></span><span class="mord mathnormal" style="margin-right:0.0359em;">σ</span></span><span style="top:-3em;"><span class="pstrut" style="height:3em;"></span><span class="accent-body" style="left:-0.25em;"><span class="mord">^</span></span></span></span></span></span></span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.7959em;"><span style="top:-2.4065em;margin-left:-0.0359em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight" style="margin-right:0.109em;">N</span><span class="mord mathnormal mtight" style="margin-right:0.1389em;">W</span></span></span></span><span style="top:-3.0448em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight">2</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.2935em;"><span></span></span></span></span></span></span><span class="mord">/</span><span class="mord mathnormal">n</span></span></span><span style="top:-2.8962em;"><span class="pstrut" style="height:3.2em;"></span><span class="hide-tail" style="min-width:1.02em;height:1.28em;"><svg xmlns="http://www.w3.org/2000/svg" width="400em" height="1.28em" viewBox="0 0 400000 1296" preserveAspectRatio="xMinYMin slice"><path d="M263,681c0.7,0,18,39.7,52,119c34,79.3,68.167,158.7,102.5,238c34.3,79.3,51.8,119.3,52.5,120c340,-704.7,510.7,-1060.3,512,-1067l0 -0c4.7,-7.3,11,-11,19,-11H40000v40H1012.3s-271.3,567,-271.3,567c-38.7,80.7,-84,175,-136,283c-52,108,-89.167,185.3,-111.5,232c-22.3,46.7,-33.8,70.3,-34.5,71c-4.7,4.7,-12.3,7,-23,7s-12,-1,-12,-1s-109,-253,-109,-253c-72.7,-168,-109.3,-252,-110,-252c-10.7,8,-22,16.7,-34,26c-22,17.3,-33.3,26,-34,26s-26,-26,-26,-26s76,-59,76,-59s76,-60,76,-60zM1001 80h400000v40h-400000z"/></svg></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.3038em;"><span></span></span></span></span></span></span></span><span style="top:-3.23em;"><span class="pstrut" style="height:3em;"></span><span class="frac-line" style="border-bottom-width:0.04em;"></span></span><span style="top:-3.677em;"><span class="pstrut" style="height:3em;"></span><span class="mord"><span class="mord accent"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.5678em;"><span style="top:-3em;"><span class="pstrut" style="height:3em;"></span><span class="mord mathnormal">x</span></span><span style="top:-3em;"><span class="pstrut" style="height:3em;"></span><span class="accent-body" style="left:-0.2222em;"><span class="mord">ˉ</span></span></span></span></span></span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:1.13em;"><span></span></span></span></span></span><span class="mclose nulldelimiter"></span></span></span></span></span></span></p><p>For portfolio returns I use the probabilistic Sharpe ratio (Bailey and López de Prado, 2012). It is the probability that the true Sharpe exceeds a threshold <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>S</mi><msup><mi>R</mi><mo>∗</mo></msup></mrow><annotation encoding="application/x-tex">SR^*</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6887em;"></span><span class="mord mathnormal" style="margin-right:0.0576em;">S</span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0077em;">R</span><span class="msupsub"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.6887em;"><span style="top:-3.063em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mbin mtight">∗</span></span></span></span></span></span></span></span></span></span></span>, given the observed per-period Sharpe <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mover accent="true"><mrow><mi>S</mi><mi>R</mi></mrow><mo stretchy="true">^</mo></mover></mrow><annotation encoding="application/x-tex">\widehat{SR}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.9833em;"></span><span class="mord accent"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.9833em;"><span style="top:-3em;"><span class="pstrut" style="height:3em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0576em;">S</span><span class="mord mathnormal" style="margin-right:0.0077em;">R</span></span></span><span class="svg-align" style="top:-3.6833em;"><span class="pstrut" style="height:3em;"></span><span style="height:0.3em;"><svg xmlns="http://www.w3.org/2000/svg" width="100%" height="0.3em" viewBox="0 0 2364 300" preserveAspectRatio="none"><path d="M1181 0h2l1171 176c6 0 10 5 10 11l-2 23c-1 6-5 10-11 10h-1L1182 67 15 220h-1c-6 0-10-4-11-10l-2-23c-1-6 4-11 10-11z"/></svg></span></span></span></span></span></span></span></span></span>, the sample length <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>n</mi></mrow><annotation encoding="application/x-tex">n</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal">n</span></span></span></span>, and the skew <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>γ</mi><mn>3</mn></msub></mrow><annotation encoding="application/x-tex">\gamma_3</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.625em;vertical-align:-0.1944em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0556em;">γ</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3011em;"><span style="top:-2.55em;margin-left:-0.0556em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight">3</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span></span></span></span> and kurtosis <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>γ</mi><mn>4</mn></msub></mrow><annotation encoding="application/x-tex">\gamma_4</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.625em;vertical-align:-0.1944em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0556em;">γ</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3011em;"><span style="top:-2.55em;margin-left:-0.0556em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight">4</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span></span></span></span> of the returns.</p><p class="katex-block"><span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mrow><mi mathvariant="normal">P</mi><mi mathvariant="normal">S</mi><mi mathvariant="normal">R</mi></mrow><mo stretchy="false">(</mo><mi>S</mi><msup><mi>R</mi><mo>∗</mo></msup><mo stretchy="false">)</mo><mo>=</mo><mi mathvariant="normal">Φ</mi><mrow><mo fence="true">(</mo><mfrac><mrow><mo stretchy="false">(</mo><mover accent="true"><mrow><mi>S</mi><mi>R</mi></mrow><mo stretchy="true">^</mo></mover><mo>−</mo><mi>S</mi><msup><mi>R</mi><mo>∗</mo></msup><mo stretchy="false">)</mo><msqrt><mrow><mi>n</mi><mo>−</mo><mn>1</mn></mrow></msqrt></mrow><msqrt><mrow><mn>1</mn><mo>−</mo><msub><mi>γ</mi><mn>3</mn></msub><mover accent="true"><mrow><mi>S</mi><mi>R</mi></mrow><mo stretchy="true">^</mo></mover><mo>+</mo><mfrac><mrow><msub><mi>γ</mi><mn>4</mn></msub><mo>−</mo><mn>1</mn></mrow><mn>4</mn></mfrac><msup><mover accent="true"><mrow><mi>S</mi><mi>R</mi></mrow><mo stretchy="true">^</mo></mover><mn>2</mn></msup></mrow></msqrt></mfrac><mo fence="true">)</mo></mrow></mrow><annotation encoding="application/x-tex">\mathrm{PSR}(SR^*) = \Phi\left(\frac{(\widehat{SR} - SR^*)\sqrt{n-1}}{\sqrt{1 - \gamma_3\widehat{SR} + \frac{\gamma_4 - 1}{4}\widehat{SR}^2}}\right)</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mord"><span class="mord mathrm">PSR</span></span><span class="mopen">(</span><span class="mord mathnormal" style="margin-right:0.0576em;">S</span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0077em;">R</span><span class="msupsub"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.7387em;"><span style="top:-3.113em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mbin mtight">∗</span></span></span></span></span></span></span></span><span class="mclose">)</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:3.78em;vertical-align:-1.73em;"></span><span class="mord">Φ</span><span class="mspace" style="margin-right:0.1667em;"></span><span class="minner"><span class="mopen"><span class="delimsizing mult"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:2.05em;"><span style="top:-4.05em;"><span class="pstrut" style="height:5.6em;"></span><span style="width:0.875em;height:3.6em;"><svg xmlns="http://www.w3.org/2000/svg" width="0.875em" height="3.6em" viewBox="0 0 875 3600"><path d="M863,9c0,-2,-2,-5,-6,-9c0,0,-17,0,-17,0c-12.7,0,-19.3,0.3,-20,1c-5.3,5.3,-10.3,11,-15,17c-242.7,294.7,-395.3,682,-458,1162c-21.3,163.3,-33.3,349,-36,557 l0,84c0.2,6,0,26,0,60c2,159.3,10,310.7,24,454c53.3,528,210,949.7,470,1265c4.7,6,9.7,11.7,15,17c0.7,0.7,7,1,19,1c0,0,18,0,18,0c4,-4,6,-7,6,-9c0,-2.7,-3.3,-8.7,-10,-18c-135.3,-192.7,-235.5,-414.3,-300.5,-665c-65,-250.7,-102.5,-544.7,-112.5,-882c-2,-104,-3,-167,-3,-189l0,-92c0,-162.7,5.7,-314,17,-454c20.7,-272,63.7,-513,129,-723c65.3,-210,155.3,-396.3,270,-559c6.7,-9.3,10,-15.3,10,-18z"/></svg></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:1.55em;"><span></span></span></span></span></span></span><span class="mord"><span class="mopen nulldelimiter"></span><span class="mfrac"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:1.6603em;"><span style="top:-2.11em;"><span class="pstrut" style="height:3.4062em;"></span><span class="mord"><span class="mord sqrt"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:1.4062em;"><span class="svg-align" style="top:-3.8em;"><span class="pstrut" style="height:3.8em;"></span><span class="mord" style="padding-left:1em;"><span class="mord">1</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">−</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0556em;">γ</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3011em;"><span style="top:-2.55em;margin-left:-0.0556em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight">3</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span><span class="mord accent"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.9833em;"><span style="top:-3em;"><span class="pstrut" style="height:3em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0576em;">S</span><span class="mord mathnormal" style="margin-right:0.0077em;">R</span></span></span><span class="svg-align" style="top:-3.6833em;"><span class="pstrut" style="height:3em;"></span><span style="height:0.3em;"><svg xmlns="http://www.w3.org/2000/svg" width="100%" height="0.3em" viewBox="0 0 2364 300" preserveAspectRatio="none"><path d="M1181 0h2l1171 176c6 0 10 5 10 11l-2 23c-1 6-5 10-11 10h-1L1182 67 15 220h-1c-6 0-10-4-11-10l-2-23c-1-6 4-11 10-11z"/></svg></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">+</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mord"><span class="mopen nulldelimiter"></span><span class="mfrac"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.8972em;"><span style="top:-2.655em;"><span class="pstrut" style="height:3em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mtight">4</span></span></span></span><span style="top:-3.23em;"><span class="pstrut" style="height:3em;"></span><span class="frac-line" style="border-bottom-width:0.04em;"></span></span><span style="top:-3.4461em;"><span class="pstrut" style="height:3em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mtight"><span class="mord mathnormal mtight" style="margin-right:0.0556em;">γ</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3173em;"><span style="top:-2.357em;margin-left:-0.0556em;margin-right:0.0714em;"><span class="pstrut" style="height:2.5em;"></span><span class="sizing reset-size3 size1 mtight"><span class="mord mtight">4</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.143em;"><span></span></span></span></span></span></span><span class="mbin mtight">−</span><span class="mord mtight">1</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.345em;"><span></span></span></span></span></span><span class="mclose nulldelimiter"></span></span><span class="mord"><span class="mord accent"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.9833em;"><span style="top:-3em;"><span class="pstrut" style="height:3em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0576em;">S</span><span class="mord mathnormal" style="margin-right:0.0077em;">R</span></span></span><span class="svg-align" style="top:-3.6833em;"><span class="pstrut" style="height:3em;"></span><span style="height:0.3em;"><svg xmlns="http://www.w3.org/2000/svg" width="100%" height="0.3em" viewBox="0 0 2364 300" preserveAspectRatio="none"><path d="M1181 0h2l1171 176c6 0 10 5 10 11l-2 23c-1 6-5 10-11 10h-1L1182 67 15 220h-1c-6 0-10-4-11-10l-2-23c-1-6 4-11 10-11z"/></svg></span></span></span></span></span></span><span class="msupsub"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:1.1873em;"><span style="top:-3.4362em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight">2</span></span></span></span></span></span></span></span></span></span><span style="top:-3.3662em;"><span class="pstrut" style="height:3.8em;"></span><span class="hide-tail" style="min-width:1.02em;height:1.88em;"><svg xmlns="http://www.w3.org/2000/svg" width="400em" height="1.88em" viewBox="0 0 400000 1944" preserveAspectRatio="xMinYMin slice"><path d="M983 90l0 -0c4,-6.7,10,-10,18,-10 H400000v40H1013.1s-83.4,268,-264.1,840c-180.7,572,-277,876.3,-289,913c-4.7,4.7,-12.7,7,-24,7s-12,0,-12,0c-1.3,-3.3,-3.7,-11.7,-7,-25c-35.3,-125.3,-106.7,-373.3,-214,-744c-10,12,-21,25,-33,39s-32,39,-32,39c-6,-5.3,-15,-14,-27,-26s25,-30,25,-30c26.7,-32.7,52,-63,76,-91s52,-60,52,-60s208,722,208,722c56,-175.3,126.3,-397.3,211,-666c84.7,-268.7,153.8,-488.2,207.5,-658.5c53.7,-170.3,84.5,-266.8,92.5,-289.5zM1001 80h400000v40h-400000z"/></svg></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.4338em;"><span></span></span></span></span></span></span></span><span style="top:-3.6362em;"><span class="pstrut" style="height:3.4062em;"></span><span class="frac-line" style="border-bottom-width:0.04em;"></span></span><span style="top:-4.0832em;"><span class="pstrut" style="height:3.4062em;"></span><span class="mord"><span class="mopen">(</span><span class="mord accent"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.9833em;"><span style="top:-3em;"><span class="pstrut" style="height:3em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0576em;">S</span><span class="mord mathnormal" style="margin-right:0.0077em;">R</span></span></span><span class="svg-align" style="top:-3.6833em;"><span class="pstrut" style="height:3em;"></span><span style="height:0.3em;"><svg xmlns="http://www.w3.org/2000/svg" width="100%" height="0.3em" viewBox="0 0 2364 300" preserveAspectRatio="none"><path d="M1181 0h2l1171 176c6 0 10 5 10 11l-2 23c-1 6-5 10-11 10h-1L1182 67 15 220h-1c-6 0-10-4-11-10l-2-23c-1-6 4-11 10-11z"/></svg></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">−</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mord mathnormal" style="margin-right:0.0576em;">S</span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0077em;">R</span><span class="msupsub"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.6887em;"><span style="top:-3.063em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mbin mtight">∗</span></span></span></span></span></span></span></span><span class="mclose">)</span><span class="mord sqrt"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.8656em;"><span class="svg-align" style="top:-3em;"><span class="pstrut" style="height:3em;"></span><span class="mord" style="padding-left:0.833em;"><span class="mord mathnormal">n</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">−</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mord">1</span></span></span><span style="top:-2.8256em;"><span class="pstrut" style="height:3em;"></span><span class="hide-tail" style="min-width:0.853em;height:1.08em;"><svg xmlns="http://www.w3.org/2000/svg" width="400em" height="1.08em" viewBox="0 0 400000 1080" preserveAspectRatio="xMinYMin slice"><path d="M95,702c-2.7,0,-7.17,-2.7,-13.5,-8c-5.8,-5.3,-9.5,-10,-9.5,-14c0,-2,0.3,-3.3,1,-4c1.3,-2.7,23.83,-20.7,67.5,-54c44.2,-33.3,65.8,-50.3,66.5,-51c1.3,-1.3,3,-2,5,-2c4.7,0,8.7,3.3,12,10s173,378,173,378c0.7,0,35.3,-71,104,-213c68.7,-142,137.5,-285,206.5,-429c69,-144,104.5,-217.7,106.5,-221l0 -0c5.3,-9.3,12,-14,20,-14H400000v40H845.2724s-225.272,467,-225.272,467s-235,486,-235,486c-2.7,4.7,-9,7,-19,7c-6,0,-10,-1,-12,-3s-194,-422,-194,-422s-65,47,-65,47zM834 80h400000v40h-400000z"/></svg></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.1744em;"><span></span></span></span></span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:1.73em;"><span></span></span></span></span></span><span class="mclose nulldelimiter"></span></span><span class="mclose"><span class="delimsizing mult"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:2.05em;"><span style="top:-4.05em;"><span class="pstrut" style="height:5.6em;"></span><span style="width:0.875em;height:3.6em;"><svg xmlns="http://www.w3.org/2000/svg" width="0.875em" height="3.6em" viewBox="0 0 875 3600"><path d="M76,0c-16.7,0,-25,3,-25,9c0,2,2,6.3,6,13c21.3,28.7,42.3,60.3,63,95c96.7,156.7,172.8,332.5,228.5,527.5c55.7,195,92.8,416.5,111.5,664.5c11.3,139.3,17,290.7,17,454c0,28,1.7,43,3.3,45l0,9c-3,4,-3.3,16.7,-3.3,38c0,162,-5.7,313.7,-17,455c-18.7,248,-55.8,469.3,-111.5,664c-55.7,194.7,-131.8,370.3,-228.5,527c-20.7,34.7,-41.7,66.3,-63,95c-2,3.3,-4,7,-6,11c0,7.3,5.7,11,17,11c0,0,11,0,11,0c9.3,0,14.3,-0.3,15,-1c5.3,-5.3,10.3,-11,15,-17c242.7,-294.7,395.3,-681.7,458,-1161c21.3,-164.7,33.3,-350.7,36,-558l0,-144c-2,-159.3,-10,-310.7,-24,-454c-53.3,-528,-210,-949.7,-470,-1265c-4.7,-6,-9.7,-11.7,-15,-17c-0.7,-0.7,-6.7,-1,-18,-1z"/></svg></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:1.55em;"><span></span></span></span></span></span></span></span></span></span></span></span></p><p>Fat tails widen the denominator, so a crypto strategy needs a longer or cleaner record than a normal-returns calculation would suggest.</p><p>For the selection step I use the deflated Sharpe ratio (2014), which is the PSR with the threshold raised to the Sharpe you would expect from the luckiest of <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>N</mi></mrow><annotation encoding="application/x-tex">N</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal" style="margin-right:0.109em;">N</span></span></span></span> worthless strategies. With <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>γ</mi></mrow><annotation encoding="application/x-tex">\gamma</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.625em;vertical-align:-0.1944em;"></span><span class="mord mathnormal" style="margin-right:0.0556em;">γ</span></span></span></span> the Euler-Mascheroni constant and <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>z</mi><mo stretchy="false">(</mo><mo>⋅</mo><mo stretchy="false">)</mo></mrow><annotation encoding="application/x-tex">z(\cdot)</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mord mathnormal" style="margin-right:0.044em;">z</span><span class="mopen">(</span><span class="mord">⋅</span><span class="mclose">)</span></span></span></span> the standard normal quantile,</p><p class="katex-block"><span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mi>S</mi><msup><mi>R</mi><mo>∗</mo></msup><mo>=</mo><msqrt><mrow><mi mathvariant="normal">Var</mi><mo>⁡</mo><mo stretchy="false">[</mo><msub><mover accent="true"><mrow><mi>S</mi><mi>R</mi></mrow><mo stretchy="true">^</mo></mover><mtext>trials</mtext></msub><mo stretchy="false">]</mo></mrow></msqrt><mtext> </mtext><mo fence="false" stretchy="true" minsize="1.8em" maxsize="1.8em">(</mo><mo stretchy="false">(</mo><mn>1</mn><mo>−</mo><mi>γ</mi><mo stretchy="false">)</mo><mtext> </mtext><mi>z</mi><mo fence="false" stretchy="true" minsize="1.2em" maxsize="1.2em">(</mo><mn>1</mn><mo>−</mo><mstyle scriptlevel="0" displaystyle="false"><mfrac><mn>1</mn><mi>N</mi></mfrac></mstyle><mo fence="false" stretchy="true" minsize="1.2em" maxsize="1.2em">)</mo><mo>+</mo><mi>γ</mi><mtext> </mtext><mi>z</mi><mo fence="false" stretchy="true" minsize="1.2em" maxsize="1.2em">(</mo><mn>1</mn><mo>−</mo><mstyle scriptlevel="0" displaystyle="false"><mfrac><mn>1</mn><mrow><mi>N</mi><mi>e</mi></mrow></mfrac></mstyle><mo fence="false" stretchy="true" minsize="1.2em" maxsize="1.2em">)</mo><mo fence="false" stretchy="true" minsize="1.8em" maxsize="1.8em">)</mo></mrow><annotation encoding="application/x-tex">SR^* = \sqrt{\operatorname{Var}[\widehat{SR}_{\text{trials}}]}\,\Big((1-\gamma)\,z\big(1 - \tfrac{1}{N}\big) + \gamma\, z\big(1 - \tfrac{1}{Ne}\big)\Big)</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.7387em;"></span><span class="mord mathnormal" style="margin-right:0.0576em;">S</span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0077em;">R</span><span class="msupsub"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.7387em;"><span style="top:-3.113em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mbin mtight">∗</span></span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:2.0506em;vertical-align:-0.65em;"></span><span class="mord sqrt"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:1.4005em;"><span class="svg-align" style="top:-3.8em;"><span class="pstrut" style="height:3.8em;"></span><span class="mord" style="padding-left:1em;"><span class="mop"><span class="mord mathrm">Var</span></span><span class="mopen">[</span><span class="mord"><span class="mord accent"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.9833em;"><span style="top:-3em;"><span class="pstrut" style="height:3em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0576em;">S</span><span class="mord mathnormal" style="margin-right:0.0077em;">R</span></span></span><span class="svg-align" style="top:-3.6833em;"><span class="pstrut" style="height:3em;"></span><span style="height:0.3em;"><svg xmlns="http://www.w3.org/2000/svg" width="100%" height="0.3em" viewBox="0 0 2364 300" preserveAspectRatio="none"><path d="M1181 0h2l1171 176c6 0 10 5 10 11l-2 23c-1 6-5 10-11 10h-1L1182 67 15 220h-1c-6 0-10-4-11-10l-2-23c-1-6 4-11 10-11z"/></svg></span></span></span></span></span></span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3361em;"><span style="top:-2.55em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord text mtight"><span class="mord mtight">trials</span></span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span><span class="mclose">]</span></span></span><span style="top:-3.3605em;"><span class="pstrut" style="height:3.8em;"></span><span class="hide-tail" style="min-width:1.02em;height:1.88em;"><svg xmlns="http://www.w3.org/2000/svg" width="400em" height="1.88em" viewBox="0 0 400000 1944" preserveAspectRatio="xMinYMin slice"><path d="M983 90l0 -0c4,-6.7,10,-10,18,-10 H400000v40H1013.1s-83.4,268,-264.1,840c-180.7,572,-277,876.3,-289,913c-4.7,4.7,-12.7,7,-24,7s-12,0,-12,0c-1.3,-3.3,-3.7,-11.7,-7,-25c-35.3,-125.3,-106.7,-373.3,-214,-744c-10,12,-21,25,-33,39s-32,39,-32,39c-6,-5.3,-15,-14,-27,-26s25,-30,25,-30c26.7,-32.7,52,-63,76,-91s52,-60,52,-60s208,722,208,722c56,-175.3,126.3,-397.3,211,-666c84.7,-268.7,153.8,-488.2,207.5,-658.5c53.7,-170.3,84.5,-266.8,92.5,-289.5zM1001 80h400000v40h-400000z"/></svg></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.4395em;"><span></span></span></span></span></span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mord"><span class="delimsizing size2">(</span></span><span class="mopen">(</span><span class="mord">1</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">−</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:1.2em;vertical-align:-0.35em;"></span><span class="mord mathnormal" style="margin-right:0.0556em;">γ</span><span class="mclose">)</span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mord mathnormal" style="margin-right:0.044em;">z</span><span class="mord"><span class="delimsizing size1">(</span></span><span class="mord">1</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">−</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:1.2em;vertical-align:-0.35em;"></span><span class="mord"><span class="mopen nulldelimiter"></span><span class="mfrac"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.8451em;"><span style="top:-2.655em;"><span class="pstrut" style="height:3em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight" style="margin-right:0.109em;">N</span></span></span></span><span style="top:-3.23em;"><span class="pstrut" style="height:3em;"></span><span class="frac-line" style="border-bottom-width:0.04em;"></span></span><span style="top:-3.394em;"><span class="pstrut" style="height:3em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mtight">1</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.345em;"><span></span></span></span></span></span><span class="mclose nulldelimiter"></span></span><span class="mord"><span class="delimsizing size1">)</span></span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">+</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:1.2em;vertical-align:-0.35em;"></span><span class="mord mathnormal" style="margin-right:0.0556em;">γ</span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mord mathnormal" style="margin-right:0.044em;">z</span><span class="mord"><span class="delimsizing size1">(</span></span><span class="mord">1</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">−</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:1.8em;vertical-align:-0.65em;"></span><span class="mord"><span class="mopen nulldelimiter"></span><span class="mfrac"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.8451em;"><span style="top:-2.655em;"><span class="pstrut" style="height:3em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight" style="margin-right:0.109em;">N</span><span class="mord mathnormal mtight">e</span></span></span></span><span style="top:-3.23em;"><span class="pstrut" style="height:3em;"></span><span class="frac-line" style="border-bottom-width:0.04em;"></span></span><span style="top:-3.394em;"><span class="pstrut" style="height:3em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mtight">1</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.345em;"><span></span></span></span></span></span><span class="mclose nulldelimiter"></span></span><span class="mord"><span class="delimsizing size1">)</span></span><span class="mord"><span class="delimsizing size2">)</span></span></span></span></span></span></p><p>The more variants you try, and the more their Sharpes disagree, the higher the bar. The implementation is five lines.</p><figure class="highlight python"><table><tr><td class="code"><pre><span class="line"><span class="keyword">def</span> <span class="title function_">expected_max_sharpe</span>(<span class="params">n_trials: <span class="built_in">int</span>, var_sr: <span class="built_in">float</span></span>) -&gt; <span class="built_in">float</span>:</span><br><span class="line">    <span class="keyword">if</span> n_trials &lt; <span class="number">2</span>:</span><br><span class="line">        <span class="keyword">return</span> <span class="number">0.0</span></span><br><span class="line">    z1 = stats.norm.ppf(<span class="number">1</span> - <span class="number">1</span> / n_trials)</span><br><span class="line">    z2 = stats.norm.ppf(<span class="number">1</span> - <span class="number">1</span> / (n_trials * math.e))</span><br><span class="line">    <span class="keyword">return</span> math.sqrt(var_sr) * ((<span class="number">1</span> - EULER_GAMMA) * z1 + EULER_GAMMA * z2)</span><br></pre></td></tr></table></figure><p>A formula like this is easy to get subtly wrong and impossible to eyeball, so <code>tests/test_metrics.py</code> checks it against a Monte Carlo estimate of the mean maximum of 50 normals over 20,000 draws and requires agreement within 5%.</p><h2 id="architecture">architecture</h2><p>The pipeline is a straight line from a public archive to a report, with two scripts that carry the protocol.</p><figure data-figure="diagram:crypto-pipeline"></figure><p>Three design choices carry most of the weight.</p><h3 id="everything-is-a-rank">everything is a rank</h3><p>Features and the training label are converted to centered cross-sectional ranks in <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mo stretchy="false">[</mo><mo>−</mo><mn>0.5</mn><mo separator="true">,</mo><mn>0.5</mn><mo stretchy="false">]</mo></mrow><annotation encoding="application/x-tex">[-0.5, 0.5]</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mopen">[</span><span class="mord">−</span><span class="mord">0.5</span><span class="mpunct">,</span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mord">0.5</span><span class="mclose">]</span></span></span></span> each day. Crypto daily returns have tails that would dominate any raw regression. The largest move inside the development universe was DOGE at +392% on 2021-01-28 (<code>results/dev/data_report.json</code>), and a squared-error fit on raw returns would spend most of its capacity on a handful of such days. The cost of ranking, which I did not appreciate until the holdout, is that the models optimize rank IC, and rank IC is not the same thing as PnL.</p><h3 id="wide-frames-for-features-a-long-frame-for-models">wide frames for features, a long frame for models</h3><p>Every rolling feature is one vectorized call on a date by symbol frame with a complete daily index, so a row is a day and a missing bar is a NaN that propagates through <code>min_periods</code>. Only universe members are stacked into the long <code>(date, symbol)</code> frame the models see. In development the wide frames are 2,373 days by 522 listing episodes (<code>results/benchmark.json</code>).</p><h3 id="the-backtest-is-two-lines-of-arithmetic">the backtest is two lines of arithmetic</h3><p>Weights formed at the close of day <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>t</mi></mrow><annotation encoding="application/x-tex">t</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6151em;"></span><span class="mord mathnormal">t</span></span></span></span> earn the return from close <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>t</mi></mrow><annotation encoding="application/x-tex">t</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6151em;"></span><span class="mord mathnormal">t</span></span></span></span> to close <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>t</mi><mo>+</mo><mn>1</mn></mrow><annotation encoding="application/x-tex">t+1</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6984em;vertical-align:-0.0833em;"></span><span class="mord mathnormal">t</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">+</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:0.6444em;"></span><span class="mord">1</span></span></span></span>. On a market that trades around the clock, the daily close at midnight UTC and the next open are the same instant, which removes a whole class of execution-timing questions that daily equity backtests have to answer.</p><figure class="highlight python"><table><tr><td class="code"><pre><span class="line">W = weights.unstack(<span class="string">&quot;symbol&quot;</span>).fillna(<span class="number">0.0</span>).sort_index()</span><br><span class="line">R = fwd_ret.unstack(<span class="string">&quot;symbol&quot;</span>).reindex_like(W).fillna(<span class="number">0.0</span>)</span><br><span class="line"><span class="keyword">if</span> lag:</span><br><span class="line">    W = W.shift(lag).fillna(<span class="number">0.0</span>)</span><br><span class="line">gross = (W * R).<span class="built_in">sum</span>(axis=<span class="number">1</span>)</span><br><span class="line">turnover = W.diff().<span class="built_in">abs</span>().<span class="built_in">sum</span>(axis=<span class="number">1</span>)</span><br></pre></td></tr></table></figure><p>Net returns at any cost are <code>gross - turnover * bps / 1e4</code>, so the whole cost grid comes from one backtest. A coin that has no bar the next day (a delisting) earns 0. The weight was chosen without knowing the bar would be missing, and the backtest must not know it either. Zero is still a choice, and an optimistic one for long positions, since a real delisting usually comes with a loss. It touches 58 of 79,190 development member rows, so the effect on the results is small, but it is not modeled.</p><h2 id="implementation">implementation</h2><p>The library is about 600 lines of numpy, pandas, scikit-learn and scipy, in eight modules under <code>src/marketpred/</code>, plus 25 pytest tests that run on synthetic panels with no network.</p><h3 id="data">data</h3><p>I list every USDT symbol in the archive (735) and download every monthly daily-kline zip from 2018-01 to 2026-08. That is 27,721 files, fetched with 32 threads in 400 seconds, parsed into 829,342 rows across 734 symbols and a 25 MB parquet (<code>results/download_log.txt</code>). Downloading every pair rather than a list of well-known coins is the main defense against survivorship bias. A hand-picked list of &quot;top coins&quot; encodes today's knowledge of which ones survived.</p><h3 id="the-universe">the universe</h3><p>On each day a symbol is eligible only if, using data up to and including that day, it has at least 60 bars of history, its base asset is not a stablecoin, fiat currency, gold token, wrapped or staked duplicate, or leveraged token, its trailing 30 day median quote volume is at least 1,000,000 USDT, and that volume ranks in the top 50 of eligible names. Days with fewer than 20 names are skipped.</p><p>The rule sounds simple and is the easiest place in the whole pipeline to leak the future. &quot;Top 50 by volume&quot; computed over the full sample, or a coin's history counted from its first bar in a later file, silently uses information from after day <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>t</mi></mrow><annotation encoding="application/x-tex">t</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6151em;"></span><span class="mord mathnormal">t</span></span></span></span>. <code>eligibility()</code> builds every quantity from cumulative or trailing operations on the wide frames, and a test perturbs every price from a date <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>T</mi></mrow><annotation encoding="application/x-tex">T</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal" style="margin-right:0.1389em;">T</span></span></span></span> on and asserts that the universe and every feature up to and including <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>T</mi><mo>−</mo><mn>1</mn></mrow><annotation encoding="application/x-tex">T-1</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.7667em;vertical-align:-0.0833em;"></span><span class="mord mathnormal" style="margin-right:0.1389em;">T</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">−</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:0.6444em;"></span><span class="mord">1</span></span></span></span> are unchanged. That test had a hole in its first version, described in <a href="#problems">problems</a>.</p><h3 id="features">features</h3><p>Nine features, computed at the close of day <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>t</mi></mrow><annotation encoding="application/x-tex">t</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6151em;"></span><span class="mord mathnormal">t</span></span></span></span> and ranked among that day's members.</p><table><thead><tr><th>Feature</th><th>Definition</th></tr></thead><tbody><tr><td><code>ret_1d</code>, <code>ret_5d</code>, <code>ret_20d</code>, <code>ret_60d</code></td><td>trailing returns, the reversal and momentum family</td></tr><tr><td><code>vol_20d</code></td><td>20 day standard deviation of log returns</td></tr><tr><td><code>dvol_20d</code></td><td>log of 20 day median quote volume</td></tr><tr><td><code>amihud_20d</code></td><td>20 day mean of absolute return over quote volume, an illiquidity measure</td></tr><tr><td><code>dist_high_20d</code></td><td>close over the 20 day high, minus 1</td></tr><tr><td><code>beta_btc_60d</code></td><td>60 day rolling beta of log returns to BTC</td></tr></tbody></table><p>A member with a missing value gets rank 0, the cross-sectional median, so a model never sees a NaN and never learns from missingness itself.</p><h3 id="models">models</h3><p>Four single factors with no fitting (<code>rev_1d</code>, <code>rev_5d</code>, <code>mom_20d</code>, <code>mom_60d</code>), ridge regression at alpha 1, 10 and 100, and scikit-learn's <code>HistGradientBoostingRegressor</code> at depth 3 and depth 6 (200 iterations, learning rate 0.05). The fitted models predict the cross-sectional rank of the next-day return from all nine ranked features.</p><h3 id="walk-forward-with-an-embargo">walk-forward with an embargo</h3><p>Development runs nine half-year test folds, 2020H1 through 2024H1, each with an expanding training window. The label on a training row dated <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>d</mi></mrow><annotation encoding="application/x-tex">d</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6944em;"></span><span class="mord mathnormal">d</span></span></span></span> is only known at the close of <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>d</mi><mo>+</mo><mn>1</mn></mrow><annotation encoding="application/x-tex">d+1</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.7778em;vertical-align:-0.0833em;"></span><span class="mord mathnormal">d</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">+</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:0.6444em;"></span><span class="mord">1</span></span></span></span>, so the embargo has to be counted from the label, not the row.</p><figure class="highlight python"><table><tr><td class="code"><pre><span class="line"><span class="meta">@property</span></span><br><span class="line"><span class="keyword">def</span> <span class="title function_">train_end</span>(<span class="params">self</span>) -&gt; pd.Timestamp:</span><br><span class="line">    <span class="comment"># A training row dated d has a label that is known at the close of d+1.</span></span><br><span class="line">    <span class="comment"># Requiring d + 1 &lt;= test_start - EMBARGO_DAYS leaves a gap of</span></span><br><span class="line">    <span class="comment"># EMBARGO_DAYS between the last label and the first test day.</span></span><br><span class="line">    <span class="keyword">return</span> <span class="variable language_">self</span>.test_start - pd.Timedelta(days=EMBARGO_DAYS + <span class="number">1</span>)</span><br></pre></td></tr></table></figure><p>Five days is more than a 1 day label strictly needs, and it costs almost nothing. A test fits a spy model that records the latest date it was shown and asserts it never passes <code>train_end</code>.</p><h3 id="the-holdout-lock">the holdout lock</h3><p>The holdout script is short and its first few lines are the protocol.</p><figure class="highlight python"><table><tr><td class="code"><pre><span class="line"><span class="keyword">if</span> LOCK.exists():</span><br><span class="line">    sys.exit(<span class="string">f&quot;refusing: holdout already evaluated, see <span class="subst">&#123;LOCK&#125;</span>&quot;</span>)</span><br><span class="line">sel = json.loads((PROJECT_ROOT / <span class="string">&quot;results&quot;</span> / <span class="string">&quot;dev&quot;</span> / <span class="string">&quot;selection.json&quot;</span>).read_text())</span><br><span class="line">prereg = (PROJECT_ROOT / <span class="string">&quot;PREREGISTRATION.md&quot;</span>).read_bytes()</span><br><span class="line">lock = &#123;<span class="string">&quot;started_utc&quot;</span>: datetime.now(timezone.utc).isoformat(),</span><br><span class="line">        <span class="string">&quot;preregistration_sha256&quot;</span>: hashlib.sha256(prereg).hexdigest(),</span><br><span class="line">        <span class="string">&quot;selection&quot;</span>: sel, <span class="string">&quot;status&quot;</span>: <span class="string">&quot;started&quot;</span>&#125;</span><br><span class="line">LOCK.write_text(json.dumps(lock, indent=<span class="number">2</span>))</span><br></pre></td></tr></table></figure><p>The lock is written before any holdout number exists, so a crash halfway through still counts as the one touch. This is not tamper-proof against me, since I could delete the file, but deleting it would be a deliberate act rather than an accident, and the hash makes any later edit to the pre-registration detectable. The lock records a start and a completion four minutes apart on 2026-09-27, with verdicts &quot;supported&quot; and &quot;not supported&quot; (<code>results/holdout/LOCK.json</code>).</p><h2 id="problems">problems</h2><p>None of these broke the build. Most of them would have produced plausible numbers.</p><h3 id="the-api-was-blocked-and-the-replacement-was-better">the API was blocked, and the replacement was better</h3><p>The Binance REST API returns HTTP 451 from the US. I switched to the static archive at data.binance.vision, which turned out to be the better research source anyway, because it keeps the monthly files of pairs that have since been delisted.</p><h3 id="timestamps-changed-units-mid-archive">timestamps changed units mid-archive</h3><p>The archive switched its kline timestamps from milliseconds to microseconds starting in 2025. Parsed naively as milliseconds, every bar from 2025 on lands tens of thousands of years in the future. I knew about the switch before writing the parser, so <code>parse_kline_csv</code> checks magnitude row by row.</p><figure class="highlight python"><table><tr><td class="code"><pre><span class="line">t = pd.to_numeric(df[<span class="string">&quot;open_time&quot;</span>]).astype(<span class="string">&quot;int64&quot;</span>).to_numpy()</span><br><span class="line">unit_us = t &gt; <span class="number">10</span>**<span class="number">14</span>   <span class="comment"># Binance spot archives use microseconds from 2025-01 onward</span></span><br><span class="line">t_ms = np.where(unit_us, t // <span class="number">1000</span>, t)</span><br></pre></td></tr></table></figure><p>A test feeds it one row of each unit. The end-to-end check is that the parsed panel ends on 2026-08-31.</p><h3 id="one-ticker-two-assets">one ticker, two assets</h3><p>LUNAUSDT was the original Terra until its collapse in May 2022, and the new Terra from late May 2022 onward. A 60 day return computed across that gap compares two unrelated assets, and the &quot;return&quot; is meaningless. The fix is to split any gap longer than 7 days into a new listing episode, <code>LUNAUSDT#2</code>, which the rest of the pipeline treats as a separate symbol. In the development data this turned 507 symbols into 522 episodes (<code>results/dev/data_report.json</code>). Both episodes show up among the five largest daily moves in the development universe, the old LUNA at -99.97% on 2022-05-12 and the new one at +168% on 2022-09-09.</p><h3 id="the-leveraged-token-filter-ate-a-real-coin">the leveraged-token filter ate a real coin</h3><p>My first rule for leveraged tokens excluded any base ending in UP, DOWN, BULL or BEAR. That catches BTCUP and ETHBEAR, and also JUP. The rule now fires only when the stripped stem is itself a listed base, and the test includes JUP and SUPER as coins that must survive.</p><h3 id="pandas-3-changed-stack">pandas 3 changed stack</h3><p>pandas 3 changed <code>DataFrame.stack()</code> to keep NaN rows. My first version of the long frame stacked the ranked feature frames directly, which under pandas 3 would have included every cell of the date by symbol grid rather than only universe members, so rows for coins outside the universe would have reached the models and the IC. I now stack the boolean mask, keep the true cells, and reindex every feature onto that index explicitly.</p><h3 id="the-run-that-hung">the run that hung</h3><p>This one cost the most. The first development run logged 14 trials, all the single factors and all the ridge models, and then sat at 0.2% CPU for more than ten minutes with no output. Nothing crashed and nothing timed out.</p><p>A standalone repro with <code>faulthandler</code> dumping stacks showed the process stuck inside <code>HistGradientBoostingRegressor</code>, in <code>_initialize_root</code>, on 80,000 random rows. With <code>OMP_NUM_THREADS=1</code> the same fit finished in 3.4 seconds. With 4 or 14 threads it did not finish within 60 seconds. scikit-learn ships its own OpenMP runtime, and the machine was also running 14 other agents' jobs at the time. I did not find the root cause. The fix is a <code>threadpool_limits(1)</code> context around every scikit-learn fit and predict.</p><figure class="highlight python"><table><tr><td class="code"><pre><span class="line"><span class="keyword">def</span> <span class="title function_">fit</span>(<span class="params">self, train: pd.DataFrame</span>):</span><br><span class="line">    t = train.dropna(subset=[<span class="string">&quot;y_rank&quot;</span>])</span><br><span class="line">    <span class="keyword">with</span> threadpool_limits(<span class="number">1</span>):</span><br><span class="line">        <span class="variable language_">self</span>.est.fit(t[FEATURES].to_numpy(np.float32), t[<span class="string">&quot;y_rank&quot;</span>].to_numpy(np.float32))</span><br><span class="line">    <span class="keyword">return</span> <span class="variable language_">self</span></span><br></pre></td></tr></table></figure><p>A depth 3 fit on 75,250 rows takes about 3 seconds single-threaded, so the price is small. The more interesting cost is on the record. The pre-registration says every trial ever evaluated in development is appended to <code>results/dev/trials.jsonl</code> and the deflated Sharpe uses its line count as <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>N</mi></mrow><annotation encoding="application/x-tex">N</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal" style="margin-right:0.109em;">N</span></span></span></span>. The hung run had already appended 14 lines, so the file has 32 lines rather than 18, and the deflated Sharpe below uses <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>N</mi><mo>=</mo><mn>32</mn></mrow><annotation encoding="application/x-tex">N = 32</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal" style="margin-right:0.109em;">N</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.6444em;"></span><span class="mord">32</span></span></span></span>. That overstates the number of distinct trials, since 14 of them are repeats, which makes the bar higher and the test more conservative. I left it that way because the rule said to count lines, and a rule you adjust when it is inconvenient is not a rule.</p><h3 id="a-bookkeeping-bug-i-found-after-the-holdout">a bookkeeping bug I found after the holdout</h3><p>After the holdout ran I noticed that the lag 1 robustness blocks in <code>holdout_report.json</code> report the IC of the unlagged forecast. The IC fields are copied from the lag 0 evaluation, while the Sharpe and return fields in those blocks are correctly lagged. Fixing it means rerunning the holdout, which the protocol forbids, so the file stays as it is and the bug is documented. The lagged IC numbers are not used anywhere in this post. The verification found one more issue in the same lag 1 blocks. <code>fwd_ret</code> only exists on rows where a coin is a universe member, so a lagged weight on a coin that left the universe the next day earns 0 instead of its actual return. This affects only the lag 1 robustness numbers, such as the -0.43 quoted below, and neither H1 nor H2.</p><h3 id="a-look-ahead-test-that-could-not-fail">a look-ahead test that could not fail</h3><p>The test that guards against look-ahead in the features changes every price from a date <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>T</mi></mrow><annotation encoding="application/x-tex">T</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal" style="margin-right:0.1389em;">T</span></span></span></span> on and checks that nothing before it moves. Its docstring said features at <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>T</mi><mo>−</mo><mn>1</mn></mrow><annotation encoding="application/x-tex">T-1</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.7667em;vertical-align:-0.0833em;"></span><span class="mord mathnormal" style="margin-right:0.1389em;">T</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">−</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:0.6444em;"></span><span class="mord">1</span></span></span></span> must not change, but the code compared only dates strictly before <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>T</mi><mo>−</mo><mn>1</mn></mrow><annotation encoding="application/x-tex">T-1</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.7667em;vertical-align:-0.0833em;"></span><span class="mord mathnormal" style="margin-right:0.1389em;">T</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">−</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:0.6444em;"></span><span class="mord">1</span></span></span></span>, with a <code>&lt;</code> where it needed <code>&lt;=</code>. That left exactly the case that matters most unguarded, a feature at the close of <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>T</mi><mo>−</mo><mn>1</mn></mrow><annotation encoding="application/x-tex">T-1</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.7667em;vertical-align:-0.0833em;"></span><span class="mord mathnormal" style="margin-right:0.1389em;">T</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">−</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:0.6444em;"></span><span class="mord">1</span></span></span></span> that quietly uses the bar at <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>T</mi></mrow><annotation encoding="application/x-tex">T</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal" style="margin-right:0.1389em;">T</span></span></span></span>.</p><p>I did not catch this. The independent verification did, with a mutation test. Replacing <code>ret_1d</code> with a one day lead, <code>close.shift(-1) / close - 1</code>, passed the original test. With <code>&lt;=</code> the lead is caught and the real features still pass, so the test now uses <code>&lt;=</code> and the suite is still 25 passing tests. The real features never had a lead, so no result changes, but for the whole study the test that was supposed to prove that could not have failed on the most likely bug. A test only counts once you have seen it fail on the thing it is meant to catch.</p><h3 id="the-one-that-shapes-the-result">the one that shapes the result</h3><p>The last problem is not a bug. High IC did not mean high PnL, and the pre-registered selection rule picked the only trial family with a negative IC. That is the story of the next two sections.</p><h2 id="experiments">experiments</h2><p>All runs use the data described above. Development is 2020-01-01 to 2024-06-30 in 9 walk-forward folds. The holdout is 2024-07-01 to 2026-08-31, 792 days. Costs are 15 bps per unit of turnover unless stated, and Sharpe ratios are annualized with 365 days. Annual returns are the daily mean times 365, not compounded.</p><h3 id="data-checks">data checks</h3><p>The development universe had a median of 50 names and a minimum of 26 on 1,643 days, drawn from 279 distinct listing episodes over the period. Of 79,190 member rows, only 58 had no next-day bar (<code>results/dev/data_report.json</code>). The holdout universe also had a median of 50 names (<code>results/holdout/holdout_report.json</code>).</p><h3 id="development-walk-forward-all-18-trials">development walk-forward, all 18 trials</h3><p>A selection of the 18 rows, from <code>results/dev/summary.csv</code>.</p><table><thead><tr><th>Trial</th><th>Mean IC</th><th>NW t</th><th>Gross Sharpe</th><th>Net Sharpe</th><th>Turnover per day</th></tr></thead><tbody><tr><td>rev_1d, rank</td><td>+0.036</td><td>6.98</td><td>0.15</td><td>-2.86</td><td>1.33</td></tr><tr><td>rev_5d, rank</td><td>+0.033</td><td>6.03</td><td>-1.13</td><td>-2.34</td><td>0.60</td></tr><tr><td>mom_20d, quintile</td><td>-0.029</td><td>-5.18</td><td>0.91</td><td>0.33</td><td>0.39</td></tr><tr><td>mom_60d, rank</td><td>-0.049</td><td>-8.51</td><td>-0.03</td><td>-0.42</td><td>0.19</td></tr><tr><td>ridge alpha 10, rank</td><td>+0.097</td><td>17.29</td><td>0.24</td><td>-0.92</td><td>0.56</td></tr><tr><td>ridge alpha 10, quintile</td><td>+0.097</td><td>17.29</td><td>0.45</td><td>-0.75</td><td>0.74</td></tr><tr><td>gbm depth 3, rank</td><td>+0.108</td><td>21.38</td><td>1.49</td><td>-0.17</td><td>0.71</td></tr><tr><td>gbm depth 3, quintile</td><td>+0.108</td><td>21.38</td><td>1.64</td><td>0.06</td><td>0.88</td></tr><tr><td>gbm depth 6, quintile</td><td>+0.099</td><td>20.77</td><td>1.77</td><td>-0.08</td><td>0.99</td></tr></tbody></table><p>The signals behaved exactly as the literature would predict at the level of ranks. Both reversal factors had positive IC in all 9 folds. Both momentum factors had negative IC, in 8 of 9 folds for <code>mom_20d</code> and all 9 for <code>mom_60d</code> (<code>results/dev/fold_ic.csv</code>). Momentum over 20 and 60 days is, in this universe, a reversal signal at a longer horizon.</p><figure data-figure="chart:projects/can-we-predict-markets/reversal-ic"></figure><p>The fitted models roughly tripled the reversal IC, to about 0.10, and they did it mostly without reversal. Ridge put its largest coefficient on low volatility, about -0.08 on <code>vol_20d</code> in every fold, and its second largest on 1 day reversal, about -0.04 on <code>ret_1d</code> (<code>results/dev/ridge_coefs.csv</code>). One possible reading, not tested here, is that this is partly a property of rank labels. On a typical day the median coin loses a little, high-volatility coins spread into both tails, and a forecast that says &quot;low-volatility coins land nearer the middle and slightly above&quot; wins a lot of small rank bets. Whether that is worth anything in dollars is exactly the question rank IC cannot answer.</p><p>Look at the last two columns of the table. The gradient boosting models had gross Sharpes of 1.49 to 1.77 and handed almost all of it to costs, because a forecast that re-ranks 50 coins daily trades most of the book daily. The depth 3 quintile portfolio earned 50.4% a year gross and 1.9% net at a turnover of 0.88 (<code>results/dev/summary.csv</code>). The single trial with a clearly positive net Sharpe was <code>mom_20d</code> with quintile weights, which had the wrong-signed IC for its name and a low turnover of 0.39.</p><h3 id="selection-and-multiple-testing">selection and multiple testing</h3><p>The frozen rule picked <code>mom_20d</code> with quintile weights, the highest development net Sharpe at 0.33, with a PSR of 0.75 (<code>results/dev/selection.json</code>). Then the deflated Sharpe ratio asked how impressive 0.33 is after 32 logged trials with the observed spread of trial Sharpes. The expected best annualized Sharpe of 32 worthless strategies is 2.01, and the deflated Sharpe ratio of the selected trial is 0.00018.</p><figure data-figure="chart:projects/can-we-predict-markets/deflated-sharpe"></figure><p>In plain words, the best development result was exactly what you would expect from the luckiest of 32 useless strategies. This was known before the holdout was opened, and the protocol said to open it anyway with the selected trial. That is the right call. The deflated Sharpe is an argument about how much to believe a number, not a license to go back and pick something else.</p><h3 id="the-holdout">the holdout</h3><p>Every model was refit once on data up to 2024-06-25 and run without refitting over the 792 holdout days. Only two evaluations are confirmatory, <code>rev_1d</code> for H1 and <code>mom_20d</code> quintile for H2. The script also evaluates the other 16 trials and writes them to a file named <code>all_trials_posthoc.csv</code>, so the name carries the warning.</p><h2 id="results">results</h2><h3 id="h1-reversal-as-a-signal-supported">H1, reversal as a signal, supported</h3><p>The holdout mean rank IC of <code>rev_1d</code> is +0.019, with a Newey-West t of 2.24 and a daily hit rate of 53.1% (<code>results/holdout/holdout_report.json</code>). That clears the pre-registered bar of t above 2, barely. It is about half the development IC of 0.036, which is what decay out of sample usually looks like, and it is still the same sign it had in all nine development folds. Short-term reversal exists in the cross-section of liquid crypto, in ranks, in a period the model never saw.</p><h3 id="reversal-as-a-strategy-dead-before-costs">reversal as a strategy, dead before costs</h3><p>The same signal traded with rank weights has a holdout gross Sharpe of -0.78 and a net Sharpe of -4.06 at 15 bps, with turnover of 1.31 a day and a maximum drawdown of -86% (<code>results/holdout/holdout_report.json</code>). It lost money at zero cost. The ordering is right slightly more often than not, and the dollars go the other way.</p><p>This is the central finding, so it is worth being precise about how it can happen. Rank IC weights every coin equally and counts only whether the order is right. PnL weights every coin by its position times its return, and in crypto the returns that matter are the few coins that move 20% in a day. A signal can win 53% of its rank comparisons and lose on the coins where it is wrong by the most. My working explanation, which I have not tested, is that the coins that just fell hardest are also the ones most likely to keep falling hard. Reversal holds on average across the cross-section and fails in the tail that sets the PnL. Testing that would mean splitting the holdout PnL by the size of the prior move, and that analysis would be post hoc on a spent holdout.</p><h3 id="h2-the-selected-portfolio-not-supported">H2, the selected portfolio, not supported</h3><p><code>mom_20d</code> with quintile weights earned a holdout gross Sharpe of 0.77, a net Sharpe of 0.08, a PSR of 0.55 and a maximum drawdown of -29% (<code>results/holdout/holdout_report.json</code>). Positive, but indistinguishable from zero, and far from the 0.95 PSR the pre-registration required. Its holdout IC was -0.004 with a t of -0.52, so it had no measurable rank skill at all. Its gross PnL came from the right tail of a few trending coins, which is the same tail effect that sank reversal, pointing the other way.</p><figure data-figure="chart:projects/can-we-predict-markets/holdout-cost"></figure><p>The cost sensitivity shows how thin the edge was. Net Sharpe falls from 0.77 at 0 bps to 0.54 at 5 bps and -0.37 at 25 bps. Delaying execution by one day takes it to -0.43 at 15 bps (<code>results/holdout/holdout_report.json</code>, the <code>H2_selected_lag1</code> block, whose Sharpe fields are correctly lagged). None of this includes the funding cost of holding the short leg as a perpetual future, which is not modeled at all and would make things worse in a rising market.</p><h3 id="post-hoc-and-not-a-claim">post hoc, and not a claim</h3><p>Because the holdout script evaluated every trial for context, I can see that the selection rule picked badly. The gradient boosting depth 3 quintile portfolio had a holdout net Sharpe of 0.71, and the three ridge quintile portfolios about 0.55, with holdout ICs near 0.10 (<code>results/holdout/all_trials_posthoc.csv</code>). The Spearman correlation between development and holdout net Sharpe across the 18 trials is 0.46 (<code>results/holdout/holdout_report.json</code>).</p><figure data-figure="chart:projects/can-we-predict-markets/dev-vs-holdout"></figure><p>It is tempting to write &quot;the machine learning models work&quot;. I cannot. I have now looked at the holdout, and picking the winner after looking is exactly the multiple-testing error the protocol exists to prevent. If I had pre-registered &quot;select on IC&quot;, I would have picked gradient boosting and this would be a positive result. I did not, and I would be writing a different rule after seeing which one would have won. Even the observation itself is weaker than it looks. None of those holdout Sharpes would pass the pre-registered PSR bar (the best, gbm depth 3 quintile, has a PSR of 0.85), none carries short-side funding costs, and all of them sit inside a set of 18 that includes four portfolios below -2. The honest status is a hypothesis for the next fresh holdout.</p><h3 id="so-can-we-predict-markets">so, can we predict markets?</h3><p>At the level of daily cross-sectional ranks in liquid crypto, yes, a little. The reversal IC was positive in nine half-years of development and in a separate two year holdout, and the fitted models found a larger and equally stable rank signal. Turning any of it into money after 15 bps per trade did not happen in the pre-registered test. The main thing the study measured is the distance between those two statements.</p><h3 id="benchmark">benchmark</h3><p>Stage timings, median of 3 runs on an M4 Pro that was heavily loaded by other jobs at the time, so treat them as rough (<code>results/benchmark.json</code>).</p><table><thead><tr><th>Stage</th><th>Time</th></tr></thead><tbody><tr><td>Build the development dataset (84,350 member rows)</td><td>11.1 s</td></tr><tr><td>Ridge fit on 75,250 rows</td><td>0.01 to 0.19 s</td></tr><tr><td>Gradient boosting fit, depth 3</td><td>3.0 s</td></tr><tr><td>Gradient boosting fit, depth 6</td><td>4.6 s</td></tr><tr><td>Daily IC over the whole development set</td><td>2.5 s</td></tr><tr><td>One backtest</td><td>0.12 s</td></tr></tbody></table><p>In the development run, the two gradient boosting walk-forwards took 76 s and 96 s of a roughly 4 minute total (<code>results/dev/timings.json</code>). The whole study, from raw archive to holdout verdict, runs in about 15 minutes of wall time on a laptop, most of it the download.</p><h2 id="what-i-would-change">what I would change</h2><h3 id="select-on-ic-or-on-something-that-knows-about-turnover">select on IC, or on something that knows about turnover</h3><p>The IC ranking of trials was stable from development to holdout. The net Sharpe ranking was noisy, and the rule chose a strategy whose only merit was a lucky right tail. A pre-registered score that combines IC with expected turnover and cost would have been more stable. I did not change the rule after seeing development results, and I should not have, but it is the first thing I would write differently.</p><h3 id="cut-turnover-inside-the-portfolio">cut turnover inside the portfolio</h3><p>The gradient boosting models had gross Sharpes near 1.5 to 1.8 in development and lost most of that to costs. Trading toward the target weights with a no-trade band, or smoothing the forecast over a few days, attacks turnover directly rather than hoping a low-turnover signal wins the selection.</p><h3 id="keep-a-second-untouched-holdout">keep a second, untouched holdout</h3><p>The post-hoc gradient boosting result deserves a real test, and there is no clean data left to give it one. Either wait for data after 2026-08, or hold out a second period or asset class from the start next time so a finding like this has somewhere to go.</p><h3 id="model-what-the-short-leg-actually-costs">model what the short leg actually costs</h3><p>Binance spot does not allow shorting, so a real implementation would short perpetual futures and pay or receive funding. That needs a separate downloader for historical funding rates. Weight drift between rebalances should also enter the turnover calculation, which currently uses target weights and slightly understates it.</p><h3 id="use-a-smaller-stricter-universe">use a smaller, stricter universe</h3><p>The names ranked 30 to 50 by volume are where both the tails and the costs live. A top 20 or top 30 universe would test whether the rank signal survives where trading is cheap.</p><h3 id="test-the-thing-that-broke">test the thing that broke</h3><p>Fix the lag 1 IC bookkeeping and the missing returns for coins that leave the universe before the next holdout, and add a test that the lag option changes every lag-dependent field. The bug was only possible because no test asserted that. More generally, run a mutation against every guard test, the way the verification did for the look-ahead test, so each one is known to fail when it should.</p><h3 id="price-delistings-as-losses">price delistings as losses</h3><p>A coin with no next bar currently earns 0. A pessimistic delisting return, or at least a sensitivity run with one, would remove a small bias in favor of long positions.</p><h2 id="reproducibility">reproducibility</h2><p>The project is a <code>uv</code> environment pinned to Python 3.12. The tests use synthetic panels, so they need no download.</p><figure class="highlight bash"><table><tr><td class="code"><pre><span class="line"><span class="built_in">cd</span> projects/09-crypto-reversal-preregistered</span><br><span class="line">uv <span class="built_in">sync</span></span><br><span class="line">uv run pytest -q</span><br></pre></td></tr></table></figure><p>The full study, in order.</p><figure class="highlight bash"><table><tr><td class="code"><pre><span class="line">uv run python scripts/download_data.py   <span class="comment"># about 7 minutes, writes data/processed/klines_1d.parquet</span></span><br><span class="line">uv run python scripts/run_dev.py         <span class="comment"># about 4 minutes, writes results/dev/</span></span><br><span class="line">uv run python scripts/run_holdout.py     <span class="comment"># one shot, refuses if results/holdout/LOCK.json exists</span></span><br><span class="line">uv run python scripts/plot_posthoc.py    <span class="comment"># dev vs holdout figure from saved CSVs only</span></span><br><span class="line">uv run python scripts/benchmark.py       <span class="comment"># writes results/benchmark.json</span></span><br></pre></td></tr></table></figure><p>The holdout has already been evaluated and the lock file is in place, so <code>run_holdout.py</code> exits with an error. That is the point. Re-running <code>run_dev.py</code> appends another 18 lines to <code>trials.jsonl</code>, which by the pre-registered rule raises <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>N</mi></mrow><annotation encoding="application/x-tex">N</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal" style="margin-right:0.109em;">N</span></span></span></span> for any future deflated Sharpe.</p><h3 id="verification">verification</h3><p>An independent reviewer who did not build the project reran it on 2026-09-26, with threads capped at 2 on a shared machine. The test suite passed (25 tests). Re-running the download from the cached archive produced an identical parquet (829,342 rows, 734 symbols). The full development study, run into a scratch folder seeded so the trial count again reached 32, reproduced <code>summary.csv</code>, <code>fold_ic.csv</code>, <code>ridge_coefs.csv</code> and <code>daily_net_returns.csv</code> exactly, with the same selection, the same deflated Sharpe of 0.00018 and the same threshold of 2.01. <code>run_holdout.py</code> refused to run because the lock exists, so the reviewer recomputed the same code path from the frozen <code>selection.json</code> into a scratch folder, which makes no new decision. Every field of <code>holdout_report.json</code> and every cell of <code>all_trials_posthoc.csv</code> matched exactly. The reviewer also confirmed the pre-registration hash against the lock and the file timestamps described above, and fixed the look-ahead test. Timing was deferred. The load average was between 9 and 60, so the benchmark was only checked to run, and the stage timings above have not been remeasured on a quiet machine.</p><p>Code, the pre-registration and every result file are in <a href="https://github.com/SadeekFarhan21/blog/tree/main/projects/09-crypto-reversal-preregistered"><code>projects/09-crypto-reversal-preregistered</code></a>.</p>]]>
    </content>
    <id>https://projects.farhansadeek.com/posts/crypto-reversal-is-real-in-ranks-and-dead-in-dollars/</id>
    <link href="https://projects.farhansadeek.com/posts/crypto-reversal-is-real-in-ranks-and-dead-in-dollars/"/>
    <published>2026-09-29T06:17:08.000Z</published>
    <summary>A pre-registered study of the 50 most liquid Binance pairs finds that one day reversal predicts next day ranks in a two year holdout, while the portfolio chosen in development earns a net Sharpe of 0.08 after 15 bps costs and fails its test.</summary>
    <title>Testing Crypto Reversal Under a Pre-Registered Protocol</title>
    <updated>2026-09-29T06:49:26.552Z</updated>
  </entry>
  <entry>
    <author>
      <name>Farhan Sadeek</name>
      <email>farhan@farhansadeek.com</email>
    </author>
    <category term="quant" scheme="https://projects.farhansadeek.com/tags/quant/"/>
    <category term="backtesting" scheme="https://projects.farhansadeek.com/tags/backtesting/"/>
    <category term="research" scheme="https://projects.farhansadeek.com/tags/research/"/>
    <content>
      <![CDATA[<p>I built qrp, a small research platform for daily crypto bars whose job is to produce numbers I would actually believe. It ingests 300 Binance USDT spot pairs from 2022-01-01 to 2026-08-31, stores them in an append-only parquet store with a point-in-time read path, computes features from a declarative spec, audits those features for look-ahead before every run, fits a model with purged walk-forward cross-validation, and backtests the resulting target weights with drift, fees, slippage and optional square-root impact. Every run lands in a local tracker you can diff from the command line.</p><p>The honest baseline, a cross-sectional ridge on seven ranked price and volume features, earns an out of sample <strong>Sharpe of 0.49</strong> over 1,338 days at 15 bps one way. Add a single feature that peeks five days ahead and switch off the audit's fail switch, and the same pipeline reports a <strong>Sharpe of 19.40</strong> with a max drawdown of 2.5%. The point of the project is the machinery that makes the second number impossible to produce by accident, and the tests that break when any piece of that machinery is wrong.</p><p>The platform is about 1,650 lines of Python with 73 tests. The compiled backtest kernel runs about 80 million asset-days per second on one thread of an Apple M4 Pro that was heavily loaded by other jobs, so every throughput number here is a lower bound.</p><p><em>Reading note.</em> The argument is that most great-looking backtests are wrong in one of five boring ways, and that each one can be turned into a written rule with a test attached. The problems section is where the rules earned their keep. Skip to <a href="#results">results</a> if you only want the numbers.</p><h2 id="table-of-contents">table of contents</h2><ul><li><a href="#what-i-wanted-to-build">what i wanted to build</a></li><li><a href="#theory">theory</a></li><li><a href="#architecture">architecture</a></li><li><a href="#implementation">implementation</a></li><li><a href="#problems">problems</a></li><li><a href="#experiments">experiments</a></li><li><a href="#results">results</a></li><li><a href="#what-i-would-change">what i would change</a></li><li><a href="#reproducibility">reproducibility</a></li></ul><h2 id="what-i-wanted-to-build">what i wanted to build</h2><p>A backtest that looks great is usually broken, and the ways it breaks are few and dull. A feature quietly uses tomorrow's data. The model trains on labels whose returns overlap the test period. The simulator trades at the same close the signal just observed. Turnover is computed as the change in target weights, ignoring that weights drift with prices, so costs come out too low. Or the data itself contains a ticker that died and came back as a different asset, and the backtester sees a 177,000x return across the gap.</p><p>None of these need a clever strategy to show up. They show up in the plumbing. So the goal for v0 was not a strategy at all. It was the smallest pipeline in which each of those five failure modes has a written contract and a test that fails when the contract is broken, fast enough that later projects can sweep thousands of variants. I left out minute bars, and US equities for a mundane reason covered in <a href="#problems">problems</a>.</p><h2 id="theory">theory</h2><p>Five ideas carry the whole design. Each one is small, and each one is violated by default in a naive research loop.</p><h3 id="one-timing-convention-used-everywhere">one timing convention, used everywhere</h3><p>Row <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>t</mi></mrow><annotation encoding="application/x-tex">t</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6151em;"></span><span class="mord mathnormal">t</span></span></span></span> means &quot;after the close of day <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>t</mi></mrow><annotation encoding="application/x-tex">t</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6151em;"></span><span class="mord mathnormal">t</span></span></span></span>&quot;. Every stage is written against that single sentence.</p><ol><li>A feature at row <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>t</mi></mrow><annotation encoding="application/x-tex">t</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6151em;"></span><span class="mord mathnormal">t</span></span></span></span> may use bars at rows <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mo>≤</mo><mi>t</mi></mrow><annotation encoding="application/x-tex">\le t</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.7719em;vertical-align:-0.136em;"></span><span class="mrel">≤</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.6151em;"></span><span class="mord mathnormal">t</span></span></span></span> only.</li><li>A target weight decided at row <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>t</mi></mrow><annotation encoding="application/x-tex">t</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6151em;"></span><span class="mord mathnormal">t</span></span></span></span> executes at the close of row <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>t</mi><mo>+</mo><mi>d</mi></mrow><annotation encoding="application/x-tex">t + d</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6984em;vertical-align:-0.0833em;"></span><span class="mord mathnormal">t</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">+</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:0.6944em;"></span><span class="mord mathnormal">d</span></span></span></span>, where <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>d</mi></mrow><annotation encoding="application/x-tex">d</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6944em;"></span><span class="mord mathnormal">d</span></span></span></span> is the execution delay (default 1).</li><li>The training label for row <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>t</mi></mrow><annotation encoding="application/x-tex">t</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6151em;"></span><span class="mord mathnormal">t</span></span></span></span> is the return from close <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>t</mi><mo>+</mo><mi>d</mi></mrow><annotation encoding="application/x-tex">t + d</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6984em;vertical-align:-0.0833em;"></span><span class="mord mathnormal">t</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">+</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:0.6944em;"></span><span class="mord mathnormal">d</span></span></span></span> to close <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>t</mi><mo>+</mo><mi>d</mi><mo>+</mo><mi>h</mi></mrow><annotation encoding="application/x-tex">t + d + h</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6984em;vertical-align:-0.0833em;"></span><span class="mord mathnormal">t</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">+</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:0.7778em;vertical-align:-0.0833em;"></span><span class="mord mathnormal">d</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">+</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:0.6944em;"></span><span class="mord mathnormal">h</span></span></span></span>, for horizon <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>h</mi></mrow><annotation encoding="application/x-tex">h</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6944em;"></span><span class="mord mathnormal">h</span></span></span></span>. Its information interval is <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mo stretchy="false">[</mo><mi>t</mi><mo separator="true">,</mo><mtext> </mtext><mi>t</mi><mo>+</mo><mi>d</mi><mo>+</mo><mi>h</mi><mo stretchy="false">]</mo></mrow><annotation encoding="application/x-tex">[t,\ t + d + h]</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mopen">[</span><span class="mord mathnormal">t</span><span class="mpunct">,</span><span class="mspace"> </span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mord mathnormal">t</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">+</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:0.7778em;vertical-align:-0.0833em;"></span><span class="mord mathnormal">d</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">+</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mord mathnormal">h</span><span class="mclose">]</span></span></span></span>.</li><li>The universe at row <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>t</mi></mrow><annotation encoding="application/x-tex">t</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6151em;"></span><span class="mord mathnormal">t</span></span></span></span> (top 100 by trailing dollar volume, at least 60 bars of history, has a bar today) is computed from rows <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mo>≤</mo><mi>t</mi></mrow><annotation encoding="application/x-tex">\le t</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.7719em;vertical-align:-0.136em;"></span><span class="mrel">≤</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.6151em;"></span><span class="mord mathnormal">t</span></span></span></span>.</li></ol><p>Delay 0 is allowed but optimistic, since it trades at the very close the signal just read.</p><h3 id="look-ahead-as-an-empirical-property">look-ahead as an empirical property</h3><p>A feature is causal if its value at row <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>t</mi></mrow><annotation encoding="application/x-tex">t</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6151em;"></span><span class="mord mathnormal">t</span></span></span></span> is a function of rows <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mo>≤</mo><mi>t</mi></mrow><annotation encoding="application/x-tex">\le t</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.7719em;vertical-align:-0.136em;"></span><span class="mrel">≤</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.6151em;"></span><span class="mord mathnormal">t</span></span></span></span> only. You can try to enforce that by reading code or by labelling operators as safe, but both rely on someone noticing. I test it instead, by changing the future and checking that the past does not move.</p><p>Pick a cut row <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>c</mi></mrow><annotation encoding="application/x-tex">c</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal">c</span></span></span></span>. The <strong>truncation probe</strong> recomputes every feature on the panel cut to rows <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mn>0..</mn><mi>c</mi></mrow><annotation encoding="application/x-tex">0..c</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6444em;"></span><span class="mord">0..</span><span class="mord mathnormal">c</span></span></span></span> and requires rows <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mo>≤</mo><mi>c</mi></mrow><annotation encoding="application/x-tex">\le c</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.7719em;vertical-align:-0.136em;"></span><span class="mrel">≤</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal">c</span></span></span></span> to be bit-for-bit unchanged. A z-score normalized with the full-sample mean and standard deviation fails this at every row, because shortening the sample changes the mean. The <strong>perturbation probe</strong> multiplies every field after <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>c</mi></mrow><annotation encoding="application/x-tex">c</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal">c</span></span></span></span> by random positive noise, recomputes, and requires the same thing. A <code>shift(-1)</code> fails it at exactly row <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>c</mi></mrow><annotation encoding="application/x-tex">c</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal">c</span></span></span></span>. A centered moving average fails both probes in the last half window before the cut.</p><p>Neither probe needs to know what an operator does.</p><h3 id="labels-overlap-so-splits-must-be-purged">labels overlap, so splits must be purged</h3><p>A 5 day forward return label at row <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>t</mi></mrow><annotation encoding="application/x-tex">t</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6151em;"></span><span class="mord mathnormal">t</span></span></span></span> contains the returns of days <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>t</mi><mo>+</mo><mn>2</mn></mrow><annotation encoding="application/x-tex">t+2</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6984em;vertical-align:-0.0833em;"></span><span class="mord mathnormal">t</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">+</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:0.6444em;"></span><span class="mord">2</span></span></span></span> through <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>t</mi><mo>+</mo><mn>6</mn></mrow><annotation encoding="application/x-tex">t+6</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6984em;vertical-align:-0.0833em;"></span><span class="mord mathnormal">t</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">+</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:0.6444em;"></span><span class="mord">6</span></span></span></span> (with delay 1). If a training row sits just before a test block, its label contains returns that are also inside the test block, and the model is partly graded on what it trained on. Purging, from chapter 7 of Lopez de Prado's <em>Advances in Financial Machine Learning</em>, drops every training row whose label interval overlaps a test label interval. Embargo drops a further band of rows after each test block, because features computed right after the block are serially correlated with the test labels.</p><p>In a walk-forward split, training data is always before the test block, so only purging matters. In k-fold, training data sits on both sides, so both matter.</p><h3 id="a-backtest-is-a-recursion-not-a-matrix-product">a backtest is a recursion, not a matrix product</h3><p>The common vectorized backtest computes turnover as <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mo>∑</mo><mi>i</mi></msub><mi mathvariant="normal">∣</mi><msub><mi>W</mi><mrow><mi>t</mi><mo separator="true">,</mo><mi>i</mi></mrow></msub><mo>−</mo><msub><mi>W</mi><mrow><mi>t</mi><mo>−</mo><mn>1</mn><mo separator="true">,</mo><mi>i</mi></mrow></msub><mi mathvariant="normal">∣</mi></mrow><annotation encoding="application/x-tex">\sum_i |W_{t,i} - W_{t-1,i}|</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1.0497em;vertical-align:-0.2997em;"></span><span class="mop"><span class="mop op-symbol small-op" style="position:relative;top:0em;">∑</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.162em;"><span style="top:-2.4003em;margin-left:0em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight">i</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.2997em;"><span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mord">∣</span><span class="mord"><span class="mord mathnormal" style="margin-right:0.1389em;">W</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3117em;"><span style="top:-2.55em;margin-left:-0.1389em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight">t</span><span class="mpunct mtight">,</span><span class="mord mathnormal mtight">i</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.2861em;"><span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">−</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:1.0361em;vertical-align:-0.2861em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.1389em;">W</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3117em;"><span style="top:-2.55em;margin-left:-0.1389em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight">t</span><span class="mbin mtight">−</span><span class="mord mtight">1</span><span class="mpunct mtight">,</span><span class="mord mathnormal mtight">i</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.2861em;"><span></span></span></span></span></span></span><span class="mord">∣</span></span></span></span>. That is wrong whenever prices move. Between rebalances, weights drift with returns, so the trade needed at day <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>t</mi></mrow><annotation encoding="application/x-tex">t</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6151em;"></span><span class="mord mathnormal">t</span></span></span></span> is the difference between today's target and the drifted weights, not between today's and yesterday's targets. And because costs reduce equity, which changes tomorrow's drifted weights, day <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>t</mi></mrow><annotation encoding="application/x-tex">t</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6151em;"></span><span class="mord mathnormal">t</span></span></span></span> depends on day <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>t</mi><mo>−</mo><mn>1</mn></mrow><annotation encoding="application/x-tex">t-1</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6984em;vertical-align:-0.0833em;"></span><span class="mord mathnormal">t</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">−</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:0.6444em;"></span><span class="mord">1</span></span></span></span> through the cost term. The loop over time cannot be vectorized away.</p><p>Per strategy and day, with equity in units of starting capital,</p><p class="katex-block"><span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><msubsup><mi>E</mi><mi>t</mi><mo>−</mo></msubsup><mo>=</mo><msub><mi>A</mi><mrow><mi>t</mi><mo>−</mo><mn>1</mn></mrow></msub><mo>+</mo><munder><mo>∑</mo><mi>i</mi></munder><msub><mi>D</mi><mrow><mi>t</mi><mo>−</mo><mn>1</mn><mo separator="true">,</mo><mi>i</mi></mrow></msub><mtext> </mtext><msub><mi>r</mi><mrow><mi>t</mi><mo separator="true">,</mo><mi>i</mi></mrow></msub><mo separator="true">,</mo><mspace width="2em"/><msub><mi>p</mi><mrow><mi>t</mi><mo separator="true">,</mo><mi>i</mi></mrow></msub><mo>=</mo><mfrac><mrow><msub><mi>D</mi><mrow><mi>t</mi><mo>−</mo><mn>1</mn><mo separator="true">,</mo><mi>i</mi></mrow></msub><mtext> </mtext><mo stretchy="false">(</mo><mn>1</mn><mo>+</mo><msub><mi>r</mi><mrow><mi>t</mi><mo separator="true">,</mo><mi>i</mi></mrow></msub><mo stretchy="false">)</mo></mrow><msubsup><mi>E</mi><mi>t</mi><mo>−</mo></msubsup></mfrac></mrow><annotation encoding="application/x-tex">E^-_t = A_{t-1} + \sum_i D_{t-1,i}\, r_{t,i}, \qquadp_{t,i} = \frac{D_{t-1,i}\,(1 + r_{t,i})}{E^-_t}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1.0683em;vertical-align:-0.247em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0576em;">E</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.8213em;"><span style="top:-2.453em;margin-left:-0.0576em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight">t</span></span></span><span style="top:-3.113em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mbin mtight">−</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.247em;"><span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.8917em;vertical-align:-0.2083em;"></span><span class="mord"><span class="mord mathnormal">A</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3011em;"><span style="top:-2.55em;margin-left:0em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight">t</span><span class="mbin mtight">−</span><span class="mord mtight">1</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.2083em;"><span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">+</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:2.3277em;vertical-align:-1.2777em;"></span><span class="mop op-limits"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:1.05em;"><span style="top:-1.8723em;margin-left:0em;"><span class="pstrut" style="height:3.05em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight">i</span></span></span><span style="top:-3.05em;"><span class="pstrut" style="height:3.05em;"></span><span><span class="mop op-symbol large-op">∑</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:1.2777em;"><span></span></span></span></span></span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0278em;">D</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3117em;"><span style="top:-2.55em;margin-left:-0.0278em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight">t</span><span class="mbin mtight">−</span><span class="mord mtight">1</span><span class="mpunct mtight">,</span><span class="mord mathnormal mtight">i</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.2861em;"><span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0278em;">r</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3117em;"><span style="top:-2.55em;margin-left:-0.0278em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight">t</span><span class="mpunct mtight">,</span><span class="mord mathnormal mtight">i</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.2861em;"><span></span></span></span></span></span></span><span class="mpunct">,</span><span class="mspace" style="margin-right:2em;"></span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mord"><span class="mord mathnormal">p</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3117em;"><span style="top:-2.55em;margin-left:0em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight">t</span><span class="mpunct mtight">,</span><span class="mord mathnormal mtight">i</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.2861em;"><span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:2.3742em;vertical-align:-0.9472em;"></span><span class="mord"><span class="mopen nulldelimiter"></span><span class="mfrac"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:1.427em;"><span style="top:-2.2985em;"><span class="pstrut" style="height:3em;"></span><span class="mord"><span class="mord"><span class="mord mathnormal" style="margin-right:0.0576em;">E</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.8115em;"><span style="top:-2.4542em;margin-left:-0.0576em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight">t</span></span></span><span style="top:-3.1031em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mbin mtight">−</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.2458em;"><span></span></span></span></span></span></span></span></span><span style="top:-3.23em;"><span class="pstrut" style="height:3em;"></span><span class="frac-line" style="border-bottom-width:0.04em;"></span></span><span style="top:-3.677em;"><span class="pstrut" style="height:3em;"></span><span class="mord"><span class="mord"><span class="mord mathnormal" style="margin-right:0.0278em;">D</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3117em;"><span style="top:-2.55em;margin-left:-0.0278em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight">t</span><span class="mbin mtight">−</span><span class="mord mtight">1</span><span class="mpunct mtight">,</span><span class="mord mathnormal mtight">i</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.2861em;"><span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mopen">(</span><span class="mord">1</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">+</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0278em;">r</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3117em;"><span style="top:-2.55em;margin-left:-0.0278em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight">t</span><span class="mpunct mtight">,</span><span class="mord mathnormal mtight">i</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.2861em;"><span></span></span></span></span></span></span><span class="mclose">)</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.9472em;"><span></span></span></span></span></span><span class="mclose nulldelimiter"></span></span></span></span></span></span></p><p class="katex-block"><span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><msub><mi>h</mi><mi>t</mi></msub><mo>=</mo><msub><mi>W</mi><mrow><mi>t</mi><mo>−</mo><mi>d</mi></mrow></msub><mo separator="true">,</mo><mspace width="2em"/><msub><mi>k</mi><mi>t</mi></msub><mo>=</mo><munder><mo>∑</mo><mi>i</mi></munder><mi mathvariant="normal">∣</mi><msub><mi>h</mi><mrow><mi>t</mi><mo separator="true">,</mo><mi>i</mi></mrow></msub><mo>−</mo><msub><mi>p</mi><mrow><mi>t</mi><mo separator="true">,</mo><mi>i</mi></mrow></msub><mi mathvariant="normal">∣</mi><mtext> </mtext><msub><mi>c</mi><mrow><mi>t</mi><mo separator="true">,</mo><mi>i</mi></mrow></msub><mo separator="true">,</mo><mspace width="2em"/><msub><mi>A</mi><mi>t</mi></msub><mo>=</mo><msubsup><mi>E</mi><mi>t</mi><mo>−</mo></msubsup><mtext> </mtext><mo stretchy="false">(</mo><mn>1</mn><mo>−</mo><msub><mi>k</mi><mi>t</mi></msub><mo stretchy="false">)</mo><mo separator="true">,</mo><mspace width="2em"/><msub><mi>D</mi><mi>t</mi></msub><mo>=</mo><msub><mi>h</mi><mi>t</mi></msub><mtext> </mtext><msubsup><mi>E</mi><mi>t</mi><mo>−</mo></msubsup></mrow><annotation encoding="application/x-tex">h_t = W_{t-d}, \qquadk_t = \sum_i |h_{t,i} - p_{t,i}|\, c_{t,i}, \qquadA_t = E^-_t\,(1 - k_t), \qquadD_t = h_t\, E^-_t</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.8444em;vertical-align:-0.15em;"></span><span class="mord"><span class="mord mathnormal">h</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.2806em;"><span style="top:-2.55em;margin-left:0em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight">t</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.9028em;vertical-align:-0.2083em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.1389em;">W</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3361em;"><span style="top:-2.55em;margin-left:-0.1389em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight">t</span><span class="mbin mtight">−</span><span class="mord mathnormal mtight">d</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.2083em;"><span></span></span></span></span></span></span><span class="mpunct">,</span><span class="mspace" style="margin-right:2em;"></span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0315em;">k</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.2806em;"><span style="top:-2.55em;margin-left:-0.0315em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight">t</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:2.3277em;vertical-align:-1.2777em;"></span><span class="mop op-limits"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:1.05em;"><span style="top:-1.8723em;margin-left:0em;"><span class="pstrut" style="height:3.05em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight">i</span></span></span><span style="top:-3.05em;"><span class="pstrut" style="height:3.05em;"></span><span><span class="mop op-symbol large-op">∑</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:1.2777em;"><span></span></span></span></span></span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mord">∣</span><span class="mord"><span class="mord mathnormal">h</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3117em;"><span style="top:-2.55em;margin-left:0em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight">t</span><span class="mpunct mtight">,</span><span class="mord mathnormal mtight">i</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.2861em;"><span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">−</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:1.0361em;vertical-align:-0.2861em;"></span><span class="mord"><span class="mord mathnormal">p</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3117em;"><span style="top:-2.55em;margin-left:0em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight">t</span><span class="mpunct mtight">,</span><span class="mord mathnormal mtight">i</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.2861em;"><span></span></span></span></span></span></span><span class="mord">∣</span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mord"><span class="mord mathnormal">c</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3117em;"><span style="top:-2.55em;margin-left:0em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight">t</span><span class="mpunct mtight">,</span><span class="mord mathnormal mtight">i</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.2861em;"><span></span></span></span></span></span></span><span class="mpunct">,</span><span class="mspace" style="margin-right:2em;"></span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mord"><span class="mord mathnormal">A</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.2806em;"><span style="top:-2.55em;margin-left:0em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight">t</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:1.0713em;vertical-align:-0.25em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0576em;">E</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.8213em;"><span style="top:-2.453em;margin-left:-0.0576em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight">t</span></span></span><span style="top:-3.113em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mbin mtight">−</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.247em;"><span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mopen">(</span><span class="mord">1</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">−</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0315em;">k</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.2806em;"><span style="top:-2.55em;margin-left:-0.0315em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight">t</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span><span class="mclose">)</span><span class="mpunct">,</span><span class="mspace" style="margin-right:2em;"></span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0278em;">D</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.2806em;"><span style="top:-2.55em;margin-left:-0.0278em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight">t</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:1.0683em;vertical-align:-0.247em;"></span><span class="mord"><span class="mord mathnormal">h</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.2806em;"><span style="top:-2.55em;margin-left:0em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight">t</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0576em;">E</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.8213em;"><span style="top:-2.453em;margin-left:-0.0576em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight">t</span></span></span><span style="top:-3.113em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mbin mtight">−</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.247em;"><span></span></span></span></span></span></span></span></span></span></span></p><p><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msup><mi>E</mi><mo>−</mo></msup></mrow><annotation encoding="application/x-tex">E^-</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.7713em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0576em;">E</span><span class="msupsub"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.7713em;"><span style="top:-3.063em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mbin mtight">−</span></span></span></span></span></span></span></span></span></span></span> is pre-trade equity after marking to market, <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>p</mi></mrow><annotation encoding="application/x-tex">p</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.625em;vertical-align:-0.1944em;"></span><span class="mord mathnormal">p</span></span></span></span> the drifted weights, <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>h</mi></mrow><annotation encoding="application/x-tex">h</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6944em;"></span><span class="mord mathnormal">h</span></span></span></span> the executed targets, <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>c</mi></mrow><annotation encoding="application/x-tex">c</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal">c</span></span></span></span> the per-unit cost, <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>A</mi></mrow><annotation encoding="application/x-tex">A</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal">A</span></span></span></span> post-trade equity and <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>D</mi></mrow><annotation encoding="application/x-tex">D</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal" style="margin-right:0.0278em;">D</span></span></span></span> dollar positions. Where an asset has no bar, <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>h</mi><mo>=</mo><mi>p</mi></mrow><annotation encoding="application/x-tex">h = p</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6944em;"></span><span class="mord mathnormal">h</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.625em;vertical-align:-0.1944em;"></span><span class="mord mathnormal">p</span></span></span></span>, so it is held and not traded.</p><p>The convention that matters is that targets are fractions of <em>pre-trade</em> equity and the cost is paid from cash. Under that convention an explicit ledger of share counts and a cash balance produces exactly the same numbers as the weight recursion, to floating point. That is what lets a slow, obviously correct ledger serve as the oracle for the fast kernels. It also guarantees three identities that the tests check directly.</p><p class="katex-block"><span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><msub><mi>A</mi><mi>t</mi></msub><mo>=</mo><msub><mi>A</mi><mrow><mi>t</mi><mo>−</mo><mn>1</mn></mrow></msub><mo>+</mo><msub><mtext>gross</mtext><mi>t</mi></msub><mo>−</mo><msub><mtext>cost</mtext><mi>t</mi></msub><mo separator="true">,</mo><mspace width="2em"/><msub><mtext>gross</mtext><mi>t</mi></msub><mo>=</mo><munder><mo>∑</mo><mi>i</mi></munder><msub><mi>D</mi><mrow><mi>t</mi><mo>−</mo><mn>1</mn><mo separator="true">,</mo><mi>i</mi></mrow></msub><mtext> </mtext><msub><mi>r</mi><mrow><mi>t</mi><mo separator="true">,</mo><mi>i</mi></mrow></msub><mo separator="true">,</mo><mspace width="2em"/><msub><mi>A</mi><mi>t</mi></msub><mo>=</mo><msub><mtext>cash</mtext><mi>t</mi></msub><mo>+</mo><munder><mo>∑</mo><mi>i</mi></munder><msub><mi>D</mi><mrow><mi>t</mi><mo separator="true">,</mo><mi>i</mi></mrow></msub></mrow><annotation encoding="application/x-tex">A_t = A_{t-1} + \text{gross}_t - \text{cost}_t, \qquad\text{gross}_t = \sum_i D_{t-1,i}\, r_{t,i}, \qquadA_t = \text{cash}_t + \sum_i D_{t,i}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.8333em;vertical-align:-0.15em;"></span><span class="mord"><span class="mord mathnormal">A</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.2806em;"><span style="top:-2.55em;margin-left:0em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight">t</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.8917em;vertical-align:-0.2083em;"></span><span class="mord"><span class="mord mathnormal">A</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3011em;"><span style="top:-2.55em;margin-left:0em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight">t</span><span class="mbin mtight">−</span><span class="mord mtight">1</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.2083em;"><span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">+</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:0.8275em;vertical-align:-0.2441em;"></span><span class="mord"><span class="mord text"><span class="mord">gross</span></span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.1864em;"><span style="top:-2.4559em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight">t</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.2441em;"><span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">−</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:0.8592em;vertical-align:-0.2441em;"></span><span class="mord"><span class="mord text"><span class="mord">cost</span></span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.2806em;"><span style="top:-2.55em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight">t</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span><span class="mpunct">,</span><span class="mspace" style="margin-right:2em;"></span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mord"><span class="mord text"><span class="mord">gross</span></span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.1864em;"><span style="top:-2.4559em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight">t</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.2441em;"><span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:2.3277em;vertical-align:-1.2777em;"></span><span class="mop op-limits"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:1.05em;"><span style="top:-1.8723em;margin-left:0em;"><span class="pstrut" style="height:3.05em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight">i</span></span></span><span style="top:-3.05em;"><span class="pstrut" style="height:3.05em;"></span><span><span class="mop op-symbol large-op">∑</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:1.2777em;"><span></span></span></span></span></span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0278em;">D</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3117em;"><span style="top:-2.55em;margin-left:-0.0278em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight">t</span><span class="mbin mtight">−</span><span class="mord mtight">1</span><span class="mpunct mtight">,</span><span class="mord mathnormal mtight">i</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.2861em;"><span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0278em;">r</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3117em;"><span style="top:-2.55em;margin-left:-0.0278em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight">t</span><span class="mpunct mtight">,</span><span class="mord mathnormal mtight">i</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.2861em;"><span></span></span></span></span></span></span><span class="mpunct">,</span><span class="mspace" style="margin-right:2em;"></span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mord"><span class="mord mathnormal">A</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.2806em;"><span style="top:-2.55em;margin-left:0em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight">t</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.8444em;vertical-align:-0.15em;"></span><span class="mord"><span class="mord text"><span class="mord">cash</span></span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.2806em;"><span style="top:-2.55em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight">t</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">+</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:2.3277em;vertical-align:-1.2777em;"></span><span class="mop op-limits"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:1.05em;"><span style="top:-1.8723em;margin-left:0em;"><span class="pstrut" style="height:3.05em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight">i</span></span></span><span style="top:-3.05em;"><span class="pstrut" style="height:3.05em;"></span><span><span class="mop op-symbol large-op">∑</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:1.2777em;"><span></span></span></span></span></span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0278em;">D</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3117em;"><span style="top:-2.55em;margin-left:-0.0278em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight">t</span><span class="mpunct mtight">,</span><span class="mord mathnormal mtight">i</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.2861em;"><span></span></span></span></span></span></span></span></span></span></span></p><h3 id="costs-and-capacity">costs and capacity</h3><p>The per-unit cost is a linear term (exchange fee plus half spread, 10 plus 5 bps by default) and an optional square-root impact term,</p><p class="katex-block"><span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><msub><mi>c</mi><mrow><mi>t</mi><mo separator="true">,</mo><mi>i</mi></mrow></msub><mo>=</mo><mi mathvariant="normal">ℓ</mi><mo>+</mo><mi>η</mi><mtext> </mtext><msub><mi>σ</mi><mrow><mi>t</mi><mo separator="true">,</mo><mi>i</mi></mrow></msub><msqrt><mfrac><mrow><mi mathvariant="normal">∣</mi><msub><mi>h</mi><mrow><mi>t</mi><mo separator="true">,</mo><mi>i</mi></mrow></msub><mo>−</mo><msub><mi>p</mi><mrow><mi>t</mi><mo separator="true">,</mo><mi>i</mi></mrow></msub><mi mathvariant="normal">∣</mi><mtext> </mtext><msubsup><mi>E</mi><mi>t</mi><mo>−</mo></msubsup><mtext> </mtext><mtext>AUM</mtext></mrow><msub><mtext>ADV</mtext><mrow><mi>t</mi><mo separator="true">,</mo><mi>i</mi></mrow></msub></mfrac></msqrt></mrow><annotation encoding="application/x-tex">c_{t,i} = \ell + \eta\, \sigma_{t,i} \sqrt{\frac{|h_{t,i} - p_{t,i}|\, E^-_t\, \text{AUM}}{\text{ADV}_{t,i}}}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.7167em;vertical-align:-0.2861em;"></span><span class="mord"><span class="mord mathnormal">c</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3117em;"><span style="top:-2.55em;margin-left:0em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight">t</span><span class="mpunct mtight">,</span><span class="mord mathnormal mtight">i</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.2861em;"><span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.7778em;vertical-align:-0.0833em;"></span><span class="mord">ℓ</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">+</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:3.04em;vertical-align:-1.1479em;"></span><span class="mord mathnormal" style="margin-right:0.0359em;">η</span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0359em;">σ</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3117em;"><span style="top:-2.55em;margin-left:-0.0359em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight">t</span><span class="mpunct mtight">,</span><span class="mord mathnormal mtight">i</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.2861em;"><span></span></span></span></span></span></span><span class="mord sqrt"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:1.8921em;"><span class="svg-align" style="top:-5em;"><span class="pstrut" style="height:5em;"></span><span class="mord" style="padding-left:1em;"><span class="mord"><span class="mopen nulldelimiter"></span><span class="mfrac"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:1.4885em;"><span style="top:-2.314em;"><span class="pstrut" style="height:3em;"></span><span class="mord"><span class="mord"><span class="mord text"><span class="mord">ADV</span></span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3117em;"><span style="top:-2.55em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight">t</span><span class="mpunct mtight">,</span><span class="mord mathnormal mtight">i</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.2861em;"><span></span></span></span></span></span></span></span></span><span style="top:-3.23em;"><span class="pstrut" style="height:3em;"></span><span class="frac-line" style="border-bottom-width:0.04em;"></span></span><span style="top:-3.677em;"><span class="pstrut" style="height:3em;"></span><span class="mord"><span class="mord">∣</span><span class="mord"><span class="mord mathnormal">h</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3117em;"><span style="top:-2.55em;margin-left:0em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight">t</span><span class="mpunct mtight">,</span><span class="mord mathnormal mtight">i</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.2861em;"><span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">−</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mord"><span class="mord mathnormal">p</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3117em;"><span style="top:-2.55em;margin-left:0em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight">t</span><span class="mpunct mtight">,</span><span class="mord mathnormal mtight">i</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.2861em;"><span></span></span></span></span></span></span><span class="mord">∣</span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0576em;">E</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.8115em;"><span style="top:-2.4542em;margin-left:-0.0576em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight">t</span></span></span><span style="top:-3.1031em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mbin mtight">−</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.2458em;"><span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mord text"><span class="mord">AUM</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.9721em;"><span></span></span></span></span></span><span class="mclose nulldelimiter"></span></span></span></span><span style="top:-3.8521em;"><span class="pstrut" style="height:5em;"></span><span class="hide-tail" style="min-width:1.02em;height:3.08em;"><svg xmlns="http://www.w3.org/2000/svg" width="400em" height="3.08em" viewBox="0 0 400000 3240" preserveAspectRatio="xMinYMin slice"><path d="M473,2793c339.3,-1799.3,509.3,-2700,510,-2702 l0 -0c3.3,-7.3,9.3,-11,18,-11 H400000v40H1017.7s-90.5,478,-276.2,1466c-185.7,988,-279.5,1483,-281.5,1485c-2,6,-10,9,-24,9c-8,0,-12,-0.7,-12,-2c0,-1.3,-5.3,-32,-16,-92c-50.7,-293.3,-119.7,-693.3,-207,-1200c0,-1.3,-5.3,8.7,-16,30c-10.7,21.3,-21.3,42.7,-32,64s-16,33,-16,33s-26,-26,-26,-26s76,-153,76,-153s77,-151,77,-151c0.7,0.7,35.7,202,105,604c67.3,400.7,102,602.7,104,606zM1001 80h400000v40H1017.7z"/></svg></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:1.1479em;"><span></span></span></span></span></span></span></span></span></span></p><p>where <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>σ</mi></mrow><annotation encoding="application/x-tex">\sigma</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal" style="margin-right:0.0359em;">σ</span></span></span></span> is trailing daily volatility and ADV trailing dollar volume, both causal. Impact makes capacity visible, since the same weights lose more per trade as the book grows.</p><h3 id="point-in-time-data">point-in-time data</h3><p>Every stored bar carries two timestamps, <code>available_at</code> (the bar's close plus 1 ms, the earliest moment its values could be known) and <code>ingested_at</code> (when the store learned them). A read as of instant <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>X</mi></mrow><annotation encoding="application/x-tex">X</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal" style="margin-right:0.0785em;">X</span></span></span></span> keeps, for each (symbol, date), the latest version with both timestamps <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mo>≤</mo><mi>X</mi></mrow><annotation encoding="application/x-tex">\le X</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.7719em;vertical-align:-0.136em;"></span><span class="mrel">≤</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal" style="margin-right:0.0785em;">X</span></span></span></span>. That is enough to reproduce any past read exactly, including across restatements, and it means a run can record precisely which data it saw.</p><h2 id="architecture">architecture</h2><p>The pipeline is a straight line of stages, each with one data structure in and one out.</p><figure data-figure="diagram:backtester-pipeline"></figure><p>The one deliberate split is between storage and compute. Storage is long parquet through polars, which is good at partition pruning and the point-in-time group-by. Compute is wide numpy arrays of shape (dates, symbols) with NaN meaning &quot;no bar&quot;, which is good at rolling and cross-sectional operations over a whole block. Nothing is forward filled at the panel layer, so <code>np.isfinite(close)</code> stays exactly &quot;this asset traded today&quot;. Partitions are by year, not symbol, which gives five files per ingest instead of 1,500 tiny ones.</p><h2 id="implementation">implementation</h2><h3 id="the-store">the store</h3><p><code>BarStore.ingest</code> validates (no duplicate keys, high at least low, positive close), stamps <code>ingested_at</code> and an ingest id, and writes one parquet file per year partition plus a row in an ingest log. Nothing is ever overwritten.</p><figure class="highlight python"><table><tr><td class="code"><pre><span class="line"><span class="keyword">if</span> as_of <span class="keyword">is</span> <span class="keyword">not</span> <span class="literal">None</span>:</span><br><span class="line">    lf = lf.<span class="built_in">filter</span>((pl.col(<span class="string">&quot;ingested_at&quot;</span>) &lt;= as_of) &amp; (pl.col(<span class="string">&quot;available_at&quot;</span>) &lt;= as_of))</span><br><span class="line"><span class="comment"># Latest known version of each bar wins (restatements).</span></span><br><span class="line">lf = (lf.sort(<span class="string">&quot;ingested_at&quot;</span>)</span><br><span class="line">        .group_by([<span class="string">&quot;symbol&quot;</span>, <span class="string">&quot;date&quot;</span>]).last()</span><br><span class="line">        .drop(<span class="string">&quot;year&quot;</span>)</span><br><span class="line">        .sort([<span class="string">&quot;symbol&quot;</span>, <span class="string">&quot;date&quot;</span>]))</span><br></pre></td></tr></table></figure><p><code>snapshot_id(as_of)</code> hashes the ingest ids visible at that instant, and every run records it. All experiments below ran on snapshot <code>417cbe2508ff</code>, which is one ingest of 466,274 bars across 300 symbols, 83 of which stopped trading before the end of the window.</p><h3 id="features-without-pandas">features without pandas</h3><p>Rolling windows are the core of almost every feature, and the usual pandas approach works one column at a time and needs care around missing values. I used masked cumulative sums over the whole (T, N) block instead. Replace NaN with zero, cumulatively sum the values, their squares and a finite mask, and difference at lag <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>w</mi></mrow><annotation encoding="application/x-tex">w</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal" style="margin-right:0.0269em;">w</span></span></span></span>.</p><figure class="highlight python"><table><tr><td class="code"><pre><span class="line">finite = np.isfinite(x)</span><br><span class="line">xz = np.where(finite, x, <span class="number">0.0</span>)</span><br><span class="line">c1 = np.concatenate([pad, np.cumsum(xz, axis=<span class="number">0</span>)])</span><br><span class="line">cn = np.concatenate([pad, np.cumsum(finite, axis=<span class="number">0</span>)])</span><br><span class="line">s[w - <span class="number">1</span>:] = c1[w:] - c1[:-w]</span><br><span class="line">n[w - <span class="number">1</span>:] = cn[w:] - cn[:-w]</span><br><span class="line"><span class="comment"># a window is valid only if all w values were finite</span></span><br><span class="line">mean = np.where(n == w, s / w, np.nan)</span><br></pre></td></tr></table></figure><p>A window with any missing bar is NaN rather than silently averaged over fewer values. Cross-sectional rank is a per-row argsort scaled to <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mo stretchy="false">[</mo><mo>−</mo><mn>0.5</mn><mo separator="true">,</mo><mn>0.5</mn><mo stretchy="false">]</mo></mrow><annotation encoding="application/x-tex">[-0.5, 0.5]</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mopen">[</span><span class="mord">−</span><span class="mord">0.5</span><span class="mpunct">,</span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mord">0.5</span><span class="mclose">]</span></span></span></span>. Specs come from TOML tables, inputs can be panel fields or other features, and the specs are topologically sorted so declaration order does not matter. There are nine causal operators and three deliberately leaky ones (a forward return, a centered moving average, a full-sample z-score) in a separate registry, so the audit has something to catch in tests and in the demo config.</p><h3 id="the-leakage-audit">the leakage audit</h3><p>The audit runs both probes at four cut points, treating NaN as equal to NaN.</p><figure class="highlight python"><table><tr><td class="code"><pre><span class="line"><span class="keyword">for</span> c <span class="keyword">in</span> cuts:</span><br><span class="line">    trunc = compute_features(panel.slice_time(c + <span class="number">1</span>), specs)</span><br><span class="line">    pert_panel = panel.copy()</span><br><span class="line">    <span class="keyword">for</span> k, v <span class="keyword">in</span> pert_panel.fields.items():</span><br><span class="line">        v[c + <span class="number">1</span>:] *= np.exp(rng.normal(<span class="number">0</span>, <span class="number">0.5</span>, v[c + <span class="number">1</span>:].shape))</span><br><span class="line">    pert = compute_features(pert_panel, specs)</span><br><span class="line">    <span class="keyword">for</span> s <span class="keyword">in</span> specs:</span><br><span class="line">        record(s.name, _same(trunc[s.name], base[s.name][: c + <span class="number">1</span>]).<span class="built_in">all</span>(axis=<span class="number">1</span>), c, <span class="string">&quot;truncation&quot;</span>)</span><br><span class="line">        record(s.name, _same(pert[s.name][: c + <span class="number">1</span>], base[s.name][: c + <span class="number">1</span>]).<span class="built_in">all</span>(axis=<span class="number">1</span>), c,</span><br><span class="line">               <span class="string">&quot;perturbation&quot;</span>)</span><br></pre></td></tr></table></figure><p>It costs nine feature computations per run (the base plus two probes at four cuts). <code>run_research</code> calls it first, on the exact spec the run uses, and raises <code>LookAheadError</code> if anything leaks. The failed run is still written to the tracker with status &quot;failed&quot;, so a refused experiment leaves a trace.</p><h3 id="three-engines-for-one-formula">three engines for one formula</h3><p>The backtester has three kernels implementing the recursion above. The reference kernel is plain Python holding a cash balance and a share count per asset. It is slow on purpose. The numpy kernel loops over days and vectorizes over (strategies, assets). The numba kernel is a compiled triple loop, parallel over strategies.</p><figure class="highlight python"><table><tr><td class="code"><pre><span class="line"><span class="keyword">for</span> s <span class="keyword">in</span> numba.prange(S):</span><br><span class="line">    A = <span class="number">1.0</span></span><br><span class="line">    D = np.zeros(N)</span><br><span class="line">    <span class="keyword">for</span> t <span class="keyword">in</span> <span class="built_in">range</span>(T):</span><br><span class="line">        gpnl = <span class="number">0.0</span></span><br><span class="line">        <span class="keyword">for</span> i <span class="keyword">in</span> <span class="built_in">range</span>(N):</span><br><span class="line">            gpnl += D[i] * r[t, i]</span><br><span class="line">        Em = A + gpnl</span><br><span class="line">        ...</span><br><span class="line">        <span class="keyword">for</span> i <span class="keyword">in</span> <span class="built_in">range</span>(N):</span><br><span class="line">            p = D[i] * (<span class="number">1.0</span> + r[t, i]) / Es</span><br><span class="line">            hi = H[s, t, i] <span class="keyword">if</span> tradeable[t, i] <span class="keyword">else</span> p</span><br><span class="line">            dw = <span class="built_in">abs</span>(hi - p)</span><br><span class="line">            unit = lin</span><br><span class="line">            <span class="keyword">if</span> impact &gt; <span class="number">0.0</span>:</span><br><span class="line">                unit = lin + impact * sigma[t, i] * np.sqrt(dw * Es * aum / adv[t, i])</span><br><span class="line">            k += dw * unit</span><br></pre></td></tr></table></figure><p>The test holding this together runs all three engines on a synthetic panel with 3% missing bars, four random strategies, three cost models (zero, linear, linear plus impact) and delays 0, 1 and 2, and requires equity, gross PnL, cost, turnover, exposures and full dollar positions to agree with the reference to a relative tolerance of 1e-10. Missing bars earn a zero return (price carried forward) and cannot be traded. If pre-trade equity reaches zero, the strategy is marked bankrupt and holds nothing afterwards.</p><p>Execution delay is a shift of the target array before the kernel, and a negative delay raises, since it would mean trading before the signal exists. A separate test changes the targets from some row <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>c</mi></mrow><annotation encoding="application/x-tex">c</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal">c</span></span></span></span> onward and checks that equity before row <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>c</mi><mo>+</mo><mi>d</mi></mrow><annotation encoding="application/x-tex">c + d</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6667em;vertical-align:-0.0833em;"></span><span class="mord mathnormal">c</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">+</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:0.6944em;"></span><span class="mord mathnormal">d</span></span></span></span> is unchanged. That is the backtester's own look-ahead test.</p><h3 id="purged-splits">purged splits</h3><p>The purge and embargo logic is a handful of boolean masks over row indices.</p><figure class="highlight python"><table><tr><td class="code"><pre><span class="line"><span class="keyword">if</span> purge:</span><br><span class="line">    <span class="comment"># Rows before the block whose labels reach into it.</span></span><br><span class="line">    train_mask &amp;= ~((rows &lt; a) &amp; (rows + label_horizon &gt;= a))</span><br><span class="line">    <span class="comment"># Rows after the block that start inside the test labels&#x27; span.</span></span><br><span class="line">    train_mask &amp;= ~((rows &gt; b) &amp; (rows &lt;= b + label_horizon))</span><br><span class="line"><span class="keyword">if</span> embargo &gt; <span class="number">0</span>:</span><br><span class="line">    h = label_horizon <span class="keyword">if</span> purge <span class="keyword">else</span> <span class="number">0</span></span><br><span class="line">    train_mask &amp;= ~((rows &gt; b) &amp; (rows &lt;= b + h + embargo))</span><br></pre></td></tr></table></figure><p>Here <code>label_horizon</code> is <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>d</mi><mo>+</mo><mi>h</mi></mrow><annotation encoding="application/x-tex">d + h</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.7778em;vertical-align:-0.0833em;"></span><span class="mord mathnormal">d</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">+</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:0.6944em;"></span><span class="mord mathnormal">h</span></span></span></span>. A separate <code>check_split</code> verifies any split against the contract independently of how it was built, and the tests run it over both split kinds with several horizons and embargoes.</p><h3 id="model-portfolio-and-tracker">model, portfolio and tracker</h3><p>The model is closed-form ridge with an unpenalized intercept, fitted on pooled (date, symbol) rows of cross-sectionally ranked features, predicting the cross-sectional rank of the forward return. Out of sample scores become dollar-neutral weights proportional to the demeaned score rank, scaled to gross exposure 1, and optionally held for <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>k</mi></mrow><annotation encoding="application/x-tex">k</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6944em;"></span><span class="mord mathnormal" style="margin-right:0.0315em;">k</span></span></span></span> days. It is a deliberately plain model. A plain model makes every downstream effect easy to attribute.</p><p>A tracked run is a directory named by timestamp, run name and config hash, holding the fully resolved config, metadata (git sha, library versions, data snapshot, duration, status), metrics and artifacts. <code>qrp runs compare a b</code> prints metrics side by side and only the config keys that differ, which is the view I actually wanted every time I compared two runs.</p><h2 id="problems">problems</h2><p>These are the real bugs and surprises, in the order I hit them. Two of them were in my own tests, and the most instructive one was in the data.</p><h3 id="the-obvious-data-sources-were-blocked">the obvious data sources were blocked</h3><p><code>api.binance.com</code> returned HTTP 451, the US geo-block, and Yahoo's chart API returned HTTP 429 on the first request. That is why v0 has no US equities. The free equity sources also lack delisted names, which would have made survivorship bias unavoidable. I switched to <code>data.binance.vision</code>, Binance's public S3 bulk archive, which serves one zip per symbol and month, has no geo-block, and includes pairs that were later delisted. The ingest lists every USDT pair through the paginated S3 XML listing, drops leveraged tokens and stablecoin pairs, and keeps the 300 with the most archived months in the window.</p><h3 id="non-ascii-tickers-crashed-the-first-ingest">non-ASCII tickers crashed the first ingest</h3><p>The S3 listing contains a few symbols with Chinese names, and <code>urllib</code> raised a <code>UnicodeEncodeError</code> when building their URLs. I now keep only ASCII alphanumeric symbols, which are the only ones in scope for this universe anyway.</p><h3 id="binance-changed-its-timestamp-unit">binance changed its timestamp unit</h3><p>Spot archive files from 2025-01-01 onward use microseconds instead of milliseconds. Read naively, every 2025 and 2026 bar would land somewhere around the year 57000, and the panel would have silently lost its last 20 months. The fix is a threshold, since no millisecond epoch before the year 5000 exceeds <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msup><mn>10</mn><mn>14</mn></msup></mrow><annotation encoding="application/x-tex">10^{14}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.8141em;"></span><span class="mord">1</span><span class="mord"><span class="mord">0</span><span class="msupsub"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.8141em;"><span style="top:-3.063em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mtight">14</span></span></span></span></span></span></span></span></span></span></span></span>.</p><figure class="highlight python"><table><tr><td class="code"><pre><span class="line"><span class="keyword">def</span> <span class="title function_">_to_millis</span>(<span class="params">ts: np.ndarray</span>) -&gt; np.ndarray:</span><br><span class="line">    ts = ts.astype(np.int64)</span><br><span class="line">    <span class="keyword">return</span> np.where(ts &gt; <span class="number">10</span>**<span class="number">14</span>, ts // <span class="number">1000</span>, ts)</span><br></pre></td></tr></table></figure><p>A test pins the conversion for a 2025-01-01 timestamp in both units.</p><h3 id="my-drift-test-was-wrong-and-the-engine-was-right">my drift test was wrong and the engine was right</h3><p>I wrote a test expecting that targets equal to the buy-and-hold drift weights would need zero turnover after the initial purchase. Without costs, that held. With 15 bps costs, day 2 still traded 0.15% of equity, and I assumed an engine bug.</p><p>It was not. The entry cost is paid from cash, so after the first trade the book holds assets worth 100% of pre-trade equity against slightly less than 100% of post-trade equity. That is a small leverage, equal to the fee. A target of 100% invested therefore has to sell a little the next day to get back to 100%, and that sale pays its own fee, which requires a smaller correction the day after, decaying geometrically. The reference ledger and the vectorized kernels agreed on this exactly, which is what convinced me the test was wrong rather than both engines. The test now asserts zero turnover without costs, about 15 bps on day 2 with costs, and a total that stays below twice that.</p><p>A second hand-built test had the position on a halted day as 0.55 when it is 0.525, because day 1 had already rebalanced back to 50% of an equity of 1.05. Writing expected values by hand for tiny cases was still worth it, since that is where the delevering effect surfaced.</p><h3 id="the-luna-relisting-blew-up-a-backtest">the luna relisting blew up a backtest</h3><p>The first purged k-fold run went to exactly zero equity on 2022-05-31, with a daily turnover above 2, which should be impossible for a book with gross exposure 1.</p><p>The data explains it. LUNAUSDT closed at $1.08 on 2022-05-11, $0.00032 on 2022-05-12 and $0.00005 on 2022-05-13. It then had no bars for 18 days, and came back on 2022-05-31 at $8.87 as LUNA 2.0, a different asset under the same ticker. The backtester forward-fills prices through missing bars (correctly, for a halted asset), so a short LUNA position that was stuck because the asset was untradeable saw a one-day return of about 177,000x. The walk-forward runs never noticed, because their test periods start in 2023.</p><p>Eight symbols in the data have gaps longer than three days.</p><table><thead><tr><th>Symbol</th><th>Gap in days</th><th>Resumes</th></tr></thead><tbody><tr><td>FTTUSDT</td><td>311</td><td>2023-09-22</td></tr><tr><td>CVCUSDT</td><td>154</td><td>2023-05-12</td></tr><tr><td>KEYUSDT</td><td>28</td><td>2023-03-10</td></tr><tr><td>LUNAUSDT</td><td>18</td><td>2022-05-31</td></tr><tr><td>VIDTUSDT</td><td>9</td><td>2022-11-09</td></tr><tr><td>STRAXUSDT</td><td>8</td><td>2024-03-28</td></tr><tr><td>BNXUSDT</td><td>6</td><td>2023-02-22</td></tr><tr><td>QUICKUSDT</td><td>4</td><td>2023-07-21</td></tr></tbody></table><p>The v0 fix is <code>split_relisted</code>, which makes any gap over three days start a new instrument.</p><figure class="highlight python"><table><tr><td class="code"><pre><span class="line">gap = pl.col(<span class="string">&quot;date&quot;</span>).diff().over(<span class="string">&quot;symbol&quot;</span>).dt.total_days()</span><br><span class="line">seg = (gap &gt; max_gap_days).fill_null(<span class="literal">False</span>).cast(pl.Int32).cum_sum().over(<span class="string">&quot;symbol&quot;</span>)</span><br></pre></td></tr></table></figure><p>The old LUNAUSDT is now delisted at its last real price and <code>LUNAUSDT#2</code> starts with no history. That took the panel from 300 to 308 instruments. A regression test builds a tiny LUNA-shaped series, shows a held short going bankrupt on the raw data, and shows the same short surviving after the split. It is a blunt rule. A redenomination without a trading pause would still produce a fake return, and a real corporate actions table is on the list below.</p><h3 id="a-high-ic-with-a-low-sharpe-looked-suspicious">a high IC with a low Sharpe looked suspicious</h3><p>The baseline ridge had an out of sample information coefficient (mean cross-sectional rank correlation between score and realized return) of 0.125. That is high for daily crypto. Meanwhile every single feature I would have guessed as the driver had an IC near zero or negative on its own. Shorting yesterday's winners had an IC of 0.000, and 30 day momentum had an IC of -0.049.</p><p>After the LUNA bug I did not trust any number better than I expected, so I added a random-score negative control. It came back with an IC of 0.0005 and a Sharpe of -1.34, the signature of trading noise through a cost model, so the IC calculation itself was not inflated. Then I computed single-feature ICs by hand. The driver is 30 day volatility, with a clearly negative IC, meaning high-volatility coins underperformed. I did not save that single-feature number to a results file, so I do not quote it here, and saving it is on the list below. The Sharpe is much lower than the IC suggests because a long-short crypto book carries large idiosyncratic volatility and a weekly holding period uses only part of a 5 day signal.</p><h3 id="the-machine-was-shared">the machine was shared</h3><p>Ten or more other jobs shared the 14-core machine for the whole session, with the load average between about 220 and 400. The experiment suite hit a 10 minute tool timeout on its first attempt and was rerun in the background, where it took 891 seconds for 17 runs. The benchmark recorded a load average of 360, 351 and 311 before it started and 231, 286 and 292 after. Every throughput number below is a lower bound, and the multithreaded ones are noisy enough that at some sizes 14 threads were slower than one.</p><h3 id="the-forward-fill-cost-more-than-the-kernel">the forward fill cost more than the kernel</h3><p>A quick profile of one strategy on 300 assets, which I did not save to the results directory, showed <code>prepare_market</code> (forward fill plus returns) taking about 18 ms of a 23 ms call under load. The compiled kernel was the cheap part. Rather than micro-optimize the forward fill, I made <code>run_backtest</code> accept a precomputed market, which is the right shape for sweeping many strategies over one panel anyway, and added a test that the result is identical. The benchmark rows labelled &quot;precomputed&quot; measure that path.</p><h2 id="experiments">experiments</h2><p>All experiments ran on snapshot <code>417cbe2508ff</code>, a panel of 1,704 daily rows from 2022-01-01 to 2026-08-31 by 308 instruments (300 symbols before the relisting split, 308 instruments after it). The panel loads in 2.88 s.</p><h3 id="the-base-configuration">the base configuration</h3><ul><li>Universe. The top 100 instruments by trailing 30 day mean dollar volume with at least 60 bars of history, recomputed every day.</li><li>Features. Cross-sectional ranks of the 1 day return, 7, 30 and 90 day momentum, 30 day volatility, 20 day moving average ratio and 20 day volume z-score.</li><li>Model. Ridge with alpha 10, predicting the rank of the 5 day forward return.</li><li>CV. Walk-forward with 8 test blocks after a 365 day minimum training period, purged on the label interval, 5 row embargo.</li><li>Portfolio. Dollar neutral, weights proportional to the demeaned score rank, gross 1, rebalanced every 5 days, executed with a 1 day delay.</li><li>Costs. 10 bps fee plus 5 bps slippage per unit traded.</li></ul><p>Metrics use only out of sample days, 1,338 of them from 2023-01-02 (the first test row, 2023-01-01, plus the 1 day delay).</p><h3 id="the-variants">the variants</h3><p>Each variant changes one thing through <code>--set</code> overrides and is its own tracked run.</p><ol><li>Costs. 0, 5, 15 (baseline), 25 and 50 bps one way, plus square-root impact with coefficient 0.5 at a $50M book.</li><li>Rebalance frequency. Daily instead of every 5 days.</li><li>Execution timing. Delay 0 against delay 1.</li><li>Controls. A random score, and two single-signal baselines with no model (short 1 day winners, long 30 day momentum).</li><li>Leakage. One leaky feature added to the model with the audit's fail switch turned off, either a 5 day forward return or an 11 day centered moving average ratio. Separately, <code>qrp audit</code> on a demo config with three honest and three leaky features, on the real data.</li><li>Purging. A 20 day label horizon, so that overlap is large, under walk-forward and k-fold, each with purge plus a 10 row embargo against neither.</li><li>Throughput. Synthetic panels with T = 1,825 days, N in {100, 300, 1000} assets and S in {1, 16, 128} strategies per call, for the numpy kernel and the numba kernel on 1 and 14 threads, plus the reference engine at a small size, the numba kernel with a precomputed market, and the real 1,704 by 300 panel. Timings are best of 5 (best of 1 for the largest numpy cases).</li></ol><h2 id="results">results</h2><p>An independent reviewer later reran the full experiment suite from a clean download and matched all 17 rows of <code>results/experiments.csv</code> exactly, and reran the leakage audit with identical output. The benchmark was not rerun, so the throughput numbers rest on the original measurement only. Two corrections from that review are applied here (the out of sample start date, and the direction of the 300 to 308 instrument split).</p><p>One caveat applies to every Sharpe below. All 17 experiments share the same 2023 to 2026-08 out of sample period, and I looked at it while debugging the IC. These are comparisons between variants on one platform, not one-shot holdout results.</p><h3 id="leakage">leakage</h3><p>The audit on the demo config flagged exactly the three leaky features on real data and passed the three honest ones. The forward return and the centered average were both first caught at row 422 against a cut at row 426, which is the correct answer, since a 5 day lead and a 5 bar half window both reach row 427 from row 422. The full-sample z-score was caught at row 30, the first row where its input exists, because a full-sample statistic contaminates every row. The command exits 1.</p><p>What the audit prevents is visible when it is switched off.</p><figure data-figure="chart:projects/quant-research-platform/leakage-sharpe"></figure><p>The 5 day forward return turns a Sharpe of 0.49 into 19.40, with an IC of 0.724, an annual return of 9,161% and a max drawdown of 2.5%. The centered average is the more instructive case. It is not an obvious bug. It is what <code>rolling(center=True)</code> does in pandas, it only sees five bars ahead through a smoothing window, and it still produces a Sharpe of 13.33, an IC of 0.438 and a max drawdown of 3.4%, ending at about 21,900 times its starting equity. Nobody would believe 19. Somebody might believe a subtler leak that added 0.3 to a Sharpe of 0.5. The audit would flag that one with the same two probes, as long as the leak lives in a feature and not in the labels or the backtester, which have their own tests.</p><h3 id="baseline">baseline</h3><p>Out of sample, the baseline earned a Sharpe of 0.490, an annual return of 7.3% at 17.6% annual volatility, a max drawdown of 29.7% and an IC of 0.125. Average daily turnover was 15.9% of equity and the annual cost drag was 8.7%. Over 1,338 days the equity grew 29.5%.</p><h3 id="costs">costs</h3><figure data-figure="chart:projects/quant-research-platform/cost-sensitivity"></figure><p>Turnover is identical across the cost runs, so Sharpe falls almost linearly with cost, from 0.988 at zero to 0.822 at 5 bps, 0.490 at 15, 0.159 at 25 and -0.659 at 50. Interpolating between the last two points puts break-even near 30 bps one way. At zero cost the same signal has a Sharpe of 0.988 and a max drawdown of 20.9% instead of 29.7%.</p><p>Adding square-root impact at a $50M book gives a Sharpe of -1.078 and a cost drag of 37.0% a year, so this signal has very little capacity. Rebalancing daily instead of weekly doubled turnover from 15.9% to 31.6% a day and cut the Sharpe from 0.490 to 0.122. The fresher signal did not earn back the extra trading.</p><h3 id="timing">timing</h3><p>Trading at the same close the signal observed (delay 0) raised the Sharpe from 0.490 to 0.692 and the IC from 0.125 to 0.133. One day of execution optimism overstates the Sharpe by about 41%. That is much harder to spot than a leaky feature, because 0.69 is a believable number.</p><h3 id="controls">controls</h3><p>The random score had an IC of 0.0005 and a Sharpe of -1.339, with a cost drag of 16.8% a year from 30.6% daily turnover. Shorting 1 day winners alone had an IC of 0.000 and a Sharpe of -0.848. Long 30 day momentum alone had an IC of -0.049 and a Sharpe of -0.324. Neither textbook signal works in this universe and period, which is why the baseline's edge needed explaining.</p><h3 id="purging">purging</h3><figure data-figure="chart:projects/quant-research-platform/purging"></figure><p>This is a real negative result. With a 20 day horizon, walk-forward gave a Sharpe of 0.842 purged against 0.847 unpurged, and k-fold gave 0.850 against 0.844. The ICs differ by less than 0.002, and the sign of the difference flips between the two split kinds.</p><p>The explanation is capacity. A seven-coefficient ridge trained on tens of thousands of pooled rows cannot memorize the 21 rows at each block boundary whose labels overlap the test block, so the leak that purging removes is too small to register. Purging stays on by default, because project 09 runs higher-capacity models on this pipeline. But this experiment does not show purging doing useful work, only that it is correct by test.</p><h3 id="throughput">throughput</h3><figure data-figure="chart:projects/quant-research-platform/throughput"></figure><p>Measured on an Apple M4 Pro with 14 cores, Python 3.12.13, numpy 2.5.3 and numba 0.67.0, with the load average between 230 and 360 during the run.</p><p>The pure-Python reference engine ran 1,386 strategy-days per second at N = 100, or 139 thousand asset-days per second. With the market precomputed, the single-threaded numba kernel ran 772 thousand strategy-days per second at N = 100, 270 thousand at N = 300 and 79 thousand at N = 1000. In asset-days that is a flat 77 to 81 million per second across a tenfold change in width, which says the kernel does constant work per cell, and it is 557 to 585 times the reference engine per asset-day. That is the price of an oracle, and it only runs in tests.</p><p>The best batched number was 2.12 million strategy-days per second, on 14 threads with S = 16 and N = 100. On the real 1,704 by 300 panel, one strategy took 10.8 ms and 64 random strategies took 0.74 s in one call, 147 thousand strategy-days per second. numpy was between 1.2 times (S = 1, N = 1000) and 20 times (S = 1, N = 100) slower than single-threaded numba, since its per-day overhead only amortizes over wide rows.</p><p>Throughput falls at the largest size. At S = 128 and N = 1000, single-threaded numba dropped to 43 thousand strategy-days per second, and the 14-thread run was slower still at 26.5 thousand. The (S, T, N) target array there is about 1.9 GB, and <code>run_backtest</code> copies it twice before the kernel starts (to lag it and to clean NaNs), so the call is dominated by memory traffic. With other jobs also contending for memory bandwidth, extra threads made it worse.</p><h3 id="tests">tests</h3><p>All 73 tests pass, on synthetic data. They cover look-ahead detection (including a one-bar peek and a leak passed through the DAG), accounting identities, cross-engine agreement, closed-form cases, purging and embargo, point-in-time reads across restatements, the relisting split, and the pipeline end to end.</p><h2 id="what-i-would-change">what i would change</h2><h3 id="lag-and-clean-the-targets-inside-the-kernel">lag and clean the targets inside the kernel</h3><p>Materializing two copies of the (S, T, N) array is the reason throughput collapses at S = 128 and N = 1000. The kernel can read <code>W[s, t - d, i]</code> directly and treat NaN as zero on the fly.</p><h3 id="rerun-the-benchmark-on-an-idle-machine">rerun the benchmark on an idle machine</h3><p>Everything here was measured at a load average between 220 and 400. I would not draw conclusions from the 14-thread rows at all.</p><h3 id="replace-the-gap-rule-with-real-reference-data">replace the gap rule with real reference data</h3><p>A symbol mapping and corporate actions table would handle ticker reuse and redenominations properly, and delisted positions should be closed at a delisting price instead of being held at the last price forever.</p><h3 id="build-the-universe-point-in-time">build the universe point in time</h3><p>The 300 pairs were chosen as the ones with the most archived months in the window, which uses today's knowledge. The daily universe filter inside that set is causal, but the set itself carries mild selection bias. Point-in-time exchange listings would remove it.</p><h3 id="show-purging-mattering">show purging mattering</h3><p>A gradient-boosted model or a much smaller sample should make the effect visible, which is the experiment to add once project 09 brings nonlinear models.</p><h3 id="save-single-feature-ics-with-every-run">save single-feature ICs with every run</h3><p>They explained the baseline faster than anything else, and saving them would have let me quote the volatility IC here.</p><h3 id="make-the-audit-cheaper">make the audit cheaper</h3><p>It recomputes every feature nine times, which is fine for 14 features on daily bars and will not be for minute bars. Caching the unchanged prefix would remove most of the truncation probe's work.</p><h2 id="reproducibility">reproducibility</h2><p>Everything runs from the project directory with <a href="https://docs.astral.sh/uv/">uv</a> and Python 3.12. The data is about 13 MB of parquet from roughly 15,400 small zip files, and is gitignored.</p><figure class="highlight bash"><table><tr><td class="code"><pre><span class="line"><span class="built_in">cd</span> projects/08-leakage-checked-backtester</span><br><span class="line">uv <span class="built_in">sync</span></span><br><span class="line">./scripts/download_data.sh                          <span class="comment"># same as: uv run qrp ingest --root data</span></span><br><span class="line">uv run qrp info                                     <span class="comment"># ingest log and symbol metadata</span></span><br><span class="line">uv run pytest                                       <span class="comment"># 73 tests, synthetic data only</span></span><br><span class="line"></span><br><span class="line">uv run qrp audit configs/leaky_demo.toml            <span class="comment"># flags the 3 leaky features, exits 1</span></span><br><span class="line">uv run qrp run configs/xs_ridge.toml                <span class="comment"># tracked baseline run</span></span><br><span class="line">uv run qrp run configs/xs_ridge.toml --<span class="built_in">set</span> costs.fee_bps=0 --<span class="built_in">set</span> portfolio.rebalance_every=1</span><br><span class="line">uv run qrp runs list --<span class="built_in">sort</span> sharpe</span><br><span class="line">uv run qrp runs compare &lt;run_id_a&gt; &lt;run_id_b&gt;</span><br><span class="line"></span><br><span class="line">uv run python scripts/run_experiments.py            <span class="comment"># results/experiments.&#123;csv,json&#125;, plots</span></span><br><span class="line">uv run python scripts/bench_backtest.py --real      <span class="comment"># results/bench_backtest.&#123;csv,json,png&#125;</span></span><br></pre></td></tr></table></figure><p>A fresh download gets a new ingest id and therefore a different snapshot id from <code>417cbe2508ff</code> (the reviewer's was <code>f2322f2cf2d1</code>), but the same 466,274 rows, and the experiment table reproduced exactly. Throughput will differ on any other machine, and should be higher on an idle one.</p><p>Code is in <code>projects/08-leakage-checked-backtester</code>. The raw outputs behind every number in this post are in its <code>results/</code> directory, <code>DESIGN.md</code> lists each invariant with the test that enforces it, and <code>DEVLOG.md</code> has the build log.</p>]]>
    </content>
    <id>https://projects.farhansadeek.com/posts/a-backtester-that-audits-itself/</id>
    <link href="https://projects.farhansadeek.com/posts/a-backtester-that-audits-itself/"/>
    <published>2026-09-29T06:17:07.000Z</published>
    <summary>A crypto research platform with a point-in-time store, purged cross-validation and a look-ahead audit that flags leaky features before they reach a backtest.</summary>
    <title>A Backtester That Refuses to Run on Leaky Features</title>
    <updated>2026-09-29T06:49:26.551Z</updated>
  </entry>
  <entry>
    <author>
      <name>Farhan Sadeek</name>
      <email>farhan@farhansadeek.com</email>
    </author>
    <category term="operating-systems" scheme="https://projects.farhansadeek.com/tags/operating-systems/"/>
    <category term="risc-v" scheme="https://projects.farhansadeek.com/tags/risc-v/"/>
    <category term="c" scheme="https://projects.farhansadeek.com/tags/c/"/>
    <content>
      <![CDATA[<p>I wrote 16-os, a small kernel for 64-bit RISC-V, in freestanding C and assembly. It boots on QEMU's <code>virt</code> board under OpenSBI, runs in supervisor mode the way Linux does, reads the machine's layout from the device tree, turns on Sv39 paging with per-section permissions and guard pages under every stack, handles traps and interrupts, and schedules preemptive kernel threads. The whole thing is <strong>3,521 lines</strong> of C, assembly, headers and linker script, and <strong>0x96ce bytes</strong> (38,606) of machine code in <code>.text</code>.</p><p>The number I trust most is not a benchmark. Every boot can run a built-in suite of <strong>22 self-tests</strong> that powers QEMU off with a pass or fail exit code, and four more boots crash on purpose and must print a readable report and exit with code 3. That makes <code>make test</code> a headless command, and on 2026-09-26 an independent clean rebuild (<code>make clean</code>, <code>make -j2</code>, <code>make test</code>) passed every self-test and all four crash reports.</p><p>Under QEMU's deterministic <code>-icount</code> mode a raw context switch costs <strong>34</strong> guest instructions, a switch through the scheduler <strong>190</strong>, a trap round trip <strong>140</strong> and a call into firmware <strong>291</strong>. These are QEMU instruction counts, not cycles on real hardware. The wall-clock timings I also took ran inside QEMU on a Mac with a load average of 129 to 151 from other jobs, so they serve only as a cross-check.</p><p>Code is in <code>projects/16-riscv-sv39-kernel</code>. v0 is kernel only, with no user mode yet.</p><p><em>Reading note.</em> The argument is that the hard part of a small kernel is not paging or the trap vector, which worked the first time they were switched on, but making every claim checkable and then finding the races that only show up when the same boot runs thirty or forty times. Skip to <a href="#problems">problems</a> for the bugs.</p><h2 id="table-of-contents">table of contents</h2><ul><li><a href="#what-i-wanted-to-build">what I wanted to build</a></li><li><a href="#theory">theory</a></li><li><a href="#architecture">architecture</a></li><li><a href="#implementation">implementation</a></li><li><a href="#problems">problems</a></li><li><a href="#experiments">experiments</a></li><li><a href="#results">results</a></li><li><a href="#what-i-would-change">what I would change</a></li><li><a href="#reproducibility">reproducibility</a></li></ul><h2 id="what-i-wanted-to-build">what I wanted to build</h2><p>I wanted a kernel small enough to read in an evening that still does the things a real one has to do. The goals for v0 were these.</p><ul><li>Boot through the standard firmware interface (SBI) rather than owning machine mode, the way Linux boots on real boards.</li><li>Hard-code nothing about the machine. RAM size, device addresses and the timer frequency come from the device tree.</li><li>Turn on paging with real permissions and an unmapped guard page below every stack.</li><li>Turn every unexpected trap into a readable report with a symbolized backtrace.</li><li>Run kernel threads that preempt each other, with spinlocks, sleep locks and wakeups that are never lost.</li><li>Check every one of those claims with a test that runs headless and fails loudly.</li></ul><p>The last goal shaped the rest. A kernel that writes to its own text segment, expects a store page fault, recovers and reports the result through QEMU's exit status proves the permission is enforced, which a boot banner does not. I read MIT's xv6-riscv and the RISC-V privileged specification for reference. The code is my own, and I note where it differs from xv6.</p><h2 id="theory">theory</h2><h3 id="three-privilege-levels-and-a-firmware-interface">three privilege levels and a firmware interface</h3><p>RISC-V has machine mode (M), supervisor mode (S) and user mode (U). On QEMU's <code>virt</code> board with <code>-bios default</code>, every hart starts in M-mode inside OpenSBI. OpenSBI configures physical memory protection, delegates most traps and interrupts down to S-mode and then jumps to the kernel at <code>0x80200000</code> in S-mode, with the hart id in <code>a0</code> and the physical address of the device tree in <code>a1</code>.</p><p>From then on the kernel asks firmware for services with <code>ecall</code>, the same instruction a user program would use to ask a kernel. The interface is the Supervisor Binary Interface, and I use four of its extensions. TIME arms the next timer interrupt, HSM reports the state of other harts, SRST powers the machine off, and DBCN gives a console before any UART driver exists.</p><h3 id="what-a-trap-does-in-hardware">what a trap does in hardware</h3><p>When an exception or interrupt is taken in S-mode, the hardware does very little. It saves the PC in <code>sepc</code>, writes the cause to <code>scause</code> (the top bit says interrupt or exception, and the low bits give the code), puts the faulting address or the faulting instruction's bits in <code>stval</code>, copies the interrupt-enable bit <code>sstatus.SIE</code> into <code>SPIE</code> and clears <code>SIE</code>, and jumps to the address in <code>stvec</code>. Every general-purpose register still holds what the interrupted code left in it, so the handler's first instruction runs with no free scratch register. Saving registers is software's job (the <code>sscratch</code> CSR exists to make room for the first one), and <code>sret</code> undoes the CSR part.</p><h3 id="sv39-in-one-diagram">Sv39 in one diagram</h3><p>Sv39 translates a 39-bit virtual address through three levels of page tables. Each table is one 4 KiB page of 512 eight-byte entries, and each level is indexed by 9 bits of the address.</p><figure data-figure="diagram:sv39-translation"></figure><p>An entry with any of R, W or X set is a leaf. One with only V set points to the next level. The A (accessed) and D (dirty) bits have a subtle rule. Without the Svadu extension enabled, hardware that finds A clear on an access, or D clear on a write, raises a page fault instead of setting the bit. QEMU's hart advertises Svadu, but I did not want correctness to depend on it, so every leaf is created with A set, and D set if it is writable.</p><h3 id="what-a-context-switch-has-to-save">what a context switch has to save</h3><p>A switch between kernel threads is an ordinary function call. The calling convention already makes the caller spill every caller-saved register it still needs, so the switch only has to save the callee-saved ones, which are <code>ra</code>, <code>sp</code> and <code>s0</code> to <code>s11</code>. That is 14 registers stored into the old thread's context and 14 loaded from the new one. The trick is that <code>ra</code> is loaded from the new context, so <code>ret</code> returns into a different thread.</p><h3 id="locking-with-interrupts">locking with interrupts</h3><p>On one hart the danger with a spinlock is an interrupt handler that wants a lock its own hart holds, which spins forever. The standard rule, which xv6 also uses, is that holding any spinlock disables interrupts locally. The first acquire remembers whether they were on, and the last release restores that.</p><p>The classic scheduling bug is the lost wakeup. A thread checks a condition, finds it false, and is about to go to sleep. Before its state changes, the condition becomes true and the wakeup fires, finds nobody sleeping and does nothing. Then the thread goes to sleep, forever. The fix is to make the check and the state change atomic with respect to <code>wakeup</code>, which in practice means both must hold the same lock.</p><h2 id="architecture">architecture</h2><p>The kernel is one image linked at <code>0x80200000</code>. Text, rodata and data each start on a page boundary so they can get different permissions. After them come the boot stack slots, each a guard page followed by a 16 KiB stack. This is the physical layout from a 128 MiB boot, with addresses taken from the boot log.</p><figure data-figure="diagram:riscv-memory-map"></figure><p>The kernel page table identity-maps all of RAM and the three devices it uses, each region with its own permissions. Thread stacks are the exception. Each lives at its own high virtual address starting at <code>0x0000_003f_0000_0000</code>, as four separately allocated frames mapped contiguously above an unmapped guard page. The boot log shows the table needed <strong>71 pages</strong>, which is mostly the cost of using 4 KiB leaves for 128 MiB of identity-mapped RAM.</p><p>Traps follow one path, and the important design choice is where the frame goes.</p><figure data-figure="diagram:riscv-trap-path"></figure><p>The frame lives on the interrupted thread's own stack, not a per-hart trap stack, and that is what makes preemption simple. The timer path calls <code>yield()</code> inside <code>kernel_trap</code>, other threads run on their own stacks, and when this thread is picked again it returns out of <code>kernel_trap</code> and <code>sret</code>s to where it was interrupted. A per-hart trap stack would be overwritten by the next trap.</p><p>The scheduler is a static table of 32 threads, a FIFO run queue under one spinlock, and one idle thread per hart made from the boot flow. Its states are these.</p><figure data-figure="diagram:riscv-thread-states"></figure><p>xv6 switches from a thread to a per-CPU scheduler context and from there to the next thread, which is two <code>swtch</code> calls per switch. I switch directly from the old thread to the new one. That halves the switches, and the price is a lock handoff that I describe below.</p><p>The kernel command line, passed through the device tree's <code>bootargs</code>, picks the first thread's job. It can be the demo with an interactive UART monitor, the self-tests, the benchmarks, or one of four deliberate crashes.</p><h2 id="implementation">implementation</h2><h3 id="the-first-instructions">the first instructions</h3><p>OpenSBI jumps to <code>_start</code> with interrupts off and the MMU off. The entry code rejects hart ids beyond its tables, runs a lottery so exactly one hart continues, points <code>sp</code> at that hart's boot stack, zeroes BSS and calls <code>kmain</code>.</p><figure class="highlight plaintext"><table><tr><td class="code"><pre><span class="line">la t0, boot_lottery</span><br><span class="line">li t1, 1</span><br><span class="line">amoswap.w.aq t1, t1, (t0)</span><br><span class="line">bnez t1, park              /* someone else is the boot hart */</span><br></pre></td></tr></table></figure><p>The lottery word lives in <code>.data</code>, because the winner's BSS clear would reset a BSS word and let a slow second hart win again. With SBI HSM the other hart never arrives (the boot log reports it as &quot;stopped&quot;), so the lottery only matters for older firmware. The entry code also clears <code>s0</code> and <code>ra</code>, so every backtrace ends at a zero frame pointer.</p><h3 id="finding-the-machine">finding the machine</h3><p>The device tree parser walks the flattened blob once. Properties always precede child nodes, so a node is finished when its first child or end token appears. One pass collects memory, the <code>reserved-memory</code> children (OpenSBI reserves itself there), the UART, the PLIC, the <code>sifive,test0</code> exit device, the timebase and <code>bootargs</code>. On the 128 MiB machine it found <strong>32,768 pages</strong> and left <strong>32,127</strong> free after the kernel, the bitmap, OpenSBI's ranges and the blob itself.</p><h3 id="a-page-allocator-that-notices-double-frees">a page allocator that notices double frees</h3><p>The physical allocator is a bitmap with one bit per 4 KiB frame, which is 4 KiB of bitmap for 128 MiB. Allocation scans 64-bit words from a rotating hint and uses count-trailing-zeros to find a free bit, so the common case touches one word.</p><p>I picked a bitmap over a free list because freeing becomes a bit test, so a double free, an unaligned free or a free outside RAM returns an error instead of corrupting a list. xv6's free list cannot detect a double free at all. One self-test allocates every free page (32,050 at that point in the suite), writes a tag into each, chains them together, checks every tag and frees them all.</p><h3 id="paging-and-proving-the-permissions">paging, and proving the permissions</h3><p>Because RAM is identity mapped, the PC after the <code>satp</code> write is still valid, so paging turns on without a jump. The mapping function sets A and D up front.</p><figure class="highlight c"><table><tr><td class="code"><pre><span class="line"><span class="comment">/* Set A (and D for writable pages) up front: without Svadu the</span></span><br><span class="line"><span class="comment"> * hardware raises a page fault instead of setting them. */</span></span><br><span class="line">u64 ad = PTE_A | ((perm &amp; PTE_W) ? PTE_D : <span class="number">0</span>);</span><br><span class="line">*pte = PA2PTE(pa + off) | perm | ad | PTE_V;</span><br></pre></td></tr></table></figure><p>A permission table is only a claim until something tries to violate it. The self-tests write to text and to rodata and expect a store page fault, execute a <code>ret</code> placed in <code>.data</code> and expect an instruction page fault, and read one byte below every stack and expect a load page fault. Doing that without killing the kernel needs fault probes, which are small assembly helpers that store a recovery address in the per-hart structure before one risky instruction.</p><figure class="highlight plaintext"><table><tr><td class="code"><pre><span class="line">probe_read64:              /* int probe_read64(u64 addr, u64 *out) */</span><br><span class="line">    la t0, 1f</span><br><span class="line">    sd t0, CPU_ONFAULT(tp)</span><br><span class="line">    ld t1, 0(a0)           /* the one instruction allowed to fault */</span><br><span class="line">    sd zero, CPU_ONFAULT(tp)</span><br><span class="line">    sd t1, 0(a1)</span><br><span class="line">    li a0, 0</span><br><span class="line">    ret</span><br><span class="line">1:  li a0, -1</span><br><span class="line">    ret</span><br></pre></td></tr></table></figure><p>If the load faults, <code>kernel_trap</code> sees <code>onfault</code> set, records the cause, points <code>sepc</code> at the recovery label and returns, and the probe returns -1. It is the idea behind Linux's exception tables, reduced to one instruction per helper.</p><h3 id="a-heap-with-checked-frees">a heap with checked frees</h3><p><code>kmalloc</code> has seven power-of-two size classes from 32 to 2,048 bytes. Each class cuts whole pages into equal blocks, and every block carries a 16-byte header with a magic value, the class and the size. So <code>kfree</code> needs no size, payloads stay 16-byte aligned, and a double free or foreign pointer fails the magic check. Larger requests take contiguous pages. A stress test runs 5,000 random operations of 1 to 3,000 bytes and checks every block's fill pattern before freeing it.</p><h3 id="the-trap-vector-checks-for-overflow-before-it-pushes">the trap vector checks for overflow before it pushes</h3><p>The obvious trap vector subtracts the frame size from <code>sp</code> and starts storing. If the trap was a stack overflow into the guard page, that first store faults again, which traps again, and the kernel loops forever without printing anything. So the vector checks first.</p><figure class="highlight plaintext"><table><tr><td class="code"><pre><span class="line">kernelvec:</span><br><span class="line">    csrw sscratch, t0</span><br><span class="line">    ld t0, CPU_KSTACK_LO(tp)</span><br><span class="line">    addi t0, t0, TF_SIZE</span><br><span class="line">    bltu sp, t0, kstack_overflow   /* pushing would fault again */</span><br><span class="line">    csrr t0, sscratch</span><br><span class="line">    addi sp, sp, -TF_SIZE</span><br><span class="line">    ...</span><br></pre></td></tr></table></figure><p><code>tp</code> points at the hart's <code>struct cpu</code>, whose <code>kstack_lo</code> the scheduler updates on every switch. On overflow the vector moves to a per-hart emergency stack and reports. In the deliberate crash the report said <code>sp</code> was 16 bytes below the stack bottom, named the faulting instruction as a store in <code>recurse+0xc</code>, and folded the recursion into one line, &quot;same frame repeated 509 more times&quot;, before ending at <code>crash_thread</code> and <code>thread_start</code>.</p><p>The vector saves all 31 registers, where xv6 saves only the caller-saved ones, which gives every fault report a complete register dump. It also writes a fake frame record holding the interrupted <code>s0</code> and <code>sepc</code>, so a backtrace walks through the trap into the faulting code.</p><h3 id="symbolized-backtraces-from-a-two-pass-link">symbolized backtraces from a two-pass link</h3><p>Raw addresses are useless in a CI log, so the kernel carries its own symbol table. The catch is that the table depends on the final addresses and adding it could move them. So the Makefile links twice. The first link has an empty table, its text symbols become a sorted C array, and the second link puts that array in <code>.rodata</code>. <code>.rodata</code> follows <code>.text</code>, so no code address can move, and every build proves it.</p><figure class="highlight make"><table><tr><td class="code"><pre><span class="line">@<span class="variable">$(NM)</span> -n <span class="variable">$(BUILD)</span>/kernel.pass1.elf | grep -i &#x27; t &#x27; &gt; <span class="variable">$(BUILD)</span>/text1.txt</span><br><span class="line">@<span class="variable">$(NM)</span> -n <span class="variable">$@</span> | grep -i &#x27; t &#x27; &gt; <span class="variable">$(BUILD)</span>/text2.txt</span><br><span class="line">@cmp -s <span class="variable">$(BUILD)</span>/text1.txt <span class="variable">$(BUILD)</span>/text2.txt || (echo <span class="string">&quot;error: text symbols moved between link passes&quot;</span>; exit 1)</span><br></pre></td></tr></table></figure><p>Backtraces walk the frame-pointer chain (<code>-fno-omit-frame-pointer</code>), where each function saves <code>ra</code> at <code>fp - 8</code> and the caller's <code>fp</code> at <code>fp - 16</code>. The lookup uses <code>ra - 1</code>, because a call to a <code>noreturn</code> function can be the last instruction of its function, which would put <code>ra</code> in the next one. That detail comes back in problem 4.</p><h3 id="the-scheduler-and-the-lock-handoff">the scheduler and the lock handoff</h3><p><code>sched()</code> pops the head of the run queue, or the idle thread if it is empty, and switches straight to it. Before switching it asserts the invariants that make this safe, which are that <code>sched_lock</code> is held and is the only spinlock held, that interrupts are off, and that the current thread is no longer <code>RUNNING</code>.</p><figure class="highlight c"><table><tr><td class="code"><pre><span class="line"><span class="type">int</span> intena = c-&gt;intena; <span class="comment">/* belongs to this thread, not to the hart */</span></span><br><span class="line">swtch(&amp;cur-&gt;ctx, &amp;next-&gt;ctx);</span><br><span class="line">mycpu()-&gt;intena = intena;</span><br></pre></td></tr></table></figure><p>The &quot;were interrupts on before the first lock&quot; flag lives in the per-hart structure but describes the thread that took the lock, so <code>sched</code> carries it across the switch in a local variable. The lock handoff is the price of switching directly. The old thread acquires <code>sched_lock</code>, and the new thread releases it, either in the caller of its own earlier <code>sched</code> or, if it has never run, in <code>thread_start</code>.</p><figure class="highlight c"><table><tr><td class="code"><pre><span class="line"><span class="type">void</span> <span class="title function_">thread_start</span><span class="params">(<span class="type">void</span>)</span> &#123;</span><br><span class="line">    spin_unlock(&amp;sched_lock); <span class="comment">/* acquired by the thread that switched to us */</span></span><br><span class="line">    intr_on();</span><br><span class="line">    <span class="class"><span class="keyword">struct</span> <span class="title">thread</span> *<span class="title">t</span> =</span> thread_current();</span><br><span class="line">    t-&gt;fn(t-&gt;arg);</span><br><span class="line">    thread_exit(<span class="number">0</span>);</span><br><span class="line">&#125;</span><br></pre></td></tr></table></figure><p><code>sleep_on</code> prevents the lost wakeup by taking <code>sched_lock</code> before it releases the caller's lock. <code>wakeup</code> needs <code>sched_lock</code> too, so it cannot run between the sleeper's check of its condition and its state becoming <code>SLEEPING</code>.</p><figure class="highlight c"><table><tr><td class="code"><pre><span class="line"><span class="type">void</span> <span class="title function_">sleep_on</span><span class="params">(<span class="type">void</span> *chan, <span class="keyword">struct</span> spinlock *lk)</span> &#123;</span><br><span class="line">    ...</span><br><span class="line">    <span class="keyword">if</span> (lk != &amp;sched_lock) &#123;</span><br><span class="line">        spin_lock(&amp;sched_lock);</span><br><span class="line">        spin_unlock(lk);</span><br><span class="line">    &#125;</span><br><span class="line">    t-&gt;chan = chan;</span><br><span class="line">    t-&gt;state = T_SLEEPING;</span><br><span class="line">    sched();</span><br><span class="line">    ...</span><br></pre></td></tr></table></figure><p>The idle loop has its own race. If an interrupt made a thread runnable between its run-queue check and its <code>wfi</code>, the hart would sleep until the next tick. So it checks with interrupts off and runs <code>wfi</code> while they are still off, which works because <code>wfi</code> wakes on a pending interrupt even when <code>SIE</code> is clear. Preemption is one 10 ms tick at 100 Hz, with deadlines on a fixed grid so ticks do not drift.</p><h3 id="exit-codes-that-qemu-can-see">exit codes that QEMU can see</h3><p>SBI's system reset can power off but cannot carry an exit code. QEMU's <code>sifive,test0</code> device can, since writing <code>(code &lt;&lt; 16) | 0x3333</code> makes QEMU exit with that code. So a pass powers off through SBI and a failure writes to the test device. The harness boots QEMU headless under a timeout, types <code>ping</code> into the UART to exercise receive interrupts, and checks the exit code and the report text. The kernel is also built without F and D and runs with <code>sstatus.FS</code> off, so a stray floating-point instruction traps instead of corrupting state that traps never save.</p><h2 id="problems">problems</h2><p>I expected the usual trouble when paging came on, and it did not happen. The first boot with <code>satp</code> written ran the whole demo. I think three choices prevented the usual failures. Every section boundary is page aligned in the linker script, A and D are set on every leaf, and the trap frame offsets are shared between C and assembly through one header that C checks with <code>_Static_assert</code>. The bugs I did hit were subtler, and three of the four were races.</p><h3 id="1-a-flaky-timer-test-and-a-fix-that-fixed-nothing">1. a flaky timer test, and a fix that fixed nothing</h3><p>The timer self-test busy-waits 100 ms and checks two things, that about 10 ticks happened, and that the global <code>ticks</code> counter and this hart's timer-interrupt count advanced by the same amount. It failed once with <code>check failed: mycpu()-&gt;ntimer - n0 == dt</code>.</p><p>My first theory was the compiler caching a counter that the interrupt handler updates, so I made the counters <code>volatile</code>, and the test passed. Later I rebuilt without <code>volatile</code> to reproduce the failure, and it still passed. The disassembly showed the counter was reloaded after an opaque <code>kprintf</code> call anyway, so <code>volatile</code> had changed nothing.</p><p>The real bug was in the test. It read <code>ticks</code> and the per-hart count as two separate loads with interrupts enabled, and at the end of the test a slow <code>kprintf</code> sat between the two loads. A tick landing in that gap made the two snapshots disagree by one. The fix takes both values in one snapshot with interrupts off.</p><figure class="highlight c"><table><tr><td class="code"><pre><span class="line"><span class="type">static</span> <span class="type">void</span> <span class="title function_">tick_snapshot</span><span class="params">(u64 *t, u64 *n)</span> &#123;</span><br><span class="line">    push_off();</span><br><span class="line">    *t = ticks;</span><br><span class="line">    *n = mycpu()-&gt;ntimer;</span><br><span class="line">    pop_off();</span><br><span class="line">&#125;</span><br></pre></td></tr></table></figure><p>To be sure this time, I booted the <code>only=ticks</code> test 30 times with each version under plain TCG. The version before the fix failed <strong>4 of 30</strong> boots, and the fixed version <strong>0 of 30</strong>. The failing boots saw 3 to 9 ticks instead of 10, so the host (load average 299) was starving QEMU's vCPU thread, which widens the window. A fix for a failure you cannot reproduce is a guess, and my first guess was wrong.</p><h3 id="2-a-store-the-compiler-was-right-to-delete">2. a store the compiler was right to delete</h3><p><code>volatile</code> was still needed, just somewhere else. The timer latency benchmark sets a <code>recording</code> flag, spins until 200 ticks pass, and clears the flag. The interrupt handler records a latency sample whenever the flag is set.</p><figure class="highlight c"><table><tr><td class="code"><pre><span class="line">timer_lat.recording = <span class="literal">true</span>;</span><br><span class="line">u64 end = ticks + <span class="number">200</span>;</span><br><span class="line"><span class="keyword">while</span> (ticks &lt; end)</span><br><span class="line">    ;</span><br><span class="line">timer_lat.recording = <span class="literal">false</span>;</span><br></pre></td></tr></table></figure><p>The loop reads only the <code>volatile</code> <code>ticks</code>. To the compiler, <code>recording = true</code> is followed by <code>recording = false</code> with nothing in between that could observe it, so it deleted the first store. The busy variant recorded no samples and printed this.</p><figure class="highlight plaintext"><table><tr><td class="code"><pre><span class="line">BENCH timer_latency_busy     n=0 min=0 p50=0 p99=0 max=0 mean=18446744073709551615 ns (resolution 100 ns)</span><br></pre></td></tr></table></figure><p>That mean is 2^64 - 1, and it is a lesson of its own. RISC-V integer division by zero does not trap. It returns all ones, so <code>sum / n</code> with <code>n = 0</code> quietly produced the largest <code>u64</code>. The idle variant worked because an opaque call to <code>thread_sleep_ticks</code> sat between its stores. Marking the latency structure <code>volatile</code> and guarding the division fixed it, and a rebuild with both fixes removed reproduces exactly this output.</p><h3 id="3-a-failing-boot-that-exited-with-0">3. a failing boot that exited with 0</h3><p>After the kernel printed a perfect page-fault report, <code>make test</code> once failed because QEMU had exited with 0 instead of 3. The report was right and the verdict was wrong.</p><p>My exit path wrote the failure value to the test device and then, &quot;just in case&quot;, fell through to an SBI shutdown call. QEMU acts on the test device write asynchronously, so the vCPU kept running for a moment. It reached the SBI call, and OpenSBI's poweroff on this board writes the <em>pass</em> value to the same device. Two exit requests were in flight, and whichever QEMU processed decided the exit code.</p><p>The fix is to never fall through. After the failure write the hart parks.</p><figure class="highlight c"><table><tr><td class="code"><pre><span class="line"><span class="keyword">if</span> (bootinfo.test_base) &#123;</span><br><span class="line">    mmio_write32(bootinfo.test_base, code == <span class="number">0</span> ? <span class="number">0x5555</span> : (((u32)code &lt;&lt; <span class="number">16</span>) | <span class="number">0x3333</span>));</span><br><span class="line">    <span class="comment">/* QEMU acts on this write asynchronously. Do not fall through to the</span></span><br><span class="line"><span class="comment">     * SRST call below ... */</span></span><br><span class="line">    <span class="keyword">for</span> (;;) wfi();</span><br><span class="line">&#125;</span><br></pre></td></tr></table></figure><p>I measured it the same way as the first bug, 20 boots each of <code>crash=pagefault</code> and <code>crash=panic</code>, expecting exit code 3. Before the fix <strong>10 of 40</strong> boots exited with the wrong code, and after it <strong>0 of 40</strong>. In a suite that boots each crash once, a one-in-four failure looks like an occasional red build that passes on rerun, which is exactly what people learn to ignore.</p><figure data-figure="chart:projects/os-from-scratch/os-from-scratch-bug-repro"></figure><h3 id="4-a-function-the-thread-never-called">4. a function the thread never called</h3><p>Every backtrace from a kernel thread ended with one bogus frame. With the fix reverted, the deliberate panic's backtrace reads like this.</p><figure class="highlight plaintext"><table><tr><td class="code"><pre><span class="line">#3  0x0000000080206204  crash_thread+0x48</span><br><span class="line">#4  0x0000000080200d94  thread_start+0x2c</span><br><span class="line">#5  0x0000000080200d68  thread_init+0x6c</span><br></pre></td></tr></table></figure><p><code>thread_init</code> runs once at boot and never calls a thread's function. A fresh thread's context had <code>ra</code> pointing at <code>thread_start</code>, so the first <code>swtch</code> &quot;returned&quot; there, and <code>thread_start</code>'s prologue saved that <code>ra</code> as its own return address. The backtrace looked up <code>ra - 1</code>, the last byte of whatever precedes <code>thread_start</code> in memory, which was <code>thread_init</code>.</p><p>The fix is a three-instruction trampoline. A new context's <code>ra</code> points at it, and it clears <code>ra</code> and <code>s0</code> before jumping to <code>thread_start</code>, so the frame record <code>thread_start</code> saves ends the chain.</p><figure class="highlight plaintext"><table><tr><td class="code"><pre><span class="line">thread_trampoline:</span><br><span class="line">    li ra, 0</span><br><span class="line">    li s0, 0</span><br><span class="line">    j thread_start</span><br></pre></td></tr></table></figure><h3 id="5-smaller-things-that-cost-time">5. smaller things that cost time</h3><ul><li><code>-std=c17</code> turns off the <code>asm</code> keyword, so the kernel builds as <code>gnu17</code>.</li><li>Clang's <code>-Winfinite-recursion</code> (an error under <code>-Werror</code>) rejected the deliberate stack-overflow function until the recursion depended on a <code>volatile</code> flag.</li><li>zsh does not split an unquoted <code>$VAR</code> into words, so my first batch of reproduction runs silently started nothing.</li><li><code>uart_init</code> resets the receive FIFO, which throws away any keystrokes QEMU delivered before the driver ran. The test harness waits for the self-test banner before it types.</li></ul><h3 id="6-the-host-was-the-noisiest-component">6. the host was the noisiest component</h3><p>Fifteen other builds shared this machine and the load average sat in the hundreds. Under plain TCG, QEMU's virtual clock follows the host clock, so a starved vCPU thread misses ticks, which are dropped rather than replayed. I loosened the tick test's lower bound and added a second pass of the suite under <code>-icount</code>, which does not depend on the host, and that is why this post's headline numbers are instruction counts.</p><h2 id="experiments">experiments</h2><p>All measurements run inside QEMU and none is a hardware number. Under plain TCG a timing includes translation, emulator bookkeeping and host scheduling, on a host heavily loaded by other jobs. Under <code>-icount shift=0,sleep=off</code>, virtual time advances one nanosecond per guest instruction, so an icount &quot;nanosecond&quot; is an instruction count. It is deterministic and ignores host load, but it treats every instruction as equal, which a real core does not. Icount numbers are the primary result here, and TCG timings a cross-check. Every benchmark loops and divides, because <code>rdtime</code> ticks at 10 MHz, a resolution of 100 ns.</p><h3 id="what-a-switch-costs">what a switch costs</h3><p>A raw <code>swtch</code> between two contexts costs <strong>34</strong> guest instructions in each direction. Most of that is the 28 loads and stores, and the rest is the return and the benchmark loop around it. A full <code>yield</code> from one thread to another through the scheduler costs <strong>190</strong>. The other 156 instructions are <code>sched_lock</code> with its interrupt masking and nesting counter, the run queue, the assertions in <code>sched</code> and bookkeeping. So the register switch is a small part of the cost and the locking around it is most of it. I did not measure xv6 on the same setup, so I cannot say how much the direct switch saves in practice.</p><h3 id="traps-against-firmware-calls">traps against firmware calls</h3><p>An <code>ebreak</code> trap into my vector and back costs <strong>140</strong> instructions. That covers the overflow check, 31 register stores, 4 CSR reads, dispatch, probe recovery, the restore and <code>sret</code>. An SBI <code>ecall</code> round trip costs <strong>291</strong>, about twice as much, since OpenSBI saves and dispatches in its own M-mode handler. Under TCG the gap nearly closes, 1,158 to 1,225 ns against 1,520 to 1,583 ns over three runs. My reading is that on this emulator most of a trap's cost is QEMU leaving translated code, which both paths pay.</p><h3 id="the-allocators">the allocators</h3><p>Allocating a page without zeroing costs <strong>105</strong> instructions and freeing one costs <strong>103</strong>. Allocating with zeroing costs <strong>2,184</strong>, so clearing 4 KiB dominates everything else in the allocator by a factor of about 20. A <code>kmalloc(64)</code> and <code>kfree</code> pair costs <strong>222</strong>. Under TCG the zeroing allocation took 1,622 to 1,645 ns, <em>less</em> than its instruction count, because QEMU runs a tight store loop faster than one instruction per nanosecond.</p><figure data-figure="chart:projects/os-from-scratch/os-from-scratch-icount-costs"></figure><h3 id="timer-interrupt-latency">timer interrupt latency</h3><p>The timer handler records the gap between the programmed deadline and the moment the C handler reads <code>rdtime</code>. I measured it with the hart idle in <code>wfi</code> and with a thread spinning, 200 ticks each.</p><p>Under icount the median was <strong>100 ns</strong> idle and busy, with a mean of 85 and 84 ns. The timebase resolution is 100 ns, so the kernel's path from deadline to handler is within one tick. Under plain TCG the median was between <strong>2.333 ms and 2.614 ms</strong> in all six runs, idle and busy, at load averages between 129 and 151. That is a steady offset, not scattered noise, and it appears even when idle. My guess, which I have not verified, is how QEMU arms its host timer on macOS. The icount numbers do show the offset is not in the kernel.</p><figure data-figure="chart:projects/os-from-scratch/os-from-scratch-timer-latency"></figure><h3 id="is-icount-actually-reproducible">is icount actually reproducible</h3><p>Two icount benchmark runs gave identical numbers in every column, and a raw diff of their benchmark lines was empty. The icount timer test saw exactly 10 ticks in 100 ms. A benchmark that gives the same answer twice can detect a one-instruction change on the switch path, which no wall-clock measurement on a loaded host could.</p><h3 id="why-the-lock-matters">why the lock matters</h3><p>Four threads each increment a shared counter, 80,000 increments in total, with a short delay between each read and each write so that preemption has a window to land in. Under the spinlock the count was exactly <strong>80,000</strong> in both passes. The unlocked control is more interesting. Under icount it lost <strong>3,605</strong> updates and ended at 76,395. Under plain TCG it lost <strong>0</strong>.</p><p>So the control is informational, not an assertion. One plausible reason for the TCG result is that a tick arriving while the lock is held is deferred until <code>spin_unlock</code> re-enables interrupts, just before the unlocked read, which biases preemption toward the edge of the racy window. I have not measured it.</p><h2 id="results">results</h2><ul><li><code>make test</code> passes. That is 22 self-tests under plain TCG, including the UART receive test with typed input, 21 under icount (the receive test needs typed input and runs only in the TCG pass), and four crash scenarios that each print the expected report and exit with code 3.</li><li>On 2026-09-26 an independent clean rebuild (<code>make clean</code>, <code>make -j2</code>, <code>make test</code>) passed every self-test and all four crash reports. That rerun checked correctness, not the benchmarks.</li><li>Under icount the kernel reaches its first thread <strong>16,061 µs</strong> of virtual time after reset. That includes OpenSBI's own startup.</li><li>The kernel page table uses <strong>71 pages</strong>, and 32,127 of 32,768 pages are free after boot.</li><li>In the demo, four workers are preempted between 4 and 7 times each, their shared tally ends at the expected 16, and the boot takes 48 context switches before it powers off.</li><li>Under icount a raw <code>swtch</code> costs <strong>34</strong> guest instructions, a scheduler switch <strong>190</strong>, a trap <strong>140</strong>, an SBI call <strong>291</strong>, and timer latency is within one 100 ns tick.</li><li>The interactive monitor works over UART receive interrupts, with 10 interrupts for the scripted session.</li></ul><h2 id="what-i-would-change">what I would change</h2><h3 id="go-higher-half-before-user-mode">go higher-half before user mode</h3><p>The identity map made v0 easy to debug, since every address in a fault report is also physical. User programs want the low half of the address space, so I would move the kernel to the top before writing any user-mode code.</p><h3 id="make-the-scheduler-smp-safe-in-the-one-place-it-is-not">make the scheduler SMP-safe in the one place it is not</h3><p><code>thread_join</code> frees a zombie's stack as soon as it sees <code>ZOMBIE</code>. On one hart the zombie has already left its stack, but with two it could still be finishing <code>sched</code> on it. It needs an &quot;on CPU&quot; flag, and each hart should get its own run queue.</p><h3 id="write-stimecmp-directly">write stimecmp directly</h3><p>The hart advertises Sstc, which lets S-mode write its timer compare register without an SBI call. That would replace an SBI round trip on every tick (the probe call I measured costs 291 instructions) with one CSR write, and I would like to know whether it also removes the TCG timer offset.</p><h3 id="tidy-the-memory-manager-and-the-console">tidy the memory manager and the console</h3><p>The 1.6 MiB between OpenSBI and the kernel image is never used, slab pages are never returned, and the contiguous allocator is a first-fit scan from zero. <code>kfree</code> on a wild pointer can fault on the header read, which a fault probe would turn into an error. UART output is polled, so a long <code>kprintf</code> keeps interrupts off while it writes.</p><h3 id="build-repeated-boot-testing-in-from-the-start">build repeated-boot testing in from the start</h3><p>The exit-code race and the tick race were both found by booting the same kernel 30 or 40 times. A single passing <code>make test</code> did not catch either, and both would have shown up as a flaky red build. Running each boot many times, preferably on a loaded host, should be a standard test mode, not a script I wrote after the fact.</p><h2 id="reproducibility">reproducibility</h2><p>The toolchain is Homebrew LLVM and QEMU on macOS. Apple's clang has no RISC-V backend. This project was built with clang 23.1.2, LLD 23.1.2 and QEMU 11.1.1, whose bundled OpenSBI (v1.8.1 in the logs) is what <code>-bios default</code> loads. From <code>projects/16-riscv-sv39-kernel</code>, the following commands build, test and reproduce every result.</p><figure class="highlight sh"><table><tr><td class="code"><pre><span class="line">brew install qemu llvm lld</span><br><span class="line">make toolchain        <span class="comment"># checks the toolchain and prints versions</span></span><br><span class="line"></span><br><span class="line">make                  <span class="comment"># build/kernel.elf (two-pass link, checks that text did not move)</span></span><br><span class="line">make <span class="built_in">test</span>             <span class="comment"># self-tests under TCG and icount, plus four crash reports</span></span><br><span class="line">make bench            <span class="comment"># 3 TCG runs and 2 icount runs, writes results/bench-summary.txt</span></span><br><span class="line">make results          <span class="comment"># regenerates everything in results/</span></span><br><span class="line"></span><br><span class="line">make run              <span class="comment"># interactive demo and monitor; quit with Ctrl-A then X</span></span><br><span class="line">make debug            <span class="comment"># QEMU halted at reset with a gdb stub on :1234</span></span><br><span class="line">make lldb             <span class="comment"># in a second terminal, attach Apple lldb to 127.0.0.1:1234</span></span><br></pre></td></tr></table></figure><p>A single test can be selected on the kernel command line, which is how the tick race was reproduced.</p><figure class="highlight sh"><table><tr><td class="code"><pre><span class="line">qemu-system-riscv64 -machine virt -cpu rv64 -smp 2 -m 128M -nographic \</span><br><span class="line">  -bios default -kernel build/kernel.elf -append <span class="string">&quot;mode=test only=ticks&quot;</span></span><br></pre></td></tr></table></figure><p>Add <code>-icount shift=0,sleep=off</code> for deterministic virtual time. Every number in this post comes from a file in <code>results/</code>, which holds the raw logs of each run, with the host load average recorded in the benchmark files.</p><p>Code is in <code>projects/16-riscv-sv39-kernel</code>, with the kernel in <code>kernel/</code>, the test and benchmark scripts in <code>tools/</code>, the memory map, trap flow and locking rules in <code>DESIGN.md</code>, and the full build log in <code>DEVLOG.md</code>.</p>]]>
    </content>
    <id>https://projects.farhansadeek.com/posts/a-risc-v-kernel-that-tests-itself/</id>
    <link href="https://projects.farhansadeek.com/posts/a-risc-v-kernel-that-tests-itself/"/>
    <published>2026-09-29T06:17:07.000Z</published>
    <summary>A small RISC-V kernel with Sv39 paging, traps and preemptive threads, plus a self-test suite that runs every time it boots.</summary>
    <title>A RISC-V Kernel That Tests Itself on Every Boot</title>
    <updated>2026-09-29T06:49:26.552Z</updated>
  </entry>
  <entry>
    <author>
      <name>Farhan Sadeek</name>
      <email>farhan@farhansadeek.com</email>
    </author>
    <category term="reinforcement-learning" scheme="https://projects.farhansadeek.com/tags/reinforcement-learning/"/>
    <category term="llm" scheme="https://projects.farhansadeek.com/tags/llm/"/>
    <category term="grpo" scheme="https://projects.farhansadeek.com/tags/grpo/"/>
    <content>
      <![CDATA[<p><strong>We have no before/after RL numbers.</strong> The pipeline runs end to end, the 300-step training run has not been launched, and we are publishing before it deliberately. Nothing below is contingent on what the final <code>pass@1</code> turns out to be.</p><p>What the pipeline did produce is a diagnosis and four post-mortems. The diagnosis is that the learning signal in GRPO is the within-group <em>spread</em> of rewards rather than their level, which makes &quot;my dataset is too easy&quot; and &quot;my dataset is too hard&quot; indistinguishable from outside the reward function. The post-mortems cover four engineering failures (a token-id collision, non-terminating rollouts, generation running in train mode, and a silently frozen adapter), none of which raised an exception, and all of which produced a training run that looked healthy.</p><p>We also measured hardware instead of assuming it. For this workload a rented Modal L4 ran about 5% faster per step than an M4 Pro laptop, because per-step time is dominated by serial decode overhead rather than arithmetic. That is one paired observation, not a benchmark.</p><p>Code is in <code>projects/math-rlvr</code>. The setup is GRPO via TRL 0.16 on <strong>Qwen2.5-0.5B-Instruct</strong>, LoRA (r=32, α=64, dropout 0.05, all projection modules) over a frozen base, <code>math_verify</code> as the reward, 8 generations per prompt, learning rate 1e-6, β 0.04, temperature 0.9, completions capped at 640 tokens.</p><p><em>Reading note.</em> The argument is the variance diagnostic and why the four bugs share a shape. Hyperparameters, the verifier comparison, and each post-mortem are collapsed. Expand only what you need.</p><h2 id="table-of-contents">table of contents</h2><ul><li><a href="#why-rlvr">why RLVR</a></li><li><a href="#grpo-and-where-the-signal-comes-from">GRPO, and where the signal comes from</a></li><li><a href="#the-diagnostic">the diagnostic</a></li><li><a href="#what-we-did-not-measure-about-difficulty">what we did not measure about difficulty</a></li><li><a href="#the-verifier-is-part-of-the-objective">the verifier is part of the objective</a></li><li><a href="#four-silent-failures">four silent failures</a></li><li><a href="#the-hardware-measurement">the hardware measurement</a></li><li><a href="#what-exists">what exists</a></li><li><a href="#conclusions">conclusions</a></li></ul><h2 id="why-rlvr">why RLVR</h2><p>Reinforcement learning from verifiable rewards is the cleanest setup in post-training. There is no reward model to train, no human preference data, and no judge model to be gamed. A math problem has a right answer, a program checks it, and the check <em>is</em> the reward. It is the recipe behind the reasoning-model results of the last two years, and at 500M it fits on one GPU.</p><p>That cleanliness is also what makes the failure modes legible. When the only moving parts are sample, score, and update, anything that goes wrong is in one of three places, which is why this project is worth writing up even without a training curve.</p><h2 id="grpo-and-where-the-signal-comes-from">GRPO, and where the signal comes from</h2><p>For each prompt, sample a group of <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>k</mi></mrow><annotation encoding="application/x-tex">k</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6944em;"></span><span class="mord mathnormal" style="margin-right:0.0315em;">k</span></span></span></span> completions and score them all. The advantage of completion <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>i</mi></mrow><annotation encoding="application/x-tex">i</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6595em;"></span><span class="mord mathnormal">i</span></span></span></span> is its reward minus the group mean, normalized by the group's spread.</p><p class="katex-block"><span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><msub><mi>A</mi><mi>i</mi></msub><mo>=</mo><mfrac><mrow><msub><mi>r</mi><mi>i</mi></msub><mo>−</mo><mi>μ</mi></mrow><mi>σ</mi></mfrac><mo separator="true">,</mo><mspace width="2em"/><mi>μ</mi><mo>=</mo><mfrac><mn>1</mn><mi>k</mi></mfrac><munderover><mo>∑</mo><mrow><mi>j</mi><mo>=</mo><mn>1</mn></mrow><mi>k</mi></munderover><msub><mi>r</mi><mi>j</mi></msub><mo separator="true">,</mo><mspace width="2em"/><mi>σ</mi><mo>=</mo><msqrt><mrow><mfrac><mn>1</mn><mi>k</mi></mfrac><munderover><mo>∑</mo><mrow><mi>j</mi><mo>=</mo><mn>1</mn></mrow><mi>k</mi></munderover><mo stretchy="false">(</mo><msub><mi>r</mi><mi>j</mi></msub><mo>−</mo><mi>μ</mi><msup><mo stretchy="false">)</mo><mn>2</mn></msup></mrow></msqrt></mrow><annotation encoding="application/x-tex">A_i = \frac{r_i - \mu}{\sigma}, \qquad\mu = \frac{1}{k}\sum_{j=1}^{k} r_j, \qquad\sigma = \sqrt{\frac{1}{k}\sum_{j=1}^{k}(r_j - \mu)^2}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.8333em;vertical-align:-0.15em;"></span><span class="mord"><span class="mord mathnormal">A</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3117em;"><span style="top:-2.55em;margin-left:0em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight">i</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:1.9463em;vertical-align:-0.686em;"></span><span class="mord"><span class="mopen nulldelimiter"></span><span class="mfrac"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:1.2603em;"><span style="top:-2.314em;"><span class="pstrut" style="height:3em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0359em;">σ</span></span></span><span style="top:-3.23em;"><span class="pstrut" style="height:3em;"></span><span class="frac-line" style="border-bottom-width:0.04em;"></span></span><span style="top:-3.677em;"><span class="pstrut" style="height:3em;"></span><span class="mord"><span class="mord"><span class="mord mathnormal" style="margin-right:0.0278em;">r</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3117em;"><span style="top:-2.55em;margin-left:-0.0278em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight">i</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">−</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mord mathnormal">μ</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.686em;"><span></span></span></span></span></span><span class="mclose nulldelimiter"></span></span><span class="mpunct">,</span><span class="mspace" style="margin-right:2em;"></span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mord mathnormal">μ</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:3.2499em;vertical-align:-1.4138em;"></span><span class="mord"><span class="mopen nulldelimiter"></span><span class="mfrac"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:1.3214em;"><span style="top:-2.314em;"><span class="pstrut" style="height:3em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0315em;">k</span></span></span><span style="top:-3.23em;"><span class="pstrut" style="height:3em;"></span><span class="frac-line" style="border-bottom-width:0.04em;"></span></span><span style="top:-3.677em;"><span class="pstrut" style="height:3em;"></span><span class="mord"><span class="mord">1</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.686em;"><span></span></span></span></span></span><span class="mclose nulldelimiter"></span></span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mop op-limits"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:1.8361em;"><span style="top:-1.8723em;margin-left:0em;"><span class="pstrut" style="height:3.05em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight" style="margin-right:0.0572em;">j</span><span class="mrel mtight">=</span><span class="mord mtight">1</span></span></span></span><span style="top:-3.05em;"><span class="pstrut" style="height:3.05em;"></span><span><span class="mop op-symbol large-op">∑</span></span></span><span style="top:-4.3em;margin-left:0em;"><span class="pstrut" style="height:3.05em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight" style="margin-right:0.0315em;">k</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:1.4138em;"><span></span></span></span></span></span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0278em;">r</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3117em;"><span style="top:-2.55em;margin-left:-0.0278em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight" style="margin-right:0.0572em;">j</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.2861em;"><span></span></span></span></span></span></span><span class="mpunct">,</span><span class="mspace" style="margin-right:2em;"></span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mord mathnormal" style="margin-right:0.0359em;">σ</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:3.4776em;vertical-align:-1.4138em;"></span><span class="mord sqrt"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:2.0639em;"><span class="svg-align" style="top:-5.4376em;"><span class="pstrut" style="height:5.4376em;"></span><span class="mord" style="padding-left:1.056em;"><span class="mord"><span class="mopen nulldelimiter"></span><span class="mfrac"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:1.3214em;"><span style="top:-2.314em;"><span class="pstrut" style="height:3em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0315em;">k</span></span></span><span style="top:-3.23em;"><span class="pstrut" style="height:3em;"></span><span class="frac-line" style="border-bottom-width:0.04em;"></span></span><span style="top:-3.677em;"><span class="pstrut" style="height:3em;"></span><span class="mord"><span class="mord">1</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.686em;"><span></span></span></span></span></span><span class="mclose nulldelimiter"></span></span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mop op-limits"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:1.8361em;"><span style="top:-1.8723em;margin-left:0em;"><span class="pstrut" style="height:3.05em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight" style="margin-right:0.0572em;">j</span><span class="mrel mtight">=</span><span class="mord mtight">1</span></span></span></span><span style="top:-3.05em;"><span class="pstrut" style="height:3.05em;"></span><span><span class="mop op-symbol large-op">∑</span></span></span><span style="top:-4.3em;margin-left:0em;"><span class="pstrut" style="height:3.05em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight" style="margin-right:0.0315em;">k</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:1.4138em;"><span></span></span></span></span></span><span class="mopen">(</span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0278em;">r</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3117em;"><span style="top:-2.55em;margin-left:-0.0278em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight" style="margin-right:0.0572em;">j</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.2861em;"><span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">−</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mord mathnormal">μ</span><span class="mclose"><span class="mclose">)</span><span class="msupsub"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.7401em;"><span style="top:-2.989em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight">2</span></span></span></span></span></span></span></span></span></span><span style="top:-4.0239em;"><span class="pstrut" style="height:5.4376em;"></span><span class="hide-tail" style="min-width:0.742em;height:3.5176em;"><svg xmlns="http://www.w3.org/2000/svg" width="400em" height="3.5176em" viewBox="0 0 400000 3517" preserveAspectRatio="xMinYMin slice"><path d="M702 80H40000040H742v3383l-4 4-4 4c-.667.7 -2 1.5-4 2.5s-4.167 1.833-6.5 2.5-5.5 1-9.5 1h-12l-28-84c-16.667-52-96.667 -294.333-240-727l-212 -643 -85 170c-4-3.333-8.333-7.667-13 -13l-13-13l77-155 77-156c66 199.333 139 419.667219 661 l218 661zM702 80H400000v40H742z"/></svg></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:1.4138em;"><span></span></span></span></span></span></span></span></span></span></p><p>The policy is then pushed toward the above-average members of its own group. There is no value network and no critic to train. The group average <em>is</em> the baseline. That is the entire simplification, and at this scale it is the reason to use GRPO rather than PPO.</p><p>Read the numerator once more, because the dominant failure mode falls directly out of it. If every completion in a group receives the same reward, then <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>r</mi><mi>i</mi></msub><mo>=</mo><mi>μ</mi></mrow><annotation encoding="application/x-tex">r_i = \mu</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.5806em;vertical-align:-0.15em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0278em;">r</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3117em;"><span style="top:-2.55em;margin-left:-0.0278em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight">i</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.625em;vertical-align:-0.1944em;"></span><span class="mord mathnormal">μ</span></span></span></span> for all <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>i</mi></mrow><annotation encoding="application/x-tex">i</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6595em;"></span><span class="mord mathnormal">i</span></span></span></span>, every advantage is zero, and the gradient contribution of that prompt is exactly zero. You spend a GPU-minute generating 8 completions and learn nothing from them.</p><p>The condition for learning is therefore <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>σ</mi><mo>&gt;</mo><mn>0</mn></mrow><annotation encoding="application/x-tex">\sigma &gt; 0</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.5782em;vertical-align:-0.0391em;"></span><span class="mord mathnormal" style="margin-right:0.0359em;">σ</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">&gt;</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.6444em;"></span><span class="mord">0</span></span></span></span> within the group, and that is a property of the <em>difficulty distribution of the prompts</em>, not of the model's competence in any absolute sense.</p><h2 id="the-diagnostic">the diagnostic</h2><figure data-figure="grpo-advantage"></figure><p>Zero within-group variance arrives from both directions, and the two are indistinguishable downstream.</p><table><thead><tr><th>Training set</th><th>What happens</th><th>Reward std</th><th>Learning signal</th></tr></thead><tbody><tr><td>GSM8K</td><td>The Instruct model aces the problem 8/8</td><td>~0</td><td>None</td></tr><tr><td>AIME</td><td>90 olympiad problems, 0/8 on nearly all</td><td>~0</td><td>None</td></tr><tr><td>MATH levels 3–5</td><td>Sometimes right, sometimes wrong</td><td>&gt; 0</td><td>Yes</td></tr></tbody></table><p>&quot;Too easy&quot; and &quot;too hard&quot; produce the <em>identical</em> symptom. Flat reward, no movement, and a run that otherwise looks completely healthy. Wall-clock burns, the loss curve is plausible, and nothing improves. Any diagnosis that only asks whether reward is increasing cannot separate them, and the two have opposite fixes.</p><p>So the reward function prints its own variance on every batch.</p><figure class="highlight python"><table><tr><td class="code"><pre><span class="line">n = <span class="built_in">len</span>(rewards)</span><br><span class="line">mean = <span class="built_in">sum</span>(rewards) / n <span class="keyword">if</span> n <span class="keyword">else</span> <span class="number">0.0</span></span><br><span class="line">var = <span class="built_in">sum</span>((r - mean) ** <span class="number">2</span> <span class="keyword">for</span> r <span class="keyword">in</span> rewards) / n <span class="keyword">if</span> n <span class="keyword">else</span> <span class="number">0.0</span></span><br><span class="line"><span class="built_in">print</span>(<span class="string">f&quot;[reward] n=<span class="subst">&#123;n&#125;</span> mean=<span class="subst">&#123;mean:<span class="number">.3</span>f&#125;</span> std=<span class="subst">&#123;var ** <span class="number">0.5</span>:<span class="number">.3</span>f&#125;</span> &quot;</span></span><br><span class="line">      <span class="string">f&quot;(std~0 =&gt; no learning signal)&quot;</span>)</span><br></pre></td></tr></table></figure><p>Ten lines of arithmetic that convert a silent failure into a visible one. It is the single highest-leverage thing in the repository. The general form is that <strong>the learning signal in policy-gradient RL comes from the spread of outcomes, not their level.</strong> Curriculum design is not a nice-to-have here, it is a precondition.</p><details class="collapsible-section"><summary><strong>Training configuration</strong></summary><table><thead><tr><th>Parameter</th><th>Value</th></tr></thead><tbody><tr><td>Base model</td><td>Qwen2.5-0.5B-Instruct</td></tr><tr><td>Adapter</td><td>LoRA, r=32, α=64, dropout 0.05, all projection modules</td></tr><tr><td>Generations per prompt (<span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>k</mi></mrow><annotation encoding="application/x-tex">k</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6944em;"></span><span class="mord mathnormal" style="margin-right:0.0315em;">k</span></span></span></span>)</td><td>8</td></tr><tr><td>Learning rate</td><td>1e-6</td></tr><tr><td>KL coefficient (β)</td><td>0.04</td></tr><tr><td>Sampling temperature</td><td>0.9</td></tr><tr><td>Max completion length</td><td>640 tokens</td></tr><tr><td>Correctness reward</td><td>1.0 if <code>math_verify</code> accepts, else 0.0</td></tr><tr><td>Format shaping</td><td>0.2 for exactly one well-formed <code>\boxed{}</code></td></tr><tr><td>Planned horizon</td><td>300 steps (not launched)</td></tr></tbody></table><p>The format shaping term exists so that early in training there is a gradient toward parseable output before there is any gradient toward correctness. It is deliberately small relative to the correctness reward, so it cannot dominate once completions parse.</p></details><h2 id="what-we-did-not-measure-about-difficulty">what we did not measure about difficulty</h2><p>The claim above is that MATH levels 3–5 produce nonzero within-group variance and the alternatives did not. That is an observation from the diagnostic, not a characterization of the difficulty band.</p><p>We did not measure where the band actually sits for this model. The principled version sweeps difficulty against measured <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mtext>pass@</mtext><mn>8</mn></mrow><annotation encoding="application/x-tex">\text{pass@}8</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.8889em;vertical-align:-0.1944em;"></span><span class="mord text"><span class="mord">pass@</span></span><span class="mord">8</span></span></span></span> and selects problems whose success probability is near <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mn>0.5</mn></mrow><annotation encoding="application/x-tex">0.5</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6444em;"></span><span class="mord">0.5</span></span></span></span>, which is where the expected within-group spread is largest. For a binary reward with success probability <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>p</mi></mrow><annotation encoding="application/x-tex">p</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.625em;vertical-align:-0.1944em;"></span><span class="mord mathnormal">p</span></span></span></span>, the group variance is maximized at <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>p</mi><mo>=</mo><mn>0.5</mn></mrow><annotation encoding="application/x-tex">p = 0.5</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.625em;vertical-align:-0.1944em;"></span><span class="mord mathnormal">p</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.6444em;"></span><span class="mord">0.5</span></span></span></span>, since</p><p class="katex-block"><span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><msup><mi>σ</mi><mn>2</mn></msup><mo>=</mo><mi>p</mi><mo stretchy="false">(</mo><mn>1</mn><mo>−</mo><mi>p</mi><mo stretchy="false">)</mo><mi mathvariant="normal">.</mi></mrow><annotation encoding="application/x-tex">\sigma^2 = p(1-p).</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.8641em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0359em;">σ</span><span class="msupsub"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.8641em;"><span style="top:-3.113em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight">2</span></span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mord mathnormal">p</span><span class="mopen">(</span><span class="mord">1</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">−</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mord mathnormal">p</span><span class="mclose">)</span><span class="mord">.</span></span></span></span></span></p><p>That sweep has not been run, so &quot;levels 3–5&quot; is a working choice supported by a binary observation, not a tuned curriculum. It is the first thing we would run given more time, and it is cheap relative to training.</p><h2 id="the-verifier-is-part-of-the-objective">the verifier is part of the objective</h2><p>The other half of RLVR is deciding whether a completion's answer matches the gold answer. Completions are prompted to end in <code>\boxed{...}</code>, so extraction is a regex on the last match. Comparison is where the subtlety lives.</p><details class="collapsible-section"><summary><strong>String equality vs. symbolic comparison</strong></summary><table><thead><tr><th>Model's boxed answer</th><th>Gold</th><th>String equality</th><th><code>math_verify</code></th></tr></thead><tbody><tr><td><code>72</code></td><td><code>72</code></td><td>✅</td><td>✅</td></tr><tr><td><code>0.5</code></td><td><code>1/2</code></td><td>❌</td><td>✅</td></tr><tr><td><code>\frac{1}{2}</code></td><td><code>1/2</code></td><td>❌</td><td>✅</td></tr><tr><td><code>40</code></td><td><code>72</code></td><td>❌</td><td>❌</td></tr><tr><td><em>(no <code>\boxed{}</code>)</em></td><td><code>72</code></td><td>none</td><td>❌</td></tr></tbody></table></details><p>Rows two and three are the reason this matters, and the reason is not noise. A naive string check does not produce a <em>noisy</em> reward, it produces a <strong>biased</strong> one. It systematically punishes correct answers for being written in a different form, and therefore trains the model toward the dataset's formatting conventions rather than toward being right. Parsing both sides symbolically and checking mathematical equality is the only version that rewards the intended thing.</p><p>Stated generally, the verifier is not an approximation of the objective, it <em>is</em> the objective. Every systematic error in it becomes a systematic pressure on the policy.</p><h2 id="four-silent-failures">four silent failures</h2><p>None of these raised an exception. All four produced a training run that appeared to work. They share one shape. Two components with defensible independent defaults must agree, and don't. That is the most common category of bug we hit in this project.</p><details class="collapsible-section"><summary><strong>1. The EOS/pad collision</strong></summary><p>Instruct models set <code>eos_token</code> to <code>&lt;|im_end|&gt;</code>. HuggingFace pads stopped rollouts with <code>pad_token</code>. TRL masks each completion at the first <code>eos</code>. If those two tokens disagree, the completion mask is wrong, and nothing tells you, because a wrong mask is still a valid mask.</p><figure class="highlight python"><table><tr><td class="code"><pre><span class="line">tokenizer = AutoTokenizer.from_pretrained(MODEL, padding_side=<span class="string">&quot;left&quot;</span>)</span><br><span class="line">tokenizer.eos_token = tokenizer.pad_token  <span class="comment"># both &lt;|endoftext|&gt;</span></span><br></pre></td></tr></table></figure><p>One line, several hours of diagnosis.</p></details><details class="collapsible-section"><summary><strong>2. Rollouts that never terminate</strong></summary><p>We started on the Qwen2.5-Math-1.5B <strong>base</strong> model. Its chat tokens are untrained, so it never emits <code>&lt;|im_end|&gt;</code> at all, so every rollout ran to <code>max_completion_length</code>. At 8 generations per prompt and 640 tokens each, nearly all of the compute was spent generating text after the answer had already been given.</p><p>Two fixes followed, switching to the 500M Instruct model, which also roughly halved per-step time, and a stopping criterion that halts a rollout as soon as it contains a closed <code>\boxed{}</code>.</p><figure class="highlight python"><table><tr><td class="code"><pre><span class="line"><span class="keyword">class</span> <span class="title class_">_StopAfterBoxed</span>(<span class="title class_ inherited__">StoppingCriteria</span>):</span><br><span class="line">    <span class="keyword">def</span> <span class="title function_">__call__</span>(<span class="params">self, input_ids, scores, **kwargs</span>):</span><br><span class="line">        texts = <span class="variable language_">self</span>.tokenizer.batch_decode(</span><br><span class="line">            input_ids[:, <span class="variable language_">self</span>.prompt_len:], skip_special_tokens=<span class="literal">True</span></span><br><span class="line">        )</span><br><span class="line">        <span class="keyword">return</span> torch.tensor([BOXED_RE.search(t) <span class="keyword">is</span> <span class="keyword">not</span> <span class="literal">None</span> <span class="keyword">for</span> t <span class="keyword">in</span> texts],</span><br><span class="line">                            device=input_ids.device)</span><br></pre></td></tr></table></figure><p>The prompt itself mentions <code>\boxed{}</code>, so the criterion has to skip the prompt prefix, the kind of detail that turns a clean idea into a debugging session.</p></details><details class="collapsible-section"><summary><strong>3. Generation running in train mode</strong></summary><p>TRL 0.16 calls <code>generate()</code> without switching the model to eval. With gradient checkpointing enabled, that forces <code>use_cache=False</code>. No KV cache means decoding is quadratic in sequence length rather than linear, on the single most expensive part of the step.</p><figure class="highlight python"><table><tr><td class="code"><pre><span class="line">was_training = model.training</span><br><span class="line">model.<span class="built_in">eval</span>()</span><br><span class="line"><span class="keyword">try</span>:</span><br><span class="line">    <span class="keyword">return</span> orig_generate(input_ids, *args, **kwargs)</span><br><span class="line"><span class="keyword">finally</span>:</span><br><span class="line">    <span class="keyword">if</span> was_training:</span><br><span class="line">        model.train()</span><br></pre></td></tr></table></figure><p>A large speedup for a small patch, and again not something that surfaces as an error. It surfaces as &quot;training is slower than I expected,&quot; which is indistinguishable from &quot;training is this slow.&quot;</p></details><details class="collapsible-section"><summary><strong>4. The adapter that loaded frozen</strong></summary><p>If an adapter already exists at the output directory, the trainer loads it with <code>is_trainable=True</code> and continues, rather than starting a fresh LoRA.</p><figure class="highlight python"><table><tr><td class="code"><pre><span class="line">model = PeftModel.from_pretrained(base, args.output_dir, is_trainable=<span class="literal">True</span>)</span><br></pre></td></tr></table></figure><p>Omitting <code>is_trainable</code> loads the adapter frozen. Training then updates nothing, reports no error, and produces a complete run with a plausible log.</p></details><p>A fifth issue is unresolved rather than fixed. TRL 0.16's vLLM integration is server-based and wants a second GPU, and single-GPU colocation requires TRL ≥ 0.18. The runs therefore use HuggingFace <code>generate</code> on one L4 and accept the throughput. That decision is what the next section measures.</p><h2 id="the-hardware-measurement">the hardware measurement</h2><table><thead><tr><th>Setup</th><th>Time per step</th></tr></thead><tbody><tr><td>M4 Pro laptop (MPS)</td><td>~77 s</td></tr><tr><td>Modal L4 (CUDA, bf16)</td><td>~73 s</td></tr></tbody></table><p>The rented datacenter GPU was roughly 5% faster than the laptop. At 300 steps that is about 6 to 6.5 hours either way.</p><p>This is a single paired observation (one configuration, one run on each device, no repetitions), so it carries no interval and should not be read as a benchmark. What it does support is a directional claim with a mechanism behind it. Per-step time is dominated by generation overhead in HuggingFace <code>generate</code>, not by matrix multiplication. Sampling 8 completions of up to 640 tokens is a long serial sequence of small, latency-bound decode steps, and a wider GPU does not shorten a serial loop.</p><p>The consequences follow from the mechanism rather than from the 5%. Larger GPUs cost proportionally more per hour for roughly the same wall-clock, and a T4 is a false economy because it lacks bf16 and loses more time than it saves in price. The real fix is not more silicon but batched inference through vLLM, continuous batching and paged attention, which is blocked on the TRL version above.</p><h2 id="what-exists">what exists</h2><table><thead><tr><th>Component</th><th>State</th></tr></thead><tbody><tr><td>Data loaders (GSM8K, MATH, AIME → <code>{prompt, answer}</code>)</td><td>Done</td></tr><tr><td>Verifier and reward-variance diagnostic</td><td>Done</td></tr><tr><td>GRPO training with LoRA, adapter-resume across runs</td><td>Done</td></tr><tr><td><code>pass@1</code> / <code>pass@k</code> evaluation, base and adapted</td><td>Done</td></tr><tr><td>Modal deployment, persistent volume, wandb</td><td>Done</td></tr><tr><td>300-step training run</td><td>Not launched</td></tr><tr><td>Before/after results</td><td>None</td></tr></tbody></table><h2 id="conclusions">conclusions</h2><p>Had we reported only that the pipeline was built, this post would have described a working RLVR stack. The honest version is that the stack is built and untested against its own objective, and that everything of value so far came out of it refusing to work.</p><ul><li><strong>Reward variance is the metric to watch, not reward mean.</strong> Printing <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>σ</mi></mrow><annotation encoding="application/x-tex">\sigma</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord mathnormal" style="margin-right:0.0359em;">σ</span></span></span></span> every batch is ten lines and converts the most likely failure from invisible to obvious.</li><li><strong>Too-easy and too-hard are the same observation from outside.</strong> Any diagnosis that checks only whether reward is increasing cannot separate them, and they have opposite fixes.</li><li><strong>The verifier is the objective.</strong> A string-equality check is a biased reward function, not an approximate one.</li><li><strong>Every failure here was silent.</strong> EOS/pad mismatch, non-terminating rollouts, generation in train mode, a frozen adapter. Four bugs, zero exceptions.</li><li><strong>Measure before renting.</strong> For this workload the GPU bought about 5%, on one paired observation, because the bottleneck is serial decode rather than arithmetic.</li><li><strong>The difficulty band is unmeasured.</strong> Levels 3–5 produce nonzero variance; we have not swept difficulty against <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mtext>pass@</mtext><mn>8</mn></mrow><annotation encoding="application/x-tex">\text{pass@}8</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.8889em;vertical-align:-0.1944em;"></span><span class="mord text"><span class="mord">pass@</span></span><span class="mord">8</span></span></span></span>, so the curriculum is a working choice rather than a tuned one.</li></ul><p>The open question is whether 300 steps on a 500M model moves <code>pass@1</code> measurably at all. The frame we will judge it against is the <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mtext>pass@</mtext><mi>k</mi><mo>≫</mo><mtext>pass@</mtext><mn>1</mn></mrow><annotation encoding="application/x-tex">\text{pass@}k \gg \text{pass@}1</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.8889em;vertical-align:-0.1944em;"></span><span class="mord text"><span class="mord">pass@</span></span><span class="mord mathnormal" style="margin-right:0.0315em;">k</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">≫</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.8889em;vertical-align:-0.1944em;"></span><span class="mord text"><span class="mord">pass@</span></span><span class="mord">1</span></span></span></span> gap. A model that solves a problem 1 time in 8 already contains the capability, and RL's job is to shift probability mass onto reasoning it can already occasionally produce rather than to teach it something new. If the gap does not narrow, the report will say that it did not.</p>]]>
    </content>
    <id>https://projects.farhansadeek.com/posts/rlvr-grpo-reward-variance/</id>
    <link href="https://projects.farhansadeek.com/posts/rlvr-grpo-reward-variance/"/>
    <published>2026-09-19T00:15:07.000Z</published>
    <summary>An RLVR pipeline on Qwen2.5-0.5B-Instruct, published before the training run. GRPO learns from the spread of rewards rather than their level, and four bugs never raised an exception.</summary>
    <title>GRPO on Qwen2.5-0.5B and the Prompts That Teach Nothing</title>
    <updated>2026-09-29T06:49:26.552Z</updated>
  </entry>
  <entry>
    <author>
      <name>Farhan Sadeek</name>
      <email>farhan@farhansadeek.com</email>
    </author>
    <category term="research" scheme="https://projects.farhansadeek.com/tags/research/"/>
    <category term="transformers" scheme="https://projects.farhansadeek.com/tags/transformers/"/>
    <category term="interpretability" scheme="https://projects.farhansadeek.com/tags/interpretability/"/>
    <content>
      <![CDATA[<p>We score all 144 attention heads of GPT-2 small on the induction diagonal and recover the five heads the literature names. They are <strong>L5H5 (0.925), L6H9 (0.915), L7H10 (0.905), L5H1 (0.902), and L7H2 (0.820)</strong>. The sixth-ranked head scores 0.529. That gap of 0.290 is what makes &quot;the induction heads&quot; a set rather than a cutoff someone chose.</p><p>The measurement is one cached forward pass over eight sequences of 101 tokens, and it runs in under a minute on an M4 Pro through MPS. No cluster, no API keys. The model is GPT-2 small (124M) via <a href="https://github.com/TransformerLensOrg/TransformerLens">TransformerLens</a>.</p><p>This is experiment 01 of six, and it is the weakest kind of evidence in the plan it opens. An induction score is a correlation between a head's attention pattern and a behavior. It cannot separate a head that performs in-context copying from a head that attends to the right position for some unrelated reason while another component does the copying. We report these scores as a validated instrument reading, not as a circuit claim. The four experiments that would make it a circuit claim are specified below and have not been run.</p><p>Code, figures, and result tables are in <code>projects/tiny-circuits</code>. Every figure and table is produced by a numbered script and regenerates with one command.</p><p><em>Reading note.</em> The argument is the gap and what it does not license. The configuration details, the TransformerLens processing notes, and the experiment plan are self-contained. Skip to <a href="#conclusions">conclusions</a> if you only want the claim and its limits.</p><h2 id="table-of-contents">table of contents</h2><ul><li><a href="#the-behavior">the behavior</a></li><li><a href="#why-a-repeated-random-sequence">why a repeated random sequence</a></li><li><a href="#the-induction-score">the induction score</a></li><li><a href="#results">results</a></li><li><a href="#what-the-gap-means">what the gap means</a></li><li><a href="#what-the-ranking-does-not-establish">what the ranking does not establish</a></li><li><a href="#the-remaining-five-experiments">the remaining five experiments</a></li><li><a href="#instrument-notes">instrument notes</a></li><li><a href="#reproducibility">reproducibility</a></li><li><a href="#conclusions">conclusions</a></li></ul><h2 id="the-behavior">the behavior</h2><p>Show a language model <code>... A B ... A</code> and it predicts <code>B</code>. It has not seen that pair in training; it learned the rule, which is to look back for where this token last occurred and emit whatever followed it. The standard argument is that this is the substrate of in-context learning generally, which is why the circuit that implements it is the usual first target.</p><p>The claimed implementation is two heads composed across layers.</p><ol><li>A <strong>previous-token head</strong>, typically in layer 0, writes &quot;the token before me was A&quot; into each position's residual stream.</li><li>An <strong>induction head</strong>, in a middle layer, queries for that signature. It finds the position whose <em>previous</em> token matches the current token, and its OV circuit copies what that position holds into the output.</li></ol><p>Two heads and one edge between them. Every part of that claim is separately testable, and this experiment tests the weakest part of it, whether heads exist whose attention goes where the story says it should.</p><details class="collapsible-section"><summary><strong>QK and OV, where a head attends versus what it moves</strong></summary><p>An attention head factors into two circuits that can be analyzed separately, and the separation is what makes &quot;which head does what&quot; a tractable question.</p><p>The <strong>QK circuit</strong> decides <em>where</em> to attend. It is the bilinear form <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>W</mi><mi>Q</mi></msub><msubsup><mi>W</mi><mi>K</mi><mi mathvariant="normal">⊤</mi></msubsup></mrow><annotation encoding="application/x-tex">W_Q W_K^{\top}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1.1352em;vertical-align:-0.2861em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.1389em;">W</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3283em;"><span style="top:-2.55em;margin-left:-0.1389em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight">Q</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.2861em;"><span></span></span></span></span></span></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.1389em;">W</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.8491em;"><span style="top:-2.4247em;margin-left:-0.1389em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight" style="margin-right:0.0715em;">K</span></span></span><span style="top:-3.063em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mtight">⊤</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.2753em;"><span></span></span></span></span></span></span></span></span></span> acting on pairs of residual-stream vectors. Given a query position and a key position, it produces the pre-softmax score. Everything about attention <em>placement</em> is in this matrix.</p><p>The <strong>OV circuit</strong> decides <em>what gets moved</em> once a position is attended to. It is <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>W</mi><mi>O</mi></msub><msub><mi>W</mi><mi>V</mi></msub></mrow><annotation encoding="application/x-tex">W_O W_V</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.8333em;vertical-align:-0.15em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.1389em;">W</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3283em;"><span style="top:-2.55em;margin-left:-0.1389em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight" style="margin-right:0.0278em;">O</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.1389em;">W</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3283em;"><span style="top:-2.55em;margin-left:-0.1389em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight" style="margin-right:0.2222em;">V</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span></span></span></span>, and composed with the embedding and unembedding it becomes <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>W</mi><mi>U</mi></msub><msub><mi>W</mi><mrow><mi>O</mi><mi>V</mi></mrow></msub><msub><mi>W</mi><mi>E</mi></msub></mrow><annotation encoding="application/x-tex">W_U W_{OV} W_E</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.8333em;vertical-align:-0.15em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.1389em;">W</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3283em;"><span style="top:-2.55em;margin-left:-0.1389em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight" style="margin-right:0.109em;">U</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.1389em;">W</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3283em;"><span style="top:-2.55em;margin-left:-0.1389em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight" style="margin-right:0.0278em;">O</span><span class="mord mathnormal mtight" style="margin-right:0.2222em;">V</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.1389em;">W</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3283em;"><span style="top:-2.55em;margin-left:-0.1389em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight" style="margin-right:0.0576em;">E</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span></span></span></span>, a map from &quot;the token at the attended position&quot; to &quot;the change in output logits.&quot; A head that copies is one whose OV circuit is approximately diagonal in token space. Attending to token <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>t</mi></mrow><annotation encoding="application/x-tex">t</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6151em;"></span><span class="mord mathnormal">t</span></span></span></span> raises the logit of <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>t</mi></mrow><annotation encoding="application/x-tex">t</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6151em;"></span><span class="mord mathnormal">t</span></span></span></span>.</p><p>The split matters for this project because an induction score only measures the QK side. It says the head looks in the right place. It says nothing about whether the OV side moves anything useful, which is why experiment 05 examines <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>W</mi><mi>U</mi></msub><msub><mi>W</mi><mrow><mi>O</mi><mi>V</mi></mrow></msub><msub><mi>W</mi><mi>E</mi></msub></mrow><annotation encoding="application/x-tex">W_U W_{OV} W_E</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.8333em;vertical-align:-0.15em;"></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.1389em;">W</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3283em;"><span style="top:-2.55em;margin-left:-0.1389em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight" style="margin-right:0.109em;">U</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.1389em;">W</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3283em;"><span style="top:-2.55em;margin-left:-0.1389em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight" style="margin-right:0.0278em;">O</span><span class="mord mathnormal mtight" style="margin-right:0.2222em;">V</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.1389em;">W</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3283em;"><span style="top:-2.55em;margin-left:-0.1389em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight" style="margin-right:0.0576em;">E</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span></span></span></span> directly from the weights.</p></details><h2 id="why-a-repeated-random-sequence">why a repeated random sequence</h2><p>Measuring induction on natural text conflates the thing we want with three things we don't. A head that attends &quot;correctly&quot; on English may be exploiting grammar, semantics, or a memorized n-gram rather than repeat structure, and the attention pattern looks the same in all four cases.</p><p>The input removes every alternative by construction.</p><figure class="highlight python"><table><tr><td class="code"><pre><span class="line">rand = torch.randint(<span class="number">0</span>, vocab, (batch, seq_len), generator=g)</span><br><span class="line">tokens = torch.cat([bos, rand, rand], dim=<span class="number">1</span>)</span><br></pre></td></tr></table></figure><p><code>[BOS][50 random tokens][the same 50 tokens]</code>. There is no grammar, no semantics, and no n-gram statistics worth exploiting. The only structure available is that the second half repeats the first, so a head scoring highly here cannot be scoring highly for a reason we failed to think of.</p><p>This is the load-bearing design decision in the experiment. It buys a clean reading. The score means one thing. It costs external validity. Every conclusion below is about behavior on inputs GPT-2 never saw in training, and a head specialized for this diagonal on random tokens is not yet shown to do anything on English.</p><h2 id="the-induction-score">the induction score</h2><p>Under that construction, a head doing induction at position <code>i</code> in the second copy must attend to position <code>i - 50 + 1</code>, the token that followed the previous occurrence of the current token. That is one fixed diagonal of the attention pattern, so the score is the mean weight along it.</p><p>For a head with attention weights <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msubsup><mi>A</mi><mrow><mi>q</mi><mo separator="true">,</mo><mi>k</mi></mrow><mrow><mo stretchy="false">(</mo><mi>b</mi><mo stretchy="false">)</mo></mrow></msubsup></mrow><annotation encoding="application/x-tex">A^{(b)}_{q,k}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1.4822em;vertical-align:-0.4374em;"></span><span class="mord"><span class="mord mathnormal">A</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:1.0448em;"><span style="top:-2.3987em;margin-left:0em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight" style="margin-right:0.0359em;">q</span><span class="mpunct mtight">,</span><span class="mord mathnormal mtight" style="margin-right:0.0315em;">k</span></span></span></span><span style="top:-3.2198em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mopen mtight">(</span><span class="mord mathnormal mtight">b</span><span class="mclose mtight">)</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.4374em;"><span></span></span></span></span></span></span></span></span></span> on sequence <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>b</mi></mrow><annotation encoding="application/x-tex">b</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6944em;"></span><span class="mord mathnormal">b</span></span></span></span> of length <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mn>2</mn><mi>N</mi><mo>+</mo><mn>1</mn></mrow><annotation encoding="application/x-tex">2N+1</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.7667em;vertical-align:-0.0833em;"></span><span class="mord">2</span><span class="mord mathnormal" style="margin-right:0.109em;">N</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">+</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:0.6444em;"></span><span class="mord">1</span></span></span></span>, the induction score is the mean weight on the induction diagonal, taken over the batch <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>B</mi></mrow><annotation encoding="application/x-tex">B</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal" style="margin-right:0.0502em;">B</span></span></span></span> and over the <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>N</mi></mrow><annotation encoding="application/x-tex">N</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6833em;"></span><span class="mord mathnormal" style="margin-right:0.109em;">N</span></span></span></span> query positions in the second copy.</p><p class="katex-block"><span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mtext>score</mtext><mo>=</mo><mfrac><mn>1</mn><mrow><mi mathvariant="normal">∣</mi><mi>B</mi><mi mathvariant="normal">∣</mi><mtext> </mtext><mi>N</mi></mrow></mfrac><munder><mo>∑</mo><mrow><mi>b</mi><mo>∈</mo><mi>B</mi></mrow></munder><munderover><mo>∑</mo><mrow><mi>j</mi><mo>=</mo><mn>0</mn></mrow><mrow><mi>N</mi><mo>−</mo><mn>1</mn></mrow></munderover><msubsup><mi>A</mi><mrow><mtext> </mtext><mn>1</mn><mo>+</mo><mi>N</mi><mo>+</mo><mi>j</mi><mo separator="true">,</mo><mtext>  </mtext><mn>2</mn><mo>+</mo><mi>j</mi></mrow><mrow><mo stretchy="false">(</mo><mi>b</mi><mo stretchy="false">)</mo></mrow></msubsup></mrow><annotation encoding="application/x-tex">\text{score} = \frac{1}{|B|\,N}\sum_{b \in B} \sum_{j=0}^{N-1} A^{(b)}_{\,1+N+j,\; 2+j}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.4306em;"></span><span class="mord text"><span class="mord">score</span></span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:3.2421em;vertical-align:-1.4138em;"></span><span class="mord"><span class="mopen nulldelimiter"></span><span class="mfrac"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:1.3214em;"><span style="top:-2.314em;"><span class="pstrut" style="height:3em;"></span><span class="mord"><span class="mord">∣</span><span class="mord mathnormal" style="margin-right:0.0502em;">B</span><span class="mord">∣</span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mord mathnormal" style="margin-right:0.109em;">N</span></span></span><span style="top:-3.23em;"><span class="pstrut" style="height:3em;"></span><span class="frac-line" style="border-bottom-width:0.04em;"></span></span><span style="top:-3.677em;"><span class="pstrut" style="height:3em;"></span><span class="mord"><span class="mord">1</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.936em;"><span></span></span></span></span></span><span class="mclose nulldelimiter"></span></span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mop op-limits"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:1.05em;"><span style="top:-1.8479em;margin-left:0em;"><span class="pstrut" style="height:3.05em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight">b</span><span class="mrel mtight">∈</span><span class="mord mathnormal mtight" style="margin-right:0.0502em;">B</span></span></span></span><span style="top:-3.05em;"><span class="pstrut" style="height:3.05em;"></span><span><span class="mop op-symbol large-op">∑</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:1.3295em;"><span></span></span></span></span></span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mop op-limits"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:1.8283em;"><span style="top:-1.8723em;margin-left:0em;"><span class="pstrut" style="height:3.05em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight" style="margin-right:0.0572em;">j</span><span class="mrel mtight">=</span><span class="mord mtight">0</span></span></span></span><span style="top:-3.05em;"><span class="pstrut" style="height:3.05em;"></span><span><span class="mop op-symbol large-op">∑</span></span></span><span style="top:-4.3em;margin-left:0em;"><span class="pstrut" style="height:3.05em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight" style="margin-right:0.109em;">N</span><span class="mbin mtight">−</span><span class="mord mtight">1</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:1.4138em;"><span></span></span></span></span></span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mord"><span class="mord mathnormal">A</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:1.0448em;"><span style="top:-2.4065em;margin-left:0em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mspace mtight" style="margin-right:0.1952em;"></span><span class="mord mtight">1</span><span class="mbin mtight">+</span><span class="mord mathnormal mtight" style="margin-right:0.109em;">N</span><span class="mbin mtight">+</span><span class="mord mathnormal mtight" style="margin-right:0.0572em;">j</span><span class="mpunct mtight">,</span><span class="mspace mtight" style="margin-right:0.3253em;"></span><span class="mord mtight">2</span><span class="mbin mtight">+</span><span class="mord mathnormal mtight" style="margin-right:0.0572em;">j</span></span></span></span><span style="top:-3.2198em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mopen mtight">(</span><span class="mord mathnormal mtight">b</span><span class="mclose mtight">)</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.4296em;"><span></span></span></span></span></span></span></span></span></span></span></p><p>Query <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mn>1</mn><mo>+</mo><mi>N</mi><mo>+</mo><mi>j</mi></mrow><annotation encoding="application/x-tex">1+N+j</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.7278em;vertical-align:-0.0833em;"></span><span class="mord">1</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">+</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:0.7667em;vertical-align:-0.0833em;"></span><span class="mord mathnormal" style="margin-right:0.109em;">N</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">+</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:0.854em;vertical-align:-0.1944em;"></span><span class="mord mathnormal" style="margin-right:0.0572em;">j</span></span></span></span> sits at offset <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>j</mi></mrow><annotation encoding="application/x-tex">j</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.854em;vertical-align:-0.1944em;"></span><span class="mord mathnormal" style="margin-right:0.0572em;">j</span></span></span></span> into the repeated half; it holds the same token as position <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mn>1</mn><mo>+</mo><mi>j</mi></mrow><annotation encoding="application/x-tex">1+j</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.7278em;vertical-align:-0.0833em;"></span><span class="mord">1</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">+</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:0.854em;vertical-align:-0.1944em;"></span><span class="mord mathnormal" style="margin-right:0.0572em;">j</span></span></span></span>, so the key it must attend to is the token that <em>followed</em> that first occurrence, at position <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mn>2</mn><mo>+</mo><mi>j</mi></mrow><annotation encoding="application/x-tex">2+j</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.7278em;vertical-align:-0.0833em;"></span><span class="mord">2</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">+</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:0.854em;vertical-align:-0.1944em;"></span><span class="mord mathnormal" style="margin-right:0.0572em;">j</span></span></span></span>. The difference between those indices is constant, which is why the whole measurement is one diagonal at offset <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mo>−</mo><mo stretchy="false">(</mo><mi>N</mi><mo>−</mo><mn>1</mn><mo stretchy="false">)</mo></mrow><annotation encoding="application/x-tex">-(N-1)</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mord">−</span><span class="mopen">(</span><span class="mord mathnormal" style="margin-right:0.109em;">N</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">−</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mord">1</span><span class="mclose">)</span></span></span></span>.</p><figure class="highlight python"><table><tr><td class="code"><pre><span class="line">pattern = cache[<span class="string">&quot;pattern&quot;</span>, layer]                     <span class="comment"># [batch, head, query, key]</span></span><br><span class="line">diag = pattern.diagonal(offset=-(seq_len - <span class="number">1</span>), dim1=-<span class="number">2</span>, dim2=-<span class="number">1</span>)</span><br><span class="line">diag = diag[..., -seq_len:]                           <span class="comment"># second-half queries only</span></span><br><span class="line">scores[layer] = diag.mean(dim=(<span class="number">0</span>, -<span class="number">1</span>))                <span class="comment"># mean over batch and diagonal</span></span><br></pre></td></tr></table></figure><p>The configuration is sequence length 50, batch 8, seed 1337, one cached forward pass, 144 numbers out.</p><p>We do not report an interval on these scores. With a single seed and a batch of 8, the run supports the ranking and the size of the rank-5 to rank-6 gap; it does not support the third decimal place. The per-head scores are stable across seeds in informal checks, but a seed sweep is a ten-line change that has not been run, so that stability is an impression rather than a measurement. Read the table as an ordering.</p><h2 id="results">results</h2><figure data-figure="induction-heatmap"></figure><table><thead><tr><th>Head</th><th>Induction score</th></tr></thead><tbody><tr><td>L5H5</td><td>0.925</td></tr><tr><td>L6H9</td><td>0.915</td></tr><tr><td>L7H10</td><td>0.905</td></tr><tr><td>L5H1</td><td>0.902</td></tr><tr><td>L7H2</td><td>0.820</td></tr><tr><td>L10H1</td><td>0.529</td></tr><tr><td>L9H6</td><td>0.509</td></tr><tr><td>L9H9</td><td>0.500</td></tr><tr><td>L10H7</td><td>0.459</td></tr><tr><td>L5H0</td><td>0.436</td></tr></tbody></table><p><em>Correction, September 26, 2026.</em> An earlier version averaged the diagonal over two extra first-half query positions, which have no earlier copy to look back to. That diluted every score by about 2% and put L5H1 above L7H10. The five-head set and the size of the gap did not change. The formula above also had an off-by-one in the key index, now fixed; the code was already correct on that point.</p><figure data-figure="induction-ranked"></figure><h2 id="what-the-gap-means">what the gap means</h2><p>Attention is a probability distribution over the entire 101-token context. A score of 0.925 means L5H5 places about 92% of its attention mass on a single position, and that position is determined purely by a structural rule in an input with no content to attend to. This is not a tendency toward induction. It is a head that does one job.</p><p>Five heads sit above 0.80 and the sixth sits at 0.53. The 0.290 gap between rank 5 and rank 6 is the reason the set is well defined. No threshold has to be chosen, because the data separates. Had the scores decayed smoothly from 0.9 to 0.4, &quot;the induction heads&quot; would name whatever cutoff we picked, and every downstream experiment would inherit that arbitrariness.</p><p>The second feature of the table is location. The five heads sit in layers 5 through 7 of a 12-layer model, downstream of the layer-0 previous-token heads the circuit needs as input, with enough depth remaining to influence the output. The theory predicts where these heads should live, and that is where they are. A prediction that could have failed and did not is worth more than the scores themselves, because the scores were always going to be high for <em>something</em>.</p><h2 id="what-the-ranking-does-not-establish">what the ranking does not establish</h2><p>Everything above is consistent with the five heads being passengers. A head can attend to exactly the right position for an unrelated reason while some other component performs the copy, and no attention-pattern measurement can tell the difference. Four objections remain open.</p><table><thead><tr><th>Objection</th><th>Experiment</th><th>Method</th></tr></thead><tbody><tr><td>Correlation, not causation</td><td>03</td><td>Activation patching, clean vs. corrupted repeated sequences</td></tr><tr><td>Composition is assumed, not shown</td><td>04</td><td>Path patching on the prev-token → induction edge</td></tr><tr><td>The behavior may be input-specific</td><td>05</td><td>Eigenvalues and diagonal dominance of <code>W_U W_OV W_E</code></td></tr><tr><td>Other heads may be redundant backups</td><td>06</td><td>Zero and mean ablation, head knockout sweep</td></tr></tbody></table><p>The distinction between 03 and 05 is the one that matters most. Activation patching is a statement about this model on these inputs; showing that the OV circuit is approximately a copy matrix is a statement about the <em>parameters</em>, independent of whatever inputs we happened to construct. Input-level evidence and weight-level evidence fail in different ways, and a circuit claim is only strong when both hold.</p><p>Experiment 06 is the one most likely to produce an awkward result. Ablation studies in this literature routinely find that knocking out a &quot;necessary&quot; component degrades behavior far less than the discovery evidence suggests, because other heads absorb the loss. If that happens here it is a finding about the circuit rather than a defect in the measurement, but it would complicate the clean two-head story considerably. We would rather name that possibility now than after the fact.</p><h2 id="the-remaining-five-experiments">the remaining five experiments</h2><p>The plan, with the evidence each step is meant to produce.</p><table><thead><tr><th>Experiment</th><th>Question</th><th>Evidence it produces</th></tr></thead><tbody><tr><td>02</td><td>Does the attention pattern look the way the score implies?</td><td>Per-head attention visualization on the repeated sequence</td></tr><tr><td>03</td><td>Do these heads cause the copying?</td><td>Recovery of logit difference under activation patching</td></tr><tr><td>04</td><td>Is the prev-token → induction edge real?</td><td>Path patching isolating that specific edge</td></tr><tr><td>05</td><td>Is the OV circuit a copy matrix?</td><td>Spectrum and diagonal dominance of <code>W_U W_OV W_E</code></td></tr><tr><td>06</td><td>Are the heads necessary?</td><td>Behavior collapse under zero and mean ablation</td></tr></tbody></table><p>Only 02 is a refinement of what is already measured. Experiments 03 through 06 are the ones that can overturn the reading above, and none of them has been run.</p><h2 id="instrument-notes">instrument notes</h2><details class="collapsible-section"><summary><strong>What TransformerLens changes about the model before you measure it</strong></summary><p>TransformerLens loads GPT-2 with LayerNorm folding, centered writing weights, and centered unembedding on by default. These change the parameterization without changing the function the model computes, and they are what make residual-stream and logit-lens analysis clean. A score computed on an unprocessed model is not being compared against the same object the published results describe.</p></details><p>This is why reproducing a known result was the point of experiment 01 rather than a formality. L5H5 and L6H9 are in the literature; we did not find them, we recovered them, with our own metric implementation, our own input construction, and our own weight processing. A pipeline that recovers the canonical heads from scratch is a pipeline that can be pointed at questions with no answer key. That validation step is usually invisible in the write-up, and it is the step that decides whether any later number means anything.</p><h2 id="reproducibility">reproducibility</h2><p>Every experiment is a self-contained numbered script with a documented research question. Figures and result tables are committed, are generated only by those scripts, and regenerate with this command.</p><figure class="highlight bash"><table><tr><td class="code"><pre><span class="line">uv run python experiments/01_induction_scores.py</span><br></pre></td></tr></table></figure><p>Optional <code>--wandb</code> logs the run, the figure, and the score table. The environment is a <code>uv</code> project pinned to Python 3.12; the system Python is 3.14, which the interpretability stack does not yet support.</p><p>The constraint that every figure must come from a committed script has already caught two errors that a notebook would have hidden.</p><h2 id="conclusions">conclusions</h2><p>Had we stopped at the score table, this post would have told a circuit-discovery story. Recovering L5H5 and L6H9 with an independent implementation is a validated instrument reading and nothing more. It establishes that the measurement is correct on a case with a published answer, which is the precondition for trusting it anywhere else.</p><ul><li>The five-head set is sharply defined. The 0.290 gap between rank 5 and rank 6 is in the data, not in a chosen threshold.</li><li>The heads sit in layers 5 through 7, where the composition story says they must. That prediction could have failed.</li><li>Nothing here separates &quot;these heads participate in induction&quot; from &quot;these heads cause induction.&quot; Patching and ablation separate them, and they are not run.</li><li>The reported scores support an ordering, not three decimal places. No seed sweep has been run.</li><li>Laptop-scale interpretability on a real pretrained model is genuinely accessible. The whole experiment is one cached forward pass over eight sequences of 101 tokens.</li></ul><p><em>Next.</em> The four experiments are now run, in <a href="/posts/testing-gpt-2s-induction-heads/">Testing GPT-2's Induction Heads and the Backups Behind Them</a>.</p>]]>
    </content>
    <id>https://projects.farhansadeek.com/posts/reverse-engineering-gpt-2s-induction-circuit/</id>
    <link href="https://projects.farhansadeek.com/posts/reverse-engineering-gpt-2s-induction-circuit/"/>
    <published>2026-09-18T22:33:41.000Z</published>
    <summary>All 144 attention heads of GPT-2 small scored on the induction diagonal in one cached forward pass on a laptop, recovering the five canonical heads. An induction score is a correlation, not a cause.</summary>
    <title>Finding GPT-2's Induction Heads in One Forward Pass</title>
    <updated>2026-09-29T06:49:26.552Z</updated>
  </entry>
  <entry>
    <author>
      <name>Farhan Sadeek</name>
      <email>farhan@farhansadeek.com</email>
    </author>
    <category term="transformers" scheme="https://projects.farhansadeek.com/tags/transformers/"/>
    <category term="pytorch" scheme="https://projects.farhansadeek.com/tags/pytorch/"/>
    <content>
      <![CDATA[<p>We trained a decoder-only GPT written from scratch, with no <code>nn.Transformer</code>, no <code>F.scaled_dot_product_attention</code> and no HuggingFace model code, to a final cross-entropy of <strong>2.14</strong> on TinyStories, in 20,000 steps on a single Apple Silicon laptop. The shipped checkpoint is <strong>5,263,848 parameters</strong>, configured as 6 pre-LN blocks, 8 heads, a 256-dimensional residual stream, a 64-token context, and a 1,000-token byte-level BPE vocabulary trained on the corpus itself.</p><p>The loss is not a result. It is a receipt, evidence that every stage of the pipeline, from tokenizer and cache through loader, model, loss and sampler, is wired correctly end to end, which was the entire goal. The run consumed 20,000 × 8 × 64 = 10.2M tokens against a cached corpus of roughly 632M tokens, so the model saw about 1.6% of the available data in a single pass. It is undertrained by construction, and no comparison to any published loss is meaningful.</p><p>What the project did produce is a clear view of which lines are load-bearing. Two constants decide silently whether the model trains at all, and two bugs cost more time than the model code did.</p><p>Code is in <code>projects/gpt-from-scratch</code>, which holds the tokenizer, streaming data pipeline, model, training loop, sampler, chat REPL, and a Modal script for the GPU runs.</p><p><em>Reading note.</em> The argument is that the naive implementation makes the mechanism visible, and that the expensive failures were all silent. Architecture tables, the cache-invalidation logic, and the Modal harness are collapsed.</p><h2 id="table-of-contents">table of contents</h2><ul><li><a href="#why-write-it-out-at-all">why write it out at all</a></li><li><a href="#the-model">the model</a></li><li><a href="#attention-is-three-linear-maps-and-a-mask">attention is three linear maps and a mask</a></li><li><a href="#the-scale-factor-is-not-a-detail">the scale factor is not a detail</a></li><li><a href="#multi-head-implemented-the-slow-way-on-purpose">multi-head, implemented the slow way on purpose</a></li><li><a href="#the-residual-stream-is-the-actual-abstraction">the residual stream is the actual abstraction</a></li><li><a href="#the-tokenizer-and-the-data-pipeline">the tokenizer and the data pipeline</a></li><li><a href="#training-and-how-much-data-the-run-actually-saw">training, and how much data the run actually saw</a></li><li><a href="#sampling">sampling</a></li><li><a href="#scaling-to-a-gpu-and-what-it-bought">scaling to a GPU, and what it bought</a></li><li><a href="#what-cost-the-most-time">what cost the most time</a></li><li><a href="#conclusions">conclusions</a></li></ul><h2 id="why-write-it-out-at-all">why write it out at all</h2><p>Reading the transformer paper and reading a reference implementation both leave the same gap. You can follow every line and still not know which lines are <em>load-bearing</em>. Typing it out closes that gap by force. Every constant omitted and every shape gotten wrong produces either a crash or, worse, a model that trains to nothing while looking fine.</p><h2 id="the-model">the model</h2><figure data-figure="gpt-parameters"></figure><details class="collapsible-section"><summary><strong>Full architecture, as configured in the shipped checkpoint</strong></summary><table><thead><tr><th>Component</th><th>Value</th></tr></thead><tbody><tr><td>Blocks</td><td>6</td></tr><tr><td>Attention heads</td><td>8</td></tr><tr><td>Residual width (<code>n_embed</code>)</td><td>256</td></tr><tr><td>Head size</td><td>32 (<code>n_embed / num_heads</code>)</td></tr><tr><td>Context (<code>block_size</code>)</td><td>64</td></tr><tr><td>Vocabulary</td><td>1,000 (byte-level BPE)</td></tr><tr><td>MLP</td><td>4× expansion, ReLU</td></tr><tr><td>Position encoding</td><td>Learned embeddings</td></tr><tr><td>Normalization</td><td>Pre-LN, two <code>LayerNorm</code>s per block</td></tr><tr><td>Parameters</td><td>5,263,848</td></tr></tbody></table><p>The parameter count breaks down as 272,384 in the embeddings (256,000 token + 16,384 positional), 4,733,952 across the six blocks, 512 in the final <code>LayerNorm</code>, and 257,000 in the unembedding. The causal masks are registered buffers and are not counted, which is why the number printed by <code>sum(p.numel() for p in model.parameters())</code> excludes the 196,608 mask entries in the state dict.</p></details><h2 id="attention-is-three-linear-maps-and-a-mask">attention is three linear maps and a mask</h2><p>The single-head forward pass in full.</p><figure class="highlight python"><table><tr><td class="code"><pre><span class="line">keys, queries, values = <span class="variable language_">self</span>.key(x), <span class="variable language_">self</span>.query(x), <span class="variable language_">self</span>.value(x)</span><br><span class="line">scores = queries @ keys.transpose(-<span class="number">2</span>, -<span class="number">1</span>)</span><br><span class="line">scores = scores * (<span class="variable language_">self</span>.head_size ** -<span class="number">0.5</span>)</span><br><span class="line">scores = scores.masked_fill(<span class="variable language_">self</span>.mask[:T, :T] == <span class="number">0</span>, <span class="built_in">float</span>(<span class="string">&quot;-inf&quot;</span>))</span><br><span class="line"><span class="keyword">return</span> F.softmax(scores, dim=-<span class="number">1</span>) @ values</span><br></pre></td></tr></table></figure><p>Four operations. A token's query is a question, every other token's key is an advertisement, the dot product scores the match, and the value is what actually gets moved. The causal mask is not a deep architectural property. It is <code>torch.tril(torch.ones(block_size, block_size))</code> registered as a buffer and a <code>masked_fill</code>.</p><h2 id="the-scale-factor-is-not-a-detail">the scale factor is not a detail</h2><p><code>self.head_size ** -0.5</code> looks like a detail, and it decides whether the model trains.</p><p>Take a query <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>q</mi></mrow><annotation encoding="application/x-tex">q</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.625em;vertical-align:-0.1944em;"></span><span class="mord mathnormal" style="margin-right:0.0359em;">q</span></span></span></span> and key <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>k</mi></mrow><annotation encoding="application/x-tex">k</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6944em;"></span><span class="mord mathnormal" style="margin-right:0.0315em;">k</span></span></span></span> in <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msup><mi mathvariant="double-struck">R</mi><mi>d</mi></msup></mrow><annotation encoding="application/x-tex">\mathbb{R}^{d}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.8491em;"></span><span class="mord"><span class="mord mathbb">R</span><span class="msupsub"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.8491em;"><span style="top:-3.063em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight">d</span></span></span></span></span></span></span></span></span></span></span></span> with independent, zero-mean, unit-variance entries. Their dot product is a sum of <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>d</mi></mrow><annotation encoding="application/x-tex">d</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6944em;"></span><span class="mord mathnormal">d</span></span></span></span> independent terms, so</p><p class="katex-block"><span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mi mathvariant="normal">Var</mi><mo>⁡</mo><mo stretchy="false">(</mo><mi>q</mi><mo>⋅</mo><mi>k</mi><mo stretchy="false">)</mo><mo>=</mo><munderover><mo>∑</mo><mrow><mi>i</mi><mo>=</mo><mn>1</mn></mrow><mi>d</mi></munderover><mi mathvariant="normal">Var</mi><mo>⁡</mo><mo stretchy="false">(</mo><msub><mi>q</mi><mi>i</mi></msub><msub><mi>k</mi><mi>i</mi></msub><mo stretchy="false">)</mo><mo>=</mo><mi>d</mi><mo separator="true">,</mo></mrow><annotation encoding="application/x-tex">\operatorname{Var}(q \cdot k) = \sum_{i=1}^{d} \operatorname{Var}(q_i k_i) = d,</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mop"><span class="mord mathrm">Var</span></span><span class="mopen">(</span><span class="mord mathnormal" style="margin-right:0.0359em;">q</span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">⋅</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:1em;vertical-align:-0.25em;"></span><span class="mord mathnormal" style="margin-right:0.0315em;">k</span><span class="mclose">)</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:3.1138em;vertical-align:-1.2777em;"></span><span class="mop op-limits"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:1.8361em;"><span style="top:-1.8723em;margin-left:0em;"><span class="pstrut" style="height:3.05em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight">i</span><span class="mrel mtight">=</span><span class="mord mtight">1</span></span></span></span><span style="top:-3.05em;"><span class="pstrut" style="height:3.05em;"></span><span><span class="mop op-symbol large-op">∑</span></span></span><span style="top:-4.3em;margin-left:0em;"><span class="pstrut" style="height:3.05em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight">d</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:1.2777em;"><span></span></span></span></span></span><span class="mspace" style="margin-right:0.1667em;"></span><span class="mop"><span class="mord mathrm">Var</span></span><span class="mopen">(</span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0359em;">q</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3117em;"><span style="top:-2.55em;margin-left:-0.0359em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight">i</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0315em;">k</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3117em;"><span style="top:-2.55em;margin-left:-0.0315em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight">i</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span><span class="mclose">)</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.8889em;vertical-align:-0.1944em;"></span><span class="mord mathnormal">d</span><span class="mpunct">,</span></span></span></span></span></p><p>and the logits entering the softmax have standard deviation <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msqrt><mi>d</mi></msqrt></mrow><annotation encoding="application/x-tex">\sqrt{d}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1.04em;vertical-align:-0.1078em;"></span><span class="mord sqrt"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.9322em;"><span class="svg-align" style="top:-3em;"><span class="pstrut" style="height:3em;"></span><span class="mord" style="padding-left:0.833em;"><span class="mord mathnormal">d</span></span></span><span style="top:-2.8922em;"><span class="pstrut" style="height:3em;"></span><span class="hide-tail" style="min-width:0.853em;height:1.08em;"><svg xmlns="http://www.w3.org/2000/svg" width="400em" height="1.08em" viewBox="0 0 400000 1080" preserveAspectRatio="xMinYMin slice"><path d="M95,702c-2.7,0,-7.17,-2.7,-13.5,-8c-5.8,-5.3,-9.5,-10,-9.5,-14c0,-2,0.3,-3.3,1,-4c1.3,-2.7,23.83,-20.7,67.5,-54c44.2,-33.3,65.8,-50.3,66.5,-51c1.3,-1.3,3,-2,5,-2c4.7,0,8.7,3.3,12,10s173,378,173,378c0.7,0,35.3,-71,104,-213c68.7,-142,137.5,-285,206.5,-429c69,-144,104.5,-217.7,106.5,-221l0 -0c5.3,-9.3,12,-14,20,-14H400000v40H845.2724s-225.272,467,-225.272,467s-235,486,-235,486c-2.7,4.7,-9,7,-19,7c-6,0,-10,-1,-12,-3s-194,-422,-194,-422s-65,47,-65,47zM834 80h400000v40h-400000z"/></svg></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.1078em;"><span></span></span></span></span></span></span></span></span>. Dividing by <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msqrt><mi>d</mi></msqrt></mrow><annotation encoding="application/x-tex">\sqrt{d}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1.04em;vertical-align:-0.1078em;"></span><span class="mord sqrt"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.9322em;"><span class="svg-align" style="top:-3em;"><span class="pstrut" style="height:3em;"></span><span class="mord" style="padding-left:0.833em;"><span class="mord mathnormal">d</span></span></span><span style="top:-2.8922em;"><span class="pstrut" style="height:3em;"></span><span class="hide-tail" style="min-width:0.853em;height:1.08em;"><svg xmlns="http://www.w3.org/2000/svg" width="400em" height="1.08em" viewBox="0 0 400000 1080" preserveAspectRatio="xMinYMin slice"><path d="M95,702c-2.7,0,-7.17,-2.7,-13.5,-8c-5.8,-5.3,-9.5,-10,-9.5,-14c0,-2,0.3,-3.3,1,-4c1.3,-2.7,23.83,-20.7,67.5,-54c44.2,-33.3,65.8,-50.3,66.5,-51c1.3,-1.3,3,-2,5,-2c4.7,0,8.7,3.3,12,10s173,378,173,378c0.7,0,35.3,-71,104,-213c68.7,-142,137.5,-285,206.5,-429c69,-144,104.5,-217.7,106.5,-221l0 -0c5.3,-9.3,12,-14,20,-14H400000v40H845.2724s-225.272,467,-225.272,467s-235,486,-235,486c-2.7,4.7,-9,7,-19,7c-6,0,-10,-1,-12,-3s-194,-422,-194,-422s-65,47,-65,47zM834 80h400000v40h-400000z"/></svg></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.1078em;"><span></span></span></span></span></span></span></span></span> restores unit variance and keeps the softmax in a regime where it is not saturated.</p><p>Without the scale, logits grow with head size, the softmax collapses toward one-hot, and the gradient through it vanishes, because <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="normal">∂</mi><mtext>softmax</mtext></mrow><annotation encoding="application/x-tex">\partial \text{softmax}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6944em;"></span><span class="mord" style="margin-right:0.0556em;">∂</span><span class="mord text"><span class="mord">softmax</span></span></span></span></span> is proportional to <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>p</mi><mi>i</mi></msub><mo stretchy="false">(</mo><msub><mi>δ</mi><mrow><mi>i</mi><mi>j</mi></mrow></msub><mo>−</mo><msub><mi>p</mi><mi>j</mi></msub><mo stretchy="false">)</mo></mrow><annotation encoding="application/x-tex">p_i(\delta_{ij} - p_j)</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1.0361em;vertical-align:-0.2861em;"></span><span class="mord"><span class="mord mathnormal">p</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3117em;"><span style="top:-2.55em;margin-left:0em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight">i</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span><span class="mopen">(</span><span class="mord"><span class="mord mathnormal" style="margin-right:0.0379em;">δ</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3117em;"><span style="top:-2.55em;margin-left:-0.0379em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mathnormal mtight" style="margin-right:0.0572em;">ij</span></span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.2861em;"><span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">−</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:1.0361em;vertical-align:-0.2861em;"></span><span class="mord"><span class="mord mathnormal">p</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3117em;"><span style="top:-2.55em;margin-left:0em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight" style="margin-right:0.0572em;">j</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.2861em;"><span></span></span></span></span></span></span><span class="mclose">)</span></span></span></span>, which goes to zero as any <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>p</mi><mi>i</mi></msub><mo>→</mo><mn>1</mn></mrow><annotation encoding="application/x-tex">p_i \to 1</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.625em;vertical-align:-0.1944em;"></span><span class="mord"><span class="mord mathnormal">p</span><span class="msupsub"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.3117em;"><span style="top:-2.55em;margin-left:0em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mathnormal mtight">i</span></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.15em;"><span></span></span></span></span></span></span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">→</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.6444em;"></span><span class="mord">1</span></span></span></span>. The model still trains, the loss barely moves, and nothing in the code looks wrong. That combination is what makes it dangerous. The failure has no error message and no obviously guilty line.</p><p>At <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>d</mi><mo>=</mo><mn>32</mn></mrow><annotation encoding="application/x-tex">d = 32</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6944em;"></span><span class="mord mathnormal">d</span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.6444em;"></span><span class="mord">32</span></span></span></span> the factor is <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mn>1</mn><mi mathvariant="normal">/</mi><msqrt><mn>32</mn></msqrt><mo>≈</mo><mn>0.177</mn></mrow><annotation encoding="application/x-tex">1/\sqrt{32} \approx 0.177</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1.1572em;vertical-align:-0.25em;"></span><span class="mord">1/</span><span class="mord sqrt"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.9072em;"><span class="svg-align" style="top:-3em;"><span class="pstrut" style="height:3em;"></span><span class="mord" style="padding-left:0.833em;"><span class="mord">32</span></span></span><span style="top:-2.8672em;"><span class="pstrut" style="height:3em;"></span><span class="hide-tail" style="min-width:0.853em;height:1.08em;"><svg xmlns="http://www.w3.org/2000/svg" width="400em" height="1.08em" viewBox="0 0 400000 1080" preserveAspectRatio="xMinYMin slice"><path d="M95,702c-2.7,0,-7.17,-2.7,-13.5,-8c-5.8,-5.3,-9.5,-10,-9.5,-14c0,-2,0.3,-3.3,1,-4c1.3,-2.7,23.83,-20.7,67.5,-54c44.2,-33.3,65.8,-50.3,66.5,-51c1.3,-1.3,3,-2,5,-2c4.7,0,8.7,3.3,12,10s173,378,173,378c0.7,0,35.3,-71,104,-213c68.7,-142,137.5,-285,206.5,-429c69,-144,104.5,-217.7,106.5,-221l0 -0c5.3,-9.3,12,-14,20,-14H400000v40H845.2724s-225.272,467,-225.272,467s-235,486,-235,486c-2.7,4.7,-9,7,-19,7c-6,0,-10,-1,-12,-3s-194,-422,-194,-422s-65,47,-65,47zM834 80h400000v40h-400000z"/></svg></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.1328em;"><span></span></span></span></span></span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">≈</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.6444em;"></span><span class="mord">0.177</span></span></span></span>, so the unscaled logits would be roughly 5.7× larger, enough to saturate and not enough to look absurd if printed.</p><h2 id="multi-head-implemented-the-slow-way-on-purpose">multi-head, implemented the slow way on purpose</h2><figure class="highlight python"><table><tr><td class="code"><pre><span class="line"><span class="variable language_">self</span>.heads = nn.ModuleList([SelfAttention(...) <span class="keyword">for</span> _ <span class="keyword">in</span> <span class="built_in">range</span>(num_heads)])</span><br><span class="line"><span class="keyword">return</span> <span class="variable language_">self</span>.proj(torch.cat([head(x) <span class="keyword">for</span> head <span class="keyword">in</span> <span class="variable language_">self</span>.heads], dim=-<span class="number">1</span>))</span><br></pre></td></tr></table></figure><p>Production implementations fold all heads into one batched matmul. This one keeps them as independent modules and concatenates. That is measurably slower, and we kept it. Heads genuinely are independent subspaces, and writing them as separate objects makes that structural fact impossible to forget. Fusing is an optimization, not a concept, and the independence is exactly the property that later interpretability work depends on. The <a href="/posts/reverse-engineering-gpt-2s-induction-circuit">induction-circuit project</a> scores heads individually, which only means something because they <em>are</em> individual.</p><p>At 3.8M–5.3M parameters on a laptop the trade is free. At any serious scale it is not, and the right move is to fuse and keep a slow reference implementation for testing against.</p><h2 id="the-residual-stream-is-the-actual-abstraction">the residual stream is the actual abstraction</h2><figure class="highlight python"><table><tr><td class="code"><pre><span class="line">x = x + <span class="variable language_">self</span>.attention(<span class="variable language_">self</span>.layer_norm1(x))</span><br><span class="line">x = x + <span class="variable language_">self</span>.feed_forward(<span class="variable language_">self</span>.layer_norm2(x))</span><br></pre></td></tr></table></figure><p>Two <code>x + ...</code> lines, and they are the reason depth works at all. Each block <em>proposes an edit</em> to a running representation rather than replacing it. Normalization goes before the sublayer (pre-LN), so the residual path from input to output is unnormalized and gradients reach layer 0 intact.</p><p>Once the residual stream reads as a shared bus that every layer writes to and reads from, the later literature of logit lens, activation patching and circuits stops being exotic and starts being the obvious next question. That framing is what sent this project toward induction heads next.</p><h2 id="the-tokenizer-and-the-data-pipeline">the tokenizer and the data pipeline</h2><p>The least interesting part took the most iterations.</p><p>Byte-level BPE with a ByteLevel pre-tokenizer and decoder means <strong>no unknown tokens are possible</strong>, since every byte sequence encodes. That sounds like a footnote and is what makes the model robust to whatever the corpus contains. The vocabulary is trained on the corpus itself rather than borrowed, which at 1,000 merges over children's stories gives a tokenizer specialized to exactly this distribution.</p><details class="collapsible-section"><summary><strong>Streaming tokenization and cache invalidation</strong></summary><p>The corpus is 1.9 GB of text and stopped fitting comfortably in memory. Tokenization became a streaming pass that writes a <code>uint16</code> token cache to disk, invalidated by comparing modification times against both the source text and the tokenizer.</p><figure class="highlight python"><table><tr><td class="code"><pre><span class="line">cache_is_stale = <span class="keyword">not</span> token_cache.exists() <span class="keyword">or</span> token_cache.stat().st_mtime_ns &lt; <span class="built_in">max</span>(</span><br><span class="line">    args.data.stat().st_mtime_ns,</span><br><span class="line">    tokenizer_path.stat().st_mtime_ns,</span><br><span class="line">)</span><br></pre></td></tr></table></figure><p><code>uint16</code> because a 1,000-token vocabulary fits in 16 bits, which halves the cache versus <code>int32</code>. The resulting cache is 1.26 GB, or roughly 632M tokens.</p><p>The staleness check exists because the tokenizer was retrained once, and an hour of training then ran against tokens from the <em>previous</em> vocabulary, a corruption that produces no error, just a model learning a scrambled language.</p></details><h2 id="training-and-how-much-data-the-run-actually-saw">training, and how much data the run actually saw</h2><p>AdamW at its default weight decay of 0.01, learning rate 3e-4, batch size 8, context 64, seed 1337, 20,000 steps, checkpointing every 1,000. Checkpoints store the model state, optimizer state, the full config, the training arguments, and the final loss, so a resume restores the run rather than approximating it.</p><p>No learning-rate schedule, no warmup, no gradient clipping. At this scale none were necessary, and leaving them out kept the loop readable. At larger scale each is the next thing to add.</p><p>The number worth stating plainly is the token budget.</p><p class="katex-block"><span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mn>20,000</mn><mtext> steps</mtext><mo>×</mo><mn>8</mn><mtext> sequences</mtext><mo>×</mo><mn>64</mn><mtext> tokens</mtext><mo>=</mo><mn>10.24</mn><mtext>M tokens</mtext><mo separator="true">,</mo></mrow><annotation encoding="application/x-tex">20{,}000 \text{ steps} \times 8 \text{ sequences} \times 64 \text{ tokens} = 10.24\text{M tokens},</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.8389em;vertical-align:-0.1944em;"></span><span class="mord">20</span><span class="mord"><span class="mpunct">,</span></span><span class="mord">000</span><span class="mord text"><span class="mord"> steps</span></span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">×</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:0.8389em;vertical-align:-0.1944em;"></span><span class="mord">8</span><span class="mord text"><span class="mord"> sequences</span></span><span class="mspace" style="margin-right:0.2222em;"></span><span class="mbin">×</span><span class="mspace" style="margin-right:0.2222em;"></span></span><span class="base"><span class="strut" style="height:0.6944em;"></span><span class="mord">64</span><span class="mord text"><span class="mord"> tokens</span></span><span class="mspace" style="margin-right:0.2778em;"></span><span class="mrel">=</span><span class="mspace" style="margin-right:0.2778em;"></span></span><span class="base"><span class="strut" style="height:0.8889em;vertical-align:-0.1944em;"></span><span class="mord">10.24</span><span class="mord text"><span class="mord">M tokens</span></span><span class="mpunct">,</span></span></span></span></span></p><p>against a cache of roughly 632M. The run therefore saw about <strong>1.6%</strong> of the corpus, once. The final loss of 2.14 is a number from a model that is data-starved by design, and the correct reading of it is &quot;the pipeline learns&quot; rather than &quot;the model is good.&quot; Anyone comparing it to a published TinyStories loss is comparing against a different experiment.</p><h2 id="sampling">sampling</h2><p>Greedy decoding on a model this size produces loops almost immediately. Temperature and top-<span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>k</mi></mrow><annotation encoding="application/x-tex">k</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.6944em;"></span><span class="mord mathnormal" style="margin-right:0.0315em;">k</span></span></span></span> were about thirty lines and changed the output more than any architectural change we made.</p><figure class="highlight python"><table><tr><td class="code"><pre><span class="line">logits = logits[:, -<span class="number">1</span>, :] / temperature</span><br><span class="line"><span class="keyword">if</span> top_k <span class="keyword">is</span> <span class="keyword">not</span> <span class="literal">None</span>:</span><br><span class="line">    cutoff = torch.topk(logits, <span class="built_in">min</span>(top_k, logits.size(-<span class="number">1</span>))).values[:, -<span class="number">1</span>, <span class="literal">None</span>]</span><br><span class="line">    logits = logits.masked_fill(logits &lt; cutoff, <span class="built_in">float</span>(<span class="string">&quot;-inf&quot;</span>))</span><br><span class="line">next_token = torch.multinomial(F.softmax(logits, dim=-<span class="number">1</span>), num_samples=<span class="number">1</span>)</span><br></pre></td></tr></table></figure><p>The thing worth internalizing is that the model is a distribution, not a speaker. Every &quot;the model said X&quot; is really &quot;this decoding strategy, at this temperature, said X.&quot; Two of the qualitative judgments we made early about model quality turned out to be judgments about decoding.</p><h2 id="scaling-to-a-gpu-and-what-it-bought">scaling to a GPU, and what it bought</h2><p>Once the laptop run worked, training moved to a Modal L4 with a larger configuration of 8 layers, 320-dimensional residual, 128-token context, mixed precision, and an auto batch-size search that doubles the batch until CUDA runs out of memory.</p><p>The observation worth recording is how little of that work was model code. The same <code>train.py</code> runs on all three devices behind a <code>--device auto</code> flag. What the GPU run required was infrastructure.</p><details class="collapsible-section"><summary><strong>The supervising loop</strong></summary><p>A persistent volume so a preempted container does not lose the checkpoint, a <code>--resume</code> path pointed at the same file as <code>--output</code>, and a loop that commits the volume every 60 seconds.</p><figure class="highlight python"><table><tr><td class="code"><pre><span class="line"><span class="keyword">while</span> <span class="literal">True</span>:</span><br><span class="line">    <span class="keyword">try</span>:</span><br><span class="line">        return_code = process.wait(timeout=<span class="number">60</span>)</span><br><span class="line">        ...</span><br><span class="line">    <span class="keyword">except</span> subprocess.TimeoutExpired:</span><br><span class="line">        storage.commit()</span><br></pre></td></tr></table></figure><p>Six hours of GPU time is worth nothing if the container dies at hour five with the checkpoint in a tmpfs.</p></details><p><code>model.py</code> is 182 lines. The tokenizer, data pipeline, training loop, sampler, chat REPL, and Modal harness together are roughly four times that, and that is where all of the bugs were. The transformer was the easy half.</p><h2 id="what-cost-the-most-time">what cost the most time</h2><table><thead><tr><th>Problem</th><th>Symptom</th><th>Cause</th></tr></thead><tbody><tr><td>Stale token cache</td><td>Model trains, output is gibberish, loss looks plausible</td><td>Tokenizer retrained without invalidating the cached token file</td></tr><tr><td>Dataset target alignment</td><td>Off-by-one, loss plateaus higher than expected</td><td>Dataset pre-shifts targets; the loss shifts them again</td></tr></tbody></table><p>Both are silent. Neither raises. Both are the same category, a pipeline that is internally consistent and describing the wrong thing. That category is, we now believe, the dominant failure mode in small-scale ML work, and it is why the receipt matters more than the number on it.</p><h2 id="conclusions">conclusions</h2><p>Had we reported only the final loss, this post would have described a trained model. The accurate version is that it describes a <em>validated pipeline</em> with a model attached that has seen 1.6% of its corpus once.</p><ul><li><strong>The naive version is the point.</strong> Every optimization skipped is a concept left visible. Unfused multi-head attention costs throughput and buys understanding, and at this scale that trade is free.</li><li><strong>The scale factor and the mask are not details.</strong> <span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mn>1</mn><mi mathvariant="normal">/</mi><msqrt><mi>d</mi></msqrt></mrow><annotation encoding="application/x-tex">1/\sqrt{d}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:1.1822em;vertical-align:-0.25em;"></span><span class="mord">1/</span><span class="mord sqrt"><span class="vlist-t vlist-t2"><span class="vlist-r"><span class="vlist" style="height:0.9322em;"><span class="svg-align" style="top:-3em;"><span class="pstrut" style="height:3em;"></span><span class="mord" style="padding-left:0.833em;"><span class="mord mathnormal">d</span></span></span><span style="top:-2.8922em;"><span class="pstrut" style="height:3em;"></span><span class="hide-tail" style="min-width:0.853em;height:1.08em;"><svg xmlns="http://www.w3.org/2000/svg" width="400em" height="1.08em" viewBox="0 0 400000 1080" preserveAspectRatio="xMinYMin slice"><path d="M95,702c-2.7,0,-7.17,-2.7,-13.5,-8c-5.8,-5.3,-9.5,-10,-9.5,-14c0,-2,0.3,-3.3,1,-4c1.3,-2.7,23.83,-20.7,67.5,-54c44.2,-33.3,65.8,-50.3,66.5,-51c1.3,-1.3,3,-2,5,-2c4.7,0,8.7,3.3,12,10s173,378,173,378c0.7,0,35.3,-71,104,-213c68.7,-142,137.5,-285,206.5,-429c69,-144,104.5,-217.7,106.5,-221l0 -0c5.3,-9.3,12,-14,20,-14H400000v40H845.2724s-225.272,467,-225.272,467s-235,486,-235,486c-2.7,4.7,-9,7,-19,7c-6,0,-10,-1,-12,-3s-194,-422,-194,-422s-65,47,-65,47zM834 80h400000v40h-400000z"/></svg></span></span></span><span class="vlist-s">​</span></span><span class="vlist-r"><span class="vlist" style="height:0.1078em;"><span></span></span></span></span></span></span></span></span> is the difference between a model that trains and one that silently does not, and the failure has no error message.</li><li><strong>Most of the code is not the model.</strong> 182 lines of model against roughly four times that in everything else, and the bugs were all in the everything else.</li><li><strong>Silent correctness bugs dominate.</strong> Nothing in this project crashed in an informative way. Both expensive bugs produced a running system quietly learning the wrong function.</li><li><strong>The loss number is not comparable to anything.</strong> One pass over 1.6% of the data is not a training run anyone should benchmark against.</li></ul>]]>
    </content>
    <id>https://projects.farhansadeek.com/posts/building-a-gpt-from-scratch/</id>
    <link href="https://projects.farhansadeek.com/posts/building-a-gpt-from-scratch/"/>
    <published>2026-09-18T19:04:52.000Z</published>
    <summary>A 5.26M-parameter GPT built from scratch in PyTorch and trained on TinyStories to a cross-entropy of 2.14 on a laptop, where the loss is a receipt that the pipeline works rather than a result.</summary>
    <title>Every Line of a Small GPT, Written by Hand</title>
    <updated>2026-09-29T06:49:26.552Z</updated>
  </entry>
</feed>
