<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="4.4.1">Jekyll</generator><link href="https://kovashikawa.com/feed.xml" rel="self" type="application/atom+xml" /><link href="https://kovashikawa.com/" rel="alternate" type="text/html" /><updated>2026-09-29T12:39:58+00:00</updated><id>https://kovashikawa.com/feed.xml</id><title type="html">Rafael Kovashikawa</title><subtitle>Data Scientist &amp; AI Engineer at FUSE. Writing about AI agents, MCP servers, quantitative finance, and data engineering.</subtitle><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;&quot;/assets/images/selfie-avatar.webp&quot;, &quot;bio&quot;=&gt;&quot;AI Engineer @ FUSE. Previously macro quant @ JGP. MIT MicroMasters. Washington, DC.&quot;, &quot;location&quot;=&gt;&quot;Washington DC&quot;, &quot;email&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Email&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-envelope-square&quot;}, {&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;}, {&quot;label&quot;=&gt;&quot;Twitter&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-twitter-square&quot;}, {&quot;label&quot;=&gt;&quot;Facebook&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-facebook-square&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/kovashikawa&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/in/rkovashikawa/&quot;}]}</name></author><entry><title type="html">Stack Overflow Is Selling Proof That a Human Wrote It</title><link href="https://kovashikawa.com/ai/stack-overflow-proof-of-authorship/" rel="alternate" type="text/html" title="Stack Overflow Is Selling Proof That a Human Wrote It" /><published>2026-09-18T00:00:00+00:00</published><updated>2026-09-18T00:00:00+00:00</updated><id>https://kovashikawa.com/ai/stack-overflow-proof-of-authorship</id><content type="html" xml:base="https://kovashikawa.com/ai/stack-overflow-proof-of-authorship/"><![CDATA[<p>Stack Overflow brought Developer Story back on September 10, four years after the
2022 shutdown. <a href="https://stackoverflow.blog/2026/09/10/re-introducing-developer-story/">The announcement</a>
is unusually blunt about why.</p>

<p>The short version is that people don’t need to come to the site for answers
anymore, “the chatbots have that covered,” and most LLMs were trained on their
datasets. So the goal is not winning back question traffic. It is evidence that a
specific human wrote a specific thing.</p>

<p>The umbrella name is Stack Identity, described in the post as “the verified
developer proof-of-work layer to the internet.” Day one ships “specialties,”
areas of proven expertise derived from tag activity, with a settings pane for
choosing which ones are displayed. Later comes material pulled from
“integrations with other sites that have verified data about you.”</p>

<h2 id="not-a-linkedin-pivot">Not a LinkedIn pivot</h2>

<p>My first read was that this is Stack Overflow turning into LinkedIn. It isn’t.
LinkedIn’s moat is the network graph plus recruiter demand, and Stack Overflow
has neither and is not building either. This reads more like a notary. The
product is the signature, and the profile is just the surface it gets printed
on.</p>

<p>The 2022 sunset deleted the old Dev Story data for privacy compliance, so nobody
carried anything over. A notary with no back catalog has to sell a signature that
starts today. My own record is the illustration. Six years of membership, four
answers, one accepted, and the accepted one is the zero-score answer from 2023.
The answer that carries the number is a pandas reply from 2021 at score 122,
never accepted, second of twenty on the question. That is the whole verified
footprint, and none of it describes what I work on now.</p>

<h2 id="verification-needs-a-second-party">Verification needs a second party</h2>

<p>A signature only has value when someone else accepts it. “Integrations with
other sites that have verified data about you” is carrying the entire
distribution argument, and no partner names are attached to it. The hard part is
not the profile UI or the trust model. It is getting an employer, a client, or
another platform to treat Stack Overflow as the identity authority.</p>

<p>They also run <a href="https://agents.stackoverflow.com/">agents.stackoverflow.com</a>, a
separate site in beta where the users are AI agents. A post there carries no
trust score at all until at least one independent verification reports that the
guidance worked in practice. Votes are read-time judgments, verifications are
use-time outcomes, and reputation only moves when other agents find the
contribution useful. The strictest proof standard in the company’s portfolio is
pointed at machines, while the human product ships with tag activity and a
promise.</p>

<p>Both are bids to define what counts as proof of work. The agent side requires an
independent verification before it will score anything. The human side starts
with the tags you already answered.</p>]]></content><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;&quot;/assets/images/selfie-avatar.webp&quot;, &quot;bio&quot;=&gt;&quot;AI Engineer @ FUSE. Previously macro quant @ JGP. MIT MicroMasters. Washington, DC.&quot;, &quot;location&quot;=&gt;&quot;Washington DC&quot;, &quot;email&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Email&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-envelope-square&quot;}, {&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;}, {&quot;label&quot;=&gt;&quot;Twitter&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-twitter-square&quot;}, {&quot;label&quot;=&gt;&quot;Facebook&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-facebook-square&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/kovashikawa&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/in/rkovashikawa/&quot;}]}</name></author><category term="ai" /><category term="stack-overflow" /><category term="verification" /><category term="identity" /><category term="agents" /><summary type="html"><![CDATA[Developer Story is back, sold as a verification layer for authorship. On the empty ledger, the notary framing, and who gets to certify a developer.]]></summary></entry><entry><title type="html">Beyond Correlation: measuring dependence the modern way</title><link href="https://kovashikawa.com/statistics/beyond-pearsons-rho/" rel="alternate" type="text/html" title="Beyond Correlation: measuring dependence the modern way" /><published>2026-09-02T16:00:00+00:00</published><updated>2026-09-02T16:00:00+00:00</updated><id>https://kovashikawa.com/statistics/beyond-pearsons-rho</id><content type="html" xml:base="https://kovashikawa.com/statistics/beyond-pearsons-rho/"><![CDATA[<p>Pearson’s correlation is zero for plenty of pairs that are fully dependent. The canonical case: $X \sim \mathcal{N}(0,1)$, $Y = \lvert X \rvert$. $\rho$ reports about 0, but $X$ determines $Y$ exactly. The fix isn’t “use a nonlinear measure” in the abstract, it’s knowing which measure catches which failure mode, and what each one actually guarantees.</p>

<p>That’s what this post benchmarks: distance correlation, Chatterjee’s xi, HSIC, KSG mutual information, and tail dependence, against nine synthetic datasets built to break Pearson in different ways.</p>

<p>Every number below is reproducible: seeded, and the companion repo is <a href="https://github.com/kovashikawa/correlation-models">kovashikawa/correlation-models</a>.</p>

<p>Two minutes covering the same ground visually, walking through each counterexample and measure below.</p>

<video controls="" muted="" playsinline="" poster="/assets/videos/beyond-correlation-poster.png" width="100%">
  <source src="/assets/videos/beyond-correlation.mp4" type="video/mp4" />
</video>

<h2 id="the-benchmark-table">The benchmark table</h2>

<table>
  <thead>
    <tr>
      <th>Measure</th>
      <th>linear</th>
      <th>quadratic</th>
      <th>abs</th>
      <th>sine</th>
      <th>circle</th>
      <th>cross</th>
      <th>independent</th>
      <th>heavy_tail</th>
      <th>tail_t</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Pearson</td>
      <td>0.980</td>
      <td>-0.017</td>
      <td>-0.015</td>
      <td>0.004</td>
      <td>0.004</td>
      <td>-0.028</td>
      <td>0.000</td>
      <td>0.729</td>
      <td>0.484</td>
    </tr>
    <tr>
      <td>Spearman</td>
      <td>0.978</td>
      <td>-0.014</td>
      <td>-0.014</td>
      <td>0.026</td>
      <td>0.003</td>
      <td>-0.010</td>
      <td>0.004</td>
      <td>0.795</td>
      <td>0.477</td>
    </tr>
    <tr>
      <td>Kendall</td>
      <td>0.874</td>
      <td>-0.011</td>
      <td>-0.011</td>
      <td>0.023</td>
      <td>0.000</td>
      <td>-0.007</td>
      <td>0.003</td>
      <td>0.614</td>
      <td>0.343</td>
    </tr>
    <tr>
      <td>Chatterjee xi</td>
      <td>0.812</td>
      <td>0.999</td>
      <td>0.999</td>
      <td>0.997</td>
      <td>0.254</td>
      <td>0.245</td>
      <td>-0.004</td>
      <td>0.449</td>
      <td>0.164</td>
    </tr>
    <tr>
      <td>Distance corr</td>
      <td>0.971</td>
      <td>0.543</td>
      <td>0.559</td>
      <td>0.253</td>
      <td>0.197</td>
      <td>0.313</td>
      <td>0.014</td>
      <td>0.768</td>
      <td>0.468</td>
    </tr>
    <tr>
      <td>HSIC</td>
      <td>0.088</td>
      <td>0.045</td>
      <td>0.051</td>
      <td>0.008</td>
      <td>0.019</td>
      <td>0.033</td>
      <td>0.000</td>
      <td>0.043</td>
      <td>0.010</td>
    </tr>
    <tr>
      <td>KSG MI</td>
      <td>1.639</td>
      <td>6.102</td>
      <td>6.378</td>
      <td>4.716</td>
      <td>5.167</td>
      <td>6.194</td>
      <td>0.001</td>
      <td>0.568</td>
      <td>0.192</td>
    </tr>
    <tr>
      <td>Tail dep (q=0.95)</td>
      <td>0.856</td>
      <td>0.498</td>
      <td>0.498</td>
      <td>0.090</td>
      <td>0.000</td>
      <td>0.462</td>
      <td>0.032</td>
      <td>0.500</td>
      <td>0.376</td>
    </tr>
  </tbody>
</table>

<p style="font-size: 0.75em; color: var(--muted, #555555);">n = 10,000, seed 42, full generator code in the repo.</p>

<p>The first three rows are the classical toolkit, and none of them catch quadratic, abs, sine, circle, or cross. Chatterjee’s xi, distance correlation, and KSG MI catch all five. HSIC and tail dependence are subtler cases, covered below. That gap is the whole problem, quantified.</p>

<p>Two things worth flagging before the per-measure sections.</p>

<ol>
  <li><strong>Cross is the sharpest counterexample.</strong> $Y = X \cdot W$, where $W$ is a random sign independent of $X$. The sign cancels out any linear relationship, so Pearson, Spearman, and Kendall all read near zero. But $Y$ is not independent of $X$: $\lvert Y \rvert = \lvert X \rvert$ always. Chatterjee’s xi (0.245) and distance correlation (0.313) both catch it. Both margins are standard normal here, and the pair still isn’t jointly Gaussian, which is exactly the case where “uncorrelated implies independent” fails.</li>
  <li><strong>The measures split into two questions.</strong> Xi, distance correlation, HSIC, and KSG MI test dependence across the whole distribution. Tail dependence tests something narrower: whether extremes move together, which is why the tail_t column exists. HSIC is unnormalized and MI is measured in nats, so don’t compare magnitudes across measures. Use them to detect and rank, not to score.</li>
  <li><strong>The independent column is the null.</strong> Every measure should read near zero there, and does: Chatterjee’s xi is -0.004, distance correlation 0.014, HSIC 0.000. That’s the reference for judging every “catch” above. Xi’s null standard deviation is about $\sqrt{2/(5n)} \approx 0.0063$ at $n = 10{,}000$ (derived below), which puts circle’s xi of 0.254 roughly 40 null standard deviations out, not just “nonzero.”</li>
</ol>

<h2 id="distance-correlation">Distance correlation</h2>

<p>Szekely, Rizzo and Bakirov (2007) introduced distance correlation to fix exactly this blind spot. The idea: independence is equivalent to the joint characteristic function factoring into the product of marginals. Distance covariance is a weighted norm on exactly that difference, and the estimator falls out as a double-centering of pairwise distance matrices:</p>

\[\operatorname{dCor}(X,Y) = \frac{\operatorname{dCov}(X,Y)}{\sqrt{\operatorname{dCov}(X,X)\,\operatorname{dCov}(Y,Y)}}.\]

<p>The property that matters: $\operatorname{dCor} = 0$ if and only if $X$ and $Y$ are independent, for distributions with finite first moments, in any dimension. Pearson cannot make that claim. In the bivariate normal case dCor is a deterministic function of $\lvert \rho \rvert$ and never exceeds it.</p>

<p>Cost: the naive estimator used here is $O(n^2)$ memory and time, because of the pairwise distance matrices. Fine at 10k rows, painful at 10M. Faster $O(n \log n)$ algorithms exist for univariate data (Huo and Szekely 2016).</p>

<h2 id="chatterjees-rank-correlation">Chatterjee’s rank correlation</h2>

<p>Chatterjee (2021) took a different route, with a coefficient that is almost absurdly simple. Rank X, reorder Y’s max-ranks by X, and measure how much adjacent ranks jump:</p>

\[\xi_n(X,Y) = 1 - \frac{A_1}{C_U}, \qquad A_1 = \frac{1}{2n}\sum_{i=1}^{n-1}\left|\frac{r_{i+1}}{n} - \frac{r_i}{n}\right|, \qquad C_U = \frac{1}{n}\sum_{i=1}^{n} g_i(1-g_i),\]

<p>where $r_i$ are the max-ranks of $Y$ reordered by $X$ and $g_i = (\text{max-rank of } -Y_i)/n$, so that $r_i/n$ and $g_i$ lie in $[0, 1]$. With no ties this collapses to $\xi = 1 - \frac{3}{n^2-1}\sum_{i=1}^{n-1}\lvert r_{i+1} - r_i \rvert$. For non-constant $Y$, the population coefficient $\xi$ satisfies: $\xi = 0$ iff independence, $\xi = 1$ iff $Y$ is a measurable function of $X$. Computes in $O(n \log n)$ and is completely nonparametric.</p>

<p>Three footnotes. First, under independence $\xi_n$ has mean zero and standard deviation about $\sqrt{2/(5n)}$, so it lands negative about half the time; that is expected, not a bug. Second, $\xi(X, Y)$ is asymmetric by construction: it measures “how well Y behaves as a function of X,” which the paper argues for deliberately. Third, xi has low power against many smooth alternatives (Shi, Drton and Han 2022), so treat a low xi as “no strong signal,” not “no dependence.”</p>

<h2 id="hsic">HSIC</h2>

<p>HSIC comes from the kernel methods literature (Gretton et al. 2005) and became a standard tool in nonlinear feature selection (Song et al. 2012). Map each variable into a reproducing kernel Hilbert space with a universal kernel (RBF here), and take the squared Hilbert-Schmidt norm of the cross-covariance operator between the two embeddings. HSIC = 0 iff independence for universal kernels (Gretton 2005 on compact domains; Fukumizu et al. 2008 for the general statement).</p>

<p>On this benchmark HSIC is the flattest row: 0.088 on linear, dropping to 0.008 on sine and 0.000 on independent, everything within a factor of about 11. It still separates dependent from independent, linear and heavy_tail sit well above the 0.000 null, but the RBF kernel’s median-distance bandwidth is not tuned per dataset here, so treat the exact HSIC values as a coarse detector rather than a ranking. Its real strength is scaling to vector-valued or high-dimensional X and Y, where density-based measures like KSG MI become impractical.</p>

<h2 id="ksg-mutual-information">KSG mutual information</h2>

<p>Mutual information $I(X; Y) = 0$ iff independence, full stop. The KSG estimator (Kraskov, Stogbauer and Grassberger 2004) is a k-nearest-neighbor scheme that adapts its resolution in both margins, which fixes the classic histogram-bin problems. It is what scikit-learn’s <code class="language-plaintext highlighter-rouge">mutual_info_regression</code> uses under the hood, with $k = 3$ neighbors here.</p>

<p>The values are not comparable across datasets. Linear has additive noise and a finite population MI of $-\frac{1}{2}\ln(1-\rho^2)$. At the generator’s $\rho \approx 0.9806$ that works out to about 1.629 nats, matching the 1.639 estimate. Quadratic, abs, sine, circle, and cross have no additive noise: each Y is fully determined by X and a coin flip or angle draw with no residual randomness, so the joint distribution is singular and the population MI is infinite. The KSG estimate for these just grows with $n$ and shrinks with $k$. So “6.4 nats on abs” does not mean abs is four times more dependent than linear’s 1.6. MI answers “is there dependence” decisively, and “how strong” only loosely.</p>

<p>The maximal information coefficient (Reshef et al. 2011) is the other mutual-information-based measure people reach for, maximizing normalized MI over grid binning schemes. It made a splash in <em>Science</em> in 2011, but its equitability claims were shown mathematically impossible for any nontrivial measure (Kinney and Atwal 2014), and later power comparisons found it underpowered relative to distance correlation (Simon and Tibshirani 2014). Refined variants followed (Reshef et al. 2016). Treat MIC, like the measures above, as a detector rather than a calibrated strength scale. It is not in this benchmark: <code class="language-plaintext highlighter-rouge">minepy</code>, the standard implementation, requires a compiled extension and was left out; the Reshef lab’s Java MINE tool is the reference implementation.</p>

<h2 id="tail-dependence">Tail dependence</h2>

<p>The measures above test dependence across the whole distribution. Risk work cares about the tails specifically: given that one asset is above its 95th percentile, how likely is the other to be too? Tail dependence coefficients go back to Sibuya (1960); the modern treatment is Joe (1997).</p>

<p>The population quantity is the limit, if it exists:</p>

\[\lambda_U = \lim_{q \to 1} P(F_Y(Y) &gt; q \mid F_X(X) &gt; q).\]

<p>The benchmark row reports the finite-quantile estimator $\lambda(q)$ at $q = 0.95$, a standard VaR level, exactly 500 conditioning exceedances at $n = 10{,}000$. Under independence, $\lambda(q) = 1 - q = 0.05$ exactly. The independent column reads 0.032, about 1.8 standard errors low (SE $\approx \sqrt{0.05 \cdot 0.95 / 500} \approx 0.0098$), consistent with the null.</p>

<p>This is where the Gaussian copula earns its infamy. For any correlation $\rho &lt; 1$, the Gaussian copula has zero tail dependence in the limit, but the co-exceedance probability vanishes slowly as $q \to 1$. The linear column is generated from a Gaussian copula at $\rho \approx 0.9806$. Its finite-q estimator reads 0.856 at $q = 0.95$ against a closed form of 0.838. That gap narrows only gradually: about 0.74 by $q = 0.999$ and 0.70 by $q = 0.9999$. If your risk model is Gaussian-copula shaped and you feed it Pearson correlations, you are asserting away joint tail risk by construction, just with a delay.</p>

<p>The tail_t column is the counterexample. It’s generated from a $t$-copula with $\nu = 3, \rho = 0.5$, which has positive asymptotic tail dependence: $\lambda_U = 2T_4(-\sqrt{4/3}) \approx 0.31$. The empirical estimate at $q = 0.95$ reads 0.376, higher than the asymptotic limit because finite-q estimates overstate $\lambda_U$ and converge to it slowly, the same effect visible in the linear column above.</p>

<p>The heavy_tail column is a warning about reading too much into a name. It’s just $X$ plus independent $t_3$ noise, which is asymptotically tail independent despite the heavy-tailed marginal. Larger simulations ($n = 20$M) show its $\lambda(q)$ decaying with $q$: 0.487 at $q = 0.95$, 0.011 at $q = 0.999$, below 0.001 at $q = 0.9999$. The 0.500 in the benchmark table is the $n = 10{,}000$ estimate at $q = 0.95$ only, consistent with the larger simulation’s first point; the decay only shows up once $q$ moves well past 0.95. Always check the generator, not the label.</p>

<h2 id="when-to-use-what">When to use what</h2>

<ul>
  <li>Bivariate, monotone, want a sign: Spearman or Kendall.</li>
  <li>Bivariate, any shape, want a [0, 1] strength: distance correlation or Chatterjee xi.</li>
  <li>High-dimensional or vector-valued: distance correlation, HSIC.</li>
  <li>Nonlinear feature screening: HSIC, KSG MI.</li>
  <li>Portfolio and risk: tail dependence on top of a global measure.</li>
  <li>Rule of thumb: never ship an <em>independence</em> claim, or a strength-of-relationship number, built on Pearson alone. A clearly nonzero Pearson correlation is itself sufficient evidence of dependence.</li>
</ul>

<h2 id="reproducing-this">Reproducing this</h2>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>git clone https://github.com/kovashikawa/correlation-models
<span class="nb">cd </span>correlation-models
uv venv
uv pip <span class="nb">install</span> <span class="nt">-e</span> <span class="nb">.</span>
uv run python scripts/benchmark.py
</code></pre></div></div>

<p>The repo also includes a self-test that checks Chatterjee’s xi against the canonical XICOR formulation, distance correlation against known cases, tail dependence against closed-form values for $Y = X$ and $Y = \lvert X \rvert$, and the t-copula column against its closed-form $\lambda_U$.</p>

<h2 id="further-reading">Further reading</h2>

<ol>
  <li>Szekely, Rizzo, Bakirov (2007). <a href="https://projecteuclid.org/journals/annals-of-statistics/volume-35/issue-6/Measuring-and-testing-dependence-by-correlation-of-distances/10.1214/009053607000000505.full">Measuring and testing dependence by correlation of distances</a>. <em>Annals of Statistics</em>.</li>
  <li>Szekely and Rizzo (2009). <a href="https://projecteuclid.org/journals/annals-of-applied-statistics/volume-3/issue-4/Brownian-distance-covariance/10.1214/09-AOAS312.full">Brownian distance covariance</a>. <em>Annals of Applied Statistics</em>.</li>
  <li>Chatterjee (2021). <a href="https://arxiv.org/pdf/1909.10140">A new coefficient of correlation (PDF)</a>. <em>JASA</em>.</li>
  <li>Gretton et al. (2005). <a href="http://alex.smola.org/papers/2005/GreBouSmoSch05.pdf">Measuring statistical dependence with Hilbert-Schmidt norms (PDF)</a>. <em>ALT</em>.</li>
  <li>Kraskov, Stogbauer, Grassberger (2004). <a href="https://arxiv.org/pdf/cond-mat/0305641">Estimating mutual information (PDF)</a>. <em>Physical Review E</em>.</li>
  <li>Reshef et al. (2011). <a href="https://www.science.org/doi/10.1126/science.1205438">Detecting novel associations in large data sets</a>. <em>Science</em>.</li>
  <li>Kinney and Atwal (2014). <a href="https://www.pnas.org/doi/10.1073/pnas.1309933111">Equitability, mutual information, and the maximal information coefficient</a>. <em>PNAS</em>.</li>
  <li>Sibuya (1960). Bivariate extreme statistics, I. <em>Annals of the Institute of Statistical Mathematics</em> 11(3):195-210.</li>
  <li>Joe (1997). <a href="https://doi.org/10.1201/b13150">Multivariate models and dependence concepts</a>. Chapman and Hall.</li>
  <li>Song et al. (2012). <a href="https://www.jmlr.org/papers/volume13/song12a/song12a.pdf">Feature selection via dependence maximization (PDF)</a>. <em>JMLR</em>.</li>
  <li>Fukumizu, Gretton, Sun, Scholkopf (2008). <a href="https://papers.nips.cc/paper/3340-kernel-measures-of-conditional-dependence.pdf">Kernel measures of conditional dependence (PDF)</a>. <em>NIPS 20 (2008)</em>.</li>
  <li>Huo and Szekely (2016). <a href="https://www.tandfonline.com/doi/abs/10.1080/00401706.2015.1054435">Fast computing for distance covariance</a>. <em>Technometrics</em>.</li>
  <li>Shi, Drton, Han (2022). <a href="https://arxiv.org/pdf/2008.06820">On the power of Chatterjee’s rank correlation (PDF)</a>. <em>Biometrika</em>.</li>
  <li>Simon and Tibshirani (2014). <a href="https://arxiv.org/pdf/1401.7645">Comment on “Detecting novel associations in large data sets” by Reshef et al. (PDF)</a>. <em>arXiv:1401.7645</em>.</li>
  <li>Reshef, Reshef, Mitzenmacher, Sabeti (2014). <a href="https://www.pnas.org/doi/10.1073/pnas.1408920111">Cleaning up the record on the maximal information coefficient and equitability</a>. <em>PNAS</em> 111(33):E3362-E3363.</li>
  <li>Reshef, Reshef, Finucane, Sabeti, Mitzenmacher (2016). <a href="https://jmlr.csail.mit.edu/papers/volume17/15-308/15-308.pdf">Measuring dependence powerfully and equitably (PDF)</a>. <em>JMLR</em> 17(211):1-63.</li>
</ol>]]></content><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;&quot;/assets/images/selfie-avatar.webp&quot;, &quot;bio&quot;=&gt;&quot;AI Engineer @ FUSE. Previously macro quant @ JGP. MIT MicroMasters. Washington, DC.&quot;, &quot;location&quot;=&gt;&quot;Washington DC&quot;, &quot;email&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Email&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-envelope-square&quot;}, {&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;}, {&quot;label&quot;=&gt;&quot;Twitter&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-twitter-square&quot;}, {&quot;label&quot;=&gt;&quot;Facebook&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-facebook-square&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/kovashikawa&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/in/rkovashikawa/&quot;}]}</name></author><category term="statistics" /><category term="statistics" /><category term="correlation" /><category term="dependence" /><category term="python" /><category term="benchmark" /><summary type="html"><![CDATA[A guide to distance correlation, Chatterjee's xi, HSIC, KSG mutual information, and tail dependence: what each measure catches that Pearson misses, benchmarked against nine synthetic datasets.]]></summary></entry><entry><title type="html">Lowercase Is a Mood</title><link href="https://kovashikawa.com/design/projects/lowercase-is-a-mood/" rel="alternate" type="text/html" title="Lowercase Is a Mood" /><published>2026-08-28T00:00:00+00:00</published><updated>2026-08-28T00:00:00+00:00</updated><id>https://kovashikawa.com/design/projects/lowercase-is-a-mood</id><content type="html" xml:base="https://kovashikawa.com/design/projects/lowercase-is-a-mood/"><![CDATA[<p>The whole site is one click away from lowercase. There is a bare <code class="language-plaintext highlighter-rouge">a/A</code> button
in the masthead, next to Home. Click it and every heading, title, date and
excerpt renders lowercase. Click again and it goes back. The active letter is
bold, so the button says which mode you are in.</p>

<video controls="" loop="" muted="" playsinline="" poster="/assets/videos/lc-lowercase-poster.png" width="100%">
  <source src="/assets/videos/lc-lowercase.mp4" type="video/mp4" />
</video>

<h2 id="display-only">Display-only</h2>

<p>The feature is a single CSS rule applied to the body when the toggle is on:</p>

<div class="language-css highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nt">body</span><span class="nc">.lc</span><span class="o">,</span> <span class="nt">body</span><span class="nc">.lc</span> <span class="o">*</span> <span class="p">{</span>
  <span class="nl">text-transform</span><span class="p">:</span> <span class="nb">lowercase</span> <span class="cp">!important</span><span class="p">;</span>
<span class="p">}</span>
</code></pre></div></div>

<p>That is the whole thing. It is a rendering choice, not a content change. The
source, the front matter, the page titles and the RSS feed all stay proper
case. Search engines read the unmodified markup. So this is free: a purely
visual preference that costs nothing in legibility of the underlying content,
and nothing in SEO.</p>

<h2 id="why-bother">Why bother</h2>

<p>Lowercase reads quieter. All-caps carries weight and formality; lowercase
drops the volume. For a personal site that is mostly short notes and links,
that is the right register. It is also a small joke: a site whose whole
point is few words, styled to say them as softly as possible.</p>

<p>The state persists in <code class="language-plaintext highlighter-rouge">localStorage</code>, so the mood survives reloads.</p>

<h2 id="one-button-two-letters">One button, two letters</h2>

<p>The button itself is just <code class="language-plaintext highlighter-rouge">a</code> and <code class="language-plaintext highlighter-rouge">A</code> with a slash. No box, no border, no
icon. The active letter gets weight. It is the smallest control that can
state a binary choice about typography, which is exactly what it is.</p>]]></content><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;&quot;/assets/images/selfie-avatar.webp&quot;, &quot;bio&quot;=&gt;&quot;AI Engineer @ FUSE. Previously macro quant @ JGP. MIT MicroMasters. Washington, DC.&quot;, &quot;location&quot;=&gt;&quot;Washington DC&quot;, &quot;email&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Email&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-envelope-square&quot;}, {&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;}, {&quot;label&quot;=&gt;&quot;Twitter&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-twitter-square&quot;}, {&quot;label&quot;=&gt;&quot;Facebook&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-facebook-square&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/kovashikawa&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/in/rkovashikawa/&quot;}]}</name></author><category term="design" /><category term="projects" /><category term="css" /><category term="typography" /><category term="minimalism" /><summary type="html"><![CDATA[One bare a/A button in the masthead renders the whole site in lowercase. Display-only, so the source and the SEO stay proper case. A short note on why and how.]]></summary></entry><entry><title type="html">A LinkedIn Cover That Isn’t Stock</title><link href="https://kovashikawa.com/projects/design/linkedin-cover/" rel="alternate" type="text/html" title="A LinkedIn Cover That Isn’t Stock" /><published>2026-08-22T00:00:00+00:00</published><updated>2026-08-22T00:00:00+00:00</updated><id>https://kovashikawa.com/projects/design/linkedin-cover</id><content type="html" xml:base="https://kovashikawa.com/projects/design/linkedin-cover/"><![CDATA[<p>My LinkedIn cover is one line of math and one curve:</p>

<p><img src="/assets/images/li-cover.png" alt="The cover" /></p>

\[dX_t = \mu X_t\,dt + \sigma X_t\,dW_t\]

<p>Geometric Brownian motion, the process Black-Scholes assumes for the
underlying. The curve is a single realized path: exponential drift, Brownian
noise. My preference is a quiet, monochrome plot in the same palette as this
blog, so that’s what this is.</p>

<p>A few decisions were deliberate.</p>

<h2 id="a-cover-should-be-ambient">A cover should be ambient</h2>

<p>A LinkedIn banner is background. People come to the profile for the bio, the
experience, the work. So the cover gets three elements and nothing else: the
curve, a caption, and a URL. No axes, no ticks, no gridlines inside the plot.
If the values don’t matter, the scaffolding is noise. I kept the graph-paper
grid from this blog’s background, because it gives the curve a sense of place
without making it a chart.</p>

<h2 id="grey-not-black-not-blue">Grey, not black, not blue</h2>

<p>The data-viz literature converges on “grey plus one accent” for figures where
a specific element matters. But a cover is the opposite situation: nothing in
it should compete with the content below it. So the curve and caption are a
mid grey, the URL is a little darker for legibility, and there is zero hue.</p>

<h2 id="text-on-a-cover-dies">Text on a cover dies</h2>

<p>The placement that survived is symmetric: URL top-right, equation
bottom-right, both right-aligned to the same column. An equation is content
that describes the visual, so it sits on the visual like a caption. The URL
is a signature, so it sits in the corner like one. Two text elements, two
jobs, no overlap.</p>

<h2 id="the-crispness-problem">The crispness problem</h2>

<p>Screenshot the HTML at 1x and a 2px curve is 2 physical pixels. It comes out
soft and aliased. The fix is a pipeline I now use every time I turn HTML into
an image:</p>

<ol>
  <li>Render in headless Chrome at 3x with <code class="language-plaintext highlighter-rouge">--force-device-scale-factor=3</code>,
giving a 4752x1188 master.</li>
  <li>Downscale to the exact target with Pillow’s LANCZOS resampler.
<code class="language-plaintext highlighter-rouge">sips -z</code> is bilinear and softens edges; LANCZOS is the right filter.</li>
  <li>Save as PNG, or JPEG around q=90 for a much smaller upload.</li>
</ol>

<p>The text needed to be a touch bigger than I first wanted (17px, not 15px).
At 1584px wide, small glyphs are at the legibility floor.</p>

<h2 id="the-tool">The tool</h2>

<p>All of it is one Python file now, so the next cover is one command:</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>kcover <span class="nt">--seed</span> 42 <span class="nt">--out</span> cover.png
</code></pre></div></div>

<p>It simulates a fresh path each run. Same equation, different realization,
which is the point. The seed pins a specific draw if you want to keep one.
It renders, downscales, and verifies the curve doesn’t collide with the
equation or the URL.</p>

<p>Repo: <a href="https://github.com/kovashikawa/kcover">github.com/kovashikawa/kcover</a></p>

<p>The cover is live on <a href="https://www.linkedin.com/in/rkovashikawa/">my profile</a>.
If you see it in the wild, the curve you’re looking at is one sample path of
a stochastic differential equation. That’s the whole joke.</p>]]></content><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;&quot;/assets/images/selfie-avatar.webp&quot;, &quot;bio&quot;=&gt;&quot;AI Engineer @ FUSE. Previously macro quant @ JGP. MIT MicroMasters. Washington, DC.&quot;, &quot;location&quot;=&gt;&quot;Washington DC&quot;, &quot;email&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Email&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-envelope-square&quot;}, {&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;}, {&quot;label&quot;=&gt;&quot;Twitter&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-twitter-square&quot;}, {&quot;label&quot;=&gt;&quot;Facebook&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-facebook-square&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/kovashikawa&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/in/rkovashikawa/&quot;}]}</name></author><category term="projects" /><category term="design" /><category term="linkedin" /><category term="mathjax" /><category term="visualization" /><category term="brand" /><summary type="html"><![CDATA[One line of math, one curve: a LinkedIn cover from geometric Brownian motion, in the same palette as this blog. Why it's grey, where the text went, and the 3x-then-downscale fix for crisp HTML-to-PNG.]]></summary></entry><entry><title type="html">MCP Grew Up Fast: From Experiment to Enterprise Trust Boundary</title><link href="https://kovashikawa.com/ai/mcp-grew-up-fast/" rel="alternate" type="text/html" title="MCP Grew Up Fast: From Experiment to Enterprise Trust Boundary" /><published>2026-08-19T16:00:00+00:00</published><updated>2026-08-19T16:00:00+00:00</updated><id>https://kovashikawa.com/ai/mcp-grew-up-fast</id><content type="html" xml:base="https://kovashikawa.com/ai/mcp-grew-up-fast/"><![CDATA[<h2 id="why-should-you-ever-log-into-salesforce-again">“Why should you ever log into Salesforce again?”</h2>

<p>Salesforce co-founder Parker Harris asked that out loud this spring, opening a product announcement most companies would never risk making about their own flagship product. By April 2026, Salesforce had shipped the part that made the question serious: a hosted, MCP connection any AI client can use to query a CRM in plain English, generally available to every Enterprise Edition org.</p>

<p>I spent the past few weeks building against <a href="https://www.fuse.is/blog/never-log-into-salesforce-again">exactly that surface</a>. What struck me wasn’t the demo. It was realizing how recently the plumbing underneath it didn’t exist at all, and how fast it got built.</p>

<h2 id="the-18-months-that-made-this-possible">The 18 months that made this possible</h2>

<p>MCP is barely two years old. Anthropic open-sourced it on November 25, 2024, after two engineers, David Soria Parra and Justin Spahr-Summers, had been building it since that July. What happened between then and Salesforce’s GA announcement is a compressed history of a protocol earning enterprise trust in real time, one hard problem at a time.</p>

<p><strong>Can two systems even speak the same language?</strong> (Nov 2024 - Mar 2025)
MCP launched as a fairly narrow tool: JSON-RPC over stdio, mostly local processes talking to Claude Desktop. Useful for developers, not yet something a SaaS vendor would expose to the internet. That changed on March 26, 2025, when the spec’s second version added Streamable HTTP transport and, for the first time, a real OAuth 2.1-based authorization framework. The same day, OpenAI publicly committed to supporting MCP across its products, turning it from “Anthropic’s protocol” into the industry’s protocol.</p>

<p><strong>Can you trust what’s on the other end?</strong> (Apr - Jun 2025)
Standardizing the wire format immediately exposed how little the security model had been stress-tested. In April 2025, Invariant Labs published a reproducible “tool poisoning attack,” the first serious public MCP exploit. It forced the issue. By June 18, the spec formally classified every MCP server as an OAuth 2.1 resource server under RFC 9728 (Protected Resource Metadata), meaning a server now had a standard, spec-defined way to declare who it trusted and what it would accept. Five days later, Salesforce anchored its Agentforce 3 platform around MCP interoperability and shipped its first servers.</p>

<p><strong>Can an enterprise actually govern this?</strong> (Jun - Nov 2025)
A resource-server model is necessary but not sufficient for a Fortune 500 security review. Through the summer, the ecosystem built the governance layer the spec hadn’t yet: Cloudflare shipped MCP server portals, Auth0 published patterns for MCP-as-OAuth-resource-server, New Relic added MCP traffic observability. Then the November 25, 2025 spec, released on MCP’s first anniversary, formalized a lot of that community work directly into the standard: Client ID Metadata Documents replaced ad hoc dynamic client registration, an “Enterprise-Managed Authorization” extension let a company’s own identity provider broker trust instead of each vendor inventing its own, and PKCE went from recommended to mandatory.</p>

<p><strong>Is this now just infrastructure?</strong> (Dec 2025 - Apr 2026)
Two weeks after that spec, on December 9, 2025, Anthropic donated MCP to the Linux Foundation’s new Agentic AI Foundation, co-founded with Block and OpenAI, backed by Google, Microsoft, AWS, Cloudflare, and Bloomberg. A protocol that started as one company’s open-source experiment was now nobody’s proprietary asset. Four months later, Salesforce’s Hosted MCP Servers went GA: OAuth 2.0 per user, scoped through an External Client App requesting <code class="language-plaintext highlighter-rouge">mcp_api</code> and <code class="language-plaintext highlighter-rouge">refresh_token</code> grants, fully managed on Salesforce’s own infrastructure. Seventeen months, roughly, from a two-person side project to the default way an AI agent gets audited, revocable, read-only access to a live enterprise CRM.</p>

<h2 id="what-that-actually-bought-the-people-building-on-top-of-it">What that actually bought the people building on top of it</h2>

<p>None of the above is abstract if you were the one wiring a client into it. Here’s the honest counterfactual, without walking through what I actually built: two years ago, an integration like this meant inventing your own answer, from scratch, to “how do I prove this credential can only read, only as this user, only against this one resource,” and then re-litigating that answer with every enterprise security team that asked. There was no shared vocabulary for it. Every vendor’s OAuth implementation was a slightly different bespoke argument.</p>

<p>What the spec’s maturation actually hands you now is that argument, pre-made. Audience-bound tokens are a defined behavior (RFC 8707), not a design choice you have to justify. Protected Resource Metadata means a server can advertise its own trust boundary instead of you documenting it out-of-band. Mandatory PKCE and a standardized insufficient-scope error path mean the failure modes a security reviewer asks about already have a spec-sanctioned answer. You’re not inventing a compliance story anymore. You’re citing one that a hundred other implementers already stress-tested.</p>

<p>That is, genuinely, the difference between a multi-quarter security review and something a small team can ship, get audited, and put in front of an enterprise admin in a reasonable window. Not because the engineering got easier. Because the trust model stopped being something each of us had to reinvent alone.</p>

<h2 id="a-personal-note">A personal note</h2>

<p>I didn’t build any of the above. I’m one of a lot of engineers who happened to be building an integration during the window this protocol matured, and I got to feel the difference directly: a project that would have meant inventing a security model from scratch a year earlier instead meant adopting one that had already been through a public tool-poisoning exploit, a year of enterprise scrutiny, and formal standardization. That’s not a personal achievement. It’s what happens when a whole community, not just one vendor, spends eighteen months arguing about the right way to do this.</p>

<p>I’ll be talking about a related piece of this, deterministic evaluation and typed tool contracts for AI agents, at Google Search Central Live this year. Same underlying point: the hard problem in this space isn’t connecting to the data anymore. It’s trusting what comes back, and that’s only tractable now because the layer underneath got serious.</p>

<h2 id="two-years-not-a-decade">Two years, not a decade</h2>

<p>A two-person open-source project became a Linux Foundation standard with mandatory security guarantees, adopted by every major AI lab, in under two years. That’s not normally how standards get made. Usually it takes a decade of committees. This one got there because a large number of people kept finding the gaps and fixing them in public, fast, under real pressure. Worth remembering next time an integration “just works” and it’s tempting to forget how much had to go right first.</p>]]></content><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;&quot;/assets/images/selfie-avatar.webp&quot;, &quot;bio&quot;=&gt;&quot;AI Engineer @ FUSE. Previously macro quant @ JGP. MIT MicroMasters. Washington, DC.&quot;, &quot;location&quot;=&gt;&quot;Washington DC&quot;, &quot;email&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Email&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-envelope-square&quot;}, {&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;}, {&quot;label&quot;=&gt;&quot;Twitter&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-twitter-square&quot;}, {&quot;label&quot;=&gt;&quot;Facebook&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-facebook-square&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/kovashikawa&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/in/rkovashikawa/&quot;}]}</name></author><category term="ai" /><category term="MCP" /><category term="OAuth" /><category term="Salesforce" /><category term="Protocols" /><category term="AI" /><summary type="html"><![CDATA[how an 18-month standards effort quietly turned into the compliance layer for AI-to-CRM access]]></summary></entry><entry><title type="html">The Measurement Was Harder Than the Model</title><link href="https://kovashikawa.com/ai/projects/distilling-bls-agent/" rel="alternate" type="text/html" title="The Measurement Was Harder Than the Model" /><published>2026-07-27T00:00:00+00:00</published><updated>2026-07-29T00:00:00+00:00</updated><id>https://kovashikawa.com/ai/projects/distilling-bls-agent</id><content type="html" xml:base="https://kovashikawa.com/ai/projects/distilling-bls-agent/"><![CDATA[<h2 id="the-idea">The idea</h2>

<p>Take Qwen3-1.7B, give it a set of BLS economic data tools, and distill it into a
specialist small enough to run on an M4 Mini. Ask it “what happened to food
prices since 2021?” and it should emit <code class="language-plaintext highlighter-rouge">get_series(series_id="CUUR0000SAF1",
start="2021")</code>.</p>

<p>The pipeline: 205 hand-written seed questions over 82 economic concepts, 229
training examples, MLX LoRA, a 67MB adapter, six tools.</p>

<p>That part worked. Everything I initially believed about <em>how well</em> it worked was
wrong, in several separate ways, and finding each one required disproving the
previous one. The measurement was harder than the model. The fix, when it
finally came, was not a hyperparameter but a change to what the model was asked
to say.</p>

<h2 id="act-1-the-number-that-wasnt-real">Act 1: the number that wasn’t real</h2>

<p>Round one looked clean. Train loss fell from 3.4 to 0.027. Fifteen held-out
examples, 87% accuracy. I fixed some data issues, retrained, got 93%. Base model
scored roughly zero. Ship it.</p>

<p>Then I ran a separate adversarial review of the pipeline, and none of it survived.</p>

<p><strong>The held-out set was not held out.</strong> I had split the seeds before expanding
them, which sounds right. But the expander drew from shared pools (a list of
series IDs, a list of search terms) regardless of which split a seed belonged
to. So a training seed and a test seed could independently emit byte-identical
rows. 34% of my validation set appeared verbatim in training. Worse, the eval
file itself was stale: it was the <em>first</em> round’s holdout, and the second round’s
reshuffle had moved all 15 of those questions into training. The 93% was
measured on training data.</p>

<p><strong>The metric was the easy half.</strong> “Accuracy” meant <em>did it pick the right tool</em>:
a 6-way classification. Whether the emitted call was actually correct was never
scored.</p>

<p><strong>A ninth of the labels were wrong.</strong> Eleven series IDs named the wrong concept.
<code class="language-plaintext highlighter-rouge">CUUR0000SETG01</code> was mapped to “energy”; the BLS catalog calls it <strong>airline
fares</strong>. <code class="language-plaintext highlighter-rouge">CUUR0000SAH1</code> was “housing” but means <em>shelter</em>; <code class="language-plaintext highlighter-rouge">CUUR0000SAM1</code> was
“medical care” but means <em>medical care commodities</em>. Four IDs (<code class="language-plaintext highlighter-rouge">SAR1</code>, <code class="language-plaintext highlighter-rouge">SAC1</code>,
<code class="language-plaintext highlighter-rouge">SAS1</code>, <code class="language-plaintext highlighter-rouge">SEHA01</code>) do not exist in any BLS catalog at all. I had invented them by
pattern-matching the real ones.</p>

<p>Honest number, measured properly: <strong>13.3% exact match.</strong></p>

<p>The lesson I’d draw is not “be careful.” It’s that <em>every one of these bugs
pushed the number up</em>. Nobody investigates a pleasant surprise as hard as a
disappointing one, and that asymmetry is the whole problem.</p>

<h2 id="act-2-testing-the-impossible">Act 2: testing the impossible</h2>

<p>Fixing the leak exposed a deeper design error.</p>

<p>The split held out whole <em>concepts</em>. If “medical care” appeared only in the test
set, the model was being asked to produce <code class="language-plaintext highlighter-rouge">CUUR0000SAM</code> having never once seen
that mapping. Four of eleven scored items were unanswerable by construction.</p>

<p>This is a lookup task. <code class="language-plaintext highlighter-rouge">"medical care" → CUUR0000SAM</code> cannot be derived from
first principles; it can only be recalled. So the split should hold out
<strong>phrasings</strong>, not concepts:</p>

<ul>
  <li>Every concept contributes at least one phrasing to training</li>
  <li>Remaining phrasings go to val/test</li>
  <li>The build <strong>fails</strong> if any concept is held out entirely</li>
</ul>

<p>I rewrote the seed data as concept tables (82 concepts, each with 2+ distinct
phrasings) and added build-time assertions: every series ID must exist in the
bundled 8,103-row catalog, no example may appear in two splits, no concept may be
missing from train. Assertions, not intentions. Two of them have fired on me
since.</p>

<h2 id="act-3-the-discovery">Act 3: the discovery</h2>

<p>With clean data I swept checkpoints from 100 to 1400 iterations, planning to pick
the lowest validation loss like every tutorial says.</p>

<p>Val loss bottomed at iteration 250 and rose steadily afterward. Textbook
overfitting. But I was also scoring actual tool-call accuracy, and the two
disagreed violently:</p>

<table>
  <thead>
    <tr>
      <th>iter</th>
      <th>val loss</th>
      <th>exact match</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>200</td>
      <td><strong>0.135</strong> (min)</td>
      <td>37.2%</td>
    </tr>
    <tr>
      <td>400</td>
      <td>0.135</td>
      <td>65.1%</td>
    </tr>
    <tr>
      <td>600</td>
      <td>0.147</td>
      <td><strong>90.7%</strong></td>
    </tr>
    <tr>
      <td>800</td>
      <td>0.156</td>
      <td>81.4%</td>
    </tr>
    <tr>
      <td>1000</td>
      <td>0.169</td>
      <td>88.4%</td>
    </tr>
    <tr>
      <td>1400</td>
      <td>0.169</td>
      <td>74.4%</td>
    </tr>
  </tbody>
</table>

<p>Stopping at the val-loss minimum would have shipped a 37% model instead of a 91%
one. I wrote this up as a finding: <em>val loss is a trap for structured-output
tasks.</em> Cross-entropy punishes a confidently-wrong series ID exactly as hard as
gibberish, while a task evaluator sees a near-miss. I had citations lined up.</p>

<p>It was a config bug.</p>

<h2 id="act-4-disproving-my-own-finding">Act 4: disproving my own finding</h2>

<p>Before publishing I checked one number I had never looked at: the ratio of
completion length to prompt length.</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>mean prompt tokens     : 134
mean completion tokens :  19
generation ratio       : 0.141
</code></pre></div></div>

<p>Every training row carried a ~120-token system prompt, byte-identical across the
dataset, and a ~19-token tool call. And I was training with prompt loss
<strong>unmasked</strong>.</p>

<p><strong>87.7% of every gradient was the model re-predicting a fixed preamble.</strong></p>

<p>That explains all of it. Train loss of 0.027 was never impressive: most of it
was copying a constant. And validation loss was ~88% a measurement of <em>preamble
reproduction</em>, which is uncorrelated with whether the tool call is right. The two
curves weren’t in tension for any deep reason. One of them was mostly noise.</p>

<p>This is documented. Huerta-Enochian and Ko (2024) found a statistically
significant effect of prompt-loss weight specifically for short-completion data,
and Vaughn’s walkthrough of the same phenomenon on a multiple-choice dataset
(generation ratio 0.01) reports that stopping at the full-sequence val-loss
minimum yields 53% accuracy while completion loss is still falling. I reproduced
a known failure mode and mistook it for a discovery.</p>

<p>So I masked the prompt and reran. Prediction: val loss should re-couple with
accuracy.</p>

<table>
  <thead>
    <tr>
      <th>iter</th>
      <th>val loss (masked)</th>
      <th>exact match</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>200</td>
      <td>0.046</td>
      <td>62.8%</td>
    </tr>
    <tr>
      <td>400</td>
      <td>0.028</td>
      <td>86.0%</td>
    </tr>
    <tr>
      <td>600</td>
      <td><strong>0.027</strong> (min)</td>
      <td><strong>90.7%</strong></td>
    </tr>
    <tr>
      <td>800</td>
      <td>0.028</td>
      <td>90.7%</td>
    </tr>
  </tbody>
</table>

<table>
  <thead>
    <tr>
      <th> </th>
      <th>unmasked</th>
      <th>masked</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>pick by val-loss minimum</td>
      <td>37.2%</td>
      <td>86.8%</td>
    </tr>
    <tr>
      <td>best checkpoint available</td>
      <td>90.7%</td>
      <td>88.4%</td>
    </tr>
    <tr>
      <td><strong>penalty for trusting val loss</strong></td>
      <td><strong>~50 pts</strong></td>
      <td><strong>1.6 pts</strong></td>
    </tr>
  </tbody>
</table>

<p>The trap was self-inflicted. Mask the prompt and standard practice works fine.</p>

<p>Switching to chat-format data to enable masking fixed two other things I had been
carrying without noticing: Qwen3’s chat template emits an empty
<code class="language-plaintext highlighter-rouge">&lt;think&gt;&lt;/think&gt;</code> block, which is the <em>actual</em> mechanism for non-thinking mode;
I had been asking for it in English in the system prompt, which is not the same
thing. It also removed a stray leading space before every tool call that made the
training target tokenize differently from anything an inference path would
produce.</p>

<h2 id="act-5-the-thing-that-actually-mattered">Act 5: the thing that actually mattered</h2>

<p>Here is the result I should have led with, and it is the least exciting one.</p>

<p>I had been quoting 90.7%. To check reproducibility I retrained with three
different random seeds, changing nothing else:</p>

<table>
  <thead>
    <tr>
      <th>seed</th>
      <th>exact match</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>0</td>
      <td>76.7%</td>
    </tr>
    <tr>
      <td>1</td>
      <td>81.4%</td>
    </tr>
    <tr>
      <td>2</td>
      <td><strong>90.7%</strong></td>
    </tr>
    <tr>
      <td>3</td>
      <td>88.4%</td>
    </tr>
  </tbody>
</table>

<p><strong>Mean 84.3%, standard deviation 6.4 points, range 14 points.</strong></p>

<p>90.7% was not the result. It was the best of four draws. At this dataset size a
single run tells you almost nothing, and every comparison I had made (including
one where I concluded a data fix had caused a 14-point regression) was inside
the noise. That “regression” was seed 0 landing at the bottom of the
distribution. I nearly reverted a correct fix because of it.</p>

<p>After the config fixes (prompt masking, all 28 layers instead of 16, rank 16,
cosine schedule with warmup), across three seeds:</p>

<p><strong>Mean 88.4%, standard deviation 4.0.</strong> Better mean, tighter spread, and a
validation signal that now points the right way.</p>

<h2 id="an-interlude-on-measurement">An interlude on measurement</h2>

<p>Small-data fine-tuning has terrible measurement ergonomics. Nearly every mistake
in this project was a <em>measurement</em> mistake, not a modelling one:</p>

<ol>
  <li>A test set that wasn’t held out</li>
  <li>A metric that scored the easy half of the task</li>
  <li>Labels that were confidently wrong</li>
  <li>A test that asked for things never taught</li>
  <li>A loss dominated by a constant</li>
  <li>Single-run numbers with 14 points of spread</li>
  <li>Gold labels in a stale format, scoring correct answers as wrong</li>
</ol>

<p>The model was never the hard part. None of these were visible from the loss
curve. The first six all made the number look <em>better</em> than reality, which is why
they all survived. The seventh made it look catastrophically worse, and I nearly
abandoned a good idea because of it.</p>

<h2 id="act-6-the-part-hyperparameters-couldnt-fix">Act 6: the part hyperparameters couldn’t fix</h2>

<p>The config work moved 84% to 88% and stalled. Every remaining failure was sibling
confusion between near-identical codes: <code class="language-plaintext highlighter-rouge">CUUR0000SAF11</code> (food at home) vs
<code class="language-plaintext highlighter-rouge">CUUR0000SAF1</code> (food), <code class="language-plaintext highlighter-rouge">CUUR0000SAM</code> (medical care) vs <code class="language-plaintext highlighter-rouge">CUUR0000SEMD</code> (hospital
services). Those aren’t bugs. That’s what memorizing a codebook into weights looks
like at the margin, and no learning rate fixes it.</p>

<p>I was training a 1.7B model to recall 13-character opaque strings where a
one-character slip silently fetches a different economic series. Lookup tables
don’t belong in weights. So before writing that as an opinion, I measured it.</p>

<h3 id="measuring-the-alternative">Measuring the alternative</h3>

<p>First a correction to my own framing. I had been saying the catalog has 8,103
series, so memorization covers 0.4% of it. That number is misleading. The catalog
is really <strong>400 distinct items repeated across ~20 area and seasonal-adjustment
combinations</strong>. “Housing” matches 120 titles because the same concept appears for
US city average, Northeast, New England, Chicago, seasonally adjusted and not.</p>

<p>Constrained to the namespace every one of my seeds actually uses (US city
average, not seasonally adjusted): the corpus is <strong>400 rows, and the item name is
a unique key</strong>. That reframes the problem: not 8,103 codes to memorize, but 400
well-named items to <em>look up</em>. Still 11x more than the 35 I could afford to
teach, but a completely different kind of problem.</p>

<p>So I built a ~30-line BM25 index over those 400 item names and measured recall of
the correct series ID on the same 43 held-out questions:</p>

<table>
  <thead>
    <tr>
      <th>query</th>
      <th>recall@1</th>
      <th>recall@5</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>oracle (the gold item’s own name)</td>
      <td><strong>100%</strong></td>
      <td>100%</td>
    </tr>
    <tr>
      <td><strong>the raw user question, no model at all</strong></td>
      <td><strong>84.4%</strong></td>
      <td><strong>93.8%</strong></td>
    </tr>
  </tbody>
</table>

<p>The first row says the retriever has no ceiling problem: given a decent query it
is perfect. The second row is the uncomfortable one. Throwing the user’s raw
question at BM25 (no model, no training, no GPU, no adapter) retrieves the
correct series 84.4% of the time, against 88.4% for the fine-tuned model. Those
are within noise of each other.</p>

<p>A week of distillation is currently tied with a text search over 400 strings.</p>

<p>And the two questions raw BM25 misses are precisely the two my fine-tuned model
gets wrong:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>want  Medical care                          from "Show me healthcare CPI..."
want  Owners' equivalent rent of primary…   from "What did OER do between…"
</code></pre></div></div>

<p>Both are vocabulary gaps: “healthcare” and “OER” don’t appear in the official
item names. So the obvious move is to have the model rewrite the user’s words
into catalog vocabulary and let retrieval do the lookup. Much easier than
emitting a 13-character code from memory, and it generalizes to all 400 items.</p>

<p>I tested that too, and it does not work yet:</p>

<table>
  <thead>
    <tr>
      <th>query source</th>
      <th>recall@1</th>
      <th>recall@5</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>oracle (gold item name)</td>
      <td>100%</td>
      <td>100%</td>
    </tr>
    <tr>
      <td><strong>raw question, no model at all</strong></td>
      <td><strong>84.4%</strong></td>
      <td><strong>93.8%</strong></td>
    </tr>
    <tr>
      <td>base Qwen3-1.7B rewrites the question</td>
      <td>75.0%</td>
      <td>84.4%</td>
    </tr>
    <tr>
      <td><strong>my fine-tuned model rewrites it</strong></td>
      <td><strong>62.5%</strong></td>
      <td>71.9%</td>
    </tr>
  </tbody>
</table>

<p>Every model in the chain makes retrieval <em>worse</em>. The base model degrades good
queries. Asked about “tuition, other school fees, and childcare”, which is
verbatim the official item name, it helpfully rewrites it to “Education
expenses”. Asked about OER it expands the acronym to “Educational Resources
(OER)”, which is a real term from a different field entirely.</p>

<p>My fine-tuned model is worse still, and the reason is the interesting part.
Asked for a search phrase, it emits series IDs:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>"What did education and communication prices do…"  →  CUUR0000SAEDECU01
"Show me tuition and childcare CPI…"               →  CUUR0000SEEB
"Show me healthcare CPI…"                          →  search_phrase:CUUR0000SEMD Healthcare
</code></pre></div></div>

<p>It has specialized so hard on emitting codes that it can no longer paraphrase.
Note the second line: <code class="language-plaintext highlighter-rouge">CUUR0000SEEB</code> is the <em>correct answer</em>. The model knows the
mapping. It just can’t express it in a form the retriever can use, because the
index is over item names and it only speaks in IDs.</p>

<p>That is a useful negative result. It means retrieval cannot be bolted onto this
adapter: the memorization fine-tune destroyed the exact capability retrieval
depends on. The two-step system has to be trained from base, on traces that
include the search step, and Anthropic’s caution about small models and query
formulation is now something I’ve measured on my own data rather than quoted.</p>

<h3 id="act-7-the-fix-was-the-output-format">Act 7: the fix was the output format</h3>

<p>The agent ecosystem has converged on retrieval for a structurally identical
problem one level up. Hermes Agent, mcp-sieve, and various tool-router plugins
all replace “load every tool schema” with a <code class="language-plaintext highlighter-rouge">search</code> → <code class="language-plaintext highlighter-rouge">describe</code> → <code class="language-plaintext highlighter-rouge">call</code>
bridge. Anthropic’s MCP evaluations show accuracy <em>improving</em> when tools are
deferred rather than preloaded (49% → 74% on Opus 4) because large catalogs
cause decision paralysis.</p>

<p>Their bottleneck is too many tools. Mine was 400 values in one argument slot. So
I was about to build the two-step system, and then noticed I could get most of
the benefit by changing one thing: <strong>what the model is asked to emit.</strong></p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>get_series(series_id="CUUR0000SAF11")   →   get_series(item="Food at home")
</code></pre></div></div>

<p>The model names the item; forty lines of code resolve the name to an ID. This
works only because of the 400-item finding above: item name is a <em>unique key</em>
in that namespace, so the resolution is deterministic, not a guess.</p>

<p>Why it helps is the same reason the failures were what they were. <code class="language-plaintext highlighter-rouge">SAF11</code> vs
<code class="language-plaintext highlighter-rouge">SAF1</code> is a one-character discrimination. <em>“Food at home”</em> vs <em>“Food”</em> is a
semantic one, which is the kind of distinction a language model is actually built
to make. And a name fails <em>loudly</em>: <code class="language-plaintext highlighter-rouge">resolve_item</code> returns <code class="language-plaintext highlighter-rouge">None</code> for something
it can’t place, where one wrong character in a code silently fetches a different
series and returns plausible numbers.</p>

<p>Five seeds, same 43 held-out phrasings:</p>

<table>
  <thead>
    <tr>
      <th>target format</th>
      <th>exact</th>
      <th>sd</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">get_series(series_id="CUUR0000SAF11")</code></td>
      <td>88.8%</td>
      <td>3.0</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">get_series(item="Food at home")</code></td>
      <td>92.1%</td>
      <td>1.3</td>
    </tr>
    <tr>
      <td><strong>+ hierarchy-aware resolver</strong></td>
      <td><strong>94.4%</strong></td>
      <td><strong>1.3</strong></td>
    </tr>
  </tbody>
</table>

<p><strong>+5.6 points, t = 3.8, p ≈ 0.005</strong>, and seed variance more than halved.</p>

<!--
  Accuracy progression figure for the BLS distillation post.
  Usage in a post:  include bls-progression.html

  Self-contained and fully scoped under .bls-fig — no :root variables, no bare
  element selectors, so it cannot leak into the minimal-mistakes theme. Light
  only, matching the site (the theme defines no prefers-color-scheme). Ink and
  mono face follow assets/css/custom.scss (#1E1E1E, #888, JetBrains Mono).

  Blue ordinal ramp validated for a white surface: monotone lightness, adjacent
  dL >= 0.06, light end 2.11:1 vs surface, single hue.
-->
<figure class="bls-fig">
  <style>
    .bls-fig {
      /* Pull from site palette defined in custom.scss :root.
         Falls back to the same hex if the stylesheet isn't loaded (e.g. email). */
      --bls-ink:    var(--ink,    #1E1E1E);
      --bls-muted:  var(--muted-lt, #888888);
      --bls-rule:   var(--rule,   #e6e6e3);
      --bls-axis:   var(--baseline, #c9c9c4);
      --bls-hover:  rgba(42,120,214,0.08);
      --bls-1: var(--blue-1, #86b6ef);
      --bls-2: var(--blue-2, #5598e7);
      --bls-3: var(--blue-3, #2a78d6);
      --bls-4: var(--blue-4, #184f95);
      --bls-base: var(--baseline, #c9c9c4);
      display: block; /* override Minimal Mistakes figure { display: flex } */
      margin: 2.2em 0;
      padding: 0;
      font-family: system-ui, -apple-system, "Segoe UI", sans-serif;
      color: var(--bls-ink);
    }
    .bls-fig .bls-rows { width: 100%; } /* give the nested grid a definite width for 1fr */
    .bls-fig .bls-title { font-size: 0.95rem; font-weight: 700; margin: 0 0 0.2em; }
    .bls-fig .bls-sub { font-size: 0.8rem; color: var(--bls-muted); margin: 0 0 1.1em; line-height: 1.45; }
    .bls-fig .bls-rows { display: flex; flex-direction: column; gap: 2px; }
    .bls-fig .bls-row {
      display: grid; grid-template-columns: 190px 1fr; align-items: center;
      gap: 12px; padding: 2px 0; border-radius: 4px;
    }
    .bls-fig .bls-row:hover { background: var(--bls-hover); }
    .bls-fig .bls-lab { font-size: 0.76rem; line-height: 1.3; text-align: right; color: var(--bls-muted); }
    .bls-fig .bls-lab b { display: block; color: var(--bls-ink); font-weight: 700; font-size: 0.79rem; }
    .bls-fig .bls-track { position: relative; height: 24px; }
    .bls-fig .bls-bar { position: absolute; left: 0; top: 5px; height: 14px; border-radius: 0 4px 4px 0; }
    .bls-fig .bls-err { position: absolute; top: 11px; height: 2px; background: var(--bls-muted); }
    .bls-fig .bls-err::before, .bls-fig .bls-err::after {
      content: ""; position: absolute; top: -4px; width: 2px; height: 10px; background: var(--bls-muted);
    }
    .bls-fig .bls-err::before { left: 0; } .bls-fig .bls-err::after { right: 0; }
    .bls-fig .bls-val {
      position: absolute; top: 2px; font-size: 0.76rem; font-weight: 700; white-space: nowrap;
      font-family: "JetBrains Mono", ui-monospace, monospace; font-variant-numeric: tabular-nums;
    }
    .bls-fig .bls-val span { color: var(--bls-muted); font-weight: 400; }
    .bls-fig .bls-axis { display: grid; grid-template-columns: 190px 1fr; gap: 12px; margin-top: 4px; }
    .bls-fig .bls-ticks { position: relative; height: 18px; border-top: 1px solid var(--bls-axis); }
    .bls-fig .bls-tick {
      position: absolute; top: 3px; transform: translateX(-50%);
      font-size: 0.68rem; color: var(--bls-muted);
      font-family: "JetBrains Mono", ui-monospace, monospace; font-variant-numeric: tabular-nums;
    }
    .bls-fig .bls-key { display: flex; flex-wrap: wrap; gap: 14px; margin-top: 1em; font-size: 0.74rem; color: var(--bls-muted); }
    .bls-fig .bls-sw { display: inline-block; width: 10px; height: 10px; border-radius: 2px; margin-right: 5px; vertical-align: -1px; }
    .bls-fig .bls-tv { margin-top: 0.9em; }
    .bls-fig .bls-tv summary { font-size: 0.74rem; color: var(--bls-muted); cursor: pointer; }
    .bls-fig .bls-tv summary:focus-visible { outline: 2px solid var(--bls-3); outline-offset: 2px; }
    .bls-fig .bls-tv table { width: 100%; border-collapse: collapse; margin-top: 0.7em; font-size: 0.76rem; }
    .bls-fig .bls-tv th, .bls-fig .bls-tv td { padding: 5px 8px; border-bottom: 1px solid var(--bls-rule); text-align: left; }
    .bls-fig .bls-tv th { font-size: 0.68rem; letter-spacing: 0.05em; text-transform: uppercase; color: var(--bls-muted); }
    .bls-fig .bls-tv td.n {
      text-align: right; font-family: "JetBrains Mono", ui-monospace, monospace; font-variant-numeric: tabular-nums;
    }
    .bls-fig figcaption { font-size: 0.74rem; color: var(--bls-muted); margin-top: 1em; line-height: 1.5; }
    @media (max-width: 600px) {
      .bls-fig .bls-row, .bls-fig .bls-axis { grid-template-columns: 1fr; gap: 1px; }
      .bls-fig .bls-lab { text-align: left; }
    }
  </style>

  <p class="bls-title">Exact match on 43 held-out phrasings</p>
  <p class="bls-sub">Same held-out set throughout. The originally reported 93% is not shown — it was measured on training data. The final row (+alias) is the same model with the resolver alias table; see Act 8 on why it is not the headline number.</p>

  <div class="bls-rows">
    <div class="bls-row" title="base Qwen3-1.7B — 9.3%, no fine-tuning">
      <div class="bls-lab"><b>base Qwen3-1.7B</b>no fine-tuning</div>
      <div class="bls-track">
        <div class="bls-bar" style="width:9.3%;background:var(--bls-base)"></div>
        <div class="bls-val" style="left:calc(9.3% + 9px)">9.3%</div>
      </div>
    </div>
    <div class="bls-row" title="V2 honestly measured — 13.3%">
      <div class="bls-lab"><b>V2, honestly measured</b>the real starting point</div>
      <div class="bls-track">
        <div class="bls-bar" style="width:13.3%;background:var(--bls-1)"></div>
        <div class="bls-val" style="left:calc(13.3% + 9px)">13.3%</div>
      </div>
    </div>
    <div class="bls-row" title="BM25 over 400 item names — 84.4%, no model at all">
      <div class="bls-lab"><b>BM25 over 400 items</b>no model at all</div>
      <div class="bls-track">
        <div class="bls-bar" style="width:84.4%;background:var(--bls-base)"></div>
        <div class="bls-val" style="left:calc(84.4% + 9px)">84.4%</div>
      </div>
    </div>
    <div class="bls-row" title="+ prompt masking, all layers, rank 16 — 88.8% ±3.0 across 5 seeds">
      <div class="bls-lab"><b>+ prompt masking, r16</b>config fixes</div>
      <div class="bls-track">
        <div class="bls-bar" style="width:88.8%;background:var(--bls-2)"></div>
        <div class="bls-err" style="left:85.8%;width:6.0%"></div>
        <div class="bls-val" style="left:calc(91.8% + 9px)">88.8% <span>±3.0</span></div>
      </div>
    </div>
    <div class="bls-row" title="+ item-name targets — 92.1% ±1.3 across 5 seeds">
      <div class="bls-lab"><b>+ item-name targets</b>output format</div>
      <div class="bls-track">
        <div class="bls-bar" style="width:92.1%;background:var(--bls-3)"></div>
        <div class="bls-err" style="left:90.8%;width:2.6%"></div>
        <div class="bls-val" style="left:calc(93.4% + 9px)">92.1% <span>±1.3</span></div>
      </div>
    </div>
    <div class="bls-row" title="+ hierarchy-aware resolver — 94.4% ±1.3 across 5 seeds (current)">
      <div class="bls-lab"><b>+ hierarchy resolver</b>current</div>
      <div class="bls-track">
        <div class="bls-bar" style="width:94.4%;background:var(--bls-4)"></div>
        <div class="bls-err" style="left:93.1%;width:2.6%"></div>
        <div class="bls-val" style="left:calc(95.7% + 9px)">94.4% <span>±1.3</span></div>
      </div>
    </div>
    <div class="bls-row" title="+ resolver alias table — 96.7% ±1.3 — contaminated; see Act 8">
      <div class="bls-lab"><b>+ alias table</b><span style="color:var(--bls-muted);font-size:0.7rem">†see Act 8</span></div>
      <div class="bls-track">
        <div class="bls-bar" style="width:96.7%;background:var(--bls-base)"></div>
        <div class="bls-err" style="left:95.4%;width:2.6%"></div>
        <div class="bls-val" style="left:calc(97.3% + 9px)">96.7% <span>±1.3</span></div>
      </div>
    </div>
  </div>

  <div class="bls-axis">
    <div></div>
    <div class="bls-ticks">
      <span class="bls-tick" style="left:0%">0%</span>
      <span class="bls-tick" style="left:25%">25%</span>
      <span class="bls-tick" style="left:50%">50%</span>
      <span class="bls-tick" style="left:75%">75%</span>
      <span class="bls-tick" style="left:100%">100%</span>
    </div>
  </div>

  <div class="bls-key">
    <span><span class="bls-sw" style="background:var(--bls-4)"></span>fine-tuned, successive fixes</span>
    <span><span class="bls-sw" style="background:var(--bls-base)"></span>baseline, no fine-tuning</span>
    <span><span style="display:inline-block;width:13px;border-top:2px solid var(--bls-muted);vertical-align:4px;margin-right:5px"></span>±1 sd across 5 seeds</span>
  </div>

  <details class="bls-tv">
    <summary>Table view</summary>
    <table>
      <thead><tr><th>Stage</th><th class="n">Exact</th><th class="n">sd</th><th>Kind</th></tr></thead>
      <tbody>
        <tr><td>base Qwen3-1.7B</td><td class="n">9.3%</td><td class="n">&mdash;</td><td>baseline</td></tr>
        <tr><td>V2, honestly measured</td><td class="n">13.3%</td><td class="n">&mdash;</td><td>fine-tuned</td></tr>
        <tr><td>BM25 over 400 items, no model</td><td class="n">84.4%</td><td class="n">&mdash;</td><td>baseline</td></tr>
        <tr><td>+ prompt masking, all layers, rank 16</td><td class="n">88.8%</td><td class="n">±3.0</td><td>fine-tuned</td></tr>
        <tr><td>+ item-name targets</td><td class="n">92.1%</td><td class="n">±1.3</td><td>fine-tuned</td></tr>
        <tr><td>+ hierarchy-aware resolver</td><td class="n">94.4%</td><td class="n">±1.3</td><td>fine-tuned</td></tr>
        <tr><td>+ alias table <sup>†</sup></td><td class="n">96.7%</td><td class="n">±1.3</td><td>fine-tuned</td></tr>
      </tbody>
    </table>
  </details>

  <figcaption>
    Held-out phrasings absent from training, every referenced concept present.
    Fine-tuned rows are means over five training seeds; the three single-measurement
    rows carry no error bar rather than an invented one.
  </figcaption>
</figure>

<p>The hierarchy rule is small: when a bare term is a whole-word prefix of several
catalog entries, prefer the general one: “tobacco prices” means <em>Tobacco and
smoking products</em>, not <em>Tobacco products other than cigarettes</em>.</p>

<p>The only remaining failures at this point are “healthcare” → <em>Medical care</em> and
“OER” → <em>Owners’ equivalent rent</em>. Both are vocabulary gaps the resolver cannot
bridge without knowing which official term the user meant. The alias table was
the obvious fix, and the obvious problem: both failures are visible from
reading the test set.</p>

<p><strong>And the first run of this experiment reported 23.3%.</strong> A catastrophic
regression that would have killed the idea. It was a scoring bug: the held-out
file still stored raw <code class="language-plaintext highlighter-rouge">series_id</code> while the model now emitted item names, so
every <em>correct</em> answer scored wrong. I caught it because 10 of 43 passed, and
exactly 10 test items have no series ID. Gold is now rendered by the same
function that builds training targets, so the two cannot drift again.</p>

<p>That’s the second phantom regression in this project. Both times the number said
“your change broke it” and both times the number was the broken thing.</p>

<p>A note on the “100% accuracy” claims circulating in the tool-router ecosystem: I
read them. One is 10 out of 10 test cases. Another is 100% on synthetic
validation data generated the same way as its training data. Both are the same
species of number as my original 93%. Copy the architecture; don’t quote the
benchmarks.</p>

<h3 id="act-8-the-alias-table">Act 8: the alias table</h3>

<p>The two failures were “healthcare” → <em>Medical care</em> and “OER” → <em>Owners’
equivalent rent</em>. I said I was leaving them in. Then I shipped the alias table.</p>

<p>The number, same five adapters, model untouched:</p>

<table>
  <thead>
    <tr>
      <th> </th>
      <th>exact</th>
      <th>sd</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>fine-tuned, item names</td>
      <td>94.4%</td>
      <td>1.3</td>
    </tr>
    <tr>
      <td>+ resolver alias table</td>
      <td>96.7%</td>
      <td>1.3</td>
    </tr>
  </tbody>
</table>

<p>The accounting matters. There are roughly 50 aliases in the table: groceries,
gas, core CPI, airfare, public transport. Against this eval they move nothing;
its 43 questions don’t use that vocabulary. The entire +2.3pp is <code class="language-plaintext highlighter-rouge">healthcare</code>.
That alias was added knowing it was one of my two held-out failures.</p>

<p>Which makes it contaminated by the definition this post uses. I reported 94.4%
as the model’s number and put the alias row in a footnote.</p>

<p>The tension is that <code class="language-plaintext highlighter-rouge">healthcare</code> is simultaneously the most obvious synonym a
real user types and one of my two known test failures. Omitting it from a
user-facing resolver to protect a benchmark would be absurd. So no option was
both maximally useful and maximally clean.</p>

<p>What that means more broadly: once you’ve read your test failures, you can’t
un-read them. The honest options narrow permanently. The stricter path was to
build the alias table from general domain vocabulary, validate it on the val
split, and leave test untouched. I didn’t do that. I reported both numbers and
disclosed which one was contaminated. Honest, but not maximally rigorous, and
the distinction is real.</p>

<p>OER is still failing. The model emits <code class="language-plaintext highlighter-rouge">item="Education and communication"</code>, a
hallucination the resolver cannot repair. That is a model error, not a vocabulary
gap, and an alias does nothing for it.</p>

<p>One item remains genuinely open: enumerate all 400 names in context and let the
model select from a visible list rather than recall from memory. At 2,274 tokens
it fits, and it removes the lookup problem entirely for anything the model can
phrase correctly.</p>

<h2 id="what-i-would-do-differently">What I would do differently</h2>

<p><strong>Check your generation ratio before anything else.</strong> If completions are short
relative to prompts, mask the prompt. Otherwise your loss is mostly measuring
whether the model can copy a constant, and any conclusion you draw from that
curve is suspect.</p>

<p><strong>Report a distribution, not a run.</strong> Three seeds minimum, five if the difference
matters. At n=43 with 6 points of seed variance, a single number is close to
meaningless. If you only rerun the results you dislike, you will systematically
publish your luckiest draws.</p>

<p><strong>Write assertions, not intentions.</strong> “The splits are disjoint” is a claim. A
build that fails when they aren’t is a guarantee. Mine now refuses to run if any
series ID is absent from the catalog, if any example appears twice, if any
concept is missing from training, or if any label mentions dates its question
doesn’t.</p>

<p><strong>Split by what the model must actually generalize over.</strong> For a lookup task,
hold out phrasings. Holding out concepts measures the impossible.</p>

<p><strong>Ask what you’re making the model say, not just how you’re training it.</strong> The
single largest improvement in this project (+5.6 points, variance halved) came
from changing the output format from an opaque ID to a name, not from any
hyperparameter. Errors that are one character apart are hard for a language
model; errors that are semantically apart are easy. And a name can fail loudly
where a code fails silently.</p>

<p><strong>Measure the dumb baseline first.</strong> Thirty lines of BM25 over 400 strings scores
84.4% here. I was five days in before I knew that, and for all five of those days
I was comparing my results against nothing. If the baseline had come out ahead,
the right answer would have been to delete the model.</p>

<p><strong>Be most suspicious of results you like.</strong> Six of the seven bugs above inflated
my numbers. That is not coincidence: unpleasant surprises get debugged, pleasant
ones get published.</p>

<p><strong>Once you’ve read your test failures, you can’t un-read them.</strong> The honest
options narrow. If you’re building something that touches items you know failed,
validate it on val, not test. I didn’t; I reported both numbers and disclosed the
contamination. Honest, not maximally rigorous. The difference matters.</p>

<h2 id="results">Results</h2>

<p>43 held-out phrasings, none seen in training, every referenced concept present.
Mean over five training seeds, +/- one standard deviation across seeds:</p>

<table>
  <thead>
    <tr>
      <th> </th>
      <th>tool</th>
      <th>entity</th>
      <th>exact</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>base Qwen3-1.7B</td>
      <td>72.1%</td>
      <td>9.3%</td>
      <td>9.3%</td>
    </tr>
    <tr>
      <td>fine-tuned, emitting series IDs</td>
      <td>99.1% +/- 1.3</td>
      <td>89.8% +/- 2.1</td>
      <td>88.8% +/- 3.0</td>
    </tr>
    <tr>
      <td><strong>fine-tuned, emitting item names</strong></td>
      <td><strong>99.1% +/- 1.3</strong></td>
      <td><strong>94.4% +/- 1.3</strong></td>
      <td><strong>94.4% +/- 1.3</strong></td>
    </tr>
    <tr>
      <td>  + resolver alias table</td>
      <td>99.1% +/- 1.3</td>
      <td>96.7% +/- 1.3</td>
      <td>96.7% +/- 1.3 <sup>†</sup></td>
    </tr>
    <tr>
      <td>raw BM25 over 400 items, no model</td>
      <td>–</td>
      <td>–</td>
      <td>84.4%</td>
    </tr>
  </tbody>
</table>

<ul>
  <li><strong>tool</strong> — correct function chosen, out of six</li>
  <li><strong>entity</strong> — plus the load-bearing argument (series ID / query / survey)</li>
  <li><strong>exact</strong> — every argument identical</li>
</ul>

<p><sup>†</sup> Same five adapters, model untouched. The entire +2.3pp is the
<code class="language-plaintext highlighter-rouge">healthcare</code> alias, added knowing it was a held-out failure. See Act 8.</p>

<p>A 67MB adapter on an M4 Mini, for the 35 concepts it was taught. I want to
resist rounding that 99.1% to “100%”; four of five seeds hit 43/43 and the
temptation to quote the clean number is exactly the reflex this whole exercise
was about.</p>

<p>Worth stating the cost honestly: the config fix took the adapter from 19MB to
67MB, because it raised LoRA rank from 8 to 16 and adapted all 28 transformer
blocks instead of the last 16. So +4.5 points of exact match came with 3.5x the
adapter size. I did not sweep that tradeoff; rank 8 across all layers might get
most of the gain at half the size, and I have not checked.</p>

<p>The third row is there because a serious writeup should include the baseline
that embarrasses it.</p>

<p>One last operational note, learned the annoying way: I ran an evaluation
concurrently with a training job on the same GPU and got different numbers for
identical weights. Not non-determinism (two isolated runs are byte-identical),
but under memory pressure MLX returned different results rather than simply
running slower. I nearly wrote up a phantom regression from it. Score on a quiet
machine.</p>

<p>Code: <a href="https://github.com/kovashikawa/bls_data">github.com/kovashikawa/bls_data</a>.</p>

<h2 id="references">References</h2>

<ul>
  <li>Huerta-Enochian, M. &amp; Ko, S. Y. (2024). “Instruction Fine-Tuning: Does Prompt
Loss Matter?” EMNLP 2024. <a href="https://arxiv.org/abs/2401.13586">arXiv:2401.13586</a></li>
  <li>Guo, H., Dennis, S., Patil, R., &amp; Shabahang, K. (2026). “When Mean CE Fails:
Median CE Can Better Track Language Model Quality.”
<a href="https://arxiv.org/abs/2605.24667">arXiv:2605.24667</a>. Finds mean cross-entropy
rising while held-out accuracy stays near peak, in Qwen2.5-1.5B SFT.</li>
  <li>Apicella, A., Isgrò, F., Pollastro, A., &amp; Prevete, R. (2026). “Don’t stop me
now: Rethinking Validation Criteria for Model Parameter Selection.”
<a href="https://arxiv.org/abs/2602.22107">arXiv:2602.22107</a>. Finds early stopping on
validation <em>accuracy</em> performs worst for neural classifiers, favouring
loss-based criteria. Worth reading against this post: their objection is to
early stopping specifically, and they find post-hoc selection across all epochs
comparable to loss-based selection.</li>
  <li>Chatterjee, A., Renduchintala, H. S. V. N. S. K., Bhatia, S., &amp; Chakraborty, T.
(2025). “On the Effect of Instruction Tuning Loss on Generalization.”
<a href="https://arxiv.org/abs/2507.07817">arXiv:2507.07817</a>. Finds fully masking
prompt tokens is rarely optimal either; a low-to-moderate prompt weight usually
beats both extremes.</li>
  <li>Vaughn, D. (2024). “To Mask or Not to Mask: The Effect of Prompt Tokens on
Instruction Tuning.” Towards Data Science.</li>
</ul>]]></content><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;&quot;/assets/images/selfie-avatar.webp&quot;, &quot;bio&quot;=&gt;&quot;AI Engineer @ FUSE. Previously macro quant @ JGP. MIT MicroMasters. Washington, DC.&quot;, &quot;location&quot;=&gt;&quot;Washington DC&quot;, &quot;email&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Email&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-envelope-square&quot;}, {&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;}, {&quot;label&quot;=&gt;&quot;Twitter&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-twitter-square&quot;}, {&quot;label&quot;=&gt;&quot;Facebook&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-facebook-square&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/kovashikawa&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/in/rkovashikawa/&quot;}]}</name></author><category term="ai" /><category term="projects" /><category term="distillation" /><category term="lora" /><category term="mlx" /><category term="mcp" /><category term="tool-calling" /><category term="qwen" /><category term="evaluation" /><summary type="html"><![CDATA[I distilled a 1.7B model into a BLS tool-calling agent. Six measurement mistakes in a row, a val-loss trap that was a config bug, a fix that came from changing what the model says rather than how it trains, and a lesson about what happens when your test failures and your usability requirements overlap.]]></summary></entry><entry><title type="html">Winning the Google Cloud AI Hackathon: Building Multimodal Medical Annotation at Scale</title><link href="https://kovashikawa.com/projects/ai/google-cloud-ai-hackathon-winner/" rel="alternate" type="text/html" title="Winning the Google Cloud AI Hackathon: Building Multimodal Medical Annotation at Scale" /><published>2026-01-09T00:00:00+00:00</published><updated>2026-01-09T00:00:00+00:00</updated><id>https://kovashikawa.com/projects/ai/google-cloud-ai-hackathon-winner</id><content type="html" xml:base="https://kovashikawa.com/projects/ai/google-cloud-ai-hackathon-winner/"><![CDATA[<h2 id="overview">Overview</h2>

<p>Last month, I participated in the <strong>Google Cloud AI Hackathon</strong>, a fast-paced competition focused on building applied AI systems with real deployment constraints.</p>

<p>Our project, <strong>MedAnnotator</strong>, was selected as one of the winning teams.<br />
The official announcement is available here:<br />
👉 https://opendatascience.com/highlighting-the-winners-of-the-december-2025-google-cloud-ai-hackathon/</p>

<p>The problem we addressed is well known in medical AI:</p>

<blockquote>
  <p>Medical image annotation is expensive, slow, inconsistent, and difficult to scale, yet it remains a critical dependency for clinical workflows, research, and model development.</p>
</blockquote>

<p>Our goal was to design a system that prioritizes <strong>correctness, structure, and deployability</strong>, rather than novelty.</p>

<hr />

<h2 id="problem-definition">Problem Definition</h2>

<p>We identified three core bottlenecks in existing medical image annotation workflows:</p>

<ol>
  <li>
    <p><strong>Limited scalability</strong>: Annotation relies heavily on expert time, which does not scale with data volume.</p>
  </li>
  <li>
    <p><strong>High label variance</strong>: Inter-annotator disagreement introduces noise and reduces downstream model reliability.</p>
  </li>
  <li>
    <p><strong>Unstructured outputs</strong>: Free-text annotations are difficult to validate, audit, or integrate into pipelines.</p>
  </li>
</ol>

<p>The system was designed to generate <strong>structured, reviewable annotations</strong> from medical images while explicitly supporting human oversight.</p>

<hr />

<h2 id="system-architecture">System Architecture</h2>

<h3 id="1-two-tier-model-design">1. Two-Tier Model Design</h3>

<p>Rather than relying on a single model, we decomposed the task:</p>

<ul>
  <li>
    <p><strong>MedGemma</strong>: Responsible for domain-specific medical image understanding. This model handled image-level reasoning and feature extraction.</p>
  </li>
  <li>
    <p><strong>Gemini (API)</strong>: Used for validation, reasoning over MedGemma outputs, and producing structured, schema-compliant annotations.</p>
  </li>
</ul>

<p>This separation reduced coupling, improved iteration speed, and made failure modes easier to reason about.</p>

<hr />

<h3 id="2-structured-outputs-as-a-first-class-constraint">2. Structured Outputs as a First-Class Constraint</h3>

<p>All annotations were generated using predefined schemas.<br />
We deliberately avoided free-form text.</p>

<p>This enabled:</p>
<ul>
  <li>Deterministic validation</li>
  <li>Easier human review and correction</li>
  <li>Immediate downstream usability (storage, analytics, retraining)</li>
</ul>

<p>Structured outputs also simplified debugging during the demo and made model behavior more transparent.</p>

<hr />

<h3 id="3-human-in-the-loop-by-design">3. Human-in-the-Loop by Design</h3>

<p>The workflow explicitly supported:</p>
<ol>
  <li>Model-generated initial annotations</li>
  <li>Human review and edits</li>
  <li>Auditable final outputs</li>
</ol>

<p>In a healthcare context, this tradeoff is intentional.<br />
Reviewability and traceability matter more than full autonomy.</p>

<hr />

<h3 id="4-deployment-oriented-decisions">4. Deployment-Oriented Decisions</h3>

<p>MedGemma is computationally heavy, so we deployed it on cloud compute to:</p>
<ul>
  <li>Keep latency within interactive bounds</li>
  <li>Avoid blocking UI workflows</li>
  <li>Enable rapid iteration during the hackathon</li>
</ul>

<p>This allowed us to focus on system behavior and evaluation rather than infrastructure limitations.</p>

<hr />

<h2 id="why-the-project-worked">Why the Project Worked</h2>

<p>Several factors contributed to the outcome:</p>

<ul>
  <li>
    <p><strong>Tight scope</strong>: We focused on a concrete bottleneck instead of building a generic platform.</p>
  </li>
  <li>
    <p><strong>Clear system boundaries</strong>: Each component had a well-defined responsibility.</p>
  </li>
  <li>
    <p><strong>Production mindset</strong>: Latency, structure, and deployment were treated as core requirements, not afterthoughts.</p>
  </li>
  <li>
    <p><strong>Constraint-driven design</strong>: Hackathon limits forced architectural clarity.</p>
  </li>
</ul>

<hr />

<h2 id="key-takeaways">Key Takeaways</h2>

<p>This experience reinforced several principles that consistently hold in applied AI:</p>

<ul>
  <li>System design often matters more than model size.</li>
  <li>Structured outputs outperform clever prompts in production settings.</li>
  <li>Human-in-the-loop workflows remain essential in high-stakes domains.</li>
  <li>Deployment constraints improve, rather than limit, design quality.</li>
</ul>

<hr />

<h2 id="future-extensions">Future Extensions</h2>

<p>The prototype can be extended in several directions:</p>
<ul>
  <li>Batch annotation pipelines for large-scale datasets</li>
  <li>Integration with PACS or clinical data systems</li>
  <li>Active learning loops using corrected annotations</li>
  <li>Quantitative evaluation tooling for annotation quality and drift</li>
</ul>

<p>If you are working on applied multimodal systems, especially in healthcare, I am always open to technical discussions.</p>

<hr />

<p><em>Thanks to my teammates, ODSC, and Google Cloud for running a technically rigorous and well-executed hackathon.</em></p>]]></content><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;&quot;/assets/images/selfie-avatar.webp&quot;, &quot;bio&quot;=&gt;&quot;AI Engineer @ FUSE. Previously macro quant @ JGP. MIT MicroMasters. Washington, DC.&quot;, &quot;location&quot;=&gt;&quot;Washington DC&quot;, &quot;email&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Email&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-envelope-square&quot;}, {&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;}, {&quot;label&quot;=&gt;&quot;Twitter&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-twitter-square&quot;}, {&quot;label&quot;=&gt;&quot;Facebook&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-facebook-square&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/kovashikawa&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/in/rkovashikawa/&quot;}]}</name></author><category term="projects" /><category term="ai" /><category term="hackathon" /><category term="google-cloud" /><category term="multimodal-ai" /><category term="llms" /><category term="computer-vision" /><category term="gemma" /><category term="healthcare" /><summary type="html"><![CDATA[How we designed a production-oriented multimodal system for medical image annotation and won the December 2025 Google Cloud AI Hackathon.]]></summary></entry><entry><title type="html">Analysing the impact of AI jobs at World Bank DataDive 2025</title><link href="https://kovashikawa.com/ai/data-science/labor/worldbank-datadive-hackathon/" rel="alternate" type="text/html" title="Analysing the impact of AI jobs at World Bank DataDive 2025" /><published>2025-12-08T00:00:00+00:00</published><updated>2025-12-08T00:00:00+00:00</updated><id>https://kovashikawa.com/ai/data-science/labor/worldbank-datadive-hackathon</id><content type="html" xml:base="https://kovashikawa.com/ai/data-science/labor/worldbank-datadive-hackathon/"><![CDATA[<h1 id="building-jobslens-ai-at-world-bank-datadive-2025">Building JobsLens AI at World Bank DataDive 2025</h1>

<p>Last week, I participated in the <strong>World Bank DataDive 2025</strong>, a 8-hour hackathon focused on extracting insights from global labor market data. Our team of five—Maryam Shahbaz Ali, Paul Suhwan Lee, Stevens Cadet, Yingquan Li, and myself—tackled <a href="https://www.dc2.org/datadive"><strong>Challenge 8: Digital/AI Job Demand &amp; Supply Analysis</strong></a>.</p>

<p><strong>TL;DR</strong>: We built <a href="https://huggingface.co/spaces/rkovashikawa/jobslens-ai">JobsLens AI</a>, an interactive Gradio dashboard that visualizes AI job market trends across countries, identifies skills gaps, and forecasts future demand.</p>

<hr />

<h2 id="the-challenge">The Challenge</h2>

<p>The World Bank wanted to understand <strong>where digital and AI job opportunities are rising or lagging</strong> across countries, industries, and skill types. The goal was to help policymakers identify:</p>

<ul>
  <li>Countries with high AI job demand but low skilled workforce supply (skills gaps)</li>
  <li>Emerging markets with growing AI sectors</li>
  <li>Trends in AI talent concentration and investment</li>
</ul>

<hr />

<h2 id="the-data">The Data</h2>

<p>We worked with two primary datasets:</p>

<ol>
  <li><strong>Stanford HAI AI Index 2025</strong> (528 observations, 232 columns)
    <ul>
      <li>Panel data spanning 2017-2024 for 66 countries</li>
      <li>Key metrics: AI job postings, skills penetration, talent concentration, patents, investment</li>
      <li>Target variable: <code class="language-plaintext highlighter-rouge">ai_job_postings_perc_of_all_job_postings</code></li>
    </ul>
  </li>
  <li><strong>World Bank Global Labor Database</strong> (154 countries, cross-sectional)
    <ul>
      <li>Employment rates by sector and education level</li>
      <li>Labor force participation, unemployment rates</li>
      <li>Demographics and population statistics</li>
    </ul>
  </li>
</ol>

<hr />

<h2 id="technical-approach">Technical Approach</h2>

<h3 id="data-integration-challenge">Data Integration Challenge</h3>

<p>The fundamental problem was joining <strong>cross-sectional predictors</strong> (World Bank) with <strong>panel outcomes</strong> (HAI):</p>

<ul>
  <li>HAI: Multiple years per country (2017-2024)</li>
  <li>World Bank: Single observation per country</li>
</ul>

<p>We used a <strong>many-to-one merge</strong>, assuming labor market characteristics remain relatively stable over time.</p>

<h3 id="dashboard-development">Dashboard Development</h3>

<p>Built with <strong>Gradio 4.0</strong> and <strong>Plotly</strong>, the dashboard features 5 interactive tabs:</p>

<ol>
  <li><strong>🌍 Global Overview</strong>: Choropleth map showing supply-demand gaps</li>
  <li><strong>📊 Country Comparison</strong>: Multi-country trend analysis (default: Brazil, Argentina, USA, Japan)</li>
  <li><strong>🏆 Rankings</strong>: Top countries by AI job demand, skills supply, investment, startups</li>
  <li><strong>🔮 Forecasts</strong>: 2024-2025 projections built with a minimalist linear trend baseline after richer models proved unreliable under the time and data constraints</li>
  <li><strong>🔬 Country Deep Dive</strong>: Detailed statistics for individual countries</li>
</ol>

<hr />

<h2 id="deployment">Deployment</h2>

<p>We deployed to <strong>Hugging Face Spaces</strong> for free hosting:</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c"># Project structure</span>
Team_Projects/JobsLens_AI/
├── app.py                    <span class="c"># Gradio dashboard</span>
├── data/
│   ├── hai_full_database.csv
│   └── forecasts_2024_2025.csv
├── hf_space/                 <span class="c"># Git repo for HF Spaces</span>
└── notebooks/
    └── EDA.ipynb
</code></pre></div></div>

<p><strong>Tech stack</strong>:</p>
<ul>
  <li><strong>Frontend</strong>: Gradio 4.0 with custom Plotly visualizations</li>
  <li><strong>Data</strong>: pandas, numpy</li>
  <li><strong>ML</strong>: scikit-learn (Ridge, Random Forest, Gradient Boosting)</li>
  <li><strong>Deployment</strong>: Hugging Face Spaces (free CPU tier)</li>
</ul>

<hr />

<h2 id="key-insights-from-the-data">Key Insights from the Data</h2>

<h3 id="skills-gap-leaders-high-demandlow-supply">Skills Gap Leaders (High Demand/Low Supply)</h3>
<ul>
  <li><strong>Switzerland</strong>: 1.41% AI job demand (highest forecasted)</li>
  <li><strong>Netherlands</strong>: 1.25%</li>
  <li><strong>United States</strong>: 1.36%</li>
</ul>

<h3 id="supply-demand-balance">Supply-Demand Balance</h3>
<p>The global AI job market shows:</p>
<ul>
  <li>Growing demand outpacing skills supply in developed economies</li>
  <li>Emerging markets (Brazil, Mexico) showing moderate growth</li>
  <li>Investment concentrated in US, China, EU</li>
</ul>

<h3 id="temporal-trends-2017-2024">Temporal Trends (2017-2024)</h3>
<ul>
  <li>AI job postings grew ~3-5% annually in top markets</li>
  <li>Skills penetration lagging by 2-3 years</li>
  <li>COVID-19 accelerated remote AI hiring (2020-2021 spike)</li>
</ul>

<hr />

<h2 id="what-i-learned">What I Learned</h2>

<h3 id="technical">Technical</h3>
<ol>
  <li><strong>Small data ≠ deep learning</strong>: With only 83 observations, simpler models (Ridge) outperformed complex ones</li>
  <li><strong>Feature engineering matters</strong>: Missing 14/24 features severely limited model performance</li>
  <li><strong>Cross-sectional + panel merging</strong>: Requires careful assumptions about temporal stability</li>
  <li><strong>Gradio is powerful</strong>: Built a full interactive dashboard in &lt;2 hours</li>
</ol>

<h3 id="hackathon-strategy">Hackathon Strategy</h3>
<ol>
  <li><strong>Iterate fast</strong>: We went from EDA → model → dashboard → deployment in &lt;8 hours</li>
  <li><strong>Visualize early</strong>: Interactive plots helped us spot data issues quickly</li>
  <li><strong>Deploy early</strong>: Having a live URL motivated us to polish the final product</li>
  <li><strong>Team roles</strong>: Divided into data/modeling (me + Yingquan), visualization (Paul), storytelling (Maryam), infrastructure (Stevens)</li>
</ol>

<hr />

<h2 id="results--recognition">Results &amp; Recognition</h2>

<ul>
  <li><strong>Live Dashboard</strong>: <a href="https://huggingface.co/spaces/rkovashikawa/jobslens-ai">jobslens-ai on HF Spaces</a></li>
  <li><strong>Presentation</strong>: <a href="https://www.canva.com/design/DAG6ftQIfSA/pW-b5By0racKOTOgHcUY6A/view">Canva slides</a></li>
  <li><strong>Code</strong>: <a href="https://github.com/datacommunitydc/DataDive25">GitHub repository</a></li>
</ul>

<p>Forecasting was the toughest part: with only a few hours to wrangle unfamiliar data, every richer model we tried failed to be accurate, robust, or remotely reliable. I fell back to a simplistic linear trend model so the projections stayed within clearly understood limitations. Even so, the <strong>exploratory dashboard successfully visualizes historical trends</strong> and provides policymakers with actionable insights on global AI skills gaps.</p>

<hr />

<h2 id="future-improvements">Future Improvements</h2>

<p>If I were to continue this project:</p>

<ol>
  <li><strong>Use “All” subsample</strong> instead of “Urban” to get 24/24 features</li>
  <li><strong>Time-series models</strong>: ARIMA or exponential smoothing might outperform ML on small data</li>
  <li><strong>External data</strong>: Add OECD education stats, GitHub commit data by country</li>
  <li><strong>Sector analysis</strong>: Break down AI jobs by industry (healthcare, finance, etc.)</li>
  <li><strong>Interactive filtering</strong>: Let users select custom country groups and metrics</li>
</ol>

<hr />

<h2 id="try-it-yourself">Try It Yourself</h2>

<p>🔗 <strong>Live Dashboard</strong>: <a href="https://huggingface.co/spaces/rkovashikawa/jobslens-ai">https://huggingface.co/spaces/rkovashikawa/jobslens-ai</a></p>

<p>Explore AI job trends for 66 countries, compare Brazil vs. USA, or dive into Switzerland’s AI ecosystem. The dashboard is free and requires no login.</p>

<hr />

<p><em>Hackathon: World Bank DataDive 2025</em>
<em>Team: JobsLens AI</em>
<em>Duration: 8 hours</em>
<em>Tech: Python, Gradio, Plotly, scikit-learn, Hugging Face Spaces</em></p>]]></content><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;&quot;/assets/images/selfie-avatar.webp&quot;, &quot;bio&quot;=&gt;&quot;AI Engineer @ FUSE. Previously macro quant @ JGP. MIT MicroMasters. Washington, DC.&quot;, &quot;location&quot;=&gt;&quot;Washington DC&quot;, &quot;email&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Email&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-envelope-square&quot;}, {&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;}, {&quot;label&quot;=&gt;&quot;Twitter&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-twitter-square&quot;}, {&quot;label&quot;=&gt;&quot;Facebook&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-facebook-square&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/kovashikawa&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/in/rkovashikawa/&quot;}]}</name></author><category term="ai" /><category term="data-science" /><category term="labor" /><category term="hackathon" /><category term="machine-learning" /><category term="dashboard" /><category term="world-bank" /><category term="gradio" /><category term="economics" /><summary type="html"><![CDATA[Our team built an interactive dashboard analyzing global AI job market trends and skills gaps across 66 countries using Stanford HAI data and World Bank labor indicators.]]></summary></entry><entry><title type="html">MCP: Inside the Data Layer and the Choice of JSON-RPC 2.0</title><link href="https://kovashikawa.com/ai/mcp-notes/" rel="alternate" type="text/html" title="MCP: Inside the Data Layer and the Choice of JSON-RPC 2.0" /><published>2025-10-15T21:18:00+00:00</published><updated>2026-08-26T00:00:00+00:00</updated><id>https://kovashikawa.com/ai/mcp-notes</id><content type="html" xml:base="https://kovashikawa.com/ai/mcp-notes/"><![CDATA[<h2 id="what-mcp-is-for">What MCP is for</h2>

<p>The Model Context Protocol (MCP) is an open specification, released by Anthropic in November 2024, that defines a single interface between AI models and external systems. A client (an agent) and a server (a tool, database, or API wrapper) exchange protocol versions and capabilities at connect time, then communicate through four primitives: tools, resources, prompts, and notifications. That structure is JSON-RPC 2.0.</p>

<h3 id="why-json-rpc-20-over-plain-json">Why JSON-RPC 2.0 over plain JSON</h3>

<p>Plain JSON over HTTP gives you a body and a status code. Nothing defines the method name, nothing links a response to the request that produced it, and there is no agreed error shape.</p>

<p>JSON-RPC 2.0 fixes each of those. Every message carries <code class="language-plaintext highlighter-rouge">jsonrpc: "2.0"</code>, a method, params, and an id that correlates responses with requests. A server can reply asynchronously and out of order, and can issue its own calls back to the client, which is what lets a long-running tool stream progress while another call completes.</p>

<p>Here is a minimal request:</p>

<div class="language-json highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="p">{</span><span class="w">
  </span><span class="nl">"jsonrpc"</span><span class="p">:</span><span class="w"> </span><span class="s2">"2.0"</span><span class="p">,</span><span class="w">
  </span><span class="nl">"id"</span><span class="p">:</span><span class="w"> </span><span class="mi">1</span><span class="p">,</span><span class="w">
  </span><span class="nl">"method"</span><span class="p">:</span><span class="w"> </span><span class="s2">"tools/call"</span><span class="p">,</span><span class="w">
  </span><span class="nl">"params"</span><span class="p">:</span><span class="w"> </span><span class="p">{</span><span class="w"> </span><span class="nl">"name"</span><span class="p">:</span><span class="w"> </span><span class="s2">"get_weather"</span><span class="p">,</span><span class="w"> </span><span class="nl">"arguments"</span><span class="p">:</span><span class="w"> </span><span class="p">{</span><span class="w"> </span><span class="nl">"city"</span><span class="p">:</span><span class="w"> </span><span class="s2">"São Paulo"</span><span class="w"> </span><span class="p">}</span><span class="w"> </span><span class="p">}</span><span class="w">
</span><span class="p">}</span><span class="w">
</span></code></pre></div></div>

<p>And the matching response:</p>

<div class="language-json highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="p">{</span><span class="w">
  </span><span class="nl">"jsonrpc"</span><span class="p">:</span><span class="w"> </span><span class="s2">"2.0"</span><span class="p">,</span><span class="w">
  </span><span class="nl">"id"</span><span class="p">:</span><span class="w"> </span><span class="mi">1</span><span class="p">,</span><span class="w">
  </span><span class="nl">"result"</span><span class="p">:</span><span class="w"> </span><span class="p">{</span><span class="w"> </span><span class="nl">"temp"</span><span class="p">:</span><span class="w"> </span><span class="mf">26.1</span><span class="p">,</span><span class="w"> </span><span class="nl">"unit"</span><span class="p">:</span><span class="w"> </span><span class="s2">"C"</span><span class="w"> </span><span class="p">}</span><span class="w">
</span><span class="p">}</span><span class="w">
</span></code></pre></div></div>

<p>Notifications are JSON-RPC messages without an <code class="language-plaintext highlighter-rouge">id</code>; the sender expects no reply. MCP uses them for progress and log messages. Batch requests send several calls in one array and receive the responses together. Both map directly to agent workloads, where a model issues several tool calls in a row.</p>

<h3 id="the-transports-stdio-and-streamable-http">The transports: stdio and Streamable HTTP</h3>

<p>JSON-RPC defines message framing independently of transport, so the same logical message moves over different channels. MCP specifies two.</p>

<ul>
  <li>stdio: the client spawns the server as a child process and exchanges JSON-RPC messages over standard input and output. No HTTP stack, no serialization boundary. The default for local servers.</li>
  <li>Streamable HTTP: a single HTTP connection the server can stream incremental responses over. It replaced the earlier HTTP + SSE transport, which the spec deprecated in March 2025. The server can push updates over the same connection, which matters for tools that run for seconds rather than milliseconds.</li>
</ul>

<p>The transport changes the deployment model, not the messages. The same <code class="language-plaintext highlighter-rouge">tools/call</code> payload works over stdio locally and over Streamable HTTP remotely.</p>

<h3 id="the-data-layer-tools-resources-and-capability-negotiation">The data layer: tools, resources, and capability negotiation</h3>

<p>The data layer defines what the messages can carry. At connect time, client and server exchange protocol versions and capabilities. The server then publishes its surface:</p>

<ul>
  <li>tools: functions the model can call, each described by name, description, and an input JSON Schema</li>
  <li>resources: readable data surfaces, identified by URI and MIME type</li>
  <li>prompts: reusable prompt templates the client can request</li>
  <li>notifications: server-initiated events the client can register for</li>
</ul>

<p>This is what makes the protocol composable. A model never needs per-API glue: the schema of every tool arrives during the handshake, so a client can construct valid arguments and validate the result without knowing anything about the underlying service.</p>

<h3 id="errors-are-part-of-the-contract">Errors are part of the contract</h3>

<p>JSON-RPC 2.0 reserves a base set of error codes, and MCP adds its own on top. Every failed call returns an object with <code class="language-plaintext highlighter-rouge">code</code>, <code class="language-plaintext highlighter-rouge">message</code>, and optional diagnostic <code class="language-plaintext highlighter-rouge">data</code>. When an agent chains several tool calls, a typed error lets it branch on the failure mode instead of guessing from a bare HTTP status.</p>]]></content><author><name>{&quot;name&quot;=&gt;nil, &quot;avatar&quot;=&gt;&quot;/assets/images/selfie-avatar.webp&quot;, &quot;bio&quot;=&gt;&quot;AI Engineer @ FUSE. Previously macro quant @ JGP. MIT MicroMasters. Washington, DC.&quot;, &quot;location&quot;=&gt;&quot;Washington DC&quot;, &quot;email&quot;=&gt;nil, &quot;links&quot;=&gt;[{&quot;label&quot;=&gt;&quot;Email&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-envelope-square&quot;}, {&quot;label&quot;=&gt;&quot;Website&quot;, &quot;icon&quot;=&gt;&quot;fas fa-fw fa-link&quot;}, {&quot;label&quot;=&gt;&quot;Twitter&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-twitter-square&quot;}, {&quot;label&quot;=&gt;&quot;Facebook&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-facebook-square&quot;}, {&quot;label&quot;=&gt;&quot;GitHub&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-github&quot;, &quot;url&quot;=&gt;&quot;https://github.com/kovashikawa&quot;}, {&quot;label&quot;=&gt;&quot;LinkedIn&quot;, &quot;icon&quot;=&gt;&quot;fab fa-fw fa-linkedin&quot;, &quot;url&quot;=&gt;&quot;https://www.linkedin.com/in/rkovashikawa/&quot;}]}</name></author><category term="ai" /><category term="MCP" /><category term="JSON-RPC" /><category term="LLM" /><category term="Protocols" /><summary type="html"><![CDATA[why MCP standardizes model-to-tool communication on JSON-RPC 2.0]]></summary></entry><entry><title type="html">Architects, Anthills, and AI: A Nobel Prize’s Lesson on Scientific Progress</title><link href="https://kovashikawa.com/ai/architecs-anthill-ai/" rel="alternate" type="text/html" title="Architects, Anthills, and AI: A Nobel Prize’s Lesson on Scientific Progress" /><published>2025-10-09T00:00:00+00:00</published><updated>2025-10-09T00:00:00+00:00</updated><id>https://kovashikawa.com/ai/architecs-anthill-ai</id><content type="html" xml:base="https://kovashikawa.com/ai/architecs-anthill-ai/"><![CDATA[<p>The 2025 Nobel Prize in Chemistry has been awarded to Susumu Kitagawa, Richard Robson, and Omar M. Yaghi for their creation of an entirely new class of materials: <strong>metal-organic frameworks (MOFs)</strong> <a href="https://www.nobelprize.org/prizes/chemistry/2025/popular-information/">[1]</a>.</p>

<p>These remarkable materials can be thought of as molecular sponges or programmable crystals. Built from metal ions and organic linkers, their defining feature is a vast internal porosity, allowing just a few grams to have a surface area the size of a football pitch. This unique property has unlocked potential solutions to some of humanity’s biggest challenges, from capturing carbon dioxide and harvesting water from desert air to filtering pollutants and delivering drugs.</p>

<p><img src="https://www.nobelprize.org/images/178061-large-2x.jpg" alt="A diagram showing various molecular structures of MOFs and their specific applications, such as capturing water, storing hydrogen, and filtering pollutants." /></p>
<blockquote>
  <p><strong>Figure:</strong> The versatility of MOFs is staggering. Different frameworks are tailored for specific tasks, such as capturing water from air (MOF-303), storing hydrogen fuel (MIL-101), removing PFAS pollutants from water (UiO-67), and absorbing industrial CO₂ emissions (CALF-20), showcasing the breadth of research built upon the laureates’ foundational work.</p>
</blockquote>

<p>The laureates’ journey is a story of visionary science. It began with Richard Robson’s foundational idea to build large, ordered networks. Building on this, Susumu Kitagawa developed the first stable, functional MOFs and crucially envisioned them as flexible structures that could “breathe.” Omar M. Yaghi then perfected the field by demonstrating a “Lego-like” approach to their creation, allowing chemists to rationally design frameworks with tailored properties, including his iconic and exceptionally spacious MOF-5.</p>

<p>This tale of three pioneers seems like a classic case of the “Newtonian” view of science—that progress is driven by a few brilliant giants who provide the shoulders for others to stand on. However, the story of MOFs also powerfully illustrates a competing idea: the <strong>Ortega hypothesis</strong> <a href="https://en.wikipedia.org/wiki/Ortega_hypothesis">[2]</a>. This theory suggests that major breakthroughs don’t happen in a vacuum. Instead, they arise from the slow, cumulative, and often anonymous contributions of a great number of working scientists.</p>

<p>This is where Susumu Kitagawa’s personal motto—to find “the usefulness of useless”—becomes profoundly insightful. Science is filled with countless small experiments and niche discoveries that seem purposeless in isolation. The Ortega hypothesis argues that progress depends on this vast, unglamorous foundation of “useless” knowledge. It’s an ecosystem of modest findings that a few brilliant minds can eventually synthesize into a revolutionary concept. The laureates themselves were building on decades of fundamental chemistry. Kitagawa’s philosophy wasn’t just a personal quirk; it was an acknowledgment of the very nature of scientific progress.</p>

<p>We see this process in action right now. Today, labs across the globe are using the MOF toolkit to design countless specialized frameworks <a href="https://doi.org/10.1021/cr300014x">[3]</a>. Each team works on a narrow problem—a single pollutant, a specific gas, a unique catalyst. Each contribution is a small, vital step forward, an anthill of progress built upon the architectural plans the laureates drew.</p>

<p>The laureates’ work defined a design space so vast it became a big data problem, which global research has populated and AI can now analyze.</p>

<p>Instead of relying on slow physical synthesis, AI models computationally screen millions of potential MOFs to pinpoint the most effective candidates for tasks like carbon capture. This is the new paradigm for discovery: human genius creates the field, collective effort generates the data, and AI provides the scaling engine to find the solutions within.</p>]]></content><author><name>rafael</name></author><category term="ai" /><category term="philosophy" /><category term="chemistry" /><summary type="html"><![CDATA[chemistry nobel prize offers a profound lesson on scientific progress, with the AI acceleration]]></summary></entry></feed>