0. Bottom line
- What this is about. An influence function estimates, without retraining, how much a model's behaviour on a query would change if one training example were removed: −∇f(θ*)ᵀ H⁻¹ ∇L(z, θ*), where the inverse Hessian rescales the two gradients by the curvature of the loss so that generic high-norm directions and facts repeated across many examples are discounted. Everything below is about making that inverse cheap and accurate (the Hessian and the solver), making the scan over training examples affordable (sketching), and whether the user's Gaussian factoring helps with the first.
- What exists. A complete MNIST influence lab on the 11.97M-parameter Ciresan MLP (exact GGN products, FP64-validated EK-FAC, PCG reference, ASTRA, 24 retrains, $0.34), four live tutorial pages including a cost calculator, an atlas, a laptop mini-lab, and four analysis reports. The research is done; the blog post and the application it was for are not, and the lab still has no git history.
- Hessians. Grosse's team and Hong, Eschenhagen, Mlodozeniec and Turner (2025/2026) agree that better curvature raises LDS (the linear datamodeling score: the Spearman correlation between predicted and measured loss changes over random retrained data subsets). The actionable finding, from the mini-lab and confirmed today on the lab's real last layer, is that at convergence K-FAC's within-layer error is a basis problem: K-FAC 2.63, EK-FAC 0.87 with 76% of the block's mass unreachable by any eigenvalue correction, optimal two-term Kronecker sum 0.37 (relative Frobenius). EK-FAC's diagonal correction applied in the optimal rank-1 Kronecker basis instead of K-FAC's marginal basis gives conditioning 1.67 and iHVP cosine 0.996 at EK-FAC cost, in sample. Hong's thesis asks for exactly this rung; nobody has tested it for influence. Out of sample the picture changes: at 4,096 calibration examples only the reweighted single term and EK-FAC survive (the optimal-basis rung is then worse than EK-FAC); at 25,000 all Kronecker-type models land within 0.79–0.87 of a held-out block (basis enrichment buys about 5% over EK-FAC, not 2.4×), and the only estimators that do better keep the highest-curvature examples exactly (0.43, better than the sample block itself at 0.47). The block of a converged layer is dominated by a few thousand examples, which argues for spending the estimation phase on all training examples with a Kronecker bulk plus an exact head, not on richer Kronecker structure fitted to a small calibration set (§2.3).
- ASTRA and the lab. ASTRA improved the solve (rank agreement with PCG 0.380 → 0.751) but 24 single-seed half-subset retrains could not see any method (all LDS intervals include zero; two LOO seeds agree at −0.007). The lab's damping also made the whole system well conditioned (κ ≤ 9.5), so its solver comparison says nothing about preconditioners. Ground truth must be ELSO-style subsets (expected leave-some-out: many random half-subsets, each retrained with many seeds and averaged) or
dattri's pre-retrained models. - Cost. Three phases: a one-time fit, a per-query iHVP, a per-candidate scan. At Llama-3-8B shape the EK-FAC iHVP is about four gradient computations, ASTRA is 4,841× that, a 10M-candidate scan is 702 H100-hours, a full-corpus scan equals pretraining. Once retrieval is a sketch lookup, the iHVP is more than 99.9% of the per-query bill, so preconditioner quality becomes the line item. Two-term Kronecker sums and the user's T-Sylvester inverse keep iHVPs at K-FAC price; anything with more than two terms needs an iterative solve per query.
- Sketching. The whitening argument against low-dimensional projections is now a theorem with a size (Hu, Hu, Ma, Zhao, Feb 2026): sketch dimension must scale with the effective dimension of the curvature at the damping scale. LoGra's 6,500× is the per-query phase only and is bought with 3.5 TB of disk; every deployed projection method pays in quality; no public method states a top-k recall guarantee (a not-found from five searches, and the step from score bounds to recall is known from the nearest-neighbour literature), which is the gap Grosse described. Sketching also pays on the estimation side, where random probes let structured curvature be fit to exact GGN products.
- The weak (Gaussian) factoring idea. As stated it adds nothing: the cross-covariance it introduces is the layer's mean gradient: identically zero in expectation for the sampled-label Fisher (today's 0.3% is one-sample Monte-Carlo noise), and 0.08% (mini-lab) or 0.07% (today) for the empirical Fisher at convergence; Martens & Grosse 2015 Appendix A plus Lemma 4 already contain the cumulant argument. Applied conditionally it derives a Kronecker-sum family, and today's experiment says which conditioning variable matters: curvature mass, not class or confidence. Exact terms for the 1% of examples carrying 69% of the curvature plus a curvature-weighted K-FAC for the rest reaches 0.19 with a rank-360 Woodbury correction (the top 5%: 0.066, and conditioning 1.4 at λ = 0.001 against 14.7 unpreconditioned); the weighted single Kronecker term alone beats EK-FAC (0.80). Out of sample at 4,096 examples the heads do no better than the raw sample block while the weighted single term survives (0.95 against EK-FAC's 0.99); at 25,000 the heads are the best estimators tested (0.43) and the only ones that improve conditioning at λ = 0.001 (κ 2.8 against 5.7 unpreconditioned). The user's Sylvester/Woodbury toolkit is the inverse primitive for this family, with one measured warning: the two-term closed form with diagonally approximated damping returns a useless vector (cosine 0.004), while with factored Tikhonov damping folded into the first pair it recovers the exactly damped solve on one split (cosine 0.9999) but not on the curvature-mass split (cosine 0.93–0.97), so the closed form serves as a preconditioner rather than as an answer.
- Caveats that bound all of this. Frobenius error, iHVP cosine and Hong's inverse-application residual rank the same methods differently at small damping; only LDS arbitrates. Everything measured today is one layer of one converged MLP.
- Next, in the restart summary's order.
git init(ten minutes); apply to the Research Engineer / Scientist, Alignment posting (twenty minutes; it does not need the post); send the post raw in Roger's "catch up?" thread, framed as in §5.5. Only then the per-layer, per-epoch sweep with both metrics and the new rungs (weighted K-FAC, EK-FAC in the optimal basis, exact-top-examples plus K-FAC), scored out of sample and against ELSO ground truth.
1. What exists after eight days
Two sittings (Wed 9 Sep, Sat 12 Sep) plus a scheduled overnight pass (Mon 14 → Tue 15 Sep) and three analysis passes on Thu 17 Sep produced the following. The research side is finished; the blog post and the application it was for are not (see §6).
| Artifact | Where | What it is |
|---|---|---|
| MNIST ASTRA lab | ~/git/influence-functions (this directory): tutorial.md, results/analysis.md, curvature.py, runner.py; W&B project yaroslavvb/Influence |
Full 11,972,510-parameter Ciresan MLP (784–2500–2000–1500–1000–500–10) from gradient-dissent; exact softmax-GGN products; FP64-validated EK-FAC; deterministic PCG reference on the full 50,000-example GGN; ASTRA; 24 half-subset retrains. Seven A100 invocations, $0.34 metered. No git history. |
| Curvature-subsampling study | results/subsample-analysis/analysis.md |
Same solvers against a 2,048-example GGN; measures what curvature subsampling costs |
| Method reviews | astra_method_review.md, ciresan_reuse_review.md, PROTOCOL.md |
Operator definitions, sign conventions, validation ladder, frozen protocol and amendments |
| Background notes | ~/Google Drive/context/2026/week-37/09sep26-wed-flush/influence_functions_notes.md, astra_notes.md |
Grosse et al. 2023 and ASTRA digested; ten ways to spend compute on the Hessian; where the Wick machinery could fit |
| Tutorial pages (live) | https://yaroslavvb.github.io/influence-functions-tutorial/ (Kaczmarz Way), price-of-influence.html, response-curve.html; https://yaroslavvb.github.io/linear-influence/ |
Interactive 2-D intuition; the general derivation with a live cost calculator at Llama-3-8B scale; the full response curve; the earlier Kaczmarz linear-influence page |
| MNIST influence atlas | mnist-influence-explorer/ (git repo, 1 commit); https://mnist-influence-atlas.yaroslavvb.chatgpt.site |
Top-5 training digits for the 100 most ambiguous test digits, CNN at 0.39% test error; never validated against retraining |
| Mini-lab report (15 Sep) | ~/spacesheep/yaroslavvb/gyro-answered-2026-09-15/2026-09-15-influence-functions-minilab.html |
Laptop measurements on a converged 784–32–10 MLP: cross-covariance, Wick vs K-FAC, EK-FAC vs Kronecker sums, conditioning, LOO ground truth; size arithmetic; inverse cost table; tool survey |
| Wick–KFAC plan (9 Sep, superseded in part) | ~/Library/CloudStorage/Dropbox/git0/gyroscope/reports/2026-09/2026-09-09-grosse-influence-functions-wick-kfac.md |
The call summary, the closed-form inverse for the three-term factorization, the falsification test |
| Eschenhagen and Hong passes (17 Sep) | https://spacesheep.dev/@yaroslavvb/runa-eschenhagen-influence-functions, https://spacesheep.dev/@yaroslavvb/better-hessians-matter-paper-and-authors | Who is on Grosse's team; what "Better Hessians Matter" and its thesis version actually claim |
| Toy scripts | ~/.codex/attachments/b4345e40-7c7b-48e5-8e07-9ad065ef981f/wick_vs_kfac_check.py, astra_preconditioner_toy.py (copies now in scratch/) |
Sampled-label vs observed-label cross terms; preconditioner conditioning toy |
| Prior code | ~/Library/CloudStorage/Dropbox/git0/autotune/autotune/autograd_lib.py (SymmetricFourthOrderCov, approx='isserlis'), newton/wicks-factoring.nb, wicks-cumulants.nb |
The 2020–2022 implementation of the Gaussian factoring and its T-Sylvester/Woodbury solver |
What the work found, in order of importance:
- Better solves did not buy better counterfactual predictions at this scale. On the Ciresan network, ASTRA raised mean score-rank agreement with the converged full-data PCG solve from 0.380 (EK-FAC) to 0.751, but across 24 single-seed half-subset retrains the mean LDS was 0.014 (EK-FAC), −0.033 (ASTRA), 0.030 (PCG), every 95% interval including zero.
- The ground truth was noise. In the mini-lab, two leave-one-out passes over the same 120 points with different seeds agree at Spearman −0.007; the observed loss changes have 14× the spread of anything influence could predict. This is Basu et al.'s fragility result and Bae et al.'s warm-start gap reproduced in ten minutes. Hong et al. avoid it with ELSO (50% subsets, 100 subsets, 50 seeds each);
dattriships 5,000 pre-retrained MNIST models. - The Wick/Isserlis factorization is neither new nor large (§5): the cross-covariance it adds is the layer's mean gradient, zero for the GGN and 0.08% at convergence for the empirical Fisher.
- What survives, in sample, is a basis problem, not an eigenvalue problem (§2): on the mini-lab's converged 784–32–10 last layer, EK-FAC's relative error is 0.671 while the best single Kronecker product is 0.557 and a two-term Kronecker sum is 0.434, with 45% of the block's mass where any EK-FAC-type correction cannot reach; on the lab's real Ciresan last layer the same four numbers are 0.87, 0.64, 0.37 and 76% (§2.3). Hong's thesis reaches the same conclusion in words and asks for this fix. Out of sample the Kronecker-type gains shrink to about 5% at 25,000 calibration examples, and what transfers is the exact contribution of the few thousand highest-curvature examples (§2.3).
- The 9 Sep falsification test was a trap as written: MNIST's constant border pixels make the whitened cross-covariance look large (ρ₁ ≈ 0.55) at every stage of training.
- Numerics decide outcomes. The lab's first EK-FAC factorization in FP32 produced a basis with orthogonality error 1.22 (transposing it does not invert it); FP64 moments and eigendecompositions fixed it (3.7×10⁻⁸). Grosse on the call: "a lot of the things that go wrong are just … numerical issues".
2. Improving the Hessian: where the error actually is
2.1 Why the Hessian is the lever
The influence of training example z on a measurement f is −∇f(θ*)ᵀ H⁻¹ ∇L(z, θ*). In practice H is the damped Gauss–Newton/Fisher matrix G + λI (Bae et al. 2022: the exact derivative of the proximal Bregman response function, well defined for non-converged, non-convex models). Three readings of H⁻¹ recur across the notes and pages: a Newton step (force divided by stiffness), a whitening (a dot product after both gradients are rescaled by curvature, so generic high-norm directions are discounted), and explaining-away (a fact present in a thousand documents contributes a large eigenvalue that H⁻¹ divides down). Dropping H⁻¹ gives gradient similarity, which in the lab ranked 0.359 against the PCG solve versus 0.751 for ASTRA.
Grosse on the 11 Jun call: "better estimators actually translate into qualitatively better results … we have this metric called linear datamodeling scores … As you improve the Hessian estimates, those scores go up, and then you look at the results and they seem to make more sense." And the constraint: "it would have to be efficient for computing the IHVPs. But there is extra free time to be used up in the Hessian estimation phase."
2.2 The approximation ladder and what "Better Hessians Matter" says
Hong, Eschenhagen, Mlodozeniec and Turner (arXiv 2509.23437, v2 Feb 2026; the 12-page version of Hong's Cambridge MPhil thesis) compare, on tanh MLPs over the 8×8 digits data with ELSO ground truth, the ladder
exact Hessian ≳ GGN > block-diagonal GGN > EK-FAC > K-FAC,
stable across depth, width and epochs. Their decomposition of the Hessian→K-FAC gap (epoch 100, depth 8) is Hessian→GGN 2.59%, GGN→block-GGN 12.55%, block-GGN→EK-FAC 26.80%, EK-FAC→K-FAC 58.06%; the thesis summary across settings gives the merged Kronecker step 70–94% of the gap, with EK-FAC's eigenvalue correction recovering 40–65% of the total gap (roughly 45–69% of that step). Their eigenvalue-mismatch share falls with training, 60.2% → 58.1% → 41.0% at 10 → 100 → 1000 epochs, and they note that "gains narrowed near convergence".
Four details matter for what to do next:
- Metric. Their error is an inverse-application residual over training-gradient directions, (1/N) Σ ‖H Ĥ⁻¹v − v‖²/‖v‖², not a matrix distance. Inversion amplifies small-eigenvalue error, which is exactly what EK-FAC repairs, so the metric alone can move the eigenvalue-versus-basis verdict. Any sweep must report both.
- Basis. Every rung below block-GGN keeps the same Kronecker eigenbasis. Both the paper (Appendix C.4) and the thesis say the residual is "persistent off diagonal mass that no method constrained to be diagonal in the Kronecker factored eigenbasis can remove"; the thesis concludes that "much of the practical loss in influence quality stems from the basis differences and from spectral under-calibration rather than from linearisation alone" and closes (thesis only) by calling improvements "that correct within the eigenbasis or capture low-rank cross-layer terms" promising. Nothing in the ladder enriches the basis: the complete set of curvature objects tested is exact Hessian, block Hessian, GGN, block-GGN, EK-FAC, K-FAC.
- Damping. A thesis chapter absent from the paper finds per-block damping beats a single global value under block heterogeneity. This is the textbook explanation for the mini-lab's finding that K-FAC preconditioning made conditioning worse than nothing at λ = 10⁻³ (κ 137.8 vs 97.6) but fine at 10⁻² (12.8 vs 10.7).
- Share is not effect. Error share and LDS loss are not proportional: at 10 epochs the Hessian→GGN step is 3.5% of the error but costs 71.6 LDS points, and the block-GGN→EK-FAC step there raises LDS by 2.3 points. Which rung to improve is decided by LDS, not by the error decomposition.
Verified today against arXiv v1, v2 (15 Feb 2026) and the thesis PDF: the ordering, the 2.59/12.55/26.80/58.06 table, the 60.2/58.1/41.0 shares, the metric formula, the ELSO settings (α = 0.5, K = 100, R = 50), and the per-block-damping chapter are all as stated above. Corrections to the local reports: the paper's title says "Studying", the thesis's "Quantifying"; the "spotlight" designation does not appear in either document; the 58.06% figure is a paper number, while the thesis uses a three-step decomposition in which the merged block-GGN→K-FAC step is 84.84% at that setting and 70–94% across settings, with EK-FAC recovering 40–65% of the total gap; the paper's Table 1 lists widths {32, 64, 128} while its figure captions fix width 16; and the hessian_influence repository cited in the local reports returns 404, neither the paper nor the thesis carries a code-availability statement, and of the account's 23 public repositories the only related ones are forks of basic_influence and kronfluence. v2 exists because Felix Dangel caught an error in v1's claim that the GGN equals the Hessian for piecewise-linear networks; v2 corrects it to block-diagonal equality only.
2.3 Measurements: the converged last layer
Mini-lab, Tue 15 Sep, on the exact 330×330 last-layer GGN block of a converged 784–32–10 float64 MLP (N = 8,000 for statistics):
| Approximation of the block | Relative Frobenius error |
|---|---|
| K-FAC, E[H] ⊗ E[aaᵀ] | 0.7685 |
| Wick/Isserlis (K-FAC + cross terms) | 0.7685 (identical: C = 0) |
| EK-FAC (exact diagonal in the Kronecker eigenbasis) | 0.6709 |
| Best rank-1 Kronecker product (Van Loan–Pitsianis) | 0.5566 |
| Rank-2 / 3 / 5 / 10 Kronecker sums | 0.4342 / 0.3542 / 0.2403 / 0.1197 |
| Mass of ‖F‖² off the Kronecker-eigenbasis diagonal | 45% |
Conditioning and inverse quality on the same block: at λ = 10⁻³ the preconditioned condition number is 97.6 (none), 137.8 (K-FAC), 108.2 (EK-FAC); at 10⁻² it is 10.7 / 12.8 / 10.7. Mean cosine between approximate and exact iHVPs is 0.803 (K-FAC) and 0.829 (EK-FAC) at λ = 10⁻³, 0.926 / 0.941 at 10⁻². The best rank-1 Kronecker product has fewer free numbers than EK-FAC (d_in² + d_out² versus that plus d_in·d_out) and still beats it, which says the constraint that binds is K-FAC's choice of factors, not the eigenvalues placed on them.
Measured today on the lab's real checkpoint. scratch/weak_factoring_check.py and its follow-ups (listed in scratch/README.md) form the exact softmax-GGN block of the trained Ciresan MLP's last layer (500 inputs plus bias × 10 outputs, 5,010 parameters, a dense 5,010×5,010 float64 matrix) over 4,096 fitting examples (subset seed 20260917). The checkpoint reproduces the recorded 98.37% validation and 98.55% test accuracy; on the subset the training accuracy is 100% and the mean cross-entropy 0.00067. The block has Frobenius norm 0.018, trace 0.049, largest eigenvalue 0.0137, nullity at least 501 (softmax shift invariance, H_n·1 = 0), and a median positive eigenvalue of 10⁻¹⁰. At the lab's damping, λ = 0.01, the damped block has condition number 2.4, so on this layer the damping dominated the curvature and the solver comparison of §2.4 was insensitive to it. Curvature mass is extremely concentrated: the top 1% of examples carry 69% of Σ_n tr(H_n), the top 5% carry 92%, the top 20% carry 99.3%.
| Approximation of the real last-layer block | Relative Frobenius error |
|---|---|
| K-FAC, E[H] ⊗ E[aaᵀ] (exact factors) | 2.63 |
| K-FAC with one sampled label per example for S | 4.34 |
| Wick/Isserlis with sampled-label cross-covariance | 4.33 (the extra terms are 0.3% of the target) |
| Wick vs K-FAC against the empirical Fisher (observed labels) | 1.4590 vs 1.4593 (the extra terms are 0.07% of the target) |
| EK-FAC, exact diagonal in the Kronecker eigenbasis | 0.87; 76% of ‖F‖² lies off that diagonal |
| EK-FAC's exact diagonal taken in the optimal rank-1 Kronecker basis instead | 0.44; 20% of ‖F‖² lies off that diagonal |
| Same, in the curvature-weighted basis | 0.72; 52% off-diagonal |
| K-FAC with the activation factor weighted by tr(H_n), one Kronecker term | 0.80 |
| Mixture by predicted class, r = 10 | 2.38 |
| Mixture by predictive-entropy bins, r = 2 / 3 / 4 | 1.17 / 0.98 / 0.94 |
| Mixture by k-means on the probability vectors, k = 4 / 8 | 2.91 / 2.58 |
| Mixture by equal-curvature-mass bins, r = 2 / 4 / 8 / 16 | 1.44 / 0.96 / 0.52 / 0.38 |
| Top 5% of examples by curvature split by predicted class, plus the rest, r = 11 | 0.56 |
| Exact terms for the top-m examples by curvature plus K-FAC for the rest, m = 10 / 40 / 204 / 819 | 1.32 / 0.75 / 0.21 / 0.020 (correction ranks 90 / 360 / 1,828 / 3,603) |
| Same, with the curvature-weighted K-FAC for the rest | 0.30 / 0.19 / 0.066 / 0.011 |
| Optimal Kronecker sums (Van Loan–Pitsianis), r = 1 / 2 / 3 / 4 / 5 / 10 | 0.64 / 0.37 / 0.24 / 0.19 / 0.14 / 0.034 |
Inverse-side metrics on 200 per-example last-layer gradients (observed labels): condition number of F̂_λ⁻¹F_λ, mean cosine between approximate and exact damped iHVPs, and Hong et al.'s residual (1/N)Σ‖F_λ F̂_λ⁻¹g − g‖²/‖g‖². The unpreconditioned condition number is 2.37 at λ = 0.01 and 14.7 at λ = 0.001. (The two scripts drew different 200-gradient samples; the shared K-FAC row agrees to three digits on κ and cosine but differs by up to 9% on the Hong residual, 0.40 vs 0.43 at λ = 0.001, so residual differences below about 10% across rows are noise.)
| Method | κ, λ = 0.01 | cos | Hong residual | κ, λ = 0.001 | cos | Hong residual |
|---|---|---|---|---|---|---|
| None (F̂ = 0, iHVP = g/λ) | 2.37 | 0.979 | 0.309 | 14.7 | 0.837 | 30.9 |
| Diagonal of F | 2.41 | 0.980 | 0.296 | 16.9 | 0.845 | 24.6 |
| K-FAC | 5.11 | 0.966 | 0.062 | 70.6 | 0.838 | 0.40 |
| EK-FAC | 2.56 | 0.986 | 0.123 | 22.3 | 0.893 | 2.64 |
| K-FAC, activations weighted by tr(H_n) | 2.42 | 0.989 | 0.091 | 15.3 | 0.930 | 1.63 |
| Mixture, entropy bins r = 2 | 3.20 | 0.979 | 0.059 | 32.9 | 0.880 | 0.72 |
| Mixture, predicted class r = 10 | 3.28 | 0.973 | 0.158 | 22.4 | 0.807 | 0.42 |
| Mixture, curvature-mass bins r = 2 | 3.79 | 0.979 | 0.066 | 37.6 | 0.887 | 0.67 |
| Mixture, curvature-mass bins r = 8 | 1.78 | 0.996 | 0.013 | 8.36 | 0.946 | 0.27 |
| Exact top-40 examples + K-FAC / + weighted K-FAC | 1.83 / 1.23 | 0.995 / 0.999 | 0.0093 / 0.0057 | 8.92 / 2.81 | 0.950 / 0.983 | 0.065 / 0.151 |
| Exact top-204 examples + K-FAC / + weighted K-FAC | 1.24 / 1.07 | 0.999 / 0.9999 | 0.0011 / 0.0006 | 2.84 / 1.44 | 0.989 / 0.998 | 0.0114 / 0.0118 |
| Optimal Kronecker sum r = 1 | 2.14 | 0.992 | 0.100 | 19.6 | 0.889 | 7.62 |
| Optimal Kronecker sum r = 2 | 1.57 | 0.997 | 0.034 | 8.31 | 0.933 | 2.20 |
| K-FAC rescaled by its optimal scalar (0.14) | 2.73 | 0.982 | 0.181 | 28.8 | 0.869 | 6.51 |
| EK-FAC's diagonal correction in the optimal rank-1 Kronecker basis | 1.67 | 0.996 | 0.031 | 8.50 | 0.933 | 1.27 |
The held-out check, which bounds everything above. scratch/weak_factoring_heldout.py fits every estimator on the 4,096-example subset A and scores it against the exact block of a disjoint 4,096-example subset B. The two blocks themselves differ by 120% in relative Frobenius norm (‖F_A − F_B‖/‖F_B‖ = 1.20; the research agent's version with a 16,384-example B gives 0.96): at this level of convergence the block is dominated by a handful of examples, so a 4,096-example block is a noisy estimate of the population block. That distance is a scale marker, not a floor: it contains the sampling error of both blocks, and an estimator that denoises A can be closer to F_B than F_A is (at 25,000 examples the exact-head estimators are, 0.43 against an A–B distance of 0.47); the true floor is the B-side sampling error, which is not measured here. Out of sample, the ranking inverts. Against F_B the relative Frobenius errors are K-FAC 4.37, curvature-weighted K-FAC 0.95, EK-FAC 0.99, K-FAC times its optimal scalar 0.91 (agent's run), optimal Kronecker sums r = 1 / 2 1.35 / 1.37, curvature-mass mixture r = 8 1.28, exact top-40 / top-204 heads plus weighted K-FAC 1.16 / 1.17, and F_A itself 1.20. The inverse-side metrics agree: at λ = 0.01 the weighted K-FAC (κ 2.15, cosine 0.993) and EK-FAC (2.10, 0.992) are best and every richer estimator matches F_A itself (2.6, 0.987); at λ = 0.001 every estimator built on A conditions F_B worse than no preconditioner at all (7.3): top-204 head 15.7, mass-bins r = 8 16.5, top-40 head 17.4, EK-FAC 20.6, weighted K-FAC 23.8, K-FAC 40.6, Kronecker sums 34. So with 4,096 calibration examples the in-sample basis structure is mostly sample-specific, and the estimators that survive are the simplest ones.
At the realistic calibration size, 25,000 examples fit against the other 25,000 (scratch/weak_factoring_heldout25k.py), the A–B distance drops to 0.47 (consistent with 1/√N from 1.20 at 4,096) and the picture settles. Every Kronecker-type estimator lands within 0.79–0.87 of the held-out block: optimal r = 2 sum 0.79, EK-FAC in the optimal basis 0.80, optimal r = 1 0.81, weighted K-FAC 0.82, EK-FAC 0.84, rescaled K-FAC 0.87, and plain K-FAC 4.68. Basis enrichment therefore buys about 5% over EK-FAC on this layer, not the 2.4× seen in sample. The estimators that keep the highest-curvature examples exactly (the top 1% or 5% of A's examples plus a weighted K-FAC bulk) score 0.45 and 0.43, better than A's entire block (0.47), and on the inverse side at λ = 0.001 they are the only estimators that improve on no preconditioning (κ 3.0 and 2.8 against 5.7; cosine 0.974 and 0.980), while every Kronecker-type estimator makes conditioning worse (7.0–9.3; K-FAC 22.9). At λ = 0.01 nothing matters much (unpreconditioned 1.47; heads 1.34; Kronecker-type 1.53–1.69).
The reading: the block of this converged layer is dominated by its few thousand highest-curvature examples. A Kronecker model in any basis captures the bulk and saturates near 0.8; the head captures the rest and is what needs many examples to estimate. For the influence-function target, which is the block over the fixed training set rather than over a population, this argues for spending the estimation phase's free time on all 50,000 examples with a Kronecker bulk plus an exact head, not on richer Kronecker structure fitted to a 2,048-example calibration set (the lab's own subsample study, §2.4, is the same lesson from the solver side). One qualification on cost: the block's rank is at most 4,509, so a head saturates it once it holds about 500 examples (measured ranks 90 / 360 / 1,828 / 3,603 at m = 10 / 40 / 204 / 819), and the 5% head at 25,000 examples (m = 1,250) is a dense solve in disguise, while the 1% head (m = 250, rank 2,250) still scores 0.45. The Woodbury saving therefore needs m capped independently of the calibration size, and it exists only on wider layers, where a few-thousand-example head has rank in the tens of thousands against blocks of hundreds of thousands of parameters.
Six readings of these numbers.
- K-FAC's error is heteroscedasticity. K-FAC overestimates here: its trace is 5.5× the true trace and its optimal rescaling is 0.14 (a relative Frobenius error above 1 by itself would only say the two matrices are badly aligned). It spreads curvature over every example's activations while the true block gets its curvature from the few uncertain examples; E[H ⊗ aaᵀ] ≠ E[H] ⊗ E[aaᵀ] because H_n and a_n a_nᵀ are correlated through the input, and the correlation is not in the activation norm (corr(tr H_n, ‖a_n‖²) = −0.09) but in which activations the high-curvature examples use. Rescaling K-FAC by its optimal scalar (0.14) alone brings the conditioning to 2.73. One Kronecker term whose activation factor is weighted by each example's output curvature repairs most of the rest (2.63 → 0.80) at exactly K-FAC's cost, and beats EK-FAC in sample on Frobenius error and on all three inverse-side metrics at both dampings (κ 2.42 vs 2.56 and 15.3 vs 22.3; cosine 0.989 vs 0.986 and 0.930 vs 0.893; Hong residual 0.091 vs 0.123 and 1.63 vs 2.64).
- The basis problem is larger than in the mini-lab. 76% of the block's mass sits off the Kronecker-eigenbasis diagonal (45% on the small net), so EK-FAC's ceiling is 0.87 here. The optimal rank-2 Kronecker sum has 0.37 error and the best conditioning at both dampings of any fixed-structure method (κ 1.57 / 8.31). EK-FAC's diagonal taken in the optimal rank-1 basis reaches 0.44 in sample with only 20% of the mass off-diagonal, and at λ = 0.001 (κ 8.50, cosine 0.933) matches the two-term sum at EK-FAC cost. Only the exact-head constructions of reading 3 do better, and they do not survive the held-out check. Note the baseline rows: at λ = 0.01 plain K-FAC preconditioning is worse than no preconditioner at all on conditioning (5.11 vs 2.37) and cosine (0.966 vs 0.979), the real-checkpoint version of the mini-lab's κ result.
- A mixture is only as good as its latent. Clusters by class, entropy or k-means give 0.94–2.9; clusters by equal curvature mass give 0.52 at r = 8 and 0.38 at r = 16, with the lowest Hong residual of any fixed-structure method at λ = 0.01 at r = 8 (0.013) and conditioning (κ 1.78) just behind the optimal rank-2 sum (1.57) and EK-FAC in the optimal basis (1.67). The top bins are single examples, so the construction converges to exact terms for the few examples that carry the curvature plus one Kronecker term for the bulk: a Kronecker-plus-low-rank matrix (each example's term has rank at most 9) with a Woodbury inverse at K-FAC price plus O(9m·d), cheap only while 9m is far below the block dimension. Measured directly: exact terms for the top 10 / 40 / 204 examples (0.2% / 1% / 5% of the sample, carrying 45% / 69% / 92% of the curvature) plus a curvature-weighted K-FAC term for the rest give 0.30 / 0.19 / 0.066 relative Frobenius error with correction ranks 90 / 360 / 1,828; with plain K-FAC for the rest the same heads give 1.32 / 0.75 / 0.21, so the weighted bulk term and the exact head are both needed. At λ = 0.001 the r = 8 curvature-mass mixture has the best cosine (0.946) and Hong residual (0.27) of any fixed-structure method, with κ 8.4 tied with the optimal rank-2 sum (8.3) against 14.7 unpreconditioned, and the exact-head constructions go further: the top 5% of examples plus the weighted K-FAC bulk give κ 1.07 / cosine 0.9999 at λ = 0.01 and κ 1.44 / cosine 0.998 at λ = 0.001 against 14.7 unpreconditioned, with a Hong residual of 0.012, the best numbers in every column except that residual, where the plain-K-FAC bulk is marginally lower (0.0114). These are in-sample numbers; the held-out check above shows that at 4,096 examples they are memorization, while at 25,000 the exact head is the one construction that transfers (0.43, better than the 0.47 A–B distance) and the rank-r sums tie with EK-FAC.
- Estimation variance, not model class, is the binding constraint. The lab calibrated EK-FAC on 2,048 examples and the paper's recipe uses 2,048–10,000; on this converged layer a 4,096-example block is 120% away from another 4,096-example block, and even two 25,000-example halves differ by 47%. The levers that survive out of sample at small sizes are the ones that add no parameters (a scalar, a per-example weighting) or average over the sample (EK-FAC's diagonal); at large sizes the Kronecker-type models tie and the exact head wins. This is the notes' first lever ("more samples") measured, with a twist: the extra examples are worth most when their contribution is kept exactly rather than folded into a Kronecker average.
- Metrics disagree where it matters. At λ = 0.001, Frobenius ranks optimal-r=1 < EK-FAC < K-FAC, cosine ranks EK-FAC ≈ optimal-r=1 > K-FAC, and Hong's residual ranks K-FAC < EK-FAC < optimal-r=1. The residual punishes underestimated curvature (EK-FAC's diagonal is near zero over most of the 5,010 directions, with 124 entries negative at the 10⁻¹⁸ level, so its damped inverse acts as 1/λ there and over-corrects; truncated Kronecker sums underestimate in the discarded directions), while Frobenius punishes K-FAC's overestimate. Only LDS can arbitrate.
- The closed-form two-term inverse works only with factored damping, and then as a preconditioner. On this block the generalized eigenproblem is ill-posed as written: E[H] has rank 9 of 10 and the within-cluster activation moment is near-singular from dead ReLU units, so a small jitter is needed. With it the generalized eigenvalues are d_A ∈ [0, 1.8×10¹⁰] (one at −6×10⁻⁷ from the jitter) and d_S ∈ [2×10³, 7×10⁴], so 1 + d_A d_S is positive, as it must be for two PSD terms. The damping choice then decides everything: in the non-orthogonal basis V ⊗ U the damping enters as λWᵀW, whose off-diagonal part is 35% of its norm and whose diagonal exceeds the curvature term by a median factor of 7×10⁹. Measured at λ = 0.01 against the exact dense damped solve of the same two-term sum: diagonally approximated damping (the choice in the 9 Sep toy script) gives cosine 0.004, useless; exact damping through a dense 4,509-dimensional capacitance solve is exact but slower than a dense solve of the original block (2.2 s vs 1.1 s); factored Tikhonov damping folded into the first Kronecker pair before the generalized eigenproblems (Martens & Grosse's Appendix B route) gives cosine 0.99985 and relative difference 0.019 on the entropy-bin split. The approximation is split-dependent: on the curvature-mass split (13 / 4,083 examples) the same construction gives cosine 0.93–0.97 against the exactly damped mixture with 48–64% relative error, and cosine 0.88 (λ = 0.01) / 0.79 (λ = 0.001) against the true damped block, below the no-preconditioner baseline of 0.979 / 0.837 (the skeptic reviewer's run, kept as
scratch/factored_damping_check.py). So factored damping makes the two-term closed form usable as a preconditioner, where iteration removes its bias, not as a closed-form answer. Two more qualifications on cost: the 0.11 s apply time measured today is a dense apply of the materialized V ⊗ U, and the K-FAC-price claim in §3 rests on the structured apply (two Kronecker-factored matmuls, about 9×10⁶ FLOPs here against 8×10⁷ dense), which is analytically right but unbenchmarked.
2.4 ASTRA: solver against preconditioner
ASTRA (Wang, Nguyen, Yang, Bae, McIlraith, Grosse, arXiv 2507.14740) runs stochastic Neumann iterations preconditioned by the EK-FAC inverse and initialized at the EK-FAC answer; one step from zero reproduces EK-FAC exactly, and the fixed point is the true damped iHVP whatever the preconditioner. Truncating the iteration after J steps of size α is the same as raising the damping by 1/(αJ), which lands on the low-curvature directions that the paper's Section 6 shows carry much of the attribution signal.
Lab results on the Ciresan network (four illustrative queries, damping 0.01, GGN over all 50,000 fitting examples):
| Method | Mean relative vector error vs PCG | Mean score Spearman vs PCG | Mean solve time |
|---|---|---|---|
| Gradient similarity | — | 0.359 | — |
| EK-FAC | 45.71% | 0.380 | — |
| SNI (stochastic Neumann iteration), unpreconditioned, 200 steps | 28.93% | 0.704 | 3.27 s |
| ASTRA, paper step, 200 steps | 25.36% | 0.751 | 5.45 s |
| ASTRA, adapted step | 35.68% | 0.602 | 5.45 s |
| PCG, 26–28 iterations to 10⁻⁴ | — | 1.000 | 9.31 s |
Two cautions the lab itself records. ASTRA's final relative residuals were still 0.57–1.16, so 200 updates had not solved the system to PCG's tolerance; and solving a 2,048-example subsampled operator accurately recovered full-data rankings only at Spearman 0.62–0.78, so the curvature dataset matters as much as the solver. The 24-retrain LDS test could not distinguish any method (§1, item 1).
A third caution follows from the lab's own recorded step sizes (0.025 divided by a 20-step power estimate of the top eigenvalue; a power estimate is a Rayleigh quotient, hence a lower bound). The unpreconditioned damped operator's estimate was 0.095 at λ = 0.01; the GGN is singular (12M parameters, rank at most 450,000), so its smallest damped eigenvalue is exactly λ and its condition number is at least 9.5, and about 9.5 if twenty power steps converged. The EK-FAC-preconditioned operator's estimate was 9.1, and its smallest eigenvalue is at most 1 (for any u in the GGN's null space the preconditioned Rayleigh quotient is λ/(uᵀĜu/‖u‖² + λ) ≤ 1), so its condition number is at least 9.1. If the estimates are tight, EK-FAC preconditioning shrank the condition number by at most 4%: at this damping the system was already well conditioned and the preconditioner had almost nothing to do. That is consistent with plain SNI (0.704) nearly matching ASTRA (0.751), but it does not isolate preconditioning as the cause, because the arms ran at very different effective step sizes: ASTRA's paper step (0.1λ = 0.001) is 220× below the 2/λ_max ≈ 0.22 stability limit of the preconditioned operator, so its 200 halved steps sum to about 0.09 and could only partially correct the EK-FAC initialization, which by itself explains final residuals near 1. The solver comparison therefore says little about preconditioner quality. The regime where it matters is smaller damping (ASTRA used 0.0052 on MNIST and 10⁻⁴ for its low-curvature analysis), and there the lab's subsample study could not converge PCG in 240 iterations at λ = 0.001.
2.5 Where to spend Hessian compute, and what each costs per query
The 9 Sep notes list ten ways to spend the "free time" in the estimation phase. Below, condensed, plus two levers measured today (curvature-weighted K-FAC and EK-FAC in the optimal rank-1 Kronecker basis), with the one constraint that matters (does the inverse stay closed-form after a one-time factorization?):
| Lever | Keeps closed-form iHVP? | Comment |
|---|---|---|
| More samples for A, S and especially Λ | yes | Cheapest lever; variance in small eigenvalues is what the inverse amplifies; yields error bars for per-eigenvalue damping |
| Cover attention, embeddings, LayerNorm | yes | Eschenhagen et al. 2023 give K-FAC for weight-sharing layers (expand/reduce); Eschenhagen is on Grosse's team |
| Between-token correlations as a Kronecker sum Σ A_tt′ ⊗ S_tt′ | yes for two terms | Two SPD Kronecker terms diagonalize jointly via generalized eigenproblems; damping is no longer diagonal in that basis |
| Heteroscedastic mixtures Σ_c π_c A_c ⊗ S_c | yes for two clusters; iterative beyond; Woodbury when the extra terms are a few exact examples | Backprop covariance depends on the input; this is the conditional-Gaussian reading of the user's idea, and today's measurement says the clusters must be chosen by curvature mass (§2.3, §5.4) |
| Curvature-weighted K-FAC (activation factor weighted by each example's output curvature) | yes, at K-FAC cost | One Kronecker term; repairs most of K-FAC's overestimate on the converged last layer (§2.3) |
| EK-FAC's diagonal correction in the optimal rank-1 Kronecker basis (K-FOC/KPSVD factors by power iteration) | yes, at EK-FAC cost | Fixes the basis rather than the eigenvalues; measured today at conditioning 1.67 against EK-FAC's 2.56 in sample; out of sample it is worse than EK-FAC at 4,096 calibration examples (1.14 vs 0.93) and about 5% better at 25,000, so it is gated on calibration size (§2.3) |
| Low-rank Woodbury corrections, fit matrix-free | yes | The only cheap route to cross-layer curvature |
| Off-diagonal blocks in the Kronecker eigenbasis | yes (banded/blocked) | Captures some of the off-diagonal mass (45% on the mini-lab net, 76% on the real last layer) |
| Matrix-free fitting to exact GGN-vector products (Hutchinson probes) | n/a | Turns estimation into an arbitrarily scalable compute sink; sidesteps sample-covariance noise |
| Per-query CG refinement (ASTRA) | no | Spends per query; fine for a monitoring loop with few queries |
| Per-layer or per-eigenvalue damping | yes | Hong's thesis chapter; the mini-lab κ result |
| Warm starts across RL checkpoints | yes | Consecutive models are close; subspace iteration and incremental Λ |
2.6 What is new in 2025–2026
Every item below was fetched today (abstract page, and the PDF where a detail is quoted). The picture: the field agrees curvature quality matters for attribution, all shipped tooling stops at EK-FAC, and richer-than-Kronecker structure exists only in the optimization literature.
| Work | What it changes | Cost profile | Bearing on the basis question |
|---|---|---|---|
| Hong, Eschenhagen, Mlodozeniec, Turner, Better Hessians Matter (arXiv 2509.23437, v2 Feb 2026) | The ladder; ELSO ground truth; per-block damping (thesis) | Exact pseudo-inverses at toy scale | Diagnoses basis mismatch, tests no fix (§2.2) |
| Wang, Nguyen, Yang, Bae, McIlraith, Grosse, ASTRA (2507.14740, Jul 2025; NeurIPS 2025 per Semantic Scholar; arXiv v1 only, no public code) | EK-FAC as preconditioner inside stochastic Neumann iterations | A few hundred iterations per query | Fixes the residual by iterating, not by modelling it (§2.4) |
| Bae, Lin, Lorraine, Grosse, SOURCE (2405.12186, NeurIPS 2024) | Segment-wise damped iHVPs for non-converged, multi-stage training; damping 1/(ηK) derived | ~3 closed-form solves per query; in kronfluence | Its segment sum is itself a sum of Kronecker-factored terms |
| Mlodozeniec, Eschenhagen, Bae, Immer, Krueger, Turner (2410.13850, ICLR 2025) | K-FAC/GGN constructions tailored to the diffusion objective; a ready LDS harness | Closed form | Shows the curvature construction is the design choice |
| Mlodozeniec, Reid, Power, Krueger, Erdogdu, Turner, Grosse, Distributional TDA (2506.12965, NeurIPS 2025) | Influence functions as the limit of unrolled differentiation over training-run randomness, no convexity | Sampling over runs | Changes what the operator means at non-convergence |
| McKinney, Thudi, Bae, Rezaei, Papernot, McIlraith, Grosse, Gauss-Newton Unlearning (2602.10568, Feb 2026) | A few uphill K-FAC Gauss–Newton steps for unlearning at LLM scale | A few closed-form solves | Retain-set damage is bounded by curvature quality |
| Coalson, Bae, Carlini, Hong, IF-Guide (2506.01790, 2025) | Token-level influence for detoxification | A 7.5× smaller proxy model suffices for the scores | An orthogonal cost lever: shrink the model, not the curvature |
| Kreer, Wu, Adam, Furman, Hoogland, Bayesian Influence Functions (2509.26544, v2 Feb 2026) | Drops the Hessian entirely; loss-covariance statistics from SGMCMC | Sampling cost, claims billion-parameter scale | The counter-programme to better Hessians |
| Artemev, Xia, Boyd, Yu, Dangel, et al., Exploiting weight-space symmetries (2606.00442, May 2026) | Structured Hessians by averaging over loss-invariant group actions; the group is the accuracy/cost dial | Designed to be invertible | A 2026 analytic route to structure, names TDA as a target |
| Yusupov, Cherniuk, Frolov, Scalable Kronecker-Fisher Approximation (2609.02451, Sep 2026) | Kronecker Fisher with cross-layer blocks for LLM compression | Inverse cost not stated | Freshest evidence on the block-diagonal assumption |
| Lin, Lowe, Dangel, Eschenhagen, Xu, Grosse, KL-Shampoo/KL-SOAP (2509.03378, v10 Jun 2026) | Fit structured second moments by KL, not Frobenius | Unchanged structure | The Frobenius objective is the wrong one; relevant to metric choice |
| Lin, Dangel, Eschenhagen, et al., SINGD (2312.05705, ICML 2024) | Inverse-free K-FAC with sparse factor structures, stable in half precision | No decomposition at all | The opposite trade: less expressive factors, cheaper inverse |
| Dangel, Eschenhagen, et al., Position: curvature matrices via linear operators (2501.19183, ICML 2025); curvlinops (last push Jul 2026) | K-FAC, EK-FAC, GGN, Hessian operators; CG, LSMR, Neumann inverses | — | No Kronecker sums; a new operator would have to be written |
| Dangel, Mucsányi, Weber, Eschenhagen, KFAC From Scratch (2507.05127, Jul 2025) | Math and code side by side, with exactness test cases | — | The test suite any new factorization should pass |
| Eschenhagen, Immer, Turner, Schneider, Hennig (2311.00636, NeurIPS 2023) | K-FAC-expand and K-FAC-reduce for weight-sharing layers, exact for deep linear nets in their setting | Closed form | What K-FAC means for attention |
| Koroko et al., KPSVD (2201.10285, 2022) | Rank-r sums of Kronecker products by rearrangement SVD, deflation, Lanczos; inverse via Martens & Grosse's Appendix B | Same cost as K-FAC at r = 1; O(L d³) beyond | The structural half of the Kronecker-sum direction, for optimization on deep autoencoders |
| Koroko et al., two-level K-FAC (2303.18083, 2023) | Coarse-space cross-layer corrections | Small extra solve | Negative: cross-layer terms "do not really improve" K-FAC |
| Schnaus, Lee, Triebel, K-FOC (NeurIPS 2021 BDL workshop) | K-FAC recast as optimization over sums of Kronecker products; rank-1 optimum by power iteration | Same as K-FAC | Independent confirmation that the marginal factors are not the optimal ones |
| Tselepidis, Kohler, Orvieto, two-level K-FAC (2011.00573, 2020) | Low-rank coarse-space correction on top of block-diagonal K-FAC | One small dense solve | The Woodbury pattern, applied across layers |
| Tang et al., SKFAC (CVPR 2021) | Woodbury inversion of low-rank Kronecker factors | O(M³) instead of O(d³) | The numerics for any low-rank correction |
| Sweeney, FisherSketch (2606.27242, ICML 2026) | Exact three-factor identity for readout Fisher alignment; measures non-separability δ ≈ 0.34 relative Frobenius across six MLPs; sketches the joint fourth moment with factored random features | Sketch cost | Quantifies what the independence assumption loses at LLM heads |
| Lee, Li, Yin, Panda, KronQ (2607.07964, COLM 2026) | Post-training quantization with H ≈ H_X ⊗ H_G under the K-FAC independence assumption | — | The assumption is still being adopted for tractability |
| kronfluence (last commit 23 May 2026) | Strategies identity / diagonal / kfac / ekfac; Linear and Conv2d only | — | Nothing beyond EK-FAC is shipped; PyPI release is two years behind git |
Corrections to the local reports that came out of this sweep: Koroko et al. 2022 contains no statement about CNNs (all three experiments are fully-connected autoencoders), so the mini-lab's quoted "negative on CNNs" should not be repeated; the unlearning paper is McKinney et al., not Bae et al.; IF-Guide is Coalson et al.; the distributional-attribution paper has seven authors including Grosse; and the KL-Shampoo venue could not be confirmed from arXiv.
3. Computational complexity: the three phases and who pays
3.1 The phases
Everything in an influence pipeline falls into three buckets, and each Hessian improvement is judged by which bucket it moves cost into:
| Phase | Paid | Cost driver |
|---|---|---|
| Curvature estimation (fit A, S; eigendecompose; fit Λ) | once per checkpoint | ≈ 3 sequence-gradients per calibration sequence; memory a² + b² per a×b matrix (larger than the weights for MLP blocks) |
| Inverse-Hessian-vector product | once per query | EK-FAC: Σ 4ab(a+b) FLOPs, a few gradient-equivalents; ASTRA: J minibatch GGN products on top |
| Scoring / retrieval | once per (query batch, candidate) | one gradient per candidate; a full-corpus scan costs 6PD FLOPs = pretraining |
| Validation | once per method | hundreds of retrains; only affordable at fine-tuning scale |
The published "Price of Influence" page prices these for a Llama-3-8B-shaped model (32 blocks, width 4096, MLP width 14336, influence over MLP weights) with one unit, the sequence gradient C_g ≈ 6PT FLOPs. Its numbers are rendered by JavaScript; recomputed today from its formulas and default inputs:
Inputs, all the page's defaults: P = 8×10⁹, T = 2048 tokens per sequence, 32 blocks, d = 4096, h = 14336 (three d×h matrices per block, 96 in all), Q = 100 queries per batch, S = 10,000 calibration sequences, J = 300 ASTRA iterations at minibatch B = 32, N = 10⁷ candidate sequences, D = 15×10¹² corpus tokens, k = 4096 sketch dimension, M = 300 validation retrains of 10⁸ tokens each, one H100 at 989 TFLOP/s × 0.4 utilization = 3.96×10¹⁴ FLOP/s. Every H100-hour below is a pure-FLOP lower bound at that utilization. One sequence gradient is C_g = 6PT = 9.8×10¹³ FLOPs, 0.25 s; one fp32 gradient vector is 32 GB.
| Step | Formula | FLOPs | H100-hours | Paid |
|---|---|---|---|---|
| Query gradients, one batch | Q·C_g | 9.8×10¹⁵ | 0.007 | per query batch |
| EK-FAC statistics | 3·S·C_g | 2.9×10¹⁸ | 2.07 | per checkpoint |
| Eigendecompositions | 96 × 10(d³ + h³) | 2.9×10¹⁵ | 0.002 | per checkpoint |
| Factor memory + Λ, fp32 | 96(d² + h²)·4 B; 96·dh·4 B | 85 GB + 23 GB | — | resident (larger than the MLP weights) |
| iHVP, EK-FAC | 96 × 4dh(d + h) | 4.2×10¹⁴ | 0.0003 (1.1 s, ≈ 4 gradients) | per query |
| iHVP, ASTRA | J·(2B·C_g + solve) | 2.0×10¹⁸ | 1.41 (85 min) | per query |
| Scan of N candidates | N·C_g + N·Q·2P | 1.0×10²¹ | 702 | per query batch |
| Scan of the full corpus | 6PD | 7.2×10²³ | 5.1×10⁵ (58 H100-years) | equals pretraining |
| Raw gradient storage | N·P·2 B (bf16) | 160 PB | — | not storable |
| Sketch build | N·(C_g + 2Pk) | 1.6×10²¹ | 1,150 (1.6× one scan) | once |
| Sketch storage / lookup | N·k·2 B; 2Nk | 82 GB; 8×10¹⁰ | 0.2 ms | per query |
| Validation retrains | M·6P·D_ft/2 | 7.2×10²⁰ | 506 | per method |
Ratios that matter: ASTRA's iHVP is 4,841× EK-FAC's, J·(1 + 2B·C_g/C_solve) with C_g/C_solve = 0.24 at these defaults (at T = 512 the ratio would be 1,436). Amortized over the batch with a scan, one query costs 8.4 H100-hours, 83% scan and 17% ASTRA. Once retrieval is a sketch lookup, one query costs 1.41 H100-hours, of which ASTRA is more than 99.9%. The dot products against 100 query vectors are 1.6% of a scan, so a scan is simply "recompute every training gradient". Validating one method (506 hours) costs 72% of one scan: certifying a curvature improvement is roughly as expensive as using it once.
3.2 Families of curvature approximation and their inverse cost
Notation: an a×b weight matrix with the bias absorbed; D parameters in total; J iterations; B minibatch; k a rank or sketch dimension; r Kronecker terms. Evaluated numbers are for the 96 MLP matrices of the Llama-3-8B shape above; rows marked "estimate" are order-of-magnitude extrapolations, not measurements.
| Family | Setup, once per checkpoint | Per iHVP | Memory | Closed form after setup? |
|---|---|---|---|---|
| Dense exact | O(D³) | O(D²) | O(D²): 573 TB for the 12M-parameter lab net | yes, but only below D ≈ 5×10⁴ (§3.4) |
| K-FAC | eig(A), eig(B): O(a³ + b³) | 2 matmuls, 2ab(a + b): 0.5 s | a² + b²: 85 GB | yes |
| EK-FAC | K-FAC plus one pass for Λ: 0.69 h of the 2.07 h total fit | 4ab(a + b): 1.05 s ≈ 4 gradients | plus ab for Λ: 108 GB | yes; damping elementwise on Λ |
| Two-term Kronecker sum | two generalized eigenproblems per side, O(a³ + b³), 2–3× the constant (estimate) | 2 matmuls, K-FAC price (structured apply, unbenchmarked) | 2(a² + b²) | yes, with factored damping, which perturbs the operator by √λπ S₁⊗I + (√λ/π) I⊗A₁ |
| Kronecker sum, r > 2 | PSD-constrained ALS fit (unbuilt) | PCG preconditioned by the r = 1 term: ~100× EK-FAC at r = 10 (estimate) | r(a² + b²) | no |
| EK-FAC + rank-k Woodbury | O(k²D) plus matrix-free HVP probes | + 4kD: +32% at k = 4096, +0.4% at k = 100 | + kD: 65 TB at k = 4096 and D = 8×10⁹, so k ≲ 100 or per layer | yes |
| Wick / T-Sylvester (the user's) | eig(A) + eig(B) + SVD of the whitened cross-covariance: +11% | 2 matmuls + 1 transposed matmul + 2 dot products, 2ab(a+b) + 4ab·min(a,b): ≈ 0.7–1.3× EK-FAC depending on association order (estimate) | plus ab for C | yes |
| Mixture of Kronecker products by class, r = 10 | r K-FAC fits | PCG or nested Woodbury: 40–100× EK-FAC (estimate) | r(a² + b²) | no; closed form only at r = 2 |
| LiSSA / SNI | none | J HVPs with J ~ σ_max/λ, 10³–10⁶: 44 H100-h per query at J = 10⁴ | a few vectors | no |
| PCG | none | O(√κ log 1/ε) full-batch HVPs; does not tolerate stochastic products | a few vectors | no |
| ASTRA | EK-FAC plus a spectral estimate | J preconditioned stochastic HVPs: 1.41 h at J = 300, B = 32 | EK-FAC state + 4 vectors | no, but J is 10²–10³ instead of 10⁴–10⁶ (Koh & Liang's LiSSA budget was 5,000) |
| TRAK / LoGra projection | N(C_g + 2Pk) once: 1,150 h | O(Nk) lookup: 0.2 ms | N·k: 82 GB | retrieval, not inversion |
Only two non-K-FAC rows keep a closed-form inverse at K-FAC's per-iHVP order: the two-term Kronecker sum and the Wick/T-Sylvester inverse. Both push their extra cost into the estimation phase. The first does so with factored damping and, on the measurements of §2.3, as a preconditioner rather than an answer; the second has the right cost profile and the wrong target (§5).
3.3 How Hessian improvements move the bill
| Improvement | One-time phase | Per-iHVP phase | Scan | Meets Grosse's constraint? |
|---|---|---|---|---|
| More Monte-Carlo samples for A, S, Λ | linear in S (10⁴ → 10⁵ sequences: 2 → 21 H100-h) | unchanged | — | yes; the cheapest lever |
| Cover attention, embeddings, LayerNorm | ~2–3× | ~2–3× | larger gradients | yes, but it raises every phase |
| Two-term Kronecker sum (cross-token, or a two-component mixture) | 2–3× the eigendecomposition constant | unchanged | — | yes, exactly |
| Wick / T-Sylvester | +11% | ≈ 0.7–1.3× | — | yes on cost; no value on the GGN (§5) |
| Kronecker sum or mixture with r > 2 | r fits; ALS unbuilt | iterative, 40–100× | — | no, as stated |
| EK-FAC + rank-k Woodbury | O(k²D) plus probes | + 4kD (+32% at k = 4096, +0.4% at k = 100) | — | yes on FLOPs; memory-bound (k ≲ 100 or per layer) |
| Banded or blocked corrections in the Kronecker eigenbasis | modest | grows with block size | — | yes |
| Matrix-free fitting to exact GGN products | an unbounded sink | unchanged | — | yes; the cleanest realization |
| Per-block or per-eigenvalue damping | needs validation runs to set | unchanged | — | yes on the iHVP; validation is the cost |
| Warm starts across RL checkpoints | lower | unchanged | — | yes |
| ASTRA refinement | plus a spectral estimate (13.4 s in the lab) | 4,841× EK-FAC at the defaults | — | no; per-query spend, tolerable at small Q |
| Sketching for retrieval | 1.6× one scan, once | unchanged | 702 h → 0.2 ms per query | the enabler that turns preconditioner quality into the bill |
Before sketching, halving J saves 8% of a query; after sketching it saves 50%. The counter-pressure belongs next to this table: Bae et al. (2022, Table 5, MNIST vision models) decompose the gap between influence estimates and retraining into warm-start 4.4–20.7, proximity 0.001–14.3, non-convergence 0.1–11.4, linearization 0.000 and solver 0.001–0.307. Better curvature improves the two smallest columns. The cost analysis says where spend is cheap; Grosse's call and Hong et al. say it still pays on LDS; the lab's 24 retrains could not see it.
The qualitative story that survives the arithmetic: with EK-FAC the inverse is essentially free and the scan is the bill, which is why Grosse et al. 2023 prefiltered with TF-IDF and batched queries. Once retrieval becomes a sketch lookup (§4), the iHVP is the whole per-query bill, and the iteration count of an ASTRA-style solve, hence preconditioner quality, is what you pay. That is the cost story for a better closed-form preconditioner: it is paid once, in the phase Grosse said has free time, and it shrinks the phase that will dominate.
3.4 Measured numbers
Lab, A100-SXM4-40GB, 11.97M-parameter Ciresan MLP, 50,000-example GGN, λ = 0.01. Every value was re-derived today from the results JSON and matches the tutorial.
| Quantity | Value |
|---|---|
| Training, 24 epochs on 50,000 images | 5.84 s |
| EK-FAC fit on 2,048 examples, FP64 | 1.20 s |
| Spectral-radius estimate, 20 power steps (excluded from every published solve time) | 13.4 s, the largest solver-adjacent cost in the study |
| PCG, mean per query, 26–28 iterations to 10⁻⁴ | 9.31 s, 0.34 s per full-batch GGN product |
| ASTRA, 200 steps at batch 256 | 5.45 s, 27.3 ms per step |
| SNI, 200 steps at batch 256 | 3.27 s, 16.4 ms per step; so one EK-FAC solve costs 10.9 ms |
| One half-data retrain; all 24 | 2.59 s; 62 s |
| Seven runs, A100 wall clock; metered cost | 486 s; $0.335 ($4.77 reserved for the seven-run study, $7.40 in the ledger today after two later CNN runs) |
| Dense FP32 GGN; all 50,000 per-example gradients | 573 TB; 2.4 TB |
The solver seconds are not equal-accuracy comparisons: ASTRA's 200 steps ended at relative residuals 0.57–1.16 while PCG reached 7.6×10⁻⁵, and a PCG iteration is a 50,000-example product against ASTRA's 256-example minibatch.
Mini-lab, Intel i9 laptop, FP64 on CPU: eigh at d = 1,000 / 3,000 / 6,000 takes 0.07 / 0.86 / 11.4 s; a 6,370-dimensional exact Hessian by autograd 18.1 s; one 784–32–10 model on 5,000 examples 1.63 s, so 1,000 LDS retrains are about 27 minutes on the laptop. From the 9 Sep notes rather than the mini-lab: the Wick T-Sylvester inverse of a 100×100 weight runs in 65 ms in Mathematica, round-trip verified to 10⁻⁹. Dense-eigendecomposition ceilings, d ≤ √(memory / 3·bytes): an 80 GB A100 reaches 81,649 (fp32) / 57,735 (fp64); the 64 GB laptop 73,030 / 51,640; a 784–128–10 MLP (101,770 parameters) fits nowhere in fp64.
4. Sketching
4.1 Why retrieval is the bottleneck, and why naive projections lose
Scoring one query against the pretraining corpus costs one backward pass per sequence, which is the price of pretraining, per query. Grosse et al. 2023 escaped this two ways: TF-IDF prefiltering to ~10⁴ candidates (cheap, but biased toward token overlap, hiding the abstract generalization the method exists to find), and query batching (hold hundreds of rank-32 preconditioned query gradients and amortize training-gradient computation across them; ≥10⁷ sequences scanned per experiment). The obvious amortization, store every training gradient once and answer queries with dot products, needs N·P·4 bytes, petabytes for a modest candidate set; random projection to k dimensions (TRAK, LoGra) brings storage to N·k·2 bytes and a query to a k-dimensional lookup.
Two facts limit the projections. First, (G+λI)⁻¹ whitens the space: after preconditioning, most directions matter roughly equally, so a low-dimensional projection that keeps the top of the spectrum discards most of the signal (Grosse et al. 2023, Sec. 6). Second, ASTRA's Section 6 shows the low-curvature tail carries much of the attribution signal (projecting the iHVP onto high-curvature eigen-bins degrades LDS), and notes that methods which project gradients into lower-dimensional subspaces (it names LoGra and Arnoldi; TRAK's random projection is grouped with LoGra elsewhere in the paper) may discard those directions, as does an iterative solver that stops early. Grosse et al. 2023 never use the word "sketch"; their Section 6 discusses approximate nearest-neighbour search and argues against low-dimensional projections on exactly the whitening grounds above. Grosse's own retrieval work, which he said took most of his last year, is "a generic technique where we can compute sketches of the training data that let us retrieve the top influential examples with pretty strong guarantees on the recall"; it is unpublished, and nothing public today matches that description (five searches today found no paper from Grosse's group or Anthropic on sketches with recall guarantees for top-influence retrieval; the closest public results are score-level guarantees, §4.2).
The whitening argument was sharpened in February 2026 by Hu, Hu, Ma and Zhao, A Unified Theory of Random Projection for Influence Functions (arXiv 2602.10449). Without damping, a sketch preserves influence scores exactly only if it is injective on the range of F, which forces the sketch dimension to be at least rank(F), about the dataset size in overparameterized models, and below that no multiplicative approximation is possible. With damping λ, the requirement is governed instead by the effective dimension d_λ(F) = tr(F(F+λI)⁻¹): directions with curvature far below λ need not be preserved. So "projection loses signal" is right as a mechanism but the quantity that decides how much is the effective dimension at the damping scale, which couples sketch size to damping. It also reconciles the two Grosse-group results: ASTRA's low-curvature signal lives between λ and the top of the spectrum, and a small λ means a large d_λ, hence a large sketch. Their general guarantee (Theorem 2, sketch size m = Ω(d_λ(F)/ε²) with a matching lower bound) applies to any PSD curvature under an oblivious dense sketch, for query gradients in the range of F; their cheap factored sketch P_A ⊗ P_E is proved only for a single-Kronecker curvature F = A ⊗ E. The paper says EK-FAC's eigenvalue correction "fundamentally alters the spectral structure" and is not covered either, so every deployed pipeline has already left the factored-sketch guarantee behind. It is an argument for staying inside the Kronecker class if that guarantee is wanted, and a reason to check whether a two-term Kronecker sum inherits it; it is not an argument that EK-FAC is safe and richer structure is not.
4.2 Published sketching and projection methods, and what each trades
All items fetched today. The projection literature never runs MNIST (TRAK, LoGra, TrackStar, RISE and LoRIF all start at CIFAR-2 or above), so none of it bears directly on the lab; it bears on the cost model.
| Method | What is stored or sketched | One-time cost | Per query | What it trades |
|---|---|---|---|---|
| Grosse et al. 2023, TF-IDF filter + query batching (2308.03296) | Rank-32 preconditioned query gradients in memory; 10⁴ TF-IDF candidates | none | one scan per batch of queries | Recall of the prefilter is unquantified; gradients to disk were slower than recompute at their scale |
| Schioppa et al. 2022, Arnoldi (2112.03052, AAAI 2022) | Top-k Krylov eigen-subspace of H | m HVPs and O(mP) memory for the basis | O(Pk) projection + O(Nk) scan | Truncates exactly the low-curvature directions; OOMs at billion scale per LoGra |
| TRAK 2023 (2303.14186, ICML 2023) | JL projection to k = 1,000–4,000; inverts Φᵀ Φ in the sketched space; M models | M trainings + M·N projections, O(MNk) storage | O(k + Nk) | LDS is non-monotone in k (their Appendix E.1); no MNIST experiment |
| LoGra 2024 (2405.13954, NeurIPS 2024) | Two-sided Kronecker projection per layer, d₁×d₂; projected gradients to disk | On Llama-3-8B: 4,696 tok/s vs 1,740 tok/s for the EK-FAC fit (2.7×); 3.5 TB storage vs 89 GB | 79,004 pairs/s vs 12.2 (6,476×) at test batch 256; 131× at batch 4 | Throughput bought with disk and with projection loss; the "6,500×" is the per-query phase only |
| RapidIn 2024 (2405.11724, ACL 2024) | Gradients compressed 200,000× | one backprop per example + compression | scan of the cache; 6,326× speedup | Far below the effective dimension, so no approximation guarantee; works for top-k in practice |
| Schioppa 2024, Efficient Sketches for TDA (2402.03994, NeurIPS 2024) | Kronecker-structured sketches (AFFD, QK) with proofs | constant memory; 64× lower peak memory than Fastfood | — | Shows sketch cost on accelerators is memory-access-bound, not FLOP-bound |
| TrackStar 2025 (2410.17413, ICLR 2025) | Full scan of a 160B-token corpus for an 8B model with projection, optimizer-state correction, unit-normalized encodings | one projected-gradient pass over the corpus | O(ND) | BM25 still beats it at finding passages that state the fact |
| GraSS / FactGraSS 2025 (2505.18976, NeurIPS 2025) | Sparsified, sparsely projected gradients; sub-linear in P | — | up to 165% higher throughput than LoGra-class methods | Attacks the cost of touching the gradient at all |
| LoRIF 2026 (2601.21929) | Rank-c factorization of each projected gradient (O(c√D) storage) and a truncated-SVD Woodbury inverse (O(Dr)) | O(NDc + NDr) | 1.3–20× lower latency than LoGra; up to 20× less storage | For a fixed byte budget, raising the projection dimension beats raising the rank |
| RISE 2026 (2604.16197) | CountSketch of the readout layer only, two channels | index 112× smaller than RapidIn; scales to 32B | scan of a small index | Discards every layer but the last |
| Hu, Hu, Ma, Zhao 2026, Unified theory of random projection (2602.10449) | Theory: sketch size m ≥ rank(F) without damping; m = Θ(d_λ(F)/ε²) with damping; Kronecker sketches valid for Kronecker curvature | — | — | Score-level guarantees, not top-k recall |
| Sketchy 2023 (2302.03764, NeurIPS 2023) | Frequent-directions sketch of Kronecker factors | O(dk) memory | — | Guarantee with error in the bottom d − k eigenvalues, exactly where influence is sensitive |
| Nyström PCG (Frangella, Tropp, Udell, 2110.02820) | Randomized Nyström preconditioner from k HVPs | O(Pk² + k³) | CG converges in O(1) iterations once k ≈ d_λ(F) | The right use of a rank-k sketch is as a preconditioner for an exact solve, not a replacement |
| Klochkov & Liu 2024 (2409.17357) | LiSSA hyperparameters from the Hessian's trace and top eigenvalue; sketching as a convergence diagnostic | — | J HVPs | The counterpoint to ASTRA: tuned LiSSA may be competitive |
| DataInf 2024 (2310.00902, ICLR 2024) | Swaps inverse and sum; closed form | O(N) | none | Error grows with per-layer width; the LoRA regime only |
| In-Run Data Shapley (2406.11011, ICLR 2025 oral) | Accumulates dot-product terms during one training run | negligible over training | none, but queries must be fixed before training | No index, no inverse, no whitening; not an interactive tool |
| Sweeney, FisherSketch (2606.27242, ICML 2026) | Factored random-feature sketch of the joint fourth moment E[(a⊗e)(a⊗e)ᵀ] | sketch cost | — | Sketching on the estimation side; measures non-separability directly |
| Wang et al., Taming hyperparameter sensitivity (2505.24261, NeurIPS 2025) | Retraining-free selection of the damping | — | — | The hidden cost of influence is tuning, because evaluation needs retraining |
The takeaways. Sketching moves influence from a recompute-per-query problem to a one-time index plus an I/O-bound scan, and every published version pays in attribution quality in a way that is now partly understood (the effective dimension at the damping scale). No published method states a top-k recall guarantee for influence retrieval, with two qualifications: this is a "not found" from five web searches run while the arXiv API was unavailable, not a proof of absence; and the missing step is small, since Guo et al. 2019 (arXiv 1903.08690, Proposition 4) convert an inner-product error bound into a recall-at-k guarantee given a gap between the k-th and a later score, and Hu et al. supply exactly such an error bound at m = Θ(d_λ(F)/ε²). The closest public results are score-level bounds (Hu et al.), regret bounds with error in the bottom eigenvalues (Sketchy), and norm preservation (Schioppa 2024). The hole Grosse described is that nobody has composed the two for influence or measured the score gap that the composition needs.
4.3 Sketching on the estimation side
Sketching is not only a retrieval tool. For any structured family Ĝ(φ) with a cheap forward product, one can fit φ by minimizing Σ_k ‖G v_k − Ĝ(φ) v_k‖² over random probes v_k, where G v_k is an exact GGN-vector product (Hutchinson-style). This turns Hessian estimation into an arbitrarily scalable compute sink that never estimates a term from a sample covariance which is zero in expectation (the failure mode of the Wick cross terms), and it works for Kronecker sums, Woodbury corrections and banded eigenbasis corrections alike. The only requirement is that the family also supports a cheap inverse product. The astra_preconditioner_toy.py variant that adds a low-rank correction only when it lowers an HVP-probe residual reproduced EK-FAC exactly when the structure was absent and cut iterations 1.5–3× when it was present.
4.4 The memory-movement framing
The user's standing thesis (every scaling advance is a re-factorization to keep the working set close) applies cleanly: the iHVP is a memory-movement problem, not a FLOPs one. Whiten and precompute every training gradient once, sketch it down, and each query becomes a nearest-neighbour lookup; that is why TRAK and LoGra trade quality for throughput and why the sketching work exists. A richer closed-form preconditioner is the thesis applied: its extra structure is nearly free because it lands in the estimation phase, which is compute-bound and has spare time, while the retrieval phase, which is bandwidth-bound, is untouched. One paragraph in the post and one sentence in the email, not a second experiment.
5. The Gaussian ("weak") factoring idea: verdict
5.1 What the idea is, and what already exists
K-FAC approximates a layer's Fisher block F_(ki),(lj) = E[a_i a_j b_k b_l] (a = augmented input, b = pre-activation backprop) by assuming activations and backprops independent: F ≈ E[bbᵀ] ⊗ E[aaᵀ]. The weak version assumes only that (a, b) is jointly Gaussian; Isserlis' theorem then gives, exactly,
E[a_i a_j b_k b_l] = A_ij B_kl + C_ik C_jl + C_il C_jk (+ mean corrections), with C = E[abᵀ],
i.e. K-FAC plus two cross-covariance terms. This was told to Grosse on the call as "three-matrix KFAC … assume that they're jointly Gaussian … Wick factoring". It is not just a Stack Overflow answer: autograd_lib.py in autotune implements it (SymmetricFourthOrderCov, rank 3 = Isserlis, rank 2 = Kronecker, rank 4 = exact; the identity Mijkl = Mij·Mkl + Mik·Mjl + Mil·Mjk − 2·Mi·Mj·Mk·Ml), train_ciresan_new.py has a --curv {zero_order,kfac,isserlis,full} ablation, and wicks-factoring.nb builds the operator and its inverse (a T-Sylvester solve reduced to Sylvester and eigendecompositions, rank-one terms by Sherman–Morrison; round-trip 10⁻⁹, 65 ms for a 100×100 weight). The 9 Sep plan added a cleaner inverse: whiten by the K-FAC factors, SVD the whitened cross-covariance Ã⁻½CB̃⁻½ (its singular values ρ are canonical correlations between activations and backprops), in that basis the operator is 2×2-block-diagonal with eigenvalues 1 ± ρ_kρ_i, and the two rank-one terms are a 2×2 Woodbury correction. Setup: eig(A) + eig(B) + one SVD; per iHVP: K-FAC price plus one transposed matmul and two dot products. That is exactly the shape Grosse asked for.
On provenance: no public Stack Overflow, Math.SE, Cross Validated or MathOverflow post stating the three-term factoring was found today. The earliest public trace is a Cross Validated question by the user from 4 April 2017, "Covariance of Kronecker product?" (question 271905), which asks for Cov(a ⊗ b) in terms of A and B in order to whiten a ⊗ b; the accepted answer derives E[(a⊗b)(a⊗b)ᵀ] = E[aaᵀ] ⊗ E[bbᵀ] and flags that independence is used. If the idea is ever written up, the record starts there, not on Stack Overflow.
5.2 Why it adds nothing for the matrix the field uses
Three facts, each measured in the mini-lab, re-derived today, and re-measured on the lab's real checkpoint (on the real last layer the cross terms are 0.3% of the GGN target and 0.07% of the empirical-Fisher target; the mean last-layer gradient has norm 1.4×10⁻³ against a shrinkage fixed-point prediction of 4.7×10⁻³):
- The cross-covariance is the mean gradient. The per-example weight gradient is baᵀ, so E[abᵀ] is the transpose of the layer's average gradient. Measured: ‖Cᵀ − ∇_W L‖_F / ‖∇_W L‖_F = 2.7×10⁻¹⁶.
- For the GGN / true Fisher it is identically zero. With labels sampled from the model, E[b | x] = 0 by the score identity at every checkpoint, converged or not; hence E[b] = 0 and E[abᵀ] = 0. The Monte-Carlo estimate of ‖C‖ falls as 1/√S (1.63×10⁻² → 2.0×10⁻³ from S = 1 to 64 samples), the signature of noise. Grosse et al. 2023 and ASTRA both insist on the sampled-label Fisher and reject the empirical Fisher.
- For the empirical Fisher it vanishes at convergence. The mean gradient is what training drives to zero (to −wd·W under weight decay): ‖C‖ on layer 1 fell from 0.424 at initialization to 0.0025 converged (170×). Wick versus K-FAC on the exact empirical Fisher of the converged net: relative error 0.8720 → 0.8713, a 0.08% change.
And the fact that closes the argument: if (a, b) are jointly Gaussian with zero centered cross-covariance, Cov(a, b) = E[abᵀ] − E[a]E[b]ᵀ = 0, they are independent for any means, so the fourth moment factorizes exactly as E[a_i a_j] E[b_k b_l]; for the sampled-label Fisher E[b] = 0, so this is the same condition as C = E[abᵀ] = 0. Under the assumption the idea adds, K-FAC is already exact.
What Martens & Grosse 2015 (arXiv 1503.05671, v7 re-read today) actually do in Appendix A ("Derivation of the expression for the approximation from Section 3.1"): they expand E[ā⁽¹⁾ā⁽²⁾g⁽¹⁾g⁽²⁾] by the general moment–cumulant formula into the 15 set-partition terms, then use their Lemma 4 (any forward-pass quantity is uncorrelated with any backward-pass derivative when labels are sampled from the model's predictive distribution; the proof is the score identity) to eliminate ten of them. What remains is the K-FAC term plus three terms that are pure third- and fourth-order cumulants, κ(ā⁽¹⁾,ā⁽²⁾,g⁽¹⁾,g⁽²⁾) + κ(ā⁽¹⁾)κ(ā⁽²⁾,g⁽¹⁾,g⁽²⁾) + κ(ā⁽²⁾)κ(ā⁽¹⁾,g⁽¹⁾,g⁽²⁾). They do not use the words "Isserlis" or "Wick" (zero occurrences), and they do not claim exactness under Gaussianity; they say the error bound "will be small insofar as the joint distribution … is well approximated by a multivariate Gaussian" while noting the block-level Kronecker approximation "likely won't become exact under any realistic set of assumptions". Exactness of the per-pair identity under joint Gaussianity plus Lemma 4 is an immediate inference from their expression (every remaining term is a cumulant of order three or four), not their stated claim; the report attributes it that way. Their footnote that with training labels "Lemma 4 would break down" marks precisely the empirical-Fisher regime where the cross terms are non-zero (§5.3). Their one quantitative check is on MNIST: for their Figure 2 network the total approximation error over the middle four layers is 2894.4 against a cumulant-based upper bound of 4134.6, their evidence that higher cumulants predict K-FAC's error.
One more trap, so the negative result is not mis-measured: plotting singular values of A⁻½CB⁻½ on MNIST's first layer gives ρ₁ ≈ 0.55 at every stage of training for both Fishers, an artifact of whitening constant border pixels (A has eigenvalues at −6.6×10⁻¹⁸; 90% of the variance in 52 of 784 directions). Truncating A to its top-k eigenvectors gives ρ₁ = 0.13 (k = 10) rising to 0.23 (k = 200), growing with k, which is noise.
5.3 Where the cross terms are not zero
- Empirical-Fisher curvature at non-converged checkpoints (the mean gradient is large early in training). But the field's target is the GGN, SOURCE's segment Hessians are GGNs at averaged segment weights, and the empirical Fisher is rejected in both Grosse papers.
- Objectives that are not matching losses (RL fine-tuning, where the "backprop" is advantage-weighted), if a pipeline used empirical curvature there. Worth a question, not a project.
- Both cross terms are low-rank structures (vec(C)vec(C)ᵀ is rank one; the twisted term has rank ≤ rank(C)²), so if they were ever needed the Woodbury route absorbs them at O(r·d) per iHVP. The machinery is cheap; the term is empty.
5.4 The rescue: Gaussian conditionally, not globally
K-FAC's real error on the GGN is that E[H ⊗ aaᵀ] ≠ E[H] ⊗ E[aaᵀ]: the softmax output Hessian H_n = diag(p_n) − p_n p_nᵀ depends on the input, so the conditional second moment of the backprops moves with the activations. In cumulant language the error is third- and fourth-order (μ_i κ₃(ã_k, b_j, b_l) and κ₄ terms, Martens & Grosse's Eq. 3); a joint-Gaussian assumption has no cumulants above second order and therefore cannot represent it. That is why the fix is a sum of Kronecker products, not a Wick term. ASTRA's Appendix B.5 names this same assumption ("it treats the activations as independent from the pre-activation pseudo-gradients") as one of the two K-FAC errors its iterations correct, and defines EK-FAC's Λ as the exact diagonal of the true second moment in the Kronecker eigenbasis (its Eq. 28); so any correction that only moves diagonal mass in that basis is already subsumed by EK-FAC, and the value of anything new has to come from off-diagonal structure.
The Gaussian idea does survive in a weaker form, and it is the form the data favors. Assume (a, b) jointly Gaussian given a discrete latent c (a mixture of Gaussians, with c a function of the input: predicted class, a predictive-entropy bin, a token position or type). Within each component the cross-covariance still vanishes for the GGN (the score identity holds conditionally on anything computed from x), so under the conditional-Gaussian assumption each component factorizes and the block is
F = Σ_c π_c E_c[H] ⊗ E_c[aaᵀ],
a Kronecker sum with r = number of components. The assumption does the work: within a cluster E_c[H ⊗ aaᵀ] equals E_c[H] ⊗ E_c[aaᵀ] only if H_n and a_n a_nᵀ are uncorrelated there, and the measurements below show how far from true that is for every latent tried. This is precisely the family the mini-lab measured (rank-r sums beating EK-FAC from r = 1), the family Hong's thesis calls promising, and the family Koroko et al. (2022, KPSVD, arXiv 2201.10285) fit by rearrangement SVD, deflation and Lanczos for optimization on deep autoencoders, finding higher rank better and inverting the two-term case with Martens & Grosse's own Appendix B algorithm (whiten by the first pair, eigendecompose, divide). (The mini-lab's quoted "no gain on CNNs" sentence is not in that paper; its experiments are fully connected. The same group's 2023 two-level paper, arXiv 2303.18083, is the negative result, and it is about cross-layer corrections, not within-block rank.) The nearest precedent to "assume Gaussian, keep the cross terms, invert in closed form" is Martens, Ba and Johnson's ICLR 2018 K-FAC for recurrent networks, which models the gradient contributions across time steps with a chain-structured linear-Gaussian graphical model, sums the cross-moments, and inverts in closed form; the Gaussian assumption is applied along time rather than between activations and backprops. Tensor Normal Training (Ren & Goldfarb 2021) goes the other way and assumes a Gaussian gradient with a separable Kronecker covariance. Sweeney's FisherSketch (ICML 2026) measures how much the independence assumption loses, a relative Frobenius non-separability averaging 0.34 across six MLP architectures (its Table 12), shows on ViT-B/16 that the separable proxy mis-ranks transfer relative to the exact readout Fisher, and estimates the joint fourth moment by factored random features instead of by cross-covariance factors. The user's contribution is then not a new factorization but (i) a principled, data-conditioned way to choose the terms, and (ii) the inverse: two SPD Kronecker terms diagonalize jointly by two generalized eigenproblems, giving K-FAC-price iHVPs after an O(d_in³ + d_out³) setup; beyond two terms the T-Sylvester/Woodbury toolkit and an EK-FAC-preconditioned PCG are the primitives.
Measured today (§2.3 has the full tables). The construction works, but only with the right latent, and the latent is not class or confidence: it is curvature mass. On the real last layer, mixtures by predicted class, entropy bins or k-means on the probability vectors reach only 0.94–2.9 relative Frobenius error, worse than EK-FAC's 0.87 (with the mixture's own weights, which inherit K-FAC's scale overestimate; with one least-squares scalar per component the class mixture reaches 0.50 and a class-by-entropy mixture 0.41 against 0.34 for curvature-mass bins at r = 8, so curvature mass wins by about 1.5× at matched treatment rather than by the raw factor, and a scalar per component is mandatory for any latent, although least-squares coefficients can go negative and are not a deployable estimator); mixtures by equal-curvature-mass bins reach 0.52 at r = 8 and 0.38 at r = 16, against 0.37 for the SVD-optimal two-term sum, and at r = 8 give the lowest Hong residual of any fixed-structure method at the lab's damping (0.013) with conditioning (κ 1.78) just behind the optimal rank-2 sum (1.57); the exact-head construction below is better still in sample (κ 1.07, residual 0.0006), and at 25,000 calibration examples it is the one construction that survives the held-out check of §2.3. The reason is that 69% of this converged layer's curvature comes from 1% of the examples, so the natural components are a few individual high-curvature examples plus one term for everything else; measured directly, exact terms for the top 40 examples plus a curvature-weighted K-FAC for the rest reach 0.19, and for the top 204, 0.066, with correction ranks 360 and 1,828, so the inverse is K-FAC plus a small Woodbury solve. The cheapest member of the family, a single Kronecker product whose activation factor weights each example by its output curvature, already beats EK-FAC at K-FAC cost (0.80 vs 0.87). So the Gaussian assumption is useful as the derivation of a Kronecker-sum family and of which conditioning variable to use; it is not useful as a way to fit the terms, where an SVD or a matrix-free fit to exact GGN products does better than any hand-chosen latent. The user's inverse toolkit is the right primitive for the family, with the caveats below.
Four caveats to carry in. Every in-sample gain above must be re-earned out of sample. Against the block of a disjoint subset, at 4,096 examples only the reweighted single term and EK-FAC hold up; at 25,000 the Kronecker-type models tie within 5% of each other (0.79–0.87) and the exact-head estimators are the only ones that beat them (0.43) or improve conditioning at λ = 0.001 (§2.3). So the mixture family's practical value is its limit, an exact head plus a Kronecker bulk, and its cost advantage exists only on layers where the head's rank (9 per example) is far below the block dimension. Truncated-SVD Kronecker terms are not individually PSD, so a PSD-constrained fit is required for r > 1 (unbuilt); mixture components and exact-example terms are PSD by construction. λI is not diagonal in the non-orthogonal joint eigenbasis (V ⊗ U): approximating it diagonally gives a useless inverse (cosine 0.004), while factored Tikhonov damping folded into the first pair gives cosine 0.9999 on the entropy-bin split but only 0.93–0.97 on the curvature-mass split (§2.3, point 5). Factored damping changes the operator being inverted by an amount that depends on the split; that is routine for a preconditioner, where ASTRA makes bias irrelevant, and not acceptable for a closed-form answer without saying so. And everything measured today is one layer of one converged MLP at two damping values, judged by matrix and inverse-side metrics, not by LDS.
5.5 What to run, and how to frame it
The mini-lab's protocol, amended by the 17 Sep passes, is about two and a half hours on the laptop:
- Commit the lab (
git init), including the two toy scripts andscratch/. - Structure pass: on the saved Ciresan checkpoint, per layer where the block fits (last layer always) and per epoch, report K-FAC / EK-FAC / curvature-weighted K-FAC / EK-FAC in the optimal Kronecker basis / rank-r Kronecker sums / curvature-mass mixtures / exact-top-examples-plus-K-FAC errors in both metrics (Frobenius and Hong's inverse-application residual), the eigenvalue-share versus basis-share of the gap, off-eigenbasis mass, κ(preconditioned) with per-block damping, cos(iHVP), and Spearman of scores against the exact inverse.
scratch/weak_factoring_check.pyand its two follow-ups already do this for the last layer at the final checkpoint; the extension is per layer (the 500×1000 block is 501,000-dimensional and needs matrix-free products) and per epoch (the lab saved only the final checkpoint, so earlier ones must be retrained). - Ground truth: ELSO with α = 0.5 and many seeds, or
dattri's pre-retrained MNIST models; report LOO only alongside its seed-to-seed ceiling. Do not repeat the 24-subset single-seed arm. - Mislabel retrieval as the cheap convincing demo (rank by self-influence; report AUC and recall at 20% inspected).
- LDS only if step 2 shows a gap.
Framing for Grosse and Eschenhagen: not "three-matrix K-FAC" (the first question back is the score identity), and not "your 60/58/41 is the early end of a curve" (the thesis already agrees), but "your thesis says corrections beyond the Kronecker eigenbasis are promising; here is the first measurement of one, with a closed-form inverse for the two-term case and a cost table that lands the extra work in the estimation phase." Lead with κ and cos(iHVP), not Frobenius error (George et al.'s footnote: a better Frobenius fit does not guarantee a better natural-gradient direction).
6. Open items and risks
- The lab has no git history. 2.8 GB on one laptop, ~250 KB of real source, flagged daily since 15 Sep. Ten minutes; highest-value action in the thread.
- Ground truth. Cold-start retraining at MNIST scale does not reproduce; every LDS claim needs ELSO-style multi-seed subsets or pre-retrained models (
dattri), and warm-start or PBRF (proximal Bregman response function, §2.1) targets where retraining is unaffordable. - In-sample versus out-of-sample. The rich structure measured on a 4,096-example block is largely sample-specific: two disjoint 4,096-example blocks differ by 120%, two 25,000-example halves by 47%, and against a held-out block the Kronecker-type gains shrink to about 5% while only the exact-head estimator holds. Every claim about a richer estimator needs a held-out or full-training-set test at a realistic calibration size, and the mini-lab's basis numbers (also in-sample) carry the same caveat.
- Metric mismatch. Frobenius error and inverse-application residual can rank the ladder differently; report both, and lead with the inverse-side quantities.
- PSD and damping for Kronecker sums. Rank-r SVD terms are not PSD; mixture terms are. Damping is not diagonal in the joint generalized eigenbasis; factored damping is the honest fix.
- Scale transfer. Every measurement so far is on MLPs (the lab, the mini-lab, Hong et al.). Convolutions (where EK-FAC is weakest and ASTRA gains most) and attention (where K-FAC's expand/reduce choice is unsettled) are untested.
- Publishing window. Grosse said Hessian details are "a little bit ambiguous" to publish from inside; anything to be published goes out before joining.
- Tooling.
kronfluenceon PyPI is two years behind its git; ASTRA has no public code;dattriandcurvlinopsare the maintained layers to build on.
Appendix A. Artifact map
Paths not already listed in §1.
- Gyroscope reports (source markdown and HTML):
~/Library/CloudStorage/Dropbox/git0/gyroscope/reports/2026-09/2026-09-09-grosse-influence-functions-wick-kfac.md,2026-09-09-influence-functions-author-histories.md,2026-09-09-grosse-prior-engagement-and-astra-authors.md,2026-09-17-directions.md; the restart, Eschenhagen and "Better Hessians Matter" passes are the spacesheep spaces linked in §1. - The 11 Jun 2026 call transcript:
~/drive/context/2026/week-24/11jun26-thu-flush/(spc, roger grosse) roger grosse 11jun26.txt(29 minutes; the quotes in §2.1 and §4.1 are from it). - Background notes:
~/Google Drive/context/2026/week-37/09sep26-wed-flush/influence_functions_notes.mdandastra_notes.md(copies in~/.codex/attachments/303a8877-…/next to the two paper PDFs). - Prior code:
~/Library/CloudStorage/Dropbox/git0/autotune/autotune/autograd_lib.py(classSymmetricFourthOrderCovat line 967; the Isserlis identity at line 1450;approx='isserlis'branches at 1463 and 1490, the second marked as running out of memory);autograd_lib_test.py:991–1014(the test asserts both the Isserlis and the K-FAC norms within 10⁻⁴ of exact on a tiny net, so it does not discriminate between them);upstream/gradient-dissent/experiments/ciresan_stochastic_depth/reference/train_ciresan_new.py:110(--curv {zero_order,kfac,isserlis,full});~/Library/CloudStorage/Dropbox/git0/newton/wicks-factoring.nb(sections: Sylvester and T-Sylvester solver, Woodbury, forward/backward checks, timing),wicks-cumulants.nb,wicks-forum.nb. - New today:
scratch/weak_factoring_check.py(the real-checkpoint experiment),check2.py(least-squares refits, alternative bases, damping options),check3.py(why K-FAC overestimates; two cheap fixes),check4.py(curvature-mass bins),check5.py(held-out block of 16,384 examples),followup.py(curvature-mass clusters, weighted K-FAC),followup2.py(exact heads),heldout.pyandheldout25k.py(fit on one subset, score on a disjoint one), each with its.out;part1_checks.py(synthetic checks of the analytic claims); copies of the two toy scripts with their outputs; andscratch/data/(the MNIST download, 10 MB). - Lab data:
results/checkpoints/base.pt(model, split and query indices),ekfac-fp64-v2.pt(validated factors),results/*.json(every run's traces),results/wandb_manifest.json.
Appendix B. Claims verified against primary sources today
Three agents re-read the primary sources today (Martens & Grosse 2015 v7 from arXiv; Hong et al. v1, v2 and the thesis PDF; the local PDFs of ASTRA v1 and Grosse et al. 2023) and checked every load-bearing claim in the local reports.
| Claim in the local reports | Verdict | Detail |
|---|---|---|
| K-FAC's assumption is stated as independence between products of activities and products of derivatives (M&G §3.1, Eq. 2) | confirmed | quoted verbatim |
| M&G Appendix A expands the fourth moment into 15 moment–cumulant terms and Lemma 4 kills 10 | confirmed | Lemma 4 needs labels sampled from the model; the proof is the score identity |
| M&G say K-FAC is exact under joint Gaussianity | partially | They bound the error by third- and fourth-order cumulants and say it "will be small insofar as" the joint distribution is near Gaussian; exactness of the per-pair identity is an inference from their Eq. 3, and they say the block-level approximation "likely won't become exact under any realistic set of assumptions" |
| "Isserlis" or "Wick" appear in M&G 2015 | refuted | zero occurrences |
| M&G discuss the empirical Fisher and the mean-gradient term | partially | Empirical Fisher rejected in four places, with "Lemma 4 would break down" under training labels; no mean-gradient rank-one term is discussed |
| Hong et al. ordering Hessian ≳ GGN > block-GGN > EK-FAC > K-FAC; the 2.59/12.55/26.80/58.06 table; the 60.2/58.1/41.0 shares | confirmed | paper numbers (v1 = v2); the thesis uses a different three-step decomposition |
| Their metric is an inverse-application residual, not Frobenius | confirmed | Eq. 9, squared norms |
| ELSO with α = 0.5, K = 100, R = 50; LOO rejected | confirmed | Table 1; thesis §2.3.2 |
| Digits 1,617/179, tanh, lr 0.03, batch 32, ε = 10⁻⁴, λ = 0 | partially | width table vs figure captions disagree; the heterogeneity chapter uses ReLU |
| The "no method constrained to be diagonal…" and "low-rank cross-layer terms are promising" quotes | confirmed | the first is in both paper (App. C.4) and thesis; the second is thesis-only |
| Per-block damping chapter | confirmed | thesis §4.2 only; hidden layers [64, 64, 32], c ∈ {1.0, …, 2.0} |
| v2 thanks Felix Dangel for a mistake; code repository public | confirmed / refuted | v1's claim that GGN = Hessian for piecewise-linear nets was corrected to block-diagonal equality; the repository returns 404 |
| Hong et al. test any basis-enriching approximation | refuted | the ladder tops out at EK-FAC |
| ASTRA = EK-FAC-preconditioned stochastic Neumann iteration initialized at EK-FAC; one unit step from zero equals EK-FAC-IF | confirmed | Eq. 5, footnote 11, Appendix B.1 |
| ASTRA recipe | partially | 200 iterations (300 for GPT-2), batches 256/128/16, α = 0.1λ / 0.01λ / λ, momentum 0.9, damping from SOURCE's 1/(ηK); the GPT-2 schedule decays by 0.9 every 100 iterations, not by half every 50 |
| ResNet-9 LDS 0.25 → 0.6 ensembled; "MLP gains modest, GPT-2 smallest" | partially | the first is stated in the introduction; the paper has no numeric LDS table and does not assert the ordering (it singles out FashionMNIST's ensembling gain) |
| Truncation adds damping 1/(αJ) | confirmed | Section 6 and Appendix G, additive |
| Section 6: low-curvature directions carry signal; TRAK, LoGra, Arnoldi discard them | partially | the section names LoGra and Arnoldi; TRAK is grouped with LoGra elsewhere |
| Empirical Fisher rejected | confirmed | Appendix B.5, citing Kunstner et al. |
| LDS protocol: 100 half-subsets; 10,000 UCI retrains, 1,000 GPT-2 fine-tunes, 5,000 for MNIST; MNIST MLP 512/256/128 on 6,144 examples | confirmed | Table 2 |
| ASTRA's per-iteration cost and wall-clock | partially | one minibatch GGN product plus one EK-FAC solve, O(BD); the paper reports no wall-clock or FLOP numbers |
| ASTRA versions, venue, code | — | arXiv v1 only, footer "Preprint. Under review."; NeurIPS 2025 per Semantic Scholar metadata; no public code found |
| Grosse 2023 EK-FAC memory M/P + P/M + 1 per parameter; 52B block split | confirmed | Eq. 27; "5.25×" is the notes' arithmetic |
| Rank-32 query batching, TF-IDF top 10,000, ≥ 10M sequences scanned | confirmed | Sections 3.2.2, 3.2.1, 5.2; Figure 3 correlation 0.995 |
| λ_ℓ = 0.1 · mean(Λ_ℓ) | confirmed | Section 5 |
| EK-FAC ≈ LiSSA at an order of magnitude less time | partially | "orders of magnitude faster" (§5.1) vs "at least an order of magnitude" (§6); Figure 7 has no numbers |
| MLP weights only; tokens i.i.d. for factors, exact token sums for Λ | confirmed | Section 3.1, footnote 4 |
| Grosse 2023 discuss sketching | partially | the word never appears; Section 6 argues against low-dimensional projections on whitening grounds |
Appendix C. Recent literature, verified today
Sixty-odd items were fetched today across three sweeps; the ones that changed the report are in the tables of §2.6 and §4.2. Items consulted but not tabulated above, all verified by fetching the abstract or PDF:
- Martens & Grosse 2015 (1503.05671, v7): Appendix A, the cumulant expansion and Lemma 4; Appendix B, the closed-form inverse of A⊗B ± C⊗D by whitening and eigendecomposition.
- Grosse & Martens 2016, KFC (1602.01407): the conv-layer factorization derived from three named assumptions, IAD (independent activations and derivatives), SH and SUD.
- Botev, Ritter, Barber 2017, KFRA/KFLR (1706.03662): no distributional assumption; expected curvature propagated recursively.
- George et al. 2018, EK-FAC (1806.03884): the Frobenius-optimal diagonal in the Kronecker eigenbasis; footnote 3 on direction versus norm.
- Ren & Goldfarb 2021, Tensor Normal Training (2106.02925): Gaussian gradient with separable Kronecker covariance.
- Benzing 2022 (2201.12250): for optimization, K-FAC outperforms exact second-order updates, so "more faithful curvature" does not transfer from attribution to training.
- Voet 2023/2024 (2307.07884): preconditioners for generalized multiterm Sylvester equations, the r > 2 Kronecker-sum case.
- Bae, Ng, Lo, Ghassemi, Grosse 2022 (NeurIPS 2022): the PBRF and the error decomposition quoted in §3.3 (via the notes; not re-fetched today).
- Koh & Liang 2017; Cook 1977; Pruthi et al. 2020 (TracIn); Agarwal, Bullins, Hazan 2017 (LiSSA): background, cited in the tutorial pages.
- Stack Exchange: Cross Validated question 271905 (2017) by the user; MathOverflow answer 436131 applying Wick's theorem to a fourth moment. No post with the three-term K-FAC factoring exists.
- Tooling: kronfluence (pomonam, last commit 2026-05-23; PyPI v1.0.1 from July 2024), dattri (TRAIS-Lab, 0.3.0, March 2026; ships pre-retrained MNIST models), curvlinops (f-dangel, last push 2026-07-17), quanda, torch-influence (cold since 2022), logix (dead), TRAK (stale). ASTRA has no public code; Hong's
hessian_influencerepository returns 404.