Does correlation between names change how many dollars of loss a harvested portfolio produces? No. It changes only the spread.
Read the paper on SSRN ↗ CiteKey results
- Under a fixed share count, expected dollar harvest is flat across correlation from 0 to 0.95; the largest paired deviation across eight specifications is 0.34% of the level.
- Correlation inflates the harvest's standard deviation by at most sqrt(N); the measured ratio divided by sqrt(N) averages 1.005 and never leaves [0.997, 1.019].
- Raising single-name volatility lifts the mean harvest by +46.7%; cutting correlation with volatility fixed moves it by -0.03%.
- Because deductions are capped by available gains, correlation lowers expected usable losses about 21% at a realistic gain budget while the gross harvest changes -1.1%.
- On 2,395 days of US equity returns, dollars harvested per $1 differ by +0.0008 (paired t +0.69) while yield differs with a paired t of +5.37.
Summary
The question
Anyone sizing a tax-loss harvesting program wants one number first: how many dollars of loss will this portfolio realise? Published studies put the answer somewhere between half a percent and two percent of value a year. Where a given portfolio lands inside that range depends on its traits, and on one of them the guidance conflicts. Practitioner work tilts direct-indexing targets toward higher dispersion, on the logic that when names disagree there are more losers to sell. A widely read primer argues the other way, that lower correlation raises the value of harvesting, because names that move together eventually leave no losers once the market rises.
We came to this through a mistake of our own. An early cross-sectional model of harvest yield gave average correlation a positive coefficient, with a weight near 0.29. It was an artifact. Finding out why became the paper.
Why correlation cannot move the average
Hold N names with a share count that never changes. Whenever a lot trades below its basis, sell it, book the loss, and buy a replacement so the basis resets down. The loss on any one name depends on that name's own price history and on nothing else. Total loss is a sum of such terms, and the expectation of a sum is the sum of the expectations, so no covariance or factor structure enters anywhere. Hold each name's own behaviour fixed and the expected dollar harvest cannot depend on how the names move together. That is true at every horizon, not only for a new account.
Could something as messy as fat tails break it? We checked with a simulation that recombines identical random shocks at different correlations, so any difference is dependence and not noise. Mean dollar harvest stayed flat from 0 to 0.95 under Student-t returns with five degrees of freedom, under jumps with skewness of -0.86 and kurtosis of 7.0, in a bear market with drift of -4%, at 60% single-name volatility, with 400 names, with monthly monitoring and over twenty years. In Table 1 the largest deviation is 0.34% of the level. Figure 1 draws the dollar harvest per $1 from an 8,000-path run as a flat line.
What correlation does instead
Dependence sets the spread. When all names share the same marginal behaviour, the harvest is an average of N contributions, and its variance lies between Var(g)/N under independence and Var(g) when every name moves in lockstep. The standard deviation can grow by at most sqrt(N). Table 2 tests that ceiling for breadth from 10 to 400: the measured ratio divided by sqrt(N) averages 1.005 and never leaves [0.997, 1.019] (Figure 2).
Numbers make this concrete. Take a hundred-name book at 35% volatility. With independent names, the middle 80% of ten-year outcomes runs from 42% to 49% of initial value; with nearly correlated names it runs from 11% to 80%. The average is the same. Someone who quotes that average for a highly correlated book can miss by a factor of four.
Reconciling the two claims
Both camps had hold of something real. Losses offset gains dollar for dollar, but only up to the gains available, and the rest carries forward. Usable losses are therefore concave in the gross harvest, so a wider spread lowers the expected usable amount even when the gross mean is unchanged. In Table 4 gross harvest moves -1.1% as correlation rises from 0 to 0.95, while the deductible amount falls by about 21% at a realistic gain budget (Figure 4). The primer's sign is right for deductions. Its reason, a shortage of losers, is not.
Measurement supplies the second channel. Harvest yield divides dollars by the account's value over the same period, and that value falls just as losses are realised. So yield picks up a dependence the dollars lack: portfolios with a flat dollar harvest show a 95% rise in yield.
What about dispersion? Table 3 raises it two ways. Raising volatility lifts the mean harvest +46.7%, and cutting correlation at fixed volatility moves it -0.03%. Tilting toward dispersion works only when volatility is the thing that rises.
The real data and what remains to predict
A simulation is only as good as its return model. So we also ran a permutation test on 2,395 days of US equity returns (Figure 5), one that needs no model at all: shift each name's history by its own random offset and every name keeps its own behaviour while the links between names vanish, whereas a common shift keeps both. Table 5 compares the arms. Dollars per $1 differ by +0.0008, a paired t of +0.69. Harvest over average NAV differs by +0.0046, a paired t of +5.37. Flat dollars, rising yield, on real prices.
There is a large caveat. The result needs a fixed share count, and most live direct-indexing products rebalance to an index; under quarterly rebalancing the mean harvest falls 53.2% as correlation rises. For those products the theorem is a benchmark and no more.
What is left to forecast? If the conditional mean is pinned down by each name's own behaviour, the useful target becomes a floor, meaning a low quantile of the outcome. We estimated quantiles with quantile regression, gradient boosting and a deep network, then calibrated them by conformalized quantile regression. On 100 held-out configurations (Table 6), gradient boosting reached a pinball loss of 0.0315 and calibrated coverage of 0.907 for the nominal 90 percent interval.
Who this is for: Investors, CPAs and tax professionals, researchers and quants, and journalists who write about direct indexing and tax-loss harvesting.
Figures
Tables
| Specification | E[L] at ρ=0 | at ρ=0.9 | change | paired t | sd ratio |
|---|---|---|---|---|---|
| Baseline | 0.45599 | 0.45256 | -0.75% | -1.99 | 9.23 |
| Student-t(5) marginals | 0.45422 | 0.45509 | +0.19% | +0.50 | 9.17 |
| Jump marginals | 0.46180 | 0.46335 | +0.34% | +0.89 | 9.18 |
| Bear market, μ = -4% | 0.68705 | 0.68885 | +0.26% | +1.16 | 9.19 |
| Single-name volatility 60% | 0.74661 | 0.74837 | +0.24% | +1.05 | 9.18 |
| N = 400 names | 0.45579 | 0.45297 | -0.62% | -0.91 | 18.32 |
| Monthly harvesting | 0.44115 | 0.44229 | +0.26% | +0.64 | 9.18 |
| Twenty-year horizon | 0.53755 | 0.53656 | -0.19% | -0.29 | 9.23 |
| Quarterly rebalancing | 1.31685 | 0.61620 | -53.21% | -385.06 | 4.55 |
| Capacity limit, deepest 10% | 0.45434 | 0.44225 | -2.66% | -6.87 | 9.34 |
| N | sd, independent | sd, comonotone | ratio | √N | ratio /√N |
|---|---|---|---|---|---|
| 10 | 0.0831 | 0.2632 | 3.17 | 3.16 | 1.001 |
| 25 | 0.0521 | 0.2609 | 5.01 | 5.00 | 1.002 |
| 50 | 0.0373 | 0.2642 | 7.09 | 7.07 | 1.003 |
| 100 | 0.0264 | 0.2629 | 9.97 | 10.00 | 0.997 |
| 200 | 0.0184 | 0.2625 | 14.28 | 14.14 | 1.009 |
| 400 | 0.0129 | 0.2638 | 20.39 | 20.00 | 1.019 |
| Intervention | realized cross-sectional dispersion | mean harvest | change | t |
|---|---|---|---|---|
| Baseline | 0.221 | 0.3775 | ||
| Raise single-name volatility to 42%, ρ fixed | 0.309 | 0.5538 | +46.7% | +553.2 |
| Cut ρ to 0.05, single-name volatility fixed | 0.290 | 0.3774 | -0.03% | -0.1 |
| Quantity | Change |
|---|---|
| Gross harvest | -1.1% |
| Deductible, gains to offset =5% of NAV | -3.4% |
| Deductible, gains to offset =10% of NAV | -5.5% |
| Deductible, gains to offset =20% of NAV | -10.3% |
| Deductible, gains to offset =40% of NAV | -20.9% |
| Present value of deductions, gains =2% of NAV per year | -12.9% |
| Present value of deductions, gains =5% of NAV per year | -19.9% |
| Quantity | dependence kept | dependence broken | difference | paired t |
|---|---|---|---|---|
| Dollars harvested per $1 | 0.3005 | 0.2998 | +0.0008 | +0.69 |
| Harvest divided by average NAV | 0.1666 | 0.1619 | +0.0046 | +5.37 |
| Model | pinball | coverage, raw | coverage, calibrated | width, raw | width, calibrated | Winkler |
|---|---|---|---|---|---|---|
| Theory benchmark | 0.0532 | 0.963 | 0.914 | 1.226 | 1.103 | 1.247 |
| Linear quantile regression | 0.0317 | 0.875 | 0.888 | 0.431 | 0.442 | 0.589 |
| Gradient boosting | 0.0315 | 0.896 | 0.907 | 0.446 | 0.455 | 0.588 |
| Deep quantile network | 0.0307 | 0.883 | 0.899 | 0.417 | 0.431 | 0.562 |
| Model | ρ < 0.3 | 0.3 ≤ ρ ≤ 0.6 | ρ > 0.6 |
|---|---|---|---|
| Theory benchmark | 0.858 | 0.950 | 0.940 |
| Linear quantile regression | 0.872 | 0.897 | 0.896 |
| Gradient boosting | 0.922 | 0.911 | 0.876 |
| Deep quantile network | 0.894 | 0.905 | 0.899 |
| τ | quantile-regression coefficient on ρ | boosting importance of ρ |
|---|---|---|
| 0.05 | -0.288 | 0.078 |
| 0.10 | -0.240 | 0.057 |
| 0.25 | -0.123 | 0.028 |
| 0.50 | +0.020 | 0.013 |
| 0.75 | +0.127 | 0.049 |
| 0.90 | +0.194 | 0.056 |
| 0.95 | +0.220 | 0.058 |
Abstract
An advisor sizing a tax-managed mandate needs to know how many dollars of harvestable loss a portfolio will produce, and the literature disagrees about one input. Practitioner work recommends tilting a direct-indexing target toward higher return dispersion; a widely read primer reports that lower correlation raises the value of harvesting.
Under a lot-level harvesting rule with a fixed share count, the expected realized loss is invariant to the dependence structure of returns once marginals are held fixed. Each name's cost basis is a functional of that name's own price path, so the expected harvest is a sum of single-name marginal expectations into which no cross-moment enters. Across eight specifications spanning fat tails, jumps, negative drift, breadth and monitoring frequency, the largest paired deviation measured is 0.34% of the level.
Correlation instead governs the second moment, with a sqrt(N) ceiling on the inflation of the harvest's standard deviation that simulation attains to within 2%. That reconciles the disagreement in two steps. Because the deduction is capped by the gains available to absorb it, expected usable losses are concave in the gross harvest and fall about 21% at a realistic gain budget, recovering the reported sign from the variance channel rather than from a shortage of losers. And because harvest yield divides by a contemporaneous NAV, it inherits a correlation dependence that is an artifact of normalization: portfolios whose dollar harvest is flat show a 95% rise in yield. Raising single-name volatility lifts the harvest 47%; raising dispersion by an equal amount through lower correlation moves it 0.03%.
A permutation test on 2,395 days of US equity returns reproduces this with no return model. Since the conditional mean is fixed by the marginals, the predictable object is the conditional quantile, estimated by quantile regression, gradient boosting and a deep quantile network and calibrated by conformalized quantile regression.
Keywords: tax-loss harvesting, direct indexing, dependence invariance, comonotonicity, convex order, conformal prediction, quantile regression, harvest yield, after-tax return, separately managed accounts
How to cite
Majumdar, A. (2026). Dependence Invariance in Tax-Loss Harvesting: An Exact Result and the Distributional Cross-Section of Harvest Yield. SSRN Working Paper No. 7337678. https://ssrn.com/abstract=7337678
@techreport{majumdar_cross_section_tax_alpha,
author={Majumdar, Anirban},
title={Dependence Invariance in Tax-Loss Harvesting: An Exact Result and the Distributional Cross-Section of Harvest Yield},
institution={SSRN},
number={7337678},
year={2026},
url={https://ssrn.com/abstract=7337678}}