Working paper · SSRN 7337678 · · 20 pages

Does correlation between names change how many dollars of loss a harvested portfolio produces? No. It changes only the spread.

Read the paper on SSRN ↗ Cite
Markets & asset pricingTax-aware investingAdvisers & investorsResearchers & quantsJournalistsCPAs & tax professionals

Key results

  • Under a fixed share count, expected dollar harvest is flat across correlation from 0 to 0.95; the largest paired deviation across eight specifications is 0.34% of the level.
  • Correlation inflates the harvest's standard deviation by at most sqrt(N); the measured ratio divided by sqrt(N) averages 1.005 and never leaves [0.997, 1.019].
  • Raising single-name volatility lifts the mean harvest by +46.7%; cutting correlation with volatility fixed moves it by -0.03%.
  • Because deductions are capped by available gains, correlation lowers expected usable losses about 21% at a realistic gain budget while the gross harvest changes -1.1%.
  • On 2,395 days of US equity returns, dollars harvested per $1 differ by +0.0008 (paired t +0.69) while yield differs with a paired t of +5.37.

Summary

The question

Anyone sizing a tax-loss harvesting program wants one number first: how many dollars of loss will this portfolio realise? Published studies put the answer somewhere between half a percent and two percent of value a year. Where a given portfolio lands inside that range depends on its traits, and on one of them the guidance conflicts. Practitioner work tilts direct-indexing targets toward higher dispersion, on the logic that when names disagree there are more losers to sell. A widely read primer argues the other way, that lower correlation raises the value of harvesting, because names that move together eventually leave no losers once the market rises.

We came to this through a mistake of our own. An early cross-sectional model of harvest yield gave average correlation a positive coefficient, with a weight near 0.29. It was an artifact. Finding out why became the paper.

Why correlation cannot move the average

Hold N names with a share count that never changes. Whenever a lot trades below its basis, sell it, book the loss, and buy a replacement so the basis resets down. The loss on any one name depends on that name's own price history and on nothing else. Total loss is a sum of such terms, and the expectation of a sum is the sum of the expectations, so no covariance or factor structure enters anywhere. Hold each name's own behaviour fixed and the expected dollar harvest cannot depend on how the names move together. That is true at every horizon, not only for a new account.

Could something as messy as fat tails break it? We checked with a simulation that recombines identical random shocks at different correlations, so any difference is dependence and not noise. Mean dollar harvest stayed flat from 0 to 0.95 under Student-t returns with five degrees of freedom, under jumps with skewness of -0.86 and kurtosis of 7.0, in a bear market with drift of -4%, at 60% single-name volatility, with 400 names, with monthly monitoring and over twenty years. In Table 1 the largest deviation is 0.34% of the level. Figure 1 draws the dollar harvest per $1 from an 8,000-path run as a flat line.

Figure 1. The same 8,000-path experiment, measured two ways. Left, dollars of realised loss per $1 of initial value, with 95% confidence intervals; the axis spans six standard errors so a real effect would be unmissable. Right, the identical harvest divided by average NAV over the same period.
Figure 1. The same 8,000-path experiment, measured two ways. Left, dollars of realised loss per $1 of initial value, with 95% confidence intervals; the axis spans six standard errors so a real effect would be unmissable. Right, the identical harvest divided by average NAV over the same period.

What correlation does instead

Dependence sets the spread. When all names share the same marginal behaviour, the harvest is an average of N contributions, and its variance lies between Var(g)/N under independence and Var(g) when every name moves in lockstep. The standard deviation can grow by at most sqrt(N). Table 2 tests that ceiling for breadth from 10 to 400: the measured ratio divided by sqrt(N) averages 1.005 and never leaves [0.997, 1.019] (Figure 2).

Figure 2. Left, the standard deviation of the realised ten-year harvest against correlation. Right, the mean with the tenth to ninetieth percentile band; the mean is the flat line from Figure 1.
Figure 2. Left, the standard deviation of the realised ten-year harvest against correlation. Right, the mean with the tenth to ninetieth percentile band; the mean is the flat line from Figure 1.

Numbers make this concrete. Take a hundred-name book at 35% volatility. With independent names, the middle 80% of ten-year outcomes runs from 42% to 49% of initial value; with nearly correlated names it runs from 11% to 80%. The average is the same. Someone who quotes that average for a highly correlated book can miss by a factor of four.

Reconciling the two claims

Both camps had hold of something real. Losses offset gains dollar for dollar, but only up to the gains available, and the rest carries forward. Usable losses are therefore concave in the gross harvest, so a wider spread lowers the expected usable amount even when the gross mean is unchanged. In Table 4 gross harvest moves -1.1% as correlation rises from 0 to 0.95, while the deductible amount falls by about 21% at a realistic gain budget (Figure 4). The primer's sign is right for deductions. Its reason, a shortage of losers, is not.

Figure 4. Left, the change in deductible losses as correlation rises, at four levels of the gains available to absorb them; the dashed line is the gross harvest. Right, the penalty at ρ=0.95 as a function of the gain budget.
Figure 4. Left, the change in deductible losses as correlation rises, at four levels of the gains available to absorb them; the dashed line is the gross harvest. Right, the penalty at ρ=0.95 as a function of the gain budget.

Measurement supplies the second channel. Harvest yield divides dollars by the account's value over the same period, and that value falls just as losses are realised. So yield picks up a dependence the dollars lack: portfolios with a flat dollar harvest show a 95% rise in yield.

What about dispersion? Table 3 raises it two ways. Raising volatility lifts the mean harvest +46.7%, and cutting correlation at fixed volatility moves it -0.03%. Tilting toward dispersion works only when volatility is the thing that rises.

The real data and what remains to predict

A simulation is only as good as its return model. So we also ran a permutation test on 2,395 days of US equity returns (Figure 5), one that needs no model at all: shift each name's history by its own random offset and every name keeps its own behaviour while the links between names vanish, whereas a common shift keeps both. Table 5 compares the arms. Dollars per $1 differ by +0.0008, a paired t of +0.69. Harvest over average NAV differs by +0.0046, a paired t of +5.37. Flat dollars, rising yield, on real prices.

Figure 5. The permutation test on US equity returns. Dependence is destroyed by independent circular shifts and preserved by a common shift, with marginals identical in both arms.
Figure 5. The permutation test on US equity returns. Dependence is destroyed by independent circular shifts and preserved by a common shift, with marginals identical in both arms.

There is a large caveat. The result needs a fixed share count, and most live direct-indexing products rebalance to an index; under quarterly rebalancing the mean harvest falls 53.2% as correlation rises. For those products the theorem is a benchmark and no more.

What is left to forecast? If the conditional mean is pinned down by each name's own behaviour, the useful target becomes a floor, meaning a low quantile of the outcome. We estimated quantiles with quantile regression, gradient boosting and a deep network, then calibrated them by conformalized quantile regression. On 100 held-out configurations (Table 6), gradient boosting reached a pinball loss of 0.0315 and calibrated coverage of 0.907 for the nominal 90 percent interval.

Who this is for: Investors, CPAs and tax professionals, researchers and quants, and journalists who write about direct indexing and tax-loss harvesting.

Figures

Figure 3. The measured inflation of the harvest's standard deviation against the √N ceiling, on log axes.
Figure 3. The measured inflation of the harvest's standard deviation against the √N ceiling, on log axes.
Figure 6. Coverage of the nominal 90% interval within correlation buckets, before and after conformal calibration. The dashed line is nominal.
Figure 6. Coverage of the nominal 90% interval within correlation buckets, before and after conformal calibration. The dashed line is nominal.
Figure 7. Weight placed on correlation by quantile. Left, the linear quantile-regression coefficient, crossing zero at the median. Right, the gradient-boosting gain share, minimised at the median.
Figure 7. Weight placed on correlation by quantile. Left, the linear quantile-regression coefficient, crossing zero at the median. Right, the gradient-boosting gain share, minimised at the median.

Tables

Table 1. Paired invariance test. Expected dollar harvest per $1 of initial value, at ρ = 0 and ρ = 0.9 under a common-random-number coupling. The last two rows violate the theorem's hypothesis on purpose.
SpecificationE[L] at ρ=0at ρ=0.9changepaired tsd ratio
Baseline0.455990.45256-0.75%-1.999.23
Student-t(5) marginals0.454220.45509+0.19%+0.509.17
Jump marginals0.461800.46335+0.34%+0.899.18
Bear market, μ = -4%0.687050.68885+0.26%+1.169.19
Single-name volatility 60%0.746610.74837+0.24%+1.059.18
N = 400 names0.455790.45297-0.62%-0.9118.32
Monthly harvesting0.441150.44229+0.26%+0.649.18
Twenty-year horizon0.537550.53656-0.19%-0.299.23
Quarterly rebalancing1.316850.61620-53.21%-385.064.55
Capacity limit, deepest 10%0.454340.44225-2.66%-6.879.34
Table 2. The √N dispersion bound. Ratio of the comonotone to the independent standard deviation of the harvest, at fixed marginals.
Nsd, independentsd, comonotoneratio√Nratio /√N
100.08310.26323.173.161.001
250.05210.26095.015.001.002
500.03730.26427.097.071.003
1000.02640.26299.9710.000.997
2000.01840.262514.2814.141.009
4000.01290.263820.3920.001.019
Table 3. Two ways to raise cross-sectional dispersion, from a baseline of 30% single-name volatility and ρ = 0.45.
Interventionrealized cross-sectional dispersionmean harvestchanget
Baseline0.2210.3775
Raise single-name volatility to 42%, ρ fixed0.3090.5538+46.7%+553.2
Cut ρ to 0.05, single-name volatility fixed0.2900.3774-0.03%-0.1
Table 4. Effect of dependence on what the investor can deduct. Change from ρ = 0 to ρ = 0.95, marginals fixed.
QuantityChange
Gross harvest-1.1%
Deductible, gains to offset =5% of NAV-3.4%
Deductible, gains to offset =10% of NAV-5.5%
Deductible, gains to offset =20% of NAV-10.3%
Deductible, gains to offset =40% of NAV-20.9%
Present value of deductions, gains =2% of NAV per year-12.9%
Present value of deductions, gains =5% of NAV per year-19.9%
Table 5. Permutation test on US equity returns. 400 random 100-name portfolios, eight replicates, paired by portfolio.
Quantitydependence keptdependence brokendifferencepaired t
Dollars harvested per $10.30050.2998+0.0008+0.69
Harvest divided by average NAV0.16660.1619+0.0046+5.37
Table 6. Predicting the distribution of the ten-year harvest on 100 held-out configurations. Pinball loss averaged over seven quantiles; coverage and width for the nominal 90% interval, before and after conformal calibration; Winkler score after calibration.
Modelpinballcoverage, rawcoverage, calibratedwidth, rawwidth, calibratedWinkler
Theory benchmark0.05320.9630.9141.2261.1031.247
Linear quantile regression0.03170.8750.8880.4310.4420.589
Gradient boosting0.03150.8960.9070.4460.4550.588
Deep quantile network0.03070.8830.8990.4170.4310.562
Table 7. Coverage of the calibrated 90% interval within correlation buckets.
Modelρ < 0.30.3 ≤ ρ ≤ 0.6ρ > 0.6
Theory benchmark0.8580.9500.940
Linear quantile regression0.8720.8970.896
Gradient boosting0.9220.9110.876
Deep quantile network0.8940.9050.899
Table 8. Weight on average pairwise correlation, by quantile. Theorem 1 requires no weight at the median.
τquantile-regression coefficient on ρboosting importance of ρ
0.05-0.2880.078
0.10-0.2400.057
0.25-0.1230.028
0.50+0.0200.013
0.75+0.1270.049
0.90+0.1940.056
0.95+0.2200.058

Abstract

An advisor sizing a tax-managed mandate needs to know how many dollars of harvestable loss a portfolio will produce, and the literature disagrees about one input. Practitioner work recommends tilting a direct-indexing target toward higher return dispersion; a widely read primer reports that lower correlation raises the value of harvesting.

Under a lot-level harvesting rule with a fixed share count, the expected realized loss is invariant to the dependence structure of returns once marginals are held fixed. Each name's cost basis is a functional of that name's own price path, so the expected harvest is a sum of single-name marginal expectations into which no cross-moment enters. Across eight specifications spanning fat tails, jumps, negative drift, breadth and monitoring frequency, the largest paired deviation measured is 0.34% of the level.

Correlation instead governs the second moment, with a sqrt(N) ceiling on the inflation of the harvest's standard deviation that simulation attains to within 2%. That reconciles the disagreement in two steps. Because the deduction is capped by the gains available to absorb it, expected usable losses are concave in the gross harvest and fall about 21% at a realistic gain budget, recovering the reported sign from the variance channel rather than from a shortage of losers. And because harvest yield divides by a contemporaneous NAV, it inherits a correlation dependence that is an artifact of normalization: portfolios whose dollar harvest is flat show a 95% rise in yield. Raising single-name volatility lifts the harvest 47%; raising dispersion by an equal amount through lower correlation moves it 0.03%.

A permutation test on 2,395 days of US equity returns reproduces this with no return model. Since the conditional mean is fixed by the marginals, the predictable object is the conditional quantile, estimated by quantile regression, gradient boosting and a deep quantile network and calibrated by conformalized quantile regression.

Keywords: tax-loss harvesting, direct indexing, dependence invariance, comonotonicity, convex order, conformal prediction, quantile regression, harvest yield, after-tax return, separately managed accounts

How to cite

Majumdar, A. (2026). Dependence Invariance in Tax-Loss Harvesting: An Exact Result and the Distributional Cross-Section of Harvest Yield. SSRN Working Paper No. 7337678. https://ssrn.com/abstract=7337678

Educational only. Not investment, tax, or legal advice, and not an offer of advisory services. InnovationStrat Wealth, LLC is not yet registered as an investment adviser. Questions about a paper? Use the contact form.  ·  ← All papers