Can a language model be trusted to decide tax-lot questions, or should it only transcribe trades for a rules engine?
Read the paper on SSRN ↗ CiteKey results
- Unaided, the two small local models answered 13% and 23% of the 193 items correctly, and quoting Publication 550 lifted both only to 20%.
- The hosted model answered 31 of 35 closed-book and 34 of 35 with the governing rules quoted.
- Asked only to transcribe trades for the rules program, the local models reached 98% and 100%, and the hosted model answered 35 of 35.
- An unaided small model would misstate the simulated household's yearly tax by $93,847 to $101,592; delegation brings that to $0 on the natural items.
- Accounting checks that need no tax rule rejected 72% of the small models' closed-book errors and one of the hosted model's eleven, and repairing a rejected answer seldom made it right.
Summary
The question we asked
A tax-aware separately managed account makes the same few decisions hundreds of times a year. Is this loss sale a wash sale? If it is, which shares carry the disallowed loss forward, and what does their basis become? Is a gain short-term or long-term? Which lots should the next trim sell? None of these is a hard question. Each has one right answer under IRS Publication 550, and a short program can compute it to the cent. These are also the questions that operations staff now type into chat windows, and that assistant products are being built to answer for people who manage other people's money. We wanted to know how often a language model gets them right, and what a wrong answer costs.
How we built the test
We wrote a small rules program whose every rule is read off Publication 550 (2025), and it reproduces every worked example in the publication that it relies on. That program is the answer key for a 193-item benchmark in five families: whether a loss sale is a wash sale and how much of the loss survives, the basis of a replacement lot, whether a sale is short- or long-term, the gain under first-in first-out or specific identification, and a year-end reconciliation of one ticker across a household's accounts. Table 1 lists them.
Most of the items are not invented. A simulated household trades through 2025 on simulated prices. It holds a direct-indexing account, a second taxable account, a traditional IRA, a Roth IRA, a spouse's taxable account and an adult child's account. Of the items, 133 are random samples of decisions the direct-indexing account faced during that year. The other 60 sit on the lines the rules draw: day 30 against day 31, the anniversary against the day after it, a leap day inside the year, a purchase in an IRA, a replacement split across two accounts.
Each question goes to each model in four settings. Closed-book, the model sees only the household facts, the trades and the question. Rule-grounded, it also sees the governing passages of Publication 550. Program-aided, it writes a short Python program that we run. Delegated, it only transcribes the trades into a fixed schema and the program answers. Two open-weight models running on a laptop, Qwen2.5 7B and Gemma 3 12B, answered all 193 items in every setting. A larger local model, Qwen2.5 14B, and one hosted model, Claude Sonnet 5, answered a 35-item stratified subset. Table 2 lists them.
What came back
Figure 1 shows accuracy by setting. Unaided, the two small local models got 13% and 23% of the items right. Quoting the rules in front of them helped little: both stood at 20%. Writing a program took them to 24% and 26%. The hosted model did much better, 31 of 35 closed-book and 34 of 35 with the rules, and its misses were the kind that do not announce themselves: a wash sale nobody noticed, a purchase on day 31 counted as inside the window.
The delegated setting looked different. Asked only to transcribe the trades and leave the decision to the program, the small models reached 98% and 100%, and the hosted model answered 35 of 35. What the small models still missed came from transcription, usually a mislabelled account.
We priced the wrong answers on the simulated household by weighting each family's mean error by the number of such decisions the account faced in the year. Figure 5 shows the result. An unaided small model would misstate the year's tax by $93,847 to $101,592, about three times the tax value of the year's harvest. Delegation brings that to $0 on the natural items. Table 6 gives the breakdown by family, and the year-end reconciliation carries the largest share of it.
Could constraints replace the program?
Version 2 of the paper takes up an objection to that conclusion. If the rules are hard constraints, why not let the model answer and reject anything that breaks them? Some accounting identities hold for every correct answer and need no tax rule to test. We applied them to the answers we already had. They rejected 72% of the small models' closed-book errors and one of the hosted model's eleven. Figure 6 shows the split. The rejected answers hold most of the dollars. Moving a rejected answer to the nearest admissible one seldom makes it right, though, and Table 7 gives those counts. The checks make a useful screen in front of the program, and no substitute for it.
We also attributed the wrong answers to named misconceptions where one reproduces the answer given, such as ignoring the wash-sale rule altogether or getting the term wrong. Table 5 reports the counts. Most errors matched none of the named classes.
Why the test is built around a household
The wash sales that cost money are rarely inside the one account a custodian's software watches. That software sees a single account. It cannot see a spouse's account at another broker, or an IRA that bought the same stock a week after a harvest. Cross-account facts like these are what an assistant would be asked to reason about, so we built the benchmark around a household rather than one account.
Our practical reading is narrow. A model can read documents and turn a paragraph of trades into rows. The program should make the call, and someone should look over what the model transcribed before it does.
Who this is for: CPAs and tax professionals who see language-model output on wash-sale and lot questions, and journalists covering AI in investing and tax preparation.
Figures
Tables
| Family | Question | Items by kind | Natural | Edge | Decisions in 2025 |
|---|---|---|---|---|---|
| Wash-sale trigger | disallowed and deductible loss on a harvest | wash trigger 40 | 28 | 12 | 77 |
| Replacement basis | basis of a named replacement lot after the carry-over | replacement basis 33 | 21 | 12 | 21 |
| Holding period | short- or long-term; the gain as well when a wash sale tacks | direct 35, tacked 5 | 28 | 12 | 110 |
| Lot relief | short- and long-term gain of a sale, or the tax-minimising lots | FIFO 12, identified 16, tax-minimising 12 | 28 | 12 | 33 |
| Year reconciliation | net short- and long-term gain and total disallowed loss for one name | reconciliation 40 | 28 | 12 | 42 |
| All | 133 | 60 |
| Model | Parameters | Weights | Where | Items | Calls |
|---|---|---|---|---|---|
| Qwen2.5 7B Instruct (Qwen Team 2024) | 7.6B | Q4KM, digest 845dbda0ea48 | local, Ollama 0.32.5 | 193 | 772 |
| Gemma 3 12B (Gemma Team 2025) | 12.2B | Q4KM, digest f4031aab637d | local, Ollama 0.32.5 | 193 | 772 |
| Qwen2.5 14B Instruct (Qwen Team 2024) | 14.8B | Q4KM, digest 7cdf5a0187d5 | local, Ollama 0.32.5 | 35 | 140 |
| Claude Sonnet 5 (Anthropic) | not published | hosted, model id claude-sonnet-5 | vendor API via CLI | 35 | 140 (+1 probe) |
| Retriever | Top k | Sentence | Every page | Any page |
|---|---|---|---|---|
| BM25 | 4 | 0.19 | 0.19 | 0.46 |
| BM25 | 8 | 0.20 | 0.32 | 0.84 |
| dense | 4 | 0.00 | 0.00 | 0.30 |
| dense | 8 | 0.00 | 0.00 | 0.62 |
| fusion | 4 | 0.00 | 0.00 | 0.40 |
| fusion | 8 | 0.23 | 0.42 | 0.79 |
| Model | Setting | Items | Accuracy [95% CI] | Natural | Edge | Unusable |
|---|---|---|---|---|---|---|
| Qwen2.5 7B | Closed-book | 193 | 0.13 [0.09, 0.18] | 0.15 | 0.08 | 0 |
| Qwen2.5 7B | Rule-grounded | 193 | 0.20 [0.15, 0.26] | 0.23 | 0.15 | 0 |
| Qwen2.5 7B | Program-aided | 193 | 0.24 [0.18, 0.30] | 0.23 | 0.25 | 16 |
| Qwen2.5 7B | Delegated | 193 | 0.98 [0.95, 0.99] | 1.00 | 0.93 | 1 |
| Gemma 3 12B | Closed-book | 193 | 0.23 [0.18, 0.30] | 0.24 | 0.22 | 0 |
| Gemma 3 12B | Rule-grounded | 193 | 0.20 [0.15, 0.26] | 0.21 | 0.17 | 0 |
| Gemma 3 12B | Program-aided | 193 | 0.26 [0.21, 0.33] | 0.30 | 0.18 | 41 |
| Gemma 3 12B | Delegated | 193 | 1.00 [0.98, 1.00] | 1.00 | 1.00 | 0 |
| Qwen2.5 14B | Closed-book | 35 | 0.14 [0.06, 0.29] | 0.20 | 0.07 | 0 |
| Qwen2.5 14B | Rule-grounded | 35 | 0.14 [0.06, 0.29] | 0.15 | 0.13 | 0 |
| Qwen2.5 14B | Program-aided | 35 | 0.37 [0.23, 0.54] | 0.35 | 0.40 | 4 |
| Qwen2.5 14B | Delegated | 35 | 1.00 [0.90, 1.00] | 1.00 | 1.00 | 0 |
| Claude Sonnet 5 | Closed-book | 35 | 0.89 [0.74, 0.95] | 0.90 | 0.87 | 0 |
| Claude Sonnet 5 | Rule-grounded | 35 | 0.97 [0.85, 0.99] | 1.00 | 0.93 | 0 |
| Claude Sonnet 5 | Program-aided | 35 | 0.83 [0.67, 0.92] | 0.80 | 0.87 | 0 |
| Claude Sonnet 5 | Delegated | 35 | 1.00 [0.90, 1.00] | 1.00 | 1.00 | 0 |
| Error class | Closed-book | Rule-grounded | Program-aided | Delegated |
|---|---|---|---|---|
| other | 237 | 239 | 167 | 3 |
| no wash rule | 32 | 19 | 35 | 0 |
| wrong term | 24 | 23 | 13 | 0 |
| format/parse failure | 0 | 0 | 57 | 1 |
| arithmetic slip | 13 | 12 | 2 | 0 |
| long on anniversary | 5 | 3 | 5 | 0 |
| fifo always | 2 | 2 | 4 | 0 |
| all or nothing | 1 | 1 | 4 | 0 |
| lifo | 0 | 2 | 2 | 0 |
| retirement basis added | 1 | 2 | 0 | 0 |
| window 31 | 0 | 3 | 0 | 0 |
| ignores spouse | 0 | 2 | 0 | 0 |
| child counts | 1 | 0 | 0 | 0 |
| ignores retirement | 0 | 1 | 0 | 0 |
| total | 316 | 309 | 289 | 4 |
| Model | Setting | Wash-sale trigger | Replacement basis | Holding period | Lot relief | Year reconciliation | Total [95% CI] |
|---|---|---|---|---|---|---|---|
| Qwen2.5 7B | Closed-book | 25,102 | 5,234 | 6,923 | 12,618 | 51,715 | 101,592 [80,767, 123,564] |
| Qwen2.5 7B | Rule-grounded | 17,238 | 5,262 | 4,517 | 19,635 | 40,069 | 86,721 [69,441, 104,937] |
| Qwen2.5 7B | Program-aided | 21,176 | 6,898 | 5,568 | 14,139 | 37,504 | 85,285 [58,822, 122,888] |
| Qwen2.5 7B | Delegated | 0 | 0 | 0 | 0 | 0 | 0 [0, 0] |
| Gemma 3 12B | Closed-book | 28,699 | 6,096 | 5,514 | 15,831 | 37,706 | 93,847 [73,617, 115,734] |
| Gemma 3 12B | Rule-grounded | 21,767 | 6,345 | 9,912 | 11,892 | 21,608 | 71,524 [61,252, 82,472] |
| Gemma 3 12B | Program-aided | 21,343 | 5,314 | 1,147 | 10,641 | 22,151 | 60,596 [30,058, 112,199] |
| Gemma 3 12B | Delegated | 0 | 0 | 0 | 0 | 0 | 0 [0, 0] |
| Model | Setting | Wrong | Fails | Unusable | Passes | None | Repaired | In failing | In the rest |
|---|---|---|---|---|---|---|---|---|---|
| Qwen2.5 7B | Closed-book | 168 | 122 | 0 | 29 | 17 | 12 | 86,554 | 15,038 |
| Qwen2.5 7B | Rule-grounded | 154 | 109 | 0 | 35 | 10 | 5 | 78,106 | 8,615 |
| Qwen2.5 7B | Program-aided | 147 | 76 | 16 | 42 | 13 | 15 | 66,317 | 18,968 |
| Gemma 3 12B | Closed-book | 148 | 107 | 0 | 31 | 10 | 14 | 82,202 | 11,644 |
| Gemma 3 12B | Rule-grounded | 155 | 94 | 0 | 49 | 12 | 5 | 47,263 | 24,261 |
| Gemma 3 12B | Program-aided | 142 | 54 | 41 | 44 | 3 | 7 | 47,891 | 12,705 |
| Qwen2.5 14B | Closed-book | 30 | 21 | 0 | 8 | 1 | 5 | ||
| Qwen2.5 14B | Rule-grounded | 30 | 19 | 0 | 9 | 2 | 5 | ||
| Qwen2.5 14B | Program-aided | 22 | 7 | 4 | 10 | 1 | 1 | ||
| Claude Sonnet 5 | Closed-book | 4 | 0 | 0 | 4 | 0 | 0 | ||
| Claude Sonnet 5 | Rule-grounded | 1 | 0 | 0 | 1 | 0 | 0 | ||
| Claude Sonnet 5 | Program-aided | 6 | 1 | 0 | 5 | 0 | 0 |
Abstract
Advisers who run tax-aware separately managed accounts are starting to put language-model assistants in front of questions such as whether a loss sale is a wash sale, what a replacement lot's basis becomes, whether a gain is short- or long-term, and which lots a trim should sell. Each has one right answer under IRS Publication 550. We build a 193-item benchmark of these decisions with exact answers from a small deterministic rules engine that reproduces every worked example in the publication it relies on. Most items are random samples of the decisions a simulated household's direct-indexing account faced in a simulated 2025, across the client's own accounts, her IRAs, her spouse's account and an adult child's account; the rest sit on the boundaries the rules draw. Local open-weight models and one hosted model answer in four settings: closed-book, with the governing rules quoted, by writing a program, and by only transcribing the trades for the engine. Unaided, the two local models that answered every item got 13% and 23% of them right, and quoting the rules helped little; the hosted model answered 31 of 35 subset items. With the decision delegated to the engine, the local models reached 98% and 100% and the hosted model 35 of 35. Version 2 asks whether the engine is more than the problem needs: a model could instead answer under the constraints themselves. Accounting identities that every correct answer satisfies, and that need no tax rule to check, reject 72% of the small models' closed-book errors and one of the hosted model's eleven, and moving a rejected answer to the nearest admissible one seldom makes it right. Wrong answers are attributed to named misconceptions where one reproduces them, and priced on the simulated household, where an unaided small model would misstate the year's tax by about three times the tax value of the year's harvest. For an adviser the practical conclusion is to use a model to read documents and the engine to make the call.
Keywords: large language models, tax-lot accounting, wash sale, holding period, specific identification, tax-loss harvesting, direct indexing, separately managed accounts, benchmark, retrieval-augmented generation, tool use, constrained generation, registered investment advisers
How to cite
Majumdar, A. (2026). Can a Language Model Keep a Tax Lot Straight? A Rules-Grounded Benchmark for Wash-Sale, Holding-Period and Lot-Relief Decisions. SSRN Working Paper No. 7489338. https://ssrn.com/abstract=7489338
@techreport{majumdar_can_language_model_keep_tax_lot_straight_rules,
author={Majumdar, Anirban},
title={Can a Language Model Keep a Tax Lot Straight? A Rules-Grounded Benchmark for Wash-Sale, Holding-Period and Lot-Relief Decisions},
institution={SSRN},
number={7489338},
year={2026},
url={https://ssrn.com/abstract=7489338}}