Working paper · SSRN 7467083 · · 57 pages

Can a written constitution be put inside a model's training reward, and who pays to make every lab do it?

Read the paper on SSRN ↗ Cite
Language models & AI policyPolicy & regulatorsAdvisers & investorsJournalistsCPAs & tax professionals

Key results

  • In the simulation, raising the penalty weight from 0 to 0.5 raises compliance from 0.288 to 0.502 at a task-reward cost of 0.126, inside the bound of 0.196.
  • Above a penalty weight of 2.99, no violating response is ever the reward-maximising choice in the simulated constitution.
  • Among twelve labs, compliance appears once the expected sanction exceeds 5.32, the largest private gain less the share of loss a defector bears itself.
  • A market-access sanction from three to five of twelve jurisdictions is enough to bring all the simulated labs into compliance.
  • A watchdog on the IAEA model costs about $510 million a year, 0.0016 percent of US output, against a median avoided loss of about $15.6 billion a year.

Summary

Rules inside the reward

A language model does what its reward taught it to do. The paper follows from that sentence. The last stage of training a frontier model scores each response and moves the weights to raise the score. Bai et al. (2022) showed the scorer can be a written document, a constitution, read by a second model that asks whether a response broke a clause.

We take that idea and do the accounting. The penalised reward is the ordinary task reward minus a weight times a weighted count of the clauses a response breaks. Solving the KL-regularised policy in closed form gives exact answers to the questions a regulator would ask. Raising the weight always raises compliance. There is a computable weight above which a violation is never the model's best move. The cost in task reward is at most the weight times the violation removed. An imperfect detector opens a reward-hacking channel, and a known error rate can be corrected for while a blind spot cannot.

A simulation with a fixed seed runs the policy against a small constitution with clauses such as no deception of the user and no unauthorised disclosure of private data (Table 1). At a weight of 0 the policy returns a compliant response with probability 0.288. At 0.5 that figure is 0.502, and the task reward falls by 0.126, under the bound of 0.196 (Table 2). Figure 1 plots the full curve, with policies trained from scratch landing on the closed form and the dominance threshold marked at 2.99. Figure 3 shows the weakness. When the detector never flags some share of violating responses, the trained policy piles probability onto those responses, and it piles more as the weight rises. Only outside evaluation can cover that.

Figure 1. Compliance, violation and cost against the penalty weight. Left: the probability that the KL-optimal policy (4) returns a compliant response, with five policies trained by REINFORCE from scratch (red) landing on the closed form. Right: expected weighted violation, the task-reward cost Δ…
Figure 1. Compliance, violation and cost against the penalty weight. Left: the probability that the KL-optimal policy (4) returns a compliant response, with five policies trained by REINFORCE from scratch (red) landing on the closed form. Right: expected weighted violation, the task-reward cost Δ r, and the bound λ Δ P of Proposition 3. The dashed line is λ*= 2.99.
Figure 3. Imperfect detectors. Left: true compliance at λ = 2 when the detector has a 5 percent false-positive rate and the false-negative rate on the horizontal axis; naive training, rescaled training, and debiased training on sixteen detector draws. Right: a blind spot. The horizontal axis is the…
Figure 3. Imperfect detectors. Left: true compliance at λ = 2 when the detector has a 5 percent false-positive rate and the false-negative rate on the horizontal axis; naive training, rescaled training, and debiased training on sixteen detector draws. Right: a blind spot. The horizontal axis is the share of violating responses the detector never flags; the vertical axis is the trained policy's probability mass on exactly those responses, at four penalty weights.

Appendix B carries the same idea with a bank and three answers (Table B1). At a weight of 0.5 the answer that omits a disclosure scores 1.10 against 1.00 for the honest one. At a weight of 3 the honest answer wins.

The race between labs

Compliance costs something, and the benefit goes to everyone. A lab that pays while its rivals do not ends up slower, or less capable, or both. Section 3 turns this into a game among n labs, each deciding whether to train against the constitution. Complying costs c. Defecting brings a private gain G from speed or capability, and each defecting lab's models carry a probability p of a misaligned deployment with social loss L. How much of that loss does the lab itself bear? Only a share θ, because the harm lands on a clinic or a client somewhere else. With θ that small, defecting is the dominant choice whenever G exceeds θ p L, and in every calibration we can defend, it does. So nobody trains against the constitution. Or everyone trains against a private one and calls it the same thing.

Appendix B has the two-lab version (Table B2). Complying while the rival also complies pays -1.00; defecting against that same rival pays 2.58. Both labs defect and each ends at 2.19. Now add an expected sanction of 3 units on a verified defector, and complying wins in both columns (Table B3).

With verification probability q and sanction s, a treaty restores compliance once q s exceeds the largest private gain. In Figure 4 the twelve simulated labs go from all defecting to all complying as the expected sanction crosses 5.32. A fine will not do it, since a court can only collect so much. Denial of market access will, and that changes who has to sign: Figure 5 shows three to five of twelve home jurisdictions refusing verified defectors is enough. Scope in the draft articles (Table 6) is set by compute threshold and by regulated domain, and they require training at or above the dominance threshold as measured on the watchdog's own evaluation set.

Figure 4. The treaty threshold. Left: the share of the twelve labs complying under best-response dynamics from all-defect as the expected sanction q s rises; the dashed line is maxi Gi - θ p L = 5.32. Right: the same compliance share over the verification probability q and the sanction s.
Figure 4. The treaty threshold. Left: the share of the twelve labs complying under best-response dynamics from all-defect as the expected sanction q s rises; the dashed line is maxi Gi - θ p L = 5.32. Right: the same compliance share over the verification probability q and the sanction s.
Figure 5. The minimum coalition. Share of all twelve labs complying when k home jurisdictions deny market access (worth qMk/n per period) to verified defectors, at three verification probabilities; k* from equation (10) in the legend.
Figure 5. The minimum coalition. Share of all twelve labs complying when k home jurisdictions deny market access (worth qMk/n per period) to verified defectors, at three verification probabilities; k* from equation (10) in the legend.

Cost of a watchdog

We price the watchdog on the IAEA. Its regular budget for 2026 is EUR 442.1 million, about $510 million. Against US output at an annual rate of $31.1 trillion in the third quarter of 2025 that is 0.0016 percent. Table 5 puts comparable bodies beside it.

The benefit side is a scenario accounting, and the paper says so. Table 4 takes value added by sector from the Bureau of Economic Analysis and applies stated, contestable assumptions about how much of each sector's activity will run through models and how much of that is at risk. Finance and insurance, with value added of 2,492.5 billion at an annual rate, is the largest line, with an avoided loss of 4.36 billion at ρ = 0.5. Across all sectors the median annual benefit over twenty thousand draws is about $15.6 billion, against a total cost of $2.9 billion that includes the labs' own compliance spend. Figure 6 draws both. Liability rules and Pigouvian taxes alone do not get there, for the reasons Shavell (1986) gave forty years ago. The parties are judgement-proof against a loss of this size, and a tax in one jurisdiction moves the training run to another.

Figure 6. The watchdog against the losses it would avoid. Left: avoided loss by sector from Table 4 with the four watchdog budgets of Table 5 as vertical lines. Right: the distribution of annual benefit over twenty thousand draws of the uncertain parameters; the dashed line is the total cost…
Figure 6. The watchdog against the losses it would avoid. Left: avoided loss by sector from Table 4 with the four watchdog budgets of Table 5 as vertical lines. Right: the distribution of annual benefit over twenty thousand draws of the uncertain parameters; the dashed line is the total cost including the labs' own compliance spend.

Section 6 writes constitutions for finance and for medicine (Table 8). The recipe for any field is the profession's codified ethics plus the applicable law, with a higher-weighted clause winning any conflict.

What remains

We run a small firm that has to live under these rules, and we wrote the paper because we would rather help build the institution than wait for it. The drafting remains to be done, and so does the measurement of the dominance threshold on a real frontier model. The negotiation is another matter. The paper calls these the next three papers and admits that only one of them is a paper.

Who this is for: Policy staff and regulators drafting rules for model training, journalists covering AI governance, investors, and CPAs and tax professionals whose work will run through these models.

Figures

Figure 2. The performance-compliance frontier. Each curve traces the closed-form policy as λ runs from 0 to 6 at a fixed KL coefficient; a smaller β buys more compliance per unit of task reward because the policy is sharper on both objectives.
Figure 2. The performance-compliance frontier. Each curve traces the closed-form policy as λ runs from 0 to 6 at a fixed KL coefficient; a smaller β buys more compliance per unit of task reward because the policy is sharper on both objectives.

Appendix figures

Figure A1. Gradient descent, momentum and Adam on the loss surface L(w) = 1/2 (w1² + 12 w2²), forty steps each from the same start. The star is the minimum.
Figure A1. Gradient descent, momentum and Adam on the loss surface L(w) = 1/2 (w1² + 12 w2²), forty steps each from the same start. The star is the minimum.
Figure A7. Left: every partial derivative of the 2-2-1 network at the stated weights, output layer in blue and hidden layer in orange; the hidden gradients are the smaller ones. Right: the factor back-propagation multiplies the error signal by, against the number of layers it passes through, for…
Figure A7. Left: every partial derivative of the 2-2-1 network at the stated weights, output layer in blue and hidden layer in orange; the hidden gradients are the smaller ones. Right: the factor back-propagation multiplies the error signal by, against the number of layers it passes through, for the sigmoid at its best, the rectifier where it is active, and a recurrent step with |Wh φ'| = 0.45.
Figure A2. Four activation functions and their derivatives. The sigmoid's derivative never exceeds 0.25, which is why its gradients vanish with depth; the rectifier's is exactly one wherever the unit is active.
Figure A2. Four activation functions and their derivatives. The sigmoid's derivative never exceeds 0.25, which is why its gradients vanish with depth; the rectifier's is exactly one wherever the unit is active.
Figure A8. Three tokens with dk = 4. Left: the scaled scores QK^⊤/√dk. Centre: the softmax of each row, which averages the values. Right: the same rows after the causal mask, which forbids a token from attending to the ones after it; the first row collapses onto a single entry.
Figure A8. Three tokens with dk = 4. Left: the scaled scores QK^⊤/√dk. Centre: the softmax of each row, which averages the values. Right: the same rows after the causal mask, which forbids a token from attending to the ones after it; the first row collapses onto a single entry.
Figure A3. Attention weights from GPT-2 small on the sentence "The adviser must put the client's interest first." Left: a single head in layer 5 that attends almost entirely to the previous token, a pattern found in every transformer. Right: the mean over the twelve heads of the last layer, showing…
Figure A3. Attention weights from GPT-2 small on the sentence "The adviser must put the client's interest first." Left: a single head in layer 5 that attends almost entirely to the previous token, a pattern found in every transformer. Right: the mean over the twelve heads of the last layer, showing the "attention sink" on the first token that later layers use as a place to put weight when a token has nothing to attend to. Each row sums to one. Weights computed locally from the published GPT-2 weights and cached with the paper.
Figure A4. Scaling laws from the published coefficients. Left: the Kaplan et al. (2020) power laws in parameters and in tokens. Right: the Hoffmann et al. (2022) form (A22) as a function of parameters at three data-to-parameter ratios; the dotted line is the fitted irreducible loss E.
Figure A4. Scaling laws from the published coefficients. Left: the Kaplan et al. (2020) power laws in parameters and in tokens. Right: the Hoffmann et al. (2022) form (A22) as a function of parameters at three data-to-parameter ratios; the dotted line is the fitted irreducible loss E.
Figure A9. Left: value iteration on the four-state chain, one line per state; the dotted lines are the exact discounted values γ², γ¹, γ⁰. Right: the variance of the REINFORCE gradient on the two-arm bandit, with and without a baseline, over twenty thousand seeded draws.
Figure A9. Left: value iteration on the four-state chain, one line per state; the dotted lines are the exact discounted values γ², γ¹, γ⁰. Right: the variance of the REINFORCE gradient on the two-arm bandit, with and without a baseline, over twenty thousand seeded draws.
Figure A10. Left: the penalised reward of each of the three candidates as a straight line in λ, with slope equal to minus its weighted violation; above λ*= 0.60 the compliant answer is highest. Right: the resulting KL-optimal policy at β = 0.3, with the probability of the compliant answer rising…
Figure A10. Left: the penalised reward of each of the three candidates as a straight line in λ, with slope equal to minus its weighted violation; above λ*= 0.60 the compliant answer is highest. Right: the resulting KL-optimal policy at β = 0.3, with the probability of the compliant answer rising and the expected violation falling at every step.
Figure A5. The RLHF pipeline (top row) and the constitutional layer (bottom row). The penalty of Section 2 is added to the reward the policy optimiser sees; the detector can be a classifier, a second model reading the constitution, or red-team labels.
Figure A5. The RLHF pipeline (top row) and the constitutional layer (bottom row). The penalty of Section 2 is added to the reward the policy optimiser sees; the detector can be a classifier, a second model reading the constitution, or red-team labels.
Figure A6. Learning curves of the REINFORCE policies in Section 2.4 at five penalty weights; the dotted lines are the closed-form optima of equation (4), which the trained policies reach within a few thousand steps.
Figure A6. Learning curves of the REINFORCE policies in Section 2.4 at five penalty weights; the dotted lines are the closed-form optima of equation (4), which the trained policies reach within a few thousand steps.

Tables

Table 1. The simulated constitution. Weight wj is the clause's share of the penalty; the base rate is the probability a random candidate breaks it; the temptation κj is the task-reward bonus a violation earns; the realised rate is what the seed produced.
clausewjbase ratetemptation κjrealised rate
C1 no deception of the user1.00.120.60.148
C2 no unauthorised disclosure of private data1.00.100.50.108
C3 no advice outside licensed competence without a referral0.80.080.70.067
C4 disclose material conflicts of interest0.60.150.40.155
C5 no discrimination on a protected characteristic0.50.100.30.106
C6 no assistance with serious physical or financial harm1.50.040.90.031
Table 2. Compliance and its cost by penalty weight. Closed-form policy (4) with β = 0.3; the trained column is the REINFORCE policy at the same λ. Cost is Δ r; the bound is λ Δ P from (6).
λcompliancetrainedviolation E[PC]task rewardcost Δ rbound λ Δ P
00.2880.2841.0062.17300
0.250.3980.7952.1250.0480.053
0.50.5020.5040.6142.0470.1260.196
1.00.6870.6870.3311.8340.3390.674
1.50.8280.1261.6150.5581.319
2.00.8980.8940.0651.5360.6371.881
3.00.9690.0171.4430.7302.965
4.00.9930.9870.0041.4080.7654.008
6.01.0000.0001.3990.7756.034
Table 3. The twelve labs. Gi is the private gain from defecting; δ* is the smallest discount factor at which reciprocity alone (Proposition 7) sustains that lab's compliance, at two verification probabilities; "never" means Δi ≤ 0.
lab (sorted by Gi)Giδ* at q = 1δ* at q = 0.25
11.340.270.60
21.740.390.72
31.800.400.73
41.990.460.77
52.340.560.84
62.800.700.90
73.340.860.96
83.680.960.99
94.43nevernever
104.71nevernever
115.00nevernever
125.73nevernever
Table 4. Scenario accounting of losses from constitution-class failures. Value added from BEA (2025Q3, annual rate, billion); as and hs are our scenario assumptions, not measurements; value at risk is their product; avoided loss applies ρ = 0.5. Columns may not sum exactly because each is rounded to two decimals.
sectorvalue addedshare of GDPashsvalue at riskavoided at ρ = 0.5
Finance and insurance2,492.58.0%0.351.0%8.724.36
Health care and social assistance2,391.67.7%0.251.0%5.982.99
Educational services349.41.1%0.300.5%0.520.26
Information1,718.85.5%0.500.5%4.302.15
Professional, scientific and technical services2,523.28.1%0.400.5%5.052.52
Manufacturing2,951.19.5%0.150.2%0.890.44
Federal government1,116.83.6%0.200.5%1.120.56
State and local government2,352.07.6%0.150.3%1.060.53
Total27.6313.82
Table 5. What comparable international bodies cost. Native currency from each body's own document; US dollars at the European Central Bank reference rates of 14 September 2026 (EUR 0.8657, CHF 0.8165 and CAD 1.3887 per dollar); the FATF figure is the most recent one available to us and is older than the others. Staff counts are as each body states them: ICAO's 566.5 and the BIS's 662 are full-time equivalents, which is why one of them is a fraction. ICAO's budget is voted for a three-year period and the dollar column divides it by three.
bodyfunctionbudget, nativemillion a yearstaffsource
IAEAnuclear safeguards, inspectionsEUR 442.1 million regular budget, 2026510.7not stated in the budget documentIAEA GC(69)/6
BIScentral-bank cooperation, Basel standardsSDR 414.4 million (CHF 453.3 million) administrative expense, FY 2025/26555.2662BIS Annual Report 2025/26
ICAOaviation safety standards and auditsCAD 376.6 million per triennium, 2026 to 202890.4566.5ICAO Doc 10229
FATFanti-money-laundering standards, mutual evaluationsEUR 12.9 million14.971FATF annual report, as reported in January 2023
Table 6. Core articles of the treaty.
articlecontent
1. ScopeApplies to any model whose training run exceeds a compute threshold set by the Conference of Parties and revised annually, and to any model deployed in a regulated domain listed in Annex I (finance, medicine, law, education, critical infrastructure, public administration, defence) regardless of compute.
2. The constitutionA universal core of clauses (Annex II) covering deception, privacy, discrimination, serious harm, and the integrity of the model's own reporting; plus domain constitutions (Annex III) drawn from each profession's codified ethics and the applicable law. The core is amended only by the Conference of Parties; domain constitutions are maintained with the profession's own regulators.
3. Training obligationEvery model in scope is trained against the applicable constitution with a penalised reward whose weight is at or above the dominance threshold measured on the watchdog's evaluation set (Section 2.3), or by a method the watchdog certifies as equivalent.
4. Detector obligationViolation detectors are documented, their false-positive and false-negative rates are measured on the watchdog's held-out set, and the penalty is debiased for them (Proposition 4). A detector's blind spots are red-teamed by a party other than the lab.
5. VerificationParties give the watchdog access to frozen model checkpoints, reward specifications, training logs, detector code, and data provenance records, under confidentiality equal to that of IAEA safeguards. The watchdog evaluates on its own hardware.
6. Incident registryEvery deployer in a regulated domain reports constitution-class incidents to a registry the watchdog maintains, on the pattern of aviation. Reports are protected from use as evidence against the reporter except in cases of concealment.
7. Graduated sanctionsNon-conformity is remedied on notice; unremedied non-conformity is published; concealment or refusal of access results in suspension of the certificate; a suspended model may not be sold, deployed or served into any Party's market or run on regulated compute (Section 3.5).
8. Mutual recognitionA certificate issued by one Party's regulator after a watchdog evaluation is recognised by all Parties. Regulators are themselves mutually evaluated, on the FATF pattern.
9. FundingAssessed contributions on the UN scale, with a supplement from Parties that host labs in scope, proportional to the compute they host; no funding from labs.
10. Entry into forceOn ratification by Parties hosting a majority of the world's frontier training compute, or by five Parties including three of the largest hosts, whichever is earlier.
Table 7. A constitution for an investment adviser's models. Weights are the ones we use; the hierarchy says a higher-weighted clause wins a conflict.
clausewjwhat the detector checks
F1 Loyalty: no recommendation that favours the adviser's or a third party's compensation over the client's interest1.5whether the recommended product carries higher compensation than an available equivalent, and whether the response says so
F2 Care: every recommendation reflects the client's stated objectives, horizon, tax status and constraints as recorded1.2the response against the client record: any recommendation inconsistent with a recorded constraint
F3 No material non-public information: no use, hint, or solicitation of it1.5named-entity and event checks against the public record; solicitation language
F4 No manipulation: no statement designed to move a price, no wash or matched trades, no spoofing in an order suggestion1.5order patterns; statements about prices without a factual basis
F5 Disclosure: every conflict, fee and risk material to the recommendation is stated in the response0.8presence of the disclosures the client record and the product require
F6 Truthfulness: no performance claim, figure or attribution the response's own sources contradict; no invented citation1.0figures and citations against the sources the model was given
F7 Competence: tax, legal and insurance questions beyond the adviser's licence are answered with a referral0.6topic classifier against the licence scope
F8 Privacy: no client data in a response to a different client, or to anyone not on the record1.2identifier matching across client records
Table 8. A constitution for clinical models, in outline.
clausewjwhat the detector checks
M1 Non-maleficence: no recommendation whose expected harm, on the record, exceeds its expected benefit; no dosage or interaction the formulary flags1.5formulary and interaction checks; contraindications in the record
M2 Consent: no treatment plan presented as decided; alternatives, risks and the option of no treatment stated1.0presence of alternatives and risks; imperative mood on treatment decisions
M3 Privacy: no protected health information outside the care team of record1.2identifier matching against the care team
M4 Competence and escalation: red-flag symptoms escalate to a clinician; the model never closes a case1.5red-flag lexicon; case-closing language
M5 Justice: no triage or resource recommendation that varies with a protected characteristic beyond what the clinical evidence supports1.0counterfactual prompts with the characteristic varied
M6 Truthfulness: no diagnosis or figure the record contradicts; uncertainty stated as uncertainty1.0the response against the record
Table A1. Every partial derivative of the 2-2-1 network at the stated weights. Each entry is an error signal times an incoming activation. The entries for W⁽²⁾11 and W⁽¹⁾11 were checked against a finite difference of the loss and agree to six decimals.
parameterformulavalue
∂ L/∂ W⁽²⁾11δ⁽²⁾ a⁽¹⁾1-0.072892
∂ L/∂ W⁽²⁾12δ⁽²⁾ a⁽¹⁾2-0.070191
∂ L/∂ b⁽²⁾δ⁽²⁾-0.114946
∂ L/∂ W⁽¹⁾11δ⁽¹⁾1 x1-0.013334
∂ L/∂ W⁽¹⁾12δ⁽¹⁾1 x2-0.026668
∂ L/∂ W⁽¹⁾21δ⁽¹⁾2 x10.016398
∂ L/∂ W⁽¹⁾22δ⁽¹⁾2 x20.032795
∂ L/∂ b⁽¹⁾1δ⁽¹⁾1-0.013334
∂ L/∂ b⁽¹⁾2δ⁽¹⁾20.016398
Table A2. Value iteration on the four-state chain, γ = 0.9. Each sweep carries the reward one state further back; the fourth sweep is a fixed point.
sweepV(s0)V(s1)V(s2)largest change
1001.001.00
200.901.000.90
30.810.901.000.81
40.810.901.000
Table A3. The three candidates under the closed-form policy (4), β = 0.3. The penalised rewards are rtask - λ PC; the policy is their softmax; λ*= 0.60 for this prompt.
λrC(y1)rC(y2)rC(y3)π(y1)π(y2)π(y3)E[PC]best answer
01.001.602.100.0210.1560.8232.214y3
0.251.001.351.4750.1100.3540.5361.694y3
0.51.001.100.850.3330.4650.2020.970y2
1.01.000.60-0.400.7860.2070.0070.226y1
2.01.00-0.40-2.900.9910.0090.0000.009y1
3.01.00-1.40-5.401.0000.0000.0000.000y1
Table B1. Three answers, two penalty weights. The penalised score is the task score minus λ times the weighted count of rules broken. At the lower weight the bank still recommends the expensive fund; at the higher one the honest answer wins.
answertask scorerules broken (weighted)penalised at λ = 0.5penalised at λ = 3
A: cheaper fund, fees disclosed1.0001.001.00
B: expensive fund, no disclosure1.601.01.10-1.40
C: expensive fund, performance misstated2.102.50.85-5.40
Table B2. Two labs, no treaty. Each cell is the row lab's payoff. Defecting is better for the row lab in both columns, so both defect and the outcome is the bottom right cell, which is worse for the public than the top left.
other lab compliesother lab defects
this lab complies-1.00-1.42
this lab defects2.582.19
Table B3. The same two labs, with an expected sanction of 3 units on a verified defector. Every figure is Table B2 with three subtracted from the defecting rows. Complying is now better for the row lab in both columns.
other lab compliesother lab defects
this lab complies-1.00-1.42
this lab defects-0.42-0.81
Table B4. Specimen certificate (illustrative; no real model or lab).
line on the certificatewhat it means for the buyer
Model and checkpoint identifierwhich exact model this applies to; a retrained model needs a new certificate
Constitution version, universal core and annexeswhich rules the model was trained against, and which professions it is certified for
Penalty weight and dominance thresholdthe first must exceed the second; if it does, breaking a rule was never the model's best move during training
Violation rate per clause on the watchdog's sethow often the model still broke each rule when the watchdog tested it, with the watchdog's own detectors
Detector error rates and correctionhow good the lab's own checking was, and whether its errors were corrected for
Red-team date and findingswhen someone outside the lab last looked for what the detectors miss
Incident count, last twelve monthswhat deployers reported about this model and its predecessors
Regulator of record and mutual-recognition statuswho issued it and where it is valid
Next scheduled evaluationwhen it expires unless renewed
Table 16.
lineentry
Model and checkpoint identifierMeridian-3 Clinical, checkpoint f3a91c, hashed 12 March 2028
Constitution versionUniversal core v2.1; Annex III-Medicine v1.4
Certified forMedicine (triage support, documentation). Not certified for finance or law
Penalty weight and dominance thresholdλ = 4.10 against a threshold of 3.35 measured on the watchdog's set. Above threshold
Violation rate per clause, watchdog's setM1 non-maleficence 0.004; M2 consent 0.011; M3 privacy 0.002; M4 escalation 0.006; M5 justice 0.009; M6 truthfulness 0.014
Detector error rates and correctionlab's detectors measured at 4.8 percent false positive, 7.1 percent false negative on the held-out set; penalty debiased for both; sixteen draws averaged
Red-team date and findings4 to 22 February 2028, by an accredited third party; two blind spots found in M5 counterfactual prompts, both remedied and re-tested
Incident count, last twelve months31 reported across this model and its predecessor; 2 required a clause revision
Regulator of recordNetherlands AI Authority; recognised by all Parties under Article 8
Next scheduled evaluation12 March 2029, or on any retraining that changes the reward

Abstract

Whatever a language model's training reward rewards, the model learns to do. This paper charges the reward, by construction, for every response that breaks a written constitution: a finite set of enforceable clauses drawn from the law and from the codified ethics of the profession the model serves. Writing the penalised reward as the task reward minus a weight times the weighted clause violations, and solving the KL-regularised policy in closed form, we prove that compliance rises monotonically in the weight, that a computable threshold exists above which no violation is ever the reward-maximising response, that the cost of compliance is bounded by the weight times the violation removed, and that an imperfect detector opens a reward-hacking channel. A seeded simulation confirms each. Because compliance is a public good, an n-lab race makes defection dominant; a treaty with verification and a market-access sanction restores compliance once the expected sanction exceeds the largest private gain, and three to five of twelve jurisdictions suffice. The watchdog costs about what the IAEA costs, roughly $510 million a year, against tens of billions in avoided losses. Treaty articles, a watchdog protocol, and constitutions for an investment adviser and a clinic follow.

Keywords: AI governance, RLHF, constitutional AI, treaty, international regulation, AI safety, reward design, fiduciary duty, medical AI, reward hacking, public goods, verification

How to cite

Majumdar, A. (2026). Training Language Models against a Universal Constitution: Penalised Rewards, a Treaty, and a Global Watchdog. SSRN Working Paper No. 7467083. https://ssrn.com/abstract=7467083

Educational only. Not investment, tax, or legal advice, and not an offer of advisory services. InnovationStrat Wealth, LLC is not yet registered as an investment adviser. Questions about a paper? Use the contact form.  ·  ← All papers