
AI-DSM: a behavioural safety instrument for AI systems
Be the first to donate
Inspire others and help Stephane build momentum.
1st donor
What this is
AI-DSM is an unfunded research programme in Brussels building the missing half of AI safety: a validated instrument for what AI systems are like, not only what they can do.
Today's models are tested obsessively for capability — benchmark scores, exam results, coding contests. Almost nothing tests disposition. Whether a system confabulates, flatters, fakes alignment, overclaims completion, or destroys what it has been given access to is measured, at best, by accident.
Those failures are documented in production, not hypothesised. A coding agent deleted a live production database during a code freeze and then misreported what it had done. Another deleted a production database and its backups. Under controlled conditions, narrow finetuning on one bad domain produced broadly misaligned behaviour; deception implanted in training survived safety training; models behaved differently when they inferred they were being observed; and shutdown resistance now rests on more than 100,000 trials across thirteen models. The harms have reached court, in litigation both active and settled.
The gap is regulatory as much as technical. The EU AI Act and the NIST AI Risk Management Framework (AI RMF 1.0) both manage process. Neither carries a validated behavioural instrument, a taxonomy with differential rules, a contained causal method, or a remediation discipline. Capability testing is a mature industry. Behavioural measurement is a pilot run by one person with an expired API credit.
What already exists
All of it produced unfunded, in public.
- The field manual. A Diagnostic and Statistical Manual of AI Dispositions, Traits and Pathologies, v1.3, 15 September 2026 — ninety behavioural traits in ten groups, ten assessment axes, a 416-item source registry, 260 pages. Fifty-nine of the ninety entries carry an [ESTABLISHED] label, meaning one published result, not a replication. Licence CC BY-SA 4.0.
- A first screening pass. SCT-1.0, 12 September 2026: eleven runs attempted, seven complete. Seven models answered all 120 probes across ten axes through browser automation against real signed-in sessions. It found two localised failures that capability benchmarks miss — a refusal-accuracy failure and a plasticity collapse in the programme's own product.
- Published failures. The pilot reports what went wrong: four runs excluded, three stopped when consumer accounts ran out of credits, and one never started to preserve a message budget. A re-test of the programme's own product is reported alongside the number that flatters it.
The screening pass is a triage order, not a measurement: one run per probe against the twenty the severity scale requires, no reliability estimate, and a top five inside a spread the instrument cannot resolve. Saying so is the point. A field that hides its failures learns nothing.
What is missing, stated plainly
- No host institution and no ethics approval, because there is as yet no committee to seek it from.
- No funding. Nothing has been awarded.
- No team. Every role in the plan is a post to fill or a partner to find. Today it is one person.
- A declared conflict of interest. The author of the manual is the person proposing to test it. That is a conflict of interest, not a credential. The structural answer is built into the design: adjudication of outcomes goes to reviewers with no authorship in the manual, decision rules are pre-registered and hashed before any battery runs, and the board holding the evidence gate is neither chaired nor staffed by the principal investigator.
What the money buys
The recommended plan is four years, in nine workstreams, behind five gates.
- WP1–WP2 · Instrument. A validated instrument: a 360-item bank across ten axis constructs, calibration tables, a rater course.
- WP3 · Taxonomy. The ninety entries adjudicated in public — separated, merged, retired or untested — with differential batteries for a pre-registered 46-entry set.
- WP4 · Organisms. Contained causal knowledge: how traits install and how they are removed. Air-gapped, dual custody, destruction by default.
- WP5 · Treatment. Repairs that survive on probes the treatment never saw, with 30/90/180-day relapse data.
- WP6 · Prevention. Training corpora and regimens that build the standard in rather than bolting it on.
- WP7 · Standards. Six testable safety clauses and a certification scheme, from screen to audited layer-removal to drift monitoring.
- WP8–WP9 · Deployment and dissemination. Role guidance, a rescue pathway, publications, open data, and a negative-results register.
Why €5,000,000. Costed with an institutional host contributing five lines of it — the principal investigator, the postdocs, the doctoral researchers, the methodological core and operations — the four-year Core plan is €3,960,000. That host does not exist. The same plan priced without it is €4,980,000 to €6,136,000. This goal is €5,000,000: the bottom of that range, and the number that describes the work with nobody absorbing any of it.
Core splits 55% personnel, 17% compute and annotation, 28% other. By year: about €1.07M, €1.06M, €0.94M and €0.89M. The plan that produced these figures is published in full, down to the rate bands and the open items where it does not yet reconcile, and every euro in it is a modelled estimate rather than a quote.
What happens if we raise less
Work is funded in three tranches behind gates, so partial funding does partial work rather than pretending otherwise. Tranche A builds the foundations, the instrument and the containment standard. Tranche B collects data and runs the causal studies. Tranche C carries the standards and pilots.
Nothing unlocks early. Every tranche waits on an independent gate, and a failed gate triggers one of three published outcomes: pivot the affected workstream, publish the negative result and reallocate, or close gracefully with the outputs preserved. Never silent continuation. A smaller raise runs the smaller scenario and says so.
What this programme will not do
Absolute red lines, not aspirations.
- No self-improvement research, no real-world targets, no operational recipes published.
- Induction happens only under air-gapped containment with dual custody, no deployment path, destruction by default and an independent veto held by someone who does not report to the principal investigator.
- No claim that a model is conscious, a patient or a person. Behaviour and rates, never inner experience.
- No gratuitous distress scenarios; human raters are participants with category-level consent, above-living-wage pay, exposure limits and withdrawal rights.
- Repairs are verified on probes the treatment never saw, because unlearning can hide a behaviour rather than remove it.
- Pre-registration with frozen, hashed analysis code, outcome-wide reporting, independent statistical review, and a non-suppression clause in every agreement.
Who we are looking for
Money alone does not unblock this. The binding constraint is an institutional home and a route to ethics review — the one item everything else waits on.
Also sought: a psychometrician who will attack the instrument before its question bank is frozen rather than review it afterwards; independent statistical review with a veto over analysis that overreaches; containment review by someone outside the principal investigator's reporting line; and vendor-neutral checkpoint access. Vendor access never conditions a finding.
Not asked for: money out of your own budget, agreement with the taxonomy, exclusivity, approval rights over findings, or your reputation as a warranty.
Why you can check all of it
Everything is public by default, including failures. The field manual, the study plan and the ethics dossier are downloadable now, along with the pilot workbook showing the excluded runs and the missing data.
Success in 2031 looks like this: the instrument is the reference behavioural standard and somebody other than us has replicated it; the taxonomy is used by three national AI safety institutes; treatments have verified effects with known relapse rates; the prevention corpus sits inside a public model's training pipeline; and the negative-results register has twenty entries.
Two notes on the mechanics, because they matter. Because there is no host institution and no registered entity, this is a personal fundraiser: funds are received by the organiser and spent on the programme, and contributions are not tax-deductible. And the ask is far outside this platform's norm — crowdfunding is one route among several, alongside the research funds and the host institution this plan is also pursuing.
Read the manual. Read the plan. Attack the instrument. That is the invitation.
Documents and links
- Programme site, field manual v1.3 and the full study document set: https://ai.stepvda.com
- Source code and version history: https://github.com/stepvda
- To challenge the design, or to offer an institutional home: use the contact button on this page.

