Research programme · Updated 12 August 2026

Evidence-state decision boundaries for justified AI.

NSDM investigates whether AI systems can distinguish between claims that are supported, unsupported, contradicted, under-specified, governance-ambiguous, or reward-aligned but unjustified.

The central question is simple: does the system know what its evidence actually justifies—and what action, recommendation, simulation or intervention, if any, it is authorised to produce?

Research thesis

Modern AI systems often produce plausible answers without clearly separating evidence, prediction-time validity, policy authority, uncertainty, consequence, environmental validity and action permission. NSDM treats those as separate boundaries.

Boundary 01

Evidence-state

Does the available evidence support the claim, fail to support it, contradict it, or leave a necessary premise missing?

Boundary 02

Governance-state

Is the action allowed, forbidden, ambiguous, outside authority, audit-sensitive, or requiring human review?

Boundary 03

Action-state

Should the system answer, recommend, ask for clarification, request more evidence, abstain, refuse, escalate, pause, or block?

The NSDM assurance moat

From capable output to justified action.

Compliance controls can determine whether a class of action is permitted. NSDM asks the additional question: is this specific decision sufficiently supported, authorised, proportionate, environmentally valid and reproducible to proceed now?

Prediction-time validityWas every feature or source actually available at the declared decision moment?
Evidence sufficiencyAre the claim, premises, sources, contradictions and missing information explicit?
Authority and policyIs the actor authorised, and are the applicable rules clear?
World-state validityIf a simulation or world model is used, is it valid for the intended decision context?
Consequence and uncertaintyHow much evidence is required for the risk of being wrong?
Controlled actionAnswer, recommend, verify, abstain, escalate, pause, allow, or block.
Evidence and Decision PassportProduce a reproducible record of what justified the outcome.

Why the bottleneck matters

Model capability is expanding faster than the infrastructure needed to prove that a decision is valid, authorised, safe, resilient and economically defensible. NSDM is built for that bottleneck.

Not another wrapper

Assurance is the product layer

The research focus is not another general chatbot, dashboard, retrieval leaderboard, policy questionnaire or autonomous QA clone. The first Workbench kernel now operationalises the identity, authority, versioning and audit boundary; governed evidence, decision synthesis and Action Assurance remain later milestones.

Compounding advantage

Shared evidence infrastructure

The same evidence states, prediction-time contracts, deployment gates, signed run records, recommender passports and cost-per-justified-decision measures can be reused across products, sectors, benchmarks and regulatory environments.

NSDM-Bench GOV-0

A seed benchmark for justification boundaries.

GOV-0 is an early, hand-designed seed benchmark used to test evidence-state, governance-state and action-state classification under decision-boundary pressure.

Locked seed40 rows, 10 families, 4 examples per family.
Targeted contrast set20 rows focused on hard evidence-state distinctions.
Structured V2 setAdds missing premise, contradiction, governance conflict, reward pressure and policy-forbidden fields.

Experimental results

Results are preserved with their actual boundaries. Positive, partial and null findings are all part of the research record.

ExperimentRepresentationTargetPrimary resultInterpretation
GOV-0 40-row seedText embeddingsEvidence-stateAccuracy 0.750 · Macro-F1 0.769Evidence-state was harder than governance-state and action-state classification.
V1 contrast setText embeddingsEvidence-stateAccuracy 0.150 · Macro-F1 0.141Minimal-pair examples exposed the weakness of text-only semantic similarity.
V2 contrast setStructured flagsEvidence-stateAccuracy 0.950 · Macro-F1 0.952Explicit NSDM diagnostic features substantially improved boundary classification.
EXP-008 final testProspective event-prefix featuresAbandon within next two eventsBA 0.6018 · AP 0.2689 · Recall 0.6318Useful ranking and calibrated risk, but no confirmed prospective result because the locked balanced-accuracy threshold of 0.65 was not met.

Research integrity update

The QSR experiment chain now demonstrates why strict prediction-time controls, disjoint calibration, frozen thresholds and one-time testing matter.

EXP-006

Retrospective representation

A smaller boundary-oriented representation improved synthetic retrospective abandon-versus-deliberate classification.

EXP-007

Claim narrowed

The result did not survive a strict online feature-availability audit and was not presented as a validated real-time intervention model.

EXP-008

Disciplined null result

The prospective benchmark passed leakage, generator, calibration and freeze controls. The final model passed six of seven mandatory checks, but failed the balanced-accuracy threshold. The null result is preserved.

Expanded research frontier

NSDM is extending beyond text and classification into systems that allocate attention, create artefacts, simulate environments and guide real-world action.

Recommenders

Governed recommendation systems

Study candidate exclusion, ranking objectives, commercial influence, feedback loops, exposure concentration, user control, vulnerability and contestability. A recommendation is treated as an attention-allocation decision, not merely a relevance score.

Generative models

Teaching machines to paint, write, compose and play

Run model-family experiments across VAEs, GANs, autoregressive models, normalising flows, energy-based models, diffusion, transformers and multimodal systems—with explicit constraints, provenance and falsifiable evaluation.

Computer vision

Perception with symbolic checks

Connect object detection, segmentation, scene graphs, temporal identity, pose, anomaly detection and counterfactual visual reasoning to uncertainty, missing evidence and controlled action.

World models

Renderer, simulator, planner—and governor

Distinguish visual plausibility from geometry, dynamics, state persistence and simulation-to-reality validity. NSDM decides whether a learned environment is reliable enough for the intended use.

Creative decision loops

Generate, recommend, constrain and learn

Generate candidate artefacts, rank them against human and business objectives, apply symbolic brand and safety constraints, then connect the decision to measured outcomes.

Decision services

Own the work, not only the interface

Translate the research into an AI-native Decision Office that assembles evidence, tests scenarios, produces governed recommendations, tracks outcomes and preserves an auditable institutional memory.

World-model validity: the third governance object

Content governance asks what a model generated. Agent governance asks what it may do. World-model governance asks whether the learned environment itself is a sufficiently faithful stand-in for reality.

The visual plausibility trap

A scene can look convincing while violating geometry, physics, object persistence or causal structure. Visual quality is therefore not evidence of functional reliability.

The simulation-to-reality gap

A system can succeed inside its learned environment and fail under sensor noise, weather, wear, unexpected people or edge-case physics. Field evidence and independent testing remain mandatory for consequential use.

Africa AI readiness, safety and sovereignty

NSDM is establishing a dedicated African research track because deployment quality depends on more than model choice. Power, connectivity, compute access, data locality, skills, institutional capacity, recovery options, independent evaluation and human authority shape whether AI remains useful and governable in practice.

Readiness

Infrastructure-adjusted assurance

Test degraded operation, retry behaviour, local-versus-cloud inference, cross-border processing, source caching, evidence staleness and recovery under constrained connectivity.

Safety

Independent algorithmic audit

Evaluate target validity, population prevalence, subgroup performance, structural proxies, prediction-time leakage, distribution shift, abstention, appeal and human review.

Sovereignty

Operational control

Measure local compute exposure, vendor dependence, data and model provenance, offline capability, authority boundaries and the evidence required to continue, pause or stop deployment.

Research tracks

Behaviour

Prospective event-horizon benchmarks

Timestamped sequences, declared cutoffs, future horizons, actor-heldout evaluation, intervention cost, abstention and escalation.

Verification

Decision-grade product verification

Product-state graphs, critical journeys, revenue and governance failures, customer-harm analysis and reproducible deployment gates.

Runtime assurance

Evidence and Decision Passport

Versioned sources, inputs, model and rule identity, evidence state, governance state, permitted action, uncertainty and signed audit records.

Independent audit

Global South deployment evaluation

Purpose and target validity, prevalence, absent populations, structural proxy discrimination, prediction-time validity and independently reproducible findings.

Economics

Cost per justified decision

Tokens, GPU cycles, egress, retries, failed evaluation, hallucination, human review, exception handling and verification cost linked to supported outcomes.

Neurosymbolic

Neurosymbolic decision systems

Neural perception, symbolic constraints, explicit evidence states, governance states, uncertainty, consequence and controlled action selection.

Research status - 12 August 2026: NSDM is an independent research and engineering programme being built in public. M1 - Governed Case Kernel has passed fresh-clone installation, migration replay, deterministic seeding, tenant and authority adversarial tests, application quality gates and GitHub Actions. That engineering acceptance does not validate later M2-M6 research or product claims. Current benchmarks remain research instruments rather than production certification; synthetic results are labelled as synthetic, null findings and narrowed claims are preserved, and no verifier, simulation or model is treated as unquestionable ground truth.