Research primers · 8 October 2026

Six concepts behind a defensible AI decision.

Short educational notes for reading our experiments critically. Each starts with a distinction, follows a hypothetical example and ends with a test. These notes are not peer-reviewed papers, clinical guidance or completed experiments.

01 · Evidence is not confidence

Confidence describes a model’s score or expressed certainty. Evidence is the material that bears on a claim: its source, availability, relevance and relationship to the premises. A persuasive answer may lack a necessary premise; a low-confidence answer may still identify a genuine contradiction.

A thought experiment

A system recommends a supplier using a historical compliance certificate. The certificate is authentic, but the current licence is missing. Familiarity with the supplier cannot fill that gap. Preserve the missing premise and request the current record.

A failure to watch

Treating semantic similarity, model confidence or source prestige as sufficient support.

Next test: Create paired cases with the same wording but different source dates, missing premises and explicit contradictions. Register the required evidence-state and permitted action before evaluation.

Related external reading →

02 · Calibration and the option to abstain

A calibrated probability estimates correctness frequency over an evaluated population. It does not establish that this particular answer is justified. Abstention introduces a decision rule: which cases should be deferred, with what review cost and residual error?

A thought experiment

Suppose a classifier assigns confidence near 0.8 to a group of cases. Compare their actual correctness frequency on held-out data. Then separately inspect whether excluded cases or changed deployment conditions invalidate the comparison. No measurements are asserted in this example.

A failure to watch

Calibrating on the test set, reporting only accepted-case accuracy, or mistaking marginal coverage for certainty about every individual.

Next test: Evaluate a fixed threshold on disjoint data. Report coverage, errors among accepted cases, rejected-case composition and human review burden. Stress-test distribution shift; conformal methods have assumptions that must be checked.

Related external reading →

03 · A faster computation is a candidate

An optimization changes an execution path. Numerical equivalence concerns specified tolerances; semantic equivalence concerns the states and decisions that matter to the task. Passing one does not automatically pass the other.

A thought experiment

Imagine an evidence score close to an action threshold. A small precision change could satisfy a global numerical tolerance yet cross that threshold. The appropriate comparison includes threshold-adjacent fixtures and the resulting action state.

A failure to watch

Accepting throughput improvement before checking semantics, or requiring identical text for every stochastic system without a task-specific criterion.

Next test: Freeze continuous-score tolerances and exact deterministic state requirements. Compare one transformation at a time. For stochastic outputs, predefine task-level distributional or behavioural acceptance criteria rather than inventing them after testing.

Related external reading →

04 · Latency, throughput and locality

Latency is elapsed time for an operation; throughput is completed work per unit time. Independent memory requests may overlap. A dependent next-address chain restricts that overlap. Data movement and operational intensity can constrain useful compute.

A thought experiment

Two loops visit the same nominal data size. One accesses predictable indices; the other visits a random order. A third obtains each next address from the previous load. Their time per access need not measure the same hardware property.

A failure to watch

Calling throughput-derived nanoseconds per access raw DRAM latency, overlooking index storage, or naming an exact cache boundary from a broad transition.

Next test: Inspect the measured CPU graph and its source. Compare matched indexed conditions for locality, then distinguish dependency effects. Record combined footprint and competing cache, translation and scheduling explanations.

Related external reading →

05 · Account for work and permission

Reconciliation tests whether a workload keeps an auditable identity through transitions. Request, semantic, resource and authority checks are candidate ACE invariants. Each requires a defined scope: logical requests differ from retry attempts, and declared capacity differs from physical allocation.

A thought experiment

A request is retried after a timeout. Two attempts may appear in the log, but the logical request needs an explicit outcome. A scheduler may declare eviction before the process stops; capacity release must reflect the chosen accounting policy and observed execution.

A failure to watch

Equating lifecycle counters without draining pending work, releasing resources twice, or treating object discovery as mutation authority.

Next test: Use unique logical and attempt IDs, explicit terminal states, idempotent allocation/release records and actor–target–operation–time authorization checks. Test delegation, revocation and delayed physical release. Policy derives effective demand; defaults are not universally additive.

Related external reading →

06 · A proxy is not the person

Behavioural measurement connects an observable signal to a defined construct. Attention, preference, comprehension, trust and purchase are different targets. A sensor or model can produce a repeatable signal without establishing the interpretation attached to it.

A thought experiment

Longer gaze at packaging could reflect interest, difficulty reading it or confusion. A single attention score cannot distinguish those explanations. Combine appropriate task measures and context rather than declaring intent.

A failure to watch

Inferring emotion or purchase certainty from a proxy, ignoring consent and population differences, or substituting synthetic personas for observed participants.

Next test: Define the construct, alternatives, consent and data-use boundaries. Choose independent behavioural measures, appropriate controls and a held-out evaluation. Report where interpretation fails. NSDM governs decision justification; neuromarketing is one application context.

Related external reading →

Keep the distinction visible.

Definitions are starting points. Useful research specifies the task, collects suitable evidence and exposes a claim to failure. Our measured graphs, proposed tests and external readings keep their separate statuses.

Existing results → · Open questions → · The NeuroAlchemist →