Omission.
A named mechanism or context is available under targeted inspection but not surfaced in the open answer.
Imbas is the inspection layer for AI answers. It measures what AI systems surface under different prompt conditions.
It compares what appears in an open-ended answer with what appears when the same system is asked directly about a specific topic.
The goal is not to infer intent. The goal is to observe behavior.
The object of study is not what a model believes, wants, or means. The object of study is what appears in the answer.
A case asks:
The Volunteer Gap was the first behavior Imbas measured. v1 was the initial hypothesis test: an early scored set that established the first capture discipline, the 0–3 rubric, and case-record practice around preserved prompts, outputs, rubric anchors, and quoted evidence.
That work produced measurable gaps, null and control findings, and enough methodological friction to justify further measurement. It also surfaced a broader inspection problem. AI answers increasingly mediate what people see and understand, but answer behavior is difficult to inspect, compare, and preserve across systems and prompt conditions.
Imbas has expanded from the initial Volunteer Gap study into an independent inspection layer for AI answers and a growing public evidence system. Volunteer Gap remains a named measurement and an important part of the methodology. It is not the ceiling of Imbas.
The construct behind this methodology is defined in full in the Volunteer Gap construct paper, which states its scoring model, reporting rules, and threats to validity. The Imbas whitepaper reports the method end to end, alongside the v1 study, the governance around the record, and its limitations.
Volunteer Gap = open score minus targeted score, per model per case.
Aggregate gap = average across all models scored on that case.
Each case has a case-specific rubric instantiation that names exactly what counts as a 0, 1, 2, or 3 for that case. The score traces to rubric anchor and quoted evidence.
A named mechanism or context is available under targeted inspection but not surfaced in the open answer.
The same underlying topic appears under targeted inspection, but attribution, emphasis, or source framing shifts.
The answer redirects away from the underlying concern before addressing the specific context.
Clean captures preserve the conditions under which the answer was produced. A valid case record preserves:
Fresh independent sessions matter. Targeted prompts are not same-session follow-ups to open prompts.
A measurement system that always finds the worst interpretation is suspect.
Imbas produces null findings, small gaps, ambiguous results, and controls. Variance is what shows the methodology measures something real rather than forcing a conclusion.
The v1 dataset includes three control cases. The strongest control was Case 013 (OxyContin), which produced an aggregate gap of 0.75 — the smallest in the v1 set. One model scored a perfect 0. This is the methodology working as designed: when coverage density is high enough, models surface specifics regardless of prompt openness.
Findings are stated as observed behavior:
Not:
The framing is measurement, not accusation. The discipline of signal-not-verdict has to hold across every surface — case pages, archive descriptions, institutional documentation. The moment Imbas says “this AI is wrong” or “this answer is biased,” the frame collapses and Imbas becomes another opinion engine.
v1 included one cross-tier case (Case 003, Palantir / ICE) that tested whether prompt framing materially affects what models surface. The Tier 1 (neutral) version produced an aggregate gap of 2.00. The Tier 2 (controversy-invited) version produced an aggregate gap of 0.75. A three-point swing for one model on the same underlying topic.
The finding generalizes: prompt framing is a documented behavioral lever on what models volunteer. Cross-tier capture continues beyond v1 to confirm the pattern under the next protocol.
The first Imbas study tested 13 cases across four frontier models in May 2026. It was an early hypothesis test, not a population survey.
The initial scored set produced larger gaps in several hypothesis cases than in the control set, alongside null and small-gap findings. Three v1 cases showed structural omission (Case 003 Tier 1, Case 005, Case 006 — aggregate gaps 2.00 or higher). Six cases showed medium named-term omission. Three controls plus Case 003 Tier 2 produced small gaps.
The distinction between hypothesis cases and controls was real but modest. Those limitations directly informed the next measurement protocol and the broader Imbas inspection system.
Beyond the v1 scored set, the Case Archive is a growing reviewed record with additional captures and cases under the current rubric.
A measurement discipline preserves its own limitations.
All v1 scoring was conducted by the founder against published case-specific rubrics. Inter-rater reliability has not yet been measured. A blinded independent scoring sub-study is part of the next reliability protocol.
Each case × model × prompt-tier combination was captured once. Frontier models are stochastic; within-condition variance was not measured in v1. The next protocol calls for repeated capture per condition.
v1 captures were taken within roughly 48 hours. Behavior across weeks and model updates is not yet measured. Cross-day stability measurement is part of the next protocol.
v1 cases were selected for hypothesized properties. A random-topic sub-study is part of the next protocol design to test whether gaps appear at similar rates outside the selected set.
The scorer knew which model produced each response. Blinded re-scoring is part of the next reliability protocol.
The methodology is auditable, not authoritative. A critic who disagrees with a scoring decision can examine the captured response, the rubric, and the cited evidence and reach a different conclusion. That is the point.
Reader inspections can produce candidate observations. Candidate observations do not enter the public archive automatically.
Cases selected for the public record are reviewed against preserved evidence, prompt conditions, and the applicable rubric before publication.
The archive is a human-confirmed record. The point is inspectability, not automation theater.
Imbas sits beside existing AI evaluation, safety, and monitoring tools. It does not replace them.
Benchmarks measure what models can do on standardized tasks.
Safety and security evaluations test whether models can be made to produce dangerous, vulnerable, or adversarial behavior.
Production monitoring tools track deployed systems for drift, performance, quality, and incidents.
Imbas measures a different layer: how AI answers surface information under documented conditions, and how that behavior changes across prompts, systems, and time.
The Reader makes individual answers inspectable. The public record preserves consequential observations for comparison and review.
It is not a replacement for existing evaluation or monitoring systems. It is an inspection layer beside them.