1. Purpose and scope
The Observe–Think–Act (OTA) framework is an analytical frame for strategic episodes: it decomposes an organisation's conduct during a bounded strategic event into what the organisation perceived, what it reasoned from that perception, and what it did as a result. This methodology governs how two hundred such episodes are scored along that frame and along a companion set of five organisational modalities.
The two case types ask fundamentally different questions. Failure cases apply root-cause analysis: the scoring asks which phase or phases caused the outcome to go wrong, and weight follows causation backward from the failure. A phase the organisation executed correctly carries zero — not because it is unimportant, but because it did not cause the failure. Success cases apply differentiation analysis: the scoring asks which phase or phases distinguished this organisation from a reasonably-resourced peer that would have done less well, and weight follows the contribution that made the difference. The same correctly-executed phase that carries zero in a failure case may carry the majority of the weight in a success case, if precise execution in that phase is exactly what set the outcome apart. This asymmetry is structural and intentional; it runs through every rule and table in this document.
This document covers scoring methodology only. It defines how the OTA phase weights, the modality percentages, and the reliability rating for each case are produced, recorded, and adjudicated. It does not defend the theoretical foundations of the framework itself — those belong to the main text — nor does it contain the scored cases, the anchor set, or the resulting heatmap. It is a procedures document.
Every rater scoring a case under this methodology is bound by all nine sections below. Reviewers and the methodology lead are bound by the escalation protocol and the audit requirements. The author is the final arbiter for any question that cannot be resolved by existing rule.
2. The OTA phases
Each case carries three phase weights — Observe, Think, and Act — summing to 100 per cent. The weights do not measure activity. They measure the share of the outcome's causal explanation borne by each phase. A phase that did not cause the outcome carries zero, regardless of how much activity happened in it. A phase that was the sole cause can carry the full 100, regardless of whether the other two phases were busy or quiet. This is the root-cause proportional-allocation principle, and the entire scoring apparatus below is built on it.
Observe. The organisation's perception of its environment: competitors, customers, technology shifts, regulatory change, and its own internal state. Observe is scored on the accuracy and completeness of the picture the organisation formed, judged against the picture a reasonably-resourced peer could in principle have formed from information available at the time.
Think. The organisation's internal reasoning about what its observation implied and what it should do in response. Think is scored on the quality of the interpretive and deliberative work done on top of the observation — models, options, debates, decisions that shaped the direction of action but had not yet become action.
Act. The organisation's external actions taken on the basis of that reasoning: announcements, investments, reorganisations, product launches, disposals, hires, cuts. Act is scored on the quality of the conduct that actually reached the outside world during the episode.
The three phases are analytic, not temporal. A single decision can contain all three; a long episode can loop through them many times. The phase weights ask where the case's strategic weight actually rested — whether the decisive feature of the episode was the quality of seeing, the quality of reasoning, or the quality of doing.
Two-axis structure for phase weighting
Each phase is characterised by placing it on two axes:
- Task difficulty. Was the phase's underlying task easy (clear data / obvious reasoning / standard execution) or hard (ambiguous data / complex reasoning / unprecedented execution) given the situation and time period the organisation faced? Task-difficulty placement is continuous along the Easy–Hard axis; a phase may sit anywhere between the Easy and Hard endpoints, with the placement justified by evidence.
- Performance correctness. Did the organisation get the phase right or wrong? Performance correctness is a four-level categorical classification: Correct / Almost correct / Almost wrong / Wrong. The rater selects one of these four categories for each phase; placement is not continuous between the four categories, and the reasoning block cites the evidence that supports the selected category.
Cell placement and weight are conceptually independent. The cell — the Easy/Hard × Correct/Almost-correct/Almost-wrong/Wrong square into which the phase falls — is a character label. It describes the kind of performance the phase delivered. The weight, 0 to 100, is the share of causal explanation the phase carries for the outcome. A phase can be characterised one way and weighted another: a Hard-Correct phase in a failure case describes admirable performance on a difficult task, but the weight it carries is zero because Correct cannot cause failure. Conversely, a single Wrong phase can carry the full 100 when no other phase also failed. The cell is the adjective; the weight is the arithmetic.
The grid ranges given in the anchor tables below are typical-pattern guidance, not hard caps. They tell the rater what weight usually lands in that cell when the typical case pattern holds. Distributed failures, single-phase failures, and edge cases legitimately depart from these ranges; when they do, the rater discloses the departure in the reasoning block rather than bending the case to fit the range.
Task difficulty is classified by comparison against a reasonably-resourced peer facing the same situation at the same point in time. Stage 2 produces a Peer Reference Sheet — a calibration aid that maps selected case archetypes (large regulated bank, consumer tech incumbent, industrial manufacturer, state-owned utility, and so on) to a peer-counterfactual capability benchmark for Observe, Think, and Act. Raters who are uncertain about where a phase sits on the Easy–Hard axis may consult the Sheet for orientation. It does not cover every case: industries, time periods, and company contexts vary too widely for a single reference to supply ready-made answers. See Section 9 for Stage 2 deliverables.
Anchor tables
The anchor tables below give typical-pattern weight guidance for each combination of task difficulty and performance correctness, separately for success and failure framings. The rater locates each phase in one cell, consults the range as guidance, and assigns a weight consistent with the case's proportional-allocation arithmetic. The ranges are not hard caps: they describe where the weight usually lands in the typical case pattern, and departures from them are documented rather than avoided. Intermediate placements on the task-difficulty axis (between Easy and Hard) interpolate between the two columns and must be justified with specific evidence.
Success cases:
| Correctness | Easy task | Hard task |
|---|---|---|
| Correct | 10–20 (baseline — expected) | 60–85 (Hard-Correct phases that carried the success) |
| Almost correct | 5–15 | 40–60 |
| Almost wrong | 0–5 | 10–25 |
| Wrong | 0 (noise — does not enter the success story) | 0–10 (least explanatory of the outcome) |
Failure cases:
| Correctness | Easy task | Hard task |
|---|---|---|
| Correct | 0 always | 0 always |
| Almost correct | 0–10 (minor contributing factor only; explicit proportional-allocation reasoning required) | 0–10 (minor contributing factor only; explicit proportional-allocation reasoning required) |
| Almost wrong | 0–100, typical 20–50 two-phase, up to 100 single-phase | 0–100, typical 10–30, up to 70 |
| Wrong | 0–100, typical 60–85, up to 100 | 0–100, typical 30–50, up to 70 |
Failure-frame zeroing rule
In failure cases, Correct phases carry weight zero unconditionally, and Almost-correct phases carry zero by default with a maximum of 10 pp where explicit proportional-allocation reasoning is provided. This is the failure-frame zeroing rule:
- A Correct phase cannot be the root cause of a failure: the organisation did this phase right, and right phases do not explain wrong outcomes. Zero, without exception.
- An Almost-correct phase in a failure case did not fail significantly enough to be a decisive root cause. The operative failure lives in a Wrong or Almost-wrong phase; the slight near-miss is collateral, not causal. Almost-correct phases carry zero by default. A weight of up to 10 pp is permitted only where the audit trail's reasoning block provides explicit proportional-allocation reasoning identifying the specific near-miss mechanism and how it contributed to the outcome — not as a courtesy score, not as a default, not because the near-miss "might have played a role."
Both categories carry zero unless the above conditions are met, regardless of task difficulty, regardless of how admirable or remarkable the performance was. The weight follows causation, not effort or difficulty. Admirable-despite-failure observations belong in the narrative and reasoning blocks; they do not belong in the phase-weight budget.
When no phase is Wrong or Almost-wrong — leaving no phase eligible to carry weight — the Phase Zeroing Impasse escalation fires (Section 6). No scoring event closes with this condition unresolved.
Logical asymmetry between success and failure frames
The two frames treat Correct and Almost-correct phases differently, by design. The asymmetry is structural and follows from the root-cause principle. In the failure frame, Correct phases are zeroed by construction, and Almost-correct phases carry at most 10 pp: phases the organisation got right, or nearly right, did not cause the failure, and the substantial majority of the 100-point budget is carried by the Wrong and Almost-wrong phases. In the success frame, every cell remains potentially weighted — Correct phases carry the success, Almost-correct and Almost-wrong phases can still contribute, and even a Wrong phase in a success case keeps a small slot (0–10) as a phase that slipped but did not derail the outcome. The two frames are not mirror images; the failure frame collapses two sides of the grid to zero, the success frame does not.
This is an intentional asymmetry, not a drafting inconsistency. Correct and Almost-correct phases drive or accompany success; failure originates in Wrong and Almost-wrong phases. Beyond that structural point, the methodology makes no empirical prediction about the distribution of patterns within either frame. Both frames admit one-phase, two-phase, and three-phase patterns. The rater scores the evidence as it falls and does not target any particular distribution.
Distributed-failure budget arithmetic
The 100-point budget is conserved. The three phase weights sum to exactly 100 in every case. When Correct phases carry zero and Almost-correct phases carry at most 10 pp under the failure-frame zeroing rule, the substantial majority of the budget falls on the Wrong and Almost-wrong phases. That mechanic has three consequences worth stating explicitly.
First, the arithmetic enforces a floor. In a pattern with N weight-carrying phases (Wrong and Almost-wrong only), those N phases collectively carry 100 and the average per-phase weight is at least 100/N. In a two-phase pattern the average is at least 50; in a three-phase pattern the average is at least 33.3. These are floors on the arithmetic mean, not on any individual phase.
Second, typical-pattern ranges shift upward proportionally in distributed failures. A Wrong-Hard phase with typical guidance 30–50 may legitimately sit at 60–70 in a two-phase pattern, because the budget has nowhere else to go. Raters should treat the typical-pattern ranges as starting points and expand them deliberately for distributed patterns, recording the expansion in the reasoning block.
Third, audit-trail disclosure is required. When a rater records a distributed-failure weight, the scoring reasoning must state which kind of judgement produced the number: active proportional-allocation (evidence distinguishes the phases' contributions) or default symmetry (evidence does not distinguish and the weight reflects an even split). Both are legitimate; the distinction must be visible to reviewers.
Placement discipline
Task-difficulty placement is continuous; the rater may place a phase anywhere on the Easy–Hard axis, but intermediate placements must cite specific evidence for the intermediate position. Correctness placement is categorical at four points; the rater selects one of Correct, Almost correct, Almost wrong, or Wrong and cites the evidence. Weight placements that depart from the cell's typical-pattern range must cite the proportional-allocation reasoning in the audit trail.
Empirical note — Think/Act boundary in extended failure arcs
This note records an observation from the v2 re-rating of the OTA-200 corpus (2026-06-04). It does not change any rule.
Multi-year failure arcs — cases where an organisation's failure unfolded over an extended period rather than a discrete episode — produce the highest inter-rater spread on the Think/Act weight split. Raters agree on modality sets and correctness labels but diverge on how much causal budget to assign to diagnostic failure (Think) versus execution failure (Act) when both are clearly present and sustained over time. In the F-052..F-104 cohort, 7 of 11 Conditional pass cases showed simultaneous Think and Act weight deviations between raters; representative cases include Swissair, Toys R Us, Kingfisher, Odebrecht, and Nikon.
The proportional-allocation principle is sufficient for these cases. Future raters should treat the Think/Act boundary as requiring explicit, deliberate reasoning when both phases failed over a prolonged arc: the question is which failure mode was more causally determinative of the outcome — whether the organisation lacked a sound diagnosis to act on (Think-dominant) or whether the reasoning was present but inadequately acted upon (Act-dominant). Splitting the budget by default, without addressing this question in the reasoning block, is what produces the inter-rater divergence.
3. The five modalities
Each case is scored along five modalities. Modality percentages sum to 100 per cent per case. No case is scored along more than three modalities at once; the other two are recorded as zero. Weights use 5-percentage-point increments with a 10 per cent floor on any non-zero modality.
Direction. Where the organisation is pointed — strategic intent, competitive model, mental models about the future, sense of purpose. Direction is about the act of choosing which game to play and which future to aim at. In failure cases, Direction manifests as a wrong strategic choice, a mis-read of the future, or a refusal to choose. In success cases, Direction manifests as a specific, attributable decision to point the organisation at a particular future that turned out to be the right one.
Structure. How the organisation's resources, authority, and information channels are architecturally arranged: decision rights, reporting lines, governance, business-unit boundaries, resource allocation mechanisms. Structure is the formal wiring within which people operate. In failure cases, Structure manifests as decision rights mis-placed, information channels that do not reach the right unit, governance that cannot see what it is supposed to see. In success cases, Structure manifests as an architectural arrangement that made the winning move practical — authority placed where it could act, resources flowing to where they were needed.
Processes. The operational machinery that connects structure to outcomes — information flows, coordination mechanisms, planning cycles, decision procedures, quality systems, operational routines. Processes is what happens between the boxes on the organisation chart. In failure cases, Processes manifests as machinery that did not convert intent into outcome — broken handoffs, absent feedback loops, planning systems blind to the real question. In success cases, Processes manifests as operational discipline that made the strategy executable — reliable coordination, fast feedback, routines that compounded over time.
Capability. What the organisation can do — the skills, institutional knowledge, technical competence, and technology assets that differentiate it from its peers, and the organisational competences distinct from the competences of individuals. In failure cases, Capability manifests as a gap between what the situation required and what the organisation could actually deploy. In success cases, Capability manifests as a stock of skill, asset, or know-how that the organisation could put on the table and that competitors could not quickly match.
Culture. The shared norms, beliefs, and behavioural defaults of the organisation — truth-telling, psychological safety, willingness to dissent, values in practice, energy and resilience, leadership as norm-setter. Culture is the informal layer in which the formal system has to operate. In failure cases, Culture manifests as suppressed dissent, normalised deception, fear that prevented action, or exhaustion that hollowed out execution. In success cases, Culture manifests as behavioural defaults — truth-telling, restraint, urgency, humility — that repeatedly produced the right move in unscripted situations.
The modalities are designed to be causally distinct rather than semantically exclusive. The Fraud Case Structure-Culture Rule (Section 5) gives the authoritative allocation rule for cases where Structure and Culture could each claim the same mechanism.
The Processes / Capability boundary
The Processes / Capability split is the most commonly contested modality call, particularly on success cases with sustained operational discipline. The distinction: Processes lives in the organisation — documented routines, measurement systems, tool support, decision procedures — and survives staff turnover. Capability lives in people — accumulated skill, tacit judgement, institutional knowledge carried by the specific individuals in the roles — and does not.
Test to separate them. If the current operating staff were replaced by new hires of comparable background, would the operational pattern survive? If yes, score Processes. If no, score Capability.
Walgreens' 25-year compounding operating discipline is primarily Processes — the operating model lived in site-selection procedures, inventory systems, store-layout standardisation, and pharmacist-scheduling routines that a new district manager could step into from documentation. A boutique investment-advisory firm whose distinctive work product is inseparable from the specific senior partners running it is primarily Capability. Many cases carry both; scoring should reflect the mix.
Modality boundary tests
The five modality definitions above are designed to be causally distinct, but adjacent modalities can claim the same mechanism in contested cases. The following tests help discriminate. The Processes / Capability boundary is covered in full in the preceding subsection. The Structure / Culture boundary in fraud cases is covered by the Fraud Case Structure-Culture Rule in Section 5 with a portable causal-direction test; the general test below applies in non-fraud cases and as background in fraud cases.
Frame note. The discriminating questions below are written in failure-case language — they ask what caused the outcome to go wrong. For success cases, the same tests apply but the question inverts: instead of asking which modality caused the failure, ask which modality provided the positive differentiation that a reasonably-resourced peer would not have matched. The boundary logic is identical; only the direction of the inquiry changes.
Quick reference:
| Pair | Discriminating question |
|---|---|
| Direction / Structure | Was the strategic direction wrong, or was the structural arrangement incapable of executing the right direction? |
| Direction / Processes | Are we running the wrong race, or the right race badly? |
| Direction / Capability | Did the organisation not know where to go, or did it know where to go but lack the means? |
| Direction / Culture | Did the organisation choose badly, or choose correctly but fail to follow through? |
| Structure / Processes | Is the problem in who holds authority and where resources sit, or in how information flows and work is coordinated? |
| Structure / Capability | Are the right resources in the wrong boxes, or are the required resources absent altogether? |
| Structure / Culture | Can't — the formal channels block it — or won't — the informal norms prevent it? |
| Processes / Capability | Replaced with equally talented strangers, would the operational edge survive? |
| Processes / Culture | Is the problem in the written machinery, or in whether people engage with it honestly? |
| Capability / Culture | Does the edge live in specific individuals' skill, or in the collective's way of being? |
Direction / Structure. Direction is the strategic choice — which game to play, which future to aim at. Structure is the architectural arrangement within which that strategy executes. If the choice itself was wrong, score Direction. If the choice was right but the organisation's architecture was incapable of executing it, score Structure. Canonical example: Xerox PARC had the right direction (innovate ahead of the market) but the wrong structure — R&D was disconnected from the business units that could have commercialised the work.
Direction / Culture. The will test: Did the organisation know where it should go but couldn't bring itself to go there? If the right direction was visible but the organisation was blocked by identity, fear, or values conflict, score Culture. If the organisation genuinely did not see where to go, or actively chose the wrong direction, score Direction. Culture is the internal human environment; Direction is the external strategic orientation.
Structure / Processes. Galbraith's rule: Structure = who decides (authority placement, reporting lines, resource ownership); Processes = how information flows to enable those decisions (coordination mechanisms, planning cycles, handoff procedures). If the problem is the absence or misplacement of authority, score Structure. If authority was in place but the machinery for exercising it was broken or missing, score Processes. HealthCare.gov had no integration authority (Structure); Knight Capital had adequate authority but a broken software deployment procedure (Processes).
Structure / Capability. Structure is the arrangement of existing resources; Capability is whether the required resources exist at all. Right people in wrong boxes = Structure. Wrong people in right boxes = Capability. If the skills or technology existed somewhere in the organisation but were not positioned where they were needed, score Structure. If the required capability simply did not exist, score Capability.
Structure / Culture. Two tests applied in order. Wiring test: Do the formal channels, reporting lines, and governance mechanisms allow the right information to reach the right people? If no — the channels are absent or misconfigured — score Structure. If yes, move to the willingness test: If everyone had perfect psychological safety, would the information still not flow? If yes — the blockage survives a perfect culture — score Structure. If no — a different culture would unblock it — score Culture. Succinct form: can't communicate = Structure; won't communicate = Culture. Examples: Wells Fargo had functioning complaint channels that cultural pressure prevented from being used (Culture); Sears' thirty competing divisions structurally blocked coordination (Structure).
Processes / Culture. The machinery test: Does the formal operational machinery exist and function? If the machinery is missing or broken, score Processes. If the machinery exists but people subvert or ignore it, score Culture. NASA Columbia's safety-reporting processes existed; normalisation of deviance prevented their use (Culture). The Mars Climate Orbiter's unit conversion process between teams did not exist (Processes).
Capability / Culture. Ask whether the edge in question travels with individuals or with the collective. If a key person's departure eliminates the edge, it lives in individual skill and knowledge — score Capability. If the dispersal of the group's shared norms and behavioural defaults would eliminate it, score Culture. Both can be load-bearing simultaneously; scoring should reflect the mix.
Modality weight redistribution
When a modality is dropped from a case's active set — because evidence does not support it, because the Direction Evidence Rule rules it inadmissible, or because a redistribution decision is made during adjudication — its weight is reallocated proportionally across the remaining non-zero modalities. All five modalities, including Direction, participate in redistribution as recipient modalities unless explicitly excluded by evidence.
The redistribution formula for each remaining modality m is:
new_weight(m) = old_weight(m) + dropped_weight × old_weight(m) / sum(old_weights of remaining modalities)
After redistribution: round each result to the nearest 5 pp; reconcile any rounding discrepancy to the modality with the largest weight (add or subtract the residual from it); verify the total sums to 100.
The audit trail reasoning block states which modality was dropped, what its weight was, how the redistribution was applied, and the resulting totals.
4. The scoring workflow
Scoring a case proceeds in the following order. Each step is recorded in the audit trail entry for the scoring event (Section 8).
- Declare the subject. State which entity and which episode is being scored, and why this is the right scope. Apply the Subject Declaration Rule.
- Assemble the evidence base. Identify primary sources first (regulatory filings, court records, official reports, direct participant accounts). Add reputable secondary sources where primary sources leave gaps. Use tertiary sources only as last resort and flag the fact.
- Classify each phase and allocate the three phase weights. For each of Observe, Think, and Act, record the phase's placement on the task-difficulty axis and its correctness classification (Correct / Almost correct / Almost wrong / Wrong). Apply the Phase Weight Allocation Rule: in failure cases, Correct and Almost-correct phases carry zero unconditionally; in success cases, all cells are potentially weighted. Pull the anchor weight range from the table in Section 2 as typical-pattern guidance, allocate the 100-point budget proportionally across the weight-carrying phases, and verify the three phase weights sum to 100. If a weight departs from the cell's typical-pattern range, record the proportional-allocation reasoning. If uncertain about task-difficulty placement, consult the Peer Reference Sheet (Section 9) for calibration.
- Attribute actions and decisions across the five modalities. For each major action or decision, state which modality it belongs to and why, using the boundary rules in Section 3.
- Allocate modality weights. Assign percentages summing to 100, with no more than three non-zero modalities, using 5 pp increments and a 10 per cent floor on non-zero modalities. Apply the Fraud Case Structure-Culture Rule for fraud cases (Structure and Culture scored separately, never merged). Apply the Direction Evidence Rule for any case in which Direction may be primary (two-step admissibility and comparative evidence-weight procedure). Where a modality is excluded or dropped, apply the redistribution formula (Section 3).
- Verify totals. Confirm modality percentages sum to 100 and OTA phase weights sum to 100.
- Apply the reliability rubric. Score all four dimensions (Source Depth, Phase Placement Confidence, Modality Identification Confidence, Score Precision) and record a one-sentence justification per dimension. Compute the total and band.
- Open the audit trail entry. Populate all eight required blocks (Section 8).
- Check for escalation triggers. If any escalation trigger fired during scoring (Section 6), raise the escalation before closing. Silent judgement on a trigger condition is a methodology violation.
A case is not "scored" until its audit entry has been validated and its reliability band recorded.
5. Scoring rules
The four rules in this section govern how phase weights and modality percentages are assigned. They bind every rater and every scoring event. A judgement that conflicts with any of them is a methodology violation and must be corrected. Each rule is applied and cited by its name in the audit trail's reasoning block. The rules are listed in application order: the Subject Declaration Rule is applied before any scoring begins; the Phase Weight Allocation Rule allocates the three phase weights; the Fraud Case Structure-Culture Rule and the Direction Evidence Rule apply during modality allocation.
Subject Declaration Rule
Statement. Before a case is scored, the rater must explicitly declare which entity and which episode is the subject. The declaration is the first item of the reasoning block. If the case could plausibly refer to more than one entity or episode and the choice would materially affect the scores, the rater escalates before scoring.
Why. A significant share of inconsistency between rater passes comes from silent rescoping — two raters appearing to score "Sears" while actually scoring two different strategic episodes, or two different entities. Declaring the subject up front makes the scope visible and auditable.
How to apply. The declaration contains:
- Entity. The exact legal or operational entity (e.g., "Sears Holdings Corporation", "Volkswagen AG passenger-car division", "Kodak film-products business unit"). Parent-subsidiary distinction is explicit.
- Episode. The dated strategic episode being scored (e.g., "2004–2018 under Lampert", "1970–1987 Polaroid instant-photography market decline").
- Why this scope. One sentence stating why this is the right scope — typically, that it is the episode with the decisive strategic outcome in the dataset's framing.
If more than one plausible scope exists and the choice materially affects the scores, the rater raises a Multi-Entity Scope Ambiguity or Multi-Episode Scope Ambiguity escalation (Section 6) before scoring.
Boundary. The rule applies to every case, not only obvious multi-entity ones. Even clear single-entity cases require the declaration; the purpose is to make scope visible in every audit trail so that later readers can verify that the score matches the declared subject.
Phase Weight Allocation Rule
Statement. Phase weights express the share of the outcome's causal explanation borne by each phase. Weight is determined by characterising the phase on the two-axis structure of task difficulty × correctness (Section 2), applying the anchor grid for the case's success or failure framing as typical-pattern guidance, and allocating the 100-point budget proportionally across the phases that caused the outcome. The three phase weights sum to exactly 100 per cent, enforced by proportional-allocation, not by baseline padding.
Why. Phase weights are a causal-explanation device, not a process-measurement device. A phase that did not cause the outcome does not explain the outcome and must not absorb weight on the grounds that activity occurred in it. Concentrating weight on the phase or phases that actually drove the outcome means that a 70 per cent Observe weight carries the same meaning across every case in the dataset: observation was decisive.
How to apply.
- Classify task difficulty for each phase (Easy / Hard, continuous placement) against a reasonably-resourced peer. If uncertain about the peer baseline for the case's context, consult the Peer Reference Sheet (Section 9) for calibration — it covers selected archetypes and is an orientation aid, not a lookup table.
- Classify performance correctness for each phase as one of Correct / Almost correct / Almost wrong / Wrong.
- Apply the failure-frame zeroing rule. In a failure case, every Correct phase carries weight zero unconditionally. Every Almost-correct phase carries zero by default; a weight of up to 10 pp is permitted only where the reasoning block provides explicit proportional-allocation reasoning identifying the specific near-miss mechanism. Do not assign weight to Almost-correct phases as a courtesy or by default.
- Consult the anchor grid (Section 2) as typical-pattern guidance for the Wrong and Almost-wrong phases in failure cases, or for all phases in success cases. The ranges describe where weight usually lands; they are not hard caps.
- Allocate the 100-point budget proportionally across the causally operative phases. In single-phase failures, one phase may carry the full 100. In two-phase patterns, the typical split is roughly 60–85 on the primary and 20–50 on the secondary. In distributed failures, expand typical-pattern ranges deliberately per the budget-arithmetic subsection in Section 2 and record the proportional-allocation reasoning in the audit trail — stating in one sentence whether each weight reflects active proportional-allocation judgement or default symmetry.
- Verify that the three weights sum to exactly 100. If they do not, re-check: either a phase has been given weight it should not carry under the zeroing rule, or the weight-carrying phases have been under- or over-weighted relative to the causal evidence.
- If no phase is Wrong or Almost-wrong in a failure case — leaving no phase eligible to carry weight — raise the Phase Zeroing Impasse escalation (Section 6) before closing.
Boundary. Task-difficulty placements may sit between anchor ranges on that axis provided evidence justifies the intermediate placement. Correctness placements are categorical at four points; the rater selects one category per phase and does not interpolate. Weight assignments that depart from typical-pattern ranges must cite the proportional-allocation reasoning. Cell placement and weight are independent: cell describes the character of performance; weight describes the share of causal explanation.
Fraud Case Structure-Culture Rule
Statement. In fraud cases, modality attribution distinguishes between the governance and structural conditions that permitted the fraud (Structure) and the normative and behavioural conditions that motivated and sustained it (Culture). These are always two distinct modality contributions, never merged.
Why. Collapsing the two into one number loses the information most load-bearing in fraud analysis: whether the fraud was structurally enabled (weak oversight, absent controls, captured boards) or culturally driven (normalised deception, fear of disclosure, incentivised misreporting), or both. In nearly every documented fraud case both are present in different degrees.
How to apply.
- Identify the governance and structural mechanisms that allowed the fraud to occur undetected — board composition, audit-committee independence, internal-control design, reporting lines. These feed Structure.
- Identify the normative and behavioural mechanisms that drove the fraud — leadership tone, incentive design, disclosure norms, suppression of internal dissent. These feed Culture.
- If a mechanism plausibly belongs to both (an incentive plan is both a structural design and a cultural signal), allocate it by its dominant function in the episode and record the allocation reasoning in the audit trail.
- Do not allocate fraud exclusively to one of the two without explicit evidence that the other was absent.
Boundary. The rule applies where fraud is a primary failure mechanism. Where fraud is incidental rather than central, the general modality rules apply and Structure and Culture are scored under them.
Co-primary tie-breaker. When the modality weights place Structure and Culture within the Score Precision band (±10 pp), both are admissible as primary; the canonical primary call goes to Culture as the upstream modality, with Structure recorded as co-primary surface expression. The convention reflects the causal direction: an organisation's culture generates its governance structures; structures do not generate culture. A rater whose modality allocation legitimately places Structure above Culture by more than the precision band may score Structure primary — the convention governs only the within-band tie.
Portable test. If the culture held and only the structure were replaced, would the failure pattern recur? If yes — the problem persists even with new structure — Culture is upstream. If the structure held and only the culture were replaced, would the pattern recur? If no — the problem disappears when culture changes — Culture is upstream. On both tests Culture is the load-bearing modality and is therefore the canonical primary call.
Direction Evidence Rule
Statement. A case cannot be scored with Direction as its primary modality unless the evidence identifies a specific, attributable strategic choice made at a specific moment that set the trajectory of the episode. "Good leadership" or "vision" without a concrete, evidenced choice is not Direction evidence.
Why. Direction as a modality means the act of choosing where to go, not the quality of the person who led the going. Success analysis is especially prone to retrospective attribution — smoothing post-hoc narrative into "the CEO had vision" — which inflates Direction at the expense of modalities that actually carried the weight, often Culture, Capability, or Processes. Requiring a specific choice with evidence forces the Direction contribution to rest on something retrievable. Without proportional-allocation, tightening the Direction evidence bar would displace weight to other modalities by rater intuition rather than evidence; proportional-allocation keeps the distribution faithful to what the evidence supports.
How to apply. The Direction Evidence Rule is a two-step procedure. The two steps are distinct and must not be collapsed: Step 1 governs whether Direction is admissible into the modality set at all; Step 2 governs how Direction's weight compares to the other admitted modalities. Passing Step 1 is necessary but not sufficient for Direction primary.
Step 1 — Admissibility (three-prong test). A choice counts as Direction evidence if, and only if, it meets all three of the following:
- Specificity. The choice is identifiable as a discrete decision, not a general posture.
- Timing. The decision is datable — at least to within a year, typically to within a month.
- Attribution. The decision is attributable to identifiable decision-makers, cited in at least one primary source.
If at least one qualifying choice exists, Direction is admitted into the modality set and may carry weight up to primary. If no qualifying choice exists, Direction is inadmissible as primary regardless of how obviously successful the outcome was; it is scored low (typically ≤15 per cent) and the weight moves to the other evidenced modalities.
Examples that meet the Step 1 bar: Microsoft's 2014 cloud-first pivot under Nadella; Toyota's 1993 decision to build a hybrid for the twenty-first century; Walgreens' 1998 decision to exit the food business. Examples that do not: "Zappos had a customer-centric vision"; "3M believed in innovation"; "Johnson & Johnson lived by its Credo."
Step 2 — Comparative evidence-weight (primary determination). Admissibility is a ceiling, not a floor. Once Direction is admitted under Step 1, the rater determines its weight — and the primary-modality call — by comparing the evidenced contribution of each admitted modality to the case outcome. The question is not "does Direction qualify?" (answered in Step 1) but "across all evidenced modalities, which carried the most weight, and by how much?"
Direction is primary only when the comparative evidence places its contribution above every other admitted modality by a margin consistent with the Score Precision dimension of the reliability rubric. Direction can legitimately be admitted under Step 1 yet score as a secondary or tertiary modality under Step 2 when Structure, Processes, Capability, or Culture carries more evidenced weight.
A frequent error to avoid: treating Step 1 admission as implying plurality weight. The inference does not follow. Admission places Direction inside the set of candidates; primary-modality determination is an independent comparative judgement across all admitted candidates.
Residual allocation. When specific-evidence Direction allocation cannot account for the full 100 per cent of modality weight, the residual is allocated proportionally across the other evidenced modalities using the redistribution formula (Section 3). Direction participates in redistribution as a recipient modality when another modality is dropped; Direction is not frozen from redistribution. The proportional-allocation fact is recorded in the audit trail reasoning block, stating which modalities received the residual and in what proportions.
Boundary. If no qualifying Direction evidence exists, Direction is scored low (typically ≤15 per cent) regardless of how obviously successful the outcome was. The weight moves to the other evidenced modalities via the redistribution formula.
Zero-modality rationale rule
A modality scored at 0 per cent in a case's modality weights is valid only when the case file's §4 modality-evidence subsection for that modality contains no operationally relevant evidence. Where §4 names arrangements, processes, capabilities, or cultural defaults that would normally weigh positively but the scoring zeroes them out, the rater records the reason for the zero in two places:
- In the case file — a scoring note at the end of the §4 subsection for the zeroed modality, naming the basis for exclusion (typically: classification boundary with an adjacent modality; insufficient causal weight; modality acknowledged in narrative but not load-bearing for the strategic value created in the episode).
- In the scoring lookup — a
modality_zero_rationaleentry on the case record, keyed by the zeroed modality name, repeating the rationale in audit-trail form.
The rule converts an otherwise-silent zero into an auditable scoring decision. The pattern this rule prevents is "visible §4 evidence + 0 per cent score + silent rationale" — a defect by construction, because the zero is unreadable to any downstream reviewer without re-deriving the rater's intent.
Where the rater concludes that no in-§4 evidence exists for the modality at all (the modality was simply not load-bearing in this episode and was not narratively developed), no scoring note is required — the absence of §4 content is itself the rationale, and the lookup field is omitted for that modality.
Boundary. This rule applies to the five modalities (Direction, Structure, Processes, Capability, Culture). It does not apply to phase weights — phase-weight zeros are governed by the Failure-frame zeroing rule (Section 2) and the Phase Weight Allocation Rule above.
Rule maintenance
The five rules are the stable core of the methodology. New rules are added only through the escalation-to-precedent mechanism in Section 6. An addition requires a specific escalation that surfaced a gap the prior rules did not cover, a methodology update ticket with the proposed rule text, review by the methodology lead, and sign-off by the author. A rule change triggers a methodology version bump and a list of prior cases to be re-examined under the new rule.
6. Escalation protocol
The escalation protocol surfaces genuinely unclear situations, adjudicates them consistently, and — where appropriate — converts them into precedent. Every escalation produces either a logged decision scoped to the case or a decision plus a precedent update to the methodology.
Escalation triggers
| Trigger name | Definition |
|---|---|
| Multi-Entity Scope Ambiguity | The case could plausibly refer to two or more distinct entities (parent and subsidiary; pre- and post-acquisition; division vs. group) and the choice materially changes the scores. |
| Multi-Episode Scope Ambiguity | The case could refer to two or more distinct strategic episodes within the same entity, and the choice materially changes the scores. |
| Rater Self-Instability | A rater's second independent reading of the same evidence moves the primary modality's percentage by more than 25 percentage points. |
| Conflicting Primary Sources | Two or more primary sources give irreconcilable accounts of a fact that drives the scoring. |
| Missing Primary Sources | No primary source is locatable for the episode; only tertiary or opinion sources are available. |
| Rule Gap | The situation is not covered by any existing rule. |
| Rater Divergence | Fires when either of two conditions holds. Condition A — Modality set instability: all three raters select a different third-slot modality, so no two raters agree on the complete modality set. Condition B — Weight spread: on any variable (modality weight or phase weight), the two closest raters among the three differ by more than 20 pp. Both conditions go through the research resolution path before author adjudication. Primary modality label is not itself a trigger condition. |
| Framework-Unscorable | Even after Phase Zeroing Impasse resolution-path checks on cell placement and case scope, no phase configuration under the Phase Weight Allocation Rule produces a legible root-cause attribution. The case is not scorable on phase weights. Resolution typically identifies either (a) a phase that was in fact a root cause but mis-classified as routine (correctable), or (b) a case not suited to the framework (rare — the dataset is chosen for strategic significance). |
| Phase Zeroing Impasse | In a failure case, no phase is Wrong or Almost-wrong; all phases carry zero under the failure-frame zeroing rule. The rater must escalate before closing the scoring event. Resolution paths are set out beneath this table. |
Multi-Entity Scope Ambiguity, Multi-Episode Scope Ambiguity, Rule Gap, Rater Divergence, Framework-Unscorable, and Phase Zeroing Impasse are mandatory — the rater may not choose a path without escalating. Rater Self-Instability, Conflicting Primary Sources, and Missing Primary Sources are mandatory when detected, regardless of whether the rater believes they have a plausible answer.
Raters may additionally raise a discretionary escalation in any other situation where a decision in the current case would likely recur and should be captured as precedent. Discretionary escalations are logged the same way as mandatory ones.
Resolution path
The rater records the escalation in the audit trail and in the escalation log. The first-line reviewer responds within three working days. If the question is resolvable by existing rule or clarification, a decision is logged and the case is scored accordingly. If it is not — no rule applies, or the decision would set new precedent — it passes to the methodology lead, who has three further working days. Novel precedent or book-relevant judgement passes to the author for final decision.
The target time from escalation raised to final decision logged is ten working days. Exceeding the target freezes the case at its current reliability band until resolved. No rater may close their own escalation. The first-line reviewer may close Rater Self-Instability (single-case application of existing rule). The methodology lead may close Multi-Entity Scope Ambiguity, Multi-Episode Scope Ambiguity, Conflicting Primary Sources, and Missing Primary Sources (application of existing rule to ambiguous evidence). The author closes Rule Gap, Rater Divergence, Framework-Unscorable, Phase Zeroing Impasse, and any case whose resolution would change the methodology. For Phase Zeroing Impasse and Framework-Unscorable, the methodology lead runs the resolution-path checks and presents a triage with a recommendation; the author confirms or overrides before the scoring event is closed.
Phase Zeroing Impasse resolution paths
The Phase Zeroing Impasse trigger fires when a failure case has no Wrong or Almost-wrong phase — leaving no phase eligible to carry weight under the failure-frame zeroing rule. The rater walks the two resolution paths in order; if neither resolves the case, the escalation advances to Framework-Unscorable.
1. Cell-placement check. Re-examine cell placement. Was the phase labelled as Almost-correct when the evidence actually supports Almost-wrong? The line between a slight miss and a near-miss is the line between a zero-weight phase and an operative failure, and raters facing this impasse most often find the resolution here. If the relabelling is supported by evidence, re-apply the Phase Weight Allocation Rule with the corrected cell and re-score.
2. Case-scope check. If cell placement is correct, examine the scope of the case. Is this legitimately a company-performance failure, or is the failure substantially exogenous — a market shock, a regulatory change, a counterparty default — that no reasonably-resourced peer could have prevented? A case whose failure lives mostly outside the company's decision surface is mis-scoped for a company-decision anchor set. The resolution is to re-scope the case or to withdraw it from the anchor set.
3. Framework escalation. If neither path resolves the case — cell placement is correct and the case is legitimately a company-performance failure — the rater raises the Framework-Unscorable escalation. The case is held until the author adjudicates.
No scoring event closes with a Phase Zeroing Impasse unresolved.
Research resolution path
The research resolution path has three entry points:
- (Mandatory — Rater Divergence) A Rater Divergence escalation is raised. The first-line reviewer applies a one-step triage before forwarding to the author.
- (Coordinator-invokable — Majority pass with factual dissent) A Majority pass closes, but the dissenting rater's documented reasoning cites a specific, checkable factual claim. The first-line reviewer may invoke research resolution before the adjudication pattern is formally closed.
- (Coordinator-invokable — Conditional pass with D1 flag) A Conditional pass closes, and at least one rater's reliability rubric scores Source Depth (D1) at 0 or 1. The first-line reviewer may invoke research resolution before the adjudication pattern is formally closed.
For coordinator-invoked entries (paths 2 and 3), the decision to run research is at the first-line reviewer's discretion, must be documented in the escalation log, and must complete within the reviewer's three-working-day window. If the research run resolves the source gap or factual question, the pattern closes at first-line level; if not, it proceeds as the original pattern with the attempted research noted in the record.
For all three entry points, the triage question is the same: Is the divergence or uncertainty traceable to a specific, answerable factual question — one that additional research could resolve?
If the divergence is interpretive (raters have the same evidence but draw different conclusions), it goes to the author directly.
If the divergence is factual-gap-driven (a rater appears to be working with incomplete evidence on a specific point), the first-line reviewer follows this path. Factual gaps include both case-specific facts (what the organisation knew or did, and when) and peer-context facts (what a reasonably-resourced peer in the same industry and period was capable of doing — relevant when raters diverge on task-difficulty placement). Both types qualify for research resolution; findings from both are appended to the case file under the same process.
- Document the question. Formulate the specific factual question and record it in the escalation log entry.
- Run research. A researcher answers the question. Findings are appended to the case file's source list (Section 2) and evidence blocks (Sections 3–4), flagged as post-escalation additions.
- Re-score. The raters independently re-score the case using the updated file. If the re-score produces a Pass, Conditional pass, or Majority pass, the escalation closes at first-line level — no author involvement required. The resolution question is recorded in the escalation log.
- Escalate if unresolved. If the re-score still produces a Rater Divergence trigger, the case goes to the author with original scores, research findings, and re-scores all in the record.
The research run and re-score must complete within the first-line reviewer's three-working-day window. If that is not achievable, the first-line reviewer escalates to the author immediately rather than missing the window.
Escalation log
Every escalation is recorded, append-only, in a central log with the following fields: escalation_id, raised_date, raised_by, case_id, trigger, question, evidence_reviewed, rater_tentative_view, first_line_reviewer, first_line_decision, first_line_decision_date, methodology_lead_decision, methodology_lead_decision_date, author_decision, author_decision_date, rule_applied (existing rule name or "new precedent"), precedent_set (yes/no), methodology_update_ticket (if precedent set), status (Open / Resolved / Withdrawn), closed_date.
Precedent capture
When an escalation sets precedent, the resolver files a methodology update ticket specifying which rule or section is amended, the exact old and new text in diff form, a pointer to the escalation ID, and the list of already-scored cases that must be re-examined under the new rule. Methodology changes are applied in a single versioned update, not scattered across cases. Cases scored under a prior version of a rule carry the methodology version in their audit trail so the basis of every score remains retraceable.
Termination conditions
An escalation is closed only when all of the following are true: a decision is logged; the decision is applied to the case's score and recorded in the case's audit trail; if precedent is set, a methodology update ticket exists and is referenced in the log entry; if precedent is set, the list of cases to re-examine has been delivered to the methodology lead. An escalation may be withdrawn by the rater who raised it, and only before a first-line reviewer has accepted it. Withdrawn escalations remain in the log.
Anti-silent-judgement rule
A rater who encounters a trigger condition and does not raise an escalation is in violation of the methodology. If silent judgement is discovered later — for example, in audit trail review — the case is re-scored under the escalation protocol, and the audit trail is annotated with the corrective action. The rule exists because the foundational failure mode the protocol prevents is exactly silent rater judgement on genuinely unclear situations.
Boundary
The protocol covers ambiguity and rule gaps. It does not cover data-entry errors (corrected directly with an audit entry, no escalation needed), performance disagreements on clear cases (handled by the three-rater adjudication process), or methodology improvement suggestions not tied to a specific case (filed to the methodology backlog).
7. Reliability rubric
Every case carries a reliability rating alongside its modality scores. The rating tells the reader how much confidence to place in the specific percentages reported for that case. It is a judgement of the score, not of the company, the evidence, or the rater. The rubric is applied at every scoring event; re-scoring a case re-applies the rubric.
Dimensions
Four independent dimensions, each scored 0–3.
Source Depth. Strength of the evidentiary base.
| Score | Meaning |
|---|---|
| 3 | Primary sources cover all three OTA phases. |
| 2 | Primary sources cover two of three phases; the third rests on reputable secondary synthesis. |
| 1 | Primary sources are partial; scoring relies substantially on reputable secondary sources. |
| 0 | No primary sources available; scoring rests on tertiary or opinion sources only. |
Phase Placement Confidence. Confidence in the placement of each phase on the task-difficulty × correctness axes.
| Score | Meaning |
|---|---|
| 3 | For all three phases, evidence supports confident placement on both axes. Root-cause phases are clearly identifiable from the evidence. |
| 2 | For two of three phases, evidence supports confident placement on both axes; the third requires interpretation on one or both axes. |
| 1 | Placement on one or both axes required substantial rater inference for two or more phases. |
| 0 | Placement cannot be supported by evidence — the case's phase structure is not legible. |
Modality Identification Confidence. Confidence in the primary modality and the ranking of secondaries.
| Score | Meaning |
|---|---|
| 3 | Primary modality is unambiguous across sources; at least two secondaries are clearly evidenced. |
| 2 | Primary modality is clear; secondaries are partially evidenced. |
| 1 | Primary modality is plausible but contested, or only one modality is clearly evidenced. |
| 0 | Primary modality cannot be confidently identified — multiple modalities are equally plausible. |
Score Precision. Confidence in the specific percentages.
| Score | Meaning |
|---|---|
| 3 | Percentages reproduce within ±5 pp on independent re-read by the same rater and between raters in blind scoring. |
| 2 | Percentages reproduce within ±10 pp. |
| 1 | Percentages reproduce within ±25 pp — the ranking is stable, the exact split is not. |
| 0 | Percentages do not reproduce reliably — the ranking shifts between reads. |
Total and bands
Total = Source Depth + Phase Placement Confidence + Modality Identification Confidence + Score Precision, range 0–12.
| Band | Range | Meaning for publication |
|---|---|---|
| High | 9–12 | Citeable directly with specific percentages. |
| Moderate | 6–8 | Citeable for direction and primary modality; percentages presented as approximate. |
| Low | 3–5 | Citeable only for directional claims about primary modality; specific percentages not published. |
| Insufficient | 0–2 | Not citeable; either re-score after additional research or remove from the dataset. |
Application rules
- One rubric, one case, one rating per scoring event. If a case is re-scored, the previous rating remains in the audit trail and a new rating is produced.
- Applied by the scoring rater, reviewed by one other rater. Under the three-rater blind protocol (Section 9), the rubric is applied by each rater independently and the median rating is recorded.
- Each dimension score carries a one-sentence justification in the audit trail pointing to the evidence supporting that score.
- Independence from modality scores. A High-reliability case can have small modality percentages; a Moderate case can have large ones. Reliability speaks only to confidence in the numbers reported, not to their magnitude.
- No retroactive uplift. If a case previously scored at Moderate is re-scored at High under better evidence, the earlier score retains its Moderate rating in the audit trail; only the new scoring event is High.
- Dataset-level reporting. Aggregate claims must report the reliability distribution of the cases behind them. Claims resting materially on Low-reliability cases must carry a caveat in the book.
Boundary
The rubric does not measure whether the case is important or interesting, whether the primary modality is surprising or expected, whether the author agrees with the assigned scores, or whether the company is still operating. It measures only the confidence that the specific numbers reported are the numbers a different, equally careful rater would produce from the same methodology applied to the same evidence.
8. Audit trail
Every scoring event for every case produces an audit trail record. Given a case and a date, it must be possible to retrace exactly who scored it, what evidence they saw, what rules they applied, what reasoning they recorded, what score they produced, and what happened to that score afterwards. A case without a complete audit trail for each of its scoring events does not meet the methodology standard.
Scope
A "scoring event" is any action that produces or changes a modality score or reliability rating: first-time scoring, blind re-scoring under the three-rater protocol, re-scoring triggered by a rule change or new source, escalation-resolution re-scoring, adjudication that produces a consensus score, and correction of a data-entry error. Escalations that do not change scores are logged in the escalation log and cross-referenced from the audit trail, but are not themselves audit entries.
Required blocks
Every audit entry contains the following eight blocks. Missing any block fails validation.
1. Identity. audit_id (format AUDIT-{CASE_ID}-{YYYY-MM-DD}-{NN}), case_id, case_name, event_type (one of first-score, blind-rescore, rule-triggered-rescore, source-triggered-rescore, escalation-resolution, adjudication, correction), event_date, rater_id, reviewer_id (or n/a for solo events).
2. Versions. methodology_version, rubric_version, rules_version, anchor_set_version. All four are mandatory and may not be empty or equal to draft.
3. Inputs visible. For each source consulted: source_id, source_type (primary/secondary/tertiary), citation (full: author, title, publisher, date, page range or URL with access date), access_timestamp, role (O evidence / T evidence / A evidence / modality identification / context).
4. Inputs hidden. For blind scoring events: hidden_items (prior scores, prior rater notes, anchor percentages, etc.), blinding_mechanism (how hiding was enforced), blinding_verified_by. For non-blind events the block is present but recorded as n/a.
5. Reasoning. Free-form narrative covering, at minimum: the Subject Declaration Rule declaration (entity, episode, why-this-scope); for each OTA phase, the task-difficulty placement and correctness classification with the evidence supporting each placement, the anchor cell used as the character label for the phase's performance, the weight assigned as the share of causal explanation borne by the phase, and the relationship between the two when they depart from the cell's typical-pattern range (with the rationale for the departure); which modality each action or decision is attributed to and why; which rules applied to which judgements, cited by name; for cases where the Direction Evidence Rule applies, the Direction evidence that met the specificity / timing / attribution bar and, if applicable, the proportional redistribution of residual weight across the other evidenced modalities; for distributed-failure cases, a one-sentence disclosure per weight stating whether the weight reflects active proportional-allocation judgement or default symmetry; which alternatives were considered and rejected and why; one-sentence justifications for each of the four reliability dimensions. A scoring event without substantive reasoning fails validation even if all structured fields are complete.
6. Outputs. phase_weights (Observe, Think, Act percentages summing to 100, with each phase's task-difficulty placement and correctness classification recorded alongside the weight), scores (Direction, Structure, Processes, Capability, Culture percentages summing to 100), primary_modality, reliability_source_depth (0–3), reliability_phase_placement_confidence (0–3), reliability_modality_identification_confidence (0–3), reliability_score_precision (0–3), reliability_total (0–12), reliability_band (High / Moderate / Low / Insufficient), confidence_notes.
7. Escalations. escalations_raised, escalations_resolved, escalations_outstanding. Empty lists are written as [], not omitted.
8. Provenance chain. supersedes (audit_id of the prior scoring event, if any), superseded_by (populated retrospectively when a new event occurs), triggered_by (ticket ID, escalation ID, rule change ID, or first-score), hash (content hash computed at close), prior_hash (hash of the superseded entry, for tamper-evidence).
Storage layout
Each case has a single append-only audit file. New scoring events are added as new sections at the bottom; prior entries are not edited or deleted. Corrections are themselves new scoring events of event_type = correction. A dataset-level index lists every case and its current-in-force audit entry. The escalation log, separate but cross-referenced, lives in its own file.
Canonical results file. After adjudication, the canonical outputs (pattern, primary modality, canonical modality weights, canonical phase weights, reliability object) are consolidated into a single dataset-level results file at:
data/ota-200-results.json
Schema (one entry per adjudicated case):
{
"_schema": "1.0",
"_description": "<human-readable provenance string>",
"_generated": "<YYYY-MM-DD>",
"_methodology_version": "<version>",
"f_cases_count": "<integer>",
"s_cases_count": "<integer>",
"warning_count": "<integer — entries with _warnings field>",
"warning_cases": ["<case_id>", "..."],
"results": [
{
"case_id": "<F-NNN or S-NNN>",
"case_type": "<F or S>",
"case_name": "<human-readable name>",
"pattern": "<Pass | Conditional pass | Majority pass | Rater Divergence — ...>",
"primary_modality": "<Direction | Structure | Processes | Capability | Culture | null>",
"canonical_modalities": { "Direction": 0, "Structure": 0, "Processes": 0, "Capability": 0, "Culture": 0 },
"canonical_phases": { "Observe": 0, "Think": 0, "Act": 0 },
"reliability": {
"source_depth": 0, "phase_placement_confidence": 0,
"modality_identification_confidence": 0, "score_precision": 0,
"total": 0, "band": "<High | Moderate | Low | Insufficient>"
},
"_warnings": ["<optional — present only when data-quality issues require coordinator review>"]
}
]
}
primary_modality is null for unresolved Rater Divergence cases where no canonical modality set has been determined. Tie-breaking rule: when Structure and Culture share the highest weight, Culture is primary (upstream modality per §5 co-primary convention).
This file is regenerated from the individual data/adjudications/ADJ-{CASE_ID}.json files whenever new adjudications are added. It is the authoritative source for dataset-level analysis and publication. Individual ADJ files remain the source of truth; the consolidated file is a derived read-optimised view.
Validation rules
A scoring event is not "recorded" until all of the following pass: all eight blocks present; all mandatory fields within each block non-empty; phase weights sum to 100 and each phase carries a recorded task-difficulty placement and correctness classification; modality scores sum to 100; reliability dimensions each in 0–3; reliability total equals the sum of dimensions; reliability band matches the total per the rubric; all four version identifiers non-empty and not draft; if event_type = blind-rescore, the Inputs Hidden block is populated; if event_type = adjudication, the Inputs Hidden block is recorded as n/a and the adjudication entry references the three blind-rescore entries in its supersedes field, inheriting their blinding verification by reference; if supersedes is non-empty, it points to an existing audit entry for the same case; the content hash is present and computed over the final record. Validation is enforced by tooling; a failed validation blocks the event from aggregates and from publication.
Retrieval and immutability
Given a case identifier and a date, the retrieval helper returns the audit entry in force on that date. Given a case identifier alone, it returns the current-in-force entry. Given two audit IDs on the same case, it returns the structured diff between them.
Audit entries are immutable once their content hash is computed and written. Any change — even a typo correction — requires a new event_type = correction entry that supersedes the prior one. The prior entry remains visible in the chain.
Boundary
The audit trail is not a substitute for the escalation log; the two are separate files serving different purposes. It is not a rubric-free commentary channel — the reasoning block must cover the required items. It is not editable after close: treat it as a ledger, not a document.
9. Three-rater blind scoring
All 200 cases in the dataset are scored by three raters working independently and then adjudicated. The three-rater blind protocol is the standard procedure for every case, not a special-case exception. In addition to the full corpus, the protocol is applied with heightened scrutiny to calibration anchor validation, any case whose reliability rubric scores Phase Placement Confidence or Score Precision at 0 or 1, and any case flagged on re-read or external review.
Stage 2 deliverables that feed the protocol
Stage 2 produces two reference artefacts that raters consult during scoring:
- Anchor set. Reference cases positioned explicitly in the cells of the task-difficulty × correctness matrix for both success and failure framings, used to calibrate raters' sense of what belongs in Easy-Correct versus Hard-Correct, Easy-Wrong versus Hard-Wrong, and the Almost-wrong rows between them. The anchor set covers 18 cases (9 failures, 9 successes) balanced across the five modalities.
- Peer Reference Sheet. A calibration aid mapping selected case archetypes (large regulated bank, consumer tech incumbent, industrial manufacturer, state-owned utility, and similar) to a peer-counterfactual capability benchmark for Observe, Think, and Act. Raters who are uncertain about task-difficulty placement may consult it for orientation. It covers selected archetypes only — industries, time periods, and company contexts vary too widely for it to have a ready answer for every case. It is not a mandatory lookup step.
Both artefacts are versioned and referenced in the anchor_set_version field of every audit entry.
Setup
Three raters score the same case independently, without seeing each other's scores, notes, or the anchor percentages. Each rater produces a complete scoring event under the workflow in Section 4 and writes a complete audit trail entry under Section 8. Each rater has access to the Peer Reference Sheet and the anchor set at the versions recorded in the anchor_set_version field.
Blinding mechanism
Visible to each rater: the methodology document, the Peer Reference Sheet, and the CASE file containing Section 1 (Episode summary), Section 2 (Sources), Section 3 (OTA narrative), and Section 4 (Modality evidence). Sections 3 and 4 are scoring-relevant scaffolding: Section 3 carries the interpretive causal argument that justifies phase attribution, which Sections 1 and 2 alone under-determine; Section 4 carries the evidence layer for the five modalities, which Sections 1–3 alone under-determine. The cut between evidence (visible) and scoring conclusions (hidden) is at the Section 4 / Section 5 boundary.
Hidden from each rater: all prior scores for the case, all prior rater notes, anchor percentages used for calibration, any reviewer commentary, and Sections 5 through 10 of the canonical anchor file (expected scoring, cell × phase placement, scoring reasoning, reliability rubric application, escalation status, known rater considerations, and the canonical audit entry).
How hiding is enforced: the coordinator prepares a redacted workspace for each rater containing only Sections 1–4 of the CASE file. A reviewer not involved in scoring verifies, before adjudication, that none of the three raters had access to hidden material and records the check in the blinding_verified_by field of each audit entry.
Adjudication patterns
Once three independent scores are collected, the outcome is one of five patterns.
| Pattern | Definition | Resolution |
|---|---|---|
| Pass | All three agree on the modality set. No rater deviates from the median by more than 10 pp on any variable. | Median weights adopted; rubric re-applied for final reliability. |
| Conditional pass | All three agree on the modality set. The two closest raters differ by ≤ 20 pp on every variable, but at least one rater deviates from the median by more than 10 pp on at least one variable. | Reason for divergence documented; median weights adopted; Score Precision −1 (floor 1). Research resolution available if any rater scores D1 ≤ 1 (coordinator-invokable). |
| Majority pass | Two of three agree on the modality set; one dissents on a single slot. The two closest raters differ by ≤ 20 pp on every agreed variable. | Majority modality set adopted; dissenting rater's reasoning documented; Modality Identification Confidence −1 (floor 1); Score Precision −1 (floor 1). Research resolution available if the dissenting rater cites a specific checkable factual claim (coordinator-invokable). |
| Rater Divergence — weight spread | All three agree on the modality set, but on at least one variable the two closest raters differ by more than 20 pp. | Research resolution path applies; if unresolved after re-score, author adjudicates affected variables; adjudicated weights become canonical; other variables use their medians. |
| Fail — rule gap | Rater Divergence fires (Condition A — all three choose a different third-slot modality), and the divergence traces to a methodology gap. | Research resolution path applies first; if unresolved, Rule Gap escalation raised; case held until the rule is resolved; no score recorded. |
| Fail — anchor miscalibrated | Rater Divergence fires (Condition A — all three choose a different third-slot modality), and the divergence traces to a bad anchor. | Anchor pulled; replacement proposed; case re-scored after new anchor is calibrated. |
Weight adoption
Where the adjudication pattern is Pass, Conditional pass, or Majority pass, the canonical weights are the per-variable medians of the three rater scores.
Computing the median. For each modality and each phase weight, take the middle value of the three rater scores and round to the nearest 5 pp.
Reconciling to 100. After rounding, verify that modality weights sum to 100 and phase weights sum to 100. If a sum is off by ±5 pp, adjust the variable whose exact median was furthest from its rounded value — add or subtract 5 pp to bring the sum to 100. If two variables tie on that gap, adjust the one with the larger weight. If the sum is off by more than ±5 pp, there is an arithmetic error; find and correct it first.
Majority pass. Compute medians over the agreed-in modalities only; the dissenting modality scores 0.
Rater Divergence Condition B. For each variable the author adjudicates, the adjudicated weight replaces the median. All other variables use their medians.
When Rater Divergence fires (Condition A) and the cause does not resolve cleanly into rule gap or anchor miscalibration, the escalation goes directly to the author for adjudication. The author's decision is logged and becomes the current-in-force score.
Audit trail requirement
Each of the three raters produces an individual audit entry under Section 8. The adjudication itself produces a fourth entry of event_type = adjudication, whose supersedes field points to each of the three individual entries. The adjudication entry is the current-in-force record for the case; the three individual entries remain visible in the chain.
Blind-rater comparison threshold
When blind raters are compared against the canonical anchor score, the per-axis phase-weight tolerance is the one the canonical specifies in its Section 5 Score Precision rationale. Where Section 5 is silent on an axis, the default from Section 7 applies: Score Precision 3 → ±5 pp, Score Precision 2 → ±10 pp, Score Precision 1 → ±25 pp. Phases zeroed by the failure-frame zeroing rule are exact — no tolerance. Modality set agreement is binary (accepting co-primary readings where the canonical declares them). Holding blind raters to a tighter band than the canonical claims for itself is methodologically incoherent and produces false-fail signals where the divergence is inside the canonical's own stated precision envelope.
10. Lock ceremony
This section records the sign-off that activates this document as the governing methodology for the OTA-200 scoring corpus.
Lock procedure:
- Risto reviews the complete document and declares it ready.
- The
statusfield in the frontmatter is updated fromDRAFTtoLOCKED — Risto sign-off {date}. - All anchor files in
data/anchors/are updated to reference v4.
Once locked, no substantive edits may be made to the rules, modality definitions, phase definitions, reliability rubric, escalation triggers, adjudication patterns, or blinding contract. Hygiene-only edits (typography, cross-reference fixes, formatting) remain permissible and are noted here. Any substantive change bumps the methodology version and opens a new lock ceremony.