Guide

Developing Good Scenarios

How to build a set of futures that changes what your reader believes - and keep it honest afterward.

Analysis foundations

13 minute read

Hinsley can generate a complete scenario set - a small number of distinct, plausible futures, each with a cited narrative, resolution criteria, and a likelihood estimate - in minutes, grounded directly in an analysis's decomposition and source material. That likelihood is then reassessed automatically as new evidence arrives, and scored for accuracy once the outcome is known. Speed alone does not make a scenario set useful, however, and this guide explains the method behind Hinsley's Scenario Builder: what to configure before generating a set, what Hinsley does automatically once generation begins, and what to check to confirm the result holds up under scrutiny.

Before building a scenario set, the research question itself should be bounded, forward-leaning, and neutrally framed. If it is not, that should be addressed first, since the steps described below depend on it (Authoring Good Analysis Questions covers that step).

The correct approach at this stage is to construct several plausible outcomes rather than work toward a single expected answer. This guide explains how to do that in Hinsley.

Two traditions inform this method. It originated at RAND in the 1950s, where Herman Kahn built structured narratives of nuclear war to require planners to consider outcomes they preferred not to address - "thinking about the unthinkable" - and introduced the word scenario into planning with the 1967 publication of The Year 2000. From RAND, the method developed along two paths. Within government, it became structured analytic tradecraft, the discipline behind techniques such as Alternative Futures Analysis and the standards that govern how intelligence analysts express uncertainty. In the private sector, Pierre Wack's team at Royal Dutch/Shell adapted it in the early 1970s, followed by Peter Schwartz and the Global Business Network, Kees van der Heijden, and more recently the Oxford Scenario Planning Approach.

The two traditions have since diverged and disagree on several points. Hinsley's Scenario Builder incorporates both, and this guide addresses where they agree, where they differ, and how to reconcile the two in practice.

What scenarios are for

A common misconception treats scenarios as forecasts, hedged by producing several of them. This is not accurate. Wack, describing how Shell's scenarios functioned, was explicit about their purpose: "we no longer saw our task as producing a documented view of the future business environment five or ten years ahead. Our real target was the microcosms of our decision makers." The scenarios functioned as an instrument; the intended change was in the decision-maker's understanding. "Unless the corporate microcosm changes," he wrote, "managerial behavior will not change."

This gives scenarios three practical purposes, each with an associated test.

Make the assumed future explicit. Every organization operates, implicitly, on one assumed future - the future that budgets, hiring plans, and roadmaps quietly reflect. In Shell's tradition this is called the official future, and it is rarely documented. Making it explicit, as one option among several, is often the single most valuable outcome of scenario work. The relevant test: if none of the scenarios resembles what leadership currently believes, that assumption has not been captured. If all of them do, the set has not moved past it.

Widen the range of futures under active consideration. This is not intended to encourage creativity for its own sake, but reflects that the range of futures an organization actively considers is typically narrower than the range that is plausible. A relevant test is whether at least one scenario describes a future that would currently be dismissed in a planning discussion.

Establish a shared vocabulary for what is being monitored. Once several futures have names, linked drivers, and defined resolution criteria, a disagreement about whether conditions are worsening becomes a disagreement about which named scenario the evidence supports. This is a more productive basis for discussion, and it is what Hinsley's likelihood tracking is designed to support.

A scenario set that fulfills none of these three purposes provides limited analytical value, regardless of the quality of the writing.

What a scenario set is

In Hinsley, a scenario is one possible future outcome for an analysis topic. It has a title, a short description, a longer AI-generated narrative explaining how it would come about, resolution criteria describing what would indicate it had happened, and a likelihood.

A scenario set is a named group of scenarios inside a single analysis - the futures to be considered together, as alternatives to one another. Every analysis starts with one set, marked primary, and additional sets can be added to explore different framings of the same question. The primary set is the one that feeds into the resulting written output.

This distinction is significant. Most of what determines whether scenario work is effective or ineffective occurs at the level of the set, not the individual scenario, which is addressed in the next section.

What makes a scenario set good

Hinsley's generation process is designed to produce a set with these qualities by default. The following is what to confirm when reviewing a generated set, or when authoring scenarios by hand: a set of four well-written, individually plausible narratives can still be a poor scenario set, so the set should be evaluated as a whole before its individual scenarios are evaluated.

  • Different in kind, not in degree. Paul Schoemaker's formulation is a useful standard: scenarios "should describe generically different futures rather than variations on one theme." High growth, medium growth, and low growth represent one scenario with a variable adjusted, not three distinct scenarios. Any two scenarios in a set should be tested for whether they imply different decisions, rather than merely different figures.
  • Plausible, not merely possible. Each scenario requires a traceable path from the present to that outcome. Plausibility is a demanding standard: it requires that the narrative could be followed step by step without requiring an actor to behave inconsistently with established incentives.
  • Balanced, with no predetermined outcome. If one scenario is presented as the evident answer and the others exist to provide the appearance of balance, the set contains strawman scenarios. Readers typically identify this quickly, which reduces the credibility of the set as a whole.
  • Consistent in level and time horizon. A scenario concerning global regulatory realignment should not be presented alongside a scenario concerning a single supplier's quarterly announcement.
  • Complete enough for the decision. The standard is not "covers all possible futures," which is unachievable, but "covers the futures that would change the resulting course of action."
  • Distinctly named. Scenario names carry into subsequent organizational discussion. A scenario named "Scenario B" is unlikely to be referenced later; a name such as "Hard Stop" is more likely to persist in later conversation.

Number of scenarios. Established guidance from scenario planning practice recommends not fewer than two and not more than four, which is why Hinsley offers a choice of 2, 3, or 4. Four is the natural result of a two-axis framing and is typically the appropriate default. Two is appropriate when the decision presents a clear fork. Three carries a specific risk: readers tend to treat the middle scenario as the expected outcome and the other two as bounds around it, which reduces the exercise to a single prediction with additional narrative framing. If three scenarios are used, none should be positioned as the moderate case.

Probability and plausibility

This disagreement becomes apparent as soon as the likelihood estimates on a scenario are examined, and is worth addressing directly.

The dominant school of corporate scenario planning - Wack and Schwartz, van der Heijden, and the Oxford approach of Rafael Ramírez and Angela Wilkinson - holds that probabilities should not be assigned to scenarios. The reasoning is substantive: a genuinely novel future is a one-off event with no relevant historical base rate, so a probability assigned to it implies a precision that does not exist. It also invites the outcome the method is designed to prevent: a reader identifies the highest number and treats it as the forecast.

The intelligence tradition adopted the opposite approach and standardized it. Both the UK's PHIA probability yardstick and the US intelligence community's equivalent lexicon under ICD 203 divide the probability scale into seven named bands, so that "likely" carries the same meaning for every reader and every analyst. Hinsley's seven-point likelihood scale - almost no chance, very unlikely, unlikely, roughly even chance, likely, very likely, almost certain - follows that convention directly. ICD 203 also requires analysts to explain the uncertainty behind a judgment and to incorporate analysis of alternatives, which corresponds closely to scenario work. Within scenario planning itself, Stephen Millett and Michel Godet have argued the opposing case: assigning a probability requires analysts to articulate their reasoning more precisely and surfaces assumptions that narrative alone can leave unstated.

Both positions are correct, in that they address different aspects of the same problem.

The scenario itself remains narrative. This is what the decision-maker reads, evaluates, and plans against - the causal account, not the figure. The probability is not the output of this exercise; it functions as an instrument. It serves two purposes that narrative alone does not: it allows a change in belief to be detected before it would otherwise be noticed, and it allows performance to be assessed once the outcome is known. Calibration demonstrates that the method used to update the scenarios is sound. It is not the primary content the reader engages with.

Two principles follow from this:

A probability should not be presented without its supporting narrative. A briefing slide that states "Scenario C: Likely (65%)" without an accompanying causal account carries the objections raised against assigning probabilities, without the corresponding benefit.

The change in probability over time is more informative than its absolute level. The absolute probability assigned to a novel future is inherently uncertain. The change in that probability, and the evidence responsible for it, provides more reliable information.

Building a scenario set in Hinsley

Step 1: Frame the scenarios

The framing decision has the greatest influence on the quality of the resulting set, and it is made before scenario generation begins.

Framing starts with a decomposition. A decomposition breaks the analysis topic into drivers - the major factors that will shape how it plays out - and indicators, the observable signals that show where a driver is heading.

When scenarios are generated, Hinsley evaluates the decomposition's drivers directly and selects the ones that are both most important to the research question, and most uncertain, building the scenario axes from those. This selection is automatic and does not require the drivers to be sorted manually beforehand. Two important-and-uncertain drivers produce a two-by-two grid and four scenarios.

Four configuration choices follow in Hinsley's Scenario Builder:

  • Number of scenarios (2, 3, or 4). See the guidance above; the choice of three warrants particular care.
  • Mutually exclusive and/or exhaustive. This selection is not merely descriptive - it changes how the AI generates and calibrates the set. Mutually exclusive and exhaustive (MECE) is appropriate when the situation resolves to a single outcome, when probabilities should sum to roughly 100%, and when the set will later be resolved and scored. Non-MECE is appropriate when exploring overlapping futures that could co-occur, or when enforcing exclusivity would distort the analysis to fit the frame. The number of axes follows the number of scenarios selected: MECE at four scenarios produces the two-axis grid described above - two independent critical uncertainties, two states each, one scenario per cell. MECE at two or three scenarios is instead built along a single critical uncertainty, with that many distinct states rather than two crossed axes.
  • Replace or append. Append when extending an existing frame; replace when changing it.
  • What to ground the generation in. Generation can draw on decomposition drivers, specific source documents, free-text instructions, and optionally the web search that is being conducted more broadly in association with the research question. The free-text field is where the two axes can be specified directly, for cases where the framing has already been decided rather than left to Hinsley to identify - when specifying axes this way, they should be the two most important and uncertain drivers, not simply the two most interesting ones, since selecting on interest alone typically produces a set that does not inform a decision. Using search results should be enabled when the situation is developing on an ongoing basis (almost always,) and left disabled when scenarios should be grounded strictly in the sources you're curating manually.

Scenarios are not required to originate from AI generation. They can also be created manually - title and description, with optional narrative and resolution criteria - and manually authored and generated scenarios can coexist within the same set. The scenarios can be reordered by drag and drop and edited inline at any time.

Step 2: Generate the set

Once you've asked Hinsley to create the scenarios, generation runs as a four-step workflow: draft the scenarios (titles, descriptions, cited narratives, and the driver axes used to differentiate them), validate, generate resolution criteria for the set, generate likelihoods for the set.

For non-MECE sets, one scenario is consistently anchored on current trajectories persisting - the closest equivalent to the official future, named as an option rather than left as an unstated default - with the remaining scenarios built as structurally distinct alternatives rather than intensity variants of the same outcome. For MECE sets there is no forced continuity scenario; one appears only if it genuinely corresponds to one of the grid cells or states, which is the case with Quiet Continuity in the worked example below. Every scenario is also drafted to be distinguishable from the others by an observable end-state condition, consistent with the standard applied to resolution criteria in Step 4, so that falsifiability is established at the point of generation rather than added afterward.

For MECE sets, the validation step performs a check that warrants close attention. It examines pairwise overlap between scenarios, identifies coverage gaps, and flags inconsistent levels of abstraction and time horizons - the set-level failures described above - and where these are found, it revises the scenarios to address them. Selecting the status badge opens the Relationship Validation Report: pairwise overlap severity, coverage gaps, a plain-language summary, and the full history across every round.

The report should be reviewed before the scenarios themselves, since it indicates whether the set is structurally sound. The scenarios can then be reviewed to assess their individual quality.

If the status badge shows Issues remain, the healing process has exhausted its attempts without resolving the structural problem. This is typically a signal about the framing rather than the writing - commonly, two axes that are not actually independent, or a question that does not decompose as assumed. In this case, revisiting Step 1 is more effective than revising the scenario text.

Step 3: Build a second set

A scenario set addresses the question "what could happen, given a particular framing," it does not indicate whether the frame itself was, exclusively, the only relevant framing. A significant critique of current scenario practice is that a two-axis matrix too often can be presented as the only viable alternatives, when there may be other plausible scenario sets based on different framing to consider.

In Hinsley, this is straightforward to address. A second set can be built on deliberately different axes and compared with the first. If both sets produce recognizably similar futures under different names, the original framing is supported. If they produce substantially different outcomes, this indicates that the choice of axes was itself a significant analytic judgment, which should be stated explicitly in any analysis you produce. Cloning a set copies its scenarios, links, and current likelihoods into an independent new set, which is the appropriate method for changing one variable without affecting the original. Any set can be promoted to primary, renamed, or, if not primary, deleted.

The chat interface provides an additional method for exploring beyond an established frame. A conversational assistant can develop scenarios in discussion and add or replace them directly, with proposals presented as cards that are explicitly accepted or rejected. This is most useful when a gap in the set is apparent but its framing is not yet clear. Some example prompts we often see analysts leverage:

  • "Generate two additional scenarios that would be extremely impactful but extremely low probability." This corresponds to the intelligence community's high-impact / low-probability technique. The purpose is not necessarily to add these scenarios to the set, but to determine whether the existing scenarios have excluded outcomes with material negative consequences.
  • "What would have to be true for Scenario C to become the most likely outcome?" If the answer involves a chain of conditions considered implausible, this supports the original estimate. If some of those conditions are already partially satisfied, this indicates relevant evidence that had not been accounted for.
  • "What future does this set leave no room for?"

Step 4: Confirm the resolution criteria

Resolution criteria specify what would count as a given scenario having occurred, and are what distinguish a scenario set from a set of narrative essays. Hinsley generates them automatically for every scenario as part of generation (Step 2). A completeness indicator flags any scenario without one - most commonly a scenario that was created by hand rather than generated.

What to check: resolution criteria should meet the standard the decision analyst Ronald Howard described as the clarity test, generally illustrated through a clairvoyant. Could an all-knowing observer determine whether this scenario occurred, without requiring further clarification? "Regional supply chains deteriorate significantly" does not meet this standard. "At least one of the three named suppliers publicly announces a capacity reduction or force-majeure allocation on the affected node before 30 June, and our landed unit cost rises more than 25% year over year" does.

If criteria are missing or need revision - for a hand-authored scenario, or because the generated version does not meet the clarity test - they can be written per scenario or regenerated for the entire set at once. Bulk regeneration applies only if a result is produced for every scenario in the set. Regenerating criteria also offers to regenerate likelihoods at the same time, since a revised definition of the outcome changes the associated probability.

The other thing worth confirming is diagnosticity. An indicator that would hold under every scenario in the set provides no information about which one is occurring. Each criterion should be checked against the other scenarios: if it would also be true under another scenario, it does not distinguish between them. "Component prices rise" is typically a weak indicator. "Supplier A files for a specific license exemption" is typically a stronger one.

The scenario narratives address the same requirement from a different angle, and are generated the same way. Every AI-generated scenario includes a multi-paragraph cited narrative explaining the causal path to that outcome and its downstream effects, with numbered citations to a Sources panel. Manually created scenarios can have one generated on request, grounded in the analysis's most relevant source documents. These narratives are editable, and reviewing them is worthwhile, since the causal path is where independent judgment contributes the most.

Step 5: Review the driver and forecasting question links

Scenarios that are not connected to the rest of the analysis are a common way in which an otherwise sound set loses analytical value over time. Hinsley proposes both of the links that prevent this automatically: a single pass over the analysis suggests driver and forecasting-question links for every scenario at once, running after scenario generation and after decomposition changes.

Decomposition driver links carry an influence weight - Weak, Minor, Moderate, Strong, or Decisive. This makes explicit why a scenario is more or less likely, and allows the affected scenarios to be identified when a driver's assessment changes.

Forecasting question links are the mechanism that makes a scenario responsive to evidence: AI and human forecasts on narrow, resolvable questions inform belief in the broader, unresolvable scenario.

These suggestions should be treated as a first draft rather than a final answer. Manually created links are always preserved; only AI-generated links are replaced when this runs again. The influence weights in particular are worth reviewing directly, since they are a judgment call Hinsley makes on limited information.

Step 6: Estimate likelihood

Every scenario displays three estimates side by side, and the difference between them is informative.

Hinsley's estimate includes a written rationale and a seven-day trend indicator. It is produced by a structured reasoning process - base rates, evidence summary, arguments for and against, temporal analysis, an assumptions check, and a final probability - with every step preserved and available for review by selecting "Why?" For mutually exclusive and exhaustive sets, the model is instructed to make probabilities sum to roughly 100%.

The user's estimate is an independent rating, updatable at any time. It should be recorded before Hinsley's rationale is reviewed. A substantial gap between the two is one of the more informative signals available: it indicates either that the model has identified evidence not otherwise considered, or that the user is applying a judgment the model does not have access to. Either case is worth resolving explicitly rather than averaging without further review.

The aggregate is calculated automatically once two or more individuals have submitted a rating. Submission access can be restricted to collaborators only, extended to anyone in the account, or made available via a public link. Forecasters can also be invited by email with a personalized pre-filled link, or the set can be opened to guest forecasters without a Hinsley account through a public link.

Regarding who to invite: for aggregation purposes, breadth of perspective is generally more valuable than depth of credential. Several individuals with the same training and the same sources will tend to produce a consistent view that is not necessarily more accurate than any one of their individual estimates. Participants who would approach the problem differently should be included.

Step 7: Red-team the set

A scenario set is a structured artifact, and structured artifacts are subject to characteristic weaknesses. Red teaming can be applied directly to a scenario set, through a choice of lens:

Lens What it looks for
General A broad first pass when you don't yet know where the weakness is
Contrarian Consensus positions that lack strong evidence, historical analogies that don't actually apply, or minority viewpoints that deserve more weight
Cognitive bias detection Anchoring, confirmation bias, availability bias, or groupthink shaping the read of the situation
Assumption surfacing Unstated assumptions about actor motivations, causal relationships between drivers, or continuity of current conditions
Evidence gaps Claims without supporting sources, over-reliance on a single source, or likelihood assessments that aren't well-grounded in evidence
Structural completeness Missing scenarios or drivers, drivers that overlap instead of being independent, or gaps in the causal chain from drivers to scenarios

Hinsley's analysis produces findings - discrete, trackable issues that a human then moves through the states open, addressed, accepted risk, or dismissed. That resolution step is as significant as the session itself. A red-team session in which every finding is dismissed has not served its intended analytical function. A set with several findings marked "accepted risk" reflects a transparent account of the set's known limitations.

This step should be completed before publication rather than after. Conducting it once a set has already been presented changes its function from an independent analytic check to a defense of conclusions already reached.

Step 8: Keep the set current

This step distinguishes scenario work that remains useful over time from scenario work that is not revisited after its initial creation, and it is the area in which Hinsley differs most from how scenario planning has traditionally been practiced since its origin at RAND.

For most of the method's history, a scenario set functioned as a one-time deliverable: a team produced it, presented it, and the document was not revisited. It was not systematically updated as circumstances changed, and its quality was not subsequently assessed, which is part of why the field has not developed consistent standards for scenario quality, and why the question of whether scenario planning is effective has remained unresolved for decades. This approach lacked a feedback loop. Hinsley is designed to provide one.

Monitoring. Automated weekly research monitoring reassesses every scenario's likelihood against new evidence and documents what changed, reflected in trend indicators and a periodic email digest. This process runs automatically regardless of user activity; the corresponding responsibility is reviewing its output. The stated explanation is central to this feature - a probability change without a documented reason provides limited analytical value.

History. Every rating change by every estimator is retained and never overwritten. A trend chart plots the aggregate, Hinsley's estimate, and each individual forecaster over time. This provides a record of when a belief changed and why, which organizations otherwise find difficult to reconstruct.

Resolution and scoring. Once the question is settled, the set can be resolved: the scenario that occurred is selected, with an optional record of how this was determined. Finalizing a resolution automatically scores every forecaster - human and AI alike - for accuracy across the full forecasting period. This is the basis for an evidentiary track record, and is what allows a claim of calibration to be substantiated rather than asserted. A resolution can be reversed if finalized prematurely, though doing so clears the accuracy scores derived from it.

Beyond this point, scenarios appear as nodes on the Analysis Lifecycle Map with likelihood rank, trend, coverage, and citation count; the primary set feeds the combined analysis write-up and is referenced by generated Action Memos and One-Pagers; any scenario or set can be embedded as a block inside another document; and a set can be published to an external stakeholder Panel for a defined audience to review and respond to. (Outputs and publishing covers the delivery side.)

A worked example

Continuing the question from the companion guide:

"How are recent semiconductor export controls likely to reshape the viability and pricing of our top three Asian suppliers over the next 12-18 months?"

This set is configured mutually exclusive and exhaustive at four scenarios, so Hinsley builds it as the two-axis grid described in Step 1: two independent critical uncertainties, two states each, one scenario per cell.

How Hinsley ranks the drivers. Demand for the end product matters but is relatively predictable over 18 months, so it is treated as a predetermined element - shared background across all four scenarios. Two drivers are both important and genuinely uncertain, so Hinsley selects them as the axes:

  • Enforcement intensity - strict and expanding, versus porous and carve-out heavy
  • Supplier adaptation speed - rapid retooling and rerouting, versus slow and capital constrained

The four scenarios:

Rapid adaptation Slow adaptation
Strict enforcement Costly Rerouting - supply continues through new intermediaries; landed costs rise sharply; margin, not availability, is the problem Hard Stop - one or more suppliers exit the affected node; allocation and spot pricing; availability is the problem
Porous enforcement Quiet Continuity - workarounds absorb the shock; minimal observable disruption Drift and Consolidation - no acute shock, but capability concentrates among suppliers who can carry compliance overhead

Quiet Continuity is likely to represent the official future - the outcome the organization is implicitly planning for. Naming it as one option among four, rather than treating it as an unstated baseline, is a significant benefit of this exercise.

These four scenarios differ in kind, because each implies a different decision. Costly Rerouting is a pricing and contract problem. Hard Stop is a qualification-of-alternate-suppliers problem that would need to begin now. Drift and Consolidation is a supplier-strategy problem on an 18-month timeline. Quiet Continuity implies no immediate action. A high / medium / low set would not meet this standard.

A resolution criterion for Hard Stop: "Before 30 June, at least one of Suppliers A, B, or C publicly announces capacity reduction, exit, or force-majeure allocation on the affected node, and our contracted volume for that node is filled below 80% in any single month."

A diagnosticity check: "component spot prices rise" would hold true under three of these four scenarios, which makes it a weak indicator. "Supplier B applies for a named license exemption" distinguishes the strict-enforcement column from the porous one, which makes it a stronger indicator.

Likelihoods, for a set configured as MECE, are generated to sum to roughly 100%: Costly Rerouting 30%, Hard Stop 15%, Quiet Continuity 25%, Drift and Consolidation 30%. These figures function as an instrument for tracking belief over time. The analytical product is the four named scenarios and the course of action each would imply.

Common failure modes (and the fix)

Failure mode What it looks like Fix
Variations on a theme High / medium / low growth; fast / slow / no change Check whether each scenario implies a different decision. If none does, confirm the decomposition has genuinely divergent important-and-uncertain drivers, or re-specify the axes in the free-text field, and regenerate.
The middle becomes the plan Three scenarios, and everyone treats the moderate one as the forecast Use two or four. If you must use three, make sure none is positioned as the moderate case.
Strawmen One scenario is obviously right; the others exist for balance Red-team with the Contrarian lens. If no serious case can be made for a scenario, it isn't plausible enough to be in the set.
Mixed levels A global-realignment scenario sitting beside a single-quarter announcement On MECE sets, the validation step normally catches and revises this automatically. If it persists after healing, revisit the framing in Step 1 rather than editing the scenario text directly.
Unfalsifiable criteria "Conditions deteriorate meaningfully" Check the generated criteria against the clairvoyant standard. If they fall short, regenerate for the set or revise directly, naming actors, thresholds, and dates.
Non-diagnostic indicators Criteria that would be satisfied in three of four scenarios Before accepting the generated criteria, test each one against every other scenario.
The official future in disguise Every scenario is a mood variant of current expectations Ask Hinsley, via chat or the free-text field, to generate a scenario that would embarrass current strategy, then check whether it is genuinely implausible or merely unwelcome.
Probability theater The number gets briefed; the narrative doesn't Never strip out the rationale Hinsley generates for each estimate when briefing or publishing a likelihood.
Set and forget The set is built, briefed, and never revisited The weekly reassessment runs on its own. Read it, watch the trend chart, and resolve the set when the world settles it.
Frame lock-in One set, one framing, never questioned Build a second set on different axes and compare.

Before you publish

The set, rather than its individual scenarios, should be evaluated against the following:

  • Kind, not degree. Does each scenario imply a different decision?
  • Official future. Is the organization's current implicit assumption represented in the set, named as one option among several?
  • Discomfort. Would at least one of these scenarios be dismissed without serious consideration in a meeting today? If not, the range may be too narrow.
  • No predetermined outcome. Could a reasonable skeptic argue seriously for each one?
  • Level and horizon. Do all scenarios share the same level of abstraction and the same time horizon?
  • Falsifiable. Could a clairvoyant determine which scenario occurred, using only the resolution criteria?
  • Diagnostic. Does each criterion distinguish its scenario from the others?
  • Linked. Are drivers and forecasting questions linked, with the AI-suggested influence weights reviewed rather than left unchecked?
  • Challenged. Has a red-team session been run, and have its findings been resolved rather than dismissed?
  • Current. Is the weekly reassessment being reviewed, and is it clear when this set will be resolved?

A scenario set that satisfies these criteria fulfills the objective Wack described: it does not predict a single outcome, but changes the range of outcomes the reader considers possible, and provides a basis for assessing accuracy once the outcome is known.

This guide was written with the help of AI, but was reviewed and edited by a human.

Put the method into practice

Ready to build a scenario set that holds up?