24  Robustness and Sensitivity

WarningUnder development

This chapter is part of a book in active development and has not yet been through the author’s review. Content may change as the review advances.

Open In Colab

The research decision. You decide the full list of robustness and sensitivity checks you will run, and you decide it before you look at a single result, then report every one of them. Two things separate a defensible finding from a lucky slice: each version has to be a valid analysis on its own, and the looking has to come after the deciding. Pre-listing protects the second; nothing rescues the first. A check you think of later is still worth running, as long as you label it as one you added after seeing results and report it with the rest.

24.1 Why this decision matters

The decision on the table: which alternative versions of your analysis you commit to running and reporting, chosen before you see any of their answers.

“Everyone shows me the version of the analysis that worked. I assume that one exists. My real question is what happened to all the reasonable versions you could have run instead, and whether you looked at them before or after you saw this answer.” — a journal reviewer, opening the first serious read your work will get

That reviewer is not accusing you of cheating. They are naming a plain fact about analysis. Almost every real estimate depends on choices you could have made differently, and any single choice can flatter the result. If you report only the version that worked, you have handed over a number with no way to tell a delicate truth from a lucky slice (Simmons et al. 2011). This chapter gives you the habit that answers the question before it gets asked.

24.2 The concept

An estimate is not one number. It is one number that depends on several choices, and a critic attacks the choices. So you attack them first. Two named attacks do this work, and people mix them up constantly.

A robustness check is re-running the same finding under a different but equally defensible choice, to see whether the answer stays. Example: you reported an average order value, so you also compute it after dropping the one account that turned out to be a wholesale buyer, and confirm the story holds. A sensitivity check, which some fields call a sensitivity analysis, is deliberately changing an assumption to see how far the answer moves (Rosenbaum 2002). Example: you assumed the checkout system logged every purchase, so you ask how much your result would shift if it quietly missed two percent of them.

Both attacks aim at the same target, the headline estimate: the single number that stands in for your whole finding. You attack it from four handles, and any honest plan turns one or more of them. The sample is which data you keep. The measurement is how you turn a concept into a number. The specification is the bundle of modeling choices behind the estimate. The metric is the yardstick you report.

Turning every handle at once and plotting the results gives a specification curve: a picture of one estimate under many defensible choices side by side (Simonsohn et al. 2020) (Steegen et al. 2016).

Before anything goes on that curve, it has to pass a gate with three parts. First, every version must test the same substantive claim, the one research question the curve is about, such as “does the offer raise shoppers’ spending?” Second, the dots inside one panel must be commensurable: measured in the same units, about the same group of people, for the same kind of quantity, so that reading them side by side means something. A choice that changes who is counted, what the outcome means, or what scale it is on does not disqualify the analysis, but it belongs in its own labeled panel, not mixed silently into this one. Third, one rule is strict: no version on this curve may define its group by something the treatment itself could have changed. “Average order value among the shoppers who actually bought something” fails that test, because whether someone bought is downstream of the offer. That comparison may be worth reporting, in its own place, clearly labeled for what it is. It is not another answer to the offer question, and folding it into the curve makes the curve unreadable. (Specialists do define careful causal quantities for post-treatment subgroups, but they need different assumptions and different machinery than a robustness curve; what is ruled out here is treating a raw survivor-style contrast as one more dot.)

Two more things the curve does not do. A curve that stays entirely on one side of zero shows that your answer is stable across the choices you listed. It cannot see a flaw that every one of those choices shares. If your comparison group was wrong to begin with, all sixteen versions inherit that error and lean the same way together, and the picture looks reassuring precisely when it should not. Shared bias is that common error running underneath an entire curve.

And the distance from the lowest dot to the highest answers only one question: how much did your analysis choices move the number? Call it the specification spread, the span your answer covered across the versions you ran. It says nothing about chance. Every single dot still carries sampling wobble of its own, the uncertainty you already learned to state with a standard error and an interval, and the spread is not a substitute for it. Report the spread as what your choices did, and keep the uncertainty of your lead estimate a separate, labeled thing.

The danger has a name. Specification searching is running many analyses and reporting only the one that gave the result you wanted, without disclosing the search. Its most notorious version, hunting for a statistically significant result, is what researchers call p-hacking. Its opposite is pre-listed checks: writing down every check you will run, and committing to report all of them, before you see any result.

That distinction is under new pressure, and you should know why. Working with an AI assistant means working in a loop: prompt, read, interrogate, refine, run again, with agentic tools now spinning those cycles on their own. Each turn of the loop can change a filter, a variable, or a model, which means each turn is a new specification. Twenty turns is twenty analyses. If you kept re-prompting because the number was disappointing, you ran a specification search and the tool kept no record of it for you. Keep your own log of the cycles, and your pre-listed plan survives contact with the loop.

24.3 A worked example

Your team randomly assigns a loyalty discount to shoppers and asks one question: does the offer raise spending? Your target, fixed before anything ran, is 30-day spending per shopper you assigned, counting the shoppers who spent nothing. Your headline: shoppers offered the discount spent about twelve percent more than those who were not (the numbers in this example are constructed). Before you defend that number, you attack it across the four handles.

  • Sample. All shoppers, or all shoppers except the one account that turned out to be a wholesale buyer placing bulk orders through the retail site, which is a legitimate exclusion on your definition of a customer, not a convenient one.
  • Measurement. Order value read as the total amount charged, or read with shipping and taxes stripped out.
  • Specification. Spending averaged with no adjustment, or adjusted for how much each shopper spent in the weeks before the offer.
  • Metric. The gain in dollars, or the gain as a percent.

Notice what is missing. Averaging only over shoppers who actually placed an order does not belong here, because the discount can change who places an order. That comparison answers a different question about a different group, so it goes in its own section with its own label, not on this curve.

You pre-list the combinations, run them, and sort them into panels before you plot anything. The adjustment choice keeps the same people, the same outcome, and the same units, so its versions share one panel. Dropping the wholesale account changes who is counted, so it opens a second panel. Stripping shipping and taxes changes what “spending” means, so it opens a third. Dollars versus percent is one comparison shown on two scales, not two pieces of evidence, so it never adds a dot at all.

Now the reporting rule that follows from those panels: one span per panel, never a span across them. In the main panel, all versions come back positive and the estimates run about plus eight to plus twelve percent. Say it that way, panel by panel: “In our primary panel, spending was higher under the offer in every version we pre-listed, with estimates spanning eight to twelve percent; our lead estimate is twelve percent, and its own sampling uncertainty is stated with it. Excluding the wholesale account, the same versions ran ten to fifteen percent.” A single eight-to-fifteen range stretched across all three panels would compare a net-of-tax number to a gross one, over different sets of shoppers, and mean nothing. The span says what your choices did. The uncertainty statement, built with the interval tools you already have, says what chance could do. Collapsing them into “the honest range is eight to fifteen” quietly claims a precision nobody computed.

Then you run one more attack, and its name matters, because three different checks get muddled together in practice.

An assignment-based null check, which statisticians call randomization inference, re-runs the assignment your study actually performed. Keep every enrolled shopper and their observed outcomes, re-draw who is “offered” using the exact procedure the real study used, with the same group sizes and the same blocking, and recompute the gap. Do that many times. One redraw proves nothing, because two randomly formed groups always differ a little by chance, so what you build is a pile of fake gaps: what your machinery produces when the offer does nothing at all. This move needs a real assignment procedure to copy, so it belongs to randomized studies.

A permutation check shuffles labels that were not randomly assigned. It answers something only if you can argue that the groups were exchangeable to begin with, and in most observational data you cannot. Shuffling anyway builds a pile describing a world you never had.

A negative control, sometimes called a placebo, is different from both: it points your analysis at an exposure, outcome, or period that should be causally null while still running through the same measurement and bias pathways. Example: check for an “effect” of the offer in the weeks before it launched. That is what catches a pipeline inventing signal, and it works whether or not anything was randomized.

For this randomized trial, use the first. Comparing your real gap with that pile is the beginning of a formal argument the uncertainty chapter finishes; here it is a smell test, and a real gap the no-effect pile produces routinely is a warning that your machinery, not the discount, may be doing the work.

That the curve’s rows must all estimate the SAME quantity, or the spread means nothing, is the condition the method rests on (Simonsohn et al. 2020).

The block below turns all four handles and prints the whole grid, plus the one comparison that must stay off it. Check that every row answers the same question before you read any of the numbers.

import numpy as np, pandas as pd
SEED = 464
rng = np.random.default_rng(SEED)

n = 4000
offered = rng.random(n) < 0.5
prior = rng.gamma(2.0, 30.0, size=n)                    # spending before the offer
ordered = rng.random(n) < (0.42 + 0.025 * offered)      # the offer moves WHO orders
spend = np.where(ordered, rng.gamma(2.0, 33.0, size=n) * (1 + 0.12 * offered)
                 + 0.25 * prior, 0.0)                   # past spending carries over
ship_tax = np.where(ordered, 6.0 + 0.07 * spend, 0.0)
gross = spend + ship_tax
wholesale = np.zeros(n, dtype=bool); wholesale[0] = True
gross[0] = 9_400.0                                      # the bulk buyer

df = pd.DataFrame({"offered": offered, "gross": gross, "net": spend,
                   "prior": prior, "ordered": ordered, "wholesale": wholesale})

def pct_gain(d, col, adjust=False):
    v = d[col] - (np.polyval(np.polyfit(d.prior, d[col], 1), d.prior)
                  if adjust else 0)
    a, b = v[d.offered].mean(), v[~d.offered].mean()
    base = d.loc[~d.offered, col].mean()
    return (a - b) / base * 100

grid = []
for s_label, d in (("all shoppers", df), ("minus the wholesale account",
                                          df[~df.wholesale])):
    for m_label, col in (("charged total", "gross"), ("net of shipping/tax", "net")):
        for sp_label, adj in (("unadjusted", False), ("adj. prior spend", True)):
            grid.append({"sample": s_label, "measurement": m_label,
                         "specification": sp_label,
                         "gain %": round(pct_gain(d, col, adj), 1)})
g = pd.DataFrame(grid)
print(g.to_string(index=False))
print(f"\nall eight share ONE estimand: spending per shopper ASSIGNED.")
print(f"range {g['gain %'].min():.1f}% to {g['gain %'].max():.1f}% — a spread of "
      f"defensible answers, not a confidence interval")

# the specification that does NOT belong: conditioning on ordering
bad = pct_gain(df[df.ordered], "gross")
print(f"\naveraging over ORDERERS only: {bad:.1f}% — a different estimand, on a")
print("group the offer itself selected. it does not belong on the curve")

24.4 An AI failure case

You paste your four-handle grid and ask an AI to write your robustness section. Back comes a long, well-organized passage that runs the sample handle, the specification handle, and the metric handle, and closes with “the twelve percent result is robust across all standard specifications.” It reads as thorough and complete. The trap is that it never listed the measurement handle, the total amount charged versus shipping and taxes stripped out, and that is precisely the choice that moves your estimate the most. The section looks exhaustive while omitting the one check a reviewer would reach for first.

You catch it because you pre-listed all four handles yourself before opening the tool. Reading its section against your own list, you notice measurement is missing, run that comparison, and watch the estimate slide toward the low end of your range. A passage that looks complete is not the same as one that is complete. You check its list against yours, not against how finished it sounds.

24.5 It is your turn

You are working inside Studio 8: Stress-test and adjudicate. Keep what you write here; the studio’s milestone chapter is where it joins the other lessons’ pieces into one artifact you can defend.

Your project has an analysis you ran with an assistant and a headline number you re-derived yourself. This step tries to break that number before a reviewer gets the chance.

The hands-on half of this section lives in the chapter’s companion notebook: open it in Colab with the badge at the top, and work the steps there.

Commit your own answer first, then delegate. Each prompt is a checkable job, not a request for a verdict.

ImportantDo not delegate

Three calls stay yours. You decide which checks you commit to before you look, which flagged issues the data actually confirms, and the final claim, with its per-panel span and its boundary, that you defend. A reviewer, human or AI, proposes. You verify against your own data, and the evidence decides.

  1. Write your headline estimate in one sentence. Then list its four handles: the sample you kept, the measurement you chose, the specification you fit, and the metric you report. For each, name one alternative you could defend just as well as the one you picked.

  2. Pre-list the checks you will run, at least two, aimed at the handles a hostile reader would attack first. Write them down before you run any of them, and commit now to reporting all of them however they come out.

    Locate the standard checks.

    Act as a methods assistant for a two-group comparison of average customer spending.
    Before any code, name the standard families of robustness and sensitivity checks for
    this design and the four things I can vary to stress the estimate. Cite a real methods
    reference you are confident exists; if you are unsure a source is real, say so.

    After running, verify: open the reference and confirm the checks and the source exist. Counters confident fabrication (an invented method or citation arrives as confidently as a real one).

    A second angle, optional:

    Delegate, with a list to verify.

    Here is my pre-listed robustness grid for one estimate of the offer's effect on
    30-day spending per assigned shopper: I vary the sample (all shoppers / drop the
    wholesale account), the measurement (total charged / shipping and taxes stripped
    out, in its own labeled panel), and the specification (no adjustment / adjusted for
    each shopper's pre-offer spending). Name up to three additional, equally defensible
    versions I did NOT list, flag any of my versions that changes the outcome
    definition or conditions on something the offer could have caused, and for each
    suggestion say which direction you expect it to push the estimate.

    After running, verify: check each suggestion against the gate (same claim, commensurable within its panel, nothing conditioned on a post-treatment outcome) and add only the ones you can actually run on your own data. Counters illusion of completeness (a tidy grid that quietly omits the one choice that would break the claim).

  3. Before you run anything, run your list through the gate. Every version tests your one substantive claim; versions that change the outcome’s definition or scale get their own labeled panel; and any version that keeps only units whose status your treatment could have changed comes off the curve entirely and is reported separately for what it is.

  4. Run every check on your list. Report two things, not one. First, the direction and a span computed within each commensurable panel, never across them (“in the primary panel, positive in all four versions, between eight and twelve percent”), with any panel that changed the population, the outcome definition, or the scale given its own labeled span. Second, your lead estimate with its own uncertainty stated beside it, built with the tools the uncertainty chapter gives you. The span measures your choices, not chance, and neither alone is an honest report.

    Red-team your span.

    Here is my claim: "higher spending in every version we pre-listed; within my primary
    panel, which holds population, outcome and units fixed, the estimates spanned 8 to 12
    percent." Act as a hostile reviewer of a retail pricing study. Name the single worst
    way this could be specification searching rather than robustness, and the one check
    that would expose it. Also check whether my span mixes versions that changed the
    population, the outcome definition, or the scale; those belong in separate labeled
    panels with their own spans. Do not rewrite the claim for me.

    After running, verify: run the check it names against your own data; the flaw is real only if your numbers confirm it. Counters sycophantic agreement (praise that reviews your ego, not your evidence).

  5. Ask what every version on your list has in common. If one flaw sits underneath all of them, such as a comparison group chosen badly, the whole curve leans together and agreement proves nothing. Name that shared threat in a sentence, since your curve cannot test it.

  6. Add one null check, and name it with the chapter’s three labels, because each label comes with its own procedure and its own reading. If your study randomized, run an assignment-based null check: re-draw the assignment the way the study actually did it, many times, and collect the fake gaps. That pile is what nothing-going-on looks like through your machinery, and it is not a pile of zeros. Ask whether your real number is ordinary in that pile or stands far outside it, and treat a number the pile produces routinely as a warning about your machinery. If your data are observational, you have no assignment procedure to copy. Shuffle labels only if you can defend in writing why your groups were exchangeable under the null; call it a permutation check and state that assumption beside the result. Otherwise use a negative control: an outcome your cause could not touch, or a period before it existed, running through the same bias pathways as your real analysis. A negative control does not build a pile for your headline number to sit in. It produces its own estimate, which should be near zero, so estimate that contrast the same way you estimated your headline and report how far it lands from zero given its own uncertainty. A clearly nonzero negative control is your machinery inventing signal, and that concern attaches to your headline too.

  7. Read back your AI cycle log and mark any re-prompt that followed a disappointing number. Every version you actually ran belongs in the report, including the ones that did not help you.

  8. Log the pre-listed plan and its full results in your AI Research Ledger, and verify at least one output with a named method from the Verification Guide. A counterexample hunt suits a robustness claim: try honestly to build one defensible specification that flips the sign, and if you find it, that specification goes in your report. An AI reviewer may run the check with you; the decision to accept or reject stays yours.

References

Rosenbaum, Paul R. 2002. Observational Studies. 2nd ed. Springer. https://doi.org/10.1007/978-1-4757-3692-2.
Simmons, Joseph P., Leif D. Nelson, and Uri Simonsohn. 2011. “False-Positive Psychology: Undisclosed Flexibility in Data Collection and Analysis Allows Presenting Anything as Significant.” Psychological Science 22 (11): 1359–66. https://doi.org/10.1177/0956797611417632.
Simonsohn, Uri, Joseph P. Simmons, and Leif D. Nelson. 2020. “Specification Curve Analysis.” Nature Human Behaviour 4: 1208–14. https://doi.org/10.1038/s41562-020-0912-z.
Steegen, Sara, Francis Tuerlinckx, Andrew Gelman, and Wolf Vanpaemel. 2016. “Increasing Transparency Through a Multiverse Analysis.” Perspectives on Psychological Science 11 (5): 702–12. https://doi.org/10.1177/1745691616658637.
opens in a new tab