Scalable Delphi: evidence flows into an LLM expert panel, producing structured probability estimates

Scalable Delphi

Large language models for structured risk estimation

Calibrated, auditable risk estimates through structured deliberation in minutes instead of months. Every estimate is traceable to its reasoning and easy to update when evidence changes.

Risk models in high-stakes domains like cybersecurity, nuclear safety, or climate science depend on quantities that cannot be measured directly, such as how much AI assistance raises a ransomware group's odds of escalating privileges inside a network. The standard way to estimate them is a Delphi study, in which experts refine their estimates over several rounds. These studies take months; large ones like NUREG-1150 for U.S. nuclear plants took years and hundreds of experts. Few organizations can afford that, and finished studies are rarely updated when the evidence changes.

Scalable Delphi adapts the classical Delphi protocol for LLM agents: diverse expert personas estimate independently, a mediator synthesizes anonymized feedback, and panelists refine their estimates through structured deliberation.

01
Expert panel
Diverse LLM personas estimate independently
→
02
Mediation
Anonymized feedback & rationale summary
→
03
Refinement
Revised estimates after deliberation
r = 0.95
Mean correlation with ground truth (two frontier models, three benchmarks)
3 studies
Agreement with expert estimates on most quantities
Minutes
vs. months for traditional Delphi

We evaluated Scalable Delphi on three cybersecurity benchmarks, predicting held-out agent success rates, and against three published expert elicitations: two on AI-enabled cyber risk and one on ice-sheet contributions to sea-level rise.

Calibration
Predicted vs. actual success rate (%). Each dot is one held-out agent (agent × task type on BountyBench). Dashed line = perfect prediction.
Expert alignment
LLM panels compared to two independent human expert panels from Murray et al. (2025). Tasks ordered by difficulty.
Tobias Lorenz, Mario Fritz · CISPA Helmholtz Center for Information Security · 2026

Scalable Delphi runs structured expert elicitation with LLM panels. The estimates track benchmark ground truth (r=0.91–0.98), agree with human experts on most quantities in three reproduced studies, and can be updated in minutes when the evidence changes.

BibTeX
@misc{lorenz2026scalabledelphi,
    title         = {Scalable Delphi: Large Language Models for Structured Risk Estimation},
    author        = {Tobias Lorenz and Mario Fritz},
    year          = {2026},
    eprint        = {2602.08889},
    archivePrefix = {arXiv},
    primaryClass  = {cs.AI},
    url           = {https://arxiv.org/abs/2602.08889}
}