PR

Probability

Decision Tree Confidence Calculator

Calculate a Wilson confidence interval for an observed success rate and compare it with the analytical success probability from a two-branch decision tree.

TREE SUCCESS CONFIDENCE

Compare observed success uncertainty with the probability declared by the tree

This calculator places a Wilson score interval around observed binary successes and compares that interval with the analytical success probability implied by a two-branch chance tree. Validation teams can use the comparison as a calibration diagnostic, provided trials are comparable and the tree was specified independently of the same outcomes.

Observed success rate-
Wilson lower bound-
Wilson upper bound-
Analytical tree success-
Tree probability status-
Critical z value-

TREE SUCCESS CONFIDENCE

Wilson interval and tree-comparison ledger

Use interval coverage as evidence about compatibility between observed success frequency and the declared tree, not as automatic proof that the tree is correct or incorrect.

Editorial illustration of an observed success marker and an analytical tree marker checked against a measured interval bracket
The interval summarizes sampling uncertainty in observed success; the tree marker is a separate analytical claim being checked for compatibility.
Wilson interval and tree-comparison ledgerCurrent unrounded calculation path
Live detail from current inputs
QuantityObserved value / centerAdjustment / multiplierCalculated valueUnit

CURRENT CALCULATION PROCESS

Formula, substitution, intermediate values, and reconciliation

p-hat = x/n; d = 1 + z^2/n; center = (p-hat + z^2/(2n))/d; half = z*sqrt(p-hat*(1-p-hat)/n + z^2/(4n^2))/d

The observed proportion is adjusted with the Wilson score construction rather than a simple normal interval. The analytical tree success rate is calculated from branch-weighted conditional success probabilities, then classified as inside or outside the observed interval without changing either quantity.

    HOW TO USE THIS MODEL

    Use observed trials as an external check on the tree

    1. Freeze the tree probabilities and the definition of success before reviewing the validation outcomes.
    2. Enter the number of independent comparable trials and the successes observed under that exact definition and follow-up period.
    3. Choose a confidence level required by the validation plan rather than changing it after seeing whether the tree rate is included.
    4. Review the observed rate and Wilson endpoints before the inside-or-outside label, because distance and interval width affect interpretation.
    5. If the analytical rate is outside, investigate branch mix, conditional rates, dependence, drift, and outcome coding before revising the model.

    TREE SUCCESS CONFIDENCE FUNDAMENTALS

    Confidence concepts for validating a success tree

    Observed proportion
    The success count divided by the independent trial count; it is a sample estimate rather than the analytical tree probability.
    Wilson score interval
    A bounded interval for a binomial proportion that behaves better near 0 or 1 and at moderate sample sizes than the unadjusted Wald interval.
    Analytical tree rate
    The branch-weighted success probability implied by the declared chance tree before observing this validation sample.
    Compatibility check
    An analytical value inside the interval is not contradicted at the chosen resolution, but many other values may also be compatible.
    Calibration evidence
    Repeated comparisons across representative conditions provide stronger model evidence than one aggregate interval.

    MODEL AND FORMULA

    Why the Wilson interval is used for binary success evidence

    p-hat = x/n; d = 1 + z^2/n; center = (p-hat + z^2/(2n))/d; half = z*sqrt(p-hat*(1-p-hat)/n + z^2/(4n^2))/d

    The observed proportion is adjusted with the Wilson score construction rather than a simple normal interval. The analytical tree success rate is calculated from branch-weighted conditional success probabilities, then classified as inside or outside the observed interval without changing either quantity.

    DEEPER ANALYSIS

    Interpretation issues beyond the interval endpoints

    Inside is not validated

    An analytical rate can fall inside a wide interval because the validation sample is small. Coverage means the data are not sharply inconsistent at that level; it does not prove branch probabilities, structure, or transportability.

    Outside does not identify the defective input

    The mismatch can arise from branch mix, conditional success rates, success coding, time drift, dependence, or selection. The interval flags inconsistency but cannot allocate blame among model components.

    Aggregate agreement can hide branch errors

    Overstating success in A and understating it in B may cancel in the total. Validate branch-specific rates and routing proportions as well as overall success.

    WORKED DECISION CASES

    Two different outcomes from the same comparison rule

    Analytical rate inside a precise interval

    The default tree implies 60% success. A validation set of 1,000 trials with 612 successes yields a Wilson interval that includes 60%. The team records compatibility while still reviewing branch-level calibration and representativeness.

    Unexpected operating regime

    A validation sample collected after a routing-policy change places the old analytical success rate outside its interval. The result is not used to tune a probability immediately; the team first checks whether the branch-A share and conditional populations changed.

    TECHNICAL LANGUAGE

    Binary validation and interval language

    Binary outcome
    A trial classified into exactly one of two states under a frozen rule, here success or non-success.
    Observed success rate
    The empirical fraction x/n in the validation sample.
    Wilson center
    The adjusted center of the Wilson score interval, generally different from the raw observed rate.
    Half-width
    The distance from the Wilson center to an interval endpoint before clipping to the probability range.
    Analytical probability
    A probability calculated from model inputs rather than estimated directly from the validation count.
    Calibration
    Agreement between stated probabilities and observed frequencies across comparable cases and probability levels.

    EVIDENCE AND DATA LINEAGE

    Retain trial-level outcomes and a pre-specified validation plan

    Keep trial identifiers, branch assignments, success coding, observation window, exclusions, missing outcomes, timestamps, operating condition, and the version and freeze date of the tree. Record whether trials are independent and whether the validation set was used to fit any entered probability. If data were reused for fitting, label the comparison in-sample rather than independent validation.

    LIMITS AND EXCLUSIONS

    What interval coverage cannot demonstrate

    • The Wilson calculation assumes binomial-style independent trials with a common success probability over the assessed population.
    • The comparison does not validate terminal payoffs, causal effects, branch completeness, or calibration within each branch.
    • A confidence interval is sensitive to the chosen confidence level and sample size; inclusion is not an equivalence test.
    • Repeated unplanned comparisons inflate the chance of at least one apparent mismatch and require a defined multiplicity strategy.

    RELIABLE SOURCES

    References for this model and its decision limits

    FREQUENTLY ASKED QUESTIONS

    Questions about Wilson intervals and tree validation

    Why not use observed rate plus or minus z times its standard error?

    The simple Wald interval can perform poorly, especially near probability boundaries or with limited samples. The Wilson score interval offers more reliable coverage in many binomial settings.

    Does inside interval mean the model passes?

    No. It means the aggregate analytical rate is compatible with this sample at the chosen interval resolution. Structural, branch-level, and external-validity checks remain necessary.

    What if observed successes are zero or equal all trials?

    Wilson bounds remain within 0% and 100% and still express uncertainty. A zero-width certainty claim would be inappropriate.

    Can trials collected over time be treated as independent?

    Only with evidence. Shared environments, learning, maintenance, drift, and repeated units can create dependence and make the nominal trial count overstate information.

    Should I adjust the confidence level until the tree value is included?

    No. Select the level in the validation protocol. Post hoc adjustment changes the decision rule to fit the observed result.

    How can overall agreement hide model defects?

    Errors in branch weights or conditional success rates can offset one another. Compare observed and modeled quantities at the branch level whenever sample size permits.

    IMPORTANT NOTE

    Compatibility is one validation result, not model approval

    This page supplies a Wilson interval and an aggregate analytical comparison. It does not certify the decision tree, establish equivalence, or replace a validation plan covering branch calibration, dependence, drift, outcome quality, and the consequences of model error.