Why Distribution Choice Matters
A distribution is a mathematical model of how probability is allocated across possible outcomes. It influences what appears typical, how quickly probability declines away from the center, how much remains in the tails, and whether unusual values are treated as rare or reasonably plausible.
Choosing the right probability distribution starts with checking whether the assumed model is suitable for the observed data and intended analysis. The NIST Engineering Statistics Handbook explains that distributional assumptions should be verified before they are used for statistical intervals, hypothesis tests, or simulations.
In E&P, it also determines how often a Monte Carlo model samples small, central, or large values for the geological, petrophysical, PVT, recovery, and production parameters used in in-place and reserves calculations. The stakes are not abstract, the same choice applied to Area, thickness/HC height, porosity, water saturation, FVF, net-to-gross, and recovery factor directly changes probabilistic in-place volumes (STOIIP, GIIP) and reserves at P10, P50, and P90.
That is why distribution choice is not a cosmetic decision. It may change P10, P50, and P90, the probability of crossing a threshold, simulated extremes, reliability estimates, and the apparent level of risk.
The Most Common Distribution Mistake
Everyone knows the first mistake: assuming every dataset should be Normal. Here is the sneakier one, assuming that two curves which look almost identical near the center must behave the same way everywhere else. They often do not, and the difference shows up exactly where it matters most, i.e. the tails.
The Brill-Viz Probability Distribution Library
Both Grouped and Smart-Stats, in Brill-Viz, contain a broad and analytically meaningful set of continuous distributions:
The current library covers symmetric, skewed, bounded-range, and progressively heavy-tailed behavior. It is broad enough for a strong continuous-distribution workflow without pretending that any fixed library contains every possible model.
Visual Reference: The Distributions Available in Brill-Viz
The following figure represents the current Smart and Grouped Stats probability distributions:
Why Similar-Looking Distributions Are Not Interchangeable?
1. Positive and right-skewed: Lognormal, Gamma, Weibull, and Exponential
All four can occupy the positive axis and decline to the right, but they represent different mechanisms and can produce different upper-tail probabilities.
2. Symmetric but increasingly heavy-tailed: Normal, Student’s t, Laplace, and Cauchy
These distributions can all be centered at the same value, yet they differ sharply in how much probability they allocate to extreme deviations.
3. Same bounds, different assumptions: Uniform vs Triangular
Both are bounded by a minimum and maximum. Uniform treats every value in the interval as equally likely. Triangular gives more probability near a specified most-likely value. Their ranges can be identical while their percentiles and simulated central concentration differ.
A Real Modeling Choice: Auto Says Lognormal, the Analyst Selects Gamma
Suppose Auto (AIC) identifies Lognormal as the strongest candidate, but the analyst selects Gamma. The Gamma curve may still look acceptable around the main body of the histogram. That visual similarity can hide a meaningful difference in the upper tail.
A fitted Lognormal may preserve more probability in very high values than a comparable Gamma model. Selecting Gamma could therefore reduce an upper percentile or the probability of exceeding a high threshold. The exact direction and magnitude depend on the fitted parameters, so the effect must be compared rather than assumed.
Oil and Gas Exploration and Production: Distribution Choice Can Move Volumes
In exploration and production, distribution choice enters directly into probabilistic volumetrics, development forecasts, and reserves ranges. The arithmetic may appear simple, but the model is usually multiplicative: modest upper- or lower-tail assumptions on several inputs can combine into a materially different low, best, or high case.
A distribution that fits the center reasonably well can still misstate the range that drives acreage value, appraisal decisions, facility sizing, well count, recovery planning, and economic screening.
The parameter list is wider than a single textbook equation. A practical uncertainty model may include mapped or connected area, gross rock volume, hydrocarbon contact, hydrocarbon height, gross and net thickness, net-to-gross, porosity, water or hydrocarbon saturation, oil or gas formation volume factor (Bo or Bg), fluid properties, recovery factor, well productivity, decline parameters, EUR, well count, uptime, facility availability, compression or pressure constraints, development schedule, abandonment or economic limit, and other project-specific factors.
Each needs a distribution or scenario that matches its physical support, evidence, scale, and dependency structure.
Exploration And Appraisal Perspective
Early in the asset life, area or gross rock volume, contact depth, hydrocarbon height, net reservoir development, fluid fill, and reservoir quality may dominate the range. The data are usually sparse and spatially biased: a few well penetrations are observations at specific locations, not automatically a complete description of field-scale uncertainty.
Alternative fault-block connections, contacts, facies concepts, or charge cases may be better represented as discrete geological scenarios, each with conditional parameter distributions, rather than forcing everything into one smooth curve.
Chance of discovery or chance of development should also be kept conceptually separate from the conditional volume distribution unless a clearly defined risked-volume calculation is being performed. A small conditional discovery may still have a high chance of success, while a very large conditional high case may have a low chance of occurrence.
Blending these concepts inside one input distribution makes the result difficult to audit.
Development And Production Perspective
As appraisal, testing, pressure, PVT, completion, production, and surveillance data accumulate, the distributions should be updated rather than inherited unchanged from exploration.
'Static' uncertainty increasingly interacts with 'dynamic' uncertainty: recovery factor, pressure support, water or gas movement, deliverability, decline behavior, type-well or well-group EUR, uptime, workovers, facility constraints, compression, development sequence, and economic limit.
Grouped distributions by reservoir, facies, completion, well type, or operating regime are often more informative than a single field-wide fit.
A probability distribution does not by itself classify resources as reserves. Technical recoverability, project definition, commerciality, approvals, timing, economics, and the applicable reporting framework still matter. Distribution analysis supports the range; it does not replace subsurface interpretation, reservoir engineering, development planning, or reserves governance.
The Hidden Multiplier: Correlation And Geological Consistency
Even a well-selected marginal distribution can produce a poor volumetric model if dependencies are ignored. Area and hydrocarbon height may be linked through the mapped area-depth relationship; porosity, permeability, net-to-gross, and water saturation may vary together by facies; Bo or Bg belongs to a coherent pressure-temperature-PVT case; recovery factor depends on reservoir quality, drive mechanism, development concept, and operating constraints; and rates, well count, uptime, and facility capacity interact through time.
Independent random sampling can create combinations that are statistically possible one variable at a time but geologically or operationally impossible together.
The opposite error is also possible: applying excessive correlation can artificially narrow the output range. Correlation should therefore come from geological geometry, petrophysical relationships, engineering physics, analog evidence, or explicit scenarios. not from a convenient default.
Brill-Viz can help diagnose and compare the individual input shapes; those assumptions can then inform Brill-Volumes or another audited Monte Carlo workflow that preserves bounds, dependencies, and scenarios.
AIC Is A Starting Point, Not A Reservoir Model
Auto (AIC) can rank the supported distributions for the observed sample, but subsurface data often include few wells, clustered measurements, truncation, detection limits, mixed facies, analogs, interpreted maps, and expert scenarios.
A statistically strong fit to plug porosity, log-derived saturation, well thickness, or analog field size may not be the correct field-scale uncertainty model.
The analyst must ask whether the sample represents the decision variable, whether physical limits are enforced, whether groups should be separated, and whether the selected model remains coherent with the geological and engineering concept.
Same Evidence, Different Volume Range: A Simple Example
Suppose seismic interpretation defines a firm minimum and maximum closure area and the team also has a defensible most-likely area. Uniform, Triangular, and an unbounded Lognormal assumption can all look plausible over part of that range, yet they allocate probability very differently.
Uniform treats every mapped area as equally likely; Triangular concentrates probability near the most-likely interpretation; Lognormal can preserve a longer high-side tail and may exceed the mapped closure unless it is truncated.
Once area is multiplied by net thickness, porosity, hydrocarbon saturation, and the FVF adjustment, the difference can materially move the petroleum P90, P50, and P10 in-place range. Adding recovery-factor and production uncertainty can widen or reshape the reserves range again.
Common Overrides and Their Possible Consequences
Every override choice above ultimately serves the same goal: turning uncertainty into a measurable basis for decisions. (The distinction between uncertainty and risk is covered in the companion article, “Beyond the Average".) The remaining question is how to choose, in practice, between the automatic recommendation and a manual override.
How Automatic Fitting and Manual Selection Differ
When automatic fit and process knowledge agree, confidence increases. When they disagree, the difference should be investigated, compared, and reported.
Statistical fit should be combined with knowledge of the process being analyzed. NIST’s guidance on choosing an appropriate distribution model recommends selecting models that make practical sense, fit the data, and have a plausible theoretical justification where possible.
What “Best Fit” Actually Means
Best fit means that one candidate performs better than the other supported candidates under the selected comparison method and observed sample. AIC ranks relative fit while penalizing unnecessary complexity; lower AIC generally indicates a better balance among the candidates compared.
In practice, Auto (AIC) fits every supported distribution to the data and selects whichever one produces the lowest AIC score, i.e. the entire mechanism behind the “Auto” recommendation.
It does not mean that the distribution is permanently proven to be the true process. Results can be affected by sample size, outliers, truncation, mixtures of groups, measurement limits, missing extremes, and the candidate families included.
How Brill-Viz Turns the Theory into a Practical Workflow
Understanding the theory is one thing; applying it fast, correctly, and repeatably is another. This is exactly where Brill-Viz's Smart and Grouped Stats take over, i.e. fitting the observed distribution, running an automatic AIC comparison, allowing manual overrides, layering grouped comparisons, and generating percentiles, interpretations, and publication-ready outputs in one pass.
Before Overriding Auto-Fit: Seven Questions
Common Mistakes
Key Takeaways
Closing Remarks
A probability distribution is not decorative mathematics. It is a model of where outcomes concentrate, how far they can spread, and how much probability remains in the tails. Those properties directly influence what appears likely, conservative, optimistic, safe, or risky.
In oil and gas, they can also change the estimated low, best, and high cases for in-place hydrocarbons, recoverable resources, production, and reserves.
Brill-Viz Grouped Stats already provides a strong and coherent distribution library. Its value lies not in offering every distribution ever developed, but in combining a meaningful candidate set with automatic AIC comparison, user-controlled alternatives, grouped analysis, and transparent reporting.
For E&P, that makes it a practical diagnostic layer for examining reservoir, petrophysical, PVT, reliability, rate, and EUR uncertainty before those assumptions are embedded in a volumetric, production, or reserves model.
The best practice is neither blind automation nor unrestricted manual choice. It is a disciplined combination: let the data speak, let geology, petrophysics, reservoir engineering, production evidence, and physical constraints challenge or confirm the fit, preserve material dependencies, compare the volumetric and decision consequences, and report the assumptions honestly.