Direct answer
An average is a useful summary only after the population, exposure, operating context, and distribution are understood. By itself, it can make an inconsistent process look acceptable, reward a small or favorable sample, and reverse the apparent ranking when data is grouped differently.
For drilling performance, show the distribution and denominator first. Then use the average—or preferably several summaries—as one part of the evidence.
The same average can describe different operations
Consider two synthetic connection-time samples:
Rig A: 4.8, 4.9, 5.0, 5.1, 5.2 minutes
Rig B: 2.0, 2.5, 5.0, 7.5, 8.0 minutes
Both average five minutes. Rig A is consistent. Rig B combines very fast and very slow outcomes with a much larger tail. A manager deciding where to investigate would reach different conclusions after seeing the distributions.
The same issue appears in ROP, slide performance, stand duration, trip speed, model residuals, and alarm frequency. The mean answers “where is the center under this weighting?” It does not answer “how repeatable is the process?” or “how often does the tail create operational risk?”
Begin with the observation unit
Before comparing performance, define what one observation represents:
- one sensor sample;
- one foot or depth bucket;
- one connection phase;
- one stand;
- one run;
- one formation interval;
- one well.
Treating one-second rows as independent observations can make a long interval appear more statistically certain simply because it contains more rows. Treating every well equally can hide that one well contributed far more footage or connections. Neither weighting is universally correct.
State the unit and the weighting rule beside the metric.
Use a compact distribution set
A practical drilling comparison does not need every statistical measure. Start with:
| Measure | What it contributes |
|---|---|
| Count and exposure | Shows how much evidence exists: events, footage, or time. |
| Median | Describes a robust center for skewed data. |
| Mean | Preserves total-rate relationships when the weighting is appropriate. |
| P10/P90 or IQR | Shows spread and tails. |
| Minimum/maximum | Helps inspect extremes but should not rank typical performance. |
| Distribution or ECDF | Reveals multimodality, skew, and overlap. |
| Confidence interval | Shows uncertainty in an estimated difference. |
Percentile direction must be explained. A low connection time is favorable; a high ROP may be favorable only within quality and operational constraints. “P90” without a direction and definition is ambiguous.
Comparable cohorts come before statistics
No statistic repairs an invalid cohort. A raw ROP comparison can mix vertical, curve, and lateral footage; different formations; rotary and slide intervals; BHAs; rigs; bit sizes; and operating limits.
A defensible workflow is hierarchical:
- Restrict to the same metric definition and units.
- Separate relevant hole sections and operating states.
- Match formation, trajectory, BHA, rig, or plan context where necessary.
- Require minimum exposure and quality coverage.
- Show the remaining differences and missing context.
- Calculate distributions and uncertainty.
- Let an engineer inspect the underlying intervals before interpreting causes.
The goal is not to remove every difference. It is to make the comparison's limitations visible.
Beware Simpson's paradox
A rig can appear faster overall while being slower within each hole section if its work mix contains more footage in an inherently faster section. When the sections are combined, the mix drives the average.
Always compare the aggregate with important strata:
- hole section;
- formation or target;
- operation/rig state;
- BHA/run;
- well or pad sequence;
- day/night or crew only when the data and governance support it.
If the ranking changes after stratification, the aggregate should not be presented as a simple performance conclusion.
ROP needs exposure and state definitions
Several values can be called average ROP:
- average of instantaneous ROP samples;
- total drilled footage divided by on-bottom drilling time;
- total footage divided by elapsed section time;
- average of one-foot bucket ROP;
- average of stand-level ROP.
They answer different questions. The metric definition should state the numerator, denominator, eligible states, depth interval, unit, and treatment of zero or missing values.
For offset comparison, raw ROP should usually be stratified or normalized before ranking. The distributions should retain footage, state coverage, and sample size so that one short high-ROP interval does not outrank a repeatable longer run without explanation.
Connection performance is a phase problem
A connection is not one indivisible duration. Slip-to-slip and weight-to-weight definitions capture different boundaries, and the useful diagnosis may be in one phase:
- coming off bottom;
- pumps down;
- setting/slips or handling pipe;
- making the connection;
- pumps up;
- returning to bottom and stabilizing.
Show the total distribution and the phase distributions. A stable total can hide one phase improving while another degrades.
Quantify uncertainty before ranking
Small samples are common in drilling. A run may contain few comparable events; a pad may have only several wells. A point estimate can move materially with one observation.
Bootstrap intervals are often useful because they make fewer distribution assumptions, but the resampling unit must respect dependence. Resampling one-second rows when the real unit is a connection or well creates false precision.
Report effect size and interval, not only a winner. If the distributions overlap substantially or the cohort is small, “insufficient evidence to distinguish performance” is an analytically useful result.
A practical comparison card
For each well, run, rig, or cohort, present:
- metric definition and favorable direction;
- comparable-context filters;
- count plus time/footage exposure;
- median and mean;
- percentile range;
- distribution plot;
- data-quality coverage;
- difference and uncertainty versus baseline;
- links to the contributing intervals.
This format supports action without pretending the ranking is more certain than the data.
Final checklist
- Define the observation unit and weighting.
- Preserve numerator, denominator, and exposure.
- Match operating context before ranking.
- Show distributions and tails, not only averages.
- Stratify important confounders and check for reversals.
- Use uncertainty appropriate to the real sampling unit.
- Avoid causal claims from observational differences alone.
- Link summaries back to the underlying intervals.
Drilling performance is not one leaderboard. It is a set of contextual distributions that help engineers decide where a difference is real, repeatable, and worth investigating.