Environment: DEV DEV e2e-parallel presubmits, including Tide batches. PR-regression-filtered means runs on the final commit of a PR that subsequently merged, plus batch retests of PRs that passed their individual checks. Each run counts once. This measures observed reliability, not test effectiveness. Individual-PR eligibility becomes known after merge, so recent filtered points can change. Batches qualify immediately.
Presubmit overall job success Raw includes all DEV e2e-parallel runs, including Tide batches. Filtered includes post-good runs plus batches, counted once. Hover or focus on points for counts and 95% Wilson confidence intervals. These describe sample-size uncertainty, not a guarantee; retests and shared incidents are correlated. Hollow markers indicate fewer than 30 eligible runs or a partial period. Missing or invalid samples remain gaps. Whiskers show the intervals. Raw is offset left and filtered right of each date for readability.
Raw success rate
PR-regression-filtered success rate
All runs Provision success excludes Other failures and counts runs that passed provisioning even if E2E later failed. E2E success excludes provision and Other failures. Overall success includes all runs. These rates have different denominators and are not additive. Weekly rates use summed counts; weeks start Monday UTC. Hover or focus on points for counts and 95% Wilson confidence intervals. These describe sample-size uncertainty, not a guarantee; retests and shared incidents are correlated. Hollow markers indicate fewer than 30 eligible runs or a partial period. Missing or invalid samples remain gaps. Series are slightly offset around each date for readability. Intervals are in tooltips only.
Provision success
E2E success
Overall success
Other failures (% of all runs) Other failures divided by runs in this chart's population; lower is better. Includes build failures, CI infrastructure failures, and unclassified failures, not a separate pipeline stage. The post-good + batches strip uses only that subset's failures and run count. Both strips use a 0–100% scale. Hollow bars mark low samples or partial periods; hover or focus for counts and confidence intervals.
Post-good + batches Provision success excludes Other failures and counts runs that passed provisioning even if E2E later failed. E2E success excludes provision and Other failures. Overall success includes all runs. These rates have different denominators and are not additive. Weekly rates use summed counts; weeks start Monday UTC. Hover or focus on points for counts and 95% Wilson confidence intervals. These describe sample-size uncertainty, not a guarantee; retests and shared incidents are correlated. Hollow markers indicate fewer than 30 eligible runs or a partial period. Missing or invalid samples remain gaps. Series are slightly offset around each date for readability. Intervals are in tooltips only.
Provision success
E2E success
Overall success