v2. Bullets written with "*" unfold one at a time; bullets written with "-" appear together with the slide. Ordered steps written "1)" unfold, "1." do not. Slides open with the question and the argument arrives as you walk through it.
Careful with the statistics. Bars that merely fail to overlap are NOT proof of a difference: for two independent means with equal standard errors, bars that just touch correspond to only p ~ 0.16. Paired t-tests over the ten folds, against the best (depth 4): depth 2 vs 4 p = 0.054 (bars nowhere near overlapping) depth 3 vs 4 p = 0.78 depth 5 vs 4 p = 0.19 depth 6 vs 4 p = 0.062 (bars do not overlap) depth 7 vs 4 p = 0.012 So depth 2 looks obviously worse and still does not clear p < 0.05. The dotted line is the one-standard-error rule, a parsimony heuristic, not a significance test. A proper comparison uses a paired test on the per-fold differences, and even that is optimistic because the folds share training data.
Both figures verified against the papers. Obermeyer, Powers, Vogeli and Mullainathan, "Dissecting racial bias in an algorithm used to manage the health of populations", Science 366(6464): 447-453, 2019. doi:10.1126/science.aax2342. "at a given risk score, Black patients are considerably sicker than White patients." Remedying it "would increase the percentage of Black patients receiving additional help from 17.7 to 46.5%." The cause: the algorithm predicted health care costs rather than illness, and less is spent caring for Black patients at the same level of need. Wong, Otles, Donnelly et al., "External Validation of a Widely Implemented Proprietary Sepsis Prediction Model in Hospitalized Patients", JAMA Internal Medicine 181(8):1065-1070, 2021. doi:10.1001/jamainternmed.2021.2626. 27,697 patients, 38,455 hospitalizations, sepsis in 2,552 (6.6%). AUC 0.63 (95% CI 0.62-0.64), "substantially worse than that reported by Epic Systems (AUC, 0.76-0.83)." Sensitivity 33%; "did not identify 1709 patients with sepsis (67%)" while alerting on 18% of hospitalizations.