Faster Results from a Smarter Schedule: Reframing Collegiate Cross Country through Analysis of the National Running Club Database

작성자

카테고리:

← 피드로
arXiv cs.AI · Jonathan A. Karr Jr, Ryan M. Fryer, Nitesh V. Chawla · 2026-08-12 AI

[Submitted on 12 Sep 2025 (v1), last revised 10 Aug 2026 (this version, v5)]

View PDF HTML (experimental)

Abstract:Collegiate cross country teams often build their season schedules on intuition rather than evidence, partly because large-scale performance datasets were not publicly accessible prior to the National Running Club Database (NRCD). We analyze the comprehensive-era Cross Country subset of NRCD, 23,355 results from 7,083 athletes (2023-2025; >97% course/weather coverage). Under leakage control and temporal validation, race-result features do not support out-of-year forecasting of individual improvement (best men’s R^2=0.043; women’s -0.029), capturing only a small fraction of the outcome’s reliability ceiling (~0.23-0.28). Against this null, team race frequency associates with nationals placement (pooled RR =2.09; GEE OR =2.56/SD). Program-wide opportunity (roster depth; Effective Racing Opportunity) outranks a single workhorse’s max race count cross-sectionally, but overall team depth for race count is controlled. `Converted Only’ times (not adjusted for weather and elevation) overstate mean first-to-last gains by 15-21 seconds relative to `Standardized’. These results challenge coaching practices that treat schedule design as purely anecdotal and show how NRCD enables evidence-based decision-making in collegiate cross country.

Submission history

From: Jonathan Karr Jr [view email]
[v1] Fri, 12 Sep 2025 17:50:23 UTC (631 KB)
[v2] Tue, 16 Sep 2025 07:27:20 UTC (623 KB)
[v3] Thu, 11 Dec 2025 17:28:17 UTC (2,984 KB)
[v4] Mon, 15 Dec 2025 17:45:12 UTC (2,984 KB)
[v5] Mon, 10 Aug 2026 19:02:14 UTC (228 KB)

원문에서 계속 ↗

추출 본문 · 출처: arxiv.org · https://arxiv.org/abs/2509.10600

코멘트

답글 남기기