Researchers have developed a new method to extract an analytical parametrization of the non-perturbative transverse-momentum-dependent (TMD) parton distribution function for unpolarized quarks. This advance is crucial for understanding the internal structure of hadrons, as TMDs describe the probability of finding a parton (quark or gluon) with specific longitudinal and transverse momentum inside a proton or neutron. The novelty lies in combining neural networks and symbolic regression to obtain a compact and precise analytical expression from experimental data.
The process began with training a factorized neural network using cross-section data from fixed-target, Tevatron, RHIC, and LHC Drell-Yan experiments. These data were analyzed at next-to-next-to-next-to-leading logarithmic (NNNLL) accuracy. Subsequently, symbolic regression was applied to each network component. This step allowed for the discovery of compact analytical expressions that describe the complex relationships observed by the neural network, translating the implicit knowledge of AI into explicit formulas.
The final formula was selected from a Pareto front, balancing expression complexity and experimental data fit (χ²). The result is a closed-form non-perturbative function with only nine free numerical constants, achieving a χ²/ndf of 1.040 over 482 data points. A significant finding is the retention of a non-trivial $x$-$b_T$ cross term, even under a sparsity prior that biases it towards zero. This indicates a genuine, albeit mild, correlation between the longitudinal momentum fraction and the transverse momentum of the parton.
This work demonstrates the viability of symbolic regression as a tool to bridge flexible machine-learning fits and interpretable analytical TMD parametrizations. It opens a systematic path toward data-driven discovery of specific features of non-perturbative Quantum Chromodynamics (QCD), which is fundamental for advancing our understanding of strong interactions at low energies and the structure of the proton.