Soutenance de thèse Dimitrios Tzivrailis

Quand

06/11/2026    
14:00 - 17:30

Centre CEA PARIS-SACLAY, Digiteo
salle Amphithéâtre de DIGITEO, 91190 Gif-sur-Yvette

Type d’évènement

Carte non disponible

AI-Driven Monte Carlo Simulations with Learned Energy Surrogates: Uncertainty Quantification and Energy-Based Regularization

Dimitrios Tzivrailis

 

Lieu de la soutenance : Centre CEA PARIS-SACLAY. Digiteo. 91190 Gif-sur-Yvette.
salle Amphithéâtre de DIGITEO

 

Markov Chain Monte Carlo (MCMC) methods are the standard tool for estimating thermodynamic observables in statistical physics and computational chemistry, but their cost is dominated by repeated evaluations of the system’s energy and forces — a cost that becomes prohibitive when these quantities require expensive first-principles calculations such as density functional theory. Machine learning surrogates promise to remove this bottleneck by replacing exact evaluations with fast, learned approximations. This thesis shows that such surrogates introduce two distinct failure modes that undermine the statistical correctness of the resulting simulation, and develops one targeted method to address each. The first failure mode is epistemic uncertainty: a surrogate trained on a finite dataset carries prediction noise that grows as the Markov chain explores regions poorly represented in training. Because the Metropolis acceptance rule is a nonlinear function of the energy difference, even zero-mean noise breaks detailed balance and biases the sampled distribution — a failure invisible to standard regression metrics. To address this, we introduce the Penalty Ensemble Method (PEM), which combines a deep ensemble of surrogates with a noise-penalty correction derived from the framework of Ceperley and Dewing, converting the ensemble’s predictive variance into a term that suppresses acceptance in unreliable regions. Validated on the two-dimensional ϕ4 lattice field theory across its ferromagnetic, paramagnetic, and critical phases, PEM restores the correct stationary distribution while introducing only a modest computational overhead. The second failure mode is rooted in the training objective itself: standard mean-squared-error (MSE) training carries no thermodynamic content and leaves the energy landscape unconstrained outside the training distribution, allowing spurious low-energy minima to form and trap the Markov chain in unphysical configurations. To address this, we introduce Contrastive Regularization for MSE (CRMSE), a training-time correction inspired by Persistent Contrastive Divergence that augments the MSE loss with a contrastive term penalizing energies at configurations visited by a failed surrogate-driven chain. Validated on two molecules from the rMD17 benchmark — ethanol and aspirin — CRMSE eliminates the spurious minima responsible for sampling failure, preserves held-out predictive accuracy, and recovers correct interatomic-distance distributions and dihedral free-energy profiles, including in a data-scarce regime; a complementary validation on the ϕ4 model, presented in an appendix, confirms that the same mechanism corrects the sampling bias in a simpler, analytically tractable setting. PEM and CRMSE form a single progression: PEM establishes that an inference-time correction is viable and shows where it breaks down, while CRMSE removes that dependency altogether and extends the approach to real molecular systems — together making AI-accelerated Monte Carlo sampling both computationally efficient and statistically reliable.

Jury : Claudio Attaccalite (rapporteurs), Aurélien Decelle, Michel Ferrero (rapporteurs), Eiji Kawasaki (co-encadrant de thèse), Alberto Rosso (directeur de thèse),  Véronique Terras, Julien Tranchida

Retour en haut