How Far Can You Grow? Characterizing the Extrapolation Frontier of Graph Generative Models for Materials Science


Polat C., Serpedin E., KURBAN M., Kurban H.

32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.1, KDD 2026, Jeju Island, Güney Kore, 9 - 13 Ağustos 2026, cilt.2-B, ss.11821-11832, (Tam Metin Bildiri)

  • Yayın Türü: Bildiri / Tam Metin Bildiri
  • Cilt numarası: 2-B
  • Doi Numarası: 10.1145/3770855.3818872
  • Basıldığı Şehir: Jeju Island
  • Basıldığı Ülke: Güney Kore
  • Sayfa Sayıları: ss.11821-11832
  • Anahtar Kelimeler: benchmark, crystal generation, equivariant architectures, graph neural networks, quantum chemistry
  • Ankara Üniversitesi Adresli: Evet

Özet

Every generative model for crystalline materials harbors a critical structure size beyond which its outputs quietly become unreliable; we call this the extrapolation frontier. Despite its direct consequences for nanomaterial design, this frontier has never been systematically measured. We introduce RADII, a radius-resolved benchmark of ∼75,000 crystal-derived nanoparticle structures (33 - 11,298 atoms) that treats radius as a continuous scaling knob to trace generation quality from in-distribution to out-of-distribution regimes under leakage-free splits. Each model is conditioned on the target composition and atom count, isolating geometric extrapolation as the evaluation variable. RADII provides frontier-specific diagnostics: per-radius error profiles pinpoint each architecture's scaling ceiling, surface - interior decomposition tests whether failures originate at boundaries or in bulk, and cross-metric failure sequencing reveals which aspect of structural fidelity breaks first. Benchmarking five state-of-the-art architectures, we find that: (i) well-behaved models degrade by ∼13% in global positional error beyond training radii, while divergent models exhibit poor absolute fidelity across scales; local bond fidelity varies sharply across architectures, from negligible degradation to more than 2× error growth; (ii) no two architectures share the same failure sequence, revealing the frontier as a multi-dimensional surface shaped by model family; and (iii) well-behaved models conform to the expected geometric scaling exponent α ≈1/3 whose in-distribution fit accurately predicts out-of-distribution error, making their frontiers quantitatively forecastable. Scaling MatterGen to its published parameter count stabilizes sampling but does not close the extrapolation frontier, while DiffCSP remains unstable even at published scale. These findings establish output scale as a first-class evaluation axis for geometric generative models. The dataset, generation pipeline, and implementations are available at https://github.com/KurbanIntelligenceLab/RADII.