Native Graph Neural Network Training in MillenniumDB

Abstract

Training Graph Neural Networks on graphs that exceed main memory often requires exporting data from graph database systems to external machine learning frameworks. We present a database-native, out-of-core GraphSAGE training pipeline integrated into MillenniumDB and exposed through four GQL procedures. The pipeline materializes reusable mini-batches, distributes features across GPU memory, pinned host memory, and NVMe storage based on access frequency, applies frequency-based tiering to the graph topology as well, and writes the learned embeddings back to MillenniumDB for graph pattern and similarity queries. We evaluate on Cora, ogbn-arxiv, ogbn-products, and ogbn-papers100M (111M nodes, 1.6B directed edges). On a commodity desktop (32 GiB RAM, 16 GB GPU), PyG, DGL, and Neo4j GDS run out of memory on ogbn-papers100M. MillenniumDB completes 50 training epochs in 41.2 minutes after a one-time 95.7-minute preparation, matching the accuracy of the state-of-the-art out-of-core system DiskGNN. An ablation shows that topology tiering speeds up offline sampling by 3.81× without altering the sampled mini-batches.

Publication
In 17th Alberto Mendelzon International Workshop on Foundations of Data Management
Sebastián Ferrada
Sebastián Ferrada
Assistant Professor

Research. Coffee. Lifting.