BiGRU-Transformer GPU Memory Requirement Prediction for Transformer Training Workloads on GPUMemNet

Authors

  • Jordan Smith Computer Science, University of Maryland, College Park, MD, USA Author

DOI:

https://doi.org/10.69987/AIMLR.2026.70110

Keywords:

GPU memory estimation, Transformer training, BiGRU, self-attention, regression, GPUMemNet, resource scheduling, admission control

Abstract

Accurate GPU memory estimation supports scheduling, admission control, and out-of-memory prevention for neural-network training. This study evaluates memory-requirement regression on the GPUMemNet MLP, CNN, and Transformer datasets, with the principal analysis centered on Transformer workloads. The target is peak GPU memory in MiB, and the predictors are restricted to architecture-level information available before execution. The evaluation includes 3,000 MLP records, 8,920 CNN records after removing 80 non-positive target rows, and 5,011 Transformer records. A BiGRU-Transformer feature-token regressor is compared with Ridge, K-nearest neighbors, decision tree, random forest, Extra Trees, and gradient boosting under a fixed 70/15/15 train/validation/test split with seed 42 and target-quantile stratification. On Transformer workloads, the BiGRU-Transformer achieves 1,191.3 MiB MAE, 1,827.7 MiB RMSE, R2 = 0.975, and 86.6% accuracy for 8-GB memory bins. Extra Trees gives the lowest Transformer RMSE at 1,248.8 MiB and also leads on MLP and CNN, indicating that the aggregated GPUMemNet descriptors strongly favor randomized tree ensembles. The neural estimator remains competitive, offers a structured mechanism for modeling interactions among workload descriptors, and provides a practical foundation for future layer-sequence inputs.

Author Biography

  • Jordan Smith, Computer Science, University of Maryland, College Park, MD, USA

     

     

     

Downloads

Published

2026-01-21

How to Cite

Jordan Smith. (2026). BiGRU-Transformer GPU Memory Requirement Prediction for Transformer Training Workloads on GPUMemNet. Artificial Intelligence and Machine Learning Review , 7(1), 136-151. https://doi.org/10.69987/AIMLR.2026.70110

Share