Neural Network Training Acceleration Based on Hybrid Data and Model Parallelism

  • José Javier García
    Innovation Strategies (Spain) jose_javier.garcia_aranda[at]nokia.com
  • Juan Ramos
    Innovation Strategies (Spain)
  • Juan José Guerrero
    Innovation Strategies (Spain)
  • Andrés Bustos
    Centro de Investigaciones Energéticas, Medioambientales y Tecnológicas
  • Rafael Mayo-García
    Centro de Investigaciones Energéticas, Medioambientales y Tecnológicas

Abstract

Nowadays, deep learning models are quickly increasing in complexity and size, making the training or inference phases unsuitable for most office computers or even high-performance computing servers. Examples of such models include, but are not limited to, large language models and image diffusion models, where the number of parameters to adjust can reach hundreds of billions. Although deep learning models are extremely useful tools, the need for powerful computers to train and use them might compromise the obtained results and also reduce accessibility for many researchers. In this work, we propose a general method to reduce the training time and computing resources (central processing unit and memory usage) by the division of a large-size model into several smaller submodels that are easier to handle but do not affect the final performance. This division depends on the original model and architecture, and hence we propose specific strategies for regression and classification problems. The main result of this work is the development of a public Python library called Skynnet to offer this alternative to deep learning computations.

  • Referencias
  • Cómo citar
  • Del mismo autor
  • Métricas
Abadi, M., Agarwal, A., Barham, P., Brevdo, E., Chen, Z., Citro, C., Corrado, G.S., Davis, A., Dean, J., Devin, M., Ghemawat, S., Goodfellow, I., Harp, A., Irving, G., Isard, M., Jia, Y., Jozefowicz, R., Kaiser, L., Kudlur, M., Levenberg, J., Mane, D., Monga, R., Moore, S., Murray, D., Olah, C., Schuster, M., Shlens, J., Steiner, B., Sutskever, I., Talwar, K., Tucker, P., Vanho Cke, V., Vasudevan, V., Viegas, F., Vinyals, O., Warden, P., Wattenberg, M., Wicke, M., Yu, Y., & Zheng, X. (2015). TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems. Software available from tensorflow.org. https://www.tensorflow.org/

Ahmed, N. (1991). How I came up with the discrete cosine transform. Digit. Signal Process, 1, 4-5. https://www.cse.iitd.ac.in/~pkalra/siv864-2019/assignment/DCT-History.pdf

Besta, M., & Hoefler, T. (2024). Parallel and Distributed Graph Neural Networks: An In-Depth Concurrency Analysis. IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(5), 2584-2606. https://doi.org/10.1109/TPAMI.2023.3303431

Brakel, F., Odyurt, U., & Varbanescu, A. L. (2024). Model Parallelism on Distributed Infrastructure: A Literature Review from Theory to LLM Case-Studies. ArXiv. https://arxiv.org/abs/2403.03699

Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., & Amodei, D. (2020). Language models are few-shot learners. In: Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M.F., Lin, H. (eds.) Advances in Neural Information Processing Systems, 33, 1877-1901. https://proceedings.neurips.cc/paper/2020/file/1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf

Corchado, J. M., López, S., Núñez, J. M., Garcia, R., & Chamoso, P. (2023). Generative Artificial Intelligence: Fundamentals. ADCAIJ: Advances in Distributed Computing and Artificial Intelligence Journal, 12(1), e31704. https://doi.org/10.14201/adcaij.31704

Dean, J., Corrado, G., Monga, R., Chen, K., Devin, M., Mao, M., Ranzato, M. A., Senior, A., Tucker, P., Yang, K., Le, Q., & Ng, A. (2012). Large scale distributed deep networks. In: Pereira, F., Burges, C. J., Bottou, L., Weinberger, K. Q. (eds.) Advances in Neural Information Processing Systems, 25. https://papers.nips.cc/paper_files/paper/2012/file/6aca97005c68f1206823815f66102863-Paper.pdf

Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., & Fei-Fei, L. (2009). Imagenet: A largescale hierarchical image database. In: 2009 IEEE Conference on Computer Vision and Pattern Recognition, 248-255. https://doi.org/10.1109/CVPR.2009.5206848

Deng, L. (2012). The mnist database of handwritten digit images for machine learning research [best of the web]. IEEE Signal Processing Magazine, 29(6), 141-142. https://doi.org/10.1109/MSP.2012.2211477

Fiesler, E., Choudry, A., & Caulfield, H. J. (1990). Weight discretization paradigm for optical neural networks. In: SPIE Proceedings, 1281, 164-173. https://doi.org/10.1117/12.20700

Garcia-Aranda, J. J., Ramos-Diaz, J., Molina-Cardin, S., Larriva-Novo, X., Bustos, A., Galindo, L. A., & Mayo-Garcia, R. (2021). Dynamically distributing tasks from an unattended parallel compiler with cloudbook. In: Nesmachnow, S., Castro, H., Tchernykh, A. (eds.) High Performance Computing, 3-17. Springer, Cham.

Geng, J., Li, D., & Wang, S. (2019). Horizontal or vertical? A hybrid approach to large-scale distributed machine learning. In: Proceedings of the 10th Workshop on Scientific Cloud Computing. ScienceCloud, 19, 1-4. https://doi.org/10.1145/3322795.3331461

Geron, A. (2019). Hands-on Machine Learning with Scikit-Learn, Keras, and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems. O'Reilly Media, Incorporated. https://books.google.es/books?id=OCS1twEACAAJ

Halevy, A., Norvig, P., & Pereira, F. (2009). The unreasonable effectiveness of data. IEEE Intelligent Systems, 24, 8-12. https://doi.org/10.1109/MIS.2009.36

He, Y., Zhang, X., & Sun, J. (2017). Channel pruning for accelerating very deep neural networks. In: 2017 IEEE International Conference on Computer Vision (ICCV), 1398-1406. https://doi.org/10.1109/ICCV.2017.155

Hong, H., Choi, D., Kim, N., Lee, H., Kang, B., Kang, H., & Kim, H. (2024). Survey of convolutional neural network accelerators on field-programmable gate array platforms: architectures and optimization techniques. Journal of Real-Time Image Processing, 21, 64. https://doi.org/10.1007/s11554-024-01442-8

Jacob, B., Kligys, S., Chen, B., Zhu, M., Tang, M., Howard, A., Adam, H., & Kalenichenko, D. (2018). Quantization and training of neural networks for efficient integer-arithmetic-only inference. In: 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2704-2713. https://doi.org/10.1109/CVPR.2018.00286

Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., & Amodei, D. (2020). Scaling laws for neural language models. ArXiv. https://arxiv.org/abs/2001.08361

Krizhevsky, A., Sutskever, I., & Hinton, G. E. (2012). Imagenet classification with deep convolutional neural networks. In: Pereira, F., Burges, C. J., Bottou, L., Weinberger, K. Q. (eds.) Advances in Neural Information Processing Systems, 25. https://proceedings.neurips.cc/paper_files/paper/2012/file/c399862d3b9d6b76c8436e924a68c45b-Paper.pdf

Lake, B. M., Salakhutdinov, R., & Tenenbaum, J. B. (2015). Human-level concept learning through probabilistic program induction. Science, 350(6266), 1332-1338. https://doi.org/10.1126/science.aab3050

Lecun, Y., Denker, J., & Solla, S. (1989). Optimal brain damage. Advances in Neural Information Processing Systems, 2, 598-605.

Li, F.-F., Andreeto, M., Ranzato, M., & Perona, P. (2022). Caltech 101. CaltechDATA. https://doi.org/10.22002/D1.20086

Li, M., Zhang, J., Wan, J., Ren, Y., Zhou, L., Wu, B., Yang, R., & Wang, J. (2020). Distributed machine learning load balancing strategy in cloud computing services. Wirel. Netw, 26(8), 5517-5533. https://doi.org/10.1007/s11276-019-02042-2

Liang, T., Glossner, J., Wang, L., Shi, S., & Zhang, X. (2021). Pruning and quantization for deep neural network acceleration: A survey. Neurocomputing, 461, 370-403. https://doi.org/10.1016/j.neucom.2021.07.045

Maas, A. L., Daly, R. E., Pham, P. T., Huang, D., Ng, A. Y., & Potts, C. (2011) Learning word vectors for sentiment analysis. In: Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, 142-150. Association for Computational Linguistics, Portland, Oregon, USA. https://www.aclweb.org/anthology/P11-1015

Mohaidat, T., & Khalil, K. (2024). A Survey on Neural Network Hardware Accelerators. IEEE Transactions on Artificial Intelligence, 5(8), 3801-3822. https://doi.org/10.1109/TAI.2024.3377147

Moreno-Alvarez, S., Haut, J. M., Paoletti, M. E., Rico-Gallego, J. A., Diaz-Martin, J. C., & Plaza, J. (2020). Training deep neural networks: a static load balancing approach. The Journal of Supercomputing, 76(12), 9739-9754. https://doi.org/10.1007/s11227-020-03200-6

Nagajaru, C., Ramesh, Y., & Krishna Mohan, C. (2024). A data parallel approach for distributed neural networks to achieve faster convergence. In Proceedings Volume 13072, Sixteenth International Conference on Machine Vision (ICMV 2023); 130721A. https://doi.org/10.1117/12.3023413

Narayanan, D., Shoeybi, M., Casper, J., LeGresley, P., Patwary, M., Korthikanti, V., Vainbrand, D., Kashinkunti, P., Bernauer, J., Catanzaro, B., Phanishayee, A., & Zaharia, M. (2021). Efficient large-scale language model training on gpu clusters using megatron-lm. In: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. SC '21. Association for Computing Machinery, New York, NY, USA. https://doi.org/10.1145/3458817.3476209

Sharma, R. K., & Singh, S. (2024). Harigeeta: Cic Mechanism with Euclidean Steiner Tree for Service Latency Prediction in Delay-Sensitive Cloud Services. ADCAIJ: Advances in Distributed Computing and Artificial Intelligence Journal, 13(1), e31594. https://doi.org/10.14201/adcaij.31594

Shoeybi, M., Patwary, M., Puri, R., LeGresley, P., Casper, J., & Catanzaro, B. (2019). Megatron-LM: Training multi-billion parameter language models using model parallelism. ArXiv. https://arxiv.org/abs/1909.08053

Singh, S., Ruwase, O., Awan, A. A., Rajbhandari, S., He, Y., & Bhatele, A. (2023). A Hybrid Tensor-Expert-Data Parallelism Approach to Optimize Mixture-of-Experts Training. Proceedings of the 37th International Conference on Supercomputing, 203–214. https://doi.org/10.1145/3577193.3593704

Strubell, E., Ganesh, A., & Mccallum, A. (2019). Energy and policy considerations for deep learning in NLP, In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Korhonen, A., Traum, D., Màrquez, L (eds), 3645-3650. https://doi.org/10.18653/v1/P19-1355

Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. U., & Polosukhin, I. (2017). Attention is all you need. In: Guyon, I., Luxburg, U.V., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., Garnett, R. (eds.) Advances in Neural Information Processing Systems, Vol. 30. Curran Associates, Inc. https://proceedings.neurips.cc/paper_files/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf

Villalobos, P., Sevilla, J., Besiroglu, T., Heim, L., Ho, A., & Hobbhahn, M. (2022). Machine Learning Model Sizes and the Parameter Gap. https://arxiv.org/abs/2207.02852

Yingna, Z., Mohd Daud, K., Moorthy, K., & Mohamad Nor, A. N. (2024). Filtering Approaches and Mish Activation Function Applied on Handwritten Chinese Character Recognition. ADCAIJ: Advances in Distributed Computing and Artificial Intelligence Journal, 13(1), e31218. https://doi.org/10.14201/adcaij.31218

Xu, W., Zhang, Y., & Tang, X. (2021). Parallelizing dnn training on gpus: Challenges and opportunities. In: Companion Proceedings of the Web Conference 2021, 174-178. Association for Computing Machinery, New York, NY, USA. https://doi.org/10.1145/3442442.3452055

Zhuang, Y., Zhao, H., Zheng, L., Li, Z., Xing, E. P., Ho, Q., Gonzalez, J. E., Stoica, I., & Zhang, H. (2024). On optimizing the communication of model parallelism. ArXiv. https://arxiv.org/abs/2211.05322
García, J. J., Ramos, J., José Guerrero, J., Bustos, A., & Mayo-García, R. (2026). Neural Network Training Acceleration Based on Hybrid Data and Model Parallelism. ADCAIJ: Advances in Distributed Computing and Artificial Intelligence Journal, 14, e32609. https://doi.org/10.14201/adcaij.32609

Downloads

Download data is not yet available.

Author Biographies

José Javier García

,
Innovation Strategies (Spain)

Innovation Department, NOKIA Spain, Madrid, 28050, Spain

Juan Ramos

,
Innovation Strategies (Spain)

Innovation Department, NOKIA TECSS, Madrid, 28050, Spain

Juan José Guerrero

,
Innovation Strategies (Spain)

Innovation Department, NOKIA TECSS, Madrid, 28050, Spain

Andrés Bustos

,
Centro de Investigaciones Energéticas, Medioambientales y Tecnológicas

CIEMAT, Madrid, 28040, Spain

Rafael Mayo-García

,
Centro de Investigaciones Energéticas, Medioambientales y Tecnológicas

CIEMAT, Madrid, 28040, Spain

+ −