Neural Network Training Acceleration Based on Hybrid Data and Model Parallelism
Abstract
Nowadays, deep learning models are quickly increasing in complexity and size, making the training or inference phases unsuitable for most office computers or even high-performance computing servers. Examples of such models include, but are not limited to, large language models and image diffusion models, where the number of parameters to adjust can reach hundreds of billions. Although deep learning models are extremely useful tools, the need for powerful computers to train and use them might compromise the obtained results and also reduce accessibility for many researchers. In this work, we propose a general method to reduce the training time and computing resources (central processing unit and memory usage) by the division of a large-size model into several smaller submodels that are easier to handle but do not affect the final performance. This division depends on the original model and architecture, and hence we propose specific strategies for regression and classification problems. The main result of this work is the development of a public Python library called Skynnet to offer this alternative to deep learning computations.
- Referencias
- Cómo citar
- Del mismo autor
- Métricas
Ahmed, N. (1991). How I came up with the discrete cosine transform. Digit. Signal Process, 1, 4-5. https://www.cse.iitd.ac.in/~pkalra/siv864-2019/assignment/DCT-History.pdf
Besta, M., & Hoefler, T. (2024). Parallel and Distributed Graph Neural Networks: An In-Depth Concurrency Analysis. IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(5), 2584-2606. https://doi.org/10.1109/TPAMI.2023.3303431
Brakel, F., Odyurt, U., & Varbanescu, A. L. (2024). Model Parallelism on Distributed Infrastructure: A Literature Review from Theory to LLM Case-Studies. ArXiv. https://arxiv.org/abs/2403.03699
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., & Amodei, D. (2020). Language models are few-shot learners. In: Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M.F., Lin, H. (eds.) Advances in Neural Information Processing Systems, 33, 1877-1901. https://proceedings.neurips.cc/paper/2020/file/1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf
Corchado, J. M., López, S., Núñez, J. M., Garcia, R., & Chamoso, P. (2023). Generative Artificial Intelligence: Fundamentals. ADCAIJ: Advances in Distributed Computing and Artificial Intelligence Journal, 12(1), e31704. https://doi.org/10.14201/adcaij.31704
Dean, J., Corrado, G., Monga, R., Chen, K., Devin, M., Mao, M., Ranzato, M. A., Senior, A., Tucker, P., Yang, K., Le, Q., & Ng, A. (2012). Large scale distributed deep networks. In: Pereira, F., Burges, C. J., Bottou, L., Weinberger, K. Q. (eds.) Advances in Neural Information Processing Systems, 25. https://papers.nips.cc/paper_files/paper/2012/file/6aca97005c68f1206823815f66102863-Paper.pdf
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., & Fei-Fei, L. (2009). Imagenet: A largescale hierarchical image database. In: 2009 IEEE Conference on Computer Vision and Pattern Recognition, 248-255. https://doi.org/10.1109/CVPR.2009.5206848
Deng, L. (2012). The mnist database of handwritten digit images for machine learning research [best of the web]. IEEE Signal Processing Magazine, 29(6), 141-142. https://doi.org/10.1109/MSP.2012.2211477
Fiesler, E., Choudry, A., & Caulfield, H. J. (1990). Weight discretization paradigm for optical neural networks. In: SPIE Proceedings, 1281, 164-173. https://doi.org/10.1117/12.20700
Garcia-Aranda, J. J., Ramos-Diaz, J., Molina-Cardin, S., Larriva-Novo, X., Bustos, A., Galindo, L. A., & Mayo-Garcia, R. (2021). Dynamically distributing tasks from an unattended parallel compiler with cloudbook. In: Nesmachnow, S., Castro, H., Tchernykh, A. (eds.) High Performance Computing, 3-17. Springer, Cham.
Geng, J., Li, D., & Wang, S. (2019). Horizontal or vertical? A hybrid approach to large-scale distributed machine learning. In: Proceedings of the 10th Workshop on Scientific Cloud Computing. ScienceCloud, 19, 1-4. https://doi.org/10.1145/3322795.3331461
Geron, A. (2019). Hands-on Machine Learning with Scikit-Learn, Keras, and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems. O'Reilly Media, Incorporated. https://books.google.es/books?id=OCS1twEACAAJ
Halevy, A., Norvig, P., & Pereira, F. (2009). The unreasonable effectiveness of data. IEEE Intelligent Systems, 24, 8-12. https://doi.org/10.1109/MIS.2009.36
He, Y., Zhang, X., & Sun, J. (2017). Channel pruning for accelerating very deep neural networks. In: 2017 IEEE International Conference on Computer Vision (ICCV), 1398-1406. https://doi.org/10.1109/ICCV.2017.155
Hong, H., Choi, D., Kim, N., Lee, H., Kang, B., Kang, H., & Kim, H. (2024). Survey of convolutional neural network accelerators on field-programmable gate array platforms: architectures and optimization techniques. Journal of Real-Time Image Processing, 21, 64. https://doi.org/10.1007/s11554-024-01442-8
Jacob, B., Kligys, S., Chen, B., Zhu, M., Tang, M., Howard, A., Adam, H., & Kalenichenko, D. (2018). Quantization and training of neural networks for efficient integer-arithmetic-only inference. In: 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2704-2713. https://doi.org/10.1109/CVPR.2018.00286
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., & Amodei, D. (2020). Scaling laws for neural language models. ArXiv. https://arxiv.org/abs/2001.08361
Krizhevsky, A., Sutskever, I., & Hinton, G. E. (2012). Imagenet classification with deep convolutional neural networks. In: Pereira, F., Burges, C. J., Bottou, L., Weinberger, K. Q. (eds.) Advances in Neural Information Processing Systems, 25. https://proceedings.neurips.cc/paper_files/paper/2012/file/c399862d3b9d6b76c8436e924a68c45b-Paper.pdf
Lake, B. M., Salakhutdinov, R., & Tenenbaum, J. B. (2015). Human-level concept learning through probabilistic program induction. Science, 350(6266), 1332-1338. https://doi.org/10.1126/science.aab3050
Lecun, Y., Denker, J., & Solla, S. (1989). Optimal brain damage. Advances in Neural Information Processing Systems, 2, 598-605.
Li, F.-F., Andreeto, M., Ranzato, M., & Perona, P. (2022). Caltech 101. CaltechDATA. https://doi.org/10.22002/D1.20086
Li, M., Zhang, J., Wan, J., Ren, Y., Zhou, L., Wu, B., Yang, R., & Wang, J. (2020). Distributed machine learning load balancing strategy in cloud computing services. Wirel. Netw, 26(8), 5517-5533. https://doi.org/10.1007/s11276-019-02042-2
Liang, T., Glossner, J., Wang, L., Shi, S., & Zhang, X. (2021). Pruning and quantization for deep neural network acceleration: A survey. Neurocomputing, 461, 370-403. https://doi.org/10.1016/j.neucom.2021.07.045
Maas, A. L., Daly, R. E., Pham, P. T., Huang, D., Ng, A. Y., & Potts, C. (2011) Learning word vectors for sentiment analysis. In: Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, 142-150. Association for Computational Linguistics, Portland, Oregon, USA. https://www.aclweb.org/anthology/P11-1015
Mohaidat, T., & Khalil, K. (2024). A Survey on Neural Network Hardware Accelerators. IEEE Transactions on Artificial Intelligence, 5(8), 3801-3822. https://doi.org/10.1109/TAI.2024.3377147
Moreno-Alvarez, S., Haut, J. M., Paoletti, M. E., Rico-Gallego, J. A., Diaz-Martin, J. C., & Plaza, J. (2020). Training deep neural networks: a static load balancing approach. The Journal of Supercomputing, 76(12), 9739-9754. https://doi.org/10.1007/s11227-020-03200-6
Nagajaru, C., Ramesh, Y., & Krishna Mohan, C. (2024). A data parallel approach for distributed neural networks to achieve faster convergence. In Proceedings Volume 13072, Sixteenth International Conference on Machine Vision (ICMV 2023); 130721A. https://doi.org/10.1117/12.3023413
Narayanan, D., Shoeybi, M., Casper, J., LeGresley, P., Patwary, M., Korthikanti, V., Vainbrand, D., Kashinkunti, P., Bernauer, J., Catanzaro, B., Phanishayee, A., & Zaharia, M. (2021). Efficient large-scale language model training on gpu clusters using megatron-lm. In: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. SC '21. Association for Computing Machinery, New York, NY, USA. https://doi.org/10.1145/3458817.3476209
Sharma, R. K., & Singh, S. (2024). Harigeeta: Cic Mechanism with Euclidean Steiner Tree for Service Latency Prediction in Delay-Sensitive Cloud Services. ADCAIJ: Advances in Distributed Computing and Artificial Intelligence Journal, 13(1), e31594. https://doi.org/10.14201/adcaij.31594
Shoeybi, M., Patwary, M., Puri, R., LeGresley, P., Casper, J., & Catanzaro, B. (2019). Megatron-LM: Training multi-billion parameter language models using model parallelism. ArXiv. https://arxiv.org/abs/1909.08053
Singh, S., Ruwase, O., Awan, A. A., Rajbhandari, S., He, Y., & Bhatele, A. (2023). A Hybrid Tensor-Expert-Data Parallelism Approach to Optimize Mixture-of-Experts Training. Proceedings of the 37th International Conference on Supercomputing, 203–214. https://doi.org/10.1145/3577193.3593704
Strubell, E., Ganesh, A., & Mccallum, A. (2019). Energy and policy considerations for deep learning in NLP, In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Korhonen, A., Traum, D., Màrquez, L (eds), 3645-3650. https://doi.org/10.18653/v1/P19-1355
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. U., & Polosukhin, I. (2017). Attention is all you need. In: Guyon, I., Luxburg, U.V., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., Garnett, R. (eds.) Advances in Neural Information Processing Systems, Vol. 30. Curran Associates, Inc. https://proceedings.neurips.cc/paper_files/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf
Villalobos, P., Sevilla, J., Besiroglu, T., Heim, L., Ho, A., & Hobbhahn, M. (2022). Machine Learning Model Sizes and the Parameter Gap. https://arxiv.org/abs/2207.02852
Yingna, Z., Mohd Daud, K., Moorthy, K., & Mohamad Nor, A. N. (2024). Filtering Approaches and Mish Activation Function Applied on Handwritten Chinese Character Recognition. ADCAIJ: Advances in Distributed Computing and Artificial Intelligence Journal, 13(1), e31218. https://doi.org/10.14201/adcaij.31218
Xu, W., Zhang, Y., & Tang, X. (2021). Parallelizing dnn training on gpus: Challenges and opportunities. In: Companion Proceedings of the Web Conference 2021, 174-178. Association for Computing Machinery, New York, NY, USA. https://doi.org/10.1145/3442442.3452055
Zhuang, Y., Zhao, H., Zheng, L., Li, Z., Xing, E. P., Ho, Q., Gonzalez, J. E., Stoica, I., & Zhang, H. (2024). On optimizing the communication of model parallelism. ArXiv. https://arxiv.org/abs/2211.05322
Most read articles by the same author(s)
- Alfonso González, Juan Ramos, Juan F. De Paz, A drug identification system for intoxicated drivers based on a systematic review , ADCAIJ: Advances in Distributed Computing and Artificial Intelligence Journal: Vol. 4 No. 4 (2015)
- Alfonso González, Juan Ramos, Juan F. De Paz, Juan M. Corchado, Obtaining Relevant Genes by Analysis of Expression Arrays with a Multi-Agent System , ADCAIJ: Advances in Distributed Computing and Artificial Intelligence Journal: Vol. 3 No. 3 (2014)