A Comparative Machine Learning, and Econometric Analysis of Macroeconomic Growth.

Authors

  • Muhammad Naseem Department Statistics, Quaid a Azam University. Email: naseem.22413002@stat.qau.edu.pk Author
  • Sadam Hussain Institute of Information Technology QAU. Email: sadam.bsit@gmail.com Author
  • Ahmad Mustafa Department statistics, Quaid a Azam University. Email: mustafahafeez00070@gmail.com Author
  • Junaid Abbas Hailey College of Commerce, University of Punjab. Email: junaidabbasskd@gmail.com Author
  • Baber ALI Université Paris-Est Créteil (UPEC), UPEC Faculty of Science and Technology. Email: baber7377@gmail.com Author
  • Iffat Tahir Department of Mathematics and Statistics. Email: iffattahir818@gmail.com Author
  • Baneen Zehra Department Statistics, Quaid a Azam University. Email: baneenzehra01@gmail.com Author
  • Muhammad Yasin yasinmeershigri@gmail.com Author
  • Muhammad Ashfaq Hassan Babar PhD scholar Gcuf. Email: ashfaqhassan75@gmail.com Author

DOI:

https://doi.org/10.63163/jpehss.v4i1.1715

Abstract

Globalization and macroeconomic integration have substantially transformed the economic structure of developing economies, yet accurately identifying and forecasting the determinants of economic growth remains challenging because macroeconomic relationships are often nonlinear, highly interdependent, and structurally unstable. This study investigates the major macroeconomic determinants of Pakistan’s economic performance over the period 1960–2021 and develops a comparative framework between conventional Multiple Linear Regression (MLR) and Random Forest (RF) regression. Gross Domestic Product (GDP) is considered the principal measure of economic performance, while inflation, per capita income, foreign investment, military expenditure, and gross national expenditure are incorporated as explanatory variables. The analysis combines conventional regression estimation with extensive diagnostic assessment and nonlinear machine-learning techniques to evaluate both statistical validity and predictive robustness. The MLR model produces an exceptionally high in-sample explanatory power, with an R-Squared of approximately 0.9994; however, diagnostic analysis reveals substantial heteroscedasticity and severe multicollinearity among key economic predictors, with variance inflation factors exceeding 29 for several variables. These findings demonstrated that conventional goodness-of-fit measures can provide a misleading assessment of model reliability in the presence of structural deficiencies. In contrast, the random forest (RF) model provides a more robust framework for capturing complex and nonlinear relationships among macroeconomic variables and demonstrates stronger out-of-sample stability. Variable-importance analysis identifies foreign investment, military expenditure, and gross national expenditure as the dominant predictors of GDP, while partial dependence analysis indicates positive but nonlinear relationships characterized by diminishing marginal effects. Per capita income contributes comparatively less to predictive performance. Overall, the findings investegated that predictive accuracy should be evaluated alongside structural validity and diagnostic robustness rather than relying solely on conventional fit statistics. The study highlights the value of machine-learning approaches for macroeconomic forecasting in Pakistan and provides evidence that economic growth is driven by interconnected and nonlinear economic mechanisms, with important implications for evidence-based macroeconomic policy and forecasting.

Downloads

Published

2026-03-30