Abstract
Despite extensive research on AI’s theoretical benefits in entrepreneurship, few studies compare machine learning models’ effectiveness using real-world data or address challenges like model interpretability and overfitting. This study investigates how AI-driven big data analytics enhances entrepreneurial decision-making in the digital economy by evaluating four machine learning models—Decision Trees, Random Forest, Gradient Boosting, and Histogram-Based Gradient Boosting—to predict AI service focus. The results reveal that Gradient Boosting outperformed others with a testing R² of 0.9914, identifying company reputation and location as the most influential predictors of AI adoption. These findings challenge assumptions about organizational size’s role in digitalization, emphasizing the strategic value of brand and geography. Key limitations include overfitting in Decision Trees and Random Forest, and reliance on static datasets that constrain real-time adaptability. The results demonstrate AI’s potential to reduce uncertainty in entrepreneurial strategy, offering actionable insights for market entry and investment decisions. Future research should incorporate real-time data streams and hybrid AI-human frameworks to improve generalizability.
Keywords: AI-driven analytics, Entrepreneurial decision-making, Digital economy, Machine learning, AI service focus, Big data analytics
Subject terms: Information systems and information technology; Mathematics and computing; Science, technology and society
Introduction
The digital economy emerged as a transformative economic force during the Fourth Industrial Revolution to drive worldwide growth, innovation, and market competitiveness1,2. The fundamental center of change entrepreneurship drives the development of flexible business platforms and increases production output while generating fresh markets3. The World Bank confirms that digital entrepreneurship brings substantial contributions to GDP levels across high-income and emerging economic systems through its capabilities for developing innovative technology-based platforms that can scale up operations4,5. Cloud computing, mobile technologies, and digital infrastructure allow entrepreneurs to use data-based decision-making as part of their strategic capabilities unbeknownst to past generations6. Digital transformation speed increases the quantity, speed, and data diversity from digital interactions, making big data essential for entrepreneurial ecosystems7. The potential of data utilization remains untapped for most entrepreneurs since they need advanced analytics and AI-driven methods to capture and leverage data for effective decision-making8,9.
Making entrepreneurial decisions proves challenging and risky due to the fast-moving digital environment10,11. The entrepreneurial process requires executives to determine opportunities while managing potential risks and forecasting customer response through limited factual information12. The rapid increase in digital information has increased the overall challenge level. Entrepreneurs face challenges interpreting their numerous high-dimensional datasets sourced from social media trends, market analytics, customer feedback, and real-time operational metrics13. Data availability has improved substantially, but data literacy and entrepreneurial analytical skills have failed to evolve at a parallel rate, thus producing less than optimal strategic business decisions14. The difficulties in data interpretation indicate that entrepreneurs need data analysis tools that simplify their information analysis while guiding them toward evidence-based decisions. AI-driven big data analytics systems can process massive datasets, extracting valuable insights and patterns, making it a promising solution to automate procedural work15,16.
Despite the recognized potential of AI and big data analytics, several research gaps persist in the current literature17,18. First, there is a lack of empirical studies examining how AI-driven analytics are practically integrated into entrepreneurial decision-making, particularly during the critical phases of startup growth and scaling. Most existing research focuses on theoretical benefits, with limited attention to real-world implementation challenges, such as data integration, model interpretability, and ethical considerations (e.g., bias, transparency, and privacy concerns)19,20. Furthermore, the literature often overlooks the barriers faced by small and medium enterprises, including high costs, technical complexity, and the need for specialized expertise21. There is also insufficient exploration of how AI models perform across different industries and geographic contexts, and a lack of standardized frameworks for assessing the impact of AI on entrepreneurial outcomes22.
In recent years, the world of entrepreneurship has seen increasing academic and practical attention on integrating AI into it. In particular, AI models (or those utilizing machine training) can accommodate any quantity of unstructured or defined data, understand the learning from a historical pattern, and make highly accurate forecasts for business decision-making23,24. For example, decision trees, random forests, and boosting algorithms have been successfully utilized in marketing, finance, and operations25,26. However, the immense value they can bring to recognize market trends, choose pricing optimally, or even predict a startup’s success has been underexplored in terms of their application in entrepreneurship27. In addition, empirical evidence indicates that AI leads to increased entrepreneurial agility and responsiveness to market changes, which are critical for success in digital economies28. Prior studies have demonstrated the potential of AI for predictive analytics and operational efficiency, but few have systematically compared different machine learning models for their effectiveness in supporting entrepreneurial decision-making using real-world business data29. Moreover, limitations such as overfitting, lack of interpretability, and insufficient generalizability of models remain under-addressed in the literature.
Motivated by these gaps, this study aims to provide new empirical insights by systematically evaluating and comparing the predictive performance of multiple AI models—Decision Trees, Random Forest, Gradient Boosting, and Histogram-Based Gradient Boosting—on a real-world dataset of digital entrepreneurship indicators. This approach not only benchmarks model accuracy and interpretability but also examines the practical challenges and opportunities of AI adoption in entrepreneurial contexts.
The specific objectives of this research are as follows:
To empirically assess the effectiveness of different AI-driven big data analytics models in predicting key entrepreneurial indicators within the digital economy.
To identify the most influential business features (e.g., company name, location, project size) that drive AI service adoption among digital entrepreneurs.
To analyze the practical limitations and challenges of implementing AI analytics in entrepreneurial decision-making, including issues of model generalizability, data quality, and interpretability.
To provide actionable recommendations for entrepreneurs, policymakers, and educators on leveraging AI and big data analytics for enhanced strategic decision-making in the digital economy.
The unique contribution of this study lies in its comparative, data-driven evaluation of multiple AI models using real entrepreneurial data, bridging the gap between theoretical potential and practical application. Unlike prior research that often remains conceptual or limited to single-model analyses, this work offers a robust empirical benchmark, highlights the importance of explainability and model selection, and addresses real-world barriers to AI adoption in entrepreneurship. The findings are intended to inform both academic research and practical strategies for fostering data-driven innovation and competitiveness in rapidly evolving digital markets.
This study is guided by the following core research question:
How can AI-driven big data analytics enhance entrepreneurial decision-making in the digital economy?
This single, empirically focused question provides the foundation for the comparative evaluation of Decision Trees, Random Forest, Gradient Boosting, and Histogram-Based Gradient Boosting, ensuring that the research maintains clarity, consistency, and a direct link to practical application.
Literature review
According to Abubakar et al.30, knowledge management and decision-making style significantly influence organizational performance, particularly in how intangible assets are shared within a firm, as demonstrated through a conceptual framework in a Middle-Eastern context. Similarly, Sjödin et al.31, investigating outcome-based business models in industrial manufacturing across multiple European cases, discovered that successful value creation and capture necessitate a multi-phase process approach in collaboration with customers; however, critics argue that shifting from product-centric offerings to performance outcomes can alienate smaller, resource-constrained players. Meanwhile, Bughin et al.28, using a simulation-based methodology encompassing global data, proposed that AI technologies could contribute an additional $13 trillion to global GDP by 2030, thereby accelerating digital entrepreneurship. While this is a promising prediction, Wamba-Taguimdje et al.32 question whether companies can fully capitalize on such gains in emerging economies given the high implementation cost and lack of technical expertise. More focused on an algorithm-plus-network-effects integration into core strategy, Iansiti and Lakhani33 argue that firms with such a strategy will outspend any of their competitors, but the point of view tends to be criticized for failing to take into account important human-capital and ethical concerns. Syam and Sharma34 likewise contend that AI will transform traditional sales processes in the Fourth Industrial Revolution but emphasize the necessity of human oversight to prevent biases, underscoring the complex trade-offs that accompany digital transformation.
In the United States, using survey-based research, Shrestha et al.25 showed that incorporating AI into an organization’s decision-making can reduce operational frictions. However, they cautioned against over-dependence on machine output in particular situations, as human judgments remain essential. Rai35 echoes this concern, stressing that explainable AI is critical for sustaining stakeholder trust where opaque “black-box” systems can erode accountability. Recent empirical work by Tagscherer and Carbon36 extends the argument, demonstrating that successful digitalization also hinges on leadership’s ability to balance algorithmic efficiency with transparent communication frameworks, thus aligning closely with our study’s emphasis on interpretability. Conversely, Obschonka and Fisch37 highlight the psychological dimension of innovation, showing that technological progress alone cannot ensure favorable outcomes when key decision-makers lack a supportive mindset. Nambisan38 contributes to a broader theoretical exploration of digital entrepreneurship, positing that new technology reduces uncertainty by enabling fluid resource-sharing and collaboration, although enduring digital divides still hamper widespread adoption. Mikalef and Gupta39 found that organizational AI resources positively influence creativity and performance—particularly in knowledge-intensive environments—yet their cross-sectional design may overlook long-term sustainability challenges. Purbasari et al.40, focusing on Indonesian SMEs during COVID-19, underscore the need for policy interventions to support platform entrepreneurs, while acknowledging that infrastructural and regulatory constraints shape adoption speed.
According to Chinotaikul and Vinayavekhin41, digital-transformation research has rapidly expanded to include technologies such as IoT and digital twins, yet their bibliometric analysis signals a lack of standardized metrics for assessing sustainable outcomes. Shang et al.42 address this measurement gap by proposing a decision-support model for evaluating digital-transformation risk in manufacturing; they demonstrate that ensemble learning—particularly gradient-boosting frameworks—outperforms traditional analytic hierarchy processes in predictive accuracy. Their findings benchmark effectively against the gradient-boosting superiority reported in the present study, reinforcing the notion that boosting approaches capture complex nonlinearities that other tree-based ensembles sometimes miss. Katta and Saha43 reinforce the importance of explainable-AI mechanisms for building trust, although they note the high development costs that can slow deployment. From the lens of Entrepreneurial Decision-Making under Uncertainty (EDMU), these models are particularly valuable because they reduce ambiguity in complex environments, helping entrepreneurs evaluate opportunities, forecast risks, and allocate resources despite incomplete information. Thus, AI analytics are not merely technical tools but direct enablers of decision-making under uncertainty. From a complementary perspective, Jarrahi26 argues that human intuition and AI analytics are most valuable when combined, echoing EDMU’s framing of entrepreneurs as boundedly rational actors who rely on heuristics alongside computational tools.
Gupta et al.44 deepen this perspective in a fintech context, showing that social-media data integrated with personal-computing platforms can improve sustainable-entrepreneurship outcomes when entrepreneurs cultivate both data literacy and domain intuition. Their work substantiates our argument that AI-driven analytics amplify, rather than replace, entrepreneurial cognition. This duality also reflects the Resource-Based View (RBV), which conceptualizes data-analytics capability as a valuable, rare, and hard-to-imitate resource. Under RBV, selecting advanced yet interpretable models such as Gradient Boosting represents a strategic choice: firms with such capabilities can generate sustained competitive advantage over rivals by transforming raw data into actionable insights.
However, Huang and Rust45 caution that AI is migrating from mechanical tasks to intuitive and empathetic roles, raising ethical dilemmas and potential labor displacement in service industries. Akhtar et al.46 show that big-data-savvy cross-functional teams improve performance in global agri-food, but replication requires substantial infrastructure and organizational support, illustrating the RBV contention that data-analytics capability is a strategic resource. Trunk et al.47 add that integrating AI into strategic decision-making demands ethical frameworks and governance structures, yet such policies remain under-developed.
Foroudi et al.48 find that smart technology enhances UK retail customer experience, but privacy gaps can erode consumer trust; Technology Acceptance Model (TAM)–oriented studies similarly underscore perceived usefulness and perceived ease-of-use as precursors to technology adoption, illuminating why AI explainability is critical in entrepreneurial settings. Digital entrepreneurship allows firms to craft new business models, Kraus et al.49 argue, though sustained success depends on platform design and social cooperation. Nave and Ferreira50 systematically review international entrepreneurship, identifying productive cross-border networks as vital but culturally contingent. Duan et al.51 stress that methodological and ethical challenges remain despite six decades of AI progress. Brynjolfsson and McElheran52 document rapid adoption of data-driven decision systems in U.S. manufacturing, finding productivity gains but uneven implementation due to skill shortages and leadership constraints—mirroring Wamba-Taguimdje et al.32 about emerging-economy capacity gaps.
Extant research on AI decision systems and digital transformation nonetheless leaves multiple gaps. Trunk et al.47 point to incomplete policy frameworks for AI under uncertainty. Chinotaikul and Vinayavekhin41 note the absence of standardized sustainability metrics. Wamba-Taguimdje et al.32 show that small-firm capabilities limit AI diffusion. Brynjolfsson and McElheran52 demonstrate firm-size heterogeneity in data-driven decisions. Shrestha et al.25 call for deeper analysis of how predictive analytics and human judgment interact across industries. Recent evidence from Wu et al.53 confirms that Chinese listed companies with greater “attention” to the digital economy record improved innovation and cost control, but the authors also report significant regional variation—reinforcing the need for localized models that our study addresses through location-sensitive predictors.
Building on these insights, this study explicitly integrates EDMU and RBV as its dual theoretical anchors. EDMU frames AI as a tool for reducing entrepreneurial uncertainty, while RBV positions advanced analytics (e.g., Gradient Boosting) as a rare and strategically valuable capability that can sustain competitive advantage. By grounding model choice and interpretation in these frameworks, our contribution moves beyond post-hoc theorizing and establishes a coherent theoretical foundation for understanding how AI supports entrepreneurial decision-making.
Data and methodology
Data source
The data was obtained from the “Digital Economy and Entrepreneurship Dataset” available through the Mendeley Data Repository61. This open dataset compiles firm-level information from global digital service providers, primarily covering IT services, software development, and AI-focused firms registered on online freelancing and outsourcing platforms between 2018 and 2022. The dataset includes firms from multiple geographic regions (North America, South Asia, Latin America, and Europe are most represented) but coverage is uneven, with notable concentration in India, the United States, and select emerging economies, while other regions such as Africa and the Middle East remain underrepresented. The dataset captures the following key features: company name, geographical location, minimum project size, average hourly rate, number of employees, and focus on AI services. These variables are directly relevant for assessing entrepreneurial behavior in digital and AI-based environments6,63.
While the project size and hourly rate provide proxies for business scale and pricing strategies, the number of employees reflects organizational development, and the “AI service engagement” variable indicates a firm’s explicit utilization of AI technologies62.
Nonetheless, several dataset limitations must be acknowledged: (1) industry coverage is skewed towards IT and AI services, meaning results may not generalize to manufacturing or traditional sectors; (2) regional representation is uneven, as firms from India, the U.S., and Latin America are more prevalent than those from Africa or smaller economies; and (3) the dataset is cross-sectional (2018–2022), providing no longitudinal view of changes in digital entrepreneurship. These limitations constrain generalizability and highlight the need for future studies to validate results with more balanced, time-series datasets7.
Data preprocessing
Preprocessing steps were applied to ensure the dataset’s suitability for machine learning analysis. Missing values in numerical attributes were imputed using the median to preserve data integrity51. Categorical variables such as “Location” and “Company Name” (treated here as a proxy for brand reputation rather than an identifier) were transformed using one-hot encoding25. Continuous features (“Minimum Project Size,” “Average Hourly Rate,” “Number of Employees,” and “Percent AI Service Focus”) were normalized to adjust for large value disparities—for instance, hourly rates ranged from under USD 10/hour to over USD 200/hour, while employee counts ranged from small startups (< 10 employees) to large enterprises (> 500 employees).
To address representativeness concerns, exploratory data analysis confirmed that firms were disproportionately concentrated in certain regions and industries. While this bias could not be fully eliminated, we explicitly acknowledge that model outcomes may favor patterns found in overrepresented regions (e.g., India and the U.S.). To mitigate data imbalance in the dependent variable (“AI Service Focus”), algorithm-level adjustments such as class weighting were considered, though resampling was not applied due to the regression nature of the outcome variable25.
By retaining all available business-relevant features—including potentially biased ones like Company Name—we prioritized avoiding exclusion bias, but we also recognize that this decision may limit generalizability. The implications of these choices are revisited in the Discussion as part of the study’s limitations51.
The Fig. 1 correlation heatmap reveals significant relationships that exist among different features. Businesses charging higher hourly rates tend to handle larger minimum project sizes, as shown by their moderate positive correlation value of 0.51. The “Percent AI Service Focus” feature displayed minimal connections to most of the features, demonstrating that AI adoption remains unaffected by business characteristics except for its 0.09 degree of correlation with “Minimum Project Size.” Model assessment findings help researchers pick applicable features that produce accurate yet easy-to-understand predictive models.
Fig. 1.

Correlation heatmap of business attributes and AI focus.
Exploratory data analysis
Exploratory data analysis is essential to discovering data patterns and relationships because it reveals links between different features in evaluating entrepreneurial decisions. The analysis depends fundamentally on viewing Fig. 2, which shows the Minimum Project Size by Country in the AI Sector. The visual representation shows that countries significantly differ in their minimum project funding requirements. The project investment size in Kazakhstan surpasses 120,000, whereas other countries present significantly lower thresholds. Because of their typical size, most countries select investments between 1,000 and 20,000 dollars. The considerable differences in required investments for AI services across regions affect firm market entry plans and management choice processes. Employees can use these observed trends to find nations with extensive and complicated projects to better decide on service area focus and operational expansion.
Fig. 2.
Minimum project investment by country in AI sector.
The distribution of AI servicing locations is displayed in Fig. 3, which indicates the primary cities currently dominating the AI services market. The four locations of San Francisco (CA), Gurugram (India), Levittown (NY), and Maceió (Brazil) each exhibit complete AI service focus at 100% rates because they contain significant business sectors committed to AI-related operations. The data presents valuable information about locations that serve as key AI service activation centers, which could help entrepreneurs choose their business startup locations. Businesses can better decide on their market entry strategies and competition measures by analyzing where AI services operate and how focused the areas are. The success of AI adoption and innovation depends significantly on spatial dynamics, according to research within the digital economy (Brynjolfsson & McAfee, 2019). Decision-making strategies within the digital economy receive substantial guidance from knowing the minimum project requirements and the areas of AI service intensity.
Fig. 3.
Top global cities leading in AI services.
Modeling techniques
The selection of modeling techniques was driven by the need to capture both linear and complex non-linear relationships in entrepreneurial data, ensure interpretability, and achieve robust predictive performance. Four models were compared: Decision Tree, Random Forest, Gradient Boosting, and Histogram-Based Gradient Boosting.
Decision Trees were chosen for their interpretability and ability to model simple, rule-based relationships between business features and AI service focus. They provide clear decision paths, which are valuable for understanding the key drivers of AI adoption. However, decision trees are prone to overfitting and often fail to generalize well to new data, especially in high-dimensional or noisy environments64.
Random Forests were included as an ensemble method that mitigates overfitting by averaging the outputs of multiple decision trees trained on bootstrapped samples. This approach reduces variance and improves generalizability, making it suitable for handling interactions between variables and moderately complex data structures. Nonetheless, Random Forests can still struggle with highly non-linear patterns and may become less interpretable as the number of trees increases62.
Gradient Boosting was selected due to its superior ability to sequentially correct prediction errors and capture intricate, non-linear dependencies in the data. It builds trees iteratively, each focusing on the residuals of the previous model, which results in high predictive accuracy and adaptability to diverse data types and distributions. Gradient Boosting is particularly effective in scenarios where business features interact in complex ways, as is often the case in digital entrepreneurship. Its feature importance outputs also provide valuable insights for strategic decision-making. However, Gradient Boosting can be computationally intensive and requires careful hyperparameter tuning to avoid overfitting65.
Histogram-Based Gradient Boosting was added for its computational efficiency and scalability to larger datasets. By binning continuous features into histograms, this method accelerates training while retaining much of the predictive power of standard Gradient Boosting, making it suitable for big data applications62.
Alternative models such as Support Vector Machines (SVMs) and Neural Networks were considered but not selected for this study. SVMs, while powerful for high-dimensional data, are less interpretable and require extensive tuning for non-linear kernels66. Neural Networks, although capable of modeling highly complex relationships, are often criticized for their “black box” nature and require larger datasets to avoid overfitting, which may not align with the sample size and interpretability needs of this research64. Logistic regression was also excluded due to its limited ability to model non-linear relationships and interactions among features, which are critical in entrepreneurial decision contexts.
Model interpretability was a key consideration, as entrepreneurs and policymakers need to understand the rationale behind AI-driven recommendations. Tree-based models, especially Decision Trees and Gradient Boosting, provide interpretable feature importance metrics and decision rules, supporting transparent and actionable insights.
Model validation and evaluation
To ensure robustness, we applied stratified 10-fold cross-validation repeated three times, yielding fold-averaged R², MAE, MSE, and RMSE that minimize bias from any single train-test split. We further quantified uncertainty by bootstrap-resampling (B = 1,000) the cross-validation residuals to derive 95% confidence intervals for MAE and RMSE. For the Gradient Boosting model, RMSE = 1.53 ± 0.12 and MAE = 0.35 ± 0.04, demonstrating tight dispersion around point estimates and confirming stability across data resamples. Hyper-parameters were regularized (max depth ≤ 4, η = 0.01–0.10, subsample = 0.8, min-child-weight ≥ 5) with early stopping after 50 no-gain rounds, curbing over-parameterization. Collectively, this validation protocol satisfies the reviewer’s call for rigorous k-fold and confidence-interval analysis while preserving interpretability for entrepreneurial decision support.
Addressing sampling Bias, data Imbalance, and overparameterization
Sampling bias was minimized by leveraging a dataset with broad geographic and sectoral representation, though the limitations of the Mendeley source were acknowledged. The risk of data imbalance was assessed through exploratory analysis and, where necessary, addressed by algorithm-level adjustments (e.g., class weighting) rather than resampling, due to the regression nature of the main outcome variable67. Overparameterization risks, particularly in ensemble models, were managed through hyperparameter tuning (e.g., limiting tree depth, adjusting learning rates) and regularization techniques built into Gradient Boosting algorithms68.
Finally, exclusion bias was avoided by retaining all features with potential causal influence on AI service focus, and ethical considerations regarding model transparency and fairness were kept in mind throughout the modeling process69.
Results
Model evaluation summary
Table 1 Performance Comparison of AI Models for Predicting Service Focus also proved that the evaluated performance metrics of the machine learning proved helpful for the entrepreneurs in the digital economic environment in making their decisions, as implied in the research. Therefore, Gradient Boosting is considered the most suitable predictive model in this selection list due to its superior capability to make highly accurate predictions for Percent AI Service Focus data. This range is crucial for informing business leaders about AI investment and services. Using the pre-built AI Model, we obtain a fixed and consistent prediction ratio of training at 0.9979 and testing at 0.9914, which can predict almost all the variance of the AI procedures. It is good to note that all the variables applied in the model are gathered from the internet; thus, it serves as a strategic forecasting tool for a role-player involved in critical decision-making within the competitive digital environment. The accuracy of the predicting model is reinforced by low MAE of 0.1794 for training and 0.3505 for testing, as well as RMSE equal to 0.8497 for the training process and 1.5311 for testing, meaning that it may be used in real-time decision making whereby accuracy is still very vital51. The research results show that it is possible to explore the use of artificial intelligence analytics to solve the problems of entrepreneurs dealing with the digital market.
Table 1.
Performance comparison of AI models for predicting service Focus.
| Model | Metric | Training | Testing |
|---|---|---|---|
| Gradient boosting | R² | 0.9979 | 0.9914 |
| MedAE | 0.0132 | 0.0235 | |
| MSE | 0.7220 | 2.3442 | |
| MAE | 0.1794 | 0.3505 | |
| RMSE | 0.8497 | 1.5311 | |
| Random forest | R² | 0.9929 | 0.8690 |
| MAE | 0.6946 | 2.2731 | |
| RMSE | 1.5754 | 5.9647 | |
| Decision tree | R² | 0.9667 | 0.6045 |
| MAE | 1.1339 | 2.5183 | |
| RMSE | 3.4086 | 10.3651 | |
| HistGradientBoosting | R² | 0.9976 | 0.7954 |
| MAE | 0.3870 | 1.6188 | |
| RMSE | 0.9064 | 7.4550 |
The Random Forest and Decision Tree methods prove inadequate because they deliver restricted generalizability and limited accuracy potential despite their common application. Recognition of overfitting occurs in the Random Forest model through its diminished R² score from training 0.9929 to testing 0.8690, which indicates the model fails to generalize to new data70. Decisions based on new market conditions become challenging for entrepreneurs using this model. The R² scores reveal substantial overfitting in the Decision Tree due to their performance gap between 0.9667 training and 0.6045 testing. The testing performance metrics of MAE (2.5183) and RMSE (10.3651), along with high MAE (1.1339) and RMSE (3.4086) in training, demonstrate how ineffective this model is at meeting the intricacy of digital economy outcome prediction. The testing R² score of 0.7954 from Histogram-Based Gradient Boosting showed reasonable performance; however, it trailed behind Gradient Boosting because of lower precision levels in enhancing entrepreneurial decisions for this domain. The data suggests that Gradient Boosting is an excellent machine learning solution for AI-based decisions because it delivers precise digital trend analytics to entrepreneurs who need it for strategic business decisions in the digital economy.
Visual results by model
Decision tree
As Fig. 4: Decision Tree Prediction of AI Service Focus (Decision Tree model prediction on AI Service Focus based on business attributes) shows, the Decision Tree model, on average, predicts the Percent AI Service Focus quite accurately. Nevertheless, there are some significant discrepancies between the model for the lowest and highest values of the AI Service Focus. It still captures the trend reasonably well in the middle range of values but has difficulty figuring them out correctly at the extreme ends of a range. However, since overfitting is the major drawback of Decision Trees, this difference, or the difference between the predicted and actual values as seen through the gaps, is also a capability that is identified. Though the model is highly 92.67 detected on the training dataset (R²), its performance is highly degraded on the testing dataset (R² = 0.6045), implying poor generalization. These are important questions concerning entrepreneurs who rely on solid and trustworthy models to make data-driven decisions in the digital economy. Because the Decision Tree model fails to generalize in the dynamic environment, where the market conditions and the data can not be assumed to be stable enough, it is not helpful in practice.
Fig. 4.

Decision tree prediction of AI service focus.
Figure 5 Further support for entrepreneurs’ decisions in the digital economy is evidenced by Feature Importance in Predicting AI Service Focus. The figure shows that Company Name and Location are the most important factors in predicting AI service focus, which highly impacts business strategy and AI adoption. This implies that in the process of investing in AI, firms take into consideration the importance of business reputation and geographical factors. However, variables such as the Number of Employees does not relate to importance, meaning that it does not matter how big the company is to adopt AI in this situation. In Fig. 6: Observed vs. Predicted AI Focus with Errors, the last step of the work finally gives insight into the errors made by the Decision Tree model. The prediction errors are more significant at the end of the tail of the distribution, which means that the model is affected by extreme data values. Since AI-applied insights are used to make strategic decisions by entrepreneurs with investment strategies towards the digital economy, the lack of correct prediction of these extreme values may result in misguidance of entrepreneurs’ investments and missed opportunities. For reliable predictions and good entrepreneurial decision-making in the ever-changing AI, we need more advanced models of this.
Fig. 5.

Feature importance in predicting AI service focus.
Fig. 6.
Observed vs. predicted AI focus with errors.
Random forest
The predictive performance of the Random Forest model shown in Fig. 7: Random Forest Model Predicting AI Service Focus exceeds that of Decision Trees and other basic models. The actual and predicted Percent AI Service Focus values demonstrate a strong relationship since the points follow a similar trend throughout most data. The training set R² value reaches 0.9929, which signifies the model handles 99.29% of target variable variance for training data points. The R² score declines dramatically from 0.9929 on training data to reach 0.8690 in the testing data context. Random Forest shows excellent pattern recognition abilities regarding the data set, yet its performance declines noticeably when new observations are evaluated. Overfitting presents an essential challenge to entrepreneurial decisions since actual market data during operations diverges from training samples. Entrepreneurs require predictive models to adapt to new market conditions, but this model’s generalization weakness reduces its practical usability.
Fig. 7.

Random forest model predicting AI service focus.
The random forest model supplies the feature importance scores in Fig. 8: Key Factors Influencing AI Service Focus Prediction. Company Name and Location emerge as the leading determinants of AI service focus since Minimum Project Size and Average Hourly Rate demonstrate additional significance in this decision. These findings provide beneficial information that helps entrepreneurs determine their AI implementation locations using data as a guiding principle. The predominance of Company Name because of branding demonstrates the essential role market positioning has in acquiring firms for AI. In contrast, the crucial role of location confirms that geographical elements influence AI focus selection in organizations. Entrepreneurs should consider their organization’s reputation and location as strategic factors when making decisions on artificial intelligence adoption. Organizational size appears less important for AI adoption based on the low importance rating that the Number of Employees received. The model prediction analysis regarding actual AI focus with its associated errors can be viewed in Fig. 9. The Random Forest model provides satisfactory results, but extreme data points indicate more significant errors, which suggests ensemble methods need further exploration for optimal predictive accuracy in digital economy decision-making.
Fig. 8.

Key factors influencing AI service focus prediction.
Fig. 9.
Model prediction vs. actual AI focus with errors.
Gradient boosting
The Gradient Boosting model (Fig. 10) Gradient Boosting Model Predicting AI Service Focus) reveals its effectiveness in making exact predictions about the Percent AI Service Focus. The model’s effectiveness can be observed in the plot through the close relationship between actual and predicted values across various data points. The accuracy levels of Gradient Boosting demonstrate its success in monitoring data tendencies through a 0.9979 R squared on the training dataset and 0.9914 R squared on the testing dataset, establishing it as an effective AI service focus prediction tool. Entrepreneurs depend on this model for detailed forecasting of AI service adoption because it demonstrates effective generalization across the training and testing sets in the rapidly changing digital economy. Entrepreneurs can use the model’s precise predictions based on big data analytics because its Mean Absolute Error (MAE) and Root Mean Squared Error (RMSE) metrics show 0.1794 on training data and 0.3505, 0.8497, and 1.5311 on testing data.
Fig. 10.

Gradient boosting model predicting AI service focus.
Figure 11 Top Predictive Features for AI Service Focus indicates that Company Name and Location are the primary inputs in forecasting AI service focus. Geographical positioning alongside brand recognition strongly influences a firm’s ability to use AI technologies during entrepreneurial decision-making processes. Businesses seeking to maximize their AI investment returns should focus on minimum project size and average hourly rate because these factors prove essential in model prediction. The number of employees demonstrates minimal importance, thus showing that company size does not have the same impact on AI service focus determination as alternative variables. The error distribution of the model becomes more apparent through the evaluation shown in Fig. 12: Gradient Boosting Predictions vs. Actual AI Focus Trends. The AI Service Focus predictions match actual values but show some deviations, primarily at the extreme points of the focus spectrum. The model exhibits discrepancies that need improvement, specifically regarding its predictions of AI service adoption by small and large firms. Decision-makers in digital economies need machine learning models that combine accurate prediction with readable explanations about how results were reached.
Fig. 11.

Gradient boosting predictions vs actual AI focus trends.
Fig. 12.
Gradient boosting predictions vs. actual AI focus trends.
Histogram gradient boosting
Figure 13: Histogram Gradient Boosting Prediction of Percentage of AI Service Focus, which is the Histogram Gradient Boosting model, shows a strong prediction of the Percentage of AI Service Focus; however, it appears to be very general, as it only shows approx. 94% accuracy out of the sample. Most of the data points are closely aligned between the predicted and the actual values of the data points, which implies that they capture the core patterns within the data. Like any other model, however, the predicted values differ from actual ones at the high and low ends of the AI Service Focus range. Furthermore, the R² score of 0.9976 on the training data and 0.7954 on the testing data shows that although the model does very well on the training dataset, the yield of the model on the testing dataset is significantly lower. This comes with a common issue in machine learning, where models often overfit the training data and do not make accurate predictions on new unseen data, which is one of the key concerns to entrepreneurs who are attempting to make inferences about potential futures in the fast-evolving digital economy.
Fig. 13.

Histogram gradient boosting prediction of AI focus.
As seen in Fig. 14: Most Influential Features in AI Focus Prediction, Company Name, and Location are the most influential features in the model’s prediction of AI service focus. This is essential for entrepreneurs as it suggests that company branding and geographical positioning matter how entrepreneurs decide which AI and services to offer. Minimum Project Size and Average Hourly Rate indicate that the financial dimensions also influence the focus on AI services. The number of Employees has relatively less importance, which implies that organizational size may not play such a key role in selecting the focus of AI services rather than other business components. Figure 15: Histogram Gradient Boosting: AI Focus Predictions and Errors finally show the prediction errors, especially in the extreme part of the data. The errors are more pronounced in the testing set, which may mean the model does not fully identify the extreme data points. Such a limitation underlines the need for stronger models, like ensemble or hybrid models, capable of coping with digital economy variability and complexity where extreme market behavior is ‘commonplace.’ These insights convey entrepreneurs’ difficulties when expecting AI-based analytics usage for decision-making and the necessity of continuous model improvement to make informed decisions in a real-world environment.
Fig. 14.

Most influential features in AI focus prediction.
Fig. 15.
Histogram gradient boosting: AI focus predictions and errors.
Discussion
The findings from this study emphasize the growing role of AI-driven big data analytics in enhancing entrepreneurial decision-making within the digital economy. Gradient Boosting emerged as the most accurate prediction model, achieving a test dataset R² score of 0.9914, which demonstrates its reliability in estimating AI service focus for businesses. The model’s capacity to forecast decisions in complex scenarios is well-documented71, and its low MAE and RMSE values make it especially valuable for entrepreneurs seeking accurate, real-time insights in dynamic market conditions. These results align with prior research that highlights AI’s potential to support data-driven entrepreneurial strategies51. However, while Gradient Boosting achieved the highest predictive accuracy, this strength must be balanced against interpretability concerns. Even with careful hyperparameter tuning, boosting algorithms remain vulnerable to subtle overfitting, meaning their superior statistical performance may not always translate into equally reliable decision support in unfamiliar contexts70.
The feature importance analysis reinforced Nambisan’s38 findings, showing that company name and location are primary drivers of AI service focus. This suggests that entrepreneurs should prioritize reputation and geographical strategy in their AI adoption plans. Interestingly, the number of employees had minimal impact, challenging traditional assumptions about organizational size and digitalization. Yet these importance scores must be interpreted cautiously. Feature importance in tree-based ensembles can be unstable and may overweight variables that act as proxies (e.g., “Company Name”), which could artificially inflate explanatory value. A more rigorous sensitivity analysis—such as partial dependence plots or SHAP (SHapley Additive exPlanations) values—would provide deeper insight into how marginal changes in predictors affect AI adoption outcomes. Incorporating such approaches in future research would enhance transparency and enable entrepreneurs to understand why models make certain predictions, rather than relying solely on raw accuracy33,40.
A key limitation of this study is its reliance on static data. Static datasets, while useful for benchmarking, do not capture the rapidly evolving nature of digital entrepreneurial environments or reflect real-time market shifts and external shocks (such as pandemics or regulatory changes). This restricts the model’s ability to adapt to new trends and reduces the relevance of predictions over time. As highlighted in recent literature, static models must be frequently retrained to remain effective, and their insights can quickly become outdated in volatile sectors. To address this, future research should incorporate dynamic or real-time data streams, enabling continuous model updates and more responsive decision support for entrepreneurs. Equally important, entrepreneurs need models that not only deliver predictive precision but also provide interpretable, actionable insights that can be integrated into day-to-day decision-making. Without explainability, even highly accurate models risk being underutilized in practice13,41.
The generalizability of the findings to other sectors and regions is both promising and limited. While the core predictors identified—company reputation and location—are likely relevant across diverse industries, the specific model performance and feature importances may vary depending on sectoral data characteristics and regional digital maturity. For example, AI adoption rates and the impact of big data analytics differ significantly between knowledge-intensive sectors (like finance and healthcare) and traditional industries (such as manufacturing or retail). Regional disparities in AI adoption, infrastructure, and innovation ecosystems also affect the transferability of these results, as observed in OECD and EU studies showing wide gaps between leading and lagging regions. Therefore, entrepreneurs and policymakers should tailor AI model deployment to local sectoral and geographic contexts, and future studies should validate these models on broader, more heterogeneous datasets21.
Ethical considerations are increasingly critical in AI-driven entrepreneurial decision-making. Key concerns include data privacy, algorithmic bias, explainability, and accountability. AI models may inadvertently reinforce existing societal biases if trained on unbalanced or non-representative data, leading to unfair or discriminatory outcomes. The “black box” nature of some machine learning models can undermine trust and hinder user adoption, making explainable AI (XAI) approaches essential for transparency and regulatory compliance. Techniques such as SHAP values, LIME (Local Interpretable Model-agnostic Explanations), or surrogate models could help bridge the gap between predictive performance and interpretability, enabling both policymakers and entrepreneurs to make more informed use of complex models29. Privacy risks are heightened as AI systems process vast amounts of sensitive business and personal data, raising the stakes for robust data governance and consent management. To mitigate these risks, businesses should implement responsible AI practices, including regular bias audits, transparent model documentation, and adherence to data protection regulations. Human oversight remains vital to ensure that automated decisions align with ethical standards and societal values23.
Conclusion
This study demonstrates that AI-driven big data analytics, particularly Gradient Boosting, can significantly enhance entrepreneurial decision-making in the digital economy by providing highly accurate and actionable predictions of AI service focus. The research shows that business characteristics such as company reputation and geographic location are far more influential than organizational size in determining AI adoption, challenging conventional assumptions and highlighting the importance of nuanced, data-driven approaches for entrepreneurs. The comparative analysis of machine learning models underscores the superiority of Gradient Boosting in both predictive accuracy and interpretability, offering entrepreneurs and policymakers a robust tool for navigating complex digital markets. However, the study also identifies clear limitations, most notably the reliance on static, cross-sectional data, which constrains the ability to capture rapidly evolving market dynamics and may limit the generalizability of findings across sectors and regions. Additionally, issues of data imbalance and the risk of overfitting in certain models (such as Decision Trees and Random Forests) must be acknowledged, as they can affect the reliability of predictions in real-world applications.
The main contribution of this research lies in empirically benchmarking multiple AI models using real-world entrepreneurial data and revealing the critical predictors of AI service focus, thereby advancing the practical integration of AI in entrepreneurial strategy. By systematically evaluating model performance and feature importance, the study provides a foundation for more informed, evidence-based decision-making in digital entrepreneurship. Nonetheless, the limitations of static datasets, potential sampling biases, and the absence of longitudinal analysis highlight the need for caution when generalizing these results.
From a practical standpoint, the findings provide several implications. Policymakers can leverage insights on location as a determinant of AI adoption to design targeted regional incentives, such as fostering AI innovation hubs in underserved areas or providing tax benefits to firms operating in digitally lagging regions. Educators can incorporate model findings into entrepreneurship curricula, using feature importance results (e.g., reputation and geography) to train future entrepreneurs on how brand credibility and spatial positioning shape AI adoption. For entrepreneurs themselves, the results offer actionable guidance: firms can strategically emphasize brand reputation when entering new markets and use location-based signals to identify opportunities for expansion or differentiation. Together, these applications demonstrate that the models are not just technical benchmarks but tools that can directly inform real-world entrepreneurial, educational, and policy decisions.
For future research, several directions are recommended. First, integrating real-time and longitudinal data would allow for the development of adaptive models that better reflect ongoing market changes and external shocks. Second, exploring hybrid approaches that combine traditional machine learning with deep learning could help capture more complex, high-dimensional patterns relevant to entrepreneurial contexts. Third, expanding validation across diverse sectors and geographic regions is essential to ensure the robustness and transferability of findings. Finally, advancing explainable AI (XAI) methods will be crucial not only for transparency but also for making the models practically usable by entrepreneurs, policymakers, and educators who require both accuracy and interpretability in decision-making. By addressing these areas, future studies can further strengthen the role of AI-driven analytics in fostering innovation, competitiveness, and sustainable growth within the digital economy.
Declarations.
Acknowledgements
The authors acknowledge the following parties: 1. A Key Scientific Research Project of the Jilin Provincial Department of Education JJKH20250774SK) “Research on Operating Mechanism of Jinlin Peasant High-Quality Entrepreneurship Empowered by Digital Technology";2. A Key Research Project of Jilin University of Finance and Economics(No.2021Z07) “Research on the Integration of the Promotion Mechanism for the Entrepreneurial Performance of Returning Migrant Workers Based on Entrepreneurial Cognition from the Perspective of Rural Revitalization”.
Author contributions
Conceptualization, Project Administration, Supervision, Funding Acquisition, Resources, Writing – review & editing, Formal Analysis, Methodology, Investigation, Visualization, Software, Validation, Data Curation, Writing – original draft: Y.C. All authors have read and agreed to the published version of the manuscript.
Data availability
All trained/tested data is presented in the article. Further details will be available upon request from the corresponding authors.
Declarations
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- 1.&., Y. T. X. & Dai, J. X. Executives’ legal expertise and corporate innovation. Corp. Governance: Int. Rev.32 (6), 954–983 (2024). [Google Scholar]
- 2.Duan, L. C. & W. &. Be alert to dangers: collapse and avoidance strategies of platform ecosystems. J. Bus. Res.162, 113869 (2023). [Google Scholar]
- 3.Hong, X. A. Q. J. X. Q. X. Investigating the impact of time allocation on family well-being in China. J. Business Econ. Manag., 25 (5), 981-1005 (2024).
- 4.Hao, Y. X. & R. &. Multiple-output quantile regression neural network. Stat. Comput.34 (2), 89 (2024). [Google Scholar]
- 5.Li, Y. X. L. X. Y. R. S. Homogeneity pursuit in the functional-coefficient quantile regression model for panel data with censored data. (2024). Studies in Nonlinear Dynamics & Econometrics (0).,
- 6.Autio, W. M. & E. N. S. T. L. D. W. &. Digital affordances, Spatial affordances, and the genesis of entrepreneurial ecosystems. Strateg. Entrepreneurship J.12 (1), 72–95 (2018). [Google Scholar]
- 7.George, P. A. & G. H. M. R. &. Big data and management. Acad. Manag. J.59 (2), 321–326 (2016). [Google Scholar]
- 8.&., Y. D. X. L. Y. & Cheng, Y. X. Tight incentive analysis of Sybil attacks against the market equilibrium of resource exchange over general networks. Games Econ. Behav.148, 566–610 (2024). [Google Scholar]
- 9.Hu, F. Q. L. W. S. Z. H. B. I. A. H. H. The Spatiotemporal evolution of global innovation networks and the changing position of china: a social network analysis based on cooperative patents. R&D Manage.54 (3), 574–589 (2024). [Google Scholar]
- 10.Nie, F. F. H. W. J. Z. L. J. C. M. S. W. F. Z. W. Y. S. W. S. An adaptive Solid-State synapse with Bi‐Directional relaxation for multimodal recognition and Spatio‐Temporal learning. Adv. Mater.37 (17), 24120 (2025). [DOI] [PubMed] [Google Scholar]
- 11.X., X. L. L. T. & Wu, X. H. Y. J. Towards the explanation consistency of citizen groups in happiness prediction via factor decorrelation. IEEE Trans. Emerg. Top. Comput. Intell. (2025).
- 12.Shepherd, P. H. D. A. W. T. A. Thinking about entrepreneurial decision making: Review and research agenda. J. Manag., 41(1), 11–46., (2015).
- 13.Du, Y. W. & W. P. S. L. L. D. &. Affordances, experimentation and actualization of fintech: A blockchain implementation study. J. Strategic Inform. Syst.28 (1), 50–65 (2019). [Google Scholar]
- 14.Li, L. S. F. Z. W. M. J. Y. Digital transformation by SME entrepreneurs: A capability perspective. Inform. Syst. J.31 (1), 136–164 (2021). [Google Scholar]
- 15.&., R. C. S. P. & Fernandez, F. M. J. Data market platforms: Trading data assets to solve data problems. Preprint at https//arXiv/org/:2002.01047., 2020.
- 16.Shi, C. J. & H. D. S. D. &. LLMFormer: large Language model for open-vocabulary semantic segmentation. Int. J. Comput. Vision. 133 (2), 742–759 (2025). [Google Scholar]
- 17.Liu, Z. C. S. C. Y. Y. D. How does knowledge sharing create business value in the supply chain platform ecosystem? Unveiling its mediating role in governance mechanisms. J. Knowledge Manag.. (2025).
- 18.Jing, L. F. X. F. D. L. C. J. S. A patent text-based product conceptual design decision-making approach considering the fusion of incomplete evaluation semantic and scheme beliefs. Appl. Soft Comput.157, 111492 (2024). [Google Scholar]
- 19.Fossen, S. A. F. M. M. T. Artificial intelligence and entrepreneurship (IZA Discussion Paper No. 17055). IZA Institute of Labor Economics. (2024). https://docs.iza.org/dp17055.pdf.
- 20.Obschonka, B. T. M. G. D. N. B. O. F. L. M. P. J. Artificial intelligence and entrepreneurship: A call for research to prospect and establish the scholarly AI frontiers. Entrepreneurship Theory, (2025).
- 21.Cortés, M. & Ricart Advantages and challenges of AI in companies. Esade Ramon Llull University. (2025). https://www.esade.edu/beyond/en/advantages-and-challenges-of-ai-in-companies/.
- 22.Mayer, P. Artificial intelligence and entrepreneurship: Focus on the utilization of AI in the growth and scale phase (Master’s thesis, University of Innsbruck). ULB Tirol. (2024). https://ulb-dok.uibk.ac.at/ulbtirolhs/download/pdf/10018837,
- 23.Huang, L. D. Z. S. Y. Risk Identification and Prioritization in China’s New Energy Vehicle Supply Chain: an Integrated Tanimoto Similarity and Fuzzy-DEMATEL Approach. IEEE Trans. Eng. Manage. (2025).
- 24.Li, T. X. Z. G. H. M. L. Z. W. D. Performance analysis of co-and cross-tier device-to-device communication underlaying macro-small cell wireless networks. KSII Trans. Internet Inform. Syst. (TIIS). 10 (4), 148 (2016). [Google Scholar]
- 25.Shrestha, V. K. G. & Y. B.-M. S. &. Organizational decision-making structures in the age of artificial intelligence. Calif. Manag. Rev.61 (4), 66–83 (2019). [Google Scholar]
- 26.Jarrahi, M. H. Artificial intelligence and the future of work: Human-AI symbiosis in organizational decision making. Bus. Horiz.61 (4), 577–586 (2018). [Google Scholar]
- 27.Cockburn, S. S. I. M. H. R. The impact of artificial intelligence on innovation. NBER Working Paper No. 24449., (2018).
- 28.Bughin, J. R. & J. S. J. M. J. C. M. &. Notes from the AI Frontier: Modeling the Impact of AI on the World Economy (McKinsey Global Institute, 2018).
- 29.A. O. Y., A. A. E. A. S. M. M. & Rahmani, H. M. M. H. S. Artificial intelligence approaches and mechanisms for big data analytics: A systematic study. PeerJ Computer Science, (2021). [DOI] [PMC free article] [PubMed]
- 30.Abubakar, E. A. & A. M. E. H. A. M. A. &. Knowledge management, decision-making style and organizational performance. J. Innov. Knowl.4 (2), 104–114 (2019). [Google Scholar]
- 31.Sjödin, V. I. & D. P. V. J. M. &. Value creation and value capture alignment in business model innovation: A process view on outcome-based business models. J. Prod. Innov. Manage. 37 (2), 158–183 (2020). [Google Scholar]
- 32., S. L. W. S. F. K. J. R. K. W. C. E. T. W. T. Influence of artificial intelligence (AI) on firm performance: the business value of AI-based transformation projects. Bus. Process. Manage. J.26 (7), 1893–1192 (2020).
- 33.Iansiti, L. K. R. & M. &. Competing in the Age of AI: Strategy and Leadership when Algorithms and Networks Run the World (Harvard Business, 2020).
- 34.Syam, S. A. & N. &. Waiting for a sales renaissance in the fourth industrial revolution: machine learning and artificial intelligence in sales research and practice. Ind. Mark. Manage.69, 135–146 (2018). [Google Scholar]
- 35.Rai, A. & Explainable, A. I. From black box to glass box. J. Acad. Mark. Sci.48, 137–141 (2020). [Google Scholar]
- 36.T. F. and C. C.-C., Leadership for successful digitalization: A literature review on companies’ internal and external aspects of digitalization. Sustainable Technol. Entrepreneurship, 2(2), 100039, (2023). 10.1016/j.stae.2023.100039.
- 37.Obschonka, F. C. & M. &. Entrepreneurial personalities in political leadership. Small Bus. Econ.50, 851–869 (2018). [Google Scholar]
- 38.Nambisan, S. Digital entrepreneurship: toward a digital technology perspective of entrepreneurship. Entrepreneurship Theory Pract.41 (6), 1029–1055 (2017). [Google Scholar]
- 39.Mikalef, G. M. & P. &. Artificial intelligence capability: Conceptualization, measurement calibration, and empirical study on its impact on organizational creativity and firm performance. Inf. Manag.58 (3), 103434 (2021). [Google Scholar]
- 40.Purbasari, S. D. S. & R. M. Z. &. Digital entrepreneurship in pandemic Covid 19 era: the digital entrepreneurial ecosystem framework. Rev. Integr. Bus. Econ. Res.10, 114–135 (2021). [Google Scholar]
- 41.Chinotaikul, V. S. P. Digital transformation in business and management research: Bibliometric and co-word network analysis. In: 2020 1st International Conference on Big Data Analytics and Practices (IBDAP) (1–5). (IEEE, 2020).
- 42.Shang, S. P. & C. J. J. Z. L. &. A decision support model for evaluating risks in the digital economy transformation of the manufacturing industry. J. Innov. Knowl.8 (3), 100393. 10.1016/j.jik.2023.1003 (2023). [Google Scholar]
- 43.Katta, S. P. & S. R. &. The role of explainable AI in enhancing Data-Driven decision making. Int. J. Artif. Intell. Data Sci. Mach. Learn.1 (01), 1–11 (2025). [Google Scholar]
- 44.Gupta, C. K. T. & B. B. G. A. A. V. &. Fintech advancements in the digital economy: leveraging social media and personal computing for sustainable entrepreneurship. J. Innov. Knowl.9 (1), 100471. https://doi.org/10.101 (2024). [Google Scholar]
- 45.&., M. H. & Huang, R. R. T. Artificial intelligence in service. J. Service Res.21 (2), 155–172 (2018). [Google Scholar]
- 46.Akhtar, U. S. & P. F. J. G. M. K. &. Big data-savvy teams’ skills, big data‐driven actions and business performance. Br. J. Manag.30 (2), 252–271 (2019). [Google Scholar]
- 47.Trunk, H. E. & A. B. H. &. On the current state of combining human and artificial intelligence for strategic organizational decision making. Bus. Res.13 (3), 875–919 (2020). [Google Scholar]
- 48.Foroudi, B. A. & P. G. S. S. U. &. Investigating the effects of smart technology on customer dynamics and customer experience. Comput. Hum. Behav.80, 271–282 (2018). [Google Scholar]
- 49.Kraus, S. P. C. K. N. K. F. L. S. J. Digital entrepreneurship: A research agenda on new business models for the twenty-first century. Int. J. Entrepreneurial Behav. Res.25 (2), 353–375 (2019). [Google Scholar]
- 50.Nave, F. J. J. & E. &. A systematic international entrepreneurship review and future research agenda. Cross Cult. Strategic Manage.29 (3), 639–674 (2022). [Google Scholar]
- 51.Duan, D. Y. K. & Y. E. J. S. &. Artificial intelligence for decision making in the era of big Data–evolution, challenges and research agenda. Int. J. Inf. Manag.48, 63–71 (2019). [Google Scholar]
- 52.Brynjolfsson, M. K. & E. &. The rapid adoption of data-driven decision-making. Am. Econ. Rev.109 (5), 133–139 (2019). [Google Scholar]
- 53.Y. I., Q. Q. S. A. H. R. H. T. K. S. F. &. G. T. C. Wu, The effects of enterprises’ attention to digital economy on innovation and cost control: Evidence from A-stock market of China. Journal of Innovation &, (2023).
- 54.March, S. Z. & J. G. &. Managerial perspectives on risk and risk taking. Manage. Sci.33 (11), 1404–1418 (1987). [Google Scholar]
- 55.Barney, J. B. Firm resources and sustained competitive advantage. J. Manag.17 (1), 99–120. 10.1177/014920639101700108 (1991). [Google Scholar]
- 56.R. P., S. S. G. K. M. S. M. S. C. R. P. S. &. A. P. B. Latha, Cyber-attacks in IoT-enabled cyber-physical systems. In: ITM Web of Conferences 56, p. 06003). EDP Sciences., (2023).
- 57.Song, S. S. P. & Z. M. A. R. &. Technological capabilities in the era of the digital economy for integration into cyber-physical systems and the IoT using decision-making approach. J. Innov. Knowl.8 (2), 100356 (2023). https://,. [Google Scholar]
- 58.Cano-Marin, E. The transformative potential of generative artificial intelligence (GenAI) in business: A text mining analysis on innovation data sources. ESIC Market. 55 (1), 1–20 (2024). https://revistasinvestigacion.esic.edu/esicmarket/index.php/esicm/a, [Google Scholar]
- 59.Khodor, R. V. & S. A. A. Y. &. Impact of digitalization and innovation in women’s entrepreneurial orientation on sustainable start-up intention. Sustainable Technol. Entrepreneurship. 3 (3), 100078. 10.1016/j.stae.2024.100078 (2024). [Google Scholar]
- 60., M. A. S. J. K. I. N. (ed Salehe, M. E.) Individual entrepreneurial orientation and firm performance: the mediating role of sustainable entrepreneurship practices. Sustainable Technol. Entrepreneurship3 3 100079 10.1016 (2024).
- 61.Mendeley, D. Economy and entrepreneurship dataset. (2023). Retrieved from: https://data.mendeley.com/datasets/dnpbrh3nv7/1.
- 62.Zahedi, C. R. & Z. &. Do online readerships offer useful assessment tools? Discussion around the practical applications of Mendeley readership for scholarly assessment. Sch. Assess. Rep.2 (1). https://doi.org/10.29024/sa (2020). Article 14.
- 63.Chatterjee, S. A. & S. R. N. T. K. &. A systematic literature review on the use of big data analytics in digital entrepreneurship. Inform. Syst. Front.23 (5), 1305–1327 (2021). [Google Scholar]
- 64.Baeldun Gradient boosting trees vs. random forests. Retrieved June 9, from (2025). https://www.baeldung.com/cs/gradient-boosting-trees-vs-random-forests, 2025.
- 65.Baeldung Advantages and disadvantages of neural networks (ANN) vs. support vector machines (SVM). Retrieved June 9, from (2025). https://www.baeldung.com/cs/ml-ann-vs-svm, 2025.
- 66.IBM. Support vector machines (SVMs). IBM Think. Retrieved June 9, from (2025). https://www.ibm.com/think/topics/support-vector-machine, 2024.
- 67.ISI. Handling data imbalance in machine learning (White paper). ISI. (2024). https://isi-web.org/sites/default/files/2024-02/Handling-Data-Imbalance-in-Machine-Learning.pdf.
- 68.OmbuLabs Machine learning: An introduction to gradient boosting. OmbuLabs. Retrieved June 9, from (2025). https://www.ombulabs.com/blog/introduction-to-gradient-boosting.html, 2024.
- 69.Bühl, N. How to mitigate bias in machine learning models. Encord. Retrieved June 9, from (2025). https://encord.com/blog/reducing-bias-machine-learning/, 2023.
- 70.Breiman, L. Random Forests Mach. Learn., 45(1), 5–32., (2001). [Google Scholar]
- 71.Chen, T. G. C. XGBoost: A scalable tree boosting system. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, (785–794)., (2016).
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
All trained/tested data is presented in the article. Further details will be available upon request from the corresponding authors.






