Abstract
Large language models (LLMs) are increasingly being explored and deployed across recruitment and selection processes, reshaping how hiring decisions are supported, communicated, and justified. Unlike earlier algorithmic hiring tools, LLMs operate through language-mediated interaction, influencing interpretive and evaluative layers of decision-making. This scoping review maps the academic literature on LLMs in hiring to examine (i) where and how these systems are applied across the hiring pipeline, (ii) what forms of evidence and outcomes are assessed, (iii) which risks and mitigation strategies are documented, and (iv) how disciplinary structures shape research focus. Following PRISMA-ScR guidelines, we synthesize research published between 2018 and 2026 across multiple disciplines using a transparent, lexicon-based coding approach. The results reveal a rapidly expanding but uneven literature, characterized by concentration in early hiring stages, selective outcome measurement favoring efficiency and performance, high awareness of ethical risks with limited empirical validation of controls, and structurally constrained interdisciplinarity. The review highlights key gaps and provides a foundation for future interdisciplinary and field-based research on responsible LLM use in hiring.
Keywords: algorithmic decision-making, generative AI, hiring, human resource management, large language models, recruitment, scoping review, socio-technical systems
1. Introduction
Hiring decisions occupy a central place in organizational system. They shape individual careers and influence organizational effectiveness. They also contribute to the reproduction or disruption of social inequality in labor markets. Hiring is often treated as a rational and regulated organizational function. However, several studies report that recruitment and selection outcomes remain systematically stratified by gender, age, ethnicity, and socioeconomic background (Bertrand and Mullainathan, 2004; Neumark, 2018; Pager and Shepherd, 2008; Quillian et al., 2017). These disparities suggest that hiring is not only a technical exercise. It is also a socially embedded process in which power, interpretation, and institutionalized expectations matter.
Hiring can also be seen as a socio-technical decision system in which technologies, formal procedures, human judgment, and organizational norms interact and shape outcomes (Orlikowski and Iacono, 2001; Trist and Bamforth, 1951). The Decisions are rarely the product of neutral assessment alone. They emerge through ongoing processes of interpretation and sense making. These decisions are often shaped by both formal evaluation procedures and subjective judgment in personnel selection processes (Highhouse, 2008). Organizational actors evaluate information, construct narratives about candidates, and justify choices to themselves and others (Weick, 1995). This makes hiring especially sensitive to technologies that intervene in how candidate information is produced, organized, and understood.
The earlier generations of hiring technologies mainly consisted of rule-based applicant tracking systems (ATS) or narrowly trained predictive machine learning (ML) models. LLMs differ in a fundamental way. They operate through language. They are increasingly used to parse and summarize résumés, generate interview questions, simulate candidate interactions, draft evaluative narratives, and documentation and compliance-related work (Bender et al., 2021; Bommasani et al., 2022). LLMs intervene in discursive spaces where meaning, legitimacy, and accountability are constructed. Even when final hiring authority remains with human decision-makers, these systems can influence how candidates are described, compared, and discussed. From a sense making perspective, LLMs shape the frames through which evaluators interpret candidate information (Weick, 1995). It also raises new questions about authority, responsibility, and bias in hiring.
Research on LLMs in hiring has grown rapidly (though fragmented and uneven). The reported contributions are conceptual or exploratory, rely on simulations or synthetic data, focus mainly on system design rather than organizational use in practice (Dietvorst et al., 2015). Existing empirical studies employ heterogeneous outcome measures and disciplinary traditions. This limits comparability and makes cumulative synthesis difficult. Under these conditions, a conventional systematic review or meta-analysis is not appropriate. We therefore adopt a scoping review approach following PRISMA-ScR guidelines (Tricco et al., 2018). Our aim is to map the existing body of research on LLMs in hiring. Scoping reviews are well suited to research areas in which concepts are still evolving. They are appropriate when terminology is unsettled and empirical maturity varies (Arksey and O’Malley, 2005; Munn et al., 2018). These conditions apply to scholarship on LLM-based hiring. Since the widespread diffusion of generative AI systems after 2022, relevant work has appeared across computer science, information systems, management, psychology, social sciences, medicine, and law. This literature draws on different theoretical traditions and prioritize different outcomes. They employ diverse methodological approaches. Bringing this dispersed body of work into a coherent analytical map is both challenging and necessary.
We do not seek to estimate effect sizes or establish causal claims. Instead, we identify where along the hiring pipeline LLMs are deployed and for what purposes. We characterize the types of evidence and study designs used. We document reported risks and proposed mitigation strategies. We also examine the disciplinary composition and temporal development of the field.
This review addresses four research questions.
RQ1: In which hiring stages are LLMs used, and for what tasks?
RQ2: What outcomes are assessed and which study designs dominate literature?
RQ3: What risks are documented, and which mitigation strategies are proposed or empirically tested?
RQ4: How is the literature structured across disciplines and how does disciplinary orientation shape research focus?
The review contributes to literature in three ways. First, it offers a stage-wise socio-technical map of LLM applications across the hiring pipeline, highlighting the stage where language-mediated automation is concentrated (RQ1). Second, it provides a theory-informed typology of evidence and outcomes, revealing systematic misalignments between how hiring is understood theoretically and how LLMs are evaluated empirically (RQ2–RQ3). Third, it analyzes disciplinary structure and epistemic fragmentation, demonstrating how disciplinary boundaries shape research agendas and constrain integrative theory-building (RQ4).
2. Methods
2.1. Protocol and reporting standard
The review was conducted as per PRISMA-ScR reporting guidelines (Tricco et al., 2018). A structured protocol was developed prior to analysis specifying eligibility criteria, search strategy, screening procedures, and data-charting variables. Using the PCC (population, concept and context) framework, eligibility criteria were defined across three interconnected dimensions (Pollock et al., 2023):
Population: Job applicants, candidates, recruiters, HR professionals, hiring managers, and organizations.
Concept: Large language models or generative AI systems explicitly used for hiring-related decision support or automation.
Context: Recruitment, selection, interviewing, and onboarding processes.
Studies were included if they examined, evaluated, or proposed LLM-based tools directly relevant to hiring decisions.
Search was conducted in Scopus, selected for its broad interdisciplinary coverage and consistent bibliographic metadata (Chadegani et al., 2013). The core search strategy combined terms related to LLM (e.g., large language model, generative AI, transformer) with hiring related terms (e.g. recruitment, selection, interview, resume screening) using appropriate discipline search syntax. This review included journal articles and review articles. The time window spanned from 2020 to 2025 reflecting the emergence of transformer-based architecture and their accelerated diffusion following 2022. Detailed search query, PRISMA ScR flowchart and inclusion exclusion table in Supplementary material.
2.2. Screening, charting, and coding
Records were de-duplicated across discipline-specific exports. Data were charted at the paper level, extracting bibliographic metadata (publication year, discipline, outlet), application characteristics (hiring stage, task type), outcomes assessed, risks discussed, and governance mechanisms proposed.
All analytical coding followed a lexicon-based, rule-driven approach:
Coding relied exclusively on explicit terms appearing in titles, abstracts, or author keywords.
Categories were coded using binary indicators (presence/absence) or proportional measures (percentage shares).
No machine learning, topic modeling, or latent inference techniques were applied, in line with scoping review best practice and to maximize interpretability.
Complete lexicon tables used for coding outcomes, risks, and conceptual focus are provided in the Supplementary material.
Results were synthesized using: Descriptive mapping, Cross-tabulations and heat maps, Pre/post-2022 comparisons to capture shifts following the widespread diffusion of generative AI. No causal claims were made. Findings were interpreted as patterns of research attention and emphasis, rather than evidence of effectiveness, harm or organizational impact.
3. Results and discussion
Scoping reviews describe the scope and characteristics of research rather than test causal effects (Arksey and O’Malley, 2005; Tricco et al., 2018). As presented in Figure 1, a sudden surge in number of publications can be clearly seen following the AI diffusion (post 2022). Moreover, looking at the funded studied vs. unfunded studies (based on the funding statement declared in the manuscript), the maximum number of funded studies were reported from south Korea, followed by U. K., China, and Germany (extremes represented by South-Korea and Spain).
Figure 1.
Growth of LLM in hiring research. Source Scopus data (2020–2025). Percentage indicate percentage of funded studies derived from dedication to project in the manuscript.
This trend does not follow the quantitative trend of total number publications (Extremes represented by the USA and Singapore). This difference may reflect variation in national funding priorities, government support for AI research, and institutional incentives for funded projects (Waltman, 2016; Zupic and Čater, 2015). Figure 2 presents the mapping of LLM applications across the hiring pipeline. It reveals a clear concentration of research attention in screening and interviewing stages. This pattern reflects a long-standing organizational tendency to automate tasks perceived as repetitive, high-volume, and operationally costly (Autor et al., 2003).
Figure 2.
Hiring stage at which LLM is used (Gen., generative; Class., classification; Summ., summarization; Interact., Interaction; Assess., assessment). Lexicon coding table- Supplementary file.
From a socio-technical systems perspective, LLMs are primarily positioned within the technical subsystem of hiring. In contrast, the social subsystem, including judgment, accountability, and organizational responsibility, remains comparatively under examined (Baxter and Sommerville, 2011; Trist and Bamforth, 1951). This concentration of LLM use in screening and interviewing suggests that organizations are integrating LLMs primarily to support structured information processing rather than to replace human evaluative authority. This pattern reflects bounded rationality and risk-sensitive technology adoption, where automation is introduced first in tasks perceived as operationally reversible (March and Simon, 1958; Parasuraman and Riley, 1997). Future research should examine whether LLM deployment expands into higher-consequence decision stages such as final selection and onboarding, and how organizational governance structures shape the scope and autonomy of LLM-supported hiring decisions.
Screening-stage applications focus on résumé parsing, ranking, and summarization. These practices frame candidates as structured data objects. Interview-stage applications introduce LLMs into more interactive and interpretive roles. These include question generation and conversational agents. This shift has important theoretical implications, as illustrated in Figure 3. Earlier algorithmic hiring tools mainly functioned as computational decision aids. By contrast, LLMs act as language-producing systems that shape how candidates are represented, evaluated, and discussed. Based on sense making theory, these systems participate in constructing the narratives through which hiring decisions are justified (Weick, 1995).
Figure 3.
Change in LLM applications across hiring stages and task types.
However, research attention to later stages such as selection and onboarding remains limited. This absence suggests an organizational preference to deploy LLMs where decision responsibility can remain diffuse and reversible. The Figure 4 indicates that approximately 70–75% of reported LLM applications in hiring are employer-facing, while 25–30% are candidate-facing. Employer-facing uses mainly involve résumé screening, candidate ranking, summarization, and drafting of evaluation notes. Candidate-facing uses are concentrated in interview chatbots and automated question answering. This pattern suggests that LLMs are currently adopted primarily to support internal decision efficiency rather than candidate experience. Such back-end–first adoption is consistent with socio-technical theory, which predicts that organizations introduce new technologies first in controllable internal processes (Trist and Bamforth, 1951; Orlikowski and Iacono, 2001). The limited visibility of LLMs to candidates also raises concerns for perceived fairness and transparency, which are central to applicant reactions (Gilliland, 1993; Colquitt et al., 2001).
Figure 4.
Candidate-facing vs. employer-facing LLM applications.
The Figure 5 shows that the literature is dominated by conceptual and quantitative studies, which together account for approximately 65–70% of publications. Experimental or simulation-based designs represent about 20–25%, while field-based and mixed-methods studies remain limited at roughly 10–15%. This distribution indicates that research on LLMs in hiring is still largely exploratory and design-oriented, with relatively little evidence drawn from real organizational deployments. Such patterns are typical of early-stage technological fields, where system development and proof-of-concept evaluations precede large-scale field validation (Orlikowski, 2007). The limited presence of qualitative and mixed-methods work also constrains understanding of how LLMs interact with human judgment and organizational context, which are central to hiring as a socio-technical process (Trist and Bamforth, 1951; Dipboye, 2018). The dominance of conceptual and experimental research indicates that the field remains in an early stage of empirical institutionalization, where technological capabilities are advancing faster than organizational adoption research. This pattern is consistent with information systems research trajectories in which theoretical framing and system development precede field-based validation (Orlikowski and Iacono, 2001). Future studies should prioritize longitudinal and field-based research examining how LLM-supported hiring systems operate in real organizational contexts, including their effects on decision accountability, interpretive practices, and institutional legitimacy.
Figure 5.
Study design typology of LLM research in hiring [Pre and Post AI diffusion (2022)].
The Figure 6 indicates that outcome assessment in LLM-based hiring research is concentrated on efficiency and performance, which together account for approximately 55–60% of reported outcomes. Fairness and bias-related outcomes represent about 20–25%, while candidate experience and compliance or governance outcomes remain limited at roughly 15–20% combined. This pattern reflects an instrumental orientation in which technologies are primarily evaluated in terms of speed, cost reduction, and predictive quality (March and Simon, 1958; Orlikowski and Iacono, 2001). In contrast, outcomes central to organizational justice and applicant reactions receive comparatively less attention, despite strong evidence that perceived fairness and transparency shape acceptance of selection systems (Gilliland, 1993; Colquitt et al., 2001).
Figure 6.
Outcome distribution.
The Figure 7 indicates that in conceptual and experimental studies, approximately 60–65% of reported outcomes relate to efficiency and accuracy, while only 15–20% address fairness or governance. In contrast, qualitative and mixed-methods studies devote roughly 45–50% of their outcome focus to fairness, candidate experience, and governance, with the remainder addressing performance-related outcomes. Candidate experience outcomes appear almost exclusively in qualitative and mixed-methods work and represent <10% of outcomes in quantitative or simulation-based studies. This pattern suggests that outcome selection is strongly shaped by methodological choice, with technical designs favoring measurable performance metrics and socially embedded outcomes requiring methods that capture perception and context (March and Simon, 1958; Gilliland, 1993; Trist and Bamforth, 1951). This emphasis on efficiency and performance reflects a technical orientation that may overlook the institutional and perceptual dimensions of hiring decisions. Organizational theory suggests that decision legitimacy and fairness perceptions play a critical role in shaping acceptance and trust in selection systems (Gilliland, 1993; Colquitt et al., 2001; Suchman, 1995). Future research should integrate technical performance evaluation with organizational and behavioral outcomes, including candidate perceptions, decision transparency, and the institutional consequences of LLM-mediated hiring practices.
Figure 7.
Outcome vs. study design matrix.
The Figure 8 shows that bias and discrimination risks are the most frequently discussed, accounting for approximately 35–40% of all reported risk mentions. Explainability and transparency risks represent about 20–25%, followed by privacy and security risks at roughly 15–20%. Reliability-related risks (e.g., hallucination, inconsistency) and accountability risks together account for approximately 15–20%. This distribution indicates that the literature is strongly oriented toward fairness-related concerns, while issues of responsibility and system dependability receive comparatively less attention. Such patterns mirror broader debates in algorithmic decision-making, where ethical risk identification often precedes systematic evaluation of governance and accountability mechanisms (Barocas and Selbst, 2016; Kroll et al., 2017; Raji et al., 2020).
Figure 8.
Distribution of documented risk.
These five empirically identified risk categories align closely with established governance dimensions defined in the NIST AI Risk Management Framework (NIST, 2023), supporting comparability with broader AI governance literature (Table 1). The Figure 9 indicates that most mitigation strategies discussed in the literature remain at an early or conceptual stage, with approximately 60–65% of mitigation mentions describing proposed rather than empirically tested controls. Human-in-the-loop oversight and bias auditing together account for roughly 40–45% of all reported mitigation approaches, while documentation and transparency mechanisms represent about 25–30%. Governance frameworks and organizational policies appear least frequently, at approximately 15–20%. This pattern suggests that while risk awareness is high, mitigation practices are not yet institutionalized or systematically evaluated. Similar gaps between ethical principles and operational governance have been documented in broader algorithmic accountability research (Raji et al., 2020; Veale and Borgesius, 2021).
Table 1.
Alignment between lexicon-derived risk categories and governance dimensions defined in the NIST AI Risk Management Framework (NIST, 2023).
| Risk category identified in this review | Description in hiring context | Corresponding NIST AI RMF risk dimension |
|---|---|---|
| Bias and discrimination | Systematic differences in candidate evaluation due to model behavior or training data | Fairness – Harmful bias and discrimination |
| Explainability and transparency | Limited ability to interpret, understand, or justify model-supported hiring outputs | Transparency and explainability |
| Privacy and security | Risks related to exposure, misuse, or protection of candidate personal data | Privacy and data security |
| Reliability | Inconsistent, incorrect, or hallucinated outputs affecting decision quality | Validity and reliability |
| Accountability | Lack of clear responsibility, oversight, or contestability of LLM-supported hiring decisions | Accountability and governance |
Figure 9.
The mitigation maturity in LLM based hiring.
Figure 10 shows an uneven alignment between risk categories and mitigation maturity. For bias and discrimination risks, approximately 50–55% of studies propose mitigation strategies, but only 15–20% report empirical testing of these controls. For privacy and security risks, around 40% of studies mention mitigations, with fewer than 15% providing evidence of implementation. Explainability risks show a similar pattern, with about 45% proposing transparency mechanisms and roughly 15% evaluating them empirically. Accountability risks display the weakest alignment, with <25% of studies proposing concrete governance mechanisms. This pattern indicates that mitigation efforts largely remain conceptual rather than operational. The gap between risk identification and empirically validated mitigation reflects broader challenges in governing algorithmic decision systems, where accountability mechanisms often lag technological capability (Raji et al., 2020; Veale and Borgesius, 2021). From an institutional perspective, the legitimacy of LLM-supported hiring decisions depends not only on system performance but also on demonstrable governance and oversight mechanisms (Scott, 2014; Suchman, 1995). Future research should empirically evaluate governance practices such as auditing, human oversight, and documentation to determine how they influence decision accountability and organizational trust. Consistent with algorithmic governance research, risk identification currently outpaces evidence on whether safeguards meaningfully change decision processes (Kroll et al., 2017; Raji et al., 2020; Veale and Borgesius, 2021). The governance and accountability risks identified in this review align closely with emerging regulatory frameworks governing AI use in employment. In particular, the European Union Artificial Intelligence Act classifies AI systems used in hiring and employment decisions as high-risk applications subject to requirements for risk management, documentation, transparency, and human oversight (Veale and Borgesius, 2021). These requirements correspond directly to the risks identified in this review, including transparency limitations, accountability gaps, and insufficiently validated mitigation mechanisms. This alignment indicates that the governance challenges discussed in the academic literature are also reflected in emerging regulatory priorities, highlighting the importance of empirically grounded research on governance practices for LLM-assisted hiring systems. Similarly, employment regulators such as the U. S. Equal Employment Opportunity Commission have emphasized that automated hiring tools must comply with existing anti-discrimination and accountability requirements, reinforcing the importance of transparency, fairness, and human oversight in AI-supported employment decisions (Equal Employment Opportunity Commission, 2023).
Figure 10.
Mitigation maturity vs. risk in LLM based hiring.
The Figure 11 indicates that approximately 65–70% of studies are single-disciplinary, while 20–25% draw on two disciplines, and only 5–10% involve three or more disciplines. This distribution suggests that most research on LLMs in hiring is still conducted within disciplinary silos, with limited theoretical and methodological integration. Such patterns are typical of emerging technological fields, where early work is anchored in dominant home disciplines before deeper interdisciplinarity develops (Abbott, 2001). The limited share of multi-disciplinary studies constrains the development of integrative frameworks that connect technical performance, human judgment, organizational context, and governance. These findings underscore the need for more genuinely interdisciplinary research to address hiring as a socio-technical system (Trist and Bamforth, 1951; Orlikowski, 2007).
Figure 11.
The disciplinary depth per paper.
The Figure 12 shows clear differences in conceptual emphasis across disciplines. Studies in computer science and engineering focus primarily on performance and efficiency outcomes, which account for approximately 60–65% of their conceptual framing. In contrast, research in the social sciences and psychology places greater emphasis on fairness, bias, and candidate experience, representing about 55–60% of conceptual focus. Business and management studies more frequently emphasize governance, compliance, and organizational adoption, accounting for roughly 45–50% of their conceptual framing. These patterns indicate that disciplinary traditions strongly shape what aspects of LLM-based hiring are foregrounded. While this diversity enriches the field, it also contributes to fragmented knowledge development, as different disciplines prioritize different problem definitions and success criteria (Abbott, 2001; Orlikowski and Iacono, 2001). Greater theoretical integration across disciplines is therefore needed to develop more holistic models of LLM use in hiring.
Figure 12.
Conceptual focus of LLM-based hiring research by discipline.
Figure 13 shows that the strongest disciplinary co-occurrence occurs between computer science and engineering, followed by links between computer science and social sciences. Co-occurrence involving psychology, medicine, and arts and humanities is comparatively limited. This pattern indicates that most interdisciplinary collaboration in LLM-based hiring research centers on technical development paired with general social analysis, while behavioral, clinical, and normative perspectives remain weakly integrated. Such boundary-maintaining collaboration is common in emerging technological fields, where disciplines interact but retain distinct problem framings and methods (Abbott, 2001). This limited interdisciplinary integration suggests that the theoretical development of LLM-based hiring research remains structurally fragmented. Technical research emphasizes system performance, while organizational and governance research focuses on institutional implications, resulting in parallel rather than cumulative knowledge development (Abbott, 2001; Orlikowski and Iacono, 2001). Future research should develop integrative theoretical frameworks that connect computational capability, human judgment, and organizational governance to better explain how LLMs reshape hiring as a socio-technical decision system.
Figure 13.
Interdisciplinary co-occurrence in LLM based hiring.
Across all four research questions, a consistent pattern emerges. LLMs are being introduced into hiring as technical solutions to organizational problems, while their deeper socio-technical implications remain insufficiently integrated into empirical research. The literature reflects what sociologists of technology describe as a technology-first trajectory, in which technical capabilities outpace governance arrangements and theory lags behind practice (Orlikowski, 2007; Winner, 2017). By reframing hiring as a socio-technical decision system rather than as a purely computational task, this review highlights the need for research that bridges system design, human judgment, organizational context, and regulatory accountability. Taken together, the findings suggest that existing theories of automation, decision-making, and governance require extension to account for language-based AI systems. The reviewed literature suggests that LLMs may influence how evaluative narratives are constructed and interpreted, with potential implications for interpretive authority and accountability relationships. Accordingly, this review supports a shift from viewing AI in hiring as a decision aid toward conceptualizing it as a discursive actor within organizational decision systems. Future theoretical work should integrate socio-technical systems theory, sense making, organizational justice, and algorithmic governance into unified frameworks capable of explaining how LLMs influence hiring outcomes in practice.
4. Theoretical implications
As a scoping review, this study maps patterns in existing literature rather than testing causal effects or empirically validating organizational outcomes. The reviewed studies provide empirical evidence regarding where LLMs are applied, what outcomes are measured, and what risks and mitigation strategies are discussed. However, many studies remain conceptual, simulation-based, or design-oriented, with limited field-based validation. Accordingly, interpretations of LLMs as socio-technical decision intermediaries or language-mediated actors should be understood as theory-informed inferences derived from observed research patterns rather than as empirically established organizational effects. This distinction ensures that the review separates descriptive synthesis from theoretical interpretation, consistent with scoping review methodology (Tricco et al., 2018).
This review tries to advance theoretical understanding of large language models (LLMs) in hiring by clarifying their emerging role as socio-technical decision intermediaries embedded within organizational evaluation systems. The concentration of LLM use in early hiring stages suggests incremental integration that redistributes cognitive and interpretive work while preserving human evaluative authority. This pattern aligns with socio-technical systems theory, which emphasizes that technological artifacts reshape decision processes through interaction with organizational structures rather than replacing human judgment completely (Trist and Bamforth, 1951; Baxter and Sommerville, 2011; Orlikowski, 2007). At the same time, the literature reveals a systematic gap between technological development and theoretical engagement with hiring as an institutional and socially integrated process. Many studies emphasize performance and efficiency but do not engage with established theories concerning fairness, legitimacy, and organizational decision-making (Gilliland, 1993; Colquitt et al., 2001; March and Simon, 1958). This exclusion limits theoretical understanding of how language-mediated AI systems interact with institutional norms and accountability structures that govern hiring decisions (Scott, 2014; Suchman, 1995). The review also highlights structural fragmentation across disciplinary domains. Computer science research primarily conceptualizes LLMs as computational tools, whereas organizational and governance research emphasizes socio-institutional implications, including legitimacy, accountability, and fairness risks (Abbott, 2001; Orlikowski and Iacono, 2001; Raji et al., 2020). This fragmentation constrains cumulative theory development and reflects the interdisciplinary but structurally disconnected nature of emerging technological research domains. Finally, theoretical claims regarding the transformative organizational impact of LLMs currently exceed the empirical evidence base, which remains largely conceptual or simulation-driven. This pattern is consistent with early-stage technological fields, where conceptual framing precedes empirical institutionalization and organizational embedding (Weick, 1995; Parasuraman and Riley, 1997). Collectively, these findings indicate that LLMs are best understood not simply as automation tools but as socio-technical artifacts embedded within institutional decision systems, whose implications depend on their interaction with organizational structures, governance mechanisms, and human evaluation practices. By identifying these structural theoretical gaps, disciplinary silos, and conceptual inconsistencies, this review contributes to theory development by clarifying the conceptual positioning of LLMs in hiring and establishing priorities for empirically grounded interdisciplinary research.
5. Conclusion
This scoping review mapped how large language models (LLMs) are conceptualized, studied, and evaluated in hiring, including their placement across the hiring pipeline, the evidence used to assess them, the risks and governance mechanisms discussed, and the disciplinary structures shaping the field. The findings reveal a rapidly expanding literature that remains uneven in empirical maturity and theoretical integration. The reviewed studies suggest that LLMs function not only as computational tools but as language-mediated decision-support systems that shape how candidate information is organized, interpreted, and justified. However, the literature remains heavily oriented toward efficiency and performance outcomes, while organizational, institutional, and governance implications receive comparatively less empirical attention. Although risks related to bias, transparency, and accountability are widely recognized, mitigation strategies are largely conceptual and rarely evaluated in real organizational settings. Research on LLM-based hiring is increasingly interdisciplinary, yet integration across disciplinary perspectives remains limited. Advancing the field will require theory-informed, field-based research that connects technical capability with organizational context, human judgment, and governance mechanisms. This review provides a structured foundation for understanding the current state of the literature and clarifies priorities for future interdisciplinary research on responsible LLM use in hiring.
6. Implications for theory and practice
This scoping review has implications for theory, organizational practice, and governance. From a theoretical perspective, the findings indicate that large language models function not only as computational tools but as socio-technical artifacts that shape how candidate information is interpreted and justified within organizational decision systems. This extends existing socio-technical and organizational decision-making frameworks by highlighting the role of language-mediated AI in interpretive and evaluative processes (Orlikowski and Iacono, 2001; Weick, 1995; Scott, 2014).
From an organizational perspective, the concentration of LLM use in efficiency-oriented tasks, alongside limited empirical evaluation of governance mechanisms, suggests that responsible adoption requires clear oversight, transparency, and accountability structures. Governance practices such as human review and auditing may be important for maintaining decision legitimacy and defensibility (Raji et al., 2020; Suchman, 1995).
From a policy perspective, the gap between risk identification and empirically validated mitigation highlights the need for governance frameworks supported by empirical evidence. Research that examines how accountability and oversight mechanisms function in practice can inform both organizational policy and regulatory development for LLM-assisted hiring systems.
7. Limitations
This review has several limitations. First, it is based on Scopus-indexed literature, which may underrepresent relevant studies published in non-indexed journals, practitioner outlets, preprints, and policy reports. As a result, some industry-led and proprietary evaluations of LLM-based hiring tools may not be captured. Second, coding relied on titles, abstracts, and keywords using a lexicon-based, rule-driven approach. While this supports transparency and reproducibility, it may miss nuanced discussions or implicit theoretical positions that appear only in full texts. Studies using alternative terminology may also be undercounted. Third, disciplinary classification was derived from search strategies and metadata and may not perfectly reflect the intellectual orientation of individual papers. Interdisciplinary work may therefore be unevenly represented. Fourth, the fast-evolving nature of generative AI introduces temporal limitations. The review reflects the literature up to early 2026, and some findings may change as models, regulations, and organizational practices develop. Finally, as a scoping review, this study does not assess causal effects, comparative effectiveness, or real-world impact. The findings should be interpreted as patterns of research attention rather than as evaluations of system performance or harm.
8. Future research directions
The findings of this scoping review highlight the need for research that moves beyond technical system evaluation toward understanding LLM use as an organizational and socio-technical phenomenon. A key priority is the development of empirically grounded research examining how LLM-assisted hiring systems are implemented, governed, and integrated within real organizational decision processes. Such work would strengthen the empirical foundation of the field and clarify how LLM-supported decision-making interacts with institutional norms, accountability structures, and organizational practices. Future research should also focus on developing integrative theoretical frameworks that connect computational capability with human judgment, organizational context, and governance mechanisms. Current research remains fragmented across disciplinary perspectives, limiting cumulative theory development. Integrative approaches that bridge technical, organizational, and governance perspectives are needed to explain how language-mediated AI systems influence decision interpretation, accountability, and legitimacy in hiring. Finally, as LLM capabilities and regulatory environments continue to evolve, research should examine how organizational adoption, governance practices, and institutional expectations co-evolve over time. Such work will be essential for understanding the long-term implications of LLM integration and for supporting the responsible and effective use of language-based AI in employment decision systems.
Funding Statement
The author(s) declared that financial support was received for this work and/or its publication. This work was supported by the project CZ.02.1.01/0.0/0.0/16_017/0002334 Research Infrastructure for Young Scientists, which is co-financed from the Operational Program Research, Development and Education, and also supported by grant no. IGA25-PEF-DP-014 of the Internal Grant Agency FBE MENDELU.
Footnotes
Edited by: Vasile Daniel Pavaloaia, Alexandru Ioan Cuza University, Romania
Reviewed by: Yueqi Li, Skidmore College, United States
Elham Albaroudi, University of Salford, United Kingdom
Author contributions
ArT: Data curation, Formal analysis, Writing – original draft. AnT: Visualization, Writing – original draft, Formal analysis, Investigation. FD: Conceptualization, Resources, Supervision, Writing – review & editing. PM: Formal analysis, Methodology, Supervision, Writing – review & editing.
Conflict of interest
The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
The author PM declared that they were an editorial board member of Frontiers, at the time of submission. This had no impact on the peer review process and the final decision.
Generative AI statement
The author(s) declared that Generative AI was used in the creation of this manuscript. ChatGPT was used to polish the text.
Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.
Publisher’s note
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.
Supplementary material
The Supplementary material for this article can be found online at: https://www.frontiersin.org/articles/10.3389/frai.2026.1798519/full#supplementary-material
References
- Abbott A. (2001). Chaos of Disciplines. Chicago, IL: University of Chicago Press. [Google Scholar]
- Arksey H., O’Malley L. (2005). Scoping studies: towards a methodological framework. Int. J. Soc. Res. Methodol. 8, 19–32. doi: 10.1080/1364557032000119616 [DOI] [Google Scholar]
- Autor D. H., Levy F., Murnane R. J. (2003). The skill content of recent technological change: an empirical exploration*. Q. J. Econ. 118, 1279–1333. doi: 10.1162/003355303322552801 [DOI] [Google Scholar]
- Barocas S., Selbst A. D. (2016). Big data’s disparate impact. Calif. Law Rev. 104, 671–732. doi: 10.15779/Z38BG31 [DOI] [Google Scholar]
- Baxter G., Sommerville I. (2011). Socio-technical systems: from design methods to systems engineering. Interact. Comput. 23, 4–17. doi: 10.1016/j.intcom.2010.07.003 [DOI] [Google Scholar]
- Bender E. M., Gebru T., McMillan-Major A., Shmitchell S. (2021). “On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?” in Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, 610–623.
- Bertrand M., Mullainathan S. (2004). Are Emily and Greg more employable than Lakisha and Jamal? A field experiment on labor market discrimination. Am. Econ. Rev. 94, 991–1013. doi: 10.1257/0002828042002561 [DOI] [Google Scholar]
- Bommasani R., Hudson D. A., Adeli E., Altman R., Arora S., Arx S. (2022). On the Opportunities and Risks of Foundation Models. Stanford, CA: Stanford University.
- Chadegani A. A., Salehi H., Yunus M. M., Farhadi H., Fooladi M., Farhadi M., et al. (2013). A comparison between two main academic literature collections: web of science and Scopus databases. Asian Soc. Sci. 9:18. doi: 10.5539/ass.v9n5p18 [DOI] [Google Scholar]
- Colquitt J. A., Conlon D. E., Wesson M. J., Porter C. O. L. H., Ng K. Y. (2001). Justice at the millennium: a meta-analytic review of 25 years of organizational justice research. J. Appl. Psychol. 86, 425–445. doi: 10.1037/0021-9010.86.3.425, [DOI] [PubMed] [Google Scholar]
- Dietvorst B. J., Simmons J. P., Massey C. (2015). Algorithm aversion: people erroneously avoid algorithms after seeing them err. J. Exp. Psychol. Gen. 144, 114–126. doi: 10.1037/xge0000033, [DOI] [PubMed] [Google Scholar]
- Dipboye R. L. (2018). Employee selection: How psychology can improve the hiring process. In R. N. Landers (Ed.), The Cambridge handbook of technology and employee behavior Cambridge University Press. 99–120. [Google Scholar]
- Equal Employment Opportunity Commission . (2023). Assessing adverse impact in software, algorithms, and artificial intelligence used in employment selection procedures. Washington, DC.: U.S. Equal Employment Opportunity Commission. [Google Scholar]
- Gilliland S. W. (1993). The perceived fairness of selection systems: an organizational justice perspective. Acad. Manag. Rev. 18, 694–734. doi: 10.2307/258595 [DOI] [Google Scholar]
- Highhouse S. (2008). Stubborn reliance on intuition and subjectivity in employee selection. Ind. Organ. Psychol. 1, 333–342. doi: 10.1111/j.1754-9434.2008.00058.x [DOI] [Google Scholar]
- Kroll J. A., Huey J., Barocas S., Felten E. W., Reidenberg J. R., Robinson D. G. (2017). Accountable algorithms (SSRN scholarly paper no. 2765268). Soc. Sci. Res. Netw. 165, 633–705. [Google Scholar]
- March J. G., Simon H. A. (1958). Organizations. New York, NY: Wiley. [Google Scholar]
- Munn Z., Peters M. D. J., Stern C., Tufanaru C., McArthur A., Aromataris E. (2018). Systematic review or scoping review? Guidance for authors when choosing between a systematic or scoping review approach. BMC Med. Res. Methodol. 18:143. doi: 10.1186/s12874-018-0611-x, [DOI] [PMC free article] [PubMed] [Google Scholar]
- Neumark D. (2018). Experimental research on labor market discrimination. J. Econ. Lit. 56, 799–866. doi: 10.1257/jel.20161309 [DOI] [Google Scholar]
- NIST (2023). Artificial Intelligence Risk Management Framework (AI RMF 1.0). U.S. Department of Commerce.
- Orlikowski W. J. (2007). Sociomaterial practices: exploring technology at work. Organ. Stud. 28, 1435–1448. doi: 10.1177/0170840607081138 [DOI] [Google Scholar]
- Orlikowski W. J., Iacono C. S. (2001). Research commentary: desperately seeking the “IT” in IT research—a call to theorizing the IT artifact. Inf. Syst. Res. 12, 121–134. doi: 10.1287/isre.12.2.121.9700 [DOI] [Google Scholar]
- Pager D., Shepherd H. (2008). The sociology of discrimination: racial discrimination in employment, housing, credit, and consumer markets. Annu. Rev. Sociol. 34, 181–209. doi: 10.1146/annurev.soc.33.040406.131740, [DOI] [PMC free article] [PubMed] [Google Scholar]
- Parasuraman R., Riley V. (1997). Humans and automation: use, misuse, disuse, abuse. Hum. Factors 39, 230–253. doi: 10.1518/001872097778543886 [DOI] [Google Scholar]
- Pollock D., Peters M. D. J., Khalil H., McInerney P., Alexander L., Tricco A. C., et al. (2023). Recommendations for the extraction, analysis, and presentation of results in scoping reviews. JBI Evid. Synth. 21:520. doi: 10.11124/JBIES-22-00123, [DOI] [PubMed] [Google Scholar]
- Quillian L., Pager D., Hexel O., Midtbøen A. H. (2017). Meta-analysis of field experiments shows no change in racial discrimination in hiring over time. Proc. Natl. Acad. Sci. 114, 10870–10875. doi: 10.1073/pnas.1706255114, [DOI] [PMC free article] [PubMed] [Google Scholar]
- Raji I. D., Smart A., White R. N., Mitchell M., Gebru T., Hutchinson B., (2020). “Closing the AI Accountability gap: Defining an end-to-end Framework for internal Algorithmic Auditing,” in Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, 33–44.
- Scott W. R. (2014). Institutions and Organizations. London: SAGE Publications. [Google Scholar]
- Suchman M. C. (1995). Managing legitimacy: strategic and institutional approaches. Acad. Manag. Rev. 20, 571–610. doi: 10.2307/258788 [DOI] [Google Scholar]
- Tricco A. C., Lillie E., Zarin W., O’Brien K. K., Colquhoun H., Levac D., et al. (2018). PRISMA extension for scoping reviews (PRISMA-ScR): checklist and explanation. Ann. Intern. Med. 169, 467–473. doi: 10.7326/M18-0850, [DOI] [PubMed] [Google Scholar]
- Trist E. L., Bamforth K. W. (1951). Some social and psychological consequences of the longwall method of coal-getting: an examination of the psychological situation and defences of a work group in relation to the social structure and technological content of the work system. Hum. Relat. 4, 3–38. doi: 10.1177/001872675100400101 [DOI] [Google Scholar]
- Veale M., Borgesius F. Z. (2021). Demystifying the draft EU artificial intelligence act—analysing the good, the bad, and the unclear elements of the proposed approach. Comput. Law Rev. Int. 22, 97–112. doi: 10.9785/cri-2021-220402 [DOI] [Google Scholar]
- Waltman L. (2016). A review of the literature on citation impact indicators. J. Informetr. 10, 365–391. doi: 10.1016/j.joi.2016.02.007 [DOI] [Google Scholar]
- Weick K. E. (1995). Sensemaking in Organizations. London: SAGE Publications Inc. [Google Scholar]
- Winner L. (2017). Do Artifacts Have Politics? Computer Ethics. London: Routledge. [Google Scholar]
- Zupic I., Čater T. (2015). Bibliometric methods in management and organization. Organ. Res. Methods 18, 429–472. doi: 10.1177/1094428114562629 [DOI] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.













