Abstract
Modern clinical laboratory operates within a stringent regulatory ecosystem requiring precise adherence to Standard Operating Procedures (SOPs). However, standard Retrieval-Augmented Generation (RAG) approaches frequently fail in this domain due to context fragmentation, where fixed-size segmentation arbitrarily severs critical procedural dependencies. To address this, we introduce LabSage, a domain-adapted framework implementing structural-semantic decoupling. This hierarchical architecture indexes compact search units for optimal retrieval precision while dynamically retrieving expanded context units to preserve procedural completeness during inference. Evaluated on authentic laboratory queries using Qwen-2.5-7B as backbone, LabSage achieved an Answer Accuracy of 0.780 and Context Recall of 0.909, outperforming standard RAG by 8.3% and 5.7%, respectively. These findings demonstrate that decoupling vector search from reasoning context is a critical architectural advancement in regulated medical domains to mitigate safety-critical omissions and ensure compliance.
Introduction
In the domain of clinical pathology, operational integrity relies fundamentally on the precise execution of standardized protocols. Regulatory bodies, such as the Clinical Laboratory Improvement Amendments (CLIA) and the College of American Pathologists (CAP), mandate that laboratory personnel adhere strictly to Standard Operating Procedures for every aspect of testing—from instrument calibration to critical value reporting1,2. However, the sheer volume of documentation required to maintain a modern laboratory creates a significant information retrieval bottleneck. With hundreds of active SOPs governing diverse testing disciplines, the cognitive demand of manually locating specific, actionable guidance during time-critical workflows presents a persistent challenge to both operational efficiency and regulatory compliance.
To address this operational bottleneck, recent advances in Large Language Models (LLMs) have catalyzed interest in deploying Retrieval-Augmented Generation (RAG)3 systems as intelligent information retrieval (IR) interfaces for medical documentation. However, the direct application of general-purpose RAG architectures to clinical laboratory SOPs introduces domain-specific failure modes that carry substantial safety implications. Standard RAG implementations conflate two distinct functional requirements: the granularity optimal for semantic search and the scope necessary for coherent generation. These systems typically employ a single text segmentation strategy that serves both as the unit of retrieval during search and as the context provided to the language model during generation. This architectural coupling proves fundamentally incompatible with the structural characteristics of laboratory procedural documentation, where achieving high precision in matching relevant instructions requires fine-grained indexing, yet generating complete and compliant responses demands access to broader procedural contexts that include prerequisite conditions, sequential dependencies, and safety constraints.
This paper presents LabSage, a domain-adapted question-answering framework engineered specifically for the structural and semantic requirements of clinical laboratory documentation. Our central contribution is the implementation of structural-semantic decoupling, an architectural pattern that separates the indexing phase from the inference phase of the retrieval pipeline. During indexing, documents are segmented into compact search units that enable the retrieval system to achieve high precision in identifying procedurally relevant passages. During inference, when a search unit is matched, the system retrieves its associated context unit, a substantially larger text span that encompasses the complete procedural scope surrounding the matched instruction. This hierarchical architecture resolves the fundamental tension between search granularity and generation context by optimizing each phase independently according to its distinct functional requirements.
Our evaluation employs two open-source language models, Qwen-2.5-7B and Llama-3.1-8B, selected for their modest computational requirements and their viability in privacy-sensitive healthcare deployments that preclude reliance on cloud-based commercial APIs. Testing against a curated dataset of authentic laboratory queries demonstrates that our hierarchical retrieval strategy yields consistent improvements across both retrieval quality and generation accuracy compared to standard RAG and zero-shot baselines. Additional ablation studies on retrieved context count and embedding model selection further validate the robustness and generalizability of the decoupling architecture.
The main contributions of this paper are as follows:
We identify and formalize structural-semantic decoupling as a fundamental architectural requirement for RAG systems operating on hierarchically organized procedural documentation in regulated medical domains.
We present a concrete implementation through a hierarchical retrieval architecture with distinct search units and context units, establishing a generalizable approach applicable to domains beyond clinical laboratories.
We demonstrate through rigorous evaluation that decoupling indexing granularity from generation context yields measurable improvements in both retrieval quality and answer accuracy, with direct implications for regulatory compliance and operational safety in clinical settings where incomplete procedural guidance can compromise patient care.
Background and Related Work
Retrieval-Augmented Generation in Healthcare
Retrieval-Augmented Generation has emerged as a promising approach to enhance the capabilities of LLMs in health-care settings, where accurate and contextually grounded responses are critical for clinical decision-making. In medical domains, RAG systems integrate external knowledge retrieval to mitigate hallucinations and improve factual accuracy, particularly when dealing with specialized documentation such as clinical guidelines, electronic health records, and procedural manuals. For instance, studies have explored RAG for interpreting clinical laboratory regulations, demonstrating how grounding LLM outputs in dynamically retrieved source documents can enhance flexibility, maintainability, and explainability while addressing domain-specific challenges like regulatory compliance4. Other work has applied RAG to improve local LLM performance in tasks such as radiology contrast media consultation and surgical fitness assessments, showing significant gains in accuracy and safety by incorporating retrieved evidence from medical corpora5. Recent studies emphasize the need for trustworthy RAG frameworks in healthcare, focusing on reliability, privacy, and fairness, but also note persistent issues with bias propagation from retrieval stages that can exacerbate disparities in clinical outcomes6. However, these applications often highlight limitations in handling fragmented contexts, where procedural logic in documents like Standard Operating Procedures is split arbitrarily, leading to incomplete or misleading generations.
Hierarchical and Contextual Retrieval Strategies
To overcome context fragmentation in RAG systems, particularly for hierarchically organized procedural documents, advanced retrieval strategies have been proposed that decouple search granularity from generation context. Hierarchical RAG (HierRAG) approaches structure data into multi-level representations, such as summaries at higher levels and detailed chunks at lower ones, enabling more precise retrieval while preserving broader contextual integrity7. In medical contexts, contextual retrieval-augmented methods have been adapted to augment chunks with explanatory preambles or knowledge graphs, improving accuracy in domains like urology and microbiology where procedural dependencies are crucial8. These strategies share similarities with graph-based RAG, which organizes information hierarchically to reason over complex queries in healthcare, such as patient conditions or lab protocols9. Despite these advances, challenges remain in computational overhead and ensuring completeness in regulated procedural texts, where omitting dependencies can lead to safety risks. Our work builds on these paradigms by implementing structural-semantic decoupling specifically for clinical laboratory SOPs, leveraging native document hierarchies without synthetic augmentations.
Methods
Dataset and Document Characteristics
Our evaluation corpus consists of 20 authentic SOPs obtained from a functioning clinical laboratory, encompassing three primary testing domains: Chemistry with 7 documents, Hematology with 10 documents, and Urinalysis with 3 documents. The Chemistry domain includes analyte-specific assays such as Ammonia and Lactate testing. The Hematology domain comprises Complete Blood Count calibration and linearity verification procedures across multiple cellular components. The Urinalysis domain covers automated analysis platforms used in clinical laboratories. These documents represent the full spectrum of laboratory documentation types encountered in routine clinical practice, including instrument operation procedures, quality control protocols, specimen handling requirements, and critical value reporting workflows.
To construct a robust information retrieval framework, as shown in Figure 1, we curated a dataset consisting of 685 question-answer pairs derived directly from the procedural content. Questions were formulated to reflect authentic information-seeking scenarios encountered by laboratory personnel during daily operations, spanning queries about specimen stability requirements, quality control materials, assay principles, and procedural methodologies. Ground truth answers were extracted verbatim from the corresponding SOPs to establish definitive reference standards against which system outputs could be evaluated.
Figure 1.
Data curation pipeline for the clinical laboratory Question-Answering dataset. Authentic Standard Operating Procedures from Chemistry, Hematology, and Urinalysis domains were processed to extract realistic information-seeking scenarios. Ground truth answers were extracted verbatim from source documents to establish a definitive reference standard for evaluation.
Clinical laboratory SOPs exhibit distinctive structural characteristics that differentiate them from general-domain documents. These procedures are organized hierarchically, with top-level sections addressing principles, reagents, and instrumentation, followed by detailed procedural steps, quality control requirements, and troubleshooting algorithms. Critical procedural segments such as troubleshooting decision trees or assay principle descriptions often contain dense conditional logic that cannot be meaningfully decomposed. Quality control failure protocols specify sequential actions where subsequent steps are contingent upon the outcomes of prior interventions, creating logical dependency chains that span multiple sentences. The structural analysis of our corpus revealed substantial variance in procedural paragraph lengths, invalidating fixed-size segmentation assumptions employed in standard RAG implementations.
Hierarchical Retrieval Architecture and Workflow
The LabSage framework, as shown in Figure 2, implements structural-semantic decoupling through a hierarchical architecture that maintains distinct text representations for the indexing and inference phases. This design addresses the tension between search precision, which benefits from fine-grained semantic units, and generation quality, which requires comprehensive procedural context. Our approach shares conceptual similarities with contextual retrieval techniques that augment individual chunks with LLM-generated contextual descriptions to improve retrieval accuracy. However, while contextual retrieval requires separate LLM invocations to generate context for each chunk, our frame-work exploits the inherent hierarchical structure of procedural documentation to achieve context enrichment through direct structural relationships, eliminating the computational overhead and potential inconsistencies associated with synthetic context generation.
Figure 2.
The LabSage hierarchical retrieval architecture implementing structural-semantic decoupling. The Offline Indexing Pipeline segments documents into fine-grained Search Units (S) for precise vector retrieval and broader Context Units (U) for content storage. The Online Pipeline retrieves the expanded Context Unit associated with the best-matching Search Unit via mapping ψ, ensuring the Large Language Model receives the complete procedural scope necessary for compliant generation.
Section-Aware Document Segmentation Let = {d1, d2, …, dN} denote the corpus of N laboratory procedural documents. Each document d ∈ is structured as a hierarchy of sections and subsections following standard SOP organization. We exploit this inherent structure to create dual-granularity representations: search units optimized for retrieval precision and context units optimized for generation completeness.
Each document d is first partitioned into context units corresponding to complete procedural sections:
where represents the i-th context unit aligned with major document sections such as “Specimen Requirements,” “Quality Control Procedures,” or “Troubleshooting Protocols.” Each context unit encompasses complete procedural logic including section headers, prerequisite statements, sequential instruction sets, and associated warnings.
Each context unit is subsequently subdivided into overlapping search units corresponding to subsections or dense procedural paragraphs:
where represents the j-th search unit within context unit . The relation ⊇ indicates that search units may overlap at boundaries to ensure semantic continuity, following standard chunking practices. This creates focused semantic units for vector search while maintaining a strict mapping function ψ : enabling context expansion during retrieval. The section-aware segmentation ensures that each search unit corresponds to a coherent procedural fragment while its associated context unit provides the complete section framework. The complete corpus representation consists of all search units with preserved mappings to their respective context units.
Embedding and Indexing The indexing phase operates exclusively on search units. Each search unit is encoded into a dense vector representation using a pretrained text embedding model :
where demb denotes the embedding dimension. The resulting vectors form a searchable index . Context units are not independently indexed; they exist solely as retrieval targets accessed through the mapping function ψ, analogous to how contextual retrieval maintains chunk-context pairs but without requiring LLM-generated context.
Context-Augmented Generation For a user query q, the system first encodes the query and identifies the top-k most relevant search units via cosine similarity:
For each retrieved search unit , the system performs context expansion by retrieving the associated complete section through the mapping ψ(S). The expanded context units are concatenated and provided to the language model:
where ⊕ denotes string concatenation and fLLM represents the generation model. This architecture separates the retrieval function operating on fine-grained search units from the context provision function operating on section-level context units. The structural-semantic decoupling ensures that semantic search identifies specific instructional fragments at subsection granularity while the generation model accesses complete procedural sections including prerequisite conditions, sequential dependencies, quality control checkpoints, and safety warnings that may be absent from isolated search units. Unlike contextual retrieval approaches that require separate LLM calls to generate contextual summaries for each chunk, our framework leverages the native document structure to provide this context directly, reducing computational cost while maintaining semantic coherence through the preservation of original procedural text.
Experiments
Baselines
We compared three configurations to evaluate the effectiveness of structural-semantic decoupling practice in the Lab-Sgae framework:
Base LLM (w/o RAG): Zero-shot generation without retrieval augmentation, where the language model relies solely on parametric knowledge encoded during pretraining to answer queries. This baseline establishes the extent to which laboratory procedural knowledge exists in general-purpose model parameters.
Standard RAG: Traditional retrieval-augmented generation implementing fixed-size chunking with uniform segmentation length. This represents the conventional approach where a single segmentation strategy serves both as the unit of retrieval during search and as the context provided to the language model during generation.
LabSage (HierRAG): Our adopted hierarchical retrieval architecture implementing structural-semantic decoupling with search units optimized for retrieval precision and context units optimized for generation completeness.
Evaluation Metrics
To assess the effectiveness of the LabSage framework, we adopted the RAGAS (Retrieval-Augmented Generation Assessment) evaluation protocol10. RAGAS utilizes an “LLM-as-a-judge” paradigm to evaluate two essential dimensions of a RAG system: (1) how well the framework surfaces the information needed to answer a question, and (2) how accurately the model uses that information to generate the final response. We selected four metrics for evaluation with each having a score range of [0, 1], where higher scores indicate better performance.
Context Recall In clinical laboratory workflows, missing a prerequisite condition or safety note can lead to operational errors. Context Recall acts as a safety metric by measuring whether the retrieved text contains the essential procedural elements required to answer the query. To calculate this, the ground truth answer is first decomposed into a set of atomic statements (claims) by the judge LLM11. The judge LLM then quantifies the proportion of these claims that can be semantically attributed to the retrieved context. A high recall score indicates that the system successfully retrieved the complete scope of procedural dependencies necessary for a compliant response.
Context Precision Even if a retriever captures all relevant content, it might also return unrelated or noisy text that can confuse the generation model. Context Precision evaluates how focused the retrieved context is by measuring the system’s ability to rank relevant chunks higher than irrelevant ones. When the retrieval module returns a list of chunks for a query, each chunk is judged by the judge LLM as either relevant or irrelevant. Context Precision then reflects how often and how early relevant chunks appear in the ranked list. Retrieval outputs that put relevant content at the top yield high Context Precision, whereas outputs that intermix or bury relevant content among irrelevant material yield lower scores. This metric penalizes systems that bury relevant procedural logic beneath unrelated information, ensuring the generation model receives a prioritized and signal-rich context.
Context Entity Recall Standard Operating Procedures rely on precise technical nomenclature, such as specific reagent identifiers, instrument names, and quantitative thresholds. Context Entity Recall assesses the system’s ability to preserve this domain-specific vocabulary. It is calculated by identifying the set of entities present in the reference ground truth and measuring what fraction of those entities are also present in the retrieved context. This metric distinguishes systems that simply retrieve thematically similar content from those that surface the exact technical specifications required for valid testing.
Answer Accuracy To evaluate the final response quality, Answer Accuracy measures the agreement between the model’s generated answer and the ground truth reference. It is computed via two LLM judge calls: the first rates the response against the reference on a discrete scale of {0, 2, 4}, and the second swaps their roles to ensure symmetric evaluation. The two ratings are then averaged and normalized to a single scaler value.
Together, these metrics allow us to evaluate whether LabSage retrieves the right information, avoids irrelevant content, preserves key laboratory terminology, and ultimately enables a model to produces accurate and actionable answers.
Implementation Details
We employed LangChain for the implementation of both standard RAG and LabSage framework. We used OpenAI text-embedding-3-large (3072-dim) as the primary embedding model femb, with all search unit vectors stored in a FAISS index. Unless otherwise specified, the number of retrieved units during inference is set to k = 3. We randomly selected 100 samples from all 685 curated question-answer pairs as the evaluation split. We base evaluation on two lightweight open-source models: Qwen-2.5-7B12 and Llama-3.1-8B13, both deployable in privacy-sensitive healthcare environments. For the judge model in RAGAS, we used GPT-4o throughout all experiments.
Main Experiment Results
Table 1 presents the comparative performance of LabSage against Standard RAG and zero-shot baselines across Qwen-2.5-7B and Llama-3.1-8B models, demonstrating consistent improvements through structural-semantic decoupling. LabSage achieved superior Answer Accuracy on both models: 0.780 versus 0.720 for Standard RAG on Qwen-2.5-7B, and 0.708 versus 0.607 on Llama-3.1-8B. These gains confirm that retrieving complete contexts is essential for generating compliant procedural guidance in regulated domains.
Table 1.
Quantitative performance comparison on clinical laboratory queries. LabSage demonstrates superior performance in retrieval and generation quality, validating the effcacy of the decoupling strategy.
| Model | Strategy | Context Recall | Context Precision | Context Entity Recall | Answer Accuracy |
|---|---|---|---|---|---|
| Qwen-2.5-7B | w/o RAG | – | – | – | 0.107 |
| w/ Standard RAG | 0.860 | 0.871 | 0.302 | 0.720 | |
| w/ HierRAG | 0.909 | 0.901 | 0.216 | 0.780 | |
| Llama-3.1-8B | w/o RAG | – | – | – | 0.100 |
| w/ Standard RAG | 0.860 | 0.871 | 0.302 | 0.607 | |
| w/ HierRAG | 0.909 | 0.901 | 0.216 | 0.708 |
For the safety-critical Context Recall metric, LabSage reached 0.909 compared to Standard RAG’s 0.860, indicating effective mitigation of context fragmentation. This ensures sequential dependencies and prerequisite conditions remain available during generation. LabSage also maintained superior Context Precision, demonstrating that fine-grained search units effectively identify relevant content while expanded context units provide necessary breadth without introducing excessive noise.
Notably, Standard RAG achieved higher Context Entity Recall. Qualitative analysis revealed this metric was inflated by retrieving procedurally irrelevant sections with high entity densities, introducing distractor entities that reduced Answer Accuracy. LabSage retrieved fewer total entities but from structurally relevant contexts, yielding higher generation accuracy. This underscores that procedural framework relevance outweighs raw entity enumeration in domain-specific documentation.
Ablation Study: Impact of Embedding Model Selection
To assess whether structural-semantic decoupling generalizes beyond a single embedding model, we replaced OpenAI text-embedding-3-large with MedCPT14, a biomedical retrieval model employing asymmetric dual-encoders trained on PubMed search logs via contrastive learning. Results are presented in Table 2.
Table 2.
Impact of embedding model on retrieval and generation quality. HierRAG consistently improves over Standard RAG regardless of embedding model, with larger relative gains under weaker retrieval signals.
| Embedding | Strategy | Context Recall | Context Precision | Context Entity Recall | Answer Accuracy |
|---|---|---|---|---|---|
| OpenAI | w/ Standard RAG | 0.860 | 0.871 | 0.302 | 0.720 |
| w/ HierRAG | 0.909 | 0.901 | 0.216 | 0.780 | |
| MedCPT | w/ Standard RAG | 0.399 | 0.440 | 0.190 | 0.398 |
| w/ HierRAG | 0.566 | 0.610 | 0.218 | 0.502 |
HierRAG yields consistent improvements regardless of the underlying embedding model. With MedCPT, context expansion via the mapping function ψ improved Context Recall by 41.9%, Context Precision by 38.6%, and Answer Accuracy by 26.1% over Standard RAG, substantially larger than those observed with the default embedding model. This suggests the decoupling architecture provides greater benefit when retrieval precision is lower, compensating for weaker embedding signals by ensuring complete procedural sections reach the generation model.
Ablation Study: Impact of Retrieved Context Units Count
To investigate the relationship between the number of retrieved context units and system performance, we conducted an ablation study varying the retrieval parameter k ∈ {1, 3, 5, 10} using Qwen-2.5-7B with LabSage architecture. Results are presented in Table 3.
Table 3.
Ablation study on retrieved context count using Qwen-2.5-7B with LabSage architecture. All configurations use the same model and hierarchical retrieval strategy, varying only the number of retrieved context units k.
| k | Context Recall | Context Precision | Answer Accuracy |
|---|---|---|---|
| 1 | 0.771 | 0.800 | 0.715 |
| 3 | 0.909 | 0.901 | 0.780 |
| 5 | 0.929 | 0.930 | 0.795 |
| 10 | 0.949 | 0.960 | 0.818 |
All metrics showed monotonic improvement with increasing k. Context Recall improved from 0.771 to 0.949, while Context Precision increased from 0.800 to 0.960, indicating the hierarchical architecture maintains ranking quality even with expanded retrieval scope. Answer Accuracy gained from 0.715 to 0.818, confirming broader procedural coverage enhances generation fidelity.
The performance gap between k = 1 and k = 3 was particularly significant, suggesting laboratory queries frequently require cross-referencing multiple procedural sections. Beyond k = 5, improvements continued at diminishing rates, suggesting practical deployments could balance computational cost against marginal accuracy gains.
Case Analysis
Table 4 presents representative cases illustrating how retrieval strategy impacts answer quality and regulatory compliance. For the ammonia quality control query, Standard RAG retrieved a passage describing chemical mechanisms with technical terminology but lacking QC protocol information. LabSage’s search unit matched the QC subsection, with context expansion retrieving the complete section including the correct material specification “BioRad ETH/AMM 1, 2, 3 INTELIQ”, enabling accurate generation.
Table 4.
Qualitative analysis of retrieval and generation outcomes. Representative cases comparing Standard RAG and LabSage on authentic laboratory queries. The examples illustrate how structural-semantic decoupling enables the retrieval of complete procedural sections, whereas Standard RAG suffers from context fragmentation, leading to incomplete or hallucinated responses.
| Query and Ground Truth | Standard RAG Retrieved Snippet and Model Response | HierRAG Retrieved Snippet and Model Response |
|---|---|---|
|
Query:
What quality control material is used for ammonia testing?
Ground Truth: Bio-Rad ETH/AMM 1, 2, 3 INTELIQ. |
Retrieved Snippet:
Ammonia reacts with -
ketoglutarate and reduced cofactor to form L-glutamate and the cofactor. The reaction is catalyzed by glutamate dehydrogenase. The decrease in absorbance due to the oxidation of the reduced cofactor is monitored at 340/694 nm and is proportional to the ammonia concentration.
Model response: The Atellica CH Amm assay uses non-interfering substances as quality control material. |
Retrieved Snippet:
QUALITY CONTROL Material BioRad ETH/AMM 1, 2, 3 INTELIQ. Procedure Refer to laboratory QC procedure for information on testing interval and management of results outside of acceptable limits
Model response: BioRad ETH/AMM 1, 2, 3 INTELIQ. |
|
Query:
How is pH determined by the clinitek?
Ground Truth: pH is based on a double indicator principle that gives a broad range of colors covering the entire urinary pH range. |
Retrieved Snippet:
During analysis, an aliquot of sample is dispensed onto the individual test pads and concentration of each analyte is determined by intensity of color that subsequently develops on each pad. The routine urinalysis includes the examination of: Physical characteristics, such as color, appearance, and specific gravity. Color is determined using a non-reactive white pad absorbs the sample to detect urine color.
Model response: By potentiometry using a sensor chain and voltameter, relating the potential to concentration via the Nernst equation. |
Retrieved Snippet:
Chemical characteristics, including pH, protein, glucose, ketones, occult blood, bilirubin, urobilinogen, nitrite, and leukocyte esterase. pH is based on a double indicator principle that gives a broad range of colors covering the entire urinary pH range. Protein is based on the protein-error-of-indicators principle. At a constant pH, the presence of protein causes a change in the color of the indicator.
Model response: pH is based on a double indicator principle that gives a broad range of colors covering the entire urinary pH range. |
For pH determination methodology, Standard RAG retrieved general analytical workflow descriptions discussing color determination but omitting pH measurement principles. LabSage’s hierarchical retrieval matched within the “Chemical Characteristics” subsection, with expansion retrieving the complete section enumerating all chemical analytes and measurement principles, enabling the model to extract the correct double indicator colorimetric methodology.
These outcomes demonstrate that in regulatory-compliant documentation, generation models require complete procedural sections to explicitly address the query topic.
Discussion
Structural-Semantic Decoupling as a Safety Requirement
Our primary finding establishes that for hierarchically organized procedural documentation in regulated domains, text segmentation constitutes a safety-critical architectural decision rather than an implementation detail. Standard RAG’s unified segmentation strategies force compromises between retrieval precision and contextual completeness, systematically fragmenting logical dependencies defining compliant execution. Structural-semantic decoupling eliminates this compromise, yielding measurable improvements in both retrieval quality and answer accuracy.
Comparison with Contextual Retrieval Paradigms
While our hierarchical architecture shares conceptual similarities with contextual retrieval approaches, fundamental differences exist in implementation and computational requirements. Contextual retrieval employs LLMs to generate synthetic contextual summaries for each chunk, requiring separate invocations for every corpus chunk. This incurs substantial computational cost and potentially introduces inconsistencies where generated context may misrepresent critical procedural constraints. Our framework exploits native document hierarchies, eliminating computational over-head while maintaining semantic coherence through original procedural text preservation. Future work could explore hybrid approaches combining structural retrieval with targeted synthetic context generation for ambiguous subsections, though such enhancements require careful validation to prevent inconsistencies.
Accuracy Gap and Deployment Considerations
While LabSage demonstrates consistent improvements over standard RAG across all evaluated configurations, the achieved Answer Accuracy of 0.780 with the best-performing setup does not yet meet the stringent requirements for autonomous deployment in clinical laboratory settings, where procedural errors can directly impact patient safety. This gap between current performance and the near-perfect accuracy demanded by regulated workflows warrants explicit discussion. We position LabSage as a decision-support tool operating within a human-in-the-loop framework rather than as a replacement for direct SOP consultation. In practice, the system would assist laboratory personnel in rapidly locating relevant procedural guidance while requiring human verification before any action is taken on the retrieved information. Several pathways exist for closing the accuracy gap in future work: (1) increasing the retrieval parameter k, which our ablation study shows yields monotonic improvements up to k = 10; (2) adopting higher-capacity embedding models with longer context windows that reduce chunking artifacts; (3) incorporating hybrid retrieval strategies that combine dense semantic search with sparse keyword matching for precise technical terminology; and (4) integrating confidence estimation mechanisms that flag low-certainty responses for mandatory human review. Deployment in regulated environments would additionally require audit logging of all system-generated responses, version control of the SOP corpus to ensure retrieval from current documents, and periodic revalidation against updated procedural content.
Limited Effect of Post-training on Domain-Specific Retrieval
We also conducted exploratory experiments with post-training strategies, specifically supervised fine-tuning and re-inforcement learning (preference learning). However, both post-training approaches yielded minimal performance improvements over base models. Specifically, supervised fine-tuning and preference learning improved the Answer Accuracy of Qwen-2.5-7B by only 1.3% and 0.4% respectively compared to the base model. This aligns with findings in medical domain adaptation, where fine-tuning on highly specialized healthcare data absent from pretraining shows limited changes to the model’s representational space and often fails to introduce novel knowledge15. Similarly, recent study shows that reinforcement learning refines sampling efficiency but does not expand reasoning boundaries beyond the base model16. This stems from clinical laboratory SOPs constituting highly specialized, institution-specific knowledge categorically absent from public pretraining corpora. Language models possess essentially zero parametric knowledge about specific procedures, instruments, and protocols in our evaluation queries. General LLMs often struggle with in-depth clinical content, and enhancements come more from retrieval-augmented approaches than parametric fine-tuning17. When hierarchical retrieval provides correct contexts, the model reliably extracts specifications; when retrieval fails, parametric knowledge cannot compensate.
Conclusion
In this work, we attempt to address the critical challenge of context fragmentation in clinical laboratory information retrieval by developing LabSage, a domain-adapted RAG framework implementing structural-semantic decoupling. We introduce a hierarchical architecture that maintains distinct text representations for indexing and inference phases—fine-grained search units optimized for semantic matching precision and expanded context units optimized for procedural completeness. The identified architectural requirements underscore the importance of domain-specific design considerations when deploying RAG systems in regulated medical environments. Our hierarchical retrieval methodology provides actionable insights and practical tools for advancing AI safety in healthcare information systems. The implications of this research extend beyond clinical laboratories to the broader question of how to design retrieval architectures that preserve critical procedural dependencies in regulated domains. As LLM-based systems become more prevalent in healthcare, ensuring their ability to retrieve complete and contextually appropriate information will be crucial for maintaining compliance, operational safety, and the integrity of clinical decision support.
Figures & Tables
References
- 1.Centers for Medicare & Medicaid Services Clinical laboratory improvement amendments of 1988. https://www.cms.gov/regulations-and-guidance/legislation/clia.
- 2.College of American Pathologists Cap accreditation checklists. https://www.cap.org/laboratory-improvement/accreditation/accreditation-checklists.
- 3.Lewis Patrick, Perez Ethan, Piktus Aleksandra, Petroni Fabio, Karpukhin Vladimir, Goyal Naman, Küttler Heinrich, Lewis Mike, Yih Wen-tau, Rocktäschel Tim, et al. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in neural information processing systems. 2020;33:9459–9474. [Google Scholar]
- 4.Javadi Saeedeh, Mirabi Sara, Gangar Manan, Ofoghi Bahadorreza. When evidence contradicts: Toward safer retrieval-augmented generation in healthcare. 2025.
- 5.Wada Akihiro, Tanaka Yusuke, Nishizawa Masato, et al. Retrieval-augmented generation elevates local llm quality in radiology contrast media consultation. npj Digital Medicine. 8(395):2025. doi: 10.1038/s41746-025-01802-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Ji Yuelyu, Zhang Hang, Wang Yanshan. Bias evaluation and mitigation in retrieval-augmented medical question-answering systems. AMIA Annual Symposium. 2025. [PMC free article] [PubMed]
- 7.Huang Haoyu, Huang Yongfeng, Yang Junjie, Pan Zhenyu, Chen Yongqiang, Ma Kaili, Chen Hongzhi, Cheng James. Retrieval-augmented generation with hierarchical knowledge. 2025.
- 8.Sriram A., Sundan M. N, B., Krishnamoorthy S. Context-aware retrieval-augmented generation for artificial intelligence in urology. Cureus. Jul 17 2025;17(7):e88167. doi: 10.7759/cureus.88167. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Wu Junde, Zhu Jiayuan, Qi Yunli, Chen Jingkun, Xu Min, Menolascina Filippo, Grau Vicente. Medical graph rag: Towards safe medical large language model via graph retrieval-augmented generation. 2024.
- 10.Es Shahul, James Jithin, Espinosa-Anke Luis, Schockaert Steven. Ragas: Automated evaluation of retrieval augmented generation. 2025.
- 11.Hu Xiangkun, Ru Dongyu, Qiu Lin, Guo Qipeng, Zhang Tianhang, Xu Yang, Luo Yun, Liu Pengfei, Zhang Yue, Zhang Zheng. Refchecker: Reference-based fine-grained hallucination checker and benchmark for large language models. 2024.
- 12.Qwen, Yang An, Yang Baosong, Zhang Beichen, Hui Binyuan, Zheng Bo, Yu Bowen, Li Chengyuan, Liu Dayiheng, Huang Fei, Wei Haoran, Lin Huan, Yang Jian, Tu Jianhong, Zhang Jianwei, Yang Jianxin, Yang Jiaxi, Zhou Jingren, Lin Junyang, Dang Kai, Lu Keming, Bao Keqin, Yang Kexin, Yu Le, Li Mei, Xue Mingfeng, Zhang Pei, Zhu Qin, Men Rui, Lin Runji, Li Tianhao, Tang Tianyi, Xia Tingyu, Ren Xingzhang, Ren Xuancheng, Fan Yang, Su Yang, Zhang Yichang, Wan Yu, Liu Yuqiong, Cui Zeyu, Zhang Zhenru, Qiu Zihan. Qwen2.5 technical report. 2025.
- 13.Dubey Abhimanyu, Jauhri Abhinav, Pandey Abhinav, Kadian Abhishek, Al-Dahle Ahmad, Letman Aiesha, Mathur Akhil, Schelten Alan, Yang Amy, Fan Angela, et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783. 2024.
- 14.Jin Qiao, Kim Won, Chen Qingyu, Comeau Donald C, Yeganova Lana, John Wilbur W, Lu Zhiyong. Medcpt: Contrastive pre-trained transformers with large-scale pubmed search logs for zero-shot biomedical information retrieval. Bioinformatics. November 2023;39(11) doi: 10.1093/bioinformatics/btad651. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Christophe Clément, Raha Tathagata, Maslenkova Svetlana, Salman Muhammad Umar, Kanithi Praveen K, Pimentel Marco AF, Khan Shadab. Beyond fine-tuning: Unleashing the potential of continuous pretraining for clinical llms. 2024.
- 16.Yue Yang, Chen Zhiqi, Lu Rui, Zhao Andrew, Wang Zhaokai, Yue Yang, Song Shiji, Huang Gao. Does reinforcement learning really incentivize reasoning capacity in llms beyond the base model? 2025.
- 17.Gekhman Zorik, Yona Gal, Aharoni Roee, Eyal Matan, Feder Amir, Reichart Roi, Herzig Jonathan. Does fine-tuning llms on new knowledge encourage hallucinations? 2024.


