Skip to main content
UKPMC Funders Author Manuscripts logoLink to UKPMC Funders Author Manuscripts
. Author manuscript; available in PMC: 2026 Jun 16.
Published in final edited form as: Immunoinformatics (Amst). 2025 Sep;19:100052. doi: 10.1016/j.immuno.2025.100052

AnalyzAIRR: A user-friendly guided workflow for AIRR data analysis

Vanessa Mhanna a,b, Gabriel Pires b, Grégoire Bohl-Viallefond a, Karim El Soufi b, Nicolas Tchitchek a, David Klatzmann a,b, Adrien Six a,‡, Hang P Pham c,†,‡, Encarnita Mariotti-Ferrandiz a,d,‡,*
PMCID: PMC7619189  EMSID: EMS214207  PMID: 42305137

Abstract

The analysis of bulk adaptive immune receptor repertoires (AIRR) enables the understanding of immune responses in both normal and pathological conditions. However, the complexity of AIRR calls for advanced, specialized methods to extract meaningful biological insights. These sophisticated approaches often present challenges for researchers with limited bioinformatics expertise, hindering access to comprehensive immune system analysis. To address this challenge, we developed AnalyzAIRR, an AIRR-compliant R package enabling advanced bulk AIRR sequencing data. The tool integrates state-of-the-art statistical and visualization methods applicable at various levels of granularity. It offers a platform for general data exploration, filtering and manipulation, and in-depth cross-comparisons of AIRR datasets, aimed at answering specific biological questions. We illustrate AnalyzAIRR functionalities using a published murine dataset of 18 T-cell receptor repertoires from three diferrent T cell subsets. We first detected and removed a major contaminant in a group of samples, before proceeding with the comparative analysis. Subsequent cross-sample analysis revealed differences in repertoire diversity that aligned with the respective cell phenotypes, and in repertoire convergence among the studied subsets. AnalyzAIRR’s set of analytical metrics is integrated into a Shiny web application and complemented with a tutorial to help users in their analytical strategy, making it user-friendly for biologists with little or no background in bioinformatics.

Keywords: Bulk high-throughput sequencing, AIRR, TCR, BCR, Statistical analyses, Repertoire diversity, Repertoire convergence

Introduction

B-cell receptors (BCR) and T-cell receptors (TCR) are engaged in antigen recognition by B and T cells respectively, sustaining the antigen-specificity of the adaptive immune response. They result from a sophisticated mechanism of genetic recombination of IG (for BCR) or TR (for TCR) genes, forming the adaptive immune receptor repertoire (AIRR). This extremely diverse repertoire carries the immune response capacity of every individual. Thus, AIRRs have been extensively studied to better understand the role of the adaptive immune response in various physiopathological conditions [1–3], disease progression [4–6], response to treatment [7,8] or vaccination [9,10], and maintenance of homeostasis [11–15].

High-throughput sequencing (HTS) tools have allowed over the last decade a faster and deeper profiling of AIRRs. Notably, it has permitted the generation of millions of sequences to help characterize the AIRR, mostly through bulk sequencing [16–18]. This breakthrough led to the advent of sophisticated computational tools [19]. On the one hand, AIRR pre-processing and annotation tools for TR and IG rearrangement sequencing data were developed, such as MiXCR [20], IMSEQ [21], IgBLAST [22], and pRESTO (pre-processing only) from the Immcantation suite [23]. Annotation involves mapping of V, D, and J genes, extracting the complementary determining region 3 (CDR3), and assembling and reconstructing AIRR rearrangements, known as clonotypes. On the other hand, strategies have been adapted from various fields such as ecology and information theory to describe, quantitatively and qualitatively, different aspects of the AIRR [24,25]. Such strategies, including diversity estimation, are implemented in tools such as ImmuneREF [26] and Immunarch [27], as well as in Immcantation, which is an end-to-end framework connecting multiple tools for high-throughput AIRR-seq dataset analysis [28].

Despite the wide range of analytical strategies proposed by most of these tools, they all require bioinformatics and programming skills, limiting the accessibility for non-bioinformatician researchers. Therefore, there is a need to better inform the scientific community of how and when to use the different suitable metrics for accurate AIRR studies.

Herein, we present AnalyzAIRR, an R package that proposes a systematic analytical routine to guide bioinformatics expert and non-expert users in AIRR analytical strategies. The package offers various statistical metrics and visualization methods, allowing a complete data exploration as well as cross-sample comparisons to answer defined biological questions. Advanced users can customize the visualizations generated by AnalyzAIRR in terms of style to better align with their preferences. Additionally, they can extract statistical data for use in other visualization platforms or tailor the presentation of the data to address specific research needs. Importantly, we complemented AnalyzAIRR with an intuitive R-Shiny web framework, making it accessible to biologists with limited bioinformatics expertise. This user-friendly interface allows them to explore and implement the proposed analytical strategies, bridging the gap between complex data analysis and practical research applications.

Design and implementation

Functionalities

AnalyzAIRR integrates a comprehensive set of methods for IG/TR repertoire analysis into an R package, accessible to both bioinformatics experts and non-experts. Three key functionalities are offered by AnalyzAIRR (Fig. 1).

Fig. 1. Workflow schematic of the AnalyzAIRR package.

Fig. 1

Comprehensive data exploration

AnalyzAIRR provides tools for the examination of individual samples or the entire dataset. This is achieved through the calculation of a set of descriptive statistics such as the number of sequences and the number of clones (corresponding to each unique V-CDR3aa-J combination), the V and J gene usage, as well as measures of clonal distribution and diversity. Among the diversity measures, Rényi and Hill indices are used to quantify diversity by capturing both species richness (the number of clones) and evenness (how evenly sequences are distributed among clones). The Rényi index generalizes several common diversity metrics, such as the Shannon and the Simpson indices, by adjusting a parameter to account for rare or abundant species, while the Hill numbers translate these indices into effective numbers of species [29,30]. This process facilitates the detection of potential outliers or contaminants within the data.

Advanced data manipulation

This functionality allows users to filter out specific sequences or samples based on user-defined criteria. In addition, two normalization methods are proposed to ensure comparability across samples: i) random down-sampling, and ii) exclusion of low-count sequences using Shannon entropy [31]. The latter further allows the correction of the clonotype distribution in small samples which are often over sequenced.

In-depth cross sample comparisons

AIRR comparisons between sample groups are possible through staregies such as, diversity estimation, repertoire similarity evaluation and the identification of differentially expressed genes/clonotypes between experimental groups. These comparative analyses enable the identification of significant patterns and differences, helping researchers address specific biological questions.

Data organization

Input data

AIRR sequencing datasets include a collection of V(D)J rearrangements, either IG or TR, also named clonotypes. Their analysis requires primarily an annotation step using dedicated alignment tools to transform raw sequencing data into analyzable immune repertoire information. AnalyzAIRR supports input files: (i) generated by a list of annotation tools including MiXCR and immunoSEQ; (ii) formatted following the MiAIRR standard guidelines from the AIRR Community [32]; or (iii) formatted following AnalyzAIRR requirements. A descriptive metadata providing sample information, such as experimental group, age, genetic background, sample type, and cell subset, can also be supplied allowing intra- and inter-group analyses.

AnalyzAIRR data structure

AnalyzAIRR employs an R-structured data container, known as an S4 object, that ensures efficient manipulation of the complex AIRR-seq data throughout the analysis workflow. AnalyzAIRR’s object, named RepSeqExperiment, can be created using the annotated AIRR data files and the metadata, when available, and is used as data input for all AnalyzAIRR functions. Alternatively, users can upload annotated files and metadata directly into the Shiny interface, which will automatically create and load the RepSeqExperiment object to be used in their analysis.

The RepSeqExperiment object is composed of four compartments, each designed to hold a specific type of data: (i) assayData, which encloses all the clonotype tables in the dataset; (ii) metaData, which contains the sample information provided by the user as well as a set of statistics calculated for each sample including the number of sequences and clones; (iii) otherData, which contains the output of many filtering and calculation functions further exemplified in the next section; and (iv) History, which registers all the actions performed on the RepSeqExperiment object.

AIRR-Shiny interface

Codeless data analysis

Shiny-AnalyzAIRR offers a user-friendly interface for biologists to perform comprehensive AIRR-seq data analysis, regardless of their bioinformatics expertise (S1 Fig). Users can apply the full range of analytical methods proposed by AnalyzAIRR, guided by detailed documentation that explains each function and provides relevant use cases. The Shiny interface includes additional features, such as interactive plots that reveal information that might be missed in static visualizations.

Navigating Shiny-AnalyzAirr

Shiny-AnalyzAIRR facilitates interactive exploration and visualization of AIRR-seq data. The app features a homepage with a “Getting Started” guide that assists users through the initial data loading process. A sidebar on the left provides access to key sections: “Data Loading”, “Data Manipulation”, “Exploratory Analysis”, “One-Sample Analysis”, and “Multi-Sample Analysis”. Once the data is loaded, the main panel displays a summary of the repertoires, including metrics like the number of samples and sequences. This panel also features interactive plots and tables for in-depth analysis. Each main section offers plots and tables, accessible via tabs at the top. After completing their analysis, users can compile all results, along with customized descriptions for each figure, into a report downloadable in PDF, HTML, or PowerPoint format. Its intuitive design ensures that novice users can efficiently explore and analyze their data.

Results

To illustrate the above-mentioned AnalyzAIRR functionalities, we made use of a published dataset composed of 18 T cell receptor α (TRA) repertoires of effector T cells (Teff), naïve regulatory T cells (nTreg), and activated/memory regulatory T cells (amTreg) collected from the spleen of young C57BL/6 mice [8]. The files were aligned with MiXCR (v.4.4.1) and used as input along with the corresponding metadata file. Herein, we will analyze the TRA repertoire of the 18 samples, as a RepSeqExperiment object can be generated for one chain at a time.

Data validation

AnalyzAIRR proposes a series of analyses that can be performed on each sample within the dataset, enabling an exploratory description of the dataset. These approaches can serve as a tool to perform dataset validation, an essential step in processing AIRR-seq data.

To assess the clonotype richness and the sequencing depth of a sample, the Chao1 index and rarefaction curves serve complementary roles. Rarefaction curves assess the sufficiency of sequencing depth and sampling (Fig. 2A); A plateau in the curve suggests that sampling is yielding few or no new species, whereas a steep ascending curve indicates that additional sequencing may still yield new clones. In our example data, sequencing depth was sufficient in samples that show a flatenned curve, while further sequencing could be required for ones showing a steep curve.

Fig. 2. Exploratory analysis of the full dataset.

Fig. 2

A) Rarefaction curves are plotted for each sample, illustrating the relationship between the number of sequences randomly selected from the sample and the number of clones they represent. Each sample is represented by a unique color. B) The Chao1 richness estimation is plotted as a function of the observed number of ntClones in each sample. The diagonal dashed line represents the reference line. C) Renyi values were calculated at the aaClone level for each of the pre-defined alpha values and plotted for each sample. Samples are separated based on the cell subset to which they belong. D) Scatter plots are generated for each Teff sample pair. The lower triangle shows scatter points, each representing a unique aaClone and its corresponding frequency in one sample shown on the x-axis (column) and a second sample shown on the y-axis (in row). X and Y axes use a logarithmic scale. The upper triangle shows the Pearson correlation calculated for each panel in the lower triangle: ***: p-value 〈 0.001; **: p-value < 0.01; *: p-value < 0.05 and ns.: p-value 〉 0.05. E) Curves plotting the aaClone frequencies as a function of their occurrence rank within each sample. X and Y axes use a logarithmic scale. Only the top 10 clones are plotted.

To estimate the total clonal richness, including clones that are likely present but remain undetected due to limited sampling and/or sequencing, the Chao1 index can be used (Fig. 2B). A higher Chao1 value reflects a greater estimated species richness, as the index measures undetected species by inferring their presence from the number of rare species observed. Consistently, samples with a high Chao1 estimation were ones showing a steep rarefaction curve, indicating that additional sampling would likely reveal more unique clones. Together, these two metrics provide a comprehensive understanding of species richness and the adequacy of sampling across different samples.

Furthermore, we assessed repertoire diversity using Rényi entropy profiles, a generalization of classic diversity measures that introduces a parameter alpha to emphasize different aspects of diversity. At alpha = 0, it corresponds to species richness by counting the number of distinct clones; at alpha = 1, it equals the Shannon index balancing richness and evenness. By plotting Rényi diversity profiles (curves of diversity as a function of alpha), we can compare both the richness and evenness of different repertoires: higher curves indicate greater diversity, and the shape of the curve reveals whether diversity is driven more by rare or abundant clones. Here, the Renyi curves showed a sharp decline in a group of samples, all being Teff cells (Fig. 2C), indicating that the diversity is driven by one or multiple abundant clone. This observation is not expected in such a polyclonal and diverse cell subset. In addition, nearly perfect clone frequency correlations were observed between each pair of Teff samples, caused by one shared clone present at a significant higher frequency than the other clones (Fig. 2D and E). This is once more an unusual observation in polyclonal Teff repertoires collected from different mice. Hence, we were able to detect a contaminant clone, present in high occurrences in the Teff samples.

Single sample analysis and filtering

In view of these observations, and to confirm the validaty of our results, we examined each of the problematic Teff sample individually. The analysis performed on one Teff sample revealed a decrease in its repertoire diversity, as evidenced by both the Simpson and Berger-Parker indices, compared to the other samples within the dataset (Fig. 3A). This reduced diversity can be explained by the high cumulative frequency of clones with occurrences higher than 10,000 (Fig. 3B), suggesting significant clonal expansions. Specifically, these expansions were driven by a single dominant clone, which exhibited a significantly higher count than subsequent clones within the same repertoire (Fig. 3C). This clone expressed one V-J combination, thus altering the overall distribution of V and J genes within the sample (Fig. 3D). These along with the previous results allowed the identification of the contaminant clone sequence, which was filtered out from all Teff samples. The filtering rescued the diversity and the clonal sharing in all Teff samples (S2 Fig).

Fig. 3. Single Teff sample analysis.

Fig. 3

A) Boxplots showing different diversity indices calculated on all the dataset samples. The sample of interest “tripod-26–682″ is represented by a red cross. B) Pie charts representing the distribution (left) and the cumulative frequency (right) of clones within each count interval for the sample of interest. C) A circular treemap plotting the top 1 % aaClones in the sample of interest. Each circle corresponds to a unique clone, and its size is proportional to the clone count. This plot shows the clone distribution within the sample repertoire. D) Heatmap of V and J combination usage within the sample of interest. Frequencies, represented by the color scale, are scaled column-wise. Barplots at the top and right side of the heatmap show the usage of each gene across the row or the column for the V and J genes, respectively.

Sample normalization

Comparing AIRR-seq data collected from different sources can often result in repertoires of differing sizes, due to biological or technical factors. A key feature of AnalyzAIRR is its ability to normalize data, ensuring comparable sample sizes and accurate comparisons in AIRR-seq datasets. AnalyzAIRR offers, among others, a Shannon-based down sampling method, leveraging the Shannon entropy to exclude sequences considered as noise [31]. While this normalization process resulted in a decrease in repertoire richness and diversity across all our samples (Fig. 4A and B), the pre- and post-normalization repertoires clustered per sample based on the Jaccard scores, which measures the similarity between two samples independently of the species abundance [33] (Fig. 4C). With the normalization applied; the datasets are now prepared for extensive comparative analysis.

Fig. 4. The effect of the Shannon normalization on the repertoires.

Fig. 4

A) Barplots showing the number of nt clones (V-CDR3nt-J combinations) for each of the not normalized (pink) and normalized (brown) samples. B) The Renyi values were calculated at the nt clone level at pre-defined alpha. For each of the not non-normalized and normalized groups, the mean (circles) and standard error (shade) are plotted. C) A heatmap showing the Jaccard distances calculated at the aa clone level between all pairwise samples. Samples are labeled column-wise based on the sample name and the normalization condition. Scores range from 0 (complete similarity) to 1 (complete dissimilarity). The Ward’s method was used to perform the hierarchical clustering.

Gene usage analysis

Users can leverage AnalyzAIRR to investigate biological hypotheses, by drawing a detailed comparative analysis between experimental groups and conditions. The first layer of any AIRR-seq data analysis involves comparing the gene usage between conditions. In our dataset, for visual clarity, we compared TRAV and TRAJ usages between the naïve and activated Treg cells. Comparing the gene frequencies showed significant differences in the usage of many gene families (Fig. 5A and B). Additionally, a differential expression analysis performed at the V-J combination level enabled the identification of differentially expressed gene combinations, contributing to the separation of the two subsets in the multidimensional scaling (MDS) analysis (Fig. 5C and D).

Fig. 5.

Fig. 5

Gene usage analysis between the Treg subsets. Comparison of the TRAV (A) and TRAJ (B) frequencies between amTregs (blue) and nTregs (green) samples. Lines on barplots correspond to the standard error values. A Wilcoxon test was performed and only significant values, when present, are plotted. C) A volcano plot showing differentially expressed TRAV-TRAJ combinations between amTreg and nTreg samples. Over-expressed combinations in the amTreg population are shown in red, down-expressed one in blue based on the fixed thresholds: a fold change ≥ ±2 and a pvalue ≤ 0.05. Only the 10 most differentially expressed combinations are labeled. D) Morisita-Horn distances was calculated on TRAV-TRAJ counts between all pairwise Treg samples and used to perform a multidimensional scaling analysis on a two-dimensional space. Each point corresponds to a sample and is colored based on the cell subset it belongs to; amTregs in blue and nTregs in green. A Wilcoxon test is applied and adjusted p-values using the Bonferroni method are shown. *P < 0.05; **P < 0.01; ***P < 0.001; ****P < 0.0001; ns P > 0.05.

Clonality and diversity assessment between conditions

We next compared the repertoire clonality of the three different cell subsets. Results showed higher clonality in Teffs and nTregs compared to amTregs, with a significant difference between amTregs and Teffs (Fig. 6A). We also examined any potential sex-related differences, revealing comparable clone numbers in male and female mice in each cell population (Fig. 6B). Furthermore, Renyi curves showed distinct profiles, with nTregs exhibiting the most diverse repertoire and amTregs the least diverse (Fig. 6C), in agreement with the clonality results. Finally, we confirmed this observation by computing and comparing the cumulative frequency of expanded clones, which were found to be higher in amTregs compared to the other subsets (Fig. 6D).

Fig. 6. Comparison of the repertoire diversity between the three cell subsets.

Fig. 6

A) The number of aaClones (bottom) is compared between amTreg (blue), nTreg (green) and Teff (orange) samples. B) The number of aaClones is compared between samples collected from female (green) and male (purple) mice within each cell subset. A-Boxplots represent the median across all samples belonging to the same group. A Wilcoxon test is applied and adjusted p-values using the Holm method are shown. C) The Renyi values were calculated at the aaClone level at the pre-defined alpha values. For each specified group, the mean (circles) and standard error (shade) are plotted. D) aaClones were grouped into intervals based on their frequency (for example, the interval]0.001,0.01] includes aaClones whose frequencies falling within this range). The cumulative frequency of aaClones within each frequency interval was compared between the three cell subsets. Boxplots represent the median across all samples belonging to the same group. A paired wilcoxon test is applied and adjusted p-values using the Bonferroni method are shown. *P < 0.05; **P < 0.01; ***P < 0.001; ****P < 0.0001; ns P > 0.05.

Repertoire composition comparison

The similarity metrics proposed by AnalyzAIRR enabled us to examine the repertoire composition between the different cell subsets, providing insights into repertoire convergence. For instance, hierarchical clustering based on the Jaccard distance separated amTreg samples from Teff and nTreg samples, despite the high dissimilarity scores that indicate modest differences in repertoire compositions (Fig. 7A). MDS analysis applied to the calculated Jaccard scores, also facilitated by AnalyzAIRR, confirmed these findings (Fig. 7B). In contrast, the Morisita-Horn index, which considers and is sensitive to species abundance, allowed clear discrimination between the three cell subsets in both hierarchical clustering and MDS analysis (Fig. 7C and D). Interestingly, when incorporating mouse sex information into the MDS analysis, we observed that Teff similarity seems to be linked to sex. In contrast, this association was not seen for amTregs and nTregs, highlighting an area for future investigation. Moreover, calculating the inter-sample similarity scores revealed that nTregs had the highest similarity, suggesting public repertoires, while amTregs showed the lowest inter-similarity samples, indicating that their repertoires are comparatively more private.

Fig. 7.

Fig. 7

Comparison of the repertoire composition. A heatmap showing the Jaccard (A) and Morisita-Horn (C) distances calculated at the aaClone level between all pairwise samples. Scores range from 0 (complete similarity) to 1 (complete dissimilarity). Samples are labeled column-wise based on the cell subset and sex of the mice used in the study. The Ward’s method was used to perform the hierarchical clustering. The calculated distances were used to perform a multidimensional scaling analysis on a two-dimensional space (B and D). Each point corresponds to a sample and is colored based on the cell subset it belongs to; amTregs in blue, nTregs in green, and Teffs in orange. Ellipses are drawn around points belonging the same group.

Conclusion

AnalyzAIRR equips researchers with a scalable, AIRR-compliant framework accompanied with a step-by-step documentation for indepth exploration of immune receptor repertoires. Using the source code or the Shiny web interface, users can navigate a sequential analysis pipeline, encompassing data quality control, sample exploration, and cross-sample comparisons. In this study, we introduced the various analyses offered by AnalyzAIRR through a direct application on a published TCR-seq dataset. By leveraging its functionalities, we were able to illustrate relevant biological insights, validating the tool’s capability to support both expert and non-expert users in the field in conducting comprehensive repertoire analyses. AnalyzAIRR distinguishes itself with comprehensive analytical features coupled with a user-friendly interface (Fig. 8). This combination empowers researchers, regardless of their computational background, to analyze AIRR-seq data, generate meaningful insights, and create publication-ready visualizations. While existing tools may excel in specific analytical tasks, they still require extensive coding. Shiny-AnalyzAIRR offers a broad range of functionalities within a single, accessible platform. The user-friendly design, featuring interactive plots and guided workflows, lowers the barrier to entry for biologists who are new to AIRR-seq.

Fig. 8.

Fig. 8

Overview of community-developed tools for the analysis of bulk AIRR-seq data. This table provide a comparative overview of the most used communitydeveloped tools for analyzing bulk AIRR-seq data. The list of criteria includes input data type and pre-processing, data filtering and manipulation, exploratory and cross-sample analysis, somatic hypermutation and phylogenetic (lineage tree) analysis, the availability of a tool documentation, presence of an interactive interface and easy-of-use score for biologists.

Supplementary Material

Supplementary materials

Supplementary material associated with this article can be found, in the online version, at doi:10.1016/j.immuno.2025.100052.

Figure S1, Figure S2

Acknowledgements

We thank Pierre Barennes, Valentin Quiniou and Wahiba Chaara for their contribution in conceiving the package.

Fundings

This work was funded by the TRiPoD European Research Council-Advanced EU (No. 322856) and iMAP (ANR-16-RHUS-0001) grants to DK, the iReceptorPlus (Horizon 2020No 825821) and the AIR-MI (ANR-18-ECVD-0001) and TALISMAN (ANR-23-CE17-0039) grants to EMF with additional support from Sorbonne Université, INSERM and AP-HP. ILTOO Pharma supported HPP.

Footnotes

CRediT authorship contribution statement

Vanessa Mhanna: Writing – review & editing, Writing – original draft, Software, Formal analysis. Gabriel Pires: Writing – review & editing, Software. Grégoire Bohl-Viallefond: Writing – review & editing, Software. Karim El Soufi: Writing – review & editing, Software. Nicolas Tchitchek: Writing – review & editing, Software. David Klatzmann: Writing – review & editing. Adrien Six: Writing – review & editing, Conceptualization. Hang P. Pham: Writing – review & editing, Software, Conceptualization. Encarnita Mariotti-Ferrandiz: Writing – review & editing, Writing – original draft, Conceptualization.

Declaration of competing interest

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

Data availability

The example data analysed in this paper and the code developed to generate the plots are publicly available at Zenodo under the DOI: 10.5281/zenodo.13350295.

Availability

AnalyzAIRR is implemented in R version 4.2.2 (License: GPL-3) and available for installation at: https://github.com/i3-unit/AnalyzAIRR. A vignette providing installation details, function documentation and use cases is also available at: https://i3-unit.github.io/AnalyzAIRR/. Shiny-AnalyzAIRR can be downloaded from https://github.com/i3-unit/Shiny-AnalyzAIRR and run locally or any device. Alternatively, it has been deployed on the web and can be used directly at https://i3-analyzairr.fr. The example data analysed in this paper and the code developed to generate the plots are publicly available at Zenodo under the DOI: 10.5281/zenodo.15448131.

References

  • [1].Bashford-Rogers RJM, et al. Analysis of the B cell receptor repertoire in six immune-mediated diseases. Nature. 2019 Oct;574(7776):122–6. doi: 10.1038/s41586-019-1595-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [2].Seay HR, et al. Tissue distribution and clonal diversity of the T and B cell repertoire in type 1 diabetes. JCI Insight. 2016 Dec;1(20) doi: 10.1172/jci.insight.88242. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [3].Krasik SV, et al. Systematic evaluation of intratumoral and peripheral BCR repertoires in three cancers. 2024 Jan 25; doi: 10.7554/eLife.89506. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [4].Chang C-M, Feng P, Wu T-H, Alachkar H, Lee K-Y, Chang W-C. Profiling of T cell repertoire in SARS-CoV-2-infected COVID-19 patients between mild disease and pneumonia. J Clin Immunol Aug. 2021;41(6):1131–45. doi: 10.1007/s10875-021-01045-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [5].Minervina AA, et al. Longitudinal high-throughput TCR repertoire profiling reveals the dynamics of T-cell memory formation after mild COVID-19 infection. eLife. 2021 Jan;10:e63502. doi: 10.7554/eLife.63502. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [6].Stern JNH, et al. B cells populating the multiple sclerosis brain mature in the draining cervical lymph nodes. Sci Transl Med. 2014 Aug;6(248) doi: 10.1126/scitranslmed.3008879. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [7].Amoriello R, et al. The TCR repertoire reconstitution in multiple sclerosis: comparing one-shot and continuous immunosuppressive therapies. Front Immunol. 2020 Apr;11:559. doi: 10.3389/fimmu.2020.00559. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [8].Mhanna V, et al. Impaired activated/memory regulatory T cell clonal expansion instigates diabetes in NOD mice. Diabetes. 2021 Apr;70(4):976–85. doi: 10.2337/db20-0896. [DOI] [PubMed] [Google Scholar]
  • [9].Pogorelyy MV, et al. Precise tracking of vaccine-responding T cell clones reveals convergent and personalized response in identical twins. Proc Natl Acad Sci USA. 2018 Dec;115(50):12704–9. doi: 10.1073/pnas.1809642115. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [10].Hoehn KB, et al. Human B cell lineages associated with germinal centers following influenza vaccination are measurably evolving. eLife. 2021 Nov;10:e70873. doi: 10.7554/eLife.70873. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [11].Quiniou V, et al. Human thymopoiesis produces polyspecific CD8+ α/β T cells responding to multiple viral antigens. eLife. 2023 Mar;12:e81274. doi: 10.7554/eLife.81274. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [12].Miho E, Roškar R, Greiff V, Reddy ST. Large-scale network analysis reveals the sequence space architecture of antibody repertoires. Nat Commun Mar. 2019;10(1):1321. doi: 10.1038/s41467-019-09278-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [13].Briney B, Inderbitzin A, Joyce C, Burton DR. Commonality despite exceptional diversity in the baseline human antibody repertoire. Nature. 2019 Feb;566(7744):393–7. doi: 10.1038/s41586-019-0879-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [14].Meng W, et al. An atlas of B-cell clonal distribution in the human body. Nat Biotechnol. 2017 Sep;35(9):879–84. doi: 10.1038/nbt.3942. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [15].Madi A, et al. T cell receptor repertoires of mice and humans are clustered in similarity networks around conserved public CDR3 sequences. eLife. 2017 Jul;6 doi: 10.7554/eLife.22057. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [16].Rosenfeld AM, Meng W, Luning Prak ET, Hershberg U. ImmuneDB, a novel tool for the analysis, storage, and dissemination of Immune repertoire sequencing data. Front Immunol. 2018 Sep;9:2107. doi: 10.3389/fimmu.2018.02107. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [17].Christley S, et al. VDJServer: a cloud-based analysis portal and data commons for Immune repertoire sequences and rearrangements. Front Immunol. 2018 May;9:976. doi: 10.3389/fimmu.2018.00976. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [18].Corrie BD, et al. iReceptor: a platform for querying and analyzing antibody/B-cell and T-cell receptor repertoire data across federated repositories. Immunol Rev. 2018 Jul;284(1):24–41. doi: 10.1111/imr.12666. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [19].Mhanna V, et al. Adaptive immune receptor repertoire analysis. Nat Rev Methods Primers. 2024 Feb;4(1):12. doi: 10.1038/s43586-024-00299-2. [DOI] [Google Scholar]
  • [20].Bolotin DA, et al. MiXCR: software for comprehensive adaptive immunity profiling. Nat Methods. 2015 May;12(5):380–1. doi: 10.1038/nmeth.3364. [DOI] [PubMed] [Google Scholar]
  • [21].Kuchenbecker L, et al. IMSEQ—A fast and error aware approach to immunogenetic sequence analysis. Bioinformatics. 2015 Sep;31(18):2963–71. doi: 10.1093/bioinformatics/btv309. [DOI] [PubMed] [Google Scholar]
  • [22].Ye J, Ma N, Madden TL, Ostell JM. IgBLAST: an immunoglobulin variable domain sequence analysis tool. Nucleic Acids Res. 2013 Jul;41(W1):W34–40. doi: 10.1093/nar/gkt382. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [23].Vander Heiden JA, et al. pRESTO: a toolkit for processing high-throughput sequencing raw reads of lymphocyte receptor repertoires. Bioinformatics. 2014 Jul;30(13):1930–2. doi: 10.1093/bioinformatics/btu138. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [24].Miho E, Yermanos A, Weber CR, Berger CT, Reddy ST, Greiff V. Computational strategies for dissecting the high-dimensional complexity of adaptive immune repertoires. Front Immunol. 2018 Feb;9:224. doi: 10.3389/fimmu.2018.00224. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [25].Six A, et al. The past, present, and future of Immune repertoire biology – The rise of next-generation repertoire analysis. Front Immunol. 2013;4 doi: 10.3389/fimmu.2013.00413. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [26].Weber CR, et al. Reference-based comparison of adaptive immune receptor repertoires. Cell Reports Methods. 2022 Aug;2(8):100269. doi: 10.1016/j.crmeth.2022.100269. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [27].Nazarov V, Bot Immunarch, Rumynskiy E. immunomind/immunarch: 0.6.5: basic single-cell support. Zenodo. 2020 Jun 15; doi: 10.5281/ZENODO.3367200. [DOI] [Google Scholar]
  • [28].Gupta NT, Vander Heiden JA, Uduman M, Gadala-Maria D, Yaari G, Kleinstein SH. Change-O: a toolkit for analyzing large-scale B cell immunoglobulin repertoire sequencing data: Table 1. Bioinformatics. 2015 Oct;31(20):3356–8. doi: 10.1093/bioinformatics/btv359. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [29].Chiffelle J, Genolet R, Perez MA, Coukos G, Zoete V, Harari A. T-cell repertoire analysis and metrics of diversity and clonality. Curr Opin Biotechnol. 2020 Oct;65:284–95. doi: 10.1016/j.copbio.2020.07.010. [DOI] [PubMed] [Google Scholar]
  • [30].Chao A, Chiu C-H, Jost L. In: Biodiversity conservation and phylogenetic systematics. Pellens R, Grandcolas P, editors. Vol. 14. Cham: Springer International Publishing; 2016. Phylogenetic diversity measures and their decomposition: a framework based on Hill numbers; pp. 141–72. (Topics in Biodiversity and Conservation). [DOI] [Google Scholar]
  • [31].Chaara W, Gonzalez-Tort A, Florez L-M, Klatzmann D, Mariotti-Ferrandiz E, Six A. RepSeq data representativeness and robustness assessment by Shannon Entropy. Front Immunol. 2018 May;9 doi: 10.3389/fimmu.2018.01038. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [32].Rubelt F, et al. Adaptive immune receptor repertoire community recommendations for sharing immune-repertoire sequencing data. Nat Immunol. 2017 Nov;18(12):1274–8. doi: 10.1038/ni.3873. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [33].Jaccard P. The distribution of the flora of the Alpine Zone. New Phytologist. 1912 Feb;11(2):37–50. doi: 10.1111/j.1469-8137.1912.tb05611.x. [DOI] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Figure S1, Figure S2

Data Availability Statement

The example data analysed in this paper and the code developed to generate the plots are publicly available at Zenodo under the DOI: 10.5281/zenodo.13350295.

AnalyzAIRR is implemented in R version 4.2.2 (License: GPL-3) and available for installation at: https://github.com/i3-unit/AnalyzAIRR. A vignette providing installation details, function documentation and use cases is also available at: https://i3-unit.github.io/AnalyzAIRR/. Shiny-AnalyzAIRR can be downloaded from https://github.com/i3-unit/Shiny-AnalyzAIRR and run locally or any device. Alternatively, it has been deployed on the web and can be used directly at https://i3-analyzairr.fr. The example data analysed in this paper and the code developed to generate the plots are publicly available at Zenodo under the DOI: 10.5281/zenodo.15448131.

RESOURCES