Skip to main content

This is a preprint.

It has not yet been peer reviewed by a journal.

The National Library of Medicine is running a pilot to include preprints that result from research funded by NIH in PMC and PubMed.

ArXiv logoLink to ArXiv
[Preprint]. 2026 Jan 6:arXiv:2401.00102v2. Originally published 2023 Dec 29. [Version 2]

sentropy: A Python Package for Revealing Hidden Differences in Complex Datasets

Phuc Nguyen a, Rohit Arora b, Elliot D Hill c, Jasper Braun a, Alexandra Morgan a, Liza M Quintana a, Gabrielle Mazzoni d, Ghee Rye Lee e, Rima Arnaout f, Ramy Arnaout a,*
PMCID: PMC11275705  PMID: 39070042

Abstract

Machine-learning datasets are typically characterized by measuring their size and class balance. However, there exists a richer and potentially more useful set of measures, termed S-entropy (similarity-sensitive entropy), that incorporate elements’ frequencies and between-element similarities. Although these have been available in the R and Julia programming languages for other applications, they have not been as readily available in Python, which is widely used for machine learning, and are not easily applied to machine-learning-sized datasets without special coding considerations. To address these issues, we developed sentropy, a Python package that calculates S-entropy and is tailored to large datasets. sentropy can calculate any of the frequency-sensitive measures of Hill’s D-number framework and their similarity-sensitive counterparts. sentropy also outputs measures that compare datasets. We first briefly review S-entropy, illustrating how it incorporates elements’ frequencies and elements’ pairwise similarities. We then describe sentropy’s key features and usage. We end with several examples—immunomics, metagenomics, computational pathology, and medical imaging—illustrating sentropy’s applicability across a range of dataset types and fields.

Keywords: data science, diversity, frequency, similarity, Shannon entropy, Simpson’s index, metagenomics, immunomics, computational pathology, medical imaging, machine learning, Python

1. Introduction

Assessing the size and quality of a dataset is an important part of machine learning in many fields (Table 1). How many unique elements are there, and how do they relate to each other? How do they cluster? Summary statistics provide initial answers and can form the basis for deeper questions. Historically, different statistics have featured more prominently in different areas of investigation. Examples include species richness in ecology, Shannon entropy in computer science, and Simpson’s index in biology. In several fields, the practice is to bin elements before calculating statistics. For example, in metagenomics, sequences are usually binned into operational taxonomic units (OTUs); in immunomics, antibody and T-cell receptor genes are often binned by lineage or antigen specificity. In machine learning, elements are binned by label to calculate class balance.

Table 1:

Areas and examples amenable to dataset composition analysis using sentropy.

Biological sciences Medical sciences Physical Sciences Engineering & Industry Social Sciences
Immunomes (immunological diversity and comparisons)
Microbiomes (metagenomics diversities and comparisons)
Viromes (numbers of different viruses)
Transcriptomes (diversity including splice variants)
Sequencing libraries (estimates of completeness)
Ecosystems/biomes (effective numbers of species)
Gene-interaction networks
Medical imaging datasets (for AI/ML)
Diseases (frequency and categorization)
Patient cohorts (representativeness)
Medications (targets or mechanisms)
Foods (assessments of dietary richness)
Astronomical objects (planetary types)
Geography and terrain (quantifying differences)
Chemical compounds (diversity and novelty)
Minerals (relationships to each other)
ML models/ensembles
Machine parts (complexity analysis)
Supply chains (robustness analysis)
Industries (anti-trust analysis)
Investment portfolios (diversification analysis)
Populations (population-level diversity)
Societies (diversity of opinions and beliefs)
Political organizations (diversity of opinions and positions)
Boards of directors
Artwork (assessment and comparisons of differences)
Classrooms (diversity of representation)
Admission/hiring/promotion/review committees
Social networks
Sports teams and leagues (composition, strategies)

1.1. Frequency-sensitive diversity measures

It has long been known that the summary statistics above, among others, are related by how they account for differences in the relative frequency of repeated dataset elements [1]. For example, species richness, Shannon entropy, and Simpson’s index, in that order, place an increasing emphasis on elements’ frequencies, essentially upweighting elements that appear more frequently by having them contribute more to the statistic’s final value. For example, in metagenomic datasets, species richness weights all OTUs the same regardless of how frequent they are, whereas in Simpson’s index common OTUs count more. In fact, these statistics are highly related to each other in that they are specific instances of a more general formula, known as Hill’s diversity framework, in which the key parameter is the frequency down-weighting q (sometimes known as the “viewpoint parameter”) [2]:

Dq=∑i=1npiq1/(1−q) (1)

(Note Dq also appears in the literature as (qD.) In this equation, the dataset contains n unique elements and pi is the frequency of the ith unique element in the dataset. In the limit of q = 1, Eq. 1 approaches exp−∑inpi ln pi, which the exponential of the Shannon entropy [3]. The only substantive difference between species richness, Shannon entropy, Simpson’s index, and many other summary statistics is in the value of q: q = 0 for species richness (all elements count the same), q = 1 for Shannon entropy (some down-weighting of rare elements), q = 2 for Simpson’s index (more down-weighting), and so on up to q = ∞ for the Berger-Parker index (maximum down-weighting) [4], for which the frequency of only the most common element matters.

1.2. Effective numbers

The D numbers of Hill’s framework yield the so-called effective number forms of Shannon entropy, Simpson’s index, and so on. These give the effective number of elements in a dataset, for a given weight q. For example, if two immunology datasets contain the same number of unique antibody sequences but the first dataset has a sequence that accounts for 90% of all sequences in the dataset (as can happen in leukemia), the first dataset will have a lower effective number of sequences for q > 0. Likely the effective number of species will be not much more than 1, since to a good approximation—“effectively”—the dataset consists of just the one dominant sequence. Fig. 1 illustrates this concept in a discipline-agnostic manner. The commonly used forms of the above statistics are simple mathematical transformations of the effective-number forms: for example, Shannon entropy is the logarithm of D1, while Simpson’s index is the reciprocal of D2. For the investigator, the advantages of using effective numbers are, first, effective numbers all have the same units, allowing them to be easily compared; and second, they behave nicely as datasets are combined or split apart. Non-effective-number forms generally lack these properties [5].

Figure 1:

Figure 1:

How frequency affects diversity (entropy). Each dataset contains six unique elements, so D0 = 6 for each dataset (the “0” in D0 makes D0 frequency insensitive). Thus, if we ignore frequency, the two datasets are equally diverse. However, Dataset 1a is mostly apples, while Dataset 1b has a nearly uniform distribution of different fruits. Thus, Dataset 1b is intuitively more diverse. Consistent with intuition, D1 is only 1.9 for Dataset 1a but nearly 6 for Dataset 2: for q = 1, Dataset 1a effectively has only 1.9 unique elements: one can think of this as the apple counting as one unique element and the other fruits in Dataset 1a collectively counting as nearly (0.9) the equivalent of a second unique element. Up-weighting common elements can also be thought of discounting rare ones; thus, different values of q can be thought of as different discount factors, and therefore lead to different effective numbers. For example, for q = ∞, D∞ = 1.17 for Dataset 1a and 5.83 for Dataset 2: with maximal discounting, Dataset 1a effectively has only fractionally more than a single element (effectively nearly all apples), whereas Dataset 1b still has close to all six (with a slight discount because there are only five bunches of grapes vs. six of each of the other fruits). The sentropy package can calculate any frequency-sensitive diversity measure.

The effective number of elements in a dataset (at some q) is a natural measure of dataset size, complexity, or diversity [6]. Because this number can also be thought of as a measure of how diverse the dataset is, the D numbers are also known as diversity measures (or diversity indices): “diversities of order q.” More specifically, when q > 0, they are frequency-sensitive diversity measures. Within- and between-dataset measures are known as alpha and beta diversity, respectively.

We note the intimate connection between these measures and the Rényi entropies (of which Shannon entropy is the best known): the D numbers are simply the exponentials of the Rényi entropies. For this reason, it is useful to think of diversity and entropy interchangeably, as the same fundamental concept, just expressed in different units: effective numbers for diversity and bits or nats for entropy. Effective numbers are preferred for their several advantages [6].

1.3. Similarity-sensitive diversity measures

When grouping elements is called for, such as for OTUs and antibody/TCR clusters, a similar approach to Eq. 1 can be used, provided that the similarity between each pair of elements is accounted for. Such accounting is possible through an extension of Hill’s framework that outputs similarity-sensitive versions of the D numbers, as shown by Leinster and Cobbold [2]. We refer to this as similarity-sensitive diversity, similarity-sensitive entropy, or S-entropy for short and will express them in effective-number form. For each pair of unique elements in the dataset, the investigator supplies a similarity between 0 and 1, with 0 meaning the two elements are completely different and 1 meaning they are identical. Similarities are provided as entries in a table or matrix. This is called the similarity matrix, Z, a square matrix of dimension n (where again n is the number of unique elements). In the context of network science or graph theory, Z can be thought of as an adjacency matrix with self-edges added to make the diagonal entries all equal to 1.

In practice, especially for large datasets, Z is populated computationally according to an investigator-provided formula or rule. Binning can be understood as a special case where the cell Zij that corresponds to the similarity between elements i and j is set to 1 if i and j are in the same bin and set to 0 otherwise. For example, for OTUs the similarity is be set to 1 if two 16S rRNAs have greater than a threshold-level of sequence identity (e.g. ≥ 95% or ≥ 99%), placing them in the same OTU, and 0 otherwise (but see below for a continuous alternative). However, in general similarities Zij can take any value between 0 and 1. They can be fractional (0.9, 0.327..., etc.) or even asymmetric, with element i more similar to element j than element j to element i; the decision is up to the investigator [7]. Some examples of similarity measures include the well-known structural similarity index measure (SSIM) for images [8], [9] and relative dissociation constants for proteins and other molecules [10].

Mathematically, incorporating Z yields the following equation:

DqZ=∑i=1npiZpiq−11/(1−q) (2)

Here, Zpi=∑jZijpj is the “ordinariness” of the ith unique element, i.e. the cumulative similarity of the ith unique element to all the other elements of the dataset [2]. The Z in DqZ designates these measures as being similarity-sensitive D numbers; the presence of q in Eq. 2 means that these S-entropy measures are also frequency sensitive, like the Dq numbers of 1. Fig. 2 provides a discipline-agnostic illustration. Note that interpreting Z as an adjacency matrix, Eq. 2 can be used to provide measures of network complexity in e.g. social networks or gene-interaction networks (Table 1).

Figure 2:

Figure 2:

How similarity affects diversity. Each dataset contains nine species (top), each at equal frequencies, so D0 = D1 = D2 = ⋅⋅⋅ = D∞ = 9 for each dataset. Thus, the two datasets are equally diverse using similarity insensitive measures. However, Dataset 2a is all birds, whereas Dataset 2b contains a wider variety of animals, making Dataset 2b intuitively more diverse. The similarity matrices Z are accordingly quite different (bottom). Consistent with the intuition, D0Z is only 1.06 for Dataset 2a, but 2.16 for Dataset 2b: Dataset 2a effectively has only 1.06 species (further interpreted as it contains just a single class or type—here, birds; the “extra” 0.06 reflects intragroup diversity among bird species). Meanwhile, Dataset 2b effectively has 2.16 species, corresponding roughly to the number of taxonomic phyla represented (Arthropoda and Chordata, which form two visible clusters in the similarity matrix). The sentropy package can calculate any similarity-sensitive diversity measure for any user-supplied definition of similarity.

1.4. Advantages of similarity-sensitive measures

S-entropic measures have several advantages over their similarity-insensitive counterparts. First, similarity-sensitive D numbers align better with an intuitive sense of what “diversity” and “complexity” should mean [6], [7]. For example, given two otherwise identical 16S rRNA microbiome datasets, the dataset that represents the greater number of taxa is intuitively more diverse. Consistent with this intuition, it will have the higher DqZ (see Metagenomics below). In contrast, the two datasets will have identical Dq, thus, the similarity-insensitive Dq measures will miss the main difference between these datasets.

Second, DqZ measures capture the effects of grouping without the need to choose a binning threshold (e.g. 95% vs. 99% sequence identity for OTUs [11], [12], or a radius of 1 vs. 2 vs. 3 nucleotide differences for clone membership [13]). There is also no need to place an element exclusively into a single group: multiple and/or fuzzy group membership is inherently allowed. This flexibility is especially useful for image datasets, where DqZ simultaneously captures the number of classes as well as the diversity within and between classes [14].

Third, DqZ measures provide rich mechanisms for exploring higher- and lower-level dataset structure, including a well-developed framework for partitioning within-(alpha) vs. between- (beta) dataset diversity. It also offers a ready method for investigating substructure using the ordinariness, Zpi. For example, to measure how “Enterobacterales-like” a given 16S rRNA dataset is, if the ordinariness to a given Enterobacterales sequence (or set of sequences) is high, the dataset is very Enterobacterales-like. This method could be useful for quantitative definitions and assessments of enterotypes (see Results).

We stress that the usefulness of S-entropy extends well beyond any one field. Table 1 lists several possible applications of S-entropic measures in different areas, from engineering to the natural and social sciences.

1.5. Representativeness as a measure of overlap between datasets

The S-entropic framework in Eq. 2 extends naturally to measures of between-dataset diversity, also known as beta diversity [15]. Given a pair of datasets A and B, β¯A describes how distinctive dataset A is, with a value of 1 meaning indistinguishable and 2 meaning completely distinct. (We write out “beta” when referring to beta diversity so as not to confuse it with β¯; unfortunately dual-use of this Greek letter is a convention in the literature.) Alternatively, we can use the representativeness, which is the inverse of distinctiveness and which is a more straightforward measure of over-lap: representativeness of 1 means the dataset A is completely representative of the pair, whereas a zero-overlap dataset has 0.5 representativeness, because it is precisely half of the pair. (The minimum value is 1/N for N datasets.) Mathematically, representativeness is written ρ¯ and is the reciprocal of β¯. We can also use (as an overlap measure) the diversity index R, which is the average of the representativenesses of the two datasets and which is known as the redundancy of the pair. Note, the original formulation of the framework [2] uses the terms “community” and “metacommunity” in place of “class”/“subset” and “overall”/“dataset,” reflecting its roots in ecology; we prefer the latter in the sentropy package.

1.6. Prior work and present contribution

The need for summary statistics that address the above-described issues has been well documented in metagenomics and immunology [10], [11], [12], [16]. Other mathematical frameworks similar to the one in Eq. 2 have been proposed [17]. Packages for calculating frequency- and similarity-sensitive α and β diversities in effective-number form exist for the R and Julia programming languages [18], [19], but not to our knowledge for Python, which is more widely used, especially in machine learning.

Of note, large datasets present a special challenge for measuring similarity-sensitive diversity: because Eqs. 1–2 require calculating the similarity between each pair of the n unique species in the dataset, and because these pairwise similarities are usually inputted as elements of an n × n matrix, direct implementations risk running out of computer memory when n is large. For example, assuming standard 8-byte floating point precision, a single dataset of a million unique elements would require 8TB of memory. Immunomes, microbiomes, and imaging datasets routinely contain tens or hundreds of thousands of unique elements [20], as do transcriptomes, cell atlases, and other complex datasets. Moreover, a single study can involve tens to hundreds of such datasets (e.g., one per sample/volunteer/patient) and thereby many millions of unique elements in all. The R package rdiversity and the Julia package Diversity both require the similarity matrix to be stored in memory, but for the applications above, such matrices are likely to be too large to store in working memory and may even be too large to write to disk. For this reason, sentropy instead also implements on-the-fly row-by-row calculations, to better handle such cases. The Python package cdiversity calculates Hill-number diversity but without similarity for immune repertoires [22].

In addition to the standard diversity indices in the β family, we introduce two new quantities, ρ^ and β^ (“ρ-hat” and “β-hat”), which are normalized versions of ρ¯ and β¯. The reason is that ρ¯ itself can be a very large number if there are a large number of datasets in the study. So, roughly speaking we can divide ρ¯ by the number of subsets or classes to obtain a measure of redundancy that usually ranges from 0 to 1. (Note the upper bound may not be respected in extreme or pathologic cases, but these will be the exception.)

Furthermore, our package computes similarity-sensitive analogs of the Kullback-Leibler divergence (also known as the relative entropy) as well as of its Rényi generalizations. We define the relative S-entropy between 2 abundance vectors P and Q to be:

DKLZ(P‖Q)=∑i∈supp(P)Pilog(ZP)i(ZQ)i (3)

In other words, starting with the usual definition of the relative entropy, we replace the ith component of P and Q inside the logarithm by the ordinariness vectors. Similarly, the similarity-sensitive analog of Rényi entropies away from viewpoint parameter 1 is:

DqZ(P‖Q)=1q−1log∑i∈supp(P)Pi(ZP)i(ZQ)iq−1 (4)

The two formulas above are in entropy (i.e. traditional) forms. The corresponding D-number form are simply the exponential of the traditional forms.

2. Methods

2.1. Python package

The package was developed and described as in the main text.

2.2. Immunomes

Immunomes were obtained and processed, and similarities calculated, as previously described [10].

2.3. Microbiomes

To construct each synthetic dataset, gaussian sampling to single-decimal-place precision was performed around each of a set of mean values between 0 and 100 to create a gaussian distribution of sequence frequencies around each mean. In Fig. 4, the same number of sequences was sampled around each mean. In Fig. 5, we sampled four times as many sequences around one of the means, which was selected as the dominant taxonomic unit defining the enterotype (a different dominant mean for each enterotype).

Figure 4:

Figure 4:

Metagenomics: Traditional vs. S-entropy alpha measures in synthetic samples. Sample 1 (a) and 2 (b) both have 1,000 sequences. For ease of illustration, distance along the x-axis is proportional to similarity (i.e. more similar sequences are nearer each other). In Sample 1, sequences form five highly distinct clusters. In Sample 2, sequences form 10 more similar clusters. The distribution of sequences within each cluster is Gaussian, consistent with observations from the Human Microbiome Project [11]. Using binning to account for similarity and then measuring diversity using traditional/similarity-insensitive measures (Dq), diversity depends on the binning threshold τ and frequency weighting q. (c) At τ = 97% and q = 0, the two samples are equally diverse, with D0 = 11 species; At higher q, Sample 2 is more diverse, with diversities falling to D∞ = 8.6 vs. D∞ = 5.6 species, respectively, at q = ∞). (d) At τ = 99%, both samples are much more diverse, with D0 = 32 vs. 23 species, respectively. (Note that at this τ, Sample 2 is more diverse than Sample 1 for all q.) (e) In contrast, accounting for similarity using S-entropy measures DqZ, which avoids the need for binning, the order flips: Sample 1 is now more diverse than Sample 2 (for all q), with ∼ 5 vs. ~ 3 − 4 species, respectively, reflecting both the number and the grouping of the clusters. (Here similarity sij between sequences i and j is calculated as sij = e−k∆ij, where ∆ij is the Levenshtein distance between sequences i and j and k = 0.02.)

Figure 5:

Figure 5:

Metagenomics: Beta diversities to elucidate population structure in synthetic samples. (a) Nine synthetic samples created as in Fig. 4a–b, three representing each of three enterotypes: black, Enterotype 1 (Samples 1–1 to 1–3); gray, Enterotype 2 (Samples 2–1 to 2–3); white, Enterotype 3 (Samples 3–1 to 3–3). (b)-(c) Clustering using similarity sij between sequences i and j as defined in Fig. 4 without binning, according to (b) the q = 0 representativeness ρ¯0 of the kth sample for the pair of samples indicated by the heatmap cell and (c) the average representativeness of each member of the pair R¯0.

2.4. Medical imaging

Images were selected at random from the dataset introduced and processed as previously described [23].

2.4.1. Microscopy/Computational pathology

The 89,996-image PathMNIST training set is a component of MedMNIST [24].

2.4.2. Preprocessing

The dataset was filtered to keep only the 82,698 images that contained specimen (removing e.g. nearly-all-black or -blue images, which are likely off the slide or pictures of marker on the slide). Regions surrounding tissue or corresponding to tears in the tissue were thresholded to white. Hue and lightness were obtained as the H and L channels from the standard HLS colorspace using the Python imagine library Pillow (v8.2.0).

Similarity sij between each pair of images i and j was determined as the minimum of the similarities of four component features: hue, lightness, texture (e.g. sheets of nuclei), and structure (e.g. colonic crypts), as detailed below. This expert-feature engineering approach was used as it provided rapid proof of principle based on expert knowledge with a minimal set of human-expert examples (as opposed to a more involved/more data-hungry machine-learning approach).

2.4.3. Hue component

The hue component was the overlap between hue histograms. Lightness component. The lightness component was the overlap between lightness histograms between 0.2 and 1.0, binned into 28 bins.

2.4.4. Texture component

The texture component was a function of the similarity of mottledness (e.g. nuclei, sheets of cells) and stripedness (e.g. muscle fibers). The mottledness mi of an image i was defined as the mean over all pixels of the maximum lightness difference δmaxi between a pixel and each of its eight nearest neighbors (normalized by the range of observed values, which happens to be 0.39): m¯i=δmax/0.39 (where the bar indicates averaging). The mottledness similarity mij between images i and j was then defined as mij=0.1 if mi<0.5 and mj>0.5, and mij=1−mi−mj otherwise. Stripedness of an image i was calculated in four directions: horizontal, vertical, and the two diagonals (σ↑i,σ→i,σ↖i,σ↗i, respectively). In each case, it was defined as one minus the mean over all pixels of the sum of the differences δi between a pixel i and its nearest neighbor in each direction (e.g., for horizontal stripedness, its neighbors to the left and right) times the sum δ⊥i of the neighbors in the perpendicular direction (for horizontal stripedness, its neighbors above and below), again normalized by the range of observed values: σ_i=1−δ¯iδ¯/c⊥ for each direction (where _ is replaced by ↑, →, ↖, or ↗ and the normalization constants were c↑ = 0.229, c→ = 0.240, and c↖ = c↗ = 0.226). The stripedness for the image as a whole, σi, was defined as the max of σ↑i, σ→i, σ↖i, andσ↗i. The stripedness similarity σij was defined similarly to mottledness similarity: σij = 0.1 if σi < 0.5 and σ j > 0.5, and σij = 1 − |σi − σ j| otherwise. The texture similarity was then defined as tij = min(mij, σij).

2.4.5. Structure component

A median filter was applied to each image to ignore outlier pixels. The structure component was then defined as the minimum of two subcomponents: similarity of structure size Sij and similarity of number of “holes” in the image Hij.

Sij was determined as follows. If the resulting image had nearly uniform lightness (lightness range <0.5), Sij was ignored (because meaningful structures could not be easily extracted). For remaining images, structure size was determined by segmenting to find the largest patch Pi around the darkest pixel in each image i. Thereafter, the following conditional logic was applied. If the largest structure in both images in the pair was <30 pixels, Sij was again ignored (structures were likely noise). Otherwise, the difference in lightness of the darkest pixel of each image was calculated (∆). If ∆ was >0.2 times the lightness range of image i and if Pi was <30 pixels, Sij was again ignored; if Pi was ≥ 30 pixels, Sij = 0. If instead ∆ was ≤ 0.2 times the lightness range of image, and if in addition over half the pixels in each image had lightness >0.8 (e.g. indicating fat), Sij = 0.9. Otherwise, if over 60 percent of the pixels in one of the two images had lightness >0.8 (meaning one was very light and the other very dark), Sij = 0.1. Otherwise, Sij was set to be the size ratio of the smaller to the larger patch.

Hij was determined as follows. For each image i, the mask used to find Pi was examined to count the number of light patches. (Patches that were only a single pixel in size were considered noise and ignored.) Each light patch was considered a hole. There was always at least 1 hole. Hij was defined as the ratio of the smaller to the larger number of holes for the pair of images. Structure similarity was defined as min(Sij, Hij).

2.4.6. Comparison to human experts

100 image pairs were selected representing the range of similarities according to the above similarity function (from 0 to 1) to and presented independently to two practicing pathologists. These domain experts were instructed to assign the similarity between the two images, without further instruction, on a 1 to 10 scale. R2 was calculated between their similarity scores as a measure of inter-expert agreement, and separately between each of their scores and Sij.

3. Results

3.1. Python package overview and main features

The sentropy package was developed to simplify the calculation of S-entropy (and traditional entropy) for Python users. It has the ability to calculate all frequency-and similarity-sensitive diversity measures presented in Reeve et al.’s work [15] (S-entropies), as well as their similarity-insensitive counterparts (“traditional” entropies), both in effective-number (preferred) and traditional forms. It offers command-line execution as a Python module for similarity-sensitive diversity calculations using a similarity matrix stored on disk and also provides a convenient interface when imported as a Python package.

3.2. Availability and installation

The sentropy package is available via the Python Package Index (http://pypi.org) and on GitHub at https://github.com/ArnaoutLab/sentropy. It can be installed by running

pip install sentropy

from the command line. The test suite runs successfully on Macintosh, Windows, and Unix systems. The unit tests and coverage report can also be run by

pytest --pyargs sentropy --cov sentropy

3.3. Basic usage: alpha diversities

We illustrate the basic usage of sentropy on simple datasets of fruits (Fig. 1) and animals (Fig. 2) because these are discipline-agnostic and easy to interpret. We will use the terms “subset” or (“class”) and “overall” instead of “subcommunity” and “metacommunity” [7]. (Note, when there is only a single dataset under study, the subset is the same as overall.) We format the input as pandas dataframes for clarity, but this is optional.

First, consider two datasets, each with 35 total elements. Each dataset has the same n = 6 unique elements, each a type of fruit: apples, oranges, bananas, pears, blueberries, and grapes (Fig. 1). Dataset 1a is mostly apples (Fig. 1, left), while Dataset 1b has almost identical frequencies of all the fruits (Fig. 1, right; Table 2).

Table 2:

Datasets 1a&1b: Fruits.

Dataset 1a Dataset 1a Dataset 1b
apple 30 6
orange 1 6
banana 1 6
pear 1 6
blueberry 1 6
grape 1 5
total 35 35

We wish to apply sentropy to these datasets, being sensitive to the essential difference between them: the difference in frequencies of the unique elements. To do so, we first specify a species counts table, formatted as in Table 2. Then we pass it to the sentropy function:

import pandas as pd
import numpy as np
from sentropy import sentropy
MEASURES =["alpha", "rho", "beta", "gamma", "normalized_alpha",\ "normalized_rho", "normalized_beta", "rho_hat", "beta_hat"]
counts_1a = {"dataset_1a": np.array([30, 1, 1, 1, 1, 1])}
counts_1b = {"dataset_1b": np.array([ 6, 6, 6, 6, 6, 5])}
df = sentropy(counts_1a, q=[0,1,np.inf], measure= MEASURES,\level="overall", return_dataframe=True)

Here we requested to get diversity indices for 3 different viewpoint parameters, and all families of measure that sentropy supports. If we print out the output dataframe, we get Table (3).

Table 3:

sentropy output for Dataset 1a.

level viewpoint alpha rho beta gamma normalized_alpha normalized_rho normalized_beta rho_hat beta_hat
0 overall 0.0 6.000000 1.0 1.0 6.000000 6.000000 1.0 1.0 1.0 1.0
1 overall 1.0 1.896549 1.0 1.0 1.896549 1.896549 1.0 1.0 1.0 1.0
2 overall inf 1.166667 1.0 1.0 1.166667 1.166667 1.0 1.0 1.0 1.0

Looking at the alpha column, we see that the value of D1 for this dataset is 1.90. If we do not pass anything to q, then only viewpoint 1 is computed, and if we do not pass anything to measure, then only alpha diversity is computed. Instead of outputting a dataframe, it is also possible to have sentropy output an object that we can query:

D = sentropy(counts_1a, q=[0,1,np.inf], level="both") D(which=’overall’, q=0, measure=’alpha’)

and we get the output 6, because there are 6 species of fruits. We can optionally omit the measure argument in the above, because the sentropy object only computed alpha diversity anyway. By default, level takes value ’overall’, which means only the set-level diversities will be computed. When we query the sentropy object, we can pass to the which argument the name of the subset of interest, or which="overall" if we are interested in diversity at that level. Also, if we had passed a numpy array instead of a pandas dataframe as the abundance matrix, then the names of the subsets are simply their indices. In that case we can pass an integer representing the index of the subset to which.

If furthermore we are interested in the traditional forms of the LCR diversity indices—logarithms (with base e)—instead of effective numbers, we can pass eff_no=False. By default, the eff_no argument (which stands for effective number) takes value True. So, if we run:

H = sentropy(counts_1a, q=0, eff_no=False)
print(H)

we get the answer 1.79. Note that in this case where we want to compute only one viewpoint parameter, one measure and there is one subset in the set, sentropy returns a number rather than an object, because there is only one diversity index to be computed.

If we now repeat the computation for Dataset 1b, we find that D1 ≈ 5.99 for that dataset. The larger value of D1 for Dataset 1b aligns with the intuitive sense that more balance in the frequencies of unique elements means a more diverse dataset.

sentropy can also calculate S-entropy measures for any user-supplied definition of similarity. To illustrate, we now consider a second example in which the elements of two datasets are all unique: here, animals (Fig. 2). Uniqueness means the frequency distributions of the two datasets are identical, so similarity is the only factor that can influence our sense of whether one or the other dataset is more diverse. The investigator always has a choice of similarity measure, which can be choosen to investigate the question at hand; here we consider (approximate) phylogenetic similarity. Dataset 2a consists entirely of birds, so all entries in the similarity matrix are close to 1:

labels_2a = [
"owl", "eagle", "flamingo", "swan",
"duck", "chicken", "turkey", "dodo",
"dove"
]
no_species_2a = len(labels_2a)
S_2a = np.identity(n=no_species_2a)
S_2a[0][1:9] = (0.91, 0.88, 0.88, 0.88, 0.88, 0.88, 0.88, 0.88) # owl
S_2a[1][2:9] = (0.88, 0.89, 0.88, 0.88, 0.88, 0.89, 0.88) # eagle
S_2a[2][3:9] = (0.90, 0.89, 0.88, 0.88, 0.88, 0.89) # flamingo
S_2a[3][4:9] = (0.92, 0.90, 0.89, 0.88, 0.88) # swan
S_2a[4][5:9] = (0.91, 0.89, 0.88, 0.88) # duck
S_2a[5][6:9] = (0.92, 0.88, 0.88) # chicken
S_2a[6][7:9] = (0.89, 0.88) # turkey
S_2a[7][8:9] = (0.88) # dodo
# dove
S_2a = np.maximum( S_2a, S_2a.transpose() )
S_2a = pd.DataFrame({labels_2a[i]: S_2a[i] for i in range(no_species_2a)}, index=labels_2a)

We make a DataFrame of counts in the same way as in the previous example:

counts_2a = pd.DataFrame(
{"dataset_2a": [1, 1, 1, 1, 1, 1, 1, 1, 1]}, index=labels_2a)

To compute similarity-sensitive diversity measures, we now pass the similarity matrix to the similarity argument of sentropy:

result_2a = sentropy(counts_2a, similarity=S_2a, q=0)
print(result_2a)

The output tells us that D1Z≈1.11. The fact that this number is close to 1 reflects the fact that all individuals in this dataset are phylogenetically very similar to each other: they are all birds. In contrast, Dataset 2b (Fig. 2) consists of members from two different phyla: vertebrates and invertebrates. As above, we define a similarity matrix:

labels_2b = ("ladybug", "bee", "butterfly", "lobster", "fish", "turtle", "parrot", "llama", "orangutan")
no_species_2b = len(labels_2b)
S_2b = np.identity(n=no_species_2b)
S_2b[0][1:9] = (0.60, 0.55, 0.45, 0.25, 0.22, 0.23, 0.18, 0.16) # ladybug
S_2b[1][2:9] = (0.60, 0.48, 0.22, 0.23, 0.21, 0.16, 0.14) # bee
S_2b[2][3:9] = (0.42, 0.27, 0.20, 0.22, 0.17, 0.15) # butterfly
S_2b[3][4:9] = (0.28, 0.26, 0.26, 0.20, 0.18) # lobster
S_2b[4][5:9] = (0.75, 0.70, 0.66, 0.63) # fish
S_2b[5][6:9] = (0.85, 0.70, 0.70) # turtle
S_2b[6][7:9] = (0.75, 0.72) # parrot
S_2b[7][8:9] = (0.85) # llama
#orangutan
S_2b = np.maximum( S_2b, S_2b.transpose() )
S_2b = pd.DataFrame({labels_2b[i]: S_2b[i] for i in range(no_species_2b)}, index=labels_2b)

The values of the similarity matrix indicate high similarity among the vertebrates, high similarity among the invertebrates, and low similarity between vertebrates and invertebrates. To calculate the alpha diversity (with q = 1 as above), we proceed as before:

counts_2b = pd.DataFrame({"dataset_2b": [1, 1, 1, 1, 1, 1, 1, 1, 1]}, index=labels_2b)
result_2b = sentropy(counts_2b, similarity=S_2b, q=0)
print(result_2b)

Inspecting the result, we find D1Z≈2.16. That this number is close to 2 reflects the fact that members in this dataset belong to two broad classes of animals: vertebrates and invertebrates. The small difference compared to 2 is interpreted as the contribution of the diversity within each phylum. Note that if we had instead simply placed animals into two bins, vertebrates and invertebrates, this contribution would be lost.

3.4. Basic usage: beta diversities

Recall that beta diversity is between-group diversity. We can use sentropy to compare the diversity of the vertebrates and invertebrates. To do so, we define two classes within Dataset 2b—invertebrates and vertebrates—as follows:

counts_2b_1 = pd.DataFrame(
{"invertebrates": [1, 1, 1, 1, 0, 0, 0, 0, 0], # invertebrates
"vertebrates": [0, 0, 0, 0, 1, 1, 1, 1, 1], # vertebrates
}, index=labels_2b)

We then obtain the representativeness ρ¯ of each subset, here at q = 0, as follows:

result_2b_1 = sentropy(counts_2b_1, similarity=S_2b, q=0, \
measure=[’alpha’, ’normalized_rho’], level="class")
result_2b_1(which = "invertebrates", measure=’normalized_rho’)

and

result_2b_1(which = "vertebrates", measure=’normalized_rho’)

We find the answers 0.63 for the invertebrates and 0.67 for the vertebrates. Recall ρ¯ indicates how well a class represents the overall dataset. Note we can also ask which subset is more diverse by calculating the alpha diversities of the two classes (also at q = 0, for ease of comparison):

result_2b_1(which = "invertebrates", measure=’alpha’)

and

result_2b_1(which = "vertebrates", measure=’alpha’)

We find the answers 3.54 for the invertebrates and 2.30 for the vertebrates. Thus, the invertebrates are more diverse than the vertebrates. In contrast, suppose we split Dataset 2b into two subsets at random, without regard to phylum:

counts_2b_2 = pd.DataFrame(
{
"random_subset_1": [1, 0, 1, 0, 1, 0, 1, 0, 1],
"random_subset_2": [0, 1, 0, 1, 0, 1, 0, 1, 0],
},
index=labels_2b
)

Proceeding again as above, we query the sentropy object to obtain the representativeness of the 2 subsets, with the command

result_2b_2 = sentropy(counts_2b_2, similarity=S_2b, q=0, \
measure=’normalized_rho’, level="class")
result_2b_2(which="random_subset_1")

and the command

result_2b_2(which="random_subset_2")

We find answers 0.93 and 0.92 respectively. These values are higher than the corresponding ones using abundance matrix 2b_1, which reflects the fact that each group now has a roughly equal mix of vertebrates and invertebrates.

3.5. Advanced usage: similarity matrix format

The similarity matrix format—DataFrame, memmap, filepath, or function—should be chosen based on the use case, with particular attention to dataset size. Our recommendations:

  • If the similarity matrix fits in RAM, pass it as a pandas.DataFrame or numpy.ndarray

  • If the similarity matrix does not fit in RAM but does fit on your hard drive (HD), pass it as a cvs/tsv filepath or numpy.memmap. To illustrate passing a csv file, we re-use counts_2b_1 and S_2b from above and save the latter as .csv files (note index=False, since the csv files do not contain row labels).
    S_2b.to_csv("S_2b.csv", index=False)
    
    Then we can pass the path as a string to sentropy:
    sentropy(counts_2b_1, similarity=’S_2b.csv’, chunk_size=5, \ q=0)
    

    We can optionally use the chunk_size argument to specify how many rows of the similarity matrix are read from the file at a time.

  • If the similarity matrix does not fit in either RAM or HD, pass a similarity function to the similarity argument and the feature set that will be used to calculate similarities to the sfargs argument of sentropy. Here’s an example with a set of 2 amino acid sequences:
    from polyleven import levenshtein as edit_distance
    def similarity_function(species_i, species_j):
    return 0.3**edit_distance(species_i, species_j)
    test_seqs = [’CARDYW’, ’CARDYV’]
    test_nos = [10, 1]
    test = pd.DataFrame(
    {"test_nos": test_nos},
    index=test_seqs
    )
    sentropy(
    test,
    similarity=similarity_function,
    sfargs=np.array(test_seqs),
    q=0
    )
    

    and we get the answer 1.22 for the alpha diversity. sentropy admits the optional argument chunk_size, which determines how many rows of the similarity matrix are processed at a time. Note that construction of the similarity matrix is an O(N2) operation; if your similarity function is expensive, this calculation can take time for large datasets. If the similarity callable is a symmetric function (as most similarities are), then a twofold speedup can be obtained by passing symmetric=True to sentropy.

    If we additionally pass parallelize=True, then the computation of diversity indices will be parallelized using the Ray package (all available processing cores will be utilized). In this case we can also pass a number to the argument max_inflight_tasks to specify at most how many parallel tasks should be submitted to Ray at a time (to avoid overwhelming Ray).

3.6. Advanced usage: inter-set ordinariness

The ordinariness of species i is the total abundance of all the species in the dataset that are similar to i, weighted by their similarity and their abundance. (This includes i itself, which is of course 100% similar to itself; Zii = 1.) If the dataset contains many species that are similar to i and/or if those species are very abundant, the ordinariness of species i will be high: species i will be considered unexceptional relative to these many/abundant similar species in the dataset, and in that sense is quite "ordinary" for the dataset.

There are situations in which one is interested in the ordinariness of a species j that is not in the dataset: i.e. how similar the species in the subset are to this outside member, weighted by their relative abundances and their similarity to j. (See [?] for an example, in which the investigators start with an antibody j and wish to calculate its ordinariness in antibody repertoires. In this context, the ordinariness is the repertoire’s “binding capacity” for the molecules that antibody j is best at binding.)

sentropy can calculate ordinariness, whether the species of interest is present in the subset (i above) or not (j):

from sentropy.abundance import Abundance
from sentropy.ray import IntersetSimilarityFromRayFunction
def similarity_function(species_i, species_j):
return 1 / (1 + np.linalg.norm(species_i - species_j))
counts = np.array([[1,0],[0,1],[1,1]])
abundance = Abundance(counts, subsets_names=[’1’, ’2’])
community_species = np.array([[1, 2], [3, 4], [5, 6]])
query_species = np.array([[1, −1], [3, 6]])
similarity = IntersetSimilarityFromRayFunction(
similarity_function,
query_species,
community_species)
ordinariness = similarity @ abundance

The code above computes the ordinariness of each of the 2 query species with respect to a dataset made of 3 other species. Here the similarity function is evaluated on the fly, with Ray parallelization.

3.7. Advanced usage: Kullback-Leibler divergence and its cousins

sentropy can also compute the Kullback-Leibler divergence (i.e. relative entropy) between datasets, as well as its generalizations the Rényi divergences and their S-entropic cousins. To do so, we simply pass 2 abundance arrays as the first 2 arguments of sentropy:

sentropy(counts_2b_1, counts_2b_2, similarity=S_2b, q=1, \ return_dataframe=True, level="both")

we get a tuple of 2 elements, the first of which is a float representing the superset Rényi divergence at viewpoint 1 (in this case, 1), and the second of which is a DataFrame containing Rényi divergences between pairs of subsets, in this case:

3.8. Advanced usage: PyTorch and GPU support

For heavier computations, it is possible to use PyTorch instead of numpy to obtain some acceleration. To do so, we pass backend="torch" to sentropy. (By the fault, the backend argument takes value "numpy".) In order to have the computation of the diversity indices run on the GPU, we can additionally pass device="mps" or device="cuda", depending on whether the computation runs on a Mac computer with Apple silicon, or a computer with NVIDIA CUDA. (By default, the device argument takes value cpu, which means the computation runs on the CPU.) For example, consider the following set with 100 subsets and 10000 entities:

big_counts = np.random.randint(101, size=(10000,100))
n = 10000
indices = np.triu_indices(n, k=1)
values = np.random.rand(len(indices[0])) big_sim_matrix = np.zeros((n, n)) big_sim_matrix[indices] = values
big_sim_matrix = big_sim_matrix + big_sim_matrix.T np.fill_diagonal(big_sim_matrix, 1.0)

On a Mac computer with MPS, we can call:

sentropy(big_counts, similarity=big_sim_matrix, \
q=[0,1, 1.5, np.inf], measure=MEASURES, backend=’torch’, device=’mps’)

which we found to be around 3 times faster than the CPU-bound computation without torch. If we only pass backend=’torch’ without passing device=’mps’, then the computation will utilize torch but will run on the CPU.

3.9. Command-line usage

The sentropy package can also be used from the command line as a module (via python -m). To illustrate this use of sentropy, we re-use again the example with counts_2b_1 and S_2b, now with counts_2b_1 also saved as .csv files (note again index=False):

counts_2b_1.to_csv("counts_2b_1.csv", index=False)

Then from the command line:

python -m sentropy -i counts_2b_1.csv -s S_2b.csv -qs 0 1 inf

The output is a table with all the diversity measures for q = 0, 1, and ∞. Note that while .csv or .tsv are acceptable as input, the output is always tab-delimited. The input filepath (-i) and the similarity matrix filepath (-s) can be URLs to data files hosted on the web. Also note that values of q > 100 are all calculated as q = ∞.

To compute relative entropies in the terminal, we simply pass a second csv file for the other abundance matrix:

python -m sentropy -i counts_2b_1.csv counts_2b_2.csv -s S_2b.csv -qs 1

The full list of flags is -i (for input filepath), -o (for output filepath), -s (for the filepath to the similarity matrix), -qs (for the viewpoint parameters), -ms (for the diversity measures of interest), -chunk_size (for the chunk size when reading the file into memory), -level (for whether to compute diversities at the overall/subset level, or both),-eff_no (for whether to compute effective numbers or entropies), -backend (for whether to use numpy or torch), -device (for whether to use the CPU or the GPU). For further options, consult the help:

python -m sentropy -h

3.10. Applications

We now illustrate some practical applications of sentropy using sample datasets from the fields of immunomics, metagenomics, medical imaging, and computational pathology, as examples of the many types of dataset to which this package can be usefully applied. Jupyter notebooks with these examples are available via the sentropy GitHub repository.

3.10.1. Immunomics

Antibody (IG) and T-cell receptor (TR) repertoires are famously diverse, with tens of thousands of different IG and TR gene sequences in every milliliter of human blood [16], [20], [25], [26], [27]. The frequency of a given sequence rises with exposure to antigens as cells divide, forming clones of identical or similar sequences, and falls to baseline over time as cells die. As a result, measures that are both frequency-sensitive and similarity-sensitive are useful for describing the state of the adaptive immune system [10], [16], [28].

To illustrate, we revisit a dataset of 30 IG heavy-chain (IGH) repertoires from 15 subjects taken before and after influenza vaccination [29], each subsampled to a same-sized 100,000-sequence subset [10]. Vaccination is associated with an increase in the number of sequences—D0, which is similarity insensitive—interpreted as vaccination leading to new sequences [29], a finding confirmed in a recent re-analysis of the data [10] that corrected for sampling bias (using the recon package [16]) (Fig. 3a). We can use sentropy to measure the alpha diversity for each before-and-after pair using the S-entropy measure D0Z, using binding similarity as previously described to construct the similarity matrix [10]. Unlike D0, which generally rises as antigen-responsive cells divide and accumulate mutations, D0Z falls in 12 of the 15 subjects (Fig. 3b). The combination of these two trends tells us that new sequences that are seen following vaccination are likely related, consistent with expansion of pre-existing clones. The three exceptions suggest something different is happening immunologically in these subjects, perhaps new immune responses or fewer mutations.

Figure 3:

Figure 3:

Immunomics: Similarity-insensitive vs. -sensitive measures in influenza vaccination. Diversity of IGH CDR3 immunomes according to the similarity-insensitive measure D0 (left) and its similarity-sensitive counterpart D0Z (middle), and a comparison of the two (right) before vs. after influenza vaccination. Left and middle: light/dotted lines denote subjects where diversity falls. Right: each line/pair of symbols show D0 and D0Z for the same subject. Dark line shows the one subject where vaccination was associated with fewer, more different sequences. Dashes in the margins indicate averages of D0 and D0Z.

3.10.2. Metagenomics

Measuring diversity is also of interest in metagenomics, where microbiomes from e.g. stool or soil samples are high-throughput sequenced to determine the number and frequency of species present. For practical reasons, often the sequenced material is not a complete genomes but telltale genes, such as 16S rRNA sequences, or fragments of such genes [30].

This task is challenging for three reasons. First, the concept of “species” is not as well defined for bacteria as it is for more complex organisms. Second, the similarity of a pair of species, measured in terms of phylogeny or sequence identity, can differ substantially from pair to pair, such that a pair of organisms with a given degree of sequence identity might be classified as the same species in one genus but as distinct species—or even distinct genera—in another. Third, many recovered sequences are difficult to map uniquely to known species. (Tools such as Unifrac that characterize the similarity between microbiomes according to the proportion of shared lineages still require confidence in lineage assignment [31].) These three challenges are in addition to the general challenge faced in measuring the diversity of complex populations: assuring that the diversity of the sample is an accurate representation of the diversity of the population from which the sample was drawn [16].

To address these three challenges, a common first step in measuring microbiome diversity has been to bin sequences into OTUs based on a set threshold level of sequence similarity, e.g. 95% or 97%. Unfortunately, binning is not without its own challenges [11], [12]. First is the problem of “adverse triplets” [12]: sequences i, j, and k for which the similarity between i and j is above the threshold, the similarity between j and k is also above the threshold, but the similarity between i and k is below the threshold. Binning all three into the same OTU because i-j and j-k are above the threshold results in the OTU containing i and k even though their similarity is below the threshold. A second challenge is determining what the appropriate threshold should be [11], [12]. These are general problems with binning or other forms of unique assignment. I.e. they are not specific to metagenomics: they are a necessary consequence of discretization of continuous or near-continuous variables, such as sequence similarity, into discrete bins.

S-entropy offer an alternative way to measure the diversity of microbiomes. We used sentropy to illustrate using synthetic/in silico microbiome samples reminiscent of enterotypes in human gut microbiomes [32] [33]. An enterotype is a microbiome that exhibits a certain compositional pattern, e.g. enrichment in bacteria of the genus Bacteroides [32], [34], [35], [36]. Consider the two samples in Fig. 4. Both have clusters of similar sequences; Sample 1 has five highly distinctive clusters (Fig. 4a) while Sample 2 has 10 clusters that are all fairly similar to each other (Fig. 4b). Similarity can be accounted for in two ways: by binning and then using traditional entropy or by not binning and instead using S-entropy measures.

These yield different results. Binning sequences yields diversity values that depend on the choice of binning threshold but generally show Sample 2 as having higher diversity, reflecting the larger number of clusters (Fig. 4c, d). In contrast, S-entropy shows Sample 1 as having higher diversity, reflecting the larger range that the sequences in Sample 2 span (Fig. 4e). Specifically, D0Z is a little over 5.4 in Sample 1, reflecting the five distinct clusters plus their (limited) intra-cluster diversity, whereas it is 3.6 in Sample 2, reflecting that the 10 clusters are themselves highly related (“effectively” having the same diversity as a sample with 3.6 completely distinct sequences).

The sentropy package is also useful for investigating population structure across different samples (subcommunities), for example to identify correlates of disease [35]. One way to do so is by defining the similarity sij between each pair of unique sequences i and j, treating each pair of samples as its own dataset, and then using sij to calculate how well each sample k represents the pair. This can be done either using the quantity, ρ¯k, which is the representativeness of sample k, or R¯, which is the average representativeness over both samples in the pair. ρ¯k generally differs for each sample in the pair, reflecting asymmetries in the overlap between them; in contrast R¯ is always symmetric, and so should be thought of as an attribute of the pair, not of the constituent samples. Note ρ¯k can be calculated at any q; in this example we omit the q subscript for readability and use q = 0 unless otherwise specified. sentropy readily calculates ρ¯k and R¯ for any q. Once calculated, one can use these values to investigate population structure by clustering on the resulting matrix (Fig. 5).

For example, consider the nine synthetic samples in Fig. 5a. Using the same sij as in the previous example, clustering by either rh¯ok (Fig. 5b) or R¯ (Fig. 5c) clearly shows these samples cluster into three enterotypes. Alternatively, one could instead use the OTU approach: account for similarity by choosing a sequence-similarity threshold τ, bin according to that threshold, and then use sentropy without sij to calculate and cluster on ρ¯k. However, OTU results are highly sensitive to the choice of threshold (Fig. 5d–e). The block structure is discernible only in the dendrogram in Fig. 5d, whereas there is no obvious clustering apparent in Fig. 5e. Note that both approaches require a choice about how similar two sequences are: in the former case via sij and in the latter via τ. (Recall that choosing sij = 1 for sequences whose similarity is ≤ τ and sij = 0 otherwise makes the latter a special case of the former, except without the problem of adverse triplets.) The choice of sij will depend on the context (sequence similarity, phylogenetic similarity, functional similarity, etc.).

3.10.3. Medical Imaging

Medical imaging is an active area of interest in machine learning (ML). Examples include view- and disease-classification tasks on X-rays, magnetic-resonance imaging (MRI) and positron-emission tomography (PET) scans, and ultrasounds (the most widely used and cost-effective modality for internal imaging). Training ML models on these and other types of medical images usually requires large, class-balanced medical-imaging datasets [9]. This section illustrates use of the sentropy package to characterize such datasets, using an echocardiogram dataset as an example [23].

The dataset consists of 136 cardiac ultrasound images (echocardiograms) labeled for five classes according to which of five views of the heart they show (A4C, A5C, 3VV, ABDO, and 3VT), selected for ease of illustration (Fig. 6). Unlike the previously discussed examples, in medical-imaging datasets each element generally appears only once, so all the images occur with the same frequency (1/n, where n is the total number of images in the dataset, i.e. the dataset size). As a result, image frequency is generally not very informative in characterizing the dataset. However, similarity can be, as we show.

Figure 6:

Figure 6:

Medical imaging: Dataset structure. Example images from each labeled class (a) and heatmap of the similarity matrix for all 136 images (b) with several clusters of similar images clearly visible along the diagonal, often but not always correlating with the class label. The bar above the heatmap indicates the view (class label) of each image (ABDO, A4C, A5C, 3VT, 3VV). Similarity here is based on root-mean-square error (RMSE) of pixel differences between images i and j according to sij = e−RMSEij. Identical images have a similarity of 1 and completely distinct images have a similarity approaching 0. Note that while RMSE is not invariant to rotation or translation, ultrasound images (like many medical images) have a privileged orientation and position.

We suppose that only half the images in this dataset are already labeled. We want a sense of what might be gained through the effort of labeling the other half. We know the size of our labeled dataset would double. Assuming that ML benefits from datasets that are not only large but diverse, how much actual diversity would we add? We can answer this question by asking how well each half reflects the composition of the full dataset. The preceding section illustrated the utility of ρ¯k for this purpose. In this section we will use a related measure introduced in the sentropy package: ρ^k (“rho-hat”), defined as ρ^k=ρk−1/(N−1), where N is the number of subsets. ρ^k has the convenient property of ranging from 0 to 1 (except for extreme/pathological cases), which may be more interpretable than the alternatives. A subset k that is maximally distinctive relative to the full dataset will have ρ^k=0 while a subset that is maximally redundant will have ρ^k=1. The dataset-level analog of ρ^k is defined by simply by taking the power mean of the class values of ρ^k, just like for any other diversity indices.

Similarly, we define β^k=βkN−1/(N−1), which typically ranges from 0 to 1 (except in extreme/pathological cases). The overall-dataset analog of β^k is again defined to be the power mean of the class values. In pathological cases where some of the class values are negative, the overall β^ is ill-defined, since the power mean is only defined for positive values. Finally, we note the special case N = 1. In this case, ρ and β are both 1, leading to division by zero, so we define ρ^=β^=1.

To illustrate the utility of ρ^k, consider two different splits of the dataset, both even. Call the subsets that result from the first split A and B and the subsets from the second split C and D. In each case, we assume one of the halves is already labeled. For each split, we calculate ρ^k for each subset relative to the whole dataset, using an RMSE-based similarity (Fig. 6). We find that ρ^A and ρ^B are both ∼ 1: both are as diverse as the overall dataset, even though they are each only half the size (Fig. 7). If either of these subsets is labeled, labeling the other adds very little diversity. In contrast, ρ^C and ρ^B are only 0.6–0.7: labeling the other half would increase diversity substantially.

Figure 7:

Figure 7:

Medical imaging: Representative and non-representative dataset splits. Heatmaps of similarity matrices for subsets of the dataset in Fig. 6. The dataset was split in half into subsets A (a) and B (b), and separately into subsets C (c) and D (d). Although all subsets were identically sized and indistinguishable in terms of class balance (e), ρ^k indicates that A and B are redundant while C and D are complementary, reflecting the similarity in clustering pattern in (a) vs. (b) and the difference in (c) vs. (d).

Currently, the most commonly used heuristics for dataset composition are size and class balance (measured as e.g. the Shannon entropy of the label frequencies). That A, B, C, and D are the same size (n = 68 images) and have essentially the same class balance (1.27–1.28, vs. 1.28 for the full dataset) yet differ markedly in ρ^k demonstrates that this new measure can captures information about datasets that size and class balance alone do not. This ability may be useful in other datasets [37], [38], [39] and as a new way to quantify the value of different forms of data augmentation [40].

3.10.4. Microscopy/Computational Pathology

Finally, we illustrate sentropy on a much larger ML imaging dataset of histological slide images (computational pathology). We use 89,996 training-set images of haematoxylin/eosin (H&E)-stained colon tissue sections from the PathMNIST subset of the MedMNIST database [24]. These are 28 × 28-pixel images of patches of 1,000-fold-magnified microscopic fields from each slide that can contain any of the many structures found in healthy and diseased colon tissue, including fat, muscle, and blood vessels as well as colonic epithelium (Fig. 8a).

Figure 8:

Figure 8:

Computational pathology: Capturing dataset diversity in random subsets. (a) Representative selection of images in the PathMNIST training set. (b) Pairs of high-similarity (left) and low-similarity (right) images according to an expert-defined similarity function (see Methods). For comparison, two human subject-matter experts rated the similarity of the left pair as 0.8 and 0.9, and the right pair 0.0 and 0.0. (c) ρ¯ for random subsets of images from the dataset. The x-axis indicates what percent of the full dataset each subset is; number of images is listed next to each datapoint. Error bars are s.d. for 10 independent random samples. Diagonal is the 1:1 line. Dotted line is a guide to the eye. The inflection point at 128 images indicates diminishing returns.

Continuing the theme of the previous section, we use sentropy to ask how much of the dataset’s diversity is captured by randomly chosen smaller subsets. Specifically, we created subsets of increasing size and measure ρ¯k on each of them, relative to the entire dataset. Because this dataset is much larger than the previous examples, sentropy’s on-the-fly similarity matrix option is used, as opposed to pre-calculating the similarity matrix and then loading it from memory as in the previous sections. This is done by passing sentropy a function for calculating the similarities between pairs of images sij. For this example, the function was designed to approximate human pathologists’ determination of pairwise similarity between images (Fig. 8b), reaching agreement of R2 = 0.48 with each of two experts (on 100 image pairs) vs. an inter-expert R2 of 0.57. We find that even small subsets capture a substantial fraction of the diversity of the full dataset: for example, subsets of 512 randomly chosen images—less than 1% of the full dataset—capture ≥ 95% of the diversity (p = 0.99) (Fig. 8c).

4. Discussion

Frequency- and especially similarity-sensitive diversity (S-entropy) measures are useful for characterizing complex datasets. Here we have introduced and described sentropy, a Python package for calculating them, and demonstrated various applications using examples in immunomics, metagenomics, medical imaging, and computational pathology. sentropy is available via PyPI and installed using the standard package manager (pip) for ease of use.

The recasting of simple counting (richness), Shannon entropy, Simpson’s index, and other measures as special cases of a single framework will be new to some investigators but has two advantages. First, it renders these otherwise disparate measures in the same unit—effective number of unique elements, allowing them to be directly compared. And second, it demonstrates that the only difference among them is in how they weight elements’ frequencies. However, it recasts investigators’ decision of which measure(s) to use as a decision of how much weighting is appropriate for a given investigation: i.e., which q to use. One can avoid this choice by essentially using all of them (e.g. Fig. 4c–e). This option captures much more information and avoids the possibility of inadvertently cherry-picking a measure that might miss some important feature of the dataset at hand. sentropy outputs results for q = 0 to ∞.

In contrast, a choice that cannot be avoided is the choice of similarity measure sij when calculating S-entropy measures. The same set of elements may be similar in different ways: physiologically, genetically, functionally, constructively, etc. This includes settings that can in principle be quite complex, for example when the similarity between a pair of elements is a function of elements’ frequencies, leading to complicated expressions for sij. Because calculation of S-entropy measures DqZ requires sij for all pairs of n unique elements in the dataset, the time required calculating S-entropy measures scales as O(n2). sentropy is implemented with parallelization to make use of multicore processors, allowing datasets of 105-106 elements in under a day on a standard workstation. However, there is opportunity on this rate-limiting step for algorithmic improvement.

The examples presented highlight the complementarity of traditional (similarity-insensitive) and S-entropic (similarity-sensitive) measures (immunomics), the utility of the continuous nature of S-entropy measures (metagenomics), and the value of S-entropy for dataset assembly and assessing dataset quality (machine learning). In the latter context, diversity may prove useful for dataset curation, the need for which has been recognized as a national priority [42]. Given the diversity of these examples, we expect other fields (Table 1) will also benefit from the DqZ-number Hill framework as brought to Python via the sentropy package.

Table 4:

Relative S-entropies between subsets of Dataset 2b_1 and Dataset 2b_2.

subset_2b_3 subset_2b_4
subset_2b_1 1.66 1.55
subset_2b_2 1.43 1.56

Acknowledgements

This work was supported by the Gordon and Betty Moore Foundation, the Food and Drug Administration, and the NIH under grants R01HL150394, R01HL150394-SI, R01AI148747, and R01AI148747-SI.

References

  • [1].Hill Mark O. Diversity and evenness: a unifying notation and its consequences. Ecology, 54(2):427–432, 1973. [Google Scholar]
  • [2].Leinster Tom and Cobbold Christina A.. Measuring diversity: the importance of species similarity. Ecology, 93(3):477–489, March 2012. Number: 3. [DOI] [PubMed] [Google Scholar]
  • [3].Jost Lou. Entropy and diversity. Oikos, 113(2):363–375, 2006. Number: 2 _eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1111/j.2006.0030-1299.14714.x. [Google Scholar]
  • [4].Berger W. H. and Parker F. L.. Diversity of planktonic foraminifera in deep-sea sediments. Science (New York, N.Y.), 168(3937):1345–1347, June 1970. Number: 3937. [DOI] [PubMed] [Google Scholar]
  • [5].Jost Lou. PARTITIONING DIVERSITY INTO INDEPENDENT ALPHA AND BETA COMPONENTS. Ecology, 88(10):2427–2439, October 2007. [DOI] [PubMed] [Google Scholar]
  • [6].Jost Lou. What do we mean bydiversity? The path towardsquantification. Mètode Revista de difusió de la investigació, (9), July 2018. [Google Scholar]
  • [7].Leinster Tom. Entropy and Diversity: The Axiomatic Approach. arXiv preprint arXiv:2012.02113, 2020. [Google Scholar]
  • [8].Wang Zhou, Bovik Alan C, Sheikh Hamid R, and Simoncelli Eero P. Image quality assessment: from error visibility to structural similarity. IEEE transactions on image processing, 13(4):600–612, 2004. Publisher: IEEE. [DOI] [PubMed] [Google Scholar]
  • [9].Chinn Erin, Arora Rohit, Arnaout Ramy, and Arnaout Rima. ENRICHing medical imaging training sets enables more efficient machine learning. Journal of the American Medical Informatics Association, 30(6):1079–1090, June 2023. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [10].Arora Rohit and Arnaout Ramy. Repertoire-scale measures of antigen binding. Proceedings of the National Academy of Sciences, 119(34):e2203505119, August 2022. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [11].Nguyen Nam-Phuong, Warnow Tandy, Pop Mihai, and White Bryan. A perspective on 16S rRNA operational taxonomic unit clustering using sequence similarity. npj Biofilms and Microbiomes , 2(1):16004, November 2016. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [12].Edgar Robert C. Updating the 97% identity threshold for 16S ribosomal RNA OTUs. Bioinformatics, 34(14):2371–2375, July 2018. [DOI] [PubMed] [Google Scholar]
  • [13].Kaplinsky J., Li A., Sun A., Coffre M., Koralov S. B., and Arnaout R.. Antibody repertoire deep sequencing reveals antigen-independent selection in maturing B cells. Proc Natl Acad Sci U S A, 111(25):E2622–9, June 2014. Number: 25. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [14].Couch Josiah, Arora Rohit, Braun Jasper, Kaplinsky Joesph, Hill Elliot, Li Anthony, Altschul Brett, and Arnaout Ramy. Scaling Monte-Carlo-Based Inference on Antibody and TCR Repertoires, December 2023. arXiv:2312.12525 [q-bio]. [Google Scholar]
  • [15].Reeve Richard, Leinster Tom, Cobbold Christina A., Thompson Jill, Brummitt Neil, Mitchell Sonia N., and Matthews Louise. How to partition diversity. arXiv:1404.6520 [q-bio], April 2014. arXiv: 1404.6520. [Google Scholar]
  • [16].Kaplinsky Joseph and Arnaout Ramy. Robust estimates of overall immune-repertoire diversity from high-throughput measurements on samples. Nature Communications, 7(1):11881, September 2016. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [17].Chiu Chun-Huo and Chao Anne. Distance-Based Functional Diversity Measures and Their Decomposition: A Framework Based on Hill Numbers. PLOS ONE, 9(7):e100014, July 2014. Number: 7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [18].Reeve Richard and Harris Claire. Diversity.jl - diversity measurement in Julia, 2019. [Google Scholar]
  • [19].Mitchell Sonia, Reeve Richard, and White Tom. rdiversity - Measurement and Partitioning of Similarity-Sensitive Biodiversity, 2022. [Google Scholar]
  • [20].Arnaout Ramy A., Luning Prak Eline T., Schwab Nicholas, Rubelt Florian, and the Adaptive Immune Receptor Repertoire Community. The Future of Blood Testing Is the Immunome. Frontiers in Immunology, 12:626793, March 2021. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [21].Pelissier Aurelien. cdiversity: Quantifying B-Cell Clonal Diversity In Repertoire Data. [Google Scholar]
  • [22].Pelissier Aurelien, Luo Siyuan, Stratigopoulou Maria, Guikema Jeroen E. J., and Martínez María Rodríguez. Exploring the impact of clonal definition on B-cell diversity: implications for the analysis of immune repertoires. Frontiers in Immunology, 14:1123968, 2023. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [23].Arnaout Rima, Curran Lara, Zhao Yili, Levine Jami C., Chinn Erin, and Moon-Grady Anita J.. An ensemble of neural networks provides expert-level prenatal detection of complex congenital heart disease. Nature Medicine, 27(5):882–891, May 2021. Number: 5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [24].Yang Jiancheng, Shi Rui, Wei Donglai, Liu Zequan, Zhao Lin, Ke Bilian, Pfister Hanspeter, and Ni Bingbing. MedMNIST v2 - A large-scale lightweight benchmark for 2D and 3D biomedical image classification. Scientific Data, 10(1):41, January 2023. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [25].Arnaout R., Lee W., Cahill P., Honan T., Sparrow T., Weiand M., Nusbaum C., Rajewsky K., and Koralov S. B.. High-resolution description of antibody heavy-chain repertoires in humans. PLoS One, 6(8):e22365, 2011. Number: 8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [26].Briney Bryan, Inderbitzin Anne, Joyce Collin, and Burton Dennis R.. Commonality despite exceptional diversity in the baseline human antibody repertoire. Nature, 566(7744):393–397, February 2019. Number: 7744 Publisher: Nature Publishing Group. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [27].Soto Cinque, Bombardi Robin G., Branchizio Andre, Kose Nurgun, Matta Pranathi, Sevy Alexander M., Sinkovits Robert S., Gilchuk Pavlo, Finn Jessica A., and Crowe James E.. High frequency of shared clonotypes in human B cell receptor repertoires. Nature, 566(7744):398, February 2019. Number: 7744. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [28].Chiffelle Johanna, Genolet Raphael, Marta AS Perez George Coukos, Zoete Vincent, and Harari Alexandre. T-cell repertoire analysis and metrics of diversity and clonality. Current Opinion in Biotechnology, 65:284–295, October 2020. [DOI] [PubMed] [Google Scholar]
  • [29].Vollmers Christopher, Sit Rene V., Weinstein Joshua A., Dekker Cornelia L., and Quake Stephen R.. Genetic measurement of memory B-cell recall using antibody repertoire sequencing. Proceedings of the National Academy of Sciences of the United States of America, 110(33):13463–13468, August 2013. Number: 33. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [30].Johnson Jethro S., Spakowicz Daniel J., Hong Bo-Young, Petersen Lauren M., Demkowicz Patrick, Chen Lei, Leopold Shana R., Hanson Blake M., Agresta Hanako O., Gerstein Mark, Sodergren Erica, and Weinstock George M.. Evaluation of 16S rRNA gene sequencing for species and strain-level microbiome analysis. Nature Communications, 10(1):5029, December 2019. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [31].Lozupone Catherine and Knight Rob. UniFrac: a New Phylogenetic Method for Comparing Microbial Communities. Applied and Environmental Microbiology, 71(12):8228–8235, December 2005. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [32].Arumugam Manimozhiyan, Raes Jeroen, Pelletier Eric, Denis Le Paslier Takuji Yamada, Mende Daniel R., Fernandes Gabriel R., Tap Julien, Bruls Thomas, Batto Jean-Michel, Bertalan Marcelo, Borruel Natalia, Casellas Francesc, Fernandez Leyden, Gautier Laurent, Hansen Torben, Hattori Masahira, Hayashi Tetsuya, Kleerebezem Michiel, Kurokawa Ken, Leclerc Marion, Levenez Florence, Manichanh Chaysavanh, Bjørn Nielsen H., Nielsen Trine, Pons Nicolas, Poulain Julie, Qin Junjie, Sicheritz-Ponten Thomas, Tims Sebastian, Torrents David, Ugarte Edgardo, Zoetendal Erwin G., Wang Jun, Guarner Francisco, Pedersen Oluf, de Vos Willem M., Brunak Søren, Doré Joel, Weissenbach Jean, Dusko Ehrlich S., and Bork Peer. Enterotypes of the human gut microbiome. Nature, 473(7346):174–180, May 2011. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [33].Lozupone Catherine A., Stombaugh Jesse I., Gordon Jeffrey I., Jansson Janet K., and Knight Rob. Diversity, stability and resilience of the human gut microbiota. Nature, 489(7415):220–230, September 2012. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [34].Petersen Lauren M, Bautista Eddy J, Nguyen Hoan, Hanson Blake M, Chen Lei, Lek Sai H, Sodergren Erica, and Weinstock George M. Community characteristics of the gut microbiomes of competitive cyclists. Microbiome, 5(1):1–13, 2017. Publisher: BioMed Central. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [35].Cheng Mingyue and Ning Kang. Stereotypes About Enterotype: the Old and New Ideas. Genomics, Proteomics & Bioinformatics, 17(1):4–12, February 2019. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [36].Zhou Xin, Wang Baohong, Demkowicz Patrick C., Johnson Jethro S., Chen Yanfei, Spakowicz Daniel J., Zhou Yanjiao, Dorsett Yair, Chen Lei, Sodergren Erica, Kuchel George A., and Weinstock George M.. Exploratory studies of oral and fecal microbiome in healthy human aging. Frontiers in aging, 3:1002405, 2022. Place: Switzerland. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [37].Kornblith Aaron E., Addo Newton, Dong Ruolei, Rogers Robert, Jacqueline Grupp-Phelan Atul Butte, Gupta Pavan, Callcut Rachael A., and Arnaout Rima. Development and Validation of a Deep Learning Strategy for Automated View Classification of Pediatric Focused Assessment With Sonography for Trauma. Journal of Ultrasound in Medicine: Oflcial Journal of the American Institute of Ultrasound in Medicine, 41(8):1915–1924, August 2022. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [38].Athalye C., van Nisselrooij A., Rizvi S., Haak M., Moon-Grady A. J., and Arnaout R.. Deep learning model for prenatal congenital heart disease (CHD) screening can be applied to retrospective imaging from the community setting, outperforming initial clinical detection in a well-annotated cohort. Ultrasound in obstetrics & gynecology : the oflcial journal of the International Society of Ultrasound in Obstetrics and Gynecology, September 2023. Place: England. [Google Scholar]
  • [39].Danielle Lopes Ferreira Zaynaf Salaymang, and Arnaout Rima. Label-free segmentation from cardiac ultrasound using self-supervised learning. ArXiv, abs/2210.04979, 2022. [Google Scholar]
  • [40].Athalye Chinmayee and Arnaout Rima. Domain-guided data augmentation for deep learning on medical imaging. PloS One, 18(3):e0282532, 2023. [DOI] [PMC free article] [PubMed] [Google Scholar]

Articles from ArXiv are provided here courtesy of arXiv

RESOURCES