Skip to main content
BMC Genomics logoLink to BMC Genomics
. 2026 May 6;27:567. doi: 10.1186/s12864-026-12680-4

BandHiC: a memory-efficient and user-friendly Python package for organizing and analyzing Hi-C matrices down to sub-kilobase resolution

Weibing Wang 1, Junping Li 1, Yusen Ye 1,✉, Lin Gao 1,✉
PMCID: PMC13317261  PMID: 42092759

Abstract

Background

Recent advances in high-resolution Hi-C and Micro-C technologies have enabled finer-scale characterization of 3D genome architecture. However, these improvements also introduce substantial computational challenges, as the memory requirements of Hi-C/Micro-C contact matrices scale quadratically with resolution, leading to prohibitive resource consumption.

Results

To address this, we developed BandHiC, a memory-efficient and user-friendly Python package for organizing and analyzing Hi-C matrices down to sub-kilobase resolution. BandHiC adopts a banded storage strategy that preserves only a configurable diagonal bandwidth of the dense contact matrix, reducing memory usage by up to 99% while maintaining fast random access and intuitive indexing operations. In addition, it provides flexible masking mechanisms to handle missing values, outliers, and unmappable regions, and supports efficient vectorized operations optimized with NumPy, thereby enabling scalable analysis of ultra-high-resolution Hi-C datasets.

Conclusions

BandHiC provides a memory-efficient and scalable framework that enables sub-kilobase-resolution Hi-C matrix analysis on standard hardware. Its seamless integration with the NumPy ecosystem and user-friendly design make it a practical and accessible foundation for future advances in 3D genomics. The source code of the BandHiC Python package is publicly available on GitHub (https://github.com/xdwwb/BandHiC-Master), and comprehensive documentation is provided at its website (https://xdwwb.github.io/BandHiC-Master/). Installation can be performed conveniently through Python’s pip package manager.

Supplementary Information

The online version contains supplementary material available at 10.1186/s12864-026-12680-4.

Keywords: 3D genome, Hi-C, Software, Python package, Data structure

Background

High-throughput chromosome conformation capture (Hi-C) [1–4] and its variants, such as Micro-C [5–7], have substantially advanced our understanding of genome architecture by enabling genome-wide mapping of chromatin interactions at progressively higher resolutions. Hi-C data typically measures the interaction frequency between evenly spaced chromatin segments, represented mathematically as a contact matrix, where each bin corresponds to a genomic interval and the bin size defines the resolution of the contact map. Over the past decade, improvements in sequencing throughput and experimental protocols have enabled a dramatic increase in Hi-C data resolution, from early megabase (Mb)-scale maps to kilobase (kb)- and even sub-kilobase (e.g., 500 bp, 250 bp) contact matrices [4, 6]. More recently, single-cell Micro-C has achieved contact maps at resolutions as high as 5 kb [8]. These advances in resolution have facilitated the discovery of finer-scale chromatin structures and their regulatory roles in gene expression [3, 6, 7, 9, 10].

However, higher-resolution data generated by these assays pose significant computational challenges, particularly in terms of memory consumption [11, 12]. For instance, loading a dense Hi-C matrix into Random Access Memory (RAM) at 1 kb resolution for the human genome (~ 3 billion base pairs) would require approximately Inline graphic of memory (Inline graphic 66 tebibytes), assuming double-precision floating-point representation (8 bytes per entry). Even when restricting the analysis to the longest human chromosome (chromosome 1, ~ 249 Mb), a dense matrix at 1 kb resolution would still require nearly 462 GiB of memory. In practical workflows, intermediate data structures generated during normalization and downstream analyses often lead to a significant increase in peak memory usage beyond the raw matrix itself. Such memory demands far exceed the capacities of most computational environments. This issue becomes even more pronounced at sub-kilobase or single-cell resolutions, rendering dense matrix representations impractical for many real-world applications. These limitations highlight the need for more memory-efficient data structures tailored for high-resolution Hi-C data.

Many current methods for identifying genome structural patterns, such as chromatin loops and topologically associating domains (TADs), rely heavily on dense matrix representations to facilitate rapid, random data access during computation [13]. Tools, such as TopDom [14], MSTD [15], DeTOKI [16] and SnapHiC [17], exhibit high memory consumption associated with dense matrices, which limits their scalability to higher-resolution Hi-C datasets. Although conventional sparse matrix formats can reduce memory usage, they lack efficient random-access capability, significantly slowing downstream analyses and complicating algorithm design, particularly for tasks requiring frequent element-wise access or submatrix extraction. Other methods, including Mustache [18] and Chromosight [19], attempt to circumvent this limitation by partitioning the full dense matrix into smaller blocks that are loaded into memory on demand. However, these methods introduce added implementation complexity and require careful memory management to maintain computational efficiency.

However, these methods typically leverage high-resolution Hi-C data predominantly within short-range genomic distances (typically within 10 Mb), as chromatin loops and TADs are generally constrained to such scales [3]. For example, TopDom detects TADs by computing an insulation-like signal within a near-diagonal sliding window, which targets domain-length scales on the order of a few hundred kilobases [14]. HiCCUPS reports that 98% of detected loops occur between loci < 2 Mb apart; accordingly, its CPU implementation adopts an 8 Mb default diagonal bandwidth as a practical compromise to reduce computation time while covering most candidate loop distances [3]. Consistently, Arrowhead identifies contact domains spanning 40 kb–3 Mb (median 185 kb) [3]. Other widely used approaches also operate under bounded distances: Fit-Hi-C focuses on interactions at an intermediate scale of ~ 50 kb–10 Mb [20], Peakachu (a machine-learning–based loop caller) uses a default upper genomic distance cutoff of a few megabases [21], and chromatin loop or Hi-C map prediction methods often restrict candidate pairs to ~ 1–2 Mb [22–25]. Together, these practices motivate focusing computation on band-limited regions for high-resolution cis analyses. However, efficient data-structure support for such workflows at sub-kilobase resolution remains limited, motivating the design of BandHiC.

Interactions at distances beyond a few megabases become increasingly sparse at high resolutions and are often irrelevant to local structures such as loops and TADs. Thus, focusing on short-range contacts is both computationally and biologically justified. Therefore, there is an urgent need for a novel memory scheme tailored specifically for local-range higher-resolution Hi-C data, one that dramatically reduces memory usage while retaining the fast random-access capabilities of dense matrices.

Built upon NumPy [26, 27], a fundamental library for numerical computing in Python, we developed BandHiC, a memory-efficient Python package specifically designed for organizing and analyzing short-range contacts of Hi-C data down to sub-kilobase resolutions. BandHiC adopts a banded matrix storage scheme that stores only a configurable diagonal bandwidth of the full Hi-C contact matrix. A banded matrix is a standard concept in numerical analysis, referring to matrices whose nonzero entries are confined near the main diagonal [28], and scientific computing libraries such as SciPy provide linear system solvers that operate on banded storage formats [29]. BandHiC preserves efficient random-access capabilities by employing a direct index mapping between the banded and dense matrix representations, while supporting familiar NumPy-style indexing semantics (slicing, Boolean array indexing, integer array indexing) to facilitate user-friendly and efficient data access. BandHiC also integrates masking functionality akin to NumPy’s MaskedArray module, enabling straightforward handling of gaps, outliers, and other aberrant values in Hi-C matrices. Finally, BandHiC supports diverse numerical operations optimized through NumPy’s efficient vectorized computations, thus offering both memory efficiency and high computational performance essential for practical high-resolution Hi-C data analysis.

Implementation

Data representation

To address the increasing memory demands posed by high-resolution Hi-C data, we introduce band_hic_matrix, the core class implemented in the BandHiC package. Given a Hi-C contact matrix Inline graphic at resolution Inline graphic, band_hic_matrix retains only the diagonals within a user-defined bandwidth Inline graphic, yielding a compact representation Inline graphic (Fig. 1). This format ensures that each column in Inline graphic corresponds to a fixed diagonal of Inline graphic, such that the mappingInline graphic.

Fig. 1.

Fig. 1

Data model of BandHiC. Schematic illustration of converting a dense symmetric matrix A into a banded representation consisting of a data matrix D, an element-wise mask matrix M, a row/column mask matrix X, and a default value d for out-of-band entries. Diagonal elements from A are reorganized into columns of D; M marks missing or outlier entries; X indicates masked rows or columns

The memory efficiency achieved by this strategy is substantial. When Inline graphic, the memory footprint of band_hic_matrix is reduced from Inline graphic to Inline graphic. For example, assuming a resolution of 1 kb and a bandwidth of 2 Mb (Inline graphic), the representation of chromosome 1 of the human genome (~ 249 Mb) requires 3.7 GiB of memory, less than 1% of the memory required by the dense matrix (~ 462 GiB). This compression makes high-resolution Hi-C data accessible even on commodity hardware, without compromising the efficiency of random data access.

To further enhance the flexibility of usage, band_hic_matrix supports an optional two-layer masking mechanism. An element-wise mask matrix Inline graphic allows users to selectively ignore missing or outlier contacts, enabling robust statistical estimation on unmasked subsets. Additionally, a bin-level mask Inline graphic supports the exclusion of entire rows or columns, particularly useful for removing repetitive genomic regions lacking valid Hi-C signals. These masking features facilitate downstream tasks such as estimation of average contact intensity at specific genomic distances, while preserving statistical validity.

Lastly, a scalar default value Inline graphic is defined to fill in the undefined entries of Inline graphic not covered by the banded matrix Inline graphic. This default is typically set to 0, consistent with the assumption that long-range interactions are negligibly sparse. The advantage of using default values is that, in addition to treating the band_hic_matrix as a full dense matrix for indexing and conversion with a dense matrix, it also allows out-of-band entries to participate in mathematical operations. For example, when adding 1 to the band_hic_matrix object, not only are the in-band entries incremented, but the default value representing the out-of-band entries also increases by 1. Together, the components Inline graphic, Inline graphic, Inline graphic, and Inline graphic allow for seamless reconstruction of the dense matrix Inline graphic when required. Overall, band_hic_matrix provides an efficient, flexible data representation for scalable Hi-C data analysis.

BandHiC package

BandHiC is distributed as an open-source Python package under the MIT license. It is compatible with Python version 3.8 or higher and can be deployed on Linux and macOS platforms. The BandHiC package relies primarily on NumPy and SciPy, which provide the computational backbone for the banded matrix data structure and its core operations. In addition, BandHiC wraps the file-reading functions of cooler and hic-straw, allowing it to directly read .hic, .cool, or .mcool files as inputs for creating a band_hic_matrix object (Fig. 2A).

Fig. 2.

Fig. 2

Overview of the BandHiC package. A Example of a band_hic_matrix object in the BandHiC package. B Indexing methods supported by BandHiC. C Example of computation methods supported by BandHiC

BandHiC primarily defines a matrix class, band_hic_matrix, which represents Hi-C contact data in a banded matrix representation. Each instance contains a numerical NumPy array of shape (bin_num, diag_num), together with Boolean arrays mask and mask_row_col that record element-wise and row/column-wise exclusions. Regarding the choice of diag_num, based on common Hi-C analysis workflows [3, 14], we recommend setting it to a diagonal bandwidth corresponding to ~ 2–8 Mb, depending on the resolution. In practice, 2 Mb is typically sufficient for TAD calling, whereas 2–8 Mb is recommended for loop calling. Users can also flexibly adjust this parameter according to their specific tasks and objectives.

Although the mask array has the same shape as the data array, it consumes significantly less memory. Specifically, the data array is typically stored using double-precision floating-point or int64 values (8 bytes per entry), whereas the mask array only requires a Boolean representation (1 byte per entry). As a result, the memory footprint of the mask array is only one eighth that of the data array. In contrast, the mask_row_col array is a one-dimensional vector, whose memory consumption is negligible compared with that of the data array. Elements outside the stored bandwidth are represented by a scalar default_value, allowing the matrix to behave as a dense symmetric array while avoiding redundant storage. Objects can be constructed from .hic, .mcool files, or from triplet-form contact records (rows, columns, and contact frequencies) (Fig. 2A).

The package fully supports NumPy-style indexing (Fig. 2B), including item-wise, slice, Boolean array, and integer array indexing. It also provides a series of methods and functions for constructing, manipulating, and performing computations on the band_hic_matrix objects (Fig. 2C). The design and implementation of the indexing and computational operations supported by BandHiC are described in detail in the following two subsections.

Indexing operations

Building on the direct coordinate mapping between the banded representation D and the full dense matrix A described above (Fig. 1), band_hic_matrix provides constant-time random access to all in-band entries while preserving dense-matrix semantics. In practice, users can access elements through Inline graphic, and interact with the object as if it were a standard NumPy array, without being exposed to the underlying storage scheme.

Leveraging this property, band_hic_matrix supports full NumPy-style indexing operations, including slicing, Boolean array, and integer array indexing (Fig. 2B). Slicing selects contiguous ranges of data, such as Inline graphic. Boolean array indexing extracts elements that meet specific conditions, for example Inline graphic. Integer array indexing allows arbitrary selection using arrays of integer indices, such as Inline graphic. This design allows users to easily query local chromatin contacts and provides a flexible and efficient framework for data manipulation in scientific computing. For instance, a slice operation such as Inline graphic or Inline graphic retrieves a banded submatrix. Combined with the todense operation, this enables reconstruction of the dense submatrix for downstream analysis or visualization.

For a band_hic_matrix object with the mask array, indexing operations return either a ma.MaskedArray object or the masked constant. In NumPy, a MaskedArray stores numerical data together with a Boolean mask that marks missing or invalid entries. In practice, indexing a band_hic_matrix behaves as if operating directly on a MaskedArray, thereby providing users with considerable flexibility.

Indexing operations that fall outside the predefined diagonal bandwidth do not raise errors; instead, such entries are filled with a user-specified default value (e.g., zero), further improving robustness and usability in practical applications.

Numerical computation

In addition to flexible data access, band_hic_matrix also supports a wide range of numerical operations, including element-wise mathematical operations and reduction operations (Fig. 2C). BandHiC is built on top of NumPy, which provides the foundation for efficient numerical computation in Python. NumPy offers high-performance, C-optimized universal functions (ufuncs) that perform element-wise operations with support for broadcasting and type casting. By leveraging these ufuncs, BandHiC implements 71 element-wise mathematical operations directly on the band_hic_matrix (Table 1). This design eliminates the need for explicit Python loops, ensures full compatibility with the NumPy ecosystem, and significantly improves computational speed. Consequently, BandHiC inherits the scalability and efficiency of NumPy, enabling fast and memory-efficient analysis of large genomic contact matrices. Moreover, NumPy provides interfaces for defining custom array-like objects while maintaining seamless integration with NumPy, enabling BandHiC to implement specialized matrix types efficiently and flexibly. By building on NumPy in this way, BandHiC inherits both its computational efficiency and its flexible programming model, making it well-suited for scalable analysis of large-scale Hi-C data.

Table 1.

Universal functions that BandHiC supports

Function Description Function Description
absolute Absolute value add Element-wise addition
arccos Inverse cosine arccosh Inverse hyperbolic cosine
arcsin Inverse sine arcsinh Inverse hyperbolic sine
arctan Inverse tangent arctan2 Arctangent of y/x with quadrant
arctanh Inverse hyperbolic tangent bitwise_and Element-wise bitwise AND
bitwise_or Element-wise bitwise OR bitwise_xor Element-wise bitwise XOR
cbrt Cube root conj Complex conjugate
conjugate Alias for conj cos Cosine function
cosh Hyperbolic cosine deg2rad Degrees to radians
degrees Radians to degrees divide Element-wise division
divmod Quotient and remainder equal Element-wise equality test
exp Exponential exp2 Base-2 exponential
expm1 exp(x) − 1 fabs Absolute value (float)
float_power Floating-point power floor_divide Integer division (floor)
fmod Modulo operation gcd Greatest common divisor
greater Element-wise greater-than test greater_equal Greater-than or equal test
heaviside Heaviside step function hypot Euclidean norm
invert Bitwise inversion lcm Least common multiple
left_shift Bitwise left shift less Element-wise less-than test
less_equal Less-than or equal test log Natural logarithm
log1p log(1 + x) log2 Base-2 logarithm
log10 Base-10 logarithm logaddexp log(exp(x) + exp(y))
logaddexp2 Base-2 version of logaddexp logical_and Element-wise logical AND
logical_or Element-wise logical OR logical_xor Element-wise logical XOR
maximum Element-wise maximum minimum Element-wise minimum
mod Remainder (modulo) multiply Element-wise multiplication
negative Element-wise negation not_equal Element-wise inequality test
positive Returns input unchanged power Raise to power
rad2deg Radians to degrees radians Degrees to radians
reciprocal Element-wise reciprocal remainder Modulo remainder
right_shift Bitwise right shift rint Round to the nearest integer
sign Sign of input sin Sine function
sinh Hyperbolic sine sqrt Square root
square Square of input subtract Element-wise subtraction
tan Tangent function tanh Hyperbolic tangent
true_divide Division that returns a float

Reduction operations refer to functions that aggregate multiple values into a single result. They can be applied globally to all elements of a matrix, or along specific axes to summarize rows or columns. Examples include sum, min, max, and mean. BandHiC supports ten such reduction operations (Table 2), which work along conventional axes (rows or columns) in the same way as NumPy. In addition, BandHiC extends these operations to the diagonal axis—a feature not available in NumPy. This diagonal reduction is particularly useful for Hi-C data, as it allows interaction frequencies to be summarized by genomic distance, thereby supporting distance-dependent normalization and analyses such as distance-decay profiling. All operations remain fully compatible with masked band_hic_matrix objects, ensuring robust handling of missing or low-quality data in large-scale Hi-C analysis.

Table 2.

Reduction functions that BandHiC supports

Function Description
sum Compute the sum of all elements along the specified axis
prod Compute the product of all elements along the specified axis
min Return the minimum value along the specified axis
max Return the maximum value along the specified axis
mean Compute the arithmetic mean along the specified axis
var Compute the variance (average squared deviation)
std Compute the standard deviation (square root of variance)
ptp Compute the range (max - min) of values along the axis
all Return True if all elements evaluate to True
any Return True if any element evaluates to True

The implementation of reduction operations in BandHiC is designed to behave equivalently to those on a dense matrix, but without explicitly constructing the dense matrix, which would otherwise consume substantial memory. During computation, BandHiC automatically fills out-of-band entries with the default value, symmetrizes interactions in the lower-triangular part of the matrix, and excludes entries masked by either element-wise or row/column masks.

Taken together, band_hic_matrix combines the memory efficiency of a banded storage model with the expressiveness of NumPy’s interface. By mimicking both Numpy’s ndarray and MaskedArray behaviors, it provides an intuitive and powerful interface for users, substantially lowering the barrier to adoption and enabling seamless integration into existing Hi-C data analysis pipelines. Please refer to BandHiC’s website for a detailed list of all supported functions and tutorials.

Results

Performance benchmarking

To verify correctness, we compared BandHiC outputs with NumPy arrays using an automated test suite covering core numerical operations as well as indexing and masking behaviors. Across all tests, BandHiC produces results identical to NumPy or consistent within numerical precision, ensuring user-level output equivalence. All test scripts are publicly available in the project’s GitHub repository.

To evaluate the efficiency of BandHiC for random, element-wise access to Hi-C contact matrices, we compared its memory footprint and read/write throughput with three commonly used representations: dense array, CSR (compressed sparse row), and COO (coordinate format), across resolutions from 50 kb to 500 bp on mouse embryonic stem cell (mESC) Micro-C data from chromosome 1 (Fig. 3A, B). For a fair comparison, both CSR and COO were restricted to store only interactions within the same diagonal bandwidth as BandHiC, and all random-access benchmarks were limited to entries inside this band (see Supplementary Material 2 for detailed experimental setup). We benchmarked element-wise random access on in-memory Hi-C matrices, reflecting the fine-grained local access pattern commonly used in Hi-C local interaction analysis (e.g., loop and TAD detection). Under other access patterns or operations (e.g., row-wise iteration or matrix operations), CSR and COO may show different relative performance.

Fig. 3.

Fig. 3

Evaluation of BandHiC. A Memory usage comparison of different matrix representations, including banded, dense, COO (coordinate format), and CSR (compressed sparse row), on mouse embryonic stem cell (mESC) Micro-C data (chromosome 1) at various resolutions. B Read and write throughput comparison between banded, CSR, and dense matrix representations on mESC Micro-C data (chromosome 1). Throughput is measured as the number of matrix elements accessed per second (elements/s). COO is not included because it does not support efficient access to individual matrix entries required for this benchmark. C Memory usage comparison between the banded and dense matrices when running TopDom on mESC Micro-C data (chromosome 1). D Runtime comparison for the same task. At 1,000 bp and 500 bp resolution, the dense representations failed due to memory overflow

Dense matrices serve as an ideal baseline for access performance, as their contiguous memory layout enables highly efficient traversal. When feasible, dense representations indeed achieve the highest read and write throughput, except at 5 kb resolution, where their write throughput is surpassed by the banded representation. However, their memory footprint grows quadratically with resolution, making them impractical at high resolutions: in our experiments, dense matrices could not be instantiated at 1 kb and 500 bp due to memory exhaustion. In contrast, across the tested resolutions, BandHiC achieves up to ~ 99% reduction in memory usage relative to dense representations, while retaining approximately 47–108% of the dense read/write throughput (Fig. 3A, B), providing a scalable alternative for high-resolution analyses.

The CSR representation substantially reduces memory usage by storing only nonzero entries together with row pointers, but this comes at the cost of access efficiency. Because retrieving an individual element requires pointer indirection and index searches within each row, CSR exhibits consistently lower element-wise throughput. Across all tested resolutions, BandHiC achieves approximately 1.3–13.5× higher read throughput and 4.2–16.6× higher write throughput than CSR under the same band-limited setting. For example, at 1 kb resolution, BandHiC shows approximately 13× higher read throughput and 17× higher write throughput than CSR. These results indicate that CSR performance degrades sharply under highly sparse conditions, whereas BandHiC maintains stable and high access rates.

In terms of memory usage, both BandHiC and CSR dramatically reduce storage requirements relative to dense matrices, but their relative efficiency depends on the density of interactions within the band. At coarse resolutions (e.g., 50 kb and 25 kb), where the banded region is relatively dense, BandHiC can be even more memory-efficient than CSR because it stores only the banded values without any pointer or index overhead. For example, at 50 kb resolution BandHiC requires ~ 1.19 MiB, slightly less than CSR (~ 1.27 MiB), and at 25 kb both formats use comparable memory (~ 4.77 vs. ~4.79 MiB). As the resolution increases and the banded region becomes sparser, CSR becomes more memory-efficient than BandHiC (e.g., ~ 73.4 MiB vs. ~119.2 MiB at 5 kb), reflecting the growing advantage of skipping zero entries in sparse storage.

The COO representation offers an even simpler sparse storage scheme by explicitly listing the coordinates and values of nonzero entries, resulting in low memory overhead and flexibility for constructing sparse matrices. However, it lacks an efficient interface for direct single-element access, as retrieving an entry would require scanning coordinate lists. Consequently, COO is not included in the throughput benchmark, as it is not compatible with the element-wise access pattern evaluated here. This reflects a fundamental trade-off: while COO is advantageous for matrix assembly and storage, it is ill-suited for random-access workloads in downstream Hi-C analysis.

We repeated the same benchmarking on GM12878 Hi-C data [3], and obtained consistent results that further confirm the advantages of BandHiC (Supplementary Fig. S1A, B). Overall, these results demonstrate that BandHiC provides a favorable compromise between dense and sparse representations. It dramatically improves scalability over dense matrices by reducing memory requirements by up to two orders of magnitude, while preserving a large fraction of dense access performance, and it substantially outperforms CSR in element-wise throughput under the same storage constraints. By combining compact banded storage with direct coordinate mapping and Inline graphic access semantics, BandHiC enables efficient random access to high-resolution Hi-C data at scales where existing formats become impractical.

Application: reducing TopDom memory consumption

While the previous section demonstrates the core functionality and syntax of BandHiC, here we evaluate its practical utility by integrating it with the TAD-calling algorithm TopDom [14]. TopDom is an efficient and robust algorithm for detecting topologically associating domains (TADs). It uses a sliding window approach on Hi-C contact matrices to identify local minima in contact intensity, which mark potential TAD boundaries. Owing to its robustness and reproducibility, TopDom was rated as the best-performing TAD detection method in a benchmark study [9]; however, its reliance on dense matrices makes it difficult to apply to sub-kilobase resolution Hi-C data. The original TopDom algorithm was developed in the R language [14]. In BandHiC, we provide a Python implementation of TopDom as a built-in function, ensuring functional consistency with the original method while extending support for both NumPy’s ndarray and BandHiC’s band_hic_matrix class. This integration not only facilitates direct identification of TADs from high-resolution Hi-C data within the BandHiC but also enables a fair comparison of memory usage and runtime performance between dense and banded representations.

To evaluate the effectiveness of BandHiC, we benchmarked the TopDom algorithm on chromosome 1 of mouse embryonic stem cell (mESC) Micro-C data across multiple resolutions using both dense matrix and band_hic_matrix, with the band_hic_matrix bandwidth set to 2 Mb (see Supplementary Material 2 for detailed experimental setup). As shown in Fig. 3C, BandHiC substantially reduces memory usage at all tested resolutions. At 1000 bp and 500 bp resolution, dense matrix representations could not be constructed within the available system memory (60,268 MiB RAM and 12,288 MiB swap space), and the program terminated with memory allocation errors. In contrast, the band_hic_matrix version of TopDom completes successfully, with memory usage of only 5,989 MiB and 23,902 MiB at 1000 bp and 500 bp resolution, respectively.

The banded matrix introduces a modest increase in runtime compared to the dense matrix (Fig. 3D). This overhead arises not from masking operations (which were disabled for this evaluation) but from index computation performed during element access within the band_hic_matrix object. Overall, in our benchmarks, BandHiC uses approximately 1% of the memory of dense matrices, while incurring approximately a 50% increase in runtime compared with the dense implementation. Because TopDom’s default sliding window (~ 200 kb) is fully contained within the 2 Mb bandwidth used in BandHiC, the BandHiC-based and dense implementations yield identical TAD calls; therefore, we do not further assess TopDom’s detection accuracy here and instead focus on benchmarking memory usage and runtime efficiency. We repeated the same benchmarking on GM12878 Hi-C data, and obtained consistent results that further confirm the advantages of BandHiC (Supplementary Fig. S1C, D).

These results indicate that band_hic_matrix enables TAD-calling algorithms like TopDom to process high-resolution Hi-C data on a standard personal computer, making large-scale analysis feasible without incurring high computational cost. Owing to BandHiC’s NumPy-like API, other Hi-C-based pattern identification methods can be reimplemented with minimal modifications to the original code, thereby significantly improving their scalability and adaptability to higher-resolution Hi-C datasets.

Usage examples

BandHiC can serve as an alternative to the NumPy package when managing and manipulating Hi-C matrices, aiming to address the issue of excessive memory usage caused by storing dense matrices with NumPy’s ndarray. At the same time, BandHiC supports masking operations like NumPy’s ma.MaskedArray module, with enhancements tailored for Hi-C data. Users can leverage their experience with NumPy when using the BandHiC package, so it is recommended that users have some basic knowledge of NumPy. Here are some code examples providing a quick guide and demonstration of the core functionalities of BandHiC:

graphic file with name 12864_2026_12680_Figa_HTML.jpg

The defined band_hic_matrix object mat corresponds to the example shown in Fig. 2. The example is presented in the IPython-style interactive format, in which “In [n]:” indicates input commands and “Out[n]:” indicates the corresponding outputs.

Conclusions

BandHiC alleviates the computational challenges of high-resolution Hi-C analysis through a banded storage scheme that reduces memory usage to ~ 1% of dense matrices while preserving constant-time random access. This enables domain-calling algorithms, such as TopDom, to run at sub-kilobase resolution on standard hardware. Seamless integration with the NumPy ecosystem, including support for universal functions, reductions, and Hi-C–specific diagonal operations, facilitates efficient distance-dependent analyses and easy adoption in existing pipelines. Furthermore, BandHiC provides a flexible programming model—supporting masking, user-defined default fill values, and robust handling of noisy or unmappable regions—which facilitates downstream analyses of local chromatin features such as loop and TAD detection.

BandHiC is designed for local interaction analysis in high-resolution Hi-C data, where most informative contacts fall within bounded genomic distances. By restricting storage to a user-defined interaction band, it achieves substantial memory reduction while preserving fast fine-grained random access and dense-matrix–like semantics, making it well suited for local-pattern analyses such as loop and TAD detection.

This band-limited design also defines its applicability boundaries. BandHiC is not intended for analyses that require full-matrix global interaction structures, such as A/B compartment, nor for normalization procedures such as KR normalization, which rely on whole-matrix constraints and extensive row-wise operations, where conventional sparse formats (e.g., CSR) are generally more appropriate. Overall, analyses of ultra long-range or inter-chromosomal interactions are inherently limited under the banded design, and future work may address this through hybrid banded–sparse storage strategies.

While BandHiC does not provide dedicated single-cell–specific analysis methods, single-cell Hi-C data can already be handled by loading each cell as an independent band_hic_matrix, with all supported operations identical to those used for bulk Hi-C. No specialized single-cell modeling or benchmarking is included in the current work. Looking forward, an explicit cell dimension extending the banded storage from two to three dimensions could offer more convenient and scalable support for large single-cell datasets and represents a possible direction for future development.

Taken together, BandHiC represents both a practical solution to the pressing computational challenges posed by ultra-high-resolution Hi-C data and a flexible foundation for future methodological advances. By combining memory efficiency, computational scalability, and a user-friendly interface, BandHiC has the potential to become a core component of the bioinformatics toolkit for 3D genomics.

Availability and requirements

Project name

BandHiC.

Project home page: https://github.com/xdwwb/BandHiC-Master.

Operating system(s): Linux, macOS, and Windows (requires WSL).

Programming language: Python

Other requirements: Python 3.8 or higher

License: MIT

Any restrictions to use by non-academics: None

Supplementary Information

Supplementary Material 1. (730.4KB, docx)
Supplementary Material 2. (20.6KB, docx)

Acknowledgements

We thank all the members of Prof. Gao’s lab for helpful comments.

Abbreviations

Hi-C

High-throughput Chromosome Conformation Capture

Kb

Kilobase (1,000 base pairs)

Mb

Megabase (1,000,000 base pairs)

RAM

Random Access Memory

GiB

Gibibyte

TAD

Topologically Associating Domain

Authors’ contributions

W.W. conceived the software design and implemented the BandHiC package. W.W. and J.L. tested it across different operating systems. W.W. also wrote the manuscript and developed the project homepage and tutorials. L.G., J.L., and Y.Y. contributed to editing the manuscript. All authors reviewed and approved the final version of the manuscript.

Funding

This work was supported by the National Natural Science Foundation of China [Grant Nos. 62502361 to W.W.; 62550005, 62132015 and U22A2037 to L.G.; 62573335 to Y.Y.].

Data availability

All the datasets used for this work are publicly available. The mESC Micro-C and GM12878 Hi-C data was downloaded from the Gene Expression Omnibus database (GEO; https://www.ncbi.nlm.nih.gov/geo/) under accession number GSE130275 and GSE63525.

Declarations

Ethics approval and consent to participate

Not applicable.

Consent for publication

Not applicable.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Contributor Information

Yusen Ye, Email: ysye@xidian.edu.cn.

Lin Gao, Email: lgao@mail.xidian.edu.cn.

References

  • 1.Lieberman-Aiden E, van Berkum NL, Williams L, Imakaev M, Ragoczy T, Telling A, et al. Comprehensive mapping of long-range interactions reveals folding principles of the human genome. Science. 2009;326:289–93. 10.1126/science.1181369. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Dixon JR, Selvaraj S, Yue F, Kim A, Li Y, Shen Y, et al. Topological domains in mammalian genomes identified by analysis of chromatin interactions. Nature. 2012;485:376–80. 10.1038/nature11082. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Rao SSP, Huntley MH, Durand NC, Stamenova EK, Bochkov ID, Robinson JT, et al. A 3D Map of the Human Genome at Kilobase Resolution Reveals Principles of Chromatin Looping. Cell. 2014;159:1665–80. 10.1016/j.cell.2014.11.021. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Akgol Oksuz B, Yang L, Abraham S, Venev SV, Krietenstein N, Parsi KM, et al. Systematic evaluation of chromosome conformation capture assays. Nat Methods. 2021;18:1046–55. 10.1038/s41592-021-01248-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Hsieh T-HS, Weiner A, Lajoie B, Dekker J, Friedman N, Rando OJ. Mapping Nucleosome Resolution Chromosome Folding in Yeast by Micro-C. Cell. 2015;162:108–19. 10.1016/j.cell.2015.05.048. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Krietenstein N, Abraham S, Venev SV, Abdennur N, Gibcus J, Hsieh T-HS, et al. Ultrastructural Details of Mammalian Chromosome Architecture. Mol Cell. 2020;78:554–e5657. 10.1016/j.molcel.2020.03.003. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Sun L, Zhou J, Xu X, Liu Y, Ma N, Liu Y, et al. Mapping nucleosome-resolution chromatin organization and enhancer-promoter loops in plants using Micro-C-XL. Nat Commun. 2024;15:35. 10.1038/s41467-023-44347-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Wu H, Zhang J, Tan L, Xie XS. Single-cell Micro-C profiles 3D genome structures at high resolution and characterizes multi-enhancer hubs. Nat Genet. 2025;57:1777–86. 10.1038/s41588-025-02247-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Zufferey M, Tavernari D, Oricchio E, Ciriello G. Comparison of computational methods for the identification of topologically associating domains. Genome Biol. 2018;19:217. 10.1186/s13059-018-1596-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Goel VY, Huseyin MK, Hansen AS. Region Capture Micro-C reveals coalescence of enhancers and promoters into nested microcompartments. Nat Genet. 2023;55:1048–56. 10.1038/s41588-023-01391-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Dekker J, Belmont AS, Guttman M, Leshyk VO, Lis JT, Lomvardas S, et al. The 4D nucleome project. Nature. 2017;549:219–26. 10.1038/nature23884. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Raffo A, Paulsen J. The shape of chromatin: insights from computational recognition of geometric patterns in Hi-C data. Brief Bioinform. 2023;24:bbad302. 10.1093/bib/bbad302. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Xu J, Xu X, Huang D, Luo Y, Lin L, Bai X, et al. A comprehensive benchmarking with interpretation and operational guidance for the hierarchy of topologically associating domains. Nat Commun. 2024;15:4376. 10.1038/s41467-024-48593-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Shin H, Shi Y, Dai C, Tjong H, Gong K, Alber F, et al. TopDom: an efficient and deterministic method for identifying topological domains in genomes. Nucleic Acids Res. 2016;44:e70. 10.1093/nar/gkv1505. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Ye Y, Gao L, Zhang S. MSTD: an efficient method for detecting multi-scale topological domains from symmetric and asymmetric 3D genomic maps. Nucleic Acids Res. 2019;47:e65–65. 10.1093/nar/gkz201. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Li X, Zeng G, Li A, Zhang Z. DeTOKI identifies and characterizes the dynamics of chromatin TAD-like domains in a single cell. Genome Biol. 2021;22:217. 10.1186/s13059-021-02435-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Yu M, Abnousi A, Zhang Y, Li G, Lee L, Chen Z, et al. SnapHiC: a computational pipeline to identify chromatin loops from single-cell Hi-C data. Nat Methods. 2021;18:1056–9. 10.1038/s41592-021-01231-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Roayaei Ardakany A, Gezer HT, Lonardi S, Ay F. Mustache: multi-scale detection of chromatin loops from Hi-C and Micro-C maps using scale-space representation. Genome Biol. 2020;21:256. 10.1186/s13059-020-02167-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Matthey-Doret C, Baudry L, Breuer A, Montagne R, Guiglielmoni N, Scolari V, et al. Computer vision for pattern detection in chromosome contact maps. Nat Commun. 2020;11:5795. 10.1038/s41467-020-19562-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Ay F, Bailey TL, Noble WS. Statistical confidence estimation for Hi-C data reveals regulatory chromatin contacts. Genome Res. 2014;24:999–1011. 10.1101/gr.160374.113. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Salameh TJ, Wang X, Song F, Zhang B, Wright SM, Khunsriraksakul C, et al. A supervised learning framework for chromatin loop detection in genome-wide contact maps. Nat Commun. 2020;11:3428. 10.1038/s41467-020-17239-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Whalen S, Truty RM, Pollard KS. Enhancer–promoter interactions are encoded by complex genomic signatures on looping chromatin. Nat Genet. 2016;48:488–96. 10.1038/ng.3539. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Wang W, Gao L, Ye Y, Gao Y. CCIP: predicting CTCF-mediated chromatin loops with transitivity. Bioinformatics. 2021;37:4635–42. 10.1093/bioinformatics/btab534. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Wang W, Ye Y, Gao L. Statistical modeling and significance estimation of multi-way chromatin contacts with HyperloopFinder. Brief Bioinform. 2024;25:bbae341. 10.1093/bib/bbae341. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Schwessinger R, Gosden M, Downes D, Brown RC, Oudelaar AM, Telenius J, et al. DeepC: predicting 3D genome folding using megabase-scale transfer learning. Nat Methods. 2020;17:1118–24. 10.1038/s41592-020-0960-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Van Der Walt S, Colbert SC, Varoquaux G. The NumPy Array: A Structure for Efficient Numerical Computation. Comput Sci Eng. 2011;13:22–30. 10.1109/MCSE.2011.37. [Google Scholar]
  • 27.Harris CR, Millman KJ, van der Walt SJ, Gommers R, Virtanen P, Cournapeau D, et al. Array programming with NumPy. Nature. 2020;585:357–62. 10.1038/s41586-020-2649-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Suli E, Mayers DF. An introduction to numerical analysis. Cambridge University Press; 2003.
  • 29.Virtanen P, Gommers R, Oliphant TE, Haberland M, Reddy T, Cournapeau D, et al. SciPy 1.0: fundamental algorithms for scientific computing in Python. Nat Methods. 2020;17:261–72. 10.1038/s41592-019-0686-2. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Material 1. (730.4KB, docx)
Supplementary Material 2. (20.6KB, docx)

Data Availability Statement

All the datasets used for this work are publicly available. The mESC Micro-C and GM12878 Hi-C data was downloaded from the Gene Expression Omnibus database (GEO; https://www.ncbi.nlm.nih.gov/geo/) under accession number GSE130275 and GSE63525.


Articles from BMC Genomics are provided here courtesy of BMC

RESOURCES