Abstract
Multiple genetic association studies linked variants at ARID5B with predisposition to B-cell–derived acute lymphoblastic leukemia (B-ALL) in children. Still, the molecular function of ARID5B remains largely uncharacterized. Here, we employ a combination of proteomics, genomics, and transcriptomics to describe the molecular mechanisms of ARID5B. We identify that ARID5B interacts with MIER1, C16ORF87, HDAC1, and HDAC2 forming a chromatin repressor complex. By CUT&RUN, we mapped ARID5B binding in active regions of the genome, tethering HDAC1 and HDAC2 to distal regulatory elements and promoters. Genes actively repressed by the ARID5B repressor complex are involved in B-cell proliferation and B-cell–specific signaling. Together, we describe how ARID5B assembles into a repressor complex and regulates B-cell–specific processes. Understanding its molecular mechanism will help elucidating how noncoding germline variants at ARID5B predispose to B-ALL.
Graphical Abstract
Graphical Abstract.
Introduction
Germline noncoding variants are known risk factors of pediatric B-cell acute lymphoblastic leukemia (B-ALL) [1–13]. Importantly, some risk variants are in eQTL and are associated with lineage-specific transcription factors, such as IKZF1 and CEBPE [1, 2], suggesting that dysregulation of lineage-specific transcriptional programs contributes to the predisposition to B-ALL. Among the strongest associations are risk variants localized in intronic regions of ARID5B [1, 2, 14] (AT-Rich Interaction Domain 5B), a DNA-binding protein whose function is not clearly understood. Furthermore, individuals carrying the risk allele or a heterozygous deletion of ARID5B oftentimes express lower levels of the gene [14, 15], highlighting the importance of ARID5B in the onset of leukemia.
Recently, inactivation of ARID5B in the mouse was described to affect B-cell development promoting an accumulation of pre-B cells and decrease in the immature B-cell population in the bone marrow [16]. In line with these findings, ARID5B overexpression in mice leads to a decrease of the pre-B cell (Hardy fraction D) population in the bone marrow [17]. ARID5B was further shown to promote tonic B-cell receptor (BCR) signaling, to limit cell proliferation and regulate fatty acid uptake and metabolism in B cells [16, 17].
ARID5B encodes two isoforms (short and long), with both isoforms having an ARID (AT-Rich Interaction Domain), responsible for DNA binding, and only the long isoform containing a BAH (Bromo Adjacent Homology) domain [18]. Disruption of the DNA binding of ARID5B has been shown to lead to the de-repression of transcriptional regulators of adipocyte thermogenesis [19]. In turn, the BAH domain of ARID5B is predicted to mediate protein–protein or protein–nucleic acid interactions on chromatin [18]. In fact, ARID5B has been suggested to interact with the histone demethylase PHF2, regulating the expression of metabolic genes in the liver [20] and genes involved in chondrogenesis [21].
Thus, while experiments in mice [16, 17] highlighted the implication of ARID5B in coordinating key cellular processes during B-cell development, especially at the pre-B-cell stage, the molecular mechanisms by which it regulates these processes is still unknown. In this study, we used proteomics to map the ARID5B protein network and identified transcriptional repressors HDAC1, HDAC2, MIER1, and C16ORF87 as main interactors. By targeting ARID5B protein domains, we mapped the BAH domain as the main interface mediating the ARID5B–HDAC interaction. Upon inactivation of ARID5B, we observed a significant loss of HDAC1 and HDAC2 on chromatin and concomitant gain in H3K27ac, suggesting that ARID5B is required to recruit HDAC1 and HDAC2 to chromatin. Gain of H3K27ac in these regions upon HDAC inhibition further confirms that these loci are directly repressed by HDAC1 and HDAC2. Finally, transcriptomics revealed that ARID5B and HDAC regulate genes involved in B cell proliferation and B cell-specific signaling.
Materials and methods
Cell culture
NALM6 (CVCL_0092), HEK293T (CVCL_0063), and HAP1 (CVCL_Y019) were cultured in RPMI 1640 (Gibco, Life Technologies), DMEM (Gibco, Life Technologies), and IMDM (Gibco, Life Technologies), respectively, all supplemented with 10% fetal bovine serum, 100 U/ml penicillin–streptomycin and 2 mM L-glutamine at 37°C, 5% CO2. Cells were tested weekly for mycoplasma contamination (MycoAlert, Lonza LT07-318).
Generation of NALM6 KO cells
For the KO of ARID5B, Cas9-stably expressing NALM6 cells were electroporated with RNP complexes containing sgRNA targeting exons 5 and 10 of the long isoform of ARID5B (Synthego, Supplementary Table S1—Oligonucleotides). Cells were directly seeded in 96-well plates at a density of 5 cells/well to obtain clones. KO clones were identified by genotyping followed by Sanger sequencing (see Supplementary Table S1—Oligonucleotides).
Genotyping of KO cells
For the validation of KO cells, the region targeted by the sgRNA was amplified by PCR (see Supplementary Table S1—Oligonucleotides), and editing was assessed by sanger sequencing. Briefly, cells were lysed in Pawell’s buffer [10 mM Tris–HCl (pH 9.0), 50 mM KCl, 0.45% NP-40, and 0.45% Tween-20) with 0.8 mg/ml Proteinase K (Qiagen, #19133) at 56°C for 15 min followed by 95°C for 5 min. OneTaq DNA Polymerase (NEB, #M0486) was used for polymerase chain reaction (PCR) with 1.5 µl of DNA as input.
Proliferation assay
A total of 500 000 cells/ml were seeded in 2 ml of RPMI in each well of a six-well plate in biological duplicates. Cells were counted daily for 7 days using the CASY cell counter (OMNI Life Science). Cell density was normalized to the day of seeding.
Treatment of NALM6 WT with HDAC inhibitors
Inhibition of HDAC Class I was achieved upon treating NALM6 Cas9 or ARID5B KO cells at a density of 1 million cells/ml with 500 nM Entinostat – MS-275 (Selleckchem, #S1053) or 0.005% Dimethyl sulfoxide (DMSO) for 24 h prior to RNA harvesting or CUT&RUN experiment.
Stimulation of NALM6 WT and ARID5B KO cells with Igµ
NALM6 WT and ARID5B KO cells were seeded at a density of 2 million cells/ml and stimulated with 10 µg/ml Goat F(ab')2 Anti-Human IgM-UNLB (Southern Biotech, #2022-01) for 10 min prior to protein harvesting.
Protein harvesting
Cells were lysed using NP-40 buffer [50 mM Tris–HCl (pH 8.0), 150 mM NaCl, and 1% NP-40] supplemented with 1:100 Protease Inhibitor Cocktail (Sigma, #P8340) and 0.2 mM phenylmethanesulfonyl fluoride (Sigma, #93482). Lysates were incubated for 15 min at 4°C and centrifuged for 10 min at top speed at 4°C. The supernatant was transferred to a fresh tube and samples were prepared for western blot.
Nuclear protein isolation
A total of 2 million cells were washed once with 100 µl of PBS prior to lysis with 100 µl NE buffer [20 mM HEPES (pH 7.9), 10 mM KCl, 0.1% Triton X-100, 20% glycerol, 1 mM MnCl2, 0.5 mM Spermidine, 1× cOmplete EDTA-free protease inhibitor cocktail (Roche, #11836170001)] and incubation at 4°C for 10 min. Nuclei were pelleted at 400 × g for 3 min and resuspended in 100 µl of Buffer A [20 mM Tris–HCl (pH 7.5), 100 mM NaCl, 5 mM MgCl2, 1 mM NaF, 10% glycerol, 0.2% NP-40, 20 mM β-glycerophosphatase, 0.5 mM Dithiothreitol (DTT), 0.2 mM phenylmethanesulfonyl fluoride (Sigma, #93482), 1:100 Protease Inhibitor Cocktail (Sigma, #P8340), 1:200 Benzonase (Millipore Sigma, #706643), 2.5 µg/ml RNAse A (Thermo Fisher, #12091021)]. Following incubation at 10°C for 30 min with agitation, the nuclear protein extract was centrifuged at 400 × g for 3 min and the supernatant transferred to a fresh tube. Nuclear extracts were used to assess HDAC1, HDAC2, H3K4me1, and H3K27ac levels.
Lentivirus production
HEK293T cells were seeded in 15 cm plates the day prior to transfection. When 80% confluent, cells were transfected with 8.75 µg VSV-G, 16.25 µg, pPAX2, and 25 µg of the lentivector of interest (see Supplementary Table S1—Plasmids) using 150 µg branched PEI (Sigma, #408727). The following day, the media of the cells was changed with fresh media. Lentiviral supernatant was collected 2 and 3 days after transfection, centrifuged at 2500 rpm for 5 min, filtered using a 0.45 µm filter (Thermo Scientific, #723-2545), and concentrated by ultracentrifugation at 24 000 rpm, at 4°C for 2 h, using 20% sucrose in PBS. For crude lentivirus preparation, transfected HEK293T cells were cultured in either DMEM, IMDM or RPMI, the lentivirus supernatant was collected 2 and 3 days post-transfected, centrifuged at 2500 rpm for 5 min, and filtered using a 0.45 µm filter (Thermo Scientific, #723-2545).
Lentivirus transduction
Cells were seeded at a density of 1 million cells/ml and 10 µl of purified lentivirus or an equal volume of crude lentivirus was added. The day after transduction, 10 µg/ml of Blasticidin (Gibco, #A1113903) or 1 µg/ml of Puromycin (Thermo Fisher, #A1113803) were used to select cells transduced with viruses carrying a Blasticidin or Puromycin resistance cassette, respectively. Cells transduced with constructs containing a GFP cassette were FACS sorted for GFP 3–5 days post-transduction.
Affinity purification of FLAG-Avi-tagged ARID5B followed by mass spectrometry (AP-MS)
The murine Arid5b ORF with 5′ N-terminus FLAG-AviTag was synthesized (IDT) and cloned into a lentivector under the EF1a promoter using Gibson Assembly Master Mix (NEB, #E2611S) (see Supplementary Table S1—Plasmids). HEK293T cells constitutively expressing the Escherichia coli Biotin ligase BirA were transduced with the ARID5B lentivector and selected with Puromycin for 5 days. One-step affinity purification with streptavidin beads was performed as previously reported [22, 23]. In brief, nuclear extracts were obtained from approximately 300 million cells using NE-PER Nuclear and Cytoplasmic Extraction Reagents (Thermo Fisher Scientific, #78833) and quantified using the Pierce BCA Protein Assay Kit (Thermo Fisher Scientific, #23225). Five milligrams of nuclear extract from HEK293T cells either expressing only BirA (control) or BirA plus biotin-tagged ARID5B were diluted in 10 ml of IP350 buffer [20 mM Tris–HCl (pH 7.5), 0.3% NP-40, 1 mM EDTA, 10% glycerol, 350 mM NaCl, 1 mM DTT, 0.2 phenylmethanesulfonyl fluoride (Sigma, #93 482)] supplemented with a Protease Inhibitor Cocktail (Sigma–Aldrich, #P8340-5, 1:100). Hundred microliters of Dynabeads MyOne streptavidin T1 beads (Thermo Fisher Scientific, #65601) were added and incubated with rotation overnight at 4°C. Beads were washed four times with 10 ml of IP350 buffer and eluted in 250 µl of Buffer II [50 mM HEPES (pH 8.0), 150 mM NaCl, 5 mM EDTA, and 2% SDS]. Next, eluted samples were diluted to equal amounts in lysis buffer [2% SDS, 50 mM HEPES, pH 8, supplemented with 1 mM phenylmethanesulfonyl fluoride (Sigma, #93482) and protease inhibitors] up to 100 µl. Next DTT (final concentration of 10 mM DTT) was added, and samples were incubated for 1 h at 56°C. After reduction, cysteine residues were treated with IAA (final concentration of approximately 50 mM) and incubated in the dark for 30 min at room temperature. To each sample, 8 µl of washed and equilibrated magnetic SP3 beads (SpeedBeads, GE Healthcare) were added, and proteins were precipitated by the addition of acetonitrile (ACN; final concentration of 70% v/v). The samples were incubated for 18 min before the beads were immobilized for 2 min and the supernatant was removed. Beads were rinsed twice with 70% (v/v) ethanol, followed by a wash in ACN. After the final washing step, 100 µl of 50 mM ammonium bicarbonate (NH4HCO3) was added to the beads, and proteins were digested by the addition of 1.5 µg of sequencing grade trypsin. Samples were incubated on the shaker with low agitation speed overnight (∼ 12 h) at 37°C. Next, samples were acidified to pH 2–3 with 30% TFA and cleaned up by STAGE-tip C18.
AP-MS data acquisition
MS-data were acquired on an Orbitrap Fusion Lumos Tribrid mass spectrometer coupled to a Dionex Ultimate 3000 RSLCnano system via a Nanospray Flex Ion Source interface. Peptides were loaded onto the trap column (PepMap 100 C18, 5 µm, 5 × 0.3 mm) and subsequently eluted onto the analytical column (50 cm), 75 µm inner diameter analytical column inhouse packed with ReproSil-Pur 120 C18-AQ, 3 µm, with an ESI emitter (20 µm ID × 7 cm L × 365 µm OD; Orifice ID: 10 µm, CoAnn Technologies) kept at 50°C. A 120 min analytical gradient with 230 nl/min flow rate was used, using as buffer A 0.4% FA in H2O, and as buffer B 0.4% FA in ACN. The gradient had the following scheme: scheme: 0–4 min 4% buffer B, 4–86 min to 24% B, 86–94 to 36% B, 94–95 min increase to 100% B, and 101–102 back to 4% B. The MS was operated in positive mode using a data-dependent acquisition (DDA) mode with a 3 s cycle time. The MS1 scan range was set from 375 to 1650 m/z, and the dynamic exclusion was set to 60 s with a mass tolerance of 10 ppm. MS1 spectra were recorded in the Orbitrap at a resolution of 120 000, with a maximum injection time of 80 ms and an AGC target value of 2e5. Only peptide precursor charge stages of 2–6 were selected for fragmentation. MS2 fragment spectra were recorded at resolution of 15 000 with an automatic gain control target intensity of 5e4. Fragmentation was achieved with higher energy collision energy (HCD) at 30%. The maximum injection time was limited to 100 ms. The isolation width was set to 1.6 m/z and fragments were recorded from 120 to 2000 m/z.
MS data processing
MS raw files were processed using FragPipe (version 22.0) [24, 25] and MSFragger (version 4.1) [26] with the label-free quantification with match-between-runs (LFQ-MBR) workflow. The peptide identification search was conducted against the human protein database downloaded from UniProtKB (consensus sequences, only reviewed, 2 June 2025), which included decoy sequences (reversed protein sequences) and common lab contaminants (e.g. rubber and bovine serum albumin). The default search engine settings of the LFQ-MBR workflow were used for the FragPipe analysis. Briefly, a closed search with the setting “strict trypsin” specificity, N-terminal methionine clipping, and up to two missed cleavages was performed. Up to three variable modifications were allowed, including oxidation on methionine, N-terminal acetylation, and carbamidomethylation of cysteine residues as a static modification. Peptide lengths were set to a minimum of 7 and a maximum of 50 amino acids. MSBooster was enabled, and Percolator was used for PSM validation. ProteinProphet controlled the protein FDR at 0.01. Quantification was performed with IonQuant [25] using LFQ with MBR enabled, with advanced options set to default.
AP-MS data analysis
The peptide-level outputs from FragPipe were further analyzed to determine protein abundance and perform differential analysis using R (version 4.4.1). The raw peptide intensities from each MS-injection were normalized based on the total peptide signal and scaled to the median total signal. Peptides which were sparsely identified peptides in less than three samples were excluded from the dataset. Subsequently, the peptides for each protein were summed and ranked across the entire experiment. The six highest abundant peptides (TOP6) were used to infer protein abundances for each sample. If fewer than six peptides were quantified, the protein abundance was inferred from the remaining peptides.
The mean signal of the peptides was used for protein abundance inference. Protein intensities were then log2 transformed, and missing values were imputed by sampling from a normal distribution centered around the lowest 5% quantile of all intensity values. Proteins for which >50% of the quantities were imputed were excluded from further analysis. Common lab contaminants (e.g. sheep wool) were also removed. Finally, the log2FC was calculated using the average signal quantified in ARID5B samples compared to the BirA negative controls. To assess enrichment of interaction partners, we applied a linear model to each protein using the limma package [27]. The model estimated the coefficients for each protein, with contrasts between ARID5B and BirA negative controls fitted using empirical Bayes moderation. The resulting statistics were adjusted for multiple testing using the Benjamini–Hochberg method. All proteins with a log2FC larger or equal to 2 and an adjusted P-value of smaller or equal to 0.05 were categorized as interaction partners.
Reconstruction of literature reported ARID5B PPI-network
PPIs between the ARID5B co-repressor complex (ARID5B, MIER1, C16ORF87, HDAC1, and HADC2) subunits were retrieved from the IntAct [28] and the BioGRID (version 4.4.243) [29] databases. Interactions were limited to those identified through affinity enrichment technologies, specifically “Affinity-Capture MS” and “Affinity-Capture Western blot.” Interactions derived from proximity labeling, co-fractionation, and yeast-two-hybrid were excluded. A total of 80 PPIs covering 35 manuscripts were reported among the complex subunits (Supplementary Table S2). The PPI-network was visualized using Cytoscape (version 3.10.2) [30].
Immunofluorescence
Before cell seeding, 8-well Ibidi chambers (Ibidi, #80827) were treated with Poly-d-Lysine (Thermo Fisher, #A3890401) diluted 1 in 10 with MilliQ water for 10 min at room temperature, before washing each well four times with MilliQ water. For cell seeding, cells were washed once with PBS and incubated with Accutase (Biolegend, #423201) for 3–5 min at room temperature, quenched with complete medium, and filtered once through a 35 µm nylon mesh, before cell counting. Cells were seeded at 80 000 cells per well one day prior to fixation in Ibidi chambers. At the time of fixation, cells were washed twice with PBS before fixation with PBS containing 4% methanol-free formaldehyde (Thermo Fisher Scientific, #28906) for 5 min. The fixative was removed and cells quenched using 10 mM Tris–HCl (Invitrogen, #15567027) pH 7.5 in PBS for 3 min. Cells were permeabilised with PBS containing 0.2% Triton-X-100 (Sigma–Aldrich, #X100-100 ml) for 5 min, washed once with PBS to remove residual detergent, and blocked for 30 min with 0.45 µm filtered 2% BSA (Sigma–Aldrich, #A9418-50G) in PBS (blocking buffer). Cells were incubated with Streptavidin AF568 (Thermo Fisher, #S11225) diluted in blocking buffer for 3 h at room temperature with gentle shaking, followed by 3 × 10 min washes with PBS containing 1.62 µM Hoechst 33 342 (Thermo Fisher Scientific, #H3570).
Microscopy and image processing
Confocal microscopy experiments were imaged on a custom Zeiss LSM 980 microscope fitted with an additional Airyscan2 detector, using a ×63 NA 1.4 oil DIC Plan-Aphrochromat (Zeiss) objective and ZEN 3.3 Blue 2020 software.
Image analysis
Fields of cells were analyzed using a custom automated analysis pipeline written in Python. A Gaussian blur was applied to the Hoechst channel using the skimage “filters” module and the Hoechst channel was then segmented using the “nuclei” module of cellpose [31]. Cells which touched the border of the image were excluded using the “clear_border” functionality from the skimage “segmentation” module. Individual masks were labeled and applied to each cell in the field. The area of the nuclear mask and mean fluorescence within the nuclear mask was then calculated for each segmented cell, and the data output as a csv file for each input condition. Biological replicates were normalised independently, relative to their own internal positive and negative controls and the normalized data then merged for the final figure.
Structural modeling and molecular dynamics simulations of ARID5B complexes
Structural models of the complexes were predicted using AlphaFold v3.0 (AF3) [32]. Structures were trimmed to remove disordered termini with pLDDT scores below 50. Molecular dynamics (MD) simulations were performed using Amber2024 [33]. Systems were prepared with the ff14SB force field [34] and TIP3P water model [35]. Each complex was solvated in a truncated octahedral box with a 20 Å buffer of TIP3P water and neutralized with Na⁺ and Cl⁻ ions. Energy minimization (40 000 steps) was followed by restrained heating from 0 to 300 K, NVT equilibration (500 ps), NPT equilibration (500 ps), and a production run of 500 ns. Production simulations were run with pmemd.cuda using periodic boundary conditions, SHAKE [36] constraints applied to bonds involving hydrogen atoms, and a 4 fs timestep (125M steps, coordinates saved every 100 ps; 5000 frames per trajectory) enabled by hydrogen mass repartitioning. Intermolecular contacts were analyzed using GetContacts (https://github.com/getcontacts/getcontacts), and structures were visualized with VMD 1.9.4 [37]. In all, ten protein complexes were simulated, each in two independent runs (run 1 and run 2) using the trimmed AlphaFold predicted structure. Complexes include ARID5B-C16ORF87, ARID5B-MIER1-C16ORF87, ARID5B-HDAC1, ARID5B-HDAC2, HDAC1-C16ORF87, HDAC2-C16ORF87, HDAC1-MIER1-C16ORF87, HDAC2-MIER1-C16ORF87, ARID5B-HDAC1-MIER1-C16ORF87, and ARID5B-HDAC2-MIER1-C16ORF87.
Co-IP of V5-tagged ARID5B variants
Lentiviral vectors encoding V5-tagged full-length murine Arid5b, domain deletions or GFP were obtained by PCR and Gateway cloning (Thermo Fisher Scientific, #11789020 and #11791020) into pLEX305 N-term miniturbo V5 (kind gift from Georg Winter, CeMM). Following lentiviral transduction, HAP1 cells were continuously cultured with Puromycin. Cells from one confluent 15 cm dish were washed with PBS, trypsinized, and pelleted at 400 × g for 5 min. For protein harvesting, the cell pellet was resuspended in IP Lysis Buffer [50 mM Tris–HCl (pH 7.5), 150 mM NaCl, 2 mM EDTA, 0.5% Triton X-100, 0.2 mM phenylmethanesulfonyl fluoride (Sigma, #93482)] supplemented with Benzonase (Millipore Sigma, #706643, 1:200) and Protease Inhibitor Cocktail (Sigma–Aldrich, #P8340-%, 1:100). Samples were incubated on ice for 30 min, with in between vortexing every 10 min. Following, the lysate was centrifuged at 10 000 × g for 10 min at 4°C and the protein in the supernatant quantified using the Pierce BCA Protein Assay Kit (Thermo Fisher Scientific, #23225). V5 Sepharose Beads (Merck, #A7345) were washed five times with IP Lysis Buffer. One milligram of protein in 400 µl of IP Lysis Buffer was incubated with 50 µl of washed V5 Sepharose Beads overnight at 4°C with rotation. The following day, IP samples were centrifuged at 3000 × g for 1 min, and supernatant (unbound fraction) was collected in a new tube. Beads were washed three times with 500 µl of IP Lysis Buffer and proteins eluted from the beads upon addition of 1× Laemmli Sample Buffer (BioRad, #1610747) and incubation at 95°C for 10 min. Samples were centrifuged at 3000 × g for 1 min and the supernatant (IP) transferred to a new tube. IP and input samples (18.75 µg—3.75% of original protein material used for IP) were analyzed by western blot.
Western blot
Proteins were denatured using XT Sample Buffer (BioRad, #1610791) or 4 × Laemmli Sample Buffer (BioRad, #1610747) for 10 min at 95°C. 4%–12% Criterion XT Bis-Tris Protein Gels (BioRad, #3 450 125) or 4%–20% Mini-Protean TGX Precast Protein Gels (BioRad, #4561096) were used to separate the proteins, which were then transferred to low-fluorescence PVDF membranes (BioRad, #1704275) using the High MW program of the Trans-Blot Turbo Transfer System (BioRad, #1704150). The membranes were blocked for 1 h at RT, prior to overnight primary antibody incubation, followed by 2 h RT secondary antibody incubation in EveryBlot blocking buffer (BioRad, #12010020) (see Supplementary Table S1—Antibodies). The BioRad ChemiDoc imager was used to develop the membranes.
CUT&RUN
CUT&RUN for histone marks, transcription factors, and transcription co-factors was performed as previously described [38]. First, 10 µl of ConA beads (Polysciences, #86 057) per reaction were activated by washing twice with 100 µl of Bead Activation Buffer [20 mM HEPES (pH 7.9), 10 mM KCl, 1 mM CaCl2, and 1 mM MnCl2] and finally resuspended in 10 µl of Bead Activation Buffer.
Following, 250 000 and 1 million cells were used for the profiling of histone marks and transcription factors/co-factors, respectively. When applicable, cells were treated with 500 nM Entinostat – MS-275 (Selleckchem, #S1053) for 24 h prior to nuclear extraction. Cells were pelleted in FACS tubes at 400 × g for 3 min, washed with PBS, and resuspended in 100 µl of NE Buffer [20 mM HEPES (pH 7.9), 10 mM KCl, 0.1% Triton X-100, 20% glycerol, 1 mM MnCl2, and 0.5 mM Spermidine) supplemented with 1× cOmplete EDTA-free protease inhibitor cocktail (Roche, #11836170001). Reactions were incubated on ice for 10 min and centrifuged at 400 × g for 3 min. Nuclear pellets were resuspended in 100 µl of NE Buffer supplemented with 1 × cOmplete EDTA-free protease inhibitor cocktail (Roche, #11836170001) and directly mixed with 10 µl of activated ConA beads (Polysciences, #86057). Nuclei were allowed to bind ConA beads (Polysciences, #86057) for 10 min at RT. The supernatant was removed and the nuclei-bound beads were resuspended in 50 µl of Antibody Buffer [20 mM HEPS (pH 7.5), 150 mM NaCl, 0.5 mM Spermidine, 0.02% digitonin, 2 mM EDTA, and 1× cOmplete EDTA-free protease inhibitor cocktail (Roche, #11836170001)] containing 1 µl of the respective antibody (see Supplementary Table S1—Antibodies). Tubes were nutated at 4°C overnight. The following day, beads were washed twice with 200 µl of Digitonin Buffer [20 mM HEPS (pH 7.5), 150 mM NaCl, 0.5 mM Spermidine, 0.02% digitonin, and 1× cOmplete EDTA-free protease inhibitor cocktail (Roche, #11836170 001)] and resuspended in 50 µl of Digitonin Buffer with pAG-MNase (Institute of Molecular Pathology, VBC, Vienna, Austria). Beads were incubated for 30 min at RT and then washed twice with 200 µl of Digitonin Buffer. Following, MNase was activated as beads were resuspended in 50 µl of Digitonin Buffer supplemented with 2 mM CaCl2 and incubated for 30 min at 4°C. The reaction was quenched upon addition of 33 µl of Stop Buffer (340 mM NaCl, 20 mM EDTA, 4 mM EGTA, 50 µg/ml RNaseA, and 50 mg/ml glycogen). Following incubation at 37°C for 10 min, the supernatant was transferred to a fresh tube and DNA was purified using the NEB Monarch PCR &DNA Clean-up Kit (NEB, #T1030).
Libraries were prepared as previously described [39], using the NEBNext Ultra II DNA Library Prep Kit for Illumina (NEB, #E7645) and the NEBNext Multiplex Oligos for Illumina (Dual Index Set 1) (NEB, #7600). Two nanograms of CUT&RUN DNA was used as input for library preparation using 1.5 µM adaptors, 14–16 PCR amplification cycles and a one-sided 0.9 × size selection.
CUT&RUN analysis
Sequencing reads (150 bp paired-end, NovaSeq 6000) were trimmed using TrimGalore (version 0.6.6) and aligned to human genome, hg38, using Bowtie2 (version 2.4.2) [40] (–local, –very-sensitive, –no-mixed, –no-discordant, –dovetail). SAMtools (version 1.15.1) [41] was used to fix mates of paired end reads, remove PCR duplicates of IgG sample, merge, sort, and index bam files, while a CUT&RUN-specific blacklist [42] was removed from the bam files using BEDtools (version 2.30.0) [43].
DeepTools (version 3.5.1) [44] served to build the genome coverage (bamcoverage –normalizeUsing cpm), to compare bigwig files and plot the correlation matrix across replicates (multiBigwigSummary and plotCorrelation), compute FRiP scores (plotEnrichment), compute the matrix, and plot the profile of histone marks, ARID5B, HDAC1, and HDAC2 (computeMatrix and plotHeatmap). Average signal intensities of CPM normalized bigwig files around ARID5B peaks or TSS of regulated genes (±5 kb) were calculated directly from the matrix obtained from deepTools (version 3.5.1) [44]. ggplot2 (version 3.5.1) [45] was used to calculate the spearman correlation coefficient and to plot per-replicate signal intensity concordance of ARID5B signal on the identified 4200 ARID5B peaks. Peaks of histone marks were called using SEACR (relaxed with IgG as control) [46]. For the calling of ARID5B, HDAC1, and HDAC2 peaks, IgG signal was first subtracted from the signal of the protein of interest using bdgcmp from MACS2 (version 2.2.5) (-m subtract) [47]. This was then used as an input file for callpeaks from MACS2 (version 2.2.5) (–cutoff 0.5, -l 200, -g 150; for ARID5B: –cutoff 0.4, -l 200, -g 150) [47]. For downstream analysis, BEDtools (version 2.30.0) [43] was used to intersect peaks, SeqCode (version 1.0) [48] (TSS proximal/promoter: 0–500 bp upstream from TSS, TSS distal: 0.5–2.5 kb upstream from TSS) and ChromHMM (version 1.24) [49] were employed to determine the genomic localization of ARID5B peaks, HOMER (version 4.11.1) [50], ChIP Atlas [51] were utilized to investigate TF motifs and binding on the regions of interest, and DiffBind (version 3.16.0) [52] was run for differential binding analysis (DBA_EDGER, without IgG control, 0.8 ≥ log2FC, and FDR ≤ 0.05). For the HOMER motif analysis [50], de novo motifs, which rely on detecting overrepresented sequences in our regions of interest and subsequently uncovering enriched TF motifs, were used. Genes associated with specific regions were identified using GREAT (two nearest genes, 2000 kb) [53]. Data were visualized using ggplot2 (version 3.5.1) [45].
Processing of publicly available ChIPseq data
Fastq files of publicly available H3K4me3 ChIPseq data in NALM6, and of ARID5B ChIPseq data in HepG2 and Jurkat cells were trimmed using TrimGalore and mapped to the human genome, hg38, using Bowtie2 (version 2.4.2) [40] (–local, –very-sensitive). SAMtools (version 1.15.1) [41] was used to convert SAM to BAM files, sort BAM files, remove PCR duplicates, and index BAM files. Finally, deepTools (version 3.5.1) [44] were used to build the genome coverage (bamcoverage –normalizeUsing cpm). H3K4me3 BAM files were also used as input in the ChromHMM (version 1.24) [49] analysis, as discussed in the section above. TF motif enrichment and co-localization analysis on ARID5B peaks in HepG2 and Jurkat cells was conducted using HOMER (version 4.11.1) [50] and ChIP-Atlas [51], respectively. Homer de novo motifs [50], which uncover TF motifs in the overrepresented sequences in our regions of interest, were used. SeqCode (version 1.0) [48] (TSS proximal/promoter: 0–500 bp upstream from TSS, TSS distal: 0.5–2.5 kb upstream from TSS) was used to determine the genomic distribution of ARID5B in these cell lines.
Processing of publicly available HiChIP data
The analysis pipeline from Dovetail Genomics was used for data processing, as H3K27ac HiChIP libraries were prepared using the Dovetail genomics kit. Reads were mapped to the human genome, hg38, using the bwa algorithm (version 0.7.17) [54]. Valid pairs were obtained by running HiC-Pro (version 3.1.0) [55] and considering MNase digested DNA fragments. Peak-to-all loops were called using FitHiChIP (version 11.0) [56] at a resolution of 10 kb bins, and a low and high distance threshold of 20 kb and 2 Mb. BEDtools (version 2.30.0) [43] was used to intersect loops with ARID5B peaks and regions gaining H3K27ac upon ARID5B KO.
Processing of DepMap data
Depmap gene effect data (DepMap Public 25Q3 + Score, Chronos) [57] was used to assess ARID5B dependency in cell lines of various tissues. The data were plotted using GraphPad Prism (version 10.4.1).
RNA isolation and RT-qPCR
RNA was harvested from 2 million cells using the RNeasy Plus Mini Kit (Qiagen, #74134). One microgram of RNA was used for complementary DNA (cDNA) synthesis using the iScript cDNA Synthesis Kit (BioRad, #1708891). Quantitative PCR (qPCR) was performed using the iQ5 SYBR Green Supermix (BioRad, #1708885) and ran in the CFX Connect Real Time System (BioRad). An initial denaturation at 95°C for 3 min, followed by 40 cycles of 10 s denaturation at 95°C and 30 s annealing and extension at 60°C was used (see Supplementary Table S1—Oligonucleotides).
mRNA-seq library preparation
RNA was isolated as described above and isent to Azenta (Leipzig, Germany) for QC and library preparation followed by next-generation sequencing (150 bp paired-end, NovaSeq 6000). When applicable, cells were treated with 500 nM Entinostat—MS-275 (Selleckchem, #S1053) for 24 h prior to RNA harvesting. Al mRNA-seq experiments were conducted using biological triplicates per condition.
mRNA-seq analysis
Sequencing reads were trimmed using TrimGalore (version 0.6.6), aligned to the human transcriptome, hg38, using STAR (version 2.7.9) [58], and counted with HTseq-count (version 0.11.2) [59, 60]. DESeq2 (version 1.46.0) [61] was used for differential gene expression (0.8 log2FC, basemean expression ≥ 30 and padj ≤ 0.05). Gene Set Enrichment Analysis (GSEA) (version 4.3.2) [62] was run using the normalized read counts from DEseq2 (version 1.46.0) [61] of expressed genes (normalized read counts ≥ 30) with default settings and a gene set permutation type. pheatmap and ggplot2 (version 3.5.1) [45] were used for data visualization.
Statistical analysis
R and the limma package [27] were used to identify significantly enriched proteins in the interaction proteomics dataset. P-values were calculated using empirical Bayes moderation, followed by multiple testing correction using the Benjamini-Hochberg procedure. All other statistical analyses were conducted using GraphPad Prism (version 10.4.1). The Kolmogorov–Smirnov test was employed when comparing histone mark or co-factor signal intensities on ARID5B peaks, and when comparing ARID5B signal intensity on the TSS of ARID5B and Entinostat regulated, ARID5B only and Entinostat only regulated genes, as the results are not expected to follow a normal distribution (non-parametric) and as the average occupancy on all of these sites (unpaired) was compared. Unpaired t-tests were performed when comparing gene expression between two conditions. One-way ANOVA followed by Tukey multiple comparison test was performed when comparing gene expression with more than two conditions. P ≤ 0.05 were considered statistically significant. *P ≤ 0.05, **P ≤ 0.01, ***P ≤ 0.001, ns = not significant.
Materials availability
Cell lines and expression vectors generated in this study are available upon request. Further information and requests for resources and reagents should be directed to and will be fulfilled by the lead contact (davide.seruggia@ccri.at).
Results
ARID5B assembles in a complex with HDAC1, HDAC2, MIER1, and C16ORF87
To map the protein–protein interactions of ARID5B, we conducted affinity purification followed by mass spectrometry (AP-MS) (Fig. 1A). For this, we overexpressed the long isoform of the murine ARID5B tagged with a FLAG-AviTag in HEK293T cells expressing the E. coli biotin ligase BirA (Supplementary Fig. S1A and B). In the presence of biotin, existing in low amounts in the cell culture medium, BirA biotinylates the AviTag, facilitating the affinity purification of ARID5B and its interactors using streptavidin beads, which can be quantified by mass spectrometry [23]. Notably, PHF2 was not identified, and consequently no interaction of ARID5B with PHF2 was recovered, contrary to previous reports [20, 21] (Fig. 1B and Supplementary Table S3). Instead, we detected significant interactions with HDAC1, HDAC2, MIER1, and C16ORF87. While HDAC1 and HDAC2 are components of six transcriptional repressor complexes [63, 64], only one complex has been reported to include the ELM-SANT-domain containing protein MIER1 [63, 65] and more recently, also C16ORF87 [66]. The MIER1–HDAC1/2 complex has been described to also associate with BAH-domaining proteins such as BAHD1 and histone octamers, potentially exerting a chromatin remodeling function [65]. Finally, the MIER1–HDAC1/2 complex regulates lipid metabolic genes [67, 68], as well as cell cycle genes in the regenerating liver [67]. Importantly, ARID5B is identified in the reciprocal affinity enrichment of HDAC1, HDAC2, MIER1, and C16ORF87 in publicly available PPI databases IntAct [28] and BioGRID (version 4.4.243) [29] (see methods “Reconstruction of literature reported ARID5B PPI-network” for filtering parameters). A protein interaction network based on PPI-databases covers 18 out of all 20 potential edges/interactions among ARID5B and its interaction partners (Supplementary Fig. S1C, Database resource and reference for each specific bait-prey interaction can be found in Supplementary Table S2), suggesting that ARID5B associates with HDAC1, HDAC2, MIER1, and C16ORF87 into a protein complex.
Figure 1.
ARID5B forms a protein complex with HDAC1, HDAC2, MIER1, and C16ORF87. (A) Experimental overview of the identification of ARID5B interaction partners. A construct encoding the murine ARID5B containing an N-term FLAG-AviTag was lentivirally overexpressed in HEK293T cells expressing the Biotin ligase BirA. The latter biotinylates the AviTag in the presence of Biotin, found in the cell culture medium. ARID5B and its interactors were then pulled down using streptavidin beads and quantified by mass spectrometry. (B) Scatterplot showing enriched proteins in the AP-MS of ARID5B against the Avi-Tag negative control (n = 3 biological replicates for samples and controls). Dotted lines indicate a log2FC threshold of larger or equal to 2 and adjusted P-value threshold of smaller or equal to 0.05 (linear model followed by eBayes for moderated t-statistics, and P-value adjustment with Benjamini–Hochberg procedure). Scored interaction partners (purple) and bait (pink) are labeled. (C) ipTM scores from AlphaFold3 models for all predicted complex combinations. (D and E) AlphaFold3-predicted structures of ARID5B–MIER1–C16ORF87 in complex with either HDAC1 (D) or HDAC2 (E). ARID5B is shown in blue, MIER1 in purple, C16ORF87 in orange, and HDAC1 (D) or HDAC2 (E) in green/light green, respectively. Disordered regions with low pLDDT scores are omitted for clarity. (F) ARID5B (blue) in complex with HDAC1 (green) as sampled during molecular dynamics simulations. The BAH domain of ARID5B is highlighted in red. (G) Predicted binding interface of ARID5B (blue) and HDAC1 (green), highlighting interacting residues (sticks) identified by GetContacts analysis during molecular dynamics simulations. (H) Co-immunoprecipitation of V5-tagged GFP, full-length ARID5B (FL), and mutants lacking the ARID (ΔARID) or BAH (ΔBAH) domains followed by immunoblotting of HDAC1 and HDAC2 in HAP1 cells.
To better understand the organization of ARID5B with its interactors, we performed structural predictions of complexes containing different interaction partners (Fig. 1C–E; Supplementary Fig. S1D and Supplementary Table S4). Complexes containing ARID5B, MIER1, C16ORF87, and HDAC1 or HDAC2 displayed interface predicted template modeling (ipTM) scores above 0.6, strongly supporting the assembly of ARID5B into a complex (Fig. 1C). Interestingly, the prediction of complexes containing C16ORF87 and HDAC1 or HDAC2 drastically improved upon addition of MIER1 and complexes lacking HDAC1 and HDAC2 had the lowest ipTM scores. Chain-pair ipTM scores further predict strong interactions of HDAC1 and/or HDAC2 with all subunits (Supplementary Fig. S1E and F). Notably, among the predicted complex–subunit interactions, ARID5B exhibits the highest-confidence score with HDAC1 and HDAC2, as indicated by an ipTM interaction score of 0.73 and 0.72, respectively (Supplementary Fig. S1E and F).
To further examine the predicted AlphaFold structures, we used a selection of ten structures and investigated the incorporation of C16ORF87 by relative binding affinities using MD simulations (Supplementary Fig. S1G). Across the complexes, the binding affinity for dimers remains between −90 and −185 kcal/mol (ARID5–C16ORF87, HDAC1/2–C16ORF87), but this increases to a range of −275 to −373 kcal/mol upon the addition of MIER1 to the same dimers. The consistent decrease of binding free energy upon the inclusion of MIER1 as a subunit indicates that binding of C16ORF87 is most likely facilitated by MIER1. This is in line with a recent study describing that C16ORF87 forms a stable complex with MIER1 and HDAC [66]. Together with the overall complex structure, this suggests that MIER1, C16ORF87, and HDAC are likely to form a scaffold, which enables the interaction of ARID5B.
While ARID5B consists mainly of intrinsically disordered regions, the BAH (aa 34–108) and ARID (aa 319–411) domains at the N-terminus have an average pLDDT score of 82, indicating a more confidently predicted and well-defined structure (Supplementary Fig. S2A and B). The N-terminus of ARID5B is also predicted to interact with HDAC1 and HDAC2, with residues 1–250 consistently having the lowest prediction alignment error across models with either HDAC1 or HDAC2 (Supplementary Fig. S2C and D). Molecular dynamics simulations of the ARID5B–HDAC1 complex revealed that residues Asn87, Ser88, Asp89, and Glu92 of the BAH domain, along with an α-helix of ARID5B (aa 181–193), are located at the interface and contribute to stabilizing the interaction with HDAC1 (Fig. 1F and G). Furthermore, the C-terminal region (aa 329–400) forms another interaction hotspot. To corroborate these findings, we performed molecular dynamics simulations of ARID5B–HDAC2, for which the interface shows the same three distinct regions (Supplementary Fig. S2E and F). The HDAC1 and HDAC2 interfaces with ARDI5B share 16 hotspot residues, confirming that the two HDACs engage essentially the same binding surface on ARID5B, including aa 89, 189, 193, 341, 396, 387, and 389. Among them, Arg341 is the single strongest contributor in both HDACs (∼ −5.9 kcal/mol), indicating a critical anchor point. Finally, the binding free energy is comparable for both complexes (average across both runs: ARID5B-HDAC1: −101.15 ± 16.15 kcal/mol; ARID5B–HDAC2: −93.4 ± 16.2 kcal/mol). Co-immunoprecipitation of murine V5-tagged ARID5B, and domain deletion mutants (ΔARID and ΔBAH) in HAP1 cells confirmed that the interaction of ARID5B with HDAC1 and HDAC2 is mediated by the BAH domain (Fig. 1H; Supplementary Fig. S2G and H). This further suggests that while the BAH domain of ARID5B is responsible for PPIs in the complex, ARID may remain accessible for DNA binding.
ARID5B localizes on active genomic regions
To map ARID5B binding on the genome, we performed CUT&RUN in wild-type (WT) and ARID5B KO NALM6 cells, a cell line derived from a B-ALL patient (Supplementary Fig. S3A–D). We identified 4200 genomic regions occupied by ARID5B in WT cells, respectively (Supplementary Fig. S3E and Supplementary Table S5), all showing a FRiP score enrichment over their IgG counterpart and a satisfactory replicate-level concordance (Supplementary Fig. S3F and Supplementary Table S5). Loci occupied by ARID5B enriched for motif and binding of B-cell lineage-specific transcription factors, such as IKZF1, EBF1, PAX5, and TCF3 (Fig. 2A and B). A similar analysis of ARID5B publicly available ChIPseq data in Jurkat (T cells) and HepG2 (liver cells) cells showed that ARID5B enriches for TF motifs and co-localizes with T cell and liver specific factors, such as NOTCH1, RUNX1, ETS1, MYB, and HNF4A, FOXA1, FOXA2, respectively (Supplementary Fig. S4A–D). Furthermore, only 2.67% of ARID5B peaks in NALM6 (B cells) overlapped with ARID5B peaks in Jurkat (T cells) cells (Supplementary Fig. S4E and Supplementary Table S6), confirming that ARID5B associates with lineage-specific factors and that its genomic localization is influenced by these specific factors.
Figure 2.
ARID5B localizes on active genomic regions. (A) Five most significant HOMER de novo motifs enriched on 4200 ARID5B peaks. (B) Enrichment analysis using ChIP Atlas of publicly available TFs or co-factors co-localizing with 4200 loci occupied by ARID5B in NALM6 WT cells. The analysis was restricted to factors probed in blood cells by setting the “Cell Class” filter to “Blood.” The bar plot depicts the top factors and their overlap with ARID5B occupied regions. Only data of endogenous, non-tagged, ChIP-seq in “blood” cells (depicted are overlaps of ChIP-seq peaks in NALM6, RCH-ACV, SEM, Acute Lymphoid Leukemia, RS-411, Pre-B cells, and OCI-LY1) with an FDR ≤ to 0.05 is shown. (C) Genome-wide coverage of ARID5B in WT and ARID5B KO NALM6 cells, and of active (H3K27ac, H3K4me1) and repressive (H3K27me3, H3K9me2) histone marks in NALM6 WT cells at the TSC1 GFI1B locus. The coverage of two merged biological replicates per histone mark and of ARID5B is shown. (D) Aggregate plots and heatmaps of ARID5B, active and repressive histone marks in NALM6 cells centered on the previously determined 4200 ARID5B peaks. (E) ChromHMM-defined chromatin states in NALM6 cells based on ARID5B occupancy and active (H3K27ac, H3K4me1, H3K4me3) and repressive (H3K27me3, H3K9me2) histone marks. The coverage of two merged biological replicates per histone mark was used for the analysis. (F) Genomic distribution of the 4200 ARID5B peaks called in NALM6 cells. (G and H) Volcano plots of H3K27ac (G) and H3K4me1 (H) differentially bound sites upon ARID5B loss. EdgeR-based DiffBind analysis was used to identify loci losing (blue) or gaining (red) H3K27ac or H3K4me1 occupancy in ARID5B KO cells compared to NALM6 WT counterparts. Dotted lines indicate a log2FC threshold of larger or equal to 0.8 and a FDR threshold of smaller or equal to 0.05. Two biological replicates per histone mark per genotype were used for the analysis. (I) Violin plot of H3K27ac and H3K4me1 average signal intensity on 4200 ARID5B peaks in WT and ARID5B KO NALM6 cells. The average signal intensity for each region was calculated within ±5 kb bins using compute matrix from deeptools. The median signal intensity is shown in red and quartiles are depicted by dotted lines. For statistics, the nonparametric and un-paired Kolmogorov–Smirnov test was performed. The coverage of two merged biological replicates per histone mark per genotype was used for the analysis. n = 4 200, *P ≤ 0.05, **P ≤ 0.01, ***P ≤ 0.001, ns = not significant.
Furthermore, ARID5B co-localized with the active histone marks H3K27ac and H3K4me1 and did not co-localize with the repressive histone mark, H3K9me2, the substrate of PHF2 [69, 70] (Fig. 2C and D). Chromatin state analysis by ChromHMM indicated that ARID5B localizes in active TSSs and enhancers (Fig. 2E), with 9.8% and 12.7% of ARID5B localizing to TSS proximal (0–500 bp upstream of TSS, promoter region) and distal (0.5–2.5 kb upstream of TSS) and 29.7% and 35% of ARID5B peaks occupying intronic and intergenic genomic regions, respectively (Fig. 2F). A similar ARID5B genomic distribution was observed in Jurkat and HepG2 cells, with ARID5B occupying TSS proximal and distal loci, as well as intronic and intergenic regions (Supplementary Fig. S4F and G).
To observe the consequences of ARID5B loss globally on chromatin, we extended our CUT&RUN analysis to NALM6 KO cells and compared the distribution of H3K4me1 and H3K27ac in WT and ARID5B KO NALM6 cells. While ARID5B loss did not affect global H3K27ac and H3K4me1 levels in the cell (Supplementary Fig. S5A), it led to local genomic changes with 980 loci gaining and 428 loci losing H3K27ac upon ARID5B KO, and 976 and 852 regions gaining and losing H3K4me1 in KO cells, respectively (Fig. 2G and H; Supplementary Fig. S5B and Supplementary Table S7). To investigate the direct relationship of ARID5B occupied sites with regions gaining H3K4me1 and chromatin loops gaining H3K27ac, we performed co-localization analysis. We observed very little overlap of regions occupied by ARID5B and gaining H3K4me1 upon ARID5B KO (Supplementary Fig. S5C and Supplementary Table S7). We further leveraged publicly available H3K27ac HiChIP data and observed that 6.4% of loci gaining H3K27ac either directly co-localized or interacted with ARID5B bound regions (Supplementary Fig. S5D and Supplementary Table S7). Still, overall H3K27ac signal intensity was markedly higher on ARID5B peaks in KO compared to WT cells (Fig. 2I and Supplementary Fig. S5E).
ARID5B recruits HDAC1 and HDAC2 to the genome
The fact that ARID5B interacts with histone deacetylases and that H3K27ac signal is increased on ARID5B peaks in KO cells (Fig. 2I), suggests a loss of HDAC activity at ARID5B occupied sites. Thus, to further dissect the chromatin regulatory role of the ARID5B complex, we performed CUT&RUN of HDAC1 and HDAC2. Strikingly, we observed that 76.3% (3206 peaks) and 77.4% (3250 peaks) of ARID5B bound sites are occupied by HDAC1 and HDAC2, respectively, with 2872 ARID5B peaks being bound by all three factors (Fig. 3A and Supplementary Table S8). These sites displayed a similar genomic distribution as seen before (Fig. 2F), localizing mainly at distal and proximal TSSs, as well as intronic and intergenic regions (Supplementary Fig. S6A). To further investigate the interaction of ARID5B with HDACs and the assembly of the complex on chromatin, we compared HDAC binding in WT and ARID5B KO cells. We uncovered that HDAC1 and HDAC2 occupancy on ARID5B peaks was reduced upon loss of ARID5B (Fig. 3B). Strikingly, 12 765 and 6061 sites significantly lost HDAC1 and HDAC2 signal upon ARID5B loss, respectively (Fig. 3C and D, and Supplementary Table S9). Interestingly, 46.5% (1955 peaks) of ARID5B occupied regions lost HDAC1 and/or HDAC2 binding (Supplementary Fig. S6B and Supplementary Table S9). Of those, 1501 ARID5B peaks showed a loss of HDAC1 occupancy only, while 183 and 271 ARID5B peaks overlapped with regions losing HDAC2 or HDAC1 and HDAC2 signal, respectively (Fig. 3E; Supplementary Fig. S6B and Supplementary Table S9). Together, this shows that the loss of ARID5B mainly affects the localization of HDAC1, suggesting a preference for ARID5B–HDAC1 containing complexes on chromatin (Fig. 3F). Importantly, no changes in HDAC1 and HDAC2 mRNA as well as protein levels were observed upon ARID5B KO (Fig. 3G and Supplementary Fig. S6C), indicating that ARID5B does not regulate HDAC1 and HDAC2 expression but rather tethers these factors to chromatin.
Figure 3.
ARID5B tethers HDAC1 and HDAC2 to the genome. (A) Venn Diagram of ARID5B, HDAC1, and HDAC2 peaks in NALM6 WT cells. (B) Violin plot of HDAC1 and HDAC2 average signal intensity on 4200 ARID5B peaks in WT and ARID5B KO NALM6 cells. The average signal intensity for each region was calculated within ±5 kb bins using compute matrix from deeptools. The median signal intensity is shown in red and quartiles are depicted by dotted lines. For statistics, the nonparametric and unpaired Kolmogorov–Smirnov test was performed. The coverage of two merged biological replicates per IP per genotype was used for the analysis. n = 4 200, *P ≤ 0.05, **P ≤ 0.01, ***P ≤ 0.001, ns = not significant. (C and D) Volcano plots of HDAC1 (C) and HDAC2 (D) differentially bound peaks upon ARID5B loss. EdgeR-based DiffBind analysis was used to identify loci losing (blue) or gaining (red) HDAC1 or HDAC2 occupancy in ARID5B KO compared to WT NALM6 cells. Dotted lines indicate a log2FC threshold of larger or equal to 0.8 and a FDR threshold of smaller or equal to 0.05. Two biological replicates for HDAC1 or HDAC2 per genotype were used for the analysis. (E) Heatmaps of ARID5B, HDAC1, and HDAC2 in WT and ARID5B KO NALM6 cells centered on ARID5B peaks losing HDAC1 and HDAC2 signal (271 loci), ARID5B peaks losing HDAC1 signal only (1501 loci) and ARID5B peaks losing HDAC2 signal only (183 loci). The coverage of two merged biological replicates per IP per genotype was used for the analysis. (F) Genome-wide coverage of ARID5B, HDAC1 and HDAC2 in WT and ARID5B KO NALM6 cells at the GCSAM locus. The coverage of two merged biological replicates of HDAC1, HDAC2, or ARID5B per genotype is shown. (G) HDAC1 and HDAC2 expression in NALM6 WT and ARID5B KO cells assessed by RT-qPCR. Relative expression was obtained upon normalization to the housekeeping gene, GUSB. Individual data points, mean and standard deviation are depicted by dots, bars and whiskers, respectively. For statistics, unpaired two-tailed Student’s t-test was run with n = 3 and *P ≤ 0.05, **P ≤ 0.01, ***P ≤ 0.001, ns = not significant. The qPCR was run using technical duplicates of every biological replicate (n = 3) sample.
Loss of ARID5B leads to de-repression of HDAC target genes altering B-cell–specific programs
To assess whether ARID5B mediates the repression of those sites showing loss of HDAC1 and HDAC2, we conducted H3K27ac CUT&RUN upon Entinostat treatment, a HDAC Class I inhibitor [71] (Fig. 4A and Supplementary Table S10). As HDAC1 and HDAC2 catalyze the de-acetylation of H3K27ac [72, 73], we compared peaks gaining H3K27ac upon HDAC inhibition, with peaks losing HDAC1 or HDAC2 upon ARID5B KO. Indeed, we observe that 37.6% (553 peaks) and 21.3% (313 peaks) of regions gaining H3K27ac upon Entinostat treatment lose HDAC1 or HDAC2 binding upon ARID5B loss (Fig. 4B and C, and Supplementary Table S10). Together, this suggests that a subset of HDAC-regulated loci are regulated by ARID5B in concert with HDAC1 and/or HDAC2, forming a chromatin repressor complex.
Figure 4.
ARID5B and HDACs regulate B-cell specific processes. (A) Volcano plot depicting differential occupancy of H3K27ac upon Entinostat treatment in NALM6 WT cells. EdgeR-based DiffBind analysis was used to identify loci losing (blue) or gaining (red) H3K27ac occupancy in Entinostat treated cells compared to vehicle-treated NALM6 WT counterparts. Dotted lines indicate a log2FC threshold of larger or equal to 0.8 and a FDR threshold of smaller or equal to 0.05. NALM6 WT cells were treated with vehicle or 500 nM Entinostat for 24 h prior to CUT&RUN experiment. Two biological replicates per treatment were used for the analysis. (B and C) Venn Diagram of peaks losing HDAC1 (B) or HDAC2 (C) occupancy upon ARID5B loss and peaks gaining H3K27ac upon HDAC inhibition by Entinostat in NALM6 WT cells. Differentially bound sites were identified using the DiffBind (edgeR mode) package and the following threshold were set for lost (log2FC ≤ −0.8, FDR ≤ 0.05) and gained regions (log2FC ≥ 0.8, FDR ≤ 0.05). NALM6 WT cells were treated with vehicle or 500 nM Entinostat for 24 h prior to CUT&RUN experiment. Two biological replicates per treatment and per HDAC1 or HDAC2 per genotype were used for the analysis. (D) Heatmap of up- and downregulated genes upon ARID5B loss. Depicted are the Z scores for each regulated gene across biological triplicates in NALM6 WT and ARID5B KO cells. The following thresholds were used for up [log2FC ≥ 0.8, padj ≤ 0.05, cpm (KO) ≥ 30] and downregulated [log2FC ≤ −0.8, padj ≤ 0.05, cpm (WT) ≥ 30] genes. Representative genes are shown. (E) Expression of ARID5B and Entinostat target genes. Relative expression was obtained upon normalization to the housekeeping gene, GUSB. Individual data points, mean and standard deviation are depicted by dots, bars and whiskers, respectively. For statistics, one-way ANOVA followed by Tukey multiple comparison test was run with n = 3 and *P ≤ 0.05, **P ≤ 0.01, ***P ≤ 0.001, ns = not significant. The qPCR was run using technical duplicates of every biological replicate (n = 3) sample. (F) Violin plot of ARID5B average signal intensity on TSSs of genes upregulated by ARID5B and upon Entinostat treatment and of genes upregulated only by ARID5B or upon Entinostat treatment. The average signal intensity for each region was calculated within ±5 kb bins using compute matrix from deeptools. The median signal intensity is shown in red and quartiles are depicted by dotted lines. For statistics, the nonparametric and unpaired Kolmogorov–Smirnov test was performed. The coverage of two merged biological replicates per genotype was used for the analysis. n = 305 (ARD5B & Entino), 685 (ARID5B only) and 283 (Entino only), *P ≤ 0.05, **P ≤ 0.01, ***P ≤ 0.001, ****P ≤ 0.0001 ns = not significant. (G) Top 10 GO-terms enriched in gene set enrichment analysis (GSEA – C2 collection) of ARID5B KO and Entinostat treated NALM6 cells. Depicted are the normalized enrichment scores (NES) and the number of genes contained in each gene ontology term. GSEA was run comparing NALM6 WT with ARID5B KO cells and comparing vehicle-treated with Entinostat-treated cells. Terms with an FDR smaller or equal to 0.25 were considered statistically significant. NALM6 WT cells were treated with vehicle or 500 nM Entinostat for 24 h prior to mRNAseq experiment. Biological triplicates per genotype and per treatment were used. (H) Simplified schematic of BCR and the co-receptor CD19, as well as downstream signaling intermediates, which are activated upon antigen binding to BCR. (I) Western blot depicting pAkt, Akt, pErk, Erk levels in NALM6 WT and ARID5B KO cells stimulated with 10 µ/ml Igµ for 10 min.
To investigate the effect of the ARID5B complex on transcription, we performed mRNAseq on WT and ARID5B KO cells. This analysis revealed that more genes are de-repressed and thus more highly expressed upon ARID5B loss, with 990 upregulated and 662 downregulated genes (Fig. 4D; Supplementary Fig S6A and Supplementary Table S11). Importantly, 305 genes are upregulated upon ARID5B loss and HDAC inhibition by Entinostat treatment, suggesting that these genes are regulated by the ARID5B repressor complex in a HDAC-dependent manner (Fig. 4E; Supplementary Fig. S7B and Supplementary Table S11). Interestingly, 16% of these genes are associated with regions bound by ARID5B and losing HDAC1 and/or HDAC2 occupancy (Supplementary Fig. S6B; Fig. 3F; Supplementary Fig. S7C; Supplementary Table S11), further strengthening the link that these genes are indeed actively repressed by ARID5B–HDAC1 and/or ARID5B–HDAC2 complexes. Finally, we observed that ARID5B signal intensity is higher on the TSS of the identified ARID5B regulated genes, compared to the ARID5B occupancy on the TSS of Entinostat-only upregulated genes (Fig. 4F and Supplementary Fig. S7C–F). This suggests that ARID5B only as well as ARID5B and Entinostat commonly upregulated genes are actively regulated by ARID5B, while Entinostat only regulated genes are repressed by other HDAC-containing repressor complexes.
Gene ontology analysis showed that genes upregulated only upon ARID5B loss or only upon HDAC inhibition are mainly involved in DNA replication or antigen presentation, respectively (Supplementary Fig. S8A and B, and Supplementary Table S12). On the other hand, genes upregulated in KO and upon Entinostat treatment, thus repressed by the ARID5B complex, are implicated in B-cell specific cellular processes and signaling pathways (Fig. 4G and Supplementary Table S12). Importantly, targets of the ARID5B repressor complex are involved in B-cell proliferation, calcium signaling regulation, and BCR signaling. While no difference in cellular proliferation was observed upon the loss of ARID5B in NALM6 cells or other cells of various lineages (Supplementary Fig. S8C and D), KO of ARID5B led to altered response to BCR signaling stimulation. To test the role of ARID5B in BCR signaling, we stimulated WT and ARID5B KO cells with Igμ and observed that both showed increased Erk phosphorylation upon stimulation (Fig. 4H and I). Interestingly, KO cells displayed an additional increase in phosphorylated Akt levels upon stimulation, suggesting a difference in antigen-dependent signaling (Fig. 4H and I). Together, this indicates that the ARID5B repressor complex plays a crucial role in BCR signaling regulation.
Taken together, we show that ARID5B associates with HDAC1 or HDAC2, MIER1, and C16ORF87 forming a novel co-repressor complex. ARID5B further tethers the complex to TSSs and distal regulatory elements, actively repressing B-cell–specific genes (Graphical abstract).
Discussion
Here we show that ARID5B forms a repressor complex with HDAC1 or HDAC2, MIER1, and C16ORF87. Importantly, the BAH domain of ARID5B mediates the interaction with HDAC1 and HDAC2, which is further supported by MIER1 and C16ORF87 binding. On a genome-wide level, ARID5B co-localizes with active histone marks, occupying mainly intergenic and intronic regions, while also binding TSS proximal and distal loci. ARID5B tethers HDAC1 and HDAC2 to chromatin, and loss of ARID5B leads to decreased HDAC occupancy, resulting in upregulation of a subset of target genes. Together, ARID5B and HDACs regulate B-cell–specific genes, fine-tuning BCR signaling responses.
One could speculate that the ARID5B repressor complex activity changes during B-cell development and is modulated by B-cell–specific signaling. Importantly, ARID5B expression continuously increases from the pro-B to large and small pre-B-cell stages [16], along with the maturation of BCR [74, 75]. This suggests that the ARID5B repressor complex is a crucial player in regulating pre-BCR and BCR signaling during development. Interestingly, a BCR signaling second messenger, Inositol 1,4,5-triphosphate (InsP3) [74, 76], together with its phosphorylated version, Inositol 1,4,5,6-tetraphosphate, have been shown to enhance the binding of ELM2-SANT domain proteins, such as MIER1, to HDAC and this way promote HDAC activity [77, 78]. Thus, one could speculate that not only the expression of ARID5B during B-cell development but also onset pre-BCR signaling and changes in cellular InsP3 levels, regulate and serve as feedback loops to the timely activity of the ARID5B repressor complex.
Interestingly, the expression of both ARID5B isoforms is differently associated with B-ALL outcome. While increased expression of ARID5B short isoform is associated with lower overall and event-free survival, higher levels of the ARID5B long isoform (that retains the BAH domain required to interact with HDACs), are associated with increased event-free survival [79]. This suggests that the long isoform of ARID5B plays a rather tumor suppressive role, implying that the presence of a BAH domain and consequently the activity of the ARID5B repressor complex is crucial to control transcripts associated with oncogenesis. Alternatively, the short isoform of ARID5B could exert a dominant negative function, which when overexpressed, outcompetes the ARID5B repressor complex on chromatin, thus de-repressing B-cell–specific genes.
Finally, altered activity of HDAC-containing repressor complexes, such as NCoR and NuRD, has been shown to promote the development of hematological malignancies, such as lymphomas [80, 81] and leukemias [82], respectively. While NCoR interacts with ERG to control leukemia-associated genes [82], loss of NuRD favors T-cell lymphoma by de-repressing E2A and creating a B- and T-cell progenitor imbalance [80]. In fact, Mbd3/NuRD ablation in hematopoietic stem cells leads to premature and increased open chromatin regions at B-cell lineages specific enhancers and consequent faster commitment to the B-cell lineage [80]. Still, NuRD loss leads to decreased B cells at early developmental stages, affecting especially the pro-B cell stage and the ratio of large to small pre-B cells [83]. The latter is due to the dysregulation of cell cycle genes and consequent cell cycle arrest of large pre-B cells [84]. Similarly, HDAC1 and HDAC2 have been shown to regulate cell cycle in pre-B and proliferating mature B cells [85]. Thus, HDAC and its complexes have been described to control a plethora of B cell specific processes pre-disposing to leukemia.
A limitation of this study is the high background in the ARID5B CUT&RUN data, which consequently prompted us to set more stringent peak calling parameters, such that the ARID5B peaks profiled in this study may be an underrepresentation of all ARID5B occupied loci in NALM6 cells. Furthermore, as we profile histone and HDAC changes in a constitutive ARID5B clonal cell line, we speculate that some initial changes in H3K27ac, are compensated for, and thus not detected in our analysis. To overcome this, an acute degradation of ARID5B followed by CUT&RUN experiments would be necessary.
Future work will be necessary to study the contribution of the ARID5B repressor complex to B-cell development in vivo. Of importance, is to distinguish the contribution of HDAC1 and HDAC2 in the context of ARID5B as well as other repressor complexes. Monitoring and characterizing the onset of leukemia in mice depleted from components of the ARID5B repressor complex will further provide insights into the contribution and role of the complex in leukemia predisposition. Together, this will elucidate the molecular mechanism by which ARID5B pre-disposes to B-ALL, providing new insights into potential therapeutic dependencies, such as HAT inhibitors.
Supplementary Material
Acknowledgements
The authors thank the Flow Cytometry Facility at CCRI and the Biomedical Sequencing Facility at CeMM. We thank the Proteomics team of the Molecular Discovery Platform at CeMM for mass-spectrometric data acquisition and want to especially acknowledge Thomas Hannich and Anna Tinnacher. We thank Emilia Korczmar for help with illustrations, Georg Winter (CeMM, Vienna) for the pLEX305 N-term miniturbo V5 plasmid and Zain Patel (MGH, Boston) for fruitful discussions on data analysis.
Author contributions: Ana Patricia P. Kutschat (Conceptualization [equal], Formal analysis [equal], Funding acquisition [supporting], Investigation [equal], Methodology [equal], Writing—original draft [equal], Writing—review & editing [equal]), Fabian Frommelt (Formal analysis [equal], Investigation [equal], Methodology [equal], Writing—review & editing [equal]), Brianda L. Santini (Formal analysis [equal], Investigation [equal], Methodology [equal], Writing—review & editing [equal]), Sophie Müller (Investigation [equal]), Paul Batty (Investigation [equal]), Animesh Awasthi (Data curation [supporting], Formal analysis [supporting]), Gerlinde Karbon (Investigation [equal]), Giulio Superti-Furga (Supervision [equal]), and Davide Seruggia (Conceptualization [equal], Funding acquisition [lead], Resources [lead], Supervision [lead], Writing—original draft [equal], Writing—review & editing [equal]).
Contributor Information
Ana P Kutschat, St. Anna Children’s Cancer Research Institute (CCRI), 1090 Vienna, Austria; CeMM Research Center for Molecular Medicine of the Austrian Academy of Sciences, 1090 Vienna, Austria.
Fabian Frommelt, CeMM Research Center for Molecular Medicine of the Austrian Academy of Sciences, 1090 Vienna, Austria.
Brianda L Santini, CeMM Research Center for Molecular Medicine of the Austrian Academy of Sciences, 1090 Vienna, Austria.
Sophie Müller, St. Anna Children’s Cancer Research Institute (CCRI), 1090 Vienna, Austria; CeMM Research Center for Molecular Medicine of the Austrian Academy of Sciences, 1090 Vienna, Austria.
Paul Batty, St. Anna Children’s Cancer Research Institute (CCRI), 1090 Vienna, Austria; CeMM Research Center for Molecular Medicine of the Austrian Academy of Sciences, 1090 Vienna, Austria.
Animesh Awasthi, CeMM Research Center for Molecular Medicine of the Austrian Academy of Sciences, 1090 Vienna, Austria; Medical University of Vienna, Institute of Artificial Intelligence, Center for Medical Data Science, 1090 Vienna, Austria.
Gerlinde Karbon, St. Anna Children’s Cancer Research Institute (CCRI), 1090 Vienna, Austria; CeMM Research Center for Molecular Medicine of the Austrian Academy of Sciences, 1090 Vienna, Austria.
Giulio Superti-Furga, CeMM Research Center for Molecular Medicine of the Austrian Academy of Sciences, 1090 Vienna, Austria; Center for Physiology and Pharmacology, Medical University of Vienna, 1090 Vienna, Austria.
Davide Seruggia, St. Anna Children’s Cancer Research Institute (CCRI), 1090 Vienna, Austria; CeMM Research Center for Molecular Medicine of the Austrian Academy of Sciences, 1090 Vienna, Austria.
Supplementary data
Supplementary data is available at NAR online.
Conflict of interest
G.S.-F. is a scientific founder and shareholder of Proxygen and Solgate Therapeutics and shareholder of Cellgate Therapeutics. The Superti-Furga laboratory has received research funding from Pfizer and Angelini.
Funding
This research was funded in whole or in part by theAustrian Science Fund (FWF) [10.55776/P36302]. For open access purposes, the author has applied a CC BY public copyright license to any author accepted manuscript version arising from this submission. Research in the Seruggia laboratory is further supported by the WES Leukemia Research Foundation, the Austrian Science Fund (FWF, project P36069), the European Research Council (ERC) under the Horizon 2020 research and innovation programme (grant agreement 947803) and under Marie Skłodowska–Curie (grant agreement 101061151). G.S.-F. is supported by the Austrian Academy of Sciences. Funding to pay the Open Access publication charges for this article was provided by the Austrian Science Fund (FWF).
Data availability
The next-generation sequencing data discussed in this publication have been deposited in NCBI’s Gene Expression Omnibus [86] and are accessible through GEO Series accession number GSE297401. Publicly available datasets used in this study are accessible at GSM1080936, GSM2570506, GSE169900, and GSM5688732. Scripts used in this study to analyze CUT&RUN, ChIPseq and mRNAseq datasets have been deposited on Github (https://github.com/seruggialab) and Zenodo (CUT&Run ChIPseq https://doi.org/10.5281/zenodo.20292380, mRNAseq https://doi.org/10.5281/zenodo.20292501, ChIPseq https://doi.org/10.5281/zenodo.20292410). The mass spectrometry proteomics data have been deposited to the ProteomeXchange Consortium via the PRIDE [87] partner repository with the dataset identifier PXD061867. Predicted structures of protein interactions identified in the ARID5B co-IP are available at https://doi.org/10.5281/zenodo.20292886. Uncropped western blot images have been deposited to Mendeley Data and can be accessed at https://doi.org/10.17632/xmvttmswpw.2. Microscopy data collected in this study can be obtained upon request by contacting the lead author (davide.seruggia@ccri.at).
References
- 1. Papaemmanuil E, Hosking FJ, Vijayakrishnan J et al. Loci on 7p12.2, 10q21.2 and 14q11.2 are associated with risk of childhood acute lymphoblastic leukemia. Nat Genet. 2009;41:1006–10. 10.1038/ng.430. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2. Treviño LR, Yang W, French D et al. Germline genomic variants associated with childhood acute lymphoblastic leukemia. Nat Genet. 2009;41:1001–5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3. Ellinghaus E, Stanulla M, Richter G et al. Identification of germline susceptibility loci in ETV6-RUNX1-rearranged childhood acute lymphoblastic leukemia. Leukemia. 2012;26:902–9. 10.1038/leu.2011.302. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4. Vijayakrishnan J, Qian M, Studd JB et al. Identification of four novel associations for B-cell acute lymphoblastic leukaemia risk. Nat Commun. 2019;10:1–9. 10.1038/s41467-019-13069-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5. Vijayakrishnan J, Kumar R, Henrion MYR et al. A genome-wide association study identifies risk loci for childhood acute lymphoblastic leukemia at 10q26.13 and 12q23.1. Leukemia. 2017;31:573–9. 10.1038/leu.2016.271. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6. Vijayakrishnan J, Studd J, Broderick P et al. Genome-wide association study identifies susceptibility loci for B-cell childhood acute lymphoblastic leukemia. Nat Commun. 2018;9:1–9. 10.1038/s41467-018-03178-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7. Clay-Gilmour AI, Hahn T, Preus LM et al. Genetic association with B-cell acute lymphoblastic leukemia in allogeneic transplant patients differs by age and sex. Blood Adv. 2017;1:1717–28. 10.1182/bloodadvances.2017006023. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8. Wiemels JL, Walsh KM, De Smith AJ et al. GWAS in childhood acute lymphoblastic leukemia reveals novel genetic associations at chromosomes. Nat Commun. 2018;9:1–8. 10.1038/s41467-017-02596-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9. Xu H, Yang W, Perez-Andreu V et al. Novel susceptibility variants at 10p12.31-12.2 for childhood acute lymphoblastic leukemia in ethnically diverse populations. J Natl Cancer Inst. 2013;105:733–42. 10.1093/jnci/djt042. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10. Migliorini G, Fiege B, Hosking FJ et al. Variation at 10p12.2 and 10p14 influences risk of childhood B-cell acute lymphoblastic leukemia and phenotype. Blood. 2013;122:3298–307. 10.1182/blood-2013-03-491316. [DOI] [PubMed] [Google Scholar]
- 11. Evans TJ, Milne E, Anderson D et al. Confirmation of childhood acute lymphoblastic leukemia Variants, ARID5B and IKZF1, and interaction with parental environmental exposures. PLoS One. 2014;9:e110255. 10.1371/journal.pone.0110255. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12. Perez-Andreu V, Roberts KG, Xu H et al. Lymphoid neoplasia: a genome-wide association study of susceptibility to acute lymphoblastic leukemia in adolescents and young adults. Blood. 2015;125:680–6. 10.1182/blood-2014-09-595744. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13. Archer NP, Perez-Andreu V, Stoltze U et al. Family-based exome-wide association study of childhood acute lymphoblastic leukemia among Hispanics confirms role of ARID5B in susceptibility. PLoS One. 2017;12:1–12. 10.1371/journal.pone.0180488. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14. Ragnarsson C, Yang M, Moura-Castro LH et al. Constitutional and acquired genetic variants in ARID5B in pediatric B-cell precursor acute lymphoblastic leukemia. Genes Chromosomes Cancer. 2024;63:1–7. 10.1002/gcc.23242. [DOI] [PubMed] [Google Scholar]
- 15. Studd JB, Vijayakrishnan J, Yang M et al. Genetic and regulatory mechanism of susceptibility to high-hyperdiploid acute lymphoblastic leukaemia at 10p21.2. Nat Commun. 2017;8:1–9. 10.1038/ncomms14616. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16. Chalise JP, Ehsani A, Lemecha M et al. ARID5B regulates fatty acid metabolism and proliferation at the Pre-B cell stage during B cell development. Front Immunol. 2023;14:1–14. 10.3389/fimmu.2023.1170475. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17. Goodings C, Zhao X, McKinney-Freeman S et al. ARID5B influences B-cell development and function in mouse. Haematologica. 2023;108:502–12. 10.3324/haematol.2022.281157. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18. Deák G, Cook AG. Missense variants reveal functional insights into the human ARID family of gene regulators. J Mol Biol. 2022;434:167529. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19. Claussnitzer M, Dankel SN, Kim K-H et al. FTO obesity variant circuitry and adipocyte browning in humans. N Engl J Med. 2015;373:895–907. 10.1056/NEJMoa1502214. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20. Baba A, Ohtake F, Okuno Y et al. PKA-dependent regulation of the histone lysine demethylase complex PHF2–ARID5B. Nat Cell Biol. 2011;13:668–75. 10.1038/ncb2228. [DOI] [PubMed] [Google Scholar]
- 21. Hata K, Takashima R, Amano K et al. Arid5b facilitates chondrogenesis by recruiting the histone demethylase Phf2 to Sox9-regulated genes. Nat Commun. 2013;4:1–11. 10.1038/ncomms3850. [DOI] [PubMed] [Google Scholar]
- 22. Seruggia D, Oti M, Tripathi P et al. TAF5L and TAF6L maintain self-renewal of embryonic stem cells via the MYC regulatory network. Mol Cell. 2019;74:1148–63.e7. 10.1016/j.molcel.2019.03.025. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23. Kim J, Cantor AB, Orkin SH et al. Use of in vivo biotinylation to study protein–protein and protein–DNA interactions in mouse embryonic stem cells. Nat Protoc. 2009;4:506–17. 10.1038/nprot.2009.23. [DOI] [PubMed] [Google Scholar]
- 24. Teo GC, Polasky DA, Yu F et al. Fast deisotoping algorithm and its implementation in the MSFragger search engine. J Proteome Res. 2021;20:498–505. 10.1021/acs.jproteome.0c00544. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25. Yu F, Haynes SE, Teo GC et al. Fast quantitative analysis of timsTOF PASEF Data with MSFragger and IonQuant. Mol Cell Proteomics. 2020;19:1575–85. 10.1074/mcp.TIR120.002048. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26. Kong AT, Leprevost FV, Avtonomov DM et al. MSFragger: ultrafast and comprehensive peptide identification in shotgun proteomics. Nat Methods. 2017;14:513–20. 10.1038/nmeth.4256. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27. Ritchie ME, Phipson B, Wu D et al. limma powers differential expression analyses for RNA-sequencing and microarray studies. Nucleic Acids Res. 2015;43:e47. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28. Toro N, Shrivastava A, Ragueneau E et al. The IntAct database : efficient access to fine-grained molecular interaction data. Nucleic Acids Res. 2022;50:648–53. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29. Stark C, Breitkreutz B, Reguly T et al. BioGRID : a general repository for interaction datasets. Nucleic Acids Res. 2006;34:535–9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30. Shannon P, Markiel A, Ozier O et al. Cytoscape: a Software Environment for Integrated Models of biomolecular interaction networks. Genome Res. 2003;13:2498–504. 10.1101/gr.1239303. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31. Stringer C, Wang T, Michaelos M et al. Cellpose: a generalist algorithm for cellular segmentation. Nat Methods. 2021;18:100–6. 10.1038/s41592-020-01018-x. [DOI] [PubMed] [Google Scholar]
- 32. Abramson J, Adler J, Dunger J et al. Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature. 2024;630:493–500. 10.1038/s41586-024-07487-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33. Case DA, Aktulga HM, Belfon K et al. AmberTools. J Chem Inf Model. 2023;63:6183–91. 10.1021/acs.jcim.3c01153. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34. Maier JA, Martinez C, Ksavajhala K et al. ff14SB: improving the accuracy of protein side chain and backbone parameters from ff99SB. J Chem Theory Comput. 2015;11:3696–713. 10.1021/acs.jctc.5b00255. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35. Jorgensen WL, Chandrasekhar J, Madura JD et al. Comparison of simple potential functions for simulating liquid water. J Chem Phys. 1983;79:926–35. 10.1063/1.445869. [DOI] [Google Scholar]
- 36. Ryckaert JP, Ciccotti G, Berendsen HJC. Numerical integration of the cartesian equations of motion of a system with constraints: molecular dynamics of n-alkanes. J Comput Phys. 1977;23:327–41. 10.1016/0021-9991(77)90098-5. [DOI] [Google Scholar]
- 37. Humphrey W, Dalke A, Schulten K. VMD: visual molecular dynamics. J Mol Graph. 1996;14:33–8. 10.1016/0263-7855(96)00018-5. [DOI] [PubMed] [Google Scholar]
- 38. Skene PJ, Henikoff JG, Henikoff S. Targeted in situ genome-wide profiling with high efficiency for low cell numbers. Nat Protoc. 2018;13:1006–19. 10.1038/nprot.2018.015. [DOI] [PubMed] [Google Scholar]
- 39. Batty P, Beneder H, Schätz C et al. Disruption of the SAGA CORE triggers collateral degradation of KAT2A. Nat Commun. 2026;17:3410. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40. Langmead B, Salzberg SL. Fast gapped-read alignment with Bowtie 2. Nat Methods. 2012;9:357–9. 10.1038/nmeth.1923. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41. Li H, Handsaker B, Wysoker A et al. The Sequence Alignment/Map format and SAMtools. Bioinformatics. 2009;25:2078–9. 10.1093/bioinformatics/btp352. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42. Nordin A, Zambanini G, Pagella P et al. The CUT&RUN suspect list of problematic regions of the genome. Genome Biol. 2023;24:1–18. 10.1186/s13059-023-03027-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43. Quinlan AR, Hall IM. BEDTools: a flexible suite of utilities for comparing genomic features. Bioinformatics. 2010;26:841–2. 10.1093/bioinformatics/btq033. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44. Ramírez F, Ryan DP, Grüning B et al. deepTools2: a next generation web server for deep-sequencing data analysis. Nucleic Acids Res. 2016;44:W160–5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45. Wickham H. ggplot2: Elegant Graphics for Data Analysis. New York: Springer-Verlag, 2016. [Google Scholar]
- 46. Meers MP, Tenenbaum D, Henikoff S. Peak calling by Sparse Enrichment Analysis for CUT&RUN chromatin profiling. Epigenetics Chromatin. 2019;12:1–11. 10.1186/s13072-019-0287-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47. Zhang Y, Liu T, Meyer CA et al. Model-based analysis of ChIP-Seq (MACS). Genome Biol. 2008;9:R137. 10.1186/gb-2008-9-9-r137. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48. Blanco E, González-Ramírez M, Di Croce L. Productive visualization of high-throughput sequencing data using the SeqCode open portable platform. Sci Rep. 2021;11:1–22. 10.1038/s41598-021-98889-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49. Ernst J, Kellis M. ChromHMM: automating chromatin-state discovery and characterization. Nat Methods. 2012;9:215–6. 10.1038/nmeth.1906. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50. Heinz S, Benner C, Spann N et al. Simple combinations of lineage-determining transcription factors prime cis-regulatory elements required for macrophage and B cell identities. Mol Cell. 2010;38:576–89. 10.1016/j.molcel.2010.05.004. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51. Zou Z, Ohta T, Miura F et al. ChIP-Atlas 2021 update: a data-mining suite for exploring epigenomic landscapes by fully integrating ChIP-seq, ATAC-seq and bisulfite-seq data. Nucleic Acids Res. 2022;50:W175–82. 10.1093/nar/gkac199. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 52. Ross-Innes CS, Stark R, Teschendorff AE et al. Differential oestrogen receptor binding is associated with clinical outcome in breast cancer. Nature. 2012;481:389–93. 10.1038/nature10730. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53. McLean CY, Bristor D, Hiller M et al. GREAT improves functional interpretation of cis-regulatory regions. Nat Biotechnol. 2010;28:495–501. 10.1038/nbt.1630. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54. Li H. Aligning sequence reads, clone sequences and assembly contigs with BWA-MEM. arXiv, https://arxiv.org/abs/2013:1303.3997, 26 May 2013, preprint: not peer reviewed.
- 55. Servant N, Varoquaux N, Lajoie BR et al. HiC-Pro: an optimized and flexible pipeline for Hi-C data processing. Genome Biol. 2015;16:1–11. 10.1186/s13059-015-0831-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 56. Bhattacharyya S, Chandra V, Vijayanand P et al. Identification of significant chromatin contacts from HiChIP data by FitHiChIP. Nat Commun. 2019;10:4221. 10.1038/s41467-019-11950-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 57. Dempster JM, Boyle I, Vazquez F et al. Chronos: a cell population dynamics model of CRISPR experiments that improves inference of gene fitness effects. Genome Biol. 2021;22:1–23. 10.1186/s13059-021-02540-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 58. Dobin A, Davis CA, Schlesinger F et al. STAR: ultrafast universal RNA-seq aligner. Bioinformatics. 2013;29:15–21. 10.1093/bioinformatics/bts635. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 59. Anders S, Pyl PT, Huber W. HTSeq-A Python framework to work with high-throughput sequencing data. Bioinformatics. 2015;31:166–9. 10.1093/bioinformatics/btu638. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60. Putri GH, Anders S, Pyl PT et al. Analysing high-throughput sequencing data in Python with HTSeq 2.0. Bioinformatics. 2022;38:2943–5. 10.1093/bioinformatics/btac166. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 61. Love MI, Huber W, Anders S. Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2. Genome Biol. 2014;15:550. 10.1186/s13059-014-0550-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 62. Subramanian A, Tamayo P, Mootha VK et al. Gene set enrichment analysis: a knowledge-based approach for interpreting genome-wide expression profiles. Proc Natl Acad Sci USA. 2005;102:15545–50. 10.1073/pnas.0506580102. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 63. Bantscheff M, Hopf C, Savitski MM et al. Chemoproteomics profiling of HDAC inhibitors reveals selective targeting of HDAC complexes. Nat Biotechnol. 2011;29:255–68. 10.1038/nbt.1759. [DOI] [PubMed] [Google Scholar]
- 64. Joshi P, Greco TM, Guise AJ et al. The functional interactome landscape of the human histone deacetylase family. Mol Syst Biol. 2013;9:1–21. 10.1038/msb.2013.26. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 65. Wang S, Fairall L, Pham TK et al. A potential histone-chaperone activity for the MIER1 histone deacetylase complex. Nucleic Acids Res. 2023;51:6006–19. 10.1093/nar/gkad294. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 66. Nde J, Kempf CG, Zimmermann RC et al. Integrative structural modeling of intrinsically disordered regions in a human HDAC2 chromatin remodeling complex. bioRxiv, 10.1101/2025.08.08.669391, 25 March 2026, preprint: not peer reviewed. [DOI] [Google Scholar]
- 67. Chen Y, Chen L, Wu X et al. Acute liver steatosis translationally controls the epigenetic regulator MIER1 to promote liver regeneration in a study with male mice. Nat Commun. 2023;14:1521. 10.1038/s41467-023-37247-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 68. Lakisic G, Lebreton A, Pourpre R et al. Role of the BAHD1 chromatin-repressive complex in placental development and regulation of steroid metabolism. PLoS Genet. 2016;12:1–26. 10.1371/journal.pgen.1005898. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 69. Dong Y, Hu H, Zhang X et al. Phosphorylation of PHF2 by AMPK releases the repressive H3K9me2 and inhibits cancer metastasis. Signal Transduct Target Ther. 2023;8:95. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 70. Wen H, Li J, Song T et al. Recognition of histone H3K4 trimethylation by the plant homeodomain of PHF2 modulates histone demethylation. J Biol Chem. 2010;285:9322–6. 10.1074/jbc.C109.097667. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 71. Inoue S, Mai A, Dyer MJS et al. Inhibition of histone deacetylase class I but not class II is critical for the sensitization of leukemic cells to tumor necrosis factor-related apoptosis-inducing ligand-induced apoptosis. Cancer Res. 2006;66:6785–92. 10.1158/0008-5472.CAN-05-4563. [DOI] [PubMed] [Google Scholar]
- 72. Inoue F, Ahituv N. Decoding enhancers using massively parallel reporter assays. Genomics. 2015;106:159–64. 10.1016/j.ygeno.2015.06.005. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 73. Yang X-J, Seto E. The Rpd3/Hda1 family of lysine deacetylases: from bacteria and yeast to mice and men. Nat Rev Mol Cell Biol. 2008;9:206–18. 10.1038/nrm2346. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 74. Rickert RC. New insights into pre-BCR and BCR signalling with relevance to B cell malignancies. Nat Rev Immunol. 2013;13:578–91. 10.1038/nri3487. [DOI] [PubMed] [Google Scholar]
- 75. Clark MR, Mandal M, Ochiai K et al. Orchestrating B cell lymphopoiesis through interplay of IL-7 receptor and pre-B cell receptor signalling. Nat Rev Immunol. 2014;14:69–80. 10.1038/nri3570. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 76. Tang H, Wang H, Lin Q et al. Loss of IP3 receptor–mediated Ca2+ release in mouse B cells results in abnormal B cell development and function. J Immunol. 2017;199:570–80. 10.4049/jimmunol.1700109. [DOI] [PubMed] [Google Scholar]
- 77. Millard CJ, Watson PJ, Celardo I et al. Class I HDACs share a common mechanism of regulation by inositol phosphates. Mol Cell. 2013;51:57–67. 10.1016/j.molcel.2013.05.020. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 78. Watson PJ, Fairall L, Santos GM et al. Structure of HDAC3 bound to co-repressor and inositol tetraphosphate. Nature. 2012;481:335–40. 10.1038/nature10728. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 79. Chalise JP, Hu Z, Li M et al. Identification of an alternative short ARID5B isoform associated with B-ALL survival. Biochem Biophys Res Commun. 2024;703:149659. 10.1016/j.bbrc.2024.149659. [DOI] [PubMed] [Google Scholar]
- 80. Loughran SJ, Comoglio F, Hamey FK et al. Mbd3/NuRD controls lymphoid cell fate and inhibits tumorigenesis by repressing a B cell transcriptional program. J Exp Med. 2017;214:3085–104. 10.1084/jem.20161827. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 81. Bagheri-Yarmand R, Balasenthil S, Gururaj AE et al. Metastasis-associated protein 1 transgenic mice: a new model of spontaneous B-cell lymphomas. Cancer Res. 2007;67:7062–7. 10.1158/0008-5472.CAN-07-0748. [DOI] [PubMed] [Google Scholar]
- 82. Kugler E, Madiwale S, Yong D et al. The NCOR-HDAC3 co-repressive complex modulates the leukemogenic potential of the transcription factor ERG. Nat Commun. 2023;14:5871. 10.1038/s41467-023-41067-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 83. Lu X, Chu CS, Fang T et al. MTA2/NuRD regulates B cell development and cooperates with OCA-B in controlling the Pre-B to immature B cell transition. Cell Rep. 2019;28:472–485.e5. 10.1016/j.celrep.2019.06.029. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 84. Yoshida T, Hu Y, Zhang Z et al. Chromatin restriction by the nucleosome remodeler Mi-2β and functional interplay with lineage-specific transcription regulators control B-cell differentiation. Genes Dev. 2019;33:763–81. 10.1101/gad.321901.118. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 85. Yamaguchi T, Cubizolles F, Zhang Y et al. Histone deacetylases 1 and 2 act in concert to promote the G1-to-S progression. Genes Dev. 2010;24:455–69. 10.1101/gad.552310. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 86. Edgar R, Domrachev M, Lash AE. Gene Expression Omnibus: NCBI gene expression and hybridization array data repository. Nucleic Acids Res. 2002;30:207–10. 10.1093/nar/30.1.207. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 87. Perez-Riverol Y, Bai J, Bandla C et al. The PRIDE database resources in 2022: a hub for mass spectrometry-based proteomics evidences. Nucleic Acids Res. 2022;50:D543–52. 10.1093/nar/gkab1038. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The next-generation sequencing data discussed in this publication have been deposited in NCBI’s Gene Expression Omnibus [86] and are accessible through GEO Series accession number GSE297401. Publicly available datasets used in this study are accessible at GSM1080936, GSM2570506, GSE169900, and GSM5688732. Scripts used in this study to analyze CUT&RUN, ChIPseq and mRNAseq datasets have been deposited on Github (https://github.com/seruggialab) and Zenodo (CUT&Run ChIPseq https://doi.org/10.5281/zenodo.20292380, mRNAseq https://doi.org/10.5281/zenodo.20292501, ChIPseq https://doi.org/10.5281/zenodo.20292410). The mass spectrometry proteomics data have been deposited to the ProteomeXchange Consortium via the PRIDE [87] partner repository with the dataset identifier PXD061867. Predicted structures of protein interactions identified in the ARID5B co-IP are available at https://doi.org/10.5281/zenodo.20292886. Uncropped western blot images have been deposited to Mendeley Data and can be accessed at https://doi.org/10.17632/xmvttmswpw.2. Microscopy data collected in this study can be obtained upon request by contacting the lead author (davide.seruggia@ccri.at).





