Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2025 Sep 26.
Published in final edited form as: Found Data Sci. 2025 Sep;7(3):737–758. doi: 10.3934/fods.2024048

PERSISTENT DIRECTED FLAG LAPLACIAN

Benjamin Jones 1,*, Guo-Wei Wei 1,2,3,*
PMCID: PMC12463187  NIHMSID: NIHMS2066548  PMID: 41019827

Abstract

Topological data analysis (TDA) has had enormous success in science and engineering in the past decade. Persistent topological Laplacians (PTLs) overcome some limitations of persistent homology, a key technique in TDA, and provide substantial insight to the behavior of various geometric and topological objects. This work extends PTLs to directed flag complexes, which are an exciting generalization to flag complexes, also known as clique complexes, that arise naturally in many situations. We introduce the directed flag Laplacian and show that the proposed persistent directed flag Laplacian (PDFL) is a distinct way of analyzing these flag complexes. Example calculations are provided to demonstrate the potential of the proposed PDFL in real world applications.

Key words and phrases. Persistent topological Laplacians, directed flag complex, directed flag Laplacians, clique complex, topological data analysis

2020 Mathematics Subject Classification. Primary: 55N31

1. Introduction.

Topological data analysis (TDA) is a rapidly growing field that uses tools from a multitude of disciplines to study connections between geometric, topological, directional, time-based, and functional information. TDA builds upon the fundamental objects of homology and cohomology through a filtration process, which gives rise to persistence modules. Persistence modules, including persistent (co)homology, are some of the most studied objects in TDA, and can convey information about a system in different dimensions, scales of size, and/or scales of time [37, 5]. They can be analyzed using persistent barcodes [13], persistent diagrams, persistent images [1], persistence landscapes [3], and persistent Betti numbers. These methods have been used to study protein folding [35], protein-ligand binding [4], virus mutation [8], chemistry [29], dynamical systems [25], and signal processing [24]. In fact, persistent homology owes its success to advanced machine learning models [17], particularly, topological deep learning introduced in 2017 [4].

However, persistent homology has many significant drawbacks that limit its utility in applications. For example, persistent homology cannot differentiate a five-member ring from a six-member ring, but such a difference is essential in molecular science. Additionally, persistent homology does not capture any non-topological changes which are crucial for networks. To overcome these drawbacks, Wei and colleagues introduced multiscale analysis into another fundamental object, topological Laplacians, including Hodge Laplacian and combinatorial Laplacians in 2019, resulting in evolutionary de Rham-Hodge theory, or persistent Hodge Laplacians [9], and persistent combinatorial Laplacians, also called persistent Laplacians [30]. Persistent Hodge Laplacians are built from differential forms on smooth manifolds [11, 9], while persistent combinatorial Laplacians are defined on simplicial complexes with point clouds [18, 10, 30]. They are referred to as persistent topological Laplacians (PTLs) due to their similarity in algebraic topology structure. The key characteristic of the topological Laplacian operators is that their harmonic spectra yield the Betti numbers that reveal the changes in topological invariants, while their non-harmonic spectra can convey additional information about the geometric variation [18]. By adding persistence, one can recover the persistent Betti numbers and capture the non-topological networking in the data. The development of PTLs has stimulated many new PTL formulations, including persistent sheaf Laplacians [34], persistent path Laplacians [31], persistent hyperdigraph Laplacians [6], and mathematical analysis of PTLs. Mémoli et al. [22] developed and analyzed algorithms for computing persistent topological Laplacians, and showed a related stability theorem for some of the eigenvalues involved. Persistent Laplacians have also been generalized to simplicial maps in Gülen et al. [16]. Recently, Liu et al. have developed an interleaving distance for studying persistent Laplacians and proved some stability theorems [19], establishing theoretical reliability of the results that persistent combinatorial Laplacians have seen in practice. Successful applications of persistent topological Laplacians have been reported in the literature, including protein-ligand binding [23] and the accurate prediction of SARS-CoV-2 emerging dominant variants [7]. More recently, Persistent Mayer homology and persistent Mayer Laplacian have been proposed to generalize the standard 2-chain complex into N-chain complex, leading to a new family of homology and Laplacian formulations [27]. The superiority of persistent Laplacians over classic persistent homology was documented in a large scale machine learning-guided protein engineering study using 34 benchmark datasets [26]. A software package was developed for computing persistent Laplacians [32]. Recently, quantum persistent homology in terms of Dirac [2], persistent Dirac [33], and persistent Dirac of path and hypergraph [28] have been proposed to achieve a similar performance for point cloud data.

Motivated by the practical need to develop high-performance TDA models, we construct a directed flag Laplacian and propose persistent topological Laplacian theory on a type of simplicial complex called the directed flag complex. Directed flag complexes are based on flag complexes, also commonly referred to as clique complexes, which are an abstract simplicial complex with an underlying undirected graph structure [20]. Flag complexes arise naturally in many problems, including computing the Vietoris-Rips complex [36], analyzing brain networks [20], and studying small world networks [21]. Expanding the flag complex to the case of directed graphs allows researchers to encode directional information, and one can further add context from persistence that can contain geometric or time-based information. Lügehetmann et al. developed an open-source software FLAGSER for the fast computation of persistent homology of directed flag complexes in 2020 [20]. Our implementation of the present persistent directed flag Laplacian employs the underlying directed flag complex computed by FLAGSER.

The rest of this paper is organized as follows. In Section 2 we provide context on the directed flag complex and relevant homology tools as in [20]. In Section 3, we first construct a directed flag Laplacian, and then develop a persistent directed flag Laplacian method, with some examples. The relationship of the present formalism and earlier persistent path Laplacian and persistent hyperdigraph Laplacian are analyzed in Section 4. In Section 5 we apply the proposed persistent directed flag Laplacian to real-world data for a protein-ligand complex. A conclusion is given in Section 6.

2. Background on directed flag complex homology.

To establish notations, we review the relevant graph theory and topology definitions of directed flag complexes used in [20]. First we define the graphs we want to consider and show how to build the directed flag complex, followed by the homology of the complex. Finally, we develop persistence.

2.1. Directed flag complexes.

Definition 2.1. A directed graph, or digraph, G is a pair (V,E) where V is a finite nonempty ordered set with elements called vertices and E is a finite nonempty set of ordered pairs of vertices e=vi,vj such that vivj, called edges.

Note that if vi,vjE, it is possible, but not necessarily true that vj,viE. Also note that this definition excludes self-loops of the form vi,vi from the possible edges.

For studying the topology of these graphs, we must define additional structures on them, so we introduce the notations around abstract simplicial complexes.

Definition 2.2. An abstract simplicial complex on a set X is a collection 𝒦 of nonempty finite subsets of X, called simplices, such that if σ𝒦, then every nonempty τσ is also in 𝒦. An ordered simplicial complex 𝒦 has simplices that are ordered nonempty finite subsets of X, although it is not necessary that X have an underlying order. If τσ are two simplices in a simplicial complex, we call τ a face of σ. Simplices with k+1 elements are called k-simplices.

We focus our attention on a particular abstract simplicial complex of digraphs, the directed flag complex. We must specify its simplices.

Definition 2.3. A k-clique of a directed graph G=(V,E) is an ordered subset of vertices σ=v0,v1,,vk-1 such that for i<j we have vi,vjE. The directed flag complex dFl(G) is the ordered simplicial complex on V with k-simplices the (k+1)-cliques of G.

By this definition, all vertices of the graph are 1-cliques and all edges are 2-cliques. Observe that by this definition, the edges and vertices entirely determine the set of (k+1)-cliques, and hence the entire complex. Although every directed flag complex is an ordered simplicial complex, the requirement that all (k+1)-cliques are in the directed flag complex means we can construct ordered simplicial complexes, which are not directed flag complexes, by removing cliques with more than 2 vertices.

There are many sources of directed flag complexes both in nature and in theory (the C-Elegans brain, Erdős Rényi graphs) [20]. In Example 1 we describe a few digraphs and their directed flag complexes, and in Example 2 we describe ordered simplicial complexes on the same graphs as those in Example 1, which are not directed flag complexes.

Example 1. Let G1,G2, and G3 be digraphs as in Figure 1. We visualize edges in the directed flag complex with a yellow background around the edge, and 3-cliques v0,v1,v2 by a wider, green background around the edges v0,v1 and v1,v2, omitting (v0,v2) to make visualization easier.

Figure 1.

Figure 1.

Digraphs G1,G2, and G3. The edges in the directed flag complex are visualized with a yellow background around the edge, and 3-cliques v0,v1,v2 by a wider, green background around the edges v0,v1 and v1,v2, omitting v0,v2 to make visualization easier.

The digraph G1 has vertices (1-cliques) V={a,b,c}, edges (2-cliques) E={(a,b),(a,c),(b,c)}, and the 3-clique (a,b,c). Then dFlG1 is the ordered simplicial complex with 0-simplices the vertices, 1-simplices the edges, and the 2-simplex (a,b,c).

The digraph G2 has 4 vertices V={a,b,c,d}, 6 edges E={(a,b),(a,c),(a,d)(b,c),(b,d),(c,d)}, 4 3–cliques {(a,b,d),(a,b,c),(a,c,d),(b,c,d)}, and the 4-clique (a,b,c,d). Note that this is the minimal example of a directed 4-clique. Then dFlG2 is the ordered simplicial complex with 0-simplices the vertices, 1-simplices the edges, 2-simplices the 4 3–cliques, and 3-simplex (a,b,c,d). If we let v0=a,v1=b,v2=c,v3=d, then for i<j,vi,vjE since (a,b),(a,c),(a,d),(b,c),(b,d), and (c,d)E.

The digraph G3 has 5 vertices {a,b,c,d,e}, 7 edges {(a,d),(b,a),(b,c),(c,e),(d,b),(d,c),(e,d)}, and 3-clique (d,b,c). Then dFlG3 is the ordered simplicial complex with 0-simplices the vertices, 1-simplices the edges, and 2-simplex (d,b,c). Note that neither (a,d,b) nor (c,e,d) are 3-cliques, since the edges (a,b) and (c,d) are not in G3.

Example 2. Let G1,G2, and G3 be digraphs as in Figure 1. Here we describe some ordered simplicial complexes that share a 1-skeleton with the directed flag complex on these graphs, but are distinct.

In G1, the only ordered simplicial complex with the same 1-skeleton as the directed flag complex consists of only the vertices {a,b,c} and edges {(a,b),(a,c),(b,c)}, with no 2-simplices.

In G2, we can have an ordered simplicial complex with vertices {a,b,c,d} and edges {(a,b),(a,c),(a,d),(b,c),(b,d),(c,d)}, with no 2-simplices. We can obtain additional ordered simplicial complexes by adding any number of the possible 2-simplices (a,b,d),(a,b,c),(a,c,d) or (b,c,d). The directed flag complex on G2 is the unique ordered simplicial complex containing the 3-simplex (a,b,c,d), since all ordered subsets of (a,b,c,d) must also be in the complex. By choosing 0, 1, 2, 3, or all 4 of the possible 2-simplices on G2, we can obtain 24=16 ordered simplicial complexes that are distinct from the directed flag complex.

If we further relax the condition that our ordered simplicial complex must share its 1-skeleton with the directed flag complex, we can obtain many more simplicial complexes on the vertex set, although none of them are faithful to the graph. For example, on the vertices of G2, we could have a simplicial complex with vertices {a,b,c,d}, edges {(a,b),(a,d),(d,b)}, and 2-simplex (a,d,b). Note that not all edges of G2 must be in the complex and that the edge (d,b) is not an edge of G2.

2.2. Directed flag complex homology.

We can define a boundary operator and chain complex on ordered simplicial complexes in the usual way.

Definition 2.4. Let F be a field and 𝒦 an ordered simplicial complex. For k0, let Ck(𝒦;F) be the vector space over F generated by k-simplices in 𝒦. The collection C(𝒦;F) forms a chain complex with the boundary operator dk:Ck(𝒦;F)Ck-1(𝒦;F) given by

dkx0,x1,,xk=i=0k(-1)ix0,x1,,xˆi,,xk,

where x0,x1,,xiˆ,,xk denotes the (k-1)-simplex obtained by removing xi from the k-simplex x0,x1,,xk. We also take d0:C0(𝒦;F)0 to be the 0 map. If dk(σ)=0, we call σ a cycle, and if σ=dk+1(τ), we call σ a boundary. Note that dkdk-1=0, which also implies all boundaries are cycles.

Since (C,d) forms a chain complex, the k-th homology of 𝒦 with coefficients in F is given by Hk(𝒦;F)=kerdk/Imdk+1. Let βk=dimHk(𝒦;F), which we call the k-th Betti number.

Note that in the theoretical treatment of persistent homology, the field F can be any field, but it is often taken to be Z/2Z or another finite field, as is done in [20], for which there are fast algorithms and implementations for computing persistence [12]. In this work we take F to be the real numbers R, because this allows us to compute the spectra of our Laplacian.

Example 3. Again consider the digraph G3 to be as in Figure 1. We computed the ordered simplicial complex dFlG3 in Example 1, which gives rise to the chain groups,

C0dFlG3;R=spana,b,c,d,e,C1dFlG3;R=span{(a,d),(b,a),(b,c),(c,e),(d,b),(d,c),(e,d)},C2dFlG3;R=spand,b,c.

The nonzero boundary operators can then be represented by the matrices

d1=(a,d)(b,a)(b,c)(c,e)(d,b)(d,c)(e,d)(a)(b)(c)(d)(e)(11000000110100001101010001110001001)
d2=(d,b,c)(a,d)(b,a)(b,c)(c,e)(d,b)(d,c)(e,d)(0010110).

We can then compute the kernels and images of the relevant boundary maps,

kerd0=span{(a),(b),(c),(d),(e)},
kerd1=span{(a,d)+(d,b)+(b,a),(d,b)(d,c)+(b,c),(c,e)+(e,d)+(d,c)},
Imd1=span{(a)(b),(c)(b),(d)(a),(e)(c)},
Imd2=span{(b,c)(d,c)+(d,b)}

Note that Imd0,kerd2,kerdi, and Imdi are all 0 for i3. From these, we compute the homology,

H0dFlG3;R=kerd0/Imd1R,H1dFlG3;R=kerd1/Imd2R2,H2dFlG3;R=kerd2/Imd30.

All other homologies are 0. Then the Betti numbers are β0=1,β1=2, and β2=0.

This example demonstrates a directed flag complex with a cycle in C1 that is a boundary and two that are not boundaries. It also demonstrates that as usual, β0 measure the number of connected components.

2.3. Persistent directed flag complex homology.

Now that we have established homology of directed flag complexes, we introduce the relevant details of persistence, including the persistent Betti numbers. Let (R,) be the category of real numbers with morphisms given by ab if ab. Let dFlag be the category of directed flag complexes, with morphisms given by inclusion. A filtration is a functor :(R,)dFlag, where (ab) is the inclusion ia,b:(a)(b). For clarity of notation, we denote Hka()=Hk((a),R) and the boundary map of this chain complex to be dka:HkaHk-1a. When the filtration is understood from context, we use Hka. Then for a given filtration and dimension k, we define the real (a,b)-persistent homology to be

Hka,b()=ImHkaHkb,

and the (a,b)-persistent Betti number of to be βka,b=dimHka,b(). Again, if the filtration is understood, we use Hka,b. We note that the coefficients do not need to be R in general, but for the development of the persistent Laplacian we will use R.

We similarly denote Cka()=Ck((a);R) or simply Cka if the filtration is clear from the context. The persistent Betti number βka,b corresponds to the number of nonzero classes in Hka that are nonzero in Hkb.

We now give an example where persistent directed flag complex homology is different from computing the non-persistent directed flag complex homology on a sequence of graphs.

Example 4. In Fig. 2 we show a filtration of graphs, where the persistent directed flag complex homology has β00,1=β01,2=1, while β00=β01=β02=2. This can be seen by the fact that only 1 connected component that existed in G0 also exists in G1, hence β00,1=1, despite there being 2 total connected components in G1.

Figure 2.

Figure 2.

Filtered directed flag complex with constant persistent Betti number β0a,a+1=1, while the non-persistent Betti number is constant at β0a=2. The area around edges are colored in yellow to signify that they form 1-simplices.

Alternatively, consider the induced map by inclusion i01,2:H01H02, where H01span{(0),(1)} and H02span{(0),(2)}. Then Imi01,2span{(0)}R, so β01,2=1. Similar computations apply to β02,3.

This example demonstrated that persistent Betti numbers provide a different perspective than calculating a sequence of non-persistent Betti numbers. Moreover, since the non-persistent Laplacian spectra recover the non-persistent Betti numbers, this motivates the need for the persistent Laplacian to recover the persistent Betti numbers.

3. Persistent directed flag Laplacian.

Persistent Betti numbers do not tell the full story of the persistence evolution. Another way to measure the changes with persistence is through persistent topological Laplacians like the persistent path Laplacian [31], topological hyperdigraph Laplacian[6], and persistent sheaf Laplacian [34]. In this work, we extend the persistent directed flag complex homology to a persistent directed flag Laplacian (PDFL). To do this, we construct the directed flag Laplacian and calculate its spectra. Then we construct the persistent version of the directed flag Laplacian and calculate its spectra.

3.1. Directed flag laplacian.

In this section, let G=(V,E) be a digraph, dFl(G) be the directed flag complex of G, and Sk be the set of k-cliques in G. Also write Ck=Ck(dFl(G);R). The Laplacian requires an adjoint of the boundary operator, so we define an inner product on the basis Sk of Ck,

,:Sn×SnR,

by

σ,τ=1ifσ=τ,0otherwise.

Next, this induces an inner product ·,· on the dual space Ck=Ck(dFl(G);R)=HomCk,R such that for any f,gCk,

f,g=σSkf(σ)g(σ).

With respect to these inner products, there is an adjoint dk*:Ck-1Ck of dk. Then we may define the k-th combinatorial directed flag Laplacian, or directed Flag Laplacian, by Δk:CkCk,

Δk=dk+1dk+1*+dk*dk,

where Δ0=d1d1* and if N is the largest integer for which SN is nonempty, ΔN=dN*dN. With respect to the standard basis of Ck, let Bk be the matrix representation of dk. In the unweighted, non-persistent case, as shown in [22], the matrix representation of dk* with respect to the standard basis of Ck-1 is BkT. In the computing the matrix representation of the adjoint in the persistent case, we will need more care as shown in [22]. The matrix representation of the k-th combinatorial directed flag Laplacian Δk:CkCk is

Lk=Bk+1Bk+1T+BkTBk.

Then Lk is positive semi-definite, with only real, non-negative eigenvalues, and is self-adjoint. Moreover, βk=dimkerΔk, which is the number of zero eigenvalues of Lk, as shown in [30].

We may extract more information from Lk than just the multiplicity of zero as an eigenvalue by examining the nonzero eigenvalues, called the non-harmonic spectra. If the dimension of Ck is n, denote the spectra of Δk in non-decreasing order by

SpectraLk=λ1k,λ2k,,λnk.

We now compute the Laplacians and spectra for the directed flag complex in Example 1, and show that the Betti numbers we previously computed agree with the harmonic spectra of the Laplacians.

Example 5. Using the graph and directed flag complex of G3 in Example 1, we compute the Laplacians and spectra. The nonzero boundary maps from before are,

B1=(a,d)(b,a)(b,c)(c,e)(d,b)(d,c)(e,d)(a)(b)(c)(d)(e)(11000000110100001101010001110001001)
B2=(d,b,c)(a,d)(b,a)(b,c)(c,e)(d,b)(d,c)(e,d)(0010110).

Then

L0=-11000000-1-10100001-10101000-1-11000100-1-100101-10000-110000-101010-10001-100001-1=2-10-10-13-1-100-13-1-1-1-1-14-100-1-12,L1=00101-1000101-10+-100101-10000-110000-101010-10001-100001-1-11000000-1-10100001-10101000-1-11000100-1=2-100-1-11-1210-100013-100000-120-1-1-1-10030-1-100-103-1100-1-1-12,L2=00101-1000101-10=(3).

From these matrices we can compute the spectra:

SpectraL0={0,3-2,3,3+2,5}SpectraL1={0,0,3-2,3,3,3+2,5}SpectraL2={3}.

Note that some eigenvalues are repeated. From these spectra, we can see that β0=1,β1=2, and β2=0, which agrees with our prior calculations.

3.2. Persistent directed flag laplacian.

We now move to the persistent case, where we have a filtration functor :(R,)dFlag. To do this, we repeat the process of construction as in Section 3.1 with the added structure of a filtration. For building the persistent directed flag Laplacian, we will work more directly with the chain groups than is necessary in formulating persistent homology as in Section 2.3. To work with the chain groups, we note that for k0, the inclusion of directed flag complexes induces morphisms on chains ika,b:CkaCkb, which also induces a morphism on homologies, ika,b:HkaHkb.

Let Ck+1a,b=σCk+1bdk+1bσCkaCkb and ιk+1a,b the inclusion Ck+1a,bCk+1b. Then Ck+1a,b can be described as the chains in Ck+1b with boundaries contained entirely in Cka. It would otherwise be possible that some element of Ck+1b has boundary intersecting CkbCka. Then we define the persistent boundary operator dk+1a,b=ika,b*dk+1bιk+1a,b. The following commutative diagram shows how the relevant maps fit into the context of the chain complexes.

3.2.

Define the k-th (a,b)-persistent directed flag Laplacian Δka,b:CkaCkb by

Δka,b=dk+1a,bdk+1a,b*+dka*dka.

In the literature, the terms dk+1a,bdk+1a,b* and dka*dka are sometimes called the “up” and “down” Laplacians, respectively. A detailed proof that the harmonic spectra of Δka,b correspond to the βka,b can be found in Theorem 2.7 of [22].

We need a matrix representation Bk+1a,b of dk+1a,b with respect to a basis of Ck+1a,b. The choice of basis matters, as discussed in [22], because the matrix representation of dk+1a,b* depends on the inner product structure of Ck+1a,b, which is an inner product subspace of Ck+1b, and Ck+1a,b does not necessarily have a basis that is a restriction of the standard basis of Ck+1a,b.

If Z is a matrix with columns that form a basis of Ck+1a,b in Ck+1b and Bk+1a,b is written with respect to that basis, then the matrix representation of dk+1a,b* is ZTZ-1Bk+1a,bT. If we take our basis to be orthonormal, then ZTZ is the identity matrix on Ck+1a,b, so the matrix representation of dk+1a,b* is simply the transpose Bk+1a,bT. Algorithm 3.1 of [22] provides an explicit procedure that can be used for finding Z. We slightly modify this procedure by transforming the basis Z into an orthonormal one. This gives a matrix representation of ιk+1a,b with respect to an orthonormal basis of Ck+1a,b and the standard basis of Ck+1b. The matrix representation Jka,b of ika,b is taken to be the identity on elements of Cka and 0 otherwise.

This gives us a matrix representation Bk+1a,b=Jka,bBk+1bZ of dk+1a,b and a representation Bk+1a,bT of the adjoint dk+1a,b*. Algorithm 3.1 of [22] uses a series of row and column operations that are equivalent to the matrix multiplication by Z and Jka,b. Replacing the computation of ZTZ in algorithm 3.1 of [22] with computing an orthonormal basis of Z is an asymptotically equal tradeoff in complexity. In practice, the slowest steps of computing the persistent directed flag Laplacian seem to be in building the complex and performing the column reduction over R for finding Bk+1a,b.

The current work focuses primarily on the knowledge gained from extending persistent direct flag homology to a Laplacian theory rather than the algorithmic efficiency. Because our goal is theoretical in nature, multiple avenues for improvement in speed have been left on the table. We note that Algorithm 4.1 of [22] provides an alternative method to compute the persistent Laplacian using Schur complements, which the authors explain may have a better time complexity in special cases, including that of clique (flag) complexes.

Then the matrix representation of Δka,b is

Lka,b=Bk+1a,bTBk+1a,b+BkaBkaT

and we denote its spectra in non-decreasing order as

SpectraLka,b=λ1ka,b,λ2ka,b,,λnka,b.

The smallest nonzero eigenvalue of Lka,b is often of importance, so we denote it λka,b.

In Fig. 3, we display a filtration of directed flag complexes on a graph with three vertices, and in Fig. 4, we show the persistent Betti numbers and λka,a+1. When the edge (0,2) is added in G5, the directed flag complex also gains a 2-cell, (0, 1, 2). There is a cycle created in C15,(1,2)-(0,2)+(0,1), but it is also the boundary of (0,1,2)C25. This cycle does not generate a homology class in the first homology group, so persistent Betti numbers cannot detect it, but λ1a,a+1 shows a change.

Figure 3.

Figure 3.

Filtration of a directed flag complex on a graph with three vertices. The area around edges are colored in yellow to signify that they form 1-simplices, and the area around the edges v0,v1 and v1,v2 is shaded green to highlight them as part of the 2-simplex v0,v1,v2.

Figure 4.

Figure 4.

Persistent spectra of the graph in Fig. 3. The values of βi5,- and λi5,- are βi5 and λi5, respectively, as filtered complex does not change after this point.

4. Discussion.

Here we place the directed flag Laplacian we have built in context with the path Laplacian described in [31], based on path homology as described in [14], and with the hyperdigraph Laplacian described in [6]. These other two topological Laplacians both generalize simplicial complexes, and hence generalize the directed flag complex, but in interesting ways.

4.1. Relationship with the path Laplacian.

Here we include some of the key components of the path Laplacian theory. The full details can be found in [31]. First, we define the terminology around paths. For a nonempty finite set V, with elements called vertices, and nonnegative integer p, an elementary p-path on V is any sequence i0ip of p+1 vertices. Let Λp be the vector space over a field K generated by the elementary p-paths, with elements called p-paths and the elements representing the elementary p-paths are denoted by ei0ip. The authors define the elementary (−1)–path to be the empty set, so the Λ-1K, and also define Λ-2={0}. A boundary operator defined in the usual way on Λ* produces a chain complex.

A path complex on V is a nonempty collection P of elementary paths such that

ifi0inP,theni0in-1Pandi1inP. (1)

The set of n-paths in P is denoted Pn.

The allowed n-paths are given by

𝓐n=spanKei0in:i0inPn,

which is a subspace of Λn. In order to form a good chain complex of allowed paths, one must further restrict to -invariant paths,

Ωn=v𝓐n:v𝓐n-1.

The path homology groups Hn(P) of the path complex P are given by the homology of the chain complex

ΩnΩn-1Ω00.

By taking K=R and defining an appropriate inner product, one can obtain the adjoint n* and construct a Laplacian theory for path complexes in a manner similar to what we have done with the directed flag complex. Denote the path Laplacian by ΔnP=n+1Pn+1P*+nP*nP, with corresponding matrix representation LnP=Bn+1PBn+1PT+BnPTBnp. It is pointed out in [22] that the transpose is not always the correct matrix representation for the adjoint of the persistent boundary matrix; the same is true here, even in the non-persistent case, because Ωn is an inner product subspace of 𝓐n with a basis which may not always be written as a subset of the basis for 𝓐n. As with the directed flag Laplacian, we will write BnP with respect to an orthonormal basis for Ωn, which resolves this issue.

It is noted in [31] and [15] that path complexes are designed to generalize abstract simplicial complexes, and hence any directed flag complex can be given an equivalent path complex. Note that Pn does not necessarily contain all elementary n-paths on V; it only needs to satisfy the conditions in Eq. 1. Although path complexes generalize abstract simplicial complexes, and hence the directed flag complex, the most natural path complex to use on a graph G will not necessarily agree with the directed flag complex on G.

For a digraph G=(V,E), it is most natural to consider the path complex P(G) consisting of all vertices, edges, and paths that can be built with those edges. Additional restrictions on the digraph are discussed in [31], but the present level of detail suffices to distinguish the directed flag Laplacian from the path Laplacian. This path complex will generally not agree with the directed flag complex on the same graph.

Consider the two digraphs in Fig. 5. The path Laplacians of these digraphs are computed in Table 1 and Table 2 of [31]. The calculation of Ω2 in G1 will require the orthonormal basis as discussed above, but the calculations for G1 that do not involve Ω2 are not impacted as the other matrices are written with respect to orthonormal bases. Calculations for G2 are also not impacted by this change.

Figure 5.

Figure 5.

Directed graphs G1 and G2.

We recalculate the path Laplacians and spectra for G1 in Table 1. The only difference between the spectra here and the non-normalized version in [31] is that here λ1P3=2, instead of 4, and λ2P0=2 instead of 4. Otherwise the spectra are the same, including the Betti numbers. We also record the calculations of the path Laplacians and spectra for G2 in Table 2, which are the same as in Table 2 of [31].

Table 1.

Path Laplacian and Spectra computation for graph G1 in Fig. 5.

n n=0 n=1 n=2
Ωn spane0,e1,e2,e3 spane01,e03,e12,e32 span12e032-12e012
Bn+1 e01e03e12e32e0e1e2e3(1100101000110101) 12e03212e012e01e03e12e32(12121212) 1 × 0 empty matrix
LnP 2-10-1-12-100-12-1-10-12 5212-12-121252-12-12-12-125212-12-121252 (2)
spectraLnP {0, 2, 2, 4} {2, 2, 2, 4} {2}
βn 1 0 0

Table 2.

Path Laplacian and Spectra computation for graph G2 in Fig. 5.

n n=0 n=1 n=2
Ωn spane0,e1,e2,e3 spane01,e03,e12,e32 {0}
Bn+1 e01e03e21e23e0e1e2e3(1100101000110101) 4 × 0 empty matrix 0 × 0 empty matrix
LnP 2-10-1-12-100-12-1-10-12 2110120110210112 0 × 0 empty matrix
spectraLnP {0, 2, 2, 4} {0, 2, 4, 4}
βn 1 1 0

Notice that the path Laplacian of G1 has β1=0 while G2 gives β1=1. Thus the path homology and path Laplacian can distinguish these graphs.

Now we compute the directed flag Laplacian for these graphs. The key difference between the path complexes and the directed flag complexes on these graphs is that 2-paths in a path complex do not correspond to 2-simplices (3-cliques) in a directed flag complex. In G1, the 2-paths e012 and e032 generate 𝓐2, and e012-e032 generates Ω2. However, in the directed flag complex of G2, there are no 3-cliques. We compute the directed flag Laplacian of G1 in Table 3 and of G2 in Table 4. Note that C2=0 for both. The resulting spectra of G1 and G2 are identical, although L1 has slight differences between the two. The directed flag Laplacian spectra alone do not detect the difference between these two digraphs.

Table 3.

Directed flag Laplacian calculations for graph G1 in Fig. 5.

n n=0 n=1
Cn span1,2,3,4 span1,2,1,4,2,3,4,3
Bn+1 (1,2)(1,4)(2,3)(4,3)(1)(2)(3)(4)(1100101000110101) 4 × 0 empty matrix
Ln 2-10-1-12-100-12-1-10-12 21-10120-1-10210-112
spectraLnP {0, 2, 2, 4} {0, 2, 2, 4}
βn 1 1

Table 4.

Directed flag Laplacian calculations for graph G2 in Fig. 5.

n n=0 n=1
Cn span1,2,3,4 span1,2,1,4,3,2,3,4
Bn+1 (1,2)(1,4)(3,2)(3,4)(1)(2)(3)(4)(1100101000110101) 4 × 0 empty matrix
Ln 2-10-1-12-100-12-1-10-12 2110120110210112
spectraLn {0, 2, 2, 4} {0, 2, 2, 4}
βn 1 1

However, if we consider the eigenvectors corresponding to these eigenvalues, we are able to distinguish the two digraphs. In Table 5 we show a linearly independent set of unit eigenvectors corresponding to the eigenvalues of L1j, the L1 matrix for the graphs Gj, where j=1,2. Ordering the eigenvalues as λ0=0,λ1=2,λ2=2, and λ3=4, we label the eigenvectors vij of L1 for the graph Gj corresponding to the eigenvalue λi, for j=1,2 and i=0,1,2,3. Critically, the eigenvectors vi1 are either not eigenvectors of L12 or, if they are, they may correspond to different eigenvalues. The same is true of the eigenvectors of L12. For example, v01 corresponds to λ0=0, but L12v01=2v01, i.e. it is an eigenvector with eigenvalue 2, not 0. This implies that two Laplacians have distinct kernels, which detects the difference between the graphs. We can also compute that v11,v21 are not eigenvectors of L12 and v12,v22 are not eigenvectors of L11. This is another way to distinguish the two graphs.

Table 5.

Linearly independent unit eigenvectors of the directed flag Laplacian L1 calculated in Tables 3 and 4 for graphs G1 and G2 of Fig. 5 corresponding to the eigenvalues λ{0,2,2,4}. Eigenvectors vij are from the graph Gj.

i 0 1 2 3
λi 0 2 2 4
vi1 121-11-1 16131316 -131616-13 12-1-111
vi2 12-111-1 -16-131316 13-1616-13 121111

We have shown in Tables 1, 2, 3, and 4 that the most natural path complex, the one composed of all paths in a directed graph G2, may have different spectra than the directed flag complex on the same directed graph. We can also see that they may agree, as in G1. Therefore, these are two distinct Laplacian theories for directed graphs, and one may be more appropriate for analyzing a given situation than the other. Considering only the eigenvalues, it may appear that the path Laplacian can sometimes distinguish graphs that the directed flag Laplacian cannot, but analyzing the eigenvectors of the directed flag Laplacian may provide enough insight to still detect the difference between the graphs. Their advantages in specific applications are to be further studied in the future.

4.2. Relationship with the hyperdigraph Laplacian.

It is also interesting to analyze the relation of the proposed method with the persistent hyperdigraph Laplacian [6]. We now describe the general idea of the hyperdigraph Laplacian as given in [6], so that one can better understand the proposed directed flag Laplacian.

Set V to be a nonempty finite ordered set. A directed p-hyperedge is a sequence v0v1vp of distinct elements in V. A hyperdigraph is a pair =(V,E), where E is a set of hyperedges. With appropriate restrictions, one can define a homology theory on a subset of the span of p-hyperedges, which is called the hyperdigraph homology Hp(;R). By taking coefficients in R, one can define an inner product, obtain a dual to the boundary operator, and craft a Laplacian Δp for a hyperdigraph . It has desirable properties similar to other topological Laplacians, e.g., it is self-adjoint, positive semidefinite, and kerΔpHp(;R), so this Laplacian also recovers the Betti numbers of the homology theory.

In Fig. 6, we show multiple hyperdigraphs with the same underlying digraph G. The digraph G has vertices (0), (1), and (2), edges (0, 1), (1, 2), and (0, 2). Then in 1 we add the 2-hyperedge (012), in 2 we add (120), and in 3 we add (201). Hyperdigraphs allow us to add these arbitrary 2-hyperedges, while a directed flag complex would only have the 2-clique (0, 1, 2). Note that 1 is a hyperdigraph corresponding to the directed flag complex on G. This puts the extreme flexibility of the hyperdigraph on display, because adding any of the possible 2-hyperedges on the three vertices produces a valid hyperdigraph.

Figure 6.

Figure 6.

A digraph G with three vertices and edges, and the three possible hyperdigraphs 1,2 and, 3 that have this underlying graph, but with the added hyperedges (012),(120), and (201), respectively.

Hyperdigraphs generalize abstract simplicial complexes on a finite set, so any directed flag complex can be formulated as an equivalent hyperdigraph. A directed flag complex is entirely determined by the directed edges and vertices. Unlike with path complexes, the extreme flexibility of hyperdigraphs means there is no natural hyperdigraph on a vertex set that is solely determined by vertices and directed edges. This flexibility is one of the core benefits of hyperdigraphs, since its relations could depend on higher dimensional information than pairwise connections. One could assign a p-hyperedge for each path of length p with distinct vertices, but this would become a path complex on the directed graph.

The additional structure enforced on the directed flag complex means that it tells us about how highly connected the graph is. Because a simplex is only added when a directed clique is present, this may reduce the number of higher dimensional simplices one uses in studying the dataset. We suspect this is one of the reasons the present work is able to continue the filtration in Fig. 8 to a cutoff of 8Å, whereas the approach in Ref. [6] could only reach up to 4 Å. The other possible reason is due to implementation factors, such as our use of FLAGSER to build the underlying complex, which is very fast with a mostly C++ implementation. A more comprehensive comparison of these two formalisms is needed but is out of the scope of the present work. In particular, there is a pressing need to develop more efficient computational algorithms for PTLs.

Figure 8.

Figure 8.

Spectra of protein-ligand complex with PDB ID 1a99, with restrictions on included atoms and filtered edges. Note that β2a,b uses a larger vertical axis scale than β0a,b and β1a,b.

5. Application.

In this section, we demonstrate the PDFL by analyzing protein-ligand complex 1a99 from the Protein Data Bank (PDB) as shown in Fig. 7 for visualizing the complex. The same complex was studied in the literature [6]. To analyze the protein-ligand binding, we consider only C atoms of the protein within 8 Å of the ligand, and only C, N, O, and S atoms of the ligand. We use the distance between two atoms in this restricted complex as the filtration value at which an edge is added between the two atoms. We round these distances to the nearest 10−3. We add edges for the bonds within the ligand, but not for bonds within the protein. The direction of the edge is pointing toward the atom with the higher electronegativity of the pair; in the case of a C-C pair, we add both possible edges. We then compute λka,b and βka,b, where k=0,1,2, and a is the filtration parameter, and b is the next filtration parameter greater than a. We used a cutoff value of 8Å angstrom because it was sufficiently large to allow the non-harmonic spectra to stabilize, and for a clear pattern to emerge in the harmonic spectra, but the implementation would allow for reasonable computation time to determine the spectra associated with larger cutoffs in different systems.

Figure 7.

Figure 7.

Protein-Ligand complex with PDBID 1a99.

The computed spectra are in Fig. 8. We can make several observations about this system. First, observe that β0a,b tends to 1, since the final graph is fully connected. Next, observe that β2a,b increases essentially monotonically; this is because the particular graph we produced does not have 3-simplices. It is challenging to have any 3-simplex because we do not add edges for within-protein atoms, and so there can only be 1 protein atom in any clique. To make a 3-simplex, the lone protein atom would need to have edges with 3 ligand atoms, that are all “bonded” to each other. This occurs rarely in nature.

Another notable feature of the spectra is that all are essentially constant until the filtration value is 3. From there on, β0a,b decreases steadily, which reflects that we added edges at every filtration value, and it is not common for many edges to have the same filtration value, which comes from the distance between atoms, rounded to the nearest 10−3. We can see that β1a,b rises rapidly from a filtration of 3 to 5, and then sharply decreases at 5 before continuing to increase, with some other perturbations. We can see that λ0a,b,λ1a,b, and λ2a,b all change suddenly for filtration values around 3 and 5, with a sharp change at 5, but what is interesting is that they stabilize thereafter even as β1a,b continues to change. This suggests that there are few significant interactions between the ligand molecule and protein atoms further than 5Å.

6. Conclusion.

Recently, persistent topological Laplacians (PTLs) have been proposed to overcome many drawbacks of classical persistent homology in topological data analysis (TDA), giving rise to a powerful new way to study the filtrations of simplicial complexes and other structures. One type of abstract simplicial complexes, directed flag complexes, commonly arises in studies of networks, including small world networks, neuron networks, and gene regulation networks. Therefore, a theory of persistent directed flag Laplacians (PDFL) may provide additional insight into challenging problems. This work builds such a theory and shows that it can provide a more nuanced representation of persistent changes in directed flag complexes than persistent Betti numbers alone can. Comparison is given to earlier persistent path Laplacians and persistent hyperdigraph Laplacians. We also demonstrate how PDFL can help analyze networks arising from real-world data, including a protein-ligand system. These examples illustrate the application potential of proposed PDFL to practice problems.

Acknowledgments.

This work was supported in part by NIH grants R01GM126189, R01AI164266, and R35GM148196, National Science Foundation grants DMS2052983, DMS-1761320, and IIS-1900473, NASA grant 80NSSC21M0023, Michigan State University Research Foundation, and Bristol-Myers Squibb 65109. The authors thank Dong Chen for technical assistance and Xiaoqi Wei for proof-reading.

Footnotes

Code availability. The codes are available at https://github.com/bdjones13/flagser-laplacian/.

REFERENCES

  • [1].Adams H, Emerson T, Kirby M, Neville R, Peterson C, Shipman P, Chepushtanova S, Hanson E, Motta F and Ziegelmeier L, Persistence images: A stable vector representation of persistent homology, Journal of Machine Learning Research, 18 (2017), 1–35. [Google Scholar]
  • [2].Ameneyro B, Maroulas V and Siopsis G, Quantum persistent homology, Journal of Applied and Computational Topology, 8 (2024), 1961–1980. [Google Scholar]
  • [3].Bubenik P, Statistical topological data analysis using persistence landscapes, Journal of Machine Learning Research, 16 (2015), 77–102. [Google Scholar]
  • [4].Cang Z and Wei G-W, TopologyNet: Topology based deep convolutional and multi-task neural networks for biomolecular property predictions, PLoS Computational Biology, 13 (2017),:e1005690. [Google Scholar]
  • [5].Carlsson G, Topology and data, Bulletin of the American Mathematical Society, 46 (2009), 255–308. [Google Scholar]
  • [6].Chen D, Liu J, Wu J and Wei G-W, Persistent hyperdigraph homology and persistent hyperdigraph Laplacians, Foundations of Data Science, 5 (2023), 558–588. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [7].Chen J, Qiu Y, Wang R and Wei G-W, Persistent Laplacian projected Omicron BA. 4 and BA. 5 to become new dominating variants, Computers in Biology and Medicine, 151 (2022), 106262. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [8].Chen J, Wang R, Wang M and Wei G-W, Mutations strengthened SARS-CoV-2 infectivity, Journal of Molecular Biology, 432 (2020), 5212–5226. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [9].Chen J, Zhao R, Tong Y and Wei G-W, Evolutionary de Rham-Hodge method, Discrete and Continuous Dynamical Systems - B, 26 (2021), 3785–3821. [Google Scholar]
  • [10].Chung FR and Langlands RP, A combinatorial Laplacian with vertex weights, Journal of Combinatorial Theory, Series A, 75 (1996), 316–327. [Google Scholar]
  • [11].Desbrun M, Hirani AN, Leok M and Marsden JE, Discrete exterior calculus, preprint, 2005, arXiv:math/0508341. [Google Scholar]
  • [12].Edelsbrunner H, Letscher D and Zomorodian A, Topological persistence and simplification, Discrete and Computational Geometry, 28 (2002), 511–533. [Google Scholar]
  • [13].Ghrist R, Barcodes: The persistent topology of data, Bulletin of the American Mathematical Society, 45 (2008), 61–75. [Google Scholar]
  • [14].Grigor’yan A, Jimenez R, Muranov Y and Yau S-T, Homology of path complexes and hypergraphs, Topology and its Applications, 267 (2019), 106877. [Google Scholar]
  • [15].Grigor’yan AA, Lin Y, Muranov Yu. V. and Yau S-T, Path Complexes and their Homologies, Journal of Mathematical Sciences, 248 (2020), 564–599. [Google Scholar]
  • [16].Gülen AB, Mémoli F, Wan Z and Wang Y, A Generalization of the persistent laplacian to simplicial maps, In Chambers EW and Gudmundsson J, editors, 39th International Symposium on Computational Geometry (SoCG 2023). Leibniz International Proceedings in Informatics (LIPIcs), Schloss Dagstuhl - Leibniz-Zentrum für Informatik, Dagstuhl, Germany, 258 (2023), 1–17. [Google Scholar]
  • [17].Hensel F, Moor M and Rieck B, A survey of topological machine learning methods, Frontiers in Artificial Intelligence, 4 (2021). [Google Scholar]
  • [18].Horak D and Jost J, Spectra of combinatorial Laplace operators on simplicial complexes, Advances in Mathematics, 244 (2013), 303–336. [Google Scholar]
  • [19].Liu J, Li J and Wu J, The algebraic stability for persistent Laplacians, preprint, 2023, arXiv:2302.03902. [Google Scholar]
  • [20].Lütgehetmann D, Govc D, Smith JP and Levi R, Computing persistent homology of directed flag complexes, Algorithms, 13(2020), 19. [Google Scholar]
  • [21].Masulli P and Villa AEP, The topology of the directed clique complex as a network invariant, SpringerPlus, 5 (2016), 388. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [22].Mémoli F, Wan Z and Wang Y, Persistent laplacians: Properties, algorithms and implications, SIAM Journal on Mathematics of Data Science, 4 (2022), 858–884. [Google Scholar]
  • [23].Meng Z and Xia Kelin, Persistent spectral-based machine learning (PerSpect ML) for protein-ligand binding affinity prediction, Science Advances, 7 (2021), eabc5329. [Google Scholar]
  • [24].Myers AD, Yesilli M, Tymochko S, Khasawneh F and Munch E, Teaspoon: A comprehensive python package for topological signal processing, TDA and Beyond, 2020. [Google Scholar]
  • [25].Perea JA, Topological time series analysis, Notices of the American Mathematical Society, 66 (2019), 686–694. [Google Scholar]
  • [26].Qiu Y and Wei G-W, Persistent spectral theory-guided protein engineering, Nature Computational Science, 3 (2023), 149–163. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [27].Shen L, Liu J and Wei G-W, Persistent Mayer homology and persistent Mayer Laplacian, Foundations of Data Science, 6 (2024), 584–612. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [28].Suwayyid F and Wei G-W, Persistent Dirac of paths on digraphs and hypergraphs, Foundations of Data Science, 6 (2024), 124–153. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [29].Townsend J, Micucci CP, Hymel JH, Maroulas V and Vogiatzis KD, Representation of molecular structures with persistent homology for machine learning applications in chemistry, Nature Communications, 11 (2020), 3230. [Google Scholar]
  • [30].Wang R, Nguyen DD and Wei G-W, Persistent spectral graph, Int. J. Numer. Methods Biomed. Eng, 36 (2020), 27 pp. [Google Scholar]
  • [31].Wang R and Wei G-W, Persistent path Laplacian, Foundations of Data Science, 5 (2023), 26–55. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [32].Wang R, Zhao R, Ribando-Gros E, Chen J, Tong Y and Wei G-W, HERMES: Persistent spectral graph software, Foundations of Data Science, 3(2021), 67–97. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [33].Wee J, Bianconi G and Xia K, Persistent Dirac for molecular representation, Scientific Reports, 13 (2023), 11183. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [34].Wei X and Wei G-W, Persistent sheaf Laplacians, Foundations of Data Science, 5 (2024), 26–55. [Google Scholar]
  • [35].Xia K and Wei G-W, Persistent homology analysis of protein structure, flexibility and folding, International Journal for Numerical Methods in Biomedical Engineering, 30 (2014), 814–844. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • [36].Zomorodian A, Fast construction of the Vietoris-Rips complex, Computers and Graphics, 34 (2010), 263–271. [Google Scholar]
  • [37].Zomorodian A and Carlsson G, Computing persistent homology, Proceedings of the twentieth annual symposium on Computational Geometry, 33 (2005), 249–274. [Google Scholar]

RESOURCES