Skip to main content
Wiley Open Access Collection logoLink to Wiley Open Access Collection
. 2026 Mar 19;42(4):374–400. doi: 10.1111/cla.70033

Analysis of alternative homologies with parsimony and step‐matrix recoding

Pablo A Goloboff 1,2,
PMCID: PMC13398255  PMID: 41858018

Abstract

This paper describes how prior uncertainty of the homology of different parts of organisms, and therefore of the characters contained in those parts, can be analysed by recoding the set of characters into a complex character. In this complex character, states represent combinations of states and homologies among the original characters, and transformation costs between states are set to values that maximize (weighted) homology. This allows a proper treatment of both the alternative homologies and the inapplicability of the characters contained in parts that can be absent or present. All the methods described are implemented in the program TNT, with a relatively simple syntax. This facilitates the treatment of characters in body parts of uncertain homology, including parts that can appear or disappear making their characters inapplicable. The analyses can also be combined with definitions of other types of dependences (e.g. morphofunctional constraints, inapplicabilities determined from characters outside of the parts of ambiguous homology). Although recoding into combinations of states and homologies allows handling relatively low numbers of characters per part (because of computation time and RAM requirements), TNT also implements options that facilitate creating alternative static alignments, allowing approximate analysis of larger number of characters.


In phylogenetic analysis of morphological data, a frequent problem is that the homology between body parts—and therefore of the characters contained in those parts—is unclear. A famous case is that of vertebrate digits—originally in the number of five, many groups undergo reduction of the digits, but different digits (e.g. digits IV and V, or I and V) can be lost (Shubin et al., 1997), and in some species it is unclear (from the study of anatomy alone) which digits are missing (Xu et al., 2009). This is more than a problem of just naming the parts, because the different digits may have different character states (claws, scaling, curvature, etc.). Therefore, in a species with reduced number of digits, comparing the characters of its digits with those of digits I‐II‐III, or of II‐III‐IV, may have different implications.

Among the first to discuss this problem in the context of phylogenetics was Endress (1994:401–402), a botanist, who pointed out without further formalization that floral parts of uncertain homology can be analysed by successively considering their alternative homology interpretations. Rieppel (1996) discussed alternative homology interpretations of elements of the pectoral girdle of turtles in the light of parsimony; Lee (1998) provided further analysis of the same problem. De Laet (2005) briefly discussed the problem of morphological characters of dubious homology, and proposed that “a hard work‐around would be to manually construct and analyse as many data sets as there are different combinations of different interpretations in different characters, which may be practically feasible when the number of such combinations is not too large.” The number of alternative datasets can grow very rapidly indeed, so simplifications are required in practice, for example, assuming beforehand that the parts in some taxa share the same homology. Ramírez (2007) discussed the problem of ambiguous homologies in more depth in a study of anyphaenid spiders, using different datasets to evaluate different hypotheses of homology of various sclerites in the male palpal bulb. He used TNT scripts to produce the alternative datasets, with eight blocks of taxa having similar sets of homologies, to keep the total number of alignments manageable. He analysed the datasets with standard methods (i.e. Fitch/Farris parsimony), but that still mistreats the inapplicable characters that result from the parts that bear them appearing or disappearing because Farris (1970) or Fitch (1971) optimization cannot maximize homology in sets of codependent characters, but no other option was available at the time. The difficulty of the treatment with multiple datasets is evident from the complexity of Ramírez's scripts (16 interacting scripts, totalling about 1800 lines of instructions, producing the 6561 possible alignments for the 8 blocks of taxa)—yet the script was hardcoded specifically for Ramírez's master dataset, and could not be used for other datasets unless substantially modified. Three other papers used less (from the computational point of view) exhaustive analyses and created the alternative datasets manually. Xu et al. (2009) assessed the homology of avian digits, constructing 14 datasets for alternative hypotheses of digit homology, Leardi et al. (2012) used 3 datasets representing different hypotheses of homology of the antorbital fenestra in thalattosuchian crocodyles, and Candela and Rasia (2012) evaluated crown molar structures of echimyid rodents with two datasets embodying possible homology correspondences. Assuming a specific evolutionary model instead of parsimony, King and Rücklin (2020) also used multiple datasets for a Bayesian approach, where the alternative alignments of characters were regularly switched (as “proposals”) in a Markov chain, thus calculating the “posterior probability” that one or other alignment would be correct. To my knowledge, these are the only formal analyses explicitly using the approach with multiple datasets. The only other paper I know of with a quantitative treatment of alternative morphological homologies is that of Agolin and D'Haese (2009), who examined series of setae in poduromorph collembolans, but in that case the ambiguity was (much as in DNA sequences) between the setae themselves instead of between the parts containing them. Their specific problem being rather similar to that of DNA sequence analysis, Agolin and D'Haese (2009) could use POY (a beta version of Varón et al., 2010) to analyse their dataset (but see De Laet, 2010 for some technical problems with their analysis).

The related problems of alternative homologies and inapplicable characters can both be solved by coding legitimate combinations of states for all the interrelated characters (a “complex” of characters) as morphotypes or “states” of a “supercharacter” (as proposed by e.g. Maddison, 1993, Goldman and Yang, 1994, or Wheeler, 1999, 2003), and defining appropriate costs for the transformations between the combinations. This recoding amounts to evaluating complete ancestral reconstructions (complete, at least, for the characters involved in the complex). Goloboff et al. (2021) and Goloboff and De Laet (2024) showed that the transformation costs between morphotypes can be set to reflect exactly the homoplasy in the original characters, including characters that become applicable when parts originate separately in distant parts of the tree. Recoding into morphotypes can thus be used for the analysis of inapplicable characters (Goloboff et al., 2021), as well as for characters that interact on the basis of rules specified by the user via a TNT command with a relatively simple but flexible syntax, xlinks (introduced by Goloboff and De Laet, 2024).

While the analysis of characters in parts of dubious homology has some similarities to sequence alignment, there are also some obvious differences that generally make application of standard methods for sequence analysis difficult. The most obvious one is that groups of characters belonging to the same part should be treated, together, as if they are all equivalent to the corresponding characters of one of the alternative homologous parts, or the characters of the other alternative part—moving in tandem as in a sort of translocation in sequence analysis, but not quite the same. It seems obvious that a proper recoding into morphotypes should also allow dealing with the problem of alternative homologies, but the existing xlinks implementation of Goloboff and De Laet (2024) is insufficient for this problem. This is exemplified by a recent paper (Grams et al., 2025) on malacostracan phylogeny, which extensively used xlinks , but also identified the problem of alternative homologies in some of their characters (stating that, e.g. exopod and endopod “are hierarchically higher characters to … another seven ontologically dependent characters,” and then in Bathynella “an identification of the single ramus as either endo‐ or exopod affects more than just the two character‐scorings…”). Grams et al. (2025) used xlinks to deal with inapplicabilities in their characters, but could not apply it to the problem of ambiguous homologies.

Thus, this paper presents the first implementation of a general solution to the problem of ambiguities in homology of parts and their characters. It solves the practical problems inherent to recoding alternative datasets (as previously proposed by Rieppel, 1996, De Laet, 2005, and Ramírez, 2007), taking into account that different instances of transformation between parts (in different regions of the tree) may correspond to different homologies (i.e. a different part being transformed)—which is difficult to handle with alternative datasets, unless those are multiplied enormously—as well as the problem of the treatment of inapplicable characters that result from the loss or gain of parts. The present methods can be applied relatively easily to almost any dataset and problem (unless the problem is too big, see below), defining the problem with simple enhancements of the syntax of the xlinks command. All the example datasets and scripts can be found as supplementary material at https://www.lillo.org.ar/phylogeny/published/Goloboff_2026_Alternative_Homologies.zip. The basic layout for the current implementation is different parts that could be homologous to each other, each of which bears a number of characters—thus the homology of the characters contained in the parts is subject to evaluation together with their containing parts, seeking the solution that homologizes the maximum number of parts and character states (for the given tree). As the number of possible combinations of states in the different characters can grow rapidly when many parts and characters are combined in a character complex, the main practical limitation of this method is in the number of parts and characters that can be analysed together. Another limitation is that, when some characters are inappplicable, due to the asymmetry in the step‐matrix for the resulting supercharacter and the need to force the condition in the root to that in the outgroup (i.e. the first‐splitting taxon), if the main character (the one that determines applicability of its subordinate characters) has not been observed for the outgroup, the homology calculations with this method are approximate (when all the transformation costs in the original characters are symmetric and the homoplasy is calculated as the difference between the observed and the minimum possible steps, if the root states are selected from among those in the first‐splitting taxon, the amount of homoplasy required by a tree does not depend on the location of the root, so that selecting as first‐splitting taxon a different terminal may often remedy this problem; see Goloboff et al., 2021, pp. 617–618, and Goloboff and De Laet, 2024, footnote 5). Thus, the recoding of inapplicable characters can guarantee a maximization of homology only when the outgroup (or root) taxon is known for the parts or main characters (as mentioned above). It is possible in theory to calculate the homoplasy for each of the alternative state assignments forced to the root (by calculating the difference between actual and minimum possible numbers of steps for each state forced at the root), using the lowest of those two values. This would produce the correct values of homoplasy, but has not been implemented in TNT.

Concept of parsimony

Before proceeding further, a brief clarification of the rationale for the approach used here is necessary. Sometimes, “parsimony” is equated with specific algorithms (e.g. Ebach et al., 2013; Williams and Gill, 2025), and sometimes with specific mechanics (“the minimum number of mutations along the tree to produce the given data,” e.g. Holmes, 2005). The first is plainly incorrect, and the second incomplete; and both leave the algorithms or procedures as mere recipes without a real justification. This paper follows the general justification for a parsimony criterion in phylogenetic analysis introduced by Farris (1983, 2008), later extended to the specific context of inapplicable characters (which Farris', 1983 formulation had not considered) by De Laet (2005, 2015), Goloboff et al. (2021) and Goloboff (2022). Thus, the criterion of parsimony is equated here with the idea that a phylogeny should maximize the number of observed similarities (shared character states) that can simultaneously be accounted for by common ancestry; a similarity between two taxa is attributed to common ancestry by connecting those two taxa in the tree with ancestors having that same state. Attributing to common ancestry is the same as “considering as homologous”; that is to say, a tree (contra Williams and Ebach, 2018, and Williams and Gill, 2025) is not chosen because its groups are based on characters we know to truly be homologies (which is unknowable), but instead because its groups allow us considering more similarities as homologies.

Some alternative formulations of “parsimony” have focused on aspects other than explaining observed similarities, such as the need to minimize steps or events (e.g. Kluge and Grant, 2006; Wheeler, 2023). The difference may seem subtle (and produces no difference for standard characters), but focusing on genealogical explanation of similarities or focusing on minimization of events may lead to different operational approaches—it is not just different justifications of the same criterion, it is actually different criteria, which can lead to the derivation of different methods and implementations. The difference is illustrated in De Laet and Goloboff's (2024) discussion of Wheeler's (2023) proposal to analyse inapplicable morphological characters with methods for insertion/deletion events in molecular sequences. An advantage of focusing on similarities is that it helps solving specific aspects of methods in a wide variety of circumstances, without the need to engage in difficult speculations about the correct delimitation of “events”. Whenever in a specific reconstruction of ancestral states a given condition occurs in two parts of the tree that are not connected by ancestors having that same condition, the observed coincidence is being interpreted as not resulting from common inheritance (the logical dependence between comparisons needs to be taken into account here; interdependent similar conditions are subsumed under a single instance; cf. Farris, 1983, p. 20). This can be applied also to transformations between morphotypes (as in Goloboff et al., 2021, Goloboff and De Laet, 2024): the cost of transforming from a set of character states A into a set of character states B can then be calculated by considering the number of similarities (in the original characters) not attributed to common ancestry if two changes A B occur in separate parts of the tree (this can of course take into account the weights of the original characters, if different from unity). In that case, the disconnected conditions of B constitute instances of similarities not attributed to common ancestry (=homoplasy), so that the corresponding cost can be set to reflect that exactly.

When the focus is on similarities attributable to common ancestry, it is clear that in two separate originations of a tail with a red color two similarities are not being attributed to common ancestry: the presence of a tail, and the red color. No intent of evaluating relative importance of events here—just a count of how many of the similarities (in the original characters, as delimited by the user) can or cannot be explained by common ancestry on different trees. As alternative homology schemes for parts containing different sets of characters often involve losses of parts (thus making all their subordinate characters inapplicable), the difference between minimizing number of similarities not connected by similar ancestors, and minimizing events, is also relevant when the homology of some parts (and their characters) is uncertain. This also applies to comparisons between the states observed for the characters that could be matched in different ways. When a part A with state 1 in character X A could be a homologue of part B with state 1 in character X B , or a homologue of part C with state 0 in character X C , then the difference is clear: if we choose A  =  C then the shared presence of state 1 in X A and X B cannot be explained by common ancestry (i.e. it cannot be attributed to homology). The matching of A with B is preferable, because it allows us to consider that the shared 1 between X A and X B is a result of common inheritance—which is what we want to maximize. The rationale for selecting among alternative homology schemes is, therefore, exactly the same as that laid out by Farris (1983) for standard characters.

Syntax for specifying alternative homologies

In TNT, the groups of characters of alternative homologies are specified using the xlinks command (also used to specify other types of dependences; Goloboff and De Laet, 2024, Goloboff, 2025). As TNT needs to allocate memory for the data structures used to store information about alternative homologies, the specification of alternative homologies needs to be enabled prior to reading the dataset, with an xlinks = N+ command (where N  = maximum number of links per character, with 1–32 as valid values). Once the dataset has been read into memory, the xlinks command followed by a double equal sign (i.e. xlinks== ) is used for specifying alternative character correspondences (a single equal sign is used for specifying other character interactions).

The implementation in TNT is designed for the case where body parts of dubious homology house comparable characters. This includes the case of parts that get transposed or shifted (so that the same number of parts exist, but their correspondences are unknown), or the case where one or more parts can disappear (with uncertainty as to which part remains and which disappears); both cases can coexist in a single dataset. To exemplify this, consider the hypothetical case of Fig. 1, where some animals have anterior and posterior lobes in their limbs (Fig. 1a). The lobes (Fig. 1b) can have or lack spines along the edge, marginal crenulations and short cusps. In some species (Fig. 1c), there is a single lobe, and it is not known whether it represents the anterior or the posterior lobe. In other species (Fig. 1d) the limb itself is rotated, so that one lobe can be recognized as superior and the other as inferior but (because the direction of the rotation cannot be determined) it is unclear whether superior = anterior and inferior = posterior, or the other way around, superior = posterior and inferior = anterior. Just as the homology of the parts is unclear, the same applies to the characters; in an animal with superior and posterior lobes, the spine of the superior lobe, spine sup , could actually be the same character as (=homologous to) the spine of the anterior lobe, spine ant (case in which spine inf will be a homologue of spine post ) or it could instead be the same character as the spine of the posterior lobe, spine post (case in which spine inf will be a homologue of spine ant ).

Fig. 1.

Fig. 1

An imaginary animal (a), where the limb can have two lobes. In some species the lobes are recognizable as anterior and posterior (b). In other species (c), there is a single lobe, but it is unclear (from morphology alone) whether it is an anterior or a posterior lobe. In other species (d), the lobes are (because of a rotation of the limb) superior and inferior, but (since the direction of the limb rotation is unclear) it is not known whether the superior lobe represents the anterior and the inferior represents the posterior or the other way around. The lobes can have different characters (e.g. spines, apical crenules or cusps), so that the homology of the characters found in the lobes is also unclear.

The xlinks command assumes that every part has the same characters (in this case, spines, crenules and cusps; a character may have only missing entries for a part). When an animal has a type of parts, and thus the characters of those parts are scored, the characters for all the other parts cannot be scored. Consider the example of Fig. 2: in an animal with superior and inferior parts (and thus with spines sup , spines inf , crenule sup , crenule inf , cusps sup , cusps inf all scored), the characters of the anterior and posterior parts (spines ant , spines post , crenule ant , crenule post , cusps ant , cusps post ) cannot be scored and are all indicated with a missing entry, and vice versa. In the dataset, the characters can be in any order, but grouping them together (as shown in Fig. 2, so that the types of body parts form visually recognizable blocks), will obviously facilitate visualization and error detection (as shown in Fig. 2, and done also by Ramírez, 2007).

Fig. 2.

Fig. 2

A dataset showing how to organize characters such as those shown in Fig. 1, for analysis using the new options in the xlinks command of TNT. The names of the parts are shown in the top part of the figure; colors correspond to the colors shown in Fig. 1. The first block of characters represents a type of parts (Ant and Post), and the subsequent blocks of characters represent other types, for which the parts are to be matched with the parts in the first block; every matching of parts implies a matching (=alignment) of the characters contained in the parts. For example, if Sup = Ant and Inf = Post, then spines ant  = spines sup and spines post  = spines inf (likewise for the other characters), but if Inf = Ant and Sup = Post, then spines ant  = spines inf and spines post  = spines sup . The xlinks command does not require that the characters of each part be next to each other (e.g. one could have listed all the spines, for each of the parts, followed by the crenules for each of the parts, etc.), but listing them as shown here facilitates organization (e.g. allowing a better visualization of the missing entries) and is therefore recommended.

The xlinks command reads lists of characters belonging to parts of ambiguous homology. The different types of body parts are separated by an equal sign (e.g. anterior and posterior, or superior and inferior, or single, are three different part‐types). In the implementation of xlinks , the first type of parts serves as reference and facilitates comparison with the subsequently listed part‐types; when the number of parts in subsequent part‐types is the same as that of the first, switching the order in the xlinks definition should not alter results.

Within each type of parts, the characters belonging to different parts are separated by slashes; within each part, the characters are listed by their name or number. The characters can be listed in any sequence, but the sequence (and the number of characters included, up to 32 in every part) must be the same for every part (i.e. the n th character of a part in a part‐type is a potential homologue of the n th character of parts in other part‐types). The numerical codes for the states are assumed to be comparable (i.e. a numerical code is assumed to correspond to the same condition in all of its potentially corresponding characters (doing otherwise would make the syntax of xlinks much more complicated, requiring specification of the equivalences between the numerical codes of the characters in the different types of parts). For part‐types beyond the first, the number of parts must be smaller than or equal to the preceding part‐types. This also implies that, in the xlinks implementation, a reference part (i.e. in the first defined part‐type) may have no homologues (thus being absent, if subsequent part‐types have fewer parts), but a part beyond the first will always be a homologue of one of the parts at the reference part‐type. Putting all of that together for the example of Fig. 2:

xlinks==012/345=678/91011=121314;

or alternatively,

xlinks==678/91011=012/345=121314;

but not

xlinks==121314=012/345=678/91011;

as reference part‐type has fewer parts than non‐reference types. Note that in this syntax, and in the dataset, there is no need for a separate character indicating the type of parts present; that is done automatically by the equal sign that separates types of parts. The parts may or may not be differentiable in themselves, or may differ in several individual characteristics, and thus TNT allows assigning transformations between parts a cost of zero or larger than unity (see below).

If the characters have been named (with the cnames command), those names can be used to specify the groups of characters with alternative homologies (the only requirement then being that the name of every character is uniquely recognizable); the entire alignment set can optionally be given a name, in parentheses, after the double equal symbol; the parts can also be given names, to facilitate specifying other settings and interpreting output, with the part name in braces preceding the list of characters for the part:

xlinks==(LOBES)ANTspinesantcrenulesantcuspsant/{POST}spinespostcrenulespostcuspspost=SUPspinessupcrenulessupcuspssup/INFspinesinfcrenulesinfcuspsinf={SING}spinessingcrenulessingcuspssing;

The section on Special Recoding Cases indicates how to specify other details of the interactions between parts and characters. The first concern, however, is how to recode into morphotypes simple situations like those in the syntax just exemplified.

Recoding into morphotypes: simplest cases

The simplest case to consider is when one of the parts can appear or disappear (as in Fig. 1c). For simplicity, the case where the parts have a single character (say, spines) is considered first; for the spine characters of Fig. 2, that would be

xlinks==ANT0/POST3=SING12; (DEFINITION 1)

The state combinations that need to be considered first for the complex of characters are those where both parts are present, and where both the ANT and POST parts can either have or lack spines. This is four combinations. Then follow two cases where the SING part is actually a transformed ANT part (either with or without spines), and two cases where the SING part is actually a transformed POST part. In the lists of morphotypes displayed by TNT, the original state for each character is indicated in the corresponding column and the character that is homologized to another is indicated with @ (followed by the number or name of the character it is being homologized to). That is,

ANT POST SING
Combin 0 3 12
0 0 0 ‐‐
1 0 1 ‐‐
2 1 0 ‐‐
3 1 1 ‐‐
4 0 ‐‐ @0
5 1 ‐‐ @0
6 ‐‐ 0 @3
7 ‐‐ 1 @3

This correspondence between morphotypes and original characters is shown by TNT with the vlinks! option. The most unusual aspect of such recoding is that some combinations of character states and homologies result in the same morphotype—identical morphologies that would be undistinguishable as different combinations on observation. In this paper, the term morphotype is reserved for a set of characters and states constituting an observable morphology, and the term combination is used for the added homology correspondences in the recoded supercharacter (i.e. every distinguishable morphotype represents a different combination, but not every distinct combination represents a different morphotype). Note that Goloboff and De Laet (2024) used the terms morphotype and combination interchangeably, as they did not consider the case of ambiguous homologies. With the present additions, the same morphotype may now appear in several lines of the output of the vlinks! command, as different combinations. Combinations 4 and 6 then both correspond to a SING part with no spines; any species with a morphotype of a SING part and no spines then must be recoded (internally) as possibly having both state combinations 4 and 6. This ambiguity corresponds to the alternative matchings, that is, the different ways to homologize character 12; the difference between the two combinations, 4 and 6, can only be realized when mapping a tree (i.e. selecting for the species and inner nodes the combination of states and homologies that minimizes score). Similarly, combinations 5 and 7 both correspond to a morphotype with a SING part with spines. The combinations of states and homologies that are assigned to each of the taxa in the example of Definition 1 are shown in Table 1.

Table 1.

State and homology combinations internally assigned to each of the taxa in the example of Fig. 2, for different definitions (1–3) of dependences

DEPENDENCE 1 2 3 4 5
Taxon_A 1 1 1 13 2
Taxon_B 1 1 1 15 3
Taxon_C 2 2 2 59 10
Taxon_D 2 2 2 33 6
Taxon_E 4,6 ? 12,14 163,171 40,44
Taxon_F 5,7 ? 13,15 167,175 42,46
Taxon_G 4,6 ? 12,14 162,170 39,43
Taxon_H ? 4,8 4,8 68,105 13,27
Taxon_I ? 6,9 6,9 89,126 24,37

The dependence #4 is as shown in Fig. 2 itself, for all the characters included: xlinks == 0 1 2 / 3 4 5 = 6 7 8 / 9 10 11 = 12 13 14 ; as there are no identical taxa in the dataset of Fig. 2, every taxon is assigned a different (set of) combinations. Note that the numbering of combinations is with exhaustive enumeration of all possible combinations; if eliminating combinations that may be unnecessary for obtaining optimal scores (with the methods discussed in Appendix 1), the numbering can change. This is shown in #5 ( xlinks&4;xlinks@4;xlinks!1; ), which discards 129 of the 176 combinations (with only 47 remaining, sufficient to produce the same optimal trees).

The transformation costs between these combinations must consider (a) the differences between observed states, (b) the transformations between parts and (c) the origination of parts and their characters, or their loss. By default, TNT assumes that the parts are distinguishable in themselves and gives transformations between parts a cost of 1 step: when a SING part originates from an ANT part in two distant parts of the tree, the similarity in having a SING part is not being attributed to common ancestry. Ramírez (2007), who used multiple datasets for his analysis of homologies, pointed out that the characters allowing recognition of the parts themselves had to be included in the dataset. With the present approach, that could be done, but doing so may multiply the number of possible combinations of recoded states enormously. The simplification of counting the transformation of the parts themselves serves the same function without multiplying the number of combinations (it would not be appropriate when different species display differences in the characters allowing recognition of the part). It is also possible that a SING part is in itself identical to an ANT or POST (so that a change from ANT or POST to SING does not cost anything), or that it involves several features (so that the change from ANT or POST to SING costs several steps instead of one). The syntax for specifying alternative homologies allows for both of these possibilities (see below).

A few example transformations between state combinations allow illustrating the mechanics for assigning the proper costs. Consider the number of similarities not attributed to common ancestry in two separate transformations from combination 1 (ANT with state 0 in spines, and POST with state 1) into combination 4 (SING with state 0 in spines and no POST): 1 instance of homoplasy for SING (the shared presence of SING instead of ANT is not being attributed to inheritance), none for spines (they are absent all along the path), and 1 for the loss of POST (the shared absence of POST is not being attributed to inheritance), for a total of 2. In the opposite direction (combination 4 into 1), the cost must be 3 instead: 1 instance of homoplasy for ANT, 1 for POST and 1 for state 1 in spines of POST.

Transforming from combination 1 into other combinations will have different costs, depending on the differences between the original states. For example, in two separate transformations from combination 1 into 6 the non‐homologized state 0 of spines in POST must be added, for a cost of 3, and in the opposite direction (6 into 1) the non‐homologized state 0 of spines in ANT (separated by regions of ANT absent), for a cost of 4. Using this logic for all possible transformations between combinations, the cost matrix can be filled:

from\to 0 1 2 3 4 5 6 7
0 1 1 2 2 3 2 3
1 1 2 1 2 3 3 2
2 1 2 1 3 2 2 3
3 2 1 1 3 2 3 2
4 3 3 4 4 1 3 3
5 4 4 3 3 1 3 3
6 3 4 3 4 3 3 1
7 4 3 4 3 3 3 1

These transformation costs are calculated automatically by TNT and shown with the vlinks =  option. With these costs for transformations between combinations, and the polymorphic scoring of terminal taxa having a SING part (either as 4 and 6, if having state 0 in spine, or 5 and 7, if having state 1, as shown in Table 1), having successive sister groups with both ANT and POST parts correctly identifies the expected matching of a SING part. When the spine in SING is present (morphotype corresponding to combinations 5, 7) and sister groups have a spine only in ANT (2) then SING is inferred to be a transformed ANT (combination 5: cost 2→5 is 2, less than 2→7 with 3), and vice versa for a spine only in POST (combination 1); when successive sister groups have a spine in both ANT and POST (morphotype corresponding to combination 3), then the homology cannot be resolved (as the cost from combination 3 into 5 and 7 is the same, 2 steps respectively); the same happens when successive sister groups both lack a spine (0 into 5 or 7 costs the same, 3 steps), or when the spine in SING is absent instead of present (if the spine is present in successive sister groups, cost of 3 into 4 or 6 is the same, 3 steps, and if the spine is absent in successive sister groups, cost of 0 into 4 or 6 is the same, 2 steps). The inferences of homology, as well as the inability to infer it, are then as expected. When several characters occur in every part (as in Fig. 2), there are many more combinations (176 for the case of Fig. 2), but the logic is similar: modified parts end up being considered as homologues of the parts with the fewest differences in their characters.

As mentioned above, it may well be that there is nothing in part SING itself that can tell it apart from ANT or POST; if the only difference between a species having ANT and POST and another species with SING is that one of the parts (ANT or POST) is missing but (on the basis of morphology alone) it is not possible to tell which it is, then one can (should) assign a transformation cost of 0 to changes from/into any of the reference parts and SING:

xlinks==ANT0/{POST}3={SING}0:ANT0:POST12;

With that, the state combinations themselves are the same ones as before, and the transformation costs have changes (highlighted in bold) in the entries that involve SING:

from\to 0 1 2 3 4 5 6 7
0 0 1 1 2 1 2 1 2
1 1 0 2 1 1 2 2 1
2 1 2 0 1 2 1 1 2
3 2 1 1 0 2 1 2 1
4 2 2 3 3 0 1 3 3
5 3 3 2 2 1 0 3 3
6 2 3 2 3 3 3 0 1
7 3 2 3 2 3 3 1 0

Two parts get modified

A slightly more complex case is when none of the parts get reduced, but they become modified so that (aside from the characters they bear) they cannot be recognized as the same parts (as exemplified in Fig. 1d). In the simplest situation, every part will just have a single character, say a spine; for the spine characters shown in Fig. 2 that would be

xlinks==ANT0/POST3=SUP6/INF9; (DEFINITION 2)

In this case, there are 12 combinations: the combinations of states 0 and 1 in the two characters, either as the reference part‐type (ANT and POST, combinations 0–3), or as SUP being ANT and INF being POST (combinations 4–7), and SUP being POST and INF being ANT (8–11):

POST ANT SUP INF
Combin 0 3 6 9
0 0 0 ‐‐ ‐‐
1 0 1 ‐‐ ‐‐
2 1 0 ‐‐ ‐‐
3 1 1 ‐‐ ‐‐
4 0 0 @0 @3
5 0 1 @0 @3
6 1 0 @0 @3
7 1 1 @0 @3
8 0 0 @3 @0
9 0 1 @3 @0
10 1 0 @3 @0
11 1 1 @3 @0

Four pairs of state combinations result in identical morphotypes and cannot be distinguished on a terminal: 4–8 (both SUP and INF with no spines), 7–11 (both with spines), 5–10 (SUP with no spines and INF with) and 6–9 (SUP with spines and INF without). The correspondences between state combinations and taxa for the example in Definition 2 are shown in Table 1.

The proper transformation costs between state combinations can be calculated as before (except that now there is no inapplicability, in contrast to the previous case). Note that in the previous case a SING part being in itself indistinguishable from ANT and POST would have required assigning no cost to the transformation between parts; a switch (in any direction) between a state combination with both parts and one with a single one would still have a cost (in the part vanishing or originating, depending on the direction). In the present case, it does not make sense to consider the possibility of SUP and INF being both indistinguishable from ANT and POST; otherwise the parts themselves would be considered as identical (which, clearly, they cannot be, as they are being scored and recognized!). At least a cost of one needs to be specified for the switch from the two ANT and POST parts to SUP and INF (or vice versa).

With the transformation costs between the 12 state combinations set as discussed, the inferences about homology are done as expected. Switches between a combination with both spines present in ANT and POST, and any combination with SUP and INF, cost the same, so that the ambiguity cannot be resolved in that case. Only when spines are present in ANT and absent in POST, or vice versa, is there a difference in cost for one of the members of the pairs 5–10 and 6–9. The homologies between parts (and their characters) are then recognized as expected.

Combining part modification and loss/gain

It is also possible, of course, to have in the same dataset some species where the parts become modified into SUP and INF, and other species where a part disappears and a SING part remains. This is then a case where multiple levels of possible homologies can occur, as multiple part‐types can transform into each other; if only the spines character is considered for the data in Fig. 2, the definition of dependences would then have three types of parts,

xlinks==ANT0/{POST}3=SUP6/INF9={SING}12; (DEFINITION 3)

and there are now 16 possible state and homology combinations,

ANT POST SUP INF SING
Combin 0 3 6 9 12
0 0 0 ‐‐ ‐‐ ‐‐
1 0 1 ‐‐ ‐‐ ‐‐
2 1 0 ‐‐ ‐‐ ‐‐
3 1 1 ‐‐ ‐‐ ‐‐
4 0 0 @0 @3 ‐‐
5 0 1 @0 @3 ‐‐
6 1 0 @0 @3 ‐‐
7 1 1 @0 @3 ‐‐
8 0 0 @3 @0 ‐‐
9 0 1 @3 @0 ‐‐
10 1 0 @3 @0 ‐‐
11 1 1 @3 @0 ‐‐
12 0 ‐‐ ‐‐ ‐‐ @0
13 1 ‐‐ ‐‐ ‐‐ @0
14 ‐‐ 0 ‐‐ ‐‐ @3
15 ‐‐ 1 ‐‐ ‐‐ @3

Calculation of transformation costs between these combinations now considers the costs of the two cases just discussed, simultaneously. The correspondence between taxa and state combinations, for the example of Definition 3, is shown in Table 1.

Restricting number of state combinations

It is easy to see that, with increasing number of characters per part, the number of combinations of morphotypes and homologies rapidly grows to a number difficult to manage. The use of efficient algorithms for calculating scores is then especially important (e.g. shortcuts to speed up tree searches with step‐matrix characters; Goloboff, 1998). TNT internally recodes character complexes into combinations, allowing up to 32 767 combinations, but tree calculations for more than a few thousand combinations are very slow and RAM demanding, even if using efficient algorithms for calculating scores. Some of those combinations, however, could never occur in a most parsimonious reconstruction of any tree. Appendix 1 discusses various methods to reduce the number of combinations so that only those needed for producing optimal scores are used; this is controlled with the & , @ and ! options of the xlinks command. The generation of all possible combinations in TNT (which is not the default) is invoked using the xlinks&0 option. Reductions in the number of combinations generated can be obtained with xlinks&N (with N  = 1 to 4, but 2–4 apply only to sets where some parts can disappear), xlinks@N (for uninformative states, only options 3–4 apply to ambiguous homologies) and xlinks!N (makes certain assumptions as to number of states in the characters of each part). In the main text of this paper, for generality, all examples assume that these reductions of combinations have not been enforced, but the reader should keep in mind that for real datasets it is generally advantageous to use those options (the amount of time and RAM saved will depend on the specific characteristics of the dataset).

Special recoding cases

Homology allowed only between some parts

In all the previous examples, a modified part could be considered as homologous to any of the parts in the reference part‐type. But–depending on anatomical details—a part beyond the first may have as alternative homologues only some of the parts in the reference part‐type. In the syntax for xlinks , this can be indicated with the “not” symbol ( ! ) after the list of characters for the part, followed by the list of parts in the reference part‐type that cannot be homologues (first part is 0; if parts named, names can be used in place of the number). Assume five binary characters 0–5, and the following ambiguities in homology, for animals that can either have parts A , B , C (with characters 0–2 respectively) or parts X , Y (with characters 3–4, respectively):

xlinks==A0/B1/C2=X3/Y4;

These unrestricted homologies between reference and non‐reference part‐types generate 32 possible combinations (assuming unnecessary combinations are not removed with xlinks& or xlinks! ). Suppose now that some anatomical details indicate that part X could be a homologue of parts B or C , but not A , and part Y could be a homologue of parts A or C , but not B :

xlinks==A0/B1/C2=X3!A/Y4!B;

As there are fewer possible pairings between parts, this obviously reduces the number of combinations, now producing (if the five characters are binary) 20 combinations instead of the former 32. The only requirement is that some ambiguity remains; a specification such as:

xlinks==A0/B1/C2=X3!A/Y4!BC;

will produce an error message, because part Y can only be a homologue of A (and then the character(s) in part Y are best merged with those of A and left outside the definition of alternative homologies).

A more flexible way to restrict possible homologies between parts is the creation of rules with xlinks itself, much as done by Goloboff and De Laet (2024) for other kinds of character interactions, but using the @ option instead of a state to indicate a homologization between characters, or @! to indicate non‐homology (this @ option is valid obviously only if some sets of characters have been defined already as having ambiguous homology). Assume parts A , B , C or parts X , Y ; the only restriction in how parts are to be homologized is that if X is a homologue of A , then we are certain that part Y must be a homologue of B :

xlinks==A0/B1/C2=X3/Y4;xlinks=3@0<4@1>;

with the restriction in the second line, instead of the parts XY being homologized in six possible ways ( AB , AC , BA , CA , BC , CB ), the second of those is invalid and the parts are homologized in only five possible ways ( AB , BA , CA , BC , CB ). Note that if multiple characters are included in each part, specifying a homology for just one of the characters suffices (as the other ones will follow).

Specification of transformation costs between parts

As indicated in the preceding sections, TNT by default assigns a cost of 1 to the transformation between parts (i.e. the same part originating in two separate regions of the tree constitutes one instance of a similarity not attributed to common ancestry). But if the parts are more different, differing in several recognizable details, it may be desirable to assign different costs. That can be done in two ways. First, in the xlinks command itself, when defining the possible homology correspondences, with the costs indicated in square brackets for each part (just before the list of characters for the part). For a reference part (i.e. in the first part‐type defined), the cost is a single number, indicating the cost of the part itself becoming applicable or inapplicable (the individual characters also require a cost when the part becomes applicable, of course, but that is calculated separately and depending on their weights and settings). For a part beyond the reference part‐type, the cost can be indicated for the transformation between the part and another reference part (needed only if different from unity); first the cost, followed by a colon and the number or name of the corresponding reference part; costs for transformations into/from several reference parts can be indicated, each with the cost and reference part separated by a colon. Note that this format cannot specify costs of transformation between two non‐reference parts (e.g. between a SUP and a SING part, in the example of definition 3 shown above for Fig. 2). As the cost of transformation between two non‐reference parts cannot be specified, when a non‐reference part transforms into another non‐reference part (e.g. a SUP into a SING), the cost is calculated as that needed to go from the reference part into the non‐reference parts. This can become complex and offers limited options.

A more flexible option is using the cost >> command. After defining the set N of ambiguous homologies, the cost command can be used to define the cost of transformation between any two parts beyond the first part‐type (costs between reference and non‐reference parts are specified in the syntax of xlinks== itself):

costNpartXpartY costXYpartXpartZ costXZ;

You only need to specify those costs different from unity. For this option to be available, you need to issue a cost≫; command (i.e. with no actual specification of costs) before reading the data (which enables the use of specific transformation costs between parts).

Invertible and non‐invertible correspondences

The order of the correspondences between parts, in the previous examples, is inverted in many combinations. For example, for three reference parts A C and two non‐reference parts X Y , the possible correspondences between parts would be:

X Y
0 A B
1 A C
2 B A
3 C A
4 B C
5 C B

This involves many “inversions” (e.g. combination 2, with B before A ). This may or may not be reasonable. For example, Ramírez (2007) considered three sclerites of the male palpal bulb, the primary conductor, the secondary conductor and a paramedian apophysis; the uncertainty in the homology between these three parts arose because they could shift positions and the bulb itself can be rotated and twisted; in such a case, inversions seem entirely possible. But Xu et al. (2009) considered the homology between fingers, and fingers are unlikely to switch positions—their order must be preserved. This will also be true of many cases with metameric parts, such as segments in arthropods—a first and second segment are unlikely to switch positions. So, if parts A C and X Y are non‐invertible, this can be indicated in the syntax of the xlinks command with the closing angle symbol ( > ), right after the double equal, forming an “arrow” trigraph:

xlinks==>A0/B1/C2=X3/Y4;

With such specification, TNT will generate only the combinations that present no inversions among the parts:

X Y
0 A B
1 A C
2 B C

A beneficial side‐effect of non‐invertibility is that the number of combinations can be substantially reduced relative to invertible sequences. The actual numbers depend, of course, on the numbers of characters and states in each part (and on their congruence and presence of uninformative states, if applying the xlinks& , xlinks! and xlinks@ options). The reductions are generally more substantial when there are more parts with fewer characters per part, and less so when every part contains many characters.

Adjacencies between parts

When there are several parts, the combinations generated when requiring non‐invertibility can still present “gaps” between parts. For example, with five reference parts A E and three non‐reference X Z , instead of the 60 orderings if inversions are allowed, there are 10 possible correspondences between parts without inversions:

X Y Z
0 A B C
1 A B D
2 A B E
3 A C D
4 A C E
5 A D E
6 B C D
7 B C E
8 B D E
9 C D E

Note that many skip a part; for example, combination 4 is the absence of part B in between A and C , and part D in between C and E ; combination 5 is the absence of parts B and C in between A and D . Depending on what is known about the anatomy and developmental patterns in the animals in question, this may or may not be acceptable. For example, one may well be willing to assume that intermediate abdominal segments do not disappear, only terminal ones. This can be indicated in the syntax of the xlinks command using a double angle symbol:

xlinks==>>A0/B1/C2/D3/E4=X5/Y6/Z7;

With adjacencies required there are only three correspondences (instead of the 10 for just non‐invertibility):

X Y Z
A B C
B C D
C D E

For part‐types with five and three parts, with two binary characters per part (and without application of xlinks& , xlinks! or xlinks@ ), there is a total of 4864 combinations if inversions are allowed, 1664 if the sequence is non‐invertible and only 1216 combinations if adjacencies are required.

Loss of adjacent parts

Particularly when limiting part losses or gains to adjacent parts, it may make sense to consider that several (e.g.) abdominal segments disappearing be subsumed under a single loss—the similarity being not in the loss of the individual segments, but in the collective loss of segments. Incidentally, this also shows that although considering parsimony as the maximization of similarities that can be explained by common ancestry (instead of the minimization of events) can make many decisions of how to evaluate trees or reconstructions easier, it does not solve all the problems and does not completely free the user from having to make decisions as to the nature of the observations contained in the data matrix.

To enforce this, the simultaneous loss of several adjacent parts (adjacent in the definition of the parts, that is; be careful with how you define the parts with alternative homologies) can be counted as a single step. This would (in the standard analysis of sequences) be sort of analogous to a gap “opening” cost with “extension” cost of 0. The simultaneous gain of parts must still be counted as several steps. This is just like the many steps counted when several characters become simultaneously applicable: as a matter of homologizing their alternative conditions. On the other hand, if the simultaneous gain was also counted as a single step, then violations of the triangle inequality would arise very easily (they cannot happen if only losses are merged). This does not change the combinations of states and homologies that are generated, only the step‐matrix costs. To indicate this, use an opening angle ( < ) after the double equal for the xlinks command; consider as example 6 binary characters, and the definition:

xlinks==<>>0/1/2=3/4=5;

The 22 combinations, and the corresponding step‐matrices (both counting every part loss independently, in the top matrix, and as a single loss when adjacent, in the bottom matrix) are shown in Fig. 3.

Fig. 3.

Fig. 3

Differences in transformation costs between combinations, when counting adjacent part losses separately (top cost matrix) and merging adjacent losses into a single event (bottom matrix, entries marked in grey), for 6 binary characters. The combinations (morphotypes and homologies) are identical in both cases, as shown in the left side.

Additional character interactions

For the characters of ambiguous alignment, all the options for other character interactions discussed by Goloboff and De Laet (2024) can be applied. As an example, consider Fig. 4, where a group of imaginary arthropods has two lobes, or a single one; when both present, the lobes can be ANT and POST, or SUP and INF (with the homology between both types unclear, as in previous examples); when only one lobe found, the part is named SING (see xlinks syntax in Fig. 4). The additional interaction for this example comes in that lobes can have or lack branchiae, but the branchiae can be fully absent (i.e. from both ANT and POST, SUP and INF, or the SING part) only when the animals are of small size (character 15) and have a thin cuticle (character 16), as otherwise oxygen cannot diffuse to the entire body. Note that all terminal taxa in Fig. 4 can breathe. A parsimony analysis for the dataset as shown in Fig. 4 (without considering the interaction between branchiae, cuticle and size) produces two trees, of 33 steps. However, as shown in Fig. 5, on the second tree, the node common to taxa B and C is reconstructed as being of large size, with a thick cuticle and lacking branchiae on both lobes—an ancestor that could not have breathed! With a command:

xlinks=151161<1141>;

the reconstructed ancestors are forced to have been capable of breathing, the second tree becomes one step longer and the ancestral reconstructions are somewhat changed. Note that characters 1 and 4 in the definition above stand for branchiae in the reference parts; the need to have state 1 in either character refers to the character in the reference part, or a character that has been homologized to it (remember that every character beyond the first part‐type must always be a homologue of a character in a reference part).

Fig. 4.

Fig. 4

A dataset illustrating the influence of morphofunctional dependences in addition to uncertain homologies. The homologies of the lobes are uncertain, as shown in the xlinks command. The taxa A–J can have or lack branchiae in their lobes, but the branchiae can be missing (state 0) only when the size is small (state 0) and the cuticle is thin (state 0), as otherwise the animal cannot breathe. None of the terminal taxa violate these constraints, but unconstrained analysis of the dataset (taking into account ambiguities in homology) produces two trees of 33 steps, onto one of which the common ancestor of taxa BC is reconstructed as being of large size and having a thick cuticle (thus incapable of breathing), as shown in Fig. 5. The lengths of the trees when non‐breathing ancestors are forbidden are shown in parentheses, in grey.

Fig. 5.

Fig. 5

Ancestral reconstructions for the second tree of Fig. 4 (with branchiae marked in grey). The reconstruction on the left is done allowing non‐breathing ancestors. The reconstruction on the right is done forbidding them (with the xlinks command shown above the tree), and in that case the best possible reconstruction implies one additional step over the unconstrained reconstruction. This also means that a reconstruction shifting the homology interpretation (with C having the Sup part homologous to Ant, and the Inf homologous to Post, parsimoniously reconstructing the common ancestor of BC as having branchiae) is as parsimonious as the original interpretation where Inf = Ant and Sup = Post (which implies an additional instance of homoplasy, in forcing the reconstruction of the common ancestor of BC as having branchiae).

On the second tree, when forbidding non‐breathing ancestors, the homology of the parts is now ambiguous; the SUP lobe of taxon C can now be a homologue of ANT, and INF of POST, with equal parsimony. In the reconstruction allowing non‐breathing ancestors the SUP lobe of taxon C was unambiguously homologous with POST, and INF with ANT; this is because the SUP part of taxon C shares two character states with the POST part of taxon B (spines and cusps), and a single one with the ANT part of B (branchiae); with the homology chosen when allowing non‐breathing ancestors (Fig. 5, left reconstruction), the branchiae are then present in two different parts in taxa B and C, and thus absent from their common ancestors (since successive sister groups lack branchiae in both lobes, which does not prevent them from breathing, because they are of small size and have a thin cuticle). When the ancestors are forced to be breathers (Fig. 5, right reconstruction), an additional step (thus producing a length of 34 in this tree) is produced by requiring branchiae to be present in one of the lobes in the ancestor of taxa B and C, but this is as parsimonious as shifting the interpretation of the homology between parts (with one more step in each of spines and cusps, and one less step in branchiae, thus having in total one more step of homoplasy than the unrestricted reconstruction), now interpreting branchiae as present in the POST lobe of both B and C.

Inapplicable characters, parts or sets

Two of the characters belonging to a part can also be related, so that one determines applicability of the other (and both become inapplicable if the part is absent); that requires no special considerations. A different situation is when a character outside of the parts determines applicability. In that case, different scenarios are possible, depending on whether the external character makes inapplicable one (or more) character(s) in a part, the part itself or the entire set of parts. Consider as example the situation when two parts (one of which can be absent) have a single binary character. Then, when the external character (say, number 3) makes the characters inapplicable (with 0 being the disabling state, that is, the state that indicates the dependent characters are inapplicable), this would be indicated as:

xlinks==0/1=2;xlinks=30<0><1>;

Note that (in the second line) there is no need to indicate that character 2 is inapplicable, as it is always a homologue of either 0 or 1 (for the case where a character beyond the first part‐type must be inapplicable only for certain homology correspondences, you can use the @ option to indicate “when homologous to”; for example, instead of the above, xlinks = 3 0 & 2@0 < 0 – > , character 0 is inapplicable only when character 3 has state 0 and character 2 is a homologue of 0). For both characters 0 and 1 being inapplicable when 3 is 0, there are 11 possible combinations (below, combinations of 0 and 1 states when the character is applicable are indicated as 0/1, to save space). There are three types of combinations with the disabling state 0 in character 3: the parts in their reference part‐type and the characters 0 and 1 being inapplicable (combination 0), the POST part absent and ANT part present but with character 0 inapplicable (5), and the ANT part absent and POST present but with character 1 inapplicable (8):

ANT POST SING EXT.CH.
Combin 0 1 2 3
0 ‐‐ ‐‐ ‐‐ 0
1–4 0/1 0/1 ‐‐ 1
5 ‐‐ ‐‐ @0 0
6–7 0/1 ‐‐ @0 1
8 ‐‐ ‐‐ @1 0
9–10 ‐‐ 0/1 @1 1

If instead character 3 makes both parts ANT and POST inapplicable (e.g. if character 3 is absence or presence of the limb; with no limb, there cannot be any lobes!), this can be indicated with a double angle (and the name/number of the set of characters of ambiguous homology):

xlinks==0/1=2;xlinks=30<<ALSET 0>>;

And now there are 9 combinations instead of 11, with only one type of combination with state 0 of character 3 (with both parts then absent):

ANT POST SING EXT.CH.
Combin 0 1 2 3
0 ‐‐ ‐‐ ‐‐ 0
1–4 0/1 0/1 ‐‐ 1
5–6 0/1 ‐‐ @0 1
7–8 ‐‐ 0/1 @1 1

Alternatively, if one were to consider that the external character makes inapplicable just one of the parts (say, ANT, instead of the entire set of characters of ambiguous homology), this would be indicated as before, but with a colon and indication of the part number or name:

xlinks==ANT0/{POST}1={SING}2;xlinks=30<<ALSET 0:ANT>>;

This will now yield 13 combinations, five of which correspond to character 3 with state 0: two combinations (0–1) with the part ANT inapplicable and POST with character 1 having states 0 or 1; two combinations (9–10) with the POST transformed into SING (ANT absent) and character 1 having states 0 or 1, and one combination (6) with POST part inapplicable and part ANT absent (given that the external character has state 0):

ANT POST SING EXT.CH.
Combin 0 1 2 3
0–1 ‐‐ 0/1 ‐‐ 0
2–5 0/1 0/1 ‐‐ 1
6 ‐‐ ‐‐ @0 0
7–8 0/1 ‐‐ @0 1
9–10 ‐‐ 0/1 @1 0
11–12 ‐‐ 0/1 @1 1

Depending on the anatomical, functional or developmental relations between the characters, one or other type of dependence will be best suited to express their relationships. In the latter case, combination number 6 (both parts absent) may or may not be acceptable; if none of the terminal taxa has the corresponding morphotype, having the combination among the set of possible states does no harm (as it will never be optimally assigned to any internal node, since step‐matrices do not violate the triangle inequality); if the combination is considered impossible, it can be eliminated by creating rules that make it impossible. An example would be xlinks = 3 0 & 2@0 < 2@1 >  (when character 3 is 0 and SING part is a homologue of ANT part, then the SING part must be a homologue of POST part, which can never be fulfilled). This produces only 12 combinations, now without number 6:

ANT POST SING EXT.CH.
Combin 0 1 2 3
0–1 ‐‐ 0/1 ‐‐ 0
2–5 0/1 0/1 ‐‐ 1
6–7 0/1 ‐‐ @0 1
8–9 ‐‐ 0/1 @1 0
10–11 ‐‐ 0/1 @1 1

Parts of ambiguous homology contained in parts of ambiguous homology

The options in the current implementation provide great flexibility and adequate handling of a big number of situations. Morphology being so complex, however, one can always imagine a situation where the current options are insufficient. Consider a case where the animals can have (instead of the single one of Fig. 1) two limbs (I and II), each of which can have two lobes (ANT and POST, or SUP and INF if the limb is rotated). If the lobes are modified independently in each limb, that problem would be tractable in the current implementation by creating two groups of characters of ambiguous homology, one for each limb.

Things get much more complicated if one of the limbs (Fig. 6) can be missing, and it is unclear whether the remaining limb is a I or a II. There are here two interrelated types of ambiguity: one in the limbs themselves, another in their lobes and one of the two types is contained within the other. There is no direct way to express this relationship in the syntax for xlinks —every part in a part‐type beyond the first must explicitly refer to one of the parts in the first part‐type (put differently, there cannot be more than one reference type of parts, and this situation would call for two, one for lobes, another for limbs themselves). As certain patterns of homology between the parts in different part‐types can be forced, and the costs between parts themselves can be set to any value, it is possible to reflect the situation just described with xlinks , although this is somewhat laborious. First, for xlinks itself, we need to consider the situation where the lobes of limb I are recognizable as ANT and POST, but those of limb II are rotated into SUP and INF (when two legs present, homology of I and II is always clear, obviously); this is level 2 in Fig. 6. Note that parts I‐ant2 and I‐post2 cannot be homologues of parts 2 or 3 of the reference part‐type (as they are on limb I, and parts 2 and 3 are on limb II); likewise for parts II‐sup2 and II‐inf2, which cannot be homologues of parts 0 and 1 of the reference part‐type (they are on the other limb). The reverse needs to be considered as well: lobes of limb I are rotated into SUP and INF, and those of limb II are recognizable as ANT and SUP (level 3). Level 4 is when both limbs have their lobes rotated as SUP and INF (but still recognizable as limbs I‐II, so that the same homologies are forbidden with the “!” symbol). The next level, 5, is finally when the lobes are recognizable as ANT and POST, but there is a single limb; as we do not know whether that single limb is I or II, characters 16 and 17 can be homologues of 0 and 1, or of 2 and 3 (further specification banning combinations like 16 being a homologue of 0, in limb I, and 17 a homologue of 3, in limb II, are done with a different xlinks command and use of the @ symbol, see below). The last level, 6 (with parts X‐sup6 and X‐inf6), is when the lobes are rotated into SUP and INF, and there is a single limb.

Fig. 6.

Fig. 6

Imaginary animals where the homology between the anterior and posterior lobes, with superior and inferior ones, is unclear, as well as the homology of the limb remaining when one of the two limbs is lost. The syntax of xlinks does not allow a direct coding of such a situation, but an appropriate coding can be achieved with the different options allowed by the xlinks and cost commands. See text for discussion.

Two further commands are needed to take into account the situation. The first is, again, xlinks , to ban specific combinations that could not be banned before; for example, while II‐sup2 and II‐inf2 can either be homologues of II‐ant1 or II‐post1, I‐ant2 can only be a homologue of I‐ant1—but if such restriction is imposed in the first instance of xlinks , there is no ambiguity in homology for that part. Then, the second xlinks command states that if character 4 (in part I‐ant2) is a homologue of character 1 (in part I‐post1), then character 4 must also be a homologue of 0—impossible, so that 4 ends up not being a homologue of 1 in any combination (alternatively, one might have given a condition that is always fulfilled, and then stated that 4@!1 , 4 is not a homologue of 1; either would work). Other restrictions actually enforce some possibilities, such as if character 16 is a homologue of 2, then 17 must be a homologue of 3 (i.e. if part X‐ant5 is a homologue of II‐ant1, then part X‐post5 must be a homologue of II‐post1, both on the same limb).

These restrictions of the two xlinks commands produce the correct lists of combinations (Fig. 7) for a situation like shown in Fig. 6, with a total of 168 combinations. Note in the list of combinations that characters 0–1 and 2–3 only get “transposed” into 1–0 or 3–2 in combinations involving SUP and INF; all cases of lobes having ANT and SUP correspond to 0–1 and 2–3.

Fig. 7.

Fig. 7

List of combinations obtained for the example of Fig. 6 (when all 20 characters are binary, a total of 168 combinations is produced when unnecessary combinations are not removed).

The last command in Fig. 6 specifies some costs of 0 for transformations between some parts. Note that in a change from both limbs having ANT and SUP lobes, into limb I identical but limb II having the lobes as SUP and INF, there are differences in limb II, but no actual change in limb I. By default, TNT will (as discussed above) assign a cost of 1 step to any transformation between parts, but the transformation between the parts (e.g.) I‐ant1 and I‐post1 into I‐ant2 and I‐post2 implies no cost (i.e. no homoplasy), unless of course there is a difference in the characters contained in the parts themselves (which TNT calculates separately). Then, those 0 costs for “transformations” involving identical parts must be specified (at the bottom of Fig. 6).

The series of commands needed for reflecting the situation shown in Fig. 6 is admittedly not simple, but if xlinks were not available, recoding the situation would be next to impossible. For example, using combination recoding, consider that (with two binary characters contained in every part) there are 2400 combinations (a number still manageable by the current implementation, even with 32‐bit versions of TNT) and then 2400 × 1399 transformation costs between combinations, or 5 757 600 cells to fill in the step‐matrix for combinations. Even if the commands shown in Fig. 6 are not simple, they are a lot easier than manually filling millions of entries in a step‐matrix. Additionally, with checks for compatibility and uninformative states already implemented in TNT, the number of combinations to use can (depending on characteristics of the dataset) be further reduced from 2400, thus saving some time. Manually performing those checks for elimination of unnecessary combinations would be extremely tedious.

Determining minima (i.e. cost of no homoplasy)

Goloboff et al. (2021) showed that, in the presence of inapplicables (such as parts that can come and go, in the present setting) and with appropriate step‐matrix costs, every additional step necessarily corresponds to an additional instance of homoplasy, so that trees and reconstructions are properly ranked; still, the point of origin—no homoplasy—needs to be established by reference to the appropriate number of steps. Establishing such point of reference is important both to calculate absolute values of homoplasy (otherwise, only differences in the homoplasy required by different trees or reconstructions could be calculated), and for procedures like implied weighting (which downweights characters on the basis of homoplasy; Goloboff, 1993).

Goloboff et al. (2021) discussed the problem that the minimum possible score actually achievable for the recoded supercharacter is not the appropriate basis of comparison for calculating amounts of homoplasy. Consider Dataset 1,

A 0 0 0
B 1 1 0
C 1 0 1
D 0 1 1
Dataset 1

for which the most parsimonious tree(s) have 5 steps; no trees with fewer steps exist for that dataset, but those 5‐step trees imply 2 steps of homoplasy for the original characters. Recoding the three characters in a complex (representing all combinations of states, and their transformation costs), the minimum achievable cost for the character is, again, 5 steps (this minimum achievable cost is difficult to calculate for step‐matrix characters—except by searching trees—when some states are not present in any terminal, as will often be the case for recoding into morphotypes; see Goloboff, 2022, for brief discussion, and Hoyal Cuthill and Lloyd, 2024 for a revision). However, those 5 steps cannot be used to indicate absence of homoplasy; absence of homoplasy is correctly indicated by the sum of minima for the original characters, 3 steps. With that theoretical (but unachievable on any tree) minimum of 3 steps for the recoded character, the homoplasy is now correctly evaluated as S – M (observed steps minus minimum) on the most parsimonious trees: 2 steps of homoplasy. Much the same considerations apply to character complexes containing inapplicable characters, except that then the theoretical minimum must also consider the extra cost of forcing the root to have the disabling state of the primary character, if scored for the outgroup (see Goloboff et al., 2021).

For sets of characters of alternative homologies, a theoretical minimum that indicates absence of homoplasy can also be calculated, and every step beyond that minimum correctly indicates an additional instance of homoplasy. That homoplasy may be in separate originations of states within characters, or separate originations of the parts. However, that minimum cannot be calculated solely from the minima of the characters themselves, as was the case for fixed alignments. Consider first a very simple dataset, with parts L and R (“left” and “right”), that can get transformed into X and Y , or just Z , each with a single character:

L R X Y Z
A 0 0
B 1 1
C 1 1
D 1 1
E 1
Dataset 2

The alternative homologies are then defined as,

xlinks==L0/R1=X2/Y3=Z4;

The sum of minima for the five characters is 2 steps (1 + 1 + 0 + 0 + 0), but aligning X , Y or Z with L or R must cost a step in every case (as the parts are themselves recognizable). Therefore, the number of steps indicating no homoplasy is 6, achievable for example, on the tree (A(B(C(DE)))), with two transformations 0➔1, L X , X Z , R Y and Y becoming absent. But if the data had instead been (note changed state in bold),

L R X Y Z
A 0 0
B 1 1
C 0 1
D 1 1
E 1
Dataset 3

then the minimum possible steps (i.e. in the absence of homoplasy) would continue being 6, but that score is unachievable on any tree: one must necessarily postulate scenarios with homoplasy. A possibility is a part L with state 0 transforming into an X with state 0, and two independent changes into state 1 (both in the untransformed L and the transformed X parts, with the state 1 then occurring in two separate parts of the tree, that is, homoplasy). Other combinations are possible ( L X with state 1, then two changes into 0; or R X with state 0 or 1, then two changes into 1 or 0, respectively), but any of those implies one instance of homoplasy as well. Note that there is no apparent conflict between the original characters of this second dataset, but no tree can be free of homoplasy; the conflict arises in the combination of parthood plus characters. Further modifying the dataset and adding a terminal,

L R X Y Z
A 0 0
B 1 1
C 0 1
D 1 0
E 1
F 0
Dataset 4

The score that indicates no homoplasy continues being 6, but that length is unachievable on any tree—best possible trees have 9 steps, or 3 steps of homoplasy (e.g. in 3 independent derivations of the state 0 in characters X , Y and Z , along with several other alternatives).

For calculating the correct minima, TNT considers all the possible alignments. While the separate minimum steps for characters L and X sum up to 2, the combined minimum steps for L aligned with X is now 1 (with some 0's and some 1's, a single step is required to transform from one into the other). Then, the minima to be used are those of the combined characters. The cost of transforming L X must be added as well. Likewise for parts Y and Z , producing the total of 6.

Note that in the Datasets 2–4 shown above, the combined minimum was more than the sum of minima for the individual characters, but the combined minimum may well be less than the sum of minima:

L R X Y
A 0 0
B 1 1
C 2 2
D 3 3
E 0 0
F 1 1
G 2 2
H 3 3
Dataset 5

The minimum for each separate character is 3 (e.g. 0→1→2→3), for a total sum of individual minima of 12. However, the combined minimum if defining alternative homologies

xlinks==L0/R1=X2/Y2;

is now less than 12. If aligning L with X and R with Y , each of the combined characters has states 0–3, for 6 steps combined in the absence of homoplasy; adding to this the cost of transformations L X and R Y , a total of M  = 8 steps is obtained as minimum. However, that score of 8 cannot be obtained on any actual tree; for example, if placing together every 0, 1, 2 and 3 when L is a homologue of X , then that would require three instances of homoplasy, in four separate derivations of the X parts; likewise for R , Y . Therefore, no tree for the dataset above can have fewer than 14 steps, of which S – M  = 6 are homoplasies. Those are the correct counts.

Generally speaking, when more than a single state is shared by a character C in a reference part and the character in another part that can be aligned with C , this will imply homoplasy on any tree, either as independent derivations of the same character state (if placing together the taxa having similar parts) or as independent derivations of the same part (if placing together the taxa having similar character states). If the reference minima are correctly calculated by enumeration of alignments, then every step beyond that reference value does correspond to an instance of homoplasy.

Origination of parts de novo

An interesting consequence of the approach discussed is that, at least with default weights, an origination of parts de novo can never be preferable to the alternative of modifying a preexisting part. Of course, some parts can be absent and then appear, which constitutes some kind of origination. Consider for example Dataset 6,

L R X
A 0
B 0
C 0 1
D 0 1
E 1 1
Dataset 6

and the tree (A(B(C(DE)))), onto which the characters are most parsimoniously reconstructed as having a (single) part X at the root, with X changing into L and R appearing de novo at the common node of CDE.

But when the option is between considering a part/character as homologous to another part/character, or instead homologous to nothing, the situation is different. That would have to be scored as an additional reference part (call it void, V ) with only missing entries in its characters, with the part whose de novo origination is suspected ( Y ) now being either a homologue of V or a preexisting ( PE ) part, as in Dataset 7:

PE V Y
A 1 1 1
B 0 1 1
C 1 1 1
D 1 1 1
E 1 1 1
F 0 0 0
Dataset 7

The dependences between characters need to be defined taking into account that part V itself is inapplicable when Y is not a homologue of it (i.e. when character 6 is not a homologue of 3):

xlinks==(NOVO)PE012/V345=Y678;xlinks=6@!3<<NOVO:V>>;

Clearly, any state shared between PE and Y will be explainable by common ancestry if considering that Y is a transformed PE , instead of a de novo part originating from the void V (an example would be a tree where taxon F is the sister group of B, sharing a 0 state). Therefore, the most favourable situation possible for Y being a part originated de novo is when no character states are shared between Y and the preexisting PE part in its sister taxon, as in tree (A(B(C(D(EF))))). But even in that most‐favourable case, the idea that Y in taxon F is a transformed PE part is more parsimonious than the idea of Y arising de novo. The cost from an all−1 combination with part PE changing into an all‐0 part Y is 4 (two separate such transformations will fail to homologize 3 character states, plus a part being Y instead of PE ), while the cost from an all−1 combination with part PE disappearing and a part Y appearing de novo as a “homologue” of V is 5 (two separate such transformations will fail to homologize 3 character states, plus the absence of part PE , plus the presence of part Y ). The assignment of part Y as a transformed PE part is, therefore, always preferable. The only way to prefer homologizing Y with V is by giving a priori some of the transformations a higher cost, thus pushing the decision into a de novo origination; with all the characters and parts weighted equally (i.e. the default in TNT), the de novo origination can never be preferred on the grounds of parsimony. If the part potentially originating de novo has only invariant characters, an alternative is to eliminate those from the counts (remember that invariant characters, by default, add to the cost). This can be done with the xlinks@ —option, but it seems far from obvious that it is well justified (two tails separately originating with the same invariant state still fail to homologize that state; see De Laet, 2018, Goloboff et al., 2021). Only when using this option is the alternative of Y being homologized with V preferable to PE .

In the context of unaligned sequence data, De Laet (2005, discussion of fig. 6.13; 2015:559‐562) made similar considerations, concluding (contra Kluge and Grant, 2006, Grant and Kluge, 2009, and Wheeler, 2012) that maximization of homology for unaligned sequence data can never lead to trivial alignments, alignments where sequences that a priori are thought to be putative homologues are explained as having arisen de novo. The way in which common ancestry functions as an explanatory concept, and how this can or cannot be reasonably applied to suggest the origination of new parts (instead of only to understand similarities between preexisting parts), also pervaded some early discussions about monophyly of the Arthropoda, in pre‐cladistic times. Sidnie Manton, a renowned functional morphologist, published extensively on arthropod evolution, summarizing her views in Manton (1977). She considered that the limbs of different groups of Arthropoda were so substantially different that it was “impossible” that one derived from the other, or that both derived from a common ancestor—no intermediate form could have been functional, in her opinion. Thus, she argued, the main groups of extant Arthropoda (Chelicerata, Crustacea and Uniramia) had to be derived independently from different ancestors—making the Arthropoda polyphyletic. But (as noted by James S. Farris) that argument is thoroughly self‐defeating: if the uniramous limbs of insects and miriapods differ so much from the biramous limbs of crustaceans that one could not have originated from the other… why are they more likely to have originated separately from something that was even more different? Unless there is, outside of “arthropoda”, some other animal group with limbs more closely resembling crustacean limbs, or insect limbs, Manton's idea simply makes no sense (Platnick, 1978 made a similar point). Crustacean and insect limbs must have originated from something (unless each group represents a separate origination of Life; Manton did not go as far), and it is unreasonable to assume that that something was even more different than those two. Manton's argument was not framed in numerical terms, but the general ideas discussed for Dataset 7 still apply.

Of course, that is not to say that no part ever originates de novo. What this means is that the phylogenetic criterion of parsimony (i.e. maximizing the degree to which similarities are attributable to common ancestry) cannot by itself prefer such hypotheses; a modification of a preexisting part is generally a better alternative on the basis of that criterion. The idea that some part originates de novo, instead of being a modification of an existing but potentially homologous part, must be decided on the basis of considerations other than phylogenetic parsimony (e.g. anatomical differences, developmental patterns, absence of any possible correspondence with other parts) and embraced when those considerations simply rule out the possibility that the part in question is a modification of another.

Creating “static” alignments

A drawback of the present approach is that only a relatively reduced number of characters with alternative homologies can be combined to form character complexes. Even with few parts, having large numbers of characters per part (and those characters having many states and few compatible characters) multiplies the number of combinations to the point of making the problem intractable in practice. In that case, an approach creating alternative alignments—as proposed by De Laet (2005) and Ramírez (2007)—may still be feasible, particularly if (as done by Ramírez, 2007) one can assume beforehand that some taxa share the same homologies between parts.

As mentioned already, the scripts used by Ramírez (2007) are rather complex and specific to his dataset. To facilitate generating all possible alignments, TNT includes now some scripting commands, which allow for much simpler scripts. The first of those is the  hambaln command (for “handle‐ambiguous‐alignments”), which can be used to generate static alignments. Those static alignments have the limitation that the characters in a part‐type beyond the first are aligned with the characters of the same reference part across all taxa (this would not allow e.g. a SING part to be a homologue of ANT in some taxa, but a homologue of POST in others; evaluating that possibility would require that every terminal with a SING part is coded as having a different type of SING part, with SINGA, SINGB … SINGN as multiple part‐types).

The hambaln command can be followed by the number or name of an alignment set with T types of parts and P parts in the first part‐type, and the name of an array with T columns and P rows. The array specifies the alignment to perform, with the first P values ignored (use 0 to P –1 for the first column, representing each of the groups in the reference part‐type) and the value in the next columns indicating the part of the reference part‐type that the entry is to be aligned with (with −1 indicating “part absent”). For example, with set 0 having T  = 3, P  = 4, and the array myaln :

0 2 1
1 −1 −1
2 1 −1
3 0 0

the command hambaln = 0 myaln will perform an alignment (for the set 0 of characters with uncertain homology) where the characters of part 0 of the reference part‐type are aligned with those of part 2 in the second part‐type and those of part 1 in the third; the characters of part 1 at the reference part‐type will not be aligned with any other; the characters of part 2 in the reference part‐type will be aligned with those of part 1 in the second part‐type; and the characters of part 3 in the reference part‐type will be aligned with those of the first part in the second part‐type, and those of the first part in the third part‐type. The alignment is “realized” in the dataset by substituting the missing entries in the corresponding reference characters with the state in its homologous character (remember that the dataset must have non‐overlapping missing entries, as in Fig. 2), changing the first character of each part at subsequent part‐types to indicate part membership (a 0 for what was missing, a 1 for what had an entry; the entry is transferred to the character this one is being aligned with), and deactivating characters beyond the first in each part in the second part‐type or beyond. This allows “undoing” the alignment easily, with a hambaln– command: the 0 in the first character of each non‐reference part must be reset to a missing entry, the 1 must be reset to the state in the character this first character was being aligned with and the reference characters changed into a missing entry, and the characters beyond the first part‐type in each non‐reference part must be reactivated. If desired, an alignment realized with hambaln can be saved with standard TNT commands for saving data (e.g. the xread command, see online help of TNT).

If there are parts that can disappear (i.e. if some part‐type beyond the first has fewer parts than the first), then treatment with the hambaln command is only approximate, because it cannot properly evaluate the homologies in inapplicable characters. It could be possible in principle to use a step‐matrix recoding for the aligned data, but that would create many combinations, which is precisely what the static alignments try to avoid. Remember that some alignments can be invalid (e.g. with rules defined with the xlinks== command), and then TNT provides the expression charalnreal[S] , which returns 1 if the alignment of set S is statically realized, 0 otherwise (invalid alignments are just skipped by the hambaln command, for example, when the alignment specified in the array implies an inversion of parts and the correspondences between parts have been defined as non‐invertible, see above).

As generating all possible permutations of the rows and columns in the matrix that represents an alignment can be laborious, the combine scripting command of TNT can be used to generate those, with the ^ option:

combine^TP1,P2,,P3myalnTNTcommandsendcomb

This will write all possible alignments into the array myaln with T columns, and P 1 rows, every time executing the corresponding TNT commands. Remember that T must equal the number of part‐types in the set, and P 1 the number of parts in the reference part‐type; the combine command requires as many comma‐separated values in parentheses as part‐types T (every value represents the number of parts for each part‐type).

With this, generating all the alignments for a set of characters of ambiguous homology is relatively simple. A fragment of a script, ONE_ALN , in Appendix 2, shows how to generate alignments for a given set of characters of uncertain homology (passed as first argument to ONE_ALN ), calculate the score of each (saving up to 10 optimal alignments to array optalns ; this array must be appropriately declared before calling the routine) and then undo the alignments. This script calculates the score of each static alignment on a tree (tree 0) instead of doing a search for each alignment; this is much faster than doing a search for each alignment (as did Ramírez, 2007), and can easily be used to produce a sort of successive approximation to the optimal alignment and tree.

Implied weighting

As discussed by Goloboff et al. (2021) and Goloboff and De Laet (2024), implied weighting (Goloboff, 1993) is fully compatible with the idea of maximizing similarities explained as homology, as there is no requirement that failures to homologize similarities must have the same cost for every character. A problem that arises here is that (as noted by Goloboff et al., 2021), while implied weighting of character complexes is theoretically possible, step‐matrix optimization algorithms (Sankoff and Rousseau, 1975; Goloboff, 1998) cannot be used to find reconstructions of maximum fit. As an (admittedly simplistic) work‐around, TNT calculates the score for a character complex using the score that corresponds to the average homoplasy of all the characters in the complex, multiplied by the number of characters in the complex (as in uniform weighting of sets of characters; see discussion in Goloboff, 2014). Thus, for a total homoplasy of H in the complex with N characters, the score (for concavity k ) is NxHN/k+HN. It is important to take into account that, in sets of characters of ambiguous homology, the number N must correspond to the number of characters in the reference part‐type (instead of the total number of characters included in the complex, since any character at subsequent part‐types will necessarily be a homologue of another character in a reference part) plus the number of parts that can be transformed (since transformation between parts has a cost and corresponds implicitly to another character).

Discussion and conclusions

The present approach provides a practical solution to the problem of alternative homologies between parts and their characters, by reference to combinations of states and homologies recoded as a complex supercharacter. With a relatively simple syntax, a wide variety of relations between parts and their characters can be considered. All the recoding and calculation of transformation costs between combinations is fully automatic, done internally by TNT. The alternative homology matchings between parts can also be combined with the definition of functional or biomechanical dependences (e.g. forbidding some character state combinations, such as stop‐codons, to be assigned to ancestors). The options provided with the current implementation may not cover all possible cases of uncertainty in homologies, but they probably serve to correctly handle the majority of situations that arise commonly in the analysis of morphological datasets.

While some authors (e.g. Ramírez, 2007:602–603) have doubted that analyses considering alternative homologies (instead of just settling on a potential homology scheme) could produce substantially different results, it is undeniable that the only way to seriously test that possibility is by using methods that can properly evaluate alternative homology schemes. The methods provided here are a step in that direction. Of course, in any case where morphology clearly points to one set of homologies, the implications of morphology should be followed and the present methods are irrelevant.

Given the computational cost (and additional difficulties) associated with applying the parsimony criterion to alternative homology schemes, it should probably be used sparingly, only for those special cases where morphology alone seems insufficient to establish homology of parts containing especially interesting or relevant characters. As the calculations may require too much memory and time when there are many characters per part, the current implementation also offers the possibility of easily creating alternative static alignments; although not exact (in not allowing parts to have different homology correspondences in different areas of the tree, and in the treatment of inapplicable characters, often a consequence of some parts being absent in some of the taxa), the static alignment approach may allow a faster approximation in some cases.

Acknowledgements

I thank Santiago Catalano, Jan De Laet, Martín Ramírez and Claudia Szumik, and an Anonymous Reviewer, for comments and discussion on the topic of this paper. I thank Jan De Laet's insistence to improve my terminology, which in the end led to greater clarity in the ms. This paper was presented at the XLII Annual Meeting of the Willi Hennig Society; I thank members of the audience for input and discussion. I also thank Jacqueline Silviria for facilitating additional bibliographic references.

Appendix 1. Restricting number of combinations

For comparison, the discussion in this Appendix uses the dataset in Fig. 2 but with two copies of every taxon and the full set of links (with three characters per part and two types of parts), which can generate 176 combinations of morphotypes and homologies. An exact search (using the ienum command, in the windows character‐mode version of TNT), takes 34.00 s to find 21 trees (when zero‐length branches are collapsed) of score 15. Goloboff and De Laet (2024) used two types of approaches to reduce the number of combinations, as discussed in their Appendix 2. Their shortcut of testing the pairwise compatibility between characters can be applied, with some limitations. This is the default in TNT (invoked with the xlinks&1 option). For two binary characters i and j , when states 00, 01, 10 and 11 (or their equivalents) occur among the terminal taxa, the characters are incompatible. Otherwise, it is possible to create internal links forbidding some internal reconstructions. For example, if only 00, 01 and 11 are present, the combination 10 can never occur in a most parsimonious reconstruction of any node on any tree, and can thus be eliminated by creating an internal link (invisible to the user):

xlinks=i1<j1>j0<i0>;

For a binary i and a multistate k , if all the taxa with one of the states b of the binary character i have a single state m for the multistate k , no most parsimonious reconstruction can posit the combination of b in the binary character and something other than m in the multistate, thus allowing creation of an internal link:

xlinks=ib<km>;

A problem with the approach of testing pairwise compatibility in the present context is that the characters can be aligned differently. For example, in Fig. 2, characters 1 and 2 appear to be compatible, but if 1 is a homologue of 7 and 2 of 8, they are incompatible (if 1 is a homologue of 10 and 2 of 11, the two characters are still compatible; note that 1 being a homologue of 7 and 2 of 11, or 1 of 10 and 2 of 8, are impossible combinations, because characters 1 and 2, 7 and 8, and 10 and 11 are all in the same parts). Thus, the testing of compatibility is done only between characters in reference parts (or between a character in a reference part and a character included in the complex but not in the set of characters of ambiguous homology), and making sure that each of the possible homologies still produces compatibility between the two characters. For the dataset of Fig. 2 duplicated, this reduces the number of combinations from 176 to 142.

When every part‐type has the same number of parts, then there cannot be missing parts or characters that become inapplicable when those parts are missing. The compatibility test just discussed is most fruitful in that case, as it can be applied between characters in different parts. But when some of the parts can disappear (as when there is a SING part in the dataset of Fig. 2) the situation is different: the compatibility shortcut can only be applied to characters that become inapplicable together (i.e. for the same disabling state of the same primary character). By default (i.e. under xlinks&1 ), then, when some reference parts may become inapplicable, TNT restricts application of the compatibility test to characters that belong to the same part; this is what produces the 142 combinations for Fig. 2 duplicated.

Optionally, TNT can apply the compatibility test to characters in different parts (or in one part and another character of unambiguous homology), or remove some combinations more restrictively, so as to make it less likely that the combinations eliminated will be needed for obtaining optimal scores under parsimony (although not impossible, as in rare cases optimal scores may not be achievable with the remaining combinations, for any of the options that go beyond xlinks&1 ). Two conditions can be applied here; condition (1) consists of removing combinations unexpected on the basis of compatibility, but only for the untransformed reference characters (e.g. only for the first four combinations that result from DEFINITION 3). The xlinks command allows specifying conditions for certain alignments of the parent characters, with the @ option; the links created internally for condition (1) are:

xlinks=i@i&j@j&i1<j1>i@i&j@j&j0<i0>;xlinks=i@i&k@k&ib<km>;

Condition (2) consists of applying the compatibility test only to those cases where one of the characters (or both), in all its possible alignments, has no more than a single state in one of the parts.

With xlinks&2 , the combinations are limited for characters of the same group (=part) that are compatible; for compatible characters in different groups that fulfil condition (2) only the combinations also fulfilling condition (1) are removed. For the dataset of Fig. 2 duplicated, this reduces the number of combinations (from the 142 of the previous xlinks&1 ) to 128 states. With xlinks&3 , for compatible characters in different groups that fulfil condition (2), any combination is removed on the basis of compatibility. This reduces the number of combinations to 114. With xlinks&4 , for compatible characters in different groups that fulfil condition (2), any combination is removed and for compatible characters that do not fulfil condition (2), only combinations fulfilling condition (1). This reduces the number of combinations to 111, and an exact search now takes 14.14 sec (i.e. 2.4 times faster than if not eliminating any combinations), finding (just like all the previous options) the same 21 trees of score 15.

Uninformative states in nonadditive characters (i.e. states found in a single terminal) still multiply the number of combinations. Goloboff and De Laet (2024) proposed to restrict such states to the combinations observed in the terminal having them (although they showed that this can in rare cases lead to errors). The rationale for this is that it is unlikely that uninformative states will need to be posited at internal nodes of the tree to obtain optimal scores. Sets of characters of alternative homologies also present problems in this case, since it is possible that a state is present (and uninformative) in both a reference character C , and the character X in another part that could be homologized with C . In that case, homologizing characters C and X and placing the two sister taxa with the uninformative state in C and X side by side on a tree will require the uninformative state at the common node, and possibly with combinations of states in the other characters different from those observed in either taxon. For characters of alternative homologies, then, TNT allows restricting some uninformative states under certain conditions, using internal links to specify the states and the homologies. With xlinks@3 , an uninformative state of a character C in a reference part is restricted to the observed combination of states for all the characters of the reference part‐type (but only if none of the characters that can be homologized with C has that same state), and an uninformative state of a character N in a non‐reference part is restricted to the observed combination of states for the characters in the corresponding part (but only if none of the characters it can be homologized with has that same state). With xlinks@4 , in addition to those restrictions, an uninformative state of a character C in a reference part that is also found in some other character X that can be homologized to C is restricted to the observed combinations of state for the characters in the corresponding part, only when the character is not being homologized with X . As these restrictions are done using links, it is advisable to allow for the maximum number of links, with xlinks = 32+ prior to defining the dataset. The example of Fig. 2 duplicated does not have uninformative states (all characters are binary), so this option does not reduce the number of combinations in that case, but the reduction can be substantial in other cases (depending on the number of uninformative states).

In addition to those adaptations of shortcuts originally proposed by Goloboff and De Laet (2024), another approach to reduce the number of combinations generated for sets of characters with alternative homologies is to restrict the combinations to those character states actually found in each part. By default, TNT will generate all combinations of the character states found in any homologues, but many of those occur only rarely in most parsimonious reconstructions. Restricting combinations to the states found in the parts greatly reduces the number of combinations. Consider the example of DEFINITION 3:

xlinks==ANT0/POST3=SUP6/INF9;

which generates (by default) 12 combinations. But combinations 5 and 7 have state 1 for character 9, which does not occur in the dataset of Fig. 2. Thus, one can assume that such combinations will not occur on most parsimonious reconstructions, and generate only combinations that use the observed states for each part. In TNT, this is done with the xlinks!1 option (the default is xlinks!0 , using all states). For the example just given, instead of 12 combinations, that produces 8:

POST ANT SUP INF
Combin 0 3 6 9
0 0 0 ‐‐ ‐‐
1 0 1 ‐‐ ‐‐
2 1 0 ‐‐ ‐‐
3 1 1 ‐‐ ‐‐
4 0 0 @0 @3
5 1 0 @0 @3
6 0 0 @3 @0
7 0 1 @3 @0

Observe that only state 0 is attributed to character 9 (in part INF). In specific cases, this restriction can make it impossible to calculate the optimal scores (see Fig. 8 for an example). For the dataset in Fig. 2 duplicated, this reduces the number of combinations beyond the reduction obtained with the other shortcuts (with number of combinations being 72, 68, 68, 66 and 51, respectively for xlinks&0 to xlinks&4 ). With the largest reduction ( xlinks!1 and xlinks&4 , producing 51 combinations), an exact search takes 3.73 s (9.11 times faster, relative to using all possible combinations) and still produces the exact same 21 trees.

Fig. 8.

Fig. 8

A case where application of the xlinks!1 shortcut of TNT (i.e. in every part, combining only those states observed for the part) can lead to errors (calculating a score higher than the actual minimum). The reconstruction in the left uses all possible combinations, and the optimal score (18 steps) is obtained only when assigning to some internal nodes (shown with an arrow in the tree in the right) character state combinations that have for some parts (character 5, state 1, in part Post) a state that the dataset does not display for the part in any of the taxa. When removing that combination from the list of available combinations (with the xlinks!1 option), the best score that can be obtained for the tree (with the mapping shown in the right) is 19 steps.

Appendix 2. Fragment of code, to generate “static” alignments

The full script ( heuraln.run ) can be found in the repository of TNT scripts, https://www.lillo.org.ar/phylogeny/tnt/scripts/heuraln.run. This is a simplified version (e.g. without options for verbose move, or for selecting a specific alignment). For alignment set S , the expression charalnltypes [S] returns the number of part‐types; charalnparts[S T] returns the number of parts in the part‐type T; and charalnreal[S] returns whether the alignment set has been converted into a static alignment (if not, it means that the combination passed to the last hambaln command was an invalid match between parts). The 3‐dimensional array optalns , with dimensions 10 ×  T  ×  P (where P  = number of parts at the reference part‐type) must have been declared before calling the ONE_ALN routine (so that the list of optimal alignments is accessible after running the routine); it is used for saving up to 10 optimal alignments.



     label ONE_ALN ;



     set myset %1 ;



     // Store number of part-types and parts for each type



     set ntypes charalntypes[ 'myset'];



     set nparts charalnparts[ 'myset' 0];



     var: alntry[ 10 ] partslis[ 'ntypes']numalns



          analn[ 'ntypes' 'nparts'] bestaln thisaln ;



     set bestaln 10000000 ;



     set numalns 0 ;



     // Deactivate all characters but those in this alignment set



     ccode [ . ; xgroup = 0 alnset='myset' ; ccode ]. [ { 0 } ;



     // Get numbers of groups for each type...



     loop 0 ('ntypes'-1)



       set partslis[ #1] charalnparts[ 'myset' #1];



       stop



     macfloat 0 ;



     set ntries 0 ;



     set noptalns (-1) ;



     //Find optimal alignment for the set



     combine 'ntypes' ^ ( 'partslis[0-('ntypes'-1)&44]' ) analn



        hambaln = 'myset' analn ;



        if ( charalnreal[ 'myset' ] )



            set numalns ++ ;



            set thisaln score[ 0 ] ;



            if ( 'thisaln' < 'bestaln' )



               set bestaln 'thisaln' ;



               set noptalns (-1 ) ;



               end



            if ( ( 'thisaln' == 'bestaln' ) && ( 'noptalns' < 9 ) )



               set noptalns ++ ;



               set * 'optalns[ 'noptalns' ]' analn ;



               end



            end



        endcomb



     // Now, undo this alignment and return



     if ( 'alnselect' < 0 )



        hambaln - 'myset' ; // unalign to reset ccode



        end



     proc/;



Data availability statement

Data sharing not applicable to this article as no datasets were generated or analysed during the current study.

References

  1. Agolin, M. and D'Haese, C. , 2009. An application of dynamic homology to morphological characters: direct optimization of setae sequences and phylogeny of the family Odontellidae (Poduromorpha, Collembola). Cladistics 25, 353–385. [DOI] [PubMed] [Google Scholar]
  2. Candela, A. and Rasia, L. , 2012. Tooth morphology of Echimyidae (Rodentia, Caviomorpha): homology assessments, fossils, and evolution. Zool. J. Linnean Soc. 164, 451–480. [Google Scholar]
  3. De Laet, J. and Goloboff, P. , 2024. Nothing to it: a reply to Wheeler's “much ado about nothing”. Cladistics 40, 456–467. [DOI] [PubMed] [Google Scholar]
  4. De Laet, J. , 2010. A problem in POY tree searches (and its work‐around) when some sequences are observed to be absent in some terminals. Cladistics 26, 453–455. [DOI] [PubMed] [Google Scholar]
  5. De Laet, J. , 2015. Parsimony analysis of unaligned sequence data: maximization of homology and minimization of homoplasy, not minimization of operationally defined total cost or minimization of equally weighted transformations. Cladistics 31, 550–567. [DOI] [PubMed] [Google Scholar]
  6. De Laet, J. , 2018. Anagallis: a program for parsimony analysis of character hierarchies. Computer program and documentation. https://www.anagallis.be/anagallis.
  7. De Laet, J.E. , 2005. Parsimony and the problem of inapplicables in sequence data. In: Albert, V.A. (Ed.), Parsimony, Phylogeny, and Genomics. Oxford University Press, Oxford, pp. 81–116. [Google Scholar]
  8. Ebach, M. , Williams, D. and Vanderlaan, T. , 2013. Implementation as theory, hierarchy as transformation, homology as synapomorphy. Zootaxa 3641, 587–594. [DOI] [PubMed] [Google Scholar]
  9. Endress, P. , 1994. Diversity and Evolutionary Biology of Tropical Flowers. Cambridge University Press, Cambridge, MA, p. 511. [Google Scholar]
  10. Farris, J. , 1970. Methods for computing Wagner trees. Syst. Zool. 19, 83–92. [Google Scholar]
  11. Farris, J. , 1983. The logical basis of phylogenetic analysis. In: Platnick, N.I. and Funk, V.A. (Eds.), Advances in Cladistics. Columbia University Press, New York, pp. 7–36. [Google Scholar]
  12. Farris, J. , 2008. Parsimony and explanatory power. Cladistics 24, 825–847. [Google Scholar]
  13. Fitch, W. , 1971. Toward defining the course of evolution: minimum change for a specific tree topology. Syst. Zool. 20, 406–416. [Google Scholar]
  14. Goldman, N. and Yang, Z. , 1994. A codon‐based model of nucleotide substitution for protein‐coding DNA sequences. Mol. Biol. Evol. 11, 725–736. [DOI] [PubMed] [Google Scholar]
  15. Goloboff, P. and De Laet, J. , 2024. Farewell to the requirement for character independence: phylogenetic methods to incorporate different types of dependence between characters. Cladistics 40, 209–241. [DOI] [PubMed] [Google Scholar]
  16. Goloboff, P. , 1993. Estimating character weights during tree search. Cladistics 9, 433–436. [DOI] [PubMed] [Google Scholar]
  17. Goloboff, P. , 1998. Tree searches under Sankoff parsimony. Cladistics 14, 229–237. [DOI] [PubMed] [Google Scholar]
  18. Goloboff, P. , 2014. Extended implied weighting. Cladistics 30, 260–272. [DOI] [PubMed] [Google Scholar]
  19. Goloboff, P. , 2022. Phylogenetic Analysis of Morphological Data. Vol. 1. From Observations to Optimal Phylogenetic Trees. CRC Press, Taylor and Francis Group, Species and Systematics Series, Boca Raton, FL, p. 277. [Google Scholar]
  20. Goloboff, P. , 2025. Phylogenetic analysis of characters with dependencies under maximum likelihood. Syst. Biol. 75, 277–295. 10.1093/sysbio/syaf051 [DOI] [PubMed] [Google Scholar]
  21. Goloboff, P. , De Laet, J. , Ríos‐Tamayo, D. and Szumik, C. , 2021. A reconsideration of inapplicable characters, and an approximation with step‐matrix recoding. Cladistics 37, 596–629. [DOI] [PubMed] [Google Scholar]
  22. Grams, M. , Torres, A. , Wirkner, C. and Richter, S. , 2025. A new morphological phylogeny of Malacostraca comparing the application of character dependencies and implied weighting. Cladistics 41, 283–303. [DOI] [PMC free article] [PubMed] [Google Scholar]
  23. Grant, T. and Kluge, A. , 2009. Perspective. Parsimony, explanatory power, and dynamic homology testing. Syst. Biodivers. 7, 357–363. [Google Scholar]
  24. Holmes, S. , 2005. Statistical approach to tests involving phylogenies. In: Gascuel, O. (Ed.), Mathematics of Evolution and Phylogeny. Oxford University Press, Oxford, pp. 91–120. [Google Scholar]
  25. Hoyal Cuthill, J. and Lloyd, G. , 2024. Measuring homoplasy I: comprehensive measures of maximum and minimum cost under parsimony across discrete cost matrix character types. Cladistics 41, 1–27. [DOI] [PMC free article] [PubMed] [Google Scholar]
  26. King, B. and Rücklin, M. , 2020. A Bayesian approach to dynamic homology of morphological characters and the ancestral phenotype of jawed vertebrates. elife 2020, e62374. [DOI] [PMC free article] [PubMed] [Google Scholar]
  27. Kluge, A. and Grant, T. , 2006. From conviction to anti‐superfluity: old and new justifications of parsimony in phylogenetic inference. Cladistics 22, 276–288. [Google Scholar]
  28. Leardi, J. , Pol, D. and Fernández, M. , 2012. The antorbital fenestra of Metriorhynchidae (Crocodyliformes, Thalattosuchia): testing its homology within a phylogenetic framework. J. Vertebr. Paleontol. 32, 490–494. [Google Scholar]
  29. Lee, M. , 1998. Similarity, parsimony and conjectures of homology: the chelonian shoulder girdle revisited. J. Evol. Biol. 11, 379–387. [Google Scholar]
  30. Maddison, W.P. , 1993. Missing data versus missing characters in phylogenetic analysis. Syst. Biol. 42, 576–581. [Google Scholar]
  31. Manton, S. , 1977. The Arthropoda: Habits, Functional Morphology, and Evolution. Oxford University Press, Oxford, p. 527. [Google Scholar]
  32. Platnick, N. , 1978. [Book review of] The Arthropoda: habits, functional morphology, and evolution. Syst. Zool. 27, 252–255. [Google Scholar]
  33. Ramírez, M. , 2007. Homology as a parsimony problem: a dynamic homology approach for morphological data. Cladistics 23, 588–612. [DOI] [PubMed] [Google Scholar]
  34. Rieppel, O. , 1996. Testing homology by congruence: the pectoral girdle of turtles. Proc Biol Sci. 263, 1395–1398. [DOI] [PubMed] [Google Scholar]
  35. Sankoff, D. and Rousseau, P. , 1975. Locating the vertices of a Steiner tree in an arbitrary space. Math. Program. 9, 240–246. [Google Scholar]
  36. Shubin, N. , Tabin, C. and Carroll, S. , 1997. Fossils, genes and the evolution of animal limbs. Nature 388, 639–643. [DOI] [PubMed] [Google Scholar]
  37. Varón, A. , Vinh, L. and Wheeler, W. , 2010. POY version 4: phylogenetic analysis using dynamic homologies. Cladistics 26, 72–85. [DOI] [PubMed] [Google Scholar]
  38. Wheeler, W. , 1999. Fixed character states and the optimization of molecular sequence data. Cladistics 15, 379–385. [DOI] [PubMed] [Google Scholar]
  39. Wheeler, W. , 2003. Search‐based optimization. Cladistics 19, 348–355. [PubMed] [Google Scholar]
  40. Wheeler, W. , 2012. Trivial minimization of extra‐steps under dynamic homology. Cladistics 28, 188–189. [DOI] [PubMed] [Google Scholar]
  41. Wheeler, W. , 2023. Much ado about nothing: inapplicable data as insertion–deletion events. Cladistics 39, 475–478. [DOI] [PubMed] [Google Scholar]
  42. Williams, D. and Ebach, M. , 2018. A Cladist is a systematist who seeks a natural classification: some comments on Quinn (2017). Biol. Philos. 33, 1–4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  43. Williams, D. and Gill, A. , 2025. Norman Platnick, the development of cladistics, ‘integrative’ taxonomy, and modern monography. In: Williams, D. and Wheeler, Q. (Eds.), The New Taxonomy: A science reimagined. CRC Press, Boca Ratón, pp. 14–41. [Google Scholar]
  44. Xu, X. , Clark, J. , Mo, J. , Choiniere, J. , Forster, C. , Erickson, G. , Hone, D. , Sullivan, C. , Eberth, D. , Nesbitt, S. , Zhao, Q. , Hernandez, R. , Jia, C. , Han, F. and Guo, Y. , 2009. A Jurassic ceratosaur from China helps clarify avian digital homologies. Nature 459, 940–944. [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

Data sharing not applicable to this article as no datasets were generated or analysed during the current study.


Articles from Cladistics are provided here courtesy of Wiley

RESOURCES