Skip to main content
Biophysical Journal logoLink to Biophysical Journal
. 2024 Sep 10;123(20):3558–3568. doi: 10.1016/j.bpj.2024.09.008

Blocking uncertain mispriming errors of PCR

Takumi Takahashi 1, Hiroyuki Aoyanagi 1, Simone Pigolotti 2, Shoichi Toyabe 1,
PMCID: PMC11494492  PMID: 39257000

Abstract

The polymerase chain reaction (PCR) plays a central role in genetic engineering and is routinely used in various applications, from biological and medical research to the diagnosis of viral infections. PCR is an extremely sensitive method for detecting target DNA sequences, but it is substantially error prone. In particular, the mishybridization of primers to contaminating sequences can result in false positives for virus tests. The blocker method, also called the clamping method, has been developed to suppress mishybridization errors. However, its application is limited by the requirement that the contaminating template sequence be known in advance. Here, we demonstrate that a mixture of multiple blocker sequences effectively suppresses the amplification of contaminating sequences even in the presence of uncertainty. The blocking effect was characterized by a simple model validated by experiments. Furthermore, the modeling allowed us to minimize the errors by optimizing the blocker concentrations. The results highlighted an inherent robustness of the blocker method in that fine-tuning the blocker concentrations is not necessary. Our method extends the applicability of PCR and other hybridization-based techniques, including genome editing, RNA interference, and DNA nanotechnology, by improving their fidelity.

Significance

The applications of PCR are increasing day by day, and there is a need to suppress PCR errors to improve the accuracy of PCR-based techniques and broaden their applicability. The blocker method has been developed to substantially suppress mispriming. However, the method requires prior knowledge of the contaminating sequence, which limits its applicability. We successfully demonstrate that adding a combination of multiple blocker sequences can substantially suppress PCR errors, even when we have only partial information about the contaminating sequences. We also construct a biophysical model of the blocking effect, which allows us to find the optimal blocker combinations that minimize the PCR error. Since the method targets hybridization, it is readily applicable to a wide range of biotechnologies.

Introduction

Polymerase chain reaction (PCR) is a technique for amplifying DNA strands characterized by specific sequences (1). In PCR, a pair of short strands, called primers, are mixed in the reaction solution. They specify the sequence region to be replicated and trigger DNA replication. By running several replication cycles, PCR is able to exponentially amplify a DNA sample. This exponential amplification makes PCR extremely sensitive but also error prone. On the one hand, the incorporation of incorrect nucleotides during polymerase extension can be significantly suppressed by using a high-fidelity polymerase and carefully tuning the reaction conditions. On the other hand, errors can be caused by mishybridization of primers onto contaminating strands (2) (Fig. 1 a). In practical applications, test samples are often far from being pure and contain a variety of DNA sequences. Primers readily hybridize to contaminating sequences similar to the target sequence, resulting in false positives in viral testing.

Figure 1.

Figure 1

PCR and blocker method. (a) Mishybridization of the primer strands (P) to the contaminating template strands (W) results in errors. (b) Blocker method. The sequence of the blocker (B) is designed to be complementary to the sequence of W and blocks the primer binding to W. The two-base floating end is attached at the 3′ end of B to prevent priming at this end. (c) Different sequences of Wi and Bj are labeled so that i=j when Wi and Bj are complementary. The number of mismatches changes depending on the relative positions of the mutation in W and the target mutation in B.

The blocker, or clamping, method was developed to suppress mispriming in PCR (3,4,5,6). (Fig. 1 b). In this method, a nucleic acid sequence called a blocker is added to the PCR mixture. The blocker sequence is designed to be complementary to the contaminating sequence. It suppresses errors by blocking mishybridization of the primer to the contaminating sequence (Fig. 1 c). A chimeric strand of DNA and locked nucleic acids (LNAs) (7) or DNA and peptide nucleic acids (8) are often used as blockers due to their high specificity (3,9,10). In our previous work, we established that blocking combines energetic and kinetic discrimination toward correct hybridization to the target sequence (10). Because of this mechanism, error suppression remains effective even at relatively low temperatures, and therefore, it is expected to be applicable to thermally sensitive systems such as biological cells. Since the blocker sequence is designed to be complementary to the contaminating sequence, the blocker method requires knowing the contaminating sequence in advance. This limitation can be quite severe in realistic situations where there are multiple possible contaminating sequences but we do not know which sequences are contaminating each sample.

This work extends the blocker method to situations where the information about contaminating sequences is limited. For this purpose, we demonstrate the effectiveness of the blocker combinations in suppressing errors. At the same time, we construct a simple biophysical model and validate the model through extensive experiments to characterize the error suppression by blocker combinations. Finally, we attempt to minimize the errors by optimizing the blocker concentrations.

We specifically consider the following scenario. A PCR mixture contains a target sequence to be detected, R, and a contaminating sequence, Wi, randomly sampled from N possible sequences with probability pi (1iN). The primer P is complementary to the primer-binding region of R. The contaminating sequence has a corresponding sequence region with single or multiple mutations, which cause mismatches in the hybridization with P. The blockers are labeled as Bi, such that Bi is perfectly complementary to Wi (Fig. 1 c).

Materials and methods

The experiments are performed using the same procedure as that in a previous work (10). The program code (Python) used for finding the optimal blocker concentrations is available on https://github.com/stoyabe/PCR_Blocker.

The experiments consist of two PCR steps (Fig. S1). In the first PCR step, R and W are amplified by linear PCR. In the second PCR step, the amount of the product in the first step, R¯ and W¯, is evaluated by a quantitative PCR (qPCR).

Linear PCR experiment

In linear PCR, the product amount is amplified linearly because a single primer sequence is used, which simplifies the error analysis (supporting materials and methods S1).

DNA strands were synthesized by Eurofins Genomics (sequences are listed in Table S1). DNA/LNA chimeric strands were synthesized by Aji Bio-Pharma (San Diego, CA, USA). The reaction mixture for polymerization contained 25 units/mL Hot-start Taq DNA polymerase (New England Biolabs, Ipswich, MA, USA), Taq standard reaction buffer, 2.5 nM R, 2.5 nM W, 100 nM P, and indicated amounts of blocker strands. We performed initial heating for 30 s at 95°C and then 10 cycles of 15 s at 95°C, 30 s at 60°C, and 5 s at 68°C using a PCR cycler (Bio-Rad, Hercules, CA, USA). Immediately after the cycles, the mixture was cooled down on ice to stop the enzyme reaction and diluted for quantification.

Quantification

The second qPCR step was performed on a real-time PCR cycler (Bio-Rad) after the linear PCR experiment for quantifying R¯ and W¯ (see supporting materials and methods S2). The reaction mixture contains Luna Universal qPCR Master Mix (New England Biolabs), 200 nM each of the primers, and the diluted sample. The dilution rate of the sample was 1/250 in the final concentration. The thermal cycle consisted of initial heating for 60 s at 95°C, 40 cycles of 15 s at 95°C, 30 s at 66°C, and 5 s at 72°C.

Optimization of the blocker concentrations

We want to minimize εmean and εmax given by Eq. 10 under the constraint that bi0 for i and jbj=btot. It is difficult to obtain the analytical solution under the constraints given by inequalities. Instead, we numerically find the optimum by a gradient descent method. For validating the method’s effectiveness, we verify that εmean and εmax are convex functions of {bi}.

  • 1)

    Minimization of εmean.

To minimize εmean, we first note that the domain {bk0k;jbj=btot} is convex, that is, a point between any two points in the domain lies in the same domain. Moreover, the Hessian matrix given by

Hm,k2εmeanbmbk=ipiε˜i/(Ki,mKi,k)(jbj/Ki,j+1)3 (1)

is of the form H=KTDK, where D is a diagonal matrix with positive entries. Here, K is a matrix with the entries Ki,m. This implies that the Hessian matrix is nonnegative, hence the function f is convex.

The fact that our minimization problem concerns a convex function in a convex domain ensures that the minimum is unique. We can, therefore, find this minimum very efficiently numerically using a gradient descent method. In particular, calling Fm=εmeanbm, the optimal solution is the unique stable fixed point of the replicator dynamics:

dbmdt=γbm(FmF¯),m=1,2,,N, (2)

where F¯=mbmFm/btot is the average gradient. The coefficient γ determines the update rate of {bm}. The second term in the right-hand side of Eq. 2 ensures that the normalization condition jbj=btot is preserved by the dynamics.

Substituting the explicit expression of the gradient into Eq. 2, we obtain

dbmdt=γbmi[piε˜i(jbj/Ki,j+1)2(1Ki,m1btotjbjKi,j)]. (3)

We used the replicator Eq. 3 to obtain the optimal blocker concentrations that minimize the average error fraction εmean.

  • 2)

    Minimization of εmax.

The minimization of εmax (Eq. 10) is also possible in a similar manner to εmean. We note that εmax({bk}) is a convex function as well since the maximum of convex functions is a convex function. Let i be the index with the maximum εi. Then, the replicator equation reads

dbmdt=γbmε˜i(jbj/Ki,j+1)2(1Ki,m1btotjbjKi,j),i=argmaxiεi. (4)

We note that, in the numerical integration of Eq. 5, one has to dynamically update i, i.e., compute at each time step the value of i that maximizes εi.

Results

Model

Our method is based on adding a combination of blocker sequences {Bi} that are complementary to all the possible contaminating sequences {Wi}. Throughout this paper, we use {·} to denote a set. The primer we use in this work is 20 bases long, which is typical for PCR virus tests. Therefore, there exist in principle 320 possible W sequences. However, the binding strength of P to W with two or more mismatches is negligible in standard PCR assays. Therefore, in practice, we only need to worry about the contaminating sequences with a single mismatch with respect to R, which are limited to 3×20=60 sequences.

We define a growth rate r as the number of strands copied in a PCR cycle divided by the number of template strands. With this definition, perfect PCR efficiency corresponds to r=2. We call rR and rWi the growth rates of R and Wi, respectively. Then, the error fraction is defined as (11)

εi=rWirR. (5)

Once mispriming occurs, the product works as a normal template in successive cycles. That is, the product contains the correct primer-binding sequence. Therefore, the mispriming typically happens only for the contaminating sequences and not for the PCR products. That is, the overall error of the PCR assay should show not a power-law-like but linear dependence on εi.

We call bi the concentration of Bi and εi({bk}) the error fraction associated with Wi in the presence of a blocker distribution {bk}. We evaluate the error by either the mean error fraction or the maximum error fraction, defined as

εmean({bk})=ipiεi({bk}),εmax({bk})=maxiεi({bk}), (6)

respectively, under the constraint that the total concentration btot=ibi is fixed. This constraint reflects the limitation that too many blocker strands also suppress the replication of R, as we will see (Fig. 2 a). The choice between focusing on εmean or εmax depends on the purpose of the PCR assay. One may want to minimize εmax in diagnosis to avoid false positives. For DNA cloning by PCR, reducing εmean may increase the product amount of the right sequence more than reducing εmax, although this may depend on the system.

Figure 2.

Figure 2

Error suppression by a single blocker sequence. (a) The production amounts of R and W1 per cycle in the presence of blocker sequence B1. (b–d) Dependence of the error fractions ε on blocker concentration for different combinations of W and B, where the target mutations of B are all at the same position in W. The subscript i of Wi is used for labeling different W sequences. Bi is the blocker sequence complementary to Wi. The expression, for example, W1(10AT) indicates that A, the tenth base from the 3′ end in R, is mutated to T in W1. The solid curves are fitting curves by Eq. 8 with the fitting parameter Ki,j. (e and f) Error suppression including the situations where the positions of the target mutations of B are different from the mutated position of W. The concentration of the blocker is 2 μM. Error bars indicate standard deviations. Each data point represents an average over three independent experiments.

We consider a biophysical model for evaluating εi({bk}). First, we consider the situation with a single B sequence and then extend the model to multiple B sequences. In experiments, the blocker concentration is relatively high (μM). We therefore assume that the blocker binding and dissociation are always equilibrated (10). Then, the growth rates are given by

rWi(bj)=r˜Wibj/Ki,j+1,rR(bj)r˜R. (7)

Here, Ki,j is the dissociation constant associated with the binding of Wi and Bj. The factor 1/(bj/Ki,j+1) represents the fraction of the free Wi that is not bound by Bj. We assumed that bj is not extremely large. r˜Wi is the growth rate in the absence of the blocker sequences. We assume that the primer then binds to free Wi, initiating the polymerase extension. Because the binding of Bj to R is negligible under the present condition (Fig. 2 a), the replication of Ri is not blocked by the blockers; rR(bj)r˜R.

Combining Eqs. 5 and 7, we obtain an expression for the error fraction in the presence of a blocker Bj:

εi(bj)=ε˜ibj/Ki,j+1. (8)

Here, ε˜i=r˜Wi/r˜Ri is the error fraction in the absence of blocker sequences. Thus, the model predicts that the blocking strength of Wi by Bj is essentially determined by Ki,j. In the presence of multiple blockers, Eq. 8 becomes

εi({bk})=ε˜ijbj/Ki,j+1. (9)

Since bj and Ki,j are nonnegative, Eq. 9 implies that εi({bk})ε˜i. As expected, the error is always reduced by the addition of blockers.

Experiments

We performed experiments to characterize the effect of the blockers on the error fraction, validate Eqs. 8 and 9, and evaluate ε˜i and Ki,j. We performed a linear PCR using a single primer instead of a primer pair since the error fraction is directly evaluated from the product amounts in the linear PCR (Fig. S1). The reaction mixture contains 2.5 nM R template, 2.5 nM Wi template, 100 nM primer, thermostable (Taq) DNA polymerase enzyme, and standard buffer mixture. Blocker sequences are added depending on the experiments. See the materials and methods and supporting material for details.

We use W sequences with mutations at the positions where the original nucleotide of R is either A or T. Mutations at A and T cause more frequent errors than those at C and G positions due to their lower binding stability. The reason for this choice is to focus on the worst cases for practical assays.

Dependence of PCR error on blocker concentration

We measured the amount of the products in the presence of a blocker sequence (Fig. 2 a) to evaluate the dependence of the error fractions on the blocker concentrations for different combinations of Wi and Bj (Fig. 2, bd). The error fraction εi(bj) decreased significantly with bj for the complementary pair i=j. This dependence is well fitted by Eq. 8, supporting the validity of the model.

We distinguish two cases for noncomplementary pairs of Wi and Bj (ij). If the mismatch positions of i and j are the same, then the blocker still suppresses the error (Fig. 2, bd). This is because the binding of Bj to Wi is not negligible in this case. In particular, purine-pyrimidine pairs (A and C) and (G and T) bind weakly even though they are not complementary. This cross talk effect between B and W explains the tendency of the error fractions in Fig. 2, bd. For example, in Fig. 2 b, B2 suppresses errors more than B3. This tendency is due to the stronger binding of B2, which has G at the corresponding mutation position, to W1 than B3, which has C. The error fraction without blocker is the highest in Fig. 2 d because G, the mutation in W3, and T of the primer at the corresponding position bind to some extent. Eq. 8 fits the results well for such pairs.

When the mismatch positions of i and j are different, error suppression is negligible (Fig. 2, e and f). This is because the number of mismatched nucleotides is two (see, for example, Wa and Bc in Fig. 1 c) and their hybridization is unstable. This may not be the case if the two mismatches are far apart (see Wa and Bd in Fig. 1 c). In such cases, however, the primer-binding region is not sufficiently covered by the blocker, and blocking is again not effective.

When the blocker concentration is higher than 10 μM, the replication of R is also suppressed (Fig. 2 a). Therefore, we used btot= 10 μM in the following experiments.

Estimation of dissociation constants Ki,j from sequences

Directly measuring Ki,j for all 60×60 pairs of Wi and Bj is impractical. Based on the results in Fig. 2, e and f, we set Ki,j= for noncomplementary pairs with different mismatch positions. For other pairs, we estimated the Ki,j values by the nearest-neighbor method (9,12,13) with LNA parameters including mismatches (Fig. 3). We found that the values estimated by the nearest-neighbor method are similar to those obtained by the fitting of Eq. 8 in Fig. 2, bd.

Figure 3.

Figure 3

Experimental versus predicted values of the dissociation constants Ki,j. The experimental values Ki,j are obtained by fitting Eq. 8 to the experimental curves in Fig. 2, bd. The predicted values are the estimation from the sequences by the nearest-neighbor method (9,12,13). We used the BioPython package with modifications (14). The dashed line indicates the equality between the experimental values and predicted values. The solid line is a fit by a curve linear in the log scale: lny=alnx+b with a=0.619 and b=4.38. Here, x and y are in the units of M.

Broadly, this result supports the validity of the model given by Eq. 8 for the blocking effect. However, there are small deviations between the experiments and predictions. The deviations could be due to the fact that the parameters of the nearest neighbor method are not optimized for our experimental conditions. For better estimation, we used a linear fitting in the log scale (solid red curve in Fig. 3) to calibrate the Ki,j values estimated by the nearest-neighbor method.

Estimation of error fraction without blocker ε˜i

We measured the dependence of ε˜i on the position and nucleotide species of the mismatch of W (Fig. 4). We observed a tendency for ε˜i to be small at both the 5′ and 3′ ends of the primer-binding region of W. This is because a mismatch near the ends causes the ends to flap and reduces stability more than internal mismatches. In particular, rigid hybridization at the 5′ end of the primer-binding region of W (3′ end of the primer) is necessary for the polymerase to initiate extension (15,16,17).

Figure 4.

Figure 4

Dependence of error fraction ε˜ on the mutation position and the nucleotide species of W in the absence of the blockers. Solid line is a fitting curve f(x)=p[1+tanh(q(xr))][1tanh(s(xu))] with the fitting parameters p=0.12, q=2.6, r=2.1, s=0.50, and u=16. We chose this function since it can reasonably model the observation that the error fraction is small at the edge and similar in the bulk sequences. Each symbol corresponds to a different nucleotide type in the W sequence. Error bars indicate standard deviations. The number of independent experiments is three for each point. The sequences of R and P are shown on the top for the reference.

The value of the parameter ε˜i depends on the nucleotide species (Fig. 4). However, this dependence is hard to model accurately. For example, the sequence near the mismatch position should affect ε˜i. Measuring all 60 values of ε˜i is not practical. We could calculate Ki,j from the sequences because the blocker concentration is sufficiently large, and the equilibrium is expected for the blocker binding. However, equilibrium is not expected for primer binding because the primer concentration is not large enough (10). Therefore, it is not straightforward to estimate ε˜i only from the sequences. As a reference, the free energy difference ΔΔG between the hybridization P:R and P:Wi are calculated by the nearest-neighbor method (12,13) (Fig. S3). The results show a similar tendency to ε˜i in Fig. 4. This implies that ε˜i depends largely on the hybridization stability.

Here, we assume a simple model where ε˜i depends only on the mismatch position and neglect the dependence on the nucleotide species. To account for the flat dependence on position in the internal region, we fit ε˜i with a combination of two hyperbolic tangent functions (Fig. 4). We, however, observed a significant variation around the fitted line. We will see how the variation affects the optimization of the blocker concentrations.

Suppression of errors with a specific error by blocker combination

In order to test the effectiveness of the blocker combinations and the model (Eq. 9), we contaminated one of three specific W sequences and evaluated error suppression by a combination of three blocker sequences targeting the three W sequences (Fig. 5, ac). The sum of blocker concentrations was kept at jbj= 10 μM. We only tested the W sequences with the same mutation positions since blocking is ineffective if the target position of B differs from the mutation position of W (see Fig. 2, e and f). We found a gradient of the error fraction in a ternary diagram. The steepest gradient is obtained along the axis of the blocker complementary to the contaminating sequence (i=j), as expected.

Figure 5.

Figure 5

Error suppression by combinations of multiple blocker sequences. Experimental data (top) and corresponding model calculation by Eq. 9 (bottom). The error fractions without blockers ε˜ are 0.52 (a), 0.19 (b), 0.75 (c), and 0.48 (d–f). The numbers in the circles in (a)–(c) are the mean error fractions of three independent experiments. The blockers highlighted with red are complementary to the contaminating sequences. btot= 10 μM. The plots are colored according to the error fractions normalized by the maximum error fraction in each plot.

We now evaluate εi({bj}) by Eq. 9 using Ki,j calculated by the nearest-neighbor method with calibration (Fig. 3) and ε˜i modeled in Fig. 4. We found that Eq. 9 qualitatively reproduces the experimental results under different conditions (Fig. 5, df). The absolute values show some discrepancy between the experimental results and the model calculations. The main source of the discrepancy seems to be the approximation by neglecting the dependence of ε˜i on the nucleotide species (Fig. 4).

Suppression of errors with partial information by blocker combination

We test if the blocker combinations are still effective when the information about the contaminating sequences is limited. We apply this approach to three situations with different mismatch patterns (Fig. 6), where we have partial prior knowledge of the errors. In both examples, we considered the worst case, where the contamination probability is uniform, pi=1/N. The results are summarized in Table 1. In each scenario, we calculated εmean and εmax from experimentally measured εi using Eq. 6.

  • (i)

    The mutation is located in one of the three positions. The nucleotide species is specified (Fig. 6, a and b).

Figure 6.

Figure 6

εmean and εmax values in the presence of blocker combinations obtained by experiments. Mutations are located at different positions (a and b) or same position with different nucleotides (c and d). The blockers bi are complementary to Wi. W1: 10A T, W2: 10A C, W3: 10A G, W6: 5T A, W7: 17A T. We measured εi({bk}) and plot εmean=iεi({bk})/3 and εmax=maxiεi({bk}). The circles indicated by green correspond to the uniform blocker concentrations. The stars indicate the optimal blocker concentrations calculated by the model. The corresponding error fractions are 0.038 (a), 0.069 (b), 0.072 (c), and 0.20 (d), which are calculated by a linear interpolation. See Figs. 7, a and b, and S4, a and b, for corresponding model calculations.

Table 1.

Summary of error suppression

Blocker εmean
εmax
None Uniform Optimal Ratio (none/optimal) None Uniform Optimal Ratio (none/optimal)
(i) Experimenta 0.27 0.040 0.038 7.1 0.49 0.073 0.069 7.1
Modela 0.36 0.047 0.040 9.0 0.48 0.089 0.049 10
(ii) Experimentb 0.48 0.071 0.072 6.7 0.75 0.16 0.20 3.7
Modelb 0.48 0.036 0.035 14 0.48 0.066 0.044 11
(iii) Modelc 0.33 0.20 0.18 1.8 0.48 0.47 0.34 1.4

εmean and εmax without blocker (none), with blocker combinations with the uniform concentrations (uniform, bk=btot/N), or with the optimal concentrations (optimal, bk).

a

The mutation is located in one of the three positions. The nucleotide species is specified (Fig. 6a and b).

b

The mutation is located in the specified position, but the nucleotide species is not identified (Fig. 6c and d).

c

Neither position nor nucleotide species are specified. The experimental values are obtained by interpolating the blocker concentrations in Fig. 6.

In this scenario, Bj has a negligible error suppression effect on Wi that does not match (ij), as seen in Fig. 2, e and f. Hence, the effect of the blockers can be decomposed as εi({bk})=εi(bi), leading to a simple superposition εmean=ipiε(bi).

Taking a uniform blocker concentration bi=btot/N as a reference (indicated by green circles in Fig. 6), we have a significant error reduction, with reduction ratios of 6.8 for εmean and 6.7 for εmax (Table 1). This significant error suppression shows the effectiveness of our approach.

  • (ii)

    The mutation is located in the specified position, but the nucleotide species is not identified (Fig. 6, c and d).

As shown in Fig. 2, bd, Bj can suppress the replication of Wi even if ij when the mutations of i and j are located at the same position. Hence, the error suppression is more complicated than in the case of scenario (i). With uniform blocker concentrations, we obtained significant reduction ratios of 6.8 for εmean and 4.7 for εmax (Table 1).

Thus, in these distinctive scenarios with partial information about the contaminating sequences, the addition of blocker combinations significantly reduced errors, supporting the practical effectiveness of the method.

We calculated the error pattern by the model. By substituting Eq. 9 to Eq. 6, we obtain

εmean({bk})=ipiε˜ijbj/Ki,j+1,εmax({bk})=maxiε˜ijbj/Ki,j+1. (10)

The parameters ε˜i and Kij could be calculated by the models in Figs. 3 and 4, respectively, for the given primer sequence.

Figs. 7, a and b, and S4, a and b, show the model calculations corresponding to Fig. 6, ac and b and d, respectively. The model qualitatively reproduced the error pattern obtained by experiments. We observed a quantitative discrepancy for εmax. This is possibly due to neglecting the nucleotide dependence of ε˜. In this particular scenario, this assumption could significantly affect the estimation. The estimation of the parameters can be improved for each system if necessary.

Figure 7.

Figure 7

Optimal blocker concentrations minimizing εmean given by (6) under the constraint btot= 10 μM. See Fig. S4 for the optimal blocker concentrations minimizing εmax. (a and b) There are possible three mutations at different positions (a) or same position (b). The values of εmean are indicated by colors normalized by the maximum error rate in each plot. The optimal blocker concentrations are indicated by stars. (c) Optimal blocker concentrations for the case with 60 possible errors. The sequences of R and P are shown on the top for the reference. The expressions b(N) (N is A, T, C, or G) indicate the optimal concentrations of the blockers targeting N mutation at each position. εmean without blockers are 0.36 (a), 0.48 (b), and 0.33 (c). The minimum εmean at the optimal blocker concentrations are 0.039 (a), 0.035 (b), and 0.18 (c).

Optimal blocker combinations

As seen in Fig. 6, εmean and εmax have a gradient, suggesting the possibility of minimizing errors by optimizing blocker concentrations. We theoretically derive the optimal blocker distribution {bk} that minimizes the mean or maximum error fraction based on Eq. 10 (see materials and methods). We obtain {bk} numerically under the constraints btot=jbj= 10 μM and bj0 for j. It can be shown that εmean({bk}) and εmax({bk}) are convex functions of {bk} for arbitrary {ε˜i} and {Ki,j}. This implies that we can efficiently find {bk} by a gradient descending method without risking to be stuck in a local minimum. Actually, the εmean and εmax have single minima (Fig. 7, a and b).

We apply this approach to three situations with different mismatch patterns (Figs. 7 and S4). In addition to scenarios (i) and (ii) introduced above, where we have partial prior knowledge of the errors, we tested (iii) a scenario with no prior knowledge. In all examples, we considered the worst case, where the contamination probability is uniform, pi=1/N. The results are summarized in Table 1.

In scenarios (i) and (ii), the model calculation shows significant error reduction by the optimization (Table 1), implying the effectiveness of the optimization. The obtained optimal blocker combinations are indicated by a star in Fig. 6. The estimated reduction ratio in the experiments was summarized in Table 1. The improvement by the optimization is estimated to be limited. The major reason for this is that the uniform blocker concentrations are already very effective. The addition of blockers with concentrations larger than 3 μM can effectively suppress errors (Fig. 2, bd). We used btot= 10 μM, corresponding to 3.3 μM for each blocker in scenarios (i) and (ii) for the uniform blocker concentrations. Thus, we can add a sufficient amount of blockers in the reaction buffer without affecting the amplification of the target sequences. In conclusion, an advantage of the method is that we can significantly suppress errors even without fine optimization of the blocker concentrations. Nevertheless, if an application of the method requires it, then the theory we introduced permits finely optimizing error suppression.

Finally, in order to see the limitation of the blocker method, we estimated the error suppression with a mixture of all 60 blocker sequences, assuming that there are 60 error possibilities (Figs. 7 c and S4 c). We found that the blocker method is still effective, and the optimal blocker combination reduced εmean and εmax with reduction ratios of 1.8 and 1.4, respectively (Table 1). Although scenario (iii) may not be relevant for practical virus testing, it illustrates that the blocker addition always reduces errors even in this worst-case scenario.

The optimal blocker distribution {bj} (Fig. 7 c) can be explained by the dependence of ε˜i on the position and the dependence of Ki,j on the sequence combinations. In particular, bj are small at the edges because ε˜i are smaller at the edges than at the internal positions (as shown in Fig. 4). Therefore, the priority to target the edges with blockers is low.

The concentrations {bj} optimizing εmean are strongly biased toward targeting C and G mutations (Fig. 7 c). Since C and G bind more strongly than A and T, corresponding Ki,j values are typically small. Therefore, targeting C and G mutations is more effective in reducing εmean under the constraint of a fixed btot. This tendency is quantified by the blocking efficacy αj of the blockers, defined as αj=i(1/Kij) (Fig. S5 a). bj for εmean increases with αj at small αj. The optimal concentrations bj for εmean decrease at large αj because a small amount is enough for the blockers with large αj to suppress errors.

In contrast, {bj} for εmax is strongly biased toward targeting A and T mutations (Fig. S4 c). Accordingly, bj for εmax decreases monotonically with αj (Fig. S5 b). In particular, bj is roughly proportional to 1/αj (Fig. S5 b, inset). This tendency is consistent with intuition as well. If εi of a particular template Wi is larger than other templates, we need to increase bi to reduce εmax. The error εi is large when Bi is not effective and, therefore, has a small αi value. Hence, bj for the blockers with small αj should be large to minimize εmax, which results in a uniform εi distribution.

Discussion

We demonstrated the effectiveness of the blocker method in PCR with limited information about the contaminating sequences. Our methodology is based on adding combinations of multiple blocker sequences informed by biophysical considerations. In particular, we successfully modeled the effect of the blockers and obtained the relevant parameters. With this approach, we were able to reduce the mean and maximum error fractions even in the worst-case scenario where all 60 possible contaminations are present (Table 1).

Adding blockers does not increase the error rate unless sufficiently large concentrations of blocker sequences inhibit the replication of R (Fig. 2 a). This means that even though the perfect optimization achieves the minimum error fraction, an inaccurate optimization is still effective. In fact, the average error εmean is also reduced by uniform blocker concentrations bi=btot/N (Table 1). This tolerance is an inherent advantage of the blocker method because we can reduce the experimental cost at the expense of reducing the error suppression effect. For example, we neglected the dependence of ε˜i on the nucleotide species and considered only the mismatch position (Fig. 4). The Ki,j values are estimated from the sequences by the nearest-neighbor method with a calibration (Fig. 3). While these approximations may reduce the error suppression effect, they significantly lower the burden of evaluating parameters, ε˜ and Ki,j. Thus, in applications, assuming that the model of ε˜i and the calibration curve of Ki,j obtained in this work are applicable to other primer sequences, one does not need additional calibration for each system. If necessary, one can further reduce errors by evaluating the Ki,j and ε˜i values in each system.

Optimization based on biophysical modeling has not been explored so far in genetic technologies and is itself of conceptual interest. However, for some applications, optimization could bring significant benefits. As mentioned above, the blocker method may be effective in thermally sensitive systems such as genome editing (18) and RNA interference (19) in biological cells, where temperature tuning for increasing hybridization specificity is limited. The blocker method targets nucleic acid hybridization and could apply to these gene technologies. However, the maximum blocker concentration in cells may be restricted. In this case, optimization of the blocker concentrations may be essential to suppress errors.

Data and code availability

Data are available from the authors upon reasonable request. The program code (Python) used for finding the optimal blocker concentrations is available at https://github.com/stoyabe/PCR_Blocker.

Acknowledgments

S.T. is supported by JSPS KAKENHI grant number JP23K17658 and JST ERATO grant number JPMJER2302, Japan. S.P. is supported by KAKENHI grant number JP23H01146.

Author contributions

T.T. did the experiments. T.T., H.A., S.P., and S.T. developed the theory, designed the system, and wrote the paper.

Declaration of interests

The authors declare no competing interests.

Editor: Jie Yan.

Footnotes

Supporting material can be found online at https://doi.org/10.1016/j.bpj.2024.09.008.

Supporting material

Document S1. Supporting materials and methods, Figures S1–S5, and Tables S1 and S2
mmc1.pdf (227.6KB, pdf)
Document S2. Article plus supporting material
mmc2.pdf (2.4MB, pdf)

References

  • 1.Mullis K.B., Faloona F.A. Specific synthesis of DNA in vitro via a polymerase-catalyzed chain reaction. Methods Enzymol. 1987;155:335–350. doi: 10.1016/0076-6879(87)55023-6. [DOI] [PubMed] [Google Scholar]
  • 2.Whiley D.M., Sloots T.P. Sequence variation in primer targets affects the accuracy of viral quantitative PCR. J. Clin. Virol. 2005;34:104–107. doi: 10.1016/j.jcv.2005.02.010. [DOI] [PubMed] [Google Scholar]
  • 3.Ørum H., Nielsen P.E., et al. Stanley C. Single base pair mutation analysis by PNA directed PCR clamping. Nucleic Acids Res. 1993;21:5332–5336. doi: 10.1093/nar/21.23.5332. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Li M., Yin F., et al. Xia Q. Nucleic acid tests for clinical translation. Chem. Rev. 2021;121:10469–10558. doi: 10.1021/acs.chemrev.1c00241. [DOI] [PubMed] [Google Scholar]
  • 5.Yang H., Yan M., et al. Xu H. A tailored LNA clamping design principle: Efficient, economized, specific and ultrasensitive for the detection of point mutations. Biotechnol. J. 2021;16 doi: 10.1002/biot.202100233. [DOI] [PubMed] [Google Scholar]
  • 6.Homma C., Inokuchi D., et al. Adachi M. Effectiveness of blocking primers and a peptide nucleic acid (PNA) clamp for 18s metabarcoding dietary analysis of herbivorous fish. PLoS One. 2022;17 doi: 10.1371/journal.pone.0266268. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Petersen M., Wengel J. LNA: a versatile tool for therapeutics and genomics. Trends Biotechnol. 2003;21:74–81. doi: 10.1016/s0167-7799(02)00038-0. [DOI] [PubMed] [Google Scholar]
  • 8.Pellestor F., Paulasova P. The peptide nucleic acids (PNAs), powerful tools for molecular genetics and cytogenetics. Eur. J. Hum. Genet. 2004;12:694–700. doi: 10.1038/sj.ejhg.5201226. [DOI] [PubMed] [Google Scholar]
  • 9.Owczarzy R., You Y., et al. Tataurov A.V. Stability and mismatch discrimination of locked nucleic acid–DNA duplexes. Biochemist. 2011;50:9352–9367. doi: 10.1021/bi200904e. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Aoyanagi H., Pigolotti S., et al. Toyabe S. Error-suppression mechanism of PCR by blocker strands. Biophys. J. 2023;122:1334–1341. doi: 10.1016/j.bpj.2023.02.028. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Hopfield J.J. Kinetic proofreading: A new mechanism for reducing errors in biosynthetic processes requiring high specificity. Proc. Natl. Acad. Sci. USA. 1974;71:4135–4139. doi: 10.1073/pnas.71.10.4135. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Allawi H.T., SantaLucia J. Thermodynamics and NMR of internal G ·T mismatches in DNA. Biochemist. 1997;36:10581–10594. doi: 10.1021/bi962590c. [DOI] [PubMed] [Google Scholar]
  • 13.SantaLucia J. A unified view of polymer, dumbbell, and oligonucleotide DNA nearest-neighbor thermodynamics. Proc. Natl. Acad. Sci. USA. 1998;95:1460–1465. doi: 10.1073/pnas.95.4.1460. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Cock P.J.A., Antao T., et al. de Hoon M.J.L. Biopython: freely available Python tools for computational molecular biology and bioinformatics. Bioinformatics. 2009;25:1422–1423. doi: 10.1093/bioinformatics/btp163. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Kwok S., Kellogg D.E., et al. Sninsky J.J. Effects of primer-template mismatches on the polymerase chain reaction: Human immunodeficiency virus type 1 model studies. Nucleic Acids Res. 1990;18:999–1005. doi: 10.1093/nar/18.4.999. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Huang M.-M., Arnheim N., Goodman M.F. Extension of base mispairs by Taq DNA polymerase: implications for single nucleotide discrimination in PCR. Nucleic Acids Res. 1992;20:4567–4573. doi: 10.1093/nar/20.17.4567. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Bru D., Martin-Laurent F., Philippot L. Quantification of the detrimental effect of a single primer-template mismatch by real-time PCR using the 16s rRNA gene as an example. Appl. Environ. Microbiol. 2008;74:1660–1663. doi: 10.1128/aem.02403-07. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Anzalone A.V., Koblan L.W., Liu D.R. Genome editing with CRISPR-Cas nucleases, base editors, transposases and prime editors. Nat. Biotechnol. 2020;38:824–844. doi: 10.1038/s41587-020-0561-9. [DOI] [PubMed] [Google Scholar]
  • 19.Agrawal N., Dasaradhi P.V.N., et al. Mukherjee S.K. RNA interference: Biology, mechanism, and applications. Microbiol. Mol. Biol. Rev. 2003;67:657–685. doi: 10.1128/mmbr.67.4.657-685.2003. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Document S1. Supporting materials and methods, Figures S1–S5, and Tables S1 and S2
mmc1.pdf (227.6KB, pdf)
Document S2. Article plus supporting material
mmc2.pdf (2.4MB, pdf)

Data Availability Statement

Data are available from the authors upon reasonable request. The program code (Python) used for finding the optimal blocker concentrations is available at https://github.com/stoyabe/PCR_Blocker.


Articles from Biophysical Journal are provided here courtesy of The Biophysical Society

RESOURCES