Live Stats, next update: Wed 02 Sep
Human PDBs Analysed
Confidently Wrong
Novel + Confidently Wrong
DB size
Visitors
Full statistics →
New PDB Depositions vs. Their Blind AlphaFold Predictions — A Running Test of “Is Folding Solved?”

Release week 2018-10-17

48
structures analysed (14 full · 29.2%)
00.0%
confidently wrong
12.1%
novel sequences
00.0%
novel & wrong
0.927
median TM-score

Zoomed into the red "confidently wrong" box above (pLDDT ≥ 70, TM < 0.5) — the structures AlphaFold got confidently wrong.

How to read this

Every point is one experimental protein structure. The horizontal axis is AlphaFold's own confidence in its prediction (mean pLDDT, 0–100). The vertical axis is how well that blind prediction actually matches the experiment (TM-score, 0–1; above 0.5 means the same fold, above 0.9 near-identical). Marker size grows with the FRAUD score (confidence-weighted error).

Cutoff: the shaded red box is the "confidently wrong" zone — AlphaFold was confident (mean pLDDT above 70) yet the fold is wrong (TM-score below 0.5). A predictor that had truly solved folding would leave that box empty.

Take-home: 0 of 48 structures (0.0%) are confidently wrong; median TM-score is 0.927.

Read moreShow less — how each metric is calculated

TM-score: how similar the two 3D shapes are overall, 0–1 (a random pair scores ~0.17, an identical fold ~1). Length-normalised so large and small proteins compare fairly. TM = (1/L) Σᵢ 1/(1 + (dᵢ/d₀)²) — dᵢ is the gap between the i-th aligned Cα atoms after best-fit superposition, and d₀ = 1.24(L−15)^⅓ − 1.8 sets the distance scale for length L. Ref: Zhang & Skolnick, Proteins 2004 doi:10.1002/prot.20264; computed with TM-align, doi:10.1093/nar/gki524.

pLDDT: AlphaFold's own confidence in each residue, 0–100 (higher = surer). It is the model predicting its own accuracy before ever seeing the experiment. We plot the per-structure mean. pLDDT = (1/N) Σᵢ pLDDTᵢ, where pLDDTᵢ is the network's confidence output for residue i.

FRAUD score: the headline number — how wrong the prediction was, weighted by how confident AlphaFold was, so a big error it was sure about counts most. FRAUD = (1/N) Σᵢ (pLDDTᵢ/100) · min(Δᵢ,15)/15, where Δᵢ is residue i's Cα distance from the experiment (Å) after superposition, capped at 15 Å. Runs 0 (perfect) to 1.

Novelty: how little AlphaFold had to go on — 100 minus the highest sequence identity between this protein and any structure released before the 2018-04-30 training cutoff. High = genuinely unseen (100% = nothing similar was ever in the training set). novelty = 100 − maxₚ identity(s, p) over pre-cutoff PDB chains p, with identity = 100 × (matching aligned residues)/(alignment length), from an MMseqs2 search.

Homology: the point colour. High novelty (above 70%) flags the sequence as novel (amber) — AlphaFold had no close template; otherwise a pre-cutoff homolog existed (blue) it could have learned the fold from. novel ⇔ novelty > 70%.

What the metrics mean

TM-score (0–1): overall fold match — above 0.5 is the same fold, above 0.9 near-identical. Cα-RMSD (Å): average backbone distance after best-fit superposition — lower is better (under 2 Å is excellent). lDDT (0–1): local accuracy measured without superposition — above 0.8 is good. Each bar counts how many structures fall in that range.

Take-home: median TM-score 0.927 — most predictions match the experimental fold well, with a long tail that do not.

Read moreShow less — how each metric is calculated

TM-score: overall shape match, 0–1 (above 0.5 = same fold, above 0.9 near-identical), length-normalised. TM = (1/L) Σᵢ 1/(1 + (dᵢ/d₀)²) with dᵢ the aligned-Cα gap after superposition and d₀ = 1.24(L−15)^⅓ − 1.8. Ref: Zhang & Skolnick, Proteins 2004 doi:10.1002/prot.20264.

Cα-RMSD: the average straight-line distance between matching backbone Cα atoms once the two structures are best-fit superposed (Å; lower is better, under ~2 Å excellent). RMSD = √( (1/N) Σᵢ |Pᵢ − (R·Qᵢ + t)|² ), with Pᵢ/Qᵢ the experimental/model Cα coordinates and R,t the rotation and translation from Kabsch superposition.

lDDT: local accuracy with no superposition — the fraction of short-range inter-residue distances the model reproduces, so it is not fooled by a single wrongly-placed domain. lDDT = (1/N) Σᵢ ¼ Σ_t 1[ |d_exp − d_model| < t ] over thresholds t ∈ {0.5, 1, 2, 4} Å and all residue pairs within 15 Å.

Trend over time

The blue line is the mean TM-score of the structures by their deposit month; the red bars count how many were confidently wrong. Binning by deposit date (rather than by when we processed them) gives an even timeline, tracking whether AlphaFold's accuracy on newly-deposited structures is holding steady, improving, or slipping as the PDB keeps growing. Months with fewer than 5 structures are omitted so each point is a meaningful average. A dashed trend line is drawn only if the change over time is statistically significant.

Read moreShow less — how each metric is calculated

Mean TM-score: the month's average shape-match score. TM̄ = (1/M) Σ TM over the M structures deposited that month.

Trend line: a monotonic-trend test over the monthly means — Mann-Kendall (Kendall's τ) for significance, with a robust Theil-Sen slope for the line. Rank-based, so it tolerates the skewed TM distribution and noisy low-count months. Shown only when p < 0.05; otherwise a note reports that no significant trend was found.

Confidently wrong: the count that were both confident and wrong. count( mean pLDDT > 70 AND TM < 0.5 ).

Matched structures

PDBUniProtProteinMethodÅDeposited Novelty %pLDDTTMlDDTGDT_TSRMSDFRAUDFlag
6DAD_A P0DP23 Calmodulin-1 X-ray 1.65 2018-05-01 0.70 86.40 0.50 0.80 8.85 12.13 0.62 ok
6DAE_A P0DP23 Calmodulin-1 X-ray 2.00 2018-05-01 0.70 86.17 0.51 0.80 8.79 12.03 0.62 ok
6DAF_A P0DP23 Calmodulin-1 X-ray 2.40 2018-05-01 0.70 86.40 0.53 0.87 10.76 10.89 0.57 ok
6DAH_A P0DP23 Calmodulin-1 X-ray 2.50 2018-05-01 0.70 85.77 0.56 0.91 19.01 9.30 0.43 ok
6EF3_s P07818 Model substrate polypeptide EM 4.17 2018-08-15 100.00 novel 34.08 0.24 0.70 4.61 13.64 0.27 ok
6GN7_L P00734 Prothrombin X-ray 2.80 2018-05-30 83.94 0.78 0.19 ok
6DAF_C Q13936 Voltage-dependent L-type calcium channel s X-ray 2.40 2018-05-01 61.94 0.71 0.18 ok
6EF1_s P07818 model substrate polypeptide EM 4.73 2018-08-15 35.25 0.21 0.62 30.00 6.11 0.14 ok
6EF2_s P07818 model substrate polypeptide EM 4.27 2018-08-15 35.82 0.21 0.69 34.38 5.09 0.11 ok
6DK0_A Q99720 Sigma non-opioid intracellular receptor 1 X-ray 2.90 2018-05-28 93.69 0.89 0.10 ok
6DJZ_A Q99720 Sigma non-opioid intracellular receptor 1 X-ray 3.08 2018-05-28 93.69 0.89 0.10 ok
6DK1_A Q99720 Sigma non-opioid intracellular receptor 1 X-ray 3.12 2018-05-28 93.69 0.89 0.10 ok
6EF0_s P07818 model substrate polypeptide EM 4.43 2018-08-15 32.14 0.28 0.70 35.42 4.36 0.09 ok
6DAE_C Q13936 Voltage-dependent L-type calcium channel s X-ray 2.00 2018-05-01 0.00 68.79 0.66 0.93 69.57 2.63 0.09 ok
6MA5_B P51610 Host Cell Factor 1 peptide X-ray 2.00 2018-08-25 26.96 0.41 0.88 33.93 5.05 0.08 ok
6DAD_C Q13936 Voltage-dependent L-type calcium channel s X-ray 1.65 2018-05-01 61.94 0.87 0.08 ok
6MA4_B P51610 Host Cell Factor 1 peptide X-ray 2.00 2018-08-25 26.96 0.41 0.87 33.93 5.03 0.08 ok
6MA2_B P51610 Host Cell Factor 1 peptide X-ray 2.10 2018-08-25 26.96 0.40 0.88 35.71 5.00 0.08 ok
6GN7_H P00734 Prothrombin X-ray 2.80 2018-05-30 83.94 0.91 0.08 ok
6MA3_B P51610 Host Cell Factor 1 peptide X-ray 2.00 2018-08-25 26.66 0.45 0.86 40.38 4.51 0.07 ok
6GGS_A O43353 Receptor-interacting serine/threonine-prot EM 3.94 2018-05-03 76.06 0.91 0.07 ok
6HQ9_A Q5T890 DNA excision repair protein ERCC-6-like 2 X-ray 1.98 2018-09-24 57.81 0.89 0.07 ok
6HVD_A Q9H2G2 STE20-like serine/threonine-protein kinase X-ray 1.63 2018-10-10 64.50 0.91 0.06 ok
6EDU_G P01730 T-cell surface glycoprotein CD4 EM 4.06 2018-08-11 85.25 0.94 0.05 ok
6MA1_B P51610 Host Cell Factor 1 peptide X-ray 2.75 2018-08-25 27.99 0.53 0.86 65.00 2.56 0.04 ok
6GYT_B Q09472 Histone acetyltransferase p300 X-ray 2.50 2018-07-01 53.25 0.93 0.04 ok
6MA1_A O15294 UDP-N-acetylglucosamine--peptide N-acetylg X-ray 2.75 2018-08-25 93.06 0.97 0.03 ok
6MA2_A O15294 UDP-N-acetylglucosamine--peptide N-acetylg X-ray 2.10 2018-08-25 93.06 0.97 0.03 ok
6MA4_A O15294 UDP-N-acetylglucosamine--peptide N-acetylg X-ray 2.00 2018-08-25 93.06 0.97 0.03 ok
6MA5_A O15294 UDP-N-acetylglucosamine--peptide N-acetylg X-ray 2.00 2018-08-25 93.06 0.97 0.03 ok
6MA3_A O15294 UDP-N-acetylglucosamine--peptide N-acetylg X-ray 2.00 2018-08-25 93.06 0.97 0.03 ok
6GYR_A Q09472 Histone acetyltransferase p300 X-ray 3.10 2018-07-01 53.25 0.95 0.03 ok
6GYT_A Q09472 Histone acetyltransferase p300 X-ray 2.50 2018-07-01 53.25 0.95 0.03 ok
6E3I_A Q16548 Bcl-2-related protein A1 X-ray 1.48 2018-07-14 87.31 0.97 0.02 ok
6E3J_A Q16548 Bcl-2-related protein A1 X-ray 1.48 2018-07-14 87.31 0.97 0.02 ok
6DI1_A Q06187 Tyrosine-protein kinase BTK X-ray 1.10 2018-05-22 84.44 0.98 0.02 ok
6DI0_A Q06187 Tyrosine-protein kinase BTK X-ray 1.30 2018-05-22 84.44 0.98 0.02 ok
6HAL_B P68871 Hemoglobin subunit beta X-ray 2.20 2018-08-07 97.19 0.98 0.02 ok
6GKG_A P61981 14-3-3 protein gamma X-ray 2.85 2018-05-20 94.19 0.98 0.02 ok
6GKF_A P61981 14-3-3 protein gamma X-ray 2.60 2018-05-20 94.19 0.99 0.01 ok
6HAL_A P69905 Hemoglobin subunit alpha X-ray 2.20 2018-08-07 98.06 0.99 0.01 ok
6HUE_A O60260 E3 ubiquitin-protein ligase parkin X-ray 2.85 2018-10-07 78.06 0.99 0.01 ok
6HT1_A Q03111 Protein ENL X-ray 2.10 2018-10-02 64.69 0.99 0.01 ok
6HT0_A Q03111 Protein ENL X-ray 1.80 2018-10-02 64.69 0.99 0.01 ok
6GIU_A P29218 Inositol monophosphatase 1 X-ray 1.39 2018-05-15 96.19 0.99 0.01 ok
6GJ0_A P29218 Inositol monophosphatase 1 X-ray 1.73 2018-05-15 96.19 1.00 0.00 ok
5ZUN_A Q99685 Monoglyceride lipase X-ray 1.35 2018-05-08 93.88 1.00 0.00 ok
6HRL_A Q6TDP4 Kelch-like protein 17 X-ray 2.60 2018-09-27 87.25 1.00 0.00 ok

Click a column header to sort. FRAUD = confidence-weighted error; a structure is flagged confidently wrong when mean pLDDT > 70 yet TM-score < 0.5.