Live Stats, next update: Wed 02 Sep
Human PDBs Analysed
Confidently Wrong
Novel + Confidently Wrong
DB size
Visitors
Full statistics →
New PDB Depositions vs. Their Blind AlphaFold Predictions — A Running Test of “Is Folding Solved?”

Release week 2020-11-04

60
structures analysed (27 full · 45.0%)
23.3%
confidently wrong
610.0%
novel sequences
00.0%
novel & wrong
0.967
median TM-score

Zoomed into the red "confidently wrong" box above (pLDDT ≥ 70, TM < 0.5) — the structures AlphaFold got confidently wrong.

The worst offenders ranked, one row per protein (AlphaFold has one model per sequence, so repeat depositions are collapsed — “×N” marks how many structures of that protein exist; the worst is shown). Each row joins what AlphaFold claimed (blue, pLDDT/100) to what the experiment showed (red, TM-score) — the longer the bar, the bigger the miss. Hover a dot to preview its Cα-deviation ribbon; click to open the full entry.

How to read this

Every point is one experimental protein structure. The horizontal axis is AlphaFold's own confidence in its prediction (mean pLDDT, 0–100). The vertical axis is how well that blind prediction actually matches the experiment (TM-score, 0–1; above 0.5 means the same fold, above 0.9 near-identical). Marker size grows with the FRAUD score (confidence-weighted error).

Cutoff: the shaded red box is the "confidently wrong" zone — AlphaFold was confident (mean pLDDT above 70) yet the fold is wrong (TM-score below 0.5). A predictor that had truly solved folding would leave that box empty.

Take-home: 2 of 60 structures (3.3%) are confidently wrong; median TM-score is 0.967.

Read moreShow less — how each metric is calculated

TM-score: how similar the two 3D shapes are overall, 0–1 (a random pair scores ~0.17, an identical fold ~1). Length-normalised so large and small proteins compare fairly. TM = (1/L) Σᵢ 1/(1 + (dᵢ/d₀)²) — dᵢ is the gap between the i-th aligned Cα atoms after best-fit superposition, and d₀ = 1.24(L−15)^⅓ − 1.8 sets the distance scale for length L. Ref: Zhang & Skolnick, Proteins 2004 doi:10.1002/prot.20264; computed with TM-align, doi:10.1093/nar/gki524.

pLDDT: AlphaFold's own confidence in each residue, 0–100 (higher = surer). It is the model predicting its own accuracy before ever seeing the experiment. We plot the per-structure mean. pLDDT = (1/N) Σᵢ pLDDTᵢ, where pLDDTᵢ is the network's confidence output for residue i.

FRAUD score: the headline number — how wrong the prediction was, weighted by how confident AlphaFold was, so a big error it was sure about counts most. FRAUD = (1/N) Σᵢ (pLDDTᵢ/100) · min(Δᵢ,15)/15, where Δᵢ is residue i's Cα distance from the experiment (Å) after superposition, capped at 15 Å. Runs 0 (perfect) to 1.

Novelty: how little AlphaFold had to go on — 100 minus the highest sequence identity between this protein and any structure released before the 2018-04-30 training cutoff. High = genuinely unseen (100% = nothing similar was ever in the training set). novelty = 100 − maxₚ identity(s, p) over pre-cutoff PDB chains p, with identity = 100 × (matching aligned residues)/(alignment length), from an MMseqs2 search.

Homology: the point colour. High novelty (above 70%) flags the sequence as novel (amber) — AlphaFold had no close template; otherwise a pre-cutoff homolog existed (blue) it could have learned the fold from. novel ⇔ novelty > 70%.

What the metrics mean

TM-score (0–1): overall fold match — above 0.5 is the same fold, above 0.9 near-identical. Cα-RMSD (Å): average backbone distance after best-fit superposition — lower is better (under 2 Å is excellent). lDDT (0–1): local accuracy measured without superposition — above 0.8 is good. Each bar counts how many structures fall in that range.

Take-home: median TM-score 0.967 — most predictions match the experimental fold well, with a long tail that do not.

Read moreShow less — how each metric is calculated

TM-score: overall shape match, 0–1 (above 0.5 = same fold, above 0.9 near-identical), length-normalised. TM = (1/L) Σᵢ 1/(1 + (dᵢ/d₀)²) with dᵢ the aligned-Cα gap after superposition and d₀ = 1.24(L−15)^⅓ − 1.8. Ref: Zhang & Skolnick, Proteins 2004 doi:10.1002/prot.20264.

Cα-RMSD: the average straight-line distance between matching backbone Cα atoms once the two structures are best-fit superposed (Å; lower is better, under ~2 Å excellent). RMSD = √( (1/N) Σᵢ |Pᵢ − (R·Qᵢ + t)|² ), with Pᵢ/Qᵢ the experimental/model Cα coordinates and R,t the rotation and translation from Kabsch superposition.

lDDT: local accuracy with no superposition — the fraction of short-range inter-residue distances the model reproduces, so it is not fooled by a single wrongly-placed domain. lDDT = (1/N) Σᵢ ¼ Σ_t 1[ |d_exp − d_model| < t ] over thresholds t ∈ {0.5, 1, 2, 4} Å and all residue pairs within 15 Å.

Trend over time

The blue line is the mean TM-score of the structures by their deposit month; the red bars count how many were confidently wrong. Binning by deposit date (rather than by when we processed them) gives an even timeline, tracking whether AlphaFold's accuracy on newly-deposited structures is holding steady, improving, or slipping as the PDB keeps growing. Months with fewer than 5 structures are omitted so each point is a meaningful average. A dashed trend line is drawn only if the change over time is statistically significant.

Read moreShow less — how each metric is calculated

Mean TM-score: the month's average shape-match score. TM̄ = (1/M) Σ TM over the M structures deposited that month.

Trend line: a monotonic-trend test over the monthly means — Mann-Kendall (Kendall's τ) for significance, with a robust Theil-Sen slope for the line. Rank-based, so it tolerates the skewed TM distribution and noisy low-count months. Shown only when p < 0.05; otherwise a note reports that no significant trend was found.

Confidently wrong: the count that were both confident and wrong. count( mean pLDDT > 70 AND TM < 0.5 ).

Matched structures

PDBUniProtProteinMethodÅDeposited Novelty %pLDDTTMlDDTGDT_TSRMSDFRAUDFlag
6TQL_A P07911 Uromodulin EM 3.96 2019-12-16 0.70 86.34 0.45 0.83 0.00 26.61 0.85 wrong
6TQK_A P07911 Uromodulin EM 3.35 2019-12-16 2.10 86.02 0.46 0.83 0.00 26.68 0.85 wrong
7CK6_E Q96B49 Mitochondrial import receptor subunit TOM6 EM 3.40 2020-07-15 100.00 novel 80.71 0.52 0.82 37.50 5.31 0.26 ok
6TKH_L P00734 Thrombin light chain X-ray 1.90 2019-11-28 0.00 91.98 0.63 0.79 44.44 4.81 0.23 ok
6TKL_L P00734 Prothrombin X-ray 1.30 2019-11-28 0.00 91.98 0.63 0.80 45.14 4.78 0.23 ok
6TKG_L P00734 Prothrombin X-ray 1.35 2019-11-28 0.00 91.98 0.63 0.80 44.44 4.76 0.23 ok
6TKJ_L P00734 Thrombin light chain X-ray 2.81 2019-11-28 0.00 93.08 0.69 0.83 60.48 3.73 0.17 ok
6TKI_L P00734 Thrombin light chain X-ray 1.80 2019-11-28 0.00 92.23 0.69 0.85 62.50 3.02 0.14 ok
6L6I_A P78348 Acid-sensing ion channel 1 X-ray 3.24 2019-10-29 10.60 91.64 0.92 0.92 70.09 3.60 0.13 ok
7CK6_I Q8N4H5 Mitochondrial import receptor subunit TOM5 EM 3.40 2020-07-15 88.00 0.86 0.12 ok
7CK6_C Q9NS69 Mitochondrial import receptor subunit TOM2 EM 3.40 2020-07-15 100.00 novel 88.08 0.62 0.95 64.15 2.39 0.12 ok
6M23_A Q9H2X9 Solute carrier family 12 member 5 EM 3.20 2020-02-26 100.00 novel 87.66 0.96 0.91 63.94 2.27 0.12 ok
7D3S_P P09683 Secretin EM 2.90 2020-09-20 65.88 0.82 0.12 ok
6W9K_B Q9UBK2 Peroxisome proliferator-activated receptor X-ray 1.60 2020-03-23 52.75 0.79 0.11 ok
7D3S_A P63092 Guanine nucleotide-binding protein G(s) su EM 2.90 2020-09-20 91.31 0.88 0.11 ok
7CK6_G Q9P0U1 Mitochondrial import receptor subunit TOM7 EM 3.40 2020-07-15 100.00 novel 94.71 0.69 0.86 71.94 1.92 0.11 ok
6M22_A Q9UHW9 Solute carrier family 12 member 6 EM 2.70 2020-02-26 100.00 novel 90.97 0.97 0.97 73.14 1.89 0.10 ok
6W9M_B Q15466 Nuclear receptor subfamily 0 group B membe X-ray 1.59 2020-03-23 58.42 0.69 0.85 63.64 2.62 0.09 ok
6L6P_A Q16515 Acid-sensing ion channel 2 X-ray 3.08 2019-10-29 28.10 93.08 0.96 0.90 81.97 2.23 0.09 ok
6L7K_A P12104 Fatty acid-binding protein, intestinal NMR 2019-11-01 1.60 91.32 0.88 0.82 75.76 1.74 0.09 ok
6TKH_H P00734 Thrombin heavy chain X-ray 1.90 2019-11-28 0.00 90.68 0.92 0.81 81.67 2.71 0.08 ok
6TKL_H P00734 Prothrombin X-ray 1.30 2019-11-28 0.00 90.68 0.92 0.82 81.77 2.72 0.08 ok
6TKG_H P00734 Prothrombin X-ray 1.35 2019-11-28 0.00 90.50 0.92 0.81 81.65 2.76 0.08 ok
6TKI_H P00734 Thrombin heavy chain X-ray 1.80 2019-11-28 0.00 90.50 0.92 0.81 82.96 2.70 0.08 ok
6M1Y_A Q9UHW9 Solute carrier family 12 member 6 EM 3.20 2020-02-26 100.00 novel 89.76 0.98 0.94 80.81 1.54 0.08 ok
6TKJ_H P00734 Thrombin heavy chain X-ray 2.81 2019-11-28 0.00 90.86 0.93 0.83 84.54 2.59 0.07 ok
7AAZ_A Q12866 Tyrosine-protein kinase Mer X-ray 1.85 2020-09-05 72.25 0.91 0.07 ok
6L65_A Q8IXJ6 NAD-dependent protein deacetylase sirtuin- X-ray 1.80 2019-10-28 0.40 93.17 0.97 0.92 88.15 1.85 0.06 ok
6L66_A Q8IXJ6 NAD-dependent protein deacetylase sirtuin- X-ray 2.17 2019-10-28 0.40 92.59 0.98 0.95 94.30 1.03 0.04 ok
6ZQZ_A O00408 cGMP-dependent 3',5'-cyclic phosphodiester X-ray 1.88 2020-07-10 83.69 0.95 0.04 ok
6XOZ_A P32322 Pyrroline-5-carboxylate reductase 1, mitoc X-ray 2.35 2020-07-07 89.81 0.96 0.03 ok
7CQE_A Q12866 Tyrosine-protein kinase Mer X-ray 2.69 2020-08-10 72.25 0.96 0.03 ok
6L7L_A P23921 Ribonucleoside-diphosphate reductase large X-ray 2.17 2019-11-01 0.00 96.14 1.00 0.98 97.59 0.61 0.03 ok
7CK6_A O96008 Mitochondrial import receptor subunit TOM4 EM 3.40 2020-07-15 78.38 0.96 0.03 ok
6XWD_A P31947 14-3-3 protein sigma X-ray 1.60 2020-01-23 92.88 0.97 0.03 ok
7AEW_AAA P31947 14-3-3 protein sigma X-ray 1.20 2020-09-18 92.88 0.97 0.03 ok
6W9L_B Q9UBK2 Peroxisome proliferator-activated receptor X-ray 1.45 2020-03-23 60.13 0.59 0.96 95.83 0.75 0.03 ok
6L6K_A P19793 Retinoic acid receptor RXR-alpha X-ray 1.80 2019-10-29 0.00 94.21 0.99 0.99 99.64 0.41 0.02 ok
6XXR_A Q8N8S7 Protein enabled homolog X-ray 1.48 2020-01-28 70.62 0.98 0.02 ok
6XP0_A P32322 Pyrroline-5-carboxylate reductase 1, mitoc X-ray 1.95 2020-07-07 89.81 0.98 0.01 ok
7D3S_R P47872 Secretin receptor EM 2.90 2020-09-20 76.56 0.98 0.01 ok
6XP3_A P32322 Pyrroline-5-carboxylate reductase 1, mitoc X-ray 1.93 2020-07-07 89.81 0.99 0.01 ok
6XP2_A P32322 Pyrroline-5-carboxylate reductase 1, mitoc X-ray 2.30 2020-07-07 89.81 0.99 0.01 ok
6VRU_A P11309 Serine/threonine-protein kinase pim-1 X-ray 2.07 2020-02-10 89.44 0.99 0.01 ok
6W58_A O60760 Hematopoietic prostaglandin D synthase X-ray 2.40 2020-03-12 97.31 0.99 0.01 ok
6VND_A Q02127 Dihydroorotate dehydrogenase (quinone), mi X-ray 1.97 2020-01-29 96.12 0.99 0.01 ok
6W8H_A O60760 Hematopoietic prostaglandin D synthase X-ray 1.97 2020-03-20 97.31 0.99 0.01 ok
6VRV_A P11309 Serine/threonine-protein kinase pim-1 X-ray 1.74 2020-02-10 89.44 0.99 0.00 ok
6XP1_A P32322 Pyrroline-5-carboxylate reductase 1, mitoc X-ray 1.75 2020-07-07 89.81 0.99 0.00 ok
7JNX_A P00918 Carbonic anhydrase 2 X-ray 1.29 2020-08-05 97.38 1.00 0.00 ok
7JOB_A P00918 Carbonic anhydrase 2 X-ray 1.38 2020-08-06 97.38 1.00 0.00 ok
7JNZ_A P00918 Carbonic anhydrase 2 X-ray 1.29 2020-08-05 97.38 1.00 0.00 ok
7JO1_A P00918 Carbonic anhydrase 2 X-ray 1.50 2020-08-05 97.38 1.00 0.00 ok
7JNV_A P00918 Carbonic anhydrase 2 X-ray 1.49 2020-08-05 97.38 1.00 0.00 ok
7JNR_A P00918 Carbonic anhydrase 2 X-ray 1.44 2020-08-05 97.38 1.00 0.00 ok
7JO3_A P00918 Carbonic anhydrase 2 X-ray 1.45 2020-08-05 97.38 1.00 0.00 ok
7JO2_A P00918 Carbonic anhydrase 2 X-ray 1.31 2020-08-05 97.38 1.00 0.00 ok
7JNW_A P00918 Carbonic anhydrase 2 X-ray 1.29 2020-08-05 97.38 1.00 0.00 ok
7JO0_A P00918 Carbonic anhydrase 2 X-ray 1.61 2020-08-05 97.38 1.00 0.00 ok
7JOC_A P00918 Carbonic anhydrase 2 X-ray 1.39 2020-08-06 97.38 1.00 0.00 ok

Click a column header to sort. FRAUD = confidence-weighted error; a structure is flagged confidently wrong when mean pLDDT > 70 yet TM-score < 0.5.