Live Stats, next update: Wed 02 Sep
Human PDBs Analysed
Confidently Wrong
Novel + Confidently Wrong
DB size
Visitors
Full statistics →
New PDB Depositions vs. Their Blind AlphaFold Predictions — A Running Test of “Is Folding Solved?”

Release week 2021-08-18

53
structures analysed (11 full · 20.8%)
47.5%
confidently wrong
11.9%
novel sequences
00.0%
novel & wrong
0.952
median TM-score

Zoomed into the red "confidently wrong" box above (pLDDT ≥ 70, TM < 0.5) — the structures AlphaFold got confidently wrong.

The worst offenders ranked, one row per protein (AlphaFold has one model per sequence, so repeat depositions are collapsed — “×N” marks how many structures of that protein exist; the worst is shown). Each row joins what AlphaFold claimed (blue, pLDDT/100) to what the experiment showed (red, TM-score) — the longer the bar, the bigger the miss. Hover a dot to preview its Cα-deviation ribbon; click to open the full entry.

How to read this

Every point is one experimental protein structure. The horizontal axis is AlphaFold's own confidence in its prediction (mean pLDDT, 0–100). The vertical axis is how well that blind prediction actually matches the experiment (TM-score, 0–1; above 0.5 means the same fold, above 0.9 near-identical). Marker size grows with the FRAUD score (confidence-weighted error).

Cutoff: the shaded red box is the "confidently wrong" zone — AlphaFold was confident (mean pLDDT above 70) yet the fold is wrong (TM-score below 0.5). A predictor that had truly solved folding would leave that box empty.

Take-home: 4 of 53 structures (7.5%) are confidently wrong; median TM-score is 0.952.

Read moreShow less — how each metric is calculated

TM-score: how similar the two 3D shapes are overall, 0–1 (a random pair scores ~0.17, an identical fold ~1). Length-normalised so large and small proteins compare fairly. TM = (1/L) Σᵢ 1/(1 + (dᵢ/d₀)²) — dᵢ is the gap between the i-th aligned Cα atoms after best-fit superposition, and d₀ = 1.24(L−15)^⅓ − 1.8 sets the distance scale for length L. Ref: Zhang & Skolnick, Proteins 2004 doi:10.1002/prot.20264; computed with TM-align, doi:10.1093/nar/gki524.

pLDDT: AlphaFold's own confidence in each residue, 0–100 (higher = surer). It is the model predicting its own accuracy before ever seeing the experiment. We plot the per-structure mean. pLDDT = (1/N) Σᵢ pLDDTᵢ, where pLDDTᵢ is the network's confidence output for residue i.

FRAUD score: the headline number — how wrong the prediction was, weighted by how confident AlphaFold was, so a big error it was sure about counts most. FRAUD = (1/N) Σᵢ (pLDDTᵢ/100) · min(Δᵢ,15)/15, where Δᵢ is residue i's Cα distance from the experiment (Å) after superposition, capped at 15 Å. Runs 0 (perfect) to 1.

Novelty: how little AlphaFold had to go on — 100 minus the highest sequence identity between this protein and any structure released before the 2018-04-30 training cutoff. High = genuinely unseen (100% = nothing similar was ever in the training set). novelty = 100 − maxₚ identity(s, p) over pre-cutoff PDB chains p, with identity = 100 × (matching aligned residues)/(alignment length), from an MMseqs2 search.

Homology: the point colour. High novelty (above 70%) flags the sequence as novel (amber) — AlphaFold had no close template; otherwise a pre-cutoff homolog existed (blue) it could have learned the fold from. novel ⇔ novelty > 70%.

What the metrics mean

TM-score (0–1): overall fold match — above 0.5 is the same fold, above 0.9 near-identical. Cα-RMSD (Å): average backbone distance after best-fit superposition — lower is better (under 2 Å is excellent). lDDT (0–1): local accuracy measured without superposition — above 0.8 is good. Each bar counts how many structures fall in that range.

Take-home: median TM-score 0.952 — most predictions match the experimental fold well, with a long tail that do not.

Read moreShow less — how each metric is calculated

TM-score: overall shape match, 0–1 (above 0.5 = same fold, above 0.9 near-identical), length-normalised. TM = (1/L) Σᵢ 1/(1 + (dᵢ/d₀)²) with dᵢ the aligned-Cα gap after superposition and d₀ = 1.24(L−15)^⅓ − 1.8. Ref: Zhang & Skolnick, Proteins 2004 doi:10.1002/prot.20264.

Cα-RMSD: the average straight-line distance between matching backbone Cα atoms once the two structures are best-fit superposed (Å; lower is better, under ~2 Å excellent). RMSD = √( (1/N) Σᵢ |Pᵢ − (R·Qᵢ + t)|² ), with Pᵢ/Qᵢ the experimental/model Cα coordinates and R,t the rotation and translation from Kabsch superposition.

lDDT: local accuracy with no superposition — the fraction of short-range inter-residue distances the model reproduces, so it is not fooled by a single wrongly-placed domain. lDDT = (1/N) Σᵢ ¼ Σ_t 1[ |d_exp − d_model| < t ] over thresholds t ∈ {0.5, 1, 2, 4} Å and all residue pairs within 15 Å.

Trend over time

The blue line is the mean TM-score of the structures by their deposit month; the red bars count how many were confidently wrong. Binning by deposit date (rather than by when we processed them) gives an even timeline, tracking whether AlphaFold's accuracy on newly-deposited structures is holding steady, improving, or slipping as the PDB keeps growing. Months with fewer than 5 structures are omitted so each point is a meaningful average. A dashed trend line is drawn only if the change over time is statistically significant.

Read moreShow less — how each metric is calculated

Mean TM-score: the month's average shape-match score. TM̄ = (1/M) Σ TM over the M structures deposited that month.

Trend line: a monotonic-trend test over the monthly means — Mann-Kendall (Kendall's τ) for significance, with a robust Theil-Sen slope for the line. Rank-based, so it tolerates the skewed TM distribution and noisy low-count months. Shown only when p < 0.05; otherwise a note reports that no significant trend was found.

Confidently wrong: the count that were both confident and wrong. count( mean pLDDT > 70 AND TM < 0.5 ).

Matched structures

PDBUniProtProteinMethodÅDeposited Novelty %pLDDTTMlDDTGDT_TSRMSDFRAUDFlag
7NFE_J P49917 DNA ligase 4 EM 4.29 2021-02-06 0.00 85.67 0.49 0.74 11.53 10.46 0.53 wrong
7MIS_B P0DP24 Calmodulin EM 2.80 2021-04-17 0.00 86.94 0.42 0.67 11.65 10.15 0.52 wrong
7NFC_M P49917 DNA ligase 4 EM 4.14 2021-02-05 0.00 85.67 0.49 0.76 15.02 10.12 0.51 wrong
7MIR_B P0DP24 Calmodulin-2 EM 2.50 2021-04-17 0.00 86.57 0.43 0.68 13.37 9.98 0.51 wrong
7CYM_A Q12864 Cadherin-17 X-ray 2.70 2020-09-03 67.30 91.56 0.60 0.95 21.10 8.17 0.44 ok
7NFE_H Q13426 DNA repair protein XRCC4 EM 4.29 2021-02-06 1.90 95.27 0.69 0.75 41.17 4.66 0.26 ok
7NFC_K Q13426 DNA repair protein XRCC4 EM 4.14 2021-02-05 1.90 95.27 0.70 0.76 41.54 4.56 0.26 ok
7NFC_Q Q9H9Q4 Non-homologous end-joining factor 1 EM 4.14 2021-02-05 0.00 94.51 0.68 0.68 42.30 4.61 0.26 ok
7DB6_A P63096 Guanine nucleotide-binding protein G(i) su EM 3.30 2020-10-19 93.75 0.78 0.21 ok
7NFE_F Q9H9Q4 Non-homologous end-joining factor 1 EM 4.29 2021-02-06 81.75 0.75 0.20 ok
7NFE_B P12956 X-ray repair cross-complementing protein 6 EM 4.29 2021-02-06 84.44 0.77 0.19 ok
7NFC_B P12956 X-ray repair cross-complementing protein 6 EM 4.14 2021-02-05 84.44 0.78 0.19 ok
7NFE_C P13010 X-ray repair cross-complementing protein 5 EM 4.29 2021-02-06 83.12 0.81 0.16 ok
7F16_P Q96A98 Tuberoinfundibular peptide of 39 residues EM 2.80 2021-06-08 72.81 0.80 0.15 ok
7NFC_C P13010 X-ray repair cross-complementing protein 5 EM 4.14 2021-02-05 83.12 0.82 0.15 ok
7F16_A P63092 Guanine nucleotide-binding protein G(s) su EM 2.80 2021-06-08 91.31 0.87 0.12 ok
7MJ8_C Q8WX77 Insulin-like growth factor-binding protein X-ray 1.79 2021-04-19 50.36 0.29 0.66 45.83 3.72 0.11 ok
7F9Y_C Q9UBU3 Ghrelin-28 EM 2.90 2021-07-05 100.00 novel 50.46 0.20 0.55 51.67 3.13 0.10 ok
7MJ7_C Q8WX77 Insulin-like growth factor-binding protein X-ray 1.60 2021-04-19 50.82 0.30 0.71 56.82 2.89 0.09 ok
7MRV_A P30046 D-dopachrome decarboxylase X-ray 1.57 2021-05-09 97.94 0.91 0.09 ok
7F9Z_G P59768 Guanine nucleotide-binding protein G(I)/G( EM 3.20 2021-07-05 89.56 0.91 0.08 ok
7DB6_D P48039 Melatonin receptor type 1A EM 3.30 2020-10-19 85.75 0.92 0.07 ok
7F3G_A O15075 Isoform 4 of Serine/threonine-protein kina X-ray 2.10 2021-06-16 72.19 0.91 0.07 ok
7KLZ_A O43791 Speckle-type POZ protein X-ray 3.40 2020-11-01 90.12 0.93 0.06 ok
7F9Y_G P59768 Guanine nucleotide-binding protein G(I)/G( EM 2.90 2021-07-05 89.56 0.93 0.06 ok
7F16_R P49190 Parathyroid hormone 2 receptor EM 2.80 2021-06-08 71.62 0.93 0.05 ok
7F9N_C Q6GTX8 Leukocyte-associated immunoglobulin-like r X-ray 3.00 2021-07-04 71.38 0.95 0.03 ok
7F9M_C Q6GTX8 Leukocyte-associated immunoglobulin-like r X-ray 2.90 2021-07-04 71.38 0.96 0.03 ok
7N19_A P01903 HLA class II histocompatibility antigen, D X-ray 2.38 2021-05-27 89.19 0.97 0.03 ok
7F9L_G Q6GTX8 Leukocyte-associated immunoglobulin-like r X-ray 2.70 2021-07-04 71.38 0.96 0.03 ok
7MJA_A A0A5H2UYS3 HLA class I histocompatibility antigen X-ray 1.69 2021-04-19 85.25 0.97 0.03 ok
7M6L_F P68106 Peptidyl-prolyl cis-trans isomerase FKBP1B EM 3.98 2021-03-25 94.88 0.97 0.03 ok
7MJ9_B P61769 Beta-2-microglobulin X-ray 1.75 2021-04-19 94.06 0.97 0.02 ok
7MJ7_A A0A140T913 MHC class I antigen X-ray 1.60 2021-04-19 84.62 0.97 0.02 ok
7MJA_B P61769 Beta-2-microglobulin X-ray 1.69 2021-04-19 94.06 0.98 0.01 ok
7MJ8_A A0A140T913 MHC class I antigen X-ray 1.79 2021-04-19 84.62 0.98 0.01 ok
7AYE_A Q00987 Isoform 11 of E3 ubiquitin-protein ligase X-ray 2.95 2020-11-12 62.59 0.98 0.01 ok
7MJ7_B P61769 Beta-2-microglobulin X-ray 1.60 2021-04-19 94.06 0.99 0.01 ok
7MJ6_A A0A140T913 MHC class I antigen X-ray 1.95 2021-04-19 84.62 0.98 0.01 ok
7MJ9_A A0A140T913 MHC class I antigen X-ray 1.75 2021-04-19 84.62 0.98 0.01 ok
7N19_B Q5Y7D1 HLA class II histocompatibility antigen DR X-ray 2.38 2021-05-27 86.38 0.99 0.01 ok
7MJ8_B P61769 Beta-2-microglobulin X-ray 1.79 2021-04-19 94.06 0.99 0.01 ok
7MRU_A P30046 D-dopachrome decarboxylase X-ray 1.33 2021-05-09 97.94 0.99 0.01 ok
7MJ6_B P61769 Beta-2-microglobulin X-ray 1.95 2021-04-19 94.06 0.99 0.01 ok
7MW7_A P30046 D-dopachrome decarboxylase X-ray 1.10 2021-05-15 97.94 0.99 0.01 ok
7MSE_A P30046 D-dopachrome decarboxylase X-ray 1.27 2021-05-11 97.94 0.99 0.01 ok
7F9Z_B P62873 Guanine nucleotide-binding protein G(I)/G( EM 3.20 2021-07-05 97.06 0.99 0.01 ok
6ZZ4_A P17706 Tyrosine-protein phosphatase non-receptor X-ray 2.43 2020-08-03 85.88 0.99 0.01 ok
7F9Y_B P62873 Guanine nucleotide-binding protein G(I)/G( EM 2.90 2021-07-05 97.06 0.99 0.01 ok
7CTC_A P22830 Ferrochelatase, mitochondrial X-ray 2.00 2020-08-18 86.56 1.00 0.00 ok
7BFZ_A Q04609 Glutamate carboxypeptidase 2 X-ray 1.73 2021-01-05 93.81 1.00 0.00 ok
7CT7_A P22830 Ferrochelatase, mitochondrial X-ray 2.00 2020-08-18 86.56 1.00 0.00 ok
7RRP_A P02794 Ferritin heavy chain EM 1.27 2021-08-10 95.31 1.00 0.00 ok

Click a column header to sort. FRAUD = confidence-weighted error; a structure is flagged confidently wrong when mean pLDDT > 70 yet TM-score < 0.5.