Live Stats, next update: Wed 02 Sep
Human PDBs Analysed
Confidently Wrong
Novel + Confidently Wrong
DB size
Visitors
Full statistics →
New PDB Depositions vs. Their Blind AlphaFold Predictions — A Running Test of “Is Folding Solved?”

Release week 2022-10-26

66
structures analysed (3 full · 4.5%)
11.5%
confidently wrong
00.0%
novel sequences
00.0%
novel & wrong
0.935
median TM-score

Zoomed into the red "confidently wrong" box above (pLDDT ≥ 70, TM < 0.5) — the structures AlphaFold got confidently wrong.

The worst offenders ranked, one row per protein (AlphaFold has one model per sequence, so repeat depositions are collapsed — “×N” marks how many structures of that protein exist; the worst is shown). Each row joins what AlphaFold claimed (blue, pLDDT/100) to what the experiment showed (red, TM-score) — the longer the bar, the bigger the miss. Hover a dot to preview its Cα-deviation ribbon; click to open the full entry.

How to read this

Every point is one experimental protein structure. The horizontal axis is AlphaFold's own confidence in its prediction (mean pLDDT, 0–100). The vertical axis is how well that blind prediction actually matches the experiment (TM-score, 0–1; above 0.5 means the same fold, above 0.9 near-identical). Marker size grows with the FRAUD score (confidence-weighted error).

Cutoff: the shaded red box is the "confidently wrong" zone — AlphaFold was confident (mean pLDDT above 70) yet the fold is wrong (TM-score below 0.5). A predictor that had truly solved folding would leave that box empty.

Take-home: 1 of 66 structures (1.5%) are confidently wrong; median TM-score is 0.935.

Read moreShow less — how each metric is calculated

TM-score: how similar the two 3D shapes are overall, 0–1 (a random pair scores ~0.17, an identical fold ~1). Length-normalised so large and small proteins compare fairly. TM = (1/L) Σᵢ 1/(1 + (dᵢ/d₀)²) — dᵢ is the gap between the i-th aligned Cα atoms after best-fit superposition, and d₀ = 1.24(L−15)^⅓ − 1.8 sets the distance scale for length L. Ref: Zhang & Skolnick, Proteins 2004 doi:10.1002/prot.20264; computed with TM-align, doi:10.1093/nar/gki524.

pLDDT: AlphaFold's own confidence in each residue, 0–100 (higher = surer). It is the model predicting its own accuracy before ever seeing the experiment. We plot the per-structure mean. pLDDT = (1/N) Σᵢ pLDDTᵢ, where pLDDTᵢ is the network's confidence output for residue i.

FRAUD score: the headline number — how wrong the prediction was, weighted by how confident AlphaFold was, so a big error it was sure about counts most. FRAUD = (1/N) Σᵢ (pLDDTᵢ/100) · min(Δᵢ,15)/15, where Δᵢ is residue i's Cα distance from the experiment (Å) after superposition, capped at 15 Å. Runs 0 (perfect) to 1.

Novelty: how little AlphaFold had to go on — 100 minus the highest sequence identity between this protein and any structure released before the 2018-04-30 training cutoff. High = genuinely unseen (100% = nothing similar was ever in the training set). novelty = 100 − maxₚ identity(s, p) over pre-cutoff PDB chains p, with identity = 100 × (matching aligned residues)/(alignment length), from an MMseqs2 search.

Homology: the point colour. High novelty (above 70%) flags the sequence as novel (amber) — AlphaFold had no close template; otherwise a pre-cutoff homolog existed (blue) it could have learned the fold from. novel ⇔ novelty > 70%.

What the metrics mean

TM-score (0–1): overall fold match — above 0.5 is the same fold, above 0.9 near-identical. Cα-RMSD (Å): average backbone distance after best-fit superposition — lower is better (under 2 Å is excellent). lDDT (0–1): local accuracy measured without superposition — above 0.8 is good. Each bar counts how many structures fall in that range.

Take-home: median TM-score 0.935 — most predictions match the experimental fold well, with a long tail that do not.

Read moreShow less — how each metric is calculated

TM-score: overall shape match, 0–1 (above 0.5 = same fold, above 0.9 near-identical), length-normalised. TM = (1/L) Σᵢ 1/(1 + (dᵢ/d₀)²) with dᵢ the aligned-Cα gap after superposition and d₀ = 1.24(L−15)^⅓ − 1.8. Ref: Zhang & Skolnick, Proteins 2004 doi:10.1002/prot.20264.

Cα-RMSD: the average straight-line distance between matching backbone Cα atoms once the two structures are best-fit superposed (Å; lower is better, under ~2 Å excellent). RMSD = √( (1/N) Σᵢ |Pᵢ − (R·Qᵢ + t)|² ), with Pᵢ/Qᵢ the experimental/model Cα coordinates and R,t the rotation and translation from Kabsch superposition.

lDDT: local accuracy with no superposition — the fraction of short-range inter-residue distances the model reproduces, so it is not fooled by a single wrongly-placed domain. lDDT = (1/N) Σᵢ ¼ Σ_t 1[ |d_exp − d_model| < t ] over thresholds t ∈ {0.5, 1, 2, 4} Å and all residue pairs within 15 Å.

Trend over time

The blue line is the mean TM-score of the structures by their deposit month; the red bars count how many were confidently wrong. Binning by deposit date (rather than by when we processed them) gives an even timeline, tracking whether AlphaFold's accuracy on newly-deposited structures is holding steady, improving, or slipping as the PDB keeps growing. Months with fewer than 5 structures are omitted so each point is a meaningful average. A dashed trend line is drawn only if the change over time is statistically significant.

Read moreShow less — how each metric is calculated

Mean TM-score: the month's average shape-match score. TM̄ = (1/M) Σ TM over the M structures deposited that month.

Trend line: a monotonic-trend test over the monthly means — Mann-Kendall (Kendall's τ) for significance, with a robust Theil-Sen slope for the line. Rank-based, so it tolerates the skewed TM distribution and noisy low-count months. Shown only when p < 0.05; otherwise a note reports that no significant trend was found.

Confidently wrong: the count that were both confident and wrong. count( mean pLDDT > 70 AND TM < 0.5 ).

Matched structures

PDBUniProtProteinMethodÅDeposited Novelty %pLDDTTMlDDTGDT_TSRMSDFRAUDFlag
7PGQ_F P21359 Neurofibromin EM 3.50 2021-08-15 6.30 82.55 0.46 0.78 2.56 23.99 0.72 wrong
7UI5_A P06702 Protein S100-A9 NMR 2022-03-28 0.90 94.32 0.64 0.77 16.01 13.94 0.54 ok
7T2Q_A P0DP23 Calmodulin-1 X-ray 1.95 2021-12-06 0.00 85.67 0.57 0.89 20.41 9.67 0.44 ok
7WYB_B P63096 Guanine nucleotide-binding protein G(i) su EM 2.97 2022-02-15 93.75 0.80 0.19 ok
7WYB_D P59768 Guanine nucleotide-binding protein G(I)/G( EM 2.97 2022-02-15 89.56 0.79 0.19 ok
7WZ4_A P63096 Guanine nucleotide-binding protein G(i) su EM 3.00 2022-02-17 93.75 0.83 0.16 ok
7WY8_C P59768 Guanine nucleotide-binding protein G(I)/G( EM 2.83 2022-02-15 89.56 0.83 0.15 ok
7VR1_A P33897 ATP-binding cassette sub-family D member 1 EM 3.40 2021-10-21 80.62 0.82 0.15 ok
7WY5_G P59768 Guanine nucleotide-binding protein G(I)/G( EM 2.83 2022-02-15 89.56 0.85 0.14 ok
7VRL_B Q9NWB1 RNA binding protein fox-1 homolog 1 NMR 2021-10-23 55.56 0.76 0.13 ok
7ZOT_A Q9UL19 Phospholipase A and acyltransferase 4 X-ray 1.74 2022-04-26 75.69 0.83 0.13 ok
8ADZ_J P01591 Immunoglobulin J chain EM 6.70 2022-07-12 87.06 0.89 0.10 ok
8AE3_J P01591 Immunoglobulin J chain EM 6.80 2022-07-12 87.06 0.89 0.10 ok
8AE0_J P01591 Immunoglobulin J chain EM 7.10 2022-07-12 87.06 0.89 0.10 ok
8AE2_J P01591 Immunoglobulin J chain EM 8.50 2022-07-12 87.06 0.89 0.10 ok
8AE3_C P01871 Immunoglobulin heavy constant mu EM 6.80 2022-07-12 85.44 0.89 0.09 ok
8AE2_C P01871 Immunoglobulin heavy constant mu EM 8.50 2022-07-12 85.44 0.89 0.09 ok
8AE0_C P01871 Immunoglobulin heavy constant mu EM 7.10 2022-07-12 85.44 0.89 0.09 ok
8ADY_J P01591 Immunoglobulin J chain EM 5.20 2022-07-12 87.06 0.89 0.09 ok
8ADY_C P01871 Immunoglobulin heavy constant mu EM 5.20 2022-07-12 85.44 0.89 0.09 ok
8ADZ_E P01871 Immunoglobulin heavy constant mu EM 6.70 2022-07-12 85.44 0.90 0.09 ok
7WQW_A P98073 Enteropeptidase non-catalytic heavy chain EM 3.20 2022-01-26 81.50 0.91 0.07 ok
7WZ4_G P59768 Guanine nucleotide-binding protein G(I)/G( EM 3.00 2022-02-17 89.56 0.92 0.07 ok
8AZA_C P98170 E3 ubiquitin-protein ligase XIAP EM 3.15 2022-09-05 74.25 0.90 0.07 ok
7VQG_A Q9NPG2 Neuroglobin X-ray 1.35 2021-10-19 95.19 0.93 0.07 ok
7WQZ_A P98073 Enteropeptidase non-catalytic heavy chain EM 3.70 2022-01-26 81.50 0.92 0.07 ok
7WR7_A P98073 Enteropeptidase non-catalytic heavy chain EM 3.10 2022-01-26 81.50 0.92 0.07 ok
8AZA_A O43353 Receptor-interacting serine/threonine-prot EM 3.15 2022-09-05 76.06 0.92 0.06 ok
7Y3L_A Q9BXA9 Sal-like protein 3 X-ray 2.50 2022-06-11 49.09 0.88 0.06 ok
7Y3M_A Q9UJQ4 Sal-like protein 4 X-ray 2.72 2022-06-11 51.06 0.88 0.06 ok
7WZ4_R Q9GZN0 Probable G-protein coupled receptor 88 EM 3.00 2022-02-17 75.81 0.93 0.05 ok
7WF6_A Q9Y5W8 Sorting nexin-13 X-ray 3.25 2021-12-26 75.94 0.93 0.05 ok
7QDO_A P01871 Isoform 2 of Immunoglobulin heavy constant EM 3.60 2021-11-27 85.44 0.94 0.05 ok
7F61_A Q9Y5N1 Histamine H3 receptor X-ray 2.60 2021-06-23 75.56 0.94 0.05 ok
7SI1_A P00533 Epidermal growth factor receptor X-ray 1.60 2021-10-12 75.94 0.94 0.04 ok
7XZQ_A Q9UKE5 TRAF2 and NCK-interacting protein kinase X-ray 2.09 2022-06-03 63.56 0.94 0.04 ok
7Y3I_A Q9UJQ4 Sal-like protein 4 X-ray 2.45 2022-06-10 51.06 0.92 0.04 ok
7XZR_A Q9UKE5 TRAF2 and NCK-interacting protein kinase X-ray 2.26 2022-06-03 63.56 0.94 0.04 ok
7WYB_C P62873 Guanine nucleotide-binding protein G(I)/G( EM 2.97 2022-02-15 97.06 0.97 0.03 ok
8E4X_A P78563 Double-stranded RNA-specific editase 1 X-ray 2.80 2022-08-19 76.50 0.96 0.03 ok
8E0F_A P78563 Double-stranded RNA-specific editase 1 X-ray 2.70 2022-08-09 76.50 0.96 0.03 ok
7SHV_A P15056 Serine/threonine-protein kinase B-raf X-ray 2.88 2021-10-11 66.38 0.95 0.03 ok
7WQX_A P98073 Enteropeptidase EM 2.70 2022-01-26 81.50 0.96 0.03 ok
7SFZ_A Q9NYP9 Protein Mis18-alpha X-ray 3.00 2021-10-04 79.12 0.96 0.03 ok
7Y3K_A Q9UJQ4 Sal-like protein 4 X-ray 2.50 2022-06-11 51.06 0.95 0.03 ok
8DYC_A P08684 Cytochrome P450 3A4 X-ray 2.40 2022-08-04 92.38 0.97 0.02 ok
7Z53_A P05164 Myeloperoxidase light chain X-ray 2.28 2022-03-07 89.00 0.97 0.02 ok
7UKZ_A P21127 Cyclin-dependent kinase 11B X-ray 2.60 2022-04-03 64.56 0.97 0.02 ok
7WQZ_B P98073 Enteropeptidase catalytic light chain EM 3.70 2022-01-26 81.50 0.97 0.02 ok
7QZR_A P05164 Myeloperoxidase light chain X-ray 2.18 2022-01-31 89.00 0.98 0.02 ok
7U8G_A P04839 EGFP, Cytochrome b-245 heavy chain chimera EM 3.20 2022-03-08 90.25 0.98 0.02 ok
7WY5_B P62873 Guanine nucleotide-binding protein G(I)/G( EM 2.83 2022-02-15 97.06 0.98 0.02 ok
7WY8_B P62873 Guanine nucleotide-binding protein G(I)/G( EM 2.83 2022-02-15 97.06 0.98 0.01 ok
7Z53_B P05164 Myeloperoxidase heavy chain X-ray 2.28 2022-03-07 89.00 0.98 0.01 ok
7WQW_B P98073 Enteropeptidase catalytic light chain EM 3.20 2022-01-26 81.50 0.98 0.01 ok
7QR4_A P09012 U1 small nuclear ribonucleoprotein A X-ray 2.83 2022-01-17 79.50 0.98 0.01 ok
7WR7_B P98073 Enteropeptidase catalytic light chain EM 3.10 2022-01-26 81.50 0.98 0.01 ok
7VR0_A P02768 Serum albumin X-ray 1.98 2021-10-21 92.69 0.99 0.01 ok
7QZR_B P05164 Myeloperoxidase heavy chain X-ray 2.18 2022-01-31 89.00 0.98 0.01 ok
7QZR_D P05164 Myeloperoxidase heavy chain X-ray 2.18 2022-01-31 89.00 0.99 0.01 ok
7U8G_B P13498 Cytochrome b-245 light chain EM 3.20 2022-03-08 76.88 0.98 0.01 ok
7QR3_A P09012 U1 small nuclear ribonucleoprotein A X-ray 2.18 2022-01-10 79.50 0.99 0.01 ok
7ZIT_A P63104 14-3-3 protein zeta/delta X-ray 1.79 2022-04-08 93.94 0.99 0.01 ok
7QHW_AAA Q5TCY1 Tau-tubulin kinase 1 X-ray 2.80 2021-12-14 51.06 0.98 0.01 ok
7WZ4_B P62873 Guanine nucleotide-binding protein G(I)/G( EM 3.00 2022-02-17 97.06 0.99 0.01 ok
7QFV_A Q92876 Kallikrein-6 X-ray 1.56 2021-12-06 91.75 1.00 0.00 ok

Click a column header to sort. FRAUD = confidence-weighted error; a structure is flagged confidently wrong when mean pLDDT > 70 yet TM-score < 0.5.