Verification Report

Verification Report: Methodology

Published · Updated

Verification Report

This report documents the independent verification of the Methodology research page.

Every factual claim on the methodology page was checked against its source data before publication. This is the first verification report on the site to draw on two separate bodies of source data — the methodology page combines quantitative validation results with the published record on research retractions, and each claim was checked against whichever of the two it came from. The source behind each claim, and its individual result, is listed claim by claim below.


How the claims were checked

The claims on the methodology page come from two distinct bodies of evidence:

  • Validation results — 16 claims covering the quantitative validation figures and how the methodology is described.
  • Retraction-record claims — 7 claims covering published retraction figures and their surrounding context.

Each claim was checked at the level that fits what it asserts:

  • Numerical checks confirm that a stated figure matches its source data exactly.
  • Framing checks confirm that the wording faithfully represents what the source actually says — that a summary doesn’t imply more than the underlying evidence supports.

A claim is only marked as framing-checked if its wording was actually put through that check; claims that assert a number, not a characterization, are confirmed numerically. The two checks are reported separately rather than blended, so “the figure is correct” is never overstated into “the interpretation is correct,” or the reverse.

One claim carries an openly documented correction. An earlier draft described a retraction trend as “accelerating year over year”; the framing check found that the source did not support that characterization. The wording was corrected, the claim was re-checked, and it then passed. We keep that history visible rather than removing it — it appears as a worked example on the methodology page itself, showing what the framing check catches and how a correction is recorded.


Summary

Claims checked23 total (16 validation results + 7 retraction-record)
ResultAll 23 confirmed — none unsupported, none uncheckable, none contradicted, and none left unchecked
Source dataTwo bodies of evidence: the quantitative validation results and the retraction-record sources (the exact versions checked are recorded below at build time)
Verified2026-05-21
Verified byVeritas v1.1

Claim-by-Claim Verification

Claim IDClaimSourceStrengthStatus
CLM-METHOD-001Verified cohort: 7 DMS entriesNeuroAutomata verified-cohort DMS entry count
Data source
strongCLM-METHOD-001Verified
CLM-METHOD-014Verified cohort: 6 distinct proteins (5 median-contributing + Calmodulin, held out as a disclosed weakness)NeuroAutomata verified-cohort distinct-protein count
Data source
strongCLM-METHOD-014Verified
CLM-METHOD-002Cohort median Spearman rho = 0.5298 (N=7 DMS entries)NeuroAutomata verified-cohort Spearman rho median (per-DMS-entry grain)
Data source
strongCLM-METHOD-002Verified
CLM-METHOD-003Cohort sorted Spearman values: [0.3989, 0.409, 0.484, 0.5298, 0.534, 0.634, 0.679]NeuroAutomata verified-cohort per-DMS Spearman rho values, sorted ascending
Data source
strongCLM-METHOD-003Verified
CLM-METHOD-004ESM-2 650M ProteinGym aggregate Spearman baseline = 0.414ESM-2 650M ProteinGym cross-leaderboard aggregate Spearman baseline (Average_Spearman column, accessed 2026-05-08)
Data source
strongCLM-METHOD-004Verified
CLM-METHOD-005Cohort outperforms ProteinGym ESM-2 baseline by +28.0% relativeRelative delta of cohort median measured against the ESM-2 650M cross-leaderboard aggregate baseline (cohort-frame)
Data source
strongCLM-METHOD-005Verified
CLM-METHOD-006Calmodulin Tier-2 stratum: rho = 0.2116Calmodulin (CALM1_HUMAN, Weile et al. 2017) DMS — Tier-2 named-weakness stratum, Spearman correlation
Data source
strongCLM-METHOD-006Verified
CLM-METHOD-006bCalmodulin Tier-2 stratum: 1,813 variantsCalmodulin (CALM1_HUMAN, Weile et al. 2017) DMS — Tier-2 named-weakness stratum, variant count
Data source
strongCLM-METHOD-006bVerified
CLM-METHOD-007Calmodulin Tier-2 stratum: -48.9% relative to cross-leaderboard baselineRelative delta of the Calmodulin named-weakness stratum measured against the ESM-2 650M cross-leaderboard aggregate baseline (cohort-frame)
Data source
strongCLM-METHOD-007Verified
CLM-METHOD-008Median taken at per-DMS-entry grain (N=7), not per-protein grainCohort-median grain convention — computed per-DMS-entry, not per-protein
Data source
strongCLM-METHOD-008Verified

Verified as an accurate citation — we confirmed this is reported by the cited source. We did not independently reproduce it ourselves.

CLM-METHOD-009Baseline sourced from catalog-first frozen archive; not regeneratedFrozen-archive provenance for the ESM-2 650M cross-leaderboard baseline (catalog-first, accessed 2026-05-08)
Supporting citation
strongCLM-METHOD-009Verified

Verified as an accurate citation — we confirmed this is reported by the cited source. We did not independently reproduce it ourselves.

CLM-METHOD-010Surface-class dual-field convention: cohort-frame cites cross-leaderboard, per-post cites per-proteinConvention for which baseline each page type compares against (summary pages vs individual protein pages)
Data source
strongCLM-METHOD-010Verified

Verified as an accurate citation — we confirmed this is reported by the cited source. We did not independently reproduce it ourselves.

CLM-METHOD-011Named-weakness routing: explicit disclosure, not collapsed, not framed as coverage failureDisclosure context for the openly-stated Calmodulin binding-affinity weakness
Data source
strongCLM-METHOD-011Verified

Verified as an accurate citation — we confirmed this is reported by the cited source. We did not independently reproduce it ourselves.

CLM-METHOD-012Methodology figures are verified-claim-grain, not prose-graininternal knowledge base decision record; internal knowledge base decision record
Supporting citation
strongCLM-METHOD-012Verified

Verified as an accurate citation — we confirmed this is reported by the cited source. We did not independently reproduce it ourselves.

CLM-METHOD-013pinned validation summary authored + independently verified by VeritasIndependent verification record backing the cohort substrate's verified status
Supporting citation
strongCLM-METHOD-013Verified

Verified as an accurate citation — we confirmed this is reported by the cited source. We did not independently reproduce it ourselves.

CLM-METHOD-015Methodology page = surface-class A cohort-frame only; per-protein figures live in per-post pagesReference to the page-type figure convention (which comparisons appear on summary vs per-protein pages)
Data source
strongCLM-METHOD-015Verified

Verified as an accurate citation — we confirmed this is reported by the cited source. We did not independently reproduce it ourselves.

CLM-METHOD-TRUST-001AI-related retractions surged sharply in 2023 and remained elevated into 2024Frontiers in Research Metrics and Analytics 2025 bibliometric review of AI-related retractions
Data source
strongCLM-METHOD-TRUST-001Verified
CLM-METHOD-TRUST-00251.1% of retracted articles retain above-average citations after retractionFrontiers in Research Metrics and Analytics 2025 bibliometric review of AI-related retractions
Data source
strongCLM-METHOD-TRUST-002Verified
CLM-METHOD-TRUST-003Median publication-to-retraction time: 550 days (~1.5 years)Frontiers in Research Metrics and Analytics 2025 bibliometric review of AI-related retractions
Data source
strongCLM-METHOD-TRUST-003Verified
CLM-METHOD-TRUST-004ChatGPT 4o-mini × 6,510 evaluations of retracted articles flagged zero retractionsWiley Learned Publishing 2025 — 'ChatGPT tends to ignore retractions' (empirical evaluation)
Data source
strongCLM-METHOD-TRUST-004Verified
CLM-METHOD-TRUST-005aAI-generated content with fabricated citations documented in published literature (Zhang et al. 2024 + Harvard MR audit; Retraction Watch case-studies adjacent-context)Zhang et al. 2024 *Surfaces and Interfaces* retraction (Crossref RETRACTED: prefix) + Harvard Misinformation Review GPT-fabricated-papers audit (2024). Retraction Watch case studies cited as adjacent-context only.
Aggregate citation
moderateCLM-METHOD-TRUST-005aVerified

Verified as an accurate citation — we confirmed this is reported by the cited source. We did not independently reproduce it ourselves.

CLM-METHOD-TRUST-005bIndependent verification of AI inference-tool validation pages remains uncommonObservation-from-absence — the basis is the lack of documented cases of independent-verification implementations in the inference-tool-validation-page record as of 2026-05-21 review.
Observation from absence
moderateCLM-METHOD-TRUST-005bVerified

Verified as an accurate citation — we confirmed this is reported by the cited source. We did not independently reproduce it ourselves.

CLM-METHOD-TRUST-006Internal-pipeline near-miss — initial PASS overturned on re-verification against open primary source (internal usage-statistic near-miss)Internal-pipeline audit-trail artifact
Internal audit
strongCLM-METHOD-TRUST-006Verified

Verified as an accurate citation — we confirmed this is reported by the cited source. We did not independently reproduce it ourselves.

Substantiated against the cohort-frame validation claim list (version a67dc48) and the editorial validation claim list (version ee7becd) — two immutable, commit-grain version pins captured at build time, providing a cross-checked source trail. The result was independently verified by Veritas, the verification authority.