Verification Report

Verification Report — ESM-C vs ESM-2 Zero-Shot Benchmark

Published · Updated

This is the claims register for We Ran a 10× Bigger Protein Model. It Didn't Rank Variants Any Better. — the ESM-C vs ESM-2 650M zero-shot ProteinGym benchmark. Each claim below was independently checked by Veritas against the producer's claims substrate, and the result for each is served from a hash-verified record.

Claims Register

Claim IDClaimSourceStrengthStatus
CLM-ESMC6B-0001ESMC-6B grouped-avg Spearman = 0.419 on 201 assays, below ESM-2 650M's 0.431 (paired delta -0.0081, t=-0.846, p=0.40, NS; 6B wins 50.2%) — no scaling gain.NeuroAutomata ESMC adoption benchmark, Run 2 (ESMC-6B); ProteinGym leaderboard upstream ce80571c26
Data source
moderateCLM-ESMC6B-0001Verified
CLM-ESMC6B-0002ESMC-6B (0.419) is below ESMC-600M (0.425) on 201 assays (paired delta -0.0044, t=-0.663, p=0.51, NS; 6B wins 46.8%) — scaling 10x does not help.NeuroAutomata ESMC adoption benchmark, Runs 1+2
Data source
moderateCLM-ESMC6B-0002Verified
CLM-ESMC6B-0003ESMC-6B wins only 94/201 (46.8%) paired assays vs ESMC-600M — below break-even; bigger does not beat smaller on our zero-shot task.NeuroAutomata ESMC adoption benchmark; paired per-assay win-rate over 201 common assays
Data source
moderateCLM-ESMC6B-0003Verified
CLM-ESMC6B-0004Point estimates decrease monotonically with ESMC scale (ESM-2 0.431 >= ESMC-600M 0.425 >= ESMC-6B 0.419); scaling to 6B yields no zero-shot ProteinGym gain over ESM-2 650M.NeuroAutomata ESMC adoption benchmark, Runs 1+2; three-way 201-assay comparison
Data source
moderateCLM-ESMC6B-0004Verified

Verified as an accurate citation — we confirmed this is reported by the cited source. We did not independently reproduce it ourselves.

Substantiated against the ESM-C benchmark validation claim list(version 6add4df, captured at build time). The result was independently issued by Veritas, the verification authority.