Tim Richardson, a software engineer at Genomics England in London, has taken a hard look at how AI handles gene splicing. He went back to a major dataset from a 2019 study by Chong et al. That study introduced the multiplexed functional assay of splicing using Sort-seq, or MFASS. Richardson did the reanalysis on his own, outside the original team.
In a comparative evaluation using the MFASS dataset, SpliceAI and Pangolin each identified 63–66 experimentally confirmed splicing disruptions in their top-100 ranked variants, highlighting the close performance of these leading AI predictors.
Richardson’s reanalysis, published on rewire.it, does more than crunch numbers. He checks if today’s AI tools match up with real experimental results. This kind of challenge is rare. It echoes the tough, data-driven skepticism seen in other genomics work, like the NIH single cell brain atlas project reported earlier.
Richardson’s full findings require registration to access. But the message is clear. The genomics field cannot just trust AI predictions without strong lab proof. MFASS, with its fast, fluorescence-based readout, lets scientists compare algorithms to real biological outcomes. Not just theory. A follow-up report from rewire.it gives the numbers: in a similar MFASS subset, Pangolin checked 8,301 variants and found 314 positives, with 65 hits in the top-100 and a recall of 20.70%. SpliceAI checked 8,194 variants, found 308 positives, got 64 top-100 hits, and a recall of 20.78%.
The positive label in MFASS reflects assay-defined splicing disruption, not a clinical diagnosis or a universal effect across all tissues. Evaluations of splicing predictors can be sensitive to the chosen variant population and comparison method, and as of late September 2026, rewire.it emphasized that metric recalculations may shift the ranking order, with no clear winner established.