Validation & Methods

How we measure, and where it stops.

Every number on this site comes from a run you can repeat. This page gives the denominators, the scoring rules, the comparison tools, the P-values — and the limits, in the same words we use with investors.

Benchmark samples
35 held-out · 1.19 M reads
Automated tests
288, all passing
Peer review
Two papers under revision
Last updated
September 2026
01 · The rescue count

6 of 22 — what is counted, and what isn't.

Denominator: 22 SELEX pools that had enriched but yielded no binder by abundance ranking, from the partner lab's archive and external groups, re-analysed with the current pipeline.

  • A rescue means the re-analysis produced a consistent structural family and a synthesis-ready shortlist that the requesting group took forward. Six pools met that bar.
  • For the other sixteen the pipeline found no consistent family under the stated conditions. We report that as "no candidate found". It is not evidence that no binder exists in the pool.
  • These are internal laboratory records. They are not part of the two manuscripts under revision and have not been peer-reviewed. We share the per-pool breakdown under NDA on request.
  • The count is a re-analysis outcome, not a validation rate: it says how often a stalled pool gave us something worth ordering, not how often the ordered sequences bound.
Counting rule
failed_pools.csv
› 22 pools · enriched, no binder by abundance
› 6 consistent family + shortlist taken forward
› 16 no candidate found (not "no binder")
› source internal records · not peer-reviewed
› breakdown per pool, under NDA
02 · Six-tool benchmark

Same data, same scoring, every tool.

35 held-out test samples (about 1.19 million reads). Six tools were run on the same inputs and scored with the same rule: composite = 0.6 × F1 + 0.4 × motif coverage + 0.1 × discovery reward. Differences against the top tool were tested with a paired one-sided Wilcoxon test; all are significant at P < 0.05.

ToolCompositeF1P (vs. fast variant)
AptaPilot — fast variant (pure sequence)0.7770.741
AptaPilot — structure-aware tool0.5580.4854.5 × 10⁻⁴
FASTAptamer0.4420.2382.5 × 10⁻⁵
Previous DBSCAN method0.4010.2721.7 × 10⁻⁷
FSBC / pFSBC0.3670.0392.3 × 10⁻⁶
AptaSUITE0.3030.0682.5 × 10⁻⁷

The fast variant's 0.777 has a 95% confidence interval of 0.669–0.869. Ablation: keeping the front end identical and only replacing DBSCAN with a fixed-k split moved the composite from 0.401 to 0.777, improving 28 of 35 samples (P = 1.7 × 10⁻⁷). Two things to read carefully: this benchmark scores against sequence-defined reference families, which favours pure-sequence methods; and the fast variant contains no structural component at all. The structure-aware tool is the one that ranks second here but leads on the real-ground-truth experiment below — and it is the only one that returns a folded structural motif per family, which is what truncation decisions rest on.

03 · Independent checks

Three places the answer was known in advance.

CheckAptaPilotComparisonWhat it is
Real doped-SELEX experimentF1 0.89previous method 0.72 (needing 126 clusters)Public doped-SELEX library, two spiked families. The only set with verifiable real ground truth; the structure-aware tool leads all methods here.
Simulated pool, family recoveryF1 0.76 / 0.61FSBC 0.51 · previous 0.18Eight known families; ground truth fixed by construction.
Independent public datasets29 / 30 · 48 / 48Ishida et al. (2020), run without modification; different targets and lengths.
G-quadruplex detection86.7%ViennaRNA 75.4% · seqfold 0%203 experimentally determined DNA G-quadruplexes.
ssDNA fully-correct structures52.0%ViennaRNA 36.0%Third-party independent benchmark, protein-free set, n = 25; F1 0.70 and MCC 0.70 vs. F1 0.50.
Engineering quality288 testsAll automated tests pass; eight classic aptamers folded as unintended cross-class checks without error.
04 · Limits

Where the methods do not apply.

Deeply enriched single rounds

On a pool that is already deeply enriched in a single round, family classification does not predict binding better than plain abundance ranking. Those are two different tasks; we say so before quoting.

Fixed k = 6

The fixed-k split is an engineering setting tuned for pools with 2–6 families. It is not a statistical inference of the true family count; pools with more families need k adjusted, and we flag when that happened.

Folding is trend-level

Absolute melting temperatures are still uncalibrated (best configuration RMSE 9.47 °C). AptaFold is positioned as a trend-level screening tool, not a thermodynamic instrument — which is why Kd comes from your assay, never from a prediction.

Not yet peer-reviewed

The two supporting manuscripts are under revision. The benchmarks above have not passed peer review; the rescue counts come from internal records and sit outside both papers. We update this page when that changes.

Candidates are not binders

Every Rescue and Characterize deliverable is a priority order for synthesis and testing. Nothing in them has been shown to bind until your assay says so. Only Discovery ends with a measured Kd.

Metrics mean what they measure

F1 0.89 is family recovery on a spiked library. 6 of 22 is a re-analysis outcome. Neither is a probability that your pool contains a binder, and we will not present them as one.

05 · Reproducibility

Rebuilt from raw data with one command.

  • All benchmark inputs, parameters, scoring and plotting code run end to end from a single command; the datasets and the comparison tools are public, so a third party can re-run them without us.
  • Client reports state every threshold applied and carry a checksum for each input file; a regression check confirms a rebuild reproduces the delivered figures exactly.
  • A second, independently written analysis is run over every project's data and compared with the first. On the case-study project that comparison caught a real grouping flaw in one of them.
  • Under NDA we share the per-pool rescue breakdown, the benchmark notebooks and the manuscript drafts.
Sources
benchmark: 35 held-out samples · 1.19 M reads
doped-SELEX: public library, two families
public sets: Ishida et al. 2020
G4 set: 203 experimental DNA G4s
ssDNA set: third-party, n = 25
figures: Sept 2026 · manuscripts under revision

Want the notebooks, not the summary?

Ask for the per-pool breakdown and benchmark runs under NDA.
Request the validation package →[email protected]
AptaPilot — aptamer discovery, done right · Waterloo, CanadaFounder · Xiaohan Zhang