Every number on this site comes from a run you can repeat. This page gives the denominators, the scoring rules, the comparison tools, the P-values — and the limits, in the same words we use with investors.
Denominator: 22 SELEX pools that had enriched but yielded no binder by abundance ranking, from the partner lab's archive and external groups, re-analysed with the current pipeline.
35 held-out test samples (about 1.19 million reads). Six tools were run on the same inputs and scored with the same rule: composite = 0.6 × F1 + 0.4 × motif coverage + 0.1 × discovery reward. Differences against the top tool were tested with a paired one-sided Wilcoxon test; all are significant at P < 0.05.
| Tool | Composite | F1 | P (vs. fast variant) |
|---|---|---|---|
| AptaPilot — fast variant (pure sequence) | 0.777 | 0.741 | — |
| AptaPilot — structure-aware tool | 0.558 | 0.485 | 4.5 × 10⁻⁴ |
| FASTAptamer | 0.442 | 0.238 | 2.5 × 10⁻⁵ |
| Previous DBSCAN method | 0.401 | 0.272 | 1.7 × 10⁻⁷ |
| FSBC / pFSBC | 0.367 | 0.039 | 2.3 × 10⁻⁶ |
| AptaSUITE | 0.303 | 0.068 | 2.5 × 10⁻⁷ |
The fast variant's 0.777 has a 95% confidence interval of 0.669–0.869. Ablation: keeping the front end identical and only replacing DBSCAN with a fixed-k split moved the composite from 0.401 to 0.777, improving 28 of 35 samples (P = 1.7 × 10⁻⁷). Two things to read carefully: this benchmark scores against sequence-defined reference families, which favours pure-sequence methods; and the fast variant contains no structural component at all. The structure-aware tool is the one that ranks second here but leads on the real-ground-truth experiment below — and it is the only one that returns a folded structural motif per family, which is what truncation decisions rest on.
| Check | AptaPilot | Comparison | What it is |
|---|---|---|---|
| Real doped-SELEX experiment | F1 0.89 | previous method 0.72 (needing 126 clusters) | Public doped-SELEX library, two spiked families. The only set with verifiable real ground truth; the structure-aware tool leads all methods here. |
| Simulated pool, family recovery | F1 0.76 / 0.61 | FSBC 0.51 · previous 0.18 | Eight known families; ground truth fixed by construction. |
| Independent public datasets | 29 / 30 · 48 / 48 | — | Ishida et al. (2020), run without modification; different targets and lengths. |
| G-quadruplex detection | 86.7% | ViennaRNA 75.4% · seqfold 0% | 203 experimentally determined DNA G-quadruplexes. |
| ssDNA fully-correct structures | 52.0% | ViennaRNA 36.0% | Third-party independent benchmark, protein-free set, n = 25; F1 0.70 and MCC 0.70 vs. F1 0.50. |
| Engineering quality | 288 tests | — | All automated tests pass; eight classic aptamers folded as unintended cross-class checks without error. |
On a pool that is already deeply enriched in a single round, family classification does not predict binding better than plain abundance ranking. Those are two different tasks; we say so before quoting.
The fixed-k split is an engineering setting tuned for pools with 2–6 families. It is not a statistical inference of the true family count; pools with more families need k adjusted, and we flag when that happened.
Absolute melting temperatures are still uncalibrated (best configuration RMSE 9.47 °C). AptaFold is positioned as a trend-level screening tool, not a thermodynamic instrument — which is why Kd comes from your assay, never from a prediction.
The two supporting manuscripts are under revision. The benchmarks above have not passed peer review; the rescue counts come from internal records and sit outside both papers. We update this page when that changes.
Every Rescue and Characterize deliverable is a priority order for synthesis and testing. Nothing in them has been shown to bind until your assay says so. Only Discovery ends with a measured Kd.
F1 0.89 is family recovery on a spiked library. 6 of 22 is a re-analysis outcome. Neither is a probability that your pool contains a binder, and we will not present them as one.