frozen measured-activity model cross-lab AUROC 0.80 · chance 0.50 n = 93,435 · held-out lab
GitHub
PRE-SYNTHESIS FAILURE TRIAGE

Catch the AI-designed enhancers that will fail, before the lab synthesizes them.

CisFalcon scores a designed DNA sequence against a frozen external measured-activity model and predicts whether it will misfire in the wrong cell type. It ranks a batch safest-first, so a lab spends its bench budget on the designs most likely to hold up.

0.80
CROSS-LAB AUROC
chance 0.50 0.8013 · n=93,435 · a different lab
41%
FEWER FAILURES
Rank safest-first and synthesize the safer half. Measured within one target cell and one generator, the honest conditioned reduction is 41%. Pooled across a mixed batch it reads 70% (6.29% to 1.91%), but a sequence-free rule using only stratum base rates scores about 90% pooled, so the pooled number is not per-design signal. 41% is a macro average over 14 of the 24 (cell × generator) strata, the ones with at least 10 failures and 10 passes, covering 87,573 of 93,435 designs; per stratum it runs 7% (SKNSH/sa_rep) to 91% (HepG2/hmc), so what one lab gets depends on which stratum it deploys in. Reproduce the per-stratum table with triage_conditioning_check.py.
Clinical grounding. The lineage-driver transcription factors it names are validated human disease genes: GATA1 (X-linked thrombocytopenia/anemia), HNF4A (MODY1), TAL1 (T-ALL), KLF1, FOXA2. It reads the regulatory grammar whose disruption causes disease.
Triage ranker, not a hard gate. The pooled 0.80 is the mixed-batch number; fully conditioned per-sequence it is 0.66 macro across 14 of the 24 strata (range 0.53 to 0.96), still above the 0.50 chance line.
CLOSED-LOOP VERIFICATION
FLAG
real hero · target HepG2 · a real BODA/Malinois design (Gosai/Tewhey 2024)
HepG2 target
K562 off-target
most-active cell: K562
SPECIFICITY GAP · log2FC(HepG2) − max off-target −3.0
−40 · fail line+4
Predicted to fail. Most-active cell is K562, not the HepG2 target. Specificity gap −3.0 (fail when gap ≤ 0). Wet-lab measured K562 +6.5, HepG2 −0.1 confirms it.
one model · one set of weights
Example library target HepG2 · ranked by risk
All · Flagged · Clear ·
DESIGNGAPOFF-TARGETSTATUS
loading designs…
base failure rate 6.29% · 1 in 16 safest half → 1.91%
OR CHECK YOUR OWN SEQUENCE
CLOSED-LOOP VERIFICATION / select a design
scoring the example library…
BATCH TRIAGE

Spend the synthesis budget where it holds up

Rank a batch of AI-designed enhancers safest-first, then synthesize down the ranked list. The specificity failures you actually make drop sharply, so a fixed bench budget buys more designs that work.

6.29%
BASE FAILURE RATE
about 1 in 16 designs misfires in the wrong cell type, in this one library (93,435 Gosai/Tewhey designs, cross-lineage in vitro). In-vivo cortical screens report far higher failure rates.
41%
FEWER FAILURES
macro across 14 of the 24 cell x generator strata (the deployment case), covering 87,573 of 93,435 designs. Per stratum it runs 7% to 91%, so a lab in SKNSH/sa_rep gets 7%. Pooled it reads 70%, but a sequence-free stratum-prior rule beats that at about 90%.
2.59×
RISKIEST-2% ENRICHMENT
macro across the same 14 strata (the deployment case). Three of them enrich below 1.0× (SKNSH/fsp 0.65×, K562/sa_rep 0.60×, HepG2/sa_rep 0.90×), i.e. no better than random within those strata. Pooled it reads 7.74× at PPV 0.487, on the same mixed-batch basis the 70% figure above is disowned for.
43%
FAILURES CAUGHT
flag the riskiest 10%, capture 43% (4.3×). Pooled basis, and unlike the two cards above it has no conditioned counterpart measured yet, so read it as an upper bound.
SPECIFICITY FAILURES IN WHAT YOU SYNTHESIZE
safest-first, ranked by CisFalcon
7% 2% 0 synthesize blind 6.29% safest half · 1.91% fraction of batch synthesized, safest-first →
The gap between the dashed blind line and the curve is the failures avoided by ranking. triage_curve.py
OPERATING POINTS
Synthesize the safest half
failure rate 1.91% pooled (70%); conditioned within cell x generator the honest figure is a 41% reduction
Flag the riskiest 10%
captures 43% of all failures, 4.3× the base rate
Flag the riskiest 2%
49% truly fail, 7.74× enrichment (PPV 0.487)
Each pursued design costs real materials and weeks of turnaround; enter your own cost in the tool below to see the wasted spend a safest-first triage averts on your batch.
Risk-ranked view, computed POOLED across the Gosai cross-lab population (n = 93,435) where the base rate is 6.29%. Pooled is the wrong basis for a single lab's batch: conditioned within one cell type and one generator the reduction is 41% rather than 70%, and the riskiest-2% enrichment is 2.59× rather than 7.74×. Reproduce both with triage_conditioning_check.py. CisFalcon is a triage ranker, not a hard gate.
RANK YOUR BATCH

Paste or upload a design library and rank it live

Any size. CisFalcon scores each design against the frozen external model, ranks safest-first, and returns a ranked CSV. Enter your cost per design to see the wasted synthesis a safest-half triage averts.

Designs (one per line or FASTA)
Triage ranking
Paste designs and rank them, or load the example set.
VALIDATION

Measured on a different lab's designs the model never saw

Cross-lab is the flagship: a frozen published activity model, scored on 93,435 designs from a different lab, a different design process, and generators it never saw (Gosai et al. 2024), with zero sequence overlap. Both the deployment number and the harder fully-conditioned number are reported.

0.80 AUROC chance 0.50 false positive rate
0.8013
cross-lab AUROC · design-level 95% CI [0.796, 0.807]
cluster 95% CI over all 24 cell × generator strata (3 cells × 8 generators): [0.614, 0.895], the honest one, wide because the ranker is strong on failure-prone generators and near chance on the best-optimized ones (bootstrap_ci.py)
n = 93,435 designs
held-out lab · Gosai/Tewhey 2024
zero sequence overlap
GC 0.51 length 0.50 random 0.50
WHERE THE SIGNAL LIVES
The pooled 0.80 is the mixed-batch deployment number. Remove the base rates and genuine per-sequence signal remains, above chance at every level.
pooled, mixed batch0.80
within target cell0.75
cell-prior only, no sequence0.68
fully conditioned per-sequence0.66
left edge = 0.50 chance · every level clears it. The bottom row is a macro average over the 14 of those 24 strata that carry at least 10 failures and 10 passes, running 0.53 to 0.96, not a single pooled quantity; 7 of the 14 sit below 0.59. cell_prior_baseline.py
TWO-HEAD ENSEMBLE
activity 0.8016 and accessibility 0.789, rank-averaged to 0.8064. Nothing trained, fixed before the number was read. These three come from the ensemble's own code path over design_gaps.csv (ensemble_ci.py); the 0.8013 headline elsewhere on this page is the same model on the flagship path over designed_scored.csv.
CALIBRATED, NOT ASSERTED
held-out calibration error ECE 0.0031 (isotonic; mean of both split directions, 0.0029 / 0.0032). Near the middle it is close to exact: the 0.4-0.5 bin runs predicted 0.420 against observed 0.420 on n=355. The high-risk tail is thinly measured and ECE cannot show it, since 79% of held-out designs sit in the lowest bin: 0.6-0.7 runs 0.696 predicted against 0.619 observed on n=42. Trust the ranking; treat an absolute number above ~0.6 as noisy.
BRAIN-DERIVED LINE
SKNSH cross-lab 0.73; safest half cuts failures 32% macro, conditioned within (SKNSH x generator), against an 82% sequence-free generator-prior null (58% pooled). A primary-neuron MPRA is the honest next step.
IN-VIVO WITHIN-CORTEX · NEIGHBOR RESOLUTION
The paradigm, not this tool, tested one tissue deeper: an independent cortical model (DeepBICCN2, per-subclass for astrocyte, oligodendrocyte, and neuron subtypes) scores 532 in-vivo AAV-screened enhancers (Ben-Simon et al., Cell 2025). Scored here on 204 of the 211 hard-labelled enhancers (140 On / 64 Off) whose target subclass DeepBICCN2 can score; the other 7 have no output head for their subclass. All three bars are matched on those same 204 rows and paired, so they are like-for-like.
specificity gap · sequence only0.706
measured accessibility0.687
measured accessibility-specificity (gini)0.731
AUROC 0.706, 95% CI [0.63, 0.77], excludes chance. Against the measured features on the matched 204, the sequence gap is statistically indistinguishable, not ranked: vs gini paired delta +0.026, 95% CI [-0.066, +0.115]; vs accessibility -0.018, 95% CI [-0.103, +0.067]. Both include zero. The honest claim is that the gap recovers from sequence what an assay would tell you, not that either ranks above the other. An earlier version of this panel ranked them by comparing the n=204 gap against n=211 baselines; that unmatched comparison was retracted on 2026-07-16 (see PREREG-ERRATA.md). Reproduce with python within_neighbor/matched_paired.py. Independent model, in-vivo functional labels. The honest next step is a functional within-cortex model.
HETEROGENEOUS RANKER · LEAD WITH THE RANGE
within-generator macro 0.73
FastSeqProp 0.62 (best-optimized, near coin-flip) Hamiltonian MC 0.92 (failure-prone)
Discrimination is highest exactly where failures cluster, and near chance on the best-optimized generators. Strongest on failure-prone and out-of-distribution designs.
HONEST BOUNDARIES
A triage ranker, not a hard gate. At the 6.29% base rate a hard gate has poor precision (PPV 0.24 at recall 0.5). Its value is prioritization.
The closed-loop flip is an in-silico consistency check on one design, ahead of wet-lab validation. It is not a wet-lab result.
The in-distribution number (0.896) carries a residual near-duplicate leakage caveat and is a sanity check only. The cross-lab result is the headline and is immune by construction.
Every cross-lab number re-derives to four decimals from the committed per-design scores with a plain rank-sum AUROC. reproduce_flagship.py, one-click Colab.
RESEARCH-TRACK EXTENSION

The same frozen model ranks single-base variant effects

Beyond cell-type specificity, the identical model scores single-nucleotide variants on a fully independent saturation-mutagenesis benchmark it never saw (Kircher et al. 2019). A second, disjoint test of the same triage signal, with the scorer loaded bit-identical to the flagship.

0.88
K562 LARGE-EFFECT AUROC
separates high- from low-impact variants
r 0.60
PEARSON · K562
n = 1,407 distinct variants. On all 2,814 assay rows it reads 0.58, but rows repeat a variant across elements.
5,065
DISTINCT SNVs SCORED
across 11,005 assay rows · Kircher 2019 · GSE126550
0.64
HepG2 AUROC
Pearson r 0.25 on 3,658 distinct variants. On all 8,191 assay rows it reads 0.29, but rows repeat a variant across elements.
LARGE-EFFECT VARIANT AUROC
high- vs low-impact single-base variants, held-out
K562 · r 0.580.88
HepG2 · r 0.290.64
left edge = chance 0.50right = 1.0
WHAT IT SHOWS
It separates high-impact from low-impact variants well, the same triage signal that flags failing enhancers. It does not calibrate absolute effect size, and it is near-blind on the LDLR element, reported with an honest per-element breakdown.
Model loading is bit-identical to the flagship scorer. The K562 headline reproduces from the committed per-SNV data with one command.
METHODS

A deterministic gate, an agent diagnosis, checked against external ground truth

Nothing is trained here. The ranking comes from a frozen, published measured-activity model; the agents add the redesign guidance a bare score cannot.

WHAT IT IS
A pre-synthesis failure-risk triage for AI-designed, cell-type-specific enhancers. It reads a designed DNA sequence and predicts, before any DNA is synthesized, whether the design will misfire in the wrong cell type in a wet-lab MPRA, grounded in an external measured-activity model rather than self-report.
WHAT IT IS NOT
It does not design DNA. It is not a hard kill-switch or abort gate, and it is not a wet-lab assay. It exposes no generative optimization loop and creates no new design capability. It is a defensive falsifier that prevents wasted synthesis.
HOW IT WORKS
01 · Gate
Frozen external model
The atlas DHS64-MPRA activity models (Castillo-Hair et al. 2025), three fold models ensembled. Input a 500 bp sequence, output predicted log2FC across 12 human cell lines. Nothing trained. The two-head activity-plus-accessibility ensemble (0.8064 vs 0.8016 activity alone) is a measured comparison in PREREG.md, not what this tool runs: what you are using is the activity ensemble.
02 · Label
Transparent fail rule
A design fails specificity when it is not most-active in its intended target.
gap = log2FC(target) − max(off-target)
FAIL := gap ≤ 0
03 · Diagnose
Agent redesign layer
Four Claude agents turn a bare FAIL into a mechanism: which off-target cell wins and which driver motifs to ablate or install. A measured ablation shows the agents do not improve classification over the gate; their value is the redesign guidance.
CLINICAL GROUNDING
CisFalcon attaches the Mendelian disease to each of the ten erythroid and hepatocyte driver factors it tracks. It reads regulatory grammar whose disruption causes human disease, so a specificity failure is not a cosmetic QC issue.
GATA1 · X-linked thrombocytopenia with dyserythropoietic anemia · OMIM 300367 HNF4A · MODY1 diabetes · OMIM 125850 TAL1 · T-cell acute lymphoblastic leukemia KLF1 · congenital dyserythropoietic anemia FOXA2 · congenital hyperinsulinism
In a real interneuron-specific enhancer-AAV therapy for Dravet syndrome, driving the gene in all neurons instead of only interneurons increased mortality in mice, while the cell-type-specific version was safe and corrected the seizures (Mich et al., Sci Transl Med 2025). That is the wrong-cell-type firing CisFalcon flags from sequence, before synthesis.
DATA & REPRODUCE
Cross-lab benchmark: Gosai et al. 2024 MPRA, 93,435 designs, a different lab and design process the model never saw.
One command, no GPU: reproduce_flagship.py re-derives AUROC 0.8013 in pure numpy from the committed per-design scores. One-click Colab in the browser.
We verify the verifier: a separate Claude Science session independently re-derived every number in the pre-conditioning release to the decimal from the committed data (docs/SCIENCE-AUDIT.md). The conditioned figures above (41%, 2.59×, the 91% null) post-date that audit and are reproduced instead by triage_conditioning_check.py, which asserts the audited pooled figures as a positive control. Built with Claude Code; full methodology in PREREG.md (frozen, hash-locked) plus PREREG-ERRATA.md, which carries dated corrections to two of its reported figures.
OPEN SOURCE
github.com/belumume/cisfalcon live tool · cisfalcon-lifesci.fly.dev
MIT licensed. The predictor is a public, already-released model.