Data availability
All primary data supporting the findings of this study are available within the Article and via Zenodo at https://doi.org/10.5281/zenodo.19179097 (ref. 62). Source data are provided with this paper.
Code availability
Protein language models were implemented using the code and pretrained parameters available from their official GitHub repositories, following the provided instructions. The ESM2 code and models are available via GitHub at https://github.com/facebookresearch/esm. We used XGBoost and other classical ML models (for example, logistic regression, random forest) from the RAPIDS cuML library for accelerated training on graphics processing unit (GPU). Fine-tuned DL model checkpoints and source code will be publicly released on Hugging Face, Docker and Github (https://github.com/fangzhe3/DeepSCan)57. The GitHub repository also introduces three demo scenarios using the DeepSCan web server or Docker image. The Docker image of the Genesis Quant 8,000 model enables users to process large-scale sequence analyses using GPUs. The supplementary software contains the DeepSCan web server application (version: app_v7.0_deployed_on_AWS), Genesis model training code (version: Genesis_Quant_8k_attention_pooling2_LLRD_v2), Genesis and Omni model inference codes and additional computational analyses codes. The DeepSCan models assume Python (v.3.11.5) + PyTorch (v.2.5.1 with CUDA v.12.4)-based workflow.
References
-
Overington, J. P., Al-Lazikani, B. & Hopkins, A. L. How many drug targets are there? Nat. Rev. Drug Discov. 5, 993–996 (2006).
-
Hegde, R. S. & Keenan, R. J. The mechanisms of integral membrane protein biogenesis. Nat. Rev. Mol. Cell Biol. 23, 107–124 (2022).
-
Bugge, K., Lindorff-Larsen, K. & Kragelund, B. B. Understanding single-pass transmembrane receptor signaling from a structural viewpoint—what are we missing?. FEBS J. 283, 4424–4451 (2016).
-
Pogozheva, I. D. & Lomize, A. L. Evolution and adaptation of single-pass transmembrane proteins. Biochim. Biophys. Acta Biomembr. 1860, 364–377 (2018).
-
Grupp, S. A. et al. Chimeric antigen receptor-modified T cells for acute lymphoid leukemia. N. Engl. J. Med. 368, 1509–1518 (2013).
-
Muller, F. et al. CD19 CAR T-cell therapy in autoimmune disease—a case series with follow-up. N. Engl. J. Med. 390, 687–700 (2024).
-
Kremer, J. M. et al. Tocilizumab inhibits structural joint damage in rheumatoid arthritis patients with inadequate responses to methotrexate: results from the double-blind treatment phase of a randomized placebo-controlled trial of tocilizumab safety and prevention of structural joint damage at one year. Arthritis Rheum. 63, 609–621 (2011).
-
Kishimoto, T. & Kang, S. IL-6 revisited: from rheumatoid arthritis to CAR T cell therapy and COVID-19. Annu. Rev. Immunol. 40, 323–348 (2022).
-
Topalian, S. L. et al. Safety, activity, and immune correlates of anti-PD-1 antibody in cancer. N. Engl. J. Med. 366, 2443–2454 (2012).
-
Han, Y., Liu, D. & Li, L. PD-1/PD-L1 pathway: current researches in cancer. Am. J. Cancer Res. 10, 727–742 (2020).
-
James, P. A. et al. 2014 evidence-based guideline for the management of high blood pressure in adults: report from the panel members appointed to the Eighth Joint National Committee (JNC 8). JAMA 311, 507–520 (2014).
-
Burnett, J. C. Jr. et al. Atrial natriuretic peptide elevation in congestive heart failure in the human. Science 231, 1145–1147 (1986).
-
Harayama, T. & Riezman, H. Understanding the diversity of membrane lipid composition. Nat. Rev. Mol. Cell Biol. 19, 281–296 (2018).
-
Casares, D., Escriba, P. V. & Rossello, C. A. Membrane lipid composition: effect on membrane and organelle structure, function and compartmentalization and therapeutic avenues. Int. J. Mol. Sci. 20, 2167 (2019).
-
Brandizzi, F. et al. The destination for single-pass membrane proteins is influenced markedly by the length of the hydrophobic domain. Plant Cell 14, 1077–1092 (2002).
-
Stavropoulos, I. et al. Protein disorder and short conserved motifs in disordered regions are enriched near the cytoplasmic side of single-pass transmembrane proteins. PLoS ONE 7, e44389 (2012).
-
Teese, M. G. & Langosch, D. Role of GxxxG motifs in transmembrane domain interactions. Biochemistry 54, 5125–5135 (2015).
-
Rapoport, T. A., Goder, V., Heinrich, S. U. & Matlack, K. E. Membrane-protein integration and the role of the translocation channel. Trends Cell Biol. 14, 568–575 (2004).
-
Honigmann, A. & Pralle, A. Compartmentalization of the cell membrane. J. Mol. Biol. 428, 4739–4748 (2016).
-
Lorent, J. H. et al. Structural determinants and functional consequences of protein affinity for membrane rafts. Nat. Commun. 8, 1219 (2017).
-
Diaz-Rohrer, B. B., Levental, K. R., Simons, K. & Levental, I. Membrane raft association is a determinant of plasma membrane localization. Proc. Natl Acad. Sci. USA 111, 8500–8505 (2014).
-
Levental, I., Lingwood, D., Grzybek, M., Coskun, U. & Simons, K. Palmitoylation regulates raft affinity for the majority of integral raft proteins. Proc. Natl Acad. Sci. USA 107, 22050–22054 (2010).
-
Fang, Z. et al. A modular vaccine platform for optimized lipid nanoparticle mRNA immunogenicity. Nat. Biomed. Eng. 10, 501–516 (2025).
-
Choe, J. H. et al. SynNotch-CAR T cells overcome challenges of specificity, heterogeneity, and persistence in treating glioblastoma. Sci. Transl. Med. 13, eabe7378 (2021).
-
Chung, J. B., Brudno, J. N., Borie, D. & Kochenderfer, J. N. Chimeric antigen receptor T cell therapy for autoimmune disease. Nat. Rev. Immunol. 24, 830–845 (2024).
-
Ellebrecht, C. T. et al. Reengineering chimeric antigen receptor T cells for targeted therapy of autoimmune disease. Science 353, 179–184 (2016).
-
Mitra, A. et al. From bench to bedside: the history and progress of CAR T cell therapy. Front. Immunol. 14, 1188049 (2023).
-
Arieta, C. M. et al. The T-cell-directed vaccine BNT162b4 encoding conserved non-spike antigens protects animals from severe SARS-CoV-2 infection. Cell 186, 2392–2409.e21 (2023).
-
Jumper, J. et al. Highly accurate protein structure prediction with AlphaFold. Nature 596, 583–589 (2021).
-
Watson, J. L. et al. De novo design of protein structure and function with RFdiffusion. Nature 620, 1089–1100 (2023).
-
Baek, M. et al. Accurate prediction of protein structures and interactions using a three-track neural network. Science 373, 871–876 (2021).
-
Wang, J. et al. Scaffolding protein functional sites using deep learning. Science 377, 387–394 (2022).
-
Abramson, J. et al. Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature 630, 493–500 (2024).
-
Sunny, S., Prakash, P. B., Gopakumar, G. & Jayaraj, P. B. DeepBindPPI: protein–protein binding site prediction using attention based graph convolutional network. Protein J. 42, 276–287 (2023).
-
Yang, X., Yang, S., Li, Q., Wuchty, S. & Zhang, Z. Prediction of human–virus protein–protein interactions through a sequence embedding-based machine learning method. Comput. Struct. Biotechnol. J. 18, 153–161 (2020).
-
Almagro Armenteros, J. J., Sonderby, C. K., Sonderby, S. K., Nielsen, H. & Winther, O. DeepLoc: prediction of protein subcellular localization using deep learning. Bioinformatics 33, 4049 (2017).
-
Thumuluri, V., Almagro Armenteros, J. J., Johansen, A. R., Nielsen, H. & Winther, O. DeepLoc 2.0: multi-label subcellular localization prediction using protein language models. Nucleic Acids Res. 50, W228–W234 (2022).
-
Jiang, Y. et al. MULocDeep: a deep-learning framework for protein subcellular and suborganellar localization prediction with residue-level interpretation. Comput. Struct. Biotechnol. J. 19, 4825–4839 (2021).
-
Parnamaa, T. & Parts, L. Accurate classification of protein subcellular localization from high-throughput microscopy images using deep learning. G3 7, 1385–1392 (2017).
-
Xiao, H., Zou, Y., Wang, J. & Wan, S. A review for artificial intelligence based protein subcellular localization. Biomolecules 14, 409 (2024).
-
Lu, Z. et al. Predicting subcellular localization of proteins using machine-learned classifiers. Bioinformatics 20, 547–556 (2004).
-
Kobayashi, H., Cheveralls, K. C., Leonetti, M. D. & Royer, L. A. Self-supervised deep learning encodes high-resolution features of protein subcellular localization. Nat. Methods 19, 995–1003 (2022).
-
Moreno, J., Nielsen, H., Winther, O. & Teufel, F. Predicting the subcellular location of prokaryotic proteins with DeepLocPro. Bioinformatics 40, btae677 (2024).
-
Lin, Z. et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science 379, 1123–1130 (2023).
-
Hayes, T. et al. Simulating 500 million years of evolution with a language model. Science 387, 850–858 (2025).
-
Ingraham, J. B. et al. Illuminating protein space with a programmable generative model. Nature 623, 1070–1078 (2023).
-
Ferruz, N., Schmidt, S. & Hocker, B. ProtGPT2 is a deep unsupervised language model for protein design. Nat. Commun. 13, 4348 (2022).
-
Wu, K. E. et al. Protein structure generation via folding diffusion. Nat. Commun. 15, 1059 (2024).
-
Shin, J. E. et al. Protein design and variant prediction using autoregressive generative models. Nat. Commun. 12, 2403 (2021).
-
Madani, A. et al. Large language models generate functional protein sequences across diverse families. Nat. Biotechnol. 41, 1099–1106 (2023).
-
Anishchenko, I. et al. De novo protein design by deep network hallucination. Nature 600, 547–552 (2021).
-
UniProt, C. UniProt: the Universal Protein Knowledgebase in 2023. Nucleic Acids Res. 51, D523–D531 (2023).
-
He, Y. et al. Antibodies to the A27 protein of vaccinia virus neutralize and protect against infection but represent a minor component of Dryvax vaccine-induced immunity. J. Infect. Dis. 196, 1026–1032 (2007).
-
Kaever, T. et al. Linear epitopes in vaccinia virus A27 are targets of protective antibodies induced by vaccination against smallpox. J. Virol. 90, 4334–4345 (2016).
-
Li, M. et al. Three neutralizing mAbs induced by MPXV A29L protein recognizing different epitopes act synergistically against orthopoxvirus. Emerg. Microbes Infect. 12, 2223669 (2023).
-
Fang, Z. et al. Polyvalent mRNA vaccination elicited potent immune response to monkeypox virus surface antigens. Cell Res. 33, 407–410 (2023).
-
Fang, Z. DeepSCan – an AI platform for discovery and design of potent cell surface display elements. Source code. GitHub https://github.com/fangzhe3/DeepSCan (2026).
-
Sun, Y. et al. Out-of-distribution detection with deep nearest neighbors. In Proc. 39th International Conference on Machine Learning (eds Chaudhuri, K. et al.) Vol. 162, 20827–20840 (PMLR, 2022).
-
Ye, L. et al. A genome-scale gain-of-function CRISPR screen in CD8 T cells identifies proline metabolism as a means to enhance CAR-T therapy. Cell Metab. 34, 595–614.e14 (2022).
-
Peng, L. et al. In vivo AAV-SB-CRISPR screens of tumor-infiltrating primary NK cells identify genetic checkpoints of CAR-NK therapy. Nat. Biotechnol. 43, 752–761 (2025).
-
Boutet, E. et al. UniProtKB/Swiss-Prot, the manually annotated section of the UniProt KnowledgeBase: how to use the entry view. Methods Mol. Biol. 1374, 23–54 (2016).
-
Yale University et al. DeepSCan – an AI platform for discovery and design of potent cell surface display elements. Zenodo https://doi.org/10.5281/zenodo.19179097 (2026).
Acknowledgements
We thank all members of the Chen Laboratory and various colleagues in the Department of Genetics, Systems Biology Institute, Yale Cancer Center (YCC), Yale Stem Cell Center, High Performance Computing, West Campus Analytical Chemistry Core, Biomedical Informatics and Data Sciences at Yale for assistance and/or discussion. S.C. is supported by a Cancer Research Institute Lloyd J. Old STAR Award and CLIP Award (CRI4964), the NIH/NCI (R33CA281702), the DoD (HT9425-23-1-0860), the Alliance for Cancer Gene Therapy (ACGT), the Pershing Square Sohn Cancer Research Alliance, the Sontag Foundation, D. Lu and B. Sperry. Z.F. is supported by the Canadian Institutes of Health Research (funding reference number 194053). J.S. is supported by a Yale MSTP training grant from the NIH (T32GM136651). C.Z. is supported by a Yale PhD training grant from the NIH (T32HD007149). S.-H.L. is supported by a Korean fellowship and a Leslie Warner fellowship.
Ethics declarations
Competing interests
An invention disclosure related to this work has been made to Yale University, which filed a patent application. S.C. is a (co)founder of EvolveImmune Tx, Cellinfinity Bio, MagicTime Med and Chen Consulting. The other authors declare no competing interests.
Peer review
Peer review information
Nature Biotechnology thanks Han Liang and the other, anonymous, reviewer(s) for their contribution to the peer review of this work. Peer reviewer reports are available.
Additional information
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Extended data
Extended Data Fig. 1 Workflow for identifying single-pass transmembrane protein on cell membrane (SPTM-CM).
This flow chart illustrates step-by-step process used to identify SPTM-CMs among >570,000 proteins in Swiss-Prot database (retrieved in Dec. 2023). To define the membrane orientation and domain arrangement of SPTM-CMs, we extracted information regarding topology, transmembrane domain boundaries and subcellular location in the database. The identified SPTM-CMs are classified into two categories based on their topology, namely type I (ECD at N-term) and type II (ECD at C-term) SPTM-CMs. SPTM-CMs are clustered based on their sequence identity and proteins in the human genome are highlighted. In addition, SPTM-CMs’ distribution in taxonomic superkingdoms and their associated biological taxons are summarized and displayed in the chart.
Extended Data Fig. 2 Predicting and discovering potential SPTM-CMs with ambiguous subcellular localization labels in UniProt database.
a, DeepScan ID (per-protein inference) and Topo (per-residue inference) models identified potential SPTM-CMs among proteins with ambiguous location labels. The ~220 predicted double label 1 (DL1) SPTM-CMs have >50% sequence identity with existing SPTM-CMs. The known SPTM-CMs are used as positive controls to evaluate DeepSCan model’s accuracy. b, Sequence-identity clusters (50% identity cutoff) are found in the identity matrix of predicted DL1 and known SPTM-CMs. c, t-SNE plot reduces data dimension of sequence identity matrix and visualizes the sequence distances of 222 predicted (DL1), 211 positive SPTM-CMs and 200 negative controls (DL0) on the 2-D plane. d, GO annotation analysis of ~700 DL1 proteins in UniProt database shows that >400 proteins have plasma membrane localization labels (indicated by red arrows), confirming their SPTM-CM identity. Schematic in a created in BioRender; Fang, Z. https://biorender.com/etn7ty8 (2026).
Extended Data Fig. 3 ~3500 type I CSD elements from SPTM-CMs are clustered based on their pair-wise identity matrix and visualized on t-SNE plot.
a-b, type I SPTM-CM clusters are color coded based on their identity clusters (a) or taxonomic superkingdoms (b). c, Histogram showing sequence length distribution of human type I CSD modules. d, Cluster-representative human CSD elements less than 201aa are highlighted in the t-SNE plot.
Extended Data Fig. 4 Representative flow plots showing the gating strategies to define PE or PE-high positive 293T cells that overexpressed flag-tagged A29L fused to N-term [Spike SP] and C-term [type I CSDs].
Cells were surface stained with PE anti-flag antibody. The SARS-CoV-2 XBB Spike full length (orange) and [Spike SP]-A29L-[HKU9 Spike CSD] (Green) were served as two reference points to identify CSD modules mediating higher A29 antigen surface expression. The A29L with endomembrane protein type I CSDs (blue), including NDUFA4 and ACP2, are served as negative controls. The CSD modules that outperformed reference points were highlighted depending on which quadrants they fall into in Extended Data Fig. 5d.
Extended Data Fig. 5 Amino acid usage frequency analysis revealed unique features of SPTM-CM, while A29-[type I CSD] screen identified potent CSD modules.
a-c, Amino acid usage frequency analysis of TMD and adjacent regions in type I SPTM-CM (a), type II SPTM-CM (b) and SPTM-nonCM (c) revealed a more frequent use of basic residues at the intracellular domain (ICD) and TMD interface of SPTM-CM. Proteins that have 21aa transmembrane domains in UniProt are used in this analysis and their N-term of TMD is defined as position 1. d, Flow surface staining and live/dead staining of 293 T cells overexpressing ~110 A29-[type I CSD] identified type I CSD modules that potently translocated A29 antigen to cell surface (left) and maintained cell viability comparable to negative control cells (right).
Extended Data Fig. 6 Amino acid usage frequency analysis of major protein domains uncovered conserved sequence features associated with potent CSD (CST3).
Potent CSD sequences from experimental screens (left) and DeepSCan model predictions (right) showed lower basic residue frequency in extracellular hinge (a-d), more diversified hydrophobic residues in TMD (e-f) and higher basic residue frequency in ICD (g-h). The DeepSCan-predicted CST3 SPTM-CM (right) recapitulated the sequence features found in the screened CSD modules (left). Data are presented as mean ± SEM. Pair-wise Tukey’s comparison test was used to determine statistical significance. Sample number n was shown in each group.
Extended Data Fig. 7 Representative flow plots showing the gating strategies to define PE or PE-high positive 293T cells that overexpressed flag-tagged A29L fused to N-term [Spike SP] and C-term type I natural (top) or generative (bottom) CSDs.
Cells were surface stained with PE anti-flag antibody. The SARS-CoV-2 XBB Spike full length (Orange), A29-HLA (Green) and A29-Spike (Green) were served as reference points to identify CSD modules mediating higher A29 antigen surface expression. The A29L alone and A29L with endomembrane protein type I NDUFA4 CSD (blue) were served as a negative control. The most potent CSD modules were highlighted in purple, as outlined in Fig. 2a.
Extended Data Fig. 8 Genesis Quant regressive model benchmarking revealed a high-attention region at the TMD and ICD interface, likely mediating surface translocation of potent CSD modules, including CXCL16-69.
a Eight ESM2-based deep learning models with different output layers and two XGBoost-based machine learning models were trained by augmented 4.5k dataset and benchmarked using the independent test set (n = 5 checkpoints). Compared to XGBoost-based approaches, test training showed superior performance of deep learning models trained using CLS tokens or layer-dependent learning rate decay (LLRD). b, Genesis Quant 8k attention model uncovered a high-attention-weight region at the TMD and ICD interface of potent gCSDs including CXCL16-69. c, Multiple sequence alignment of CXCL16 WT and gCSDs highlights the conserved features and unique variations associated with gCSDs. Schematic in a created in BioRender; Fang, Z. https://biorender.com/xv22ugk (2026).
Extended Data Fig. 9 Genesis Dawn single-task model benchmarking revealed optimal model architectures that serve as foundation for multi-task Omni models and achieved higher accuracy than widely used deep learning methods in SPTM-CM feature predictions.
a, Compared to first-generation DeepSCan models, Genesis single-task models were benchmarked and optimized using different output layers and training datasets, establishing a strong basis for multi-task Omni models. b, Genesis CST model accuracy for independent test set (Accuracy 2) peaked and outperformed XGBoost models when trained with the attention pooling layer and Quant 6.6k dataset (n = 5 checkpoints). c, Genesis ID models showed relatively stable prediction accuracy of ~97% across different output layers and training datasets (n = 5 checkpoints). d, The Genesis Quant and Genesis ID datasets were merged to generate Omni 1399 test set and 8 training sets, which were used to train Genesis Topo, TM and Omni models. e, Genesis Topo models achieved highest accuracy when trained with 81k dataset and attention pooling layer (n = 5 checkpoints). f, Relative to widely used existing methods, Genesis ID and Topo models demonstrate superior performance in accuracy of SPTM-CM identity and topology predictions. Schematics created in BioRender: a, Fang, Z. https://biorender.com/o15o11j (2026); d, Fang, Z. https://biorender.com/frueuj5 (2026).
Extended Data Fig. 10 Similar distribution of Genesis Quant 8k training and Query 177k sequences in embedding space increased confidence that model predictions fall within expected accuracy range observed during validation.
a, Among all SPTM-nonCMs, GO annotated SPTM-CMs were more enriched in the Genesis Quant predicted CST1-5 categories, confirming model’s ability to identify CST signals in ambiguously labeled sequences. b, The averaged logits of the top 3 Genesis Quant models yielded lower MSE2 (Quant 101) than any individual Genesis Quant model. Performance of the best 3 Genesis Quant models with attention pooling, multilayer perceptron (MLP) combined with gaussian error linear unit activation function (GELU), or MLP + rectified linear unit (ReLU) output layers were evaluated using MSE and CST accuracy on Quant 101 and all ~170,000 sequences. The results were also compared with the average logits of best three models. c-e, PCA analysis of training 8k set (c), query 170k set (d) and test 1399 set (e) sequences in embedding space. Sequences were colored based on their CST categories. f, kNN out of distribution (OOD) distances of training 8k, query 170k and test 1399 sets. The threshold for OOD samples is top 5% in test set, 0.286. g, Test 1399 and query 170k sequence identity with nearest neighbor in training 8k set.
Supplementary information
Source data
Rights and permissions
Springer Nature or its licensor (e.g. a society or other partner) holds exclusive rights to this article under a publishing agreement with the author(s) or other rightsholder(s); author self-archiving of the accepted manuscript version of this article is solely governed by the terms of such publishing agreement and applicable law.
About this article
Cite this article
Fang, Z., Saskin, J., Lee, SH. et al. Discovery and design of potent cell surface display elements.
Nat Biotechnol (2026). https://doi.org/10.1038/s41587-026-03144-x
-
Received:
-
Accepted:
-
Published:
-
Version of record:
-
DOI: https://doi.org/10.1038/s41587-026-03144-x