Discovery and design of potent cell surface display elements

Data availability

All primary data supporting the findings of this study are available within the Article and via Zenodo at https://doi.org/10.5281/zenodo.19179097 (ref. 62). Source data are provided with this paper.

Code availability

Protein language models were implemented using the code and pretrained parameters available from their official GitHub repositories, following the provided instructions. The ESM2 code and models are available via GitHub at https://github.com/facebookresearch/esm. We used XGBoost and other classical ML models (for example, logistic regression, random forest) from the RAPIDS cuML library for accelerated training on graphics processing unit (GPU). Fine-tuned DL model checkpoints and source code will be publicly released on Hugging Face, Docker and Github (https://github.com/fangzhe3/DeepSCan)57. The GitHub repository also introduces three demo scenarios using the DeepSCan web server or Docker image. The Docker image of the Genesis Quant 8,000 model enables users to process large-scale sequence analyses using GPUs. The supplementary software contains the DeepSCan web server application (version: app_v7.0_deployed_on_AWS), Genesis model training code (version: Genesis_Quant_8k_attention_pooling2_LLRD_v2), Genesis and Omni model inference codes and additional computational analyses codes. The DeepSCan models assume Python (v.3.11.5) + PyTorch (v.2.5.1 with CUDA v.12.4)-based workflow.

References

  1. Overington, J. P., Al-Lazikani, B. & Hopkins, A. L. How many drug targets are there? Nat. Rev. Drug Discov. 5, 993–996 (2006).

    Article 
    CAS 
    PubMed 

    Google Scholar
     

  2. Hegde, R. S. & Keenan, R. J. The mechanisms of integral membrane protein biogenesis. Nat. Rev. Mol. Cell Biol. 23, 107–124 (2022).

    Article 
    CAS 
    PubMed 

    Google Scholar
     

  3. Bugge, K., Lindorff-Larsen, K. & Kragelund, B. B. Understanding single-pass transmembrane receptor signaling from a structural viewpoint—what are we missing?. FEBS J. 283, 4424–4451 (2016).

    Article 
    CAS 
    PubMed 

    Google Scholar
     

  4. Pogozheva, I. D. & Lomize, A. L. Evolution and adaptation of single-pass transmembrane proteins. Biochim. Biophys. Acta Biomembr. 1860, 364–377 (2018).

    Article 
    CAS 
    PubMed 

    Google Scholar
     

  5. Grupp, S. A. et al. Chimeric antigen receptor-modified T cells for acute lymphoid leukemia. N. Engl. J. Med. 368, 1509–1518 (2013).

    Article 
    CAS 
    PubMed 
    PubMed Central 

    Google Scholar
     

  6. Muller, F. et al. CD19 CAR T-cell therapy in autoimmune disease—a case series with follow-up. N. Engl. J. Med. 390, 687–700 (2024).

    Article 
    PubMed 

    Google Scholar
     

  7. Kremer, J. M. et al. Tocilizumab inhibits structural joint damage in rheumatoid arthritis patients with inadequate responses to methotrexate: results from the double-blind treatment phase of a randomized placebo-controlled trial of tocilizumab safety and prevention of structural joint damage at one year. Arthritis Rheum. 63, 609–621 (2011).

    Article 
    CAS 
    PubMed 

    Google Scholar
     

  8. Kishimoto, T. & Kang, S. IL-6 revisited: from rheumatoid arthritis to CAR T cell therapy and COVID-19. Annu. Rev. Immunol. 40, 323–348 (2022).

    Article 
    CAS 
    PubMed 

    Google Scholar
     

  9. Topalian, S. L. et al. Safety, activity, and immune correlates of anti-PD-1 antibody in cancer. N. Engl. J. Med. 366, 2443–2454 (2012).

    Article 
    CAS 
    PubMed 
    PubMed Central 

    Google Scholar
     

  10. Han, Y., Liu, D. & Li, L. PD-1/PD-L1 pathway: current researches in cancer. Am. J. Cancer Res. 10, 727–742 (2020).

    CAS 
    PubMed 
    PubMed Central 

    Google Scholar
     

  11. James, P. A. et al. 2014 evidence-based guideline for the management of high blood pressure in adults: report from the panel members appointed to the Eighth Joint National Committee (JNC 8). JAMA 311, 507–520 (2014).

    Article 
    CAS 
    PubMed 

    Google Scholar
     

  12. Burnett, J. C. Jr. et al. Atrial natriuretic peptide elevation in congestive heart failure in the human. Science 231, 1145–1147 (1986).

    Article 
    PubMed 

    Google Scholar
     

  13. Harayama, T. & Riezman, H. Understanding the diversity of membrane lipid composition. Nat. Rev. Mol. Cell Biol. 19, 281–296 (2018).

    Article 
    CAS 
    PubMed 

    Google Scholar
     

  14. Casares, D., Escriba, P. V. & Rossello, C. A. Membrane lipid composition: effect on membrane and organelle structure, function and compartmentalization and therapeutic avenues. Int. J. Mol. Sci. 20, 2167 (2019).

    Article 
    CAS 
    PubMed 
    PubMed Central 

    Google Scholar
     

  15. Brandizzi, F. et al. The destination for single-pass membrane proteins is influenced markedly by the length of the hydrophobic domain. Plant Cell 14, 1077–1092 (2002).

    Article 
    CAS 
    PubMed 
    PubMed Central 

    Google Scholar
     

  16. Stavropoulos, I. et al. Protein disorder and short conserved motifs in disordered regions are enriched near the cytoplasmic side of single-pass transmembrane proteins. PLoS ONE 7, e44389 (2012).

    Article 
    CAS 
    PubMed 
    PubMed Central 

    Google Scholar
     

  17. Teese, M. G. & Langosch, D. Role of GxxxG motifs in transmembrane domain interactions. Biochemistry 54, 5125–5135 (2015).

    Article 
    CAS 
    PubMed 

    Google Scholar
     

  18. Rapoport, T. A., Goder, V., Heinrich, S. U. & Matlack, K. E. Membrane-protein integration and the role of the translocation channel. Trends Cell Biol. 14, 568–575 (2004).

    Article 
    CAS 
    PubMed 

    Google Scholar
     

  19. Honigmann, A. & Pralle, A. Compartmentalization of the cell membrane. J. Mol. Biol. 428, 4739–4748 (2016).

    Article 
    CAS 
    PubMed 

    Google Scholar
     

  20. Lorent, J. H. et al. Structural determinants and functional consequences of protein affinity for membrane rafts. Nat. Commun. 8, 1219 (2017).

    Article 
    PubMed 
    PubMed Central 

    Google Scholar
     

  21. Diaz-Rohrer, B. B., Levental, K. R., Simons, K. & Levental, I. Membrane raft association is a determinant of plasma membrane localization. Proc. Natl Acad. Sci. USA 111, 8500–8505 (2014).

    Article 
    CAS 
    PubMed 
    PubMed Central 

    Google Scholar
     

  22. Levental, I., Lingwood, D., Grzybek, M., Coskun, U. & Simons, K. Palmitoylation regulates raft affinity for the majority of integral raft proteins. Proc. Natl Acad. Sci. USA 107, 22050–22054 (2010).

    Article 
    CAS 
    PubMed 
    PubMed Central 

    Google Scholar
     

  23. Fang, Z. et al. A modular vaccine platform for optimized lipid nanoparticle mRNA immunogenicity. Nat. Biomed. Eng. 10, 501–516 (2025).

    Article 
    PubMed 

    Google Scholar
     

  24. Choe, J. H. et al. SynNotch-CAR T cells overcome challenges of specificity, heterogeneity, and persistence in treating glioblastoma. Sci. Transl. Med. 13, eabe7378 (2021).

    Article 
    CAS 
    PubMed 
    PubMed Central 

    Google Scholar
     

  25. Chung, J. B., Brudno, J. N., Borie, D. & Kochenderfer, J. N. Chimeric antigen receptor T cell therapy for autoimmune disease. Nat. Rev. Immunol. 24, 830–845 (2024).

    Article 
    CAS 
    PubMed 
    PubMed Central 

    Google Scholar
     

  26. Ellebrecht, C. T. et al. Reengineering chimeric antigen receptor T cells for targeted therapy of autoimmune disease. Science 353, 179–184 (2016).

    Article 
    CAS 
    PubMed 
    PubMed Central 

    Google Scholar
     

  27. Mitra, A. et al. From bench to bedside: the history and progress of CAR T cell therapy. Front. Immunol. 14, 1188049 (2023).

    Article 
    CAS 
    PubMed 
    PubMed Central 

    Google Scholar
     

  28. Arieta, C. M. et al. The T-cell-directed vaccine BNT162b4 encoding conserved non-spike antigens protects animals from severe SARS-CoV-2 infection. Cell 186, 2392–2409.e21 (2023).

    Article 
    CAS 
    PubMed 
    PubMed Central 

    Google Scholar
     

  29. Jumper, J. et al. Highly accurate protein structure prediction with AlphaFold. Nature 596, 583–589 (2021).

    Article 
    CAS 
    PubMed 
    PubMed Central 

    Google Scholar
     

  30. Watson, J. L. et al. De novo design of protein structure and function with RFdiffusion. Nature 620, 1089–1100 (2023).

    Article 
    CAS 
    PubMed 
    PubMed Central 

    Google Scholar
     

  31. Baek, M. et al. Accurate prediction of protein structures and interactions using a three-track neural network. Science 373, 871–876 (2021).

    Article 
    CAS 
    PubMed 
    PubMed Central 

    Google Scholar
     

  32. Wang, J. et al. Scaffolding protein functional sites using deep learning. Science 377, 387–394 (2022).

    Article 
    CAS 
    PubMed 
    PubMed Central 

    Google Scholar
     

  33. Abramson, J. et al. Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature 630, 493–500 (2024).

    Article 
    CAS 
    PubMed 
    PubMed Central 

    Google Scholar
     

  34. Sunny, S., Prakash, P. B., Gopakumar, G. & Jayaraj, P. B. DeepBindPPI: protein–protein binding site prediction using attention based graph convolutional network. Protein J. 42, 276–287 (2023).

    Article 
    CAS 
    PubMed 
    PubMed Central 

    Google Scholar
     

  35. Yang, X., Yang, S., Li, Q., Wuchty, S. & Zhang, Z. Prediction of human–virus protein–protein interactions through a sequence embedding-based machine learning method. Comput. Struct. Biotechnol. J. 18, 153–161 (2020).

    Article 
    CAS 
    PubMed 

    Google Scholar
     

  36. Almagro Armenteros, J. J., Sonderby, C. K., Sonderby, S. K., Nielsen, H. & Winther, O. DeepLoc: prediction of protein subcellular localization using deep learning. Bioinformatics 33, 4049 (2017).

    Article 
    PubMed 

    Google Scholar
     

  37. Thumuluri, V., Almagro Armenteros, J. J., Johansen, A. R., Nielsen, H. & Winther, O. DeepLoc 2.0: multi-label subcellular localization prediction using protein language models. Nucleic Acids Res. 50, W228–W234 (2022).

    Article 
    CAS 
    PubMed 
    PubMed Central 

    Google Scholar
     

  38. Jiang, Y. et al. MULocDeep: a deep-learning framework for protein subcellular and suborganellar localization prediction with residue-level interpretation. Comput. Struct. Biotechnol. J. 19, 4825–4839 (2021).

    Article 
    CAS 
    PubMed 
    PubMed Central 

    Google Scholar
     

  39. Parnamaa, T. & Parts, L. Accurate classification of protein subcellular localization from high-throughput microscopy images using deep learning. G3 7, 1385–1392 (2017).

    Article 
    CAS 
    PubMed 
    PubMed Central 

    Google Scholar
     

  40. Xiao, H., Zou, Y., Wang, J. & Wan, S. A review for artificial intelligence based protein subcellular localization. Biomolecules 14, 409 (2024).

    Article 
    CAS 
    PubMed 
    PubMed Central 

    Google Scholar
     

  41. Lu, Z. et al. Predicting subcellular localization of proteins using machine-learned classifiers. Bioinformatics 20, 547–556 (2004).

    Article 
    CAS 
    PubMed 

    Google Scholar
     

  42. Kobayashi, H., Cheveralls, K. C., Leonetti, M. D. & Royer, L. A. Self-supervised deep learning encodes high-resolution features of protein subcellular localization. Nat. Methods 19, 995–1003 (2022).

    Article 
    CAS 
    PubMed 
    PubMed Central 

    Google Scholar
     

  43. Moreno, J., Nielsen, H., Winther, O. & Teufel, F. Predicting the subcellular location of prokaryotic proteins with DeepLocPro. Bioinformatics 40, btae677 (2024).

    Article 
    CAS 
    PubMed 
    PubMed Central 

    Google Scholar
     

  44. Lin, Z. et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science 379, 1123–1130 (2023).

    Article 
    CAS 
    PubMed 

    Google Scholar
     

  45. Hayes, T. et al. Simulating 500 million years of evolution with a language model. Science 387, 850–858 (2025).

    Article 
    CAS 
    PubMed 

    Google Scholar
     

  46. Ingraham, J. B. et al. Illuminating protein space with a programmable generative model. Nature 623, 1070–1078 (2023).

    Article 
    CAS 
    PubMed 
    PubMed Central 

    Google Scholar
     

  47. Ferruz, N., Schmidt, S. & Hocker, B. ProtGPT2 is a deep unsupervised language model for protein design. Nat. Commun. 13, 4348 (2022).

    Article 
    CAS 
    PubMed 
    PubMed Central 

    Google Scholar
     

  48. Wu, K. E. et al. Protein structure generation via folding diffusion. Nat. Commun. 15, 1059 (2024).

    Article 
    CAS 
    PubMed 
    PubMed Central 

    Google Scholar
     

  49. Shin, J. E. et al. Protein design and variant prediction using autoregressive generative models. Nat. Commun. 12, 2403 (2021).

    Article 
    CAS 
    PubMed 
    PubMed Central 

    Google Scholar
     

  50. Madani, A. et al. Large language models generate functional protein sequences across diverse families. Nat. Biotechnol. 41, 1099–1106 (2023).

    Article 
    CAS 
    PubMed 
    PubMed Central 

    Google Scholar
     

  51. Anishchenko, I. et al. De novo protein design by deep network hallucination. Nature 600, 547–552 (2021).

    Article 
    CAS 
    PubMed 
    PubMed Central 

    Google Scholar
     

  52. UniProt, C. UniProt: the Universal Protein Knowledgebase in 2023. Nucleic Acids Res. 51, D523–D531 (2023).

    Article 

    Google Scholar
     

  53. He, Y. et al. Antibodies to the A27 protein of vaccinia virus neutralize and protect against infection but represent a minor component of Dryvax vaccine-induced immunity. J. Infect. Dis. 196, 1026–1032 (2007).

    Article 
    CAS 
    PubMed 

    Google Scholar
     

  54. Kaever, T. et al. Linear epitopes in vaccinia virus A27 are targets of protective antibodies induced by vaccination against smallpox. J. Virol. 90, 4334–4345 (2016).

    Article 
    CAS 
    PubMed 
    PubMed Central 

    Google Scholar
     

  55. Li, M. et al. Three neutralizing mAbs induced by MPXV A29L protein recognizing different epitopes act synergistically against orthopoxvirus. Emerg. Microbes Infect. 12, 2223669 (2023).

    Article 
    PubMed 
    PubMed Central 

    Google Scholar
     

  56. Fang, Z. et al. Polyvalent mRNA vaccination elicited potent immune response to monkeypox virus surface antigens. Cell Res. 33, 407–410 (2023).

    Article 
    CAS 
    PubMed 
    PubMed Central 

    Google Scholar
     

  57. Fang, Z. DeepSCan – an AI platform for discovery and design of potent cell surface display elements. Source code. GitHub https://github.com/fangzhe3/DeepSCan (2026).

  58. Sun, Y. et al. Out-of-distribution detection with deep nearest neighbors. In Proc. 39th International Conference on Machine Learning (eds Chaudhuri, K. et al.) Vol. 162, 20827–20840 (PMLR, 2022).

  59. Ye, L. et al. A genome-scale gain-of-function CRISPR screen in CD8 T cells identifies proline metabolism as a means to enhance CAR-T therapy. Cell Metab. 34, 595–614.e14 (2022).

    Article 
    CAS 
    PubMed 
    PubMed Central 

    Google Scholar
     

  60. Peng, L. et al. In vivo AAV-SB-CRISPR screens of tumor-infiltrating primary NK cells identify genetic checkpoints of CAR-NK therapy. Nat. Biotechnol. 43, 752–761 (2025).

    Article 
    CAS 
    PubMed 

    Google Scholar
     

  61. Boutet, E. et al. UniProtKB/Swiss-Prot, the manually annotated section of the UniProt KnowledgeBase: how to use the entry view. Methods Mol. Biol. 1374, 23–54 (2016).

    Article 
    CAS 
    PubMed 

    Google Scholar
     

  62. Yale University et al. DeepSCan – an AI platform for discovery and design of potent cell surface display elements. Zenodo https://doi.org/10.5281/zenodo.19179097 (2026).

Download references

Acknowledgements

We thank all members of the Chen Laboratory and various colleagues in the Department of Genetics, Systems Biology Institute, Yale Cancer Center (YCC), Yale Stem Cell Center, High Performance Computing, West Campus Analytical Chemistry Core, Biomedical Informatics and Data Sciences at Yale for assistance and/or discussion. S.C. is supported by a Cancer Research Institute Lloyd J. Old STAR Award and CLIP Award (CRI4964), the NIH/NCI (R33CA281702), the DoD (HT9425-23-1-0860), the Alliance for Cancer Gene Therapy (ACGT), the Pershing Square Sohn Cancer Research Alliance, the Sontag Foundation, D. Lu and B. Sperry. Z.F. is supported by the Canadian Institutes of Health Research (funding reference number 194053). J.S. is supported by a Yale MSTP training grant from the NIH (T32GM136651). C.Z. is supported by a Yale PhD training grant from the NIH (T32HD007149). S.-H.L. is supported by a Korean fellowship and a Leslie Warner fellowship.

Author information

Author notes

  1. These authors contributed equally: Zhenhao Fang, Joshua Saskin.

Authors and Affiliations

  1. Department of Genetics, Yale University School of Medicine, New Haven, CT, USA

    Zhenhao Fang, Joshua Saskin, Seok-Hoon Lee, Charles Zou, Shan Xin, Xiaoyu Huang, Chuanpeng Dong, Ardavan Abiri, Yanzhi Feng & Sidi Chen

  2. Systems Biology Institute, Yale University, West Haven, CT, USA

    Zhenhao Fang, Joshua Saskin, Seok-Hoon Lee, Charles Zou, Shan Xin, Xiaoyu Huang, Chuanpeng Dong, Ardavan Abiri, Yanzhi Feng & Sidi Chen

  3. Center for Cancer Systems Biology, Yale University, West Haven, CT, USA

    Zhenhao Fang, Joshua Saskin, Seok-Hoon Lee, Charles Zou, Shan Xin, Xiaoyu Huang, Chuanpeng Dong, Ardavan Abiri, Yanzhi Feng & Sidi Chen

  4. Yale MD–PhD Program, Yale University, New Haven, CT, USA

    Joshua Saskin & Sidi Chen

  5. Combined Program in the Biological and Biomedical Sciences, Yale University, New Haven, CT, USA

    Joshua Saskin, Charles Zou, Ardavan Abiri & Sidi Chen

  6. Molecular Cell Biology, Genetics, and Development Program, Yale University, New Haven, CT, USA

    Joshua Saskin, Charles Zou & Sidi Chen

  7. Department of Neurosurgery, Baylor College of Medicine, Temple, TX, USA

    Xingxin Pan, Nidhi Sahni & S. Stephen Yi

  8. Center for Biomedical Data Science, Yale University School of Medicine, New Haven, CT, USA

    Ardavan Abiri & Sidi Chen

  9. Computational Biology and Biomedical Informatics Program, Yale University, New Haven, CT, USA

    Ardavan Abiri & Sidi Chen

  10. Immunobiology Program, Yale University, West Haven, CT, USA

    Yanzhi Feng & Sidi Chen

  11. Quantitative and Computational Biosciences Program, Baylor College of Medicine, Houston, TX, USA

    Nidhi Sahni

  12. Dan L Duncan Comprehensive Cancer Center, and Dan L Duncan Institute for Clinical and Translational Research, Baylor College of Medicine, Houston, TX, USA

    S. Stephen Yi

  13. Department of Neurosurgery, Yale University School of Medicine, New Haven, CT, USA

    Lei Peng & Sidi Chen

  14. Yale College, New Haven, CT, USA

    Sidi Chen

  15. Yale Cancer Center, Yale University School of Medicine, New Haven, CT, USA

    Sidi Chen

  16. Stem Cell Center, Yale University School of Medicine, New Haven, CT, USA

    Sidi Chen

Authors

  1. Zhenhao Fang
  2. Joshua Saskin
  3. Seok-Hoon Lee
  4. Charles Zou
  5. Shan Xin
  6. Xiaoyu Huang
  7. Xingxin Pan
  8. Chuanpeng Dong
  9. Ardavan Abiri
  10. Yanzhi Feng
  11. Nidhi Sahni
  12. S. Stephen Yi
  13. Lei Peng
  14. Sidi Chen

Contributions

Z.F. and S.C. conceived the study and designed the experiments. Z.F. and J.S. performed most of the experiments with assistance from various co-authors. Z.F. designed the CSD libraries and performed the CSD screens. J.S. independently validated CSD screen results in multiple immunological and functional assays. Z.F. developed the DeepSCan model, web server and Docker image. ZF performed most of the computational analyses. S.-H.L., C.Z., S.X., X.H., X.P., C.D., A.A., Y.F. and L.P. assisted with various experiments and analyses. N.S. and S.S.Y. supervised trainees and provided support. Z.F., J.S. and S.C. prepared the paper with input from all authors. S.C. secured funding and provided overall supervision of the project.

Corresponding author

Correspondence to
Sidi Chen.

Ethics declarations

Competing interests

An invention disclosure related to this work has been made to Yale University, which filed a patent application. S.C. is a (co)founder of EvolveImmune Tx, Cellinfinity Bio, MagicTime Med and Chen Consulting. The other authors declare no competing interests.

Peer review

Peer review information

Nature Biotechnology thanks Han Liang and the other, anonymous, reviewer(s) for their contribution to the peer review of this work. Peer reviewer reports are available.

Additional information

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Extended data

Extended Data Fig. 1 Workflow for identifying single-pass transmembrane protein on cell membrane (SPTM-CM).

This flow chart illustrates step-by-step process used to identify SPTM-CMs among >570,000 proteins in Swiss-Prot database (retrieved in Dec. 2023). To define the membrane orientation and domain arrangement of SPTM-CMs, we extracted information regarding topology, transmembrane domain boundaries and subcellular location in the database. The identified SPTM-CMs are classified into two categories based on their topology, namely type I (ECD at N-term) and type II (ECD at C-term) SPTM-CMs. SPTM-CMs are clustered based on their sequence identity and proteins in the human genome are highlighted. In addition, SPTM-CMs’ distribution in taxonomic superkingdoms and their associated biological taxons are summarized and displayed in the chart.

Extended Data Fig. 2 Predicting and discovering potential SPTM-CMs with ambiguous subcellular localization labels in UniProt database.

a, DeepScan ID (per-protein inference) and Topo (per-residue inference) models identified potential SPTM-CMs among proteins with ambiguous location labels. The ~220 predicted double label 1 (DL1) SPTM-CMs have >50% sequence identity with existing SPTM-CMs. The known SPTM-CMs are used as positive controls to evaluate DeepSCan model’s accuracy. b, Sequence-identity clusters (50% identity cutoff) are found in the identity matrix of predicted DL1 and known SPTM-CMs. c, t-SNE plot reduces data dimension of sequence identity matrix and visualizes the sequence distances of 222 predicted (DL1), 211 positive SPTM-CMs and 200 negative controls (DL0) on the 2-D plane. d, GO annotation analysis of ~700 DL1 proteins in UniProt database shows that >400 proteins have plasma membrane localization labels (indicated by red arrows), confirming their SPTM-CM identity. Schematic in a created in BioRender; Fang, Z. https://biorender.com/etn7ty8 (2026).

Extended Data Fig. 3 ~3500 type I CSD elements from SPTM-CMs are clustered based on their pair-wise identity matrix and visualized on t-SNE plot.

a-b, type I SPTM-CM clusters are color coded based on their identity clusters (a) or taxonomic superkingdoms (b). c, Histogram showing sequence length distribution of human type I CSD modules. d, Cluster-representative human CSD elements less than 201aa are highlighted in the t-SNE plot.

Source data

Extended Data Fig. 4 Representative flow plots showing the gating strategies to define PE or PE-high positive 293T cells that overexpressed flag-tagged A29L fused to N-term [Spike SP] and C-term [type I CSDs].

Cells were surface stained with PE anti-flag antibody. The SARS-CoV-2 XBB Spike full length (orange) and [Spike SP]-A29L-[HKU9 Spike CSD] (Green) were served as two reference points to identify CSD modules mediating higher A29 antigen surface expression. The A29L with endomembrane protein type I CSDs (blue), including NDUFA4 and ACP2, are served as negative controls. The CSD modules that outperformed reference points were highlighted depending on which quadrants they fall into in Extended Data Fig. 5d.

Extended Data Fig. 5 Amino acid usage frequency analysis revealed unique features of SPTM-CM, while A29-[type I CSD] screen identified potent CSD modules.

a-c, Amino acid usage frequency analysis of TMD and adjacent regions in type I SPTM-CM (a), type II SPTM-CM (b) and SPTM-nonCM (c) revealed a more frequent use of basic residues at the intracellular domain (ICD) and TMD interface of SPTM-CM. Proteins that have 21aa transmembrane domains in UniProt are used in this analysis and their N-term of TMD is defined as position 1. d, Flow surface staining and live/dead staining of 293 T cells overexpressing ~110 A29-[type I CSD] identified type I CSD modules that potently translocated A29 antigen to cell surface (left) and maintained cell viability comparable to negative control cells (right).

Source data

Extended Data Fig. 6 Amino acid usage frequency analysis of major protein domains uncovered conserved sequence features associated with potent CSD (CST3).

Potent CSD sequences from experimental screens (left) and DeepSCan model predictions (right) showed lower basic residue frequency in extracellular hinge (a-d), more diversified hydrophobic residues in TMD (e-f) and higher basic residue frequency in ICD (g-h). The DeepSCan-predicted CST3 SPTM-CM (right) recapitulated the sequence features found in the screened CSD modules (left). Data are presented as mean ± SEM. Pair-wise Tukey’s comparison test was used to determine statistical significance. Sample number n was shown in each group.

Source data

Extended Data Fig. 7 Representative flow plots showing the gating strategies to define PE or PE-high positive 293T cells that overexpressed flag-tagged A29L fused to N-term [Spike SP] and C-term type I natural (top) or generative (bottom) CSDs.

Cells were surface stained with PE anti-flag antibody. The SARS-CoV-2 XBB Spike full length (Orange), A29-HLA (Green) and A29-Spike (Green) were served as reference points to identify CSD modules mediating higher A29 antigen surface expression. The A29L alone and A29L with endomembrane protein type I NDUFA4 CSD (blue) were served as a negative control. The most potent CSD modules were highlighted in purple, as outlined in Fig. 2a.

Extended Data Fig. 8 Genesis Quant regressive model benchmarking revealed a high-attention region at the TMD and ICD interface, likely mediating surface translocation of potent CSD modules, including CXCL16-69.

a Eight ESM2-based deep learning models with different output layers and two XGBoost-based machine learning models were trained by augmented 4.5k dataset and benchmarked using the independent test set (n = 5 checkpoints). Compared to XGBoost-based approaches, test training showed superior performance of deep learning models trained using CLS tokens or layer-dependent learning rate decay (LLRD). b, Genesis Quant 8k attention model uncovered a high-attention-weight region at the TMD and ICD interface of potent gCSDs including CXCL16-69. c, Multiple sequence alignment of CXCL16 WT and gCSDs highlights the conserved features and unique variations associated with gCSDs. Schematic in a created in BioRender; Fang, Z. https://biorender.com/xv22ugk (2026).

Source data

Extended Data Fig. 9 Genesis Dawn single-task model benchmarking revealed optimal model architectures that serve as foundation for multi-task Omni models and achieved higher accuracy than widely used deep learning methods in SPTM-CM feature predictions.

a, Compared to first-generation DeepSCan models, Genesis single-task models were benchmarked and optimized using different output layers and training datasets, establishing a strong basis for multi-task Omni models. b, Genesis CST model accuracy for independent test set (Accuracy 2) peaked and outperformed XGBoost models when trained with the attention pooling layer and Quant 6.6k dataset (n = 5 checkpoints). c, Genesis ID models showed relatively stable prediction accuracy of ~97% across different output layers and training datasets (n = 5 checkpoints). d, The Genesis Quant and Genesis ID datasets were merged to generate Omni 1399 test set and 8 training sets, which were used to train Genesis Topo, TM and Omni models. e, Genesis Topo models achieved highest accuracy when trained with 81k dataset and attention pooling layer (n = 5 checkpoints). f, Relative to widely used existing methods, Genesis ID and Topo models demonstrate superior performance in accuracy of SPTM-CM identity and topology predictions. Schematics created in BioRender: a, Fang, Z. https://biorender.com/o15o11j (2026); d, Fang, Z. https://biorender.com/frueuj5 (2026).

Source data

Extended Data Fig. 10 Similar distribution of Genesis Quant 8k training and Query 177k sequences in embedding space increased confidence that model predictions fall within expected accuracy range observed during validation.

a, Among all SPTM-nonCMs, GO annotated SPTM-CMs were more enriched in the Genesis Quant predicted CST1-5 categories, confirming model’s ability to identify CST signals in ambiguously labeled sequences. b, The averaged logits of the top 3 Genesis Quant models yielded lower MSE2 (Quant 101) than any individual Genesis Quant model. Performance of the best 3 Genesis Quant models with attention pooling, multilayer perceptron (MLP) combined with gaussian error linear unit activation function (GELU), or MLP + rectified linear unit (ReLU) output layers were evaluated using MSE and CST accuracy on Quant 101 and all ~170,000 sequences. The results were also compared with the average logits of best three models. c-e, PCA analysis of training 8k set (c), query 170k set (d) and test 1399 set (e) sequences in embedding space. Sequences were colored based on their CST categories. f, kNN out of distribution (OOD) distances of training 8k, query 170k and test 1399 sets. The threshold for OOD samples is top 5% in test set, 0.286. g, Test 1399 and query 170k sequence identity with nearest neighbor in training 8k set.

Source data

Supplementary information

Source data

About this article

Cite this article

Fang, Z., Saskin, J., Lee, SH. et al. Discovery and design of potent cell surface display elements.
Nat Biotechnol (2026). https://doi.org/10.1038/s41587-026-03144-x

Download citation

  • Received:

  • Accepted:

  • Published:

  • Version of record:

  • DOI: https://doi.org/10.1038/s41587-026-03144-x

Read More

Related posts

How Nigerians can buy Dangote Refinery as company submits applications

China to tighten entry and exit regulations for travellers from Sept 15 , China News

The Hydrogen Stream: Commissioning begins for 2.2 GW Neom H2 project in Saudi Arabia