AI Insight
Researchers have developed nanorepertoire, a standardized computational pipeline for analyzing nanobody repertoires from camelid heavy-chain antibodies using high-throughput sequencing data. The pipeline automates the complete analysis workflow from raw sequencing files through quality control, clonotyping, and CDR3 annotation, producing comprehensive reports on clonal diversity and repertoire characteristics. When tested on llama antibody libraries selected against SARS-CoV-2, the pipeline processed 4.8 million reads in under one hour and successfully identified over 41,000 distinct antibody binding sites.
Why it matters
This tool addresses a critical gap in nanobody research by providing a reproducible, standardized approach to analyzing antibody repertoires, which is essential for therapeutic antibody development and understanding immune responses. The pipeline's portability across different computing platforms and inclusion of carbon footprint tracking makes it accessible to laboratories worldwide while promoting environmentally conscious computational practices.
Understand the Science
⚠️ Preprint – Noch nicht peer-reviewed
Dieser Artikel wurde noch nicht von unabhängigen Experten begutachtet. Die Ergebnisse sind vorläufig und sollten mit Vorsicht interpretiert werden.
Camelid heavy-chain antibodies, and particularly their variable domains known as nanobodies or VHHs, combine full antigen-binding capacity with a compact and highly stable scaffold, which makes them attractive for both fundamental immunology and therapeutic development. High-throughput adaptive immune receptor repertoire sequencing (AIRR-seq) allows nanobody repertoires to be profiled at great depth, but the analyses applied to VHH data are typically assembled ad hoc from standalone scripts, which limits standardisation and reproducibility across laboratories. Here we present nanorepertoire, an end-to-end Nextflow DSL2 pipeline dedicated to camelid VHH repertoires. It takes paired-end AIRR-seq FASTQ files through quality control, adapter trimming, read merging, in-silico translation, CD-HIT clonotyping and deep-learning CDR3 annotation with nanoCDR-X (Bagordo et al., 2026), and returns an interactive HTML report describing clonal architecture, CDR3 length and amino-acid composition, intra-clonal homogeneity and repertoire diversity, together with the computational carbon footprint of the run. Applied to two publicly available SARS-CoV-2 RBD-selected llama libraries sampled before and after phage-display enrichment (4.8 million paired-end reads in total), the pipeline completed in 59 min on a 16-vCPU cloud instance and recovered 41,363 distinct CDR3 paratopes, reproducing the expected contraction of clonal diversity upon selection. nanorepertoire is open source under the MIT licence at https://github.com/lescailab/nanorepertoire, is archived on Zenodo, and runs unchanged on local, HPC and cloud infrastructures.
Source: nanorepertoire: an end-to-end Nextflow pipeline for nanobody repertoire analysis