Biology

AI Model Decodes RNA’s Complex Language for Better Disease Treatment

How the science connects

Artificial intelli…Computational biol…RNA

AI Insight

RIBOSPAN is a 1.61-billion-parameter RNA foundation model designed to process complete RNA transcripts up to 10,240 nucleotides in length, addressing limitations of existing models that cannot handle full-length messenger RNAs. The model uses bidirectional self-attention and single-nucleotide tokenization to achieve high-resolution modeling of long RNA sequences, demonstrating strong performance in nucleotide reconstruction, context-aware representation learning, and classification across diverse RNA types. Additionally, the researchers developed a diffusion-based framework built on RIBOSPAN for generating and redesigning full-length mRNA sequences, including protein-preserving optimization through synonymous codon substitution.


This advancement could significantly improve computational tools for RNA-based therapeutics, vaccine design, and synthetic biology by enabling accurate modeling and generation of complete RNA transcripts. The ability to optimize coding sequences while preserving protein function has direct applications in mRNA drug development and gene therapy optimization.


⚠️ Preprint – Noch nicht peer-reviewed

Dieser Artikel wurde noch nicht von unabhängigen Experten begutachtet. Die Ergebnisse sind vorläufig und sollten mit Vorsicht interpretiert werden.

Abstract: Full-length RNAs, particularly messenger RNAs, often exceed the context lengths used to pretrain existing RNA foundation models, limiting complete-transcript modeling at single-nucleotide resolution. We present RIBOSPAN, a 1.61-billion-parameter bidirectional RNA foundation model natively pretrained with context lengths up to 10,240 nt. RIBOSPAN combines dense bidirectional self-attention, single-nucleotide tokenization, and attention-isolated sequence packing to enable high-resolution modeling of complete long RNAs. We evaluate the model through nucleotide reconstruction, a controlled long-context representation benchmark, and frozen RNA-type representation analysis. Native 10K pretraining preserves strong reconstruction at 10,240 tokens, while continued pretraining with 40% masking improves recovery under heavy corruption while preserving representation quality. The long-context benchmark further shows that native 10K models maintain strong contextual responsiveness and context-specific representation separation while keeping perturbation-induced representation changes highly localized. Inference-time YaRN scaling recovers much of the contextual organization lost by direct extrapolation of short-context models, but induces substantially greater distal representation diffusion. Frozen-representation evaluations further demonstrate state-of-the-art RNA representation quality, with RIBOSPAN achieving the strongest overall performance across diverse RNA types and retaining a clear advantage on long RNAs. Building on the same backbone, we develop a multidimensionally conditioned discrete-diffusion framework for full-length mRNA generation and redesign, including synonymous-codon diffusion for protein-preserving CDS optimization. Together, RIBOSPAN establishes a powerful long-context foundation for transferable RNA representation learning and full-transcript mRNA design.

Source: RIBOSPAN: A Long-Context RNA Foundation Model for Versatile RNA Modeling