Biology

Scientists develop better way to identify genetic targets for disease treatment

How the science connects

Machine learningCausal inferenceGene expression re…

AI Insight

This study reveals that current methods for identifying which genes drive cellular state changes are systematically overestimating their performance due to data leakage problems. The researchers developed a new evaluation framework called Cause-excluded Effect Evaluation (CEE) that eliminates these biases by separating training and test data properly and removing the perturbed gene's own expression signal. Testing across 16 single-cell datasets showed that simple structure-based methods actually outperform complex deep learning models when evaluated correctly, indicating that perturbation responses follow low-rank, additive patterns rather than complex nonlinear relationships.


This framework addresses a critical problem in drug target discovery and cellular engineering, where identifying the right genes to manipulate is essential for developing new therapies. By exposing flaws in current evaluation practices and showing that simpler methods work better, this research could redirect computational efforts toward more reliable and interpretable approaches for predicting genetic intervention targets.


Understand the Science

Machine learning 327 articles Explore Concept → Causal inference Concept coming soon Gene expression regulation Concept coming soon

⚠️ Preprint – Noch nicht peer-reviewed

Dieser Artikel wurde noch nicht von unabhängigen Experten begutachtet. Die Ergebnisse sind vorläufig und sollten mit Vorsicht interpretiert werden.

Identifying which gene drives a desired cell-state transition is a foundational inverse problem in target discovery and cell-state engineering, yet far less developed than forward prediction. We show that prevailing benchmarks inflate performance through two artefacts: training leakage, where test-perturbation cells are seen during fitting, and evaluation leakage, where the perturbed gene’s own expression change exposes the answer. Here, we introduce Cause-excluded Effect Evaluation (CEE), a framework that separates train and test perturbation targets, removes target-gene self-effects from expression profiles and scores candidates in downstream response space. With CEE, we find that simple structure-based methods consistently outperform complex deep learning models across 16 single-cell perturbation datasets, revealing that the recoverable signal is governed by low-rank and additive structure in perturbation responses rather than by model complexity. Our work disentangles evaluation artefacts from genuine predictive capacity and offers a design principle for developing perturbation target identification methods.

Source: Cause-excluded Effect Evaluation: a framework for the faithful assessment of genetic perturbation target identification