Study Questions Whether Causal Gene Relationships Can Be Reliably Recovered from Bulk Gene Expression Data
Two preprint studies independently identify critical limitations in standard computational approaches to analyzing gene expression data. The first shows that recovering causal relationships between genes from bulk RNA data is theoretically valid only under strict linearity conditions that empirical data rarely satisfy, while the second warns that widely used batch correction methods can distort gene-to-gene correlation structures. Together, the findings cast doubt on a broad class of commonly used bioinformatics workflows that underpin disease research and drug discovery.
A study posted to arXiv formalizes the conditions under which causal gene regulatory relationships can be recovered from bulk gene expression data — which pools RNA across many cells — and finds that recoverability holds only when both the aggregation process is linear and the underlying biological system follows affine structural equations. Analyses of four bulk and four single-cell datasets showed that pairwise regulatory functions between genes consistently deviate from linearity in practice, undermining the theoretical basis for causal inference from bulk data. Separately, a bioRxiv preprint highlights that embedding-based batch correction methods, among the most popular tools for harmonizing multi-dataset gene expression studies, can systematically distort feature-to-feature relationships such as gene co-expression patterns. The bioRxiv authors introduce a new metric to quantify this distortion, suggesting the problem has gone largely unmeasured in prior work. Taken together, the two studies identify compounding risks: bulk aggregation may obscure true causal structure, and a standard preprocessing step intended to improve data quality may further corrupt the relational signals that downstream analyses depend on.
What's missing
Both studies are preprints and have not yet undergone formal peer review. Neither study proposes fully validated alternative methods that overcome the identified limitations, leaving open the practical question of how researchers should proceed when bulk data or batch correction is unavoidable. The extent to which these distortions have materially affected published biological conclusions in the literature remains unquantified.
What different sources said
- bioRxivCenter
When batch correction corrupts gene expression: uncovering distortions in correlation structures
- arXiv cs.LGCenter
On the Recoverability of Causal Relations from Bulk Gene Expression Data
Related
Gut Bacteria Enzyme Found to Break Down Heat-Processed Food Compounds, Producing Novel Biogenic Amines
Researchers have discovered that an enzyme in common gut bacteria can degrade N-epsilon-carboxymethyllysine (CML), a compound formed during thermal food processing, producing previously unknown biogenic amines. The enzyme, ornithine decarboxylase SpeC from enterobacteria, acts on CML and related modified lysine derivatives through a low-level 'underground' catalytic activity. This finding suggests a previously unrecognized communication axis between thermally processed dietary compounds and gut microbial physiology, with potential implications for host health.
Full-Length Gene Sequencing Reveals Two Distinct Bacterial Communities in Black-Legged Ticks Expanding Into Canada
Researchers used Oxford Nanopore full-length 16S rRNA gene sequencing to characterize the microbiome of Ixodes scapularis black-legged ticks collected in Nova Scotia, Canada, distinguishing between tick-adapted bacteria and environmentally acquired bacteria. The study comes as I. scapularis — the primary vector of Lyme disease — is rapidly expanding northward into Canada due to climate change. The findings suggest that environmentally derived bacteria in tick microbiomes are not mere contamination, which has implications for how tick microbiome data is collected and interpreted across surveillance studies.
Study Identifies Metabolic Link Between Cell Envelope Stress and Biofilm Formation in Bacteria
Researchers have discovered that the metabolite acetyl-CoA directly inhibits enzymes that degrade the bacterial signaling molecule c-di-GMP, connecting cell envelope biosynthesis stress to biofilm formation in Pseudomonas aeruginosa. The study found that sub-inhibitory concentrations of antibiotics targeting early peptidoglycan biosynthesis — but not other antibiotic classes — elevate c-di-GMP levels by reducing phosphodiesterase activity, with acetyl-CoA competing for the enzyme active site. Because the relevant enzyme domain is broadly conserved across bacterial species, this checkpoint mechanism may be widespread and could have implications for understanding antibiotic-induced biofilm responses.