Neural Networks Maintain Robustness When Learning From Heavily Corrupted Input Data
Researchers have found that multi-layer perceptron neural networks maintain well-above-chance classification accuracy even when more than 90% of their input data is corrupted. The study uses a mean-field-inspired mathematical analysis of infinite-width networks to show that under heavy corruption, networks effectively implement a 'nearest-class-mean' rule, assigning inputs to whichever class has the closest average training example. This finding offers an interpretable explanation for a surprising robustness property of neural networks and has implications for understanding machine learning reliability in noisy real-world settings.
A new preprint posted to arXiv investigates how neural networks cope with heavily corrupted input data — a problem known as attribute noise — where inputs are damaged but class labels remain intact. Through experiments with multi-layer perceptrons (MLPs) on corrupted classification datasets, the authors demonstrate that networks sustain meaningful accuracy even when inputs exceed 90% corruption, a level far beyond human recognition ability. To explain this phenomenon, the researchers analyzed infinite-width networks in the heavy-corruption regime using a mean-field-inspired theoretical framework. Their analysis yields a leading-order decision rule: the network behaves as a prototype classifier, assigning each test point to the class whose training-set centroid (average) it most closely resembles — a rule known as nearest-class-mean. Crucially, this centroid mechanism is shown to be universal across a wide range of MLP architectures, holding regardless of network depth, activation function, or noise distribution type (additive or replacement). The theoretical predictions also closely match observed finite-width network behavior in experiments, lending empirical support to the analytical framework. The work bridges practical robustness questions with fundamental learnability theory, providing a tractable account of why learning can succeed even when individual training examples carry almost no signal.
What's missing
The study focuses on multi-layer perceptrons and does not directly address whether the nearest-class-mean mechanism generalizes to other widely used architectures such as transformers or convolutional neural networks. The study also does not evaluate performance on standard real-world benchmarks under corruption, leaving open questions about practical applicability.
What different sources said
- arXiv cs.LGCenter
Learning from almost nothing: How neural networks survive heavy input corruption
Related
Gut Bacteria Enzyme Found to Break Down Heat-Processed Food Compounds, Producing Novel Biogenic Amines
Researchers have discovered that an enzyme in common gut bacteria can degrade N-epsilon-carboxymethyllysine (CML), a compound formed during thermal food processing, producing previously unknown biogenic amines. The enzyme, ornithine decarboxylase SpeC from enterobacteria, acts on CML and related modified lysine derivatives through a low-level 'underground' catalytic activity. This finding suggests a previously unrecognized communication axis between thermally processed dietary compounds and gut microbial physiology, with potential implications for host health.
Full-Length Gene Sequencing Reveals Two Distinct Bacterial Communities in Black-Legged Ticks Expanding Into Canada
Researchers used Oxford Nanopore full-length 16S rRNA gene sequencing to characterize the microbiome of Ixodes scapularis black-legged ticks collected in Nova Scotia, Canada, distinguishing between tick-adapted bacteria and environmentally acquired bacteria. The study comes as I. scapularis — the primary vector of Lyme disease — is rapidly expanding northward into Canada due to climate change. The findings suggest that environmentally derived bacteria in tick microbiomes are not mere contamination, which has implications for how tick microbiome data is collected and interpreted across surveillance studies.
Study Identifies Metabolic Link Between Cell Envelope Stress and Biofilm Formation in Bacteria
Researchers have discovered that the metabolite acetyl-CoA directly inhibits enzymes that degrade the bacterial signaling molecule c-di-GMP, connecting cell envelope biosynthesis stress to biofilm formation in Pseudomonas aeruginosa. The study found that sub-inhibitory concentrations of antibiotics targeting early peptidoglycan biosynthesis — but not other antibiotic classes — elevate c-di-GMP levels by reducing phosphodiesterase activity, with acetyl-CoA competing for the enzyme active site. Because the relevant enzyme domain is broadly conserved across bacterial species, this checkpoint mechanism may be widespread and could have implications for understanding antibiotic-induced biofilm responses.