Study Reveals Large Language Models Vulnerable to Misleading Medical Information
A new benchmark study shows that large language models (LLMs) drop from 71.1% to 38.0% accuracy on medical questions when misleading context is injected, even when they originally answered those questions correctly. The research introduces MedMisBench, a dataset of nearly 11,000 medical question items designed to measure what the authors call 'epistemic resilience' — a model's ability to maintain correct judgment under adversarial pressure. The findings challenge the widespread assumption that high scores on medical licensing exams indicate safe clinical judgment.
Researchers from multiple institutions have introduced MedMisBench, a benchmark containing 10,932 medical question items and 48,889 misleading context-option pairs, to evaluate whether LLMs can maintain correct medical reasoning when exposed to false or manipulative information. Across 11 model configurations tested, mean accuracy fell from 71.1% on original questions to 38.0% when focused misleading context was introduced, representing a 51.5% attack success rate. The most effective adversarial injections were formal, rule-like fabrications: authority-framed falsehoods achieved a 69.5% attack success rate, while exception-poisoning claims reached 64.1%. A 14-member international clinical panel reviewed a subset of cases and identified serious potential harm in 38.2% of them. The study argues that existing medical AI benchmarks measure factual knowledge but fail to assess whether models can resist plausible-sounding misinformation — a critical gap as patients increasingly turn to LLMs for health advice. The authors frame epistemic resilience as a structural blind spot in current LLM evaluation methodology.
What's missing
The paper does not specify which 11 LLM configurations were tested, making it difficult to assess how findings generalize across model families or generations. It is also unclear whether any mitigation strategies (e.g., prompt engineering, retrieval-augmented generation) were evaluated for their ability to improve epistemic resilience. The clinical panel review methodology — including how cases were selected for review — is not detailed in the abstract.
What different sources said
- arXiv cs.CLCenter
Measuring Epistemic Resilience of LLMs Under Misleading Medical Context
Related
Gut Bacteria Enzyme Found to Break Down Heat-Processed Food Compounds, Producing Novel Biogenic Amines
Researchers have discovered that an enzyme in common gut bacteria can degrade N-epsilon-carboxymethyllysine (CML), a compound formed during thermal food processing, producing previously unknown biogenic amines. The enzyme, ornithine decarboxylase SpeC from enterobacteria, acts on CML and related modified lysine derivatives through a low-level 'underground' catalytic activity. This finding suggests a previously unrecognized communication axis between thermally processed dietary compounds and gut microbial physiology, with potential implications for host health.
Full-Length Gene Sequencing Reveals Two Distinct Bacterial Communities in Black-Legged Ticks Expanding Into Canada
Researchers used Oxford Nanopore full-length 16S rRNA gene sequencing to characterize the microbiome of Ixodes scapularis black-legged ticks collected in Nova Scotia, Canada, distinguishing between tick-adapted bacteria and environmentally acquired bacteria. The study comes as I. scapularis — the primary vector of Lyme disease — is rapidly expanding northward into Canada due to climate change. The findings suggest that environmentally derived bacteria in tick microbiomes are not mere contamination, which has implications for how tick microbiome data is collected and interpreted across surveillance studies.
Study Identifies Metabolic Link Between Cell Envelope Stress and Biofilm Formation in Bacteria
Researchers have discovered that the metabolite acetyl-CoA directly inhibits enzymes that degrade the bacterial signaling molecule c-di-GMP, connecting cell envelope biosynthesis stress to biofilm formation in Pseudomonas aeruginosa. The study found that sub-inhibitory concentrations of antibiotics targeting early peptidoglycan biosynthesis — but not other antibiotic classes — elevate c-di-GMP levels by reducing phosphodiesterase activity, with acetyl-CoA competing for the enzyme active site. Because the relevant enzyme domain is broadly conserved across bacterial species, this checkpoint mechanism may be widespread and could have implications for understanding antibiotic-induced biofilm responses.