Study finds AI generates empathic responses using formulaic templates that people rate highly
A large-scale evaluation of nine open-source code-generating AI models on 2,707 LeetCode problems across 12 programming languages found the best-performing model reached only 23.64% mean correctness, compared to a 57.2% human acceptance baseline. The study, comprising over 325,000 problem-model-language test runs, also found that compile errors account for nearly two-thirds of failed submissions, meaning many models fail before semantic correctness can even be assessed. The findings suggest that standard single-language, pass-rate-only benchmarks obscure significant performance variation and overstate model capability.
Researchers from arXiv's cs.AI community conducted a large-scale, execution-grounded evaluation of nine openly accessible large language models (LLMs) specialized for coding, testing them on 2,707 free LeetCode problems across 12 programming languages, generating 325,343 individual test jobs. The top-performing model, Yi-Coder-9B-Chat, achieved a mean correctness of 23.64%, falling well short of the 57.2% human acceptance baseline used as a reference point. Model rankings were found to be metric- and slice-dependent: Qwen2.5-Coder-14B-Instruct led on hard problems and distinct-problem coverage, while Gemma-2-27B-IT posted the highest all-language lint pass rate. A failure analysis revealed that compile errors account for 63.25% of non-accepted best submissions, indicating a fundamental gap in syntactic reliability before semantic quality can even be evaluated. The study also found that static code quality metrics diverge from functional correctness, meaning a model can produce clean-looking code that nonetheless fails execution. The authors argue that multilingual, artifact-preserving evaluation frameworks expose tradeoffs and weaknesses that aggregate leaderboards systematically hide.
What's missing
The study uses LeetCode problems as its sole benchmark domain, which may not generalize to real-world software engineering tasks such as debugging, code completion in existing codebases, or open-ended development. The study does not include proprietary models (e.g., GPT-4, Claude, Gemini), limiting comparisons to the open-source ecosystem. It is also unclear how prompt design choices may have influenced model performance across languages.
What different sources said
- arXiv cs.CLCenter
Discourse-Role Labels as Presentation-Time Variables for Context Use in Language Models
Related
Gut Bacteria Enzyme Found to Break Down Heat-Processed Food Compounds, Producing Novel Biogenic Amines
Researchers have discovered that an enzyme in common gut bacteria can degrade N-epsilon-carboxymethyllysine (CML), a compound formed during thermal food processing, producing previously unknown biogenic amines. The enzyme, ornithine decarboxylase SpeC from enterobacteria, acts on CML and related modified lysine derivatives through a low-level 'underground' catalytic activity. This finding suggests a previously unrecognized communication axis between thermally processed dietary compounds and gut microbial physiology, with potential implications for host health.
Full-Length Gene Sequencing Reveals Two Distinct Bacterial Communities in Black-Legged Ticks Expanding Into Canada
Researchers used Oxford Nanopore full-length 16S rRNA gene sequencing to characterize the microbiome of Ixodes scapularis black-legged ticks collected in Nova Scotia, Canada, distinguishing between tick-adapted bacteria and environmentally acquired bacteria. The study comes as I. scapularis — the primary vector of Lyme disease — is rapidly expanding northward into Canada due to climate change. The findings suggest that environmentally derived bacteria in tick microbiomes are not mere contamination, which has implications for how tick microbiome data is collected and interpreted across surveillance studies.
Study Identifies Metabolic Link Between Cell Envelope Stress and Biofilm Formation in Bacteria
Researchers have discovered that the metabolite acetyl-CoA directly inhibits enzymes that degrade the bacterial signaling molecule c-di-GMP, connecting cell envelope biosynthesis stress to biofilm formation in Pseudomonas aeruginosa. The study found that sub-inhibitory concentrations of antibiotics targeting early peptidoglycan biosynthesis — but not other antibiotic classes — elevate c-di-GMP levels by reducing phosphodiesterase activity, with acetyl-CoA competing for the enzyme active site. Because the relevant enzyme domain is broadly conserved across bacterial species, this checkpoint mechanism may be widespread and could have implications for understanding antibiotic-induced biofilm responses.