New Benchmark Tests LLMs' Ability to Identify Code Errors in Competitive Programming
Researchers have introduced UOJ-Bench, a benchmark evaluating large language models on code generation, hacking, and repair tasks drawn from real competitive programming submissions. The benchmark reveals that even top models identify errors in fewer than 50% of incorrect submissions under standard one-shot evaluation, though test-time scaling pushes success rates above 90% at significant computational cost. The findings highlight both the promise and current limitations of LLMs as educational tools in programming contexts.
A new benchmark called UOJ-Bench has been proposed to assess large language models (LLMs) not just on solving competitive programming problems, but on their ability to detect and fix errors in human-written code — tasks central to programming education. The benchmark draws on real submissions from the Universal Online Judge (UOJ) platform and evaluates models across three tasks: code generation, code hacking (identifying bugs in others' code), and code repair. Under one-shot evaluation, even the strongest frontier models fail to catch errors in more than half of submissions already flagged as incorrect by the judging system. Test-time scaling — allocating more compute at inference time — improves detection rates to above 90%, but the associated computational costs make large-scale deployment impractical. Notably, the best-performing models under test-time scaling can surface errors in over 5% of full-score submissions across roughly 30 problems, suggesting LLMs may offer complementary signals beyond what standard automated judges provide. The work points to a gap between LLMs' well-documented problem-solving capabilities and their readiness to serve as reliable educational assistants in competitive programming.
What's missing
The study does not address how UOJ-Bench performance might generalize to other online judge platforms or programming languages beyond those represented in UOJ submissions. The problem set of 'roughly 30 problems' used for full-score hacking analysis is relatively small, limiting the statistical robustness of that finding.
What different sources said
- arXiv cs.AICenter
Beyond Problem Solving: UOJ-Bench for Evaluating Code Generation, Hacking, and Repair in Competitive Programming
Related
Gut Bacteria Enzyme Found to Break Down Heat-Processed Food Compounds, Producing Novel Biogenic Amines
Researchers have discovered that an enzyme in common gut bacteria can degrade N-epsilon-carboxymethyllysine (CML), a compound formed during thermal food processing, producing previously unknown biogenic amines. The enzyme, ornithine decarboxylase SpeC from enterobacteria, acts on CML and related modified lysine derivatives through a low-level 'underground' catalytic activity. This finding suggests a previously unrecognized communication axis between thermally processed dietary compounds and gut microbial physiology, with potential implications for host health.
Full-Length Gene Sequencing Reveals Two Distinct Bacterial Communities in Black-Legged Ticks Expanding Into Canada
Researchers used Oxford Nanopore full-length 16S rRNA gene sequencing to characterize the microbiome of Ixodes scapularis black-legged ticks collected in Nova Scotia, Canada, distinguishing between tick-adapted bacteria and environmentally acquired bacteria. The study comes as I. scapularis — the primary vector of Lyme disease — is rapidly expanding northward into Canada due to climate change. The findings suggest that environmentally derived bacteria in tick microbiomes are not mere contamination, which has implications for how tick microbiome data is collected and interpreted across surveillance studies.
Study Identifies Metabolic Link Between Cell Envelope Stress and Biofilm Formation in Bacteria
Researchers have discovered that the metabolite acetyl-CoA directly inhibits enzymes that degrade the bacterial signaling molecule c-di-GMP, connecting cell envelope biosynthesis stress to biofilm formation in Pseudomonas aeruginosa. The study found that sub-inhibitory concentrations of antibiotics targeting early peptidoglycan biosynthesis — but not other antibiotic classes — elevate c-di-GMP levels by reducing phosphodiesterase activity, with acetyl-CoA competing for the enzyme active site. Because the relevant enzyme domain is broadly conserved across bacterial species, this checkpoint mechanism may be widespread and could have implications for understanding antibiotic-induced biofilm responses.