SoftMatcha 2: New Algorithm Enables Fast Semantic Search Across Trillion-Token Corpora
Researchers have developed SoftMatcha 2, a search algorithm capable of querying trillion-scale natural language corpora in under 0.3 seconds while supporting semantic variations such as substitutions, insertions, and deletions. The method combines suffix array-based string matching with word vector representations, using dynamic corpus-aware pruning and a disk-aware design to prevent combinatorial explosion. The system outperforms existing tools like infini-gram and the original SoftMatcha, and has practical applications in detecting benchmark contamination in AI training data, information retrieval, and paraphrase detection.
SoftMatcha 2 is an ultra-fast, semantically flexible text search algorithm presented by researchers and accepted at ICML 2026. It can search corpora of up to 1.4 trillion tokens—such as FineWeb-Edu—in under 0.3 seconds, substantially outperforming prior methods including infini-gram, infini-gram mini, and the original SoftMatcha. The algorithm is built on suffix arrays for scalable string matching and uses word vector representations to enable semantic flexibility, allowing queries to match paraphrases and near-duplicates through substitution, insertion, and deletion operations. Two core algorithmic innovations—dynamic corpus-aware pruning and fast exact lookup via a disk-aware design—are used to theoretically and empirically mitigate the exponential growth in search space that semantic relaxation would otherwise cause. A key practical application is uncovering benchmark contamination in large language model training corpora that existing tools fail to detect, which has direct implications for the validity of AI model evaluations. The researchers also provide an online demo supporting soft search across corpora in seven languages, making the tool broadly accessible. Source code and a project page are publicly available.
What's missing
The paper does not detail the hardware configuration required to achieve sub-0.3-second latency at trillion-token scale, which is relevant for assessing real-world deployability. The scope of benchmark contamination cases discovered is not quantified in the abstract.
What different sources said
- arXiv cs.LGCenter
SoftMatcha 2: A Fast and Soft Pattern Matcher for Trillion-Scale Corpora
Related
Gut Bacteria Enzyme Found to Break Down Heat-Processed Food Compounds, Producing Novel Biogenic Amines
Researchers have discovered that an enzyme in common gut bacteria can degrade N-epsilon-carboxymethyllysine (CML), a compound formed during thermal food processing, producing previously unknown biogenic amines. The enzyme, ornithine decarboxylase SpeC from enterobacteria, acts on CML and related modified lysine derivatives through a low-level 'underground' catalytic activity. This finding suggests a previously unrecognized communication axis between thermally processed dietary compounds and gut microbial physiology, with potential implications for host health.
Full-Length Gene Sequencing Reveals Two Distinct Bacterial Communities in Black-Legged Ticks Expanding Into Canada
Researchers used Oxford Nanopore full-length 16S rRNA gene sequencing to characterize the microbiome of Ixodes scapularis black-legged ticks collected in Nova Scotia, Canada, distinguishing between tick-adapted bacteria and environmentally acquired bacteria. The study comes as I. scapularis — the primary vector of Lyme disease — is rapidly expanding northward into Canada due to climate change. The findings suggest that environmentally derived bacteria in tick microbiomes are not mere contamination, which has implications for how tick microbiome data is collected and interpreted across surveillance studies.
Study Identifies Metabolic Link Between Cell Envelope Stress and Biofilm Formation in Bacteria
Researchers have discovered that the metabolite acetyl-CoA directly inhibits enzymes that degrade the bacterial signaling molecule c-di-GMP, connecting cell envelope biosynthesis stress to biofilm formation in Pseudomonas aeruginosa. The study found that sub-inhibitory concentrations of antibiotics targeting early peptidoglycan biosynthesis — but not other antibiotic classes — elevate c-di-GMP levels by reducing phosphodiesterase activity, with acetyl-CoA competing for the enzyme active site. Because the relevant enzyme domain is broadly conserved across bacterial species, this checkpoint mechanism may be widespread and could have implications for understanding antibiotic-induced biofilm responses.