HuggingFace released Carbon-A on October 8, 2026, saying it predicts protein-coding regions directly from DNA using a single model across diverse organisms. It is the company's first major model update since its initial bioinformatics work in 2024.
HuggingFace reported a macro-averaged nucleotide F1 of 0.944, measured on 42 benchmark genomes. That compares with the baseline models like AUGUSTUS and Tiberius, which had lower scores.
Carbon-A is built on a 98,304-base-pair context window and targets genome annotation, which is the bottleneck in understanding biological sequences. Availability begins with the Carbon Annotation Database, initially for researchers and biologists.
"We've used this model to discover 566 million new gene candidates across 22,617 species," said Georgia Channing, HuggingFaceBio community lead. The model's ability to generalize across species is a key innovation.
The announcement follows growing interest in AI for biology, with major labs preparing to tackle the field. HuggingFace said it aims to keep the open science ecosystem going by making models and insights accessible.
HuggingFace did not say how many of the newly predicted genes are functional, and the open question is whether these RNAs are translated into proteins. The next step is to look for evidence of translation using ribosome profiling.
Source: huggingface