,

A gene-finding model and a database of 566 million predicted loci

Hugging Face’s biology research group has released Carbon-Annotator, a 1.2-billion-parameter model that predicts protein-coding regions directly from DNA sequence, and the database it produced by running it across public genome assemblies.

The model reads a context window of 98,304 base pairs and predicts on both strands, covering mammals, other vertebrates, invertebrates, plants, fungi and protists. The accompanying database holds 566 million predicted protein-coding loci across 22,617 taxa and 48,167 assemblies, roughly 27 trillion base pairs. That is about eleven times more taxa and nine times more sequence than the curated corpus the model was trained on, which is the point: most sequenced genomes have never been annotated to the standard the well-studied ones enjoy.

Reported accuracy is a macro-averaged nucleotide F1 of 0.944 across 42 benchmark genomes, which the authors say beats seven existing tools at nucleotide, exon and gene level.

The caveats are theirs and they are substantial. Training labels came from a curated reference set that is itself imperfect and changes over time, so the claim that the model generalises past its training annotations is their interpretation. Long-read transcript data supports where exons are, not that a protein is made. The decoder emits one consensus path per locus, so alternative isoforms are not resolved. About half the target genomes are annotated so far, and no licence is stated.

Source: Hugging Face Bio, 8 October 2026.


Related