← Back to Research Radar
Academic Publication Academic Publication

BRAKER3: Fully automated genome annotation using RNA-seq and protein evidence with GeneMark-ETP, AUGUSTUS, and TSEBRA

516
Citations
May 1, 2024
Published Date

Research Abstract & Technology Focus

Gene prediction has remained an active area of bioinformatics research for a long time. Still, gene prediction in large eukaryotic genomes presents a challenge that must be addressed by new algorithms. The amount and significance of the evidence available from transcriptomes and proteomes vary across genomes, between genes, and even along a single gene. User-friendly and accurate annotation pipelines that can cope with such data heterogeneity are needed. The previously developed annotation pipelines BRAKER1 and BRAKER2 use RNA-seq or protein data, respectively, but not both. A further significant performance improvement integrating all three data types was made by the recently released GeneMark-ETP. We here present the BRAKER3 pipeline that builds on GeneMark-ETP and AUGUSTUS, and further improves accuracy using the TSEBRA combiner. BRAKER3 annotates protein-coding genes in eukaryotic genomes using both short-read RNA-seq and a large protein database, along with statistical models learned iteratively and specifically for the target genome. We benchmarked the new pipeline on genomes of 11 species under an assumed level of relatedness of the target species proteome to available proteomes. BRAKER3 outperforms BRAKER1 and BRAKER2. The average transcript-level F1-score is increased by about 20 percentage points on average, whereas the difference is most pronounced for species with large and complex genomes. BRAKER3 also outperforms other existing tools, MAKER2, Funannotate, and FINDER. The code of BRAKER3 is available on GitHub and as a ready-to-run Docker container for execution with Docker or Singularity. Overall, BRAKER3 is an accurate, easy-to-use tool for eukaryotic genome annotation.
Read Full Literature

AI Semantic Synergy Context

Connecting this academic literature to real-world market discussions and products.

crossref.org › academic paper
0%

Leveraging functional genomic annotations and genome coverage to improve polygenic prediction of complex traits within and between ancestries

Abstract We develop a method, SBayesRC, that integrates genome-wide association study (GWAS) summary statistics with functional genomic annotations to improve polygenic prediction...

crossref.org › academic paper
0%

edgeR v4: powerful differential analysis of sequencing data with expanded functionality and improved support for small counts and larger datasets

Abstract edgeR is an R/Bioconductor software package for differential analyses of sequencing data in the form of read counts for genes or genomic features. Over the past 15 years,...

crossref.org › academic paper
0%

ChIP-Atlas 3.0: a data-mining suite to explore chromosome architecture together with large-scale regulome data

Abstract ChIP-Atlas (https://chip-atlas.org/) presents a suite of data-mining tools for analyzing epigenomic landscapes, powered by the comprehensive integration of over 376 000 publ...

crossref.org › academic paper
0%

AlphaGenome: advancing regulatory variant effect prediction with a unified DNA sequence model

Deep learning models that predict functional genomic measurements from DNA sequence are powerful tools for deciphering the genetic regulatory code. Existing methods trade off between input sequence...

crossref.org › academic paper
0%

InterPro: the protein sequence classification resource in 2025

Abstract InterPro (https://www.ebi.ac.uk/interpro) is a freely accessible resource for the classification of protein sequences into families. It integrates predictive models, known a...

Frequently Asked Questions (FAQ)

Curated market intelligence mapped to this research.

What is the core focus of the research titled 'BRAKER3: Fully automated genome annotation using RNA-seq and protein evidence with GeneMark-ETP, AUGUSTUS, and TSEBRA'?

This literature focuses on: Gene prediction has remained an active area of bioinformatics research for a long time. Still, gene prediction in large eukaryotic genomes presents a challenge that must be addressed by new algorithms. The amount and significance of the evidence a...

Are there open-source GitHub repositories related to BRAKER3: Fully automated genome annotation using RNA-seq and protein evidence with GeneMark-ETP, AUGUSTUS, and TSEBRA?

Yes, open-source projects like aiming-lab/AutoResearchClaw (Fully autonomous & self-evolving research from idea to paper. Chat an Idea. Get a Paper. 🦞) are actively building upon these concepts.

Which startups are commercializing the technology behind BRAKER3: Fully automated genome annotation using RNA-seq and protein evidence with GeneMark-ETP, AUGUSTUS, and TSEBRA?

Products like OpenOwl are bringing this to market. Their focus is: Automate what APIs can't, one prompt, fully handled.

What other academic literature is closely related to 'BRAKER3: Fully automated genome annotation using RNA-seq and protein evidence with GeneMark-ETP, AUGUSTUS, and TSEBRA'?

Yes, highly correlated activity was mapped. An entry titled 'Leveraging functional genomic annotations and genome coverage to improve polygenic prediction of complex traits within and between ancestries' discusses this: Abstract We develop a method, SBayesRC, that integrates genome-wide association study (GWAS) summary statistics with functional g...

Cite this Market Intelligence Report

Reference our AI-mapped synergy between this research and the commercial market to instantly build authority.

Commercial Realization

Startups and Open Source tools heavily associated with the concepts explored in this paper.

Associated Media Narrative