What biochemistry, genetics and molecular biology is citing right now

The 40 most-cited open-access biochemistry, genetics and molecular biology papers published since 2023. Every one is free to read, and each links to the journal that published it — so you can see what that venue costs to publish in while you are there.

Ranked by citations recorded in OpenAlex. The leading paper here has 15,144. Citation counts favour older papers, so treat this as what the field is building on, not what is newest.

Accurate structure prediction of biomolecular interactions with AlphaFold 3

Josh Abramson, Jonas Adler, Jack Dunger, Richard Evans and 44 more · 2024

Abstract The introduction of AlphaFold 21 has spurred a revolution in modelling the structure of proteins and their interactions, enabling a huge range of applications in protein modelling and design2–6. Here we describe our AlphaFold 3 model with a substantially updated diffusion-based architecture that is capable of predicting the joint structure of complexes including proteins, nucleic acids, small molecules, ions and modified residues. The new AlphaFold model demonstrates substantially improved accuracy over many previous specialized tools: far greater accuracy for protein–ligand interacti…

Nature15,144 citationsDOI

Targeted Branching for the Maximum Independent Set Problem Using Graph Neural Networks

Silva, Gabriel, Rodrigues, Mário, Teixeira, António, Amorim, Marlene · 2024

Identifying a maximum independent set is a fundamental NP-hard problem. This problem has several real-world applications and requires finding the largest possible set of vertices not adjacent to each other in an undirected graph. Over the past few years, branch-and-bound and branch-and-reduce algorithms have emerged as some of the most effective methods for solving the problem exactly. Specifically, the branch-and-reduce approach, which combines branch-and-bound principles with reduction rules, has proven particularly successful in tackling previously unmanageable real-world instances. This pr…

DROPS (Schloss Dagstuhl – Leibniz Center for Informatics)5,410 citationsDOI

Evolutionary-scale prediction of atomic-level protein structure with a language model

Zeming Lin, Halil Akin, Roshan Rao, Brian Hie and 11 more · 2023

Recent advances in machine learning have leveraged evolutionary information in multiple sequence alignments to predict protein structure. We demonstrate direct inference of full atomic-level protein structure from primary sequence using a large language model. As language models of protein sequences are scaled up to 15 billion parameters, an atomic-resolution picture of protein structure emerges in the learned representations. This results in an order-of-magnitude acceleration of high-resolution structure prediction, which enables large-scale structural characterization of metagenomic proteins…

Science5,320 citationsDOI

Interactive Tree of Life (iTOL) v6: recent updates to the phylogenetic tree display and annotation tool

Ivica Letunić, Peer Bork · 2024

The Interactive Tree Of Life (https://itol.embl.de) is an online tool for the management, display, annotation and manipulation of phylogenetic and other trees. It is freely available and open to everyone. iTOL version 6 introduces a modernized and completely rewritten user interface, together with numerous new features. A new dataset type has been introduced (colored/labeled ranges), greatly upgrading the functionality of the previous simple colored range annotation function. Additional annotation options have been implemented for several existing dataset types. Dataset template files now supp…

Nucleic Acids Research4,275 citationsDOI

FinnGen provides genetic insights from a well-phenotyped isolated population

Mitja Kurki, Juha Karjalainen, Priit Palta, Timo P. Sipilä and 96 more · 2023

Abstract Population isolates such as those in Finland benefit genetic research because deleterious alleles are often concentrated on a small number of low-frequency variants (0.1% ≤ minor allele frequency < 5%). These variants survived the founding bottleneck rather than being distributed over a large number of ultrarare variants. Although this effect is well established in Mendelian genetics, its value in common disease genetics is less explored 1,2 . FinnGen aims to study the genome and national health register data of 500,000 Finnish individuals. Given the relatively high median age of part…

Nature4,248 citationsDOI

Minimal information for studies of extracellular vesicles (MISEV2023): From basic to advanced approaches

Joshua A Welsh, Deborah C. I. Goberdhan, Lorraine O’Driscoll, Edit I. Buzás and 68 more · 2024

Extracellular vesicles (EVs), through their complex cargo, can reflect the state of their cell of origin and change the functions and phenotypes of other cells. These features indicate strong biomarker and therapeutic potential and have generated broad interest, as evidenced by the steady year-on-year increase in the numbers of scientific publications about EVs. Important advances have been made in EV metrology and in understanding and applying EV biology. However, hurdles remain to realising the potential of EVs in domains ranging from basic biology to clinical applications due to challenges…

Journal of Extracellular Vesicles4,061 citationsDOI

The Gene Ontology knowledgebase in 2023

Suzi Aleksander, James P. Balhoff, Seth Carbon, J. Michael Cherry and 96 more · 2023

The Gene Ontology (GO) knowledgebase (http://geneontology.org) is a comprehensive resource concerning the functions of genes and gene products (proteins and noncoding RNAs). GO annotations cover genes from organisms across the tree of life as well as viruses, though most gene function knowledge currently derives from experiments carried out in a relatively small number of model organisms. Here, we provide an updated overview of the GO knowledgebase, as well as the efforts of the broad, international consortium of scientists that develops, maintains, and updates the GO knowledgebase. The GO kno…

Genetics2,813 citationsDOI

Fast and accurate protein structure search with Foldseek

Michel van Kempen, Stephanie Kim, Charlotte Tumescheit, Milot Mirdita and 4 more · 2023

As structure prediction methods are generating millions of publicly available protein structures, searching these databases is becoming a bottleneck. Foldseek aligns the structure of a query protein against a database by describing tertiary amino acid interactions within proteins as sequences over a structural alphabet. Foldseek decreases computation times by four to five orders of magnitude with 86%, 88% and 133% of the sensitivities of Dali, TM-align and CE, respectively.…

Nature Biotechnology2,519 citationsDOI

SRplot: A free online platform for data visualization and graphing

Doudou Tang, Mingjie Chen, Xinhua Huang, Guicheng Zhang and 4 more · 2023

Graphics are widely used to provide summarization of complex data in scientific publications. Although there are many tools available for drawing graphics, their use is limited by programming skills, costs, and platform specificities. Here, we presented a freely accessible easy-to-use web server named SRplot that integrated more than a hundred of commonly used data visualization and graphing functions together. It can be run easily using all Web browsers and there are no strong requirements on the computing power of users' machines. With a user-friendly graphical interface, users can simply pa…

PLoS ONE2,480 citationsDOI

KEGG: biological systems database as a model of the real world

Minoru Kanehisa, Miho Furumichi, Yoko Sato, Yuriko Matsuura and 1 more · 2024

KEGG (https://www.kegg.jp/) is a database resource for representation and analysis of biological systems. Pathway maps are the primary dataset in KEGG representing systemic functions of the cell and the organism in terms of molecular interaction and reaction networks. The KEGG Orthology (KO) system is a mechanism for linking genes and proteins to pathway maps and other molecular networks. Each KO is a generic gene identifier and each pathway map is created as a network of KO nodes. This architecture enables KEGG pathway mapping to uncover systemic features from KO assigned genomes and metageno…

Nucleic Acids Research2,367 citationsDOI

UniProt: the Universal Protein Knowledgebase in 2025

Alex Bateman, María Martin, Sandra Orchard, Michele Magrane and 95 more · 2024

The aim of the UniProt Knowledgebase (UniProtKB; https://www.uniprot.org/) is to provide users with a comprehensive, high-quality and freely accessible set of protein sequences annotated with functional information. In this publication, we describe ongoing changes to our production pipeline to limit the sequences available in UniProtKB to high-quality, non-redundant reference proteomes. We continue to manually curate the scientific literature to add the latest functional data and use machine learning techniques. We also encourage community curation to ensure key publications are not missed. We…

Nucleic Acids Research2,238 citationsDOI

MetaboAnalyst 6.0: towards a unified platform for metabolomics data processing, analysis and interpretation

Zhiqiang Pang, Yao Lü, Guangyan Zhou, Fiona Hui and 7 more · 2024

We introduce MetaboAnalyst version 6.0 as a unified platform for processing, analyzing, and interpreting data from targeted as well as untargeted metabolomics studies using liquid chromatography - mass spectrometry (LC-MS). The two main objectives in developing version 6.0 are to support tandem MS (MS2) data processing and annotation, as well as to support the analysis of data from exposomics studies and related experiments. Key features of MetaboAnalyst 6.0 include: (i) a significantly enhanced Spectra Processing module with support for MS2 data and the asari algorithm; (ii) a MS2 Peak Annota…

Nucleic Acids Research2,170 citationsDOI

Accurate proteome-wide missense variant effect prediction with AlphaMissense

Jun Cheng, Guido Novati, Joshua Pan, Clare Bycroft and 12 more · 2023

The vast majority of missense variants observed in the human genome are of unknown clinical significance. We present AlphaMissense, an adaptation of AlphaFold fine-tuned on human and primate variant population frequency databases to predict missense variant pathogenicity. By combining structural context and evolutionary conservation, our model achieves state-of-the-art results across a wide range of genetic and experimental benchmarks, all without explicitly training on such data. The average pathogenicity score of genes is also predictive for their cell essentiality, capable of identifying sh…

Science2,167 citationsDOI

De novo design of protein structure and function with RFdiffusion

Joseph L. Watson, David Juergens, Nathaniel R. Bennett, Brian L. Trippe and 24 more · 2023

Abstract There has been considerable recent progress in designing new proteins using deep-learning methods 1–9 . Despite this progress, a general deep-learning framework for protein design that enables solution of a wide range of design challenges, including de novo binder design and design of higher-order symmetric architectures, has yet to be described. Diffusion models 10,11 have had considerable success in image and language generative modelling but limited success when applied to protein modelling, probably due to the complexity of protein backbone geometry and sequence–structure relation…

Nature2,148 citationsDOI

Ultrafast one‐pass FASTQ data preprocessing, quality control, and deduplication using fastp

Shifu Chen · 2023

A large amount of sequencing data is generated and processed every day with the continuous evolution of sequencing technology and the expansion of sequencing applications. One consequence of such sequencing data explosion is the increasing cost and complexity of data processing. The preprocessing of FASTQ data, which means removing adapter contamination, filtering low-quality reads, and correcting wrongly represented bases, is an indispensable but resource intensive part of sequencing data analysis. Therefore, although a lot of software applications have been developed to solve this problem, b…

iMeta2,115 citationsDOI

PI3K/AKT/mTOR signaling transduction pathway and targeted therapies in cancer

Antonino Glaviano, Aaron Song Chuan Foo, Hiu Yan Lam, Kenneth Chun-Yong Yap and 23 more · 2023

The PI3K/AKT/mTOR (PAM) signaling pathway is a highly conserved signal transduction network in eukaryotic cells that promotes cell survival, cell growth, and cell cycle progression. Growth factor signalling to transcription factors in the PAM axis is highly regulated by multiple cross-interactions with several other signaling pathways, and dysregulation of signal transduction can predispose to cancer development. The PAM axis is the most frequently activated signaling pathway in human cancer and is often implicated in resistance to anticancer therapies. Dysfunction of components of this pathwa…

Molecular Cancer2,043 citationsDOI

AmberTools

David A. Case, Hasan Metin Aktulga, Kellon Belfon, David S. Cerutti and 43 more · 2023

High Resolution Image Download MS PowerPoint Slide AmberTools is a free and open-source collection of programs used to set up, run, and analyze molecular simulations. The newer features contained within AmberTools23 are briefly described in this Application note.…

Journal of Chemical Information and Modeling2,034 citationsDOI

NF-κB in biology and targeted therapy: new insights and translational implications

Qing Guo, Yizi Jin, Xinyu Chen, Xiaomin Ye and 5 more · 2024

NF-κB signaling has been discovered for nearly 40 years. Initially, NF-κB signaling was identified as a pivotal pathway in mediating inflammatory responses. However, with extensive and in-depth investigations, researchers have discovered that its role can be expanded to a variety of signaling mechanisms, biological processes, human diseases, and treatment options. In this review, we first scrutinize the research process of NF-κB signaling, and summarize the composition, activation, and regulatory mechanism of NF-κB signaling. We investigate the interaction of NF-κB signaling with other importa…

Signal Transduction and Targeted Therapy1,883 citationsDOI

MitoHiFi: a python pipeline for mitochondrial genome assembly from PacBio high fidelity reads

Marcela Uliano‐Silva, João Gabriel R. N. Ferreira, Ksenia Krasheninnikova, Mark Blaxter and 23 more · 2023

BACKGROUND: PacBio high fidelity (HiFi) sequencing reads are both long (15-20 kb) and highly accurate (> Q20). Because of these properties, they have revolutionised genome assembly leading to more accurate and contiguous genomes. In eukaryotes the mitochondrial genome is sequenced alongside the nuclear genome often at very high coverage. A dedicated tool for mitochondrial genome assembly using HiFi reads is still missing. RESULTS: MitoHiFi was developed within the Darwin Tree of Life Project to assemble mitochondrial genomes from the HiFi reads generated for target species. The input for MitoH…

BMC Bioinformatics1,876 citationsDOI

Proksee: in-depth characterization and visualization of bacterial genomes

Jason R. Grant, Eric Enns, Eric Marinier, Arnab Mandal and 5 more · 2023

Proksee (https://proksee.ca) provides users with a powerful, easy-to-use, and feature-rich system for assembling, annotating, analysing, and visualizing bacterial genomes. Proksee accepts Illumina sequence reads as compressed FASTQ files or pre-assembled contigs in raw, FASTA, or GenBank format. Alternatively, users can supply a GenBank accession or a previously generated Proksee map in JSON format. Proksee then performs assembly (for raw sequence data), generates a graphical map, and provides an interface for customizing the map and launching further analysis jobs. Notable features of Proksee…

Nucleic Acids Research1,840 citationsDOI

g:Profiler—interoperable web service for functional enrichment analysis and gene identifier mapping (2023 update)

Liis Kolberg, Uku Raudvere, Ivan Kuzmin, Priit Adler and 2 more · 2023

g:Profiler is a reliable and up-to-date functional enrichment analysis tool that supports various evidence types, identifier types and organisms. The toolset integrates many databases, including Gene Ontology, KEGG and TRANSFAC, to provide a comprehensive and in-depth analysis of gene lists. It also provides interactive and intuitive user interfaces and supports ordered queries and custom statistical backgrounds, among other settings. g:Profiler provides multiple programmatic interfaces to access its functionality. These can be easily integrated into custom workflows and external tools, making…

Nucleic Acids Research1,700 citationsDOI

Plasma proteomic associations with genetics and health in the UK Biobank

Benjamin B. Sun, Joshua Chiou, Matthew Traylor, Christian Benner and 60 more · 2023

Abstract The Pharma Proteomics Project is a precompetitive biopharmaceutical consortium characterizing the plasma proteomic profiles of 54,219 UK Biobank participants. Here we provide a detailed summary of this initiative, including technical and biological validations, insights into proteomic disease signatures, and prediction modelling for various demographic and health indicators. We present comprehensive protein quantitative trait locus (pQTL) mapping of 2,923 proteins that identifies 14,287 primary genetic associations, of which 81% are previously undescribed, alongside ancestry-specific…

Nature1,628 citationsDOI

Artificial Hallucinations in ChatGPT: Implications in Scientific Writing

Hussam Alkaissi, Samy I. McFarlane · 2023

While still in its infancy, ChatGPT (Generative Pretrained Transformer), introduced in November 2022, is bound to hugely impact many industries, including healthcare, medical education, biomedical research, and scientific writing. Implications of ChatGPT, that new chatbot introduced by OpenAI on academic writing, is largely unknown. In response to the Journal of Medical Science (Cureus) Turing Test - call for case reports written with the assistance of ChatGPT, we present two cases one of homocystinuria-associated osteoporosis, and the other is on late-onset Pompe disease (LOPD), a rare metabo…

Cureus1,516 citationsDOI

Hypoxic microenvironment in cancer: molecular mechanisms and therapeutic interventions

Zhou Chen, Fangfang Han, Yan Du, Huaqing Shi and 1 more · 2023

Having a hypoxic microenvironment is a common and salient feature of most solid tumors. Hypoxia has a profound effect on the biological behavior and malignant phenotype of cancer cells, mediates the effects of cancer chemotherapy, radiotherapy, and immunotherapy through complex mechanisms, and is closely associated with poor prognosis in various cancer patients. Accumulating studies have demonstrated that through normalization of the tumor vasculature, nanoparticle carriers and biocarriers can effectively increase the oxygen concentration in the tumor microenvironment, improve drug delivery an…

Signal Transduction and Targeted Therapy1,511 citationsDOI

The Reactome Pathway Knowledgebase 2024

M Orlic-Milacic, Deidre Beavers, Patrick Conley, Chuqiao Gong and 20 more · 2023

The Reactome Knowledgebase (https://reactome.org), an Elixir and GCBR core biological data resource, provides manually curated molecular details of a broad range of normal and disease-related biological processes. Processes are annotated as an ordered network of molecular transformations in a single consistent data model. Reactome thus functions both as a digital archive of manually curated human biological processes and as a tool for discovering functional relationships in data such as gene expression profiles or somatic mutation catalogs from tumor cells. Here we review progress towards anno…

Nucleic Acids Research1,456 citationsDOI

Extending and improving metagenomic taxonomic profiling with uncharacterized species using MetaPhlAn 4

Aitor Blanco‐Míguez, Francesco Beghini, Fabio Cumbo, Lauren J. McIver and 22 more · 2023

Metagenomic assembly enables new organism discovery from microbial communities, but it can only capture few abundant organisms from most metagenomes. Here we present MetaPhlAn 4, which integrates information from metagenome assemblies and microbial isolate genomes for more comprehensive metagenomic taxonomic profiling. From a curated collection of 1.01 M prokaryotic reference and metagenome-assembled genomes, we define unique marker genes for 26,970 species-level genome bins, 4,992 of them taxonomically unidentified at the species level. MetaPhlAn 4 explains ~20% more reads in most internation…

Nature Biotechnology1,445 citationsDOI

Tree Visualization By One Table (tvBOT): a web application for visualizing, modifying and annotating phylogenetic trees

Jianmin Xie, Yuerong Chen, Guanjing Cai, Runlin Cai and 2 more · 2023

tvBOT is a user-friendly and efficient web application for visualizing, modifying, and annotating phylogenetic trees. It is highly efficient in data preparation without requiring redundant style and syntax data. Tree annotations are powered by a data-driven engine that only requires practical data organized in uniform formats and saved as one table file. A layer manager is developed to manage annotation dataset layers, allowing the addition of a specific layer by selecting the columns of a corresponding annotation data file. Furthermore, tvBOT renders style adjustments in real-time and diversi…

Nucleic Acids Research1,302 citationsDOI

Angiogenic signaling pathways and anti-angiogenic therapy for cancer

Zhenling Liu, Huanhuan Chen, Lili Zheng, Li‐Ping Sun and 1 more · 2023

Angiogenesis, the formation of new blood vessels, is a complex and dynamic process regulated by various pro- and anti-angiogenic molecules, which plays a crucial role in tumor growth, invasion, and metastasis. With the advances in molecular and cellular biology, various biomolecules such as growth factors, chemokines, and adhesion factors involved in tumor angiogenesis has gradually been elucidated. Targeted therapeutic research based on these molecules has driven anti-angiogenic treatment to become a promising strategy in anti-tumor therapy. The most widely used anti-angiogenic agents include…

Signal Transduction and Targeted Therapy1,258 citationsDOI

Extracellular vesicles as tools and targets in therapy for diseases

Mudasir A. Kumar, Sadaf Khursheed Baba, Hana Q. Sadida, Sara Al Marzooqi and 10 more · 2024

Extracellular vesicles (EVs) are nano-sized, membranous structures secreted into the extracellular space. They exhibit diverse sizes, contents, and surface markers and are ubiquitously released from cells under normal and pathological conditions. Human serum is a rich source of these EVs, though their isolation from serum proteins and non-EV lipid particles poses challenges. These vesicles transport various cellular components such as proteins, mRNAs, miRNAs, DNA, and lipids across distances, influencing numerous physiological and pathological events, including those within the tumor microenvi…

Signal Transduction and Targeted Therapy1,247 citationsDOI

Next-Generation Sequencing Technology: Current Trends and Advancements

Heena Satam, Kandarp Joshi, Upasana Mangrolia, Sanober Waghoo and 7 more · 2023

The advent of next-generation sequencing (NGS) has brought about a paradigm shift in genomics research, offering unparalleled capabilities for analyzing DNA and RNA molecules in a high-throughput and cost-effective manner. This transformative technology has swiftly propelled genomics advancements across diverse domains. NGS allows for the rapid sequencing of millions of DNA fragments simultaneously, providing comprehensive insights into genome structure, genetic variations, gene expression profiles, and epigenetic modifications. The versatility of NGS platforms has expanded the scope of genomi…

Biology1,237 citationsDOI

2. Diagnosis and Classification of Diabetes: Standards of Care in Diabetes—2025

Nuha A. ElSayed, Rozalina G. McCoy, Grazia Aleppo, Kirthikaa Balapattabi and 21 more · 2024

The American Diabetes Association (ADA) "Standards of Care in Diabetes" includes the ADA's current clinical practice recommendations and is intended to provide the components of diabetes care, general treatment goals and guidelines, and tools to evaluate quality of care. Members of the ADA Professional Practice Committee, an interprofessional expert committee, are responsible for updating the Standards of Care annually, or more frequently as warranted. For a detailed description of ADA standards, statements, and reports, as well as the evidence-grading system for ADA's clinical practice recomm…

Diabetes Care1,223 citationsDOI

Mitochondrial dynamics in health and disease: mechanisms and potential targets

Wen Chen, Huakan Zhao, Yongsheng Li · 2023

Mitochondria are organelles that are able to adjust and respond to different stressors and metabolic needs within a cell, showcasing their plasticity and dynamic nature. These abilities allow them to effectively coordinate various cellular functions. Mitochondrial dynamics refers to the changing process of fission, fusion, mitophagy and transport, which is crucial for optimal function in signal transduction and metabolism. An imbalance in mitochondrial dynamics can disrupt mitochondrial function, leading to abnormal cellular fate, and a range of diseases, including neurodegenerative disorders,…

Signal Transduction and Targeted Therapy1,220 citationsDOI

Short-Chain Fatty-Acid-Producing Bacteria: Key Components of the Human Gut Microbiota

William G. Fusco, Manuel Bernabeu, Marco Cintoni, Serena Porcari and 8 more · 2023

Short-chain fatty acids (SCFAs) play a key role in health and disease, as they regulate gut homeostasis and their deficiency is involved in the pathogenesis of several disorders, including inflammatory bowel diseases, colorectal cancer, and cardiometabolic disorders. SCFAs are metabolites of specific bacterial taxa of the human gut microbiota, and their production is influenced by specific foods or food supplements, mainly prebiotics, by the direct fostering of these taxa. This Review provides an overview of SCFAs' roles and functions, and of SCFA-producing bacteria, from their microbiological…

Nutrients1,185 citationsDOI