tgindex
Chem ML/AI/Datasets

Chem ML/AI/Datasets

Статистика
@chem_mlанглийский

Daily articles and news from the field of machine learning in chemistry from the researchers of IGIC RAS @chemrussia For contact: @levkrasnov @st613laboratory @StasBezzubov

Последний пост
18:23
Последнее чтение
12 авг.
Постов за неделю
2
Всего постов
43
Тип
открытый
Язык
английский
В каталоге с
12 авг.
Подписчики
824
−3 за 3 дн.
Сутки
0
0,00%
Неделя
 
Месяц
 
Просмотров на пост
414
40 постов
Вовлечённость
50,2%
к подписчикам
Постов в день
0,3
всего 43
Упоминаний
2
каналов
Охват размещения
оценка
1/24сутки в ленте
243
1/48двое суток
278
1/72трое суток
300

Оценка по просмотрам недавних постов: пост набирает почти всё за первые сутки.

Посты

  • 18:23160158

    видео или голосовое, без подписи

  • 10 авг.2741410

    Artificial intelligence in drug discovery — what it is, where we stand and the path forward https://www.nature.com/articles/s41573-026-01496-2 In this Perspective we discuss potential reasons, including an insufficient focus on clinical translation during model development, difficulties with applying AI algorithms on conditional life science data, and insufficient problem definitions and the resulting underspecification of computational models for real-world use cases. ‘Technology push’ compared with ‘science pull’ is also likely to be an underlying factor, as well as the substantial time required to operationalize technical capabilities into systems that are sufficiently scaled and accessible for users. We provide recommendations for the development of AI in drug discovery with the aim of increasing its translational relevance. For example, benchmarking studies of AI tools in drug discovery need to move on from model validation and instead focus on their ability to improve decision making. 📕 nature reviews drug discovery (IF = 91.2)

  • 7 авг.372106

    Benchmarking and developing large language models using one million clinical trials 🔥 https://www.nature.com/articles/s41746-026-02933-7 Here, we introduce TrialPanorama, a large-scale structured resource aggregating 1.6M clinical trial records from fifteen global registries linked with biomedical ontologies and literature. Using this resource, we construct 152K training and testing samples spanning eight clinical research tasks, including systematic review, trial design, and trial optimization. Benchmarking cutting-edge large language models (LLMs) reveals limited clinical reasoning capability in generic LLMs. In contrast, an 8B LLM developed on TrialPanorama using supervised fine-tuning and reinforcement learning outperforms 70B generic counterparts across all eight tasks, with relative improvements of 73.7, 67.6, 38.4, 37.8, 26.5, 20.7, 20.0, 18.1, and 5.2%, respectively. These results demonstrate the potential of domain-adapted AI to improve evidence synthesis and clinical trial design, establishing TrialPanorama as a foundation for scaling AI in clinical research. 📕npj digital medicine (IF = 18.0)

  • 4 авг.330321

    видео или голосовое, без подписи

  • 4 авг.318108

    The past, present and future of self-driving laboratories https://www.nature.com/articles/s41570-026-00847-2 This Review traces the evolution of self-driving laboratories and examines the structural asymmetries that limit their maturation into shared scientific infrastructure. We frame the next phase of the field around three interdependent requirements: scalability, generalizability and provenance-complete experimentation. Realizing collective scientific superintelligence will require SDLs that reliably scale throughput, transfer workflows and learned models across laboratories and scientific domains and capture end-to-end experimental data and metadata from precursor preparation through synthesis, characterization and performance evaluation. Achieving this transition will depend on interoperable data and metadata standards, modular and integrable experimental hardware, and trustworthy artificial intelligence agents that reason under uncertainty within rigorous safety and ethical boundaries. 📕Nature Reviews Chemistry (IF=50.3)

  • 29 июл.4014218

    видео или голосовое, без подписи

  • Do Language Models Dream of Binding Molecules? Benchmarking LLMs under Spatial Constraints🔥 https://arxiv.org/abs/2607.18144 In this work, we systematically analyze whether current general-purpose LLMs are capable of navigating complex 3D constraints compared to established baselines such as specialized diffusion models. We consider 3D ligand generation conditioned on protein pockets together with ligand- and interaction-derived spatial constraints, including anchor fragments, pharmacophore points, and mandatory pocket-ligand interactions. To enable this evaluation, we introduce 3D-Fit - a token-efficient benchmarking strategy for assessing LLM performance on multi-conditioned spatial molecule generation. Our findings reveal a clear pattern in LLM spatial capabilities: while they still lag behind state-of-the-art approaches, they are promising and can handle multiple spatial constraints simultaneously, enabling scaling to heterogeneous setups.

  • Yield Smarter, Not Harder: Good Practices for Machine Learning of Reaction Outcomes https://doi.org/10.1021/jacs.6c02213 Reaction yield prediction is a longstanding challenge in synthetic chemistry, with broad implications for route planning, scalability, and high-throughput experimentation (HTE). While recent machine learning (ML) approaches have demonstrated promise in modeling reactivity, they often use complex descriptors or deep architectures that are computationally expensive and limit interpretability and scalability. Here, we assess how much information is stored in simpler descriptors and whether model accuracy is improved by increasing the complexity of the descriptors. Using classical ML models trained on descriptors with different complexity levels, we benchmark predictive performance on four publicly available HTE data sets covering three diverse reaction data sets: Buchwald–Hartwig (BH) amination, Suzuki–Miyaura (SM) coupling, and the silicon–amine protocol (SLAP). Our evaluation furthermore discusses (1) generalization via component-wise data splitting, (2) robustness through external validation across data sets, and (3) performance across asymmetric yield distributions characteristic of HTE data. Contrary to conventional expectations, we find that simpler models with interpretable features can achieve competitive performance under rigorous validation protocols. Based on our findings, we formulate good practices for future studies in this area. For example, comparison to low-cost baseline models should become a requirement for future ML studies for reaction-yield prediction. 📕Journal of the American Chemical Society (IF=16.6) #method

  • Data-Driven Insights into Ionic Conductivity in High-Dimensional Sodium Battery Electrolytes https://doi.org/10.1021/acsenergylett.6c01023 The discovery of advanced battery electrolytes is challenged by the vast compositional space of multi-component liquid formulations. Here, we introduce the ELectrolyte Laboratory for Integrated Experimentation (ELLIE), an automated platform that combines electrolyte formulation and impedance spectroscopy to map ionic conductivity across high-dimensional sodium electrolytes containing up to five salts and 15 solvents, generating an experimental dataset spanning nearly two orders of magnitude in conductivity. 23Na NMR, Raman spectroscopy, and viscosity measurements on a subset of electrolytes at a fixed salt concentration reveal that conductivity is jointly influenced by Na+ solvation strength, ion association, and solvent dynamics and positively correlates with inverse viscosity. Conductivity estimates based on the Nernst–Einstein relation captures broad concentration and viscosity relationships but do not extrapolate well across compositionally diverse electrolytes. Random forest modeling identifies lower solvent molecular weight as the dominant descriptor of high conductivity. Together, these results establish solvent molecular size as a physically interpretable descriptor of ion transport and demonstrate how automated experimentation can accelerate data-driven electrolyte optimization across complex compositional spaces. 📕ACS Energy Letters (IF=17.5) #method

  • 17 июл.7263512

    https://chemrxiv.org/doi/10.26434/chemrxiv.15006080/v1 Коллеги из 🏛ИОНХ РАН, 🏛ИНЭОС РАН и Университета Барселоны выложили препринт про BLIND (Bimodal Learning from Imperfect NMR Data) — трансформер, который переводит спектры ¹H и/или ¹³C ЯМР напрямую в молекулярную структуру (SMILES). Что делает работу интересной: Модель не получает ни брутто-формулу, ни набор возможных фрагментов, ни элементный состав — только то, как спектр записан в статье («7.85–7.81 (m, 3H)…»). Это гораздо честнее большинства предыдущих подходов. Учили на реальных, "грязных" данных. Стартовали с 7.5 млн (!!) записей спектров, извлечённых из литературы через базу OdanChem. После минимальной фильтрации (только CDCl₃ в качестве растворителя, удаление дубликатов и явных выбросов) осталось 5.5 млн спектров для 2.6 млн уникальных структур. Разбивка по уникальным структурам, строго без утечек между наборами. Обучающая выборка состояла из 1.82 млн уникальных SMILES, 1.89 млн ¹³C- и 2.0 млн ¹H-спектров. При этом избыточности почти нет — в среднем 1.2 спектра на молекулу, то есть модель учится обобщать буквально с одного измерения на соединение. Выборку намеренно не чистили до нейтральной органики, как это обычно бывает в хемоинформатике, и оставили редкие элементы (Se, Fe, Te…), соли и комплексы. Та же модель, обученная только на идеальном подмножестве (2 млн записей), даёт 48.4%, а обученная на всём хаосе — 56.7% на тех же тестах. Реальный шум дает разнообразие, которое помогает обобщать. 🔥Онлайн-версия доступна на https://odanchem.org/predict-multimodal-compound-search

  • Strategies for Identifying Molecules of Interest in Large Chemical Spaces 🔥 https://pubs.acs.org/doi/10.1021/acs.jcim.6c01496 Searching in ultralarge Chemical Spaces with known 2D similarity metrics, like fingerprint-based Tanimoto, substructure, or pharmacophore similarity searches, contains pitfalls due to the representation of molecules as synthons with connectivity rules. Applied to a set of almost 3000 drug-relevant queries we analyzed the ability of similarity search methods to retrieve analog compounds from Chemical Spaces, and how to best approach typical use cases in early phase drug discovery. Distinct characteristics of each similarity metric suggest orthogonal complementarity, enabling a versatile framework to diverse challenges present in hit discovery and lead expansion campaigns. 📕Journal of Chemical Information and Modeling (IF=6.4)

  • 11 июл.4293910

    И сегодня еще у нас на ChemRxiv вышел бенчмарк по растворимости для LLM, который мы пилили последние 2 месяца: Can LLMs Reason About Solubility? The SoluBench Benchmark for Pure and Mixed Solvent Systems: https://doi.org/10.26434/chemrxiv.15000632/v1 В…

  • In Silico ADMET: From Current Practices to Novel Profilers https://pubs.acs.org/doi/10.1021/acs.jmedchem.6c00049 We introduce OneADMET, a meticulously curated data set of 738,161 compounds with 1,119,719 measurements spanning 44 ADMET end points and 1 489 biological activities. We report a unified ChemProp-based MTL model capable of handling hundreds of continuous tasks simultaneously, which has practical advantages for model deployment and maintenance 📕Journal of Medicinal Chemistry (IF=7.3)

  • видео или голосовое, без подписи

  • 9 июл.410245

    Фантастика, конечно. Если бы мне кто-то в 2020 году сказал, что в 2026 можно будет одним запросом "Дай мне SMILES", получать SMILES всех комплексов с такой картинки (и без ошибок) с помощью модели, которую на это даже не файнтюнили — я бы не поверил. Про целесообразность, правда, отдельный вопрос. Потому что Opus 4.8 в low-режиме съел 6% пятичасового лимита токенов на это :)

  • видео или голосовое, без подписи

  • 9 июл.400151

    URSA: Chemistry-Aware Benchmark for Utilitarian Retrosynthesis Assessment 🔥 https://arxiv.org/abs/2607.04688 In the current study, we are introducing the URSA (Utilitarian RetroSynthesis Assessment) evaluation framework that provides the opportunity to benchmark the synthetic routes not only from a formal perspective, such as convergence to commercially available starting materials, but also from a chemical plausibility perspective, mimicking the way expert chemists evaluate the reactions and routes. The study covers a comprehensive evaluation of both conventional end-to-end retrosynthesis solutions and LLMs for the synthesis planning task on a set of novel, diverse target molecules with undisclosed synthetic routes, which represent realistic tasks in the daily drug design routine. We find that while LLMs can support high-level strategic planning, they currently underperform specialized retrosynthesis models in reliably solving synthesis planning tasks.

  • 1 июл.515146

    Impact of molecular multimodality on neural network models for prediction tasks related to drug discovery🔥 https://doi.org/10.1038/s41467-026-74487-x The number of unimodal molecule representation constantly increases, and researchers investigate how to combine them. Intuitively, multimodal representations may provide complementary information and combining them promises better performance. In this work, we systematically explore how combining multiple molecular modalities affects the performance of downstream prediction tasks, providing a baseline for informed decision making. Our study covers 7 molecular modalities and combines them using intermediate and late fusion, and 2 neural network architectures (with or without using knowledge graphs). We conduct experiments with 3 benchmarks for drug-target binding affinity, and 22 molecule property prediction. In total, we train and evaluate over 1400 models. In summary, our results show that combining multiple modalities improve the performance provided that effective fusion strategies are chosen. Knowledge-enhanced representation learning further boosts model performance. Notably, we find that even the use of simple late-fusion approaches establishes state-of-the-art results for some tasks. 📕 Nature Communications (IF=18.1)

  • видео или голосовое, без подписи

  • https://pubs.acs.org/doi/10.1021/acs.oprd.6c00057 Группа авторов из MIT выпустила обзор "Automation in Pharmaceutical Crystallization", посвященный инновациям в процессах автоматизации роста кристаллов. Фокус уделяется как разработке непосредственно алгоритмов и инструментов для проведения высокопроизводительных кристаллизаций, так и влиянию таких подходов на качество, воспроизводимость получаемых данных и перспективам разработки на их основе технологий искусственного интеллекта Также в тексте авторы отмечают, что предсказание растворимости почти всё сосредоточено на водных системах, тогда как для дизайна кристаллизации критична растворимость в неводных и смешанных растворителях — и как примеры датасетов, закрывающих этот пробел, приводят наши BigSolDB и MixtureSolDB :) 📕Organic Process Research & Development (IF=3.3)

Chem ML/AI/Datasets — tgindex