tgindex
C

Chemoinformatics papers

Статистика
@chemoinfo_papesанглийский

Papers and interesting materials on Chemoinformatics. Opinion is personal and belong to author of post.

Последний пост
10 мар.
Последнее чтение
15 авг.
Постов за неделю
0
Всего постов
25
Тип
открытый
Язык
английский
В каталоге с
13 авг.
Подписчики
980
+6 за 3 дн.
Сутки
+2
+0,20%
Неделя
 
Месяц
 
Просмотров на пост
962
24 постов
Вовлечённость
98,2%
к подписчикам
Постов в день
0,0
всего 25
Упоминаний
0
каналов
Охват размещения
оценка
1/24сутки в ленте
1/48двое суток
1/72трое суток

Оценка по просмотрам недавних постов: пост набирает почти всё за первые сутки.

Посты

  • 10 мар.8561625

    If you're teaching chemoinformatics or drug design, this could be of interest to you. UCL published tutorial and Jupyther notebooks on docking using SMINA. Bare minimal, but it looks like important information is present, including presentation. https://github.com/UCL/Open_Docking_Lab_Handbook

  • 5 мар.1 00332

    Raymond lab decided to go further after GDB-17 and decided to collect GDB-20! The size is obviously too large 32 trillion structures, so they sampled subset by “GenerativeAI”. As to me, even their GDB-17 was a big crazy idea however it helped to understand how weird are randomly generated structures. But GDB-20… looks to me as an artefact of the epoch went for good… but great that this time they provide access to 12 B subset. https://chemrxiv.org/doi/full/10.26434/chemrxiv.15000288/v1

  • 23 янв.1 27558

    Wendi Warr shared her free report from the 2025 CINF Herman Skolnik Award symposia celebrating contribution of Professor Matthias Rarey. Quite interesting reading mostly about structure-based drug design. #SBDD https://drive.google.com/file/d/1aZcHqy07mSQKaq7My-R8WVjrXD1I6YPg/view

  • 17 нояб.1 4801210

    Quite a nice and helpful open source tool - the Python code for finding pockets in proteins. Intsallable via pip and GitHub. #SBDD #bioinformatics #openscience #docking GitHub https://github.com/cch1999/pocketeer Article: https://bmcbioinformatics.biomedcentral.com/articles/10.1186/1471-2105-10-168

  • 21 окт.1 3221114

    I think this tool could be rather useful for those who work in the drug design and especially structure based drug design. PocketMaster is a flexible and automated tool for analyzing, clustering, and visualizing protein binding sites. Main Functionalities of PocketMaster ✅ Automatic structure alignment ✅ Flexible methods for defining binding sites ✅ Support for multiple RMSD methods ✅ RMSD calculation across all atoms or only Cα atoms within the binding sites ✅ Creation of clear visualizations ✅ Saving of aligned structures and analysis results ✅ Generation of summary reports with structural and sequence differences of binding sites ✅ Choice of clustering linkage methods ✅ Customizable clustering options Look useful, especially for me who does this very occasionally. #SBDD #drugdesign #bioinformatics #docking https://github.com/narek-abelyan/PocketMaster

  • Grzybowski work shows how they used rather cheap robot for doing quite fancy study of reactivity and catalyst design. Rather interesting reading as to me but more from the point of view what could be done, and how insights can be gathered. Some code was open-sourced too, which look rather new for Grzybowski lab, good direction to go! #robochemistry #chemicalspace https://www.nature.com/articles/s41586-025-09490-1

  • MIT work on prediction of solubility in mixture of solvents. No rocket science or fancy ML as to me, just a well done work. But the model and data available. https://www.nature.com/articles/s41467-025-62717-7

  • And adding to previous post. Kevin has just published (in September 2025, does he have time machine?) a paper on problems of testing LLMs. For the first time in my carrier I read it more like a scream from author's soul. Ok, there were articles like this on data reproducibility, but this one I take more personally, probably. #LLM https://www.sciencedirect.com/science/article/pii/S0927025625003842 BTW, it worth also reading his post in LinkedIN: "I think that most ML benchmarks might be measuring the wrong things entirely. We've been building evaluation tools for ML models in chemistry/materials science for a while now (ChemBench, MaCBench, MatText, etc.), and honestly: I am more and more worried about what we do as a field. You can take the exact same models, change how you aggregate scores or define your test set, and suddenly the "best" model is completely different. We showed this with ChemBench - depending on which metric you pick, the model rankings can flip around. In addition, many benchmarks are basically solved at this point: We've hit the noise floor of the underlying DFT calculations. Yet people keep using it and claiming "progress." A crucial insight is that we are somehow stuck in a datasets-as-benchmarks paradigm: we're using datasets to both define what we want to measure AND to do the measuring. It's like using a ruler to measure itself. We wrote this paper as therapy - trying to understand why evaluation feels so broken and what we might do about it. Turns out, most crucial evaluation design choices are just... hidden. All these little decisions that completely change your results, but nobody talks about them. Our small contribution: "evaluation cards" - basically forcing ourselves to document all the weird choices we made and why. " (from https://www.linkedin.com/posts/kevin-maik-jablonka_i-think-that-most-ml-benchmarks-might-be-activity-7357688705874075648-jJqA)

  • Startup Harmonic develops AI chatbot for math reasoning with the idea to develop "mathematical superintelligence" (https://www.techticia.com/2025/07/harmonic-launches-aristotle-ai-chatbot.html). It is interesting when will we come to LLM that is on par with human in chemistry reasoning? So, far works of Philippe Schwaller and Kevin Jablonka show that LLM struggle in reasoning in chemistry domain. But it is amazing, that we live in times, when asking questions like this does not mean that I know nothing about AI state of the art in chemistry (or I'm a journalist) 🤣

  • Interesting benchmark of different neural network potentials (NNPs) to predict protein-ligand interaction energy. They used NNPs trained on materials-science data (Orb-v3 and MACE-MP-0b2-L), specific models for predicting certain biological targets (Orb-v3) and six NNPs trained on molecular data (ANI-2x, AIMNet2, Egret-1, eSEN-OMol25-sm-conserving, UMA-s, and UMA-m). For comparison, semi-empirical DFT (GFN2-xTB and g-xTB) and polarizable force field was used (GFN-FF). Conclusion: 1. Among the NNPs, the models trained on OMol25 (eSEN-OMol25-sm-conserving, UMA-s, and UMA-m) are the best. 2. The materials-science models (Orb-v3 and MACE-MP-0b2-L) perform worst. 3. Semiempirical approaches (g-xTB) are the best of all. Also, they are superior from the performance point of view. 4. AIMNet2 shows quite good Spearman correlation and coefficient of determination (close to UMA and g-xTB) but is not that good in relative error. Probably need to be considered more carefully https://rowansci.com/blog/benchmarking-protein-ligand-interaction-energy

  • Really cool tool was released by Rarey group: the list of 40 000 (!) functional group SMARTS and corresponding software that gives a list of groups that present in a molecule. The application is run as backend service, which I don't really like but can be helpful in some applications. But what is great - that SMARTS and their labels are available in csv file of the GitHub. That's supercool thing. Paper: https://pubs.acs.org/doi/10.1021/acs.jcim.5c00599 GitHub: https://github.com/torbengutermuth/SmartChemist/

  • Amazing publication from Frank Noe and Microsoft Research team: generative model that predicts ensemble of peptide conformations. Basically, it is generative model that returns MD results at the costs of an hour. What a time we live in! Abstract: Following the sequence and structure revolutions, predicting functionally relevant protein structure changes at scale remains an outstanding challenge. We introduce BioEmu, a deep learning system that emulates protein equilibrium ensembles by generating thousands of statistically independent structures per hour on a single GPU. BioEmu integrates over 200 milliseconds of molecular dynamics (MD) simulations, static structures and experimental protein stabilities using novel training algorithms. It captures diverse functional motions—including cryptic pocket formation, local unfolding, and domain rearrangements—and predicts relative free energies with 1 kcal/mol accuracy compared to millisecond-scale MD and experimental data. BioEmu provides mechanistic insights by jointly modelling structural ensembles and thermodynamic properties. This approach amortizes the cost of MD and experimental data generation, demonstrating a scalable path toward understanding and designing protein function. #SBDD #moldyn #moleculardynamics #deeplearning https://www.science.org/doi/10.1126/science.adv9817

  • Humongous work was done by Kevin Jablonka's team to write review or maybe rather short textbook on training LLM models that can comprehend chemistry. They call them "general purpose models" which I find better describing them than simply "large language models in chemistry". They describe techniques for training, fine-tuning such models, and applications such models already found. The only challenge I see now with it - the space is changing daily. But I find this text extremely helpful. One of my must-read document. #LLM Here is online clickable book: https://gpmbook.lamalab.org/ At this is the PDF: https://arxiv.org/pdf/2507.07456

  • Interesting stuff - multi-agent system for writing literature review. And this time it is available in pip and as a source code! Worth trying to write a review for your article 😊 #LLM #agenticLLM https://github.com/stanford-oval/storm/

  • Interesting publication on federated learning in chemistry from Aachen, with special attention to chemical engineering. I am really interested in this topic, and I believe there is a future in this technology, however due to legal issues it cannot be widely adopted. One of cases where legal constraints and lack of trust kill the development of technology and drug development. I had a very interesting discussion with one of thought leaders from big pharma about federated learning and, according to them, one of the reason for low adoption is complexity to guarantee balance of interests of all parties: someone add more data, someone may train on low data or even fake data and benefit from others. Nonetheless, the topic itself is super interesting. #QSAR #federatedlearning https://arxiv.org/abs/2506.18525

  • Guys from a company that I never heard of developed MCP server for RDKit, so one can access some simple RDKit functions by natural text queries to LLM. Personally, rather useless stuff as such, but could be a good starting point to add the tools that one wants to use in their projects. The code seem to be rather self-explanatory, rather simple to reuse and adapt. #LLM #agenticLLM https://github.com/tandemai-inc/rdkit-mcp-server/tree/master

  • Oles Isaev team developed a new version of his ANI-type network for modelling reaction pathways. Also, some transition state search algorithm was implemented. #NNforQC #deeplearning #reaction https://chemrxiv.org/engage/chemrxiv/article-details/685505c9c1cb1ecda0f701de

  • видео или голосовое, без подписи

  • видео или голосовое, без подписи

  • видео или голосовое, без подписи

Chemoinformatics papers — tgindex