arXivpreprint
Noel Suarez-Barro · Manuel Lama · Juan C. Vidal
Drug discovery is a costly and high-risk process, where toxicity-related failures remain a major cause of attrition in both preclinical and clinical stages. As a result, accurate early prediction of chemical toxicity is essential to reduce downstream costs and improve compound prioritization. In this context, graph deep learning (GDL) has emerged as a powerful paradigm for toxicity prediction, leveraging molecular graph representations to learn directly from chemical structure with improved expressivity over traditional approaches. Despite the growing number of proposed models, current literature-based comparisons are often difficult to interpret due to inconsistencies in datasets, preprocessing pipelines, and evaluation protocols. To address this limitation, we introduce a unified and standardized benchmarking framework for GDL-based toxicity prediction. We systematically evaluate more than 20 representative approaches under consistent experimental conditions and across multiple datasets and partitioning strategies, enabling a fair and reproducible comparison of model performance. In addition, we complement this empirical study with a structured literature analysis to contextualize existing methodological trends and performance claims. Our results provide a clearer and more reliable assessment of the current state of the field, highlighting both the strengths and limitations of existing graph-based approaches. To support transparency and reproducibility, we release our benchmarking framework as open-source software https://gitlab.citius.gal/noel.suarez/benchtox, allowing the community to evaluate and compare models under consistent conditions.
arXivpreprint
Emma Granqvist · Rocío Mercado · Samuel Genheden
Agentic large language model (LLM) systems are reshaping scientific workflows in chemistry and drug discovery, but evaluating their open-ended, tool-augmented outputs remains a fundamental bottleneck. The LLM-as-a-Judge paradigm has emerged as a scalable alternative, but existing drug discovery benchmarks deploy LLM judges without validating their alignment with human experts. In this work, we present an LLM-as-a-Judge evaluation framework for ChatInvent, an agentic drug discovery assistant deployed at AstraZeneca, with five contributions. First, we define four output-quality evaluation dimensions---Completeness, Relevancy, Structural Clarity, and Scope Adherence---alongside deterministic Tool Call Correctness checks. Second, we validate the judge through a human alignment study with five expert annotators, comparing Gemini 3.1 Pro, Claude Opus 4.7, GPT-5, and Llama 3.1 70B as candidate judges. Third, we optimize the best-performing judge using few-shot demonstrations of human-annotated examples, improving alignment with the human majority vote from 0.80 to 0.86. Fourth, applying the optimized judge to 70 held-out questions, we surface concrete limitations and find no strong evidence that informal phrasing degrades output quality; it may, however, still be helpful to have the LLM rewrite the original question before querying the agent. Finally, we extend the framework to 38 adversarial questions that are ambiguous, invalid, out-of-scope or ethically sensitive, and show that the agent's refusal behavior is guided by the stated intent of a request. Our framework provides a reusable template for human-aligned evaluation of agentic systems in scientific domains.
OpenAlexarticle
Huazhang Ying · Yipin Lei · Yihai Luo +2 authors
Published in Cell Biomaterials. Open the paper details to explore the original source.
Chemical Synthesis and AnalysisComputational Drug Discovery Methodsvaccines and immunoinformatics approaches
View paper →
OpenAlexarticle
Mohamed-Amine Chadi · Amzil Asmaa · Hajar Ezzahoud +1 authors
Published in Frontiers in Pharmacology. Open the paper details to explore the original source.
Computational Drug Discovery MethodsMachine Learning in Materials ScienceProtein Degradation and Inhibitors
View paper →
OpenAlexpreprint
Sourajyoti Goswami · Puspita Roy · Shouvik Kumar Nandy
Published in Research Square. Open the paper details to explore the original source.
Computational Drug Discovery MethodsMicrotubule and mitosis dynamicsProtein Degradation and Inhibitors
View paper →
OpenAlexarticle
Sukanta Kumar Satapathy · Ganesh Patro · Santanu Shaw
Published in Recent Advances in Drug Delivery and Formulation. Open the paper details to explore the original source.
Computational Drug Discovery MethodsPharmaceutical Quality and CounterfeitingPharmacovigilance and Adverse Drug Reactions
View paper →
OpenAlexarticle
Abubakar Sadiq Musa · Muhammad Lawan Wasaram · Kashim Ibrahim MUHAMMAD +1 authors
Published in International Journal of Applied Sciences and Biotechnology. Open the paper details to explore the original source.
Artificial Intelligence in Healthcare and EducationProstate Cancer Diagnosis and TreatmentProstate Cancer Treatment and Research
View paper →
arXivpreprint
Marvellous O. Ajala · Zainab Ashimiyu-Abdusalam · Comfort Adesina
We introduce Malaria-Instruct, a curated instruction-following dataset derived from the ChEMBL Legacy Malaria corpus for Malaria virtual screening, and conduct a systematic evaluation of five open-source LLMs; Gemma-2 2B/9B, TxGemma-2B/9B, and LlaSMol-Mistral-7B, on a rigorous out-of-distribution data split. Performance was benchmarked against classical ML models (Random Forest, XGBoost) and frontier proprietary models (Gemini 2.5, OpenAI o3) under few-shot conditions. Fine-tuned LLMs substantially outperformed all baselines: TxGemma-9B achieved the highest ROC-AUC ($0.731 \pm 0.005$) and LlaSMol-Mistral-7B the best enrichment factor (EF@1\% $\approx$ 4.99). Domain-specific fine-tuning proved categorically indispensable with TxGemma-9B collapsing from ROC-AUC 0.731 to 0.499, under its best few-shot condition, and neither Gemini 2.5 (ROC-AUC $\approx$ 0.53) nor o3 (ROC-AUC $\approx$ 0.59) achieved reliable discrimination without fine-tuning. Biomedical pretraining conferred a measurable advantage at equivalent scale, while chemistry-aware pretraining yielded superior prospective enrichment. Fine-tuned open-source LLMs represent a compelling, resource-efficient paradigm for antimalarial VS, outperforming both classical pipelines and proprietary reasoning models under structurally challenging conditions.
OpenAlexarticle
Adel Alhowyan · Ahmad J. Obaidullah · Wael Ali Mahdi
Published in Frontiers in Medicine. Open the paper details to explore the original source.
Computational Drug Discovery MethodsMachine Learning in Materials SciencePhase Equilibria and Thermodynamics
View paper →
OpenAlexarticle
Qianran Sun · Xiaocong Pang · Yonghong Liu
Published in Pharmaceutics. Open the paper details to explore the original source.
Bioinformatics and Genomic NetworksComputational Drug Discovery MethodsMelanoma and MAPK Pathways
View paper →
arXivpreprint
Xuan Lin · Jingyu Sheng · Tengfei Ma +2 authors
Phenotypic drug discovery enables the discovery of functional relationships between molecular structures and cellular responses. However, existing multimodal representation learning methods often optimize cross-modal alignment without considering the intrinsic organization of chemical space, resulting in distorted molecular representations and loss of structural information. We propose \textbf{PhenMol}, a structure-preserving framework for phenotype-aware molecular representation learning. PhenMol disentangles molecular and cellular representations into shared and private components, enabling phenotype-guided alignment while preserving chemical structures through a dedicated molecular branch. This design integrates cellular phenotype information without disrupting molecular neighborhood organization. Experiments on approximately $3.04 \times 10^{4}$ molecule--cell morphology pairs demonstrate that PhenMol improves molecular property prediction across 270 bioactivity tasks, molecule--phenotype retrieval, and clinical trial outcome prediction. Moreover, ECFP4-based structural analysis shows that PhenMol better preserves molecular neighborhoods and reduces embedding distortion compared with existing multimodal alignment methods. These results highlight the importance of structure-aware constraints in multimodal molecular representation learning and provide an effective approach for integrating cellular phenotypes with chemical knowledge for drug discovery.
arXivpreprint
Thanina Hamitouch · Khadidja Henni · Abdelkrim Arie +3 authors
Understanding how drugs interact with protein targets is fundamental to drug discovery, drug repurposing and the early identification of promising therapeutic candidates before costly experimental testing. Sequence-based DTI models face three practical limitations: labelled interactions are scarce and unevenly distributed, large pretrained chemical and protein encoders are expensive to fine-tune end-to-end, and independently encoded sequences do not capture pair-specific dependencies. We present BERT4DTI, which encodes SMILES strings with ChemBERTa and amino-acid sequences with ProtBERT, applies bidirectional mutual attention between token-level representations, and classifies the resulting interaction features using convolutional layers and a multilayer perceptron. To reduce trainable size, ProtBERT is truncated to 18 retained layers and only the last two layers of each encoder are fine-tuned. On BIOSNAP, DAVIS and BindingDB, BERT4DTI is competitive, achieving the best ROC-AUC and PR-AUC on BIOSNAP and the highest sensitivity on all three benchmarks. An ablation on DAVIS shows that mutual attention improves PR-AUC and specificity. With 125M trainable parameters compared with 353M for full BERT fine-tuning, BERT4DTI provides a favourable performance-parameter trade-off for sequence-based DTI screening, while leaving runtime profiling, calibration and leakage-audited validation for future work.