Skip to content

Author

Roberto Uribe-Paredes

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Aug 2026

Data-Centric Evaluation of Protein Function Prediction Pipelines

Performance estimates in protein function prediction depend not only on model choice but also on upstream decisions that define the learning problem. Using antioxidant protein classification as a controlled case study, we evaluated how dataset harmonisation, protein representation, redundancy control, and partitioning strategy affect protein machine learning pipelines. We integrated 18,804 records from 12 publicly available dataset entries into a curated consensus dataset of 4,193 protein sequences. One-hot encoding and six pretrained protein language model representations were evaluated as model inputs and as similarity spaces for redundancy reduction and distance-aware splitting. Representation choice substantially altered dataset geometry, retained dataset size, class balance, and downstream evaluation. At representation-specific p90 thresholds, one-hot encoding retained the complete dataset, whereas pretrained embeddings retained between 5% and 25% of sequences. Distance-aware partitioning reduced apparent performance relative to random splitting by up to 0.15 MCC before redundancy control, while this difference narrowed after similarity filtering. Selected configurations nevertheless maintained high performance under stricter evaluation, reaching an MCC of 0.84. These findings show that performance estimates should be interpreted as outcomes of complete data-centric workflows rather than isolated properties of predictive models.

Nicole Soto-García, Norma Murillo-Acevedo, Julián García-Vinuesa et al. · 0 citations
Open access Jul 2026

Toward explainable and adaptive artificial intelligence systems for antimicrobial peptide discovery under resistance pressure

Therapeutic peptides have emerged as promising candidates for combating antimicrobial resistance, particularly against multidrug-resistant pathogens for which conventional antibiotics are becoming increasingly ineffective. Although artificial intelligence has accelerated antimicrobial peptide discovery through predictive modelling, generative design, and large-scale in silico screening, many current workflows remain fragmented, weakly interpretable, and only loosely connected to experimental feedback. In this Perspective, we propose a methodological framework for resistance-aware antimicrobial peptide discovery built upon three complementary principles: explainable and uncertainty-aware prediction, biologically constrained and rule-guided generative design, and iterative design–test–learn workflows capable of continuously incorporating new evidence. By organising existing methodologies into adaptive and transparent discovery systems, the proposed framework supports candidate prioritisation, optimisation, and iterative refinement under evolving resistance pressures. To demonstrate these concepts in practice, we present a workflow integrating interpretable prediction, uncertainty-aware prioritisation, rule extraction, and adaptive model updating. Collectively, these elements provide a roadmap for advancing antimicrobial peptide discovery from isolated predictive tasks toward integrated and evidence-driven discovery ecosystems.

Ana Luisa Islas-Avila, Julián García-Vinuesa, Nicole Soto-García et al. · 0 citations