Coverage Holes in Historical Chemical Repositories
Scientific humility requires exposing data absence. ChEMBL is historically optimized for approved oral small molecules and kinase/GPCR campaigns. Peptides, natural extracts, and gray-market compounds represent systematic coverage holes.
Scientific Status: Preclinical tissue-healing literature in rodents (E1); marketed heavily in gray-market wellness communities without mammalian lifespan proof (E0).
Scientific Status: Telomerase elongation claims in cultured human fibroblasts (E1); widely sold as an anti-aging elixir without replicated ITP lifespan validation (E0).
Scientific Status: Unresolved (404); correctly flagged as unmapped.
Benchmark Model Comparison — Scaffold Split vs Random Split
Evaluating L2-regularized Logistic Regression (Baseline) vs HistGradientBoosting (Contender) across 3 seeds (42, 123, 456) on 2048-bit Morgan circular fingerprints + standardized RDKit 2D descriptors on ChEMBL mTOR kinase activity (N = 561 molecules).
| Split Strategy | Model Pipeline | AUROC | AUPRC | Recall @ 5% FPR | Brier Score |
|---|---|---|---|---|---|
| Bemis-Murcko Scaffold (Honest) | Baseline (Logistic Regression) | 0.9735 ± 0.0203 | 0.9929 ± 0.0057 | 0.8496 ± 0.1207 | 0.0486 ± 0.0238 |
| Bemis-Murcko Scaffold (Honest) | Contender (HistGradientBoosting) | 0.9434 ± 0.0294 | 0.9855 ± 0.0095 | 0.7187 ± 0.1550 | 0.0820 ± 0.0392 |
| Stratified Random (Diagnostic) | Baseline (Logistic Regression) | 0.8912 ± 0.0482 | 0.9780 ± 0.0106 | 0.6528 ± 0.0856 | 0.0922 ± 0.0223 |
| Stratified Random (Diagnostic) | Contender (HistGradientBoosting) | 0.8766 ± 0.0467 | 0.9748 ± 0.0117 | 0.6667 ± 0.1623 | 0.1159 ± 0.0147 |
Key Takeaways for Hiring Managers & Reviewers:
- Linear Baseline Generalization: Regularized linear models (Logistic Regression) outperform non-linear tree ensembles on sparse circular fingerprint representations ($p = 2057$), avoiding overfitting to dominant scaffold clusters.
- Zero Scaffold Leakage: Enforcing Bemis-Murcko clustering ensures test compounds share zero scaffold topology with training compounds (Train ∩ Test = ∅).
Top Highest-Confidence False Positive Predictions
Compounds with inactive ground truth (pChEMBL < 6.0) predicted as active with highest confidence under scaffold split.
| ChEMBL ID | Predicted Prob | True pChEMBL | Murcko Scaffold | Chemical & Pharmacophore Analysis |
|---|
Constrained Molecular Generator (Genetic Algorithm)
Optimizing the frozen Phase 3 ChEMBL mTOR surrogate model while evaluating QED as an adjustable bias control (λQED) and penalizing assay interference (PAINS).
QED Bias Sensitivity Analysis (λQED)
Evaluating how the molecular generator behaves when historical small-molecule oral drug-likeness priors are removed vs enforced.
| QED Weight (λQED) | Objective Character | Mean mTOR Prob | Mean QED | Internal Diversity | Novelty Rate |
|---|---|---|---|---|---|
| λ = 0.0 | Pure Target Affinity (Unconstrained) | 1.000 | 0.059 | 0.563 | 100.0% |
| λ = 0.2 | Balanced Tradeoff (Primary) | 0.991 | 0.754 | 0.655 | 99.5% |
| λ = 0.5 | Oral Drug-Likeness Constrained | 0.985 | 0.635 | 0.670 | 96.5% |
Chemical Space Takeaway (PCA Embedding):
Principal Component projection (PC1: 13.3%, PC2: 9.1% variance) demonstrates that GA-evolved candidates explore the mTOR training active chemotype manifold adjacent to benchmark binders, rather than drifting into unphysical junk space. When unconstrained (λQED = 0.0), molecules drift toward high molecular weight macrocyclic configurations (mean QED 0.059), demonstrating the restrictive nature of historical small-molecule heuristics.