Scoring Functions & Post-docking Refinement

Rank Poses with Confidence. Refine Affinity Predictions. Reduce False Positives.
ML-Based Scoring Consensus Rescoring Post-Docking Free Energy Refinement

Docking generates poses; scoring separates signal from noise. We develop target-specific scoring functions and run consensus rescoring pipelines to improve rank-order correlation with measured affinity.

Why Scoring Functions & Post-docking Refinement Are the Critical Bridge Between Pose and Affinity?

Docking samples billions of conformations, but default force fields mis-rank 30–50% of true binders. For seed-stage biotechs, false positives burn synthesis budgets. For pharma teams, false negatives discard viable scaffolds. Our platform replaces generic scores with ML-calibrated, target-specific scoring and physics-based free energy refinement.

What Sets the Platform Apart

Target-Specific ML Scoring

ML models trained on ITC and BLI data for your target. R² > 0.75. ADMET Prediction & Modeling filters liabilities pre-ranking.

Consensus Rescoring

Multi-engine consensus (Vina + Glide + GNINA + ML) reduces false positives. Benchmarked on internal validation sets.

The Scoring & Refinement Suite

ML Scoring Function Development

Target-Specific Affinity Prediction Beyond Generic Force Fields

Graph neural network model mapping protein-ligand interaction features for affinity prediction.

Key Features:

  • Graph Neural Network & Transformer Models — GCN and attention-based architectures trained on target-specific SPR and ITC datasets to learn non-linear binding patterns.
  • Feature Engineering — Interaction fingerprints, pharmacophore matching, and desolvation terms combined with structural descriptors.
  • Ideal For — Targets where generic scoring fails (kinases, proteases, PPIs); Hit Identification enrichment; Lead Optimization analog ranking.

For virtual biotechs screening 100,000+ compounds, target-specific ML scoring improves enrichment factors 3–5× over default Vina scores. For pharma teams, models are retrained with each project's experimental data, improving accuracy across congeneric series.

Consensus Scoring & Ranking

Multi-Engine Voting to Eliminate False Positives

Multi-engine consensus scoring overlay on a single protein-ligand docking pose.

Key Features:

  • Cross-Engine Consensus — Poses rescored with Vina, Glide, GNINA CNN, and custom ML functions; outliers flagged and consensus rank calculated.
  • Enrichment Benchmarking — ROC-AUC and early enrichment factor (EF1%) calculated against known actives before prospective screening.
  • Ideal For — Large-scale Structure-Based Virtual Screening (SBVS); Fragment-based Screening (FBS) hit triage; prioritization before synthesis.

No single scoring function is perfect. Consensus scoring leverages the orthogonal strengths of physics-based and ML engines to surface true binders that any single method misses.

Post-Docking Refinement

Pose Stability and Affinity Validation

Explicit water network and energy minimization within a refined binding pocket.

Key Features:

  • MM/PBSA & MM/GBSA Rescoring — Post-docking binding free energy estimation for top 50–500 poses with explicit solvent and entropy corrections.
  • FEP/TI Calculations — Alchemical free energy perturbation for congeneric series in Lead Optimization.
  • MD Trajectory Analysis — RMSD, RMSF, and hydrogen-bond lifetime monitoring to discard unstable poses.

For Covalent Docking campaigns, post-docking refinement confirms warhead stability. For Protein-Protein/Peptide Docking (Rigid Body / Flexible Docking), MM/PBSA rescoring improves interface ranking correlation with measured Kd.

Platform Instrumentation

Software / System Core Capability
GNINA 1.0 CNN rescoring with PDBbind-trained models; gradient-based pose optimization.
Schrödinger Glide + Prime MM-GBSA rescoring and induced-fit refinement with OPLS4 force field.
AutoDock Vina 1.2.0 Batch-mode pose scoring and flexible side-chain evaluation.
GROMACS 2023 + AMBER 22 All-atom MD for pose stability and Binding Free Energy Calculation (FEP/TI, MM/PBSA).
OpenEye SZYBKI + FF MM-PBSA and solvation free energy calculations with PBSA solver.
RDKit + scikit-learn Custom ML scoring function development with Morgan fingerprints and graph descriptors.
NVIDIA A100 GPU Cluster Parallelized consensus scoring and ML inference on million-pose datasets.
PyMOL + Maestro Interaction fingerprint analysis, pose comparison, and enrichment visualization.

Standardized Workflow

Project Workflow

A standardized, milestone-driven execution system. From raw docking poses to validated, rescored hit lists—managed by a single computational project team, tracked in real time.

01 Docking Pose Review & Quality Filter Week 1
02 Scoring Function Selection & Calibration Week 1
03 Consensus Rescoring Execution Weeks 2–3
04 Post-Docking Refinement Weeks 3–4
05 Report & Handoff Week 4–5

01 Docking Pose Review & Quality Filter

  • Raw pose import from docking engines (Vina, Glide, etc.).
  • Quality filtering: RMSD clustering, strain energy, and clash detection.
  • Interaction fingerprint extraction.

Deliverable: Filtered pose dataset + quality report.

02 Scoring Function Selection & Calibration

  • Scoring function selection: generic, target-specific ML, or hybrid.
  • Retrospective benchmarking against known actives (ROC-AUC, EF1%).
  • Model calibration with project-specific experimental data.

Deliverable: Scoring function benchmark report + calibration metrics.

03 Consensus Rescoring Execution

  • Multi-engine rescoring and consensus rank calculation.
  • Outlier detection and false-positive flagging.
  • Consensus score normalization and rank aggregation.

Deliverable: Consensus-ranked pose list with outlier flags.

04 Post-Docking Refinement

  • MM/PBSA or FEP/TI rescoring for top 50–500 poses.
  • Molecular Dynamics (MD) Simulations for pose stability (RMSD, RMSF).
  • Hydrogen-bond lifetime and water network analysis.

Deliverable: Refined pose dataset with free energy and stability metrics.

05 Report & Handoff

Deliverable: Final technical report + electronic data package + transition plan to Hit to Lead or Lead Optimization.

Sample Requirements

Requirement Details
Docking poses SDF or PDB format from Vina, Glide, or other engines; include receptor structure
Known actives / decoys Confirmed binders and inactive compounds for scoring calibration and benchmarking
Experimental affinity data IC50, Kd, or pIC50 for ML model training and validation (if available)
Project scope Hit identification enrichment, lead optimization ranking, or pose validation
Prior scoring data Any existing scoring results or known false positives to guide refinement strategy

Standard Deliverables

  • Filtered and clustered pose dataset with quality metrics
  • Target-specific ML scoring function or consensus ranking table
  • Retrospective enrichment benchmarking report (ROC-AUC, EF1%)
  • MM/PBSA or FEP/TI free energy estimates for top poses (if contracted)
  • Pose stability report via Molecular Dynamics (MD) Simulations (if contracted)
  • Electronic data package formatted for Structure-Based Virtual Screening (SBVS) or Hit to Lead handoff

Frequently Asked Questions

Case Study

Case Study: Graph Convolutional Networks Improve Target-Specific Scoring for cGAS and KRAS

Published Evidence:
Li S, Li W, Li X, et al. Graph convolutional neural networks improved target-specific scoring functions for cGAS and kRAS in virtual screening. Comput Struct Biotechnol J. 2025;27:2176–2185. (Open Access)

Key Findings:

  • Target-Specific Advantage: GCN-based scoring functions outperformed generic docking scores (Vina, Glide) for both cGAS and kRAS, with improved ROC-AUC and early enrichment.
  • Extrapolation Performance: Graph convolutional networks generalized to chemically distinct test sets better than traditional ML models, capturing complex protein-ligand binding patterns.
  • Industrial Applicability: Target-specific functions demonstrated robustness in discriminating active from inactive molecules, with significant potential for structure-based virtual screening campaigns.

Industrial Translation:
For seed-stage biotechs, this paradigm confirms that target-specific ML scoring improves hit rates without requiring proprietary multi-target datasets. For pharma teams, GCN-based functions integrate into existing docking pipelines as a rescoring layer, converting generic docking output into target-enriched hit lists. Our platform operationalizes this peer-reviewed approach within an audit-ready workflow, pairing GCN model development with consensus rescoring and Binding Free Energy Calculation (FEP/TI, MM/PBSA) validation.

Comparison of NEF1% for different scoring functions on cGAS and kRAS datasets.

Figure 1. Comparison of NEF1 % for different scoring functions on cGAS and kRAS datasets. (Li S, et al., 2025)

Reference

  1. Li S, Li W, Li X, et al. Graph convolutional neural networks improved target-specific scoring functions for cGAS and kRAS in virtual screening. Comput Struct Biotechnol J. 2025;27:2176–2185.

Need validated scoring and refinement for your virtual screening pipeline? Our team can develop a target-specific scoring strategy or consensus rescoring protocol tailored to your target, docking poses, and experimental data. Contact our scientific team today to start your project.