Scoring Functions & Post-docking Refinement
Docking generates poses; scoring separates signal from noise. We develop target-specific scoring functions and run consensus rescoring pipelines to improve rank-order correlation with measured affinity.
Why Scoring Functions & Post-docking Refinement Are the Critical Bridge Between Pose and Affinity?
Docking samples billions of conformations, but default force fields mis-rank 30–50% of true binders. For seed-stage biotechs, false positives burn synthesis budgets. For pharma teams, false negatives discard viable scaffolds. Our platform replaces generic scores with ML-calibrated, target-specific scoring and physics-based free energy refinement.
What Sets the Platform Apart
Target-Specific ML Scoring
ML models trained on ITC and BLI data for your target. R² > 0.75. ADMET Prediction & Modeling filters liabilities pre-ranking.
Consensus Rescoring
Multi-engine consensus (Vina + Glide + GNINA + ML) reduces false positives. Benchmarked on internal validation sets.
Physics-Based Refinement
Binding Free Energy Calculation (FEP/TI, MM/PBSA) and Molecular Dynamics (MD) Simulations refine top poses.
The Scoring & Refinement Suite
ML Scoring Function Development
Target-Specific Affinity Prediction Beyond Generic Force Fields

Key Features:
- Graph Neural Network & Transformer Models — GCN and attention-based architectures trained on target-specific SPR and ITC datasets to learn non-linear binding patterns.
- Feature Engineering — Interaction fingerprints, pharmacophore matching, and desolvation terms combined with structural descriptors.
- Ideal For — Targets where generic scoring fails (kinases, proteases, PPIs); Hit Identification enrichment; Lead Optimization analog ranking.
For virtual biotechs screening 100,000+ compounds, target-specific ML scoring improves enrichment factors 3–5× over default Vina scores. For pharma teams, models are retrained with each project's experimental data, improving accuracy across congeneric series.
Consensus Scoring & Ranking
Multi-Engine Voting to Eliminate False Positives

Key Features:
- Cross-Engine Consensus — Poses rescored with Vina, Glide, GNINA CNN, and custom ML functions; outliers flagged and consensus rank calculated.
- Enrichment Benchmarking — ROC-AUC and early enrichment factor (EF1%) calculated against known actives before prospective screening.
- Ideal For — Large-scale Structure-Based Virtual Screening (SBVS); Fragment-based Screening (FBS) hit triage; prioritization before synthesis.
No single scoring function is perfect. Consensus scoring leverages the orthogonal strengths of physics-based and ML engines to surface true binders that any single method misses.
Post-Docking Refinement
Pose Stability and Affinity Validation

Key Features:
- MM/PBSA & MM/GBSA Rescoring — Post-docking binding free energy estimation for top 50–500 poses with explicit solvent and entropy corrections.
- FEP/TI Calculations — Alchemical free energy perturbation for congeneric series in Lead Optimization.
- MD Trajectory Analysis — RMSD, RMSF, and hydrogen-bond lifetime monitoring to discard unstable poses.
For Covalent Docking campaigns, post-docking refinement confirms warhead stability. For Protein-Protein/Peptide Docking (Rigid Body / Flexible Docking), MM/PBSA rescoring improves interface ranking correlation with measured Kd.
Platform Instrumentation
| Software / System | Core Capability |
|---|---|
| GNINA 1.0 | CNN rescoring with PDBbind-trained models; gradient-based pose optimization. |
| Schrödinger Glide + Prime | MM-GBSA rescoring and induced-fit refinement with OPLS4 force field. |
| AutoDock Vina 1.2.0 | Batch-mode pose scoring and flexible side-chain evaluation. |
| GROMACS 2023 + AMBER 22 | All-atom MD for pose stability and Binding Free Energy Calculation (FEP/TI, MM/PBSA). |
| OpenEye SZYBKI + FF | MM-PBSA and solvation free energy calculations with PBSA solver. |
| RDKit + scikit-learn | Custom ML scoring function development with Morgan fingerprints and graph descriptors. |
| NVIDIA A100 GPU Cluster | Parallelized consensus scoring and ML inference on million-pose datasets. |
| PyMOL + Maestro | Interaction fingerprint analysis, pose comparison, and enrichment visualization. |
Standardized Workflow
Project Workflow
A standardized, milestone-driven execution system. From raw docking poses to validated, rescored hit lists—managed by a single computational project team, tracked in real time.
01 Docking Pose Review & Quality Filter
- Raw pose import from docking engines (Vina, Glide, etc.).
- Quality filtering: RMSD clustering, strain energy, and clash detection.
- Interaction fingerprint extraction.
Deliverable: Filtered pose dataset + quality report.
02 Scoring Function Selection & Calibration
- Scoring function selection: generic, target-specific ML, or hybrid.
- Retrospective benchmarking against known actives (ROC-AUC, EF1%).
- Model calibration with project-specific experimental data.
Deliverable: Scoring function benchmark report + calibration metrics.
03 Consensus Rescoring Execution
- Multi-engine rescoring and consensus rank calculation.
- Outlier detection and false-positive flagging.
- Consensus score normalization and rank aggregation.
Deliverable: Consensus-ranked pose list with outlier flags.
04 Post-Docking Refinement
- MM/PBSA or FEP/TI rescoring for top 50–500 poses.
- Molecular Dynamics (MD) Simulations for pose stability (RMSD, RMSF).
- Hydrogen-bond lifetime and water network analysis.
Deliverable: Refined pose dataset with free energy and stability metrics.
05 Report & Handoff
- Final ranked hit list with confidence scores.
- Enrichment metrics and synthetic priority recommendations.
- Direct handoff to Hit Biophysical Characterization or Co-crystallization if contracted.
Deliverable: Final technical report + electronic data package + transition plan to Hit to Lead or Lead Optimization.
Sample Requirements
| Requirement | Details |
|---|---|
| Docking poses | SDF or PDB format from Vina, Glide, or other engines; include receptor structure |
| Known actives / decoys | Confirmed binders and inactive compounds for scoring calibration and benchmarking |
| Experimental affinity data | IC50, Kd, or pIC50 for ML model training and validation (if available) |
| Project scope | Hit identification enrichment, lead optimization ranking, or pose validation |
| Prior scoring data | Any existing scoring results or known false positives to guide refinement strategy |
Standard Deliverables
- Filtered and clustered pose dataset with quality metrics
- Target-specific ML scoring function or consensus ranking table
- Retrospective enrichment benchmarking report (ROC-AUC, EF1%)
- MM/PBSA or FEP/TI free energy estimates for top poses (if contracted)
- Pose stability report via Molecular Dynamics (MD) Simulations (if contracted)
- Electronic data package formatted for Structure-Based Virtual Screening (SBVS) or Hit to Lead handoff
Frequently Asked Questions
Case Study
Case Study: Graph Convolutional Networks Improve Target-Specific Scoring for cGAS and KRAS
Published Evidence:
Li S, Li W, Li X, et al. Graph convolutional neural networks improved target-specific scoring functions for cGAS and kRAS in virtual screening. Comput Struct Biotechnol J. 2025;27:2176–2185. (Open Access)
Key Findings:
- Target-Specific Advantage: GCN-based scoring functions outperformed generic docking scores (Vina, Glide) for both cGAS and kRAS, with improved ROC-AUC and early enrichment.
- Extrapolation Performance: Graph convolutional networks generalized to chemically distinct test sets better than traditional ML models, capturing complex protein-ligand binding patterns.
- Industrial Applicability: Target-specific functions demonstrated robustness in discriminating active from inactive molecules, with significant potential for structure-based virtual screening campaigns.
Industrial Translation:
For seed-stage biotechs, this paradigm confirms that target-specific ML scoring improves hit rates without requiring proprietary multi-target datasets. For pharma teams, GCN-based functions integrate into existing docking pipelines as a rescoring layer, converting generic docking output into target-enriched hit lists. Our platform operationalizes this peer-reviewed approach within an audit-ready workflow, pairing GCN model development with consensus rescoring and Binding Free Energy Calculation (FEP/TI, MM/PBSA) validation.

Figure 1. Comparison of NEF1 % for different scoring functions on cGAS and kRAS datasets. (Li S, et al., 2025)
Reference
- Li S, Li W, Li X, et al. Graph convolutional neural networks improved target-specific scoring functions for cGAS and kRAS in virtual screening. Comput Struct Biotechnol J. 2025;27:2176–2185.
Need validated scoring and refinement for your virtual screening pipeline? Our team can develop a target-specific scoring strategy or consensus rescoring protocol tailored to your target, docking poses, and experimental data. Contact our scientific team today to start your project.