Ab Initio Modeling & Co-evolutionary Analysis

From Orphan Sequence to Validated Structure. No Templates. No Homologs. Just Physics, Evolution, and AI.
Template-Free Folding Co-evolutionary Constraints Experimental Validation

Standard homology modeling collapses when no templates exist. We combine co-evolutionary contact prediction with physics-based folding and MD refinement, delivering validated structures for orphan proteins through integrated X-ray and Cryo-EM services.

Why Ab Initio Modeling Is the Critical Frontier

The druggable proteome is limited by the crystallizable proteome. ~40% of human proteins lack homologs for reliable homology modeling, including orphan receptors, novel fold enzymes, and intrinsically disordered proteins. Seed-stage biotechs are forced into expensive blind cell-based campaigns; pharma teams lose years to structural biology programs with uncertain outcomes.
Our platform breaks this dependency by combining co-evolutionary contact prediction with physics-based folding simulations, then validating through NMR, X-ray, or Cryo-EM. You receive a structural hypothesis for targets that exist outside the template universe.

What Sets the Platform Apart

Evolution + Physics

Evolutionary contacts guide Rosetta folding; MD refinement validates physical plausibility. We integrate both.

Novel Fold Recognition

Deep learning classifiers assess fold novelty and confidence, enabling functional annotation and pocket prediction even for orphan proteins.

Experimentally Validated

Every model advances to NMR validation, co-crystallography, or Cryo-EM. You receive coordinates backed by real data.

Technology Suite

Co-evolutionary Contact Prediction

Evolutionary Covariance as Folding GPS

Co-evolutionary Contact Prediction

Key Features:

  • Deep learning contact prediction (trRosetta, AlphaFold-derived) achieves >90% accuracy from MSAs.
  • Sparse MSA enhancement augments orphan proteins with metagenomic data, improving contact density 2–3×.
  • Distance distribution prediction enables precise 3D reconstruction beyond binary contacts.

Ideal For: Orphan proteins; novel fold families; intrinsically disordered proteins; metagenomic targets.

What We Offer:
Virtual biotechs access structural hypotheses for targets with no known structures. Pharma teams explore the "dark matter" of the proteome — structurally unique but therapeutically relevant proteins.

Physics-Based Folding & MD Refinement

From Contact Map to Stable Structure

Physics-Based Folding & MD Refinement

Key Features:

  • Rosetta fragment assembly guided by co-evolutionary distance restraints reduces conformational search space.
  • All-atom energy refinement optimizes side-chain packing and hydrogen bonding for realistic models.
  • Long-timescale MD equilibration (1–5 μs) tests stability and identifies deviations from physical behavior.

Ideal For: Targets with no detectable templates; novel fold enzymes; de novo designed proteins.

What We Offer:
Ab initio folding is computationally expensive and requires expertise in MSA construction, contact prediction, and energy landscape navigation. Our automated pipeline handles these complexities, delivering validated models without manual intervention.

Experimental Validation

From Prediction to Proof

Experimental Validation

Key Features:

  • NMR backbone assignment and STD-NMR confirm secondary structure and tertiary fold predictions.
  • Cryo-EM model building uses ab initio models as initial fits for single-particle analysis.
  • Construct redesign informed by quality assessment increases experimental success rates.

Ideal For: Programs requiring IND-grade evidence; orphan receptors; novel fold targets.

What We Offer:
We do not just predict; we prove. Our integrated structural biology pipeline validates ab initio models through the same experimental techniques used for any structural target.

Platform Instrumentation

MagHelix™ Core Instruments

Instrument Capability
NVIDIA DGX A100 Co-evolutionary contact prediction, Rosetta folding, deep learning fold recognition
NVIDIA RTX A6000 Cluster Real-time MD refinement and ensemble analysis
GROMACS/AMBER HPC Microsecond-scale all-atom MD; physics-based equilibration
Rosetta Suite Fragment assembly, ab initio folding, contact-guided prediction
Bruker AVANCE NEO 600 MHz NMR validation of fold predictions and disorder regions
Thermo Fisher Krios G4 Cryo-EM SPA for large complex validation

Standardized Workflow

Project Workflow

A milestone-driven execution system from sequence to validated ab initio model.

01 Target Review Week 1
02 Contact Prediction Week 1-2
03 Ab Initio Folding Week 2-4
04 MD Refinement Week 4-5
05 Validation Week 5-12

01 Target Review

  • Sequence analysis and domain annotation.
  • MSA construction with metagenomic augmentation.
  • Deliverable: MSA report + complexity assessment.

02 Contact Prediction

  • Deep learning contact prediction (trRosetta, AlphaFold-derived).
  • Distance distribution prediction and confidence scoring.
  • Deliverable: Contact map + confidence scores.

03 Ab Initio Folding

  • Rosetta fragment assembly with contact constraints.
  • QUARK physics-based folding for small domains.
  • Deliverable: Initial models + fold classification.

04 MD Refinement

  • All-atom MD equilibration and stability testing.
  • Restrained MD with contact preservation.
  • Deliverable: Refined ensemble + stability report.

05 Validation

  • NMR or X-ray/Cryo-EM validation (optional).
  • Model-to-experiment deviation analysis.
  • Deliverable: Validated model + final report.

Sample Requirements

  • Target Sequence: Amino acid sequence in FASTA format; UniProt ID acceptable
  • Known Information: Any functional data, predicted domains, or prior folding attempts
  • Prior Structural Data: Existing PDB entries or literature (for comparative analysis)
  • Project Background: Target class, disease relevance, known challenges (orphan status, disorder, complexity)

Standard Deliverables

  • Ab initio model with contact satisfaction scores and fold classification (PDB)
  • Co-evolutionary contact map and distance distributions
  • MD-refined conformational ensemble (50–200 representative PDBs)
  • Pocket analysis report with druggability scores
  • Experimental validation data (if selected): NMR, X-ray, or Cryo-EM coordinates
  • Final technical report with quality metrics and SBDD recommendations
  • Electronic data package (raw models, contact predictions, MD trajectories)

Frequently Asked Questions

Case Study

Case Study: Direct Coupling Analysis and the Attention Mechanism

Goal: Establish a parameter-efficient, attention-based DCA framework bridging transformer architectures with co-evolutionary protein modeling, enabling cross-family learning and generative sequence design.

Key Data:

  • Parameter compression: AttentionDCA achieves PlmDCA-comparable contact prediction with only 5–20% of parameters via factored attention decomposition.
  • Direct contact encoding: Attention matrices directly encode structural contacts, matching standard DCA accuracy without arbitrary Frobenius scoring.
  • Multi-family learning: Shares amino acid interaction parameters across protein families while retaining family-specific positional features.
  • Generative capability: Autoregressive model produces artificial MSAs reproducing natural two-site correlation statistics (Pearson r > 0.94).
  • Cross-family validation: Tested across nine diverse protein families with lengths of 53–202 residues and variable effective depths.

Why it matters: For drug developers targeting orphan proteins or engineering novel binders, this work demonstrates that attention-based co-evolutionary models capture structural constraints with minimal parameters while generalizing across protein families. The generative capability opens a path to in silico sequence design for therapeutic protein engineering, without building massive deep-learning infrastructure.

Contact map from Frobenius Score analysis

Figure 1. Contact map from Frobenius Score analysis of the AttentionDCA interaction tensor. Blue and red dots indicate positive and negative predictions, respectively; gray dots represent the native family structure. (Caredda F, et al., 2025)

Reference

Caredda F, Pagnani A. Direct coupling analysis and the attention mechanism. BMC Bioinformatics. 2025 Feb 6;26(1):41.

Ready to Fold the Unfoldable?
From orphan sequence to validated structure — no templates required.
Request Project Scoping →

Our technical team responds within 24 hours. All inquiries protected under NDA.