I develop computational approaches for understanding biology across molecular scales — from learning representations of individual molecules to modeling protein interactions and the biological systems those interactions reshape.
My research sits at the intersection of:
molecular representation learning · protein interactions · targeted protein degradation · molecular glues · mechanistic machine learning · biomedical knowledge graphs
Research question: How can AI move beyond memorizing observed chemical and biological space to reason about unseen molecules, interactions and biological states?
MOLECULES → REPRESENTATIONS → INTERACTIONS → BIOLOGICAL SYSTEMS → DISCOVERY
My projects address different parts of this continuum — from representation to intervention. The projects below are my own research (01–05); the final entry (06) is collaborative work.
Designing molecules that control protein fate
Why do apparently similar PROTACs produce very different degradation outcomes?
SynGlue approaches targeted protein degradation as a coupled molecular-design problem involving:
Target ligand + E3-ligase ligand + linker + ternary-complex geometry + degradation behaviour
rather than treating the warhead as the sole determinant of degrader activity.
The platform integrates computational approaches for analysing and designing PROTACs, including molecular generation, linker reasoning, degradation modelling and structure-informed prioritisation. The larger scientific question is how small molecules can be engineered to create productive interactions between proteins and redirect cellular machinery.
Research themes: Targeted Protein Degradation · PROTACs · Generative AI · Ternary Complexes · Polypharmacology
Learning representations of molecules
How should a molecule be represented when no single molecular description captures all of its biology and chemistry?
ChemicalDice / CDI explores multimodal molecular representation learning by integrating complementary molecular views into a unified latent representation. The framework brings together information derived from:
- physicochemical properties
- molecular graphs
- two-dimensional molecular representations
- bioactivity information
- quantum-chemical properties
- molecular language models
These complementary representations are integrated and subsequently distilled into a deployable molecular representation accessible from SMILES. The broader objective is to construct representations that remain useful beyond the exact chemical space observed during training.
Research themes: Multimodal Learning · Molecular Representations · Representation Distillation · Chemical Space · Drug Discovery
Resources: Code · Documentation
Mechanism-aware molecular machine learning
Can molecular ML distinguish compounds through the mechanisms by which they alter redox biology?
This work develops mechanism-aware models for analysing ROS modulators and antioxidant behaviour while connecting molecular information with biologically interpretable mechanisms.
Role: first author
Themes: Redox Biology · Mechanistic ML · Molecular Representations · Interpretability
Resources: Code
Evolution-guided ligand–GPCR prediction
Can evolutionary information improve molecular recognition models across the mammalian GPCRome?
EvOlf uses evolutionary and molecular information to model ligand–GPCR interactions across diverse mammalian receptors and species, with applications including receptor deorphanisation. The accompanying computational pipeline supports large-scale ligand–receptor screening using molecular featurisation, protein representations and deep-learning inference.
Role: first author
Themes: GPCRs · Evolutionary Biology · Protein Embeddings · Molecular Recognition
Resources: Code · Web Server · Pipeline
Thermodynamic fingerprints as biological signals
Can molecular-interaction thermodynamics encode diagnostically useful biological states?
Inertrope investigates machine-learning approaches based on thermodynamic and spectroscopic fingerprints for distinguishing biological sample states. This work reflects a broader interest in extracting predictive biological information from biophysical measurements rather than relying only on conventional molecular descriptors.
Role: first author
Themes: Biophysical ML · Thermodynamics · Diagnostics · Molecular Interactions
Resources: Code
Structure-guided molecular design
Can structural information guide the discovery of molecules that modulate protein signalling?
Gcoupler explores AI-driven structure-based molecular design with applications in GPCR–G-protein signalling and allosteric modulation. The framework combines computational molecular design, graph-based learning, structural information and bioactivity prioritisation.
Role: contributing author (collaboration)
Themes: Structure-Based Design · Graph Neural Networks · GPCR Signalling · Molecular Design
Represent
ChemicalDice / CDI — learn richer representations of molecular identity.
↓
Recognise
EvOlf · Trojan-Horses · Gcoupler — understand how molecular structure relates to biological recognition and mechanism.
↓
Reprogram
SynGlue — study molecules that create, stabilise, inhibit or redirect protein interactions (targeted degradation).
↓
Reason
Biomedical knowledge graphs · pathway models — connect molecular perturbations to larger biological systems.
↓
Discover
Develop computational strategies for molecular intervention and biological discovery.
The models above depend on molecular and biological data that are standardised, traceable and evaluation-ready. I therefore maintain a set of lightweight computational-biology tools addressing recurring infrastructure problems.
PREPARE → HARMONISE → MAP → SPLIT → STRUCTURE → AUDIT
Prepare — smiles-cleankit Canonicalise, validate and standardise molecular structures.
Harmonise — assaytablecleaner Standardise bioactivity measurements and derive comparable activity values.
Map — molidmapper Resolve and harmonise molecular and biological identifiers across databases.
Split — scaffoldsplitlab Generate leakage-aware molecular machine-learning splits.
Structure — biokg-signmapper Standardise relation semantics in biomedical knowledge graphs.
Audit — kg-stats-audit Inspect graph structure, connectivity and dataset quality.
Toolkit — CompBio Toolkit Suite A common entry point connecting the reusable components.
SciSVG is an open collection of editable scientific vector graphics designed for figures, presentations and scientific communication.
This project reflects another principle of my work: scientific outputs should be reusable — not only scientific models and datasets, but also the tools used to communicate them.
- Molecular generalisation — How do molecular models remain useful outside the chemical space on which they were trained?
- Representation complementarity — What genuinely new information does one molecular representation contribute beyond another?
- Interaction biology — How can small molecules create, stabilise, inhibit or reconfigure protein interactions?
- Mechanistic machine learning — Can predictive models provide insight into biological mechanisms rather than producing endpoint scores alone?
- Biological state — How can molecular perturbations be linked to pathways, interaction networks and system-level biological consequences?
- Reproducibility — How do we turn computational experiments into research systems that another scientist can reproduce and extend?
Python · PyTorch · RDKit · scikit-learn · pandas · NumPy
Deep Learning · Multimodal Learning · Molecular Representation Learning
Protein–Ligand Modelling · Protein Interaction Modelling
Cheminformatics · Biomedical Knowledge Graphs
Docker · Reproducible Pipelines · Scientific Software
I am interested in collaborations spanning:
- molecular representation learning
- molecular recognition
- protein-interaction modulation
- targeted protein degradation
- molecular glues
- multimodal biological AI
- biomedical knowledge graphs
- computational drug discovery
Computational Biology · IIIT Delhi








