In silico protein binder database

Search designed binder candidates by target UniProt accession or entry name

Bindome provides over 300,000 in silico protein binder candidates generated with BindCraft, covering more than 8,000 human protein targets.

Bindome logo
Computational workflow used to create the Bindome database.
Computational workflow used to create the Bindome database.

Bindome is an open database of computationally designed protein binder candidates for the human proteome. These candidates were generated to provide starting points for researchers interested in studying, detecting, or perturbing human proteins.

The current release was built using BindCraft, an experimentally validated protein binder design pipeline that generates binder candidates from protein structures. To run BindCraft at proteome scale, we developed an accelerated version of the workflow tailored for deployment on high-performance computing systems. This allowed us to scale binder design across thousands of human proteins and release the resulting candidates as a precomputed resource: the Bindome.

To build the Bindome, we started from AlphaFold models from the AlphaFold database. Since current binder design methods are best suited for compact, well-structured protein regions, we first divided full-length proteins into structurally well-defined domains. We then selected the domains most suitable for the current workflow, excluding regions that were poorly predicted, highly disordered or membrane-embedded. Candidate binders were designed against accessible domain surfaces and then checked in the context of the full-length protein model to reduce obvious structural clashes with other confidently positioned regions of the same protein.

Bindome is intended as a practical starting point. Users can search for a protein of interest, inspect predicted binder models, compare confidence metrics, and download candidates for downstream analysis or experimental testing.

Note that all entries in Bindome are computational predictions. They have not been experimentally validated, and their binding, specificity, stability, expression, and activity in biological assays are not guaranteed. Experimental testing is essential before using any candidate as a research reagent.

About this release

This release contains 306,151 in silico binder candidates for 8,296 human proteins, corresponding to 40.9% of the full human proteome and 77.8% of the targetable proteome defined in this work. Each candidate includes a designed amino-acid sequence, a predicted binder-target structure and confidence metrics.

These data are provided as a catalogue of starting points for exploration and follow-up. Users can search for candidate binders to a protein of interest, inspect their predicted binding modes, compare designs using confidence metrics, and select candidates for downstream analysis or experimental testing.

The current release should be considered a starting resource for exploring candidate binders, not a set of validated reagents. Binding, specificity, expression, stability, and activity in biological assays have not been experimentally established. Users should evaluate individual candidates according to their own downstream criteria before experimental use.

What’s next?

This first release focuses on the human proteome and on structured regions that are suitable for the current workflow. Future releases aim to expand coverage to additional proteomes and to more challenging target classes, including conformation-specific states, disordered regions, and membrane proteins with accessible domains. We also plan to add potential functional annotations for binder candidates, such as predicted protein-protein interaction disruptors, based on how they interact with their targets.

License & citation

Data are available for academic and commercial use under a Creative Commons Attribution 4.0 International license.

Citation: preprint coming soon.

Model detail

Binder model

Select a model from the search results to inspect it here.

3D model

Predicted aligned error

PAE matrix not loaded.

Average AF metrics

FAQs

Frequently asked questions

Project

Bindome is an open database of computationally designed protein binder candidates for the human proteome. Each entry links a designed binder candidate to a human target protein and includes a designed amino-acid sequence, a predicted binder-target structure, and confidence metrics to help users inspect and prioritize candidates.

The current release contains 306,151 in silico binder candidates for 8,296 human proteins, covering 40.9% of the full human proteome and 77.8% of the targetable proteome defined in this work.

No. Bindome entries are computational predictions and should be treated as candidate binders, not validated affinity reagents.

The candidates were generated with BindCraft, an experimentally validated protein binder design pipeline, using its default settings and filters. However, individual Bindome candidates have not been experimentally tested. Their binding, selectivity, expression, stability, and activity in biological assays are not guaranteed.

Experimental validation is essential before using any candidate as a research reagent.

Bindome provides starting points for researchers who want to explore candidate binders for human proteins without running large-scale binder design campaigns themselves.

Users can search for a protein of interest, compare confidence metrics, inspect predicted binder-target models, and download structures for downstream analysis or experimental testing.

Example web/API workflow: query Bindome by UniProt accession, retrieve candidates, rank by AlphaFold confidence metrics, visualize candidates, and download results.
Example web/API workflow.

Bindome can also help generate hypotheses for protein-level perturbation. For example, a candidate binder that contacts a protein-protein interaction interface may be tested as a possible interaction disruptor. Similarly, candidates predicted near ligand-binding sites, post-translational modification sites, or DNA-binding regions may help users explore whether those regions can be functionally perturbed.

The main limitation is that the binders are not experimentally validated. They were designed for target proteins, but their actual binding strength and selectivity are unknown.

This release also focuses on protein regions that are suitable for the current design workflow: compact, well-structured domains with confident AlphaFold quality metrics. As a result, many disordered regions, regions with low AlphaFold confidence, and membrane-embedded regions are not covered.

Finally, coverage should not be interpreted as complete surface coverage. Many proteins have at least one candidate binder, but only part of the total accessible protein surface is currently targeted.

In this release, the "targetable proteome" refers to the subset of human proteins that contain at least one region suitable for the current binder design workflow.

Practically, this means compact, well-structured domains that are confidently modeled by AlphaFold, are not membrane-embedded, and pass additional filters related to size, shape, and structural quality. Proteins or regions that do not meet these criteria may still be biologically important, but they were outside the scope of this release.

Bindome data can be accessed in several ways:

  • through the interactive web interface
  • through bulk downloads
  • programmatically through an API
  • through an MCP server for natural-language querying with compatible LLM tools
  • through machine-learning splits released on Hugging Face
Overview of the different ways to access Bindome data.
Overview of the different ways to access Bindome data.

Please cite:

Human Bindome: A Proteome-scale Atlas of Designed Binder Candidates. XXXX

Website

Yes. You can enter several UniProt accessions or entry names in the search bar, separated by spaces, commas, or semicolons.

The option to display the full protein structure together with binder candidates is only available when searching for a single UniProt accession.

When searching for multiple accessions at the same time, this option is disabled to keep the visualization manageable.

There are several possible reasons.

Your protein may not contain a domain that passed the target-selection filters used in this release. For example, the relevant region may have low AlphaFold confidence metrics, be highly disordered, be membrane-embedded, or be too short for the current workflow.

It is also possible that the protein contains targetable domains, but no binder candidates passed the design and filtering criteria. Bindome is a large resource, but it does not yet cover every human protein or every targetable region.

Bindome candidates were generated with the default BindCraft design workflow and without specifying predefined binding hotspots.

This means that the design process was allowed to find suitable binding sites on the selected domain surfaces, rather than being forced to cover every region of a protein. As a result, some surfaces may have many candidates, while others may have none.

In addition, only domains that passed the target-selection filters were used for design, so disordered regions, membrane-embedded regions, and regions with low AlphaFold confidence are generally not targeted in this release.

A good starting point is to compare candidates using the available confidence metrics and predicted structures. Candidates with stronger in silico confidence may be better starting points, but the metrics do not guarantee experimental success.

Users should also consider the intended application. For example, a candidate predicted to bind near a functional site may be more useful for perturbation experiments, while a candidate binding an exposed surface may be more useful for detection.

Whenever possible, we recommend prioritizing more than one candidate for experimental testing.

The confidence metrics are computational scores used to assess and rank predicted binder candidates. They are useful for prioritization, but they are not experimental measurements.

Some of the AlphaFold-derived metrics included in Bindome are:

  • pLDDT: estimates the confidence of the predicted structure at the residue level. Higher values indicate higher confidence.
  • ipLDDT: pLDDT values at the predicted binder-target interface.
  • iPAE: estimates the predicted alignment error (PAE) between the binder and the target. Lower values indicate higher confidence in the relative binder-target placement.
  • ipTM: estimates the confidence of the predicted interaction between chains in a complex. Higher values indicate higher confidence.
  • ipSAE: an interface-focused confidence score used to assess the predicted binder-target interaction. Higher values indicate higher confidence. Dunbrack 2026

Accepted candidates in Bindome passed predefined in silico confidence filters from the BindCraft workflow. These metrics should not be interpreted as direct measurements of binding affinity, selectivity, expression, or biological activity.

Connectivity

The Bindome MCP server allows compatible large language model tools to query Bindome data using natural-language requests.

Instead of manually building database or API queries, users can ask questions such as which candidates are available for a target protein or whether any candidates overlap a type of functional site. The LLM tool can then retrieve and summarize relevant Bindome results.

The Bindome MCP server URL is https://bindome.epfl.ch/mcp and no authentication is required.

In ChatGPT, MCP app support depends on your plan and workspace settings. If available in your workspace, enable Developer mode, create a new custom app or connector, add the Bindome MCP URL, and select no authentication. You can then enable the Bindome app in a new chat.

In Claude, MCP app support also depends on your plan and workspace settings. If available in your workspace, go to Customize > Plugins, choose the option to add a custom connector, and enter the Bindome MCP URL.

You can ask target-centered or candidate-centered questions. For example:

  • "Which binder candidates are available for TP53?"
  • "Show me candidates for this UniProt accession."
  • "Which candidates have the best confidence metrics for this target?"

The answers should still be interpreted as summaries of computational predictions, not experimental validation.

Check the API documentation for full details.

Yes. Bindome provides bulk download options for users who want to analyze the complete resource or predefined subsets.

In the download section you can get binder sequences, predicted structures, BindCraft metrics, PAE matrices and a summary table of the targeted proteins.

Technical details

Bindome was built by scaling the BindCraft protein binder design pipeline to the human proteome.

The workflow started from human protein structures and PAE matrices from the AlphaFold Database. Full-length proteins were divided into structured domains, and only domains suitable for the current design workflow were selected as targets.

This domain-based strategy helped in two ways. First, it reduced the size of the design problem, which is important because BindCraft runtime increases strongly with target size. Second, it focused the design process on compact, well-structured regions where current structure-based binder design methods are more likely to work.

Candidate binders were accepted using the default BindCraft in silico quality filters. Because Bindome binders were designed against isolated domains, we added an additional full-length protein check: each binder was placed back into the full protein model to reduce obvious clashes with other confidently positioned regions, based on AlphaFold PAE information.

Computational workflow used to create the Bindome database.
Computational workflow used to create the Bindome database.

Target domains were selected from full-length AlphaFold models using the AlphaFold Predicted Aligned Error (PAE) matrices.

PAE-based domain segmentation was performed using the pae_to_domains tool by Tristan Croll: github.com/tristanic/pae_to_domains. The resulting domains were filtered using several criteria, including AlphaFold pLDDT confidence, domain length, compactness, intra-domain contact density, secondary-structure composition, and overlap with membrane-embedded regions.

Only domains that passed these filters were used for binder design.

Some domains were excluded because they were not suitable for the current structure-based binder design workflow. These exclusions helped focus computational resources on protein regions where current binder design methods are more likely to produce useful candidates.

All binders in this release were designed with a fixed length of 90 amino acids for both computational and practical reasons.

Computationally, fixing the binder length helped accelerate the workflow by avoiding repeated recompilation steps during BindCraft runs. This was important for scaling binder design across thousands of target domains on high-performance computing systems.

Practically, 90-amino-acid binders are relatively small and cost-effective to produce experimentally. This helps make Bindome more accessible to the scientific community and supports experimental exploration of candidate binders in downstream applications.

Bindome was built using structural models and PAE matrices from the AlphaFold Protein Structure Database version 6.

We adapted BindCraft into an accelerated workflow tailored for high-performance computing deployment.

The main changes were:

  • Running many independent design jobs in parallel on HPC systems.
  • Using a domain-based strategy to reduce target size.
  • Fixing binder length to 90 amino acids to avoid repeated recompilation steps.

Together, these changes made it possible to run binder design across thousands of human protein domains.

Binders were filtered to reduce obvious clashes with confidently positioned regions of the full-length protein model.

However, not all apparent clashes were treated equally. If the clashing region had low AlphaFold confidence, based on pLDDT, or if its position relative to the target domain was uncertain, based on PAE, the candidate was not automatically discarded.

The reason is that these regions may adopt a different arrangement from the one shown in the AlphaFold model. In these cases, an apparent clash may reflect uncertainty in the model rather than a true physical incompatibility.

Representative target structure with a binder that overlaps the full-length target yet was retained, owing to low predicted aligned error (PAE) over the overlapping target region.
Representative target structure with a binder that overlaps the full-length target yet was retained, owing to low predicted aligned error (PAE) over the overlapping target region.

Bindome used the default BindCraft in silico confidence filters without modification. These filters evaluate properties such as binder structure confidence and predicted binder-target complex quality. More details are available in the BindCraft publication and GitHub repository.

In addition, Bindome applied a full-length protein clash check because candidates were designed against individual domains and then placed back into the context of the full protein model.

Binder selectivity has not been experimentally validated.

Each candidate was designed for a chosen target protein or target domain, but this does not prove that it binds only that protein. Some candidates may not bind their intended target, and some may bind other proteins or surfaces.

Users should experimentally test both binding and selectivity for their intended application.

Bindome includes predefined train, validation, and test splits for machine-learning use. These splits can be downloaded through Hugging Face: huggingface.co/datasets/wjulius/HumanBindome

The splitting strategy was inspired by the PINDER approach for reducing information leakage in protein-interface benchmarks.

Bindome complexes were grouped by similarity of their binding interfaces. Related interfaces were assigned to the same partition where possible. This makes the splits more suitable for benchmarking models that should generalize to new binder-target interfaces, rather than memorizing similar examples from the training set.

Downloads

Bulk downloads

Download Bindome data in bulk. Available files include predicted structures, binder sequences, BindCraft metrics, PAE matrices, and a summary table of targeted proteins.

Available datasets

Loading available downloads…
Organism Subset Binders Targets Domains Download links Last update
Loading available downloads…