Supplementary Data for NIPS Publication: Protein Interface Prediction using Graph Convolutional Networks.
These data sets can be used to re-run the experiments from our paper, Protein Interface Prediction using Graph Convolutional Networks. The data are derived from protein complexes in the docking benchmark dataset v. 5.0. Each file is a python tuple that has been saved using cPickle and compre
These data sets can be used to re-run the experiments from our paper, Protein Interface Prediction using Graph Convolutional Networks. The data are derived from protein complexes in the docking benchmark dataset v. 5.0. Each file is a python tuple that has been saved using cPickle and compressed using gzip.
Links:
Paper: https://papers.nips.cc/paper/7231-protein-interface-prediction-using-graph-convolutional-networks
Poster: https://zenodo.org/record/1134154
Code: https://github.com/fouticus/pipgcn
File Descriptions:
train.cpkl.gz and test.cpkl.gz have the data formatted for neighborhood based graph convolutions. The diffc_ files are the same data formatted for the diffusion convolutional neural networks that we compare against.
train.cpkl.gz is a tuple of length 2:
- element 0 is a list of length 175 containing the PDB codes from the docking benchmark dataset
- element 1 is a list of length 175 containing features for each protein. Each element is a dictionary containing the following keys:
- r_vertex: vertex (residue) features for the receptor. numpy array of shape (x, 70) where x is the number of residues in the receptor and 70 is the number of features.
- l_vertex: vertex (residue) features for the ligand. analogous to above, with shape (y, 70) where y is the number of residues in the ligand.
- complex_code: PDB code of the complex. matches the list of codes described above.
- l_edge: edge features for the neighborhood around each residue in the ligand. numpy array of shape (y, 20, 2) where y is defined as above. the second dimension is the edges to the 20 nearest neighboring residues, ordered by decreasing distance. The third dimension allows for two features per edge.
- r_edge: edge features for the neighborhood around each residue in the receptor. numpy array of shape (x, 20, 2) where x is as above.
- l_hood_indices: the index of the 20 closest residues to each residue, ordered by decreasing distance. numpy array of shape (y, 20, 1). "Index" means which row in l_vertex gives the vertex features for the closest neighbor, second closest neighbor, etc.
- r_hood_indices: analogous to above, shape (x, 20, 1).
- label: 1 or -1 label for each residue pair. numpy array of shape (x*y, 3). Each row looks like (i, j, k) where i is the index of the ligand residue, j is the index of the receptor residue, and k is either -1 (negative example) or 1 (positive example).
test.cpkl.gz matches the structure of train.cpkl.gz except it has the test set of 55 complexes.
Descriptions of the vertex and edge features can be found in Appendix A of this.
diffc_g2_p2_train.cpkl.gz is a tuple of le
📤 Share this page
Found this useful? Share it with your network.
Files are hosted on the source repository. Click download to access the full dataset.