{
  "id": 506985,
  "title": "What should we do with the input of protein names?",
  "url": "/competitions/leash-BELKA/discussion/506985",
  "author_name": "",
  "post_date": "2024-05-24T02:46:59.502326200Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I am a computer science major, and there are only two biology students in my group. So I lack some background knowledge.</p>\n<p>Only the name of the protein is given in the dataset, do we need to take the property site of the protein in question as input? Or do we just use the name as the embedding input?</p>\n<p>Thanks for your advice.</p>",
  "messages": [
    {
      "id": "2832989",
      "postDate": "05/24/2024 02:46:59",
      "content": "<p>I am a computer science major, and there are only two biology students in my group. So I lack some background knowledge.</p>\n<p>Only the name of the protein is given in the dataset, do we need to take the property site of the protein in question as input? Or do we just use the name as the embedding input?</p>\n<p>Thanks for your advice.</p>",
      "rawMarkdown": "I am a computer science major, and there are only two biology students in my group. So I lack some background knowledge.\n\nOnly the name of the protein is given in the dataset, do we need to take the property site of the protein in question as input? Or do we just use the name as the embedding input?\n\nThanks for your advice.",
      "votes": null
    },
    {
      "id": "2834082",
      "postDate": "05/24/2024 14:23:23",
      "content": "<p>Is just to reference to what protein refers the binding column. Since there are binding information of the 3 targets for each of the samples on train. Is not usable information at all. You could use the structure as reinforcement information. But as I said. I don't think will be usefull use a variable that actually doesn't vary.</p>",
      "rawMarkdown": "Is just to reference to what protein refers the binding column. Since there are binding information of the 3 targets for each of the samples on train. Is not usable information at all. You could use the structure as reinforcement information. But as I said. I don't think will be usefull use a variable that actually doesn't vary.",
      "votes": null
    },
    {
      "id": "2841361",
      "postDate": "05/28/2024 14:54:31",
      "content": "<p>From the data tab of the competition</p>\n<blockquote>\n  <p>We screened EPHX2/sEH purchased from Cayman Chemical, a life sciences commercial vendor. For those contestants wishing to incorporate protein structural information in their submissions, the amino sequence is positions 2-555 from UniProt entry P34913, the crystal structure can be found in PDB entry 3i28, and predicted structure can be found in AlphaFold2 entry 34913. Additional EPHX2/sEH crystal structures with ligands bound can be found in PDB.<br>\n  We screened BRD4 purchased from Active Motif, a life sciences commercial vendor. For those contestants wishing to incorporate protein structural information in their submissions, the amino acid sequence is positions 44-460 from UniProt entry O60885-1, the crystal structure (for a single domain) can be found in PDB entry 7USK and predicted structure can be found in AlphaFold2 entry O60885. Additional BRD4 crystal structures with ligands bound can be found in PDB.<br>\n  We screened ALB purchased from Active Motif. For those contestants wishing to incorporate protein structural information in their submissions, the amino acid sequence is positions 25 to 609 from UniProt entry P02768, the crystal structure can be found in PDB entry 1AO6, and predicted structure can be found in AlphaFold2 entry P02768. Additional ALB crystal structures with ligands bound can be found in PDB.</p>\n</blockquote>\n<p>You can have some more meaningful preprocessing using this information</p>",
      "rawMarkdown": "From the data tab of the competition\n>We screened EPHX2/sEH purchased from Cayman Chemical, a life sciences commercial vendor. For those contestants wishing to incorporate protein structural information in their submissions, the amino sequence is positions 2-555 from UniProt entry P34913, the crystal structure can be found in PDB entry 3i28, and predicted structure can be found in AlphaFold2 entry 34913. Additional EPHX2/sEH crystal structures with ligands bound can be found in PDB.\n>We screened BRD4 purchased from Active Motif, a life sciences commercial vendor. For those contestants wishing to incorporate protein structural information in their submissions, the amino acid sequence is positions 44-460 from UniProt entry O60885-1, the crystal structure (for a single domain) can be found in PDB entry 7USK and predicted structure can be found in AlphaFold2 entry O60885. Additional BRD4 crystal structures with ligands bound can be found in PDB.\n>We screened ALB purchased from Active Motif. For those contestants wishing to incorporate protein structural information in their submissions, the amino acid sequence is positions 25 to 609 from UniProt entry P02768, the crystal structure can be found in PDB entry 1AO6, and predicted structure can be found in AlphaFold2 entry P02768. Additional ALB crystal structures with ligands bound can be found in PDB.\n\nYou can have some more meaningful preprocessing using this information",
      "votes": null
    },
    {
      "id": "2841439",
      "postDate": "05/28/2024 15:21:31",
      "content": "<p>There is a clear difference in the percentage of molecules that bind to each of the three proteins.  (Create an EDA to explore the size of the difference.)</p>\n<p>Many of the shared notebooks build individual models and predictions for each protein separately.  The expectation for that design would be that you can find molecule features that explain why one protein is able to bind to more molecules than another.  </p>\n<p>If you create a model that includes the protein as a feature you will likely end up with a different submission (and LB score), but a simple one-shot encoding of the protein will likely not explain much - so you will want to build features on the protein's.  A number of research papers and code have been disclosed in discussion topics.  The code I have looked at in some detail wants to generalize for both molecules and proteins and therefore follows this path.</p>\n<p>The test file includes no new proteins but does include \"new\" molecules.  </p>\n<p>Will features based on the proteins help predict binding on new molecules - who knows ??   So try it and find out.</p>",
      "rawMarkdown": "There is a clear difference in the percentage of molecules that bind to each of the three proteins.  (Create an EDA to explore the size of the difference.)\n\nMany of the shared notebooks build individual models and predictions for each protein separately.  The expectation for that design would be that you can find molecule features that explain why one protein is able to bind to more molecules than another.  \n\nIf you create a model that includes the protein as a feature you will likely end up with a different submission (and LB score), but a simple one-shot encoding of the protein will likely not explain much - so you will want to build features on the protein's.  A number of research papers and code have been disclosed in discussion topics.  The code I have looked at in some detail wants to generalize for both molecules and proteins and therefore follows this path.\n\nThe test file includes no new proteins but does include \"new\" molecules.  \n\nWill features based on the proteins help predict binding on new molecules - who knows ??   So try it and find out.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2834082,
      "author_name": "sacuscreed",
      "author_url": "",
      "post_date": "05/24/2024 14:23:23",
      "content": "<p>Is just to reference to what protein refers the binding column. Since there are binding information of the 3 targets for each of the samples on train. Is not usable information at all. You could use the structure as reinforcement information. But as I said. I don't think will be usefull use a variable that actually doesn't vary.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2841361,
      "author_name": "giorgiomicaletto",
      "author_url": "",
      "post_date": "05/28/2024 14:54:31",
      "content": "<p>From the data tab of the competition</p>\n<blockquote>\n  <p>We screened EPHX2/sEH purchased from Cayman Chemical, a life sciences commercial vendor. For those contestants wishing to incorporate protein structural information in their submissions, the amino sequence is positions 2-555 from UniProt entry P34913, the crystal structure can be found in PDB entry 3i28, and predicted structure can be found in AlphaFold2 entry 34913. Additional EPHX2/sEH crystal structures with ligands bound can be found in PDB.<br>\n  We screened BRD4 purchased from Active Motif, a life sciences commercial vendor. For those contestants wishing to incorporate protein structural information in their submissions, the amino acid sequence is positions 44-460 from UniProt entry O60885-1, the crystal structure (for a single domain) can be found in PDB entry 7USK and predicted structure can be found in AlphaFold2 entry O60885. Additional BRD4 crystal structures with ligands bound can be found in PDB.<br>\n  We screened ALB purchased from Active Motif. For those contestants wishing to incorporate protein structural information in their submissions, the amino acid sequence is positions 25 to 609 from UniProt entry P02768, the crystal structure can be found in PDB entry 1AO6, and predicted structure can be found in AlphaFold2 entry P02768. Additional ALB crystal structures with ligands bound can be found in PDB.</p>\n</blockquote>\n<p>You can have some more meaningful preprocessing using this information</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2841439,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "05/28/2024 15:21:31",
      "content": "<p>There is a clear difference in the percentage of molecules that bind to each of the three proteins.  (Create an EDA to explore the size of the difference.)</p>\n<p>Many of the shared notebooks build individual models and predictions for each protein separately.  The expectation for that design would be that you can find molecule features that explain why one protein is able to bind to more molecules than another.  </p>\n<p>If you create a model that includes the protein as a feature you will likely end up with a different submission (and LB score), but a simple one-shot encoding of the protein will likely not explain much - so you will want to build features on the protein's.  A number of research papers and code have been disclosed in discussion topics.  The code I have looked at in some detail wants to generalize for both molecules and proteins and therefore follows this path.</p>\n<p>The test file includes no new proteins but does include \"new\" molecules.  </p>\n<p>Will features based on the proteins help predict binding on new molecules - who knows ??   So try it and find out.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2832989": "I am a computer science major, and there are only two biology students in my group. So I lack some background knowledge.\n\nOnly the name of the protein is given in the dataset, do we need to take the property site of the protein in question as input? Or do we just use the name as the embedding input?\n\nThanks for your advice.",
    "2834082": "Is just to reference to what protein refers the binding column. Since there are binding information of the 3 targets for each of the samples on train. Is not usable information at all. You could use the structure as reinforcement information. But as I said. I don't think will be usefull use a variable that actually doesn't vary.",
    "2841361": "From the data tab of the competition\n>We screened EPHX2/sEH purchased from Cayman Chemical, a life sciences commercial vendor. For those contestants wishing to incorporate protein structural information in their submissions, the amino sequence is positions 2-555 from UniProt entry P34913, the crystal structure can be found in PDB entry 3i28, and predicted structure can be found in AlphaFold2 entry 34913. Additional EPHX2/sEH crystal structures with ligands bound can be found in PDB.\n>We screened BRD4 purchased from Active Motif, a life sciences commercial vendor. For those contestants wishing to incorporate protein structural information in their submissions, the amino acid sequence is positions 44-460 from UniProt entry O60885-1, the crystal structure (for a single domain) can be found in PDB entry 7USK and predicted structure can be found in AlphaFold2 entry O60885. Additional BRD4 crystal structures with ligands bound can be found in PDB.\n>We screened ALB purchased from Active Motif. For those contestants wishing to incorporate protein structural information in their submissions, the amino acid sequence is positions 25 to 609 from UniProt entry P02768, the crystal structure can be found in PDB entry 1AO6, and predicted structure can be found in AlphaFold2 entry P02768. Additional ALB crystal structures with ligands bound can be found in PDB.\n\nYou can have some more meaningful preprocessing using this information",
    "2841439": "There is a clear difference in the percentage of molecules that bind to each of the three proteins.  (Create an EDA to explore the size of the difference.)\n\nMany of the shared notebooks build individual models and predictions for each protein separately.  The expectation for that design would be that you can find molecule features that explain why one protein is able to bind to more molecules than another.  \n\nIf you create a model that includes the protein as a feature you will likely end up with a different submission (and LB score), but a simple one-shot encoding of the protein will likely not explain much - so you will want to build features on the protein's.  A number of research papers and code have been disclosed in discussion topics.  The code I have looked at in some detail wants to generalize for both molecules and proteins and therefore follows this path.\n\nThe test file includes no new proteins but does include \"new\" molecules.  \n\nWill features based on the proteins help predict binding on new molecules - who knows ??   So try it and find out."
  },
  "source": "meta"
}