{
  "id": 506782,
  "title": "112M molecule measurements ",
  "url": "/competitions/leash-BELKA/discussion/506782",
  "author_name": "",
  "post_date": "2024-05-23T07:24:39.492100600Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I have a question regarding the structure of the training datasets. The introduction states that the available data consists of 3.6 billion measurements (3 targets * 3 rounds of selection * 3 replicates * 133M molecules).  Does each molecule for each target have 9 measurements (3 rounds of selection * 3 replicates)?  I am not sure about this information. If each molecule has 9 measurements, shouldn't the total number of unique SMILES be 133 million instead of 98415610?  How can we distinguish between the same molecules that are measured under different rounds of selection and different replicates?</p>",
  "messages": [
    {
      "id": "2830432",
      "postDate": "05/23/2024 07:24:39",
      "content": "<p>I have a question regarding the structure of the training datasets. The introduction states that the available data consists of 3.6 billion measurements (3 targets * 3 rounds of selection * 3 replicates * 133M molecules).  Does each molecule for each target have 9 measurements (3 rounds of selection * 3 replicates)?  I am not sure about this information. If each molecule has 9 measurements, shouldn't the total number of unique SMILES be 133 million instead of 98415610?  How can we distinguish between the same molecules that are measured under different rounds of selection and different replicates?</p>",
      "rawMarkdown": "I have a question regarding the structure of the training datasets. The introduction states that the available data consists of 3.6 billion measurements (3 targets * 3 rounds of selection * 3 replicates * 133M molecules).  Does each molecule for each target have 9 measurements (3 rounds of selection * 3 replicates)?  I am not sure about this information. If each molecule has 9 measurements, shouldn't the total number of unique SMILES be 133 million instead of 98415610?  How can we distinguish between the same molecules that are measured under different rounds of selection and different replicates?",
      "votes": null
    },
    {
      "id": "2831360",
      "postDate": "05/23/2024 16:31:45",
      "content": "<p>We don't get the 3x3 measurement data, only the final result (binds = 1 or binds = 0 for each protein)</p>\n<p>With shared and non shared building blocks, if you added the non-shared building blocks together with the building blocks in train, presumably you would get 133 million possible combinations. I haven't done the math. However, a small number of building blocks were held out of train data, and partial overlaps (some in train bb, some in test bb) were excluded completely from the data provided to us. This decreases the total data given to us by a lot, because there's many more partial overlap (around 30 million?) than no overlap (tens of thousands) given how bb combinations work. </p>",
      "rawMarkdown": "We don't get the 3x3 measurement data, only the final result (binds = 1 or binds = 0 for each protein)\n\nWith shared and non shared building blocks, if you added the non-shared building blocks together with the building blocks in train, presumably you would get 133 million possible combinations. I haven't done the math. However, a small number of building blocks were held out of train data, and partial overlaps (some in train bb, some in test bb) were excluded completely from the data provided to us. This decreases the total data given to us by a lot, because there's many more partial overlap (around 30 million?) than no overlap (tens of thousands) given how bb combinations work.",
      "votes": null
    },
    {
      "id": "2831375",
      "postDate": "05/23/2024 16:36:24",
      "content": "<p>Oh, and the separate library of non-triazines, while they limited the test data to a sample of 500,000, is around 18 million combinations I recall. So maybe it's around 115million triazine, of which 1 million of in test data, and ~15 million excluded, plus 18 million from the separate library</p>",
      "rawMarkdown": "Oh, and the separate library of non-triazines, while they limited the test data to a sample of 500,000, is around 18 million combinations I recall. So maybe it's around 115million triazine, of which 1 million of in test data, and ~15 million excluded, plus 18 million from the separate library",
      "votes": null
    },
    {
      "id": "2832984",
      "postDate": "05/24/2024 02:38:04",
      "content": "<p>Robert, Thanks very much for your detailed explanation. I am clear now. </p>",
      "rawMarkdown": "Robert, Thanks very much for your detailed explanation. I am clear now.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2831360,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "05/23/2024 16:31:45",
      "content": "<p>We don't get the 3x3 measurement data, only the final result (binds = 1 or binds = 0 for each protein)</p>\n<p>With shared and non shared building blocks, if you added the non-shared building blocks together with the building blocks in train, presumably you would get 133 million possible combinations. I haven't done the math. However, a small number of building blocks were held out of train data, and partial overlaps (some in train bb, some in test bb) were excluded completely from the data provided to us. This decreases the total data given to us by a lot, because there's many more partial overlap (around 30 million?) than no overlap (tens of thousands) given how bb combinations work. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2831375,
          "author_name": "roberthatch",
          "author_url": "",
          "post_date": "05/23/2024 16:36:24",
          "content": "<p>Oh, and the separate library of non-triazines, while they limited the test data to a sample of 500,000, is around 18 million combinations I recall. So maybe it's around 115million triazine, of which 1 million of in test data, and ~15 million excluded, plus 18 million from the separate library</p>",
          "votes": null,
          "replies": [
            {
              "id": 2832984,
              "author_name": "minwangcnn",
              "author_url": "",
              "post_date": "05/24/2024 02:38:04",
              "content": "<p>Robert, Thanks very much for your detailed explanation. I am clear now. </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2830432": "I have a question regarding the structure of the training datasets. The introduction states that the available data consists of 3.6 billion measurements (3 targets * 3 rounds of selection * 3 replicates * 133M molecules).  Does each molecule for each target have 9 measurements (3 rounds of selection * 3 replicates)?  I am not sure about this information. If each molecule has 9 measurements, shouldn't the total number of unique SMILES be 133 million instead of 98415610?  How can we distinguish between the same molecules that are measured under different rounds of selection and different replicates?",
    "2831360": "We don't get the 3x3 measurement data, only the final result (binds = 1 or binds = 0 for each protein)\n\nWith shared and non shared building blocks, if you added the non-shared building blocks together with the building blocks in train, presumably you would get 133 million possible combinations. I haven't done the math. However, a small number of building blocks were held out of train data, and partial overlaps (some in train bb, some in test bb) were excluded completely from the data provided to us. This decreases the total data given to us by a lot, because there's many more partial overlap (around 30 million?) than no overlap (tens of thousands) given how bb combinations work.",
    "2831375": "Oh, and the separate library of non-triazines, while they limited the test data to a sample of 500,000, is around 18 million combinations I recall. So maybe it's around 115million triazine, of which 1 million of in test data, and ~15 million excluded, plus 18 million from the separate library",
    "2832984": "Robert, Thanks very much for your detailed explanation. I am clear now."
  },
  "source": "meta"
}