{
  "id": 518729,
  "title": "Models that could lead to shakeup",
  "url": "/competitions/leash-BELKA/discussion/518729",
  "author_name": "",
  "post_date": "2024-07-08T02:08:28.974416900Z",
  "votes": 6,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I am assuming there would be a significant drop in the scores on the private LB:</p>\n<ol>\n<li>Majority of the models are not trained to handle non-triazine molecules. This could be one of the major factors lowering the scores. </li>\n<li>Models trained on subsets that are closely representative of public lb distribution might be overfitting (if private lb's distribution turns out to be explicitly different )</li>\n<li>Models not incorporating CV might be over-estimating. </li>\n<li>Ignoring data diversity (very prominent discussions there), might have been a mistake. <br>\nNot all negative samples are useless after-all, indeed! XP</li>\n</ol>\n<p>Correct me, if I am wrong thou! :)</p>",
  "messages": [
    {
      "id": "2910882",
      "postDate": "07/08/2024 02:08:28",
      "content": "<p>I am assuming there would be a significant drop in the scores on the private LB:</p>\n<ol>\n<li>Majority of the models are not trained to handle non-triazine molecules. This could be one of the major factors lowering the scores. </li>\n<li>Models trained on subsets that are closely representative of public lb distribution might be overfitting (if private lb's distribution turns out to be explicitly different )</li>\n<li>Models not incorporating CV might be over-estimating. </li>\n<li>Ignoring data diversity (very prominent discussions there), might have been a mistake. <br>\nNot all negative samples are useless after-all, indeed! XP</li>\n</ol>\n<p>Correct me, if I am wrong thou! :)</p>",
      "rawMarkdown": "I am assuming there would be a significant drop in the scores on the private LB:\n\n1. Majority of the models are not trained to handle non-triazine molecules. This could be one of the major factors lowering the scores. \n2. Models trained on subsets that are closely representative of public lb distribution might be overfitting (if private lb's distribution turns out to be explicitly different )\n3. Models not incorporating CV might be over-estimating. \n4. Ignoring data diversity (very prominent discussions there), might have been a mistake. \nNot all negative samples are useless after-all, indeed! XP\n\nCorrect me, if I am wrong thou! :)",
      "votes": null
    },
    {
      "id": "2911124",
      "postDate": "07/08/2024 05:45:19",
      "content": "<p>I didn't train with any external data (including non-triazine molecules).  I guess I would see the drop.  I am done with the competition after using up all my last day submissions.  </p>",
      "rawMarkdown": "I didn't train with any external data (including non-triazine molecules).  I guess I would see the drop.  I am done with the competition after using up all my last day submissions.",
      "votes": null
    },
    {
      "id": "2911172",
      "postDate": "07/08/2024 06:07:09",
      "content": "<p>I guess, majority of the submissions didn't train for non-triazine molecules, including mine. And, that would lead to a significant drop for everybody. For some of the top solutions, I feel are the result of training on a subset (described in few discussion posts) rather than the entire dataset. Given, there are 2/3rd unseen molecules in private lb, I doubt if some of these scores would survive the shakeup. </p>",
      "rawMarkdown": "I guess, majority of the submissions didn't train for non-triazine molecules, including mine. And, that would lead to a significant drop for everybody. For some of the top solutions, I feel are the result of training on a subset (described in few discussion posts) rather than the entire dataset. Given, there are 2/3rd unseen molecules in private lb, I doubt if some of these scores would survive the shakeup.",
      "votes": null
    },
    {
      "id": "2911207",
      "postDate": "07/08/2024 06:31:10",
      "content": "<p>Good to know.  I only separated positive samples from negative samples. I only sampled the negatives at 3.0X, 2.5X, 2.0X, 1.5X, 1.0X and 0.5X ratio relative to positive samples.  </p>",
      "rawMarkdown": "Good to know.  I only separated positive samples from negative samples. I only sampled the negatives at 3.0X, 2.5X, 2.0X, 1.5X, 1.0X and 0.5X ratio relative to positive samples.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2911124,
      "author_name": "joejeo1",
      "author_url": "",
      "post_date": "07/08/2024 05:45:19",
      "content": "<p>I didn't train with any external data (including non-triazine molecules).  I guess I would see the drop.  I am done with the competition after using up all my last day submissions.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 2911172,
          "author_name": "ahsuna123",
          "author_url": "",
          "post_date": "07/08/2024 06:07:09",
          "content": "<p>I guess, majority of the submissions didn't train for non-triazine molecules, including mine. And, that would lead to a significant drop for everybody. For some of the top solutions, I feel are the result of training on a subset (described in few discussion posts) rather than the entire dataset. Given, there are 2/3rd unseen molecules in private lb, I doubt if some of these scores would survive the shakeup. </p>",
          "votes": null,
          "replies": [
            {
              "id": 2911207,
              "author_name": "joejeo1",
              "author_url": "",
              "post_date": "07/08/2024 06:31:10",
              "content": "<p>Good to know.  I only separated positive samples from negative samples. I only sampled the negatives at 3.0X, 2.5X, 2.0X, 1.5X, 1.0X and 0.5X ratio relative to positive samples.  </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2910882": "I am assuming there would be a significant drop in the scores on the private LB:\n\n1. Majority of the models are not trained to handle non-triazine molecules. This could be one of the major factors lowering the scores. \n2. Models trained on subsets that are closely representative of public lb distribution might be overfitting (if private lb's distribution turns out to be explicitly different )\n3. Models not incorporating CV might be over-estimating. \n4. Ignoring data diversity (very prominent discussions there), might have been a mistake. \nNot all negative samples are useless after-all, indeed! XP\n\nCorrect me, if I am wrong thou! :)",
    "2911124": "I didn't train with any external data (including non-triazine molecules).  I guess I would see the drop.  I am done with the competition after using up all my last day submissions.",
    "2911172": "I guess, majority of the submissions didn't train for non-triazine molecules, including mine. And, that would lead to a significant drop for everybody. For some of the top solutions, I feel are the result of training on a subset (described in few discussion posts) rather than the entire dataset. Given, there are 2/3rd unseen molecules in private lb, I doubt if some of these scores would survive the shakeup.",
    "2911207": "Good to know.  I only separated positive samples from negative samples. I only sampled the negatives at 3.0X, 2.5X, 2.0X, 1.5X, 1.0X and 0.5X ratio relative to positive samples."
  },
  "source": "meta"
}