{
  "id": 102342,
  "title": "Seed change resulting in drastic changes in LB",
  "url": "/competitions/aptos2019-blindness-detection/discussion/102342",
  "author_name": "",
  "post_date": "2019-08-01T14:21:33.646685300Z",
  "votes": 2,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Just by changing the seed, and by not changing anything, my LB changed from 0.674 to 0.757</p>\n\n<p>How should I interpret this change? Is this much change normal? Based on this, can I infer anything about the model or about the public test set?</p>",
  "messages": [
    {
      "id": "589889",
      "postDate": "08/01/2019 14:21:33",
      "content": "<p>Just by changing the seed, and by not changing anything, my LB changed from 0.674 to 0.757</p>\n\n<p>How should I interpret this change? Is this much change normal? Based on this, can I infer anything about the model or about the public test set?</p>",
      "rawMarkdown": "Just by changing the seed, and by not changing anything, my LB changed from 0.674 to 0.757\n\nHow should I interpret this change? Is this much change normal? Based on this, can I infer anything about the model or about the public test set?",
      "votes": null
    },
    {
      "id": "589944",
      "postDate": "08/01/2019 15:36:19",
      "content": "<p>Wow thanks for info</p>",
      "rawMarkdown": "Wow thanks for info",
      "votes": null
    },
    {
      "id": "589962",
      "postDate": "08/01/2019 16:16:29",
      "content": "<p>I'm guessing you submit models based on the epoch where they scored best on a single validation set during training (I do)? If that's the case, a few potential explanations : \n- the weight initialization, which depends on the seed, is the main culprit (possible but I doubt it)\n- the train/validation split, which also depends on the seed, is the main culprit (that's what I believe right now)</p>\n\n<p>I think our models overfit on our validation set during training, and the distribution between classes of this validation set may be vastly different from the distribution of the public test set (which can itself have also a very different distribution from the private test set). For this reason, it's probably time to rethink this kind of validation scheme, in order to avoid a big surprise when the private leaderboard is revealed.</p>",
      "rawMarkdown": "I'm guessing you submit models based on the epoch where they scored best on a single validation set during training (I do)? If that's the case, a few potential explanations : \n- the weight initialization, which depends on the seed, is the main culprit (possible but I doubt it)\n- the train/validation split, which also depends on the seed, is the main culprit (that's what I believe right now)\n\nI think our models overfit on our validation set during training, and the distribution between classes of this validation set may be vastly different from the distribution of the public test set (which can itself have also a very different distribution from the private test set). For this reason, it's probably time to rethink this kind of validation scheme, in order to avoid a big surprise when the private leaderboard is revealed.",
      "votes": null
    },
    {
      "id": "590075",
      "postDate": "08/01/2019 19:17:33",
      "content": "<p>thanks <a href=\"/juliencs\">@juliencs</a> for your comments.</p>\n\n<p>Yes, I am submitting the models that scored best on a single validation set. Should I use k-fold CV to see more stable results ?</p>\n\n<ul>\n<li>Yes, I am also suspecting the weight initialization.</li>\n<li>Regarding the train/validation split, the validation set is created before the training and does not change based on the seed. This is one way for me to really measure the impact of the seeds. I am also not doing any augmentation on the validation set to avoid the any randomness there.</li>\n</ul>\n\n<p>At least I have a deterministic way to measure the randomness now. Glad that I moved from Keras to PyTorch.</p>\n\n<p>If you have any thoughts on coming up with a proper validation scheme, pls share them.</p>",
      "rawMarkdown": "thanks @juliencs for your comments.\n\nYes, I am submitting the models that scored best on a single validation set. Should I use k-fold CV to see more stable results ?\n\n- Yes, I am also suspecting the weight initialization.\n- Regarding the train/validation split, the validation set is created before the training and does not change based on the seed. This is one way for me to really measure the impact of the seeds. I am also not doing any augmentation on the validation set to avoid the any randomness there.\n\nAt least I have a deterministic way to measure the randomness now. Glad that I moved from Keras to PyTorch.\n\nIf you have any thoughts on coming up with a proper validation scheme, pls share them.",
      "votes": null
    },
    {
      "id": "611737",
      "postDate": "08/29/2019 12:06:20",
      "content": "<p>Hi Ravi,\nDid you find an optimal seed for the dataset?</p>",
      "rawMarkdown": "Hi Ravi,\nDid you find an optimal seed for the dataset?",
      "votes": null
    },
    {
      "id": "613000",
      "postDate": "08/30/2019 06:44:28",
      "content": "<p>Apart from <a href=\"/juliencs\">@juliencs</a> topics affected by the random seed, the augmentations change also with the random seed</p>\n\n<p>I'd bid for the random split, since it's an unbalanced dataset, and there are different distributions between trainset and public restart classes.</p>\n\n<p>But in the past, I noticed big changes having fixed the k-fold splits (not affected by the random seed).  I suspected then about the augmentations</p>",
      "rawMarkdown": "Apart from @juliencs topics affected by the random seed, the augmentations change also with the random seed\n\nI'd bid for the random split, since it's an unbalanced dataset, and there are different distributions between trainset and public restart classes.\n\nBut in the past, I noticed big changes having fixed the k-fold splits (not affected by the random seed).  I suspected then about the augmentations",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 589944,
      "author_name": "vaishvik25",
      "author_url": "",
      "post_date": "08/01/2019 15:36:19",
      "content": "<p>Wow thanks for info</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 589962,
      "author_name": "juliencs",
      "author_url": "",
      "post_date": "08/01/2019 16:16:29",
      "content": "<p>I'm guessing you submit models based on the epoch where they scored best on a single validation set during training (I do)? If that's the case, a few potential explanations : \n- the weight initialization, which depends on the seed, is the main culprit (possible but I doubt it)\n- the train/validation split, which also depends on the seed, is the main culprit (that's what I believe right now)</p>\n\n<p>I think our models overfit on our validation set during training, and the distribution between classes of this validation set may be vastly different from the distribution of the public test set (which can itself have also a very different distribution from the private test set). For this reason, it's probably time to rethink this kind of validation scheme, in order to avoid a big surprise when the private leaderboard is revealed.</p>",
      "votes": null,
      "replies": [
        {
          "id": 590075,
          "author_name": "ravivadapalli",
          "author_url": "",
          "post_date": "08/01/2019 19:17:33",
          "content": "<p>thanks <a href=\"/juliencs\">@juliencs</a> for your comments.</p>\n\n<p>Yes, I am submitting the models that scored best on a single validation set. Should I use k-fold CV to see more stable results ?</p>\n\n<ul>\n<li>Yes, I am also suspecting the weight initialization.</li>\n<li>Regarding the train/validation split, the validation set is created before the training and does not change based on the seed. This is one way for me to really measure the impact of the seeds. I am also not doing any augmentation on the validation set to avoid the any randomness there.</li>\n</ul>\n\n<p>At least I have a deterministic way to measure the randomness now. Glad that I moved from Keras to PyTorch.</p>\n\n<p>If you have any thoughts on coming up with a proper validation scheme, pls share them.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 611737,
          "author_name": "abyaadrafid",
          "author_url": "",
          "post_date": "08/29/2019 12:06:20",
          "content": "<p>Hi Ravi,\nDid you find an optimal seed for the dataset?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 613000,
      "author_name": "virilo",
      "author_url": "",
      "post_date": "08/30/2019 06:44:28",
      "content": "<p>Apart from <a href=\"/juliencs\">@juliencs</a> topics affected by the random seed, the augmentations change also with the random seed</p>\n\n<p>I'd bid for the random split, since it's an unbalanced dataset, and there are different distributions between trainset and public restart classes.</p>\n\n<p>But in the past, I noticed big changes having fixed the k-fold splits (not affected by the random seed).  I suspected then about the augmentations</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "589889": "Just by changing the seed, and by not changing anything, my LB changed from 0.674 to 0.757\n\nHow should I interpret this change? Is this much change normal? Based on this, can I infer anything about the model or about the public test set?",
    "589944": "Wow thanks for info",
    "589962": "I'm guessing you submit models based on the epoch where they scored best on a single validation set during training (I do)? If that's the case, a few potential explanations : \n- the weight initialization, which depends on the seed, is the main culprit (possible but I doubt it)\n- the train/validation split, which also depends on the seed, is the main culprit (that's what I believe right now)\n\nI think our models overfit on our validation set during training, and the distribution between classes of this validation set may be vastly different from the distribution of the public test set (which can itself have also a very different distribution from the private test set). For this reason, it's probably time to rethink this kind of validation scheme, in order to avoid a big surprise when the private leaderboard is revealed.",
    "590075": "thanks @juliencs for your comments.\n\nYes, I am submitting the models that scored best on a single validation set. Should I use k-fold CV to see more stable results ?\n\n- Yes, I am also suspecting the weight initialization.\n- Regarding the train/validation split, the validation set is created before the training and does not change based on the seed. This is one way for me to really measure the impact of the seeds. I am also not doing any augmentation on the validation set to avoid the any randomness there.\n\nAt least I have a deterministic way to measure the randomness now. Glad that I moved from Keras to PyTorch.\n\nIf you have any thoughts on coming up with a proper validation scheme, pls share them.",
    "611737": "Hi Ravi,\nDid you find an optimal seed for the dataset?",
    "613000": "Apart from @juliencs topics affected by the random seed, the augmentations change also with the random seed\n\nI'd bid for the random split, since it's an unbalanced dataset, and there are different distributions between trainset and public restart classes.\n\nBut in the past, I noticed big changes having fixed the k-fold splits (not affected by the random seed).  I suspected then about the augmentations"
  },
  "source": "meta"
}