{
  "id": 273748,
  "title": "What is a lucky run?Is it good to pick a lucky run notebook for final 2 submissions?",
  "url": "/competitions/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/273748",
  "author_name": "",
  "post_date": "2021-09-22T10:05:13.261147300Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>In some discussions I found people saying that their high LB score is just a product of a \"lucky run, the model does not learn anything when initialized with different seeds.\", specifically in this notebook (<a href=\"https://www.kaggle.com/rluethy/efficientnet3d-with-one-mri-type\" target=\"_blank\">https://www.kaggle.com/rluethy/efficientnet3d-with-one-mri-type</a>) changing seeds make a huge difference in LB, so when picking the best 2 submissions for the final should I pick a lucky notebook? Or should I believe in my other submissions which work well even when seeds are different?</p>\n<p>Just another question: Suppose you are tasked to deploy a model for a real-time prediction and you have two models one is of the lucky run which gives better accuracy and lesser loss and the other one works well for different seeds but with less accuracy and higher loss which one would you pick?</p>",
  "messages": [
    {
      "id": "1520400",
      "postDate": "09/22/2021 10:05:13",
      "content": "<p>In some discussions I found people saying that their high LB score is just a product of a \"lucky run, the model does not learn anything when initialized with different seeds.\", specifically in this notebook (<a href=\"https://www.kaggle.com/rluethy/efficientnet3d-with-one-mri-type\" target=\"_blank\">https://www.kaggle.com/rluethy/efficientnet3d-with-one-mri-type</a>) changing seeds make a huge difference in LB, so when picking the best 2 submissions for the final should I pick a lucky notebook? Or should I believe in my other submissions which work well even when seeds are different?</p>\n<p>Just another question: Suppose you are tasked to deploy a model for a real-time prediction and you have two models one is of the lucky run which gives better accuracy and lesser loss and the other one works well for different seeds but with less accuracy and higher loss which one would you pick?</p>",
      "rawMarkdown": "In some discussions I found people saying that their high LB score is just a product of a \"lucky run, the model does not learn anything when initialized with different seeds.\", specifically in this notebook (https://www.kaggle.com/rluethy/efficientnet3d-with-one-mri-type) changing seeds make a huge difference in LB, so when picking the best 2 submissions for the final should I pick a lucky notebook? Or should I believe in my other submissions which work well even when seeds are different?\n\nJust another question: Suppose you are tasked to deploy a model for a real-time prediction and you have two models one is of the lucky run which gives better accuracy and lesser loss and the other one works well for different seeds but with less accuracy and higher loss which one would you pick?",
      "votes": null
    },
    {
      "id": "1521095",
      "postDate": "09/22/2021 21:58:34",
      "content": "<p>My recommendation, having experience in previous competitions, is don't pick the lucky run. That notebook has over 20 versions, and he likely submitted it multiple times on different seeds/hyperparams. You are only seeing the best score across all the submissions.</p>\n<p>You want to choose the most robust model that will generalize well to the private LB</p>",
      "rawMarkdown": "My recommendation, having experience in previous competitions, is don't pick the lucky run. That notebook has over 20 versions, and he likely submitted it multiple times on different seeds/hyperparams. You are only seeing the best score across all the submissions.\n\nYou want to choose the most robust model that will generalize well to the private LB",
      "votes": null
    },
    {
      "id": "1521343",
      "postDate": "09/23/2021 06:32:59",
      "content": "<p>I agree. A \"lucky\" run on the public LB doesn't necessarily translate to private LB result. What if it's a lucky run on the public LB but not so lucky on the private LB?</p>\n<p>But then again, the CV score is not reliable either because the dataset is too small…</p>",
      "rawMarkdown": "I agree. A \"lucky\" run on the public LB doesn't necessarily translate to private LB result. What if it's a lucky run on the public LB but not so lucky on the private LB?\n\nBut then again, the CV score is not reliable either because the dataset is too small...",
      "votes": null
    },
    {
      "id": "1524239",
      "postDate": "09/26/2021 09:36:12",
      "content": "<p>I think you shouldn't because in some cases you can general and choose the best fit model is better than lucky choosing</p>",
      "rawMarkdown": "I think you shouldn't because in some cases you can general and choose the best fit model is better than lucky choosing",
      "votes": null
    },
    {
      "id": "1525229",
      "postDate": "09/27/2021 09:06:45",
      "content": "<p>It's not just robustness - it should have a logical ground. The mentioned notebook just picks 64 images from the middle (what is the rationale?) and what it predicts is around 0.5. I saw another notebook that scores more than 0.7 - by looking at the submission, I understood the reason: all the predictions were around 0.29. (probably only 20% of the test set is positive).<br>\nTo me, a good model shall deliver results close to 0.0 and 1.0</p>",
      "rawMarkdown": "It's not just robustness - it should have a logical ground. The mentioned notebook just picks 64 images from the middle (what is the rationale?) and what it predicts is around 0.5. I saw another notebook that scores more than 0.7 - by looking at the submission, I understood the reason: all the predictions were around 0.29. (probably only 20% of the test set is positive).\nTo me, a good model shall deliver results close to 0.0 and 1.0",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1521095,
      "author_name": "returnofsputnik",
      "author_url": "",
      "post_date": "09/22/2021 21:58:34",
      "content": "<p>My recommendation, having experience in previous competitions, is don't pick the lucky run. That notebook has over 20 versions, and he likely submitted it multiple times on different seeds/hyperparams. You are only seeing the best score across all the submissions.</p>\n<p>You want to choose the most robust model that will generalize well to the private LB</p>",
      "votes": null,
      "replies": [
        {
          "id": 1521343,
          "author_name": "nanguyen",
          "author_url": "",
          "post_date": "09/23/2021 06:32:59",
          "content": "<p>I agree. A \"lucky\" run on the public LB doesn't necessarily translate to private LB result. What if it's a lucky run on the public LB but not so lucky on the private LB?</p>\n<p>But then again, the CV score is not reliable either because the dataset is too small…</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1524239,
      "author_name": "huyquoctrinh",
      "author_url": "",
      "post_date": "09/26/2021 09:36:12",
      "content": "<p>I think you shouldn't because in some cases you can general and choose the best fit model is better than lucky choosing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1525229,
      "author_name": "jhasanov",
      "author_url": "",
      "post_date": "09/27/2021 09:06:45",
      "content": "<p>It's not just robustness - it should have a logical ground. The mentioned notebook just picks 64 images from the middle (what is the rationale?) and what it predicts is around 0.5. I saw another notebook that scores more than 0.7 - by looking at the submission, I understood the reason: all the predictions were around 0.29. (probably only 20% of the test set is positive).<br>\nTo me, a good model shall deliver results close to 0.0 and 1.0</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1520400": "In some discussions I found people saying that their high LB score is just a product of a \"lucky run, the model does not learn anything when initialized with different seeds.\", specifically in this notebook (https://www.kaggle.com/rluethy/efficientnet3d-with-one-mri-type) changing seeds make a huge difference in LB, so when picking the best 2 submissions for the final should I pick a lucky notebook? Or should I believe in my other submissions which work well even when seeds are different?\n\nJust another question: Suppose you are tasked to deploy a model for a real-time prediction and you have two models one is of the lucky run which gives better accuracy and lesser loss and the other one works well for different seeds but with less accuracy and higher loss which one would you pick?",
    "1521095": "My recommendation, having experience in previous competitions, is don't pick the lucky run. That notebook has over 20 versions, and he likely submitted it multiple times on different seeds/hyperparams. You are only seeing the best score across all the submissions.\n\nYou want to choose the most robust model that will generalize well to the private LB",
    "1521343": "I agree. A \"lucky\" run on the public LB doesn't necessarily translate to private LB result. What if it's a lucky run on the public LB but not so lucky on the private LB?\n\nBut then again, the CV score is not reliable either because the dataset is too small...",
    "1524239": "I think you shouldn't because in some cases you can general and choose the best fit model is better than lucky choosing",
    "1525229": "It's not just robustness - it should have a logical ground. The mentioned notebook just picks 64 images from the middle (what is the rationale?) and what it predicts is around 0.5. I saw another notebook that scores more than 0.7 - by looking at the submission, I understood the reason: all the predictions were around 0.29. (probably only 20% of the test set is positive).\nTo me, a good model shall deliver results close to 0.0 and 1.0"
  },
  "source": "meta"
}