{
  "id": 168253,
  "title": "How big is going to be the shake up?",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/168253",
  "author_name": "Martin Kovacevic Buvinic",
  "post_date": "2020-07-19T21:23:07.910000",
  "votes": 19,
  "comment_count": 34,
  "views": 0,
  "content": "<p>Some public kernels have a big difference between the public leadearboard and their cross validation. Using the same scripts and modify the seed gives +- 0.02 roc auc in the public leadearbord. In my opinion and experience, we should not trust in the public leadearboard and work in a stable cross validation strategy.  Understanding the problem, the distributions in the training and testing data and how they differ is the key to construct a good cv strategy.</p>\n\n<p>One way to do this is change the seed and run your experiments multiple times and check your cross validation and public leader board score, if they don't differ that much, your model is stable and hopefully you will not recieve that disturbing surprise in the privite leaderboard when the competition ends. </p>\n\n<p>If you want to post your analysis and estimate how big is going to be the shake up, you are welcome.</p>\n\n<p>My opininion is that their is going to be a big shake up, distribution between the training set and the test set are different (classic kaggle :) so a solid model that generalize well will get the gold medal. My estimation are in the range of +- 0.02 roc auc difference, unless folks make a solid cv strategy.</p>",
  "messages": [
    {
      "id": 937467,
      "postDate": "2020-07-21T03:35:51.990Z",
      "content": "<p><img src=\"https://scontent.fhan2-4.fna.fbcdn.net/v/t1.15752-9/93803848_3592968004063432_882792869419548672_n.jpg?_nc_cat=110&amp;_nc_sid=b96e70&amp;_nc_ohc=t3pI-isEIHMAX85VH1a&amp;_nc_ht=scontent.fhan2-4.fna&amp;oh=87cc570ec7ab8c10e974c3f304d40033&amp;oe=5F3B8133\" alt=\"shakeup\"></p>",
      "rawMarkdown": "![shakeup](https://scontent.fhan2-4.fna.fbcdn.net/v/t1.15752-9/93803848_3592968004063432_882792869419548672_n.jpg?_nc_cat=110&amp;_nc_sid=b96e70&amp;_nc_ohc=t3pI-isEIHMAX85VH1a&amp;_nc_ht=scontent.fhan2-4.fna&amp;oh=87cc570ec7ab8c10e974c3f304d40033&amp;oe=5F3B8133)",
      "votes": 19,
      "replies": [
        {
          "id": 937495,
          "postDate": "2020-07-21T03:54:15.947Z",
          "rawMarkdown": "",
          "votes": 3,
          "isDeleted": true
        },
        {
          "id": 937520,
          "postDate": "2020-07-21T04:18:01.933Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 937540,
          "postDate": "2020-07-21T04:30:58.777Z",
          "content": "<p>I love it when an image can capture my world view that might take 10 minutes and a shot of booze to express.</p>\n\n<p>Nice one !</p>",
          "rawMarkdown": "I love it when an image can capture my world view that might take 10 minutes and a shot of booze to express.\n\nNice one !",
          "votes": 2
        }
      ]
    },
    {
      "id": 935983,
      "postDate": "2020-07-19T21:23:07.910Z",
      "content": "<p>Some public kernels have a big difference between the public leadearboard and their cross validation. Using the same scripts and modify the seed gives +- 0.02 roc auc in the public leadearbord. In my opinion and experience, we should not trust in the public leadearboard and work in a stable cross validation strategy.  Understanding the problem, the distributions in the training and testing data and how they differ is the key to construct a good cv strategy.</p>\n\n<p>One way to do this is change the seed and run your experiments multiple times and check your cross validation and public leader board score, if they don't differ that much, your model is stable and hopefully you will not recieve that disturbing surprise in the privite leaderboard when the competition ends. </p>\n\n<p>If you want to post your analysis and estimate how big is going to be the shake up, you are welcome.</p>\n\n<p>My opininion is that their is going to be a big shake up, distribution between the training set and the test set are different (classic kaggle :) so a solid model that generalize well will get the gold medal. My estimation are in the range of +- 0.02 roc auc difference, unless folks make a solid cv strategy.</p>",
      "rawMarkdown": "Some public kernels have a big difference between the public leadearboard and their cross validation. Using the same scripts and modify the seed gives +- 0.02 roc auc in the public leadearbord. In my opinion and experience, we should not trust in the public leadearboard and work in a stable cross validation strategy.  Understanding the problem, the distributions in the training and testing data and how they differ is the key to construct a good cv strategy.\n\nOne way to do this is change the seed and run your experiments multiple times and check your cross validation and public leader board score, if they don't differ that much, your model is stable and hopefully you will not recieve that disturbing surprise in the privite leaderboard when the competition ends. \n\nIf you want to post your analysis and estimate how big is going to be the shake up, you are welcome.\n\nMy opininion is that their is going to be a big shake up, distribution between the training set and the test set are different (classic kaggle :) so a solid model that generalize well will get the gold medal. My estimation are in the range of +- 0.02 roc auc difference, unless folks make a solid cv strategy.",
      "votes": 18
    },
    {
      "id": 936297,
      "postDate": "2020-07-20T06:15:36.117Z",
      "content": "<p>It will be a lottery.</p>",
      "rawMarkdown": "It will be a lottery.",
      "votes": 8,
      "replies": [
        {
          "id": 936382,
          "postDate": "2020-07-20T07:17:45.743Z",
          "content": "<p><a href=\"https://www.kaggle.com/aerdem4\" target=\"_blank\">@aerdem4</a> not saying that I disagree with you but could you elaborate a bit more on what makes you say that?</p>",
          "rawMarkdown": "@aerdem4 not saying that I disagree with you but could you elaborate a bit more on what makes you say that?",
          "votes": 1
        },
        {
          "id": 936386,
          "postDate": "2020-07-20T07:20:08.537Z",
          "content": "<p>Number of positive classes is very few for both public and private. Also I have tried the public best kernel locally, CV and LB seem to be very different.</p>",
          "rawMarkdown": "Number of positive classes is very few for both public and private. Also I have tried the public best kernel locally, CV and LB seem to be very different.",
          "votes": 6
        },
        {
          "id": 942975,
          "postDate": "2020-07-24T04:54:15.413Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 943420,
          "postDate": "2020-07-24T10:54:57.093Z",
          "content": "<p><a href=\"/aerdem4\">@aerdem4</a>   <code>Also I have tried the public best kernel locally, CV and LB seem to be very different.</code></p>\n\n<p>i see you haven't made any submission in this competition so far,how you know LB then?</p>",
          "rawMarkdown": "@aerdem4   `Also I have tried the public best kernel locally, CV and LB seem to be very different.`\n\ni see you haven't made any submission in this competition so far,how you know LB then?",
          "replies": [
            {
              "id": 943531,
              "postDate": "2020-07-24T12:18:11.167Z",
              "content": "<p><a href=\"/mobassir\">@mobassir</a> Ahmet knows LB score because he used a <strong>public kernel</strong>, so presumably the kernel already has a score before he forked it.</p>",
              "rawMarkdown": "@mobassir Ahmet knows LB score because he used a **public kernel**, so presumably the kernel already has a score before he forked it."
            }
          ]
        },
        {
          "id": 943564,
          "postDate": "2020-07-24T12:45:54.190Z",
          "content": "<p>Because public kernel has its public score:)</p>",
          "rawMarkdown": "Because public kernel has its public score:)",
          "votes": 3
        }
      ]
    },
    {
      "id": 941966,
      "postDate": "2020-07-23T14:20:44.740Z",
      "content": "<p>if public\n<img src=\"https://sun9-71.userapi.com/c853524/v853524193/217353/eQmIMHFmXBk.jpg\" alt=\"\"></p>\n\n<p>and private\n<img src=\"https://sun9-51.userapi.com/c853524/v853524193/21735b/qc_bUvfg-x4.jpg\" alt=\"\"></p>",
      "rawMarkdown": "if public\n![](https://sun9-71.userapi.com/c853524/v853524193/217353/eQmIMHFmXBk.jpg)\n\nand private\n![](https://sun9-51.userapi.com/c853524/v853524193/21735b/qc_bUvfg-x4.jpg)",
      "votes": 5,
      "replies": [
        {
          "id": 942344,
          "postDate": "2020-07-23T17:49:10.440Z",
          "content": "<p>good one !</p>",
          "rawMarkdown": "good one !",
          "votes": 1
        }
      ]
    },
    {
      "id": 936016,
      "postDate": "2020-07-19T22:24:15.833Z",
      "content": "<p>I don't see a huge shakeup for two reasons : \n- The majority are using <a href=\"/cdeotte\">@cdeotte</a> 's stratified k-fold with no leaks. In my opinion, this is the right validation strategy.\n- Almost all people are ensembling a lot but still we have to be careful about the weights.</p>",
      "rawMarkdown": "I don't see a huge shakeup for two reasons : \n- The majority are using @cdeotte 's stratified k-fold with no leaks. In my opinion, this is the right validation strategy.\n- Almost all people are ensembling a lot but still we have to be careful about the weights.",
      "votes": 6,
      "replies": [
        {
          "id": 936398,
          "postDate": "2020-07-20T07:27:41.857Z",
          "content": "<p>If people are choosing weights based on their <strong>PublicLB</strong> scores, then there is a chance that they will badly overfit.</p>",
          "rawMarkdown": "If people are choosing weights based on their **PublicLB** scores, then there is a chance that they will badly overfit.",
          "votes": 8
        },
        {
          "id": 936468,
          "postDate": "2020-07-20T08:37:00.100Z",
          "rawMarkdown": "",
          "votes": 5,
          "isDeleted": true
        },
        {
          "id": 937161,
          "postDate": "2020-07-20T19:14:46.980Z",
          "content": "<p>When you ensemble the gap will reduce a lot, try that. When I ensemble my models I have around 0.940CV.</p>",
          "rawMarkdown": "When you ensemble the gap will reduce a lot, try that. When I ensemble my models I have around 0.940CV."
        },
        {
          "id": 937166,
          "postDate": "2020-07-20T19:18:03.180Z",
          "content": "<p><a href=\"/meemr5\">@meemr5</a> I never choose weights to maximize my LB score, and read my second point, I said we need to be careful about the weights</p>",
          "rawMarkdown": "@meemr5 I never choose weights to maximize my LB score, and read my second point, I said we need to be careful about the weights"
        },
        {
          "id": 937174,
          "postDate": "2020-07-20T19:24:20.073Z",
          "content": "<p><a href=\"/seif95\">@seif95</a> I’m just adding something in what you mentioned above. 😅\nI was talking about some people not specifically you. \nSorry if you get me wrong! 😊</p>",
          "rawMarkdown": "@seif95 I’m just adding something in what you mentioned above. 😅\nI was talking about some people not specifically you. \nSorry if you get me wrong! 😊\n",
          "votes": 1
        },
        {
          "id": 937179,
          "postDate": "2020-07-20T19:29:27.250Z",
          "content": "<p>I didn't mean to be aggressive : ) sorry for that </p>",
          "rawMarkdown": "I didn't mean to be aggressive : ) sorry for that "
        }
      ]
    },
    {
      "id": 937009,
      "postDate": "2020-07-20T16:55:02.503Z",
      "content": "<p>I quote Trump : 'It's going to be huuuuuuuge!'</p>",
      "rawMarkdown": "I quote Trump : 'It's going to be huuuuuuuge!'",
      "votes": 3
    },
    {
      "id": 936277,
      "postDate": "2020-07-20T05:50:32.973Z",
      "content": "<p>On a recent competition I jumped 1000 places up into a metal, as I had started late and ran out of time before I had the chance to over fit  my models :)</p>\n\n<p>I think those who ensemble their way into the medals right now are at huge risk of falling out.  Ensemble is giving too big a boost to not be tempted to throw all my hats into that ring.</p>\n\n<p>Those who are using <a href=\"https://www.kaggle.com/cdeotte\">chris's</a> stratified tfrecords in a single model should feel like their position is on solid ground.</p>",
      "rawMarkdown": "On a recent competition I jumped 1000 places up into a metal, as I had started late and ran out of time before I had the chance to over fit  my models :)\n\nI think those who ensemble their way into the medals right now are at huge risk of falling out.  Ensemble is giving too big a boost to not be tempted to throw all my hats into that ring.\n\nThose who are using [chris's](https://www.kaggle.com/cdeotte) stratified tfrecords in a single model should feel like their position is on solid ground.\n\n",
      "votes": 2,
      "replies": [
        {
          "id": 936410,
          "postDate": "2020-07-20T07:44:54.033Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        },
        {
          "id": 936429,
          "postDate": "2020-07-20T07:58:48.680Z",
          "content": "<p>Yes - that is the evil temptation that I am thinking about :)  At least one of my canidate's for final evaluation will be an ensemble.</p>\n\n<p>If you play with weights and adjust based on resulting leader board score than \"DANGER WILL ROBINSON\".</p>\n\n<p>I do not plan to play with weights - my kernel for generating the submission does not include weights - each version gets an equal share.  So I think my ensemble submission will be a minimum risk (but not at zero risk).  I am sure if I played with weights I could increase my LB score today.  But for this data set with such a small number of positive's,  LB adjustments of weights seems like a sure fire method to over fit the public test set.</p>\n\n<p>UPDATE:  I realize that I am a liar.  I am playing with weights.  I have a model that I ran at 32, 64, 96, 128, 224, 256, 384, and 512.  I lack the courage to include the 32 to 128 models in my final because their inclusion DID REDUCE my LB score.</p>",
          "rawMarkdown": "Yes - that is the evil temptation that I am thinking about :)  At least one of my canidate's for final evaluation will be an ensemble.\n\nIf you play with weights and adjust based on resulting leader board score than \"DANGER WILL ROBINSON\".\n\nI do not plan to play with weights - my kernel for generating the submission does not include weights - each version gets an equal share.  So I think my ensemble submission will be a minimum risk (but not at zero risk).  I am sure if I played with weights I could increase my LB score today.  But for this data set with such a small number of positive's,  LB adjustments of weights seems like a sure fire method to over fit the public test set.\n\nUPDATE:  I realize that I am a liar.  I am playing with weights.  I have a model that I ran at 32, 64, 96, 128, 224, 256, 384, and 512.  I lack the courage to include the 32 to 128 models in my final because their inclusion DID REDUCE my LB score.",
          "votes": 5
        },
        {
          "id": 936444,
          "postDate": "2020-07-20T08:19:10.710Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        },
        {
          "id": 936446,
          "postDate": "2020-07-20T08:20:28.523Z",
          "content": "<p>wow</p>",
          "rawMarkdown": "wow",
          "votes": 1
        },
        {
          "id": 936448,
          "postDate": "2020-07-20T08:23:43.023Z",
          "content": "<p><a href=\"/pcjimmmy\">@pcjimmmy</a> Use progressive training starting from smaller images,\nbut in the end use only 384+ sizes for higher score</p>",
          "rawMarkdown": "@pcjimmmy Use progressive training starting from smaller images,\nbut in the end use only 384+ sizes for higher score",
          "votes": 3
        },
        {
          "id": 936451,
          "postDate": "2020-07-20T08:25:42.797Z",
          "content": "<p><a href=\"/x2t2cx\">@x2t2cx</a> Weights being referred are 0.1 x sub1 + 0.9 x sub2 etc\ninstead of simple (sub1 + sub2)/2 (i.e. each submission gets a weight of 1/n)</p>",
          "rawMarkdown": "@x2t2cx Weights being referred are 0.1 x sub1 + 0.9 x sub2 etc\ninstead of simple (sub1 + sub2)/2 (i.e. each submission gets a weight of 1/n)",
          "votes": 3
        },
        {
          "id": 936613,
          "postDate": "2020-07-20T11:05:24.807Z",
          "rawMarkdown": "",
          "votes": 3,
          "isDeleted": true
        },
        {
          "id": 936837,
          "postDate": "2020-07-20T14:55:14.940Z",
          "content": "<p>Nothing wrong with unequal weights IF THEY ARE derived from logic and thinking about the model and the versions.  </p>\n\n<p>They become a danger when they are established based on the leader board score and adjustments to maximize by weights.  </p>\n\n<p>More of a danger when the test set distribution appears to be different than the training.  In a recent competition I was very comfortable tweaking the weights to maximize the LB score because the train and test distributions looked very similar (a rare event on Kaggle).   This dataset does have a different train vs test distribution in a couple of ways.  Tweaking weights might work but the subject of this post was the shake up.   Shake ups happen for lots of reasons - overfitting the biggest.  </p>",
          "rawMarkdown": "Nothing wrong with unequal weights IF THEY ARE derived from logic and thinking about the model and the versions.  \n\nThey become a danger when they are established based on the leader board score and adjustments to maximize by weights.  \n\nMore of a danger when the test set distribution appears to be different than the training.  In a recent competition I was very comfortable tweaking the weights to maximize the LB score because the train and test distributions looked very similar (a rare event on Kaggle).   This dataset does have a different train vs test distribution in a couple of ways.  Tweaking weights might work but the subject of this post was the shake up.   Shake ups happen for lots of reasons - overfitting the biggest.  ",
          "votes": 2
        },
        {
          "id": 937498,
          "postDate": "2020-07-21T03:56:30.160Z",
          "content": "<p><a href=\"/pcjimmmy\">@pcjimmmy</a> san, what if the weights are derived from OOF CV scores? Is that a good logic?\nEdit: I think it also depends on the usage of external data, CV can't be trusted when validation set doesn't follow the distributions in test set.</p>",
          "rawMarkdown": "@pcjimmmy san, what if the weights are derived from OOF CV scores? Is that a good logic?\nEdit: I think it also depends on the usage of external data, CV can't be trusted when validation set doesn't follow the distributions in test set."
        },
        {
          "id": 937534,
          "postDate": "2020-07-21T04:27:53.453Z",
          "content": "<p>It's one I have used for sure.  Anytime you do not use the leader board score to adjust the weights than I believe you have avoided the danger zone of over fitting to the LB.</p>\n\n<p>Since I am an old guy I have a 73 year old notion that methods I want to learn and practice should be things I could use in a real world model.  In the real world I don't get to write an App that predicts melanoma AFTER I get the results of the biopsy from the lab.   </p>\n\n<p>So I like to ask myself - Does this method work if an entire new set of test samples is presented to my model.</p>",
          "rawMarkdown": "It's one I have used for sure.  Anytime you do not use the leader board score to adjust the weights than I believe you have avoided the danger zone of over fitting to the LB.\n\nSince I am an old guy I have a 73 year old notion that methods I want to learn and practice should be things I could use in a real world model.  In the real world I don't get to write an App that predicts melanoma AFTER I get the results of the biopsy from the lab.   \n\nSo I like to ask myself - Does this method work if an entire new set of test samples is presented to my model.",
          "votes": 3
        }
      ]
    },
    {
      "id": 944040,
      "postDate": "2020-07-24T19:03:36.910Z",
      "content": "<p>Great piece of info !</p>",
      "rawMarkdown": "Great piece of info !"
    },
    {
      "id": 936421,
      "postDate": "2020-07-20T07:51:32.093Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 937467,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2020-07-21T03:35:51.990000",
      "content": "<p><img src=\"https://scontent.fhan2-4.fna.fbcdn.net/v/t1.15752-9/93803848_3592968004063432_882792869419548672_n.jpg?_nc_cat=110&amp;_nc_sid=b96e70&amp;_nc_ohc=t3pI-isEIHMAX85VH1a&amp;_nc_ht=scontent.fhan2-4.fna&amp;oh=87cc570ec7ab8c10e974c3f304d40033&amp;oe=5F3B8133\" alt=\"shakeup\"></p>",
      "votes": 19,
      "replies": [
        {
          "id": 937495,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-21T03:54:15.947000",
          "content": "",
          "votes": 3,
          "replies": []
        },
        {
          "id": 937520,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-21T04:18:01.933000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 937540,
          "author_name": "PC Jimmmy",
          "author_url": "",
          "post_date": "2020-07-21T04:30:58.777000",
          "content": "<p>I love it when an image can capture my world view that might take 10 minutes and a shot of booze to express.</p>\n\n<p>Nice one !</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 936297,
      "author_name": "Ahmet Erdem",
      "author_url": "",
      "post_date": "2020-07-20T06:15:36.117000",
      "content": "<p>It will be a lottery.</p>",
      "votes": 8,
      "replies": [
        {
          "id": 936382,
          "author_name": "Optimo",
          "author_url": "",
          "post_date": "2020-07-20T07:17:45.743000",
          "content": "<p><a href=\"https://www.kaggle.com/aerdem4\" target=\"_blank\">@aerdem4</a> not saying that I disagree with you but could you elaborate a bit more on what makes you say that?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 936386,
          "author_name": "Ahmet Erdem",
          "author_url": "",
          "post_date": "2020-07-20T07:20:08.537000",
          "content": "<p>Number of positive classes is very few for both public and private. Also I have tried the public best kernel locally, CV and LB seem to be very different.</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 942975,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-24T04:54:15.413000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 943420,
          "author_name": "Mobassir",
          "author_url": "",
          "post_date": "2020-07-24T10:54:57.093000",
          "content": "<p><a href=\"/aerdem4\">@aerdem4</a>   <code>Also I have tried the public best kernel locally, CV and LB seem to be very different.</code></p>\n\n<p>i see you haven't made any submission in this competition so far,how you know LB then?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 943531,
              "author_name": "Sirish Somanchi",
              "author_url": "",
              "post_date": "2020-07-24T12:18:11.167000",
              "content": "<p><a href=\"/mobassir\">@mobassir</a> Ahmet knows LB score because he used a <strong>public kernel</strong>, so presumably the kernel already has a score before he forked it.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 943564,
          "author_name": "Ahmet Erdem",
          "author_url": "",
          "post_date": "2020-07-24T12:45:54.190000",
          "content": "<p>Because public kernel has its public score:)</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 941966,
      "author_name": "Innat",
      "author_url": "",
      "post_date": "2020-07-23T14:20:44.740000",
      "content": "<p>if public\n<img src=\"https://sun9-71.userapi.com/c853524/v853524193/217353/eQmIMHFmXBk.jpg\" alt=\"\"></p>\n\n<p>and private\n<img src=\"https://sun9-51.userapi.com/c853524/v853524193/21735b/qc_bUvfg-x4.jpg\" alt=\"\"></p>",
      "votes": 5,
      "replies": [
        {
          "id": 942344,
          "author_name": "Alin Cijov",
          "author_url": "",
          "post_date": "2020-07-23T17:49:10.440000",
          "content": "<p>good one !</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 936016,
      "author_name": "Seifeddine Fezzani",
      "author_url": "",
      "post_date": "2020-07-19T22:24:15.833000",
      "content": "<p>I don't see a huge shakeup for two reasons : \n- The majority are using <a href=\"/cdeotte\">@cdeotte</a> 's stratified k-fold with no leaks. In my opinion, this is the right validation strategy.\n- Almost all people are ensembling a lot but still we have to be careful about the weights.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 936398,
          "author_name": "Meet Ranoliya",
          "author_url": "",
          "post_date": "2020-07-20T07:27:41.857000",
          "content": "<p>If people are choosing weights based on their <strong>PublicLB</strong> scores, then there is a chance that they will badly overfit.</p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 936468,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-20T08:37:00.100000",
          "content": "",
          "votes": 5,
          "replies": []
        },
        {
          "id": 937161,
          "author_name": "Seifeddine Fezzani",
          "author_url": "",
          "post_date": "2020-07-20T19:14:46.980000",
          "content": "<p>When you ensemble the gap will reduce a lot, try that. When I ensemble my models I have around 0.940CV.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 937166,
          "author_name": "Seifeddine Fezzani",
          "author_url": "",
          "post_date": "2020-07-20T19:18:03.180000",
          "content": "<p><a href=\"/meemr5\">@meemr5</a> I never choose weights to maximize my LB score, and read my second point, I said we need to be careful about the weights</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 937174,
          "author_name": "Meet Ranoliya",
          "author_url": "",
          "post_date": "2020-07-20T19:24:20.073000",
          "content": "<p><a href=\"/seif95\">@seif95</a> I’m just adding something in what you mentioned above. 😅\nI was talking about some people not specifically you. \nSorry if you get me wrong! 😊</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 937179,
          "author_name": "Seifeddine Fezzani",
          "author_url": "",
          "post_date": "2020-07-20T19:29:27.250000",
          "content": "<p>I didn't mean to be aggressive : ) sorry for that </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 937009,
      "author_name": "Alin Cijov",
      "author_url": "",
      "post_date": "2020-07-20T16:55:02.503000",
      "content": "<p>I quote Trump : 'It's going to be huuuuuuuge!'</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 936277,
      "author_name": "PC Jimmmy",
      "author_url": "",
      "post_date": "2020-07-20T05:50:32.973000",
      "content": "<p>On a recent competition I jumped 1000 places up into a metal, as I had started late and ran out of time before I had the chance to over fit  my models :)</p>\n\n<p>I think those who ensemble their way into the medals right now are at huge risk of falling out.  Ensemble is giving too big a boost to not be tempted to throw all my hats into that ring.</p>\n\n<p>Those who are using <a href=\"https://www.kaggle.com/cdeotte\">chris's</a> stratified tfrecords in a single model should feel like their position is on solid ground.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 936410,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-20T07:44:54.033000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 936429,
          "author_name": "PC Jimmmy",
          "author_url": "",
          "post_date": "2020-07-20T07:58:48.680000",
          "content": "<p>Yes - that is the evil temptation that I am thinking about :)  At least one of my canidate's for final evaluation will be an ensemble.</p>\n\n<p>If you play with weights and adjust based on resulting leader board score than \"DANGER WILL ROBINSON\".</p>\n\n<p>I do not plan to play with weights - my kernel for generating the submission does not include weights - each version gets an equal share.  So I think my ensemble submission will be a minimum risk (but not at zero risk).  I am sure if I played with weights I could increase my LB score today.  But for this data set with such a small number of positive's,  LB adjustments of weights seems like a sure fire method to over fit the public test set.</p>\n\n<p>UPDATE:  I realize that I am a liar.  I am playing with weights.  I have a model that I ran at 32, 64, 96, 128, 224, 256, 384, and 512.  I lack the courage to include the 32 to 128 models in my final because their inclusion DID REDUCE my LB score.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 936444,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-20T08:19:10.710000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 936446,
          "author_name": "masoud parpanchi",
          "author_url": "",
          "post_date": "2020-07-20T08:20:28.523000",
          "content": "<p>wow</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 936448,
          "author_name": "Sirish Somanchi",
          "author_url": "",
          "post_date": "2020-07-20T08:23:43.023000",
          "content": "<p><a href=\"/pcjimmmy\">@pcjimmmy</a> Use progressive training starting from smaller images,\nbut in the end use only 384+ sizes for higher score</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 936451,
          "author_name": "Sirish Somanchi",
          "author_url": "",
          "post_date": "2020-07-20T08:25:42.797000",
          "content": "<p><a href=\"/x2t2cx\">@x2t2cx</a> Weights being referred are 0.1 x sub1 + 0.9 x sub2 etc\ninstead of simple (sub1 + sub2)/2 (i.e. each submission gets a weight of 1/n)</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 936613,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-20T11:05:24.807000",
          "content": "",
          "votes": 3,
          "replies": []
        },
        {
          "id": 936837,
          "author_name": "PC Jimmmy",
          "author_url": "",
          "post_date": "2020-07-20T14:55:14.940000",
          "content": "<p>Nothing wrong with unequal weights IF THEY ARE derived from logic and thinking about the model and the versions.  </p>\n\n<p>They become a danger when they are established based on the leader board score and adjustments to maximize by weights.  </p>\n\n<p>More of a danger when the test set distribution appears to be different than the training.  In a recent competition I was very comfortable tweaking the weights to maximize the LB score because the train and test distributions looked very similar (a rare event on Kaggle).   This dataset does have a different train vs test distribution in a couple of ways.  Tweaking weights might work but the subject of this post was the shake up.   Shake ups happen for lots of reasons - overfitting the biggest.  </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 937498,
          "author_name": "spoon spoon",
          "author_url": "",
          "post_date": "2020-07-21T03:56:30.160000",
          "content": "<p><a href=\"/pcjimmmy\">@pcjimmmy</a> san, what if the weights are derived from OOF CV scores? Is that a good logic?\nEdit: I think it also depends on the usage of external data, CV can't be trusted when validation set doesn't follow the distributions in test set.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 937534,
          "author_name": "PC Jimmmy",
          "author_url": "",
          "post_date": "2020-07-21T04:27:53.453000",
          "content": "<p>It's one I have used for sure.  Anytime you do not use the leader board score to adjust the weights than I believe you have avoided the danger zone of over fitting to the LB.</p>\n\n<p>Since I am an old guy I have a 73 year old notion that methods I want to learn and practice should be things I could use in a real world model.  In the real world I don't get to write an App that predicts melanoma AFTER I get the results of the biopsy from the lab.   </p>\n\n<p>So I like to ask myself - Does this method work if an entire new set of test samples is presented to my model.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 944040,
      "author_name": "Nitesh Chaudhry",
      "author_url": "",
      "post_date": "2020-07-24T19:03:36.910000",
      "content": "<p>Great piece of info !</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 936421,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-20T07:51:32.093000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "937467": "![shakeup](https://scontent.fhan2-4.fna.fbcdn.net/v/t1.15752-9/93803848_3592968004063432_882792869419548672_n.jpg?_nc_cat=110&amp;_nc_sid=b96e70&amp;_nc_ohc=t3pI-isEIHMAX85VH1a&amp;_nc_ht=scontent.fhan2-4.fna&amp;oh=87cc570ec7ab8c10e974c3f304d40033&amp;oe=5F3B8133)",
    "935983": "Some public kernels have a big difference between the public leadearboard and their cross validation. Using the same scripts and modify the seed gives +- 0.02 roc auc in the public leadearbord. In my opinion and experience, we should not trust in the public leadearboard and work in a stable cross validation strategy.  Understanding the problem, the distributions in the training and testing data and how they differ is the key to construct a good cv strategy.\n\nOne way to do this is change the seed and run your experiments multiple times and check your cross validation and public leader board score, if they don't differ that much, your model is stable and hopefully you will not recieve that disturbing surprise in the privite leaderboard when the competition ends. \n\nIf you want to post your analysis and estimate how big is going to be the shake up, you are welcome.\n\nMy opininion is that their is going to be a big shake up, distribution between the training set and the test set are different (classic kaggle :) so a solid model that generalize well will get the gold medal. My estimation are in the range of +- 0.02 roc auc difference, unless folks make a solid cv strategy.",
    "936297": "It will be a lottery.",
    "941966": "if public\n![](https://sun9-71.userapi.com/c853524/v853524193/217353/eQmIMHFmXBk.jpg)\n\nand private\n![](https://sun9-51.userapi.com/c853524/v853524193/21735b/qc_bUvfg-x4.jpg)",
    "936016": "I don't see a huge shakeup for two reasons : \n- The majority are using @cdeotte 's stratified k-fold with no leaks. In my opinion, this is the right validation strategy.\n- Almost all people are ensembling a lot but still we have to be careful about the weights.",
    "937009": "I quote Trump : 'It's going to be huuuuuuuge!'",
    "936277": "On a recent competition I jumped 1000 places up into a metal, as I had started late and ran out of time before I had the chance to over fit  my models :)\n\nI think those who ensemble their way into the medals right now are at huge risk of falling out.  Ensemble is giving too big a boost to not be tempted to throw all my hats into that ring.\n\nThose who are using [chris's](https://www.kaggle.com/cdeotte) stratified tfrecords in a single model should feel like their position is on solid ground.\n\n",
    "944040": "Great piece of info !",
    "936421": ""
  }
}