{
  "id": 171525,
  "title": "You May Choose 3 Final Submissions",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/171525",
  "author_name": "Chris Deotte",
  "post_date": "2020-08-01T08:26:17.557000",
  "votes": 46,
  "comment_count": 50,
  "views": 0,
  "content": "<p>I would like to point out that this competition allows us to select 3 final submissions <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/submissions\">here</a>\n&gt; You may select up to 3 submissions to be used to count towards your final leaderboard score. </p>\n\n<p>Keep this in mind as it affects final model building strategy and how we should invest our time during the last 2 weeks. Good luck everyone!</p>",
  "messages": [
    {
      "id": 953900,
      "postDate": "2020-08-01T08:26:17.557Z",
      "content": "<p>I would like to point out that this competition allows us to select 3 final submissions <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/submissions\">here</a>\n&gt; You may select up to 3 submissions to be used to count towards your final leaderboard score. </p>\n\n<p>Keep this in mind as it affects final model building strategy and how we should invest our time during the last 2 weeks. Good luck everyone!</p>",
      "rawMarkdown": "I would like to point out that this competition allows us to select 3 final submissions [here][1]\n&gt; You may select up to 3 submissions to be used to count towards your final leaderboard score. \n\nKeep this in mind as it affects final model building strategy and how we should invest our time during the last 2 weeks. Good luck everyone!\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/submissions",
      "votes": 46
    },
    {
      "id": 954129,
      "postDate": "2020-08-01T12:09:02.697Z",
      "content": "<ul>\n<li>Submission 1 : Without patient_id context: Triple-Stratified-CV ensemble with EfficientNet-B5/6/7, FixEfficientNet-B5/6/7, random augmentations, seeds</li>\n<li>Submission 2 : With patient_id context: Submission 1 + tabular data (NN/TabNet/LGBM/XGB)</li>\n<li>Submission 3 : Without patient_id context: Experimental custom NN using only images</li>\n</ul>",
      "rawMarkdown": "- Submission 1 : Without patient_id context: Triple-Stratified-CV ensemble with EfficientNet-B5/6/7, FixEfficientNet-B5/6/7, random augmentations, seeds\n- Submission 2 : With patient_id context: Submission 1 + tabular data (NN/TabNet/LGBM/XGB)\n- Submission 3 : Without patient_id context: Experimental custom NN using only images",
      "votes": 5,
      "replies": [
        {
          "id": 954916,
          "postDate": "2020-08-02T07:38:51.860Z",
          "content": "<p><a href=\"/sirishks\">@sirishks</a> what is this FixEfficientNet??Could you explain?\nWhere can i find the implementation??</p>",
          "rawMarkdown": "@sirishks what is this FixEfficientNet??Could you explain?\nWhere can i find the implementation??"
        },
        {
          "id": 955269,
          "postDate": "2020-08-02T14:00:57.260Z",
          "content": "<p>Your scores getting more and more unreal :) I am excited to see how they will perform on the private lb - good luck!</p>",
          "rawMarkdown": "Your scores getting more and more unreal :) I am excited to see how they will perform on the private lb - good luck!",
          "votes": 2
        },
        {
          "id": 958233,
          "postDate": "2020-08-04T21:21:08.303Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 955402,
      "postDate": "2020-08-02T15:52:09.100Z",
      "content": "<p>Choose 3 Final Submissions, one more option for overfitting with public LB (increase the probability of winning the lottery) :))</p>",
      "rawMarkdown": "Choose 3 Final Submissions, one more option for overfitting with public LB (increase the probability of winning the lottery) :))",
      "votes": 2
    },
    {
      "id": 954395,
      "postDate": "2020-08-01T17:45:06.277Z",
      "content": "<p>I have a question: <strong>Are submissions made of public kernels qualified as legit at the end of competition?</strong> It seems to me that if the answer is yes, then it is kind of unfair. I always thought that public kernels are for fair exchange of ideas and should not be used to make submissions except for the author of kernel.</p>",
      "rawMarkdown": "I have a question: **Are submissions made of public kernels qualified as legit at the end of competition?** It seems to me that if the answer is yes, then it is kind of unfair. I always thought that public kernels are for fair exchange of ideas and should not be used to make submissions except for the author of kernel.",
      "votes": 3,
      "replies": [
        {
          "id": 954398,
          "postDate": "2020-08-01T17:47:24.880Z",
          "content": "<p>They are fine to use, they are released under an open source license.</p>",
          "rawMarkdown": "They are fine to use, they are released under an open source license.",
          "votes": 5
        },
        {
          "id": 954399,
          "postDate": "2020-08-01T17:47:36.583Z",
          "content": "<p>Using public kernels is allowed. You can either use the <code>submission.csv</code> file directly, or you can fork the notebook and run the exact same code again and use the new <code>submission.csv</code>.</p>",
          "rawMarkdown": "Using public kernels is allowed. You can either use the `submission.csv` file directly, or you can fork the notebook and run the exact same code again and use the new `submission.csv`.",
          "votes": 2
        },
        {
          "id": 954421,
          "postDate": "2020-08-01T18:05:36.410Z",
          "content": "<p>Thank you <a href=\"/fchmiel\">@fchmiel</a> and <a href=\"/cdeotte\">@cdeotte</a>  for clarifying it for me. Still strange that it is allowed though.</p>",
          "rawMarkdown": "Thank you @fchmiel and @cdeotte  for clarifying it for me. Still strange that it is allowed though.",
          "votes": 2
        }
      ]
    },
    {
      "id": 953937,
      "postDate": "2020-08-01T08:55:32.480Z",
      "content": "<p>Maybe, the hosts realized beforehand that the shakeup is inevitable (with two submissions), so they are giving participants more chances.</p>",
      "rawMarkdown": "Maybe, the hosts realized beforehand that the shakeup is inevitable (with two submissions), so they are giving participants more chances.",
      "votes": 1,
      "replies": [
        {
          "id": 954504,
          "postDate": "2020-08-01T19:29:39.603Z",
          "content": "<p>They are also giving more hope which may lead to more disheartening after loosing ranks in shakeup.😅 😂 </p>",
          "rawMarkdown": "They are also giving more hope which may lead to more disheartening after loosing ranks in shakeup.😅 😂 ",
          "votes": 1
        }
      ]
    },
    {
      "id": 953955,
      "postDate": "2020-08-01T09:14:03.360Z",
      "content": "<p>3 submissions are necessary to manage LB + 2 special prizes (Top model With and Without Context)</p>\n\n<blockquote>\n  <p>Eligible submissions for these Special Prizes must be a selected submission, for which each team is permitted 3 in this competition. As such, teams may consider diversifying their submission selections, if they also wish to be considered for either special prize.</p>\n</blockquote>\n\n<p><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/overview/prizes\">https://www.kaggle.com/c/siim-isic-melanoma-classification/overview/prizes</a></p>",
      "rawMarkdown": "3 submissions are necessary to manage LB + 2 special prizes (Top model With and Without Context)\n\n&gt; Eligible submissions for these Special Prizes must be a selected submission, for which each team is permitted 3 in this competition. As such, teams may consider diversifying their submission selections, if they also wish to be considered for either special prize.\n\nhttps://www.kaggle.com/c/siim-isic-melanoma-classification/overview/prizes",
      "votes": 2
    },
    {
      "id": 953934,
      "postDate": "2020-08-01T08:51:39.343Z",
      "content": "<p>Wow, how fo you feel about it?</p>",
      "rawMarkdown": "Wow, how fo you feel about it?",
      "votes": 2,
      "replies": [
        {
          "id": 954036,
          "postDate": "2020-08-01T10:23:48.597Z",
          "content": "<p>It will add more luck to the final results. </p>\n\n<p>This comp will have an interesting final leaderboard. I think many strong competitors are carefully maximizing a triple stratified CV. My overall CV score is around AUC 0.950. There will be a cluster of private LB scores from these careful models.</p>\n\n<p>The question is, where will all the random public notebook ensembles finish? How many will finish higher than a careful CV and how many lower? I'm not sure. With everyone having 3 final submissions, a careful CV model may not be able to reach the final top 50.</p>",
          "rawMarkdown": "It will add more luck to the final results. \n\nThis comp will have an interesting final leaderboard. I think many strong competitors are carefully maximizing a triple stratified CV. My overall CV score is around AUC 0.950. There will be a cluster of private LB scores from these careful models.\n\nThe question is, where will all the random public notebook ensembles finish? How many will finish higher than a careful CV and how many lower? I'm not sure. With everyone having 3 final submissions, a careful CV model may not be able to reach the final top 50.",
          "votes": 10
        },
        {
          "id": 954100,
          "postDate": "2020-08-01T11:36:38.607Z",
          "content": "<p>Yes, I am not really happy about that. It makes it really easy to overfit public LB and choose it as one of the three final subs. This means that there is a higher chance in the end that some top scores are rather lucky ones, which might be a disadvantage to robust solutions as you say.</p>",
          "rawMarkdown": "Yes, I am not really happy about that. It makes it really easy to overfit public LB and choose it as one of the three final subs. This means that there is a higher chance in the end that some top scores are rather lucky ones, which might be a disadvantage to robust solutions as you say.",
          "votes": 4
        },
        {
          "id": 954264,
          "postDate": "2020-08-01T15:13:31.920Z",
          "content": "<p>This is an interesting conversation and question that seems to come up often. It got me thinking about how the competition design might make it more or less susceptible to \"luck\". Initially, I thought that since everyone is given the same number of submissions- then the variance in results per participant would decrease as the number of final submissions increases, which would decrease \"luck\". I know M5 having a 1 submission limit made many complained afterwards that the results were more a result of luck.</p>\n\n<p>All this led me down the rabbit hole this morning and found <a href=\"https://pdfs.semanticscholar.org/5e31/2f4f821eb0f0da538cebb18ff0cbe29f9735.pdf?_ga=2.232360031.1105651826.1596292464-934602416.1596292464\">this paper studying kaggle tournament design: </a> they discuss the impacts of the financial reward and size of the competition- but I don't see anywhere that they specifically discuss the impact of submission limits. It might be a fun thing to investigate using the meta kaggle dataset.</p>\n\n<p>I guess the definition of \"luck\" here is key. Some might say if there is a big shakeup then the competition involved more luck, others say that if the public LB can be overfit and is closely correlated to private then a lucky submission is one based only on LB and not CV. I'm still not convinced of either. I think a solution based solely on public LB is foolish because it may be overfit, but a submission based only on CV might ignore important signal that can be gained from the public LB. This especially is important when the test set isn't a random sample of the same data as the provided training data. For instance the \"mystery\" images in this competition means that information can be gained from the public LB that can't be gained from the training data alone.</p>\n\n<p>I don't think having 3 submissions instead of 2 will really change the number of people submitting public kernel blends. There will always be a substantial number of people that will do that regardless of the limit.</p>",
          "rawMarkdown": "This is an interesting conversation and question that seems to come up often. It got me thinking about how the competition design might make it more or less susceptible to \"luck\". Initially, I thought that since everyone is given the same number of submissions- then the variance in results per participant would decrease as the number of final submissions increases, which would decrease \"luck\". I know M5 having a 1 submission limit made many complained afterwards that the results were more a result of luck.\n\nAll this led me down the rabbit hole this morning and found [this paper studying kaggle tournament design: ](https://pdfs.semanticscholar.org/5e31/2f4f821eb0f0da538cebb18ff0cbe29f9735.pdf?_ga=2.232360031.1105651826.1596292464-934602416.1596292464) they discuss the impacts of the financial reward and size of the competition- but I don't see anywhere that they specifically discuss the impact of submission limits. It might be a fun thing to investigate using the meta kaggle dataset.\n\nI guess the definition of \"luck\" here is key. Some might say if there is a big shakeup then the competition involved more luck, others say that if the public LB can be overfit and is closely correlated to private then a lucky submission is one based only on LB and not CV. I'm still not convinced of either. I think a solution based solely on public LB is foolish because it may be overfit, but a submission based only on CV might ignore important signal that can be gained from the public LB. This especially is important when the test set isn't a random sample of the same data as the provided training data. For instance the \"mystery\" images in this competition means that information can be gained from the public LB that can't be gained from the training data alone.\n\nI don't think having 3 submissions instead of 2 will really change the number of people submitting public kernel blends. There will always be a substantial number of people that will do that regardless of the limit.",
          "votes": 3
        },
        {
          "id": 954282,
          "postDate": "2020-08-01T15:44:25.773Z",
          "content": "<p>In general more submissions doesn't necessarily add more luck, but this competition is different (because of high model variance).</p>\n\n<p>Because the metric is AUC and because there are few positives, AUC fluctuates a lot. Imagine everyone has a model with average LB 0.950 and if they train and submit the same model 100 times, the LB may range from 0.935 to 0.965 (i.e. STD = 0.005). In other competitions the LB may only range from 0.949 to 0.951 (low model variance). Since STD is 0.005, they have a 1% chance (based on z score table) of scoring 0.963+.</p>\n\n<p>If 3000 participants have 3 subs each, that's 9000 subs. If everyone repeatedly trains this same model, then 90 subs would score over 0.963+. </p>\n\n<p>Now let's say one person has a model with average LB 0.955 (and STD 0.005). It appears to be a better model, but they would need 0.965+ to place in top 50. Therefore they only have 2.5% chance (based on z score table) of placing in top 50 even though they have a better model with 1 sub and 7.3% chance with 3 subs.</p>\n\n<p>If all participants had 2 subs each, then only 60 subs would be over 0.963+ and the better model would have 10% chance of being in top 50 with 1 sub and 27% chance with 3 subs.</p>\n\n<p>EDIT: This argument is only valid when the better model also has lower STD than the repeatedly submitted model.</p>",
          "rawMarkdown": "In general more submissions doesn't necessarily add more luck, but this competition is different (because of high model variance).\n\nBecause the metric is AUC and because there are few positives, AUC fluctuates a lot. Imagine everyone has a model with average LB 0.950 and if they train and submit the same model 100 times, the LB may range from 0.935 to 0.965 (i.e. STD = 0.005). In other competitions the LB may only range from 0.949 to 0.951 (low model variance). Since STD is 0.005, they have a 1% chance (based on z score table) of scoring 0.963+.\n\nIf 3000 participants have 3 subs each, that's 9000 subs. If everyone repeatedly trains this same model, then 90 subs would score over 0.963+. \n\nNow let's say one person has a model with average LB 0.955 (and STD 0.005). It appears to be a better model, but they would need 0.965+ to place in top 50. Therefore they only have 2.5% chance (based on z score table) of placing in top 50 even though they have a better model with 1 sub and 7.3% chance with 3 subs.\n\nIf all participants had 2 subs each, then only 60 subs would be over 0.963+ and the better model would have 10% chance of being in top 50 with 1 sub and 27% chance with 3 subs.\n\nEDIT: This argument is only valid when the better model also has lower STD than the repeatedly submitted model.",
          "votes": 8
        },
        {
          "id": 954328,
          "postDate": "2020-08-01T16:27:29.247Z",
          "content": "<p>🤔 Interesting <a href=\"/cdeotte\">@cdeotte</a> - I'll have to think about that. I follow you up until this point:</p>\n\n<blockquote>\n  <p>Now let's say one person has a model with average LB 0.955 (and STD 0.05). It appears to be a better model, but they would need 0.965+ to place in top 50. Therefore they only have 2.5% chance (based on z score table) of placing in top 50 even though they have a better model.</p>\n</blockquote>\n\n<p>Because the person with average LB model of 0.955 and STD 0.05 gets 3 subs also. I might try to simulate this 😄 </p>",
          "rawMarkdown": "🤔 Interesting @cdeotte - I'll have to think about that. I follow you up until this point:\n\n&gt;  Now let's say one person has a model with average LB 0.955 (and STD 0.05). It appears to be a better model, but they would need 0.965+ to place in top 50. Therefore they only have 2.5% chance (based on z score table) of placing in top 50 even though they have a better model.\n\nBecause the person with average LB model of 0.955 and STD 0.05 gets 3 subs also. I might try to simulate this 😄 ",
          "votes": 1
        },
        {
          "id": 954330,
          "postDate": "2020-08-01T16:31:51.383Z",
          "content": "<p><a href=\"/robikscube\">@robikscube</a>\n&gt; Because the person with average LB model of 0.955 and STD 0.05 gets 3 subs also</p>\n\n<p>Yes, i forgot about their 3 subs. The better model has a 7.3% chance of placing in top 50 with 3 subs. (because they have <code>97.5 ** 3 = 92.7%</code> chance of not placing in top 50 with 3 subs).</p>",
          "rawMarkdown": "@robikscube\n&gt; Because the person with average LB model of 0.955 and STD 0.05 gets 3 subs also\n\nYes, i forgot about their 3 subs. The better model has a 7.3% chance of placing in top 50 with 3 subs. (because they have `97.5 ** 3 = 92.7%` chance of not placing in top 50 with 3 subs)."
        },
        {
          "id": 954332,
          "postDate": "2020-08-01T16:34:14.657Z",
          "content": "<p>But the question is about if 1 vs 2 vs 3 subs makes a difference right? I honestly have no idea, putting together a simulation right now to see what I find.</p>",
          "rawMarkdown": "But the question is about if 1 vs 2 vs 3 subs makes a difference right? I honestly have no idea, putting together a simulation right now to see what I find."
        },
        {
          "id": 954366,
          "postDate": "2020-08-01T17:20:38.167Z",
          "content": "<p>I ran a simulation with some assumptions to see what I would find.\n- 1010 participants, 1000 use public kernels, 10 are \"smart\" solutions\n- Submissions have the same std of 0.005\n- \"public kernel\" submissions average lb score is 0.950\n- \"smart\" submission average lb score is 0.955\n- Simulate 1, 2, or 3 submissions allowed, select the best submission and rank\n- bootstrap this simulation 1000x\n- look at the average results.</p>\n\n<p>What I found was:\nWith 1 submission allowed the average \"smart\" solution ranks 241st place\nWith 2 submissions: 212th\nWith 3 submissions: 191st!</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F644036%2Fb1b09fc3081bd6abb8f973d36d282b73%2Fdownload.png?generation=1596302347362455&amp;alt=media\" alt=\"\"></p>\n\n<p><a href=\"https://www.kaggle.com/robikscube/simulate-the-impact-of-submission-count\">I made it into a kernel here</a></p>",
          "rawMarkdown": "I ran a simulation with some assumptions to see what I would find.\n- 1010 participants, 1000 use public kernels, 10 are \"smart\" solutions\n- Submissions have the same std of 0.005\n- \"public kernel\" submissions average lb score is 0.950\n- \"smart\" submission average lb score is 0.955\n- Simulate 1, 2, or 3 submissions allowed, select the best submission and rank\n- bootstrap this simulation 1000x\n- look at the average results.\n\nWhat I found was:\nWith 1 submission allowed the average \"smart\" solution ranks 241st place\nWith 2 submissions: 212th\nWith 3 submissions: 191st!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F644036%2Fb1b09fc3081bd6abb8f973d36d282b73%2Fdownload.png?generation=1596302347362455&amp;alt=media)\n\n[I made it into a kernel here](https://www.kaggle.com/robikscube/simulate-the-impact-of-submission-count)",
          "votes": 12,
          "replies": [
            {
              "id": 954375,
              "postDate": "2020-08-01T17:29:26.033Z",
              "content": "<p>Nice work! </p>\n\n<p>It seems reasonable that 'smart' solutions would have a smaller standard deviation, which would likely effect the outcome.</p>",
              "rawMarkdown": "Nice work! \n\nIt seems reasonable that 'smart' solutions would have a smaller standard deviation, which would likely effect the outcome."
            }
          ]
        },
        {
          "id": 954402,
          "postDate": "2020-08-01T17:49:08.073Z",
          "content": "<p>Great simulation <a href=\"/robikscube\">@robikscube</a> . The result surprises me. Your simulation suggests that more submissions removes luck. I'll need to think more about this as it contradicts my intuition. Thanks for posting.</p>",
          "rawMarkdown": "Great simulation @robikscube . The result surprises me. Your simulation suggests that more submissions removes luck. I'll need to think more about this as it contradicts my intuition. Thanks for posting.",
          "votes": 1
        },
        {
          "id": 954414,
          "postDate": "2020-08-01T17:57:58.187Z",
          "content": "<p>Ok, i realized where my intuition went wrong. If the \"smart\" solution removes variance from their model then more submissions means more luck. And the \"smart\" solution will have a more difficult chance achieving top 50.</p>\n\n<p>Imagine that public notebooks have average LB 0.950 and STD 0.005. Now imagine that a \"smart\" solution has average LB 0.955 and STD 0.000. In this case, the \"smart\" solution is hurt by more submissions.</p>\n\n<p>(This is also true if smart solution has STD 0.001, but using STD 0.000 makes the argument clearer).</p>",
          "rawMarkdown": "Ok, i realized where my intuition went wrong. If the \"smart\" solution removes variance from their model then more submissions means more luck. And the \"smart\" solution will have a more difficult chance achieving top 50.\n\nImagine that public notebooks have average LB 0.950 and STD 0.005. Now imagine that a \"smart\" solution has average LB 0.955 and STD 0.000. In this case, the \"smart\" solution is hurt by more submissions.\n\n(This is also true if smart solution has STD 0.001, but using STD 0.000 makes the argument clearer).",
          "votes": 2
        },
        {
          "id": 954430,
          "postDate": "2020-08-01T18:10:11.523Z",
          "content": "<p><a href=\"/robikscube\">@robikscube</a> Here is your simulation if I change the STD of \"smart solution\" to 0.001 instead of 0.005. Now having more submissions lowers the rank of the \"smart solution\", thus introducing more luck:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fd4e704928178e75146e7006717abcbad%2FScreen%20Shot%202020-08-01%20at%2011.07.43%20AM.png?generation=1596305356280633&amp;alt=media\" alt=\"\"></p>\n\n<p>Now the result is:\nWith 1 submission allowed the average \"smart\" solution ranks 168th place!\nWith 2 submissions: 258th\nWith 3 submissions: 331st</p>",
          "rawMarkdown": "@robikscube Here is your simulation if I change the STD of \"smart solution\" to 0.001 instead of 0.005. Now having more submissions lowers the rank of the \"smart solution\", thus introducing more luck:\n  \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fd4e704928178e75146e7006717abcbad%2FScreen%20Shot%202020-08-01%20at%2011.07.43%20AM.png?generation=1596305356280633&amp;alt=media)\n\nNow the result is:\nWith 1 submission allowed the average \"smart\" solution ranks 168th place!\nWith 2 submissions: 258th\nWith 3 submissions: 331st\n",
          "votes": 3
        },
        {
          "id": 954440,
          "postDate": "2020-08-01T18:20:14.017Z",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> - I see your point. If you add the assumption that \"smart\" submissions have little or no randomness (close to zero std) and public kernels do have randomness, then \"smart\" submissions are at a disadvantage if more submissions are allowed. This is because the \"smart\" solutions are essentially 3 identical submissions. 👍 </p>\n\n<p>But... I don't know if I agree that the standard deviation is that drastic between public kernels and \"smart\" subs. Is there a reason why we would assume public kernels would have 5x the stdev of a \"smart\" solution? If anything I would think they are equal- but I have no data to support that.</p>\n\n<p>This is a very fun and interesting conversation for me- though I'm still skeptical that 3 subs in some way increases the luck of the private LB.</p>",
          "rawMarkdown": "@cdeotte - I see your point. If you add the assumption that \"smart\" submissions have little or no randomness (close to zero std) and public kernels do have randomness, then \"smart\" submissions are at a disadvantage if more submissions are allowed. This is because the \"smart\" solutions are essentially 3 identical submissions. 👍 \n\nBut... I don't know if I agree that the standard deviation is that drastic between public kernels and \"smart\" subs. Is there a reason why we would assume public kernels would have 5x the stdev of a \"smart\" solution? If anything I would think they are equal- but I have no data to support that.\n\nThis is a very fun and interesting conversation for me- though I'm still skeptical that 3 subs in some way increases the luck of the private LB."
        },
        {
          "id": 954451,
          "postDate": "2020-08-01T18:29:29.140Z",
          "content": "<p>This comp is about increasing AUC. I think that many recent research papers suggesting how to increase AUC will also lower it's variance. In this Kaggle comp, we need to increase AUC and keep it's variance.</p>",
          "rawMarkdown": "This comp is about increasing AUC. I think that many recent research papers suggesting how to increase AUC will also lower it's variance. In this Kaggle comp, we need to increase AUC and keep it's variance.",
          "votes": 1
        },
        {
          "id": 954460,
          "postDate": "2020-08-01T18:40:07.617Z",
          "content": "<p>&gt; increase AUC will also lower it's variance.</p>\n\n<p>Great point! I didn't consider this. I think you've convinced me.</p>\n\n<p>The cool thing is that after the competition is over we will be able to see all the submission data- and can re-calculate the LB based on if only 2 submissions were allowed based on public LB score.</p>",
          "rawMarkdown": "&gt; increase AUC will also lower it's variance.\n \nGreat point! I didn't consider this. I think you've convinced me.\n\nThe cool thing is that after the competition is over we will be able to see all the submission data- and can re-calculate the LB based on if only 2 submissions were allowed based on public LB score.",
          "votes": 1
        },
        {
          "id": 955302,
          "postDate": "2020-08-02T14:35:41.867Z",
          "content": "<p>Your notebooks and datasets are awesome <a href=\"/cdeotte\">@cdeotte</a> \nYour over all CV  AUC  (0.95 that you mentioned)is pretty good.I am getting a 0.926 on my best model but it does not score well on the leader board LB 0.934 ,but the one having less CV is giving better rank on LB.\nI have followed your triple stratified CV and my doubt  is how much should I believe the CV AUC and should it be valued less than the LB score and how should I interpret  the good AUC that I get . </p>",
          "rawMarkdown": "Your notebooks and datasets are awesome @cdeotte \nYour over all CV  AUC  (0.95 that you mentioned)is pretty good.I am getting a 0.926 on my best model but it does not score well on the leader board LB 0.934 ,but the one having less CV is giving better rank on LB.\nI have followed your triple stratified CV and my doubt  is how much should I believe the CV AUC and should it be valued less than the LB score and how should I interpret  the good AUC that I get . ",
          "votes": 1
        },
        {
          "id": 956045,
          "postDate": "2020-08-03T07:28:16.090Z",
          "content": "<blockquote>\n  <p>my doubt is how much should I believe the CV AUC and should it be valued less than the LB score</p>\n</blockquote>\n\n<p>This is always a difficult question. Public test has 3000 images while local validation has 33000 images. This seems to imply that CV is 10x more important. However, the public test may be more similar to private test than train. So this increases public LB's importance. It is hard to know what to do.</p>",
          "rawMarkdown": "&gt; my doubt is how much should I believe the CV AUC and should it be valued less than the LB score\n\nThis is always a difficult question. Public test has 3000 images while local validation has 33000 images. This seems to imply that CV is 10x more important. However, the public test may be more similar to private test than train. So this increases public LB's importance. It is hard to know what to do."
        },
        {
          "id": 956712,
          "postDate": "2020-08-03T17:54:28.167Z",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> I think the large model variance arise from the randomness of the large TTA steps within each trial. This is the first time for me to see random TTA being applied at the inference phase.</p>",
          "rawMarkdown": "@cdeotte I think the large model variance arise from the randomness of the large TTA steps within each trial. This is the first time for me to see random TTA being applied at the inference phase."
        }
      ]
    },
    {
      "id": 966489,
      "postDate": "2020-08-11T13:06:22.763Z",
      "content": "<p>Anyone planning to include nested blend of blends as one of your submissions? A lot of people are expecting a large shake up due to public submissions but I am wondering if a lottery may instead occur and blend of blends actually end up in top submissions lol. I personally will focus on submissions with best CV and reasonable LB</p>",
      "rawMarkdown": "Anyone planning to include nested blend of blends as one of your submissions? A lot of people are expecting a large shake up due to public submissions but I am wondering if a lottery may instead occur and blend of blends actually end up in top submissions lol. I personally will focus on submissions with best CV and reasonable LB"
    },
    {
      "id": 966383,
      "postDate": "2020-08-11T11:36:15.820Z",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> should I need to submit only test data set prediction CVS file? Or also I need to submit the model? Is it necessary to make the notebook publicly available on the Kaggle? Sorry for this question but I am new to Kaggle and this is my first competition.</p>",
      "rawMarkdown": "@cdeotte should I need to submit only test data set prediction CVS file? Or also I need to submit the model? Is it necessary to make the notebook publicly available on the Kaggle? Sorry for this question but I am new to Kaggle and this is my first competition.",
      "replies": [
        {
          "id": 966627,
          "postDate": "2020-08-11T15:26:47.530Z",
          "content": "<p>Yes, you need only to submit test data set prediction CSV file. (So for example, you can do all your modeling on your home computer and then just submit the output CSV file).</p>\n<p>If you win a cash prize (i.e. there are 5 cash prizes), then you will need to email your code to Kaggle (after winning the prize) for them to review to verify that you followed the competition rules before you get your cash prize.</p>",
          "rawMarkdown": "Yes, you need only to submit test data set prediction CSV file. (So for example, you can do all your modeling on your home computer and then just submit the output CSV file).\n\nIf you win a cash prize (i.e. there are 5 cash prizes), then you will need to email your code to Kaggle (after winning the prize) for them to review to verify that you followed the competition rules before you get your cash prize."
        }
      ]
    },
    {
      "id": 960295,
      "postDate": "2020-08-06T09:38:09.267Z",
      "content": "<p>It had passed me by totally, what a great surprise and present, and with six months until Christmas Eve 👍 🙏 🎉 \nA fun surprise for those who have not seen it, like me, and too bad not use that opportunity, unusual, easy to miss, thank you!</p>",
      "rawMarkdown": "It had passed me by totally, what a great surprise and present, and with six months until Christmas Eve 👍 🙏 🎉 \nA fun surprise for those who have not seen it, like me, and too bad not use that opportunity, unusual, easy to miss, thank you!"
    },
    {
      "id": 960107,
      "postDate": "2020-08-06T06:25:20.790Z",
      "content": "<p>3 submissions mean that the organizer has foreseen the big shakeup. 3 submissions to me are fairer than 2. </p>",
      "rawMarkdown": "3 submissions mean that the organizer has foreseen the big shakeup. 3 submissions to me are fairer than 2. "
    },
    {
      "id": 954410,
      "postDate": "2020-08-01T17:54:33.133Z",
      "content": "<p>I started a topic recently on the same issue but more abstractly. I, actulally, don't understand the done votes. It was about possible solution. Why not, while still having private and public, evaluate the results at the end of competition based on both. E.g. the higher both are and the less discrepancy they have the better. If evaluating competition like this, one can find more generalizable solutions.\nIt is more like a thought process, how can Kaggle do better.</p>",
      "rawMarkdown": "I started a topic recently on the same issue but more abstractly. I, actulally, don't understand the done votes. It was about possible solution. Why not, while still having private and public, evaluate the results at the end of competition based on both. E.g. the higher both are and the less discrepancy they have the better. If evaluating competition like this, one can find more generalizable solutions.\nIt is more like a thought process, how can Kaggle do better.",
      "replies": [
        {
          "id": 954422,
          "postDate": "2020-08-01T18:05:41.257Z",
          "content": "<p>Your suggestion (automatically selecting highest private LB sub) removes human choice. Part of a Kaggle competition is <strong>choosing</strong> (determining) which of your multiple models is better at generalizing (predicting unseen data).</p>\n\n<p>(Because in the real world, you will need to <strong>choose</strong> which model to put into production <strong>before</strong> having results from production).</p>",
          "rawMarkdown": "Your suggestion (automatically selecting highest private LB sub) removes human choice. Part of a Kaggle competition is **choosing** (determining) which of your multiple models is better at generalizing (predicting unseen data).\n\n(Because in the real world, you will need to **choose** which model to put into production **before** having results from production).",
          "votes": 4
        },
        {
          "id": 954427,
          "postDate": "2020-08-01T18:08:30.207Z",
          "content": "<p>I meant you still choose the best solution yourself beforehand according the local validation set results and some other intuitions or knowledge you may have on the project. The result is just something like harmonized mean of private and public scores for those chosen by you solutions</p>\n\n<p>You chose before you know private (how well it is in \"production\")</p>",
          "rawMarkdown": "I meant you still choose the best solution yourself beforehand according the local validation set results and some other intuitions or knowledge you may have on the project. The result is just something like harmonized mean of private and public scores for those chosen by you solutions\n\nYou chose before you know private (how well it is in \"production\")",
          "replies": [
            {
              "id": 954441,
              "postDate": "2020-08-01T18:20:27.457Z",
              "content": "<p>The problem is in a competition the public score can be easily overfit. It is similar to saying you should evaluate the performance of a model by its training score, important but an awful metric for testing generalisation of a model as you can trivially make a perfect classifier on the training data (e.g. a fully developed decision tree).</p>",
              "rawMarkdown": "The problem is in a competition the public score can be easily overfit. It is similar to saying you should evaluate the performance of a model by its training score, important but an awful metric for testing generalisation of a model as you can trivially make a perfect classifier on the training data (e.g. a fully developed decision tree)."
            }
          ]
        },
        {
          "id": 954435,
          "postDate": "2020-08-01T18:16:41.480Z",
          "content": "<p>Thats an interesting idea. The idea being: giving some weight to public LB score in the final standings. It would make shakeups less drastic when they occur.</p>\n\n<p>However, it would probably encourage bad data science though. People would be making lots of submissions just trying to overfit public LB. I'm not sure that its a good thing to encourage that.</p>",
          "rawMarkdown": "Thats an interesting idea. The idea being: giving some weight to public LB score in the final standings. It would make shakeups less drastic when they occur.\n\nHowever, it would probably encourage bad data science though. People would be making lots of submissions just trying to overfit public LB. I'm not sure that its a good thing to encourage that.",
          "votes": 1
        },
        {
          "id": 954439,
          "postDate": "2020-08-01T18:19:09.900Z",
          "content": "<p>As I tell, more like a thought process. I belive  Kaggle can try it as one timer and see if it make sense )</p>",
          "rawMarkdown": "As I tell, more like a thought process. I belive  Kaggle can try it as one timer and see if it make sense )",
          "replies": [
            {
              "id": 954457,
              "postDate": "2020-08-01T18:37:06.523Z",
              "content": "<p>Actually, my recommendation would be to make the entire competition completely like real life, where you only have training data and no access to test data at all.</p>\n\n<ol>\n<li>Do not provide either public test data or private test data to participants (so even Public LB cannot be guessed)</li>\n<li>Completely hide private test data, and ask participants to just submit the <strong>model file</strong> directly 🙏 \nSee <a href=\"https://www.kaggle.com/c/landmark-retrieval-2020/overview/evaluation\">Google Landmark Retrieval 2020</a></li>\n</ol>\n\n<p>Then only the <strong>best model</strong> wins! (Production deployments don't use ensemble solutions like Kaggle anyway)\nNo more Overfitting, no LB probing, no Leak hunting!\n(However, this will likely <strong>increase</strong> the element of luck for everyone, just like real-world implementations!)</p>",
              "rawMarkdown": "Actually, my recommendation would be to make the entire competition completely like real life, where you only have training data and no access to test data at all.\n\n1. Do not provide either public test data or private test data to participants (so even Public LB cannot be guessed)\n2. Completely hide private test data, and ask participants to just submit the **model file** directly 🙏 \nSee [Google Landmark Retrieval 2020](https://www.kaggle.com/c/landmark-retrieval-2020/overview/evaluation)\n\nThen only the **best model** wins! (Production deployments don't use ensemble solutions like Kaggle anyway)\nNo more Overfitting, no LB probing, no Leak hunting!\n(However, this will likely **increase** the element of luck for everyone, just like real-world implementations!)",
              "votes": 1
            }
          ]
        },
        {
          "id": 954445,
          "postDate": "2020-08-01T18:22:42Z",
          "content": "<p>Shakeups are heart breaking, but after you experience a few, you begin to know when they will occur and are less surprised by them. Also you learn the important skill of making more general models to avoid them.</p>",
          "rawMarkdown": "Shakeups are heart breaking, but after you experience a few, you begin to know when they will occur and are less surprised by them. Also you learn the important skill of making more general models to avoid them.",
          "votes": 1
        },
        {
          "id": 954447,
          "postDate": "2020-08-01T18:26:03.140Z",
          "content": "<p>Good points :) Thanks a lot for all your inspiring work here, by the way!</p>",
          "rawMarkdown": "Good points :) Thanks a lot for all your inspiring work here, by the way!",
          "votes": 1
        },
        {
          "id": 966851,
          "postDate": "2020-08-11T17:35:05.693Z",
          "content": "<p>I should be in favor of using a combination of public and private LB.  Indeed, I would have won several competitions if final standing was a mix of public and private LB, namely the ones where our team survived shakeup, for instance LANL or Malware.  It would remove the lucky random subs from private LB.  </p>\n<p>But we know it would encourage public LB overfitting a lot, which would change the outcome.  Current setting discourages public LB overfitting.  I much prefer it that way.  It is some guarantee that our models generalize correctly.</p>",
          "rawMarkdown": "I should be in favor of using a combination of public and private LB.  Indeed, I would have won several competitions if final standing was a mix of public and private LB, namely the ones where our team survived shakeup, for instance LANL or Malware.  It would remove the lucky random subs from private LB.  \n\nBut we know it would encourage public LB overfitting a lot, which would change the outcome.  Current setting discourages public LB overfitting.  I much prefer it that way.  It is some guarantee that our models generalize correctly.",
          "votes": 1
        }
      ]
    },
    {
      "id": 955244,
      "postDate": "2020-08-02T13:34:55.867Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 954757,
      "postDate": "2020-08-02T04:17:28.470Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 958632,
      "postDate": "2020-08-05T05:02:02.127Z",
      "content": "<p>Thanks</p>",
      "rawMarkdown": "Thanks",
      "votes": 1
    },
    {
      "id": 955296,
      "postDate": "2020-08-02T14:31:40.900Z",
      "content": "<p>THANKS</p>",
      "rawMarkdown": "THANKS",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 954129,
      "author_name": "Sirish Somanchi",
      "author_url": "",
      "post_date": "2020-08-01T12:09:02.697000",
      "content": "<ul>\n<li>Submission 1 : Without patient_id context: Triple-Stratified-CV ensemble with EfficientNet-B5/6/7, FixEfficientNet-B5/6/7, random augmentations, seeds</li>\n<li>Submission 2 : With patient_id context: Submission 1 + tabular data (NN/TabNet/LGBM/XGB)</li>\n<li>Submission 3 : Without patient_id context: Experimental custom NN using only images</li>\n</ul>",
      "votes": 5,
      "replies": [
        {
          "id": 954916,
          "author_name": "MhdSharuk",
          "author_url": "",
          "post_date": "2020-08-02T07:38:51.860000",
          "content": "<p><a href=\"/sirishks\">@sirishks</a> what is this FixEfficientNet??Could you explain?\nWhere can i find the implementation??</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 955269,
          "author_name": "Roman Weilguny",
          "author_url": "",
          "post_date": "2020-08-02T14:00:57.260000",
          "content": "<p>Your scores getting more and more unreal :) I am excited to see how they will perform on the private lb - good luck!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 958233,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-04T21:21:08.303000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 955402,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2020-08-02T15:52:09.100000",
      "content": "<p>Choose 3 Final Submissions, one more option for overfitting with public LB (increase the probability of winning the lottery) :))</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 954395,
      "author_name": "Anatoly Pavlov",
      "author_url": "",
      "post_date": "2020-08-01T17:45:06.277000",
      "content": "<p>I have a question: <strong>Are submissions made of public kernels qualified as legit at the end of competition?</strong> It seems to me that if the answer is yes, then it is kind of unfair. I always thought that public kernels are for fair exchange of ideas and should not be used to make submissions except for the author of kernel.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 954398,
          "author_name": "FChmiel",
          "author_url": "",
          "post_date": "2020-08-01T17:47:24.880000",
          "content": "<p>They are fine to use, they are released under an open source license.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 954399,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-01T17:47:36.583000",
          "content": "<p>Using public kernels is allowed. You can either use the <code>submission.csv</code> file directly, or you can fork the notebook and run the exact same code again and use the new <code>submission.csv</code>.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 954421,
          "author_name": "Anatoly Pavlov",
          "author_url": "",
          "post_date": "2020-08-01T18:05:36.410000",
          "content": "<p>Thank you <a href=\"/fchmiel\">@fchmiel</a> and <a href=\"/cdeotte\">@cdeotte</a>  for clarifying it for me. Still strange that it is allowed though.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 953937,
      "author_name": "Indranil Bhattacharya",
      "author_url": "",
      "post_date": "2020-08-01T08:55:32.480000",
      "content": "<p>Maybe, the hosts realized beforehand that the shakeup is inevitable (with two submissions), so they are giving participants more chances.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 954504,
          "author_name": "Nischay Dhankhar",
          "author_url": "",
          "post_date": "2020-08-01T19:29:39.603000",
          "content": "<p>They are also giving more hope which may lead to more disheartening after loosing ranks in shakeup.😅 😂 </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 953955,
      "author_name": "steubk",
      "author_url": "",
      "post_date": "2020-08-01T09:14:03.360000",
      "content": "<p>3 submissions are necessary to manage LB + 2 special prizes (Top model With and Without Context)</p>\n\n<blockquote>\n  <p>Eligible submissions for these Special Prizes must be a selected submission, for which each team is permitted 3 in this competition. As such, teams may consider diversifying their submission selections, if they also wish to be considered for either special prize.</p>\n</blockquote>\n\n<p><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/overview/prizes\">https://www.kaggle.com/c/siim-isic-melanoma-classification/overview/prizes</a></p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 953934,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2020-08-01T08:51:39.343000",
      "content": "<p>Wow, how fo you feel about it?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 954036,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-01T10:23:48.597000",
          "content": "<p>It will add more luck to the final results. </p>\n\n<p>This comp will have an interesting final leaderboard. I think many strong competitors are carefully maximizing a triple stratified CV. My overall CV score is around AUC 0.950. There will be a cluster of private LB scores from these careful models.</p>\n\n<p>The question is, where will all the random public notebook ensembles finish? How many will finish higher than a careful CV and how many lower? I'm not sure. With everyone having 3 final submissions, a careful CV model may not be able to reach the final top 50.</p>",
          "votes": 10,
          "replies": []
        },
        {
          "id": 954100,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-08-01T11:36:38.607000",
          "content": "<p>Yes, I am not really happy about that. It makes it really easy to overfit public LB and choose it as one of the three final subs. This means that there is a higher chance in the end that some top scores are rather lucky ones, which might be a disadvantage to robust solutions as you say.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 954264,
          "author_name": "Rob Mulla",
          "author_url": "",
          "post_date": "2020-08-01T15:13:31.920000",
          "content": "<p>This is an interesting conversation and question that seems to come up often. It got me thinking about how the competition design might make it more or less susceptible to \"luck\". Initially, I thought that since everyone is given the same number of submissions- then the variance in results per participant would decrease as the number of final submissions increases, which would decrease \"luck\". I know M5 having a 1 submission limit made many complained afterwards that the results were more a result of luck.</p>\n\n<p>All this led me down the rabbit hole this morning and found <a href=\"https://pdfs.semanticscholar.org/5e31/2f4f821eb0f0da538cebb18ff0cbe29f9735.pdf?_ga=2.232360031.1105651826.1596292464-934602416.1596292464\">this paper studying kaggle tournament design: </a> they discuss the impacts of the financial reward and size of the competition- but I don't see anywhere that they specifically discuss the impact of submission limits. It might be a fun thing to investigate using the meta kaggle dataset.</p>\n\n<p>I guess the definition of \"luck\" here is key. Some might say if there is a big shakeup then the competition involved more luck, others say that if the public LB can be overfit and is closely correlated to private then a lucky submission is one based only on LB and not CV. I'm still not convinced of either. I think a solution based solely on public LB is foolish because it may be overfit, but a submission based only on CV might ignore important signal that can be gained from the public LB. This especially is important when the test set isn't a random sample of the same data as the provided training data. For instance the \"mystery\" images in this competition means that information can be gained from the public LB that can't be gained from the training data alone.</p>\n\n<p>I don't think having 3 submissions instead of 2 will really change the number of people submitting public kernel blends. There will always be a substantial number of people that will do that regardless of the limit.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 954282,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-01T15:44:25.773000",
          "content": "<p>In general more submissions doesn't necessarily add more luck, but this competition is different (because of high model variance).</p>\n\n<p>Because the metric is AUC and because there are few positives, AUC fluctuates a lot. Imagine everyone has a model with average LB 0.950 and if they train and submit the same model 100 times, the LB may range from 0.935 to 0.965 (i.e. STD = 0.005). In other competitions the LB may only range from 0.949 to 0.951 (low model variance). Since STD is 0.005, they have a 1% chance (based on z score table) of scoring 0.963+.</p>\n\n<p>If 3000 participants have 3 subs each, that's 9000 subs. If everyone repeatedly trains this same model, then 90 subs would score over 0.963+. </p>\n\n<p>Now let's say one person has a model with average LB 0.955 (and STD 0.005). It appears to be a better model, but they would need 0.965+ to place in top 50. Therefore they only have 2.5% chance (based on z score table) of placing in top 50 even though they have a better model with 1 sub and 7.3% chance with 3 subs.</p>\n\n<p>If all participants had 2 subs each, then only 60 subs would be over 0.963+ and the better model would have 10% chance of being in top 50 with 1 sub and 27% chance with 3 subs.</p>\n\n<p>EDIT: This argument is only valid when the better model also has lower STD than the repeatedly submitted model.</p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 954328,
          "author_name": "Rob Mulla",
          "author_url": "",
          "post_date": "2020-08-01T16:27:29.247000",
          "content": "<p>🤔 Interesting <a href=\"/cdeotte\">@cdeotte</a> - I'll have to think about that. I follow you up until this point:</p>\n\n<blockquote>\n  <p>Now let's say one person has a model with average LB 0.955 (and STD 0.05). It appears to be a better model, but they would need 0.965+ to place in top 50. Therefore they only have 2.5% chance (based on z score table) of placing in top 50 even though they have a better model.</p>\n</blockquote>\n\n<p>Because the person with average LB model of 0.955 and STD 0.05 gets 3 subs also. I might try to simulate this 😄 </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 954330,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-01T16:31:51.383000",
          "content": "<p><a href=\"/robikscube\">@robikscube</a>\n&gt; Because the person with average LB model of 0.955 and STD 0.05 gets 3 subs also</p>\n\n<p>Yes, i forgot about their 3 subs. The better model has a 7.3% chance of placing in top 50 with 3 subs. (because they have <code>97.5 ** 3 = 92.7%</code> chance of not placing in top 50 with 3 subs).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 954332,
          "author_name": "Rob Mulla",
          "author_url": "",
          "post_date": "2020-08-01T16:34:14.657000",
          "content": "<p>But the question is about if 1 vs 2 vs 3 subs makes a difference right? I honestly have no idea, putting together a simulation right now to see what I find.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 954366,
          "author_name": "Rob Mulla",
          "author_url": "",
          "post_date": "2020-08-01T17:20:38.167000",
          "content": "<p>I ran a simulation with some assumptions to see what I would find.\n- 1010 participants, 1000 use public kernels, 10 are \"smart\" solutions\n- Submissions have the same std of 0.005\n- \"public kernel\" submissions average lb score is 0.950\n- \"smart\" submission average lb score is 0.955\n- Simulate 1, 2, or 3 submissions allowed, select the best submission and rank\n- bootstrap this simulation 1000x\n- look at the average results.</p>\n\n<p>What I found was:\nWith 1 submission allowed the average \"smart\" solution ranks 241st place\nWith 2 submissions: 212th\nWith 3 submissions: 191st!</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F644036%2Fb1b09fc3081bd6abb8f973d36d282b73%2Fdownload.png?generation=1596302347362455&amp;alt=media\" alt=\"\"></p>\n\n<p><a href=\"https://www.kaggle.com/robikscube/simulate-the-impact-of-submission-count\">I made it into a kernel here</a></p>",
          "votes": 12,
          "replies": [
            {
              "id": 954375,
              "author_name": "FChmiel",
              "author_url": "",
              "post_date": "2020-08-01T17:29:26.033000",
              "content": "<p>Nice work! </p>\n\n<p>It seems reasonable that 'smart' solutions would have a smaller standard deviation, which would likely effect the outcome.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 954402,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-01T17:49:08.073000",
          "content": "<p>Great simulation <a href=\"/robikscube\">@robikscube</a> . The result surprises me. Your simulation suggests that more submissions removes luck. I'll need to think more about this as it contradicts my intuition. Thanks for posting.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 954414,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-01T17:57:58.187000",
          "content": "<p>Ok, i realized where my intuition went wrong. If the \"smart\" solution removes variance from their model then more submissions means more luck. And the \"smart\" solution will have a more difficult chance achieving top 50.</p>\n\n<p>Imagine that public notebooks have average LB 0.950 and STD 0.005. Now imagine that a \"smart\" solution has average LB 0.955 and STD 0.000. In this case, the \"smart\" solution is hurt by more submissions.</p>\n\n<p>(This is also true if smart solution has STD 0.001, but using STD 0.000 makes the argument clearer).</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 954430,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-01T18:10:11.523000",
          "content": "<p><a href=\"/robikscube\">@robikscube</a> Here is your simulation if I change the STD of \"smart solution\" to 0.001 instead of 0.005. Now having more submissions lowers the rank of the \"smart solution\", thus introducing more luck:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fd4e704928178e75146e7006717abcbad%2FScreen%20Shot%202020-08-01%20at%2011.07.43%20AM.png?generation=1596305356280633&amp;alt=media\" alt=\"\"></p>\n\n<p>Now the result is:\nWith 1 submission allowed the average \"smart\" solution ranks 168th place!\nWith 2 submissions: 258th\nWith 3 submissions: 331st</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 954440,
          "author_name": "Rob Mulla",
          "author_url": "",
          "post_date": "2020-08-01T18:20:14.017000",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> - I see your point. If you add the assumption that \"smart\" submissions have little or no randomness (close to zero std) and public kernels do have randomness, then \"smart\" submissions are at a disadvantage if more submissions are allowed. This is because the \"smart\" solutions are essentially 3 identical submissions. 👍 </p>\n\n<p>But... I don't know if I agree that the standard deviation is that drastic between public kernels and \"smart\" subs. Is there a reason why we would assume public kernels would have 5x the stdev of a \"smart\" solution? If anything I would think they are equal- but I have no data to support that.</p>\n\n<p>This is a very fun and interesting conversation for me- though I'm still skeptical that 3 subs in some way increases the luck of the private LB.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 954451,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-01T18:29:29.140000",
          "content": "<p>This comp is about increasing AUC. I think that many recent research papers suggesting how to increase AUC will also lower it's variance. In this Kaggle comp, we need to increase AUC and keep it's variance.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 954460,
          "author_name": "Rob Mulla",
          "author_url": "",
          "post_date": "2020-08-01T18:40:07.617000",
          "content": "<p>&gt; increase AUC will also lower it's variance.</p>\n\n<p>Great point! I didn't consider this. I think you've convinced me.</p>\n\n<p>The cool thing is that after the competition is over we will be able to see all the submission data- and can re-calculate the LB based on if only 2 submissions were allowed based on public LB score.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 955302,
          "author_name": "adigamer970",
          "author_url": "",
          "post_date": "2020-08-02T14:35:41.867000",
          "content": "<p>Your notebooks and datasets are awesome <a href=\"/cdeotte\">@cdeotte</a> \nYour over all CV  AUC  (0.95 that you mentioned)is pretty good.I am getting a 0.926 on my best model but it does not score well on the leader board LB 0.934 ,but the one having less CV is giving better rank on LB.\nI have followed your triple stratified CV and my doubt  is how much should I believe the CV AUC and should it be valued less than the LB score and how should I interpret  the good AUC that I get . </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 956045,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-03T07:28:16.090000",
          "content": "<blockquote>\n  <p>my doubt is how much should I believe the CV AUC and should it be valued less than the LB score</p>\n</blockquote>\n\n<p>This is always a difficult question. Public test has 3000 images while local validation has 33000 images. This seems to imply that CV is 10x more important. However, the public test may be more similar to private test than train. So this increases public LB's importance. It is hard to know what to do.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 956712,
          "author_name": "Ahmed Sabry",
          "author_url": "",
          "post_date": "2020-08-03T17:54:28.167000",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> I think the large model variance arise from the randomness of the large TTA steps within each trial. This is the first time for me to see random TTA being applied at the inference phase.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 966489,
      "author_name": "Alan Choon",
      "author_url": "",
      "post_date": "2020-08-11T13:06:22.763000",
      "content": "<p>Anyone planning to include nested blend of blends as one of your submissions? A lot of people are expecting a large shake up due to public submissions but I am wondering if a lottery may instead occur and blend of blends actually end up in top submissions lol. I personally will focus on submissions with best CV and reasonable LB</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 966383,
      "author_name": "SANKET PATEL108784",
      "author_url": "",
      "post_date": "2020-08-11T11:36:15.820000",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> should I need to submit only test data set prediction CVS file? Or also I need to submit the model? Is it necessary to make the notebook publicly available on the Kaggle? Sorry for this question but I am new to Kaggle and this is my first competition.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 966627,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-11T15:26:47.530000",
          "content": "<p>Yes, you need only to submit test data set prediction CSV file. (So for example, you can do all your modeling on your home computer and then just submit the output CSV file).</p>\n<p>If you win a cash prize (i.e. there are 5 cash prizes), then you will need to email your code to Kaggle (after winning the prize) for them to review to verify that you followed the competition rules before you get your cash prize.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 960295,
      "author_name": "Kirderf",
      "author_url": "",
      "post_date": "2020-08-06T09:38:09.267000",
      "content": "<p>It had passed me by totally, what a great surprise and present, and with six months until Christmas Eve 👍 🙏 🎉 \nA fun surprise for those who have not seen it, like me, and too bad not use that opportunity, unusual, easy to miss, thank you!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 960107,
      "author_name": "Kha Vo",
      "author_url": "",
      "post_date": "2020-08-06T06:25:20.790000",
      "content": "<p>3 submissions mean that the organizer has foreseen the big shakeup. 3 submissions to me are fairer than 2. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 954410,
      "author_name": "A.Demyanchuk",
      "author_url": "",
      "post_date": "2020-08-01T17:54:33.133000",
      "content": "<p>I started a topic recently on the same issue but more abstractly. I, actulally, don't understand the done votes. It was about possible solution. Why not, while still having private and public, evaluate the results at the end of competition based on both. E.g. the higher both are and the less discrepancy they have the better. If evaluating competition like this, one can find more generalizable solutions.\nIt is more like a thought process, how can Kaggle do better.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 954422,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-01T18:05:41.257000",
          "content": "<p>Your suggestion (automatically selecting highest private LB sub) removes human choice. Part of a Kaggle competition is <strong>choosing</strong> (determining) which of your multiple models is better at generalizing (predicting unseen data).</p>\n\n<p>(Because in the real world, you will need to <strong>choose</strong> which model to put into production <strong>before</strong> having results from production).</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 954427,
          "author_name": "A.Demyanchuk",
          "author_url": "",
          "post_date": "2020-08-01T18:08:30.207000",
          "content": "<p>I meant you still choose the best solution yourself beforehand according the local validation set results and some other intuitions or knowledge you may have on the project. The result is just something like harmonized mean of private and public scores for those chosen by you solutions</p>\n\n<p>You chose before you know private (how well it is in \"production\")</p>",
          "votes": 0,
          "replies": [
            {
              "id": 954441,
              "author_name": "FChmiel",
              "author_url": "",
              "post_date": "2020-08-01T18:20:27.457000",
              "content": "<p>The problem is in a competition the public score can be easily overfit. It is similar to saying you should evaluate the performance of a model by its training score, important but an awful metric for testing generalisation of a model as you can trivially make a perfect classifier on the training data (e.g. a fully developed decision tree).</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 954435,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-01T18:16:41.480000",
          "content": "<p>Thats an interesting idea. The idea being: giving some weight to public LB score in the final standings. It would make shakeups less drastic when they occur.</p>\n\n<p>However, it would probably encourage bad data science though. People would be making lots of submissions just trying to overfit public LB. I'm not sure that its a good thing to encourage that.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 954439,
          "author_name": "A.Demyanchuk",
          "author_url": "",
          "post_date": "2020-08-01T18:19:09.900000",
          "content": "<p>As I tell, more like a thought process. I belive  Kaggle can try it as one timer and see if it make sense )</p>",
          "votes": 0,
          "replies": [
            {
              "id": 954457,
              "author_name": "Sirish Somanchi",
              "author_url": "",
              "post_date": "2020-08-01T18:37:06.523000",
              "content": "<p>Actually, my recommendation would be to make the entire competition completely like real life, where you only have training data and no access to test data at all.</p>\n\n<ol>\n<li>Do not provide either public test data or private test data to participants (so even Public LB cannot be guessed)</li>\n<li>Completely hide private test data, and ask participants to just submit the <strong>model file</strong> directly 🙏 \nSee <a href=\"https://www.kaggle.com/c/landmark-retrieval-2020/overview/evaluation\">Google Landmark Retrieval 2020</a></li>\n</ol>\n\n<p>Then only the <strong>best model</strong> wins! (Production deployments don't use ensemble solutions like Kaggle anyway)\nNo more Overfitting, no LB probing, no Leak hunting!\n(However, this will likely <strong>increase</strong> the element of luck for everyone, just like real-world implementations!)</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 954445,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-01T18:22:42",
          "content": "<p>Shakeups are heart breaking, but after you experience a few, you begin to know when they will occur and are less surprised by them. Also you learn the important skill of making more general models to avoid them.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 954447,
          "author_name": "A.Demyanchuk",
          "author_url": "",
          "post_date": "2020-08-01T18:26:03.140000",
          "content": "<p>Good points :) Thanks a lot for all your inspiring work here, by the way!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 966851,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-08-11T17:35:05.693000",
          "content": "<p>I should be in favor of using a combination of public and private LB.  Indeed, I would have won several competitions if final standing was a mix of public and private LB, namely the ones where our team survived shakeup, for instance LANL or Malware.  It would remove the lucky random subs from private LB.  </p>\n<p>But we know it would encourage public LB overfitting a lot, which would change the outcome.  Current setting discourages public LB overfitting.  I much prefer it that way.  It is some guarantee that our models generalize correctly.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 955244,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-02T13:34:55.867000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 954757,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-02T04:17:28.470000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 958632,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-05T05:02:02.127000",
      "content": "<p>Thanks</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 955296,
      "author_name": "FABIO PANCOTTI MORENTE",
      "author_url": "",
      "post_date": "2020-08-02T14:31:40.900000",
      "content": "<p>THANKS</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "953900": "I would like to point out that this competition allows us to select 3 final submissions [here][1]\n&gt; You may select up to 3 submissions to be used to count towards your final leaderboard score. \n\nKeep this in mind as it affects final model building strategy and how we should invest our time during the last 2 weeks. Good luck everyone!\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/submissions",
    "954129": "- Submission 1 : Without patient_id context: Triple-Stratified-CV ensemble with EfficientNet-B5/6/7, FixEfficientNet-B5/6/7, random augmentations, seeds\n- Submission 2 : With patient_id context: Submission 1 + tabular data (NN/TabNet/LGBM/XGB)\n- Submission 3 : Without patient_id context: Experimental custom NN using only images",
    "955402": "Choose 3 Final Submissions, one more option for overfitting with public LB (increase the probability of winning the lottery) :))",
    "954395": "I have a question: **Are submissions made of public kernels qualified as legit at the end of competition?** It seems to me that if the answer is yes, then it is kind of unfair. I always thought that public kernels are for fair exchange of ideas and should not be used to make submissions except for the author of kernel.",
    "953937": "Maybe, the hosts realized beforehand that the shakeup is inevitable (with two submissions), so they are giving participants more chances.",
    "953955": "3 submissions are necessary to manage LB + 2 special prizes (Top model With and Without Context)\n\n&gt; Eligible submissions for these Special Prizes must be a selected submission, for which each team is permitted 3 in this competition. As such, teams may consider diversifying their submission selections, if they also wish to be considered for either special prize.\n\nhttps://www.kaggle.com/c/siim-isic-melanoma-classification/overview/prizes",
    "953934": "Wow, how fo you feel about it?",
    "966489": "Anyone planning to include nested blend of blends as one of your submissions? A lot of people are expecting a large shake up due to public submissions but I am wondering if a lottery may instead occur and blend of blends actually end up in top submissions lol. I personally will focus on submissions with best CV and reasonable LB",
    "966383": "@cdeotte should I need to submit only test data set prediction CVS file? Or also I need to submit the model? Is it necessary to make the notebook publicly available on the Kaggle? Sorry for this question but I am new to Kaggle and this is my first competition.",
    "960295": "It had passed me by totally, what a great surprise and present, and with six months until Christmas Eve 👍 🙏 🎉 \nA fun surprise for those who have not seen it, like me, and too bad not use that opportunity, unusual, easy to miss, thank you!",
    "960107": "3 submissions mean that the organizer has foreseen the big shakeup. 3 submissions to me are fairer than 2. ",
    "954410": "I started a topic recently on the same issue but more abstractly. I, actulally, don't understand the done votes. It was about possible solution. Why not, while still having private and public, evaluate the results at the end of competition based on both. E.g. the higher both are and the less discrepancy they have the better. If evaluating competition like this, one can find more generalizable solutions.\nIt is more like a thought process, how can Kaggle do better.",
    "955244": "",
    "954757": "",
    "958632": "Thanks",
    "955296": "THANKS"
  }
}