{
  "id": 77295,
  "title": "Sob story: anyone else trust their CV and miss out on gold?",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/77295",
  "author_name": "",
  "post_date": "2019-01-11T07:32:23.619313200Z",
  "votes": 5,
  "comment_count": 19,
  "views": 0,
  "content": "<p>[Firstly: <strong>congratulations to all the winners/top scorers!</strong>]</p>\n\n<p>Yep, that was us...and I'm not upset (even a little bit). ;)</p>\n\n<p>It turns out the models we had with the highest public LB also had the highest private LB score (would have won gold based on current scores), correlation ~86%. However we chose models performing (slightly) worse on the public LB but that had been validated on a much bigger validation set (one was the entire training data 5 fold, one was half of it). </p>\n\n<p>Basically it turns out using the public LB as the validation set beat validating models on one's own CV. I'm a little shocked, we were fairly confident the whole public LB was massively overfit (small, unstable metric, huge class imbalance etc...).</p>\n\n<p>What was everyone else's strategy for choosing a final model? Anyone else have a similar experience?</p>\n\n<p>Mark</p>\n\n<p>P.S. Obviously 'missing out on gold' is conditional on everyone else's final selection remaining fixed.</p>",
  "messages": [
    {
      "id": "454166",
      "postDate": "01/11/2019 07:32:23",
      "content": "<p>[Firstly: <strong>congratulations to all the winners/top scorers!</strong>]</p>\n\n<p>Yep, that was us...and I'm not upset (even a little bit). ;)</p>\n\n<p>It turns out the models we had with the highest public LB also had the highest private LB score (would have won gold based on current scores), correlation ~86%. However we chose models performing (slightly) worse on the public LB but that had been validated on a much bigger validation set (one was the entire training data 5 fold, one was half of it). </p>\n\n<p>Basically it turns out using the public LB as the validation set beat validating models on one's own CV. I'm a little shocked, we were fairly confident the whole public LB was massively overfit (small, unstable metric, huge class imbalance etc...).</p>\n\n<p>What was everyone else's strategy for choosing a final model? Anyone else have a similar experience?</p>\n\n<p>Mark</p>\n\n<p>P.S. Obviously 'missing out on gold' is conditional on everyone else's final selection remaining fixed.</p>",
      "rawMarkdown": "[Firstly: **congratulations to all the winners/top scorers!**]\n\nYep, that was us...and I'm not upset (even a little bit). ;)\n\nIt turns out the models we had with the highest public LB also had the highest private LB score (would have won gold based on current scores), correlation ~86%. However we chose models performing (slightly) worse on the public LB but that had been validated on a much bigger validation set (one was the entire training data 5 fold, one was half of it). \n\nBasically it turns out using the public LB as the validation set beat validating models on one's own CV. I'm a little shocked, we were fairly confident the whole public LB was massively overfit (small, unstable metric, huge class imbalance etc...).\n\nWhat was everyone else's strategy for choosing a final model? Anyone else have a similar experience?\n\nMark\n\nP.S. Obviously 'missing out on gold' is conditional on everyone else's final selection remaining fixed.",
      "votes": null
    },
    {
      "id": "454191",
      "postDate": "01/11/2019 08:17:40",
      "content": "<p>I saw exactly the opposite.</p>",
      "rawMarkdown": "I saw exactly the opposite.",
      "votes": null
    },
    {
      "id": "454259",
      "postDate": "01/11/2019 10:11:04",
      "content": "<p>I know that feel bro!</p>",
      "rawMarkdown": "I know that feel bro!",
      "votes": null
    },
    {
      "id": "454310",
      "postDate": "01/11/2019 12:08:17",
      "content": "<p>Hi <a href=\"/tcapelle\">@tcapelle</a> - care to elaborate? Looking at all our public to private scores the correlation is about 0.9 =&gt; public LB is a 'better' validation set than the much larger training data. Interested to hear what you did.</p>",
      "rawMarkdown": "Hi @tcapelle - care to elaborate? Looking at all our public to private scores the correlation is about 0.9 =&gt; public LB is a 'better' validation set than the much larger training data. Interested to hear what you did.",
      "votes": null
    },
    {
      "id": "454338",
      "postDate": "01/11/2019 13:21:20",
      "content": "<p>I am just a noob here, but I had huge variance. My best performing model public LB was 0.553 scored 0.494, but another model performing 0.552 would have scored 0.509. Also our validations scores where high, between 0.55 and 0.6 but hey scored low in public LB, but good on Private... </p>",
      "rawMarkdown": "I am just a noob here, but I had huge variance. My best performing model public LB was 0.553 scored 0.494, but another model performing 0.552 would have scored 0.509. Also our validations scores where high, between 0.55 and 0.6 but hey scored low in public LB, but good on Private...",
      "votes": null
    },
    {
      "id": "454346",
      "postDate": "01/11/2019 13:34:27",
      "content": "<p>Might depend on if you always added the leak file in or not.</p>",
      "rawMarkdown": "Might depend on if you always added the leak file in or not.",
      "votes": null
    },
    {
      "id": "454362",
      "postDate": "01/11/2019 14:15:23",
      "content": "<p>My choices of submissions missed the best Private LB score too. I think building a robust cv set is one of the key problems in this competition. A solid cv set could be helpful not only for selecting the best model&amp;submission but also helpful for tuning the thresholds as well. wonder how bestfitting did this :)</p>",
      "rawMarkdown": "My choices of submissions missed the best Private LB score too. I think building a robust cv set is one of the key problems in this competition. A solid cv set could be helpful not only for selecting the best model&amp;submission but also helpful for tuning the thresholds as well. wonder how bestfitting did this :)",
      "votes": null
    },
    {
      "id": "454540",
      "postDate": "01/11/2019 20:02:40",
      "content": "<p>There were duplicates with different brightness in both train and external data (and between them). This is why validation score was 0.7-0.8 while LB score was about 0.55-0.6 or a bit higher.  So we have actually overfitted the train set.</p>\n\n<p>Not sure why private LB score is generally lower than the public one, though.</p>",
      "rawMarkdown": "There were duplicates with different brightness in both train and external data (and between them). This is why validation score was 0.7-0.8 while LB score was about 0.55-0.6 or a bit higher.  So we have actually overfitted the train set.\n\nNot sure why private LB score is generally lower than the public one, though.",
      "votes": null
    },
    {
      "id": "454544",
      "postDate": "01/11/2019 20:08:03",
      "content": "<p>I excluded duplicates. Also, I am referring to models not using hpa in the validation set that scored highly with poor LB correspondence.</p>",
      "rawMarkdown": "I excluded duplicates. Also, I am referring to models not using hpa in the validation set that scored highly with poor LB correspondence.",
      "votes": null
    },
    {
      "id": "454548",
      "postDate": "01/11/2019 20:22:02",
      "content": "<p>Sorry to hear this. I’ve seen it suggested that you should use the best model on your local CV and the best model on the public LB as your two submissions. That could help you hedge your bets and avoid these kind of situations. It’s also likely that your best model on local CV is only marginally better than your second-best, so mitigating the risk might have more reward. Still a tough call in the end though, especially if you’re near the top of the LB fighting for each 0.001</p>",
      "rawMarkdown": "Sorry to hear this. I’ve seen it suggested that you should use the best model on your local CV and the best model on the public LB as your two submissions. That could help you hedge your bets and avoid these kind of situations. It’s also likely that your best model on local CV is only marginally better than your second-best, so mitigating the risk might have more reward. Still a tough call in the end though, especially if you’re near the top of the LB fighting for each 0.001",
      "votes": null
    },
    {
      "id": "454562",
      "postDate": "01/11/2019 20:38:41",
      "content": "<p>My best private submission was worse both on public LB and in my local validation...</p>",
      "rawMarkdown": "My best private submission was worse both on public LB and in my local validation...",
      "votes": null
    },
    {
      "id": "454564",
      "postDate": "01/11/2019 20:42:11",
      "content": "<p>Thanks William! Ordinarily that's what I'd do but the HPA added a 3rd dimension so we went with a model not using HPA at all which trusted the kaggle data (i.e. ignoring the public LB which we didn't think reliable) and then a good scoring model on public LB that we had validated locally too expecting it to be more robust (rather than our best pure public LB tuned models - which obviously with hindsight did best).</p>",
      "rawMarkdown": "Thanks William! Ordinarily that's what I'd do but the HPA added a 3rd dimension so we went with a model not using HPA at all which trusted the kaggle data (i.e. ignoring the public LB which we didn't think reliable) and then a good scoring model on public LB that we had validated locally too expecting it to be more robust (rather than our best pure public LB tuned models - which obviously with hindsight did best).",
      "votes": null
    },
    {
      "id": "454589",
      "postDate": "01/11/2019 21:15:43",
      "content": "<p>Mark, my sympathies as you seem to have done everything in the most rational way. I was leery of the CV (mostly due to comments by you and others and by the mismatch) and spent my time on improvements that would benefit all measures cv/private/public. For example, in the last week I bit the bullet and started running 768-HPA. It was not as hard as I thought it would be (not many epochs) and I ran resnet34 and 50 in that time - previously I ran only resnet34. My best score came from an ensemble of resnet34x2 and resnet50x1. \nFor thresholds, I chose using the public LB and I found that to be actually quite stable (I used a class-variable curve with two free parameters to keep the choice simple and the parameters didn't vary much over the range of models I trained).</p>",
      "rawMarkdown": "Mark, my sympathies as you seem to have done everything in the most rational way. I was leery of the CV (mostly due to comments by you and others and by the mismatch) and spent my time on improvements that would benefit all measures cv/private/public. For example, in the last week I bit the bullet and started running 768-HPA. It was not as hard as I thought it would be (not many epochs) and I ran resnet34 and 50 in that time - previously I ran only resnet34. My best score came from an ensemble of resnet34x2 and resnet50x1. \nFor thresholds, I chose using the public LB and I found that to be actually quite stable (I used a class-variable curve with two free parameters to keep the choice simple and the parameters didn't vary much over the range of models I trained).",
      "votes": null
    },
    {
      "id": "454595",
      "postDate": "01/11/2019 21:28:00",
      "content": "<p>Thanks Pete, glad you placed well and nice to hear a little of what you did - you definitely bring a lot to the discussions. </p>\n\n<p>Yeah, guess it's just one of those things. I feel doubly bad as it was me pushing my team to use one model without HPA as a final submission...I just couldn't shake the feeling of how silly I'd feel if I listened to a ~3.5k public LB score vs. models doing well on almost 9x that data (esp. considering class imbalance,  metric etc..). Lesson 1 of kaggle is meant to be live and die by your CV...well, I just lost a life.</p>",
      "rawMarkdown": "Thanks Pete, glad you placed well and nice to hear a little of what you did - you definitely bring a lot to the discussions. \n\nYeah, guess it's just one of those things. I feel doubly bad as it was me pushing my team to use one model without HPA as a final submission...I just couldn't shake the feeling of how silly I'd feel if I listened to a ~3.5k public LB score vs. models doing well on almost 9x that data (esp. considering class imbalance,  metric etc..). Lesson 1 of kaggle is meant to be live and die by your CV...well, I just lost a life.",
      "votes": null
    },
    {
      "id": "454636",
      "postDate": "01/11/2019 22:57:48",
      "content": "<p>My best sub on Private LB scored 0.50979.  Unfortunately it was not one of my final selections, one of which was based on my best on Public LB and scored 0.50481 on Private LB, and the other which involved some \"clever\" thresholding based on validation set performance but scored only 0.49623 on Private LB.  But I'm pleased with my result, since I went up 89 places from Public to Private and had no aspirations for anything beyond my bronze medal.   I made no use of the external data, having shied away from it because of the leaks and other complications.\nI think it's not worth agonizing over the strategic choices that led to a loss of gold or other precious metals.  Because of the small numbers of examples of the rare classes, I believe that just plain luck played a significant role in determining the final rankings.  To paraphrase Cassius in Shakespeare’s play \"Julius Caesar\" (but reversing the meaning): \"The fault is not in ourselves but in our stars\".</p>",
      "rawMarkdown": "My best sub on Private LB scored 0.50979.  Unfortunately it was not one of my final selections, one of which was based on my best on Public LB and scored 0.50481 on Private LB, and the other which involved some \"clever\" thresholding based on validation set performance but scored only 0.49623 on Private LB.  But I'm pleased with my result, since I went up 89 places from Public to Private and had no aspirations for anything beyond my bronze medal.   I made no use of the external data, having shied away from it because of the leaks and other complications.\nI think it's not worth agonizing over the strategic choices that led to a loss of gold or other precious metals.  Because of the small numbers of examples of the rare classes, I believe that just plain luck played a significant role in determining the final rankings.  To paraphrase Cassius in Shakespeare’s play \"Julius Caesar\" (but reversing the meaning): \"The fault is not in ourselves but in our stars\".",
      "votes": null
    },
    {
      "id": "454883",
      "postDate": "01/12/2019 12:31:17",
      "content": "<p>Why do you think you have correctly detected all duplicates? Btw, ignoring them completely seems a bit wasteful. What's your validation score?</p>",
      "rawMarkdown": "Why do you think you have correctly detected all duplicates? Btw, ignoring them completely seems a bit wasteful. What's your validation score?",
      "votes": null
    },
    {
      "id": "454949",
      "postDate": "01/12/2019 15:44:00",
      "content": "<p>Sorry for you.  One strategy used by many is to select two subs that way: select the best CV score, and select the best LB score.  This minimizes risk.  I have also seen people computing a compound score between CV and LB and using that to select their subs.</p>\n\n<p>I am more like you, trusting CV blindly, unless there is a proven difference between train and test data, like in the last two competitions I entered.  In that case public LB is probably a better yardstick.</p>\n\n<p>I didn't enter this one, hence I have absolutely no clue about train vs test difference.  I might have selected exactly like you.</p>",
      "rawMarkdown": "Sorry for you.  One strategy used by many is to select two subs that way: select the best CV score, and select the best LB score.  This minimizes risk.  I have also seen people computing a compound score between CV and LB and using that to select their subs.\n\nI am more like you, trusting CV blindly, unless there is a proven difference between train and test data, like in the last two competitions I entered.  In that case public LB is probably a better yardstick.\n\nI didn't enter this one, hence I have absolutely no clue about train vs test difference.  I might have selected exactly like you.",
      "votes": null
    },
    {
      "id": "454991",
      "postDate": "01/12/2019 17:46:59",
      "content": "<p><a href=\"/cpmpml\">@cpmpml</a> - yeah, I usually do that but there was a 3rd dimension here with the huge external data that was basically essential to get a good score - I still firmly think trusting one's CV is the better long-term strategy. Unhealthily (or not, depending on your view) there are a lot of kagglers who have just had a huge reinforcement of the view that trusting the public LB is the right thing to do (sometimes yes, but never blindly) and might infer the wrong lessons. Simply, here, the public LB was about 9% of the size of the training data, the competition metric was highly punishing for few examples and there was huge class imbalance =&gt; public LB could be completely misleading. I'm still staggered it held up so well and still haven't fully understood why people (everyone, not just me) witnessed a large drop from scoring well on CV (without external data) to public LB.</p>\n\n<p><a href=\"/artyomp\">@artyomp</a> - I'm sure I didn't correctly identify all the duplicates and I wrote elsewhere that throwing all of them away was likely to be overly wasteful but favoured losing a little information to gain the robustness needed to trust one's CV. Particularly as I suspected some people might not be as careful...at the time I had starting using the hpa data and witnessed how unstable it could be (even though it did give a large public LB bump - this could easily be explained by leakage). CV on just kaggle data was ~0.78 (public LB ~0.5), CV on kaggle + hpa was much lower, maybe 0.6 with public LB ~0.55+.</p>",
      "rawMarkdown": "cpmpml - yeah, I usually do that but there was a 3rd dimension here with the huge external data that was basically essential to get a good score - I still firmly think trusting one's CV is the better long-term strategy. Unhealthily (or not, depending on your view) there are a lot of kagglers who have just had a huge reinforcement of the view that trusting the public LB is the right thing to do (sometimes yes, but never blindly) and might infer the wrong lessons. Simply, here, the public LB was about 9% of the size of the training data, the competition metric was highly punishing for few examples and there was huge class imbalance =&gt; public LB could be completely misleading. I'm still staggered it held up so well and still haven't fully understood why people (everyone, not just me) witnessed a large drop from scoring well on CV (without external data) to public LB.\n\n@artyomp - I'm sure I didn't correctly identify all the duplicates and I wrote elsewhere that throwing all of them away was likely to be overly wasteful but favoured losing a little information to gain the robustness needed to trust one's CV. Particularly as I suspected some people might not be as careful...at the time I had starting using the hpa data and witnessed how unstable it could be (even though it did give a large public LB bump - this could easily be explained by leakage). CV on just kaggle data was ~0.78 (public LB ~0.5), CV on kaggle + hpa was much lower, maybe 0.6 with public LB ~0.55+.",
      "votes": null
    },
    {
      "id": "455023",
      "postDate": "01/12/2019 19:12:48",
      "content": "<blockquote>\n  <p>a lot of kagglers who have just had a huge reinforcement of the view that trusting the public LB is the right thing to do</p>\n</blockquote>\n\n<p>The more, the better, it makes it easier to fare well ;)  </p>",
      "rawMarkdown": "&gt;  a lot of kagglers who have just had a huge reinforcement of the view that trusting the public LB is the right thing to do\n\nThe more, the better, it makes it easier to fare well ;)",
      "votes": null
    },
    {
      "id": "455307",
      "postDate": "01/13/2019 15:03:02",
      "content": "<p>I joined the game very late, I did not have a lot subs to exploit the leaderboard. Before I crank up augmentations I did see that cv and public board are not correlated, but it became directionally aligned after I added more augmentations. I believe random cropping is very useful. With that said, I used cv to select threshold, which still fails.</p>",
      "rawMarkdown": "I joined the game very late, I did not have a lot subs to exploit the leaderboard. Before I crank up augmentations I did see that cv and public board are not correlated, but it became directionally aligned after I added more augmentations. I believe random cropping is very useful. With that said, I used cv to select threshold, which still fails.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 454191,
      "author_name": "tcapelle",
      "author_url": "",
      "post_date": "01/11/2019 08:17:40",
      "content": "<p>I saw exactly the opposite.</p>",
      "votes": null,
      "replies": [
        {
          "id": 454310,
          "author_name": "maw501",
          "author_url": "",
          "post_date": "01/11/2019 12:08:17",
          "content": "<p>Hi <a href=\"/tcapelle\">@tcapelle</a> - care to elaborate? Looking at all our public to private scores the correlation is about 0.9 =&gt; public LB is a 'better' validation set than the much larger training data. Interested to hear what you did.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 454338,
          "author_name": "tcapelle",
          "author_url": "",
          "post_date": "01/11/2019 13:21:20",
          "content": "<p>I am just a noob here, but I had huge variance. My best performing model public LB was 0.553 scored 0.494, but another model performing 0.552 would have scored 0.509. Also our validations scores where high, between 0.55 and 0.6 but hey scored low in public LB, but good on Private... </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 454346,
          "author_name": "maw501",
          "author_url": "",
          "post_date": "01/11/2019 13:34:27",
          "content": "<p>Might depend on if you always added the leak file in or not.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 454259,
      "author_name": "mathormad",
      "author_url": "",
      "post_date": "01/11/2019 10:11:04",
      "content": "<p>I know that feel bro!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 454362,
      "author_name": "atm584",
      "author_url": "",
      "post_date": "01/11/2019 14:15:23",
      "content": "<p>My choices of submissions missed the best Private LB score too. I think building a robust cv set is one of the key problems in this competition. A solid cv set could be helpful not only for selecting the best model&amp;submission but also helpful for tuning the thresholds as well. wonder how bestfitting did this :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 454540,
      "author_name": "artyomp",
      "author_url": "",
      "post_date": "01/11/2019 20:02:40",
      "content": "<p>There were duplicates with different brightness in both train and external data (and between them). This is why validation score was 0.7-0.8 while LB score was about 0.55-0.6 or a bit higher.  So we have actually overfitted the train set.</p>\n\n<p>Not sure why private LB score is generally lower than the public one, though.</p>",
      "votes": null,
      "replies": [
        {
          "id": 454544,
          "author_name": "maw501",
          "author_url": "",
          "post_date": "01/11/2019 20:08:03",
          "content": "<p>I excluded duplicates. Also, I am referring to models not using hpa in the validation set that scored highly with poor LB correspondence.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 454589,
          "author_name": "petewills",
          "author_url": "",
          "post_date": "01/11/2019 21:15:43",
          "content": "<p>Mark, my sympathies as you seem to have done everything in the most rational way. I was leery of the CV (mostly due to comments by you and others and by the mismatch) and spent my time on improvements that would benefit all measures cv/private/public. For example, in the last week I bit the bullet and started running 768-HPA. It was not as hard as I thought it would be (not many epochs) and I ran resnet34 and 50 in that time - previously I ran only resnet34. My best score came from an ensemble of resnet34x2 and resnet50x1. \nFor thresholds, I chose using the public LB and I found that to be actually quite stable (I used a class-variable curve with two free parameters to keep the choice simple and the parameters didn't vary much over the range of models I trained).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 454595,
          "author_name": "maw501",
          "author_url": "",
          "post_date": "01/11/2019 21:28:00",
          "content": "<p>Thanks Pete, glad you placed well and nice to hear a little of what you did - you definitely bring a lot to the discussions. </p>\n\n<p>Yeah, guess it's just one of those things. I feel doubly bad as it was me pushing my team to use one model without HPA as a final submission...I just couldn't shake the feeling of how silly I'd feel if I listened to a ~3.5k public LB score vs. models doing well on almost 9x that data (esp. considering class imbalance,  metric etc..). Lesson 1 of kaggle is meant to be live and die by your CV...well, I just lost a life.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 454883,
          "author_name": "artyomp",
          "author_url": "",
          "post_date": "01/12/2019 12:31:17",
          "content": "<p>Why do you think you have correctly detected all duplicates? Btw, ignoring them completely seems a bit wasteful. What's your validation score?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 454548,
      "author_name": "hortonhearsafoo",
      "author_url": "",
      "post_date": "01/11/2019 20:22:02",
      "content": "<p>Sorry to hear this. I’ve seen it suggested that you should use the best model on your local CV and the best model on the public LB as your two submissions. That could help you hedge your bets and avoid these kind of situations. It’s also likely that your best model on local CV is only marginally better than your second-best, so mitigating the risk might have more reward. Still a tough call in the end though, especially if you’re near the top of the LB fighting for each 0.001</p>",
      "votes": null,
      "replies": [
        {
          "id": 454564,
          "author_name": "maw501",
          "author_url": "",
          "post_date": "01/11/2019 20:42:11",
          "content": "<p>Thanks William! Ordinarily that's what I'd do but the HPA added a 3rd dimension so we went with a model not using HPA at all which trusted the kaggle data (i.e. ignoring the public LB which we didn't think reliable) and then a good scoring model on public LB that we had validated locally too expecting it to be more robust (rather than our best pure public LB tuned models - which obviously with hindsight did best).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 454562,
      "author_name": "spsancti",
      "author_url": "",
      "post_date": "01/11/2019 20:38:41",
      "content": "<p>My best private submission was worse both on public LB and in my local validation...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 454636,
      "author_name": "dslate",
      "author_url": "",
      "post_date": "01/11/2019 22:57:48",
      "content": "<p>My best sub on Private LB scored 0.50979.  Unfortunately it was not one of my final selections, one of which was based on my best on Public LB and scored 0.50481 on Private LB, and the other which involved some \"clever\" thresholding based on validation set performance but scored only 0.49623 on Private LB.  But I'm pleased with my result, since I went up 89 places from Public to Private and had no aspirations for anything beyond my bronze medal.   I made no use of the external data, having shied away from it because of the leaks and other complications.\nI think it's not worth agonizing over the strategic choices that led to a loss of gold or other precious metals.  Because of the small numbers of examples of the rare classes, I believe that just plain luck played a significant role in determining the final rankings.  To paraphrase Cassius in Shakespeare’s play \"Julius Caesar\" (but reversing the meaning): \"The fault is not in ourselves but in our stars\".</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 454949,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "01/12/2019 15:44:00",
      "content": "<p>Sorry for you.  One strategy used by many is to select two subs that way: select the best CV score, and select the best LB score.  This minimizes risk.  I have also seen people computing a compound score between CV and LB and using that to select their subs.</p>\n\n<p>I am more like you, trusting CV blindly, unless there is a proven difference between train and test data, like in the last two competitions I entered.  In that case public LB is probably a better yardstick.</p>\n\n<p>I didn't enter this one, hence I have absolutely no clue about train vs test difference.  I might have selected exactly like you.</p>",
      "votes": null,
      "replies": [
        {
          "id": 454991,
          "author_name": "maw501",
          "author_url": "",
          "post_date": "01/12/2019 17:46:59",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> - yeah, I usually do that but there was a 3rd dimension here with the huge external data that was basically essential to get a good score - I still firmly think trusting one's CV is the better long-term strategy. Unhealthily (or not, depending on your view) there are a lot of kagglers who have just had a huge reinforcement of the view that trusting the public LB is the right thing to do (sometimes yes, but never blindly) and might infer the wrong lessons. Simply, here, the public LB was about 9% of the size of the training data, the competition metric was highly punishing for few examples and there was huge class imbalance =&gt; public LB could be completely misleading. I'm still staggered it held up so well and still haven't fully understood why people (everyone, not just me) witnessed a large drop from scoring well on CV (without external data) to public LB.</p>\n\n<p><a href=\"/artyomp\">@artyomp</a> - I'm sure I didn't correctly identify all the duplicates and I wrote elsewhere that throwing all of them away was likely to be overly wasteful but favoured losing a little information to gain the robustness needed to trust one's CV. Particularly as I suspected some people might not be as careful...at the time I had starting using the hpa data and witnessed how unstable it could be (even though it did give a large public LB bump - this could easily be explained by leakage). CV on just kaggle data was ~0.78 (public LB ~0.5), CV on kaggle + hpa was much lower, maybe 0.6 with public LB ~0.55+.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 455023,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "01/12/2019 19:12:48",
          "content": "<blockquote>\n  <p>a lot of kagglers who have just had a huge reinforcement of the view that trusting the public LB is the right thing to do</p>\n</blockquote>\n\n<p>The more, the better, it makes it easier to fare well ;)  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 455307,
      "author_name": "ryanzhang",
      "author_url": "",
      "post_date": "01/13/2019 15:03:02",
      "content": "<p>I joined the game very late, I did not have a lot subs to exploit the leaderboard. Before I crank up augmentations I did see that cv and public board are not correlated, but it became directionally aligned after I added more augmentations. I believe random cropping is very useful. With that said, I used cv to select threshold, which still fails.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "454166": "[Firstly: **congratulations to all the winners/top scorers!**]\n\nYep, that was us...and I'm not upset (even a little bit). ;)\n\nIt turns out the models we had with the highest public LB also had the highest private LB score (would have won gold based on current scores), correlation ~86%. However we chose models performing (slightly) worse on the public LB but that had been validated on a much bigger validation set (one was the entire training data 5 fold, one was half of it). \n\nBasically it turns out using the public LB as the validation set beat validating models on one's own CV. I'm a little shocked, we were fairly confident the whole public LB was massively overfit (small, unstable metric, huge class imbalance etc...).\n\nWhat was everyone else's strategy for choosing a final model? Anyone else have a similar experience?\n\nMark\n\nP.S. Obviously 'missing out on gold' is conditional on everyone else's final selection remaining fixed.",
    "454191": "I saw exactly the opposite.",
    "454259": "I know that feel bro!",
    "454310": "Hi @tcapelle - care to elaborate? Looking at all our public to private scores the correlation is about 0.9 =&gt; public LB is a 'better' validation set than the much larger training data. Interested to hear what you did.",
    "454338": "I am just a noob here, but I had huge variance. My best performing model public LB was 0.553 scored 0.494, but another model performing 0.552 would have scored 0.509. Also our validations scores where high, between 0.55 and 0.6 but hey scored low in public LB, but good on Private...",
    "454346": "Might depend on if you always added the leak file in or not.",
    "454362": "My choices of submissions missed the best Private LB score too. I think building a robust cv set is one of the key problems in this competition. A solid cv set could be helpful not only for selecting the best model&amp;submission but also helpful for tuning the thresholds as well. wonder how bestfitting did this :)",
    "454540": "There were duplicates with different brightness in both train and external data (and between them). This is why validation score was 0.7-0.8 while LB score was about 0.55-0.6 or a bit higher.  So we have actually overfitted the train set.\n\nNot sure why private LB score is generally lower than the public one, though.",
    "454544": "I excluded duplicates. Also, I am referring to models not using hpa in the validation set that scored highly with poor LB correspondence.",
    "454548": "Sorry to hear this. I’ve seen it suggested that you should use the best model on your local CV and the best model on the public LB as your two submissions. That could help you hedge your bets and avoid these kind of situations. It’s also likely that your best model on local CV is only marginally better than your second-best, so mitigating the risk might have more reward. Still a tough call in the end though, especially if you’re near the top of the LB fighting for each 0.001",
    "454562": "My best private submission was worse both on public LB and in my local validation...",
    "454564": "Thanks William! Ordinarily that's what I'd do but the HPA added a 3rd dimension so we went with a model not using HPA at all which trusted the kaggle data (i.e. ignoring the public LB which we didn't think reliable) and then a good scoring model on public LB that we had validated locally too expecting it to be more robust (rather than our best pure public LB tuned models - which obviously with hindsight did best).",
    "454589": "Mark, my sympathies as you seem to have done everything in the most rational way. I was leery of the CV (mostly due to comments by you and others and by the mismatch) and spent my time on improvements that would benefit all measures cv/private/public. For example, in the last week I bit the bullet and started running 768-HPA. It was not as hard as I thought it would be (not many epochs) and I ran resnet34 and 50 in that time - previously I ran only resnet34. My best score came from an ensemble of resnet34x2 and resnet50x1. \nFor thresholds, I chose using the public LB and I found that to be actually quite stable (I used a class-variable curve with two free parameters to keep the choice simple and the parameters didn't vary much over the range of models I trained).",
    "454595": "Thanks Pete, glad you placed well and nice to hear a little of what you did - you definitely bring a lot to the discussions. \n\nYeah, guess it's just one of those things. I feel doubly bad as it was me pushing my team to use one model without HPA as a final submission...I just couldn't shake the feeling of how silly I'd feel if I listened to a ~3.5k public LB score vs. models doing well on almost 9x that data (esp. considering class imbalance,  metric etc..). Lesson 1 of kaggle is meant to be live and die by your CV...well, I just lost a life.",
    "454636": "My best sub on Private LB scored 0.50979.  Unfortunately it was not one of my final selections, one of which was based on my best on Public LB and scored 0.50481 on Private LB, and the other which involved some \"clever\" thresholding based on validation set performance but scored only 0.49623 on Private LB.  But I'm pleased with my result, since I went up 89 places from Public to Private and had no aspirations for anything beyond my bronze medal.   I made no use of the external data, having shied away from it because of the leaks and other complications.\nI think it's not worth agonizing over the strategic choices that led to a loss of gold or other precious metals.  Because of the small numbers of examples of the rare classes, I believe that just plain luck played a significant role in determining the final rankings.  To paraphrase Cassius in Shakespeare’s play \"Julius Caesar\" (but reversing the meaning): \"The fault is not in ourselves but in our stars\".",
    "454883": "Why do you think you have correctly detected all duplicates? Btw, ignoring them completely seems a bit wasteful. What's your validation score?",
    "454949": "Sorry for you.  One strategy used by many is to select two subs that way: select the best CV score, and select the best LB score.  This minimizes risk.  I have also seen people computing a compound score between CV and LB and using that to select their subs.\n\nI am more like you, trusting CV blindly, unless there is a proven difference between train and test data, like in the last two competitions I entered.  In that case public LB is probably a better yardstick.\n\nI didn't enter this one, hence I have absolutely no clue about train vs test difference.  I might have selected exactly like you.",
    "454991": "cpmpml - yeah, I usually do that but there was a 3rd dimension here with the huge external data that was basically essential to get a good score - I still firmly think trusting one's CV is the better long-term strategy. Unhealthily (or not, depending on your view) there are a lot of kagglers who have just had a huge reinforcement of the view that trusting the public LB is the right thing to do (sometimes yes, but never blindly) and might infer the wrong lessons. Simply, here, the public LB was about 9% of the size of the training data, the competition metric was highly punishing for few examples and there was huge class imbalance =&gt; public LB could be completely misleading. I'm still staggered it held up so well and still haven't fully understood why people (everyone, not just me) witnessed a large drop from scoring well on CV (without external data) to public LB.\n\n@artyomp - I'm sure I didn't correctly identify all the duplicates and I wrote elsewhere that throwing all of them away was likely to be overly wasteful but favoured losing a little information to gain the robustness needed to trust one's CV. Particularly as I suspected some people might not be as careful...at the time I had starting using the hpa data and witnessed how unstable it could be (even though it did give a large public LB bump - this could easily be explained by leakage). CV on just kaggle data was ~0.78 (public LB ~0.5), CV on kaggle + hpa was much lower, maybe 0.6 with public LB ~0.55+.",
    "455023": "&gt;  a lot of kagglers who have just had a huge reinforcement of the view that trusting the public LB is the right thing to do\n\nThe more, the better, it makes it easier to fare well ;)",
    "455307": "I joined the game very late, I did not have a lot subs to exploit the leaderboard. Before I crank up augmentations I did see that cv and public board are not correlated, but it became directionally aligned after I added more augmentations. I believe random cropping is very useful. With that said, I used cv to select threshold, which still fails."
  },
  "source": "meta"
}