{
  "id": 189198,
  "title": "How I missed the Gold: Overfitting to CV was NOT good here",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/189198",
  "author_name": "",
  "post_date": "2020-10-07T00:56:44.178473300Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Thanks for organizers to host this great competition, and good job for those who survived a big shake-up! </p>\n<p>This was a competition with medical data of a rare disease, in which the number of patients is relatively small. So the final standing was hard to predict, as the training sample might represent nothing about the private dataset.</p>\n<p>My concern was exactly this: <strong>Trusting LB is out of the question but should I really trust CV?</strong></p>\n<p>Eventually I trusted my CV, and got the solo Bronze medal, which is alright. My solution used both tabular and image data, and first predicting FVC with many models (2DCNN, MLP, GBDT, linear…) then optimized confidence also with many models. </p>\n<p>Eventually it turned out that I also have <strong>a gold medal submission, which I could not choose because its LB as well as CV were not great</strong>. Basically the main difference between the two was </p>\n<ul>\n<li>stacking with linear model (Bayesian Ridge) --&gt; gold (CV: -6.9986, public: -6.9613, private: -6.8379)</li>\n<li>multilayer stacking (MLP+XGB -&gt; Bayesian Ridge) --&gt; bronze (CV: -6.9415, public: -6.9822, private: -6.8680)</li>\n</ul>\n<p>Apparently the latter improves CV but also overfits to the training samples. The former had relatively worse CV but better LB.</p>\n<p>So how I missed my first gold medal? </p>\n<p>The answer seems to me that <strong>I overfitted the training samples and trusted my CV too much</strong>. </p>\n<p>I should have taken into account the danger of overfitting to the training samples in this kind of datasets with small samples, and should have picked up a relatively robust way of ensemble such as stacking with linear model. </p>",
  "messages": [
    {
      "id": "1040056",
      "postDate": "10/07/2020 00:56:44",
      "content": "<p>Thanks for organizers to host this great competition, and good job for those who survived a big shake-up! </p>\n<p>This was a competition with medical data of a rare disease, in which the number of patients is relatively small. So the final standing was hard to predict, as the training sample might represent nothing about the private dataset.</p>\n<p>My concern was exactly this: <strong>Trusting LB is out of the question but should I really trust CV?</strong></p>\n<p>Eventually I trusted my CV, and got the solo Bronze medal, which is alright. My solution used both tabular and image data, and first predicting FVC with many models (2DCNN, MLP, GBDT, linear…) then optimized confidence also with many models. </p>\n<p>Eventually it turned out that I also have <strong>a gold medal submission, which I could not choose because its LB as well as CV were not great</strong>. Basically the main difference between the two was </p>\n<ul>\n<li>stacking with linear model (Bayesian Ridge) --&gt; gold (CV: -6.9986, public: -6.9613, private: -6.8379)</li>\n<li>multilayer stacking (MLP+XGB -&gt; Bayesian Ridge) --&gt; bronze (CV: -6.9415, public: -6.9822, private: -6.8680)</li>\n</ul>\n<p>Apparently the latter improves CV but also overfits to the training samples. The former had relatively worse CV but better LB.</p>\n<p>So how I missed my first gold medal? </p>\n<p>The answer seems to me that <strong>I overfitted the training samples and trusted my CV too much</strong>. </p>\n<p>I should have taken into account the danger of overfitting to the training samples in this kind of datasets with small samples, and should have picked up a relatively robust way of ensemble such as stacking with linear model. </p>",
      "rawMarkdown": "Thanks for organizers to host this great competition, and good job for those who survived a big shake-up! \n\nThis was a competition with medical data of a rare disease, in which the number of patients is relatively small. So the final standing was hard to predict, as the training sample might represent nothing about the private dataset.\n\nMy concern was exactly this: **Trusting LB is out of the question but should I really trust CV?**\n\nEventually I trusted my CV, and got the solo Bronze medal, which is alright. My solution used both tabular and image data, and first predicting FVC with many models (2DCNN, MLP, GBDT, linear...) then optimized confidence also with many models. \n\nEventually it turned out that I also have **a gold medal submission, which I could not choose because its LB as well as CV were not great**. Basically the main difference between the two was \n\n- stacking with linear model (Bayesian Ridge) --> gold (CV: -6.9986, public: -6.9613, private: -6.8379)\n- multilayer stacking (MLP+XGB -> Bayesian Ridge) --> bronze (CV: -6.9415, public: -6.9822, private: -6.8680)\n\nApparently the latter improves CV but also overfits to the training samples. The former had relatively worse CV but better LB.\n\nSo how I missed my first gold medal? \n\nThe answer seems to me that **I overfitted the training samples and trusted my CV too much**. \n\nI should have taken into account the danger of overfitting to the training samples in this kind of datasets with small samples, and should have picked up a relatively robust way of ensemble such as stacking with linear model.",
      "votes": null
    },
    {
      "id": "1040139",
      "postDate": "10/07/2020 01:58:33",
      "content": "<p>Hi, can i know the details of your validation set in cv? as in was containing all the weeks or just latest 3 weeks datapoints?</p>",
      "rawMarkdown": "Hi, can i know the details of your validation set in cv? as in was containing all the weeks or just latest 3 weeks datapoints?",
      "votes": null
    },
    {
      "id": "1040160",
      "postDate": "10/07/2020 02:17:27",
      "content": "<ul>\n<li>For FVC prediction: all weeks (stratified group kfold based on the degree of linear decay and patient as a group)</li>\n<li>For optimized Confidence prediction: last 3 datapoints (same above but replacing non last 3 datapoints with randomly chosen last 3 datapoints)</li>\n</ul>\n<p>See my solution (please have a look at <strong>Version 1</strong>) <a href=\"https://www.kaggle.com/code1110/osic-cv-tab-cnn-ensemble-ocne-nop-lasts#Final-CV-score\" target=\"_blank\">here</a>. </p>",
      "rawMarkdown": "For FVC prediction: all weeks (stratified group kfold based on the degree of linear decay and patient as a group)\n- For optimized Confidence prediction: last 3 datapoints (same above but replacing non last 3 datapoints with randomly chosen last 3 datapoints)\n\nSee my solution (please have a look at **Version 1**) [here](https://www.kaggle.com/code1110/osic-cv-tab-cnn-ensemble-ocne-nop-lasts#Final-CV-score).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1040139,
      "author_name": "mugdhahardikar",
      "author_url": "",
      "post_date": "10/07/2020 01:58:33",
      "content": "<p>Hi, can i know the details of your validation set in cv? as in was containing all the weeks or just latest 3 weeks datapoints?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1040160,
          "author_name": "code1110",
          "author_url": "",
          "post_date": "10/07/2020 02:17:27",
          "content": "<ul>\n<li>For FVC prediction: all weeks (stratified group kfold based on the degree of linear decay and patient as a group)</li>\n<li>For optimized Confidence prediction: last 3 datapoints (same above but replacing non last 3 datapoints with randomly chosen last 3 datapoints)</li>\n</ul>\n<p>See my solution (please have a look at <strong>Version 1</strong>) <a href=\"https://www.kaggle.com/code1110/osic-cv-tab-cnn-ensemble-ocne-nop-lasts#Final-CV-score\" target=\"_blank\">here</a>. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1040056": "Thanks for organizers to host this great competition, and good job for those who survived a big shake-up! \n\nThis was a competition with medical data of a rare disease, in which the number of patients is relatively small. So the final standing was hard to predict, as the training sample might represent nothing about the private dataset.\n\nMy concern was exactly this: **Trusting LB is out of the question but should I really trust CV?**\n\nEventually I trusted my CV, and got the solo Bronze medal, which is alright. My solution used both tabular and image data, and first predicting FVC with many models (2DCNN, MLP, GBDT, linear...) then optimized confidence also with many models. \n\nEventually it turned out that I also have **a gold medal submission, which I could not choose because its LB as well as CV were not great**. Basically the main difference between the two was \n\n- stacking with linear model (Bayesian Ridge) --> gold (CV: -6.9986, public: -6.9613, private: -6.8379)\n- multilayer stacking (MLP+XGB -> Bayesian Ridge) --> bronze (CV: -6.9415, public: -6.9822, private: -6.8680)\n\nApparently the latter improves CV but also overfits to the training samples. The former had relatively worse CV but better LB.\n\nSo how I missed my first gold medal? \n\nThe answer seems to me that **I overfitted the training samples and trusted my CV too much**. \n\nI should have taken into account the danger of overfitting to the training samples in this kind of datasets with small samples, and should have picked up a relatively robust way of ensemble such as stacking with linear model.",
    "1040139": "Hi, can i know the details of your validation set in cv? as in was containing all the weeks or just latest 3 weeks datapoints?",
    "1040160": "For FVC prediction: all weeks (stratified group kfold based on the degree of linear decay and patient as a group)\n- For optimized Confidence prediction: last 3 datapoints (same above but replacing non last 3 datapoints with randomly chosen last 3 datapoints)\n\nSee my solution (please have a look at **Version 1**) [here](https://www.kaggle.com/code1110/osic-cv-tab-cnn-ensemble-ocne-nop-lasts#Final-CV-score)."
  },
  "source": "meta"
}