{
  "id": 189189,
  "title": "Silver medal in three days of work",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/writeups/seifeddine-fezzani-silver-medal-in-three-days-of-w",
  "author_name": "",
  "post_date": "2020-10-07T08:11:05.640Z",
  "votes": 4,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Congrats to everyone and thank you Kaggle for hosting this competition.<br>\nThis contest was by far one of the most realistic ones. </p>\n<p>I decided to join late this competition after I saw the public kernels, it was obvious that people were overfitting the public leaderboard. Then I decided to find the best validation strategy and use it to find new features and tune hyper-parameters.</p>\n<h2>Validation strategy</h2>\n<p>I used GroupKFold to create a K-fold partition (K == 3) with non-overlapping groups (here the Patient ID) and the distribution of patients across the folds is close to being uniform. (explanation and code took from <a href=\"https://www.kaggle.com/rftexas\" target=\"_blank\">@rftexas</a> 's notebook.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1803801%2Fff68dceee4b8e5de81f1bcb5cc7dc1ca%2FScreen%20Shot%202020-10-07%20at%201.17.31%20AM.png?generation=1602030140932166&amp;alt=media\" alt=\"\"></p>\n<p>Then like I mentioned in this discussion, I didn't include the Percent feature in my validation folds and I only included the last three measurements for each patient. Why ? because I want to mimic exactly the testing dataset, since we don't have too many features, using Percent improved CV and Lb scores. BUT we only have initial Percent values in test. So I decided to create init_Percent feature that I used in my validation folds.</p>\n<p>With that I was able to have a good correlation between CV and LB scores, with that I was more confident.</p>\n<h2>Features used</h2>\n<p>I only used the tabular dataset and I kept features that improved my validation score. I didn't include weeks and base_week. <br>\nI created init_FVC and init_Percent features, there are the initial values of FVC and Percent for each patient.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1803801%2F3f5a07f25b7a753c48d1807f173ae01b%2FScreen%20Shot%202020-10-07%20at%201.26.32%20AM.png?generation=1602030491883502&amp;alt=media\" alt=\"\"></p>\n<p>I added Sex_SmokingStatus <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1803801%2Fc3965f9c27c8278dfd6c2fed72ff62ae%2FScreen%20Shot%202020-10-07%20at%201.26.36%20AM.png?generation=1602030521844243&amp;alt=media\" alt=\"\"></p>\n<p>I used group statistics features<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1803801%2Ff93fb7f94f4636fe054e1c8133a4f94f%2FScreen%20Shot%202020-10-07%20at%201.26.26%20AM.png?generation=1602030547099121&amp;alt=media\" alt=\"\"></p>\n<h2>Model used</h2>\n<p>I used a simple neural network with pinball loss for multiple quantiles using [0.2, 0.5, 0.8].</p>\n<h2>What didn't work</h2>\n<p>I tried to extract features from images but I wasn't able to improve my CV score. <br>\nI also tried others models like linear regression and lightgbm but they performed poorly. Even ensembling them didn't improve my results.</p>\n<p>I didn't spent a lot of time in this competition, I didn't find time to improve this solution. </p>\n<h2>Key for this competition</h2>\n<p>Find the validation strategy that gives you a correlation between CV and LB scores !!!</p>\n<p>Again, thank you for this competition and for reading my solution, it is only a top 100 solution but it means a lot to me because I will be a kaggle competition master :p </p>",
  "messages": [
    {
      "id": "1040014",
      "postDate": "10/07/2020 00:35:11",
      "content": "<p>Congrats to everyone and thank you Kaggle for hosting this competition.<br>\nThis contest was by far one of the most realistic ones. </p>\n<p>I decided to join late this competition after I saw the public kernels, it was obvious that people were overfitting the public leaderboard. Then I decided to find the best validation strategy and use it to find new features and tune hyper-parameters.</p>\n<h2>Validation strategy</h2>\n<p>I used GroupKFold to create a K-fold partition (K == 3) with non-overlapping groups (here the Patient ID) and the distribution of patients across the folds is close to being uniform. (explanation and code took from <a href=\"https://www.kaggle.com/rftexas\" target=\"_blank\">@rftexas</a> 's notebook.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1803801%2Fff68dceee4b8e5de81f1bcb5cc7dc1ca%2FScreen%20Shot%202020-10-07%20at%201.17.31%20AM.png?generation=1602030140932166&amp;alt=media\" alt=\"\"></p>\n<p>Then like I mentioned in this discussion, I didn't include the Percent feature in my validation folds and I only included the last three measurements for each patient. Why ? because I want to mimic exactly the testing dataset, since we don't have too many features, using Percent improved CV and Lb scores. BUT we only have initial Percent values in test. So I decided to create init_Percent feature that I used in my validation folds.</p>\n<p>With that I was able to have a good correlation between CV and LB scores, with that I was more confident.</p>\n<h2>Features used</h2>\n<p>I only used the tabular dataset and I kept features that improved my validation score. I didn't include weeks and base_week. <br>\nI created init_FVC and init_Percent features, there are the initial values of FVC and Percent for each patient.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1803801%2F3f5a07f25b7a753c48d1807f173ae01b%2FScreen%20Shot%202020-10-07%20at%201.26.32%20AM.png?generation=1602030491883502&amp;alt=media\" alt=\"\"></p>\n<p>I added Sex_SmokingStatus <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1803801%2Fc3965f9c27c8278dfd6c2fed72ff62ae%2FScreen%20Shot%202020-10-07%20at%201.26.36%20AM.png?generation=1602030521844243&amp;alt=media\" alt=\"\"></p>\n<p>I used group statistics features<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1803801%2Ff93fb7f94f4636fe054e1c8133a4f94f%2FScreen%20Shot%202020-10-07%20at%201.26.26%20AM.png?generation=1602030547099121&amp;alt=media\" alt=\"\"></p>\n<h2>Model used</h2>\n<p>I used a simple neural network with pinball loss for multiple quantiles using [0.2, 0.5, 0.8].</p>\n<h2>What didn't work</h2>\n<p>I tried to extract features from images but I wasn't able to improve my CV score. <br>\nI also tried others models like linear regression and lightgbm but they performed poorly. Even ensembling them didn't improve my results.</p>\n<p>I didn't spent a lot of time in this competition, I didn't find time to improve this solution. </p>\n<h2>Key for this competition</h2>\n<p>Find the validation strategy that gives you a correlation between CV and LB scores !!!</p>\n<p>Again, thank you for this competition and for reading my solution, it is only a top 100 solution but it means a lot to me because I will be a kaggle competition master :p </p>",
      "rawMarkdown": "Congrats to everyone and thank you Kaggle for hosting this competition.\nThis contest was by far one of the most realistic ones. \n\nI decided to join late this competition after I saw the public kernels, it was obvious that people were overfitting the public leaderboard. Then I decided to find the best validation strategy and use it to find new features and tune hyper-parameters.\n\n##Validation strategy\n\nI used GroupKFold to create a K-fold partition (K == 3) with non-overlapping groups (here the Patient ID) and the distribution of patients across the folds is close to being uniform. (explanation and code took from @rftexas 's notebook.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1803801%2Fff68dceee4b8e5de81f1bcb5cc7dc1ca%2FScreen%20Shot%202020-10-07%20at%201.17.31%20AM.png?generation=1602030140932166&alt=media)\n\nThen like I mentioned in this discussion, I didn't include the Percent feature in my validation folds and I only included the last three measurements for each patient. Why ? because I want to mimic exactly the testing dataset, since we don't have too many features, using Percent improved CV and Lb scores. BUT we only have initial Percent values in test. So I decided to create init_Percent feature that I used in my validation folds.\n\n\nWith that I was able to have a good correlation between CV and LB scores, with that I was more confident.\n\n##Features used\n\nI only used the tabular dataset and I kept features that improved my validation score. I didn't include weeks and base_week. \nI created init_FVC and init_Percent features, there are the initial values of FVC and Percent for each patient.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1803801%2F3f5a07f25b7a753c48d1807f173ae01b%2FScreen%20Shot%202020-10-07%20at%201.26.32%20AM.png?generation=1602030491883502&alt=media)\n\nI added Sex_SmokingStatus \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1803801%2Fc3965f9c27c8278dfd6c2fed72ff62ae%2FScreen%20Shot%202020-10-07%20at%201.26.36%20AM.png?generation=1602030521844243&alt=media)\n\n\nI used group statistics features\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1803801%2Ff93fb7f94f4636fe054e1c8133a4f94f%2FScreen%20Shot%202020-10-07%20at%201.26.26%20AM.png?generation=1602030547099121&alt=media)\n\n## Model used\nI used a simple neural network with pinball loss for multiple quantiles using [0.2, 0.5, 0.8].\n\n##What didn't work\nI tried to extract features from images but I wasn't able to improve my CV score. \nI also tried others models like linear regression and lightgbm but they performed poorly. Even ensembling them didn't improve my results.\n\nI didn't spent a lot of time in this competition, I didn't find time to improve this solution. \n\n##Key for this competition\nFind the validation strategy that gives you a correlation between CV and LB scores !!!\n\nAgain, thank you for this competition and for reading my solution, it is only a top 100 solution but it means a lot to me because I will be a kaggle competition master :p",
      "votes": null
    },
    {
      "id": "1040030",
      "postDate": "10/07/2020 00:45:22",
      "content": "<p>Congratulation. </p>",
      "rawMarkdown": "Congratulation.",
      "votes": null
    },
    {
      "id": "1040166",
      "postDate": "10/07/2020 02:23:30",
      "content": "<p>Congrats &amp; Well Done !</p>",
      "rawMarkdown": "Congrats & Well Done !",
      "votes": null
    },
    {
      "id": "1040190",
      "postDate": "10/07/2020 02:51:22",
      "content": "<p>Congratulations! looks like people who spent more and more time they ended up overfitting!</p>",
      "rawMarkdown": "Congratulations! looks like people who spent more and more time they ended up overfitting!",
      "votes": null
    },
    {
      "id": "1040203",
      "postDate": "10/07/2020 02:59:31",
      "content": "<p>All teams who focused on improving the CV would have done well. I feel the public LB is a double edge sword, especially in competitions like these with only 15% in public test. It gives competitors a false sense of progress when they overfit.</p>",
      "rawMarkdown": "All teams who focused on improving the CV would have done well. I feel the public LB is a double edge sword, especially in competitions like these with only 15% in public test. It gives competitors a false sense of progress when they overfit.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1040030,
      "author_name": "koalasheep",
      "author_url": "",
      "post_date": "10/07/2020 00:45:22",
      "content": "<p>Congratulation. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1040166,
      "author_name": "sidneyng",
      "author_url": "",
      "post_date": "10/07/2020 02:23:30",
      "content": "<p>Congrats &amp; Well Done !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1040190,
      "author_name": "prasunmishra",
      "author_url": "",
      "post_date": "10/07/2020 02:51:22",
      "content": "<p>Congratulations! looks like people who spent more and more time they ended up overfitting!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1040203,
          "author_name": "abhishekgbhat",
          "author_url": "",
          "post_date": "10/07/2020 02:59:31",
          "content": "<p>All teams who focused on improving the CV would have done well. I feel the public LB is a double edge sword, especially in competitions like these with only 15% in public test. It gives competitors a false sense of progress when they overfit.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1040014": "Congrats to everyone and thank you Kaggle for hosting this competition.\nThis contest was by far one of the most realistic ones. \n\nI decided to join late this competition after I saw the public kernels, it was obvious that people were overfitting the public leaderboard. Then I decided to find the best validation strategy and use it to find new features and tune hyper-parameters.\n\n##Validation strategy\n\nI used GroupKFold to create a K-fold partition (K == 3) with non-overlapping groups (here the Patient ID) and the distribution of patients across the folds is close to being uniform. (explanation and code took from @rftexas 's notebook.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1803801%2Fff68dceee4b8e5de81f1bcb5cc7dc1ca%2FScreen%20Shot%202020-10-07%20at%201.17.31%20AM.png?generation=1602030140932166&alt=media)\n\nThen like I mentioned in this discussion, I didn't include the Percent feature in my validation folds and I only included the last three measurements for each patient. Why ? because I want to mimic exactly the testing dataset, since we don't have too many features, using Percent improved CV and Lb scores. BUT we only have initial Percent values in test. So I decided to create init_Percent feature that I used in my validation folds.\n\n\nWith that I was able to have a good correlation between CV and LB scores, with that I was more confident.\n\n##Features used\n\nI only used the tabular dataset and I kept features that improved my validation score. I didn't include weeks and base_week. \nI created init_FVC and init_Percent features, there are the initial values of FVC and Percent for each patient.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1803801%2F3f5a07f25b7a753c48d1807f173ae01b%2FScreen%20Shot%202020-10-07%20at%201.26.32%20AM.png?generation=1602030491883502&alt=media)\n\nI added Sex_SmokingStatus \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1803801%2Fc3965f9c27c8278dfd6c2fed72ff62ae%2FScreen%20Shot%202020-10-07%20at%201.26.36%20AM.png?generation=1602030521844243&alt=media)\n\n\nI used group statistics features\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1803801%2Ff93fb7f94f4636fe054e1c8133a4f94f%2FScreen%20Shot%202020-10-07%20at%201.26.26%20AM.png?generation=1602030547099121&alt=media)\n\n## Model used\nI used a simple neural network with pinball loss for multiple quantiles using [0.2, 0.5, 0.8].\n\n##What didn't work\nI tried to extract features from images but I wasn't able to improve my CV score. \nI also tried others models like linear regression and lightgbm but they performed poorly. Even ensembling them didn't improve my results.\n\nI didn't spent a lot of time in this competition, I didn't find time to improve this solution. \n\n##Key for this competition\nFind the validation strategy that gives you a correlation between CV and LB scores !!!\n\nAgain, thank you for this competition and for reading my solution, it is only a top 100 solution but it means a lot to me because I will be a kaggle competition master :p",
    "1040030": "Congratulation.",
    "1040166": "Congrats & Well Done !",
    "1040190": "Congratulations! looks like people who spent more and more time they ended up overfitting!",
    "1040203": "All teams who focused on improving the CV would have done well. I feel the public LB is a double edge sword, especially in competitions like these with only 15% in public test. It gives competitors a false sense of progress when they overfit."
  },
  "source": "meta"
}