{
  "id": 190029,
  "title": "Methods and ideas used in the competition",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/190029",
  "author_name": "",
  "post_date": "2020-10-09T19:11:01.221402400Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hallo,</p>\n<p>First of all I would like to thank the organizers, as well as Kaggle team, for giving us the opportunity to take part in this competition.</p>\n<p>I would like to share the strategies I used.</p>\n<p>I used different models : a neural network, gradient boosting algorithms, as well as <a href=\"https://www.kaggle.com/ulrich07/osic-multiple-quantile-regression-starter\" target=\"_blank\">this notebook</a>. For the last one, I tried several batch sizes and I found the best result was obtained for batch_size=64.</p>\n<p>On all of them, I tried these strategies : </p>\n<ol>\n<li>I did data augmentation, namely predicting values for FVC for new weeks,  using Facebook Prophet library. I gave more details in <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/186448\" target=\"_blank\">this discussion</a>. I tried different settings and chose the best result for each sample. This method did not provide a good result.</li>\n<li>I worked on the images and performed lung segmentation, got masks for each image, then parsed the volumes for left and right lung, then parsed the mean, min and max value, for each patient. I added these as new features, namely : <code>volume_mean_left</code>, <code>volume_mean_right</code>, <code>volume_min_left</code>, <code>volume_min_right</code>, <code>volume_max_left</code>, <code>volume_max_right</code>. I parsed the mean, max and min values only for the masks which were not empty. This strategy did not provide a good result neither.</li>\n<li>I performed target encoding, with respect to FVC, for the IDs of the patients.</li>\n<li>Using different models, I predicted FVC first, then Percent. Well, Predict is an uncertainty measure with respect to FVC, therefore I find a little weird to predict with a regression algorithm. Yet I made this test too. </li>\n</ol>",
  "messages": [
    {
      "id": "1044445",
      "postDate": "10/09/2020 19:11:01",
      "content": "<p>Hallo,</p>\n<p>First of all I would like to thank the organizers, as well as Kaggle team, for giving us the opportunity to take part in this competition.</p>\n<p>I would like to share the strategies I used.</p>\n<p>I used different models : a neural network, gradient boosting algorithms, as well as <a href=\"https://www.kaggle.com/ulrich07/osic-multiple-quantile-regression-starter\" target=\"_blank\">this notebook</a>. For the last one, I tried several batch sizes and I found the best result was obtained for batch_size=64.</p>\n<p>On all of them, I tried these strategies : </p>\n<ol>\n<li>I did data augmentation, namely predicting values for FVC for new weeks,  using Facebook Prophet library. I gave more details in <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/186448\" target=\"_blank\">this discussion</a>. I tried different settings and chose the best result for each sample. This method did not provide a good result.</li>\n<li>I worked on the images and performed lung segmentation, got masks for each image, then parsed the volumes for left and right lung, then parsed the mean, min and max value, for each patient. I added these as new features, namely : <code>volume_mean_left</code>, <code>volume_mean_right</code>, <code>volume_min_left</code>, <code>volume_min_right</code>, <code>volume_max_left</code>, <code>volume_max_right</code>. I parsed the mean, max and min values only for the masks which were not empty. This strategy did not provide a good result neither.</li>\n<li>I performed target encoding, with respect to FVC, for the IDs of the patients.</li>\n<li>Using different models, I predicted FVC first, then Percent. Well, Predict is an uncertainty measure with respect to FVC, therefore I find a little weird to predict with a regression algorithm. Yet I made this test too. </li>\n</ol>",
      "rawMarkdown": "Hallo,\n\nFirst of all I would like to thank the organizers, as well as Kaggle team, for giving us the opportunity to take part in this competition.\n\nI would like to share the strategies I used.\n\nI used different models : a neural network, gradient boosting algorithms, as well as [this notebook](https://www.kaggle.com/ulrich07/osic-multiple-quantile-regression-starter). For the last one, I tried several batch sizes and I found the best result was obtained for batch_size=64.\n\nOn all of them, I tried these strategies : \n\n1. I did data augmentation, namely predicting values for FVC for new weeks,  using Facebook Prophet library. I gave more details in [this discussion](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/186448). I tried different settings and chose the best result for each sample. This method did not provide a good result.\n2. I worked on the images and performed lung segmentation, got masks for each image, then parsed the volumes for left and right lung, then parsed the mean, min and max value, for each patient. I added these as new features, namely : `volume_mean_left`, `volume_mean_right`, `volume_min_left`, `volume_min_right`, `volume_max_left`, `volume_max_right`. I parsed the mean, max and min values only for the masks which were not empty. This strategy did not provide a good result neither.\n3. I performed target encoding, with respect to FVC, for the IDs of the patients.\n4. Using different models, I predicted FVC first, then Percent. Well, Predict is an uncertainty measure with respect to FVC, therefore I find a little weird to predict with a regression algorithm. Yet I made this test too.",
      "votes": null
    },
    {
      "id": "1047525",
      "postDate": "10/12/2020 17:28:50",
      "content": "<p>Thanks for sharing your experience on that very interesting competition!</p>",
      "rawMarkdown": "Thanks for sharing your experience on that very interesting competition!",
      "votes": null
    },
    {
      "id": "1047545",
      "postDate": "10/12/2020 17:47:13",
      "content": "<p>I hope it might be useful for other people and further work. I did not get a good score, though. I must have missed something, somewhere. </p>",
      "rawMarkdown": "I hope it might be useful for other people and further work. I did not get a good score, though. I must have missed something, somewhere.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1047525,
      "author_name": "douglaskgaraujo",
      "author_url": "",
      "post_date": "10/12/2020 17:28:50",
      "content": "<p>Thanks for sharing your experience on that very interesting competition!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1047545,
          "author_name": "catadanna",
          "author_url": "",
          "post_date": "10/12/2020 17:47:13",
          "content": "<p>I hope it might be useful for other people and further work. I did not get a good score, though. I must have missed something, somewhere. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1044445": "Hallo,\n\nFirst of all I would like to thank the organizers, as well as Kaggle team, for giving us the opportunity to take part in this competition.\n\nI would like to share the strategies I used.\n\nI used different models : a neural network, gradient boosting algorithms, as well as [this notebook](https://www.kaggle.com/ulrich07/osic-multiple-quantile-regression-starter). For the last one, I tried several batch sizes and I found the best result was obtained for batch_size=64.\n\nOn all of them, I tried these strategies : \n\n1. I did data augmentation, namely predicting values for FVC for new weeks,  using Facebook Prophet library. I gave more details in [this discussion](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/186448). I tried different settings and chose the best result for each sample. This method did not provide a good result.\n2. I worked on the images and performed lung segmentation, got masks for each image, then parsed the volumes for left and right lung, then parsed the mean, min and max value, for each patient. I added these as new features, namely : `volume_mean_left`, `volume_mean_right`, `volume_min_left`, `volume_min_right`, `volume_max_left`, `volume_max_right`. I parsed the mean, max and min values only for the masks which were not empty. This strategy did not provide a good result neither.\n3. I performed target encoding, with respect to FVC, for the IDs of the patients.\n4. Using different models, I predicted FVC first, then Percent. Well, Predict is an uncertainty measure with respect to FVC, therefore I find a little weird to predict with a regression algorithm. Yet I made this test too.",
    "1047525": "Thanks for sharing your experience on that very interesting competition!",
    "1047545": "I hope it might be useful for other people and further work. I did not get a good score, though. I must have missed something, somewhere."
  },
  "source": "meta"
}