{
  "id": 189365,
  "title": "My first Silver Medal on Kaggle! (65th on Private LB)",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/writeups/naresh-omega-my-first-silver-medal-on-kaggle-65th-",
  "author_name": "",
  "post_date": "2020-10-07T12:01:52.713773Z",
  "votes": 3,
  "comment_count": 1,
  "views": 0,
  "content": "<p>My hearty congratulations to all the winners of this competition. Quite honestly I have to admit that I gave up any hopes of winning this competition. I simply tried tinkering around with models &amp; ideas. And when I found that I had secured my first silver at my first ever serious kaggle competition, my joy knew no bounds! I had a public LB rank of 1480, the last time I checked, and I woke up to find a rank of 65. </p>\n<p>Here's a link to my quick inference Notebook: <a href=\"https://www.kaggle.com/doctorkael/osic-inference?scriptVersionId=44167954\" target=\"_blank\">https://www.kaggle.com/doctorkael/osic-inference?scriptVersionId=44167954</a></p>\n<p>My winning solution can be summarized as follows:</p>\n<p><strong>1. Perform data augmentation:</strong> Create synthetic week FVC from other patients data who have similar <em>cosine similarity</em>. For this purpose, we calculate the various features to ensure that augmented data is as similar as possible to the actual patients data. I also tried another type of augmentation with simple interpolation methods such as akima, linear, slinear, etc and used cross validation to see which performed the best at run time and use it to create models which would make the test predictions.  </p>\n<p><strong>2. Some more augmentations</strong> (sort of): We treat each week's FVC and Percent as <em>Baseline</em> values from which  we try to predict data for other weeks. This idea was borrowed from <a href=\"https://www.kaggle.com/yasufuminakama/osic-lgb-baseline\" target=\"_blank\">here</a>. It helps creating models that are more robust and helps when we have to predict for weeks -12 and the like when all the model would have seen otherwise is week -5. Using <code>Week_Offset</code> as feature instead of <code>Weeks</code> really helped here.</p>\n<p><strong>3. The model:</strong> The secret sauce was to use a simple Gradient boosting regressor to predict the confidence intervals. Methods on how to accomplish this can be found <a href=\"https://scikit-learn.org/stable/auto_examples/ensemble/plot_gradient_boosting_quantile.html#sphx-glr-auto-examples-ensemble-plot-gradient-boosting-quantile-py\" target=\"_blank\">here</a>. Choosing an <code>alpha</code> value of 0.75 worked best. We create models on all the different augmented data samples, cross validate them to see which perform best with least error in patient scores (since we only have 3 final predictions which would actually count in scoring, we need to minimize this).</p>\n<p><strong>4. Making Predictions:</strong> The maximum (for upper bound) and minimum (for lower bound) of all the predictions made for that particular patient and week is calculated, subtracted to obtain the the <code>Confidence</code> values.</p>\n<p>I would soon publish my notebook containing all the ideas I tinkered around with soon as it is still in an unedited format. </p>",
  "messages": [
    {
      "id": "1040855",
      "postDate": "10/07/2020 12:01:52",
      "content": "<p>My hearty congratulations to all the winners of this competition. Quite honestly I have to admit that I gave up any hopes of winning this competition. I simply tried tinkering around with models &amp; ideas. And when I found that I had secured my first silver at my first ever serious kaggle competition, my joy knew no bounds! I had a public LB rank of 1480, the last time I checked, and I woke up to find a rank of 65. </p>\n<p>Here's a link to my quick inference Notebook: <a href=\"https://www.kaggle.com/doctorkael/osic-inference?scriptVersionId=44167954\" target=\"_blank\">https://www.kaggle.com/doctorkael/osic-inference?scriptVersionId=44167954</a></p>\n<p>My winning solution can be summarized as follows:</p>\n<p><strong>1. Perform data augmentation:</strong> Create synthetic week FVC from other patients data who have similar <em>cosine similarity</em>. For this purpose, we calculate the various features to ensure that augmented data is as similar as possible to the actual patients data. I also tried another type of augmentation with simple interpolation methods such as akima, linear, slinear, etc and used cross validation to see which performed the best at run time and use it to create models which would make the test predictions.  </p>\n<p><strong>2. Some more augmentations</strong> (sort of): We treat each week's FVC and Percent as <em>Baseline</em> values from which  we try to predict data for other weeks. This idea was borrowed from <a href=\"https://www.kaggle.com/yasufuminakama/osic-lgb-baseline\" target=\"_blank\">here</a>. It helps creating models that are more robust and helps when we have to predict for weeks -12 and the like when all the model would have seen otherwise is week -5. Using <code>Week_Offset</code> as feature instead of <code>Weeks</code> really helped here.</p>\n<p><strong>3. The model:</strong> The secret sauce was to use a simple Gradient boosting regressor to predict the confidence intervals. Methods on how to accomplish this can be found <a href=\"https://scikit-learn.org/stable/auto_examples/ensemble/plot_gradient_boosting_quantile.html#sphx-glr-auto-examples-ensemble-plot-gradient-boosting-quantile-py\" target=\"_blank\">here</a>. Choosing an <code>alpha</code> value of 0.75 worked best. We create models on all the different augmented data samples, cross validate them to see which perform best with least error in patient scores (since we only have 3 final predictions which would actually count in scoring, we need to minimize this).</p>\n<p><strong>4. Making Predictions:</strong> The maximum (for upper bound) and minimum (for lower bound) of all the predictions made for that particular patient and week is calculated, subtracted to obtain the the <code>Confidence</code> values.</p>\n<p>I would soon publish my notebook containing all the ideas I tinkered around with soon as it is still in an unedited format. </p>",
      "rawMarkdown": "My hearty congratulations to all the winners of this competition. Quite honestly I have to admit that I gave up any hopes of winning this competition. I simply tried tinkering around with models & ideas. And when I found that I had secured my first silver at my first ever serious kaggle competition, my joy knew no bounds! I had a public LB rank of 1480, the last time I checked, and I woke up to find a rank of 65. \n\nHere's a link to my quick inference Notebook: https://www.kaggle.com/doctorkael/osic-inference?scriptVersionId=44167954\n\nMy winning solution can be summarized as follows:\n\n**1. Perform data augmentation:** Create synthetic week FVC from other patients data who have similar *cosine similarity*. For this purpose, we calculate the various features to ensure that augmented data is as similar as possible to the actual patients data. I also tried another type of augmentation with simple interpolation methods such as akima, linear, slinear, etc and used cross validation to see which performed the best at run time and use it to create models which would make the test predictions.  \n\n**2. Some more augmentations** (sort of): We treat each week's FVC and Percent as *Baseline* values from which  we try to predict data for other weeks. This idea was borrowed from [here](https://www.kaggle.com/yasufuminakama/osic-lgb-baseline). It helps creating models that are more robust and helps when we have to predict for weeks -12 and the like when all the model would have seen otherwise is week -5. Using `Week_Offset` as feature instead of `Weeks` really helped here.\n\n**3. The model:** The secret sauce was to use a simple Gradient boosting regressor to predict the confidence intervals. Methods on how to accomplish this can be found [here](https://scikit-learn.org/stable/auto_examples/ensemble/plot_gradient_boosting_quantile.html#sphx-glr-auto-examples-ensemble-plot-gradient-boosting-quantile-py). Choosing an `alpha` value of 0.75 worked best. We create models on all the different augmented data samples, cross validate them to see which perform best with least error in patient scores (since we only have 3 final predictions which would actually count in scoring, we need to minimize this).\n\n**4. Making Predictions:** The maximum (for upper bound) and minimum (for lower bound) of all the predictions made for that particular patient and week is calculated, subtracted to obtain the the `Confidence` values.\n\nI would soon publish my notebook containing all the ideas I tinkered around with soon as it is still in an unedited format.",
      "votes": null
    },
    {
      "id": "1044749",
      "postDate": "10/10/2020 04:57:08",
      "content": "<p>I have documented my experiments in this Notebook <a href=\"https://www.kaggle.com/doctorkael/osic-incrementally-improved-models\" target=\"_blank\">here</a>. Do check it out if you are interested :)</p>",
      "rawMarkdown": "I have documented my experiments in this Notebook [here](https://www.kaggle.com/doctorkael/osic-incrementally-improved-models). Do check it out if you are interested :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1044749,
      "author_name": "doctorkael",
      "author_url": "",
      "post_date": "10/10/2020 04:57:08",
      "content": "<p>I have documented my experiments in this Notebook <a href=\"https://www.kaggle.com/doctorkael/osic-incrementally-improved-models\" target=\"_blank\">here</a>. Do check it out if you are interested :)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1040855": "My hearty congratulations to all the winners of this competition. Quite honestly I have to admit that I gave up any hopes of winning this competition. I simply tried tinkering around with models & ideas. And when I found that I had secured my first silver at my first ever serious kaggle competition, my joy knew no bounds! I had a public LB rank of 1480, the last time I checked, and I woke up to find a rank of 65. \n\nHere's a link to my quick inference Notebook: https://www.kaggle.com/doctorkael/osic-inference?scriptVersionId=44167954\n\nMy winning solution can be summarized as follows:\n\n**1. Perform data augmentation:** Create synthetic week FVC from other patients data who have similar *cosine similarity*. For this purpose, we calculate the various features to ensure that augmented data is as similar as possible to the actual patients data. I also tried another type of augmentation with simple interpolation methods such as akima, linear, slinear, etc and used cross validation to see which performed the best at run time and use it to create models which would make the test predictions.  \n\n**2. Some more augmentations** (sort of): We treat each week's FVC and Percent as *Baseline* values from which  we try to predict data for other weeks. This idea was borrowed from [here](https://www.kaggle.com/yasufuminakama/osic-lgb-baseline). It helps creating models that are more robust and helps when we have to predict for weeks -12 and the like when all the model would have seen otherwise is week -5. Using `Week_Offset` as feature instead of `Weeks` really helped here.\n\n**3. The model:** The secret sauce was to use a simple Gradient boosting regressor to predict the confidence intervals. Methods on how to accomplish this can be found [here](https://scikit-learn.org/stable/auto_examples/ensemble/plot_gradient_boosting_quantile.html#sphx-glr-auto-examples-ensemble-plot-gradient-boosting-quantile-py). Choosing an `alpha` value of 0.75 worked best. We create models on all the different augmented data samples, cross validate them to see which perform best with least error in patient scores (since we only have 3 final predictions which would actually count in scoring, we need to minimize this).\n\n**4. Making Predictions:** The maximum (for upper bound) and minimum (for lower bound) of all the predictions made for that particular patient and week is calculated, subtracted to obtain the the `Confidence` values.\n\nI would soon publish my notebook containing all the ideas I tinkered around with soon as it is still in an unedited format.",
    "1044749": "I have documented my experiments in this Notebook [here](https://www.kaggle.com/doctorkael/osic-incrementally-improved-models). Do check it out if you are interested :)"
  },
  "source": "meta"
}