{
  "id": 189496,
  "title": "23rd Place Solution - Single Tabnet model",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/writeups/gauthamkumaran-23rd-place-solution-single-tabnet-m",
  "author_name": "",
  "post_date": "2020-10-07T19:12:36.633452700Z",
  "votes": 14,
  "comment_count": 9,
  "views": 0,
  "content": "<p>This was my first Kaggle silver medal, super happy about that 😄 This was a great learning experience for me. Kudos to all the winners and thanks to the entire Kaggle ecosystem.</p>\n<h3><strong>Solution Overview</strong></h3>\n<p>Tabular features and features obtained from CT scans were used in a <a href=\"https://github.com/dreamquark-ai/tabnet\" target=\"_blank\">Tabnet model</a>, trained with pinball loss. As a form of data augmentation, every FVC score available for each person was assumed to be the first FVC test score, and other features were built accordingly.</p>\n<p><strong>Training Notebook</strong> - <a href=\"https://www.kaggle.com/gautham11/quantile-regression-pytorch-lightning-training?scriptVersionId=44239361\" target=\"_blank\">https://www.kaggle.com/gautham11/quantile-regression-pytorch-lightning-training?scriptVersionId=44239361</a></p>\n<p><strong>Inference Notebook</strong> - <a href=\"https://www.kaggle.com/gautham11/quantile-regression-pytorch-lightning-inference/data?scriptVersionId=43574381\" target=\"_blank\">https://www.kaggle.com/gautham11/quantile-regression-pytorch-lightning-inference/data?scriptVersionId=43574381</a></p>\n<h3><strong>Tabular Features</strong></h3>\n<p>This is fairly straightforward.</p>\n<p><code>first_test_fvc</code> - FVC score of the assumed first visit<br>\n<code>predict_week</code> - Weeks from the CT scan, <code>Week</code> feature from the train data<br>\n<code>weeks_from_first_visit</code> - number of weeks from the first FVC test.<br>\n<code>expected_fvc</code> - <code>fvc * percent</code> this is a constant for each patient.<br>\n<code>percent</code> - Percent of the <code>first_test_fvc</code> when compared to the <code>expected_fvc</code><br>\n<code>age</code> - Age of the patient</p>\n<p>One-hot encoded values of <code>Sex</code> and <code>SmokingStatus</code></p>\n<h3><strong>CT Scan Features</strong></h3>\n<p>3D rescaled segmented lung model was generated for each patient from which the image features were obtained.</p>\n<p><strong>Segmentation</strong> - The lung segmentation was performed using <a href=\"https://www.kaggle.com/aadhavvignesh/lung-segmentation-by-marker-controlled-watershed\" target=\"_blank\">Marker-controlled watershed segmentation</a>, thanks to the amazing work by <a href=\"https://www.kaggle.com/aadhavvignesh\" target=\"_blank\">Aadhav Vignesh</a></p>\n<p><strong>Rescaling the voxel pixels</strong> - The rescaling was done using the <code>resample</code> method from <a href=\"https://www.kaggle.com/allunia/pulmonary-dicom-preprocessing\" target=\"_blank\">Pulmonary Dicom Preprocessing</a>. Thanks <a href=\"https://www.kaggle.com/allunia\" target=\"_blank\">Laura Fink</a> for the great kernel.</p>\n<p>The only issue I faced with this approach was to stay within the 4 hours submission runtime while generating the 3D models for private dataset in the inference notebook. I was able to overcome this by switching <code>scipy.ndimage.zoom</code> in <code>resample</code> function with <code>torch.nn.functional.interpolate</code>, which was much faster.</p>\n<p><strong>Image features notebook</strong> - <a href=\"https://www.kaggle.com/gautham11/lung-volume-lung-height-hu-values-image-features\" target=\"_blank\">https://www.kaggle.com/gautham11/lung-volume-lung-height-hu-values-image-features</a></p>\n<p><strong>train Image features dataset</strong> - <a href=\"https://www.kaggle.com/gautham11/lung-image-features\" target=\"_blank\">https://www.kaggle.com/gautham11/lung-image-features</a></p>\n<p>Image features generated are</p>\n<p><code>lung_volume_in_liters</code> - Number of segmented lung voxel pixels in 3d model / 1e6.<br>\n<code>lung_height</code> - Difference between the first and last slice with more that 1000 lung voxel pixels.<br>\n<code>lung_mean_hu</code> - Mean of the HU values.<br>\n<code>lung_skew</code> - skew of the HU values.<br>\n<code>lung_kurtosis</code> - Kurtosis of the HU values.</p>\n<p>And histogram bin values of the HU values in the 3D model were used as features - <code>bin_-900</code>, <code>bin-800</code>…<code>bin_200</code>.</p>\n<p>The inspiration for most of these features was from the <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165727\" target=\"_blank\">Domain expert's insights</a> post by <a href=\"https://www.kaggle.com/sandorkonya\" target=\"_blank\">Dr. Konya</a></p>\n<p>These features gave me a good improvement in my CV score.</p>\n<h3><strong>Model Training</strong></h3>\n<p>The data was min-max scaled. The data was split into 5 folds, stratified on the <code>expected_fvc</code>, intuition being that split could be based on the Patients' characteristics, not sure if this was the best method. For each split, the Tabnet model was trained with <code>SGD</code> optimizer and One-cycle LR scheduler for 350 epochs.</p>\n<p><strong>Model weights dataset</strong> - <a href=\"https://www.kaggle.com/gautham11/gautham-quantmodel\" target=\"_blank\">https://www.kaggle.com/gautham11/gautham-quantmodel</a></p>\n<p>I gave Pytorch-Lightning a try for the first time in this competition, and I'm fairly certain that I'll use it in every project from here on.</p>\n<p>Finally, on selecting my submission, I selected my best CV score model and my best Public LB model. In retrospect, selecting the best Public LB model was a poor decision because it got a bad score in private LB, all the more affirmation to always <strong>trust the CV</strong></p>",
  "messages": [
    {
      "id": "1041491",
      "postDate": "10/07/2020 19:12:36",
      "content": "<p>This was my first Kaggle silver medal, super happy about that 😄 This was a great learning experience for me. Kudos to all the winners and thanks to the entire Kaggle ecosystem.</p>\n<h3><strong>Solution Overview</strong></h3>\n<p>Tabular features and features obtained from CT scans were used in a <a href=\"https://github.com/dreamquark-ai/tabnet\" target=\"_blank\">Tabnet model</a>, trained with pinball loss. As a form of data augmentation, every FVC score available for each person was assumed to be the first FVC test score, and other features were built accordingly.</p>\n<p><strong>Training Notebook</strong> - <a href=\"https://www.kaggle.com/gautham11/quantile-regression-pytorch-lightning-training?scriptVersionId=44239361\" target=\"_blank\">https://www.kaggle.com/gautham11/quantile-regression-pytorch-lightning-training?scriptVersionId=44239361</a></p>\n<p><strong>Inference Notebook</strong> - <a href=\"https://www.kaggle.com/gautham11/quantile-regression-pytorch-lightning-inference/data?scriptVersionId=43574381\" target=\"_blank\">https://www.kaggle.com/gautham11/quantile-regression-pytorch-lightning-inference/data?scriptVersionId=43574381</a></p>\n<h3><strong>Tabular Features</strong></h3>\n<p>This is fairly straightforward.</p>\n<p><code>first_test_fvc</code> - FVC score of the assumed first visit<br>\n<code>predict_week</code> - Weeks from the CT scan, <code>Week</code> feature from the train data<br>\n<code>weeks_from_first_visit</code> - number of weeks from the first FVC test.<br>\n<code>expected_fvc</code> - <code>fvc * percent</code> this is a constant for each patient.<br>\n<code>percent</code> - Percent of the <code>first_test_fvc</code> when compared to the <code>expected_fvc</code><br>\n<code>age</code> - Age of the patient</p>\n<p>One-hot encoded values of <code>Sex</code> and <code>SmokingStatus</code></p>\n<h3><strong>CT Scan Features</strong></h3>\n<p>3D rescaled segmented lung model was generated for each patient from which the image features were obtained.</p>\n<p><strong>Segmentation</strong> - The lung segmentation was performed using <a href=\"https://www.kaggle.com/aadhavvignesh/lung-segmentation-by-marker-controlled-watershed\" target=\"_blank\">Marker-controlled watershed segmentation</a>, thanks to the amazing work by <a href=\"https://www.kaggle.com/aadhavvignesh\" target=\"_blank\">Aadhav Vignesh</a></p>\n<p><strong>Rescaling the voxel pixels</strong> - The rescaling was done using the <code>resample</code> method from <a href=\"https://www.kaggle.com/allunia/pulmonary-dicom-preprocessing\" target=\"_blank\">Pulmonary Dicom Preprocessing</a>. Thanks <a href=\"https://www.kaggle.com/allunia\" target=\"_blank\">Laura Fink</a> for the great kernel.</p>\n<p>The only issue I faced with this approach was to stay within the 4 hours submission runtime while generating the 3D models for private dataset in the inference notebook. I was able to overcome this by switching <code>scipy.ndimage.zoom</code> in <code>resample</code> function with <code>torch.nn.functional.interpolate</code>, which was much faster.</p>\n<p><strong>Image features notebook</strong> - <a href=\"https://www.kaggle.com/gautham11/lung-volume-lung-height-hu-values-image-features\" target=\"_blank\">https://www.kaggle.com/gautham11/lung-volume-lung-height-hu-values-image-features</a></p>\n<p><strong>train Image features dataset</strong> - <a href=\"https://www.kaggle.com/gautham11/lung-image-features\" target=\"_blank\">https://www.kaggle.com/gautham11/lung-image-features</a></p>\n<p>Image features generated are</p>\n<p><code>lung_volume_in_liters</code> - Number of segmented lung voxel pixels in 3d model / 1e6.<br>\n<code>lung_height</code> - Difference between the first and last slice with more that 1000 lung voxel pixels.<br>\n<code>lung_mean_hu</code> - Mean of the HU values.<br>\n<code>lung_skew</code> - skew of the HU values.<br>\n<code>lung_kurtosis</code> - Kurtosis of the HU values.</p>\n<p>And histogram bin values of the HU values in the 3D model were used as features - <code>bin_-900</code>, <code>bin-800</code>…<code>bin_200</code>.</p>\n<p>The inspiration for most of these features was from the <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165727\" target=\"_blank\">Domain expert's insights</a> post by <a href=\"https://www.kaggle.com/sandorkonya\" target=\"_blank\">Dr. Konya</a></p>\n<p>These features gave me a good improvement in my CV score.</p>\n<h3><strong>Model Training</strong></h3>\n<p>The data was min-max scaled. The data was split into 5 folds, stratified on the <code>expected_fvc</code>, intuition being that split could be based on the Patients' characteristics, not sure if this was the best method. For each split, the Tabnet model was trained with <code>SGD</code> optimizer and One-cycle LR scheduler for 350 epochs.</p>\n<p><strong>Model weights dataset</strong> - <a href=\"https://www.kaggle.com/gautham11/gautham-quantmodel\" target=\"_blank\">https://www.kaggle.com/gautham11/gautham-quantmodel</a></p>\n<p>I gave Pytorch-Lightning a try for the first time in this competition, and I'm fairly certain that I'll use it in every project from here on.</p>\n<p>Finally, on selecting my submission, I selected my best CV score model and my best Public LB model. In retrospect, selecting the best Public LB model was a poor decision because it got a bad score in private LB, all the more affirmation to always <strong>trust the CV</strong></p>",
      "rawMarkdown": "This was my first Kaggle silver medal, super happy about that 😄 This was a great learning experience for me. Kudos to all the winners and thanks to the entire Kaggle ecosystem.\n\n### **Solution Overview**\n\nTabular features and features obtained from CT scans were used in a [Tabnet model](https://github.com/dreamquark-ai/tabnet), trained with pinball loss. As a form of data augmentation, every FVC score available for each person was assumed to be the first FVC test score, and other features were built accordingly.\n\n**Training Notebook** - https://www.kaggle.com/gautham11/quantile-regression-pytorch-lightning-training?scriptVersionId=44239361\n\n**Inference Notebook** - https://www.kaggle.com/gautham11/quantile-regression-pytorch-lightning-inference/data?scriptVersionId=43574381\n\n### **Tabular Features**\n\nThis is fairly straightforward.\n\n`first_test_fvc` - FVC score of the assumed first visit\n`predict_week` - Weeks from the CT scan, `Week` feature from the train data\n`weeks_from_first_visit` - number of weeks from the first FVC test.\n`expected_fvc` - `fvc * percent` this is a constant for each patient.\n`percent` - Percent of the `first_test_fvc` when compared to the `expected_fvc`\n`age` - Age of the patient\n\nOne-hot encoded values of `Sex` and `SmokingStatus`\n\n### **CT Scan Features**\n\n3D rescaled segmented lung model was generated for each patient from which the image features were obtained.\n\n**Segmentation** - The lung segmentation was performed using [Marker-controlled watershed segmentation](https://www.kaggle.com/aadhavvignesh/lung-segmentation-by-marker-controlled-watershed), thanks to the amazing work by [Aadhav Vignesh](https://www.kaggle.com/aadhavvignesh)\n\n**Rescaling the voxel pixels** - The rescaling was done using the `resample` method from [Pulmonary Dicom Preprocessing](https://www.kaggle.com/allunia/pulmonary-dicom-preprocessing). Thanks [Laura Fink](https://www.kaggle.com/allunia) for the great kernel.\n\nThe only issue I faced with this approach was to stay within the 4 hours submission runtime while generating the 3D models for private dataset in the inference notebook. I was able to overcome this by switching `scipy.ndimage.zoom` in `resample` function with `torch.nn.functional.interpolate`, which was much faster.\n\n**Image features notebook** - https://www.kaggle.com/gautham11/lung-volume-lung-height-hu-values-image-features\n\n**train Image features dataset** - https://www.kaggle.com/gautham11/lung-image-features\n\nImage features generated are\n\n`lung_volume_in_liters` - Number of segmented lung voxel pixels in 3d model / 1e6.\n`lung_height` - Difference between the first and last slice with more that 1000 lung voxel pixels.\n`lung_mean_hu` - Mean of the HU values.\n`lung_skew` - skew of the HU values.\n`lung_kurtosis` - Kurtosis of the HU values.\n\nAnd histogram bin values of the HU values in the 3D model were used as features - `bin_-900`, `bin-800`...`bin_200`.\n\nThe inspiration for most of these features was from the [Domain expert's insights](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165727) post by [Dr. Konya](https://www.kaggle.com/sandorkonya)\n\nThese features gave me a good improvement in my CV score.\n\n### **Model Training**\n\nThe data was min-max scaled. The data was split into 5 folds, stratified on the `expected_fvc`, intuition being that split could be based on the Patients' characteristics, not sure if this was the best method. For each split, the Tabnet model was trained with `SGD` optimizer and One-cycle LR scheduler for 350 epochs.\n\n**Model weights dataset** - https://www.kaggle.com/gautham11/gautham-quantmodel\n\nI gave Pytorch-Lightning a try for the first time in this competition, and I'm fairly certain that I'll use it in every project from here on.\n\nFinally, on selecting my submission, I selected my best CV score model and my best Public LB model. In retrospect, selecting the best Public LB model was a poor decision because it got a bad score in private LB, all the more affirmation to always **trust the CV**",
      "votes": null
    },
    {
      "id": "1042040",
      "postDate": "10/08/2020 03:10:04",
      "content": "<p><a href=\"https://www.kaggle.com/gautham11\" target=\"_blank\">@gautham11</a> Congratulations on the silver! +1 for PyTorch-Lightning :P</p>",
      "rawMarkdown": "gautham11 Congratulations on the silver! +1 for PyTorch-Lightning :P",
      "votes": null
    },
    {
      "id": "1042194",
      "postDate": "10/08/2020 06:01:54",
      "content": "<p>Thanks, <a href=\"https://www.kaggle.com/aadhavvignesh\" target=\"_blank\">@aadhavvignesh</a>! Good to see a fellow pytorch-lightning fan :P</p>",
      "rawMarkdown": "Thanks, @aadhavvignesh! Good to see a fellow pytorch-lightning fan :P",
      "votes": null
    },
    {
      "id": "1042333",
      "postDate": "10/08/2020 07:18:22",
      "content": "<p>Congratulations ! <br>\nMy single tabnet model with PyTorch got top 8% only with tabular features. I do believe that tabnet is a robust solution for this challenge. Does pytorch lighthing provide a hyperparameter tuning for tabnet ? </p>",
      "rawMarkdown": "Congratulations ! \nMy single tabnet model with PyTorch got top 8% only with tabular features. I do believe that tabnet is a robust solution for this challenge. Does pytorch lighthing provide a hyperparameter tuning for tabnet ?",
      "votes": null
    },
    {
      "id": "1042461",
      "postDate": "10/08/2020 08:40:16",
      "content": "<p>Thanks! My tabnet model with only tabular features also scored in the same range.<br>\nGreat question! It seems pytorch-lightning supports <a href=\"https://pytorch-lightning.readthedocs.io/en/latest/hyperparameters.html#hyperparameter-optimization\" target=\"_blank\">hyperparameter tuning by integrating with other libraries</a>. I didn't try it in this competition, but we should be able to tune tabnet parameters this way.</p>",
      "rawMarkdown": "Thanks! My tabnet model with only tabular features also scored in the same range.\nGreat question! It seems pytorch-lightning supports [hyperparameter tuning by integrating with other libraries](https://pytorch-lightning.readthedocs.io/en/latest/hyperparameters.html#hyperparameter-optimization). I didn't try it in this competition, but we should be able to tune tabnet parameters this way.",
      "votes": null
    },
    {
      "id": "1042697",
      "postDate": "10/08/2020 12:04:41",
      "content": "<p><a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a> Told you Tabnet will do wonders<br>\n<a href=\"https://www.kaggle.com/gautham11\" target=\"_blank\">@gautham11</a> Thanks for the post , also congratulations on the medal </p>",
      "rawMarkdown": "optimo Told you Tabnet will do wonders\n@gautham11 Thanks for the post , also congratulations on the medal",
      "votes": null
    },
    {
      "id": "1042704",
      "postDate": "10/08/2020 12:15:42",
      "content": "<p><a href=\"https://www.kaggle.com/gautham11\" target=\"_blank\">@gautham11</a> congratulations! will you make your solution public? I'd like to see how easy it was for you to use the library (I don't want to spoil, but some nice improvements are coming to make pytorch-tabnet easier to customize and so even more competitive! :))</p>\n<p><a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a> thanks for the mention, more wonders to come hopefully!</p>",
      "rawMarkdown": "gautham11 congratulations! will you make your solution public? I'd like to see how easy it was for you to use the library (I don't want to spoil, but some nice improvements are coming to make pytorch-tabnet easier to customize and so even more competitive! :))\n\n@tanulsingh077 thanks for the mention, more wonders to come hopefully!",
      "votes": null
    },
    {
      "id": "1042974",
      "postDate": "10/08/2020 15:30:45",
      "content": "<p><a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a> Very thankful for the work you've done, can't wait for the updates. <br>\nThe training code is available in this notebook - <a href=\"https://www.kaggle.com/gautham11/quantile-regression-pytorch-lightning-training?scriptVersionId=44239361\" target=\"_blank\">https://www.kaggle.com/gautham11/quantile-regression-pytorch-lightning-training?scriptVersionId=44239361</a><br>\nI didn't need the sklearn wrapper, so used <code>TabNet</code> directly.</p>",
      "rawMarkdown": "optimo Very thankful for the work you've done, can't wait for the updates. \nThe training code is available in this notebook - https://www.kaggle.com/gautham11/quantile-regression-pytorch-lightning-training?scriptVersionId=44239361\nI didn't need the sklearn wrapper, so used `TabNet` directly.",
      "votes": null
    },
    {
      "id": "1042975",
      "postDate": "10/08/2020 15:31:47",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a> :)</p>",
      "rawMarkdown": "Thanks @tanulsingh077 :)",
      "votes": null
    },
    {
      "id": "1043550",
      "postDate": "10/09/2020 04:14:14",
      "content": "<p>Congratulations on the medal. Thank you sharing your experiance</p>",
      "rawMarkdown": "Congratulations on the medal. Thank you sharing your experiance",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1042040,
      "author_name": "aadhavvignesh",
      "author_url": "",
      "post_date": "10/08/2020 03:10:04",
      "content": "<p><a href=\"https://www.kaggle.com/gautham11\" target=\"_blank\">@gautham11</a> Congratulations on the silver! +1 for PyTorch-Lightning :P</p>",
      "votes": null,
      "replies": [
        {
          "id": 1042194,
          "author_name": "gautham11",
          "author_url": "",
          "post_date": "10/08/2020 06:01:54",
          "content": "<p>Thanks, <a href=\"https://www.kaggle.com/aadhavvignesh\" target=\"_blank\">@aadhavvignesh</a>! Good to see a fellow pytorch-lightning fan :P</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1042333,
      "author_name": "alexj21",
      "author_url": "",
      "post_date": "10/08/2020 07:18:22",
      "content": "<p>Congratulations ! <br>\nMy single tabnet model with PyTorch got top 8% only with tabular features. I do believe that tabnet is a robust solution for this challenge. Does pytorch lighthing provide a hyperparameter tuning for tabnet ? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1042461,
          "author_name": "gautham11",
          "author_url": "",
          "post_date": "10/08/2020 08:40:16",
          "content": "<p>Thanks! My tabnet model with only tabular features also scored in the same range.<br>\nGreat question! It seems pytorch-lightning supports <a href=\"https://pytorch-lightning.readthedocs.io/en/latest/hyperparameters.html#hyperparameter-optimization\" target=\"_blank\">hyperparameter tuning by integrating with other libraries</a>. I didn't try it in this competition, but we should be able to tune tabnet parameters this way.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1042697,
      "author_name": "tanulsingh077",
      "author_url": "",
      "post_date": "10/08/2020 12:04:41",
      "content": "<p><a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a> Told you Tabnet will do wonders<br>\n<a href=\"https://www.kaggle.com/gautham11\" target=\"_blank\">@gautham11</a> Thanks for the post , also congratulations on the medal </p>",
      "votes": null,
      "replies": [
        {
          "id": 1042704,
          "author_name": "optimo",
          "author_url": "",
          "post_date": "10/08/2020 12:15:42",
          "content": "<p><a href=\"https://www.kaggle.com/gautham11\" target=\"_blank\">@gautham11</a> congratulations! will you make your solution public? I'd like to see how easy it was for you to use the library (I don't want to spoil, but some nice improvements are coming to make pytorch-tabnet easier to customize and so even more competitive! :))</p>\n<p><a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a> thanks for the mention, more wonders to come hopefully!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1042974,
          "author_name": "gautham11",
          "author_url": "",
          "post_date": "10/08/2020 15:30:45",
          "content": "<p><a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a> Very thankful for the work you've done, can't wait for the updates. <br>\nThe training code is available in this notebook - <a href=\"https://www.kaggle.com/gautham11/quantile-regression-pytorch-lightning-training?scriptVersionId=44239361\" target=\"_blank\">https://www.kaggle.com/gautham11/quantile-regression-pytorch-lightning-training?scriptVersionId=44239361</a><br>\nI didn't need the sklearn wrapper, so used <code>TabNet</code> directly.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1042975,
          "author_name": "gautham11",
          "author_url": "",
          "post_date": "10/08/2020 15:31:47",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a> :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1043550,
      "author_name": "",
      "author_url": "",
      "post_date": "10/09/2020 04:14:14",
      "content": "<p>Congratulations on the medal. Thank you sharing your experiance</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1041491": "This was my first Kaggle silver medal, super happy about that 😄 This was a great learning experience for me. Kudos to all the winners and thanks to the entire Kaggle ecosystem.\n\n### **Solution Overview**\n\nTabular features and features obtained from CT scans were used in a [Tabnet model](https://github.com/dreamquark-ai/tabnet), trained with pinball loss. As a form of data augmentation, every FVC score available for each person was assumed to be the first FVC test score, and other features were built accordingly.\n\n**Training Notebook** - https://www.kaggle.com/gautham11/quantile-regression-pytorch-lightning-training?scriptVersionId=44239361\n\n**Inference Notebook** - https://www.kaggle.com/gautham11/quantile-regression-pytorch-lightning-inference/data?scriptVersionId=43574381\n\n### **Tabular Features**\n\nThis is fairly straightforward.\n\n`first_test_fvc` - FVC score of the assumed first visit\n`predict_week` - Weeks from the CT scan, `Week` feature from the train data\n`weeks_from_first_visit` - number of weeks from the first FVC test.\n`expected_fvc` - `fvc * percent` this is a constant for each patient.\n`percent` - Percent of the `first_test_fvc` when compared to the `expected_fvc`\n`age` - Age of the patient\n\nOne-hot encoded values of `Sex` and `SmokingStatus`\n\n### **CT Scan Features**\n\n3D rescaled segmented lung model was generated for each patient from which the image features were obtained.\n\n**Segmentation** - The lung segmentation was performed using [Marker-controlled watershed segmentation](https://www.kaggle.com/aadhavvignesh/lung-segmentation-by-marker-controlled-watershed), thanks to the amazing work by [Aadhav Vignesh](https://www.kaggle.com/aadhavvignesh)\n\n**Rescaling the voxel pixels** - The rescaling was done using the `resample` method from [Pulmonary Dicom Preprocessing](https://www.kaggle.com/allunia/pulmonary-dicom-preprocessing). Thanks [Laura Fink](https://www.kaggle.com/allunia) for the great kernel.\n\nThe only issue I faced with this approach was to stay within the 4 hours submission runtime while generating the 3D models for private dataset in the inference notebook. I was able to overcome this by switching `scipy.ndimage.zoom` in `resample` function with `torch.nn.functional.interpolate`, which was much faster.\n\n**Image features notebook** - https://www.kaggle.com/gautham11/lung-volume-lung-height-hu-values-image-features\n\n**train Image features dataset** - https://www.kaggle.com/gautham11/lung-image-features\n\nImage features generated are\n\n`lung_volume_in_liters` - Number of segmented lung voxel pixels in 3d model / 1e6.\n`lung_height` - Difference between the first and last slice with more that 1000 lung voxel pixels.\n`lung_mean_hu` - Mean of the HU values.\n`lung_skew` - skew of the HU values.\n`lung_kurtosis` - Kurtosis of the HU values.\n\nAnd histogram bin values of the HU values in the 3D model were used as features - `bin_-900`, `bin-800`...`bin_200`.\n\nThe inspiration for most of these features was from the [Domain expert's insights](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165727) post by [Dr. Konya](https://www.kaggle.com/sandorkonya)\n\nThese features gave me a good improvement in my CV score.\n\n### **Model Training**\n\nThe data was min-max scaled. The data was split into 5 folds, stratified on the `expected_fvc`, intuition being that split could be based on the Patients' characteristics, not sure if this was the best method. For each split, the Tabnet model was trained with `SGD` optimizer and One-cycle LR scheduler for 350 epochs.\n\n**Model weights dataset** - https://www.kaggle.com/gautham11/gautham-quantmodel\n\nI gave Pytorch-Lightning a try for the first time in this competition, and I'm fairly certain that I'll use it in every project from here on.\n\nFinally, on selecting my submission, I selected my best CV score model and my best Public LB model. In retrospect, selecting the best Public LB model was a poor decision because it got a bad score in private LB, all the more affirmation to always **trust the CV**",
    "1042040": "gautham11 Congratulations on the silver! +1 for PyTorch-Lightning :P",
    "1042194": "Thanks, @aadhavvignesh! Good to see a fellow pytorch-lightning fan :P",
    "1042333": "Congratulations ! \nMy single tabnet model with PyTorch got top 8% only with tabular features. I do believe that tabnet is a robust solution for this challenge. Does pytorch lighthing provide a hyperparameter tuning for tabnet ?",
    "1042461": "Thanks! My tabnet model with only tabular features also scored in the same range.\nGreat question! It seems pytorch-lightning supports [hyperparameter tuning by integrating with other libraries](https://pytorch-lightning.readthedocs.io/en/latest/hyperparameters.html#hyperparameter-optimization). I didn't try it in this competition, but we should be able to tune tabnet parameters this way.",
    "1042697": "optimo Told you Tabnet will do wonders\n@gautham11 Thanks for the post , also congratulations on the medal",
    "1042704": "gautham11 congratulations! will you make your solution public? I'd like to see how easy it was for you to use the library (I don't want to spoil, but some nice improvements are coming to make pytorch-tabnet easier to customize and so even more competitive! :))\n\n@tanulsingh077 thanks for the mention, more wonders to come hopefully!",
    "1042974": "optimo Very thankful for the work you've done, can't wait for the updates. \nThe training code is available in this notebook - https://www.kaggle.com/gautham11/quantile-regression-pytorch-lightning-training?scriptVersionId=44239361\nI didn't need the sklearn wrapper, so used `TabNet` directly.",
    "1042975": "Thanks @tanulsingh077 :)",
    "1043550": "Congratulations on the medal. Thank you sharing your experiance"
  },
  "source": "meta"
}