{
  "id": 189217,
  "title": "10th Place Solution",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/writeups/amed-10th-place-solution",
  "author_name": "",
  "post_date": "2022-08-26T21:41:02.427Z",
  "votes": 42,
  "comment_count": 16,
  "views": 0,
  "content": "<p>Congratz to the winners !</p>\n<p>Many thanks to kaggle team for hosting this competition,hope the winning solutions will bring useful insight to the problem.</p>\n<p>My solution is based only on tabular data.</p>\n<h1>Data Augmentation</h1>\n<p>The strategy has been shared in many public kernels.<br>\nYou can find it <a href=\"https://www.kaggle.com/ttahara/osic-baseline-lgbm-with-custom-metric\" target=\"_blank\">here</a>.</p>\n<h1>Validation</h1>\n<p>I already shared my validation strategy <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/168610\" target=\"_blank\">here </a>during the competition.</p>\n<h1>Features</h1>\n<p>I use the following features : <code>'base_FVC', 'base_Percent', 'base_Age', 'Week_passed', 'Sex', 'SmokingStatus'</code></p>\n<h1>FVC Prediction</h1>\n<p>Final fvc prediction is  a mean of the following regressor : </p>\n<ul>\n<li>SVM</li>\n<li>KNN</li>\n<li>NN</li>\n<li>Quantile Regressor (0.5)</li>\n<li>RF</li>\n<li>LM</li>\n<li>HuberRegressor</li>\n<li>ElasticNet</li>\n<li>Lgbm</li>\n</ul>\n<h1>Confidence Prediction</h1>\n<p>1 - Compute the optimal Confidence value for final fvc,which is actually:<br>\n<code>Confidence = np.sqrt(2)*np.abs(FVC_true - FVC_pred)</code><br>\n2 - Use the same models in FVC  prediction to train on the Confidence this time<br>\n3 - Train a binary classifier to know if <code>Confidence&lt;=100</code><br>\n4 - Use prediction from (2) and post-process them by (3)</p>\n<h1>Things that didn't work for me</h1>\n<ul>\n<li>Everything with image data</li>\n<li>Linear Augmentation on tabular data</li>\n</ul>",
  "messages": [
    {
      "id": "1040174",
      "postDate": "10/07/2020 02:38:33",
      "content": "<p>Congratz to the winners !</p>\n<p>Many thanks to kaggle team for hosting this competition,hope the winning solutions will bring useful insight to the problem.</p>\n<p>My solution is based only on tabular data.</p>\n<h1>Data Augmentation</h1>\n<p>The strategy has been shared in many public kernels.<br>\nYou can find it <a href=\"https://www.kaggle.com/ttahara/osic-baseline-lgbm-with-custom-metric\" target=\"_blank\">here</a>.</p>\n<h1>Validation</h1>\n<p>I already shared my validation strategy <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/168610\" target=\"_blank\">here </a>during the competition.</p>\n<h1>Features</h1>\n<p>I use the following features : <code>'base_FVC', 'base_Percent', 'base_Age', 'Week_passed', 'Sex', 'SmokingStatus'</code></p>\n<h1>FVC Prediction</h1>\n<p>Final fvc prediction is  a mean of the following regressor : </p>\n<ul>\n<li>SVM</li>\n<li>KNN</li>\n<li>NN</li>\n<li>Quantile Regressor (0.5)</li>\n<li>RF</li>\n<li>LM</li>\n<li>HuberRegressor</li>\n<li>ElasticNet</li>\n<li>Lgbm</li>\n</ul>\n<h1>Confidence Prediction</h1>\n<p>1 - Compute the optimal Confidence value for final fvc,which is actually:<br>\n<code>Confidence = np.sqrt(2)*np.abs(FVC_true - FVC_pred)</code><br>\n2 - Use the same models in FVC  prediction to train on the Confidence this time<br>\n3 - Train a binary classifier to know if <code>Confidence&lt;=100</code><br>\n4 - Use prediction from (2) and post-process them by (3)</p>\n<h1>Things that didn't work for me</h1>\n<ul>\n<li>Everything with image data</li>\n<li>Linear Augmentation on tabular data</li>\n</ul>",
      "rawMarkdown": "Congratz to the winners !\n\nMany thanks to kaggle team for hosting this competition,hope the winning solutions will bring useful insight to the problem.\n\nMy solution is based only on tabular data.\n\n#  Data Augmentation\nThe strategy has been shared in many public kernels.\nYou can find it [here](https://www.kaggle.com/ttahara/osic-baseline-lgbm-with-custom-metric).\n# Validation \nI already shared my validation strategy [here ](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/168610)during the competition.\n\n# Features \nI use the following features : `'base_FVC', 'base_Percent', 'base_Age', 'Week_passed', 'Sex', 'SmokingStatus'`\n# FVC Prediction\nFinal fvc prediction is  a mean of the following regressor : \n- SVM\n- KNN\n- NN\n- Quantile Regressor (0.5)\n- RF\n- LM\n- HuberRegressor\n- ElasticNet\n- Lgbm\n\n# Confidence Prediction\n\n1 - Compute the optimal Confidence value for final fvc,which is actually:\n`Confidence = np.sqrt(2)*np.abs(FVC_true - FVC_pred)`\n2 - Use the same models in FVC  prediction to train on the Confidence this time\n3 - Train a binary classifier to know if `Confidence<=100`\n4 - Use prediction from (2) and post-process them by (3)\n\n# Things that didn't work for me\n- Everything with image data\n- Linear Augmentation on tabular data",
      "votes": null
    },
    {
      "id": "1040185",
      "postDate": "10/07/2020 02:49:08",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/amedprof\" target=\"_blank\">@amedprof</a> on the solo gold! It seems that your validation strategy helped you to get a good CV/LB correlation!</p>",
      "rawMarkdown": "Congratulations @amedprof on the solo gold! It seems that your validation strategy helped you to get a good CV/LB correlation!",
      "votes": null
    },
    {
      "id": "1040362",
      "postDate": "10/07/2020 05:15:32",
      "content": "<p>Great framework. Congratulations on your gold medal.</p>",
      "rawMarkdown": "Great framework. Congratulations on your gold medal.",
      "votes": null
    },
    {
      "id": "1040572",
      "postDate": "10/07/2020 08:23:22",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/imoore\" target=\"_blank\">@imoore</a> 😊</p>",
      "rawMarkdown": "Thanks @imoore 😊",
      "votes": null
    },
    {
      "id": "1040573",
      "postDate": "10/07/2020 08:23:47",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/aadhavvignesh\" target=\"_blank\">@aadhavvignesh</a></p>",
      "rawMarkdown": "Thanks @aadhavvignesh",
      "votes": null
    },
    {
      "id": "1042119",
      "postDate": "10/08/2020 04:31:33",
      "content": "<p>Congratulations! I'm impressed by your clean yet strong tabular only approach.</p>",
      "rawMarkdown": "Congratulations! I'm impressed by your clean yet strong tabular only approach.",
      "votes": null
    },
    {
      "id": "1042272",
      "postDate": "10/08/2020 06:41:08",
      "content": "<p>Congratulations <a href=\"/amedprof\">@amedprof</a> on the solo gold!</p>",
      "rawMarkdown": "Congratulations @amedprof on the solo gold!",
      "votes": null
    },
    {
      "id": "1042578",
      "postDate": "10/08/2020 10:24:25",
      "content": "<p>Félicitations ! You cv strategy is definitively interesting ; I still have many things to learn in that matter, I'm one of those who have been fooled by the public testset overfitting … ;)</p>\n<p>I have a couple of questions regarding your work :</p>\n<p><strong>Dataset preparation</strong></p>\n<ul>\n<li><a href=\"https://www.kaggle.com/ttahara/osic-baseline-lgbm-with-custom-metric\" target=\"_blank\">The mentioned kernel</a> suggests that you've used label encoding for <code>Sex</code> and <code>SmokingStatus</code>, right ? (not OHE ?)</li>\n<li>Did you scale <code>BaseFVC</code>, <code>BasePercent</code>, <code>Age</code>, and <code>WeekPassed</code> in some way ?</li>\n<li>Did you directly predict the <code>FVC</code> ? Or some percent of the <code>BaseFVC</code> ?</li>\n</ul>\n<p><strong>Training policy &amp; regularization</strong></p>\n<ul>\n<li>How did you manage to avoid overfitting in your models (in particular for the NN) : Did you rely on regularization (dropout, also BN as a side effect) ? What was your end-of-training policy ?</li>\n</ul>\n<p><strong>Post-processing &amp; Cross-validation</strong></p>\n<ul>\n<li>Can you elaborate on your post-processing strategy for the confidence prediction ? I don't understand the role of the final binary classifier.</li>\n<li>Your cross validation strategy is based on this great K=2 clustering you've described in <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/168610\" target=\"_blank\">this topic</a>. I assume you relied OOF validation on your 2 folds ? Or did you stick to a single train/set split, relying on several random seeds running to mitigate the randomness ? I guess the predictions from both folds had pretty close distribution …</li>\n</ul>\n<p>Congratz again, I look forward to your answer ! :)</p>",
      "rawMarkdown": "Félicitations ! You cv strategy is definitively interesting ; I still have many things to learn in that matter, I'm one of those who have been fooled by the public testset overfitting ... ;)\n\nI have a couple of questions regarding your work :\n\n\n**Dataset preparation**\n- [The mentioned kernel](https://www.kaggle.com/ttahara/osic-baseline-lgbm-with-custom-metric) suggests that you've used label encoding for `Sex` and `SmokingStatus`, right ? (not OHE ?)\n- Did you scale `BaseFVC`, `BasePercent`, `Age`, and `WeekPassed` in some way ?\n- Did you directly predict the `FVC` ? Or some percent of the `BaseFVC` ?\n\n**Training policy & regularization**\n- How did you manage to avoid overfitting in your models (in particular for the NN) : Did you rely on regularization (dropout, also BN as a side effect) ? What was your end-of-training policy ?\n\n**Post-processing & Cross-validation**\n- Can you elaborate on your post-processing strategy for the confidence prediction ? I don't understand the role of the final binary classifier.\n- Your cross validation strategy is based on this great K=2 clustering you've described in [this topic](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/168610). I assume you relied OOF validation on your 2 folds ? Or did you stick to a single train/set split, relying on several random seeds running to mitigate the randomness ? I guess the predictions from both folds had pretty close distribution ...\n\n\nCongratz again, I look forward to your answer ! :)",
      "votes": null
    },
    {
      "id": "1042592",
      "postDate": "10/08/2020 10:46:44",
      "content": "<p>You achieved the results using only the tabular features?</p>",
      "rawMarkdown": "You achieved the results using only the tabular features?",
      "votes": null
    },
    {
      "id": "1043830",
      "postDate": "10/09/2020 09:24:11",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/johannhuber\" target=\"_blank\">@johannhuber</a> .</p>\n<p><strong>Dataset preparation</strong></p>\n<ul>\n<li>I used ordinal encoder. ( eg: Male is more important than Female in Sex variable …..).</li>\n<li>No scaling</li>\n<li>Yes direct prediction of FVC.<br>\nI don't think the 3 points above has a huge impact on the result. The most important was the augmentation trick.</li>\n</ul>\n<p><strong>Training policy &amp; CV</strong><br>\nCV and validation set were the key.<br>\nI randomly remove 20% of patient in train and add them to test. (Based on k=2 clustering ) <br>\nSo 20% of Patient will not be part of CV they will be a validation set i'll not train on them.<br>\nNN : I used Mlpregressor from sklearn and I fixed params based on oof and validation predictions<br>\nRF: fixed params according to oof and validation predictions</p>\n<p><strong>Post-Processing</strong><br>\nBecause of the metric you can ajust Confidence to get better score.<br>\nThe binary classifier tell you for which observation you should do that.</p>",
      "rawMarkdown": "Thanks @johannhuber .\n\n**Dataset preparation**\n\n- I used ordinal encoder. ( eg: Male is more important than Female in Sex variable .....).\n- No scaling\n- Yes direct prediction of FVC.\nI don't think the 3 points above has a huge impact on the result. The most important was the augmentation trick.\n\n**Training policy & CV**\nCV and validation set were the key.\nI randomly remove 20% of patient in train and add them to test. (Based on k=2 clustering ) \nSo 20% of Patient will not be part of CV they will be a validation set i'll not train on them.\nNN : I used Mlpregressor from sklearn and I fixed params based on oof and validation predictions\nRF: fixed params according to oof and validation predictions\n\n\n\n**Post-Processing**\nBecause of the metric you can ajust Confidence to get better score.\nThe binary classifier tell you for which observation you should do that.",
      "votes": null
    },
    {
      "id": "1043843",
      "postDate": "10/09/2020 09:32:44",
      "content": "<p>Yes only tabular data !</p>",
      "rawMarkdown": "Yes only tabular data !",
      "votes": null
    },
    {
      "id": "1043844",
      "postDate": "10/09/2020 09:33:02",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/hypothermia15\" target=\"_blank\">@hypothermia15</a> </p>",
      "rawMarkdown": "Thanks @hypothermia15",
      "votes": null
    },
    {
      "id": "1043845",
      "postDate": "10/09/2020 09:33:24",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/code1110\" target=\"_blank\">@code1110</a> </p>",
      "rawMarkdown": "Thanks @code1110",
      "votes": null
    },
    {
      "id": "1044008",
      "postDate": "10/09/2020 12:19:24",
      "content": "<p>Awesome, thank you for those additional details !</p>",
      "rawMarkdown": "Awesome, thank you for those additional details !",
      "votes": null
    },
    {
      "id": "1045016",
      "postDate": "10/10/2020 09:31:31",
      "content": "<p>Congratulations….and Thanks for sharing..</p>",
      "rawMarkdown": "Congratulations....and Thanks for sharing..",
      "votes": null
    },
    {
      "id": "1045106",
      "postDate": "10/10/2020 10:56:33",
      "content": "<p><a href=\"https://www.kaggle.com/amedprof\" target=\"_blank\">@amedprof</a> Congrats on the solo gold!!</p>",
      "rawMarkdown": "amedprof Congrats on the solo gold!!",
      "votes": null
    },
    {
      "id": "1046819",
      "postDate": "10/12/2020 03:12:43",
      "content": "<p>Good work!!</p>",
      "rawMarkdown": "Good work!!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1040185,
      "author_name": "aadhavvignesh",
      "author_url": "",
      "post_date": "10/07/2020 02:49:08",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/amedprof\" target=\"_blank\">@amedprof</a> on the solo gold! It seems that your validation strategy helped you to get a good CV/LB correlation!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1040573,
          "author_name": "amedprof",
          "author_url": "",
          "post_date": "10/07/2020 08:23:47",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/aadhavvignesh\" target=\"_blank\">@aadhavvignesh</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1040362,
      "author_name": "imoore",
      "author_url": "",
      "post_date": "10/07/2020 05:15:32",
      "content": "<p>Great framework. Congratulations on your gold medal.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1040572,
          "author_name": "amedprof",
          "author_url": "",
          "post_date": "10/07/2020 08:23:22",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/imoore\" target=\"_blank\">@imoore</a> 😊</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1042119,
      "author_name": "code1110",
      "author_url": "",
      "post_date": "10/08/2020 04:31:33",
      "content": "<p>Congratulations! I'm impressed by your clean yet strong tabular only approach.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1043845,
          "author_name": "amedprof",
          "author_url": "",
          "post_date": "10/09/2020 09:33:24",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/code1110\" target=\"_blank\">@code1110</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1042578,
      "author_name": "johannhuber",
      "author_url": "",
      "post_date": "10/08/2020 10:24:25",
      "content": "<p>Félicitations ! You cv strategy is definitively interesting ; I still have many things to learn in that matter, I'm one of those who have been fooled by the public testset overfitting … ;)</p>\n<p>I have a couple of questions regarding your work :</p>\n<p><strong>Dataset preparation</strong></p>\n<ul>\n<li><a href=\"https://www.kaggle.com/ttahara/osic-baseline-lgbm-with-custom-metric\" target=\"_blank\">The mentioned kernel</a> suggests that you've used label encoding for <code>Sex</code> and <code>SmokingStatus</code>, right ? (not OHE ?)</li>\n<li>Did you scale <code>BaseFVC</code>, <code>BasePercent</code>, <code>Age</code>, and <code>WeekPassed</code> in some way ?</li>\n<li>Did you directly predict the <code>FVC</code> ? Or some percent of the <code>BaseFVC</code> ?</li>\n</ul>\n<p><strong>Training policy &amp; regularization</strong></p>\n<ul>\n<li>How did you manage to avoid overfitting in your models (in particular for the NN) : Did you rely on regularization (dropout, also BN as a side effect) ? What was your end-of-training policy ?</li>\n</ul>\n<p><strong>Post-processing &amp; Cross-validation</strong></p>\n<ul>\n<li>Can you elaborate on your post-processing strategy for the confidence prediction ? I don't understand the role of the final binary classifier.</li>\n<li>Your cross validation strategy is based on this great K=2 clustering you've described in <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/168610\" target=\"_blank\">this topic</a>. I assume you relied OOF validation on your 2 folds ? Or did you stick to a single train/set split, relying on several random seeds running to mitigate the randomness ? I guess the predictions from both folds had pretty close distribution …</li>\n</ul>\n<p>Congratz again, I look forward to your answer ! :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1043830,
          "author_name": "amedprof",
          "author_url": "",
          "post_date": "10/09/2020 09:24:11",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/johannhuber\" target=\"_blank\">@johannhuber</a> .</p>\n<p><strong>Dataset preparation</strong></p>\n<ul>\n<li>I used ordinal encoder. ( eg: Male is more important than Female in Sex variable …..).</li>\n<li>No scaling</li>\n<li>Yes direct prediction of FVC.<br>\nI don't think the 3 points above has a huge impact on the result. The most important was the augmentation trick.</li>\n</ul>\n<p><strong>Training policy &amp; CV</strong><br>\nCV and validation set were the key.<br>\nI randomly remove 20% of patient in train and add them to test. (Based on k=2 clustering ) <br>\nSo 20% of Patient will not be part of CV they will be a validation set i'll not train on them.<br>\nNN : I used Mlpregressor from sklearn and I fixed params based on oof and validation predictions<br>\nRF: fixed params according to oof and validation predictions</p>\n<p><strong>Post-Processing</strong><br>\nBecause of the metric you can ajust Confidence to get better score.<br>\nThe binary classifier tell you for which observation you should do that.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1044008,
          "author_name": "johannhuber",
          "author_url": "",
          "post_date": "10/09/2020 12:19:24",
          "content": "<p>Awesome, thank you for those additional details !</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1042592,
      "author_name": "nur988",
      "author_url": "",
      "post_date": "10/08/2020 10:46:44",
      "content": "<p>You achieved the results using only the tabular features?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1043843,
          "author_name": "amedprof",
          "author_url": "",
          "post_date": "10/09/2020 09:32:44",
          "content": "<p>Yes only tabular data !</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1045016,
      "author_name": "hiteshtripathi",
      "author_url": "",
      "post_date": "10/10/2020 09:31:31",
      "content": "<p>Congratulations….and Thanks for sharing..</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1045106,
      "author_name": "geekysaint",
      "author_url": "",
      "post_date": "10/10/2020 10:56:33",
      "content": "<p><a href=\"https://www.kaggle.com/amedprof\" target=\"_blank\">@amedprof</a> Congrats on the solo gold!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1042272,
      "author_name": "hypothermia15",
      "author_url": "",
      "post_date": "10/08/2020 06:41:08",
      "content": "<p>Congratulations <a href=\"/amedprof\">@amedprof</a> on the solo gold!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1043844,
          "author_name": "amedprof",
          "author_url": "",
          "post_date": "10/09/2020 09:33:02",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/hypothermia15\" target=\"_blank\">@hypothermia15</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1046819,
      "author_name": "jasleenkaur2020",
      "author_url": "",
      "post_date": "10/12/2020 03:12:43",
      "content": "<p>Good work!!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1040174": "Congratz to the winners !\n\nMany thanks to kaggle team for hosting this competition,hope the winning solutions will bring useful insight to the problem.\n\nMy solution is based only on tabular data.\n\n#  Data Augmentation\nThe strategy has been shared in many public kernels.\nYou can find it [here](https://www.kaggle.com/ttahara/osic-baseline-lgbm-with-custom-metric).\n# Validation \nI already shared my validation strategy [here ](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/168610)during the competition.\n\n# Features \nI use the following features : `'base_FVC', 'base_Percent', 'base_Age', 'Week_passed', 'Sex', 'SmokingStatus'`\n# FVC Prediction\nFinal fvc prediction is  a mean of the following regressor : \n- SVM\n- KNN\n- NN\n- Quantile Regressor (0.5)\n- RF\n- LM\n- HuberRegressor\n- ElasticNet\n- Lgbm\n\n# Confidence Prediction\n\n1 - Compute the optimal Confidence value for final fvc,which is actually:\n`Confidence = np.sqrt(2)*np.abs(FVC_true - FVC_pred)`\n2 - Use the same models in FVC  prediction to train on the Confidence this time\n3 - Train a binary classifier to know if `Confidence<=100`\n4 - Use prediction from (2) and post-process them by (3)\n\n# Things that didn't work for me\n- Everything with image data\n- Linear Augmentation on tabular data",
    "1040185": "Congratulations @amedprof on the solo gold! It seems that your validation strategy helped you to get a good CV/LB correlation!",
    "1040362": "Great framework. Congratulations on your gold medal.",
    "1040572": "Thanks @imoore 😊",
    "1040573": "Thanks @aadhavvignesh",
    "1042119": "Congratulations! I'm impressed by your clean yet strong tabular only approach.",
    "1042272": "Congratulations @amedprof on the solo gold!",
    "1042578": "Félicitations ! You cv strategy is definitively interesting ; I still have many things to learn in that matter, I'm one of those who have been fooled by the public testset overfitting ... ;)\n\nI have a couple of questions regarding your work :\n\n\n**Dataset preparation**\n- [The mentioned kernel](https://www.kaggle.com/ttahara/osic-baseline-lgbm-with-custom-metric) suggests that you've used label encoding for `Sex` and `SmokingStatus`, right ? (not OHE ?)\n- Did you scale `BaseFVC`, `BasePercent`, `Age`, and `WeekPassed` in some way ?\n- Did you directly predict the `FVC` ? Or some percent of the `BaseFVC` ?\n\n**Training policy & regularization**\n- How did you manage to avoid overfitting in your models (in particular for the NN) : Did you rely on regularization (dropout, also BN as a side effect) ? What was your end-of-training policy ?\n\n**Post-processing & Cross-validation**\n- Can you elaborate on your post-processing strategy for the confidence prediction ? I don't understand the role of the final binary classifier.\n- Your cross validation strategy is based on this great K=2 clustering you've described in [this topic](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/168610). I assume you relied OOF validation on your 2 folds ? Or did you stick to a single train/set split, relying on several random seeds running to mitigate the randomness ? I guess the predictions from both folds had pretty close distribution ...\n\n\nCongratz again, I look forward to your answer ! :)",
    "1042592": "You achieved the results using only the tabular features?",
    "1043830": "Thanks @johannhuber .\n\n**Dataset preparation**\n\n- I used ordinal encoder. ( eg: Male is more important than Female in Sex variable .....).\n- No scaling\n- Yes direct prediction of FVC.\nI don't think the 3 points above has a huge impact on the result. The most important was the augmentation trick.\n\n**Training policy & CV**\nCV and validation set were the key.\nI randomly remove 20% of patient in train and add them to test. (Based on k=2 clustering ) \nSo 20% of Patient will not be part of CV they will be a validation set i'll not train on them.\nNN : I used Mlpregressor from sklearn and I fixed params based on oof and validation predictions\nRF: fixed params according to oof and validation predictions\n\n\n\n**Post-Processing**\nBecause of the metric you can ajust Confidence to get better score.\nThe binary classifier tell you for which observation you should do that.",
    "1043843": "Yes only tabular data !",
    "1043844": "Thanks @hypothermia15",
    "1043845": "Thanks @code1110",
    "1044008": "Awesome, thank you for those additional details !",
    "1045016": "Congratulations....and Thanks for sharing..",
    "1045106": "amedprof Congrats on the solo gold!!",
    "1046819": "Good work!!"
  },
  "source": "meta"
}