{
  "id": 189220,
  "title": "6th place solution",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/writeups/y-nakama-6th-place-solution",
  "author_name": "",
  "post_date": "2020-10-08T22:01:59.227Z",
  "votes": 79,
  "comment_count": 41,
  "views": 0,
  "content": "<p>I would like to thank kaggle &amp; host for the interesting competition and to all the participants for giving me a lot of ideas. And congrats to winners!</p>\n<p>My work is based on </p>\n<ul>\n<li>My tabular approach notebooks<ul>\n<li><a href=\"https://www.kaggle.com/yasufuminakama/osic-lgb-baseline\" target=\"_blank\">https://www.kaggle.com/yasufuminakama/osic-lgb-baseline</a> </li>\n<li><a href=\"https://www.kaggle.com/yasufuminakama/osic-ridge-baseline\" target=\"_blank\">https://www.kaggle.com/yasufuminakama/osic-ridge-baseline</a> </li></ul></li>\n<li>Image approach notebooks<ul>\n<li><a href=\"https://www.kaggle.com/miklgr500/linear-decay-based-on-resnet-cnn\" target=\"_blank\">https://www.kaggle.com/miklgr500/linear-decay-based-on-resnet-cnn</a></li>\n<li><a href=\"https://www.kaggle.com/khoongweihao/k-fold-tf-efficientnet-models-training\" target=\"_blank\">https://www.kaggle.com/khoongweihao/k-fold-tf-efficientnet-models-training</a></li></ul></li>\n</ul>\n<h1>Solution Overview</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1695531%2F1f5423ba0fc5e5bb3bc182c0b8a75b5a%2FOSIC-solution.png?generation=1602040115132884&amp;alt=media\" alt=\"OSIC-solution-overview\"></p>\n<h1>How to construct input data</h1>\n<p>As I showed in <a href=\"https://www.kaggle.com/yasufuminakama/osic-lgb-baseline\" target=\"_blank\">my notebook</a>, I constructed input data by treating every single measurement as if it were a “baseline” measurement. And “Week” is used to create “Week_passed” for each measurements. We don’t know which “Week” is given in test data, so this input data construction gives us robustness to build models.</p>\n<h1>How to treat image data</h1>\n<p>Based on image approach notebooks, I used efficientnet-b0 for image size 320x320. <br>\nI tried other efficientnet models but there are not so much different results among them when I fix quantile value as 0.5, so I used the smallest one.<br>\nThis efficientnet-b0 output is used for FVC &amp; Confidence model training with tabular features.</p>\n<h1>FVC prediction</h1>\n<p>FVC prediction is almost same as my notebooks, the difference is I prepared 5 models and blend them automatically using <code>sp.optimize.minimize</code>, weight is [Lasso, Ridge, ElasticNet, SVM, NN] = [0.68573749, 0., 0., 0.07551167, 0.23750526].</p>\n<h1>Confidence prediction</h1>\n<p>Confidence labels are made using FVC oof &amp; <code>sp.optimize.minimize</code> as shown in my notebooks.<br>\nConfidence prediction is almost same as my notebooks too, the difference is I prepared 5 models and blend them automatically using <code>sp.optimize.minimize</code>, weight is [Lasso, Ridge, ElasticNet, SVM, NN] = [0.22062125, 0., 0., 0., 0.80819966].</p>\n<h1>About shake up</h1>\n<p>There are so many public notebooks which overfits Public LB.<br>\nPublic LB has not so much data, so you don't need to care about Public LB so much.</p>\n<h1>CV and LB Transition</h1>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>CV</th>\n<th>Public</th>\n<th>Private</th>\n<th>Medal</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>LGB</td>\n<td>-6.85094</td>\n<td>-6.9605</td>\n<td>-7.0037</td>\n<td>None</td>\n</tr>\n<tr>\n<td>Ridge</td>\n<td>-6.73738</td>\n<td>-6.9357</td>\n<td>-6.8562</td>\n<td>Bronze</td>\n</tr>\n<tr>\n<td>Efficientnet-b0 + Ridge</td>\n<td>-6.58651</td>\n<td>-6.8921</td>\n<td>-6.8443</td>\n<td>Silver</td>\n</tr>\n<tr>\n<td>Efficientnet-b0 + Blend models v1</td>\n<td>-6.54917</td>\n<td>-6.8922</td>\n<td>-6.8421</td>\n<td>Top Silver</td>\n</tr>\n<tr>\n<td>Efficientnet-b0 + Blend models v2</td>\n<td>-6.53815</td>\n<td>-6.8972</td>\n<td>-6.8363</td>\n<td>Gold</td>\n</tr>\n</tbody>\n</table>\n<h1>Final result</h1>\n<p>I selected my best CV, and it was the best Private LB submission.<br>\nCV (4fold): -6.53815<br>\nPrivate LB: -6.8363</p>",
  "messages": [
    {
      "id": "1040186",
      "postDate": "10/07/2020 02:49:25",
      "content": "<p>I would like to thank kaggle &amp; host for the interesting competition and to all the participants for giving me a lot of ideas. And congrats to winners!</p>\n<p>My work is based on </p>\n<ul>\n<li>My tabular approach notebooks<ul>\n<li><a href=\"https://www.kaggle.com/yasufuminakama/osic-lgb-baseline\" target=\"_blank\">https://www.kaggle.com/yasufuminakama/osic-lgb-baseline</a> </li>\n<li><a href=\"https://www.kaggle.com/yasufuminakama/osic-ridge-baseline\" target=\"_blank\">https://www.kaggle.com/yasufuminakama/osic-ridge-baseline</a> </li></ul></li>\n<li>Image approach notebooks<ul>\n<li><a href=\"https://www.kaggle.com/miklgr500/linear-decay-based-on-resnet-cnn\" target=\"_blank\">https://www.kaggle.com/miklgr500/linear-decay-based-on-resnet-cnn</a></li>\n<li><a href=\"https://www.kaggle.com/khoongweihao/k-fold-tf-efficientnet-models-training\" target=\"_blank\">https://www.kaggle.com/khoongweihao/k-fold-tf-efficientnet-models-training</a></li></ul></li>\n</ul>\n<h1>Solution Overview</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1695531%2F1f5423ba0fc5e5bb3bc182c0b8a75b5a%2FOSIC-solution.png?generation=1602040115132884&amp;alt=media\" alt=\"OSIC-solution-overview\"></p>\n<h1>How to construct input data</h1>\n<p>As I showed in <a href=\"https://www.kaggle.com/yasufuminakama/osic-lgb-baseline\" target=\"_blank\">my notebook</a>, I constructed input data by treating every single measurement as if it were a “baseline” measurement. And “Week” is used to create “Week_passed” for each measurements. We don’t know which “Week” is given in test data, so this input data construction gives us robustness to build models.</p>\n<h1>How to treat image data</h1>\n<p>Based on image approach notebooks, I used efficientnet-b0 for image size 320x320. <br>\nI tried other efficientnet models but there are not so much different results among them when I fix quantile value as 0.5, so I used the smallest one.<br>\nThis efficientnet-b0 output is used for FVC &amp; Confidence model training with tabular features.</p>\n<h1>FVC prediction</h1>\n<p>FVC prediction is almost same as my notebooks, the difference is I prepared 5 models and blend them automatically using <code>sp.optimize.minimize</code>, weight is [Lasso, Ridge, ElasticNet, SVM, NN] = [0.68573749, 0., 0., 0.07551167, 0.23750526].</p>\n<h1>Confidence prediction</h1>\n<p>Confidence labels are made using FVC oof &amp; <code>sp.optimize.minimize</code> as shown in my notebooks.<br>\nConfidence prediction is almost same as my notebooks too, the difference is I prepared 5 models and blend them automatically using <code>sp.optimize.minimize</code>, weight is [Lasso, Ridge, ElasticNet, SVM, NN] = [0.22062125, 0., 0., 0., 0.80819966].</p>\n<h1>About shake up</h1>\n<p>There are so many public notebooks which overfits Public LB.<br>\nPublic LB has not so much data, so you don't need to care about Public LB so much.</p>\n<h1>CV and LB Transition</h1>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>CV</th>\n<th>Public</th>\n<th>Private</th>\n<th>Medal</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>LGB</td>\n<td>-6.85094</td>\n<td>-6.9605</td>\n<td>-7.0037</td>\n<td>None</td>\n</tr>\n<tr>\n<td>Ridge</td>\n<td>-6.73738</td>\n<td>-6.9357</td>\n<td>-6.8562</td>\n<td>Bronze</td>\n</tr>\n<tr>\n<td>Efficientnet-b0 + Ridge</td>\n<td>-6.58651</td>\n<td>-6.8921</td>\n<td>-6.8443</td>\n<td>Silver</td>\n</tr>\n<tr>\n<td>Efficientnet-b0 + Blend models v1</td>\n<td>-6.54917</td>\n<td>-6.8922</td>\n<td>-6.8421</td>\n<td>Top Silver</td>\n</tr>\n<tr>\n<td>Efficientnet-b0 + Blend models v2</td>\n<td>-6.53815</td>\n<td>-6.8972</td>\n<td>-6.8363</td>\n<td>Gold</td>\n</tr>\n</tbody>\n</table>\n<h1>Final result</h1>\n<p>I selected my best CV, and it was the best Private LB submission.<br>\nCV (4fold): -6.53815<br>\nPrivate LB: -6.8363</p>",
      "rawMarkdown": "I would like to thank kaggle & host for the interesting competition and to all the participants for giving me a lot of ideas. And congrats to winners!\n\nMy work is based on \n- My tabular approach notebooks\n  - https://www.kaggle.com/yasufuminakama/osic-lgb-baseline \n  - https://www.kaggle.com/yasufuminakama/osic-ridge-baseline \n- Image approach notebooks\n  - https://www.kaggle.com/miklgr500/linear-decay-based-on-resnet-cnn\n  - https://www.kaggle.com/khoongweihao/k-fold-tf-efficientnet-models-training\n\n# Solution Overview\n![OSIC-solution-overview](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1695531%2F1f5423ba0fc5e5bb3bc182c0b8a75b5a%2FOSIC-solution.png?generation=1602040115132884&alt=media)\n\n# How to construct input data\nAs I showed in [my notebook](https://www.kaggle.com/yasufuminakama/osic-lgb-baseline), I constructed input data by treating every single measurement as if it were a “baseline” measurement. And “Week” is used to create “Week_passed” for each measurements. We don’t know which “Week” is given in test data, so this input data construction gives us robustness to build models.\n\n# How to treat image data\nBased on image approach notebooks, I used efficientnet-b0 for image size 320x320. \nI tried other efficientnet models but there are not so much different results among them when I fix quantile value as 0.5, so I used the smallest one.\nThis efficientnet-b0 output is used for FVC & Confidence model training with tabular features.\n\n# FVC prediction\nFVC prediction is almost same as my notebooks, the difference is I prepared 5 models and blend them automatically using ```sp.optimize.minimize```, weight is [Lasso, Ridge, ElasticNet, SVM, NN] = [0.68573749, 0., 0., 0.07551167, 0.23750526].\n\n# Confidence prediction\nConfidence labels are made using FVC oof & ```sp.optimize.minimize``` as shown in my notebooks.\nConfidence prediction is almost same as my notebooks too, the difference is I prepared 5 models and blend them automatically using ```sp.optimize.minimize```, weight is [Lasso, Ridge, ElasticNet, SVM, NN] = [0.22062125, 0., 0., 0., 0.80819966].\n\n# About shake up\nThere are so many public notebooks which overfits Public LB.\nPublic LB has not so much data, so you don't need to care about Public LB so much.\n\n# CV and LB Transition\n|  Model |  CV  |  Public  | Private | Medal |\n| ---- | ---- | ---- | ---- | ---- |\n|  LGB | -6.85094 | -6.9605 | -7.0037 | None |\n|  Ridge | -6.73738 | -6.9357 | -6.8562 | Bronze |\n|  Efficientnet-b0 + Ridge | -6.58651 | -6.8921 | -6.8443 | Silver |\n|  Efficientnet-b0 + Blend models v1 | -6.54917 | -6.8922 | -6.8421 | Top Silver |\n|  Efficientnet-b0 + Blend models v2 | -6.53815 | -6.8972 | -6.8363 | Gold |\n\n# Final result \nI selected my best CV, and it was the best Private LB submission.\nCV (4fold): -6.53815\nPrivate LB: -6.8363",
      "votes": null
    },
    {
      "id": "1040188",
      "postDate": "10/07/2020 02:50:51",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> on the solo gold! </p>",
      "rawMarkdown": "Congratulations @yasufuminakama on the solo gold!",
      "votes": null
    },
    {
      "id": "1040279",
      "postDate": "10/07/2020 04:13:46",
      "content": "<p>Congratulations and great work <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> . I am curious about the contribution from CT dicom files in your solution. Since some solutions get good score with tabular data only, like <a href=\"url\" target=\"_blank\">https://www.kaggle.com/lhagiimn/solution-for-the-first-place-but-we-didn-t-select</a> and <a href=\"url\" target=\"_blank\">https://www.kaggle.com/ulrich07/osic-basic-tabular-data-augmentation-nn?scriptVersionId=39788369</a>. </p>",
      "rawMarkdown": "Congratulations and great work @yasufuminakama . I am curious about the contribution from CT dicom files in your solution. Since some solutions get good score with tabular data only, like [https://www.kaggle.com/lhagiimn/solution-for-the-first-place-but-we-didn-t-select](url) and [https://www.kaggle.com/ulrich07/osic-basic-tabular-data-augmentation-nn?scriptVersionId=39788369](url).",
      "votes": null
    },
    {
      "id": "1040307",
      "postDate": "10/07/2020 04:40:15",
      "content": "<p>Congrats for your solo gold 🎉 I'm really impressed with your strategy. Thank you for sharing.</p>",
      "rawMarkdown": "Congrats for your solo gold 🎉 I'm really impressed with your strategy. Thank you for sharing.",
      "votes": null
    },
    {
      "id": "1040313",
      "postDate": "10/07/2020 04:44:46",
      "content": "<p>Thanks!<br>\nI have tabular only submission that scores private -6.8366 for gold, but CV using CT dicom files is much better than tabular only, so IMHO tabular only submission can get gold with some luck, but using CT dicom files leads to more solid result.</p>",
      "rawMarkdown": "Thanks!\nI have tabular only submission that scores private -6.8366 for gold, but CV using CT dicom files is much better than tabular only, so IMHO tabular only submission can get gold with some luck, but using CT dicom files leads to more solid result.",
      "votes": null
    },
    {
      "id": "1040316",
      "postDate": "10/07/2020 04:46:41",
      "content": "<p>Thanks! :)</p>",
      "rawMarkdown": "Thanks! :)",
      "votes": null
    },
    {
      "id": "1040317",
      "postDate": "10/07/2020 04:46:59",
      "content": "<p>Thank you! :) </p>",
      "rawMarkdown": "Thank you! :)",
      "votes": null
    },
    {
      "id": "1040332",
      "postDate": "10/07/2020 04:57:29",
      "content": "<p>Congratulations!!!!!</p>",
      "rawMarkdown": "Congratulations!!!!!",
      "votes": null
    },
    {
      "id": "1040337",
      "postDate": "10/07/2020 04:59:37",
      "content": "<p>Thank you for your reply. I am agree with your opinion. </p>",
      "rawMarkdown": "Thank you for your reply. I am agree with your opinion.",
      "votes": null
    },
    {
      "id": "1040374",
      "postDate": "10/07/2020 05:23:07",
      "content": "<p>Your solution is simple and cool.<br>\nThanks for your sharing!</p>",
      "rawMarkdown": "Your solution is simple and cool.\nThanks for your sharing!",
      "votes": null
    },
    {
      "id": "1040394",
      "postDate": "10/07/2020 05:41:18",
      "content": "<p>Congratulations for your 6th place!<br>\nThanks for sharing great works and ideas. </p>",
      "rawMarkdown": "Congratulations for your 6th place!\nThanks for sharing great works and ideas.",
      "votes": null
    },
    {
      "id": "1040435",
      "postDate": "10/07/2020 06:17:28",
      "content": "<p>Thanks you!</p>",
      "rawMarkdown": "Thanks you!",
      "votes": null
    },
    {
      "id": "1040543",
      "postDate": "10/07/2020 07:55:52",
      "content": "<p>Congratulations you deserve it ! thank you for your baseline kernels btw</p>",
      "rawMarkdown": "Congratulations you deserve it ! thank you for your baseline kernels btw",
      "votes": null
    },
    {
      "id": "1040578",
      "postDate": "10/07/2020 08:26:20",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!",
      "votes": null
    },
    {
      "id": "1040617",
      "postDate": "10/07/2020 09:01:34",
      "content": "<p>Congratulations , I wouldn't have achive gold medal without your kernels.</p>",
      "rawMarkdown": "Congratulations , I wouldn't have achive gold medal without your kernels.",
      "votes": null
    },
    {
      "id": "1040646",
      "postDate": "10/07/2020 09:14:01",
      "content": "<p>Thank you! Glad to hear that, Congrats to your gold too :)</p>",
      "rawMarkdown": "Thank you! Glad to hear that, Congrats to your gold too :)",
      "votes": null
    },
    {
      "id": "1040655",
      "postDate": "10/07/2020 09:18:47",
      "content": "<p>Congrats! What is your last3 cv? And did CV and private score correlate well?<br>\nMaybe data augmentation of tabular data is key to survive the shake up…</p>",
      "rawMarkdown": "Congrats! What is your last3 cv? And did CV and private score correlate well?\nMaybe data augmentation of tabular data is key to survive the shake up...",
      "votes": null
    },
    {
      "id": "1040671",
      "postDate": "10/07/2020 09:29:40",
      "content": "<p>Thanks!</p>\n<p>I don't think we should monitor last3 cv because it's smaller to trust, so I haven't checked it.</p>\n<p>My recent submissions are as follows, seems CV and Private correlate well.</p>\n<ul>\n<li>CV: -6.58651 Public: -6.8921 Private: -6.8443</li>\n<li>CV: -6.54917 Public: -6.8922 Private: -6.8421</li>\n<li>CV: -6.53815 Public: -6.8972 Private: -6.8363</li>\n</ul>\n<p>Yes, I think my input data construction strategy is important.</p>",
      "rawMarkdown": "Thanks!\n\nI don't think we should monitor last3 cv because it's smaller to trust, so I haven't checked it.\n\nMy recent submissions are as follows, seems CV and Private correlate well.\n- CV: -6.58651 Public: -6.8921 Private: -6.8443\n- CV: -6.54917 Public: -6.8922 Private: -6.8421\n- CV: -6.53815 Public: -6.8972 Private: -6.8363\n\nYes, I think my input data construction strategy is important.",
      "votes": null
    },
    {
      "id": "1040679",
      "postDate": "10/07/2020 09:34:15",
      "content": "<p>Congratulations to your solo Gold!</p>\n<p>My best decision in this competition was to make my solution based on your public kernel.  Thanks for your contribution <br>\nto the community, and I'm looking forward to the day of you becoming the GM!</p>",
      "rawMarkdown": "Congratulations to your solo Gold!\n\nMy best decision in this competition was to make my solution based on your public kernel.  Thanks for your contribution \nto the community, and I'm looking forward to the day of you becoming the GM!",
      "votes": null
    },
    {
      "id": "1040685",
      "postDate": "10/07/2020 09:39:39",
      "content": "<p>Thanks you!<br>\nGlad to hear that, I'll do my best :)</p>",
      "rawMarkdown": "Thanks you!\nGlad to hear that, I'll do my best :)",
      "votes": null
    },
    {
      "id": "1040726",
      "postDate": "10/07/2020 10:07:03",
      "content": "<p>Oh OK. Last 3 CV is untrustworthy, then I was totally wrong from the start point…<br>\nIf I used your data augmentation strategy!!!!!<br>\nAnd you have high ensemble weight on Lasso, which is similar to me…<br>\nMy solution:<br>\n<a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/189341\" target=\"_blank\">https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/189341</a></p>\n<p>My inspiration that Lasso is effective was proven to be true…</p>",
      "rawMarkdown": "Oh OK. Last 3 CV is untrustworthy, then I was totally wrong from the start point...\nIf I used your data augmentation strategy!!!!!\nAnd you have high ensemble weight on Lasso, which is similar to me...\nMy solution:\nhttps://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/189341\n\nMy inspiration that Lasso is effective was proven to be true...",
      "votes": null
    },
    {
      "id": "1040805",
      "postDate": "10/07/2020 11:05:12",
      "content": "<p>Well done on getting 6th place; I find your notebooks are very helpful! </p>\n<p>You mention your cv was -6.53815 - but is this for all data or just last 3 points? </p>",
      "rawMarkdown": "Well done on getting 6th place; I find your notebooks are very helpful! \n\nYou mention your cv was -6.53815 - but is this for all data or just last 3 points?",
      "votes": null
    },
    {
      "id": "1040866",
      "postDate": "10/07/2020 12:12:36",
      "content": "<p>Your idea really helped me getting my first ever Silver medal on Kaggle:</p>\n<pre><code>Treating every single measurement as if it were a “baseline” measurement. And “Week” is used to create “Week_passed” for each measurements. We don’t know which “Week” is given in test data, so this input data construction gives us robustness to build models.\n</code></pre>\n<p>Using this idea I performed some more data augmentations and a GBR to create my predictions. Many congrats to your Solo Gold!</p>",
      "rawMarkdown": "Your idea really helped me getting my first ever Silver medal on Kaggle:\n\n    Treating every single measurement as if it were a “baseline” measurement. And “Week” is used to create “Week_passed” for each measurements. We don’t know which “Week” is given in test data, so this input data construction gives us robustness to build models.\n\nUsing this idea I performed some more data augmentations and a GBR to create my predictions. Many congrats to your Solo Gold!",
      "votes": null
    },
    {
      "id": "1041015",
      "postDate": "10/07/2020 13:55:24",
      "content": "<p>Thank you! Glad to hear that :)<br>\nIt's for all data, I don't think we should monitor last3 cv because it's smaller to trust.</p>",
      "rawMarkdown": "Thank you! Glad to hear that :)\nIt's for all data, I don't think we should monitor last3 cv because it's smaller to trust.",
      "votes": null
    },
    {
      "id": "1041017",
      "postDate": "10/07/2020 13:56:34",
      "content": "<p>Thank you &amp; Congrats to your effort and result!</p>",
      "rawMarkdown": "Thank you & Congrats to your effort and result!",
      "votes": null
    },
    {
      "id": "1041025",
      "postDate": "10/07/2020 14:03:12",
      "content": "<p>Congrats! How did you actually use the image data? From what I experimented using images did not help with the predictions much, I randomly sampled 2~3 images as channels around the middle of all the scans the patient has but the performance was the same/lower than using tabular data alone. And I tried out the public kernel which uses efficient, it just randomly samples from any scans (if i understand correctly), I changed the images size to smaller number (e.g. 128x128) or even commented out all the cnn branch, it didn't change the score (even improved the local score in that kernel a bit). So im really curious how you used images. </p>\n<p>It has a few issues IMHO: 1) what set of images to use? there are patients with many number of scans while some only has 14 scans (also pixel spacing is different, its probably arguable if its needed to scale it, plus slice thickness is also different). 2) the image position (i.e. head or feet first) affects the order of the lung slices (e.g. bottom of lung appear early or later), reorder is probably needed based on the imagePosition attribute, but I also observed at least a few imagePosition value is incorrect (i.e. head first images labelled as FFS etc). 3) there were a few patients in train doesnt have complete lung scans, and one patient has two lungs over all his/her scan, although this might be okay given its a small portion 4) the number of training data is small, i mean the unique number of patients is small and the slices that actually carry useful pattern is small (i.e. probably mainly the bottom of the lung)</p>\n<p>so i switched from using image directly to extract features, such as skews and kurtosis and total volumes, these features actually shows high correlation with FVC (at least in our train set…., e.g. lung_volume as ~0.84 Pearson correlation with FVC) and others has large magnitude coefficient in my bayesian ridge regression models just trailing behind base_FVC and base_Percent. But this didnt help us at all in the private test set…. </p>",
      "rawMarkdown": "Congrats! How did you actually use the image data? From what I experimented using images did not help with the predictions much, I randomly sampled 2~3 images as channels around the middle of all the scans the patient has but the performance was the same/lower than using tabular data alone. And I tried out the public kernel which uses efficient, it just randomly samples from any scans (if i understand correctly), I changed the images size to smaller number (e.g. 128x128) or even commented out all the cnn branch, it didn't change the score (even improved the local score in that kernel a bit). So im really curious how you used images. \n\nIt has a few issues IMHO: 1) what set of images to use? there are patients with many number of scans while some only has 14 scans (also pixel spacing is different, its probably arguable if its needed to scale it, plus slice thickness is also different). 2) the image position (i.e. head or feet first) affects the order of the lung slices (e.g. bottom of lung appear early or later), reorder is probably needed based on the imagePosition attribute, but I also observed at least a few imagePosition value is incorrect (i.e. head first images labelled as FFS etc). 3) there were a few patients in train doesnt have complete lung scans, and one patient has two lungs over all his/her scan, although this might be okay given its a small portion 4) the number of training data is small, i mean the unique number of patients is small and the slices that actually carry useful pattern is small (i.e. probably mainly the bottom of the lung)\n\nso i switched from using image directly to extract features, such as skews and kurtosis and total volumes, these features actually shows high correlation with FVC (at least in our train set...., e.g. lung_volume as ~0.84 Pearson correlation with FVC) and others has large magnitude coefficient in my bayesian ridge regression models just trailing behind base_FVC and base_Percent. But this didnt help us at all in the private test set....",
      "votes": null
    },
    {
      "id": "1041065",
      "postDate": "10/07/2020 14:30:40",
      "content": "<p>Thanks!</p>\n<ul>\n<li>How did you actually use the image data?<ul>\n<li>Training is almost same method as public notebooks I wrote above. I simply used efficientnet-b0 oof predictions for each patient to train other models. Quantile value is fixed to 0.5 when I get predictions. <strong>Linear decay for weeks is not applied.</strong></li></ul></li>\n<li>From what I experimented using images did not help with the predictions much<ul>\n<li>I'm not sure about your CV, but in my experiments for example my single ridge improved CV around 0.15 by using image predictions.</li></ul></li>\n</ul>",
      "rawMarkdown": "Thanks!\n- How did you actually use the image data?\n  - Training is almost same method as public notebooks I wrote above. I simply used efficientnet-b0 oof predictions for each patient to train other models. Quantile value is fixed to 0.5 when I get predictions. **Linear decay for weeks is not applied.**\n- From what I experimented using images did not help with the predictions much\n  - I'm not sure about your CV, but in my experiments for example my single ridge improved CV around 0.15 by using image predictions.",
      "votes": null
    },
    {
      "id": "1041080",
      "postDate": "10/07/2020 14:46:27",
      "content": "<p>Thank you for sharing your wonderful notebook, and congratulations on Solo Gold.<br>\nI added more data by doing the same thing with your notebook. But I made one change and improved the CV.</p>\n<p>In your notebook, <code>Week_passed</code> is roughly -60 to 60 range. Since the predictions are always after <code>base_Week</code>, I thought negative values weren't important, so I did the following.</p>\n<pre><code>train = output[(-12&lt;=output['Week_passed']) &amp; (output['Week_passed']! =0)].reset_index(drop=True)\n</code></pre>\n<p>A value of -12 gave me the highest CV.<br>\nMaybe I just accidentally got a higher CV, but try it out when you have the time!<br>\nAnyway, Thanks to you, I got my first medal! Thank you.</p>",
      "rawMarkdown": "Thank you for sharing your wonderful notebook, and congratulations on Solo Gold.\nI added more data by doing the same thing with your notebook. But I made one change and improved the CV.\n\nIn your notebook, `Week_passed` is roughly -60 to 60 range. Since the predictions are always after `base_Week`, I thought negative values weren't important, so I did the following.\n````\ntrain = output[(-12<=output['Week_passed']) & (output['Week_passed']! =0)].reset_index(drop=True)\n````\nA value of -12 gave me the highest CV.\nMaybe I just accidentally got a higher CV, but try it out when you have the time!\nAnyway, Thanks to you, I got my first medal! Thank you.",
      "votes": null
    },
    {
      "id": "1041104",
      "postDate": "10/07/2020 15:04:32",
      "content": "<p>I think the reason you got higher CV is that it's not same Cross-Validation before you apply <code>-12&lt;=output['Week_passed']</code>, you can't compare them.<br>\nThanks &amp; Congrats to your result!</p>",
      "rawMarkdown": "I think the reason you got higher CV is that it's not same Cross-Validation before you apply `-12<=output['Week_passed']`, you can't compare them.\nThanks & Congrats to your result!",
      "votes": null
    },
    {
      "id": "1041206",
      "postDate": "10/07/2020 16:01:35",
      "content": "<p>Indeed. I was wrong.<br>\nI'm sorry for missing the point.</p>",
      "rawMarkdown": "Indeed. I was wrong.\nI'm sorry for missing the point.",
      "votes": null
    },
    {
      "id": "1041309",
      "postDate": "10/07/2020 17:17:55",
      "content": "<p>Super… such a useful Information </p>",
      "rawMarkdown": "Super... such a useful Information",
      "votes": null
    },
    {
      "id": "1041311",
      "postDate": "10/07/2020 17:18:20",
      "content": "<p>Thanks for sharing this</p>",
      "rawMarkdown": "Thanks for sharing this",
      "votes": null
    },
    {
      "id": "1041769",
      "postDate": "10/07/2020 22:55:48",
      "content": "<p>Very Interesting </p>",
      "rawMarkdown": "Very Interesting",
      "votes": null
    },
    {
      "id": "1041952",
      "postDate": "10/08/2020 01:45:09",
      "content": "<p>Now, you became Grandmaster with the result of OpenVaccine. Congratulations again 🎉🎉🎉<br>\n<a href=\"https://www.kaggle.com/c/stanford-covid-vaccine\" target=\"_blank\">https://www.kaggle.com/c/stanford-covid-vaccine</a></p>",
      "rawMarkdown": "Now, you became Grandmaster with the result of OpenVaccine. Congratulations again 🎉🎉🎉\nhttps://www.kaggle.com/c/stanford-covid-vaccine",
      "votes": null
    },
    {
      "id": "1041986",
      "postDate": "10/08/2020 02:10:52",
      "content": "<p>Thanks again!</p>",
      "rawMarkdown": "Thanks again!",
      "votes": null
    },
    {
      "id": "1042095",
      "postDate": "10/08/2020 04:01:39",
      "content": "<p>Congratulations and thanks for sharing :) </p>",
      "rawMarkdown": "Congratulations and thanks for sharing :)",
      "votes": null
    },
    {
      "id": "1042612",
      "postDate": "10/08/2020 11:06:28",
      "content": "<p>Updated CV and LB Transition, CV and Private correlate well :)</p>",
      "rawMarkdown": "Updated CV and LB Transition, CV and Private correlate well :)",
      "votes": null
    },
    {
      "id": "1046317",
      "postDate": "10/11/2020 14:45:43",
      "content": "<p>Hi, I don’t know whether I have missed something, but actually I have the similar idea to you.<br>\nI did combine the test tabular and the train tabular together and after that, since the test data is 730 rows and have many rows with week_passed less than zero, I delete all of them and then do the whole minmax scalar to both the train and the test. </p>",
      "rawMarkdown": "Hi, I don’t know whether I have missed something, but actually I have the similar idea to you.\nI did combine the test tabular and the train tabular together and after that, since the test data is 730 rows and have many rows with week_passed less than zero, I delete all of them and then do the whole minmax scalar to both the train and the test.",
      "votes": null
    },
    {
      "id": "1048386",
      "postDate": "10/13/2020 12:48:42",
      "content": "<p>Congratulations and thank for sharing your solution!</p>",
      "rawMarkdown": "Congratulations and thank for sharing your solution!",
      "votes": null
    },
    {
      "id": "1048534",
      "postDate": "10/13/2020 15:18:29",
      "content": "<p>Congratulations!! Thanks for sharing</p>",
      "rawMarkdown": "Congratulations!! Thanks for sharing",
      "votes": null
    },
    {
      "id": "1048643",
      "postDate": "10/13/2020 17:13:56",
      "content": "<p>Congratulations and Thanks for sharing : ))</p>",
      "rawMarkdown": "Congratulations and Thanks for sharing : ))",
      "votes": null
    },
    {
      "id": "1387413",
      "postDate": "07/14/2021 06:56:55",
      "content": "<p>For the image part, I am not understanding how you are not using a 3D conv network.<br>\nAs each datapoint has lots of images, and together they form a 3d data, so are you treating each channel of this 3D data as an image separately? and feeding it to a pretrained conv network?</p>",
      "rawMarkdown": "For the image part, I am not understanding how you are not using a 3D conv network.\nAs each datapoint has lots of images, and together they form a 3d data, so are you treating each channel of this 3D data as an image separately? and feeding it to a pretrained conv network?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1040188,
      "author_name": "aadhavvignesh",
      "author_url": "",
      "post_date": "10/07/2020 02:50:51",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> on the solo gold! </p>",
      "votes": null,
      "replies": [
        {
          "id": 1040317,
          "author_name": "yasufuminakama",
          "author_url": "",
          "post_date": "10/07/2020 04:46:59",
          "content": "<p>Thank you! :) </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1040279,
      "author_name": "koalasheep",
      "author_url": "",
      "post_date": "10/07/2020 04:13:46",
      "content": "<p>Congratulations and great work <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> . I am curious about the contribution from CT dicom files in your solution. Since some solutions get good score with tabular data only, like <a href=\"url\" target=\"_blank\">https://www.kaggle.com/lhagiimn/solution-for-the-first-place-but-we-didn-t-select</a> and <a href=\"url\" target=\"_blank\">https://www.kaggle.com/ulrich07/osic-basic-tabular-data-augmentation-nn?scriptVersionId=39788369</a>. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1040313,
          "author_name": "yasufuminakama",
          "author_url": "",
          "post_date": "10/07/2020 04:44:46",
          "content": "<p>Thanks!<br>\nI have tabular only submission that scores private -6.8366 for gold, but CV using CT dicom files is much better than tabular only, so IMHO tabular only submission can get gold with some luck, but using CT dicom files leads to more solid result.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1040337,
          "author_name": "koalasheep",
          "author_url": "",
          "post_date": "10/07/2020 04:59:37",
          "content": "<p>Thank you for your reply. I am agree with your opinion. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1040307,
      "author_name": "sishihara",
      "author_url": "",
      "post_date": "10/07/2020 04:40:15",
      "content": "<p>Congrats for your solo gold 🎉 I'm really impressed with your strategy. Thank you for sharing.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1040316,
          "author_name": "yasufuminakama",
          "author_url": "",
          "post_date": "10/07/2020 04:46:41",
          "content": "<p>Thanks! :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1041952,
          "author_name": "sishihara",
          "author_url": "",
          "post_date": "10/08/2020 01:45:09",
          "content": "<p>Now, you became Grandmaster with the result of OpenVaccine. Congratulations again 🎉🎉🎉<br>\n<a href=\"https://www.kaggle.com/c/stanford-covid-vaccine\" target=\"_blank\">https://www.kaggle.com/c/stanford-covid-vaccine</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1041986,
          "author_name": "yasufuminakama",
          "author_url": "",
          "post_date": "10/08/2020 02:10:52",
          "content": "<p>Thanks again!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1040332,
      "author_name": "amrut11",
      "author_url": "",
      "post_date": "10/07/2020 04:57:29",
      "content": "<p>Congratulations!!!!!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1040435,
          "author_name": "yasufuminakama",
          "author_url": "",
          "post_date": "10/07/2020 06:17:28",
          "content": "<p>Thanks you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1040374,
      "author_name": "loto610",
      "author_url": "",
      "post_date": "10/07/2020 05:23:07",
      "content": "<p>Your solution is simple and cool.<br>\nThanks for your sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1040394,
      "author_name": "harutot",
      "author_url": "",
      "post_date": "10/07/2020 05:41:18",
      "content": "<p>Congratulations for your 6th place!<br>\nThanks for sharing great works and ideas. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1040543,
      "author_name": "alexj21",
      "author_url": "",
      "post_date": "10/07/2020 07:55:52",
      "content": "<p>Congratulations you deserve it ! thank you for your baseline kernels btw</p>",
      "votes": null,
      "replies": [
        {
          "id": 1040578,
          "author_name": "yasufuminakama",
          "author_url": "",
          "post_date": "10/07/2020 08:26:20",
          "content": "<p>Thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1040617,
      "author_name": "amedprof",
      "author_url": "",
      "post_date": "10/07/2020 09:01:34",
      "content": "<p>Congratulations , I wouldn't have achive gold medal without your kernels.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1040646,
          "author_name": "yasufuminakama",
          "author_url": "",
          "post_date": "10/07/2020 09:14:01",
          "content": "<p>Thank you! Glad to hear that, Congrats to your gold too :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1040655,
      "author_name": "resistance0108",
      "author_url": "",
      "post_date": "10/07/2020 09:18:47",
      "content": "<p>Congrats! What is your last3 cv? And did CV and private score correlate well?<br>\nMaybe data augmentation of tabular data is key to survive the shake up…</p>",
      "votes": null,
      "replies": [
        {
          "id": 1040671,
          "author_name": "yasufuminakama",
          "author_url": "",
          "post_date": "10/07/2020 09:29:40",
          "content": "<p>Thanks!</p>\n<p>I don't think we should monitor last3 cv because it's smaller to trust, so I haven't checked it.</p>\n<p>My recent submissions are as follows, seems CV and Private correlate well.</p>\n<ul>\n<li>CV: -6.58651 Public: -6.8921 Private: -6.8443</li>\n<li>CV: -6.54917 Public: -6.8922 Private: -6.8421</li>\n<li>CV: -6.53815 Public: -6.8972 Private: -6.8363</li>\n</ul>\n<p>Yes, I think my input data construction strategy is important.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1040726,
          "author_name": "resistance0108",
          "author_url": "",
          "post_date": "10/07/2020 10:07:03",
          "content": "<p>Oh OK. Last 3 CV is untrustworthy, then I was totally wrong from the start point…<br>\nIf I used your data augmentation strategy!!!!!<br>\nAnd you have high ensemble weight on Lasso, which is similar to me…<br>\nMy solution:<br>\n<a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/189341\" target=\"_blank\">https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/189341</a></p>\n<p>My inspiration that Lasso is effective was proven to be true…</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1040679,
      "author_name": "code1110",
      "author_url": "",
      "post_date": "10/07/2020 09:34:15",
      "content": "<p>Congratulations to your solo Gold!</p>\n<p>My best decision in this competition was to make my solution based on your public kernel.  Thanks for your contribution <br>\nto the community, and I'm looking forward to the day of you becoming the GM!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1040685,
          "author_name": "yasufuminakama",
          "author_url": "",
          "post_date": "10/07/2020 09:39:39",
          "content": "<p>Thanks you!<br>\nGlad to hear that, I'll do my best :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1040805,
      "author_name": "lgreig",
      "author_url": "",
      "post_date": "10/07/2020 11:05:12",
      "content": "<p>Well done on getting 6th place; I find your notebooks are very helpful! </p>\n<p>You mention your cv was -6.53815 - but is this for all data or just last 3 points? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1041015,
          "author_name": "yasufuminakama",
          "author_url": "",
          "post_date": "10/07/2020 13:55:24",
          "content": "<p>Thank you! Glad to hear that :)<br>\nIt's for all data, I don't think we should monitor last3 cv because it's smaller to trust.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1040866,
      "author_name": "doctorkael",
      "author_url": "",
      "post_date": "10/07/2020 12:12:36",
      "content": "<p>Your idea really helped me getting my first ever Silver medal on Kaggle:</p>\n<pre><code>Treating every single measurement as if it were a “baseline” measurement. And “Week” is used to create “Week_passed” for each measurements. We don’t know which “Week” is given in test data, so this input data construction gives us robustness to build models.\n</code></pre>\n<p>Using this idea I performed some more data augmentations and a GBR to create my predictions. Many congrats to your Solo Gold!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1041017,
          "author_name": "yasufuminakama",
          "author_url": "",
          "post_date": "10/07/2020 13:56:34",
          "content": "<p>Thank you &amp; Congrats to your effort and result!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1041025,
      "author_name": "samshipengs",
      "author_url": "",
      "post_date": "10/07/2020 14:03:12",
      "content": "<p>Congrats! How did you actually use the image data? From what I experimented using images did not help with the predictions much, I randomly sampled 2~3 images as channels around the middle of all the scans the patient has but the performance was the same/lower than using tabular data alone. And I tried out the public kernel which uses efficient, it just randomly samples from any scans (if i understand correctly), I changed the images size to smaller number (e.g. 128x128) or even commented out all the cnn branch, it didn't change the score (even improved the local score in that kernel a bit). So im really curious how you used images. </p>\n<p>It has a few issues IMHO: 1) what set of images to use? there are patients with many number of scans while some only has 14 scans (also pixel spacing is different, its probably arguable if its needed to scale it, plus slice thickness is also different). 2) the image position (i.e. head or feet first) affects the order of the lung slices (e.g. bottom of lung appear early or later), reorder is probably needed based on the imagePosition attribute, but I also observed at least a few imagePosition value is incorrect (i.e. head first images labelled as FFS etc). 3) there were a few patients in train doesnt have complete lung scans, and one patient has two lungs over all his/her scan, although this might be okay given its a small portion 4) the number of training data is small, i mean the unique number of patients is small and the slices that actually carry useful pattern is small (i.e. probably mainly the bottom of the lung)</p>\n<p>so i switched from using image directly to extract features, such as skews and kurtosis and total volumes, these features actually shows high correlation with FVC (at least in our train set…., e.g. lung_volume as ~0.84 Pearson correlation with FVC) and others has large magnitude coefficient in my bayesian ridge regression models just trailing behind base_FVC and base_Percent. But this didnt help us at all in the private test set…. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1041065,
          "author_name": "yasufuminakama",
          "author_url": "",
          "post_date": "10/07/2020 14:30:40",
          "content": "<p>Thanks!</p>\n<ul>\n<li>How did you actually use the image data?<ul>\n<li>Training is almost same method as public notebooks I wrote above. I simply used efficientnet-b0 oof predictions for each patient to train other models. Quantile value is fixed to 0.5 when I get predictions. <strong>Linear decay for weeks is not applied.</strong></li></ul></li>\n<li>From what I experimented using images did not help with the predictions much<ul>\n<li>I'm not sure about your CV, but in my experiments for example my single ridge improved CV around 0.15 by using image predictions.</li></ul></li>\n</ul>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1041080,
      "author_name": "yutoaoki",
      "author_url": "",
      "post_date": "10/07/2020 14:46:27",
      "content": "<p>Thank you for sharing your wonderful notebook, and congratulations on Solo Gold.<br>\nI added more data by doing the same thing with your notebook. But I made one change and improved the CV.</p>\n<p>In your notebook, <code>Week_passed</code> is roughly -60 to 60 range. Since the predictions are always after <code>base_Week</code>, I thought negative values weren't important, so I did the following.</p>\n<pre><code>train = output[(-12&lt;=output['Week_passed']) &amp; (output['Week_passed']! =0)].reset_index(drop=True)\n</code></pre>\n<p>A value of -12 gave me the highest CV.<br>\nMaybe I just accidentally got a higher CV, but try it out when you have the time!<br>\nAnyway, Thanks to you, I got my first medal! Thank you.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1041104,
          "author_name": "yasufuminakama",
          "author_url": "",
          "post_date": "10/07/2020 15:04:32",
          "content": "<p>I think the reason you got higher CV is that it's not same Cross-Validation before you apply <code>-12&lt;=output['Week_passed']</code>, you can't compare them.<br>\nThanks &amp; Congrats to your result!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1041206,
          "author_name": "yutoaoki",
          "author_url": "",
          "post_date": "10/07/2020 16:01:35",
          "content": "<p>Indeed. I was wrong.<br>\nI'm sorry for missing the point.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1046317,
          "author_name": "shenjiaxjtlu",
          "author_url": "",
          "post_date": "10/11/2020 14:45:43",
          "content": "<p>Hi, I don’t know whether I have missed something, but actually I have the similar idea to you.<br>\nI did combine the test tabular and the train tabular together and after that, since the test data is 730 rows and have many rows with week_passed less than zero, I delete all of them and then do the whole minmax scalar to both the train and the test. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1041309,
      "author_name": "rjmanoj",
      "author_url": "",
      "post_date": "10/07/2020 17:17:55",
      "content": "<p>Super… such a useful Information </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1041311,
      "author_name": "rjmanoj",
      "author_url": "",
      "post_date": "10/07/2020 17:18:20",
      "content": "<p>Thanks for sharing this</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1041769,
      "author_name": "hardikmirani",
      "author_url": "",
      "post_date": "10/07/2020 22:55:48",
      "content": "<p>Very Interesting </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1042612,
      "author_name": "yasufuminakama",
      "author_url": "",
      "post_date": "10/08/2020 11:06:28",
      "content": "<p>Updated CV and LB Transition, CV and Private correlate well :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1048386,
      "author_name": "emanueleguzzi",
      "author_url": "",
      "post_date": "10/13/2020 12:48:42",
      "content": "<p>Congratulations and thank for sharing your solution!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1048534,
      "author_name": "andreasstavrou",
      "author_url": "",
      "post_date": "10/13/2020 15:18:29",
      "content": "<p>Congratulations!! Thanks for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1048643,
      "author_name": "crapvv",
      "author_url": "",
      "post_date": "10/13/2020 17:13:56",
      "content": "<p>Congratulations and Thanks for sharing : ))</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1387413,
      "author_name": "mrinath",
      "author_url": "",
      "post_date": "07/14/2021 06:56:55",
      "content": "<p>For the image part, I am not understanding how you are not using a 3D conv network.<br>\nAs each datapoint has lots of images, and together they form a 3d data, so are you treating each channel of this 3D data as an image separately? and feeding it to a pretrained conv network?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1042095,
      "author_name": "deeprakesh",
      "author_url": "",
      "post_date": "10/08/2020 04:01:39",
      "content": "<p>Congratulations and thanks for sharing :) </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1040186": "I would like to thank kaggle & host for the interesting competition and to all the participants for giving me a lot of ideas. And congrats to winners!\n\nMy work is based on \n- My tabular approach notebooks\n  - https://www.kaggle.com/yasufuminakama/osic-lgb-baseline \n  - https://www.kaggle.com/yasufuminakama/osic-ridge-baseline \n- Image approach notebooks\n  - https://www.kaggle.com/miklgr500/linear-decay-based-on-resnet-cnn\n  - https://www.kaggle.com/khoongweihao/k-fold-tf-efficientnet-models-training\n\n# Solution Overview\n![OSIC-solution-overview](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1695531%2F1f5423ba0fc5e5bb3bc182c0b8a75b5a%2FOSIC-solution.png?generation=1602040115132884&alt=media)\n\n# How to construct input data\nAs I showed in [my notebook](https://www.kaggle.com/yasufuminakama/osic-lgb-baseline), I constructed input data by treating every single measurement as if it were a “baseline” measurement. And “Week” is used to create “Week_passed” for each measurements. We don’t know which “Week” is given in test data, so this input data construction gives us robustness to build models.\n\n# How to treat image data\nBased on image approach notebooks, I used efficientnet-b0 for image size 320x320. \nI tried other efficientnet models but there are not so much different results among them when I fix quantile value as 0.5, so I used the smallest one.\nThis efficientnet-b0 output is used for FVC & Confidence model training with tabular features.\n\n# FVC prediction\nFVC prediction is almost same as my notebooks, the difference is I prepared 5 models and blend them automatically using ```sp.optimize.minimize```, weight is [Lasso, Ridge, ElasticNet, SVM, NN] = [0.68573749, 0., 0., 0.07551167, 0.23750526].\n\n# Confidence prediction\nConfidence labels are made using FVC oof & ```sp.optimize.minimize``` as shown in my notebooks.\nConfidence prediction is almost same as my notebooks too, the difference is I prepared 5 models and blend them automatically using ```sp.optimize.minimize```, weight is [Lasso, Ridge, ElasticNet, SVM, NN] = [0.22062125, 0., 0., 0., 0.80819966].\n\n# About shake up\nThere are so many public notebooks which overfits Public LB.\nPublic LB has not so much data, so you don't need to care about Public LB so much.\n\n# CV and LB Transition\n|  Model |  CV  |  Public  | Private | Medal |\n| ---- | ---- | ---- | ---- | ---- |\n|  LGB | -6.85094 | -6.9605 | -7.0037 | None |\n|  Ridge | -6.73738 | -6.9357 | -6.8562 | Bronze |\n|  Efficientnet-b0 + Ridge | -6.58651 | -6.8921 | -6.8443 | Silver |\n|  Efficientnet-b0 + Blend models v1 | -6.54917 | -6.8922 | -6.8421 | Top Silver |\n|  Efficientnet-b0 + Blend models v2 | -6.53815 | -6.8972 | -6.8363 | Gold |\n\n# Final result \nI selected my best CV, and it was the best Private LB submission.\nCV (4fold): -6.53815\nPrivate LB: -6.8363",
    "1040188": "Congratulations @yasufuminakama on the solo gold!",
    "1040279": "Congratulations and great work @yasufuminakama . I am curious about the contribution from CT dicom files in your solution. Since some solutions get good score with tabular data only, like [https://www.kaggle.com/lhagiimn/solution-for-the-first-place-but-we-didn-t-select](url) and [https://www.kaggle.com/ulrich07/osic-basic-tabular-data-augmentation-nn?scriptVersionId=39788369](url).",
    "1040307": "Congrats for your solo gold 🎉 I'm really impressed with your strategy. Thank you for sharing.",
    "1040313": "Thanks!\nI have tabular only submission that scores private -6.8366 for gold, but CV using CT dicom files is much better than tabular only, so IMHO tabular only submission can get gold with some luck, but using CT dicom files leads to more solid result.",
    "1040316": "Thanks! :)",
    "1040317": "Thank you! :)",
    "1040332": "Congratulations!!!!!",
    "1040337": "Thank you for your reply. I am agree with your opinion.",
    "1040374": "Your solution is simple and cool.\nThanks for your sharing!",
    "1040394": "Congratulations for your 6th place!\nThanks for sharing great works and ideas.",
    "1040435": "Thanks you!",
    "1040543": "Congratulations you deserve it ! thank you for your baseline kernels btw",
    "1040578": "Thank you!",
    "1040617": "Congratulations , I wouldn't have achive gold medal without your kernels.",
    "1040646": "Thank you! Glad to hear that, Congrats to your gold too :)",
    "1040655": "Congrats! What is your last3 cv? And did CV and private score correlate well?\nMaybe data augmentation of tabular data is key to survive the shake up...",
    "1040671": "Thanks!\n\nI don't think we should monitor last3 cv because it's smaller to trust, so I haven't checked it.\n\nMy recent submissions are as follows, seems CV and Private correlate well.\n- CV: -6.58651 Public: -6.8921 Private: -6.8443\n- CV: -6.54917 Public: -6.8922 Private: -6.8421\n- CV: -6.53815 Public: -6.8972 Private: -6.8363\n\nYes, I think my input data construction strategy is important.",
    "1040679": "Congratulations to your solo Gold!\n\nMy best decision in this competition was to make my solution based on your public kernel.  Thanks for your contribution \nto the community, and I'm looking forward to the day of you becoming the GM!",
    "1040685": "Thanks you!\nGlad to hear that, I'll do my best :)",
    "1040726": "Oh OK. Last 3 CV is untrustworthy, then I was totally wrong from the start point...\nIf I used your data augmentation strategy!!!!!\nAnd you have high ensemble weight on Lasso, which is similar to me...\nMy solution:\nhttps://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/189341\n\nMy inspiration that Lasso is effective was proven to be true...",
    "1040805": "Well done on getting 6th place; I find your notebooks are very helpful! \n\nYou mention your cv was -6.53815 - but is this for all data or just last 3 points?",
    "1040866": "Your idea really helped me getting my first ever Silver medal on Kaggle:\n\n    Treating every single measurement as if it were a “baseline” measurement. And “Week” is used to create “Week_passed” for each measurements. We don’t know which “Week” is given in test data, so this input data construction gives us robustness to build models.\n\nUsing this idea I performed some more data augmentations and a GBR to create my predictions. Many congrats to your Solo Gold!",
    "1041015": "Thank you! Glad to hear that :)\nIt's for all data, I don't think we should monitor last3 cv because it's smaller to trust.",
    "1041017": "Thank you & Congrats to your effort and result!",
    "1041025": "Congrats! How did you actually use the image data? From what I experimented using images did not help with the predictions much, I randomly sampled 2~3 images as channels around the middle of all the scans the patient has but the performance was the same/lower than using tabular data alone. And I tried out the public kernel which uses efficient, it just randomly samples from any scans (if i understand correctly), I changed the images size to smaller number (e.g. 128x128) or even commented out all the cnn branch, it didn't change the score (even improved the local score in that kernel a bit). So im really curious how you used images. \n\nIt has a few issues IMHO: 1) what set of images to use? there are patients with many number of scans while some only has 14 scans (also pixel spacing is different, its probably arguable if its needed to scale it, plus slice thickness is also different). 2) the image position (i.e. head or feet first) affects the order of the lung slices (e.g. bottom of lung appear early or later), reorder is probably needed based on the imagePosition attribute, but I also observed at least a few imagePosition value is incorrect (i.e. head first images labelled as FFS etc). 3) there were a few patients in train doesnt have complete lung scans, and one patient has two lungs over all his/her scan, although this might be okay given its a small portion 4) the number of training data is small, i mean the unique number of patients is small and the slices that actually carry useful pattern is small (i.e. probably mainly the bottom of the lung)\n\nso i switched from using image directly to extract features, such as skews and kurtosis and total volumes, these features actually shows high correlation with FVC (at least in our train set...., e.g. lung_volume as ~0.84 Pearson correlation with FVC) and others has large magnitude coefficient in my bayesian ridge regression models just trailing behind base_FVC and base_Percent. But this didnt help us at all in the private test set....",
    "1041065": "Thanks!\n- How did you actually use the image data?\n  - Training is almost same method as public notebooks I wrote above. I simply used efficientnet-b0 oof predictions for each patient to train other models. Quantile value is fixed to 0.5 when I get predictions. **Linear decay for weeks is not applied.**\n- From what I experimented using images did not help with the predictions much\n  - I'm not sure about your CV, but in my experiments for example my single ridge improved CV around 0.15 by using image predictions.",
    "1041080": "Thank you for sharing your wonderful notebook, and congratulations on Solo Gold.\nI added more data by doing the same thing with your notebook. But I made one change and improved the CV.\n\nIn your notebook, `Week_passed` is roughly -60 to 60 range. Since the predictions are always after `base_Week`, I thought negative values weren't important, so I did the following.\n````\ntrain = output[(-12<=output['Week_passed']) & (output['Week_passed']! =0)].reset_index(drop=True)\n````\nA value of -12 gave me the highest CV.\nMaybe I just accidentally got a higher CV, but try it out when you have the time!\nAnyway, Thanks to you, I got my first medal! Thank you.",
    "1041104": "I think the reason you got higher CV is that it's not same Cross-Validation before you apply `-12<=output['Week_passed']`, you can't compare them.\nThanks & Congrats to your result!",
    "1041206": "Indeed. I was wrong.\nI'm sorry for missing the point.",
    "1041309": "Super... such a useful Information",
    "1041311": "Thanks for sharing this",
    "1041769": "Very Interesting",
    "1041952": "Now, you became Grandmaster with the result of OpenVaccine. Congratulations again 🎉🎉🎉\nhttps://www.kaggle.com/c/stanford-covid-vaccine",
    "1041986": "Thanks again!",
    "1042095": "Congratulations and thanks for sharing :)",
    "1042612": "Updated CV and LB Transition, CV and Private correlate well :)",
    "1046317": "Hi, I don’t know whether I have missed something, but actually I have the similar idea to you.\nI did combine the test tabular and the train tabular together and after that, since the test data is 730 rows and have many rows with week_passed less than zero, I delete all of them and then do the whole minmax scalar to both the train and the test.",
    "1048386": "Congratulations and thank for sharing your solution!",
    "1048534": "Congratulations!! Thanks for sharing",
    "1048643": "Congratulations and Thanks for sharing : ))",
    "1387413": "For the image part, I am not understanding how you are not using a 3D conv network.\nAs each datapoint has lots of images, and together they form a 3d data, so are you treating each channel of this 3D data as an image separately? and feeding it to a pretrained conv network?"
  },
  "source": "meta"
}