{
  "id": 189251,
  "title": "top 9 solution",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/writeups/risers-top-9-solution",
  "author_name": "",
  "post_date": "2020-10-07T15:41:18.223Z",
  "votes": 57,
  "comment_count": 49,
  "views": 0,
  "content": "<p>Congratulations to all participants and thanks to organizers for making this competition possible. Despite it didn't go as expected, I hope many people learned something new here. Also, I would like to express my gratitude to my teammates for working together with me on this challenge. Special thanks to <a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> for not giving up on this competition and motivating our team to keep going, without him we would not go that far. Though, getting the medal I have a little bit bitter feeling. I'll take this chance and describe some of our ideas: hopefully they could help organizers to fight fibrosis. </p>\n<h3>CT model: Concatenate Tile Pooling</h3>\n<p>Working on PANDA competition I have proposed <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169205\" target=\"_blank\">Concatenate Tile Pooling</a> method, which also appeared to be applicable quite well for CT scan data. The idea of this method is illustrated in image below. Instead of assigning labels, like FVC decay and confidence to each CT layer, which may be difficult to predict based on a single image, why not to assign it to all images together?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2Fde95b636611aa0e65432f4dcf6756269%2Ftile.png?generation=1602038549896215&amp;alt=media\" alt=\"\"><br>\n<strong>why not 3D conv?</strong></p>\n<ul>\n<li>CT scans have different spacing, which deteriorates performance of 3D models. It can be viewed as training and performing inference with random stretch of images along one dimension.</li>\n<li>Many accurate and well-optimized 2D pretrained models.</li>\n<li>Not possible to train a good 3D model from scratch for so limited data</li>\n</ul>\n<p><strong>Important details:</strong></p>\n<ul>\n<li>Since the provided data is very limited, just ~150 samples, train the model first on masked lungs only (masks are produced with <a href=\"https://github.com/JoHof/lungmask/\" target=\"_blank\">lungmask</a>) with following finetuning on the original images to be able to perform inference without lungmask (So I could do inference of a single 4 fold model within just 10-20 min at kaggle). Check the image below.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2F31a1d24eb0da1882deb565777c1281f2%2Fmasks.png?generation=1602039389832792&amp;alt=media\" alt=\"\"></li>\n<li>For CT models it's better to predict percent rather than FCV. Image augmentation, like zoom, as well as use only of a limited number of slides makes it difficult to compute the volume based on CT. Meanwhile, the model could more reliably say what is the percent of the entire volume is affected by fibrosis by looking just at several scans.</li>\n<li>Assume linear dependence of the FVC and confidence on the week and predict the slope and the initial value.</li>\n<li>LLL_loss (Laplace Log Likelihood) works but for convergence one needs to ensure that the output at the beginning of training is close to the gt FVC. I used the following: <code>FVC = V0*(0.01*a*(w-w0)/134 + b + 0.01*p0)</code>, <code>sigma = V0*softplus(c*w/134 + d)</code>, were a,b,c,d are model predictions, V0 is the full lung volume computed as V0 = 100*FVC/percent. So, initially the predicted FVCs for model before training are quite close to the expected FVC, and the loss converges nicely.  I have computed the loss based on all FVC measurements for a given patient (dropping the first one), so there was no need to evaluate gt slops explicitly. An alternative could be doing something similar to public kernels: model tries to predict FVC directly with using mloss, followed by model finetuning with LLL_loss (which is essentially the metric in this competition).</li>\n<li><a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169205\" target=\"_blank\">Tile cutout</a> - random selection of CT layers during training could be viewed as a cutout in 3d image.</li>\n<li>When work with the data try to plot it to see what is really going on. For a number of patients I saw incorrect offsets, so the level of signal must be changed. Also I clipped signal at [-1024,600]. Finally, some images are not square, so I needed to take a crop to prevent deformation of images when rescale.   </li>\n</ul>\n<p><strong>Additional details:</strong><br>\n64 layers x 256 x 256 setup<br>\nResNeXt50 and ResNet18 backbones<br>\nStandard augmentation like rotation, zoom, horizontal flip at the pretraining stage.<br>\nUnfortunately, this beautiful approach didn't work well giving only ~6.85/6.90 at private/public LB.</p>\n<h3>Tabular models</h3>\n<p>This is quite standard, and one can find many examples in public kernels. <strong>The important thing: always do proper validation and exclude all kinds of leaks.</strong> So percent must not be used as a feature unless it is set to be equal to the value at the first visit (and the first visit must be excluded from training and validation). Also, we have computed validation based on last 3 visits, similar to the competition metric. We have 2 kinds of models: based only on tabular and tabular+CT features (with some additional variations). <br>\nSome thoughts about <strong>percent madness</strong> people used in public kernels: It seems that by chance public LB contains mostly cases with low slope (FVC decay rate), so anything that artificially reduces slope could boost public LB. If one is training with percent feature and then set it to the constant, it is equivalent to pivoting the prediction in such a way that reduces the slope. </p>\n<h3>Final submission</h3>\n<p>Our final submission is an ensemble of CT models + TAB NN based models + <a href=\"https://www.kaggle.com/jeabat/osic-bayesian-ridge-regression\" target=\"_blank\">Bayesian Ridge Regression</a> based on maximizing CV. This submission got 6.8385/6.8884 at private/public LB. It is surprising for so large shake up: we have chosen our best private LB sub.</p>",
  "messages": [
    {
      "id": "1040287",
      "postDate": "10/07/2020 04:26:05",
      "content": "<p>Congratulations to all participants and thanks to organizers for making this competition possible. Despite it didn't go as expected, I hope many people learned something new here. Also, I would like to express my gratitude to my teammates for working together with me on this challenge. Special thanks to <a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> for not giving up on this competition and motivating our team to keep going, without him we would not go that far. Though, getting the medal I have a little bit bitter feeling. I'll take this chance and describe some of our ideas: hopefully they could help organizers to fight fibrosis. </p>\n<h3>CT model: Concatenate Tile Pooling</h3>\n<p>Working on PANDA competition I have proposed <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169205\" target=\"_blank\">Concatenate Tile Pooling</a> method, which also appeared to be applicable quite well for CT scan data. The idea of this method is illustrated in image below. Instead of assigning labels, like FVC decay and confidence to each CT layer, which may be difficult to predict based on a single image, why not to assign it to all images together?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2Fde95b636611aa0e65432f4dcf6756269%2Ftile.png?generation=1602038549896215&amp;alt=media\" alt=\"\"><br>\n<strong>why not 3D conv?</strong></p>\n<ul>\n<li>CT scans have different spacing, which deteriorates performance of 3D models. It can be viewed as training and performing inference with random stretch of images along one dimension.</li>\n<li>Many accurate and well-optimized 2D pretrained models.</li>\n<li>Not possible to train a good 3D model from scratch for so limited data</li>\n</ul>\n<p><strong>Important details:</strong></p>\n<ul>\n<li>Since the provided data is very limited, just ~150 samples, train the model first on masked lungs only (masks are produced with <a href=\"https://github.com/JoHof/lungmask/\" target=\"_blank\">lungmask</a>) with following finetuning on the original images to be able to perform inference without lungmask (So I could do inference of a single 4 fold model within just 10-20 min at kaggle). Check the image below.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2F31a1d24eb0da1882deb565777c1281f2%2Fmasks.png?generation=1602039389832792&amp;alt=media\" alt=\"\"></li>\n<li>For CT models it's better to predict percent rather than FCV. Image augmentation, like zoom, as well as use only of a limited number of slides makes it difficult to compute the volume based on CT. Meanwhile, the model could more reliably say what is the percent of the entire volume is affected by fibrosis by looking just at several scans.</li>\n<li>Assume linear dependence of the FVC and confidence on the week and predict the slope and the initial value.</li>\n<li>LLL_loss (Laplace Log Likelihood) works but for convergence one needs to ensure that the output at the beginning of training is close to the gt FVC. I used the following: <code>FVC = V0*(0.01*a*(w-w0)/134 + b + 0.01*p0)</code>, <code>sigma = V0*softplus(c*w/134 + d)</code>, were a,b,c,d are model predictions, V0 is the full lung volume computed as V0 = 100*FVC/percent. So, initially the predicted FVCs for model before training are quite close to the expected FVC, and the loss converges nicely.  I have computed the loss based on all FVC measurements for a given patient (dropping the first one), so there was no need to evaluate gt slops explicitly. An alternative could be doing something similar to public kernels: model tries to predict FVC directly with using mloss, followed by model finetuning with LLL_loss (which is essentially the metric in this competition).</li>\n<li><a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169205\" target=\"_blank\">Tile cutout</a> - random selection of CT layers during training could be viewed as a cutout in 3d image.</li>\n<li>When work with the data try to plot it to see what is really going on. For a number of patients I saw incorrect offsets, so the level of signal must be changed. Also I clipped signal at [-1024,600]. Finally, some images are not square, so I needed to take a crop to prevent deformation of images when rescale.   </li>\n</ul>\n<p><strong>Additional details:</strong><br>\n64 layers x 256 x 256 setup<br>\nResNeXt50 and ResNet18 backbones<br>\nStandard augmentation like rotation, zoom, horizontal flip at the pretraining stage.<br>\nUnfortunately, this beautiful approach didn't work well giving only ~6.85/6.90 at private/public LB.</p>\n<h3>Tabular models</h3>\n<p>This is quite standard, and one can find many examples in public kernels. <strong>The important thing: always do proper validation and exclude all kinds of leaks.</strong> So percent must not be used as a feature unless it is set to be equal to the value at the first visit (and the first visit must be excluded from training and validation). Also, we have computed validation based on last 3 visits, similar to the competition metric. We have 2 kinds of models: based only on tabular and tabular+CT features (with some additional variations). <br>\nSome thoughts about <strong>percent madness</strong> people used in public kernels: It seems that by chance public LB contains mostly cases with low slope (FVC decay rate), so anything that artificially reduces slope could boost public LB. If one is training with percent feature and then set it to the constant, it is equivalent to pivoting the prediction in such a way that reduces the slope. </p>\n<h3>Final submission</h3>\n<p>Our final submission is an ensemble of CT models + TAB NN based models + <a href=\"https://www.kaggle.com/jeabat/osic-bayesian-ridge-regression\" target=\"_blank\">Bayesian Ridge Regression</a> based on maximizing CV. This submission got 6.8385/6.8884 at private/public LB. It is surprising for so large shake up: we have chosen our best private LB sub.</p>",
      "rawMarkdown": "Congratulations to all participants and thanks to organizers for making this competition possible. Despite it didn't go as expected, I hope many people learned something new here. Also, I would like to express my gratitude to my teammates for working together with me on this challenge. Special thanks to @jaideepvalani for not giving up on this competition and motivating our team to keep going, without him we would not go that far. Though, getting the medal I have a little bit bitter feeling. I'll take this chance and describe some of our ideas: hopefully they could help organizers to fight fibrosis. \n\n### CT model: Concatenate Tile Pooling\nWorking on PANDA competition I have proposed [Concatenate Tile Pooling](https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169205) method, which also appeared to be applicable quite well for CT scan data. The idea of this method is illustrated in image below. Instead of assigning labels, like FVC decay and confidence to each CT layer, which may be difficult to predict based on a single image, why not to assign it to all images together?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2Fde95b636611aa0e65432f4dcf6756269%2Ftile.png?generation=1602038549896215&alt=media)\n**why not 3D conv?**\n- CT scans have different spacing, which deteriorates performance of 3D models. It can be viewed as training and performing inference with random stretch of images along one dimension.\n- Many accurate and well-optimized 2D pretrained models.\n- Not possible to train a good 3D model from scratch for so limited data\n\n**Important details:**\n- Since the provided data is very limited, just ~150 samples, train the model first on masked lungs only (masks are produced with [lungmask](https://github.com/JoHof/lungmask/)) with following finetuning on the original images to be able to perform inference without lungmask (So I could do inference of a single 4 fold model within just 10-20 min at kaggle). Check the image below.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2F31a1d24eb0da1882deb565777c1281f2%2Fmasks.png?generation=1602039389832792&alt=media)\n- For CT models it's better to predict percent rather than FCV. Image augmentation, like zoom, as well as use only of a limited number of slides makes it difficult to compute the volume based on CT. Meanwhile, the model could more reliably say what is the percent of the entire volume is affected by fibrosis by looking just at several scans.\n- Assume linear dependence of the FVC and confidence on the week and predict the slope and the initial value.\n- LLL_loss (Laplace Log Likelihood) works but for convergence one needs to ensure that the output at the beginning of training is close to the gt FVC. I used the following: `FVC = V0*(0.01*a*(w-w0)/134 + b + 0.01*p0)`, `sigma = V0*softplus(c*w/134 + d)`, were a,b,c,d are model predictions, V0 is the full lung volume computed as V0 = 100*FVC/percent. So, initially the predicted FVCs for model before training are quite close to the expected FVC, and the loss converges nicely.  I have computed the loss based on all FVC measurements for a given patient (dropping the first one), so there was no need to evaluate gt slops explicitly. An alternative could be doing something similar to public kernels: model tries to predict FVC directly with using mloss, followed by model finetuning with LLL_loss (which is essentially the metric in this competition).\n- [Tile cutout](https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169205) - random selection of CT layers during training could be viewed as a cutout in 3d image.\n- When work with the data try to plot it to see what is really going on. For a number of patients I saw incorrect offsets, so the level of signal must be changed. Also I clipped signal at [-1024,600]. Finally, some images are not square, so I needed to take a crop to prevent deformation of images when rescale.   \n\n**Additional details:**\n64 layers x 256 x 256 setup\nResNeXt50 and ResNet18 backbones\nStandard augmentation like rotation, zoom, horizontal flip at the pretraining stage.\nUnfortunately, this beautiful approach didn't work well giving only ~6.85/6.90 at private/public LB.\n\n### Tabular models\nThis is quite standard, and one can find many examples in public kernels. **The important thing: always do proper validation and exclude all kinds of leaks.** So percent must not be used as a feature unless it is set to be equal to the value at the first visit (and the first visit must be excluded from training and validation). Also, we have computed validation based on last 3 visits, similar to the competition metric. We have 2 kinds of models: based only on tabular and tabular+CT features (with some additional variations). \nSome thoughts about **percent madness** people used in public kernels: It seems that by chance public LB contains mostly cases with low slope (FVC decay rate), so anything that artificially reduces slope could boost public LB. If one is training with percent feature and then set it to the constant, it is equivalent to pivoting the prediction in such a way that reduces the slope. \n\n### Final submission\nOur final submission is an ensemble of CT models + TAB NN based models + [Bayesian Ridge Regression](https://www.kaggle.com/jeabat/osic-bayesian-ridge-regression) based on maximizing CV. This submission got 6.8385/6.8884 at private/public LB. It is surprising for so large shake up: we have chosen our best private LB sub.",
      "votes": null
    },
    {
      "id": "1040301",
      "postDate": "10/07/2020 04:34:30",
      "content": "<p>Wow, your solution is really good! We never thought of using Concatenate Tile Pooling in our model :(</p>\n<p>Anyways, congratulations on the gold! <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> <a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> <a href=\"https://www.kaggle.com/robikscube\" target=\"_blank\">@robikscube</a> <a href=\"https://www.kaggle.com/poteman\" target=\"_blank\">@poteman</a> <a href=\"https://www.kaggle.com/mathurinache\" target=\"_blank\">@mathurinache</a> </p>",
      "rawMarkdown": "Wow, your solution is really good! We never thought of using Concatenate Tile Pooling in our model :(\n\nAnyways, congratulations on the gold! @iafoss @jaideepvalani @robikscube @poteman @mathurinache",
      "votes": null
    },
    {
      "id": "1040306",
      "postDate": "10/07/2020 04:39:09",
      "content": "<p>Great approach. The only gold winning team I could see with similar public and private LB score. Congrats <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> 😀</p>",
      "rawMarkdown": "Great approach. The only gold winning team I could see with similar public and private LB score. Congrats @iafoss 😀",
      "votes": null
    },
    {
      "id": "1040308",
      "postDate": "10/07/2020 04:40:41",
      "content": "<p>Thank you so much. Actually <a href=\"https://www.kaggle.com/robikscube\" target=\"_blank\">@robikscube</a> today won 2 gold medals and become a 3xGM.</p>",
      "rawMarkdown": "Thank you so much. Actually @robikscube today won 2 gold medals and become a 3xGM.",
      "votes": null
    },
    {
      "id": "1040311",
      "postDate": "10/07/2020 04:43:25",
      "content": "<p><a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> Yeah, I know that :) Congrats Rob!</p>\n<p><a href=\"https://www.kaggle.com/robikscube\" target=\"_blank\">@robikscube</a> Time to post a dataset now ;)</p>",
      "rawMarkdown": "iafoss Yeah, I know that :) Congrats Rob!\n\n@robikscube Time to post a dataset now ;)",
      "votes": null
    },
    {
      "id": "1040314",
      "postDate": "10/07/2020 04:45:00",
      "content": "<p>Thank you so much, <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> . This competition was quite difficult for all. </p>",
      "rawMarkdown": "Thank you so much, @nischaydnk . This competition was quite difficult for all.",
      "votes": null
    },
    {
      "id": "1040318",
      "postDate": "10/07/2020 04:47:33",
      "content": "<p>Congratulations!!!!!!!</p>",
      "rawMarkdown": "Congratulations!!!!!!!",
      "votes": null
    },
    {
      "id": "1040325",
      "postDate": "10/07/2020 04:53:29",
      "content": "<p>Thanks       </p>",
      "rawMarkdown": "Thanks",
      "votes": null
    },
    {
      "id": "1040329",
      "postDate": "10/07/2020 04:55:15",
      "content": "<p>Congrats ! You are the only team survived the shake up.</p>",
      "rawMarkdown": "Congrats ! You are the only team survived the shake up.",
      "votes": null
    },
    {
      "id": "1040339",
      "postDate": "10/07/2020 05:00:44",
      "content": "<p>Thank you. It was indeed quite difficult.</p>",
      "rawMarkdown": "Thank you. It was indeed quite difficult.",
      "votes": null
    },
    {
      "id": "1040357",
      "postDate": "10/07/2020 05:13:33",
      "content": "<p>Comprehensive work. Congratulations on your medal.</p>",
      "rawMarkdown": "Comprehensive work. Congratulations on your medal.",
      "votes": null
    },
    {
      "id": "1040372",
      "postDate": "10/07/2020 05:21:48",
      "content": "<p>Thank you  </p>",
      "rawMarkdown": "Thank you",
      "votes": null
    },
    {
      "id": "1040405",
      "postDate": "10/07/2020 05:55:13",
      "content": "<p>Congrats, truly amazing efforts. Definitely skill, not luck !</p>",
      "rawMarkdown": "Congrats, truly amazing efforts. Definitely skill, not luck !",
      "votes": null
    },
    {
      "id": "1040414",
      "postDate": "10/07/2020 06:02:50",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!",
      "votes": null
    },
    {
      "id": "1040431",
      "postDate": "10/07/2020 06:14:32",
      "content": "<p><a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a>  it was your presence that mattered most towards  the ending days in addition to team effort which was very much needed to keep moving in competition when many participants  already might have left race  for obvious reasons . <br>\nYou are person deserving to be in top 3 ,matter  of just some more time . <br>\nI will spend time now in carefully going through detailed approach. </p>",
      "rawMarkdown": "iafoss  it was your presence that mattered most towards  the ending days in addition to team effort which was very much needed to keep moving in competition when many participants  already might have left race  for obvious reasons . \nYou are person deserving to be in top 3 ,matter  of just some more time . \nI will spend time now in carefully going through detailed approach.",
      "votes": null
    },
    {
      "id": "1040490",
      "postDate": "10/07/2020 07:03:11",
      "content": "<p>Some highlight on Tab features that we used in tab model for ensemble </p>\n<p>1)  Usage of Volume and Skew features extracted from CT using   segmentation model in this kernel . .<br>\n<a href=\"https://www.kaggle.com/hfutybx/osic-feature-extract-from-ct\" target=\"_blank\">https://www.kaggle.com/hfutybx/osic-feature-extract-from-ct</a> .Standalone Tab model on the top of public kernel using these feature score quite well   on a public lb . Thanks to <a href=\"https://www.kaggle.com/sandorkonya\" target=\"_blank\">@sandorkonya</a> for his insightful post highlighting the usage and importance of it . (Mean,kurtosis dint help ,it worsened the fold cv score )</p>\n<p>2) other features   added in tab model are below, where  bins are quantiles  for age ,first week fvc percent <br>\n<code>['Age_bins','min_pct_bin','min_Percent','week','BASE' ,'Volume','Skew' ].</code><br>\nadding quantiles improved cv  and lb earlier </p>\n<p>3) mloss tuning  using 60 pct of quantile and 40 pct of  lll loss ,max loss of quantile to 800.</p>\n<p>I dint pay much attention over  CV splits  for my tab model so was not relying on CV score for private lb  ,just was checking  how right the approach used, based on lb score  between 6.85 to 6.90 ,any thing less than this was almost impossible  unless using Percent feature which was tightly coupled with FVC  ,fvc=k*percent, where k is constant for Gender,Smooking,Age group. hence causing leak of target to model.</p>\n<p>last thing that i wanted to try but due to lack of time i couldnt try<br>\n  Sampling of batches such  that all patients are in same batch  as its time series data </p>",
      "rawMarkdown": "Some highlight on Tab features that we used in tab model for ensemble \n\n1)  Usage of Volume and Skew features extracted from CT using   segmentation model in this kernel . .\nhttps://www.kaggle.com/hfutybx/osic-feature-extract-from-ct .Standalone Tab model on the top of public kernel using these feature score quite well   on a public lb . Thanks to @sandorkonya for his insightful post highlighting the usage and importance of it . (Mean,kurtosis dint help ,it worsened the fold cv score )\n  \n2) other features   added in tab model are below, where  bins are quantiles  for age ,first week fvc percent \n`['Age_bins','min_pct_bin','min_Percent','week','BASE' ,'Volume','Skew' ].`\nadding quantiles improved cv  and lb earlier \n\n3) mloss tuning  using 60 pct of quantile and 40 pct of  lll loss ,max loss of quantile to 800.\n\nI dint pay much attention over  CV splits  for my tab model so was not relying on CV score for private lb  ,just was checking  how right the approach used, based on lb score  between 6.85 to 6.90 ,any thing less than this was almost impossible  unless using Percent feature which was tightly coupled with FVC  ,fvc=k*percent, where k is constant for Gender,Smooking,Age group. hence causing leak of target to model.\n\nlast thing that i wanted to try but due to lack of time i couldnt try\n  Sampling of batches such  that all patients are in same batch  as its time series data",
      "votes": null
    },
    {
      "id": "1040497",
      "postDate": "10/07/2020 07:07:33",
      "content": "<p>Congrats to you and team on results.Thanks for sharing solution <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> </p>",
      "rawMarkdown": "Congrats to you and team on results.Thanks for sharing solution @iafoss",
      "votes": null
    },
    {
      "id": "1040534",
      "postDate": "10/07/2020 07:48:36",
      "content": "<p>Congrats to you and the team for surviving shakeup! Concat tile pooling together with the finetuning are very clever ideas</p>",
      "rawMarkdown": "Congrats to you and the team for surviving shakeup! Concat tile pooling together with the finetuning are very clever ideas",
      "votes": null
    },
    {
      "id": "1040536",
      "postDate": "10/07/2020 07:52:47",
      "content": "<p>Really nice job, very interesting approach</p>",
      "rawMarkdown": "Really nice job, very interesting approach",
      "votes": null
    },
    {
      "id": "1040539",
      "postDate": "10/07/2020 07:54:01",
      "content": "<p>Congratulations to Risers team. Great work!</p>",
      "rawMarkdown": "Congratulations to Risers team. Great work!",
      "votes": null
    },
    {
      "id": "1040541",
      "postDate": "10/07/2020 07:55:13",
      "content": "<p>Congratulations you deserve it! <br>\nI used your tiling method for PANDA challenge and thought about using it here.. it's interesting to know that it works ! </p>",
      "rawMarkdown": "Congratulations you deserve it! \nI used your tiling method for PANDA challenge and thought about using it here.. it's interesting to know that it works !",
      "votes": null
    },
    {
      "id": "1040676",
      "postDate": "10/07/2020 09:31:42",
      "content": "<p>Congratulations on the strong finish. From what I see you're (almost) the only team in the top 50 that didn't jump &gt;500~places, which is really impressive.</p>",
      "rawMarkdown": "Congratulations on the strong finish. From what I see you're (almost) the only team in the top 50 that didn't jump >500~places, which is really impressive.",
      "votes": null
    },
    {
      "id": "1041019",
      "postDate": "10/07/2020 13:57:36",
      "content": "<p>Congrats to you and your team for the solid solution!! Y'all put in so much work and deserved it!</p>",
      "rawMarkdown": "Congrats to you and your team for the solid solution!! Y'all put in so much work and deserved it!",
      "votes": null
    },
    {
      "id": "1041122",
      "postDate": "10/07/2020 15:16:14",
      "content": "<p>Thank you. I joined the competition because I got this idea. Unfortunately, it worked not completely as expected, but I think it's the best way to make use of CT.</p>",
      "rawMarkdown": "Thank you. I joined the competition because I got this idea. Unfortunately, it worked not completely as expected, but I think it's the best way to make use of CT.",
      "votes": null
    },
    {
      "id": "1041125",
      "postDate": "10/07/2020 15:16:52",
      "content": "<p>Thank you     </p>",
      "rawMarkdown": "Thank you",
      "votes": null
    },
    {
      "id": "1041132",
      "postDate": "10/07/2020 15:20:03",
      "content": "<p>Thank you so much</p>",
      "rawMarkdown": "Thank you so much",
      "votes": null
    },
    {
      "id": "1041156",
      "postDate": "10/07/2020 15:28:30",
      "content": "<p><a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> , Thank you so much. Public LB was was not very representative because many participants were exploiting  percent feature: train with FVC percent given at each week and do inference with percent only from the first week, which by coincidence appeared to work really well at public LB, but not making much sense.</p>",
      "rawMarkdown": "theoviel , Thank you so much. Public LB was was not very representative because many participants were exploiting  percent feature: train with FVC percent given at each week and do inference with percent only from the first week, which by coincidence appeared to work really well at public LB, but not making much sense.",
      "votes": null
    },
    {
      "id": "1041157",
      "postDate": "10/07/2020 15:28:58",
      "content": "<p>Thank you so much.</p>",
      "rawMarkdown": "Thank you so much.",
      "votes": null
    },
    {
      "id": "1041158",
      "postDate": "10/07/2020 15:29:20",
      "content": "<p>Thanks         </p>",
      "rawMarkdown": "Thanks",
      "votes": null
    },
    {
      "id": "1041161",
      "postDate": "10/07/2020 15:29:56",
      "content": "<p>Thanks, and you are welcome.</p>",
      "rawMarkdown": "Thanks, and you are welcome.",
      "votes": null
    },
    {
      "id": "1041182",
      "postDate": "10/07/2020 15:40:55",
      "content": "<p>Thank you so much. </p>",
      "rawMarkdown": "Thank you so much.",
      "votes": null
    },
    {
      "id": "1041193",
      "postDate": "10/07/2020 15:48:34",
      "content": "<p>Thank you so much, I really like those ideas too. I hoped that they would work better, but fortunately they made a strong contribution to our ensemble.</p>",
      "rawMarkdown": "Thank you so much, I really like those ideas too. I hoped that they would work better, but fortunately they made a strong contribution to our ensemble.",
      "votes": null
    },
    {
      "id": "1041449",
      "postDate": "10/07/2020 18:42:15",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> , <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> and the rest of the team!</p>\n<p>Sadly, we did not survive the shakeup,<br>\ni'll share our segmentation and image features soon.</p>",
      "rawMarkdown": "Congrats @jaideepvalani , @iafoss and the rest of the team!\n\nSadly, we did not survive the shakeup,\ni'll share our segmentation and image features soon.",
      "votes": null
    },
    {
      "id": "1041871",
      "postDate": "10/08/2020 00:31:09",
      "content": "<p>Good Job!</p>",
      "rawMarkdown": "Good Job!",
      "votes": null
    },
    {
      "id": "1041940",
      "postDate": "10/08/2020 01:41:00",
      "content": "<p>Thanks      </p>",
      "rawMarkdown": "Thanks",
      "votes": null
    },
    {
      "id": "1041946",
      "postDate": "10/08/2020 01:42:43",
      "content": "<p>Congratulations and thank you for sharing. I'm really impressed with your deep understanding and use of CT scan data.</p>",
      "rawMarkdown": "Congratulations and thank you for sharing. I'm really impressed with your deep understanding and use of CT scan data.",
      "votes": null
    },
    {
      "id": "1041962",
      "postDate": "10/08/2020 01:52:30",
      "content": "<p>Thanks, I really appreciate your words.</p>",
      "rawMarkdown": "Thanks, I really appreciate your words.",
      "votes": null
    },
    {
      "id": "1041965",
      "postDate": "10/08/2020 01:57:54",
      "content": "<p>Amazing Ideas <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> . I've followed this insights of yours in the PANDA Competition but didn't think to use it in this one. Congratulations on the well deserved gold.</p>",
      "rawMarkdown": "Amazing Ideas @iafoss . I've followed this insights of yours in the PANDA Competition but didn't think to use it in this one. Congratulations on the well deserved gold.",
      "votes": null
    },
    {
      "id": "1041967",
      "postDate": "10/08/2020 02:00:59",
      "content": "<p>Thank you so much <a href=\"https://www.kaggle.com/sandorkonya\" target=\"_blank\">@sandorkonya</a> , I saw your hand-labeled dataset posted, really impressive work and the amount of efforts you and your team contributed. I'm so sorry that you didn't win a reward proportional to your contribution.</p>",
      "rawMarkdown": "Thank you so much @sandorkonya , I saw your hand-labeled dataset posted, really impressive work and the amount of efforts you and your team contributed. I'm so sorry that you didn't win a reward proportional to your contribution.",
      "votes": null
    },
    {
      "id": "1042014",
      "postDate": "10/08/2020 02:42:38",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/ronaldokun\" target=\"_blank\">@ronaldokun</a></p>",
      "rawMarkdown": "Thank you @ronaldokun",
      "votes": null
    },
    {
      "id": "1043149",
      "postDate": "10/08/2020 18:11:26",
      "content": "<p>Congrats on another gold medal! You really deserve it!</p>",
      "rawMarkdown": "Congrats on another gold medal! You really deserve it!",
      "votes": null
    },
    {
      "id": "1043198",
      "postDate": "10/08/2020 18:54:54",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> , u2 congratulations with solo gold in Vaccine competition, well done. Also, nice presentation at the workshop today.</p>",
      "rawMarkdown": "Thanks @shujun717 , u2 congratulations with solo gold in Vaccine competition, well done. Also, nice presentation at the workshop today.",
      "votes": null
    },
    {
      "id": "1043301",
      "postDate": "10/08/2020 21:32:29",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a>. I enjoyed your presentation as well. The OpenVaccine competition turned out to be lucky for me in the end. They had to rescore on the last day (Monday), and because of that they delayed the deadline and hence I was able to get more models out during my ensembling. Also I believed in the idea of median avg which would have landed my team a gold in PANDA</p>\n<p>As for the attention masks I showed, they do look concerning. I can't speak for the ISUP grade attention weight distribution, but the ones for majority and minority gleason scores are definitely wrong all the time since the predictions aren't even close even when a lot of times ISUP grades were predicted correctly. But this is the nice thing about using something like self-attention: it is somewhat interpretible and we can make improvements logically based on what we can see from the attention weights.</p>\n<p>As for your tile method, the competition would have been so much more difficult without it, so i think it's still good you shared it. And congrats on best innovative method! That was really well deserved</p>",
      "rawMarkdown": "Thanks @iafoss. I enjoyed your presentation as well. The OpenVaccine competition turned out to be lucky for me in the end. They had to rescore on the last day (Monday), and because of that they delayed the deadline and hence I was able to get more models out during my ensembling. Also I believed in the idea of median avg which would have landed my team a gold in PANDA\n\nAs for the attention masks I showed, they do look concerning. I can't speak for the ISUP grade attention weight distribution, but the ones for majority and minority gleason scores are definitely wrong all the time since the predictions aren't even close even when a lot of times ISUP grades were predicted correctly. But this is the nice thing about using something like self-attention: it is somewhat interpretible and we can make improvements logically based on what we can see from the attention weights.\n\nAs for your tile method, the competition would have been so much more difficult without it, so i think it's still good you shared it. And congrats on best innovative method! That was really well deserved",
      "votes": null
    },
    {
      "id": "1043412",
      "postDate": "10/09/2020 00:53:07",
      "content": "<p>Thank you so much, congratulations to u too for running up reward, attention maps was really good idea. But I'm really worrying about the model interpretability at this point( Hopefully organizers will figure it out how to get all out of the method, especially since it was not that bad at the external data. </p>",
      "rawMarkdown": "Thank you so much, congratulations to u too for running up reward, attention maps was really good idea. But I'm really worrying about the model interpretability at this point( Hopefully organizers will figure it out how to get all out of the method, especially since it was not that bad at the external data.",
      "votes": null
    },
    {
      "id": "1043714",
      "postDate": "10/09/2020 07:28:18",
      "content": "<p>Very interesting approach. Thanks for sharing.</p>",
      "rawMarkdown": "Very interesting approach. Thanks for sharing.",
      "votes": null
    },
    {
      "id": "1043729",
      "postDate": "10/09/2020 07:50:29",
      "content": "<p><a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> <br>\nit was really difficult choice to make for private lb submission  based on public lb score thanks to  <a href=\"https://www.kaggle.com/Iafoss\" target=\"_blank\">@Iafoss</a> for his acumen  towards closing hours . <br>\n to give you an eg o   . one of my late submission scored yesterday on public lb   6.8827 , now just imagine how competitive does this score looks on a public lb  but it got only 6.85 on private lb ,while   one of the kernel made public    had  public lb score  of 6.8810  but private lb score 6.8347 .</p>",
      "rawMarkdown": "gunesevitan \nit was really difficult choice to make for private lb submission  based on public lb score thanks to  @Iafoss for his acumen  towards closing hours . \n\n  to give you an eg o   . one of my late submission scored yesterday on public lb   6.8827 , now just imagine how competitive does this score looks on a public lb  but it got only 6.85 on private lb ,while   one of the kernel made public    had  public lb score  of 6.8810  but private lb score 6.8347 .",
      "votes": null
    },
    {
      "id": "1043745",
      "postDate": "10/09/2020 08:03:36",
      "content": "<p>I agree external data performance was not bad at all! I'm looking forward to how the organizers will organize the results from the competition in the paper</p>",
      "rawMarkdown": "I agree external data performance was not bad at all! I'm looking forward to how the organizers will organize the results from the competition in the paper",
      "votes": null
    },
    {
      "id": "1043771",
      "postDate": "10/09/2020 08:28:11",
      "content": "<p><a href=\"https://www.kaggle.com/sandorkonya\" target=\"_blank\">@sandorkonya</a>  would you help understanding   <a href=\"https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection\" target=\"_blank\">https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection</a><br>\nbasically the labels of it.  and why do we have so much classification.</p>",
      "rawMarkdown": "sandorkonya  would you help understanding   https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection\nbasically the labels of it.  and why do we have so much classification.",
      "votes": null
    },
    {
      "id": "1043971",
      "postDate": "10/09/2020 11:31:29",
      "content": "<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> ,<br>\n<a href=\"https://www.kaggle.com/redwankarimsony\" target=\"_blank\">@redwankarimsony</a> already did a <a href=\"https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/183850\" target=\"_blank\">very nice summary</a> about it!</p>",
      "rawMarkdown": "jaideepvalani ,\n@redwankarimsony already did a [very nice summary](https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/183850) about it!",
      "votes": null
    },
    {
      "id": "1044161",
      "postDate": "10/09/2020 15:03:27",
      "content": "<p>You are welcome</p>",
      "rawMarkdown": "You are welcome",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1040301,
      "author_name": "aadhavvignesh",
      "author_url": "",
      "post_date": "10/07/2020 04:34:30",
      "content": "<p>Wow, your solution is really good! We never thought of using Concatenate Tile Pooling in our model :(</p>\n<p>Anyways, congratulations on the gold! <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> <a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> <a href=\"https://www.kaggle.com/robikscube\" target=\"_blank\">@robikscube</a> <a href=\"https://www.kaggle.com/poteman\" target=\"_blank\">@poteman</a> <a href=\"https://www.kaggle.com/mathurinache\" target=\"_blank\">@mathurinache</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 1040308,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "10/07/2020 04:40:41",
          "content": "<p>Thank you so much. Actually <a href=\"https://www.kaggle.com/robikscube\" target=\"_blank\">@robikscube</a> today won 2 gold medals and become a 3xGM.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1040311,
          "author_name": "aadhavvignesh",
          "author_url": "",
          "post_date": "10/07/2020 04:43:25",
          "content": "<p><a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> Yeah, I know that :) Congrats Rob!</p>\n<p><a href=\"https://www.kaggle.com/robikscube\" target=\"_blank\">@robikscube</a> Time to post a dataset now ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1040306,
      "author_name": "nischaydnk",
      "author_url": "",
      "post_date": "10/07/2020 04:39:09",
      "content": "<p>Great approach. The only gold winning team I could see with similar public and private LB score. Congrats <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> 😀</p>",
      "votes": null,
      "replies": [
        {
          "id": 1040314,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "10/07/2020 04:45:00",
          "content": "<p>Thank you so much, <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> . This competition was quite difficult for all. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1040318,
      "author_name": "amrut11",
      "author_url": "",
      "post_date": "10/07/2020 04:47:33",
      "content": "<p>Congratulations!!!!!!!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1040325,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "10/07/2020 04:53:29",
          "content": "<p>Thanks       </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1040329,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "10/07/2020 04:55:15",
      "content": "<p>Congrats ! You are the only team survived the shake up.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1040339,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "10/07/2020 05:00:44",
          "content": "<p>Thank you. It was indeed quite difficult.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1043729,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "10/09/2020 07:50:29",
          "content": "<p><a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> <br>\nit was really difficult choice to make for private lb submission  based on public lb score thanks to  <a href=\"https://www.kaggle.com/Iafoss\" target=\"_blank\">@Iafoss</a> for his acumen  towards closing hours . <br>\n to give you an eg o   . one of my late submission scored yesterday on public lb   6.8827 , now just imagine how competitive does this score looks on a public lb  but it got only 6.85 on private lb ,while   one of the kernel made public    had  public lb score  of 6.8810  but private lb score 6.8347 .</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1040357,
      "author_name": "imoore",
      "author_url": "",
      "post_date": "10/07/2020 05:13:33",
      "content": "<p>Comprehensive work. Congratulations on your medal.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1040372,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "10/07/2020 05:21:48",
          "content": "<p>Thank you  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1040431,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "10/07/2020 06:14:32",
          "content": "<p><a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a>  it was your presence that mattered most towards  the ending days in addition to team effort which was very much needed to keep moving in competition when many participants  already might have left race  for obvious reasons . <br>\nYou are person deserving to be in top 3 ,matter  of just some more time . <br>\nI will spend time now in carefully going through detailed approach. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1040405,
      "author_name": "sidneyng",
      "author_url": "",
      "post_date": "10/07/2020 05:55:13",
      "content": "<p>Congrats, truly amazing efforts. Definitely skill, not luck !</p>",
      "votes": null,
      "replies": [
        {
          "id": 1041182,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "10/07/2020 15:40:55",
          "content": "<p>Thank you so much. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1040414,
      "author_name": "freddd1",
      "author_url": "",
      "post_date": "10/07/2020 06:02:50",
      "content": "<p>Congratulations!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1041158,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "10/07/2020 15:29:20",
          "content": "<p>Thanks         </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1040490,
      "author_name": "jaideepvalani",
      "author_url": "",
      "post_date": "10/07/2020 07:03:11",
      "content": "<p>Some highlight on Tab features that we used in tab model for ensemble </p>\n<p>1)  Usage of Volume and Skew features extracted from CT using   segmentation model in this kernel . .<br>\n<a href=\"https://www.kaggle.com/hfutybx/osic-feature-extract-from-ct\" target=\"_blank\">https://www.kaggle.com/hfutybx/osic-feature-extract-from-ct</a> .Standalone Tab model on the top of public kernel using these feature score quite well   on a public lb . Thanks to <a href=\"https://www.kaggle.com/sandorkonya\" target=\"_blank\">@sandorkonya</a> for his insightful post highlighting the usage and importance of it . (Mean,kurtosis dint help ,it worsened the fold cv score )</p>\n<p>2) other features   added in tab model are below, where  bins are quantiles  for age ,first week fvc percent <br>\n<code>['Age_bins','min_pct_bin','min_Percent','week','BASE' ,'Volume','Skew' ].</code><br>\nadding quantiles improved cv  and lb earlier </p>\n<p>3) mloss tuning  using 60 pct of quantile and 40 pct of  lll loss ,max loss of quantile to 800.</p>\n<p>I dint pay much attention over  CV splits  for my tab model so was not relying on CV score for private lb  ,just was checking  how right the approach used, based on lb score  between 6.85 to 6.90 ,any thing less than this was almost impossible  unless using Percent feature which was tightly coupled with FVC  ,fvc=k*percent, where k is constant for Gender,Smooking,Age group. hence causing leak of target to model.</p>\n<p>last thing that i wanted to try but due to lack of time i couldnt try<br>\n  Sampling of batches such  that all patients are in same batch  as its time series data </p>",
      "votes": null,
      "replies": [
        {
          "id": 1041449,
          "author_name": "sandorkonya",
          "author_url": "",
          "post_date": "10/07/2020 18:42:15",
          "content": "<p>Congrats <a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> , <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> and the rest of the team!</p>\n<p>Sadly, we did not survive the shakeup,<br>\ni'll share our segmentation and image features soon.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1041967,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "10/08/2020 02:00:59",
          "content": "<p>Thank you so much <a href=\"https://www.kaggle.com/sandorkonya\" target=\"_blank\">@sandorkonya</a> , I saw your hand-labeled dataset posted, really impressive work and the amount of efforts you and your team contributed. I'm so sorry that you didn't win a reward proportional to your contribution.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1043771,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "10/09/2020 08:28:11",
          "content": "<p><a href=\"https://www.kaggle.com/sandorkonya\" target=\"_blank\">@sandorkonya</a>  would you help understanding   <a href=\"https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection\" target=\"_blank\">https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection</a><br>\nbasically the labels of it.  and why do we have so much classification.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1043971,
          "author_name": "sandorkonya",
          "author_url": "",
          "post_date": "10/09/2020 11:31:29",
          "content": "<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> ,<br>\n<a href=\"https://www.kaggle.com/redwankarimsony\" target=\"_blank\">@redwankarimsony</a> already did a <a href=\"https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/183850\" target=\"_blank\">very nice summary</a> about it!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1040497,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "10/07/2020 07:07:33",
      "content": "<p>Congrats to you and team on results.Thanks for sharing solution <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 1041161,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "10/07/2020 15:29:56",
          "content": "<p>Thanks, and you are welcome.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1040534,
      "author_name": "rafiko1",
      "author_url": "",
      "post_date": "10/07/2020 07:48:36",
      "content": "<p>Congrats to you and the team for surviving shakeup! Concat tile pooling together with the finetuning are very clever ideas</p>",
      "votes": null,
      "replies": [
        {
          "id": 1041193,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "10/07/2020 15:48:34",
          "content": "<p>Thank you so much, I really like those ideas too. I hoped that they would work better, but fortunately they made a strong contribution to our ensemble.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1040536,
      "author_name": "lukereijnen",
      "author_url": "",
      "post_date": "10/07/2020 07:52:47",
      "content": "<p>Really nice job, very interesting approach</p>",
      "votes": null,
      "replies": [
        {
          "id": 1041157,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "10/07/2020 15:28:58",
          "content": "<p>Thank you so much.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1040539,
      "author_name": "ivanilyushchenko",
      "author_url": "",
      "post_date": "10/07/2020 07:54:01",
      "content": "<p>Congratulations to Risers team. Great work!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1041125,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "10/07/2020 15:16:52",
          "content": "<p>Thank you     </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1040541,
      "author_name": "alexj21",
      "author_url": "",
      "post_date": "10/07/2020 07:55:13",
      "content": "<p>Congratulations you deserve it! <br>\nI used your tiling method for PANDA challenge and thought about using it here.. it's interesting to know that it works ! </p>",
      "votes": null,
      "replies": [
        {
          "id": 1041122,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "10/07/2020 15:16:14",
          "content": "<p>Thank you. I joined the competition because I got this idea. Unfortunately, it worked not completely as expected, but I think it's the best way to make use of CT.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1040676,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "10/07/2020 09:31:42",
      "content": "<p>Congratulations on the strong finish. From what I see you're (almost) the only team in the top 50 that didn't jump &gt;500~places, which is really impressive.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1041156,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "10/07/2020 15:28:30",
          "content": "<p><a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> , Thank you so much. Public LB was was not very representative because many participants were exploiting  percent feature: train with FVC percent given at each week and do inference with percent only from the first week, which by coincidence appeared to work really well at public LB, but not making much sense.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1041019,
      "author_name": "greatgamedota",
      "author_url": "",
      "post_date": "10/07/2020 13:57:36",
      "content": "<p>Congrats to you and your team for the solid solution!! Y'all put in so much work and deserved it!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1041132,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "10/07/2020 15:20:03",
          "content": "<p>Thank you so much</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1041946,
      "author_name": "sishihara",
      "author_url": "",
      "post_date": "10/08/2020 01:42:43",
      "content": "<p>Congratulations and thank you for sharing. I'm really impressed with your deep understanding and use of CT scan data.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1041962,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "10/08/2020 01:52:30",
          "content": "<p>Thanks, I really appreciate your words.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1041965,
      "author_name": "ronaldokun",
      "author_url": "",
      "post_date": "10/08/2020 01:57:54",
      "content": "<p>Amazing Ideas <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> . I've followed this insights of yours in the PANDA Competition but didn't think to use it in this one. Congratulations on the well deserved gold.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1042014,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "10/08/2020 02:42:38",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/ronaldokun\" target=\"_blank\">@ronaldokun</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1043149,
      "author_name": "shujun717",
      "author_url": "",
      "post_date": "10/08/2020 18:11:26",
      "content": "<p>Congrats on another gold medal! You really deserve it!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1043198,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "10/08/2020 18:54:54",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> , u2 congratulations with solo gold in Vaccine competition, well done. Also, nice presentation at the workshop today.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1043301,
          "author_name": "shujun717",
          "author_url": "",
          "post_date": "10/08/2020 21:32:29",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a>. I enjoyed your presentation as well. The OpenVaccine competition turned out to be lucky for me in the end. They had to rescore on the last day (Monday), and because of that they delayed the deadline and hence I was able to get more models out during my ensembling. Also I believed in the idea of median avg which would have landed my team a gold in PANDA</p>\n<p>As for the attention masks I showed, they do look concerning. I can't speak for the ISUP grade attention weight distribution, but the ones for majority and minority gleason scores are definitely wrong all the time since the predictions aren't even close even when a lot of times ISUP grades were predicted correctly. But this is the nice thing about using something like self-attention: it is somewhat interpretible and we can make improvements logically based on what we can see from the attention weights.</p>\n<p>As for your tile method, the competition would have been so much more difficult without it, so i think it's still good you shared it. And congrats on best innovative method! That was really well deserved</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1043412,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "10/09/2020 00:53:07",
          "content": "<p>Thank you so much, congratulations to u too for running up reward, attention maps was really good idea. But I'm really worrying about the model interpretability at this point( Hopefully organizers will figure it out how to get all out of the method, especially since it was not that bad at the external data. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1043745,
          "author_name": "shujun717",
          "author_url": "",
          "post_date": "10/09/2020 08:03:36",
          "content": "<p>I agree external data performance was not bad at all! I'm looking forward to how the organizers will organize the results from the competition in the paper</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1043714,
      "author_name": "umarzubair95",
      "author_url": "",
      "post_date": "10/09/2020 07:28:18",
      "content": "<p>Very interesting approach. Thanks for sharing.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1044161,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "10/09/2020 15:03:27",
          "content": "<p>You are welcome</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1041871,
      "author_name": "vicentemantero",
      "author_url": "",
      "post_date": "10/08/2020 00:31:09",
      "content": "<p>Good Job!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1041940,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "10/08/2020 01:41:00",
          "content": "<p>Thanks      </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1040287": "Congratulations to all participants and thanks to organizers for making this competition possible. Despite it didn't go as expected, I hope many people learned something new here. Also, I would like to express my gratitude to my teammates for working together with me on this challenge. Special thanks to @jaideepvalani for not giving up on this competition and motivating our team to keep going, without him we would not go that far. Though, getting the medal I have a little bit bitter feeling. I'll take this chance and describe some of our ideas: hopefully they could help organizers to fight fibrosis. \n\n### CT model: Concatenate Tile Pooling\nWorking on PANDA competition I have proposed [Concatenate Tile Pooling](https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169205) method, which also appeared to be applicable quite well for CT scan data. The idea of this method is illustrated in image below. Instead of assigning labels, like FVC decay and confidence to each CT layer, which may be difficult to predict based on a single image, why not to assign it to all images together?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2Fde95b636611aa0e65432f4dcf6756269%2Ftile.png?generation=1602038549896215&alt=media)\n**why not 3D conv?**\n- CT scans have different spacing, which deteriorates performance of 3D models. It can be viewed as training and performing inference with random stretch of images along one dimension.\n- Many accurate and well-optimized 2D pretrained models.\n- Not possible to train a good 3D model from scratch for so limited data\n\n**Important details:**\n- Since the provided data is very limited, just ~150 samples, train the model first on masked lungs only (masks are produced with [lungmask](https://github.com/JoHof/lungmask/)) with following finetuning on the original images to be able to perform inference without lungmask (So I could do inference of a single 4 fold model within just 10-20 min at kaggle). Check the image below.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1212661%2F31a1d24eb0da1882deb565777c1281f2%2Fmasks.png?generation=1602039389832792&alt=media)\n- For CT models it's better to predict percent rather than FCV. Image augmentation, like zoom, as well as use only of a limited number of slides makes it difficult to compute the volume based on CT. Meanwhile, the model could more reliably say what is the percent of the entire volume is affected by fibrosis by looking just at several scans.\n- Assume linear dependence of the FVC and confidence on the week and predict the slope and the initial value.\n- LLL_loss (Laplace Log Likelihood) works but for convergence one needs to ensure that the output at the beginning of training is close to the gt FVC. I used the following: `FVC = V0*(0.01*a*(w-w0)/134 + b + 0.01*p0)`, `sigma = V0*softplus(c*w/134 + d)`, were a,b,c,d are model predictions, V0 is the full lung volume computed as V0 = 100*FVC/percent. So, initially the predicted FVCs for model before training are quite close to the expected FVC, and the loss converges nicely.  I have computed the loss based on all FVC measurements for a given patient (dropping the first one), so there was no need to evaluate gt slops explicitly. An alternative could be doing something similar to public kernels: model tries to predict FVC directly with using mloss, followed by model finetuning with LLL_loss (which is essentially the metric in this competition).\n- [Tile cutout](https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/169205) - random selection of CT layers during training could be viewed as a cutout in 3d image.\n- When work with the data try to plot it to see what is really going on. For a number of patients I saw incorrect offsets, so the level of signal must be changed. Also I clipped signal at [-1024,600]. Finally, some images are not square, so I needed to take a crop to prevent deformation of images when rescale.   \n\n**Additional details:**\n64 layers x 256 x 256 setup\nResNeXt50 and ResNet18 backbones\nStandard augmentation like rotation, zoom, horizontal flip at the pretraining stage.\nUnfortunately, this beautiful approach didn't work well giving only ~6.85/6.90 at private/public LB.\n\n### Tabular models\nThis is quite standard, and one can find many examples in public kernels. **The important thing: always do proper validation and exclude all kinds of leaks.** So percent must not be used as a feature unless it is set to be equal to the value at the first visit (and the first visit must be excluded from training and validation). Also, we have computed validation based on last 3 visits, similar to the competition metric. We have 2 kinds of models: based only on tabular and tabular+CT features (with some additional variations). \nSome thoughts about **percent madness** people used in public kernels: It seems that by chance public LB contains mostly cases with low slope (FVC decay rate), so anything that artificially reduces slope could boost public LB. If one is training with percent feature and then set it to the constant, it is equivalent to pivoting the prediction in such a way that reduces the slope. \n\n### Final submission\nOur final submission is an ensemble of CT models + TAB NN based models + [Bayesian Ridge Regression](https://www.kaggle.com/jeabat/osic-bayesian-ridge-regression) based on maximizing CV. This submission got 6.8385/6.8884 at private/public LB. It is surprising for so large shake up: we have chosen our best private LB sub.",
    "1040301": "Wow, your solution is really good! We never thought of using Concatenate Tile Pooling in our model :(\n\nAnyways, congratulations on the gold! @iafoss @jaideepvalani @robikscube @poteman @mathurinache",
    "1040306": "Great approach. The only gold winning team I could see with similar public and private LB score. Congrats @iafoss 😀",
    "1040308": "Thank you so much. Actually @robikscube today won 2 gold medals and become a 3xGM.",
    "1040311": "iafoss Yeah, I know that :) Congrats Rob!\n\n@robikscube Time to post a dataset now ;)",
    "1040314": "Thank you so much, @nischaydnk . This competition was quite difficult for all.",
    "1040318": "Congratulations!!!!!!!",
    "1040325": "Thanks",
    "1040329": "Congrats ! You are the only team survived the shake up.",
    "1040339": "Thank you. It was indeed quite difficult.",
    "1040357": "Comprehensive work. Congratulations on your medal.",
    "1040372": "Thank you",
    "1040405": "Congrats, truly amazing efforts. Definitely skill, not luck !",
    "1040414": "Congratulations!",
    "1040431": "iafoss  it was your presence that mattered most towards  the ending days in addition to team effort which was very much needed to keep moving in competition when many participants  already might have left race  for obvious reasons . \nYou are person deserving to be in top 3 ,matter  of just some more time . \nI will spend time now in carefully going through detailed approach.",
    "1040490": "Some highlight on Tab features that we used in tab model for ensemble \n\n1)  Usage of Volume and Skew features extracted from CT using   segmentation model in this kernel . .\nhttps://www.kaggle.com/hfutybx/osic-feature-extract-from-ct .Standalone Tab model on the top of public kernel using these feature score quite well   on a public lb . Thanks to @sandorkonya for his insightful post highlighting the usage and importance of it . (Mean,kurtosis dint help ,it worsened the fold cv score )\n  \n2) other features   added in tab model are below, where  bins are quantiles  for age ,first week fvc percent \n`['Age_bins','min_pct_bin','min_Percent','week','BASE' ,'Volume','Skew' ].`\nadding quantiles improved cv  and lb earlier \n\n3) mloss tuning  using 60 pct of quantile and 40 pct of  lll loss ,max loss of quantile to 800.\n\nI dint pay much attention over  CV splits  for my tab model so was not relying on CV score for private lb  ,just was checking  how right the approach used, based on lb score  between 6.85 to 6.90 ,any thing less than this was almost impossible  unless using Percent feature which was tightly coupled with FVC  ,fvc=k*percent, where k is constant for Gender,Smooking,Age group. hence causing leak of target to model.\n\nlast thing that i wanted to try but due to lack of time i couldnt try\n  Sampling of batches such  that all patients are in same batch  as its time series data",
    "1040497": "Congrats to you and team on results.Thanks for sharing solution @iafoss",
    "1040534": "Congrats to you and the team for surviving shakeup! Concat tile pooling together with the finetuning are very clever ideas",
    "1040536": "Really nice job, very interesting approach",
    "1040539": "Congratulations to Risers team. Great work!",
    "1040541": "Congratulations you deserve it! \nI used your tiling method for PANDA challenge and thought about using it here.. it's interesting to know that it works !",
    "1040676": "Congratulations on the strong finish. From what I see you're (almost) the only team in the top 50 that didn't jump >500~places, which is really impressive.",
    "1041019": "Congrats to you and your team for the solid solution!! Y'all put in so much work and deserved it!",
    "1041122": "Thank you. I joined the competition because I got this idea. Unfortunately, it worked not completely as expected, but I think it's the best way to make use of CT.",
    "1041125": "Thank you",
    "1041132": "Thank you so much",
    "1041156": "theoviel , Thank you so much. Public LB was was not very representative because many participants were exploiting  percent feature: train with FVC percent given at each week and do inference with percent only from the first week, which by coincidence appeared to work really well at public LB, but not making much sense.",
    "1041157": "Thank you so much.",
    "1041158": "Thanks",
    "1041161": "Thanks, and you are welcome.",
    "1041182": "Thank you so much.",
    "1041193": "Thank you so much, I really like those ideas too. I hoped that they would work better, but fortunately they made a strong contribution to our ensemble.",
    "1041449": "Congrats @jaideepvalani , @iafoss and the rest of the team!\n\nSadly, we did not survive the shakeup,\ni'll share our segmentation and image features soon.",
    "1041871": "Good Job!",
    "1041940": "Thanks",
    "1041946": "Congratulations and thank you for sharing. I'm really impressed with your deep understanding and use of CT scan data.",
    "1041962": "Thanks, I really appreciate your words.",
    "1041965": "Amazing Ideas @iafoss . I've followed this insights of yours in the PANDA Competition but didn't think to use it in this one. Congratulations on the well deserved gold.",
    "1041967": "Thank you so much @sandorkonya , I saw your hand-labeled dataset posted, really impressive work and the amount of efforts you and your team contributed. I'm so sorry that you didn't win a reward proportional to your contribution.",
    "1042014": "Thank you @ronaldokun",
    "1043149": "Congrats on another gold medal! You really deserve it!",
    "1043198": "Thanks @shujun717 , u2 congratulations with solo gold in Vaccine competition, well done. Also, nice presentation at the workshop today.",
    "1043301": "Thanks @iafoss. I enjoyed your presentation as well. The OpenVaccine competition turned out to be lucky for me in the end. They had to rescore on the last day (Monday), and because of that they delayed the deadline and hence I was able to get more models out during my ensembling. Also I believed in the idea of median avg which would have landed my team a gold in PANDA\n\nAs for the attention masks I showed, they do look concerning. I can't speak for the ISUP grade attention weight distribution, but the ones for majority and minority gleason scores are definitely wrong all the time since the predictions aren't even close even when a lot of times ISUP grades were predicted correctly. But this is the nice thing about using something like self-attention: it is somewhat interpretible and we can make improvements logically based on what we can see from the attention weights.\n\nAs for your tile method, the competition would have been so much more difficult without it, so i think it's still good you shared it. And congrats on best innovative method! That was really well deserved",
    "1043412": "Thank you so much, congratulations to u too for running up reward, attention maps was really good idea. But I'm really worrying about the model interpretability at this point( Hopefully organizers will figure it out how to get all out of the method, especially since it was not that bad at the external data.",
    "1043714": "Very interesting approach. Thanks for sharing.",
    "1043729": "gunesevitan \nit was really difficult choice to make for private lb submission  based on public lb score thanks to  @Iafoss for his acumen  towards closing hours . \n\n  to give you an eg o   . one of my late submission scored yesterday on public lb   6.8827 , now just imagine how competitive does this score looks on a public lb  but it got only 6.85 on private lb ,while   one of the kernel made public    had  public lb score  of 6.8810  but private lb score 6.8347 .",
    "1043745": "I agree external data performance was not bad at all! I'm looking forward to how the organizers will organize the results from the competition in the paper",
    "1043771": "sandorkonya  would you help understanding   https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection\nbasically the labels of it.  and why do we have so much classification.",
    "1043971": "jaideepvalani ,\n@redwankarimsony already did a [very nice summary](https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/183850) about it!",
    "1044161": "You are welcome"
  },
  "source": "meta"
}