{
  "id": 182653,
  "title": "Just 3 weeks to go - still just one fabulous idea how to use IMG-Data?",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/182653",
  "author_name": "from coffee import *",
  "post_date": "2020-09-13T19:44:19.793000",
  "votes": 11,
  "comment_count": 16,
  "views": 0,
  "content": "<p>Dear fellow Kagglers,</p>\n<p>I found that there are still very very few notebooks which extensively use IMG-Data, but there are a lot of ideas for tabular data or for hyperparameter optimization (Tab-Data only):</p>\n<h3>Notebooks using Tab-Data only</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/chrisden/6-82-quantile-reg-lr-schedulers-checkpoints\" target=\"_blank\">Quantile-Regression, LR Schedulers, Checkpoints</a></li>\n<li><a href=\"https://www.kaggle.com/chrisden/6-82x-tabulardata-find-best-hyperparameter-w-cv\" target=\"_blank\">Hyperparameter-Grid- Search Optimization &amp; CV</a></li>\n<li><a href=\"https://www.kaggle.com/ulrich07/osic-multiple-quantile-regression-starter\" target=\"_blank\">Multiple Quantile Regression Starter, AMAZING Notebook!</a></li>\n</ul>\n<p>----- Sidenote ------<br>\nbtw, check out this <a href=\"https://www.kaggle.com/c/lish-moa\" target=\"_blank\">MOA-COMPETITION!</a> It uses the beautiful log-loss and you can use all the notebooks of above. I already adopted one:<br>\n<a href=\"https://www.kaggle.com/chrisden/tf-keras-multi-label-neuralnets-easy-to-config\" target=\"_blank\">TF Keras: Neural Nets Starter</a><br>\n-----End of Sidenote ------</p>\n<h3>Notebooks using IMG-Data</h3>\n<p>But: Where are the cool Notebooks using IMG Data?<br>\nBest Notebook using IMG Data for me so far is clearly:<br>\n<a href=\"https://www.kaggle.com/carlossouza/end-to-end-model-ct-scans-tabular\" target=\"_blank\">End2End-model-ct-scans-tabular</a></p>\n<p>My team and I are working on an idea recently and I saw several cool ideas in the discussions, like:</p>\n<ol>\n<li>Re-costruct as 3d IMG and calculate the lung volume or</li>\n<li>Segment the lung-area and train a Neural Network based on segmented lung-areas</li>\n</ol>\n<p>But: Did I miss some of the amazing IMG-Data using Notebooks or is it really mostly <a href=\"https://www.kaggle.com/carlossouza\" target=\"_blank\">@carlossouza</a>  amazing notebook?</p>\n<p>Thanks for any insights!</p>",
  "messages": [
    {
      "id": 1009259,
      "postDate": "2020-09-13T19:44:19.793Z",
      "content": "<p>Dear fellow Kagglers,</p>\n<p>I found that there are still very very few notebooks which extensively use IMG-Data, but there are a lot of ideas for tabular data or for hyperparameter optimization (Tab-Data only):</p>\n<h3>Notebooks using Tab-Data only</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/chrisden/6-82-quantile-reg-lr-schedulers-checkpoints\" target=\"_blank\">Quantile-Regression, LR Schedulers, Checkpoints</a></li>\n<li><a href=\"https://www.kaggle.com/chrisden/6-82x-tabulardata-find-best-hyperparameter-w-cv\" target=\"_blank\">Hyperparameter-Grid- Search Optimization &amp; CV</a></li>\n<li><a href=\"https://www.kaggle.com/ulrich07/osic-multiple-quantile-regression-starter\" target=\"_blank\">Multiple Quantile Regression Starter, AMAZING Notebook!</a></li>\n</ul>\n<p>----- Sidenote ------<br>\nbtw, check out this <a href=\"https://www.kaggle.com/c/lish-moa\" target=\"_blank\">MOA-COMPETITION!</a> It uses the beautiful log-loss and you can use all the notebooks of above. I already adopted one:<br>\n<a href=\"https://www.kaggle.com/chrisden/tf-keras-multi-label-neuralnets-easy-to-config\" target=\"_blank\">TF Keras: Neural Nets Starter</a><br>\n-----End of Sidenote ------</p>\n<h3>Notebooks using IMG-Data</h3>\n<p>But: Where are the cool Notebooks using IMG Data?<br>\nBest Notebook using IMG Data for me so far is clearly:<br>\n<a href=\"https://www.kaggle.com/carlossouza/end-to-end-model-ct-scans-tabular\" target=\"_blank\">End2End-model-ct-scans-tabular</a></p>\n<p>My team and I are working on an idea recently and I saw several cool ideas in the discussions, like:</p>\n<ol>\n<li>Re-costruct as 3d IMG and calculate the lung volume or</li>\n<li>Segment the lung-area and train a Neural Network based on segmented lung-areas</li>\n</ol>\n<p>But: Did I miss some of the amazing IMG-Data using Notebooks or is it really mostly <a href=\"https://www.kaggle.com/carlossouza\" target=\"_blank\">@carlossouza</a>  amazing notebook?</p>\n<p>Thanks for any insights!</p>",
      "rawMarkdown": "Dear fellow Kagglers,\n\nI found that there are still very very few notebooks which extensively use IMG-Data, but there are a lot of ideas for tabular data or for hyperparameter optimization (Tab-Data only):\n\n### Notebooks using Tab-Data only\n- [Quantile-Regression, LR Schedulers, Checkpoints](https://www.kaggle.com/chrisden/6-82-quantile-reg-lr-schedulers-checkpoints)\n- [Hyperparameter-Grid- Search Optimization & CV](https://www.kaggle.com/chrisden/6-82x-tabulardata-find-best-hyperparameter-w-cv)\n- [Multiple Quantile Regression Starter, AMAZING Notebook!](https://www.kaggle.com/ulrich07/osic-multiple-quantile-regression-starter)\n\n----- Sidenote ------\nbtw, check out this [MOA-COMPETITION!](https://www.kaggle.com/c/lish-moa) It uses the beautiful log-loss and you can use all the notebooks of above. I already adopted one:\n[TF Keras: Neural Nets Starter](https://www.kaggle.com/chrisden/tf-keras-multi-label-neuralnets-easy-to-config)\n-----End of Sidenote ------\n\n### Notebooks using IMG-Data\nBut: Where are the cool Notebooks using IMG Data?\nBest Notebook using IMG Data for me so far is clearly:\n[End2End-model-ct-scans-tabular](https://www.kaggle.com/carlossouza/end-to-end-model-ct-scans-tabular)\n\nMy team and I are working on an idea recently and I saw several cool ideas in the discussions, like:\n1.  Re-costruct as 3d IMG and calculate the lung volume or\n2. Segment the lung-area and train a Neural Network based on segmented lung-areas\n\nBut: Did I miss some of the amazing IMG-Data using Notebooks or is it really mostly @carlossouza  amazing notebook?\n\nThanks for any insights!",
      "votes": 10
    },
    {
      "id": 1012297,
      "postDate": "2020-09-16T02:41:05.630Z",
      "content": "<p>I have just made public the approach I was pursuing that made use of image data, I have to stop working on it now because I am super busy until after the competition ends, but I thought the idea I had was pretty cool, and it seems to work fairly well. I made a discussion explaining it <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/183235\" target=\"_blank\">here</a> if anyone wants to check it out, I really hope someone picks up from where I left off and makes it work!!</p>",
      "rawMarkdown": "I have just made public the approach I was pursuing that made use of image data, I have to stop working on it now because I am super busy until after the competition ends, but I thought the idea I had was pretty cool, and it seems to work fairly well. I made a discussion explaining it [here](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/183235) if anyone wants to check it out, I really hope someone picks up from where I left off and makes it work!!",
      "votes": 1,
      "replies": [
        {
          "id": 1012310,
          "postDate": "2020-09-16T02:51:28.797Z",
          "content": "<p>Seems like a really cool idea! </p>",
          "rawMarkdown": "Seems like a really cool idea! ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1011376,
      "postDate": "2020-09-15T12:20:25.257Z",
      "content": "<p>I'm playing around with predicting only the FVC (without confidence) and evaluate using mean average error.<br>\nUsing groupKfold with 5 folds and groups respecting patientID, I achieve a mean MAE of 143 across the fold using only tabular data. Adding the (processed) CT scans, I achieve a mean MAE of 139. However, this does not seem significant to me and can (despite cross validation) very well be overfitting I guess. :/</p>",
      "rawMarkdown": "I'm playing around with predicting only the FVC (without confidence) and evaluate using mean average error.\nUsing groupKfold with 5 folds and groups respecting patientID, I achieve a mean MAE of 143 across the fold using only tabular data. Adding the (processed) CT scans, I achieve a mean MAE of 139. However, this does not seem significant to me and can (despite cross validation) very well be overfitting I guess. :/",
      "votes": 1
    },
    {
      "id": 1009433,
      "postDate": "2020-09-14T02:30:52.763Z",
      "content": "<p><a href=\"https://www.kaggle.com/chrisden\" target=\"_blank\">@chrisden</a> I've been asking the same thing with regards to image data: I have attempted to make use of PyTorch-XLA and train on the data but so far it's been erratic at best. And the CV in itself is not so great as a tabular model (only -6.9xxx) so we had to move on from that project to working on a better tabular model.</p>",
      "rawMarkdown": "@chrisden I've been asking the same thing with regards to image data: I have attempted to make use of PyTorch-XLA and train on the data but so far it's been erratic at best. And the CV in itself is not so great as a tabular model (only -6.9xxx) so we had to move on from that project to working on a better tabular model.",
      "votes": 1,
      "replies": [
        {
          "id": 1009507,
          "postDate": "2020-09-14T03:42:34.600Z",
          "content": "<p>Yes, we tried to extract all possible features using CT scans (some were really leaky, had to drop them). On integrating them with the model, the CV didn't improve and LB was erratic (-7.10xxx to -7.32xxx).</p>\n<p>We are still focusing on working on a better tabular model. </p>",
          "rawMarkdown": "Yes, we tried to extract all possible features using CT scans (some were really leaky, had to drop them). On integrating them with the model, the CV didn't improve and LB was erratic (-7.10xxx to -7.32xxx).\n\nWe are still focusing on working on a better tabular model. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1024445,
      "postDate": "2020-09-23T21:47:04.367Z",
      "content": "<p>I have also analysed slopes of FVC and Percents. These can be predicted by the images. Also very useful approach would be to include the uncertainty of the FVC and Percent to the fit. This might sharpen the separation between different illness dynamics.<br>\nMy rough thoughts are summarized in the <a href=\"https://www.kaggle.com/gannadolinska/are-all-the-patients-in-the-data-ill\" target=\"_blank\">notebook</a></p>",
      "rawMarkdown": "I have also analysed slopes of FVC and Percents. These can be predicted by the images. Also very useful approach would be to include the uncertainty of the FVC and Percent to the fit. This might sharpen the separation between different illness dynamics.\nMy rough thoughts are summarized in the [notebook](https://www.kaggle.com/gannadolinska/are-all-the-patients-in-the-data-ill)",
      "votes": 2
    },
    {
      "id": 1012372,
      "postDate": "2020-09-16T03:52:15.590Z",
      "content": "<p>I'm wondering whether or not there is enough data for the models to get any useful data from the images. I took a quick peek at the images and the slopes of their FVC values and the ones with the larger slopes (i.e. quicker declines) didn't have lungs too different from the average until you got to like the bottom 10% of slopes. While I realize I have two untrained eyes, the typical components of honeycombing and such seem very obvious in the really bad lungs while the more typical lungs have little to no difference (and there are even some bad looking lungs that fare fine).</p>\n<p>I have tried training a CNN to recognize a lung that leads to a real bad decline, but that obviously comes with the problem of overfitting on such a small dataset (where there are like 15 visually obvious \"bad lungs\") and I didn't get anywhere past ~6.8 CV. Right now I'm toying with visual attention in the CNN but I'm still getting similar results. Visualizing layer outputs just shows me that the CNN recognizes the fact that it is indeed a lung but doesn't recognize some of the lung structures that signify a quicker decline (e.g. honeycombing). Might not help that I'm using the raw images though, segmentation is still in the to-do list.</p>\n<p>Something I haven't used though is using the DICOM data to calculate lung volumes and use those. Thinking about just using this competition to refine my regression knowledge and tech though because I'm quickly losing hope on the image files being of huge benefit, but I'm also pretty new to this so maybe I'm missing something obvious.</p>",
      "rawMarkdown": "I'm wondering whether or not there is enough data for the models to get any useful data from the images. I took a quick peek at the images and the slopes of their FVC values and the ones with the larger slopes (i.e. quicker declines) didn't have lungs too different from the average until you got to like the bottom 10% of slopes. While I realize I have two untrained eyes, the typical components of honeycombing and such seem very obvious in the really bad lungs while the more typical lungs have little to no difference (and there are even some bad looking lungs that fare fine).\n\nI have tried training a CNN to recognize a lung that leads to a real bad decline, but that obviously comes with the problem of overfitting on such a small dataset (where there are like 15 visually obvious \"bad lungs\") and I didn't get anywhere past ~6.8 CV. Right now I'm toying with visual attention in the CNN but I'm still getting similar results. Visualizing layer outputs just shows me that the CNN recognizes the fact that it is indeed a lung but doesn't recognize some of the lung structures that signify a quicker decline (e.g. honeycombing). Might not help that I'm using the raw images though, segmentation is still in the to-do list.\n\nSomething I haven't used though is using the DICOM data to calculate lung volumes and use those. Thinking about just using this competition to refine my regression knowledge and tech though because I'm quickly losing hope on the image files being of huge benefit, but I'm also pretty new to this so maybe I'm missing something obvious.",
      "votes": 2
    },
    {
      "id": 1009699,
      "postDate": "2020-09-14T07:17:35.187Z",
      "content": "<p>Dear fellows,<br>\nsame experience here so far: IMG Data is not yet leading to a better score. We need to check if it leads to a better CV score, as the LB currently seems to be strongly overfitted.</p>",
      "rawMarkdown": "Dear fellows,\nsame experience here so far: IMG Data is not yet leading to a better score. We need to check if it leads to a better CV score, as the LB currently seems to be strongly overfitted.",
      "votes": 2
    },
    {
      "id": 1025016,
      "postDate": "2020-09-24T08:58:44.357Z",
      "content": "<p>Currently I'm working a lot with Graph Neural Networks.<br>\nI thought about converting the 3D segmented representation to a Graph - mesh and then using Graph Neural Networks to learn the representation of the lung. <br>\nBut I'm not entirely sure how to convert the 3D \"image\" to a graph. </p>\n<p>Generally GNN's perform representation learning and I think that's what we look for - an accurate representation of the lung. But probably it will be too few data points anyway… :)</p>",
      "rawMarkdown": "Currently I'm working a lot with Graph Neural Networks.\nI thought about converting the 3D segmented representation to a Graph - mesh and then using Graph Neural Networks to learn the representation of the lung. \nBut I'm not entirely sure how to convert the 3D \"image\" to a graph. \n\nGenerally GNN's perform representation learning and I think that's what we look for - an accurate representation of the lung. But probably it will be too few data points anyway... :)"
    },
    {
      "id": 1011640,
      "postDate": "2020-09-15T15:52:24.573Z",
      "content": "<p>Let me know if there are any cool public notebooks using IMG-Data, then I will add them to the top-post!</p>",
      "rawMarkdown": "Let me know if there are any cool public notebooks using IMG-Data, then I will add them to the top-post!"
    },
    {
      "id": 1009957,
      "postDate": "2020-09-14T11:55:57.077Z",
      "content": "<p>Are you using the raw image or are you using the segmented images already?<br>\nI guess from the raw images the NN can't really learn anything, especially given the 2h GPU Runtime-LImit, which is not allowing you 5 folds * 1000 Epochs</p>",
      "rawMarkdown": "Are you using the raw image or are you using the segmented images already?\nI guess from the raw images the NN can't really learn anything, especially given the 2h GPU Runtime-LImit, which is not allowing you 5 folds * 1000 Epochs"
    },
    {
      "id": 1009940,
      "postDate": "2020-09-14T11:41:07.090Z",
      "content": "<p>I wish it were a problem of overfit! But what I see is that the models are <strong>not able to learn any useful features</strong> </p>\n<p>To validate the usefulness of image data, I'm setting it as a regression problem, predicting delta, with just the image data (general idea from <a href=\"https://www.kaggle.com/sandorkonya\" target=\"_blank\">@sandorkonya</a> 's post: <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165727\" target=\"_blank\">https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165727</a>)</p>\n<p>I don't see it significantly outperform the dumb baseline (mean delta of train set).<br>\nThere is high variance in the results in CV.</p>",
      "rawMarkdown": "I wish it were a problem of overfit! But what I see is that the models are **not able to learn any useful features** \n\nTo validate the usefulness of image data, I'm setting it as a regression problem, predicting delta, with just the image data (general idea from @sandorkonya 's post: https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165727)\n\nI don't see it significantly outperform the dumb baseline (mean delta of train set).\nThere is high variance in the results in CV."
    },
    {
      "id": 1015039,
      "postDate": "2020-09-17T22:13:30.610Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 1015048,
          "postDate": "2020-09-17T22:21:39.467Z",
          "content": "<p>Sorry I don’t get it. What do you want to say?</p>",
          "rawMarkdown": "Sorry I don’t get it. What do you want to say?"
        }
      ]
    },
    {
      "id": 1012774,
      "postDate": "2020-09-16T09:28:04.250Z",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing."
    },
    {
      "id": 1012442,
      "postDate": "2020-09-16T05:18:35.880Z",
      "content": "<p>Thanks for sharing them all.</p>",
      "rawMarkdown": "Thanks for sharing them all."
    }
  ],
  "comments": [
    {
      "id": 1012297,
      "author_name": "Sam Klein",
      "author_url": "",
      "post_date": "2020-09-16T02:41:05.630000",
      "content": "<p>I have just made public the approach I was pursuing that made use of image data, I have to stop working on it now because I am super busy until after the competition ends, but I thought the idea I had was pretty cool, and it seems to work fairly well. I made a discussion explaining it <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/183235\" target=\"_blank\">here</a> if anyone wants to check it out, I really hope someone picks up from where I left off and makes it work!!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1012310,
          "author_name": "Aadhav Vignesh",
          "author_url": "",
          "post_date": "2020-09-16T02:51:28.797000",
          "content": "<p>Seems like a really cool idea! </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1011376,
      "author_name": "Tobias Tesch",
      "author_url": "",
      "post_date": "2020-09-15T12:20:25.257000",
      "content": "<p>I'm playing around with predicting only the FVC (without confidence) and evaluate using mean average error.<br>\nUsing groupKfold with 5 folds and groups respecting patientID, I achieve a mean MAE of 143 across the fold using only tabular data. Adding the (processed) CT scans, I achieve a mean MAE of 139. However, this does not seem significant to me and can (despite cross validation) very well be overfitting I guess. :/</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1009433,
      "author_name": "Trigram",
      "author_url": "",
      "post_date": "2020-09-14T02:30:52.763000",
      "content": "<p><a href=\"https://www.kaggle.com/chrisden\" target=\"_blank\">@chrisden</a> I've been asking the same thing with regards to image data: I have attempted to make use of PyTorch-XLA and train on the data but so far it's been erratic at best. And the CV in itself is not so great as a tabular model (only -6.9xxx) so we had to move on from that project to working on a better tabular model.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1009507,
          "author_name": "Aadhav Vignesh",
          "author_url": "",
          "post_date": "2020-09-14T03:42:34.600000",
          "content": "<p>Yes, we tried to extract all possible features using CT scans (some were really leaky, had to drop them). On integrating them with the model, the CV didn't improve and LB was erratic (-7.10xxx to -7.32xxx).</p>\n<p>We are still focusing on working on a better tabular model. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1024445,
      "author_name": "Hannah Koenig",
      "author_url": "",
      "post_date": "2020-09-23T21:47:04.367000",
      "content": "<p>I have also analysed slopes of FVC and Percents. These can be predicted by the images. Also very useful approach would be to include the uncertainty of the FVC and Percent to the fit. This might sharpen the separation between different illness dynamics.<br>\nMy rough thoughts are summarized in the <a href=\"https://www.kaggle.com/gannadolinska/are-all-the-patients-in-the-data-ill\" target=\"_blank\">notebook</a></p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1012372,
      "author_name": "Joseph Tan",
      "author_url": "",
      "post_date": "2020-09-16T03:52:15.590000",
      "content": "<p>I'm wondering whether or not there is enough data for the models to get any useful data from the images. I took a quick peek at the images and the slopes of their FVC values and the ones with the larger slopes (i.e. quicker declines) didn't have lungs too different from the average until you got to like the bottom 10% of slopes. While I realize I have two untrained eyes, the typical components of honeycombing and such seem very obvious in the really bad lungs while the more typical lungs have little to no difference (and there are even some bad looking lungs that fare fine).</p>\n<p>I have tried training a CNN to recognize a lung that leads to a real bad decline, but that obviously comes with the problem of overfitting on such a small dataset (where there are like 15 visually obvious \"bad lungs\") and I didn't get anywhere past ~6.8 CV. Right now I'm toying with visual attention in the CNN but I'm still getting similar results. Visualizing layer outputs just shows me that the CNN recognizes the fact that it is indeed a lung but doesn't recognize some of the lung structures that signify a quicker decline (e.g. honeycombing). Might not help that I'm using the raw images though, segmentation is still in the to-do list.</p>\n<p>Something I haven't used though is using the DICOM data to calculate lung volumes and use those. Thinking about just using this competition to refine my regression knowledge and tech though because I'm quickly losing hope on the image files being of huge benefit, but I'm also pretty new to this so maybe I'm missing something obvious.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1009699,
      "author_name": "from coffee import *",
      "author_url": "",
      "post_date": "2020-09-14T07:17:35.187000",
      "content": "<p>Dear fellows,<br>\nsame experience here so far: IMG Data is not yet leading to a better score. We need to check if it leads to a better CV score, as the LB currently seems to be strongly overfitted.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1025016,
      "author_name": "Flogrammer",
      "author_url": "",
      "post_date": "2020-09-24T08:58:44.357000",
      "content": "<p>Currently I'm working a lot with Graph Neural Networks.<br>\nI thought about converting the 3D segmented representation to a Graph - mesh and then using Graph Neural Networks to learn the representation of the lung. <br>\nBut I'm not entirely sure how to convert the 3D \"image\" to a graph. </p>\n<p>Generally GNN's perform representation learning and I think that's what we look for - an accurate representation of the lung. But probably it will be too few data points anyway… :)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1011640,
      "author_name": "from coffee import *",
      "author_url": "",
      "post_date": "2020-09-15T15:52:24.573000",
      "content": "<p>Let me know if there are any cool public notebooks using IMG-Data, then I will add them to the top-post!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1009957,
      "author_name": "from coffee import *",
      "author_url": "",
      "post_date": "2020-09-14T11:55:57.077000",
      "content": "<p>Are you using the raw image or are you using the segmented images already?<br>\nI guess from the raw images the NN can't really learn anything, especially given the 2h GPU Runtime-LImit, which is not allowing you 5 folds * 1000 Epochs</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1009940,
      "author_name": "JohnT",
      "author_url": "",
      "post_date": "2020-09-14T11:41:07.090000",
      "content": "<p>I wish it were a problem of overfit! But what I see is that the models are <strong>not able to learn any useful features</strong> </p>\n<p>To validate the usefulness of image data, I'm setting it as a regression problem, predicting delta, with just the image data (general idea from <a href=\"https://www.kaggle.com/sandorkonya\" target=\"_blank\">@sandorkonya</a> 's post: <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165727\" target=\"_blank\">https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165727</a>)</p>\n<p>I don't see it significantly outperform the dumb baseline (mean delta of train set).<br>\nThere is high variance in the results in CV.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1015039,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-17T22:13:30.610000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1015048,
          "author_name": "from coffee import *",
          "author_url": "",
          "post_date": "2020-09-17T22:21:39.467000",
          "content": "<p>Sorry I don’t get it. What do you want to say?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1012774,
      "author_name": "R. Joseph Manoj, PhD",
      "author_url": "",
      "post_date": "2020-09-16T09:28:04.250000",
      "content": "<p>Thanks for sharing.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1012442,
      "author_name": "Taral Sarvagod",
      "author_url": "",
      "post_date": "2020-09-16T05:18:35.880000",
      "content": "<p>Thanks for sharing them all.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1009259": "Dear fellow Kagglers,\n\nI found that there are still very very few notebooks which extensively use IMG-Data, but there are a lot of ideas for tabular data or for hyperparameter optimization (Tab-Data only):\n\n### Notebooks using Tab-Data only\n- [Quantile-Regression, LR Schedulers, Checkpoints](https://www.kaggle.com/chrisden/6-82-quantile-reg-lr-schedulers-checkpoints)\n- [Hyperparameter-Grid- Search Optimization & CV](https://www.kaggle.com/chrisden/6-82x-tabulardata-find-best-hyperparameter-w-cv)\n- [Multiple Quantile Regression Starter, AMAZING Notebook!](https://www.kaggle.com/ulrich07/osic-multiple-quantile-regression-starter)\n\n----- Sidenote ------\nbtw, check out this [MOA-COMPETITION!](https://www.kaggle.com/c/lish-moa) It uses the beautiful log-loss and you can use all the notebooks of above. I already adopted one:\n[TF Keras: Neural Nets Starter](https://www.kaggle.com/chrisden/tf-keras-multi-label-neuralnets-easy-to-config)\n-----End of Sidenote ------\n\n### Notebooks using IMG-Data\nBut: Where are the cool Notebooks using IMG Data?\nBest Notebook using IMG Data for me so far is clearly:\n[End2End-model-ct-scans-tabular](https://www.kaggle.com/carlossouza/end-to-end-model-ct-scans-tabular)\n\nMy team and I are working on an idea recently and I saw several cool ideas in the discussions, like:\n1.  Re-costruct as 3d IMG and calculate the lung volume or\n2. Segment the lung-area and train a Neural Network based on segmented lung-areas\n\nBut: Did I miss some of the amazing IMG-Data using Notebooks or is it really mostly @carlossouza  amazing notebook?\n\nThanks for any insights!",
    "1012297": "I have just made public the approach I was pursuing that made use of image data, I have to stop working on it now because I am super busy until after the competition ends, but I thought the idea I had was pretty cool, and it seems to work fairly well. I made a discussion explaining it [here](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/183235) if anyone wants to check it out, I really hope someone picks up from where I left off and makes it work!!",
    "1011376": "I'm playing around with predicting only the FVC (without confidence) and evaluate using mean average error.\nUsing groupKfold with 5 folds and groups respecting patientID, I achieve a mean MAE of 143 across the fold using only tabular data. Adding the (processed) CT scans, I achieve a mean MAE of 139. However, this does not seem significant to me and can (despite cross validation) very well be overfitting I guess. :/",
    "1009433": "@chrisden I've been asking the same thing with regards to image data: I have attempted to make use of PyTorch-XLA and train on the data but so far it's been erratic at best. And the CV in itself is not so great as a tabular model (only -6.9xxx) so we had to move on from that project to working on a better tabular model.",
    "1024445": "I have also analysed slopes of FVC and Percents. These can be predicted by the images. Also very useful approach would be to include the uncertainty of the FVC and Percent to the fit. This might sharpen the separation between different illness dynamics.\nMy rough thoughts are summarized in the [notebook](https://www.kaggle.com/gannadolinska/are-all-the-patients-in-the-data-ill)",
    "1012372": "I'm wondering whether or not there is enough data for the models to get any useful data from the images. I took a quick peek at the images and the slopes of their FVC values and the ones with the larger slopes (i.e. quicker declines) didn't have lungs too different from the average until you got to like the bottom 10% of slopes. While I realize I have two untrained eyes, the typical components of honeycombing and such seem very obvious in the really bad lungs while the more typical lungs have little to no difference (and there are even some bad looking lungs that fare fine).\n\nI have tried training a CNN to recognize a lung that leads to a real bad decline, but that obviously comes with the problem of overfitting on such a small dataset (where there are like 15 visually obvious \"bad lungs\") and I didn't get anywhere past ~6.8 CV. Right now I'm toying with visual attention in the CNN but I'm still getting similar results. Visualizing layer outputs just shows me that the CNN recognizes the fact that it is indeed a lung but doesn't recognize some of the lung structures that signify a quicker decline (e.g. honeycombing). Might not help that I'm using the raw images though, segmentation is still in the to-do list.\n\nSomething I haven't used though is using the DICOM data to calculate lung volumes and use those. Thinking about just using this competition to refine my regression knowledge and tech though because I'm quickly losing hope on the image files being of huge benefit, but I'm also pretty new to this so maybe I'm missing something obvious.",
    "1009699": "Dear fellows,\nsame experience here so far: IMG Data is not yet leading to a better score. We need to check if it leads to a better CV score, as the LB currently seems to be strongly overfitted.",
    "1025016": "Currently I'm working a lot with Graph Neural Networks.\nI thought about converting the 3D segmented representation to a Graph - mesh and then using Graph Neural Networks to learn the representation of the lung. \nBut I'm not entirely sure how to convert the 3D \"image\" to a graph. \n\nGenerally GNN's perform representation learning and I think that's what we look for - an accurate representation of the lung. But probably it will be too few data points anyway... :)",
    "1011640": "Let me know if there are any cool public notebooks using IMG-Data, then I will add them to the top-post!",
    "1009957": "Are you using the raw image or are you using the segmented images already?\nI guess from the raw images the NN can't really learn anything, especially given the 2h GPU Runtime-LImit, which is not allowing you 5 folds * 1000 Epochs",
    "1009940": "I wish it were a problem of overfit! But what I see is that the models are **not able to learn any useful features** \n\nTo validate the usefulness of image data, I'm setting it as a regression problem, predicting delta, with just the image data (general idea from @sandorkonya 's post: https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165727)\n\nI don't see it significantly outperform the dumb baseline (mean delta of train set).\nThere is high variance in the results in CV.",
    "1015039": "",
    "1012774": "Thanks for sharing.",
    "1012442": "Thanks for sharing them all."
  }
}