{
  "id": 185077,
  "title": "Results of experiments with different IMG-derived features",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/185077",
  "author_name": "",
  "post_date": "2020-09-19T09:44:48.929491600Z",
  "votes": 25,
  "comment_count": 44,
  "views": 0,
  "content": "<p>Dear fellow Kagglers,</p>\n<p>my team &amp; I have tried various experiments with <strong>newly derived features from Image data</strong>, like mentioned in Laura's <a href=\"https://www.kaggle.com/allunia/pulmonary-dicom-preprocessing\" target=\"_blank\">Notebook: pulmonary-dicom-preprocessing.</a></p>\n<p>Some of those features were:</p>\n<ul>\n<li>img resolution</li>\n<li>slice-thickness</li>\n<li>pixelspacing_r</li>\n<li>pixelspacing_c</li>\n<li>pixelspacing_area</li>\n<li>window_width</li>\n<li>image_area_cm2</li>\n<li>slice_volume_cm3</li>\n</ul>\n<p>Our current finding is, that those features sadly DO NOT improve CV/LB.</p>\n<p>Do you have share the same experience &amp; get similar results, or are any of those features working for you?<br>\nDid I miss important features?</p>\n<p>To be more precise:<br>\nThe above mentioned features do NOT have any causal relationship with the actual medical condition or prognosis of the patients. It's basically a test if there is a correlation: we couldn't find any. Even if there is a correlation, this would only help in the competition, but obviously not in the real-world examples.</p>\n<p>Using features mentioned by <a href=\"https://www.kaggle.com/aadhavvignesh\" target=\"_blank\">@aadhavvignesh</a> and <a href=\"https://www.kaggle.com/abhishekgbhat\" target=\"_blank\">@abhishekgbhat</a> <strong>might</strong> have a causal relationship with the actual prognosis and the FVC:</p>\n<ul>\n<li>Lung Volume in cm^3</li>\n<li>Lung Area in cm^2</li>\n<li>Average of (Tissue area)/(Lung Area) across all images for a particular patient and FVC</li>\n</ul>\n<p>Thank you all for adding some insights on fighting Pulmonary Fibrosis!</p>\n<p><a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a> feel free to add!</p>",
  "messages": [
    {
      "id": "1017848",
      "postDate": "09/19/2020 09:44:48",
      "content": "<p>Dear fellow Kagglers,</p>\n<p>my team &amp; I have tried various experiments with <strong>newly derived features from Image data</strong>, like mentioned in Laura's <a href=\"https://www.kaggle.com/allunia/pulmonary-dicom-preprocessing\" target=\"_blank\">Notebook: pulmonary-dicom-preprocessing.</a></p>\n<p>Some of those features were:</p>\n<ul>\n<li>img resolution</li>\n<li>slice-thickness</li>\n<li>pixelspacing_r</li>\n<li>pixelspacing_c</li>\n<li>pixelspacing_area</li>\n<li>window_width</li>\n<li>image_area_cm2</li>\n<li>slice_volume_cm3</li>\n</ul>\n<p>Our current finding is, that those features sadly DO NOT improve CV/LB.</p>\n<p>Do you have share the same experience &amp; get similar results, or are any of those features working for you?<br>\nDid I miss important features?</p>\n<p>To be more precise:<br>\nThe above mentioned features do NOT have any causal relationship with the actual medical condition or prognosis of the patients. It's basically a test if there is a correlation: we couldn't find any. Even if there is a correlation, this would only help in the competition, but obviously not in the real-world examples.</p>\n<p>Using features mentioned by <a href=\"https://www.kaggle.com/aadhavvignesh\" target=\"_blank\">@aadhavvignesh</a> and <a href=\"https://www.kaggle.com/abhishekgbhat\" target=\"_blank\">@abhishekgbhat</a> <strong>might</strong> have a causal relationship with the actual prognosis and the FVC:</p>\n<ul>\n<li>Lung Volume in cm^3</li>\n<li>Lung Area in cm^2</li>\n<li>Average of (Tissue area)/(Lung Area) across all images for a particular patient and FVC</li>\n</ul>\n<p>Thank you all for adding some insights on fighting Pulmonary Fibrosis!</p>\n<p><a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a> feel free to add!</p>",
      "rawMarkdown": "Dear fellow Kagglers,\n\nmy team & I have tried various experiments with **newly derived features from Image data**, like mentioned in Laura's [Notebook: pulmonary-dicom-preprocessing.](https://www.kaggle.com/allunia/pulmonary-dicom-preprocessing)\n\nSome of those features were:\n\n- img resolution\n- slice-thickness\n- pixelspacing_r\n- pixelspacing_c\n- pixelspacing_area\n- window_width\n- image_area_cm2\n- slice_volume_cm3\n\nOur current finding is, that those features sadly DO NOT improve CV/LB.\n\nDo you have share the same experience & get similar results, or are any of those features working for you?\nDid I miss important features?\n\nTo be more precise:\nThe above mentioned features do NOT have any causal relationship with the actual medical condition or prognosis of the patients. It's basically a test if there is a correlation: we couldn't find any. Even if there is a correlation, this would only help in the competition, but obviously not in the real-world examples.\n\nUsing features mentioned by @aadhavvignesh and @abhishekgbhat **might** have a causal relationship with the actual prognosis and the FVC:\n- Lung Volume in cm^3\n- Lung Area in cm^2\n- Average of (Tissue area)/(Lung Area) across all images for a particular patient and FVC\n\n\nThank you all for adding some insights on fighting Pulmonary Fibrosis!\n\n@tanulsingh077 feel free to add!",
      "votes": null
    },
    {
      "id": "1017994",
      "postDate": "09/19/2020 11:28:17",
      "content": "<p>I have tried including the following features you mentioned:</p>\n<ul>\n<li>Slice Thickness </li>\n<li>Pixel Spacing R</li>\n<li>Pixel Spacing C<br>\nNone of the above features had any correlation with FVC or decline in FVC over time.</li>\n</ul>\n<p>But I did observe some good correlation between:</p>\n<ul>\n<li>Lung volume and FVC</li>\n<li>Average of (Tissue area)/(Lung Area) across all images for a particular patient and FVC</li>\n<li>Average tissue area and decline in FVC over time</li>\n</ul>\n<p>You can consider including some of the above features in your experiment. <br>\nAlso how do you define - image_area_cm2, slice_volume_cm3?</p>",
      "rawMarkdown": "I have tried including the following features you mentioned:\n- Slice Thickness \n- Pixel Spacing R\n- Pixel Spacing C\nNone of the above features had any correlation with FVC or decline in FVC over time.\n\nBut I did observe some good correlation between:\n- Lung volume and FVC\n- Average of (Tissue area)/(Lung Area) across all images for a particular patient and FVC\n- Average tissue area and decline in FVC over time\n\nYou can consider including some of the above features in your experiment. \nAlso how do you define - image_area_cm2, slice_volume_cm3?",
      "votes": null
    },
    {
      "id": "1018139",
      "postDate": "09/19/2020 13:22:50",
      "content": "<p>i agree these features didnt helped me to get a better CV/LB .\nbtw would you like to merge team with me as i dont have a team ,</p>",
      "rawMarkdown": "i agree these features didnt helped me to get a better CV/LB .\nbtw would you like to merge team with me as i dont have a team ,",
      "votes": null
    },
    {
      "id": "1018277",
      "postDate": "09/19/2020 15:15:39",
      "content": "<p><strong>What I got to know from Image Data:</strong></p>\n<p>I had already mentioned that these features didn't work:</p>\n<ul>\n<li>Lung Volume <code>in cm^3</code></li>\n<li>Lung Area <code>in cm^2</code></li>\n<li>Image Resolution/Dimensions</li>\n</ul>\n<p>Segmented lung images didn't work as expected. Poor CV/LB still persists. </p>\n<p>This has led our team to put less focus on image data, and more on the available tabular data. I unfortunately can't reveal anything beyond this, because every team would love to keep their sauce 'secret' :P</p>\n<p><strong>About LB and rankings:</strong></p>\n<p>I personally feel that the LB is <strong>NOT</strong> a good indicator of your team's results. I guess that I can safely say that our team has a stable model (my instinct says so :P), and our rankings aren't good. A shake-up is imminent, so I do feel teams should not be looking at the LB scores, but indeed should have a good CV set up.</p>",
      "rawMarkdown": "**What I got to know from Image Data:**\n\nI had already mentioned that these features didn't work:\n\n- Lung Volume `in cm^3`\n- Lung Area `in cm^2`\n- Image Resolution/Dimensions\n\nSegmented lung images didn't work as expected. Poor CV/LB still persists. \n\nThis has led our team to put less focus on image data, and more on the available tabular data. I unfortunately can't reveal anything beyond this, because every team would love to keep their sauce 'secret' :P\n\n**About LB and rankings:**\n\nI personally feel that the LB is **NOT** a good indicator of your team's results. I guess that I can safely say that our team has a stable model (my instinct says so :P), and our rankings aren't good. A shake-up is imminent, so I do feel teams should not be looking at the LB scores, but indeed should have a good CV set up.",
      "votes": null
    },
    {
      "id": "1018303",
      "postDate": "09/19/2020 15:32:43",
      "content": "<p>it doesn't help for CV/LB. </p>",
      "rawMarkdown": "it doesn't help for CV/LB.",
      "votes": null
    },
    {
      "id": "1018351",
      "postDate": "09/19/2020 16:18:50",
      "content": "<p>hey adhav does your team needs a member ? , i am currently seraching for a team , my current rank is 26</p>",
      "rawMarkdown": "hey adhav does your team needs a member ? , i am currently seraching for a team , my current rank is 26",
      "votes": null
    },
    {
      "id": "1018991",
      "postDate": "09/20/2020 06:02:46",
      "content": "<p>Tried image resolution , augmentation while training , DenseNet and Efficientnet ensembling but all in vain</p>",
      "rawMarkdown": "Tried image resolution , augmentation while training , DenseNet and Efficientnet ensembling but all in vain",
      "votes": null
    },
    {
      "id": "1019284",
      "postDate": "09/20/2020 10:28:07",
      "content": "<p>Thanks for sharing your findings. Even if they improve cv score, I wouldn't use them anyway. Since FVC is not dependent to them, it would be huge gamble to use them in your final submissions.  </p>",
      "rawMarkdown": "Thanks for sharing your findings. Even if they improve cv score, I wouldn't use them anyway. Since FVC is not dependent to them, it would be huge gamble to use them in your final submissions.",
      "votes": null
    },
    {
      "id": "1019757",
      "postDate": "09/20/2020 16:45:34",
      "content": "<p>These features contain diagnostic information for the CT scanner setup, any correlation they have with lung function will be completely incidental anyway. </p>",
      "rawMarkdown": "These features contain diagnostic information for the CT scanner setup, any correlation they have with lung function will be completely incidental anyway.",
      "votes": null
    },
    {
      "id": "1020565",
      "postDate": "09/21/2020 09:07:22",
      "content": "<p>The holdout predictions for private LB is 85% where as public LB is scored on 15% of the predictions, although features related pixel spacing, slice thickness are useless, others like lung volume, lung are useful. </p>\n<p>Wait for a huge LB shakeup once private LB is announced.</p>",
      "rawMarkdown": "The holdout predictions for private LB is 85% where as public LB is scored on 15% of the predictions, although features related pixel spacing, slice thickness are useless, others like lung volume, lung are useful. \n\nWait for a huge LB shakeup once private LB is announced.",
      "votes": null
    },
    {
      "id": "1022098",
      "postDate": "09/22/2020 10:07:01",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/abhishekgbhat\" target=\"_blank\">@abhishekgbhat</a>, I just saw your post yesterday night, sorry for the delay.<br>\nThe feautes are derived from Laura's Notebook <a href=\"https://www.kaggle.com/allunia/pulmonary-dicom-preprocessing\" target=\"_blank\">pulmonary-dicom-preprocessing</a>. E.g. img_area_cm2 is derived by the resolution.</p>\n<p>How do you derive \"Average of (Tissue area)/(Lung Area)\"?<br>\nIs there a notebook available? Sounds like a lot of work using Hounsfield values &amp; segmentation.</p>",
      "rawMarkdown": "Hi @abhishekgbhat, I just saw your post yesterday night, sorry for the delay.\nThe feautes are derived from Laura's Notebook [pulmonary-dicom-preprocessing](https://www.kaggle.com/allunia/pulmonary-dicom-preprocessing). E.g. img_area_cm2 is derived by the resolution.\n\nHow do you derive \"Average of (Tissue area)/(Lung Area)\"?\nIs there a notebook available? Sounds like a lot of work using Hounsfield values & segmentation.",
      "votes": null
    },
    {
      "id": "1022100",
      "postDate": "09/22/2020 10:08:33",
      "content": "<p>Thank you for the feedback, that supports our findings.</p>",
      "rawMarkdown": "Thank you for the feedback, that supports our findings.",
      "votes": null
    },
    {
      "id": "1022140",
      "postDate": "09/22/2020 10:37:55",
      "content": "<p>I'm really sorry <a href=\"https://www.kaggle.com/subzeroop\" target=\"_blank\">@subzeroop</a>, our team is full. I hope you can find a team soon :)</p>",
      "rawMarkdown": "I'm really sorry @subzeroop, our team is full. I hope you can find a team soon :)",
      "votes": null
    },
    {
      "id": "1022812",
      "postDate": "09/22/2020 18:25:35",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/ChristianDenich\" target=\"_blank\">@ChristianDenich</a>, thanks for the clarification. </p>\n<blockquote>\n  <p>How do you derive \"Average of (Tissue area)/(Lung Area)\"?</p>\n</blockquote>\n<p>No, I haven't seen any notebook with this feature. From the limited research I have done, the fiberosis tissue plays an important role in the severity of the disease. So I found out a way to segment the tissues within the lung. From this segmented image I obtained the number of tissue pixels and multiplied it with Pixel Spacing R and C. This gives the tissue area in each image. Assuming you have the lung area for each image, you can get this ratio for each image. Finally, if you have this ratio for all images of a particular patient you can simply take the average. The tough part is to segment the tissues effectively. </p>",
      "rawMarkdown": "Hey @ChristianDenich, thanks for the clarification. \n\n> How do you derive \"Average of (Tissue area)/(Lung Area)\"?\n\nNo, I haven't seen any notebook with this feature. From the limited research I have done, the fiberosis tissue plays an important role in the severity of the disease. So I found out a way to segment the tissues within the lung. From this segmented image I obtained the number of tissue pixels and multiplied it with Pixel Spacing R and C. This gives the tissue area in each image. Assuming you have the lung area for each image, you can get this ratio for each image. Finally, if you have this ratio for all images of a particular patient you can simply take the average. The tough part is to segment the tissues effectively.",
      "votes": null
    },
    {
      "id": "1023338",
      "postDate": "09/23/2020 06:25:40",
      "content": "<p>So from a maths point of view, this should be the same as considering (Tissue volume)/(Lung volume) I guess. </p>\n<p>I considered this feature and thought it gave me a slight improvement, but after changing some hyperparameters etc. I tried again without the feature and got a higher score than with the feature. </p>\n<p>Maybe it works better with a better segmentation technique though. I just used the ranges of the HU scale for segmentation.</p>",
      "rawMarkdown": "So from a maths point of view, this should be the same as considering (Tissue volume)/(Lung volume) I guess. \n\nI considered this feature and thought it gave me a slight improvement, but after changing some hyperparameters etc. I tried again without the feature and got a higher score than with the feature. \n\nMaybe it works better with a better segmentation technique though. I just used the ranges of the HU scale for segmentation.",
      "votes": null
    },
    {
      "id": "1024017",
      "postDate": "09/23/2020 15:35:13",
      "content": "<p>This is not exactly the volume because im not considering the third dimension(i.e slice thickness) here. This feature tells us the following:  For a particular patient, on an average what percent of the lung area does the tissue take up </p>",
      "rawMarkdown": "This is not exactly the volume because im not considering the third dimension(i.e slice thickness) here. This feature tells us the following:  For a particular patient, on an average what percent of the lung area does the tissue take up",
      "votes": null
    },
    {
      "id": "1024030",
      "postDate": "09/23/2020 15:46:04",
      "content": "<p>Yes, but to get the volume of lung, you would have to multiply the average lung area by the length of the lung (3rd dimension). Analoguously, you get the volume of tissue, when you multiply the average tissue area by the length of the lung. When you divide both things, the length cancels out, such that<br>\nvolume of lung / volume of tissue =  average area of lung / average area of tissue.<br>\nDo you agree?</p>",
      "rawMarkdown": "Yes, but to get the volume of lung, you would have to multiply the average lung area by the length of the lung (3rd dimension). Analoguously, you get the volume of tissue, when you multiply the average tissue area by the length of the lung. When you divide both things, the length cancels out, such that\nvolume of lung / volume of tissue =  average area of lung / average area of tissue.\nDo you agree?",
      "votes": null
    },
    {
      "id": "1024276",
      "postDate": "09/23/2020 18:19:18",
      "content": "<p>We tried:</p>\n<p>CenterCrop, Resampled to Pixel Spacing and SliceThickness equal to 1, Mask the Lung and added the following</p>\n<p>Pixel Voxel Statistics: Mean, STD, Kurtosis &amp; Skew Distribution<br>\nPercentage in Lung Window<br>\nChest Circunference<br>\nLung Height<br>\nLung Volume in cm3 or liters (dm3)<br>\nHU Bins Distribution from 2 to 14 bins ( Just adding the HU distribution in bins ).</p>\n<p>Any addition or combination of those items just made it worse in the LB, sometimes better CV but always worse LB.</p>",
      "rawMarkdown": "We tried:\n\nCenterCrop, Resampled to Pixel Spacing and SliceThickness equal to 1, Mask the Lung and added the following\n\nPixel Voxel Statistics: Mean, STD, Kurtosis & Skew Distribution\nPercentage in Lung Window\nChest Circunference\nLung Height\nLung Volume in cm3 or liters (dm3)\nHU Bins Distribution from 2 to 14 bins ( Just adding the HU distribution in bins ).\n\nAny addition or combination of those items just made it worse in the LB, sometimes better CV but always worse LB.",
      "votes": null
    },
    {
      "id": "1024639",
      "postDate": "09/24/2020 03:33:35",
      "content": "<p>I agree with this. We've tried all possible combinations and on integrating them with the model gave us worse LB and worse CV in almost every case.</p>\n<p>Does your team know what exactly the problem is, because at this point making use of scans seems pointless. (not meant to be conclusive, but including them is making CV/LB worse)</p>\n<p>Do you also think that the winners' solution would help the hosts if they don't use image data for the final model, as I guess the primary objective of the hosts would be to make use of the scans for determining pulmonary fibrosis?</p>",
      "rawMarkdown": "I agree with this. We've tried all possible combinations and on integrating them with the model gave us worse LB and worse CV in almost every case.\n\nDoes your team know what exactly the problem is, because at this point making use of scans seems pointless. (not meant to be conclusive, but including them is making CV/LB worse)\n\nDo you also think that the winners' solution would help the hosts if they don't use image data for the final model, as I guess the primary objective of the hosts would be to make use of the scans for determining pulmonary fibrosis?",
      "votes": null
    },
    {
      "id": "1024676",
      "postDate": "09/24/2020 04:15:49",
      "content": "<p>We are not sure what is wrong, but we think that the release of just the baseline week of the images and also the test tabular data was a fundamental flaw in this competition. We are supposed to predict 146 weeks - time series data - and we don't have a time series data to train on.<br>\nMedical data is already extremely heterogenous, it seems we have different distributions of patients in the training set ( this notebook by <a href=\"https://www.kaggle.com/gannadolinska\" target=\"_blank\">@gannadolinska</a> shows we can consider at least 3 different distributions of patients, those which decrease lung capacity, those steady and those which get better. ) <br>\nThe tabular data has the percent feature, another controversial addition given that is highly correlated with the target, and we also don't have it in the test data. <br>\nIn an ideal case the image in the baseline would have enough information to guide the model to the different distributions of patients and the evolution of them but we failed to extract any<br>\nWe pretty much gave up using the images entirely, it was a very frustrating couple of months.  <br>\nEither someone has a bag of tricks no one else has implemented using the images or the organizers are going to pay for a baseline tabular model, overfitted and completely useless medically. </p>",
      "rawMarkdown": "We are not sure what is wrong, but we think that the release of just the baseline week of the images and also the test tabular data was a fundamental flaw in this competition. We are supposed to predict 146 weeks - time series data - and we don't have a time series data to train on.\n\nMedical data is already extremely heterogenous, it seems we have different distributions of patients in the training set ( this notebook by @gannadolinska shows we can consider at least 3 different distributions of patients, those which decrease lung capacity, those steady and those which get better. ) \n\nThe tabular data has the percent feature, another controversial addition given that is highly correlated with the target, and we also don't have it in the test data. \n\nIn an ideal case the image in the baseline would have enough information to guide the model to the different distributions of patients and the evolution of them but we failed to extract any\n\n We pretty much gave up using the images entirely, it was a very frustrating couple of months.  \n\nEither someone has a bag of tricks no one else has implemented using the images or the organizers are going to pay for a baseline tabular model, overfitted and completely useless medically.",
      "votes": null
    },
    {
      "id": "1025134",
      "postDate": "09/24/2020 10:53:06",
      "content": "<p>Yes, we had initially planned to extract some useful features using image data, and I even saw some teams trying to manually honeycomb lungs and trying to make it work. Eventually, everyone who had used images became frustrated due to poor results.</p>\n<p>I too fear that the organizers are going to end up with a baseline tabular model, rendering the objective of the competition to be useless.</p>\n<p>Hope we get to see a novel approach from the winners, but for now I guess most of the LB is saturated with overfitting models.</p>",
      "rawMarkdown": "Yes, we had initially planned to extract some useful features using image data, and I even saw some teams trying to manually honeycomb lungs and trying to make it work. Eventually, everyone who had used images became frustrated due to poor results.\n\nI too fear that the organizers are going to end up with a baseline tabular model, rendering the objective of the competition to be useless.\n\nHope we get to see a novel approach from the winners, but for now I guess most of the LB is saturated with overfitting models.",
      "votes": null
    },
    {
      "id": "1025632",
      "postDate": "09/24/2020 17:21:10",
      "content": "<p>Hi, I've been noticing too  that the images are frustratingly not helping the public LB position but I'm going to keep focussing on them anyway.  My hope  is that the CT scans must have some value and that a big shakeup will occur where those models that make good use of the scans perform better on the private LB.   I'm not looking at the technial aspects of the ct-scan (like pixel size, slice thickness) but I'm confident that the shapes, areas and textures of the slices themselves can be useful.</p>\n<p>I must admit though that this hope is more an ensemble of sunk-cost fallacy + dumb-optimism + \"what is point of this competition if not\"…. :)</p>",
      "rawMarkdown": "Hi, I've been noticing too  that the images are frustratingly not helping the public LB position but I'm going to keep focussing on them anyway.  My hope  is that the CT scans must have some value and that a big shakeup will occur where those models that make good use of the scans perform better on the private LB.   I'm not looking at the technial aspects of the ct-scan (like pixel size, slice thickness) but I'm confident that the shapes, areas and textures of the slices themselves can be useful.\n\nI must admit though that this hope is more an ensemble of sunk-cost fallacy + dumb-optimism + \"what is point of this competition if not\".... :)",
      "votes": null
    },
    {
      "id": "1025705",
      "postDate": "09/24/2020 18:07:31",
      "content": "<p>I understand the logic though: If a Radiologist can see and characterize the prognosis of decline by just looking into a baseline CT, then a properly designed model would be able to do it too because the information is there. Either we failed to extract it properly or the images are problematic and is very difficult to do it or the info is not really there.</p>",
      "rawMarkdown": "I understand the logic though: If a Radiologist can see and characterize the prognosis of decline by just looking into a baseline CT, then a properly designed model would be able to do it too because the information is there. Either we failed to extract it properly or the images are problematic and is very difficult to do it or the info is not really there.",
      "votes": null
    },
    {
      "id": "1025706",
      "postDate": "09/24/2020 18:08:22",
      "content": "<p>I've insisted on the previous 2 months with the same hope…</p>",
      "rawMarkdown": "I've insisted on the previous 2 months with the same hope...",
      "votes": null
    },
    {
      "id": "1025914",
      "postDate": "09/24/2020 21:46:13",
      "content": "<p>In general, my impression is that this task is quite specific.<br>\nFirst of all, fighting for picture technical characteristics (pixels, slices etc) has no correlation with actual image content. And really cool that the experiment showed that no correlation has been found. If some correlation would have been found, I would bet on some bias in the data sample, which one definitely should not rely on as there would be no guarantee that the test data contains the same bias.</p>\n<p>But what actually resulted the features like Lung Volume and Lung Area?<br>\nTo be honest, I think the task is quite complicated and has many internal interdependencies (and for that too few data ;) ).<br>\nAs I have mentioned before, there are clients with different dynamics of the disease --&gt; not a good idea to mix them up in creating a model.<br>\nMoreover, I see even deeper patterns in the data. There is even more variety in the dynamics of the disease caused by patient individual characteristics. Like two clients might be recovering, but one much faster than the other, because this patient has better \"input\" - higher lung volume and non-smoker and a female. Thus, I would first try to split the data to the meaningful groups with \"similar\" dynamics of the FVC inside and let image \"predict\" or \"provide\" one or several grouping criteria.</p>\n<p>As long as the final model will be kept simple (linear model, for example), one still has a lot of space to reduce the amount of data which is enough for training a model. This will allow to create more homogenous groups and train the models in these groups.</p>",
      "rawMarkdown": "In general, my impression is that this task is quite specific.\nFirst of all, fighting for picture technical characteristics (pixels, slices etc) has no correlation with actual image content. And really cool that the experiment showed that no correlation has been found. If some correlation would have been found, I would bet on some bias in the data sample, which one definitely should not rely on as there would be no guarantee that the test data contains the same bias.\n\nBut what actually resulted the features like Lung Volume and Lung Area?\nTo be honest, I think the task is quite complicated and has many internal interdependencies (and for that too few data ;) ).\nAs I have mentioned before, there are clients with different dynamics of the disease --> not a good idea to mix them up in creating a model.\nMoreover, I see even deeper patterns in the data. There is even more variety in the dynamics of the disease caused by patient individual characteristics. Like two clients might be recovering, but one much faster than the other, because this patient has better \"input\" - higher lung volume and non-smoker and a female. Thus, I would first try to split the data to the meaningful groups with \"similar\" dynamics of the FVC inside and let image \"predict\" or \"provide\" one or several grouping criteria.\n\nAs long as the final model will be kept simple (linear model, for example), one still has a lot of space to reduce the amount of data which is enough for training a model. This will allow to create more homogenous groups and train the models in these groups.",
      "votes": null
    },
    {
      "id": "1025934",
      "postDate": "09/24/2020 22:25:54",
      "content": "<p>Hello Hannah, thanks for extending the discussion.</p>\n<p>I didn't think of spliting the data by the FVC pattern because I thought I would be leaking the target information. But I guess it's like a feature engineering if we add this information just for the training set and not for the validation set. Did I get that right?</p>\n<p>My doubt was what to do for the test set, but rereading the ending insights of your notebook it's clear what you mean.</p>\n<p>I'll inform if I had any promising results.</p>",
      "rawMarkdown": "Hello Hannah, thanks for extending the discussion.\n\nI didn't think of spliting the data by the FVC pattern because I thought I would be leaking the target information. But I guess it's like a feature engineering if we add this information just for the training set and not for the validation set. Did I get that right?\n\nMy doubt was what to do for the test set, but rereading the ending insights of your notebook it's clear what you mean.\n\nI'll inform if I had any promising results.",
      "votes": null
    },
    {
      "id": "1026011",
      "postDate": "09/25/2020 01:34:46",
      "content": "<p>I agree in principle that different clusters of patients could be better analysed by different models, but my concern is that there simply isn't enough data to make generalisation possible in that case. We really could have done with more samples in this dataset…</p>",
      "rawMarkdown": "I agree in principle that different clusters of patients could be better analysed by different models, but my concern is that there simply isn't enough data to make generalisation possible in that case. We really could have done with more samples in this dataset...",
      "votes": null
    },
    {
      "id": "1026063",
      "postDate": "09/25/2020 03:58:43",
      "content": "<p><a href=\"https://www.kaggle.com/cascadenite\" target=\"_blank\">@cascadenite</a> The image features didn't work out really well (if you're thinking it would), as our team did integrate it with the model, but it resulted in poor CV, which I guess could be the indicator that the scans won't work out really good on the private LB.</p>\n<p>Even though the competition's emphasis would be make use of the scans, but technically at this point, I'm not that confident that scans would be useful.</p>\n<p>I might be wrong, but currently it feels that scans do not add up much value in improving LB/CV.</p>",
      "rawMarkdown": "cascadenite The image features didn't work out really well (if you're thinking it would), as our team did integrate it with the model, but it resulted in poor CV, which I guess could be the indicator that the scans won't work out really good on the private LB.\n\nEven though the competition's emphasis would be make use of the scans, but technically at this point, I'm not that confident that scans would be useful.\n\nI might be wrong, but currently it feels that scans do not add up much value in improving LB/CV.",
      "votes": null
    },
    {
      "id": "1026285",
      "postDate": "09/25/2020 07:41:24",
      "content": "<p>Yes, I agree. One might have also taken it as a feature engineering and adding the info to the targeting set. However, we have kind of not that much data to my taste to train a model which will be able to catch all the patterns itself, I fear. That is why I was more thinking of an approach where I overtake some tasks of the model :)</p>",
      "rawMarkdown": "Yes, I agree. One might have also taken it as a feature engineering and adding the info to the targeting set. However, we have kind of not that much data to my taste to train a model which will be able to catch all the patterns itself, I fear. That is why I was more thinking of an approach where I overtake some tasks of the model :)",
      "votes": null
    },
    {
      "id": "1026579",
      "postDate": "09/25/2020 12:28:43",
      "content": "<p>Apologies for the delayed response. I'm not convinced with your definition of lung volume, please do correct me if you meant something else. I'm going to write down mathematically what i have understood from your explaination about lung volume.</p>\n<p>Let \\( I_1, … I_n \\) denote the CT scan images for a patient \\(P\\) and \\(A_1,…,A_n\\) denote the corresponding area for segmented lung in each of the images. Let \\(l\\) denote the slice thickness for each image, then the approx. length of the lung can be obtained as \\(L = l*n\\).</p>\n<ul>\n<li><p>Your Definition of lung volume:<br>\n\\(Volume = (A_1+…+A_n)/n * L\\)</p></li>\n<li><p>Ideal definition of lung volume:<br>\n\\(Volume = {A_1}*l +…+{A_n} * l = (A_1+…+A_n)*l \\)</p></li>\n</ul>\n<p>Both the definitions are different. Does this make sense?</p>",
      "rawMarkdown": "Apologies for the delayed response. I'm not convinced with your definition of lung volume, please do correct me if you meant something else. I'm going to write down mathematically what i have understood from your explaination about lung volume.\n\nLet \\\\( I_1, ... I_n \\\\) denote the CT scan images for a patient \\\\(P\\\\) and \\\\(A_1,...,A_n\\\\) denote the corresponding area for segmented lung in each of the images. Let \\\\(l\\\\) denote the slice thickness for each image, then the approx. length of the lung can be obtained as \\\\(L = l*n\\\\).\n\n- Your Definition of lung volume:\n\\\\(Volume = (A_1+...+A_n)/n * L\\\\)\n\n- Ideal definition of lung volume:\n\\\\(Volume = {A_1}*l +...+{A_n} * l = (A_1+...+A_n)*l \\\\)\n\nBoth the definitions are different. Does this make sense?",
      "votes": null
    },
    {
      "id": "1026596",
      "postDate": "09/25/2020 12:42:52",
      "content": "<p>The two definitions you wrote down are identical?<br>\n$$ (A_1+\\ldots + A_n)/n * L = (A_1+\\ldots + A_n)/n * \\ell * n = (A_1 + \\ldots + A_n) * \\ell $$ </p>",
      "rawMarkdown": "The two definitions you wrote down are identical?\n$$ (A_1+\\ldots + A_n)/n * L = (A_1+\\ldots + A_n)/n * \\ell * n = (A_1 + \\ldots + A_n) * \\ell $$",
      "votes": null
    },
    {
      "id": "1026651",
      "postDate": "09/25/2020 13:21:31",
      "content": "<p>Thank you all for your contributions, that's exactly the technical discussion I was looking for!</p>\n<p>In summary the winning model will probably be a not-too-complex Tab-Data model, which is only using some additional features, but from Images probably only a tissue-amount/lung-volume value.</p>",
      "rawMarkdown": "Thank you all for your contributions, that's exactly the technical discussion I was looking for!\n\nIn summary the winning model will probably be a not-too-complex Tab-Data model, which is only using some additional features, but from Images probably only a tissue-amount/lung-volume value.",
      "votes": null
    },
    {
      "id": "1026665",
      "postDate": "09/25/2020 13:31:40",
      "content": "<p><a href=\"https://www.kaggle.com/tobiit\" target=\"_blank\">@tobiit</a>, thanks a ton! I completely missed that, my bad. Now I completely agree with:</p>\n<blockquote>\n  <p>volume of lung / volume of tissue = average area of lung / average area of tissue</p>\n</blockquote>\n<p>I think the quality tissue segmentation is the key here. </p>",
      "rawMarkdown": "tobiit, thanks a ton! I completely missed that, my bad. Now I completely agree with:\n\n> volume of lung / volume of tissue = average area of lung / average area of tissue\n\nI think the quality tissue segmentation is the key here.",
      "votes": null
    },
    {
      "id": "1030375",
      "postDate": "09/28/2020 16:07:52",
      "content": "<p>What <a href=\"https://www.kaggle.com/ronaldokun\" target=\"_blank\">@ronaldokun</a> said about the logic makes sense to me. But to my understanding, part of the problem is precisely because clinicians are not able to offer a reasonably certain prognosis even with imaging.</p>",
      "rawMarkdown": "What @ronaldokun said about the logic makes sense to me. But to my understanding, part of the problem is precisely because clinicians are not able to offer a reasonably certain prognosis even with imaging.",
      "votes": null
    },
    {
      "id": "1030456",
      "postDate": "09/28/2020 17:19:47",
      "content": "<p>Yes, I think the whole point of this competition is to fill this gap. It seems a lot harder for this problem judging by the results so far.</p>",
      "rawMarkdown": "Yes, I think the whole point of this competition is to fill this gap. It seems a lot harder for this problem judging by the results so far.",
      "votes": null
    },
    {
      "id": "1031748",
      "postDate": "09/29/2020 16:58:11",
      "content": "<p><a href=\"https://www.kaggle.com/abhishekgbhat\" target=\"_blank\">@abhishekgbhat</a>   CT report provided to us , has got many slices. Are they different CT reports or they part of Single CT scan report only ?</p>",
      "rawMarkdown": "abhishekgbhat   CT report provided to us , has got many slices. Are they different CT reports or they part of Single CT scan report only ?",
      "votes": null
    },
    {
      "id": "1031836",
      "postDate": "09/29/2020 17:55:42",
      "content": "<p><a href=\"https://www.kaggle.com/ronaldokun\" target=\"_blank\">@ronaldokun</a>  to your point of Percent feature. <br>\nWhen we use it for training and hence prediction we get better lb inspite of fact that it means for all patient visits we are saying patients FVC is same of expected FVC for given segment of patient,please correct me if m wrong  in contrary when we try to correct this by not including FVC then we get,worse lb </p>\n<p>.<br>\nIn short is your current score is making use of Percent in Train ?</p>",
      "rawMarkdown": "ronaldokun  to your point of Percent feature. \nWhen we use it for training and hence prediction we get better lb inspite of fact that it means for all patient visits we are saying patients FVC is same of expected FVC for given segment of patient,please correct me if m wrong  in contrary when we try to correct this by not including FVC then we get,worse lb \n \n.\nIn short is your current score is making use of Percent in Train ?",
      "votes": null
    },
    {
      "id": "1032016",
      "postDate": "09/29/2020 21:15:52",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> </p>\n<p>Yes, you are absolutely right. Our current best score makes use of percent but by the point made above, IMHO it seems conceptually wrong.</p>\n<p>But yes, we are using the percent the same way as yours in our submission because it gave way better scores in CV and LB but we are probably overfitting a lot to the Public LB.  </p>",
      "rawMarkdown": "Hello @jaideepvalani \n\nYes, you are absolutely right. Our current best score makes use of percent but by the point made above, IMHO it seems conceptually wrong.\n\nBut yes, we are using the percent the same way as yours in our submission because it gave way better scores in CV and LB but we are probably overfitting a lot to the Public LB.",
      "votes": null
    },
    {
      "id": "1032160",
      "postDate": "09/30/2020 02:27:39",
      "content": "<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> CT scan for a particular patient is a collection of slices and all these belong to the same report. So to answer your question, for each patient we have just one CT report with many slices.</p>",
      "rawMarkdown": "jaideepvalani CT scan for a particular patient is a collection of slices and all these belong to the same report. So to answer your question, for each patient we have just one CT report with many slices.",
      "votes": null
    },
    {
      "id": "1036519",
      "postDate": "10/03/2020 21:57:17",
      "content": "<p>Hi everyone,</p>\n<p>we use image-derived features, and it helped to get  +0.015 in LB above the tabular results.</p>\n<p>I manually segmented 100+ CT scans for the lung itself, and it delivers perfect results in lung segmentation.</p>\n<p>Also made ~ 40k markups for segmentation of honeycombing, reticulation, GGO/GGR but the heterogenity of the dataset… well… kind-of makes the work pointless. <br>\nThe unets are generalising pretty poorly on tissue features like reticulation or GGO, but honeycombing is not bad.</p>\n<p><img src=\"http://drkonya.com/projects/kaggle/lung3ds.jpg\" alt=\"Lung\"><br>\n<a href=\"http://drkonya.com/projects/kaggle/lung3d.gif\" target=\"_blank\">3D Lung</a></p>",
      "rawMarkdown": "Hi everyone,\n\nwe use image-derived features, and it helped to get  +0.015 in LB above the tabular results.\n\nI manually segmented 100+ CT scans for the lung itself, and it delivers perfect results in lung segmentation.\n\nAlso made ~ 40k markups for segmentation of honeycombing, reticulation, GGO/GGR but the heterogenity of the dataset... well... kind-of makes the work pointless. \nThe unets are generalising pretty poorly on tissue features like reticulation or GGO, but honeycombing is not bad.\n\n![Lung](http://drkonya.com/projects/kaggle/lung3ds.jpg)\n[3D Lung](http://drkonya.com/projects/kaggle/lung3d.gif)",
      "votes": null
    },
    {
      "id": "1037057",
      "postDate": "10/04/2020 15:00:11",
      "content": "<p>I hope that after the competition your work on the manual segmentation doesn't go to waste. I'm sure there'll be other applications that would really benefit - deserves more than + 0.015</p>",
      "rawMarkdown": "I hope that after the competition your work on the manual segmentation doesn't go to waste. I'm sure there'll be other applications that would really benefit - deserves more than + 0.015",
      "votes": null
    },
    {
      "id": "1037089",
      "postDate": "10/04/2020 15:44:02",
      "content": "<p><a href=\"https://www.kaggle.com/sandorkonya\" target=\"_blank\">@sandorkonya</a>  even if do lung seg for all the patients ct report,average volume of all the lungs should stand same,every one has got same lungs structure except that the tissues in it could be varrying but we lose that while we do segmentation,so is segmentation still useful ?</p>",
      "rawMarkdown": "sandorkonya  even if do lung seg for all the patients ct report,average volume of all the lungs should stand same,every one has got same lungs structure except that the tissues in it could be varrying but we lose that while we do segmentation,so is segmentation still useful ?",
      "votes": null
    },
    {
      "id": "1037278",
      "postDate": "10/04/2020 19:13:55",
      "content": "<p><a href=\"https://www.kaggle.com/jameschapman19\" target=\"_blank\">@jameschapman19</a>  it won't go waste… i can give back something to the community =)</p>",
      "rawMarkdown": "jameschapman19  it won't go waste... i can give back something to the community =)",
      "votes": null
    },
    {
      "id": "1037281",
      "postDate": "10/04/2020 19:18:28",
      "content": "<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> <br>\nthe average lung volume is not the same! It depends from the constitution (height, BMI, sex) of the patient… and it also depends on how far the fibrosis is developed. The fibrosis itself makes a volumen depletion of the lung volume (makes it to shrink).<br>\nWhy would we lose the \"tissue of the lung\" while we segment?</p>",
      "rawMarkdown": "jaideepvalani \nthe average lung volume is not the same! It depends from the constitution (height, BMI, sex) of the patient... and it also depends on how far the fibrosis is developed. The fibrosis itself makes a volumen depletion of the lung volume (makes it to shrink).\nWhy would we lose the \"tissue of the lung\" while we segment?",
      "votes": null
    },
    {
      "id": "1038768",
      "postDate": "10/06/2020 03:38:33",
      "content": "<p><a href=\"https://www.kaggle.com/sandorkonya\" target=\"_blank\">@sandorkonya</a>  hope you get paid off by your domain expertise..<br>\nHow can we determine height from ct … any one trying to do </p>",
      "rawMarkdown": "sandorkonya  hope you get paid off by your domain expertise..\nHow can we determine height from ct ... any one trying to do",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1017994,
      "author_name": "abhishekgbhat",
      "author_url": "",
      "post_date": "09/19/2020 11:28:17",
      "content": "<p>I have tried including the following features you mentioned:</p>\n<ul>\n<li>Slice Thickness </li>\n<li>Pixel Spacing R</li>\n<li>Pixel Spacing C<br>\nNone of the above features had any correlation with FVC or decline in FVC over time.</li>\n</ul>\n<p>But I did observe some good correlation between:</p>\n<ul>\n<li>Lung volume and FVC</li>\n<li>Average of (Tissue area)/(Lung Area) across all images for a particular patient and FVC</li>\n<li>Average tissue area and decline in FVC over time</li>\n</ul>\n<p>You can consider including some of the above features in your experiment. <br>\nAlso how do you define - image_area_cm2, slice_volume_cm3?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1022098,
          "author_name": "ChristianDenich",
          "author_url": "",
          "post_date": "09/22/2020 10:07:01",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/abhishekgbhat\" target=\"_blank\">@abhishekgbhat</a>, I just saw your post yesterday night, sorry for the delay.<br>\nThe feautes are derived from Laura's Notebook <a href=\"https://www.kaggle.com/allunia/pulmonary-dicom-preprocessing\" target=\"_blank\">pulmonary-dicom-preprocessing</a>. E.g. img_area_cm2 is derived by the resolution.</p>\n<p>How do you derive \"Average of (Tissue area)/(Lung Area)\"?<br>\nIs there a notebook available? Sounds like a lot of work using Hounsfield values &amp; segmentation.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1022812,
          "author_name": "abhishekgbhat",
          "author_url": "",
          "post_date": "09/22/2020 18:25:35",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/ChristianDenich\" target=\"_blank\">@ChristianDenich</a>, thanks for the clarification. </p>\n<blockquote>\n  <p>How do you derive \"Average of (Tissue area)/(Lung Area)\"?</p>\n</blockquote>\n<p>No, I haven't seen any notebook with this feature. From the limited research I have done, the fiberosis tissue plays an important role in the severity of the disease. So I found out a way to segment the tissues within the lung. From this segmented image I obtained the number of tissue pixels and multiplied it with Pixel Spacing R and C. This gives the tissue area in each image. Assuming you have the lung area for each image, you can get this ratio for each image. Finally, if you have this ratio for all images of a particular patient you can simply take the average. The tough part is to segment the tissues effectively. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1023338,
          "author_name": "tobiit",
          "author_url": "",
          "post_date": "09/23/2020 06:25:40",
          "content": "<p>So from a maths point of view, this should be the same as considering (Tissue volume)/(Lung volume) I guess. </p>\n<p>I considered this feature and thought it gave me a slight improvement, but after changing some hyperparameters etc. I tried again without the feature and got a higher score than with the feature. </p>\n<p>Maybe it works better with a better segmentation technique though. I just used the ranges of the HU scale for segmentation.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1024017,
          "author_name": "abhishekgbhat",
          "author_url": "",
          "post_date": "09/23/2020 15:35:13",
          "content": "<p>This is not exactly the volume because im not considering the third dimension(i.e slice thickness) here. This feature tells us the following:  For a particular patient, on an average what percent of the lung area does the tissue take up </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1024030,
          "author_name": "tobiit",
          "author_url": "",
          "post_date": "09/23/2020 15:46:04",
          "content": "<p>Yes, but to get the volume of lung, you would have to multiply the average lung area by the length of the lung (3rd dimension). Analoguously, you get the volume of tissue, when you multiply the average tissue area by the length of the lung. When you divide both things, the length cancels out, such that<br>\nvolume of lung / volume of tissue =  average area of lung / average area of tissue.<br>\nDo you agree?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1026579,
          "author_name": "abhishekgbhat",
          "author_url": "",
          "post_date": "09/25/2020 12:28:43",
          "content": "<p>Apologies for the delayed response. I'm not convinced with your definition of lung volume, please do correct me if you meant something else. I'm going to write down mathematically what i have understood from your explaination about lung volume.</p>\n<p>Let \\( I_1, … I_n \\) denote the CT scan images for a patient \\(P\\) and \\(A_1,…,A_n\\) denote the corresponding area for segmented lung in each of the images. Let \\(l\\) denote the slice thickness for each image, then the approx. length of the lung can be obtained as \\(L = l*n\\).</p>\n<ul>\n<li><p>Your Definition of lung volume:<br>\n\\(Volume = (A_1+…+A_n)/n * L\\)</p></li>\n<li><p>Ideal definition of lung volume:<br>\n\\(Volume = {A_1}*l +…+{A_n} * l = (A_1+…+A_n)*l \\)</p></li>\n</ul>\n<p>Both the definitions are different. Does this make sense?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1026596,
          "author_name": "tobiit",
          "author_url": "",
          "post_date": "09/25/2020 12:42:52",
          "content": "<p>The two definitions you wrote down are identical?<br>\n$$ (A_1+\\ldots + A_n)/n * L = (A_1+\\ldots + A_n)/n * \\ell * n = (A_1 + \\ldots + A_n) * \\ell $$ </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1026665,
          "author_name": "abhishekgbhat",
          "author_url": "",
          "post_date": "09/25/2020 13:31:40",
          "content": "<p><a href=\"https://www.kaggle.com/tobiit\" target=\"_blank\">@tobiit</a>, thanks a ton! I completely missed that, my bad. Now I completely agree with:</p>\n<blockquote>\n  <p>volume of lung / volume of tissue = average area of lung / average area of tissue</p>\n</blockquote>\n<p>I think the quality tissue segmentation is the key here. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1031748,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "09/29/2020 16:58:11",
          "content": "<p><a href=\"https://www.kaggle.com/abhishekgbhat\" target=\"_blank\">@abhishekgbhat</a>   CT report provided to us , has got many slices. Are they different CT reports or they part of Single CT scan report only ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1032160,
          "author_name": "abhishekgbhat",
          "author_url": "",
          "post_date": "09/30/2020 02:27:39",
          "content": "<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> CT scan for a particular patient is a collection of slices and all these belong to the same report. So to answer your question, for each patient we have just one CT report with many slices.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1018277,
      "author_name": "aadhavvignesh",
      "author_url": "",
      "post_date": "09/19/2020 15:15:39",
      "content": "<p><strong>What I got to know from Image Data:</strong></p>\n<p>I had already mentioned that these features didn't work:</p>\n<ul>\n<li>Lung Volume <code>in cm^3</code></li>\n<li>Lung Area <code>in cm^2</code></li>\n<li>Image Resolution/Dimensions</li>\n</ul>\n<p>Segmented lung images didn't work as expected. Poor CV/LB still persists. </p>\n<p>This has led our team to put less focus on image data, and more on the available tabular data. I unfortunately can't reveal anything beyond this, because every team would love to keep their sauce 'secret' :P</p>\n<p><strong>About LB and rankings:</strong></p>\n<p>I personally feel that the LB is <strong>NOT</strong> a good indicator of your team's results. I guess that I can safely say that our team has a stable model (my instinct says so :P), and our rankings aren't good. A shake-up is imminent, so I do feel teams should not be looking at the LB scores, but indeed should have a good CV set up.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1018351,
          "author_name": "subzeroop",
          "author_url": "",
          "post_date": "09/19/2020 16:18:50",
          "content": "<p>hey adhav does your team needs a member ? , i am currently seraching for a team , my current rank is 26</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1022140,
          "author_name": "aadhavvignesh",
          "author_url": "",
          "post_date": "09/22/2020 10:37:55",
          "content": "<p>I'm really sorry <a href=\"https://www.kaggle.com/subzeroop\" target=\"_blank\">@subzeroop</a>, our team is full. I hope you can find a team soon :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1025632,
          "author_name": "cascadenite",
          "author_url": "",
          "post_date": "09/24/2020 17:21:10",
          "content": "<p>Hi, I've been noticing too  that the images are frustratingly not helping the public LB position but I'm going to keep focussing on them anyway.  My hope  is that the CT scans must have some value and that a big shakeup will occur where those models that make good use of the scans perform better on the private LB.   I'm not looking at the technial aspects of the ct-scan (like pixel size, slice thickness) but I'm confident that the shapes, areas and textures of the slices themselves can be useful.</p>\n<p>I must admit though that this hope is more an ensemble of sunk-cost fallacy + dumb-optimism + \"what is point of this competition if not\"…. :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1025706,
          "author_name": "ronaldokun",
          "author_url": "",
          "post_date": "09/24/2020 18:08:22",
          "content": "<p>I've insisted on the previous 2 months with the same hope…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1026063,
          "author_name": "aadhavvignesh",
          "author_url": "",
          "post_date": "09/25/2020 03:58:43",
          "content": "<p><a href=\"https://www.kaggle.com/cascadenite\" target=\"_blank\">@cascadenite</a> The image features didn't work out really well (if you're thinking it would), as our team did integrate it with the model, but it resulted in poor CV, which I guess could be the indicator that the scans won't work out really good on the private LB.</p>\n<p>Even though the competition's emphasis would be make use of the scans, but technically at this point, I'm not that confident that scans would be useful.</p>\n<p>I might be wrong, but currently it feels that scans do not add up much value in improving LB/CV.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1018303,
      "author_name": "yimacs",
      "author_url": "",
      "post_date": "09/19/2020 15:32:43",
      "content": "<p>it doesn't help for CV/LB. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1018991,
      "author_name": "ajay19",
      "author_url": "",
      "post_date": "09/20/2020 06:02:46",
      "content": "<p>Tried image resolution , augmentation while training , DenseNet and Efficientnet ensembling but all in vain</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1019284,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "09/20/2020 10:28:07",
      "content": "<p>Thanks for sharing your findings. Even if they improve cv score, I wouldn't use them anyway. Since FVC is not dependent to them, it would be huge gamble to use them in your final submissions.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1019757,
      "author_name": "bigironsphere",
      "author_url": "",
      "post_date": "09/20/2020 16:45:34",
      "content": "<p>These features contain diagnostic information for the CT scanner setup, any correlation they have with lung function will be completely incidental anyway. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1020565,
      "author_name": "prk007",
      "author_url": "",
      "post_date": "09/21/2020 09:07:22",
      "content": "<p>The holdout predictions for private LB is 85% where as public LB is scored on 15% of the predictions, although features related pixel spacing, slice thickness are useless, others like lung volume, lung are useful. </p>\n<p>Wait for a huge LB shakeup once private LB is announced.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1024276,
      "author_name": "ronaldokun",
      "author_url": "",
      "post_date": "09/23/2020 18:19:18",
      "content": "<p>We tried:</p>\n<p>CenterCrop, Resampled to Pixel Spacing and SliceThickness equal to 1, Mask the Lung and added the following</p>\n<p>Pixel Voxel Statistics: Mean, STD, Kurtosis &amp; Skew Distribution<br>\nPercentage in Lung Window<br>\nChest Circunference<br>\nLung Height<br>\nLung Volume in cm3 or liters (dm3)<br>\nHU Bins Distribution from 2 to 14 bins ( Just adding the HU distribution in bins ).</p>\n<p>Any addition or combination of those items just made it worse in the LB, sometimes better CV but always worse LB.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1024639,
          "author_name": "aadhavvignesh",
          "author_url": "",
          "post_date": "09/24/2020 03:33:35",
          "content": "<p>I agree with this. We've tried all possible combinations and on integrating them with the model gave us worse LB and worse CV in almost every case.</p>\n<p>Does your team know what exactly the problem is, because at this point making use of scans seems pointless. (not meant to be conclusive, but including them is making CV/LB worse)</p>\n<p>Do you also think that the winners' solution would help the hosts if they don't use image data for the final model, as I guess the primary objective of the hosts would be to make use of the scans for determining pulmonary fibrosis?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1024676,
          "author_name": "ronaldokun",
          "author_url": "",
          "post_date": "09/24/2020 04:15:49",
          "content": "<p>We are not sure what is wrong, but we think that the release of just the baseline week of the images and also the test tabular data was a fundamental flaw in this competition. We are supposed to predict 146 weeks - time series data - and we don't have a time series data to train on.<br>\nMedical data is already extremely heterogenous, it seems we have different distributions of patients in the training set ( this notebook by <a href=\"https://www.kaggle.com/gannadolinska\" target=\"_blank\">@gannadolinska</a> shows we can consider at least 3 different distributions of patients, those which decrease lung capacity, those steady and those which get better. ) <br>\nThe tabular data has the percent feature, another controversial addition given that is highly correlated with the target, and we also don't have it in the test data. <br>\nIn an ideal case the image in the baseline would have enough information to guide the model to the different distributions of patients and the evolution of them but we failed to extract any<br>\nWe pretty much gave up using the images entirely, it was a very frustrating couple of months.  <br>\nEither someone has a bag of tricks no one else has implemented using the images or the organizers are going to pay for a baseline tabular model, overfitted and completely useless medically. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1025134,
          "author_name": "aadhavvignesh",
          "author_url": "",
          "post_date": "09/24/2020 10:53:06",
          "content": "<p>Yes, we had initially planned to extract some useful features using image data, and I even saw some teams trying to manually honeycomb lungs and trying to make it work. Eventually, everyone who had used images became frustrated due to poor results.</p>\n<p>I too fear that the organizers are going to end up with a baseline tabular model, rendering the objective of the competition to be useless.</p>\n<p>Hope we get to see a novel approach from the winners, but for now I guess most of the LB is saturated with overfitting models.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1025705,
          "author_name": "ronaldokun",
          "author_url": "",
          "post_date": "09/24/2020 18:07:31",
          "content": "<p>I understand the logic though: If a Radiologist can see and characterize the prognosis of decline by just looking into a baseline CT, then a properly designed model would be able to do it too because the information is there. Either we failed to extract it properly or the images are problematic and is very difficult to do it or the info is not really there.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1030375,
          "author_name": "douglaskgaraujo",
          "author_url": "",
          "post_date": "09/28/2020 16:07:52",
          "content": "<p>What <a href=\"https://www.kaggle.com/ronaldokun\" target=\"_blank\">@ronaldokun</a> said about the logic makes sense to me. But to my understanding, part of the problem is precisely because clinicians are not able to offer a reasonably certain prognosis even with imaging.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1030456,
          "author_name": "ronaldokun",
          "author_url": "",
          "post_date": "09/28/2020 17:19:47",
          "content": "<p>Yes, I think the whole point of this competition is to fill this gap. It seems a lot harder for this problem judging by the results so far.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1031836,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "09/29/2020 17:55:42",
          "content": "<p><a href=\"https://www.kaggle.com/ronaldokun\" target=\"_blank\">@ronaldokun</a>  to your point of Percent feature. <br>\nWhen we use it for training and hence prediction we get better lb inspite of fact that it means for all patient visits we are saying patients FVC is same of expected FVC for given segment of patient,please correct me if m wrong  in contrary when we try to correct this by not including FVC then we get,worse lb </p>\n<p>.<br>\nIn short is your current score is making use of Percent in Train ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1032016,
          "author_name": "ronaldokun",
          "author_url": "",
          "post_date": "09/29/2020 21:15:52",
          "content": "<p>Hello <a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> </p>\n<p>Yes, you are absolutely right. Our current best score makes use of percent but by the point made above, IMHO it seems conceptually wrong.</p>\n<p>But yes, we are using the percent the same way as yours in our submission because it gave way better scores in CV and LB but we are probably overfitting a lot to the Public LB.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1025914,
      "author_name": "gannadolinska",
      "author_url": "",
      "post_date": "09/24/2020 21:46:13",
      "content": "<p>In general, my impression is that this task is quite specific.<br>\nFirst of all, fighting for picture technical characteristics (pixels, slices etc) has no correlation with actual image content. And really cool that the experiment showed that no correlation has been found. If some correlation would have been found, I would bet on some bias in the data sample, which one definitely should not rely on as there would be no guarantee that the test data contains the same bias.</p>\n<p>But what actually resulted the features like Lung Volume and Lung Area?<br>\nTo be honest, I think the task is quite complicated and has many internal interdependencies (and for that too few data ;) ).<br>\nAs I have mentioned before, there are clients with different dynamics of the disease --&gt; not a good idea to mix them up in creating a model.<br>\nMoreover, I see even deeper patterns in the data. There is even more variety in the dynamics of the disease caused by patient individual characteristics. Like two clients might be recovering, but one much faster than the other, because this patient has better \"input\" - higher lung volume and non-smoker and a female. Thus, I would first try to split the data to the meaningful groups with \"similar\" dynamics of the FVC inside and let image \"predict\" or \"provide\" one or several grouping criteria.</p>\n<p>As long as the final model will be kept simple (linear model, for example), one still has a lot of space to reduce the amount of data which is enough for training a model. This will allow to create more homogenous groups and train the models in these groups.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1025934,
          "author_name": "ronaldokun",
          "author_url": "",
          "post_date": "09/24/2020 22:25:54",
          "content": "<p>Hello Hannah, thanks for extending the discussion.</p>\n<p>I didn't think of spliting the data by the FVC pattern because I thought I would be leaking the target information. But I guess it's like a feature engineering if we add this information just for the training set and not for the validation set. Did I get that right?</p>\n<p>My doubt was what to do for the test set, but rereading the ending insights of your notebook it's clear what you mean.</p>\n<p>I'll inform if I had any promising results.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1026011,
          "author_name": "bigironsphere",
          "author_url": "",
          "post_date": "09/25/2020 01:34:46",
          "content": "<p>I agree in principle that different clusters of patients could be better analysed by different models, but my concern is that there simply isn't enough data to make generalisation possible in that case. We really could have done with more samples in this dataset…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1026285,
          "author_name": "gannadolinska",
          "author_url": "",
          "post_date": "09/25/2020 07:41:24",
          "content": "<p>Yes, I agree. One might have also taken it as a feature engineering and adding the info to the targeting set. However, we have kind of not that much data to my taste to train a model which will be able to catch all the patterns itself, I fear. That is why I was more thinking of an approach where I overtake some tasks of the model :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1026651,
      "author_name": "ChristianDenich",
      "author_url": "",
      "post_date": "09/25/2020 13:21:31",
      "content": "<p>Thank you all for your contributions, that's exactly the technical discussion I was looking for!</p>\n<p>In summary the winning model will probably be a not-too-complex Tab-Data model, which is only using some additional features, but from Images probably only a tissue-amount/lung-volume value.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1036519,
      "author_name": "sandorkonya",
      "author_url": "",
      "post_date": "10/03/2020 21:57:17",
      "content": "<p>Hi everyone,</p>\n<p>we use image-derived features, and it helped to get  +0.015 in LB above the tabular results.</p>\n<p>I manually segmented 100+ CT scans for the lung itself, and it delivers perfect results in lung segmentation.</p>\n<p>Also made ~ 40k markups for segmentation of honeycombing, reticulation, GGO/GGR but the heterogenity of the dataset… well… kind-of makes the work pointless. <br>\nThe unets are generalising pretty poorly on tissue features like reticulation or GGO, but honeycombing is not bad.</p>\n<p><img src=\"http://drkonya.com/projects/kaggle/lung3ds.jpg\" alt=\"Lung\"><br>\n<a href=\"http://drkonya.com/projects/kaggle/lung3d.gif\" target=\"_blank\">3D Lung</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1037057,
          "author_name": "jameschapman19",
          "author_url": "",
          "post_date": "10/04/2020 15:00:11",
          "content": "<p>I hope that after the competition your work on the manual segmentation doesn't go to waste. I'm sure there'll be other applications that would really benefit - deserves more than + 0.015</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1037089,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "10/04/2020 15:44:02",
          "content": "<p><a href=\"https://www.kaggle.com/sandorkonya\" target=\"_blank\">@sandorkonya</a>  even if do lung seg for all the patients ct report,average volume of all the lungs should stand same,every one has got same lungs structure except that the tissues in it could be varrying but we lose that while we do segmentation,so is segmentation still useful ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1037278,
          "author_name": "sandorkonya",
          "author_url": "",
          "post_date": "10/04/2020 19:13:55",
          "content": "<p><a href=\"https://www.kaggle.com/jameschapman19\" target=\"_blank\">@jameschapman19</a>  it won't go waste… i can give back something to the community =)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1037281,
          "author_name": "sandorkonya",
          "author_url": "",
          "post_date": "10/04/2020 19:18:28",
          "content": "<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> <br>\nthe average lung volume is not the same! It depends from the constitution (height, BMI, sex) of the patient… and it also depends on how far the fibrosis is developed. The fibrosis itself makes a volumen depletion of the lung volume (makes it to shrink).<br>\nWhy would we lose the \"tissue of the lung\" while we segment?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1038768,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "10/06/2020 03:38:33",
          "content": "<p><a href=\"https://www.kaggle.com/sandorkonya\" target=\"_blank\">@sandorkonya</a>  hope you get paid off by your domain expertise..<br>\nHow can we determine height from ct … any one trying to do </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1018139,
      "author_name": "subzeroop",
      "author_url": "",
      "post_date": "09/19/2020 13:22:50",
      "content": "<p>i agree these features didnt helped me to get a better CV/LB .\nbtw would you like to merge team with me as i dont have a team ,</p>",
      "votes": null,
      "replies": [
        {
          "id": 1022100,
          "author_name": "ChristianDenich",
          "author_url": "",
          "post_date": "09/22/2020 10:08:33",
          "content": "<p>Thank you for the feedback, that supports our findings.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1017848": "Dear fellow Kagglers,\n\nmy team & I have tried various experiments with **newly derived features from Image data**, like mentioned in Laura's [Notebook: pulmonary-dicom-preprocessing.](https://www.kaggle.com/allunia/pulmonary-dicom-preprocessing)\n\nSome of those features were:\n\n- img resolution\n- slice-thickness\n- pixelspacing_r\n- pixelspacing_c\n- pixelspacing_area\n- window_width\n- image_area_cm2\n- slice_volume_cm3\n\nOur current finding is, that those features sadly DO NOT improve CV/LB.\n\nDo you have share the same experience & get similar results, or are any of those features working for you?\nDid I miss important features?\n\nTo be more precise:\nThe above mentioned features do NOT have any causal relationship with the actual medical condition or prognosis of the patients. It's basically a test if there is a correlation: we couldn't find any. Even if there is a correlation, this would only help in the competition, but obviously not in the real-world examples.\n\nUsing features mentioned by @aadhavvignesh and @abhishekgbhat **might** have a causal relationship with the actual prognosis and the FVC:\n- Lung Volume in cm^3\n- Lung Area in cm^2\n- Average of (Tissue area)/(Lung Area) across all images for a particular patient and FVC\n\n\nThank you all for adding some insights on fighting Pulmonary Fibrosis!\n\n@tanulsingh077 feel free to add!",
    "1017994": "I have tried including the following features you mentioned:\n- Slice Thickness \n- Pixel Spacing R\n- Pixel Spacing C\nNone of the above features had any correlation with FVC or decline in FVC over time.\n\nBut I did observe some good correlation between:\n- Lung volume and FVC\n- Average of (Tissue area)/(Lung Area) across all images for a particular patient and FVC\n- Average tissue area and decline in FVC over time\n\nYou can consider including some of the above features in your experiment. \nAlso how do you define - image_area_cm2, slice_volume_cm3?",
    "1018139": "i agree these features didnt helped me to get a better CV/LB .\nbtw would you like to merge team with me as i dont have a team ,",
    "1018277": "**What I got to know from Image Data:**\n\nI had already mentioned that these features didn't work:\n\n- Lung Volume `in cm^3`\n- Lung Area `in cm^2`\n- Image Resolution/Dimensions\n\nSegmented lung images didn't work as expected. Poor CV/LB still persists. \n\nThis has led our team to put less focus on image data, and more on the available tabular data. I unfortunately can't reveal anything beyond this, because every team would love to keep their sauce 'secret' :P\n\n**About LB and rankings:**\n\nI personally feel that the LB is **NOT** a good indicator of your team's results. I guess that I can safely say that our team has a stable model (my instinct says so :P), and our rankings aren't good. A shake-up is imminent, so I do feel teams should not be looking at the LB scores, but indeed should have a good CV set up.",
    "1018303": "it doesn't help for CV/LB.",
    "1018351": "hey adhav does your team needs a member ? , i am currently seraching for a team , my current rank is 26",
    "1018991": "Tried image resolution , augmentation while training , DenseNet and Efficientnet ensembling but all in vain",
    "1019284": "Thanks for sharing your findings. Even if they improve cv score, I wouldn't use them anyway. Since FVC is not dependent to them, it would be huge gamble to use them in your final submissions.",
    "1019757": "These features contain diagnostic information for the CT scanner setup, any correlation they have with lung function will be completely incidental anyway.",
    "1020565": "The holdout predictions for private LB is 85% where as public LB is scored on 15% of the predictions, although features related pixel spacing, slice thickness are useless, others like lung volume, lung are useful. \n\nWait for a huge LB shakeup once private LB is announced.",
    "1022098": "Hi @abhishekgbhat, I just saw your post yesterday night, sorry for the delay.\nThe feautes are derived from Laura's Notebook [pulmonary-dicom-preprocessing](https://www.kaggle.com/allunia/pulmonary-dicom-preprocessing). E.g. img_area_cm2 is derived by the resolution.\n\nHow do you derive \"Average of (Tissue area)/(Lung Area)\"?\nIs there a notebook available? Sounds like a lot of work using Hounsfield values & segmentation.",
    "1022100": "Thank you for the feedback, that supports our findings.",
    "1022140": "I'm really sorry @subzeroop, our team is full. I hope you can find a team soon :)",
    "1022812": "Hey @ChristianDenich, thanks for the clarification. \n\n> How do you derive \"Average of (Tissue area)/(Lung Area)\"?\n\nNo, I haven't seen any notebook with this feature. From the limited research I have done, the fiberosis tissue plays an important role in the severity of the disease. So I found out a way to segment the tissues within the lung. From this segmented image I obtained the number of tissue pixels and multiplied it with Pixel Spacing R and C. This gives the tissue area in each image. Assuming you have the lung area for each image, you can get this ratio for each image. Finally, if you have this ratio for all images of a particular patient you can simply take the average. The tough part is to segment the tissues effectively.",
    "1023338": "So from a maths point of view, this should be the same as considering (Tissue volume)/(Lung volume) I guess. \n\nI considered this feature and thought it gave me a slight improvement, but after changing some hyperparameters etc. I tried again without the feature and got a higher score than with the feature. \n\nMaybe it works better with a better segmentation technique though. I just used the ranges of the HU scale for segmentation.",
    "1024017": "This is not exactly the volume because im not considering the third dimension(i.e slice thickness) here. This feature tells us the following:  For a particular patient, on an average what percent of the lung area does the tissue take up",
    "1024030": "Yes, but to get the volume of lung, you would have to multiply the average lung area by the length of the lung (3rd dimension). Analoguously, you get the volume of tissue, when you multiply the average tissue area by the length of the lung. When you divide both things, the length cancels out, such that\nvolume of lung / volume of tissue =  average area of lung / average area of tissue.\nDo you agree?",
    "1024276": "We tried:\n\nCenterCrop, Resampled to Pixel Spacing and SliceThickness equal to 1, Mask the Lung and added the following\n\nPixel Voxel Statistics: Mean, STD, Kurtosis & Skew Distribution\nPercentage in Lung Window\nChest Circunference\nLung Height\nLung Volume in cm3 or liters (dm3)\nHU Bins Distribution from 2 to 14 bins ( Just adding the HU distribution in bins ).\n\nAny addition or combination of those items just made it worse in the LB, sometimes better CV but always worse LB.",
    "1024639": "I agree with this. We've tried all possible combinations and on integrating them with the model gave us worse LB and worse CV in almost every case.\n\nDoes your team know what exactly the problem is, because at this point making use of scans seems pointless. (not meant to be conclusive, but including them is making CV/LB worse)\n\nDo you also think that the winners' solution would help the hosts if they don't use image data for the final model, as I guess the primary objective of the hosts would be to make use of the scans for determining pulmonary fibrosis?",
    "1024676": "We are not sure what is wrong, but we think that the release of just the baseline week of the images and also the test tabular data was a fundamental flaw in this competition. We are supposed to predict 146 weeks - time series data - and we don't have a time series data to train on.\n\nMedical data is already extremely heterogenous, it seems we have different distributions of patients in the training set ( this notebook by @gannadolinska shows we can consider at least 3 different distributions of patients, those which decrease lung capacity, those steady and those which get better. ) \n\nThe tabular data has the percent feature, another controversial addition given that is highly correlated with the target, and we also don't have it in the test data. \n\nIn an ideal case the image in the baseline would have enough information to guide the model to the different distributions of patients and the evolution of them but we failed to extract any\n\n We pretty much gave up using the images entirely, it was a very frustrating couple of months.  \n\nEither someone has a bag of tricks no one else has implemented using the images or the organizers are going to pay for a baseline tabular model, overfitted and completely useless medically.",
    "1025134": "Yes, we had initially planned to extract some useful features using image data, and I even saw some teams trying to manually honeycomb lungs and trying to make it work. Eventually, everyone who had used images became frustrated due to poor results.\n\nI too fear that the organizers are going to end up with a baseline tabular model, rendering the objective of the competition to be useless.\n\nHope we get to see a novel approach from the winners, but for now I guess most of the LB is saturated with overfitting models.",
    "1025632": "Hi, I've been noticing too  that the images are frustratingly not helping the public LB position but I'm going to keep focussing on them anyway.  My hope  is that the CT scans must have some value and that a big shakeup will occur where those models that make good use of the scans perform better on the private LB.   I'm not looking at the technial aspects of the ct-scan (like pixel size, slice thickness) but I'm confident that the shapes, areas and textures of the slices themselves can be useful.\n\nI must admit though that this hope is more an ensemble of sunk-cost fallacy + dumb-optimism + \"what is point of this competition if not\".... :)",
    "1025705": "I understand the logic though: If a Radiologist can see and characterize the prognosis of decline by just looking into a baseline CT, then a properly designed model would be able to do it too because the information is there. Either we failed to extract it properly or the images are problematic and is very difficult to do it or the info is not really there.",
    "1025706": "I've insisted on the previous 2 months with the same hope...",
    "1025914": "In general, my impression is that this task is quite specific.\nFirst of all, fighting for picture technical characteristics (pixels, slices etc) has no correlation with actual image content. And really cool that the experiment showed that no correlation has been found. If some correlation would have been found, I would bet on some bias in the data sample, which one definitely should not rely on as there would be no guarantee that the test data contains the same bias.\n\nBut what actually resulted the features like Lung Volume and Lung Area?\nTo be honest, I think the task is quite complicated and has many internal interdependencies (and for that too few data ;) ).\nAs I have mentioned before, there are clients with different dynamics of the disease --> not a good idea to mix them up in creating a model.\nMoreover, I see even deeper patterns in the data. There is even more variety in the dynamics of the disease caused by patient individual characteristics. Like two clients might be recovering, but one much faster than the other, because this patient has better \"input\" - higher lung volume and non-smoker and a female. Thus, I would first try to split the data to the meaningful groups with \"similar\" dynamics of the FVC inside and let image \"predict\" or \"provide\" one or several grouping criteria.\n\nAs long as the final model will be kept simple (linear model, for example), one still has a lot of space to reduce the amount of data which is enough for training a model. This will allow to create more homogenous groups and train the models in these groups.",
    "1025934": "Hello Hannah, thanks for extending the discussion.\n\nI didn't think of spliting the data by the FVC pattern because I thought I would be leaking the target information. But I guess it's like a feature engineering if we add this information just for the training set and not for the validation set. Did I get that right?\n\nMy doubt was what to do for the test set, but rereading the ending insights of your notebook it's clear what you mean.\n\nI'll inform if I had any promising results.",
    "1026011": "I agree in principle that different clusters of patients could be better analysed by different models, but my concern is that there simply isn't enough data to make generalisation possible in that case. We really could have done with more samples in this dataset...",
    "1026063": "cascadenite The image features didn't work out really well (if you're thinking it would), as our team did integrate it with the model, but it resulted in poor CV, which I guess could be the indicator that the scans won't work out really good on the private LB.\n\nEven though the competition's emphasis would be make use of the scans, but technically at this point, I'm not that confident that scans would be useful.\n\nI might be wrong, but currently it feels that scans do not add up much value in improving LB/CV.",
    "1026285": "Yes, I agree. One might have also taken it as a feature engineering and adding the info to the targeting set. However, we have kind of not that much data to my taste to train a model which will be able to catch all the patterns itself, I fear. That is why I was more thinking of an approach where I overtake some tasks of the model :)",
    "1026579": "Apologies for the delayed response. I'm not convinced with your definition of lung volume, please do correct me if you meant something else. I'm going to write down mathematically what i have understood from your explaination about lung volume.\n\nLet \\\\( I_1, ... I_n \\\\) denote the CT scan images for a patient \\\\(P\\\\) and \\\\(A_1,...,A_n\\\\) denote the corresponding area for segmented lung in each of the images. Let \\\\(l\\\\) denote the slice thickness for each image, then the approx. length of the lung can be obtained as \\\\(L = l*n\\\\).\n\n- Your Definition of lung volume:\n\\\\(Volume = (A_1+...+A_n)/n * L\\\\)\n\n- Ideal definition of lung volume:\n\\\\(Volume = {A_1}*l +...+{A_n} * l = (A_1+...+A_n)*l \\\\)\n\nBoth the definitions are different. Does this make sense?",
    "1026596": "The two definitions you wrote down are identical?\n$$ (A_1+\\ldots + A_n)/n * L = (A_1+\\ldots + A_n)/n * \\ell * n = (A_1 + \\ldots + A_n) * \\ell $$",
    "1026651": "Thank you all for your contributions, that's exactly the technical discussion I was looking for!\n\nIn summary the winning model will probably be a not-too-complex Tab-Data model, which is only using some additional features, but from Images probably only a tissue-amount/lung-volume value.",
    "1026665": "tobiit, thanks a ton! I completely missed that, my bad. Now I completely agree with:\n\n> volume of lung / volume of tissue = average area of lung / average area of tissue\n\nI think the quality tissue segmentation is the key here.",
    "1030375": "What @ronaldokun said about the logic makes sense to me. But to my understanding, part of the problem is precisely because clinicians are not able to offer a reasonably certain prognosis even with imaging.",
    "1030456": "Yes, I think the whole point of this competition is to fill this gap. It seems a lot harder for this problem judging by the results so far.",
    "1031748": "abhishekgbhat   CT report provided to us , has got many slices. Are they different CT reports or they part of Single CT scan report only ?",
    "1031836": "ronaldokun  to your point of Percent feature. \nWhen we use it for training and hence prediction we get better lb inspite of fact that it means for all patient visits we are saying patients FVC is same of expected FVC for given segment of patient,please correct me if m wrong  in contrary when we try to correct this by not including FVC then we get,worse lb \n \n.\nIn short is your current score is making use of Percent in Train ?",
    "1032016": "Hello @jaideepvalani \n\nYes, you are absolutely right. Our current best score makes use of percent but by the point made above, IMHO it seems conceptually wrong.\n\nBut yes, we are using the percent the same way as yours in our submission because it gave way better scores in CV and LB but we are probably overfitting a lot to the Public LB.",
    "1032160": "jaideepvalani CT scan for a particular patient is a collection of slices and all these belong to the same report. So to answer your question, for each patient we have just one CT report with many slices.",
    "1036519": "Hi everyone,\n\nwe use image-derived features, and it helped to get  +0.015 in LB above the tabular results.\n\nI manually segmented 100+ CT scans for the lung itself, and it delivers perfect results in lung segmentation.\n\nAlso made ~ 40k markups for segmentation of honeycombing, reticulation, GGO/GGR but the heterogenity of the dataset... well... kind-of makes the work pointless. \nThe unets are generalising pretty poorly on tissue features like reticulation or GGO, but honeycombing is not bad.\n\n![Lung](http://drkonya.com/projects/kaggle/lung3ds.jpg)\n[3D Lung](http://drkonya.com/projects/kaggle/lung3d.gif)",
    "1037057": "I hope that after the competition your work on the manual segmentation doesn't go to waste. I'm sure there'll be other applications that would really benefit - deserves more than + 0.015",
    "1037089": "sandorkonya  even if do lung seg for all the patients ct report,average volume of all the lungs should stand same,every one has got same lungs structure except that the tissues in it could be varrying but we lose that while we do segmentation,so is segmentation still useful ?",
    "1037278": "jameschapman19  it won't go waste... i can give back something to the community =)",
    "1037281": "jaideepvalani \nthe average lung volume is not the same! It depends from the constitution (height, BMI, sex) of the patient... and it also depends on how far the fibrosis is developed. The fibrosis itself makes a volumen depletion of the lung volume (makes it to shrink).\nWhy would we lose the \"tissue of the lung\" while we segment?",
    "1038768": "sandorkonya  hope you get paid off by your domain expertise..\nHow can we determine height from ct ... any one trying to do"
  },
  "source": "meta"
}