{
  "id": 166123,
  "title": "Domain expert's insights #2",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/166123",
  "author_name": "",
  "post_date": "2020-07-11T20:54:31.096851700Z",
  "votes": 130,
  "comment_count": 38,
  "views": 0,
  "content": "<p>If you haven't done yet, please read my first insight <strong><a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165727\">here</a></strong>, in this one</p>\n\n<h3>i'll show you how to get more baseline variables that had been shown to determine the progression of the disease!</h3>\n\n<p>As mentioned in the first insight, almost every approach that derives features from the lung tissue is based on lung tissue segmentation, so I assume you have your CT images already segmented into a lung segmentation mask &amp; background. You are now diligently working on generating features based on the images but you get a plethora of features and ... well only 178 patient... </p>\n\n<p>Before you get highly overfitted to the frustration (pun intended), i will try to give you some hints for manual feature engineering.</p>\n\n<p><img src=\"http://drkonya.com/projects/kaggle/fibrosis_6.png\" alt=\"slices\">\n*image courtesy of <a href=\"https://www.kaggle.com/sainatarajan7\"> Sai</a></p>\n\n<h3>First some theory:</h3>\n\n<p>&gt; Restrictive disorders are characterised by a loss in lung volume. [...] This occurs in pulmonary fibrosis, pleural disease, chest wall disorders (kyphoscoliosis), neuromuscular disorders, pneumonectomy, pulmonary oedema and obesity, to name a few.  - from <a href=\"https://breathe.ersjournals.com/content/8/3/232#ref-7\">here</a></p>\n\n<p>These disorders show themselves through the changes in spirometric (and many other) tests.</p>\n\n<h3>Results obtained from lung function tests have no meaning -</h3>\n\n<p>unless they are compared against reference values (predicted values) , and exactly this is the way we can tell that some values are reduced!</p>\n\n<p><strong>Forced vital capacity (FVC)</strong>,  - the maximum amount of air that can be exhaled when blowing out as fast as possible - is given in mililitre and as a percentage in this contest, the percentage is the percentage of the predicted value. \nIn adults, age, height, sex and race are the main determinants of the reference values for spirometric measurements.</p>\n\n<p>&gt; A restrictive ventilatory defect, defined by a reduction in static (TLC) and/or operating (VC) lung volumes, is typical in patients with IPF as in other ILDs. Reduction of lung compliance is key to restriction because both chest wall compliance and respiratory muscle strength, as assessed by measurements of transdiaphragmatic pressure and maximal inspiratory pressure at the mouth , are mostly preserved.\nRestriction is often absent at the time of diagnosis. In 96 patients with biopsy-confirmed IPF, forced vital capacity (FVC) ranged from 26% to 112% pred, while TLC ranged from 42% to 125% pred. In recent clinical trials, mean FVC was close to 80% pred, consistent with half of patients having normal operating volumes. These elements indicate poor sensitivity of lung volume measurements for the diagnosis of IPF. Although restriction of operating lung volumes is consistently associated with an increased risk of death, it correlates weakly with dyspnoea or an altered quality of life in IPF, consistent with other physiological alterations also playing key roles in clinical expression of the disease. - from <a href=\"https://err.ersjournals.com/content/27/147/170062\">here</a></p>\n\n<p>(--&gt; to summarise the above text: chest wall is not important, only lung parenchyma, FVC is thought not to be a good determinant for the DIAGNOSIS - not the follow up!)</p>\n\n<p>If we assume that in this competition a european cohort had been diagnosed (see OSIC website), we may use the <a href=\"https://erj.ersjournals.com/content/31/3/687.2\">spirometric references of the European Coal and Steel Community</a>  through which - since we already know the gender and age of the patient -  <strong>we can derive the height of the patient</strong> . </p>\n\n<p><strong>Body mass index (BMI)</strong> was furthermore investigated as one of the determinants of the progression (see previous insight) as baseline characteristic and we may try to make assumptions on this, see <a href=\"https://pubs.rsna.org/doi/abs/10.1148/radiol.2283020095?uritype=cgi&amp;journalCode=radiology\">this study</a> or <a href=\"http://www.mipg.upenn.edu/yubing/Tong_PLOS2.pdf\">this one</a>, but the most simple approaches are usually the best:</p>\n\n<h3>By determining the chest area,</h3>\n\n<p>we can categorize the patients into different BMI categories - <a href=\"https://www.researchgate.net/publication/51461844_Direct_chest_area_measurement_A_potential_anthropometric_replacement_for_BMI_to_inform_cardiac_CT_dose_parameters\">see reference here</a>. Moreover the chest circumference showed also correlation to BMI.  </p>\n\n<p>You may ask where should you determine the chest area... that is relatively simple to standardize if you have the lung segmented already: take the position of the first and the last image where lung is segmented and get the midposition. Chose the slice nearest to this position for your chest area. Instead of this, you may measure at 3 distinct levels (also derived from the topmost and bottom points) and make an average for your value.</p>\n\n<p><img src=\"http://drkonya.com/projects/kaggle/fibrosis_5.png\" alt=\"midposition\"></p>\n\n<p>Now we have: <br>\n- total lung volume (see <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165727\">previous insight</a>)\n- height of the patient\n- BMI (estimate) of the patient\n- chest circumference as standalone criteria</p>\n\n<p>In the next insight i will write about vessels and airways and their possible role in the progression determination.</p>\n\n<p><img src=\"http://drkonya.com/projects/kaggle/inspi.jpg\" alt=\"inspiration\"></p>",
  "messages": [
    {
      "id": "925140",
      "postDate": "07/11/2020 20:54:31",
      "content": "<p>If you haven't done yet, please read my first insight <strong><a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165727\">here</a></strong>, in this one</p>\n\n<h3>i'll show you how to get more baseline variables that had been shown to determine the progression of the disease!</h3>\n\n<p>As mentioned in the first insight, almost every approach that derives features from the lung tissue is based on lung tissue segmentation, so I assume you have your CT images already segmented into a lung segmentation mask &amp; background. You are now diligently working on generating features based on the images but you get a plethora of features and ... well only 178 patient... </p>\n\n<p>Before you get highly overfitted to the frustration (pun intended), i will try to give you some hints for manual feature engineering.</p>\n\n<p><img src=\"http://drkonya.com/projects/kaggle/fibrosis_6.png\" alt=\"slices\">\n*image courtesy of <a href=\"https://www.kaggle.com/sainatarajan7\"> Sai</a></p>\n\n<h3>First some theory:</h3>\n\n<p>&gt; Restrictive disorders are characterised by a loss in lung volume. [...] This occurs in pulmonary fibrosis, pleural disease, chest wall disorders (kyphoscoliosis), neuromuscular disorders, pneumonectomy, pulmonary oedema and obesity, to name a few.  - from <a href=\"https://breathe.ersjournals.com/content/8/3/232#ref-7\">here</a></p>\n\n<p>These disorders show themselves through the changes in spirometric (and many other) tests.</p>\n\n<h3>Results obtained from lung function tests have no meaning -</h3>\n\n<p>unless they are compared against reference values (predicted values) , and exactly this is the way we can tell that some values are reduced!</p>\n\n<p><strong>Forced vital capacity (FVC)</strong>,  - the maximum amount of air that can be exhaled when blowing out as fast as possible - is given in mililitre and as a percentage in this contest, the percentage is the percentage of the predicted value. \nIn adults, age, height, sex and race are the main determinants of the reference values for spirometric measurements.</p>\n\n<p>&gt; A restrictive ventilatory defect, defined by a reduction in static (TLC) and/or operating (VC) lung volumes, is typical in patients with IPF as in other ILDs. Reduction of lung compliance is key to restriction because both chest wall compliance and respiratory muscle strength, as assessed by measurements of transdiaphragmatic pressure and maximal inspiratory pressure at the mouth , are mostly preserved.\nRestriction is often absent at the time of diagnosis. In 96 patients with biopsy-confirmed IPF, forced vital capacity (FVC) ranged from 26% to 112% pred, while TLC ranged from 42% to 125% pred. In recent clinical trials, mean FVC was close to 80% pred, consistent with half of patients having normal operating volumes. These elements indicate poor sensitivity of lung volume measurements for the diagnosis of IPF. Although restriction of operating lung volumes is consistently associated with an increased risk of death, it correlates weakly with dyspnoea or an altered quality of life in IPF, consistent with other physiological alterations also playing key roles in clinical expression of the disease. - from <a href=\"https://err.ersjournals.com/content/27/147/170062\">here</a></p>\n\n<p>(--&gt; to summarise the above text: chest wall is not important, only lung parenchyma, FVC is thought not to be a good determinant for the DIAGNOSIS - not the follow up!)</p>\n\n<p>If we assume that in this competition a european cohort had been diagnosed (see OSIC website), we may use the <a href=\"https://erj.ersjournals.com/content/31/3/687.2\">spirometric references of the European Coal and Steel Community</a>  through which - since we already know the gender and age of the patient -  <strong>we can derive the height of the patient</strong> . </p>\n\n<p><strong>Body mass index (BMI)</strong> was furthermore investigated as one of the determinants of the progression (see previous insight) as baseline characteristic and we may try to make assumptions on this, see <a href=\"https://pubs.rsna.org/doi/abs/10.1148/radiol.2283020095?uritype=cgi&amp;journalCode=radiology\">this study</a> or <a href=\"http://www.mipg.upenn.edu/yubing/Tong_PLOS2.pdf\">this one</a>, but the most simple approaches are usually the best:</p>\n\n<h3>By determining the chest area,</h3>\n\n<p>we can categorize the patients into different BMI categories - <a href=\"https://www.researchgate.net/publication/51461844_Direct_chest_area_measurement_A_potential_anthropometric_replacement_for_BMI_to_inform_cardiac_CT_dose_parameters\">see reference here</a>. Moreover the chest circumference showed also correlation to BMI.  </p>\n\n<p>You may ask where should you determine the chest area... that is relatively simple to standardize if you have the lung segmented already: take the position of the first and the last image where lung is segmented and get the midposition. Chose the slice nearest to this position for your chest area. Instead of this, you may measure at 3 distinct levels (also derived from the topmost and bottom points) and make an average for your value.</p>\n\n<p><img src=\"http://drkonya.com/projects/kaggle/fibrosis_5.png\" alt=\"midposition\"></p>\n\n<p>Now we have: <br>\n- total lung volume (see <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165727\">previous insight</a>)\n- height of the patient\n- BMI (estimate) of the patient\n- chest circumference as standalone criteria</p>\n\n<p>In the next insight i will write about vessels and airways and their possible role in the progression determination.</p>\n\n<p><img src=\"http://drkonya.com/projects/kaggle/inspi.jpg\" alt=\"inspiration\"></p>",
      "rawMarkdown": "If you haven't done yet, please read my first insight **[here](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165727)**, in this one\n###i'll show you how to get more baseline variables that had been shown to determine the progression of the disease!\n\nAs mentioned in the first insight, almost every approach that derives features from the lung tissue is based on lung tissue segmentation, so I assume you have your CT images already segmented into a lung segmentation mask &amp; background. You are now diligently working on generating features based on the images but you get a plethora of features and ... well only 178 patient... \n\nBefore you get highly overfitted to the frustration (pun intended), i will try to give you some hints for manual feature engineering.\n\n![slices](http://drkonya.com/projects/kaggle/fibrosis_6.png)\n*image courtesy of [ Sai](https://www.kaggle.com/sainatarajan7)\n\n###First some theory: \n\n&gt; Restrictive disorders are characterised by a loss in lung volume. [...] This occurs in pulmonary fibrosis, pleural disease, chest wall disorders (kyphoscoliosis), neuromuscular disorders, pneumonectomy, pulmonary oedema and obesity, to name a few.  - from [here](https://breathe.ersjournals.com/content/8/3/232#ref-7)\n\nThese disorders show themselves through the changes in spirometric (and many other) tests.\n\n###Results obtained from lung function tests have no meaning -\n\nunless they are compared against reference values (predicted values) , and exactly this is the way we can tell that some values are reduced!\n\n **Forced vital capacity (FVC)**,  - the maximum amount of air that can be exhaled when blowing out as fast as possible - is given in mililitre and as a percentage in this contest, the percentage is the percentage of the predicted value. \nIn adults, age, height, sex and race are the main determinants of the reference values for spirometric measurements.\n\n&gt; A restrictive ventilatory defect, defined by a reduction in static (TLC) and/or operating (VC) lung volumes, is typical in patients with IPF as in other ILDs. Reduction of lung compliance is key to restriction because both chest wall compliance and respiratory muscle strength, as assessed by measurements of transdiaphragmatic pressure and maximal inspiratory pressure at the mouth , are mostly preserved.\nRestriction is often absent at the time of diagnosis. In 96 patients with biopsy-confirmed IPF, forced vital capacity (FVC) ranged from 26% to 112% pred, while TLC ranged from 42% to 125% pred. In recent clinical trials, mean FVC was close to 80% pred, consistent with half of patients having normal operating volumes. These elements indicate poor sensitivity of lung volume measurements for the diagnosis of IPF. Although restriction of operating lung volumes is consistently associated with an increased risk of death, it correlates weakly with dyspnoea or an altered quality of life in IPF, consistent with other physiological alterations also playing key roles in clinical expression of the disease. - from [here](https://err.ersjournals.com/content/27/147/170062)\n\n(--&gt; to summarise the above text: chest wall is not important, only lung parenchyma, FVC is thought not to be a good determinant for the DIAGNOSIS - not the follow up!)\n\nIf we assume that in this competition a european cohort had been diagnosed (see OSIC website), we may use the [spirometric references of the European Coal and Steel Community]( https://erj.ersjournals.com/content/31/3/687.2)  through which - since we already know the gender and age of the patient -  **we can derive the height of the patient** . \n\n**Body mass index (BMI)** was furthermore investigated as one of the determinants of the progression (see previous insight) as baseline characteristic and we may try to make assumptions on this, see [this study](https://pubs.rsna.org/doi/abs/10.1148/radiol.2283020095?uritype=cgi&amp;journalCode=radiology) or [this one](http://www.mipg.upenn.edu/yubing/Tong_PLOS2.pdf), but the most simple approaches are usually the best:\n###By determining the chest area, \nwe can categorize the patients into different BMI categories - [see reference here](https://www.researchgate.net/publication/51461844_Direct_chest_area_measurement_A_potential_anthropometric_replacement_for_BMI_to_inform_cardiac_CT_dose_parameters). Moreover the chest circumference showed also correlation to BMI.  \n\nYou may ask where should you determine the chest area... that is relatively simple to standardize if you have the lung segmented already: take the position of the first and the last image where lung is segmented and get the midposition. Chose the slice nearest to this position for your chest area. Instead of this, you may measure at 3 distinct levels (also derived from the topmost and bottom points) and make an average for your value.\n\n![midposition](http://drkonya.com/projects/kaggle/fibrosis_5.png)\n\n\nNow we have:  \n- total lung volume (see [previous insight](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165727))\n- height of the patient\n- BMI (estimate) of the patient\n- chest circumference as standalone criteria\n\n\nIn the next insight i will write about vessels and airways and their possible role in the progression determination.\n\n![inspiration](http://drkonya.com/projects/kaggle/inspi.jpg)",
      "votes": null
    },
    {
      "id": "925172",
      "postDate": "07/11/2020 21:28:19",
      "content": "<p>A newbie clarification for the data preprocessing: what exactly do you mean by </p>\n\n<blockquote>\n  <p>I assume you have your CT images segmented already...\n  I understand the lung segmentation mask part, just want to clarify this one.</p>\n</blockquote>\n\n<p>Thank you </p>",
      "rawMarkdown": "A newbie clarification for the data preprocessing: what exactly do you mean by \n&gt; I assume you have your CT images segmented already...\nI understand the lung segmentation mask part, just want to clarify this one.\n\nThank you",
      "votes": null
    },
    {
      "id": "925182",
      "postDate": "07/11/2020 21:40:38",
      "content": "<p>I am sorry if i was not clear, </p>\n\n<p>i was referring to my first insight, where i mentioned that almost every approach that derives features from the lung tissue is based on lung tissue segmentation (from the chest CT). </p>\n\n<p>Since there are multiple repositories on Github (<a href=\"https://github.com/JoHof/lungmask\">like this</a>) that do this out of the box -pretty well also in case of severely diseased lung - i assumed that everyone in the competition is done with the segmentation =)</p>\n\n<p>UPTADE: i changed the wording.</p>",
      "rawMarkdown": "I am sorry if i was not clear, \n\ni was referring to my first insight, where i mentioned that almost every approach that derives features from the lung tissue is based on lung tissue segmentation (from the chest CT). \n\nSince there are multiple repositories on Github ([like this](https://github.com/JoHof/lungmask)) that do this out of the box -pretty well also in case of severely diseased lung - i assumed that everyone in the competition is done with the segmentation =)\n\nUPTADE: i changed the wording.",
      "votes": null
    },
    {
      "id": "925188",
      "postDate": "07/11/2020 21:54:07",
      "content": "<p>Perfect! Thank you</p>",
      "rawMarkdown": "Perfect! Thank you",
      "votes": null
    },
    {
      "id": "925224",
      "postDate": "07/11/2020 23:47:55",
      "content": "<p>You've lost me at \"<strong>manual</strong> feature engineering\"... :)\nJust kidding, great post, thanks for sharing! </p>\n\n<p>(But isn't the point of Deep Learning, which powered all recent breakthroughs in Computer Vision, exactly to AVOID manual handcrafted feature engineering? :))</p>",
      "rawMarkdown": "You've lost me at \"**manual** feature engineering\"... :)\nJust kidding, great post, thanks for sharing! \n\n(But isn't the point of Deep Learning, which powered all recent breakthroughs in Computer Vision, exactly to AVOID manual handcrafted feature engineering? :))",
      "votes": null
    },
    {
      "id": "925255",
      "postDate": "07/12/2020 00:24:47",
      "content": "<p>Thank you for these amazing insights. I am a little overwhelmed but excited at the same time to try these. </p>",
      "rawMarkdown": "Thank you for these amazing insights. I am a little overwhelmed but excited at the same time to try these.",
      "votes": null
    },
    {
      "id": "925417",
      "postDate": "07/12/2020 04:29:33",
      "content": "<p>Thanks for the amazing insights! </p>\n\n<p>Can you tell me how FVC and chest circumference are related to smoking status? I've been trying to find a correlation, but this is a bit confusing for an individual with no medical background :)</p>",
      "rawMarkdown": "Thanks for the amazing insights! \n\nCan you tell me how FVC and chest circumference are related to smoking status? I've been trying to find a correlation, but this is a bit confusing for an individual with no medical background :)",
      "votes": null
    },
    {
      "id": "925699",
      "postDate": "07/12/2020 08:45:00",
      "content": "<p>The correlation between the variables FVC and chest circumference is not studied yet - as far i know.</p>\n\n<p>There are some points to consider:\n- the chest circumference is probably higher correlated with BMI (the fat layer sorrounding the rib cage can be sometimes pretty thick). \n- smoking is correlated to the effect of obstructive and restrictive changes in the lung, in this study we are investigating the restrictive effects.</p>\n\n<p>The correlations of smoking status on FVC are usually combined in composite scores (along with gender and age).</p>\n\n<p>One good work on this is probably this:\n<a href=\"https://pubmed.ncbi.nlm.nih.gov/27767347/\">Idiopathic Pulmonary Fibrosis: The Association Between the Adaptive Multiple Features Method and Fibrosis Outcomes</a>  --- (hint: they have correlation between image morphologic changes and progression to!)</p>",
      "rawMarkdown": "The correlation between the variables FVC and chest circumference is not studied yet - as far i know.\n\nThere are some points to consider:\n- the chest circumference is probably higher correlated with BMI (the fat layer sorrounding the rib cage can be sometimes pretty thick). \n- smoking is correlated to the effect of obstructive and restrictive changes in the lung, in this study we are investigating the restrictive effects.\n\nThe correlations of smoking status on FVC are usually combined in composite scores (along with gender and age).\n\nOne good work on this is probably this:\n[Idiopathic Pulmonary Fibrosis: The Association Between the Adaptive Multiple Features Method and Fibrosis Outcomes](https://pubmed.ncbi.nlm.nih.gov/27767347/)  --- (hint: they have correlation between image morphologic changes and progression to!)",
      "votes": null
    },
    {
      "id": "925703",
      "postDate": "07/12/2020 08:48:49",
      "content": "<p>Thank you for the nice words!</p>",
      "rawMarkdown": "Thank you for the nice words!",
      "votes": null
    },
    {
      "id": "925709",
      "postDate": "07/12/2020 08:53:53",
      "content": "<p>Hi there, </p>\n\n<p>yes, you're absolutely right about the point of DL.\nHowever, with mere 178 Patient's 178 * (30 - 350) image, each having a 512*512 pixel, good luck finding correlations =).\nThis is the point where apriori knowledge is essential... but maybe i'm wrong about this.</p>",
      "rawMarkdown": "Hi there, \n\nyes, you're absolutely right about the point of DL.\nHowever, with mere 178 Patient's 178 * (30 - 350) image, each having a 512*512 pixel, good luck finding correlations =).\nThis is the point where apriori knowledge is essential... but maybe i'm wrong about this.",
      "votes": null
    },
    {
      "id": "926081",
      "postDate": "07/12/2020 13:32:28",
      "content": "<p>When I saw 176 patients, at first I thought the same... but then I read the following in a paper from DeepMind:</p>\n\n<p>\"In many biomedical applications, only very few images are required to train a network that generalizes reasonably well. This is because each image already comprises repetitive structures with corresponding variation. In volumetric images, this effect is further pronounced, such that <strong>we can train a network on just two volumetric images in order to generalize to a third one</strong>. A weighted loss function and special data augmentation enable us to train the network with only few manually annotated slices, i.e., from sparsely annotated training data.\"</p>\n\n<p>Çiçek, Ö., Abdulkadir, A., Lienkamp, S. S., Brox, T., &amp; Ronneberger, O. (2016, October). 3D U-Net: learning dense volumetric segmentation from sparse annotation. In <em>International conference on medical image computing and computer-assisted intervention</em> (pp. 424-432). Springer, Cham</p>",
      "rawMarkdown": "When I saw 176 patients, at first I thought the same... but then I read the following in a paper from DeepMind:\n\n\"In many biomedical applications, only very few images are required to train a network that generalizes reasonably well. This is because each image already comprises repetitive structures with corresponding variation. In volumetric images, this effect is further pronounced, such that **we can train a network on just two volumetric images in order to generalize to a third one**. A weighted loss function and special data augmentation enable us to train the network with only few manually annotated slices, i.e., from sparsely annotated training data.\"\n\nÇiçek, Ö., Abdulkadir, A., Lienkamp, S. S., Brox, T., &amp; Ronneberger, O. (2016, October). 3D U-Net: learning dense volumetric segmentation from sparse annotation. In *International conference on medical image computing and computer-assisted intervention* (pp. 424-432). Springer, Cham",
      "votes": null
    },
    {
      "id": "926118",
      "postDate": "07/12/2020 14:02:44",
      "content": "<p>It is not about segmenting the diseased lung from the chest CT... that is no problem... </p>\n\n<p>The problem is to derive the good features from this segmented volumetric data and correlate them to the overall progression of the disease.</p>",
      "rawMarkdown": "It is not about segmenting the diseased lung from the chest CT... that is no problem... \n\nThe problem is to derive the good features from this segmented volumetric data and correlate them to the overall progression of the disease.",
      "votes": null
    },
    {
      "id": "926133",
      "postDate": "07/12/2020 14:18:21",
      "content": "<p>Thank you very much <a href=\"/sandorkonya\">@sandorkonya</a> for a deeper insight into the practical dataset. </p>\n\n<p>Though I am no medical diagnosis expert, I made a simple introduction here <a href=\"https://www.kaggle.com/redwankarimsony/pulmonary-fibrosis-for-non-med-people\">Pulmonary fibrosis for Non-Med People</a> to Pulmonary fibrosis collecting bits and pieces from here and there on the internet. Can you please have a look and point out if I have something wrong? </p>\n\n<p>Thanks in advance. </p>\n\n<p>NB: This is not a publicity request. </p>",
      "rawMarkdown": "Thank you very much @sandorkonya for a deeper insight into the practical dataset. \n\nThough I am no medical diagnosis expert, I made a simple introduction here [Pulmonary fibrosis for Non-Med People](https://www.kaggle.com/redwankarimsony/pulmonary-fibrosis-for-non-med-people) to Pulmonary fibrosis collecting bits and pieces from here and there on the internet. Can you please have a look and point out if I have something wrong? \n\nThanks in advance. \n\nNB: This is not a publicity request.",
      "votes": null
    },
    {
      "id": "926137",
      "postDate": "07/12/2020 14:20:23",
      "content": "<p>Finding out this correlation is the main task of the ML or Data Science Engineers I guess. </p>",
      "rawMarkdown": "Finding out this correlation is the main task of the ML or Data Science Engineers I guess.",
      "votes": null
    },
    {
      "id": "926155",
      "postDate": "07/12/2020 14:35:25",
      "content": "<p>Dear <a href=\"/redwankarimsony\">@redwankarimsony</a> , nice starter with relevant information,\nkeep up the good work!</p>",
      "rawMarkdown": "Dear @redwankarimsony , nice starter with relevant information,\nkeep up the good work!",
      "votes": null
    },
    {
      "id": "926161",
      "postDate": "07/12/2020 14:39:25",
      "content": "<p>I know the problem is not segmentation, the problem here is different. That is not the point. The point I am making is that, in the cited paper, they proved they can use a ConvNet to extract latent features from very few samples of 3D medical images. Once you have latent features, you can tweak the net to tackle any problem you want (classification, segmentation, etc). Or just use the features ;)</p>\n\n<p>Just uploaded 12 GB of pre-processed tensors to a Kaggle dataset. Let's see if I'm lucky in training a ConvNet to find such features...</p>",
      "rawMarkdown": "I know the problem is not segmentation, the problem here is different. That is not the point. The point I am making is that, in the cited paper, they proved they can use a ConvNet to extract latent features from very few samples of 3D medical images. Once you have latent features, you can tweak the net to tackle any problem you want (classification, segmentation, etc). Or just use the features ;)\n\nJust uploaded 12 GB of pre-processed tensors to a Kaggle dataset. Let's see if I'm lucky in training a ConvNet to find such features...",
      "votes": null
    },
    {
      "id": "926168",
      "postDate": "07/12/2020 14:45:10",
      "content": "<p>Thanks <a href=\"/sandorkonya\">@sandorkonya</a> for your words of advice.. 👍 </p>",
      "rawMarkdown": "Thanks @sandorkonya for your words of advice.. 👍",
      "votes": null
    },
    {
      "id": "932092",
      "postDate": "07/16/2020 17:32:11",
      "content": "<p><a href=\"/carlossouza\">@carlossouza</a> Were you lucky in training the model to learn the features?</p>",
      "rawMarkdown": "carlossouza Were you lucky in training the model to learn the features?",
      "votes": null
    },
    {
      "id": "932132",
      "postDate": "07/16/2020 18:21:45",
      "content": "<p>Yes... The model learns latent features. Check the Auto Encoder inputs/outputs below:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F915913%2F762097554f8bb3808687aab5e3eb2393%2Fsample.gif?generation=1594923576774616&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F915913%2F237b23ecf1eb6e61163f0951b0daf984%2Foutput.gif?generation=1594923596122095&amp;alt=media\" alt=\"\"></p>\n\n<p>This was my latest experiment, using properly masked lungs as inputs, instead of raw CT scan images. Now, next steps are:\n1. Implement a baseline model using only tabular data\n2. Implement a model combining tabular data and the latent features learned\n3. Evaluate how much the latent features are improving the baseline model</p>",
      "rawMarkdown": "Yes... The model learns latent features. Check the Auto Encoder inputs/outputs below:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F915913%2F762097554f8bb3808687aab5e3eb2393%2Fsample.gif?generation=1594923576774616&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F915913%2F237b23ecf1eb6e61163f0951b0daf984%2Foutput.gif?generation=1594923596122095&amp;alt=media)\n\nThis was my latest experiment, using properly masked lungs as inputs, instead of raw CT scan images. Now, next steps are:\n1. Implement a baseline model using only tabular data\n2. Implement a model combining tabular data and the latent features learned\n3. Evaluate how much the latent features are improving the baseline model",
      "votes": null
    },
    {
      "id": "932144",
      "postDate": "07/16/2020 18:34:47",
      "content": "<p><a href=\"https://www.kaggle.com/carlossouza\" target=\"_blank\">@carlossouza</a>, nice!<br>\nKeep up the good work &amp; keep us informed!</p>",
      "rawMarkdown": "carlossouza, nice!\nKeep up the good work &amp; keep us informed!",
      "votes": null
    },
    {
      "id": "932165",
      "postDate": "07/16/2020 19:12:28",
      "content": "<p>Those masks of yours seem very granular and detailed. I also resampled the images to make all slice_thickness and pixel spacings as 1mm and created masks with u-net from them with a fixed threshold of -500. Did some post-processing to remove noise. My masks have the \"right\" contours but are not detailed in the inside of the lung as yours. Can you give us some hints of what you did differently?</p>\n\n<p>Thanks.</p>",
      "rawMarkdown": "Those masks of yours seem very granular and detailed. I also resampled the images to make all slice_thickness and pixel spacings as 1mm and created masks with u-net from them with a fixed threshold of -500. Did some post-processing to remove noise. My masks have the \"right\" contours but are not detailed in the inside of the lung as yours. Can you give us some hints of what you did differently?\n\nThanks.",
      "votes": null
    },
    {
      "id": "932195",
      "postDate": "07/16/2020 19:50:34",
      "content": "<p>Sure! Pre-processing was very painful. Getting the transformations right, then get the right order of transformations made a big difference. Here's a code snippet from my codebase that illustrates my overall process:</p>\n\n<p><code>\ndata = CTScansDataset(\n   root_dir='../data/train',\n   transform=transforms.Compose([\n      CropBoundingBox(),\n      ConvertToHU(),\n      Resize((40, 256, 256)),\n      Clip(bounds=(-1000, 500)),\n      Mask(method=MaskMethod.MORPHOLOGICAL, threshold=-500),\n      Normalize(bounds=(-1000, -500)),\n      ToTensor()\n   ]))\n</code></p>\n\n<p>First, CTScanDataset reads all subfolders in the root dir, stacks all pixel_arrays from all DICOM files within each subfolder in a 3D numpy array for each patient, with varied dimensions (median is 90 x 512 x 512, but varies a lot). Then, the preprocessing begins, following the order you see in the code. They are straightforward. On masking, the 5th step, I tried 2 methods: a morphological one, following <a href=\"https://www.kaggle.com/miklgr500/unsupervise-lung-detection\">this great notebook</a> by <a href=\"/miklgr500\">@miklgr500</a> , and a deep learning one, using <a href=\"https://github.com/JoHof/lungmask\">this pre-trained model</a>. The problem with the deep learning model is that it is extremely slow. For the sake of speed, I'm sticking to the morphological approach.</p>\n\n<p>Things that were not trivial: figuring out the data types to avoid loss on conversions; figuring out the best method to resize/interpolate; and putting everything together in the right order.</p>\n\n<p>PS: If you resample using the wrong method/parameters, you can either i) end up with very poor resolution images, or ii) end up with huge tensors that are not in the same dimensions/impossible to load in GPU memory..</p>\n\n<p>Hope it helps!</p>",
      "rawMarkdown": "Sure! Pre-processing was very painful. Getting the transformations right, then get the right order of transformations made a big difference. Here's a code snippet from my codebase that illustrates my overall process:\n\n```\ndata = CTScansDataset(\n   root_dir='../data/train',\n   transform=transforms.Compose([\n      CropBoundingBox(),\n      ConvertToHU(),\n      Resize((40, 256, 256)),\n      Clip(bounds=(-1000, 500)),\n      Mask(method=MaskMethod.MORPHOLOGICAL, threshold=-500),\n      Normalize(bounds=(-1000, -500)),\n      ToTensor()\n   ]))\n```\n\nFirst, CTScanDataset reads all subfolders in the root dir, stacks all pixel_arrays from all DICOM files within each subfolder in a 3D numpy array for each patient, with varied dimensions (median is 90 x 512 x 512, but varies a lot). Then, the preprocessing begins, following the order you see in the code. They are straightforward. On masking, the 5th step, I tried 2 methods: a morphological one, following [this great notebook](https://www.kaggle.com/miklgr500/unsupervise-lung-detection) by @miklgr500 , and a deep learning one, using [this pre-trained model](https://github.com/JoHof/lungmask). The problem with the deep learning model is that it is extremely slow. For the sake of speed, I'm sticking to the morphological approach.\n\nThings that were not trivial: figuring out the data types to avoid loss on conversions; figuring out the best method to resize/interpolate; and putting everything together in the right order.\n\nPS: If you resample using the wrong method/parameters, you can either i) end up with very poor resolution images, or ii) end up with huge tensors that are not in the same dimensions/impossible to load in GPU memory..\n\nHope it helps!",
      "votes": null
    },
    {
      "id": "932211",
      "postDate": "07/16/2020 20:20:39",
      "content": "<p>It definitely helped a lot. I've used the same deep learning model you mentioned. What I did was to run and save the binary masks generated from the resampled images ( I've ran the DL model just once for the whole dataset in ~40mins) and use them to index the resampled arrays and them transform to HU. Your approach with a pipeline of transforms is a lot more flexible. My resampled images are a lot more compact than the original ones but don't seem to have lost a lot of information. ( Visual inspection is not a very robust metric ).</p>\n\n<p>I'll definitely reevaluate them given your insight. </p>\n\n<p>Thank you very much</p>",
      "rawMarkdown": "It definitely helped a lot. I've used the same deep learning model you mentioned. What I did was to run and save the binary masks generated from the resampled images ( I've ran the DL model just once for the whole dataset in ~40mins) and use them to index the resampled arrays and them transform to HU. Your approach with a pipeline of transforms is a lot more flexible. My resampled images are a lot more compact than the original ones but don't seem to have lost a lot of information. ( Visual inspection is not a very robust metric ).\n\n I'll definitely reevaluate them given your insight. \n\nThank you very much",
      "votes": null
    },
    {
      "id": "977789",
      "postDate": "08/19/2020 17:51:55",
      "content": "<p><a href=\"https://www.kaggle.com/sandorkonya\" target=\"_blank\">@sandorkonya</a>  many thanks for insight however few questions<br>\nin some cases see even too low FVC like less than 2500,2000  are conisdered 90 to 100 above normal values so a normal value for reference cases is just 1500… shall we consider such low absolute value to be indicative of normal state of Lungs breathing capacity ?</p>",
      "rawMarkdown": "sandorkonya  many thanks for insight however few questions\nin some cases see even too low FVC like less than 2500,2000  are conisdered 90 to 100 above normal values so a normal value for reference cases is just 1500... shall we consider such low absolute value to be indicative of normal state of Lungs breathing capacity ?",
      "votes": null
    },
    {
      "id": "978203",
      "postDate": "08/20/2020 03:00:16",
      "content": "<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> thanks for pointing this out. I will try to explore more about this and post my findings here. </p>",
      "rawMarkdown": "jaideepvalani thanks for pointing this out. I will try to explore more about this and post my findings here.",
      "votes": null
    },
    {
      "id": "978204",
      "postDate": "08/20/2020 03:01:36",
      "content": "<p><a href=\"https://www.kaggle.com/ronaldokun\" target=\"_blank\">@ronaldokun</a> Thanks for the appreciation. I am overwhelmed too cause this is the first competition for me of this sort. </p>",
      "rawMarkdown": "ronaldokun Thanks for the appreciation. I am overwhelmed too cause this is the first competition for me of this sort.",
      "votes": null
    },
    {
      "id": "979191",
      "postDate": "08/20/2020 17:23:13",
      "content": "<p><a href=\"https://www.kaggle.com/redwankarimsony\" target=\"_blank\">@redwankarimsony</a>   <br>\n\"we already know the gender and age of the patient - we can derive the height of the patient .\"</p>\n<p>How do we determine height here, height is kind of random distribution or gausian distribution..<br>\nCan any field of meta data help ?</p>",
      "rawMarkdown": "redwankarimsony   \n\"we already know the gender and age of the patient - we can derive the height of the patient .\"\n\nHow do we determine height here, height is kind of random distribution or gausian distribution..\nCan any field of meta data help ?",
      "votes": null
    },
    {
      "id": "979208",
      "postDate": "08/20/2020 17:36:29",
      "content": "<p><a href=\"https://www.kaggle.com/ronaldokun\" target=\"_blank\">@ronaldokun</a> team up.. my current score is addition of very useful non img features ,i think your approach and mine should help go further</p>",
      "rawMarkdown": "ronaldokun team up.. my current score is addition of very useful non img features ,i think your approach and mine should help go further",
      "votes": null
    },
    {
      "id": "979945",
      "postDate": "08/21/2020 08:19:34",
      "content": "<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> see <a href=\"https://erj.ersjournals.com/content/erj/11/6/1354.full.pdf\" target=\"_blank\">here</a> table 2 for the equations. I've tried but did not manage to get good results</p>",
      "rawMarkdown": "jaideepvalani see [here](https://erj.ersjournals.com/content/erj/11/6/1354.full.pdf) table 2 for the equations. I've tried but did not manage to get good results",
      "votes": null
    },
    {
      "id": "980483",
      "postDate": "08/21/2020 16:16:46",
      "content": "<p>i think what we need is actual height not a height chart based on age to build a better correlation for model .</p>",
      "rawMarkdown": "i think what we need is actual height not a height chart based on age to build a better correlation for model .",
      "votes": null
    },
    {
      "id": "982711",
      "postDate": "08/23/2020 15:47:01",
      "content": "<p><a href=\"https://www.kaggle.com/alexj21\" target=\"_blank\">@alexj21</a> this is actually exactly what i meant for the height calculation.</p>\n<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> is correct, the actual height would be ideal… there should be a correlation between the size of the  vertebrae and height, so this is something that one could actually derive from the scans.</p>",
      "rawMarkdown": "alexj21 this is actually exactly what i meant for the height calculation.\n\n@jaideepvalani is correct, the actual height would be ideal... there should be a correlation between the size of the  vertebrae and height, so this is something that one could actually derive from the scans.",
      "votes": null
    },
    {
      "id": "982721",
      "postDate": "08/23/2020 15:54:40",
      "content": "<p><a href=\"https://www.kaggle.com/sandorkonya\" target=\"_blank\">@sandorkonya</a>  yes i agree but there would be a deviation from actual one based on heredity of person. <br>\nEg. If my father and mother are tall,then my adult height could be more than same person of my age if persons parents have average height or short height. </p>\n<p>Scans do have this field but unfortunately this is not entered info in scan. </p>",
      "rawMarkdown": "sandorkonya  yes i agree but there would be a deviation from actual one based on heredity of person. \nEg. If my father and mother are tall,then my adult height could be more than same person of my age if persons parents have average height or short height. \n\nScans do have this field but unfortunately this is not entered info in scan.",
      "votes": null
    },
    {
      "id": "982738",
      "postDate": "08/23/2020 16:08:49",
      "content": "<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> ,<br>\nyes. many important information is missing. </p>",
      "rawMarkdown": "jaideepvalani ,\nyes. many important information is missing.",
      "votes": null
    },
    {
      "id": "982832",
      "postDate": "08/23/2020 17:38:56",
      "content": "<p>1) how can we get chest area/vol <br>\n2) voxal vol same as chest volume?</p>",
      "rawMarkdown": "1) how can we get chest area/vol \n2) voxal vol same as chest volume?",
      "votes": null
    },
    {
      "id": "987572",
      "postDate": "08/27/2020 11:00:18",
      "content": "<p><a href=\"https://www.kaggle.com/sandorkonya\" target=\"_blank\">@sandorkonya</a> <br>\ncould u help me understand below terminology<br>\n1) slice thickness<br>\n2) slice spacings <br>\n3) voxel volumes  (i understand this is volume of a pixel )    ,how does it propertional to above 2  directly or inversely</p>",
      "rawMarkdown": "sandorkonya \ncould u help me understand below terminology\n1) slice thickness\n2) slice spacings \n3) voxel volumes  (i understand this is volume of a pixel )    ,how does it propertional to above 2  directly or inversely",
      "votes": null
    },
    {
      "id": "1004976",
      "postDate": "09/10/2020 07:14:07",
      "content": "<p>Nice.</p>",
      "rawMarkdown": "Nice.",
      "votes": null
    },
    {
      "id": "1028703",
      "postDate": "09/27/2020 05:59:10",
      "content": "<p>Thanks you for these great posts  - helps a newbie like me very much</p>",
      "rawMarkdown": "Thanks you for these great posts  - helps a newbie like me very much",
      "votes": null
    },
    {
      "id": "1038568",
      "postDate": "10/05/2020 21:52:37",
      "content": "<p><a href=\"https://www.kaggle.com/vishram6\" target=\"_blank\">@vishram6</a> ,<br>\nyou're welcome!</p>",
      "rawMarkdown": "vishram6 ,\nyou're welcome!",
      "votes": null
    },
    {
      "id": "1041656",
      "postDate": "10/07/2020 21:30:23",
      "content": "<p>UPDATE</p>\n<p>Dear fellow kagglers, </p>\n<p>i posted <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/189528\" target=\"_blank\">segmentation masks</a> after the competition.</p>",
      "rawMarkdown": "UPDATE\n\nDear fellow kagglers, \n\ni posted [segmentation masks](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/189528) after the competition.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 977789,
      "author_name": "jaideepvalani",
      "author_url": "",
      "post_date": "08/19/2020 17:51:55",
      "content": "<p><a href=\"https://www.kaggle.com/sandorkonya\" target=\"_blank\">@sandorkonya</a>  many thanks for insight however few questions<br>\nin some cases see even too low FVC like less than 2500,2000  are conisdered 90 to 100 above normal values so a normal value for reference cases is just 1500… shall we consider such low absolute value to be indicative of normal state of Lungs breathing capacity ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 978203,
          "author_name": "redwankarimsony",
          "author_url": "",
          "post_date": "08/20/2020 03:00:16",
          "content": "<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> thanks for pointing this out. I will try to explore more about this and post my findings here. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 979191,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "08/20/2020 17:23:13",
          "content": "<p><a href=\"https://www.kaggle.com/redwankarimsony\" target=\"_blank\">@redwankarimsony</a>   <br>\n\"we already know the gender and age of the patient - we can derive the height of the patient .\"</p>\n<p>How do we determine height here, height is kind of random distribution or gausian distribution..<br>\nCan any field of meta data help ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 979945,
          "author_name": "alexj21",
          "author_url": "",
          "post_date": "08/21/2020 08:19:34",
          "content": "<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> see <a href=\"https://erj.ersjournals.com/content/erj/11/6/1354.full.pdf\" target=\"_blank\">here</a> table 2 for the equations. I've tried but did not manage to get good results</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 980483,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "08/21/2020 16:16:46",
          "content": "<p>i think what we need is actual height not a height chart based on age to build a better correlation for model .</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 982711,
          "author_name": "sandorkonya",
          "author_url": "",
          "post_date": "08/23/2020 15:47:01",
          "content": "<p><a href=\"https://www.kaggle.com/alexj21\" target=\"_blank\">@alexj21</a> this is actually exactly what i meant for the height calculation.</p>\n<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> is correct, the actual height would be ideal… there should be a correlation between the size of the  vertebrae and height, so this is something that one could actually derive from the scans.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 982721,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "08/23/2020 15:54:40",
          "content": "<p><a href=\"https://www.kaggle.com/sandorkonya\" target=\"_blank\">@sandorkonya</a>  yes i agree but there would be a deviation from actual one based on heredity of person. <br>\nEg. If my father and mother are tall,then my adult height could be more than same person of my age if persons parents have average height or short height. </p>\n<p>Scans do have this field but unfortunately this is not entered info in scan. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 982738,
          "author_name": "sandorkonya",
          "author_url": "",
          "post_date": "08/23/2020 16:08:49",
          "content": "<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> ,<br>\nyes. many important information is missing. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 982832,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "08/23/2020 17:38:56",
          "content": "<p>1) how can we get chest area/vol <br>\n2) voxal vol same as chest volume?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 987572,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "08/27/2020 11:00:18",
          "content": "<p><a href=\"https://www.kaggle.com/sandorkonya\" target=\"_blank\">@sandorkonya</a> <br>\ncould u help me understand below terminology<br>\n1) slice thickness<br>\n2) slice spacings <br>\n3) voxel volumes  (i understand this is volume of a pixel )    ,how does it propertional to above 2  directly or inversely</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1028703,
      "author_name": "vishram6",
      "author_url": "",
      "post_date": "09/27/2020 05:59:10",
      "content": "<p>Thanks you for these great posts  - helps a newbie like me very much</p>",
      "votes": null,
      "replies": [
        {
          "id": 1038568,
          "author_name": "sandorkonya",
          "author_url": "",
          "post_date": "10/05/2020 21:52:37",
          "content": "<p><a href=\"https://www.kaggle.com/vishram6\" target=\"_blank\">@vishram6</a> ,<br>\nyou're welcome!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1041656,
      "author_name": "sandorkonya",
      "author_url": "",
      "post_date": "10/07/2020 21:30:23",
      "content": "<p>UPDATE</p>\n<p>Dear fellow kagglers, </p>\n<p>i posted <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/189528\" target=\"_blank\">segmentation masks</a> after the competition.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 925172,
      "author_name": "ronaldokun",
      "author_url": "",
      "post_date": "07/11/2020 21:28:19",
      "content": "<p>A newbie clarification for the data preprocessing: what exactly do you mean by </p>\n\n<blockquote>\n  <p>I assume you have your CT images segmented already...\n  I understand the lung segmentation mask part, just want to clarify this one.</p>\n</blockquote>\n\n<p>Thank you </p>",
      "votes": null,
      "replies": [
        {
          "id": 925182,
          "author_name": "sandorkonya",
          "author_url": "",
          "post_date": "07/11/2020 21:40:38",
          "content": "<p>I am sorry if i was not clear, </p>\n\n<p>i was referring to my first insight, where i mentioned that almost every approach that derives features from the lung tissue is based on lung tissue segmentation (from the chest CT). </p>\n\n<p>Since there are multiple repositories on Github (<a href=\"https://github.com/JoHof/lungmask\">like this</a>) that do this out of the box -pretty well also in case of severely diseased lung - i assumed that everyone in the competition is done with the segmentation =)</p>\n\n<p>UPTADE: i changed the wording.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 925188,
          "author_name": "ronaldokun",
          "author_url": "",
          "post_date": "07/11/2020 21:54:07",
          "content": "<p>Perfect! Thank you</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 925224,
      "author_name": "carlossouza",
      "author_url": "",
      "post_date": "07/11/2020 23:47:55",
      "content": "<p>You've lost me at \"<strong>manual</strong> feature engineering\"... :)\nJust kidding, great post, thanks for sharing! </p>\n\n<p>(But isn't the point of Deep Learning, which powered all recent breakthroughs in Computer Vision, exactly to AVOID manual handcrafted feature engineering? :))</p>",
      "votes": null,
      "replies": [
        {
          "id": 925709,
          "author_name": "sandorkonya",
          "author_url": "",
          "post_date": "07/12/2020 08:53:53",
          "content": "<p>Hi there, </p>\n\n<p>yes, you're absolutely right about the point of DL.\nHowever, with mere 178 Patient's 178 * (30 - 350) image, each having a 512*512 pixel, good luck finding correlations =).\nThis is the point where apriori knowledge is essential... but maybe i'm wrong about this.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 926081,
          "author_name": "carlossouza",
          "author_url": "",
          "post_date": "07/12/2020 13:32:28",
          "content": "<p>When I saw 176 patients, at first I thought the same... but then I read the following in a paper from DeepMind:</p>\n\n<p>\"In many biomedical applications, only very few images are required to train a network that generalizes reasonably well. This is because each image already comprises repetitive structures with corresponding variation. In volumetric images, this effect is further pronounced, such that <strong>we can train a network on just two volumetric images in order to generalize to a third one</strong>. A weighted loss function and special data augmentation enable us to train the network with only few manually annotated slices, i.e., from sparsely annotated training data.\"</p>\n\n<p>Çiçek, Ö., Abdulkadir, A., Lienkamp, S. S., Brox, T., &amp; Ronneberger, O. (2016, October). 3D U-Net: learning dense volumetric segmentation from sparse annotation. In <em>International conference on medical image computing and computer-assisted intervention</em> (pp. 424-432). Springer, Cham</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 926118,
          "author_name": "sandorkonya",
          "author_url": "",
          "post_date": "07/12/2020 14:02:44",
          "content": "<p>It is not about segmenting the diseased lung from the chest CT... that is no problem... </p>\n\n<p>The problem is to derive the good features from this segmented volumetric data and correlate them to the overall progression of the disease.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 926137,
          "author_name": "redwankarimsony",
          "author_url": "",
          "post_date": "07/12/2020 14:20:23",
          "content": "<p>Finding out this correlation is the main task of the ML or Data Science Engineers I guess. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 926161,
          "author_name": "carlossouza",
          "author_url": "",
          "post_date": "07/12/2020 14:39:25",
          "content": "<p>I know the problem is not segmentation, the problem here is different. That is not the point. The point I am making is that, in the cited paper, they proved they can use a ConvNet to extract latent features from very few samples of 3D medical images. Once you have latent features, you can tweak the net to tackle any problem you want (classification, segmentation, etc). Or just use the features ;)</p>\n\n<p>Just uploaded 12 GB of pre-processed tensors to a Kaggle dataset. Let's see if I'm lucky in training a ConvNet to find such features...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 932092,
          "author_name": "rsreeram17",
          "author_url": "",
          "post_date": "07/16/2020 17:32:11",
          "content": "<p><a href=\"/carlossouza\">@carlossouza</a> Were you lucky in training the model to learn the features?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 932132,
          "author_name": "carlossouza",
          "author_url": "",
          "post_date": "07/16/2020 18:21:45",
          "content": "<p>Yes... The model learns latent features. Check the Auto Encoder inputs/outputs below:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F915913%2F762097554f8bb3808687aab5e3eb2393%2Fsample.gif?generation=1594923576774616&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F915913%2F237b23ecf1eb6e61163f0951b0daf984%2Foutput.gif?generation=1594923596122095&amp;alt=media\" alt=\"\"></p>\n\n<p>This was my latest experiment, using properly masked lungs as inputs, instead of raw CT scan images. Now, next steps are:\n1. Implement a baseline model using only tabular data\n2. Implement a model combining tabular data and the latent features learned\n3. Evaluate how much the latent features are improving the baseline model</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 932144,
          "author_name": "sandorkonya",
          "author_url": "",
          "post_date": "07/16/2020 18:34:47",
          "content": "<p><a href=\"https://www.kaggle.com/carlossouza\" target=\"_blank\">@carlossouza</a>, nice!<br>\nKeep up the good work &amp; keep us informed!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 932165,
          "author_name": "ronaldokun",
          "author_url": "",
          "post_date": "07/16/2020 19:12:28",
          "content": "<p>Those masks of yours seem very granular and detailed. I also resampled the images to make all slice_thickness and pixel spacings as 1mm and created masks with u-net from them with a fixed threshold of -500. Did some post-processing to remove noise. My masks have the \"right\" contours but are not detailed in the inside of the lung as yours. Can you give us some hints of what you did differently?</p>\n\n<p>Thanks.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 932195,
          "author_name": "carlossouza",
          "author_url": "",
          "post_date": "07/16/2020 19:50:34",
          "content": "<p>Sure! Pre-processing was very painful. Getting the transformations right, then get the right order of transformations made a big difference. Here's a code snippet from my codebase that illustrates my overall process:</p>\n\n<p><code>\ndata = CTScansDataset(\n   root_dir='../data/train',\n   transform=transforms.Compose([\n      CropBoundingBox(),\n      ConvertToHU(),\n      Resize((40, 256, 256)),\n      Clip(bounds=(-1000, 500)),\n      Mask(method=MaskMethod.MORPHOLOGICAL, threshold=-500),\n      Normalize(bounds=(-1000, -500)),\n      ToTensor()\n   ]))\n</code></p>\n\n<p>First, CTScanDataset reads all subfolders in the root dir, stacks all pixel_arrays from all DICOM files within each subfolder in a 3D numpy array for each patient, with varied dimensions (median is 90 x 512 x 512, but varies a lot). Then, the preprocessing begins, following the order you see in the code. They are straightforward. On masking, the 5th step, I tried 2 methods: a morphological one, following <a href=\"https://www.kaggle.com/miklgr500/unsupervise-lung-detection\">this great notebook</a> by <a href=\"/miklgr500\">@miklgr500</a> , and a deep learning one, using <a href=\"https://github.com/JoHof/lungmask\">this pre-trained model</a>. The problem with the deep learning model is that it is extremely slow. For the sake of speed, I'm sticking to the morphological approach.</p>\n\n<p>Things that were not trivial: figuring out the data types to avoid loss on conversions; figuring out the best method to resize/interpolate; and putting everything together in the right order.</p>\n\n<p>PS: If you resample using the wrong method/parameters, you can either i) end up with very poor resolution images, or ii) end up with huge tensors that are not in the same dimensions/impossible to load in GPU memory..</p>\n\n<p>Hope it helps!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 932211,
          "author_name": "ronaldokun",
          "author_url": "",
          "post_date": "07/16/2020 20:20:39",
          "content": "<p>It definitely helped a lot. I've used the same deep learning model you mentioned. What I did was to run and save the binary masks generated from the resampled images ( I've ran the DL model just once for the whole dataset in ~40mins) and use them to index the resampled arrays and them transform to HU. Your approach with a pipeline of transforms is a lot more flexible. My resampled images are a lot more compact than the original ones but don't seem to have lost a lot of information. ( Visual inspection is not a very robust metric ).</p>\n\n<p>I'll definitely reevaluate them given your insight. </p>\n\n<p>Thank you very much</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 979208,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "08/20/2020 17:36:29",
          "content": "<p><a href=\"https://www.kaggle.com/ronaldokun\" target=\"_blank\">@ronaldokun</a> team up.. my current score is addition of very useful non img features ,i think your approach and mine should help go further</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 925255,
      "author_name": "ronaldokun",
      "author_url": "",
      "post_date": "07/12/2020 00:24:47",
      "content": "<p>Thank you for these amazing insights. I am a little overwhelmed but excited at the same time to try these. </p>",
      "votes": null,
      "replies": [
        {
          "id": 925703,
          "author_name": "sandorkonya",
          "author_url": "",
          "post_date": "07/12/2020 08:48:49",
          "content": "<p>Thank you for the nice words!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 978204,
          "author_name": "redwankarimsony",
          "author_url": "",
          "post_date": "08/20/2020 03:01:36",
          "content": "<p><a href=\"https://www.kaggle.com/ronaldokun\" target=\"_blank\">@ronaldokun</a> Thanks for the appreciation. I am overwhelmed too cause this is the first competition for me of this sort. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 925417,
      "author_name": "aadhavvignesh",
      "author_url": "",
      "post_date": "07/12/2020 04:29:33",
      "content": "<p>Thanks for the amazing insights! </p>\n\n<p>Can you tell me how FVC and chest circumference are related to smoking status? I've been trying to find a correlation, but this is a bit confusing for an individual with no medical background :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 925699,
          "author_name": "sandorkonya",
          "author_url": "",
          "post_date": "07/12/2020 08:45:00",
          "content": "<p>The correlation between the variables FVC and chest circumference is not studied yet - as far i know.</p>\n\n<p>There are some points to consider:\n- the chest circumference is probably higher correlated with BMI (the fat layer sorrounding the rib cage can be sometimes pretty thick). \n- smoking is correlated to the effect of obstructive and restrictive changes in the lung, in this study we are investigating the restrictive effects.</p>\n\n<p>The correlations of smoking status on FVC are usually combined in composite scores (along with gender and age).</p>\n\n<p>One good work on this is probably this:\n<a href=\"https://pubmed.ncbi.nlm.nih.gov/27767347/\">Idiopathic Pulmonary Fibrosis: The Association Between the Adaptive Multiple Features Method and Fibrosis Outcomes</a>  --- (hint: they have correlation between image morphologic changes and progression to!)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 926133,
      "author_name": "redwankarimsony",
      "author_url": "",
      "post_date": "07/12/2020 14:18:21",
      "content": "<p>Thank you very much <a href=\"/sandorkonya\">@sandorkonya</a> for a deeper insight into the practical dataset. </p>\n\n<p>Though I am no medical diagnosis expert, I made a simple introduction here <a href=\"https://www.kaggle.com/redwankarimsony/pulmonary-fibrosis-for-non-med-people\">Pulmonary fibrosis for Non-Med People</a> to Pulmonary fibrosis collecting bits and pieces from here and there on the internet. Can you please have a look and point out if I have something wrong? </p>\n\n<p>Thanks in advance. </p>\n\n<p>NB: This is not a publicity request. </p>",
      "votes": null,
      "replies": [
        {
          "id": 926155,
          "author_name": "sandorkonya",
          "author_url": "",
          "post_date": "07/12/2020 14:35:25",
          "content": "<p>Dear <a href=\"/redwankarimsony\">@redwankarimsony</a> , nice starter with relevant information,\nkeep up the good work!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 926168,
          "author_name": "redwankarimsony",
          "author_url": "",
          "post_date": "07/12/2020 14:45:10",
          "content": "<p>Thanks <a href=\"/sandorkonya\">@sandorkonya</a> for your words of advice.. 👍 </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1004976,
      "author_name": "bhaguanand",
      "author_url": "",
      "post_date": "09/10/2020 07:14:07",
      "content": "<p>Nice.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "925140": "If you haven't done yet, please read my first insight **[here](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165727)**, in this one\n###i'll show you how to get more baseline variables that had been shown to determine the progression of the disease!\n\nAs mentioned in the first insight, almost every approach that derives features from the lung tissue is based on lung tissue segmentation, so I assume you have your CT images already segmented into a lung segmentation mask &amp; background. You are now diligently working on generating features based on the images but you get a plethora of features and ... well only 178 patient... \n\nBefore you get highly overfitted to the frustration (pun intended), i will try to give you some hints for manual feature engineering.\n\n![slices](http://drkonya.com/projects/kaggle/fibrosis_6.png)\n*image courtesy of [ Sai](https://www.kaggle.com/sainatarajan7)\n\n###First some theory: \n\n&gt; Restrictive disorders are characterised by a loss in lung volume. [...] This occurs in pulmonary fibrosis, pleural disease, chest wall disorders (kyphoscoliosis), neuromuscular disorders, pneumonectomy, pulmonary oedema and obesity, to name a few.  - from [here](https://breathe.ersjournals.com/content/8/3/232#ref-7)\n\nThese disorders show themselves through the changes in spirometric (and many other) tests.\n\n###Results obtained from lung function tests have no meaning -\n\nunless they are compared against reference values (predicted values) , and exactly this is the way we can tell that some values are reduced!\n\n **Forced vital capacity (FVC)**,  - the maximum amount of air that can be exhaled when blowing out as fast as possible - is given in mililitre and as a percentage in this contest, the percentage is the percentage of the predicted value. \nIn adults, age, height, sex and race are the main determinants of the reference values for spirometric measurements.\n\n&gt; A restrictive ventilatory defect, defined by a reduction in static (TLC) and/or operating (VC) lung volumes, is typical in patients with IPF as in other ILDs. Reduction of lung compliance is key to restriction because both chest wall compliance and respiratory muscle strength, as assessed by measurements of transdiaphragmatic pressure and maximal inspiratory pressure at the mouth , are mostly preserved.\nRestriction is often absent at the time of diagnosis. In 96 patients with biopsy-confirmed IPF, forced vital capacity (FVC) ranged from 26% to 112% pred, while TLC ranged from 42% to 125% pred. In recent clinical trials, mean FVC was close to 80% pred, consistent with half of patients having normal operating volumes. These elements indicate poor sensitivity of lung volume measurements for the diagnosis of IPF. Although restriction of operating lung volumes is consistently associated with an increased risk of death, it correlates weakly with dyspnoea or an altered quality of life in IPF, consistent with other physiological alterations also playing key roles in clinical expression of the disease. - from [here](https://err.ersjournals.com/content/27/147/170062)\n\n(--&gt; to summarise the above text: chest wall is not important, only lung parenchyma, FVC is thought not to be a good determinant for the DIAGNOSIS - not the follow up!)\n\nIf we assume that in this competition a european cohort had been diagnosed (see OSIC website), we may use the [spirometric references of the European Coal and Steel Community]( https://erj.ersjournals.com/content/31/3/687.2)  through which - since we already know the gender and age of the patient -  **we can derive the height of the patient** . \n\n**Body mass index (BMI)** was furthermore investigated as one of the determinants of the progression (see previous insight) as baseline characteristic and we may try to make assumptions on this, see [this study](https://pubs.rsna.org/doi/abs/10.1148/radiol.2283020095?uritype=cgi&amp;journalCode=radiology) or [this one](http://www.mipg.upenn.edu/yubing/Tong_PLOS2.pdf), but the most simple approaches are usually the best:\n###By determining the chest area, \nwe can categorize the patients into different BMI categories - [see reference here](https://www.researchgate.net/publication/51461844_Direct_chest_area_measurement_A_potential_anthropometric_replacement_for_BMI_to_inform_cardiac_CT_dose_parameters). Moreover the chest circumference showed also correlation to BMI.  \n\nYou may ask where should you determine the chest area... that is relatively simple to standardize if you have the lung segmented already: take the position of the first and the last image where lung is segmented and get the midposition. Chose the slice nearest to this position for your chest area. Instead of this, you may measure at 3 distinct levels (also derived from the topmost and bottom points) and make an average for your value.\n\n![midposition](http://drkonya.com/projects/kaggle/fibrosis_5.png)\n\n\nNow we have:  \n- total lung volume (see [previous insight](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165727))\n- height of the patient\n- BMI (estimate) of the patient\n- chest circumference as standalone criteria\n\n\nIn the next insight i will write about vessels and airways and their possible role in the progression determination.\n\n![inspiration](http://drkonya.com/projects/kaggle/inspi.jpg)",
    "925172": "A newbie clarification for the data preprocessing: what exactly do you mean by \n&gt; I assume you have your CT images segmented already...\nI understand the lung segmentation mask part, just want to clarify this one.\n\nThank you",
    "925182": "I am sorry if i was not clear, \n\ni was referring to my first insight, where i mentioned that almost every approach that derives features from the lung tissue is based on lung tissue segmentation (from the chest CT). \n\nSince there are multiple repositories on Github ([like this](https://github.com/JoHof/lungmask)) that do this out of the box -pretty well also in case of severely diseased lung - i assumed that everyone in the competition is done with the segmentation =)\n\nUPTADE: i changed the wording.",
    "925188": "Perfect! Thank you",
    "925224": "You've lost me at \"**manual** feature engineering\"... :)\nJust kidding, great post, thanks for sharing! \n\n(But isn't the point of Deep Learning, which powered all recent breakthroughs in Computer Vision, exactly to AVOID manual handcrafted feature engineering? :))",
    "925255": "Thank you for these amazing insights. I am a little overwhelmed but excited at the same time to try these.",
    "925417": "Thanks for the amazing insights! \n\nCan you tell me how FVC and chest circumference are related to smoking status? I've been trying to find a correlation, but this is a bit confusing for an individual with no medical background :)",
    "925699": "The correlation between the variables FVC and chest circumference is not studied yet - as far i know.\n\nThere are some points to consider:\n- the chest circumference is probably higher correlated with BMI (the fat layer sorrounding the rib cage can be sometimes pretty thick). \n- smoking is correlated to the effect of obstructive and restrictive changes in the lung, in this study we are investigating the restrictive effects.\n\nThe correlations of smoking status on FVC are usually combined in composite scores (along with gender and age).\n\nOne good work on this is probably this:\n[Idiopathic Pulmonary Fibrosis: The Association Between the Adaptive Multiple Features Method and Fibrosis Outcomes](https://pubmed.ncbi.nlm.nih.gov/27767347/)  --- (hint: they have correlation between image morphologic changes and progression to!)",
    "925703": "Thank you for the nice words!",
    "925709": "Hi there, \n\nyes, you're absolutely right about the point of DL.\nHowever, with mere 178 Patient's 178 * (30 - 350) image, each having a 512*512 pixel, good luck finding correlations =).\nThis is the point where apriori knowledge is essential... but maybe i'm wrong about this.",
    "926081": "When I saw 176 patients, at first I thought the same... but then I read the following in a paper from DeepMind:\n\n\"In many biomedical applications, only very few images are required to train a network that generalizes reasonably well. This is because each image already comprises repetitive structures with corresponding variation. In volumetric images, this effect is further pronounced, such that **we can train a network on just two volumetric images in order to generalize to a third one**. A weighted loss function and special data augmentation enable us to train the network with only few manually annotated slices, i.e., from sparsely annotated training data.\"\n\nÇiçek, Ö., Abdulkadir, A., Lienkamp, S. S., Brox, T., &amp; Ronneberger, O. (2016, October). 3D U-Net: learning dense volumetric segmentation from sparse annotation. In *International conference on medical image computing and computer-assisted intervention* (pp. 424-432). Springer, Cham",
    "926118": "It is not about segmenting the diseased lung from the chest CT... that is no problem... \n\nThe problem is to derive the good features from this segmented volumetric data and correlate them to the overall progression of the disease.",
    "926133": "Thank you very much @sandorkonya for a deeper insight into the practical dataset. \n\nThough I am no medical diagnosis expert, I made a simple introduction here [Pulmonary fibrosis for Non-Med People](https://www.kaggle.com/redwankarimsony/pulmonary-fibrosis-for-non-med-people) to Pulmonary fibrosis collecting bits and pieces from here and there on the internet. Can you please have a look and point out if I have something wrong? \n\nThanks in advance. \n\nNB: This is not a publicity request.",
    "926137": "Finding out this correlation is the main task of the ML or Data Science Engineers I guess.",
    "926155": "Dear @redwankarimsony , nice starter with relevant information,\nkeep up the good work!",
    "926161": "I know the problem is not segmentation, the problem here is different. That is not the point. The point I am making is that, in the cited paper, they proved they can use a ConvNet to extract latent features from very few samples of 3D medical images. Once you have latent features, you can tweak the net to tackle any problem you want (classification, segmentation, etc). Or just use the features ;)\n\nJust uploaded 12 GB of pre-processed tensors to a Kaggle dataset. Let's see if I'm lucky in training a ConvNet to find such features...",
    "926168": "Thanks @sandorkonya for your words of advice.. 👍",
    "932092": "carlossouza Were you lucky in training the model to learn the features?",
    "932132": "Yes... The model learns latent features. Check the Auto Encoder inputs/outputs below:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F915913%2F762097554f8bb3808687aab5e3eb2393%2Fsample.gif?generation=1594923576774616&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F915913%2F237b23ecf1eb6e61163f0951b0daf984%2Foutput.gif?generation=1594923596122095&amp;alt=media)\n\nThis was my latest experiment, using properly masked lungs as inputs, instead of raw CT scan images. Now, next steps are:\n1. Implement a baseline model using only tabular data\n2. Implement a model combining tabular data and the latent features learned\n3. Evaluate how much the latent features are improving the baseline model",
    "932144": "carlossouza, nice!\nKeep up the good work &amp; keep us informed!",
    "932165": "Those masks of yours seem very granular and detailed. I also resampled the images to make all slice_thickness and pixel spacings as 1mm and created masks with u-net from them with a fixed threshold of -500. Did some post-processing to remove noise. My masks have the \"right\" contours but are not detailed in the inside of the lung as yours. Can you give us some hints of what you did differently?\n\nThanks.",
    "932195": "Sure! Pre-processing was very painful. Getting the transformations right, then get the right order of transformations made a big difference. Here's a code snippet from my codebase that illustrates my overall process:\n\n```\ndata = CTScansDataset(\n   root_dir='../data/train',\n   transform=transforms.Compose([\n      CropBoundingBox(),\n      ConvertToHU(),\n      Resize((40, 256, 256)),\n      Clip(bounds=(-1000, 500)),\n      Mask(method=MaskMethod.MORPHOLOGICAL, threshold=-500),\n      Normalize(bounds=(-1000, -500)),\n      ToTensor()\n   ]))\n```\n\nFirst, CTScanDataset reads all subfolders in the root dir, stacks all pixel_arrays from all DICOM files within each subfolder in a 3D numpy array for each patient, with varied dimensions (median is 90 x 512 x 512, but varies a lot). Then, the preprocessing begins, following the order you see in the code. They are straightforward. On masking, the 5th step, I tried 2 methods: a morphological one, following [this great notebook](https://www.kaggle.com/miklgr500/unsupervise-lung-detection) by @miklgr500 , and a deep learning one, using [this pre-trained model](https://github.com/JoHof/lungmask). The problem with the deep learning model is that it is extremely slow. For the sake of speed, I'm sticking to the morphological approach.\n\nThings that were not trivial: figuring out the data types to avoid loss on conversions; figuring out the best method to resize/interpolate; and putting everything together in the right order.\n\nPS: If you resample using the wrong method/parameters, you can either i) end up with very poor resolution images, or ii) end up with huge tensors that are not in the same dimensions/impossible to load in GPU memory..\n\nHope it helps!",
    "932211": "It definitely helped a lot. I've used the same deep learning model you mentioned. What I did was to run and save the binary masks generated from the resampled images ( I've ran the DL model just once for the whole dataset in ~40mins) and use them to index the resampled arrays and them transform to HU. Your approach with a pipeline of transforms is a lot more flexible. My resampled images are a lot more compact than the original ones but don't seem to have lost a lot of information. ( Visual inspection is not a very robust metric ).\n\n I'll definitely reevaluate them given your insight. \n\nThank you very much",
    "977789": "sandorkonya  many thanks for insight however few questions\nin some cases see even too low FVC like less than 2500,2000  are conisdered 90 to 100 above normal values so a normal value for reference cases is just 1500... shall we consider such low absolute value to be indicative of normal state of Lungs breathing capacity ?",
    "978203": "jaideepvalani thanks for pointing this out. I will try to explore more about this and post my findings here.",
    "978204": "ronaldokun Thanks for the appreciation. I am overwhelmed too cause this is the first competition for me of this sort.",
    "979191": "redwankarimsony   \n\"we already know the gender and age of the patient - we can derive the height of the patient .\"\n\nHow do we determine height here, height is kind of random distribution or gausian distribution..\nCan any field of meta data help ?",
    "979208": "ronaldokun team up.. my current score is addition of very useful non img features ,i think your approach and mine should help go further",
    "979945": "jaideepvalani see [here](https://erj.ersjournals.com/content/erj/11/6/1354.full.pdf) table 2 for the equations. I've tried but did not manage to get good results",
    "980483": "i think what we need is actual height not a height chart based on age to build a better correlation for model .",
    "982711": "alexj21 this is actually exactly what i meant for the height calculation.\n\n@jaideepvalani is correct, the actual height would be ideal... there should be a correlation between the size of the  vertebrae and height, so this is something that one could actually derive from the scans.",
    "982721": "sandorkonya  yes i agree but there would be a deviation from actual one based on heredity of person. \nEg. If my father and mother are tall,then my adult height could be more than same person of my age if persons parents have average height or short height. \n\nScans do have this field but unfortunately this is not entered info in scan.",
    "982738": "jaideepvalani ,\nyes. many important information is missing.",
    "982832": "1) how can we get chest area/vol \n2) voxal vol same as chest volume?",
    "987572": "sandorkonya \ncould u help me understand below terminology\n1) slice thickness\n2) slice spacings \n3) voxel volumes  (i understand this is volume of a pixel )    ,how does it propertional to above 2  directly or inversely",
    "1004976": "Nice.",
    "1028703": "Thanks you for these great posts  - helps a newbie like me very much",
    "1038568": "vishram6 ,\nyou're welcome!",
    "1041656": "UPDATE\n\nDear fellow kagglers, \n\ni posted [segmentation masks](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/189528) after the competition."
  },
  "source": "meta"
}