{
  "id": 165727,
  "title": "Domain expert's insights",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/165727",
  "author_name": "dr. Konya",
  "post_date": "2020-07-10T22:18:46.358000",
  "votes": 264,
  "comment_count": 59,
  "views": 0,
  "content": "<h3>Disclaimer</h3>\n\n<pre><code>Although the described insights are derived mainly from empirical data, there are some brain-storm parts that can bring you deep in the rabbit hole and make you work for nothing... or may lead to groundbreaking results.\n</code></pre>\n\n<h3>Structure</h3>\n\n<pre><code>1 - Goals of the present competition\n2 - Factors known to determine the outcome of the Idiopathic Pulmonary Fibrosis (IPF)\n3 - Image based features\n\n###Goals of the present competition\n</code></pre>\n\n<p>Let’s take some <strong>epidemiological details</strong> (gender, age, smoking habits), some measured parameters (Forced Vital Capacity) along with <strong>diagnostic imaging</strong> (CT) in a follow up period up to hundred(and some more) weeks and try to predict the changes of the measured lung function parameter based given a baseline value and a baseline CT.</p>\n\n<h3> </h3>\n\n<p>It is hypothetized that the <strong>image data</strong> has valuable information that not only correlate to the changes in the lung function tests but <strong>can predict the progression</strong>. \n    This connection has been already studied extensively (see references below).</p>\n\n<pre><code>![changes](http://drkonya.com/projects/kaggle/fibrosis_1.png)\n</code></pre>\n\n<h3> </h3>\n\n<p>In the Analysis of Diffuse Lung Disease we have several methods that ca be used to quantitatively analyse the images and find correlations between the derived information and the measured lung function parameters. Many of them (but not all!) rely on the segmentation of the lung tissue from the images.</p>\n\n<h3> </h3>\n\n<p>The <strong>segmentation of the lung tissue by fibrosing lung diseases is challenging</strong>: in contrary to the normal lung tissue, that has a low attenuation ranging -750-950 HU. The diseased lungs in this cohort have fibrous tissue (hence the name) that does not allow a \"simple\" threshold based segmentation of the lung parenchyma. In the previous Kaggle challenges there are many notebooks on lung segmentation (and many readily trained models on Github), so i do not go into details about it here, we will assume that you already segmented the lung.</p>\n\n<p><img src=\"http://drkonya.com/projects/kaggle/fibrosis_2.png\" alt=\"threshold\"></p>\n\n<h3>Factors known to determine the outcome of the Idiopathic Pulmonary Fibrosis (IPF)</h3>\n\n<p>Before getting to the image analysis, few words about the course of the diffuse interstitial lung disease. The diffuse ILDs are characterized by infiltration of the interstitial compartment of the lung with varying degrees of inflammation and fibrosis - this is what we can see on the CT scan.</p>\n\n<h3> </h3>\n\n<p>Several studies have investigated the relationship between biomarkers and outcomes at a single point in time. Baseline characteristics like: Male gender, age &gt; 70, diagnosis delay, grade of dyspnea, cardiovascular comorbidities, fibrosis extent on CT, pulmonary hypertension, blood test results (antibodies) have a negative effect on morbidity.  </p>\n\n<h3> </h3>\n\n<h3> </h3>\n\n<p>The decrease in FVC % is the parameter of lung function that best predicts mortality. Changes between 6 or 12 months in the percentage of FVC [...] define the worsening, stability or improvement of the disease. [...] Some studies have shown a good correlation between the pulmonary function tests and the extension of IPF (abbreviation of intestinal pulmonal disease, in this context ~ ILD)  in HRCT findings.  <a href=\"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6024649/\">Prognosis and Follow-Up of Idiopathic Pulmonary Fibrosis</a></p>\n\n<h3> </h3>\n\n<p>There are several EDA-s to this competition already that show the changes of epidemiological factors. These have to be taken into account as independent ( -or better to some degree dependent - ) variables to the features derived from the images.</p>\n\n<h3> </h3>\n\n<h3>Image based features</h3>\n\n<p>Let's assume we have our segmented lung, we want to derive information about the tissue there.\nThe first information should be the <strong>volume of the lung</strong> (let's now just forget the larger airways), since the CT is made in maximal inspiration, the so called total lung capacity can be estimated. It can be calculated as: pixel number on your lung segmentation map * pixel spacing (from DICOM header information, usually 0.625 mm/px ) * slice distance.  Slice distance ( ~= slice thickness) is derived from the difference between 2 consecutive slice's Image Position DICOM attribute - usually ranging from 1 mm to 7 mm. This gives you the volume of the lung in mm3. This information could be correlated with the FVC during the follow ups.</p>\n\n<h3> </h3>\n\n<p>Next, one could analyse the lung with simple methods, like:</p>\n\n<h3> </h3>\n\n<pre><code>**Threshold based analysis:**\nSum of pixels below / above of a given threshold in your region of interest - the more fibrous tissue present, the higher this value gets for high values.\n</code></pre>\n\n<h3> </h3>\n\n<pre><code>Further with the **analysis of the histogram:** Mean, Skew, Kurthosis.\n\n**Mean** - average value is higher if fibrous tisue is present - the more fibrous tissue, the higher is the average; don't forget, we are talking about attenuation values here, where air has -1000 and water 0, fibrous tissue up to 50-70 (when calcified then much more). \n</code></pre>\n\n<h3> </h3>\n\n<pre><code>**Skew** - normal lung is skewed to left (much more low attenuation values) whereas fibrous lung is skewed to the right (much more high attenuation values).\n</code></pre>\n\n<h3> </h3>\n\n<pre><code>**Kurthosis** - the peak of the low attenuation pixels is much much lower (since we have more higher attenuating areas instead of it).\n</code></pre>\n\n<h3> </h3>\n\n<pre><code>![histogram](http://drkonya.com/projects/kaggle/fibrosis_3.png)\n</code></pre>\n\n<h3> </h3>\n\n<pre><code>In a work from 2003, [Quantitative CT Indexes in Idiopathic Pulmonary Fibrosis: Relationship With Physiologic Impairment](https://pubmed.ncbi.nlm.nih.gov/12802000/) the authors claim that:\n</code></pre>\n\n<h3> </h3>\n\n<p>\"The lungs were isolated by using a semiautomated thresholding technique, with an upper threshold of -200 HU. [...] Pulmonary function tests (PFTs) included forced vital capacity, total lung capacity, forced expiratory volume in 1 second, and diffusing lung capacity. Moderate correlations existed between histogram features and PFT results. <strong>Kurtosis showed the greatest degree of correlation with physiologic abnormality</strong>\"  </p>\n\n<h3> </h3>\n\n<p>So <strong><em>why not to calculate these histogram features</em></strong> first - for the whole lung or later for smaller segments of the lung? But how to get smaller segments? You do not have to keep these segments restricted to anatomical constrains (lobes of the lung), you can for example define the topmost point the lung and beginning from that point divide your whole lung volume into smaller, for example 1 x 1 x 1 cm cubicles (blocks) and analyse these.</p>\n\n<h3> </h3>\n\n<pre><code>![block](http://drkonya.com/projects/kaggle/fibrosis_4.png)\n</code></pre>\n\n<h3> </h3>\n\n<p>If we assume that similar tissue changes of the lung on different position in the lung itself contribute differently to the change in the spirometry results, then we need to know some information about the position/localisation of these changes.</p>\n\n<p>One slice however contains multiple areas of different tissue changes == multiple different textures (normal tissue, altered tissue [like: honeycombing, septal thickening, emphysema, ground glass...]), so to divide these areas on one slice further down in regions is advisable. </p>\n\n<h3> </h3>\n\n<p>This further categorisation can be made with the help of trained NN (see the part on MedGift database) but also without NN:\n The idea of creating blocks comes initially <a href=\"https://arxiv.org/pdf/1811.02651.pdf\">from here</a>, and we could use this technique to analyse the blocks and get information about the texture of different regions.</p>\n\n<h3> </h3>\n\n<p>Additionaly we can obtain an additional position/localisation value of each block.</p>\n\n<h3> </h3>\n\n<p>In order to get a localisation, we need a good definable point, and i chose for this the topmost point of the lung (called apex) since it is very simple to obtain: take the first slice that contains mask of your lung, and take the midpoint of this area.  The distance between this point and the midpoint of each block is then the distance i visioned. Other points could be set to, like the division of the trachea (called carina) but obtaining that point on the CT requires more segmentation work.</p>\n\n<h3> </h3>\n\n<pre><code>![block](http://drkonya.com/projects/kaggle/fibrosis_7.png)\n</code></pre>\n\n<p>Since every people is different, we have to normalize somehow this distance (normalized distance), one possibility is to take the \"height\" of the whole lung and normalize the distance with this value. The height of the lung is simply the difference of the position of the first slice containing lung mask and position of the last slice containing lung mask.</p>\n\n<h3> </h3>\n\n<p>The changes of the \"diffuse\" lung disease are usually not \"diffuse\", they have localisation preferences (like by so called UIP pattern where subpleural and basal predominance is seen) and there are different patterns we radiologists describe, these are texture based features. These typical features are: honeycombing, reticular opacities, traction bronchiektasis, ground-glass-opacities (detailed desc for each <a href=\"https://radiopaedia.org/articles/usual-interstitial-pneumonia?lang=gb\">here</a> ).</p>\n\n<h3> </h3>\n\n<p>There are studies where texture based analysis of the lung had already been performed, for example: <a href=\"https://www.hindawi.com/journals/bmri/2019/2045432/\">Automatic Lung Segmentation Based on Texture and Deep Features of HRCT Images with Interstitial Lung Disease</a> </p>\n\n<h3> </h3>\n\n<p>Here a very important database, the <a href=\"http://medgift.hevs.ch/wordpress/databases/ild-database/\">MedGift database</a> is described,  that is a collection of manually annotated CT's of 108 patients - a very good starting point for texture based studies  - one can train a semantic segmentation model to characterise the lung tissue (without even having to segment the lung itself beforehand) and compare these changes (like percentage of honeycombing or reticular pattern of the whole lung tissue and it's correlation with FVC).</p>\n\n<h3> </h3>\n\n<p>I would start to analyse the changes in the extent of honeycombing. Why the honeycombing? Studies have shown that is has the largest influenceon lung functional tests:</p>\n\n<h3> </h3>\n\n<p>Several studies have investigated the relationship between extent of fibrosis on CT and outcomes at a single point in time. Baseline patterns, in particular honeycombing [...] using non-volumetric CT, found that the overall extent of fibrosis, defined as the extent of reticulation and honeycombing on CT, was an independent predictor of mortality in IPF. Similarly, studies  [..] found that the overall extent of fibrosis indicated a poor prognosis in populations that contained not only patients with IPF but other fibrotic IIPs. The results from these studies indicate that increasing disease extent on CT is undeniably related to adverse prognosis in IPF. [...] In one study [...] it was shown that baseline honeycombing &gt;25%, fibrosis score &gt;30% and traction bronchiectasis in all four lobes were indicators of poor outcome.  - from <a href=\"https://err.ersjournals.com/content/26/145/170051\">Evaluating disease severity in idiopathic pulmonary fibrosis</a></p>\n\n<h3> </h3>\n\n<p>Radiomics based analysis will generate you quite much data, so chose wisely how big area you are analysing with it, probably dividing your lung parenchyma into 1x1x1 cm cubicles (block building) is a good starting point.</p>\n\n<h3> </h3>\n\n<p>Importance of Airway- and vessel segmentations, coming soon.</p>\n\n<pre><code>![inspiration](http://drkonya.com/projects/kaggle/inspi.jpg)\n</code></pre>",
  "messages": [
    {
      "id": 923472,
      "postDate": "2020-07-10T22:18:46.360Z",
      "content": "<h3>Disclaimer</h3>\n\n<pre><code>Although the described insights are derived mainly from empirical data, there are some brain-storm parts that can bring you deep in the rabbit hole and make you work for nothing... or may lead to groundbreaking results.\n</code></pre>\n\n<h3>Structure</h3>\n\n<pre><code>1 - Goals of the present competition\n2 - Factors known to determine the outcome of the Idiopathic Pulmonary Fibrosis (IPF)\n3 - Image based features\n\n###Goals of the present competition\n</code></pre>\n\n<p>Let’s take some <strong>epidemiological details</strong> (gender, age, smoking habits), some measured parameters (Forced Vital Capacity) along with <strong>diagnostic imaging</strong> (CT) in a follow up period up to hundred(and some more) weeks and try to predict the changes of the measured lung function parameter based given a baseline value and a baseline CT.</p>\n\n<h3> </h3>\n\n<p>It is hypothetized that the <strong>image data</strong> has valuable information that not only correlate to the changes in the lung function tests but <strong>can predict the progression</strong>. \n    This connection has been already studied extensively (see references below).</p>\n\n<pre><code>![changes](http://drkonya.com/projects/kaggle/fibrosis_1.png)\n</code></pre>\n\n<h3> </h3>\n\n<p>In the Analysis of Diffuse Lung Disease we have several methods that ca be used to quantitatively analyse the images and find correlations between the derived information and the measured lung function parameters. Many of them (but not all!) rely on the segmentation of the lung tissue from the images.</p>\n\n<h3> </h3>\n\n<p>The <strong>segmentation of the lung tissue by fibrosing lung diseases is challenging</strong>: in contrary to the normal lung tissue, that has a low attenuation ranging -750-950 HU. The diseased lungs in this cohort have fibrous tissue (hence the name) that does not allow a \"simple\" threshold based segmentation of the lung parenchyma. In the previous Kaggle challenges there are many notebooks on lung segmentation (and many readily trained models on Github), so i do not go into details about it here, we will assume that you already segmented the lung.</p>\n\n<p><img src=\"http://drkonya.com/projects/kaggle/fibrosis_2.png\" alt=\"threshold\"></p>\n\n<h3>Factors known to determine the outcome of the Idiopathic Pulmonary Fibrosis (IPF)</h3>\n\n<p>Before getting to the image analysis, few words about the course of the diffuse interstitial lung disease. The diffuse ILDs are characterized by infiltration of the interstitial compartment of the lung with varying degrees of inflammation and fibrosis - this is what we can see on the CT scan.</p>\n\n<h3> </h3>\n\n<p>Several studies have investigated the relationship between biomarkers and outcomes at a single point in time. Baseline characteristics like: Male gender, age &gt; 70, diagnosis delay, grade of dyspnea, cardiovascular comorbidities, fibrosis extent on CT, pulmonary hypertension, blood test results (antibodies) have a negative effect on morbidity.  </p>\n\n<h3> </h3>\n\n<h3> </h3>\n\n<p>The decrease in FVC % is the parameter of lung function that best predicts mortality. Changes between 6 or 12 months in the percentage of FVC [...] define the worsening, stability or improvement of the disease. [...] Some studies have shown a good correlation between the pulmonary function tests and the extension of IPF (abbreviation of intestinal pulmonal disease, in this context ~ ILD)  in HRCT findings.  <a href=\"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6024649/\">Prognosis and Follow-Up of Idiopathic Pulmonary Fibrosis</a></p>\n\n<h3> </h3>\n\n<p>There are several EDA-s to this competition already that show the changes of epidemiological factors. These have to be taken into account as independent ( -or better to some degree dependent - ) variables to the features derived from the images.</p>\n\n<h3> </h3>\n\n<h3>Image based features</h3>\n\n<p>Let's assume we have our segmented lung, we want to derive information about the tissue there.\nThe first information should be the <strong>volume of the lung</strong> (let's now just forget the larger airways), since the CT is made in maximal inspiration, the so called total lung capacity can be estimated. It can be calculated as: pixel number on your lung segmentation map * pixel spacing (from DICOM header information, usually 0.625 mm/px ) * slice distance.  Slice distance ( ~= slice thickness) is derived from the difference between 2 consecutive slice's Image Position DICOM attribute - usually ranging from 1 mm to 7 mm. This gives you the volume of the lung in mm3. This information could be correlated with the FVC during the follow ups.</p>\n\n<h3> </h3>\n\n<p>Next, one could analyse the lung with simple methods, like:</p>\n\n<h3> </h3>\n\n<pre><code>**Threshold based analysis:**\nSum of pixels below / above of a given threshold in your region of interest - the more fibrous tissue present, the higher this value gets for high values.\n</code></pre>\n\n<h3> </h3>\n\n<pre><code>Further with the **analysis of the histogram:** Mean, Skew, Kurthosis.\n\n**Mean** - average value is higher if fibrous tisue is present - the more fibrous tissue, the higher is the average; don't forget, we are talking about attenuation values here, where air has -1000 and water 0, fibrous tissue up to 50-70 (when calcified then much more). \n</code></pre>\n\n<h3> </h3>\n\n<pre><code>**Skew** - normal lung is skewed to left (much more low attenuation values) whereas fibrous lung is skewed to the right (much more high attenuation values).\n</code></pre>\n\n<h3> </h3>\n\n<pre><code>**Kurthosis** - the peak of the low attenuation pixels is much much lower (since we have more higher attenuating areas instead of it).\n</code></pre>\n\n<h3> </h3>\n\n<pre><code>![histogram](http://drkonya.com/projects/kaggle/fibrosis_3.png)\n</code></pre>\n\n<h3> </h3>\n\n<pre><code>In a work from 2003, [Quantitative CT Indexes in Idiopathic Pulmonary Fibrosis: Relationship With Physiologic Impairment](https://pubmed.ncbi.nlm.nih.gov/12802000/) the authors claim that:\n</code></pre>\n\n<h3> </h3>\n\n<p>\"The lungs were isolated by using a semiautomated thresholding technique, with an upper threshold of -200 HU. [...] Pulmonary function tests (PFTs) included forced vital capacity, total lung capacity, forced expiratory volume in 1 second, and diffusing lung capacity. Moderate correlations existed between histogram features and PFT results. <strong>Kurtosis showed the greatest degree of correlation with physiologic abnormality</strong>\"  </p>\n\n<h3> </h3>\n\n<p>So <strong><em>why not to calculate these histogram features</em></strong> first - for the whole lung or later for smaller segments of the lung? But how to get smaller segments? You do not have to keep these segments restricted to anatomical constrains (lobes of the lung), you can for example define the topmost point the lung and beginning from that point divide your whole lung volume into smaller, for example 1 x 1 x 1 cm cubicles (blocks) and analyse these.</p>\n\n<h3> </h3>\n\n<pre><code>![block](http://drkonya.com/projects/kaggle/fibrosis_4.png)\n</code></pre>\n\n<h3> </h3>\n\n<p>If we assume that similar tissue changes of the lung on different position in the lung itself contribute differently to the change in the spirometry results, then we need to know some information about the position/localisation of these changes.</p>\n\n<p>One slice however contains multiple areas of different tissue changes == multiple different textures (normal tissue, altered tissue [like: honeycombing, septal thickening, emphysema, ground glass...]), so to divide these areas on one slice further down in regions is advisable. </p>\n\n<h3> </h3>\n\n<p>This further categorisation can be made with the help of trained NN (see the part on MedGift database) but also without NN:\n The idea of creating blocks comes initially <a href=\"https://arxiv.org/pdf/1811.02651.pdf\">from here</a>, and we could use this technique to analyse the blocks and get information about the texture of different regions.</p>\n\n<h3> </h3>\n\n<p>Additionaly we can obtain an additional position/localisation value of each block.</p>\n\n<h3> </h3>\n\n<p>In order to get a localisation, we need a good definable point, and i chose for this the topmost point of the lung (called apex) since it is very simple to obtain: take the first slice that contains mask of your lung, and take the midpoint of this area.  The distance between this point and the midpoint of each block is then the distance i visioned. Other points could be set to, like the division of the trachea (called carina) but obtaining that point on the CT requires more segmentation work.</p>\n\n<h3> </h3>\n\n<pre><code>![block](http://drkonya.com/projects/kaggle/fibrosis_7.png)\n</code></pre>\n\n<p>Since every people is different, we have to normalize somehow this distance (normalized distance), one possibility is to take the \"height\" of the whole lung and normalize the distance with this value. The height of the lung is simply the difference of the position of the first slice containing lung mask and position of the last slice containing lung mask.</p>\n\n<h3> </h3>\n\n<p>The changes of the \"diffuse\" lung disease are usually not \"diffuse\", they have localisation preferences (like by so called UIP pattern where subpleural and basal predominance is seen) and there are different patterns we radiologists describe, these are texture based features. These typical features are: honeycombing, reticular opacities, traction bronchiektasis, ground-glass-opacities (detailed desc for each <a href=\"https://radiopaedia.org/articles/usual-interstitial-pneumonia?lang=gb\">here</a> ).</p>\n\n<h3> </h3>\n\n<p>There are studies where texture based analysis of the lung had already been performed, for example: <a href=\"https://www.hindawi.com/journals/bmri/2019/2045432/\">Automatic Lung Segmentation Based on Texture and Deep Features of HRCT Images with Interstitial Lung Disease</a> </p>\n\n<h3> </h3>\n\n<p>Here a very important database, the <a href=\"http://medgift.hevs.ch/wordpress/databases/ild-database/\">MedGift database</a> is described,  that is a collection of manually annotated CT's of 108 patients - a very good starting point for texture based studies  - one can train a semantic segmentation model to characterise the lung tissue (without even having to segment the lung itself beforehand) and compare these changes (like percentage of honeycombing or reticular pattern of the whole lung tissue and it's correlation with FVC).</p>\n\n<h3> </h3>\n\n<p>I would start to analyse the changes in the extent of honeycombing. Why the honeycombing? Studies have shown that is has the largest influenceon lung functional tests:</p>\n\n<h3> </h3>\n\n<p>Several studies have investigated the relationship between extent of fibrosis on CT and outcomes at a single point in time. Baseline patterns, in particular honeycombing [...] using non-volumetric CT, found that the overall extent of fibrosis, defined as the extent of reticulation and honeycombing on CT, was an independent predictor of mortality in IPF. Similarly, studies  [..] found that the overall extent of fibrosis indicated a poor prognosis in populations that contained not only patients with IPF but other fibrotic IIPs. The results from these studies indicate that increasing disease extent on CT is undeniably related to adverse prognosis in IPF. [...] In one study [...] it was shown that baseline honeycombing &gt;25%, fibrosis score &gt;30% and traction bronchiectasis in all four lobes were indicators of poor outcome.  - from <a href=\"https://err.ersjournals.com/content/26/145/170051\">Evaluating disease severity in idiopathic pulmonary fibrosis</a></p>\n\n<h3> </h3>\n\n<p>Radiomics based analysis will generate you quite much data, so chose wisely how big area you are analysing with it, probably dividing your lung parenchyma into 1x1x1 cm cubicles (block building) is a good starting point.</p>\n\n<h3> </h3>\n\n<p>Importance of Airway- and vessel segmentations, coming soon.</p>\n\n<pre><code>![inspiration](http://drkonya.com/projects/kaggle/inspi.jpg)\n</code></pre>",
      "rawMarkdown": "\n\n\n### Disclaimer\n\n\tAlthough the described insights are derived mainly from empirical data, there are some brain-storm parts that can bring you deep in the rabbit hole and make you work for nothing... or may lead to groundbreaking results.\n\n### Structure\n\n\t1 - Goals of the present competition\n\t2 - Factors known to determine the outcome of the Idiopathic Pulmonary Fibrosis (IPF)\n\t3 - Image based features\n\n\t###Goals of the present competition\n\nLet’s take some **epidemiological details** (gender, age, smoking habits), some measured parameters (Forced Vital Capacity) along with **diagnostic imaging** (CT) in a follow up period up to hundred(and some more) weeks and try to predict the changes of the measured lung function parameter based given a baseline value and a baseline CT.\n### \n\nIt is hypothetized that the **image data** has valuable information that not only correlate to the changes in the lung function tests but **can predict the progression**. \n\tThis connection has been already studied extensively (see references below).\n\n\t![changes](http://drkonya.com/projects/kaggle/fibrosis_1.png)\n### \n\nIn the Analysis of Diffuse Lung Disease we have several methods that ca be used to quantitatively analyse the images and find correlations between the derived information and the measured lung function parameters. Many of them (but not all!) rely on the segmentation of the lung tissue from the images.\n### \n\nThe **segmentation of the lung tissue by fibrosing lung diseases is challenging**: in contrary to the normal lung tissue, that has a low attenuation ranging -750-950 HU. The diseased lungs in this cohort have fibrous tissue (hence the name) that does not allow a \"simple\" threshold based segmentation of the lung parenchyma. In the previous Kaggle challenges there are many notebooks on lung segmentation (and many readily trained models on Github), so i do not go into details about it here, we will assume that you already segmented the lung.\n\n![threshold](http://drkonya.com/projects/kaggle/fibrosis_2.png)\n\n###Factors known to determine the outcome of the Idiopathic Pulmonary Fibrosis (IPF)\n\n\nBefore getting to the image analysis, few words about the course of the diffuse interstitial lung disease. The diffuse ILDs are characterized by infiltration of the interstitial compartment of the lung with varying degrees of inflammation and fibrosis - this is what we can see on the CT scan.\n\n### \n\nSeveral studies have investigated the relationship between biomarkers and outcomes at a single point in time. Baseline characteristics like: Male gender, age &gt; 70, diagnosis delay, grade of dyspnea, cardiovascular comorbidities, fibrosis extent on CT, pulmonary hypertension, blood test results (antibodies) have a negative effect on morbidity.  \n### \n### \nThe decrease in FVC % is the parameter of lung function that best predicts mortality. Changes between 6 or 12 months in the percentage of FVC [...] define the worsening, stability or improvement of the disease. [...] Some studies have shown a good correlation between the pulmonary function tests and the extension of IPF (abbreviation of intestinal pulmonal disease, in this context ~ ILD)  in HRCT findings. \t[Prognosis and Follow-Up of Idiopathic Pulmonary Fibrosis](https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6024649/)\n\n### \nThere are several EDA-s to this competition already that show the changes of epidemiological factors. These have to be taken into account as independent ( -or better to some degree dependent - ) variables to the features derived from the images.\n### \n\n###Image based features\n\nLet's assume we have our segmented lung, we want to derive information about the tissue there.\nThe first information should be the **volume of the lung** (let's now just forget the larger airways), since the CT is made in maximal inspiration, the so called total lung capacity can be estimated. It can be calculated as: pixel number on your lung segmentation map * pixel spacing (from DICOM header information, usually 0.625 mm/px ) * slice distance.  Slice distance ( ~= slice thickness) is derived from the difference between 2 consecutive slice's Image Position DICOM attribute - usually ranging from 1 mm to 7 mm. This gives you the volume of the lung in mm3. This information could be correlated with the FVC during the follow ups.\n### \n\nNext, one could analyse the lung with simple methods, like:\n### \n\n\t**Threshold based analysis:**\n\tSum of pixels below / above of a given threshold in your region of interest - the more fibrous tissue present, the higher this value gets for high values.\n### \n\n\tFurther with the **analysis of the histogram:** Mean, Skew, Kurthosis.\n\n\t**Mean** - average value is higher if fibrous tisue is present - the more fibrous tissue, the higher is the average; don't forget, we are talking about attenuation values here, where air has -1000 and water 0, fibrous tissue up to 50-70 (when calcified then much more). \n### \n\n\t**Skew** - normal lung is skewed to left (much more low attenuation values) whereas fibrous lung is skewed to the right (much more high attenuation values).\n### \n\n\t**Kurthosis** - the peak of the low attenuation pixels is much much lower (since we have more higher attenuating areas instead of it).\n### \n\n\t![histogram](http://drkonya.com/projects/kaggle/fibrosis_3.png)\n\n### \n\tIn a work from 2003, [Quantitative CT Indexes in Idiopathic Pulmonary Fibrosis: Relationship With Physiologic Impairment](https://pubmed.ncbi.nlm.nih.gov/12802000/) the authors claim that:\n### \n\n\"The lungs were isolated by using a semiautomated thresholding technique, with an upper threshold of -200 HU. [...] Pulmonary function tests (PFTs) included forced vital capacity, total lung capacity, forced expiratory volume in 1 second, and diffusing lung capacity. Moderate correlations existed between histogram features and PFT results. **Kurtosis showed the greatest degree of correlation with physiologic abnormality**\"  \n### \n\nSo ***why not to calculate these histogram features*** first - for the whole lung or later for smaller segments of the lung? But how to get smaller segments? You do not have to keep these segments restricted to anatomical constrains (lobes of the lung), you can for example define the topmost point the lung and beginning from that point divide your whole lung volume into smaller, for example 1 x 1 x 1 cm cubicles (blocks) and analyse these.\n\n### \n\n\t![block](http://drkonya.com/projects/kaggle/fibrosis_4.png)\n\n### \nIf we assume that similar tissue changes of the lung on different position in the lung itself contribute differently to the change in the spirometry results, then we need to know some information about the position/localisation of these changes.\n\nOne slice however contains multiple areas of different tissue changes == multiple different textures (normal tissue, altered tissue [like: honeycombing, septal thickening, emphysema, ground glass...]), so to divide these areas on one slice further down in regions is advisable. \n### \n\nThis further categorisation can be made with the help of trained NN (see the part on MedGift database) but also without NN:\n The idea of creating blocks comes initially [from here](https://arxiv.org/pdf/1811.02651.pdf), and we could use this technique to analyse the blocks and get information about the texture of different regions.\n### \n\n\nAdditionaly we can obtain an additional position/localisation value of each block.\n### \n\nIn order to get a localisation, we need a good definable point, and i chose for this the topmost point of the lung (called apex) since it is very simple to obtain: take the first slice that contains mask of your lung, and take the midpoint of this area.  The distance between this point and the midpoint of each block is then the distance i visioned. Other points could be set to, like the division of the trachea (called carina) but obtaining that point on the CT requires more segmentation work.\n### \n\t![block](http://drkonya.com/projects/kaggle/fibrosis_7.png)\n\nSince every people is different, we have to normalize somehow this distance (normalized distance), one possibility is to take the \"height\" of the whole lung and normalize the distance with this value. The height of the lung is simply the difference of the position of the first slice containing lung mask and position of the last slice containing lung mask.\n### \n\nThe changes of the \"diffuse\" lung disease are usually not \"diffuse\", they have localisation preferences (like by so called UIP pattern where subpleural and basal predominance is seen) and there are different patterns we radiologists describe, these are texture based features. These typical features are: honeycombing, reticular opacities, traction bronchiektasis, ground-glass-opacities (detailed desc for each [here](https://radiopaedia.org/articles/usual-interstitial-pneumonia?lang=gb) ).\n\n\n### \n\nThere are studies where texture based analysis of the lung had already been performed, for example: [Automatic Lung Segmentation Based on Texture and Deep Features of HRCT Images with Interstitial Lung Disease](https://www.hindawi.com/journals/bmri/2019/2045432/  ) \n\n### \n\nHere a very important database, the [MedGift database](http://medgift.hevs.ch/wordpress/databases/ild-database/) is described,  that is a collection of manually annotated CT's of 108 patients - a very good starting point for texture based studies  - one can train a semantic segmentation model to characterise the lung tissue (without even having to segment the lung itself beforehand) and compare these changes (like percentage of honeycombing or reticular pattern of the whole lung tissue and it's correlation with FVC).\n\n### \n\nI would start to analyse the changes in the extent of honeycombing. Why the honeycombing? Studies have shown that is has the largest influenceon lung functional tests:\n### \n\nSeveral studies have investigated the relationship between extent of fibrosis on CT and outcomes at a single point in time. Baseline patterns, in particular honeycombing [...] using non-volumetric CT, found that the overall extent of fibrosis, defined as the extent of reticulation and honeycombing on CT, was an independent predictor of mortality in IPF. Similarly, studies  [..] found that the overall extent of fibrosis indicated a poor prognosis in populations that contained not only patients with IPF but other fibrotic IIPs. The results from these studies indicate that increasing disease extent on CT is undeniably related to adverse prognosis in IPF. [...] In one study [...] it was shown that baseline honeycombing &gt;25%, fibrosis score &gt;30% and traction bronchiectasis in all four lobes were indicators of poor outcome.  - from [Evaluating disease severity in idiopathic pulmonary fibrosis](https://err.ersjournals.com/content/26/145/170051)\n\n### \n\nRadiomics based analysis will generate you quite much data, so chose wisely how big area you are analysing with it, probably dividing your lung parenchyma into 1x1x1 cm cubicles (block building) is a good starting point.\n### \n\nImportance of Airway- and vessel segmentations, coming soon.\n\n\t![inspiration](http://drkonya.com/projects/kaggle/inspi.jpg)\n\n",
      "votes": 263
    },
    {
      "id": 1041652,
      "postDate": "2020-10-07T21:28:35.170Z",
      "content": "<p>UPDATE</p>\n<p>Dear fellow kagglers, </p>\n<p>i posted <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/189528\" target=\"_blank\">segmentation masks</a> after the competition.</p>",
      "rawMarkdown": "UPDATE\n\nDear fellow kagglers, \n\ni posted [segmentation masks](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/189528) after the competition.",
      "votes": 1
    },
    {
      "id": 925161,
      "postDate": "2020-07-11T21:15:53.363Z",
      "content": "<p>Very insightful. Thank you!</p>",
      "rawMarkdown": "Very insightful. Thank you!",
      "votes": 1,
      "replies": [
        {
          "id": 925162,
          "postDate": "2020-07-11T21:17:07.883Z",
          "content": "<p>Thank you, be sure you read my <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/166123\">second one</a> to!</p>",
          "rawMarkdown": "Thank you, be sure you read my [second one](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/166123) to!",
          "votes": 1
        },
        {
          "id": 925165,
          "postDate": "2020-07-11T21:19:03.893Z",
          "content": "<p>I'm on it right now👊 </p>",
          "rawMarkdown": "I'm on it right now👊 "
        }
      ]
    },
    {
      "id": 924666,
      "postDate": "2020-07-11T15:10:46.033Z",
      "content": "<p>Very informative. Seems like the first step is the segmentation, a task that still seems to have a precision not quite high enough as what would be wished in the field</p>",
      "rawMarkdown": "Very informative. Seems like the first step is the segmentation, a task that still seems to have a precision not quite high enough as what would be wished in the field",
      "votes": 1,
      "replies": [
        {
          "id": 924741,
          "postDate": "2020-07-11T15:55:11.700Z",
          "content": "<p>We suppose (or better presume) that the lung parenchyma is more responsible for the changes in the pulmonary function tests than other parts of the chest (muscles, ribs, fat), so the segmentation is the first  dimensionality reduction. \nThere are more information (about BMI, motion constraints of the chest's bony cage due to degeneration of the ribs, shape of diaphragm, calcifications as sign for cardiovascular comorbidity) that are then not taken into account, but they should be addressed.</p>\n\n<p>So while segmentatiin of the lungs is one way, it helps by crafting manual features that lead to further dimensionality reduction... and maybe we can derive clinically relevant (and correlating) variables but not the only way one can try.</p>",
          "rawMarkdown": "We suppose (or better presume) that the lung parenchyma is more responsible for the changes in the pulmonary function tests than other parts of the chest (muscles, ribs, fat), so the segmentation is the first  dimensionality reduction. \nThere are more information (about BMI, motion constraints of the chest's bony cage due to degeneration of the ribs, shape of diaphragm, calcifications as sign for cardiovascular comorbidity) that are then not taken into account, but they should be addressed.\n\nSo while segmentatiin of the lungs is one way, it helps by crafting manual features that lead to further dimensionality reduction... and maybe we can derive clinically relevant (and correlating) variables but not the only way one can try.",
          "votes": 1
        },
        {
          "id": 927150,
          "postDate": "2020-07-13T08:22:35.360Z",
          "content": "<p>Thank you! Just a question: when calculating the volume, shouldn't pixel_spacing be pixel_spacing^2 (assuming that row spacing equals column spacing)?</p>",
          "rawMarkdown": "Thank you! Just a question: when calculating the volume, shouldn't pixel_spacing be pixel_spacing^2 (assuming that row spacing equals column spacing)?",
          "votes": 1
        },
        {
          "id": 929242,
          "postDate": "2020-07-14T14:51:48.150Z",
          "content": "<p><a href=\"/nickgm\">@nickgm</a> ,\nyes, you're obviously right!\nThe CT slices have usually equal row &amp; col spacing.</p>",
          "rawMarkdown": "@nickgm ,\nyes, you're obviously right!\nThe CT slices have usually equal row &amp; col spacing."
        }
      ]
    },
    {
      "id": 958201,
      "postDate": "2020-08-04T20:50:12.490Z",
      "content": "<p>Thanks for your amazing insights. I implemented some of them in my <a href=\"https://www.kaggle.com/unforgiven/osic-comprehensive-eda\">EDA notebook</a>, if you want to take a look.</p>",
      "rawMarkdown": "Thanks for your amazing insights. I implemented some of them in my [EDA notebook](https://www.kaggle.com/unforgiven/osic-comprehensive-eda), if you want to take a look.",
      "votes": 2,
      "replies": [
        {
          "id": 962858,
          "postDate": "2020-08-08T13:50:49.563Z",
          "content": "<p><a href=\"https://www.kaggle.com/unforgiven\" target=\"_blank\">@unforgiven</a>,<br>\ni took a look and it is 👍👍, also left a comment.</p>",
          "rawMarkdown": "@unforgiven,\ni took a look and it is 👍👍, also left a comment."
        }
      ]
    },
    {
      "id": 925141,
      "postDate": "2020-07-11T20:57:12.617Z",
      "content": "<h3>UPDATE see the second part of my insights here:</h3>\n\n<p><a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/166123\">https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/166123</a></p>",
      "rawMarkdown": "###UPDATE see the second part of my insights here: \nhttps://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/166123",
      "votes": 2,
      "replies": [
        {
          "id": 960593,
          "postDate": "2020-08-06T14:24:23.667Z",
          "content": "<p>Thank you for the insightful thread!</p>\n\n<p>Did you post a thread discussing airway and vessel segmentation?</p>",
          "rawMarkdown": "Thank you for the insightful thread!\n\n Did you post a thread discussing airway and vessel segmentation?"
        },
        {
          "id": 962842,
          "postDate": "2020-08-08T13:32:51.953Z",
          "content": "<p><a href=\"https://www.kaggle.com/niksapraljak\" target=\"_blank\">@niksapraljak</a> ,<br>\nwe discussed it in the team and decided not to do it as for now. <br>\nNearing the end of the competition however i plan to make an in-depth thread about it.</p>",
          "rawMarkdown": "@niksapraljak ,\nwe discussed it in the team and decided not to do it as for now. \nNearing the end of the competition however i plan to make an in-depth thread about it."
        }
      ]
    },
    {
      "id": 1028702,
      "postDate": "2020-09-27T05:57:36.597Z",
      "content": "<p>Thank you for your informative post - Really helps for someone new like me.</p>",
      "rawMarkdown": "Thank you for your informative post - Really helps for someone new like me."
    },
    {
      "id": 1015532,
      "postDate": "2020-09-18T08:44:50.847Z",
      "content": "<p>Thanks for very useful insights!<br>\nHowever, I have two questions for you.</p>\n<p>1.<br>\n    You wrote \"It (the volume of the lung) can be calculated as: pixel number on your lung segmentation map * pixel spacing (from DICOM header information, usually 0.625 mm/px ) * slice distance\".<br>\n    But, I think it is calculated as: pixel number on your lung segmentation map * pixel spacing (from DICOM header information, usually 0.625 mm/px )** 2 * slice distance\". (pixel spacing is multiplyed twice)<br>\n    Is this wrong?</p>\n<p>2.<br>\n    You wrote \"Slice distance ( ~= slice thickness) is derived from the difference between 2 consecutive slice's Image Position DICOM attribute - usually ranging from 1 mm to 7 mm. \"<br>\n    How to get position of DICOM?</p>\n<p>I would be grateful if you could answer to these questions.</p>",
      "rawMarkdown": "Thanks for very useful insights!\nHowever, I have two questions for you.\n\n1.\n    You wrote \"It (the volume of the lung) can be calculated as: pixel number on your lung segmentation map * pixel spacing (from DICOM header information, usually 0.625 mm/px ) * slice distance\".\n    But, I think it is calculated as: pixel number on your lung segmentation map * pixel spacing (from DICOM header information, usually 0.625 mm/px )** 2 * slice distance\". (pixel spacing is multiplyed twice)\n    Is this wrong?\n\n2.\n    You wrote \"Slice distance ( ~= slice thickness) is derived from the difference between 2 consecutive slice's Image Position DICOM attribute - usually ranging from 1 mm to 7 mm. \"\n    How to get position of DICOM?\n\nI would be grateful if you could answer to these questions.",
      "replies": [
        {
          "id": 1016358,
          "postDate": "2020-09-18T21:57:43.953Z",
          "content": "<p><a href=\"https://www.kaggle.com/koichirokamada\" target=\"_blank\">@koichirokamada</a> ,<br>\nyou are right pixel spacing is squared, check the commen below for <a href=\"https://www.kaggle.com/ronaldokun\" target=\"_blank\">@ronaldokun</a>.<br>\nSlice location comes from this DICOM tag:<br>\n<a href=\"https://dicom.innolitics.com/ciods/ct-image/image-plane/00201041\" target=\"_blank\">https://dicom.innolitics.com/ciods/ct-image/image-plane/00201041</a><br>\nYou are interested in the third, [2], z-position. <br>\nGood luck!</p>",
          "rawMarkdown": "@koichirokamada ,\nyou are right pixel spacing is squared, check the commen below for @ronaldokun.\nSlice location comes from this DICOM tag:\nhttps://dicom.innolitics.com/ciods/ct-image/image-plane/00201041\nYou are interested in the third, [2], z-position. \nGood luck!",
          "votes": 1
        },
        {
          "id": 1016501,
          "postDate": "2020-09-19T03:14:46.523Z",
          "content": "<p>I see.<br>\nThank you very much!</p>",
          "rawMarkdown": "I see.\nThank you very much!"
        },
        {
          "id": 1019883,
          "postDate": "2020-09-20T18:20:17.653Z",
          "content": "<p><a href=\"https://www.kaggle.com/sandorkonya\" target=\"_blank\">@sandorkonya</a> So if I understood correctly, then if we \"resample\" the scan so that it has a new spacing of  1x1x1mm, then calculating the volume would simply require to sum the number of pixels of the segmented lungs, correct?</p>",
          "rawMarkdown": "@sandorkonya So if I understood correctly, then if we \"resample\" the scan so that it has a new spacing of  1x1x1mm, then calculating the volume would simply require to sum the number of pixels of the segmented lungs, correct?"
        }
      ]
    },
    {
      "id": 1000721,
      "postDate": "2020-09-06T18:31:14.240Z",
      "content": "<p>Thanks again for the insights, <a href=\"https://www.kaggle.com/sandorkonya\" target=\"_blank\">@sandorkonya</a> !</p>\n<p>I'd love to pick your brain on ideas to move forward. So far, the most promising approach IMHO is derived from your ideas. </p>\n<p>Let's assume FVC declines are lines, from your first image, which alphas (intercepts) and betas (slope, i.e. the severity of the decline). Thus, the objective of the competition could be restated as predicting the beta for each patient from the baseline measurements &amp; CT scan. If we have alphas &amp; betas well predicted for each patient, we can infer FVCs for every possible week.</p>\n<p>Using a Bayesian approach, we can learn that the 176 test patients have a distribution of betas normally distributed with center in -4.2 and std dev of 5 (all approx). This very simple to obtain. Then, if I sample a beta from this distribution for each test patient, and make their lines pass through their baseline FVC, and calculate the sigmas properly (I demonstrated in my last notebook), this yields a score of -6.90. That's our minimum score to beat. Anything worse than that is pure nonsense.</p>\n<p>Now, enter the CT scans. As you correctly pointed out, the job is to predict the severity of decline in FVC (betas) using features derived from the CT scans. So far, I implemented several features you suggested (volume in L, chest circumference in mm, height, HU histogram moments (mean, stddev, skew, kurtosis), % of lung volume in different ranges of HU, etc). I also tried deriving latent features learned with unsupervised learning (using a Variational Autoencoder, not public yet). </p>\n<p>The problem is that this dataset with 176 CT scans with 176 betas is too small. There's too much noise: variance is very high, such that the gains for using all this noisy data are marginal in comparison to assuming all test patients have betas sampled from the same global distribution (i.e. all have the same severity). Even if I transform the problem into a classification problem (patients with betas lower than distribution median are \"high severity\", and higher than median are \"low severity\"), best I get is ~60-70% accuracy, loo low.</p>\n<p>You mentioned the blocks idea. I'd love to use the MedGift database to classify every voxel of our patients, but one of the organizers said it is off-limits because of their license agreement.</p>\n<p>Anyway, I'm getting out of ideas on how to improve…</p>",
      "rawMarkdown": "Thanks again for the insights, @sandorkonya !\n\nI'd love to pick your brain on ideas to move forward. So far, the most promising approach IMHO is derived from your ideas. \n\nLet's assume FVC declines are lines, from your first image, which alphas (intercepts) and betas (slope, i.e. the severity of the decline). Thus, the objective of the competition could be restated as predicting the beta for each patient from the baseline measurements & CT scan. If we have alphas & betas well predicted for each patient, we can infer FVCs for every possible week.\n\nUsing a Bayesian approach, we can learn that the 176 test patients have a distribution of betas normally distributed with center in -4.2 and std dev of 5 (all approx). This very simple to obtain. Then, if I sample a beta from this distribution for each test patient, and make their lines pass through their baseline FVC, and calculate the sigmas properly (I demonstrated in my last notebook), this yields a score of -6.90. That's our minimum score to beat. Anything worse than that is pure nonsense.\n\nNow, enter the CT scans. As you correctly pointed out, the job is to predict the severity of decline in FVC (betas) using features derived from the CT scans. So far, I implemented several features you suggested (volume in L, chest circumference in mm, height, HU histogram moments (mean, stddev, skew, kurtosis), % of lung volume in different ranges of HU, etc). I also tried deriving latent features learned with unsupervised learning (using a Variational Autoencoder, not public yet). \n\nThe problem is that this dataset with 176 CT scans with 176 betas is too small. There's too much noise: variance is very high, such that the gains for using all this noisy data are marginal in comparison to assuming all test patients have betas sampled from the same global distribution (i.e. all have the same severity). Even if I transform the problem into a classification problem (patients with betas lower than distribution median are \"high severity\", and higher than median are \"low severity\"), best I get is ~60-70% accuracy, loo low.\n \nYou mentioned the blocks idea. I'd love to use the MedGift database to classify every voxel of our patients, but one of the organizers said it is off-limits because of their license agreement.\n\nAnyway, I'm getting out of ideas on how to improve...",
      "replies": [
        {
          "id": 1000861,
          "postDate": "2020-09-06T21:00:16.767Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/carlossouza\" target=\"_blank\">@carlossouza</a> ,</p>\n<p>you're right, we were also struggling with the high heterogenity of the dataset. It is no secret that the correct cropping &amp; resampling of the dataset took up to 90% our time.<br>\nWe fancied the idea of tackling this as a classification problem to, but no image derived feature seems to be strong enough (yet) to be a good classifier.</p>\n<p>I managed to manually segment 100 lungs (this enables us to get almost perfect segmentation every time) , moreover I made &gt;30k markups for reticulation, honeycombing, airways, normal tissue… we are integrating these features into our existing pipeline, we will see how it performs.</p>\n<p><img src=\"http://www.drkonya.com/projects/kaggle/fibrosis_9.JPG\" alt=\"\"></p>",
          "rawMarkdown": "Hi @carlossouza ,\n\nyou're right, we were also struggling with the high heterogenity of the dataset. It is no secret that the correct cropping & resampling of the dataset took up to 90% our time.\nWe fancied the idea of tackling this as a classification problem to, but no image derived feature seems to be strong enough (yet) to be a good classifier.\n\nI managed to manually segment 100 lungs (this enables us to get almost perfect segmentation every time) , moreover I made >30k markups for reticulation, honeycombing, airways, normal tissue... we are integrating these features into our existing pipeline, we will see how it performs.\n\n![](http://www.drkonya.com/projects/kaggle/fibrosis_9.JPG)"
        }
      ]
    },
    {
      "id": 977444,
      "postDate": "2020-08-19T13:33:48.437Z",
      "content": "<p>Super interesting and helpful. Can't wait for the part on airway- and vessel segmentation!</p>",
      "rawMarkdown": "Super interesting and helpful. Can't wait for the part on airway- and vessel segmentation!"
    },
    {
      "id": 936228,
      "postDate": "2020-07-20T04:55:36.340Z",
      "content": "<p>Thank you for comment. It was very insightful.</p>",
      "rawMarkdown": "Thank you for comment. It was very insightful."
    },
    {
      "id": 936078,
      "postDate": "2020-07-20T00:37:14.577Z",
      "content": "<p>Like <a href=\"/simonlfwalsh\">@simonlfwalsh</a>  has pointed out \"Change in traction bronchiectasis severity is a measure of disease progression that could be used to help resolve the clinical importance of marginal FVC declines.\"<a href=\"https://thorax.bmj.com/content/75/8/648\">source</a>. I think its warranted to try to exploit this from the CT images provided in aiding the model to make accurate predictions.</p>",
      "rawMarkdown": "Like @simonlfwalsh  has pointed out \"Change in traction bronchiectasis severity is a measure of disease progression that could be used to help resolve the clinical importance of marginal FVC declines.\"[source](https://thorax.bmj.com/content/75/8/648). I think its warranted to try to exploit this from the CT images provided in aiding the model to make accurate predictions.",
      "replies": [
        {
          "id": 962864,
          "postDate": "2020-08-08T13:54:33.080Z",
          "content": "<p><a href=\"https://www.kaggle.com/drelias15\" target=\"_blank\">@drelias15</a>,  indeed,<br>\nthe <strong>change</strong> … but we have a single scan. There are patients among the least declining ones who have severe multilobar traction bronchiectasis and still do not decline further (even get better in the follow up period).</p>",
          "rawMarkdown": "@drelias15,  indeed,\nthe **change** ... but we have a single scan. There are patients among the least declining ones who have severe multilobar traction bronchiectasis and still do not decline further (even get better in the follow up period).\n"
        }
      ]
    },
    {
      "id": 934130,
      "postDate": "2020-07-18T08:51:06.530Z",
      "content": "<p>Very informative. Thank you very much for such great information. </p>",
      "rawMarkdown": "Very informative. Thank you very much for such great information. "
    },
    {
      "id": 931172,
      "postDate": "2020-07-16T03:27:43.080Z",
      "content": "<p>Amazing article describing ways to extract useful feature from images.</p>",
      "rawMarkdown": "Amazing article describing ways to extract useful feature from images."
    },
    {
      "id": 929868,
      "postDate": "2020-07-15T03:19:19.770Z",
      "content": "<p>Very Interesting, well done!</p>",
      "rawMarkdown": "Very Interesting, well done!",
      "replies": [
        {
          "id": 930098,
          "postDate": "2020-07-15T07:36:47.217Z",
          "content": "<p><a href=\"/brandonmcmanus\">@brandonmcmanus</a> thank you. If you liked this, read the second one to!</p>",
          "rawMarkdown": "@brandonmcmanus thank you. If you liked this, read the second one to!"
        }
      ]
    },
    {
      "id": 929814,
      "postDate": "2020-07-15T01:27:04.573Z",
      "content": "<p>Another newbie question:\nI have my masks. In order to calculate the <strong>volume of the lung</strong>: The <code>pixel number</code> mentioned is the count of pixels in my mask right, assuming the mask is correctly segmenting the lung. For instance, if a patient has a mask with the form (31, 512, 512) - 31 slices of width and height 512. The pixel number would be 31 * 512 * 512.</p>\n\n<p>Sorry for such elementary questions, it's just to ensure I don't waste time on silly mistakes.</p>\n\n<p>Thank you.</p>",
      "rawMarkdown": "Another newbie question:\nI have my masks. In order to calculate the **volume of the lung**: The `pixel number` mentioned is the count of pixels in my mask right, assuming the mask is correctly segmenting the lung. For instance, if a patient has a mask with the form (31, 512, 512) - 31 slices of width and height 512. The pixel number would be 31 * 512 * 512.\n\nSorry for such elementary questions, it's just to ensure I don't waste time on silly mistakes.\n\nThank you.",
      "replies": [
        {
          "id": 930125,
          "postDate": "2020-07-15T08:03:25.070Z",
          "content": "<p><a href=\"/ronaldokun\">@ronaldokun</a> ,</p>\n\n<p>31 * 512 * 512 gives you all the pixels (in this case voxels) in your entire volume! But you don't need that, you only need the volume of the lung itself.</p>\n\n<p>In the best case your lung label mask contains 2 labels (background and lung), you only have to count the pixels (voxels) labelled as lung.</p>\n\n<p><img src=\"http://drkonya.com/projects/kaggle/fibrosis_8.png\" alt=\"lung segmentation\"></p>\n\n<p>The first image is the unsegmented lung, the second is the mask and the third as overlay. You need the volume of the \"blue\" areas from the mask. The area between the white vertical lines is 512 x 512 (upscaled for this image but this only for presentation purposes) the entire mask. Labelled as lung is however only a portion of these pixels.</p>\n\n<p>So if your mask is in a numpy array with the pixels 0 when segmented as background and 1 if segmented as lung, the pixels you need from the image equals: numpy.count_nonzero(). ( you can perform morphological preprocessing or add a lowest count threshold to avoid counting noise though).</p>\n\n<p>Do not forget, these are the pixel values. In order to get the volumen in metric values, you have to multiply it with the pixelspacing^2 (this changes from scan to scan however usually 0.625 mm / pixel, meaning a 10 pixel wide and high area equals 39.0625 mm2) - see the comment above from <a href=\"/nickgm\">@nickgm</a> to - and then multiply with the distance between the slices (usually equals slice thickness) and it is already in mm. </p>\n\n<p>Then you end up with the volume of your segmented lung in mm3.</p>\n\n<p>Hope it helps.</p>",
          "rawMarkdown": "@ronaldokun ,\n\n31 * 512 * 512 gives you all the pixels (in this case voxels) in your entire volume! But you don't need that, you only need the volume of the lung itself.\n\nIn the best case your lung label mask contains 2 labels (background and lung), you only have to count the pixels (voxels) labelled as lung.\n\n![lung segmentation](http://drkonya.com/projects/kaggle/fibrosis_8.png)\n\nThe first image is the unsegmented lung, the second is the mask and the third as overlay. You need the volume of the \"blue\" areas from the mask. The area between the white vertical lines is 512 x 512 (upscaled for this image but this only for presentation purposes) the entire mask. Labelled as lung is however only a portion of these pixels.\n\n\nSo if your mask is in a numpy array with the pixels 0 when segmented as background and 1 if segmented as lung, the pixels you need from the image equals: numpy.count_nonzero(). ( you can perform morphological preprocessing or add a lowest count threshold to avoid counting noise though).\n\n\nDo not forget, these are the pixel values. In order to get the volumen in metric values, you have to multiply it with the pixelspacing^2 (this changes from scan to scan however usually 0.625 mm / pixel, meaning a 10 pixel wide and high area equals 39.0625 mm2) - see the comment above from @nickgm to - and then multiply with the distance between the slices (usually equals slice thickness) and it is already in mm. \n\nThen you end up with the volume of your segmented lung in mm3.\n\nHope it helps.\n",
          "votes": 4
        },
        {
          "id": 930399,
          "postDate": "2020-07-15T12:41:01.150Z",
          "content": "<p>Thank you a lot for the extensive explanation. Very much clear now.</p>",
          "rawMarkdown": "Thank you a lot for the extensive explanation. Very much clear now."
        },
        {
          "id": 992885,
          "postDate": "2020-08-31T13:58:40.030Z",
          "content": "<p>First of all, thank you so much for the informative explanation and insights, <a href=\"https://www.kaggle.com/sandorkonya\" target=\"_blank\">@sandorkonya</a> . I joined the competition quite late, so I'm still processing the information you posted. I have a question would like to ask if you don't mind. From you explanation, the volume of lung can be estimated from pixel count of lung * pixelspacing^2 * slice thickness. But since there are multiple ct scans from same patient, so the pixel count of lung should be different between scans, right? Does this mean we might get several lung volume information from same patient?</p>\n<p>Thanks again for the great posts.</p>",
          "rawMarkdown": "First of all, thank you so much for the informative explanation and insights, @sandorkonya . I joined the competition quite late, so I'm still processing the information you posted. I have a question would like to ask if you don't mind. From you explanation, the volume of lung can be estimated from pixel count of lung * pixelspacing^2 * slice thickness. But since there are multiple ct scans from same patient, so the pixel count of lung should be different between scans, right? Does this mean we might get several lung volume information from same patient?\n\nThanks again for the great posts."
        },
        {
          "id": 999680,
          "postDate": "2020-09-05T21:53:50.983Z",
          "content": "<p><a href=\"https://www.kaggle.com/xiejialun\" target=\"_blank\">@xiejialun</a> ,<br>\nEvery Patient has only one CT scan. There are multiple slices in one scan. <br>\nSum pixels on every slice.<br>\nNote, that the slice thickness is not in every case is the same as the difference between the slice localisations. </p>",
          "rawMarkdown": "@xiejialun ,\nEvery Patient has only one CT scan. There are multiple slices in one scan. \nSum pixels on every slice.\nNote, that the slice thickness is not in every case is the same as the difference between the slice localisations. ",
          "votes": 1
        },
        {
          "id": 1000912,
          "postDate": "2020-09-06T22:34:28.457Z",
          "content": "<p><a href=\"https://www.kaggle.com/sandorkonya\" target=\"_blank\">@sandorkonya</a> Thank you so much for the answer, I already manage to get lung volume information from my segmentation model and metadata in CT scans. Now I need to figure out how to use this feature well. Thanks again for all the contributions in this competition, good luck in private score :)</p>",
          "rawMarkdown": "@sandorkonya Thank you so much for the answer, I already manage to get lung volume information from my segmentation model and metadata in CT scans. Now I need to figure out how to use this feature well. Thanks again for all the contributions in this competition, good luck in private score :)"
        }
      ]
    },
    {
      "id": 929219,
      "postDate": "2020-07-14T14:38:38.520Z",
      "content": "<p>Just as a tip; honeycombing certainly is a predictor of outcome in these patients. However, traction bronchiectasis may be a much more sensitive predictor and so far, no one has successfully developed any automated approach to feature segmentation that can reliably quantify traction bronchiectasis. There is a fairly large body of literature based on visual assessments that suggests this and automated airway extraction specifically in fibrotic lung disease is considered an open research question.....</p>\n\n<p>[basic densitometry and related metrics has been investigated many times before, many years ago with reasonable albeit not earth-shattering results]</p>\n\n<p>Now get to work....!</p>",
      "rawMarkdown": "Just as a tip; honeycombing certainly is a predictor of outcome in these patients. However, traction bronchiectasis may be a much more sensitive predictor and so far, no one has successfully developed any automated approach to feature segmentation that can reliably quantify traction bronchiectasis. There is a fairly large body of literature based on visual assessments that suggests this and automated airway extraction specifically in fibrotic lung disease is considered an open research question.....\n\n[basic densitometry and related metrics has been investigated many times before, many years ago with reasonable albeit not earth-shattering results]\n\nNow get to work....!",
      "replies": [
        {
          "id": 929255,
          "postDate": "2020-07-14T15:01:59.033Z",
          "content": "<p>Hi <a href=\"/simonlfwalsh\">@simonlfwalsh</a> ,\ni totally agree with you, this is why i wrote in the last sentence:</p>\n\n<blockquote>\n  <p>Importance of Airway- and vessel segmentations, coming soon.</p>\n</blockquote>\n\n<p>The traction bronciectasis is the result of the volumen depletion and traction (hence the name) of the sorrounding, fibrotic lung parenchyma, and I strongly believe that reflects much more the mechanical changes within the lung as honeycombing or reticulation (these are just different phenotypes of the fibrosis). \nAnd as we all know, where no airflow, shunting happens, this is why the local and global vessel diameter (volume or density, whatever we call it) will be probably also a very good predictor ( google CLIPERS for this).\nWhether these visually detectable changes hide any prediction regarding the future progression... let's see! =)</p>",
          "rawMarkdown": "Hi @simonlfwalsh ,\ni totally agree with you, this is why i wrote in the last sentence:\n\n&gt; Importance of Airway- and vessel segmentations, coming soon.\n\nThe traction bronciectasis is the result of the volumen depletion and traction (hence the name) of the sorrounding, fibrotic lung parenchyma, and I strongly believe that reflects much more the mechanical changes within the lung as honeycombing or reticulation (these are just different phenotypes of the fibrosis). \nAnd as we all know, where no airflow, shunting happens, this is why the local and global vessel diameter (volume or density, whatever we call it) will be probably also a very good predictor ( google CLIPERS for this).\nWhether these visually detectable changes hide any prediction regarding the future progression... let's see! =)\n"
        },
        {
          "id": 929494,
          "postDate": "2020-07-14T17:51:39.833Z",
          "content": "<p>I think you mean CALIPER. Keep in mind that it is not clear what CALIPER is actually measuring. It could be vessel volume, it could be vessel volume and perivascular fibrosis, it could be something else. I encourage you to think outside the box for this challenge. That’s why we posted it!</p>\n\n<p>Other references worth looking at are DTA (a deep learning tool for fibrosis quantification - Steve Humphries) or some of our own work. </p>",
          "rawMarkdown": "I think you mean CALIPER. Keep in mind that it is not clear what CALIPER is actually measuring. It could be vessel volume, it could be vessel volume and perivascular fibrosis, it could be something else. I encourage you to think outside the box for this challenge. That’s why we posted it!\n\nOther references worth looking at are DTA (a deep learning tool for fibrosis quantification - Steve Humphries) or some of our own work. "
        },
        {
          "id": 929560,
          "postDate": "2020-07-14T18:44:47.500Z",
          "content": "<p><a href=\"/simonlfwalsh\">@simonlfwalsh</a>,</p>\n\n<p>yes, i meant CALIPER (shame on you predictive text on mobile phone).</p>\n\n<p>I already came across some of your work, for example <a href=\"https://pubmed.ncbi.nlm.nih.gov/27262146/\">this one</a>, there is an image with an overlay from the software in there. Well, after seeing that, i think i would vote for \"something else\"... there is quite much space left for improvement, for the trained eye pretty obvious non-vascular elements are labelled as \"vessel\" BUT it was a huge achievement in 2016.</p>\n\n<p>Thank you for the reference, i'll check it definitely!!</p>\n\n<p>As per out of the box thinking: if my first topic would be about \" we should take a closer look at the structural determinants of ventilation by finding any connection between supra-regional and/or regional airway heterogeneity and textural changes of the structural variability of the regional lung parenchyma correlated with the FCV\"... well, it would get no reads and a downvote =) But if it is the 5th in the row... someone would probably read it... one may implements it... who knows.</p>",
          "rawMarkdown": "@simonlfwalsh,\n\nyes, i meant CALIPER (shame on you predictive text on mobile phone).\n\nI already came across some of your work, for example [this one](https://pubmed.ncbi.nlm.nih.gov/27262146/), there is an image with an overlay from the software in there. Well, after seeing that, i think i would vote for \"something else\"... there is quite much space left for improvement, for the trained eye pretty obvious non-vascular elements are labelled as \"vessel\" BUT it was a huge achievement in 2016.\n\nThank you for the reference, i'll check it definitely!!\n\nAs per out of the box thinking: if my first topic would be about \" we should take a closer look at the structural determinants of ventilation by finding any connection between supra-regional and/or regional airway heterogeneity and textural changes of the structural variability of the regional lung parenchyma correlated with the FCV\"... well, it would get no reads and a downvote =) But if it is the 5th in the row... someone would probably read it... one may implements it... who knows."
        },
        {
          "id": 929678,
          "postDate": "2020-07-14T20:58:20.240Z",
          "content": "<p>Excellent; I look forward to seeing your solutions! </p>",
          "rawMarkdown": "Excellent; I look forward to seeing your solutions! ",
          "votes": 1
        },
        {
          "id": 1041655,
          "postDate": "2020-10-07T21:29:44.973Z",
          "content": "<p><a href=\"https://www.kaggle.com/simonlfwalsh\" target=\"_blank\">@simonlfwalsh</a> ,<br>\nunexpected results. Can we expect a second round? :)</p>",
          "rawMarkdown": "@simonlfwalsh ,\nunexpected results. Can we expect a second round? :)"
        }
      ]
    },
    {
      "id": 928396,
      "postDate": "2020-07-14T00:32:31.697Z",
      "content": "<p>Very informative and explained with clarity ... will definitely be implementing these insights into my work, thank you! 🙂 </p>",
      "rawMarkdown": "Very informative and explained with clarity ... will definitely be implementing these insights into my work, thank you! 🙂 ",
      "replies": [
        {
          "id": 929258,
          "postDate": "2020-07-14T15:03:06.153Z",
          "content": "<p><a href=\"/sdlee94\">@sdlee94</a> , thank you, let me know if any of them works! =)</p>",
          "rawMarkdown": "@sdlee94 , thank you, let me know if any of them works! =)"
        }
      ]
    },
    {
      "id": 926450,
      "postDate": "2020-07-12T18:22:42.290Z",
      "content": "<p>Domain knowledge is key. Thank you for your insights!</p>",
      "rawMarkdown": "Domain knowledge is key. Thank you for your insights!",
      "replies": [
        {
          "id": 926473,
          "postDate": "2020-07-12T18:42:35.420Z",
          "content": "<p><a href=\"/correax\">@correax</a> , thx,\nDomain knowledge... or brute force =)</p>",
          "rawMarkdown": "@correax , thx,\nDomain knowledge... or brute force =)"
        },
        {
          "id": 926697,
          "postDate": "2020-07-12T22:30:49.763Z",
          "content": "<p>That works too! :P</p>",
          "rawMarkdown": "That works too! :P"
        }
      ]
    },
    {
      "id": 926224,
      "postDate": "2020-07-12T15:29:44.937Z",
      "content": "<p>I really appreciate your expert insight, it has given me so many ideas for making use of the image data. Thanks!</p>",
      "rawMarkdown": "I really appreciate your expert insight, it has given me so many ideas for making use of the image data. Thanks!",
      "replies": [
        {
          "id": 926478,
          "postDate": "2020-07-12T18:46:13.653Z",
          "content": "<p><a href=\"/andypenrose\">@andypenrose</a> , this means i achieved my goal... </p>\n\n<p>&gt; It is like the seed put in the soil - the more one sows, the greater the harvest.\n<em>Orison Swett Marden</em></p>",
          "rawMarkdown": "@andypenrose , this means i achieved my goal... \n\n&gt; It is like the seed put in the soil - the more one sows, the greater the harvest.\n*Orison Swett Marden*",
          "votes": 1
        }
      ]
    },
    {
      "id": 925415,
      "postDate": "2020-07-12T04:25:21.040Z",
      "content": "<p>I just read both of your insights. It is really helpful!</p>\n\n<p>I can't seem to comprehend the meaning of this statement:</p>\n\n<blockquote>\n  <p>define the topmost point the lung and beginning from that point divide your whole lung volume into smaller, for example 1 x 1 x 1 cm cubicles (blocks) and analyse these - along with their distance as value to the topmost point - so not to lose localisation information. </p>\n</blockquote>\n\n<p>Why is block segmentation needed for the mask? Did I miss something you wrote?</p>",
      "rawMarkdown": "I just read both of your insights. It is really helpful!\n\nI can't seem to comprehend the meaning of this statement:\n\n&gt; define the topmost point the lung and beginning from that point divide your whole lung volume into smaller, for example 1 x 1 x 1 cm cubicles (blocks) and analyse these - along with their distance as value to the topmost point - so not to lose localisation information. \n\nWhy is block segmentation needed for the mask? Did I miss something you wrote?",
      "replies": [
        {
          "id": 925736,
          "postDate": "2020-07-12T09:19:59.733Z",
          "content": "<p>Thank you <a href=\"/aadhavvignesh\">@aadhavvignesh</a> , </p>\n\n<p>i should clarify it further, you're right:</p>\n\n<p>if we assume that similar tissue changes of the lung on different position in the lung itself contribute differently to the change in the spirometry results, then we need to know some information about the position/localisation of these changes. \nOne slice however contains multiple areas of different tissue changes == multiple different textures (normal tissue, altered tissue [like: honeycombing, septal thickening, emphysema, ground glass...]), so to divide these areas on one slice further down in regions is advisable. </p>\n\n<p>This further categorisation can be made with the help of trained NN (see the part on MedGift database) but also without NN:\n The idea of creating blocks comes initially <a href=\"https://arxiv.org/pdf/1811.02651.pdf\">from here</a>, and we could use this technique to analyse the blocks and get information about the texture of different regions.</p>\n\n<p>Additionaly we can obtain an additional position/localisation value of each block.</p>\n\n<p>In order to get a localisation, we need a good definable point, and i chose for this the topmost point of the lung (called apex) since it is very simple to obtain: take the first slice that contains mask of your lung, and take the midpoint of this area.  The distance between this point and the midpoint of each block is then the distance i visioned. Other points could be set to, like the division of the trachea (called carina) but obtaining that point on the CT requires more segmentation work.</p>\n\n<p>Since every people is different, we have to normalize somehow this distance, one possibility is to take the \"height\" of the whole lung and normalize the distance with this value. The height of the lung is simply the difference of the position of the first slice containing lung mask and position of the last slice containing lung mask.</p>\n\n<p>I hope i made my thoughts clear.</p>",
          "rawMarkdown": "Thank you @aadhavvignesh , \n\ni should clarify it further, you're right:\n\nif we assume that similar tissue changes of the lung on different position in the lung itself contribute differently to the change in the spirometry results, then we need to know some information about the position/localisation of these changes. \nOne slice however contains multiple areas of different tissue changes == multiple different textures (normal tissue, altered tissue [like: honeycombing, septal thickening, emphysema, ground glass...]), so to divide these areas on one slice further down in regions is advisable. \n\nThis further categorisation can be made with the help of trained NN (see the part on MedGift database) but also without NN:\n The idea of creating blocks comes initially [from here](https://arxiv.org/pdf/1811.02651.pdf), and we could use this technique to analyse the blocks and get information about the texture of different regions.\n\nAdditionaly we can obtain an additional position/localisation value of each block.\n\nIn order to get a localisation, we need a good definable point, and i chose for this the topmost point of the lung (called apex) since it is very simple to obtain: take the first slice that contains mask of your lung, and take the midpoint of this area.  The distance between this point and the midpoint of each block is then the distance i visioned. Other points could be set to, like the division of the trachea (called carina) but obtaining that point on the CT requires more segmentation work.\n\nSince every people is different, we have to normalize somehow this distance, one possibility is to take the \"height\" of the whole lung and normalize the distance with this value. The height of the lung is simply the difference of the position of the first slice containing lung mask and position of the last slice containing lung mask.\n\nI hope i made my thoughts clear.",
          "votes": 3
        },
        {
          "id": 925741,
          "postDate": "2020-07-12T09:25:33.393Z",
          "content": "<p>Thanks for the clarification! This seems a bit overwhelming for me, but I'll try to come up with something :)</p>\n\n<p>Thank you for such an interesting read!</p>",
          "rawMarkdown": "Thanks for the clarification! This seems a bit overwhelming for me, but I'll try to come up with something :)\n\nThank you for such an interesting read!"
        },
        {
          "id": 925955,
          "postDate": "2020-07-12T11:52:52.810Z",
          "content": "<p>Thank you for sharing such valuable information! It surely helps a lot!</p>",
          "rawMarkdown": "Thank you for sharing such valuable information! It surely helps a lot!"
        },
        {
          "id": 926479,
          "postDate": "2020-07-12T18:47:07.587Z",
          "content": "<p><a href=\"/shikha130vv\">@shikha130vv</a> ,\nYou're welcome.</p>",
          "rawMarkdown": "@shikha130vv ,\nYou're welcome.",
          "votes": 1
        }
      ]
    },
    {
      "id": 924310,
      "postDate": "2020-07-11T11:16:31.393Z",
      "content": "<p>cool</p>",
      "rawMarkdown": "cool",
      "replies": [
        {
          "id": 924359,
          "postDate": "2020-07-11T11:53:03.273Z",
          "content": "<p>thank you.</p>",
          "rawMarkdown": "thank you."
        }
      ]
    },
    {
      "id": 928898,
      "postDate": "2020-07-14T10:07:48.370Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 929257,
          "postDate": "2020-07-14T15:02:37.937Z",
          "content": "<p><a href=\"/adityavineeth\">@adityavineeth</a>  thank you!</p>",
          "rawMarkdown": "@adityavineeth  thank you!"
        }
      ]
    },
    {
      "id": 1022260,
      "postDate": "2020-09-22T12:14:47.443Z",
      "content": "<p>Thank you. It is great!</p>",
      "rawMarkdown": "Thank you. It is great!"
    },
    {
      "id": 933912,
      "postDate": "2020-07-18T05:48:09.783Z",
      "content": "<p>thanks for the insight!</p>",
      "rawMarkdown": "thanks for the insight!"
    },
    {
      "id": 931728,
      "postDate": "2020-07-16T12:10:13.680Z",
      "content": "<p>Thank you dr. Konya! it's Interesting.</p>",
      "rawMarkdown": "Thank you dr. Konya! it's Interesting."
    }
  ],
  "comments": [
    {
      "id": 1041652,
      "author_name": "dr. Konya",
      "author_url": "",
      "post_date": "2020-10-07T21:28:35.170000",
      "content": "<p>UPDATE</p>\n<p>Dear fellow kagglers, </p>\n<p>i posted <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/189528\" target=\"_blank\">segmentation masks</a> after the competition.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 925161,
      "author_name": "Ronaldo S.A. Batista",
      "author_url": "",
      "post_date": "2020-07-11T21:15:53.363000",
      "content": "<p>Very insightful. Thank you!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 925162,
          "author_name": "dr. Konya",
          "author_url": "",
          "post_date": "2020-07-11T21:17:07.883000",
          "content": "<p>Thank you, be sure you read my <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/166123\">second one</a> to!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 925165,
          "author_name": "Ronaldo S.A. Batista",
          "author_url": "",
          "post_date": "2020-07-11T21:19:03.893000",
          "content": "<p>I'm on it right now👊 </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 924666,
      "author_name": "NickGraaf",
      "author_url": "",
      "post_date": "2020-07-11T15:10:46.033000",
      "content": "<p>Very informative. Seems like the first step is the segmentation, a task that still seems to have a precision not quite high enough as what would be wished in the field</p>",
      "votes": 1,
      "replies": [
        {
          "id": 924741,
          "author_name": "dr. Konya",
          "author_url": "",
          "post_date": "2020-07-11T15:55:11.700000",
          "content": "<p>We suppose (or better presume) that the lung parenchyma is more responsible for the changes in the pulmonary function tests than other parts of the chest (muscles, ribs, fat), so the segmentation is the first  dimensionality reduction. \nThere are more information (about BMI, motion constraints of the chest's bony cage due to degeneration of the ribs, shape of diaphragm, calcifications as sign for cardiovascular comorbidity) that are then not taken into account, but they should be addressed.</p>\n\n<p>So while segmentatiin of the lungs is one way, it helps by crafting manual features that lead to further dimensionality reduction... and maybe we can derive clinically relevant (and correlating) variables but not the only way one can try.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 927150,
          "author_name": "NickGraaf",
          "author_url": "",
          "post_date": "2020-07-13T08:22:35.360000",
          "content": "<p>Thank you! Just a question: when calculating the volume, shouldn't pixel_spacing be pixel_spacing^2 (assuming that row spacing equals column spacing)?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 929242,
          "author_name": "dr. Konya",
          "author_url": "",
          "post_date": "2020-07-14T14:51:48.150000",
          "content": "<p><a href=\"/nickgm\">@nickgm</a> ,\nyes, you're obviously right!\nThe CT slices have usually equal row &amp; col spacing.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 958201,
      "author_name": "Stefano Pelliccioli",
      "author_url": "",
      "post_date": "2020-08-04T20:50:12.490000",
      "content": "<p>Thanks for your amazing insights. I implemented some of them in my <a href=\"https://www.kaggle.com/unforgiven/osic-comprehensive-eda\">EDA notebook</a>, if you want to take a look.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 962858,
          "author_name": "dr. Konya",
          "author_url": "",
          "post_date": "2020-08-08T13:50:49.563000",
          "content": "<p><a href=\"https://www.kaggle.com/unforgiven\" target=\"_blank\">@unforgiven</a>,<br>\ni took a look and it is 👍👍, also left a comment.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 925141,
      "author_name": "dr. Konya",
      "author_url": "",
      "post_date": "2020-07-11T20:57:12.617000",
      "content": "<h3>UPDATE see the second part of my insights here:</h3>\n\n<p><a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/166123\">https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/166123</a></p>",
      "votes": 2,
      "replies": [
        {
          "id": 960593,
          "author_name": "Nikša Praljak",
          "author_url": "",
          "post_date": "2020-08-06T14:24:23.667000",
          "content": "<p>Thank you for the insightful thread!</p>\n\n<p>Did you post a thread discussing airway and vessel segmentation?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 962842,
          "author_name": "dr. Konya",
          "author_url": "",
          "post_date": "2020-08-08T13:32:51.953000",
          "content": "<p><a href=\"https://www.kaggle.com/niksapraljak\" target=\"_blank\">@niksapraljak</a> ,<br>\nwe discussed it in the team and decided not to do it as for now. <br>\nNearing the end of the competition however i plan to make an in-depth thread about it.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1028702,
      "author_name": "VSR",
      "author_url": "",
      "post_date": "2020-09-27T05:57:36.597000",
      "content": "<p>Thank you for your informative post - Really helps for someone new like me.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1015532,
      "author_name": "Kuruton",
      "author_url": "",
      "post_date": "2020-09-18T08:44:50.847000",
      "content": "<p>Thanks for very useful insights!<br>\nHowever, I have two questions for you.</p>\n<p>1.<br>\n    You wrote \"It (the volume of the lung) can be calculated as: pixel number on your lung segmentation map * pixel spacing (from DICOM header information, usually 0.625 mm/px ) * slice distance\".<br>\n    But, I think it is calculated as: pixel number on your lung segmentation map * pixel spacing (from DICOM header information, usually 0.625 mm/px )** 2 * slice distance\". (pixel spacing is multiplyed twice)<br>\n    Is this wrong?</p>\n<p>2.<br>\n    You wrote \"Slice distance ( ~= slice thickness) is derived from the difference between 2 consecutive slice's Image Position DICOM attribute - usually ranging from 1 mm to 7 mm. \"<br>\n    How to get position of DICOM?</p>\n<p>I would be grateful if you could answer to these questions.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1016358,
          "author_name": "dr. Konya",
          "author_url": "",
          "post_date": "2020-09-18T21:57:43.953000",
          "content": "<p><a href=\"https://www.kaggle.com/koichirokamada\" target=\"_blank\">@koichirokamada</a> ,<br>\nyou are right pixel spacing is squared, check the commen below for <a href=\"https://www.kaggle.com/ronaldokun\" target=\"_blank\">@ronaldokun</a>.<br>\nSlice location comes from this DICOM tag:<br>\n<a href=\"https://dicom.innolitics.com/ciods/ct-image/image-plane/00201041\" target=\"_blank\">https://dicom.innolitics.com/ciods/ct-image/image-plane/00201041</a><br>\nYou are interested in the third, [2], z-position. <br>\nGood luck!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1016501,
          "author_name": "Kuruton",
          "author_url": "",
          "post_date": "2020-09-19T03:14:46.523000",
          "content": "<p>I see.<br>\nThank you very much!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1019883,
          "author_name": "Vasileios Kagklis",
          "author_url": "",
          "post_date": "2020-09-20T18:20:17.653000",
          "content": "<p><a href=\"https://www.kaggle.com/sandorkonya\" target=\"_blank\">@sandorkonya</a> So if I understood correctly, then if we \"resample\" the scan so that it has a new spacing of  1x1x1mm, then calculating the volume would simply require to sum the number of pixels of the segmented lungs, correct?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1000721,
      "author_name": "Carlos Souza",
      "author_url": "",
      "post_date": "2020-09-06T18:31:14.240000",
      "content": "<p>Thanks again for the insights, <a href=\"https://www.kaggle.com/sandorkonya\" target=\"_blank\">@sandorkonya</a> !</p>\n<p>I'd love to pick your brain on ideas to move forward. So far, the most promising approach IMHO is derived from your ideas. </p>\n<p>Let's assume FVC declines are lines, from your first image, which alphas (intercepts) and betas (slope, i.e. the severity of the decline). Thus, the objective of the competition could be restated as predicting the beta for each patient from the baseline measurements &amp; CT scan. If we have alphas &amp; betas well predicted for each patient, we can infer FVCs for every possible week.</p>\n<p>Using a Bayesian approach, we can learn that the 176 test patients have a distribution of betas normally distributed with center in -4.2 and std dev of 5 (all approx). This very simple to obtain. Then, if I sample a beta from this distribution for each test patient, and make their lines pass through their baseline FVC, and calculate the sigmas properly (I demonstrated in my last notebook), this yields a score of -6.90. That's our minimum score to beat. Anything worse than that is pure nonsense.</p>\n<p>Now, enter the CT scans. As you correctly pointed out, the job is to predict the severity of decline in FVC (betas) using features derived from the CT scans. So far, I implemented several features you suggested (volume in L, chest circumference in mm, height, HU histogram moments (mean, stddev, skew, kurtosis), % of lung volume in different ranges of HU, etc). I also tried deriving latent features learned with unsupervised learning (using a Variational Autoencoder, not public yet). </p>\n<p>The problem is that this dataset with 176 CT scans with 176 betas is too small. There's too much noise: variance is very high, such that the gains for using all this noisy data are marginal in comparison to assuming all test patients have betas sampled from the same global distribution (i.e. all have the same severity). Even if I transform the problem into a classification problem (patients with betas lower than distribution median are \"high severity\", and higher than median are \"low severity\"), best I get is ~60-70% accuracy, loo low.</p>\n<p>You mentioned the blocks idea. I'd love to use the MedGift database to classify every voxel of our patients, but one of the organizers said it is off-limits because of their license agreement.</p>\n<p>Anyway, I'm getting out of ideas on how to improve…</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1000861,
          "author_name": "dr. Konya",
          "author_url": "",
          "post_date": "2020-09-06T21:00:16.767000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/carlossouza\" target=\"_blank\">@carlossouza</a> ,</p>\n<p>you're right, we were also struggling with the high heterogenity of the dataset. It is no secret that the correct cropping &amp; resampling of the dataset took up to 90% our time.<br>\nWe fancied the idea of tackling this as a classification problem to, but no image derived feature seems to be strong enough (yet) to be a good classifier.</p>\n<p>I managed to manually segment 100 lungs (this enables us to get almost perfect segmentation every time) , moreover I made &gt;30k markups for reticulation, honeycombing, airways, normal tissue… we are integrating these features into our existing pipeline, we will see how it performs.</p>\n<p><img src=\"http://www.drkonya.com/projects/kaggle/fibrosis_9.JPG\" alt=\"\"></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 977444,
      "author_name": "Frederik Laubisch",
      "author_url": "",
      "post_date": "2020-08-19T13:33:48.437000",
      "content": "<p>Super interesting and helpful. Can't wait for the part on airway- and vessel segmentation!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 936228,
      "author_name": "Sagyam Thapa",
      "author_url": "",
      "post_date": "2020-07-20T04:55:36.340000",
      "content": "<p>Thank you for comment. It was very insightful.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 936078,
      "author_name": "zedreng",
      "author_url": "",
      "post_date": "2020-07-20T00:37:14.577000",
      "content": "<p>Like <a href=\"/simonlfwalsh\">@simonlfwalsh</a>  has pointed out \"Change in traction bronchiectasis severity is a measure of disease progression that could be used to help resolve the clinical importance of marginal FVC declines.\"<a href=\"https://thorax.bmj.com/content/75/8/648\">source</a>. I think its warranted to try to exploit this from the CT images provided in aiding the model to make accurate predictions.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 962864,
          "author_name": "dr. Konya",
          "author_url": "",
          "post_date": "2020-08-08T13:54:33.080000",
          "content": "<p><a href=\"https://www.kaggle.com/drelias15\" target=\"_blank\">@drelias15</a>,  indeed,<br>\nthe <strong>change</strong> … but we have a single scan. There are patients among the least declining ones who have severe multilobar traction bronchiectasis and still do not decline further (even get better in the follow up period).</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 934130,
      "author_name": "Kuldeep Singh Chouhan",
      "author_url": "",
      "post_date": "2020-07-18T08:51:06.530000",
      "content": "<p>Very informative. Thank you very much for such great information. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 931172,
      "author_name": "Vivek Gopal Ramaswamy",
      "author_url": "",
      "post_date": "2020-07-16T03:27:43.080000",
      "content": "<p>Amazing article describing ways to extract useful feature from images.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 929868,
      "author_name": "Brandon McManus",
      "author_url": "",
      "post_date": "2020-07-15T03:19:19.770000",
      "content": "<p>Very Interesting, well done!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 930098,
          "author_name": "dr. Konya",
          "author_url": "",
          "post_date": "2020-07-15T07:36:47.217000",
          "content": "<p><a href=\"/brandonmcmanus\">@brandonmcmanus</a> thank you. If you liked this, read the second one to!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 929814,
      "author_name": "Ronaldo S.A. Batista",
      "author_url": "",
      "post_date": "2020-07-15T01:27:04.573000",
      "content": "<p>Another newbie question:\nI have my masks. In order to calculate the <strong>volume of the lung</strong>: The <code>pixel number</code> mentioned is the count of pixels in my mask right, assuming the mask is correctly segmenting the lung. For instance, if a patient has a mask with the form (31, 512, 512) - 31 slices of width and height 512. The pixel number would be 31 * 512 * 512.</p>\n\n<p>Sorry for such elementary questions, it's just to ensure I don't waste time on silly mistakes.</p>\n\n<p>Thank you.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 930125,
          "author_name": "dr. Konya",
          "author_url": "",
          "post_date": "2020-07-15T08:03:25.070000",
          "content": "<p><a href=\"/ronaldokun\">@ronaldokun</a> ,</p>\n\n<p>31 * 512 * 512 gives you all the pixels (in this case voxels) in your entire volume! But you don't need that, you only need the volume of the lung itself.</p>\n\n<p>In the best case your lung label mask contains 2 labels (background and lung), you only have to count the pixels (voxels) labelled as lung.</p>\n\n<p><img src=\"http://drkonya.com/projects/kaggle/fibrosis_8.png\" alt=\"lung segmentation\"></p>\n\n<p>The first image is the unsegmented lung, the second is the mask and the third as overlay. You need the volume of the \"blue\" areas from the mask. The area between the white vertical lines is 512 x 512 (upscaled for this image but this only for presentation purposes) the entire mask. Labelled as lung is however only a portion of these pixels.</p>\n\n<p>So if your mask is in a numpy array with the pixels 0 when segmented as background and 1 if segmented as lung, the pixels you need from the image equals: numpy.count_nonzero(). ( you can perform morphological preprocessing or add a lowest count threshold to avoid counting noise though).</p>\n\n<p>Do not forget, these are the pixel values. In order to get the volumen in metric values, you have to multiply it with the pixelspacing^2 (this changes from scan to scan however usually 0.625 mm / pixel, meaning a 10 pixel wide and high area equals 39.0625 mm2) - see the comment above from <a href=\"/nickgm\">@nickgm</a> to - and then multiply with the distance between the slices (usually equals slice thickness) and it is already in mm. </p>\n\n<p>Then you end up with the volume of your segmented lung in mm3.</p>\n\n<p>Hope it helps.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 930399,
          "author_name": "Ronaldo S.A. Batista",
          "author_url": "",
          "post_date": "2020-07-15T12:41:01.150000",
          "content": "<p>Thank you a lot for the extensive explanation. Very much clear now.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 992885,
          "author_name": "Tsai29",
          "author_url": "",
          "post_date": "2020-08-31T13:58:40.030000",
          "content": "<p>First of all, thank you so much for the informative explanation and insights, <a href=\"https://www.kaggle.com/sandorkonya\" target=\"_blank\">@sandorkonya</a> . I joined the competition quite late, so I'm still processing the information you posted. I have a question would like to ask if you don't mind. From you explanation, the volume of lung can be estimated from pixel count of lung * pixelspacing^2 * slice thickness. But since there are multiple ct scans from same patient, so the pixel count of lung should be different between scans, right? Does this mean we might get several lung volume information from same patient?</p>\n<p>Thanks again for the great posts.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 999680,
          "author_name": "dr. Konya",
          "author_url": "",
          "post_date": "2020-09-05T21:53:50.983000",
          "content": "<p><a href=\"https://www.kaggle.com/xiejialun\" target=\"_blank\">@xiejialun</a> ,<br>\nEvery Patient has only one CT scan. There are multiple slices in one scan. <br>\nSum pixels on every slice.<br>\nNote, that the slice thickness is not in every case is the same as the difference between the slice localisations. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1000912,
          "author_name": "Tsai29",
          "author_url": "",
          "post_date": "2020-09-06T22:34:28.457000",
          "content": "<p><a href=\"https://www.kaggle.com/sandorkonya\" target=\"_blank\">@sandorkonya</a> Thank you so much for the answer, I already manage to get lung volume information from my segmentation model and metadata in CT scans. Now I need to figure out how to use this feature well. Thanks again for all the contributions in this competition, good luck in private score :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 929219,
      "author_name": "SimonWalsh",
      "author_url": "",
      "post_date": "2020-07-14T14:38:38.520000",
      "content": "<p>Just as a tip; honeycombing certainly is a predictor of outcome in these patients. However, traction bronchiectasis may be a much more sensitive predictor and so far, no one has successfully developed any automated approach to feature segmentation that can reliably quantify traction bronchiectasis. There is a fairly large body of literature based on visual assessments that suggests this and automated airway extraction specifically in fibrotic lung disease is considered an open research question.....</p>\n\n<p>[basic densitometry and related metrics has been investigated many times before, many years ago with reasonable albeit not earth-shattering results]</p>\n\n<p>Now get to work....!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 929255,
          "author_name": "dr. Konya",
          "author_url": "",
          "post_date": "2020-07-14T15:01:59.033000",
          "content": "<p>Hi <a href=\"/simonlfwalsh\">@simonlfwalsh</a> ,\ni totally agree with you, this is why i wrote in the last sentence:</p>\n\n<blockquote>\n  <p>Importance of Airway- and vessel segmentations, coming soon.</p>\n</blockquote>\n\n<p>The traction bronciectasis is the result of the volumen depletion and traction (hence the name) of the sorrounding, fibrotic lung parenchyma, and I strongly believe that reflects much more the mechanical changes within the lung as honeycombing or reticulation (these are just different phenotypes of the fibrosis). \nAnd as we all know, where no airflow, shunting happens, this is why the local and global vessel diameter (volume or density, whatever we call it) will be probably also a very good predictor ( google CLIPERS for this).\nWhether these visually detectable changes hide any prediction regarding the future progression... let's see! =)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 929494,
          "author_name": "SimonWalsh",
          "author_url": "",
          "post_date": "2020-07-14T17:51:39.833000",
          "content": "<p>I think you mean CALIPER. Keep in mind that it is not clear what CALIPER is actually measuring. It could be vessel volume, it could be vessel volume and perivascular fibrosis, it could be something else. I encourage you to think outside the box for this challenge. That’s why we posted it!</p>\n\n<p>Other references worth looking at are DTA (a deep learning tool for fibrosis quantification - Steve Humphries) or some of our own work. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 929560,
          "author_name": "dr. Konya",
          "author_url": "",
          "post_date": "2020-07-14T18:44:47.500000",
          "content": "<p><a href=\"/simonlfwalsh\">@simonlfwalsh</a>,</p>\n\n<p>yes, i meant CALIPER (shame on you predictive text on mobile phone).</p>\n\n<p>I already came across some of your work, for example <a href=\"https://pubmed.ncbi.nlm.nih.gov/27262146/\">this one</a>, there is an image with an overlay from the software in there. Well, after seeing that, i think i would vote for \"something else\"... there is quite much space left for improvement, for the trained eye pretty obvious non-vascular elements are labelled as \"vessel\" BUT it was a huge achievement in 2016.</p>\n\n<p>Thank you for the reference, i'll check it definitely!!</p>\n\n<p>As per out of the box thinking: if my first topic would be about \" we should take a closer look at the structural determinants of ventilation by finding any connection between supra-regional and/or regional airway heterogeneity and textural changes of the structural variability of the regional lung parenchyma correlated with the FCV\"... well, it would get no reads and a downvote =) But if it is the 5th in the row... someone would probably read it... one may implements it... who knows.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 929678,
          "author_name": "SimonWalsh",
          "author_url": "",
          "post_date": "2020-07-14T20:58:20.240000",
          "content": "<p>Excellent; I look forward to seeing your solutions! </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1041655,
          "author_name": "dr. Konya",
          "author_url": "",
          "post_date": "2020-10-07T21:29:44.973000",
          "content": "<p><a href=\"https://www.kaggle.com/simonlfwalsh\" target=\"_blank\">@simonlfwalsh</a> ,<br>\nunexpected results. Can we expect a second round? :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 928396,
      "author_name": "Stephen Lee",
      "author_url": "",
      "post_date": "2020-07-14T00:32:31.697000",
      "content": "<p>Very informative and explained with clarity ... will definitely be implementing these insights into my work, thank you! 🙂 </p>",
      "votes": 0,
      "replies": [
        {
          "id": 929258,
          "author_name": "dr. Konya",
          "author_url": "",
          "post_date": "2020-07-14T15:03:06.153000",
          "content": "<p><a href=\"/sdlee94\">@sdlee94</a> , thank you, let me know if any of them works! =)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 926450,
      "author_name": "Fabio Correa",
      "author_url": "",
      "post_date": "2020-07-12T18:22:42.290000",
      "content": "<p>Domain knowledge is key. Thank you for your insights!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 926473,
          "author_name": "dr. Konya",
          "author_url": "",
          "post_date": "2020-07-12T18:42:35.420000",
          "content": "<p><a href=\"/correax\">@correax</a> , thx,\nDomain knowledge... or brute force =)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 926697,
          "author_name": "Fabio Correa",
          "author_url": "",
          "post_date": "2020-07-12T22:30:49.763000",
          "content": "<p>That works too! :P</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 926224,
      "author_name": "Andy Penrose",
      "author_url": "",
      "post_date": "2020-07-12T15:29:44.937000",
      "content": "<p>I really appreciate your expert insight, it has given me so many ideas for making use of the image data. Thanks!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 926478,
          "author_name": "dr. Konya",
          "author_url": "",
          "post_date": "2020-07-12T18:46:13.653000",
          "content": "<p><a href=\"/andypenrose\">@andypenrose</a> , this means i achieved my goal... </p>\n\n<p>&gt; It is like the seed put in the soil - the more one sows, the greater the harvest.\n<em>Orison Swett Marden</em></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 925415,
      "author_name": "Aadhav Vignesh",
      "author_url": "",
      "post_date": "2020-07-12T04:25:21.040000",
      "content": "<p>I just read both of your insights. It is really helpful!</p>\n\n<p>I can't seem to comprehend the meaning of this statement:</p>\n\n<blockquote>\n  <p>define the topmost point the lung and beginning from that point divide your whole lung volume into smaller, for example 1 x 1 x 1 cm cubicles (blocks) and analyse these - along with their distance as value to the topmost point - so not to lose localisation information. </p>\n</blockquote>\n\n<p>Why is block segmentation needed for the mask? Did I miss something you wrote?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 925736,
          "author_name": "dr. Konya",
          "author_url": "",
          "post_date": "2020-07-12T09:19:59.733000",
          "content": "<p>Thank you <a href=\"/aadhavvignesh\">@aadhavvignesh</a> , </p>\n\n<p>i should clarify it further, you're right:</p>\n\n<p>if we assume that similar tissue changes of the lung on different position in the lung itself contribute differently to the change in the spirometry results, then we need to know some information about the position/localisation of these changes. \nOne slice however contains multiple areas of different tissue changes == multiple different textures (normal tissue, altered tissue [like: honeycombing, septal thickening, emphysema, ground glass...]), so to divide these areas on one slice further down in regions is advisable. </p>\n\n<p>This further categorisation can be made with the help of trained NN (see the part on MedGift database) but also without NN:\n The idea of creating blocks comes initially <a href=\"https://arxiv.org/pdf/1811.02651.pdf\">from here</a>, and we could use this technique to analyse the blocks and get information about the texture of different regions.</p>\n\n<p>Additionaly we can obtain an additional position/localisation value of each block.</p>\n\n<p>In order to get a localisation, we need a good definable point, and i chose for this the topmost point of the lung (called apex) since it is very simple to obtain: take the first slice that contains mask of your lung, and take the midpoint of this area.  The distance between this point and the midpoint of each block is then the distance i visioned. Other points could be set to, like the division of the trachea (called carina) but obtaining that point on the CT requires more segmentation work.</p>\n\n<p>Since every people is different, we have to normalize somehow this distance, one possibility is to take the \"height\" of the whole lung and normalize the distance with this value. The height of the lung is simply the difference of the position of the first slice containing lung mask and position of the last slice containing lung mask.</p>\n\n<p>I hope i made my thoughts clear.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 925741,
          "author_name": "Aadhav Vignesh",
          "author_url": "",
          "post_date": "2020-07-12T09:25:33.393000",
          "content": "<p>Thanks for the clarification! This seems a bit overwhelming for me, but I'll try to come up with something :)</p>\n\n<p>Thank you for such an interesting read!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 925955,
          "author_name": "Shikha Agrawal",
          "author_url": "",
          "post_date": "2020-07-12T11:52:52.810000",
          "content": "<p>Thank you for sharing such valuable information! It surely helps a lot!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 926479,
          "author_name": "dr. Konya",
          "author_url": "",
          "post_date": "2020-07-12T18:47:07.587000",
          "content": "<p><a href=\"/shikha130vv\">@shikha130vv</a> ,\nYou're welcome.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 924310,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-11T11:16:31.393000",
      "content": "<p>cool</p>",
      "votes": 0,
      "replies": [
        {
          "id": 924359,
          "author_name": "dr. Konya",
          "author_url": "",
          "post_date": "2020-07-11T11:53:03.273000",
          "content": "<p>thank you.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 928898,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-14T10:07:48.370000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 929257,
          "author_name": "dr. Konya",
          "author_url": "",
          "post_date": "2020-07-14T15:02:37.937000",
          "content": "<p><a href=\"/adityavineeth\">@adityavineeth</a>  thank you!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1022260,
      "author_name": "monday",
      "author_url": "",
      "post_date": "2020-09-22T12:14:47.443000",
      "content": "<p>Thank you. It is great!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 933912,
      "author_name": "khailash santhakumar",
      "author_url": "",
      "post_date": "2020-07-18T05:48:09.783000",
      "content": "<p>thanks for the insight!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 931728,
      "author_name": "bayesian monk",
      "author_url": "",
      "post_date": "2020-07-16T12:10:13.680000",
      "content": "<p>Thank you dr. Konya! it's Interesting.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "923472": "\n\n\n### Disclaimer\n\n\tAlthough the described insights are derived mainly from empirical data, there are some brain-storm parts that can bring you deep in the rabbit hole and make you work for nothing... or may lead to groundbreaking results.\n\n### Structure\n\n\t1 - Goals of the present competition\n\t2 - Factors known to determine the outcome of the Idiopathic Pulmonary Fibrosis (IPF)\n\t3 - Image based features\n\n\t###Goals of the present competition\n\nLet’s take some **epidemiological details** (gender, age, smoking habits), some measured parameters (Forced Vital Capacity) along with **diagnostic imaging** (CT) in a follow up period up to hundred(and some more) weeks and try to predict the changes of the measured lung function parameter based given a baseline value and a baseline CT.\n### \n\nIt is hypothetized that the **image data** has valuable information that not only correlate to the changes in the lung function tests but **can predict the progression**. \n\tThis connection has been already studied extensively (see references below).\n\n\t![changes](http://drkonya.com/projects/kaggle/fibrosis_1.png)\n### \n\nIn the Analysis of Diffuse Lung Disease we have several methods that ca be used to quantitatively analyse the images and find correlations between the derived information and the measured lung function parameters. Many of them (but not all!) rely on the segmentation of the lung tissue from the images.\n### \n\nThe **segmentation of the lung tissue by fibrosing lung diseases is challenging**: in contrary to the normal lung tissue, that has a low attenuation ranging -750-950 HU. The diseased lungs in this cohort have fibrous tissue (hence the name) that does not allow a \"simple\" threshold based segmentation of the lung parenchyma. In the previous Kaggle challenges there are many notebooks on lung segmentation (and many readily trained models on Github), so i do not go into details about it here, we will assume that you already segmented the lung.\n\n![threshold](http://drkonya.com/projects/kaggle/fibrosis_2.png)\n\n###Factors known to determine the outcome of the Idiopathic Pulmonary Fibrosis (IPF)\n\n\nBefore getting to the image analysis, few words about the course of the diffuse interstitial lung disease. The diffuse ILDs are characterized by infiltration of the interstitial compartment of the lung with varying degrees of inflammation and fibrosis - this is what we can see on the CT scan.\n\n### \n\nSeveral studies have investigated the relationship between biomarkers and outcomes at a single point in time. Baseline characteristics like: Male gender, age &gt; 70, diagnosis delay, grade of dyspnea, cardiovascular comorbidities, fibrosis extent on CT, pulmonary hypertension, blood test results (antibodies) have a negative effect on morbidity.  \n### \n### \nThe decrease in FVC % is the parameter of lung function that best predicts mortality. Changes between 6 or 12 months in the percentage of FVC [...] define the worsening, stability or improvement of the disease. [...] Some studies have shown a good correlation between the pulmonary function tests and the extension of IPF (abbreviation of intestinal pulmonal disease, in this context ~ ILD)  in HRCT findings. \t[Prognosis and Follow-Up of Idiopathic Pulmonary Fibrosis](https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6024649/)\n\n### \nThere are several EDA-s to this competition already that show the changes of epidemiological factors. These have to be taken into account as independent ( -or better to some degree dependent - ) variables to the features derived from the images.\n### \n\n###Image based features\n\nLet's assume we have our segmented lung, we want to derive information about the tissue there.\nThe first information should be the **volume of the lung** (let's now just forget the larger airways), since the CT is made in maximal inspiration, the so called total lung capacity can be estimated. It can be calculated as: pixel number on your lung segmentation map * pixel spacing (from DICOM header information, usually 0.625 mm/px ) * slice distance.  Slice distance ( ~= slice thickness) is derived from the difference between 2 consecutive slice's Image Position DICOM attribute - usually ranging from 1 mm to 7 mm. This gives you the volume of the lung in mm3. This information could be correlated with the FVC during the follow ups.\n### \n\nNext, one could analyse the lung with simple methods, like:\n### \n\n\t**Threshold based analysis:**\n\tSum of pixels below / above of a given threshold in your region of interest - the more fibrous tissue present, the higher this value gets for high values.\n### \n\n\tFurther with the **analysis of the histogram:** Mean, Skew, Kurthosis.\n\n\t**Mean** - average value is higher if fibrous tisue is present - the more fibrous tissue, the higher is the average; don't forget, we are talking about attenuation values here, where air has -1000 and water 0, fibrous tissue up to 50-70 (when calcified then much more). \n### \n\n\t**Skew** - normal lung is skewed to left (much more low attenuation values) whereas fibrous lung is skewed to the right (much more high attenuation values).\n### \n\n\t**Kurthosis** - the peak of the low attenuation pixels is much much lower (since we have more higher attenuating areas instead of it).\n### \n\n\t![histogram](http://drkonya.com/projects/kaggle/fibrosis_3.png)\n\n### \n\tIn a work from 2003, [Quantitative CT Indexes in Idiopathic Pulmonary Fibrosis: Relationship With Physiologic Impairment](https://pubmed.ncbi.nlm.nih.gov/12802000/) the authors claim that:\n### \n\n\"The lungs were isolated by using a semiautomated thresholding technique, with an upper threshold of -200 HU. [...] Pulmonary function tests (PFTs) included forced vital capacity, total lung capacity, forced expiratory volume in 1 second, and diffusing lung capacity. Moderate correlations existed between histogram features and PFT results. **Kurtosis showed the greatest degree of correlation with physiologic abnormality**\"  \n### \n\nSo ***why not to calculate these histogram features*** first - for the whole lung or later for smaller segments of the lung? But how to get smaller segments? You do not have to keep these segments restricted to anatomical constrains (lobes of the lung), you can for example define the topmost point the lung and beginning from that point divide your whole lung volume into smaller, for example 1 x 1 x 1 cm cubicles (blocks) and analyse these.\n\n### \n\n\t![block](http://drkonya.com/projects/kaggle/fibrosis_4.png)\n\n### \nIf we assume that similar tissue changes of the lung on different position in the lung itself contribute differently to the change in the spirometry results, then we need to know some information about the position/localisation of these changes.\n\nOne slice however contains multiple areas of different tissue changes == multiple different textures (normal tissue, altered tissue [like: honeycombing, septal thickening, emphysema, ground glass...]), so to divide these areas on one slice further down in regions is advisable. \n### \n\nThis further categorisation can be made with the help of trained NN (see the part on MedGift database) but also without NN:\n The idea of creating blocks comes initially [from here](https://arxiv.org/pdf/1811.02651.pdf), and we could use this technique to analyse the blocks and get information about the texture of different regions.\n### \n\n\nAdditionaly we can obtain an additional position/localisation value of each block.\n### \n\nIn order to get a localisation, we need a good definable point, and i chose for this the topmost point of the lung (called apex) since it is very simple to obtain: take the first slice that contains mask of your lung, and take the midpoint of this area.  The distance between this point and the midpoint of each block is then the distance i visioned. Other points could be set to, like the division of the trachea (called carina) but obtaining that point on the CT requires more segmentation work.\n### \n\t![block](http://drkonya.com/projects/kaggle/fibrosis_7.png)\n\nSince every people is different, we have to normalize somehow this distance (normalized distance), one possibility is to take the \"height\" of the whole lung and normalize the distance with this value. The height of the lung is simply the difference of the position of the first slice containing lung mask and position of the last slice containing lung mask.\n### \n\nThe changes of the \"diffuse\" lung disease are usually not \"diffuse\", they have localisation preferences (like by so called UIP pattern where subpleural and basal predominance is seen) and there are different patterns we radiologists describe, these are texture based features. These typical features are: honeycombing, reticular opacities, traction bronchiektasis, ground-glass-opacities (detailed desc for each [here](https://radiopaedia.org/articles/usual-interstitial-pneumonia?lang=gb) ).\n\n\n### \n\nThere are studies where texture based analysis of the lung had already been performed, for example: [Automatic Lung Segmentation Based on Texture and Deep Features of HRCT Images with Interstitial Lung Disease](https://www.hindawi.com/journals/bmri/2019/2045432/  ) \n\n### \n\nHere a very important database, the [MedGift database](http://medgift.hevs.ch/wordpress/databases/ild-database/) is described,  that is a collection of manually annotated CT's of 108 patients - a very good starting point for texture based studies  - one can train a semantic segmentation model to characterise the lung tissue (without even having to segment the lung itself beforehand) and compare these changes (like percentage of honeycombing or reticular pattern of the whole lung tissue and it's correlation with FVC).\n\n### \n\nI would start to analyse the changes in the extent of honeycombing. Why the honeycombing? Studies have shown that is has the largest influenceon lung functional tests:\n### \n\nSeveral studies have investigated the relationship between extent of fibrosis on CT and outcomes at a single point in time. Baseline patterns, in particular honeycombing [...] using non-volumetric CT, found that the overall extent of fibrosis, defined as the extent of reticulation and honeycombing on CT, was an independent predictor of mortality in IPF. Similarly, studies  [..] found that the overall extent of fibrosis indicated a poor prognosis in populations that contained not only patients with IPF but other fibrotic IIPs. The results from these studies indicate that increasing disease extent on CT is undeniably related to adverse prognosis in IPF. [...] In one study [...] it was shown that baseline honeycombing &gt;25%, fibrosis score &gt;30% and traction bronchiectasis in all four lobes were indicators of poor outcome.  - from [Evaluating disease severity in idiopathic pulmonary fibrosis](https://err.ersjournals.com/content/26/145/170051)\n\n### \n\nRadiomics based analysis will generate you quite much data, so chose wisely how big area you are analysing with it, probably dividing your lung parenchyma into 1x1x1 cm cubicles (block building) is a good starting point.\n### \n\nImportance of Airway- and vessel segmentations, coming soon.\n\n\t![inspiration](http://drkonya.com/projects/kaggle/inspi.jpg)\n\n",
    "1041652": "UPDATE\n\nDear fellow kagglers, \n\ni posted [segmentation masks](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/189528) after the competition.",
    "925161": "Very insightful. Thank you!",
    "924666": "Very informative. Seems like the first step is the segmentation, a task that still seems to have a precision not quite high enough as what would be wished in the field",
    "958201": "Thanks for your amazing insights. I implemented some of them in my [EDA notebook](https://www.kaggle.com/unforgiven/osic-comprehensive-eda), if you want to take a look.",
    "925141": "###UPDATE see the second part of my insights here: \nhttps://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/166123",
    "1028702": "Thank you for your informative post - Really helps for someone new like me.",
    "1015532": "Thanks for very useful insights!\nHowever, I have two questions for you.\n\n1.\n    You wrote \"It (the volume of the lung) can be calculated as: pixel number on your lung segmentation map * pixel spacing (from DICOM header information, usually 0.625 mm/px ) * slice distance\".\n    But, I think it is calculated as: pixel number on your lung segmentation map * pixel spacing (from DICOM header information, usually 0.625 mm/px )** 2 * slice distance\". (pixel spacing is multiplyed twice)\n    Is this wrong?\n\n2.\n    You wrote \"Slice distance ( ~= slice thickness) is derived from the difference between 2 consecutive slice's Image Position DICOM attribute - usually ranging from 1 mm to 7 mm. \"\n    How to get position of DICOM?\n\nI would be grateful if you could answer to these questions.",
    "1000721": "Thanks again for the insights, @sandorkonya !\n\nI'd love to pick your brain on ideas to move forward. So far, the most promising approach IMHO is derived from your ideas. \n\nLet's assume FVC declines are lines, from your first image, which alphas (intercepts) and betas (slope, i.e. the severity of the decline). Thus, the objective of the competition could be restated as predicting the beta for each patient from the baseline measurements & CT scan. If we have alphas & betas well predicted for each patient, we can infer FVCs for every possible week.\n\nUsing a Bayesian approach, we can learn that the 176 test patients have a distribution of betas normally distributed with center in -4.2 and std dev of 5 (all approx). This very simple to obtain. Then, if I sample a beta from this distribution for each test patient, and make their lines pass through their baseline FVC, and calculate the sigmas properly (I demonstrated in my last notebook), this yields a score of -6.90. That's our minimum score to beat. Anything worse than that is pure nonsense.\n\nNow, enter the CT scans. As you correctly pointed out, the job is to predict the severity of decline in FVC (betas) using features derived from the CT scans. So far, I implemented several features you suggested (volume in L, chest circumference in mm, height, HU histogram moments (mean, stddev, skew, kurtosis), % of lung volume in different ranges of HU, etc). I also tried deriving latent features learned with unsupervised learning (using a Variational Autoencoder, not public yet). \n\nThe problem is that this dataset with 176 CT scans with 176 betas is too small. There's too much noise: variance is very high, such that the gains for using all this noisy data are marginal in comparison to assuming all test patients have betas sampled from the same global distribution (i.e. all have the same severity). Even if I transform the problem into a classification problem (patients with betas lower than distribution median are \"high severity\", and higher than median are \"low severity\"), best I get is ~60-70% accuracy, loo low.\n \nYou mentioned the blocks idea. I'd love to use the MedGift database to classify every voxel of our patients, but one of the organizers said it is off-limits because of their license agreement.\n\nAnyway, I'm getting out of ideas on how to improve...",
    "977444": "Super interesting and helpful. Can't wait for the part on airway- and vessel segmentation!",
    "936228": "Thank you for comment. It was very insightful.",
    "936078": "Like @simonlfwalsh  has pointed out \"Change in traction bronchiectasis severity is a measure of disease progression that could be used to help resolve the clinical importance of marginal FVC declines.\"[source](https://thorax.bmj.com/content/75/8/648). I think its warranted to try to exploit this from the CT images provided in aiding the model to make accurate predictions.",
    "934130": "Very informative. Thank you very much for such great information. ",
    "931172": "Amazing article describing ways to extract useful feature from images.",
    "929868": "Very Interesting, well done!",
    "929814": "Another newbie question:\nI have my masks. In order to calculate the **volume of the lung**: The `pixel number` mentioned is the count of pixels in my mask right, assuming the mask is correctly segmenting the lung. For instance, if a patient has a mask with the form (31, 512, 512) - 31 slices of width and height 512. The pixel number would be 31 * 512 * 512.\n\nSorry for such elementary questions, it's just to ensure I don't waste time on silly mistakes.\n\nThank you.",
    "929219": "Just as a tip; honeycombing certainly is a predictor of outcome in these patients. However, traction bronchiectasis may be a much more sensitive predictor and so far, no one has successfully developed any automated approach to feature segmentation that can reliably quantify traction bronchiectasis. There is a fairly large body of literature based on visual assessments that suggests this and automated airway extraction specifically in fibrotic lung disease is considered an open research question.....\n\n[basic densitometry and related metrics has been investigated many times before, many years ago with reasonable albeit not earth-shattering results]\n\nNow get to work....!",
    "928396": "Very informative and explained with clarity ... will definitely be implementing these insights into my work, thank you! 🙂 ",
    "926450": "Domain knowledge is key. Thank you for your insights!",
    "926224": "I really appreciate your expert insight, it has given me so many ideas for making use of the image data. Thanks!",
    "925415": "I just read both of your insights. It is really helpful!\n\nI can't seem to comprehend the meaning of this statement:\n\n&gt; define the topmost point the lung and beginning from that point divide your whole lung volume into smaller, for example 1 x 1 x 1 cm cubicles (blocks) and analyse these - along with their distance as value to the topmost point - so not to lose localisation information. \n\nWhy is block segmentation needed for the mask? Did I miss something you wrote?",
    "924310": "cool",
    "928898": "",
    "1022260": "Thank you. It is great!",
    "933912": "thanks for the insight!",
    "931728": "Thank you dr. Konya! it's Interesting."
  }
}