{
  "id": 64979,
  "title": "How to get more labeled X-ray. A radiologist's proposal.",
  "url": "/competitions/rsna-pneumonia-detection-challenge/discussion/64979",
  "author_name": "",
  "post_date": "2018-09-04T22:35:08.941443300Z",
  "votes": 48,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I have a <strong>proposal to create hundreds of thousands of labeled image data</strong> to train your model.</p>\n\n<p><strong>The basic idea</strong> is</p>\n\n<blockquote>\n  <p>to use chest-CT volumetric data  <strong>to create pseudo X-Ray</strong> images by\n  3d to 2d projection for training models.</p>\n</blockquote>\n\n<p>Before you say (nah, low-res, supine position instead of standing, etc.) read my short rationale, we'll come back to these issues later.</p>\n\n<p><strong>1. Lung segmentation based solely on X-ray is not optimal.</strong></p>\n\n<p>In the analysis of Chest X-rays you would like to evaluate the changes of the pulmonal parenchyma. Everything else that summarizes over it, creating a final image, is disturbing factor (containing usually otherwise important information,  but not for this challenge). \nThis means that to train, ideally, you should segment the lungs first and use only that data; the upper abdomen, shoulders, lower neck is not of importance. Segmenting the lungs is a task that can be solved fairly simply (see example from the host <a href=\"https://github.com/mdai/ml-lessons/blob/master/lesson2-lung-xrays-segmentation.ipynb\">here</a>) - but their method is erroneous, even their demo image shows that a fair amount of the basal lung segments are cut off,  here opacies may hide  (and believe me they do), not to mention mediastinum that makes a challenge in only one projection (when no lateral projection is available)</p>\n\n<p><img src=\"http://drkonya.com/projects/kaggle/xray1.jpg\" alt=\"X-ray cut off\"></p>\n\n<p><strong>2. not every X-ray is orthograd projected</strong></p>\n\n<p>The assumption that every X-ray is orthograd projected is simply put: wrong and this has to be taken into consideration when analysing chest X-rays. The patients may move and rotate, resulting in an oblique projectied image of the chest and have silhouettes that may imply as pathology. On the image below the jugulum doesn't project onto the spine showing a rotation of the patient. </p>\n\n<p><img src=\"http://drkonya.com/projects/kaggle/xray2.jpg\" alt=\"off-center\"></p>\n\n<p><strong>So how could we overcome these problems?</strong></p>\n\n<p>There are many external databases that contain chest-CTs. Think of Data science bowl 2017 - with the 1M dollar challenge of analysing Nodules on CT. That data is (sadly) not available, but several others are (for example <a href=\"https://www.kaggle.com/c/data-science-bowl-2017/discussion/27666\">here</a>). If you take a CT volumetric data with high enough resolution (not more than 1-2 mm slice thickness) <strong>you can create a projected image of that volumetric data</strong> by 3d to 2d projection like this:</p>\n\n<p><img src=\"http://drkonya.com/projects/kaggle/xray3.jpg\" alt=\"enter image description here\"></p>\n\n<p>One would iterate through all the pixels of the pseudo x-ray image and trace the ray back to the source to create an intensity map calculated on all the voxel HU-values the beam surpassed. This is a (way) oversimplified version of the physic of the x-ray , but for now it could be enough (later on adding scattering, modelling different spectrum etc. may help).</p>\n\n<p><strong>This method will create not only one projected image!</strong> You can move the \"X-ray tube\" in space to create another projection of the same model, even slight off aligned. Let's assume 5 mm steps in all dimensions and allow X 3 cm, Y 6 cm, Z 15-20 cm - this is at least 36 x 15 projections AND we did not even rotate the body itself to get the oblique views. If we rotate with 2 degree steps up to 15 degree  in both directions = 15 steps.  <strong>This adds up to ~ 8000 projections</strong> , which is a fairly good amount of projected images <strong>from one CT-Volume</strong>. I've already found a <a href=\"https://gist.github.com/salmonmoose/2760072\">script on GitHub</a> that would could do the job (with some modification)</p>\n\n<p><img src=\"http://drkonya.com/projects/kaggle/xray4.jpg\" alt=\"3d to 2d\"></p>\n\n<p>Moreover, if you segment the lungs (for example with <a href=\"https://www.kaggle.com/ankasor/improved-lung-segmentation-using-watershed\">this</a> method from last year) you can create labeled images with <strong>correct</strong> segmentation of the lung parenchyma. Note the red areas overlapping the diaphragm.</p>\n\n<p><img src=\"http://drkonya.com/projects/kaggle/xray5.jpg\" alt=\"lung sementation\"></p>\n\n<p>Or even try to make better predictions by labeling the bones / trachea / hearth / vessels with the same method!</p>\n\n<p><strong>Just a random thought:</strong>\nWhat if we say, we would like to get all the signal originating not from the lung parenchyma and <strong>train a model</strong>  based on pseudo x-rays <strong>with the lung parenchyma completely removed</strong> so the end image only contains chest wall signal: bones and skin &amp; fat and the (great) vessels. That model could help differentiating a suspected area on a diseased image.</p>\n\n<p><strong>... but let's be more creative! Try to cheat! ;)</strong></p>\n\n<p>Why not take a step further, at least theoretically?\nWe know how a pneumonia looks like in a chest CT  - if not, take a look at <a href=\"https://radiopaedia.org/articles/lobar-pneumonia\">Radiopedia</a>. A lobar <strong>pneumonia can be theoretically</strong> relatively simply \"faked\" or better said <strong>generated computationally</strong>: Take a segment in the lung, and fill the voxels that do not belong to a bronchus or vessel with a higher value than air (= tissue), with respect to the anatomical boundaries (pleura). Add some shrinking of the filled volume with proportionally expanding the not affected lung areas of neighboring segments (not affecting the bronchus or vessel diameters) and there you are - again a <strong>3D voxel data that can generate thousands of pseudo x-rays</strong>. Make more affected areas with the same CT, you'll get more data to play with.The histogram of voxel densities in a diseased lung in a CT could be serve as basis for the HU values to be used by the creation of new pseudo-diseased lung like below. </p>\n\n<p><img src=\"http://drkonya.com/projects/kaggle/xray6.jpg\" alt=\"disease generation\"></p>\n\n<p>Oh, wait, <strong>where is my bounding box and label</strong> you may ask... the thing is, since you know which lung segment is affected by the \"disease generation\", you can simply <strong>project only the changed pixels onto your pseudo x-ray</strong> and there you go, that represents the affected area you can simply put a bounding box around.</p>\n\n<p>Not enough? Than my <strong>ultimate mind blowing idea</strong> at last:\nOne could try a modified style transfer algorithm that takes the pneumonia \"style\" from the diseased lung and creates a style transfer to a healthy lung segment. How about that?</p>\n\n<p><strong>Now back to earth.</strong> </p>\n\n<p>The brainstorming above was based on the assumption that good enough 3d -&gt; 2d projection can be achieved. If we take into account that the training images are usually downscaled, then our pseudo x-rays doesn't have to be UHD resolution. A 300 x 400 pixel should be enough (at least to try).</p>\n\n<p><strong>There are other issues that may affect the result</strong>: \nThe prone position of the patient - itself alone lead to changes for example: in the location and configuration of heart, diaphragm, vessel diameter - just to mention few. If there is a pleural effusion, the configuration will not be the same as if the patient would stay. The scapulae are rotated outwards (just because of the arm positioning in the prone position) and we could continue with several further example why this method could not work... but the real question is:</p>\n\n<h2>what do you think? Could this work?</h2>\n\n<p><img src=\"http://drkonya.com/projects/kaggle/inspi.jpg\" alt=\"inspiration\"></p>",
  "messages": [
    {
      "id": "381641",
      "postDate": "09/04/2018 22:35:08",
      "content": "<p>I have a <strong>proposal to create hundreds of thousands of labeled image data</strong> to train your model.</p>\n\n<p><strong>The basic idea</strong> is</p>\n\n<blockquote>\n  <p>to use chest-CT volumetric data  <strong>to create pseudo X-Ray</strong> images by\n  3d to 2d projection for training models.</p>\n</blockquote>\n\n<p>Before you say (nah, low-res, supine position instead of standing, etc.) read my short rationale, we'll come back to these issues later.</p>\n\n<p><strong>1. Lung segmentation based solely on X-ray is not optimal.</strong></p>\n\n<p>In the analysis of Chest X-rays you would like to evaluate the changes of the pulmonal parenchyma. Everything else that summarizes over it, creating a final image, is disturbing factor (containing usually otherwise important information,  but not for this challenge). \nThis means that to train, ideally, you should segment the lungs first and use only that data; the upper abdomen, shoulders, lower neck is not of importance. Segmenting the lungs is a task that can be solved fairly simply (see example from the host <a href=\"https://github.com/mdai/ml-lessons/blob/master/lesson2-lung-xrays-segmentation.ipynb\">here</a>) - but their method is erroneous, even their demo image shows that a fair amount of the basal lung segments are cut off,  here opacies may hide  (and believe me they do), not to mention mediastinum that makes a challenge in only one projection (when no lateral projection is available)</p>\n\n<p><img src=\"http://drkonya.com/projects/kaggle/xray1.jpg\" alt=\"X-ray cut off\"></p>\n\n<p><strong>2. not every X-ray is orthograd projected</strong></p>\n\n<p>The assumption that every X-ray is orthograd projected is simply put: wrong and this has to be taken into consideration when analysing chest X-rays. The patients may move and rotate, resulting in an oblique projectied image of the chest and have silhouettes that may imply as pathology. On the image below the jugulum doesn't project onto the spine showing a rotation of the patient. </p>\n\n<p><img src=\"http://drkonya.com/projects/kaggle/xray2.jpg\" alt=\"off-center\"></p>\n\n<p><strong>So how could we overcome these problems?</strong></p>\n\n<p>There are many external databases that contain chest-CTs. Think of Data science bowl 2017 - with the 1M dollar challenge of analysing Nodules on CT. That data is (sadly) not available, but several others are (for example <a href=\"https://www.kaggle.com/c/data-science-bowl-2017/discussion/27666\">here</a>). If you take a CT volumetric data with high enough resolution (not more than 1-2 mm slice thickness) <strong>you can create a projected image of that volumetric data</strong> by 3d to 2d projection like this:</p>\n\n<p><img src=\"http://drkonya.com/projects/kaggle/xray3.jpg\" alt=\"enter image description here\"></p>\n\n<p>One would iterate through all the pixels of the pseudo x-ray image and trace the ray back to the source to create an intensity map calculated on all the voxel HU-values the beam surpassed. This is a (way) oversimplified version of the physic of the x-ray , but for now it could be enough (later on adding scattering, modelling different spectrum etc. may help).</p>\n\n<p><strong>This method will create not only one projected image!</strong> You can move the \"X-ray tube\" in space to create another projection of the same model, even slight off aligned. Let's assume 5 mm steps in all dimensions and allow X 3 cm, Y 6 cm, Z 15-20 cm - this is at least 36 x 15 projections AND we did not even rotate the body itself to get the oblique views. If we rotate with 2 degree steps up to 15 degree  in both directions = 15 steps.  <strong>This adds up to ~ 8000 projections</strong> , which is a fairly good amount of projected images <strong>from one CT-Volume</strong>. I've already found a <a href=\"https://gist.github.com/salmonmoose/2760072\">script on GitHub</a> that would could do the job (with some modification)</p>\n\n<p><img src=\"http://drkonya.com/projects/kaggle/xray4.jpg\" alt=\"3d to 2d\"></p>\n\n<p>Moreover, if you segment the lungs (for example with <a href=\"https://www.kaggle.com/ankasor/improved-lung-segmentation-using-watershed\">this</a> method from last year) you can create labeled images with <strong>correct</strong> segmentation of the lung parenchyma. Note the red areas overlapping the diaphragm.</p>\n\n<p><img src=\"http://drkonya.com/projects/kaggle/xray5.jpg\" alt=\"lung sementation\"></p>\n\n<p>Or even try to make better predictions by labeling the bones / trachea / hearth / vessels with the same method!</p>\n\n<p><strong>Just a random thought:</strong>\nWhat if we say, we would like to get all the signal originating not from the lung parenchyma and <strong>train a model</strong>  based on pseudo x-rays <strong>with the lung parenchyma completely removed</strong> so the end image only contains chest wall signal: bones and skin &amp; fat and the (great) vessels. That model could help differentiating a suspected area on a diseased image.</p>\n\n<p><strong>... but let's be more creative! Try to cheat! ;)</strong></p>\n\n<p>Why not take a step further, at least theoretically?\nWe know how a pneumonia looks like in a chest CT  - if not, take a look at <a href=\"https://radiopaedia.org/articles/lobar-pneumonia\">Radiopedia</a>. A lobar <strong>pneumonia can be theoretically</strong> relatively simply \"faked\" or better said <strong>generated computationally</strong>: Take a segment in the lung, and fill the voxels that do not belong to a bronchus or vessel with a higher value than air (= tissue), with respect to the anatomical boundaries (pleura). Add some shrinking of the filled volume with proportionally expanding the not affected lung areas of neighboring segments (not affecting the bronchus or vessel diameters) and there you are - again a <strong>3D voxel data that can generate thousands of pseudo x-rays</strong>. Make more affected areas with the same CT, you'll get more data to play with.The histogram of voxel densities in a diseased lung in a CT could be serve as basis for the HU values to be used by the creation of new pseudo-diseased lung like below. </p>\n\n<p><img src=\"http://drkonya.com/projects/kaggle/xray6.jpg\" alt=\"disease generation\"></p>\n\n<p>Oh, wait, <strong>where is my bounding box and label</strong> you may ask... the thing is, since you know which lung segment is affected by the \"disease generation\", you can simply <strong>project only the changed pixels onto your pseudo x-ray</strong> and there you go, that represents the affected area you can simply put a bounding box around.</p>\n\n<p>Not enough? Than my <strong>ultimate mind blowing idea</strong> at last:\nOne could try a modified style transfer algorithm that takes the pneumonia \"style\" from the diseased lung and creates a style transfer to a healthy lung segment. How about that?</p>\n\n<p><strong>Now back to earth.</strong> </p>\n\n<p>The brainstorming above was based on the assumption that good enough 3d -&gt; 2d projection can be achieved. If we take into account that the training images are usually downscaled, then our pseudo x-rays doesn't have to be UHD resolution. A 300 x 400 pixel should be enough (at least to try).</p>\n\n<p><strong>There are other issues that may affect the result</strong>: \nThe prone position of the patient - itself alone lead to changes for example: in the location and configuration of heart, diaphragm, vessel diameter - just to mention few. If there is a pleural effusion, the configuration will not be the same as if the patient would stay. The scapulae are rotated outwards (just because of the arm positioning in the prone position) and we could continue with several further example why this method could not work... but the real question is:</p>\n\n<h2>what do you think? Could this work?</h2>\n\n<p><img src=\"http://drkonya.com/projects/kaggle/inspi.jpg\" alt=\"inspiration\"></p>",
      "rawMarkdown": "I have a **proposal to create hundreds of thousands of labeled image data** to train your model.\n\n**The basic idea** is\n\n&gt; to use chest-CT volumetric data  **to create pseudo X-Ray** images by\n&gt; 3d to 2d projection for training models.\n\nBefore you say (nah, low-res, supine position instead of standing, etc.) read my short rationale, we'll come back to these issues later.\n\n**1. Lung segmentation based solely on X-ray is not optimal.**\n\nIn the analysis of Chest X-rays you would like to evaluate the changes of the pulmonal parenchyma. Everything else that summarizes over it, creating a final image, is disturbing factor (containing usually otherwise important information,  but not for this challenge). \nThis means that to train, ideally, you should segment the lungs first and use only that data; the upper abdomen, shoulders, lower neck is not of importance. Segmenting the lungs is a task that can be solved fairly simply (see example from the host [here][1]) - but their method is erroneous, even their demo image shows that a fair amount of the basal lung segments are cut off,  here opacies may hide  (and believe me they do), not to mention mediastinum that makes a challenge in only one projection (when no lateral projection is available)\n\n![X-ray cut off][2]\n\n\n**2. not every X-ray is orthograd projected**\n\nThe assumption that every X-ray is orthograd projected is simply put: wrong and this has to be taken into consideration when analysing chest X-rays. The patients may move and rotate, resulting in an oblique projectied image of the chest and have silhouettes that may imply as pathology. On the image below the jugulum doesn't project onto the spine showing a rotation of the patient. \n\n![off-center][3]\n\n**So how could we overcome these problems?**\n\nThere are many external databases that contain chest-CTs. Think of Data science bowl 2017 - with the 1M dollar challenge of analysing Nodules on CT. That data is (sadly) not available, but several others are (for example [here][4]). If you take a CT volumetric data with high enough resolution (not more than 1-2 mm slice thickness) **you can create a projected image of that volumetric data** by 3d to 2d projection like this:\n\n![enter image description here][5]\n\nOne would iterate through all the pixels of the pseudo x-ray image and trace the ray back to the source to create an intensity map calculated on all the voxel HU-values the beam surpassed. This is a (way) oversimplified version of the physic of the x-ray , but for now it could be enough (later on adding scattering, modelling different spectrum etc. may help).\n\n**This method will create not only one projected image!** You can move the \"X-ray tube\" in space to create another projection of the same model, even slight off aligned. Let's assume 5 mm steps in all dimensions and allow X 3 cm, Y 6 cm, Z 15-20 cm - this is at least 36 x 15 projections AND we did not even rotate the body itself to get the oblique views. If we rotate with 2 degree steps up to 15 degree  in both directions = 15 steps.  **This adds up to ~ 8000 projections** , which is a fairly good amount of projected images **from one CT-Volume**. I've already found a [script on GitHub][6] that would could do the job (with some modification)\n\n![3d to 2d][7]\n\nMoreover, if you segment the lungs (for example with [this][8] method from last year) you can create labeled images with **correct** segmentation of the lung parenchyma. Note the red areas overlapping the diaphragm.\n\n![lung sementation][9]\n\nOr even try to make better predictions by labeling the bones / trachea / hearth / vessels with the same method!\n\n**Just a random thought:**\nWhat if we say, we would like to get all the signal originating not from the lung parenchyma and **train a model**  based on pseudo x-rays **with the lung parenchyma completely removed** so the end image only contains chest wall signal: bones and skin &amp; fat and the (great) vessels. That model could help differentiating a suspected area on a diseased image.\n\n**... but let's be more creative! Try to cheat! ;)**\n\nWhy not take a step further, at least theoretically?\nWe know how a pneumonia looks like in a chest CT  - if not, take a look at [Radiopedia][10]. A lobar **pneumonia can be theoretically** relatively simply \"faked\" or better said **generated computationally**: Take a segment in the lung, and fill the voxels that do not belong to a bronchus or vessel with a higher value than air (= tissue), with respect to the anatomical boundaries (pleura). Add some shrinking of the filled volume with proportionally expanding the not affected lung areas of neighboring segments (not affecting the bronchus or vessel diameters) and there you are - again a **3D voxel data that can generate thousands of pseudo x-rays**. Make more affected areas with the same CT, you'll get more data to play with.The histogram of voxel densities in a diseased lung in a CT could be serve as basis for the HU values to be used by the creation of new pseudo-diseased lung like below. \n\n![disease generation][11]\n\nOh, wait, **where is my bounding box and label** you may ask... the thing is, since you know which lung segment is affected by the \"disease generation\", you can simply **project only the changed pixels onto your pseudo x-ray** and there you go, that represents the affected area you can simply put a bounding box around.\n\nNot enough? Than my **ultimate mind blowing idea** at last:\nOne could try a modified style transfer algorithm that takes the pneumonia \"style\" from the diseased lung and creates a style transfer to a healthy lung segment. How about that?\n\n**Now back to earth.** \n\nThe brainstorming above was based on the assumption that good enough 3d -&gt; 2d projection can be achieved. If we take into account that the training images are usually downscaled, then our pseudo x-rays doesn't have to be UHD resolution. A 300 x 400 pixel should be enough (at least to try).\n\n **There are other issues that may affect the result**: \nThe prone position of the patient - itself alone lead to changes for example: in the location and configuration of heart, diaphragm, vessel diameter - just to mention few. If there is a pleural effusion, the configuration will not be the same as if the patient would stay. The scapulae are rotated outwards (just because of the arm positioning in the prone position) and we could continue with several further example why this method could not work... but the real question is:\n## what do you think? Could this work?\n\n![inspiration][12]\n\n\n  [1]: https://github.com/mdai/ml-lessons/blob/master/lesson2-lung-xrays-segmentation.ipynb\n  [2]: http://drkonya.com/projects/kaggle/xray1.jpg\n  [3]: http://drkonya.com/projects/kaggle/xray2.jpg\n  [4]: https://www.kaggle.com/c/data-science-bowl-2017/discussion/27666\n  [5]: http://drkonya.com/projects/kaggle/xray3.jpg\n  [6]: https://gist.github.com/salmonmoose/2760072\n  [7]: http://drkonya.com/projects/kaggle/xray4.jpg\n  [8]: https://www.kaggle.com/ankasor/improved-lung-segmentation-using-watershed\n  [9]: http://drkonya.com/projects/kaggle/xray5.jpg\n  [10]: https://radiopaedia.org/articles/lobar-pneumonia\n  [11]: http://drkonya.com/projects/kaggle/xray6.jpg\n  [12]: http://drkonya.com/projects/kaggle/inspi.jpg",
      "votes": null
    },
    {
      "id": "574917",
      "postDate": "07/14/2019 17:30:59",
      "content": "<p>absolute gem!</p>",
      "rawMarkdown": "absolute gem!",
      "votes": null
    },
    {
      "id": "610026",
      "postDate": "08/28/2019 10:25:02",
      "content": "<p>yeah !!</p>",
      "rawMarkdown": "yeah !!",
      "votes": null
    },
    {
      "id": "616111",
      "postDate": "09/02/2019 17:45:22",
      "content": "<p>Me and <a href=\"/sainatarajan7\">@sainatarajan7</a> are already working on this and got some neat results that we will publish later!</p>",
      "rawMarkdown": "Me and @sainatarajan7 are already working on this and got some neat results that we will publish later!",
      "votes": null
    },
    {
      "id": "1057527",
      "postDate": "10/22/2020 18:15:19",
      "content": "<p>Great idea!</p>",
      "rawMarkdown": "Great idea!",
      "votes": null
    },
    {
      "id": "1066826",
      "postDate": "11/02/2020 04:55:01",
      "content": "<p>Hi Konya, do you have any code examples on how to create pseudo X-Ray images?</p>",
      "rawMarkdown": "Hi Konya, do you have any code examples on how to create pseudo X-Ray images?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1057527,
      "author_name": "bardiakh",
      "author_url": "",
      "post_date": "10/22/2020 18:15:19",
      "content": "<p>Great idea!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1066826,
      "author_name": "yuanthony",
      "author_url": "",
      "post_date": "11/02/2020 04:55:01",
      "content": "<p>Hi Konya, do you have any code examples on how to create pseudo X-Ray images?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 574917,
      "author_name": "ekhtiar",
      "author_url": "",
      "post_date": "07/14/2019 17:30:59",
      "content": "<p>absolute gem!</p>",
      "votes": null,
      "replies": [
        {
          "id": 610026,
          "author_name": "yashvi1502",
          "author_url": "",
          "post_date": "08/28/2019 10:25:02",
          "content": "<p>yeah !!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 616111,
          "author_name": "sandorkonya",
          "author_url": "",
          "post_date": "09/02/2019 17:45:22",
          "content": "<p>Me and <a href=\"/sainatarajan7\">@sainatarajan7</a> are already working on this and got some neat results that we will publish later!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "381641": "I have a **proposal to create hundreds of thousands of labeled image data** to train your model.\n\n**The basic idea** is\n\n&gt; to use chest-CT volumetric data  **to create pseudo X-Ray** images by\n&gt; 3d to 2d projection for training models.\n\nBefore you say (nah, low-res, supine position instead of standing, etc.) read my short rationale, we'll come back to these issues later.\n\n**1. Lung segmentation based solely on X-ray is not optimal.**\n\nIn the analysis of Chest X-rays you would like to evaluate the changes of the pulmonal parenchyma. Everything else that summarizes over it, creating a final image, is disturbing factor (containing usually otherwise important information,  but not for this challenge). \nThis means that to train, ideally, you should segment the lungs first and use only that data; the upper abdomen, shoulders, lower neck is not of importance. Segmenting the lungs is a task that can be solved fairly simply (see example from the host [here][1]) - but their method is erroneous, even their demo image shows that a fair amount of the basal lung segments are cut off,  here opacies may hide  (and believe me they do), not to mention mediastinum that makes a challenge in only one projection (when no lateral projection is available)\n\n![X-ray cut off][2]\n\n\n**2. not every X-ray is orthograd projected**\n\nThe assumption that every X-ray is orthograd projected is simply put: wrong and this has to be taken into consideration when analysing chest X-rays. The patients may move and rotate, resulting in an oblique projectied image of the chest and have silhouettes that may imply as pathology. On the image below the jugulum doesn't project onto the spine showing a rotation of the patient. \n\n![off-center][3]\n\n**So how could we overcome these problems?**\n\nThere are many external databases that contain chest-CTs. Think of Data science bowl 2017 - with the 1M dollar challenge of analysing Nodules on CT. That data is (sadly) not available, but several others are (for example [here][4]). If you take a CT volumetric data with high enough resolution (not more than 1-2 mm slice thickness) **you can create a projected image of that volumetric data** by 3d to 2d projection like this:\n\n![enter image description here][5]\n\nOne would iterate through all the pixels of the pseudo x-ray image and trace the ray back to the source to create an intensity map calculated on all the voxel HU-values the beam surpassed. This is a (way) oversimplified version of the physic of the x-ray , but for now it could be enough (later on adding scattering, modelling different spectrum etc. may help).\n\n**This method will create not only one projected image!** You can move the \"X-ray tube\" in space to create another projection of the same model, even slight off aligned. Let's assume 5 mm steps in all dimensions and allow X 3 cm, Y 6 cm, Z 15-20 cm - this is at least 36 x 15 projections AND we did not even rotate the body itself to get the oblique views. If we rotate with 2 degree steps up to 15 degree  in both directions = 15 steps.  **This adds up to ~ 8000 projections** , which is a fairly good amount of projected images **from one CT-Volume**. I've already found a [script on GitHub][6] that would could do the job (with some modification)\n\n![3d to 2d][7]\n\nMoreover, if you segment the lungs (for example with [this][8] method from last year) you can create labeled images with **correct** segmentation of the lung parenchyma. Note the red areas overlapping the diaphragm.\n\n![lung sementation][9]\n\nOr even try to make better predictions by labeling the bones / trachea / hearth / vessels with the same method!\n\n**Just a random thought:**\nWhat if we say, we would like to get all the signal originating not from the lung parenchyma and **train a model**  based on pseudo x-rays **with the lung parenchyma completely removed** so the end image only contains chest wall signal: bones and skin &amp; fat and the (great) vessels. That model could help differentiating a suspected area on a diseased image.\n\n**... but let's be more creative! Try to cheat! ;)**\n\nWhy not take a step further, at least theoretically?\nWe know how a pneumonia looks like in a chest CT  - if not, take a look at [Radiopedia][10]. A lobar **pneumonia can be theoretically** relatively simply \"faked\" or better said **generated computationally**: Take a segment in the lung, and fill the voxels that do not belong to a bronchus or vessel with a higher value than air (= tissue), with respect to the anatomical boundaries (pleura). Add some shrinking of the filled volume with proportionally expanding the not affected lung areas of neighboring segments (not affecting the bronchus or vessel diameters) and there you are - again a **3D voxel data that can generate thousands of pseudo x-rays**. Make more affected areas with the same CT, you'll get more data to play with.The histogram of voxel densities in a diseased lung in a CT could be serve as basis for the HU values to be used by the creation of new pseudo-diseased lung like below. \n\n![disease generation][11]\n\nOh, wait, **where is my bounding box and label** you may ask... the thing is, since you know which lung segment is affected by the \"disease generation\", you can simply **project only the changed pixels onto your pseudo x-ray** and there you go, that represents the affected area you can simply put a bounding box around.\n\nNot enough? Than my **ultimate mind blowing idea** at last:\nOne could try a modified style transfer algorithm that takes the pneumonia \"style\" from the diseased lung and creates a style transfer to a healthy lung segment. How about that?\n\n**Now back to earth.** \n\nThe brainstorming above was based on the assumption that good enough 3d -&gt; 2d projection can be achieved. If we take into account that the training images are usually downscaled, then our pseudo x-rays doesn't have to be UHD resolution. A 300 x 400 pixel should be enough (at least to try).\n\n **There are other issues that may affect the result**: \nThe prone position of the patient - itself alone lead to changes for example: in the location and configuration of heart, diaphragm, vessel diameter - just to mention few. If there is a pleural effusion, the configuration will not be the same as if the patient would stay. The scapulae are rotated outwards (just because of the arm positioning in the prone position) and we could continue with several further example why this method could not work... but the real question is:\n## what do you think? Could this work?\n\n![inspiration][12]\n\n\n  [1]: https://github.com/mdai/ml-lessons/blob/master/lesson2-lung-xrays-segmentation.ipynb\n  [2]: http://drkonya.com/projects/kaggle/xray1.jpg\n  [3]: http://drkonya.com/projects/kaggle/xray2.jpg\n  [4]: https://www.kaggle.com/c/data-science-bowl-2017/discussion/27666\n  [5]: http://drkonya.com/projects/kaggle/xray3.jpg\n  [6]: https://gist.github.com/salmonmoose/2760072\n  [7]: http://drkonya.com/projects/kaggle/xray4.jpg\n  [8]: https://www.kaggle.com/ankasor/improved-lung-segmentation-using-watershed\n  [9]: http://drkonya.com/projects/kaggle/xray5.jpg\n  [10]: https://radiopaedia.org/articles/lobar-pneumonia\n  [11]: http://drkonya.com/projects/kaggle/xray6.jpg\n  [12]: http://drkonya.com/projects/kaggle/inspi.jpg",
    "574917": "absolute gem!",
    "610026": "yeah !!",
    "616111": "Me and @sainatarajan7 are already working on this and got some neat results that we will publish later!",
    "1057527": "Great idea!",
    "1066826": "Hi Konya, do you have any code examples on how to create pseudo X-Ray images?"
  },
  "source": "meta"
}