{
  "id": 64310,
  "title": "Extracting age and gender from DICOM?",
  "url": "/competitions/rsna-pneumonia-detection-challenge/discussion/64310",
  "author_name": "",
  "post_date": "2018-08-28T03:32:52.335203200Z",
  "votes": 4,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I think it was this TWIML episode where the Googler said that he observed that a DNN can perform better if you give it more stuff to learn <a href=\"https://twimlai.com/twiml-talk-122-predicting-cardiovascular-risk-factors-eye-images-ryan-poplin/\">https://twimlai.com/twiml-talk-122-predicting-cardiovascular-risk-factors-eye-images-ryan-poplin/</a></p>\n\n<p>In other words, instead of just training it to output a probability of pheumonia + bounding box, if you also train the network to output other factors from the image, such as age, gender, etc., the network can perform better at the core task you're after (in this case, probability of pneumonia with bounding boxes). I'm curious to experiment with that.</p>\n\n<p>My question: I hadn't worked with DICOM files before. Using <code>pydicom</code>, it's trivial to parse the file with:</p>\n\n<pre><code>dcm_data = pydicom.read_file(dcm_file)\n</code></pre>\n\n<p>But from there, how do I access the individual lines in the meta description? For example, output of the above might look like:</p>\n\n<pre><code>print(dcm_data)\n(0008, 0005) Specific Character Set              CS: 'ISO_IR 100'\n(0008, 0020) Study Date                          DA: '19010101'\n(0008, 0030) Study Time                          TM: '000000.00'\n(0008, 0050) Accession Number                    SH: ''\n(0008, 0060) Modality                            CS: 'CR'\n(0008, 0064) Conversion Type                     CS: 'WSD'\n(0008, 0090) Referring Physician's Name          PN: ''\n(0008, 103e) Series Description                  LO: 'view: PA'\n(0010, 0010) Patient's Name                      PN: '0004cfab-14fd'\n(0010, 0020) Patient ID                          LO: '0004cfab-14fd'\n(0010, 0030) Patient's Birth Date                DA: ''\n(0010, 0040) Patient's Sex                       CS: 'F'\n(0010, 1010) Patient's Age                       AS: '51'\n(0018, 0015) Body Part Examined                  CS: 'CHEST'\n</code></pre>\n\n<p>Is there a better way to get the value of <code>(0010, 0040) =&amp;gt; Patient's Sex</code> from <code>dcm_data</code>? </p>\n\n<p>Thanks.</p>",
  "messages": [
    {
      "id": "376774",
      "postDate": "08/28/2018 03:32:52",
      "content": "<p>I think it was this TWIML episode where the Googler said that he observed that a DNN can perform better if you give it more stuff to learn <a href=\"https://twimlai.com/twiml-talk-122-predicting-cardiovascular-risk-factors-eye-images-ryan-poplin/\">https://twimlai.com/twiml-talk-122-predicting-cardiovascular-risk-factors-eye-images-ryan-poplin/</a></p>\n\n<p>In other words, instead of just training it to output a probability of pheumonia + bounding box, if you also train the network to output other factors from the image, such as age, gender, etc., the network can perform better at the core task you're after (in this case, probability of pneumonia with bounding boxes). I'm curious to experiment with that.</p>\n\n<p>My question: I hadn't worked with DICOM files before. Using <code>pydicom</code>, it's trivial to parse the file with:</p>\n\n<pre><code>dcm_data = pydicom.read_file(dcm_file)\n</code></pre>\n\n<p>But from there, how do I access the individual lines in the meta description? For example, output of the above might look like:</p>\n\n<pre><code>print(dcm_data)\n(0008, 0005) Specific Character Set              CS: 'ISO_IR 100'\n(0008, 0020) Study Date                          DA: '19010101'\n(0008, 0030) Study Time                          TM: '000000.00'\n(0008, 0050) Accession Number                    SH: ''\n(0008, 0060) Modality                            CS: 'CR'\n(0008, 0064) Conversion Type                     CS: 'WSD'\n(0008, 0090) Referring Physician's Name          PN: ''\n(0008, 103e) Series Description                  LO: 'view: PA'\n(0010, 0010) Patient's Name                      PN: '0004cfab-14fd'\n(0010, 0020) Patient ID                          LO: '0004cfab-14fd'\n(0010, 0030) Patient's Birth Date                DA: ''\n(0010, 0040) Patient's Sex                       CS: 'F'\n(0010, 1010) Patient's Age                       AS: '51'\n(0018, 0015) Body Part Examined                  CS: 'CHEST'\n</code></pre>\n\n<p>Is there a better way to get the value of <code>(0010, 0040) =&amp;gt; Patient's Sex</code> from <code>dcm_data</code>? </p>\n\n<p>Thanks.</p>",
      "rawMarkdown": "I think it was this TWIML episode where the Googler said that he observed that a DNN can perform better if you give it more stuff to learn https://twimlai.com/twiml-talk-122-predicting-cardiovascular-risk-factors-eye-images-ryan-poplin/\n\nIn other words, instead of just training it to output a probability of pheumonia + bounding box, if you also train the network to output other factors from the image, such as age, gender, etc., the network can perform better at the core task you're after (in this case, probability of pneumonia with bounding boxes). I'm curious to experiment with that.\n\nMy question: I hadn't worked with DICOM files before. Using `pydicom`, it's trivial to parse the file with:\n\n    dcm_data = pydicom.read_file(dcm_file)\n\nBut from there, how do I access the individual lines in the meta description? For example, output of the above might look like:\n\n    print(dcm_data)\n    (0008, 0005) Specific Character Set              CS: 'ISO_IR 100'\n    (0008, 0020) Study Date                          DA: '19010101'\n    (0008, 0030) Study Time                          TM: '000000.00'\n    (0008, 0050) Accession Number                    SH: ''\n    (0008, 0060) Modality                            CS: 'CR'\n    (0008, 0064) Conversion Type                     CS: 'WSD'\n    (0008, 0090) Referring Physician's Name          PN: ''\n    (0008, 103e) Series Description                  LO: 'view: PA'\n    (0010, 0010) Patient's Name                      PN: '0004cfab-14fd'\n    (0010, 0020) Patient ID                          LO: '0004cfab-14fd'\n    (0010, 0030) Patient's Birth Date                DA: ''\n    (0010, 0040) Patient's Sex                       CS: 'F'\n    (0010, 1010) Patient's Age                       AS: '51'\n    (0018, 0015) Body Part Examined                  CS: 'CHEST'\n\nIs there a better way to get the value of `(0010, 0040) =&gt; Patient's Sex` from `dcm_data`? \n\nThanks.",
      "votes": null
    },
    {
      "id": "376897",
      "postDate": "08/28/2018 08:48:52",
      "content": "<p>Hi, formigone.\nI have the same idea and made a short notebook. How about my approach?\n<a href=\"https://www.kaggle.com/masato/patient-data-in-dicom-file\">https://www.kaggle.com/masato/patient-data-in-dicom-file</a></p>",
      "rawMarkdown": "Hi, formigone.\nI have the same idea and made a short notebook. How about my approach?\nhttps://www.kaggle.com/masato/patient-data-in-dicom-file",
      "votes": null
    },
    {
      "id": "376903",
      "postDate": "08/28/2018 09:01:33",
      "content": "<p>Is this what you're looking for ?  (My first time using this file format as well)</p>\n\n<p>dcm_data[0x00100040].value </p>\n\n<p>--&gt; 'F'</p>",
      "rawMarkdown": "Is this what you're looking for ?  (My first time using this file format as well)\n\n dcm_data[0x00100040].value \n\n--&gt; 'F'",
      "votes": null
    },
    {
      "id": "377226",
      "postDate": "08/28/2018 19:14:06",
      "content": "<p>Thanks. I'll take a look. My fear is that <code>0x00100040</code> will not consistently map to that property. What you suggest is a good place to start, though. Thanks.</p>",
      "rawMarkdown": "Thanks. I'll take a look. My fear is that `0x00100040` will not consistently map to that property. What you suggest is a good place to start, though. Thanks.",
      "votes": null
    },
    {
      "id": "377227",
      "postDate": "08/28/2018 19:16:12",
      "content": "<p>That looks awesome. Great work.</p>",
      "rawMarkdown": "That looks awesome. Great work.",
      "votes": null
    },
    {
      "id": "377254",
      "postDate": "08/28/2018 20:35:31",
      "content": "<p>Because the images conform to the DICOM standard, that hex address always corresponds to Patient Sex. Here’s a list of other tags, although many of them have been modified or stripped in the anonimization process: <a href=\"https://sno.phy.queensu.ca/~phil/exiftool/TagNames/DICOM.html\">https://sno.phy.queensu.ca/~phil/exiftool/TagNames/DICOM.html</a></p>",
      "rawMarkdown": "Because the images conform to the DICOM standard, that hex address always corresponds to Patient Sex. Here’s a list of other tags, although many of them have been modified or stripped in the anonimization process: https://sno.phy.queensu.ca/~phil/exiftool/TagNames/DICOM.html",
      "votes": null
    },
    {
      "id": "379844",
      "postDate": "09/01/2018 03:23:24",
      "content": "<p>Hi, formigone. When viewed in DICOM image viewer there are two values C and W at bottom left of the image. Do you have any information what these values could be or be useful in making prediction?</p>",
      "rawMarkdown": "Hi, formigone. When viewed in DICOM image viewer there are two values C and W at bottom left of the image. Do you have any information what these values could be or be useful in making prediction?",
      "votes": null
    },
    {
      "id": "380118",
      "postDate": "09/01/2018 19:02:04",
      "content": "<p>W = Window\nC = Center</p>\n\n<p>This happens because the images contain more greyscale values than can be displayed at a single time.  So - the setting of a center and a window allows you to map the pixel values to greyscale values.</p>\n\n<p>Most DICOM images come with default values for W and C, but they are essentially suggestions based on the acquisition device.</p>\n\n<p>It's much more relevant for displaying CT images as described in this radiopedia article:</p>\n\n<p><a href=\"https://radiopaedia.org/articles/windowing-ct\">https://radiopaedia.org/articles/windowing-ct</a></p>\n\n<p>Cheers!</p>",
      "rawMarkdown": "W = Window\nC = Center\n\nThis happens because the images contain more greyscale values than can be displayed at a single time.  So - the setting of a center and a window allows you to map the pixel values to greyscale values.\n\nMost DICOM images come with default values for W and C, but they are essentially suggestions based on the acquisition device.\n\nIt's much more relevant for displaying CT images as described in this radiopedia article:\n\nhttps://radiopaedia.org/articles/windowing-ct\n\nCheers!",
      "votes": null
    },
    {
      "id": "380206",
      "postDate": "09/02/2018 03:40:40",
      "content": "<p>Hey formigone, I have a notebook that has a method that (I think) reliably parses all the keywords and their values (as well as constructs a map from the DICOM tag to the keyword). <a href=\"https://www.kaggle.com/jtlowery/intro-eda-with-dicom-metadata/\">notebook - click the 'Parsing Metadata from DICOM Object' to skip right to it</a></p>",
      "rawMarkdown": "Hey formigone, I have a notebook that has a method that (I think) reliably parses all the keywords and their values (as well as constructs a map from the DICOM tag to the keyword). [notebook - click the 'Parsing Metadata from DICOM Object' to skip right to it][1]\n\n\n  [1]: https://www.kaggle.com/jtlowery/intro-eda-with-dicom-metadata/",
      "votes": null
    },
    {
      "id": "381115",
      "postDate": "09/04/2018 04:36:07",
      "content": "<p>Hi formigone,</p>\n\n<p>You might find the following kernels and dataset useful on how to extract DICOM attributes:  <a href=\"https://www.kaggle.com/kanwalinder/rsna-preprocessed-nonimage-inputs\">RSNA Preprocessed Non-Image Inputs Dataset</a>, generated by <a href=\"https://www.kaggle.com/kanwalinder/preprocessing-rsna-non-image-inputs\">Preprocessing RSNA Non-Image Inputs</a>, and used by <a href=\"https://www.kaggle.com/kanwalinder/rsna-review-inputs\">RSNA Review Inputs</a>.</p>",
      "rawMarkdown": "Hi formigone,\n\nYou might find the following kernels and dataset useful on how to extract DICOM attributes:  [RSNA Preprocessed Non-Image Inputs Dataset][1], generated by [Preprocessing RSNA Non-Image Inputs][2], and used by [RSNA Review Inputs][3].\n\n\n  [1]: https://www.kaggle.com/kanwalinder/rsna-preprocessed-nonimage-inputs\n  [2]: https://www.kaggle.com/kanwalinder/preprocessing-rsna-non-image-inputs\n  [3]: https://www.kaggle.com/kanwalinder/rsna-review-inputs",
      "votes": null
    },
    {
      "id": "381674",
      "postDate": "09/05/2018 01:15:18",
      "content": "<p>Looks like this may have been answered in other ways. Just in case though, here's my simple approach:</p>\n\n<p>dcm_data = pydicom.read_file(dcm_file)</p>\n\n<p>age = dcm_data.PatientAge</p>\n\n<p>sex = dcm_data.PatientSex</p>",
      "rawMarkdown": "Looks like this may have been answered in other ways. Just in case though, here's my simple approach:\n\ndcm_data = pydicom.read_file(dcm_file)\n\nage = dcm_data.PatientAge\n\nsex = dcm_data.PatientSex",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 376897,
      "author_name": "masato",
      "author_url": "",
      "post_date": "08/28/2018 08:48:52",
      "content": "<p>Hi, formigone.\nI have the same idea and made a short notebook. How about my approach?\n<a href=\"https://www.kaggle.com/masato/patient-data-in-dicom-file\">https://www.kaggle.com/masato/patient-data-in-dicom-file</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 377227,
          "author_name": "formigone",
          "author_url": "",
          "post_date": "08/28/2018 19:16:12",
          "content": "<p>That looks awesome. Great work.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 376903,
      "author_name": "nyleve",
      "author_url": "",
      "post_date": "08/28/2018 09:01:33",
      "content": "<p>Is this what you're looking for ?  (My first time using this file format as well)</p>\n\n<p>dcm_data[0x00100040].value </p>\n\n<p>--&gt; 'F'</p>",
      "votes": null,
      "replies": [
        {
          "id": 377226,
          "author_name": "formigone",
          "author_url": "",
          "post_date": "08/28/2018 19:14:06",
          "content": "<p>Thanks. I'll take a look. My fear is that <code>0x00100040</code> will not consistently map to that property. What you suggest is a good place to start, though. Thanks.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 377254,
          "author_name": "marckohli",
          "author_url": "",
          "post_date": "08/28/2018 20:35:31",
          "content": "<p>Because the images conform to the DICOM standard, that hex address always corresponds to Patient Sex. Here’s a list of other tags, although many of them have been modified or stripped in the anonimization process: <a href=\"https://sno.phy.queensu.ca/~phil/exiftool/TagNames/DICOM.html\">https://sno.phy.queensu.ca/~phil/exiftool/TagNames/DICOM.html</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 379844,
      "author_name": "navinprashath",
      "author_url": "",
      "post_date": "09/01/2018 03:23:24",
      "content": "<p>Hi, formigone. When viewed in DICOM image viewer there are two values C and W at bottom left of the image. Do you have any information what these values could be or be useful in making prediction?</p>",
      "votes": null,
      "replies": [
        {
          "id": 380118,
          "author_name": "marckohli",
          "author_url": "",
          "post_date": "09/01/2018 19:02:04",
          "content": "<p>W = Window\nC = Center</p>\n\n<p>This happens because the images contain more greyscale values than can be displayed at a single time.  So - the setting of a center and a window allows you to map the pixel values to greyscale values.</p>\n\n<p>Most DICOM images come with default values for W and C, but they are essentially suggestions based on the acquisition device.</p>\n\n<p>It's much more relevant for displaying CT images as described in this radiopedia article:</p>\n\n<p><a href=\"https://radiopaedia.org/articles/windowing-ct\">https://radiopaedia.org/articles/windowing-ct</a></p>\n\n<p>Cheers!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 380206,
      "author_name": "jtlowery",
      "author_url": "",
      "post_date": "09/02/2018 03:40:40",
      "content": "<p>Hey formigone, I have a notebook that has a method that (I think) reliably parses all the keywords and their values (as well as constructs a map from the DICOM tag to the keyword). <a href=\"https://www.kaggle.com/jtlowery/intro-eda-with-dicom-metadata/\">notebook - click the 'Parsing Metadata from DICOM Object' to skip right to it</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 381115,
      "author_name": "kanwalinder",
      "author_url": "",
      "post_date": "09/04/2018 04:36:07",
      "content": "<p>Hi formigone,</p>\n\n<p>You might find the following kernels and dataset useful on how to extract DICOM attributes:  <a href=\"https://www.kaggle.com/kanwalinder/rsna-preprocessed-nonimage-inputs\">RSNA Preprocessed Non-Image Inputs Dataset</a>, generated by <a href=\"https://www.kaggle.com/kanwalinder/preprocessing-rsna-non-image-inputs\">Preprocessing RSNA Non-Image Inputs</a>, and used by <a href=\"https://www.kaggle.com/kanwalinder/rsna-review-inputs\">RSNA Review Inputs</a>.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 381674,
      "author_name": "agudmundson",
      "author_url": "",
      "post_date": "09/05/2018 01:15:18",
      "content": "<p>Looks like this may have been answered in other ways. Just in case though, here's my simple approach:</p>\n\n<p>dcm_data = pydicom.read_file(dcm_file)</p>\n\n<p>age = dcm_data.PatientAge</p>\n\n<p>sex = dcm_data.PatientSex</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "376774": "I think it was this TWIML episode where the Googler said that he observed that a DNN can perform better if you give it more stuff to learn https://twimlai.com/twiml-talk-122-predicting-cardiovascular-risk-factors-eye-images-ryan-poplin/\n\nIn other words, instead of just training it to output a probability of pheumonia + bounding box, if you also train the network to output other factors from the image, such as age, gender, etc., the network can perform better at the core task you're after (in this case, probability of pneumonia with bounding boxes). I'm curious to experiment with that.\n\nMy question: I hadn't worked with DICOM files before. Using `pydicom`, it's trivial to parse the file with:\n\n    dcm_data = pydicom.read_file(dcm_file)\n\nBut from there, how do I access the individual lines in the meta description? For example, output of the above might look like:\n\n    print(dcm_data)\n    (0008, 0005) Specific Character Set              CS: 'ISO_IR 100'\n    (0008, 0020) Study Date                          DA: '19010101'\n    (0008, 0030) Study Time                          TM: '000000.00'\n    (0008, 0050) Accession Number                    SH: ''\n    (0008, 0060) Modality                            CS: 'CR'\n    (0008, 0064) Conversion Type                     CS: 'WSD'\n    (0008, 0090) Referring Physician's Name          PN: ''\n    (0008, 103e) Series Description                  LO: 'view: PA'\n    (0010, 0010) Patient's Name                      PN: '0004cfab-14fd'\n    (0010, 0020) Patient ID                          LO: '0004cfab-14fd'\n    (0010, 0030) Patient's Birth Date                DA: ''\n    (0010, 0040) Patient's Sex                       CS: 'F'\n    (0010, 1010) Patient's Age                       AS: '51'\n    (0018, 0015) Body Part Examined                  CS: 'CHEST'\n\nIs there a better way to get the value of `(0010, 0040) =&gt; Patient's Sex` from `dcm_data`? \n\nThanks.",
    "376897": "Hi, formigone.\nI have the same idea and made a short notebook. How about my approach?\nhttps://www.kaggle.com/masato/patient-data-in-dicom-file",
    "376903": "Is this what you're looking for ?  (My first time using this file format as well)\n\n dcm_data[0x00100040].value \n\n--&gt; 'F'",
    "377226": "Thanks. I'll take a look. My fear is that `0x00100040` will not consistently map to that property. What you suggest is a good place to start, though. Thanks.",
    "377227": "That looks awesome. Great work.",
    "377254": "Because the images conform to the DICOM standard, that hex address always corresponds to Patient Sex. Here’s a list of other tags, although many of them have been modified or stripped in the anonimization process: https://sno.phy.queensu.ca/~phil/exiftool/TagNames/DICOM.html",
    "379844": "Hi, formigone. When viewed in DICOM image viewer there are two values C and W at bottom left of the image. Do you have any information what these values could be or be useful in making prediction?",
    "380118": "W = Window\nC = Center\n\nThis happens because the images contain more greyscale values than can be displayed at a single time.  So - the setting of a center and a window allows you to map the pixel values to greyscale values.\n\nMost DICOM images come with default values for W and C, but they are essentially suggestions based on the acquisition device.\n\nIt's much more relevant for displaying CT images as described in this radiopedia article:\n\nhttps://radiopaedia.org/articles/windowing-ct\n\nCheers!",
    "380206": "Hey formigone, I have a notebook that has a method that (I think) reliably parses all the keywords and their values (as well as constructs a map from the DICOM tag to the keyword). [notebook - click the 'Parsing Metadata from DICOM Object' to skip right to it][1]\n\n\n  [1]: https://www.kaggle.com/jtlowery/intro-eda-with-dicom-metadata/",
    "381115": "Hi formigone,\n\nYou might find the following kernels and dataset useful on how to extract DICOM attributes:  [RSNA Preprocessed Non-Image Inputs Dataset][1], generated by [Preprocessing RSNA Non-Image Inputs][2], and used by [RSNA Review Inputs][3].\n\n\n  [1]: https://www.kaggle.com/kanwalinder/rsna-preprocessed-nonimage-inputs\n  [2]: https://www.kaggle.com/kanwalinder/preprocessing-rsna-non-image-inputs\n  [3]: https://www.kaggle.com/kanwalinder/rsna-review-inputs",
    "381674": "Looks like this may have been answered in other ways. Just in case though, here's my simple approach:\n\ndcm_data = pydicom.read_file(dcm_file)\n\nage = dcm_data.PatientAge\n\nsex = dcm_data.PatientSex"
  },
  "source": "meta"
}