{
  "id": 328602,
  "title": "What's the best practice to get numpy arrays of images?",
  "url": "/competitions/unifesp-x-ray-body-part-classifier/discussion/328602",
  "author_name": "",
  "post_date": "2022-06-02T03:23:23.527329800Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>The basic way to get pixel data as numpy arrays from dcm files I've learned is using \"pixel_array\" property of dcm dataset:</p>\n<p>Example:</p>\n<pre><code>import pydicom\nimport numpy as np\nimport time\n\nds = pydicom.dcmread('../input/unifesp-x-ray-body-part-classifier/train/train/1/1.2.826.0.1.3680043.8.498.89102450329340531816015855773961083133/1.2.826.0.1.3680043.8.498.11278653404499913987623237519434199794/1.2.826.0.1.3680043.8.498.65452424240994805812717428674475343109-c.dcm')\n\nt1 = time.time()\narray = ds.pixel_array\nprint(time.time() - t1)\n</code></pre>\n<pre><code>3.8031232357025146\n</code></pre>\n<p>This piece of code runs for around 3.8 seconds on Kaggle notebooks and some similar time on my laptop. There's over 1000 images in the training set which could take around an hour in loading the whole dataset. Am I on the correct track of doing this?</p>",
  "messages": [
    {
      "id": "1808612",
      "postDate": "06/02/2022 03:23:23",
      "content": "<p>The basic way to get pixel data as numpy arrays from dcm files I've learned is using \"pixel_array\" property of dcm dataset:</p>\n<p>Example:</p>\n<pre><code>import pydicom\nimport numpy as np\nimport time\n\nds = pydicom.dcmread('../input/unifesp-x-ray-body-part-classifier/train/train/1/1.2.826.0.1.3680043.8.498.89102450329340531816015855773961083133/1.2.826.0.1.3680043.8.498.11278653404499913987623237519434199794/1.2.826.0.1.3680043.8.498.65452424240994805812717428674475343109-c.dcm')\n\nt1 = time.time()\narray = ds.pixel_array\nprint(time.time() - t1)\n</code></pre>\n<pre><code>3.8031232357025146\n</code></pre>\n<p>This piece of code runs for around 3.8 seconds on Kaggle notebooks and some similar time on my laptop. There's over 1000 images in the training set which could take around an hour in loading the whole dataset. Am I on the correct track of doing this?</p>",
      "rawMarkdown": "The basic way to get pixel data as numpy arrays from dcm files I've learned is using \"pixel_array\" property of dcm dataset:\n\nExample:\n\n```python\nimport pydicom\nimport numpy as np\nimport time\n\nds = pydicom.dcmread('../input/unifesp-x-ray-body-part-classifier/train/train/1/1.2.826.0.1.3680043.8.498.89102450329340531816015855773961083133/1.2.826.0.1.3680043.8.498.11278653404499913987623237519434199794/1.2.826.0.1.3680043.8.498.65452424240994805812717428674475343109-c.dcm')\n\nt1 = time.time()\narray = ds.pixel_array\nprint(time.time() - t1)\n```\n\n```\n3.8031232357025146\n```\nThis piece of code runs for around 3.8 seconds on Kaggle notebooks and some similar time on my laptop. There's over 1000 images in the training set which could take around an hour in loading the whole dataset. Am I on the correct track of doing this?",
      "votes": null
    },
    {
      "id": "1809125",
      "postDate": "06/02/2022 12:13:47",
      "content": "<p>Reading DICOM files is slow, specially for big images like X-rays with compression. On SSDs it should be a bit faster. One work around is to read everything once and convert to numpy files with a reduced image size. Alternatively, you can use this dataset: <a href=\"https://www.kaggle.com/datasets/felipekitamura/unifesp-xray-bodypart-classification\" target=\"_blank\">https://www.kaggle.com/datasets/felipekitamura/unifesp-xray-bodypart-classification</a> <br>\nIt had image matrix reduced already. Same images as in the competition dataset.</p>",
      "rawMarkdown": "Reading DICOM files is slow, specially for big images like X-rays with compression. On SSDs it should be a bit faster. One work around is to read everything once and convert to numpy files with a reduced image size. Alternatively, you can use this dataset: https://www.kaggle.com/datasets/felipekitamura/unifesp-xray-bodypart-classification \nIt had image matrix reduced already. Same images as in the competition dataset.",
      "votes": null
    },
    {
      "id": "1809648",
      "postDate": "06/02/2022 22:43:44",
      "content": "<p>Thank you. This alternative dataset helps a lot.</p>",
      "rawMarkdown": "Thank you. This alternative dataset helps a lot.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1809125,
      "author_name": "felipekitamura",
      "author_url": "",
      "post_date": "06/02/2022 12:13:47",
      "content": "<p>Reading DICOM files is slow, specially for big images like X-rays with compression. On SSDs it should be a bit faster. One work around is to read everything once and convert to numpy files with a reduced image size. Alternatively, you can use this dataset: <a href=\"https://www.kaggle.com/datasets/felipekitamura/unifesp-xray-bodypart-classification\" target=\"_blank\">https://www.kaggle.com/datasets/felipekitamura/unifesp-xray-bodypart-classification</a> <br>\nIt had image matrix reduced already. Same images as in the competition dataset.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1809648,
          "author_name": "ritiger",
          "author_url": "",
          "post_date": "06/02/2022 22:43:44",
          "content": "<p>Thank you. This alternative dataset helps a lot.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1808612": "The basic way to get pixel data as numpy arrays from dcm files I've learned is using \"pixel_array\" property of dcm dataset:\n\nExample:\n\n```python\nimport pydicom\nimport numpy as np\nimport time\n\nds = pydicom.dcmread('../input/unifesp-x-ray-body-part-classifier/train/train/1/1.2.826.0.1.3680043.8.498.89102450329340531816015855773961083133/1.2.826.0.1.3680043.8.498.11278653404499913987623237519434199794/1.2.826.0.1.3680043.8.498.65452424240994805812717428674475343109-c.dcm')\n\nt1 = time.time()\narray = ds.pixel_array\nprint(time.time() - t1)\n```\n\n```\n3.8031232357025146\n```\nThis piece of code runs for around 3.8 seconds on Kaggle notebooks and some similar time on my laptop. There's over 1000 images in the training set which could take around an hour in loading the whole dataset. Am I on the correct track of doing this?",
    "1809125": "Reading DICOM files is slow, specially for big images like X-rays with compression. On SSDs it should be a bit faster. One work around is to read everything once and convert to numpy files with a reduced image size. Alternatively, you can use this dataset: https://www.kaggle.com/datasets/felipekitamura/unifesp-xray-bodypart-classification \nIt had image matrix reduced already. Same images as in the competition dataset.",
    "1809648": "Thank you. This alternative dataset helps a lot."
  },
  "source": "meta"
}