{
  "id": 515702,
  "title": "Dataset preprocess, which is faster",
  "url": "/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/515702",
  "author_name": "",
  "post_date": "2024-06-29T13:13:40.042901600Z",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi guys, I'm reading directly from dicom files (pydicom &amp; pixel_array). I saw people convert dicom to png/jpg and store it as a seperate dataset. </p>\n<p>Is reading png/jpg faster than pydicom? </p>\n<p>Any help will be appreciated! :)</p>",
  "messages": [
    {
      "id": "2895929",
      "postDate": "06/29/2024 13:13:40",
      "content": "<p>Hi guys, I'm reading directly from dicom files (pydicom &amp; pixel_array). I saw people convert dicom to png/jpg and store it as a seperate dataset. </p>\n<p>Is reading png/jpg faster than pydicom? </p>\n<p>Any help will be appreciated! :)</p>",
      "rawMarkdown": "Hi guys, I'm reading directly from dicom files (pydicom & pixel_array). I saw people convert dicom to png/jpg and store it as a seperate dataset. \n\nIs reading png/jpg faster than pydicom? \n\nAny help will be appreciated! :)",
      "votes": null
    },
    {
      "id": "2895945",
      "postDate": "06/29/2024 13:33:29",
      "content": "<p>JPG compresses images too much as such, there is possibility of losing information. Therefore png is a good alternative, fast and yet doesn't lose much information </p>",
      "rawMarkdown": "JPG compresses images too much as such, there is possibility of losing information. Therefore png is a good alternative, fast and yet doesn't lose much information",
      "votes": null
    },
    {
      "id": "2896169",
      "postDate": "06/29/2024 16:25:07",
      "content": "<p>Incase of custom dataset loading is reading png/jpg faster than pydicom?? Please answer</p>",
      "rawMarkdown": "Incase of custom dataset loading is reading png/jpg faster than pydicom?? Please answer",
      "votes": null
    },
    {
      "id": "2896290",
      "postDate": "06/29/2024 18:18:39",
      "content": "<p>DICOM we are given are mostly RLE encoded and contain 16-bit integer data. While saving to PNG, we first normalize this and save (by default) to a bit depth of 8.</p>\n<p>So as there is less data stored, it is faster to store it in PNG. I would recommend however to experiment with NPY and NPZ formats (these are Numpy arrays saved with <code>np.save</code>, <code>np.savez</code> and <code>np.savez_compressed</code> functions) for PyTorch and the TFRecord format for TensorFlow (using tf.data module) if data reading is the bottleneck during your training.</p>",
      "rawMarkdown": "DICOM we are given are mostly RLE encoded and contain 16-bit integer data. While saving to PNG, we first normalize this and save (by default) to a bit depth of 8.\n\nSo as there is less data stored, it is faster to store it in PNG. I would recommend however to experiment with NPY and NPZ formats (these are Numpy arrays saved with `np.save`, `np.savez` and `np.savez_compressed` functions) for PyTorch and the TFRecord format for TensorFlow (using tf.data module) if data reading is the bottleneck during your training.",
      "votes": null
    },
    {
      "id": "2898238",
      "postDate": "07/01/2024 04:20:42",
      "content": "<p>I think jpeg is faster</p>",
      "rawMarkdown": "I think jpeg is faster",
      "votes": null
    },
    {
      "id": "2911807",
      "postDate": "07/08/2024 14:49:52",
      "content": "<p>Best way of saving the data is just save your pixel_array as a npy file. Load the entire dataset in a minute.</p>",
      "rawMarkdown": "Best way of saving the data is just save your pixel_array as a npy file. Load the entire dataset in a minute.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2895945,
      "author_name": "samu2505",
      "author_url": "",
      "post_date": "06/29/2024 13:33:29",
      "content": "<p>JPG compresses images too much as such, there is possibility of losing information. Therefore png is a good alternative, fast and yet doesn't lose much information </p>",
      "votes": null,
      "replies": [
        {
          "id": 2896169,
          "author_name": "devsya",
          "author_url": "",
          "post_date": "06/29/2024 16:25:07",
          "content": "<p>Incase of custom dataset loading is reading png/jpg faster than pydicom?? Please answer</p>",
          "votes": null,
          "replies": [
            {
              "id": 2898238,
              "author_name": "samu2505",
              "author_url": "",
              "post_date": "07/01/2024 04:20:42",
              "content": "<p>I think jpeg is faster</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2896290,
      "author_name": "coderrkj",
      "author_url": "",
      "post_date": "06/29/2024 18:18:39",
      "content": "<p>DICOM we are given are mostly RLE encoded and contain 16-bit integer data. While saving to PNG, we first normalize this and save (by default) to a bit depth of 8.</p>\n<p>So as there is less data stored, it is faster to store it in PNG. I would recommend however to experiment with NPY and NPZ formats (these are Numpy arrays saved with <code>np.save</code>, <code>np.savez</code> and <code>np.savez_compressed</code> functions) for PyTorch and the TFRecord format for TensorFlow (using tf.data module) if data reading is the bottleneck during your training.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2911807,
      "author_name": "llleeeoooh",
      "author_url": "",
      "post_date": "07/08/2024 14:49:52",
      "content": "<p>Best way of saving the data is just save your pixel_array as a npy file. Load the entire dataset in a minute.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2895929": "Hi guys, I'm reading directly from dicom files (pydicom & pixel_array). I saw people convert dicom to png/jpg and store it as a seperate dataset. \n\nIs reading png/jpg faster than pydicom? \n\nAny help will be appreciated! :)",
    "2895945": "JPG compresses images too much as such, there is possibility of losing information. Therefore png is a good alternative, fast and yet doesn't lose much information",
    "2896169": "Incase of custom dataset loading is reading png/jpg faster than pydicom?? Please answer",
    "2896290": "DICOM we are given are mostly RLE encoded and contain 16-bit integer data. While saving to PNG, we first normalize this and save (by default) to a bit depth of 8.\n\nSo as there is less data stored, it is faster to store it in PNG. I would recommend however to experiment with NPY and NPZ formats (these are Numpy arrays saved with `np.save`, `np.savez` and `np.savez_compressed` functions) for PyTorch and the TFRecord format for TensorFlow (using tf.data module) if data reading is the bottleneck during your training.",
    "2898238": "I think jpeg is faster",
    "2911807": "Best way of saving the data is just save your pixel_array as a npy file. Load the entire dataset in a minute."
  },
  "source": "meta"
}