{
  "id": 255091,
  "title": "Understanding DICOM format How to read, write, and organize medical images",
  "url": "/competitions/siim-covid19-detection/discussion/255091",
  "author_name": "ACE",
  "post_date": "2021-07-25T15:56:30.145000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>DICOM is the primary file format for storing and transferring medical images in a hospital’s database.  Besides DICOM, you may also see medical images saved in the NIFTI format (file suffix “.nii”)</p>\n<p><strong><em><em>Challenges of DICOM</em></em></strong><br>\nFor deep learning tasks, the endgame is usually to load the image data as a NumPy or other file, and in this case, DICOM can be difficult:</p>\n<p>1.DICOM saves one file per slice, so a 3D scan may have hundreds of files</p>\n<p>2.DICOM files are named with a unique identifier (UID). This makes it hard to sort files from the folder level (in some cases, file names are so long that they exceed the 256 character maximum on Windows computers which causes saving/loading issues)</p>\n<ol>\n<li>Patient and hospital information is embedded in the file header, which can make DICOM tricky to anonymize</li>\n</ol>\n<p><strong><em><em>Identifying a DICOM file</em></em></strong><br>\nEach DICOM file is designed to be standalone — all the information needed to identify the file is embedded in each header. This information is organized into 4 levels of hierarchy — patient, study, series, and instance.</p>\n<ul>\n<li><p>“Patient” is the person receiving the exam</p></li>\n<li><p>“Study” is the imaging procedure being performed, at a certain date and time, in the hospital</p></li>\n</ul>\n<p>-“Series” — Each study consists of multiple series. A series may represent the patient being physically scanned multiple times in one study (typical for MRI), or it may be virtual, where the patient is scanned once and that data is reconstructed in different ways (typical for CT)</p>\n<p>-“Instance” — every slice of a 3D image is treated as a separate instance. In this context, “instance” is synonymous with the DICOM file itself</p>\n<p><strong><em><em>Unique Identifiers: UIDs</em></em></strong><br>\nIn addition to the text descriptions, the scan is identified by the unique Patient ID (5553226), Study UID (1.2.826.0.1.3680043.2.1125.1. 38381854871216336385978062044218957), Series UID (1.2.826.0.1. 3680043.2.1125.1.68878959984837726447916707551399667), and Instance Number (20).</p>\n<p>If you were to load the very next DICOM file in this folder, the Patient ID, Study UID, and Series UID would all have the same value, and only the Instance Number would be different</p>",
  "messages": [
    {
      "id": 1399722,
      "postDate": "2021-07-25T15:56:30.147Z",
      "content": "<p>DICOM is the primary file format for storing and transferring medical images in a hospital’s database.  Besides DICOM, you may also see medical images saved in the NIFTI format (file suffix “.nii”)</p>\n<p><strong><em><em>Challenges of DICOM</em></em></strong><br>\nFor deep learning tasks, the endgame is usually to load the image data as a NumPy or other file, and in this case, DICOM can be difficult:</p>\n<p>1.DICOM saves one file per slice, so a 3D scan may have hundreds of files</p>\n<p>2.DICOM files are named with a unique identifier (UID). This makes it hard to sort files from the folder level (in some cases, file names are so long that they exceed the 256 character maximum on Windows computers which causes saving/loading issues)</p>\n<ol>\n<li>Patient and hospital information is embedded in the file header, which can make DICOM tricky to anonymize</li>\n</ol>\n<p><strong><em><em>Identifying a DICOM file</em></em></strong><br>\nEach DICOM file is designed to be standalone — all the information needed to identify the file is embedded in each header. This information is organized into 4 levels of hierarchy — patient, study, series, and instance.</p>\n<ul>\n<li><p>“Patient” is the person receiving the exam</p></li>\n<li><p>“Study” is the imaging procedure being performed, at a certain date and time, in the hospital</p></li>\n</ul>\n<p>-“Series” — Each study consists of multiple series. A series may represent the patient being physically scanned multiple times in one study (typical for MRI), or it may be virtual, where the patient is scanned once and that data is reconstructed in different ways (typical for CT)</p>\n<p>-“Instance” — every slice of a 3D image is treated as a separate instance. In this context, “instance” is synonymous with the DICOM file itself</p>\n<p><strong><em><em>Unique Identifiers: UIDs</em></em></strong><br>\nIn addition to the text descriptions, the scan is identified by the unique Patient ID (5553226), Study UID (1.2.826.0.1.3680043.2.1125.1. 38381854871216336385978062044218957), Series UID (1.2.826.0.1. 3680043.2.1125.1.68878959984837726447916707551399667), and Instance Number (20).</p>\n<p>If you were to load the very next DICOM file in this folder, the Patient ID, Study UID, and Series UID would all have the same value, and only the Instance Number would be different</p>",
      "rawMarkdown": "DICOM is the primary file format for storing and transferring medical images in a hospital’s database.  Besides DICOM, you may also see medical images saved in the NIFTI format (file suffix “.nii”)\n\n ****Challenges of DICOM****\nFor deep learning tasks, the endgame is usually to load the image data as a NumPy or other file, and in this case, DICOM can be difficult:\n\n1.DICOM saves one file per slice, so a 3D scan may have hundreds of files\n\n2.DICOM files are named with a unique identifier (UID). This makes it hard to sort files from the folder level (in some cases, file names are so long that they exceed the 256 character maximum on Windows computers which causes saving/loading issues)\n\n3. Patient and hospital information is embedded in the file header, which can make DICOM tricky to anonymize\n\n****Identifying a DICOM file****\nEach DICOM file is designed to be standalone — all the information needed to identify the file is embedded in each header. This information is organized into 4 levels of hierarchy — patient, study, series, and instance.\n\n- “Patient” is the person receiving the exam\n\n- “Study” is the imaging procedure being performed, at a certain date and time, in the hospital\n\n-“Series” — Each study consists of multiple series. A series may represent the patient being physically scanned multiple times in one study (typical for MRI), or it may be virtual, where the patient is scanned once and that data is reconstructed in different ways (typical for CT)\n\n-“Instance” — every slice of a 3D image is treated as a separate instance. In this context, “instance” is synonymous with the DICOM file itself\n\n****Unique Identifiers: UIDs****\nIn addition to the text descriptions, the scan is identified by the unique Patient ID (5553226), Study UID (1.2.826.0.1.3680043.2.1125.1. 38381854871216336385978062044218957), Series UID (1.2.826.0.1. 3680043.2.1125.1.68878959984837726447916707551399667), and Instance Number (20).\n\nIf you were to load the very next DICOM file in this folder, the Patient ID, Study UID, and Series UID would all have the same value, and only the Instance Number would be different",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1399722": "DICOM is the primary file format for storing and transferring medical images in a hospital’s database.  Besides DICOM, you may also see medical images saved in the NIFTI format (file suffix “.nii”)\n\n ****Challenges of DICOM****\nFor deep learning tasks, the endgame is usually to load the image data as a NumPy or other file, and in this case, DICOM can be difficult:\n\n1.DICOM saves one file per slice, so a 3D scan may have hundreds of files\n\n2.DICOM files are named with a unique identifier (UID). This makes it hard to sort files from the folder level (in some cases, file names are so long that they exceed the 256 character maximum on Windows computers which causes saving/loading issues)\n\n3. Patient and hospital information is embedded in the file header, which can make DICOM tricky to anonymize\n\n****Identifying a DICOM file****\nEach DICOM file is designed to be standalone — all the information needed to identify the file is embedded in each header. This information is organized into 4 levels of hierarchy — patient, study, series, and instance.\n\n- “Patient” is the person receiving the exam\n\n- “Study” is the imaging procedure being performed, at a certain date and time, in the hospital\n\n-“Series” — Each study consists of multiple series. A series may represent the patient being physically scanned multiple times in one study (typical for MRI), or it may be virtual, where the patient is scanned once and that data is reconstructed in different ways (typical for CT)\n\n-“Instance” — every slice of a 3D image is treated as a separate instance. In this context, “instance” is synonymous with the DICOM file itself\n\n****Unique Identifiers: UIDs****\nIn addition to the text descriptions, the scan is identified by the unique Patient ID (5553226), Study UID (1.2.826.0.1.3680043.2.1125.1. 38381854871216336385978062044218957), Series UID (1.2.826.0.1. 3680043.2.1125.1.68878959984837726447916707551399667), and Instance Number (20).\n\nIf you were to load the very next DICOM file in this folder, the Patient ID, Study UID, and Series UID would all have the same value, and only the Instance Number would be different"
  }
}