{
  "id": 266679,
  "title": "Task 1 Dataset Uploaded",
  "url": "/competitions/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/266679",
  "author_name": "",
  "post_date": "2021-08-20T00:56:42.884540Z",
  "votes": 41,
  "comment_count": 21,
  "views": 0,
  "content": "<p>Hi there.</p>\n<p>I uploaded the task one dataset (<a href=\"https://www.kaggle.com/dschettler8845/brats-2021-task1\" target=\"_blank\"><strong>here</strong></a>) and made a basic notebook (<a href=\"https://www.kaggle.com/dschettler8845/how-to-load-basic-data-exploration\" target=\"_blank\"><strong>here</strong></a>) showing how to <strong><code>untar</code></strong> and use a basic library to convert to NumPy.</p>\n<p>Please ensure that you register for Task 1 via the Synapse website prior to using the dataset. Also please ensure you read all of the rules and whatnot before using the data.<br>\n  -&gt; <a href=\"https://www.synapse.org/#!Synapse:syn25829067/wiki/\" target=\"_blank\"><strong>SYNAPSE LINK HERE</strong></a></p>\n<p>I believe having this data public on the Kaggle platform is allowed… if not please let me know and I can take it down.</p>\n<hr>\n<p>NOTE: Please see <a href=\"https://www.kaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/253488\" target=\"_blank\"><strong>this post</strong></a> on notes/concerns regarding using this as external data.</p>\n<hr>\n<p>I hope this helps!</p>",
  "messages": [
    {
      "id": "1482233",
      "postDate": "08/20/2021 00:56:42",
      "content": "<p>Hi there.</p>\n<p>I uploaded the task one dataset (<a href=\"https://www.kaggle.com/dschettler8845/brats-2021-task1\" target=\"_blank\"><strong>here</strong></a>) and made a basic notebook (<a href=\"https://www.kaggle.com/dschettler8845/how-to-load-basic-data-exploration\" target=\"_blank\"><strong>here</strong></a>) showing how to <strong><code>untar</code></strong> and use a basic library to convert to NumPy.</p>\n<p>Please ensure that you register for Task 1 via the Synapse website prior to using the dataset. Also please ensure you read all of the rules and whatnot before using the data.<br>\n  -&gt; <a href=\"https://www.synapse.org/#!Synapse:syn25829067/wiki/\" target=\"_blank\"><strong>SYNAPSE LINK HERE</strong></a></p>\n<p>I believe having this data public on the Kaggle platform is allowed… if not please let me know and I can take it down.</p>\n<hr>\n<p>NOTE: Please see <a href=\"https://www.kaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/253488\" target=\"_blank\"><strong>this post</strong></a> on notes/concerns regarding using this as external data.</p>\n<hr>\n<p>I hope this helps!</p>",
      "rawMarkdown": "Hi there.\n\nI uploaded the task one dataset ([**here**](https://www.kaggle.com/dschettler8845/brats-2021-task1)) and made a basic notebook ([**here**](https://www.kaggle.com/dschettler8845/how-to-load-basic-data-exploration)) showing how to **`untar`** and use a basic library to convert to NumPy.\n\nPlease ensure that you register for Task 1 via the Synapse website prior to using the dataset. Also please ensure you read all of the rules and whatnot before using the data.\n  -> [**SYNAPSE LINK HERE**](https://www.synapse.org/#!Synapse:syn25829067/wiki/)\n\nI believe having this data public on the Kaggle platform is allowed... if not please let me know and I can take it down.\n\n---\n\nNOTE: Please see [**this post**](https://www.kaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/253488) on notes/concerns regarding using this as external data.\n\n---\n\nI hope this helps!",
      "votes": null
    },
    {
      "id": "1482291",
      "postDate": "08/20/2021 02:20:46",
      "content": "<p>Thank you. I wonder if we can resample task 2's data into the task 1's by ourselves. I read from the paper they contains several steps such as: convert to RAI system, co-register on SRI24, resample spacing to 1mm x 1mm x 1mm. The host used the CaPTk tools to process their task1 data. But we cannot install such software on Kaggle. </p>",
      "rawMarkdown": "Thank you. I wonder if we can resample task 2's data into the task 1's by ourselves. I read from the paper they contains several steps such as: convert to RAI system, co-register on SRI24, resample spacing to 1mm x 1mm x 1mm. The host used the CaPTk tools to process their task1 data. But we cannot install such software on Kaggle.",
      "votes": null
    },
    {
      "id": "1482402",
      "postDate": "08/20/2021 04:30:11",
      "content": "<p>your notebook link is broken</p>",
      "rawMarkdown": "your notebook link is broken",
      "votes": null
    },
    {
      "id": "1482993",
      "postDate": "08/20/2021 11:41:47",
      "content": "<p>Ah sorry. It was still private. Should be fixed now.</p>",
      "rawMarkdown": "Ah sorry. It was still private. Should be fixed now.",
      "votes": null
    },
    {
      "id": "1483237",
      "postDate": "08/20/2021 14:21:09",
      "content": "<p>in the data description of brats2021 its written <br>\n\"pathologically confirmed diagnosis and available MGMT promoter methylation status, are used as the training, validation, and testing data for this year’s BraTS challenge.\"<br>\nwhat does it mean, any idea?</p>",
      "rawMarkdown": "in the data description of brats2021 its written \n\"pathologically confirmed diagnosis and available MGMT promoter methylation status, are used as the training, validation, and testing data for this year’s BraTS challenge.\"\nwhat does it mean, any idea?",
      "votes": null
    },
    {
      "id": "1483456",
      "postDate": "08/20/2021 16:50:03",
      "content": "<p>That's what I will be doing… or I will look for outside tools to accomplish the same that can be accessed within Kaggle.</p>",
      "rawMarkdown": "That's what I will be doing... or I will look for outside tools to accomplish the same that can be accessed within Kaggle.",
      "votes": null
    },
    {
      "id": "1485654",
      "postDate": "08/22/2021 10:24:13",
      "content": "<p>Thanks. how many images exist in this task1 dataset?<br>\nall of this images contain tumor. isn't it?<br>\nIt is useful for our model to use it or not?</p>",
      "rawMarkdown": "Thanks. how many images exist in this task1 dataset?\nall of this images contain tumor. isn't it?\nIt is useful for our model to use it or not?",
      "votes": null
    },
    {
      "id": "1485656",
      "postDate": "08/22/2021 10:26:52",
      "content": "<p>please share this paper that you said . <br>\nwhat is CaPTK tools and how we can use it ? any notebook is available ?</p>",
      "rawMarkdown": "please share this paper that you said . \nwhat is CaPTK tools and how we can use it ? any notebook is available ?",
      "votes": null
    },
    {
      "id": "1486126",
      "postDate": "08/22/2021 17:12:28",
      "content": "<p><a href=\"https://arxiv.org/abs/2107.02314#:~:text=The%20RSNA%2DASNR%2DMICCAI%20BraTS%202021%20challenge%20targets%20the%20evaluation,mpMRI%20data%20from%202%2C000%20patients\" target=\"_blank\">https://arxiv.org/abs/2107.02314#:~:text=The%20RSNA%2DASNR%2DMICCAI%20BraTS%202021%20challenge%20targets%20the%20evaluation,mpMRI%20data%20from%202%2C000%20patients</a></p>",
      "rawMarkdown": "https://arxiv.org/abs/2107.02314#:~:text=The%20RSNA%2DASNR%2DMICCAI%20BraTS%202021%20challenge%20targets%20the%20evaluation,mpMRI%20data%20from%202%2C000%20patients",
      "votes": null
    },
    {
      "id": "1490312",
      "postDate": "08/25/2021 15:21:48",
      "content": "<p>Hi Darien,</p>\n<p>Thanks a lot for uploading the dataset for Task1.</p>\n<p>While exploring the data, I saw that the #Channels for all the modalities across patients is constant i.e. 155. But in the dataset we have on kaggle (Task 2) we have variable set of dicom images.</p>\n<p>Is this inconsistency due to the pre processing steps they did while preparing the task1 data?</p>\n<p>How should we tackle this while building a pipeline which contains Segmentation -&gt; Classification?</p>\n<p>Thanks</p>",
      "rawMarkdown": "Hi Darien,\n\nThanks a lot for uploading the dataset for Task1.\n\nWhile exploring the data, I saw that the #Channels for all the modalities across patients is constant i.e. 155. But in the dataset we have on kaggle (Task 2) we have variable set of dicom images.\n\nIs this inconsistency due to the pre processing steps they did while preparing the task1 data?\n\nHow should we tackle this while building a pipeline which contains Segmentation -> Classification?\n\nThanks",
      "votes": null
    },
    {
      "id": "1490800",
      "postDate": "08/25/2021 21:41:37",
      "content": "<p>Is it safe to assume that the subject IDs represent the same patients across Task 1 and Task 2 data? So, if I were to look at Task 1's 00000 subject the NIFTI files will have all characteristics in it for it to be designated as MGMT_1?</p>",
      "rawMarkdown": "Is it safe to assume that the subject IDs represent the same patients across Task 1 and Task 2 data? So, if I were to look at Task 1's 00000 subject the NIFTI files will have all characteristics in it for it to be designated as MGMT_1?",
      "votes": null
    },
    {
      "id": "1490928",
      "postDate": "08/26/2021 02:28:30",
      "content": "<p>If it is, it will be a huge data leakage.</p>",
      "rawMarkdown": "If it is, it will be a huge data leakage.",
      "votes": null
    },
    {
      "id": "1490942",
      "postDate": "08/26/2021 03:05:30",
      "content": "<p>hmm… why do you say that? Task1's data is only going to give us tumor segmentations and not any information about presence of methylation</p>",
      "rawMarkdown": "hmm... why do you say that? Task1's data is only going to give us tumor segmentations and not any information about presence of methylation",
      "votes": null
    },
    {
      "id": "1491332",
      "postDate": "08/26/2021 10:02:40",
      "content": "<p>Thanks. I have two question about task1 dataset.How many pictures are there in this dataset and are those different whit this kaggle dataset?<br>\nIn this dataset exist pictures with tumor and without tumor or not ?and how many different pictures are there exist?</p>",
      "rawMarkdown": "Thanks. I have two question about task1 dataset.How many pictures are there in this dataset and are those different whit this kaggle dataset?\nIn this dataset exist pictures with tumor and without tumor or not ?and how many different pictures are there exist?",
      "votes": null
    },
    {
      "id": "1491482",
      "postDate": "08/26/2021 12:23:56",
      "content": "<p>Yes. From what I can see the BraTSIDs overlap between the two training datasets with only a fraction of the IDs present in Task 2 not being present in Task 1.</p>",
      "rawMarkdown": "Yes. From what I can see the BraTSIDs overlap between the two training datasets with only a fraction of the IDs present in Task 2 not being present in Task 1.",
      "votes": null
    },
    {
      "id": "1491483",
      "postDate": "08/26/2021 12:24:53",
      "content": "<p>The images are preprocessed differently in Task 1 v. Task 2.</p>\n<p>To use the Task 1 images effectively, you may need to convert them so that they are similarly processed/aligned/registered.</p>",
      "rawMarkdown": "The images are preprocessed differently in Task 1 v. Task 2.\n\nTo use the Task 1 images effectively, you may need to convert them so that they are similarly processed/aligned/registered.",
      "votes": null
    },
    {
      "id": "1491485",
      "postDate": "08/26/2021 12:27:57",
      "content": "<p><em>how many images exist in this task1 dataset?</em></p>\n<p>I believe there are 1262 patients/BraTSIDs in the Task 1 dataset. With each patient/ID containing 155 channels. This could be interpreted as 195610 slices or images if you like.</p>\n<p><br></p>\n<p><em>all of this images contain tumor. isn't it?</em></p>\n<p>All of the patient's MRI scans contain tumours. However, not all slices/channels within a patient will have the tumour present.</p>\n<p><br></p>\n<p><em>It is useful for our model to use it or not?</em></p>\n<p>That's up to you to decide.</p>",
      "rawMarkdown": "*how many images exist in this task1 dataset?*\n\nI believe there are 1262 patients/BraTSIDs in the Task 1 dataset. With each patient/ID containing 155 channels. This could be interpreted as 195610 slices or images if you like.\n\n<br>\n\n*all of this images contain tumor. isn't it?*\n\nAll of the patient's MRI scans contain tumours. However, not all slices/channels within a patient will have the tumour present.\n\n<br>\n\n*It is useful for our model to use it or not?*\n\nThat's up to you to decide.",
      "votes": null
    },
    {
      "id": "1491491",
      "postDate": "08/26/2021 12:32:27",
      "content": "<p>A few quick definitions first…</p>\n<hr>\n<p>The definition of <strong>\"pathologically confirmed diagnosis\"</strong> is the following:</p>\n<blockquote>\n  <p>Determination of the cause or causes of an illness by examining fluids and tissues from the patient before or after death. The examination may be performed on blood, plasma, microscopic tissue samples, or gross specimens.</p>\n</blockquote>\n<p>And MGMT promoter methylation status is determined usually using the following techniques:</p>\n<blockquote>\n  <p>Quantitative PCR, methylation-specific PCR, pyrosequencing, real time PCR with high resolution melt, or the infinitum methylation EPIC beadChip</p>\n</blockquote>\n<hr>\n<p><strong>To answer your question, I think they are just saying that they determined the labels (tumour segmentation maps and MGMT promoter methylation status) based on a combination of autopsy and genetic investigation techniques.</strong></p>",
      "rawMarkdown": "A few quick definitions first...\n\n---\n\nThe definition of **\"pathologically confirmed diagnosis\"** is the following:\n\n> Determination of the cause or causes of an illness by examining fluids and tissues from the patient before or after death. The examination may be performed on blood, plasma, microscopic tissue samples, or gross specimens.\n\nAnd MGMT promoter methylation status is determined usually using the following techniques:\n\n> Quantitative PCR, methylation-specific PCR, pyrosequencing, real time PCR with high resolution melt, or the infinitum methylation EPIC beadChip\n\n---\n\n**To answer your question, I think they are just saying that they determined the labels (tumour segmentation maps and MGMT promoter methylation status) based on a combination of autopsy and genetic investigation techniques.**",
      "votes": null
    },
    {
      "id": "1491493",
      "postDate": "08/26/2021 12:32:56",
      "content": "<p>Answered in the above comment.</p>",
      "rawMarkdown": "Answered in the above comment.",
      "votes": null
    },
    {
      "id": "1491550",
      "postDate": "08/26/2021 13:04:23",
      "content": "<p>But there is no MGMT status for the data in that task right</p>",
      "rawMarkdown": "But there is no MGMT status for the data in that task right",
      "votes": null
    },
    {
      "id": "1491585",
      "postDate": "08/26/2021 13:21:44",
      "content": "<p>You are correct. Technically no MGMT status is given for Task 1.</p>\n<p>However, both Task 1 and Task 2 use BraTSID numbers indicating a particular patient. The patient is the same in both datasets and therefore the MGMT status from Task 2 can be assumed for Task 1… and similarly, the segmentation maps (if aligned properly) can be considered for Task 2.</p>\n<p>If you read through the <a href=\"https://arxiv.org/pdf/2107.02314.pdf\" target=\"_blank\">paper authored by the competition hosts</a>, you can see that the data goes through a pipeline roughly like (p. 5 and 6).</p>\n<p><br></p>\n<p><strong>TO OBTAIN THE TASK 1 IMAGES</strong><br>\n<strong>Standardized pre-processing has been applied to all the ORIGINAL BraTS mpMRI scans</strong> </p>\n<ul>\n<li>Specifically, the applied pre-processing routines include conversion of the DICOM files to the NIFTI file format</li>\n<li>re-orientation to a common orientation system (i.e., RAI)</li>\n<li>co-registration to the same anatomical template (SRI24)</li>\n<li>resampling to a uniform isotropic resolution (1mm3)</li>\n<li>skull-stripping</li>\n</ul>\n<p><strong><em>NOTE:</em></strong> <em>The preprocessing pipeline is publicly available through the Cancer Imaging Phenomics Toolkit (CaPTk) and Federated Tumor Segmentation (FeTS) tool 1</em></p>\n<p><br></p>\n<p><strong>TO OBTAIN THE TASK 2 IMAGES</strong><br>\n<strong>All the imaging volumes were converted from NIFTI to DICOM files, while ensuring that the original patient space is preserved</strong></p>\n<ul>\n<li>To make this conversion both the skull-stripped brain volume in NIFTI format of each MRI sequence and its corresponding original DICOM scan in the patient space are required</li>\n<li>The DICOM volume is read as an ITK image and the skull-stripped volume is rigidly registered to it, providing a transformation matrix that defines the spatial mapping between the 2 volumes.<ul>\n<li>This transformation matrix is applied to the skull-stripped volume and to the corresponding segmentation labels, in order to translate them both to the patient space. </li></ul></li>\n<li>These transformed volumes are then passed through CaPTk’s NIFTI to DICOM conversion engine to generate DICOM image volumes for the skullstripped image. </li>\n<li>Once all MRI sequences were converted back to the DICOM file format, further de-identification took place based on a two-step process. <ul>\n<li>The first step used the RSNA CTP (Clinical Trials Processor) Anonymizer 2 with the standard built-in script. </li>\n<li>Step two then consisted of whitelisting the DICOM files from step 1. </li>\n<li>The whitelisting process removes all non-essential tags from the DICOM header. </li>\n<li>This last process ensures there are no protected health information (PHI) entries left in the DICOM header.</li></ul></li>\n</ul>\n<hr>\n<p>I hope this clears things up!</p>",
      "rawMarkdown": "You are correct. Technically no MGMT status is given for Task 1.\n\nHowever, both Task 1 and Task 2 use BraTSID numbers indicating a particular patient. The patient is the same in both datasets and therefore the MGMT status from Task 2 can be assumed for Task 1... and similarly, the segmentation maps (if aligned properly) can be considered for Task 2.\n\nIf you read through the [paper authored by the competition hosts](https://arxiv.org/pdf/2107.02314.pdf), you can see that the data goes through a pipeline roughly like (p. 5 and 6).\n\n<br>\n\n**TO OBTAIN THE TASK 1 IMAGES**\n**Standardized pre-processing has been applied to all the ORIGINAL BraTS mpMRI scans** \n  - Specifically, the applied pre-processing routines include conversion of the DICOM files to the NIFTI file format\n  - re-orientation to a common orientation system (i.e., RAI)\n  - co-registration to the same anatomical template (SRI24)\n  - resampling to a uniform isotropic resolution (1mm3)\n  - skull-stripping\n\n***NOTE:*** *The preprocessing pipeline is publicly available through the Cancer Imaging Phenomics Toolkit (CaPTk) and Federated Tumor Segmentation (FeTS) tool 1*\n\n<br>\n\n**TO OBTAIN THE TASK 2 IMAGES**\n**All the imaging volumes were converted from NIFTI to DICOM files, while ensuring that the original patient space is preserved**\n* To make this conversion both the skull-stripped brain volume in NIFTI format of each MRI sequence and its corresponding original DICOM scan in the patient space are required\n* The DICOM volume is read as an ITK image and the skull-stripped volume is rigidly registered to it, providing a transformation matrix that defines the spatial mapping between the 2 volumes.\n  - This transformation matrix is applied to the skull-stripped volume and to the corresponding segmentation labels, in order to translate them both to the patient space. \n* These transformed volumes are then passed through CaPTk’s NIFTI to DICOM conversion engine to generate DICOM image volumes for the skullstripped image. \n* Once all MRI sequences were converted back to the DICOM file format, further de-identification took place based on a two-step process. \n  - The first step used the RSNA CTP (Clinical Trials Processor) Anonymizer 2 with the standard built-in script. \n  - Step two then consisted of whitelisting the DICOM files from step 1. \n  - The whitelisting process removes all non-essential tags from the DICOM header. \n  - This last process ensures there are no protected health information (PHI) entries left in the DICOM header.\n\n---\n\nI hope this clears things up!",
      "votes": null
    },
    {
      "id": "1496239",
      "postDate": "08/30/2021 08:08:56",
      "content": "<p>Thank you Darien for your hard work!! Can I ask how to convert this nii.gz to dcm? I tried to find it but I couldn't convert them… and if so is it going to be similar to the data that is from task2 data?</p>",
      "rawMarkdown": "Thank you Darien for your hard work!! Can I ask how to convert this nii.gz to dcm? I tried to find it but I couldn't convert them... and if so is it going to be similar to the data that is from task2 data?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1482291,
      "author_name": "namgalielei",
      "author_url": "",
      "post_date": "08/20/2021 02:20:46",
      "content": "<p>Thank you. I wonder if we can resample task 2's data into the task 1's by ourselves. I read from the paper they contains several steps such as: convert to RAI system, co-register on SRI24, resample spacing to 1mm x 1mm x 1mm. The host used the CaPTk tools to process their task1 data. But we cannot install such software on Kaggle. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1483456,
          "author_name": "dschettler8845",
          "author_url": "",
          "post_date": "08/20/2021 16:50:03",
          "content": "<p>That's what I will be doing… or I will look for outside tools to accomplish the same that can be accessed within Kaggle.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1485656,
          "author_name": "mohammadhosein1998",
          "author_url": "",
          "post_date": "08/22/2021 10:26:52",
          "content": "<p>please share this paper that you said . <br>\nwhat is CaPTK tools and how we can use it ? any notebook is available ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1486126,
          "author_name": "dschettler8845",
          "author_url": "",
          "post_date": "08/22/2021 17:12:28",
          "content": "<p><a href=\"https://arxiv.org/abs/2107.02314#:~:text=The%20RSNA%2DASNR%2DMICCAI%20BraTS%202021%20challenge%20targets%20the%20evaluation,mpMRI%20data%20from%202%2C000%20patients\" target=\"_blank\">https://arxiv.org/abs/2107.02314#:~:text=The%20RSNA%2DASNR%2DMICCAI%20BraTS%202021%20challenge%20targets%20the%20evaluation,mpMRI%20data%20from%202%2C000%20patients</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1491332,
          "author_name": "mohammadhosein1998",
          "author_url": "",
          "post_date": "08/26/2021 10:02:40",
          "content": "<p>Thanks. I have two question about task1 dataset.How many pictures are there in this dataset and are those different whit this kaggle dataset?<br>\nIn this dataset exist pictures with tumor and without tumor or not ?and how many different pictures are there exist?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1491493,
          "author_name": "dschettler8845",
          "author_url": "",
          "post_date": "08/26/2021 12:32:56",
          "content": "<p>Answered in the above comment.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1482402,
      "author_name": "mrinath",
      "author_url": "",
      "post_date": "08/20/2021 04:30:11",
      "content": "<p>your notebook link is broken</p>",
      "votes": null,
      "replies": [
        {
          "id": 1482993,
          "author_name": "dschettler8845",
          "author_url": "",
          "post_date": "08/20/2021 11:41:47",
          "content": "<p>Ah sorry. It was still private. Should be fixed now.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1483237,
      "author_name": "mrinath",
      "author_url": "",
      "post_date": "08/20/2021 14:21:09",
      "content": "<p>in the data description of brats2021 its written <br>\n\"pathologically confirmed diagnosis and available MGMT promoter methylation status, are used as the training, validation, and testing data for this year’s BraTS challenge.\"<br>\nwhat does it mean, any idea?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1491491,
          "author_name": "dschettler8845",
          "author_url": "",
          "post_date": "08/26/2021 12:32:27",
          "content": "<p>A few quick definitions first…</p>\n<hr>\n<p>The definition of <strong>\"pathologically confirmed diagnosis\"</strong> is the following:</p>\n<blockquote>\n  <p>Determination of the cause or causes of an illness by examining fluids and tissues from the patient before or after death. The examination may be performed on blood, plasma, microscopic tissue samples, or gross specimens.</p>\n</blockquote>\n<p>And MGMT promoter methylation status is determined usually using the following techniques:</p>\n<blockquote>\n  <p>Quantitative PCR, methylation-specific PCR, pyrosequencing, real time PCR with high resolution melt, or the infinitum methylation EPIC beadChip</p>\n</blockquote>\n<hr>\n<p><strong>To answer your question, I think they are just saying that they determined the labels (tumour segmentation maps and MGMT promoter methylation status) based on a combination of autopsy and genetic investigation techniques.</strong></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1491550,
          "author_name": "mrinath",
          "author_url": "",
          "post_date": "08/26/2021 13:04:23",
          "content": "<p>But there is no MGMT status for the data in that task right</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1491585,
          "author_name": "dschettler8845",
          "author_url": "",
          "post_date": "08/26/2021 13:21:44",
          "content": "<p>You are correct. Technically no MGMT status is given for Task 1.</p>\n<p>However, both Task 1 and Task 2 use BraTSID numbers indicating a particular patient. The patient is the same in both datasets and therefore the MGMT status from Task 2 can be assumed for Task 1… and similarly, the segmentation maps (if aligned properly) can be considered for Task 2.</p>\n<p>If you read through the <a href=\"https://arxiv.org/pdf/2107.02314.pdf\" target=\"_blank\">paper authored by the competition hosts</a>, you can see that the data goes through a pipeline roughly like (p. 5 and 6).</p>\n<p><br></p>\n<p><strong>TO OBTAIN THE TASK 1 IMAGES</strong><br>\n<strong>Standardized pre-processing has been applied to all the ORIGINAL BraTS mpMRI scans</strong> </p>\n<ul>\n<li>Specifically, the applied pre-processing routines include conversion of the DICOM files to the NIFTI file format</li>\n<li>re-orientation to a common orientation system (i.e., RAI)</li>\n<li>co-registration to the same anatomical template (SRI24)</li>\n<li>resampling to a uniform isotropic resolution (1mm3)</li>\n<li>skull-stripping</li>\n</ul>\n<p><strong><em>NOTE:</em></strong> <em>The preprocessing pipeline is publicly available through the Cancer Imaging Phenomics Toolkit (CaPTk) and Federated Tumor Segmentation (FeTS) tool 1</em></p>\n<p><br></p>\n<p><strong>TO OBTAIN THE TASK 2 IMAGES</strong><br>\n<strong>All the imaging volumes were converted from NIFTI to DICOM files, while ensuring that the original patient space is preserved</strong></p>\n<ul>\n<li>To make this conversion both the skull-stripped brain volume in NIFTI format of each MRI sequence and its corresponding original DICOM scan in the patient space are required</li>\n<li>The DICOM volume is read as an ITK image and the skull-stripped volume is rigidly registered to it, providing a transformation matrix that defines the spatial mapping between the 2 volumes.<ul>\n<li>This transformation matrix is applied to the skull-stripped volume and to the corresponding segmentation labels, in order to translate them both to the patient space. </li></ul></li>\n<li>These transformed volumes are then passed through CaPTk’s NIFTI to DICOM conversion engine to generate DICOM image volumes for the skullstripped image. </li>\n<li>Once all MRI sequences were converted back to the DICOM file format, further de-identification took place based on a two-step process. <ul>\n<li>The first step used the RSNA CTP (Clinical Trials Processor) Anonymizer 2 with the standard built-in script. </li>\n<li>Step two then consisted of whitelisting the DICOM files from step 1. </li>\n<li>The whitelisting process removes all non-essential tags from the DICOM header. </li>\n<li>This last process ensures there are no protected health information (PHI) entries left in the DICOM header.</li></ul></li>\n</ul>\n<hr>\n<p>I hope this clears things up!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1485654,
      "author_name": "mohammadhosein1998",
      "author_url": "",
      "post_date": "08/22/2021 10:24:13",
      "content": "<p>Thanks. how many images exist in this task1 dataset?<br>\nall of this images contain tumor. isn't it?<br>\nIt is useful for our model to use it or not?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1491485,
          "author_name": "dschettler8845",
          "author_url": "",
          "post_date": "08/26/2021 12:27:57",
          "content": "<p><em>how many images exist in this task1 dataset?</em></p>\n<p>I believe there are 1262 patients/BraTSIDs in the Task 1 dataset. With each patient/ID containing 155 channels. This could be interpreted as 195610 slices or images if you like.</p>\n<p><br></p>\n<p><em>all of this images contain tumor. isn't it?</em></p>\n<p>All of the patient's MRI scans contain tumours. However, not all slices/channels within a patient will have the tumour present.</p>\n<p><br></p>\n<p><em>It is useful for our model to use it or not?</em></p>\n<p>That's up to you to decide.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1490312,
      "author_name": "neelambuj717",
      "author_url": "",
      "post_date": "08/25/2021 15:21:48",
      "content": "<p>Hi Darien,</p>\n<p>Thanks a lot for uploading the dataset for Task1.</p>\n<p>While exploring the data, I saw that the #Channels for all the modalities across patients is constant i.e. 155. But in the dataset we have on kaggle (Task 2) we have variable set of dicom images.</p>\n<p>Is this inconsistency due to the pre processing steps they did while preparing the task1 data?</p>\n<p>How should we tackle this while building a pipeline which contains Segmentation -&gt; Classification?</p>\n<p>Thanks</p>",
      "votes": null,
      "replies": [
        {
          "id": 1491483,
          "author_name": "dschettler8845",
          "author_url": "",
          "post_date": "08/26/2021 12:24:53",
          "content": "<p>The images are preprocessed differently in Task 1 v. Task 2.</p>\n<p>To use the Task 1 images effectively, you may need to convert them so that they are similarly processed/aligned/registered.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1490800,
      "author_name": "pankajattri",
      "author_url": "",
      "post_date": "08/25/2021 21:41:37",
      "content": "<p>Is it safe to assume that the subject IDs represent the same patients across Task 1 and Task 2 data? So, if I were to look at Task 1's 00000 subject the NIFTI files will have all characteristics in it for it to be designated as MGMT_1?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1490928,
          "author_name": "bcghost",
          "author_url": "",
          "post_date": "08/26/2021 02:28:30",
          "content": "<p>If it is, it will be a huge data leakage.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1490942,
          "author_name": "pankajattri",
          "author_url": "",
          "post_date": "08/26/2021 03:05:30",
          "content": "<p>hmm… why do you say that? Task1's data is only going to give us tumor segmentations and not any information about presence of methylation</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1491482,
          "author_name": "dschettler8845",
          "author_url": "",
          "post_date": "08/26/2021 12:23:56",
          "content": "<p>Yes. From what I can see the BraTSIDs overlap between the two training datasets with only a fraction of the IDs present in Task 2 not being present in Task 1.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1496239,
      "author_name": "sungjangwon",
      "author_url": "",
      "post_date": "08/30/2021 08:08:56",
      "content": "<p>Thank you Darien for your hard work!! Can I ask how to convert this nii.gz to dcm? I tried to find it but I couldn't convert them… and if so is it going to be similar to the data that is from task2 data?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1482233": "Hi there.\n\nI uploaded the task one dataset ([**here**](https://www.kaggle.com/dschettler8845/brats-2021-task1)) and made a basic notebook ([**here**](https://www.kaggle.com/dschettler8845/how-to-load-basic-data-exploration)) showing how to **`untar`** and use a basic library to convert to NumPy.\n\nPlease ensure that you register for Task 1 via the Synapse website prior to using the dataset. Also please ensure you read all of the rules and whatnot before using the data.\n  -> [**SYNAPSE LINK HERE**](https://www.synapse.org/#!Synapse:syn25829067/wiki/)\n\nI believe having this data public on the Kaggle platform is allowed... if not please let me know and I can take it down.\n\n---\n\nNOTE: Please see [**this post**](https://www.kaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/253488) on notes/concerns regarding using this as external data.\n\n---\n\nI hope this helps!",
    "1482291": "Thank you. I wonder if we can resample task 2's data into the task 1's by ourselves. I read from the paper they contains several steps such as: convert to RAI system, co-register on SRI24, resample spacing to 1mm x 1mm x 1mm. The host used the CaPTk tools to process their task1 data. But we cannot install such software on Kaggle.",
    "1482402": "your notebook link is broken",
    "1482993": "Ah sorry. It was still private. Should be fixed now.",
    "1483237": "in the data description of brats2021 its written \n\"pathologically confirmed diagnosis and available MGMT promoter methylation status, are used as the training, validation, and testing data for this year’s BraTS challenge.\"\nwhat does it mean, any idea?",
    "1483456": "That's what I will be doing... or I will look for outside tools to accomplish the same that can be accessed within Kaggle.",
    "1485654": "Thanks. how many images exist in this task1 dataset?\nall of this images contain tumor. isn't it?\nIt is useful for our model to use it or not?",
    "1485656": "please share this paper that you said . \nwhat is CaPTK tools and how we can use it ? any notebook is available ?",
    "1486126": "https://arxiv.org/abs/2107.02314#:~:text=The%20RSNA%2DASNR%2DMICCAI%20BraTS%202021%20challenge%20targets%20the%20evaluation,mpMRI%20data%20from%202%2C000%20patients",
    "1490312": "Hi Darien,\n\nThanks a lot for uploading the dataset for Task1.\n\nWhile exploring the data, I saw that the #Channels for all the modalities across patients is constant i.e. 155. But in the dataset we have on kaggle (Task 2) we have variable set of dicom images.\n\nIs this inconsistency due to the pre processing steps they did while preparing the task1 data?\n\nHow should we tackle this while building a pipeline which contains Segmentation -> Classification?\n\nThanks",
    "1490800": "Is it safe to assume that the subject IDs represent the same patients across Task 1 and Task 2 data? So, if I were to look at Task 1's 00000 subject the NIFTI files will have all characteristics in it for it to be designated as MGMT_1?",
    "1490928": "If it is, it will be a huge data leakage.",
    "1490942": "hmm... why do you say that? Task1's data is only going to give us tumor segmentations and not any information about presence of methylation",
    "1491332": "Thanks. I have two question about task1 dataset.How many pictures are there in this dataset and are those different whit this kaggle dataset?\nIn this dataset exist pictures with tumor and without tumor or not ?and how many different pictures are there exist?",
    "1491482": "Yes. From what I can see the BraTSIDs overlap between the two training datasets with only a fraction of the IDs present in Task 2 not being present in Task 1.",
    "1491483": "The images are preprocessed differently in Task 1 v. Task 2.\n\nTo use the Task 1 images effectively, you may need to convert them so that they are similarly processed/aligned/registered.",
    "1491485": "*how many images exist in this task1 dataset?*\n\nI believe there are 1262 patients/BraTSIDs in the Task 1 dataset. With each patient/ID containing 155 channels. This could be interpreted as 195610 slices or images if you like.\n\n<br>\n\n*all of this images contain tumor. isn't it?*\n\nAll of the patient's MRI scans contain tumours. However, not all slices/channels within a patient will have the tumour present.\n\n<br>\n\n*It is useful for our model to use it or not?*\n\nThat's up to you to decide.",
    "1491491": "A few quick definitions first...\n\n---\n\nThe definition of **\"pathologically confirmed diagnosis\"** is the following:\n\n> Determination of the cause or causes of an illness by examining fluids and tissues from the patient before or after death. The examination may be performed on blood, plasma, microscopic tissue samples, or gross specimens.\n\nAnd MGMT promoter methylation status is determined usually using the following techniques:\n\n> Quantitative PCR, methylation-specific PCR, pyrosequencing, real time PCR with high resolution melt, or the infinitum methylation EPIC beadChip\n\n---\n\n**To answer your question, I think they are just saying that they determined the labels (tumour segmentation maps and MGMT promoter methylation status) based on a combination of autopsy and genetic investigation techniques.**",
    "1491493": "Answered in the above comment.",
    "1491550": "But there is no MGMT status for the data in that task right",
    "1491585": "You are correct. Technically no MGMT status is given for Task 1.\n\nHowever, both Task 1 and Task 2 use BraTSID numbers indicating a particular patient. The patient is the same in both datasets and therefore the MGMT status from Task 2 can be assumed for Task 1... and similarly, the segmentation maps (if aligned properly) can be considered for Task 2.\n\nIf you read through the [paper authored by the competition hosts](https://arxiv.org/pdf/2107.02314.pdf), you can see that the data goes through a pipeline roughly like (p. 5 and 6).\n\n<br>\n\n**TO OBTAIN THE TASK 1 IMAGES**\n**Standardized pre-processing has been applied to all the ORIGINAL BraTS mpMRI scans** \n  - Specifically, the applied pre-processing routines include conversion of the DICOM files to the NIFTI file format\n  - re-orientation to a common orientation system (i.e., RAI)\n  - co-registration to the same anatomical template (SRI24)\n  - resampling to a uniform isotropic resolution (1mm3)\n  - skull-stripping\n\n***NOTE:*** *The preprocessing pipeline is publicly available through the Cancer Imaging Phenomics Toolkit (CaPTk) and Federated Tumor Segmentation (FeTS) tool 1*\n\n<br>\n\n**TO OBTAIN THE TASK 2 IMAGES**\n**All the imaging volumes were converted from NIFTI to DICOM files, while ensuring that the original patient space is preserved**\n* To make this conversion both the skull-stripped brain volume in NIFTI format of each MRI sequence and its corresponding original DICOM scan in the patient space are required\n* The DICOM volume is read as an ITK image and the skull-stripped volume is rigidly registered to it, providing a transformation matrix that defines the spatial mapping between the 2 volumes.\n  - This transformation matrix is applied to the skull-stripped volume and to the corresponding segmentation labels, in order to translate them both to the patient space. \n* These transformed volumes are then passed through CaPTk’s NIFTI to DICOM conversion engine to generate DICOM image volumes for the skullstripped image. \n* Once all MRI sequences were converted back to the DICOM file format, further de-identification took place based on a two-step process. \n  - The first step used the RSNA CTP (Clinical Trials Processor) Anonymizer 2 with the standard built-in script. \n  - Step two then consisted of whitelisting the DICOM files from step 1. \n  - The whitelisting process removes all non-essential tags from the DICOM header. \n  - This last process ensures there are no protected health information (PHI) entries left in the DICOM header.\n\n---\n\nI hope this clears things up!",
    "1496239": "Thank you Darien for your hard work!! Can I ask how to convert this nii.gz to dcm? I tried to find it but I couldn't convert them... and if so is it going to be similar to the data that is from task2 data?"
  },
  "source": "meta"
}