{
  "id": 269164,
  "title": "Why Is The Task 2 Dataset Different Than The Task 1 Dataset?",
  "url": "/competitions/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/269164",
  "author_name": "",
  "post_date": "2021-08-30T15:12:08.279059100Z",
  "votes": 30,
  "comment_count": 10,
  "views": 0,
  "content": "<p>This question is primarily for the competition hosts and/or Kaggle staff. <a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> <a href=\"https://www.kaggle.com/sbakas\" target=\"_blank\">@sbakas</a> <a href=\"https://www.kaggle.com/ujjwalbaid\" target=\"_blank\">@ujjwalbaid</a> <a href=\"https://www.kaggle.com/cdcarr\" target=\"_blank\">@cdcarr</a> </p>\n<hr>\n<p>Many of us in this competition are spending a significant period of time trying to convert the Task 2 Dataset to the Task 1 Dataset <a href=\"https://cbica.github.io/CaPTk/preprocessing_brats.html\" target=\"_blank\">structure/style</a></p>\n<p>This is for a number of reasons including, but not limited to:</p>\n<ul>\n<li><strong>Being able to use the GT segmentation masks from the Task 1 Dataset</strong></li>\n<li><strong>Having a consistent number of images/orientation for each patient and modality</strong></li>\n<li>N4 Bias Correction already performed</li>\n<li>etc.</li>\n</ul>\n<hr>\n<p>Is there a particular reason the data was provided in the format it was? </p>\n<p>Am I missing something? i.e.</p>\n<ul>\n<li>Are we not supposed to use segmentation as an intermediary step towards predicting MGMT Promoter Methylation Status? </li>\n<li>Is part of this competition specifically to identify faster (more end-to-end) methods of going from raw DICOM to a label? </li>\n</ul>\n<p>To me, this seems to show that the hosts believe that this task would be trivial if the data was provided in a similar/same format as in Task 1. However, I don't believe this is the case, as in most recent research papers, performance has been limited to under 90% AUC (sometimes even lower). </p>\n<hr>\n<p><br></p>\n<p><strong><em>This might be an impossible ask… but:</em></strong></p>\n<p><strong>Would it be possible to get a secondary competition dataset added that contains the same files, but with the appropriate Task 1 Preprocessing?</strong> </p>\n<p><em>This secondary dataset would function similar to the current folders (full train dataset and the test dataset would be swapped from public to private upon submission).</em></p>\n<p><br></p>\n<hr>\n<p><em>Sorry for my ignorance and thanks in advance.</em></p>",
  "messages": [
    {
      "id": "1496698",
      "postDate": "08/30/2021 15:12:08",
      "content": "<p>This question is primarily for the competition hosts and/or Kaggle staff. <a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> <a href=\"https://www.kaggle.com/sbakas\" target=\"_blank\">@sbakas</a> <a href=\"https://www.kaggle.com/ujjwalbaid\" target=\"_blank\">@ujjwalbaid</a> <a href=\"https://www.kaggle.com/cdcarr\" target=\"_blank\">@cdcarr</a> </p>\n<hr>\n<p>Many of us in this competition are spending a significant period of time trying to convert the Task 2 Dataset to the Task 1 Dataset <a href=\"https://cbica.github.io/CaPTk/preprocessing_brats.html\" target=\"_blank\">structure/style</a></p>\n<p>This is for a number of reasons including, but not limited to:</p>\n<ul>\n<li><strong>Being able to use the GT segmentation masks from the Task 1 Dataset</strong></li>\n<li><strong>Having a consistent number of images/orientation for each patient and modality</strong></li>\n<li>N4 Bias Correction already performed</li>\n<li>etc.</li>\n</ul>\n<hr>\n<p>Is there a particular reason the data was provided in the format it was? </p>\n<p>Am I missing something? i.e.</p>\n<ul>\n<li>Are we not supposed to use segmentation as an intermediary step towards predicting MGMT Promoter Methylation Status? </li>\n<li>Is part of this competition specifically to identify faster (more end-to-end) methods of going from raw DICOM to a label? </li>\n</ul>\n<p>To me, this seems to show that the hosts believe that this task would be trivial if the data was provided in a similar/same format as in Task 1. However, I don't believe this is the case, as in most recent research papers, performance has been limited to under 90% AUC (sometimes even lower). </p>\n<hr>\n<p><br></p>\n<p><strong><em>This might be an impossible ask… but:</em></strong></p>\n<p><strong>Would it be possible to get a secondary competition dataset added that contains the same files, but with the appropriate Task 1 Preprocessing?</strong> </p>\n<p><em>This secondary dataset would function similar to the current folders (full train dataset and the test dataset would be swapped from public to private upon submission).</em></p>\n<p><br></p>\n<hr>\n<p><em>Sorry for my ignorance and thanks in advance.</em></p>",
      "rawMarkdown": "This question is primarily for the competition hosts and/or Kaggle staff. @juliaelliott @sbakas @ujjwalbaid @cdcarr \n\n---\n\nMany of us in this competition are spending a significant period of time trying to convert the Task 2 Dataset to the Task 1 Dataset [structure/style](https://cbica.github.io/CaPTk/preprocessing_brats.html)\n\nThis is for a number of reasons including, but not limited to:\n- **Being able to use the GT segmentation masks from the Task 1 Dataset**\n- **Having a consistent number of images/orientation for each patient and modality**\n- N4 Bias Correction already performed\n- etc.\n\n---\n\nIs there a particular reason the data was provided in the format it was? \n\nAm I missing something? i.e.\n* Are we not supposed to use segmentation as an intermediary step towards predicting MGMT Promoter Methylation Status? \n* Is part of this competition specifically to identify faster (more end-to-end) methods of going from raw DICOM to a label? \n\nTo me, this seems to show that the hosts believe that this task would be trivial if the data was provided in a similar/same format as in Task 1. However, I don't believe this is the case, as in most recent research papers, performance has been limited to under 90% AUC (sometimes even lower). \n\n---\n\n<br>\n\n***This might be an impossible ask... but:***\n\n**Would it be possible to get a secondary competition dataset added that contains the same files, but with the appropriate Task 1 Preprocessing?** \n\n*This secondary dataset would function similar to the current folders (full train dataset and the test dataset would be swapped from public to private upon submission).*\n\n<br>\n\n---\n\n*Sorry for my ignorance and thanks in advance.*",
      "votes": null
    },
    {
      "id": "1496886",
      "postDate": "08/30/2021 17:40:07",
      "content": "<p>+1 for all the notes from <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a> above. </p>\n<p>Earlier I also asked a related question about releasing the pre-processing scripts.<br>\n<a href=\"https://www.kaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/267818\" target=\"_blank\">https://www.kaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/267818</a></p>",
      "rawMarkdown": "1 for all the notes from @dschettler8845 above. \n\nEarlier I also asked a related question about releasing the pre-processing scripts.\nhttps://www.kaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/267818",
      "votes": null
    },
    {
      "id": "1496919",
      "postDate": "08/30/2021 18:07:06",
      "content": "<p>Yes. I saw this.</p>\n<p>I believe they use <a href=\"https://github.com/CBICA/CaPTk\" target=\"_blank\"><strong>CaPTk</strong></a> and/or <a href=\"https://github.com/FETS-AI/Front-End/\" target=\"_blank\"><strong>FeTS</strong></a> to generate the Task 1 dataset.</p>\n<p>This is from the <a href=\"https://arxiv.org/pdf/2107.02314.pdf\" target=\"_blank\"><strong>BraTS 2021 Publication</strong></a></p>\n<blockquote>\n  <p>The preprocessing pipeline is publicly available through the Cancer Imaging Phenomics Toolkit (CaPTk) and Federated Tumor Segmentation (FeTS) tool</p>\n</blockquote>",
      "rawMarkdown": "Yes. I saw this.\n\nI believe they use [**CaPTk**](https://github.com/CBICA/CaPTk) and/or [**FeTS**](https://github.com/FETS-AI/Front-End/) to generate the Task 1 dataset.\n\nThis is from the [**BraTS 2021 Publication**](https://arxiv.org/pdf/2107.02314.pdf)\n> The preprocessing pipeline is publicly available through the Cancer Imaging Phenomics Toolkit (CaPTk) and Federated Tumor Segmentation (FeTS) tool",
      "votes": null
    },
    {
      "id": "1497638",
      "postDate": "08/31/2021 10:35:44",
      "content": "<p>Hey, just checking, I think N4 bias correction has not been applied on the images from task 1. For what I understand they apply it only before the registration, but then they use the transformation matrix they get to register the biased images.</p>\n<p><a href=\"https://cbica.github.io/CaPTk/preprocessing_brats.html\" target=\"_blank\">https://cbica.github.io/CaPTk/preprocessing_brats.html</a> </p>",
      "rawMarkdown": "Hey, just checking, I think N4 bias correction has not been applied on the images from task 1. For what I understand they apply it only before the registration, but then they use the transformation matrix they get to register the biased images.\n\nhttps://cbica.github.io/CaPTk/preprocessing_brats.html",
      "votes": null
    },
    {
      "id": "1497642",
      "postDate": "08/31/2021 10:39:41",
      "content": "<p>Yeah, I believe your right.</p>\n<p>However, we still have to run that part to successfully convert between tasks I think?</p>",
      "rawMarkdown": "Yeah, I believe your right.\n\nHowever, we still have to run that part to successfully convert between tasks I think?",
      "votes": null
    },
    {
      "id": "1497745",
      "postDate": "08/31/2021 11:59:12",
      "content": "<p>+1 for getting some clarification on this and if possible a compatible dataset with the segmentation information from task 1.</p>\n<p>Many people are spending a lot of time data figuring out aspects of data compatibility between the two Tasks and this can impact the amount left for focusing on modeling bits. So it would be really beneficial to get a clear answer on this.</p>",
      "rawMarkdown": "1 for getting some clarification on this and if possible a compatible dataset with the segmentation information from task 1.\n\nMany people are spending a lot of time data figuring out aspects of data compatibility between the two Tasks and this can impact the amount left for focusing on modeling bits. So it would be really beneficial to get a clear answer on this.",
      "votes": null
    },
    {
      "id": "1499133",
      "postDate": "09/01/2021 13:12:41",
      "content": "<p>Well, it just \"improves\" the image. It is like harmonizing the light of a normal image. You can avoid it but at your own risk (worse data worse result). If you are performing a data augmentation step, you might use torchio to add data with the different bias field artifact. I guess that would help during inference. </p>\n<p><a href=\"https://github.com/fepegar/torchio\" target=\"_blank\">https://github.com/fepegar/torchio</a></p>",
      "rawMarkdown": "Well, it just \"improves\" the image. It is like harmonizing the light of a normal image. You can avoid it but at your own risk (worse data worse result). If you are performing a data augmentation step, you might use torchio to add data with the different bias field artifact. I guess that would help during inference. \n\nhttps://github.com/fepegar/torchio",
      "votes": null
    },
    {
      "id": "1499228",
      "postDate": "09/01/2021 14:05:24",
      "content": "<p>Interesting! Thanks for the information!</p>",
      "rawMarkdown": "Interesting! Thanks for the information!",
      "votes": null
    },
    {
      "id": "1512142",
      "postDate": "09/14/2021 02:58:12",
      "content": "<p>As for the training part, I have just checked the rules and it says that 'C. External Data. You may use data other than the Competition Data (“External Data”) to develop and test your Submissions'. Since the data for Task 1 is also publicly available, I think we can at least get the data for task 1 and use it for training first.</p>",
      "rawMarkdown": "As for the training part, I have just checked the rules and it says that 'C. External Data. You may use data other than the Competition Data (“External Data”) to develop and test your Submissions'. Since the data for Task 1 is also publicly available, I think we can at least get the data for task 1 and use it for training first.",
      "votes": null
    },
    {
      "id": "1512756",
      "postDate": "09/14/2021 14:48:33",
      "content": "<p>That's what I'm currently doing, however, it offers no benefit to train on Task 1 data if we can't transform the Task 2 <em>test</em> data into a similar format.</p>",
      "rawMarkdown": "That's what I'm currently doing, however, it offers no benefit to train on Task 1 data if we can't transform the Task 2 *test* data into a similar format.",
      "votes": null
    },
    {
      "id": "1534357",
      "postDate": "10/04/2021 20:06:00",
      "content": "<p><a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a> et al did you successfully performed preprocessing (registration, etc)+ inference in the Kaggle kernel within the time limit (9 hours CPU run time). It seems preprocessing a single subject takes about 5-10 minutes. How is it feasible to complete the inference of all test subjects within the time limit?</p>",
      "rawMarkdown": "dschettler8845 et al did you successfully performed preprocessing (registration, etc)+ inference in the Kaggle kernel within the time limit (9 hours CPU run time). It seems preprocessing a single subject takes about 5-10 minutes. How is it feasible to complete the inference of all test subjects within the time limit?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1496886,
      "author_name": "mpsampat",
      "author_url": "",
      "post_date": "08/30/2021 17:40:07",
      "content": "<p>+1 for all the notes from <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a> above. </p>\n<p>Earlier I also asked a related question about releasing the pre-processing scripts.<br>\n<a href=\"https://www.kaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/267818\" target=\"_blank\">https://www.kaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/267818</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1496919,
          "author_name": "dschettler8845",
          "author_url": "",
          "post_date": "08/30/2021 18:07:06",
          "content": "<p>Yes. I saw this.</p>\n<p>I believe they use <a href=\"https://github.com/CBICA/CaPTk\" target=\"_blank\"><strong>CaPTk</strong></a> and/or <a href=\"https://github.com/FETS-AI/Front-End/\" target=\"_blank\"><strong>FeTS</strong></a> to generate the Task 1 dataset.</p>\n<p>This is from the <a href=\"https://arxiv.org/pdf/2107.02314.pdf\" target=\"_blank\"><strong>BraTS 2021 Publication</strong></a></p>\n<blockquote>\n  <p>The preprocessing pipeline is publicly available through the Cancer Imaging Phenomics Toolkit (CaPTk) and Federated Tumor Segmentation (FeTS) tool</p>\n</blockquote>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1497638,
      "author_name": "ranafago",
      "author_url": "",
      "post_date": "08/31/2021 10:35:44",
      "content": "<p>Hey, just checking, I think N4 bias correction has not been applied on the images from task 1. For what I understand they apply it only before the registration, but then they use the transformation matrix they get to register the biased images.</p>\n<p><a href=\"https://cbica.github.io/CaPTk/preprocessing_brats.html\" target=\"_blank\">https://cbica.github.io/CaPTk/preprocessing_brats.html</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 1497642,
          "author_name": "dschettler8845",
          "author_url": "",
          "post_date": "08/31/2021 10:39:41",
          "content": "<p>Yeah, I believe your right.</p>\n<p>However, we still have to run that part to successfully convert between tasks I think?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1499133,
          "author_name": "ranafago",
          "author_url": "",
          "post_date": "09/01/2021 13:12:41",
          "content": "<p>Well, it just \"improves\" the image. It is like harmonizing the light of a normal image. You can avoid it but at your own risk (worse data worse result). If you are performing a data augmentation step, you might use torchio to add data with the different bias field artifact. I guess that would help during inference. </p>\n<p><a href=\"https://github.com/fepegar/torchio\" target=\"_blank\">https://github.com/fepegar/torchio</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1499228,
          "author_name": "dschettler8845",
          "author_url": "",
          "post_date": "09/01/2021 14:05:24",
          "content": "<p>Interesting! Thanks for the information!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1497745,
      "author_name": "smoschou55",
      "author_url": "",
      "post_date": "08/31/2021 11:59:12",
      "content": "<p>+1 for getting some clarification on this and if possible a compatible dataset with the segmentation information from task 1.</p>\n<p>Many people are spending a lot of time data figuring out aspects of data compatibility between the two Tasks and this can impact the amount left for focusing on modeling bits. So it would be really beneficial to get a clear answer on this.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1512142,
      "author_name": "formtyan",
      "author_url": "",
      "post_date": "09/14/2021 02:58:12",
      "content": "<p>As for the training part, I have just checked the rules and it says that 'C. External Data. You may use data other than the Competition Data (“External Data”) to develop and test your Submissions'. Since the data for Task 1 is also publicly available, I think we can at least get the data for task 1 and use it for training first.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1512756,
          "author_name": "dschettler8845",
          "author_url": "",
          "post_date": "09/14/2021 14:48:33",
          "content": "<p>That's what I'm currently doing, however, it offers no benefit to train on Task 1 data if we can't transform the Task 2 <em>test</em> data into a similar format.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1534357,
      "author_name": "saruaralam",
      "author_url": "",
      "post_date": "10/04/2021 20:06:00",
      "content": "<p><a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a> et al did you successfully performed preprocessing (registration, etc)+ inference in the Kaggle kernel within the time limit (9 hours CPU run time). It seems preprocessing a single subject takes about 5-10 minutes. How is it feasible to complete the inference of all test subjects within the time limit?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1496698": "This question is primarily for the competition hosts and/or Kaggle staff. @juliaelliott @sbakas @ujjwalbaid @cdcarr \n\n---\n\nMany of us in this competition are spending a significant period of time trying to convert the Task 2 Dataset to the Task 1 Dataset [structure/style](https://cbica.github.io/CaPTk/preprocessing_brats.html)\n\nThis is for a number of reasons including, but not limited to:\n- **Being able to use the GT segmentation masks from the Task 1 Dataset**\n- **Having a consistent number of images/orientation for each patient and modality**\n- N4 Bias Correction already performed\n- etc.\n\n---\n\nIs there a particular reason the data was provided in the format it was? \n\nAm I missing something? i.e.\n* Are we not supposed to use segmentation as an intermediary step towards predicting MGMT Promoter Methylation Status? \n* Is part of this competition specifically to identify faster (more end-to-end) methods of going from raw DICOM to a label? \n\nTo me, this seems to show that the hosts believe that this task would be trivial if the data was provided in a similar/same format as in Task 1. However, I don't believe this is the case, as in most recent research papers, performance has been limited to under 90% AUC (sometimes even lower). \n\n---\n\n<br>\n\n***This might be an impossible ask... but:***\n\n**Would it be possible to get a secondary competition dataset added that contains the same files, but with the appropriate Task 1 Preprocessing?** \n\n*This secondary dataset would function similar to the current folders (full train dataset and the test dataset would be swapped from public to private upon submission).*\n\n<br>\n\n---\n\n*Sorry for my ignorance and thanks in advance.*",
    "1496886": "1 for all the notes from @dschettler8845 above. \n\nEarlier I also asked a related question about releasing the pre-processing scripts.\nhttps://www.kaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/267818",
    "1496919": "Yes. I saw this.\n\nI believe they use [**CaPTk**](https://github.com/CBICA/CaPTk) and/or [**FeTS**](https://github.com/FETS-AI/Front-End/) to generate the Task 1 dataset.\n\nThis is from the [**BraTS 2021 Publication**](https://arxiv.org/pdf/2107.02314.pdf)\n> The preprocessing pipeline is publicly available through the Cancer Imaging Phenomics Toolkit (CaPTk) and Federated Tumor Segmentation (FeTS) tool",
    "1497638": "Hey, just checking, I think N4 bias correction has not been applied on the images from task 1. For what I understand they apply it only before the registration, but then they use the transformation matrix they get to register the biased images.\n\nhttps://cbica.github.io/CaPTk/preprocessing_brats.html",
    "1497642": "Yeah, I believe your right.\n\nHowever, we still have to run that part to successfully convert between tasks I think?",
    "1497745": "1 for getting some clarification on this and if possible a compatible dataset with the segmentation information from task 1.\n\nMany people are spending a lot of time data figuring out aspects of data compatibility between the two Tasks and this can impact the amount left for focusing on modeling bits. So it would be really beneficial to get a clear answer on this.",
    "1499133": "Well, it just \"improves\" the image. It is like harmonizing the light of a normal image. You can avoid it but at your own risk (worse data worse result). If you are performing a data augmentation step, you might use torchio to add data with the different bias field artifact. I guess that would help during inference. \n\nhttps://github.com/fepegar/torchio",
    "1499228": "Interesting! Thanks for the information!",
    "1512142": "As for the training part, I have just checked the rules and it says that 'C. External Data. You may use data other than the Competition Data (“External Data”) to develop and test your Submissions'. Since the data for Task 1 is also publicly available, I think we can at least get the data for task 1 and use it for training first.",
    "1512756": "That's what I'm currently doing, however, it offers no benefit to train on Task 1 data if we can't transform the Task 2 *test* data into a similar format.",
    "1534357": "dschettler8845 et al did you successfully performed preprocessing (registration, etc)+ inference in the Kaggle kernel within the time limit (9 hours CPU run time). It seems preprocessing a single subject takes about 5-10 minutes. How is it feasible to complete the inference of all test subjects within the time limit?"
  },
  "source": "meta"
}