{
  "id": 280448,
  "title": "Top 30 finish with 1 submission",
  "url": "/competitions/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/280448",
  "author_name": "ImBczVr",
  "post_date": "2021-10-21T14:58:04.865000",
  "votes": 7,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I checked the results this morning and was surprised to see a top 30 finish. TBH I did not expect it because my public score was only 0.54 and going by the early public scores of 0.7+ I was not expecting any better results. I want to highlight that this wasn't a fluke and I really did spend a lot of time on this but ended up submitting only one version. I wanted to try combining traditional models with CNNs but there wasn't enough time in the end. For the only submission that I was able to put in I used decision trees based models (RF, XGBoost etc.) and used Radiomics for extracting features related to tumor shapes. I am sharing my approach here hoping it will be of help to some:</p>\n<p>Use <strong>Task 1 Dataset</strong> to create segmentation model (thank you <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a> for creating the dataset)</p>\n<ol>\n<li>Convert Task 1 Nifti files to Dicom files using plastimatch. This was done to align with Task 2 file formats. I only used FLAIR and Segmentation files. </li>\n<li>Use pydicom to return 3D voxels from Dicom files. Randomly select one of the three axis for a subject and convert to TFrecords for training a VGG based U-Net with nodes aggregation segmentation model. I did not predict on multiple classes… model only predicts tumorous cells. I also did not refine the model too much and was satisfied with a CV score of about 0.92.</li>\n</ol>\n<p><strong>Task 2 Dataset:</strong></p>\n<ol>\n<li>Resample Task 2 dataset to correct for inconsistent spacing between slices &amp; the image shapes.</li>\n<li>Predict tumor segments using model trained in step 1. Choose 154 slices for each subject with the middle slice having the max tumor area.</li>\n<li>Use Pyradiomics to extract features by passing predicted masks and the corresponding slices from #2. I only extracted the firstorder, shape, and grey length features… totaling 58 features. Use these features to train a XGboost model.</li>\n<li>Pass Test images from the same pipeline and predict prob of MGMT class.</li>\n</ol>\n<p>As you can see there is a lot of scope for improvement in this approach. By using true masks from Task 1 dataset I was getting CV scores of 0.64+. The competition hosts can use true masks for Task 2 dataset, and also use T1/T2 type scans. I had hoped to also extract some features using a simple CNN model using the predicted masks but did not get time. </p>\n<p>Hopefully this explanation was useful to some. I did learn a lot from this competition so a shout out to the competition hosts.</p>",
  "messages": [
    {
      "id": 1552624,
      "postDate": "2021-10-21T14:58:04.867Z",
      "content": "<p>I checked the results this morning and was surprised to see a top 30 finish. TBH I did not expect it because my public score was only 0.54 and going by the early public scores of 0.7+ I was not expecting any better results. I want to highlight that this wasn't a fluke and I really did spend a lot of time on this but ended up submitting only one version. I wanted to try combining traditional models with CNNs but there wasn't enough time in the end. For the only submission that I was able to put in I used decision trees based models (RF, XGBoost etc.) and used Radiomics for extracting features related to tumor shapes. I am sharing my approach here hoping it will be of help to some:</p>\n<p>Use <strong>Task 1 Dataset</strong> to create segmentation model (thank you <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a> for creating the dataset)</p>\n<ol>\n<li>Convert Task 1 Nifti files to Dicom files using plastimatch. This was done to align with Task 2 file formats. I only used FLAIR and Segmentation files. </li>\n<li>Use pydicom to return 3D voxels from Dicom files. Randomly select one of the three axis for a subject and convert to TFrecords for training a VGG based U-Net with nodes aggregation segmentation model. I did not predict on multiple classes… model only predicts tumorous cells. I also did not refine the model too much and was satisfied with a CV score of about 0.92.</li>\n</ol>\n<p><strong>Task 2 Dataset:</strong></p>\n<ol>\n<li>Resample Task 2 dataset to correct for inconsistent spacing between slices &amp; the image shapes.</li>\n<li>Predict tumor segments using model trained in step 1. Choose 154 slices for each subject with the middle slice having the max tumor area.</li>\n<li>Use Pyradiomics to extract features by passing predicted masks and the corresponding slices from #2. I only extracted the firstorder, shape, and grey length features… totaling 58 features. Use these features to train a XGboost model.</li>\n<li>Pass Test images from the same pipeline and predict prob of MGMT class.</li>\n</ol>\n<p>As you can see there is a lot of scope for improvement in this approach. By using true masks from Task 1 dataset I was getting CV scores of 0.64+. The competition hosts can use true masks for Task 2 dataset, and also use T1/T2 type scans. I had hoped to also extract some features using a simple CNN model using the predicted masks but did not get time. </p>\n<p>Hopefully this explanation was useful to some. I did learn a lot from this competition so a shout out to the competition hosts.</p>",
      "rawMarkdown": "I checked the results this morning and was surprised to see a top 30 finish. TBH I did not expect it because my public score was only 0.54 and going by the early public scores of 0.7+ I was not expecting any better results. I want to highlight that this wasn't a fluke and I really did spend a lot of time on this but ended up submitting only one version. I wanted to try combining traditional models with CNNs but there wasn't enough time in the end. For the only submission that I was able to put in I used decision trees based models (RF, XGBoost etc.) and used Radiomics for extracting features related to tumor shapes. I am sharing my approach here hoping it will be of help to some:\n\nUse **Task 1 Dataset** to create segmentation model (thank you @dschettler8845 for creating the dataset)\n1. Convert Task 1 Nifti files to Dicom files using plastimatch. This was done to align with Task 2 file formats. I only used FLAIR and Segmentation files. \n2. Use pydicom to return 3D voxels from Dicom files. Randomly select one of the three axis for a subject and convert to TFrecords for training a VGG based U-Net with nodes aggregation segmentation model. I did not predict on multiple classes... model only predicts tumorous cells. I also did not refine the model too much and was satisfied with a CV score of about 0.92.\n\n**Task 2 Dataset:**\n1. Resample Task 2 dataset to correct for inconsistent spacing between slices & the image shapes.\n2. Predict tumor segments using model trained in step 1. Choose 154 slices for each subject with the middle slice having the max tumor area.\n3. Use Pyradiomics to extract features by passing predicted masks and the corresponding slices from #2. I only extracted the firstorder, shape, and grey length features... totaling 58 features. Use these features to train a XGboost model.\n4. Pass Test images from the same pipeline and predict prob of MGMT class.\n\nAs you can see there is a lot of scope for improvement in this approach. By using true masks from Task 1 dataset I was getting CV scores of 0.64+. The competition hosts can use true masks for Task 2 dataset, and also use T1/T2 type scans. I had hoped to also extract some features using a simple CNN model using the predicted masks but did not get time. \n\nHopefully this explanation was useful to some. I did learn a lot from this competition so a shout out to the competition hosts.\n",
      "votes": 7
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1552624": "I checked the results this morning and was surprised to see a top 30 finish. TBH I did not expect it because my public score was only 0.54 and going by the early public scores of 0.7+ I was not expecting any better results. I want to highlight that this wasn't a fluke and I really did spend a lot of time on this but ended up submitting only one version. I wanted to try combining traditional models with CNNs but there wasn't enough time in the end. For the only submission that I was able to put in I used decision trees based models (RF, XGBoost etc.) and used Radiomics for extracting features related to tumor shapes. I am sharing my approach here hoping it will be of help to some:\n\nUse **Task 1 Dataset** to create segmentation model (thank you @dschettler8845 for creating the dataset)\n1. Convert Task 1 Nifti files to Dicom files using plastimatch. This was done to align with Task 2 file formats. I only used FLAIR and Segmentation files. \n2. Use pydicom to return 3D voxels from Dicom files. Randomly select one of the three axis for a subject and convert to TFrecords for training a VGG based U-Net with nodes aggregation segmentation model. I did not predict on multiple classes... model only predicts tumorous cells. I also did not refine the model too much and was satisfied with a CV score of about 0.92.\n\n**Task 2 Dataset:**\n1. Resample Task 2 dataset to correct for inconsistent spacing between slices & the image shapes.\n2. Predict tumor segments using model trained in step 1. Choose 154 slices for each subject with the middle slice having the max tumor area.\n3. Use Pyradiomics to extract features by passing predicted masks and the corresponding slices from #2. I only extracted the firstorder, shape, and grey length features... totaling 58 features. Use these features to train a XGboost model.\n4. Pass Test images from the same pipeline and predict prob of MGMT class.\n\nAs you can see there is a lot of scope for improvement in this approach. By using true masks from Task 1 dataset I was getting CV scores of 0.64+. The competition hosts can use true masks for Task 2 dataset, and also use T1/T2 type scans. I had hoped to also extract some features using a simple CNN model using the predicted masks but did not get time. \n\nHopefully this explanation was useful to some. I did learn a lot from this competition so a shout out to the competition hosts.\n"
  }
}