{
  "id": 269383,
  "title": "How to prepare the data for  model",
  "url": "/competitions/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/269383",
  "author_name": "",
  "post_date": "2021-08-31T12:31:30.931647Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi All,<br>\nI am newbie to medical image analysis, I just started this competition today .<br>\nI want to through the some of the notebooks, but I didn't understand .<br>\nplease share me any blog or video for data preparation for model.<br>\nplease help me , I would like to learn this things.</p>",
  "messages": [
    {
      "id": "1497776",
      "postDate": "08/31/2021 12:31:30",
      "content": "<p>Hi All,<br>\nI am newbie to medical image analysis, I just started this competition today .<br>\nI want to through the some of the notebooks, but I didn't understand .<br>\nplease share me any blog or video for data preparation for model.<br>\nplease help me , I would like to learn this things.</p>",
      "rawMarkdown": "Hi All,\nI am newbie to medical image analysis, I just started this competition today .\nI want to through the some of the notebooks, but I didn't understand .\nplease share me any blog or video for data preparation for model.\nplease help me , I would like to learn this things.",
      "votes": null
    },
    {
      "id": "1498065",
      "postDate": "08/31/2021 16:17:44",
      "content": "<p>I'm unsure what others have been doing so this is definitely not the best method but you can consider adding zero-padding the scans to equalize the shape across the training dataset. I looped over the dataset and found the smallest sizes you could squeeze the images by removing zero values which were:<br>\n<code>target_shapes = {\n        'FLAIR': (538, 477),\n        'T1w': (445, 394),\n        'T1wCE': (424, 388),\n        'T2w': (797, 616)\n    }</code></p>\n<p>I know this isn't the best method, but it's a starting point if you're eager to get a benchmark for what a simple 2D CNN could do using a single scan (turns out not much better than guessing…).</p>\n<p>And the padding algorithm:</p>\n<pre><code>def normalize_array(array, target_shape):\n    while np.all(array.shape != target_shape):\n        bool_array = array.shape &lt; np.array(target_shape)\n        for i, small in enumerate(bool_array):\n            # if axis i is smaller than the target shape, padding should be applied across axis i\n            if small:\n                dx = target_shape[i] - array.shape[i]\n                # if adding zeros to dimension 0, then padding vectors should match along dimension 1\n                if i == 0:\n                    array = np.concatenate([np.zeros((dx // 2, array.shape[1])), array,\n                                            np.zeros((dx // 2 + int(dx % 2 != 0), array.shape[1]))], axis=i)\n                # and vice versa\n                elif i == 1:\n                    array = np.concatenate([np.zeros((array.shape[0], dx // 2)), array,\n                                            np.zeros((array.shape[0], dx // 2 + int(dx % 2 != 0)))], axis=i)\n\n            # if axis i is larger than the target shape, cut off the shape\n            elif not small:\n                dx = array.shape[i] - target_shape[i]\n                if i == 0:\n                    array = array[dx // 2:-(dx // 2 + int(dx % 2 != 0))]\n                else:\n                    array = array[:, dx // 2:-(dx // 2 + int(dx % 2 != 0))]\n    return array\n</code></pre>\n<p>Other ideas such as factoring the orientation (coronal, transverse, sagittal) in as a feature are also ideas worth considering. </p>",
      "rawMarkdown": "I'm unsure what others have been doing so this is definitely not the best method but you can consider adding zero-padding the scans to equalize the shape across the training dataset. I looped over the dataset and found the smallest sizes you could squeeze the images by removing zero values which were:\n`target_shapes = {\n        'FLAIR': (538, 477),\n        'T1w': (445, 394),\n        'T1wCE': (424, 388),\n        'T2w': (797, 616)\n    }`\n\nI know this isn't the best method, but it's a starting point if you're eager to get a benchmark for what a simple 2D CNN could do using a single scan (turns out not much better than guessing...).\n\nAnd the padding algorithm:\n```\ndef normalize_array(array, target_shape):\n    while np.all(array.shape != target_shape):\n        bool_array = array.shape < np.array(target_shape)\n        for i, small in enumerate(bool_array):\n            # if axis i is smaller than the target shape, padding should be applied across axis i\n            if small:\n                dx = target_shape[i] - array.shape[i]\n                # if adding zeros to dimension 0, then padding vectors should match along dimension 1\n                if i == 0:\n                    array = np.concatenate([np.zeros((dx // 2, array.shape[1])), array,\n                                            np.zeros((dx // 2 + int(dx % 2 != 0), array.shape[1]))], axis=i)\n                # and vice versa\n                elif i == 1:\n                    array = np.concatenate([np.zeros((array.shape[0], dx // 2)), array,\n                                            np.zeros((array.shape[0], dx // 2 + int(dx % 2 != 0)))], axis=i)\n\n            # if axis i is larger than the target shape, cut off the shape\n            elif not small:\n                dx = array.shape[i] - target_shape[i]\n                if i == 0:\n                    array = array[dx // 2:-(dx // 2 + int(dx % 2 != 0))]\n                else:\n                    array = array[:, dx // 2:-(dx // 2 + int(dx % 2 != 0))]\n    return array\n```\n\nOther ideas such as factoring the orientation (coronal, transverse, sagittal) in as a feature are also ideas worth considering.",
      "votes": null
    },
    {
      "id": "1498723",
      "postDate": "09/01/2021 06:44:46",
      "content": "<p>Thanks Jasper</p>",
      "rawMarkdown": "Thanks Jasper",
      "votes": null
    },
    {
      "id": "1500301",
      "postDate": "09/02/2021 09:10:15",
      "content": "<p>HI suri, this notebook: <a href=\"https://www.kaggle.com/experienceinai/aivo-vii-nopea\" target=\"_blank\">https://www.kaggle.com/experienceinai/aivo-vii-nopea</a> includes step-by-step description about data preparation and modelling. Including pictures and also possibility to show additional information if you need. </p>",
      "rawMarkdown": "HI suri, this notebook: https://www.kaggle.com/experienceinai/aivo-vii-nopea includes step-by-step description about data preparation and modelling. Including pictures and also possibility to show additional information if you need.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1498065,
      "author_name": "jasperbutcher",
      "author_url": "",
      "post_date": "08/31/2021 16:17:44",
      "content": "<p>I'm unsure what others have been doing so this is definitely not the best method but you can consider adding zero-padding the scans to equalize the shape across the training dataset. I looped over the dataset and found the smallest sizes you could squeeze the images by removing zero values which were:<br>\n<code>target_shapes = {\n        'FLAIR': (538, 477),\n        'T1w': (445, 394),\n        'T1wCE': (424, 388),\n        'T2w': (797, 616)\n    }</code></p>\n<p>I know this isn't the best method, but it's a starting point if you're eager to get a benchmark for what a simple 2D CNN could do using a single scan (turns out not much better than guessing…).</p>\n<p>And the padding algorithm:</p>\n<pre><code>def normalize_array(array, target_shape):\n    while np.all(array.shape != target_shape):\n        bool_array = array.shape &lt; np.array(target_shape)\n        for i, small in enumerate(bool_array):\n            # if axis i is smaller than the target shape, padding should be applied across axis i\n            if small:\n                dx = target_shape[i] - array.shape[i]\n                # if adding zeros to dimension 0, then padding vectors should match along dimension 1\n                if i == 0:\n                    array = np.concatenate([np.zeros((dx // 2, array.shape[1])), array,\n                                            np.zeros((dx // 2 + int(dx % 2 != 0), array.shape[1]))], axis=i)\n                # and vice versa\n                elif i == 1:\n                    array = np.concatenate([np.zeros((array.shape[0], dx // 2)), array,\n                                            np.zeros((array.shape[0], dx // 2 + int(dx % 2 != 0)))], axis=i)\n\n            # if axis i is larger than the target shape, cut off the shape\n            elif not small:\n                dx = array.shape[i] - target_shape[i]\n                if i == 0:\n                    array = array[dx // 2:-(dx // 2 + int(dx % 2 != 0))]\n                else:\n                    array = array[:, dx // 2:-(dx // 2 + int(dx % 2 != 0))]\n    return array\n</code></pre>\n<p>Other ideas such as factoring the orientation (coronal, transverse, sagittal) in as a feature are also ideas worth considering. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1498723,
          "author_name": "sureshsuri",
          "author_url": "",
          "post_date": "09/01/2021 06:44:46",
          "content": "<p>Thanks Jasper</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1500301,
      "author_name": "experienceinai",
      "author_url": "",
      "post_date": "09/02/2021 09:10:15",
      "content": "<p>HI suri, this notebook: <a href=\"https://www.kaggle.com/experienceinai/aivo-vii-nopea\" target=\"_blank\">https://www.kaggle.com/experienceinai/aivo-vii-nopea</a> includes step-by-step description about data preparation and modelling. Including pictures and also possibility to show additional information if you need. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1497776": "Hi All,\nI am newbie to medical image analysis, I just started this competition today .\nI want to through the some of the notebooks, but I didn't understand .\nplease share me any blog or video for data preparation for model.\nplease help me , I would like to learn this things.",
    "1498065": "I'm unsure what others have been doing so this is definitely not the best method but you can consider adding zero-padding the scans to equalize the shape across the training dataset. I looped over the dataset and found the smallest sizes you could squeeze the images by removing zero values which were:\n`target_shapes = {\n        'FLAIR': (538, 477),\n        'T1w': (445, 394),\n        'T1wCE': (424, 388),\n        'T2w': (797, 616)\n    }`\n\nI know this isn't the best method, but it's a starting point if you're eager to get a benchmark for what a simple 2D CNN could do using a single scan (turns out not much better than guessing...).\n\nAnd the padding algorithm:\n```\ndef normalize_array(array, target_shape):\n    while np.all(array.shape != target_shape):\n        bool_array = array.shape < np.array(target_shape)\n        for i, small in enumerate(bool_array):\n            # if axis i is smaller than the target shape, padding should be applied across axis i\n            if small:\n                dx = target_shape[i] - array.shape[i]\n                # if adding zeros to dimension 0, then padding vectors should match along dimension 1\n                if i == 0:\n                    array = np.concatenate([np.zeros((dx // 2, array.shape[1])), array,\n                                            np.zeros((dx // 2 + int(dx % 2 != 0), array.shape[1]))], axis=i)\n                # and vice versa\n                elif i == 1:\n                    array = np.concatenate([np.zeros((array.shape[0], dx // 2)), array,\n                                            np.zeros((array.shape[0], dx // 2 + int(dx % 2 != 0)))], axis=i)\n\n            # if axis i is larger than the target shape, cut off the shape\n            elif not small:\n                dx = array.shape[i] - target_shape[i]\n                if i == 0:\n                    array = array[dx // 2:-(dx // 2 + int(dx % 2 != 0))]\n                else:\n                    array = array[:, dx // 2:-(dx // 2 + int(dx % 2 != 0))]\n    return array\n```\n\nOther ideas such as factoring the orientation (coronal, transverse, sagittal) in as a feature are also ideas worth considering.",
    "1498723": "Thanks Jasper",
    "1500301": "HI suri, this notebook: https://www.kaggle.com/experienceinai/aivo-vii-nopea includes step-by-step description about data preparation and modelling. Including pictures and also possibility to show additional information if you need."
  },
  "source": "meta"
}