{
  "id": 252909,
  "title": "How to load the data to make a model on it",
  "url": "/competitions/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/252909",
  "author_name": "",
  "post_date": "2021-07-14T07:20:42.498210400Z",
  "votes": 2,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hello, Kagglers</p>\n<p>I'm a beginner starting out my Deep learning journey. <br>\nThis dataset looks promising and firstly I can't able to load the data into a PyTorch loader to make a simple MLP on the data.</p>\n<p>As the data is structured into more subfolders how to use the specific folder and target associated with it to make predictions.</p>\n<p>EX: As in the train data, the first data point in <strong>00000</strong> as per the target.csv, but the <strong>00000</strong> itself contains four subfolders in it namely: <strong>FLAIR</strong>, <strong>T1w</strong>, <strong>T1wCE</strong>, <strong>T2w</strong> .</p>\n<p>How to even merge these all folders into one datapoint for an MLP or any other model. </p>",
  "messages": [
    {
      "id": "1387436",
      "postDate": "07/14/2021 07:20:42",
      "content": "<p>Hello, Kagglers</p>\n<p>I'm a beginner starting out my Deep learning journey. <br>\nThis dataset looks promising and firstly I can't able to load the data into a PyTorch loader to make a simple MLP on the data.</p>\n<p>As the data is structured into more subfolders how to use the specific folder and target associated with it to make predictions.</p>\n<p>EX: As in the train data, the first data point in <strong>00000</strong> as per the target.csv, but the <strong>00000</strong> itself contains four subfolders in it namely: <strong>FLAIR</strong>, <strong>T1w</strong>, <strong>T1wCE</strong>, <strong>T2w</strong> .</p>\n<p>How to even merge these all folders into one datapoint for an MLP or any other model. </p>",
      "rawMarkdown": "Hello, Kagglers\n\nI'm a beginner starting out my Deep learning journey. \nThis dataset looks promising and firstly I can't able to load the data into a PyTorch loader to make a simple MLP on the data.\n\nAs the data is structured into more subfolders how to use the specific folder and target associated with it to make predictions.\n\nEX: As in the train data, the first data point in **00000** as per the target.csv, but the **00000** itself contains four subfolders in it namely: **FLAIR**, **T1w**, **T1wCE**, **T2w** .\n\nHow to even merge these all folders into one datapoint for an MLP or any other model.",
      "votes": null
    },
    {
      "id": "1387492",
      "postDate": "07/14/2021 08:25:18",
      "content": "<p>You shouldn't merge all the data. Instead if you are going to use a DL model you should create a dataset generator. In the <strong>getitem</strong>() method you can pass the direction of the folder and load it in that part. That way you are only loading in RAM the needed data and no more. You should also preprocess the data in that method, that way you don't have to duplicate the dataset.</p>",
      "rawMarkdown": "You shouldn't merge all the data. Instead if you are going to use a DL model you should create a dataset generator. In the __getitem__() method you can pass the direction of the folder and load it in that part. That way you are only loading in RAM the needed data and no more. You should also preprocess the data in that method, that way you don't have to duplicate the dataset.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1387492,
      "author_name": "josepc",
      "author_url": "",
      "post_date": "07/14/2021 08:25:18",
      "content": "<p>You shouldn't merge all the data. Instead if you are going to use a DL model you should create a dataset generator. In the <strong>getitem</strong>() method you can pass the direction of the folder and load it in that part. That way you are only loading in RAM the needed data and no more. You should also preprocess the data in that method, that way you don't have to duplicate the dataset.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1387436": "Hello, Kagglers\n\nI'm a beginner starting out my Deep learning journey. \nThis dataset looks promising and firstly I can't able to load the data into a PyTorch loader to make a simple MLP on the data.\n\nAs the data is structured into more subfolders how to use the specific folder and target associated with it to make predictions.\n\nEX: As in the train data, the first data point in **00000** as per the target.csv, but the **00000** itself contains four subfolders in it namely: **FLAIR**, **T1w**, **T1wCE**, **T2w** .\n\nHow to even merge these all folders into one datapoint for an MLP or any other model.",
    "1387492": "You shouldn't merge all the data. Instead if you are going to use a DL model you should create a dataset generator. In the __getitem__() method you can pass the direction of the folder and load it in that part. That way you are only loading in RAM the needed data and no more. You should also preprocess the data in that method, that way you don't have to duplicate the dataset."
  },
  "source": "meta"
}