{
  "id": 240709,
  "title": "Where can I analyse 119Gb of data?",
  "url": "/competitions/siim-covid19-detection/discussion/240709",
  "author_name": "",
  "post_date": "2021-05-21T05:06:05.625725700Z",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi everyone,<br>\nExcuse me the question but I'm very happy if you can recommend me how can I attack these big data set (which cloud or desktop tools?).</p>\n<p>Excuse me,</p>\n<p>Thanks.</p>",
  "messages": [
    {
      "id": "1317010",
      "postDate": "05/21/2021 05:06:05",
      "content": "<p>Hi everyone,<br>\nExcuse me the question but I'm very happy if you can recommend me how can I attack these big data set (which cloud or desktop tools?).</p>\n<p>Excuse me,</p>\n<p>Thanks.</p>",
      "rawMarkdown": "Hi everyone,\nExcuse me the question but I'm very happy if you can recommend me how can I attack these big data set (which cloud or desktop tools?).\n\nExcuse me,\n\nThanks.",
      "votes": null
    },
    {
      "id": "1317274",
      "postDate": "05/21/2021 09:08:25",
      "content": "<p>Hey ChristianHidalgo,</p>\n<p>you can check out <a href=\"https://gradient.paperspace.com/\" target=\"_blank\">https://gradient.paperspace.com/</a>. They're offering 1-Click-Jupyter Notebooks with a variety of specifications, some of them are free some for a small amount of money. I'm not sure if the free ones can handle such large amounts of data. You should try it out. </p>\n<p>Best regards</p>",
      "rawMarkdown": "Hey ChristianHidalgo,\n\nyou can check out https://gradient.paperspace.com/. They're offering 1-Click-Jupyter Notebooks with a variety of specifications, some of them are free some for a small amount of money. I'm not sure if the free ones can handle such large amounts of data. You should try it out. \n\nBest regards",
      "votes": null
    },
    {
      "id": "1317278",
      "postDate": "05/21/2021 09:15:25",
      "content": "<p><a href=\"https://www.kaggle.com/cristianhidalgo\" target=\"_blank\">@cristianhidalgo</a>  I have heard that this big data problems are solved using hadoop which uses map reduce concept .</p>",
      "rawMarkdown": "cristianhidalgo  I have heard that this big data problems are solved using hadoop which uses map reduce concept .",
      "votes": null
    },
    {
      "id": "1317789",
      "postDate": "05/21/2021 17:43:14",
      "content": "<p>It's not as bad as it may look at first sight. You don't have to load the entire dataset into memory at once (but if you have &gt; 128–160 GB of RAM in your system, sure, go ahead =). Instead you'll operate on just a few images at a time, the exact number depending on what resources your system have. If you're using one of the major deep learning frameworks, it is easy to implement data loaders and then the frameworks take care of loading and unloading data as needed.</p>\n<p>The main thing is to have enough secondary storage, ideally an SSD but you can get away with a decent hard drive too.</p>\n<p>To further help things, you'll probably want to preprocess the data set before training your models, doing things such as lung segmentation and scaling the resulting images, further reducing the size of the dataset. This can be done right here on Kaggle and saved as a new dataset. People often post various kinds of smaller preprocessed datasets, so you might want to check if there already are some that will suit your needs.</p>",
      "rawMarkdown": "It's not as bad as it may look at first sight. You don't have to load the entire dataset into memory at once (but if you have > 128–160 GB of RAM in your system, sure, go ahead =). Instead you'll operate on just a few images at a time, the exact number depending on what resources your system have. If you're using one of the major deep learning frameworks, it is easy to implement data loaders and then the frameworks take care of loading and unloading data as needed.\n\nThe main thing is to have enough secondary storage, ideally an SSD but you can get away with a decent hard drive too.\n\nTo further help things, you'll probably want to preprocess the data set before training your models, doing things such as lung segmentation and scaling the resulting images, further reducing the size of the dataset. This can be done right here on Kaggle and saved as a new dataset. People often post various kinds of smaller preprocessed datasets, so you might want to check if there already are some that will suit your needs.",
      "votes": null
    },
    {
      "id": "1319591",
      "postDate": "05/23/2021 11:03:27",
      "content": "<p>Hi Cristian<br>\nActually the dataset is not that big, since it contains images.<br>\nImages are in DICOM format, that is a standard for Digital Radiology. <br>\nThe first step is to extract images from DICOM files and store in a compressed format (JPEG). In that step you can choose a smaller resolution (images are till to 4000x400, too big for Deep learning. For example you could start with 384x384.<br>\nI would use the TF TFrecord format. Then you need to use Kaggle TPU.</p>\n<p>To create the TFRecord you can do also on a local machine… it will take time. I would suggest using Linux Ubuntu. If you have a MacBook it is ok, don't know Windows.</p>",
      "rawMarkdown": "Hi Cristian\nActually the dataset is not that big, since it contains images.\nImages are in DICOM format, that is a standard for Digital Radiology. \nThe first step is to extract images from DICOM files and store in a compressed format (JPEG). In that step you can choose a smaller resolution (images are till to 4000x400, too big for Deep learning. For example you could start with 384x384.\nI would use the TF TFrecord format. Then you need to use Kaggle TPU.\n\nTo create the TFRecord you can do also on a local machine... it will take time. I would suggest using Linux Ubuntu. If you have a MacBook it is ok, don't know Windows.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1317274,
      "author_name": "schmidde",
      "author_url": "",
      "post_date": "05/21/2021 09:08:25",
      "content": "<p>Hey ChristianHidalgo,</p>\n<p>you can check out <a href=\"https://gradient.paperspace.com/\" target=\"_blank\">https://gradient.paperspace.com/</a>. They're offering 1-Click-Jupyter Notebooks with a variety of specifications, some of them are free some for a small amount of money. I'm not sure if the free ones can handle such large amounts of data. You should try it out. </p>\n<p>Best regards</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1317278,
      "author_name": "gauravsarkar",
      "author_url": "",
      "post_date": "05/21/2021 09:15:25",
      "content": "<p><a href=\"https://www.kaggle.com/cristianhidalgo\" target=\"_blank\">@cristianhidalgo</a>  I have heard that this big data problems are solved using hadoop which uses map reduce concept .</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1317789,
      "author_name": "christoffer",
      "author_url": "",
      "post_date": "05/21/2021 17:43:14",
      "content": "<p>It's not as bad as it may look at first sight. You don't have to load the entire dataset into memory at once (but if you have &gt; 128–160 GB of RAM in your system, sure, go ahead =). Instead you'll operate on just a few images at a time, the exact number depending on what resources your system have. If you're using one of the major deep learning frameworks, it is easy to implement data loaders and then the frameworks take care of loading and unloading data as needed.</p>\n<p>The main thing is to have enough secondary storage, ideally an SSD but you can get away with a decent hard drive too.</p>\n<p>To further help things, you'll probably want to preprocess the data set before training your models, doing things such as lung segmentation and scaling the resulting images, further reducing the size of the dataset. This can be done right here on Kaggle and saved as a new dataset. People often post various kinds of smaller preprocessed datasets, so you might want to check if there already are some that will suit your needs.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1319591,
      "author_name": "luigisaetta",
      "author_url": "",
      "post_date": "05/23/2021 11:03:27",
      "content": "<p>Hi Cristian<br>\nActually the dataset is not that big, since it contains images.<br>\nImages are in DICOM format, that is a standard for Digital Radiology. <br>\nThe first step is to extract images from DICOM files and store in a compressed format (JPEG). In that step you can choose a smaller resolution (images are till to 4000x400, too big for Deep learning. For example you could start with 384x384.<br>\nI would use the TF TFrecord format. Then you need to use Kaggle TPU.</p>\n<p>To create the TFRecord you can do also on a local machine… it will take time. I would suggest using Linux Ubuntu. If you have a MacBook it is ok, don't know Windows.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1317010": "Hi everyone,\nExcuse me the question but I'm very happy if you can recommend me how can I attack these big data set (which cloud or desktop tools?).\n\nExcuse me,\n\nThanks.",
    "1317274": "Hey ChristianHidalgo,\n\nyou can check out https://gradient.paperspace.com/. They're offering 1-Click-Jupyter Notebooks with a variety of specifications, some of them are free some for a small amount of money. I'm not sure if the free ones can handle such large amounts of data. You should try it out. \n\nBest regards",
    "1317278": "cristianhidalgo  I have heard that this big data problems are solved using hadoop which uses map reduce concept .",
    "1317789": "It's not as bad as it may look at first sight. You don't have to load the entire dataset into memory at once (but if you have > 128–160 GB of RAM in your system, sure, go ahead =). Instead you'll operate on just a few images at a time, the exact number depending on what resources your system have. If you're using one of the major deep learning frameworks, it is easy to implement data loaders and then the frameworks take care of loading and unloading data as needed.\n\nThe main thing is to have enough secondary storage, ideally an SSD but you can get away with a decent hard drive too.\n\nTo further help things, you'll probably want to preprocess the data set before training your models, doing things such as lung segmentation and scaling the resulting images, further reducing the size of the dataset. This can be done right here on Kaggle and saved as a new dataset. People often post various kinds of smaller preprocessed datasets, so you might want to check if there already are some that will suit your needs.",
    "1319591": "Hi Cristian\nActually the dataset is not that big, since it contains images.\nImages are in DICOM format, that is a standard for Digital Radiology. \nThe first step is to extract images from DICOM files and store in a compressed format (JPEG). In that step you can choose a smaller resolution (images are till to 4000x400, too big for Deep learning. For example you could start with 384x384.\nI would use the TF TFrecord format. Then you need to use Kaggle TPU.\n\nTo create the TFRecord you can do also on a local machine... it will take time. I would suggest using Linux Ubuntu. If you have a MacBook it is ok, don't know Windows."
  },
  "source": "meta"
}