{
  "id": 167696,
  "title": "How to handle such large datasets?",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/167696",
  "author_name": "",
  "post_date": "2020-07-17T14:46:50.236152700Z",
  "votes": 5,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Attention Please!</p>\n\n<p>While participating in any of the competition over Kaggle first thing I look, is the size of the dataset. If the size is big enough like this competition then I simply Discard, even I want to participate. </p>\n\n<p>Like to know from, you guys about your experience, How you handle such a big Dataset. If you are using external machines, then what configurations are you using. I know it's bit tedious job to explain out everything but I request you, invest your little time and let me know so I can too participate and contribute.  </p>\n\n<h1>kagglecommunity</h1>",
  "messages": [
    {
      "id": "933178",
      "postDate": "07/17/2020 14:46:50",
      "content": "<p>Attention Please!</p>\n\n<p>While participating in any of the competition over Kaggle first thing I look, is the size of the dataset. If the size is big enough like this competition then I simply Discard, even I want to participate. </p>\n\n<p>Like to know from, you guys about your experience, How you handle such a big Dataset. If you are using external machines, then what configurations are you using. I know it's bit tedious job to explain out everything but I request you, invest your little time and let me know so I can too participate and contribute.  </p>\n\n<h1>kagglecommunity</h1>",
      "rawMarkdown": "Attention Please!\n\nWhile participating in any of the competition over Kaggle first thing I look, is the size of the dataset. If the size is big enough like this competition then I simply Discard, even I want to participate. \n\nLike to know from, you guys about your experience, How you handle such a big Dataset. If you are using external machines, then what configurations are you using. I know it's bit tedious job to explain out everything but I request you, invest your little time and let me know so I can too participate and contribute.  \n\n#kagglecommunity",
      "votes": null
    },
    {
      "id": "933197",
      "postDate": "07/17/2020 15:00:36",
      "content": "<p>You can use a notebook so that you don't have to download the data.\nIf run time is an issue, then I would recommend training on subsets. If there are tons of columns, I think dimensionality reduction is your friend.</p>",
      "rawMarkdown": "You can use a notebook so that you don't have to download the data.\nIf run time is an issue, then I would recommend training on subsets. If there are tons of columns, I think dimensionality reduction is your friend.",
      "votes": null
    },
    {
      "id": "933205",
      "postDate": "07/17/2020 15:05:29",
      "content": "<p>What issue I am facing while training a model it consumed my most of GPU hours within 2-3 days and I don't have an externel GPU. So it's only me who is facing this issue of others are also going through it.</p>",
      "rawMarkdown": "What issue I am facing while training a model it consumed my most of GPU hours within 2-3 days and I don't have an externel GPU. So it's only me who is facing this issue of others are also going through it.",
      "votes": null
    },
    {
      "id": "934400",
      "postDate": "07/18/2020 12:13:24",
      "content": "<p>I have also faced a similar issue. Yet isn't 60 hours provided by Kaggle enough(w/ TPU + GPU) enough, also you can use other free tiers in google Colabs and other platforms.</p>",
      "rawMarkdown": "I have also faced a similar issue. Yet isn't 60 hours provided by Kaggle enough(w/ TPU + GPU) enough, also you can use other free tiers in google Colabs and other platforms.",
      "votes": null
    },
    {
      "id": "934859",
      "postDate": "07/18/2020 20:13:24",
      "content": "<p>The first thing you can do is to optimize your algorithm in terms of speed. While optimizing use only a small portion of the data not to consume too much GPU. <br>\nAnother approach is to apply feature engineering on the data and extract the useful features. So the input of your model would be much smaller.<br>\nWhile running your model or EDA note the elapsed time for the first subset of the data and approximate the total required time. If it is satisfying keep going otherwise inspect your model and modify it.</p>",
      "rawMarkdown": "The first thing you can do is to optimize your algorithm in terms of speed. While optimizing use only a small portion of the data not to consume too much GPU. \nAnother approach is to apply feature engineering on the data and extract the useful features. So the input of your model would be much smaller.\nWhile running your model or EDA note the elapsed time for the first subset of the data and approximate the total required time. If it is satisfying keep going otherwise inspect your model and modify it.",
      "votes": null
    },
    {
      "id": "934898",
      "postDate": "07/18/2020 21:43:29",
      "content": "<p>Take time to look at your dataloader/datagenerator if you're using torch/tf. You can run loops of these without running the model to see how long it's taking and find out whether the bottleneck is there (more a cpu or ram issue) or in the model and backpropagation (then you'll have to look at some of the options here like using colab).</p>",
      "rawMarkdown": "Take time to look at your dataloader/datagenerator if you're using torch/tf. You can run loops of these without running the model to see how long it's taking and find out whether the bottleneck is there (more a cpu or ram issue) or in the model and backpropagation (then you'll have to look at some of the options here like using colab).",
      "votes": null
    },
    {
      "id": "935332",
      "postDate": "07/19/2020 09:41:19",
      "content": "<p>Already using Colab but feels it is very much slower compared to the Kaggle GPU.</p>",
      "rawMarkdown": "Already using Colab but feels it is very much slower compared to the Kaggle GPU.",
      "votes": null
    },
    {
      "id": "936209",
      "postDate": "07/20/2020 04:33:46",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "936210",
      "postDate": "07/20/2020 04:34:51",
      "content": "<p>Great piece of advice, surely gonna apply them in my notebook.</p>",
      "rawMarkdown": "Great piece of advice, surely gonna apply them in my notebook.",
      "votes": null
    },
    {
      "id": "936211",
      "postDate": "07/20/2020 04:35:44",
      "content": "<p>Could you suggest some better learning sources...</p>",
      "rawMarkdown": "Could you suggest some better learning sources...",
      "votes": null
    },
    {
      "id": "936411",
      "postDate": "07/20/2020 07:45:30",
      "content": "<p>Hi. I suggest to use JPG images : <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/164958\">https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/164958</a>\nthe size is less than 3GB so it's OK for colab (I don't know if using jpeg images has an impact on the accuracy for this competition). Each time you run a kernel on colab you can either download the dataset from kaggle or copy/paste from your drive. It will be quick, 3GB is not too big. \nRemember that colab free is limited to 12h per session and colab pro is 24h and cost 10$/months. \nPlus, colab pro generally provides faster GPUs than colab free. That is why the training is slower when using colab free</p>",
      "rawMarkdown": "Hi. I suggest to use JPG images : https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/164958\nthe size is less than 3GB so it's OK for colab (I don't know if using jpeg images has an impact on the accuracy for this competition). Each time you run a kernel on colab you can either download the dataset from kaggle or copy/paste from your drive. It will be quick, 3GB is not too big. \nRemember that colab free is limited to 12h per session and colab pro is 24h and cost 10$/months. \nPlus, colab pro generally provides faster GPUs than colab free. That is why the training is slower when using colab free",
      "votes": null
    },
    {
      "id": "936417",
      "postDate": "07/20/2020 07:50:24",
      "content": "<p><code>Colab pro</code> this sounds interesting surely gonna give it a try..</p>",
      "rawMarkdown": "```Colab pro``` this sounds interesting surely gonna give it a try..",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 934859,
      "author_name": "umeyrkl",
      "author_url": "",
      "post_date": "07/18/2020 20:13:24",
      "content": "<p>The first thing you can do is to optimize your algorithm in terms of speed. While optimizing use only a small portion of the data not to consume too much GPU. <br>\nAnother approach is to apply feature engineering on the data and extract the useful features. So the input of your model would be much smaller.<br>\nWhile running your model or EDA note the elapsed time for the first subset of the data and approximate the total required time. If it is satisfying keep going otherwise inspect your model and modify it.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 936209,
      "author_name": "vin1234",
      "author_url": "",
      "post_date": "07/20/2020 04:33:46",
      "content": "",
      "votes": null,
      "replies": []
    },
    {
      "id": 936210,
      "author_name": "vin1234",
      "author_url": "",
      "post_date": "07/20/2020 04:34:51",
      "content": "<p>Great piece of advice, surely gonna apply them in my notebook.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 933197,
      "author_name": "ttalbitt",
      "author_url": "",
      "post_date": "07/17/2020 15:00:36",
      "content": "<p>You can use a notebook so that you don't have to download the data.\nIf run time is an issue, then I would recommend training on subsets. If there are tons of columns, I think dimensionality reduction is your friend.</p>",
      "votes": null,
      "replies": [
        {
          "id": 933205,
          "author_name": "vin1234",
          "author_url": "",
          "post_date": "07/17/2020 15:05:29",
          "content": "<p>What issue I am facing while training a model it consumed my most of GPU hours within 2-3 days and I don't have an externel GPU. So it's only me who is facing this issue of others are also going through it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 934400,
          "author_name": "kurianbenoy",
          "author_url": "",
          "post_date": "07/18/2020 12:13:24",
          "content": "<p>I have also faced a similar issue. Yet isn't 60 hours provided by Kaggle enough(w/ TPU + GPU) enough, also you can use other free tiers in google Colabs and other platforms.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 935332,
          "author_name": "vin1234",
          "author_url": "",
          "post_date": "07/19/2020 09:41:19",
          "content": "<p>Already using Colab but feels it is very much slower compared to the Kaggle GPU.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 934898,
      "author_name": "jameschapman19",
      "author_url": "",
      "post_date": "07/18/2020 21:43:29",
      "content": "<p>Take time to look at your dataloader/datagenerator if you're using torch/tf. You can run loops of these without running the model to see how long it's taking and find out whether the bottleneck is there (more a cpu or ram issue) or in the model and backpropagation (then you'll have to look at some of the options here like using colab).</p>",
      "votes": null,
      "replies": [
        {
          "id": 936211,
          "author_name": "vin1234",
          "author_url": "",
          "post_date": "07/20/2020 04:35:44",
          "content": "<p>Could you suggest some better learning sources...</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 936411,
      "author_name": "alexj21",
      "author_url": "",
      "post_date": "07/20/2020 07:45:30",
      "content": "<p>Hi. I suggest to use JPG images : <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/164958\">https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/164958</a>\nthe size is less than 3GB so it's OK for colab (I don't know if using jpeg images has an impact on the accuracy for this competition). Each time you run a kernel on colab you can either download the dataset from kaggle or copy/paste from your drive. It will be quick, 3GB is not too big. \nRemember that colab free is limited to 12h per session and colab pro is 24h and cost 10$/months. \nPlus, colab pro generally provides faster GPUs than colab free. That is why the training is slower when using colab free</p>",
      "votes": null,
      "replies": [
        {
          "id": 936417,
          "author_name": "vin1234",
          "author_url": "",
          "post_date": "07/20/2020 07:50:24",
          "content": "<p><code>Colab pro</code> this sounds interesting surely gonna give it a try..</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "933178": "Attention Please!\n\nWhile participating in any of the competition over Kaggle first thing I look, is the size of the dataset. If the size is big enough like this competition then I simply Discard, even I want to participate. \n\nLike to know from, you guys about your experience, How you handle such a big Dataset. If you are using external machines, then what configurations are you using. I know it's bit tedious job to explain out everything but I request you, invest your little time and let me know so I can too participate and contribute.  \n\n#kagglecommunity",
    "933197": "You can use a notebook so that you don't have to download the data.\nIf run time is an issue, then I would recommend training on subsets. If there are tons of columns, I think dimensionality reduction is your friend.",
    "933205": "What issue I am facing while training a model it consumed my most of GPU hours within 2-3 days and I don't have an externel GPU. So it's only me who is facing this issue of others are also going through it.",
    "934400": "I have also faced a similar issue. Yet isn't 60 hours provided by Kaggle enough(w/ TPU + GPU) enough, also you can use other free tiers in google Colabs and other platforms.",
    "934859": "The first thing you can do is to optimize your algorithm in terms of speed. While optimizing use only a small portion of the data not to consume too much GPU. \nAnother approach is to apply feature engineering on the data and extract the useful features. So the input of your model would be much smaller.\nWhile running your model or EDA note the elapsed time for the first subset of the data and approximate the total required time. If it is satisfying keep going otherwise inspect your model and modify it.",
    "934898": "Take time to look at your dataloader/datagenerator if you're using torch/tf. You can run loops of these without running the model to see how long it's taking and find out whether the bottleneck is there (more a cpu or ram issue) or in the model and backpropagation (then you'll have to look at some of the options here like using colab).",
    "935332": "Already using Colab but feels it is very much slower compared to the Kaggle GPU.",
    "936209": "",
    "936210": "Great piece of advice, surely gonna apply them in my notebook.",
    "936211": "Could you suggest some better learning sources...",
    "936411": "Hi. I suggest to use JPG images : https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/164958\nthe size is less than 3GB so it's OK for colab (I don't know if using jpeg images has an impact on the accuracy for this competition). Each time you run a kernel on colab you can either download the dataset from kaggle or copy/paste from your drive. It will be quick, 3GB is not too big. \nRemember that colab free is limited to 12h per session and colab pro is 24h and cost 10$/months. \nPlus, colab pro generally provides faster GPUs than colab free. That is why the training is slower when using colab free",
    "936417": "```Colab pro``` this sounds interesting surely gonna give it a try.."
  },
  "source": "meta"
}