{
  "id": 91408,
  "title": "Cloud computing",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/91408",
  "author_name": "",
  "post_date": "2019-05-04T11:15:52.957843700Z",
  "votes": 1,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I think for vast majority of competitors computing resources are rather limited. \nUsual go to (atleast for me) is AWS, since they are almost free for the first year. But they have 10GB limit which does not allow to upload everything. I tried to look around for some alternatives but its only worse.</p>\n\n<p>Now, you do get what you pay for but my question is, are there some <strong>cheap alternatives to AWS out there (excluding google, ibm etc) that have sufficient space?</strong></p>",
  "messages": [
    {
      "id": "527009",
      "postDate": "05/04/2019 11:15:52",
      "content": "<p>I think for vast majority of competitors computing resources are rather limited. \nUsual go to (atleast for me) is AWS, since they are almost free for the first year. But they have 10GB limit which does not allow to upload everything. I tried to look around for some alternatives but its only worse.</p>\n\n<p>Now, you do get what you pay for but my question is, are there some <strong>cheap alternatives to AWS out there (excluding google, ibm etc) that have sufficient space?</strong></p>",
      "rawMarkdown": "I think for vast majority of competitors computing resources are rather limited. \nUsual go to (atleast for me) is AWS, since they are almost free for the first year. But they have 10GB limit which does not allow to upload everything. I tried to look around for some alternatives but its only worse.\n\nNow, you do get what you pay for but my question is, are there some **cheap alternatives to AWS out there (excluding google, ibm etc) that have sufficient space?**",
      "votes": null
    },
    {
      "id": "527085",
      "postDate": "05/04/2019 14:48:32",
      "content": "<p>First cheap alternative is to compress the training data set.  I like to pickle the file - it reads so much faster on all those runs I make and it's smaller.  So short script on local to create pickle versions of large files, using float32 or smaller when you can.  </p>\n\n<p>Than put the pickle on AWS.  My pickle training file is 4,915,200 KB.</p>",
      "rawMarkdown": "First cheap alternative is to compress the training data set.  I like to pickle the file - it reads so much faster on all those runs I make and it's smaller.  So short script on local to create pickle versions of large files, using float32 or smaller when you can.  \n\nThan put the pickle on AWS.  My pickle training file is 4,915,200 KB.",
      "votes": null
    },
    {
      "id": "527162",
      "postDate": "05/04/2019 18:31:25",
      "content": "<p>Yup, 3.5 GB</p>\n\n<p>To complete the question, just use </p>\n\n<p>df.to_pickle('train.pickle')</p>",
      "rawMarkdown": "Yup, 3.5 GB\n\nTo complete the question, just use \n\ndf.to_pickle('train.pickle')",
      "votes": null
    },
    {
      "id": "527234",
      "postDate": "05/04/2019 21:14:54",
      "content": "<p>Google Colab works fine.</p>",
      "rawMarkdown": "Google Colab works fine.",
      "votes": null
    },
    {
      "id": "527258",
      "postDate": "05/05/2019 00:10:23",
      "content": "<p>In the Microsoft Malware competition I used over $150 of Google Cloud's free $300 trial for training large models that wouldn't fit on Kaggle. It was very tempting due to the power available online. But ultimately, my scores on these models were worse. What I learned is that our time is better spent refining CV strategy, selecting features and trying to understand the data better. Only when you are completely sure of what you need to do, and why, and when you are certain that Kaggle is not an option, should you use cloud resources.</p>\n\n<p>I'm not lecturing you - this is just the lesson I learned from that experience. As the saying goes, work smart - not hard.</p>",
      "rawMarkdown": "In the Microsoft Malware competition I used over $150 of Google Cloud's free $300 trial for training large models that wouldn't fit on Kaggle. It was very tempting due to the power available online. But ultimately, my scores on these models were worse. What I learned is that our time is better spent refining CV strategy, selecting features and trying to understand the data better. Only when you are completely sure of what you need to do, and why, and when you are certain that Kaggle is not an option, should you use cloud resources.\n\nI'm not lecturing you - this is just the lesson I learned from that experience. As the saying goes, work smart - not hard.",
      "votes": null
    },
    {
      "id": "527427",
      "postDate": "05/05/2019 13:25:57",
      "content": "<p>I posted an alternate to pickling the train through feather format. </p>\n\n<p><a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/90330\">https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/90330</a></p>",
      "rawMarkdown": "I posted an alternate to pickling the train through feather format. \n\nhttps://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/90330",
      "votes": null
    },
    {
      "id": "527497",
      "postDate": "05/05/2019 17:02:41",
      "content": "<p>use this UNDERRATED kernel <a href=\"https://www.kaggle.com/friedchips/how-to-reduce-the-training-data-to-400mb\">https://www.kaggle.com/friedchips/how-to-reduce-the-training-data-to-400mb</a> from <a href=\"/friedchips\">@friedchips</a></p>",
      "rawMarkdown": "use this UNDERRATED kernel https://www.kaggle.com/friedchips/how-to-reduce-the-training-data-to-400mb from @friedchips",
      "votes": null
    },
    {
      "id": "527585",
      "postDate": "05/05/2019 21:53:42",
      "content": "<p>Great advice. I also try to make things work in Kaggle kernels first. That forces me to be efficient in feature selection/engineering and also in computing resource management. Once things are working well in Kaggle kernels, I then move into the cloud for more power. </p>",
      "rawMarkdown": "Great advice. I also try to make things work in Kaggle kernels first. That forces me to be efficient in feature selection/engineering and also in computing resource management. Once things are working well in Kaggle kernels, I then move into the cloud for more power.",
      "votes": null
    },
    {
      "id": "527602",
      "postDate": "05/05/2019 23:04:17",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> are u gonna attend this competition? </p>",
      "rawMarkdown": "cdeotte are u gonna attend this competition?",
      "votes": null
    },
    {
      "id": "527607",
      "postDate": "05/05/2019 23:25:09",
      "content": "<p>Probably not. I will be traveling a lot this month and away from a computer for some of it too.</p>",
      "rawMarkdown": "Probably not. I will be traveling a lot this month and away from a computer for some of it too.",
      "votes": null
    },
    {
      "id": "527610",
      "postDate": "05/05/2019 23:33:00",
      "content": "<p>Ah ok. I learned a lot from you in the Santander competition. I wish you could attend this one too. Anyways, good luck and I hope you have a wonderful month.</p>",
      "rawMarkdown": "Ah ok. I learned a lot from you in the Santander competition. I wish you could attend this one too. Anyways, good luck and I hope you have a wonderful month.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 527085,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "05/04/2019 14:48:32",
      "content": "<p>First cheap alternative is to compress the training data set.  I like to pickle the file - it reads so much faster on all those runs I make and it's smaller.  So short script on local to create pickle versions of large files, using float32 or smaller when you can.  </p>\n\n<p>Than put the pickle on AWS.  My pickle training file is 4,915,200 KB.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 527162,
      "author_name": "zikazika",
      "author_url": "",
      "post_date": "05/04/2019 18:31:25",
      "content": "<p>Yup, 3.5 GB</p>\n\n<p>To complete the question, just use </p>\n\n<p>df.to_pickle('train.pickle')</p>",
      "votes": null,
      "replies": [
        {
          "id": 527427,
          "author_name": "teeyee314",
          "author_url": "",
          "post_date": "05/05/2019 13:25:57",
          "content": "<p>I posted an alternate to pickling the train through feather format. </p>\n\n<p><a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/90330\">https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/90330</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 527234,
      "author_name": "hmcranbercourt",
      "author_url": "",
      "post_date": "05/04/2019 21:14:54",
      "content": "<p>Google Colab works fine.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 527258,
      "author_name": "bigironsphere",
      "author_url": "",
      "post_date": "05/05/2019 00:10:23",
      "content": "<p>In the Microsoft Malware competition I used over $150 of Google Cloud's free $300 trial for training large models that wouldn't fit on Kaggle. It was very tempting due to the power available online. But ultimately, my scores on these models were worse. What I learned is that our time is better spent refining CV strategy, selecting features and trying to understand the data better. Only when you are completely sure of what you need to do, and why, and when you are certain that Kaggle is not an option, should you use cloud resources.</p>\n\n<p>I'm not lecturing you - this is just the lesson I learned from that experience. As the saying goes, work smart - not hard.</p>",
      "votes": null,
      "replies": [
        {
          "id": 527585,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "05/05/2019 21:53:42",
          "content": "<p>Great advice. I also try to make things work in Kaggle kernels first. That forces me to be efficient in feature selection/engineering and also in computing resource management. Once things are working well in Kaggle kernels, I then move into the cloud for more power. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 527602,
          "author_name": "mhviraf",
          "author_url": "",
          "post_date": "05/05/2019 23:04:17",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> are u gonna attend this competition? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 527607,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "05/05/2019 23:25:09",
          "content": "<p>Probably not. I will be traveling a lot this month and away from a computer for some of it too.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 527610,
          "author_name": "mhviraf",
          "author_url": "",
          "post_date": "05/05/2019 23:33:00",
          "content": "<p>Ah ok. I learned a lot from you in the Santander competition. I wish you could attend this one too. Anyways, good luck and I hope you have a wonderful month.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 527497,
      "author_name": "mhviraf",
      "author_url": "",
      "post_date": "05/05/2019 17:02:41",
      "content": "<p>use this UNDERRATED kernel <a href=\"https://www.kaggle.com/friedchips/how-to-reduce-the-training-data-to-400mb\">https://www.kaggle.com/friedchips/how-to-reduce-the-training-data-to-400mb</a> from <a href=\"/friedchips\">@friedchips</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "527009": "I think for vast majority of competitors computing resources are rather limited. \nUsual go to (atleast for me) is AWS, since they are almost free for the first year. But they have 10GB limit which does not allow to upload everything. I tried to look around for some alternatives but its only worse.\n\nNow, you do get what you pay for but my question is, are there some **cheap alternatives to AWS out there (excluding google, ibm etc) that have sufficient space?**",
    "527085": "First cheap alternative is to compress the training data set.  I like to pickle the file - it reads so much faster on all those runs I make and it's smaller.  So short script on local to create pickle versions of large files, using float32 or smaller when you can.  \n\nThan put the pickle on AWS.  My pickle training file is 4,915,200 KB.",
    "527162": "Yup, 3.5 GB\n\nTo complete the question, just use \n\ndf.to_pickle('train.pickle')",
    "527234": "Google Colab works fine.",
    "527258": "In the Microsoft Malware competition I used over $150 of Google Cloud's free $300 trial for training large models that wouldn't fit on Kaggle. It was very tempting due to the power available online. But ultimately, my scores on these models were worse. What I learned is that our time is better spent refining CV strategy, selecting features and trying to understand the data better. Only when you are completely sure of what you need to do, and why, and when you are certain that Kaggle is not an option, should you use cloud resources.\n\nI'm not lecturing you - this is just the lesson I learned from that experience. As the saying goes, work smart - not hard.",
    "527427": "I posted an alternate to pickling the train through feather format. \n\nhttps://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/90330",
    "527497": "use this UNDERRATED kernel https://www.kaggle.com/friedchips/how-to-reduce-the-training-data-to-400mb from @friedchips",
    "527585": "Great advice. I also try to make things work in Kaggle kernels first. That forces me to be efficient in feature selection/engineering and also in computing resource management. Once things are working well in Kaggle kernels, I then move into the cloud for more power.",
    "527602": "cdeotte are u gonna attend this competition?",
    "527607": "Probably not. I will be traveling a lot this month and away from a computer for some of it too.",
    "527610": "Ah ok. I learned a lot from you in the Santander competition. I wish you could attend this one too. Anyways, good luck and I hope you have a wonderful month."
  },
  "source": "meta"
}