{
  "id": 56729,
  "title": "whats the cost-efficient way to run this big GB data?",
  "url": "/competitions/avito-demand-prediction/discussion/56729",
  "author_name": "Stephanie0001",
  "post_date": "2018-05-14T07:13:32.781000",
  "votes": 1,
  "comment_count": 6,
  "views": 0,
  "content": "<p>whats the cost-efficient way to run this big GB data? Do you use the Amazon or what else service to upload the data and run the script to compute? I don't even have sufficient room to download data on my mac. </p>",
  "messages": [
    {
      "id": 328750,
      "postDate": "2018-05-15T02:15:34.467Z",
      "content": "<ol>\n<li>As mentioned by everyone first use Kernel to explore. You can learn a lot and get a really good score without even using image data (e.g. my current score). 16gb RAM provided by kernels should mostly be fine for most modelling purposes in this competition.</li>\n<li>If you are running out of disk space, best option would be to buy a external hard drive (trust me they are really cheap compared to cloud services)</li>\n</ol>",
      "rawMarkdown": "1. As mentioned by everyone first use Kernel to explore. You can learn a lot and get a really good score without even using image data (e.g. my current score). 16gb RAM provided by kernels should mostly be fine for most modelling purposes in this competition.\n2. If you are running out of disk space, best option would be to buy a external hard drive (trust me they are really cheap compared to cloud services)",
      "votes": 1
    },
    {
      "id": 328389,
      "postDate": "2018-05-14T07:13:32.780Z",
      "content": "<p>whats the cost-efficient way to run this big GB data? Do you use the Amazon or what else service to upload the data and run the script to compute? I don't even have sufficient room to download data on my mac. </p>",
      "rawMarkdown": "whats the cost-efficient way to run this big GB data? Do you use the Amazon or what else service to upload the data and run the script to compute? I don't even have sufficient room to download data on my mac. ",
      "votes": 1
    },
    {
      "id": 328569,
      "postDate": "2018-05-14T16:18:58.813Z",
      "content": "<p>I'd recommend using kernels for your EDA. Once you're ready to train a model, Google Cloud offers $300 of free credits for new users. <a href=\"https://cloud.google.com/free/docs/frequently-asked-questions\">https://cloud.google.com/free/docs/frequently-asked-questions</a></p>",
      "rawMarkdown": "I'd recommend using kernels for your EDA. Once you're ready to train a model, Google Cloud offers $300 of free credits for new users. https://cloud.google.com/free/docs/frequently-asked-questions",
      "votes": 2
    },
    {
      "id": 329979,
      "postDate": "2018-05-17T18:31:58.503Z",
      "content": "<p>I also came across this site shared by a fellow Kaggle user - <a href=\"https://www.onepanel.io/\">OnePanel</a>. They offer 4 hours of GPU usage in beta.\nAzure offers free credits as well. </p>\n\n<p>Also, you could just start with a kernel. I started off recently on Kaggle so I know how this can be. :)</p>\n\n<p>As a specific suggestion, for this competition, you can start with just the basic \"train\" csv and build some basic models with it before trying to use \"train_periods\", \"train_active\"  </p>",
      "rawMarkdown": "I also came across this site shared by a fellow Kaggle user - [OnePanel][1]. They offer 4 hours of GPU usage in beta.\nAzure offers free credits as well. \n\nAlso, you could just start with a kernel. I started off recently on Kaggle so I know how this can be. :)\n\nAs a specific suggestion, for this competition, you can start with just the basic \"train\" csv and build some basic models with it before trying to use \"train_periods\", \"train_active\"  \n\n  [1]: https://www.onepanel.io/"
    },
    {
      "id": 329686,
      "postDate": "2018-05-17T01:36:14.200Z",
      "content": "<p>Don't download the data- use a kernel to run a pre-trained image model on the image set and download the high level activations. Then, download the activations and use those during your training. </p>",
      "rawMarkdown": "Don't download the data- use a kernel to run a pre-trained image model on the image set and download the high level activations. Then, download the activations and use those during your training. "
    },
    {
      "id": 328492,
      "postDate": "2018-05-14T13:17:19.937Z",
      "content": "<p>I think Amazon is the right thing. <br>\nYou could try AWS S3 to store the data and AWS SageMaker to run Jupyter Notebooks remotely, which you can securely access from your laptop.  </p>",
      "rawMarkdown": "I think Amazon is the right thing.  \nYou could try AWS S3 to store the data and AWS SageMaker to run Jupyter Notebooks remotely, which you can securely access from your laptop.  ",
      "isDeleted": true
    },
    {
      "id": 328460,
      "postDate": "2018-05-14T11:51:53.910Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 328750,
      "author_name": "Mohsin hasan",
      "author_url": "",
      "post_date": "2018-05-15T02:15:34.467000",
      "content": "<ol>\n<li>As mentioned by everyone first use Kernel to explore. You can learn a lot and get a really good score without even using image data (e.g. my current score). 16gb RAM provided by kernels should mostly be fine for most modelling purposes in this competition.</li>\n<li>If you are running out of disk space, best option would be to buy a external hard drive (trust me they are really cheap compared to cloud services)</li>\n</ol>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 328569,
      "author_name": "Sohier Dane",
      "author_url": "",
      "post_date": "2018-05-14T16:18:58.813000",
      "content": "<p>I'd recommend using kernels for your EDA. Once you're ready to train a model, Google Cloud offers $300 of free credits for new users. <a href=\"https://cloud.google.com/free/docs/frequently-asked-questions\">https://cloud.google.com/free/docs/frequently-asked-questions</a></p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 329979,
      "author_name": "Shanth",
      "author_url": "",
      "post_date": "2018-05-17T18:31:58.503000",
      "content": "<p>I also came across this site shared by a fellow Kaggle user - <a href=\"https://www.onepanel.io/\">OnePanel</a>. They offer 4 hours of GPU usage in beta.\nAzure offers free credits as well. </p>\n\n<p>Also, you could just start with a kernel. I started off recently on Kaggle so I know how this can be. :)</p>\n\n<p>As a specific suggestion, for this competition, you can start with just the basic \"train\" csv and build some basic models with it before trying to use \"train_periods\", \"train_active\"  </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 329686,
      "author_name": "Vannak",
      "author_url": "",
      "post_date": "2018-05-17T01:36:14.200000",
      "content": "<p>Don't download the data- use a kernel to run a pre-trained image model on the image set and download the high level activations. Then, download the activations and use those during your training. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 328492,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-14T13:17:19.937000",
      "content": "<p>I think Amazon is the right thing. <br>\nYou could try AWS S3 to store the data and AWS SageMaker to run Jupyter Notebooks remotely, which you can securely access from your laptop.  </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 328460,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-14T11:51:53.910000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "328750": "1. As mentioned by everyone first use Kernel to explore. You can learn a lot and get a really good score without even using image data (e.g. my current score). 16gb RAM provided by kernels should mostly be fine for most modelling purposes in this competition.\n2. If you are running out of disk space, best option would be to buy a external hard drive (trust me they are really cheap compared to cloud services)",
    "328389": "whats the cost-efficient way to run this big GB data? Do you use the Amazon or what else service to upload the data and run the script to compute? I don't even have sufficient room to download data on my mac. ",
    "328569": "I'd recommend using kernels for your EDA. Once you're ready to train a model, Google Cloud offers $300 of free credits for new users. https://cloud.google.com/free/docs/frequently-asked-questions",
    "329979": "I also came across this site shared by a fellow Kaggle user - [OnePanel][1]. They offer 4 hours of GPU usage in beta.\nAzure offers free credits as well. \n\nAlso, you could just start with a kernel. I started off recently on Kaggle so I know how this can be. :)\n\nAs a specific suggestion, for this competition, you can start with just the basic \"train\" csv and build some basic models with it before trying to use \"train_periods\", \"train_active\"  \n\n  [1]: https://www.onepanel.io/",
    "329686": "Don't download the data- use a kernel to run a pre-trained image model on the image set and download the high level activations. Then, download the activations and use those during your training. ",
    "328492": "I think Amazon is the right thing.  \nYou could try AWS S3 to store the data and AWS SageMaker to run Jupyter Notebooks remotely, which you can securely access from your laptop.  ",
    "328460": ""
  }
}