{
  "id": 51245,
  "title": "Do you run code on cloud?",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/51245",
  "author_name": "",
  "post_date": "2018-03-07T01:11:00.679435Z",
  "votes": 10,
  "comment_count": 16,
  "views": 0,
  "content": "<p>Hi guys,</p>\n\n<p>The size of train data is quite big and I'm highly doubt I could implement any ml algorithms on such a big train data set on normal laptops. Where do you guys run your code? On cloud or powerful desktop? Thank you very much for you help!</p>",
  "messages": [
    {
      "id": "291838",
      "postDate": "03/07/2018 01:11:00",
      "content": "<p>Hi guys,</p>\n\n<p>The size of train data is quite big and I'm highly doubt I could implement any ml algorithms on such a big train data set on normal laptops. Where do you guys run your code? On cloud or powerful desktop? Thank you very much for you help!</p>",
      "rawMarkdown": "Hi guys,\n\nThe size of train data is quite big and I'm highly doubt I could implement any ml algorithms on such a big train data set on normal laptops. Where do you guys run your code? On cloud or powerful desktop? Thank you very much for you help!",
      "votes": null
    },
    {
      "id": "291945",
      "postDate": "03/07/2018 05:49:01",
      "content": "<p>Second the question.  I'm debating if I should even try this one, as I only have a macbook at my disposal...</p>",
      "rawMarkdown": "Second the question.  I'm debating if I should even try this one, as I only have a macbook at my disposal...",
      "votes": null
    },
    {
      "id": "291946",
      "postDate": "03/07/2018 05:50:50",
      "content": "<p>I am building my model with Spark ML and Scala. </p>",
      "rawMarkdown": "I am building my model with Spark ML and Scala.",
      "votes": null
    },
    {
      "id": "292000",
      "postDate": "03/07/2018 07:43:55",
      "content": "<p>I'm in the same position, early 2011 MBP with 16 GB RAM. I used the train sample set to get a feel for the data, and I'm now in the process of scaling that up a bit using the <code>train</code> set, but only importing part of it:</p>\n\n<p><code>train &lt;- read.csv(\"train.csv\", nrows = 1000000)</code></p>\n\n<p>I then plan on doing the next stage of model building in the cloud.</p>",
      "rawMarkdown": "I'm in the same position, early 2011 MBP with 16 GB RAM. I used the train sample set to get a feel for the data, and I'm now in the process of scaling that up a bit using the `train` set, but only importing part of it:\n\n`train &lt;- read.csv(\"train.csv\", nrows = 1000000)`\n\nI then plan on doing the next stage of model building in the cloud.",
      "votes": null
    },
    {
      "id": "292124",
      "postDate": "03/07/2018 13:57:07",
      "content": "<p>Thank you very much!</p>",
      "rawMarkdown": "Thank you very much!",
      "votes": null
    },
    {
      "id": "292125",
      "postDate": "03/07/2018 13:57:19",
      "content": "<p>Thank you thank you!</p>",
      "rawMarkdown": "Thank you thank you!",
      "votes": null
    },
    {
      "id": "292169",
      "postDate": "03/07/2018 15:10:54",
      "content": "<p>Does anyone know given the size of train data(say 8 GB), what is the rough requirement on memory to run common algorithms on 8GB data such as logistic regression, xgboost, lightgbm. etc. ?</p>",
      "rawMarkdown": "Does anyone know given the size of train data(say 8 GB), what is the rough requirement on memory to run common algorithms on 8GB data such as logistic regression, xgboost, lightgbm. etc. ?",
      "votes": null
    },
    {
      "id": "292303",
      "postDate": "03/07/2018 19:30:36",
      "content": "<p>I haven't worked with the data yet, but I'd definitely recommend subsetting the data for preliminary feature engineering and modeling (say to ~1 million records). Maybe oversample the rare class (you'll likely want to do this anyway) but take a small sample from the majority class. This way you can get a great feel for the data, build/validate a feature engineering pipeline, and get faster insight into what models are likely to work well before scaling to massive compute power. Hyperparameters that work for a representative sample of the data will likely work well for the entire data, so it may not be worth it to waste time running a ton of expensive searches on the full data set.</p>",
      "rawMarkdown": "I haven't worked with the data yet, but I'd definitely recommend subsetting the data for preliminary feature engineering and modeling (say to ~1 million records). Maybe oversample the rare class (you'll likely want to do this anyway) but take a small sample from the majority class. This way you can get a great feel for the data, build/validate a feature engineering pipeline, and get faster insight into what models are likely to work well before scaling to massive compute power. Hyperparameters that work for a representative sample of the data will likely work well for the entire data, so it may not be worth it to waste time running a ton of expensive searches on the full data set.",
      "votes": null
    },
    {
      "id": "292339",
      "postDate": "03/07/2018 20:54:33",
      "content": "<p>Here is a great <a href=\"https://www.kaggle.com/arjanso/reducing-dataframe-memory-size-by-65/notebook\">kernel</a><a href=\"https://www.kaggle.com/arjanso\">ArjanGroen</a> on how to reduce data-frame size by 65% when loading with pandas</p>",
      "rawMarkdown": "Here is a great [kernel][1][ArjanGroen][2] on how to reduce data-frame size by 65% when loading with pandas\n\n\n  [1]: https://www.kaggle.com/arjanso/reducing-dataframe-memory-size-by-65/notebook\n  [2]: https://www.kaggle.com/arjanso",
      "votes": null
    },
    {
      "id": "292402",
      "postDate": "03/07/2018 22:27:59",
      "content": "<p>You could look into cloud based virtual machines for running your code. Paperpace, AWS or Google Cloud all have great support for Ubuntu + Jupyter Notebook solutions.</p>",
      "rawMarkdown": "You could look into cloud based virtual machines for running your code. Paperpace, AWS or Google Cloud all have great support for Ubuntu + Jupyter Notebook solutions.",
      "votes": null
    },
    {
      "id": "292619",
      "postDate": "03/08/2018 09:13:06",
      "content": "<p>What if I try AWS server, they provide cloud computing service for trial purpose.</p>",
      "rawMarkdown": "What if I try AWS server, they provide cloud computing service for trial purpose.",
      "votes": null
    },
    {
      "id": "292855",
      "postDate": "03/08/2018 19:28:22",
      "content": "<p>Google Cloud gives you $300 credit. You can get something like 24 cores and 120 GB of ram, if you want. Just make sure not to leave it idle.</p>",
      "rawMarkdown": "Google Cloud gives you $300 credit. You can get something like 24 cores and 120 GB of ram, if you want. Just make sure not to leave it idle.",
      "votes": null
    },
    {
      "id": "292947",
      "postDate": "03/08/2018 23:47:42",
      "content": "<p>It seems more than 16G is required. I ran xgb &amp; lgb(full data, 4 features) on the kernal(17G+), and oops! OOM...</p>",
      "rawMarkdown": "It seems more than 16G is required. I ran xgb &amp; lgb(full data, 4 features) on the kernal(17G+), and oops! OOM...",
      "votes": null
    },
    {
      "id": "292960",
      "postDate": "03/09/2018 00:12:46",
      "content": "<p>My approach - for now - is with online models with Vowpal Wabbit. If you fancy giving VW a shot, here's a script for preparing the input file in the necessary format:</p>\n\n<p><a href=\"https://www.kaggle.com/konradb/vowpal-wabbit-input-file-preparation\">https://www.kaggle.com/konradb/vowpal-wabbit-input-file-preparation</a></p>",
      "rawMarkdown": "My approach - for now - is with online models with Vowpal Wabbit. If you fancy giving VW a shot, here's a script for preparing the input file in the necessary format:\n\nhttps://www.kaggle.com/konradb/vowpal-wabbit-input-file-preparation",
      "votes": null
    },
    {
      "id": "292967",
      "postDate": "03/09/2018 00:28:51",
      "content": "<p>You might wanna try <strong><a href=\"https://databricks.com/aws\">databricks</a></strong> . They will take care of the underlying RAM requirements and all that you have to worry about are the two ends, the data and your code. I hope this helps. Cheers!</p>",
      "rawMarkdown": "You might wanna try **[databricks](https://databricks.com/aws)** . They will take care of the underlying RAM requirements and all that you have to worry about are the two ends, the data and your code. I hope this helps. Cheers!",
      "votes": null
    },
    {
      "id": "311310",
      "postDate": "04/09/2018 20:09:05",
      "content": "<p>I did try Google Collab . No cookie there . I think that the best bet is sampling the data and then averaging models.</p>",
      "rawMarkdown": "I did try Google Collab . No cookie there . I think that the best bet is sampling the data and then averaging models.",
      "votes": null
    },
    {
      "id": "311346",
      "postDate": "04/09/2018 22:23:48",
      "content": "<p>Thanks for this great reference, have you tried using this in your work ? If yes what is the eventual size of train based on this implementation.</p>",
      "rawMarkdown": "Thanks for this great reference, have you tried using this in your work ? If yes what is the eventual size of train based on this implementation.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 291945,
      "author_name": "yuliagm",
      "author_url": "",
      "post_date": "03/07/2018 05:49:01",
      "content": "<p>Second the question.  I'm debating if I should even try this one, as I only have a macbook at my disposal...</p>",
      "votes": null,
      "replies": [
        {
          "id": 292619,
          "author_name": "litgupta",
          "author_url": "",
          "post_date": "03/08/2018 09:13:06",
          "content": "<p>What if I try AWS server, they provide cloud computing service for trial purpose.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 311310,
          "author_name": "mayanksoni",
          "author_url": "",
          "post_date": "04/09/2018 20:09:05",
          "content": "<p>I did try Google Collab . No cookie there . I think that the best bet is sampling the data and then averaging models.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 291946,
      "author_name": "jiegzhan",
      "author_url": "",
      "post_date": "03/07/2018 05:50:50",
      "content": "<p>I am building my model with Spark ML and Scala. </p>",
      "votes": null,
      "replies": [
        {
          "id": 292124,
          "author_name": "wythhh",
          "author_url": "",
          "post_date": "03/07/2018 13:57:07",
          "content": "<p>Thank you very much!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 292000,
      "author_name": "chrisbow",
      "author_url": "",
      "post_date": "03/07/2018 07:43:55",
      "content": "<p>I'm in the same position, early 2011 MBP with 16 GB RAM. I used the train sample set to get a feel for the data, and I'm now in the process of scaling that up a bit using the <code>train</code> set, but only importing part of it:</p>\n\n<p><code>train &lt;- read.csv(\"train.csv\", nrows = 1000000)</code></p>\n\n<p>I then plan on doing the next stage of model building in the cloud.</p>",
      "votes": null,
      "replies": [
        {
          "id": 292125,
          "author_name": "wythhh",
          "author_url": "",
          "post_date": "03/07/2018 13:57:19",
          "content": "<p>Thank you thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 292169,
      "author_name": "wythhh",
      "author_url": "",
      "post_date": "03/07/2018 15:10:54",
      "content": "<p>Does anyone know given the size of train data(say 8 GB), what is the rough requirement on memory to run common algorithms on 8GB data such as logistic regression, xgboost, lightgbm. etc. ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 292947,
          "author_name": "laevatein",
          "author_url": "",
          "post_date": "03/08/2018 23:47:42",
          "content": "<p>It seems more than 16G is required. I ran xgb &amp; lgb(full data, 4 features) on the kernal(17G+), and oops! OOM...</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 292303,
      "author_name": "aquatic",
      "author_url": "",
      "post_date": "03/07/2018 19:30:36",
      "content": "<p>I haven't worked with the data yet, but I'd definitely recommend subsetting the data for preliminary feature engineering and modeling (say to ~1 million records). Maybe oversample the rare class (you'll likely want to do this anyway) but take a small sample from the majority class. This way you can get a great feel for the data, build/validate a feature engineering pipeline, and get faster insight into what models are likely to work well before scaling to massive compute power. Hyperparameters that work for a representative sample of the data will likely work well for the entire data, so it may not be worth it to waste time running a ton of expensive searches on the full data set.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 292339,
      "author_name": "sohaibomar",
      "author_url": "",
      "post_date": "03/07/2018 20:54:33",
      "content": "<p>Here is a great <a href=\"https://www.kaggle.com/arjanso/reducing-dataframe-memory-size-by-65/notebook\">kernel</a><a href=\"https://www.kaggle.com/arjanso\">ArjanGroen</a> on how to reduce data-frame size by 65% when loading with pandas</p>",
      "votes": null,
      "replies": [
        {
          "id": 311346,
          "author_name": "shivrajp",
          "author_url": "",
          "post_date": "04/09/2018 22:23:48",
          "content": "<p>Thanks for this great reference, have you tried using this in your work ? If yes what is the eventual size of train based on this implementation.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 292402,
      "author_name": "factorwonk",
      "author_url": "",
      "post_date": "03/07/2018 22:27:59",
      "content": "<p>You could look into cloud based virtual machines for running your code. Paperpace, AWS or Google Cloud all have great support for Ubuntu + Jupyter Notebook solutions.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 292855,
      "author_name": "bridgeport",
      "author_url": "",
      "post_date": "03/08/2018 19:28:22",
      "content": "<p>Google Cloud gives you $300 credit. You can get something like 24 cores and 120 GB of ram, if you want. Just make sure not to leave it idle.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 292960,
      "author_name": "konradb",
      "author_url": "",
      "post_date": "03/09/2018 00:12:46",
      "content": "<p>My approach - for now - is with online models with Vowpal Wabbit. If you fancy giving VW a shot, here's a script for preparing the input file in the necessary format:</p>\n\n<p><a href=\"https://www.kaggle.com/konradb/vowpal-wabbit-input-file-preparation\">https://www.kaggle.com/konradb/vowpal-wabbit-input-file-preparation</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 292967,
      "author_name": "zerocoder",
      "author_url": "",
      "post_date": "03/09/2018 00:28:51",
      "content": "<p>You might wanna try <strong><a href=\"https://databricks.com/aws\">databricks</a></strong> . They will take care of the underlying RAM requirements and all that you have to worry about are the two ends, the data and your code. I hope this helps. Cheers!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "291838": "Hi guys,\n\nThe size of train data is quite big and I'm highly doubt I could implement any ml algorithms on such a big train data set on normal laptops. Where do you guys run your code? On cloud or powerful desktop? Thank you very much for you help!",
    "291945": "Second the question.  I'm debating if I should even try this one, as I only have a macbook at my disposal...",
    "291946": "I am building my model with Spark ML and Scala.",
    "292000": "I'm in the same position, early 2011 MBP with 16 GB RAM. I used the train sample set to get a feel for the data, and I'm now in the process of scaling that up a bit using the `train` set, but only importing part of it:\n\n`train &lt;- read.csv(\"train.csv\", nrows = 1000000)`\n\nI then plan on doing the next stage of model building in the cloud.",
    "292124": "Thank you very much!",
    "292125": "Thank you thank you!",
    "292169": "Does anyone know given the size of train data(say 8 GB), what is the rough requirement on memory to run common algorithms on 8GB data such as logistic regression, xgboost, lightgbm. etc. ?",
    "292303": "I haven't worked with the data yet, but I'd definitely recommend subsetting the data for preliminary feature engineering and modeling (say to ~1 million records). Maybe oversample the rare class (you'll likely want to do this anyway) but take a small sample from the majority class. This way you can get a great feel for the data, build/validate a feature engineering pipeline, and get faster insight into what models are likely to work well before scaling to massive compute power. Hyperparameters that work for a representative sample of the data will likely work well for the entire data, so it may not be worth it to waste time running a ton of expensive searches on the full data set.",
    "292339": "Here is a great [kernel][1][ArjanGroen][2] on how to reduce data-frame size by 65% when loading with pandas\n\n\n  [1]: https://www.kaggle.com/arjanso/reducing-dataframe-memory-size-by-65/notebook\n  [2]: https://www.kaggle.com/arjanso",
    "292402": "You could look into cloud based virtual machines for running your code. Paperpace, AWS or Google Cloud all have great support for Ubuntu + Jupyter Notebook solutions.",
    "292619": "What if I try AWS server, they provide cloud computing service for trial purpose.",
    "292855": "Google Cloud gives you $300 credit. You can get something like 24 cores and 120 GB of ram, if you want. Just make sure not to leave it idle.",
    "292947": "It seems more than 16G is required. I ran xgb &amp; lgb(full data, 4 features) on the kernal(17G+), and oops! OOM...",
    "292960": "My approach - for now - is with online models with Vowpal Wabbit. If you fancy giving VW a shot, here's a script for preparing the input file in the necessary format:\n\nhttps://www.kaggle.com/konradb/vowpal-wabbit-input-file-preparation",
    "292967": "You might wanna try **[databricks](https://databricks.com/aws)** . They will take care of the underlying RAM requirements and all that you have to worry about are the two ends, the data and your code. I hope this helps. Cheers!",
    "311310": "I did try Google Collab . No cookie there . I think that the best bet is sampling the data and then averaging models.",
    "311346": "Thanks for this great reference, have you tried using this in your work ? If yes what is the eventual size of train based on this implementation."
  },
  "source": "meta"
}