{
  "id": 51504,
  "title": "AWS EC2 advice for running algorithms on whole data",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/51504",
  "author_name": "",
  "post_date": "2018-03-09T16:24:13.036875600Z",
  "votes": 2,
  "comment_count": 16,
  "views": 0,
  "content": "<p>From the advice given under this topic, our group selected AWS EC2 instance to have a trial play with the whole data. We choose r4.8xlarge windows instance which as 32 CPU and 244GB memory with $3.6/hour.\nWe extract hour and minute from click_time and remove click_time and attributed_time and we noticed from spyder IDE that the memory usage is around 26%(almost 64GB) when we run 5-fold cv using lightgbm. I think this would be a kind of benchmark of memory usage for people who wants to run algorithms on whole data set. Using whole data set may not be a good idea but if you want, then 32GB or even 64GB is not enough for some complicated algorithms.</p>",
  "messages": [
    {
      "id": "293320",
      "postDate": "03/09/2018 16:24:13",
      "content": "<p>From the advice given under this topic, our group selected AWS EC2 instance to have a trial play with the whole data. We choose r4.8xlarge windows instance which as 32 CPU and 244GB memory with $3.6/hour.\nWe extract hour and minute from click_time and remove click_time and attributed_time and we noticed from spyder IDE that the memory usage is around 26%(almost 64GB) when we run 5-fold cv using lightgbm. I think this would be a kind of benchmark of memory usage for people who wants to run algorithms on whole data set. Using whole data set may not be a good idea but if you want, then 32GB or even 64GB is not enough for some complicated algorithms.</p>",
      "rawMarkdown": "From the advice given under this topic, our group selected AWS EC2 instance to have a trial play with the whole data. We choose r4.8xlarge windows instance which as 32 CPU and 244GB memory with $3.6/hour.\nWe extract hour and minute from click_time and remove click_time and attributed_time and we noticed from spyder IDE that the memory usage is around 26%(almost 64GB) when we run 5-fold cv using lightgbm. I think this would be a kind of benchmark of memory usage for people who wants to run algorithms on whole data set. Using whole data set may not be a good idea but if you want, then 32GB or even 64GB is not enough for some complicated algorithms.",
      "votes": null
    },
    {
      "id": "293527",
      "postDate": "03/10/2018 03:51:50",
      "content": "<p>Hi, Is there any other option $3.6/hr is costly for me and i want to use whole dataset.</p>",
      "rawMarkdown": "Hi, Is there any other option $3.6/hr is costly for me and i want to use whole dataset.",
      "votes": null
    },
    {
      "id": "293530",
      "postDate": "03/10/2018 04:08:47",
      "content": "<p>I'm using a 16GB RAM Mac doing some feature engineering and training a lightgbm model using 50% of the trainset.</p>",
      "rawMarkdown": "I'm using a 16GB RAM Mac doing some feature engineering and training a lightgbm model using 50% of the trainset.",
      "votes": null
    },
    {
      "id": "293533",
      "postDate": "03/10/2018 04:13:27",
      "content": "<p>It means you are on top using lightgbm and only 50% trainset. AMAZING<br></p>",
      "rawMarkdown": "It means you are on top using lightgbm and only 50% trainset. AMAZING<br>",
      "votes": null
    },
    {
      "id": "293535",
      "postDate": "03/10/2018 04:20:41",
      "content": "<p>Well, that's interesting and could I ask how is the memory usage when you run lightgbm? Did it use almost 100% of your RAM? Actually I forget the memory usage when I train lightgbm using whole dataset. But when I did 5 fold cv lightgbm on whole data set, the memory usage is around 64GB.</p>",
      "rawMarkdown": "Well, that's interesting and could I ask how is the memory usage when you run lightgbm? Did it use almost 100% of your RAM? Actually I forget the memory usage when I train lightgbm using whole dataset. But when I did 5 fold cv lightgbm on whole data set, the memory usage is around 64GB.",
      "votes": null
    },
    {
      "id": "294046",
      "postDate": "03/11/2018 06:39:50",
      "content": "<p>Amazing! However, for further FE or more complicated algorithms, 16GB RAM could be enough?</p>",
      "rawMarkdown": "Amazing! However, for further FE or more complicated algorithms, 16GB RAM could be enough?",
      "votes": null
    },
    {
      "id": "294211",
      "postDate": "03/11/2018 15:17:09",
      "content": "<p>You could try google cloud engine, they always have $300 coupons for starters.</p>",
      "rawMarkdown": "You could try google cloud engine, they always have $300 coupons for starters.",
      "votes": null
    },
    {
      "id": "294297",
      "postDate": "03/11/2018 17:46:34",
      "content": "<p>I trained using only 1 oof. Remember this is a time series challenge you can't do random kFold. I'm using all my RAM available and a huge SWAP file (most of my dataset goes to swap file). Since Mac SSDs are really  fast I can train a model in about 1 hour</p>",
      "rawMarkdown": "I trained using only 1 oof. Remember this is a time series challenge you can't do random kFold. I'm using all my RAM available and a huge SWAP file (most of my dataset goes to swap file). Since Mac SSDs are really  fast I can train a model in about 1 hour",
      "votes": null
    },
    {
      "id": "296776",
      "postDate": "03/15/2018 18:40:49",
      "content": "<pre><code>using 50% of the trainset.\n</code></pre>\n\n<p>May I get a hint on how did you select that 50% of train set? Thanks</p>",
      "rawMarkdown": "using 50% of the trainset.\n\nMay I get a hint on how did you select that 50% of train set? Thanks",
      "votes": null
    },
    {
      "id": "296972",
      "postDate": "03/16/2018 03:24:03",
      "content": "<p>I think ec2 spot instance is a better choice, much cheaper ... and I built my environment and download data on an external EBS to keep it long time availabel.</p>",
      "rawMarkdown": "I think ec2 spot instance is a better choice, much cheaper ... and I built my environment and download data on an external EBS to keep it long time availabel.",
      "votes": null
    },
    {
      "id": "298020",
      "postDate": "03/18/2018 14:54:26",
      "content": "<p>Amazing!!</p>",
      "rawMarkdown": "Amazing!!",
      "votes": null
    },
    {
      "id": "298287",
      "postDate": "03/19/2018 09:11:54",
      "content": "<p>Are you using R or Python? And most important: who have any hint about a tutorial/guide to use SWAP file in R using windows 10? Very interesting discovery!</p>",
      "rawMarkdown": "Are you using R or Python? And most important: who have any hint about a tutorial/guide to use SWAP file in R using windows 10? Very interesting discovery!",
      "votes": null
    },
    {
      "id": "298848",
      "postDate": "03/20/2018 04:44:59",
      "content": "<p>I am using a 32GB RAM and 64G virtual memory from my SSD (WIN10, python). Training 50%-60% of the train dataset may take around 1 hour.  </p>",
      "rawMarkdown": "I am using a 32GB RAM and 64G virtual memory from my SSD (WIN10, python). Training 50%-60% of the train dataset may take around 1 hour.",
      "votes": null
    },
    {
      "id": "300415",
      "postDate": "03/21/2018 15:46:43",
      "content": "<p>I'm thinking if we can change the data type from int64 to int32 (or int16), I mean except the time column. Then we could get a somehow smaller data set.</p>",
      "rawMarkdown": "I'm thinking if we can change the data type from int64 to int32 (or int16), I mean except the time column. Then we could get a somehow smaller data set.",
      "votes": null
    },
    {
      "id": "300421",
      "postDate": "03/21/2018 15:51:09",
      "content": "<p>actually you can drop the time column, only left day, hour in int</p>",
      "rawMarkdown": "actually you can drop the time column, only left day, hour in int",
      "votes": null
    },
    {
      "id": "300432",
      "postDate": "03/21/2018 15:54:50",
      "content": "<p>Good point, these are two days' data.</p>",
      "rawMarkdown": "Good point, these are two days' data.",
      "votes": null
    },
    {
      "id": "300492",
      "postDate": "03/21/2018 16:31:10",
      "content": "<p>Have you heard about <a href=\"https://medium.com/deep-learning-turkey/google-colab-free-gpu-tutorial-e113627b9f5d\">this</a>`?\n=)</p>",
      "rawMarkdown": "Have you heard about [this][1]`?\n=)\n\n\n  [1]: https://medium.com/deep-learning-turkey/google-colab-free-gpu-tutorial-e113627b9f5d",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 293527,
      "author_name": "blasteraj",
      "author_url": "",
      "post_date": "03/10/2018 03:51:50",
      "content": "<p>Hi, Is there any other option $3.6/hr is costly for me and i want to use whole dataset.</p>",
      "votes": null,
      "replies": [
        {
          "id": 294211,
          "author_name": "wythhh",
          "author_url": "",
          "post_date": "03/11/2018 15:17:09",
          "content": "<p>You could try google cloud engine, they always have $300 coupons for starters.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 296972,
          "author_name": "zyczyc",
          "author_url": "",
          "post_date": "03/16/2018 03:24:03",
          "content": "<p>I think ec2 spot instance is a better choice, much cheaper ... and I built my environment and download data on an external EBS to keep it long time availabel.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 293530,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "03/10/2018 04:08:47",
      "content": "<p>I'm using a 16GB RAM Mac doing some feature engineering and training a lightgbm model using 50% of the trainset.</p>",
      "votes": null,
      "replies": [
        {
          "id": 293533,
          "author_name": "sohaibomar",
          "author_url": "",
          "post_date": "03/10/2018 04:13:27",
          "content": "<p>It means you are on top using lightgbm and only 50% trainset. AMAZING<br></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 293535,
          "author_name": "wythhh",
          "author_url": "",
          "post_date": "03/10/2018 04:20:41",
          "content": "<p>Well, that's interesting and could I ask how is the memory usage when you run lightgbm? Did it use almost 100% of your RAM? Actually I forget the memory usage when I train lightgbm using whole dataset. But when I did 5 fold cv lightgbm on whole data set, the memory usage is around 64GB.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 294046,
          "author_name": "xiumugengmu",
          "author_url": "",
          "post_date": "03/11/2018 06:39:50",
          "content": "<p>Amazing! However, for further FE or more complicated algorithms, 16GB RAM could be enough?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 294297,
          "author_name": "titericz",
          "author_url": "",
          "post_date": "03/11/2018 17:46:34",
          "content": "<p>I trained using only 1 oof. Remember this is a time series challenge you can't do random kFold. I'm using all my RAM available and a huge SWAP file (most of my dataset goes to swap file). Since Mac SSDs are really  fast I can train a model in about 1 hour</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 296776,
          "author_name": "sohaibomar",
          "author_url": "",
          "post_date": "03/15/2018 18:40:49",
          "content": "<pre><code>using 50% of the trainset.\n</code></pre>\n\n<p>May I get a hint on how did you select that 50% of train set? Thanks</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 298020,
          "author_name": "aakashnain",
          "author_url": "",
          "post_date": "03/18/2018 14:54:26",
          "content": "<p>Amazing!!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 298287,
          "author_name": "enricospada",
          "author_url": "",
          "post_date": "03/19/2018 09:11:54",
          "content": "<p>Are you using R or Python? And most important: who have any hint about a tutorial/guide to use SWAP file in R using windows 10? Very interesting discovery!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 298848,
          "author_name": "hfsong",
          "author_url": "",
          "post_date": "03/20/2018 04:44:59",
          "content": "<p>I am using a 32GB RAM and 64G virtual memory from my SSD (WIN10, python). Training 50%-60% of the train dataset may take around 1 hour.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 300415,
      "author_name": "taojiang",
      "author_url": "",
      "post_date": "03/21/2018 15:46:43",
      "content": "<p>I'm thinking if we can change the data type from int64 to int32 (or int16), I mean except the time column. Then we could get a somehow smaller data set.</p>",
      "votes": null,
      "replies": [
        {
          "id": 300421,
          "author_name": "hfsong",
          "author_url": "",
          "post_date": "03/21/2018 15:51:09",
          "content": "<p>actually you can drop the time column, only left day, hour in int</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 300432,
          "author_name": "taojiang",
          "author_url": "",
          "post_date": "03/21/2018 15:54:50",
          "content": "<p>Good point, these are two days' data.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 300492,
      "author_name": "drhabib",
      "author_url": "",
      "post_date": "03/21/2018 16:31:10",
      "content": "<p>Have you heard about <a href=\"https://medium.com/deep-learning-turkey/google-colab-free-gpu-tutorial-e113627b9f5d\">this</a>`?\n=)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "293320": "From the advice given under this topic, our group selected AWS EC2 instance to have a trial play with the whole data. We choose r4.8xlarge windows instance which as 32 CPU and 244GB memory with $3.6/hour.\nWe extract hour and minute from click_time and remove click_time and attributed_time and we noticed from spyder IDE that the memory usage is around 26%(almost 64GB) when we run 5-fold cv using lightgbm. I think this would be a kind of benchmark of memory usage for people who wants to run algorithms on whole data set. Using whole data set may not be a good idea but if you want, then 32GB or even 64GB is not enough for some complicated algorithms.",
    "293527": "Hi, Is there any other option $3.6/hr is costly for me and i want to use whole dataset.",
    "293530": "I'm using a 16GB RAM Mac doing some feature engineering and training a lightgbm model using 50% of the trainset.",
    "293533": "It means you are on top using lightgbm and only 50% trainset. AMAZING<br>",
    "293535": "Well, that's interesting and could I ask how is the memory usage when you run lightgbm? Did it use almost 100% of your RAM? Actually I forget the memory usage when I train lightgbm using whole dataset. But when I did 5 fold cv lightgbm on whole data set, the memory usage is around 64GB.",
    "294046": "Amazing! However, for further FE or more complicated algorithms, 16GB RAM could be enough?",
    "294211": "You could try google cloud engine, they always have $300 coupons for starters.",
    "294297": "I trained using only 1 oof. Remember this is a time series challenge you can't do random kFold. I'm using all my RAM available and a huge SWAP file (most of my dataset goes to swap file). Since Mac SSDs are really  fast I can train a model in about 1 hour",
    "296776": "using 50% of the trainset.\n\nMay I get a hint on how did you select that 50% of train set? Thanks",
    "296972": "I think ec2 spot instance is a better choice, much cheaper ... and I built my environment and download data on an external EBS to keep it long time availabel.",
    "298020": "Amazing!!",
    "298287": "Are you using R or Python? And most important: who have any hint about a tutorial/guide to use SWAP file in R using windows 10? Very interesting discovery!",
    "298848": "I am using a 32GB RAM and 64G virtual memory from my SSD (WIN10, python). Training 50%-60% of the train dataset may take around 1 hour.",
    "300415": "I'm thinking if we can change the data type from int64 to int32 (or int16), I mean except the time column. Then we could get a somehow smaller data set.",
    "300421": "actually you can drop the time column, only left day, hour in int",
    "300432": "Good point, these are two days' data.",
    "300492": "Have you heard about [this][1]`?\n=)\n\n\n  [1]: https://medium.com/deep-learning-turkey/google-colab-free-gpu-tutorial-e113627b9f5d"
  },
  "source": "meta"
}