{
  "id": 54072,
  "title": "How to groupby more than one column with minimum RAM",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/54072",
  "author_name": "",
  "post_date": "2018-04-09T08:37:09.530770400Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I am trying to do a groupby on more than one column, but It seems that it consumes too much RAM. Is there a smart way to do it with minimum RAM. Here is the piece of code I use: </p>\n\n<pre><code>ip_count1 = merge.groupby(['ip', 'dateclick']).size().reset_index().rename(columns={0:'nbclick'})\n</code></pre>",
  "messages": [
    {
      "id": "311043",
      "postDate": "04/09/2018 08:37:09",
      "content": "<p>I am trying to do a groupby on more than one column, but It seems that it consumes too much RAM. Is there a smart way to do it with minimum RAM. Here is the piece of code I use: </p>\n\n<pre><code>ip_count1 = merge.groupby(['ip', 'dateclick']).size().reset_index().rename(columns={0:'nbclick'})\n</code></pre>",
      "rawMarkdown": "I am trying to do a groupby on more than one column, but It seems that it consumes too much RAM. Is there a smart way to do it with minimum RAM. Here is the piece of code I use: \n\n    ip_count1 = merge.groupby(['ip', 'dateclick']).size().reset_index().rename(columns={0:'nbclick'})",
      "votes": null
    },
    {
      "id": "311088",
      "postDate": "04/09/2018 10:57:50",
      "content": "<p>How about trying a simple groupby and seeing how much time that takes:\nip_count1= merge.groupby(['ip','dateclick']).reset_index()</p>",
      "rawMarkdown": "How about trying a simple groupby and seeing how much time that takes:\nip_count1= merge.groupby(['ip','dateclick']).reset_index()",
      "votes": null
    },
    {
      "id": "311112",
      "postDate": "04/09/2018 12:17:24",
      "content": "<p>Cast the output type to the smallest type that can contain it.  Also rescale counts if they are large.  A good example of how to do it is given in this great kernel from @anttip : <a href=\"https://www.kaggle.com/anttip/talkingdata-wordbatch-fm-ftrl-lb-0-9711\">https://www.kaggle.com/anttip/talkingdata-wordbatch-fm-ftrl-lb-0-9711</a></p>",
      "rawMarkdown": "Cast the output type to the smallest type that can contain it.  Also rescale counts if they are large.  A good example of how to do it is given in this great kernel from @anttip : https://www.kaggle.com/anttip/talkingdata-wordbatch-fm-ftrl-lb-0-9711",
      "votes": null
    },
    {
      "id": "311147",
      "postDate": "04/09/2018 13:46:25",
      "content": "<p>You can collapse multiple columns into a single column, as I did <a href=\"https://www.kaggle.com/aharless/talkingdata-time-deltas\">in this kernel</a>.</p>",
      "rawMarkdown": "You can collapse multiple columns into a single column, as I did [in this kernel][1].\n\n [1]: https://www.kaggle.com/aharless/talkingdata-time-deltas",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 311088,
      "author_name": "sahibachopra",
      "author_url": "",
      "post_date": "04/09/2018 10:57:50",
      "content": "<p>How about trying a simple groupby and seeing how much time that takes:\nip_count1= merge.groupby(['ip','dateclick']).reset_index()</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 311112,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "04/09/2018 12:17:24",
      "content": "<p>Cast the output type to the smallest type that can contain it.  Also rescale counts if they are large.  A good example of how to do it is given in this great kernel from @anttip : <a href=\"https://www.kaggle.com/anttip/talkingdata-wordbatch-fm-ftrl-lb-0-9711\">https://www.kaggle.com/anttip/talkingdata-wordbatch-fm-ftrl-lb-0-9711</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 311147,
      "author_name": "aharless",
      "author_url": "",
      "post_date": "04/09/2018 13:46:25",
      "content": "<p>You can collapse multiple columns into a single column, as I did <a href=\"https://www.kaggle.com/aharless/talkingdata-time-deltas\">in this kernel</a>.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "311043": "I am trying to do a groupby on more than one column, but It seems that it consumes too much RAM. Is there a smart way to do it with minimum RAM. Here is the piece of code I use: \n\n    ip_count1 = merge.groupby(['ip', 'dateclick']).size().reset_index().rename(columns={0:'nbclick'})",
    "311088": "How about trying a simple groupby and seeing how much time that takes:\nip_count1= merge.groupby(['ip','dateclick']).reset_index()",
    "311112": "Cast the output type to the smallest type that can contain it.  Also rescale counts if they are large.  A good example of how to do it is given in this great kernel from @anttip : https://www.kaggle.com/anttip/talkingdata-wordbatch-fm-ftrl-lb-0-9711",
    "311147": "You can collapse multiple columns into a single column, as I did [in this kernel][1].\n\n [1]: https://www.kaggle.com/aharless/talkingdata-time-deltas"
  },
  "source": "meta"
}