{
  "id": 55593,
  "title": "Using SQL For Complex GROUP BY Features",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/55593",
  "author_name": "",
  "post_date": "2018-04-29T06:45:11.166528200Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Lately I was trying to do some complex GROUP BY operations and even with 32GB RAM and around 70GB Swap space I had seen memory error. One work around that I found for these complex operations is to use SQL. With SQL the GROUP BY operations are taking way less time.. like few minutes compared to few hours using pandas. Once the GROUP BY is done we can create a CSV file and use it to merge in to the original dataframe.\nWe could also do the merging using SQL and use the final csv file for building the model.</p>\n\n<p>I find this method very useful when I had to test on 4-5 new features....</p>",
  "messages": [
    {
      "id": "320590",
      "postDate": "04/29/2018 06:45:11",
      "content": "<p>Lately I was trying to do some complex GROUP BY operations and even with 32GB RAM and around 70GB Swap space I had seen memory error. One work around that I found for these complex operations is to use SQL. With SQL the GROUP BY operations are taking way less time.. like few minutes compared to few hours using pandas. Once the GROUP BY is done we can create a CSV file and use it to merge in to the original dataframe.\nWe could also do the merging using SQL and use the final csv file for building the model.</p>\n\n<p>I find this method very useful when I had to test on 4-5 new features....</p>",
      "rawMarkdown": "Lately I was trying to do some complex GROUP BY operations and even with 32GB RAM and around 70GB Swap space I had seen memory error. One work around that I found for these complex operations is to use SQL. With SQL the GROUP BY operations are taking way less time.. like few minutes compared to few hours using pandas. Once the GROUP BY is done we can create a CSV file and use it to merge in to the original dataframe.\nWe could also do the merging using SQL and use the final csv file for building the model.\n\nI find this method very useful when I had to test on 4-5 new features....",
      "votes": null
    },
    {
      "id": "320728",
      "postDate": "04/29/2018 16:21:08",
      "content": "<p>Hi, </p>\n\n<p>Are you calling SQL queries in Python? If so could you share the library you are using. </p>\n\n<p>I have been using Spark SQL and I gained speed from the data partitioning. </p>",
      "rawMarkdown": "Hi, \n\nAre you calling SQL queries in Python? If so could you share the library you are using. \n\nI have been using Spark SQL and I gained speed from the data partitioning.",
      "votes": null
    },
    {
      "id": "320730",
      "postDate": "04/29/2018 16:25:43",
      "content": "<p>You can use this kernel of mine - <a href=\"https://www.kaggle.com/samratp/talkingdata-basic-eda-using-sqlite-full-data\">https://www.kaggle.com/samratp/talkingdata-basic-eda-using-sqlite-full-data</a></p>",
      "rawMarkdown": "You can use this kernel of mine - https://www.kaggle.com/samratp/talkingdata-basic-eda-using-sqlite-full-data",
      "votes": null
    },
    {
      "id": "320732",
      "postDate": "04/29/2018 16:31:47",
      "content": "<p>Thanks, exactly what I was looking for. </p>",
      "rawMarkdown": "Thanks, exactly what I was looking for.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 320728,
      "author_name": "rdekou",
      "author_url": "",
      "post_date": "04/29/2018 16:21:08",
      "content": "<p>Hi, </p>\n\n<p>Are you calling SQL queries in Python? If so could you share the library you are using. </p>\n\n<p>I have been using Spark SQL and I gained speed from the data partitioning. </p>",
      "votes": null,
      "replies": [
        {
          "id": 320730,
          "author_name": "samratp",
          "author_url": "",
          "post_date": "04/29/2018 16:25:43",
          "content": "<p>You can use this kernel of mine - <a href=\"https://www.kaggle.com/samratp/talkingdata-basic-eda-using-sqlite-full-data\">https://www.kaggle.com/samratp/talkingdata-basic-eda-using-sqlite-full-data</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 320732,
          "author_name": "rdekou",
          "author_url": "",
          "post_date": "04/29/2018 16:31:47",
          "content": "<p>Thanks, exactly what I was looking for. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "320590": "Lately I was trying to do some complex GROUP BY operations and even with 32GB RAM and around 70GB Swap space I had seen memory error. One work around that I found for these complex operations is to use SQL. With SQL the GROUP BY operations are taking way less time.. like few minutes compared to few hours using pandas. Once the GROUP BY is done we can create a CSV file and use it to merge in to the original dataframe.\nWe could also do the merging using SQL and use the final csv file for building the model.\n\nI find this method very useful when I had to test on 4-5 new features....",
    "320728": "Hi, \n\nAre you calling SQL queries in Python? If so could you share the library you are using. \n\nI have been using Spark SQL and I gained speed from the data partitioning.",
    "320730": "You can use this kernel of mine - https://www.kaggle.com/samratp/talkingdata-basic-eda-using-sqlite-full-data",
    "320732": "Thanks, exactly what I was looking for."
  },
  "source": "meta"
}