{
  "id": 44280,
  "title": "Feature Engineering in User_Logs",
  "url": "/competitions/kkbox-churn-prediction-challenge/discussion/44280",
  "author_name": "",
  "post_date": "2017-11-26T14:53:37.878321300Z",
  "votes": null,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I am trying to create features in user_log but am finding that I am having a few performance issues. Are there any ways to more quickly open the user_logs file? When both versions are consolidated, the file is ~32gb and often crashes the kernel. I am also trying to create features clustering around specific users and find that performance is also really slow. Are there any tips to  improve performance on such a large file? </p>",
  "messages": [
    {
      "id": "248587",
      "postDate": "11/26/2017 14:53:37",
      "content": "<p>I am trying to create features in user_log but am finding that I am having a few performance issues. Are there any ways to more quickly open the user_logs file? When both versions are consolidated, the file is ~32gb and often crashes the kernel. I am also trying to create features clustering around specific users and find that performance is also really slow. Are there any tips to  improve performance on such a large file? </p>",
      "rawMarkdown": "I am trying to create features in user_log but am finding that I am having a few performance issues. Are there any ways to more quickly open the user_logs file? When both versions are consolidated, the file is ~32gb and often crashes the kernel. I am also trying to create features clustering around specific users and find that performance is also really slow. Are there any tips to  improve performance on such a large file?",
      "votes": null
    },
    {
      "id": "248641",
      "postDate": "11/26/2017 17:10:16",
      "content": "<p>For user_logs, you can read it in chunks and then group those chunks and then append next chunk onto the previous one. You should also group the final DF once the file has been read. Be sure to use use_col in pd.read_csv to read only specific columns. I found only three columns to be of any use. </p>",
      "rawMarkdown": "For user_logs, you can read it in chunks and then group those chunks and then append next chunk onto the previous one. You should also group the final DF once the file has been read. Be sure to use use_col in pd.read_csv to read only specific columns. I found only three columns to be of any use.",
      "votes": null
    },
    {
      "id": "248643",
      "postDate": "11/26/2017 17:11:48",
      "content": "<p>Also, let me know if you would want to team up. </p>",
      "rawMarkdown": "Also, let me know if you would want to team up.",
      "votes": null
    },
    {
      "id": "248750",
      "postDate": "11/26/2017 22:13:44",
      "content": "<p>I'm currently loading it in chunks, but will try loading a few columns to see if that speeds things up. I'l love to team up and compare notes. </p>",
      "rawMarkdown": "I'm currently loading it in chunks, but will try loading a few columns to see if that speeds things up. I'l love to team up and compare notes.",
      "votes": null
    },
    {
      "id": "248768",
      "postDate": "11/26/2017 23:07:34",
      "content": "<p>I have used chunk size = 1 million rows and it is working fine. </p>\n\n<p>Add me...</p>",
      "rawMarkdown": "I have used chunk size = 1 million rows and it is working fine. \n\nAdd me...",
      "votes": null
    },
    {
      "id": "248783",
      "postDate": "11/27/2017 00:12:01",
      "content": "<p>Sent the invite. Message me to get your contact details so we can more easily communicate. </p>",
      "rawMarkdown": "Sent the invite. Message me to get your contact details so we can more easily communicate.",
      "votes": null
    },
    {
      "id": "248807",
      "postDate": "11/27/2017 01:51:37",
      "content": "<p>I can't use the messaging service as well. </p>",
      "rawMarkdown": "I can't use the messaging service as well.",
      "votes": null
    },
    {
      "id": "248987",
      "postDate": "11/27/2017 13:23:50",
      "content": "<p>send me a friend request on Facebook. James Chartouni </p>",
      "rawMarkdown": "send me a friend request on Facebook. James Chartouni",
      "votes": null
    },
    {
      "id": "255485",
      "postDate": "12/09/2017 07:01:04",
      "content": "<p>I just posted a log processing script in the kernels section that I used to summarize all of the logs.  Ran it overnight on a MacBook pro laptop and it worked fine. Maybe something in there will help.</p>",
      "rawMarkdown": "I just posted a log processing script in the kernels section that I used to summarize all of the logs.  Ran it overnight on a MacBook pro laptop and it worked fine. Maybe something in there will help.",
      "votes": null
    },
    {
      "id": "256122",
      "postDate": "12/11/2017 07:51:26",
      "content": "<p>One suggestion is to separate the log file to several sub files based on the date. I myself separate the file to 27 sub files. (by month)</p>",
      "rawMarkdown": "One suggestion is to separate the log file to several sub files based on the date. I myself separate the file to 27 sub files. (by month)",
      "votes": null
    },
    {
      "id": "256444",
      "postDate": "12/11/2017 23:50:17",
      "content": "<p>Another suggestion is to map the msno's (which have length 44) to unique integers, which requires significantly less space.  You just have to convert the id's back to msno's for submissions. </p>",
      "rawMarkdown": "Another suggestion is to map the msno's (which have length 44) to unique integers, which requires significantly less space.  You just have to convert the id's back to msno's for submissions.",
      "votes": null
    },
    {
      "id": "257078",
      "postDate": "12/13/2017 11:00:26",
      "content": "<p>@ SecondTimeAround</p>\n\n<p>this is what I did to avoid the massive disk trashing of the msnos..... for those of us with limited resources</p>\n\n<ol>\n<li>I moved into gzipped files</li>\n<li>get all (unique) msnos of all files</li>\n</ol>\n\n<p>zcat train.csv.gz sample_submission_zero.csv.gz transactions.csv.gz user_logs.csv.gz members.csv.gz | cut -d \",\" -f 1 | sort -u -T . --parallel=8 | gzip &gt; msno.gz &amp;\nLoad into pandas and copy the index to a second column. Set msno as the new index, save on hdf/format table</p>\n\n<ol>\n<li><p>Get all files sorted</p>\n\n<p>by msno and other interesting fields. This will help later, as all transactions/usage per user are collated together  in the dataframes and in the files,  making partitioning   very easy and fast. No need for groupbys. It's all done already.</p></li>\n</ol>\n\n<p>zcat transactions.csv.gz | sort --key=1,1 --key=7,7 --key=8,8 -t \",\" -T . --parallel=8 | gzip -c &gt; transactions_s178.gz </p>\n\n<ol>\n<li>Make a join of the msno table, on msno, with these new sorted files</li>\n</ol>\n\n<p>with these new sorted files and save on hdf/format table. Don't forget to list columns you want to use as file index later ( i.e.  msnoIndex, transaction_date, membership.....). This way you can extract  interesting date ranges.</p>\n\n<p>Once  you have the all this, you can fly over the data very easy and fast.  The whole thing on ssd and corei7 took about an hour.</p>\n\n<p>When you are ready for your submission, simple make the inverse join to get again the msno strings.......</p>\n\n<p>EDIT: Make sure you export LC_ALL=C in your  shell to make sure sorting is consistent with pandas sorting. Be mindful as well if you have headers. Strip them out with grep.</p>",
      "rawMarkdown": "SecondTimeAround\n\nthis is what I did to avoid the massive disk trashing of the msnos..... for those of us with limited resources\n\n 1. I moved into gzipped files\n 2. get all (unique) msnos of all files\n\n\nzcat train.csv.gz sample_submission_zero.csv.gz transactions.csv.gz user_logs.csv.gz members.csv.gz | cut -d \",\" -f 1 | sort -u -T . --parallel=8 | gzip &gt; msno.gz &amp;\nLoad into pandas and copy the index to a second column. Set msno as the new index, save on hdf/format table\n\n 3. Get all files sorted\n\n by msno and other interesting fields. This will help later, as all transactions/usage per user are collated together  in the dataframes and in the files,  making partitioning   very easy and fast. No need for groupbys. It's all done already.\n\nzcat transactions.csv.gz | sort --key=1,1 --key=7,7 --key=8,8 -t \",\" -T . --parallel=8 | gzip -c &gt; transactions_s178.gz \n\n 5. Make a join of the msno table, on msno, with these new sorted files\n\nwith these new sorted files and save on hdf/format table. Don't forget to list columns you want to use as file index later ( i.e.  msnoIndex, transaction_date, membership.....). This way you can extract  interesting date ranges.\n\nOnce  you have the all this, you can fly over the data very easy and fast.  The whole thing on ssd and corei7 took about an hour.\n\nWhen you are ready for your submission, simple make the inverse join to get again the msno strings.......\n\nEDIT: Make sure you export LC_ALL=C in your  shell to make sure sorting is consistent with pandas sorting. Be mindful as well if you have headers. Strip them out with grep.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 248641,
      "author_name": "aroragaurav",
      "author_url": "",
      "post_date": "11/26/2017 17:10:16",
      "content": "<p>For user_logs, you can read it in chunks and then group those chunks and then append next chunk onto the previous one. You should also group the final DF once the file has been read. Be sure to use use_col in pd.read_csv to read only specific columns. I found only three columns to be of any use. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 248643,
      "author_name": "aroragaurav",
      "author_url": "",
      "post_date": "11/26/2017 17:11:48",
      "content": "<p>Also, let me know if you would want to team up. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 248750,
      "author_name": "jameschartouni",
      "author_url": "",
      "post_date": "11/26/2017 22:13:44",
      "content": "<p>I'm currently loading it in chunks, but will try loading a few columns to see if that speeds things up. I'l love to team up and compare notes. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 248768,
      "author_name": "aroragaurav",
      "author_url": "",
      "post_date": "11/26/2017 23:07:34",
      "content": "<p>I have used chunk size = 1 million rows and it is working fine. </p>\n\n<p>Add me...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 248783,
      "author_name": "jameschartouni",
      "author_url": "",
      "post_date": "11/27/2017 00:12:01",
      "content": "<p>Sent the invite. Message me to get your contact details so we can more easily communicate. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 248807,
      "author_name": "aroragaurav",
      "author_url": "",
      "post_date": "11/27/2017 01:51:37",
      "content": "<p>I can't use the messaging service as well. </p>",
      "votes": null,
      "replies": [
        {
          "id": 248987,
          "author_name": "jameschartouni",
          "author_url": "",
          "post_date": "11/27/2017 13:23:50",
          "content": "<p>send me a friend request on Facebook. James Chartouni </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 255485,
      "author_name": "ecodan",
      "author_url": "",
      "post_date": "12/09/2017 07:01:04",
      "content": "<p>I just posted a log processing script in the kernels section that I used to summarize all of the logs.  Ran it overnight on a MacBook pro laptop and it worked fine. Maybe something in there will help.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 256122,
      "author_name": "infinitewing",
      "author_url": "",
      "post_date": "12/11/2017 07:51:26",
      "content": "<p>One suggestion is to separate the log file to several sub files based on the date. I myself separate the file to 27 sub files. (by month)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 256444,
      "author_name": "",
      "author_url": "",
      "post_date": "12/11/2017 23:50:17",
      "content": "<p>Another suggestion is to map the msno's (which have length 44) to unique integers, which requires significantly less space.  You just have to convert the id's back to msno's for submissions. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 257078,
      "author_name": "julcol",
      "author_url": "",
      "post_date": "12/13/2017 11:00:26",
      "content": "<p>@ SecondTimeAround</p>\n\n<p>this is what I did to avoid the massive disk trashing of the msnos..... for those of us with limited resources</p>\n\n<ol>\n<li>I moved into gzipped files</li>\n<li>get all (unique) msnos of all files</li>\n</ol>\n\n<p>zcat train.csv.gz sample_submission_zero.csv.gz transactions.csv.gz user_logs.csv.gz members.csv.gz | cut -d \",\" -f 1 | sort -u -T . --parallel=8 | gzip &gt; msno.gz &amp;\nLoad into pandas and copy the index to a second column. Set msno as the new index, save on hdf/format table</p>\n\n<ol>\n<li><p>Get all files sorted</p>\n\n<p>by msno and other interesting fields. This will help later, as all transactions/usage per user are collated together  in the dataframes and in the files,  making partitioning   very easy and fast. No need for groupbys. It's all done already.</p></li>\n</ol>\n\n<p>zcat transactions.csv.gz | sort --key=1,1 --key=7,7 --key=8,8 -t \",\" -T . --parallel=8 | gzip -c &gt; transactions_s178.gz </p>\n\n<ol>\n<li>Make a join of the msno table, on msno, with these new sorted files</li>\n</ol>\n\n<p>with these new sorted files and save on hdf/format table. Don't forget to list columns you want to use as file index later ( i.e.  msnoIndex, transaction_date, membership.....). This way you can extract  interesting date ranges.</p>\n\n<p>Once  you have the all this, you can fly over the data very easy and fast.  The whole thing on ssd and corei7 took about an hour.</p>\n\n<p>When you are ready for your submission, simple make the inverse join to get again the msno strings.......</p>\n\n<p>EDIT: Make sure you export LC_ALL=C in your  shell to make sure sorting is consistent with pandas sorting. Be mindful as well if you have headers. Strip them out with grep.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "248587": "I am trying to create features in user_log but am finding that I am having a few performance issues. Are there any ways to more quickly open the user_logs file? When both versions are consolidated, the file is ~32gb and often crashes the kernel. I am also trying to create features clustering around specific users and find that performance is also really slow. Are there any tips to  improve performance on such a large file?",
    "248641": "For user_logs, you can read it in chunks and then group those chunks and then append next chunk onto the previous one. You should also group the final DF once the file has been read. Be sure to use use_col in pd.read_csv to read only specific columns. I found only three columns to be of any use.",
    "248643": "Also, let me know if you would want to team up.",
    "248750": "I'm currently loading it in chunks, but will try loading a few columns to see if that speeds things up. I'l love to team up and compare notes.",
    "248768": "I have used chunk size = 1 million rows and it is working fine. \n\nAdd me...",
    "248783": "Sent the invite. Message me to get your contact details so we can more easily communicate.",
    "248807": "I can't use the messaging service as well.",
    "248987": "send me a friend request on Facebook. James Chartouni",
    "255485": "I just posted a log processing script in the kernels section that I used to summarize all of the logs.  Ran it overnight on a MacBook pro laptop and it worked fine. Maybe something in there will help.",
    "256122": "One suggestion is to separate the log file to several sub files based on the date. I myself separate the file to 27 sub files. (by month)",
    "256444": "Another suggestion is to map the msno's (which have length 44) to unique integers, which requires significantly less space.  You just have to convert the id's back to msno's for submissions.",
    "257078": "SecondTimeAround\n\nthis is what I did to avoid the massive disk trashing of the msnos..... for those of us with limited resources\n\n 1. I moved into gzipped files\n 2. get all (unique) msnos of all files\n\n\nzcat train.csv.gz sample_submission_zero.csv.gz transactions.csv.gz user_logs.csv.gz members.csv.gz | cut -d \",\" -f 1 | sort -u -T . --parallel=8 | gzip &gt; msno.gz &amp;\nLoad into pandas and copy the index to a second column. Set msno as the new index, save on hdf/format table\n\n 3. Get all files sorted\n\n by msno and other interesting fields. This will help later, as all transactions/usage per user are collated together  in the dataframes and in the files,  making partitioning   very easy and fast. No need for groupbys. It's all done already.\n\nzcat transactions.csv.gz | sort --key=1,1 --key=7,7 --key=8,8 -t \",\" -T . --parallel=8 | gzip -c &gt; transactions_s178.gz \n\n 5. Make a join of the msno table, on msno, with these new sorted files\n\nwith these new sorted files and save on hdf/format table. Don't forget to list columns you want to use as file index later ( i.e.  msnoIndex, transaction_date, membership.....). This way you can extract  interesting date ranges.\n\nOnce  you have the all this, you can fly over the data very easy and fast.  The whole thing on ssd and corei7 took about an hour.\n\nWhen you are ready for your submission, simple make the inverse join to get again the msno strings.......\n\nEDIT: Make sure you export LC_ALL=C in your  shell to make sure sorting is consistent with pandas sorting. Be mindful as well if you have headers. Strip them out with grep."
  },
  "source": "meta"
}