{
  "id": 53450,
  "title": "C++ Time Delta Code",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/53450",
  "author_name": "",
  "post_date": "2018-03-30T19:32:45.349350500Z",
  "votes": 23,
  "comment_count": 19,
  "views": 0,
  "content": "<p>Here is the code - call it a reward for CPMP saving me loads of time</p>\n\n<p>Don't forget to use all the <strong><em>correct</em></strong> data and time order it ;)</p>",
  "messages": [
    {
      "id": "306646",
      "postDate": "03/30/2018 19:32:45",
      "content": "<p>Here is the code - call it a reward for CPMP saving me loads of time</p>\n\n<p>Don't forget to use all the <strong><em>correct</em></strong> data and time order it ;)</p>",
      "rawMarkdown": "Here is the code - call it a reward for CPMP saving me loads of time\n\nDon't forget to use all the ***correct*** data and time order it ;)",
      "votes": null
    },
    {
      "id": "308620",
      "postDate": "04/03/2018 19:54:28",
      "content": "<pre><code>Don't forget to use all the correct data and time order it ;)\n</code></pre>\n\n<p>means to sort train and test file by click_time and the find delta, am I right?</p>",
      "rawMarkdown": "Don't forget to use all the correct data and time order it ;)\n\n\n\nmeans to sort train and test file by click_time and the find delta, am I right?",
      "votes": null
    },
    {
      "id": "308627",
      "postDate": "04/03/2018 20:13:42",
      "content": "<p>Indeed you are measure times between <em>successive</em> clicks</p>",
      "rawMarkdown": "Indeed you are measure times between *successive* clicks",
      "votes": null
    },
    {
      "id": "308636",
      "postDate": "04/03/2018 20:36:02",
      "content": "<blockquote>\n  <p>call it a reward for CPMP saving me loads of time</p>\n</blockquote>\n\n<p>@Scirpus, thank you. Wish I had your ability to solve all my programming problems in C++.</p>\n\n<p>@CPMP, better post something here so we can reward you as well.</p>",
      "rawMarkdown": "&gt; call it a reward for CPMP saving me loads of time\n\n@Scirpus, thank you. Wish I had your ability to solve all my programming problems in C++.\n\n@CPMP, better post something here so we can reward you as well.",
      "votes": null
    },
    {
      "id": "308653",
      "postDate": "04/03/2018 21:20:30",
      "content": "<p>I posted there: <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53361#306625\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53361#306625</a></p>",
      "rawMarkdown": "I posted there: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53361#306625",
      "votes": null
    },
    {
      "id": "308859",
      "postDate": "04/04/2018 08:02:22",
      "content": "<p>How to run this on linux(For c/c++ noobs like me ):<br>\ndownload CMAKE file by @Scirpus from here <a href=\"https://storage.googleapis.com/kaggle-forum-message-attachments/142921/5288/CMakeLists.txt\">here</a> and copy it in to the directory with main.cpp and CMakeCache file. And do the:</p>\n\n<pre><code>$ sudo apt-get install cmake    \n$ cmake\n$ make\n$ ./GP\n</code></pre>\n\n<p><strong>Note:</strong> You might need to change directory path in main.cpp and CMakeCache files if you have different directory structure.<br>\nThanks</p>",
      "rawMarkdown": "How to run this on linux(For c/c++ noobs like me ):<br>\ndownload CMAKE file by @Scirpus from here [here][1] and copy it in to the directory with main.cpp and CMakeCache file. And do the:\n\n    $ sudo apt-get install cmake    \n    $ cmake\n    $ make\n    $ ./GP\n\n**Note:** You might need to change directory path in main.cpp and CMakeCache files if you have different directory structure.<br>\nThanks\n\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/142921/5288/CMakeLists.txt",
      "votes": null
    },
    {
      "id": "308953",
      "postDate": "04/04/2018 12:27:27",
      "content": "<p>Hi Scirpus, I'm totally noob in C. Can you tell me against what columns are you grouping click-time to calculate the delta? \nFor eg; for every ip and app, you sort clicktime and calculate the delta? Or its something else?</p>\n\n<p>Self edit : Saw the older Time Delta thread. Got it.</p>",
      "rawMarkdown": "Hi Scirpus, I'm totally noob in C. Can you tell me against what columns are you grouping click-time to calculate the delta? \nFor eg; for every ip and app, you sort clicktime and calculate the delta? Or its something else?\n\nSelf edit : Saw the older Time Delta thread. Got it.",
      "votes": null
    },
    {
      "id": "308958",
      "postDate": "04/04/2018 12:37:02",
      "content": "<p>Using R data.table goes pretty fast as well (on the whole train+test approx 10 min).</p>\n\n<p>If I get it right (let me know otherwise), the time delta @Scirpus is computing is equivalent to the following:</p>\n\n<pre><code>  dt[,tt:=as.numeric(click_time)]\n  setorder(dt, tt)\n\n  dt[,delta_ip     :=tt-shift(tt), by=.(ip)]\n  dt[,delta_ip_dev :=tt-shift(tt), by=.(ip,device)]\n  dt[,delta_ip_app :=tt-shift(tt), by=.(ip,app)]\n  dt[,delta_ip_chan:=tt-shift(tt), by=.(ip,channel)]\n\n  dt[,tt:=NULL]\n\n  vars &lt;- dt %&gt;% names %gv% \"^delta_\" \n  for(jj in vars){\n    set(dt,i= which(is.na(dt[[jj]])), j=jj, value=-9999)\n  }\n</code></pre>\n\n<p>Where %gv% is a custom function to grep values</p>",
      "rawMarkdown": "Using R data.table goes pretty fast as well (on the whole train+test approx 10 min).\n \nIf I get it right (let me know otherwise), the time delta @Scirpus is computing is equivalent to the following:\n\n\n      dt[,tt:=as.numeric(click_time)]\n      setorder(dt, tt)\n      \n      dt[,delta_ip     :=tt-shift(tt), by=.(ip)]\n      dt[,delta_ip_dev :=tt-shift(tt), by=.(ip,device)]\n      dt[,delta_ip_app :=tt-shift(tt), by=.(ip,app)]\n      dt[,delta_ip_chan:=tt-shift(tt), by=.(ip,channel)]\n     \n      dt[,tt:=NULL]\n      \n      vars &lt;- dt %&gt;% names %gv% \"^delta_\" \n      for(jj in vars){\n        set(dt,i= which(is.na(dt[[jj]])), j=jj, value=-9999)\n      }\n\n  Where %gv% is a custom function to grep values",
      "votes": null
    },
    {
      "id": "309038",
      "postDate": "04/04/2018 14:40:05",
      "content": "<p>So it would be </p>\n\n<pre><code>df['time_delta'] = df.sort_values(['click_time']).groupby('ip', 'app', 'device', os', 'channel')['click_time'].diff()\n</code></pre>\n\n<p>In pandas right?</p>",
      "rawMarkdown": "So it would be \n\n    df['time_delta'] = df.sort_values(['click_time']).groupby('ip', 'app', 'device', os', 'channel')['click_time'].diff()\n\nIn pandas right?",
      "votes": null
    },
    {
      "id": "309041",
      "postDate": "04/04/2018 14:50:00",
      "content": "<p>Pandas one is very very very slow...</p>",
      "rawMarkdown": "Pandas one is very very very slow...",
      "votes": null
    },
    {
      "id": "309083",
      "postDate": "04/04/2018 16:04:14",
      "content": "<p>Yeah just wanted to see if I understood</p>",
      "rawMarkdown": "Yeah just wanted to see if I understood",
      "votes": null
    },
    {
      "id": "309153",
      "postDate": "04/04/2018 18:19:13",
      "content": "<p>You only need to sort by click_time, that's faster.</p>",
      "rawMarkdown": "You only need to sort by click_time, that's faster.",
      "votes": null
    },
    {
      "id": "309398",
      "postDate": "04/05/2018 08:24:33",
      "content": "<p>Could someone upload the results to kaggle? Would like to test it but it keeps crashing after 53million rows for me.</p>",
      "rawMarkdown": "Could someone upload the results to kaggle? Would like to test it but it keeps crashing after 53million rows for me.",
      "votes": null
    },
    {
      "id": "309513",
      "postDate": "04/05/2018 13:54:10",
      "content": "<p>Pandas diff requires a ridiculous amount of memory (probably why it's crashing). We could do it in chunks, I guess.</p>",
      "rawMarkdown": "Pandas diff requires a ridiculous amount of memory (probably why it's crashing). We could do it in chunks, I guess.",
      "votes": null
    },
    {
      "id": "309681",
      "postDate": "04/05/2018 19:48:56",
      "content": "<p>OK, I've tried to <a href=\"https://www.kaggle.com/aharless/time-deltas\">do it in Python</a>, rather kludgily, I'm afraid, but it runs in reasonable time, and (after a great deal of difficulty) doesn't breach the memory limit.  I haven't tested the result, so if someone wants to check it against what Scirpus has, that would be useful.  (The output is in pickle format with NaN values for the first occurrence of each combined category. Obviously, assuming it's correct, one could reuse the code for different aggregations and so on.)</p>",
      "rawMarkdown": "OK, I've tried to [do it in Python][1], rather kludgily, I'm afraid, but it runs in reasonable time, and (after a great deal of difficulty) doesn't breach the memory limit.  I haven't tested the result, so if someone wants to check it against what Scirpus has, that would be useful.  (The output is in pickle format with NaN values for the first occurrence of each combined category. Obviously, assuming it's correct, one could reuse the code for different aggregations and so on.)\n\n [1]: https://www.kaggle.com/aharless/time-deltas",
      "votes": null
    },
    {
      "id": "316772",
      "postDate": "04/19/2018 21:18:43",
      "content": "<p>Here's a Makefile for a simple build without CMake. @Scirpus Did you intend to attach CMakeLists.txt instead of CMakeCache.txt? </p>",
      "rawMarkdown": "Here's a Makefile for a simple build without CMake. @Scirpus Did you intend to attach CMakeLists.txt instead of CMakeCache.txt?",
      "votes": null
    },
    {
      "id": "316857",
      "postDate": "04/20/2018 04:52:01",
      "content": "<p>Quote:Don't forget to use all the correct data and time order it ;)\nYou mean to say data is not already sorted?</p>",
      "rawMarkdown": "Quote:Don't forget to use all the correct data and time order it ;)\nYou mean to say data is not already sorted?",
      "votes": null
    },
    {
      "id": "316866",
      "postDate": "04/20/2018 05:10:49",
      "content": "<p>Perhaps it has to do if you plan to do it with merged train and test files. </p>",
      "rawMarkdown": "Perhaps it has to do if you plan to do it with merged train and test files.",
      "votes": null
    },
    {
      "id": "316902",
      "postDate": "04/20/2018 07:01:47",
      "content": "<p>I just added this as a reminder to check - especially when using the test supplement</p>",
      "rawMarkdown": "I just added this as a reminder to check - especially when using the test supplement",
      "votes": null
    },
    {
      "id": "316903",
      "postDate": "04/20/2018 07:02:11",
      "content": "<p>Add CMake File - thanks</p>",
      "rawMarkdown": "Add CMake File - thanks",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 308620,
      "author_name": "sohaibomar",
      "author_url": "",
      "post_date": "04/03/2018 19:54:28",
      "content": "<pre><code>Don't forget to use all the correct data and time order it ;)\n</code></pre>\n\n<p>means to sort train and test file by click_time and the find delta, am I right?</p>",
      "votes": null,
      "replies": [
        {
          "id": 308627,
          "author_name": "scirpus",
          "author_url": "",
          "post_date": "04/03/2018 20:13:42",
          "content": "<p>Indeed you are measure times between <em>successive</em> clicks</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 308953,
          "author_name": "nuhsikander",
          "author_url": "",
          "post_date": "04/04/2018 12:27:27",
          "content": "<p>Hi Scirpus, I'm totally noob in C. Can you tell me against what columns are you grouping click-time to calculate the delta? \nFor eg; for every ip and app, you sort clicktime and calculate the delta? Or its something else?</p>\n\n<p>Self edit : Saw the older Time Delta thread. Got it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 308636,
      "author_name": "tilii7",
      "author_url": "",
      "post_date": "04/03/2018 20:36:02",
      "content": "<blockquote>\n  <p>call it a reward for CPMP saving me loads of time</p>\n</blockquote>\n\n<p>@Scirpus, thank you. Wish I had your ability to solve all my programming problems in C++.</p>\n\n<p>@CPMP, better post something here so we can reward you as well.</p>",
      "votes": null,
      "replies": [
        {
          "id": 308653,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/03/2018 21:20:30",
          "content": "<p>I posted there: <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53361#306625\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53361#306625</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 308859,
      "author_name": "sohaibomar",
      "author_url": "",
      "post_date": "04/04/2018 08:02:22",
      "content": "<p>How to run this on linux(For c/c++ noobs like me ):<br>\ndownload CMAKE file by @Scirpus from here <a href=\"https://storage.googleapis.com/kaggle-forum-message-attachments/142921/5288/CMakeLists.txt\">here</a> and copy it in to the directory with main.cpp and CMakeCache file. And do the:</p>\n\n<pre><code>$ sudo apt-get install cmake    \n$ cmake\n$ make\n$ ./GP\n</code></pre>\n\n<p><strong>Note:</strong> You might need to change directory path in main.cpp and CMakeCache files if you have different directory structure.<br>\nThanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 308958,
      "author_name": "juansillero",
      "author_url": "",
      "post_date": "04/04/2018 12:37:02",
      "content": "<p>Using R data.table goes pretty fast as well (on the whole train+test approx 10 min).</p>\n\n<p>If I get it right (let me know otherwise), the time delta @Scirpus is computing is equivalent to the following:</p>\n\n<pre><code>  dt[,tt:=as.numeric(click_time)]\n  setorder(dt, tt)\n\n  dt[,delta_ip     :=tt-shift(tt), by=.(ip)]\n  dt[,delta_ip_dev :=tt-shift(tt), by=.(ip,device)]\n  dt[,delta_ip_app :=tt-shift(tt), by=.(ip,app)]\n  dt[,delta_ip_chan:=tt-shift(tt), by=.(ip,channel)]\n\n  dt[,tt:=NULL]\n\n  vars &lt;- dt %&gt;% names %gv% \"^delta_\" \n  for(jj in vars){\n    set(dt,i= which(is.na(dt[[jj]])), j=jj, value=-9999)\n  }\n</code></pre>\n\n<p>Where %gv% is a custom function to grep values</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 309038,
      "author_name": "michaelsnell",
      "author_url": "",
      "post_date": "04/04/2018 14:40:05",
      "content": "<p>So it would be </p>\n\n<pre><code>df['time_delta'] = df.sort_values(['click_time']).groupby('ip', 'app', 'device', os', 'channel')['click_time'].diff()\n</code></pre>\n\n<p>In pandas right?</p>",
      "votes": null,
      "replies": [
        {
          "id": 309041,
          "author_name": "cczaixian",
          "author_url": "",
          "post_date": "04/04/2018 14:50:00",
          "content": "<p>Pandas one is very very very slow...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 309083,
          "author_name": "michaelsnell",
          "author_url": "",
          "post_date": "04/04/2018 16:04:14",
          "content": "<p>Yeah just wanted to see if I understood</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 309153,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/04/2018 18:19:13",
          "content": "<p>You only need to sort by click_time, that's faster.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 309513,
          "author_name": "aharless",
          "author_url": "",
          "post_date": "04/05/2018 13:54:10",
          "content": "<p>Pandas diff requires a ridiculous amount of memory (probably why it's crashing). We could do it in chunks, I guess.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 309398,
      "author_name": "michaelsnell",
      "author_url": "",
      "post_date": "04/05/2018 08:24:33",
      "content": "<p>Could someone upload the results to kaggle? Would like to test it but it keeps crashing after 53million rows for me.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 309681,
      "author_name": "aharless",
      "author_url": "",
      "post_date": "04/05/2018 19:48:56",
      "content": "<p>OK, I've tried to <a href=\"https://www.kaggle.com/aharless/time-deltas\">do it in Python</a>, rather kludgily, I'm afraid, but it runs in reasonable time, and (after a great deal of difficulty) doesn't breach the memory limit.  I haven't tested the result, so if someone wants to check it against what Scirpus has, that would be useful.  (The output is in pickle format with NaN values for the first occurrence of each combined category. Obviously, assuming it's correct, one could reuse the code for different aggregations and so on.)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 316772,
      "author_name": "pliptor",
      "author_url": "",
      "post_date": "04/19/2018 21:18:43",
      "content": "<p>Here's a Makefile for a simple build without CMake. @Scirpus Did you intend to attach CMakeLists.txt instead of CMakeCache.txt? </p>",
      "votes": null,
      "replies": [
        {
          "id": 316903,
          "author_name": "scirpus",
          "author_url": "",
          "post_date": "04/20/2018 07:02:11",
          "content": "<p>Add CMake File - thanks</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 316857,
      "author_name": "adityakumarsinha",
      "author_url": "",
      "post_date": "04/20/2018 04:52:01",
      "content": "<p>Quote:Don't forget to use all the correct data and time order it ;)\nYou mean to say data is not already sorted?</p>",
      "votes": null,
      "replies": [
        {
          "id": 316866,
          "author_name": "pliptor",
          "author_url": "",
          "post_date": "04/20/2018 05:10:49",
          "content": "<p>Perhaps it has to do if you plan to do it with merged train and test files. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 316902,
          "author_name": "scirpus",
          "author_url": "",
          "post_date": "04/20/2018 07:01:47",
          "content": "<p>I just added this as a reminder to check - especially when using the test supplement</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "306646": "Here is the code - call it a reward for CPMP saving me loads of time\n\nDon't forget to use all the ***correct*** data and time order it ;)",
    "308620": "Don't forget to use all the correct data and time order it ;)\n\n\n\nmeans to sort train and test file by click_time and the find delta, am I right?",
    "308627": "Indeed you are measure times between *successive* clicks",
    "308636": "&gt; call it a reward for CPMP saving me loads of time\n\n@Scirpus, thank you. Wish I had your ability to solve all my programming problems in C++.\n\n@CPMP, better post something here so we can reward you as well.",
    "308653": "I posted there: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53361#306625",
    "308859": "How to run this on linux(For c/c++ noobs like me ):<br>\ndownload CMAKE file by @Scirpus from here [here][1] and copy it in to the directory with main.cpp and CMakeCache file. And do the:\n\n    $ sudo apt-get install cmake    \n    $ cmake\n    $ make\n    $ ./GP\n\n**Note:** You might need to change directory path in main.cpp and CMakeCache files if you have different directory structure.<br>\nThanks\n\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/142921/5288/CMakeLists.txt",
    "308953": "Hi Scirpus, I'm totally noob in C. Can you tell me against what columns are you grouping click-time to calculate the delta? \nFor eg; for every ip and app, you sort clicktime and calculate the delta? Or its something else?\n\nSelf edit : Saw the older Time Delta thread. Got it.",
    "308958": "Using R data.table goes pretty fast as well (on the whole train+test approx 10 min).\n \nIf I get it right (let me know otherwise), the time delta @Scirpus is computing is equivalent to the following:\n\n\n      dt[,tt:=as.numeric(click_time)]\n      setorder(dt, tt)\n      \n      dt[,delta_ip     :=tt-shift(tt), by=.(ip)]\n      dt[,delta_ip_dev :=tt-shift(tt), by=.(ip,device)]\n      dt[,delta_ip_app :=tt-shift(tt), by=.(ip,app)]\n      dt[,delta_ip_chan:=tt-shift(tt), by=.(ip,channel)]\n     \n      dt[,tt:=NULL]\n      \n      vars &lt;- dt %&gt;% names %gv% \"^delta_\" \n      for(jj in vars){\n        set(dt,i= which(is.na(dt[[jj]])), j=jj, value=-9999)\n      }\n\n  Where %gv% is a custom function to grep values",
    "309038": "So it would be \n\n    df['time_delta'] = df.sort_values(['click_time']).groupby('ip', 'app', 'device', os', 'channel')['click_time'].diff()\n\nIn pandas right?",
    "309041": "Pandas one is very very very slow...",
    "309083": "Yeah just wanted to see if I understood",
    "309153": "You only need to sort by click_time, that's faster.",
    "309398": "Could someone upload the results to kaggle? Would like to test it but it keeps crashing after 53million rows for me.",
    "309513": "Pandas diff requires a ridiculous amount of memory (probably why it's crashing). We could do it in chunks, I guess.",
    "309681": "OK, I've tried to [do it in Python][1], rather kludgily, I'm afraid, but it runs in reasonable time, and (after a great deal of difficulty) doesn't breach the memory limit.  I haven't tested the result, so if someone wants to check it against what Scirpus has, that would be useful.  (The output is in pickle format with NaN values for the first occurrence of each combined category. Obviously, assuming it's correct, one could reuse the code for different aggregations and so on.)\n\n [1]: https://www.kaggle.com/aharless/time-deltas",
    "316772": "Here's a Makefile for a simple build without CMake. @Scirpus Did you intend to attach CMakeLists.txt instead of CMakeCache.txt?",
    "316857": "Quote:Don't forget to use all the correct data and time order it ;)\nYou mean to say data is not already sorted?",
    "316866": "Perhaps it has to do if you plan to do it with merged train and test files.",
    "316902": "I just added this as a reminder to check - especially when using the test supplement",
    "316903": "Add CMake File - thanks"
  },
  "source": "meta"
}