{
  "id": 56065,
  "title": "Question : How to create Nextclick features in R?",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/56065",
  "author_name": "LegenDaD",
  "post_date": "2018-05-05T11:29:12.237000",
  "votes": 1,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Hello, everyone. <br>\nI am recently started to learn R and I find it really interesting. I need your help on creating 'Next Click' feature in R. The code below is what I made for it.</p>\n\n<pre><code>tdiff &lt;- function(tdiff) {\n  vc &lt;- NULL\n  for (x in 1:length(tdiff)) {\n    interval &lt;- interval(tdiff[x], tdiff[x+1])\n    diffsecond &lt;- as.integer(seconds(interval))\n    vc &lt;- append(vc, diffsecond)\n  }\n  vc &lt;- ifelse(is.na(vc), 0, vc)\n  vc &lt;- append(vc, 0)\n  return(vc)\n}\n\nadt[, clicker_Next := tdiff(click_time), by = list(ip, device, os)]\n</code></pre>\n\n<p>When I try to run this code through Train_Sample, it is working but takes too long about more than 15 mins. But when it comes to Train CSV, this code doesn't work at all. If you could share your experty on this, it would be much appreciated.  </p>\n\n<p>Thank you for your help in advance.</p>",
  "messages": [
    {
      "id": 323580,
      "postDate": "2018-05-05T15:33:27.970Z",
      "content": "<p><code>shift</code> is known to be slow when too many groups are involved in data.table (because the <code>:=</code> operator is not <code>gforce</code>-optimized yet in data.table)</p>\n\n<p>You can do like the following for even faster lag/lead deltas:</p>\n\n<pre><code>adt[, deltaUApp_nextclick1 := c(click_time_num[-1], NA), by = list(ip, app, device, os)]\nadt[, deltaUApp_nextclick1 := deltaUApp_nextclick1 - click_time_num]\nadt[is.na(deltaUApp_nextclick1), deltaUApp_nextclick1 := 0]\n\nadt[, deltaUApp_prevclick1 := c(NA, click_time_num[-.N]), by = list(ip, app, device, os)]\nadt[, deltaUApp_prevclick1 := deltaUApp_prevclick1 - click_time_num]\nadt[is.na(deltaUApp_prevclick1), deltaUApp_prevclick1 := 0]\n</code></pre>\n\n<p>It assumes you have a integer/numeric time called <code>click_time_num</code>.</p>\n\n<p>Takes less than 5 minutes to do it (i7-7700K, 5.0 GHz) for:</p>\n\n<ul>\n<li>all the 5 groups used on the top kernels</li>\n<li>with ALL data (train + test_supplement)</li>\n<li>under 16GB peak (OS not included)</li>\n</ul>\n\n<p>text</p>\n\n<pre><code>data[, click_time_num := as.numeric(click_time)]\n\ndata[, by_ipappdeviceoschannel_nextclick1 := c(click_time_num[-1], NA), by = c(\"ip\", \"app\", \"device\", \"os\", \"channel\")]\ndata[, by_ipappdeviceoschannel_nextclick1 := by_ipappdeviceoschannel_nextclick1 - click_time_num]\ndata[is.na(by_ipappdeviceoschannel_nextclick1), by_ipappdeviceoschannel_nextclick1 := 0]\n\ndata[, by_iposdevice_nextclick1 := c(click_time_num[-1], NA), by = c(\"ip\", \"os\", \"device\")]\ndata[, by_iposdevice_nextclick1 := by_iposdevice_nextclick1 - click_time_num]\ndata[is.na(by_iposdevice_nextclick1), by_iposdevice_nextclick1 := 0]\n\ndata[, by_iposdeviceapp_nextclick1 := c(click_time_num[-1], NA), by = c(\"ip\", \"os\", \"device\", \"app\")]\ndata[, by_iposdeviceapp_nextclick1 := by_iposdeviceapp_nextclick1 - click_time_num]\ndata[is.na(by_iposdeviceapp_nextclick1), by_iposdeviceapp_nextclick1 := 0]\n\ndata[, by_ipchannel_prevclick1 := c(NA, click_time_num[-.N]), by = c(\"ip\", \"channel\")]\ndata[, by_ipchannel_prevclick1 := click_time_num - by_ipchannel_prevclick1]\ndata[is.na(by_ipchannel_prevclick1), by_ipchannel_prevclick1 := 0]\n\ndata[, by_ipos_prevclick1 := c(NA, click_time_num[-.N]), by = c(\"ip\", \"os\")]\ndata[, by_ipos_prevclick1 := click_time_num - by_ipos_prevclick1]\ndata[is.na(by_ipos_prevclick1), by_ipos_prevclick1 := 0]\n\n      user  system elapsed \n    281.92    6.17  288.11 \n</code></pre>\n\n<p>If computed in a parallel fashion (requires 80GB), it takes approx 80 seconds for all 5. In my code, NA conversion is unnecessary, better use 0 instead of NA in that combine (c) function to speed it even more (will require approx 1GB less RAM in that case).</p>",
      "rawMarkdown": "`shift` is known to be slow when too many groups are involved in data.table (because the `:=` operator is not `gforce`-optimized yet in data.table)\n\nYou can do like the following for even faster lag/lead deltas:\n\n    adt[, deltaUApp_nextclick1 := c(click_time_num[-1], NA), by = list(ip, app, device, os)]\n    adt[, deltaUApp_nextclick1 := deltaUApp_nextclick1 - click_time_num]\n    adt[is.na(deltaUApp_nextclick1), deltaUApp_nextclick1 := 0]\n    \n    adt[, deltaUApp_prevclick1 := c(NA, click_time_num[-.N]), by = list(ip, app, device, os)]\n    adt[, deltaUApp_prevclick1 := deltaUApp_prevclick1 - click_time_num]\n    adt[is.na(deltaUApp_prevclick1), deltaUApp_prevclick1 := 0]\n\nIt assumes you have a integer/numeric time called `click_time_num`.\n\nTakes less than 5 minutes to do it (i7-7700K, 5.0 GHz) for:\n\n* all the 5 groups used on the top kernels\n* with ALL data (train + test_supplement)\n* under 16GB peak (OS not included)\n\ntext\n\n    data[, click_time_num := as.numeric(click_time)]\n\n    data[, by_ipappdeviceoschannel_nextclick1 := c(click_time_num[-1], NA), by = c(\"ip\", \"app\", \"device\", \"os\", \"channel\")]\n    data[, by_ipappdeviceoschannel_nextclick1 := by_ipappdeviceoschannel_nextclick1 - click_time_num]\n    data[is.na(by_ipappdeviceoschannel_nextclick1), by_ipappdeviceoschannel_nextclick1 := 0]\n    \n    data[, by_iposdevice_nextclick1 := c(click_time_num[-1], NA), by = c(\"ip\", \"os\", \"device\")]\n    data[, by_iposdevice_nextclick1 := by_iposdevice_nextclick1 - click_time_num]\n    data[is.na(by_iposdevice_nextclick1), by_iposdevice_nextclick1 := 0]\n    \n    data[, by_iposdeviceapp_nextclick1 := c(click_time_num[-1], NA), by = c(\"ip\", \"os\", \"device\", \"app\")]\n    data[, by_iposdeviceapp_nextclick1 := by_iposdeviceapp_nextclick1 - click_time_num]\n    data[is.na(by_iposdeviceapp_nextclick1), by_iposdeviceapp_nextclick1 := 0]\n    \n    data[, by_ipchannel_prevclick1 := c(NA, click_time_num[-.N]), by = c(\"ip\", \"channel\")]\n    data[, by_ipchannel_prevclick1 := click_time_num - by_ipchannel_prevclick1]\n    data[is.na(by_ipchannel_prevclick1), by_ipchannel_prevclick1 := 0]\n    \n    data[, by_ipos_prevclick1 := c(NA, click_time_num[-.N]), by = c(\"ip\", \"os\")]\n    data[, by_ipos_prevclick1 := click_time_num - by_ipos_prevclick1]\n    data[is.na(by_ipos_prevclick1), by_ipos_prevclick1 := 0]\n\n          user  system elapsed \n        281.92    6.17  288.11 \n\nIf computed in a parallel fashion (requires 80GB), it takes approx 80 seconds for all 5. In my code, NA conversion is unnecessary, better use 0 instead of NA in that combine (c) function to speed it even more (will require approx 1GB less RAM in that case).",
      "votes": 7,
      "replies": [
        {
          "id": 323585,
          "postDate": "2018-05-05T15:58:29.173Z",
          "content": "<p>I didn't exptect to get this much of help from people and I really feel that I am lucky enough to have a great adviser. Thank you so much for your help. Have a good evening sir. :)</p>",
          "rawMarkdown": "I didn't exptect to get this much of help from people and I really feel that I am lucky enough to have a great adviser. Thank you so much for your help. Have a good evening sir. :)",
          "votes": 1
        },
        {
          "id": 323596,
          "postDate": "2018-05-05T16:34:37.037Z",
          "content": "<p>I use below syntax (similarly for other grouping) after converting converting click_time to  fastPOSIXct (GMT)\ndf[, ip_os_dev_prevclick:=click_time - shift(click_time, 1, fill = 0, \"lag\"),by = list(ip, os, device)]</p>\n\n<p>Let me know if this works as fast as suggestion given above. </p>",
          "rawMarkdown": "I use below syntax (similarly for other grouping) after converting converting click_time to  fastPOSIXct (GMT)\ndf[, ip_os_dev_prevclick:=click_time - shift(click_time, 1, fill = 0, \"lag\"),by = list(ip, os, device)]\n\nLet me know if this works as fast as suggestion given above. "
        },
        {
          "id": 323597,
          "postDate": "2018-05-05T16:39:36.183Z",
          "content": "<p><a href=\"/laurae2\">@laurae2</a>, is this parallel implementation available out of the box. The only time I see multiple cores being used from data.table is fwrite (I may be wong). From my experience, fread and other calculations using data.table always use single thread only even if I set setDTthreads() to change the numbers</p>",
          "rawMarkdown": "@laurae2, is this parallel implementation available out of the box. The only time I see multiple cores being used from data.table is fwrite (I may be wong). From my experience, fread and other calculations using data.table always use single thread only even if I set setDTthreads() to change the numbers"
        },
        {
          "id": 323602,
          "postDate": "2018-05-05T16:53:02.130Z",
          "content": "<p><a href=\"/lalthan\">@lalthan</a> You need to fork threads (Linux) or create a cluster (Windows), and create specific code to parallelize either asynchronously (futures/promises) or synchronously for the group of operation you want to do.</p>\n\n<p>Using fork it will take significantly less than 80GB (Windows requires copies which explodes RAM quickly).</p>",
          "rawMarkdown": "@lalthan You need to fork threads (Linux) or create a cluster (Windows), and create specific code to parallelize either asynchronously (futures/promises) or synchronously for the group of operation you want to do.\n\nUsing fork it will take significantly less than 80GB (Windows requires copies which explodes RAM quickly).",
          "votes": 2
        }
      ]
    },
    {
      "id": 323532,
      "postDate": "2018-05-05T11:53:42.400Z",
      "content": "<p>Convert click_time to an integer and  use the  <em>shift</em> function.   Assuming adt is a data.table it's quite fast, less than 5 minutes for the whole train+test dataset.</p>\n\n<pre><code>time&lt;-fast_strptime(adt$click_time, format = \"%Y-%m-%d %H:%M:%S\")\ni_time&lt;-as.numeric(time)\nadt$i_time&lt;-i_time\nadt[, deltaUApp:=shift(i_time,1,type = 'lead',fill=0)-i_time,by=list(ip, app, device, os)]) \n</code></pre>\n\n<p>Hope it works for you</p>",
      "rawMarkdown": "Convert click_time to an integer and  use the  *shift* function.   Assuming adt is a data.table it's quite fast, less than 5 minutes for the whole train+test dataset.\n\n    time&lt;-fast_strptime(adt$click_time, format = \"%Y-%m-%d %H:%M:%S\")\n    i_time&lt;-as.numeric(time)\n    adt$i_time&lt;-i_time\n    adt[, deltaUApp:=shift(i_time,1,type = 'lead',fill=0)-i_time,by=list(ip, app, device, os)]) \n\n\nHope it works for you\n\n   \n  \n",
      "votes": 3,
      "replies": [
        {
          "id": 323539,
          "postDate": "2018-05-05T12:30:44.237Z",
          "content": "<p>Thank you for your advise. I think it is really valid one, So I want to play with this tonight. I wish I could share some advise with other new comers in the future just like you did for me. Thanks a lot. :)</p>",
          "rawMarkdown": "Thank you for your advise. I think it is really valid one, So I want to play with this tonight. I wish I could share some advise with other new comers in the future just like you did for me. Thanks a lot. :)",
          "votes": 2
        }
      ]
    },
    {
      "id": 323583,
      "postDate": "2018-05-05T15:55:50.673Z",
      "content": "<p>If you want to check multiple combos, the following does work for me on 16gb mbp...</p>\n\n<pre><code>xtr &lt;- fread(paste('../input/train.csv', sep = ''))\nxtr[ ,click_time := parse_date_time(click_time, orders = 'YmdHMS')]\nxcols &lt;- c('ip', 'app', 'device', 'os', 'channel')\n\nfor (nelem in 1:4)\n{\n  xcomb &lt;- combn(xcols, nelem)\n\n  for (ii in 1:ncol(xcomb))\n  {\n\n    # next \n    xnam &lt;- paste('nxt', paste(str_sub(xcomb[,ii],1,1), sep = '', collapse = ''), sep = '_')\n    xtr[ ,next_time := shift(click_time, n= 1, type = 'lead'), by = c(xcomb[,ii])]\n    xtr[ ,x := as.numeric(next_time - click_time, units=  'secs')]\n    xtr[ ,next_time := NULL, with = TRUE]\n    xc &lt;- c('click_id', 'x'); xmat &lt;- xtr[, ..xc, with = TRUE]; xtr[, x:= NULL, with = TRUE]\n    xmat$x &lt;- round(log1p(xmat$x)); \n    xmat$x[is.na(xmat$x)] &lt;- -1; setnames(xmat, 'x', xnam)\n\n      save(xmat, file = paste('../input/', paste(xnam, '_f',fnam, sep = ''), '.RData', sep = ''))\n\n    rm(xmat)\n\n    # last \n    xnam &amp;lt;- paste('lst', paste(str_sub(xcomb[,ii],1,1), sep = '', collapse = ''), sep = '_')\n    xtr[ ,next_time := shift(click_time, n= 1, type = 'lag'), by = c(xcomb[,ii])]\n    xtr[ ,x := as.numeric(click_time - next_time, units=  'secs')]\n    xtr[ ,next_time := NULL, with = TRUE]\n    xc &lt;- c('click_id', 'x'); xmat &amp;lt;- xtr[, ..xc, with = TRUE]; xtr[, x:= NULL, with = TRUE]\n    xmat$x &lt;- round(log1p(xmat$x),2); xmat$x[is.na(xmat$x)] &lt;- -1; setnames(xmat, 'x', xnam)\n\n      save(xmat, file = paste('../input/', paste(xnam, '_f',fnam, sep = ''), '.RData', sep = ''))\n\n    rm(xmat)\n\n  }\n  print(paste(nelem,'way combos:done', sep = '-'))\n}\n</code></pre>",
      "rawMarkdown": "If you want to check multiple combos, the following does work for me on 16gb mbp...\n\n    xtr &lt;- fread(paste('../input/train.csv', sep = ''))\n    xtr[ ,click_time := parse_date_time(click_time, orders = 'YmdHMS')]\n    xcols &lt;- c('ip', 'app', 'device', 'os', 'channel')\n    \n    for (nelem in 1:4)\n    {\n      xcomb &lt;- combn(xcols, nelem)\n      \n      for (ii in 1:ncol(xcomb))\n      {\n        \n        # next \n        xnam &lt;- paste('nxt', paste(str_sub(xcomb[,ii],1,1), sep = '', collapse = ''), sep = '_')\n        xtr[ ,next_time := shift(click_time, n= 1, type = 'lead'), by = c(xcomb[,ii])]\n        xtr[ ,x := as.numeric(next_time - click_time, units=  'secs')]\n        xtr[ ,next_time := NULL, with = TRUE]\n        xc &lt;- c('click_id', 'x'); xmat &lt;- xtr[, ..xc, with = TRUE]; xtr[, x:= NULL, with = TRUE]\n        xmat$x &lt;- round(log1p(xmat$x)); \n        xmat$x[is.na(xmat$x)] &lt;- -1; setnames(xmat, 'x', xnam)\n        \n          save(xmat, file = paste('../input/', paste(xnam, '_f',fnam, sep = ''), '.RData', sep = ''))\n    \n        rm(xmat)\n        \n        # last \n        xnam &lt;- paste('lst', paste(str_sub(xcomb[,ii],1,1), sep = '', collapse = ''), sep = '_')\n        xtr[ ,next_time := shift(click_time, n= 1, type = 'lag'), by = c(xcomb[,ii])]\n        xtr[ ,x := as.numeric(click_time - next_time, units=  'secs')]\n        xtr[ ,next_time := NULL, with = TRUE]\n        xc &lt;- c('click_id', 'x'); xmat &lt;- xtr[, ..xc, with = TRUE]; xtr[, x:= NULL, with = TRUE]\n        xmat$x &lt;- round(log1p(xmat$x),2); xmat$x[is.na(xmat$x)] &lt;- -1; setnames(xmat, 'x', xnam)\n        \n          save(xmat, file = paste('../input/', paste(xnam, '_f',fnam, sep = ''), '.RData', sep = ''))\n        \n        rm(xmat)\n        \n      }\n      print(paste(nelem,'way combos:done', sep = '-'))\n    }\n\n",
      "votes": 1
    },
    {
      "id": 323529,
      "postDate": "2018-05-05T11:29:12.237Z",
      "content": "<p>Hello, everyone. <br>\nI am recently started to learn R and I find it really interesting. I need your help on creating 'Next Click' feature in R. The code below is what I made for it.</p>\n\n<pre><code>tdiff &lt;- function(tdiff) {\n  vc &lt;- NULL\n  for (x in 1:length(tdiff)) {\n    interval &lt;- interval(tdiff[x], tdiff[x+1])\n    diffsecond &lt;- as.integer(seconds(interval))\n    vc &lt;- append(vc, diffsecond)\n  }\n  vc &lt;- ifelse(is.na(vc), 0, vc)\n  vc &lt;- append(vc, 0)\n  return(vc)\n}\n\nadt[, clicker_Next := tdiff(click_time), by = list(ip, device, os)]\n</code></pre>\n\n<p>When I try to run this code through Train_Sample, it is working but takes too long about more than 15 mins. But when it comes to Train CSV, this code doesn't work at all. If you could share your experty on this, it would be much appreciated.  </p>\n\n<p>Thank you for your help in advance.</p>",
      "rawMarkdown": "Hello, everyone.  \nI am recently started to learn R and I find it really interesting. I need your help on creating 'Next Click' feature in R. The code below is what I made for it.\n\n    tdiff &lt;- function(tdiff) {\n      vc &lt;- NULL\n      for (x in 1:length(tdiff)) {\n        interval &lt;- interval(tdiff[x], tdiff[x+1])\n        diffsecond &lt;- as.integer(seconds(interval))\n        vc &lt;- append(vc, diffsecond)\n      }\n      vc &lt;- ifelse(is.na(vc), 0, vc)\n      vc &lt;- append(vc, 0)\n      return(vc)\n    }\n    \n    adt[, clicker_Next := tdiff(click_time), by = list(ip, device, os)]\n\n\nWhen I try to run this code through Train_Sample, it is working but takes too long about more than 15 mins. But when it comes to Train CSV, this code doesn't work at all. If you could share your experty on this, it would be much appreciated.  \n\nThank you for your help in advance.\n\n",
      "votes": 1
    },
    {
      "id": 323531,
      "postDate": "2018-05-05T11:40:19.207Z",
      "content": "<p>Computing this is also quite long with pandas in Python.  You may just want to let your code run longer.  If your code breaks, which is worse, then I'm afraid you'll need someone more expert than me on R to help.</p>",
      "rawMarkdown": "Computing this is also quite long with pandas in Python.  You may just want to let your code run longer.  If your code breaks, which is worse, then I'm afraid you'll need someone more expert than me on R to help.",
      "replies": [
        {
          "id": 323537,
          "postDate": "2018-05-05T12:29:18.907Z",
          "content": "<p>Thank you for your attention on my question. Have a great day!</p>",
          "rawMarkdown": "Thank you for your attention on my question. Have a great day!",
          "votes": 2
        }
      ]
    },
    {
      "id": 323606,
      "postDate": "2018-05-05T17:02:04.527Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 323580,
      "author_name": "Laurae",
      "author_url": "",
      "post_date": "2018-05-05T15:33:27.970000",
      "content": "<p><code>shift</code> is known to be slow when too many groups are involved in data.table (because the <code>:=</code> operator is not <code>gforce</code>-optimized yet in data.table)</p>\n\n<p>You can do like the following for even faster lag/lead deltas:</p>\n\n<pre><code>adt[, deltaUApp_nextclick1 := c(click_time_num[-1], NA), by = list(ip, app, device, os)]\nadt[, deltaUApp_nextclick1 := deltaUApp_nextclick1 - click_time_num]\nadt[is.na(deltaUApp_nextclick1), deltaUApp_nextclick1 := 0]\n\nadt[, deltaUApp_prevclick1 := c(NA, click_time_num[-.N]), by = list(ip, app, device, os)]\nadt[, deltaUApp_prevclick1 := deltaUApp_prevclick1 - click_time_num]\nadt[is.na(deltaUApp_prevclick1), deltaUApp_prevclick1 := 0]\n</code></pre>\n\n<p>It assumes you have a integer/numeric time called <code>click_time_num</code>.</p>\n\n<p>Takes less than 5 minutes to do it (i7-7700K, 5.0 GHz) for:</p>\n\n<ul>\n<li>all the 5 groups used on the top kernels</li>\n<li>with ALL data (train + test_supplement)</li>\n<li>under 16GB peak (OS not included)</li>\n</ul>\n\n<p>text</p>\n\n<pre><code>data[, click_time_num := as.numeric(click_time)]\n\ndata[, by_ipappdeviceoschannel_nextclick1 := c(click_time_num[-1], NA), by = c(\"ip\", \"app\", \"device\", \"os\", \"channel\")]\ndata[, by_ipappdeviceoschannel_nextclick1 := by_ipappdeviceoschannel_nextclick1 - click_time_num]\ndata[is.na(by_ipappdeviceoschannel_nextclick1), by_ipappdeviceoschannel_nextclick1 := 0]\n\ndata[, by_iposdevice_nextclick1 := c(click_time_num[-1], NA), by = c(\"ip\", \"os\", \"device\")]\ndata[, by_iposdevice_nextclick1 := by_iposdevice_nextclick1 - click_time_num]\ndata[is.na(by_iposdevice_nextclick1), by_iposdevice_nextclick1 := 0]\n\ndata[, by_iposdeviceapp_nextclick1 := c(click_time_num[-1], NA), by = c(\"ip\", \"os\", \"device\", \"app\")]\ndata[, by_iposdeviceapp_nextclick1 := by_iposdeviceapp_nextclick1 - click_time_num]\ndata[is.na(by_iposdeviceapp_nextclick1), by_iposdeviceapp_nextclick1 := 0]\n\ndata[, by_ipchannel_prevclick1 := c(NA, click_time_num[-.N]), by = c(\"ip\", \"channel\")]\ndata[, by_ipchannel_prevclick1 := click_time_num - by_ipchannel_prevclick1]\ndata[is.na(by_ipchannel_prevclick1), by_ipchannel_prevclick1 := 0]\n\ndata[, by_ipos_prevclick1 := c(NA, click_time_num[-.N]), by = c(\"ip\", \"os\")]\ndata[, by_ipos_prevclick1 := click_time_num - by_ipos_prevclick1]\ndata[is.na(by_ipos_prevclick1), by_ipos_prevclick1 := 0]\n\n      user  system elapsed \n    281.92    6.17  288.11 \n</code></pre>\n\n<p>If computed in a parallel fashion (requires 80GB), it takes approx 80 seconds for all 5. In my code, NA conversion is unnecessary, better use 0 instead of NA in that combine (c) function to speed it even more (will require approx 1GB less RAM in that case).</p>",
      "votes": 7,
      "replies": [
        {
          "id": 323585,
          "author_name": "LegenDaD",
          "author_url": "",
          "post_date": "2018-05-05T15:58:29.173000",
          "content": "<p>I didn't exptect to get this much of help from people and I really feel that I am lucky enough to have a great adviser. Thank you so much for your help. Have a good evening sir. :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 323596,
          "author_name": "lalthan",
          "author_url": "",
          "post_date": "2018-05-05T16:34:37.037000",
          "content": "<p>I use below syntax (similarly for other grouping) after converting converting click_time to  fastPOSIXct (GMT)\ndf[, ip_os_dev_prevclick:=click_time - shift(click_time, 1, fill = 0, \"lag\"),by = list(ip, os, device)]</p>\n\n<p>Let me know if this works as fast as suggestion given above. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 323597,
          "author_name": "lalthan",
          "author_url": "",
          "post_date": "2018-05-05T16:39:36.183000",
          "content": "<p><a href=\"/laurae2\">@laurae2</a>, is this parallel implementation available out of the box. The only time I see multiple cores being used from data.table is fwrite (I may be wong). From my experience, fread and other calculations using data.table always use single thread only even if I set setDTthreads() to change the numbers</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 323602,
          "author_name": "Laurae",
          "author_url": "",
          "post_date": "2018-05-05T16:53:02.130000",
          "content": "<p><a href=\"/lalthan\">@lalthan</a> You need to fork threads (Linux) or create a cluster (Windows), and create specific code to parallelize either asynchronously (futures/promises) or synchronously for the group of operation you want to do.</p>\n\n<p>Using fork it will take significantly less than 80GB (Windows requires copies which explodes RAM quickly).</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 323532,
      "author_name": "Vicens Gaitan",
      "author_url": "",
      "post_date": "2018-05-05T11:53:42.400000",
      "content": "<p>Convert click_time to an integer and  use the  <em>shift</em> function.   Assuming adt is a data.table it's quite fast, less than 5 minutes for the whole train+test dataset.</p>\n\n<pre><code>time&lt;-fast_strptime(adt$click_time, format = \"%Y-%m-%d %H:%M:%S\")\ni_time&lt;-as.numeric(time)\nadt$i_time&lt;-i_time\nadt[, deltaUApp:=shift(i_time,1,type = 'lead',fill=0)-i_time,by=list(ip, app, device, os)]) \n</code></pre>\n\n<p>Hope it works for you</p>",
      "votes": 3,
      "replies": [
        {
          "id": 323539,
          "author_name": "LegenDaD",
          "author_url": "",
          "post_date": "2018-05-05T12:30:44.237000",
          "content": "<p>Thank you for your advise. I think it is really valid one, So I want to play with this tonight. I wish I could share some advise with other new comers in the future just like you did for me. Thanks a lot. :)</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 323583,
      "author_name": "Konrad Banachewicz",
      "author_url": "",
      "post_date": "2018-05-05T15:55:50.673000",
      "content": "<p>If you want to check multiple combos, the following does work for me on 16gb mbp...</p>\n\n<pre><code>xtr &lt;- fread(paste('../input/train.csv', sep = ''))\nxtr[ ,click_time := parse_date_time(click_time, orders = 'YmdHMS')]\nxcols &lt;- c('ip', 'app', 'device', 'os', 'channel')\n\nfor (nelem in 1:4)\n{\n  xcomb &lt;- combn(xcols, nelem)\n\n  for (ii in 1:ncol(xcomb))\n  {\n\n    # next \n    xnam &lt;- paste('nxt', paste(str_sub(xcomb[,ii],1,1), sep = '', collapse = ''), sep = '_')\n    xtr[ ,next_time := shift(click_time, n= 1, type = 'lead'), by = c(xcomb[,ii])]\n    xtr[ ,x := as.numeric(next_time - click_time, units=  'secs')]\n    xtr[ ,next_time := NULL, with = TRUE]\n    xc &lt;- c('click_id', 'x'); xmat &lt;- xtr[, ..xc, with = TRUE]; xtr[, x:= NULL, with = TRUE]\n    xmat$x &lt;- round(log1p(xmat$x)); \n    xmat$x[is.na(xmat$x)] &lt;- -1; setnames(xmat, 'x', xnam)\n\n      save(xmat, file = paste('../input/', paste(xnam, '_f',fnam, sep = ''), '.RData', sep = ''))\n\n    rm(xmat)\n\n    # last \n    xnam &amp;lt;- paste('lst', paste(str_sub(xcomb[,ii],1,1), sep = '', collapse = ''), sep = '_')\n    xtr[ ,next_time := shift(click_time, n= 1, type = 'lag'), by = c(xcomb[,ii])]\n    xtr[ ,x := as.numeric(click_time - next_time, units=  'secs')]\n    xtr[ ,next_time := NULL, with = TRUE]\n    xc &lt;- c('click_id', 'x'); xmat &amp;lt;- xtr[, ..xc, with = TRUE]; xtr[, x:= NULL, with = TRUE]\n    xmat$x &lt;- round(log1p(xmat$x),2); xmat$x[is.na(xmat$x)] &lt;- -1; setnames(xmat, 'x', xnam)\n\n      save(xmat, file = paste('../input/', paste(xnam, '_f',fnam, sep = ''), '.RData', sep = ''))\n\n    rm(xmat)\n\n  }\n  print(paste(nelem,'way combos:done', sep = '-'))\n}\n</code></pre>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 323531,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2018-05-05T11:40:19.207000",
      "content": "<p>Computing this is also quite long with pandas in Python.  You may just want to let your code run longer.  If your code breaks, which is worse, then I'm afraid you'll need someone more expert than me on R to help.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 323537,
          "author_name": "LegenDaD",
          "author_url": "",
          "post_date": "2018-05-05T12:29:18.907000",
          "content": "<p>Thank you for your attention on my question. Have a great day!</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 323606,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-05T17:02:04.527000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "323580": "`shift` is known to be slow when too many groups are involved in data.table (because the `:=` operator is not `gforce`-optimized yet in data.table)\n\nYou can do like the following for even faster lag/lead deltas:\n\n    adt[, deltaUApp_nextclick1 := c(click_time_num[-1], NA), by = list(ip, app, device, os)]\n    adt[, deltaUApp_nextclick1 := deltaUApp_nextclick1 - click_time_num]\n    adt[is.na(deltaUApp_nextclick1), deltaUApp_nextclick1 := 0]\n    \n    adt[, deltaUApp_prevclick1 := c(NA, click_time_num[-.N]), by = list(ip, app, device, os)]\n    adt[, deltaUApp_prevclick1 := deltaUApp_prevclick1 - click_time_num]\n    adt[is.na(deltaUApp_prevclick1), deltaUApp_prevclick1 := 0]\n\nIt assumes you have a integer/numeric time called `click_time_num`.\n\nTakes less than 5 minutes to do it (i7-7700K, 5.0 GHz) for:\n\n* all the 5 groups used on the top kernels\n* with ALL data (train + test_supplement)\n* under 16GB peak (OS not included)\n\ntext\n\n    data[, click_time_num := as.numeric(click_time)]\n\n    data[, by_ipappdeviceoschannel_nextclick1 := c(click_time_num[-1], NA), by = c(\"ip\", \"app\", \"device\", \"os\", \"channel\")]\n    data[, by_ipappdeviceoschannel_nextclick1 := by_ipappdeviceoschannel_nextclick1 - click_time_num]\n    data[is.na(by_ipappdeviceoschannel_nextclick1), by_ipappdeviceoschannel_nextclick1 := 0]\n    \n    data[, by_iposdevice_nextclick1 := c(click_time_num[-1], NA), by = c(\"ip\", \"os\", \"device\")]\n    data[, by_iposdevice_nextclick1 := by_iposdevice_nextclick1 - click_time_num]\n    data[is.na(by_iposdevice_nextclick1), by_iposdevice_nextclick1 := 0]\n    \n    data[, by_iposdeviceapp_nextclick1 := c(click_time_num[-1], NA), by = c(\"ip\", \"os\", \"device\", \"app\")]\n    data[, by_iposdeviceapp_nextclick1 := by_iposdeviceapp_nextclick1 - click_time_num]\n    data[is.na(by_iposdeviceapp_nextclick1), by_iposdeviceapp_nextclick1 := 0]\n    \n    data[, by_ipchannel_prevclick1 := c(NA, click_time_num[-.N]), by = c(\"ip\", \"channel\")]\n    data[, by_ipchannel_prevclick1 := click_time_num - by_ipchannel_prevclick1]\n    data[is.na(by_ipchannel_prevclick1), by_ipchannel_prevclick1 := 0]\n    \n    data[, by_ipos_prevclick1 := c(NA, click_time_num[-.N]), by = c(\"ip\", \"os\")]\n    data[, by_ipos_prevclick1 := click_time_num - by_ipos_prevclick1]\n    data[is.na(by_ipos_prevclick1), by_ipos_prevclick1 := 0]\n\n          user  system elapsed \n        281.92    6.17  288.11 \n\nIf computed in a parallel fashion (requires 80GB), it takes approx 80 seconds for all 5. In my code, NA conversion is unnecessary, better use 0 instead of NA in that combine (c) function to speed it even more (will require approx 1GB less RAM in that case).",
    "323532": "Convert click_time to an integer and  use the  *shift* function.   Assuming adt is a data.table it's quite fast, less than 5 minutes for the whole train+test dataset.\n\n    time&lt;-fast_strptime(adt$click_time, format = \"%Y-%m-%d %H:%M:%S\")\n    i_time&lt;-as.numeric(time)\n    adt$i_time&lt;-i_time\n    adt[, deltaUApp:=shift(i_time,1,type = 'lead',fill=0)-i_time,by=list(ip, app, device, os)]) \n\n\nHope it works for you\n\n   \n  \n",
    "323583": "If you want to check multiple combos, the following does work for me on 16gb mbp...\n\n    xtr &lt;- fread(paste('../input/train.csv', sep = ''))\n    xtr[ ,click_time := parse_date_time(click_time, orders = 'YmdHMS')]\n    xcols &lt;- c('ip', 'app', 'device', 'os', 'channel')\n    \n    for (nelem in 1:4)\n    {\n      xcomb &lt;- combn(xcols, nelem)\n      \n      for (ii in 1:ncol(xcomb))\n      {\n        \n        # next \n        xnam &lt;- paste('nxt', paste(str_sub(xcomb[,ii],1,1), sep = '', collapse = ''), sep = '_')\n        xtr[ ,next_time := shift(click_time, n= 1, type = 'lead'), by = c(xcomb[,ii])]\n        xtr[ ,x := as.numeric(next_time - click_time, units=  'secs')]\n        xtr[ ,next_time := NULL, with = TRUE]\n        xc &lt;- c('click_id', 'x'); xmat &lt;- xtr[, ..xc, with = TRUE]; xtr[, x:= NULL, with = TRUE]\n        xmat$x &lt;- round(log1p(xmat$x)); \n        xmat$x[is.na(xmat$x)] &lt;- -1; setnames(xmat, 'x', xnam)\n        \n          save(xmat, file = paste('../input/', paste(xnam, '_f',fnam, sep = ''), '.RData', sep = ''))\n    \n        rm(xmat)\n        \n        # last \n        xnam &lt;- paste('lst', paste(str_sub(xcomb[,ii],1,1), sep = '', collapse = ''), sep = '_')\n        xtr[ ,next_time := shift(click_time, n= 1, type = 'lag'), by = c(xcomb[,ii])]\n        xtr[ ,x := as.numeric(click_time - next_time, units=  'secs')]\n        xtr[ ,next_time := NULL, with = TRUE]\n        xc &lt;- c('click_id', 'x'); xmat &lt;- xtr[, ..xc, with = TRUE]; xtr[, x:= NULL, with = TRUE]\n        xmat$x &lt;- round(log1p(xmat$x),2); xmat$x[is.na(xmat$x)] &lt;- -1; setnames(xmat, 'x', xnam)\n        \n          save(xmat, file = paste('../input/', paste(xnam, '_f',fnam, sep = ''), '.RData', sep = ''))\n        \n        rm(xmat)\n        \n      }\n      print(paste(nelem,'way combos:done', sep = '-'))\n    }\n\n",
    "323529": "Hello, everyone.  \nI am recently started to learn R and I find it really interesting. I need your help on creating 'Next Click' feature in R. The code below is what I made for it.\n\n    tdiff &lt;- function(tdiff) {\n      vc &lt;- NULL\n      for (x in 1:length(tdiff)) {\n        interval &lt;- interval(tdiff[x], tdiff[x+1])\n        diffsecond &lt;- as.integer(seconds(interval))\n        vc &lt;- append(vc, diffsecond)\n      }\n      vc &lt;- ifelse(is.na(vc), 0, vc)\n      vc &lt;- append(vc, 0)\n      return(vc)\n    }\n    \n    adt[, clicker_Next := tdiff(click_time), by = list(ip, device, os)]\n\n\nWhen I try to run this code through Train_Sample, it is working but takes too long about more than 15 mins. But when it comes to Train CSV, this code doesn't work at all. If you could share your experty on this, it would be much appreciated.  \n\nThank you for your help in advance.\n\n",
    "323531": "Computing this is also quite long with pandas in Python.  You may just want to let your code run longer.  If your code breaks, which is worse, then I'm afraid you'll need someone more expert than me on R to help.",
    "323606": ""
  }
}