{
  "id": 53361,
  "title": "Time Deltas",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/53361",
  "author_name": "",
  "post_date": "2018-03-29T17:14:28.155735Z",
  "votes": 18,
  "comment_count": 38,
  "views": 0,
  "content": "<p>I looked at time deltas grouped by ip,app,device,os,channel on the data - looks like quite a different distribution between zeros and ones</p>",
  "messages": [
    {
      "id": "305960",
      "postDate": "03/29/2018 17:14:28",
      "content": "<p>I looked at time deltas grouped by ip,app,device,os,channel on the data - looks like quite a different distribution between zeros and ones</p>",
      "rawMarkdown": "I looked at time deltas grouped by ip,app,device,os,channel on the data - looks like quite a different distribution between zeros and ones",
      "votes": null
    },
    {
      "id": "306037",
      "postDate": "03/29/2018 19:19:56",
      "content": "<p>What do you mean by \"time deltas\"? The time between successive clicks from a given group?</p>",
      "rawMarkdown": "What do you mean by \"time deltas\"? The time between successive clicks from a given group?",
      "votes": null
    },
    {
      "id": "306060",
      "postDate": "03/29/2018 20:09:06",
      "content": "<p>Indeed</p>",
      "rawMarkdown": "Indeed",
      "votes": null
    },
    {
      "id": "306129",
      "postDate": "03/29/2018 22:28:25",
      "content": "<p>Hi, Scirpus, could we have some explanation of this figure? like what's the meaning of x and y axis and what does color blue and yellow present? Thank you very much!</p>",
      "rawMarkdown": "Hi, Scirpus, could we have some explanation of this figure? like what's the meaning of x and y axis and what does color blue and yellow present? Thank you very much!",
      "votes": null
    },
    {
      "id": "306273",
      "postDate": "03/30/2018 06:52:04",
      "content": "<p>\"looks like quite a different distribution between zeros and ones\" tells you the colors of the plots represent the is_attributed  zeros and ones</p>\n\n<p>x axis is \"time deltas grouped by ip,app,device,os,channel on the data\" and the y axis is number of incidences per x value</p>\n\n<p>Hope this helps</p>",
      "rawMarkdown": "\"looks like quite a different distribution between zeros and ones\" tells you the colors of the plots represent the is_attributed  zeros and ones\n\nx axis is \"time deltas grouped by ip,app,device,os,channel on the data\" and the y axis is number of incidences per x value\n\nHope this helps",
      "votes": null
    },
    {
      "id": "306369",
      "postDate": "03/30/2018 10:07:55",
      "content": "<p>So this chart is grouped by all 5 categories at once?</p>",
      "rawMarkdown": "So this chart is grouped by all 5 categories at once?",
      "votes": null
    },
    {
      "id": "306371",
      "postDate": "03/30/2018 10:13:39",
      "content": "<p>That is true - I am just trying it now to see if it improves things.</p>\n\n<p>I wrote a little c++ app to do it as pandas was <strong><em>really really slow</em></strong> on all the data (time-ordered, concatenated on train,test and test supplement) do you want a copy?</p>",
      "rawMarkdown": "That is true - I am just trying it now to see if it improves things.\n\nI wrote a little c++ app to do it as pandas was ***really really slow*** on all the data (time-ordered, concatenated on train,test and test supplement) do you want a copy?",
      "votes": null
    },
    {
      "id": "306624",
      "postDate": "03/30/2018 19:00:43",
      "content": "<p>How do you define it when there is a single observation in the group?</p>",
      "rawMarkdown": "How do you define it when there is a single observation in the group?",
      "votes": null
    },
    {
      "id": "306625",
      "postDate": "03/30/2018 19:01:58",
      "content": "<p>test is included in test_supplement, therefore you should concatenate train and test_supplement only.</p>",
      "rawMarkdown": "test is included in test_supplement, therefore you should concatenate train and test_supplement only.",
      "votes": null
    },
    {
      "id": "306630",
      "postDate": "03/30/2018 19:10:22",
      "content": "<p>I chose NAN</p>",
      "rawMarkdown": "I chose NAN",
      "votes": null
    },
    {
      "id": "306635",
      "postDate": "03/30/2018 19:16:22",
      "content": "<p>Thanks for this I completely missed this owing to the fact that I read supplement and assumed kaggle was using the word supplement to actually mean something.  Why on earth didn't they can it \"originaltestset\" goddamn it! </p>",
      "rawMarkdown": "Thanks for this I completely missed this owing to the fact that I read supplement and assumed kaggle was using the word supplement to actually mean something.  Why on earth didn't they can it \"originaltestset\" goddamn it!",
      "votes": null
    },
    {
      "id": "306642",
      "postDate": "03/30/2018 19:25:23",
      "content": "<p>No pb.  Sometimes a fresh eye is good ;)</p>",
      "rawMarkdown": "No pb.  Sometimes a fresh eye is good ;)",
      "votes": null
    },
    {
      "id": "306648",
      "postDate": "03/30/2018 19:35:05",
      "content": "<p>Thanks for stopping me going down a rabbit hole!  I have just published the code in the discussion.  It is pretty generic and may come in handy for you in future competitions!</p>",
      "rawMarkdown": "Thanks for stopping me going down a rabbit hole!  I have just published the code in the discussion.  It is pretty generic and may come in handy for you in future competitions!",
      "votes": null
    },
    {
      "id": "306661",
      "postDate": "03/30/2018 19:52:33",
      "content": "<p>Is time deltas difference in time of successive clicks of each group?</p>",
      "rawMarkdown": "Is time deltas difference in time of successive clicks of each group?",
      "votes": null
    },
    {
      "id": "306712",
      "postDate": "03/30/2018 21:20:52",
      "content": "<p>Yes. (See the top of this thread.)</p>",
      "rawMarkdown": "Yes. (See the top of this thread.)",
      "votes": null
    },
    {
      "id": "306774",
      "postDate": "03/31/2018 01:35:30",
      "content": "<p>Time Deltas on what? Any explanation thanks a lot </p>",
      "rawMarkdown": "Time Deltas on what? Any explanation thanks a lot",
      "votes": null
    },
    {
      "id": "306883",
      "postDate": "03/31/2018 08:35:13",
      "content": "<p>Look at the code in the <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53450\">link</a></p>",
      "rawMarkdown": "Look at the code in the [link][1]\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53450",
      "votes": null
    },
    {
      "id": "308242",
      "postDate": "04/03/2018 07:36:11",
      "content": "<p>this also looked like a good feature to me. but when i tried it didnt help the score. it made it even worse. i used different grouping like ip alone and ip, device, os and so on.</p>",
      "rawMarkdown": "this also looked like a good feature to me. but when i tried it didnt help the score. it made it even worse. i used different grouping like ip alone and ip, device, os and so on.",
      "votes": null
    },
    {
      "id": "308393",
      "postDate": "04/03/2018 13:03:00",
      "content": "<p>Did you time order the data ?</p>",
      "rawMarkdown": "Did you time order the data ?",
      "votes": null
    },
    {
      "id": "308425",
      "postDate": "04/03/2018 13:49:34",
      "content": "<p>Yes I did. For example sorting by ip and click time and then doing a diff row wise. Also tried sort by [ip, device,os, click_time] . </p>",
      "rawMarkdown": "Yes I did. For example sorting by ip and click time and then doing a diff row wise. Also tried sort by [ip, device,os, click_time] .",
      "votes": null
    },
    {
      "id": "308430",
      "postDate": "04/03/2018 13:58:10",
      "content": "<p>And you used test supplement rather than test?</p>",
      "rawMarkdown": "And you used test supplement rather than test?",
      "votes": null
    },
    {
      "id": "308444",
      "postDate": "04/03/2018 14:25:40",
      "content": "<p>I only tried with cv I think. I have to recheck maybe. Did it help your score a lot?</p>",
      "rawMarkdown": "I only tried with cv I think. I have to recheck maybe. Did it help your score a lot?",
      "votes": null
    },
    {
      "id": "308453",
      "postDate": "04/03/2018 14:35:51",
      "content": "<p>It helped a little (.002) moving me up to the top 50 </p>",
      "rawMarkdown": "It helped a little (.002) moving me up to the top 50",
      "votes": null
    },
    {
      "id": "309738",
      "postDate": "04/05/2018 22:07:49",
      "content": "<p>Jumping in on the thread cause I can't find answer anywhere else and it seems relevant here, as I understand you are concatenating old test and train... </p>\n\n<p>What do you think it means that the old test and train overlap?  There is about an hour and a half of data that covers both train and test.  Does that mean that that hour of train is missing data or did they duplicate it in test (didn't look that way on cursory exploration)?    That hour actually corresponds to the time on the real test, and if it's really missing, it can make those hours misleading in a model (i.e. the counts in those times were actually higher than what we are training on)...  Any thoughts on what happened there or the effect?</p>\n\n<p>Specifically  train ends on 2017:11:09 at 16:00:00, while old test starts on 2017-11-06 14:32:21.  So the time between 14:32:21 and 16:00:00 on the 9th some clicks are in train and some on test.  </p>\n\n<p>PS:  I was working on similar time_delta feature as well and one thing to keep in mind is that there are many duplicate clicks, which means that if you just sort them by click_time they may not stay in their original order, and thus a time_delta of a set of 2-3 clicks that just launched could have wrong click assigned 0 and the longer delta...  </p>",
      "rawMarkdown": "Jumping in on the thread cause I can't find answer anywhere else and it seems relevant here, as I understand you are concatenating old test and train... \n\nWhat do you think it means that the old test and train overlap?  There is about an hour and a half of data that covers both train and test.  Does that mean that that hour of train is missing data or did they duplicate it in test (didn't look that way on cursory exploration)?    That hour actually corresponds to the time on the real test, and if it's really missing, it can make those hours misleading in a model (i.e. the counts in those times were actually higher than what we are training on)...  Any thoughts on what happened there or the effect?\n\nSpecifically  train ends on 2017:11:09 at 16:00:00, while old test starts on 2017-11-06 14:32:21.  So the time between 14:32:21 and 16:00:00 on the 9th some clicks are in train and some on test.  \n\nPS:  I was working on similar time_delta feature as well and one thing to keep in mind is that there are many duplicate clicks, which means that if you just sort them by click_time they may not stay in their original order, and thus a time_delta of a set of 2-3 clicks that just launched could have wrong click assigned 0 and the longer delta...",
      "votes": null
    },
    {
      "id": "309743",
      "postDate": "04/05/2018 22:30:38",
      "content": "<p>Re: your last point - nice observation, I think we can just leave the clicks in the original order for time delta computations because they're already time sorted?</p>",
      "rawMarkdown": "Re: your last point - nice observation, I think we can just leave the clicks in the original order for time delta computations because they're already time sorted?",
      "votes": null
    },
    {
      "id": "309750",
      "postDate": "04/05/2018 22:44:54",
      "content": "<p>It's just a 'be careful' kinda note because if people sort/resort/concatenate data for various groupings, the original flow can get lost.  Something like that happened to me in one of my tries and took me some time to realize it was happening.  </p>\n\n<p>I'm doing this in pandas and it's taking forever, so these kind of errors have literally cost me days to untangle.  Currently in the process of fixing something along those lines and really wishing had better coding skills/faster computer...  </p>\n\n<p>Also the issue is the blending of train/test times that overlap.  a) how to understand it b) if accept it and get clicks that are duplicates in both, which came before/after?  splicing the two together really affects the time range after 14:00 which can hurt the model, but won't be noticed on the public leader board.</p>",
      "rawMarkdown": "It's just a 'be careful' kinda note because if people sort/resort/concatenate data for various groupings, the original flow can get lost.  Something like that happened to me in one of my tries and took me some time to realize it was happening.  \n\nI'm doing this in pandas and it's taking forever, so these kind of errors have literally cost me days to untangle.  Currently in the process of fixing something along those lines and really wishing had better coding skills/faster computer...  \n\nAlso the issue is the blending of train/test times that overlap.  a) how to understand it b) if accept it and get clicks that are duplicates in both, which came before/after?  splicing the two together really affects the time range after 14:00 which can hurt the model, but won't be noticed on the public leader board.",
      "votes": null
    },
    {
      "id": "309942",
      "postDate": "04/06/2018 08:33:46",
      "content": "<p>the amount of transactions in test test in the 1.5 hour of overlap is too little (&lt; 1000) and i think could be just simply ignored.</p>",
      "rawMarkdown": "the amount of transactions in test test in the 1.5 hour of overlap is too little (&lt; 1000) and i think could be just simply ignored.",
      "votes": null
    },
    {
      "id": "310454",
      "postDate": "04/07/2018 15:33:10",
      "content": "<p>I've done the calculation <a href=\"https://www.kaggle.com/aharless/talkingdata-time-deltas\">in Pandas</a> (for just the training set in this case).  It runs in reasonable time (a significant part of the kernel time is just writing to disk), but I had to jump through a lot of hoops to avoid breaching the Kaggle memory constraint, and the disk space constraint didn't permit me to write out the whole training file with deltas (although one could write just the deltas and do the combination later when reading).</p>",
      "rawMarkdown": "I've done the calculation [in Pandas][1] (for just the training set in this case).  It runs in reasonable time (a significant part of the kernel time is just writing to disk), but I had to jump through a lot of hoops to avoid breaching the Kaggle memory constraint, and the disk space constraint didn't permit me to write out the whole training file with deltas (although one could write just the deltas and do the combination later when reading).\n\n [1]: https://www.kaggle.com/aharless/talkingdata-time-deltas",
      "votes": null
    },
    {
      "id": "310969",
      "postDate": "04/09/2018 05:53:26",
      "content": "<p>I added a hashing Python version that works in linear time and memory to my kernel (<a href=\"https://www.kaggle.com/anttip/talkingdata-wordbatch-fm-ftrl-lb-0-9711\">https://www.kaggle.com/anttip/talkingdata-wordbatch-fm-ftrl-lb-0-9711</a>). It does around a million rows per 4 seconds per thread in Kaggle kernel. LB score jumped from 0.9682 to 0.9711.</p>\n\n<p>This is a really powerful and almost a leak feature, since it uses future information for each click. It should still be useful to the organizer, if they do prediction per batches, or with some latency.</p>",
      "rawMarkdown": "I added a hashing Python version that works in linear time and memory to my kernel (https://www.kaggle.com/anttip/talkingdata-wordbatch-fm-ftrl-lb-0-9711). It does around a million rows per 4 seconds per thread in Kaggle kernel. LB score jumped from 0.9682 to 0.9711.\n\nThis is a really powerful and almost a leak feature, since it uses future information for each click. It should still be useful to the organizer, if they do prediction per batches, or with some latency.",
      "votes": null
    },
    {
      "id": "311158",
      "postDate": "04/09/2018 14:08:46",
      "content": "<p>@anttip So you're doing forward time deltas (as in <a href=\"https://www.kaggle.com/nanomathias/feature-engineering-importance-testing\">NanoMathias' kernel</a>), which are probably more powerful than backward time deltas (what I assumed we've been mostly talking about).  I suppose one could use both.  Does your code deal with batch boundaries, or just treat them like the end of the file?</p>",
      "rawMarkdown": "anttip So you're doing forward time deltas (as in [NanoMathias' kernel][1]), which are probably more powerful than backward time deltas (what I assumed we've been mostly talking about).  I suppose one could use both.  Does your code deal with batch boundaries, or just treat them like the end of the file?\n\n [1]: https://www.kaggle.com/nanomathias/feature-engineering-importance-testing",
      "votes": null
    },
    {
      "id": "311174",
      "postDate": "04/09/2018 14:44:22",
      "content": "<p>Out of principle I tried to avoid doing this. However, your reasons FOR it being ok make a lot of sense to me. Thanks anttip for the justification :)</p>",
      "rawMarkdown": "Out of principle I tried to avoid doing this. However, your reasons FOR it being ok make a lot of sense to me. Thanks anttip for the justification :)",
      "votes": null
    },
    {
      "id": "311382",
      "postDate": "04/10/2018 00:21:27",
      "content": "<p>Hi @Andy, could I ask is 'forward time deltas' = time gap between current click and next click and 'backward time deltas' = time gap between current click and previous click?</p>",
      "rawMarkdown": "Hi @Andy, could I ask is 'forward time deltas' = time gap between current click and next click and 'backward time deltas' = time gap between current click and previous click?",
      "votes": null
    },
    {
      "id": "311513",
      "postDate": "04/10/2018 07:23:28",
      "content": "<p>@andy those are approximated as end of file. However, the features are log-binned, and batch size is 10M, so the proportion of cases that are affected should be tiny</p>",
      "rawMarkdown": "andy those are approximated as end of file. However, the features are log-binned, and batch size is 10M, so the proportion of cases that are affected should be tiny",
      "votes": null
    },
    {
      "id": "311981",
      "postDate": "04/11/2018 02:39:50",
      "content": "<p>@andy In your kernel, you do the sorting first then calculate time delta. But in anttip's monster kernel, he didn't do it. Have you tested on both? PS: my sorted time delta val score is lower than un sorted one. This really makes me confused.</p>",
      "rawMarkdown": "andy In your kernel, you do the sorting first then calculate time delta. But in anttip's monster kernel, he didn't do it. Have you tested on both? PS: my sorted time delta val score is lower than un sorted one. This really makes me confused.",
      "votes": null
    },
    {
      "id": "312268",
      "postDate": "04/11/2018 13:42:37",
      "content": "<p>I haven't tested sorting vs. not sorting. The data should be in order by time already, so sorting may not be necessary: I just did so to be safe, because I'm not sure whether Pandas promises to keep the original time sequence through my other operations. OTOH, maybe it is better <em>not</em> to sort by time, because sorting by time may lose information about the order of duplicate records (i.e. rows that are identical in <code>click_time</code> and all other raw features but are sometimes different in <code>is_attributed</code> and/or <code>attributed_time</code>). </p>",
      "rawMarkdown": "I haven't tested sorting vs. not sorting. The data should be in order by time already, so sorting may not be necessary: I just did so to be safe, because I'm not sure whether Pandas promises to keep the original time sequence through my other operations. OTOH, maybe it is better *not* to sort by time, because sorting by time may lose information about the order of duplicate records (i.e. rows that are identical in `click_time` and all other raw features but are sometimes different in `is_attributed` and/or `attributed_time`).",
      "votes": null
    },
    {
      "id": "312364",
      "postDate": "04/11/2018 16:23:00",
      "content": "<p>I'm going to change it and do my sort by index instead of <code>click_time</code>.  I'm not sure whether sorting by index is necessary. (Will Pandas automatically keep the index order within categories?)  But anyhow I am now thinking this is the correct way to sort. If two records are identical (except for possibly <code>is_attributed</code> and/or <code>attributed_time</code>), then I think the best assumptions are (1) that they are two separate clicks and (2) that the one that appears first in the file happened first (less than a second earlier, hence having the same apparent timestamp).</p>\n\n<p>@KALE Under this assumption it makes sense that your sorted version performed worse, since it may have contained less accurate information, sometimes reversing the order of duplicates.</p>",
      "rawMarkdown": "I'm going to change it and do my sort by index instead of `click_time`.  I'm not sure whether sorting by index is necessary. (Will Pandas automatically keep the index order within categories?)  But anyhow I am now thinking this is the correct way to sort. If two records are identical (except for possibly `is_attributed` and/or `attributed_time`), then I think the best assumptions are (1) that they are two separate clicks and (2) that the one that appears first in the file happened first (less than a second earlier, hence having the same apparent timestamp).\n\n@KALE Under this assumption it makes sense that your sorted version performed worse, since it may have contained less accurate information, sometimes reversing the order of duplicates.",
      "votes": null
    },
    {
      "id": "312376",
      "postDate": "04/11/2018 16:50:42",
      "content": "<p>I think it's safer to not sort, since the data is already in time order. Other aggregations should not mess up the ordering, especially if you're using merges to add features to the collective dataframe. Also, pandas sorting is not stable by default (quicksort) and will not preserve the original order of duplicates. </p>\n\n<p>Think of it this way - for duplicate timestamps, there's no way to distinguish order except for the original row order in the dataframe. I wouldn't want to touch that, pandas will almost certainly mess it up unless you're very careful, and even then I think you're doing unnecessary work.</p>",
      "rawMarkdown": "I think it's safer to not sort, since the data is already in time order. Other aggregations should not mess up the ordering, especially if you're using merges to add features to the collective dataframe. Also, pandas sorting is not stable by default (quicksort) and will not preserve the original order of duplicates. \n\nThink of it this way - for duplicate timestamps, there's no way to distinguish order except for the original row order in the dataframe. I wouldn't want to touch that, pandas will almost certainly mess it up unless you're very careful, and even then I think you're doing unnecessary work.",
      "votes": null
    },
    {
      "id": "312430",
      "postDate": "04/11/2018 18:15:37",
      "content": "<p>I've revised mine now. In the latest version, I sort by index instead of time. (Part of my memory-saving kludge involves sorting by category, so I actually sort by category*(2**32)+index, so as to keep records in the right order within categories.) I'm allowing that Pandas might mess up the order within categories if \nI sorted by category alone.</p>",
      "rawMarkdown": "I've revised mine now. In the latest version, I sort by index instead of time. (Part of my memory-saving kludge involves sorting by category, so I actually sort by category*(2**32)+index, so as to keep records in the right order within categories.) I'm allowing that Pandas might mess up the order within categories if \nI sorted by category alone.",
      "votes": null
    },
    {
      "id": "316557",
      "postDate": "04/19/2018 10:28:58",
      "content": "<p><a href=\"https://github.com/SudalaiRajkumar/ML/blob/master/AV_LordOfTheMachines/Explorations.ipynb\">https://github.com/SudalaiRajkumar/ML/blob/master/AV_LordOfTheMachines/Explorations.ipynb</a>\ncan we use this to create forward time deltas?</p>",
      "rawMarkdown": "https://github.com/SudalaiRajkumar/ML/blob/master/AV_LordOfTheMachines/Explorations.ipynb\ncan we use this to create forward time deltas?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 306037,
      "author_name": "aharless",
      "author_url": "",
      "post_date": "03/29/2018 19:19:56",
      "content": "<p>What do you mean by \"time deltas\"? The time between successive clicks from a given group?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 306060,
      "author_name": "scirpus",
      "author_url": "",
      "post_date": "03/29/2018 20:09:06",
      "content": "<p>Indeed</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 306129,
      "author_name": "wythhh",
      "author_url": "",
      "post_date": "03/29/2018 22:28:25",
      "content": "<p>Hi, Scirpus, could we have some explanation of this figure? like what's the meaning of x and y axis and what does color blue and yellow present? Thank you very much!</p>",
      "votes": null,
      "replies": [
        {
          "id": 306273,
          "author_name": "scirpus",
          "author_url": "",
          "post_date": "03/30/2018 06:52:04",
          "content": "<p>\"looks like quite a different distribution between zeros and ones\" tells you the colors of the plots represent the is_attributed  zeros and ones</p>\n\n<p>x axis is \"time deltas grouped by ip,app,device,os,channel on the data\" and the y axis is number of incidences per x value</p>\n\n<p>Hope this helps</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 306369,
          "author_name": "aharless",
          "author_url": "",
          "post_date": "03/30/2018 10:07:55",
          "content": "<p>So this chart is grouped by all 5 categories at once?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 306371,
          "author_name": "scirpus",
          "author_url": "",
          "post_date": "03/30/2018 10:13:39",
          "content": "<p>That is true - I am just trying it now to see if it improves things.</p>\n\n<p>I wrote a little c++ app to do it as pandas was <strong><em>really really slow</em></strong> on all the data (time-ordered, concatenated on train,test and test supplement) do you want a copy?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 306625,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "03/30/2018 19:01:58",
          "content": "<p>test is included in test_supplement, therefore you should concatenate train and test_supplement only.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 306635,
          "author_name": "scirpus",
          "author_url": "",
          "post_date": "03/30/2018 19:16:22",
          "content": "<p>Thanks for this I completely missed this owing to the fact that I read supplement and assumed kaggle was using the word supplement to actually mean something.  Why on earth didn't they can it \"originaltestset\" goddamn it! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 306642,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "03/30/2018 19:25:23",
          "content": "<p>No pb.  Sometimes a fresh eye is good ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 306648,
          "author_name": "scirpus",
          "author_url": "",
          "post_date": "03/30/2018 19:35:05",
          "content": "<p>Thanks for stopping me going down a rabbit hole!  I have just published the code in the discussion.  It is pretty generic and may come in handy for you in future competitions!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 306624,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "03/30/2018 19:00:43",
      "content": "<p>How do you define it when there is a single observation in the group?</p>",
      "votes": null,
      "replies": [
        {
          "id": 306630,
          "author_name": "scirpus",
          "author_url": "",
          "post_date": "03/30/2018 19:10:22",
          "content": "<p>I chose NAN</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 306661,
      "author_name": "sohaibomar",
      "author_url": "",
      "post_date": "03/30/2018 19:52:33",
      "content": "<p>Is time deltas difference in time of successive clicks of each group?</p>",
      "votes": null,
      "replies": [
        {
          "id": 306712,
          "author_name": "aharless",
          "author_url": "",
          "post_date": "03/30/2018 21:20:52",
          "content": "<p>Yes. (See the top of this thread.)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 306774,
      "author_name": "baomengjiao",
      "author_url": "",
      "post_date": "03/31/2018 01:35:30",
      "content": "<p>Time Deltas on what? Any explanation thanks a lot </p>",
      "votes": null,
      "replies": [
        {
          "id": 306883,
          "author_name": "scirpus",
          "author_url": "",
          "post_date": "03/31/2018 08:35:13",
          "content": "<p>Look at the code in the <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53450\">link</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 308242,
      "author_name": "samehif",
      "author_url": "",
      "post_date": "04/03/2018 07:36:11",
      "content": "<p>this also looked like a good feature to me. but when i tried it didnt help the score. it made it even worse. i used different grouping like ip alone and ip, device, os and so on.</p>",
      "votes": null,
      "replies": [
        {
          "id": 308393,
          "author_name": "scirpus",
          "author_url": "",
          "post_date": "04/03/2018 13:03:00",
          "content": "<p>Did you time order the data ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 308425,
          "author_name": "samehif",
          "author_url": "",
          "post_date": "04/03/2018 13:49:34",
          "content": "<p>Yes I did. For example sorting by ip and click time and then doing a diff row wise. Also tried sort by [ip, device,os, click_time] . </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 308430,
          "author_name": "scirpus",
          "author_url": "",
          "post_date": "04/03/2018 13:58:10",
          "content": "<p>And you used test supplement rather than test?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 308444,
          "author_name": "samehif",
          "author_url": "",
          "post_date": "04/03/2018 14:25:40",
          "content": "<p>I only tried with cv I think. I have to recheck maybe. Did it help your score a lot?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 308453,
          "author_name": "scirpus",
          "author_url": "",
          "post_date": "04/03/2018 14:35:51",
          "content": "<p>It helped a little (.002) moving me up to the top 50 </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 309738,
      "author_name": "yuliagm",
      "author_url": "",
      "post_date": "04/05/2018 22:07:49",
      "content": "<p>Jumping in on the thread cause I can't find answer anywhere else and it seems relevant here, as I understand you are concatenating old test and train... </p>\n\n<p>What do you think it means that the old test and train overlap?  There is about an hour and a half of data that covers both train and test.  Does that mean that that hour of train is missing data or did they duplicate it in test (didn't look that way on cursory exploration)?    That hour actually corresponds to the time on the real test, and if it's really missing, it can make those hours misleading in a model (i.e. the counts in those times were actually higher than what we are training on)...  Any thoughts on what happened there or the effect?</p>\n\n<p>Specifically  train ends on 2017:11:09 at 16:00:00, while old test starts on 2017-11-06 14:32:21.  So the time between 14:32:21 and 16:00:00 on the 9th some clicks are in train and some on test.  </p>\n\n<p>PS:  I was working on similar time_delta feature as well and one thing to keep in mind is that there are many duplicate clicks, which means that if you just sort them by click_time they may not stay in their original order, and thus a time_delta of a set of 2-3 clicks that just launched could have wrong click assigned 0 and the longer delta...  </p>",
      "votes": null,
      "replies": [
        {
          "id": 309743,
          "author_name": "aquatic",
          "author_url": "",
          "post_date": "04/05/2018 22:30:38",
          "content": "<p>Re: your last point - nice observation, I think we can just leave the clicks in the original order for time delta computations because they're already time sorted?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 309750,
          "author_name": "yuliagm",
          "author_url": "",
          "post_date": "04/05/2018 22:44:54",
          "content": "<p>It's just a 'be careful' kinda note because if people sort/resort/concatenate data for various groupings, the original flow can get lost.  Something like that happened to me in one of my tries and took me some time to realize it was happening.  </p>\n\n<p>I'm doing this in pandas and it's taking forever, so these kind of errors have literally cost me days to untangle.  Currently in the process of fixing something along those lines and really wishing had better coding skills/faster computer...  </p>\n\n<p>Also the issue is the blending of train/test times that overlap.  a) how to understand it b) if accept it and get clicks that are duplicates in both, which came before/after?  splicing the two together really affects the time range after 14:00 which can hurt the model, but won't be noticed on the public leader board.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 309942,
          "author_name": "samehif",
          "author_url": "",
          "post_date": "04/06/2018 08:33:46",
          "content": "<p>the amount of transactions in test test in the 1.5 hour of overlap is too little (&lt; 1000) and i think could be just simply ignored.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 310454,
      "author_name": "aharless",
      "author_url": "",
      "post_date": "04/07/2018 15:33:10",
      "content": "<p>I've done the calculation <a href=\"https://www.kaggle.com/aharless/talkingdata-time-deltas\">in Pandas</a> (for just the training set in this case).  It runs in reasonable time (a significant part of the kernel time is just writing to disk), but I had to jump through a lot of hoops to avoid breaching the Kaggle memory constraint, and the disk space constraint didn't permit me to write out the whole training file with deltas (although one could write just the deltas and do the combination later when reading).</p>",
      "votes": null,
      "replies": [
        {
          "id": 310969,
          "author_name": "anttip",
          "author_url": "",
          "post_date": "04/09/2018 05:53:26",
          "content": "<p>I added a hashing Python version that works in linear time and memory to my kernel (<a href=\"https://www.kaggle.com/anttip/talkingdata-wordbatch-fm-ftrl-lb-0-9711\">https://www.kaggle.com/anttip/talkingdata-wordbatch-fm-ftrl-lb-0-9711</a>). It does around a million rows per 4 seconds per thread in Kaggle kernel. LB score jumped from 0.9682 to 0.9711.</p>\n\n<p>This is a really powerful and almost a leak feature, since it uses future information for each click. It should still be useful to the organizer, if they do prediction per batches, or with some latency.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 311158,
          "author_name": "aharless",
          "author_url": "",
          "post_date": "04/09/2018 14:08:46",
          "content": "<p>@anttip So you're doing forward time deltas (as in <a href=\"https://www.kaggle.com/nanomathias/feature-engineering-importance-testing\">NanoMathias' kernel</a>), which are probably more powerful than backward time deltas (what I assumed we've been mostly talking about).  I suppose one could use both.  Does your code deal with batch boundaries, or just treat them like the end of the file?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 311174,
          "author_name": "datadugong",
          "author_url": "",
          "post_date": "04/09/2018 14:44:22",
          "content": "<p>Out of principle I tried to avoid doing this. However, your reasons FOR it being ok make a lot of sense to me. Thanks anttip for the justification :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 311382,
          "author_name": "wythhh",
          "author_url": "",
          "post_date": "04/10/2018 00:21:27",
          "content": "<p>Hi @Andy, could I ask is 'forward time deltas' = time gap between current click and next click and 'backward time deltas' = time gap between current click and previous click?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 311513,
          "author_name": "anttip",
          "author_url": "",
          "post_date": "04/10/2018 07:23:28",
          "content": "<p>@andy those are approximated as end of file. However, the features are log-binned, and batch size is 10M, so the proportion of cases that are affected should be tiny</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 311981,
          "author_name": "cczaixian",
          "author_url": "",
          "post_date": "04/11/2018 02:39:50",
          "content": "<p>@andy In your kernel, you do the sorting first then calculate time delta. But in anttip's monster kernel, he didn't do it. Have you tested on both? PS: my sorted time delta val score is lower than un sorted one. This really makes me confused.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 312268,
          "author_name": "aharless",
          "author_url": "",
          "post_date": "04/11/2018 13:42:37",
          "content": "<p>I haven't tested sorting vs. not sorting. The data should be in order by time already, so sorting may not be necessary: I just did so to be safe, because I'm not sure whether Pandas promises to keep the original time sequence through my other operations. OTOH, maybe it is better <em>not</em> to sort by time, because sorting by time may lose information about the order of duplicate records (i.e. rows that are identical in <code>click_time</code> and all other raw features but are sometimes different in <code>is_attributed</code> and/or <code>attributed_time</code>). </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 312364,
          "author_name": "aharless",
          "author_url": "",
          "post_date": "04/11/2018 16:23:00",
          "content": "<p>I'm going to change it and do my sort by index instead of <code>click_time</code>.  I'm not sure whether sorting by index is necessary. (Will Pandas automatically keep the index order within categories?)  But anyhow I am now thinking this is the correct way to sort. If two records are identical (except for possibly <code>is_attributed</code> and/or <code>attributed_time</code>), then I think the best assumptions are (1) that they are two separate clicks and (2) that the one that appears first in the file happened first (less than a second earlier, hence having the same apparent timestamp).</p>\n\n<p>@KALE Under this assumption it makes sense that your sorted version performed worse, since it may have contained less accurate information, sometimes reversing the order of duplicates.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 312376,
          "author_name": "aquatic",
          "author_url": "",
          "post_date": "04/11/2018 16:50:42",
          "content": "<p>I think it's safer to not sort, since the data is already in time order. Other aggregations should not mess up the ordering, especially if you're using merges to add features to the collective dataframe. Also, pandas sorting is not stable by default (quicksort) and will not preserve the original order of duplicates. </p>\n\n<p>Think of it this way - for duplicate timestamps, there's no way to distinguish order except for the original row order in the dataframe. I wouldn't want to touch that, pandas will almost certainly mess it up unless you're very careful, and even then I think you're doing unnecessary work.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 312430,
          "author_name": "aharless",
          "author_url": "",
          "post_date": "04/11/2018 18:15:37",
          "content": "<p>I've revised mine now. In the latest version, I sort by index instead of time. (Part of my memory-saving kludge involves sorting by category, so I actually sort by category*(2**32)+index, so as to keep records in the right order within categories.) I'm allowing that Pandas might mess up the order within categories if \nI sorted by category alone.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 316557,
      "author_name": "nick7hill",
      "author_url": "",
      "post_date": "04/19/2018 10:28:58",
      "content": "<p><a href=\"https://github.com/SudalaiRajkumar/ML/blob/master/AV_LordOfTheMachines/Explorations.ipynb\">https://github.com/SudalaiRajkumar/ML/blob/master/AV_LordOfTheMachines/Explorations.ipynb</a>\ncan we use this to create forward time deltas?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "305960": "I looked at time deltas grouped by ip,app,device,os,channel on the data - looks like quite a different distribution between zeros and ones",
    "306037": "What do you mean by \"time deltas\"? The time between successive clicks from a given group?",
    "306060": "Indeed",
    "306129": "Hi, Scirpus, could we have some explanation of this figure? like what's the meaning of x and y axis and what does color blue and yellow present? Thank you very much!",
    "306273": "\"looks like quite a different distribution between zeros and ones\" tells you the colors of the plots represent the is_attributed  zeros and ones\n\nx axis is \"time deltas grouped by ip,app,device,os,channel on the data\" and the y axis is number of incidences per x value\n\nHope this helps",
    "306369": "So this chart is grouped by all 5 categories at once?",
    "306371": "That is true - I am just trying it now to see if it improves things.\n\nI wrote a little c++ app to do it as pandas was ***really really slow*** on all the data (time-ordered, concatenated on train,test and test supplement) do you want a copy?",
    "306624": "How do you define it when there is a single observation in the group?",
    "306625": "test is included in test_supplement, therefore you should concatenate train and test_supplement only.",
    "306630": "I chose NAN",
    "306635": "Thanks for this I completely missed this owing to the fact that I read supplement and assumed kaggle was using the word supplement to actually mean something.  Why on earth didn't they can it \"originaltestset\" goddamn it!",
    "306642": "No pb.  Sometimes a fresh eye is good ;)",
    "306648": "Thanks for stopping me going down a rabbit hole!  I have just published the code in the discussion.  It is pretty generic and may come in handy for you in future competitions!",
    "306661": "Is time deltas difference in time of successive clicks of each group?",
    "306712": "Yes. (See the top of this thread.)",
    "306774": "Time Deltas on what? Any explanation thanks a lot",
    "306883": "Look at the code in the [link][1]\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53450",
    "308242": "this also looked like a good feature to me. but when i tried it didnt help the score. it made it even worse. i used different grouping like ip alone and ip, device, os and so on.",
    "308393": "Did you time order the data ?",
    "308425": "Yes I did. For example sorting by ip and click time and then doing a diff row wise. Also tried sort by [ip, device,os, click_time] .",
    "308430": "And you used test supplement rather than test?",
    "308444": "I only tried with cv I think. I have to recheck maybe. Did it help your score a lot?",
    "308453": "It helped a little (.002) moving me up to the top 50",
    "309738": "Jumping in on the thread cause I can't find answer anywhere else and it seems relevant here, as I understand you are concatenating old test and train... \n\nWhat do you think it means that the old test and train overlap?  There is about an hour and a half of data that covers both train and test.  Does that mean that that hour of train is missing data or did they duplicate it in test (didn't look that way on cursory exploration)?    That hour actually corresponds to the time on the real test, and if it's really missing, it can make those hours misleading in a model (i.e. the counts in those times were actually higher than what we are training on)...  Any thoughts on what happened there or the effect?\n\nSpecifically  train ends on 2017:11:09 at 16:00:00, while old test starts on 2017-11-06 14:32:21.  So the time between 14:32:21 and 16:00:00 on the 9th some clicks are in train and some on test.  \n\nPS:  I was working on similar time_delta feature as well and one thing to keep in mind is that there are many duplicate clicks, which means that if you just sort them by click_time they may not stay in their original order, and thus a time_delta of a set of 2-3 clicks that just launched could have wrong click assigned 0 and the longer delta...",
    "309743": "Re: your last point - nice observation, I think we can just leave the clicks in the original order for time delta computations because they're already time sorted?",
    "309750": "It's just a 'be careful' kinda note because if people sort/resort/concatenate data for various groupings, the original flow can get lost.  Something like that happened to me in one of my tries and took me some time to realize it was happening.  \n\nI'm doing this in pandas and it's taking forever, so these kind of errors have literally cost me days to untangle.  Currently in the process of fixing something along those lines and really wishing had better coding skills/faster computer...  \n\nAlso the issue is the blending of train/test times that overlap.  a) how to understand it b) if accept it and get clicks that are duplicates in both, which came before/after?  splicing the two together really affects the time range after 14:00 which can hurt the model, but won't be noticed on the public leader board.",
    "309942": "the amount of transactions in test test in the 1.5 hour of overlap is too little (&lt; 1000) and i think could be just simply ignored.",
    "310454": "I've done the calculation [in Pandas][1] (for just the training set in this case).  It runs in reasonable time (a significant part of the kernel time is just writing to disk), but I had to jump through a lot of hoops to avoid breaching the Kaggle memory constraint, and the disk space constraint didn't permit me to write out the whole training file with deltas (although one could write just the deltas and do the combination later when reading).\n\n [1]: https://www.kaggle.com/aharless/talkingdata-time-deltas",
    "310969": "I added a hashing Python version that works in linear time and memory to my kernel (https://www.kaggle.com/anttip/talkingdata-wordbatch-fm-ftrl-lb-0-9711). It does around a million rows per 4 seconds per thread in Kaggle kernel. LB score jumped from 0.9682 to 0.9711.\n\nThis is a really powerful and almost a leak feature, since it uses future information for each click. It should still be useful to the organizer, if they do prediction per batches, or with some latency.",
    "311158": "anttip So you're doing forward time deltas (as in [NanoMathias' kernel][1]), which are probably more powerful than backward time deltas (what I assumed we've been mostly talking about).  I suppose one could use both.  Does your code deal with batch boundaries, or just treat them like the end of the file?\n\n [1]: https://www.kaggle.com/nanomathias/feature-engineering-importance-testing",
    "311174": "Out of principle I tried to avoid doing this. However, your reasons FOR it being ok make a lot of sense to me. Thanks anttip for the justification :)",
    "311382": "Hi @Andy, could I ask is 'forward time deltas' = time gap between current click and next click and 'backward time deltas' = time gap between current click and previous click?",
    "311513": "andy those are approximated as end of file. However, the features are log-binned, and batch size is 10M, so the proportion of cases that are affected should be tiny",
    "311981": "andy In your kernel, you do the sorting first then calculate time delta. But in anttip's monster kernel, he didn't do it. Have you tested on both? PS: my sorted time delta val score is lower than un sorted one. This really makes me confused.",
    "312268": "I haven't tested sorting vs. not sorting. The data should be in order by time already, so sorting may not be necessary: I just did so to be safe, because I'm not sure whether Pandas promises to keep the original time sequence through my other operations. OTOH, maybe it is better *not* to sort by time, because sorting by time may lose information about the order of duplicate records (i.e. rows that are identical in `click_time` and all other raw features but are sometimes different in `is_attributed` and/or `attributed_time`).",
    "312364": "I'm going to change it and do my sort by index instead of `click_time`.  I'm not sure whether sorting by index is necessary. (Will Pandas automatically keep the index order within categories?)  But anyhow I am now thinking this is the correct way to sort. If two records are identical (except for possibly `is_attributed` and/or `attributed_time`), then I think the best assumptions are (1) that they are two separate clicks and (2) that the one that appears first in the file happened first (less than a second earlier, hence having the same apparent timestamp).\n\n@KALE Under this assumption it makes sense that your sorted version performed worse, since it may have contained less accurate information, sometimes reversing the order of duplicates.",
    "312376": "I think it's safer to not sort, since the data is already in time order. Other aggregations should not mess up the ordering, especially if you're using merges to add features to the collective dataframe. Also, pandas sorting is not stable by default (quicksort) and will not preserve the original order of duplicates. \n\nThink of it this way - for duplicate timestamps, there's no way to distinguish order except for the original row order in the dataframe. I wouldn't want to touch that, pandas will almost certainly mess it up unless you're very careful, and even then I think you're doing unnecessary work.",
    "312430": "I've revised mine now. In the latest version, I sort by index instead of time. (Part of my memory-saving kludge involves sorting by category, so I actually sort by category*(2**32)+index, so as to keep records in the right order within categories.) I'm allowing that Pandas might mess up the order within categories if \nI sorted by category alone.",
    "316557": "https://github.com/SudalaiRajkumar/ML/blob/master/AV_LordOfTheMachines/Explorations.ipynb\ncan we use this to create forward time deltas?"
  },
  "source": "meta"
}