{
  "id": 56053,
  "title": "Next_click",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/56053",
  "author_name": "",
  "post_date": "2018-05-05T03:45:30.743471600Z",
  "votes": null,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I referred  <a href=\"https://www.kaggle.com/bk0000/non-blending-lightgbm-model-lb-0-977\">kernel</a> . Kindly help me to understand   </p>\n\n<ul>\n<li>Why should we need to use \"3000000000\" here? </li>\n<li>Why should we divide cilck_time with 10 ** 9?  </li>\n</ul>\n\n<blockquote>\n<pre><code>D = 2 ** 26\ntrain_df['category'] = (train_df['ip'].astype(str) + \"_\" + train_df['app'].astype(str) + \"_\" + \n                 train_df['device'].astype( str) + \"_\" + train_df['os'].astype(str)).apply(hash) % D\nclick_buffer = np.full(D, 3000000000, dtype=np.uint32)\ntrain_df['epochtime'] = train_df['click_time'].astype(np.int64) // 10 ** 9\nnext_clicks = []\nfor category, t in zip(reversed(train_df['category'].values), reversed(train_df['epochtime'].values)):\n            next_clicks.append(click_buffer[category] - t)\n            click_buffer[category] = t\n</code></pre>\n</blockquote>",
  "messages": [
    {
      "id": "323406",
      "postDate": "05/05/2018 03:45:30",
      "content": "<p>I referred  <a href=\"https://www.kaggle.com/bk0000/non-blending-lightgbm-model-lb-0-977\">kernel</a> . Kindly help me to understand   </p>\n\n<ul>\n<li>Why should we need to use \"3000000000\" here? </li>\n<li>Why should we divide cilck_time with 10 ** 9?  </li>\n</ul>\n\n<blockquote>\n<pre><code>D = 2 ** 26\ntrain_df['category'] = (train_df['ip'].astype(str) + \"_\" + train_df['app'].astype(str) + \"_\" + \n                 train_df['device'].astype( str) + \"_\" + train_df['os'].astype(str)).apply(hash) % D\nclick_buffer = np.full(D, 3000000000, dtype=np.uint32)\ntrain_df['epochtime'] = train_df['click_time'].astype(np.int64) // 10 ** 9\nnext_clicks = []\nfor category, t in zip(reversed(train_df['category'].values), reversed(train_df['epochtime'].values)):\n            next_clicks.append(click_buffer[category] - t)\n            click_buffer[category] = t\n</code></pre>\n</blockquote>",
      "rawMarkdown": "I referred  [kernel][1] . Kindly help me to understand   \n\n - Why should we need to use \"3000000000\" here? \n - Why should we divide cilck_time with 10 ** 9?  \n\n&gt;     D = 2 ** 26\n&gt;     train_df['category'] = (train_df['ip'].astype(str) + \"_\" + train_df['app'].astype(str) + \"_\" + \n&gt;                      train_df['device'].astype( str) + \"_\" + train_df['os'].astype(str)).apply(hash) % D\n&gt;     click_buffer = np.full(D, 3000000000, dtype=np.uint32)\n&gt;     train_df['epochtime'] = train_df['click_time'].astype(np.int64) // 10 ** 9\n&gt;     next_clicks = []\n&gt;     for category, t in zip(reversed(train_df['category'].values), reversed(train_df['epochtime'].values)):\n&gt;                 next_clicks.append(click_buffer[category] - t)\n&gt;                 click_buffer[category] = t\n\n  \n\n\n  [1]: https://www.kaggle.com/bk0000/non-blending-lightgbm-model-lb-0-977",
      "votes": null
    },
    {
      "id": "323420",
      "postDate": "05/05/2018 05:23:46",
      "content": "<p>Use <a href=\"https://www.kaggle.com/asydorchuk\">this kernel</a> by Andrii Sydorchuk for FE'ing nextclick instead. It's a lot faster at building this feature and you don't have to worry about the remotely minute potential for hash conflicts. But to answer your Qs, the 300xxx number is just a large placeholder (think of it has fillna placeholder that fits in uint32 but doesn't occur in our samples). And the 10 ** 9 is to get you down to seconds, which is the highest resolution we're provided in the dataset anyway.</p>",
      "rawMarkdown": "Use [this kernel][1] by Andrii Sydorchuk for FE'ing nextclick instead. It's a lot faster at building this feature and you don't have to worry about the remotely minute potential for hash conflicts. But to answer your Qs, the 300xxx number is just a large placeholder (think of it has fillna placeholder that fits in uint32 but doesn't occur in our samples). And the 10 ** 9 is to get you down to seconds, which is the highest resolution we're provided in the dataset anyway.\n\n\n  [1]: https://www.kaggle.com/asydorchuk",
      "votes": null
    },
    {
      "id": "323424",
      "postDate": "05/05/2018 05:49:39",
      "content": "<p>thank you</p>",
      "rawMarkdown": "thank you",
      "votes": null
    },
    {
      "id": "323436",
      "postDate": "05/05/2018 06:35:51",
      "content": "<p>little curious. Kindly help me.</p>\n\n<ul>\n<li>Why should we do reverse? <br>\n   <code>python\n      zip(reversed(train_df['category'].values), reversed(train_df['epochtime'].values))\n</code>  </li>\n<li>Why should we minus click_time from click_buffer placeholder value(3000xxx)? <br>\n <code>python\n  next_clicks.append(click_buffer[category] - t)\n</code></li>\n</ul>",
      "rawMarkdown": "little curious. Kindly help me.\n\n - Why should we do reverse?  \n       ```python\n          zip(reversed(train_df['category'].values), reversed(train_df['epochtime'].values))\n      ```  \n - Why should we minus click_time from click_buffer placeholder value(3000xxx)?  \n     ```python\n      next_clicks.append(click_buffer[category] - t)\n   ```",
      "votes": null
    },
    {
      "id": "323665",
      "postDate": "05/05/2018 20:29:45",
      "content": "<p>@ Pallavi </p>\n\n<p>I'l answer your second question first and that leads into the first one </p>\n\n<p>This click time can be calculated without the hashing trick as you may know but will create a NAN value for the last record in every group since there is no record you can compare it against.  A default value is therefore used which is the 3 billion from which the click_time of the last record is subtracted </p>\n\n<p>Why Reversed - Records are in chronological order of click time. So you get the last click of every category and calculate the time delta using the 3B value and then go on to the clicks earlier in time. </p>\n\n<p>Hope this helps \nRegards\nShanth</p>",
      "rawMarkdown": "Pallavi \n\nI'l answer your second question first and that leads into the first one \n\nThis click time can be calculated without the hashing trick as you may know but will create a NAN value for the last record in every group since there is no record you can compare it against.  A default value is therefore used which is the 3 billion from which the click_time of the last record is subtracted \n\nWhy Reversed - Records are in chronological order of click time. So you get the last click of every category and calculate the time delta using the 3B value and then go on to the clicks earlier in time. \n\nHope this helps \nRegards\nShanth",
      "votes": null
    },
    {
      "id": "323727",
      "postDate": "05/06/2018 02:26:14",
      "content": "<p>Thanks for response. Its useful.\nI agree with your answer for 2nd question. But, Why only 3B ? Why not any other number?</p>",
      "rawMarkdown": "Thanks for response. Its useful.\nI agree with your answer for 2nd question. But, Why only 3B ? Why not any other number?",
      "votes": null
    },
    {
      "id": "323758",
      "postDate": "05/06/2018 05:27:28",
      "content": "<p>When you notice the value generated for the column epoch time, its a very large INT. Choosing 3 Billion could have to do with the max value of the epoch time field. </p>",
      "rawMarkdown": "When you notice the value generated for the column epoch time, its a very large INT. Choosing 3 Billion could have to do with the max value of the epoch time field.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 323420,
      "author_name": "authman",
      "author_url": "",
      "post_date": "05/05/2018 05:23:46",
      "content": "<p>Use <a href=\"https://www.kaggle.com/asydorchuk\">this kernel</a> by Andrii Sydorchuk for FE'ing nextclick instead. It's a lot faster at building this feature and you don't have to worry about the remotely minute potential for hash conflicts. But to answer your Qs, the 300xxx number is just a large placeholder (think of it has fillna placeholder that fits in uint32 but doesn't occur in our samples). And the 10 ** 9 is to get you down to seconds, which is the highest resolution we're provided in the dataset anyway.</p>",
      "votes": null,
      "replies": [
        {
          "id": 323424,
          "author_name": "pallaviroyal",
          "author_url": "",
          "post_date": "05/05/2018 05:49:39",
          "content": "<p>thank you</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 323436,
      "author_name": "pallaviroyal",
      "author_url": "",
      "post_date": "05/05/2018 06:35:51",
      "content": "<p>little curious. Kindly help me.</p>\n\n<ul>\n<li>Why should we do reverse? <br>\n   <code>python\n      zip(reversed(train_df['category'].values), reversed(train_df['epochtime'].values))\n</code>  </li>\n<li>Why should we minus click_time from click_buffer placeholder value(3000xxx)? <br>\n <code>python\n  next_clicks.append(click_buffer[category] - t)\n</code></li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 323665,
          "author_name": "shanth84",
          "author_url": "",
          "post_date": "05/05/2018 20:29:45",
          "content": "<p>@ Pallavi </p>\n\n<p>I'l answer your second question first and that leads into the first one </p>\n\n<p>This click time can be calculated without the hashing trick as you may know but will create a NAN value for the last record in every group since there is no record you can compare it against.  A default value is therefore used which is the 3 billion from which the click_time of the last record is subtracted </p>\n\n<p>Why Reversed - Records are in chronological order of click time. So you get the last click of every category and calculate the time delta using the 3B value and then go on to the clicks earlier in time. </p>\n\n<p>Hope this helps \nRegards\nShanth</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 323727,
          "author_name": "pallaviroyal",
          "author_url": "",
          "post_date": "05/06/2018 02:26:14",
          "content": "<p>Thanks for response. Its useful.\nI agree with your answer for 2nd question. But, Why only 3B ? Why not any other number?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 323758,
          "author_name": "shanth84",
          "author_url": "",
          "post_date": "05/06/2018 05:27:28",
          "content": "<p>When you notice the value generated for the column epoch time, its a very large INT. Choosing 3 Billion could have to do with the max value of the epoch time field. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "323406": "I referred  [kernel][1] . Kindly help me to understand   \n\n - Why should we need to use \"3000000000\" here? \n - Why should we divide cilck_time with 10 ** 9?  \n\n&gt;     D = 2 ** 26\n&gt;     train_df['category'] = (train_df['ip'].astype(str) + \"_\" + train_df['app'].astype(str) + \"_\" + \n&gt;                      train_df['device'].astype( str) + \"_\" + train_df['os'].astype(str)).apply(hash) % D\n&gt;     click_buffer = np.full(D, 3000000000, dtype=np.uint32)\n&gt;     train_df['epochtime'] = train_df['click_time'].astype(np.int64) // 10 ** 9\n&gt;     next_clicks = []\n&gt;     for category, t in zip(reversed(train_df['category'].values), reversed(train_df['epochtime'].values)):\n&gt;                 next_clicks.append(click_buffer[category] - t)\n&gt;                 click_buffer[category] = t\n\n  \n\n\n  [1]: https://www.kaggle.com/bk0000/non-blending-lightgbm-model-lb-0-977",
    "323420": "Use [this kernel][1] by Andrii Sydorchuk for FE'ing nextclick instead. It's a lot faster at building this feature and you don't have to worry about the remotely minute potential for hash conflicts. But to answer your Qs, the 300xxx number is just a large placeholder (think of it has fillna placeholder that fits in uint32 but doesn't occur in our samples). And the 10 ** 9 is to get you down to seconds, which is the highest resolution we're provided in the dataset anyway.\n\n\n  [1]: https://www.kaggle.com/asydorchuk",
    "323424": "thank you",
    "323436": "little curious. Kindly help me.\n\n - Why should we do reverse?  \n       ```python\n          zip(reversed(train_df['category'].values), reversed(train_df['epochtime'].values))\n      ```  \n - Why should we minus click_time from click_buffer placeholder value(3000xxx)?  \n     ```python\n      next_clicks.append(click_buffer[category] - t)\n   ```",
    "323665": "Pallavi \n\nI'l answer your second question first and that leads into the first one \n\nThis click time can be calculated without the hashing trick as you may know but will create a NAN value for the last record in every group since there is no record you can compare it against.  A default value is therefore used which is the 3 billion from which the click_time of the last record is subtracted \n\nWhy Reversed - Records are in chronological order of click time. So you get the last click of every category and calculate the time delta using the 3B value and then go on to the clicks earlier in time. \n\nHope this helps \nRegards\nShanth",
    "323727": "Thanks for response. Its useful.\nI agree with your answer for 2nd question. But, Why only 3B ? Why not any other number?",
    "323758": "When you notice the value generated for the column epoch time, its a very large INT. Choosing 3 Billion could have to do with the max value of the epoch time field."
  },
  "source": "meta"
}