{
  "id": 53378,
  "title": "Test and test supplement click ids are not the same",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/53378",
  "author_name": "",
  "post_date": "2018-03-29T19:56:04.935986Z",
  "votes": 17,
  "comment_count": 35,
  "views": 0,
  "content": "<p>The data description says that the competition test set is a subset of the test supplement (the old test set). This is not strictly true, because the click_id column differs between the two. In particular, the click ids in the new test set are used for completely different clicks in the supplement.</p>\n\n<p>I found this out the hard way when I submitted predictions on a version of the supplement filtered by the overlapping click ids. I was rewarded with a nice .4998 AUC. Maybe this was obvious to others, but it was not obvious to me :)</p>\n\n<p>Luckily, Alexander Firsov has already provided a mapping from supplement click ids (old) to proper test click ids (new) that you can use to filter. Check out his <a href=\"https://www.kaggle.com/alexfir/mapping-between-old-test-and-new-test\">great and so far under-appreciated kernel here</a></p>",
  "messages": [
    {
      "id": "306050",
      "postDate": "03/29/2018 19:56:04",
      "content": "<p>The data description says that the competition test set is a subset of the test supplement (the old test set). This is not strictly true, because the click_id column differs between the two. In particular, the click ids in the new test set are used for completely different clicks in the supplement.</p>\n\n<p>I found this out the hard way when I submitted predictions on a version of the supplement filtered by the overlapping click ids. I was rewarded with a nice .4998 AUC. Maybe this was obvious to others, but it was not obvious to me :)</p>\n\n<p>Luckily, Alexander Firsov has already provided a mapping from supplement click ids (old) to proper test click ids (new) that you can use to filter. Check out his <a href=\"https://www.kaggle.com/alexfir/mapping-between-old-test-and-new-test\">great and so far under-appreciated kernel here</a></p>",
      "rawMarkdown": "The data description says that the competition test set is a subset of the test supplement (the old test set). This is not strictly true, because the click_id column differs between the two. In particular, the click ids in the new test set are used for completely different clicks in the supplement.\n\nI found this out the hard way when I submitted predictions on a version of the supplement filtered by the overlapping click ids. I was rewarded with a nice .4998 AUC. Maybe this was obvious to others, but it was not obvious to me :)\n\nLuckily, Alexander Firsov has already provided a mapping from supplement click ids (old) to proper test click ids (new) that you can use to filter. Check out his [great and so far under-appreciated kernel here][1]\n\n\n  [1]: https://www.kaggle.com/alexfir/mapping-between-old-test-and-new-test",
      "votes": null
    },
    {
      "id": "306201",
      "postDate": "03/30/2018 02:53:03",
      "content": "<p>I'm using a join to map test_supplement to test, this is safe, and not so slow:</p>\n\n<pre><code>print(\"Predicting the submission data...\")\n\ntest_supplement['is_attributed'] = lgb_model.predict(test_supplement[predictors], num_iteration=lgb_model.best_iteration)\n\nprint('projecting prediction onto test')\n\njoin_cols = ['ip', 'app', 'device', 'os', 'channel', 'click_time']\nall_cols = join_cols + ['is_attributed']\n\ntest = test.merge(test_supplement[all_cols], how='left', on=join_cols)\n\ntest = test.drop_duplicates(subset=['click_id'])\n\nprint(\"Writing the submission data into a csv file...\")\n\ntest[['click_id', 'is_attributed']].to_csv('sub.csv', index=False)\n\nprint(\"All done...\")\n</code></pre>",
      "rawMarkdown": "I'm using a join to map test_supplement to test, this is safe, and not so slow:\n\n    print(\"Predicting the submission data...\")\n    \n    test_supplement['is_attributed'] = lgb_model.predict(test_supplement[predictors], num_iteration=lgb_model.best_iteration)\n    \n    print('projecting prediction onto test')\n    \n    join_cols = ['ip', 'app', 'device', 'os', 'channel', 'click_time']\n    all_cols = join_cols + ['is_attributed']\n    \n    test = test.merge(test_supplement[all_cols], how='left', on=join_cols)\n    \n    test = test.drop_duplicates(subset=['click_id'])\n    \n    print(\"Writing the submission data into a csv file...\")\n    \n    test[['click_id', 'is_attributed']].to_csv('sub.csv', index=False)\n    \n    print(\"All done...\")",
      "votes": null
    },
    {
      "id": "306208",
      "postDate": "03/30/2018 03:15:57",
      "content": "<p>By a similar procedure, I find that there is <a href=\"https://www.kaggle.com/aharless/two-clicks-in-both-train-and-test-supplement\">some overlap</a> between the training data and test_supplement (maybe distinct but indistinguishable clicks).</p>",
      "rawMarkdown": "By a similar procedure, I find that there is [some overlap][1] between the training data and test_supplement (maybe distinct but indistinguishable clicks).\n\n\n [1]: https://www.kaggle.com/aharless/two-clicks-in-both-train-and-test-supplement",
      "votes": null
    },
    {
      "id": "306328",
      "postDate": "03/30/2018 08:48:41",
      "content": "<p>What are the advantages of predicting the entire test set? The way I see it, generating features (aggregations and things that needs the entire test set) on the full test set and only predict on the real test set is less expensive and should give the same predictions, no ?</p>",
      "rawMarkdown": "What are the advantages of predicting the entire test set? The way I see it, generating features (aggregations and things that needs the entire test set) on the full test set and only predict on the real test set is less expensive and should give the same predictions, no ?",
      "votes": null
    },
    {
      "id": "306364",
      "postDate": "03/30/2018 09:46:10",
      "content": "<p>Agreed! Lagging features... Everything Time Series motivated. ConvNets would profit from the whole test set obviously. I can also see some use augmenting the training set though pseudo-labelling.</p>",
      "rawMarkdown": "Agreed! Lagging features... Everything Time Series motivated. ConvNets would profit from the whole test set obviously. I can also see some use augmenting the training set though pseudo-labelling.",
      "votes": null
    },
    {
      "id": "306386",
      "postDate": "03/30/2018 11:00:45",
      "content": "<blockquote>\n  <p>What are the advantages of predicting the entire test set?</p>\n</blockquote>\n\n<p>I found it simpler as  I do feature engineering on train + test_supplement.  There is no need to bother about extracting test with engineered features.</p>",
      "rawMarkdown": "&gt; What are the advantages of predicting the entire test set?\n\nI found it simpler as  I do feature engineering on train + test_supplement.  There is no need to bother about extracting test with engineered features.",
      "votes": null
    },
    {
      "id": "306387",
      "postDate": "03/30/2018 11:07:08",
      "content": "<p>ok, thanks</p>",
      "rawMarkdown": "ok, thanks",
      "votes": null
    },
    {
      "id": "306390",
      "postDate": "03/30/2018 11:15:39",
      "content": "<pre><code>I do feature engineering on train + test_supplement\n</code></pre>\n\n<p>May I ask how do you avoid leakage if you aggregate on train + test(test_supplement)?</p>",
      "rawMarkdown": "I do feature engineering on train + test_supplement\n\nMay I ask how do you avoid leakage if you aggregate on train + test(test_supplement)?",
      "votes": null
    },
    {
      "id": "306391",
      "postDate": "03/30/2018 11:20:18",
      "content": "<p>Define leakage please.</p>",
      "rawMarkdown": "Define leakage please.",
      "votes": null
    },
    {
      "id": "306401",
      "postDate": "03/30/2018 11:31:26",
      "content": "<pre><code>Define leakage please.\n</code></pre>\n\n<p>Influence of future on present. @Joe explained it like this <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53325#306098\">here</a></p>\n\n<pre><code>For example, if you count all the clicks for an ip across all 4 days, your model might not do a good job comparing this with the total count for an ip that only shows up on the last day but might be just as spammy on a percentage basis.\n</code></pre>",
      "rawMarkdown": "Define leakage please.\n\nInfluence of future on present. @Joe explained it like this [here][1]\n\n    For example, if you count all the clicks for an ip across all 4 days, your model might not do a good job comparing this with the total count for an ip that only shows up on the last day but might be just as spammy on a percentage basis.\n\n\n\n\n\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53325#306098",
      "votes": null
    },
    {
      "id": "306403",
      "postDate": "03/30/2018 11:32:23",
      "content": "<p>Thanks, I was asking because usually leakage is defined as including some of the target value in the training features.</p>\n\n<p>You can do FE on all data that without doing what you quote above.  And in my view one has to check if this is harmful or not.  If there is a clear explanation of why this would always be bad then I'd like to see it.</p>",
      "rawMarkdown": "Thanks, I was asking because usually leakage is defined as including some of the target value in the training features.\n\nYou can do FE on all data that without doing what you quote above.  And in my view one has to check if this is harmful or not.  If there is a clear explanation of why this would always be bad then I'd like to see it.",
      "votes": null
    },
    {
      "id": "306406",
      "postDate": "03/30/2018 11:34:46",
      "content": "<p>Thanks for the clarification.</p>",
      "rawMarkdown": "Thanks for the clarification.",
      "votes": null
    },
    {
      "id": "306420",
      "postDate": "03/30/2018 12:01:59",
      "content": "<pre><code>You can do FE on all data that without doing what you quote above.\n</code></pre>\n\n<p>I tried different interaction based features like(concat app_channel then encode them with encoder) but they did not help me so I dropped them. And I don't know about any other leakage free feature engineering techniques which can be use full on all data set(train+test). </p>\n\n<pre><code>And in my view one has to check if this is harmful or not.\n</code></pre>\n\n<p>I haven't done frequency based aggregation on all dataset(train+test), it's on my to-do list now. <br> Feature engineering for me takes a lot of hit and trials. I need to get a lot better at this. </p>",
      "rawMarkdown": "You can do FE on all data that without doing what you quote above.\n\nI tried different interaction based features like(concat app_channel then encode them with encoder) but they did not help me so I dropped them. And I don't know about any other leakage free feature engineering techniques which can be use full on all data set(train+test). \n\n    And in my view one has to check if this is harmful or not.\n\nI haven't done frequency based aggregation on all dataset(train+test), it's on my to-do list now. <br> Feature engineering for me takes a lot of hit and trials. I need to get a lot better at this.",
      "votes": null
    },
    {
      "id": "306425",
      "postDate": "03/30/2018 12:05:58",
      "content": "<p>Read solutions to previous time series competitions, you'll get ideas.  Some will work, some won't.  But you're doing the right thing: try things, be happy when it works, learn when it doesn't.</p>",
      "rawMarkdown": "Read solutions to previous time series competitions, you'll get ideas.  Some will work, some won't.  But you're doing the right thing: try things, be happy when it works, learn when it doesn't.",
      "votes": null
    },
    {
      "id": "306428",
      "postDate": "03/30/2018 12:15:13",
      "content": "<blockquote>\n  <p>try things, be happy when it works, learn when it doesn't.</p>\n</blockquote>\n\n<p>I'm learning a lot at the moment ... :)</p>",
      "rawMarkdown": "&gt; try things, be happy when it works, learn when it doesn't.\n\nI'm learning a lot at the moment ... :)",
      "votes": null
    },
    {
      "id": "306433",
      "postDate": "03/30/2018 12:24:21",
      "content": "<blockquote>\n  <p>I'm learning a lot at the moment ... :)</p>\n</blockquote>\n\n<p>LOL.</p>\n\n<p>Me too ;)</p>\n\n<p>edit: for some reason I cannot upvote you.  Weird.  Will try later.</p>",
      "rawMarkdown": "&gt;    I'm learning a lot at the moment ... :)\n\nLOL.\n\nMe too ;)\n\nedit: for some reason I cannot upvote you.  Weird.  Will try later.",
      "votes": null
    },
    {
      "id": "306518",
      "postDate": "03/30/2018 15:09:17",
      "content": "<blockquote>\n  <p><strong>CPMP wrote</strong></p>\n  \n  <blockquote>\n    <p>Thanks, I was asking because usually leakage is defined as including some of the target value in the training features.</p>\n  </blockquote>\n  \n  <p>You can do FE on all data that without doing what you quote above.  And in my view one has to check if this is harmful or not.  If there is a clear explanation of why this would always be bad then I'd like to see it.</p>\n</blockquote>\n\n<p>Yeah, I have been abusing some terminology here to use \"leaky\" to refer to both leaking the target and leaking information from the future that wouldn't be available at prediction time. </p>\n\n<p>I agree it's always a good idea check with validation, but my strong instinct is that features should only describe characteristics that are known at prediction time. For example, clicks throughout the entire day is known at prediction time, but clicks through day + 3 more days into the future is not known.</p>",
      "rawMarkdown": "&gt; **CPMP wrote**\n&gt; \n&gt; &gt; Thanks, I was asking because usually leakage is defined as including some of the target value in the training features.\n&gt; \n&gt; You can do FE on all data that without doing what you quote above.  And in my view one has to check if this is harmful or not.  If there is a clear explanation of why this would always be bad then I'd like to see it.\n\nYeah, I have been abusing some terminology here to use \"leaky\" to refer to both leaking the target and leaking information from the future that wouldn't be available at prediction time. \n\nI agree it's always a good idea check with validation, but my strong instinct is that features should only describe characteristics that are known at prediction time. For example, clicks throughout the entire day is known at prediction time, but clicks through day + 3 more days into the future is not known.",
      "votes": null
    },
    {
      "id": "306526",
      "postDate": "03/30/2018 15:23:26",
      "content": "<p>@Joe, If I was developing a model for a real production use case then I would not use future info in training that would not be available at prediction time.  I therefore agree with you in general.  </p>\n\n<p>The general case is that our model will be used to make prediction on data unknown at training time.</p>\n\n<p>But here were are not in that case:  we know all the test data. I've seen case where using future features was very helpful (eg in Caesars).</p>",
      "rawMarkdown": "Joe, If I was developing a model for a real production use case then I would not use future info in training that would not be available at prediction time.  I therefore agree with you in general.  \n\nThe general case is that our model will be used to make prediction on data unknown at training time.\n\nBut here were are not in that case:  we know all the test data. I've seen case where using future features was very helpful (eg in Caesars).",
      "votes": null
    },
    {
      "id": "307119",
      "postDate": "03/31/2018 21:07:54",
      "content": "<p>@CPMP I think we're on the same page - I agree that for this competition we should use the whole test data. Here prediction time is essentially the end of the day, looking back to predict all the clicks that happened that day. That definitely feels unrealistic to me, but for the competition it'd be a mistake not to exploit it.  </p>\n\n<p>When I talk about future leaking features here, I mean features that are derived from future days, not later in the same day. I.e. if we train on day 8 using counts partially derived from day 9, that aligns poorly with what we have at prediction time for day 10 (since we don't have day 11). </p>\n\n<p>For the uses of future features you mention, do they extend to training on future features unavailable at training time? Sounds interesting. I don't seem to be able to find anything on the Caesars competition (maybe because it's masters only?)</p>",
      "rawMarkdown": "CPMP I think we're on the same page - I agree that for this competition we should use the whole test data. Here prediction time is essentially the end of the day, looking back to predict all the clicks that happened that day. That definitely feels unrealistic to me, but for the competition it'd be a mistake not to exploit it.  \n\nWhen I talk about future leaking features here, I mean features that are derived from future days, not later in the same day. I.e. if we train on day 8 using counts partially derived from day 9, that aligns poorly with what we have at prediction time for day 10 (since we don't have day 11). \n\nFor the uses of future features you mention, do they extend to training on future features unavailable at training time? Sounds interesting. I don't seem to be able to find anything on the Caesars competition (maybe because it's masters only?)",
      "votes": null
    },
    {
      "id": "307236",
      "postDate": "04/01/2018 06:47:55",
      "content": "<p>Yes, Caesars was master only.    Some features there included history info, so the features values in the future were including target info for current period.  Using future feature was so foreign to me that I missed that.  Some others didn't...  We don't have history based features here, so this doe snot apply directly.  </p>",
      "rawMarkdown": "Yes, Caesars was master only.    Some features there included history info, so the features values in the future were including target info for current period.  Using future feature was so foreign to me that I missed that.  Some others didn't...  We don't have history based features here, so this doe snot apply directly.",
      "votes": null
    },
    {
      "id": "307241",
      "postDate": "04/01/2018 07:00:23",
      "content": "<p>I believe that the test_supplement will prove to be useful. After all if (and you'd better be) using any test data for feature engineering, I can see only advantages in using MORE test data for the same features :) Next to that the \"small\" test set introduces gaps in the information (hours missing) so one more argument to use the \"full\" test in my opinion.</p>",
      "rawMarkdown": "I believe that the test_supplement will prove to be useful. After all if (and you'd better be) using any test data for feature engineering, I can see only advantages in using MORE test data for the same features :) Next to that the \"small\" test set introduces gaps in the information (hours missing) so one more argument to use the \"full\" test in my opinion.",
      "votes": null
    },
    {
      "id": "309332",
      "postDate": "04/05/2018 04:09:12",
      "content": "<blockquote>\n  <p><strong>Joe Eddy wrote</strong></p>\n  \n  <p>When I talk about future leaking features here, I mean features that are derived from future days, not later in the same day. I.e. if we train on day 8 using counts partially derived from day 9, that aligns poorly with what we have at prediction time for day 10 (since we don't have day 11). </p>\n</blockquote>\n\n<p>As long as you adjust datetimes for China time being UTC+8, yes.</p>",
      "rawMarkdown": "&gt; **Joe Eddy wrote**\n&gt; \n&gt;\n&gt; When I talk about future leaking features here, I mean features that are derived from future days, not later in the same day. I.e. if we train on day 8 using counts partially derived from day 9, that aligns poorly with what we have at prediction time for day 10 (since we don't have day 11). \n\nAs long as you adjust datetimes for China time being UTC+8, yes.",
      "votes": null
    },
    {
      "id": "309356",
      "postDate": "04/05/2018 05:51:36",
      "content": "<p>@Stephen You mean to avoid leakage one should not convert \"click_time\" to UTC+8, right?</p>",
      "rawMarkdown": "Stephen You mean to avoid leakage one should not convert \"click_time\" to UTC+8, right?",
      "votes": null
    },
    {
      "id": "311963",
      "postDate": "04/11/2018 01:43:57",
      "content": "<p>@CPMP Do you use the subset option in <code>test = test.drop_duplicates(subset=['click_id'])</code> to make it work faster?</p>",
      "rawMarkdown": "CPMP Do you use the subset option in `test = test.drop_duplicates(subset=['click_id'])` to make it work faster?",
      "votes": null
    },
    {
      "id": "312086",
      "postDate": "04/11/2018 06:59:28",
      "content": "<p>Let me answer, drop_duplicates is important here as sometimes there are equal records with different click_ids. For example, test has equal records with click_ids = [15, 20], corresponding click_ids from test_supplement is [21290892, 21290894].</p>\n\n<p>Join will produce 4 records:</p>\n\n<pre><code>(15,21290892) -&gt; 15\n(15,21290894) -&gt; 15\n(20,21290892) -&gt; 20\n(20,21290894) -&gt; 20\n</code></pre>",
      "rawMarkdown": "Let me answer, drop_duplicates is important here as sometimes there are equal records with different click_ids. For example, test has equal records with click_ids = [15, 20], corresponding click_ids from test_supplement is [21290892, 21290894].\n\nJoin will produce 4 records:\n\n    (15,21290892) -&gt; 15\n    (15,21290894) -&gt; 15\n    (20,21290892) -&gt; 20\n    (20,21290894) -&gt; 20",
      "votes": null
    },
    {
      "id": "312138",
      "postDate": "04/11/2018 08:56:19",
      "content": "<p>I'm not using drop duplicates but I did notice indeed that there are duplicates with different target values.  How to treat these is open to me yet.  I'm sorry for the delay in responding, but I'm busy with DSB competition.  Will be back here more actively in few days.</p>",
      "rawMarkdown": "I'm not using drop duplicates but I did notice indeed that there are duplicates with different target values.  How to treat these is open to me yet.  I'm sorry for the delay in responding, but I'm busy with DSB competition.  Will be back here more actively in few days.",
      "votes": null
    },
    {
      "id": "312194",
      "postDate": "04/11/2018 11:02:46",
      "content": "<p>Alexander and CPMP -- thank you for clarifying it.</p>",
      "rawMarkdown": "Alexander and CPMP -- thank you for clarifying it.",
      "votes": null
    },
    {
      "id": "323367",
      "postDate": "05/04/2018 23:57:46",
      "content": "<p>I try this and it cannot merge on click_time.  I've tried with the original test.csv and test_supplement.csv click_times and the merge still does not work.  They have the same datatype too.</p>",
      "rawMarkdown": "I try this and it cannot merge on click_time.  I've tried with the original test.csv and test_supplement.csv click_times and the merge still does not work.  They have the same datatype too.",
      "votes": null
    },
    {
      "id": "323418",
      "postDate": "05/05/2018 05:11:18",
      "content": "<p>What do you mean by 'the merge still does not work' ?  It works fine for me at least.</p>",
      "rawMarkdown": "What do you mean by 'the merge still does not work' ?  It works fine for me at least.",
      "votes": null
    },
    {
      "id": "323695",
      "postDate": "05/05/2018 22:33:10",
      "content": "<p>I mean that when it merges all is_attributed is NaN.  </p>",
      "rawMarkdown": "I mean that when it merges all is_attributed is NaN.",
      "votes": null
    },
    {
      "id": "323698",
      "postDate": "05/05/2018 23:13:38",
      "content": "<p>You probably have a type mismatch. Check the types of your columns.</p>",
      "rawMarkdown": "You probably have a type mismatch. Check the types of your columns.",
      "votes": null
    },
    {
      "id": "323699",
      "postDate": "05/05/2018 23:23:15",
      "content": "<p>I check the datatypes of click_time in test and test_supp and they are both datetime64[ns], still no luck.</p>",
      "rawMarkdown": "I check the datatypes of click_time in test and test_supp and they are both datetime64[ns], still no luck.",
      "votes": null
    },
    {
      "id": "323700",
      "postDate": "05/05/2018 23:26:36",
      "content": "<p>You use the code I shared above?</p>",
      "rawMarkdown": "You use the code I shared above?",
      "votes": null
    },
    {
      "id": "323702",
      "postDate": "05/05/2018 23:30:23",
      "content": "<p>I copied the merge and cols exactly.  I even tried to merge with the original test_supplement.csv and test.csv files with no success.</p>\n\n<pre><code>test_cols = ['ip', 'app', 'device', 'os', 'click_time', 'channel', 'click_id']\ntest = pd.read_csv(path+'test.csv', usecols=test_cols, dtype=dtypes)\ntest_supp = pd.read_csv('submits/lgbm_supp.csv', usecols=list(set(test_cols)-{'click_id'})+['is_attributed'], dtype=dtypes)\n#test_supp = pd.read_csv(path+'test_supplement.csv', usecols=test_cols)\n#test_supp = test_supp.dropna()\n#test_supp = test_supp.astype(dtypes)\ntest['click_time'] = pd.to_datetime(test.click_time)\ntest_supp['click_time'] = pd.to_datetime(test_supp.click_time)\n\njoin_cols = ['ip', 'app', 'device', 'os', 'channel', 'click_time']\n#join_cols = ['click_time']\nall_cols = join_cols + ['is_attributed']\n\nprint('Test:')\nprint(test.click_time.dtype, test_supp.click_time.dtype)\n\ntest = test.merge(test_supp[all_cols], how='left', on=join_cols)\ntest = test.drop_duplicates(subset=['click_id'])\n</code></pre>",
      "rawMarkdown": "I copied the merge and cols exactly.  I even tried to merge with the original test_supplement.csv and test.csv files with no success.\n\n    test_cols = ['ip', 'app', 'device', 'os', 'click_time', 'channel', 'click_id']\n    test = pd.read_csv(path+'test.csv', usecols=test_cols, dtype=dtypes)\n    test_supp = pd.read_csv('submits/lgbm_supp.csv', usecols=list(set(test_cols)-{'click_id'})+['is_attributed'], dtype=dtypes)\n    #test_supp = pd.read_csv(path+'test_supplement.csv', usecols=test_cols)\n    #test_supp = test_supp.dropna()\n    #test_supp = test_supp.astype(dtypes)\n    test['click_time'] = pd.to_datetime(test.click_time)\n    test_supp['click_time'] = pd.to_datetime(test_supp.click_time)\n    \n    join_cols = ['ip', 'app', 'device', 'os', 'channel', 'click_time']\n    #join_cols = ['click_time']\n    all_cols = join_cols + ['is_attributed']\n    \n    print('Test:')\n    print(test.click_time.dtype, test_supp.click_time.dtype)\n    \n    test = test.merge(test_supp[all_cols], how='left', on=join_cols)\n    test = test.drop_duplicates(subset=['click_id'])",
      "votes": null
    },
    {
      "id": "323972",
      "postDate": "05/06/2018 20:10:26",
      "content": "<p>Did you solve your issues Callum?\nI had a similar issue where merging didn't work, but my dtypes were different.\nI tried your code and I got all 'NaT' in the 'click_time' column, so I added 'click_time': 'object' to the dtypes. I also had to set usecols to a simple array:</p>\n\n<pre><code>dtypes = {\n    'ip' :'uint32',\n    'app' :'uint16',\n    'device': 'uint16',\n    'os' :'uint16',\n    'channel': 'uint16',\n    'is_attributed': 'uint8',\n    'click_id': 'uint32',\n    'click_time': 'object'\n}\n\ntest_cols = ['ip', 'app', 'device', 'os', 'click_time', 'channel', 'click_id']\ntest_supp_cols = ['ip', 'app', 'device', 'os', 'click_time', 'channel']\ntest = pd.read_csv(path+'/test.csv', usecols=test_cols, dtype=dtypes)\ntest_supp = pd.read_csv(path+'/test_supplement.csv', usecols=test_supp_cols, dtype=dtypes)\ntest['click_time'] = pd.to_datetime(test.click_time)\ntest_supp['click_time'] = pd.to_datetime(test_supp.click_time)\ntest_supp['is_attributed'] = 0.5\njoin_cols = ['ip', 'app', 'device', 'os', 'channel', 'click_time']\nall_cols = join_cols + ['is_attributed']\n\nprint('Test:')\nprint(test.click_time.dtype, test_supp.click_time.dtype)\n\ntest = test.merge(test_supp[all_cols], how='left', on=join_cols)\ntest = test.drop_duplicates(subset=['click_id'])\nprint(test['is_attributed'].mean())\n</code></pre>",
      "rawMarkdown": "Did you solve your issues Callum?\nI had a similar issue where merging didn't work, but my dtypes were different.\nI tried your code and I got all 'NaT' in the 'click_time' column, so I added 'click_time': 'object' to the dtypes. I also had to set usecols to a simple array:\n\n    \n    dtypes = {\n        'ip' :'uint32',\n        'app' :'uint16',\n        'device': 'uint16',\n        'os' :'uint16',\n        'channel': 'uint16',\n        'is_attributed': 'uint8',\n        'click_id': 'uint32',\n        'click_time': 'object'\n    }\n    \n    test_cols = ['ip', 'app', 'device', 'os', 'click_time', 'channel', 'click_id']\n    test_supp_cols = ['ip', 'app', 'device', 'os', 'click_time', 'channel']\n    test = pd.read_csv(path+'/test.csv', usecols=test_cols, dtype=dtypes)\n    test_supp = pd.read_csv(path+'/test_supplement.csv', usecols=test_supp_cols, dtype=dtypes)\n    test['click_time'] = pd.to_datetime(test.click_time)\n    test_supp['click_time'] = pd.to_datetime(test_supp.click_time)\n    test_supp['is_attributed'] = 0.5\n    join_cols = ['ip', 'app', 'device', 'os', 'channel', 'click_time']\n    all_cols = join_cols + ['is_attributed']\n    \n    print('Test:')\n    print(test.click_time.dtype, test_supp.click_time.dtype)\n    \n    test = test.merge(test_supp[all_cols], how='left', on=join_cols)\n    test = test.drop_duplicates(subset=['click_id'])\n    print(test['is_attributed'].mean())",
      "votes": null
    },
    {
      "id": "323981",
      "postDate": "05/06/2018 20:34:27",
      "content": "<p>No I didn't.  I try using your code and it fails because it find NaN values in certain columns of test_supp.</p>",
      "rawMarkdown": "No I didn't.  I try using your code and it fails because it find NaN values in certain columns of test_supp.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 306201,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "03/30/2018 02:53:03",
      "content": "<p>I'm using a join to map test_supplement to test, this is safe, and not so slow:</p>\n\n<pre><code>print(\"Predicting the submission data...\")\n\ntest_supplement['is_attributed'] = lgb_model.predict(test_supplement[predictors], num_iteration=lgb_model.best_iteration)\n\nprint('projecting prediction onto test')\n\njoin_cols = ['ip', 'app', 'device', 'os', 'channel', 'click_time']\nall_cols = join_cols + ['is_attributed']\n\ntest = test.merge(test_supplement[all_cols], how='left', on=join_cols)\n\ntest = test.drop_duplicates(subset=['click_id'])\n\nprint(\"Writing the submission data into a csv file...\")\n\ntest[['click_id', 'is_attributed']].to_csv('sub.csv', index=False)\n\nprint(\"All done...\")\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 306208,
          "author_name": "aharless",
          "author_url": "",
          "post_date": "03/30/2018 03:15:57",
          "content": "<p>By a similar procedure, I find that there is <a href=\"https://www.kaggle.com/aharless/two-clicks-in-both-train-and-test-supplement\">some overlap</a> between the training data and test_supplement (maybe distinct but indistinguishable clicks).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 306328,
          "author_name": "comartel",
          "author_url": "",
          "post_date": "03/30/2018 08:48:41",
          "content": "<p>What are the advantages of predicting the entire test set? The way I see it, generating features (aggregations and things that needs the entire test set) on the full test set and only predict on the real test set is less expensive and should give the same predictions, no ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 306364,
          "author_name": "welcomeworld",
          "author_url": "",
          "post_date": "03/30/2018 09:46:10",
          "content": "<p>Agreed! Lagging features... Everything Time Series motivated. ConvNets would profit from the whole test set obviously. I can also see some use augmenting the training set though pseudo-labelling.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 306386,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "03/30/2018 11:00:45",
          "content": "<blockquote>\n  <p>What are the advantages of predicting the entire test set?</p>\n</blockquote>\n\n<p>I found it simpler as  I do feature engineering on train + test_supplement.  There is no need to bother about extracting test with engineered features.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 306387,
          "author_name": "comartel",
          "author_url": "",
          "post_date": "03/30/2018 11:07:08",
          "content": "<p>ok, thanks</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 306390,
          "author_name": "sohaibomar",
          "author_url": "",
          "post_date": "03/30/2018 11:15:39",
          "content": "<pre><code>I do feature engineering on train + test_supplement\n</code></pre>\n\n<p>May I ask how do you avoid leakage if you aggregate on train + test(test_supplement)?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 306391,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "03/30/2018 11:20:18",
          "content": "<p>Define leakage please.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 306401,
          "author_name": "sohaibomar",
          "author_url": "",
          "post_date": "03/30/2018 11:31:26",
          "content": "<pre><code>Define leakage please.\n</code></pre>\n\n<p>Influence of future on present. @Joe explained it like this <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53325#306098\">here</a></p>\n\n<pre><code>For example, if you count all the clicks for an ip across all 4 days, your model might not do a good job comparing this with the total count for an ip that only shows up on the last day but might be just as spammy on a percentage basis.\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 306403,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "03/30/2018 11:32:23",
          "content": "<p>Thanks, I was asking because usually leakage is defined as including some of the target value in the training features.</p>\n\n<p>You can do FE on all data that without doing what you quote above.  And in my view one has to check if this is harmful or not.  If there is a clear explanation of why this would always be bad then I'd like to see it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 306406,
          "author_name": "sohaibomar",
          "author_url": "",
          "post_date": "03/30/2018 11:34:46",
          "content": "<p>Thanks for the clarification.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 306420,
          "author_name": "sohaibomar",
          "author_url": "",
          "post_date": "03/30/2018 12:01:59",
          "content": "<pre><code>You can do FE on all data that without doing what you quote above.\n</code></pre>\n\n<p>I tried different interaction based features like(concat app_channel then encode them with encoder) but they did not help me so I dropped them. And I don't know about any other leakage free feature engineering techniques which can be use full on all data set(train+test). </p>\n\n<pre><code>And in my view one has to check if this is harmful or not.\n</code></pre>\n\n<p>I haven't done frequency based aggregation on all dataset(train+test), it's on my to-do list now. <br> Feature engineering for me takes a lot of hit and trials. I need to get a lot better at this. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 306425,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "03/30/2018 12:05:58",
          "content": "<p>Read solutions to previous time series competitions, you'll get ideas.  Some will work, some won't.  But you're doing the right thing: try things, be happy when it works, learn when it doesn't.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 306428,
          "author_name": "comartel",
          "author_url": "",
          "post_date": "03/30/2018 12:15:13",
          "content": "<blockquote>\n  <p>try things, be happy when it works, learn when it doesn't.</p>\n</blockquote>\n\n<p>I'm learning a lot at the moment ... :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 306433,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "03/30/2018 12:24:21",
          "content": "<blockquote>\n  <p>I'm learning a lot at the moment ... :)</p>\n</blockquote>\n\n<p>LOL.</p>\n\n<p>Me too ;)</p>\n\n<p>edit: for some reason I cannot upvote you.  Weird.  Will try later.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 306518,
          "author_name": "aquatic",
          "author_url": "",
          "post_date": "03/30/2018 15:09:17",
          "content": "<blockquote>\n  <p><strong>CPMP wrote</strong></p>\n  \n  <blockquote>\n    <p>Thanks, I was asking because usually leakage is defined as including some of the target value in the training features.</p>\n  </blockquote>\n  \n  <p>You can do FE on all data that without doing what you quote above.  And in my view one has to check if this is harmful or not.  If there is a clear explanation of why this would always be bad then I'd like to see it.</p>\n</blockquote>\n\n<p>Yeah, I have been abusing some terminology here to use \"leaky\" to refer to both leaking the target and leaking information from the future that wouldn't be available at prediction time. </p>\n\n<p>I agree it's always a good idea check with validation, but my strong instinct is that features should only describe characteristics that are known at prediction time. For example, clicks throughout the entire day is known at prediction time, but clicks through day + 3 more days into the future is not known.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 306526,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "03/30/2018 15:23:26",
          "content": "<p>@Joe, If I was developing a model for a real production use case then I would not use future info in training that would not be available at prediction time.  I therefore agree with you in general.  </p>\n\n<p>The general case is that our model will be used to make prediction on data unknown at training time.</p>\n\n<p>But here were are not in that case:  we know all the test data. I've seen case where using future features was very helpful (eg in Caesars).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 307119,
          "author_name": "aquatic",
          "author_url": "",
          "post_date": "03/31/2018 21:07:54",
          "content": "<p>@CPMP I think we're on the same page - I agree that for this competition we should use the whole test data. Here prediction time is essentially the end of the day, looking back to predict all the clicks that happened that day. That definitely feels unrealistic to me, but for the competition it'd be a mistake not to exploit it.  </p>\n\n<p>When I talk about future leaking features here, I mean features that are derived from future days, not later in the same day. I.e. if we train on day 8 using counts partially derived from day 9, that aligns poorly with what we have at prediction time for day 10 (since we don't have day 11). </p>\n\n<p>For the uses of future features you mention, do they extend to training on future features unavailable at training time? Sounds interesting. I don't seem to be able to find anything on the Caesars competition (maybe because it's masters only?)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 307236,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/01/2018 06:47:55",
          "content": "<p>Yes, Caesars was master only.    Some features there included history info, so the features values in the future were including target info for current period.  Using future feature was so foreign to me that I missed that.  Some others didn't...  We don't have history based features here, so this doe snot apply directly.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 309332,
          "author_name": "smcinerney",
          "author_url": "",
          "post_date": "04/05/2018 04:09:12",
          "content": "<blockquote>\n  <p><strong>Joe Eddy wrote</strong></p>\n  \n  <p>When I talk about future leaking features here, I mean features that are derived from future days, not later in the same day. I.e. if we train on day 8 using counts partially derived from day 9, that aligns poorly with what we have at prediction time for day 10 (since we don't have day 11). </p>\n</blockquote>\n\n<p>As long as you adjust datetimes for China time being UTC+8, yes.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 309356,
          "author_name": "sohaibomar",
          "author_url": "",
          "post_date": "04/05/2018 05:51:36",
          "content": "<p>@Stephen You mean to avoid leakage one should not convert \"click_time\" to UTC+8, right?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 311963,
          "author_name": "graf10a",
          "author_url": "",
          "post_date": "04/11/2018 01:43:57",
          "content": "<p>@CPMP Do you use the subset option in <code>test = test.drop_duplicates(subset=['click_id'])</code> to make it work faster?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 312086,
          "author_name": "alexfir",
          "author_url": "",
          "post_date": "04/11/2018 06:59:28",
          "content": "<p>Let me answer, drop_duplicates is important here as sometimes there are equal records with different click_ids. For example, test has equal records with click_ids = [15, 20], corresponding click_ids from test_supplement is [21290892, 21290894].</p>\n\n<p>Join will produce 4 records:</p>\n\n<pre><code>(15,21290892) -&gt; 15\n(15,21290894) -&gt; 15\n(20,21290892) -&gt; 20\n(20,21290894) -&gt; 20\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 312138,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/11/2018 08:56:19",
          "content": "<p>I'm not using drop duplicates but I did notice indeed that there are duplicates with different target values.  How to treat these is open to me yet.  I'm sorry for the delay in responding, but I'm busy with DSB competition.  Will be back here more actively in few days.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 312194,
          "author_name": "graf10a",
          "author_url": "",
          "post_date": "04/11/2018 11:02:46",
          "content": "<p>Alexander and CPMP -- thank you for clarifying it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 323367,
          "author_name": "cgundlach",
          "author_url": "",
          "post_date": "05/04/2018 23:57:46",
          "content": "<p>I try this and it cannot merge on click_time.  I've tried with the original test.csv and test_supplement.csv click_times and the merge still does not work.  They have the same datatype too.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 323418,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/05/2018 05:11:18",
          "content": "<p>What do you mean by 'the merge still does not work' ?  It works fine for me at least.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 323695,
          "author_name": "cgundlach",
          "author_url": "",
          "post_date": "05/05/2018 22:33:10",
          "content": "<p>I mean that when it merges all is_attributed is NaN.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 323698,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/05/2018 23:13:38",
          "content": "<p>You probably have a type mismatch. Check the types of your columns.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 323699,
          "author_name": "cgundlach",
          "author_url": "",
          "post_date": "05/05/2018 23:23:15",
          "content": "<p>I check the datatypes of click_time in test and test_supp and they are both datetime64[ns], still no luck.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 323700,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/05/2018 23:26:36",
          "content": "<p>You use the code I shared above?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 323702,
          "author_name": "cgundlach",
          "author_url": "",
          "post_date": "05/05/2018 23:30:23",
          "content": "<p>I copied the merge and cols exactly.  I even tried to merge with the original test_supplement.csv and test.csv files with no success.</p>\n\n<pre><code>test_cols = ['ip', 'app', 'device', 'os', 'click_time', 'channel', 'click_id']\ntest = pd.read_csv(path+'test.csv', usecols=test_cols, dtype=dtypes)\ntest_supp = pd.read_csv('submits/lgbm_supp.csv', usecols=list(set(test_cols)-{'click_id'})+['is_attributed'], dtype=dtypes)\n#test_supp = pd.read_csv(path+'test_supplement.csv', usecols=test_cols)\n#test_supp = test_supp.dropna()\n#test_supp = test_supp.astype(dtypes)\ntest['click_time'] = pd.to_datetime(test.click_time)\ntest_supp['click_time'] = pd.to_datetime(test_supp.click_time)\n\njoin_cols = ['ip', 'app', 'device', 'os', 'channel', 'click_time']\n#join_cols = ['click_time']\nall_cols = join_cols + ['is_attributed']\n\nprint('Test:')\nprint(test.click_time.dtype, test_supp.click_time.dtype)\n\ntest = test.merge(test_supp[all_cols], how='left', on=join_cols)\ntest = test.drop_duplicates(subset=['click_id'])\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 323972,
          "author_name": "sjb1988",
          "author_url": "",
          "post_date": "05/06/2018 20:10:26",
          "content": "<p>Did you solve your issues Callum?\nI had a similar issue where merging didn't work, but my dtypes were different.\nI tried your code and I got all 'NaT' in the 'click_time' column, so I added 'click_time': 'object' to the dtypes. I also had to set usecols to a simple array:</p>\n\n<pre><code>dtypes = {\n    'ip' :'uint32',\n    'app' :'uint16',\n    'device': 'uint16',\n    'os' :'uint16',\n    'channel': 'uint16',\n    'is_attributed': 'uint8',\n    'click_id': 'uint32',\n    'click_time': 'object'\n}\n\ntest_cols = ['ip', 'app', 'device', 'os', 'click_time', 'channel', 'click_id']\ntest_supp_cols = ['ip', 'app', 'device', 'os', 'click_time', 'channel']\ntest = pd.read_csv(path+'/test.csv', usecols=test_cols, dtype=dtypes)\ntest_supp = pd.read_csv(path+'/test_supplement.csv', usecols=test_supp_cols, dtype=dtypes)\ntest['click_time'] = pd.to_datetime(test.click_time)\ntest_supp['click_time'] = pd.to_datetime(test_supp.click_time)\ntest_supp['is_attributed'] = 0.5\njoin_cols = ['ip', 'app', 'device', 'os', 'channel', 'click_time']\nall_cols = join_cols + ['is_attributed']\n\nprint('Test:')\nprint(test.click_time.dtype, test_supp.click_time.dtype)\n\ntest = test.merge(test_supp[all_cols], how='left', on=join_cols)\ntest = test.drop_duplicates(subset=['click_id'])\nprint(test['is_attributed'].mean())\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 323981,
          "author_name": "cgundlach",
          "author_url": "",
          "post_date": "05/06/2018 20:34:27",
          "content": "<p>No I didn't.  I try using your code and it fails because it find NaN values in certain columns of test_supp.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 307241,
      "author_name": "asparuhhristov",
      "author_url": "",
      "post_date": "04/01/2018 07:00:23",
      "content": "<p>I believe that the test_supplement will prove to be useful. After all if (and you'd better be) using any test data for feature engineering, I can see only advantages in using MORE test data for the same features :) Next to that the \"small\" test set introduces gaps in the information (hours missing) so one more argument to use the \"full\" test in my opinion.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "306050": "The data description says that the competition test set is a subset of the test supplement (the old test set). This is not strictly true, because the click_id column differs between the two. In particular, the click ids in the new test set are used for completely different clicks in the supplement.\n\nI found this out the hard way when I submitted predictions on a version of the supplement filtered by the overlapping click ids. I was rewarded with a nice .4998 AUC. Maybe this was obvious to others, but it was not obvious to me :)\n\nLuckily, Alexander Firsov has already provided a mapping from supplement click ids (old) to proper test click ids (new) that you can use to filter. Check out his [great and so far under-appreciated kernel here][1]\n\n\n  [1]: https://www.kaggle.com/alexfir/mapping-between-old-test-and-new-test",
    "306201": "I'm using a join to map test_supplement to test, this is safe, and not so slow:\n\n    print(\"Predicting the submission data...\")\n    \n    test_supplement['is_attributed'] = lgb_model.predict(test_supplement[predictors], num_iteration=lgb_model.best_iteration)\n    \n    print('projecting prediction onto test')\n    \n    join_cols = ['ip', 'app', 'device', 'os', 'channel', 'click_time']\n    all_cols = join_cols + ['is_attributed']\n    \n    test = test.merge(test_supplement[all_cols], how='left', on=join_cols)\n    \n    test = test.drop_duplicates(subset=['click_id'])\n    \n    print(\"Writing the submission data into a csv file...\")\n    \n    test[['click_id', 'is_attributed']].to_csv('sub.csv', index=False)\n    \n    print(\"All done...\")",
    "306208": "By a similar procedure, I find that there is [some overlap][1] between the training data and test_supplement (maybe distinct but indistinguishable clicks).\n\n\n [1]: https://www.kaggle.com/aharless/two-clicks-in-both-train-and-test-supplement",
    "306328": "What are the advantages of predicting the entire test set? The way I see it, generating features (aggregations and things that needs the entire test set) on the full test set and only predict on the real test set is less expensive and should give the same predictions, no ?",
    "306364": "Agreed! Lagging features... Everything Time Series motivated. ConvNets would profit from the whole test set obviously. I can also see some use augmenting the training set though pseudo-labelling.",
    "306386": "&gt; What are the advantages of predicting the entire test set?\n\nI found it simpler as  I do feature engineering on train + test_supplement.  There is no need to bother about extracting test with engineered features.",
    "306387": "ok, thanks",
    "306390": "I do feature engineering on train + test_supplement\n\nMay I ask how do you avoid leakage if you aggregate on train + test(test_supplement)?",
    "306391": "Define leakage please.",
    "306401": "Define leakage please.\n\nInfluence of future on present. @Joe explained it like this [here][1]\n\n    For example, if you count all the clicks for an ip across all 4 days, your model might not do a good job comparing this with the total count for an ip that only shows up on the last day but might be just as spammy on a percentage basis.\n\n\n\n\n\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53325#306098",
    "306403": "Thanks, I was asking because usually leakage is defined as including some of the target value in the training features.\n\nYou can do FE on all data that without doing what you quote above.  And in my view one has to check if this is harmful or not.  If there is a clear explanation of why this would always be bad then I'd like to see it.",
    "306406": "Thanks for the clarification.",
    "306420": "You can do FE on all data that without doing what you quote above.\n\nI tried different interaction based features like(concat app_channel then encode them with encoder) but they did not help me so I dropped them. And I don't know about any other leakage free feature engineering techniques which can be use full on all data set(train+test). \n\n    And in my view one has to check if this is harmful or not.\n\nI haven't done frequency based aggregation on all dataset(train+test), it's on my to-do list now. <br> Feature engineering for me takes a lot of hit and trials. I need to get a lot better at this.",
    "306425": "Read solutions to previous time series competitions, you'll get ideas.  Some will work, some won't.  But you're doing the right thing: try things, be happy when it works, learn when it doesn't.",
    "306428": "&gt; try things, be happy when it works, learn when it doesn't.\n\nI'm learning a lot at the moment ... :)",
    "306433": "&gt;    I'm learning a lot at the moment ... :)\n\nLOL.\n\nMe too ;)\n\nedit: for some reason I cannot upvote you.  Weird.  Will try later.",
    "306518": "&gt; **CPMP wrote**\n&gt; \n&gt; &gt; Thanks, I was asking because usually leakage is defined as including some of the target value in the training features.\n&gt; \n&gt; You can do FE on all data that without doing what you quote above.  And in my view one has to check if this is harmful or not.  If there is a clear explanation of why this would always be bad then I'd like to see it.\n\nYeah, I have been abusing some terminology here to use \"leaky\" to refer to both leaking the target and leaking information from the future that wouldn't be available at prediction time. \n\nI agree it's always a good idea check with validation, but my strong instinct is that features should only describe characteristics that are known at prediction time. For example, clicks throughout the entire day is known at prediction time, but clicks through day + 3 more days into the future is not known.",
    "306526": "Joe, If I was developing a model for a real production use case then I would not use future info in training that would not be available at prediction time.  I therefore agree with you in general.  \n\nThe general case is that our model will be used to make prediction on data unknown at training time.\n\nBut here were are not in that case:  we know all the test data. I've seen case where using future features was very helpful (eg in Caesars).",
    "307119": "CPMP I think we're on the same page - I agree that for this competition we should use the whole test data. Here prediction time is essentially the end of the day, looking back to predict all the clicks that happened that day. That definitely feels unrealistic to me, but for the competition it'd be a mistake not to exploit it.  \n\nWhen I talk about future leaking features here, I mean features that are derived from future days, not later in the same day. I.e. if we train on day 8 using counts partially derived from day 9, that aligns poorly with what we have at prediction time for day 10 (since we don't have day 11). \n\nFor the uses of future features you mention, do they extend to training on future features unavailable at training time? Sounds interesting. I don't seem to be able to find anything on the Caesars competition (maybe because it's masters only?)",
    "307236": "Yes, Caesars was master only.    Some features there included history info, so the features values in the future were including target info for current period.  Using future feature was so foreign to me that I missed that.  Some others didn't...  We don't have history based features here, so this doe snot apply directly.",
    "307241": "I believe that the test_supplement will prove to be useful. After all if (and you'd better be) using any test data for feature engineering, I can see only advantages in using MORE test data for the same features :) Next to that the \"small\" test set introduces gaps in the information (hours missing) so one more argument to use the \"full\" test in my opinion.",
    "309332": "&gt; **Joe Eddy wrote**\n&gt; \n&gt;\n&gt; When I talk about future leaking features here, I mean features that are derived from future days, not later in the same day. I.e. if we train on day 8 using counts partially derived from day 9, that aligns poorly with what we have at prediction time for day 10 (since we don't have day 11). \n\nAs long as you adjust datetimes for China time being UTC+8, yes.",
    "309356": "Stephen You mean to avoid leakage one should not convert \"click_time\" to UTC+8, right?",
    "311963": "CPMP Do you use the subset option in `test = test.drop_duplicates(subset=['click_id'])` to make it work faster?",
    "312086": "Let me answer, drop_duplicates is important here as sometimes there are equal records with different click_ids. For example, test has equal records with click_ids = [15, 20], corresponding click_ids from test_supplement is [21290892, 21290894].\n\nJoin will produce 4 records:\n\n    (15,21290892) -&gt; 15\n    (15,21290894) -&gt; 15\n    (20,21290892) -&gt; 20\n    (20,21290894) -&gt; 20",
    "312138": "I'm not using drop duplicates but I did notice indeed that there are duplicates with different target values.  How to treat these is open to me yet.  I'm sorry for the delay in responding, but I'm busy with DSB competition.  Will be back here more actively in few days.",
    "312194": "Alexander and CPMP -- thank you for clarifying it.",
    "323367": "I try this and it cannot merge on click_time.  I've tried with the original test.csv and test_supplement.csv click_times and the merge still does not work.  They have the same datatype too.",
    "323418": "What do you mean by 'the merge still does not work' ?  It works fine for me at least.",
    "323695": "I mean that when it merges all is_attributed is NaN.",
    "323698": "You probably have a type mismatch. Check the types of your columns.",
    "323699": "I check the datatypes of click_time in test and test_supp and they are both datetime64[ns], still no luck.",
    "323700": "You use the code I shared above?",
    "323702": "I copied the merge and cols exactly.  I even tried to merge with the original test_supplement.csv and test.csv files with no success.\n\n    test_cols = ['ip', 'app', 'device', 'os', 'click_time', 'channel', 'click_id']\n    test = pd.read_csv(path+'test.csv', usecols=test_cols, dtype=dtypes)\n    test_supp = pd.read_csv('submits/lgbm_supp.csv', usecols=list(set(test_cols)-{'click_id'})+['is_attributed'], dtype=dtypes)\n    #test_supp = pd.read_csv(path+'test_supplement.csv', usecols=test_cols)\n    #test_supp = test_supp.dropna()\n    #test_supp = test_supp.astype(dtypes)\n    test['click_time'] = pd.to_datetime(test.click_time)\n    test_supp['click_time'] = pd.to_datetime(test_supp.click_time)\n    \n    join_cols = ['ip', 'app', 'device', 'os', 'channel', 'click_time']\n    #join_cols = ['click_time']\n    all_cols = join_cols + ['is_attributed']\n    \n    print('Test:')\n    print(test.click_time.dtype, test_supp.click_time.dtype)\n    \n    test = test.merge(test_supp[all_cols], how='left', on=join_cols)\n    test = test.drop_duplicates(subset=['click_id'])",
    "323972": "Did you solve your issues Callum?\nI had a similar issue where merging didn't work, but my dtypes were different.\nI tried your code and I got all 'NaT' in the 'click_time' column, so I added 'click_time': 'object' to the dtypes. I also had to set usecols to a simple array:\n\n    \n    dtypes = {\n        'ip' :'uint32',\n        'app' :'uint16',\n        'device': 'uint16',\n        'os' :'uint16',\n        'channel': 'uint16',\n        'is_attributed': 'uint8',\n        'click_id': 'uint32',\n        'click_time': 'object'\n    }\n    \n    test_cols = ['ip', 'app', 'device', 'os', 'click_time', 'channel', 'click_id']\n    test_supp_cols = ['ip', 'app', 'device', 'os', 'click_time', 'channel']\n    test = pd.read_csv(path+'/test.csv', usecols=test_cols, dtype=dtypes)\n    test_supp = pd.read_csv(path+'/test_supplement.csv', usecols=test_supp_cols, dtype=dtypes)\n    test['click_time'] = pd.to_datetime(test.click_time)\n    test_supp['click_time'] = pd.to_datetime(test_supp.click_time)\n    test_supp['is_attributed'] = 0.5\n    join_cols = ['ip', 'app', 'device', 'os', 'channel', 'click_time']\n    all_cols = join_cols + ['is_attributed']\n    \n    print('Test:')\n    print(test.click_time.dtype, test_supp.click_time.dtype)\n    \n    test = test.merge(test_supp[all_cols], how='left', on=join_cols)\n    test = test.drop_duplicates(subset=['click_id'])\n    print(test['is_attributed'].mean())",
    "323981": "No I didn't.  I try using your code and it fails because it find NaN values in certain columns of test_supp."
  },
  "source": "meta"
}