{
  "id": 53384,
  "title": "proper hour aggregations require test_supplement",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/53384",
  "author_name": "",
  "post_date": "2018-03-29T21:07:27.536071400Z",
  "votes": 6,
  "comment_count": 12,
  "views": 0,
  "content": "<p>The official test dataset contains one second of data from each of hours 6, 11, and 15.  If you aggregate these hours using the official test data, you will get mostly nonsense (only a little bit of nonsense, mind you: one second might not have much impact on your results).  It might be better to recode these hours to the previous hour if you're using only the official test dataset.  To handle them in a consistent way, you need to aggregate using the supplement data, which contains the full hours.  This may not give the best results either, because the first second of an hour is not a typical second.  But at least it's internally consistent.</p>",
  "messages": [
    {
      "id": "306094",
      "postDate": "03/29/2018 21:07:27",
      "content": "<p>The official test dataset contains one second of data from each of hours 6, 11, and 15.  If you aggregate these hours using the official test data, you will get mostly nonsense (only a little bit of nonsense, mind you: one second might not have much impact on your results).  It might be better to recode these hours to the previous hour if you're using only the official test dataset.  To handle them in a consistent way, you need to aggregate using the supplement data, which contains the full hours.  This may not give the best results either, because the first second of an hour is not a typical second.  But at least it's internally consistent.</p>",
      "rawMarkdown": "The official test dataset contains one second of data from each of hours 6, 11, and 15.  If you aggregate these hours using the official test data, you will get mostly nonsense (only a little bit of nonsense, mind you: one second might not have much impact on your results).  It might be better to recode these hours to the previous hour if you're using only the official test dataset.  To handle them in a consistent way, you need to aggregate using the supplement data, which contains the full hours.  This may not give the best results either, because the first second of an hour is not a typical second.  But at least it's internally consistent.",
      "votes": null
    },
    {
      "id": "306167",
      "postDate": "03/30/2018 01:34:54",
      "content": "<p>Also, none of these weird first seconds of the hour is in the public test data.  So the only way we can gauge these effects is from validation data.</p>",
      "rawMarkdown": "Also, none of these weird first seconds of the hour is in the public test data.  So the only way we can gauge these effects is from validation data.",
      "votes": null
    },
    {
      "id": "306286",
      "postDate": "03/30/2018 07:23:03",
      "content": "<p>You are not wrong!  The same goes with doing any time delta stuff</p>",
      "rawMarkdown": "You are not wrong!  The same goes with doing any time delta stuff",
      "votes": null
    },
    {
      "id": "306940",
      "postDate": "03/31/2018 11:24:22",
      "content": "<p>@Andy, thanks for the heads up !</p>",
      "rawMarkdown": "Andy, thanks for the heads up !",
      "votes": null
    },
    {
      "id": "307014",
      "postDate": "03/31/2018 16:00:31",
      "content": "<p>@Andy, Thank you very much for sharing this observation. It indeed seems to prove useful if one uses only the smaller \"official\" test data.</p>",
      "rawMarkdown": "Andy, Thank you very much for sharing this observation. It indeed seems to prove useful if one uses only the smaller \"official\" test data.",
      "votes": null
    },
    {
      "id": "307036",
      "postDate": "03/31/2018 17:13:57",
      "content": "<p>My LB dropped 0.0006 when used test_supplement.csv for feature aggregation instead of test.csv? Has any one experienced this?</p>",
      "rawMarkdown": "My LB dropped 0.0006 when used test_supplement.csv for feature aggregation instead of test.csv? Has any one experienced this?",
      "votes": null
    },
    {
      "id": "307046",
      "postDate": "03/31/2018 17:48:10",
      "content": "<blockquote>\n  <p>My LB dropped 0.0006 when used test_supplement.csv for feature aggregation instead of test.csv? Has any one experienced this?</p>\n</blockquote>\n\n<p>I believe a technical term for that is \"random variation.\"</p>",
      "rawMarkdown": "&gt; My LB dropped 0.0006 when used test_supplement.csv for feature aggregation instead of test.csv? Has any one experienced this?\n\nI believe a technical term for that is \"random variation.\"",
      "votes": null
    },
    {
      "id": "307060",
      "postDate": "03/31/2018 18:17:08",
      "content": "<p>How may I know what caused this random variation? Using same model(I load the saved model) when predict on test.csv I get same(better) result  vs when predicting on test_supplement.csv(0.0006 low).  I am using @CPMP proposed <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53378#306201\">method</a> to map test_supplement predictions to test.</p>\n\n<p>EDIT: I make aggregation features separately for each day.</p>",
      "rawMarkdown": "How may I know what caused this random variation? Using same model(I load the saved model) when predict on test.csv I get same(better) result  vs when predicting on test_supplement.csv(0.0006 low).  I am using @CPMP proposed [method][1] to map test_supplement predictions to test.\n\nEDIT: I make aggregation features separately for each day.\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53378#306201",
      "votes": null
    },
    {
      "id": "307092",
      "postDate": "03/31/2018 19:27:30",
      "content": "<blockquote>\n  <p>How may I know what caused this random variation?</p>\n</blockquote>\n\n<p>@Sohaib Omar I should have been more explicit. I have no idea what causes this difference. My point was that I would consider a LB difference of 0.0006 to be relatively unimportant given that it reflects only 18% of test data.</p>",
      "rawMarkdown": "&gt; How may I know what caused this random variation?\n\n@Sohaib Omar I should have been more explicit. I have no idea what causes this difference. My point was that I would consider a LB difference of 0.0006 to be relatively unimportant given that it reflects only 18% of test data.",
      "votes": null
    },
    {
      "id": "307493",
      "postDate": "04/01/2018 20:07:33",
      "content": "<pre><code>It indeed seems to prove useful if one uses only the smaller \"official\" test data.\n</code></pre>\n\n<p>My LB decreased when I aggregated using test_supplement.csv. still trying to figure out if I should use test.csv or test_supplement.csv.</p>",
      "rawMarkdown": "It indeed seems to prove useful if one uses only the smaller \"official\" test data.\n\nMy LB decreased when I aggregated using test_supplement.csv. still trying to figure out if I should use test.csv or test_supplement.csv.",
      "votes": null
    },
    {
      "id": "307503",
      "postDate": "04/01/2018 20:56:16",
      "content": "<p>My guess would be that it depends on the feature in question.</p>",
      "rawMarkdown": "My guess would be that it depends on the feature in question.",
      "votes": null
    },
    {
      "id": "307861",
      "postDate": "04/02/2018 15:58:54",
      "content": "<p>For hour aggregations, I don't think it should make any difference in the <em>public</em> LB score — except random variation, given the way, e.g., gradient boosting algorithms are often touchy about irrelevant details.  (You could run with different random seeds and see if this result is in the same range of differences.)  The public LB scores are based on one hour of data, which should be the same in both test_supplement and the official test data.  It might make a difference with aggregations that don't include hour (e.g., if you count the total number of clicks by IP address across the whole data set).  In theory, it should still be better to use the supplement data, since you should get more accurate aggregations that way, but since we're not really clear on how the data were selected, I'm a little uncertain about that.</p>",
      "rawMarkdown": "For hour aggregations, I don't think it should make any difference in the *public* LB score — except random variation, given the way, e.g., gradient boosting algorithms are often touchy about irrelevant details.  (You could run with different random seeds and see if this result is in the same range of differences.)  The public LB scores are based on one hour of data, which should be the same in both test_supplement and the official test data.  It might make a difference with aggregations that don't include hour (e.g., if you count the total number of clicks by IP address across the whole data set).  In theory, it should still be better to use the supplement data, since you should get more accurate aggregations that way, but since we're not really clear on how the data were selected, I'm a little uncertain about that.",
      "votes": null
    },
    {
      "id": "307866",
      "postDate": "04/02/2018 16:11:48",
      "content": "<p>Thanks @Andy, I will run same model on test_supplement with different seeds and share the results.<br> I think this variation could be due to hour 6, 11 and 15, as hour aggregation features will have different range for these hours in test_supplement. for ex: in test.csv only hour 6,11,15 contains clicks of only one second(as found by you). </p>",
      "rawMarkdown": "Thanks @Andy, I will run same model on test_supplement with different seeds and share the results.<br> I think this variation could be due to hour 6, 11 and 15, as hour aggregation features will have different range for these hours in test_supplement. for ex: in test.csv only hour 6,11,15 contains clicks of only one second(as found by you).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 306167,
      "author_name": "aharless",
      "author_url": "",
      "post_date": "03/30/2018 01:34:54",
      "content": "<p>Also, none of these weird first seconds of the hour is in the public test data.  So the only way we can gauge these effects is from validation data.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 306286,
      "author_name": "scirpus",
      "author_url": "",
      "post_date": "03/30/2018 07:23:03",
      "content": "<p>You are not wrong!  The same goes with doing any time delta stuff</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 306940,
      "author_name": "ogrellier",
      "author_url": "",
      "post_date": "03/31/2018 11:24:22",
      "content": "<p>@Andy, thanks for the heads up !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 307014,
      "author_name": "asparuhhristov",
      "author_url": "",
      "post_date": "03/31/2018 16:00:31",
      "content": "<p>@Andy, Thank you very much for sharing this observation. It indeed seems to prove useful if one uses only the smaller \"official\" test data.</p>",
      "votes": null,
      "replies": [
        {
          "id": 307493,
          "author_name": "sohaibomar",
          "author_url": "",
          "post_date": "04/01/2018 20:07:33",
          "content": "<pre><code>It indeed seems to prove useful if one uses only the smaller \"official\" test data.\n</code></pre>\n\n<p>My LB decreased when I aggregated using test_supplement.csv. still trying to figure out if I should use test.csv or test_supplement.csv.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 307503,
          "author_name": "asparuhhristov",
          "author_url": "",
          "post_date": "04/01/2018 20:56:16",
          "content": "<p>My guess would be that it depends on the feature in question.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 307861,
          "author_name": "aharless",
          "author_url": "",
          "post_date": "04/02/2018 15:58:54",
          "content": "<p>For hour aggregations, I don't think it should make any difference in the <em>public</em> LB score — except random variation, given the way, e.g., gradient boosting algorithms are often touchy about irrelevant details.  (You could run with different random seeds and see if this result is in the same range of differences.)  The public LB scores are based on one hour of data, which should be the same in both test_supplement and the official test data.  It might make a difference with aggregations that don't include hour (e.g., if you count the total number of clicks by IP address across the whole data set).  In theory, it should still be better to use the supplement data, since you should get more accurate aggregations that way, but since we're not really clear on how the data were selected, I'm a little uncertain about that.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 307866,
          "author_name": "sohaibomar",
          "author_url": "",
          "post_date": "04/02/2018 16:11:48",
          "content": "<p>Thanks @Andy, I will run same model on test_supplement with different seeds and share the results.<br> I think this variation could be due to hour 6, 11 and 15, as hour aggregation features will have different range for these hours in test_supplement. for ex: in test.csv only hour 6,11,15 contains clicks of only one second(as found by you). </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 307036,
      "author_name": "sohaibomar",
      "author_url": "",
      "post_date": "03/31/2018 17:13:57",
      "content": "<p>My LB dropped 0.0006 when used test_supplement.csv for feature aggregation instead of test.csv? Has any one experienced this?</p>",
      "votes": null,
      "replies": [
        {
          "id": 307046,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "03/31/2018 17:48:10",
          "content": "<blockquote>\n  <p>My LB dropped 0.0006 when used test_supplement.csv for feature aggregation instead of test.csv? Has any one experienced this?</p>\n</blockquote>\n\n<p>I believe a technical term for that is \"random variation.\"</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 307060,
          "author_name": "sohaibomar",
          "author_url": "",
          "post_date": "03/31/2018 18:17:08",
          "content": "<p>How may I know what caused this random variation? Using same model(I load the saved model) when predict on test.csv I get same(better) result  vs when predicting on test_supplement.csv(0.0006 low).  I am using @CPMP proposed <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53378#306201\">method</a> to map test_supplement predictions to test.</p>\n\n<p>EDIT: I make aggregation features separately for each day.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 307092,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "03/31/2018 19:27:30",
          "content": "<blockquote>\n  <p>How may I know what caused this random variation?</p>\n</blockquote>\n\n<p>@Sohaib Omar I should have been more explicit. I have no idea what causes this difference. My point was that I would consider a LB difference of 0.0006 to be relatively unimportant given that it reflects only 18% of test data.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "306094": "The official test dataset contains one second of data from each of hours 6, 11, and 15.  If you aggregate these hours using the official test data, you will get mostly nonsense (only a little bit of nonsense, mind you: one second might not have much impact on your results).  It might be better to recode these hours to the previous hour if you're using only the official test dataset.  To handle them in a consistent way, you need to aggregate using the supplement data, which contains the full hours.  This may not give the best results either, because the first second of an hour is not a typical second.  But at least it's internally consistent.",
    "306167": "Also, none of these weird first seconds of the hour is in the public test data.  So the only way we can gauge these effects is from validation data.",
    "306286": "You are not wrong!  The same goes with doing any time delta stuff",
    "306940": "Andy, thanks for the heads up !",
    "307014": "Andy, Thank you very much for sharing this observation. It indeed seems to prove useful if one uses only the smaller \"official\" test data.",
    "307036": "My LB dropped 0.0006 when used test_supplement.csv for feature aggregation instead of test.csv? Has any one experienced this?",
    "307046": "&gt; My LB dropped 0.0006 when used test_supplement.csv for feature aggregation instead of test.csv? Has any one experienced this?\n\nI believe a technical term for that is \"random variation.\"",
    "307060": "How may I know what caused this random variation? Using same model(I load the saved model) when predict on test.csv I get same(better) result  vs when predicting on test_supplement.csv(0.0006 low).  I am using @CPMP proposed [method][1] to map test_supplement predictions to test.\n\nEDIT: I make aggregation features separately for each day.\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53378#306201",
    "307092": "&gt; How may I know what caused this random variation?\n\n@Sohaib Omar I should have been more explicit. I have no idea what causes this difference. My point was that I would consider a LB difference of 0.0006 to be relatively unimportant given that it reflects only 18% of test data.",
    "307493": "It indeed seems to prove useful if one uses only the smaller \"official\" test data.\n\nMy LB decreased when I aggregated using test_supplement.csv. still trying to figure out if I should use test.csv or test_supplement.csv.",
    "307503": "My guess would be that it depends on the feature in question.",
    "307861": "For hour aggregations, I don't think it should make any difference in the *public* LB score — except random variation, given the way, e.g., gradient boosting algorithms are often touchy about irrelevant details.  (You could run with different random seeds and see if this result is in the same range of differences.)  The public LB scores are based on one hour of data, which should be the same in both test_supplement and the official test data.  It might make a difference with aggregations that don't include hour (e.g., if you count the total number of clicks by IP address across the whole data set).  In theory, it should still be better to use the supplement data, since you should get more accurate aggregations that way, but since we're not really clear on how the data were selected, I'm a little uncertain about that.",
    "307866": "Thanks @Andy, I will run same model on test_supplement with different seeds and share the results.<br> I think this variation could be due to hour 6, 11 and 15, as hour aggregation features will have different range for these hours in test_supplement. for ex: in test.csv only hour 6,11,15 contains clicks of only one second(as found by you)."
  },
  "source": "meta"
}