{
  "id": 51877,
  "title": "Test data has only half of clicks",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/51877",
  "author_name": "",
  "post_date": "2018-03-13T22:06:57.074722400Z",
  "votes": 54,
  "comment_count": 19,
  "views": 0,
  "content": "<p>I have counted number of records/clicks between 12:00 and 23:00 for each day from train and compared with test data. Interval 12:00-23:00 is a time range of test data (local Chinese time)</p>\n\n<p>Result is the following:</p>\n\n<ul>\n<li>2017-11-07: 34 308 844 - train</li>\n<li>2017-11-08: 36 475 438 - train</li>\n<li>2017-11-09: 37 169 180 - train</li>\n<li>2017-11-10: 18 790 469 - test</li>\n</ul>\n\n<p>It looks like test data has only half of clicks comparing to  any day from train and the same time range as in test.</p>\n\n<p>I have a question to organizers, what is the reason for such a big difference? \nHave you removed half of clicks or this is natural decrease due to Friday evening or something else?</p>",
  "messages": [
    {
      "id": "295596",
      "postDate": "03/13/2018 22:06:57",
      "content": "<p>I have counted number of records/clicks between 12:00 and 23:00 for each day from train and compared with test data. Interval 12:00-23:00 is a time range of test data (local Chinese time)</p>\n\n<p>Result is the following:</p>\n\n<ul>\n<li>2017-11-07: 34 308 844 - train</li>\n<li>2017-11-08: 36 475 438 - train</li>\n<li>2017-11-09: 37 169 180 - train</li>\n<li>2017-11-10: 18 790 469 - test</li>\n</ul>\n\n<p>It looks like test data has only half of clicks comparing to  any day from train and the same time range as in test.</p>\n\n<p>I have a question to organizers, what is the reason for such a big difference? \nHave you removed half of clicks or this is natural decrease due to Friday evening or something else?</p>",
      "rawMarkdown": "I have counted number of records/clicks between 12:00 and 23:00 for each day from train and compared with test data. Interval 12:00-23:00 is a time range of test data (local Chinese time)\n\nResult is the following:\n\n* 2017-11-07: 34 308 844 - train\n* 2017-11-08: 36 475 438 - train\n* 2017-11-09: 37 169 180 - train\n* 2017-11-10: 18 790 469 - test\n\nIt looks like test data has only half of clicks comparing to  any day from train and the same time range as in test.\n\nI have a question to organizers, what is the reason for such a big difference? \nHave you removed half of clicks or this is natural decrease due to Friday evening or something else?",
      "votes": null
    },
    {
      "id": "295642",
      "postDate": "03/14/2018 00:38:17",
      "content": "<p>If you check the hour of the timestamps in the test set:</p>\n\n<p><code>test_df.click_time.str[11:13].value_counts().sort_index()</code></p>\n\n<p>... the records are not missing at random, there are three distinct bands of time with records:</p>\n\n<pre><code>04    3344125\n05    2858427\n06        381\n09    2984808\n10    3127993\n11        413\n13    3212566\n14    3261257\n15        499\nName: click_time, dtype: int64\n</code></pre>\n\n<p>(There are no rows at all with hours 07, 08, 12). So I'd say it's a biased sample of data, not a natural decrease in traffic. I would like to hear the organizers reasons too...</p>",
      "rawMarkdown": "If you check the hour of the timestamps in the test set:\n\n`test_df.click_time.str[11:13].value_counts().sort_index()`\n\n... the records are not missing at random, there are three distinct bands of time with records:\n\n    04    3344125\n    05    2858427\n    06        381\n    09    2984808\n    10    3127993\n    11        413\n    13    3212566\n    14    3261257\n    15        499\n    Name: click_time, dtype: int64\n\n(There are no rows at all with hours 07, 08, 12). So I'd say it's a biased sample of data, not a natural decrease in traffic. I would like to hear the organizers reasons too...",
      "votes": null
    },
    {
      "id": "295675",
      "postDate": "03/14/2018 02:17:21",
      "content": "<p>I guess this means our validation data should be weighted by hour, or just exclude certain times (assuming one follows <a href=\"https://www.kaggle.com/konradb/validation-set/comments#latest-295404\">this kind of validation approach</a>)</p>",
      "rawMarkdown": "I guess this means our validation data should be weighted by hour, or just exclude certain times (assuming one follows [this kind of validation approach][1])\n\n [1]: https://www.kaggle.com/konradb/validation-set/comments#latest-295404",
      "votes": null
    },
    {
      "id": "295739",
      "postDate": "03/14/2018 05:06:53",
      "content": "<p>Perhaps organizers want to limit approaches based on time series.  It may be reasonable considering amount of data.</p>",
      "rawMarkdown": "Perhaps organizers want to limit approaches based on time series.  It may be reasonable considering amount of data.",
      "votes": null
    },
    {
      "id": "295772",
      "postDate": "03/14/2018 06:42:30",
      "content": "<p>The reason why evaluation data is missing records by specific hours is that the original evaluation data is too large and kaggle's platform has a maximum size of evaluation data uploading, that's why we have to truncate the dataset by hours. Sorry for the inconvenience. <br>\nWe choose three ranges of time, as we believed they would play an more significant role in Chinese daily life. <br>\nThe ranges of time are:\n12pm-14pm, 17pm-19pm, 21pm-23pm <br>\nNote: the data's timestamp is in UTC+0, and Chinese timezone is UTC+8</p>",
      "rawMarkdown": "The reason why evaluation data is missing records by specific hours is that the original evaluation data is too large and kaggle's platform has a maximum size of evaluation data uploading, that's why we have to truncate the dataset by hours. Sorry for the inconvenience.   \nWe choose three ranges of time, as we believed they would play an more significant role in Chinese daily life.  \nThe ranges of time are:\n12pm-14pm, 17pm-19pm, 21pm-23pm  \nNote: the data's timestamp is in UTC+0, and Chinese timezone is UTC+8",
      "votes": null
    },
    {
      "id": "295786",
      "postDate": "03/14/2018 07:27:44",
      "content": "<p>Thanks for the clarification. One thing is still a bit unclear though - what do you mean by \"more significant role\"...?</p>",
      "rawMarkdown": "Thanks for the clarification. One thing is still a bit unclear though - what do you mean by \"more significant role\"...?",
      "votes": null
    },
    {
      "id": "295788",
      "postDate": "03/14/2018 07:28:51",
      "content": "<p>Thank you for sharing this insight! </p>",
      "rawMarkdown": "Thank you for sharing this insight!",
      "votes": null
    },
    {
      "id": "295798",
      "postDate": "03/14/2018 07:47:46",
      "content": "<p>I made the same analysis for the old test data - everything seems to be there (in the old one), the clicks per hour follow the same distribution as the one from the training data. Here is a snippet for the hours you provided:</p>\n\n<p>04    3344571</p>\n\n<p>05    2858427</p>\n\n<p>06    2621978</p>\n\n<p>07    2683972</p>\n\n<p>08    2878913</p>\n\n<p>09    2985213</p>\n\n<p>10    3127993</p>\n\n<p>11     3249395</p>\n\n<p>12    3068579</p>\n\n<p>13    3213032</p>\n\n<p>14    3261257</p>\n\n<p>15    2953455</p>",
      "rawMarkdown": "I made the same analysis for the old test data - everything seems to be there (in the old one), the clicks per hour follow the same distribution as the one from the training data. Here is a snippet for the hours you provided:\n\n04    3344571\n\n05    2858427\n\n06    2621978\n\n07    2683972\n\n08    2878913\n\n09    2985213\n\n10    3127993\n\n11     3249395\n\n12    3068579\n\n13    3213032\n\n14    3261257\n\n15    2953455",
      "votes": null
    },
    {
      "id": "295814",
      "postDate": "03/14/2018 08:48:04",
      "content": "<p>If this is the case then I think the old test data should be made available to everyone. Those who have access to it may have an advantage knowing how many clicks certain <code>ips</code>, <code>apps</code>, etc. were clicked during those missing hours.</p>",
      "rawMarkdown": "If this is the case then I think the old test data should be made available to everyone. Those who have access to it may have an advantage knowing how many clicks certain `ips`, `apps`, etc. were clicked during those missing hours.",
      "votes": null
    },
    {
      "id": "295821",
      "postDate": "03/14/2018 08:58:25",
      "content": "<p>Completely agree with you! <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51506\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51506</a></p>",
      "rawMarkdown": "Completely agree with you! https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51506",
      "votes": null
    },
    {
      "id": "295822",
      "postDate": "03/14/2018 09:02:40",
      "content": "<p>omg .. I'll quit this</p>",
      "rawMarkdown": "omg .. I'll quit this",
      "votes": null
    },
    {
      "id": "295824",
      "postDate": "03/14/2018 09:05:29",
      "content": "<p>Thanks Aspurah. I already had it, but good to see it's available to everyone. Should be on the data page, IMO.</p>",
      "rawMarkdown": "Thanks Aspurah. I already had it, but good to see it's available to everyone. Should be on the data page, IMO.",
      "votes": null
    },
    {
      "id": "295998",
      "postDate": "03/14/2018 15:14:21",
      "content": "<p>I believe the records from hours 6, 11, and 15 are all from the first second of the hour.  So it's essentially 4:00:00 to 6:00:00 inclusive, and so on.</p>",
      "rawMarkdown": "I believe the records from hours 6, 11, and 15 are all from the first second of the hour.  So it's essentially 4:00:00 to 6:00:00 inclusive, and so on.",
      "votes": null
    },
    {
      "id": "296106",
      "postDate": "03/14/2018 18:11:38",
      "content": "<p>Agreed, please release the old data. Otherwise those who have it (like myself) have an unfair advantage.</p>",
      "rawMarkdown": "Agreed, please release the old data. Otherwise those who have it (like myself) have an unfair advantage.",
      "votes": null
    },
    {
      "id": "296110",
      "postDate": "03/14/2018 18:20:28",
      "content": "<p>Thanks for sharing! You are a nice person!</p>",
      "rawMarkdown": "Thanks for sharing! You are a nice person!",
      "votes": null
    },
    {
      "id": "296347",
      "postDate": "03/15/2018 04:24:35",
      "content": "<p>@anokas Here it is. <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51506\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51506</a></p>",
      "rawMarkdown": "anokas Here it is. https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51506",
      "votes": null
    },
    {
      "id": "296448",
      "postDate": "03/15/2018 08:19:54",
      "content": "<p>People should be more aware of this. </p>",
      "rawMarkdown": "People should be more aware of this.",
      "votes": null
    },
    {
      "id": "297739",
      "postDate": "03/17/2018 21:10:40",
      "content": "<blockquote>\n  <p><strong>Aaron Yin wrote</strong></p>\n  \n  <blockquote>\n    <p>The reason why evaluation data is missing records by specific hours is that the original evaluation data is too large and kaggle's platform has a maximum size of evaluation data uploading, that's why we have to truncate the dataset by hours. Sorry for the inconvenience.   </p>\n  </blockquote>\n</blockquote>\n\n<p>Wouldn't it be the more reasonable option to have Kaggle adjust their platform limits instead and bring back the original test data? Just saying... it'll save everyone a headache and it'll make the competition results more valuable for you.</p>",
      "rawMarkdown": "&gt; **Aaron Yin wrote**\n&gt; \n&gt; &gt; The reason why evaluation data is missing records by specific hours is that the original evaluation data is too large and kaggle's platform has a maximum size of evaluation data uploading, that's why we have to truncate the dataset by hours. Sorry for the inconvenience.   \n\nWouldn't it be the more reasonable option to have Kaggle adjust their platform limits instead and bring back the original test data? Just saying... it'll save everyone a headache and it'll make the competition results more valuable for you.",
      "votes": null
    },
    {
      "id": "297750",
      "postDate": "03/17/2018 21:55:24",
      "content": "<p>@Toby -</p>\n\n<p>It would require significant platform changes. Eventually, once we're 100% migrated from Azure to GCP, we'll have a lot more flexibility. Until then, it's not a viable option to rework functionality that will be replaced in the short-term anyway.</p>",
      "rawMarkdown": "Toby -\n\nIt would require significant platform changes. Eventually, once we're 100% migrated from Azure to GCP, we'll have a lot more flexibility. Until then, it's not a viable option to rework functionality that will be replaced in the short-term anyway.",
      "votes": null
    },
    {
      "id": "299685",
      "postDate": "03/21/2018 02:01:57",
      "content": "<p>Are we allowed to use old test data for feature engineering or is it against rules?</p>",
      "rawMarkdown": "Are we allowed to use old test data for feature engineering or is it against rules?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 295642,
      "author_name": "jtrotman",
      "author_url": "",
      "post_date": "03/14/2018 00:38:17",
      "content": "<p>If you check the hour of the timestamps in the test set:</p>\n\n<p><code>test_df.click_time.str[11:13].value_counts().sort_index()</code></p>\n\n<p>... the records are not missing at random, there are three distinct bands of time with records:</p>\n\n<pre><code>04    3344125\n05    2858427\n06        381\n09    2984808\n10    3127993\n11        413\n13    3212566\n14    3261257\n15        499\nName: click_time, dtype: int64\n</code></pre>\n\n<p>(There are no rows at all with hours 07, 08, 12). So I'd say it's a biased sample of data, not a natural decrease in traffic. I would like to hear the organizers reasons too...</p>",
      "votes": null,
      "replies": [
        {
          "id": 295739,
          "author_name": "alexfir",
          "author_url": "",
          "post_date": "03/14/2018 05:06:53",
          "content": "<p>Perhaps organizers want to limit approaches based on time series.  It may be reasonable considering amount of data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 295772,
          "author_name": "aaronyin",
          "author_url": "",
          "post_date": "03/14/2018 06:42:30",
          "content": "<p>The reason why evaluation data is missing records by specific hours is that the original evaluation data is too large and kaggle's platform has a maximum size of evaluation data uploading, that's why we have to truncate the dataset by hours. Sorry for the inconvenience. <br>\nWe choose three ranges of time, as we believed they would play an more significant role in Chinese daily life. <br>\nThe ranges of time are:\n12pm-14pm, 17pm-19pm, 21pm-23pm <br>\nNote: the data's timestamp is in UTC+0, and Chinese timezone is UTC+8</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 295786,
          "author_name": "asparuhhristov",
          "author_url": "",
          "post_date": "03/14/2018 07:27:44",
          "content": "<p>Thanks for the clarification. One thing is still a bit unclear though - what do you mean by \"more significant role\"...?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 295798,
          "author_name": "asparuhhristov",
          "author_url": "",
          "post_date": "03/14/2018 07:47:46",
          "content": "<p>I made the same analysis for the old test data - everything seems to be there (in the old one), the clicks per hour follow the same distribution as the one from the training data. Here is a snippet for the hours you provided:</p>\n\n<p>04    3344571</p>\n\n<p>05    2858427</p>\n\n<p>06    2621978</p>\n\n<p>07    2683972</p>\n\n<p>08    2878913</p>\n\n<p>09    2985213</p>\n\n<p>10    3127993</p>\n\n<p>11     3249395</p>\n\n<p>12    3068579</p>\n\n<p>13    3213032</p>\n\n<p>14    3261257</p>\n\n<p>15    2953455</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 295814,
          "author_name": "brandenkmurray",
          "author_url": "",
          "post_date": "03/14/2018 08:48:04",
          "content": "<p>If this is the case then I think the old test data should be made available to everyone. Those who have access to it may have an advantage knowing how many clicks certain <code>ips</code>, <code>apps</code>, etc. were clicked during those missing hours.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 295821,
          "author_name": "asparuhhristov",
          "author_url": "",
          "post_date": "03/14/2018 08:58:25",
          "content": "<p>Completely agree with you! <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51506\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51506</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 295822,
          "author_name": "mjahrer",
          "author_url": "",
          "post_date": "03/14/2018 09:02:40",
          "content": "<p>omg .. I'll quit this</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 295824,
          "author_name": "brandenkmurray",
          "author_url": "",
          "post_date": "03/14/2018 09:05:29",
          "content": "<p>Thanks Aspurah. I already had it, but good to see it's available to everyone. Should be on the data page, IMO.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 295998,
          "author_name": "aharless",
          "author_url": "",
          "post_date": "03/14/2018 15:14:21",
          "content": "<p>I believe the records from hours 6, 11, and 15 are all from the first second of the hour.  So it's essentially 4:00:00 to 6:00:00 inclusive, and so on.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 296106,
          "author_name": "anokas",
          "author_url": "",
          "post_date": "03/14/2018 18:11:38",
          "content": "<p>Agreed, please release the old data. Otherwise those who have it (like myself) have an unfair advantage.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 296347,
          "author_name": "onodera",
          "author_url": "",
          "post_date": "03/15/2018 04:24:35",
          "content": "<p>@anokas Here it is. <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51506\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51506</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 297739,
          "author_name": "tobycheese",
          "author_url": "",
          "post_date": "03/17/2018 21:10:40",
          "content": "<blockquote>\n  <p><strong>Aaron Yin wrote</strong></p>\n  \n  <blockquote>\n    <p>The reason why evaluation data is missing records by specific hours is that the original evaluation data is too large and kaggle's platform has a maximum size of evaluation data uploading, that's why we have to truncate the dataset by hours. Sorry for the inconvenience.   </p>\n  </blockquote>\n</blockquote>\n\n<p>Wouldn't it be the more reasonable option to have Kaggle adjust their platform limits instead and bring back the original test data? Just saying... it'll save everyone a headache and it'll make the competition results more valuable for you.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 297750,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "03/17/2018 21:55:24",
          "content": "<p>@Toby -</p>\n\n<p>It would require significant platform changes. Eventually, once we're 100% migrated from Azure to GCP, we'll have a lot more flexibility. Until then, it's not a viable option to rework functionality that will be replaced in the short-term anyway.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 299685,
          "author_name": "yuliagm",
          "author_url": "",
          "post_date": "03/21/2018 02:01:57",
          "content": "<p>Are we allowed to use old test data for feature engineering or is it against rules?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 295675,
      "author_name": "aharless",
      "author_url": "",
      "post_date": "03/14/2018 02:17:21",
      "content": "<p>I guess this means our validation data should be weighted by hour, or just exclude certain times (assuming one follows <a href=\"https://www.kaggle.com/konradb/validation-set/comments#latest-295404\">this kind of validation approach</a>)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 295788,
      "author_name": "asparuhhristov",
      "author_url": "",
      "post_date": "03/14/2018 07:28:51",
      "content": "<p>Thank you for sharing this insight! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 296110,
      "author_name": "wythhh",
      "author_url": "",
      "post_date": "03/14/2018 18:20:28",
      "content": "<p>Thanks for sharing! You are a nice person!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 296448,
      "author_name": "muhammadalfiansyah",
      "author_url": "",
      "post_date": "03/15/2018 08:19:54",
      "content": "<p>People should be more aware of this. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "295596": "I have counted number of records/clicks between 12:00 and 23:00 for each day from train and compared with test data. Interval 12:00-23:00 is a time range of test data (local Chinese time)\n\nResult is the following:\n\n* 2017-11-07: 34 308 844 - train\n* 2017-11-08: 36 475 438 - train\n* 2017-11-09: 37 169 180 - train\n* 2017-11-10: 18 790 469 - test\n\nIt looks like test data has only half of clicks comparing to  any day from train and the same time range as in test.\n\nI have a question to organizers, what is the reason for such a big difference? \nHave you removed half of clicks or this is natural decrease due to Friday evening or something else?",
    "295642": "If you check the hour of the timestamps in the test set:\n\n`test_df.click_time.str[11:13].value_counts().sort_index()`\n\n... the records are not missing at random, there are three distinct bands of time with records:\n\n    04    3344125\n    05    2858427\n    06        381\n    09    2984808\n    10    3127993\n    11        413\n    13    3212566\n    14    3261257\n    15        499\n    Name: click_time, dtype: int64\n\n(There are no rows at all with hours 07, 08, 12). So I'd say it's a biased sample of data, not a natural decrease in traffic. I would like to hear the organizers reasons too...",
    "295675": "I guess this means our validation data should be weighted by hour, or just exclude certain times (assuming one follows [this kind of validation approach][1])\n\n [1]: https://www.kaggle.com/konradb/validation-set/comments#latest-295404",
    "295739": "Perhaps organizers want to limit approaches based on time series.  It may be reasonable considering amount of data.",
    "295772": "The reason why evaluation data is missing records by specific hours is that the original evaluation data is too large and kaggle's platform has a maximum size of evaluation data uploading, that's why we have to truncate the dataset by hours. Sorry for the inconvenience.   \nWe choose three ranges of time, as we believed they would play an more significant role in Chinese daily life.  \nThe ranges of time are:\n12pm-14pm, 17pm-19pm, 21pm-23pm  \nNote: the data's timestamp is in UTC+0, and Chinese timezone is UTC+8",
    "295786": "Thanks for the clarification. One thing is still a bit unclear though - what do you mean by \"more significant role\"...?",
    "295788": "Thank you for sharing this insight!",
    "295798": "I made the same analysis for the old test data - everything seems to be there (in the old one), the clicks per hour follow the same distribution as the one from the training data. Here is a snippet for the hours you provided:\n\n04    3344571\n\n05    2858427\n\n06    2621978\n\n07    2683972\n\n08    2878913\n\n09    2985213\n\n10    3127993\n\n11     3249395\n\n12    3068579\n\n13    3213032\n\n14    3261257\n\n15    2953455",
    "295814": "If this is the case then I think the old test data should be made available to everyone. Those who have access to it may have an advantage knowing how many clicks certain `ips`, `apps`, etc. were clicked during those missing hours.",
    "295821": "Completely agree with you! https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51506",
    "295822": "omg .. I'll quit this",
    "295824": "Thanks Aspurah. I already had it, but good to see it's available to everyone. Should be on the data page, IMO.",
    "295998": "I believe the records from hours 6, 11, and 15 are all from the first second of the hour.  So it's essentially 4:00:00 to 6:00:00 inclusive, and so on.",
    "296106": "Agreed, please release the old data. Otherwise those who have it (like myself) have an unfair advantage.",
    "296110": "Thanks for sharing! You are a nice person!",
    "296347": "anokas Here it is. https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51506",
    "296448": "People should be more aware of this.",
    "297739": "&gt; **Aaron Yin wrote**\n&gt; \n&gt; &gt; The reason why evaluation data is missing records by specific hours is that the original evaluation data is too large and kaggle's platform has a maximum size of evaluation data uploading, that's why we have to truncate the dataset by hours. Sorry for the inconvenience.   \n\nWouldn't it be the more reasonable option to have Kaggle adjust their platform limits instead and bring back the original test data? Just saying... it'll save everyone a headache and it'll make the competition results more valuable for you.",
    "297750": "Toby -\n\nIt would require significant platform changes. Eventually, once we're 100% migrated from Azure to GCP, we'll have a lot more flexibility. Until then, it's not a viable option to rework functionality that will be replaced in the short-term anyway.",
    "299685": "Are we allowed to use old test data for feature engineering or is it against rules?"
  },
  "source": "meta"
}