{
  "id": 52373,
  "title": "Big Test / Small Test - Evaluating Next Steps",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/52373",
  "author_name": "",
  "post_date": "2018-03-19T16:58:29.208923100Z",
  "votes": 8,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Hi All - </p>\n\n<p>First of all, let me express sincere apologies for the confusion caused by an old (larger) version of the Test file getting re-copied to the data page. As has been mentioned previously, we're working on migrating our platform from Azure to GCP. On occasion, this causes some issues.</p>\n\n<p>I also want to acknowledge my slowness in addressing this issue. While it often takes time to figure out the best course of action, it would have been better to have communicated the status up-front. </p>\n\n<p>It might be a day or two as we assess the pros / cons of next steps and coordinate with the competition sponsor. Any course of action we take will have some downside, and we need to make sure we think through them to come up with the best option at this point. I will keep everyone posted.</p>\n\n<p>One thing that will be helpful in the future - If data is inadvertently made available, please do not take it upon yourself to repost it to another location (even if on Kaggle Datasets). It ends up creating more confusion and creates additional situations we need to think through to get everything back on track. </p>",
  "messages": [
    {
      "id": "298523",
      "postDate": "03/19/2018 16:58:29",
      "content": "<p>Hi All - </p>\n\n<p>First of all, let me express sincere apologies for the confusion caused by an old (larger) version of the Test file getting re-copied to the data page. As has been mentioned previously, we're working on migrating our platform from Azure to GCP. On occasion, this causes some issues.</p>\n\n<p>I also want to acknowledge my slowness in addressing this issue. While it often takes time to figure out the best course of action, it would have been better to have communicated the status up-front. </p>\n\n<p>It might be a day or two as we assess the pros / cons of next steps and coordinate with the competition sponsor. Any course of action we take will have some downside, and we need to make sure we think through them to come up with the best option at this point. I will keep everyone posted.</p>\n\n<p>One thing that will be helpful in the future - If data is inadvertently made available, please do not take it upon yourself to repost it to another location (even if on Kaggle Datasets). It ends up creating more confusion and creates additional situations we need to think through to get everything back on track. </p>",
      "rawMarkdown": "Hi All - \n\nFirst of all, let me express sincere apologies for the confusion caused by an old (larger) version of the Test file getting re-copied to the data page. As has been mentioned previously, we're working on migrating our platform from Azure to GCP. On occasion, this causes some issues.\n\nI also want to acknowledge my slowness in addressing this issue. While it often takes time to figure out the best course of action, it would have been better to have communicated the status up-front. \n\nIt might be a day or two as we assess the pros / cons of next steps and coordinate with the competition sponsor. Any course of action we take will have some downside, and we need to make sure we think through them to come up with the best option at this point. I will keep everyone posted.\n\nOne thing that will be helpful in the future - If data is inadvertently made available, please do not take it upon yourself to repost it to another location (even if on Kaggle Datasets). It ends up creating more confusion and creates additional situations we need to think through to get everything back on track.",
      "votes": null
    },
    {
      "id": "298715",
      "postDate": "03/19/2018 23:48:11",
      "content": "<p>Please act quickly, please please. \nI think you could re-post the old(big) test data and let kagglers use it, but the prediction file submitted to kaggle could just be the new(small) one because your platform couldn't handle the evaluation for the big one. This is internet era and people could always find a way to get access to old test data. Even you could forbidden people to use that, but there is always a way to get through.</p>",
      "rawMarkdown": "Please act quickly, please please. \nI think you could re-post the old(big) test data and let kagglers use it, but the prediction file submitted to kaggle could just be the new(small) one because your platform couldn't handle the evaluation for the big one. This is internet era and people could always find a way to get access to old test data. Even you could forbidden people to use that, but there is always a way to get through.",
      "votes": null
    },
    {
      "id": "299071",
      "postDate": "03/20/2018 12:35:55",
      "content": "<p>Thanks for the update - hopefully whatever is decided it will destroy all existing blends ;)</p>",
      "rawMarkdown": "Thanks for the update - hopefully whatever is decided it will destroy all existing blends ;)",
      "votes": null
    },
    {
      "id": "299081",
      "postDate": "03/20/2018 12:53:19",
      "content": "<p>I am bit confused: does that mean the data will be changed? I am utilizing the test set version downloaded on 27.02 - should I do sth about that? I would appreciate a 'talk to me like I'm a 5 year old' explanation :-)</p>",
      "rawMarkdown": "I am bit confused: does that mean the data will be changed? I am utilizing the test set version downloaded on 27.02 - should I do sth about that? I would appreciate a 'talk to me like I'm a 5 year old' explanation :-)",
      "votes": null
    },
    {
      "id": "299103",
      "postDate": "03/20/2018 13:34:42",
      "content": "<p>Kaggle posted a new version of test data two weeks ago and this one should be much smaller compared with the one you downloaded on 27.02.  Kaggle did this because the platform is not able to handle the evaluation on big test file.</p>",
      "rawMarkdown": "Kaggle posted a new version of test data two weeks ago and this one should be much smaller compared with the one you downloaded on 27.02.  Kaggle did this because the platform is not able to handle the evaluation on big test file.",
      "votes": null
    },
    {
      "id": "299105",
      "postDate": "03/20/2018 13:39:07",
      "content": "<p>Thanks a lot :-) presumably it's a subset of the 'big' one - the submissions I've been preparing keep getting correct evaluations, mostly inline with validation.</p>",
      "rawMarkdown": "Thanks a lot :-) presumably it's a subset of the 'big' one - the submissions I've been preparing keep getting correct evaluations, mostly inline with validation.",
      "votes": null
    },
    {
      "id": "299116",
      "postDate": "03/20/2018 14:01:55",
      "content": "<p>Hi Konrad -</p>\n\n<p>It depends . . . there was an issue with a mirroring script. The intended Test was uploaded, but then written over by a previous (pre-launch) version. You may have downloaded the correct files.</p>\n\n<p>The correct <code>test.csv</code> file should be:</p>\n\n<ul>\n<li>md5sum: 8f27a6d1b1f5bcd96c9183654863df98</li>\n<li>rows (including header): 18,790,470</li>\n<li>test.csv.zip should be ~160 Mb (not ~500 Mb)</li>\n</ul>",
      "rawMarkdown": "Hi Konrad -\n\nIt depends . . . there was an issue with a mirroring script. The intended Test was uploaded, but then written over by a previous (pre-launch) version. You may have downloaded the correct files.\n\nThe correct `test.csv` file should be:\n\n - md5sum: 8f27a6d1b1f5bcd96c9183654863df98\n - rows (including header): 18,790,470\n - test.csv.zip should be ~160 Mb (not ~500 Mb)",
      "votes": null
    },
    {
      "id": "299124",
      "postDate": "03/20/2018 14:14:42",
      "content": "<p>Hi, could I ask when could we get the information of next steps? </p>",
      "rawMarkdown": "Hi, could I ask when could we get the information of next steps?",
      "votes": null
    },
    {
      "id": "299161",
      "postDate": "03/20/2018 15:55:33",
      "content": "<p>@Inversion: thanks for the quick reaction and detailed info - it seems like I've got the right one after all.</p>",
      "rawMarkdown": "Inversion: thanks for the quick reaction and detailed info - it seems like I've got the right one after all.",
      "votes": null
    },
    {
      "id": "299166",
      "postDate": "03/20/2018 16:07:35",
      "content": "<p>@Snorlax: By tomorrow at the latest.</p>",
      "rawMarkdown": "Snorlax: By tomorrow at the latest.",
      "votes": null
    },
    {
      "id": "299880",
      "postDate": "03/21/2018 05:23:39",
      "content": "<p>The current test data is subsample of old test data. So we should use old test data for feature engineering. But the data tab only provides new test data. This might cause unfairness.</p>\n\n<p>If you(admin) want to reduce data size, my solution for this problem is</p>\n\n<ul>\n<li>provides train, test data with new date ranges, and renumber each category for preventing from using current data</li>\n</ul>",
      "rawMarkdown": "The current test data is subsample of old test data. So we should use old test data for feature engineering. But the data tab only provides new test data. This might cause unfairness.\n\n\nIf you(admin) want to reduce data size, my solution for this problem is\n\n - provides train, test data with new date ranges, and renumber each category for preventing from using current data",
      "votes": null
    },
    {
      "id": "300516",
      "postDate": "03/21/2018 17:02:07",
      "content": "<p>I completely agree with you. \nIf admins don't re-post the old test data and forbid people to use it, they need to check everyone's code to confirm he/she doesn't use the old test data, which is clearly impossible and nonsense.</p>",
      "rawMarkdown": "I completely agree with you. \nIf admins don't re-post the old test data and forbid people to use it, they need to check everyone's code to confirm he/she doesn't use the old test data, which is clearly impossible and nonsense.",
      "votes": null
    },
    {
      "id": "300702",
      "postDate": "03/21/2018 20:44:27",
      "content": "<p>I waited some comment in this thread regarding old test data, It is officially <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52658\">released</a>. </p>\n\n<p>File is the same as old test</p>\n\n<p><code>&gt;md5sum old_test.csv\n8898d8cfa6a62de21d4e34e4a72a8889 *old_test.csv</code></p>\n\n<p><code>&gt;md5sum test_supplement.csv\n8898d8cfa6a62de21d4e34e4a72a8889 *test_supplement.csv</code></p>",
      "rawMarkdown": "I waited some comment in this thread regarding old test data, It is officially [released][1]. \n\nFile is the same as old test\n\n```&gt;md5sum old_test.csv\n8898d8cfa6a62de21d4e34e4a72a8889 *old_test.csv```\n\n```&gt;md5sum test_supplement.csv\n8898d8cfa6a62de21d4e34e4a72a8889 *test_supplement.csv```\n\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52658",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 298715,
      "author_name": "wythhh",
      "author_url": "",
      "post_date": "03/19/2018 23:48:11",
      "content": "<p>Please act quickly, please please. \nI think you could re-post the old(big) test data and let kagglers use it, but the prediction file submitted to kaggle could just be the new(small) one because your platform couldn't handle the evaluation for the big one. This is internet era and people could always find a way to get access to old test data. Even you could forbidden people to use that, but there is always a way to get through.</p>",
      "votes": null,
      "replies": [
        {
          "id": 300516,
          "author_name": "mamasinkgs",
          "author_url": "",
          "post_date": "03/21/2018 17:02:07",
          "content": "<p>I completely agree with you. \nIf admins don't re-post the old test data and forbid people to use it, they need to check everyone's code to confirm he/she doesn't use the old test data, which is clearly impossible and nonsense.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 299071,
      "author_name": "scirpus",
      "author_url": "",
      "post_date": "03/20/2018 12:35:55",
      "content": "<p>Thanks for the update - hopefully whatever is decided it will destroy all existing blends ;)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 299081,
      "author_name": "konradb",
      "author_url": "",
      "post_date": "03/20/2018 12:53:19",
      "content": "<p>I am bit confused: does that mean the data will be changed? I am utilizing the test set version downloaded on 27.02 - should I do sth about that? I would appreciate a 'talk to me like I'm a 5 year old' explanation :-)</p>",
      "votes": null,
      "replies": [
        {
          "id": 299103,
          "author_name": "wythhh",
          "author_url": "",
          "post_date": "03/20/2018 13:34:42",
          "content": "<p>Kaggle posted a new version of test data two weeks ago and this one should be much smaller compared with the one you downloaded on 27.02.  Kaggle did this because the platform is not able to handle the evaluation on big test file.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 299105,
          "author_name": "konradb",
          "author_url": "",
          "post_date": "03/20/2018 13:39:07",
          "content": "<p>Thanks a lot :-) presumably it's a subset of the 'big' one - the submissions I've been preparing keep getting correct evaluations, mostly inline with validation.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 299116,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "03/20/2018 14:01:55",
          "content": "<p>Hi Konrad -</p>\n\n<p>It depends . . . there was an issue with a mirroring script. The intended Test was uploaded, but then written over by a previous (pre-launch) version. You may have downloaded the correct files.</p>\n\n<p>The correct <code>test.csv</code> file should be:</p>\n\n<ul>\n<li>md5sum: 8f27a6d1b1f5bcd96c9183654863df98</li>\n<li>rows (including header): 18,790,470</li>\n<li>test.csv.zip should be ~160 Mb (not ~500 Mb)</li>\n</ul>",
          "votes": null,
          "replies": []
        },
        {
          "id": 299124,
          "author_name": "wythhh",
          "author_url": "",
          "post_date": "03/20/2018 14:14:42",
          "content": "<p>Hi, could I ask when could we get the information of next steps? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 299161,
          "author_name": "konradb",
          "author_url": "",
          "post_date": "03/20/2018 15:55:33",
          "content": "<p>@Inversion: thanks for the quick reaction and detailed info - it seems like I've got the right one after all.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 299166,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "03/20/2018 16:07:35",
          "content": "<p>@Snorlax: By tomorrow at the latest.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 299880,
      "author_name": "onodera",
      "author_url": "",
      "post_date": "03/21/2018 05:23:39",
      "content": "<p>The current test data is subsample of old test data. So we should use old test data for feature engineering. But the data tab only provides new test data. This might cause unfairness.</p>\n\n<p>If you(admin) want to reduce data size, my solution for this problem is</p>\n\n<ul>\n<li>provides train, test data with new date ranges, and renumber each category for preventing from using current data</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 300702,
      "author_name": "alexfir",
      "author_url": "",
      "post_date": "03/21/2018 20:44:27",
      "content": "<p>I waited some comment in this thread regarding old test data, It is officially <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52658\">released</a>. </p>\n\n<p>File is the same as old test</p>\n\n<p><code>&gt;md5sum old_test.csv\n8898d8cfa6a62de21d4e34e4a72a8889 *old_test.csv</code></p>\n\n<p><code>&gt;md5sum test_supplement.csv\n8898d8cfa6a62de21d4e34e4a72a8889 *test_supplement.csv</code></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "298523": "Hi All - \n\nFirst of all, let me express sincere apologies for the confusion caused by an old (larger) version of the Test file getting re-copied to the data page. As has been mentioned previously, we're working on migrating our platform from Azure to GCP. On occasion, this causes some issues.\n\nI also want to acknowledge my slowness in addressing this issue. While it often takes time to figure out the best course of action, it would have been better to have communicated the status up-front. \n\nIt might be a day or two as we assess the pros / cons of next steps and coordinate with the competition sponsor. Any course of action we take will have some downside, and we need to make sure we think through them to come up with the best option at this point. I will keep everyone posted.\n\nOne thing that will be helpful in the future - If data is inadvertently made available, please do not take it upon yourself to repost it to another location (even if on Kaggle Datasets). It ends up creating more confusion and creates additional situations we need to think through to get everything back on track.",
    "298715": "Please act quickly, please please. \nI think you could re-post the old(big) test data and let kagglers use it, but the prediction file submitted to kaggle could just be the new(small) one because your platform couldn't handle the evaluation for the big one. This is internet era and people could always find a way to get access to old test data. Even you could forbidden people to use that, but there is always a way to get through.",
    "299071": "Thanks for the update - hopefully whatever is decided it will destroy all existing blends ;)",
    "299081": "I am bit confused: does that mean the data will be changed? I am utilizing the test set version downloaded on 27.02 - should I do sth about that? I would appreciate a 'talk to me like I'm a 5 year old' explanation :-)",
    "299103": "Kaggle posted a new version of test data two weeks ago and this one should be much smaller compared with the one you downloaded on 27.02.  Kaggle did this because the platform is not able to handle the evaluation on big test file.",
    "299105": "Thanks a lot :-) presumably it's a subset of the 'big' one - the submissions I've been preparing keep getting correct evaluations, mostly inline with validation.",
    "299116": "Hi Konrad -\n\nIt depends . . . there was an issue with a mirroring script. The intended Test was uploaded, but then written over by a previous (pre-launch) version. You may have downloaded the correct files.\n\nThe correct `test.csv` file should be:\n\n - md5sum: 8f27a6d1b1f5bcd96c9183654863df98\n - rows (including header): 18,790,470\n - test.csv.zip should be ~160 Mb (not ~500 Mb)",
    "299124": "Hi, could I ask when could we get the information of next steps?",
    "299161": "Inversion: thanks for the quick reaction and detailed info - it seems like I've got the right one after all.",
    "299166": "Snorlax: By tomorrow at the latest.",
    "299880": "The current test data is subsample of old test data. So we should use old test data for feature engineering. But the data tab only provides new test data. This might cause unfairness.\n\n\nIf you(admin) want to reduce data size, my solution for this problem is\n\n - provides train, test data with new date ranges, and renumber each category for preventing from using current data",
    "300516": "I completely agree with you. \nIf admins don't re-post the old test data and forbid people to use it, they need to check everyone's code to confirm he/she doesn't use the old test data, which is clearly impossible and nonsense.",
    "300702": "I waited some comment in this thread regarding old test data, It is officially [released][1]. \n\nFile is the same as old test\n\n```&gt;md5sum old_test.csv\n8898d8cfa6a62de21d4e34e4a72a8889 *old_test.csv```\n\n```&gt;md5sum test_supplement.csv\n8898d8cfa6a62de21d4e34e4a72a8889 *test_supplement.csv```\n\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52658"
  },
  "source": "meta"
}