{
  "id": 51210,
  "title": "Please Check / Re-download test.csv",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/51210",
  "author_name": "",
  "post_date": "2018-03-06T16:10:54.325808900Z",
  "votes": 10,
  "comment_count": 12,
  "views": 0,
  "content": "<p>There was an issue with mirroring that caused an older (and larger) test file to download from the data page. The issue is fixed. If you've already downloaded <code>test.csv</code> before seeing this, please re-download the file. </p>\n\n<p>The <code>test.csv</code> file should be:</p>\n\n<ul>\n<li>md5sum: 8f27a6d1b1f5bcd96c9183654863df98</li>\n<li>rows (including header): 18,790,470</li>\n<li>test.csv.zip should be ~160 Mb (not ~500 Mb)</li>\n</ul>\n\n<p>Note: The file in kernels is unaffected.</p>",
  "messages": [
    {
      "id": "291668",
      "postDate": "03/06/2018 16:10:54",
      "content": "<p>There was an issue with mirroring that caused an older (and larger) test file to download from the data page. The issue is fixed. If you've already downloaded <code>test.csv</code> before seeing this, please re-download the file. </p>\n\n<p>The <code>test.csv</code> file should be:</p>\n\n<ul>\n<li>md5sum: 8f27a6d1b1f5bcd96c9183654863df98</li>\n<li>rows (including header): 18,790,470</li>\n<li>test.csv.zip should be ~160 Mb (not ~500 Mb)</li>\n</ul>\n\n<p>Note: The file in kernels is unaffected.</p>",
      "rawMarkdown": "There was an issue with mirroring that caused an older (and larger) test file to download from the data page. The issue is fixed. If you've already downloaded `test.csv` before seeing this, please re-download the file. \n\nThe `test.csv` file should be:\n\n - md5sum: 8f27a6d1b1f5bcd96c9183654863df98\n - rows (including header): 18,790,470\n - test.csv.zip should be ~160 Mb (not ~500 Mb)\n\nNote: The file in kernels is unaffected.",
      "votes": null
    },
    {
      "id": "292487",
      "postDate": "03/08/2018 02:42:43",
      "content": "<p>I get \"The kernel appears to have died. It will restart automatically \" repeatedly when I try to load train data. Is this expected behavior?</p>",
      "rawMarkdown": "I get \"The kernel appears to have died. It will restart automatically \" repeatedly when I try to load train data. Is this expected behavior?",
      "votes": null
    },
    {
      "id": "292999",
      "postDate": "03/09/2018 02:39:08",
      "content": "<p>Hi,</p>\n\n<p>I don't see attributed_time column in the test.csv dataset. As per the data description, this column should have been present. Can you please clarify?</p>",
      "rawMarkdown": "Hi,\n\nI don't see attributed_time column in the test.csv dataset. As per the data description, this column should have been present. Can you please clarify?",
      "votes": null
    },
    {
      "id": "293231",
      "postDate": "03/09/2018 13:22:53",
      "content": "<p>For anyone else with the same question, the attributed_time is not useful because the aim is click-through-rate prediction rather than click-fraud detection. Thus, the test set doesn't have the attributed_time column (as this would be a give-away of whether the app was installed or not).</p>",
      "rawMarkdown": "For anyone else with the same question, the attributed_time is not useful because the aim is click-through-rate prediction rather than click-fraud detection. Thus, the test set doesn't have the attributed_time column (as this would be a give-away of whether the app was installed or not).",
      "votes": null
    },
    {
      "id": "293258",
      "postDate": "03/09/2018 14:38:55",
      "content": "<p>Can we have the old csv too?\nI assume the old csv has relevant information, and it would be an unfair advantage to those who couldn't download the old one in time.\nOr is it just something like the file format messed up?</p>",
      "rawMarkdown": "Can we have the old csv too?\nI assume the old csv has relevant information, and it would be an unfair advantage to those who couldn't download the old one in time.\nOr is it just something like the file format messed up?",
      "votes": null
    },
    {
      "id": "293327",
      "postDate": "03/09/2018 16:35:23",
      "content": "<p>sample_submission.csv.zip still has  57537506 rows.</p>",
      "rawMarkdown": "sample_submission.csv.zip still has  57537506 rows.",
      "votes": null
    },
    {
      "id": "293362",
      "postDate": "03/09/2018 17:56:46",
      "content": "<p>Can you do a browser refresh and let me know if you still see the old file? Thx.</p>",
      "rawMarkdown": "Can you do a browser refresh and let me know if you still see the old file? Thx.",
      "votes": null
    },
    {
      "id": "293364",
      "postDate": "03/09/2018 18:06:02",
      "content": "<p>Thx. It was fixed. Maybe I forgot to refresh the page or fixed about an hour ago. <br>\n(time zone is JST (UTC+9) below.)</p>\n\n<pre><code>-rwxrwxrwx 1 root root 128924534 Mar  6 07:12 sample_submission.csv (1).zip\n-rwxrwxrwx 1 root root 128924534 Mar 10 01:32 sample_submission.csv (2).zip\n-rwxrwxrwx 1 root root 128924534 Mar 10 01:33 sample_submission.csv (3).zip\n-rwxrwxrwx 1 root root  42058441 Mar 10 02:57 sample_submission.csv (4).zip\n</code></pre>",
      "rawMarkdown": "Thx. It was fixed. Maybe I forgot to refresh the page or fixed about an hour ago.   \n(time zone is JST (UTC+9) below.)\n\n    -rwxrwxrwx 1 root root 128924534 Mar  6 07:12 sample_submission.csv (1).zip\n    -rwxrwxrwx 1 root root 128924534 Mar 10 01:32 sample_submission.csv (2).zip\n    -rwxrwxrwx 1 root root 128924534 Mar 10 01:33 sample_submission.csv (3).zip\n    -rwxrwxrwx 1 root root  42058441 Mar 10 02:57 sample_submission.csv (4).zip",
      "votes": null
    },
    {
      "id": "293367",
      "postDate": "03/09/2018 18:13:24",
      "content": "<p>I uploaded it as <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51506.\">a Kaggle dataset</a>, but I don't know it's okay or not. I hope Kaggle team will mention it.</p>",
      "rawMarkdown": "I uploaded it as [a Kaggle dataset][1], but I don't know it's okay or not. I hope Kaggle team will mention it.\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51506.",
      "votes": null
    },
    {
      "id": "293494",
      "postDate": "03/10/2018 01:34:03",
      "content": "<p>Wait, I actually had the older file... maybe my advantage slipped away :) <br>\nAnyways, thank you tkm-san. I hope this levels the field.</p>",
      "rawMarkdown": "Wait, I actually had the older file... maybe my advantage slipped away :)  \nAnyways, thank you tkm-san. I hope this levels the field.",
      "votes": null
    },
    {
      "id": "298312",
      "postDate": "03/19/2018 10:30:46",
      "content": "<p>May we use the old test data file for semi-supervised learning, or this is forbidden?</p>",
      "rawMarkdown": "May we use the old test data file for semi-supervised learning, or this is forbidden?",
      "votes": null
    },
    {
      "id": "298474",
      "postDate": "03/19/2018 15:59:25",
      "content": "<p>Hi inversion,</p>\n\n<p>Can you clarify that we can officially use the original test set if we want to? </p>",
      "rawMarkdown": "Hi inversion,\n\nCan you clarify that we can officially use the original test set if we want to?",
      "votes": null
    },
    {
      "id": "312014",
      "postDate": "04/11/2018 03:49:20",
      "content": "<p>hi, could you please check this post ? There is always error after I uploaded my submission file. I have not find any errors before I uploaded. \n<a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/54160#311675\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/54160#311675</a>\nThanks</p>",
      "rawMarkdown": "hi, could you please check this post ? There is always error after I uploaded my submission file. I have not find any errors before I uploaded. \nhttps://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/54160#311675\nThanks",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 292487,
      "author_name": "bgopalakrishnan",
      "author_url": "",
      "post_date": "03/08/2018 02:42:43",
      "content": "<p>I get \"The kernel appears to have died. It will restart automatically \" repeatedly when I try to load train data. Is this expected behavior?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 292999,
      "author_name": "unstructured",
      "author_url": "",
      "post_date": "03/09/2018 02:39:08",
      "content": "<p>Hi,</p>\n\n<p>I don't see attributed_time column in the test.csv dataset. As per the data description, this column should have been present. Can you please clarify?</p>",
      "votes": null,
      "replies": [
        {
          "id": 293231,
          "author_name": "unstructured",
          "author_url": "",
          "post_date": "03/09/2018 13:22:53",
          "content": "<p>For anyone else with the same question, the attributed_time is not useful because the aim is click-through-rate prediction rather than click-fraud detection. Thus, the test set doesn't have the attributed_time column (as this would be a give-away of whether the app was installed or not).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 293258,
      "author_name": "pocketsuteado",
      "author_url": "",
      "post_date": "03/09/2018 14:38:55",
      "content": "<p>Can we have the old csv too?\nI assume the old csv has relevant information, and it would be an unfair advantage to those who couldn't download the old one in time.\nOr is it just something like the file format messed up?</p>",
      "votes": null,
      "replies": [
        {
          "id": 293367,
          "author_name": "tkm2261",
          "author_url": "",
          "post_date": "03/09/2018 18:13:24",
          "content": "<p>I uploaded it as <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51506.\">a Kaggle dataset</a>, but I don't know it's okay or not. I hope Kaggle team will mention it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 293494,
          "author_name": "pocketsuteado",
          "author_url": "",
          "post_date": "03/10/2018 01:34:03",
          "content": "<p>Wait, I actually had the older file... maybe my advantage slipped away :) <br>\nAnyways, thank you tkm-san. I hope this levels the field.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 293327,
      "author_name": "tkm2261",
      "author_url": "",
      "post_date": "03/09/2018 16:35:23",
      "content": "<p>sample_submission.csv.zip still has  57537506 rows.</p>",
      "votes": null,
      "replies": [
        {
          "id": 293362,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "03/09/2018 17:56:46",
          "content": "<p>Can you do a browser refresh and let me know if you still see the old file? Thx.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 293364,
          "author_name": "tkm2261",
          "author_url": "",
          "post_date": "03/09/2018 18:06:02",
          "content": "<p>Thx. It was fixed. Maybe I forgot to refresh the page or fixed about an hour ago. <br>\n(time zone is JST (UTC+9) below.)</p>\n\n<pre><code>-rwxrwxrwx 1 root root 128924534 Mar  6 07:12 sample_submission.csv (1).zip\n-rwxrwxrwx 1 root root 128924534 Mar 10 01:32 sample_submission.csv (2).zip\n-rwxrwxrwx 1 root root 128924534 Mar 10 01:33 sample_submission.csv (3).zip\n-rwxrwxrwx 1 root root  42058441 Mar 10 02:57 sample_submission.csv (4).zip\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 298312,
          "author_name": "nedelko",
          "author_url": "",
          "post_date": "03/19/2018 10:30:46",
          "content": "<p>May we use the old test data file for semi-supervised learning, or this is forbidden?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 298474,
      "author_name": "scirpus",
      "author_url": "",
      "post_date": "03/19/2018 15:59:25",
      "content": "<p>Hi inversion,</p>\n\n<p>Can you clarify that we can officially use the original test set if we want to? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 312014,
      "author_name": "",
      "author_url": "",
      "post_date": "04/11/2018 03:49:20",
      "content": "<p>hi, could you please check this post ? There is always error after I uploaded my submission file. I have not find any errors before I uploaded. \n<a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/54160#311675\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/54160#311675</a>\nThanks</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "291668": "There was an issue with mirroring that caused an older (and larger) test file to download from the data page. The issue is fixed. If you've already downloaded `test.csv` before seeing this, please re-download the file. \n\nThe `test.csv` file should be:\n\n - md5sum: 8f27a6d1b1f5bcd96c9183654863df98\n - rows (including header): 18,790,470\n - test.csv.zip should be ~160 Mb (not ~500 Mb)\n\nNote: The file in kernels is unaffected.",
    "292487": "I get \"The kernel appears to have died. It will restart automatically \" repeatedly when I try to load train data. Is this expected behavior?",
    "292999": "Hi,\n\nI don't see attributed_time column in the test.csv dataset. As per the data description, this column should have been present. Can you please clarify?",
    "293231": "For anyone else with the same question, the attributed_time is not useful because the aim is click-through-rate prediction rather than click-fraud detection. Thus, the test set doesn't have the attributed_time column (as this would be a give-away of whether the app was installed or not).",
    "293258": "Can we have the old csv too?\nI assume the old csv has relevant information, and it would be an unfair advantage to those who couldn't download the old one in time.\nOr is it just something like the file format messed up?",
    "293327": "sample_submission.csv.zip still has  57537506 rows.",
    "293362": "Can you do a browser refresh and let me know if you still see the old file? Thx.",
    "293364": "Thx. It was fixed. Maybe I forgot to refresh the page or fixed about an hour ago.   \n(time zone is JST (UTC+9) below.)\n\n    -rwxrwxrwx 1 root root 128924534 Mar  6 07:12 sample_submission.csv (1).zip\n    -rwxrwxrwx 1 root root 128924534 Mar 10 01:32 sample_submission.csv (2).zip\n    -rwxrwxrwx 1 root root 128924534 Mar 10 01:33 sample_submission.csv (3).zip\n    -rwxrwxrwx 1 root root  42058441 Mar 10 02:57 sample_submission.csv (4).zip",
    "293367": "I uploaded it as [a Kaggle dataset][1], but I don't know it's okay or not. I hope Kaggle team will mention it.\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51506.",
    "293494": "Wait, I actually had the older file... maybe my advantage slipped away :)  \nAnyways, thank you tkm-san. I hope this levels the field.",
    "298312": "May we use the old test data file for semi-supervised learning, or this is forbidden?",
    "298474": "Hi inversion,\n\nCan you clarify that we can officially use the original test set if we want to?",
    "312014": "hi, could you please check this post ? There is always error after I uploaded my submission file. I have not find any errors before I uploaded. \nhttps://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/54160#311675\nThanks"
  },
  "source": "meta"
}