{
  "id": 51168,
  "title": "Unusual train.csv.zip structure",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/51168",
  "author_name": "",
  "post_date": "2018-03-06T03:15:37.172796300Z",
  "votes": 3,
  "comment_count": 18,
  "views": 0,
  "content": "<p>Given some issues with past competitions, I want to be extra careful and wanted to confirm the train is as expected as this is the first time I see a folder structure like this. \n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/291388/8691/zipContent.jpeg\" alt=\"enter image description here\"></p>\n\n<p>Thanks.</p>",
  "messages": [
    {
      "id": "291388",
      "postDate": "03/06/2018 03:15:37",
      "content": "<p>Given some issues with past competitions, I want to be extra careful and wanted to confirm the train is as expected as this is the first time I see a folder structure like this. \n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/291388/8691/zipContent.jpeg\" alt=\"enter image description here\"></p>\n\n<p>Thanks.</p>",
      "rawMarkdown": "Given some issues with past competitions, I want to be extra careful and wanted to confirm the train is as expected as this is the first time I see a folder structure like this. \n![enter image description here][1]\n\n\nThanks.\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/291388/8691/zipContent.jpeg",
      "votes": null
    },
    {
      "id": "291389",
      "postDate": "03/06/2018 03:17:56",
      "content": "<p>Can you re-download and check again? (Regardless, the file should be fine. zip just got excited during the compression.)</p>",
      "rawMarkdown": "Can you re-download and check again? (Regardless, the file should be fine. zip just got excited during the compression.)",
      "votes": null
    },
    {
      "id": "291393",
      "postDate": "03/06/2018 03:22:53",
      "content": "<p>I am lazy to re-download, but I trust you, if this is it, it is.</p>\n\n<p>For reference (md5sum): <br>\n8898d8cfa6a62de21d4e34e4a72a8889  test.csv <br>\n36da1e8fec8d6765e56d894096a1462e  train.csv</p>",
      "rawMarkdown": "I am lazy to re-download, but I trust you, if this is it, it is.\n\nFor reference (md5sum):  \n8898d8cfa6a62de21d4e34e4a72a8889  test.csv  \n36da1e8fec8d6765e56d894096a1462e  train.csv",
      "votes": null
    },
    {
      "id": "291398",
      "postDate": "03/06/2018 03:41:20",
      "content": "<p>Please re-download test.csv. (You're good to go with Train)</p>\n\n<p>Should be: 8f27a6d1b1f5bcd96c9183654863df98  test.csv</p>",
      "rawMarkdown": "Please re-download test.csv. (You're good to go with Train)\n\nShould be: 8f27a6d1b1f5bcd96c9183654863df98  test.csv",
      "votes": null
    },
    {
      "id": "291407",
      "postDate": "03/06/2018 04:08:48",
      "content": "<p>I cant match the md5sum(test.csv), I just downloaded it again, it still has the full-path structure but I guess that's OK as train matched. </p>\n\n<p>Unless otherwise notified, I will guess I am just being paranoic here and continue.</p>",
      "rawMarkdown": "I cant match the md5sum(test.csv), I just downloaded it again, it still has the full-path structure but I guess that's OK as train matched. \n\nUnless otherwise notified, I will guess I am just being paranoic here and continue.",
      "votes": null
    },
    {
      "id": "291479",
      "postDate": "03/06/2018 07:59:02",
      "content": "<p>NB The test dataset on the download page is still an old version with 50M samples - until kaggle wake up the only option is to download it from kernels: <a href=\"https://www.kaggle.com/anokas/getting-fixed-test-data\">https://www.kaggle.com/anokas/getting-fixed-test-data</a></p>",
      "rawMarkdown": "NB The test dataset on the download page is still an old version with 50M samples - until kaggle wake up the only option is to download it from kernels: https://www.kaggle.com/anokas/getting-fixed-test-data",
      "votes": null
    },
    {
      "id": "291506",
      "postDate": "03/06/2018 08:54:36",
      "content": "<p>Thank you for the kind notice.</p>",
      "rawMarkdown": "Thank you for the kind notice.",
      "votes": null
    },
    {
      "id": "291670",
      "postDate": "03/06/2018 16:13:48",
      "content": "<p>Please see this thread:</p>\n\n<p><a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51210\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51210</a></p>",
      "rawMarkdown": "Please see this thread:\n\nhttps://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51210",
      "votes": null
    },
    {
      "id": "295254",
      "postDate": "03/13/2018 10:56:35",
      "content": "<p>Didn't the last talkingdata competition have leakage? Judging from the scores, are we sure this one doesn't too?</p>",
      "rawMarkdown": "Didn't the last talkingdata competition have leakage? Judging from the scores, are we sure this one doesn't too?",
      "votes": null
    },
    {
      "id": "295258",
      "postDate": "03/13/2018 11:12:37",
      "content": "<p>What do you mean by judging from the scores?</p>",
      "rawMarkdown": "What do you mean by judging from the scores?",
      "votes": null
    },
    {
      "id": "295318",
      "postDate": "03/13/2018 13:27:38",
      "content": "<p>Very high AUC</p>",
      "rawMarkdown": "Very high AUC",
      "votes": null
    },
    {
      "id": "295320",
      "postDate": "03/13/2018 13:35:56",
      "content": "<p>Haven't checked in that much detail yet, sorting the scores on IP gives an AUC ~ 0.7 =&gt; it does smell funny indeed...</p>",
      "rawMarkdown": "Haven't checked in that much detail yet, sorting the scores on IP gives an AUC ~ 0.7 =&gt; it does smell funny indeed...",
      "votes": null
    },
    {
      "id": "295324",
      "postDate": "03/13/2018 13:41:36",
      "content": "<p>I don't think that indicates a leak IMO. It makes logical sense that an IP address (user) who is more likely to download an app in the training set is more likely to download an app a few days later (in the test set) - some people download every app they see while others never download anything. </p>\n\n<p>Same with apps - certain apps will have higher likelihoods of being downloaded (they may be more appealing etc.) - and this is why you can get 0.9+ just by ordering the apps by their conversion rate in the training set.</p>",
      "rawMarkdown": "I don't think that indicates a leak IMO. It makes logical sense that an IP address (user) who is more likely to download an app in the training set is more likely to download an app a few days later (in the test set) - some people download every app they see while others never download anything. \n\nSame with apps - certain apps will have higher likelihoods of being downloaded (they may be more appealing etc.) - and this is why you can get 0.9+ just by ordering the apps by their conversion rate in the training set.",
      "votes": null
    },
    {
      "id": "295326",
      "postDate": "03/13/2018 13:49:50",
      "content": "<p>In principle I agree, except there have been quite a few competitions where sorting by that sort of thing lead to massive jumps (which under random encoding shouldn't be possible). But I haven't done a rigorous analysis, so it's pure speculation at this point.</p>",
      "rawMarkdown": "In principle I agree, except there have been quite a few competitions where sorting by that sort of thing lead to massive jumps (which under random encoding shouldn't be possible). But I haven't done a rigorous analysis, so it's pure speculation at this point.",
      "votes": null
    },
    {
      "id": "295328",
      "postDate": "03/13/2018 13:53:59",
      "content": "<p>Oh, apologies. I thought you meant sorting the IPs based on their target values in the training set. I agree that just sorting using the random encoding provided shouldn't give &gt;0.5.</p>\n\n<p>However, I don't think the high AUC <em>on it's own</em> is an indicator of a leak. Some problems are just easy to predict.</p>",
      "rawMarkdown": "Oh, apologies. I thought you meant sorting the IPs based on their target values in the training set. I agree that just sorting using the random encoding provided shouldn't give &gt;0.5.\n\nHowever, I don't think the high AUC _on it's own_ is an indicator of a leak. Some problems are just easy to predict.",
      "votes": null
    },
    {
      "id": "295330",
      "postDate": "03/13/2018 13:58:31",
      "content": "<p>No problem - always good to bounce one's thoughts off somebody else.</p>",
      "rawMarkdown": "No problem - always good to bounce one's thoughts off somebody else.",
      "votes": null
    },
    {
      "id": "295347",
      "postDate": "03/13/2018 14:29:50",
      "content": "<p>and there was a similar ID/ sorting leak in the last competition, if I remember correctly</p>",
      "rawMarkdown": "and there was a similar ID/ sorting leak in the last competition, if I remember correctly",
      "votes": null
    },
    {
      "id": "295348",
      "postDate": "03/13/2018 14:31:07",
      "content": "<p>and if I was the company and could easily get 96 AUC I wouldn't pay $$ to outsource the problem</p>",
      "rawMarkdown": "and if I was the company and could easily get 96 AUC I wouldn't pay $$ to outsource the problem",
      "votes": null
    },
    {
      "id": "295356",
      "postDate": "03/13/2018 14:50:27",
      "content": "<p>Well, that's a more general phenomenon, case in point: the Toxic Comments contest. Given a NB-SVM benchmark scoring ~ .982 AUC while running on a laptop, would any business care about an advanced deep learning model achieving .988 on a cluster?</p>",
      "rawMarkdown": "Well, that's a more general phenomenon, case in point: the Toxic Comments contest. Given a NB-SVM benchmark scoring ~ .982 AUC while running on a laptop, would any business care about an advanced deep learning model achieving .988 on a cluster?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 291389,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "03/06/2018 03:17:56",
      "content": "<p>Can you re-download and check again? (Regardless, the file should be fine. zip just got excited during the compression.)</p>",
      "votes": null,
      "replies": [
        {
          "id": 291393,
          "author_name": "carloshuertas",
          "author_url": "",
          "post_date": "03/06/2018 03:22:53",
          "content": "<p>I am lazy to re-download, but I trust you, if this is it, it is.</p>\n\n<p>For reference (md5sum): <br>\n8898d8cfa6a62de21d4e34e4a72a8889  test.csv <br>\n36da1e8fec8d6765e56d894096a1462e  train.csv</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 291398,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "03/06/2018 03:41:20",
          "content": "<p>Please re-download test.csv. (You're good to go with Train)</p>\n\n<p>Should be: 8f27a6d1b1f5bcd96c9183654863df98  test.csv</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 291407,
          "author_name": "carloshuertas",
          "author_url": "",
          "post_date": "03/06/2018 04:08:48",
          "content": "<p>I cant match the md5sum(test.csv), I just downloaded it again, it still has the full-path structure but I guess that's OK as train matched. </p>\n\n<p>Unless otherwise notified, I will guess I am just being paranoic here and continue.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 291670,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "03/06/2018 16:13:48",
          "content": "<p>Please see this thread:</p>\n\n<p><a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51210\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51210</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 291479,
      "author_name": "anokas",
      "author_url": "",
      "post_date": "03/06/2018 07:59:02",
      "content": "<p>NB The test dataset on the download page is still an old version with 50M samples - until kaggle wake up the only option is to download it from kernels: <a href=\"https://www.kaggle.com/anokas/getting-fixed-test-data\">https://www.kaggle.com/anokas/getting-fixed-test-data</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 291506,
          "author_name": "wenpengwei",
          "author_url": "",
          "post_date": "03/06/2018 08:54:36",
          "content": "<p>Thank you for the kind notice.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 295254,
      "author_name": "domcastro",
      "author_url": "",
      "post_date": "03/13/2018 10:56:35",
      "content": "<p>Didn't the last talkingdata competition have leakage? Judging from the scores, are we sure this one doesn't too?</p>",
      "votes": null,
      "replies": [
        {
          "id": 295258,
          "author_name": "anokas",
          "author_url": "",
          "post_date": "03/13/2018 11:12:37",
          "content": "<p>What do you mean by judging from the scores?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 295318,
          "author_name": "domcastro",
          "author_url": "",
          "post_date": "03/13/2018 13:27:38",
          "content": "<p>Very high AUC</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 295320,
          "author_name": "konradb",
          "author_url": "",
          "post_date": "03/13/2018 13:35:56",
          "content": "<p>Haven't checked in that much detail yet, sorting the scores on IP gives an AUC ~ 0.7 =&gt; it does smell funny indeed...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 295324,
          "author_name": "anokas",
          "author_url": "",
          "post_date": "03/13/2018 13:41:36",
          "content": "<p>I don't think that indicates a leak IMO. It makes logical sense that an IP address (user) who is more likely to download an app in the training set is more likely to download an app a few days later (in the test set) - some people download every app they see while others never download anything. </p>\n\n<p>Same with apps - certain apps will have higher likelihoods of being downloaded (they may be more appealing etc.) - and this is why you can get 0.9+ just by ordering the apps by their conversion rate in the training set.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 295326,
          "author_name": "konradb",
          "author_url": "",
          "post_date": "03/13/2018 13:49:50",
          "content": "<p>In principle I agree, except there have been quite a few competitions where sorting by that sort of thing lead to massive jumps (which under random encoding shouldn't be possible). But I haven't done a rigorous analysis, so it's pure speculation at this point.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 295328,
          "author_name": "anokas",
          "author_url": "",
          "post_date": "03/13/2018 13:53:59",
          "content": "<p>Oh, apologies. I thought you meant sorting the IPs based on their target values in the training set. I agree that just sorting using the random encoding provided shouldn't give &gt;0.5.</p>\n\n<p>However, I don't think the high AUC <em>on it's own</em> is an indicator of a leak. Some problems are just easy to predict.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 295330,
          "author_name": "konradb",
          "author_url": "",
          "post_date": "03/13/2018 13:58:31",
          "content": "<p>No problem - always good to bounce one's thoughts off somebody else.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 295347,
          "author_name": "domcastro",
          "author_url": "",
          "post_date": "03/13/2018 14:29:50",
          "content": "<p>and there was a similar ID/ sorting leak in the last competition, if I remember correctly</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 295348,
          "author_name": "domcastro",
          "author_url": "",
          "post_date": "03/13/2018 14:31:07",
          "content": "<p>and if I was the company and could easily get 96 AUC I wouldn't pay $$ to outsource the problem</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 295356,
          "author_name": "konradb",
          "author_url": "",
          "post_date": "03/13/2018 14:50:27",
          "content": "<p>Well, that's a more general phenomenon, case in point: the Toxic Comments contest. Given a NB-SVM benchmark scoring ~ .982 AUC while running on a laptop, would any business care about an advanced deep learning model achieving .988 on a cluster?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "291388": "Given some issues with past competitions, I want to be extra careful and wanted to confirm the train is as expected as this is the first time I see a folder structure like this. \n![enter image description here][1]\n\n\nThanks.\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/291388/8691/zipContent.jpeg",
    "291389": "Can you re-download and check again? (Regardless, the file should be fine. zip just got excited during the compression.)",
    "291393": "I am lazy to re-download, but I trust you, if this is it, it is.\n\nFor reference (md5sum):  \n8898d8cfa6a62de21d4e34e4a72a8889  test.csv  \n36da1e8fec8d6765e56d894096a1462e  train.csv",
    "291398": "Please re-download test.csv. (You're good to go with Train)\n\nShould be: 8f27a6d1b1f5bcd96c9183654863df98  test.csv",
    "291407": "I cant match the md5sum(test.csv), I just downloaded it again, it still has the full-path structure but I guess that's OK as train matched. \n\nUnless otherwise notified, I will guess I am just being paranoic here and continue.",
    "291479": "NB The test dataset on the download page is still an old version with 50M samples - until kaggle wake up the only option is to download it from kernels: https://www.kaggle.com/anokas/getting-fixed-test-data",
    "291506": "Thank you for the kind notice.",
    "291670": "Please see this thread:\n\nhttps://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51210",
    "295254": "Didn't the last talkingdata competition have leakage? Judging from the scores, are we sure this one doesn't too?",
    "295258": "What do you mean by judging from the scores?",
    "295318": "Very high AUC",
    "295320": "Haven't checked in that much detail yet, sorting the scores on IP gives an AUC ~ 0.7 =&gt; it does smell funny indeed...",
    "295324": "I don't think that indicates a leak IMO. It makes logical sense that an IP address (user) who is more likely to download an app in the training set is more likely to download an app a few days later (in the test set) - some people download every app they see while others never download anything. \n\nSame with apps - certain apps will have higher likelihoods of being downloaded (they may be more appealing etc.) - and this is why you can get 0.9+ just by ordering the apps by their conversion rate in the training set.",
    "295326": "In principle I agree, except there have been quite a few competitions where sorting by that sort of thing lead to massive jumps (which under random encoding shouldn't be possible). But I haven't done a rigorous analysis, so it's pure speculation at this point.",
    "295328": "Oh, apologies. I thought you meant sorting the IPs based on their target values in the training set. I agree that just sorting using the random encoding provided shouldn't give &gt;0.5.\n\nHowever, I don't think the high AUC _on it's own_ is an indicator of a leak. Some problems are just easy to predict.",
    "295330": "No problem - always good to bounce one's thoughts off somebody else.",
    "295347": "and there was a similar ID/ sorting leak in the last competition, if I remember correctly",
    "295348": "and if I was the company and could easily get 96 AUC I wouldn't pay $$ to outsource the problem",
    "295356": "Well, that's a more general phenomenon, case in point: the Toxic Comments contest. Given a NB-SVM benchmark scoring ~ .982 AUC while running on a laptop, would any business care about an advanced deep learning model achieving .988 on a cluster?"
  },
  "source": "meta"
}