{
  "id": 53668,
  "title": "Be careful - Ids are not sorted",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/53668",
  "author_name": "Μαριος Μιχαηλιδης KazAnova",
  "post_date": "2018-04-03T09:52:23.126000",
  "votes": 29,
  "comment_count": 24,
  "views": 0,
  "content": "<p>I assumed the ids were sorted after looking at them and never put them in memory (bad practice!) . I was looping the predictions' array to generate the index for submission. </p>\n\n<p>I post to make certain people dont waste time with the same mistake :).</p>\n\n<p>E.g</p>\n\n<pre><code>18790435,0\n18790436,0\n18790437,0\n18790439,0&lt;-----------------\n18790438,0 \n</code></pre>\n\n<p>That was tricky because you will get a decline in score, but not enough to understand that there is problem with the ordering. </p>",
  "messages": [
    {
      "id": 308299,
      "postDate": "2018-04-03T09:52:23.127Z",
      "content": "<p>I assumed the ids were sorted after looking at them and never put them in memory (bad practice!) . I was looping the predictions' array to generate the index for submission. </p>\n\n<p>I post to make certain people dont waste time with the same mistake :).</p>\n\n<p>E.g</p>\n\n<pre><code>18790435,0\n18790436,0\n18790437,0\n18790439,0&lt;-----------------\n18790438,0 \n</code></pre>\n\n<p>That was tricky because you will get a decline in score, but not enough to understand that there is problem with the ordering. </p>",
      "rawMarkdown": "I assumed the ids were sorted after looking at them and never put them in memory (bad practice!) . I was looping the predictions' array to generate the index for submission. \n\nI post to make certain people dont waste time with the same mistake :).\n\nE.g\n\n    18790435,0\n    18790436,0\n    18790437,0\n    18790439,0&lt;-----------------\n    18790438,0 \n\nThat was tricky because you will get a decline in score, but not enough to understand that there is problem with the ordering. ",
      "votes": 29
    },
    {
      "id": 308379,
      "postDate": "2018-04-03T12:46:20.080Z",
      "content": "<p>TakingData competitions have a great reputation for creative use of IDs for predictive purposes. It’s one of the less well understood and underappreciated ML techniques. I look forward to finding out what fun surprises this one has in store. Hopefully something major will be revealed in the last week of the competition.</p>",
      "rawMarkdown": "TakingData competitions have a great reputation for creative use of IDs for predictive purposes. It’s one of the less well understood and underappreciated ML techniques. I look forward to finding out what fun surprises this one has in store. Hopefully something major will be revealed in the last week of the competition.",
      "votes": 8,
      "replies": [
        {
          "id": 308382,
          "postDate": "2018-04-03T12:57:08.213Z",
          "content": "<p>I found one already: <a href=\"https://www.kaggle.com/cpmpml/ip-download-rates\">https://www.kaggle.com/cpmpml/ip-download-rates</a> after it was hinted at: <a href=\"https://www.kaggle.com/yuliagm/be-careful-about-ips-as-a-signal\">https://www.kaggle.com/yuliagm/be-careful-about-ips-as-a-signal</a></p>\n\n<p>It was further refined: <a href=\"https://www.kaggle.com/its7171/ip-encoding-looks-ok-data-is-clean\">https://www.kaggle.com/its7171/ip-encoding-looks-ok-data-is-clean</a></p>",
          "rawMarkdown": "I found one already: https://www.kaggle.com/cpmpml/ip-download-rates after it was hinted at: https://www.kaggle.com/yuliagm/be-careful-about-ips-as-a-signal\n\nIt was further refined: https://www.kaggle.com/its7171/ip-encoding-looks-ok-data-is-clean",
          "votes": 2
        },
        {
          "id": 308388,
          "postDate": "2018-04-03T13:01:44.400Z",
          "content": "<p>CPMP, Bojan doesn't say IP but ID:) I think if there was a leak on ID, it could be found already and we could see people with very high scores.</p>",
          "rawMarkdown": "CPMP, Bojan doesn't say IP but ID:) I think if there was a leak on ID, it could be found already and we could see people with very high scores.",
          "votes": 9
        },
        {
          "id": 308401,
          "postDate": "2018-04-03T13:06:37.033Z",
          "content": "<p>@AhmetErdem, thanks.  My bad.  There is no id in the training data however.</p>",
          "rawMarkdown": "@AhmetErdem, thanks.  My bad.  There is no id in the training data however."
        },
        {
          "id": 308432,
          "postDate": "2018-04-03T14:09:24.170Z",
          "content": "<blockquote>\n  <p>I think if there was a leak on ID, it could be found already and we could see people with very high scores.</p>\n</blockquote>\n\n<p>That's <strong>exactly</strong> something that someone who's already found a leak would say. </p>",
          "rawMarkdown": "&gt;I think if there was a leak on ID, it could be found already and we could see people with very high scores.\n\nThat's **exactly** something that someone who's already found a leak would say. \n\n",
          "votes": 3
        },
        {
          "id": 308452,
          "postDate": "2018-04-03T14:34:43.210Z",
          "content": "<p>Since you noticed, I will reveal it. (IP % 27) gives you 0.9806 already. The other leak which makes it 0.9807, I will keep it for myself.</p>",
          "rawMarkdown": "Since you noticed, I will reveal it. (IP % 27) gives you 0.9806 already. The other leak which makes it 0.9807, I will keep it for myself.",
          "votes": 11
        },
        {
          "id": 308486,
          "postDate": "2018-04-03T15:08:59.033Z",
          "content": "<p><img src=\"http://noahtravisphillips.com/HTMLCSSNTP/Weather3ntp/img/leaky%20bucket%20dollar%20ntp.gif\" alt=\"leaky\"></p>",
          "rawMarkdown": "![leaky][1]\n\n\n  [1]: http://noahtravisphillips.com/HTMLCSSNTP/Weather3ntp/img/leaky%20bucket%20dollar%20ntp.gif",
          "votes": 1
        },
        {
          "id": 308491,
          "postDate": "2018-04-03T15:18:23.217Z",
          "content": "<p>@Bojan: paranoia - excellent survival trait</p>\n\n<p><a href=\"http://cheezburger.com/8425753600\">http://cheezburger.com/8425753600</a></p>",
          "rawMarkdown": "@Bojan: paranoia - excellent survival trait\n\nhttp://cheezburger.com/8425753600",
          "votes": 2
        },
        {
          "id": 308494,
          "postDate": "2018-04-03T15:25:06.970Z",
          "content": "<p>@Konrad - I think you are just trying to throw me off and confuse me with that comment. </p>",
          "rawMarkdown": "@Konrad - I think you are just trying to throw me off and confuse me with that comment. ",
          "votes": 1
        },
        {
          "id": 308504,
          "postDate": "2018-04-03T15:49:35.580Z",
          "content": "<p>Here I am, trying to give moral support to a fellow paranoid - and this is the thanks I get... Truly, no good deed goes unpunished :-)</p>",
          "rawMarkdown": "Here I am, trying to give moral support to a fellow paranoid - and this is the thanks I get... Truly, no good deed goes unpunished :-)",
          "votes": 1
        },
        {
          "id": 308524,
          "postDate": "2018-04-03T16:30:26.180Z",
          "content": "<p>Hi @AhmetErdem, could we have more explanation on (IP % 27)? Is it a joke or do you mean there is a leakage feature which value is IP mod 27? I just tried IP mod 27 and found the distribution of is_attributed is almost the same with each (IP mod 27). Thank you.</p>",
          "rawMarkdown": " Hi @AhmetErdem, could we have more explanation on (IP % 27)? Is it a joke or do you mean there is a leakage feature which value is IP mod 27? I just tried IP mod 27 and found the distribution of is_attributed is almost the same with each (IP mod 27). Thank you.\n",
          "votes": 2
        },
        {
          "id": 308525,
          "postDate": "2018-04-03T16:35:03.820Z",
          "content": "<p>I think that was a joke. </p>\n\n<p>In any case, I would not rule out a leak. </p>",
          "rawMarkdown": "I think that was a joke. \n\nIn any case, I would not rule out a leak. ",
          "votes": 3
        },
        {
          "id": 308540,
          "postDate": "2018-04-03T17:01:50.980Z",
          "content": "<p>Hi @Snorlax, I am really sorry for making you try that. I just wanted to answer @Bojan's joke with a joke:)</p>",
          "rawMarkdown": "Hi @Snorlax, I am really sorry for making you try that. I just wanted to answer @Bojan's joke with a joke:)",
          "votes": 1
        },
        {
          "id": 308571,
          "postDate": "2018-04-03T17:35:32.870Z",
          "content": "<p>Hi @AhmetErdem, no need to say sorry. I volunteered to try that, lol.</p>",
          "rawMarkdown": "Hi @AhmetErdem, no need to say sorry. I volunteered to try that, lol.",
          "votes": 1
        },
        {
          "id": 308588,
          "postDate": "2018-04-03T17:57:53.603Z",
          "content": "<blockquote>\n  <p>I just wanted to answer @Bojan's joke with a joke:)</p>\n</blockquote>\n\n<p>Or is it? Hmmmm ...</p>",
          "rawMarkdown": "&gt; I just wanted to answer @Bojan's joke with a joke:)\n\nOr is it? Hmmmm ..."
        },
        {
          "id": 309147,
          "postDate": "2018-04-04T18:13:26.793Z",
          "content": "<p>@KazAnova: Not sure if there is a leak or not but there are some magical features for sure. Someone from top teams already mentioned that with only 4 crafted features they could achieve 0.976x</p>",
          "rawMarkdown": "@KazAnova: Not sure if there is a leak or not but there are some magical features for sure. Someone from top teams already mentioned that with only 4 crafted features they could achieve 0.976x",
          "votes": 1
        },
        {
          "id": 309195,
          "postDate": "2018-04-04T19:32:45.207Z",
          "content": "<p>I don't think you need a leak to explain that.  With buggy data prep I could get 0.9736 with a single model and few features only.</p>\n\n<p>Maybe I'll find that my bug is uncovering a leak when I'll see that with fixed data prep my score will be way lower....</p>",
          "rawMarkdown": "I don't think you need a leak to explain that.  With buggy data prep I could get 0.9736 with a single model and few features only.\n\nMaybe I'll find that my bug is uncovering a leak when I'll see that with fixed data prep my score will be way lower...."
        }
      ]
    },
    {
      "id": 308320,
      "postDate": "2018-04-03T10:42:13.553Z",
      "content": "<p>People are bumping into this again and again, see for instance <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52354#298395\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52354#298395</a></p>",
      "rawMarkdown": "People are bumping into this again and again, see for instance https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52354#298395",
      "votes": 1,
      "replies": [
        {
          "id": 308333,
          "postDate": "2018-04-03T11:15:40.280Z",
          "content": "<p>Yeah - I missed that :(</p>\n\n<p>Worth making its own thread I think. Very misleading!  Kaggle should have highlighted this, because on a quick glance it looks sorted.</p>",
          "rawMarkdown": "Yeah - I missed that :(\n\nWorth making its own thread I think. Very misleading!  Kaggle should have highlighted this, because on a quick glance it looks sorted.",
          "votes": 2
        },
        {
          "id": 308338,
          "postDate": "2018-04-03T11:28:04.537Z",
          "content": "<p>I agree it is worth a thread.  And coming from you people will read it!</p>\n\n<p>When I tried to explain it to someone in this forum I was dismissed... ;)</p>",
          "rawMarkdown": "I agree it is worth a thread.  And coming from you people will read it!\n\nWhen I tried to explain it to someone in this forum I was dismissed... ;)",
          "votes": 1
        }
      ]
    },
    {
      "id": 308303,
      "postDate": "2018-04-03T10:00:49.157Z",
      "content": "<p>It's probably my caffeine deficit speaking, but I don't see how it can impact the score - the only case I can think of is if you are trying to combine the predictions on the holdout set (for stacking or sth) and concatenate them as opposed to joining.</p>",
      "rawMarkdown": "It's probably my caffeine deficit speaking, but I don't see how it can impact the score - the only case I can think of is if you are trying to combine the predictions on the holdout set (for stacking or sth) and concatenate them as opposed to joining.",
      "replies": [
        {
          "id": 308305,
          "postDate": "2018-04-03T10:04:39.150Z",
          "content": "<p>Or...</p>\n\n<p>without changing the ordering of the test data and you get a <code>preds</code> array that hold your predictions and you generate submission like I did!</p>\n\n<pre><code>for i in range (0,len(preds)):\n        file.write(str(i) + \",\" + str(preds[i]) )\n</code></pre>\n\n<p>When I say score , I mean the <code>AUC</code></p>",
          "rawMarkdown": "Or...\n\nwithout changing the ordering of the test data and you get a `preds` array that hold your predictions and you generate submission like I did!\n\n    for i in range (0,len(preds)):\n            file.write(str(i) + \",\" + str(preds[i]) )\n\nWhen I say score , I mean the `AUC`\n",
          "votes": 1
        },
        {
          "id": 308307,
          "postDate": "2018-04-03T10:06:42.987Z",
          "content": "<p>Ok, so I wasn't that far off :-)</p>",
          "rawMarkdown": "Ok, so I wasn't that far off :-)",
          "votes": 1
        }
      ]
    },
    {
      "id": 309165,
      "postDate": "2018-04-04T18:52:55.237Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 308379,
      "author_name": "Bojan Tunguz",
      "author_url": "",
      "post_date": "2018-04-03T12:46:20.080000",
      "content": "<p>TakingData competitions have a great reputation for creative use of IDs for predictive purposes. It’s one of the less well understood and underappreciated ML techniques. I look forward to finding out what fun surprises this one has in store. Hopefully something major will be revealed in the last week of the competition.</p>",
      "votes": 8,
      "replies": [
        {
          "id": 308382,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-04-03T12:57:08.213000",
          "content": "<p>I found one already: <a href=\"https://www.kaggle.com/cpmpml/ip-download-rates\">https://www.kaggle.com/cpmpml/ip-download-rates</a> after it was hinted at: <a href=\"https://www.kaggle.com/yuliagm/be-careful-about-ips-as-a-signal\">https://www.kaggle.com/yuliagm/be-careful-about-ips-as-a-signal</a></p>\n\n<p>It was further refined: <a href=\"https://www.kaggle.com/its7171/ip-encoding-looks-ok-data-is-clean\">https://www.kaggle.com/its7171/ip-encoding-looks-ok-data-is-clean</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 308388,
          "author_name": "Ahmet Erdem",
          "author_url": "",
          "post_date": "2018-04-03T13:01:44.400000",
          "content": "<p>CPMP, Bojan doesn't say IP but ID:) I think if there was a leak on ID, it could be found already and we could see people with very high scores.</p>",
          "votes": 9,
          "replies": []
        },
        {
          "id": 308401,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-04-03T13:06:37.033000",
          "content": "<p>@AhmetErdem, thanks.  My bad.  There is no id in the training data however.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 308432,
          "author_name": "Bojan Tunguz",
          "author_url": "",
          "post_date": "2018-04-03T14:09:24.170000",
          "content": "<blockquote>\n  <p>I think if there was a leak on ID, it could be found already and we could see people with very high scores.</p>\n</blockquote>\n\n<p>That's <strong>exactly</strong> something that someone who's already found a leak would say. </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 308452,
          "author_name": "Ahmet Erdem",
          "author_url": "",
          "post_date": "2018-04-03T14:34:43.210000",
          "content": "<p>Since you noticed, I will reveal it. (IP % 27) gives you 0.9806 already. The other leak which makes it 0.9807, I will keep it for myself.</p>",
          "votes": 11,
          "replies": []
        },
        {
          "id": 308486,
          "author_name": "Bojan Tunguz",
          "author_url": "",
          "post_date": "2018-04-03T15:08:59.033000",
          "content": "<p><img src=\"http://noahtravisphillips.com/HTMLCSSNTP/Weather3ntp/img/leaky%20bucket%20dollar%20ntp.gif\" alt=\"leaky\"></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 308491,
          "author_name": "Konrad Banachewicz",
          "author_url": "",
          "post_date": "2018-04-03T15:18:23.217000",
          "content": "<p>@Bojan: paranoia - excellent survival trait</p>\n\n<p><a href=\"http://cheezburger.com/8425753600\">http://cheezburger.com/8425753600</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 308494,
          "author_name": "Bojan Tunguz",
          "author_url": "",
          "post_date": "2018-04-03T15:25:06.970000",
          "content": "<p>@Konrad - I think you are just trying to throw me off and confuse me with that comment. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 308504,
          "author_name": "Konrad Banachewicz",
          "author_url": "",
          "post_date": "2018-04-03T15:49:35.580000",
          "content": "<p>Here I am, trying to give moral support to a fellow paranoid - and this is the thanks I get... Truly, no good deed goes unpunished :-)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 308524,
          "author_name": "Snorlax",
          "author_url": "",
          "post_date": "2018-04-03T16:30:26.180000",
          "content": "<p>Hi @AhmetErdem, could we have more explanation on (IP % 27)? Is it a joke or do you mean there is a leakage feature which value is IP mod 27? I just tried IP mod 27 and found the distribution of is_attributed is almost the same with each (IP mod 27). Thank you.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 308525,
          "author_name": "Μαριος Μιχαηλιδης KazAnova",
          "author_url": "",
          "post_date": "2018-04-03T16:35:03.820000",
          "content": "<p>I think that was a joke. </p>\n\n<p>In any case, I would not rule out a leak. </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 308540,
          "author_name": "Ahmet Erdem",
          "author_url": "",
          "post_date": "2018-04-03T17:01:50.980000",
          "content": "<p>Hi @Snorlax, I am really sorry for making you try that. I just wanted to answer @Bojan's joke with a joke:)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 308571,
          "author_name": "Snorlax",
          "author_url": "",
          "post_date": "2018-04-03T17:35:32.870000",
          "content": "<p>Hi @AhmetErdem, no need to say sorry. I volunteered to try that, lol.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 308588,
          "author_name": "Bojan Tunguz",
          "author_url": "",
          "post_date": "2018-04-03T17:57:53.603000",
          "content": "<blockquote>\n  <p>I just wanted to answer @Bojan's joke with a joke:)</p>\n</blockquote>\n\n<p>Or is it? Hmmmm ...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 309147,
          "author_name": "Sameh Faidi",
          "author_url": "",
          "post_date": "2018-04-04T18:13:26.793000",
          "content": "<p>@KazAnova: Not sure if there is a leak or not but there are some magical features for sure. Someone from top teams already mentioned that with only 4 crafted features they could achieve 0.976x</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 309195,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-04-04T19:32:45.207000",
          "content": "<p>I don't think you need a leak to explain that.  With buggy data prep I could get 0.9736 with a single model and few features only.</p>\n\n<p>Maybe I'll find that my bug is uncovering a leak when I'll see that with fixed data prep my score will be way lower....</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 308320,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2018-04-03T10:42:13.553000",
      "content": "<p>People are bumping into this again and again, see for instance <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52354#298395\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52354#298395</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 308333,
          "author_name": "Μαριος Μιχαηλιδης KazAnova",
          "author_url": "",
          "post_date": "2018-04-03T11:15:40.280000",
          "content": "<p>Yeah - I missed that :(</p>\n\n<p>Worth making its own thread I think. Very misleading!  Kaggle should have highlighted this, because on a quick glance it looks sorted.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 308338,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-04-03T11:28:04.537000",
          "content": "<p>I agree it is worth a thread.  And coming from you people will read it!</p>\n\n<p>When I tried to explain it to someone in this forum I was dismissed... ;)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 308303,
      "author_name": "Konrad Banachewicz",
      "author_url": "",
      "post_date": "2018-04-03T10:00:49.157000",
      "content": "<p>It's probably my caffeine deficit speaking, but I don't see how it can impact the score - the only case I can think of is if you are trying to combine the predictions on the holdout set (for stacking or sth) and concatenate them as opposed to joining.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 308305,
          "author_name": "Μαριος Μιχαηλιδης KazAnova",
          "author_url": "",
          "post_date": "2018-04-03T10:04:39.150000",
          "content": "<p>Or...</p>\n\n<p>without changing the ordering of the test data and you get a <code>preds</code> array that hold your predictions and you generate submission like I did!</p>\n\n<pre><code>for i in range (0,len(preds)):\n        file.write(str(i) + \",\" + str(preds[i]) )\n</code></pre>\n\n<p>When I say score , I mean the <code>AUC</code></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 308307,
          "author_name": "Konrad Banachewicz",
          "author_url": "",
          "post_date": "2018-04-03T10:06:42.987000",
          "content": "<p>Ok, so I wasn't that far off :-)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 309165,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-04-04T18:52:55.237000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "308299": "I assumed the ids were sorted after looking at them and never put them in memory (bad practice!) . I was looping the predictions' array to generate the index for submission. \n\nI post to make certain people dont waste time with the same mistake :).\n\nE.g\n\n    18790435,0\n    18790436,0\n    18790437,0\n    18790439,0&lt;-----------------\n    18790438,0 \n\nThat was tricky because you will get a decline in score, but not enough to understand that there is problem with the ordering. ",
    "308379": "TakingData competitions have a great reputation for creative use of IDs for predictive purposes. It’s one of the less well understood and underappreciated ML techniques. I look forward to finding out what fun surprises this one has in store. Hopefully something major will be revealed in the last week of the competition.",
    "308320": "People are bumping into this again and again, see for instance https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52354#298395",
    "308303": "It's probably my caffeine deficit speaking, but I don't see how it can impact the score - the only case I can think of is if you are trying to combine the predictions on the holdout set (for stacking or sth) and concatenate them as opposed to joining.",
    "309165": ""
  }
}