{
  "id": 52374,
  "title": "False leak?",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/52374",
  "author_name": "Andy Harless",
  "post_date": "2018-03-19T17:04:51.184000",
  "votes": 19,
  "comment_count": 24,
  "views": 0,
  "content": "<p>EDIT 8: <em>See <a href=\"https://www.kaggle.com/its7171/ip-encoding-looks-ok-data-is-clean\">CuteChibiko's kernel</a>. The IP codes were apparently assigned starting with the test day. There is a \"false leak\" in the sense I meant, but it is entirely false. Do <strong>not</strong> use <strong>numerical</strong> IP codes as a feature.  (You can use IP codes if you hide the original numerical order, e.g., by frequency encoding, target encoding, one-hot encoding, just assigning arbitrary random new codes, or using a model that ignores the order.)  Other than that, there's nothing you really need to know.</em></p>\n\n<p>EDIT 7: <em>In the comments below, several people report having found that IP codes &gt; 126413 are present in the training set but not the test set and that IP codes in that range are more likely to download.  See <a href=\"https://www.kaggle.com/cpmpml/ip-download-rates\">CPMP's kernel</a>, for example.  (Ignore the rest of this post unless you're curious.)</em></p>\n\n<p>EDIT 5: <em>See yulia's comment in the comments section below [EDIT 6: including second comment with link to her kernel].  In summary, numeric IP codes are informative in the training data but not in the test data, so you have to be careful with validation.</em></p>\n\n<p>There is something weird going on with the way the IP addresses are coded.  If you look at ranges of coded IP addresses (excluding where they match exactly), they have considerable predictive power across splits of the training set.  For example, Joe Eddy finds (as reported <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51236#292536\">here</a>) that a simple logistic regression on the numeric IP codes predicts well (AUC around .7) out of sample (for IP addresses not in the sample used for training) on the training data.  I got <a href=\"https://www.kaggle.com/aharless/predicting-with-just-ip-ranges-as-coded\">similar results</a> using a simple single-feature decision tree.  But this approach does not work for predicting the test data, as I found when I submitted the kernel in the second link above.  <strike>In fact (assuming there is no bug in my code), it does so badly (AUC around .3) it's difficult to attribute the result to chance.  (It could maybe be an artifact of the relationship between IPs that do and don't appear in the training set, since I set the predictions for the former to the global training mean.)</strike>. EDIT3: OK, revised version (which removes influence of global mean) scores around .5 on the test data, so it's not as weird as it seemed at first, but still, the predictability seems to exist across splits of the training data but not between training and test data, so it's still a little weird.</p>\n\n<p>Anybody have an idea what might be going on here? Is the sponsor playing with us?</p>\n\n<p>EDIT: <strike>Further experimentation seems to indicate that the difference between test and validation is probably not an issue about the predictability of the target for new IP addresses but about the relationship between old and new IP addresses.</strike>  In my validation data, new IP addresses seem to be more likely to download, <strike>whereas apparently, in the test data, new IP addresses are less likely to download</strike>. </p>\n\n<p>EDIT2: <strike>No, that's not it either.</strike>  New IP addresses seem to be more likely to download in <a href=\"https://www.kaggle.com/aharless/old-and-new-ips\">both</a> the validation data and the test data.  Perhaps the particular ranges of IP addresses in the test data led to lower predictions. (EDIT3 also: Yes, that seems to be it.) (EDIT4: I mean, that explains why my original result was so very strange, but it doesn't explain why the predictability exists across training data splits but not between training and test.)</p>",
  "messages": [
    {
      "id": 298531,
      "postDate": "2018-03-19T17:04:51.183Z",
      "content": "<p>EDIT 8: <em>See <a href=\"https://www.kaggle.com/its7171/ip-encoding-looks-ok-data-is-clean\">CuteChibiko's kernel</a>. The IP codes were apparently assigned starting with the test day. There is a \"false leak\" in the sense I meant, but it is entirely false. Do <strong>not</strong> use <strong>numerical</strong> IP codes as a feature.  (You can use IP codes if you hide the original numerical order, e.g., by frequency encoding, target encoding, one-hot encoding, just assigning arbitrary random new codes, or using a model that ignores the order.)  Other than that, there's nothing you really need to know.</em></p>\n\n<p>EDIT 7: <em>In the comments below, several people report having found that IP codes &gt; 126413 are present in the training set but not the test set and that IP codes in that range are more likely to download.  See <a href=\"https://www.kaggle.com/cpmpml/ip-download-rates\">CPMP's kernel</a>, for example.  (Ignore the rest of this post unless you're curious.)</em></p>\n\n<p>EDIT 5: <em>See yulia's comment in the comments section below [EDIT 6: including second comment with link to her kernel].  In summary, numeric IP codes are informative in the training data but not in the test data, so you have to be careful with validation.</em></p>\n\n<p>There is something weird going on with the way the IP addresses are coded.  If you look at ranges of coded IP addresses (excluding where they match exactly), they have considerable predictive power across splits of the training set.  For example, Joe Eddy finds (as reported <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51236#292536\">here</a>) that a simple logistic regression on the numeric IP codes predicts well (AUC around .7) out of sample (for IP addresses not in the sample used for training) on the training data.  I got <a href=\"https://www.kaggle.com/aharless/predicting-with-just-ip-ranges-as-coded\">similar results</a> using a simple single-feature decision tree.  But this approach does not work for predicting the test data, as I found when I submitted the kernel in the second link above.  <strike>In fact (assuming there is no bug in my code), it does so badly (AUC around .3) it's difficult to attribute the result to chance.  (It could maybe be an artifact of the relationship between IPs that do and don't appear in the training set, since I set the predictions for the former to the global training mean.)</strike>. EDIT3: OK, revised version (which removes influence of global mean) scores around .5 on the test data, so it's not as weird as it seemed at first, but still, the predictability seems to exist across splits of the training data but not between training and test data, so it's still a little weird.</p>\n\n<p>Anybody have an idea what might be going on here? Is the sponsor playing with us?</p>\n\n<p>EDIT: <strike>Further experimentation seems to indicate that the difference between test and validation is probably not an issue about the predictability of the target for new IP addresses but about the relationship between old and new IP addresses.</strike>  In my validation data, new IP addresses seem to be more likely to download, <strike>whereas apparently, in the test data, new IP addresses are less likely to download</strike>. </p>\n\n<p>EDIT2: <strike>No, that's not it either.</strike>  New IP addresses seem to be more likely to download in <a href=\"https://www.kaggle.com/aharless/old-and-new-ips\">both</a> the validation data and the test data.  Perhaps the particular ranges of IP addresses in the test data led to lower predictions. (EDIT3 also: Yes, that seems to be it.) (EDIT4: I mean, that explains why my original result was so very strange, but it doesn't explain why the predictability exists across training data splits but not between training and test.)</p>",
      "rawMarkdown": "EDIT 8: *See [CuteChibiko's kernel][5]. The IP codes were apparently assigned starting with the test day. There is a \"false leak\" in the sense I meant, but it is entirely false. Do **not** use **numerical** IP codes as a feature.  (You can use IP codes if you hide the original numerical order, e.g., by frequency encoding, target encoding, one-hot encoding, just assigning arbitrary random new codes, or using a model that ignores the order.)  Other than that, there's nothing you really need to know.*\n\nEDIT 7: *In the comments below, several people report having found that IP codes &gt; 126413 are present in the training set but not the test set and that IP codes in that range are more likely to download.  See [CPMP's kernel][1], for example.  (Ignore the rest of this post unless you're curious.)*\n\nEDIT 5: *See yulia's comment in the comments section below [EDIT 6: including second comment with link to her kernel].  In summary, numeric IP codes are informative in the training data but not in the test data, so you have to be careful with validation.*\n\nThere is something weird going on with the way the IP addresses are coded.  If you look at ranges of coded IP addresses (excluding where they match exactly), they have considerable predictive power across splits of the training set.  For example, Joe Eddy finds (as reported [here][2]) that a simple logistic regression on the numeric IP codes predicts well (AUC around .7) out of sample (for IP addresses not in the sample used for training) on the training data.  I got [similar results][3] using a simple single-feature decision tree.  But this approach does not work for predicting the test data, as I found when I submitted the kernel in the second link above.  <strike>In fact (assuming there is no bug in my code), it does so badly (AUC around .3) it's difficult to attribute the result to chance.  (It could maybe be an artifact of the relationship between IPs that do and don't appear in the training set, since I set the predictions for the former to the global training mean.)</strike>. EDIT3: OK, revised version (which removes influence of global mean) scores around .5 on the test data, so it's not as weird as it seemed at first, but still, the predictability seems to exist across splits of the training data but not between training and test data, so it's still a little weird.\n\nAnybody have an idea what might be going on here? Is the sponsor playing with us?\n\nEDIT: <strike>Further experimentation seems to indicate that the difference between test and validation is probably not an issue about the predictability of the target for new IP addresses but about the relationship between old and new IP addresses.</strike>  In my validation data, new IP addresses seem to be more likely to download, <strike>whereas apparently, in the test data, new IP addresses are less likely to download</strike>. \n\nEDIT2: <strike>No, that's not it either.</strike>  New IP addresses seem to be more likely to download in [both][4] the validation data and the test data.  Perhaps the particular ranges of IP addresses in the test data led to lower predictions. (EDIT3 also: Yes, that seems to be it.) (EDIT4: I mean, that explains why my original result was so very strange, but it doesn't explain why the predictability exists across training data splits but not between training and test.)\n\n [1]: https://www.kaggle.com/cpmpml/ip-download-rates\n [2]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51236#292536\n [3]: https://www.kaggle.com/aharless/predicting-with-just-ip-ranges-as-coded\n [4]: https://www.kaggle.com/aharless/old-and-new-ips\n [5]: https://www.kaggle.com/its7171/ip-encoding-looks-ok-data-is-clean",
      "votes": 19
    },
    {
      "id": 299247,
      "postDate": "2018-03-20T17:48:48.127Z",
      "content": "<p>Based on train_sample.csv and my own subsamples (i haven't been able to run a test on full train set), it appears that lower number IPs are strongly associated with higher number of clicks.  i.e. numbers in the range 1 through 125000 have substantially more clicks than IPs in the range 300000 and up.  It's almost as if when the ip values were generated to mask real ips, the data was pre-sorted by the number of clicks per IP, in a few major chunks.  (So they gathered most frequent group, and assigned numbers from 1 to 125000, than next group and another bulk of numbers, than next, etc).  I see about 4 bands of frequencies.</p>\n\n<p>However, based on a few test subsamples I ran, the pattern does not repeat in test data.  The ips over test seem to be mapped truly randomly, and if anything have consistent click density.</p>\n\n<p>Also, (and somebody please check these calculations!) there appear to be the following distribution of IPs:</p>\n\n<p>Overall the number of IPs (test OR train): 333168<br>\nNumber of IPs that are in both (test AND train): 38164<br>\nNumber of IPs that are in Train and NOT in Test: 239232<br>\nNumber of IPs that are in Test and NOT in Train: 55772<br></p>\n\n<p>That means that way over half of IPs in test do not follow the mapping rules of train data.</p>\n\n<p>Hense I think need to be careful in validation.  The pattern of IP assignment in Train does not mimic the one in test.  If you go by using IP value as a signal in train, your final test results will be off substantially.</p>",
      "rawMarkdown": "Based on train_sample.csv and my own subsamples (i haven't been able to run a test on full train set), it appears that lower number IPs are strongly associated with higher number of clicks.  i.e. numbers in the range 1 through 125000 have substantially more clicks than IPs in the range 300000 and up.  It's almost as if when the ip values were generated to mask real ips, the data was pre-sorted by the number of clicks per IP, in a few major chunks.  (So they gathered most frequent group, and assigned numbers from 1 to 125000, than next group and another bulk of numbers, than next, etc).  I see about 4 bands of frequencies.\n\nHowever, based on a few test subsamples I ran, the pattern does not repeat in test data.  The ips over test seem to be mapped truly randomly, and if anything have consistent click density.\n\nAlso, (and somebody please check these calculations!) there appear to be the following distribution of IPs:\n\nOverall the number of IPs (test OR train): 333168<br>\nNumber of IPs that are in both (test AND train): 38164<br>\nNumber of IPs that are in Train and NOT in Test: 239232<br>\nNumber of IPs that are in Test and NOT in Train: 55772<br>\n\nThat means that way over half of IPs in test do not follow the mapping rules of train data.\n\nHense I think need to be careful in validation.  The pattern of IP assignment in Train does not mimic the one in test.  If you go by using IP value as a signal in train, your final test results will be off substantially.\n",
      "votes": 11,
      "replies": [
        {
          "id": 299299,
          "postDate": "2018-03-20T19:09:20.437Z",
          "content": "<p>I calculated these in this notebook : <a>Explore</a></p>\n\n<ul>\n<li>Number of distinct IPs in train:  277396</li>\n<li>Number of distinct IPs in test:  93936</li>\n<li>% IPs in test that are in train as well:  40.63%</li>\n<li>% IPs in train that are in test as well:  13.76%</li>\n</ul>\n\n<p>&gt; So, there are 40% of test IPs present in train and only 13% of train IPs in test.</p>\n\n<p><img src=\"https://image.ibb.co/mUamSx/Capture.jpg\" alt=\"Relation\"></p>",
          "rawMarkdown": "I calculated these in this notebook : [Explore][1]\n\n - Number of distinct IPs in train:  277396\n - Number of distinct IPs in test:  93936\n - % IPs in test that are in train as well:  40.63%\n - % IPs in train that are in test as well:  13.76%\n\n&gt; So, there are 40% of test IPs present in train and only 13% of train IPs in test.\n\n![Relation][2]\n\n\n  [1]: http://From%20my%20notebook%20:%20https://github.com/kimkartavyavimudh/Kaggle-TalkingData/blob/master/Exploration%2Bon%2BFull%2BSet.ipynb\n  [2]: https://image.ibb.co/mUamSx/Capture.jpg",
          "votes": 4
        },
        {
          "id": 299312,
          "postDate": "2018-03-20T19:28:24.837Z",
          "content": "<p>Yep... getting the same patterns...</p>",
          "rawMarkdown": "Yep... getting the same patterns..."
        },
        {
          "id": 299976,
          "postDate": "2018-03-21T06:56:41.903Z",
          "content": "<p>@Yulia and the others, I read your kernel. I think what happended is they try to filter IP in test set only to include 0 - 125000. If we see only in this zone (0-125000) we see the click distribution holds in test and train set. CMIIW. </p>",
          "rawMarkdown": "@Yulia and the others, I read your kernel. I think what happended is they try to filter IP in test set only to include 0 - 125000. If we see only in this zone (0-125000) we see the click distribution holds in test and train set. CMIIW. "
        },
        {
          "id": 300239,
          "postDate": "2018-03-21T12:52:26.707Z",
          "content": "<p>Is it because they filtered it that way, or it was the first group they assigned numbers to?  I'm assuming the clicks in test are all inclusive for times given.  </p>\n\n<p>So maybe they took all IPs on test date and train day 1 and assigned some values at same time. Then all new ids on day 2 of train, than new from day 3 and so on?  This would explain different densities based on @Alexander's chart above.  Except they started on last day.  </p>\n\n<p>Did the old test data cup off as well?</p>",
          "rawMarkdown": "Is it because they filtered it that way, or it was the first group they assigned numbers to?  I'm assuming the clicks in test are all inclusive for times given.  \n\nSo maybe they took all IPs on test date and train day 1 and assigned some values at same time. Then all new ids on day 2 of train, than new from day 3 and so on?  This would explain different densities based on @Alexander's chart above.  Except they started on last day.  \n\nDid the old test data cup off as well?"
        },
        {
          "id": 301353,
          "postDate": "2018-03-22T16:51:10.563Z",
          "content": "<p>@yulia, you are right, they used old test / test supplement to assign ID to IPs.\nOld test data has exactly 126414 unique ips from 0 to 126413.  Ids assigned randomly as I can judge.</p>",
          "rawMarkdown": "@yulia, you are right, they used old test / test supplement to assign ID to IPs.\nOld test data has exactly 126414 unique ips from 0 to 126413.  Ids assigned randomly as I can judge.\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 299274,
      "postDate": "2018-03-20T18:34:44.407Z",
      "content": "<p>I added a Kernel showing the distributions: <a href=\"https://www.kaggle.com/yuliagm/be-careful-about-ips-as-a-signal\">https://www.kaggle.com/yuliagm/be-careful-about-ips-as-a-signal</a></p>\n\n<p>What are your thoughts on how it can impact the validation process?  What's a better setup?\n(@Andy I was using the one you shared in a kernel before...)</p>",
      "rawMarkdown": "I added a Kernel showing the distributions: https://www.kaggle.com/yuliagm/be-careful-about-ips-as-a-signal\n\nWhat are your thoughts on how it can impact the validation process?  What's a better setup?\n(@Andy I was using the one you shared in a kernel before...)",
      "votes": 5,
      "replies": [
        {
          "id": 299291,
          "postDate": "2018-03-20T18:50:27.387Z",
          "content": "<p>I guess we just have to be careful about how we use IP as a feature (in particular, for example, try to avoid using the original codes as input to a tree model, but recode using something like target or frequency encoding, which ignores the original numerical order).  I can't think of a validation method that would be robust to misuse of the IP feature, so for now I'm sticking with what's in my pickle kernel.</p>",
          "rawMarkdown": "I guess we just have to be careful about how we use IP as a feature (in particular, for example, try to avoid using the original codes as input to a tree model, but recode using something like target or frequency encoding, which ignores the original numerical order).  I can't think of a validation method that would be robust to misuse of the IP feature, so for now I'm sticking with what's in my pickle kernel.",
          "votes": 3
        }
      ]
    },
    {
      "id": 299307,
      "postDate": "2018-03-20T19:23:44.177Z",
      "content": "<p>I made a chart of new ips vs time on train data + old test covering period 7 Nov 00:00 - 11 Nov 00:00. \nX axis is a day/time, where </p>\n\n<p>7.0 is Nov 7, 00:00, </p>\n\n<p>7.5 is Nov 7, 12:00</p>\n\n<p>Y axis - number of  ips.</p>\n\n<p>Blue line is number of new ips starting from 7 Nov 00:00. In case of orange line, known ips are ones that used during last hour, all others are consider new. Green - total number of unique IPs during this hour. Red - total number of known IPs.</p>\n\n<p>As one can see there is a pattern, and new ip flow is permanent. There is a spike at midnight, and one can assume that this is some kind of automatic IP reassignment.</p>",
      "rawMarkdown": "I made a chart of new ips vs time on train data + old test covering period 7 Nov 00:00 - 11 Nov 00:00. \nX axis is a day/time, where \n\n7.0 is Nov 7, 00:00, \n\n7.5 is Nov 7, 12:00\n\nY axis - number of  ips.\n\nBlue line is number of new ips starting from 7 Nov 00:00. In case of orange line, known ips are ones that used during last hour, all others are consider new. Green - total number of unique IPs during this hour. Red - total number of known IPs.\n\nAs one can see there is a pattern, and new ip flow is permanent. There is a spike at midnight, and one can assume that this is some kind of automatic IP reassignment.",
      "votes": 6,
      "replies": [
        {
          "id": 299459,
          "postDate": "2018-03-20T22:05:41.347Z",
          "content": "<p>Is the spike at midnight in UTC time or Chinese Standard Time?</p>\n\n<p>Also, what do the dips in Green mean?  Having harder time wrapping my head around that one...</p>",
          "rawMarkdown": "Is the spike at midnight in UTC time or Chinese Standard Time?\n\nAlso, what do the dips in Green mean?  Having harder time wrapping my head around that one..."
        },
        {
          "id": 300010,
          "postDate": "2018-03-21T07:37:01.127Z",
          "content": "<p>@Alexander Firsov Your chart deserves its own time analysis kernel. Great Job!</p>",
          "rawMarkdown": "@Alexander Firsov Your chart deserves its own time analysis kernel. Great Job!"
        },
        {
          "id": 300043,
          "postDate": "2018-03-21T08:02:42.667Z",
          "content": "<p>Spikes are at midnight in Chinese Standard Time.\nGreen is total number of unique IP addresses during a hour. Lowering number of unique IPs during night is natuaral, I suppose. Attached is the same chart but without known ips (red line), it has more details.\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/300043/8834/Figure_2.png\" alt=\"Unique IPs per hour\"></p>",
          "rawMarkdown": "Spikes are at midnight in Chinese Standard Time.\nGreen is total number of unique IP addresses during a hour. Lowering number of unique IPs during night is natuaral, I suppose. Attached is the same chart but without known ips (red line), it has more details.\n![Unique IPs per hour][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/300043/8834/Figure_2.png",
          "votes": 6
        },
        {
          "id": 300223,
          "postDate": "2018-03-21T12:38:03.737Z",
          "content": "<p>Got it! Thanks!  It mimics the clicks pattern.  I think I didn't recognize it when it looked flatter :)</p>",
          "rawMarkdown": "Got it! Thanks!  It mimics the clicks pattern.  I think I didn't recognize it when it looked flatter :)"
        }
      ]
    },
    {
      "id": 303906,
      "postDate": "2018-03-26T18:55:04.737Z",
      "content": "<p>I think it's from the order in which the encoding was done.\n<a href=\"https://www.kaggle.com/its7171/ip-encoding-looks-ok-data-is-clean\">https://www.kaggle.com/its7171/ip-encoding-looks-ok-data-is-clean</a></p>",
      "rawMarkdown": "I think it's from the order in which the encoding was done.\nhttps://www.kaggle.com/its7171/ip-encoding-looks-ok-data-is-clean",
      "votes": 3,
      "replies": [
        {
          "id": 303977,
          "postDate": "2018-03-26T20:45:09.090Z",
          "content": "<p>I think this analysis fully answers the open questions and shows that actually nothing weird is happening.</p>",
          "rawMarkdown": "I think this analysis fully answers the open questions and shows that actually nothing weird is happening."
        }
      ]
    },
    {
      "id": 298658,
      "postDate": "2018-03-19T21:13:47.967Z",
      "content": "<p>Maximizing AUC is defining an order of instances, so instances with higher probability are grouped together. One can think about probability in case of AUC as some kind of index/position to be used for sorting. Assigning the same probability to a number of instances creates some problem for AUC calculation, as there is no explicit order.\nIf there is no order, algorithm will use <em>some</em> order, for example order of instances.</p>\n\n<p>I reversed train set using sort_index and run your experiment. Validation score was 0.5855305874221562 instead of 0.6516781789687467 (original train set order).</p>\n\n<p>I may assume that the effect you observed is not really related to IP nature.</p>\n\n<p>Update: I am not sure my experiment with reversing order is very convincing as DecisionTreeClassifier had no fixed random state. But I still think that problem exists for equal preditions. </p>\n\n<p>Update2: it seems like my understanding of how AUC works was wrong. I could not show change in ROC AUC depending on order of input data in case of equal predictions. My knowledge was based on Gini calculation (Gini = 2*AUC-1) from Porto competition. There were custom Gini implementations that had this flaw, but roc_auc_score has the same result regardless of the order, so my conclusion was incorrect.</p>",
      "rawMarkdown": "Maximizing AUC is defining an order of instances, so instances with higher probability are grouped together. One can think about probability in case of AUC as some kind of index/position to be used for sorting. Assigning the same probability to a number of instances creates some problem for AUC calculation, as there is no explicit order.\nIf there is no order, algorithm will use _some_ order, for example order of instances.\n\nI reversed train set using sort_index and run your experiment. Validation score was 0.5855305874221562 instead of 0.6516781789687467 (original train set order).\n\nI may assume that the effect you observed is not really related to IP nature.\n\nUpdate: I am not sure my experiment with reversing order is very convincing as DecisionTreeClassifier had no fixed random state. But I still think that problem exists for equal preditions. \n\nUpdate2: it seems like my understanding of how AUC works was wrong. I could not show change in ROC AUC depending on order of input data in case of equal predictions. My knowledge was based on Gini calculation (Gini = 2*AUC-1) from Porto competition. There were custom Gini implementations that had this flaw, but roc_auc_score has the same result regardless of the order, so my conclusion was incorrect.",
      "votes": 3
    },
    {
      "id": 298589,
      "postDate": "2018-03-19T19:21:37.317Z",
      "content": "<p>I think .3 AUC is much stranger than just being worse than chance. It means that the predictive signal flips around since you could define a new classifier that takes 1 - p as the probability and get 1 - AUC as your new AUC score. </p>\n\n<p>So the predictive meaning of the ip encoding essentially reverses from the train data to the test data? Strange indeed.</p>",
      "rawMarkdown": "I think .3 AUC is much stranger than just being worse than chance. It means that the predictive signal flips around since you could define a new classifier that takes 1 - p as the probability and get 1 - AUC as your new AUC score. \n\nSo the predictive meaning of the ip encoding essentially reverses from the train data to the test data? Strange indeed.",
      "votes": 3
    },
    {
      "id": 301605,
      "postDate": "2018-03-23T01:49:28.703Z",
      "content": "<p>Maybe this can help: <a href=\"https://www.kaggle.com/cpmpml/ip-download-rates\">https://www.kaggle.com/cpmpml/ip-download-rates</a></p>",
      "rawMarkdown": "Maybe this can help: https://www.kaggle.com/cpmpml/ip-download-rates"
    },
    {
      "id": 298590,
      "postDate": "2018-03-19T19:24:14.123Z",
      "content": "<p>Maybe I'm missing something, but what's the difference between 0.7 AUC and 0.3 AUC?  Wouldn't you convert 0.3 AUC to 0.7 AUC if you just negate or invert your predictions?</p>",
      "rawMarkdown": "Maybe I'm missing something, but what's the difference between 0.7 AUC and 0.3 AUC?  Wouldn't you convert 0.3 AUC to 0.7 AUC if you just negate or invert your predictions?",
      "replies": [
        {
          "id": 298594,
          "postDate": "2018-03-19T19:32:08.573Z",
          "content": "<p>Yes.  But I didn't invert the predictions. The same model that produced 0.7 for the validation data produced 0.3 for the test data.  It's as if the meaning of the feature inverted.  Apparently the IP ranges that were associated with more downloads in the training and validation data were associated with fewer downloads in the test data.</p>",
          "rawMarkdown": "Yes.  But I didn't invert the predictions. The same model that produced 0.7 for the validation data produced 0.3 for the test data.  It's as if the meaning of the feature inverted.  Apparently the IP ranges that were associated with more downloads in the training and validation data were associated with fewer downloads in the test data.",
          "votes": 4
        }
      ]
    },
    {
      "id": 298576,
      "postDate": "2018-03-19T18:44:06.817Z",
      "content": "<p>IP is dynamic in China, when you restart the route, you are likely to get a new IP.\nThere are 93936 unique IPS in the test set, but only 38164 are in the training set, this may affect it?</p>",
      "rawMarkdown": "IP is dynamic in China, when you restart the route, you are likely to get a new IP.\nThere are 93936 unique IPS in the test set, but only 38164 are in the training set, this may affect it?"
    },
    {
      "id": 298534,
      "postDate": "2018-03-19T17:09:24.690Z",
      "content": "<p>Maybe IPs are dynamic?</p>",
      "rawMarkdown": "Maybe IPs are dynamic?",
      "replies": [
        {
          "id": 298539,
          "postDate": "2018-03-19T17:17:23.010Z",
          "content": "<p>That might explain the pattern in the training data if the IPs are coded in a way that preserves subnet information, but in that case you would expect the predictions for the test set to be better than chance rather than worse.</p>",
          "rawMarkdown": "That might explain the pattern in the training data if the IPs are coded in a way that preserves subnet information, but in that case you would expect the predictions for the test set to be better than chance rather than worse."
        }
      ]
    },
    {
      "id": 299872,
      "postDate": "2018-03-21T05:17:39.103Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 299247,
      "author_name": "yulia",
      "author_url": "",
      "post_date": "2018-03-20T17:48:48.127000",
      "content": "<p>Based on train_sample.csv and my own subsamples (i haven't been able to run a test on full train set), it appears that lower number IPs are strongly associated with higher number of clicks.  i.e. numbers in the range 1 through 125000 have substantially more clicks than IPs in the range 300000 and up.  It's almost as if when the ip values were generated to mask real ips, the data was pre-sorted by the number of clicks per IP, in a few major chunks.  (So they gathered most frequent group, and assigned numbers from 1 to 125000, than next group and another bulk of numbers, than next, etc).  I see about 4 bands of frequencies.</p>\n\n<p>However, based on a few test subsamples I ran, the pattern does not repeat in test data.  The ips over test seem to be mapped truly randomly, and if anything have consistent click density.</p>\n\n<p>Also, (and somebody please check these calculations!) there appear to be the following distribution of IPs:</p>\n\n<p>Overall the number of IPs (test OR train): 333168<br>\nNumber of IPs that are in both (test AND train): 38164<br>\nNumber of IPs that are in Train and NOT in Test: 239232<br>\nNumber of IPs that are in Test and NOT in Train: 55772<br></p>\n\n<p>That means that way over half of IPs in test do not follow the mapping rules of train data.</p>\n\n<p>Hense I think need to be careful in validation.  The pattern of IP assignment in Train does not mimic the one in test.  If you go by using IP value as a signal in train, your final test results will be off substantially.</p>",
      "votes": 11,
      "replies": [
        {
          "id": 299299,
          "author_name": "Utkarsh",
          "author_url": "",
          "post_date": "2018-03-20T19:09:20.437000",
          "content": "<p>I calculated these in this notebook : <a>Explore</a></p>\n\n<ul>\n<li>Number of distinct IPs in train:  277396</li>\n<li>Number of distinct IPs in test:  93936</li>\n<li>% IPs in test that are in train as well:  40.63%</li>\n<li>% IPs in train that are in test as well:  13.76%</li>\n</ul>\n\n<p>&gt; So, there are 40% of test IPs present in train and only 13% of train IPs in test.</p>\n\n<p><img src=\"https://image.ibb.co/mUamSx/Capture.jpg\" alt=\"Relation\"></p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 299312,
          "author_name": "yulia",
          "author_url": "",
          "post_date": "2018-03-20T19:28:24.837000",
          "content": "<p>Yep... getting the same patterns...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 299976,
          "author_name": "Muhammad Alfiansyah",
          "author_url": "",
          "post_date": "2018-03-21T06:56:41.903000",
          "content": "<p>@Yulia and the others, I read your kernel. I think what happended is they try to filter IP in test set only to include 0 - 125000. If we see only in this zone (0-125000) we see the click distribution holds in test and train set. CMIIW. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 300239,
          "author_name": "yulia",
          "author_url": "",
          "post_date": "2018-03-21T12:52:26.707000",
          "content": "<p>Is it because they filtered it that way, or it was the first group they assigned numbers to?  I'm assuming the clicks in test are all inclusive for times given.  </p>\n\n<p>So maybe they took all IPs on test date and train day 1 and assigned some values at same time. Then all new ids on day 2 of train, than new from day 3 and so on?  This would explain different densities based on @Alexander's chart above.  Except they started on last day.  </p>\n\n<p>Did the old test data cup off as well?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 301353,
          "author_name": "Alexander Firsov",
          "author_url": "",
          "post_date": "2018-03-22T16:51:10.563000",
          "content": "<p>@yulia, you are right, they used old test / test supplement to assign ID to IPs.\nOld test data has exactly 126414 unique ips from 0 to 126413.  Ids assigned randomly as I can judge.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 299274,
      "author_name": "yulia",
      "author_url": "",
      "post_date": "2018-03-20T18:34:44.407000",
      "content": "<p>I added a Kernel showing the distributions: <a href=\"https://www.kaggle.com/yuliagm/be-careful-about-ips-as-a-signal\">https://www.kaggle.com/yuliagm/be-careful-about-ips-as-a-signal</a></p>\n\n<p>What are your thoughts on how it can impact the validation process?  What's a better setup?\n(@Andy I was using the one you shared in a kernel before...)</p>",
      "votes": 5,
      "replies": [
        {
          "id": 299291,
          "author_name": "Andy Harless",
          "author_url": "",
          "post_date": "2018-03-20T18:50:27.387000",
          "content": "<p>I guess we just have to be careful about how we use IP as a feature (in particular, for example, try to avoid using the original codes as input to a tree model, but recode using something like target or frequency encoding, which ignores the original numerical order).  I can't think of a validation method that would be robust to misuse of the IP feature, so for now I'm sticking with what's in my pickle kernel.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 299307,
      "author_name": "Alexander Firsov",
      "author_url": "",
      "post_date": "2018-03-20T19:23:44.177000",
      "content": "<p>I made a chart of new ips vs time on train data + old test covering period 7 Nov 00:00 - 11 Nov 00:00. \nX axis is a day/time, where </p>\n\n<p>7.0 is Nov 7, 00:00, </p>\n\n<p>7.5 is Nov 7, 12:00</p>\n\n<p>Y axis - number of  ips.</p>\n\n<p>Blue line is number of new ips starting from 7 Nov 00:00. In case of orange line, known ips are ones that used during last hour, all others are consider new. Green - total number of unique IPs during this hour. Red - total number of known IPs.</p>\n\n<p>As one can see there is a pattern, and new ip flow is permanent. There is a spike at midnight, and one can assume that this is some kind of automatic IP reassignment.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 299459,
          "author_name": "yulia",
          "author_url": "",
          "post_date": "2018-03-20T22:05:41.347000",
          "content": "<p>Is the spike at midnight in UTC time or Chinese Standard Time?</p>\n\n<p>Also, what do the dips in Green mean?  Having harder time wrapping my head around that one...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 300010,
          "author_name": "Muhammad Alfiansyah",
          "author_url": "",
          "post_date": "2018-03-21T07:37:01.127000",
          "content": "<p>@Alexander Firsov Your chart deserves its own time analysis kernel. Great Job!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 300043,
          "author_name": "Alexander Firsov",
          "author_url": "",
          "post_date": "2018-03-21T08:02:42.667000",
          "content": "<p>Spikes are at midnight in Chinese Standard Time.\nGreen is total number of unique IP addresses during a hour. Lowering number of unique IPs during night is natuaral, I suppose. Attached is the same chart but without known ips (red line), it has more details.\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/300043/8834/Figure_2.png\" alt=\"Unique IPs per hour\"></p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 300223,
          "author_name": "yulia",
          "author_url": "",
          "post_date": "2018-03-21T12:38:03.737000",
          "content": "<p>Got it! Thanks!  It mimics the clicks pattern.  I think I didn't recognize it when it looked flatter :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 303906,
      "author_name": "tito",
      "author_url": "",
      "post_date": "2018-03-26T18:55:04.737000",
      "content": "<p>I think it's from the order in which the encoding was done.\n<a href=\"https://www.kaggle.com/its7171/ip-encoding-looks-ok-data-is-clean\">https://www.kaggle.com/its7171/ip-encoding-looks-ok-data-is-clean</a></p>",
      "votes": 3,
      "replies": [
        {
          "id": 303977,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2018-03-26T20:45:09.090000",
          "content": "<p>I think this analysis fully answers the open questions and shows that actually nothing weird is happening.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 298658,
      "author_name": "Alexander Firsov",
      "author_url": "",
      "post_date": "2018-03-19T21:13:47.967000",
      "content": "<p>Maximizing AUC is defining an order of instances, so instances with higher probability are grouped together. One can think about probability in case of AUC as some kind of index/position to be used for sorting. Assigning the same probability to a number of instances creates some problem for AUC calculation, as there is no explicit order.\nIf there is no order, algorithm will use <em>some</em> order, for example order of instances.</p>\n\n<p>I reversed train set using sort_index and run your experiment. Validation score was 0.5855305874221562 instead of 0.6516781789687467 (original train set order).</p>\n\n<p>I may assume that the effect you observed is not really related to IP nature.</p>\n\n<p>Update: I am not sure my experiment with reversing order is very convincing as DecisionTreeClassifier had no fixed random state. But I still think that problem exists for equal preditions. </p>\n\n<p>Update2: it seems like my understanding of how AUC works was wrong. I could not show change in ROC AUC depending on order of input data in case of equal predictions. My knowledge was based on Gini calculation (Gini = 2*AUC-1) from Porto competition. There were custom Gini implementations that had this flaw, but roc_auc_score has the same result regardless of the order, so my conclusion was incorrect.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 298589,
      "author_name": "Joe Eddy",
      "author_url": "",
      "post_date": "2018-03-19T19:21:37.317000",
      "content": "<p>I think .3 AUC is much stranger than just being worse than chance. It means that the predictive signal flips around since you could define a new classifier that takes 1 - p as the probability and get 1 - AUC as your new AUC score. </p>\n\n<p>So the predictive meaning of the ip encoding essentially reverses from the train data to the test data? Strange indeed.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 301605,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2018-03-23T01:49:28.703000",
      "content": "<p>Maybe this can help: <a href=\"https://www.kaggle.com/cpmpml/ip-download-rates\">https://www.kaggle.com/cpmpml/ip-download-rates</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 298590,
      "author_name": "Dmitriy Guller",
      "author_url": "",
      "post_date": "2018-03-19T19:24:14.123000",
      "content": "<p>Maybe I'm missing something, but what's the difference between 0.7 AUC and 0.3 AUC?  Wouldn't you convert 0.3 AUC to 0.7 AUC if you just negate or invert your predictions?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 298594,
          "author_name": "Andy Harless",
          "author_url": "",
          "post_date": "2018-03-19T19:32:08.573000",
          "content": "<p>Yes.  But I didn't invert the predictions. The same model that produced 0.7 for the validation data produced 0.3 for the test data.  It's as if the meaning of the feature inverted.  Apparently the IP ranges that were associated with more downloads in the training and validation data were associated with fewer downloads in the test data.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 298576,
      "author_name": "jyyx",
      "author_url": "",
      "post_date": "2018-03-19T18:44:06.817000",
      "content": "<p>IP is dynamic in China, when you restart the route, you are likely to get a new IP.\nThere are 93936 unique IPS in the test set, but only 38164 are in the training set, this may affect it?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 298534,
      "author_name": "Ahmet Erdem",
      "author_url": "",
      "post_date": "2018-03-19T17:09:24.690000",
      "content": "<p>Maybe IPs are dynamic?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 298539,
          "author_name": "Andy Harless",
          "author_url": "",
          "post_date": "2018-03-19T17:17:23.010000",
          "content": "<p>That might explain the pattern in the training data if the IPs are coded in a way that preserves subnet information, but in that case you would expect the predictions for the test set to be better than chance rather than worse.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 299872,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-03-21T05:17:39.103000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "298531": "EDIT 8: *See [CuteChibiko's kernel][5]. The IP codes were apparently assigned starting with the test day. There is a \"false leak\" in the sense I meant, but it is entirely false. Do **not** use **numerical** IP codes as a feature.  (You can use IP codes if you hide the original numerical order, e.g., by frequency encoding, target encoding, one-hot encoding, just assigning arbitrary random new codes, or using a model that ignores the order.)  Other than that, there's nothing you really need to know.*\n\nEDIT 7: *In the comments below, several people report having found that IP codes &gt; 126413 are present in the training set but not the test set and that IP codes in that range are more likely to download.  See [CPMP's kernel][1], for example.  (Ignore the rest of this post unless you're curious.)*\n\nEDIT 5: *See yulia's comment in the comments section below [EDIT 6: including second comment with link to her kernel].  In summary, numeric IP codes are informative in the training data but not in the test data, so you have to be careful with validation.*\n\nThere is something weird going on with the way the IP addresses are coded.  If you look at ranges of coded IP addresses (excluding where they match exactly), they have considerable predictive power across splits of the training set.  For example, Joe Eddy finds (as reported [here][2]) that a simple logistic regression on the numeric IP codes predicts well (AUC around .7) out of sample (for IP addresses not in the sample used for training) on the training data.  I got [similar results][3] using a simple single-feature decision tree.  But this approach does not work for predicting the test data, as I found when I submitted the kernel in the second link above.  <strike>In fact (assuming there is no bug in my code), it does so badly (AUC around .3) it's difficult to attribute the result to chance.  (It could maybe be an artifact of the relationship between IPs that do and don't appear in the training set, since I set the predictions for the former to the global training mean.)</strike>. EDIT3: OK, revised version (which removes influence of global mean) scores around .5 on the test data, so it's not as weird as it seemed at first, but still, the predictability seems to exist across splits of the training data but not between training and test data, so it's still a little weird.\n\nAnybody have an idea what might be going on here? Is the sponsor playing with us?\n\nEDIT: <strike>Further experimentation seems to indicate that the difference between test and validation is probably not an issue about the predictability of the target for new IP addresses but about the relationship between old and new IP addresses.</strike>  In my validation data, new IP addresses seem to be more likely to download, <strike>whereas apparently, in the test data, new IP addresses are less likely to download</strike>. \n\nEDIT2: <strike>No, that's not it either.</strike>  New IP addresses seem to be more likely to download in [both][4] the validation data and the test data.  Perhaps the particular ranges of IP addresses in the test data led to lower predictions. (EDIT3 also: Yes, that seems to be it.) (EDIT4: I mean, that explains why my original result was so very strange, but it doesn't explain why the predictability exists across training data splits but not between training and test.)\n\n [1]: https://www.kaggle.com/cpmpml/ip-download-rates\n [2]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51236#292536\n [3]: https://www.kaggle.com/aharless/predicting-with-just-ip-ranges-as-coded\n [4]: https://www.kaggle.com/aharless/old-and-new-ips\n [5]: https://www.kaggle.com/its7171/ip-encoding-looks-ok-data-is-clean",
    "299247": "Based on train_sample.csv and my own subsamples (i haven't been able to run a test on full train set), it appears that lower number IPs are strongly associated with higher number of clicks.  i.e. numbers in the range 1 through 125000 have substantially more clicks than IPs in the range 300000 and up.  It's almost as if when the ip values were generated to mask real ips, the data was pre-sorted by the number of clicks per IP, in a few major chunks.  (So they gathered most frequent group, and assigned numbers from 1 to 125000, than next group and another bulk of numbers, than next, etc).  I see about 4 bands of frequencies.\n\nHowever, based on a few test subsamples I ran, the pattern does not repeat in test data.  The ips over test seem to be mapped truly randomly, and if anything have consistent click density.\n\nAlso, (and somebody please check these calculations!) there appear to be the following distribution of IPs:\n\nOverall the number of IPs (test OR train): 333168<br>\nNumber of IPs that are in both (test AND train): 38164<br>\nNumber of IPs that are in Train and NOT in Test: 239232<br>\nNumber of IPs that are in Test and NOT in Train: 55772<br>\n\nThat means that way over half of IPs in test do not follow the mapping rules of train data.\n\nHense I think need to be careful in validation.  The pattern of IP assignment in Train does not mimic the one in test.  If you go by using IP value as a signal in train, your final test results will be off substantially.\n",
    "299274": "I added a Kernel showing the distributions: https://www.kaggle.com/yuliagm/be-careful-about-ips-as-a-signal\n\nWhat are your thoughts on how it can impact the validation process?  What's a better setup?\n(@Andy I was using the one you shared in a kernel before...)",
    "299307": "I made a chart of new ips vs time on train data + old test covering period 7 Nov 00:00 - 11 Nov 00:00. \nX axis is a day/time, where \n\n7.0 is Nov 7, 00:00, \n\n7.5 is Nov 7, 12:00\n\nY axis - number of  ips.\n\nBlue line is number of new ips starting from 7 Nov 00:00. In case of orange line, known ips are ones that used during last hour, all others are consider new. Green - total number of unique IPs during this hour. Red - total number of known IPs.\n\nAs one can see there is a pattern, and new ip flow is permanent. There is a spike at midnight, and one can assume that this is some kind of automatic IP reassignment.",
    "303906": "I think it's from the order in which the encoding was done.\nhttps://www.kaggle.com/its7171/ip-encoding-looks-ok-data-is-clean",
    "298658": "Maximizing AUC is defining an order of instances, so instances with higher probability are grouped together. One can think about probability in case of AUC as some kind of index/position to be used for sorting. Assigning the same probability to a number of instances creates some problem for AUC calculation, as there is no explicit order.\nIf there is no order, algorithm will use _some_ order, for example order of instances.\n\nI reversed train set using sort_index and run your experiment. Validation score was 0.5855305874221562 instead of 0.6516781789687467 (original train set order).\n\nI may assume that the effect you observed is not really related to IP nature.\n\nUpdate: I am not sure my experiment with reversing order is very convincing as DecisionTreeClassifier had no fixed random state. But I still think that problem exists for equal preditions. \n\nUpdate2: it seems like my understanding of how AUC works was wrong. I could not show change in ROC AUC depending on order of input data in case of equal predictions. My knowledge was based on Gini calculation (Gini = 2*AUC-1) from Porto competition. There were custom Gini implementations that had this flaw, but roc_auc_score has the same result regardless of the order, so my conclusion was incorrect.",
    "298589": "I think .3 AUC is much stranger than just being worse than chance. It means that the predictive signal flips around since you could define a new classifier that takes 1 - p as the probability and get 1 - AUC as your new AUC score. \n\nSo the predictive meaning of the ip encoding essentially reverses from the train data to the test data? Strange indeed.",
    "301605": "Maybe this can help: https://www.kaggle.com/cpmpml/ip-download-rates",
    "298590": "Maybe I'm missing something, but what's the difference between 0.7 AUC and 0.3 AUC?  Wouldn't you convert 0.3 AUC to 0.7 AUC if you just negate or invert your predictions?",
    "298576": "IP is dynamic in China, when you restart the route, you are likely to get a new IP.\nThere are 93936 unique IPS in the test set, but only 38164 are in the training set, this may affect it?",
    "298534": "Maybe IPs are dynamic?",
    "299872": ""
  }
}