{
  "id": 51236,
  "title": "Data encoding - is it completely random?",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/51236",
  "author_name": "",
  "post_date": "2018-03-06T21:29:23.955471500Z",
  "votes": 2,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I've a question on the method used to encode the data. Let's take 'ip' as an example.</p>\n\n<p>Do I get it correct that 'ip' is a completely random id assigned for every unique real IP address?\nThat is we cannot expect any similarity for ip's which correspond to similar  real IPs?</p>",
  "messages": [
    {
      "id": "291797",
      "postDate": "03/06/2018 21:29:23",
      "content": "<p>I've a question on the method used to encode the data. Let's take 'ip' as an example.</p>\n\n<p>Do I get it correct that 'ip' is a completely random id assigned for every unique real IP address?\nThat is we cannot expect any similarity for ip's which correspond to similar  real IPs?</p>",
      "rawMarkdown": "I've a question on the method used to encode the data. Let's take 'ip' as an example.\n\nDo I get it correct that 'ip' is a completely random id assigned for every unique real IP address?\nThat is we cannot expect any similarity for ip's which correspond to similar  real IPs?",
      "votes": null
    },
    {
      "id": "291875",
      "postDate": "03/07/2018 02:42:05",
      "content": "<p>All columns except channel is converted to integer as an index. They are supposed to be categorical.\nI don't quite understand your second question, could you please be more specific?</p>",
      "rawMarkdown": "All columns except channel is converted to integer as an index. They are supposed to be categorical.\nI don't quite understand your second question, could you please be more specific?",
      "votes": null
    },
    {
      "id": "292091",
      "postDate": "03/07/2018 12:18:10",
      "content": "<p>He's asking if you can figure out the value of the raw data from the encoding, being able to do so could lead to leakage. Like if I had the original data points \"X\" and the encoded points \"e(X)\", that x_1 - x_0 is not equal or similar to e(x_1) - e(x_0).</p>",
      "rawMarkdown": "He's asking if you can figure out the value of the raw data from the encoding, being able to do so could lead to leakage. Like if I had the original data points \"X\" and the encoded points \"e(X)\", that x_1 - x_0 is not equal or similar to e(x_1) - e(x_0).",
      "votes": null
    },
    {
      "id": "292126",
      "postDate": "03/07/2018 13:58:11",
      "content": "<p>This is also important as to know if there is continuity information. E.g. if 170.10.20.30 become 12345, does 170.10.20.31 become something like 12346? Not necessarily that, but is the encoding function bijective or monotonous?</p>",
      "rawMarkdown": "This is also important as to know if there is continuity information. E.g. if 170.10.20.30 become 12345, does 170.10.20.31 become something like 12346? Not necessarily that, but is the encoding function bijective or monotonous?",
      "votes": null
    },
    {
      "id": "292509",
      "postDate": "03/08/2018 04:04:34",
      "content": "<p>I believe <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51348\">this discussion</a> is the answer to the question.</p>",
      "rawMarkdown": "I believe [this discussion][1] is the answer to the question.\n\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51348",
      "votes": null
    },
    {
      "id": "292536",
      "postDate": "03/08/2018 05:07:15",
      "content": "<p>I'm pretty sure it's far from random. I've only looked at the sample, but on the sample you can get held-out AUC of &gt; .7 by running a logistic regression on the ip column, even if you remove all the duplicate ips. If the encoding is random then the ordering would be random and the raw ip column would be useless for prediction. The ip encoding is at least a proxy for something with signal.</p>\n\n<p>From what Aaron Yin has said, I don't think this means that the encoded IP addresses are necessarily close together as real ips. I think it's more likely that the encoding scheme is based on some other similarity between true ips.</p>",
      "rawMarkdown": "I'm pretty sure it's far from random. I've only looked at the sample, but on the sample you can get held-out AUC of &gt; .7 by running a logistic regression on the ip column, even if you remove all the duplicate ips. If the encoding is random then the ordering would be random and the raw ip column would be useless for prediction. The ip encoding is at least a proxy for something with signal.\n\nFrom what Aaron Yin has said, I don't think this means that the encoded IP addresses are necessarily close together as real ips. I think it's more likely that the encoding scheme is based on some other similarity between true ips.",
      "votes": null
    },
    {
      "id": "294600",
      "postDate": "03/12/2018 07:57:56",
      "content": "<p>Excuse me, but I couldn't follow your conclusion. Why IP would be useless for prediction if the encoding is random? Example: I (unique IP, encoded or not) like downloading apps because I believe in advertisments and therefore given that I converted already once, it is likely that I convert again. Bots that just click and don't download have never downloaded so far and won't do in the future either.\nAlso what additional information (besides proximity, which I also don't know how would be indicative) would you expect from having the \"real\" IPs instead of the encoded ones?</p>",
      "rawMarkdown": "Excuse me, but I couldn't follow your conclusion. Why IP would be useless for prediction if the encoding is random? Example: I (unique IP, encoded or not) like downloading apps because I believe in advertisments and therefore given that I converted already once, it is likely that I convert again. Bots that just click and don't download have never downloaded so far and won't do in the future either.\nAlso what additional information (besides proximity, which I also don't know how would be indicative) would you expect from having the \"real\" IPs instead of the encoded ones?",
      "votes": null
    },
    {
      "id": "294744",
      "postDate": "03/12/2018 15:04:47",
      "content": "<p>My conclusion is specific to using the original ip encoding column as a feature without first deriving other features from it. I definitely agree that you can derive past behavior features as you suggest, regardless of whether the encoding is random. My point is that if the encoding is random, the <strong>ordering</strong> of the encoded ip column should not be useful for prediction - but the success of a simple logistic regression on the original column argues that the ordering is useful, and that therefore the encoding is not random.</p>\n\n<p>Hope that clears it up! The second question is more open-ended - I do think proximity would be the main signal in the \"real\" ip. My point with the encoding is that we're told it's not preserving proximity, yet there is clear signal in ordering. So if we believe the first part (that proximity isn't preserved), we should think that some other signal is getting baked into the encoding.</p>",
      "rawMarkdown": "My conclusion is specific to using the original ip encoding column as a feature without first deriving other features from it. I definitely agree that you can derive past behavior features as you suggest, regardless of whether the encoding is random. My point is that if the encoding is random, the **ordering** of the encoded ip column should not be useful for prediction - but the success of a simple logistic regression on the original column argues that the ordering is useful, and that therefore the encoding is not random.\n\nHope that clears it up! The second question is more open-ended - I do think proximity would be the main signal in the \"real\" ip. My point with the encoding is that we're told it's not preserving proximity, yet there is clear signal in ordering. So if we believe the first part (that proximity isn't preserved), we should think that some other signal is getting baked into the encoding.",
      "votes": null
    },
    {
      "id": "294746",
      "postDate": "03/12/2018 15:08:38",
      "content": "<p>Thanks for the clarification!</p>",
      "rawMarkdown": "Thanks for the clarification!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 291875,
      "author_name": "aaronyin",
      "author_url": "",
      "post_date": "03/07/2018 02:42:05",
      "content": "<p>All columns except channel is converted to integer as an index. They are supposed to be categorical.\nI don't quite understand your second question, could you please be more specific?</p>",
      "votes": null,
      "replies": [
        {
          "id": 292091,
          "author_name": "jessiecarnegie7777",
          "author_url": "",
          "post_date": "03/07/2018 12:18:10",
          "content": "<p>He's asking if you can figure out the value of the raw data from the encoding, being able to do so could lead to leakage. Like if I had the original data points \"X\" and the encoded points \"e(X)\", that x_1 - x_0 is not equal or similar to e(x_1) - e(x_0).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 292509,
          "author_name": "aaronyin",
          "author_url": "",
          "post_date": "03/08/2018 04:04:34",
          "content": "<p>I believe <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51348\">this discussion</a> is the answer to the question.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 292126,
      "author_name": "ajallooeian",
      "author_url": "",
      "post_date": "03/07/2018 13:58:11",
      "content": "<p>This is also important as to know if there is continuity information. E.g. if 170.10.20.30 become 12345, does 170.10.20.31 become something like 12346? Not necessarily that, but is the encoding function bijective or monotonous?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 292536,
      "author_name": "aquatic",
      "author_url": "",
      "post_date": "03/08/2018 05:07:15",
      "content": "<p>I'm pretty sure it's far from random. I've only looked at the sample, but on the sample you can get held-out AUC of &gt; .7 by running a logistic regression on the ip column, even if you remove all the duplicate ips. If the encoding is random then the ordering would be random and the raw ip column would be useless for prediction. The ip encoding is at least a proxy for something with signal.</p>\n\n<p>From what Aaron Yin has said, I don't think this means that the encoded IP addresses are necessarily close together as real ips. I think it's more likely that the encoding scheme is based on some other similarity between true ips.</p>",
      "votes": null,
      "replies": [
        {
          "id": 294600,
          "author_name": "asparuhhristov",
          "author_url": "",
          "post_date": "03/12/2018 07:57:56",
          "content": "<p>Excuse me, but I couldn't follow your conclusion. Why IP would be useless for prediction if the encoding is random? Example: I (unique IP, encoded or not) like downloading apps because I believe in advertisments and therefore given that I converted already once, it is likely that I convert again. Bots that just click and don't download have never downloaded so far and won't do in the future either.\nAlso what additional information (besides proximity, which I also don't know how would be indicative) would you expect from having the \"real\" IPs instead of the encoded ones?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 294744,
          "author_name": "aquatic",
          "author_url": "",
          "post_date": "03/12/2018 15:04:47",
          "content": "<p>My conclusion is specific to using the original ip encoding column as a feature without first deriving other features from it. I definitely agree that you can derive past behavior features as you suggest, regardless of whether the encoding is random. My point is that if the encoding is random, the <strong>ordering</strong> of the encoded ip column should not be useful for prediction - but the success of a simple logistic regression on the original column argues that the ordering is useful, and that therefore the encoding is not random.</p>\n\n<p>Hope that clears it up! The second question is more open-ended - I do think proximity would be the main signal in the \"real\" ip. My point with the encoding is that we're told it's not preserving proximity, yet there is clear signal in ordering. So if we believe the first part (that proximity isn't preserved), we should think that some other signal is getting baked into the encoding.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 294746,
          "author_name": "asparuhhristov",
          "author_url": "",
          "post_date": "03/12/2018 15:08:38",
          "content": "<p>Thanks for the clarification!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "291797": "I've a question on the method used to encode the data. Let's take 'ip' as an example.\n\nDo I get it correct that 'ip' is a completely random id assigned for every unique real IP address?\nThat is we cannot expect any similarity for ip's which correspond to similar  real IPs?",
    "291875": "All columns except channel is converted to integer as an index. They are supposed to be categorical.\nI don't quite understand your second question, could you please be more specific?",
    "292091": "He's asking if you can figure out the value of the raw data from the encoding, being able to do so could lead to leakage. Like if I had the original data points \"X\" and the encoded points \"e(X)\", that x_1 - x_0 is not equal or similar to e(x_1) - e(x_0).",
    "292126": "This is also important as to know if there is continuity information. E.g. if 170.10.20.30 become 12345, does 170.10.20.31 become something like 12346? Not necessarily that, but is the encoding function bijective or monotonous?",
    "292509": "I believe [this discussion][1] is the answer to the question.\n\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51348",
    "292536": "I'm pretty sure it's far from random. I've only looked at the sample, but on the sample you can get held-out AUC of &gt; .7 by running a logistic regression on the ip column, even if you remove all the duplicate ips. If the encoding is random then the ordering would be random and the raw ip column would be useless for prediction. The ip encoding is at least a proxy for something with signal.\n\nFrom what Aaron Yin has said, I don't think this means that the encoded IP addresses are necessarily close together as real ips. I think it's more likely that the encoding scheme is based on some other similarity between true ips.",
    "294600": "Excuse me, but I couldn't follow your conclusion. Why IP would be useless for prediction if the encoding is random? Example: I (unique IP, encoded or not) like downloading apps because I believe in advertisments and therefore given that I converted already once, it is likely that I convert again. Bots that just click and don't download have never downloaded so far and won't do in the future either.\nAlso what additional information (besides proximity, which I also don't know how would be indicative) would you expect from having the \"real\" IPs instead of the encoded ones?",
    "294744": "My conclusion is specific to using the original ip encoding column as a feature without first deriving other features from it. I definitely agree that you can derive past behavior features as you suggest, regardless of whether the encoding is random. My point is that if the encoding is random, the **ordering** of the encoded ip column should not be useful for prediction - but the success of a simple logistic regression on the original column argues that the ordering is useful, and that therefore the encoding is not random.\n\nHope that clears it up! The second question is more open-ended - I do think proximity would be the main signal in the \"real\" ip. My point with the encoding is that we're told it's not preserving proximity, yet there is clear signal in ordering. So if we believe the first part (that proximity isn't preserved), we should think that some other signal is getting baked into the encoding.",
    "294746": "Thanks for the clarification!"
  },
  "source": "meta"
}