{
  "id": 51348,
  "title": "Encoded Data Questions",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/51348",
  "author_name": "",
  "post_date": "2018-03-07T22:40:16.896219800Z",
  "votes": 7,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Noticed that some of the data fields are \"encoded\":\n - ip: ip address of click, app: app id for marketing\n - device: device type id of user mobile phone (e.g., iphone 6 plus, iphone 7, huawei mate 7, etc.)\n - os: os version id of user mobile phone\n - channel: channel id of mobile ad publisher</p>\n\n<ol>\n<li><p>Are all of the encoding methods one-to-one encoding? For example, every \"device\" of type \"A\" was encoded if and only if it was the same \"device\"?</p></li>\n<li><p>Did the encoding methods destroy \"nearness\" information? For example, IP addresses have specific groupings based on specific allocations of ranges of IP addresses. Are two encoded IP addresses \"near\" each other in the same way the original IP addresses were \"near\" each other?</p></li>\n</ol>",
  "messages": [
    {
      "id": "292407",
      "postDate": "03/07/2018 22:40:16",
      "content": "<p>Noticed that some of the data fields are \"encoded\":\n - ip: ip address of click, app: app id for marketing\n - device: device type id of user mobile phone (e.g., iphone 6 plus, iphone 7, huawei mate 7, etc.)\n - os: os version id of user mobile phone\n - channel: channel id of mobile ad publisher</p>\n\n<ol>\n<li><p>Are all of the encoding methods one-to-one encoding? For example, every \"device\" of type \"A\" was encoded if and only if it was the same \"device\"?</p></li>\n<li><p>Did the encoding methods destroy \"nearness\" information? For example, IP addresses have specific groupings based on specific allocations of ranges of IP addresses. Are two encoded IP addresses \"near\" each other in the same way the original IP addresses were \"near\" each other?</p></li>\n</ol>",
      "rawMarkdown": "Noticed that some of the data fields are \"encoded\":\n - ip: ip address of click, app: app id for marketing\n - device: device type id of user mobile phone (e.g., iphone 6 plus, iphone 7, huawei mate 7, etc.)\n - os: os version id of user mobile phone\n - channel: channel id of mobile ad publisher\n\n1. Are all of the encoding methods one-to-one encoding? For example, every \"device\" of type \"A\" was encoded if and only if it was the same \"device\"?\n\n2. Did the encoding methods destroy \"nearness\" information? For example, IP addresses have specific groupings based on specific allocations of ranges of IP addresses. Are two encoded IP addresses \"near\" each other in the same way the original IP addresses were \"near\" each other?",
      "votes": null
    },
    {
      "id": "292498",
      "postDate": "03/08/2018 03:28:47",
      "content": "<ol>\n<li>All of them are one-to-one encoding</li>\n<li>IP address did have <strong>a lot of</strong> nearness information. But we prefer kagglers to build a model without nearness information, because in China, IP address is highly unstable, user's IP can vary from time to time.</li>\n</ol>",
      "rawMarkdown": "1. All of them are one-to-one encoding\n2. IP address did have **a lot of** nearness information. But we prefer kagglers to build a model without nearness information, because in China, IP address is highly unstable, user's IP can vary from time to time.",
      "votes": null
    },
    {
      "id": "292537",
      "postDate": "03/08/2018 05:08:44",
      "content": "<p>Here is my interpretation of the dataset and the answers to your questions:</p>\n\n<ol>\n<li>Yes. I believe there is a one-to-one relationship between the actual entity and its encoding integer. </li>\n<li>No. I do not believe the encoding scheme that the data engineering team who produced and published the data to kaggle platform would leave a data leak like that. I guess the ordering of IP encoding integers do not depict the 'nearness' within IP addresses. </li>\n</ol>\n\n<p>Also, I do not think that one user would have only one IP throughout the dataset. Meaning, IP addresses are dynamic and depends on the network provider. Hence it is harder to analyze the dedicated user activity in the dataset as well. </p>\n\n<p>I hope these answer your question. </p>",
      "rawMarkdown": "Here is my interpretation of the dataset and the answers to your questions:\n\n1. Yes. I believe there is a one-to-one relationship between the actual entity and its encoding integer. \n2. No. I do not believe the encoding scheme that the data engineering team who produced and published the data to kaggle platform would leave a data leak like that. I guess the ordering of IP encoding integers do not depict the 'nearness' within IP addresses. \n\nAlso, I do not think that one user would have only one IP throughout the dataset. Meaning, IP addresses are dynamic and depends on the network provider. Hence it is harder to analyze the dedicated user activity in the dataset as well. \n\nI hope these answer your question.",
      "votes": null
    },
    {
      "id": "294601",
      "postDate": "03/12/2018 07:59:07",
      "content": "<p>Thank you for confirming this.</p>",
      "rawMarkdown": "Thank you for confirming this.",
      "votes": null
    },
    {
      "id": "301742",
      "postDate": "03/23/2018 07:21:02",
      "content": "<blockquote>\n  <p>Also, I do not think that one user would have only one IP throughout the dataset. Meaning, IP addresses are dynamic and depends on the network provider. Hence it is harder to analyze the dedicated user activity in the dataset as well. </p>\n</blockquote>\n\n<p>I figure if you picked a very rare combination of <code>(device,os)</code> you might see it appear on multiple ips. (Maybe we should hash <code>device x os</code>). Mind you we only have 2+1 days' worth of data.</p>",
      "rawMarkdown": "&gt; Also, I do not think that one user would have only one IP throughout the dataset. Meaning, IP addresses are dynamic and depends on the network provider. Hence it is harder to analyze the dedicated user activity in the dataset as well. \n\nI figure if you picked a very rare combination of `(device,os)` you might see it appear on multiple ips. (Maybe we should hash `device x os`). Mind you we only have 2+1 days' worth of data.",
      "votes": null
    },
    {
      "id": "306383",
      "postDate": "03/30/2018 10:49:17",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!",
      "votes": null
    },
    {
      "id": "310181",
      "postDate": "04/06/2018 18:22:11",
      "content": "<p>I agree. If we were to find a rare pair of <code>(device,os)</code>, we might see it appearing on multiple IPs, hence confirming my previous argument. However, that depends how you define <code>very rare combination</code>. But since it is a rare occurrence, we would not be able to prove / derive anything from that since that will be statistically insignificant, don't you think. </p>\n\n<p>May be if we had some sort of domain knowledge or set of rules to interpret the encoding that they have done, we can analyze the data more effectively. But again as you mentioned we have 2+1 days worth of data. It is a limitation on its own. </p>",
      "rawMarkdown": "I agree. If we were to find a rare pair of `(device,os)`, we might see it appearing on multiple IPs, hence confirming my previous argument. However, that depends how you define `very rare combination`. But since it is a rare occurrence, we would not be able to prove / derive anything from that since that will be statistically insignificant, don't you think. \n\nMay be if we had some sort of domain knowledge or set of rules to interpret the encoding that they have done, we can analyze the data more effectively. But again as you mentioned we have 2+1 days worth of data. It is a limitation on its own.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 292498,
      "author_name": "aaronyin",
      "author_url": "",
      "post_date": "03/08/2018 03:28:47",
      "content": "<ol>\n<li>All of them are one-to-one encoding</li>\n<li>IP address did have <strong>a lot of</strong> nearness information. But we prefer kagglers to build a model without nearness information, because in China, IP address is highly unstable, user's IP can vary from time to time.</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 294601,
          "author_name": "asparuhhristov",
          "author_url": "",
          "post_date": "03/12/2018 07:59:07",
          "content": "<p>Thank you for confirming this.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 306383,
          "author_name": "nitishsingh41",
          "author_url": "",
          "post_date": "03/30/2018 10:49:17",
          "content": "<p>Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 292537,
      "author_name": "zerocoder",
      "author_url": "",
      "post_date": "03/08/2018 05:08:44",
      "content": "<p>Here is my interpretation of the dataset and the answers to your questions:</p>\n\n<ol>\n<li>Yes. I believe there is a one-to-one relationship between the actual entity and its encoding integer. </li>\n<li>No. I do not believe the encoding scheme that the data engineering team who produced and published the data to kaggle platform would leave a data leak like that. I guess the ordering of IP encoding integers do not depict the 'nearness' within IP addresses. </li>\n</ol>\n\n<p>Also, I do not think that one user would have only one IP throughout the dataset. Meaning, IP addresses are dynamic and depends on the network provider. Hence it is harder to analyze the dedicated user activity in the dataset as well. </p>\n\n<p>I hope these answer your question. </p>",
      "votes": null,
      "replies": [
        {
          "id": 301742,
          "author_name": "smcinerney",
          "author_url": "",
          "post_date": "03/23/2018 07:21:02",
          "content": "<blockquote>\n  <p>Also, I do not think that one user would have only one IP throughout the dataset. Meaning, IP addresses are dynamic and depends on the network provider. Hence it is harder to analyze the dedicated user activity in the dataset as well. </p>\n</blockquote>\n\n<p>I figure if you picked a very rare combination of <code>(device,os)</code> you might see it appear on multiple ips. (Maybe we should hash <code>device x os</code>). Mind you we only have 2+1 days' worth of data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 310181,
          "author_name": "zerocoder",
          "author_url": "",
          "post_date": "04/06/2018 18:22:11",
          "content": "<p>I agree. If we were to find a rare pair of <code>(device,os)</code>, we might see it appearing on multiple IPs, hence confirming my previous argument. However, that depends how you define <code>very rare combination</code>. But since it is a rare occurrence, we would not be able to prove / derive anything from that since that will be statistically insignificant, don't you think. </p>\n\n<p>May be if we had some sort of domain knowledge or set of rules to interpret the encoding that they have done, we can analyze the data more effectively. But again as you mentioned we have 2+1 days worth of data. It is a limitation on its own. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "292407": "Noticed that some of the data fields are \"encoded\":\n - ip: ip address of click, app: app id for marketing\n - device: device type id of user mobile phone (e.g., iphone 6 plus, iphone 7, huawei mate 7, etc.)\n - os: os version id of user mobile phone\n - channel: channel id of mobile ad publisher\n\n1. Are all of the encoding methods one-to-one encoding? For example, every \"device\" of type \"A\" was encoded if and only if it was the same \"device\"?\n\n2. Did the encoding methods destroy \"nearness\" information? For example, IP addresses have specific groupings based on specific allocations of ranges of IP addresses. Are two encoded IP addresses \"near\" each other in the same way the original IP addresses were \"near\" each other?",
    "292498": "1. All of them are one-to-one encoding\n2. IP address did have **a lot of** nearness information. But we prefer kagglers to build a model without nearness information, because in China, IP address is highly unstable, user's IP can vary from time to time.",
    "292537": "Here is my interpretation of the dataset and the answers to your questions:\n\n1. Yes. I believe there is a one-to-one relationship between the actual entity and its encoding integer. \n2. No. I do not believe the encoding scheme that the data engineering team who produced and published the data to kaggle platform would leave a data leak like that. I guess the ordering of IP encoding integers do not depict the 'nearness' within IP addresses. \n\nAlso, I do not think that one user would have only one IP throughout the dataset. Meaning, IP addresses are dynamic and depends on the network provider. Hence it is harder to analyze the dedicated user activity in the dataset as well. \n\nI hope these answer your question.",
    "294601": "Thank you for confirming this.",
    "301742": "&gt; Also, I do not think that one user would have only one IP throughout the dataset. Meaning, IP addresses are dynamic and depends on the network provider. Hence it is harder to analyze the dedicated user activity in the dataset as well. \n\nI figure if you picked a very rare combination of `(device,os)` you might see it appear on multiple ips. (Maybe we should hash `device x os`). Mind you we only have 2+1 days' worth of data.",
    "306383": "Thanks!",
    "310181": "I agree. If we were to find a rare pair of `(device,os)`, we might see it appearing on multiple IPs, hence confirming my previous argument. However, that depends how you define `very rare combination`. But since it is a rare occurrence, we would not be able to prove / derive anything from that since that will be statistically insignificant, don't you think. \n\nMay be if we had some sort of domain knowledge or set of rules to interpret the encoding that they have done, we can analyze the data more effectively. But again as you mentioned we have 2+1 days worth of data. It is a limitation on its own."
  },
  "source": "meta"
}