{
  "id": 53299,
  "title": "does correlation effect?",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/53299",
  "author_name": "",
  "post_date": "2018-03-29T06:29:07.852067500Z",
  "votes": null,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I took ip and ip count feature and it gave better prediction than the model with just ip count feature. I was wondering as ip and ip count feature are dependent on each other then shouldn't it decrease my model accuracy?</p>",
  "messages": [
    {
      "id": "305610",
      "postDate": "03/29/2018 06:29:07",
      "content": "<p>I took ip and ip count feature and it gave better prediction than the model with just ip count feature. I was wondering as ip and ip count feature are dependent on each other then shouldn't it decrease my model accuracy?</p>",
      "rawMarkdown": "I took ip and ip count feature and it gave better prediction than the model with just ip count feature. I was wondering as ip and ip count feature are dependent on each other then shouldn't it decrease my model accuracy?",
      "votes": null
    },
    {
      "id": "305629",
      "postDate": "03/29/2018 07:01:05",
      "content": "<p>You might wanna look into <a href=\"https://www.kaggle.com/yuliagm/be-careful-about-ips-as-a-signal\">this</a>, <a href=\"https://www.kaggle.com/its7171/ip-encoding-looks-ok-data-is-clean\">this</a> and <a href=\"https://www.kaggle.com/cpmpml/ip-download-rates\">this</a> kernel to know more about why not to use IP as a feature.</p>",
      "rawMarkdown": "You might wanna look into [this][1], [this][2] and [this][3] kernel to know more about why not to use IP as a feature.\n\n\n  [1]: https://www.kaggle.com/yuliagm/be-careful-about-ips-as-a-signal\n  [2]: https://www.kaggle.com/its7171/ip-encoding-looks-ok-data-is-clean\n  [3]: https://www.kaggle.com/cpmpml/ip-download-rates",
      "votes": null
    },
    {
      "id": "305675",
      "postDate": "03/29/2018 08:55:16",
      "content": "<p>I meant how to decide if we should add an encoded column or replace the original column?</p>",
      "rawMarkdown": "I meant how to decide if we should add an encoded column or replace the original column?",
      "votes": null
    },
    {
      "id": "305901",
      "postDate": "03/29/2018 15:39:22",
      "content": "<p>If you're using a tree-based algorithm (like LightGBM or XGBoost), then there is no problem in general with using more than one encoding of a high cardinality feature, because the algorithm will split them in different places.  However, in the particular case of IP, the original coding provided in these data (see @SohaibOmar's links) will overfit the training set.</p>",
      "rawMarkdown": "If you're using a tree-based algorithm (like LightGBM or XGBoost), then there is no problem in general with using more than one encoding of a high cardinality feature, because the algorithm will split them in different places.  However, in the particular case of IP, the original coding provided in these data (see @SohaibOmar's links) will overfit the training set.",
      "votes": null
    },
    {
      "id": "305920",
      "postDate": "03/29/2018 16:10:17",
      "content": "<p>so how can we reduce this overfitting?</p>",
      "rawMarkdown": "so how can we reduce this overfitting?",
      "votes": null
    },
    {
      "id": "306006",
      "postDate": "03/29/2018 18:21:46",
      "content": "<p>Just don't use the original coding for IP. You can try various other codings (such as frequency encoding, which is what the IP count feature does), either individually or in combination.</p>",
      "rawMarkdown": "Just don't use the original coding for IP. You can try various other codings (such as frequency encoding, which is what the IP count feature does), either individually or in combination.",
      "votes": null
    },
    {
      "id": "306025",
      "postDate": "03/29/2018 19:01:15",
      "content": "<p>why will it overfit?</p>",
      "rawMarkdown": "why will it overfit?",
      "votes": null
    },
    {
      "id": "306035",
      "postDate": "03/29/2018 19:16:54",
      "content": "<p>Because many of the IPs that don't appear in the test set have higher numeric codes and are more likely to download. If you use the original coding for IP, it will misattribute download rates to IP code range when the real issue is whether a particular IP is new or not.  See the links from @SohaibOmar above (including the comments).</p>",
      "rawMarkdown": "Because many of the IPs that don't appear in the test set have higher numeric codes and are more likely to download. If you use the original coding for IP, it will misattribute download rates to IP code range when the real issue is whether a particular IP is new or not.  See the links from @SohaibOmar above (including the comments).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 305629,
      "author_name": "sohaibomar",
      "author_url": "",
      "post_date": "03/29/2018 07:01:05",
      "content": "<p>You might wanna look into <a href=\"https://www.kaggle.com/yuliagm/be-careful-about-ips-as-a-signal\">this</a>, <a href=\"https://www.kaggle.com/its7171/ip-encoding-looks-ok-data-is-clean\">this</a> and <a href=\"https://www.kaggle.com/cpmpml/ip-download-rates\">this</a> kernel to know more about why not to use IP as a feature.</p>",
      "votes": null,
      "replies": [
        {
          "id": 305675,
          "author_name": "nitishsingh41",
          "author_url": "",
          "post_date": "03/29/2018 08:55:16",
          "content": "<p>I meant how to decide if we should add an encoded column or replace the original column?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 305901,
      "author_name": "aharless",
      "author_url": "",
      "post_date": "03/29/2018 15:39:22",
      "content": "<p>If you're using a tree-based algorithm (like LightGBM or XGBoost), then there is no problem in general with using more than one encoding of a high cardinality feature, because the algorithm will split them in different places.  However, in the particular case of IP, the original coding provided in these data (see @SohaibOmar's links) will overfit the training set.</p>",
      "votes": null,
      "replies": [
        {
          "id": 305920,
          "author_name": "nitishsingh41",
          "author_url": "",
          "post_date": "03/29/2018 16:10:17",
          "content": "<p>so how can we reduce this overfitting?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 306006,
          "author_name": "aharless",
          "author_url": "",
          "post_date": "03/29/2018 18:21:46",
          "content": "<p>Just don't use the original coding for IP. You can try various other codings (such as frequency encoding, which is what the IP count feature does), either individually or in combination.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 306025,
          "author_name": "nitishsingh41",
          "author_url": "",
          "post_date": "03/29/2018 19:01:15",
          "content": "<p>why will it overfit?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 306035,
          "author_name": "aharless",
          "author_url": "",
          "post_date": "03/29/2018 19:16:54",
          "content": "<p>Because many of the IPs that don't appear in the test set have higher numeric codes and are more likely to download. If you use the original coding for IP, it will misattribute download rates to IP code range when the real issue is whether a particular IP is new or not.  See the links from @SohaibOmar above (including the comments).</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "305610": "I took ip and ip count feature and it gave better prediction than the model with just ip count feature. I was wondering as ip and ip count feature are dependent on each other then shouldn't it decrease my model accuracy?",
    "305629": "You might wanna look into [this][1], [this][2] and [this][3] kernel to know more about why not to use IP as a feature.\n\n\n  [1]: https://www.kaggle.com/yuliagm/be-careful-about-ips-as-a-signal\n  [2]: https://www.kaggle.com/its7171/ip-encoding-looks-ok-data-is-clean\n  [3]: https://www.kaggle.com/cpmpml/ip-download-rates",
    "305675": "I meant how to decide if we should add an encoded column or replace the original column?",
    "305901": "If you're using a tree-based algorithm (like LightGBM or XGBoost), then there is no problem in general with using more than one encoding of a high cardinality feature, because the algorithm will split them in different places.  However, in the particular case of IP, the original coding provided in these data (see @SohaibOmar's links) will overfit the training set.",
    "305920": "so how can we reduce this overfitting?",
    "306006": "Just don't use the original coding for IP. You can try various other codings (such as frequency encoding, which is what the IP count feature does), either individually or in combination.",
    "306025": "why will it overfit?",
    "306035": "Because many of the IPs that don't appear in the test set have higher numeric codes and are more likely to download. If you use the original coding for IP, it will misattribute download rates to IP code range when the real issue is whether a particular IP is new or not.  See the links from @SohaibOmar above (including the comments)."
  },
  "source": "meta"
}