{
  "id": 53136,
  "title": "IP as a predictor or not?",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/53136",
  "author_name": "",
  "post_date": "2018-03-27T13:38:41.011297Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi all,\nI tried to run some deep learning models and wanted to clarify a few things:</p>\n\n<p>When I don't drop IP and use it as a feature, the model overfits and reaches a training accuracy of 0.9974. It outputs all 0s and hence I get a AUC score of 0.5. </p>\n\n<p>On the other hand, I saw some gradient boosting kernels using IP as one of the features. In fact, it helped them improve the LB score.</p>\n\n<p>Does anyone have an idea why the 'IP' feature is behaving this way?</p>",
  "messages": [
    {
      "id": "304377",
      "postDate": "03/27/2018 13:38:41",
      "content": "<p>Hi all,\nI tried to run some deep learning models and wanted to clarify a few things:</p>\n\n<p>When I don't drop IP and use it as a feature, the model overfits and reaches a training accuracy of 0.9974. It outputs all 0s and hence I get a AUC score of 0.5. </p>\n\n<p>On the other hand, I saw some gradient boosting kernels using IP as one of the features. In fact, it helped them improve the LB score.</p>\n\n<p>Does anyone have an idea why the 'IP' feature is behaving this way?</p>",
      "rawMarkdown": "Hi all,\nI tried to run some deep learning models and wanted to clarify a few things:\n\nWhen I don't drop IP and use it as a feature, the model overfits and reaches a training accuracy of 0.9974. It outputs all 0s and hence I get a AUC score of 0.5. \n\nOn the other hand, I saw some gradient boosting kernels using IP as one of the features. In fact, it helped them improve the LB score.\n\nDoes anyone have an idea why the 'IP' feature is behaving this way?",
      "votes": null
    },
    {
      "id": "304637",
      "postDate": "03/27/2018 20:19:30",
      "content": "<p>There are several discussion posts and kernels on this topic, most of which have \"IP\" in the title.  The main conclusion is that IP might be a useful feature but only if you treat it as purely categorical and make sure that your model is not using the implicit ordering in the numeric values of the codes.  Otherwise it will overfit.</p>",
      "rawMarkdown": "There are several discussion posts and kernels on this topic, most of which have \"IP\" in the title.  The main conclusion is that IP might be a useful feature but only if you treat it as purely categorical and make sure that your model is not using the implicit ordering in the numeric values of the codes.  Otherwise it will overfit.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 304637,
      "author_name": "aharless",
      "author_url": "",
      "post_date": "03/27/2018 20:19:30",
      "content": "<p>There are several discussion posts and kernels on this topic, most of which have \"IP\" in the title.  The main conclusion is that IP might be a useful feature but only if you treat it as purely categorical and make sure that your model is not using the implicit ordering in the numeric values of the codes.  Otherwise it will overfit.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "304377": "Hi all,\nI tried to run some deep learning models and wanted to clarify a few things:\n\nWhen I don't drop IP and use it as a feature, the model overfits and reaches a training accuracy of 0.9974. It outputs all 0s and hence I get a AUC score of 0.5. \n\nOn the other hand, I saw some gradient boosting kernels using IP as one of the features. In fact, it helped them improve the LB score.\n\nDoes anyone have an idea why the 'IP' feature is behaving this way?",
    "304637": "There are several discussion posts and kernels on this topic, most of which have \"IP\" in the title.  The main conclusion is that IP might be a useful feature but only if you treat it as purely categorical and make sure that your model is not using the implicit ordering in the numeric values of the codes.  Otherwise it will overfit."
  },
  "source": "meta"
}