{
  "id": 41266,
  "title": "why the test loss is much worse than validation loss?",
  "url": "/competitions/kkbox-churn-prediction-challenge/discussion/41266",
  "author_name": "",
  "post_date": "2017-10-15T15:01:49.990864300Z",
  "votes": 3,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I propose two model A and B, where A model performs about 0.05 and model B performs 0.11 in the validation data, but the test loss is 0.23 in model A while 0.20 in model B. Can someone explain why the losses between validation and test data can be much different. May be something wrong in the data processing?</p>",
  "messages": [
    {
      "id": "231641",
      "postDate": "10/15/2017 15:01:49",
      "content": "<p>I propose two model A and B, where A model performs about 0.05 and model B performs 0.11 in the validation data, but the test loss is 0.23 in model A while 0.20 in model B. Can someone explain why the losses between validation and test data can be much different. May be something wrong in the data processing?</p>",
      "rawMarkdown": "I propose two model A and B, where A model performs about 0.05 and model B performs 0.11 in the validation data, but the test loss is 0.23 in model A while 0.20 in model B. Can someone explain why the losses between validation and test data can be much different. May be something wrong in the data processing?",
      "votes": null
    },
    {
      "id": "233404",
      "postDate": "10/20/2017 05:13:08",
      "content": "<p>i seen similar behavior in the facebook fraud detection when using IP address from former fraudster as a feature.  in validation it was awesome, on the leader board it was terrible, because fraudsters ended up having to use a new IP when they were banned.  If you check the mean of your prediction before you submit, I bet its much lower than y_train.</p>\n\n<p>I am having the same issue btw, haven't figured out which features are causing this</p>",
      "rawMarkdown": "i seen similar behavior in the facebook fraud detection when using IP address from former fraudster as a feature.  in validation it was awesome, on the leader board it was terrible, because fraudsters ended up having to use a new IP when they were banned.  If you check the mean of your prediction before you submit, I bet its much lower than y_train.\n\nI am having the same issue btw, haven't figured out which features are causing this",
      "votes": null
    },
    {
      "id": "239647",
      "postDate": "11/04/2017 01:55:49",
      "content": "<p>I may be wrong here, but I believe you may be overfitting the validation data.  This can be mitigated by  a technique called early stopping.</p>",
      "rawMarkdown": "I may be wrong here, but I believe you may be overfitting the validation data.  This can be mitigated by  a technique called early stopping.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 233404,
      "author_name": "autokad",
      "author_url": "",
      "post_date": "10/20/2017 05:13:08",
      "content": "<p>i seen similar behavior in the facebook fraud detection when using IP address from former fraudster as a feature.  in validation it was awesome, on the leader board it was terrible, because fraudsters ended up having to use a new IP when they were banned.  If you check the mean of your prediction before you submit, I bet its much lower than y_train.</p>\n\n<p>I am having the same issue btw, haven't figured out which features are causing this</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 239647,
      "author_name": "stravinsky",
      "author_url": "",
      "post_date": "11/04/2017 01:55:49",
      "content": "<p>I may be wrong here, but I believe you may be overfitting the validation data.  This can be mitigated by  a technique called early stopping.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "231641": "I propose two model A and B, where A model performs about 0.05 and model B performs 0.11 in the validation data, but the test loss is 0.23 in model A while 0.20 in model B. Can someone explain why the losses between validation and test data can be much different. May be something wrong in the data processing?",
    "233404": "i seen similar behavior in the facebook fraud detection when using IP address from former fraudster as a feature.  in validation it was awesome, on the leader board it was terrible, because fraudsters ended up having to use a new IP when they were banned.  If you check the mean of your prediction before you submit, I bet its much lower than y_train.\n\nI am having the same issue btw, haven't figured out which features are causing this",
    "239647": "I may be wrong here, but I believe you may be overfitting the validation data.  This can be mitigated by  a technique called early stopping."
  },
  "source": "meta"
}