{
  "id": 33752,
  "title": "Validation loss and Test loss completely out of sync",
  "url": "/competitions/intel-mobileodt-cervical-cancer-screening/discussion/33752",
  "author_name": "",
  "post_date": "2017-05-29T04:45:03.390629700Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I am struggling to debug my simple CNN model (3 conv2D and 2 Dense). Validation loss on given data goes down pretty good while training (saturates at 0.5 or less after 50 epochs). However, when results on test data is submitted on kaggle site, the error is far off the value (2-4). It seems my model is overfitting but I don't understand why will it behave so differently on two subsets of same dataset (given that test dataset belongs to the same dataset). Any pointers to it will be greatly helpful. Besides, has anyone faced this issue as well? </p>",
  "messages": [
    {
      "id": "186718",
      "postDate": "05/29/2017 04:45:03",
      "content": "<p>I am struggling to debug my simple CNN model (3 conv2D and 2 Dense). Validation loss on given data goes down pretty good while training (saturates at 0.5 or less after 50 epochs). However, when results on test data is submitted on kaggle site, the error is far off the value (2-4). It seems my model is overfitting but I don't understand why will it behave so differently on two subsets of same dataset (given that test dataset belongs to the same dataset). Any pointers to it will be greatly helpful. Besides, has anyone faced this issue as well? </p>",
      "rawMarkdown": "I am struggling to debug my simple CNN model (3 conv2D and 2 Dense). Validation loss on given data goes down pretty good while training (saturates at 0.5 or less after 50 epochs). However, when results on test data is submitted on kaggle site, the error is far off the value (2-4). It seems my model is overfitting but I don't understand why will it behave so differently on two subsets of same dataset (given that test dataset belongs to the same dataset). Any pointers to it will be greatly helpful. Besides, has anyone faced this issue as well?",
      "votes": null
    },
    {
      "id": "186740",
      "postDate": "05/29/2017 06:41:23",
      "content": "<p>Do you have a local CV? Did you train on the whole dataset or did you leave a part of it out for validation? So after the 50 epochs, is it the <strong>training error</strong> that is 0.5, or is it the error on the validation part?</p>",
      "rawMarkdown": "Do you have a local CV? Did you train on the whole dataset or did you leave a part of it out for validation? So after the 50 epochs, is it the **training error** that is 0.5, or is it the error on the validation part?",
      "votes": null
    },
    {
      "id": "186779",
      "postDate": "05/29/2017 08:49:10",
      "content": "<p>My mistake. I was augmenting data by adding more replicas of same data, and then splitting data into train and validation set. Obviously, my train set had data from validation set as well (due to previous replication) and therefore was converging very well. No so after fixing the error. :-( </p>",
      "rawMarkdown": "My mistake. I was augmenting data by adding more replicas of same data, and then splitting data into train and validation set. Obviously, my train set had data from validation set as well (due to previous replication) and therefore was converging very well. No so after fixing the error. :-(",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 186740,
      "author_name": "oysteijo",
      "author_url": "",
      "post_date": "05/29/2017 06:41:23",
      "content": "<p>Do you have a local CV? Did you train on the whole dataset or did you leave a part of it out for validation? So after the 50 epochs, is it the <strong>training error</strong> that is 0.5, or is it the error on the validation part?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 186779,
      "author_name": "sbharti",
      "author_url": "",
      "post_date": "05/29/2017 08:49:10",
      "content": "<p>My mistake. I was augmenting data by adding more replicas of same data, and then splitting data into train and validation set. Obviously, my train set had data from validation set as well (due to previous replication) and therefore was converging very well. No so after fixing the error. :-( </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "186718": "I am struggling to debug my simple CNN model (3 conv2D and 2 Dense). Validation loss on given data goes down pretty good while training (saturates at 0.5 or less after 50 epochs). However, when results on test data is submitted on kaggle site, the error is far off the value (2-4). It seems my model is overfitting but I don't understand why will it behave so differently on two subsets of same dataset (given that test dataset belongs to the same dataset). Any pointers to it will be greatly helpful. Besides, has anyone faced this issue as well?",
    "186740": "Do you have a local CV? Did you train on the whole dataset or did you leave a part of it out for validation? So after the 50 epochs, is it the **training error** that is 0.5, or is it the error on the validation part?",
    "186779": "My mistake. I was augmenting data by adding more replicas of same data, and then splitting data into train and validation set. Obviously, my train set had data from validation set as well (due to previous replication) and therefore was converging very well. No so after fixing the error. :-("
  },
  "source": "meta"
}