{
  "id": 107461,
  "title": "What generally causes 'val_loss' to be lower than 'train_loss'?",
  "url": "/competitions/understanding_cloud_organization/discussion/107461",
  "author_name": "",
  "post_date": "2019-09-04T13:48:04.855067400Z",
  "votes": 3,
  "comment_count": 15,
  "views": 0,
  "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3177784%2F2941d3b8c87597eb119ae96a68f3b9c0%2F20190904144701.png?generation=1567604875942707&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "617798",
      "postDate": "09/04/2019 13:48:04",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3177784%2F2941d3b8c87597eb119ae96a68f3b9c0%2F20190904144701.png?generation=1567604875942707&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3177784%2F2941d3b8c87597eb119ae96a68f3b9c0%2F20190904144701.png?generation=1567604875942707&amp;alt=media)",
      "votes": null
    },
    {
      "id": "617809",
      "postDate": "09/04/2019 13:54:19",
      "content": "<p>In this case, the model gives worse results. Is it called underfitting, overfitting or somewhat unknown fitting?</p>",
      "rawMarkdown": "In this case, the model gives worse results. Is it called underfitting, overfitting or somewhat unknown fitting?",
      "votes": null
    },
    {
      "id": "617866",
      "postDate": "09/04/2019 14:47:23",
      "content": "<p>as from andrew ng's course if i can remember correctly it calls underfitting when the graph looks like yours and when the train and valid curve produces large gap between them or the validation curve is quickly fluctuating means we have high variance,actually i can strongly remember andrew ng said \"it is very difficult to understand overfitting and underfitting sometimes\" but from our curve it looks like underfitting,correct me if i am wrong!!! thanks</p>",
      "rawMarkdown": "as from andrew ng's course if i can remember correctly it calls underfitting when the graph looks like yours and when the train and valid curve produces large gap between them or the validation curve is quickly fluctuating means we have high variance,actually i can strongly remember andrew ng said \"it is very difficult to understand overfitting and underfitting sometimes\" but from our curve it looks like underfitting,correct me if i am wrong!!! thanks",
      "votes": null
    },
    {
      "id": "617869",
      "postDate": "09/04/2019 14:53:43",
      "content": "<p>for better understanding you can download my <a href=\"https://github.com/mobassir94/Machine-Learning-Note-Khata-\">handwritten notes</a> collected from andew ng's course,i had to spend 77 days to complete that coursera machine learning course and i strongly believe i plotted such graphs(like yours) in my note-khata by hearing lectures of andrew ng,i still remember he said it's underfitting,you can re check for further clarification</p>",
      "rawMarkdown": "for better understanding you can download my [handwritten notes](https://github.com/mobassir94/Machine-Learning-Note-Khata-) collected from andew ng's course,i had to spend 77 days to complete that coursera machine learning course and i strongly believe i plotted such graphs(like yours) in my note-khata by hearing lectures of andrew ng,i still remember he said it's underfitting,you can re check for further clarification",
      "votes": null
    },
    {
      "id": "617878",
      "postDate": "09/04/2019 15:05:30",
      "content": "<p>Thank you very much for your link! </p>",
      "rawMarkdown": "Thank you very much for your link!",
      "votes": null
    },
    {
      "id": "618037",
      "postDate": "09/04/2019 18:49:34",
      "content": "<p>I guess the validation set may contain some pattern that the model starts to learn at the end, this pattern exists a lot in the validation set but not in the training set. or the validation is too small to validate the model's performance.</p>",
      "rawMarkdown": "I guess the validation set may contain some pattern that the model starts to learn at the end, this pattern exists a lot in the validation set but not in the training set. or the validation is too small to validate the model's performance.",
      "votes": null
    },
    {
      "id": "618042",
      "postDate": "09/04/2019 18:54:20",
      "content": "<p>i guess  the validation is too small to validate the model's performance is the right answer!!! thank you <a href=\"/mariammohamed\">@mariammohamed</a> </p>",
      "rawMarkdown": "i guess  the validation is too small to validate the model's performance is the right answer!!! thank you @mariammohamed",
      "votes": null
    },
    {
      "id": "618739",
      "postDate": "09/05/2019 12:54:24",
      "content": "<p>If the validation loss is like this, how do you decide the number of epochs?</p>",
      "rawMarkdown": "If the validation loss is like this, how do you decide the number of epochs?",
      "votes": null
    },
    {
      "id": "618772",
      "postDate": "09/05/2019 13:33:22",
      "content": "<p>I use <code>earlystopping</code> callback. If the <code>val_dice_coef</code> is not increasing, the model will stop training.</p>",
      "rawMarkdown": "I use `earlystopping` callback. If the `val_dice_coef` is not increasing, the model will stop training.",
      "votes": null
    },
    {
      "id": "618788",
      "postDate": "09/05/2019 13:46:49",
      "content": "<p>I mean in the early steps, my validation loss goes up and down between a range of values, so I can not decide when to stop, it does not converge or go parallel to the training line.</p>",
      "rawMarkdown": "I mean in the early steps, my validation loss goes up and down between a range of values, so I can not decide when to stop, it does not converge or go parallel to the training line.",
      "votes": null
    },
    {
      "id": "618790",
      "postDate": "09/05/2019 13:51:32",
      "content": "<p>Perhaps this is because you are using a large learning rate or a small batch size.</p>",
      "rawMarkdown": "Perhaps this is because you are using a large learning rate or a small batch size.",
      "votes": null
    },
    {
      "id": "618801",
      "postDate": "09/05/2019 14:02:16",
      "content": "<p>yes, that is true, thanks.</p>",
      "rawMarkdown": "yes, that is true, thanks.",
      "votes": null
    },
    {
      "id": "620217",
      "postDate": "09/07/2019 06:53:40",
      "content": "<p><a href=\"/gogo827jz\">@gogo827jz</a> Do you do a lot of heavy augmentation? This sometimes indicate that the training data is much more difficult to learn than that of val, might also require removing early stopping</p>",
      "rawMarkdown": "gogo827jz Do you do a lot of heavy augmentation? This sometimes indicate that the training data is much more difficult to learn than that of val, might also require removing early stopping",
      "votes": null
    },
    {
      "id": "620363",
      "postDate": "09/07/2019 11:30:09",
      "content": "<p>Thanks. I think the augmentation is the reason why train loss is a little bit higher. Not a lot, this figure comes from my public kernel Version 9.</p>",
      "rawMarkdown": "Thanks. I think the augmentation is the reason why train loss is a little bit higher. Not a lot, this figure comes from my public kernel Version 9.",
      "votes": null
    },
    {
      "id": "620566",
      "postDate": "09/07/2019 17:20:55",
      "content": "<p>I have been ranting in couple of other discussion posts about how poor the resolution of the measurement system is - poor definition of what is a fish, variation between folks creating the \"truth\" - while ranting I forgot that the result of having a data set with large variation in the truth is that you need to increase the size of the validation data.  In the real world one would hope to increase the number number of evaluations to allow for larger validation set.  Here we have to balance training needs with validation need.  </p>\n\n<p>So thanks for the reminder - increasing the size of the validation set in all my models.</p>",
      "rawMarkdown": "I have been ranting in couple of other discussion posts about how poor the resolution of the measurement system is - poor definition of what is a fish, variation between folks creating the \"truth\" - while ranting I forgot that the result of having a data set with large variation in the truth is that you need to increase the size of the validation data.  In the real world one would hope to increase the number number of evaluations to allow for larger validation set.  Here we have to balance training needs with validation need.  \n\nSo thanks for the reminder - increasing the size of the validation set in all my models.",
      "votes": null
    },
    {
      "id": "620632",
      "postDate": "09/07/2019 18:55:05",
      "content": "<p>I see, thanks for making it a public kernel</p>",
      "rawMarkdown": "I see, thanks for making it a public kernel",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 617809,
      "author_name": "gogo827jz",
      "author_url": "",
      "post_date": "09/04/2019 13:54:19",
      "content": "<p>In this case, the model gives worse results. Is it called underfitting, overfitting or somewhat unknown fitting?</p>",
      "votes": null,
      "replies": [
        {
          "id": 617866,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "09/04/2019 14:47:23",
          "content": "<p>as from andrew ng's course if i can remember correctly it calls underfitting when the graph looks like yours and when the train and valid curve produces large gap between them or the validation curve is quickly fluctuating means we have high variance,actually i can strongly remember andrew ng said \"it is very difficult to understand overfitting and underfitting sometimes\" but from our curve it looks like underfitting,correct me if i am wrong!!! thanks</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 617869,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "09/04/2019 14:53:43",
          "content": "<p>for better understanding you can download my <a href=\"https://github.com/mobassir94/Machine-Learning-Note-Khata-\">handwritten notes</a> collected from andew ng's course,i had to spend 77 days to complete that coursera machine learning course and i strongly believe i plotted such graphs(like yours) in my note-khata by hearing lectures of andrew ng,i still remember he said it's underfitting,you can re check for further clarification</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 617878,
          "author_name": "gogo827jz",
          "author_url": "",
          "post_date": "09/04/2019 15:05:30",
          "content": "<p>Thank you very much for your link! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 620217,
          "author_name": "joonl04",
          "author_url": "",
          "post_date": "09/07/2019 06:53:40",
          "content": "<p><a href=\"/gogo827jz\">@gogo827jz</a> Do you do a lot of heavy augmentation? This sometimes indicate that the training data is much more difficult to learn than that of val, might also require removing early stopping</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 620363,
          "author_name": "gogo827jz",
          "author_url": "",
          "post_date": "09/07/2019 11:30:09",
          "content": "<p>Thanks. I think the augmentation is the reason why train loss is a little bit higher. Not a lot, this figure comes from my public kernel Version 9.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 620632,
          "author_name": "joonl04",
          "author_url": "",
          "post_date": "09/07/2019 18:55:05",
          "content": "<p>I see, thanks for making it a public kernel</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 618037,
      "author_name": "mariammohamed",
      "author_url": "",
      "post_date": "09/04/2019 18:49:34",
      "content": "<p>I guess the validation set may contain some pattern that the model starts to learn at the end, this pattern exists a lot in the validation set but not in the training set. or the validation is too small to validate the model's performance.</p>",
      "votes": null,
      "replies": [
        {
          "id": 618042,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "09/04/2019 18:54:20",
          "content": "<p>i guess  the validation is too small to validate the model's performance is the right answer!!! thank you <a href=\"/mariammohamed\">@mariammohamed</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 620566,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "09/07/2019 17:20:55",
          "content": "<p>I have been ranting in couple of other discussion posts about how poor the resolution of the measurement system is - poor definition of what is a fish, variation between folks creating the \"truth\" - while ranting I forgot that the result of having a data set with large variation in the truth is that you need to increase the size of the validation data.  In the real world one would hope to increase the number number of evaluations to allow for larger validation set.  Here we have to balance training needs with validation need.  </p>\n\n<p>So thanks for the reminder - increasing the size of the validation set in all my models.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 618739,
      "author_name": "mariammohamed",
      "author_url": "",
      "post_date": "09/05/2019 12:54:24",
      "content": "<p>If the validation loss is like this, how do you decide the number of epochs?</p>",
      "votes": null,
      "replies": [
        {
          "id": 618772,
          "author_name": "gogo827jz",
          "author_url": "",
          "post_date": "09/05/2019 13:33:22",
          "content": "<p>I use <code>earlystopping</code> callback. If the <code>val_dice_coef</code> is not increasing, the model will stop training.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 618788,
          "author_name": "mariammohamed",
          "author_url": "",
          "post_date": "09/05/2019 13:46:49",
          "content": "<p>I mean in the early steps, my validation loss goes up and down between a range of values, so I can not decide when to stop, it does not converge or go parallel to the training line.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 618790,
          "author_name": "gogo827jz",
          "author_url": "",
          "post_date": "09/05/2019 13:51:32",
          "content": "<p>Perhaps this is because you are using a large learning rate or a small batch size.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 618801,
          "author_name": "mariammohamed",
          "author_url": "",
          "post_date": "09/05/2019 14:02:16",
          "content": "<p>yes, that is true, thanks.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "617798": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3177784%2F2941d3b8c87597eb119ae96a68f3b9c0%2F20190904144701.png?generation=1567604875942707&amp;alt=media)",
    "617809": "In this case, the model gives worse results. Is it called underfitting, overfitting or somewhat unknown fitting?",
    "617866": "as from andrew ng's course if i can remember correctly it calls underfitting when the graph looks like yours and when the train and valid curve produces large gap between them or the validation curve is quickly fluctuating means we have high variance,actually i can strongly remember andrew ng said \"it is very difficult to understand overfitting and underfitting sometimes\" but from our curve it looks like underfitting,correct me if i am wrong!!! thanks",
    "617869": "for better understanding you can download my [handwritten notes](https://github.com/mobassir94/Machine-Learning-Note-Khata-) collected from andew ng's course,i had to spend 77 days to complete that coursera machine learning course and i strongly believe i plotted such graphs(like yours) in my note-khata by hearing lectures of andrew ng,i still remember he said it's underfitting,you can re check for further clarification",
    "617878": "Thank you very much for your link!",
    "618037": "I guess the validation set may contain some pattern that the model starts to learn at the end, this pattern exists a lot in the validation set but not in the training set. or the validation is too small to validate the model's performance.",
    "618042": "i guess  the validation is too small to validate the model's performance is the right answer!!! thank you @mariammohamed",
    "618739": "If the validation loss is like this, how do you decide the number of epochs?",
    "618772": "I use `earlystopping` callback. If the `val_dice_coef` is not increasing, the model will stop training.",
    "618788": "I mean in the early steps, my validation loss goes up and down between a range of values, so I can not decide when to stop, it does not converge or go parallel to the training line.",
    "618790": "Perhaps this is because you are using a large learning rate or a small batch size.",
    "618801": "yes, that is true, thanks.",
    "620217": "gogo827jz Do you do a lot of heavy augmentation? This sometimes indicate that the training data is much more difficult to learn than that of val, might also require removing early stopping",
    "620363": "Thanks. I think the augmentation is the reason why train loss is a little bit higher. Not a lot, this figure comes from my public kernel Version 9.",
    "620566": "I have been ranting in couple of other discussion posts about how poor the resolution of the measurement system is - poor definition of what is a fish, variation between folks creating the \"truth\" - while ranting I forgot that the result of having a data set with large variation in the truth is that you need to increase the size of the validation data.  In the real world one would hope to increase the number number of evaluations to allow for larger validation set.  Here we have to balance training needs with validation need.  \n\nSo thanks for the reminder - increasing the size of the validation set in all my models.",
    "620632": "I see, thanks for making it a public kernel"
  },
  "source": "meta"
}