{
  "id": 481558,
  "title": "Fluctuation in validation loss",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/481558",
  "author_name": "",
  "post_date": "2024-03-04T05:54:14.208911Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi all, <br>\nI'm encountering a fluctuating validation loss in my model training. I'm not really sure about why there is such a sudden rise and fall. The following is the log of my model training. As you can see in epoch 8, there is a sharp increase in validation loss from 8.4387 to 13.2108, accompanied by a high KL divergence value of 23.0834. Could it be due to the high learning rate or the anomalies in data? By the way, I'm using StratifiedGroupKFold with n_splits = 5. Many thanks.</p>\n<pre><code>Epoch : LearningRateScheduler setting learning rate to .\nEpoch /\n/ [==============================] - 223s 520ms/step - loss:  - kl_divergence_categorical:  - val_loss:  - val_kl_divergence_categorical:  - lr: \n\nEpoch : LearningRateScheduler setting learning rate to .\nEpoch /\n/ [==============================] - 223s 521ms/step - loss:  - kl_divergence_categorical:  - val_loss:  - val_kl_divergence_categorical:  - lr: \n\nEpoch : LearningRateScheduler setting learning rate to .\nEpoch /\n/ [==============================] - 223s 522ms/step - loss:  - kl_divergence_categorical:  - val_loss:  - val_kl_divergence_categorical:  - lr: \n\nEpoch : LearningRateScheduler setting learning rate to .\nEpoch /\n/ [==============================] - 224s 524ms/step - loss:  - kl_divergence_categorical:  - val_loss:  - val_kl_divergence_categorical:  - lr: \n</code></pre>",
  "messages": [
    {
      "id": "2680373",
      "postDate": "03/04/2024 05:54:14",
      "content": "<p>Hi all, <br>\nI'm encountering a fluctuating validation loss in my model training. I'm not really sure about why there is such a sudden rise and fall. The following is the log of my model training. As you can see in epoch 8, there is a sharp increase in validation loss from 8.4387 to 13.2108, accompanied by a high KL divergence value of 23.0834. Could it be due to the high learning rate or the anomalies in data? By the way, I'm using StratifiedGroupKFold with n_splits = 5. Many thanks.</p>\n<pre><code>Epoch : LearningRateScheduler setting learning rate to .\nEpoch /\n/ [==============================] - 223s 520ms/step - loss:  - kl_divergence_categorical:  - val_loss:  - val_kl_divergence_categorical:  - lr: \n\nEpoch : LearningRateScheduler setting learning rate to .\nEpoch /\n/ [==============================] - 223s 521ms/step - loss:  - kl_divergence_categorical:  - val_loss:  - val_kl_divergence_categorical:  - lr: \n\nEpoch : LearningRateScheduler setting learning rate to .\nEpoch /\n/ [==============================] - 223s 522ms/step - loss:  - kl_divergence_categorical:  - val_loss:  - val_kl_divergence_categorical:  - lr: \n\nEpoch : LearningRateScheduler setting learning rate to .\nEpoch /\n/ [==============================] - 224s 524ms/step - loss:  - kl_divergence_categorical:  - val_loss:  - val_kl_divergence_categorical:  - lr: \n</code></pre>",
      "rawMarkdown": "Hi all, \nI'm encountering a fluctuating validation loss in my model training. I'm not really sure about why there is such a sudden rise and fall. The following is the log of my model training. As you can see in epoch 8, there is a sharp increase in validation loss from 8.4387 to 13.2108, accompanied by a high KL divergence value of 23.0834. Could it be due to the high learning rate or the anomalies in data? By the way, I'm using StratifiedGroupKFold with n_splits = 5. Many thanks.\n\n```python\nEpoch 6: LearningRateScheduler setting learning rate to 0.0009251834592969424.\nEpoch 6/20\n428/428 [==============================] - 223s 520ms/step - loss: 4.8026 - kl_divergence_categorical: 1.2343 - val_loss: 5.6506 - val_kl_divergence_categorical: 2.1202 - lr: 9.2518e-04\n\nEpoch 7: LearningRateScheduler setting learning rate to 0.0008696349541517193.\nEpoch 7/20\n428/428 [==============================] - 223s 521ms/step - loss: 4.3777 - kl_divergence_categorical: 1.2163 - val_loss: 4.7210 - val_kl_divergence_categorical: 1.4947 - lr: 8.6963e-04\n\nEpoch 8: LearningRateScheduler setting learning rate to 0.0008015160008714387.\nEpoch 8/20\n428/428 [==============================] - 223s 522ms/step - loss: 3.9699 - kl_divergence_categorical: 1.1799 - val_loss: 13.2108 - val_kl_divergence_categorical: 23.0834 - lr: 8.0152e-04\n\nEpoch 9: LearningRateScheduler setting learning rate to 0.0007231463087103811.\nEpoch 9/20\n428/428 [==============================] - 224s 524ms/step - loss: 3.5515 - kl_divergence_categorical: 1.1558 - val_loss: 8.4387 - val_kl_divergence_categorical: 6.0689 - lr: 7.2315e-04\n```",
      "votes": null
    },
    {
      "id": "2680482",
      "postDate": "03/04/2024 08:12:51",
      "content": "<p>Hi. I've experienced very fluctuation starting trainings. In my case I think because I was using the same dataset random generator for training and validation. Once I fixed it for validation has been reduced for the training end but still with some weird peaks, again, at start. I think is because the sensibility of log in KL toguether with some statistical barriers along the way. In my case starts with a fast convergenge to 1. (I think equivalent to 1/6 vote everywhere). From here the model starts to discriminate majority classes with more effort and some features works better on training than in validation. Is not till this first plateau is passed that validation stabilizes. Often I see some more plateaus along the way with the same behavior. Reduction on training convergence with validation fluctuations.<br>\nAnd in my experience I obtain best results with higher lr and fluctuant training starts that with lower lr and smoother training.</p>",
      "rawMarkdown": "Hi. I've experienced very fluctuation starting trainings. In my case I think because I was using the same dataset random generator for training and validation. Once I fixed it for validation has been reduced for the training end but still with some weird peaks, again, at start. I think is because the sensibility of log in KL toguether with some statistical barriers along the way. In my case starts with a fast convergenge to 1. (I think equivalent to 1/6 vote everywhere). From here the model starts to discriminate majority classes with more effort and some features works better on training than in validation. Is not till this first plateau is passed that validation stabilizes. Often I see some more plateaus along the way with the same behavior. Reduction on training convergence with validation fluctuations.\nAnd in my experience I obtain best results with higher lr and fluctuant training starts that with lower lr and smoother training.",
      "votes": null
    },
    {
      "id": "2683424",
      "postDate": "03/06/2024 02:42:59",
      "content": "<p>I will try using the higher learning rate. Thanks</p>",
      "rawMarkdown": "I will try using the higher learning rate. Thanks",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2680482,
      "author_name": "sacuscreed",
      "author_url": "",
      "post_date": "03/04/2024 08:12:51",
      "content": "<p>Hi. I've experienced very fluctuation starting trainings. In my case I think because I was using the same dataset random generator for training and validation. Once I fixed it for validation has been reduced for the training end but still with some weird peaks, again, at start. I think is because the sensibility of log in KL toguether with some statistical barriers along the way. In my case starts with a fast convergenge to 1. (I think equivalent to 1/6 vote everywhere). From here the model starts to discriminate majority classes with more effort and some features works better on training than in validation. Is not till this first plateau is passed that validation stabilizes. Often I see some more plateaus along the way with the same behavior. Reduction on training convergence with validation fluctuations.<br>\nAnd in my experience I obtain best results with higher lr and fluctuant training starts that with lower lr and smoother training.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2683424,
          "author_name": "kaungmyatkyaw",
          "author_url": "",
          "post_date": "03/06/2024 02:42:59",
          "content": "<p>I will try using the higher learning rate. Thanks</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2680373": "Hi all, \nI'm encountering a fluctuating validation loss in my model training. I'm not really sure about why there is such a sudden rise and fall. The following is the log of my model training. As you can see in epoch 8, there is a sharp increase in validation loss from 8.4387 to 13.2108, accompanied by a high KL divergence value of 23.0834. Could it be due to the high learning rate or the anomalies in data? By the way, I'm using StratifiedGroupKFold with n_splits = 5. Many thanks.\n\n```python\nEpoch 6: LearningRateScheduler setting learning rate to 0.0009251834592969424.\nEpoch 6/20\n428/428 [==============================] - 223s 520ms/step - loss: 4.8026 - kl_divergence_categorical: 1.2343 - val_loss: 5.6506 - val_kl_divergence_categorical: 2.1202 - lr: 9.2518e-04\n\nEpoch 7: LearningRateScheduler setting learning rate to 0.0008696349541517193.\nEpoch 7/20\n428/428 [==============================] - 223s 521ms/step - loss: 4.3777 - kl_divergence_categorical: 1.2163 - val_loss: 4.7210 - val_kl_divergence_categorical: 1.4947 - lr: 8.6963e-04\n\nEpoch 8: LearningRateScheduler setting learning rate to 0.0008015160008714387.\nEpoch 8/20\n428/428 [==============================] - 223s 522ms/step - loss: 3.9699 - kl_divergence_categorical: 1.1799 - val_loss: 13.2108 - val_kl_divergence_categorical: 23.0834 - lr: 8.0152e-04\n\nEpoch 9: LearningRateScheduler setting learning rate to 0.0007231463087103811.\nEpoch 9/20\n428/428 [==============================] - 224s 524ms/step - loss: 3.5515 - kl_divergence_categorical: 1.1558 - val_loss: 8.4387 - val_kl_divergence_categorical: 6.0689 - lr: 7.2315e-04\n```",
    "2680482": "Hi. I've experienced very fluctuation starting trainings. In my case I think because I was using the same dataset random generator for training and validation. Once I fixed it for validation has been reduced for the training end but still with some weird peaks, again, at start. I think is because the sensibility of log in KL toguether with some statistical barriers along the way. In my case starts with a fast convergenge to 1. (I think equivalent to 1/6 vote everywhere). From here the model starts to discriminate majority classes with more effort and some features works better on training than in validation. Is not till this first plateau is passed that validation stabilizes. Often I see some more plateaus along the way with the same behavior. Reduction on training convergence with validation fluctuations.\nAnd in my experience I obtain best results with higher lr and fluctuant training starts that with lower lr and smoother training.",
    "2683424": "I will try using the higher learning rate. Thanks"
  },
  "source": "meta"
}