{
  "id": 586575,
  "title": "Divergence of the val_loss",
  "url": "/competitions/waveform-inversion/discussion/586575",
  "author_name": "",
  "post_date": "2025-06-27T08:26:27.967452200Z",
  "votes": 3,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi everyone<br>\nWhen I tried training bartley's models (caformer + convnext) i experienced divergence for val_loss after 2-3 epochs (it keeps increasing after each epoch, e.g. 500 --&gt; 610 --&gt; 830 --&gt;…). I increased the batch size to 96 instead of 16 but results are almost same. The training loss is decreasing very well, but not the validation one. <br>\nDo u know what the problem could be? tried both bfloat16 and float32. Tried manipulating the learning rate scheduler. Tried even training on the full data (train+val) and still i get bad score for val??</p>",
  "messages": [
    {
      "id": "3233844",
      "postDate": "06/27/2025 08:26:27",
      "content": "<p>Hi everyone<br>\nWhen I tried training bartley's models (caformer + convnext) i experienced divergence for val_loss after 2-3 epochs (it keeps increasing after each epoch, e.g. 500 --&gt; 610 --&gt; 830 --&gt;…). I increased the batch size to 96 instead of 16 but results are almost same. The training loss is decreasing very well, but not the validation one. <br>\nDo u know what the problem could be? tried both bfloat16 and float32. Tried manipulating the learning rate scheduler. Tried even training on the full data (train+val) and still i get bad score for val??</p>",
      "rawMarkdown": "Hi everyone\nWhen I tried training bartley's models (caformer + convnext) i experienced divergence for val_loss after 2-3 epochs (it keeps increasing after each epoch, e.g. 500 --> 610 --> 830 -->...). I increased the batch size to 96 instead of 16 but results are almost same. The training loss is decreasing very well, but not the validation one. \nDo u know what the problem could be? tried both bfloat16 and float32. Tried manipulating the learning rate scheduler. Tried even training on the full data (train+val) and still i get bad score for val??",
      "votes": null
    },
    {
      "id": "3233875",
      "postDate": "06/27/2025 09:13:59",
      "content": "<p>Generally, train loss decreasing, val metric increasing, is a sign of overfitting to the train set.</p>\n<p>But the numbers you mention are too high to start with. Plus, I am not sure if anyone on any reasonable compute was able to overfit to this massive train set.</p>\n<p>My best guess would be that you might have an issue with how you process the validation set. Meaning, any normalization you do to train, you should do to val, and for calculating the loss you might need to reverse the normalization you apply before feeding data to your model.</p>\n<p>But really hard to say.</p>\n<p>Also, are you training on the full train set? Meaning, the full train fold of the train set? It is quite unlikely you are overfitting to the train set, but if you are iterating over a small portion of the train set, say 200 examples, you might see the symptoms you mention.</p>\n<p>If you still don't see anything that would stand out, maybe grab the entirety of your code and ask the best ChatGPT llm you can put your hands on 🙂 I was surprised to what extent they are capable of dissecting and discussing deep learning code.</p>\n<p>Either way, best of luck in getting to the root of this!</p>",
      "rawMarkdown": "Generally, train loss decreasing, val metric increasing, is a sign of overfitting to the train set.\n\nBut the numbers you mention are too high to start with. Plus, I am not sure if anyone on any reasonable compute was able to overfit to this massive train set.\n\nMy best guess would be that you might have an issue with how you process the validation set. Meaning, any normalization you do to train, you should do to val, and for calculating the loss you might need to reverse the normalization you apply before feeding data to your model.\n\nBut really hard to say.\n\nAlso, are you training on the full train set? Meaning, the full train fold of the train set? It is quite unlikely you are overfitting to the train set, but if you are iterating over a small portion of the train set, say 200 examples, you might see the symptoms you mention.\n\nIf you still don't see anything that would stand out, maybe grab the entirety of your code and ask the best ChatGPT llm you can put your hands on 🙂 I was surprised to what extent they are capable of dissecting and discussing deep learning code.\n\nEither way, best of luck in getting to the root of this!",
      "votes": null
    },
    {
      "id": "3234048",
      "postDate": "06/27/2025 13:05:20",
      "content": "<p>You should train in fp32 or bf16. Thee is a comment about it in the notebook.<br>\nIf you still have divergence then decrease the learning rate maybe.</p>",
      "rawMarkdown": "You should train in fp32 or bf16. Thee is a comment about it in the notebook.\nIf you still have divergence then decrease the learning rate maybe.",
      "votes": null
    },
    {
      "id": "3234060",
      "postDate": "06/27/2025 13:25:15",
      "content": "<p><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> <br>\nThanks all, it worked properly after i disabled EMA. Not sure why it causes problems though?</p>",
      "rawMarkdown": "radek1 @cpmpml \nThanks all, it worked properly after i disabled EMA. Not sure why it causes problems though?",
      "votes": null
    },
    {
      "id": "3234154",
      "postDate": "06/27/2025 15:02:12",
      "content": "<p>It means you do not update the ema model correctly.</p>",
      "rawMarkdown": "It means you do not update the ema model correctly.",
      "votes": null
    },
    {
      "id": "3234426",
      "postDate": "06/27/2025 19:51:08",
      "content": "<p>Thanks for the help guys! We have 3 days for the final push! Good thing our goal was not sub 15 this time 😂</p>",
      "rawMarkdown": "Thanks for the help guys! We have 3 days for the final push! Good thing our goal was not sub 15 this time 😂",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3233875,
      "author_name": "radek1",
      "author_url": "",
      "post_date": "06/27/2025 09:13:59",
      "content": "<p>Generally, train loss decreasing, val metric increasing, is a sign of overfitting to the train set.</p>\n<p>But the numbers you mention are too high to start with. Plus, I am not sure if anyone on any reasonable compute was able to overfit to this massive train set.</p>\n<p>My best guess would be that you might have an issue with how you process the validation set. Meaning, any normalization you do to train, you should do to val, and for calculating the loss you might need to reverse the normalization you apply before feeding data to your model.</p>\n<p>But really hard to say.</p>\n<p>Also, are you training on the full train set? Meaning, the full train fold of the train set? It is quite unlikely you are overfitting to the train set, but if you are iterating over a small portion of the train set, say 200 examples, you might see the symptoms you mention.</p>\n<p>If you still don't see anything that would stand out, maybe grab the entirety of your code and ask the best ChatGPT llm you can put your hands on 🙂 I was surprised to what extent they are capable of dissecting and discussing deep learning code.</p>\n<p>Either way, best of luck in getting to the root of this!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3234048,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "06/27/2025 13:05:20",
      "content": "<p>You should train in fp32 or bf16. Thee is a comment about it in the notebook.<br>\nIf you still have divergence then decrease the learning rate maybe.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3234060,
      "author_name": "mohammad2012191",
      "author_url": "",
      "post_date": "06/27/2025 13:25:15",
      "content": "<p><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> <br>\nThanks all, it worked properly after i disabled EMA. Not sure why it causes problems though?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3234154,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/27/2025 15:02:12",
          "content": "<p>It means you do not update the ema model correctly.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3234426,
              "author_name": "cody11null",
              "author_url": "",
              "post_date": "06/27/2025 19:51:08",
              "content": "<p>Thanks for the help guys! We have 3 days for the final push! Good thing our goal was not sub 15 this time 😂</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3233844": "Hi everyone\nWhen I tried training bartley's models (caformer + convnext) i experienced divergence for val_loss after 2-3 epochs (it keeps increasing after each epoch, e.g. 500 --> 610 --> 830 -->...). I increased the batch size to 96 instead of 16 but results are almost same. The training loss is decreasing very well, but not the validation one. \nDo u know what the problem could be? tried both bfloat16 and float32. Tried manipulating the learning rate scheduler. Tried even training on the full data (train+val) and still i get bad score for val??",
    "3233875": "Generally, train loss decreasing, val metric increasing, is a sign of overfitting to the train set.\n\nBut the numbers you mention are too high to start with. Plus, I am not sure if anyone on any reasonable compute was able to overfit to this massive train set.\n\nMy best guess would be that you might have an issue with how you process the validation set. Meaning, any normalization you do to train, you should do to val, and for calculating the loss you might need to reverse the normalization you apply before feeding data to your model.\n\nBut really hard to say.\n\nAlso, are you training on the full train set? Meaning, the full train fold of the train set? It is quite unlikely you are overfitting to the train set, but if you are iterating over a small portion of the train set, say 200 examples, you might see the symptoms you mention.\n\nIf you still don't see anything that would stand out, maybe grab the entirety of your code and ask the best ChatGPT llm you can put your hands on 🙂 I was surprised to what extent they are capable of dissecting and discussing deep learning code.\n\nEither way, best of luck in getting to the root of this!",
    "3234048": "You should train in fp32 or bf16. Thee is a comment about it in the notebook.\nIf you still have divergence then decrease the learning rate maybe.",
    "3234060": "radek1 @cpmpml \nThanks all, it worked properly after i disabled EMA. Not sure why it causes problems though?",
    "3234154": "It means you do not update the ema model correctly.",
    "3234426": "Thanks for the help guys! We have 3 days for the final push! Good thing our goal was not sub 15 this time 😂"
  },
  "source": "meta"
}