{
  "id": 583815,
  "title": "[Convent Train Failed] Has anyone else encountered the same issue?",
  "url": "/competitions/waveform-inversion/discussion/583815",
  "author_name": "Ray",
  "post_date": "2025-06-09T15:33:58.846000",
  "votes": 2,
  "comment_count": 18,
  "views": 0,
  "content": "<p>I am trying to reproduce the results of the public scheme: ConvNeXt - Full Resolution Baseline. However, the training stopped early at Epoch 33. I have tried multiple times, and the MAE on the validation set remains around 77 and cannot be reduced further. Here is the information from the last training run.</p>\n<p>2025-06-09 23:20:53,898 Epoch 33:     Train MAE: 85.97     Val MAE: 76.19     Time: 0:18:18     Step: 10801/14376<br>\n2025-06-09 23:21:03,955 Epoch 33:     Train MAE: 89.58     Val MAE: 76.19     Time: 0:18:29     Step: 10901/14376<br>\n2025-06-09 23:21:13,790 Epoch 33:     Train MAE: 90.62     Val MAE: 76.19     Time: 0:18:38     Step: 11001/14376<br>\n2025-06-09 23:21:23,642 Epoch 33:     Train MAE: 86.11     Val MAE: 76.19     Time: 0:18:48     Step: 11101/14376<br>\n2025-06-09 23:21:33,597 Epoch 33:     Train MAE: 83.13     Val MAE: 76.19     Time: 0:18:58     Step: 11201/14376<br>\n2025-06-09 23:21:43,519 Epoch 33:     Train MAE: 84.91     Val MAE: 76.19     Time: 0:19:08     Step: 11301/14376<br>\n2025-06-09 23:21:53,448 Epoch 33:     Train MAE: 77.74     Val MAE: 76.19     Time: 0:19:18     Step: 11401/14376<br>\n2025-06-09 23:22:03,344 Epoch 33:     Train MAE: 87.30     Val MAE: 76.19     Time: 0:19:28     Step: 11501/14376<br>\n2025-06-09 23:22:13,626 Epoch 33:     Train MAE: 84.35     Val MAE: 76.19     Time: 0:19:38     Step: 11601/14376<br>\n2025-06-09 23:22:23,514 Epoch 33:     Train MAE: 81.85     Val MAE: 76.19     Time: 0:19:48     Step: 11701/14376<br>\n2025-06-09 23:22:33,330 Epoch 33:     Train MAE: 84.13     Val MAE: 76.19     Time: 0:19:58     Step: 11801/14376<br>\n2025-06-09 23:22:43,207 Epoch 33:     Train MAE: 89.64     Val MAE: 76.19     Time: 0:20:08     Step: 11901/14376<br>\n2025-06-09 23:22:53,112 Epoch 33:     Train MAE: 91.88     Val MAE: 76.19     Time: 0:20:18     Step: 12001/14376<br>\n2025-06-09 23:23:03,022 Epoch 33:     Train MAE: 91.37     Val MAE: 76.19     Time: 0:20:28     Step: 12101/14376<br>\n2025-06-09 23:23:13,069 Epoch 33:     Train MAE: 85.62     Val MAE: 76.19     Time: 0:20:38     Step: 12201/14376<br>\n2025-06-09 23:23:23,202 Epoch 33:     Train MAE: 87.78     Val MAE: 76.19     Time: 0:20:48     Step: 12301/14376<br>\n2025-06-09 23:23:33,118 Epoch 33:     Train MAE: 87.34     Val MAE: 76.19     Time: 0:20:58     Step: 12401/14376<br>\n2025-06-09 23:23:43,239 Epoch 33:     Train MAE: 84.86     Val MAE: 76.19     Time: 0:21:08     Step: 12501/14376<br>\n2025-06-09 23:23:53,208 Epoch 33:     Train MAE: 83.54     Val MAE: 76.19     Time: 0:21:18     Step: 12601/14376<br>\n2025-06-09 23:24:03,054 Epoch 33:     Train MAE: 104.47     Val MAE: 76.19     Time: 0:21:28     Step: 12701/14376<br>\n2025-06-09 23:24:13,211 Epoch 33:     Train MAE: 99.03     Val MAE: 76.19     Time: 0:21:38     Step: 12801/14376<br>\n2025-06-09 23:24:24,264 Epoch 33:     Train MAE: 87.60     Val MAE: 76.19     Time: 0:21:49     Step: 12901/14376<br>\n2025-06-09 23:24:34,385 Epoch 33:     Train MAE: 83.16     Val MAE: 76.19     Time: 0:21:59     Step: 13001/14376<br>\n2025-06-09 23:24:44,347 Epoch 33:     Train MAE: 82.10     Val MAE: 76.19     Time: 0:22:09     Step: 13101/14376<br>\n2025-06-09 23:24:54,260 Epoch 33:     Train MAE: 82.50     Val MAE: 76.19     Time: 0:22:19     Step: 13201/14376<br>\n2025-06-09 23:25:04,202 Epoch 33:     Train MAE: 81.22     Val MAE: 76.19     Time: 0:22:29     Step: 13301/14376<br>\n2025-06-09 23:25:14,303 Epoch 33:     Train MAE: 82.55     Val MAE: 76.19     Time: 0:22:39     Step: 13401/14376<br>\n2025-06-09 23:25:25,281 Epoch 33:     Train MAE: 95.99     Val MAE: 76.19     Time: 0:22:50     Step: 13501/14376<br>\n2025-06-09 23:25:35,251 Epoch 33:     Train MAE: 86.35     Val MAE: 76.19     Time: 0:23:00     Step: 13601/14376<br>\n2025-06-09 23:25:45,164 Epoch 33:     Train MAE: 86.44     Val MAE: 76.19     Time: 0:23:10     Step: 13701/14376<br>\n2025-06-09 23:25:55,359 Epoch 33:     Train MAE: 83.41     Val MAE: 76.19     Time: 0:23:20     Step: 13801/14376<br>\n2025-06-09 23:26:05,459 Epoch 33:     Train MAE: 91.97     Val MAE: 76.19     Time: 0:23:30     Step: 13901/14376<br>\n2025-06-09 23:26:15,826 Epoch 33:     Train MAE: 93.03     Val MAE: 76.19     Time: 0:23:40     Step: 14001/14376<br>\n2025-06-09 23:26:26,370 Epoch 33:     Train MAE: 86.71     Val MAE: 76.19     Time: 0:23:51     Step: 14101/14376<br>\n2025-06-09 23:26:36,151 Epoch 33:     Train MAE: 85.28     Val MAE: 76.19     Time: 0:24:01     Step: 14201/14376<br>\n2025-06-09 23:26:46,053 Epoch 33:     Train MAE: 81.70     Val MAE: 76.19     Time: 0:24:11     Step: 14301/14376<br>\n100%|█████████████████████████████████████████| 313/313 [00:09&lt;00:00, 33.08it/s]<br>\n2025-06-09 23:27:05,892 Ending training (early_stopping).</p>",
  "messages": [
    {
      "id": 3220609,
      "postDate": "2025-06-09T15:33:58.847Z",
      "content": "<p>I am trying to reproduce the results of the public scheme: ConvNeXt - Full Resolution Baseline. However, the training stopped early at Epoch 33. I have tried multiple times, and the MAE on the validation set remains around 77 and cannot be reduced further. Here is the information from the last training run.</p>\n<p>2025-06-09 23:20:53,898 Epoch 33:     Train MAE: 85.97     Val MAE: 76.19     Time: 0:18:18     Step: 10801/14376<br>\n2025-06-09 23:21:03,955 Epoch 33:     Train MAE: 89.58     Val MAE: 76.19     Time: 0:18:29     Step: 10901/14376<br>\n2025-06-09 23:21:13,790 Epoch 33:     Train MAE: 90.62     Val MAE: 76.19     Time: 0:18:38     Step: 11001/14376<br>\n2025-06-09 23:21:23,642 Epoch 33:     Train MAE: 86.11     Val MAE: 76.19     Time: 0:18:48     Step: 11101/14376<br>\n2025-06-09 23:21:33,597 Epoch 33:     Train MAE: 83.13     Val MAE: 76.19     Time: 0:18:58     Step: 11201/14376<br>\n2025-06-09 23:21:43,519 Epoch 33:     Train MAE: 84.91     Val MAE: 76.19     Time: 0:19:08     Step: 11301/14376<br>\n2025-06-09 23:21:53,448 Epoch 33:     Train MAE: 77.74     Val MAE: 76.19     Time: 0:19:18     Step: 11401/14376<br>\n2025-06-09 23:22:03,344 Epoch 33:     Train MAE: 87.30     Val MAE: 76.19     Time: 0:19:28     Step: 11501/14376<br>\n2025-06-09 23:22:13,626 Epoch 33:     Train MAE: 84.35     Val MAE: 76.19     Time: 0:19:38     Step: 11601/14376<br>\n2025-06-09 23:22:23,514 Epoch 33:     Train MAE: 81.85     Val MAE: 76.19     Time: 0:19:48     Step: 11701/14376<br>\n2025-06-09 23:22:33,330 Epoch 33:     Train MAE: 84.13     Val MAE: 76.19     Time: 0:19:58     Step: 11801/14376<br>\n2025-06-09 23:22:43,207 Epoch 33:     Train MAE: 89.64     Val MAE: 76.19     Time: 0:20:08     Step: 11901/14376<br>\n2025-06-09 23:22:53,112 Epoch 33:     Train MAE: 91.88     Val MAE: 76.19     Time: 0:20:18     Step: 12001/14376<br>\n2025-06-09 23:23:03,022 Epoch 33:     Train MAE: 91.37     Val MAE: 76.19     Time: 0:20:28     Step: 12101/14376<br>\n2025-06-09 23:23:13,069 Epoch 33:     Train MAE: 85.62     Val MAE: 76.19     Time: 0:20:38     Step: 12201/14376<br>\n2025-06-09 23:23:23,202 Epoch 33:     Train MAE: 87.78     Val MAE: 76.19     Time: 0:20:48     Step: 12301/14376<br>\n2025-06-09 23:23:33,118 Epoch 33:     Train MAE: 87.34     Val MAE: 76.19     Time: 0:20:58     Step: 12401/14376<br>\n2025-06-09 23:23:43,239 Epoch 33:     Train MAE: 84.86     Val MAE: 76.19     Time: 0:21:08     Step: 12501/14376<br>\n2025-06-09 23:23:53,208 Epoch 33:     Train MAE: 83.54     Val MAE: 76.19     Time: 0:21:18     Step: 12601/14376<br>\n2025-06-09 23:24:03,054 Epoch 33:     Train MAE: 104.47     Val MAE: 76.19     Time: 0:21:28     Step: 12701/14376<br>\n2025-06-09 23:24:13,211 Epoch 33:     Train MAE: 99.03     Val MAE: 76.19     Time: 0:21:38     Step: 12801/14376<br>\n2025-06-09 23:24:24,264 Epoch 33:     Train MAE: 87.60     Val MAE: 76.19     Time: 0:21:49     Step: 12901/14376<br>\n2025-06-09 23:24:34,385 Epoch 33:     Train MAE: 83.16     Val MAE: 76.19     Time: 0:21:59     Step: 13001/14376<br>\n2025-06-09 23:24:44,347 Epoch 33:     Train MAE: 82.10     Val MAE: 76.19     Time: 0:22:09     Step: 13101/14376<br>\n2025-06-09 23:24:54,260 Epoch 33:     Train MAE: 82.50     Val MAE: 76.19     Time: 0:22:19     Step: 13201/14376<br>\n2025-06-09 23:25:04,202 Epoch 33:     Train MAE: 81.22     Val MAE: 76.19     Time: 0:22:29     Step: 13301/14376<br>\n2025-06-09 23:25:14,303 Epoch 33:     Train MAE: 82.55     Val MAE: 76.19     Time: 0:22:39     Step: 13401/14376<br>\n2025-06-09 23:25:25,281 Epoch 33:     Train MAE: 95.99     Val MAE: 76.19     Time: 0:22:50     Step: 13501/14376<br>\n2025-06-09 23:25:35,251 Epoch 33:     Train MAE: 86.35     Val MAE: 76.19     Time: 0:23:00     Step: 13601/14376<br>\n2025-06-09 23:25:45,164 Epoch 33:     Train MAE: 86.44     Val MAE: 76.19     Time: 0:23:10     Step: 13701/14376<br>\n2025-06-09 23:25:55,359 Epoch 33:     Train MAE: 83.41     Val MAE: 76.19     Time: 0:23:20     Step: 13801/14376<br>\n2025-06-09 23:26:05,459 Epoch 33:     Train MAE: 91.97     Val MAE: 76.19     Time: 0:23:30     Step: 13901/14376<br>\n2025-06-09 23:26:15,826 Epoch 33:     Train MAE: 93.03     Val MAE: 76.19     Time: 0:23:40     Step: 14001/14376<br>\n2025-06-09 23:26:26,370 Epoch 33:     Train MAE: 86.71     Val MAE: 76.19     Time: 0:23:51     Step: 14101/14376<br>\n2025-06-09 23:26:36,151 Epoch 33:     Train MAE: 85.28     Val MAE: 76.19     Time: 0:24:01     Step: 14201/14376<br>\n2025-06-09 23:26:46,053 Epoch 33:     Train MAE: 81.70     Val MAE: 76.19     Time: 0:24:11     Step: 14301/14376<br>\n100%|█████████████████████████████████████████| 313/313 [00:09&lt;00:00, 33.08it/s]<br>\n2025-06-09 23:27:05,892 Ending training (early_stopping).</p>",
      "rawMarkdown": "I am trying to reproduce the results of the public scheme: ConvNeXt - Full Resolution Baseline. However, the training stopped early at Epoch 33. I have tried multiple times, and the MAE on the validation set remains around 77 and cannot be reduced further. Here is the information from the last training run.\n\n2025-06-09 23:20:53,898 Epoch 33:     Train MAE: 85.97     Val MAE: 76.19     Time: 0:18:18     Step: 10801/14376\n2025-06-09 23:21:03,955 Epoch 33:     Train MAE: 89.58     Val MAE: 76.19     Time: 0:18:29     Step: 10901/14376\n2025-06-09 23:21:13,790 Epoch 33:     Train MAE: 90.62     Val MAE: 76.19     Time: 0:18:38     Step: 11001/14376\n2025-06-09 23:21:23,642 Epoch 33:     Train MAE: 86.11     Val MAE: 76.19     Time: 0:18:48     Step: 11101/14376\n2025-06-09 23:21:33,597 Epoch 33:     Train MAE: 83.13     Val MAE: 76.19     Time: 0:18:58     Step: 11201/14376\n2025-06-09 23:21:43,519 Epoch 33:     Train MAE: 84.91     Val MAE: 76.19     Time: 0:19:08     Step: 11301/14376\n2025-06-09 23:21:53,448 Epoch 33:     Train MAE: 77.74     Val MAE: 76.19     Time: 0:19:18     Step: 11401/14376\n2025-06-09 23:22:03,344 Epoch 33:     Train MAE: 87.30     Val MAE: 76.19     Time: 0:19:28     Step: 11501/14376\n2025-06-09 23:22:13,626 Epoch 33:     Train MAE: 84.35     Val MAE: 76.19     Time: 0:19:38     Step: 11601/14376\n2025-06-09 23:22:23,514 Epoch 33:     Train MAE: 81.85     Val MAE: 76.19     Time: 0:19:48     Step: 11701/14376\n2025-06-09 23:22:33,330 Epoch 33:     Train MAE: 84.13     Val MAE: 76.19     Time: 0:19:58     Step: 11801/14376\n2025-06-09 23:22:43,207 Epoch 33:     Train MAE: 89.64     Val MAE: 76.19     Time: 0:20:08     Step: 11901/14376\n2025-06-09 23:22:53,112 Epoch 33:     Train MAE: 91.88     Val MAE: 76.19     Time: 0:20:18     Step: 12001/14376\n2025-06-09 23:23:03,022 Epoch 33:     Train MAE: 91.37     Val MAE: 76.19     Time: 0:20:28     Step: 12101/14376\n2025-06-09 23:23:13,069 Epoch 33:     Train MAE: 85.62     Val MAE: 76.19     Time: 0:20:38     Step: 12201/14376\n2025-06-09 23:23:23,202 Epoch 33:     Train MAE: 87.78     Val MAE: 76.19     Time: 0:20:48     Step: 12301/14376\n2025-06-09 23:23:33,118 Epoch 33:     Train MAE: 87.34     Val MAE: 76.19     Time: 0:20:58     Step: 12401/14376\n2025-06-09 23:23:43,239 Epoch 33:     Train MAE: 84.86     Val MAE: 76.19     Time: 0:21:08     Step: 12501/14376\n2025-06-09 23:23:53,208 Epoch 33:     Train MAE: 83.54     Val MAE: 76.19     Time: 0:21:18     Step: 12601/14376\n2025-06-09 23:24:03,054 Epoch 33:     Train MAE: 104.47     Val MAE: 76.19     Time: 0:21:28     Step: 12701/14376\n2025-06-09 23:24:13,211 Epoch 33:     Train MAE: 99.03     Val MAE: 76.19     Time: 0:21:38     Step: 12801/14376\n2025-06-09 23:24:24,264 Epoch 33:     Train MAE: 87.60     Val MAE: 76.19     Time: 0:21:49     Step: 12901/14376\n2025-06-09 23:24:34,385 Epoch 33:     Train MAE: 83.16     Val MAE: 76.19     Time: 0:21:59     Step: 13001/14376\n2025-06-09 23:24:44,347 Epoch 33:     Train MAE: 82.10     Val MAE: 76.19     Time: 0:22:09     Step: 13101/14376\n2025-06-09 23:24:54,260 Epoch 33:     Train MAE: 82.50     Val MAE: 76.19     Time: 0:22:19     Step: 13201/14376\n2025-06-09 23:25:04,202 Epoch 33:     Train MAE: 81.22     Val MAE: 76.19     Time: 0:22:29     Step: 13301/14376\n2025-06-09 23:25:14,303 Epoch 33:     Train MAE: 82.55     Val MAE: 76.19     Time: 0:22:39     Step: 13401/14376\n2025-06-09 23:25:25,281 Epoch 33:     Train MAE: 95.99     Val MAE: 76.19     Time: 0:22:50     Step: 13501/14376\n2025-06-09 23:25:35,251 Epoch 33:     Train MAE: 86.35     Val MAE: 76.19     Time: 0:23:00     Step: 13601/14376\n2025-06-09 23:25:45,164 Epoch 33:     Train MAE: 86.44     Val MAE: 76.19     Time: 0:23:10     Step: 13701/14376\n2025-06-09 23:25:55,359 Epoch 33:     Train MAE: 83.41     Val MAE: 76.19     Time: 0:23:20     Step: 13801/14376\n2025-06-09 23:26:05,459 Epoch 33:     Train MAE: 91.97     Val MAE: 76.19     Time: 0:23:30     Step: 13901/14376\n2025-06-09 23:26:15,826 Epoch 33:     Train MAE: 93.03     Val MAE: 76.19     Time: 0:23:40     Step: 14001/14376\n2025-06-09 23:26:26,370 Epoch 33:     Train MAE: 86.71     Val MAE: 76.19     Time: 0:23:51     Step: 14101/14376\n2025-06-09 23:26:36,151 Epoch 33:     Train MAE: 85.28     Val MAE: 76.19     Time: 0:24:01     Step: 14201/14376\n2025-06-09 23:26:46,053 Epoch 33:     Train MAE: 81.70     Val MAE: 76.19     Time: 0:24:11     Step: 14301/14376\n100%|█████████████████████████████████████████| 313/313 [00:09<00:00, 33.08it/s]\n2025-06-09 23:27:05,892 Ending training (early_stopping).",
      "votes": 2
    },
    {
      "id": 3221146,
      "postDate": "2025-06-10T13:23:41.520Z",
      "content": "<p>I would like to share the HP where my Convnext experiment produced results close to the published results <strong>[Val MAE: 28.95, LB: 33.2]</strong>. <br>\nI hope it helps you.</p>\n<ul>\n<li>My Convnext Exp Result<ul>\n<li><strong>Val MAE: 30.07</strong></li>\n<li><strong>LB:35.1</strong></li></ul></li>\n</ul>\n<pre><code>\n.epochs = \n.batch_size = \n.batch_size_val = \n.early_stopping = {: , : }\n\n(model.parameters(), lr=e-, weight_decay=e-)\n</code></pre>",
      "rawMarkdown": "I would like to share the HP where my Convnext experiment produced results close to the published results **[Val MAE: 28.95, LB: 33.2]**. \nI hope it helps you.\n\n* My Convnext Exp Result\n    * **Val MAE: 30.07**\n    * **LB:35.1**\n```\n# HP\ncfg.epochs = 150\ncfg.batch_size = 16\ncfg.batch_size_val = 16\ncfg.early_stopping = {\"patience\": 3, \"streak\": 0}\n\nAdamW(model.parameters(), lr=2e-4, weight_decay=2e-5)\n```",
      "replies": [
        {
          "id": 3221203,
          "postDate": "2025-06-10T15:01:32.763Z",
          "content": "<p>Thank you very much for your suggestion. I will give it a try!</p>",
          "rawMarkdown": "Thank you very much for your suggestion. I will give it a try!"
        }
      ]
    },
    {
      "id": 3220635,
      "postDate": "2025-06-09T16:24:30.947Z",
      "content": "<p>Yes, I tried continue training from the public model but no improved. After 25 epochs, the loss on valid set still higher current model.</p>",
      "rawMarkdown": "Yes, I tried continue training from the public model but no improved. After 25 epochs, the loss on valid set still higher current model.",
      "replies": [
        {
          "id": 3220638,
          "postDate": "2025-06-09T16:28:33.250Z",
          "content": "<p>May I ask what was the lowest MAE on the validation set when you trained using this approach?</p>",
          "rawMarkdown": "May I ask what was the lowest MAE on the validation set when you trained using this approach?",
          "replies": [
            {
              "id": 3220665,
              "postDate": "2025-06-09T17:22:07.230Z",
              "content": "<p>Start from 32.63 to 31.24 at epoch 25. The public model is 30.62.</p>\n<p>I replace it with 1 convnext in current best notebook and recieved score are 28.9.</p>",
              "rawMarkdown": "Start from 32.63 to 31.24 at epoch 25. The public model is 30.62.\n\nI replace it with 1 convnext in current best notebook and recieved score are 28.9."
            },
            {
              "id": 3220801,
              "postDate": "2025-06-10T02:06:58.527Z",
              "content": "<p>May I ask if it is one of the two fine-tuned weights you use for subsequent training, or one of the three weights originally written by the author?</p>\n<p>Then is your current score the average of the two models, CAF and your own convnext?</p>",
              "rawMarkdown": "May I ask if it is one of the two fine-tuned weights you use for subsequent training, or one of the three weights originally written by the author?\n\nThen is your current score the average of the two models, CAF and your own convnext?"
            },
            {
              "id": 3220823,
              "postDate": "2025-06-10T03:03:41.470Z",
              "content": "<p>Yes, Yes.</p>\n<p>weight ensemble are 8-1-1 like the public notebook.</p>",
              "rawMarkdown": "Yes, Yes.\n\nweight ensemble are 8-1-1 like the public notebook."
            },
            {
              "id": 3220982,
              "postDate": "2025-06-10T08:22:50.877Z",
              "content": "<p>His initial Val MAE:30.29. Why is it even higher when you continue training</p>",
              "rawMarkdown": "His initial Val MAE:30.29. Why is it even higher when you continue training"
            },
            {
              "id": 3221013,
              "postDate": "2025-06-10T09:19:40.767Z",
              "content": "<blockquote>\n  <p>His initial Val MAE:30.29. Why is it even higher when you continue training  </p>\n</blockquote>\n<p>Probably because you only load model weights.  <br>\nTo properly continue training you need also the optimizer weights.  <br>\nWithout optimizer weights, the phenomena you describe is the expected one.  </p>",
              "rawMarkdown": "> His initial Val MAE:30.29. Why is it even higher when you continue training  \n\nProbably because you only load model weights.  \nTo properly continue training you need also the optimizer weights.  \nWithout optimizer weights, the phenomena you describe is the expected one.  "
            },
            {
              "id": 3221142,
              "postDate": "2025-06-10T13:20:53.300Z",
              "content": "<p>If the parameters such as the optimizer, scheduler, and ema are not saved, will they still definitely increase even if I have preheated the training?</p>",
              "rawMarkdown": "If the parameters such as the optimizer, scheduler, and ema are not saved, will they still definitely increase even if I have preheated the training?"
            }
          ]
        }
      ]
    },
    {
      "id": 3220621,
      "postDate": "2025-06-09T15:52:27.370Z",
      "content": "<blockquote>\n  <p>i'm struggling with nan issues, did you come across nan while training? <a href=\"https://www.kaggle.com/faykudbq\" target=\"_blank\">@faykudbq</a> </p>\n</blockquote>\n<hr>\n<p><a href=\"https://www.kaggle.com/competitions/waveform-inversion/discussion/579841#3206037\" target=\"_blank\">from training logs</a> - maybe we need to increase early stop &gt; 3 ?</p>",
      "rawMarkdown": "> i'm struggling with nan issues, did you come across nan while training? @faykudbq \n\n---\n\n[from training logs](https://www.kaggle.com/competitions/waveform-inversion/discussion/579841#3206037) - maybe we need to increase early stop > 3 ?",
      "replies": [
        {
          "id": 3220623,
          "postDate": "2025-06-09T16:01:42.380Z",
          "content": "<p>Actually, a solution to this problem has already been provided in previous posts. The issue occurs because you are using float16 precision. Switching to bfloat16 or float32 will resolve it. For example:</p>\n<p><code>with autocast(device_type=cfg.device.type, dtype=torch.bfloat16):</code></p>",
          "rawMarkdown": "Actually, a solution to this problem has already been provided in previous posts. The issue occurs because you are using float16 precision. Switching to bfloat16 or float32 will resolve it. For example:\n\n`with autocast(device_type=cfg.device.type, dtype=torch.bfloat16):`",
          "votes": 1,
          "replies": [
            {
              "id": 3220767,
              "postDate": "2025-06-09T22:37:21.530Z",
              "content": "<p>Thanks <a href=\"https://www.kaggle.com/faykudbq\" target=\"_blank\">@faykudbq</a> , my <strong>NaN issue resolved</strong></p>",
              "rawMarkdown": "Thanks @faykudbq , my **NaN issue resolved**"
            },
            {
              "id": 3220783,
              "postDate": "2025-06-10T01:08:39.470Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 3220802,
              "postDate": "2025-06-10T02:08:33.607Z",
              "content": "<p>In the original code, I only modified the learning rate 1e-4. It seems that nan won't appear. Nothing else was changed</p>",
              "rawMarkdown": "In the original code, I only modified the learning rate 1e-4. It seems that nan won't appear. Nothing else was changed"
            },
            {
              "id": 3220816,
              "postDate": "2025-06-10T02:43:10.057Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 3220817,
              "postDate": "2025-06-10T02:43:26.010Z",
              "content": "<p>May I ask, under this setup, what are your training results? What is the MAE on the validation set? Can your training results reach the results claimed by the author? CV MAE: 31.2.</p>",
              "rawMarkdown": "May I ask, under this setup, what are your training results? What is the MAE on the validation set? Can your training results reach the results claimed by the author? CV MAE: 31.2."
            },
            {
              "id": 3220901,
              "postDate": "2025-06-10T05:58:32.940Z",
              "content": "<p>I continued the training under fine-tuning the weights and fine-tuned the model structure at the same time. For a single model, epoch 44 val: 31.07</p>",
              "rawMarkdown": "I continued the training under fine-tuning the weights and fine-tuned the model structure at the same time. For a single model, epoch 44 val: 31.07"
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3221146,
      "author_name": "yukiZ",
      "author_url": "",
      "post_date": "2025-06-10T13:23:41.520000",
      "content": "<p>I would like to share the HP where my Convnext experiment produced results close to the published results <strong>[Val MAE: 28.95, LB: 33.2]</strong>. <br>\nI hope it helps you.</p>\n<ul>\n<li>My Convnext Exp Result<ul>\n<li><strong>Val MAE: 30.07</strong></li>\n<li><strong>LB:35.1</strong></li></ul></li>\n</ul>\n<pre><code>\n.epochs = \n.batch_size = \n.batch_size_val = \n.early_stopping = {: , : }\n\n(model.parameters(), lr=e-, weight_decay=e-)\n</code></pre>",
      "votes": 0,
      "replies": [
        {
          "id": 3221203,
          "author_name": "Ray",
          "author_url": "",
          "post_date": "2025-06-10T15:01:32.763000",
          "content": "<p>Thank you very much for your suggestion. I will give it a try!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3220635,
      "author_name": "Ayaka Kirima",
      "author_url": "",
      "post_date": "2025-06-09T16:24:30.947000",
      "content": "<p>Yes, I tried continue training from the public model but no improved. After 25 epochs, the loss on valid set still higher current model.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3220638,
          "author_name": "Ray",
          "author_url": "",
          "post_date": "2025-06-09T16:28:33.250000",
          "content": "<p>May I ask what was the lowest MAE on the validation set when you trained using this approach?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3220665,
              "author_name": "Ayaka Kirima",
              "author_url": "",
              "post_date": "2025-06-09T17:22:07.230000",
              "content": "<p>Start from 32.63 to 31.24 at epoch 25. The public model is 30.62.</p>\n<p>I replace it with 1 convnext in current best notebook and recieved score are 28.9.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3220801,
              "author_name": "water joe",
              "author_url": "",
              "post_date": "2025-06-10T02:06:58.527000",
              "content": "<p>May I ask if it is one of the two fine-tuned weights you use for subsequent training, or one of the three weights originally written by the author?</p>\n<p>Then is your current score the average of the two models, CAF and your own convnext?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3220823,
              "author_name": "Ayaka Kirima",
              "author_url": "",
              "post_date": "2025-06-10T03:03:41.470000",
              "content": "<p>Yes, Yes.</p>\n<p>weight ensemble are 8-1-1 like the public notebook.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3220982,
              "author_name": "water joe",
              "author_url": "",
              "post_date": "2025-06-10T08:22:50.877000",
              "content": "<p>His initial Val MAE:30.29. Why is it even higher when you continue training</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3221013,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2025-06-10T09:19:40.767000",
              "content": "<blockquote>\n  <p>His initial Val MAE:30.29. Why is it even higher when you continue training  </p>\n</blockquote>\n<p>Probably because you only load model weights.  <br>\nTo properly continue training you need also the optimizer weights.  <br>\nWithout optimizer weights, the phenomena you describe is the expected one.  </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3221142,
              "author_name": "water joe",
              "author_url": "",
              "post_date": "2025-06-10T13:20:53.300000",
              "content": "<p>If the parameters such as the optimizer, scheduler, and ema are not saved, will they still definitely increase even if I have preheated the training?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3220621,
      "author_name": "SeshuRaju 🧘‍♂️",
      "author_url": "",
      "post_date": "2025-06-09T15:52:27.370000",
      "content": "<blockquote>\n  <p>i'm struggling with nan issues, did you come across nan while training? <a href=\"https://www.kaggle.com/faykudbq\" target=\"_blank\">@faykudbq</a> </p>\n</blockquote>\n<hr>\n<p><a href=\"https://www.kaggle.com/competitions/waveform-inversion/discussion/579841#3206037\" target=\"_blank\">from training logs</a> - maybe we need to increase early stop &gt; 3 ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3220623,
          "author_name": "Ray",
          "author_url": "",
          "post_date": "2025-06-09T16:01:42.380000",
          "content": "<p>Actually, a solution to this problem has already been provided in previous posts. The issue occurs because you are using float16 precision. Switching to bfloat16 or float32 will resolve it. For example:</p>\n<p><code>with autocast(device_type=cfg.device.type, dtype=torch.bfloat16):</code></p>",
          "votes": 1,
          "replies": [
            {
              "id": 3220767,
              "author_name": "SeshuRaju 🧘‍♂️",
              "author_url": "",
              "post_date": "2025-06-09T22:37:21.530000",
              "content": "<p>Thanks <a href=\"https://www.kaggle.com/faykudbq\" target=\"_blank\">@faykudbq</a> , my <strong>NaN issue resolved</strong></p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3220783,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-06-10T01:08:39.470000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3220802,
              "author_name": "water joe",
              "author_url": "",
              "post_date": "2025-06-10T02:08:33.607000",
              "content": "<p>In the original code, I only modified the learning rate 1e-4. It seems that nan won't appear. Nothing else was changed</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3220816,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-06-10T02:43:10.057000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3220817,
              "author_name": "Ray",
              "author_url": "",
              "post_date": "2025-06-10T02:43:26.010000",
              "content": "<p>May I ask, under this setup, what are your training results? What is the MAE on the validation set? Can your training results reach the results claimed by the author? CV MAE: 31.2.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3220901,
              "author_name": "water joe",
              "author_url": "",
              "post_date": "2025-06-10T05:58:32.940000",
              "content": "<p>I continued the training under fine-tuning the weights and fine-tuned the model structure at the same time. For a single model, epoch 44 val: 31.07</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3220609": "I am trying to reproduce the results of the public scheme: ConvNeXt - Full Resolution Baseline. However, the training stopped early at Epoch 33. I have tried multiple times, and the MAE on the validation set remains around 77 and cannot be reduced further. Here is the information from the last training run.\n\n2025-06-09 23:20:53,898 Epoch 33:     Train MAE: 85.97     Val MAE: 76.19     Time: 0:18:18     Step: 10801/14376\n2025-06-09 23:21:03,955 Epoch 33:     Train MAE: 89.58     Val MAE: 76.19     Time: 0:18:29     Step: 10901/14376\n2025-06-09 23:21:13,790 Epoch 33:     Train MAE: 90.62     Val MAE: 76.19     Time: 0:18:38     Step: 11001/14376\n2025-06-09 23:21:23,642 Epoch 33:     Train MAE: 86.11     Val MAE: 76.19     Time: 0:18:48     Step: 11101/14376\n2025-06-09 23:21:33,597 Epoch 33:     Train MAE: 83.13     Val MAE: 76.19     Time: 0:18:58     Step: 11201/14376\n2025-06-09 23:21:43,519 Epoch 33:     Train MAE: 84.91     Val MAE: 76.19     Time: 0:19:08     Step: 11301/14376\n2025-06-09 23:21:53,448 Epoch 33:     Train MAE: 77.74     Val MAE: 76.19     Time: 0:19:18     Step: 11401/14376\n2025-06-09 23:22:03,344 Epoch 33:     Train MAE: 87.30     Val MAE: 76.19     Time: 0:19:28     Step: 11501/14376\n2025-06-09 23:22:13,626 Epoch 33:     Train MAE: 84.35     Val MAE: 76.19     Time: 0:19:38     Step: 11601/14376\n2025-06-09 23:22:23,514 Epoch 33:     Train MAE: 81.85     Val MAE: 76.19     Time: 0:19:48     Step: 11701/14376\n2025-06-09 23:22:33,330 Epoch 33:     Train MAE: 84.13     Val MAE: 76.19     Time: 0:19:58     Step: 11801/14376\n2025-06-09 23:22:43,207 Epoch 33:     Train MAE: 89.64     Val MAE: 76.19     Time: 0:20:08     Step: 11901/14376\n2025-06-09 23:22:53,112 Epoch 33:     Train MAE: 91.88     Val MAE: 76.19     Time: 0:20:18     Step: 12001/14376\n2025-06-09 23:23:03,022 Epoch 33:     Train MAE: 91.37     Val MAE: 76.19     Time: 0:20:28     Step: 12101/14376\n2025-06-09 23:23:13,069 Epoch 33:     Train MAE: 85.62     Val MAE: 76.19     Time: 0:20:38     Step: 12201/14376\n2025-06-09 23:23:23,202 Epoch 33:     Train MAE: 87.78     Val MAE: 76.19     Time: 0:20:48     Step: 12301/14376\n2025-06-09 23:23:33,118 Epoch 33:     Train MAE: 87.34     Val MAE: 76.19     Time: 0:20:58     Step: 12401/14376\n2025-06-09 23:23:43,239 Epoch 33:     Train MAE: 84.86     Val MAE: 76.19     Time: 0:21:08     Step: 12501/14376\n2025-06-09 23:23:53,208 Epoch 33:     Train MAE: 83.54     Val MAE: 76.19     Time: 0:21:18     Step: 12601/14376\n2025-06-09 23:24:03,054 Epoch 33:     Train MAE: 104.47     Val MAE: 76.19     Time: 0:21:28     Step: 12701/14376\n2025-06-09 23:24:13,211 Epoch 33:     Train MAE: 99.03     Val MAE: 76.19     Time: 0:21:38     Step: 12801/14376\n2025-06-09 23:24:24,264 Epoch 33:     Train MAE: 87.60     Val MAE: 76.19     Time: 0:21:49     Step: 12901/14376\n2025-06-09 23:24:34,385 Epoch 33:     Train MAE: 83.16     Val MAE: 76.19     Time: 0:21:59     Step: 13001/14376\n2025-06-09 23:24:44,347 Epoch 33:     Train MAE: 82.10     Val MAE: 76.19     Time: 0:22:09     Step: 13101/14376\n2025-06-09 23:24:54,260 Epoch 33:     Train MAE: 82.50     Val MAE: 76.19     Time: 0:22:19     Step: 13201/14376\n2025-06-09 23:25:04,202 Epoch 33:     Train MAE: 81.22     Val MAE: 76.19     Time: 0:22:29     Step: 13301/14376\n2025-06-09 23:25:14,303 Epoch 33:     Train MAE: 82.55     Val MAE: 76.19     Time: 0:22:39     Step: 13401/14376\n2025-06-09 23:25:25,281 Epoch 33:     Train MAE: 95.99     Val MAE: 76.19     Time: 0:22:50     Step: 13501/14376\n2025-06-09 23:25:35,251 Epoch 33:     Train MAE: 86.35     Val MAE: 76.19     Time: 0:23:00     Step: 13601/14376\n2025-06-09 23:25:45,164 Epoch 33:     Train MAE: 86.44     Val MAE: 76.19     Time: 0:23:10     Step: 13701/14376\n2025-06-09 23:25:55,359 Epoch 33:     Train MAE: 83.41     Val MAE: 76.19     Time: 0:23:20     Step: 13801/14376\n2025-06-09 23:26:05,459 Epoch 33:     Train MAE: 91.97     Val MAE: 76.19     Time: 0:23:30     Step: 13901/14376\n2025-06-09 23:26:15,826 Epoch 33:     Train MAE: 93.03     Val MAE: 76.19     Time: 0:23:40     Step: 14001/14376\n2025-06-09 23:26:26,370 Epoch 33:     Train MAE: 86.71     Val MAE: 76.19     Time: 0:23:51     Step: 14101/14376\n2025-06-09 23:26:36,151 Epoch 33:     Train MAE: 85.28     Val MAE: 76.19     Time: 0:24:01     Step: 14201/14376\n2025-06-09 23:26:46,053 Epoch 33:     Train MAE: 81.70     Val MAE: 76.19     Time: 0:24:11     Step: 14301/14376\n100%|█████████████████████████████████████████| 313/313 [00:09<00:00, 33.08it/s]\n2025-06-09 23:27:05,892 Ending training (early_stopping).",
    "3221146": "I would like to share the HP where my Convnext experiment produced results close to the published results **[Val MAE: 28.95, LB: 33.2]**. \nI hope it helps you.\n\n* My Convnext Exp Result\n    * **Val MAE: 30.07**\n    * **LB:35.1**\n```\n# HP\ncfg.epochs = 150\ncfg.batch_size = 16\ncfg.batch_size_val = 16\ncfg.early_stopping = {\"patience\": 3, \"streak\": 0}\n\nAdamW(model.parameters(), lr=2e-4, weight_decay=2e-5)\n```",
    "3220635": "Yes, I tried continue training from the public model but no improved. After 25 epochs, the loss on valid set still higher current model.",
    "3220621": "> i'm struggling with nan issues, did you come across nan while training? @faykudbq \n\n---\n\n[from training logs](https://www.kaggle.com/competitions/waveform-inversion/discussion/579841#3206037) - maybe we need to increase early stop > 3 ?"
  }
}