{
  "id": 583392,
  "title": "Why does fine-tuning make mae larger",
  "url": "/competitions/waveform-inversion/discussion/583392",
  "author_name": "water joe",
  "post_date": "2025-06-06T14:17:44.663000",
  "votes": 2,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Epoch 4:     Train MAE: 22.33     Val MAE: 30.67     Time: 00:00:21     Step: 1/451<br>\nEpoch 4:     Train MAE: 21.83     Val MAE: 30.67     Time: 00:08:20     Step: 101/451<br>\nEpoch 4:     Train MAE: 21.77     Val MAE: 30.67     Time: 00:16:18     Step: 201/451<br>\nEpoch 4:     Train MAE: 21.89     Val MAE: 30.67     Time: 00:24:24     Step: 301/451<br>\nEpoch 4:     Train MAE: 21.71     Val MAE: 30.67     Time: 00:32:23     Step: 401/451<br>\n100%|█████████████████████████████████████████| 313/313 [01:08&lt;00:00,  4.58it/s]<br>\nEnding training (early_stopping).</p>\n<h1>my last round of training. The val here is very small, but why does the following test become larger? Please ask everyone🤔</h1>\n<p>CurveFault_A 720.66<br>\nCurveFault_B 751.34Why does fine-tuning make mae larger<br>\nCurveVel_A 650.84<br>\nCurveVel_B 749.36<br>\nFlatFault_A 722.53<br>\nFlatFault_B 732.67<br>\nFlatVel_A 658.38<br>\nFlatVel_B 751.96<br>\nStyle_A 532.39<br>\nStyle_B 488.19<br>\n= = = = = = = = = = = = = = = = = = = = = = = = =<br>\nVal MAE: 675.83<br>\n= = = = = = = = = = = = = = = = = = = = = = = = =The backbone was frozen, and then only a joint training was added</p>",
  "messages": [
    {
      "id": 3219225,
      "postDate": "2025-06-07T10:34:45.013Z",
      "content": "<p>It looks you use a model without fine tuning. Do you really use the output of your fine tuning/</p>",
      "rawMarkdown": "It looks you use a model without fine tuning. Do you really use the output of your fine tuning/",
      "votes": 2,
      "replies": [
        {
          "id": 3219228,
          "postDate": "2025-06-07T10:45:27.750Z",
          "content": "<p>Thank you for your reply. This is my way of load the network，</p>\n<p>`import glob<br>\nimport torch<br>\nimport torch.nn as nn<br>\nimport torch.nn.functional as F<br>\nfrom _cfg import cfg<br>\nfrom _model import Net, EnsembleModel</p>\n<p>if RUN_VALID or RUN_TEST:<br>\n    # Load pretrained models<br>\n    models = []<br>\n    m = Net(<br>\n                backbone=\"convnext_small.fb_in22k_ft_in1k\",<br>\n                pretrained=False,<br>\n            )<br>\n    state_dict= torch.load('/kaggle/working/best_model_42.pt', map_location=cfg.device, weights_only=True)<br>\n    models.append(m)<br>\n    # Combine<br>\n    model = EnsembleModel(models)<br>\n    model = model.to(cfg.device)<br>\n    model = model.eval()<br>\n    print(\"n_models: {:_}\".format(len(models)))`</p>\n<p>The following is the same as the previous notebook ConvNeXt.</p>",
          "rawMarkdown": "Thank you for your reply. This is my way of load the network，\n\n`import glob\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\nfrom _cfg import cfg\nfrom _model import Net, EnsembleModel\n\nif RUN_VALID or RUN_TEST:\n    # Load pretrained models\n    models = []\n    m = Net(\n                backbone=\"convnext_small.fb_in22k_ft_in1k\",\n                pretrained=False,\n            )\n    state_dict= torch.load('/kaggle/working/best_model_42.pt', map_location=cfg.device, weights_only=True)\n    models.append(m)\n    # Combine\n    model = EnsembleModel(models)\n    model = model.to(cfg.device)\n    model = model.eval()\n    print(\"n_models: {:_}\".format(len(models)))`\n\nThe following is the same as the previous notebook ConvNeXt.",
          "replies": [
            {
              "id": 3219238,
              "postDate": "2025-06-07T11:00:51.067Z",
              "content": "<p>why do you use ensemble model given you have only one model?</p>",
              "rawMarkdown": "why do you use ensemble model given you have only one model?",
              "votes": 1
            },
            {
              "id": 3219243,
              "postDate": "2025-06-07T11:09:00.263Z",
              "content": "<p>Yes, of course your answer is correct. I think but this code should be able to handle the output of a single model. It's just a little efficiency issue。</p>\n<p>`class EnsembleModel(nn.Module):<br>\n    def <strong>init</strong>(self, models):<br>\n        super().<strong>init</strong>()<br>\n        self.models = nn.ModuleList(models).eval()</p>\n<pre><code> ():\n    output = \n     m  .models:\n        out = m(x)  \n         (out, ):  \n            logits = out[]\n        : \n            logits = out\n\n         output  :\n            output = logits\n        :\n            output += logits\n    output /= (.models)\n     output,   `\n</code></pre>",
              "rawMarkdown": "Yes, of course your answer is correct. I think but this code should be able to handle the output of a single model. It's just a little efficiency issue。\n\n`class EnsembleModel(nn.Module):\n    def __init__(self, models):\n        super().__init__()\n        self.models = nn.ModuleList(models).eval()\n            \n    def forward(self, x):\n        output = None\n        for m in self.models:\n            out = m(x)  \n            if isinstance(out, tuple):  \n                logits = out[0]\n            else: \n                logits = out\n            \n            if output is None:\n                output = logits\n            else:\n                output += logits\n        output /= len(self.models)\n        return output, None  `"
            },
            {
              "id": 3219271,
              "postDate": "2025-06-07T11:57:14.513Z",
              "content": "<p><a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/waterjoe/convnext-crossentropyloss</a></p>\n<p>I would be most grateful if you had time to take a look at my notebook</p>",
              "rawMarkdown": "[https://www.kaggle.com/code/waterjoe/convnext-crossentropyloss](url)\n\nI would be most grateful if you had time to take a look at my notebook"
            },
            {
              "id": 3219306,
              "postDate": "2025-06-07T13:07:22.423Z",
              "content": "<p><a href=\"https://www.kaggle.com/waterjoe\" target=\"_blank\">@waterjoe</a> you don't load trained state dict to model, predict are random, use <code>m.load_state_dict(state_dict)</code> before append to models</p>",
              "rawMarkdown": "@waterjoe you don't load trained state dict to model, predict are random, use `m.load_state_dict(state_dict)` before append to models",
              "votes": 1
            },
            {
              "id": 3219320,
              "postDate": "2025-06-07T13:54:31.873Z",
              "content": "<blockquote>\n  <p>you don't load trained state dict to model, predict are random</p>\n</blockquote>\n<p>That's what I said above 😀</p>",
              "rawMarkdown": ">  you don't load trained state dict to model, predict are random\n\nThat's what I said above 😀",
              "votes": 1
            },
            {
              "id": 3219323,
              "postDate": "2025-06-07T13:56:26.637Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 3219326,
              "postDate": "2025-06-07T13:57:07.120Z",
              "content": "<p>Thank you. I think my eyes are blurry. Hahaha</p>",
              "rawMarkdown": "Thank you. I think my eyes are blurry. Hahaha"
            },
            {
              "id": 3219329,
              "postDate": "2025-06-07T13:58:04.793Z",
              "content": "<p>Thank you. I think my eyes are blurry. Hahaha, I haven't reacted yet 🥹</p>",
              "rawMarkdown": "Thank you. I think my eyes are blurry. Hahaha, I haven't reacted yet 🥹"
            }
          ]
        }
      ]
    },
    {
      "id": 3218660,
      "postDate": "2025-06-06T14:17:44.663Z",
      "content": "<p>Epoch 4:     Train MAE: 22.33     Val MAE: 30.67     Time: 00:00:21     Step: 1/451<br>\nEpoch 4:     Train MAE: 21.83     Val MAE: 30.67     Time: 00:08:20     Step: 101/451<br>\nEpoch 4:     Train MAE: 21.77     Val MAE: 30.67     Time: 00:16:18     Step: 201/451<br>\nEpoch 4:     Train MAE: 21.89     Val MAE: 30.67     Time: 00:24:24     Step: 301/451<br>\nEpoch 4:     Train MAE: 21.71     Val MAE: 30.67     Time: 00:32:23     Step: 401/451<br>\n100%|█████████████████████████████████████████| 313/313 [01:08&lt;00:00,  4.58it/s]<br>\nEnding training (early_stopping).</p>\n<h1>my last round of training. The val here is very small, but why does the following test become larger? Please ask everyone🤔</h1>\n<p>CurveFault_A 720.66<br>\nCurveFault_B 751.34Why does fine-tuning make mae larger<br>\nCurveVel_A 650.84<br>\nCurveVel_B 749.36<br>\nFlatFault_A 722.53<br>\nFlatFault_B 732.67<br>\nFlatVel_A 658.38<br>\nFlatVel_B 751.96<br>\nStyle_A 532.39<br>\nStyle_B 488.19<br>\n= = = = = = = = = = = = = = = = = = = = = = = = =<br>\nVal MAE: 675.83<br>\n= = = = = = = = = = = = = = = = = = = = = = = = =The backbone was frozen, and then only a joint training was added</p>",
      "rawMarkdown": "Epoch 4:     Train MAE: 22.33     Val MAE: 30.67     Time: 00:00:21     Step: 1/451\nEpoch 4:     Train MAE: 21.83     Val MAE: 30.67     Time: 00:08:20     Step: 101/451\nEpoch 4:     Train MAE: 21.77     Val MAE: 30.67     Time: 00:16:18     Step: 201/451\nEpoch 4:     Train MAE: 21.89     Val MAE: 30.67     Time: 00:24:24     Step: 301/451\nEpoch 4:     Train MAE: 21.71     Val MAE: 30.67     Time: 00:32:23     Step: 401/451\n100%|█████████████████████████████████████████| 313/313 [01:08<00:00,  4.58it/s]\nEnding training (early_stopping).\n\nmy last round of training. The val here is very small, but why does the following test become larger? Please ask everyone🤔\n=========================\nCurveFault_A 720.66\nCurveFault_B 751.34Why does fine-tuning make mae larger\nCurveVel_A 650.84\nCurveVel_B 749.36\nFlatFault_A 722.53\nFlatFault_B 732.67\nFlatVel_A 658.38\nFlatVel_B 751.96\nStyle_A 532.39\nStyle_B 488.19\n= = = = = = = = = = = = = = = = = = = = = = = = =\nVal MAE: 675.83\n= = = = = = = = = = = = = = = = = = = = = = = = =The backbone was frozen, and then only a joint training was added",
      "votes": 2
    },
    {
      "id": 3219868,
      "postDate": "2025-06-08T12:31:08.757Z",
      "content": "<p>same issue</p>",
      "rawMarkdown": "same issue"
    },
    {
      "id": 3219217,
      "postDate": "2025-06-07T10:17:59.827Z",
      "content": "<p>Me too, don't know why</p>",
      "rawMarkdown": "Me too, don't know why"
    }
  ],
  "comments": [
    {
      "id": 3219225,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2025-06-07T10:34:45.013000",
      "content": "<p>It looks you use a model without fine tuning. Do you really use the output of your fine tuning/</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3219228,
          "author_name": "water joe",
          "author_url": "",
          "post_date": "2025-06-07T10:45:27.750000",
          "content": "<p>Thank you for your reply. This is my way of load the network，</p>\n<p>`import glob<br>\nimport torch<br>\nimport torch.nn as nn<br>\nimport torch.nn.functional as F<br>\nfrom _cfg import cfg<br>\nfrom _model import Net, EnsembleModel</p>\n<p>if RUN_VALID or RUN_TEST:<br>\n    # Load pretrained models<br>\n    models = []<br>\n    m = Net(<br>\n                backbone=\"convnext_small.fb_in22k_ft_in1k\",<br>\n                pretrained=False,<br>\n            )<br>\n    state_dict= torch.load('/kaggle/working/best_model_42.pt', map_location=cfg.device, weights_only=True)<br>\n    models.append(m)<br>\n    # Combine<br>\n    model = EnsembleModel(models)<br>\n    model = model.to(cfg.device)<br>\n    model = model.eval()<br>\n    print(\"n_models: {:_}\".format(len(models)))`</p>\n<p>The following is the same as the previous notebook ConvNeXt.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3219238,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2025-06-07T11:00:51.067000",
              "content": "<p>why do you use ensemble model given you have only one model?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3219243,
              "author_name": "water joe",
              "author_url": "",
              "post_date": "2025-06-07T11:09:00.263000",
              "content": "<p>Yes, of course your answer is correct. I think but this code should be able to handle the output of a single model. It's just a little efficiency issue。</p>\n<p>`class EnsembleModel(nn.Module):<br>\n    def <strong>init</strong>(self, models):<br>\n        super().<strong>init</strong>()<br>\n        self.models = nn.ModuleList(models).eval()</p>\n<pre><code> ():\n    output = \n     m  .models:\n        out = m(x)  \n         (out, ):  \n            logits = out[]\n        : \n            logits = out\n\n         output  :\n            output = logits\n        :\n            output += logits\n    output /= (.models)\n     output,   `\n</code></pre>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3219271,
              "author_name": "water joe",
              "author_url": "",
              "post_date": "2025-06-07T11:57:14.513000",
              "content": "<p><a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/waterjoe/convnext-crossentropyloss</a></p>\n<p>I would be most grateful if you had time to take a look at my notebook</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3219306,
              "author_name": "Nguyen",
              "author_url": "",
              "post_date": "2025-06-07T13:07:22.423000",
              "content": "<p><a href=\"https://www.kaggle.com/waterjoe\" target=\"_blank\">@waterjoe</a> you don't load trained state dict to model, predict are random, use <code>m.load_state_dict(state_dict)</code> before append to models</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3219320,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2025-06-07T13:54:31.873000",
              "content": "<blockquote>\n  <p>you don't load trained state dict to model, predict are random</p>\n</blockquote>\n<p>That's what I said above 😀</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3219323,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-06-07T13:56:26.637000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3219326,
              "author_name": "water joe",
              "author_url": "",
              "post_date": "2025-06-07T13:57:07.120000",
              "content": "<p>Thank you. I think my eyes are blurry. Hahaha</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3219329,
              "author_name": "water joe",
              "author_url": "",
              "post_date": "2025-06-07T13:58:04.793000",
              "content": "<p>Thank you. I think my eyes are blurry. Hahaha, I haven't reacted yet 🥹</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3219868,
      "author_name": "Quang Hưng",
      "author_url": "",
      "post_date": "2025-06-08T12:31:08.757000",
      "content": "<p>same issue</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3219217,
      "author_name": "Ren_si_yi0905",
      "author_url": "",
      "post_date": "2025-06-07T10:17:59.827000",
      "content": "<p>Me too, don't know why</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3219225": "It looks you use a model without fine tuning. Do you really use the output of your fine tuning/",
    "3218660": "Epoch 4:     Train MAE: 22.33     Val MAE: 30.67     Time: 00:00:21     Step: 1/451\nEpoch 4:     Train MAE: 21.83     Val MAE: 30.67     Time: 00:08:20     Step: 101/451\nEpoch 4:     Train MAE: 21.77     Val MAE: 30.67     Time: 00:16:18     Step: 201/451\nEpoch 4:     Train MAE: 21.89     Val MAE: 30.67     Time: 00:24:24     Step: 301/451\nEpoch 4:     Train MAE: 21.71     Val MAE: 30.67     Time: 00:32:23     Step: 401/451\n100%|█████████████████████████████████████████| 313/313 [01:08<00:00,  4.58it/s]\nEnding training (early_stopping).\n\nmy last round of training. The val here is very small, but why does the following test become larger? Please ask everyone🤔\n=========================\nCurveFault_A 720.66\nCurveFault_B 751.34Why does fine-tuning make mae larger\nCurveVel_A 650.84\nCurveVel_B 749.36\nFlatFault_A 722.53\nFlatFault_B 732.67\nFlatVel_A 658.38\nFlatVel_B 751.96\nStyle_A 532.39\nStyle_B 488.19\n= = = = = = = = = = = = = = = = = = = = = = = = =\nVal MAE: 675.83\n= = = = = = = = = = = = = = = = = = = = = = = = =The backbone was frozen, and then only a joint training was added",
    "3219868": "same issue",
    "3219217": "Me too, don't know why"
  }
}