{
  "id": 584137,
  "title": "CAFormer 8Core TPU Training(Pytorch/XLA)",
  "url": "/competitions/waveform-inversion/discussion/584137",
  "author_name": "Nakanishi Wataru",
  "post_date": "2025-06-12T00:37:52.709000",
  "votes": 7,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I've adapted a <a href=\"https://www.kaggle.com/code/haruiig/train-unet-with-tpu-pytorch-xla\" target=\"_blank\">notebook</a> for training U-Net with TPUs (PyTorch/XLA) to work with Bartley's CAFormer model. Although it takes over 8 hours to train one epoch with 460,000 samples, it might be useful for those with remaining TPU runtime. </p>\n<p>Notebook is <a href=\"https://www.kaggle.com/code/nakanishiwataru/train-caformer-with-tpu-pytorch-xla\" target=\"_blank\">here</a></p>\n<p>Any advice or suggestions would be appreciated.</p>",
  "messages": [
    {
      "id": 3222174,
      "postDate": "2025-06-12T00:37:52.710Z",
      "content": "<p>I've adapted a <a href=\"https://www.kaggle.com/code/haruiig/train-unet-with-tpu-pytorch-xla\" target=\"_blank\">notebook</a> for training U-Net with TPUs (PyTorch/XLA) to work with Bartley's CAFormer model. Although it takes over 8 hours to train one epoch with 460,000 samples, it might be useful for those with remaining TPU runtime. </p>\n<p>Notebook is <a href=\"https://www.kaggle.com/code/nakanishiwataru/train-caformer-with-tpu-pytorch-xla\" target=\"_blank\">here</a></p>\n<p>Any advice or suggestions would be appreciated.</p>",
      "rawMarkdown": "I've adapted a [notebook](https://www.kaggle.com/code/haruiig/train-unet-with-tpu-pytorch-xla) for training U-Net with TPUs (PyTorch/XLA) to work with Bartley's CAFormer model. Although it takes over 8 hours to train one epoch with 460,000 samples, it might be useful for those with remaining TPU runtime. \n\nNotebook is [here](https://www.kaggle.com/code/nakanishiwataru/train-caformer-with-tpu-pytorch-xla)\n\nAny advice or suggestions would be appreciated.",
      "votes": 7
    },
    {
      "id": 3222433,
      "postDate": "2025-06-12T07:40:22.823Z",
      "content": "<p>The open-source model achieved a validation MAE of 24.15, but after training, the model's validation MAE increased to 27.95 — it seems the performance actually degraded…</p>",
      "rawMarkdown": "The open-source model achieved a validation MAE of 24.15, but after training, the model's validation MAE increased to 27.95 — it seems the performance actually degraded...",
      "votes": 2,
      "replies": [
        {
          "id": 3222476,
          "postDate": "2025-06-12T08:08:55.743Z",
          "content": "<p>Thanks for your comment. I haven't checked my validation MAE yet, but if performance degrades, then training with a TPU might not be the best approach.</p>\n<p>In my notebook, I changed the learning rate (lr) from 1e-4 to 1e-6 and cos_pct from 0.5 to 0.7. Could these hyperparameter changes be affecting the MAE? Or is this simply a limitation of using a TPU?</p>",
          "rawMarkdown": "Thanks for your comment. I haven't checked my validation MAE yet, but if performance degrades, then training with a TPU might not be the best approach.\n\nIn my notebook, I changed the learning rate (lr) from 1e-4 to 1e-6 and cos_pct from 0.5 to 0.7. Could these hyperparameter changes be affecting the MAE? Or is this simply a limitation of using a TPU?",
          "replies": [
            {
              "id": 3222522,
              "postDate": "2025-06-12T08:36:11.520Z",
              "content": "<p>Thank you for sharing your notebook! As for your question, I'm not sure how to answer it at the moment. I'd like to run some experiments first and get back to you afterward.</p>",
              "rawMarkdown": "Thank you for sharing your notebook! As for your question, I'm not sure how to answer it at the moment. I'd like to run some experiments first and get back to you afterward.",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 3222758,
      "postDate": "2025-06-12T13:26:12.507Z",
      "content": "<p>CurveFault_A    9.34<br>\nCurveFault_B    73.58<br>\nCurveVel_A      13.92<br>\nCurveVel_B      40.71<br>\nFlatFault_A     8.30<br>\nFlatFault_B     27.55<br>\nFlatVel_A       7.80<br>\nFlatVel_B       11.22<br>\nStyle_A         36.86<br>\nStyle_B         50.25</p>\n<p>Val MAE: 27.95</p>\n<p>The validation MAE worsened after training for one epoch (460k samples) with a very low learning rate of 1e-6. This suggests that TPUs might not be suitable for fine-tuning with low learning rates.</p>",
      "rawMarkdown": "CurveFault_A    9.34\nCurveFault_B    73.58\nCurveVel_A      13.92\nCurveVel_B      40.71\nFlatFault_A     8.30\nFlatFault_B     27.55\nFlatVel_A       7.80\nFlatVel_B       11.22\nStyle_A         36.86\nStyle_B         50.25\n\nVal MAE: 27.95\n\n\nThe validation MAE worsened after training for one epoch (460k samples) with a very low learning rate of 1e-6. This suggests that TPUs might not be suitable for fine-tuning with low learning rates.",
      "replies": [
        {
          "id": 3222779,
          "postDate": "2025-06-12T14:05:07.887Z",
          "content": "<p>With GPU it's good?<br>\nWhen you finetune with GPU you use float32?<br>\nTPU use bfloat16. So you need larger batches. At least 64, better 128 in my experience. If the batch was smaller it may have been the issue.</p>",
          "rawMarkdown": "With GPU it's good?\nWhen you finetune with GPU you use float32?\nTPU use bfloat16. So you need larger batches. At least 64, better 128 in my experience. If the batch was smaller it may have been the issue.",
          "votes": 3,
          "replies": [
            {
              "id": 3222894,
              "postDate": "2025-06-12T16:17:36.640Z",
              "content": "<p>Thank you for your advice. I admit I hadn't been paying much attention to the batch size. I will try increasing it and run the training again!</p>",
              "rawMarkdown": "Thank you for your advice. I admit I hadn't been paying much attention to the batch size. I will try increasing it and run the training again!"
            }
          ]
        }
      ]
    },
    {
      "id": 3222555,
      "postDate": "2025-06-12T08:58:39.513Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3222433,
      "author_name": "bestwater",
      "author_url": "",
      "post_date": "2025-06-12T07:40:22.823000",
      "content": "<p>The open-source model achieved a validation MAE of 24.15, but after training, the model's validation MAE increased to 27.95 — it seems the performance actually degraded…</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3222476,
          "author_name": "Nakanishi Wataru",
          "author_url": "",
          "post_date": "2025-06-12T08:08:55.743000",
          "content": "<p>Thanks for your comment. I haven't checked my validation MAE yet, but if performance degrades, then training with a TPU might not be the best approach.</p>\n<p>In my notebook, I changed the learning rate (lr) from 1e-4 to 1e-6 and cos_pct from 0.5 to 0.7. Could these hyperparameter changes be affecting the MAE? Or is this simply a limitation of using a TPU?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3222522,
              "author_name": "bestwater",
              "author_url": "",
              "post_date": "2025-06-12T08:36:11.520000",
              "content": "<p>Thank you for sharing your notebook! As for your question, I'm not sure how to answer it at the moment. I'd like to run some experiments first and get back to you afterward.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3222758,
      "author_name": "Nakanishi Wataru",
      "author_url": "",
      "post_date": "2025-06-12T13:26:12.507000",
      "content": "<p>CurveFault_A    9.34<br>\nCurveFault_B    73.58<br>\nCurveVel_A      13.92<br>\nCurveVel_B      40.71<br>\nFlatFault_A     8.30<br>\nFlatFault_B     27.55<br>\nFlatVel_A       7.80<br>\nFlatVel_B       11.22<br>\nStyle_A         36.86<br>\nStyle_B         50.25</p>\n<p>Val MAE: 27.95</p>\n<p>The validation MAE worsened after training for one epoch (460k samples) with a very low learning rate of 1e-6. This suggests that TPUs might not be suitable for fine-tuning with low learning rates.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3222779,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2025-06-12T14:05:07.887000",
          "content": "<p>With GPU it's good?<br>\nWhen you finetune with GPU you use float32?<br>\nTPU use bfloat16. So you need larger batches. At least 64, better 128 in my experience. If the batch was smaller it may have been the issue.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 3222894,
              "author_name": "Nakanishi Wataru",
              "author_url": "",
              "post_date": "2025-06-12T16:17:36.640000",
              "content": "<p>Thank you for your advice. I admit I hadn't been paying much attention to the batch size. I will try increasing it and run the training again!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3222555,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-06-12T08:58:39.513000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3222174": "I've adapted a [notebook](https://www.kaggle.com/code/haruiig/train-unet-with-tpu-pytorch-xla) for training U-Net with TPUs (PyTorch/XLA) to work with Bartley's CAFormer model. Although it takes over 8 hours to train one epoch with 460,000 samples, it might be useful for those with remaining TPU runtime. \n\nNotebook is [here](https://www.kaggle.com/code/nakanishiwataru/train-caformer-with-tpu-pytorch-xla)\n\nAny advice or suggestions would be appreciated.",
    "3222433": "The open-source model achieved a validation MAE of 24.15, but after training, the model's validation MAE increased to 27.95 — it seems the performance actually degraded...",
    "3222758": "CurveFault_A    9.34\nCurveFault_B    73.58\nCurveVel_A      13.92\nCurveVel_B      40.71\nFlatFault_A     8.30\nFlatFault_B     27.55\nFlatVel_A       7.80\nFlatVel_B       11.22\nStyle_A         36.86\nStyle_B         50.25\n\nVal MAE: 27.95\n\n\nThe validation MAE worsened after training for one epoch (460k samples) with a very low learning rate of 1e-6. This suggests that TPUs might not be suitable for fine-tuning with low learning rates.",
    "3222555": ""
  }
}