{
  "id": 578305,
  "title": "HGNet-V2 Encoder - [CV 55.6 LB 60.9]",
  "url": "/competitions/waveform-inversion/discussion/578305",
  "author_name": "",
  "post_date": "2025-05-09T23:31:42.582352700Z",
  "votes": 45,
  "comment_count": 22,
  "views": 0,
  "content": "<p>Excited to share a notebook that builds on the great starter notebook from <a href=\"https://www.kaggle.com/egortrushin\" target=\"_blank\">@egortrushin</a>. The updated notebook showcases training on 2x T4 GPUs to maximize our GPU quota!</p>\n<p>Notebook <a href=\"https://www.kaggle.com/code/brendanartley/hgnet-v2-starter-cv-65-2-lb-69-7\" target=\"_blank\">here</a><br>\nDataset <a href=\"https://www.kaggle.com/datasets/brendanartley/openfwi-preprocessed-72x72\" target=\"_blank\">here</a></p>\n<p>I also provide a few trained checkpoints for a HGNet-V2-Unet model, flip augmentations, EMA, and more. Here are the CV scores by dataset type. Experiment away! 😀</p>\n<pre><code>+--------------+--------+\n| Dataset      | Score  |\n+--------------+--------+\n| CurveFault_A |   |\n| CurveFault_B |  |\n| CurveVel_A   |   |\n| CurveVel_B   |   |\n| FlatFault_A  |   |\n| FlatFault_B  |   |\n| FlatVel_A    |   |\n| FlatVel_B    |   |\n| Style_A      |   |\n| Style_B      |   |\n+--------------+--------+\n</code></pre>",
  "messages": [
    {
      "id": "3198769",
      "postDate": "05/09/2025 23:31:42",
      "content": "<p>Excited to share a notebook that builds on the great starter notebook from <a href=\"https://www.kaggle.com/egortrushin\" target=\"_blank\">@egortrushin</a>. The updated notebook showcases training on 2x T4 GPUs to maximize our GPU quota!</p>\n<p>Notebook <a href=\"https://www.kaggle.com/code/brendanartley/hgnet-v2-starter-cv-65-2-lb-69-7\" target=\"_blank\">here</a><br>\nDataset <a href=\"https://www.kaggle.com/datasets/brendanartley/openfwi-preprocessed-72x72\" target=\"_blank\">here</a></p>\n<p>I also provide a few trained checkpoints for a HGNet-V2-Unet model, flip augmentations, EMA, and more. Here are the CV scores by dataset type. Experiment away! 😀</p>\n<pre><code>+--------------+--------+\n| Dataset      | Score  |\n+--------------+--------+\n| CurveFault_A |   |\n| CurveFault_B |  |\n| CurveVel_A   |   |\n| CurveVel_B   |   |\n| FlatFault_A  |   |\n| FlatFault_B  |   |\n| FlatVel_A    |   |\n| FlatVel_B    |   |\n| Style_A      |   |\n| Style_B      |   |\n+--------------+--------+\n</code></pre>",
      "rawMarkdown": "Excited to share a notebook that builds on the great starter notebook from @egortrushin. The updated notebook showcases training on 2x T4 GPUs to maximize our GPU quota!\n\nNotebook [here](https://www.kaggle.com/code/brendanartley/hgnet-v2-starter-cv-65-2-lb-69-7)\nDataset [here](https://www.kaggle.com/datasets/brendanartley/openfwi-preprocessed-72x72)\n\nI also provide a few trained checkpoints for a HGNet-V2-Unet model, flip augmentations, EMA, and more. Here are the CV scores by dataset type. Experiment away! 😀\n\n```python\n+--------------+--------+\n| Dataset      | Score  |\n+--------------+--------+\n| CurveFault_A |  15.47 |\n| CurveFault_B | 132.66 |\n| CurveVel_A   |  37.10 |\n| CurveVel_B   |  96.18 |\n| FlatFault_A  |  14.07 |\n| FlatFault_B  |  71.20 |\n| FlatVel_A    |  19.20 |\n| FlatVel_B    |  48.95 |\n| Style_A      |  51.15 |\n| Style_B      |  70.72 |\n+--------------+--------+\n\n```",
      "votes": null
    },
    {
      "id": "3199046",
      "postDate": "05/10/2025 11:29:17",
      "content": "<p><a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> <br>\nThank you for sharing your notebook.<br>\nIt seems that the CV scores differ significantly between the different datasets, do you have any new insights or advice on the differences?For example, should we focus on a particular dataset to address it, etc.<br>\nBest regards.</p>",
      "rawMarkdown": "brendanartley \nThank you for sharing your notebook.\nIt seems that the CV scores differ significantly between the different datasets, do you have any new insights or advice on the differences?For example, should we focus on a particular dataset to address it, etc.\nBest regards.",
      "votes": null
    },
    {
      "id": "3199141",
      "postDate": "05/10/2025 14:18:44",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/ov104ss\" target=\"_blank\">@ov104ss</a>,</p>\n<p>A 2-stage approach could work well. For example, you could train a model to predict which class each data point belongs to. Then, you could train a family specific model for the velocity-map reconstruction step. It may be that some model architectures work well for some families but not for others. Definitely worth trying!</p>\n<p>Maybe others can chime in if they have already compared 1-stage vs 2-stage approaches?</p>",
      "rawMarkdown": "Hi @ov104ss,\n\nA 2-stage approach could work well. For example, you could train a model to predict which class each data point belongs to. Then, you could train a family specific model for the velocity-map reconstruction step. It may be that some model architectures work well for some families but not for others. Definitely worth trying!\n\nMaybe others can chime in if they have already compared 1-stage vs 2-stage approaches?",
      "votes": null
    },
    {
      "id": "3199157",
      "postDate": "05/10/2025 14:35:45",
      "content": "<p><a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> <br>\nThank you for your reply. That idea is worth a try. I was thinking that it would be difficult to cover all families with a single model, so I will consider it. Thank you very much.</p>",
      "rawMarkdown": "brendanartley \nThank you for your reply. That idea is worth a try. I was thinking that it would be difficult to cover all families with a single model, so I will consider it. Thank you very much.",
      "votes": null
    },
    {
      "id": "3199190",
      "postDate": "05/10/2025 15:36:18",
      "content": "<p>DDP does appear that its the way of the future for running multiple GPU's, but I am finding it pretty difficult to debug errors when making changes to the code.  Any suggestions (besides lots of print code) to ease this issue using Visual Code on local machine?</p>\n<p>Another thing I knew but forgot and relearned the hard way, its going to write over any existing config, utils.py, etc when you run it local.  If your like me and playing with lots of the shared notebooks in the competition your going to be unhappy if you don't get into a habit of separate folders per shared notebook on your local machine.</p>",
      "rawMarkdown": "DDP does appear that its the way of the future for running multiple GPU's, but I am finding it pretty difficult to debug errors when making changes to the code.  Any suggestions (besides lots of print code) to ease this issue using Visual Code on local machine?\n\nAnother thing I knew but forgot and relearned the hard way, its going to write over any existing config, utils.py, etc when you run it local.  If your like me and playing with lots of the shared notebooks in the competition your going to be unhappy if you don't get into a habit of separate folders per shared notebook on your local machine.",
      "votes": null
    },
    {
      "id": "3199327",
      "postDate": "05/10/2025 19:54:24",
      "content": "<p>I wonder if a MoE type of approach would work for it ?, Cause the task is to recreate the velocity maps in all the types, but the scores vary, as far as I'm aware this indicates a change in distribution, a router based mixture of experts type of thing would work pretty well in my opinion.</p>",
      "rawMarkdown": "I wonder if a MoE type of approach would work for it ?, Cause the task is to recreate the velocity maps in all the types, but the scores vary, as far as I'm aware this indicates a change in distribution, a router based mixture of experts type of thing would work pretty well in my opinion.",
      "votes": null
    },
    {
      "id": "3199388",
      "postDate": "05/10/2025 21:57:38",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a>, thanks for the comment.</p>\n<p>One thing you could do is extract pieces of this notebook and insert them into an existing pipeline that works for you (models, augs, snippets, etc). This might help pinpoint issues in your environment. Hope this helps!</p>",
      "rawMarkdown": "Hi @pcjimmmy, thanks for the comment.\n\nOne thing you could do is extract pieces of this notebook and insert them into an existing pipeline that works for you (models, augs, snippets, etc). This might help pinpoint issues in your environment. Hope this helps!",
      "votes": null
    },
    {
      "id": "3201496",
      "postDate": "05/14/2025 00:38:44",
      "content": "<p><strong>Update 13/5/2025</strong></p>\n<p>CV: 65.2 -&gt; 55.6<br>\nLB: 69.7 -&gt; 60.9</p>\n<p>Thanks everyone for waiting, the training script now works on Kaggle. I also released 3x stronger pretrained checkpoints (trained for 150 epochs), and added more customizable decoder. Cheers!</p>",
      "rawMarkdown": "**Update 13/5/2025**\n\nCV: 65.2 -> 55.6\nLB: 69.7 -> 60.9\n\nThanks everyone for waiting, the training script now works on Kaggle. I also released 3x stronger pretrained checkpoints (trained for 150 epochs), and added more customizable decoder. Cheers!",
      "votes": null
    },
    {
      "id": "3201517",
      "postDate": "05/14/2025 01:49:22",
      "content": "<p>May I ask what's your training set MAE after 50/150 epochs? In my case train/CV/LB : 25/36/36 after 40 epochs, I am considering whether it has fully converged.</p>",
      "rawMarkdown": "May I ask what's your training set MAE after 50/150 epochs? In my case train/CV/LB : 25/36/36 after 40 epochs, I am considering whether it has fully converged.",
      "votes": null
    },
    {
      "id": "3201530",
      "postDate": "05/14/2025 02:29:01",
      "content": "<p>Sure, here are the loss curves for a 150 epoch training run.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F4ae987ace74101119d0c0ec8de06971f%2Fimage.jpg?generation=1747189628091945&amp;alt=media\" alt=\"LossImages\"></p>",
      "rawMarkdown": "Sure, here are the loss curves for a 150 epoch training run.\n\n![LossImages](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F4ae987ace74101119d0c0ec8de06971f%2Fimage.jpg?generation=1747189628091945&alt=media)",
      "votes": null
    },
    {
      "id": "3201534",
      "postDate": "05/14/2025 02:37:13",
      "content": "<p>Crazy epochs! Out of curiosity, how long did the 250 epochs take?</p>",
      "rawMarkdown": "Crazy epochs! Out of curiosity, how long did the 250 epochs take?",
      "votes": null
    },
    {
      "id": "3201538",
      "postDate": "05/14/2025 03:00:12",
      "content": "<p>Thanks for sharing! The improvement is mainly due to the changes in pretrained model, right? And may I ask about your training data for pretrained model, with all OpenFWI dataset?</p>",
      "rawMarkdown": "Thanks for sharing! The improvement is mainly due to the changes in pretrained model, right? And may I ask about your training data for pretrained model, with all OpenFWI dataset?",
      "votes": null
    },
    {
      "id": "3201548",
      "postDate": "05/14/2025 03:14:10",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/chenxin1991\" target=\"_blank\">@chenxin1991</a>, thanks for the comment.</p>\n<p>The pre-trained model is actually the same from V1 -&gt; V4, there is just more customizability now. The models are trained on the same train/valid set as the public notebook (~98% of all OpenFWI), and the CV/LB improvement came purely from increasing the training epochs!</p>",
      "rawMarkdown": "Hi @chenxin1991, thanks for the comment.\n\nThe pre-trained model is actually the same from V1 -> V4, there is just more customizability now. The models are trained on the same train/valid set as the public notebook (~98% of all OpenFWI), and the CV/LB improvement came purely from increasing the training epochs!",
      "votes": null
    },
    {
      "id": "3201550",
      "postDate": "05/14/2025 03:31:06",
      "content": "<p>Hello, what are the differences between these three models?</p>",
      "rawMarkdown": "Hello, what are the differences between these three models?",
      "votes": null
    },
    {
      "id": "3201573",
      "postDate": "05/14/2025 04:30:54",
      "content": "<p>It takes ~24hrs on 2x T4 GPUs for 150 epochs, though it is much faster on more modern GPUs.</p>",
      "rawMarkdown": "It takes ~24hrs on 2x T4 GPUs for 150 epochs, though it is much faster on more modern GPUs.",
      "votes": null
    },
    {
      "id": "3201600",
      "postDate": "05/14/2025 05:57:39",
      "content": "<p>Your dataset is excellent—it eliminates the bottleneck caused by disk loading.</p>",
      "rawMarkdown": "Your dataset is excellent—it eliminates the bottleneck caused by disk loading.",
      "votes": null
    },
    {
      "id": "3201686",
      "postDate": "05/14/2025 08:54:23",
      "content": "<p>May I ask what's the difference among checkpoint best0~2 during training? Since they can all be loaded by the same model structure, it suggests the architecture is identical. Is the difference due to 3-fold cross-validation?</p>",
      "rawMarkdown": "May I ask what's the difference among checkpoint best0~2 during training? Since they can all be loaded by the same model structure, it suggests the architecture is identical. Is the difference due to 3-fold cross-validation?",
      "votes": null
    },
    {
      "id": "3201822",
      "postDate": "05/14/2025 12:46:23",
      "content": "<p><a href=\"https://www.kaggle.com/cqrcqr\" target=\"_blank\">@cqrcqr</a>, <a href=\"https://www.kaggle.com/wayne127\" target=\"_blank\">@wayne127</a> thanks for your comments.</p>\n<p>The only difference between the three models is their seed. They were trained and evaluated on the same fold as in the public notebook.</p>",
      "rawMarkdown": "cqrcqr, @wayne127 thanks for your comments.\n\nThe only difference between the three models is their seed. They were trained and evaluated on the same fold as in the public notebook.",
      "votes": null
    },
    {
      "id": "3202350",
      "postDate": "05/15/2025 09:01:25",
      "content": "<p>Thank you for sharing. I noticed that you used data augmentation. May I ask why the width of the flip speed graph corresponds to the flipping of num_Sources and num_deceivers? I don't quite understand the relationship between them. I hope you can help me🥹</p>",
      "rawMarkdown": "Thank you for sharing. I noticed that you used data augmentation. May I ask why the width of the flip speed graph corresponds to the flipping of num_Sources and num_deceivers? I don't quite understand the relationship between them. I hope you can help me🥹",
      "votes": null
    },
    {
      "id": "3202468",
      "postDate": "05/15/2025 12:24:11",
      "content": "<p><a href=\"https://www.kaggle.com/code/brendanartley/hgnet-v2-starter/comments#3200039\" target=\"_blank\">https://www.kaggle.com/code/brendanartley/hgnet-v2-starter/comments#3200039</a> 🌚</p>",
      "rawMarkdown": "https://www.kaggle.com/code/brendanartley/hgnet-v2-starter/comments#3200039 🌚",
      "votes": null
    },
    {
      "id": "3202524",
      "postDate": "05/15/2025 14:14:52",
      "content": "<p>Bartley, thank you for sharing your work. I'm learning a lot from it. </p>\n<p>Your network seems to be composed of many different blocks. Is there a methodology to the way you pick the building blocks? Or is it experience and/or intuition?</p>",
      "rawMarkdown": "Bartley, thank you for sharing your work. I'm learning a lot from it. \n\nYour network seems to be composed of many different blocks. Is there a methodology to the way you pick the building blocks? Or is it experience and/or intuition?",
      "votes": null
    },
    {
      "id": "3202551",
      "postDate": "05/15/2025 14:44:12",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/mgoksu\" target=\"_blank\">@mgoksu</a>, most of the blocks are there for customizability. </p>\n<p>For me, choosing which blocks to use is a mix of intuition and cross validation (mostly the latter). It is common for my intuition to be wrong, but a robust cross validation scheme rarely lies.</p>",
      "rawMarkdown": "Hi @mgoksu, most of the blocks are there for customizability. \n\nFor me, choosing which blocks to use is a mix of intuition and cross validation (mostly the latter). It is common for my intuition to be wrong, but a robust cross validation scheme rarely lies.",
      "votes": null
    },
    {
      "id": "3202696",
      "postDate": "05/15/2025 19:12:56",
      "content": "<p>thanks 嘴爷😍</p>",
      "rawMarkdown": "thanks 嘴爷😍",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3199046,
      "author_name": "ov104ss",
      "author_url": "",
      "post_date": "05/10/2025 11:29:17",
      "content": "<p><a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> <br>\nThank you for sharing your notebook.<br>\nIt seems that the CV scores differ significantly between the different datasets, do you have any new insights or advice on the differences?For example, should we focus on a particular dataset to address it, etc.<br>\nBest regards.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3199141,
          "author_name": "brendanartley",
          "author_url": "",
          "post_date": "05/10/2025 14:18:44",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/ov104ss\" target=\"_blank\">@ov104ss</a>,</p>\n<p>A 2-stage approach could work well. For example, you could train a model to predict which class each data point belongs to. Then, you could train a family specific model for the velocity-map reconstruction step. It may be that some model architectures work well for some families but not for others. Definitely worth trying!</p>\n<p>Maybe others can chime in if they have already compared 1-stage vs 2-stage approaches?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3199157,
              "author_name": "ov104ss",
              "author_url": "",
              "post_date": "05/10/2025 14:35:45",
              "content": "<p><a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> <br>\nThank you for your reply. That idea is worth a try. I was thinking that it would be difficult to cover all families with a single model, so I will consider it. Thank you very much.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3199190,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "05/10/2025 15:36:18",
      "content": "<p>DDP does appear that its the way of the future for running multiple GPU's, but I am finding it pretty difficult to debug errors when making changes to the code.  Any suggestions (besides lots of print code) to ease this issue using Visual Code on local machine?</p>\n<p>Another thing I knew but forgot and relearned the hard way, its going to write over any existing config, utils.py, etc when you run it local.  If your like me and playing with lots of the shared notebooks in the competition your going to be unhappy if you don't get into a habit of separate folders per shared notebook on your local machine.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3199388,
          "author_name": "brendanartley",
          "author_url": "",
          "post_date": "05/10/2025 21:57:38",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a>, thanks for the comment.</p>\n<p>One thing you could do is extract pieces of this notebook and insert them into an existing pipeline that works for you (models, augs, snippets, etc). This might help pinpoint issues in your environment. Hope this helps!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3199327,
      "author_name": "pshikk",
      "author_url": "",
      "post_date": "05/10/2025 19:54:24",
      "content": "<p>I wonder if a MoE type of approach would work for it ?, Cause the task is to recreate the velocity maps in all the types, but the scores vary, as far as I'm aware this indicates a change in distribution, a router based mixture of experts type of thing would work pretty well in my opinion.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3201496,
      "author_name": "brendanartley",
      "author_url": "",
      "post_date": "05/14/2025 00:38:44",
      "content": "<p><strong>Update 13/5/2025</strong></p>\n<p>CV: 65.2 -&gt; 55.6<br>\nLB: 69.7 -&gt; 60.9</p>\n<p>Thanks everyone for waiting, the training script now works on Kaggle. I also released 3x stronger pretrained checkpoints (trained for 150 epochs), and added more customizable decoder. Cheers!</p>",
      "votes": null,
      "replies": [
        {
          "id": 3201517,
          "author_name": "w5833946",
          "author_url": "",
          "post_date": "05/14/2025 01:49:22",
          "content": "<p>May I ask what's your training set MAE after 50/150 epochs? In my case train/CV/LB : 25/36/36 after 40 epochs, I am considering whether it has fully converged.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3201530,
              "author_name": "brendanartley",
              "author_url": "",
              "post_date": "05/14/2025 02:29:01",
              "content": "<p>Sure, here are the loss curves for a 150 epoch training run.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F4ae987ace74101119d0c0ec8de06971f%2Fimage.jpg?generation=1747189628091945&amp;alt=media\" alt=\"LossImages\"></p>",
              "votes": null,
              "replies": [
                {
                  "id": 3201534,
                  "author_name": "zy1343930734",
                  "author_url": "",
                  "post_date": "05/14/2025 02:37:13",
                  "content": "<p>Crazy epochs! Out of curiosity, how long did the 250 epochs take?</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3201573,
                      "author_name": "brendanartley",
                      "author_url": "",
                      "post_date": "05/14/2025 04:30:54",
                      "content": "<p>It takes ~24hrs on 2x T4 GPUs for 150 epochs, though it is much faster on more modern GPUs.</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        },
        {
          "id": 3201538,
          "author_name": "chenxin1991",
          "author_url": "",
          "post_date": "05/14/2025 03:00:12",
          "content": "<p>Thanks for sharing! The improvement is mainly due to the changes in pretrained model, right? And may I ask about your training data for pretrained model, with all OpenFWI dataset?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3201548,
              "author_name": "brendanartley",
              "author_url": "",
              "post_date": "05/14/2025 03:14:10",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/chenxin1991\" target=\"_blank\">@chenxin1991</a>, thanks for the comment.</p>\n<p>The pre-trained model is actually the same from V1 -&gt; V4, there is just more customizability now. The models are trained on the same train/valid set as the public notebook (~98% of all OpenFWI), and the CV/LB improvement came purely from increasing the training epochs!</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 3201686,
          "author_name": "wayne127",
          "author_url": "",
          "post_date": "05/14/2025 08:54:23",
          "content": "<p>May I ask what's the difference among checkpoint best0~2 during training? Since they can all be loaded by the same model structure, it suggests the architecture is identical. Is the difference due to 3-fold cross-validation?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3201550,
      "author_name": "cqrcqr",
      "author_url": "",
      "post_date": "05/14/2025 03:31:06",
      "content": "<p>Hello, what are the differences between these three models?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3201822,
          "author_name": "brendanartley",
          "author_url": "",
          "post_date": "05/14/2025 12:46:23",
          "content": "<p><a href=\"https://www.kaggle.com/cqrcqr\" target=\"_blank\">@cqrcqr</a>, <a href=\"https://www.kaggle.com/wayne127\" target=\"_blank\">@wayne127</a> thanks for your comments.</p>\n<p>The only difference between the three models is their seed. They were trained and evaluated on the same fold as in the public notebook.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3201600,
      "author_name": "bibanh",
      "author_url": "",
      "post_date": "05/14/2025 05:57:39",
      "content": "<p>Your dataset is excellent—it eliminates the bottleneck caused by disk loading.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3202350,
      "author_name": "peilwang",
      "author_url": "",
      "post_date": "05/15/2025 09:01:25",
      "content": "<p>Thank you for sharing. I noticed that you used data augmentation. May I ask why the width of the flip speed graph corresponds to the flipping of num_Sources and num_deceivers? I don't quite understand the relationship between them. I hope you can help me🥹</p>",
      "votes": null,
      "replies": [
        {
          "id": 3202468,
          "author_name": "zui0711",
          "author_url": "",
          "post_date": "05/15/2025 12:24:11",
          "content": "<p><a href=\"https://www.kaggle.com/code/brendanartley/hgnet-v2-starter/comments#3200039\" target=\"_blank\">https://www.kaggle.com/code/brendanartley/hgnet-v2-starter/comments#3200039</a> 🌚</p>",
          "votes": null,
          "replies": [
            {
              "id": 3202696,
              "author_name": "peilwang",
              "author_url": "",
              "post_date": "05/15/2025 19:12:56",
              "content": "<p>thanks 嘴爷😍</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3202524,
      "author_name": "mgoksu",
      "author_url": "",
      "post_date": "05/15/2025 14:14:52",
      "content": "<p>Bartley, thank you for sharing your work. I'm learning a lot from it. </p>\n<p>Your network seems to be composed of many different blocks. Is there a methodology to the way you pick the building blocks? Or is it experience and/or intuition?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3202551,
          "author_name": "brendanartley",
          "author_url": "",
          "post_date": "05/15/2025 14:44:12",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/mgoksu\" target=\"_blank\">@mgoksu</a>, most of the blocks are there for customizability. </p>\n<p>For me, choosing which blocks to use is a mix of intuition and cross validation (mostly the latter). It is common for my intuition to be wrong, but a robust cross validation scheme rarely lies.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3198769": "Excited to share a notebook that builds on the great starter notebook from @egortrushin. The updated notebook showcases training on 2x T4 GPUs to maximize our GPU quota!\n\nNotebook [here](https://www.kaggle.com/code/brendanartley/hgnet-v2-starter-cv-65-2-lb-69-7)\nDataset [here](https://www.kaggle.com/datasets/brendanartley/openfwi-preprocessed-72x72)\n\nI also provide a few trained checkpoints for a HGNet-V2-Unet model, flip augmentations, EMA, and more. Here are the CV scores by dataset type. Experiment away! 😀\n\n```python\n+--------------+--------+\n| Dataset      | Score  |\n+--------------+--------+\n| CurveFault_A |  15.47 |\n| CurveFault_B | 132.66 |\n| CurveVel_A   |  37.10 |\n| CurveVel_B   |  96.18 |\n| FlatFault_A  |  14.07 |\n| FlatFault_B  |  71.20 |\n| FlatVel_A    |  19.20 |\n| FlatVel_B    |  48.95 |\n| Style_A      |  51.15 |\n| Style_B      |  70.72 |\n+--------------+--------+\n\n```",
    "3199046": "brendanartley \nThank you for sharing your notebook.\nIt seems that the CV scores differ significantly between the different datasets, do you have any new insights or advice on the differences?For example, should we focus on a particular dataset to address it, etc.\nBest regards.",
    "3199141": "Hi @ov104ss,\n\nA 2-stage approach could work well. For example, you could train a model to predict which class each data point belongs to. Then, you could train a family specific model for the velocity-map reconstruction step. It may be that some model architectures work well for some families but not for others. Definitely worth trying!\n\nMaybe others can chime in if they have already compared 1-stage vs 2-stage approaches?",
    "3199157": "brendanartley \nThank you for your reply. That idea is worth a try. I was thinking that it would be difficult to cover all families with a single model, so I will consider it. Thank you very much.",
    "3199190": "DDP does appear that its the way of the future for running multiple GPU's, but I am finding it pretty difficult to debug errors when making changes to the code.  Any suggestions (besides lots of print code) to ease this issue using Visual Code on local machine?\n\nAnother thing I knew but forgot and relearned the hard way, its going to write over any existing config, utils.py, etc when you run it local.  If your like me and playing with lots of the shared notebooks in the competition your going to be unhappy if you don't get into a habit of separate folders per shared notebook on your local machine.",
    "3199327": "I wonder if a MoE type of approach would work for it ?, Cause the task is to recreate the velocity maps in all the types, but the scores vary, as far as I'm aware this indicates a change in distribution, a router based mixture of experts type of thing would work pretty well in my opinion.",
    "3199388": "Hi @pcjimmmy, thanks for the comment.\n\nOne thing you could do is extract pieces of this notebook and insert them into an existing pipeline that works for you (models, augs, snippets, etc). This might help pinpoint issues in your environment. Hope this helps!",
    "3201496": "**Update 13/5/2025**\n\nCV: 65.2 -> 55.6\nLB: 69.7 -> 60.9\n\nThanks everyone for waiting, the training script now works on Kaggle. I also released 3x stronger pretrained checkpoints (trained for 150 epochs), and added more customizable decoder. Cheers!",
    "3201517": "May I ask what's your training set MAE after 50/150 epochs? In my case train/CV/LB : 25/36/36 after 40 epochs, I am considering whether it has fully converged.",
    "3201530": "Sure, here are the loss curves for a 150 epoch training run.\n\n![LossImages](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F4ae987ace74101119d0c0ec8de06971f%2Fimage.jpg?generation=1747189628091945&alt=media)",
    "3201534": "Crazy epochs! Out of curiosity, how long did the 250 epochs take?",
    "3201538": "Thanks for sharing! The improvement is mainly due to the changes in pretrained model, right? And may I ask about your training data for pretrained model, with all OpenFWI dataset?",
    "3201548": "Hi @chenxin1991, thanks for the comment.\n\nThe pre-trained model is actually the same from V1 -> V4, there is just more customizability now. The models are trained on the same train/valid set as the public notebook (~98% of all OpenFWI), and the CV/LB improvement came purely from increasing the training epochs!",
    "3201550": "Hello, what are the differences between these three models?",
    "3201573": "It takes ~24hrs on 2x T4 GPUs for 150 epochs, though it is much faster on more modern GPUs.",
    "3201600": "Your dataset is excellent—it eliminates the bottleneck caused by disk loading.",
    "3201686": "May I ask what's the difference among checkpoint best0~2 during training? Since they can all be loaded by the same model structure, it suggests the architecture is identical. Is the difference due to 3-fold cross-validation?",
    "3201822": "cqrcqr, @wayne127 thanks for your comments.\n\nThe only difference between the three models is their seed. They were trained and evaluated on the same fold as in the public notebook.",
    "3202350": "Thank you for sharing. I noticed that you used data augmentation. May I ask why the width of the flip speed graph corresponds to the flipping of num_Sources and num_deceivers? I don't quite understand the relationship between them. I hope you can help me🥹",
    "3202468": "https://www.kaggle.com/code/brendanartley/hgnet-v2-starter/comments#3200039 🌚",
    "3202524": "Bartley, thank you for sharing your work. I'm learning a lot from it. \n\nYour network seems to be composed of many different blocks. Is there a methodology to the way you pick the building blocks? Or is it experience and/or intuition?",
    "3202551": "Hi @mgoksu, most of the blocks are there for customizability. \n\nFor me, choosing which blocks to use is a mix of intuition and cross validation (mostly the latter). It is common for my intuition to be wrong, but a robust cross validation scheme rarely lies.",
    "3202696": "thanks 嘴爷😍"
  },
  "source": "meta"
}