{
  "id": 587511,
  "title": "15th place: Two-stage model",
  "url": "/competitions/waveform-inversion/writeups/jun-koda-15th-place-two-stage-model",
  "author_name": "",
  "post_date": "2025-07-01T10:24:16.880Z",
  "votes": 30,
  "comment_count": 12,
  "views": 0,
  "content": "<p>I thank the organizer and Kaggle for hosting this interesting physics competition.</p>\n<p>My solution is a 2-stage model:</p>\n<ol>\n<li>I use Bartley's caformer model, as the 1st stage, to predict the target sound speed field y.</li>\n<li>Then I simulate signal x1 = Foward[y1] with the forward model using the 1st-stage prediction y1.</li>\n<li>The input for the 2nd stage is 10-channel, concatenating the original signal x, and<br>\nthe deviation of the simulated signal from the original: δx = x - x1.</li>\n<li>The other input is the predicted sound speed field y1.</li>\n<li>The two inputs are added up after stem convolutions and enter the caformer model in the Bartley model.</li>\n<li>The target is the deviation from the input: δy = y - y1.</li>\n</ol>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F954117%2Fa7025b78275f2249ecb5832b59de7125%2Fgeophysical2.001.jpeg?generation=1751365434963517&amp;alt=media\" alt=\"\"></p>\n<p>My imagination behind is that the sound speed field, y, is crucial for mapping time to depth, and the model must want to know it from the beginning.</p>\n<p>The residual signal, δx, tells the discrepancy between the true sound speed and the estimated sound speed, and it should be informative how the model improve the previously estimated field y1.</p>\n<p>I also had the diffusion model in mind, and I wanted to estimate the sound speed field iteratively: y_n -&gt; y_n+1, but I could not make it work:</p>\n<ul>\n<li>I could not generate perturbed data y_n for training that generalize to the 1st-stage output</li>\n<li>So I use the out-of-fold predictions of the 5-fold 1st-stage model as the input for the 2nd stage</li>\n<li>I also tried the 3rd stage, but that did not improve the prediction. Maybe y1 is sufficiant and slightly better y2 does not help, but I am not convinced.</li>\n</ul>\n<ol>\n<li>The 1st-stage model yields a validation score 27 (or ~30 for proper weight imitating the test set) and public score 27.9 for 5-fold median. </li>\n<li>The second-stage model reduces the error to Local CV 13 for a single prediction,  11 when I use 5 outputs from the 1st stage, and Public score of 13. I was convinced that this is a gold model solution, 6 am (Japan time) this morning, 9 am is the deadline, but to my surprise, there were 1-submission Grandmasters, not only one but three!!!</li>\n</ol>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F954117%2F0160d8af31e941b1e50eac18fd7b4470%2Fscore.png?generation=1751364097322562&amp;alt=media\" alt=\"\"></p>\n<p>1st stage is trained for 120 epochs and 2nd stage is 125 epochs. The 2nd-stage output δy=0 corresponds to the 1st-stage solution, so the 2nd-stage model starts from the 1st-stage score by construction, although the model starts from new weights.</p>\n<p>True flip augmentation:<br>\nI did not like the flip augmentation because the source at the center (channel 2) is at 34 and flip maps to 69 - 34 = 35, one pixel shifted. I simulated the signal for the flipped field with the same source position at 34 and randomly selected one of them during training.</p>",
  "messages": [
    {
      "id": "3237802",
      "postDate": "07/01/2025 10:11:36",
      "content": "<p>I thank the organizer and Kaggle for hosting this interesting physics competition.</p>\n<p>My solution is a 2-stage model:</p>\n<ol>\n<li>I use Bartley's caformer model, as the 1st stage, to predict the target sound speed field y.</li>\n<li>Then I simulate signal x1 = Foward[y1] with the forward model using the 1st-stage prediction y1.</li>\n<li>The input for the 2nd stage is 10-channel, concatenating the original signal x, and<br>\nthe deviation of the simulated signal from the original: δx = x - x1.</li>\n<li>The other input is the predicted sound speed field y1.</li>\n<li>The two inputs are added up after stem convolutions and enter the caformer model in the Bartley model.</li>\n<li>The target is the deviation from the input: δy = y - y1.</li>\n</ol>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F954117%2Fa7025b78275f2249ecb5832b59de7125%2Fgeophysical2.001.jpeg?generation=1751365434963517&amp;alt=media\" alt=\"\"></p>\n<p>My imagination behind is that the sound speed field, y, is crucial for mapping time to depth, and the model must want to know it from the beginning.</p>\n<p>The residual signal, δx, tells the discrepancy between the true sound speed and the estimated sound speed, and it should be informative how the model improve the previously estimated field y1.</p>\n<p>I also had the diffusion model in mind, and I wanted to estimate the sound speed field iteratively: y_n -&gt; y_n+1, but I could not make it work:</p>\n<ul>\n<li>I could not generate perturbed data y_n for training that generalize to the 1st-stage output</li>\n<li>So I use the out-of-fold predictions of the 5-fold 1st-stage model as the input for the 2nd stage</li>\n<li>I also tried the 3rd stage, but that did not improve the prediction. Maybe y1 is sufficiant and slightly better y2 does not help, but I am not convinced.</li>\n</ul>\n<ol>\n<li>The 1st-stage model yields a validation score 27 (or ~30 for proper weight imitating the test set) and public score 27.9 for 5-fold median. </li>\n<li>The second-stage model reduces the error to Local CV 13 for a single prediction,  11 when I use 5 outputs from the 1st stage, and Public score of 13. I was convinced that this is a gold model solution, 6 am (Japan time) this morning, 9 am is the deadline, but to my surprise, there were 1-submission Grandmasters, not only one but three!!!</li>\n</ol>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F954117%2F0160d8af31e941b1e50eac18fd7b4470%2Fscore.png?generation=1751364097322562&amp;alt=media\" alt=\"\"></p>\n<p>1st stage is trained for 120 epochs and 2nd stage is 125 epochs. The 2nd-stage output δy=0 corresponds to the 1st-stage solution, so the 2nd-stage model starts from the 1st-stage score by construction, although the model starts from new weights.</p>\n<p>True flip augmentation:<br>\nI did not like the flip augmentation because the source at the center (channel 2) is at 34 and flip maps to 69 - 34 = 35, one pixel shifted. I simulated the signal for the flipped field with the same source position at 34 and randomly selected one of them during training.</p>",
      "rawMarkdown": "I thank the organizer and Kaggle for hosting this interesting physics competition.\n\nMy solution is a 2-stage model:\n\n1. I use Bartley's caformer model, as the 1st stage, to predict the target sound speed field y.\n2. Then I simulate signal x1 = Foward[y1] with the forward model using the 1st-stage prediction y1.\n3. The input for the 2nd stage is 10-channel, concatenating the original signal x, and\nthe deviation of the simulated signal from the original: δx = x - x1.\n4. The other input is the predicted sound speed field y1.\n5. The two inputs are added up after stem convolutions and enter the caformer model in the Bartley model.\n6. The target is the deviation from the input: δy = y - y1.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F954117%2Fa7025b78275f2249ecb5832b59de7125%2Fgeophysical2.001.jpeg?generation=1751365434963517&alt=media)\n\nMy imagination behind is that the sound speed field, y, is crucial for mapping time to depth, and the model must want to know it from the beginning.\n\nThe residual signal, δx, tells the discrepancy between the true sound speed and the estimated sound speed, and it should be informative how the model improve the previously estimated field y1.\n\nI also had the diffusion model in mind, and I wanted to estimate the sound speed field iteratively: y_n -> y_n+1, but I could not make it work:\n\n- I could not generate perturbed data y_n for training that generalize to the 1st-stage output\n- So I use the out-of-fold predictions of the 5-fold 1st-stage model as the input for the 2nd stage\n- I also tried the 3rd stage, but that did not improve the prediction. Maybe y1 is sufficiant and slightly better y2 does not help, but I am not convinced.\n\n1. The 1st-stage model yields a validation score 27 (or ~30 for proper weight imitating the test set) and public score 27.9 for 5-fold median. \n2. The second-stage model reduces the error to Local CV 13 for a single prediction,  11 when I use 5 outputs from the 1st stage, and Public score of 13. I was convinced that this is a gold model solution, 6 am (Japan time) this morning, 9 am is the deadline, but to my surprise, there were 1-submission Grandmasters, not only one but three!!!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F954117%2F0160d8af31e941b1e50eac18fd7b4470%2Fscore.png?generation=1751364097322562&alt=media)\n\n1st stage is trained for 120 epochs and 2nd stage is 125 epochs. The 2nd-stage output δy=0 corresponds to the 1st-stage solution, so the 2nd-stage model starts from the 1st-stage score by construction, although the model starts from new weights.\n\nTrue flip augmentation:\nI did not like the flip augmentation because the source at the center (channel 2) is at 34 and flip maps to 69 - 34 = 35, one pixel shifted. I simulated the signal for the flipped field with the same source position at 34 and randomly selected one of them during training.",
      "votes": null
    },
    {
      "id": "3237814",
      "postDate": "07/01/2025 10:25:01",
      "content": "<p>Your stage 2 is something I wanted to implement myself, great to see that it can yield such improvements! Thank you for publishing your solution!</p>",
      "rawMarkdown": "Your stage 2 is something I wanted to implement myself, great to see that it can yield such improvements! Thank you for publishing your solution!",
      "votes": null
    },
    {
      "id": "3237824",
      "postDate": "07/01/2025 10:29:50",
      "content": "<p>Thanks! Congratulations for the victory and the amaizin score!!</p>",
      "rawMarkdown": "Thanks! Congratulations for the victory and the amaizin score!!",
      "votes": null
    },
    {
      "id": "3237834",
      "postDate": "07/01/2025 10:34:32",
      "content": "<p>So you trained only on the original data, without augmentin and/or use test data?<br>\nReally amazing!</p>",
      "rawMarkdown": "So you trained only on the original data, without augmentin and/or use test data?\nReally amazing!",
      "votes": null
    },
    {
      "id": "3237840",
      "postDate": "07/01/2025 10:37:24",
      "content": "<p>This approach is like training 2x depth bartley's model with residual connection under the first 1x depth weight is frozen (with richer features).<br>\nIt seems smart approach witch fits less computing resource.<br>\nThanks you for sharing!</p>",
      "rawMarkdown": "This approach is like training 2x depth bartley's model with residual connection under the first 1x depth weight is frozen (with richer features).\nIt seems smart approach witch fits less computing resource.\nThanks you for sharing!",
      "votes": null
    },
    {
      "id": "3237866",
      "postDate": "07/01/2025 10:57:21",
      "content": "<p>Just x2 for flip. Your improving speed in score was amazing too! In hindsight, having training val gap, more data was obvious, but I was rushing in the final week…</p>",
      "rawMarkdown": "Just x2 for flip. Your improving speed in score was amazing too! In hindsight, having training val gap, more data was obvious, but I was rushing in the final week...",
      "votes": null
    },
    {
      "id": "3237875",
      "postDate": "07/01/2025 11:00:40",
      "content": "<p>Yes, since your method is so strong only with the original data, I really do wonder how effective it could be if trained with more data+test predictions.<br>\nIt maybe can even crash 1st place! </p>",
      "rawMarkdown": "Yes, since your method is so strong only with the original data, I really do wonder how effective it could be if trained with more data+test predictions.\nIt maybe can even crash 1st place!",
      "votes": null
    },
    {
      "id": "3237881",
      "postDate": "07/01/2025 11:03:14",
      "content": "<p>I see, maybe I should have combined the two and trained in one step. I started from the idea of the diffusion model, so thinking using one single model for multiple steps, no idea combining finite steps.</p>\n<p>[ADD] Oh, the problem is that I didn't have fast forward simulation yn -&gt; xn for computation during training.</p>",
      "rawMarkdown": "I see, maybe I should have combined the two and trained in one step. I started from the idea of the diffusion model, so thinking using one single model for multiple steps, no idea combining finite steps.\n\n[ADD] Oh, the problem is that I didn't have fast forward simulation yn -> xn for computation during training.",
      "votes": null
    },
    {
      "id": "3237916",
      "postDate": "07/01/2025 11:35:56",
      "content": "<blockquote>\n  <p>I see, maybe I should have combined the two and trained in one step.</p>\n</blockquote>\n<p>It might some improvement if computation and memory consumption allows. But I love your solution since it leverages as much as possible using public resources (which I failed during competition.)</p>\n<blockquote>\n  <p>[ADD] Oh, the problem is that I didn't have fast forward simulation yn -&gt; xn for computation during training.</p>\n</blockquote>\n<p>Forward model is very expensive to compute. I think some effort to reduce computation will be required to train end to end.</p>\n<p>[Edit]</p>\n<p>If residual seismic feature is crucial for this approach, end-to-end approach is costly comparing to gain.</p>\n<p>On the other hand, if the first model's velocity feature is rather crucial, maybe end to end approach might have some gain (with dropping seismic residual feature.)</p>",
      "rawMarkdown": "> I see, maybe I should have combined the two and trained in one step.\n\nIt might some improvement if computation and memory consumption allows. But I love your solution since it leverages as much as possible using public resources (which I failed during competition.)\n\n> [ADD] Oh, the problem is that I didn't have fast forward simulation yn -> xn for computation during training.\n\nForward model is very expensive to compute. I think some effort to reduce computation will be required to train end to end.\n\n[Edit]\n\nIf residual seismic feature is crucial for this approach, end-to-end approach is costly comparing to gain.\n\nOn the other hand, if the first model's velocity feature is rather crucial, maybe end to end approach might have some gain (with dropping seismic residual feature.)",
      "votes": null
    },
    {
      "id": "3237983",
      "postDate": "07/01/2025 12:38:21",
      "content": "<p>No synthetic data! Really amazing!</p>",
      "rawMarkdown": "No synthetic data! Really amazing!",
      "votes": null
    },
    {
      "id": "3238579",
      "postDate": "07/02/2025 02:09:56",
      "content": "<p>You have done a great job, by sharing this it will help many people and they will also get to learn from it.</p>",
      "rawMarkdown": "You have done a great job, by sharing this it will help many people and they will also get to learn from it.",
      "votes": null
    },
    {
      "id": "3238863",
      "postDate": "07/02/2025 09:05:27",
      "content": "<p>Thanks for sharing this! And congratulations with a great work! </p>\n<p>I attempted something similar. <br>\nI didn’t make it work though. I was getting improvements around 0.05 MAE on the val per epoch for the first couple of epochs, so I stoped to push this direction.</p>\n<p>The other differences were:<br>\n1) In addition to stem, I’ve added two independent first stages of the decoder for the velocity model.<br>\n2) I played with both DeepWave and more accurate PyTorch model, saw no difference in the training, only the batch time changed.<br>\n3) I also attempted to train more than two stages at the same time (specifically I was training two extra stages at the same time). The extra stage gave some additional minor improvement on the val.<br>\n4) I’ve used the simulated x as input instead of delta x.<br>\n5) I was greedy for the compute, so I used the pre-trained model from Bartley for both the first stage and to initialize my second stage network.<br>\n6) And maybe I got some extra bugs.</p>",
      "rawMarkdown": "Thanks for sharing this! And congratulations with a great work! \n\nI attempted something similar. \nI didn’t make it work though. I was getting improvements around 0.05 MAE on the val per epoch for the first couple of epochs, so I stoped to push this direction.\n\nThe other differences were:\n1) In addition to stem, I’ve added two independent first stages of the decoder for the velocity model.\n2) I played with both DeepWave and more accurate PyTorch model, saw no difference in the training, only the batch time changed.\n3) I also attempted to train more than two stages at the same time (specifically I was training two extra stages at the same time). The extra stage gave some additional minor improvement on the val.\n4) I’ve used the simulated x as input instead of delta x.\n5) I was greedy for the compute, so I used the pre-trained model from Bartley for both the first stage and to initialize my second stage network.\n6) And maybe I got some extra bugs.",
      "votes": null
    },
    {
      "id": "3238918",
      "postDate": "07/02/2025 10:19:39",
      "content": "<p>Thanks! Congratulations for the 1-submission gold! 🤯🤯🤯</p>",
      "rawMarkdown": "Thanks! Congratulations for the 1-submission gold! 🤯🤯🤯",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3237814,
      "author_name": "harshitsheoran",
      "author_url": "",
      "post_date": "07/01/2025 10:25:01",
      "content": "<p>Your stage 2 is something I wanted to implement myself, great to see that it can yield such improvements! Thank you for publishing your solution!</p>",
      "votes": null,
      "replies": [
        {
          "id": 3237824,
          "author_name": "junkoda",
          "author_url": "",
          "post_date": "07/01/2025 10:29:50",
          "content": "<p>Thanks! Congratulations for the victory and the amaizin score!!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3237834,
      "author_name": "shlomoron",
      "author_url": "",
      "post_date": "07/01/2025 10:34:32",
      "content": "<p>So you trained only on the original data, without augmentin and/or use test data?<br>\nReally amazing!</p>",
      "votes": null,
      "replies": [
        {
          "id": 3237866,
          "author_name": "junkoda",
          "author_url": "",
          "post_date": "07/01/2025 10:57:21",
          "content": "<p>Just x2 for flip. Your improving speed in score was amazing too! In hindsight, having training val gap, more data was obvious, but I was rushing in the final week…</p>",
          "votes": null,
          "replies": [
            {
              "id": 3237875,
              "author_name": "shlomoron",
              "author_url": "",
              "post_date": "07/01/2025 11:00:40",
              "content": "<p>Yes, since your method is so strong only with the original data, I really do wonder how effective it could be if trained with more data+test predictions.<br>\nIt maybe can even crash 1st place! </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3237840,
      "author_name": "tatamikenn",
      "author_url": "",
      "post_date": "07/01/2025 10:37:24",
      "content": "<p>This approach is like training 2x depth bartley's model with residual connection under the first 1x depth weight is frozen (with richer features).<br>\nIt seems smart approach witch fits less computing resource.<br>\nThanks you for sharing!</p>",
      "votes": null,
      "replies": [
        {
          "id": 3237881,
          "author_name": "junkoda",
          "author_url": "",
          "post_date": "07/01/2025 11:03:14",
          "content": "<p>I see, maybe I should have combined the two and trained in one step. I started from the idea of the diffusion model, so thinking using one single model for multiple steps, no idea combining finite steps.</p>\n<p>[ADD] Oh, the problem is that I didn't have fast forward simulation yn -&gt; xn for computation during training.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3237916,
              "author_name": "tatamikenn",
              "author_url": "",
              "post_date": "07/01/2025 11:35:56",
              "content": "<blockquote>\n  <p>I see, maybe I should have combined the two and trained in one step.</p>\n</blockquote>\n<p>It might some improvement if computation and memory consumption allows. But I love your solution since it leverages as much as possible using public resources (which I failed during competition.)</p>\n<blockquote>\n  <p>[ADD] Oh, the problem is that I didn't have fast forward simulation yn -&gt; xn for computation during training.</p>\n</blockquote>\n<p>Forward model is very expensive to compute. I think some effort to reduce computation will be required to train end to end.</p>\n<p>[Edit]</p>\n<p>If residual seismic feature is crucial for this approach, end-to-end approach is costly comparing to gain.</p>\n<p>On the other hand, if the first model's velocity feature is rather crucial, maybe end to end approach might have some gain (with dropping seismic residual feature.)</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3237983,
      "author_name": "hydantess",
      "author_url": "",
      "post_date": "07/01/2025 12:38:21",
      "content": "<p>No synthetic data! Really amazing!</p>",
      "votes": null,
      "replies": [
        {
          "id": 3238918,
          "author_name": "junkoda",
          "author_url": "",
          "post_date": "07/02/2025 10:19:39",
          "content": "<p>Thanks! Congratulations for the 1-submission gold! 🤯🤯🤯</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3238579,
      "author_name": "prashantkumaryt",
      "author_url": "",
      "post_date": "07/02/2025 02:09:56",
      "content": "<p>You have done a great job, by sharing this it will help many people and they will also get to learn from it.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3238863,
      "author_name": "dolphininacoma",
      "author_url": "",
      "post_date": "07/02/2025 09:05:27",
      "content": "<p>Thanks for sharing this! And congratulations with a great work! </p>\n<p>I attempted something similar. <br>\nI didn’t make it work though. I was getting improvements around 0.05 MAE on the val per epoch for the first couple of epochs, so I stoped to push this direction.</p>\n<p>The other differences were:<br>\n1) In addition to stem, I’ve added two independent first stages of the decoder for the velocity model.<br>\n2) I played with both DeepWave and more accurate PyTorch model, saw no difference in the training, only the batch time changed.<br>\n3) I also attempted to train more than two stages at the same time (specifically I was training two extra stages at the same time). The extra stage gave some additional minor improvement on the val.<br>\n4) I’ve used the simulated x as input instead of delta x.<br>\n5) I was greedy for the compute, so I used the pre-trained model from Bartley for both the first stage and to initialize my second stage network.<br>\n6) And maybe I got some extra bugs.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3237802": "I thank the organizer and Kaggle for hosting this interesting physics competition.\n\nMy solution is a 2-stage model:\n\n1. I use Bartley's caformer model, as the 1st stage, to predict the target sound speed field y.\n2. Then I simulate signal x1 = Foward[y1] with the forward model using the 1st-stage prediction y1.\n3. The input for the 2nd stage is 10-channel, concatenating the original signal x, and\nthe deviation of the simulated signal from the original: δx = x - x1.\n4. The other input is the predicted sound speed field y1.\n5. The two inputs are added up after stem convolutions and enter the caformer model in the Bartley model.\n6. The target is the deviation from the input: δy = y - y1.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F954117%2Fa7025b78275f2249ecb5832b59de7125%2Fgeophysical2.001.jpeg?generation=1751365434963517&alt=media)\n\nMy imagination behind is that the sound speed field, y, is crucial for mapping time to depth, and the model must want to know it from the beginning.\n\nThe residual signal, δx, tells the discrepancy between the true sound speed and the estimated sound speed, and it should be informative how the model improve the previously estimated field y1.\n\nI also had the diffusion model in mind, and I wanted to estimate the sound speed field iteratively: y_n -> y_n+1, but I could not make it work:\n\n- I could not generate perturbed data y_n for training that generalize to the 1st-stage output\n- So I use the out-of-fold predictions of the 5-fold 1st-stage model as the input for the 2nd stage\n- I also tried the 3rd stage, but that did not improve the prediction. Maybe y1 is sufficiant and slightly better y2 does not help, but I am not convinced.\n\n1. The 1st-stage model yields a validation score 27 (or ~30 for proper weight imitating the test set) and public score 27.9 for 5-fold median. \n2. The second-stage model reduces the error to Local CV 13 for a single prediction,  11 when I use 5 outputs from the 1st stage, and Public score of 13. I was convinced that this is a gold model solution, 6 am (Japan time) this morning, 9 am is the deadline, but to my surprise, there were 1-submission Grandmasters, not only one but three!!!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F954117%2F0160d8af31e941b1e50eac18fd7b4470%2Fscore.png?generation=1751364097322562&alt=media)\n\n1st stage is trained for 120 epochs and 2nd stage is 125 epochs. The 2nd-stage output δy=0 corresponds to the 1st-stage solution, so the 2nd-stage model starts from the 1st-stage score by construction, although the model starts from new weights.\n\nTrue flip augmentation:\nI did not like the flip augmentation because the source at the center (channel 2) is at 34 and flip maps to 69 - 34 = 35, one pixel shifted. I simulated the signal for the flipped field with the same source position at 34 and randomly selected one of them during training.",
    "3237814": "Your stage 2 is something I wanted to implement myself, great to see that it can yield such improvements! Thank you for publishing your solution!",
    "3237824": "Thanks! Congratulations for the victory and the amaizin score!!",
    "3237834": "So you trained only on the original data, without augmentin and/or use test data?\nReally amazing!",
    "3237840": "This approach is like training 2x depth bartley's model with residual connection under the first 1x depth weight is frozen (with richer features).\nIt seems smart approach witch fits less computing resource.\nThanks you for sharing!",
    "3237866": "Just x2 for flip. Your improving speed in score was amazing too! In hindsight, having training val gap, more data was obvious, but I was rushing in the final week...",
    "3237875": "Yes, since your method is so strong only with the original data, I really do wonder how effective it could be if trained with more data+test predictions.\nIt maybe can even crash 1st place!",
    "3237881": "I see, maybe I should have combined the two and trained in one step. I started from the idea of the diffusion model, so thinking using one single model for multiple steps, no idea combining finite steps.\n\n[ADD] Oh, the problem is that I didn't have fast forward simulation yn -> xn for computation during training.",
    "3237916": "> I see, maybe I should have combined the two and trained in one step.\n\nIt might some improvement if computation and memory consumption allows. But I love your solution since it leverages as much as possible using public resources (which I failed during competition.)\n\n> [ADD] Oh, the problem is that I didn't have fast forward simulation yn -> xn for computation during training.\n\nForward model is very expensive to compute. I think some effort to reduce computation will be required to train end to end.\n\n[Edit]\n\nIf residual seismic feature is crucial for this approach, end-to-end approach is costly comparing to gain.\n\nOn the other hand, if the first model's velocity feature is rather crucial, maybe end to end approach might have some gain (with dropping seismic residual feature.)",
    "3237983": "No synthetic data! Really amazing!",
    "3238579": "You have done a great job, by sharing this it will help many people and they will also get to learn from it.",
    "3238863": "Thanks for sharing this! And congratulations with a great work! \n\nI attempted something similar. \nI didn’t make it work though. I was getting improvements around 0.05 MAE on the val per epoch for the first couple of epochs, so I stoped to push this direction.\n\nThe other differences were:\n1) In addition to stem, I’ve added two independent first stages of the decoder for the velocity model.\n2) I played with both DeepWave and more accurate PyTorch model, saw no difference in the training, only the batch time changed.\n3) I also attempted to train more than two stages at the same time (specifically I was training two extra stages at the same time). The extra stage gave some additional minor improvement on the val.\n4) I’ve used the simulated x as input instead of delta x.\n5) I was greedy for the compute, so I used the pre-trained model from Bartley for both the first stage and to initialize my second stage network.\n6) And maybe I got some extra bugs.",
    "3238918": "Thanks! Congratulations for the 1-submission gold! 🤯🤯🤯"
  },
  "source": "meta"
}