{
  "id": 585539,
  "title": "[Let's Collaborate] Generating `CurveVel` & `CurveFault` for Augmentation",
  "url": "/competitions/waveform-inversion/discussion/585539",
  "author_name": "",
  "post_date": "2025-06-21T05:21:32.908605800Z",
  "votes": 8,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>Like many of you, I've noticed that certain datasets like <code>CurveFault_B</code>, <code>CurveVel_B</code>, and <code>Style_B</code> are particularly challenging and seem to be a bottleneck for improving our scores.</p>\n<p>I believe that creating a robust data augmentation pipeline for these difficult cases could be a key to success. To that end, I've been working on reproducing the velocity map generation process described in Appendix B of the OpenFWI paper.</p>\n<h3>What I've Done So Far</h3>\n<p>I've managed to create a notebook that partially replicates the generation process for the <code>CurveVel</code> and <code>CurveFault</code> families based on the formulas in the paper.</p>\n<p><strong>You can find all the code and a detailed explanation in my notebook here:</strong></p>\n<p>Data Augmentation Notebook: <a href=\"https://www.kaggle.com/code/haruiig/fwi-velocity-map-augmentation-for-curvefault-vel#%E3%81%AF%E3%81%98%E3%82%81%E3%81%AB\" target=\"_blank\">here</a></p>\n<p>Here’s a quick look at the results. The generation for <code>CurveVel</code> looks quite promising:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22004535%2Fccfa101f5b2263fb758229fdd363e2fa%2FScreenshot%20from%202025-06-21%2014-17-05.png?generation=1750483117816171&amp;alt=media\" alt=\"Transformed Vel Map to generate CurveVel\"></p>\n<p>Original Vel Map:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22004535%2F6a75b063509748c442d655544337f73d%2FScreenshot%20from%202025-06-21%2014-19-27.png?generation=1750483192841609&amp;alt=media\" alt=\"\"></p>\n<p>The <code>CurveFault</code> generation is a bit more complex due to the larger number of parameters, and it's still a work-in-progress, but it's a start:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22004535%2Fdc74e47208477c8db0b44bde8aacd9f9%2FScreenshot%20from%202025-06-21%2014-20-22.png?generation=1750483236072461&amp;alt=media\" alt=\"\"></p>\n<h3>The Challenge &amp; How You Can Help</h3>\n<p>While this is a good first step, there are still many unknowns to create a truly effective augmentation pipeline. This is where I'd love to get the community's help!</p>\n<p>The main open questions are:</p>\n<ul>\n<li><strong>Parameter Distributions:</strong> What are the appropriate ranges and distributions for sampling the parameters (<code>a, k, s, s'</code>, etc.)?</li>\n<li><strong>Number of Iterations:</strong> How many times are these transformations recursively applied to create the final complex maps?</li>\n<li><strong>Initial Flat Maps:</strong> How are the initial flat-layered maps generated in the first place?</li>\n</ul>\n<h3>An Idea for Reverse-Engineering</h3>\n<p>One idea I had for <code>CurveVel</code> is to <strong>reverse-engineer the parameters from the existing training data</strong>. For example, we could potentially apply a <strong>Fourier Transform</strong> to the provided <code>CurveVel</code> maps to estimate the distribution of the key wave parameters (<code>a</code> and <code>k</code>).</p>\n<p><strong>If anyone has experience with this or would be interested in trying to implement it, it would be a huge contribution!</strong> Or, if you have a better idea, I'm all ears!</p>\n<p>Let's work together to decode this generation process and build a powerful augmentation strategy that benefits everyone. I look forward to your thoughts and ideas</p>",
  "messages": [
    {
      "id": "3229182",
      "postDate": "06/21/2025 05:21:32",
      "content": "<p>Hi everyone,</p>\n<p>Like many of you, I've noticed that certain datasets like <code>CurveFault_B</code>, <code>CurveVel_B</code>, and <code>Style_B</code> are particularly challenging and seem to be a bottleneck for improving our scores.</p>\n<p>I believe that creating a robust data augmentation pipeline for these difficult cases could be a key to success. To that end, I've been working on reproducing the velocity map generation process described in Appendix B of the OpenFWI paper.</p>\n<h3>What I've Done So Far</h3>\n<p>I've managed to create a notebook that partially replicates the generation process for the <code>CurveVel</code> and <code>CurveFault</code> families based on the formulas in the paper.</p>\n<p><strong>You can find all the code and a detailed explanation in my notebook here:</strong></p>\n<p>Data Augmentation Notebook: <a href=\"https://www.kaggle.com/code/haruiig/fwi-velocity-map-augmentation-for-curvefault-vel#%E3%81%AF%E3%81%98%E3%82%81%E3%81%AB\" target=\"_blank\">here</a></p>\n<p>Here’s a quick look at the results. The generation for <code>CurveVel</code> looks quite promising:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22004535%2Fccfa101f5b2263fb758229fdd363e2fa%2FScreenshot%20from%202025-06-21%2014-17-05.png?generation=1750483117816171&amp;alt=media\" alt=\"Transformed Vel Map to generate CurveVel\"></p>\n<p>Original Vel Map:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22004535%2F6a75b063509748c442d655544337f73d%2FScreenshot%20from%202025-06-21%2014-19-27.png?generation=1750483192841609&amp;alt=media\" alt=\"\"></p>\n<p>The <code>CurveFault</code> generation is a bit more complex due to the larger number of parameters, and it's still a work-in-progress, but it's a start:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22004535%2Fdc74e47208477c8db0b44bde8aacd9f9%2FScreenshot%20from%202025-06-21%2014-20-22.png?generation=1750483236072461&amp;alt=media\" alt=\"\"></p>\n<h3>The Challenge &amp; How You Can Help</h3>\n<p>While this is a good first step, there are still many unknowns to create a truly effective augmentation pipeline. This is where I'd love to get the community's help!</p>\n<p>The main open questions are:</p>\n<ul>\n<li><strong>Parameter Distributions:</strong> What are the appropriate ranges and distributions for sampling the parameters (<code>a, k, s, s'</code>, etc.)?</li>\n<li><strong>Number of Iterations:</strong> How many times are these transformations recursively applied to create the final complex maps?</li>\n<li><strong>Initial Flat Maps:</strong> How are the initial flat-layered maps generated in the first place?</li>\n</ul>\n<h3>An Idea for Reverse-Engineering</h3>\n<p>One idea I had for <code>CurveVel</code> is to <strong>reverse-engineer the parameters from the existing training data</strong>. For example, we could potentially apply a <strong>Fourier Transform</strong> to the provided <code>CurveVel</code> maps to estimate the distribution of the key wave parameters (<code>a</code> and <code>k</code>).</p>\n<p><strong>If anyone has experience with this or would be interested in trying to implement it, it would be a huge contribution!</strong> Or, if you have a better idea, I'm all ears!</p>\n<p>Let's work together to decode this generation process and build a powerful augmentation strategy that benefits everyone. I look forward to your thoughts and ideas</p>",
      "rawMarkdown": "Hi everyone,\n\nLike many of you, I've noticed that certain datasets like `CurveFault_B`, `CurveVel_B`, and `Style_B` are particularly challenging and seem to be a bottleneck for improving our scores.\n\nI believe that creating a robust data augmentation pipeline for these difficult cases could be a key to success. To that end, I've been working on reproducing the velocity map generation process described in Appendix B of the OpenFWI paper.\n\n### What I've Done So Far\n\nI've managed to create a notebook that partially replicates the generation process for the `CurveVel` and `CurveFault` families based on the formulas in the paper.\n\n**You can find all the code and a detailed explanation in my notebook here:**\n\nData Augmentation Notebook: [here](https://www.kaggle.com/code/haruiig/fwi-velocity-map-augmentation-for-curvefault-vel#%E3%81%AF%E3%81%98%E3%82%81%E3%81%AB)\n\nHere’s a quick look at the results. The generation for `CurveVel` looks quite promising:\n\n![Transformed Vel Map to generate CurveVel](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22004535%2Fccfa101f5b2263fb758229fdd363e2fa%2FScreenshot%20from%202025-06-21%2014-17-05.png?generation=1750483117816171&alt=media)\n\nOriginal Vel Map:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22004535%2F6a75b063509748c442d655544337f73d%2FScreenshot%20from%202025-06-21%2014-19-27.png?generation=1750483192841609&alt=media)\n\nThe `CurveFault` generation is a bit more complex due to the larger number of parameters, and it's still a work-in-progress, but it's a start:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22004535%2Fdc74e47208477c8db0b44bde8aacd9f9%2FScreenshot%20from%202025-06-21%2014-20-22.png?generation=1750483236072461&alt=media)\n\n\n### The Challenge & How You Can Help\n\nWhile this is a good first step, there are still many unknowns to create a truly effective augmentation pipeline. This is where I'd love to get the community's help!\n\nThe main open questions are:\n\n*   **Parameter Distributions:** What are the appropriate ranges and distributions for sampling the parameters (`a, k, s, s'`, etc.)?\n*   **Number of Iterations:** How many times are these transformations recursively applied to create the final complex maps?\n*   **Initial Flat Maps:** How are the initial flat-layered maps generated in the first place?\n\n### An Idea for Reverse-Engineering\n\nOne idea I had for `CurveVel` is to **reverse-engineer the parameters from the existing training data**. For example, we could potentially apply a **Fourier Transform** to the provided `CurveVel` maps to estimate the distribution of the key wave parameters (`a` and `k`).\n\n**If anyone has experience with this or would be interested in trying to implement it, it would be a huge contribution!** Or, if you have a better idea, I'm all ears!\n\nLet's work together to decode this generation process and build a powerful augmentation strategy that benefits everyone. I look forward to your thoughts and ideas",
      "votes": null
    },
    {
      "id": "3229201",
      "postDate": "06/21/2025 06:25:37",
      "content": "<p>thanks for the code. i am interested to know which is the winner here:<br>\n1) 20x more data  + unet + mae loss<br>\n2) 1x data  + unet + mae loss + physics loss (FWM)</p>\n<p>it may turn out the at the equivalent point, training time is the same</p>",
      "rawMarkdown": "thanks for the code. i am interested to know which is the winner here:\n1) 20x more data  + unet + mae loss\n2) 1x data  + unet + mae loss + physics loss (FWM)\n\nit may turn out the at the equivalent point, training time is the same",
      "votes": null
    },
    {
      "id": "3229315",
      "postDate": "06/21/2025 09:49:19",
      "content": "<p>i believe 20x more data beats phyiscs loss + 1x data! (if we can reproduce the vel-map generation process correctly)<br>\nusing more data, ml model can learn constraints of the model map(like continuity of velocity, the number of layers, how curvy velocity map be) precisely.<br>\nless data and phyiscs loss cannot incorporate the constraints well, i believe.<br>\nOf course, it's biased😂</p>",
      "rawMarkdown": "i believe 20x more data beats phyiscs loss + 1x data! (if we can reproduce the vel-map generation process correctly)\nusing more data, ml model can learn constraints of the model map(like continuity of velocity, the number of layers, how curvy velocity map be) precisely.\nless data and phyiscs loss cannot incorporate the constraints well, i believe.\nOf course, it's biased😂",
      "votes": null
    },
    {
      "id": "3229331",
      "postDate": "06/21/2025 10:18:24",
      "content": "<p>why don't do this:</p>\n<pre><code> given test semsic, initialise , k, s, ... = \n\n velocity x70 = Draw(, k, s, ....)\n\n semsic = FWM(velocity)\n\n mae loss = |semsic- test semsic|\n\n\n\n  gradient descend  update , k, s, ...\n</code></pre>\n<p>instead of trial and error, you can uncover the parameters by:</p>\n<pre><code>)    MLP -&gt; param =a, k, s, ...\n)  D velocity1\n\nloss =|| velocity1 - train velocty||\ndist loss = KL diverge (latent, standard guassian)\n\n back propagate ...  you get distrubution of predicted param ...\n\n# loss = GAN loss ...  you need to train discrminator\n</code></pre>\n<p>`</p>",
      "rawMarkdown": "why don't do this:\n\n```\n0. given test semsic, initialise a, k, s, ... = random\n\n1. velocity 70x70 = Draw(a, k, s, ....)\n\n2. semsic = FWM(velocity)\n\n3. mae loss = |semsic- test semsic|\n\n#or GAN loss if you want to just generate samples\n\n4. do gradient descend and update a, k, s, ...\n\n```\n\ninstead of trial and error, you can uncover the parameters by:\n\n```\n1) train velocty -> encoder -> latent -> MLP -> param =a, k, s, ...\n2) param -> Draw -> velocity1\n\nloss =|| velocity1 - train velocty||\ndist loss = KL diverge (latent, standard guassian)\n\n back propagate ... then you get distrubution of predicted param ...\n\n#or loss = GAN loss ... then you need to train discrminator\n````",
      "votes": null
    },
    {
      "id": "3229600",
      "postDate": "06/21/2025 18:50:33",
      "content": "<p>hi, what is FWM loss?</p>",
      "rawMarkdown": "hi, what is FWM loss?",
      "votes": null
    },
    {
      "id": "3229621",
      "postDate": "06/21/2025 19:42:27",
      "content": "<p>Either the physical loss has been non - converging, or the top team has been continuously generating synthetic data. Otherwise, it is difficult to explain why the score can keep improving.</p>",
      "rawMarkdown": "Either the physical loss has been non - converging, or the top team has been continuously generating synthetic data. Otherwise, it is difficult to explain why the score can keep improving.",
      "votes": null
    },
    {
      "id": "3229686",
      "postDate": "06/22/2025 01:03:00",
      "content": "<p>20x more data??? Hahaha I might have to tap out of this one. </p>",
      "rawMarkdown": "20x more data??? Hahaha I might have to tap out of this one.",
      "votes": null
    },
    {
      "id": "3229699",
      "postDate": "06/22/2025 01:38:14",
      "content": "<p>The physic loss is non convex. You good solver.</p>",
      "rawMarkdown": "The physic loss is non convex. You good solver.",
      "votes": null
    },
    {
      "id": "3229916",
      "postDate": "06/22/2025 09:46:20",
      "content": "<p>The impact of synthetic data doesn't seem to be that significant (?), in the end, it still comes down to \"patiently solving the problem.\"</p>",
      "rawMarkdown": "The impact of synthetic data doesn't seem to be that significant (?), in the end, it still comes down to \"patiently solving the problem.\"",
      "votes": null
    },
    {
      "id": "3229917",
      "postDate": "06/22/2025 09:50:22",
      "content": "<p>Depends. If u can generate test alike data, then you win</p>",
      "rawMarkdown": "Depends. If u can generate test alike data, then you win",
      "votes": null
    },
    {
      "id": "3229925",
      "postDate": "06/22/2025 10:12:22",
      "content": "<blockquote>\n  <p>Depends. If u can generate test alike data, then you win  </p>\n</blockquote>\n<p>Ain't there a public notebook that generate pretty accurate test alike data? Does not seems too difficult. </p>",
      "rawMarkdown": ">Depends. If u can generate test alike data, then you win  \n\nAin't there a public notebook that generate pretty accurate test alike data? Does not seems too difficult.",
      "votes": null
    },
    {
      "id": "3229930",
      "postDate": "06/22/2025 10:17:32",
      "content": "<p>All data science competitions are about</p>\n<ol>\n<li>A model that is large enough to eat and crunch available data</li>\n<li>Training data that is close to the hidden test</li>\n<li>Modeling less related to data/ML but more to methods … aka domain knowledge, physics eqn/ simulation, retrieval and matching, new modality/view, pre/post processing and novel<br>\nsolution steps/flow/pipeline</li>\n</ol>\n<p>If you are not getting test lb score to your training loss, the data are not similar enough</p>\n<p>—-</p>\n<p>Current tip solution is below 10. So I assume that his validation loss is similar and the train error is much lower. If validation error has reached a point that is low enough, his model basically has duplicated the whole physics process. </p>\n<p>He either has very good model or architecture that is same os physics equation or enough train data such that now test data are cinterpolation or extrapolation of the train. </p>",
      "rawMarkdown": "All data science competitions are about\n1. A model that is large enough to eat and crunch available data\n2. Training data that is close to the hidden test\n3. Modeling less related to data/ML but more to methods … aka domain knowledge, physics eqn/ simulation, retrieval and matching, new modality/view, pre/post processing and novel\n solution steps/flow/pipeline\n\nIf you are not getting test lb score to your training loss, the data are not similar enough\n\n—-\n\nCurrent tip solution is below 10. So I assume that his validation loss is similar and the train error is much lower. If validation error has reached a point that is low enough, his model basically has duplicated the whole physics process. \n\nHe either has very good model or architecture that is same os physics equation or enough train data such that now test data are cinterpolation or extrapolation of the train.",
      "votes": null
    },
    {
      "id": "3232012",
      "postDate": "06/25/2025 09:20:47",
      "content": "<p>how to draw velocity from param?</p>",
      "rawMarkdown": "how to draw velocity from param?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3229201,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "06/21/2025 06:25:37",
      "content": "<p>thanks for the code. i am interested to know which is the winner here:<br>\n1) 20x more data  + unet + mae loss<br>\n2) 1x data  + unet + mae loss + physics loss (FWM)</p>\n<p>it may turn out the at the equivalent point, training time is the same</p>",
      "votes": null,
      "replies": [
        {
          "id": 3229315,
          "author_name": "haruiig",
          "author_url": "",
          "post_date": "06/21/2025 09:49:19",
          "content": "<p>i believe 20x more data beats phyiscs loss + 1x data! (if we can reproduce the vel-map generation process correctly)<br>\nusing more data, ml model can learn constraints of the model map(like continuity of velocity, the number of layers, how curvy velocity map be) precisely.<br>\nless data and phyiscs loss cannot incorporate the constraints well, i believe.<br>\nOf course, it's biased😂</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3229600,
          "author_name": "jamalsaeedi",
          "author_url": "",
          "post_date": "06/21/2025 18:50:33",
          "content": "<p>hi, what is FWM loss?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3229621,
          "author_name": "guoooooooss",
          "author_url": "",
          "post_date": "06/21/2025 19:42:27",
          "content": "<p>Either the physical loss has been non - converging, or the top team has been continuously generating synthetic data. Otherwise, it is difficult to explain why the score can keep improving.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3229699,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "06/22/2025 01:38:14",
              "content": "<p>The physic loss is non convex. You good solver.</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 3229686,
          "author_name": "cody11null",
          "author_url": "",
          "post_date": "06/22/2025 01:03:00",
          "content": "<p>20x more data??? Hahaha I might have to tap out of this one. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3229916,
          "author_name": "atom1231",
          "author_url": "",
          "post_date": "06/22/2025 09:46:20",
          "content": "<p>The impact of synthetic data doesn't seem to be that significant (?), in the end, it still comes down to \"patiently solving the problem.\"</p>",
          "votes": null,
          "replies": [
            {
              "id": 3229917,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "06/22/2025 09:50:22",
              "content": "<p>Depends. If u can generate test alike data, then you win</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3229925,
                  "author_name": "shlomoron",
                  "author_url": "",
                  "post_date": "06/22/2025 10:12:22",
                  "content": "<blockquote>\n  <p>Depends. If u can generate test alike data, then you win  </p>\n</blockquote>\n<p>Ain't there a public notebook that generate pretty accurate test alike data? Does not seems too difficult. </p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3229930,
                      "author_name": "hengck23",
                      "author_url": "",
                      "post_date": "06/22/2025 10:17:32",
                      "content": "<p>All data science competitions are about</p>\n<ol>\n<li>A model that is large enough to eat and crunch available data</li>\n<li>Training data that is close to the hidden test</li>\n<li>Modeling less related to data/ML but more to methods … aka domain knowledge, physics eqn/ simulation, retrieval and matching, new modality/view, pre/post processing and novel<br>\nsolution steps/flow/pipeline</li>\n</ol>\n<p>If you are not getting test lb score to your training loss, the data are not similar enough</p>\n<p>—-</p>\n<p>Current tip solution is below 10. So I assume that his validation loss is similar and the train error is much lower. If validation error has reached a point that is low enough, his model basically has duplicated the whole physics process. </p>\n<p>He either has very good model or architecture that is same os physics equation or enough train data such that now test data are cinterpolation or extrapolation of the train. </p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3229331,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "06/21/2025 10:18:24",
      "content": "<p>why don't do this:</p>\n<pre><code> given test semsic, initialise , k, s, ... = \n\n velocity x70 = Draw(, k, s, ....)\n\n semsic = FWM(velocity)\n\n mae loss = |semsic- test semsic|\n\n\n\n  gradient descend  update , k, s, ...\n</code></pre>\n<p>instead of trial and error, you can uncover the parameters by:</p>\n<pre><code>)    MLP -&gt; param =a, k, s, ...\n)  D velocity1\n\nloss =|| velocity1 - train velocty||\ndist loss = KL diverge (latent, standard guassian)\n\n back propagate ...  you get distrubution of predicted param ...\n\n# loss = GAN loss ...  you need to train discrminator\n</code></pre>\n<p>`</p>",
      "votes": null,
      "replies": [
        {
          "id": 3232012,
          "author_name": "jamalsaeedi",
          "author_url": "",
          "post_date": "06/25/2025 09:20:47",
          "content": "<p>how to draw velocity from param?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3229182": "Hi everyone,\n\nLike many of you, I've noticed that certain datasets like `CurveFault_B`, `CurveVel_B`, and `Style_B` are particularly challenging and seem to be a bottleneck for improving our scores.\n\nI believe that creating a robust data augmentation pipeline for these difficult cases could be a key to success. To that end, I've been working on reproducing the velocity map generation process described in Appendix B of the OpenFWI paper.\n\n### What I've Done So Far\n\nI've managed to create a notebook that partially replicates the generation process for the `CurveVel` and `CurveFault` families based on the formulas in the paper.\n\n**You can find all the code and a detailed explanation in my notebook here:**\n\nData Augmentation Notebook: [here](https://www.kaggle.com/code/haruiig/fwi-velocity-map-augmentation-for-curvefault-vel#%E3%81%AF%E3%81%98%E3%82%81%E3%81%AB)\n\nHere’s a quick look at the results. The generation for `CurveVel` looks quite promising:\n\n![Transformed Vel Map to generate CurveVel](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22004535%2Fccfa101f5b2263fb758229fdd363e2fa%2FScreenshot%20from%202025-06-21%2014-17-05.png?generation=1750483117816171&alt=media)\n\nOriginal Vel Map:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22004535%2F6a75b063509748c442d655544337f73d%2FScreenshot%20from%202025-06-21%2014-19-27.png?generation=1750483192841609&alt=media)\n\nThe `CurveFault` generation is a bit more complex due to the larger number of parameters, and it's still a work-in-progress, but it's a start:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22004535%2Fdc74e47208477c8db0b44bde8aacd9f9%2FScreenshot%20from%202025-06-21%2014-20-22.png?generation=1750483236072461&alt=media)\n\n\n### The Challenge & How You Can Help\n\nWhile this is a good first step, there are still many unknowns to create a truly effective augmentation pipeline. This is where I'd love to get the community's help!\n\nThe main open questions are:\n\n*   **Parameter Distributions:** What are the appropriate ranges and distributions for sampling the parameters (`a, k, s, s'`, etc.)?\n*   **Number of Iterations:** How many times are these transformations recursively applied to create the final complex maps?\n*   **Initial Flat Maps:** How are the initial flat-layered maps generated in the first place?\n\n### An Idea for Reverse-Engineering\n\nOne idea I had for `CurveVel` is to **reverse-engineer the parameters from the existing training data**. For example, we could potentially apply a **Fourier Transform** to the provided `CurveVel` maps to estimate the distribution of the key wave parameters (`a` and `k`).\n\n**If anyone has experience with this or would be interested in trying to implement it, it would be a huge contribution!** Or, if you have a better idea, I'm all ears!\n\nLet's work together to decode this generation process and build a powerful augmentation strategy that benefits everyone. I look forward to your thoughts and ideas",
    "3229201": "thanks for the code. i am interested to know which is the winner here:\n1) 20x more data  + unet + mae loss\n2) 1x data  + unet + mae loss + physics loss (FWM)\n\nit may turn out the at the equivalent point, training time is the same",
    "3229315": "i believe 20x more data beats phyiscs loss + 1x data! (if we can reproduce the vel-map generation process correctly)\nusing more data, ml model can learn constraints of the model map(like continuity of velocity, the number of layers, how curvy velocity map be) precisely.\nless data and phyiscs loss cannot incorporate the constraints well, i believe.\nOf course, it's biased😂",
    "3229331": "why don't do this:\n\n```\n0. given test semsic, initialise a, k, s, ... = random\n\n1. velocity 70x70 = Draw(a, k, s, ....)\n\n2. semsic = FWM(velocity)\n\n3. mae loss = |semsic- test semsic|\n\n#or GAN loss if you want to just generate samples\n\n4. do gradient descend and update a, k, s, ...\n\n```\n\ninstead of trial and error, you can uncover the parameters by:\n\n```\n1) train velocty -> encoder -> latent -> MLP -> param =a, k, s, ...\n2) param -> Draw -> velocity1\n\nloss =|| velocity1 - train velocty||\ndist loss = KL diverge (latent, standard guassian)\n\n back propagate ... then you get distrubution of predicted param ...\n\n#or loss = GAN loss ... then you need to train discrminator\n````",
    "3229600": "hi, what is FWM loss?",
    "3229621": "Either the physical loss has been non - converging, or the top team has been continuously generating synthetic data. Otherwise, it is difficult to explain why the score can keep improving.",
    "3229686": "20x more data??? Hahaha I might have to tap out of this one.",
    "3229699": "The physic loss is non convex. You good solver.",
    "3229916": "The impact of synthetic data doesn't seem to be that significant (?), in the end, it still comes down to \"patiently solving the problem.\"",
    "3229917": "Depends. If u can generate test alike data, then you win",
    "3229925": ">Depends. If u can generate test alike data, then you win  \n\nAin't there a public notebook that generate pretty accurate test alike data? Does not seems too difficult.",
    "3229930": "All data science competitions are about\n1. A model that is large enough to eat and crunch available data\n2. Training data that is close to the hidden test\n3. Modeling less related to data/ML but more to methods … aka domain knowledge, physics eqn/ simulation, retrieval and matching, new modality/view, pre/post processing and novel\n solution steps/flow/pipeline\n\nIf you are not getting test lb score to your training loss, the data are not similar enough\n\n—-\n\nCurrent tip solution is below 10. So I assume that his validation loss is similar and the train error is much lower. If validation error has reached a point that is low enough, his model basically has duplicated the whole physics process. \n\nHe either has very good model or architecture that is same os physics equation or enough train data such that now test data are cinterpolation or extrapolation of the train.",
    "3232012": "how to draw velocity from param?"
  },
  "source": "meta"
}