{
  "id": 585815,
  "title": "Sharing a few ideas (+ torch vel_to_seis)",
  "url": "/competitions/waveform-inversion/discussion/585815",
  "author_name": "gguillard",
  "post_date": "2025-06-23T09:10:52.063000",
  "votes": 8,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Since I won’t have enough time to build, test and validate my ideas before the competition ends (too bad I wasted days <a href=\"https://www.kaggle.com/discussions/product-feedback/574221\" target=\"_blank\">debugging TPUs</a>…), I decided to share the ones I like best here, maybe someone will find some useful.  I’m eager to see if some of these strategies are already parts of some leaderboard leaders’ workflows.</p>\n<p>I’m open to teaming if anyone is interested — I can spare 2 or 3 days and still have 29&nbsp;h of Kaggle GPU left this week, FWIW.  Also, I have starter code for most (if not all) of the strategies listed below.  (NB : I didn’t submit using any of these strategies, except for some of the \"FlatVels are easy\" ones, which only marginally improved <a href=\"https://www.kaggle.com/code/overvalueawareness/caformer-full-resolution-improved\" target=\"_blank\">Overvalue Awareness’s public notebook</a> : same rounded LB score.)</p>\n<p>Before starting, there seems to be interest in a Torch version of the vel_to_seis function, <a href=\"https://www.kaggle.com/code/gguillard/torch-vel-to-seis/\" target=\"_blank\">here is one implementation</a> (8&nbsp;s -&gt; 0.5&nbsp;s, global MAE 1e-5).</p>\n<p>Note that most of the ideas below focus on CurveFault_B, since this is the MAE bottleneck when you start from <a href=\"https://www.kaggle.com/code/brendanartley/caformer-full-resolution-improved\" target=\"_blank\">Bartley’s last notebook</a> or its derivatives : it would be more effective to gain only 10&nbsp;% on CurveFault_B’s MAE than to e.g. completely suppress FlatVel_B’s reconstruction error.</p>\n<p>Also, I’m sorry many of these ideas are unfortunately more aimed at gaming the system than improving the state of the art of waveform inversion… (: </p>\n<h1>Observations</h1>\n<h2>Velocity distribution depends on the category</h2>\n<p>While Style_A/B maps have a continuous distribution of velocities, the other categories have a discrete distribution of velocities, with peculiar features which makes categorization somehow easy.</p>\n<h2>CAFormer’s MAE is not homogeneous</h2>\n<p>The MAE at the center of the map is globally 7 times higher than on the edges.  This information could be used to fine-tune the model with a weighted map.</p>\n<h2>The CurveFault_B jigsaw puzzle</h2>\n<p>This one would require a lot of work, but…  CurveFault_B maps seem to be patchworks of simpler maps, mostly CurveVel_A/B.  Maybe one could train a model to identify the puzzle pieces ?</p>\n<h2>Less data at high velocities</h2>\n<p>The global distribution of velocities is flat until about 3750&nbsp;m/s, then breaks down.  Data augmentation to make it flat is straightforward to implement.</p>\n<h1>Strategies</h1>\n<h2>FlatVels are easy</h2>\n<p>Flat maps can be identified from the sources symmetry.  Then a single source is enough to reconstruct the whole map.  In addition, the speed and depth of the first layer can be extracted numerically.</p>\n<p>However the improvement is very small, only of interest in the case of ties (I didn’t check but I guess this information can also probably improve the reconstruction on other categories, though).</p>\n<h2>Find  peaks</h2>\n<p>Since velocities are discrete, one should be able to improve the resolution by identify peaks in the reconstructed map, or to constrain the model.</p>\n<h2>Focus on badly reconstructed maps</h2>\n<p>There are proxies allowing to estimate the reconstruction error without knowing the true map.  With such a quality estimator, and since the larger the individual MAE, the more it contributes to the global MAE, one can focus on improving the maps with the largest error (I mean on the test set directly).</p>\n<h2>Image processing</h2>\n<p>I found this idea very promising, but again it would require a lot of work to make a robust implementation : apply edge detection, check if the edges are regular curves, and correct where they’re not.</p>\n<h2>Inter-source interpolation</h2>\n<p>One could simulate the signal for a number of sources between the known 5 sources, and train the main model on it.  Then another model would be trained to interpolate the new sources signals from the 5 known sources.  Finally, the interpolated sources are added to the test data before inference.</p>\n<p>Feedback welcome, especially if you already tried some of these strategies and it didn’t work out. ;)</p>",
  "messages": [
    {
      "id": 3230626,
      "postDate": "2025-06-23T09:10:52.063Z",
      "content": "<p>Since I won’t have enough time to build, test and validate my ideas before the competition ends (too bad I wasted days <a href=\"https://www.kaggle.com/discussions/product-feedback/574221\" target=\"_blank\">debugging TPUs</a>…), I decided to share the ones I like best here, maybe someone will find some useful.  I’m eager to see if some of these strategies are already parts of some leaderboard leaders’ workflows.</p>\n<p>I’m open to teaming if anyone is interested — I can spare 2 or 3 days and still have 29&nbsp;h of Kaggle GPU left this week, FWIW.  Also, I have starter code for most (if not all) of the strategies listed below.  (NB : I didn’t submit using any of these strategies, except for some of the \"FlatVels are easy\" ones, which only marginally improved <a href=\"https://www.kaggle.com/code/overvalueawareness/caformer-full-resolution-improved\" target=\"_blank\">Overvalue Awareness’s public notebook</a> : same rounded LB score.)</p>\n<p>Before starting, there seems to be interest in a Torch version of the vel_to_seis function, <a href=\"https://www.kaggle.com/code/gguillard/torch-vel-to-seis/\" target=\"_blank\">here is one implementation</a> (8&nbsp;s -&gt; 0.5&nbsp;s, global MAE 1e-5).</p>\n<p>Note that most of the ideas below focus on CurveFault_B, since this is the MAE bottleneck when you start from <a href=\"https://www.kaggle.com/code/brendanartley/caformer-full-resolution-improved\" target=\"_blank\">Bartley’s last notebook</a> or its derivatives : it would be more effective to gain only 10&nbsp;% on CurveFault_B’s MAE than to e.g. completely suppress FlatVel_B’s reconstruction error.</p>\n<p>Also, I’m sorry many of these ideas are unfortunately more aimed at gaming the system than improving the state of the art of waveform inversion… (: </p>\n<h1>Observations</h1>\n<h2>Velocity distribution depends on the category</h2>\n<p>While Style_A/B maps have a continuous distribution of velocities, the other categories have a discrete distribution of velocities, with peculiar features which makes categorization somehow easy.</p>\n<h2>CAFormer’s MAE is not homogeneous</h2>\n<p>The MAE at the center of the map is globally 7 times higher than on the edges.  This information could be used to fine-tune the model with a weighted map.</p>\n<h2>The CurveFault_B jigsaw puzzle</h2>\n<p>This one would require a lot of work, but…  CurveFault_B maps seem to be patchworks of simpler maps, mostly CurveVel_A/B.  Maybe one could train a model to identify the puzzle pieces ?</p>\n<h2>Less data at high velocities</h2>\n<p>The global distribution of velocities is flat until about 3750&nbsp;m/s, then breaks down.  Data augmentation to make it flat is straightforward to implement.</p>\n<h1>Strategies</h1>\n<h2>FlatVels are easy</h2>\n<p>Flat maps can be identified from the sources symmetry.  Then a single source is enough to reconstruct the whole map.  In addition, the speed and depth of the first layer can be extracted numerically.</p>\n<p>However the improvement is very small, only of interest in the case of ties (I didn’t check but I guess this information can also probably improve the reconstruction on other categories, though).</p>\n<h2>Find  peaks</h2>\n<p>Since velocities are discrete, one should be able to improve the resolution by identify peaks in the reconstructed map, or to constrain the model.</p>\n<h2>Focus on badly reconstructed maps</h2>\n<p>There are proxies allowing to estimate the reconstruction error without knowing the true map.  With such a quality estimator, and since the larger the individual MAE, the more it contributes to the global MAE, one can focus on improving the maps with the largest error (I mean on the test set directly).</p>\n<h2>Image processing</h2>\n<p>I found this idea very promising, but again it would require a lot of work to make a robust implementation : apply edge detection, check if the edges are regular curves, and correct where they’re not.</p>\n<h2>Inter-source interpolation</h2>\n<p>One could simulate the signal for a number of sources between the known 5 sources, and train the main model on it.  Then another model would be trained to interpolate the new sources signals from the 5 known sources.  Finally, the interpolated sources are added to the test data before inference.</p>\n<p>Feedback welcome, especially if you already tried some of these strategies and it didn’t work out. ;)</p>",
      "rawMarkdown": "Since I won’t have enough time to build, test and validate my ideas before the competition ends (too bad I wasted days [debugging TPUs](https://www.kaggle.com/discussions/product-feedback/574221)…), I decided to share the ones I like best here, maybe someone will find some useful.  I’m eager to see if some of these strategies are already parts of some leaderboard leaders’ workflows.\n\nI’m open to teaming if anyone is interested — I can spare 2 or 3 days and still have 29 h of Kaggle GPU left this week, FWIW.  Also, I have starter code for most (if not all) of the strategies listed below.  (NB : I didn’t submit using any of these strategies, except for some of the \"FlatVels are easy\" ones, which only marginally improved [Overvalue Awareness’s public notebook](https://www.kaggle.com/code/overvalueawareness/caformer-full-resolution-improved) : same rounded LB score.)\n\nBefore starting, there seems to be interest in a Torch version of the vel_to_seis function, [here is one implementation](https://www.kaggle.com/code/gguillard/torch-vel-to-seis/) (8 s -> 0.5 s, global MAE 1e-5).\n\nNote that most of the ideas below focus on CurveFault_B, since this is the MAE bottleneck when you start from [Bartley’s last notebook](https://www.kaggle.com/code/brendanartley/caformer-full-resolution-improved) or its derivatives : it would be more effective to gain only 10 % on CurveFault_B’s MAE than to e.g. completely suppress FlatVel_B’s reconstruction error.\n\nAlso, I’m sorry many of these ideas are unfortunately more aimed at gaming the system than improving the state of the art of waveform inversion… (: \n\n# Observations\n\n## Velocity distribution depends on the category\n\nWhile Style_A/B maps have a continuous distribution of velocities, the other categories have a discrete distribution of velocities, with peculiar features which makes categorization somehow easy.\n\n## CAFormer’s MAE is not homogeneous\n\nThe MAE at the center of the map is globally 7 times higher than on the edges.  This information could be used to fine-tune the model with a weighted map.\n\n## The CurveFault_B jigsaw puzzle\n\nThis one would require a lot of work, but…  CurveFault_B maps seem to be patchworks of simpler maps, mostly CurveVel_A/B.  Maybe one could train a model to identify the puzzle pieces ?\n\n## Less data at high velocities\n\nThe global distribution of velocities is flat until about 3750 m/s, then breaks down.  Data augmentation to make it flat is straightforward to implement.\n\n# Strategies\n\n## FlatVels are easy\n\nFlat maps can be identified from the sources symmetry.  Then a single source is enough to reconstruct the whole map.  In addition, the speed and depth of the first layer can be extracted numerically.\n\nHowever the improvement is very small, only of interest in the case of ties (I didn’t check but I guess this information can also probably improve the reconstruction on other categories, though).\n\n## Find  peaks\n\nSince velocities are discrete, one should be able to improve the resolution by identify peaks in the reconstructed map, or to constrain the model.\n\n## Focus on badly reconstructed maps\n\nThere are proxies allowing to estimate the reconstruction error without knowing the true map.  With such a quality estimator, and since the larger the individual MAE, the more it contributes to the global MAE, one can focus on improving the maps with the largest error (I mean on the test set directly).\n\n## Image processing\n\nI found this idea very promising, but again it would require a lot of work to make a robust implementation : apply edge detection, check if the edges are regular curves, and correct where they’re not.\n\n## Inter-source interpolation\n\nOne could simulate the signal for a number of sources between the known 5 sources, and train the main model on it.  Then another model would be trained to interpolate the new sources signals from the 5 known sources.  Finally, the interpolated sources are added to the test data before inference.\n\nFeedback welcome, especially if you already tried some of these strategies and it didn’t work out. ;)",
      "votes": 8
    },
    {
      "id": 3230738,
      "postDate": "2025-06-23T11:55:38.970Z",
      "content": "<p>Wow, I wasn't expecting so many downvotes (at least 3 at this time).  Is there anything wrong with my post&nbsp;? 🤨</p>",
      "rawMarkdown": "Wow, I wasn't expecting so many downvotes (at least 3 at this time).  Is there anything wrong with my post ? 🤨",
      "votes": 1,
      "replies": [
        {
          "id": 3230757,
          "postDate": "2025-06-23T12:19:31.053Z",
          "content": "<p>Some people just don't want there to be so many shares in the final week, as it could affect the rankings.</p>",
          "rawMarkdown": "Some people just don't want there to be so many shares in the final week, as it could affect the rankings.",
          "votes": 2,
          "replies": [
            {
              "id": 3230781,
              "postDate": "2025-06-23T12:36:50.090Z",
              "content": "<p>Thanks, that makes sense, I get it (although technically, \"shares [that] could affect the rankings\" occur throughout the whole competition, sometimes <a href=\"https://www.kaggle.com/competitions/waveform-inversion/discussion/579841#3206525\" target=\"_blank\">dramatically</a>…).</p>",
              "rawMarkdown": "Thanks, that makes sense, I get it (although technically, \"shares [that] could affect the rankings\" occur throughout the whole competition, sometimes [dramatically](https://www.kaggle.com/competitions/waveform-inversion/discussion/579841#3206525)…)."
            },
            {
              "id": 3230799,
              "postDate": "2025-06-23T12:57:11.713Z",
              "content": "<p><a href=\"https://www.kaggle.com/discussions/general/291540\" target=\"_blank\">Last week is different</a>.  <br>\nTechnically, you are still good since there are several hours to last week (this is why I don't downvote), but it's very close to the deadline so yeah, downvotes are expected.</p>",
              "rawMarkdown": "[Last week is different](https://www.kaggle.com/discussions/general/291540).  \nTechnically, you are still good since there are several hours to last week (this is why I don't downvote), but it's very close to the deadline so yeah, downvotes are expected."
            },
            {
              "id": 3230834,
              "postDate": "2025-06-23T14:00:36.247Z",
              "content": "<p>Thanks for the link and feedback.  I didn’t share any code though (besides the torch version of vel_to_seis, which isn’t a game changer and for which other implementations already exist anyway).</p>\n<p>It’s much easier to throw (and read) ideas than to implement them.</p>",
              "rawMarkdown": "Thanks for the link and feedback.  I didn’t share any code though (besides the torch version of vel_to_seis, which isn’t a game changer and for which other implementations already exist anyway).\n\nIt’s much easier to throw (and read) ideas than to implement them.",
              "votes": 1
            },
            {
              "id": 3230839,
              "postDate": "2025-06-23T14:10:09.947Z",
              "content": "<blockquote>\n  <blockquote>\n    <p>I didn’t share any code though  </p>\n  </blockquote>\n</blockquote>\n<p>Doesn't matter. Yes the rules refer specifically to code, however, the community frown upon any meaningful last-week sharing and will show it with downvotes.</p>",
              "rawMarkdown": ">>I didn’t share any code though  \n\nDoesn't matter. Yes the rules refer specifically to code, however, the community frown upon any meaningful last-week sharing and will show it with downvotes.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3235975,
      "postDate": "2025-06-29T20:29:48.030Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3230738,
      "author_name": "gguillard",
      "author_url": "",
      "post_date": "2025-06-23T11:55:38.970000",
      "content": "<p>Wow, I wasn't expecting so many downvotes (at least 3 at this time).  Is there anything wrong with my post&nbsp;? 🤨</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3230757,
          "author_name": "atom1231",
          "author_url": "",
          "post_date": "2025-06-23T12:19:31.053000",
          "content": "<p>Some people just don't want there to be so many shares in the final week, as it could affect the rankings.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3230781,
              "author_name": "gguillard",
              "author_url": "",
              "post_date": "2025-06-23T12:36:50.090000",
              "content": "<p>Thanks, that makes sense, I get it (although technically, \"shares [that] could affect the rankings\" occur throughout the whole competition, sometimes <a href=\"https://www.kaggle.com/competitions/waveform-inversion/discussion/579841#3206525\" target=\"_blank\">dramatically</a>…).</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3230799,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2025-06-23T12:57:11.713000",
              "content": "<p><a href=\"https://www.kaggle.com/discussions/general/291540\" target=\"_blank\">Last week is different</a>.  <br>\nTechnically, you are still good since there are several hours to last week (this is why I don't downvote), but it's very close to the deadline so yeah, downvotes are expected.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3230834,
              "author_name": "gguillard",
              "author_url": "",
              "post_date": "2025-06-23T14:00:36.247000",
              "content": "<p>Thanks for the link and feedback.  I didn’t share any code though (besides the torch version of vel_to_seis, which isn’t a game changer and for which other implementations already exist anyway).</p>\n<p>It’s much easier to throw (and read) ideas than to implement them.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3230839,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2025-06-23T14:10:09.947000",
              "content": "<blockquote>\n  <blockquote>\n    <p>I didn’t share any code though  </p>\n  </blockquote>\n</blockquote>\n<p>Doesn't matter. Yes the rules refer specifically to code, however, the community frown upon any meaningful last-week sharing and will show it with downvotes.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3235975,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-06-29T20:29:48.030000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3230626": "Since I won’t have enough time to build, test and validate my ideas before the competition ends (too bad I wasted days [debugging TPUs](https://www.kaggle.com/discussions/product-feedback/574221)…), I decided to share the ones I like best here, maybe someone will find some useful.  I’m eager to see if some of these strategies are already parts of some leaderboard leaders’ workflows.\n\nI’m open to teaming if anyone is interested — I can spare 2 or 3 days and still have 29 h of Kaggle GPU left this week, FWIW.  Also, I have starter code for most (if not all) of the strategies listed below.  (NB : I didn’t submit using any of these strategies, except for some of the \"FlatVels are easy\" ones, which only marginally improved [Overvalue Awareness’s public notebook](https://www.kaggle.com/code/overvalueawareness/caformer-full-resolution-improved) : same rounded LB score.)\n\nBefore starting, there seems to be interest in a Torch version of the vel_to_seis function, [here is one implementation](https://www.kaggle.com/code/gguillard/torch-vel-to-seis/) (8 s -> 0.5 s, global MAE 1e-5).\n\nNote that most of the ideas below focus on CurveFault_B, since this is the MAE bottleneck when you start from [Bartley’s last notebook](https://www.kaggle.com/code/brendanartley/caformer-full-resolution-improved) or its derivatives : it would be more effective to gain only 10 % on CurveFault_B’s MAE than to e.g. completely suppress FlatVel_B’s reconstruction error.\n\nAlso, I’m sorry many of these ideas are unfortunately more aimed at gaming the system than improving the state of the art of waveform inversion… (: \n\n# Observations\n\n## Velocity distribution depends on the category\n\nWhile Style_A/B maps have a continuous distribution of velocities, the other categories have a discrete distribution of velocities, with peculiar features which makes categorization somehow easy.\n\n## CAFormer’s MAE is not homogeneous\n\nThe MAE at the center of the map is globally 7 times higher than on the edges.  This information could be used to fine-tune the model with a weighted map.\n\n## The CurveFault_B jigsaw puzzle\n\nThis one would require a lot of work, but…  CurveFault_B maps seem to be patchworks of simpler maps, mostly CurveVel_A/B.  Maybe one could train a model to identify the puzzle pieces ?\n\n## Less data at high velocities\n\nThe global distribution of velocities is flat until about 3750 m/s, then breaks down.  Data augmentation to make it flat is straightforward to implement.\n\n# Strategies\n\n## FlatVels are easy\n\nFlat maps can be identified from the sources symmetry.  Then a single source is enough to reconstruct the whole map.  In addition, the speed and depth of the first layer can be extracted numerically.\n\nHowever the improvement is very small, only of interest in the case of ties (I didn’t check but I guess this information can also probably improve the reconstruction on other categories, though).\n\n## Find  peaks\n\nSince velocities are discrete, one should be able to improve the resolution by identify peaks in the reconstructed map, or to constrain the model.\n\n## Focus on badly reconstructed maps\n\nThere are proxies allowing to estimate the reconstruction error without knowing the true map.  With such a quality estimator, and since the larger the individual MAE, the more it contributes to the global MAE, one can focus on improving the maps with the largest error (I mean on the test set directly).\n\n## Image processing\n\nI found this idea very promising, but again it would require a lot of work to make a robust implementation : apply edge detection, check if the edges are regular curves, and correct where they’re not.\n\n## Inter-source interpolation\n\nOne could simulate the signal for a number of sources between the known 5 sources, and train the main model on it.  Then another model would be trained to interpolate the new sources signals from the 5 known sources.  Finally, the interpolated sources are added to the test data before inference.\n\nFeedback welcome, especially if you already tried some of these strategies and it didn’t work out. ;)",
    "3230738": "Wow, I wasn't expecting so many downvotes (at least 3 at this time).  Is there anything wrong with my post ? 🤨",
    "3235975": ""
  }
}