{
  "id": 587498,
  "title": "9th place - Custom CUDA kernel for wave propagation",
  "url": "/competitions/waveform-inversion/writeups/dieter-9th-place-custom-cuda-kernel-for-wave-propa",
  "author_name": "",
  "post_date": "2025-07-01T12:06:54.270Z",
  "votes": 59,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Thanks to kaggle and everyone involved for hosting this exciting competition. It was a great learning experience.</p>\n<h2>Introduction</h2>\n<p>I joined the competition only 2 weeks before end and wanted to make the most of the time in terms of learning experience. Coming from two computer vision competitions before, I was not very interesting in spending a lot of time on the modeling side. So I basically took over the two main models and hyperparameters from <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> and focused on other aspects of the competition. My main focus was to do efficient experimentation under time pressure when training takes a long time. I also wanted to try an \"idle speculation\" to just add some spice to this competition. </p>\n<h2>Cross validation</h2>\n<p>For cross-validation I simple split the data by file into 8 folds. In the first week I only trained and evaluated on CurveFault_B since this seemed to be the hardest, yet representative class.</p>\n<h2>Model</h2>\n<p>As said in the introduction, my models are copied from <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a>. I used the <code>caformer_b36.sail_in22k_ft_in1k</code> and same convnext set-up, although adjusted to run a bigger version, <code>convnextv2_huge.fcmae</code> . </p>\n<h2>Training</h2>\n<h3>General setup</h3>\n<p>The main difference to other competitors is probably the training part. On the one hand I wanted to recyle training runs as much as possible as MAE converges quite slowly. So I kept the lr flat at 0.0005 so I can just continue on the best experiment when trying new ideas. </p>\n<h3>The main magic</h3>\n<p>The main boost in score comes from augmentations. In order to augment the data, you basically need to transform the target and then do proper wave propagation to generate the corresponding seismogram. A starting code to do that closely enough to how organizers generated the seismograms was shared <a href=\"https://www.kaggle.com/code/manatoyo/improved-vel-to-seis\" target=\"_blank\">here</a> thanks to <a href=\"https://www.kaggle.com/manatoyo\" target=\"_blank\">@manatoyo</a> and <a href=\"https://www.kaggle.com/jaewook704\" target=\"_blank\">@jaewook704</a> . But that generation took 3s per target and was not viable to run on the fly while training. So, just for learning purpose I tried speed up the wave propagation as much as possible using GPUs. I spent days for improving the speed. First, I converted the wave propagation to pytorch on GPU and make it run with torch compile. Then I implemented a batched version of it. Finally, I created  a custom CUDA kernel for batched wave propagation. (very good learning for me). That was fast enough to directly use it in my model to recalculate seismograms whenever the target has been augmented. Speed-up compared to the original CPU implementation is about 100x. I used shift, scale, rotate augmentations and an own implementation for shifting intensities. </p>\n<h3>Inference</h3>\n<p>On the last 2 days, I took the few checkpoints I had (which were run on 0.0005 lr) and added another 150 epochs with cosine schedule, giving more sample weights to difficult classes. I saved test predictions every 3 epochs and took median of the last 10 saved test predictions, to get a single prediction per run. Then simply took mean over the runs. For time reasons I only trained around 6 models spread over 3 of my 8 folds.</p>\n<h2>Idle speculation</h2>\n<p>I asked 2 people for teaming, but was rejected. So I thought, why not try an <a href=\"https://www.kaggle.com/competitions/expedia-hotel-recommendations/leaderboard\" target=\"_blank\">idle speculation</a>, i.e. try to win with a single sub. That was a learning experience of its own, especially under time pressure. You need to manage training as long as possible, but also hedge the risk of failing platforms (compute/ data host plattform  as well as kaggle). So 2 days before end I already had a \"backup\" submission ready as csv file in case anything crashes. My solution is definitely worse in score because I only used a single submission approach. For example I only realized 2 days before end, that you only need to predict every 2nd pixel. If I realized that earlier I would have been able to train twice as fast. Also how models and folds are selected and weighted was just by gut feeling. Although I see high potential of my solution especially if I would have put any effort in modelling, I am very happy that my submission worked out and I achieved solo gold in this restricted time window. </p>\n<p>Thanks for reading. </p>",
  "messages": [
    {
      "id": "3237746",
      "postDate": "07/01/2025 09:21:48",
      "content": "<p>Thanks to kaggle and everyone involved for hosting this exciting competition. It was a great learning experience.</p>\n<h2>Introduction</h2>\n<p>I joined the competition only 2 weeks before end and wanted to make the most of the time in terms of learning experience. Coming from two computer vision competitions before, I was not very interesting in spending a lot of time on the modeling side. So I basically took over the two main models and hyperparameters from <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> and focused on other aspects of the competition. My main focus was to do efficient experimentation under time pressure when training takes a long time. I also wanted to try an \"idle speculation\" to just add some spice to this competition. </p>\n<h2>Cross validation</h2>\n<p>For cross-validation I simple split the data by file into 8 folds. In the first week I only trained and evaluated on CurveFault_B since this seemed to be the hardest, yet representative class.</p>\n<h2>Model</h2>\n<p>As said in the introduction, my models are copied from <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a>. I used the <code>caformer_b36.sail_in22k_ft_in1k</code> and same convnext set-up, although adjusted to run a bigger version, <code>convnextv2_huge.fcmae</code> . </p>\n<h2>Training</h2>\n<h3>General setup</h3>\n<p>The main difference to other competitors is probably the training part. On the one hand I wanted to recyle training runs as much as possible as MAE converges quite slowly. So I kept the lr flat at 0.0005 so I can just continue on the best experiment when trying new ideas. </p>\n<h3>The main magic</h3>\n<p>The main boost in score comes from augmentations. In order to augment the data, you basically need to transform the target and then do proper wave propagation to generate the corresponding seismogram. A starting code to do that closely enough to how organizers generated the seismograms was shared <a href=\"https://www.kaggle.com/code/manatoyo/improved-vel-to-seis\" target=\"_blank\">here</a> thanks to <a href=\"https://www.kaggle.com/manatoyo\" target=\"_blank\">@manatoyo</a> and <a href=\"https://www.kaggle.com/jaewook704\" target=\"_blank\">@jaewook704</a> . But that generation took 3s per target and was not viable to run on the fly while training. So, just for learning purpose I tried speed up the wave propagation as much as possible using GPUs. I spent days for improving the speed. First, I converted the wave propagation to pytorch on GPU and make it run with torch compile. Then I implemented a batched version of it. Finally, I created  a custom CUDA kernel for batched wave propagation. (very good learning for me). That was fast enough to directly use it in my model to recalculate seismograms whenever the target has been augmented. Speed-up compared to the original CPU implementation is about 100x. I used shift, scale, rotate augmentations and an own implementation for shifting intensities. </p>\n<h3>Inference</h3>\n<p>On the last 2 days, I took the few checkpoints I had (which were run on 0.0005 lr) and added another 150 epochs with cosine schedule, giving more sample weights to difficult classes. I saved test predictions every 3 epochs and took median of the last 10 saved test predictions, to get a single prediction per run. Then simply took mean over the runs. For time reasons I only trained around 6 models spread over 3 of my 8 folds.</p>\n<h2>Idle speculation</h2>\n<p>I asked 2 people for teaming, but was rejected. So I thought, why not try an <a href=\"https://www.kaggle.com/competitions/expedia-hotel-recommendations/leaderboard\" target=\"_blank\">idle speculation</a>, i.e. try to win with a single sub. That was a learning experience of its own, especially under time pressure. You need to manage training as long as possible, but also hedge the risk of failing platforms (compute/ data host plattform  as well as kaggle). So 2 days before end I already had a \"backup\" submission ready as csv file in case anything crashes. My solution is definitely worse in score because I only used a single submission approach. For example I only realized 2 days before end, that you only need to predict every 2nd pixel. If I realized that earlier I would have been able to train twice as fast. Also how models and folds are selected and weighted was just by gut feeling. Although I see high potential of my solution especially if I would have put any effort in modelling, I am very happy that my submission worked out and I achieved solo gold in this restricted time window. </p>\n<p>Thanks for reading. </p>",
      "rawMarkdown": "Thanks to kaggle and everyone involved for hosting this exciting competition. It was a great learning experience.\n\n## Introduction\nI joined the competition only 2 weeks before end and wanted to make the most of the time in terms of learning experience. Coming from two computer vision competitions before, I was not very interesting in spending a lot of time on the modeling side. So I basically took over the two main models and hyperparameters from @brendanartley and focused on other aspects of the competition. My main focus was to do efficient experimentation under time pressure when training takes a long time. I also wanted to try an \"idle speculation\" to just add some spice to this competition. \n\n## Cross validation\nFor cross-validation I simple split the data by file into 8 folds. In the first week I only trained and evaluated on CurveFault_B since this seemed to be the hardest, yet representative class.\n\n## Model\n\nAs said in the introduction, my models are copied from @brendanartley. I used the `caformer_b36.sail_in22k_ft_in1k` and same convnext set-up, although adjusted to run a bigger version, `convnextv2_huge.fcmae` . \n\n## Training\n\n### General setup\nThe main difference to other competitors is probably the training part. On the one hand I wanted to recyle training runs as much as possible as MAE converges quite slowly. So I kept the lr flat at 0.0005 so I can just continue on the best experiment when trying new ideas. \n\n### The main magic\nThe main boost in score comes from augmentations. In order to augment the data, you basically need to transform the target and then do proper wave propagation to generate the corresponding seismogram. A starting code to do that closely enough to how organizers generated the seismograms was shared [here](https://www.kaggle.com/code/manatoyo/improved-vel-to-seis) thanks to @manatoyo and @jaewook704 . But that generation took 3s per target and was not viable to run on the fly while training. So, just for learning purpose I tried speed up the wave propagation as much as possible using GPUs. I spent days for improving the speed. First, I converted the wave propagation to pytorch on GPU and make it run with torch compile. Then I implemented a batched version of it. Finally, I created  a custom CUDA kernel for batched wave propagation. (very good learning for me). That was fast enough to directly use it in my model to recalculate seismograms whenever the target has been augmented. Speed-up compared to the original CPU implementation is about 100x. I used shift, scale, rotate augmentations and an own implementation for shifting intensities. \n\n### Inference\nOn the last 2 days, I took the few checkpoints I had (which were run on 0.0005 lr) and added another 150 epochs with cosine schedule, giving more sample weights to difficult classes. I saved test predictions every 3 epochs and took median of the last 10 saved test predictions, to get a single prediction per run. Then simply took mean over the runs. For time reasons I only trained around 6 models spread over 3 of my 8 folds.\n\n\n## Idle speculation\nI asked 2 people for teaming, but was rejected. So I thought, why not try an [idle speculation](https://www.kaggle.com/competitions/expedia-hotel-recommendations/leaderboard), i.e. try to win with a single sub. That was a learning experience of its own, especially under time pressure. You need to manage training as long as possible, but also hedge the risk of failing platforms (compute/ data host plattform  as well as kaggle). So 2 days before end I already had a \"backup\" submission ready as csv file in case anything crashes. My solution is definitely worse in score because I only used a single submission approach. For example I only realized 2 days before end, that you only need to predict every 2nd pixel. If I realized that earlier I would have been able to train twice as fast. Also how models and folds are selected and weighted was just by gut feeling. Although I see high potential of my solution especially if I would have put any effort in modelling, I am very happy that my submission worked out and I achieved solo gold in this restricted time window. \n\nThanks for reading.",
      "votes": null
    },
    {
      "id": "3237788",
      "postDate": "07/01/2025 10:02:19",
      "content": "<p>Deepseek told me to write cuda kernel also. However, my coding skills aren't that strong, so I gave up after it failed a few times. Being able to pull that off would have been a total game-changer for this competition. Two weeks was definitely too tight!</p>",
      "rawMarkdown": "Deepseek told me to write cuda kernel also. However, my coding skills aren't that strong, so I gave up after it failed a few times. Being able to pull that off would have been a total game-changer for this competition. Two weeks was definitely too tight!",
      "votes": null
    },
    {
      "id": "3237801",
      "postDate": "07/01/2025 10:11:11",
      "content": "<blockquote>\n  <blockquote>\n    <p>If I realized that earlier I would have been able to train twice as fast.   </p>\n  </blockquote>\n</blockquote>\n<p>Not necessarily, it may (my gut feeling is it would) have hurt accuracy too much (did not try).</p>",
      "rawMarkdown": ">> If I realized that earlier I would have been able to train twice as fast.   \n\nNot necessarily, it may (my gut feeling is it would) have hurt accuracy too much (did not try).",
      "votes": null
    },
    {
      "id": "3237822",
      "postDate": "07/01/2025 10:29:38",
      "content": "<p>Ok, this approach and solution is on another level. Holly smokes</p>\n<p>Kaggle (or more broadly, problem solving) at its best. What a joy to read 🙏🏻</p>",
      "rawMarkdown": "Ok, this approach and solution is on another level. Holly smokes\n\nKaggle (or more broadly, problem solving) at its best. What a joy to read 🙏🏻",
      "votes": null
    },
    {
      "id": "3237838",
      "postDate": "07/01/2025 10:35:57",
      "content": "<p>I quickly tried on last day, it works without accuracy loss. Just downsample axis by 2 in in stem and predict 70x35. </p>",
      "rawMarkdown": "I quickly tried on last day, it works without accuracy loss. Just downsample axis by 2 in in stem and predict 70x35.",
      "votes": null
    },
    {
      "id": "3237856",
      "postDate": "07/01/2025 10:50:01",
      "content": "<p>Interesting! Now lot of people probably feel stupid for not doing it. I am. 🤣</p>",
      "rawMarkdown": "Interesting! Now lot of people probably feel stupid for not doing it. I am. 🤣",
      "votes": null
    },
    {
      "id": "3237909",
      "postDate": "07/01/2025 11:20:51",
      "content": "<p>I did not realize that I should reuse weight until 2 days before the end of the competition … </p>",
      "rawMarkdown": "I did not realize that I should reuse weight until 2 days before the end of the competition ...",
      "votes": null
    },
    {
      "id": "3238338",
      "postDate": "07/01/2025 18:47:22",
      "content": "<p>Amazing you managed gold in such a short time! Would you be willing to share the cuda kernel you wrote? </p>",
      "rawMarkdown": "Amazing you managed gold in such a short time! Would you be willing to share the cuda kernel you wrote?",
      "votes": null
    },
    {
      "id": "3238583",
      "postDate": "07/02/2025 02:16:04",
      "content": "<p>Good experience with one submission. And another confirmation that data in general and learning tricks is more important than model architecture, such results inspire freshes to think more in the direction of working with data and organizing the entire data pipeline not only modern ml modeling. Great instance and job! </p>",
      "rawMarkdown": "Good experience with one submission. And another confirmation that data in general and learning tricks is more important than model architecture, such results inspire freshes to think more in the direction of working with data and organizing the entire data pipeline not only modern ml modeling. Great instance and job!",
      "votes": null
    },
    {
      "id": "3238964",
      "postDate": "07/02/2025 11:08:27",
      "content": "<p>I wonder why rejected.<br>\nBecause having you who is consistently rank no.1 is like almost guarantee good things will happen.</p>",
      "rawMarkdown": "I wonder why rejected.\nBecause having you who is consistently rank no.1 is like almost guarantee good things will happen.",
      "votes": null
    },
    {
      "id": "3239032",
      "postDate": "07/02/2025 12:14:14",
      "content": "<p>Sometimes people prefer to be solo for solo medal, or solo/small team for larger price share ^^ (I was not among the two people, but I did reject another very good offer because I was hungry for price haha)</p>",
      "rawMarkdown": "Sometimes people prefer to be solo for solo medal, or solo/small team for larger price share ^^ (I was not among the two people, but I did reject another very good offer because I was hungry for price haha)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3237788,
      "author_name": "hydantess",
      "author_url": "",
      "post_date": "07/01/2025 10:02:19",
      "content": "<p>Deepseek told me to write cuda kernel also. However, my coding skills aren't that strong, so I gave up after it failed a few times. Being able to pull that off would have been a total game-changer for this competition. Two weeks was definitely too tight!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3237801,
      "author_name": "shlomoron",
      "author_url": "",
      "post_date": "07/01/2025 10:11:11",
      "content": "<blockquote>\n  <blockquote>\n    <p>If I realized that earlier I would have been able to train twice as fast.   </p>\n  </blockquote>\n</blockquote>\n<p>Not necessarily, it may (my gut feeling is it would) have hurt accuracy too much (did not try).</p>",
      "votes": null,
      "replies": [
        {
          "id": 3237838,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "07/01/2025 10:35:57",
          "content": "<p>I quickly tried on last day, it works without accuracy loss. Just downsample axis by 2 in in stem and predict 70x35. </p>",
          "votes": null,
          "replies": [
            {
              "id": 3237856,
              "author_name": "shlomoron",
              "author_url": "",
              "post_date": "07/01/2025 10:50:01",
              "content": "<p>Interesting! Now lot of people probably feel stupid for not doing it. I am. 🤣</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3237822,
      "author_name": "radek1",
      "author_url": "",
      "post_date": "07/01/2025 10:29:38",
      "content": "<p>Ok, this approach and solution is on another level. Holly smokes</p>\n<p>Kaggle (or more broadly, problem solving) at its best. What a joy to read 🙏🏻</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3237909,
      "author_name": "yuanzhezhou",
      "author_url": "",
      "post_date": "07/01/2025 11:20:51",
      "content": "<p>I did not realize that I should reuse weight until 2 days before the end of the competition … </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3238338,
      "author_name": "snehalverma10",
      "author_url": "",
      "post_date": "07/01/2025 18:47:22",
      "content": "<p>Amazing you managed gold in such a short time! Would you be willing to share the cuda kernel you wrote? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3238583,
      "author_name": "sggpls",
      "author_url": "",
      "post_date": "07/02/2025 02:16:04",
      "content": "<p>Good experience with one submission. And another confirmation that data in general and learning tricks is more important than model architecture, such results inspire freshes to think more in the direction of working with data and organizing the entire data pipeline not only modern ml modeling. Great instance and job! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3238964,
      "author_name": "overvalueawareness",
      "author_url": "",
      "post_date": "07/02/2025 11:08:27",
      "content": "<p>I wonder why rejected.<br>\nBecause having you who is consistently rank no.1 is like almost guarantee good things will happen.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3239032,
          "author_name": "shlomoron",
          "author_url": "",
          "post_date": "07/02/2025 12:14:14",
          "content": "<p>Sometimes people prefer to be solo for solo medal, or solo/small team for larger price share ^^ (I was not among the two people, but I did reject another very good offer because I was hungry for price haha)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3237746": "Thanks to kaggle and everyone involved for hosting this exciting competition. It was a great learning experience.\n\n## Introduction\nI joined the competition only 2 weeks before end and wanted to make the most of the time in terms of learning experience. Coming from two computer vision competitions before, I was not very interesting in spending a lot of time on the modeling side. So I basically took over the two main models and hyperparameters from @brendanartley and focused on other aspects of the competition. My main focus was to do efficient experimentation under time pressure when training takes a long time. I also wanted to try an \"idle speculation\" to just add some spice to this competition. \n\n## Cross validation\nFor cross-validation I simple split the data by file into 8 folds. In the first week I only trained and evaluated on CurveFault_B since this seemed to be the hardest, yet representative class.\n\n## Model\n\nAs said in the introduction, my models are copied from @brendanartley. I used the `caformer_b36.sail_in22k_ft_in1k` and same convnext set-up, although adjusted to run a bigger version, `convnextv2_huge.fcmae` . \n\n## Training\n\n### General setup\nThe main difference to other competitors is probably the training part. On the one hand I wanted to recyle training runs as much as possible as MAE converges quite slowly. So I kept the lr flat at 0.0005 so I can just continue on the best experiment when trying new ideas. \n\n### The main magic\nThe main boost in score comes from augmentations. In order to augment the data, you basically need to transform the target and then do proper wave propagation to generate the corresponding seismogram. A starting code to do that closely enough to how organizers generated the seismograms was shared [here](https://www.kaggle.com/code/manatoyo/improved-vel-to-seis) thanks to @manatoyo and @jaewook704 . But that generation took 3s per target and was not viable to run on the fly while training. So, just for learning purpose I tried speed up the wave propagation as much as possible using GPUs. I spent days for improving the speed. First, I converted the wave propagation to pytorch on GPU and make it run with torch compile. Then I implemented a batched version of it. Finally, I created  a custom CUDA kernel for batched wave propagation. (very good learning for me). That was fast enough to directly use it in my model to recalculate seismograms whenever the target has been augmented. Speed-up compared to the original CPU implementation is about 100x. I used shift, scale, rotate augmentations and an own implementation for shifting intensities. \n\n### Inference\nOn the last 2 days, I took the few checkpoints I had (which were run on 0.0005 lr) and added another 150 epochs with cosine schedule, giving more sample weights to difficult classes. I saved test predictions every 3 epochs and took median of the last 10 saved test predictions, to get a single prediction per run. Then simply took mean over the runs. For time reasons I only trained around 6 models spread over 3 of my 8 folds.\n\n\n## Idle speculation\nI asked 2 people for teaming, but was rejected. So I thought, why not try an [idle speculation](https://www.kaggle.com/competitions/expedia-hotel-recommendations/leaderboard), i.e. try to win with a single sub. That was a learning experience of its own, especially under time pressure. You need to manage training as long as possible, but also hedge the risk of failing platforms (compute/ data host plattform  as well as kaggle). So 2 days before end I already had a \"backup\" submission ready as csv file in case anything crashes. My solution is definitely worse in score because I only used a single submission approach. For example I only realized 2 days before end, that you only need to predict every 2nd pixel. If I realized that earlier I would have been able to train twice as fast. Also how models and folds are selected and weighted was just by gut feeling. Although I see high potential of my solution especially if I would have put any effort in modelling, I am very happy that my submission worked out and I achieved solo gold in this restricted time window. \n\nThanks for reading.",
    "3237788": "Deepseek told me to write cuda kernel also. However, my coding skills aren't that strong, so I gave up after it failed a few times. Being able to pull that off would have been a total game-changer for this competition. Two weeks was definitely too tight!",
    "3237801": ">> If I realized that earlier I would have been able to train twice as fast.   \n\nNot necessarily, it may (my gut feeling is it would) have hurt accuracy too much (did not try).",
    "3237822": "Ok, this approach and solution is on another level. Holly smokes\n\nKaggle (or more broadly, problem solving) at its best. What a joy to read 🙏🏻",
    "3237838": "I quickly tried on last day, it works without accuracy loss. Just downsample axis by 2 in in stem and predict 70x35.",
    "3237856": "Interesting! Now lot of people probably feel stupid for not doing it. I am. 🤣",
    "3237909": "I did not realize that I should reuse weight until 2 days before the end of the competition ...",
    "3238338": "Amazing you managed gold in such a short time! Would you be willing to share the cuda kernel you wrote?",
    "3238583": "Good experience with one submission. And another confirmation that data in general and learning tricks is more important than model architecture, such results inspire freshes to think more in the direction of working with data and organizing the entire data pipeline not only modern ml modeling. Great instance and job!",
    "3238964": "I wonder why rejected.\nBecause having you who is consistently rank no.1 is like almost guarantee good things will happen.",
    "3239032": "Sometimes people prefer to be solo for solo medal, or solo/small team for larger price share ^^ (I was not among the two people, but I did reject another very good offer because I was hungry for price haha)"
  },
  "source": "meta"
}