{
  "id": 572334,
  "title": "is all OpenFWI Dataset ~ 670G vs Kaggle Dataset ~ Train 14TB (dataset) / Pre-trained Models(dataset)",
  "url": "/competitions/waveform-inversion/discussion/572334",
  "author_name": "SeshuRaju 🧘‍♂️",
  "post_date": "2025-04-09T01:54:52.390000",
  "votes": 21,
  "comment_count": 7,
  "views": 0,
  "content": "<blockquote>\n  <p><strong>OpenFWI</strong> is a collection of <strong>large-scale, multi-structural benchmark datasets</strong> for <strong>machine learning-driven seismic FWI</strong>.  </p>\n  <ul>\n  <li>We release <strong>twelve datasets</strong> synthesized from <strong>different priors</strong>, including <strong>one 3D dataset</strong>.  </li>\n  <li>We also provide <strong>baseline experimental results</strong> with four deep learning methods:  <strong>InversionNet</strong>, <strong>VelocityGAN</strong>, <strong>UPFWI</strong>, and <strong>InversionNet3D</strong>.  </li>\n  <li><strong>OpenFWI</strong> is the <strong>first open-source platform</strong> to facilitate <strong>data-driven FWI research</strong>.  <br>\n  It will be <strong>actively developed</strong>, and the <strong>datasets are expected to evolve</strong>.</li>\n  </ul>\n</blockquote>\n<hr>\n<h1><a href=\"https://openfwi-lanl.github.io/docs/data.html\" target=\"_blank\">OpenFWI Dataset</a></h1>\n<blockquote>\n  <p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F0a5b195c9d21dd418b4ab60314acb173%2FScreenshot%202025-04-09%20at%206.50.53AM.png?generation=1744161752730004&amp;alt=media\" alt=\"\"></p>\n</blockquote>\n<table>\n<thead>\n<tr>\n<th>Dataset Name</th>\n<th>OpenFWI Size</th>\n<th>Kaggle Size</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>FlatVel-A</td>\n<td>43G</td>\n<td>1.4G</td>\n</tr>\n<tr>\n<td>FlatVel-B</td>\n<td>43G</td>\n<td>1.4G</td>\n</tr>\n<tr>\n<td>FlatFault-A</td>\n<td>77G</td>\n<td>1.4G</td>\n</tr>\n<tr>\n<td>FlatFault-B</td>\n<td>77G</td>\n<td>1.4G</td>\n</tr>\n<tr>\n<td><strong>Flat</strong></td>\n<td><strong>240G</strong></td>\n<td><strong>5.6G</strong></td>\n</tr>\n<tr>\n<td>CurveVel-A</td>\n<td>43G</td>\n<td>1.4G</td>\n</tr>\n<tr>\n<td>CurveVel-B</td>\n<td>43G</td>\n<td>1.4G</td>\n</tr>\n<tr>\n<td>CurveFault-A</td>\n<td>77G</td>\n<td>1.4G</td>\n</tr>\n<tr>\n<td>CurveFault-B</td>\n<td>77G</td>\n<td>1.4G</td>\n</tr>\n<tr>\n<td><strong>Curve</strong></td>\n<td><strong>240G</strong></td>\n<td><strong>5.6G</strong></td>\n</tr>\n<tr>\n<td>Style-A</td>\n<td>95G</td>\n<td>1.4G</td>\n</tr>\n<tr>\n<td>Style-B</td>\n<td>95G</td>\n<td>1.4G</td>\n</tr>\n<tr>\n<td><strong>Style</strong></td>\n<td><strong>190G</strong></td>\n<td><strong>2.4G</strong></td>\n</tr>\n<tr>\n<td>Kimberlina-CO2</td>\n<td>93G</td>\n<td>- (not application to this competition as host suggested )</td>\n</tr>\n<tr>\n<td>3D Kimberlina-V1</td>\n<td>1.4T</td>\n<td>- (not application to this competition as host suggested)</td>\n</tr>\n<tr>\n<td></td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td><strong>Total</strong></td>\n<td><strong>670G</strong></td>\n<td><strong>14G</strong></td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h1>Test Set</h1>\n<blockquote>\n  <ul>\n  <li>To clarify, the <strong>test set</strong> used in this competition follows a <strong>similar distribution</strong> as the original <strong>OpenFWI dataset</strong>. It was generated using the <strong>same simulation code and hyperparameters</strong> as the training data.</li>\n  <li>However, to encourage <strong>robustness</strong> and <strong>broad generalization</strong>, the test set is composed of a <strong>mixture of all 10 OpenFWI subsets</strong> (e.g., <strong>FlatVel_A</strong>, <strong>CurveFault_B</strong>, <strong>Style_A</strong>, etc.). While there is <strong>no deliberate domain shift</strong> introduced, the <strong>diversity across these subsets</strong> means the challenge inherently tests a model’s ability to <strong>generalize across different geological features and styles</strong>.</li>\n  </ul>\n  <p>by author <a href=\"https://www.kaggle.com/hanchenwang114\" target=\"_blank\">@hanchenwang114</a>  <a href=\"https://www.kaggle.com/competitions/waveform-inversion/discussion/572334#3175071\" target=\"_blank\">here</a></p>\n</blockquote>\n<hr>\n<h1>Tutorials</h1>\n<blockquote>\n  <h2><a href=\"https://colab.research.google.com/drive/17s5JmVs9ABl8MpmFlhWMSslj9_d5Atfx?usp=sharing#scrollTo=oB54haGtkrRt\" target=\"_blank\">Colab Notebook - Tutorial </a></h2>\n  <h2><a href=\"https://openfwi-lanl.github.io/tutorial/#/\" target=\"_blank\">OpenFWI Documentation</a></h2>\n</blockquote>\n<hr>\n<h1>Kaggle Datasets</h1>\n<blockquote>\n  <h2><a href=\"https://www.kaggle.com/datasets/seshurajup/waveform-inversion-train\" target=\"_blank\">Kaggle  Dataset - 14G Train Dataset</a></h2>\n  <h2><a href=\"https://www.kaggle.com/datasets/seshurajup/waveform-inversion-models\" target=\"_blank\">Kaggle Dataset - OpenFWI Pre-trained Models</a> - from <a href=\"https://smileunc.github.io/projects/openfwi/resources\" target=\"_blank\">Resources</a></h2>\n  <p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2Fa0fe31da0cb4203bf2d233b476e46545%2FScreenshot%202025-04-09%20at%208.22.46AM.png?generation=1744167249231075&amp;alt=media\" alt=\"\"><br>\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F85e565191e3639c746d4c7b423dc25d0%2FScreenshot%202025-04-09%20at%208.28.08AM.png?generation=1744167529971173&amp;alt=media\" alt=\"\"><br>\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F26023415244e3d2d6d2f240db4e9b6c2%2FScreenshot%202025-04-09%20at%208.28.23AM.png?generation=1744167538794863&amp;alt=media\" alt=\"\"></p>\n</blockquote>",
  "messages": [
    {
      "id": 3174325,
      "postDate": "2025-04-09T01:54:52.390Z",
      "content": "<blockquote>\n  <p><strong>OpenFWI</strong> is a collection of <strong>large-scale, multi-structural benchmark datasets</strong> for <strong>machine learning-driven seismic FWI</strong>.  </p>\n  <ul>\n  <li>We release <strong>twelve datasets</strong> synthesized from <strong>different priors</strong>, including <strong>one 3D dataset</strong>.  </li>\n  <li>We also provide <strong>baseline experimental results</strong> with four deep learning methods:  <strong>InversionNet</strong>, <strong>VelocityGAN</strong>, <strong>UPFWI</strong>, and <strong>InversionNet3D</strong>.  </li>\n  <li><strong>OpenFWI</strong> is the <strong>first open-source platform</strong> to facilitate <strong>data-driven FWI research</strong>.  <br>\n  It will be <strong>actively developed</strong>, and the <strong>datasets are expected to evolve</strong>.</li>\n  </ul>\n</blockquote>\n<hr>\n<h1><a href=\"https://openfwi-lanl.github.io/docs/data.html\" target=\"_blank\">OpenFWI Dataset</a></h1>\n<blockquote>\n  <p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F0a5b195c9d21dd418b4ab60314acb173%2FScreenshot%202025-04-09%20at%206.50.53AM.png?generation=1744161752730004&amp;alt=media\" alt=\"\"></p>\n</blockquote>\n<table>\n<thead>\n<tr>\n<th>Dataset Name</th>\n<th>OpenFWI Size</th>\n<th>Kaggle Size</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>FlatVel-A</td>\n<td>43G</td>\n<td>1.4G</td>\n</tr>\n<tr>\n<td>FlatVel-B</td>\n<td>43G</td>\n<td>1.4G</td>\n</tr>\n<tr>\n<td>FlatFault-A</td>\n<td>77G</td>\n<td>1.4G</td>\n</tr>\n<tr>\n<td>FlatFault-B</td>\n<td>77G</td>\n<td>1.4G</td>\n</tr>\n<tr>\n<td><strong>Flat</strong></td>\n<td><strong>240G</strong></td>\n<td><strong>5.6G</strong></td>\n</tr>\n<tr>\n<td>CurveVel-A</td>\n<td>43G</td>\n<td>1.4G</td>\n</tr>\n<tr>\n<td>CurveVel-B</td>\n<td>43G</td>\n<td>1.4G</td>\n</tr>\n<tr>\n<td>CurveFault-A</td>\n<td>77G</td>\n<td>1.4G</td>\n</tr>\n<tr>\n<td>CurveFault-B</td>\n<td>77G</td>\n<td>1.4G</td>\n</tr>\n<tr>\n<td><strong>Curve</strong></td>\n<td><strong>240G</strong></td>\n<td><strong>5.6G</strong></td>\n</tr>\n<tr>\n<td>Style-A</td>\n<td>95G</td>\n<td>1.4G</td>\n</tr>\n<tr>\n<td>Style-B</td>\n<td>95G</td>\n<td>1.4G</td>\n</tr>\n<tr>\n<td><strong>Style</strong></td>\n<td><strong>190G</strong></td>\n<td><strong>2.4G</strong></td>\n</tr>\n<tr>\n<td>Kimberlina-CO2</td>\n<td>93G</td>\n<td>- (not application to this competition as host suggested )</td>\n</tr>\n<tr>\n<td>3D Kimberlina-V1</td>\n<td>1.4T</td>\n<td>- (not application to this competition as host suggested)</td>\n</tr>\n<tr>\n<td></td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td><strong>Total</strong></td>\n<td><strong>670G</strong></td>\n<td><strong>14G</strong></td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h1>Test Set</h1>\n<blockquote>\n  <ul>\n  <li>To clarify, the <strong>test set</strong> used in this competition follows a <strong>similar distribution</strong> as the original <strong>OpenFWI dataset</strong>. It was generated using the <strong>same simulation code and hyperparameters</strong> as the training data.</li>\n  <li>However, to encourage <strong>robustness</strong> and <strong>broad generalization</strong>, the test set is composed of a <strong>mixture of all 10 OpenFWI subsets</strong> (e.g., <strong>FlatVel_A</strong>, <strong>CurveFault_B</strong>, <strong>Style_A</strong>, etc.). While there is <strong>no deliberate domain shift</strong> introduced, the <strong>diversity across these subsets</strong> means the challenge inherently tests a model’s ability to <strong>generalize across different geological features and styles</strong>.</li>\n  </ul>\n  <p>by author <a href=\"https://www.kaggle.com/hanchenwang114\" target=\"_blank\">@hanchenwang114</a>  <a href=\"https://www.kaggle.com/competitions/waveform-inversion/discussion/572334#3175071\" target=\"_blank\">here</a></p>\n</blockquote>\n<hr>\n<h1>Tutorials</h1>\n<blockquote>\n  <h2><a href=\"https://colab.research.google.com/drive/17s5JmVs9ABl8MpmFlhWMSslj9_d5Atfx?usp=sharing#scrollTo=oB54haGtkrRt\" target=\"_blank\">Colab Notebook - Tutorial </a></h2>\n  <h2><a href=\"https://openfwi-lanl.github.io/tutorial/#/\" target=\"_blank\">OpenFWI Documentation</a></h2>\n</blockquote>\n<hr>\n<h1>Kaggle Datasets</h1>\n<blockquote>\n  <h2><a href=\"https://www.kaggle.com/datasets/seshurajup/waveform-inversion-train\" target=\"_blank\">Kaggle  Dataset - 14G Train Dataset</a></h2>\n  <h2><a href=\"https://www.kaggle.com/datasets/seshurajup/waveform-inversion-models\" target=\"_blank\">Kaggle Dataset - OpenFWI Pre-trained Models</a> - from <a href=\"https://smileunc.github.io/projects/openfwi/resources\" target=\"_blank\">Resources</a></h2>\n  <p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2Fa0fe31da0cb4203bf2d233b476e46545%2FScreenshot%202025-04-09%20at%208.22.46AM.png?generation=1744167249231075&amp;alt=media\" alt=\"\"><br>\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F85e565191e3639c746d4c7b423dc25d0%2FScreenshot%202025-04-09%20at%208.28.08AM.png?generation=1744167529971173&amp;alt=media\" alt=\"\"><br>\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F26023415244e3d2d6d2f240db4e9b6c2%2FScreenshot%202025-04-09%20at%208.28.23AM.png?generation=1744167538794863&amp;alt=media\" alt=\"\"></p>\n</blockquote>",
      "rawMarkdown": "> **OpenFWI** is a collection of **large-scale, multi-structural benchmark datasets** for **machine learning-driven seismic FWI**.  \n- We release **twelve datasets** synthesized from **different priors**, including **one 3D dataset**.  \n- We also provide **baseline experimental results** with four deep learning methods:  **InversionNet**, **VelocityGAN**, **UPFWI**, and **InversionNet3D**.  \n- **OpenFWI** is the **first open-source platform** to facilitate **data-driven FWI research**.  \nIt will be **actively developed**, and the **datasets are expected to evolve**.\n\n---\n# [OpenFWI Dataset](https://openfwi-lanl.github.io/docs/data.html)\n> ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F0a5b195c9d21dd418b4ab60314acb173%2FScreenshot%202025-04-09%20at%206.50.53AM.png?generation=1744161752730004&alt=media)\n\n\n| Dataset Name        | OpenFWI Size  | Kaggle Size |\n|---------------------|-------|-------|\n| FlatVel-A           | 43G   | 1.4G |\n| FlatVel-B           | 43G   | 1.4G |\n| FlatFault-A         | 77G   | 1.4G |\n| FlatFault-B         | 77G   | 1.4G |\n| **Flat** | **240G** | **5.6G** |\n| CurveVel-A          | 43G   | 1.4G |\n| CurveVel-B          | 43G   | 1.4G |\n| CurveFault-A        | 77G   | 1.4G |\n| CurveFault-B        | 77G   | 1.4G \n| **Curve** | **240G** | **5.6G** |\n| Style-A             | 95G   | 1.4G |\n| Style-B             | 95G   | 1.4G |\n| **Style** | **190G** | **2.4G** |\n| Kimberlina-CO2      | 93G   | - (not application to this competition as host suggested ) |\n| 3D Kimberlina-V1    | 1.4T  | - (not application to this competition as host suggested) |\n| ~~**Total**~~ | ~~**2.15T**~~ | ~~**14G**~~ |\n| **Total** | **670G** | **14G** |\n\n---\n# Test Set\n> - To clarify, the **test set** used in this competition follows a **similar distribution** as the original **OpenFWI dataset**. It was generated using the **same simulation code and hyperparameters** as the training data.\n> - However, to encourage **robustness** and **broad generalization**, the test set is composed of a **mixture of all 10 OpenFWI subsets** (e.g., **FlatVel_A**, **CurveFault_B**, **Style_A**, etc.). While there is **no deliberate domain shift** introduced, the **diversity across these subsets** means the challenge inherently tests a model’s ability to **generalize across different geological features and styles**.\n\n> by author @hanchenwang114  [here](https://www.kaggle.com/competitions/waveform-inversion/discussion/572334#3175071)\n\n---\n# Tutorials\n>## [Colab Notebook - Tutorial ](https://colab.research.google.com/drive/17s5JmVs9ABl8MpmFlhWMSslj9_d5Atfx?usp=sharing#scrollTo=oB54haGtkrRt)\n> ## [OpenFWI Documentation](https://openfwi-lanl.github.io/tutorial/#/)\n\n---\n# Kaggle Datasets\n>## [Kaggle  Dataset - 14G Train Dataset](https://www.kaggle.com/datasets/seshurajup/waveform-inversion-train)\n>## [Kaggle Dataset - OpenFWI Pre-trained Models](https://www.kaggle.com/datasets/seshurajup/waveform-inversion-models) - from [Resources](https://smileunc.github.io/projects/openfwi/resources)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2Fa0fe31da0cb4203bf2d233b476e46545%2FScreenshot%202025-04-09%20at%208.22.46AM.png?generation=1744167249231075&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F85e565191e3639c746d4c7b423dc25d0%2FScreenshot%202025-04-09%20at%208.28.08AM.png?generation=1744167529971173&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F26023415244e3d2d6d2f240db4e9b6c2%2FScreenshot%202025-04-09%20at%208.28.23AM.png?generation=1744167538794863&alt=media)",
      "votes": 21
    },
    {
      "id": 3174335,
      "postDate": "2025-04-09T02:07:16.780Z",
      "content": "<p>Hello there,</p>\n<p>We present a few samples from each OpenFWI subset on Kaggle as the training examples. You are welcome to download our original OpenFWI dataset as much as you want to be your training set from the official OpenFWI website. </p>\n<p>For the purposes of this competition, we recommend ignoring the Kimberlina family of OpenFWI dataset, as it was designed for a different geophysical task—time-lapse CO₂ monitoring—which falls outside the scope of this competition. </p>\n<p>Best,<br>\nThe Waveform Inversion Team </p>",
      "rawMarkdown": "Hello there,\n\nWe present a few samples from each OpenFWI subset on Kaggle as the training examples. You are welcome to download our original OpenFWI dataset as much as you want to be your training set from the official OpenFWI website. \n\nFor the purposes of this competition, we recommend ignoring the Kimberlina family of OpenFWI dataset, as it was designed for a different geophysical task—time-lapse CO₂ monitoring—which falls outside the scope of this competition. \n\nBest,\nThe Waveform Inversion Team ",
      "votes": 5,
      "replies": [
        {
          "id": 3174355,
          "postDate": "2025-04-09T02:33:56.503Z",
          "content": "<p>Thanks for the quick reply <a href=\"https://www.kaggle.com/hanchenwang114\" target=\"_blank\">@hanchenwang114</a> updated the topic, </p>\n<blockquote>\n  <p>Could we expect the distribution of the test set having mix of 10 groups with same distribution as Kaggle Training dataset i.e equal or same distribution as OpenFWI dataset?</p>\n</blockquote>\n<hr>\n<blockquote>\n  <p> miss understood it is notebook competition</p>\n</blockquote>",
          "rawMarkdown": "Thanks for the quick reply @hanchenwang114 updated the topic, \n\n> Could we expect the distribution of the test set having mix of 10 groups with same distribution as Kaggle Training dataset i.e equal or same distribution as OpenFWI dataset?\n\n---\n\n> ~~How many private test files or size to estimate inference submission time!~~ miss understood it is notebook competition",
          "votes": 2,
          "replies": [
            {
              "id": 3174403,
              "postDate": "2025-04-09T04:28:20.953Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 3174587,
              "postDate": "2025-04-09T08:30:58.087Z",
              "content": "<blockquote>\n  <p>However, to encourage robustness and broad generalization, the test set is composed of a mixture of all 10 OpenFWI subsets (e.g., FlatVel_A, CurveFault_B, Style_A, etc.). While there is no deliberate domain shift introduced, the diversity across these subsets means the challenge inherently tests a model’s ability to generalize across different geological features and styles. - by author <a href=\"https://www.kaggle.com/hanchenwang114\" target=\"_blank\">@hanchenwang114</a> here</p>\n</blockquote>\n<p>This comment was deleted by the host. I see there was another comment here that was deleted, but I didn't see what it was.<br>\nStrange things are happening, for sure.</p>",
              "rawMarkdown": ">However, to encourage robustness and broad generalization, the test set is composed of a mixture of all 10 OpenFWI subsets (e.g., FlatVel_A, CurveFault_B, Style_A, etc.). While there is no deliberate domain shift introduced, the diversity across these subsets means the challenge inherently tests a model’s ability to generalize across different geological features and styles. - by author @hanchenwang114 here\n\nThis comment was deleted by the host. I see there was another comment here that was deleted, but I didn't see what it was.\nStrange things are happening, for sure.",
              "votes": 1
            },
            {
              "id": 3174607,
              "postDate": "2025-04-09T08:56:13.577Z",
              "content": "<blockquote>\n  <p><a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a> i deleted the my comment which i posted about it i.e it having <strong>65,818*70</strong> = 4607260 ( after i understood  it is not notebook competition ). i'm position my deleted comment again as it dump question for csv submission.</p>\n</blockquote>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F418fb5f6f81a5280e8d5f0f74bf4625c%2FScreenshot%202025-04-09%20at%209.56.10AM.png?generation=1744188784478736&amp;alt=media\" alt=\"\"></p>",
              "rawMarkdown": "> @shlomoron i deleted the my comment which i posted about it i.e it having **65,818*70** = 4607260 ( after i understood  it is not notebook competition ). i'm position my deleted comment again as it dump question for csv submission.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F418fb5f6f81a5280e8d5f0f74bf4625c%2FScreenshot%202025-04-09%20at%209.56.10AM.png?generation=1744188784478736&alt=media)",
              "votes": 1
            },
            {
              "id": 3175069,
              "postDate": "2025-04-09T17:30:59.793Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 3175071,
              "postDate": "2025-04-09T17:31:40.220Z",
              "content": "<p>Hi there,</p>\n<p>Thanks for raising this important point!</p>\n<p>To clarify, the test set used in this competition <strong>follows the similar distribution as the original OpenFWI dataset</strong>. It was generated using the same simulation code and hyperparameters as the training data.</p>\n<p>However, to encourage robustness and broad generalization, the test set is composed of <strong>a mixture of all 10 OpenFWI subsets (e.g., FlatVel_A, CurveFault_B, Style_A, etc.)</strong>. While there is no deliberate domain shift introduced, the diversity across these subsets means the challenge inherently tests a model’s ability to generalize across different geological features and styles.</p>\n<p>We appreciate your thoughtful question and hope this provides clarity.</p>\n<p>Best regards,<br>\nThe Waveform Inversion Team</p>",
              "rawMarkdown": "Hi there,\n\nThanks for raising this important point!\n\nTo clarify, the test set used in this competition **follows the similar distribution as the original OpenFWI dataset**. It was generated using the same simulation code and hyperparameters as the training data.\n\nHowever, to encourage robustness and broad generalization, the test set is composed of **a mixture of all 10 OpenFWI subsets (e.g., FlatVel_A, CurveFault_B, Style_A, etc.)**. While there is no deliberate domain shift introduced, the diversity across these subsets means the challenge inherently tests a model’s ability to generalize across different geological features and styles.\n\nWe appreciate your thoughtful question and hope this provides clarity.\n\nBest regards,\nThe Waveform Inversion Team",
              "votes": 10
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3174335,
      "author_name": "Hanchen Wang",
      "author_url": "",
      "post_date": "2025-04-09T02:07:16.780000",
      "content": "<p>Hello there,</p>\n<p>We present a few samples from each OpenFWI subset on Kaggle as the training examples. You are welcome to download our original OpenFWI dataset as much as you want to be your training set from the official OpenFWI website. </p>\n<p>For the purposes of this competition, we recommend ignoring the Kimberlina family of OpenFWI dataset, as it was designed for a different geophysical task—time-lapse CO₂ monitoring—which falls outside the scope of this competition. </p>\n<p>Best,<br>\nThe Waveform Inversion Team </p>",
      "votes": 5,
      "replies": [
        {
          "id": 3174355,
          "author_name": "SeshuRaju 🧘‍♂️",
          "author_url": "",
          "post_date": "2025-04-09T02:33:56.503000",
          "content": "<p>Thanks for the quick reply <a href=\"https://www.kaggle.com/hanchenwang114\" target=\"_blank\">@hanchenwang114</a> updated the topic, </p>\n<blockquote>\n  <p>Could we expect the distribution of the test set having mix of 10 groups with same distribution as Kaggle Training dataset i.e equal or same distribution as OpenFWI dataset?</p>\n</blockquote>\n<hr>\n<blockquote>\n  <p> miss understood it is notebook competition</p>\n</blockquote>",
          "votes": 2,
          "replies": [
            {
              "id": 3174403,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-04-09T04:28:20.953000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3174587,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2025-04-09T08:30:58.087000",
              "content": "<blockquote>\n  <p>However, to encourage robustness and broad generalization, the test set is composed of a mixture of all 10 OpenFWI subsets (e.g., FlatVel_A, CurveFault_B, Style_A, etc.). While there is no deliberate domain shift introduced, the diversity across these subsets means the challenge inherently tests a model’s ability to generalize across different geological features and styles. - by author <a href=\"https://www.kaggle.com/hanchenwang114\" target=\"_blank\">@hanchenwang114</a> here</p>\n</blockquote>\n<p>This comment was deleted by the host. I see there was another comment here that was deleted, but I didn't see what it was.<br>\nStrange things are happening, for sure.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3174607,
              "author_name": "SeshuRaju 🧘‍♂️",
              "author_url": "",
              "post_date": "2025-04-09T08:56:13.577000",
              "content": "<blockquote>\n  <p><a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a> i deleted the my comment which i posted about it i.e it having <strong>65,818*70</strong> = 4607260 ( after i understood  it is not notebook competition ). i'm position my deleted comment again as it dump question for csv submission.</p>\n</blockquote>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F418fb5f6f81a5280e8d5f0f74bf4625c%2FScreenshot%202025-04-09%20at%209.56.10AM.png?generation=1744188784478736&amp;alt=media\" alt=\"\"></p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3175069,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-04-09T17:30:59.793000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3175071,
              "author_name": "Hanchen Wang",
              "author_url": "",
              "post_date": "2025-04-09T17:31:40.220000",
              "content": "<p>Hi there,</p>\n<p>Thanks for raising this important point!</p>\n<p>To clarify, the test set used in this competition <strong>follows the similar distribution as the original OpenFWI dataset</strong>. It was generated using the same simulation code and hyperparameters as the training data.</p>\n<p>However, to encourage robustness and broad generalization, the test set is composed of <strong>a mixture of all 10 OpenFWI subsets (e.g., FlatVel_A, CurveFault_B, Style_A, etc.)</strong>. While there is no deliberate domain shift introduced, the diversity across these subsets means the challenge inherently tests a model’s ability to generalize across different geological features and styles.</p>\n<p>We appreciate your thoughtful question and hope this provides clarity.</p>\n<p>Best regards,<br>\nThe Waveform Inversion Team</p>",
              "votes": 10,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3174325": "> **OpenFWI** is a collection of **large-scale, multi-structural benchmark datasets** for **machine learning-driven seismic FWI**.  \n- We release **twelve datasets** synthesized from **different priors**, including **one 3D dataset**.  \n- We also provide **baseline experimental results** with four deep learning methods:  **InversionNet**, **VelocityGAN**, **UPFWI**, and **InversionNet3D**.  \n- **OpenFWI** is the **first open-source platform** to facilitate **data-driven FWI research**.  \nIt will be **actively developed**, and the **datasets are expected to evolve**.\n\n---\n# [OpenFWI Dataset](https://openfwi-lanl.github.io/docs/data.html)\n> ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F0a5b195c9d21dd418b4ab60314acb173%2FScreenshot%202025-04-09%20at%206.50.53AM.png?generation=1744161752730004&alt=media)\n\n\n| Dataset Name        | OpenFWI Size  | Kaggle Size |\n|---------------------|-------|-------|\n| FlatVel-A           | 43G   | 1.4G |\n| FlatVel-B           | 43G   | 1.4G |\n| FlatFault-A         | 77G   | 1.4G |\n| FlatFault-B         | 77G   | 1.4G |\n| **Flat** | **240G** | **5.6G** |\n| CurveVel-A          | 43G   | 1.4G |\n| CurveVel-B          | 43G   | 1.4G |\n| CurveFault-A        | 77G   | 1.4G |\n| CurveFault-B        | 77G   | 1.4G \n| **Curve** | **240G** | **5.6G** |\n| Style-A             | 95G   | 1.4G |\n| Style-B             | 95G   | 1.4G |\n| **Style** | **190G** | **2.4G** |\n| Kimberlina-CO2      | 93G   | - (not application to this competition as host suggested ) |\n| 3D Kimberlina-V1    | 1.4T  | - (not application to this competition as host suggested) |\n| ~~**Total**~~ | ~~**2.15T**~~ | ~~**14G**~~ |\n| **Total** | **670G** | **14G** |\n\n---\n# Test Set\n> - To clarify, the **test set** used in this competition follows a **similar distribution** as the original **OpenFWI dataset**. It was generated using the **same simulation code and hyperparameters** as the training data.\n> - However, to encourage **robustness** and **broad generalization**, the test set is composed of a **mixture of all 10 OpenFWI subsets** (e.g., **FlatVel_A**, **CurveFault_B**, **Style_A**, etc.). While there is **no deliberate domain shift** introduced, the **diversity across these subsets** means the challenge inherently tests a model’s ability to **generalize across different geological features and styles**.\n\n> by author @hanchenwang114  [here](https://www.kaggle.com/competitions/waveform-inversion/discussion/572334#3175071)\n\n---\n# Tutorials\n>## [Colab Notebook - Tutorial ](https://colab.research.google.com/drive/17s5JmVs9ABl8MpmFlhWMSslj9_d5Atfx?usp=sharing#scrollTo=oB54haGtkrRt)\n> ## [OpenFWI Documentation](https://openfwi-lanl.github.io/tutorial/#/)\n\n---\n# Kaggle Datasets\n>## [Kaggle  Dataset - 14G Train Dataset](https://www.kaggle.com/datasets/seshurajup/waveform-inversion-train)\n>## [Kaggle Dataset - OpenFWI Pre-trained Models](https://www.kaggle.com/datasets/seshurajup/waveform-inversion-models) - from [Resources](https://smileunc.github.io/projects/openfwi/resources)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2Fa0fe31da0cb4203bf2d233b476e46545%2FScreenshot%202025-04-09%20at%208.22.46AM.png?generation=1744167249231075&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F85e565191e3639c746d4c7b423dc25d0%2FScreenshot%202025-04-09%20at%208.28.08AM.png?generation=1744167529971173&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F26023415244e3d2d6d2f240db4e9b6c2%2FScreenshot%202025-04-09%20at%208.28.23AM.png?generation=1744167538794863&alt=media)",
    "3174335": "Hello there,\n\nWe present a few samples from each OpenFWI subset on Kaggle as the training examples. You are welcome to download our original OpenFWI dataset as much as you want to be your training set from the official OpenFWI website. \n\nFor the purposes of this competition, we recommend ignoring the Kimberlina family of OpenFWI dataset, as it was designed for a different geophysical task—time-lapse CO₂ monitoring—which falls outside the scope of this competition. \n\nBest,\nThe Waveform Inversion Team "
  }
}