{
  "id": 587450,
  "title": "OpenFWI UINT8+1 quantized dataset",
  "url": "/competitions/waveform-inversion/discussion/587450",
  "author_name": "Bilzard",
  "post_date": "2025-07-01T05:38:31.712000",
  "votes": 2,
  "comment_count": 1,
  "views": 0,
  "content": "<p>One of the challenging issues in this competition is the size of the dataset.<br>\nAlthough it contains only 470K samples, it consumes approximately <strong>311GB</strong> when quantized to float16.</p>\n<p>Can we further reduce the dataset size?</p>\n<p>As you know, to quantize the huge weights of recent LLMs, some sophisticated quantization techniques are used. So I tested whether these techniques are also transferable to the <strong>dataset domain</strong>.</p>\n<p>I tried block quantization, which is adopted in Llama.cpp’s GGUF format [2].</p>\n<ul>\n<li>[2] <a href=\"https://github.com/ggml-org/llama.cpp/discussions/5063\" target=\"_blank\">https://github.com/ggml-org/llama.cpp/discussions/5063</a></li>\n</ul>\n<p>Briefly speaking, it basically uses UTF-8 quantization for value encoding. However, the 0–255 quantization levels lose a lot of information from the input data. To mitigate this issue, we can scale the ranges of the input dataset to fully utilize the limited data levels.<br>\nBlock quantization is one method to scale the data range block-wise.</p>\n<p>I used a 4×4 block, which results in (1/4 × 1/4) × 16 = 1.0 bit/pixel to encode the block-wise data range. Thus, the final bit rate is 8.0 + 1.0 = <strong>9.0 bit/pixel</strong>.</p>\n<p>The resulting dataset was compressed to <strong>173GB</strong>, and the quantization error was MAE=~4.84e-03.</p>\n<p><strong>Limitations:</strong></p>\n<p>I used this dataset for inversion of the forward model [1] and confirmed that it does not affect the final results.<br>\nHowever, I am not sure whether this quantization error will affect training the backward model (seismic → velocity).<br>\nTherefore, <strong>there is a risk of harmful results due to quantization error</strong>.</p>\n<ul>\n<li>[1] <a href=\"https://www.kaggle.com/competitions/waveform-inversion/discussion/587422\" target=\"_blank\">https://www.kaggle.com/competitions/waveform-inversion/discussion/587422</a></li>\n</ul>\n<p><strong>Availability:</strong></p>\n<p></p>\n<p>-&gt; now I uploaded the dataset and usage example notebook.</p>\n<ul>\n<li>Dataset: <a href=\"https://www.kaggle.com/datasets/tatamikenn/openfwi-uint8\" target=\"_blank\">https://www.kaggle.com/datasets/tatamikenn/openfwi-uint8</a></li>\n<li>Notebook: <a href=\"https://www.kaggle.com/code/tatamikenn/ywi-usage-example-of-openfwi-uint8-1-dataset?scriptVersionId=248278227\" target=\"_blank\">https://www.kaggle.com/code/tatamikenn/ywi-usage-example-of-openfwi-uint8-1-dataset?scriptVersionId=248278227</a></li>\n</ul>\n<p>Hopefully, someone will try using this dataset to test whether it is also effective for training the backward model.</p>\n<p><strong>Sample Image</strong></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fc4fcc53d5b693114e0a6e4675e165d53%2Fmu-law.jpeg?generation=1751348205912731&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 3237460,
      "postDate": "2025-07-01T05:38:31.713Z",
      "content": "<p>One of the challenging issues in this competition is the size of the dataset.<br>\nAlthough it contains only 470K samples, it consumes approximately <strong>311GB</strong> when quantized to float16.</p>\n<p>Can we further reduce the dataset size?</p>\n<p>As you know, to quantize the huge weights of recent LLMs, some sophisticated quantization techniques are used. So I tested whether these techniques are also transferable to the <strong>dataset domain</strong>.</p>\n<p>I tried block quantization, which is adopted in Llama.cpp’s GGUF format [2].</p>\n<ul>\n<li>[2] <a href=\"https://github.com/ggml-org/llama.cpp/discussions/5063\" target=\"_blank\">https://github.com/ggml-org/llama.cpp/discussions/5063</a></li>\n</ul>\n<p>Briefly speaking, it basically uses UTF-8 quantization for value encoding. However, the 0–255 quantization levels lose a lot of information from the input data. To mitigate this issue, we can scale the ranges of the input dataset to fully utilize the limited data levels.<br>\nBlock quantization is one method to scale the data range block-wise.</p>\n<p>I used a 4×4 block, which results in (1/4 × 1/4) × 16 = 1.0 bit/pixel to encode the block-wise data range. Thus, the final bit rate is 8.0 + 1.0 = <strong>9.0 bit/pixel</strong>.</p>\n<p>The resulting dataset was compressed to <strong>173GB</strong>, and the quantization error was MAE=~4.84e-03.</p>\n<p><strong>Limitations:</strong></p>\n<p>I used this dataset for inversion of the forward model [1] and confirmed that it does not affect the final results.<br>\nHowever, I am not sure whether this quantization error will affect training the backward model (seismic → velocity).<br>\nTherefore, <strong>there is a risk of harmful results due to quantization error</strong>.</p>\n<ul>\n<li>[1] <a href=\"https://www.kaggle.com/competitions/waveform-inversion/discussion/587422\" target=\"_blank\">https://www.kaggle.com/competitions/waveform-inversion/discussion/587422</a></li>\n</ul>\n<p><strong>Availability:</strong></p>\n<p></p>\n<p>-&gt; now I uploaded the dataset and usage example notebook.</p>\n<ul>\n<li>Dataset: <a href=\"https://www.kaggle.com/datasets/tatamikenn/openfwi-uint8\" target=\"_blank\">https://www.kaggle.com/datasets/tatamikenn/openfwi-uint8</a></li>\n<li>Notebook: <a href=\"https://www.kaggle.com/code/tatamikenn/ywi-usage-example-of-openfwi-uint8-1-dataset?scriptVersionId=248278227\" target=\"_blank\">https://www.kaggle.com/code/tatamikenn/ywi-usage-example-of-openfwi-uint8-1-dataset?scriptVersionId=248278227</a></li>\n</ul>\n<p>Hopefully, someone will try using this dataset to test whether it is also effective for training the backward model.</p>\n<p><strong>Sample Image</strong></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fc4fcc53d5b693114e0a6e4675e165d53%2Fmu-law.jpeg?generation=1751348205912731&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "One of the challenging issues in this competition is the size of the dataset.\nAlthough it contains only 470K samples, it consumes approximately **311GB** when quantized to float16.\n\nCan we further reduce the dataset size?\n\nAs you know, to quantize the huge weights of recent LLMs, some sophisticated quantization techniques are used. So I tested whether these techniques are also transferable to the **dataset domain**.\n\nI tried block quantization, which is adopted in Llama.cpp’s GGUF format [2].\n\n- [2] https://github.com/ggml-org/llama.cpp/discussions/5063\n\nBriefly speaking, it basically uses UTF-8 quantization for value encoding. However, the 0–255 quantization levels lose a lot of information from the input data. To mitigate this issue, we can scale the ranges of the input dataset to fully utilize the limited data levels.\nBlock quantization is one method to scale the data range block-wise.\n\nI used a 4×4 block, which results in (1/4 × 1/4) × 16 = 1.0 bit/pixel to encode the block-wise data range. Thus, the final bit rate is 8.0 + 1.0 = **9.0 bit/pixel**.\n\nThe resulting dataset was compressed to **173GB**, and the quantization error was MAE=\\~4.84e-03.\n\n**Limitations:**\n\nI used this dataset for inversion of the forward model \\[1] and confirmed that it does not affect the final results.\nHowever, I am not sure whether this quantization error will affect training the backward model (seismic → velocity).\nTherefore, **there is a risk of harmful results due to quantization error**.\n\n* \\[1] [https://www.kaggle.com/competitions/waveform-inversion/discussion/587422](https://www.kaggle.com/competitions/waveform-inversion/discussion/587422)\n\n**Availability:**\n\n~~I am currently uploading this dataset to Kaggle Datasets.~~\n\n-> now I uploaded the dataset and usage example notebook.\n\n- Dataset: https://www.kaggle.com/datasets/tatamikenn/openfwi-uint8\n- Notebook: https://www.kaggle.com/code/tatamikenn/ywi-usage-example-of-openfwi-uint8-1-dataset?scriptVersionId=248278227\n\nHopefully, someone will try using this dataset to test whether it is also effective for training the backward model.\n\n**Sample Image**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fc4fcc53d5b693114e0a6e4675e165d53%2Fmu-law.jpeg?generation=1751348205912731&alt=media)",
      "votes": 2
    },
    {
      "id": 3237735,
      "postDate": "2025-07-01T09:16:47.010Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3237735,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-07-01T09:16:47.010000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3237460": "One of the challenging issues in this competition is the size of the dataset.\nAlthough it contains only 470K samples, it consumes approximately **311GB** when quantized to float16.\n\nCan we further reduce the dataset size?\n\nAs you know, to quantize the huge weights of recent LLMs, some sophisticated quantization techniques are used. So I tested whether these techniques are also transferable to the **dataset domain**.\n\nI tried block quantization, which is adopted in Llama.cpp’s GGUF format [2].\n\n- [2] https://github.com/ggml-org/llama.cpp/discussions/5063\n\nBriefly speaking, it basically uses UTF-8 quantization for value encoding. However, the 0–255 quantization levels lose a lot of information from the input data. To mitigate this issue, we can scale the ranges of the input dataset to fully utilize the limited data levels.\nBlock quantization is one method to scale the data range block-wise.\n\nI used a 4×4 block, which results in (1/4 × 1/4) × 16 = 1.0 bit/pixel to encode the block-wise data range. Thus, the final bit rate is 8.0 + 1.0 = **9.0 bit/pixel**.\n\nThe resulting dataset was compressed to **173GB**, and the quantization error was MAE=\\~4.84e-03.\n\n**Limitations:**\n\nI used this dataset for inversion of the forward model \\[1] and confirmed that it does not affect the final results.\nHowever, I am not sure whether this quantization error will affect training the backward model (seismic → velocity).\nTherefore, **there is a risk of harmful results due to quantization error**.\n\n* \\[1] [https://www.kaggle.com/competitions/waveform-inversion/discussion/587422](https://www.kaggle.com/competitions/waveform-inversion/discussion/587422)\n\n**Availability:**\n\n~~I am currently uploading this dataset to Kaggle Datasets.~~\n\n-> now I uploaded the dataset and usage example notebook.\n\n- Dataset: https://www.kaggle.com/datasets/tatamikenn/openfwi-uint8\n- Notebook: https://www.kaggle.com/code/tatamikenn/ywi-usage-example-of-openfwi-uint8-1-dataset?scriptVersionId=248278227\n\nHopefully, someone will try using this dataset to test whether it is also effective for training the backward model.\n\n**Sample Image**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fc4fcc53d5b693114e0a6e4675e165d53%2Fmu-law.jpeg?generation=1751348205912731&alt=media)",
    "3237735": ""
  }
}