{
  "id": 491472,
  "title": "The train set is bloated. After shrinking it can be fully loaded in Kaggle.",
  "url": "/competitions/leash-BELKA/discussion/491472",
  "author_name": "greySnow",
  "post_date": "2024-04-06T00:52:13.137000",
  "votes": 90,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Shrinking strategy:</p>\n<ol>\n<li>No ID column.</li>\n<li>binds columns saved in bytes.</li>\n<li>buildingblock1_smiles/buildingblock2_smiles/buildingblock3_smiles columns saved as int16, with encoded index of the building blobks. I saved the building blocks and their indices in separate dictionaries.</li>\n<li>I transformed the protein/label columns into three columns of labels per protein, shrinking the dataset length by three. (The other columns have identical values for each three consecutive rows).</li>\n</ol>\n<p>Original size: over 50GB.<br>\nShrunken size: ~10GB.</p>\n<p><a href=\"https://www.kaggle.com/code/shlomoron/belka-shrinking-the-dataset\" target=\"_blank\">Shrinking notebook</a>.<br>\n<a href=\"https://www.kaggle.com/datasets/shlomoron/belka-shrunken-train-set\" target=\"_blank\">Shrunken dataset</a>.<br>\n<a href=\"https://www.kaggle.com/code/shlomoron/belka-shrunken-train-set-loading\" target=\"_blank\">Shrunken dataset loading notebook</a>.</p>\n<p>Good luck and happy kaggling.</p>",
  "messages": [
    {
      "id": 2737828,
      "postDate": "2024-04-06T00:52:13.137Z",
      "content": "<p>Shrinking strategy:</p>\n<ol>\n<li>No ID column.</li>\n<li>binds columns saved in bytes.</li>\n<li>buildingblock1_smiles/buildingblock2_smiles/buildingblock3_smiles columns saved as int16, with encoded index of the building blobks. I saved the building blocks and their indices in separate dictionaries.</li>\n<li>I transformed the protein/label columns into three columns of labels per protein, shrinking the dataset length by three. (The other columns have identical values for each three consecutive rows).</li>\n</ol>\n<p>Original size: over 50GB.<br>\nShrunken size: ~10GB.</p>\n<p><a href=\"https://www.kaggle.com/code/shlomoron/belka-shrinking-the-dataset\" target=\"_blank\">Shrinking notebook</a>.<br>\n<a href=\"https://www.kaggle.com/datasets/shlomoron/belka-shrunken-train-set\" target=\"_blank\">Shrunken dataset</a>.<br>\n<a href=\"https://www.kaggle.com/code/shlomoron/belka-shrunken-train-set-loading\" target=\"_blank\">Shrunken dataset loading notebook</a>.</p>\n<p>Good luck and happy kaggling.</p>",
      "rawMarkdown": "Shrinking strategy:\n1. No ID column.\n2. binds columns saved in bytes.\n3. buildingblock1_smiles/buildingblock2_smiles/buildingblock3_smiles columns saved as int16, with encoded index of the building blobks. I saved the building blocks and their indices in separate dictionaries.\n4. I transformed the protein/label columns into three columns of labels per protein, shrinking the dataset length by three. (The other columns have identical values for each three consecutive rows).\n\nOriginal size: over 50GB.\nShrunken size: ~10GB.\n\n[Shrinking notebook](https://www.kaggle.com/code/shlomoron/belka-shrinking-the-dataset).\n[Shrunken dataset](https://www.kaggle.com/datasets/shlomoron/belka-shrunken-train-set).\n[Shrunken dataset loading notebook](https://www.kaggle.com/code/shlomoron/belka-shrunken-train-set-loading).\n\nGood luck and happy kaggling.",
      "votes": 87
    },
    {
      "id": 2740127,
      "postDate": "2024-04-07T15:41:53.487Z",
      "content": "<p>Thanks for the work! I upvoted your Dataset</p>",
      "rawMarkdown": "Thanks for the work! I upvoted your Dataset",
      "votes": 1
    },
    {
      "id": 2740783,
      "postDate": "2024-04-07T23:33:06.800Z",
      "content": "<p>Thanks! Couldn't even load the train parquet with polars on a 128Gb machine, with this dataset I imagine competition will attract way more attention :]</p>",
      "rawMarkdown": "Thanks! Couldn't even load the train parquet with polars on a 128Gb machine, with this dataset I imagine competition will attract way more attention :]",
      "votes": 2,
      "replies": [
        {
          "id": 2740978,
          "postDate": "2024-04-08T04:25:11.517Z",
          "content": "<p>If linux - create a large swap file - you can than load the train and do this or your version of shrinkage.   I use 200GB swap.</p>",
          "rawMarkdown": "If linux - create a large swap file - you can than load the train and do this or your version of shrinkage.   I use 200GB swap.\n",
          "votes": 1,
          "replies": [
            {
              "id": 2743981,
              "postDate": "2024-04-09T17:30:37.223Z",
              "content": "<p>You can create a lazyframe and save it to disk using streaming, so you can process datasets that are bigger than the available ram.</p>\n<p><a href=\"https://docs.pola.rs/py-polars/html/reference/api/polars.LazyFrame.sink_parquet.html\" target=\"_blank\">polars sink_parquet</a></p>",
              "rawMarkdown": "You can create a lazyframe and save it to disk using streaming, so you can process datasets that are bigger than the available ram.\n\n[polars sink_parquet](https://docs.pola.rs/py-polars/html/reference/api/polars.LazyFrame.sink_parquet.html)"
            }
          ]
        }
      ]
    },
    {
      "id": 2741140,
      "postDate": "2024-04-08T06:57:46.940Z",
      "content": "<p>WellDone sir</p>",
      "rawMarkdown": "WellDone sir"
    },
    {
      "id": 2741134,
      "postDate": "2024-04-08T06:56:40.410Z",
      "content": "<p>Amazing that's great!</p>",
      "rawMarkdown": "Amazing that's great!"
    },
    {
      "id": 2740176,
      "postDate": "2024-04-07T16:43:43.210Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2740127,
      "author_name": "Abish Pius",
      "author_url": "",
      "post_date": "2024-04-07T15:41:53.487000",
      "content": "<p>Thanks for the work! I upvoted your Dataset</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2740783,
      "author_name": "slime",
      "author_url": "",
      "post_date": "2024-04-07T23:33:06.800000",
      "content": "<p>Thanks! Couldn't even load the train parquet with polars on a 128Gb machine, with this dataset I imagine competition will attract way more attention :]</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2740978,
          "author_name": "PC Jimmmy",
          "author_url": "",
          "post_date": "2024-04-08T04:25:11.517000",
          "content": "<p>If linux - create a large swap file - you can than load the train and do this or your version of shrinkage.   I use 200GB swap.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2743981,
              "author_name": "Antonio Félix",
              "author_url": "",
              "post_date": "2024-04-09T17:30:37.223000",
              "content": "<p>You can create a lazyframe and save it to disk using streaming, so you can process datasets that are bigger than the available ram.</p>\n<p><a href=\"https://docs.pola.rs/py-polars/html/reference/api/polars.LazyFrame.sink_parquet.html\" target=\"_blank\">polars sink_parquet</a></p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2741140,
      "author_name": "Hina Ismail",
      "author_url": "",
      "post_date": "2024-04-08T06:57:46.940000",
      "content": "<p>WellDone sir</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2741134,
      "author_name": "Hina Ismail",
      "author_url": "",
      "post_date": "2024-04-08T06:56:40.410000",
      "content": "<p>Amazing that's great!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2740176,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-07T16:43:43.210000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2737828": "Shrinking strategy:\n1. No ID column.\n2. binds columns saved in bytes.\n3. buildingblock1_smiles/buildingblock2_smiles/buildingblock3_smiles columns saved as int16, with encoded index of the building blobks. I saved the building blocks and their indices in separate dictionaries.\n4. I transformed the protein/label columns into three columns of labels per protein, shrinking the dataset length by three. (The other columns have identical values for each three consecutive rows).\n\nOriginal size: over 50GB.\nShrunken size: ~10GB.\n\n[Shrinking notebook](https://www.kaggle.com/code/shlomoron/belka-shrinking-the-dataset).\n[Shrunken dataset](https://www.kaggle.com/datasets/shlomoron/belka-shrunken-train-set).\n[Shrunken dataset loading notebook](https://www.kaggle.com/code/shlomoron/belka-shrunken-train-set-loading).\n\nGood luck and happy kaggling.",
    "2740127": "Thanks for the work! I upvoted your Dataset",
    "2740783": "Thanks! Couldn't even load the train parquet with polars on a 128Gb machine, with this dataset I imagine competition will attract way more attention :]",
    "2741140": "WellDone sir",
    "2741134": "Amazing that's great!",
    "2740176": ""
  }
}