{
  "id": 491402,
  "title": "Huge train dataset file.",
  "url": "/competitions/leash-BELKA/discussion/491402",
  "author_name": "",
  "post_date": "2024-04-05T18:35:31.412981200Z",
  "votes": 10,
  "comment_count": 7,
  "views": 0,
  "content": "<p>The training csv dataset is huge : 59GB unzipped. </p>\n<p>Definitely, we do not have enough RAM on Kaggle notebooks in order to manage this dataset. </p>\n<p>This is how  I solved the problem:</p>\n<ol>\n<li>I downloaded the dataset on my computer (32Gb RAM)  and cut the file into smaller csv files. </li>\n<li>I made a model per small file, and then use the models for the final submission. </li>\n<li>On the other hand, if we want to encode the data, you need all the unique values in the columns (we have only categorical columns). We can see in the dataset in Kaggle the number of unique values. I iterated on the small files in order to get all the unique values in each column, excepting the one with more than 1000000 unique values, which I deleted.</li>\n</ol>\n<p>I am still working on this, it takes time …</p>",
  "messages": [
    {
      "id": "2737389",
      "postDate": "04/05/2024 18:35:31",
      "content": "<p>The training csv dataset is huge : 59GB unzipped. </p>\n<p>Definitely, we do not have enough RAM on Kaggle notebooks in order to manage this dataset. </p>\n<p>This is how  I solved the problem:</p>\n<ol>\n<li>I downloaded the dataset on my computer (32Gb RAM)  and cut the file into smaller csv files. </li>\n<li>I made a model per small file, and then use the models for the final submission. </li>\n<li>On the other hand, if we want to encode the data, you need all the unique values in the columns (we have only categorical columns). We can see in the dataset in Kaggle the number of unique values. I iterated on the small files in order to get all the unique values in each column, excepting the one with more than 1000000 unique values, which I deleted.</li>\n</ol>\n<p>I am still working on this, it takes time …</p>",
      "rawMarkdown": "The training csv dataset is huge : 59GB unzipped. \n\nDefinitely, we do not have enough RAM on Kaggle notebooks in order to manage this dataset. \n\nThis is how  I solved the problem:\n\n1. I downloaded the dataset on my computer (32Gb RAM)  and cut the file into smaller csv files. \n2. I made a model per small file, and then use the models for the final submission. \n3. On the other hand, if we want to encode the data, you need all the unique values in the columns (we have only categorical columns). We can see in the dataset in Kaggle the number of unique values. I iterated on the small files in order to get all the unique values in each column, excepting the one with more than 1000000 unique values, which I deleted.\n\nI am still working on this, it takes time ...",
      "votes": null
    },
    {
      "id": "2737502",
      "postDate": "04/05/2024 19:27:38",
      "content": "<p><a href=\"https://www.kaggle.com/catadanna\" target=\"_blank\">@catadanna</a> you may consider using pyspark for this if you want <br>\nPandas is not suitable in my opinion for such a dataset </p>\n<p>Also parquet/ feather format is a better one that csv</p>",
      "rawMarkdown": "catadanna you may consider using pyspark for this if you want \nPandas is not suitable in my opinion for such a dataset \n\nAlso parquet/ feather format is a better one that csv",
      "votes": null
    },
    {
      "id": "2737533",
      "postDate": "04/05/2024 19:47:08",
      "content": "<p>Good idea, pyspark ! </p>",
      "rawMarkdown": "Good idea, pyspark !",
      "votes": null
    },
    {
      "id": "2737710",
      "postDate": "04/05/2024 22:47:33",
      "content": "<p>The actual size of the train set is ~10GB and is sorely bloated; see my <a href=\"https://www.kaggle.com/code/shlomoron/belka-shrinking-the-dataset\" target=\"_blank\">notebook here</a>. More details to follow soon.</p>",
      "rawMarkdown": "The actual size of the train set is ~10GB and is sorely bloated; see my [notebook here](https://www.kaggle.com/code/shlomoron/belka-shrinking-the-dataset). More details to follow soon.",
      "votes": null
    },
    {
      "id": "2738088",
      "postDate": "04/06/2024 05:31:48",
      "content": "<p><a href=\"https://www.kaggle.com/catadanna\" target=\"_blank\">@catadanna</a> more importantly, using parquet/ feather to store files and saving the dataset with memory reduction is important for easier usage subsequently.</p>",
      "rawMarkdown": "catadanna more importantly, using parquet/ feather to store files and saving the dataset with memory reduction is important for easier usage subsequently.",
      "votes": null
    },
    {
      "id": "2741305",
      "postDate": "04/08/2024 08:42:03",
      "content": "<p>yes！！！it is so huge!!!</p>",
      "rawMarkdown": "yes！！！it is so huge!!!",
      "votes": null
    },
    {
      "id": "2759240",
      "postDate": "04/18/2024 14:54:55",
      "content": "<p>Isn't pyspark for parallel computing (With several machines)?<br>\nIsn't Vaex or Dask more suitable here?</p>",
      "rawMarkdown": "Isn't pyspark for parallel computing (With several machines)?\nIsn't Vaex or Dask more suitable here?",
      "votes": null
    },
    {
      "id": "2873864",
      "postDate": "06/15/2024 20:02:12",
      "content": "<p>You can use it on your computer and choose to use as many machines as the number of your cores. Anyway Spark is meant for bigdata so for huge datasets.</p>",
      "rawMarkdown": "You can use it on your computer and choose to use as many machines as the number of your cores. Anyway Spark is meant for bigdata so for huge datasets.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2737502,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "04/05/2024 19:27:38",
      "content": "<p><a href=\"https://www.kaggle.com/catadanna\" target=\"_blank\">@catadanna</a> you may consider using pyspark for this if you want <br>\nPandas is not suitable in my opinion for such a dataset </p>\n<p>Also parquet/ feather format is a better one that csv</p>",
      "votes": null,
      "replies": [
        {
          "id": 2737533,
          "author_name": "catadanna",
          "author_url": "",
          "post_date": "04/05/2024 19:47:08",
          "content": "<p>Good idea, pyspark ! </p>",
          "votes": null,
          "replies": [
            {
              "id": 2738088,
              "author_name": "ravi20076",
              "author_url": "",
              "post_date": "04/06/2024 05:31:48",
              "content": "<p><a href=\"https://www.kaggle.com/catadanna\" target=\"_blank\">@catadanna</a> more importantly, using parquet/ feather to store files and saving the dataset with memory reduction is important for easier usage subsequently.</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 2759240,
          "author_name": "gyulamaloveczky4",
          "author_url": "",
          "post_date": "04/18/2024 14:54:55",
          "content": "<p>Isn't pyspark for parallel computing (With several machines)?<br>\nIsn't Vaex or Dask more suitable here?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2873864,
              "author_name": "catadanna",
              "author_url": "",
              "post_date": "06/15/2024 20:02:12",
              "content": "<p>You can use it on your computer and choose to use as many machines as the number of your cores. Anyway Spark is meant for bigdata so for huge datasets.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2737710,
      "author_name": "shlomoron",
      "author_url": "",
      "post_date": "04/05/2024 22:47:33",
      "content": "<p>The actual size of the train set is ~10GB and is sorely bloated; see my <a href=\"https://www.kaggle.com/code/shlomoron/belka-shrinking-the-dataset\" target=\"_blank\">notebook here</a>. More details to follow soon.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2741305,
      "author_name": "qwerrtyty",
      "author_url": "",
      "post_date": "04/08/2024 08:42:03",
      "content": "<p>yes！！！it is so huge!!!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2737389": "The training csv dataset is huge : 59GB unzipped. \n\nDefinitely, we do not have enough RAM on Kaggle notebooks in order to manage this dataset. \n\nThis is how  I solved the problem:\n\n1. I downloaded the dataset on my computer (32Gb RAM)  and cut the file into smaller csv files. \n2. I made a model per small file, and then use the models for the final submission. \n3. On the other hand, if we want to encode the data, you need all the unique values in the columns (we have only categorical columns). We can see in the dataset in Kaggle the number of unique values. I iterated on the small files in order to get all the unique values in each column, excepting the one with more than 1000000 unique values, which I deleted.\n\nI am still working on this, it takes time ...",
    "2737502": "catadanna you may consider using pyspark for this if you want \nPandas is not suitable in my opinion for such a dataset \n\nAlso parquet/ feather format is a better one that csv",
    "2737533": "Good idea, pyspark !",
    "2737710": "The actual size of the train set is ~10GB and is sorely bloated; see my [notebook here](https://www.kaggle.com/code/shlomoron/belka-shrinking-the-dataset). More details to follow soon.",
    "2738088": "catadanna more importantly, using parquet/ feather to store files and saving the dataset with memory reduction is important for easier usage subsequently.",
    "2741305": "yes！！！it is so huge!!!",
    "2759240": "Isn't pyspark for parallel computing (With several machines)?\nIsn't Vaex or Dask more suitable here?",
    "2873864": "You can use it on your computer and choose to use as many machines as the number of your cores. Anyway Spark is meant for bigdata so for huge datasets."
  },
  "source": "meta"
}