{
  "id": 483682,
  "title": "What data format should be used?",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/483682",
  "author_name": "Andreas Bisiadis",
  "post_date": "2024-03-13T14:27:20.842000",
  "votes": 1,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hello there,</p>\n<p>This competition serves as a great learning experience for me to get familiar working with tabular data. However, I am having some concerns regarding data loading and preprocessing. Specifically:</p>\n<ol>\n<li>The train and test data are both present is <code>.csv</code> and <code>.parquet</code> format. Is important which data format to use during training and inference?</li>\n<li>My assumption may be false, but most training notebook do not use all files for training. Considering this code snippet from the <a href=\"https://www.kaggle.com/code/greysky/home-credit-baseline\" target=\"_blank\">Home Credit Baseline</a> notebook:</li>\n</ol>\n<pre><code>data_store = {\n    : read_file(TRAIN_DIR / ),\n    : [\n        read_file(TRAIN_DIR / ),\n        read_files(TRAIN_DIR / ),\n    ],\n    : [\n        read_files(TRAIN_DIR / , ),\n        read_file(TRAIN_DIR / , ),\n        read_file(TRAIN_DIR / , ),\n        read_file(TRAIN_DIR / , ),\n        read_file(TRAIN_DIR / , ),\n        read_file(TRAIN_DIR / , ),\n        read_file(TRAIN_DIR / , ),\n        read_file(TRAIN_DIR / , ),\n        read_file(TRAIN_DIR / , ),\n    ],\n    : [\n        read_file(TRAIN_DIR / , ),\n    ]\n}\n</code></pre>\n<p>Only 13 out of the 32 train files are used. Why is this the case? Are the remaining files secondary?</p>",
  "messages": [
    {
      "id": 2696332,
      "postDate": "2024-03-14T08:39:02.593Z",
      "content": "<p>Hey I am new to this competition and I think choosing between CSV and Parquet is simple, always go for parquet. It is preferred for large datasets and distributed computing environments. CSV as the host mentioned is for new users.</p>\n<p>As for exclusions, I think file exclusion is mostly because of the resource constrain. Eventually as the competition goes on, we will probably figure out the best combination of files to use to save resources and still get the relevant information. Till then I would focus on the algorithm itself for whatever files you end up selecting and improve that.</p>",
      "rawMarkdown": "Hey I am new to this competition and I think choosing between CSV and Parquet is simple, always go for parquet. It is preferred for large datasets and distributed computing environments. CSV as the host mentioned is for new users.\n\nAs for exclusions, I think file exclusion is mostly because of the resource constrain. Eventually as the competition goes on, we will probably figure out the best combination of files to use to save resources and still get the relevant information. Till then I would focus on the algorithm itself for whatever files you end up selecting and improve that.",
      "votes": 3,
      "replies": [
        {
          "id": 2696358,
          "postDate": "2024-03-14T09:10:18.363Z",
          "content": "<p>I agree, Polars and parquet is a great EDA combination <a href=\"https://www.kaggle.com/propriyam\" target=\"_blank\">@propriyam</a> </p>",
          "rawMarkdown": "I agree, Polars and parquet is a great EDA combination @propriyam ",
          "votes": 2
        }
      ]
    },
    {
      "id": 2698288,
      "postDate": "2024-03-15T11:48:45.127Z",
      "content": "<blockquote>\n  <p>Only 13 out of the 32 train files are used.</p>\n</blockquote>\n<p>Hi,</p>\n<pre><code>read_files(TRAIN_DIR / ),\n</code></pre>\n<p>The * is a <a href=\"https://en.wikipedia.org/wiki/Wildcard_character\" target=\"_blank\">wildcard</a>, any file that starts with 'train_static_0_' and ends with '.parquet'</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14573416%2Ff69a42a89ddd665471d8a1e1b397755f%2FScreenshot%202024-03-15%20at%2012.12.44PM.png?generation=1710501198975449&amp;alt=media\"></p>\n<p>For example, the file train_credit_bureau_a_2_ has eleven parts that can be combined to create the complete file, but is just that, one file.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14573416%2F0878f04ce35b87b16a5ea10cbaa2c72d%2FScreenshot%202024-03-15%20at%2012.16.22PM.png?generation=1710501633321918&amp;alt=media\"></p>\n<p><a href=\"https://docs.pola.rs/user-guide/io/multiple/\" target=\"_blank\">Polars examples.</a></p>\n<p>The code in your post is missing five files I think.</p>",
      "rawMarkdown": ">Only 13 out of the 32 train files are used.\n\nHi,\n\n```python\nread_files(TRAIN_DIR / \"train_static_0_*.parquet\"),\n```\n\nThe * is a [wildcard](https://en.wikipedia.org/wiki/Wildcard_character), any file that starts with 'train_static_0_' and ends with '.parquet'\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14573416%2Ff69a42a89ddd665471d8a1e1b397755f%2FScreenshot%202024-03-15%20at%2012.12.44PM.png?generation=1710501198975449&alt=media)\n\nFor example, the file train_credit_bureau_a_2_ has eleven parts that can be combined to create the complete file, but is just that, one file.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14573416%2F0878f04ce35b87b16a5ea10cbaa2c72d%2FScreenshot%202024-03-15%20at%2012.16.22PM.png?generation=1710501633321918&alt=media)\n\n[Polars examples.](https://docs.pola.rs/user-guide/io/multiple/)\n\nThe code in your post is missing five files I think.",
      "votes": 1,
      "replies": [
        {
          "id": 2698911,
          "postDate": "2024-03-15T17:12:47.457Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 2695503,
      "postDate": "2024-03-13T17:50:46.150Z",
      "content": "<p>1) the file format is not important<br>\nbut working with parquet is much faster</p>\n<p>2) Not everything is used, because they weigh a lot and you need to somehow process it all in memory. On the contrary, for the most part, what is not in public notebooks will be more informative</p>",
      "rawMarkdown": "1) the file format is not important\nbut working with parquet is much faster\n\n2) Not everything is used, because they weigh a lot and you need to somehow process it all in memory. On the contrary, for the most part, what is not in public notebooks will be more informative",
      "votes": 1,
      "replies": [
        {
          "id": 2695527,
          "postDate": "2024-03-13T18:02:14.867Z",
          "rawMarkdown": "",
          "isDeleted": true,
          "replies": [
            {
              "id": 2696625,
              "postDate": "2024-03-14T12:57:12.060Z",
              "content": "<p>The file format will not matter when you load data into pandas, for example.<br>\nThe advantage of parquet is that it can be quickly saved/loaded + it contains meta-information on the columns (let’s say their types)<br>\nIn this competition, you won’t be able to take all the files, load them into memory and work quietly. You need to process it piece by piece, draw your own conclusions, what needs to be left, what is not needed.</p>",
              "rawMarkdown": "The file format will not matter when you load data into pandas, for example.\nThe advantage of parquet is that it can be quickly saved/loaded + it contains meta-information on the columns (let’s say their types)\nIn this competition, you won’t be able to take all the files, load them into memory and work quietly. You need to process it piece by piece, draw your own conclusions, what needs to be left, what is not needed.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2695164,
      "postDate": "2024-03-13T14:27:20.843Z",
      "content": "<p>Hello there,</p>\n<p>This competition serves as a great learning experience for me to get familiar working with tabular data. However, I am having some concerns regarding data loading and preprocessing. Specifically:</p>\n<ol>\n<li>The train and test data are both present is <code>.csv</code> and <code>.parquet</code> format. Is important which data format to use during training and inference?</li>\n<li>My assumption may be false, but most training notebook do not use all files for training. Considering this code snippet from the <a href=\"https://www.kaggle.com/code/greysky/home-credit-baseline\" target=\"_blank\">Home Credit Baseline</a> notebook:</li>\n</ol>\n<pre><code>data_store = {\n    : read_file(TRAIN_DIR / ),\n    : [\n        read_file(TRAIN_DIR / ),\n        read_files(TRAIN_DIR / ),\n    ],\n    : [\n        read_files(TRAIN_DIR / , ),\n        read_file(TRAIN_DIR / , ),\n        read_file(TRAIN_DIR / , ),\n        read_file(TRAIN_DIR / , ),\n        read_file(TRAIN_DIR / , ),\n        read_file(TRAIN_DIR / , ),\n        read_file(TRAIN_DIR / , ),\n        read_file(TRAIN_DIR / , ),\n        read_file(TRAIN_DIR / , ),\n    ],\n    : [\n        read_file(TRAIN_DIR / , ),\n    ]\n}\n</code></pre>\n<p>Only 13 out of the 32 train files are used. Why is this the case? Are the remaining files secondary?</p>",
      "rawMarkdown": "Hello there,\n\nThis competition serves as a great learning experience for me to get familiar working with tabular data. However, I am having some concerns regarding data loading and preprocessing. Specifically:\n\n1. The train and test data are both present is `.csv` and `.parquet` format. Is important which data format to use during training and inference?\n2. My assumption may be false, but most training notebook do not use all files for training. Considering this code snippet from the [Home Credit Baseline](https://www.kaggle.com/code/greysky/home-credit-baseline) notebook:\n\n```python\ndata_store = {\n    \"df_base\": read_file(TRAIN_DIR / \"train_base.parquet\"),\n    \"depth_0\": [\n        read_file(TRAIN_DIR / \"train_static_cb_0.parquet\"),\n        read_files(TRAIN_DIR / \"train_static_0_*.parquet\"),\n    ],\n    \"depth_1\": [\n        read_files(TRAIN_DIR / \"train_applprev_1_*.parquet\", 1),\n        read_file(TRAIN_DIR / \"train_tax_registry_a_1.parquet\", 1),\n        read_file(TRAIN_DIR / \"train_tax_registry_b_1.parquet\", 1),\n        read_file(TRAIN_DIR / \"train_tax_registry_c_1.parquet\", 1),\n        read_file(TRAIN_DIR / \"train_credit_bureau_b_1.parquet\", 1),\n        read_file(TRAIN_DIR / \"train_other_1.parquet\", 1),\n        read_file(TRAIN_DIR / \"train_person_1.parquet\", 1),\n        read_file(TRAIN_DIR / \"train_deposit_1.parquet\", 1),\n        read_file(TRAIN_DIR / \"train_debitcard_1.parquet\", 1),\n    ],\n    \"depth_2\": [\n        read_file(TRAIN_DIR / \"train_credit_bureau_b_2.parquet\", 2),\n    ]\n}\n```\n\nOnly 13 out of the 32 train files are used. Why is this the case? Are the remaining files secondary?",
      "votes": 1
    },
    {
      "id": 2695552,
      "postDate": "2024-03-13T18:18:46.133Z",
      "content": "<p>We were hesitant whether to include .csv as well, but we thought new kagglers that want to enter datasci and will participate as a learning experience, might find that helpful. </p>",
      "rawMarkdown": "We were hesitant whether to include .csv as well, but we thought new kagglers that want to enter datasci and will participate as a learning experience, might find that helpful. ",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 2696332,
      "author_name": "ProPriyam",
      "author_url": "",
      "post_date": "2024-03-14T08:39:02.593000",
      "content": "<p>Hey I am new to this competition and I think choosing between CSV and Parquet is simple, always go for parquet. It is preferred for large datasets and distributed computing environments. CSV as the host mentioned is for new users.</p>\n<p>As for exclusions, I think file exclusion is mostly because of the resource constrain. Eventually as the competition goes on, we will probably figure out the best combination of files to use to save resources and still get the relevant information. Till then I would focus on the algorithm itself for whatever files you end up selecting and improve that.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2696358,
          "author_name": "Ravi Ramakrishnan",
          "author_url": "",
          "post_date": "2024-03-14T09:10:18.363000",
          "content": "<p>I agree, Polars and parquet is a great EDA combination <a href=\"https://www.kaggle.com/propriyam\" target=\"_blank\">@propriyam</a> </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2698288,
      "author_name": "Antonio Félix",
      "author_url": "",
      "post_date": "2024-03-15T11:48:45.127000",
      "content": "<blockquote>\n  <p>Only 13 out of the 32 train files are used.</p>\n</blockquote>\n<p>Hi,</p>\n<pre><code>read_files(TRAIN_DIR / ),\n</code></pre>\n<p>The * is a <a href=\"https://en.wikipedia.org/wiki/Wildcard_character\" target=\"_blank\">wildcard</a>, any file that starts with 'train_static_0_' and ends with '.parquet'</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14573416%2Ff69a42a89ddd665471d8a1e1b397755f%2FScreenshot%202024-03-15%20at%2012.12.44PM.png?generation=1710501198975449&amp;alt=media\"></p>\n<p>For example, the file train_credit_bureau_a_2_ has eleven parts that can be combined to create the complete file, but is just that, one file.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14573416%2F0878f04ce35b87b16a5ea10cbaa2c72d%2FScreenshot%202024-03-15%20at%2012.16.22PM.png?generation=1710501633321918&amp;alt=media\"></p>\n<p><a href=\"https://docs.pola.rs/user-guide/io/multiple/\" target=\"_blank\">Polars examples.</a></p>\n<p>The code in your post is missing five files I think.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2698911,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-03-15T17:12:47.457000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2695503,
      "author_name": "dima",
      "author_url": "",
      "post_date": "2024-03-13T17:50:46.150000",
      "content": "<p>1) the file format is not important<br>\nbut working with parquet is much faster</p>\n<p>2) Not everything is used, because they weigh a lot and you need to somehow process it all in memory. On the contrary, for the most part, what is not in public notebooks will be more informative</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2695527,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-03-13T18:02:14.867000",
          "content": "",
          "votes": 0,
          "replies": [
            {
              "id": 2696625,
              "author_name": "dima",
              "author_url": "",
              "post_date": "2024-03-14T12:57:12.060000",
              "content": "<p>The file format will not matter when you load data into pandas, for example.<br>\nThe advantage of parquet is that it can be quickly saved/loaded + it contains meta-information on the columns (let’s say their types)<br>\nIn this competition, you won’t be able to take all the files, load them into memory and work quietly. You need to process it piece by piece, draw your own conclusions, what needs to be left, what is not needed.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2695552,
      "author_name": "Daniel Herman",
      "author_url": "",
      "post_date": "2024-03-13T18:18:46.133000",
      "content": "<p>We were hesitant whether to include .csv as well, but we thought new kagglers that want to enter datasci and will participate as a learning experience, might find that helpful. </p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2696332": "Hey I am new to this competition and I think choosing between CSV and Parquet is simple, always go for parquet. It is preferred for large datasets and distributed computing environments. CSV as the host mentioned is for new users.\n\nAs for exclusions, I think file exclusion is mostly because of the resource constrain. Eventually as the competition goes on, we will probably figure out the best combination of files to use to save resources and still get the relevant information. Till then I would focus on the algorithm itself for whatever files you end up selecting and improve that.",
    "2698288": ">Only 13 out of the 32 train files are used.\n\nHi,\n\n```python\nread_files(TRAIN_DIR / \"train_static_0_*.parquet\"),\n```\n\nThe * is a [wildcard](https://en.wikipedia.org/wiki/Wildcard_character), any file that starts with 'train_static_0_' and ends with '.parquet'\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14573416%2Ff69a42a89ddd665471d8a1e1b397755f%2FScreenshot%202024-03-15%20at%2012.12.44PM.png?generation=1710501198975449&alt=media)\n\nFor example, the file train_credit_bureau_a_2_ has eleven parts that can be combined to create the complete file, but is just that, one file.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F14573416%2F0878f04ce35b87b16a5ea10cbaa2c72d%2FScreenshot%202024-03-15%20at%2012.16.22PM.png?generation=1710501633321918&alt=media)\n\n[Polars examples.](https://docs.pola.rs/user-guide/io/multiple/)\n\nThe code in your post is missing five files I think.",
    "2695503": "1) the file format is not important\nbut working with parquet is much faster\n\n2) Not everything is used, because they weigh a lot and you need to somehow process it all in memory. On the contrary, for the most part, what is not in public notebooks will be more informative",
    "2695164": "Hello there,\n\nThis competition serves as a great learning experience for me to get familiar working with tabular data. However, I am having some concerns regarding data loading and preprocessing. Specifically:\n\n1. The train and test data are both present is `.csv` and `.parquet` format. Is important which data format to use during training and inference?\n2. My assumption may be false, but most training notebook do not use all files for training. Considering this code snippet from the [Home Credit Baseline](https://www.kaggle.com/code/greysky/home-credit-baseline) notebook:\n\n```python\ndata_store = {\n    \"df_base\": read_file(TRAIN_DIR / \"train_base.parquet\"),\n    \"depth_0\": [\n        read_file(TRAIN_DIR / \"train_static_cb_0.parquet\"),\n        read_files(TRAIN_DIR / \"train_static_0_*.parquet\"),\n    ],\n    \"depth_1\": [\n        read_files(TRAIN_DIR / \"train_applprev_1_*.parquet\", 1),\n        read_file(TRAIN_DIR / \"train_tax_registry_a_1.parquet\", 1),\n        read_file(TRAIN_DIR / \"train_tax_registry_b_1.parquet\", 1),\n        read_file(TRAIN_DIR / \"train_tax_registry_c_1.parquet\", 1),\n        read_file(TRAIN_DIR / \"train_credit_bureau_b_1.parquet\", 1),\n        read_file(TRAIN_DIR / \"train_other_1.parquet\", 1),\n        read_file(TRAIN_DIR / \"train_person_1.parquet\", 1),\n        read_file(TRAIN_DIR / \"train_deposit_1.parquet\", 1),\n        read_file(TRAIN_DIR / \"train_debitcard_1.parquet\", 1),\n    ],\n    \"depth_2\": [\n        read_file(TRAIN_DIR / \"train_credit_bureau_b_2.parquet\", 2),\n    ]\n}\n```\n\nOnly 13 out of the 32 train files are used. Why is this the case? Are the remaining files secondary?",
    "2695552": "We were hesitant whether to include .csv as well, but we thought new kagglers that want to enter datasci and will participate as a learning experience, might find that helpful. "
  }
}