{
  "id": 499418,
  "title": "Files in Parquet train sets - please confirm on these columns/features?",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/499418",
  "author_name": "",
  "post_date": "2024-05-01T17:06:02.234450700Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Based on the files in the parquet train datasets, please confirm the following:</p>\n<ul>\n<li>Number of files = 32?</li>\n<li>Each file represents a separate dataset?</li>\n<li>Each dataset (dataframe) contains the same column names, which total 41 column names?</li>\n<li>I have attached an example showing 3 of the dataframes with their respective schemas. </li>\n<li>Just so I am clear we are dealing with separate dataframes I have included the total number of rows in each of the 3 dataframe examples I show in the attached image.</li>\n</ul>\n<p>I have observed there are no columns for WEEK_NUM, date_decision, MONTH and target. Am I missing something here?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16083849%2F264807eb6e49eab31ce0f9f0ddf6d452%2FExample_Dataframe_Schema_01.png?generation=1714582784969039&amp;alt=media\"></p>",
  "messages": [
    {
      "id": "2787300",
      "postDate": "05/01/2024 17:06:02",
      "content": "<p>Based on the files in the parquet train datasets, please confirm the following:</p>\n<ul>\n<li>Number of files = 32?</li>\n<li>Each file represents a separate dataset?</li>\n<li>Each dataset (dataframe) contains the same column names, which total 41 column names?</li>\n<li>I have attached an example showing 3 of the dataframes with their respective schemas. </li>\n<li>Just so I am clear we are dealing with separate dataframes I have included the total number of rows in each of the 3 dataframe examples I show in the attached image.</li>\n</ul>\n<p>I have observed there are no columns for WEEK_NUM, date_decision, MONTH and target. Am I missing something here?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16083849%2F264807eb6e49eab31ce0f9f0ddf6d452%2FExample_Dataframe_Schema_01.png?generation=1714582784969039&amp;alt=media\"></p>",
      "rawMarkdown": "Based on the files in the parquet train datasets, please confirm the following:\n- Number of files = 32?\n- Each file represents a separate dataset?\n- Each dataset (dataframe) contains the same column names, which total 41 column names?\n- I have attached an example showing 3 of the dataframes with their respective schemas. \n- Just so I am clear we are dealing with separate dataframes I have included the total number of rows in each of the 3 dataframe examples I show in the attached image.\n\nI have observed there are no columns for WEEK_NUM, date_decision, MONTH and target. Am I missing something here?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16083849%2F264807eb6e49eab31ce0f9f0ddf6d452%2FExample_Dataframe_Schema_01.png?generation=1714582784969039&alt=media)",
      "votes": null
    },
    {
      "id": "2790355",
      "postDate": "05/03/2024 06:43:50",
      "content": "<p>Each dataset (dataframe) contains the same column names, which total 41 column names?</p>\n<p>I am not very sure  what you mean by it, but columns of file 'train_applprev_1_1' are different from 'train_credit_bureau_a_1_0', although the similarity of column names will be there among 'train_credit_bureau_a_1_0' and 'train_credit_bureau_a_1_1'</p>",
      "rawMarkdown": "Each dataset (dataframe) contains the same column names, which total 41 column names?\n\nI am not very sure  what you mean by it, but columns of file 'train_applprev_1_1' are different from 'train_credit_bureau_a_1_0', although the similarity of column names will be there among 'train_credit_bureau_a_1_0' and 'train_credit_bureau_a_1_1'",
      "votes": null
    },
    {
      "id": "2790981",
      "postDate": "05/03/2024 12:12:38",
      "content": "<p>Did you have a look at the image I provided where I show the schemas for train_base, train_applprev_1_1 and train_credit_bureau_a_1_0? Please let me know if you see any differences in the names of the columns?</p>",
      "rawMarkdown": "Did you have a look at the image I provided where I show the schemas for train_base, train_applprev_1_1 and train_credit_bureau_a_1_0? Please let me know if you see any differences in the names of the columns?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2790355,
      "author_name": "shreyas9181",
      "author_url": "",
      "post_date": "05/03/2024 06:43:50",
      "content": "<p>Each dataset (dataframe) contains the same column names, which total 41 column names?</p>\n<p>I am not very sure  what you mean by it, but columns of file 'train_applprev_1_1' are different from 'train_credit_bureau_a_1_0', although the similarity of column names will be there among 'train_credit_bureau_a_1_0' and 'train_credit_bureau_a_1_1'</p>",
      "votes": null,
      "replies": [
        {
          "id": 2790981,
          "author_name": "michaelwaynetreasure",
          "author_url": "",
          "post_date": "05/03/2024 12:12:38",
          "content": "<p>Did you have a look at the image I provided where I show the schemas for train_base, train_applprev_1_1 and train_credit_bureau_a_1_0? Please let me know if you see any differences in the names of the columns?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2787300": "Based on the files in the parquet train datasets, please confirm the following:\n- Number of files = 32?\n- Each file represents a separate dataset?\n- Each dataset (dataframe) contains the same column names, which total 41 column names?\n- I have attached an example showing 3 of the dataframes with their respective schemas. \n- Just so I am clear we are dealing with separate dataframes I have included the total number of rows in each of the 3 dataframe examples I show in the attached image.\n\nI have observed there are no columns for WEEK_NUM, date_decision, MONTH and target. Am I missing something here?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16083849%2F264807eb6e49eab31ce0f9f0ddf6d452%2FExample_Dataframe_Schema_01.png?generation=1714582784969039&alt=media)",
    "2790355": "Each dataset (dataframe) contains the same column names, which total 41 column names?\n\nI am not very sure  what you mean by it, but columns of file 'train_applprev_1_1' are different from 'train_credit_bureau_a_1_0', although the similarity of column names will be there among 'train_credit_bureau_a_1_0' and 'train_credit_bureau_a_1_1'",
    "2790981": "Did you have a look at the image I provided where I show the schemas for train_base, train_applprev_1_1 and train_credit_bureau_a_1_0? Please let me know if you see any differences in the names of the columns?"
  },
  "source": "meta"
}