{
  "id": 499368,
  "title": "dtypes of train and test",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/499368",
  "author_name": "",
  "post_date": "2024-05-01T14:13:10.665908400Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Is there a way to ensure same types and order for train and test</p>",
  "messages": [
    {
      "id": "2786949",
      "postDate": "05/01/2024 14:13:10",
      "content": "<p>Is there a way to ensure same types and order for train and test</p>",
      "rawMarkdown": "Is there a way to ensure same types and order for train and test",
      "votes": null
    },
    {
      "id": "2790363",
      "postDate": "05/03/2024 06:48:56",
      "content": "<p><a href=\"https://www.kaggle.com/code/yunsuxiaozi/home-credit-inconsistent-data-types\" target=\"_blank\">https://www.kaggle.com/code/yunsuxiaozi/home-credit-inconsistent-data-types</a></p>",
      "rawMarkdown": "https://www.kaggle.com/code/yunsuxiaozi/home-credit-inconsistent-data-types",
      "votes": null
    },
    {
      "id": "2794864",
      "postDate": "05/05/2024 14:50:02",
      "content": "<p>Use below code before you split the training data to train &amp; test:</p>\n<h1>standardization of dtypes of full data</h1>\n<p>fill_cat_values = {}<br>\nfor col in cat_cols:<br>\n    cats = df_train[col].cat.categories<br>\n    fill_cat_values[col] = dict(zip(cats,range(len(cats))))<br>\n    df_train[col].replace(fill_cat_values[col],inplace=True)<br>\n    if df_train[col].isnull().sum() &gt; 0:<br>\n        df_train[col] = df_train[col].cat.add_categories([-1]).fillna(-1)</p>",
      "rawMarkdown": "Use below code before you split the training data to train & test:\n\n# standardization of dtypes of full data  \nfill_cat_values = {}\nfor col in cat_cols:\n    cats = df_train[col].cat.categories\n    fill_cat_values[col] = dict(zip(cats,range(len(cats))))\n    df_train[col].replace(fill_cat_values[col],inplace=True)\n    if df_train[col].isnull().sum() > 0:\n        df_train[col] = df_train[col].cat.add_categories([-1]).fillna(-1)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2790363,
      "author_name": "yunsuxiaozi",
      "author_url": "",
      "post_date": "05/03/2024 06:48:56",
      "content": "<p><a href=\"https://www.kaggle.com/code/yunsuxiaozi/home-credit-inconsistent-data-types\" target=\"_blank\">https://www.kaggle.com/code/yunsuxiaozi/home-credit-inconsistent-data-types</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2794864,
      "author_name": "kunduruanil",
      "author_url": "",
      "post_date": "05/05/2024 14:50:02",
      "content": "<p>Use below code before you split the training data to train &amp; test:</p>\n<h1>standardization of dtypes of full data</h1>\n<p>fill_cat_values = {}<br>\nfor col in cat_cols:<br>\n    cats = df_train[col].cat.categories<br>\n    fill_cat_values[col] = dict(zip(cats,range(len(cats))))<br>\n    df_train[col].replace(fill_cat_values[col],inplace=True)<br>\n    if df_train[col].isnull().sum() &gt; 0:<br>\n        df_train[col] = df_train[col].cat.add_categories([-1]).fillna(-1)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2786949": "Is there a way to ensure same types and order for train and test",
    "2790363": "https://www.kaggle.com/code/yunsuxiaozi/home-credit-inconsistent-data-types",
    "2794864": "Use below code before you split the training data to train & test:\n\n# standardization of dtypes of full data  \nfill_cat_values = {}\nfor col in cat_cols:\n    cats = df_train[col].cat.categories\n    fill_cat_values[col] = dict(zip(cats,range(len(cats))))\n    df_train[col].replace(fill_cat_values[col],inplace=True)\n    if df_train[col].isnull().sum() > 0:\n        df_train[col] = df_train[col].cat.add_categories([-1]).fillna(-1)"
  },
  "source": "meta"
}