{
  "id": 490837,
  "title": "Mixed Types in a columns",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/490837",
  "author_name": "",
  "post_date": "2024-04-03T17:54:48.511274100Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>In looking at the data in <code>train_static_0_0.csv</code> it appears there are mixed data types in some columns. See below:</p>\n<pre><code>ts_00 = (dataPath + )\nts_00()\n</code></pre>\n<blockquote>\n  <p>/opt/conda/lib/python3.10/site-packages/dask/dataframe/io/csv.py:195: DtypeWarning: Columns (20) have mixed types. Specify dtype option on import or set low_memory=False.</p>\n</blockquote>\n<pre><code>column_20_types = ts_00.iloc[:, ].apply(type, meta=(, )).value_counts().compute()\ncolumn_20_types\n\n\nbankacctype_710L\n&lt; ''&gt;      \n&lt; ''&gt;    \n\n</code></pre>\n<p>How have you been able to manage these columns that have mixed types in them?</p>",
  "messages": [
    {
      "id": "2733602",
      "postDate": "04/03/2024 17:54:48",
      "content": "<p>In looking at the data in <code>train_static_0_0.csv</code> it appears there are mixed data types in some columns. See below:</p>\n<pre><code>ts_00 = (dataPath + )\nts_00()\n</code></pre>\n<blockquote>\n  <p>/opt/conda/lib/python3.10/site-packages/dask/dataframe/io/csv.py:195: DtypeWarning: Columns (20) have mixed types. Specify dtype option on import or set low_memory=False.</p>\n</blockquote>\n<pre><code>column_20_types = ts_00.iloc[:, ].apply(type, meta=(, )).value_counts().compute()\ncolumn_20_types\n\n\nbankacctype_710L\n&lt; ''&gt;      \n&lt; ''&gt;    \n\n</code></pre>\n<p>How have you been able to manage these columns that have mixed types in them?</p>",
      "rawMarkdown": "In looking at the data in `train_static_0_0.csv` it appears there are mixed data types in some columns. See below:\n\n```\nts_00 = dd.read_csv(dataPath + \"csv_files/train/train_static_0_0.csv\")\nts_00.head()\n```\n\n>/opt/conda/lib/python3.10/site-packages/dask/dataframe/io/csv.py:195: DtypeWarning: Columns (20) have mixed types. Specify dtype option on import or set low_memory=False.\n\n```\ncolumn_20_types = ts_00.iloc[:, 20].apply(type, meta=('bankacctype_710L', 'object')).value_counts().compute()\ncolumn_20_types\n\n\nbankacctype_710L\n<class 'str'>      304690\n<class 'float'>    699067\nName: count, dtype: int64\n```\n\nHow have you been able to manage these columns that have mixed types in them?",
      "votes": null
    },
    {
      "id": "2736638",
      "postDate": "04/05/2024 10:11:45",
      "content": "<p>read from parquet files, they have column dtype predefined and you can force convert some types as indicated in data description</p>",
      "rawMarkdown": "read from parquet files, they have column dtype predefined and you can force convert some types as indicated in data description",
      "votes": null
    },
    {
      "id": "2736838",
      "postDate": "04/05/2024 12:33:56",
      "content": "<p><a href=\"https://www.kaggle.com/eu1234\" target=\"_blank\">@eu1234</a>  Perfect !! Thank you</p>",
      "rawMarkdown": "eu1234  Perfect !! Thank you",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2736638,
      "author_name": "eu1234",
      "author_url": "",
      "post_date": "04/05/2024 10:11:45",
      "content": "<p>read from parquet files, they have column dtype predefined and you can force convert some types as indicated in data description</p>",
      "votes": null,
      "replies": [
        {
          "id": 2736838,
          "author_name": "nadereafshar",
          "author_url": "",
          "post_date": "04/05/2024 12:33:56",
          "content": "<p><a href=\"https://www.kaggle.com/eu1234\" target=\"_blank\">@eu1234</a>  Perfect !! Thank you</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2733602": "In looking at the data in `train_static_0_0.csv` it appears there are mixed data types in some columns. See below:\n\n```\nts_00 = dd.read_csv(dataPath + \"csv_files/train/train_static_0_0.csv\")\nts_00.head()\n```\n\n>/opt/conda/lib/python3.10/site-packages/dask/dataframe/io/csv.py:195: DtypeWarning: Columns (20) have mixed types. Specify dtype option on import or set low_memory=False.\n\n```\ncolumn_20_types = ts_00.iloc[:, 20].apply(type, meta=('bankacctype_710L', 'object')).value_counts().compute()\ncolumn_20_types\n\n\nbankacctype_710L\n<class 'str'>      304690\n<class 'float'>    699067\nName: count, dtype: int64\n```\n\nHow have you been able to manage these columns that have mixed types in them?",
    "2736638": "read from parquet files, they have column dtype predefined and you can force convert some types as indicated in data description",
    "2736838": "eu1234  Perfect !! Thank you"
  },
  "source": "meta"
}