{
  "id": 586341,
  "title": "What is efficient way to handle Inf value ?",
  "url": "/competitions/drw-crypto-market-prediction/discussion/586341",
  "author_name": "KapilSuara",
  "post_date": "2025-06-26T08:11:55.077000",
  "votes": 2,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I observed that many features in dataset contain values that are too large for float64, leading to errors during preprocessing. When I replaced these infinite values with NaN and then applied imputation using the mean or median, it significantly altered the original distribution of those features. This, in turn, impacted statistical properties such as correlations and potentially model performance. I'm looking for a more robust strategy to handle these extreme or infinite values without distorting the data's underlying structure?</p>",
  "messages": [
    {
      "id": 3233139,
      "postDate": "2025-06-26T13:52:36.660Z",
      "content": "<p>Base on this code, if a column contain a <code>inf</code>, then all of its data are <code>inf</code>. I think these columns should be dropped.</p>\n<pre><code> pandas  pd\n\nfile_path = \ndf = pd.read_parquet(file_path)\n\nany_inf_columns = df.columns[df.isin([(), -()]).()]\nall_inf_columns = df.columns[df.isin([(), -()]).()]\n\n any_inf_columns.():\n    ()\n    (any_inf_columns)\n:\n    ()\n\n all_inf_columns.():\n    ()\n    (all_inf_columns)\n:\n    ()\n\n (any_inf_columns == all_inf_columns).():\n    ()\n</code></pre>\n<p>The output of this code:</p>\n<pre><code>colums that contain inf data:\n([, , , , , , , , ,\n       , , , , , , , , ,\n       , , ],\n      dtype=)\ncolums   data  inf:\n([, , , , , , , , ,\n       , , , , , , , , ,\n       , , ],\n      dtype=)\n inf columns are the same   inf columns\n</code></pre>",
      "rawMarkdown": "Base on this code, if a column contain a `inf`, then all of its data are `inf`. I think these columns should be dropped.\n```\nimport pandas as pd\n\nfile_path = 'data/train.parquet'\ndf = pd.read_parquet(file_path)\n\nany_inf_columns = df.columns[df.isin([float('inf'), -float('inf')]).any()]\nall_inf_columns = df.columns[df.isin([float('inf'), -float('inf')]).all()]\n\nif any_inf_columns.any():\n    print(\"colums that contain inf data:\")\n    print(any_inf_columns)\nelse:\n    print(\"does not contain inf data.\")\n\nif all_inf_columns.any():\n    print(\"colums where all data is inf:\")\n    print(all_inf_columns)\nelse:\n    print(\"does not contain columns where all data is inf.\")\n\nif (any_inf_columns == all_inf_columns).all():\n    print(\"any inf columns are the same as all inf columns\")\n\n```\nThe output of this code:\n```\ncolums that contain inf data:\nIndex(['X697', 'X698', 'X699', 'X700', 'X701', 'X702', 'X703', 'X704', 'X705',\n       'X706', 'X707', 'X708', 'X709', 'X710', 'X711', 'X712', 'X713', 'X714',\n       'X715', 'X716', 'X717'],\n      dtype='object')\ncolums where all data is inf:\nIndex(['X697', 'X698', 'X699', 'X700', 'X701', 'X702', 'X703', 'X704', 'X705',\n       'X706', 'X707', 'X708', 'X709', 'X710', 'X711', 'X712', 'X713', 'X714',\n       'X715', 'X716', 'X717'],\n      dtype='object')\nany inf columns are the same as all inf columns\n```",
      "votes": 1
    },
    {
      "id": 3238085,
      "postDate": "2025-07-01T14:27:43.037Z",
      "content": "<p>KNN imputer is great. You can also do a simple replace with the nearest valid value, or 2x the nearest valid value.</p>",
      "rawMarkdown": "KNN imputer is great. You can also do a simple replace with the nearest valid value, or 2x the nearest valid value.",
      "votes": 2
    },
    {
      "id": 3233994,
      "postDate": "2025-06-27T11:59:07.913Z",
      "content": "<p>KNN imputer from sklearn is more robust. </p>",
      "rawMarkdown": "KNN imputer from sklearn is more robust. ",
      "votes": 2
    },
    {
      "id": 3232793,
      "postDate": "2025-06-26T08:11:55.077Z",
      "content": "<p>I observed that many features in dataset contain values that are too large for float64, leading to errors during preprocessing. When I replaced these infinite values with NaN and then applied imputation using the mean or median, it significantly altered the original distribution of those features. This, in turn, impacted statistical properties such as correlations and potentially model performance. I'm looking for a more robust strategy to handle these extreme or infinite values without distorting the data's underlying structure?</p>",
      "rawMarkdown": "I observed that many features in dataset contain values that are too large for float64, leading to errors during preprocessing. When I replaced these infinite values with NaN and then applied imputation using the mean or median, it significantly altered the original distribution of those features. This, in turn, impacted statistical properties such as correlations and potentially model performance. I'm looking for a more robust strategy to handle these extreme or infinite values without distorting the data's underlying structure?\n\n",
      "votes": 2
    },
    {
      "id": 3233131,
      "postDate": "2025-06-26T13:49:46.110Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 3233133,
          "postDate": "2025-06-26T13:50:57.187Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3233139,
      "author_name": "jinghong liu",
      "author_url": "",
      "post_date": "2025-06-26T13:52:36.660000",
      "content": "<p>Base on this code, if a column contain a <code>inf</code>, then all of its data are <code>inf</code>. I think these columns should be dropped.</p>\n<pre><code> pandas  pd\n\nfile_path = \ndf = pd.read_parquet(file_path)\n\nany_inf_columns = df.columns[df.isin([(), -()]).()]\nall_inf_columns = df.columns[df.isin([(), -()]).()]\n\n any_inf_columns.():\n    ()\n    (any_inf_columns)\n:\n    ()\n\n all_inf_columns.():\n    ()\n    (all_inf_columns)\n:\n    ()\n\n (any_inf_columns == all_inf_columns).():\n    ()\n</code></pre>\n<p>The output of this code:</p>\n<pre><code>colums that contain inf data:\n([, , , , , , , , ,\n       , , , , , , , , ,\n       , , ],\n      dtype=)\ncolums   data  inf:\n([, , , , , , , , ,\n       , , , , , , , , ,\n       , , ],\n      dtype=)\n inf columns are the same   inf columns\n</code></pre>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3238085,
      "author_name": "Taylor S. Amarel",
      "author_url": "",
      "post_date": "2025-07-01T14:27:43.037000",
      "content": "<p>KNN imputer is great. You can also do a simple replace with the nearest valid value, or 2x the nearest valid value.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3233994,
      "author_name": "Rafał Pawłowski",
      "author_url": "",
      "post_date": "2025-06-27T11:59:07.913000",
      "content": "<p>KNN imputer from sklearn is more robust. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3233131,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-06-26T13:49:46.110000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 3233133,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-06-26T13:50:57.187000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3233139": "Base on this code, if a column contain a `inf`, then all of its data are `inf`. I think these columns should be dropped.\n```\nimport pandas as pd\n\nfile_path = 'data/train.parquet'\ndf = pd.read_parquet(file_path)\n\nany_inf_columns = df.columns[df.isin([float('inf'), -float('inf')]).any()]\nall_inf_columns = df.columns[df.isin([float('inf'), -float('inf')]).all()]\n\nif any_inf_columns.any():\n    print(\"colums that contain inf data:\")\n    print(any_inf_columns)\nelse:\n    print(\"does not contain inf data.\")\n\nif all_inf_columns.any():\n    print(\"colums where all data is inf:\")\n    print(all_inf_columns)\nelse:\n    print(\"does not contain columns where all data is inf.\")\n\nif (any_inf_columns == all_inf_columns).all():\n    print(\"any inf columns are the same as all inf columns\")\n\n```\nThe output of this code:\n```\ncolums that contain inf data:\nIndex(['X697', 'X698', 'X699', 'X700', 'X701', 'X702', 'X703', 'X704', 'X705',\n       'X706', 'X707', 'X708', 'X709', 'X710', 'X711', 'X712', 'X713', 'X714',\n       'X715', 'X716', 'X717'],\n      dtype='object')\ncolums where all data is inf:\nIndex(['X697', 'X698', 'X699', 'X700', 'X701', 'X702', 'X703', 'X704', 'X705',\n       'X706', 'X707', 'X708', 'X709', 'X710', 'X711', 'X712', 'X713', 'X714',\n       'X715', 'X716', 'X717'],\n      dtype='object')\nany inf columns are the same as all inf columns\n```",
    "3238085": "KNN imputer is great. You can also do a simple replace with the nearest valid value, or 2x the nearest valid value.",
    "3233994": "KNN imputer from sklearn is more robust. ",
    "3232793": "I observed that many features in dataset contain values that are too large for float64, leading to errors during preprocessing. When I replaced these infinite values with NaN and then applied imputation using the mean or median, it significantly altered the original distribution of those features. This, in turn, impacted statistical properties such as correlations and potentially model performance. I'm looking for a more robust strategy to handle these extreme or infinite values without distorting the data's underlying structure?\n\n",
    "3233131": ""
  }
}