{
  "id": 583124,
  "title": "ValueError: Input X contains infinity or a value too large for dtype('float64').",
  "url": "/competitions/drw-crypto-market-prediction/discussion/583124",
  "author_name": "",
  "post_date": "2025-06-04T23:09:21.554722300Z",
  "votes": 3,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi everyone<br>\nI have tried an exploratory data, and choose the best columns although PCA, or similar, but when I use standard scaler, I have the next error <em>ValueError: Input X contains infinity or a value too large for dtype('float64')</em>. I have the next code <code>print(\"¿Do you have infinity values?\", np.isinf(train_df).any())</code> to verify infinity values, but I don't find anything infinity values, also I have used the next code <code>missing_values_count = train_df.isnull().sum()</code> to verify NaN values, but I don't find anything NaN values. <br>\nBut when I want to use standard scaler although this code:<br>\n<code>scaler = StandardScaler() \nX_scaled = scaler.fit_transform(train_df)</code><br>\nI have the next error: ValueError: Input X contains infinity or a value too large for dtype('float64').<br>\nIn agree with the error, maybe exists value too large, but I don't want to change the data structure, what's the better solution for this problem.<br>\nI want to know what's your opinion?</p>",
  "messages": [
    {
      "id": "3217343",
      "postDate": "06/04/2025 23:09:21",
      "content": "<p>Hi everyone<br>\nI have tried an exploratory data, and choose the best columns although PCA, or similar, but when I use standard scaler, I have the next error <em>ValueError: Input X contains infinity or a value too large for dtype('float64')</em>. I have the next code <code>print(\"¿Do you have infinity values?\", np.isinf(train_df).any())</code> to verify infinity values, but I don't find anything infinity values, also I have used the next code <code>missing_values_count = train_df.isnull().sum()</code> to verify NaN values, but I don't find anything NaN values. <br>\nBut when I want to use standard scaler although this code:<br>\n<code>scaler = StandardScaler() \nX_scaled = scaler.fit_transform(train_df)</code><br>\nI have the next error: ValueError: Input X contains infinity or a value too large for dtype('float64').<br>\nIn agree with the error, maybe exists value too large, but I don't want to change the data structure, what's the better solution for this problem.<br>\nI want to know what's your opinion?</p>",
      "rawMarkdown": "Hi everyone\nI have tried an exploratory data, and choose the best columns although PCA, or similar, but when I use standard scaler, I have the next error *ValueError: Input X contains infinity or a value too large for dtype('float64')*. I have the next code `print(\"¿Do you have infinity values?\", np.isinf(train_df).any())` to verify infinity values, but I don't find anything infinity values, also I have used the next code `missing_values_count = train_df.isnull().sum()` to verify NaN values, but I don't find anything NaN values. \nBut when I want to use standard scaler although this code:\n`scaler = StandardScaler() \nX_scaled = scaler.fit_transform(train_df)`\nI have the next error: ValueError: Input X contains infinity or a value too large for dtype('float64').\nIn agree with the error, maybe exists value too large, but I don't want to change the data structure, what's the better solution for this problem.\nI want to know what's your opinion?",
      "votes": null
    },
    {
      "id": "3217595",
      "postDate": "06/05/2025 07:19:35",
      "content": "<p>You can use an intermediate step to convert infinity to np.nan -&gt; <code>torch.nan_to_num</code> is a good way to do this. Else, if you are using pandas, then you can simply use <code>df.replace([np.inf, -1*np.inf], np.nan])</code><br>\n<a href=\"https://www.kaggle.com/danielphys\" target=\"_blank\">@danielphys</a> </p>",
      "rawMarkdown": "You can use an intermediate step to convert infinity to np.nan -> `torch.nan_to_num` is a good way to do this. Else, if you are using pandas, then you can simply use `df.replace([np.inf, -1*np.inf], np.nan])`\n@danielphys",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3217595,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "06/05/2025 07:19:35",
      "content": "<p>You can use an intermediate step to convert infinity to np.nan -&gt; <code>torch.nan_to_num</code> is a good way to do this. Else, if you are using pandas, then you can simply use <code>df.replace([np.inf, -1*np.inf], np.nan])</code><br>\n<a href=\"https://www.kaggle.com/danielphys\" target=\"_blank\">@danielphys</a> </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3217343": "Hi everyone\nI have tried an exploratory data, and choose the best columns although PCA, or similar, but when I use standard scaler, I have the next error *ValueError: Input X contains infinity or a value too large for dtype('float64')*. I have the next code `print(\"¿Do you have infinity values?\", np.isinf(train_df).any())` to verify infinity values, but I don't find anything infinity values, also I have used the next code `missing_values_count = train_df.isnull().sum()` to verify NaN values, but I don't find anything NaN values. \nBut when I want to use standard scaler although this code:\n`scaler = StandardScaler() \nX_scaled = scaler.fit_transform(train_df)`\nI have the next error: ValueError: Input X contains infinity or a value too large for dtype('float64').\nIn agree with the error, maybe exists value too large, but I don't want to change the data structure, what's the better solution for this problem.\nI want to know what's your opinion?",
    "3217595": "You can use an intermediate step to convert infinity to np.nan -> `torch.nan_to_num` is a good way to do this. Else, if you are using pandas, then you can simply use `df.replace([np.inf, -1*np.inf], np.nan])`\n@danielphys"
  },
  "source": "meta"
}