{
  "id": 546259,
  "title": "How do you handle the Nan value?",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/546259",
  "author_name": "",
  "post_date": "2024-11-14T16:37:04.262165700Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16529018%2Fc02435b683cbe3069dc1dd153b13359d%2F2024-11-15%20003424.png?generation=1731602241716268&amp;alt=media\" alt=\"\"></p>\n<p>It seems that feature00 ~ feature04 is a kind of rolling feature, beacause they only disappeare in the first 246 dates.And There may be some correlation between features with the same nan values.How do you handle the nan value?</p>",
  "messages": [
    {
      "id": "3045645",
      "postDate": "11/14/2024 16:37:04",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16529018%2Fc02435b683cbe3069dc1dd153b13359d%2F2024-11-15%20003424.png?generation=1731602241716268&amp;alt=media\" alt=\"\"></p>\n<p>It seems that feature00 ~ feature04 is a kind of rolling feature, beacause they only disappeare in the first 246 dates.And There may be some correlation between features with the same nan values.How do you handle the nan value?</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16529018%2Fc02435b683cbe3069dc1dd153b13359d%2F2024-11-15%20003424.png?generation=1731602241716268&alt=media)\n\nIt seems that feature00 ~ feature04 is a kind of rolling feature, beacause they only disappeare in the first 246 dates.And There may be some correlation between features with the same nan values.How do you handle the nan value?",
      "votes": null
    },
    {
      "id": "3045860",
      "postDate": "11/14/2024 21:27:32",
      "content": "<p>You can use .ffill for each parition to keep within the RAM restrictions, you could also probably do fill with mean or median by computing it first in a sequential manner storing it and the fillig in the same manner as I described the ffill.</p>",
      "rawMarkdown": "You can use .ffill for each parition to keep within the RAM restrictions, you could also probably do fill with mean or median by computing it first in a sequential manner storing it and the fillig in the same manner as I described the ffill.",
      "votes": null
    },
    {
      "id": "3046202",
      "postDate": "11/15/2024 08:09:39",
      "content": "<p>Why not to leave them as NaN? Any lgbm, xgb, catboost can handle missing values</p>",
      "rawMarkdown": "Why not to leave them as NaN? Any lgbm, xgb, catboost can handle missing values",
      "votes": null
    },
    {
      "id": "3046596",
      "postDate": "11/15/2024 16:00:20",
      "content": "<p>Because I wanna try to use neural network to train the data.</p>",
      "rawMarkdown": "Because I wanna try to use neural network to train the data.",
      "votes": null
    },
    {
      "id": "3047061",
      "postDate": "11/16/2024 07:58:06",
      "content": "<p>I tried filling with mean, median and 0 and so far 0 has worked best. But I'm sure this is not the best solution (I'm using NN btw)</p>",
      "rawMarkdown": "I tried filling with mean, median and 0 and so far 0 has worked best. But I'm sure this is not the best solution (I'm using NN btw)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3045860,
      "author_name": "denthe6",
      "author_url": "",
      "post_date": "11/14/2024 21:27:32",
      "content": "<p>You can use .ffill for each parition to keep within the RAM restrictions, you could also probably do fill with mean or median by computing it first in a sequential manner storing it and the fillig in the same manner as I described the ffill.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3046202,
      "author_name": "eu1234",
      "author_url": "",
      "post_date": "11/15/2024 08:09:39",
      "content": "<p>Why not to leave them as NaN? Any lgbm, xgb, catboost can handle missing values</p>",
      "votes": null,
      "replies": [
        {
          "id": 3046596,
          "author_name": "i2nfinit3y",
          "author_url": "",
          "post_date": "11/15/2024 16:00:20",
          "content": "<p>Because I wanna try to use neural network to train the data.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3047061,
      "author_name": "natanlabarrere",
      "author_url": "",
      "post_date": "11/16/2024 07:58:06",
      "content": "<p>I tried filling with mean, median and 0 and so far 0 has worked best. But I'm sure this is not the best solution (I'm using NN btw)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3045645": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16529018%2Fc02435b683cbe3069dc1dd153b13359d%2F2024-11-15%20003424.png?generation=1731602241716268&alt=media)\n\nIt seems that feature00 ~ feature04 is a kind of rolling feature, beacause they only disappeare in the first 246 dates.And There may be some correlation between features with the same nan values.How do you handle the nan value?",
    "3045860": "You can use .ffill for each parition to keep within the RAM restrictions, you could also probably do fill with mean or median by computing it first in a sequential manner storing it and the fillig in the same manner as I described the ffill.",
    "3046202": "Why not to leave them as NaN? Any lgbm, xgb, catboost can handle missing values",
    "3046596": "Because I wanna try to use neural network to train the data.",
    "3047061": "I tried filling with mean, median and 0 and so far 0 has worked best. But I'm sure this is not the best solution (I'm using NN btw)"
  },
  "source": "meta"
}