{
  "id": 202851,
  "title": "How are you imputing missing values ?",
  "url": "/competitions/predict-volcanic-eruptions-ingv-oe/discussion/202851",
  "author_name": "Harshit Gupta",
  "post_date": "2020-12-12T10:07:57.804000",
  "votes": 0,
  "comment_count": 6,
  "views": 0,
  "content": "<p>In case of a few segment ids, all values of some sensors are missing, maybe due to sensor malfunction. I've tried imputing with single mean values of other sensors and row mean(axis=1). In both cases, the performance decreases in comparison to leaving NaNs as it is. </p>\n<p>How are you dealing with NaNs ? </p>",
  "messages": [
    {
      "id": 1110073,
      "postDate": "2020-12-12T11:59:24.397Z",
      "content": "<p>hi, <br>\nI'm still trying to impute missing values ​​right.  I am working on the calculated characteristics and first I tried KNN, I had very bad results.  Then with median and improved a little.  now i am using the mean and it improved another bit.  I'm still trying to improve this though.</p>",
      "rawMarkdown": "hi, \nI'm still trying to impute missing values ​​right.  I am working on the calculated characteristics and first I tried KNN, I had very bad results.  Then with median and improved a little.  now i am using the mean and it improved another bit.  I'm still trying to improve this though.",
      "votes": 1,
      "replies": [
        {
          "id": 1110127,
          "postDate": "2020-12-12T13:01:08.027Z",
          "content": "<p>Hey, thanks for sharing. </p>\n<p>In case where all values of a sensor column are missing, I tried imputing with row mean. The results were bad. How are you doing in cases where the entire column is NaNs ?</p>",
          "rawMarkdown": "Hey, thanks for sharing. \n\nIn case where all values of a sensor column are missing, I tried imputing with row mean. The results were bad. How are you doing in cases where the entire column is NaNs ?"
        },
        {
          "id": 1110592,
          "postDate": "2020-12-12T22:21:48.993Z",
          "content": "<p>Are you working directly on the time series or on created features?</p>\n<p>If it is about time series, I don't know if there is a way to make a good estimate for impute when all data is missing. In the case of partial losses you could do mean, median or linear regression (<a href=\"https://www.kaggle.com/juejuewang/handle-missing-values-in-time-series-for-beginners\" target=\"_blank\">handle-missing-values-in-time-series-for-beginners</a>)</p>\n<p>If it is with created characteristics, it is possible to make an impute of a data using the mean or median of the data column, also something a little more complex is to estimate with random forest or knn.<br>\nIf all the data in a column are NA, you should check how you are calculating the feature and / or delete it. Because it is strange that a column of features created from a time series returns all values as NA.</p>",
          "rawMarkdown": "Are you working directly on the time series or on created features?\n\nIf it is about time series, I don't know if there is a way to make a good estimate for impute when all data is missing. In the case of partial losses you could do mean, median or linear regression ([handle-missing-values-in-time-series-for-beginners](https://www.kaggle.com/juejuewang/handle-missing-values-in-time-series-for-beginners))\n\nIf it is with created characteristics, it is possible to make an impute of a data using the mean or median of the data column, also something a little more complex is to estimate with random forest or knn.\nIf all the data in a column are NA, you should check how you are calculating the feature and / or delete it. Because it is strange that a column of features created from a time series returns all values as NA.",
          "votes": 1
        },
        {
          "id": 1110930,
          "postDate": "2020-12-13T08:02:39.203Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1113335,
          "postDate": "2020-12-15T11:29:55.580Z",
          "content": "<p>I am working on created features. </p>\n<p>And all values are NA in case of some segments of time series data, not the created features. I tried imputing the NAs in created features with mean and median but no improvement. </p>\n<p>My cv score is in the range of 2.5M but lb is 5M+.</p>",
          "rawMarkdown": "I am working on created features. \n\nAnd all values are NA in case of some segments of time series data, not the created features. I tried imputing the NAs in created features with mean and median but no improvement. \n\nMy cv score is in the range of 2.5M but lb is 5M+."
        },
        {
          "id": 1113879,
          "postDate": "2020-12-15T18:50:24.183Z",
          "content": "<p>The problem may not necessarily just be NAs, it may also be overfitting.</p>\n<p>As the training set is so small (approx. 4000 rows) you have to check very well how you are training, the sampling method, the outliers, etc.</p>\n<p>If it is the NAs that cause you the most problems, you would have to work without these observations, but losing information.</p>\n<p>In my case I am using impute to mean, plus impute outliers, this increased the performance a bit from 5.0M to 4.7M, but I still have a lot of overfitting.</p>\n<p>Now I am fixed at 4.7M and I have not been able to get out of there.</p>",
          "rawMarkdown": "The problem may not necessarily just be NAs, it may also be overfitting.\n\nAs the training set is so small (approx. 4000 rows) you have to check very well how you are training, the sampling method, the outliers, etc.\n\nIf it is the NAs that cause you the most problems, you would have to work without these observations, but losing information.\n\nIn my case I am using impute to mean, plus impute outliers, this increased the performance a bit from 5.0M to 4.7M, but I still have a lot of overfitting.\n\nNow I am fixed at 4.7M and I have not been able to get out of there."
        }
      ]
    },
    {
      "id": 1110004,
      "postDate": "2020-12-12T10:07:57.803Z",
      "content": "<p>In case of a few segment ids, all values of some sensors are missing, maybe due to sensor malfunction. I've tried imputing with single mean values of other sensors and row mean(axis=1). In both cases, the performance decreases in comparison to leaving NaNs as it is. </p>\n<p>How are you dealing with NaNs ? </p>",
      "rawMarkdown": "In case of a few segment ids, all values of some sensors are missing, maybe due to sensor malfunction. I've tried imputing with single mean values of other sensors and row mean(axis=1). In both cases, the performance decreases in comparison to leaving NaNs as it is. \n\nHow are you dealing with NaNs ? "
    }
  ],
  "comments": [
    {
      "id": 1110073,
      "author_name": "Desareca",
      "author_url": "",
      "post_date": "2020-12-12T11:59:24.397000",
      "content": "<p>hi, <br>\nI'm still trying to impute missing values ​​right.  I am working on the calculated characteristics and first I tried KNN, I had very bad results.  Then with median and improved a little.  now i am using the mean and it improved another bit.  I'm still trying to improve this though.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1110127,
          "author_name": "Harshit Gupta",
          "author_url": "",
          "post_date": "2020-12-12T13:01:08.027000",
          "content": "<p>Hey, thanks for sharing. </p>\n<p>In case where all values of a sensor column are missing, I tried imputing with row mean. The results were bad. How are you doing in cases where the entire column is NaNs ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1110592,
          "author_name": "Desareca",
          "author_url": "",
          "post_date": "2020-12-12T22:21:48.993000",
          "content": "<p>Are you working directly on the time series or on created features?</p>\n<p>If it is about time series, I don't know if there is a way to make a good estimate for impute when all data is missing. In the case of partial losses you could do mean, median or linear regression (<a href=\"https://www.kaggle.com/juejuewang/handle-missing-values-in-time-series-for-beginners\" target=\"_blank\">handle-missing-values-in-time-series-for-beginners</a>)</p>\n<p>If it is with created characteristics, it is possible to make an impute of a data using the mean or median of the data column, also something a little more complex is to estimate with random forest or knn.<br>\nIf all the data in a column are NA, you should check how you are calculating the feature and / or delete it. Because it is strange that a column of features created from a time series returns all values as NA.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1110930,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-13T08:02:39.203000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1113335,
          "author_name": "Harshit Gupta",
          "author_url": "",
          "post_date": "2020-12-15T11:29:55.580000",
          "content": "<p>I am working on created features. </p>\n<p>And all values are NA in case of some segments of time series data, not the created features. I tried imputing the NAs in created features with mean and median but no improvement. </p>\n<p>My cv score is in the range of 2.5M but lb is 5M+.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1113879,
          "author_name": "Desareca",
          "author_url": "",
          "post_date": "2020-12-15T18:50:24.183000",
          "content": "<p>The problem may not necessarily just be NAs, it may also be overfitting.</p>\n<p>As the training set is so small (approx. 4000 rows) you have to check very well how you are training, the sampling method, the outliers, etc.</p>\n<p>If it is the NAs that cause you the most problems, you would have to work without these observations, but losing information.</p>\n<p>In my case I am using impute to mean, plus impute outliers, this increased the performance a bit from 5.0M to 4.7M, but I still have a lot of overfitting.</p>\n<p>Now I am fixed at 4.7M and I have not been able to get out of there.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1110073": "hi, \nI'm still trying to impute missing values ​​right.  I am working on the calculated characteristics and first I tried KNN, I had very bad results.  Then with median and improved a little.  now i am using the mean and it improved another bit.  I'm still trying to improve this though.",
    "1110004": "In case of a few segment ids, all values of some sensors are missing, maybe due to sensor malfunction. I've tried imputing with single mean values of other sensors and row mean(axis=1). In both cases, the performance decreases in comparison to leaving NaNs as it is. \n\nHow are you dealing with NaNs ? "
  }
}