{
  "id": 476776,
  "title": "What mean NAN exactly in finance ?",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/476776",
  "author_name": "",
  "post_date": "2024-02-13T13:20:06.243170100Z",
  "votes": null,
  "comment_count": 10,
  "views": 0,
  "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16503833%2F79bc025680753dba06e31b3609268ddb%2FScreenshot%202024-02-13%20184716.png?generation=1707830378205068&amp;alt=media\"></p>\n<p>How to deal with this ? Should us drop or replace ?</p>",
  "messages": [
    {
      "id": "2650464",
      "postDate": "02/13/2024 13:20:06",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16503833%2F79bc025680753dba06e31b3609268ddb%2FScreenshot%202024-02-13%20184716.png?generation=1707830378205068&amp;alt=media\"></p>\n<p>How to deal with this ? Should us drop or replace ?</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16503833%2F79bc025680753dba06e31b3609268ddb%2FScreenshot%202024-02-13%20184716.png?generation=1707830378205068&alt=media)\n\nHow to deal with this ? Should us drop or replace ?",
      "votes": null
    },
    {
      "id": "2650495",
      "postDate": "02/13/2024 13:40:03",
      "content": "<p>Those values were not found in the database for the combination of case_id and the column. Basically, the value was not set in our database. Dropping them would not be considered wise as the fact that the value was not set contains some information value as well. </p>",
      "rawMarkdown": "Those values were not found in the database for the combination of case_id and the column. Basically, the value was not set in our database. Dropping them would not be considered wise as the fact that the value was not set contains some information value as well.",
      "votes": null
    },
    {
      "id": "2650524",
      "postDate": "02/13/2024 13:59:45",
      "content": "<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> Thank you so much sir . But replacing with some replacing value can bias data and model on a side . There are high chances of a bias data , models may be will not perform well . Can I do if 75% of rows not have any value , then drop ? I accept that it will miss value , but we have a large data , and it will not make models and data biased ? What is your opinion ? I am not in this field so i am asking you if there is another method . </p>",
      "rawMarkdown": "jetakow Thank you so much sir . But replacing with some replacing value can bias data and model on a side . There are high chances of a bias data , models may be will not perform well . Can I do if 75% of rows not have any value , then drop ? I accept that it will miss value , but we have a large data , and it will not make models and data biased ? What is your opinion ? I am not in this field so i am asking you if there is another method .",
      "votes": null
    },
    {
      "id": "2650720",
      "postDate": "02/13/2024 17:02:45",
      "content": "<p>I think ultimately the decision is up to you. You can work with the data however you like. What I do at my work is that I try to leverage nan values instead of dropping them. You can use for example optbinning package with target mean encoding and a high number of bins. On the other hand, you will break the multivariate interactions, so there is a tradeoff. </p>",
      "rawMarkdown": "I think ultimately the decision is up to you. You can work with the data however you like. What I do at my work is that I try to leverage nan values instead of dropping them. You can use for example optbinning package with target mean encoding and a high number of bins. On the other hand, you will break the multivariate interactions, so there is a tradeoff.",
      "votes": null
    },
    {
      "id": "2650730",
      "postDate": "02/13/2024 17:09:55",
      "content": "<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> Thank you sir !! I have another question . how did you handle large datasets ? I am currently doing memory management for it . Do you have any other way ? basically I get out of memory . </p>\n<p>sir can we connect on linkedin ?</p>",
      "rawMarkdown": "jetakow Thank you sir !! I have another question . how did you handle large datasets ? I am currently doing memory management for it . Do you have any other way ? basically I get out of memory . \n\nsir can we connect on linkedin ?",
      "votes": null
    },
    {
      "id": "2650738",
      "postDate": "02/13/2024 17:16:46",
      "content": "<p>For large datasets, I suggest using polars and lazy evaluation. </p>\n<p>Sure, you can add me on LinkedIn, but I must warn you that all kaggle competition topics shall remain here in Discussion section. I won't provide any additional information regarding this competition outside kaggle website. </p>",
      "rawMarkdown": "For large datasets, I suggest using polars and lazy evaluation. \n\nSure, you can add me on LinkedIn, but I must warn you that all kaggle competition topics shall remain here in Discussion section. I won't provide any additional information regarding this competition outside kaggle website.",
      "votes": null
    },
    {
      "id": "2650744",
      "postDate": "02/13/2024 17:20:30",
      "content": "<p>Yes sir . I am following rules . I am a student , so I wish to ask you questions about careers and practices . Your way of explaining impressed me alot . Thanks for it . </p>",
      "rawMarkdown": "Yes sir . I am following rules . I am a student , so I wish to ask you questions about careers and practices . Your way of explaining impressed me alot . Thanks for it .",
      "votes": null
    },
    {
      "id": "2651019",
      "postDate": "02/13/2024 20:21:58",
      "content": "<p><a href=\"https://www.kaggle.com/ayushkhaire\" target=\"_blank\">@ayushkhaire</a> , It simply means that a value is missing. It needs to be filled in depending on the specific attribute. Sometimes it can be a median or mode, sometimes zero, but there are other options (<a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.impute.KNNImputer.html)\" target=\"_blank\">https://scikit-learn.org/stable/modules/generated/sklearn.impute.KNNImputer.html)</a>.</p>",
      "rawMarkdown": "ayushkhaire , It simply means that a value is missing. It needs to be filled in depending on the specific attribute. Sometimes it can be a median or mode, sometimes zero, but there are other options (https://scikit-learn.org/stable/modules/generated/sklearn.impute.KNNImputer.html).",
      "votes": null
    },
    {
      "id": "2651188",
      "postDate": "02/14/2024 02:51:42",
      "content": "<p><a href=\"https://www.kaggle.com/bratkovskyevgeny\" target=\"_blank\">@bratkovskyevgeny</a> thanks a lot sir.  My question was about handling it , actually there are a lot of nans.  If we replace,  our data will become more towards biased , and will definitely affect model performance.  If we drop , we have a some \" sample \" of data at the end , and we can consider it as a CV and train data . What do you think? In replacement,  if we train a model,  then it will have same values , it will repeat gradually,  and it can not able to handle the data of outstanding test data points.  I think dropping rows having 75% nan values , make a sample , and then in that sample,  we will do replacement . Is any other method in this case ?</p>",
      "rawMarkdown": "bratkovskyevgeny thanks a lot sir.  My question was about handling it , actually there are a lot of nans.  If we replace,  our data will become more towards biased , and will definitely affect model performance.  If we drop , we have a some \" sample \" of data at the end , and we can consider it as a CV and train data . What do you think? In replacement,  if we train a model,  then it will have same values , it will repeat gradually,  and it can not able to handle the data of outstanding test data points.  I think dropping rows having 75% nan values , make a sample , and then in that sample,  we will do replacement . Is any other method in this case ?",
      "votes": null
    },
    {
      "id": "2651765",
      "postDate": "02/14/2024 11:11:11",
      "content": "<p>1) In finance a nan often means 0, but this should be deducted on a per-feature basis. <br>\n2) Nowadays gbdts models are able ot handle nans as-is.</p>",
      "rawMarkdown": "1) In finance a nan often means 0, but this should be deducted on a per-feature basis. \n2) Nowadays gbdts models are able ot handle nans as-is.",
      "votes": null
    },
    {
      "id": "2651823",
      "postDate": "02/14/2024 11:47:11",
      "content": "<p><a href=\"https://www.kaggle.com/lucasmorin\" target=\"_blank\">@lucasmorin</a>  thanks a lot sir , so should I replace the nans ?</p>",
      "rawMarkdown": "lucasmorin  thanks a lot sir , so should I replace the nans ?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2650495,
      "author_name": "jetakow",
      "author_url": "",
      "post_date": "02/13/2024 13:40:03",
      "content": "<p>Those values were not found in the database for the combination of case_id and the column. Basically, the value was not set in our database. Dropping them would not be considered wise as the fact that the value was not set contains some information value as well. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2650524,
          "author_name": "ayushkhaire",
          "author_url": "",
          "post_date": "02/13/2024 13:59:45",
          "content": "<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> Thank you so much sir . But replacing with some replacing value can bias data and model on a side . There are high chances of a bias data , models may be will not perform well . Can I do if 75% of rows not have any value , then drop ? I accept that it will miss value , but we have a large data , and it will not make models and data biased ? What is your opinion ? I am not in this field so i am asking you if there is another method . </p>",
          "votes": null,
          "replies": [
            {
              "id": 2650720,
              "author_name": "jetakow",
              "author_url": "",
              "post_date": "02/13/2024 17:02:45",
              "content": "<p>I think ultimately the decision is up to you. You can work with the data however you like. What I do at my work is that I try to leverage nan values instead of dropping them. You can use for example optbinning package with target mean encoding and a high number of bins. On the other hand, you will break the multivariate interactions, so there is a tradeoff. </p>",
              "votes": null,
              "replies": [
                {
                  "id": 2650730,
                  "author_name": "ayushkhaire",
                  "author_url": "",
                  "post_date": "02/13/2024 17:09:55",
                  "content": "<p><a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> Thank you sir !! I have another question . how did you handle large datasets ? I am currently doing memory management for it . Do you have any other way ? basically I get out of memory . </p>\n<p>sir can we connect on linkedin ?</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2650738,
                      "author_name": "jetakow",
                      "author_url": "",
                      "post_date": "02/13/2024 17:16:46",
                      "content": "<p>For large datasets, I suggest using polars and lazy evaluation. </p>\n<p>Sure, you can add me on LinkedIn, but I must warn you that all kaggle competition topics shall remain here in Discussion section. I won't provide any additional information regarding this competition outside kaggle website. </p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2650744,
                          "author_name": "ayushkhaire",
                          "author_url": "",
                          "post_date": "02/13/2024 17:20:30",
                          "content": "<p>Yes sir . I am following rules . I am a student , so I wish to ask you questions about careers and practices . Your way of explaining impressed me alot . Thanks for it . </p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2651019,
      "author_name": "bratkovskyevgeny",
      "author_url": "",
      "post_date": "02/13/2024 20:21:58",
      "content": "<p><a href=\"https://www.kaggle.com/ayushkhaire\" target=\"_blank\">@ayushkhaire</a> , It simply means that a value is missing. It needs to be filled in depending on the specific attribute. Sometimes it can be a median or mode, sometimes zero, but there are other options (<a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.impute.KNNImputer.html)\" target=\"_blank\">https://scikit-learn.org/stable/modules/generated/sklearn.impute.KNNImputer.html)</a>.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2651188,
          "author_name": "ayushkhaire",
          "author_url": "",
          "post_date": "02/14/2024 02:51:42",
          "content": "<p><a href=\"https://www.kaggle.com/bratkovskyevgeny\" target=\"_blank\">@bratkovskyevgeny</a> thanks a lot sir.  My question was about handling it , actually there are a lot of nans.  If we replace,  our data will become more towards biased , and will definitely affect model performance.  If we drop , we have a some \" sample \" of data at the end , and we can consider it as a CV and train data . What do you think? In replacement,  if we train a model,  then it will have same values , it will repeat gradually,  and it can not able to handle the data of outstanding test data points.  I think dropping rows having 75% nan values , make a sample , and then in that sample,  we will do replacement . Is any other method in this case ?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2651765,
      "author_name": "lucasmorin",
      "author_url": "",
      "post_date": "02/14/2024 11:11:11",
      "content": "<p>1) In finance a nan often means 0, but this should be deducted on a per-feature basis. <br>\n2) Nowadays gbdts models are able ot handle nans as-is.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2651823,
          "author_name": "ayushkhaire",
          "author_url": "",
          "post_date": "02/14/2024 11:47:11",
          "content": "<p><a href=\"https://www.kaggle.com/lucasmorin\" target=\"_blank\">@lucasmorin</a>  thanks a lot sir , so should I replace the nans ?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2650464": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16503833%2F79bc025680753dba06e31b3609268ddb%2FScreenshot%202024-02-13%20184716.png?generation=1707830378205068&alt=media)\n\nHow to deal with this ? Should us drop or replace ?",
    "2650495": "Those values were not found in the database for the combination of case_id and the column. Basically, the value was not set in our database. Dropping them would not be considered wise as the fact that the value was not set contains some information value as well.",
    "2650524": "jetakow Thank you so much sir . But replacing with some replacing value can bias data and model on a side . There are high chances of a bias data , models may be will not perform well . Can I do if 75% of rows not have any value , then drop ? I accept that it will miss value , but we have a large data , and it will not make models and data biased ? What is your opinion ? I am not in this field so i am asking you if there is another method .",
    "2650720": "I think ultimately the decision is up to you. You can work with the data however you like. What I do at my work is that I try to leverage nan values instead of dropping them. You can use for example optbinning package with target mean encoding and a high number of bins. On the other hand, you will break the multivariate interactions, so there is a tradeoff.",
    "2650730": "jetakow Thank you sir !! I have another question . how did you handle large datasets ? I am currently doing memory management for it . Do you have any other way ? basically I get out of memory . \n\nsir can we connect on linkedin ?",
    "2650738": "For large datasets, I suggest using polars and lazy evaluation. \n\nSure, you can add me on LinkedIn, but I must warn you that all kaggle competition topics shall remain here in Discussion section. I won't provide any additional information regarding this competition outside kaggle website.",
    "2650744": "Yes sir . I am following rules . I am a student , so I wish to ask you questions about careers and practices . Your way of explaining impressed me alot . Thanks for it .",
    "2651019": "ayushkhaire , It simply means that a value is missing. It needs to be filled in depending on the specific attribute. Sometimes it can be a median or mode, sometimes zero, but there are other options (https://scikit-learn.org/stable/modules/generated/sklearn.impute.KNNImputer.html).",
    "2651188": "bratkovskyevgeny thanks a lot sir.  My question was about handling it , actually there are a lot of nans.  If we replace,  our data will become more towards biased , and will definitely affect model performance.  If we drop , we have a some \" sample \" of data at the end , and we can consider it as a CV and train data . What do you think? In replacement,  if we train a model,  then it will have same values , it will repeat gradually,  and it can not able to handle the data of outstanding test data points.  I think dropping rows having 75% nan values , make a sample , and then in that sample,  we will do replacement . Is any other method in this case ?",
    "2651765": "1) In finance a nan often means 0, but this should be deducted on a per-feature basis. \n2) Nowadays gbdts models are able ot handle nans as-is.",
    "2651823": "lucasmorin  thanks a lot sir , so should I replace the nans ?"
  },
  "source": "meta"
}