{
  "id": 478577,
  "title": "Seeking Guidance: Handling External Data Sources",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/478577",
  "author_name": "",
  "post_date": "2024-02-21T11:14:46.116264400Z",
  "votes": 3,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hello everyone,</p>\n<p>I started working with the Home Credit dataset recently, and I'm still relatively new to Kaggle and data science. I've come across a dilemma regarding the handling of external data sources. The data description mentions that some external data providers might not be available for future evaluations, which is expected. Additionally, it has been confirmed in a forum post that the whole table could be missing after a certain date, although it's also mentioned that empty tables would still be retained if all data are missing in the test sample.</p>\n<p>While exploring public notebooks for guidance on dealing with external data, I've noticed that many include external data in their processing pipelines, such as aggregating data from the static_cb_0 dataset to create the training dataset. Handling missing values seems to be addressed by filling them with column mean/mode.</p>\n<p>However, I'm somewhat concerned about this approach. I'm worried that if a feature from an external data source becomes the dominant predictor for the trained model, and the private test data doesn't include these features, the model's performance could suffer.</p>\n<p>So, what's the best approach to tackle this issue? Should we train multiple models and assess feature importance? If features from external data have high importance for a particular model, should we consider dropping that model from the final inference? Alternatively, would it be better to use external data solely to gain insights for better engineering of internal data, without directly incorporating it as a feature for the model or ensemble? Or perhaps my concern is invalid because the features from external data aren't influential enough? Or is it better to consider downweighting features incorporated from external data during model training?</p>\n<p>I'd appreciate any insights or advice on how to handle this situation effectively. Thank you!</p>\n<p>Aforementioned forum post: <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476231\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476231</a></p>",
  "messages": [
    {
      "id": "2661590",
      "postDate": "02/21/2024 11:14:46",
      "content": "<p>Hello everyone,</p>\n<p>I started working with the Home Credit dataset recently, and I'm still relatively new to Kaggle and data science. I've come across a dilemma regarding the handling of external data sources. The data description mentions that some external data providers might not be available for future evaluations, which is expected. Additionally, it has been confirmed in a forum post that the whole table could be missing after a certain date, although it's also mentioned that empty tables would still be retained if all data are missing in the test sample.</p>\n<p>While exploring public notebooks for guidance on dealing with external data, I've noticed that many include external data in their processing pipelines, such as aggregating data from the static_cb_0 dataset to create the training dataset. Handling missing values seems to be addressed by filling them with column mean/mode.</p>\n<p>However, I'm somewhat concerned about this approach. I'm worried that if a feature from an external data source becomes the dominant predictor for the trained model, and the private test data doesn't include these features, the model's performance could suffer.</p>\n<p>So, what's the best approach to tackle this issue? Should we train multiple models and assess feature importance? If features from external data have high importance for a particular model, should we consider dropping that model from the final inference? Alternatively, would it be better to use external data solely to gain insights for better engineering of internal data, without directly incorporating it as a feature for the model or ensemble? Or perhaps my concern is invalid because the features from external data aren't influential enough? Or is it better to consider downweighting features incorporated from external data during model training?</p>\n<p>I'd appreciate any insights or advice on how to handle this situation effectively. Thank you!</p>\n<p>Aforementioned forum post: <a href=\"https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476231\" target=\"_blank\">https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476231</a></p>",
      "rawMarkdown": "Hello everyone,\n\nI started working with the Home Credit dataset recently, and I'm still relatively new to Kaggle and data science. I've come across a dilemma regarding the handling of external data sources. The data description mentions that some external data providers might not be available for future evaluations, which is expected. Additionally, it has been confirmed in a forum post that the whole table could be missing after a certain date, although it's also mentioned that empty tables would still be retained if all data are missing in the test sample.\n\nWhile exploring public notebooks for guidance on dealing with external data, I've noticed that many include external data in their processing pipelines, such as aggregating data from the static_cb_0 dataset to create the training dataset. Handling missing values seems to be addressed by filling them with column mean/mode.\n\nHowever, I'm somewhat concerned about this approach. I'm worried that if a feature from an external data source becomes the dominant predictor for the trained model, and the private test data doesn't include these features, the model's performance could suffer.\n\nSo, what's the best approach to tackle this issue? Should we train multiple models and assess feature importance? If features from external data have high importance for a particular model, should we consider dropping that model from the final inference? Alternatively, would it be better to use external data solely to gain insights for better engineering of internal data, without directly incorporating it as a feature for the model or ensemble? Or perhaps my concern is invalid because the features from external data aren't influential enough? Or is it better to consider downweighting features incorporated from external data during model training?\n\nI'd appreciate any insights or advice on how to handle this situation effectively. Thank you!\n\nAforementioned forum post: https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476231",
      "votes": null
    },
    {
      "id": "2661625",
      "postDate": "02/21/2024 11:29:31",
      "content": "<p><a href=\"https://www.kaggle.com/pelinkeskin\" target=\"_blank\">@pelinkeskin</a> I am reasonably sure that data for certain dates may not be available but I think all columns should be available across the datasets. </p>",
      "rawMarkdown": "pelinkeskin I am reasonably sure that data for certain dates may not be available but I think all columns should be available across the datasets.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2661625,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "02/21/2024 11:29:31",
      "content": "<p><a href=\"https://www.kaggle.com/pelinkeskin\" target=\"_blank\">@pelinkeskin</a> I am reasonably sure that data for certain dates may not be available but I think all columns should be available across the datasets. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2661590": "Hello everyone,\n\nI started working with the Home Credit dataset recently, and I'm still relatively new to Kaggle and data science. I've come across a dilemma regarding the handling of external data sources. The data description mentions that some external data providers might not be available for future evaluations, which is expected. Additionally, it has been confirmed in a forum post that the whole table could be missing after a certain date, although it's also mentioned that empty tables would still be retained if all data are missing in the test sample.\n\nWhile exploring public notebooks for guidance on dealing with external data, I've noticed that many include external data in their processing pipelines, such as aggregating data from the static_cb_0 dataset to create the training dataset. Handling missing values seems to be addressed by filling them with column mean/mode.\n\nHowever, I'm somewhat concerned about this approach. I'm worried that if a feature from an external data source becomes the dominant predictor for the trained model, and the private test data doesn't include these features, the model's performance could suffer.\n\nSo, what's the best approach to tackle this issue? Should we train multiple models and assess feature importance? If features from external data have high importance for a particular model, should we consider dropping that model from the final inference? Alternatively, would it be better to use external data solely to gain insights for better engineering of internal data, without directly incorporating it as a feature for the model or ensemble? Or perhaps my concern is invalid because the features from external data aren't influential enough? Or is it better to consider downweighting features incorporated from external data during model training?\n\nI'd appreciate any insights or advice on how to handle this situation effectively. Thank you!\n\nAforementioned forum post: https://www.kaggle.com/competitions/home-credit-credit-risk-model-stability/discussion/476231",
    "2661625": "pelinkeskin I am reasonably sure that data for certain dates may not be available but I think all columns should be available across the datasets."
  },
  "source": "meta"
}