{
  "id": 535321,
  "title": "Key Strategies to deal with missing values in data ",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/535321",
  "author_name": "",
  "post_date": "2024-09-21T12:49:53.866704700Z",
  "votes": 8,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Dear colleagues, as you know that there is a lot of missing values, as per my observations and insight into the data, following approach may be adapted to deal with missing data. </p>\n<ul>\n<li>Focus should be on finding missing values for target columns starting with PCIAT that are actually the base for inferring sii - target variable </li>\n<li>Clustering or dimensionality reduction approach may be used to investigate the data where the target sii is missing – a kind of unsupervised approach </li>\n<li>Labeled and unlabeled data may be combined to train a model and infer missing sii values - semi-supervised learning</li>\n<li>Estimate the missing target values and/or missing features using various imputation techniques - </li>\n<li>After having inferred estimated missing labels, apply traditional supervised learning to the fully labeled data.</li>\n</ul>\n<p>Your comments please….</p>",
  "messages": [
    {
      "id": "2994773",
      "postDate": "09/21/2024 12:49:53",
      "content": "<p>Dear colleagues, as you know that there is a lot of missing values, as per my observations and insight into the data, following approach may be adapted to deal with missing data. </p>\n<ul>\n<li>Focus should be on finding missing values for target columns starting with PCIAT that are actually the base for inferring sii - target variable </li>\n<li>Clustering or dimensionality reduction approach may be used to investigate the data where the target sii is missing – a kind of unsupervised approach </li>\n<li>Labeled and unlabeled data may be combined to train a model and infer missing sii values - semi-supervised learning</li>\n<li>Estimate the missing target values and/or missing features using various imputation techniques - </li>\n<li>After having inferred estimated missing labels, apply traditional supervised learning to the fully labeled data.</li>\n</ul>\n<p>Your comments please….</p>",
      "rawMarkdown": "Dear colleagues, as you know that there is a lot of missing values, as per my observations and insight into the data, following approach may be adapted to deal with missing data. \n\n- Focus should be on finding missing values for target columns starting with PCIAT that are actually the base for inferring sii - target variable \n- Clustering or dimensionality reduction approach may be used to investigate the data where the target sii is missing – a kind of unsupervised approach \n- Labeled and unlabeled data may be combined to train a model and infer missing sii values - semi-supervised learning\n- Estimate the missing target values and/or missing features using various imputation techniques - \n- After having inferred estimated missing labels, apply traditional supervised learning to the fully labeled data.\n\nYour comments please....",
      "votes": null
    },
    {
      "id": "2995300",
      "postDate": "09/22/2024 06:35:37",
      "content": "<p>good idea </p>",
      "rawMarkdown": "good idea",
      "votes": null
    },
    {
      "id": "2997785",
      "postDate": "09/24/2024 22:32:51",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/tariqcp\" target=\"_blank\">@tariqcp</a> , Your plan is <strong>Great</strong> for handling missing data and making accurate predictions.</p>\n<p><strong>Suggestions</strong></p>\n<ol>\n<li>Evaluate Imputation Methods: Test different imputation techniques like KNN or MICE to find the best fit for your dataset.</li>\n<li>Use Ensemble Methods: Consider using ensemble models like Random Forests or Gradient Boosting for robust predictions.</li>\n</ol>",
      "rawMarkdown": "Hi @tariqcp , Your plan is **Great** for handling missing data and making accurate predictions.\n\n**Suggestions**\n1. Evaluate Imputation Methods: Test different imputation techniques like KNN or MICE to find the best fit for your dataset.\n2. Use Ensemble Methods: Consider using ensemble models like Random Forests or Gradient Boosting for robust predictions.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2995300,
      "author_name": "samanfatima7",
      "author_url": "",
      "post_date": "09/22/2024 06:35:37",
      "content": "<p>good idea </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2997785,
      "author_name": "salaheddineelkhirani",
      "author_url": "",
      "post_date": "09/24/2024 22:32:51",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/tariqcp\" target=\"_blank\">@tariqcp</a> , Your plan is <strong>Great</strong> for handling missing data and making accurate predictions.</p>\n<p><strong>Suggestions</strong></p>\n<ol>\n<li>Evaluate Imputation Methods: Test different imputation techniques like KNN or MICE to find the best fit for your dataset.</li>\n<li>Use Ensemble Methods: Consider using ensemble models like Random Forests or Gradient Boosting for robust predictions.</li>\n</ol>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2994773": "Dear colleagues, as you know that there is a lot of missing values, as per my observations and insight into the data, following approach may be adapted to deal with missing data. \n\n- Focus should be on finding missing values for target columns starting with PCIAT that are actually the base for inferring sii - target variable \n- Clustering or dimensionality reduction approach may be used to investigate the data where the target sii is missing – a kind of unsupervised approach \n- Labeled and unlabeled data may be combined to train a model and infer missing sii values - semi-supervised learning\n- Estimate the missing target values and/or missing features using various imputation techniques - \n- After having inferred estimated missing labels, apply traditional supervised learning to the fully labeled data.\n\nYour comments please....",
    "2995300": "good idea",
    "2997785": "Hi @tariqcp , Your plan is **Great** for handling missing data and making accurate predictions.\n\n**Suggestions**\n1. Evaluate Imputation Methods: Test different imputation techniques like KNN or MICE to find the best fit for your dataset.\n2. Use Ensemble Methods: Consider using ensemble models like Random Forests or Gradient Boosting for robust predictions."
  },
  "source": "meta"
}