{
  "id": 535388,
  "title": "Irregularities in BIA category values",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/535388",
  "author_name": "",
  "post_date": "2024-09-21T21:51:06.895334500Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi all,</p>\n<p>While performing EDA, I noticed irregularities many of the subcategories under BIA (bio-electrical impedance analysis). For reference, I have attached descriptive statistics associated with these features.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9669614%2Fce88d6e7e0f332e2e5ee361b06dcf249%2FScreenshot%202024-09-21%20at%2017.43.57.png?generation=1726955060883658&amp;alt=media\" alt=\"\"></p>\n<p>For example, BMC and BMR both have very large max values. Fortunately though, these irregularities are not too frequent. Can anyone recommend references where ranges of these values can be found?</p>\n<p>Thanks</p>",
  "messages": [
    {
      "id": "2995107",
      "postDate": "09/21/2024 21:51:06",
      "content": "<p>Hi all,</p>\n<p>While performing EDA, I noticed irregularities many of the subcategories under BIA (bio-electrical impedance analysis). For reference, I have attached descriptive statistics associated with these features.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9669614%2Fce88d6e7e0f332e2e5ee361b06dcf249%2FScreenshot%202024-09-21%20at%2017.43.57.png?generation=1726955060883658&amp;alt=media\" alt=\"\"></p>\n<p>For example, BMC and BMR both have very large max values. Fortunately though, these irregularities are not too frequent. Can anyone recommend references where ranges of these values can be found?</p>\n<p>Thanks</p>",
      "rawMarkdown": "Hi all,\n\nWhile performing EDA, I noticed irregularities many of the subcategories under BIA (bio-electrical impedance analysis). For reference, I have attached descriptive statistics associated with these features.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9669614%2Fce88d6e7e0f332e2e5ee361b06dcf249%2FScreenshot%202024-09-21%20at%2017.43.57.png?generation=1726955060883658&alt=media)\n\nFor example, BMC and BMR both have very large max values. Fortunately though, these irregularities are not too frequent. Can anyone recommend references where ranges of these values can be found?\n\nThanks",
      "votes": null
    },
    {
      "id": "3002380",
      "postDate": "09/30/2024 01:27:14",
      "content": "<p>It seems many of the BIA data have very unusual outliers, which could be ignored or handled in different ways. For example, I chose to clip these variables for now:</p>\n<pre><code>\n\n[] = np.clip(train[], ., .)\n\n[] = np.clip(train[], ., np.inf)\n\n[] = np.clip(train[], ., .)\n</code></pre>",
      "rawMarkdown": "It seems many of the BIA data have very unusual outliers, which could be ignored or handled in different ways. For example, I chose to clip these variables for now:\n```\n# Most BIAs have outlier issues...\n# Clip:\ntrain[\"BIA-BIA_FFMI\"] = np.clip(train[\"BIA-BIA_FFMI\"], 0.0, 25.0)\n# Clip:\ntrain[\"BIA-BIA_FMI\"] = np.clip(train[\"BIA-BIA_FMI\"], 0.0, np.inf)\n# Clip:\ntrain[\"BIA-BIA_LST\"] = np.clip(train[\"BIA-BIA_LST\"], 0.0, 250.0)\n```",
      "votes": null
    },
    {
      "id": "3002386",
      "postDate": "09/30/2024 01:52:33",
      "content": "<p>😀It seems that many BIA data have very unusual outliers that could be ignored or treated differently.</p>",
      "rawMarkdown": "😀It seems that many BIA data have very unusual outliers that could be ignored or treated differently.",
      "votes": null
    },
    {
      "id": "3006913",
      "postDate": "10/04/2024 17:03:38",
      "content": "<p>I found that the FMI and Fat columns having negative values is extremely unusual - so I have just replaced them with missing values.</p>",
      "rawMarkdown": "I found that the FMI and Fat columns having negative values is extremely unusual - so I have just replaced them with missing values.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3002380,
      "author_name": "dan3dewey",
      "author_url": "",
      "post_date": "09/30/2024 01:27:14",
      "content": "<p>It seems many of the BIA data have very unusual outliers, which could be ignored or handled in different ways. For example, I chose to clip these variables for now:</p>\n<pre><code>\n\n[] = np.clip(train[], ., .)\n\n[] = np.clip(train[], ., np.inf)\n\n[] = np.clip(train[], ., .)\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3002386,
      "author_name": "quincyyyy",
      "author_url": "",
      "post_date": "09/30/2024 01:52:33",
      "content": "<p>😀It seems that many BIA data have very unusual outliers that could be ignored or treated differently.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3006913,
      "author_name": "peterhopkinson",
      "author_url": "",
      "post_date": "10/04/2024 17:03:38",
      "content": "<p>I found that the FMI and Fat columns having negative values is extremely unusual - so I have just replaced them with missing values.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2995107": "Hi all,\n\nWhile performing EDA, I noticed irregularities many of the subcategories under BIA (bio-electrical impedance analysis). For reference, I have attached descriptive statistics associated with these features.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F9669614%2Fce88d6e7e0f332e2e5ee361b06dcf249%2FScreenshot%202024-09-21%20at%2017.43.57.png?generation=1726955060883658&alt=media)\n\nFor example, BMC and BMR both have very large max values. Fortunately though, these irregularities are not too frequent. Can anyone recommend references where ranges of these values can be found?\n\nThanks",
    "3002380": "It seems many of the BIA data have very unusual outliers, which could be ignored or handled in different ways. For example, I chose to clip these variables for now:\n```\n# Most BIAs have outlier issues...\n# Clip:\ntrain[\"BIA-BIA_FFMI\"] = np.clip(train[\"BIA-BIA_FFMI\"], 0.0, 25.0)\n# Clip:\ntrain[\"BIA-BIA_FMI\"] = np.clip(train[\"BIA-BIA_FMI\"], 0.0, np.inf)\n# Clip:\ntrain[\"BIA-BIA_LST\"] = np.clip(train[\"BIA-BIA_LST\"], 0.0, 250.0)\n```",
    "3002386": "😀It seems that many BIA data have very unusual outliers that could be ignored or treated differently.",
    "3006913": "I found that the FMI and Fat columns having negative values is extremely unusual - so I have just replaced them with missing values."
  },
  "source": "meta"
}