{
  "id": 535354,
  "title": "Some findings from the features EDA (+ data issues, outliers)",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/535354",
  "author_name": "",
  "post_date": "2024-09-21T17:03:42.608324100Z",
  "votes": 63,
  "comment_count": 19,
  "views": 0,
  "content": "<p>Hi there! Made some The EDA is in this notebook: <a href=\"https://www.kaggle.com/code/antoninadolgorukova/cmi-piu-features-eda\" target=\"_blank\">https://www.kaggle.com/code/antoninadolgorukova/cmi-piu-features-eda</a></p>\n<p>So I'd like to share my initial observations:</p>\n<p>Apparently, 40% of the participants were not affected by Internet use, 31% were not assessed, and only the minority (~10%) are moderately to severely impaired.</p>\n<p>The <strong>Internet usage</strong> data seems to be consistent with this, showing that half of the participants spend at least 1 hour or less per day on the Internet, and the maximum usage is 3 hours per day (as for me, there are no extreme Internet users in this train data set 😜).</p>\n<p>Each group of features in the data includes a \"<strong>season</strong>\" column, likely indicating when the data was collected or when participants joined the study. Seasonal changes might influence variables like fitness, physical activity, sleep patterns, and internet usage, but I haven't found any clear connections yet.</p>\n<p>I noticed that there are <strong>more males than females</strong> in most age groups, especially among the younger participants. This might help us uncover gender-specific differences in some of the variables.</p>\n<p>Now, the <strong>physical measurements</strong> seem a bit off. Many values are outside the normal range—for example, some participants weigh up to 142 kg, and there are unusual blood pressure readings. Surprisingly, a high BMI doesn't seem to be linked with high blood pressure, which is what I'd normally expect.</p>\n<p>Number of rows with values outside normal ranges:<br>\nPhysical-BMI: 2027 (67.23%)<br>\nPhysical-Height: 10 (0.33%)<br>\nPhysical-Weight: 165 (5.47%)<br>\nPhysical-Waist_Circumference: 93 (10.36%)<br>\nPhysical-Diastolic_BP: 1022 (34.61%)<br>\nPhysical-HeartRate: 350 (11.80%)<br>\nPhysical-Systolic_BP: 1078 (36.51%)</p>\n<p>From a rough estimate, about 55% of the participants are underweight, and around 11% have issues with being overweight. While most of the height and weight measurements seem reasonable, a whopping 67.23% have a BMI outside the normal range. This could mean that many participants have unusual body proportions, or maybe there were some measurement errors?</p>\n<p>The <strong>Physical Activity Questionnaire</strong> data is a bit confusing. It's divided between adolescents and children in a weird way. Participants with data in the children's columns (PAQ_C_Total) are aged 7 to 17, which overlaps with those in the adolescents' columns, who are aged 13 to 18.</p>\n<p>Endurance (<strong>FitnessGram Vitals and Treadmill</strong>) was only measured in kids aged 5 to 12, and it doesn't seem to change much with age. Interestingly, a few participants aged 7-8 showed exceptionally high endurance, while some didn't even complete the first stage (minimum score = 0).</p>\n<p>The <strong>FitnessGram Child</strong> data puzzles me too. There's overlap between the \"Healthy\" and \"Needs Improvement\" fitness zones, and these measurements were taken for participants aged 5 to 21. So why is it labeled \"Child\"? Maybe I'm misunderstanding what this data represents.</p>\n<p>Most of the <strong>bioelectrical impedance analysis</strong> data is highly skewed. The majority of participants have values at the extreme ends, with a few outliers that might be measurement errors. Some variables, like fat mass index and body fat percentage, even have implausibly negative values.</p>\n<p>Lastly, the <strong>Children's Global Assessment Scale</strong> has one outlier with a value of 999, which seems like a mistake.</p>\n<p>Of course, this is just my first look at the data (e.g. didn't yet examined actigraphy files), so I might find explanations for these oddities as I dig deeper. But I thought it would be interesting to discuss these findings with you all. If you have any thoughts or insights, I'd love to hear them 😊</p>",
  "messages": [
    {
      "id": "2994953",
      "postDate": "09/21/2024 17:03:42",
      "content": "<p>Hi there! Made some The EDA is in this notebook: <a href=\"https://www.kaggle.com/code/antoninadolgorukova/cmi-piu-features-eda\" target=\"_blank\">https://www.kaggle.com/code/antoninadolgorukova/cmi-piu-features-eda</a></p>\n<p>So I'd like to share my initial observations:</p>\n<p>Apparently, 40% of the participants were not affected by Internet use, 31% were not assessed, and only the minority (~10%) are moderately to severely impaired.</p>\n<p>The <strong>Internet usage</strong> data seems to be consistent with this, showing that half of the participants spend at least 1 hour or less per day on the Internet, and the maximum usage is 3 hours per day (as for me, there are no extreme Internet users in this train data set 😜).</p>\n<p>Each group of features in the data includes a \"<strong>season</strong>\" column, likely indicating when the data was collected or when participants joined the study. Seasonal changes might influence variables like fitness, physical activity, sleep patterns, and internet usage, but I haven't found any clear connections yet.</p>\n<p>I noticed that there are <strong>more males than females</strong> in most age groups, especially among the younger participants. This might help us uncover gender-specific differences in some of the variables.</p>\n<p>Now, the <strong>physical measurements</strong> seem a bit off. Many values are outside the normal range—for example, some participants weigh up to 142 kg, and there are unusual blood pressure readings. Surprisingly, a high BMI doesn't seem to be linked with high blood pressure, which is what I'd normally expect.</p>\n<p>Number of rows with values outside normal ranges:<br>\nPhysical-BMI: 2027 (67.23%)<br>\nPhysical-Height: 10 (0.33%)<br>\nPhysical-Weight: 165 (5.47%)<br>\nPhysical-Waist_Circumference: 93 (10.36%)<br>\nPhysical-Diastolic_BP: 1022 (34.61%)<br>\nPhysical-HeartRate: 350 (11.80%)<br>\nPhysical-Systolic_BP: 1078 (36.51%)</p>\n<p>From a rough estimate, about 55% of the participants are underweight, and around 11% have issues with being overweight. While most of the height and weight measurements seem reasonable, a whopping 67.23% have a BMI outside the normal range. This could mean that many participants have unusual body proportions, or maybe there were some measurement errors?</p>\n<p>The <strong>Physical Activity Questionnaire</strong> data is a bit confusing. It's divided between adolescents and children in a weird way. Participants with data in the children's columns (PAQ_C_Total) are aged 7 to 17, which overlaps with those in the adolescents' columns, who are aged 13 to 18.</p>\n<p>Endurance (<strong>FitnessGram Vitals and Treadmill</strong>) was only measured in kids aged 5 to 12, and it doesn't seem to change much with age. Interestingly, a few participants aged 7-8 showed exceptionally high endurance, while some didn't even complete the first stage (minimum score = 0).</p>\n<p>The <strong>FitnessGram Child</strong> data puzzles me too. There's overlap between the \"Healthy\" and \"Needs Improvement\" fitness zones, and these measurements were taken for participants aged 5 to 21. So why is it labeled \"Child\"? Maybe I'm misunderstanding what this data represents.</p>\n<p>Most of the <strong>bioelectrical impedance analysis</strong> data is highly skewed. The majority of participants have values at the extreme ends, with a few outliers that might be measurement errors. Some variables, like fat mass index and body fat percentage, even have implausibly negative values.</p>\n<p>Lastly, the <strong>Children's Global Assessment Scale</strong> has one outlier with a value of 999, which seems like a mistake.</p>\n<p>Of course, this is just my first look at the data (e.g. didn't yet examined actigraphy files), so I might find explanations for these oddities as I dig deeper. But I thought it would be interesting to discuss these findings with you all. If you have any thoughts or insights, I'd love to hear them 😊</p>",
      "rawMarkdown": "Hi there! Made some The EDA is in this notebook: https://www.kaggle.com/code/antoninadolgorukova/cmi-piu-features-eda\n\nSo I'd like to share my initial observations:\n\nApparently, 40% of the participants were not affected by Internet use, 31% were not assessed, and only the minority (~10%) are moderately to severely impaired.\n\nThe **Internet usage** data seems to be consistent with this, showing that half of the participants spend at least 1 hour or less per day on the Internet, and the maximum usage is 3 hours per day (as for me, there are no extreme Internet users in this train data set 😜).\n\nEach group of features in the data includes a \"**season**\" column, likely indicating when the data was collected or when participants joined the study. Seasonal changes might influence variables like fitness, physical activity, sleep patterns, and internet usage, but I haven't found any clear connections yet.\n\nI noticed that there are **more males than females** in most age groups, especially among the younger participants. This might help us uncover gender-specific differences in some of the variables.\n\nNow, the **physical measurements** seem a bit off. Many values are outside the normal range—for example, some participants weigh up to 142 kg, and there are unusual blood pressure readings. Surprisingly, a high BMI doesn't seem to be linked with high blood pressure, which is what I'd normally expect.\n\nNumber of rows with values outside normal ranges:\nPhysical-BMI: 2027 (67.23%)\nPhysical-Height: 10 (0.33%)\nPhysical-Weight: 165 (5.47%)\nPhysical-Waist_Circumference: 93 (10.36%)\nPhysical-Diastolic_BP: 1022 (34.61%)\nPhysical-HeartRate: 350 (11.80%)\nPhysical-Systolic_BP: 1078 (36.51%)\n\nFrom a rough estimate, about 55% of the participants are underweight, and around 11% have issues with being overweight. While most of the height and weight measurements seem reasonable, a whopping 67.23% have a BMI outside the normal range. This could mean that many participants have unusual body proportions, or maybe there were some measurement errors?\n\nThe **Physical Activity Questionnaire** data is a bit confusing. It's divided between adolescents and children in a weird way. Participants with data in the children's columns (PAQ_C_Total) are aged 7 to 17, which overlaps with those in the adolescents' columns, who are aged 13 to 18.\n\nEndurance (**FitnessGram Vitals and Treadmill**) was only measured in kids aged 5 to 12, and it doesn't seem to change much with age. Interestingly, a few participants aged 7-8 showed exceptionally high endurance, while some didn't even complete the first stage (minimum score = 0).\n\nThe **FitnessGram Child** data puzzles me too. There's overlap between the \"Healthy\" and \"Needs Improvement\" fitness zones, and these measurements were taken for participants aged 5 to 21. So why is it labeled \"Child\"? Maybe I'm misunderstanding what this data represents.\n\nMost of the **bioelectrical impedance analysis** data is highly skewed. The majority of participants have values at the extreme ends, with a few outliers that might be measurement errors. Some variables, like fat mass index and body fat percentage, even have implausibly negative values.\n\nLastly, the **Children's Global Assessment Scale** has one outlier with a value of 999, which seems like a mistake.\n\nOf course, this is just my first look at the data (e.g. didn't yet examined actigraphy files), so I might find explanations for these oddities as I dig deeper. But I thought it would be interesting to discuss these findings with you all. If you have any thoughts or insights, I'd love to hear them 😊",
      "votes": null
    },
    {
      "id": "2994962",
      "postDate": "09/21/2024 17:20:53",
      "content": "<blockquote>\n  <p>While most of the height and weight measurements seem reasonable, a whopping 67.23% have a BMI outside the normal range. This could mean that many participants have unusual body proportions, or maybe there were some measurement errors?</p>\n</blockquote>\n<p>BMI is one of the most useless metric I know of, so it makes total sense to me that 67,23% are being labeled as \"out of the ordinary\". All it does it literally taking the height of the person against his weight. It doesn't cover muscle percentage, fat percentage, fitness level, body proportions, blood cells - it doesn't cover virtually anything really. </p>\n<p>To give you an example: any olympic athlete that requires some form of strength - be it handball, rowing, spear throwing, whatever - will have an increased BMI and be mixed in with people that haven't done sports once in their life. I've had an \"extremely high BMI\" all of my life.</p>\n<p>So no, sadly I don't think it's wrongly calculated. It's just a really dense metric.</p>",
      "rawMarkdown": "> While most of the height and weight measurements seem reasonable, a whopping 67.23% have a BMI outside the normal range. This could mean that many participants have unusual body proportions, or maybe there were some measurement errors?\n\nBMI is one of the most useless metric I know of, so it makes total sense to me that 67,23% are being labeled as \"out of the ordinary\". All it does it literally taking the height of the person against his weight. It doesn't cover muscle percentage, fat percentage, fitness level, body proportions, blood cells - it doesn't cover virtually anything really. \n\nTo give you an example: any olympic athlete that requires some form of strength - be it handball, rowing, spear throwing, whatever - will have an increased BMI and be mixed in with people that haven't done sports once in their life. I've had an \"extremely high BMI\" all of my life.\n\nSo no, sadly I don't think it's wrongly calculated. It's just a really dense metric.",
      "votes": null
    },
    {
      "id": "2995061",
      "postDate": "09/21/2024 19:41:02",
      "content": "<p>I have also noticed discrepancies in the data, especially in the blood pressure columns. There are a few samples for which systolic BP are equal to or lower than the diastolic. Also, there seems to be a very high spread in the BP values. Discounting for the lower ones, there are a large number of very high BP entries. I wondered if there is any correlation between the heart rate as it may suggest that the high BP recordings were taken during exercise. There is some correlation, but one cannot be sure.</p>",
      "rawMarkdown": "I have also noticed discrepancies in the data, especially in the blood pressure columns. There are a few samples for which systolic BP are equal to or lower than the diastolic. Also, there seems to be a very high spread in the BP values. Discounting for the lower ones, there are a large number of very high BP entries. I wondered if there is any correlation between the heart rate as it may suggest that the high BP recordings were taken during exercise. There is some correlation, but one cannot be sure.",
      "votes": null
    },
    {
      "id": "2995199",
      "postDate": "09/22/2024 02:10:39",
      "content": "<p>indeed, now I see these 4 cases with SBP &lt;= DBP too </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12193662%2F32368c71d38c55edbf76bcf008d6aa8b%2FScreenshot%202024-09-22%20042305.png?generation=1726968204590596&amp;alt=media\" alt=\"now \"></p>\n<p>Also, I forgot to mention in the post, that there are other obvious artifacts  - the minimum values of 0 for measures like BMI, weight, and blood pressure.</p>\n<p>As for BP and HR, I think they were probably taken at rest, since there are no clear relationships.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12193662%2F0a63ffba213be3e7afa4831c21154005%2FScreenshot%202024-09-22%20050917.png?generation=1726970970260165&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "indeed, now I see these 4 cases with SBP <= DBP too \n\n![now ](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12193662%2F32368c71d38c55edbf76bcf008d6aa8b%2FScreenshot%202024-09-22%20042305.png?generation=1726968204590596&alt=media)\n\nAlso, I forgot to mention in the post, that there are other obvious artifacts  - the minimum values of 0 for measures like BMI, weight, and blood pressure.\n\nAs for BP and HR, I think they were probably taken at rest, since there are no clear relationships.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12193662%2F0a63ffba213be3e7afa4831c21154005%2FScreenshot%202024-09-22%20050917.png?generation=1726970970260165&alt=media)",
      "votes": null
    },
    {
      "id": "2995203",
      "postDate": "09/22/2024 02:30:56",
      "content": "<p>I think this is true for adults and athletes, where BMI may misclassify people with high muscle mass as overweight or obese. But most of the participants in this dataset are children or adolescents… My results are more likely due to the fact that I used reference values for adults, but had to use age-appropriate percentiles for children (e.g., see BMI-for-age growth charts on the CDC or WHO websites).</p>",
      "rawMarkdown": "I think this is true for adults and athletes, where BMI may misclassify people with high muscle mass as overweight or obese. But most of the participants in this dataset are children or adolescents... My results are more likely due to the fact that I used reference values for adults, but had to use age-appropriate percentiles for children (e.g., see BMI-for-age growth charts on the CDC or WHO websites).",
      "votes": null
    },
    {
      "id": "2997942",
      "postDate": "09/25/2024 03:53:00",
      "content": "<p>Absolutely correct. With datasets that include physical heights &amp; weights, I've found success dropping BMI all-together. In this particular competition, I saw it ranked low feature importance with Sklearn's RFE with Catboost. Unsurprising!</p>",
      "rawMarkdown": "Absolutely correct. With datasets that include physical heights & weights, I've found success dropping BMI all-together. In this particular competition, I saw it ranked low feature importance with Sklearn's RFE with Catboost. Unsurprising!",
      "votes": null
    },
    {
      "id": "2999156",
      "postDate": "09/26/2024 12:32:43",
      "content": "<p>Here is the actigraphy data EDA: <a href=\"https://www.kaggle.com/code/antoninadolgorukova/cmi-piu-actigraphy-data-eda\" target=\"_blank\">https://www.kaggle.com/code/antoninadolgorukova/cmi-piu-actigraphy-data-eda</a></p>\n<p>For now there are</p>\n<ul>\n<li>Main statistics for all participants (in general)</li>\n<li>No movement periods</li>\n<li>Circadian rhythm analysis</li>\n<li>Physical activity analysis</li>\n<li>Relationships between activity and light exposure</li>\n<li>Seasonal, weekly and daily trends</li>\n</ul>\n<p>I'd love to hear more ideas and hope to find time to update it and add more insights for feature engineering.</p>",
      "rawMarkdown": "Here is the actigraphy data EDA: https://www.kaggle.com/code/antoninadolgorukova/cmi-piu-actigraphy-data-eda\n\nFor now there are\n- Main statistics for all participants (in general)\n- No movement periods\n- Circadian rhythm analysis\n- Physical activity analysis\n- Relationships between activity and light exposure\n- Seasonal, weekly and daily trends\n\nI'd love to hear more ideas and hope to find time to update it and add more insights for feature engineering.",
      "votes": null
    },
    {
      "id": "3001151",
      "postDate": "09/28/2024 14:49:18",
      "content": "<p>Thanks to a nice catch by <a href=\"https://www.kaggle.com/siukeitin\" target=\"_blank\">@siukeitin</a> (sorry if sombody else also mentioned this, and i did miss it), we now also know that some of the questions can be ignored by a respondent (missing values in the PCIAT columns), but the SII score is still calculated as the sum of the non-NA values, leading to potentially invalid SII scores.</p>\n<p>I found 17 rows where the target variable was calculated incorrectly (ignoring missing responses) (just updated notebook with EDA). </p>\n<p>I suggest recalculating the SII in this way:</p>\n<pre><code> ():\n     pd.isna(row[]):\n         np.nan\n    max_possible = row[] + row[PCIAT_cols].isna().() * \n     row[] &lt;=   max_possible &lt;= :\n         \n      &lt;= row[] &lt;=   max_possible &lt;= :\n         \n      &lt;= row[] &lt;=   max_possible &lt;= :\n         \n     row[] &gt;=   max_possible &gt;= :\n         \n     np.nan\n</code></pre>",
      "rawMarkdown": "Thanks to a nice catch by @siukeitin (sorry if sombody else also mentioned this, and i did miss it), we now also know that some of the questions can be ignored by a respondent (missing values in the PCIAT columns), but the SII score is still calculated as the sum of the non-NA values, leading to potentially invalid SII scores.\n\nI found 17 rows where the target variable was calculated incorrectly (ignoring missing responses) (just updated notebook with EDA). \n\nI suggest recalculating the SII in this way:\n\n```python\ndef recalculate_sii(row):\n    if pd.isna(row['PCIAT-PCIAT_Total']):\n        return np.nan\n    max_possible = row['PCIAT-PCIAT_Total'] + row[PCIAT_cols].isna().sum() * 5\n    if row['PCIAT-PCIAT_Total'] <= 30 and max_possible <= 30:\n        return 0\n    elif 31 <= row['PCIAT-PCIAT_Total'] <= 49 and max_possible <= 49:\n        return 1\n    elif 50 <= row['PCIAT-PCIAT_Total'] <= 79 and max_possible <= 79:\n        return 2\n    elif row['PCIAT-PCIAT_Total'] >= 80 and max_possible >= 80:\n        return 3\n    return np.nan\n```",
      "votes": null
    },
    {
      "id": "3002034",
      "postDate": "09/29/2024 14:58:01",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/antoninadolgorukova\" target=\"_blank\">@antoninadolgorukova</a> - thanks for this informative notebook!<br>\nI think I'm going to select only data with battery_voltage &gt; 3800 ✔️😃</p>\n<p>I'm only 1/4 through it, so maybe the following is something you noticed already - sorry if so: Good you pointed out the many missing step/data value in <code>'0417c91e'</code> during the nights -- in comparison, the file <code>'d6cca65e'</code> does not have these dropped steps (I believe). So I wonder if this may be a processing or device mode to reduce data that is or is not present in different files?</p>\n<p>Another difference between these two data sets is the second one has very strange looking <code>'light'</code> values compared to the first, though maybe that's due to the dropped data points.</p>",
      "rawMarkdown": "Hi @antoninadolgorukova - thanks for this informative notebook!\nI think I'm going to select only data with battery_voltage > 3800 ✔️😃\n\nI'm only 1/4 through it, so maybe the following is something you noticed already - sorry if so: Good you pointed out the many missing step/data value in `'0417c91e'` during the nights -- in comparison, the file `'d6cca65e'` does not have these dropped steps (I believe). So I wonder if this may be a processing or device mode to reduce data that is or is not present in different files?\n\nAnother difference between these two data sets is the second one has very strange looking `'light'` values compared to the first, though maybe that's due to the dropped data points.",
      "votes": null
    },
    {
      "id": "3002094",
      "postDate": "09/29/2024 15:45:53",
      "content": "<p>Comments like these are very valuable to everyone, so please don't hesitate to share them (as I will)! I'm sure I've missed a lot, and I'm still adding to my analysis, so any insights are very welcome. </p>\n<p>Regarding the gaps in the data - I've been running this notebook interactively with files from different participants, and yes, the gaps aren't always there. Some have all the steps exactly every 5 seconds. It seems necessary to adjust the feature generation steps to handle both scenarios. </p>\n<p>I agree that the light exposure data seems to be the most unreliable, but I'm not quite sure how to address this yet…</p>",
      "rawMarkdown": "Comments like these are very valuable to everyone, so please don't hesitate to share them (as I will)! I'm sure I've missed a lot, and I'm still adding to my analysis, so any insights are very welcome. \n\nRegarding the gaps in the data - I've been running this notebook interactively with files from different participants, and yes, the gaps aren't always there. Some have all the steps exactly every 5 seconds. It seems necessary to adjust the feature generation steps to handle both scenarios. \n\nI agree that the light exposure data seems to be the most unreliable, but I'm not quite sure how to address this yet...",
      "votes": null
    },
    {
      "id": "3002096",
      "postDate": "09/29/2024 15:47:30",
      "content": "<p>If you don't mind me asking, why did you choose the 3800 mV threshold for battery_voltage?</p>",
      "rawMarkdown": "If you don't mind me asking, why did you choose the 3800 mV threshold for battery_voltage?",
      "votes": null
    },
    {
      "id": "3002379",
      "postDate": "09/30/2024 01:20:11",
      "content": "<p>It was based on your battery voltage graph with the data becoming strange at 35 days and 3700 mV -- I was thinking it was voltage related.  Since then, I saw your information about a low battery value of 3.1 V… and your histogram of min voltage… Maybe 3500 or 3200 or no limit is fine also 🙃</p>",
      "rawMarkdown": "It was based on your battery voltage graph with the data becoming strange at 35 days and 3700 mV -- I was thinking it was voltage related.  Since then, I saw your information about a low battery value of 3.1 V... and your histogram of min voltage... Maybe 3500 or 3200 or no limit is fine also 🙃",
      "votes": null
    },
    {
      "id": "3002848",
      "postDate": "09/30/2024 13:12:45",
      "content": "<p>P.S. I think the difference between these two types of files is whether:<br>\n<strong>non-wear_flag == 1 rows have been excluded (<code>0417c91e</code>, ~350 files) or included (<code>d6cca65e</code>)</strong>.<br>\nIt's too bad there is this difference in file contents, at least it is easy tell them apart.</p>",
      "rawMarkdown": "P.S. I think the difference between these two types of files is whether:\n**non-wear_flag == 1 rows have been excluded (`0417c91e`, ~350 files) or included (`d6cca65e`)**.\nIt's too bad there is this difference in file contents, at least it is easy tell them apart.",
      "votes": null
    },
    {
      "id": "3012869",
      "postDate": "10/09/2024 12:51:42",
      "content": "<p>This is really amazing. Thank you for contributing your analysis and insights into the data for public discourse. 😀</p>",
      "rawMarkdown": "This is really amazing. Thank you for contributing your analysis and insights into the data for public discourse. 😀",
      "votes": null
    },
    {
      "id": "3013006",
      "postDate": "10/09/2024 15:13:59",
      "content": "<p>But sii in the test set is based on the same summ so if there were missing PCIAT when test sii was calculated - we can't run away from these samples</p>",
      "rawMarkdown": "But sii in the test set is based on the same summ so if there were missing PCIAT when test sii was calculated - we can't run away from these samples",
      "votes": null
    },
    {
      "id": "3013035",
      "postDate": "10/09/2024 15:40:45",
      "content": "<p>Good point. I hadn't thought about the test as I'm not interested in modelling, so thanks for sharing! Idk, I think it's worth asking the hosts if such misleading sii exist in the test. There aren't many in the train, but it's another modelling complication that could be avoided if the data were properly prepared…</p>",
      "rawMarkdown": "Good point. I hadn't thought about the test as I'm not interested in modelling, so thanks for sharing! Idk, I think it's worth asking the hosts if such misleading sii exist in the test. There aren't many in the train, but it's another modelling complication that could be avoided if the data were properly prepared...",
      "votes": null
    },
    {
      "id": "3018873",
      "postDate": "10/16/2024 05:26:51",
      "content": "<p>Great job <a href=\"https://www.kaggle.com/antoninadolgorukova\" target=\"_blank\">@antoninadolgorukova</a>! I have some concerns about how we are categorizing the participants by age, specifically regarding the separation of children and adolescents. I found that the Children’s Global Assessment Scale (CGAS), which is detailed in this <a href=\"https://www.corc.uk.net/outcome-experience-measures/childrens-global-assessment-scale-cgas/\" target=\"_blank\">link</a> indicates that \"the Children’s Global Assessment Scale (CGAS), adapted from the Global Assessment Scale for adults, is a rating of general functioning for children and young people aged 4-16 years old.\" This means that this feature is intended for participants aged 4 to 16, while we also have data for those older than 16. This might be an error.<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20346534%2F94e9059e2fb1f2bb00e08edc61b574d6%2Fages_CGAS_score.png?generation=1729056359868115&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Great job @antoninadolgorukova! I have some concerns about how we are categorizing the participants by age, specifically regarding the separation of children and adolescents. I found that the Children’s Global Assessment Scale (CGAS), which is detailed in this [link](https://www.corc.uk.net/outcome-experience-measures/childrens-global-assessment-scale-cgas/) indicates that \"the Children’s Global Assessment Scale (CGAS), adapted from the Global Assessment Scale for adults, is a rating of general functioning for children and young people aged 4-16 years old.\" This means that this feature is intended for participants aged 4 to 16, while we also have data for those older than 16. This might be an error.![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20346534%2F94e9059e2fb1f2bb00e08edc61b574d6%2Fages_CGAS_score.png?generation=1729056359868115&alt=media)",
      "votes": null
    },
    {
      "id": "3024016",
      "postDate": "10/21/2024 08:48:04",
      "content": "<blockquote>\n  <p>Number of rows with values outside normal ranges:<br>\n  Physical-BMI: 2027 (67.23%)</p>\n</blockquote>\n<p>As <a href=\"https://www.kaggle.com/kevinbnisch\" target=\"_blank\">@kevinbnisch</a> already mentioned, BMI is not a very informative metric when it comes to health. I totally agree when looking at a single person. It can have many reasons why the BMI is not \"normal\". <br>\n But with this many values outside the \"normal\" range, I expect this to be some kind of selection bias of the population we are working with.<br>\nTherefore, I looked into the WHO standards for male and female BMI within the age groups 5-19 (data from <a href=\"https://www.who.int/tools/growth-reference-data-for-5to19-years/indicators/bmi-for-age\" target=\"_blank\">here</a>).</p>\n<p>According to the WHO deviations from the \"normal\" value (SD0 = 0 Standard Deviations from the mea) can be interpreted as follows:</p>\n<ul>\n<li>Overweight: &gt;SD1 (equivalent to BMI 25 kg/m2 at 19 years)</li>\n<li>Obesity: &gt;SD2 (equivalent to BMI 30 kg/m2 at 19 years)</li>\n<li>Thinness: &lt; SD2neg</li>\n</ul>\n<p>The population here is slightly overweight! I think this is way to consistent for explainaitions like \"they are very athletic\" or \"they have high bone density\". <br>\nIf I'm not mistaken the data comes from people from New York. And New York struggles with child obesity <a href=\"https://www.publichealth.columbia.edu/research/centers/columbia-center-childrens-environmental-health/our-research/health-effects/obesity\" target=\"_blank\">(here)</a>:</p>\n<blockquote>\n  <p>Recent studies of NYC children show that 15-19.4% of children are overweight and an additional 22-27% of children are obese.</p>\n</blockquote>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1599018%2F982bb03c028552aebc1475e37922e3b6%2Fbmi_who.png?generation=1729499077402610&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": ">Number of rows with values outside normal ranges:\nPhysical-BMI: 2027 (67.23%)\n\nAs @kevinbnisch already mentioned, BMI is not a very informative metric when it comes to health. I totally agree when looking at a single person. It can have many reasons why the BMI is not \"normal\". \n But with this many values outside the \"normal\" range, I expect this to be some kind of selection bias of the population we are working with.\nTherefore, I looked into the WHO standards for male and female BMI within the age groups 5-19 (data from [here](https://www.who.int/tools/growth-reference-data-for-5to19-years/indicators/bmi-for-age)).\n\nAccording to the WHO deviations from the \"normal\" value (SD0 = 0 Standard Deviations from the mea) can be interpreted as follows:\n- Overweight: >SD1 (equivalent to BMI 25 kg/m2 at 19 years)\n- Obesity: >SD2 (equivalent to BMI 30 kg/m2 at 19 years)\n- Thinness: < SD2neg\n\nThe population here is slightly overweight! I think this is way to consistent for explainaitions like \"they are very athletic\" or \"they have high bone density\". \nIf I'm not mistaken the data comes from people from New York. And New York struggles with child obesity [(here)](https://www.publichealth.columbia.edu/research/centers/columbia-center-childrens-environmental-health/our-research/health-effects/obesity):\n>Recent studies of NYC children show that 15-19.4% of children are overweight and an additional 22-27% of children are obese.\n\n\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1599018%2F982bb03c028552aebc1475e37922e3b6%2Fbmi_who.png?generation=1729499077402610&alt=media)",
      "votes": null
    },
    {
      "id": "3054500",
      "postDate": "11/24/2024 18:29:25",
      "content": "<p>What's your insight regarding the actigraphy data? Can any useful features be derived from that?</p>",
      "rawMarkdown": "What's your insight regarding the actigraphy data? Can any useful features be derived from that?",
      "votes": null
    },
    {
      "id": "3074934",
      "postDate": "12/18/2024 08:19:24",
      "content": "<p>It was helpful to be aware of the controversial issue surrounding BMI values.</p>",
      "rawMarkdown": "It was helpful to be aware of the controversial issue surrounding BMI values.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2994962,
      "author_name": "kevinbnisch",
      "author_url": "",
      "post_date": "09/21/2024 17:20:53",
      "content": "<blockquote>\n  <p>While most of the height and weight measurements seem reasonable, a whopping 67.23% have a BMI outside the normal range. This could mean that many participants have unusual body proportions, or maybe there were some measurement errors?</p>\n</blockquote>\n<p>BMI is one of the most useless metric I know of, so it makes total sense to me that 67,23% are being labeled as \"out of the ordinary\". All it does it literally taking the height of the person against his weight. It doesn't cover muscle percentage, fat percentage, fitness level, body proportions, blood cells - it doesn't cover virtually anything really. </p>\n<p>To give you an example: any olympic athlete that requires some form of strength - be it handball, rowing, spear throwing, whatever - will have an increased BMI and be mixed in with people that haven't done sports once in their life. I've had an \"extremely high BMI\" all of my life.</p>\n<p>So no, sadly I don't think it's wrongly calculated. It's just a really dense metric.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2995203,
          "author_name": "antoninadolgorukova",
          "author_url": "",
          "post_date": "09/22/2024 02:30:56",
          "content": "<p>I think this is true for adults and athletes, where BMI may misclassify people with high muscle mass as overweight or obese. But most of the participants in this dataset are children or adolescents… My results are more likely due to the fact that I used reference values for adults, but had to use age-appropriate percentiles for children (e.g., see BMI-for-age growth charts on the CDC or WHO websites).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2997942,
          "author_name": "exidekat",
          "author_url": "",
          "post_date": "09/25/2024 03:53:00",
          "content": "<p>Absolutely correct. With datasets that include physical heights &amp; weights, I've found success dropping BMI all-together. In this particular competition, I saw it ranked low feature importance with Sklearn's RFE with Catboost. Unsurprising!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2995061,
      "author_name": "rafidmahbub",
      "author_url": "",
      "post_date": "09/21/2024 19:41:02",
      "content": "<p>I have also noticed discrepancies in the data, especially in the blood pressure columns. There are a few samples for which systolic BP are equal to or lower than the diastolic. Also, there seems to be a very high spread in the BP values. Discounting for the lower ones, there are a large number of very high BP entries. I wondered if there is any correlation between the heart rate as it may suggest that the high BP recordings were taken during exercise. There is some correlation, but one cannot be sure.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2995199,
          "author_name": "antoninadolgorukova",
          "author_url": "",
          "post_date": "09/22/2024 02:10:39",
          "content": "<p>indeed, now I see these 4 cases with SBP &lt;= DBP too </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12193662%2F32368c71d38c55edbf76bcf008d6aa8b%2FScreenshot%202024-09-22%20042305.png?generation=1726968204590596&amp;alt=media\" alt=\"now \"></p>\n<p>Also, I forgot to mention in the post, that there are other obvious artifacts  - the minimum values of 0 for measures like BMI, weight, and blood pressure.</p>\n<p>As for BP and HR, I think they were probably taken at rest, since there are no clear relationships.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12193662%2F0a63ffba213be3e7afa4831c21154005%2FScreenshot%202024-09-22%20050917.png?generation=1726970970260165&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2999156,
      "author_name": "antoninadolgorukova",
      "author_url": "",
      "post_date": "09/26/2024 12:32:43",
      "content": "<p>Here is the actigraphy data EDA: <a href=\"https://www.kaggle.com/code/antoninadolgorukova/cmi-piu-actigraphy-data-eda\" target=\"_blank\">https://www.kaggle.com/code/antoninadolgorukova/cmi-piu-actigraphy-data-eda</a></p>\n<p>For now there are</p>\n<ul>\n<li>Main statistics for all participants (in general)</li>\n<li>No movement periods</li>\n<li>Circadian rhythm analysis</li>\n<li>Physical activity analysis</li>\n<li>Relationships between activity and light exposure</li>\n<li>Seasonal, weekly and daily trends</li>\n</ul>\n<p>I'd love to hear more ideas and hope to find time to update it and add more insights for feature engineering.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3002034,
          "author_name": "dan3dewey",
          "author_url": "",
          "post_date": "09/29/2024 14:58:01",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/antoninadolgorukova\" target=\"_blank\">@antoninadolgorukova</a> - thanks for this informative notebook!<br>\nI think I'm going to select only data with battery_voltage &gt; 3800 ✔️😃</p>\n<p>I'm only 1/4 through it, so maybe the following is something you noticed already - sorry if so: Good you pointed out the many missing step/data value in <code>'0417c91e'</code> during the nights -- in comparison, the file <code>'d6cca65e'</code> does not have these dropped steps (I believe). So I wonder if this may be a processing or device mode to reduce data that is or is not present in different files?</p>\n<p>Another difference between these two data sets is the second one has very strange looking <code>'light'</code> values compared to the first, though maybe that's due to the dropped data points.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3002094,
              "author_name": "antoninadolgorukova",
              "author_url": "",
              "post_date": "09/29/2024 15:45:53",
              "content": "<p>Comments like these are very valuable to everyone, so please don't hesitate to share them (as I will)! I'm sure I've missed a lot, and I'm still adding to my analysis, so any insights are very welcome. </p>\n<p>Regarding the gaps in the data - I've been running this notebook interactively with files from different participants, and yes, the gaps aren't always there. Some have all the steps exactly every 5 seconds. It seems necessary to adjust the feature generation steps to handle both scenarios. </p>\n<p>I agree that the light exposure data seems to be the most unreliable, but I'm not quite sure how to address this yet…</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 3002096,
              "author_name": "antoninadolgorukova",
              "author_url": "",
              "post_date": "09/29/2024 15:47:30",
              "content": "<p>If you don't mind me asking, why did you choose the 3800 mV threshold for battery_voltage?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3002379,
                  "author_name": "dan3dewey",
                  "author_url": "",
                  "post_date": "09/30/2024 01:20:11",
                  "content": "<p>It was based on your battery voltage graph with the data becoming strange at 35 days and 3700 mV -- I was thinking it was voltage related.  Since then, I saw your information about a low battery value of 3.1 V… and your histogram of min voltage… Maybe 3500 or 3200 or no limit is fine also 🙃</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            },
            {
              "id": 3002848,
              "author_name": "dan3dewey",
              "author_url": "",
              "post_date": "09/30/2024 13:12:45",
              "content": "<p>P.S. I think the difference between these two types of files is whether:<br>\n<strong>non-wear_flag == 1 rows have been excluded (<code>0417c91e</code>, ~350 files) or included (<code>d6cca65e</code>)</strong>.<br>\nIt's too bad there is this difference in file contents, at least it is easy tell them apart.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3001151,
      "author_name": "antoninadolgorukova",
      "author_url": "",
      "post_date": "09/28/2024 14:49:18",
      "content": "<p>Thanks to a nice catch by <a href=\"https://www.kaggle.com/siukeitin\" target=\"_blank\">@siukeitin</a> (sorry if sombody else also mentioned this, and i did miss it), we now also know that some of the questions can be ignored by a respondent (missing values in the PCIAT columns), but the SII score is still calculated as the sum of the non-NA values, leading to potentially invalid SII scores.</p>\n<p>I found 17 rows where the target variable was calculated incorrectly (ignoring missing responses) (just updated notebook with EDA). </p>\n<p>I suggest recalculating the SII in this way:</p>\n<pre><code> ():\n     pd.isna(row[]):\n         np.nan\n    max_possible = row[] + row[PCIAT_cols].isna().() * \n     row[] &lt;=   max_possible &lt;= :\n         \n      &lt;= row[] &lt;=   max_possible &lt;= :\n         \n      &lt;= row[] &lt;=   max_possible &lt;= :\n         \n     row[] &gt;=   max_possible &gt;= :\n         \n     np.nan\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 3013006,
          "author_name": "eu1234",
          "author_url": "",
          "post_date": "10/09/2024 15:13:59",
          "content": "<p>But sii in the test set is based on the same summ so if there were missing PCIAT when test sii was calculated - we can't run away from these samples</p>",
          "votes": null,
          "replies": [
            {
              "id": 3013035,
              "author_name": "antoninadolgorukova",
              "author_url": "",
              "post_date": "10/09/2024 15:40:45",
              "content": "<p>Good point. I hadn't thought about the test as I'm not interested in modelling, so thanks for sharing! Idk, I think it's worth asking the hosts if such misleading sii exist in the test. There aren't many in the train, but it's another modelling complication that could be avoided if the data were properly prepared…</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3012869,
      "author_name": "dehanc",
      "author_url": "",
      "post_date": "10/09/2024 12:51:42",
      "content": "<p>This is really amazing. Thank you for contributing your analysis and insights into the data for public discourse. 😀</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3018873,
      "author_name": "yannaktb",
      "author_url": "",
      "post_date": "10/16/2024 05:26:51",
      "content": "<p>Great job <a href=\"https://www.kaggle.com/antoninadolgorukova\" target=\"_blank\">@antoninadolgorukova</a>! I have some concerns about how we are categorizing the participants by age, specifically regarding the separation of children and adolescents. I found that the Children’s Global Assessment Scale (CGAS), which is detailed in this <a href=\"https://www.corc.uk.net/outcome-experience-measures/childrens-global-assessment-scale-cgas/\" target=\"_blank\">link</a> indicates that \"the Children’s Global Assessment Scale (CGAS), adapted from the Global Assessment Scale for adults, is a rating of general functioning for children and young people aged 4-16 years old.\" This means that this feature is intended for participants aged 4 to 16, while we also have data for those older than 16. This might be an error.<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20346534%2F94e9059e2fb1f2bb00e08edc61b574d6%2Fages_CGAS_score.png?generation=1729056359868115&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3024016,
      "author_name": "vulpesvulpes",
      "author_url": "",
      "post_date": "10/21/2024 08:48:04",
      "content": "<blockquote>\n  <p>Number of rows with values outside normal ranges:<br>\n  Physical-BMI: 2027 (67.23%)</p>\n</blockquote>\n<p>As <a href=\"https://www.kaggle.com/kevinbnisch\" target=\"_blank\">@kevinbnisch</a> already mentioned, BMI is not a very informative metric when it comes to health. I totally agree when looking at a single person. It can have many reasons why the BMI is not \"normal\". <br>\n But with this many values outside the \"normal\" range, I expect this to be some kind of selection bias of the population we are working with.<br>\nTherefore, I looked into the WHO standards for male and female BMI within the age groups 5-19 (data from <a href=\"https://www.who.int/tools/growth-reference-data-for-5to19-years/indicators/bmi-for-age\" target=\"_blank\">here</a>).</p>\n<p>According to the WHO deviations from the \"normal\" value (SD0 = 0 Standard Deviations from the mea) can be interpreted as follows:</p>\n<ul>\n<li>Overweight: &gt;SD1 (equivalent to BMI 25 kg/m2 at 19 years)</li>\n<li>Obesity: &gt;SD2 (equivalent to BMI 30 kg/m2 at 19 years)</li>\n<li>Thinness: &lt; SD2neg</li>\n</ul>\n<p>The population here is slightly overweight! I think this is way to consistent for explainaitions like \"they are very athletic\" or \"they have high bone density\". <br>\nIf I'm not mistaken the data comes from people from New York. And New York struggles with child obesity <a href=\"https://www.publichealth.columbia.edu/research/centers/columbia-center-childrens-environmental-health/our-research/health-effects/obesity\" target=\"_blank\">(here)</a>:</p>\n<blockquote>\n  <p>Recent studies of NYC children show that 15-19.4% of children are overweight and an additional 22-27% of children are obese.</p>\n</blockquote>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1599018%2F982bb03c028552aebc1475e37922e3b6%2Fbmi_who.png?generation=1729499077402610&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3054500,
      "author_name": "larryjiang07",
      "author_url": "",
      "post_date": "11/24/2024 18:29:25",
      "content": "<p>What's your insight regarding the actigraphy data? Can any useful features be derived from that?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3074934,
      "author_name": "honeywk",
      "author_url": "",
      "post_date": "12/18/2024 08:19:24",
      "content": "<p>It was helpful to be aware of the controversial issue surrounding BMI values.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2994953": "Hi there! Made some The EDA is in this notebook: https://www.kaggle.com/code/antoninadolgorukova/cmi-piu-features-eda\n\nSo I'd like to share my initial observations:\n\nApparently, 40% of the participants were not affected by Internet use, 31% were not assessed, and only the minority (~10%) are moderately to severely impaired.\n\nThe **Internet usage** data seems to be consistent with this, showing that half of the participants spend at least 1 hour or less per day on the Internet, and the maximum usage is 3 hours per day (as for me, there are no extreme Internet users in this train data set 😜).\n\nEach group of features in the data includes a \"**season**\" column, likely indicating when the data was collected or when participants joined the study. Seasonal changes might influence variables like fitness, physical activity, sleep patterns, and internet usage, but I haven't found any clear connections yet.\n\nI noticed that there are **more males than females** in most age groups, especially among the younger participants. This might help us uncover gender-specific differences in some of the variables.\n\nNow, the **physical measurements** seem a bit off. Many values are outside the normal range—for example, some participants weigh up to 142 kg, and there are unusual blood pressure readings. Surprisingly, a high BMI doesn't seem to be linked with high blood pressure, which is what I'd normally expect.\n\nNumber of rows with values outside normal ranges:\nPhysical-BMI: 2027 (67.23%)\nPhysical-Height: 10 (0.33%)\nPhysical-Weight: 165 (5.47%)\nPhysical-Waist_Circumference: 93 (10.36%)\nPhysical-Diastolic_BP: 1022 (34.61%)\nPhysical-HeartRate: 350 (11.80%)\nPhysical-Systolic_BP: 1078 (36.51%)\n\nFrom a rough estimate, about 55% of the participants are underweight, and around 11% have issues with being overweight. While most of the height and weight measurements seem reasonable, a whopping 67.23% have a BMI outside the normal range. This could mean that many participants have unusual body proportions, or maybe there were some measurement errors?\n\nThe **Physical Activity Questionnaire** data is a bit confusing. It's divided between adolescents and children in a weird way. Participants with data in the children's columns (PAQ_C_Total) are aged 7 to 17, which overlaps with those in the adolescents' columns, who are aged 13 to 18.\n\nEndurance (**FitnessGram Vitals and Treadmill**) was only measured in kids aged 5 to 12, and it doesn't seem to change much with age. Interestingly, a few participants aged 7-8 showed exceptionally high endurance, while some didn't even complete the first stage (minimum score = 0).\n\nThe **FitnessGram Child** data puzzles me too. There's overlap between the \"Healthy\" and \"Needs Improvement\" fitness zones, and these measurements were taken for participants aged 5 to 21. So why is it labeled \"Child\"? Maybe I'm misunderstanding what this data represents.\n\nMost of the **bioelectrical impedance analysis** data is highly skewed. The majority of participants have values at the extreme ends, with a few outliers that might be measurement errors. Some variables, like fat mass index and body fat percentage, even have implausibly negative values.\n\nLastly, the **Children's Global Assessment Scale** has one outlier with a value of 999, which seems like a mistake.\n\nOf course, this is just my first look at the data (e.g. didn't yet examined actigraphy files), so I might find explanations for these oddities as I dig deeper. But I thought it would be interesting to discuss these findings with you all. If you have any thoughts or insights, I'd love to hear them 😊",
    "2994962": "> While most of the height and weight measurements seem reasonable, a whopping 67.23% have a BMI outside the normal range. This could mean that many participants have unusual body proportions, or maybe there were some measurement errors?\n\nBMI is one of the most useless metric I know of, so it makes total sense to me that 67,23% are being labeled as \"out of the ordinary\". All it does it literally taking the height of the person against his weight. It doesn't cover muscle percentage, fat percentage, fitness level, body proportions, blood cells - it doesn't cover virtually anything really. \n\nTo give you an example: any olympic athlete that requires some form of strength - be it handball, rowing, spear throwing, whatever - will have an increased BMI and be mixed in with people that haven't done sports once in their life. I've had an \"extremely high BMI\" all of my life.\n\nSo no, sadly I don't think it's wrongly calculated. It's just a really dense metric.",
    "2995061": "I have also noticed discrepancies in the data, especially in the blood pressure columns. There are a few samples for which systolic BP are equal to or lower than the diastolic. Also, there seems to be a very high spread in the BP values. Discounting for the lower ones, there are a large number of very high BP entries. I wondered if there is any correlation between the heart rate as it may suggest that the high BP recordings were taken during exercise. There is some correlation, but one cannot be sure.",
    "2995199": "indeed, now I see these 4 cases with SBP <= DBP too \n\n![now ](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12193662%2F32368c71d38c55edbf76bcf008d6aa8b%2FScreenshot%202024-09-22%20042305.png?generation=1726968204590596&alt=media)\n\nAlso, I forgot to mention in the post, that there are other obvious artifacts  - the minimum values of 0 for measures like BMI, weight, and blood pressure.\n\nAs for BP and HR, I think they were probably taken at rest, since there are no clear relationships.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12193662%2F0a63ffba213be3e7afa4831c21154005%2FScreenshot%202024-09-22%20050917.png?generation=1726970970260165&alt=media)",
    "2995203": "I think this is true for adults and athletes, where BMI may misclassify people with high muscle mass as overweight or obese. But most of the participants in this dataset are children or adolescents... My results are more likely due to the fact that I used reference values for adults, but had to use age-appropriate percentiles for children (e.g., see BMI-for-age growth charts on the CDC or WHO websites).",
    "2997942": "Absolutely correct. With datasets that include physical heights & weights, I've found success dropping BMI all-together. In this particular competition, I saw it ranked low feature importance with Sklearn's RFE with Catboost. Unsurprising!",
    "2999156": "Here is the actigraphy data EDA: https://www.kaggle.com/code/antoninadolgorukova/cmi-piu-actigraphy-data-eda\n\nFor now there are\n- Main statistics for all participants (in general)\n- No movement periods\n- Circadian rhythm analysis\n- Physical activity analysis\n- Relationships between activity and light exposure\n- Seasonal, weekly and daily trends\n\nI'd love to hear more ideas and hope to find time to update it and add more insights for feature engineering.",
    "3001151": "Thanks to a nice catch by @siukeitin (sorry if sombody else also mentioned this, and i did miss it), we now also know that some of the questions can be ignored by a respondent (missing values in the PCIAT columns), but the SII score is still calculated as the sum of the non-NA values, leading to potentially invalid SII scores.\n\nI found 17 rows where the target variable was calculated incorrectly (ignoring missing responses) (just updated notebook with EDA). \n\nI suggest recalculating the SII in this way:\n\n```python\ndef recalculate_sii(row):\n    if pd.isna(row['PCIAT-PCIAT_Total']):\n        return np.nan\n    max_possible = row['PCIAT-PCIAT_Total'] + row[PCIAT_cols].isna().sum() * 5\n    if row['PCIAT-PCIAT_Total'] <= 30 and max_possible <= 30:\n        return 0\n    elif 31 <= row['PCIAT-PCIAT_Total'] <= 49 and max_possible <= 49:\n        return 1\n    elif 50 <= row['PCIAT-PCIAT_Total'] <= 79 and max_possible <= 79:\n        return 2\n    elif row['PCIAT-PCIAT_Total'] >= 80 and max_possible >= 80:\n        return 3\n    return np.nan\n```",
    "3002034": "Hi @antoninadolgorukova - thanks for this informative notebook!\nI think I'm going to select only data with battery_voltage > 3800 ✔️😃\n\nI'm only 1/4 through it, so maybe the following is something you noticed already - sorry if so: Good you pointed out the many missing step/data value in `'0417c91e'` during the nights -- in comparison, the file `'d6cca65e'` does not have these dropped steps (I believe). So I wonder if this may be a processing or device mode to reduce data that is or is not present in different files?\n\nAnother difference between these two data sets is the second one has very strange looking `'light'` values compared to the first, though maybe that's due to the dropped data points.",
    "3002094": "Comments like these are very valuable to everyone, so please don't hesitate to share them (as I will)! I'm sure I've missed a lot, and I'm still adding to my analysis, so any insights are very welcome. \n\nRegarding the gaps in the data - I've been running this notebook interactively with files from different participants, and yes, the gaps aren't always there. Some have all the steps exactly every 5 seconds. It seems necessary to adjust the feature generation steps to handle both scenarios. \n\nI agree that the light exposure data seems to be the most unreliable, but I'm not quite sure how to address this yet...",
    "3002096": "If you don't mind me asking, why did you choose the 3800 mV threshold for battery_voltage?",
    "3002379": "It was based on your battery voltage graph with the data becoming strange at 35 days and 3700 mV -- I was thinking it was voltage related.  Since then, I saw your information about a low battery value of 3.1 V... and your histogram of min voltage... Maybe 3500 or 3200 or no limit is fine also 🙃",
    "3002848": "P.S. I think the difference between these two types of files is whether:\n**non-wear_flag == 1 rows have been excluded (`0417c91e`, ~350 files) or included (`d6cca65e`)**.\nIt's too bad there is this difference in file contents, at least it is easy tell them apart.",
    "3012869": "This is really amazing. Thank you for contributing your analysis and insights into the data for public discourse. 😀",
    "3013006": "But sii in the test set is based on the same summ so if there were missing PCIAT when test sii was calculated - we can't run away from these samples",
    "3013035": "Good point. I hadn't thought about the test as I'm not interested in modelling, so thanks for sharing! Idk, I think it's worth asking the hosts if such misleading sii exist in the test. There aren't many in the train, but it's another modelling complication that could be avoided if the data were properly prepared...",
    "3018873": "Great job @antoninadolgorukova! I have some concerns about how we are categorizing the participants by age, specifically regarding the separation of children and adolescents. I found that the Children’s Global Assessment Scale (CGAS), which is detailed in this [link](https://www.corc.uk.net/outcome-experience-measures/childrens-global-assessment-scale-cgas/) indicates that \"the Children’s Global Assessment Scale (CGAS), adapted from the Global Assessment Scale for adults, is a rating of general functioning for children and young people aged 4-16 years old.\" This means that this feature is intended for participants aged 4 to 16, while we also have data for those older than 16. This might be an error.![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20346534%2F94e9059e2fb1f2bb00e08edc61b574d6%2Fages_CGAS_score.png?generation=1729056359868115&alt=media)",
    "3024016": ">Number of rows with values outside normal ranges:\nPhysical-BMI: 2027 (67.23%)\n\nAs @kevinbnisch already mentioned, BMI is not a very informative metric when it comes to health. I totally agree when looking at a single person. It can have many reasons why the BMI is not \"normal\". \n But with this many values outside the \"normal\" range, I expect this to be some kind of selection bias of the population we are working with.\nTherefore, I looked into the WHO standards for male and female BMI within the age groups 5-19 (data from [here](https://www.who.int/tools/growth-reference-data-for-5to19-years/indicators/bmi-for-age)).\n\nAccording to the WHO deviations from the \"normal\" value (SD0 = 0 Standard Deviations from the mea) can be interpreted as follows:\n- Overweight: >SD1 (equivalent to BMI 25 kg/m2 at 19 years)\n- Obesity: >SD2 (equivalent to BMI 30 kg/m2 at 19 years)\n- Thinness: < SD2neg\n\nThe population here is slightly overweight! I think this is way to consistent for explainaitions like \"they are very athletic\" or \"they have high bone density\". \nIf I'm not mistaken the data comes from people from New York. And New York struggles with child obesity [(here)](https://www.publichealth.columbia.edu/research/centers/columbia-center-childrens-environmental-health/our-research/health-effects/obesity):\n>Recent studies of NYC children show that 15-19.4% of children are overweight and an additional 22-27% of children are obese.\n\n\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1599018%2F982bb03c028552aebc1475e37922e3b6%2Fbmi_who.png?generation=1729499077402610&alt=media)",
    "3054500": "What's your insight regarding the actigraphy data? Can any useful features be derived from that?",
    "3074934": "It was helpful to be aware of the controversial issue surrounding BMI values."
  },
  "source": "meta"
}