{
  "id": 535525,
  "title": "Questioning the task",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/535525",
  "author_name": "Antonina Dolgorukova",
  "post_date": "2024-09-22T15:05:30.712000",
  "votes": 87,
  "comment_count": 56,
  "views": 0,
  "content": "<p>Topic aim: to understand the real aim of this competition and the practical implications of the model required. If you are struggling with this like me, please feel free to add your questions. If you can prove me wrong, then please do so!</p>\n<p><strong>TL;TR: please see the summary (based on  the original post + comments + questions from other threads) at the end</strong></p>\n<p>Is a PIU prediction model redundant with Internet use as an input?</p>\n<p>Problematic Internet use (PIU) is defined as the use of the Internet that creates psychological, social, school and/or work difficulties in a person's life (thanks <a href=\"https://www.kaggle.com/mpwolke\" target=\"_blank\">@mpwolke</a> for the <a href=\"https://www.sciencedirect.com/science/article/abs/pii/S0165178119320098\" target=\"_blank\">article</a>).</p>\n<p><strong>So you can't have PIU without excessive Internet use and without health or social problems.</strong><br>\nI wonder what exactly is the need for a Severity Impairment Index (SII) prediction model that relies on features such as \"Internet usage\" (hours spent online per day)? If we need to collect internet usage and health-related data to make predictions, then we're already measuring what we're trying to predict, or what I'm not seeing here?</p>\n<p>In practical terms:</p>\n<ul>\n<li><p>For patients with high Internet use and health problems: We can already suspect problematic Internet use without needing a model, and include in the treatment plan not only medications for the health condition(s), but also recommendations for reducing Internet use and healthy lifestyles.</p></li>\n<li><p>For patients with health problems but low Internet use: We can infer that their health problems have other causes and focus on diagnosing and treating them, again without the help of the model. Similarly, patients with low physical activity and unhealthy lifestyles need advice on both.</p></li>\n</ul>\n<p>Of course, self-reported Internet use can be inaccurate, but we don't have any other characteristics related to Internet use to train a model. So if we can't rely on reported Internet use, we can't properly diagnose PIU.</p>\n<p>Furthermore, any associations found between Internet use and health problems do not prove causation. The root cause may be poor health leading to increased internet use rather than the internet use causing health problems. </p>\n<p>I'm curious to hear your perspectives on this. Let's discuss! </p>\n<hr>\n<p>Another puzzling observation:</p>\n<p>The competition description states that the goal is \"to detect early indicators of problematic Internet and technology use.\"</p>\n<p>Given that, as noted in the comments below, PIU reflect the negative consequences of Internet use (when Internet use begins to cause problems), it seems logical to study early signs of PIU in a population of participants who use the Internet more than average.</p>\n<p>This would mean recognizing subtle shifts in behavior, physical activity, or psychological well-being that could signal the onset of problematic internet use before it leads to significant impairment.</p>\n<p>At the same time, 38.5% of the participants in the train dataset  use internet less than hour a day. Moreover, a noticable proportion of participants with high SII scores  use internet less than hour a day (21.55% of participants with SII = 2, moderately impaired  by problematic internet use, and 14.7% of those with SII = 3, severely impaired)!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12193662%2Fbc48c089630fb2973007d1501260640a%2FScreenshot%202024-09-23%20080435.png?generation=1727067903360763&amp;alt=media\" alt=\"\"></p>\n<p><a href=\"https://www.kaggle.com/code/antoninadolgorukova/cmi-piu-features-eda\" target=\"_blank\">More details are here</a></p>\n<p>If the index is intended to measure problematic internet use, it shouldn’t produce high scores for participants who  spend so little time online, should it? So is this investigator bias, unreliable self-reporting, or data collection error?</p>\n<p>And what is the point of including such a large proportion of people who  spend &lt;1h/day in the internet when the goal is to detect early signs of harmfull internet usage?</p>\n<p>In my opinion, Internet usage data should not be used as a feature for this task, but rather as a condition that participants must meet in order to be included in the study. In this way, we could develop a model to predict whether these individuals show signs of impairment and how severe that impairment is. This would be more consistent with the goals of detecting problematic Internet use.</p>\n<hr>\n<p><strong>Questioning summary:</strong></p>\n<ul>\n<li><p>Data collection for the model involves professional assessments, questionnaires and specialised equipment, adding complexity rather than simplifying the process for families.</p></li>\n<li><p>Clinicians with CGAS scores, physical health measures and internet use data have sufficient information to diagnose and assess the severity of PIU in one visit. They could also simply administer the PCIAT during assessments to obtain more accurate SII scores than a predictive model would provide (which also requires much more data and time to collect).</p></li>\n<li><p>As <a href=\"https://www.kaggle.com/expensivelunch\" target=\"_blank\">@expensivelunch</a> mentioned <a href=\"https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/538082\" target=\"_blank\">here</a>, most actigraphy data were collected after the target outcome had been measured, representing future states. This introduces an inherent inaccuracy into the model.</p></li>\n</ul>\n<p>These are not rhetorical questions, so I would really appreciate any comments from hosts  <a href=\"https://www.kaggle.com/gkiar07\" target=\"_blank\">@gkiar07</a> or other interested people.</p>",
  "messages": [
    {
      "id": 2995763,
      "postDate": "2024-09-22T15:05:30.713Z",
      "content": "<p>Topic aim: to understand the real aim of this competition and the practical implications of the model required. If you are struggling with this like me, please feel free to add your questions. If you can prove me wrong, then please do so!</p>\n<p><strong>TL;TR: please see the summary (based on  the original post + comments + questions from other threads) at the end</strong></p>\n<p>Is a PIU prediction model redundant with Internet use as an input?</p>\n<p>Problematic Internet use (PIU) is defined as the use of the Internet that creates psychological, social, school and/or work difficulties in a person's life (thanks <a href=\"https://www.kaggle.com/mpwolke\" target=\"_blank\">@mpwolke</a> for the <a href=\"https://www.sciencedirect.com/science/article/abs/pii/S0165178119320098\" target=\"_blank\">article</a>).</p>\n<p><strong>So you can't have PIU without excessive Internet use and without health or social problems.</strong><br>\nI wonder what exactly is the need for a Severity Impairment Index (SII) prediction model that relies on features such as \"Internet usage\" (hours spent online per day)? If we need to collect internet usage and health-related data to make predictions, then we're already measuring what we're trying to predict, or what I'm not seeing here?</p>\n<p>In practical terms:</p>\n<ul>\n<li><p>For patients with high Internet use and health problems: We can already suspect problematic Internet use without needing a model, and include in the treatment plan not only medications for the health condition(s), but also recommendations for reducing Internet use and healthy lifestyles.</p></li>\n<li><p>For patients with health problems but low Internet use: We can infer that their health problems have other causes and focus on diagnosing and treating them, again without the help of the model. Similarly, patients with low physical activity and unhealthy lifestyles need advice on both.</p></li>\n</ul>\n<p>Of course, self-reported Internet use can be inaccurate, but we don't have any other characteristics related to Internet use to train a model. So if we can't rely on reported Internet use, we can't properly diagnose PIU.</p>\n<p>Furthermore, any associations found between Internet use and health problems do not prove causation. The root cause may be poor health leading to increased internet use rather than the internet use causing health problems. </p>\n<p>I'm curious to hear your perspectives on this. Let's discuss! </p>\n<hr>\n<p>Another puzzling observation:</p>\n<p>The competition description states that the goal is \"to detect early indicators of problematic Internet and technology use.\"</p>\n<p>Given that, as noted in the comments below, PIU reflect the negative consequences of Internet use (when Internet use begins to cause problems), it seems logical to study early signs of PIU in a population of participants who use the Internet more than average.</p>\n<p>This would mean recognizing subtle shifts in behavior, physical activity, or psychological well-being that could signal the onset of problematic internet use before it leads to significant impairment.</p>\n<p>At the same time, 38.5% of the participants in the train dataset  use internet less than hour a day. Moreover, a noticable proportion of participants with high SII scores  use internet less than hour a day (21.55% of participants with SII = 2, moderately impaired  by problematic internet use, and 14.7% of those with SII = 3, severely impaired)!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12193662%2Fbc48c089630fb2973007d1501260640a%2FScreenshot%202024-09-23%20080435.png?generation=1727067903360763&amp;alt=media\" alt=\"\"></p>\n<p><a href=\"https://www.kaggle.com/code/antoninadolgorukova/cmi-piu-features-eda\" target=\"_blank\">More details are here</a></p>\n<p>If the index is intended to measure problematic internet use, it shouldn’t produce high scores for participants who  spend so little time online, should it? So is this investigator bias, unreliable self-reporting, or data collection error?</p>\n<p>And what is the point of including such a large proportion of people who  spend &lt;1h/day in the internet when the goal is to detect early signs of harmfull internet usage?</p>\n<p>In my opinion, Internet usage data should not be used as a feature for this task, but rather as a condition that participants must meet in order to be included in the study. In this way, we could develop a model to predict whether these individuals show signs of impairment and how severe that impairment is. This would be more consistent with the goals of detecting problematic Internet use.</p>\n<hr>\n<p><strong>Questioning summary:</strong></p>\n<ul>\n<li><p>Data collection for the model involves professional assessments, questionnaires and specialised equipment, adding complexity rather than simplifying the process for families.</p></li>\n<li><p>Clinicians with CGAS scores, physical health measures and internet use data have sufficient information to diagnose and assess the severity of PIU in one visit. They could also simply administer the PCIAT during assessments to obtain more accurate SII scores than a predictive model would provide (which also requires much more data and time to collect).</p></li>\n<li><p>As <a href=\"https://www.kaggle.com/expensivelunch\" target=\"_blank\">@expensivelunch</a> mentioned <a href=\"https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/538082\" target=\"_blank\">here</a>, most actigraphy data were collected after the target outcome had been measured, representing future states. This introduces an inherent inaccuracy into the model.</p></li>\n</ul>\n<p>These are not rhetorical questions, so I would really appreciate any comments from hosts  <a href=\"https://www.kaggle.com/gkiar07\" target=\"_blank\">@gkiar07</a> or other interested people.</p>",
      "rawMarkdown": "Topic aim: to understand the real aim of this competition and the practical implications of the model required. If you are struggling with this like me, please feel free to add your questions. If you can prove me wrong, then please do so!\n\n**TL;TR: please see the summary (based on  the original post + comments + questions from other threads) at the end**\n\nIs a PIU prediction model redundant with Internet use as an input?\n\nProblematic Internet use (PIU) is defined as the use of the Internet that creates psychological, social, school and/or work difficulties in a person's life (thanks @mpwolke for the [article](https://www.sciencedirect.com/science/article/abs/pii/S0165178119320098)).\n\n**So you can't have PIU without excessive Internet use and without health or social problems.**\nI wonder what exactly is the need for a Severity Impairment Index (SII) prediction model that relies on features such as \"Internet usage\" (hours spent online per day)? If we need to collect internet usage and health-related data to make predictions, then we're already measuring what we're trying to predict, or what I'm not seeing here?\n\nIn practical terms:\n\n- For patients with high Internet use and health problems: We can already suspect problematic Internet use without needing a model, and include in the treatment plan not only medications for the health condition(s), but also recommendations for reducing Internet use and healthy lifestyles.\n\n- For patients with health problems but low Internet use: We can infer that their health problems have other causes and focus on diagnosing and treating them, again without the help of the model. Similarly, patients with low physical activity and unhealthy lifestyles need advice on both.\n\nOf course, self-reported Internet use can be inaccurate, but we don't have any other characteristics related to Internet use to train a model. So if we can't rely on reported Internet use, we can't properly diagnose PIU.\n\nFurthermore, any associations found between Internet use and health problems do not prove causation. The root cause may be poor health leading to increased internet use rather than the internet use causing health problems. \n\nI'm curious to hear your perspectives on this. Let's discuss! \n\n****\n\nAnother puzzling observation:\n\nThe competition description states that the goal is \"to detect early indicators of problematic Internet and technology use.\"\n\nGiven that, as noted in the comments below, PIU reflect the negative consequences of Internet use (when Internet use begins to cause problems), it seems logical to study early signs of PIU in a population of participants who use the Internet more than average.\n\nThis would mean recognizing subtle shifts in behavior, physical activity, or psychological well-being that could signal the onset of problematic internet use before it leads to significant impairment.\n\nAt the same time, 38.5% of the participants in the train dataset ~~do not use the internet at all~~ use internet less than hour a day. Moreover, a noticable proportion of participants with high SII scores ~~do not use the Internet~~ use internet less than hour a day (21.55% of participants with SII = 2, moderately impaired  by problematic internet use, and 14.7% of those with SII = 3, severely impaired)!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12193662%2Fbc48c089630fb2973007d1501260640a%2FScreenshot%202024-09-23%20080435.png?generation=1727067903360763&alt=media)\n\n[More details are here](https://www.kaggle.com/code/antoninadolgorukova/cmi-piu-features-eda)\n\nIf the index is intended to measure problematic internet use, it shouldn’t produce high scores for participants who ~~don’t use the internet at all~~ spend so little time online, should it? So is this investigator bias, unreliable self-reporting, or data collection error?\n\nAnd what is the point of including such a large proportion of people who ~~do not use the Internet at all~~ spend <1h/day in the internet when the goal is to detect early signs of harmfull internet usage?\n\nIn my opinion, Internet usage data should not be used as a feature for this task, but rather as a condition that participants must meet in order to be included in the study. In this way, we could develop a model to predict whether these individuals show signs of impairment and how severe that impairment is. This would be more consistent with the goals of detecting problematic Internet use.\n\n****\n**Questioning summary:**\n- Data collection for the model involves professional assessments, questionnaires and specialised equipment, adding complexity rather than simplifying the process for families.\n\n- Clinicians with CGAS scores, physical health measures and internet use data have sufficient information to diagnose and assess the severity of PIU in one visit. They could also simply administer the PCIAT during assessments to obtain more accurate SII scores than a predictive model would provide (which also requires much more data and time to collect).\n\n- As @expensivelunch mentioned [here](https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/538082), most actigraphy data were collected after the target outcome had been measured, representing future states. This introduces an inherent inaccuracy into the model.\n\n\nThese are not rhetorical questions, so I would really appreciate any comments from hosts  @gkiar07 or other interested people.",
      "votes": 86
    },
    {
      "id": 3005187,
      "postDate": "2024-10-02T16:33:54.540Z",
      "content": "<p><a href=\"https://www.kaggle.com/antoninadolgorukova\" target=\"_blank\">@antoninadolgorukova</a> I have read two of your posts and went through some folks EDA's. <br>\nDo we even need to question the task if everything is out of hands here?</p>\n<p>Here is my summary:</p>\n<ol>\n<li>There is noise in the data, arguably more noise than the actual signal (missing values, not strictly controlled environment during data collection, and so on - you named it).</li>\n<li>We are to predict the target which is a questionnaire response aggregation (I took some tests online - I failed most of them miserably and got severe score). Some of the features can be a proxy for the responses, though we might doubt that the internet usage is a root cause here.</li>\n<li>CV-LB thread shows \"random-walking\".</li>\n</ol>\n<p>In the light of the mentioned above, we are going to see highly overfitted public leaderboard here.<br>\nThe standard \"analytical toolbox\" (deep understanding of the data, feature engineering), which I guess is your competitive advantage, might be discouraged from the very beginning. It will be hard not to fell into ensembles trap early on and junky-whatever 5 submit a day scenario. </p>\n<p>It still can be a good competition for someone to test ds/ml skills, (data cleaning, wrangling, feature engineering, modeling and inference). Someone even get into the gold zone and will be encouraged to continue his/her DS/ML journey with the money prize. </p>\n<p>Do we really help to solve a real life problem or just playing ds/ml here? Will it be deserved in the same way as if it were well established task? if you really care, this competition might not be a good fit, otherwise it can be a fun. </p>",
      "rawMarkdown": "@antoninadolgorukova I have read two of your posts and went through some folks EDA's. \nDo we even need to question the task if everything is out of hands here?\n\nHere is my summary:\n 1. There is noise in the data, arguably more noise than the actual signal (missing values, not strictly controlled environment during data collection, and so on - you named it).\n 2. We are to predict the target which is a questionnaire response aggregation (I took some tests online - I failed most of them miserably and got severe score). Some of the features can be a proxy for the responses, though we might doubt that the internet usage is a root cause here.\n 3. CV-LB thread shows \"random-walking\".\n\nIn the light of the mentioned above, we are going to see highly overfitted public leaderboard here.\nThe standard \"analytical toolbox\" (deep understanding of the data, feature engineering), which I guess is your competitive advantage, might be discouraged from the very beginning. It will be hard not to fell into ensembles trap early on and junky-whatever 5 submit a day scenario. \n\nIt still can be a good competition for someone to test ds/ml skills, (data cleaning, wrangling, feature engineering, modeling and inference). Someone even get into the gold zone and will be encouraged to continue his/her DS/ML journey with the money prize. \n\nDo we really help to solve a real life problem or just playing ds/ml here? Will it be deserved in the same way as if it were well established task? if you really care, this competition might not be a good fit, otherwise it can be a fun. ",
      "votes": 14,
      "replies": [
        {
          "id": 3005508,
          "postDate": "2024-10-03T02:50:12.777Z",
          "content": "<p>Hi! Well said, agree, this competition feels more like playing with ML than solving a real problem. The data analysis and processing is quite interesting though, so I'll stick with that part). </p>",
          "rawMarkdown": "Hi! Well said, agree, this competition feels more like playing with ML than solving a real problem. The data analysis and processing is quite interesting though, so I'll stick with that part). ",
          "votes": 7
        },
        {
          "id": 3007722,
          "postDate": "2024-10-05T16:43:40.817Z",
          "content": "<p><a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a> \"Questioning the task\" can be seen as a form of trying to better understand the data available and how it relates to the target we want to predict. I don't think it is the same as saying: \"This is pointless, I'm not working on it, and neither should you.\" 🙂 </p>",
          "rawMarkdown": "@sergiosaharovskiy \"Questioning the task\" can be seen as a form of trying to better understand the data available and how it relates to the target we want to predict. I don't think it is the same as saying: \"This is pointless, I'm not working on it, and neither should you.\" 🙂 ",
          "votes": 2,
          "replies": [
            {
              "id": 3008313,
              "postDate": "2024-10-06T13:13:51.593Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        },
        {
          "id": 3078030,
          "postDate": "2024-12-21T17:41:42.243Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/antoninadolgorukova\" target=\"_blank\">@antoninadolgorukova</a>,</p>\n<p>You’ve summed up the challenges of this competition brilliantly. The noise, overfitting risks, and reliance on ensembles definitely make it tricky, but it’s also a great way to practice dealing with messy, real-world-like data. Whether it’s a meaningful problem or just ML “play” depends on one’s goals, but there’s still value in the learning experience. Thanks for sharing your perspective!</p>",
          "rawMarkdown": "Hi @antoninadolgorukova,\n\nYou’ve summed up the challenges of this competition brilliantly. The noise, overfitting risks, and reliance on ensembles definitely make it tricky, but it’s also a great way to practice dealing with messy, real-world-like data. Whether it’s a meaningful problem or just ML “play” depends on one’s goals, but there’s still value in the learning experience. Thanks for sharing your perspective!",
          "votes": 1
        }
      ]
    },
    {
      "id": 2995901,
      "postDate": "2024-09-22T18:15:52.237Z",
      "content": "<p>Hi Antonina,</p>\n<p>Firstly,  excellent topic to be explored. I'd like to read other users approach. Their Data Science approach.</p>\n<p>The issue starts with \"The P\" (Problematic).  I'm not able to provide a Data Science point-of-view, due to my lack of knowledge. <br>\nHowever, I can only speak for myself and the many hours that I spend on Kaggle. Checking my colormap, the only day I was out (Feb, 2024), it was due that I didn't bring the computer to the hospital. </p>\n<p>That's a reflect of NOT HAVING a life cause I used to take care of someone.  Many times when I delivered more public work I was in hospital, when I had almost nothing to do, except to wait. Better stays kaggling on these times, than to think negatively.</p>\n<p>I confess that I burned many pans cause I simply forget that I was cooking something while I was kaggling. In fact, I intend to forget the WHOLE thing, which is bigger and much more Problematic than My Internet usage.</p>\n<p>On the other hand, I'm not totally unhealthy. Since, I'm retired, I can go to the beach and swim daily. Meanwhile, while I'm kaggling, I drink water and don't forget to go to the toilet. I stand-up, shake the body a little bit to change positions  too.  There are specific times that I prefer to watch  television,  to laugh a little bit and mostly when I begin to feel tired of thinking on programming languages, I start to \"abstract\" and stop to see what's the solution .</p>\n<p>Additionally, on my second specialization, the subject was \"Dental Health Promotion\" when I learned so many holistic concepts that I won't be able to forget to keep applying them in my life. </p>\n<p>The key word is Balance. Always  the \"Ancient Greek Temperance\".</p>\n<p>And for me, Kaggling is almost a synonym for The Internet. A Positive Internet usage. At least, I'm often learning substantial stuff. And mostly, keep my mind occupied with subjects that contribute to my personal growth.  </p>\n<p>I also would like to see and read more female Kagglers participations. They are so talented. And, when we advance in career, it's harder to have any time for ouselves to deliver or work on things that we really enjoy. In general, more than males.</p>\n<p>By the way, our Audience is just a bonus 😊</p>\n<p>Cheers, <br>\nMarília.</p>",
      "rawMarkdown": "Hi Antonina,\n\nFirstly,  excellent topic to be explored. I'd like to read other users approach. Their Data Science approach.\n\nThe issue starts with \"The P\" (Problematic).  I'm not able to provide a Data Science point-of-view, due to my lack of knowledge. \nHowever, I can only speak for myself and the many hours that I spend on Kaggle. Checking my colormap, the only day I was out (Feb, 2024), it was due that I didn't bring the computer to the hospital. \n\nThat's a reflect of NOT HAVING a life cause I used to take care of someone.  Many times when I delivered more public work I was in hospital, when I had almost nothing to do, except to wait. Better stays kaggling on these times, than to think negatively.\n\nI confess that I burned many pans cause I simply forget that I was cooking something while I was kaggling. In fact, I intend to forget the WHOLE thing, which is bigger and much more Problematic than My Internet usage.\n\nOn the other hand, I'm not totally unhealthy. Since, I'm retired, I can go to the beach and swim daily. Meanwhile, while I'm kaggling, I drink water and don't forget to go to the toilet. I stand-up, shake the body a little bit to change positions  too.  There are specific times that I prefer to watch  television,  to laugh a little bit and mostly when I begin to feel tired of thinking on programming languages, I start to \"abstract\" and stop to see what's the solution .\n\nAdditionally, on my second specialization, the subject was \"Dental Health Promotion\" when I learned so many holistic concepts that I won't be able to forget to keep applying them in my life. \n\nThe key word is Balance. Always  the \"Ancient Greek Temperance\".\n\nAnd for me, Kaggling is almost a synonym for The Internet. A Positive Internet usage. At least, I'm often learning substantial stuff. And mostly, keep my mind occupied with subjects that contribute to my personal growth.  \n\nI also would like to see and read more female Kagglers participations. They are so talented. And, when we advance in career, it's harder to have any time for ouselves to deliver or work on things that we really enjoy. In general, more than males.\n\nBy the way, our Audience is just a bonus 😊\n\nCheers, \nMarília.",
      "votes": 10,
      "replies": [
        {
          "id": 2996071,
          "postDate": "2024-09-23T03:19:05.023Z",
          "content": "<p>Thank you for sharing your personal experiences and thoughts. I absolutely agree with your emphasis on balance - using the internet productively, like on Kaggle, can be a positive and beneficial ☺️</p>\n<p>Personally, I spend 7-10 hours a day on my laptop, but I still find time to work out or swim. I wouldn't say I suffer from excessive internet use. However, I feel like I might not even notice when this becomes problematic… It seems like PIU is something that our relatives would be more likely to point out - similar to other mental health issues/behavioral addictions. But PIU appears to be surrounded by controversy and ambiguity:</p>\n<p>In fact, articles on this topic acknowledge the importance of the Internet in our lives, stating something like \"There is widespread agreement that the Internet can serve as a tool to enhance well-being, but…\". (here) or \"In the modern world, the Internet constitutes an important tool for work, especially during the SARS-CoV-2 pandemic. It is also used as a common pastime\" (<a href=\"https://annals-general-psychiatry.biomedcentral.com/articles/10.1186/s12991-022-00384-4\" target=\"_blank\">here</a>)… </p>\n<p>and more interesting: </p>\n<p>\"The diagnosis of PIU does not appear in any official diagnostic system, including DSM-V, and there are no widely accepted diagnostic criteria\" (<a href=\"https://www.sciencedirect.com/science/article/abs/pii/S0306460313002669\" target=\"_blank\">here</a>)</p>\n<p>\"The problem starts with terminology, as the appropriate name for the condition or behaviour often labelled “Internet addiction” is not clear. \" (<a href=\"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC2911083/\" target=\"_blank\">link</a>)</p>\n<p>\"Scientific understanding of PIU has lagged behind media attention mainly because of inconsistencies in defining PIU, disagreement about its very existence, and the variable methodological approaches used in studying it\" and </p>\n<p>\"Although there is still not enough empirical evidence to establish whether PIU is a clearly nosologically defined psychopathological entity, it is still widely accepted that it is a global mental health problem\" (<a href=\"https://www.sciencedirect.com/science/article/pii/S0160791X24001362\" target=\"_blank\">link</a>)</p>\n<p>Seems there’s ongoing debate about PIU existence, whether PIU should even be considered a distinct clinical disorder, or if it’s simply a symptom of underlying mental health issues. If we can’t even define the condition consistently, it’s hard to create standardized methods for identifying it.</p>",
          "rawMarkdown": "Thank you for sharing your personal experiences and thoughts. I absolutely agree with your emphasis on balance - using the internet productively, like on Kaggle, can be a positive and beneficial ☺️\n\nPersonally, I spend 7-10 hours a day on my laptop, but I still find time to work out or swim. I wouldn't say I suffer from excessive internet use. However, I feel like I might not even notice when this becomes problematic... It seems like PIU is something that our relatives would be more likely to point out - similar to other mental health issues/behavioral addictions. But PIU appears to be surrounded by controversy and ambiguity:\n\nIn fact, articles on this topic acknowledge the importance of the Internet in our lives, stating something like \"There is widespread agreement that the Internet can serve as a tool to enhance well-being, but...\". (here) or \"In the modern world, the Internet constitutes an important tool for work, especially during the SARS-CoV-2 pandemic. It is also used as a common pastime\" ([here](https://annals-general-psychiatry.biomedcentral.com/articles/10.1186/s12991-022-00384-4))... \n\nand more interesting: \n\n\"The diagnosis of PIU does not appear in any official diagnostic system, including DSM-V, and there are no widely accepted diagnostic criteria\" ([here](https://www.sciencedirect.com/science/article/abs/pii/S0306460313002669))\n\n\"The problem starts with terminology, as the appropriate name for the condition or behaviour often labelled “Internet addiction” is not clear. \" ([link](https://www.ncbi.nlm.nih.gov/pmc/articles/PMC2911083/))\n\n\"Scientific understanding of PIU has lagged behind media attention mainly because of inconsistencies in defining PIU, disagreement about its very existence, and the variable methodological approaches used in studying it\" and \n\n\"Although there is still not enough empirical evidence to establish whether PIU is a clearly nosologically defined psychopathological entity, it is still widely accepted that it is a global mental health problem\" ([link](https://www.sciencedirect.com/science/article/pii/S0160791X24001362))\n\nSeems there’s ongoing debate about PIU existence, whether PIU should even be considered a distinct clinical disorder, or if it’s simply a symptom of underlying mental health issues. If we can’t even define the condition consistently, it’s hard to create standardized methods for identifying it.\n",
          "votes": 2,
          "replies": [
            {
              "id": 2996354,
              "postDate": "2024-09-23T11:44:37.150Z",
              "content": "<p>For the record, people that really have addition don't get cured (My opinion). They change their additions. Which initially could seem a cure is because they aren't reproducing the same behavior. Example, someone that have gambling issues could start to have problems on another field (like shopping, sex, eating disorders).</p>\n<p>On the Internet, many are addicted to likes/votes/medals. I can understand those that  can monetize their \"performances\" (e.g. influencers). However, the ones that are glad to have \"followers\" and many likes to their irrelevant posts (many with incorrect information), these persons are just searching for validation to their thoughts/ideas. And get some kind of attention that they won't have on their personal lives. Albeit, after a certain time, they'll realize that this behavior won't lead them to better positions or even anywhere. </p>\n<p>Thank you for the many links about this subject. </p>",
              "rawMarkdown": "For the record, people that really have addition don't get cured (My opinion). They change their additions. Which initially could seem a cure is because they aren't reproducing the same behavior. Example, someone that have gambling issues could start to have problems on another field (like shopping, sex, eating disorders).\n\nOn the Internet, many are addicted to likes/votes/medals. I can understand those that  can monetize their \"performances\" (e.g. influencers). However, the ones that are glad to have \"followers\" and many likes to their irrelevant posts (many with incorrect information), these persons are just searching for validation to their thoughts/ideas. And get some kind of attention that they won't have on their personal lives. Albeit, after a certain time, they'll realize that this behavior won't lead them to better positions or even anywhere. \n\nThank you for the many links about this subject. ",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 3000006,
      "postDate": "2024-09-27T07:08:59.477Z",
      "content": "<p>I've just realised: what we're essentially doing here is training a model to predict a score derived from potentially biased or 'unfair' responses, rather than an objective measure of impairment from internet use… </p>\n<p>The SII is calculated from the PCIAT test score, so in other words, we have to predict how test participants (or their parents) would answer this questionnaire😅. That's kinda weird in practical terms, but technically one can try to predict the answer to each question, then calculate the total score and the corresponding SII.</p>",
      "rawMarkdown": "I've just realised: what we're essentially doing here is training a model to predict a score derived from potentially biased or 'unfair' responses, rather than an objective measure of impairment from internet use... \n\nThe SII is calculated from the PCIAT test score, so in other words, we have to predict how test participants (or their parents) would answer this questionnaire😅. That's kinda weird in practical terms, but technically one can try to predict the answer to each question, then calculate the total score and the corresponding SII.",
      "votes": 8,
      "replies": [
        {
          "id": 3000018,
          "postDate": "2024-09-27T07:34:34.230Z",
          "content": "<p>Yes, and the problem is how to estimate PCIAT_total.<br>\nI tried to skip those 20 PCIAT sub-questions since most of them are 0 or 1.<br>\nI tried to find some physical feature that's correlated to PCIAT_total, but not working well yet.<br>\nSome of the obvious one will be age, height, weight.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22136130%2Fe0f24b88e765fcbd9693fee36075ed1c%2FPCIAT_Fitness.png?generation=1727422040143130&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Yes, and the problem is how to estimate PCIAT_total.\nI tried to skip those 20 PCIAT sub-questions since most of them are 0 or 1.\nI tried to find some physical feature that's correlated to PCIAT_total, but not working well yet.\nSome of the obvious one will be age, height, weight.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22136130%2Fe0f24b88e765fcbd9693fee36075ed1c%2FPCIAT_Fitness.png?generation=1727422040143130&alt=media)",
          "votes": 3,
          "replies": [
            {
              "id": 3001159,
              "postDate": "2024-09-28T14:55:07.273Z",
              "content": "<p>Interesting, will this be any better if you remove rows where any of the 20 PCIAT sub-questions are missing (and PCIAT_total is actually wrong as it is a sum of non-NA values)?</p>",
              "rawMarkdown": "Interesting, will this be any better if you remove rows where any of the 20 PCIAT sub-questions are missing (and PCIAT_total is actually wrong as it is a sum of non-NA values)?"
            },
            {
              "id": 3012271,
              "postDate": "2024-10-08T20:34:46.670Z",
              "content": "<p>Indeed. If we leave out the 20 PCIAT, the maximum correlation of 'sii' is ~35% with age, height, and weight—which I think is why most of us are getting model accuracy at 47% max. </p>\n<p>I thought of predicting some PCIAT features using parquet file info, but that is not complete; it only has about ~900 IDs. Can you think of any other ways to predict PCIAT features?</p>",
              "rawMarkdown": "Indeed. If we leave out the 20 PCIAT, the maximum correlation of 'sii' is ~35% with age, height, and weight—which I think is why most of us are getting model accuracy at 47% max. \n\nI thought of predicting some PCIAT features using parquet file info, but that is not complete; it only has about ~900 IDs. Can you think of any other ways to predict PCIAT features?",
              "votes": 1
            },
            {
              "id": 3012293,
              "postDate": "2024-10-08T21:09:46.417Z",
              "content": "<p>If you look at the questions corresponding to each PCIAT feature in data dictionary, you'll see that it's hard to imagine which of the provided features could possibly serve as predictors for many of them, like phone calls, or checking emails, for example 😔. Most of the questions assess behaviour or mood… But I didn't do this analysis yet,  may be I'll find smth, but I doubt.</p>",
              "rawMarkdown": "If you look at the questions corresponding to each PCIAT feature in data dictionary, you'll see that it's hard to imagine which of the provided features could possibly serve as predictors for many of them, like phone calls, or checking emails, for example 😔. Most of the questions assess behaviour or mood... But I didn't do this analysis yet,  may be I'll find smth, but I doubt.",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 3008342,
      "postDate": "2024-10-06T13:39:25.520Z",
      "content": "<p><em>This comment has a \"questioning the task\" flavor so I'm posting it here.</em> </p>\n<p>Looking at a crosstab of my (not so well) predicted and actual sii values, I have a lot of actual sii=0 being classified as sii=1. One reason for that is that the PCIAT Total sii=0 boundary at 30 is very close to the peak in the distribution of the Total values (histogram below.)</p>\n<p>Having the boundary near the peak maximizes the confusion between the classes, and effectively cuts the around-the-peak class in two. For examples, better, data-based boundary locations on either side of the peak could be at 19-and-below, and/or 37-and-up. Though there may be other considerations in selecting the boundaries than ease of classification… 🙂</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1496178%2F0e79d497b4472182a502c8f73ecf16e0%2FDistribution_of_Total.jpg?generation=1728220007766912&amp;alt=media\" alt=\"Histogram of Total\"></p>",
      "rawMarkdown": "*This comment has a \"questioning the task\" flavor so I'm posting it here.* \n\nLooking at a crosstab of my (not so well) predicted and actual sii values, I have a lot of actual sii=0 being classified as sii=1. One reason for that is that the PCIAT Total sii=0 boundary at 30 is very close to the peak in the distribution of the Total values (histogram below.)\n\nHaving the boundary near the peak maximizes the confusion between the classes, and effectively cuts the around-the-peak class in two. For examples, better, data-based boundary locations on either side of the peak could be at 19-and-below, and/or 37-and-up. Though there may be other considerations in selecting the boundaries than ease of classification... 🙂\n\n![Histogram of Total](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1496178%2F0e79d497b4472182a502c8f73ecf16e0%2FDistribution_of_Total.jpg?generation=1728220007766912&alt=media)",
      "votes": 6,
      "replies": [
        {
          "id": 3009636,
          "postDate": "2024-10-08T06:37:57.597Z",
          "content": "<p>Interesting point! I'm not into modelling, so don't follow how it goes, but I think I've seen people in other threads saying that regression works better than classification for this task, so the boundary issue could at least partially explain that. And what if you predict PCIAT_Total directly and then map the predicted values to SII, does that reduce the confusion between adjacent SII classes?</p>",
          "rawMarkdown": "Interesting point! I'm not into modelling, so don't follow how it goes, but I think I've seen people in other threads saying that regression works better than classification for this task, so the boundary issue could at least partially explain that. And what if you predict PCIAT_Total directly and then map the predicted values to SII, does that reduce the confusion between adjacent SII classes?",
          "votes": 1
        }
      ]
    },
    {
      "id": 3002388,
      "postDate": "2024-09-30T02:03:48.060Z",
      "content": "<p>Thanks for brining up this discussion and the responses. Yes, we are trying to predict a questionaire score using as predictors sometimes hard-to-get data values (csv) as well as the wrist-worn enmo, angles, and light. It does seem problematic 🥴</p>\n<p>Taking your comments into account, maybe a different competition would be based on predicting sii level based on data easily obtained in a single office visit (height, weight, age, BP, etc.) when the participant also picked up their wrist-worn device to return in a month… Though it would be more an assessment of \"healthy activive lifestyle\" than internet per se.</p>\n<p>P.S. <a href=\"https://www.kaggle.com/waechter\" target=\"_blank\">@waechter</a> , replace 'gaming' with 'kaggle' and I'm 8 for 9 in your list -- I'm OK with 7: my wife will say as I stare into space: \"You're thinking about data again, right?\" 😃</p>",
      "rawMarkdown": "Thanks for brining up this discussion and the responses. Yes, we are trying to predict a questionaire score using as predictors sometimes hard-to-get data values (csv) as well as the wrist-worn enmo, angles, and light. It does seem problematic 🥴\n\nTaking your comments into account, maybe a different competition would be based on predicting sii level based on data easily obtained in a single office visit (height, weight, age, BP, etc.) when the participant also picked up their wrist-worn device to return in a month... Though it would be more an assessment of \"healthy activive lifestyle\" than internet per se.\n\nP.S. @waechter , replace 'gaming' with 'kaggle' and I'm 8 for 9 in your list -- I'm OK with 7: my wife will say as I stare into space: \"You're thinking about data again, right?\" 😃",
      "votes": 5,
      "replies": [
        {
          "id": 3003025,
          "postDate": "2024-09-30T15:28:03.920Z",
          "content": "<p>That's cute! I think the big difference between problematic usage and being passionnate is ashamed vs proud. I bet you don't lie to your wife about kaggle (7) and are proud of the time you spend here 😀</p>",
          "rawMarkdown": "That's cute! I think the big difference between problematic usage and being passionnate is ashamed vs proud. I bet you don't lie to your wife about kaggle (7) and are proud of the time you spend here 😀",
          "votes": 3
        }
      ]
    },
    {
      "id": 2999149,
      "postDate": "2024-09-26T12:24:32.573Z",
      "content": "<p>By the way, here's another thought on the subject: I don't think the measures (data to be used as features for modelling) provided in this competition are \"easily obtainable\". The collection of data like accelerometry, sleep patterns, and physical activity (fitness, endurance) requires not only requires specialised equipment, but can also take days or even months to collect (the contestants wore accelerometers for up to 81 days) 🥲.</p>\n<hr>\n<p>UPD: The CGAS (Children's Global Assessment Scale) and several other features can only be assessed by a clinician. This blows my mind:</p>\n<p>PROBLEM FROM THE COMPETITION OWERVIEW:</p>\n<blockquote>\n  <p>Current methods for measuring problematic internet use in children and adolescents are often complex and require professional assessments. This creates access, cultural, and linguistic barriers for many families.</p>\n</blockquote>\n<p>SOLUTION SEEMS TO BE:</p>\n<p>training a model on &gt;50 features, some of which can only be assessed by a clinician and some of which take months to collect…</p>",
      "rawMarkdown": "By the way, here's another thought on the subject: I don't think the measures (data to be used as features for modelling) provided in this competition are \"easily obtainable\". The collection of data like accelerometry, sleep patterns, and physical activity (fitness, endurance) requires not only requires specialised equipment, but can also take days or even months to collect (the contestants wore accelerometers for up to 81 days) 🥲.\n\n****\nUPD: The CGAS (Children's Global Assessment Scale) and several other features can only be assessed by a clinician. This blows my mind:\n\nPROBLEM FROM THE COMPETITION OWERVIEW:\n>Current methods for measuring problematic internet use in children and adolescents are often complex and require professional assessments. This creates access, cultural, and linguistic barriers for many families.\n\nSOLUTION SEEMS TO BE:\n\ntraining a model on >50 features, some of which can only be assessed by a clinician and some of which take months to collect...",
      "votes": 5,
      "replies": [
        {
          "id": 3018248,
          "postDate": "2024-10-15T16:22:08.670Z",
          "content": "<p>I completely agree with you this is futile. This is in no way useful</p>",
          "rawMarkdown": "I completely agree with you this is futile. This is in no way useful"
        }
      ]
    },
    {
      "id": 2995826,
      "postDate": "2024-09-22T16:28:48.767Z",
      "content": "<p>Interesting question, but I think that's a bit of a misconception. From what I understand, problematic usage/addiction is defined by negative consequences, not time or amounts spent.</p>\n<p>From <a href=\"https://en.wikipedia.org/wiki/Video_game_addiction\" target=\"_blank\">https://en.wikipedia.org/wiki/Video_game_addiction</a> there is 9 criteria to Internet Gaming Disorder:</p>\n<blockquote>\n  <ol>\n  <li>Pre-occupation. Do you spend a lot of time thinking about games even when you are not playing, or planning when you can play next?</li>\n  <li>Withdrawal. Do you feel restless, irritable, moody, angry, anxious or sad when attempting to cut down or stop gaming, or when you are unable to play?</li>\n  <li>Tolerance. Do you feel the need to play for increasing amounts of time, play more exciting games, or use more powerful equipment to get the same amount of excitement you used to get?</li>\n  <li>Reduce/stop. Do you feel that you should play less, but are unable to cut back on the amount of time you spend playing games?</li>\n  <li>Give up other activities. Do you lose interest in or reduce participation in other recreational activities due to gaming?</li>\n  <li>Continue despite problems. Do you continue to play games even though you are aware of negative consequences, such as not getting enough sleep, being late to school/work, spending too much money, having arguments with others, or neglecting important duties?</li>\n  <li>Deceive/cover up. Do you lie to family, friends or others about how much you game, or try to keep your family or friends from knowing how much you game?</li>\n  <li>Escape adverse moods. Do you game to escape from or forget about personal problems, or to relieve uncomfortable feelings such as guilt, anxiety, helplessness or depression?</li>\n  <li>Risk/lose relationships/opportunities. Do you risk or lose significant relationships, or job, educational or career opportunities because of gaming?</li>\n  </ol>\n</blockquote>\n<p>These questions are not about the time spent gaming/online, but how it affects life. </p>\n<p>For this competition we are trying to predict the <code>Severity Impairment Index</code>, which is the total of the responses from 20 questions <code>Parent-Child Internet Addiction Test</code>. Not the time spent online</p>\n<p>For example: </p>\n<ul>\n<li>Someone can be online 8h/day and have no problem, being a perfect child.</li>\n<li>And another be online way less, but with conflict with his parents and negative consequences</li>\n</ul>\n<p>Hope this make sense!</p>",
      "rawMarkdown": "Interesting question, but I think that's a bit of a misconception. From what I understand, problematic usage/addiction is defined by negative consequences, not time or amounts spent.\n\nFrom https://en.wikipedia.org/wiki/Video_game_addiction there is 9 criteria to Internet Gaming Disorder:\n> 1. Pre-occupation. Do you spend a lot of time thinking about games even when you are not playing, or planning when you can play next?\n1. Withdrawal. Do you feel restless, irritable, moody, angry, anxious or sad when attempting to cut down or stop gaming, or when you are unable to play?\n1. Tolerance. Do you feel the need to play for increasing amounts of time, play more exciting games, or use more powerful equipment to get the same amount of excitement you used to get?\n1. Reduce/stop. Do you feel that you should play less, but are unable to cut back on the amount of time you spend playing games?\n1. Give up other activities. Do you lose interest in or reduce participation in other recreational activities due to gaming?\n1. Continue despite problems. Do you continue to play games even though you are aware of negative consequences, such as not getting enough sleep, being late to school/work, spending too much money, having arguments with others, or neglecting important duties?\n1. Deceive/cover up. Do you lie to family, friends or others about how much you game, or try to keep your family or friends from knowing how much you game?\n1. Escape adverse moods. Do you game to escape from or forget about personal problems, or to relieve uncomfortable feelings such as guilt, anxiety, helplessness or depression?\n1. Risk/lose relationships/opportunities. Do you risk or lose significant relationships, or job, educational or career opportunities because of gaming?\n\nThese questions are not about the time spent gaming/online, but how it affects life. \n\nFor this competition we are trying to predict the `Severity Impairment Index`, which is the total of the responses from 20 questions `Parent-Child Internet Addiction Test`. Not the time spent online\n\nFor example: \n- Someone can be online 8h/day and have no problem, being a perfect child.\n- And another be online way less, but with conflict with his parents and negative consequences\n\nHope this make sense!",
      "votes": 6,
      "replies": [
        {
          "id": 2995846,
          "postDate": "2024-09-22T17:05:22.387Z",
          "content": "<p>Thanks! I completely agree that PIU is characterized by negative consequences rather than just the amount of time spent online. </p>\n<p>Essentially, the SII is an index that can be calculated for a patient using the questionnaire as you mentioned. We are being asked to predict it using more objective data from physical observations (although there is data from other questionnaires). If we remove internet use from these data, there is nothing left that would allow us to link the participants' condition specifically to promlematic internet use.</p>\n<p>For example, in the training data, there are children who are not entirely healthy (e.g., overweight, sleep disturbance) who, it turns out, do not use the internet at all. And there are others with problems who report using it at least an hour a day. Both groups have health issues, but only the condition of those who use the internet can somehow be linked to PIU.</p>",
          "rawMarkdown": "Thanks! I completely agree that PIU is characterized by negative consequences rather than just the amount of time spent online. \n\nEssentially, the SII is an index that can be calculated for a patient using the questionnaire as you mentioned. We are being asked to predict it using more objective data from physical observations (although there is data from other questionnaires). If we remove internet use from these data, there is nothing left that would allow us to link the participants' condition specifically to promlematic internet use.\n\nFor example, in the training data, there are children who are not entirely healthy (e.g., overweight, sleep disturbance) who, it turns out, do not use the internet at all. And there are others with problems who report using it at least an hour a day. Both groups have health issues, but only the condition of those who use the internet can somehow be linked to PIU.",
          "votes": 2,
          "replies": [
            {
              "id": 2995881,
              "postDate": "2024-09-22T17:59:12.383Z",
              "content": "<p>I understand your point, a children can be unhealthy due to other cause than internet usage.<br>\nBut I think Internet Usage=<code>PreInt_EduHx-computerinternet_hoursday</code> is misleading, because 0=Less than 1h/day, not don't use internet at all. <br>\nThere are rows in the training data with <code>ssi</code>&gt;=2 and <code>PreInt_EduHx-computerinternet_hoursday</code>==0</p>",
              "rawMarkdown": "I understand your point, a children can be unhealthy due to other cause than internet usage.\nBut I think Internet Usage=`PreInt_EduHx-computerinternet_hoursday` is misleading, because 0=Less than 1h/day, not don't use internet at all. \nThere are rows in the training data with `ssi`>=2 and `PreInt_EduHx-computerinternet_hoursday`==0",
              "votes": 2
            },
            {
              "id": 2995905,
              "postDate": "2024-09-22T18:19:50.647Z",
              "content": "<p>Yep, high SSI and no internet use seems strange, since excessive internet use is by definition expected for a person with PIU. And overall, the data in this column is misleading, since the goal of this competition was formulated as</p>\n<blockquote>\n  <p>to detect early indicators of problematic Internet and technology use</p>\n</blockquote>\n<p>How could indicators of PIU be detected in participants who did not use the Internet at all or used it up to 3 hours a day (which I would say is less than usual today) ?</p>",
              "rawMarkdown": "Yep, high SSI and no internet use seems strange, since excessive internet use is by definition expected for a person with PIU. And overall, the data in this column is misleading, since the goal of this competition was formulated as\n\n> to detect early indicators of problematic Internet and technology use\n\nHow could indicators of PIU be detected in participants who did not use the Internet at all or used it up to 3 hours a day (which I would say is less than usual today) ?",
              "votes": 1
            },
            {
              "id": 2995939,
              "postDate": "2024-09-22T19:16:07.933Z",
              "content": "<p>BTW, in the PCIAT test, based on which this SII was calculated, most questions are about internet use,  or directly related to it (but as you mentioned, not all participants use internet! how could they get high SII scores? it's rather investigator bias).</p>\n<p>I think if we imagine that we are predicting just some kind of severity index related to overall physical activity and well-being, it starts to make more sense. Suppose we get an accurate model that predicts this index. Now we can collect data (all the train features) from children in schools, for example, and predict the index. Children with high scores would need to be studied further to find the cause - it could be anything: genetics, poor nutrition, social circumstances, etc., or it could be excessive Internet use. So by analyzing this particular cohort with high SII scores and a history that indicates potential Internet addiction, we can detect PIU.</p>",
              "rawMarkdown": "BTW, in the PCIAT test, based on which this SII was calculated, most questions are about internet use,  or directly related to it (but as you mentioned, not all participants use internet! how could they get high SII scores? it's rather investigator bias).\n\nI think if we imagine that we are predicting just some kind of severity index related to overall physical activity and well-being, it starts to make more sense. Suppose we get an accurate model that predicts this index. Now we can collect data (all the train features) from children in schools, for example, and predict the index. Children with high scores would need to be studied further to find the cause - it could be anything: genetics, poor nutrition, social circumstances, etc., or it could be excessive Internet use. So by analyzing this particular cohort with high SII scores and a history that indicates potential Internet addiction, we can detect PIU.",
              "votes": 2
            },
            {
              "id": 3003335,
              "postDate": "2024-09-30T22:54:42.650Z",
              "content": "<p>It wasn't told explicitly, but I guess the goal is to use model based on physical activities/ measures to replace the questionnaire in future.<br>\nThe psychologists understand that the PICAT questionnaire is not objective enough, which is based on parent's response.<br>\nAnd it depends on the content kids spending on internet, e.g. 3 hours spending on STEM is different form 1/2 hour spending on violence video.<br>\nI guess the time-series can provide some hidden objective insights to sii, e.g. if they have a lot of sunshine (light) and activities (X, Y, Z etc) on Saturday/ Sunday, parents tend to rate lower sii score. The only problem is  I've no idea how to extract those data from the time-series 🤣</p>",
              "rawMarkdown": "It wasn't told explicitly, but I guess the goal is to use model based on physical activities/ measures to replace the questionnaire in future.\nThe psychologists understand that the PICAT questionnaire is not objective enough, which is based on parent's response.\nAnd it depends on the content kids spending on internet, e.g. 3 hours spending on STEM is different form 1/2 hour spending on violence video.\nI guess the time-series can provide some hidden objective insights to sii, e.g. if they have a lot of sunshine (light) and activities (X, Y, Z etc) on Saturday/ Sunday, parents tend to rate lower sii score. The only problem is  I've no idea how to extract those data from the time-series 🤣",
              "votes": 2
            },
            {
              "id": 3003618,
              "postDate": "2024-10-01T05:55:03.227Z",
              "content": "<p>Thanks for yur comment! </p>\n<blockquote>\n  <p>the goal is to use model based on physical activities/ measures to replace the questionnaire in future.</p>\n</blockquote>\n<p>I believe that training a model to predict a priori noisy and imprecise questionnaire responses inevitably leads to a model that is even less reliable than the questionnaire itself.</p>\n<p>Regarding actigraphy data, you may find my notebook \"CMI-PIU: Actigraphy data EDA\" usefull, although the work on feature extraction is still in progress, I have already added some code for such features as duration and number of different activity patterns (no movement, low, moderate and vigorous physical activity) and features capturing the variation in activity over the 24-hour cycle).</p>",
              "rawMarkdown": "Thanks for yur comment! \n>the goal is to use model based on physical activities/ measures to replace the questionnaire in future.\n\nI believe that training a model to predict a priori noisy and imprecise questionnaire responses inevitably leads to a model that is even less reliable than the questionnaire itself.\n\nRegarding actigraphy data, you may find my notebook \"CMI-PIU: Actigraphy data EDA\" usefull, although the work on feature extraction is still in progress, I have already added some code for such features as duration and number of different activity patterns (no movement, low, moderate and vigorous physical activity) and features capturing the variation in activity over the 24-hour cycle)."
            }
          ]
        }
      ]
    },
    {
      "id": 3017884,
      "postDate": "2024-10-15T10:16:29.643Z",
      "content": "<p>I think your question arrives because you are looking from a data analyst perspective or like a doctor or specialist who tries to deeply understand the whole bunch of data and get a final diagnosis. But that's about ML in general - to do the hard work of putting everything together and in a few seconds give a diagnosis that a specialist could take a closer look if it's of interest and save time (and $$$) by not analyzing all the lower class samples. So we (as data scientists) are spending 3 months by understanding the data and trying our best to construct a pretty decent model and then other specialists will save their time in the future, or we hope so at least…</p>",
      "rawMarkdown": "I think your question arrives because you are looking from a data analyst perspective or like a doctor or specialist who tries to deeply understand the whole bunch of data and get a final diagnosis. But that's about ML in general - to do the hard work of putting everything together and in a few seconds give a diagnosis that a specialist could take a closer look if it's of interest and save time (and $$$) by not analyzing all the lower class samples. So we (as data scientists) are spending 3 months by understanding the data and trying our best to construct a pretty decent model and then other specialists will save their time in the future, or we hope so at least...",
      "votes": 3,
      "replies": [
        {
          "id": 3018196,
          "postDate": "2024-10-15T15:29:02.553Z",
          "content": "<p>Hi! I appreciate your point about the value of machine learning (ML) in assisting specialists and potentially saving time and resources. And I'm specifically questioning the ML in the task and the practical implications of the results - the model. </p>\n<p>Look, when we have a model trained to predict human responses to the PCIAT questionnaire (which is essentially the SII), for new participants, we still need to collect all the features to run it and get predictions. The prediction will take a few seconds, agree, buy some of these features take months to collect, and some require visits to a clinician or special equipment.</p>\n<p>At the same time, a clinician could get answers to the PCIAT questionnaire directly in a single visit. So why would they need a model whose prediction is obviously less accurate and takes a lot of time?</p>",
          "rawMarkdown": "Hi! I appreciate your point about the value of machine learning (ML) in assisting specialists and potentially saving time and resources. And I'm specifically questioning the ML in the task and the practical implications of the results - the model. \n\nLook, when we have a model trained to predict human responses to the PCIAT questionnaire (which is essentially the SII), for new participants, we still need to collect all the features to run it and get predictions. The prediction will take a few seconds, agree, buy some of these features take months to collect, and some require visits to a clinician or special equipment.\n\nAt the same time, a clinician could get answers to the PCIAT questionnaire directly in a single visit. So why would they need a model whose prediction is obviously less accurate and takes a lot of time?",
          "votes": 4,
          "replies": [
            {
              "id": 3018232,
              "postDate": "2024-10-15T16:16:43.750Z",
              "content": "<p>There are some features that requires time/specialists, but there are also the time series that are very difficult and time consuming to analyze manually, but relatively easy to obtain (put the bracelet on and take it in a month), so one reason is to research what can we do with them. Another reason can be - to avoid additional time/resources for 20 tests that actually is our base target so if we could predict them or the final result - these tests will be necessary just for separate cases for specialist's use. So the overall idea is to research what can we do with all the data to cut expenses, optimize and hurry the evaluation process and maybe discover something additional on this path that could help in taking the decission.</p>",
              "rawMarkdown": "There are some features that requires time/specialists, but there are also the time series that are very difficult and time consuming to analyze manually, but relatively easy to obtain (put the bracelet on and take it in a month), so one reason is to research what can we do with them. Another reason can be - to avoid additional time/resources for 20 tests that actually is our base target so if we could predict them or the final result - these tests will be necessary just for separate cases for specialist's use. So the overall idea is to research what can we do with all the data to cut expenses, optimize and hurry the evaluation process and maybe discover something additional on this path that could help in taking the decission."
            },
            {
              "id": 3018428,
              "postDate": "2024-10-15T18:43:04.040Z",
              "content": "<p>The target in this competition is an index derived from the total scores of 20 questions. In real life, this index is calculated based on responses of parents to those 20 questions, just  this. The model you develop for this competition is trained to predict these scores. </p>\n<p>So to get this Sii,  a doctor can either ask a parent to answer 20 questions or collect &gt;50 measurements (features) plus actigraphy data and predict this Sii with a model. What would be the more accurate and easier way for both sides?</p>\n<p>May be I just don't see you point…</p>",
              "rawMarkdown": "The target in this competition is an index derived from the total scores of 20 questions. In real life, this index is calculated based on responses of parents to those 20 questions, just  this. The model you develop for this competition is trained to predict these scores. \n\nSo to get this Sii,  a doctor can either ask a parent to answer 20 questions or collect >50 measurements (features) plus actigraphy data and predict this Sii with a model. What would be the more accurate and easier way for both sides?\n\nMay be I just don't see you point...\n",
              "votes": 3
            },
            {
              "id": 3039398,
              "postDate": "2024-11-08T02:23:16.757Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 3077971,
      "postDate": "2024-12-21T16:18:18.843Z",
      "content": "<p><a href=\"https://www.kaggle.com/antoninadolgorukova\" target=\"_blank\">@antoninadolgorukova</a>  - I really appreciated your work and perspective on this competition! I looked for your submission, then I saw you only do bio-competitions - best of luck in your next one :)</p>",
      "rawMarkdown": "@antoninadolgorukova  - I really appreciated your work and perspective on this competition! I looked for your submission, then I saw you only do bio-competitions - best of luck in your next one :)",
      "votes": 1
    },
    {
      "id": 3006555,
      "postDate": "2024-10-04T10:23:17.267Z",
      "content": "<p>If we were predicting future health and social problems with current internet use, this challenge would have made more sense imo. But given these items are features, it is quite a futile exercise. </p>",
      "rawMarkdown": "If we were predicting future health and social problems with current internet use, this challenge would have made more sense imo. But given these items are features, it is quite a futile exercise. ",
      "votes": 1
    },
    {
      "id": 3004401,
      "postDate": "2024-10-01T19:19:40.100Z",
      "content": "<p>guys I found out that most of the insights from the time-series (parquet files) don't match with the original tabular data. is it expected?</p>",
      "rawMarkdown": "guys I found out that most of the insights from the time-series (parquet files) don't match with the original tabular data. is it expected?",
      "votes": 1,
      "replies": [
        {
          "id": 3005168,
          "postDate": "2024-10-02T15:56:06.583Z",
          "content": "<p>It is not typically expected for insights from time-series parquet files to deviate significantly from the original tabular data unless there are issues during data processing or transformation.</p>",
          "rawMarkdown": "It is not typically expected for insights from time-series parquet files to deviate significantly from the original tabular data unless there are issues during data processing or transformation."
        },
        {
          "id": 3005510,
          "postDate": "2024-10-03T03:04:15.983Z",
          "content": "<p>It would be interesting to hear more details about what exactly doesn't match. I haven't yet compared these 2 data sources myself, but I would expect to see some agreement, for example, between the total activity level derived from actigraphy data and, say, the Activity Summary Score from the Physical Activity Questionnaire, and for sleep patterns and duration to match the SDS features in train.csv. However, agreement may depend on how the time series data have been cleaned and aggregated. And some discrepancies might be expected if the season of data collection varied.</p>",
          "rawMarkdown": "It would be interesting to hear more details about what exactly doesn't match. I haven't yet compared these 2 data sources myself, but I would expect to see some agreement, for example, between the total activity level derived from actigraphy data and, say, the Activity Summary Score from the Physical Activity Questionnaire, and for sleep patterns and duration to match the SDS features in train.csv. However, agreement may depend on how the time series data have been cleaned and aggregated. And some discrepancies might be expected if the season of data collection varied.",
          "replies": [
            {
              "id": 3005850,
              "postDate": "2024-10-03T12:25:52.273Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 3005892,
              "postDate": "2024-10-03T13:31:35.660Z",
              "content": "<p>Is this CHatGPT answering 😱? <br>\nIf not, sorry, but what's the point? Your answer doesn't seem to be related to my comment to <a href=\"https://www.kaggle.com/godanu\" target=\"_blank\">@godanu</a> quesion or to the data we have.</p>",
              "rawMarkdown": "Is this CHatGPT answering 😱? \nIf not, sorry, but what's the point? Your answer doesn't seem to be related to my comment to @godanu quesion or to the data we have.",
              "votes": 1
            },
            {
              "id": 3005956,
              "postDate": "2024-10-03T14:49:36.860Z",
              "content": "<p>I apologize for any confusion. It seems I misunderstood the context of your comment. Let me clarify or provide more relevant information based on the specific data  you're referring to.</p>",
              "rawMarkdown": "I apologize for any confusion. It seems I misunderstood the context of your comment. Let me clarify or provide more relevant information based on the specific data  you're referring to."
            },
            {
              "id": 3006063,
              "postDate": "2024-10-03T17:27:13.727Z",
              "content": "<p>Hey ChatGPT, do you perchance have any good cookie recipes?</p>",
              "rawMarkdown": "Hey ChatGPT, do you perchance have any good cookie recipes?",
              "votes": 1
            },
            {
              "id": 3006066,
              "postDate": "2024-10-03T17:30:46.143Z",
              "content": "<p><a href=\"https://chatgpt.com/share/66fed4b9-05b4-8001-b33e-a3ab61deaf93\" target=\"_blank\">https://chatgpt.com/share/66fed4b9-05b4-8001-b33e-a3ab61deaf93</a></p>",
              "rawMarkdown": "https://chatgpt.com/share/66fed4b9-05b4-8001-b33e-a3ab61deaf93",
              "votes": 1
            }
          ]
        },
        {
          "id": 3007706,
          "postDate": "2024-10-05T16:30:00.703Z",
          "content": "<p>The parquet data are an independent source of information from the tabular data, so I don't think there is any exact matching that we'd expect.<br>\n<em>Do the parquet data have any predictive power for the target value?</em> Here's a plot of the PCIAT Total score versus a first-cut feature, Wrist-pct_active, that I calculated from the parquet files. The correlation is -0.288 which is stronger than all but 6 of the csv tabular features. So there is some information in the parquet files that is related to the target value, and including parquet-derived features could help the model.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1496178%2Fa54350668784502478b89f4296807598%2FWrist_Total_vs_pct_active.jpg?generation=1728144923016716&amp;alt=media\" alt=\"PCIAT Total score versus Wrist-pct_active\"></p>",
          "rawMarkdown": "The parquet data are an independent source of information from the tabular data, so I don't think there is any exact matching that we'd expect.\n*Do the parquet data have any predictive power for the target value?* Here's a plot of the PCIAT Total score versus a first-cut feature, Wrist-pct_active, that I calculated from the parquet files. The correlation is -0.288 which is stronger than all but 6 of the csv tabular features. So there is some information in the parquet files that is related to the target value, and including parquet-derived features could help the model.\n\n![PCIAT Total score versus Wrist-pct_active](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1496178%2Fa54350668784502478b89f4296807598%2FWrist_Total_vs_pct_active.jpg?generation=1728144923016716&alt=media)\n\n",
          "votes": 2
        }
      ]
    },
    {
      "id": 3001847,
      "postDate": "2024-09-29T10:40:36.213Z",
      "content": "<p>The approach is pretty nice but should i remove the outliers or replace it with mean/median?</p>",
      "rawMarkdown": "The approach is pretty nice but should i remove the outliers or replace it with mean/median?",
      "votes": 1
    },
    {
      "id": 3000567,
      "postDate": "2024-09-27T19:17:44.347Z",
      "content": "<p>I had not thought about this problem. thank you for the input! <br>\nI was just wondering how you would do to fix the data? </p>",
      "rawMarkdown": "I had not thought about this problem. thank you for the input! \nI was just wondering how you would do to fix the data? \n",
      "votes": 1,
      "replies": [
        {
          "id": 3000921,
          "postDate": "2024-09-28T08:53:59.687Z",
          "content": "<p>Thanks! Idk, it seems the whole competition is in needs to be fixed, haha. The inconsistencies I point out in this thread are probably also present in the hidden test data. Even if they have valid explanations, we are essentially being asked to predict human responses to a questionnaire - information that is inherently prone to inaccuracy and bias.<br>\nThis makes it inherently impossible to develop a reliable model to solve the task, but you still have a chance to model the noise and inconsistencies and get a good score, but it's more a matter of luck…</p>",
          "rawMarkdown": "Thanks! Idk, it seems the whole competition is in needs to be fixed, haha. The inconsistencies I point out in this thread are probably also present in the hidden test data. Even if they have valid explanations, we are essentially being asked to predict human responses to a questionnaire - information that is inherently prone to inaccuracy and bias.\nThis makes it inherently impossible to develop a reliable model to solve the task, but you still have a chance to model the noise and inconsistencies and get a good score, but it's more a matter of luck...",
          "votes": 3
        }
      ]
    },
    {
      "id": 3009758,
      "postDate": "2024-10-08T09:48:16.090Z",
      "content": "<p>The first post contained my thoughts after looking at the competition description at first glance. Since then, I think I have come up with a reasonable explanation for some of the points, and others have posted a lot of other relevant observations and questions. So I have decided to summarise them all in one place and add to the end of the post (see \"Questioning summary\" part). If I missed something, please let me know!</p>\n<p>As <a href=\"https://www.kaggle.com/dan3dewey\" target=\"_blank\">@dan3dewey</a> mentioned below, the aim of the discussion was to understand the true objective of this competition and the practical implications of the required model. This doesn’t mean the data are useless - quite the opposite! I find them to have great value for data science and learning!</p>",
      "rawMarkdown": "The first post contained my thoughts after looking at the competition description at first glance. Since then, I think I have come up with a reasonable explanation for some of the points, and others have posted a lot of other relevant observations and questions. So I have decided to summarise them all in one place and add to the end of the post (see \"Questioning summary\" part). If I missed something, please let me know!\n\nAs @dan3dewey mentioned below, the aim of the discussion was to understand the true objective of this competition and the practical implications of the required model. This doesn’t mean the data are useless - quite the opposite! I find them to have great value for data science and learning!"
    },
    {
      "id": 3003623,
      "postDate": "2024-10-01T06:01:37.897Z",
      "content": "<p>You're raising important concerns about the potential redundancy of a PIU prediction model that includes Internet usage as a feature. Since PIU inherently involves excessive Internet use and related health or social problems, relying on Internet usage data seems circular. We can already suspect PIU when high Internet use and health issues are present.</p>\n<p>Additionally, the dataset includes participants with minimal Internet use but high PIU scores, which raises questions about data reliability or misinterpretation. It might be more effective to focus the study on participants who use the Internet more than average and predict impairment based on other behavioral or psychological factors. This would align better with the goal of detecting early signs of harmful Internet use.</p>",
      "rawMarkdown": "You're raising important concerns about the potential redundancy of a PIU prediction model that includes Internet usage as a feature. Since PIU inherently involves excessive Internet use and related health or social problems, relying on Internet usage data seems circular. We can already suspect PIU when high Internet use and health issues are present.\n\nAdditionally, the dataset includes participants with minimal Internet use but high PIU scores, which raises questions about data reliability or misinterpretation. It might be more effective to focus the study on participants who use the Internet more than average and predict impairment based on other behavioral or psychological factors. This would align better with the goal of detecting early signs of harmful Internet use."
    },
    {
      "id": 3003303,
      "postDate": "2024-09-30T21:12:04.853Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true,
      "replies": [
        {
          "id": 3003606,
          "postDate": "2024-10-01T05:28:52.520Z",
          "content": "<p>Indeed! Thanks for pointing that out! I had to carefully read all the value labels in the data_dict. I've corrected the wording in the post. </p>\n<p>Actually 3=More than 3 hrs/day is more important here - as it appears, it could be anything, 3, 5, 10 hrs/day… so I agree, if the need is to find relationships between internet use and health issues, why is this categorisation so narrow? </p>",
          "rawMarkdown": "Indeed! Thanks for pointing that out! I had to carefully read all the value labels in the data_dict. I've corrected the wording in the post. \n\nActually 3=More than 3 hrs/day is more important here - as it appears, it could be anything, 3, 5, 10 hrs/day... so I agree, if the need is to find relationships between internet use and health issues, why is this categorisation so narrow? "
        },
        {
          "id": 3023489,
          "postDate": "2024-10-20T16:04:31.190Z",
          "content": "<blockquote>\n  <p>How can people who barely use the internet have PIU?</p>\n</blockquote>\n<p>Here's what I think could be the case. These might be very busy people in their day jobs or other work and use the internet very little in a day. But when they do, it's problematic for example use of gambling apps, watching pornography, etc. If a person watches just 20-30 mins of pornography daily he/ she can have severe social issues by just using internet less than 1hr/day. It will be surely classified as severely PIU or high sii.</p>",
          "rawMarkdown": ">How can people who barely use the internet have PIU?\n\nHere's what I think could be the case. These might be very busy people in their day jobs or other work and use the internet very little in a day. But when they do, it's problematic for example use of gambling apps, watching pornography, etc. If a person watches just 20-30 mins of pornography daily he/ she can have severe social issues by just using internet less than 1hr/day. It will be surely classified as severely PIU or high sii.",
          "replies": [
            {
              "id": 3023934,
              "postDate": "2024-10-21T07:08:56.523Z",
              "content": "<p>This refers to risky use of the Internet. Even a small amount of time of risky use (e.g. if someone is abusive, posting compromising material, cyberbullying, trolling, etc.) is enough to get a PIU. Therefore, the time spent online may not matter much.</p>",
              "rawMarkdown": "This refers to risky use of the Internet. Even a small amount of time of risky use (e.g. if someone is abusive, posting compromising material, cyberbullying, trolling, etc.) is enough to get a PIU. Therefore, the time spent online may not matter much."
            },
            {
              "id": 3024038,
              "postDate": "2024-10-21T09:13:15.193Z",
              "content": "<p>I see your point and <a href=\"https://www.kaggle.com/lordpatil\" target=\"_blank\">@lordpatil</a>'s point. However, we shouldn't subjectively interpret what a high SII score means beyond the scope of the data. In this dataset, PIU severity (SII) is defined based on the PCIAT test scores falling into specific ranges. </p>\n<p>To achieve an SII score of 3 (PCIAT total score of 80 or more), a participant would need to score an average of at least 4 out of 5 on each of 20 question.</p>\n<p>If a participant reportedly uses the internet very little, say, less than an hour a day, it's logically inconsistent for them to exhibit behaviors such as:</p>\n<p>Neglecting household chores or schoolwork to spend more time online<br>\nPreferring to spend time online over family time<br>\nReporting that grades suffer because of the allowed amount of time online<br>\nShowing withdrawal symptoms when not online (several questions about this)</p>\n<p>Who would rate this high if it's 1 hour?</p>\n<p>So high SII scores in participants who barely use the internet are likely indicative of data anomalies rather than true representations of their internet use behaviors. </p>\n<p>But since my first post, I have come up with another explanation:  These discrepancies can be explained by considering that if children had a history of excessive Internet use leading to problematic behavior, and  their parents intervened by imposing strict time limits. As a result, data collected after these interventions would show lower current Internet use, even though symptoms of problematic Internet use (PIU) might still be reflected in SII scores.</p>\n<p>The thing is, both are equally possible since we do not know the history or timing of the administration of the questionnaire/data collection.</p>",
              "rawMarkdown": "I see your point and @lordpatil's point. However, we shouldn't subjectively interpret what a high SII score means beyond the scope of the data. In this dataset, PIU severity (SII) is defined based on the PCIAT test scores falling into specific ranges. \n\nTo achieve an SII score of 3 (PCIAT total score of 80 or more), a participant would need to score an average of at least 4 out of 5 on each of 20 question.\n\nIf a participant reportedly uses the internet very little, say, less than an hour a day, it's logically inconsistent for them to exhibit behaviors such as:\n\nNeglecting household chores or schoolwork to spend more time online\nPreferring to spend time online over family time\nReporting that grades suffer because of the allowed amount of time online\nShowing withdrawal symptoms when not online (several questions about this)\n\nWho would rate this high if it's 1 hour?\n\nSo high SII scores in participants who barely use the internet are likely indicative of data anomalies rather than true representations of their internet use behaviors. \n\nBut since my first post, I have come up with another explanation:  These discrepancies can be explained by considering that if children had a history of excessive Internet use leading to problematic behavior, and  their parents intervened by imposing strict time limits. As a result, data collected after these interventions would show lower current Internet use, even though symptoms of problematic Internet use (PIU) might still be reflected in SII scores.\n\nThe thing is, both are equally possible since we do not know the history or timing of the administration of the questionnaire/data collection.",
              "votes": 3
            },
            {
              "id": 3025055,
              "postDate": "2024-10-22T10:28:58.837Z",
              "content": "<p>The neglect of household chores may not be to spend more time on the Internet, but, for example, due to trauma resulting from risky/impulsive Internet use. The same is true for not taking family vacations or getting worse grades.</p>",
              "rawMarkdown": "The neglect of household chores may not be to spend more time on the Internet, but, for example, due to trauma resulting from risky/impulsive Internet use. The same is true for not taking family vacations or getting worse grades."
            }
          ]
        }
      ]
    },
    {
      "id": 2996838,
      "postDate": "2024-09-23T23:53:31.850Z",
      "rawMarkdown": "",
      "votes": -3,
      "isDeleted": true
    },
    {
      "id": 3001840,
      "postDate": "2024-09-29T10:31:59.440Z",
      "content": "<p>thanks a lot for this!</p>",
      "rawMarkdown": "thanks a lot for this!",
      "votes": 1
    },
    {
      "id": 3002669,
      "postDate": "2024-09-30T09:18:51.100Z",
      "content": "<p>thanks a lot off counting</p>",
      "rawMarkdown": "thanks a lot off counting",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 3005187,
      "author_name": "SSS",
      "author_url": "",
      "post_date": "2024-10-02T16:33:54.540000",
      "content": "<p><a href=\"https://www.kaggle.com/antoninadolgorukova\" target=\"_blank\">@antoninadolgorukova</a> I have read two of your posts and went through some folks EDA's. <br>\nDo we even need to question the task if everything is out of hands here?</p>\n<p>Here is my summary:</p>\n<ol>\n<li>There is noise in the data, arguably more noise than the actual signal (missing values, not strictly controlled environment during data collection, and so on - you named it).</li>\n<li>We are to predict the target which is a questionnaire response aggregation (I took some tests online - I failed most of them miserably and got severe score). Some of the features can be a proxy for the responses, though we might doubt that the internet usage is a root cause here.</li>\n<li>CV-LB thread shows \"random-walking\".</li>\n</ol>\n<p>In the light of the mentioned above, we are going to see highly overfitted public leaderboard here.<br>\nThe standard \"analytical toolbox\" (deep understanding of the data, feature engineering), which I guess is your competitive advantage, might be discouraged from the very beginning. It will be hard not to fell into ensembles trap early on and junky-whatever 5 submit a day scenario. </p>\n<p>It still can be a good competition for someone to test ds/ml skills, (data cleaning, wrangling, feature engineering, modeling and inference). Someone even get into the gold zone and will be encouraged to continue his/her DS/ML journey with the money prize. </p>\n<p>Do we really help to solve a real life problem or just playing ds/ml here? Will it be deserved in the same way as if it were well established task? if you really care, this competition might not be a good fit, otherwise it can be a fun. </p>",
      "votes": 14,
      "replies": [
        {
          "id": 3005508,
          "author_name": "Antonina Dolgorukova",
          "author_url": "",
          "post_date": "2024-10-03T02:50:12.777000",
          "content": "<p>Hi! Well said, agree, this competition feels more like playing with ML than solving a real problem. The data analysis and processing is quite interesting though, so I'll stick with that part). </p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 3007722,
          "author_name": "Daniel Dewey",
          "author_url": "",
          "post_date": "2024-10-05T16:43:40.817000",
          "content": "<p><a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a> \"Questioning the task\" can be seen as a form of trying to better understand the data available and how it relates to the target we want to predict. I don't think it is the same as saying: \"This is pointless, I'm not working on it, and neither should you.\" 🙂 </p>",
          "votes": 2,
          "replies": [
            {
              "id": 3008313,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-10-06T13:13:51.593000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 3078030,
          "author_name": "Sudhansu_IISC_Bangalore",
          "author_url": "",
          "post_date": "2024-12-21T17:41:42.243000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/antoninadolgorukova\" target=\"_blank\">@antoninadolgorukova</a>,</p>\n<p>You’ve summed up the challenges of this competition brilliantly. The noise, overfitting risks, and reliance on ensembles definitely make it tricky, but it’s also a great way to practice dealing with messy, real-world-like data. Whether it’s a meaningful problem or just ML “play” depends on one’s goals, but there’s still value in the learning experience. Thanks for sharing your perspective!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2995901,
      "author_name": "Marília Prata",
      "author_url": "",
      "post_date": "2024-09-22T18:15:52.237000",
      "content": "<p>Hi Antonina,</p>\n<p>Firstly,  excellent topic to be explored. I'd like to read other users approach. Their Data Science approach.</p>\n<p>The issue starts with \"The P\" (Problematic).  I'm not able to provide a Data Science point-of-view, due to my lack of knowledge. <br>\nHowever, I can only speak for myself and the many hours that I spend on Kaggle. Checking my colormap, the only day I was out (Feb, 2024), it was due that I didn't bring the computer to the hospital. </p>\n<p>That's a reflect of NOT HAVING a life cause I used to take care of someone.  Many times when I delivered more public work I was in hospital, when I had almost nothing to do, except to wait. Better stays kaggling on these times, than to think negatively.</p>\n<p>I confess that I burned many pans cause I simply forget that I was cooking something while I was kaggling. In fact, I intend to forget the WHOLE thing, which is bigger and much more Problematic than My Internet usage.</p>\n<p>On the other hand, I'm not totally unhealthy. Since, I'm retired, I can go to the beach and swim daily. Meanwhile, while I'm kaggling, I drink water and don't forget to go to the toilet. I stand-up, shake the body a little bit to change positions  too.  There are specific times that I prefer to watch  television,  to laugh a little bit and mostly when I begin to feel tired of thinking on programming languages, I start to \"abstract\" and stop to see what's the solution .</p>\n<p>Additionally, on my second specialization, the subject was \"Dental Health Promotion\" when I learned so many holistic concepts that I won't be able to forget to keep applying them in my life. </p>\n<p>The key word is Balance. Always  the \"Ancient Greek Temperance\".</p>\n<p>And for me, Kaggling is almost a synonym for The Internet. A Positive Internet usage. At least, I'm often learning substantial stuff. And mostly, keep my mind occupied with subjects that contribute to my personal growth.  </p>\n<p>I also would like to see and read more female Kagglers participations. They are so talented. And, when we advance in career, it's harder to have any time for ouselves to deliver or work on things that we really enjoy. In general, more than males.</p>\n<p>By the way, our Audience is just a bonus 😊</p>\n<p>Cheers, <br>\nMarília.</p>",
      "votes": 10,
      "replies": [
        {
          "id": 2996071,
          "author_name": "Antonina Dolgorukova",
          "author_url": "",
          "post_date": "2024-09-23T03:19:05.023000",
          "content": "<p>Thank you for sharing your personal experiences and thoughts. I absolutely agree with your emphasis on balance - using the internet productively, like on Kaggle, can be a positive and beneficial ☺️</p>\n<p>Personally, I spend 7-10 hours a day on my laptop, but I still find time to work out or swim. I wouldn't say I suffer from excessive internet use. However, I feel like I might not even notice when this becomes problematic… It seems like PIU is something that our relatives would be more likely to point out - similar to other mental health issues/behavioral addictions. But PIU appears to be surrounded by controversy and ambiguity:</p>\n<p>In fact, articles on this topic acknowledge the importance of the Internet in our lives, stating something like \"There is widespread agreement that the Internet can serve as a tool to enhance well-being, but…\". (here) or \"In the modern world, the Internet constitutes an important tool for work, especially during the SARS-CoV-2 pandemic. It is also used as a common pastime\" (<a href=\"https://annals-general-psychiatry.biomedcentral.com/articles/10.1186/s12991-022-00384-4\" target=\"_blank\">here</a>)… </p>\n<p>and more interesting: </p>\n<p>\"The diagnosis of PIU does not appear in any official diagnostic system, including DSM-V, and there are no widely accepted diagnostic criteria\" (<a href=\"https://www.sciencedirect.com/science/article/abs/pii/S0306460313002669\" target=\"_blank\">here</a>)</p>\n<p>\"The problem starts with terminology, as the appropriate name for the condition or behaviour often labelled “Internet addiction” is not clear. \" (<a href=\"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC2911083/\" target=\"_blank\">link</a>)</p>\n<p>\"Scientific understanding of PIU has lagged behind media attention mainly because of inconsistencies in defining PIU, disagreement about its very existence, and the variable methodological approaches used in studying it\" and </p>\n<p>\"Although there is still not enough empirical evidence to establish whether PIU is a clearly nosologically defined psychopathological entity, it is still widely accepted that it is a global mental health problem\" (<a href=\"https://www.sciencedirect.com/science/article/pii/S0160791X24001362\" target=\"_blank\">link</a>)</p>\n<p>Seems there’s ongoing debate about PIU existence, whether PIU should even be considered a distinct clinical disorder, or if it’s simply a symptom of underlying mental health issues. If we can’t even define the condition consistently, it’s hard to create standardized methods for identifying it.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2996354,
              "author_name": "Marília Prata",
              "author_url": "",
              "post_date": "2024-09-23T11:44:37.150000",
              "content": "<p>For the record, people that really have addition don't get cured (My opinion). They change their additions. Which initially could seem a cure is because they aren't reproducing the same behavior. Example, someone that have gambling issues could start to have problems on another field (like shopping, sex, eating disorders).</p>\n<p>On the Internet, many are addicted to likes/votes/medals. I can understand those that  can monetize their \"performances\" (e.g. influencers). However, the ones that are glad to have \"followers\" and many likes to their irrelevant posts (many with incorrect information), these persons are just searching for validation to their thoughts/ideas. And get some kind of attention that they won't have on their personal lives. Albeit, after a certain time, they'll realize that this behavior won't lead them to better positions or even anywhere. </p>\n<p>Thank you for the many links about this subject. </p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3000006,
      "author_name": "Antonina Dolgorukova",
      "author_url": "",
      "post_date": "2024-09-27T07:08:59.477000",
      "content": "<p>I've just realised: what we're essentially doing here is training a model to predict a score derived from potentially biased or 'unfair' responses, rather than an objective measure of impairment from internet use… </p>\n<p>The SII is calculated from the PCIAT test score, so in other words, we have to predict how test participants (or their parents) would answer this questionnaire😅. That's kinda weird in practical terms, but technically one can try to predict the answer to each question, then calculate the total score and the corresponding SII.</p>",
      "votes": 8,
      "replies": [
        {
          "id": 3000018,
          "author_name": "Tom Yuen",
          "author_url": "",
          "post_date": "2024-09-27T07:34:34.230000",
          "content": "<p>Yes, and the problem is how to estimate PCIAT_total.<br>\nI tried to skip those 20 PCIAT sub-questions since most of them are 0 or 1.<br>\nI tried to find some physical feature that's correlated to PCIAT_total, but not working well yet.<br>\nSome of the obvious one will be age, height, weight.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22136130%2Fe0f24b88e765fcbd9693fee36075ed1c%2FPCIAT_Fitness.png?generation=1727422040143130&amp;alt=media\" alt=\"\"></p>",
          "votes": 3,
          "replies": [
            {
              "id": 3001159,
              "author_name": "Antonina Dolgorukova",
              "author_url": "",
              "post_date": "2024-09-28T14:55:07.273000",
              "content": "<p>Interesting, will this be any better if you remove rows where any of the 20 PCIAT sub-questions are missing (and PCIAT_total is actually wrong as it is a sum of non-NA values)?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3012271,
              "author_name": "Aditya Shukla",
              "author_url": "",
              "post_date": "2024-10-08T20:34:46.670000",
              "content": "<p>Indeed. If we leave out the 20 PCIAT, the maximum correlation of 'sii' is ~35% with age, height, and weight—which I think is why most of us are getting model accuracy at 47% max. </p>\n<p>I thought of predicting some PCIAT features using parquet file info, but that is not complete; it only has about ~900 IDs. Can you think of any other ways to predict PCIAT features?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3012293,
              "author_name": "Antonina Dolgorukova",
              "author_url": "",
              "post_date": "2024-10-08T21:09:46.417000",
              "content": "<p>If you look at the questions corresponding to each PCIAT feature in data dictionary, you'll see that it's hard to imagine which of the provided features could possibly serve as predictors for many of them, like phone calls, or checking emails, for example 😔. Most of the questions assess behaviour or mood… But I didn't do this analysis yet,  may be I'll find smth, but I doubt.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3008342,
      "author_name": "Daniel Dewey",
      "author_url": "",
      "post_date": "2024-10-06T13:39:25.520000",
      "content": "<p><em>This comment has a \"questioning the task\" flavor so I'm posting it here.</em> </p>\n<p>Looking at a crosstab of my (not so well) predicted and actual sii values, I have a lot of actual sii=0 being classified as sii=1. One reason for that is that the PCIAT Total sii=0 boundary at 30 is very close to the peak in the distribution of the Total values (histogram below.)</p>\n<p>Having the boundary near the peak maximizes the confusion between the classes, and effectively cuts the around-the-peak class in two. For examples, better, data-based boundary locations on either side of the peak could be at 19-and-below, and/or 37-and-up. Though there may be other considerations in selecting the boundaries than ease of classification… 🙂</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1496178%2F0e79d497b4472182a502c8f73ecf16e0%2FDistribution_of_Total.jpg?generation=1728220007766912&amp;alt=media\" alt=\"Histogram of Total\"></p>",
      "votes": 6,
      "replies": [
        {
          "id": 3009636,
          "author_name": "Antonina Dolgorukova",
          "author_url": "",
          "post_date": "2024-10-08T06:37:57.597000",
          "content": "<p>Interesting point! I'm not into modelling, so don't follow how it goes, but I think I've seen people in other threads saying that regression works better than classification for this task, so the boundary issue could at least partially explain that. And what if you predict PCIAT_Total directly and then map the predicted values to SII, does that reduce the confusion between adjacent SII classes?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3002388,
      "author_name": "Daniel Dewey",
      "author_url": "",
      "post_date": "2024-09-30T02:03:48.060000",
      "content": "<p>Thanks for brining up this discussion and the responses. Yes, we are trying to predict a questionaire score using as predictors sometimes hard-to-get data values (csv) as well as the wrist-worn enmo, angles, and light. It does seem problematic 🥴</p>\n<p>Taking your comments into account, maybe a different competition would be based on predicting sii level based on data easily obtained in a single office visit (height, weight, age, BP, etc.) when the participant also picked up their wrist-worn device to return in a month… Though it would be more an assessment of \"healthy activive lifestyle\" than internet per se.</p>\n<p>P.S. <a href=\"https://www.kaggle.com/waechter\" target=\"_blank\">@waechter</a> , replace 'gaming' with 'kaggle' and I'm 8 for 9 in your list -- I'm OK with 7: my wife will say as I stare into space: \"You're thinking about data again, right?\" 😃</p>",
      "votes": 5,
      "replies": [
        {
          "id": 3003025,
          "author_name": "waechter",
          "author_url": "",
          "post_date": "2024-09-30T15:28:03.920000",
          "content": "<p>That's cute! I think the big difference between problematic usage and being passionnate is ashamed vs proud. I bet you don't lie to your wife about kaggle (7) and are proud of the time you spend here 😀</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2999149,
      "author_name": "Antonina Dolgorukova",
      "author_url": "",
      "post_date": "2024-09-26T12:24:32.573000",
      "content": "<p>By the way, here's another thought on the subject: I don't think the measures (data to be used as features for modelling) provided in this competition are \"easily obtainable\". The collection of data like accelerometry, sleep patterns, and physical activity (fitness, endurance) requires not only requires specialised equipment, but can also take days or even months to collect (the contestants wore accelerometers for up to 81 days) 🥲.</p>\n<hr>\n<p>UPD: The CGAS (Children's Global Assessment Scale) and several other features can only be assessed by a clinician. This blows my mind:</p>\n<p>PROBLEM FROM THE COMPETITION OWERVIEW:</p>\n<blockquote>\n  <p>Current methods for measuring problematic internet use in children and adolescents are often complex and require professional assessments. This creates access, cultural, and linguistic barriers for many families.</p>\n</blockquote>\n<p>SOLUTION SEEMS TO BE:</p>\n<p>training a model on &gt;50 features, some of which can only be assessed by a clinician and some of which take months to collect…</p>",
      "votes": 5,
      "replies": [
        {
          "id": 3018248,
          "author_name": "Satya Sai Deepak Velagapudi",
          "author_url": "",
          "post_date": "2024-10-15T16:22:08.670000",
          "content": "<p>I completely agree with you this is futile. This is in no way useful</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2995826,
      "author_name": "waechter",
      "author_url": "",
      "post_date": "2024-09-22T16:28:48.767000",
      "content": "<p>Interesting question, but I think that's a bit of a misconception. From what I understand, problematic usage/addiction is defined by negative consequences, not time or amounts spent.</p>\n<p>From <a href=\"https://en.wikipedia.org/wiki/Video_game_addiction\" target=\"_blank\">https://en.wikipedia.org/wiki/Video_game_addiction</a> there is 9 criteria to Internet Gaming Disorder:</p>\n<blockquote>\n  <ol>\n  <li>Pre-occupation. Do you spend a lot of time thinking about games even when you are not playing, or planning when you can play next?</li>\n  <li>Withdrawal. Do you feel restless, irritable, moody, angry, anxious or sad when attempting to cut down or stop gaming, or when you are unable to play?</li>\n  <li>Tolerance. Do you feel the need to play for increasing amounts of time, play more exciting games, or use more powerful equipment to get the same amount of excitement you used to get?</li>\n  <li>Reduce/stop. Do you feel that you should play less, but are unable to cut back on the amount of time you spend playing games?</li>\n  <li>Give up other activities. Do you lose interest in or reduce participation in other recreational activities due to gaming?</li>\n  <li>Continue despite problems. Do you continue to play games even though you are aware of negative consequences, such as not getting enough sleep, being late to school/work, spending too much money, having arguments with others, or neglecting important duties?</li>\n  <li>Deceive/cover up. Do you lie to family, friends or others about how much you game, or try to keep your family or friends from knowing how much you game?</li>\n  <li>Escape adverse moods. Do you game to escape from or forget about personal problems, or to relieve uncomfortable feelings such as guilt, anxiety, helplessness or depression?</li>\n  <li>Risk/lose relationships/opportunities. Do you risk or lose significant relationships, or job, educational or career opportunities because of gaming?</li>\n  </ol>\n</blockquote>\n<p>These questions are not about the time spent gaming/online, but how it affects life. </p>\n<p>For this competition we are trying to predict the <code>Severity Impairment Index</code>, which is the total of the responses from 20 questions <code>Parent-Child Internet Addiction Test</code>. Not the time spent online</p>\n<p>For example: </p>\n<ul>\n<li>Someone can be online 8h/day and have no problem, being a perfect child.</li>\n<li>And another be online way less, but with conflict with his parents and negative consequences</li>\n</ul>\n<p>Hope this make sense!</p>",
      "votes": 6,
      "replies": [
        {
          "id": 2995846,
          "author_name": "Antonina Dolgorukova",
          "author_url": "",
          "post_date": "2024-09-22T17:05:22.387000",
          "content": "<p>Thanks! I completely agree that PIU is characterized by negative consequences rather than just the amount of time spent online. </p>\n<p>Essentially, the SII is an index that can be calculated for a patient using the questionnaire as you mentioned. We are being asked to predict it using more objective data from physical observations (although there is data from other questionnaires). If we remove internet use from these data, there is nothing left that would allow us to link the participants' condition specifically to promlematic internet use.</p>\n<p>For example, in the training data, there are children who are not entirely healthy (e.g., overweight, sleep disturbance) who, it turns out, do not use the internet at all. And there are others with problems who report using it at least an hour a day. Both groups have health issues, but only the condition of those who use the internet can somehow be linked to PIU.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2995881,
              "author_name": "waechter",
              "author_url": "",
              "post_date": "2024-09-22T17:59:12.383000",
              "content": "<p>I understand your point, a children can be unhealthy due to other cause than internet usage.<br>\nBut I think Internet Usage=<code>PreInt_EduHx-computerinternet_hoursday</code> is misleading, because 0=Less than 1h/day, not don't use internet at all. <br>\nThere are rows in the training data with <code>ssi</code>&gt;=2 and <code>PreInt_EduHx-computerinternet_hoursday</code>==0</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2995905,
              "author_name": "Antonina Dolgorukova",
              "author_url": "",
              "post_date": "2024-09-22T18:19:50.647000",
              "content": "<p>Yep, high SSI and no internet use seems strange, since excessive internet use is by definition expected for a person with PIU. And overall, the data in this column is misleading, since the goal of this competition was formulated as</p>\n<blockquote>\n  <p>to detect early indicators of problematic Internet and technology use</p>\n</blockquote>\n<p>How could indicators of PIU be detected in participants who did not use the Internet at all or used it up to 3 hours a day (which I would say is less than usual today) ?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2995939,
              "author_name": "Antonina Dolgorukova",
              "author_url": "",
              "post_date": "2024-09-22T19:16:07.933000",
              "content": "<p>BTW, in the PCIAT test, based on which this SII was calculated, most questions are about internet use,  or directly related to it (but as you mentioned, not all participants use internet! how could they get high SII scores? it's rather investigator bias).</p>\n<p>I think if we imagine that we are predicting just some kind of severity index related to overall physical activity and well-being, it starts to make more sense. Suppose we get an accurate model that predicts this index. Now we can collect data (all the train features) from children in schools, for example, and predict the index. Children with high scores would need to be studied further to find the cause - it could be anything: genetics, poor nutrition, social circumstances, etc., or it could be excessive Internet use. So by analyzing this particular cohort with high SII scores and a history that indicates potential Internet addiction, we can detect PIU.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3003335,
              "author_name": "Tom Yuen",
              "author_url": "",
              "post_date": "2024-09-30T22:54:42.650000",
              "content": "<p>It wasn't told explicitly, but I guess the goal is to use model based on physical activities/ measures to replace the questionnaire in future.<br>\nThe psychologists understand that the PICAT questionnaire is not objective enough, which is based on parent's response.<br>\nAnd it depends on the content kids spending on internet, e.g. 3 hours spending on STEM is different form 1/2 hour spending on violence video.<br>\nI guess the time-series can provide some hidden objective insights to sii, e.g. if they have a lot of sunshine (light) and activities (X, Y, Z etc) on Saturday/ Sunday, parents tend to rate lower sii score. The only problem is  I've no idea how to extract those data from the time-series 🤣</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3003618,
              "author_name": "Antonina Dolgorukova",
              "author_url": "",
              "post_date": "2024-10-01T05:55:03.227000",
              "content": "<p>Thanks for yur comment! </p>\n<blockquote>\n  <p>the goal is to use model based on physical activities/ measures to replace the questionnaire in future.</p>\n</blockquote>\n<p>I believe that training a model to predict a priori noisy and imprecise questionnaire responses inevitably leads to a model that is even less reliable than the questionnaire itself.</p>\n<p>Regarding actigraphy data, you may find my notebook \"CMI-PIU: Actigraphy data EDA\" usefull, although the work on feature extraction is still in progress, I have already added some code for such features as duration and number of different activity patterns (no movement, low, moderate and vigorous physical activity) and features capturing the variation in activity over the 24-hour cycle).</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3017884,
      "author_name": "Danu A.",
      "author_url": "",
      "post_date": "2024-10-15T10:16:29.643000",
      "content": "<p>I think your question arrives because you are looking from a data analyst perspective or like a doctor or specialist who tries to deeply understand the whole bunch of data and get a final diagnosis. But that's about ML in general - to do the hard work of putting everything together and in a few seconds give a diagnosis that a specialist could take a closer look if it's of interest and save time (and $$$) by not analyzing all the lower class samples. So we (as data scientists) are spending 3 months by understanding the data and trying our best to construct a pretty decent model and then other specialists will save their time in the future, or we hope so at least…</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3018196,
          "author_name": "Antonina Dolgorukova",
          "author_url": "",
          "post_date": "2024-10-15T15:29:02.553000",
          "content": "<p>Hi! I appreciate your point about the value of machine learning (ML) in assisting specialists and potentially saving time and resources. And I'm specifically questioning the ML in the task and the practical implications of the results - the model. </p>\n<p>Look, when we have a model trained to predict human responses to the PCIAT questionnaire (which is essentially the SII), for new participants, we still need to collect all the features to run it and get predictions. The prediction will take a few seconds, agree, buy some of these features take months to collect, and some require visits to a clinician or special equipment.</p>\n<p>At the same time, a clinician could get answers to the PCIAT questionnaire directly in a single visit. So why would they need a model whose prediction is obviously less accurate and takes a lot of time?</p>",
          "votes": 4,
          "replies": [
            {
              "id": 3018232,
              "author_name": "Danu A.",
              "author_url": "",
              "post_date": "2024-10-15T16:16:43.750000",
              "content": "<p>There are some features that requires time/specialists, but there are also the time series that are very difficult and time consuming to analyze manually, but relatively easy to obtain (put the bracelet on and take it in a month), so one reason is to research what can we do with them. Another reason can be - to avoid additional time/resources for 20 tests that actually is our base target so if we could predict them or the final result - these tests will be necessary just for separate cases for specialist's use. So the overall idea is to research what can we do with all the data to cut expenses, optimize and hurry the evaluation process and maybe discover something additional on this path that could help in taking the decission.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3018428,
              "author_name": "Antonina Dolgorukova",
              "author_url": "",
              "post_date": "2024-10-15T18:43:04.040000",
              "content": "<p>The target in this competition is an index derived from the total scores of 20 questions. In real life, this index is calculated based on responses of parents to those 20 questions, just  this. The model you develop for this competition is trained to predict these scores. </p>\n<p>So to get this Sii,  a doctor can either ask a parent to answer 20 questions or collect &gt;50 measurements (features) plus actigraphy data and predict this Sii with a model. What would be the more accurate and easier way for both sides?</p>\n<p>May be I just don't see you point…</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 3039398,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-11-08T02:23:16.757000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3077971,
      "author_name": "Daniel Dewey",
      "author_url": "",
      "post_date": "2024-12-21T16:18:18.843000",
      "content": "<p><a href=\"https://www.kaggle.com/antoninadolgorukova\" target=\"_blank\">@antoninadolgorukova</a>  - I really appreciated your work and perspective on this competition! I looked for your submission, then I saw you only do bio-competitions - best of luck in your next one :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3006555,
      "author_name": "Arindam Roy",
      "author_url": "",
      "post_date": "2024-10-04T10:23:17.267000",
      "content": "<p>If we were predicting future health and social problems with current internet use, this challenge would have made more sense imo. But given these items are features, it is quite a futile exercise. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3004401,
      "author_name": "A. Nihat Uzunalioglu",
      "author_url": "",
      "post_date": "2024-10-01T19:19:40.100000",
      "content": "<p>guys I found out that most of the insights from the time-series (parquet files) don't match with the original tabular data. is it expected?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3005168,
          "author_name": "Sumit_08",
          "author_url": "",
          "post_date": "2024-10-02T15:56:06.583000",
          "content": "<p>It is not typically expected for insights from time-series parquet files to deviate significantly from the original tabular data unless there are issues during data processing or transformation.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3005510,
          "author_name": "Antonina Dolgorukova",
          "author_url": "",
          "post_date": "2024-10-03T03:04:15.983000",
          "content": "<p>It would be interesting to hear more details about what exactly doesn't match. I haven't yet compared these 2 data sources myself, but I would expect to see some agreement, for example, between the total activity level derived from actigraphy data and, say, the Activity Summary Score from the Physical Activity Questionnaire, and for sleep patterns and duration to match the SDS features in train.csv. However, agreement may depend on how the time series data have been cleaned and aggregated. And some discrepancies might be expected if the season of data collection varied.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3005850,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-10-03T12:25:52.273000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3005892,
              "author_name": "Antonina Dolgorukova",
              "author_url": "",
              "post_date": "2024-10-03T13:31:35.660000",
              "content": "<p>Is this CHatGPT answering 😱? <br>\nIf not, sorry, but what's the point? Your answer doesn't seem to be related to my comment to <a href=\"https://www.kaggle.com/godanu\" target=\"_blank\">@godanu</a> quesion or to the data we have.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3005956,
              "author_name": "Sumit_08",
              "author_url": "",
              "post_date": "2024-10-03T14:49:36.860000",
              "content": "<p>I apologize for any confusion. It seems I misunderstood the context of your comment. Let me clarify or provide more relevant information based on the specific data  you're referring to.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3006063,
              "author_name": "Lennart Haupts",
              "author_url": "",
              "post_date": "2024-10-03T17:27:13.727000",
              "content": "<p>Hey ChatGPT, do you perchance have any good cookie recipes?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3006066,
              "author_name": "Sumit_08",
              "author_url": "",
              "post_date": "2024-10-03T17:30:46.143000",
              "content": "<p><a href=\"https://chatgpt.com/share/66fed4b9-05b4-8001-b33e-a3ab61deaf93\" target=\"_blank\">https://chatgpt.com/share/66fed4b9-05b4-8001-b33e-a3ab61deaf93</a></p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 3007706,
          "author_name": "Daniel Dewey",
          "author_url": "",
          "post_date": "2024-10-05T16:30:00.703000",
          "content": "<p>The parquet data are an independent source of information from the tabular data, so I don't think there is any exact matching that we'd expect.<br>\n<em>Do the parquet data have any predictive power for the target value?</em> Here's a plot of the PCIAT Total score versus a first-cut feature, Wrist-pct_active, that I calculated from the parquet files. The correlation is -0.288 which is stronger than all but 6 of the csv tabular features. So there is some information in the parquet files that is related to the target value, and including parquet-derived features could help the model.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1496178%2Fa54350668784502478b89f4296807598%2FWrist_Total_vs_pct_active.jpg?generation=1728144923016716&amp;alt=media\" alt=\"PCIAT Total score versus Wrist-pct_active\"></p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 3001847,
      "author_name": "m4nocha",
      "author_url": "",
      "post_date": "2024-09-29T10:40:36.213000",
      "content": "<p>The approach is pretty nice but should i remove the outliers or replace it with mean/median?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3000567,
      "author_name": "Lucca rodriguez",
      "author_url": "",
      "post_date": "2024-09-27T19:17:44.347000",
      "content": "<p>I had not thought about this problem. thank you for the input! <br>\nI was just wondering how you would do to fix the data? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 3000921,
          "author_name": "Antonina Dolgorukova",
          "author_url": "",
          "post_date": "2024-09-28T08:53:59.687000",
          "content": "<p>Thanks! Idk, it seems the whole competition is in needs to be fixed, haha. The inconsistencies I point out in this thread are probably also present in the hidden test data. Even if they have valid explanations, we are essentially being asked to predict human responses to a questionnaire - information that is inherently prone to inaccuracy and bias.<br>\nThis makes it inherently impossible to develop a reliable model to solve the task, but you still have a chance to model the noise and inconsistencies and get a good score, but it's more a matter of luck…</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 3009758,
      "author_name": "Antonina Dolgorukova",
      "author_url": "",
      "post_date": "2024-10-08T09:48:16.090000",
      "content": "<p>The first post contained my thoughts after looking at the competition description at first glance. Since then, I think I have come up with a reasonable explanation for some of the points, and others have posted a lot of other relevant observations and questions. So I have decided to summarise them all in one place and add to the end of the post (see \"Questioning summary\" part). If I missed something, please let me know!</p>\n<p>As <a href=\"https://www.kaggle.com/dan3dewey\" target=\"_blank\">@dan3dewey</a> mentioned below, the aim of the discussion was to understand the true objective of this competition and the practical implications of the required model. This doesn’t mean the data are useless - quite the opposite! I find them to have great value for data science and learning!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3003623,
      "author_name": "Sumit_08",
      "author_url": "",
      "post_date": "2024-10-01T06:01:37.897000",
      "content": "<p>You're raising important concerns about the potential redundancy of a PIU prediction model that includes Internet usage as a feature. Since PIU inherently involves excessive Internet use and related health or social problems, relying on Internet usage data seems circular. We can already suspect PIU when high Internet use and health issues are present.</p>\n<p>Additionally, the dataset includes participants with minimal Internet use but high PIU scores, which raises questions about data reliability or misinterpretation. It might be more effective to focus the study on participants who use the Internet more than average and predict impairment based on other behavioral or psychological factors. This would align better with the goal of detecting early signs of harmful Internet use.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3003303,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-09-30T21:12:04.853000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 3003606,
          "author_name": "Antonina Dolgorukova",
          "author_url": "",
          "post_date": "2024-10-01T05:28:52.520000",
          "content": "<p>Indeed! Thanks for pointing that out! I had to carefully read all the value labels in the data_dict. I've corrected the wording in the post. </p>\n<p>Actually 3=More than 3 hrs/day is more important here - as it appears, it could be anything, 3, 5, 10 hrs/day… so I agree, if the need is to find relationships between internet use and health issues, why is this categorisation so narrow? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3023489,
          "author_name": "Chirag Patil",
          "author_url": "",
          "post_date": "2024-10-20T16:04:31.190000",
          "content": "<blockquote>\n  <p>How can people who barely use the internet have PIU?</p>\n</blockquote>\n<p>Here's what I think could be the case. These might be very busy people in their day jobs or other work and use the internet very little in a day. But when they do, it's problematic for example use of gambling apps, watching pornography, etc. If a person watches just 20-30 mins of pornography daily he/ she can have severe social issues by just using internet less than 1hr/day. It will be surely classified as severely PIU or high sii.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3023934,
              "author_name": "kalakagat",
              "author_url": "",
              "post_date": "2024-10-21T07:08:56.523000",
              "content": "<p>This refers to risky use of the Internet. Even a small amount of time of risky use (e.g. if someone is abusive, posting compromising material, cyberbullying, trolling, etc.) is enough to get a PIU. Therefore, the time spent online may not matter much.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3024038,
              "author_name": "Antonina Dolgorukova",
              "author_url": "",
              "post_date": "2024-10-21T09:13:15.193000",
              "content": "<p>I see your point and <a href=\"https://www.kaggle.com/lordpatil\" target=\"_blank\">@lordpatil</a>'s point. However, we shouldn't subjectively interpret what a high SII score means beyond the scope of the data. In this dataset, PIU severity (SII) is defined based on the PCIAT test scores falling into specific ranges. </p>\n<p>To achieve an SII score of 3 (PCIAT total score of 80 or more), a participant would need to score an average of at least 4 out of 5 on each of 20 question.</p>\n<p>If a participant reportedly uses the internet very little, say, less than an hour a day, it's logically inconsistent for them to exhibit behaviors such as:</p>\n<p>Neglecting household chores or schoolwork to spend more time online<br>\nPreferring to spend time online over family time<br>\nReporting that grades suffer because of the allowed amount of time online<br>\nShowing withdrawal symptoms when not online (several questions about this)</p>\n<p>Who would rate this high if it's 1 hour?</p>\n<p>So high SII scores in participants who barely use the internet are likely indicative of data anomalies rather than true representations of their internet use behaviors. </p>\n<p>But since my first post, I have come up with another explanation:  These discrepancies can be explained by considering that if children had a history of excessive Internet use leading to problematic behavior, and  their parents intervened by imposing strict time limits. As a result, data collected after these interventions would show lower current Internet use, even though symptoms of problematic Internet use (PIU) might still be reflected in SII scores.</p>\n<p>The thing is, both are equally possible since we do not know the history or timing of the administration of the questionnaire/data collection.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 3025055,
              "author_name": "kalakagat",
              "author_url": "",
              "post_date": "2024-10-22T10:28:58.837000",
              "content": "<p>The neglect of household chores may not be to spend more time on the Internet, but, for example, due to trauma resulting from risky/impulsive Internet use. The same is true for not taking family vacations or getting worse grades.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2996838,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-09-23T23:53:31.850000",
      "content": "",
      "votes": -3,
      "replies": []
    },
    {
      "id": 3001840,
      "author_name": "m4nocha",
      "author_url": "",
      "post_date": "2024-09-29T10:31:59.440000",
      "content": "<p>thanks a lot for this!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3002669,
      "author_name": "Rust Za",
      "author_url": "",
      "post_date": "2024-09-30T09:18:51.100000",
      "content": "<p>thanks a lot off counting</p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2995763": "Topic aim: to understand the real aim of this competition and the practical implications of the model required. If you are struggling with this like me, please feel free to add your questions. If you can prove me wrong, then please do so!\n\n**TL;TR: please see the summary (based on  the original post + comments + questions from other threads) at the end**\n\nIs a PIU prediction model redundant with Internet use as an input?\n\nProblematic Internet use (PIU) is defined as the use of the Internet that creates psychological, social, school and/or work difficulties in a person's life (thanks @mpwolke for the [article](https://www.sciencedirect.com/science/article/abs/pii/S0165178119320098)).\n\n**So you can't have PIU without excessive Internet use and without health or social problems.**\nI wonder what exactly is the need for a Severity Impairment Index (SII) prediction model that relies on features such as \"Internet usage\" (hours spent online per day)? If we need to collect internet usage and health-related data to make predictions, then we're already measuring what we're trying to predict, or what I'm not seeing here?\n\nIn practical terms:\n\n- For patients with high Internet use and health problems: We can already suspect problematic Internet use without needing a model, and include in the treatment plan not only medications for the health condition(s), but also recommendations for reducing Internet use and healthy lifestyles.\n\n- For patients with health problems but low Internet use: We can infer that their health problems have other causes and focus on diagnosing and treating them, again without the help of the model. Similarly, patients with low physical activity and unhealthy lifestyles need advice on both.\n\nOf course, self-reported Internet use can be inaccurate, but we don't have any other characteristics related to Internet use to train a model. So if we can't rely on reported Internet use, we can't properly diagnose PIU.\n\nFurthermore, any associations found between Internet use and health problems do not prove causation. The root cause may be poor health leading to increased internet use rather than the internet use causing health problems. \n\nI'm curious to hear your perspectives on this. Let's discuss! \n\n****\n\nAnother puzzling observation:\n\nThe competition description states that the goal is \"to detect early indicators of problematic Internet and technology use.\"\n\nGiven that, as noted in the comments below, PIU reflect the negative consequences of Internet use (when Internet use begins to cause problems), it seems logical to study early signs of PIU in a population of participants who use the Internet more than average.\n\nThis would mean recognizing subtle shifts in behavior, physical activity, or psychological well-being that could signal the onset of problematic internet use before it leads to significant impairment.\n\nAt the same time, 38.5% of the participants in the train dataset ~~do not use the internet at all~~ use internet less than hour a day. Moreover, a noticable proportion of participants with high SII scores ~~do not use the Internet~~ use internet less than hour a day (21.55% of participants with SII = 2, moderately impaired  by problematic internet use, and 14.7% of those with SII = 3, severely impaired)!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12193662%2Fbc48c089630fb2973007d1501260640a%2FScreenshot%202024-09-23%20080435.png?generation=1727067903360763&alt=media)\n\n[More details are here](https://www.kaggle.com/code/antoninadolgorukova/cmi-piu-features-eda)\n\nIf the index is intended to measure problematic internet use, it shouldn’t produce high scores for participants who ~~don’t use the internet at all~~ spend so little time online, should it? So is this investigator bias, unreliable self-reporting, or data collection error?\n\nAnd what is the point of including such a large proportion of people who ~~do not use the Internet at all~~ spend <1h/day in the internet when the goal is to detect early signs of harmfull internet usage?\n\nIn my opinion, Internet usage data should not be used as a feature for this task, but rather as a condition that participants must meet in order to be included in the study. In this way, we could develop a model to predict whether these individuals show signs of impairment and how severe that impairment is. This would be more consistent with the goals of detecting problematic Internet use.\n\n****\n**Questioning summary:**\n- Data collection for the model involves professional assessments, questionnaires and specialised equipment, adding complexity rather than simplifying the process for families.\n\n- Clinicians with CGAS scores, physical health measures and internet use data have sufficient information to diagnose and assess the severity of PIU in one visit. They could also simply administer the PCIAT during assessments to obtain more accurate SII scores than a predictive model would provide (which also requires much more data and time to collect).\n\n- As @expensivelunch mentioned [here](https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/538082), most actigraphy data were collected after the target outcome had been measured, representing future states. This introduces an inherent inaccuracy into the model.\n\n\nThese are not rhetorical questions, so I would really appreciate any comments from hosts  @gkiar07 or other interested people.",
    "3005187": "@antoninadolgorukova I have read two of your posts and went through some folks EDA's. \nDo we even need to question the task if everything is out of hands here?\n\nHere is my summary:\n 1. There is noise in the data, arguably more noise than the actual signal (missing values, not strictly controlled environment during data collection, and so on - you named it).\n 2. We are to predict the target which is a questionnaire response aggregation (I took some tests online - I failed most of them miserably and got severe score). Some of the features can be a proxy for the responses, though we might doubt that the internet usage is a root cause here.\n 3. CV-LB thread shows \"random-walking\".\n\nIn the light of the mentioned above, we are going to see highly overfitted public leaderboard here.\nThe standard \"analytical toolbox\" (deep understanding of the data, feature engineering), which I guess is your competitive advantage, might be discouraged from the very beginning. It will be hard not to fell into ensembles trap early on and junky-whatever 5 submit a day scenario. \n\nIt still can be a good competition for someone to test ds/ml skills, (data cleaning, wrangling, feature engineering, modeling and inference). Someone even get into the gold zone and will be encouraged to continue his/her DS/ML journey with the money prize. \n\nDo we really help to solve a real life problem or just playing ds/ml here? Will it be deserved in the same way as if it were well established task? if you really care, this competition might not be a good fit, otherwise it can be a fun. ",
    "2995901": "Hi Antonina,\n\nFirstly,  excellent topic to be explored. I'd like to read other users approach. Their Data Science approach.\n\nThe issue starts with \"The P\" (Problematic).  I'm not able to provide a Data Science point-of-view, due to my lack of knowledge. \nHowever, I can only speak for myself and the many hours that I spend on Kaggle. Checking my colormap, the only day I was out (Feb, 2024), it was due that I didn't bring the computer to the hospital. \n\nThat's a reflect of NOT HAVING a life cause I used to take care of someone.  Many times when I delivered more public work I was in hospital, when I had almost nothing to do, except to wait. Better stays kaggling on these times, than to think negatively.\n\nI confess that I burned many pans cause I simply forget that I was cooking something while I was kaggling. In fact, I intend to forget the WHOLE thing, which is bigger and much more Problematic than My Internet usage.\n\nOn the other hand, I'm not totally unhealthy. Since, I'm retired, I can go to the beach and swim daily. Meanwhile, while I'm kaggling, I drink water and don't forget to go to the toilet. I stand-up, shake the body a little bit to change positions  too.  There are specific times that I prefer to watch  television,  to laugh a little bit and mostly when I begin to feel tired of thinking on programming languages, I start to \"abstract\" and stop to see what's the solution .\n\nAdditionally, on my second specialization, the subject was \"Dental Health Promotion\" when I learned so many holistic concepts that I won't be able to forget to keep applying them in my life. \n\nThe key word is Balance. Always  the \"Ancient Greek Temperance\".\n\nAnd for me, Kaggling is almost a synonym for The Internet. A Positive Internet usage. At least, I'm often learning substantial stuff. And mostly, keep my mind occupied with subjects that contribute to my personal growth.  \n\nI also would like to see and read more female Kagglers participations. They are so talented. And, when we advance in career, it's harder to have any time for ouselves to deliver or work on things that we really enjoy. In general, more than males.\n\nBy the way, our Audience is just a bonus 😊\n\nCheers, \nMarília.",
    "3000006": "I've just realised: what we're essentially doing here is training a model to predict a score derived from potentially biased or 'unfair' responses, rather than an objective measure of impairment from internet use... \n\nThe SII is calculated from the PCIAT test score, so in other words, we have to predict how test participants (or their parents) would answer this questionnaire😅. That's kinda weird in practical terms, but technically one can try to predict the answer to each question, then calculate the total score and the corresponding SII.",
    "3008342": "*This comment has a \"questioning the task\" flavor so I'm posting it here.* \n\nLooking at a crosstab of my (not so well) predicted and actual sii values, I have a lot of actual sii=0 being classified as sii=1. One reason for that is that the PCIAT Total sii=0 boundary at 30 is very close to the peak in the distribution of the Total values (histogram below.)\n\nHaving the boundary near the peak maximizes the confusion between the classes, and effectively cuts the around-the-peak class in two. For examples, better, data-based boundary locations on either side of the peak could be at 19-and-below, and/or 37-and-up. Though there may be other considerations in selecting the boundaries than ease of classification... 🙂\n\n![Histogram of Total](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1496178%2F0e79d497b4472182a502c8f73ecf16e0%2FDistribution_of_Total.jpg?generation=1728220007766912&alt=media)",
    "3002388": "Thanks for brining up this discussion and the responses. Yes, we are trying to predict a questionaire score using as predictors sometimes hard-to-get data values (csv) as well as the wrist-worn enmo, angles, and light. It does seem problematic 🥴\n\nTaking your comments into account, maybe a different competition would be based on predicting sii level based on data easily obtained in a single office visit (height, weight, age, BP, etc.) when the participant also picked up their wrist-worn device to return in a month... Though it would be more an assessment of \"healthy activive lifestyle\" than internet per se.\n\nP.S. @waechter , replace 'gaming' with 'kaggle' and I'm 8 for 9 in your list -- I'm OK with 7: my wife will say as I stare into space: \"You're thinking about data again, right?\" 😃",
    "2999149": "By the way, here's another thought on the subject: I don't think the measures (data to be used as features for modelling) provided in this competition are \"easily obtainable\". The collection of data like accelerometry, sleep patterns, and physical activity (fitness, endurance) requires not only requires specialised equipment, but can also take days or even months to collect (the contestants wore accelerometers for up to 81 days) 🥲.\n\n****\nUPD: The CGAS (Children's Global Assessment Scale) and several other features can only be assessed by a clinician. This blows my mind:\n\nPROBLEM FROM THE COMPETITION OWERVIEW:\n>Current methods for measuring problematic internet use in children and adolescents are often complex and require professional assessments. This creates access, cultural, and linguistic barriers for many families.\n\nSOLUTION SEEMS TO BE:\n\ntraining a model on >50 features, some of which can only be assessed by a clinician and some of which take months to collect...",
    "2995826": "Interesting question, but I think that's a bit of a misconception. From what I understand, problematic usage/addiction is defined by negative consequences, not time or amounts spent.\n\nFrom https://en.wikipedia.org/wiki/Video_game_addiction there is 9 criteria to Internet Gaming Disorder:\n> 1. Pre-occupation. Do you spend a lot of time thinking about games even when you are not playing, or planning when you can play next?\n1. Withdrawal. Do you feel restless, irritable, moody, angry, anxious or sad when attempting to cut down or stop gaming, or when you are unable to play?\n1. Tolerance. Do you feel the need to play for increasing amounts of time, play more exciting games, or use more powerful equipment to get the same amount of excitement you used to get?\n1. Reduce/stop. Do you feel that you should play less, but are unable to cut back on the amount of time you spend playing games?\n1. Give up other activities. Do you lose interest in or reduce participation in other recreational activities due to gaming?\n1. Continue despite problems. Do you continue to play games even though you are aware of negative consequences, such as not getting enough sleep, being late to school/work, spending too much money, having arguments with others, or neglecting important duties?\n1. Deceive/cover up. Do you lie to family, friends or others about how much you game, or try to keep your family or friends from knowing how much you game?\n1. Escape adverse moods. Do you game to escape from or forget about personal problems, or to relieve uncomfortable feelings such as guilt, anxiety, helplessness or depression?\n1. Risk/lose relationships/opportunities. Do you risk or lose significant relationships, or job, educational or career opportunities because of gaming?\n\nThese questions are not about the time spent gaming/online, but how it affects life. \n\nFor this competition we are trying to predict the `Severity Impairment Index`, which is the total of the responses from 20 questions `Parent-Child Internet Addiction Test`. Not the time spent online\n\nFor example: \n- Someone can be online 8h/day and have no problem, being a perfect child.\n- And another be online way less, but with conflict with his parents and negative consequences\n\nHope this make sense!",
    "3017884": "I think your question arrives because you are looking from a data analyst perspective or like a doctor or specialist who tries to deeply understand the whole bunch of data and get a final diagnosis. But that's about ML in general - to do the hard work of putting everything together and in a few seconds give a diagnosis that a specialist could take a closer look if it's of interest and save time (and $$$) by not analyzing all the lower class samples. So we (as data scientists) are spending 3 months by understanding the data and trying our best to construct a pretty decent model and then other specialists will save their time in the future, or we hope so at least...",
    "3077971": "@antoninadolgorukova  - I really appreciated your work and perspective on this competition! I looked for your submission, then I saw you only do bio-competitions - best of luck in your next one :)",
    "3006555": "If we were predicting future health and social problems with current internet use, this challenge would have made more sense imo. But given these items are features, it is quite a futile exercise. ",
    "3004401": "guys I found out that most of the insights from the time-series (parquet files) don't match with the original tabular data. is it expected?",
    "3001847": "The approach is pretty nice but should i remove the outliers or replace it with mean/median?",
    "3000567": "I had not thought about this problem. thank you for the input! \nI was just wondering how you would do to fix the data? \n",
    "3009758": "The first post contained my thoughts after looking at the competition description at first glance. Since then, I think I have come up with a reasonable explanation for some of the points, and others have posted a lot of other relevant observations and questions. So I have decided to summarise them all in one place and add to the end of the post (see \"Questioning summary\" part). If I missed something, please let me know!\n\nAs @dan3dewey mentioned below, the aim of the discussion was to understand the true objective of this competition and the practical implications of the required model. This doesn’t mean the data are useless - quite the opposite! I find them to have great value for data science and learning!",
    "3003623": "You're raising important concerns about the potential redundancy of a PIU prediction model that includes Internet usage as a feature. Since PIU inherently involves excessive Internet use and related health or social problems, relying on Internet usage data seems circular. We can already suspect PIU when high Internet use and health issues are present.\n\nAdditionally, the dataset includes participants with minimal Internet use but high PIU scores, which raises questions about data reliability or misinterpretation. It might be more effective to focus the study on participants who use the Internet more than average and predict impairment based on other behavioral or psychological factors. This would align better with the goal of detecting early signs of harmful Internet use.",
    "3003303": "",
    "2996838": "",
    "3001840": "thanks a lot for this!",
    "3002669": "thanks a lot off counting"
  }
}