{
  "id": 538235,
  "title": "Addressing some of the challenges of our dataset",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/538235",
  "author_name": "Greg Kiar",
  "post_date": "2024-10-07T18:27:17.890000",
  "votes": 73,
  "comment_count": 46,
  "views": 0,
  "content": "<p>Hi, Kagglers!</p>\n<p>It's wonderful to see so much engagement already in our competition. Firstly, on behalf of all of our organizing team, we just want to say thank you for joining us in this challenge!</p>\n<p>Over the last couple of weeks we've noticed a handful of discussion threads pop-up that touch on some of the same core elements of our dataset, and so I wanted to speak to some of those details directly, and provide more context about the nature of our data.</p>\n<p>In particular: <strong>yes, the data is noisy, biased, incomplete, and subject to a variety of other off-target sources of variation.</strong> This heterogeneity is at the heart of our challenge, and the heart of large-scale psychiatric research more broadly. Many elements of the phenomena we try to measure are subjective, and are influenced by both systemic bias (e.g., accessibility) and random variation (e.g., time of day, last meal). Indeed, some biases will be related directly to the signal we care about (e.g., parental report of problematic internet use will be biased by the parents' own internet use), and in other cases, they may be entirely unrelated. Indeed, many of our sample are not severely impacted, and that is consistent with our observed prevalence of this issue in our community. Overfitting will be easy, with all of these imbalances and sources of noise.</p>\n<p>The list of these sorts of relationships and the challenges they each introduce goes on and on; our question is ultimately <em>how can we make use of these data, regardless?</em> In fact, these challenges are precisely <strong><em>how our dataset represents the real-world</em></strong> context that we are trying to model and address.</p>\n<p>We deliberately did not want to ship a sanitized version of this data for you to work from, because then we would not be able to apply what you learn in practice. We want to collectively learn about the hidden relationships in our data and how they can jointly be used to model problematic internet use in children and adolescents. In order to do so, we needed to show you what we see: a vast portrait with a lot of signal, but also some noise.</p>\n<p>If you want to hear a bit more about our dataset and challenge, join for the <a href=\"https://events.dell.com/event/c87be06c-a180-4c88-b9bf-7fe6f365d43a/register\" target=\"_blank\">Webinar that Dell Technologies is hosting at the end of the week</a>.</p>\n<p><strong><em>Edit:</em></strong> <a href=\"https://vimeo.com/1020745085\" target=\"_blank\">Here is a link of the recording</a>. </p>\n<p>Thanks again for joining us, and we're excited to continue learning about our data with you! 🚀</p>\n<p>-Greg</p>",
  "messages": [
    {
      "id": 3009335,
      "postDate": "2024-10-07T18:27:17.890Z",
      "content": "<p>Hi, Kagglers!</p>\n<p>It's wonderful to see so much engagement already in our competition. Firstly, on behalf of all of our organizing team, we just want to say thank you for joining us in this challenge!</p>\n<p>Over the last couple of weeks we've noticed a handful of discussion threads pop-up that touch on some of the same core elements of our dataset, and so I wanted to speak to some of those details directly, and provide more context about the nature of our data.</p>\n<p>In particular: <strong>yes, the data is noisy, biased, incomplete, and subject to a variety of other off-target sources of variation.</strong> This heterogeneity is at the heart of our challenge, and the heart of large-scale psychiatric research more broadly. Many elements of the phenomena we try to measure are subjective, and are influenced by both systemic bias (e.g., accessibility) and random variation (e.g., time of day, last meal). Indeed, some biases will be related directly to the signal we care about (e.g., parental report of problematic internet use will be biased by the parents' own internet use), and in other cases, they may be entirely unrelated. Indeed, many of our sample are not severely impacted, and that is consistent with our observed prevalence of this issue in our community. Overfitting will be easy, with all of these imbalances and sources of noise.</p>\n<p>The list of these sorts of relationships and the challenges they each introduce goes on and on; our question is ultimately <em>how can we make use of these data, regardless?</em> In fact, these challenges are precisely <strong><em>how our dataset represents the real-world</em></strong> context that we are trying to model and address.</p>\n<p>We deliberately did not want to ship a sanitized version of this data for you to work from, because then we would not be able to apply what you learn in practice. We want to collectively learn about the hidden relationships in our data and how they can jointly be used to model problematic internet use in children and adolescents. In order to do so, we needed to show you what we see: a vast portrait with a lot of signal, but also some noise.</p>\n<p>If you want to hear a bit more about our dataset and challenge, join for the <a href=\"https://events.dell.com/event/c87be06c-a180-4c88-b9bf-7fe6f365d43a/register\" target=\"_blank\">Webinar that Dell Technologies is hosting at the end of the week</a>.</p>\n<p><strong><em>Edit:</em></strong> <a href=\"https://vimeo.com/1020745085\" target=\"_blank\">Here is a link of the recording</a>. </p>\n<p>Thanks again for joining us, and we're excited to continue learning about our data with you! 🚀</p>\n<p>-Greg</p>",
      "rawMarkdown": "Hi, Kagglers!\n\nIt's wonderful to see so much engagement already in our competition. Firstly, on behalf of all of our organizing team, we just want to say thank you for joining us in this challenge!\n\nOver the last couple of weeks we've noticed a handful of discussion threads pop-up that touch on some of the same core elements of our dataset, and so I wanted to speak to some of those details directly, and provide more context about the nature of our data.\n\nIn particular: **yes, the data is noisy, biased, incomplete, and subject to a variety of other off-target sources of variation.** This heterogeneity is at the heart of our challenge, and the heart of large-scale psychiatric research more broadly. Many elements of the phenomena we try to measure are subjective, and are influenced by both systemic bias (e.g., accessibility) and random variation (e.g., time of day, last meal). Indeed, some biases will be related directly to the signal we care about (e.g., parental report of problematic internet use will be biased by the parents' own internet use), and in other cases, they may be entirely unrelated. Indeed, many of our sample are not severely impacted, and that is consistent with our observed prevalence of this issue in our community. Overfitting will be easy, with all of these imbalances and sources of noise.\n\nThe list of these sorts of relationships and the challenges they each introduce goes on and on; our question is ultimately *how can we make use of these data, regardless?* In fact, these challenges are precisely ***how our dataset represents the real-world*** context that we are trying to model and address.\n\nWe deliberately did not want to ship a sanitized version of this data for you to work from, because then we would not be able to apply what you learn in practice. We want to collectively learn about the hidden relationships in our data and how they can jointly be used to model problematic internet use in children and adolescents. In order to do so, we needed to show you what we see: a vast portrait with a lot of signal, but also some noise.\n\nIf you want to hear a bit more about our dataset and challenge, join for the [Webinar that Dell Technologies is hosting at the end of the week](https://events.dell.com/event/c87be06c-a180-4c88-b9bf-7fe6f365d43a/register).\n\n***Edit:** [Here is a link of the recording]( https://vimeo.com/1020745085).* \n\nThanks again for joining us, and we're excited to continue learning about our data with you! :rocket:\n\n-Greg",
      "votes": 72
    },
    {
      "id": 3009735,
      "postDate": "2024-10-08T09:13:39.227Z",
      "content": "<p>Thanks for your post! I was questioning the practical implications of the model (\"Questioning the task\" thread), but now I think I can just focus on this new goal: finding out how we can make use of this data. </p>\n<p>Many of us have already realised that we need to do a lot of data cleaning. Some variables require domain knowledge to clean (e.g. unrealistic blood pressure, heart rate, BMI), but that seems to be the easiest part - finding normal ranges and removing outliers or replacing them with imputation methods. For the rest, we need more information to find normal ranges. It would also help a lot if we could better understand the nature of the data. For now, I think it would be very helpful if you could answer these questions:</p>\n<p>1) FitnessGram Vitals and Treadmill: Could you please share the methodology - what was the test? what devide was used? could you direct us to where we can read about it (maybe there is a user manual?), or just give reference values for age groups?<br>\n2) FitnessGram Child - Zones: Could you explain how the values were assigned to the zones?<br>\n3) Bioelectrical Impedance Analysis: What machine was used (would be great to see the user manual). Can you point us to some resources for finding normal ranges specific to the machine used to collect data for this competition?</p>\n<p>UPD: for BIA - What do the values represent - raw data or did you use some BIA equation models to estimate muscle mass and other parameters? If so, which equation was used?</p>\n<p>4) Physical activity questionnaires: What questions were asked?<br>\n5) Sleep disturbance scale: How was it measured and where can we get some information on its interpretation?</p>",
      "rawMarkdown": "Thanks for your post! I was questioning the practical implications of the model (\"Questioning the task\" thread), but now I think I can just focus on this new goal: finding out how we can make use of this data. \n\nMany of us have already realised that we need to do a lot of data cleaning. Some variables require domain knowledge to clean (e.g. unrealistic blood pressure, heart rate, BMI), but that seems to be the easiest part - finding normal ranges and removing outliers or replacing them with imputation methods. For the rest, we need more information to find normal ranges. It would also help a lot if we could better understand the nature of the data. For now, I think it would be very helpful if you could answer these questions:\n\n1) FitnessGram Vitals and Treadmill: Could you please share the methodology - what was the test? what devide was used? could you direct us to where we can read about it (maybe there is a user manual?), or just give reference values for age groups?\n2) FitnessGram Child - Zones: Could you explain how the values were assigned to the zones?\n3) Bioelectrical Impedance Analysis: What machine was used (would be great to see the user manual). Can you point us to some resources for finding normal ranges specific to the machine used to collect data for this competition?\n\nUPD: for BIA - What do the values represent - raw data or did you use some BIA equation models to estimate muscle mass and other parameters? If so, which equation was used?\n\n4) Physical activity questionnaires: What questions were asked?\n5) Sleep disturbance scale: How was it measured and where can we get some information on its interpretation?",
      "votes": 19,
      "replies": [
        {
          "id": 3023850,
          "postDate": "2024-10-21T05:13:40.833Z",
          "content": "<p>Good points. Thank you</p>",
          "rawMarkdown": "Good points. Thank you",
          "votes": 1
        },
        {
          "id": 3029198,
          "postDate": "2024-10-27T02:56:43.250Z",
          "content": "<p><a href=\"https://www.kaggle.com/antoninadolgorukova\" target=\"_blank\">@antoninadolgorukova</a> were any of your questions answered? I am also very curious how those values were obtained. Thanks!</p>",
          "rawMarkdown": "@antoninadolgorukova were any of your questions answered? I am also very curious how those values were obtained. Thanks!",
          "votes": 2,
          "replies": [
            {
              "id": 3029733,
              "postDate": "2024-10-27T16:56:40.393Z",
              "content": "<p>No, nobody answered these questions as well as the ones below in my separate comment…</p>",
              "rawMarkdown": "No, nobody answered these questions as well as the ones below in my separate comment..."
            }
          ]
        }
      ]
    },
    {
      "id": 3009630,
      "postDate": "2024-10-08T06:17:07.597Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/gkiar07\" target=\"_blank\">@gkiar07</a>, thank you for this clarification! There is one question many Kagglers have in this competition; perhaps you're willing to answer it: How did you split the data into Kaggle's training and test datasets? Is it a completely random split, or do training and test represent different batches of your data acquisition?</p>",
      "rawMarkdown": "Hi @gkiar07, thank you for this clarification! There is one question many Kagglers have in this competition; perhaps you're willing to answer it: How did you split the data into Kaggle's training and test datasets? Is it a completely random split, or do training and test represent different batches of your data acquisition?",
      "votes": 12,
      "replies": [
        {
          "id": 3017447,
          "postDate": "2024-10-14T23:10:06.747Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 3032429,
          "postDate": "2024-10-30T21:48:22.867Z",
          "content": "<p>Hello. I'd like to ask - why is it reported and why there's a deleted comment? Was there any leakage of information that someone could have seen already?  <a href=\"https://www.kaggle.com/gkiar07\" target=\"_blank\">@gkiar07</a></p>",
          "rawMarkdown": "Hello. I'd like to ask - why is it reported and why there's a deleted comment? Was there any leakage of information that someone could have seen already?  @gkiar07",
          "votes": 1,
          "replies": [
            {
              "id": 3073775,
              "postDate": "2024-12-16T22:37:45.713Z",
              "content": "<p>Is there a reason that the discussion of whether the train/test split was random keeps coming up. Just trying to understand how people would respond if the split were not random..</p>",
              "rawMarkdown": "Is there a reason that the discussion of whether the train/test split was random keeps coming up. Just trying to understand how people would respond if the split were not random.."
            }
          ]
        }
      ]
    },
    {
      "id": 3041343,
      "postDate": "2024-11-10T07:43:12.160Z",
      "content": "<p>It wouldn't make sense if the data wasn't this badly noisy and incomplete. That's sort of expected, and only make the challenge more exciting tbh.</p>",
      "rawMarkdown": "It wouldn't make sense if the data wasn't this badly noisy and incomplete. That's sort of expected, and only make the challenge more exciting tbh.",
      "votes": 4,
      "replies": [
        {
          "id": 3041346,
          "postDate": "2024-11-10T07:48:22.613Z",
          "content": "<p>I think same. If the data was not this badly noisy and incomplete then there was no reason for this competition to have prize money. It has prize money to attract data scientists to solve this problem. I will only say that there should have a been a little more explanation of when the data was collected, because carefully looking at data it shows that different parts of data for a single person were collected at different ages.</p>",
          "rawMarkdown": "I think same. If the data was not this badly noisy and incomplete then there was no reason for this competition to have prize money. It has prize money to attract data scientists to solve this problem. I will only say that there should have a been a little more explanation of when the data was collected, because carefully looking at data it shows that different parts of data for a single person were collected at different ages."
        },
        {
          "id": 3041478,
          "postDate": "2024-11-10T12:25:46.173Z",
          "content": "<p>Completely agree! The complexity adds to the challenge, but clearer details about the data collection timeline would have provided a better foundation for tackling the inconsistencies.</p>",
          "rawMarkdown": "Completely agree! The complexity adds to the challenge, but clearer details about the data collection timeline would have provided a better foundation for tackling the inconsistencies."
        }
      ]
    },
    {
      "id": 3109184,
      "postDate": "2025-01-28T15:24:44.720Z",
      "content": "<p>teşekkürler</p>",
      "rawMarkdown": "teşekkürler",
      "votes": 2
    },
    {
      "id": 3009986,
      "postDate": "2024-10-08T15:08:53.267Z",
      "content": "<p>Thanks for acknowleging the challenges of the data — though we’ve come to expect that from Kaggle :)</p>\n<p>Re: “We want to collectively learn about the hidden relationships in our data…” — this is a good goal and/but it is unspervised learning and so it doesn’t have a clear, objective metric. (A peers’-subjective metric might have levels: “meh”, “interesting”, “cool”, and “wow!”.)</p>\n<p>“.. and how they can jointly be used to model [predict] problematic internet use in children and adolescents.” This part is supervised learning and requires a “ground truth” target. As the “Questioning the Task” discussion bought out, the “ground truth” target here is actually a parent’s (hastily? not-at-all?) filled-out questionaire rather than some expert’s assessment based on discussions, observations, etc. So we’re really predicting one (important) feature from the others rather than predicting the “experts’ sii” from all the features.</p>\n<p>As many have said, none of that really matters to those of us with high levels of Problematic Kaggle Use —  Thanks for the data!   :)</p>",
      "rawMarkdown": "Thanks for acknowleging the challenges of the data — though we’ve come to expect that from Kaggle :)\n\nRe: “We want to collectively learn about the hidden relationships in our data…” — this is a good goal and/but it is unspervised learning and so it doesn’t have a clear, objective metric. (A peers’-subjective metric might have levels: “meh”, “interesting”, “cool”, and “wow!”.)\n\n“.. and how they can jointly be used to model [predict] problematic internet use in children and adolescents.” This part is supervised learning and requires a “ground truth” target. As the “Questioning the Task” discussion bought out, the “ground truth” target here is actually a parent’s (hastily? not-at-all?) filled-out questionaire rather than some expert’s assessment based on discussions, observations, etc. So we’re really predicting one (important) feature from the others rather than predicting the “experts’ sii” from all the features.\n\nAs many have said, none of that really matters to those of us with high levels of Problematic Kaggle Use —  Thanks for the data!   :)",
      "votes": 3,
      "replies": [
        {
          "id": 3012528,
          "postDate": "2024-10-09T06:21:35.110Z",
          "content": "<p>I like the way you put it: \"we really do predict one (important) feature from the others\". I still hope to hear justification/explanation from the hosts though .</p>\n<p>Indeed, with the target provided, we can only predict how parents would rate the severity of PIU in their children, i.e., kind of a level of parental worry or concern.</p>\n<p>Additionally, the applicability of the PCIAT questionnaire across the age range is questionable. I think all the questions in the PCIAT are much more suitable for adolescents. For example:</p>\n<ul>\n<li>A 5-7-year-old may not have household chores, as this depends on cultural norms.</li>\n<li>The question about academic impact may not apply to younger children who are not yet in school or graduated adults.</li>\n<li>Email use and receiving phone calls from \"online friends\" seem out of context for young 5-6 y.o. children too.</li>\n<li>Questions about reaction to the time allowed to spend on the internet (there are at least 3 of them) are not applicable to adults… usually :).</li>\n</ul>\n<p>Given this, if we see, for example, that adolescents seem to have the highest SII across all levels of Internet use (added to my \"Features EDA\" notebook, the assessment is rough as there is huge variability, but it illustrates my point well)… So are they more susceptible to PIU, or is this questionnaire just more sensitive to PIU in this age group? </p>\n<p>BTW, did the adults in this data fill in the questionnaire themselves, or were their parents asked 🙃?</p>",
          "rawMarkdown": "I like the way you put it: \"we really do predict one (important) feature from the others\". I still hope to hear justification/explanation from the hosts though .\n\nIndeed, with the target provided, we can only predict how parents would rate the severity of PIU in their children, i.e., kind of a level of parental worry or concern.\n\nAdditionally, the applicability of the PCIAT questionnaire across the age range is questionable. I think all the questions in the PCIAT are much more suitable for adolescents. For example:\n\n- A 5-7-year-old may not have household chores, as this depends on cultural norms.\n- The question about academic impact may not apply to younger children who are not yet in school or graduated adults.\n- Email use and receiving phone calls from \"online friends\" seem out of context for young 5-6 y.o. children too.\n- Questions about reaction to the time allowed to spend on the internet (there are at least 3 of them) are not applicable to adults... usually :).\n\nGiven this, if we see, for example, that adolescents seem to have the highest SII across all levels of Internet use (added to my \"Features EDA\" notebook, the assessment is rough as there is huge variability, but it illustrates my point well)... So are they more susceptible to PIU, or is this questionnaire just more sensitive to PIU in this age group? \n\nBTW, did the adults in this data fill in the questionnaire themselves, or were their parents asked 🙃?",
          "votes": 4,
          "replies": [
            {
              "id": 3012938,
              "postDate": "2024-10-09T14:04:58.693Z",
              "content": "<p>… and the questions feel dated, from the late 1990s or early 2000s? I was a parent then :)</p>",
              "rawMarkdown": "… and the questions feel dated, from the late 1990s or early 2000s? I was a parent then :)"
            }
          ]
        }
      ]
    },
    {
      "id": 3009749,
      "postDate": "2024-10-08T09:31:26.520Z",
      "content": "<p>it also would be great if you could comment on these:</p>\n<p>1) Participants with data in the children's Physical Activity Questionnaire columns (e.g., PAQ_C_Total) are aged 7 to 17, which overlaps with those in the adolescents' columns (PAQ_A_Total), who are aged 13 to 18. One participant, aged 13, appeared to complete both questionnaires. Could you clarify?</p>\n<p>2) Endurance measures (FitnessGram Vitals and Treadmill tests) were only conducted in children aged 5-12 years. Is there a specific reason for this, or did it just happen that way?</p>\n<p>3) Some questions in the PCIAT test appear to have been skipped by respondents (resulting in missing values in the PCIAT columns), but the Total score is still calculated as the sum of the non-missing values. This could potentially lead to invalid SII scores. </p>\n<p>4) A great question from <a href=\"https://www.kaggle.com/rafidmahbub\" target=\"_blank\">@rafidmahbub</a> in the thread I mentioned: There are a lot of high blood pressure (BP) and heart rate (HR) values. Should we interpret these as pathological, or were they taken during exercise? Basically, what do the BP and HR data represent - exercise response tests or usual resting values?</p>\n<p>This are the main issues posted in another thread (\"Some findings from the features EDA (+ data issues, outliers)\")</p>",
      "rawMarkdown": "it also would be great if you could comment on these:\n\n1) Participants with data in the children's Physical Activity Questionnaire columns (e.g., PAQ_C_Total) are aged 7 to 17, which overlaps with those in the adolescents' columns (PAQ_A_Total), who are aged 13 to 18. One participant, aged 13, appeared to complete both questionnaires. Could you clarify?\n\n2) Endurance measures (FitnessGram Vitals and Treadmill tests) were only conducted in children aged 5-12 years. Is there a specific reason for this, or did it just happen that way?\n\n3) Some questions in the PCIAT test appear to have been skipped by respondents (resulting in missing values in the PCIAT columns), but the Total score is still calculated as the sum of the non-missing values. This could potentially lead to invalid SII scores. \n\n4) A great question from @rafidmahbub in the thread I mentioned: There are a lot of high blood pressure (BP) and heart rate (HR) values. Should we interpret these as pathological, or were they taken during exercise? Basically, what do the BP and HR data represent - exercise response tests or usual resting values?\n\nThis are the main issues posted in another thread (\"Some findings from the features EDA (+ data issues, outliers)\")",
      "votes": 3,
      "replies": [
        {
          "id": 3061829,
          "postDate": "2024-12-03T02:32:53.150Z",
          "content": "<p>On your 3rd point, it may make sense to impute the missing values and then calculate an adjusted SII.</p>",
          "rawMarkdown": "On your 3rd point, it may make sense to impute the missing values and then calculate an adjusted SII."
        }
      ]
    },
    {
      "id": 3047668,
      "postDate": "2024-11-17T03:08:27.407Z",
      "content": "<p>Thanks for the information. Could you please clarify if train/test split has been done completely randomly or not?</p>",
      "rawMarkdown": "Thanks for the information. Could you please clarify if train/test split has been done completely randomly or not?",
      "votes": 1
    },
    {
      "id": 3018467,
      "postDate": "2024-10-15T19:03:02.547Z",
      "content": "<p>Thanks for engaging in this project. I appreciate the context of what you are trying to do. It also gives us data scientists who are trying to improve our craft excellent real world experience.</p>",
      "rawMarkdown": "Thanks for engaging in this project. I appreciate the context of what you are trying to do. It also gives us data scientists who are trying to improve our craft excellent real world experience.",
      "votes": 1
    },
    {
      "id": 3013798,
      "postDate": "2024-10-10T14:41:57.703Z",
      "content": "<p>Thank you for this post :)<br>\nyour data need a lot of cleansing , and this affected  the reliability of it , in some cases explain more details about method of data gathering , standard ranges and ages is helpful <br>\nadditionally, if you use any grouping boarder(some thing like obesity, fat, normal, etc..) or use a standard, please explain , I can't find nothing </p>",
      "rawMarkdown": "Thank you for this post :)\nyour data need a lot of cleansing , and this affected  the reliability of it , in some cases explain more details about method of data gathering , standard ranges and ages is helpful \nadditionally, if you use any grouping boarder(some thing like obesity, fat, normal, etc..) or use a standard, please explain , I can't find nothing ",
      "votes": 1
    },
    {
      "id": 3023845,
      "postDate": "2024-10-21T05:08:30.337Z",
      "content": "<p>Thank you for the detailed explanation.</p>\n<p>Based on the provided details it is clear that reaching a high accuracy in predictions is the real challenge so, I would like to know that how much score on the LB will be considered a high score? keeping in view the current condition of the data.</p>\n<p>I mean what score will be satisfactory to the organizers of the competition? Currently maximum scores are going around 0.49+. The models which are generating these scores are these sufficient to be used in a practical scenario or do you think a much higher score is needed for using the solution in practical scenarios.  </p>\n<p>I know \"the higher the better\" but keeping in view the current condition of data, what score will qualify to be practical?</p>\n<p>Thank you</p>",
      "rawMarkdown": "Thank you for the detailed explanation.\n\nBased on the provided details it is clear that reaching a high accuracy in predictions is the real challenge so, I would like to know that how much score on the LB will be considered a high score? keeping in view the current condition of the data.\n\nI mean what score will be satisfactory to the organizers of the competition? Currently maximum scores are going around 0.49+. The models which are generating these scores are these sufficient to be used in a practical scenario or do you think a much higher score is needed for using the solution in practical scenarios.  \n\nI know \"the higher the better\" but keeping in view the current condition of data, what score will qualify to be practical?\n\nThank you",
      "votes": 2
    },
    {
      "id": 3014239,
      "postDate": "2024-10-11T04:57:36.657Z",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/gkiar07\" target=\"_blank\">@gkiar07</a>,<br>\n I missed to attend the above mentioned webinar can we get recording of it !, it would be more helpful for persons like me.<br>\nThank you.</p>",
      "rawMarkdown": "Dear @gkiar07,\n I missed to attend the above mentioned webinar can we get recording of it !, it would be more helpful for persons like me.\nThank you.",
      "votes": 2,
      "replies": [
        {
          "id": 3022395,
          "postDate": "2024-10-19T14:12:26.987Z",
          "content": "<p><a href=\"https://vimeo.com/1020745085\" target=\"_blank\">https://vimeo.com/1020745085</a></p>",
          "rawMarkdown": "https://vimeo.com/1020745085",
          "votes": 3
        }
      ]
    },
    {
      "id": 3023871,
      "postDate": "2024-10-21T05:56:19.283Z",
      "content": "<p>In real scenarios, we have to work with noisy data. Thanks, Greg for the clarification on the dataset.</p>",
      "rawMarkdown": "In real scenarios, we have to work with noisy data. Thanks, Greg for the clarification on the dataset.",
      "votes": 1
    },
    {
      "id": 3022139,
      "postDate": "2024-10-19T10:19:31.393Z",
      "content": "<p>Thanks, Greg! I'm excited to work with this real-world data, even with all the noise and challenges. Looking forward to finding some great insights!</p>",
      "rawMarkdown": "Thanks, Greg! I'm excited to work with this real-world data, even with all the noise and challenges. Looking forward to finding some great insights!",
      "votes": 1,
      "replies": [
        {
          "id": 3023849,
          "postDate": "2024-10-21T05:13:13.563Z",
          "content": "<p>yeah! its really interesting to work with this kind of data.</p>",
          "rawMarkdown": "yeah! its really interesting to work with this kind of data.",
          "votes": 1
        }
      ]
    },
    {
      "id": 3073201,
      "postDate": "2024-12-16T06:51:46.200Z",
      "content": "<p>Emmmmmm…….maybe </p>",
      "rawMarkdown": "Emmmmmm.......maybe "
    },
    {
      "id": 3061823,
      "postDate": "2024-12-03T02:25:54.177Z",
      "content": "<p>Hi, thanks for explaination. Want to understand for the final private LB, would the test set be completely different or using part of the current data in public LB? and How big the data set would be?</p>",
      "rawMarkdown": "Hi, thanks for explaination. Want to understand for the final private LB, would the test set be completely different or using part of the current data in public LB? and How big the data set would be?",
      "replies": [
        {
          "id": 3073777,
          "postDate": "2024-12-16T22:43:05.690Z",
          "content": "<p>I think it's mentioned in the dataset description - \"Note that this is a Code Competition, in which the actual test set is hidden. In this public version, we give some sample data in the correct format to help you author your solutions. The full test set comprises about 3800 instances.\" Does this answer your question?</p>",
          "rawMarkdown": "I think it's mentioned in the dataset description - \"Note that this is a Code Competition, in which the actual test set is hidden. In this public version, we give some sample data in the correct format to help you author your solutions. The full test set comprises about 3800 instances.\" Does this answer your question?"
        }
      ]
    },
    {
      "id": 3054824,
      "postDate": "2024-11-25T08:15:46.473Z",
      "content": "<p>It's helpful!</p>",
      "rawMarkdown": "It's helpful!"
    },
    {
      "id": 3044171,
      "postDate": "2024-11-13T06:11:00.300Z",
      "content": "<p>----jhjjhjjjjjjjjj</p>",
      "rawMarkdown": "----jhjjhjjjjjjjjj"
    },
    {
      "id": 3028455,
      "postDate": "2024-10-26T05:47:17.937Z",
      "content": "<p>This discussion was useful.</p>",
      "rawMarkdown": "This discussion was useful."
    },
    {
      "id": 3022380,
      "postDate": "2024-10-19T13:54:56.500Z",
      "content": "<p>The recording of the Dell webinar <a href=\"https://vimeo.com/1020745085\" target=\"_blank\">https://vimeo.com/1020745085</a></p>",
      "rawMarkdown": "The recording of the Dell webinar https://vimeo.com/1020745085",
      "replies": [
        {
          "id": 3034946,
          "postDate": "2024-11-02T19:23:43.690Z",
          "content": "<p>Thanks for sharing webinar!</p>",
          "rawMarkdown": "Thanks for sharing webinar!"
        }
      ]
    },
    {
      "id": 3020077,
      "postDate": "2024-10-17T06:24:49.957Z",
      "content": "<p>I guess it makes the competition more engaging, since It is my second competition on Kaggle, I really had to do lots of hard work here, I am willing to find more and more competitive environments.</p>",
      "rawMarkdown": "I guess it makes the competition more engaging, since It is my second competition on Kaggle, I really had to do lots of hard work here, I am willing to find more and more competitive environments."
    },
    {
      "id": 3012441,
      "postDate": "2024-10-09T03:29:07.940Z",
      "content": "<p>some features of the training set are not found in the test set。 And I find features about the \"Parent-Child Internet Addiction Test\" had a big impact on the results。In the final grading session, does the test set have these characteristics or does it not?</p>",
      "rawMarkdown": "some features of the training set are not found in the test set。 And I find features about the \"Parent-Child Internet Addiction Test\" had a big impact on the results。In the final grading session, does the test set have these characteristics or does it not?",
      "replies": [
        {
          "id": 3018917,
          "postDate": "2024-10-16T06:11:01.310Z",
          "content": "<p>target is derived from Parent-Child Internet Addiction Test, they should not be considered as features.</p>",
          "rawMarkdown": "target is derived from Parent-Child Internet Addiction Test, they should not be considered as features.",
          "votes": 1,
          "replies": [
            {
              "id": 3059213,
              "postDate": "2024-11-30T13:02:02.280Z",
              "content": "<p>which features should not be considered?  I think modeling should be based on only common features.</p>",
              "rawMarkdown": "which features should not be considered?  I think modeling should be based on only common features."
            },
            {
              "id": 3060675,
              "postDate": "2024-12-02T00:56:12.197Z",
              "content": "<p>I directly use pd.columns to get test dataset‘s columns. And use train[test_clolums].copy(),so train and test has the same columns</p>",
              "rawMarkdown": "I directly use pd.columns to get test dataset‘s columns. And use train[test_clolums].copy(),so train and test has the same columns"
            }
          ]
        }
      ]
    },
    {
      "id": 3066045,
      "postDate": "2024-12-07T15:35:07.633Z",
      "content": "<p>在实际场景中，我们必须处理嘈杂的数据</p>",
      "rawMarkdown": "在实际场景中，我们必须处理嘈杂的数据",
      "votes": -3,
      "isDeleted": true
    },
    {
      "id": 3079520,
      "postDate": "2024-12-23T18:38:08.617Z",
      "content": "<p>Thanks for the information.</p>",
      "rawMarkdown": "Thanks for the information."
    },
    {
      "id": 3070277,
      "postDate": "2024-12-12T13:48:29.997Z",
      "content": "<p>good!!!!!!</p>",
      "rawMarkdown": "good!!!!!!"
    },
    {
      "id": 3042690,
      "postDate": "2024-11-11T17:16:30.237Z",
      "content": "<p>Thanks Greg for the clarification !</p>",
      "rawMarkdown": "Thanks Greg for the clarification !"
    },
    {
      "id": 3037738,
      "postDate": "2024-11-06T04:42:10.780Z",
      "content": "<p>Thank you for the clarification! </p>",
      "rawMarkdown": "Thank you for the clarification! "
    },
    {
      "id": 3030845,
      "postDate": "2024-10-29T01:18:05.843Z",
      "content": "<p>Let me benefit a lot! Thank you！</p>",
      "rawMarkdown": "Let me benefit a lot! Thank you！"
    },
    {
      "id": 3009627,
      "postDate": "2024-10-08T06:13:04.830Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3009735,
      "author_name": "Antonina Dolgorukova",
      "author_url": "",
      "post_date": "2024-10-08T09:13:39.227000",
      "content": "<p>Thanks for your post! I was questioning the practical implications of the model (\"Questioning the task\" thread), but now I think I can just focus on this new goal: finding out how we can make use of this data. </p>\n<p>Many of us have already realised that we need to do a lot of data cleaning. Some variables require domain knowledge to clean (e.g. unrealistic blood pressure, heart rate, BMI), but that seems to be the easiest part - finding normal ranges and removing outliers or replacing them with imputation methods. For the rest, we need more information to find normal ranges. It would also help a lot if we could better understand the nature of the data. For now, I think it would be very helpful if you could answer these questions:</p>\n<p>1) FitnessGram Vitals and Treadmill: Could you please share the methodology - what was the test? what devide was used? could you direct us to where we can read about it (maybe there is a user manual?), or just give reference values for age groups?<br>\n2) FitnessGram Child - Zones: Could you explain how the values were assigned to the zones?<br>\n3) Bioelectrical Impedance Analysis: What machine was used (would be great to see the user manual). Can you point us to some resources for finding normal ranges specific to the machine used to collect data for this competition?</p>\n<p>UPD: for BIA - What do the values represent - raw data or did you use some BIA equation models to estimate muscle mass and other parameters? If so, which equation was used?</p>\n<p>4) Physical activity questionnaires: What questions were asked?<br>\n5) Sleep disturbance scale: How was it measured and where can we get some information on its interpretation?</p>",
      "votes": 19,
      "replies": [
        {
          "id": 3023850,
          "author_name": "Taimour Nazar",
          "author_url": "",
          "post_date": "2024-10-21T05:13:40.833000",
          "content": "<p>Good points. Thank you</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 3029198,
          "author_name": "Anna Gams",
          "author_url": "",
          "post_date": "2024-10-27T02:56:43.250000",
          "content": "<p><a href=\"https://www.kaggle.com/antoninadolgorukova\" target=\"_blank\">@antoninadolgorukova</a> were any of your questions answered? I am also very curious how those values were obtained. Thanks!</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3029733,
              "author_name": "Antonina Dolgorukova",
              "author_url": "",
              "post_date": "2024-10-27T16:56:40.393000",
              "content": "<p>No, nobody answered these questions as well as the ones below in my separate comment…</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3009630,
      "author_name": "AmbrosM",
      "author_url": "",
      "post_date": "2024-10-08T06:17:07.597000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/gkiar07\" target=\"_blank\">@gkiar07</a>, thank you for this clarification! There is one question many Kagglers have in this competition; perhaps you're willing to answer it: How did you split the data into Kaggle's training and test datasets? Is it a completely random split, or do training and test represent different batches of your data acquisition?</p>",
      "votes": 12,
      "replies": [
        {
          "id": 3017447,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-10-14T23:10:06.747000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3032429,
          "author_name": "DavidS",
          "author_url": "",
          "post_date": "2024-10-30T21:48:22.867000",
          "content": "<p>Hello. I'd like to ask - why is it reported and why there's a deleted comment? Was there any leakage of information that someone could have seen already?  <a href=\"https://www.kaggle.com/gkiar07\" target=\"_blank\">@gkiar07</a></p>",
          "votes": 1,
          "replies": [
            {
              "id": 3073775,
              "author_name": "SahanaB",
              "author_url": "",
              "post_date": "2024-12-16T22:37:45.713000",
              "content": "<p>Is there a reason that the discussion of whether the train/test split was random keeps coming up. Just trying to understand how people would respond if the split were not random..</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3041343,
      "author_name": "ctrfd",
      "author_url": "",
      "post_date": "2024-11-10T07:43:12.160000",
      "content": "<p>It wouldn't make sense if the data wasn't this badly noisy and incomplete. That's sort of expected, and only make the challenge more exciting tbh.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 3041346,
          "author_name": "Taimour Nazar",
          "author_url": "",
          "post_date": "2024-11-10T07:48:22.613000",
          "content": "<p>I think same. If the data was not this badly noisy and incomplete then there was no reason for this competition to have prize money. It has prize money to attract data scientists to solve this problem. I will only say that there should have a been a little more explanation of when the data was collected, because carefully looking at data it shows that different parts of data for a single person were collected at different ages.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3041478,
          "author_name": "amir mohd",
          "author_url": "",
          "post_date": "2024-11-10T12:25:46.173000",
          "content": "<p>Completely agree! The complexity adds to the challenge, but clearer details about the data collection timeline would have provided a better foundation for tackling the inconsistencies.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3109184,
      "author_name": "Nurhat Baydağ",
      "author_url": "",
      "post_date": "2025-01-28T15:24:44.720000",
      "content": "<p>teşekkürler</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3009986,
      "author_name": "Daniel Dewey",
      "author_url": "",
      "post_date": "2024-10-08T15:08:53.267000",
      "content": "<p>Thanks for acknowleging the challenges of the data — though we’ve come to expect that from Kaggle :)</p>\n<p>Re: “We want to collectively learn about the hidden relationships in our data…” — this is a good goal and/but it is unspervised learning and so it doesn’t have a clear, objective metric. (A peers’-subjective metric might have levels: “meh”, “interesting”, “cool”, and “wow!”.)</p>\n<p>“.. and how they can jointly be used to model [predict] problematic internet use in children and adolescents.” This part is supervised learning and requires a “ground truth” target. As the “Questioning the Task” discussion bought out, the “ground truth” target here is actually a parent’s (hastily? not-at-all?) filled-out questionaire rather than some expert’s assessment based on discussions, observations, etc. So we’re really predicting one (important) feature from the others rather than predicting the “experts’ sii” from all the features.</p>\n<p>As many have said, none of that really matters to those of us with high levels of Problematic Kaggle Use —  Thanks for the data!   :)</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3012528,
          "author_name": "Antonina Dolgorukova",
          "author_url": "",
          "post_date": "2024-10-09T06:21:35.110000",
          "content": "<p>I like the way you put it: \"we really do predict one (important) feature from the others\". I still hope to hear justification/explanation from the hosts though .</p>\n<p>Indeed, with the target provided, we can only predict how parents would rate the severity of PIU in their children, i.e., kind of a level of parental worry or concern.</p>\n<p>Additionally, the applicability of the PCIAT questionnaire across the age range is questionable. I think all the questions in the PCIAT are much more suitable for adolescents. For example:</p>\n<ul>\n<li>A 5-7-year-old may not have household chores, as this depends on cultural norms.</li>\n<li>The question about academic impact may not apply to younger children who are not yet in school or graduated adults.</li>\n<li>Email use and receiving phone calls from \"online friends\" seem out of context for young 5-6 y.o. children too.</li>\n<li>Questions about reaction to the time allowed to spend on the internet (there are at least 3 of them) are not applicable to adults… usually :).</li>\n</ul>\n<p>Given this, if we see, for example, that adolescents seem to have the highest SII across all levels of Internet use (added to my \"Features EDA\" notebook, the assessment is rough as there is huge variability, but it illustrates my point well)… So are they more susceptible to PIU, or is this questionnaire just more sensitive to PIU in this age group? </p>\n<p>BTW, did the adults in this data fill in the questionnaire themselves, or were their parents asked 🙃?</p>",
          "votes": 4,
          "replies": [
            {
              "id": 3012938,
              "author_name": "Daniel Dewey",
              "author_url": "",
              "post_date": "2024-10-09T14:04:58.693000",
              "content": "<p>… and the questions feel dated, from the late 1990s or early 2000s? I was a parent then :)</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3009749,
      "author_name": "Antonina Dolgorukova",
      "author_url": "",
      "post_date": "2024-10-08T09:31:26.520000",
      "content": "<p>it also would be great if you could comment on these:</p>\n<p>1) Participants with data in the children's Physical Activity Questionnaire columns (e.g., PAQ_C_Total) are aged 7 to 17, which overlaps with those in the adolescents' columns (PAQ_A_Total), who are aged 13 to 18. One participant, aged 13, appeared to complete both questionnaires. Could you clarify?</p>\n<p>2) Endurance measures (FitnessGram Vitals and Treadmill tests) were only conducted in children aged 5-12 years. Is there a specific reason for this, or did it just happen that way?</p>\n<p>3) Some questions in the PCIAT test appear to have been skipped by respondents (resulting in missing values in the PCIAT columns), but the Total score is still calculated as the sum of the non-missing values. This could potentially lead to invalid SII scores. </p>\n<p>4) A great question from <a href=\"https://www.kaggle.com/rafidmahbub\" target=\"_blank\">@rafidmahbub</a> in the thread I mentioned: There are a lot of high blood pressure (BP) and heart rate (HR) values. Should we interpret these as pathological, or were they taken during exercise? Basically, what do the BP and HR data represent - exercise response tests or usual resting values?</p>\n<p>This are the main issues posted in another thread (\"Some findings from the features EDA (+ data issues, outliers)\")</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3061829,
          "author_name": "Lawrence Chernin",
          "author_url": "",
          "post_date": "2024-12-03T02:32:53.150000",
          "content": "<p>On your 3rd point, it may make sense to impute the missing values and then calculate an adjusted SII.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3047668,
      "author_name": "Vladislav Shakhrai",
      "author_url": "",
      "post_date": "2024-11-17T03:08:27.407000",
      "content": "<p>Thanks for the information. Could you please clarify if train/test split has been done completely randomly or not?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3018467,
      "author_name": "Lonnie Wibberding",
      "author_url": "",
      "post_date": "2024-10-15T19:03:02.547000",
      "content": "<p>Thanks for engaging in this project. I appreciate the context of what you are trying to do. It also gives us data scientists who are trying to improve our craft excellent real world experience.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3013798,
      "author_name": "Azadeh Razmi",
      "author_url": "",
      "post_date": "2024-10-10T14:41:57.703000",
      "content": "<p>Thank you for this post :)<br>\nyour data need a lot of cleansing , and this affected  the reliability of it , in some cases explain more details about method of data gathering , standard ranges and ages is helpful <br>\nadditionally, if you use any grouping boarder(some thing like obesity, fat, normal, etc..) or use a standard, please explain , I can't find nothing </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3023845,
      "author_name": "Taimour Nazar",
      "author_url": "",
      "post_date": "2024-10-21T05:08:30.337000",
      "content": "<p>Thank you for the detailed explanation.</p>\n<p>Based on the provided details it is clear that reaching a high accuracy in predictions is the real challenge so, I would like to know that how much score on the LB will be considered a high score? keeping in view the current condition of the data.</p>\n<p>I mean what score will be satisfactory to the organizers of the competition? Currently maximum scores are going around 0.49+. The models which are generating these scores are these sufficient to be used in a practical scenario or do you think a much higher score is needed for using the solution in practical scenarios.  </p>\n<p>I know \"the higher the better\" but keeping in view the current condition of data, what score will qualify to be practical?</p>\n<p>Thank you</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3014239,
      "author_name": "Prabhakar_Nimmagadda",
      "author_url": "",
      "post_date": "2024-10-11T04:57:36.657000",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/gkiar07\" target=\"_blank\">@gkiar07</a>,<br>\n I missed to attend the above mentioned webinar can we get recording of it !, it would be more helpful for persons like me.<br>\nThank you.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3022395,
          "author_name": "DAcacu",
          "author_url": "",
          "post_date": "2024-10-19T14:12:26.987000",
          "content": "<p><a href=\"https://vimeo.com/1020745085\" target=\"_blank\">https://vimeo.com/1020745085</a></p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 3023871,
      "author_name": "Anand Kumar",
      "author_url": "",
      "post_date": "2024-10-21T05:56:19.283000",
      "content": "<p>In real scenarios, we have to work with noisy data. Thanks, Greg for the clarification on the dataset.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3022139,
      "author_name": "abhijit shinde",
      "author_url": "",
      "post_date": "2024-10-19T10:19:31.393000",
      "content": "<p>Thanks, Greg! I'm excited to work with this real-world data, even with all the noise and challenges. Looking forward to finding some great insights!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3023849,
          "author_name": "Taimour Nazar",
          "author_url": "",
          "post_date": "2024-10-21T05:13:13.563000",
          "content": "<p>yeah! its really interesting to work with this kind of data.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3073201,
      "author_name": "Zachary Zhao",
      "author_url": "",
      "post_date": "2024-12-16T06:51:46.200000",
      "content": "<p>Emmmmmm…….maybe </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3061823,
      "author_name": "MJeremy",
      "author_url": "",
      "post_date": "2024-12-03T02:25:54.177000",
      "content": "<p>Hi, thanks for explaination. Want to understand for the final private LB, would the test set be completely different or using part of the current data in public LB? and How big the data set would be?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3073777,
          "author_name": "SahanaB",
          "author_url": "",
          "post_date": "2024-12-16T22:43:05.690000",
          "content": "<p>I think it's mentioned in the dataset description - \"Note that this is a Code Competition, in which the actual test set is hidden. In this public version, we give some sample data in the correct format to help you author your solutions. The full test set comprises about 3800 instances.\" Does this answer your question?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3054824,
      "author_name": "WEXHICY",
      "author_url": "",
      "post_date": "2024-11-25T08:15:46.473000",
      "content": "<p>It's helpful!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3044171,
      "author_name": "Yash",
      "author_url": "",
      "post_date": "2024-11-13T06:11:00.300000",
      "content": "<p>----jhjjhjjjjjjjjj</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3028455,
      "author_name": "Kazuya Kimura",
      "author_url": "",
      "post_date": "2024-10-26T05:47:17.937000",
      "content": "<p>This discussion was useful.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3022380,
      "author_name": "DAcacu",
      "author_url": "",
      "post_date": "2024-10-19T13:54:56.500000",
      "content": "<p>The recording of the Dell webinar <a href=\"https://vimeo.com/1020745085\" target=\"_blank\">https://vimeo.com/1020745085</a></p>",
      "votes": 0,
      "replies": [
        {
          "id": 3034946,
          "author_name": "PaPa22",
          "author_url": "",
          "post_date": "2024-11-02T19:23:43.690000",
          "content": "<p>Thanks for sharing webinar!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3020077,
      "author_name": "Shivaabhishek108",
      "author_url": "",
      "post_date": "2024-10-17T06:24:49.957000",
      "content": "<p>I guess it makes the competition more engaging, since It is my second competition on Kaggle, I really had to do lots of hard work here, I am willing to find more and more competitive environments.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3012441,
      "author_name": "Nigel Xiang",
      "author_url": "",
      "post_date": "2024-10-09T03:29:07.940000",
      "content": "<p>some features of the training set are not found in the test set。 And I find features about the \"Parent-Child Internet Addiction Test\" had a big impact on the results。In the final grading session, does the test set have these characteristics or does it not?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3018917,
          "author_name": "cggStandStill",
          "author_url": "",
          "post_date": "2024-10-16T06:11:01.310000",
          "content": "<p>target is derived from Parent-Child Internet Addiction Test, they should not be considered as features.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3059213,
              "author_name": "Maryam Shahzadi",
              "author_url": "",
              "post_date": "2024-11-30T13:02:02.280000",
              "content": "<p>which features should not be considered?  I think modeling should be based on only common features.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3060675,
              "author_name": "Nigel Xiang",
              "author_url": "",
              "post_date": "2024-12-02T00:56:12.197000",
              "content": "<p>I directly use pd.columns to get test dataset‘s columns. And use train[test_clolums].copy(),so train and test has the same columns</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3066045,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-12-07T15:35:07.633000",
      "content": "<p>在实际场景中，我们必须处理嘈杂的数据</p>",
      "votes": -3,
      "replies": []
    },
    {
      "id": 3079520,
      "author_name": "Kabir Olawale Mohammed",
      "author_url": "",
      "post_date": "2024-12-23T18:38:08.617000",
      "content": "<p>Thanks for the information.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3070277,
      "author_name": "Orelia Lu",
      "author_url": "",
      "post_date": "2024-12-12T13:48:29.997000",
      "content": "<p>good!!!!!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3042690,
      "author_name": "Sichoix Bryan",
      "author_url": "",
      "post_date": "2024-11-11T17:16:30.237000",
      "content": "<p>Thanks Greg for the clarification !</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3037738,
      "author_name": "meh",
      "author_url": "",
      "post_date": "2024-11-06T04:42:10.780000",
      "content": "<p>Thank you for the clarification! </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3030845,
      "author_name": "Mengsihan12138",
      "author_url": "",
      "post_date": "2024-10-29T01:18:05.843000",
      "content": "<p>Let me benefit a lot! Thank you！</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3009627,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-10-08T06:13:04.830000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3009335": "Hi, Kagglers!\n\nIt's wonderful to see so much engagement already in our competition. Firstly, on behalf of all of our organizing team, we just want to say thank you for joining us in this challenge!\n\nOver the last couple of weeks we've noticed a handful of discussion threads pop-up that touch on some of the same core elements of our dataset, and so I wanted to speak to some of those details directly, and provide more context about the nature of our data.\n\nIn particular: **yes, the data is noisy, biased, incomplete, and subject to a variety of other off-target sources of variation.** This heterogeneity is at the heart of our challenge, and the heart of large-scale psychiatric research more broadly. Many elements of the phenomena we try to measure are subjective, and are influenced by both systemic bias (e.g., accessibility) and random variation (e.g., time of day, last meal). Indeed, some biases will be related directly to the signal we care about (e.g., parental report of problematic internet use will be biased by the parents' own internet use), and in other cases, they may be entirely unrelated. Indeed, many of our sample are not severely impacted, and that is consistent with our observed prevalence of this issue in our community. Overfitting will be easy, with all of these imbalances and sources of noise.\n\nThe list of these sorts of relationships and the challenges they each introduce goes on and on; our question is ultimately *how can we make use of these data, regardless?* In fact, these challenges are precisely ***how our dataset represents the real-world*** context that we are trying to model and address.\n\nWe deliberately did not want to ship a sanitized version of this data for you to work from, because then we would not be able to apply what you learn in practice. We want to collectively learn about the hidden relationships in our data and how they can jointly be used to model problematic internet use in children and adolescents. In order to do so, we needed to show you what we see: a vast portrait with a lot of signal, but also some noise.\n\nIf you want to hear a bit more about our dataset and challenge, join for the [Webinar that Dell Technologies is hosting at the end of the week](https://events.dell.com/event/c87be06c-a180-4c88-b9bf-7fe6f365d43a/register).\n\n***Edit:** [Here is a link of the recording]( https://vimeo.com/1020745085).* \n\nThanks again for joining us, and we're excited to continue learning about our data with you! :rocket:\n\n-Greg",
    "3009735": "Thanks for your post! I was questioning the practical implications of the model (\"Questioning the task\" thread), but now I think I can just focus on this new goal: finding out how we can make use of this data. \n\nMany of us have already realised that we need to do a lot of data cleaning. Some variables require domain knowledge to clean (e.g. unrealistic blood pressure, heart rate, BMI), but that seems to be the easiest part - finding normal ranges and removing outliers or replacing them with imputation methods. For the rest, we need more information to find normal ranges. It would also help a lot if we could better understand the nature of the data. For now, I think it would be very helpful if you could answer these questions:\n\n1) FitnessGram Vitals and Treadmill: Could you please share the methodology - what was the test? what devide was used? could you direct us to where we can read about it (maybe there is a user manual?), or just give reference values for age groups?\n2) FitnessGram Child - Zones: Could you explain how the values were assigned to the zones?\n3) Bioelectrical Impedance Analysis: What machine was used (would be great to see the user manual). Can you point us to some resources for finding normal ranges specific to the machine used to collect data for this competition?\n\nUPD: for BIA - What do the values represent - raw data or did you use some BIA equation models to estimate muscle mass and other parameters? If so, which equation was used?\n\n4) Physical activity questionnaires: What questions were asked?\n5) Sleep disturbance scale: How was it measured and where can we get some information on its interpretation?",
    "3009630": "Hi @gkiar07, thank you for this clarification! There is one question many Kagglers have in this competition; perhaps you're willing to answer it: How did you split the data into Kaggle's training and test datasets? Is it a completely random split, or do training and test represent different batches of your data acquisition?",
    "3041343": "It wouldn't make sense if the data wasn't this badly noisy and incomplete. That's sort of expected, and only make the challenge more exciting tbh.",
    "3109184": "teşekkürler",
    "3009986": "Thanks for acknowleging the challenges of the data — though we’ve come to expect that from Kaggle :)\n\nRe: “We want to collectively learn about the hidden relationships in our data…” — this is a good goal and/but it is unspervised learning and so it doesn’t have a clear, objective metric. (A peers’-subjective metric might have levels: “meh”, “interesting”, “cool”, and “wow!”.)\n\n“.. and how they can jointly be used to model [predict] problematic internet use in children and adolescents.” This part is supervised learning and requires a “ground truth” target. As the “Questioning the Task” discussion bought out, the “ground truth” target here is actually a parent’s (hastily? not-at-all?) filled-out questionaire rather than some expert’s assessment based on discussions, observations, etc. So we’re really predicting one (important) feature from the others rather than predicting the “experts’ sii” from all the features.\n\nAs many have said, none of that really matters to those of us with high levels of Problematic Kaggle Use —  Thanks for the data!   :)",
    "3009749": "it also would be great if you could comment on these:\n\n1) Participants with data in the children's Physical Activity Questionnaire columns (e.g., PAQ_C_Total) are aged 7 to 17, which overlaps with those in the adolescents' columns (PAQ_A_Total), who are aged 13 to 18. One participant, aged 13, appeared to complete both questionnaires. Could you clarify?\n\n2) Endurance measures (FitnessGram Vitals and Treadmill tests) were only conducted in children aged 5-12 years. Is there a specific reason for this, or did it just happen that way?\n\n3) Some questions in the PCIAT test appear to have been skipped by respondents (resulting in missing values in the PCIAT columns), but the Total score is still calculated as the sum of the non-missing values. This could potentially lead to invalid SII scores. \n\n4) A great question from @rafidmahbub in the thread I mentioned: There are a lot of high blood pressure (BP) and heart rate (HR) values. Should we interpret these as pathological, or were they taken during exercise? Basically, what do the BP and HR data represent - exercise response tests or usual resting values?\n\nThis are the main issues posted in another thread (\"Some findings from the features EDA (+ data issues, outliers)\")",
    "3047668": "Thanks for the information. Could you please clarify if train/test split has been done completely randomly or not?",
    "3018467": "Thanks for engaging in this project. I appreciate the context of what you are trying to do. It also gives us data scientists who are trying to improve our craft excellent real world experience.",
    "3013798": "Thank you for this post :)\nyour data need a lot of cleansing , and this affected  the reliability of it , in some cases explain more details about method of data gathering , standard ranges and ages is helpful \nadditionally, if you use any grouping boarder(some thing like obesity, fat, normal, etc..) or use a standard, please explain , I can't find nothing ",
    "3023845": "Thank you for the detailed explanation.\n\nBased on the provided details it is clear that reaching a high accuracy in predictions is the real challenge so, I would like to know that how much score on the LB will be considered a high score? keeping in view the current condition of the data.\n\nI mean what score will be satisfactory to the organizers of the competition? Currently maximum scores are going around 0.49+. The models which are generating these scores are these sufficient to be used in a practical scenario or do you think a much higher score is needed for using the solution in practical scenarios.  \n\nI know \"the higher the better\" but keeping in view the current condition of data, what score will qualify to be practical?\n\nThank you",
    "3014239": "Dear @gkiar07,\n I missed to attend the above mentioned webinar can we get recording of it !, it would be more helpful for persons like me.\nThank you.",
    "3023871": "In real scenarios, we have to work with noisy data. Thanks, Greg for the clarification on the dataset.",
    "3022139": "Thanks, Greg! I'm excited to work with this real-world data, even with all the noise and challenges. Looking forward to finding some great insights!",
    "3073201": "Emmmmmm.......maybe ",
    "3061823": "Hi, thanks for explaination. Want to understand for the final private LB, would the test set be completely different or using part of the current data in public LB? and How big the data set would be?",
    "3054824": "It's helpful!",
    "3044171": "----jhjjhjjjjjjjjj",
    "3028455": "This discussion was useful.",
    "3022380": "The recording of the Dell webinar https://vimeo.com/1020745085",
    "3020077": "I guess it makes the competition more engaging, since It is my second competition on Kaggle, I really had to do lots of hard work here, I am willing to find more and more competitive environments.",
    "3012441": "some features of the training set are not found in the test set。 And I find features about the \"Parent-Child Internet Addiction Test\" had a big impact on the results。In the final grading session, does the test set have these characteristics or does it not?",
    "3066045": "在实际场景中，我们必须处理嘈杂的数据",
    "3079520": "Thanks for the information.",
    "3070277": "good!!!!!!",
    "3042690": "Thanks Greg for the clarification !",
    "3037738": "Thank you for the clarification! ",
    "3030845": "Let me benefit a lot! Thank you！",
    "3009627": ""
  }
}