{
  "id": 552601,
  "title": "End of Competition - Thank you!",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/552601",
  "author_name": "Greg Kiar",
  "post_date": "2024-12-20T13:36:44.802000",
  "votes": 20,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi Kagglers!</p>\n<p>We made it! On behalf of our team at the Child Mind Institute and our sponsors at the California Department of Health Care Services, Dell Technologies, and NVIDIA, thank you all SO much for participating in our competition over the last three months! We couldn’t be happier to have the participation of over 4,500 of you, submitting over 85,000 submissions to tackle our challenge! Truly, we are grateful for the time and thought you invested in solving our problem. The contributions of each and every one of you, through your submissions, interactions on the discussion board, and exploration of our dataset, are hugely valuable and we look forward to digesting them all and learning from you.</p>\n<p>In the past few days it’s been fun and insightful to read the discussion forums and see you, the community, examine the data features, visualize interesting patterns, learn from the data, and help each other with model definition and other data science aspects of the competition.</p>\n<h3>Leaderboard Shake-Up: What Happened?</h3>\n<p>We also noticed that you started to recognize a trend that we’ve observed behind-the-scenes for a while: there’s plenty of opportunity for overfitting in our competition, and as a result, a large shake-up on the leaderboard. There have been several interesting threads where you’ve each shared what were <a href=\"https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/551758\" target=\"_blank\">your signs for detecting overfitting</a>, making <a href=\"https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/551913\" target=\"_blank\">predictions about the distribution of Problematic Internet Use in the test set</a>, and <a href=\"https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/551164\" target=\"_blank\">guessing what the peak performance in the private test dataset will be</a> (which were generally more conservative than we see in practice). Across each, it’s clear that you have started to recognize the potential for overfitting by optimizing for leaderboard performance over consistency in cross-validation.</p>\n<p>There are plenty of other examples of Kaggle competitions with shake-ups (interestingly, according to this <a href=\"https://www.kaggle.com/code/jtrotman/meta-kaggle-scatter-plot-competition-shake-up\" target=\"_blank\">Meta-Kaggle post</a>, the largest shake-up in Kaggle history is also <a href=\"https://www.kaggle.com/c/icr-identify-age-related-conditions\" target=\"_blank\">in the space of healthcare</a>). </p>\n<p>To give some perspective, when I cloned our leaderboard on 12/18 (apologies if it’s slightly out of date at the time of posting), I calculated our shakeup to be <strong>0.255</strong>, which is around 15–20th all-time on Kaggle according to the Meta-Kaggle post I mentioned above — putting us right between two NFL-based competitions (<a href=\"https://www.kaggle.com/competitions/data-science-bowl-2017\" target=\"_blank\">1</a>, <a href=\"https://www.kaggle.com/c/nfl-impact-detection\" target=\"_blank\">2</a>). Now, I guess the more interesting question, <em>what is the source of our shake-up</em>?</p>\n<h3>Key Hypotheses: Overfitting or Not?</h3>\n<p>First, a few details. We stratified our sample explicitly across binned age, sex, and the presence of actigraphy data. In practice, we also ended up with a balanced split of all of our instruments (at most, the availability of an instrument differed by 4% between the public and private test set). The proportion of participants in each class was also balanced across the splits, and you can see the breakdown below:</p>\n<table>\n<thead>\n<tr>\n<th>sii</th>\n<th>Overall</th>\n<th>Train</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>0.577</td>\n<td>0.581</td>\n<td>0.587</td>\n<td>0.561</td>\n</tr>\n<tr>\n<td>1</td>\n<td>0.272</td>\n<td>0.275</td>\n<td>0.258</td>\n<td>0.273</td>\n</tr>\n<tr>\n<td>2</td>\n<td>0.14</td>\n<td>0.133</td>\n<td>0.144</td>\n<td>0.154</td>\n</tr>\n<tr>\n<td>3</td>\n<td>0.012</td>\n<td>0.012</td>\n<td>0.011</td>\n<td>0.013</td>\n</tr>\n</tbody>\n</table>\n<p>Given that our data are relatively well balanced, <em>what is going on here</em>? </p>\n<p>There are more than a few hypotheses… Let’s first assume that there wasn’t actually any overfitting of submissions in this competition, but that models had different strengths. If this were the case, it’s possible that models optimized with higher sensitivity to sii severity (i.e., were more likely to predict a higher score) were those that raised up the private leaderboard. This is possible given the small change in the distributions of sii classes between the private and public test sets. In this scenario, models that performed the best on the public set could have been doing a better job at predicting participants with an sii=0, of which there was a slightly smaller portion in the private test set.</p>\n<p>In practice, we’re not sure that this is the case. We saw some notebooks tinkering with minor model hyperparameters, thresholds, or random seeds to climb the public leaderboard. Another sign of overfitting on our public leaderboard is the submission count. If we look at the top 10 teams in the public leaderboard versus the private leaderboard, the public-leaders on average submitted 212 models, whereas the private-leaders on average submitted 64. If we look at the median, these values are 199 vs 25. This suggests that there was certainly some work to cook scores on the public leaderboard, that didn’t result in success on our held out private data.</p>\n<p>I think it’s relatively safe to rule out the hypothesis that there isn’t overfitting here, which now begs the question, <em>what about our data lends itself to overfitting or poor generalization</em>? </p>\n<p>There have been many (extremely informative!) discussion posts over the course of this competition that highlight possibilities: the possibility of batch effects across our samples; there could be data entry errors that limit reliability in the dataset; or the assessment instrument/prediction target is noisy and biased. All of these are potentially valid hypotheses, and deserve further exploration to understand. While we can evaluate the presence of batch effects in our data with tools like ComBAT, and automatically detect obvious data entry errors, evaluating the bias of our instruments is another can of worms altogether, and often impossible to do after data collection has taken place. If other assessments with similar purposes were deliberately included in data collection, this may be possible, but even in that case, the bias of the other assessments would likely need to be explored independently, perpetuating this challenge. In addition to the data challenges we’ve already mentioned, the nature of working with such heterogeneous datasets with relatively few samples is that there may simply not be strong enough statistical associations to give stable rankings — or critically, predictions when applied in real-world contexts — even in the absence of overfitting.</p>\n<h3>Real-World Data, Real-World Challenges</h3>\n<p>Batch effects, data entry errors, measurement biases, and dataset heterogeneity are all core challenges that many psychological sciences face on a regular basis. This competition reflects those “real-world” challenges, rather than hiding them or proposing sanitized data that researchers, data scientists, and clinicians rarely work with. Moreover, mental health and the definitions of “within expected ranges” or “acceptable” vary across cultures, ages, genders, specific samples (for example, clinical vs non-clinical sample), and more: variation is a challenge inherent to mental health data, and highly generalizable performance is often elusive. This competition showed itself to be no exception!</p>\n<p>Here, we aimed to offer participants the possibility of working with real data from the mental health space, to collectively learn from the data, identify gaps in data collection practices, bring innovation to the field, and hopefully have fun by participating! We feel that we’ve learned an immense amount from all of you, and are grateful for your company on this adventure. We look forward to debriefing with a handful of you at the winner’s calls soon. 🙂</p>\n<p>Sincerely,<br>\n-Greg</p>\n<pre><code>Gregory Kiar, PhD\nResearch Center for Data Analytics, Innovation, Rigor\nChild Mind </code></pre>",
  "messages": [
    {
      "id": 3077021,
      "postDate": "2024-12-20T13:36:44.803Z",
      "content": "<p>Hi Kagglers!</p>\n<p>We made it! On behalf of our team at the Child Mind Institute and our sponsors at the California Department of Health Care Services, Dell Technologies, and NVIDIA, thank you all SO much for participating in our competition over the last three months! We couldn’t be happier to have the participation of over 4,500 of you, submitting over 85,000 submissions to tackle our challenge! Truly, we are grateful for the time and thought you invested in solving our problem. The contributions of each and every one of you, through your submissions, interactions on the discussion board, and exploration of our dataset, are hugely valuable and we look forward to digesting them all and learning from you.</p>\n<p>In the past few days it’s been fun and insightful to read the discussion forums and see you, the community, examine the data features, visualize interesting patterns, learn from the data, and help each other with model definition and other data science aspects of the competition.</p>\n<h3>Leaderboard Shake-Up: What Happened?</h3>\n<p>We also noticed that you started to recognize a trend that we’ve observed behind-the-scenes for a while: there’s plenty of opportunity for overfitting in our competition, and as a result, a large shake-up on the leaderboard. There have been several interesting threads where you’ve each shared what were <a href=\"https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/551758\" target=\"_blank\">your signs for detecting overfitting</a>, making <a href=\"https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/551913\" target=\"_blank\">predictions about the distribution of Problematic Internet Use in the test set</a>, and <a href=\"https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/551164\" target=\"_blank\">guessing what the peak performance in the private test dataset will be</a> (which were generally more conservative than we see in practice). Across each, it’s clear that you have started to recognize the potential for overfitting by optimizing for leaderboard performance over consistency in cross-validation.</p>\n<p>There are plenty of other examples of Kaggle competitions with shake-ups (interestingly, according to this <a href=\"https://www.kaggle.com/code/jtrotman/meta-kaggle-scatter-plot-competition-shake-up\" target=\"_blank\">Meta-Kaggle post</a>, the largest shake-up in Kaggle history is also <a href=\"https://www.kaggle.com/c/icr-identify-age-related-conditions\" target=\"_blank\">in the space of healthcare</a>). </p>\n<p>To give some perspective, when I cloned our leaderboard on 12/18 (apologies if it’s slightly out of date at the time of posting), I calculated our shakeup to be <strong>0.255</strong>, which is around 15–20th all-time on Kaggle according to the Meta-Kaggle post I mentioned above — putting us right between two NFL-based competitions (<a href=\"https://www.kaggle.com/competitions/data-science-bowl-2017\" target=\"_blank\">1</a>, <a href=\"https://www.kaggle.com/c/nfl-impact-detection\" target=\"_blank\">2</a>). Now, I guess the more interesting question, <em>what is the source of our shake-up</em>?</p>\n<h3>Key Hypotheses: Overfitting or Not?</h3>\n<p>First, a few details. We stratified our sample explicitly across binned age, sex, and the presence of actigraphy data. In practice, we also ended up with a balanced split of all of our instruments (at most, the availability of an instrument differed by 4% between the public and private test set). The proportion of participants in each class was also balanced across the splits, and you can see the breakdown below:</p>\n<table>\n<thead>\n<tr>\n<th>sii</th>\n<th>Overall</th>\n<th>Train</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>0.577</td>\n<td>0.581</td>\n<td>0.587</td>\n<td>0.561</td>\n</tr>\n<tr>\n<td>1</td>\n<td>0.272</td>\n<td>0.275</td>\n<td>0.258</td>\n<td>0.273</td>\n</tr>\n<tr>\n<td>2</td>\n<td>0.14</td>\n<td>0.133</td>\n<td>0.144</td>\n<td>0.154</td>\n</tr>\n<tr>\n<td>3</td>\n<td>0.012</td>\n<td>0.012</td>\n<td>0.011</td>\n<td>0.013</td>\n</tr>\n</tbody>\n</table>\n<p>Given that our data are relatively well balanced, <em>what is going on here</em>? </p>\n<p>There are more than a few hypotheses… Let’s first assume that there wasn’t actually any overfitting of submissions in this competition, but that models had different strengths. If this were the case, it’s possible that models optimized with higher sensitivity to sii severity (i.e., were more likely to predict a higher score) were those that raised up the private leaderboard. This is possible given the small change in the distributions of sii classes between the private and public test sets. In this scenario, models that performed the best on the public set could have been doing a better job at predicting participants with an sii=0, of which there was a slightly smaller portion in the private test set.</p>\n<p>In practice, we’re not sure that this is the case. We saw some notebooks tinkering with minor model hyperparameters, thresholds, or random seeds to climb the public leaderboard. Another sign of overfitting on our public leaderboard is the submission count. If we look at the top 10 teams in the public leaderboard versus the private leaderboard, the public-leaders on average submitted 212 models, whereas the private-leaders on average submitted 64. If we look at the median, these values are 199 vs 25. This suggests that there was certainly some work to cook scores on the public leaderboard, that didn’t result in success on our held out private data.</p>\n<p>I think it’s relatively safe to rule out the hypothesis that there isn’t overfitting here, which now begs the question, <em>what about our data lends itself to overfitting or poor generalization</em>? </p>\n<p>There have been many (extremely informative!) discussion posts over the course of this competition that highlight possibilities: the possibility of batch effects across our samples; there could be data entry errors that limit reliability in the dataset; or the assessment instrument/prediction target is noisy and biased. All of these are potentially valid hypotheses, and deserve further exploration to understand. While we can evaluate the presence of batch effects in our data with tools like ComBAT, and automatically detect obvious data entry errors, evaluating the bias of our instruments is another can of worms altogether, and often impossible to do after data collection has taken place. If other assessments with similar purposes were deliberately included in data collection, this may be possible, but even in that case, the bias of the other assessments would likely need to be explored independently, perpetuating this challenge. In addition to the data challenges we’ve already mentioned, the nature of working with such heterogeneous datasets with relatively few samples is that there may simply not be strong enough statistical associations to give stable rankings — or critically, predictions when applied in real-world contexts — even in the absence of overfitting.</p>\n<h3>Real-World Data, Real-World Challenges</h3>\n<p>Batch effects, data entry errors, measurement biases, and dataset heterogeneity are all core challenges that many psychological sciences face on a regular basis. This competition reflects those “real-world” challenges, rather than hiding them or proposing sanitized data that researchers, data scientists, and clinicians rarely work with. Moreover, mental health and the definitions of “within expected ranges” or “acceptable” vary across cultures, ages, genders, specific samples (for example, clinical vs non-clinical sample), and more: variation is a challenge inherent to mental health data, and highly generalizable performance is often elusive. This competition showed itself to be no exception!</p>\n<p>Here, we aimed to offer participants the possibility of working with real data from the mental health space, to collectively learn from the data, identify gaps in data collection practices, bring innovation to the field, and hopefully have fun by participating! We feel that we’ve learned an immense amount from all of you, and are grateful for your company on this adventure. We look forward to debriefing with a handful of you at the winner’s calls soon. 🙂</p>\n<p>Sincerely,<br>\n-Greg</p>\n<pre><code>Gregory Kiar, PhD\nResearch Center for Data Analytics, Innovation, Rigor\nChild Mind </code></pre>",
      "rawMarkdown": "Hi Kagglers!\n\nWe made it! On behalf of our team at the Child Mind Institute and our sponsors at the California Department of Health Care Services, Dell Technologies, and NVIDIA, thank you all SO much for participating in our competition over the last three months! We couldn’t be happier to have the participation of over 4,500 of you, submitting over 85,000 submissions to tackle our challenge! Truly, we are grateful for the time and thought you invested in solving our problem. The contributions of each and every one of you, through your submissions, interactions on the discussion board, and exploration of our dataset, are hugely valuable and we look forward to digesting them all and learning from you.\n\nIn the past few days it’s been fun and insightful to read the discussion forums and see you, the community, examine the data features, visualize interesting patterns, learn from the data, and help each other with model definition and other data science aspects of the competition.\n\n### Leaderboard Shake-Up: What Happened?\nWe also noticed that you started to recognize a trend that we’ve observed behind-the-scenes for a while: there’s plenty of opportunity for overfitting in our competition, and as a result, a large shake-up on the leaderboard. There have been several interesting threads where you’ve each shared what were [your signs for detecting overfitting](https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/551758), making [predictions about the distribution of Problematic Internet Use in the test set](https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/551913), and [guessing what the peak performance in the private test dataset will be](https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/551164) (which were generally more conservative than we see in practice). Across each, it’s clear that you have started to recognize the potential for overfitting by optimizing for leaderboard performance over consistency in cross-validation.\n\nThere are plenty of other examples of Kaggle competitions with shake-ups (interestingly, according to this [Meta-Kaggle post](https://www.kaggle.com/code/jtrotman/meta-kaggle-scatter-plot-competition-shake-up), the largest shake-up in Kaggle history is also [in the space of healthcare](https://www.kaggle.com/c/icr-identify-age-related-conditions)). \n\nTo give some perspective, when I cloned our leaderboard on 12/18 (apologies if it’s slightly out of date at the time of posting), I calculated our shakeup to be **0.255**, which is around 15–20th all-time on Kaggle according to the Meta-Kaggle post I mentioned above — putting us right between two NFL-based competitions ([1](https://www.kaggle.com/competitions/data-science-bowl-2017), [2](https://www.kaggle.com/c/nfl-impact-detection)). Now, I guess the more interesting question, *what is the source of our shake-up*?\n\n### Key Hypotheses: Overfitting or Not?\nFirst, a few details. We stratified our sample explicitly across binned age, sex, and the presence of actigraphy data. In practice, we also ended up with a balanced split of all of our instruments (at most, the availability of an instrument differed by 4% between the public and private test set). The proportion of participants in each class was also balanced across the splits, and you can see the breakdown below:\n\n|sii |Overall |Train |Public |Private |\n|:--|:--|:--|:--|:--|\n| 0 | 0.577 | 0.581 | 0.587 | 0.561 |\n| 1 | 0.272 | 0.275 | 0.258 | 0.273 |\n| 2 | 0.14 | 0.133 | 0.144 | 0.154 |\n| 3 | 0.012 | 0.012 | 0.011 | 0.013 |\n\nGiven that our data are relatively well balanced, *what is going on here*? \n\nThere are more than a few hypotheses… Let’s first assume that there wasn’t actually any overfitting of submissions in this competition, but that models had different strengths. If this were the case, it’s possible that models optimized with higher sensitivity to sii severity (i.e., were more likely to predict a higher score) were those that raised up the private leaderboard. This is possible given the small change in the distributions of sii classes between the private and public test sets. In this scenario, models that performed the best on the public set could have been doing a better job at predicting participants with an sii=0, of which there was a slightly smaller portion in the private test set.\n\nIn practice, we’re not sure that this is the case. We saw some notebooks tinkering with minor model hyperparameters, thresholds, or random seeds to climb the public leaderboard. Another sign of overfitting on our public leaderboard is the submission count. If we look at the top 10 teams in the public leaderboard versus the private leaderboard, the public-leaders on average submitted 212 models, whereas the private-leaders on average submitted 64. If we look at the median, these values are 199 vs 25. This suggests that there was certainly some work to cook scores on the public leaderboard, that didn’t result in success on our held out private data.\n\nI think it’s relatively safe to rule out the hypothesis that there isn’t overfitting here, which now begs the question, *what about our data lends itself to overfitting or poor generalization*? \n\nThere have been many (extremely informative!) discussion posts over the course of this competition that highlight possibilities: the possibility of batch effects across our samples; there could be data entry errors that limit reliability in the dataset; or the assessment instrument/prediction target is noisy and biased. All of these are potentially valid hypotheses, and deserve further exploration to understand. While we can evaluate the presence of batch effects in our data with tools like ComBAT, and automatically detect obvious data entry errors, evaluating the bias of our instruments is another can of worms altogether, and often impossible to do after data collection has taken place. If other assessments with similar purposes were deliberately included in data collection, this may be possible, but even in that case, the bias of the other assessments would likely need to be explored independently, perpetuating this challenge. In addition to the data challenges we’ve already mentioned, the nature of working with such heterogeneous datasets with relatively few samples is that there may simply not be strong enough statistical associations to give stable rankings — or critically, predictions when applied in real-world contexts — even in the absence of overfitting.\n\n### Real-World Data, Real-World Challenges\nBatch effects, data entry errors, measurement biases, and dataset heterogeneity are all core challenges that many psychological sciences face on a regular basis. This competition reflects those “real-world” challenges, rather than hiding them or proposing sanitized data that researchers, data scientists, and clinicians rarely work with. Moreover, mental health and the definitions of “within expected ranges” or “acceptable” vary across cultures, ages, genders, specific samples (for example, clinical vs non-clinical sample), and more: variation is a challenge inherent to mental health data, and highly generalizable performance is often elusive. This competition showed itself to be no exception!\n\nHere, we aimed to offer participants the possibility of working with real data from the mental health space, to collectively learn from the data, identify gaps in data collection practices, bring innovation to the field, and hopefully have fun by participating! We feel that we’ve learned an immense amount from all of you, and are grateful for your company on this adventure. We look forward to debriefing with a handful of you at the winner’s calls soon. 🙂\n\n\nSincerely,\n-Greg\n\n\n```\nGregory Kiar, PhD\nResearch Scientist\nDirector, Center for Data Analytics, Innovation, and Rigor\nChild Mind Institute\n```",
      "votes": 20
    },
    {
      "id": 3077524,
      "postDate": "2024-12-21T04:53:43.517Z",
      "content": "<p>Thank you for organizing this competition! Recently, a Kaggler pointed out the striking differences between missing values proportion in train vs test data <a href=\"https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/552488#3076452\" target=\"_blank\">here</a> - this issue was not raised in this post</p>\n<p>What do you think may have caused this ?</p>",
      "rawMarkdown": "Thank you for organizing this competition! Recently, a Kaggler pointed out the striking differences between missing values proportion in train vs test data [here](https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/552488#3076452) - this issue was not raised in this post\n\nWhat do you think may have caused this ?",
      "votes": 5
    },
    {
      "id": 3077157,
      "postDate": "2024-12-20T16:19:03.883Z",
      "content": "<p>Thank you for this challenging real world data, learnt a lot and Many Congratulations to the Winners!</p>",
      "rawMarkdown": "Thank you for this challenging real world data, learnt a lot and Many Congratulations to the Winners!",
      "votes": 1
    },
    {
      "id": 3150238,
      "postDate": "2025-03-15T08:00:50.277Z",
      "content": "<p>This is really difficult for a kaggle beginner</p>",
      "rawMarkdown": "This is really difficult for a kaggle beginner"
    },
    {
      "id": 3135608,
      "postDate": "2025-02-27T16:33:20.363Z",
      "content": "<p>perfect competition</p>",
      "rawMarkdown": "perfect competition"
    },
    {
      "id": 3078498,
      "postDate": "2024-12-22T11:47:40.687Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 3083596,
      "postDate": "2024-12-29T17:45:58.300Z",
      "content": "<p>thank you </p>",
      "rawMarkdown": "thank you ",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 3077524,
      "author_name": "Geremie Yeo",
      "author_url": "",
      "post_date": "2024-12-21T04:53:43.517000",
      "content": "<p>Thank you for organizing this competition! Recently, a Kaggler pointed out the striking differences between missing values proportion in train vs test data <a href=\"https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/552488#3076452\" target=\"_blank\">here</a> - this issue was not raised in this post</p>\n<p>What do you think may have caused this ?</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 3077157,
      "author_name": "Deepak Saldanha",
      "author_url": "",
      "post_date": "2024-12-20T16:19:03.883000",
      "content": "<p>Thank you for this challenging real world data, learnt a lot and Many Congratulations to the Winners!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3150238,
      "author_name": "Suzie_xtime",
      "author_url": "",
      "post_date": "2025-03-15T08:00:50.277000",
      "content": "<p>This is really difficult for a kaggle beginner</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3135608,
      "author_name": "Đặng Đào Xuân Trúc",
      "author_url": "",
      "post_date": "2025-02-27T16:33:20.363000",
      "content": "<p>perfect competition</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3078498,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-12-22T11:47:40.687000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3083596,
      "author_name": "benkerrouche abdelbasset",
      "author_url": "",
      "post_date": "2024-12-29T17:45:58.300000",
      "content": "<p>thank you </p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3077021": "Hi Kagglers!\n\nWe made it! On behalf of our team at the Child Mind Institute and our sponsors at the California Department of Health Care Services, Dell Technologies, and NVIDIA, thank you all SO much for participating in our competition over the last three months! We couldn’t be happier to have the participation of over 4,500 of you, submitting over 85,000 submissions to tackle our challenge! Truly, we are grateful for the time and thought you invested in solving our problem. The contributions of each and every one of you, through your submissions, interactions on the discussion board, and exploration of our dataset, are hugely valuable and we look forward to digesting them all and learning from you.\n\nIn the past few days it’s been fun and insightful to read the discussion forums and see you, the community, examine the data features, visualize interesting patterns, learn from the data, and help each other with model definition and other data science aspects of the competition.\n\n### Leaderboard Shake-Up: What Happened?\nWe also noticed that you started to recognize a trend that we’ve observed behind-the-scenes for a while: there’s plenty of opportunity for overfitting in our competition, and as a result, a large shake-up on the leaderboard. There have been several interesting threads where you’ve each shared what were [your signs for detecting overfitting](https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/551758), making [predictions about the distribution of Problematic Internet Use in the test set](https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/551913), and [guessing what the peak performance in the private test dataset will be](https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/551164) (which were generally more conservative than we see in practice). Across each, it’s clear that you have started to recognize the potential for overfitting by optimizing for leaderboard performance over consistency in cross-validation.\n\nThere are plenty of other examples of Kaggle competitions with shake-ups (interestingly, according to this [Meta-Kaggle post](https://www.kaggle.com/code/jtrotman/meta-kaggle-scatter-plot-competition-shake-up), the largest shake-up in Kaggle history is also [in the space of healthcare](https://www.kaggle.com/c/icr-identify-age-related-conditions)). \n\nTo give some perspective, when I cloned our leaderboard on 12/18 (apologies if it’s slightly out of date at the time of posting), I calculated our shakeup to be **0.255**, which is around 15–20th all-time on Kaggle according to the Meta-Kaggle post I mentioned above — putting us right between two NFL-based competitions ([1](https://www.kaggle.com/competitions/data-science-bowl-2017), [2](https://www.kaggle.com/c/nfl-impact-detection)). Now, I guess the more interesting question, *what is the source of our shake-up*?\n\n### Key Hypotheses: Overfitting or Not?\nFirst, a few details. We stratified our sample explicitly across binned age, sex, and the presence of actigraphy data. In practice, we also ended up with a balanced split of all of our instruments (at most, the availability of an instrument differed by 4% between the public and private test set). The proportion of participants in each class was also balanced across the splits, and you can see the breakdown below:\n\n|sii |Overall |Train |Public |Private |\n|:--|:--|:--|:--|:--|\n| 0 | 0.577 | 0.581 | 0.587 | 0.561 |\n| 1 | 0.272 | 0.275 | 0.258 | 0.273 |\n| 2 | 0.14 | 0.133 | 0.144 | 0.154 |\n| 3 | 0.012 | 0.012 | 0.011 | 0.013 |\n\nGiven that our data are relatively well balanced, *what is going on here*? \n\nThere are more than a few hypotheses… Let’s first assume that there wasn’t actually any overfitting of submissions in this competition, but that models had different strengths. If this were the case, it’s possible that models optimized with higher sensitivity to sii severity (i.e., were more likely to predict a higher score) were those that raised up the private leaderboard. This is possible given the small change in the distributions of sii classes between the private and public test sets. In this scenario, models that performed the best on the public set could have been doing a better job at predicting participants with an sii=0, of which there was a slightly smaller portion in the private test set.\n\nIn practice, we’re not sure that this is the case. We saw some notebooks tinkering with minor model hyperparameters, thresholds, or random seeds to climb the public leaderboard. Another sign of overfitting on our public leaderboard is the submission count. If we look at the top 10 teams in the public leaderboard versus the private leaderboard, the public-leaders on average submitted 212 models, whereas the private-leaders on average submitted 64. If we look at the median, these values are 199 vs 25. This suggests that there was certainly some work to cook scores on the public leaderboard, that didn’t result in success on our held out private data.\n\nI think it’s relatively safe to rule out the hypothesis that there isn’t overfitting here, which now begs the question, *what about our data lends itself to overfitting or poor generalization*? \n\nThere have been many (extremely informative!) discussion posts over the course of this competition that highlight possibilities: the possibility of batch effects across our samples; there could be data entry errors that limit reliability in the dataset; or the assessment instrument/prediction target is noisy and biased. All of these are potentially valid hypotheses, and deserve further exploration to understand. While we can evaluate the presence of batch effects in our data with tools like ComBAT, and automatically detect obvious data entry errors, evaluating the bias of our instruments is another can of worms altogether, and often impossible to do after data collection has taken place. If other assessments with similar purposes were deliberately included in data collection, this may be possible, but even in that case, the bias of the other assessments would likely need to be explored independently, perpetuating this challenge. In addition to the data challenges we’ve already mentioned, the nature of working with such heterogeneous datasets with relatively few samples is that there may simply not be strong enough statistical associations to give stable rankings — or critically, predictions when applied in real-world contexts — even in the absence of overfitting.\n\n### Real-World Data, Real-World Challenges\nBatch effects, data entry errors, measurement biases, and dataset heterogeneity are all core challenges that many psychological sciences face on a regular basis. This competition reflects those “real-world” challenges, rather than hiding them or proposing sanitized data that researchers, data scientists, and clinicians rarely work with. Moreover, mental health and the definitions of “within expected ranges” or “acceptable” vary across cultures, ages, genders, specific samples (for example, clinical vs non-clinical sample), and more: variation is a challenge inherent to mental health data, and highly generalizable performance is often elusive. This competition showed itself to be no exception!\n\nHere, we aimed to offer participants the possibility of working with real data from the mental health space, to collectively learn from the data, identify gaps in data collection practices, bring innovation to the field, and hopefully have fun by participating! We feel that we’ve learned an immense amount from all of you, and are grateful for your company on this adventure. We look forward to debriefing with a handful of you at the winner’s calls soon. 🙂\n\n\nSincerely,\n-Greg\n\n\n```\nGregory Kiar, PhD\nResearch Scientist\nDirector, Center for Data Analytics, Innovation, and Rigor\nChild Mind Institute\n```",
    "3077524": "Thank you for organizing this competition! Recently, a Kaggler pointed out the striking differences between missing values proportion in train vs test data [here](https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/552488#3076452) - this issue was not raised in this post\n\nWhat do you think may have caused this ?",
    "3077157": "Thank you for this challenging real world data, learnt a lot and Many Congratulations to the Winners!",
    "3150238": "This is really difficult for a kaggle beginner",
    "3135608": "perfect competition",
    "3078498": "",
    "3083596": "thank you "
  }
}