{
  "id": 541007,
  "title": "Correlations with the Questions",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/541007",
  "author_name": "",
  "post_date": "2024-10-17T02:50:12.620497300Z",
  "votes": 12,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Here's a correlation heatmap that shows correlations between the 20 questions, some csv features that are highly correlated with the PCIAT Total, and a few simple features made from the parquet files, \"Wrist-…\".  (The heatmap code is based on the code in <a href=\"https://www.kaggle.com/code/docxian/cmi-problematic-internet-use-visual-starter\" target=\"_blank\">CMI Problematic Internet Use - Visual Starter</a> ) There are some interesting things:</p>\n<p>In the 20x20 square of question-question correlations, questions 4, 7, and 12 stand out has having less correlation with the other questions.</p>\n<p>In the square of csv-feature with csv-feature correlations, two really stand out as less correlated with the others: the PreInt_…_hoursday and the SDS-SDS_Total_T; perhaps also FGC-FGC_CU and FGC-FGC_PU.</p>\n<p>Looking at the correlation of the questions with the csv features and the few Wrist- features, the highest single-question correlations show for question 7 and the lowest correlations are seen for question 16. The Wrist- features are also higher correlated with question 7. Note that the Wrist- features are all negatively correlation with the questions and csv features - more activity, less internet problems.</p>\n<p>The very last row shows the correlation of Total with the questions and features - again questions 4, 7, and 12 stand out from the others as being less correlated. <br>\nOf course (linear) correlation isn't everything and perhaps the ML will find relationships by combining features that together are important for prediction.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1496178%2F89833b4bebd09831e2c8827a56b8b018%2FQuestions_Correlations.png?generation=1729130995529783&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "3019921",
      "postDate": "10/17/2024 02:50:12",
      "content": "<p>Here's a correlation heatmap that shows correlations between the 20 questions, some csv features that are highly correlated with the PCIAT Total, and a few simple features made from the parquet files, \"Wrist-…\".  (The heatmap code is based on the code in <a href=\"https://www.kaggle.com/code/docxian/cmi-problematic-internet-use-visual-starter\" target=\"_blank\">CMI Problematic Internet Use - Visual Starter</a> ) There are some interesting things:</p>\n<p>In the 20x20 square of question-question correlations, questions 4, 7, and 12 stand out has having less correlation with the other questions.</p>\n<p>In the square of csv-feature with csv-feature correlations, two really stand out as less correlated with the others: the PreInt_…_hoursday and the SDS-SDS_Total_T; perhaps also FGC-FGC_CU and FGC-FGC_PU.</p>\n<p>Looking at the correlation of the questions with the csv features and the few Wrist- features, the highest single-question correlations show for question 7 and the lowest correlations are seen for question 16. The Wrist- features are also higher correlated with question 7. Note that the Wrist- features are all negatively correlation with the questions and csv features - more activity, less internet problems.</p>\n<p>The very last row shows the correlation of Total with the questions and features - again questions 4, 7, and 12 stand out from the others as being less correlated. <br>\nOf course (linear) correlation isn't everything and perhaps the ML will find relationships by combining features that together are important for prediction.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1496178%2F89833b4bebd09831e2c8827a56b8b018%2FQuestions_Correlations.png?generation=1729130995529783&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Here's a correlation heatmap that shows correlations between the 20 questions, some csv features that are highly correlated with the PCIAT Total, and a few simple features made from the parquet files, \"Wrist-...\".  (The heatmap code is based on the code in [CMI Problematic Internet Use - Visual Starter](https://www.kaggle.com/code/docxian/cmi-problematic-internet-use-visual-starter) ) There are some interesting things:\n\nIn the 20x20 square of question-question correlations, questions 4, 7, and 12 stand out has having less correlation with the other questions.\n\nIn the square of csv-feature with csv-feature correlations, two really stand out as less correlated with the others: the PreInt_..._hoursday and the SDS-SDS_Total_T; perhaps also FGC-FGC_CU and FGC-FGC_PU.\n\nLooking at the correlation of the questions with the csv features and the few Wrist- features, the highest single-question correlations show for question 7 and the lowest correlations are seen for question 16. The Wrist- features are also higher correlated with question 7. Note that the Wrist- features are all negatively correlation with the questions and csv features - more activity, less internet problems.\n\nThe very last row shows the correlation of Total with the questions and features - again questions 4, 7, and 12 stand out from the others as being less correlated. \nOf course (linear) correlation isn't everything and perhaps the ML will find relationships by combining features that together are important for prediction.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1496178%2F89833b4bebd09831e2c8827a56b8b018%2FQuestions_Correlations.png?generation=1729130995529783&alt=media)",
      "votes": null
    },
    {
      "id": "3020153",
      "postDate": "10/17/2024 07:55:50",
      "content": "<p>Very nice graph. Thank you for sharing.<br>\nAnd we can see PICAT 3&amp;5, and 16&amp;18 pair are having high correlation. Just checked the question and they're in fact asking the same question in 2 different ways.<br>\nYes, 4 7 12 are asking some special questions that tried to raise some red flags. So the bad news is while correlation with everything else is low, it will be difficult to capture this pattern in test set. </p>",
      "rawMarkdown": "Very nice graph. Thank you for sharing.\nAnd we can see PICAT 3&5, and 16&18 pair are having high correlation. Just checked the question and they're in fact asking the same question in 2 different ways.\nYes, 4 7 12 are asking some special questions that tried to raise some red flags. So the bad news is while correlation with everything else is low, it will be difficult to capture this pattern in test set.",
      "votes": null
    },
    {
      "id": "3020444",
      "postDate": "10/17/2024 14:12:35",
      "content": "<p>Great work! Just keep in mind that in this dataset, age is a very significant confounding factor. For example, the PCIAT test is much more appropriate for adolescents, and some questions are simply not applicable to young children or adults. Also, the target variable increases with age, peaking during adolescence (forming an approximate U-shaped relationship, although the variability is huge). </p>\n<p>For example, there is a positive correlation between the target and height, weight, and waist circumference, meaning that taller and fatter individuals tend to have a higher SII. However, since these physical parameters increase with age - and we already know that SII tends to be highest in adolescents - this could indicate that they act as proxies for age (likely reflecting age-related trends). Sometimes, if you look at the correlations between a feature and the target within a particular age group, you may find that any association disappears.</p>",
      "rawMarkdown": "Great work! Just keep in mind that in this dataset, age is a very significant confounding factor. For example, the PCIAT test is much more appropriate for adolescents, and some questions are simply not applicable to young children or adults. Also, the target variable increases with age, peaking during adolescence (forming an approximate U-shaped relationship, although the variability is huge). \n\nFor example, there is a positive correlation between the target and height, weight, and waist circumference, meaning that taller and fatter individuals tend to have a higher SII. However, since these physical parameters increase with age - and we already know that SII tends to be highest in adolescents - this could indicate that they act as proxies for age (likely reflecting age-related trends). Sometimes, if you look at the correlations between a feature and the target within a particular age group, you may find that any association disappears.",
      "votes": null
    },
    {
      "id": "3022056",
      "postDate": "10/19/2024 07:53:09",
      "content": "<p>Yes, good point about the correlations with age; I get correlations of 0.371 and 0.422 for Weight and Height with PCIAT Total. Subtracting simple trends of these with age reduces their correlations to  0.138 and 0.045; the trends used are in the code and plots below. Although the de-trended variables feel like better, independent features to me, in some quick tests the ML (XGB) quality seems about the same -- either way, the ML finds what's important 😀 </p>\n<pre><code> weight_of_age(age):\n    \n     . + ((.-.)/.)*(age - .)\n height_of_age(age):\n    \n     -.*age** + .*age + .\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1496178%2Feae4b0d0ac0dd6476d9d9935596d4dd2%2Fweight_height_vs_age.jpg?generation=1729322653683664&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Yes, good point about the correlations with age; I get correlations of 0.371 and 0.422 for Weight and Height with PCIAT Total. Subtracting simple trends of these with age reduces their correlations to  0.138 and 0.045; the trends used are in the code and plots below. Although the de-trended variables feel like better, independent features to me, in some quick tests the ML (XGB) quality seems about the same -- either way, the ML finds what's important 😀 \n\n```\ndef weight_of_age(age):\n    # Line based on the median weights at 5 and 16 years:\n    return 47.3 + ((139.6-47.3)/11.0)*(age - 5.0)\ndef height_of_age(age):\n    # Use a quadratic fit to 5, 13, and 20 years:\n    return -0.0969388*age**2 + 4.19898*age + 24.7959\n```\n \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1496178%2Feae4b0d0ac0dd6476d9d9935596d4dd2%2Fweight_height_vs_age.jpg?generation=1729322653683664&alt=media)",
      "votes": null
    },
    {
      "id": "3024399",
      "postDate": "10/21/2024 15:49:05",
      "content": "<p>P.S. One way to undo/reduce the Age confounding is to use percentiles or z values for quantities that have a strong, 'natural' variation with age, like Height and Weight.  In a recent  <a href=\"https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/535354#3024016\" target=\"_blank\">discussion reply</a> <a href=\"https://www.kaggle.com/vulpesvulpes\" target=\"_blank\">@vulpesvulpes</a> pointed out that these are available, e.g., from the WHO.</p>\n<p>As a simpler approximation, I used fits to the medians with age to make simple quadratic and linear functions for Height and Weight (like the equ.s/plots I showed.) These functions can be used to make a (~standard) scaler:</p>\n<pre><code>Height = ( Height -  ) /  \n</code></pre>\n<p>Doing this reduced the correlations of Height and Weight with Age from 0.892 and 0.801, to 0.108 and 0.087.    Their correlations with PCIAT Total are reduce to 0.128 and 0.113; their feature importance in the model is also reduced, but by only a few places. </p>",
      "rawMarkdown": "P.S. One way to undo/reduce the Age confounding is to use percentiles or z values for quantities that have a strong, 'natural' variation with age, like Height and Weight.  In a recent  [discussion reply](https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/535354#3024016) @vulpesvulpes pointed out that these are available, e.g., from the WHO.\n\nAs a simpler approximation, I used fits to the medians with age to make simple quadratic and linear functions for Height and Weight (like the equ.s/plots I showed.) These functions can be used to make a (~standard) scaler:\n```\nHeight = ( Height - func(Age) ) /  func(Age)\n```\nDoing this reduced the correlations of Height and Weight with Age from 0.892 and 0.801, to 0.108 and 0.087.    Their correlations with PCIAT Total are reduce to 0.128 and 0.113; their feature importance in the model is also reduced, but by only a few places.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3020153,
      "author_name": "tomyuen",
      "author_url": "",
      "post_date": "10/17/2024 07:55:50",
      "content": "<p>Very nice graph. Thank you for sharing.<br>\nAnd we can see PICAT 3&amp;5, and 16&amp;18 pair are having high correlation. Just checked the question and they're in fact asking the same question in 2 different ways.<br>\nYes, 4 7 12 are asking some special questions that tried to raise some red flags. So the bad news is while correlation with everything else is low, it will be difficult to capture this pattern in test set. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3020444,
      "author_name": "antoninadolgorukova",
      "author_url": "",
      "post_date": "10/17/2024 14:12:35",
      "content": "<p>Great work! Just keep in mind that in this dataset, age is a very significant confounding factor. For example, the PCIAT test is much more appropriate for adolescents, and some questions are simply not applicable to young children or adults. Also, the target variable increases with age, peaking during adolescence (forming an approximate U-shaped relationship, although the variability is huge). </p>\n<p>For example, there is a positive correlation between the target and height, weight, and waist circumference, meaning that taller and fatter individuals tend to have a higher SII. However, since these physical parameters increase with age - and we already know that SII tends to be highest in adolescents - this could indicate that they act as proxies for age (likely reflecting age-related trends). Sometimes, if you look at the correlations between a feature and the target within a particular age group, you may find that any association disappears.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3022056,
          "author_name": "dan3dewey",
          "author_url": "",
          "post_date": "10/19/2024 07:53:09",
          "content": "<p>Yes, good point about the correlations with age; I get correlations of 0.371 and 0.422 for Weight and Height with PCIAT Total. Subtracting simple trends of these with age reduces their correlations to  0.138 and 0.045; the trends used are in the code and plots below. Although the de-trended variables feel like better, independent features to me, in some quick tests the ML (XGB) quality seems about the same -- either way, the ML finds what's important 😀 </p>\n<pre><code> weight_of_age(age):\n    \n     . + ((.-.)/.)*(age - .)\n height_of_age(age):\n    \n     -.*age** + .*age + .\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1496178%2Feae4b0d0ac0dd6476d9d9935596d4dd2%2Fweight_height_vs_age.jpg?generation=1729322653683664&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3024399,
          "author_name": "dan3dewey",
          "author_url": "",
          "post_date": "10/21/2024 15:49:05",
          "content": "<p>P.S. One way to undo/reduce the Age confounding is to use percentiles or z values for quantities that have a strong, 'natural' variation with age, like Height and Weight.  In a recent  <a href=\"https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/535354#3024016\" target=\"_blank\">discussion reply</a> <a href=\"https://www.kaggle.com/vulpesvulpes\" target=\"_blank\">@vulpesvulpes</a> pointed out that these are available, e.g., from the WHO.</p>\n<p>As a simpler approximation, I used fits to the medians with age to make simple quadratic and linear functions for Height and Weight (like the equ.s/plots I showed.) These functions can be used to make a (~standard) scaler:</p>\n<pre><code>Height = ( Height -  ) /  \n</code></pre>\n<p>Doing this reduced the correlations of Height and Weight with Age from 0.892 and 0.801, to 0.108 and 0.087.    Their correlations with PCIAT Total are reduce to 0.128 and 0.113; their feature importance in the model is also reduced, but by only a few places. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3019921": "Here's a correlation heatmap that shows correlations between the 20 questions, some csv features that are highly correlated with the PCIAT Total, and a few simple features made from the parquet files, \"Wrist-...\".  (The heatmap code is based on the code in [CMI Problematic Internet Use - Visual Starter](https://www.kaggle.com/code/docxian/cmi-problematic-internet-use-visual-starter) ) There are some interesting things:\n\nIn the 20x20 square of question-question correlations, questions 4, 7, and 12 stand out has having less correlation with the other questions.\n\nIn the square of csv-feature with csv-feature correlations, two really stand out as less correlated with the others: the PreInt_..._hoursday and the SDS-SDS_Total_T; perhaps also FGC-FGC_CU and FGC-FGC_PU.\n\nLooking at the correlation of the questions with the csv features and the few Wrist- features, the highest single-question correlations show for question 7 and the lowest correlations are seen for question 16. The Wrist- features are also higher correlated with question 7. Note that the Wrist- features are all negatively correlation with the questions and csv features - more activity, less internet problems.\n\nThe very last row shows the correlation of Total with the questions and features - again questions 4, 7, and 12 stand out from the others as being less correlated. \nOf course (linear) correlation isn't everything and perhaps the ML will find relationships by combining features that together are important for prediction.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1496178%2F89833b4bebd09831e2c8827a56b8b018%2FQuestions_Correlations.png?generation=1729130995529783&alt=media)",
    "3020153": "Very nice graph. Thank you for sharing.\nAnd we can see PICAT 3&5, and 16&18 pair are having high correlation. Just checked the question and they're in fact asking the same question in 2 different ways.\nYes, 4 7 12 are asking some special questions that tried to raise some red flags. So the bad news is while correlation with everything else is low, it will be difficult to capture this pattern in test set.",
    "3020444": "Great work! Just keep in mind that in this dataset, age is a very significant confounding factor. For example, the PCIAT test is much more appropriate for adolescents, and some questions are simply not applicable to young children or adults. Also, the target variable increases with age, peaking during adolescence (forming an approximate U-shaped relationship, although the variability is huge). \n\nFor example, there is a positive correlation between the target and height, weight, and waist circumference, meaning that taller and fatter individuals tend to have a higher SII. However, since these physical parameters increase with age - and we already know that SII tends to be highest in adolescents - this could indicate that they act as proxies for age (likely reflecting age-related trends). Sometimes, if you look at the correlations between a feature and the target within a particular age group, you may find that any association disappears.",
    "3022056": "Yes, good point about the correlations with age; I get correlations of 0.371 and 0.422 for Weight and Height with PCIAT Total. Subtracting simple trends of these with age reduces their correlations to  0.138 and 0.045; the trends used are in the code and plots below. Although the de-trended variables feel like better, independent features to me, in some quick tests the ML (XGB) quality seems about the same -- either way, the ML finds what's important 😀 \n\n```\ndef weight_of_age(age):\n    # Line based on the median weights at 5 and 16 years:\n    return 47.3 + ((139.6-47.3)/11.0)*(age - 5.0)\ndef height_of_age(age):\n    # Use a quadratic fit to 5, 13, and 20 years:\n    return -0.0969388*age**2 + 4.19898*age + 24.7959\n```\n \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1496178%2Feae4b0d0ac0dd6476d9d9935596d4dd2%2Fweight_height_vs_age.jpg?generation=1729322653683664&alt=media)",
    "3024399": "P.S. One way to undo/reduce the Age confounding is to use percentiles or z values for quantities that have a strong, 'natural' variation with age, like Height and Weight.  In a recent  [discussion reply](https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/535354#3024016) @vulpesvulpes pointed out that these are available, e.g., from the WHO.\n\nAs a simpler approximation, I used fits to the medians with age to make simple quadratic and linear functions for Height and Weight (like the equ.s/plots I showed.) These functions can be used to make a (~standard) scaler:\n```\nHeight = ( Height - func(Age) ) /  func(Age)\n```\nDoing this reduced the correlations of Height and Weight with Age from 0.892 and 0.801, to 0.108 and 0.087.    Their correlations with PCIAT Total are reduce to 0.128 and 0.113; their feature importance in the model is also reduced, but by only a few places."
  },
  "source": "meta"
}