{
  "id": 538082,
  "title": "Are we using the future to predict the past?",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/538082",
  "author_name": "ExpensiveLunch",
  "post_date": "2024-10-06T21:42:19.083000",
  "votes": 23,
  "comment_count": 7,
  "views": 0,
  "content": "<p>The \"relative_date_PCIAT\" field in the actigraphy data is defined as the \"number of days (integer) since the PCIAT test was administered (negative days indicate that the actigraphy data has been collected before the test was administered)\".</p>\n<p>Only 59 actigraphy files have any data with relative_date_PCIAT &lt;0, whereas most of the files have some data with relative_date_PCIAT &gt; 0. </p>\n<p>The target (sii) is derived from the field PCIAT-PCIAT_Total.  </p>\n<p>Unless I've misunderstood, if I use the relative_date_PCIAT &gt;0 data, I'm using data that was collected after sii was observed.  In other words, I'm using the future to predict the past.  I'm using movement data collected after the child's level of problematic internet usage (if any) was decided.</p>\n<p>It's up to the Child Mind Institute to decide what prediction problem is of interest to them. I can certainly see how some prediction problems which use the future to predict the past could be of interest (e.g., using college grades to predict high school grades). </p>\n<p>Nonetheless, it's a bit surprising to me that we are given predictors collected after the target was collected. Deploying a model trained on such predictors in a real-world setting might be problematic.</p>\n<p>Or have I misunderstood something ?  </p>",
  "messages": [
    {
      "id": 3008635,
      "postDate": "2024-10-06T21:42:19.083Z",
      "content": "<p>The \"relative_date_PCIAT\" field in the actigraphy data is defined as the \"number of days (integer) since the PCIAT test was administered (negative days indicate that the actigraphy data has been collected before the test was administered)\".</p>\n<p>Only 59 actigraphy files have any data with relative_date_PCIAT &lt;0, whereas most of the files have some data with relative_date_PCIAT &gt; 0. </p>\n<p>The target (sii) is derived from the field PCIAT-PCIAT_Total.  </p>\n<p>Unless I've misunderstood, if I use the relative_date_PCIAT &gt;0 data, I'm using data that was collected after sii was observed.  In other words, I'm using the future to predict the past.  I'm using movement data collected after the child's level of problematic internet usage (if any) was decided.</p>\n<p>It's up to the Child Mind Institute to decide what prediction problem is of interest to them. I can certainly see how some prediction problems which use the future to predict the past could be of interest (e.g., using college grades to predict high school grades). </p>\n<p>Nonetheless, it's a bit surprising to me that we are given predictors collected after the target was collected. Deploying a model trained on such predictors in a real-world setting might be problematic.</p>\n<p>Or have I misunderstood something ?  </p>",
      "rawMarkdown": "The \"relative_date_PCIAT\" field in the actigraphy data is defined as the \"number of days (integer) since the PCIAT test was administered (negative days indicate that the actigraphy data has been collected before the test was administered)\".\n\nOnly 59 actigraphy files have any data with relative_date_PCIAT <0, whereas most of the files have some data with relative_date_PCIAT > 0. \n\nThe target (sii) is derived from the field PCIAT-PCIAT_Total.  \n\nUnless I've misunderstood, if I use the relative_date_PCIAT >0 data, I'm using data that was collected after sii was observed.  In other words, I'm using the future to predict the past.  I'm using movement data collected after the child's level of problematic internet usage (if any) was decided.\n\nIt's up to the Child Mind Institute to decide what prediction problem is of interest to them. I can certainly see how some prediction problems which use the future to predict the past could be of interest (e.g., using college grades to predict high school grades). \n\nNonetheless, it's a bit surprising to me that we are given predictors collected after the target was collected. Deploying a model trained on such predictors in a real-world setting might be problematic.\n\nOr have I misunderstood something ?  \n\n",
      "votes": 22
    },
    {
      "id": 3017467,
      "postDate": "2024-10-15T00:51:35.210Z",
      "content": "<p>It looks like most of the PCIAT relative dates are positive, so at least whatever effect there is is common to most samples.</p>",
      "rawMarkdown": "It looks like most of the PCIAT relative dates are positive, so at least whatever effect there is is common to most samples.",
      "votes": 1,
      "replies": [
        {
          "id": 3017773,
          "postDate": "2024-10-15T07:32:29.520Z",
          "content": "<p>No one can guarantee that the hidden test data was collected in the same way with the same proportions after/before PCIAT. Also, I don't think the bias is negligible, and I agree with <a href=\"https://www.kaggle.com/gkitchen\" target=\"_blank\">@gkitchen</a> 's point: if parents become aware that their child has problematic internet use (PIU) after administering the PCIAT test, they may take immediate action to change their child's behaviour. Even if they don't know the result, the very fact that they have thought about it and the potential harm of PIU may lead them to change something, even subconsciously (such as encouraging their child to be more active, enforcing new rules around internet use, etc.). So the activity after the PCIAT test may be very different from before. </p>\n<p>And the model will still be a priori inaccurate in practical implementations (impossible to collect future data). <br>\nTrying to model something you know won't work in real life is a strange idea…</p>",
          "rawMarkdown": "No one can guarantee that the hidden test data was collected in the same way with the same proportions after/before PCIAT. Also, I don't think the bias is negligible, and I agree with @gkitchen 's point: if parents become aware that their child has problematic internet use (PIU) after administering the PCIAT test, they may take immediate action to change their child's behaviour. Even if they don't know the result, the very fact that they have thought about it and the potential harm of PIU may lead them to change something, even subconsciously (such as encouraging their child to be more active, enforcing new rules around internet use, etc.). So the activity after the PCIAT test may be very different from before. \n\nAnd the model will still be a priori inaccurate in practical implementations (impossible to collect future data). \nTrying to model something you know won't work in real life is a strange idea...",
          "votes": 3
        }
      ]
    },
    {
      "id": 3017167,
      "postDate": "2024-10-14T15:45:03.297Z",
      "content": "<p>Your understanding seems correct, at least I see this the same way… This is just another thing introducing inherent inaccuracy into the modelling… sorry I mentioned your point in my \"questioning the task\" tread but forgot to answer you … This is a good catch!</p>",
      "rawMarkdown": "Your understanding seems correct, at least I see this the same way... This is just another thing introducing inherent inaccuracy into the modelling... sorry I mentioned your point in my \"questioning the task\" tread but forgot to answer you ... This is a good catch!",
      "votes": 1
    },
    {
      "id": 3012455,
      "postDate": "2024-10-09T04:03:52.297Z",
      "content": "<p>I don't think the leakage is as significant as it may seem. Most features computable from the time series data are aggregations and in theory should be representative of the subject's daily life, regardless of when the sample is recorded.</p>",
      "rawMarkdown": "I don't think the leakage is as significant as it may seem. Most features computable from the time series data are aggregations and in theory should be representative of the subject's daily life, regardless of when the sample is recorded.",
      "votes": 1,
      "replies": [
        {
          "id": 3012849,
          "postDate": "2024-10-09T12:19:11.327Z",
          "content": "<p>Yes , You're right . <br>\nBut still it's little bit confusing to accept this . If we look at this as a traditional  machine learning approach.</p>",
          "rawMarkdown": "Yes , You're right . \nBut still it's little bit confusing to accept this . If we look at this as a traditional  machine learning approach."
        },
        {
          "id": 3012882,
          "postDate": "2024-10-09T13:07:36.210Z",
          "content": "<p>OK, thanks very much for the response. I agree that a reasonable (but not ironclad) argument can be made that the leakage is not a big deal. I mostly just wanted to confirm that I hadn't misunderstood something. </p>",
          "rawMarkdown": "OK, thanks very much for the response. I agree that a reasonable (but not ironclad) argument can be made that the leakage is not a big deal. I mostly just wanted to confirm that I hadn't misunderstood something. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 3017157,
      "postDate": "2024-10-14T15:19:39.813Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3017467,
      "author_name": "Daniel Dewey",
      "author_url": "",
      "post_date": "2024-10-15T00:51:35.210000",
      "content": "<p>It looks like most of the PCIAT relative dates are positive, so at least whatever effect there is is common to most samples.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3017773,
          "author_name": "Antonina Dolgorukova",
          "author_url": "",
          "post_date": "2024-10-15T07:32:29.520000",
          "content": "<p>No one can guarantee that the hidden test data was collected in the same way with the same proportions after/before PCIAT. Also, I don't think the bias is negligible, and I agree with <a href=\"https://www.kaggle.com/gkitchen\" target=\"_blank\">@gkitchen</a> 's point: if parents become aware that their child has problematic internet use (PIU) after administering the PCIAT test, they may take immediate action to change their child's behaviour. Even if they don't know the result, the very fact that they have thought about it and the potential harm of PIU may lead them to change something, even subconsciously (such as encouraging their child to be more active, enforcing new rules around internet use, etc.). So the activity after the PCIAT test may be very different from before. </p>\n<p>And the model will still be a priori inaccurate in practical implementations (impossible to collect future data). <br>\nTrying to model something you know won't work in real life is a strange idea…</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 3017167,
      "author_name": "Antonina Dolgorukova",
      "author_url": "",
      "post_date": "2024-10-14T15:45:03.297000",
      "content": "<p>Your understanding seems correct, at least I see this the same way… This is just another thing introducing inherent inaccuracy into the modelling… sorry I mentioned your point in my \"questioning the task\" tread but forgot to answer you … This is a good catch!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3012455,
      "author_name": "Sam",
      "author_url": "",
      "post_date": "2024-10-09T04:03:52.297000",
      "content": "<p>I don't think the leakage is as significant as it may seem. Most features computable from the time series data are aggregations and in theory should be representative of the subject's daily life, regardless of when the sample is recorded.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3012849,
          "author_name": "Eman Gope",
          "author_url": "",
          "post_date": "2024-10-09T12:19:11.327000",
          "content": "<p>Yes , You're right . <br>\nBut still it's little bit confusing to accept this . If we look at this as a traditional  machine learning approach.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3012882,
          "author_name": "ExpensiveLunch",
          "author_url": "",
          "post_date": "2024-10-09T13:07:36.210000",
          "content": "<p>OK, thanks very much for the response. I agree that a reasonable (but not ironclad) argument can be made that the leakage is not a big deal. I mostly just wanted to confirm that I hadn't misunderstood something. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3017157,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-10-14T15:19:39.813000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3008635": "The \"relative_date_PCIAT\" field in the actigraphy data is defined as the \"number of days (integer) since the PCIAT test was administered (negative days indicate that the actigraphy data has been collected before the test was administered)\".\n\nOnly 59 actigraphy files have any data with relative_date_PCIAT <0, whereas most of the files have some data with relative_date_PCIAT > 0. \n\nThe target (sii) is derived from the field PCIAT-PCIAT_Total.  \n\nUnless I've misunderstood, if I use the relative_date_PCIAT >0 data, I'm using data that was collected after sii was observed.  In other words, I'm using the future to predict the past.  I'm using movement data collected after the child's level of problematic internet usage (if any) was decided.\n\nIt's up to the Child Mind Institute to decide what prediction problem is of interest to them. I can certainly see how some prediction problems which use the future to predict the past could be of interest (e.g., using college grades to predict high school grades). \n\nNonetheless, it's a bit surprising to me that we are given predictors collected after the target was collected. Deploying a model trained on such predictors in a real-world setting might be problematic.\n\nOr have I misunderstood something ?  \n\n",
    "3017467": "It looks like most of the PCIAT relative dates are positive, so at least whatever effect there is is common to most samples.",
    "3017167": "Your understanding seems correct, at least I see this the same way... This is just another thing introducing inherent inaccuracy into the modelling... sorry I mentioned your point in my \"questioning the task\" tread but forgot to answer you ... This is a good catch!",
    "3012455": "I don't think the leakage is as significant as it may seem. Most features computable from the time series data are aggregations and in theory should be representative of the subject's daily life, regardless of when the sample is recorded.",
    "3017157": ""
  }
}