{
  "id": 246710,
  "title": "Question about evaluation phase",
  "url": "/competitions/mlb-player-digital-engagement-forecasting/discussion/246710",
  "author_name": "",
  "post_date": "2021-06-16T14:24:54.832111600Z",
  "votes": 9,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hello!</p>\n<p>I didn't find any fields in test_df for previous target_1, …, target_4. So, during evaluation phase, I can use lag features, which are based only my own prediction?</p>\n<p>For example, while evaluating I predict targets for <strong>date</strong>, after that I have to predict targets for <strong>date + 1</strong>, but at that time I will not know the real value for targets for <strong>date</strong>?</p>",
  "messages": [
    {
      "id": "1351691",
      "postDate": "06/16/2021 14:24:54",
      "content": "<p>Hello!</p>\n<p>I didn't find any fields in test_df for previous target_1, …, target_4. So, during evaluation phase, I can use lag features, which are based only my own prediction?</p>\n<p>For example, while evaluating I predict targets for <strong>date</strong>, after that I have to predict targets for <strong>date + 1</strong>, but at that time I will not know the real value for targets for <strong>date</strong>?</p>",
      "rawMarkdown": "Hello!\n\nI didn't find any fields in test_df for previous target_1, ..., target_4. So, during evaluation phase, I can use lag features, which are based only my own prediction?\n\nFor example, while evaluating I predict targets for **date**, after that I have to predict targets for **date + 1**, but at that time I will not know the real value for targets for **date**?",
      "votes": null
    },
    {
      "id": "1351700",
      "postDate": "06/16/2021 14:30:02",
      "content": "<p>Correct. In this challenge the lagged ground-truth <code>target*</code> variables are not provided within the test set time window.</p>",
      "rawMarkdown": "Correct. In this challenge the lagged ground-truth `target*` variables are not provided within the test set time window.",
      "votes": null
    },
    {
      "id": "1351705",
      "postDate": "06/16/2021 14:32:11",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!",
      "votes": null
    },
    {
      "id": "1354478",
      "postDate": "06/17/2021 15:58:15",
      "content": "<p>Good question! This is what I wanted to know!</p>",
      "rawMarkdown": "Good question! This is what I wanted to know!",
      "votes": null
    },
    {
      "id": "1354606",
      "postDate": "06/17/2021 17:49:51",
      "content": "<p><a href=\"https://www.kaggle.com/wcukierski\" target=\"_blank\">Will</a> - my memory is often suspect but it tells me that in past competitions that used a similar API approach the target values did not get provided immediately but delayed by a batch size.  The API intent was that you made the prediction for a batch and than could update with the truth for batch 1 upon arrival of batch 2.  (but you could not go back in time and \"fix\" your predictions)</p>\n<p>Your answer here and from other posts indicates that July 31 will be the last target data provided up to final LB results on 9/15.   I really hate to ask an answered question but would you please confirm that the hosts desire a model that can predict out 30 to 45 days in the future.   </p>",
      "rawMarkdown": "[Will](https://www.kaggle.com/wcukierski) - my memory is often suspect but it tells me that in past competitions that used a similar API approach the target values did not get provided immediately but delayed by a batch size.  The API intent was that you made the prediction for a batch and than could update with the truth for batch 1 upon arrival of batch 2.  (but you could not go back in time and \"fix\" your predictions)\n\nYour answer here and from other posts indicates that July 31 will be the last target data provided up to final LB results on 9/15.   I really hate to ask an answered question but would you please confirm that the hosts desire a model that can predict out 30 to 45 days in the future.",
      "votes": null
    },
    {
      "id": "1357266",
      "postDate": "06/19/2021 15:15:54",
      "content": "<p><a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> that's a good question. I'd like to know as well. </p>",
      "rawMarkdown": "pcjimmmy that's a good question. I'd like to know as well.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1351700,
      "author_name": "wcukierski",
      "author_url": "",
      "post_date": "06/16/2021 14:30:02",
      "content": "<p>Correct. In this challenge the lagged ground-truth <code>target*</code> variables are not provided within the test set time window.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1351705,
          "author_name": "petrchuikov",
          "author_url": "",
          "post_date": "06/16/2021 14:32:11",
          "content": "<p>Thank you!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1354606,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "06/17/2021 17:49:51",
          "content": "<p><a href=\"https://www.kaggle.com/wcukierski\" target=\"_blank\">Will</a> - my memory is often suspect but it tells me that in past competitions that used a similar API approach the target values did not get provided immediately but delayed by a batch size.  The API intent was that you made the prediction for a batch and than could update with the truth for batch 1 upon arrival of batch 2.  (but you could not go back in time and \"fix\" your predictions)</p>\n<p>Your answer here and from other posts indicates that July 31 will be the last target data provided up to final LB results on 9/15.   I really hate to ask an answered question but would you please confirm that the hosts desire a model that can predict out 30 to 45 days in the future.   </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1357266,
          "author_name": "crained",
          "author_url": "",
          "post_date": "06/19/2021 15:15:54",
          "content": "<p><a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> that's a good question. I'd like to know as well. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1354478,
      "author_name": "tomokikmogura",
      "author_url": "",
      "post_date": "06/17/2021 15:58:15",
      "content": "<p>Good question! This is what I wanted to know!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1351691": "Hello!\n\nI didn't find any fields in test_df for previous target_1, ..., target_4. So, during evaluation phase, I can use lag features, which are based only my own prediction?\n\nFor example, while evaluating I predict targets for **date**, after that I have to predict targets for **date + 1**, but at that time I will not know the real value for targets for **date**?",
    "1351700": "Correct. In this challenge the lagged ground-truth `target*` variables are not provided within the test set time window.",
    "1351705": "Thank you!",
    "1354478": "Good question! This is what I wanted to know!",
    "1354606": "[Will](https://www.kaggle.com/wcukierski) - my memory is often suspect but it tells me that in past competitions that used a similar API approach the target values did not get provided immediately but delayed by a batch size.  The API intent was that you made the prediction for a batch and than could update with the truth for batch 1 upon arrival of batch 2.  (but you could not go back in time and \"fix\" your predictions)\n\nYour answer here and from other posts indicates that July 31 will be the last target data provided up to final LB results on 9/15.   I really hate to ask an answered question but would you please confirm that the hosts desire a model that can predict out 30 to 45 days in the future.",
    "1357266": "pcjimmmy that's a good question. I'd like to know as well."
  },
  "source": "meta"
}