{
  "id": 339787,
  "title": "How to Aggregate the test data?",
  "url": "/competitions/amex-default-prediction/discussion/339787",
  "author_name": "",
  "post_date": "2022-07-26T11:30:29.124052500Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I've seen the test data had more than one row by customer_ID. My question is, should I submite considering just the last observation (based on S_2) or should I just do  a groupby('customer_ID').max()?</p>",
  "messages": [
    {
      "id": "1871617",
      "postDate": "07/26/2022 11:30:29",
      "content": "<p>I've seen the test data had more than one row by customer_ID. My question is, should I submite considering just the last observation (based on S_2) or should I just do  a groupby('customer_ID').max()?</p>",
      "rawMarkdown": "I've seen the test data had more than one row by customer_ID. My question is, should I submite considering just the last observation (based on S_2) or should I just do  a groupby('customer_ID').max()?",
      "votes": null
    },
    {
      "id": "1872008",
      "postDate": "07/26/2022 15:38:31",
      "content": "<p>You should aggregate the test data in the <strong>same</strong> way that you aggregated your train data during training.</p>",
      "rawMarkdown": "You should aggregate the test data in the **same** way that you aggregated your train data during training.",
      "votes": null
    },
    {
      "id": "1872139",
      "postDate": "07/26/2022 17:39:12",
      "content": "<p>But how can I know what label use at each date in training?<br>\nSome customer_ID have target 0 and 1, how the match between the sets are done?</p>",
      "rawMarkdown": "But how can I know what label use at each date in training?\nSome customer_ID have target 0 and 1, how the match between the sets are done?",
      "votes": null
    },
    {
      "id": "1872214",
      "postDate": "07/26/2022 18:10:00",
      "content": "<p>Each train customer has only 1 target in the file <code>train_labels.csv</code>. So even though train customers have up to 13 rows in <code>train_data.csv</code>, you train your model with 1 target per customer. Therefore during test inference, even though test customers have up to 13 rows in <code>test_data.csv</code>, you predict 1 target per customer.</p>",
      "rawMarkdown": "Each train customer has only 1 target in the file `train_labels.csv`. So even though train customers have up to 13 rows in `train_data.csv`, you train your model with 1 target per customer. Therefore during test inference, even though test customers have up to 13 rows in `test_data.csv`, you predict 1 target per customer.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1872008,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "07/26/2022 15:38:31",
      "content": "<p>You should aggregate the test data in the <strong>same</strong> way that you aggregated your train data during training.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1872139,
          "author_name": "alexandreg1998",
          "author_url": "",
          "post_date": "07/26/2022 17:39:12",
          "content": "<p>But how can I know what label use at each date in training?<br>\nSome customer_ID have target 0 and 1, how the match between the sets are done?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1872214,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "07/26/2022 18:10:00",
          "content": "<p>Each train customer has only 1 target in the file <code>train_labels.csv</code>. So even though train customers have up to 13 rows in <code>train_data.csv</code>, you train your model with 1 target per customer. Therefore during test inference, even though test customers have up to 13 rows in <code>test_data.csv</code>, you predict 1 target per customer.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1871617": "I've seen the test data had more than one row by customer_ID. My question is, should I submite considering just the last observation (based on S_2) or should I just do  a groupby('customer_ID').max()?",
    "1872008": "You should aggregate the test data in the **same** way that you aggregated your train data during training.",
    "1872139": "But how can I know what label use at each date in training?\nSome customer_ID have target 0 and 1, how the match between the sets are done?",
    "1872214": "Each train customer has only 1 target in the file `train_labels.csv`. So even though train customers have up to 13 rows in `train_data.csv`, you train your model with 1 target per customer. Therefore during test inference, even though test customers have up to 13 rows in `test_data.csv`, you predict 1 target per customer."
  },
  "source": "meta"
}