{
  "id": 329534,
  "title": "Duplicate Customer IDs in test data and hence in submission",
  "url": "/competitions/amex-default-prediction/discussion/329534",
  "author_name": "",
  "post_date": "2022-06-07T09:26:18.487552400Z",
  "votes": 5,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hey folks, <br>\nWhen I run the model on test data file and left join the sample submission file with the test data file results on customer_ID, I get 11363762 rows instead of 924621 rows as in the sample submission file. This is due to duplicate customer IDs in the test data file. Is it OK to submit 11363762 rows in the submission file? If not, what could be possible method to aggregate model output (11363762 rows) in test data at unique customer_ID level (924621 rows)? Any thoughts/guidance on this is most welcome.</p>",
  "messages": [
    {
      "id": "1813834",
      "postDate": "06/07/2022 09:26:18",
      "content": "<p>Hey folks, <br>\nWhen I run the model on test data file and left join the sample submission file with the test data file results on customer_ID, I get 11363762 rows instead of 924621 rows as in the sample submission file. This is due to duplicate customer IDs in the test data file. Is it OK to submit 11363762 rows in the submission file? If not, what could be possible method to aggregate model output (11363762 rows) in test data at unique customer_ID level (924621 rows)? Any thoughts/guidance on this is most welcome.</p>",
      "rawMarkdown": "Hey folks, \nWhen I run the model on test data file and left join the sample submission file with the test data file results on customer_ID, I get 11363762 rows instead of 924621 rows as in the sample submission file. This is due to duplicate customer IDs in the test data file. Is it OK to submit 11363762 rows in the submission file? If not, what could be possible method to aggregate model output (11363762 rows) in test data at unique customer_ID level (924621 rows)? Any thoughts/guidance on this is most welcome.",
      "votes": null
    },
    {
      "id": "1813839",
      "postDate": "06/07/2022 09:29:46",
      "content": "<p>Your submission must only have each customer once. So you have two options. You can either make 1 inference per test customer and save those inferences. Or you can infer each customer multiple times and then aggregate mean the predictions.</p>",
      "rawMarkdown": "Your submission must only have each customer once. So you have two options. You can either make 1 inference per test customer and save those inferences. Or you can infer each customer multiple times and then aggregate mean the predictions.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1813839,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "06/07/2022 09:29:46",
      "content": "<p>Your submission must only have each customer once. So you have two options. You can either make 1 inference per test customer and save those inferences. Or you can infer each customer multiple times and then aggregate mean the predictions.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1813834": "Hey folks, \nWhen I run the model on test data file and left join the sample submission file with the test data file results on customer_ID, I get 11363762 rows instead of 924621 rows as in the sample submission file. This is due to duplicate customer IDs in the test data file. Is it OK to submit 11363762 rows in the submission file? If not, what could be possible method to aggregate model output (11363762 rows) in test data at unique customer_ID level (924621 rows)? Any thoughts/guidance on this is most welcome.",
    "1813839": "Your submission must only have each customer once. So you have two options. You can either make 1 inference per test customer and save those inferences. Or you can infer each customer multiple times and then aggregate mean the predictions."
  },
  "source": "meta"
}