{
  "id": 338289,
  "title": "Duplicates in test dataset ",
  "url": "/competitions/amex-default-prediction/discussion/338289",
  "author_name": "",
  "post_date": "2022-07-20T01:17:14.937052900Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Can anyone share how have you dealt with duplicates in the test dataset? </p>",
  "messages": [
    {
      "id": "1862747",
      "postDate": "07/20/2022 01:17:14",
      "content": "<p>Can anyone share how have you dealt with duplicates in the test dataset? </p>",
      "rawMarkdown": "Can anyone share how have you dealt with duplicates in the test dataset?",
      "votes": null
    },
    {
      "id": "1862764",
      "postDate": "07/20/2022 01:48:43",
      "content": "<p>There are no duplicates in the test dataset. All users have multiple time points corresponding to their credit card statements from different points in time - see <code>S_2</code> variable.</p>",
      "rawMarkdown": "There are no duplicates in the test dataset. All users have multiple time points corresponding to their credit card statements from different points in time - see `S_2` variable.",
      "votes": null
    },
    {
      "id": "1862785",
      "postDate": "07/20/2022 02:17:57",
      "content": "<p>Yes, but the submission should have each customer once. I wonder how to aggregate mean for predictions??</p>",
      "rawMarkdown": "Yes, but the submission should have each customer once. I wonder how to aggregate mean for predictions??",
      "votes": null
    },
    {
      "id": "1862832",
      "postDate": "07/20/2022 03:40:02",
      "content": "<p>However you are aggregating train data to prepare them for training, you should do the same with test data. That is part of the preprocessing procedure and it should be the same for both datasets.</p>\n<p>If you haven't aggregated either train or test data, meaning you are treating every single sample in train individually, you will likely overfit during training. But if you really want to do it that way, and assuming that after prediction you got a dataframe <code>df</code> that has columns <code>['custom_ID', 'prediction']</code>:</p>\n<pre><code>df_new = df.groupby(\"customer_ID\").mean().reset_index()\n</code></pre>\n<p>That will group all the rows that have the same ID and take their mean value into new dataframe, which now can be submitted.</p>",
      "rawMarkdown": "However you are aggregating train data to prepare them for training, you should do the same with test data. That is part of the preprocessing procedure and it should be the same for both datasets.\n\nIf you haven't aggregated either train or test data, meaning you are treating every single sample in train individually, you will likely overfit during training. But if you really want to do it that way, and assuming that after prediction you got a dataframe `df` that has columns `['custom_ID', 'prediction']`:\n\n    df_new = df.groupby(\"customer_ID\").mean().reset_index()\n\nThat will group all the rows that have the same ID and take their mean value into new dataframe, which now can be submitted.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1862764,
      "author_name": "tilii7",
      "author_url": "",
      "post_date": "07/20/2022 01:48:43",
      "content": "<p>There are no duplicates in the test dataset. All users have multiple time points corresponding to their credit card statements from different points in time - see <code>S_2</code> variable.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1862785,
      "author_name": "nozamd",
      "author_url": "",
      "post_date": "07/20/2022 02:17:57",
      "content": "<p>Yes, but the submission should have each customer once. I wonder how to aggregate mean for predictions??</p>",
      "votes": null,
      "replies": [
        {
          "id": 1862832,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "07/20/2022 03:40:02",
          "content": "<p>However you are aggregating train data to prepare them for training, you should do the same with test data. That is part of the preprocessing procedure and it should be the same for both datasets.</p>\n<p>If you haven't aggregated either train or test data, meaning you are treating every single sample in train individually, you will likely overfit during training. But if you really want to do it that way, and assuming that after prediction you got a dataframe <code>df</code> that has columns <code>['custom_ID', 'prediction']</code>:</p>\n<pre><code>df_new = df.groupby(\"customer_ID\").mean().reset_index()\n</code></pre>\n<p>That will group all the rows that have the same ID and take their mean value into new dataframe, which now can be submitted.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1862747": "Can anyone share how have you dealt with duplicates in the test dataset?",
    "1862764": "There are no duplicates in the test dataset. All users have multiple time points corresponding to their credit card statements from different points in time - see `S_2` variable.",
    "1862785": "Yes, but the submission should have each customer once. I wonder how to aggregate mean for predictions??",
    "1862832": "However you are aggregating train data to prepare them for training, you should do the same with test data. That is part of the preprocessing procedure and it should be the same for both datasets.\n\nIf you haven't aggregated either train or test data, meaning you are treating every single sample in train individually, you will likely overfit during training. But if you really want to do it that way, and assuming that after prediction you got a dataframe `df` that has columns `['custom_ID', 'prediction']`:\n\n    df_new = df.groupby(\"customer_ID\").mean().reset_index()\n\nThat will group all the rows that have the same ID and take their mean value into new dataframe, which now can be submitted."
  },
  "source": "meta"
}