{
  "id": 341125,
  "title": "I can't submit correctly because of chunking !!",
  "url": "/competitions/amex-default-prediction/discussion/341125",
  "author_name": "Bahaa al-deen Kattan",
  "post_date": "2022-08-01T10:51:07.534000",
  "votes": 1,
  "comment_count": 7,
  "views": 0,
  "content": "<p>when I process test data, I process them by chunks. And this will make more than one record for a certain customer_ID if depending on the chunks they exist in. any advice to avoid making duplicates??</p>\n<p>the submission length =924621<br>\nmy submission length =924672 (around 50 more records)</p>",
  "messages": [
    {
      "id": 1880441,
      "postDate": "2022-08-01T18:07:24.830Z",
      "content": "<p>I present a solution in my notebook code cell #17 <a href=\"https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\" target=\"_blank\">here</a>. You must first read the entire column of <code>customer_ID</code>. Then you must choose row splits that do not break one group of similar customer_IDs into two groups.</p>",
      "rawMarkdown": "I present a solution in my notebook code cell #17 [here][1]. You must first read the entire column of `customer_ID`. Then you must choose row splits that do not break one group of similar customer_IDs into two groups.\n\n[1]: https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793",
      "votes": 1
    },
    {
      "id": 1879903,
      "postDate": "2022-08-01T10:51:07.533Z",
      "content": "<p>when I process test data, I process them by chunks. And this will make more than one record for a certain customer_ID if depending on the chunks they exist in. any advice to avoid making duplicates??</p>\n<p>the submission length =924621<br>\nmy submission length =924672 (around 50 more records)</p>",
      "rawMarkdown": "when I process test data, I process them by chunks. And this will make more than one record for a certain customer_ID if depending on the chunks they exist in. any advice to avoid making duplicates??\n\nthe submission length =924621\nmy submission length =924672 (around 50 more records)",
      "votes": 1
    },
    {
      "id": 1880046,
      "postDate": "2022-08-01T12:37:06.690Z",
      "content": "<p>did you do 50 chunks? I guess you did not handle it well - most likely one customer ended in a chunk and the same customer started in a following chunk. You need to handle that.</p>",
      "rawMarkdown": "did you do 50 chunks? I guess you did not handle it well - most likely one customer ended in a chunk and the same customer started in a following chunk. You need to handle that.",
      "votes": 2,
      "replies": [
        {
          "id": 1880150,
          "postDate": "2022-08-01T13:50:03.580Z",
          "content": "<p>I iterate every 200,000 records ?</p>",
          "rawMarkdown": "I iterate every 200,000 records ?",
          "votes": -2
        },
        {
          "id": 1880275,
          "postDate": "2022-08-01T15:23:16.177Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1882206,
      "postDate": "2022-08-03T06:09:27.260Z",
      "content": "<p>If you load your submission csv file into a pandas dataframe,  you can figure out which customer_ID is showing up more than once: </p>\n<p><code>df_submission.customer_ID.unique()[df_submission.customer_ID.value_counts() &gt; 1])</code></p>",
      "rawMarkdown": "If you load your submission csv file into a pandas dataframe,  you can figure out which customer_ID is showing up more than once: \n\n`df_submission.customer_ID.unique()[df_submission.customer_ID.value_counts() > 1])`"
    },
    {
      "id": 1880799,
      "postDate": "2022-08-02T04:21:10.483Z",
      "content": "<p>Use Dask and also Sorting that will give only unique customer_ID record…!!! Happy kaggling</p>",
      "rawMarkdown": "Use Dask and also Sorting that will give only unique customer_ID record...!!! Happy kaggling"
    },
    {
      "id": 1880158,
      "postDate": "2022-08-01T13:56:21.463Z",
      "content": "<p>Same as <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a>, you need to make sure you don't split a customers record between files. This is a useful function for that <code>df.groupby(\"customer_ID\").indices</code></p>",
      "rawMarkdown": "Same as @raddar, you need to make sure you don't split a customers record between files. This is a useful function for that ```df.groupby(\"customer_ID\").indices```"
    }
  ],
  "comments": [
    {
      "id": 1880441,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2022-08-01T18:07:24.830000",
      "content": "<p>I present a solution in my notebook code cell #17 <a href=\"https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\" target=\"_blank\">here</a>. You must first read the entire column of <code>customer_ID</code>. Then you must choose row splits that do not break one group of similar customer_IDs into two groups.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1880046,
      "author_name": "raddar",
      "author_url": "",
      "post_date": "2022-08-01T12:37:06.690000",
      "content": "<p>did you do 50 chunks? I guess you did not handle it well - most likely one customer ended in a chunk and the same customer started in a following chunk. You need to handle that.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1880150,
          "author_name": "Bahaa al-deen Kattan",
          "author_url": "",
          "post_date": "2022-08-01T13:50:03.580000",
          "content": "<p>I iterate every 200,000 records ?</p>",
          "votes": -2,
          "replies": []
        },
        {
          "id": 1880275,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-08-01T15:23:16.177000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1882206,
      "author_name": "ambitiousDonut",
      "author_url": "",
      "post_date": "2022-08-03T06:09:27.260000",
      "content": "<p>If you load your submission csv file into a pandas dataframe,  you can figure out which customer_ID is showing up more than once: </p>\n<p><code>df_submission.customer_ID.unique()[df_submission.customer_ID.value_counts() &gt; 1])</code></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1880799,
      "author_name": "Urban Monk 09",
      "author_url": "",
      "post_date": "2022-08-02T04:21:10.483000",
      "content": "<p>Use Dask and also Sorting that will give only unique customer_ID record…!!! Happy kaggling</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1880158,
      "author_name": "Jake",
      "author_url": "",
      "post_date": "2022-08-01T13:56:21.463000",
      "content": "<p>Same as <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a>, you need to make sure you don't split a customers record between files. This is a useful function for that <code>df.groupby(\"customer_ID\").indices</code></p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1880441": "I present a solution in my notebook code cell #17 [here][1]. You must first read the entire column of `customer_ID`. Then you must choose row splits that do not break one group of similar customer_IDs into two groups.\n\n[1]: https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793",
    "1879903": "when I process test data, I process them by chunks. And this will make more than one record for a certain customer_ID if depending on the chunks they exist in. any advice to avoid making duplicates??\n\nthe submission length =924621\nmy submission length =924672 (around 50 more records)",
    "1880046": "did you do 50 chunks? I guess you did not handle it well - most likely one customer ended in a chunk and the same customer started in a following chunk. You need to handle that.",
    "1882206": "If you load your submission csv file into a pandas dataframe,  you can figure out which customer_ID is showing up more than once: \n\n`df_submission.customer_ID.unique()[df_submission.customer_ID.value_counts() > 1])`",
    "1880799": "Use Dask and also Sorting that will give only unique customer_ID record...!!! Happy kaggling",
    "1880158": "Same as @raddar, you need to make sure you don't split a customers record between files. This is a useful function for that ```df.groupby(\"customer_ID\").indices```"
  }
}