{
  "id": 376191,
  "title": "How to fix \"Duplicate bug\" in current baseline notebooks",
  "url": "/competitions/otto-recommender-system/discussion/376191",
  "author_name": "",
  "post_date": "2023-01-05T05:55:55.002543600Z",
  "votes": 8,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I found that current baseline notebooks include \"duplicate item\" in output CSV.<br>\nAt first, I checked the output submission of the following notebook which generates the highest score (0.577).</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/utm529fg/otto-tuning-candidate-rerank-model-lb-0-577\" target=\"_blank\">https://www.kaggle.com/code/utm529fg/otto-tuning-candidate-rerank-model-lb-0-577</a></li>\n</ul>\n<p>I checked the duplicate line as follows:</p>\n<pre><code>df = pd.read_csv(f'{PATH}/submission.csv')\ndf['labels'] = df['labels'].str.split(' ')\ndf.head()\n</code></pre>\n<pre><code>    session_type    labels\n0    12899779_clicks [59625, 1253524, 737445, 438191, 731692, 17907...\n1    12899780_clicks [1142000, 736515, 973453, 582732, 889686, 4871...\n2    12899781_clicks [918667, 199008, 194067, 57315, 141736, 146057...\n3    12899782_clicks [834354, 740494, 987399, 889671, 779477, 12740...\n4    12899783_clicks [1817895, 607638, 1754419, 1216820, 1729553, 3...\n</code></pre>\n<pre><code>ex_df = df.explode(\"labels\").reset_index(drop=True)\nex_df[ex_df.duplicated()]\n</code></pre>\n<pre><code>    session_type    labels\n839    12899820_clicks 986164\n1918    12899874_clicks 108125\n2298    12899893_clicks 108125\n4276    12899992_clicks 1460571\n6156    12900086_clicks 1460571\n...    ... ...\n97088356    14410590_carts  554660\n97535437    14432944_carts  1460571\n98239751    14468160_carts  1116095\n98983436    14505344_carts  1022566\n98995955    14505970_carts  986164\n40944 rows × 2 columns\n</code></pre>\n<p>Obviously, these output file include the duplicate line which means that the same aid is shown in the same session.</p>\n<p>I think this problem caused from the following line which is included in the most of the baseline notebooks:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/utm529fg/otto-tuning-candidate-rerank-model-lb-0-577\" target=\"_blank\">https://www.kaggle.com/code/utm529fg/otto-tuning-candidate-rerank-model-lb-0-577</a></li>\n<li><a href=\"https://www.kaggle.com/code/tuongkhang/otto-pipeline2-lb-0-576\" target=\"_blank\">https://www.kaggle.com/code/tuongkhang/otto-pipeline2-lb-0-576</a></li>\n<li><a href=\"https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575</a></li>\n</ul>\n<pre><code>    # USE TOP20 TEST CLICKS\n    return result + list(top_clicks)[:20-len(result)]\n</code></pre>\n<p>I changed this line to</p>\n<pre><code>    #return result + list(top_clicks)[:20-len(result)]\n    # FIXED BY tetsuro731, remove duplicate\n    set_result = set(result)\n    return result + [i for i in top_clicks if i not in set_result][:20 - len(result)]\n</code></pre>\n<p>My revised notebook: <a href=\"https://www.kaggle.com/code/tetsuro731/duplicate-fix-otto-tuning-pipeline2-lb-0-577\" target=\"_blank\">https://www.kaggle.com/code/tetsuro731/duplicate-fix-otto-tuning-pipeline2-lb-0-577</a></p>\n<p>Strangely, the score has not improved even though I fixed the duplicate bugs.<br>\nI'm not sure the reason.<br>\nHowever, I'm wondering the number of duplicate lines are so small that the score is not changed by this update.</p>",
  "messages": [
    {
      "id": "2086873",
      "postDate": "01/05/2023 05:55:55",
      "content": "<p>I found that current baseline notebooks include \"duplicate item\" in output CSV.<br>\nAt first, I checked the output submission of the following notebook which generates the highest score (0.577).</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/utm529fg/otto-tuning-candidate-rerank-model-lb-0-577\" target=\"_blank\">https://www.kaggle.com/code/utm529fg/otto-tuning-candidate-rerank-model-lb-0-577</a></li>\n</ul>\n<p>I checked the duplicate line as follows:</p>\n<pre><code>df = pd.read_csv(f'{PATH}/submission.csv')\ndf['labels'] = df['labels'].str.split(' ')\ndf.head()\n</code></pre>\n<pre><code>    session_type    labels\n0    12899779_clicks [59625, 1253524, 737445, 438191, 731692, 17907...\n1    12899780_clicks [1142000, 736515, 973453, 582732, 889686, 4871...\n2    12899781_clicks [918667, 199008, 194067, 57315, 141736, 146057...\n3    12899782_clicks [834354, 740494, 987399, 889671, 779477, 12740...\n4    12899783_clicks [1817895, 607638, 1754419, 1216820, 1729553, 3...\n</code></pre>\n<pre><code>ex_df = df.explode(\"labels\").reset_index(drop=True)\nex_df[ex_df.duplicated()]\n</code></pre>\n<pre><code>    session_type    labels\n839    12899820_clicks 986164\n1918    12899874_clicks 108125\n2298    12899893_clicks 108125\n4276    12899992_clicks 1460571\n6156    12900086_clicks 1460571\n...    ... ...\n97088356    14410590_carts  554660\n97535437    14432944_carts  1460571\n98239751    14468160_carts  1116095\n98983436    14505344_carts  1022566\n98995955    14505970_carts  986164\n40944 rows × 2 columns\n</code></pre>\n<p>Obviously, these output file include the duplicate line which means that the same aid is shown in the same session.</p>\n<p>I think this problem caused from the following line which is included in the most of the baseline notebooks:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/utm529fg/otto-tuning-candidate-rerank-model-lb-0-577\" target=\"_blank\">https://www.kaggle.com/code/utm529fg/otto-tuning-candidate-rerank-model-lb-0-577</a></li>\n<li><a href=\"https://www.kaggle.com/code/tuongkhang/otto-pipeline2-lb-0-576\" target=\"_blank\">https://www.kaggle.com/code/tuongkhang/otto-pipeline2-lb-0-576</a></li>\n<li><a href=\"https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575</a></li>\n</ul>\n<pre><code>    # USE TOP20 TEST CLICKS\n    return result + list(top_clicks)[:20-len(result)]\n</code></pre>\n<p>I changed this line to</p>\n<pre><code>    #return result + list(top_clicks)[:20-len(result)]\n    # FIXED BY tetsuro731, remove duplicate\n    set_result = set(result)\n    return result + [i for i in top_clicks if i not in set_result][:20 - len(result)]\n</code></pre>\n<p>My revised notebook: <a href=\"https://www.kaggle.com/code/tetsuro731/duplicate-fix-otto-tuning-pipeline2-lb-0-577\" target=\"_blank\">https://www.kaggle.com/code/tetsuro731/duplicate-fix-otto-tuning-pipeline2-lb-0-577</a></p>\n<p>Strangely, the score has not improved even though I fixed the duplicate bugs.<br>\nI'm not sure the reason.<br>\nHowever, I'm wondering the number of duplicate lines are so small that the score is not changed by this update.</p>",
      "rawMarkdown": "I found that current baseline notebooks include \"duplicate item\" in output CSV.\nAt first, I checked the output submission of the following notebook which generates the highest score (0.577).\n\n- https://www.kaggle.com/code/utm529fg/otto-tuning-candidate-rerank-model-lb-0-577\n\nI checked the duplicate line as follows:\n\n```\ndf = pd.read_csv(f'{PATH}/submission.csv')\ndf['labels'] = df['labels'].str.split(' ')\ndf.head()\n```\n\n```\n\tsession_type\tlabels\n0\t12899779_clicks\t[59625, 1253524, 737445, 438191, 731692, 17907...\n1\t12899780_clicks\t[1142000, 736515, 973453, 582732, 889686, 4871...\n2\t12899781_clicks\t[918667, 199008, 194067, 57315, 141736, 146057...\n3\t12899782_clicks\t[834354, 740494, 987399, 889671, 779477, 12740...\n4\t12899783_clicks\t[1817895, 607638, 1754419, 1216820, 1729553, 3...\n```\n\n```\nex_df = df.explode(\"labels\").reset_index(drop=True)\nex_df[ex_df.duplicated()]\n```\n\n```\n\tsession_type\tlabels\n839\t12899820_clicks\t986164\n1918\t12899874_clicks\t108125\n2298\t12899893_clicks\t108125\n4276\t12899992_clicks\t1460571\n6156\t12900086_clicks\t1460571\n...\t...\t...\n97088356\t14410590_carts\t554660\n97535437\t14432944_carts\t1460571\n98239751\t14468160_carts\t1116095\n98983436\t14505344_carts\t1022566\n98995955\t14505970_carts\t986164\n40944 rows × 2 columns\n```\n\nObviously, these output file include the duplicate line which means that the same aid is shown in the same session.\n\nI think this problem caused from the following line which is included in the most of the baseline notebooks:\n- https://www.kaggle.com/code/utm529fg/otto-tuning-candidate-rerank-model-lb-0-577\n- https://www.kaggle.com/code/tuongkhang/otto-pipeline2-lb-0-576\n- https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575\n\n```\n    # USE TOP20 TEST CLICKS\n    return result + list(top_clicks)[:20-len(result)]\n```\n\nI changed this line to\n```\n    #return result + list(top_clicks)[:20-len(result)]\n    # FIXED BY tetsuro731, remove duplicate\n    set_result = set(result)\n    return result + [i for i in top_clicks if i not in set_result][:20 - len(result)]\n```\nMy revised notebook: https://www.kaggle.com/code/tetsuro731/duplicate-fix-otto-tuning-pipeline2-lb-0-577\n\nStrangely, the score has not improved even though I fixed the duplicate bugs.\nI'm not sure the reason.\nHowever, I'm wondering the number of duplicate lines are so small that the score is not changed by this update.",
      "votes": null
    },
    {
      "id": "2086885",
      "postDate": "01/05/2023 06:14:19",
      "content": "<p>Submitting or removing duplicate aids for the same session wouldn't push the metric. Normann answered <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363973#2029245\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "Submitting or removing duplicate aids for the same session wouldn't push the metric. Normann answered [here](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363973#2029245)",
      "votes": null
    },
    {
      "id": "2086925",
      "postDate": "01/05/2023 06:51:36",
      "content": "<p>Thank you for your nice information!<br>\nI understand removing duplicate <code>aid</code> itself doesn't improve the metric.<br>\nHowever, we can chose another item if we remove duplicated aid, which can improve the metrics.<br>\nFor example:</p>\n<p>Case1: No duplicate</p>\n<ul>\n<li>session1_click, aid1, aid2, aid3, …, aid19, aid20<br>\n-&gt; We can select 20 aids</li>\n</ul>\n<p>Case2: Duplicate (aid1 is shown 2 times)</p>\n<ul>\n<li>session1_click, aid1, aid2, aid3, …, aid19, aid1<br>\n-&gt; We can select 19 aids</li>\n</ul>\n<p>Therefore, I think we should remove duplicated aid and we should chose distinct 20 aids for each session to get a good score.</p>",
      "rawMarkdown": "Thank you for your nice information!\nI understand removing duplicate `aid` itself doesn't improve the metric.\nHowever, we can chose another item if we remove duplicated aid, which can improve the metrics.\nFor example:\n\nCase1: No duplicate\n- session1_click, aid1, aid2, aid3, ..., aid19, aid20\n-> We can select 20 aids\n\nCase2: Duplicate (aid1 is shown 2 times)\n- session1_click, aid1, aid2, aid3, ..., aid19, aid1\n-> We can select 19 aids\n\nTherefore, I think we should remove duplicated aid and we should chose distinct 20 aids for each session to get a good score.",
      "votes": null
    },
    {
      "id": "2086939",
      "postDate": "01/05/2023 07:08:29",
      "content": "<p>I think of 2 possibilities, one is that the aids you added are not good enough to increase the metric, and the other is that the session &lt; 20 aids/or is already covered by previous predicted items.</p>",
      "rawMarkdown": "I think of 2 possibilities, one is that the aids you added are not good enough to increase the metric, and the other is that the session < 20 aids/or is already covered by previous predicted items.",
      "votes": null
    },
    {
      "id": "2086948",
      "postDate": "01/05/2023 07:24:38",
      "content": "<p>Yes.<br>\nI totally agree with your ideas.</p>\n<p>In addition, I'm also wondering we should be careful if we choose the candidate items by similar procedures which are used for the next ranking phase.<br>\nI think most of the people just copied these notebook without understanding the fact that the results include the duplicated aids.</p>\n<p>Anyway, my correction have not improve the score so this issue does not cause the critical problem for now.</p>",
      "rawMarkdown": "Yes.\nI totally agree with your ideas.\n\nIn addition, I'm also wondering we should be careful if we choose the candidate items by similar procedures which are used for the next ranking phase.\nI think most of the people just copied these notebook without understanding the fact that the results include the duplicated aids.\n\nAnyway, my correction have not improve the score so this issue does not cause the critical problem for now.",
      "votes": null
    },
    {
      "id": "2087226",
      "postDate": "01/05/2023 12:52:44",
      "content": "<p>You are correct that removing duplicate aid entries and choosing distinct products for each session can improve the model's performance. This is because having duplicate aid entries in the same session can reduce the number of unique products that the model can choose from, which may limit its ability to accurately predict the correct products for that session.</p>\n<p>By removing duplicate aid entries and ensuring that each session has a set of distinct products, you can give the model a better chance to make accurate predictions. This can potentially improve the overall score, depending on the specific metric being used to evaluate the model's performance.</p>\n<p>It is always a good idea to ensure that the data used to train and evaluate a model is as clean and accurate as possible. Removing duplicate entries and addressing other issues with the data can help improve the model's performance and increase its reliability.</p>",
      "rawMarkdown": "You are correct that removing duplicate aid entries and choosing distinct products for each session can improve the model's performance. This is because having duplicate aid entries in the same session can reduce the number of unique products that the model can choose from, which may limit its ability to accurately predict the correct products for that session.\n\nBy removing duplicate aid entries and ensuring that each session has a set of distinct products, you can give the model a better chance to make accurate predictions. This can potentially improve the overall score, depending on the specific metric being used to evaluate the model's performance.\n\nIt is always a good idea to ensure that the data used to train and evaluate a model is as clean and accurate as possible. Removing duplicate entries and addressing other issues with the data can help improve the model's performance and increase its reliability.",
      "votes": null
    },
    {
      "id": "2087409",
      "postDate": "01/05/2023 15:21:03",
      "content": "<p>Thank you for your opinion.<br>\nThat is just what I wanted to say!<br>\nI think I have a deeper understanding of this idea now.</p>",
      "rawMarkdown": "Thank you for your opinion.\nThat is just what I wanted to say!\nI think I have a deeper understanding of this idea now.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2086885,
      "author_name": "bibanh",
      "author_url": "",
      "post_date": "01/05/2023 06:14:19",
      "content": "<p>Submitting or removing duplicate aids for the same session wouldn't push the metric. Normann answered <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363973#2029245\" target=\"_blank\">here</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2086925,
          "author_name": "tetsuro731",
          "author_url": "",
          "post_date": "01/05/2023 06:51:36",
          "content": "<p>Thank you for your nice information!<br>\nI understand removing duplicate <code>aid</code> itself doesn't improve the metric.<br>\nHowever, we can chose another item if we remove duplicated aid, which can improve the metrics.<br>\nFor example:</p>\n<p>Case1: No duplicate</p>\n<ul>\n<li>session1_click, aid1, aid2, aid3, …, aid19, aid20<br>\n-&gt; We can select 20 aids</li>\n</ul>\n<p>Case2: Duplicate (aid1 is shown 2 times)</p>\n<ul>\n<li>session1_click, aid1, aid2, aid3, …, aid19, aid1<br>\n-&gt; We can select 19 aids</li>\n</ul>\n<p>Therefore, I think we should remove duplicated aid and we should chose distinct 20 aids for each session to get a good score.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2086939,
              "author_name": "bibanh",
              "author_url": "",
              "post_date": "01/05/2023 07:08:29",
              "content": "<p>I think of 2 possibilities, one is that the aids you added are not good enough to increase the metric, and the other is that the session &lt; 20 aids/or is already covered by previous predicted items.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2086948,
                  "author_name": "tetsuro731",
                  "author_url": "",
                  "post_date": "01/05/2023 07:24:38",
                  "content": "<p>Yes.<br>\nI totally agree with your ideas.</p>\n<p>In addition, I'm also wondering we should be careful if we choose the candidate items by similar procedures which are used for the next ranking phase.<br>\nI think most of the people just copied these notebook without understanding the fact that the results include the duplicated aids.</p>\n<p>Anyway, my correction have not improve the score so this issue does not cause the critical problem for now.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            },
            {
              "id": 2087226,
              "author_name": "aviralmishra1998",
              "author_url": "",
              "post_date": "01/05/2023 12:52:44",
              "content": "<p>You are correct that removing duplicate aid entries and choosing distinct products for each session can improve the model's performance. This is because having duplicate aid entries in the same session can reduce the number of unique products that the model can choose from, which may limit its ability to accurately predict the correct products for that session.</p>\n<p>By removing duplicate aid entries and ensuring that each session has a set of distinct products, you can give the model a better chance to make accurate predictions. This can potentially improve the overall score, depending on the specific metric being used to evaluate the model's performance.</p>\n<p>It is always a good idea to ensure that the data used to train and evaluate a model is as clean and accurate as possible. Removing duplicate entries and addressing other issues with the data can help improve the model's performance and increase its reliability.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2087409,
                  "author_name": "tetsuro731",
                  "author_url": "",
                  "post_date": "01/05/2023 15:21:03",
                  "content": "<p>Thank you for your opinion.<br>\nThat is just what I wanted to say!<br>\nI think I have a deeper understanding of this idea now.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2086873": "I found that current baseline notebooks include \"duplicate item\" in output CSV.\nAt first, I checked the output submission of the following notebook which generates the highest score (0.577).\n\n- https://www.kaggle.com/code/utm529fg/otto-tuning-candidate-rerank-model-lb-0-577\n\nI checked the duplicate line as follows:\n\n```\ndf = pd.read_csv(f'{PATH}/submission.csv')\ndf['labels'] = df['labels'].str.split(' ')\ndf.head()\n```\n\n```\n\tsession_type\tlabels\n0\t12899779_clicks\t[59625, 1253524, 737445, 438191, 731692, 17907...\n1\t12899780_clicks\t[1142000, 736515, 973453, 582732, 889686, 4871...\n2\t12899781_clicks\t[918667, 199008, 194067, 57315, 141736, 146057...\n3\t12899782_clicks\t[834354, 740494, 987399, 889671, 779477, 12740...\n4\t12899783_clicks\t[1817895, 607638, 1754419, 1216820, 1729553, 3...\n```\n\n```\nex_df = df.explode(\"labels\").reset_index(drop=True)\nex_df[ex_df.duplicated()]\n```\n\n```\n\tsession_type\tlabels\n839\t12899820_clicks\t986164\n1918\t12899874_clicks\t108125\n2298\t12899893_clicks\t108125\n4276\t12899992_clicks\t1460571\n6156\t12900086_clicks\t1460571\n...\t...\t...\n97088356\t14410590_carts\t554660\n97535437\t14432944_carts\t1460571\n98239751\t14468160_carts\t1116095\n98983436\t14505344_carts\t1022566\n98995955\t14505970_carts\t986164\n40944 rows × 2 columns\n```\n\nObviously, these output file include the duplicate line which means that the same aid is shown in the same session.\n\nI think this problem caused from the following line which is included in the most of the baseline notebooks:\n- https://www.kaggle.com/code/utm529fg/otto-tuning-candidate-rerank-model-lb-0-577\n- https://www.kaggle.com/code/tuongkhang/otto-pipeline2-lb-0-576\n- https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575\n\n```\n    # USE TOP20 TEST CLICKS\n    return result + list(top_clicks)[:20-len(result)]\n```\n\nI changed this line to\n```\n    #return result + list(top_clicks)[:20-len(result)]\n    # FIXED BY tetsuro731, remove duplicate\n    set_result = set(result)\n    return result + [i for i in top_clicks if i not in set_result][:20 - len(result)]\n```\nMy revised notebook: https://www.kaggle.com/code/tetsuro731/duplicate-fix-otto-tuning-pipeline2-lb-0-577\n\nStrangely, the score has not improved even though I fixed the duplicate bugs.\nI'm not sure the reason.\nHowever, I'm wondering the number of duplicate lines are so small that the score is not changed by this update.",
    "2086885": "Submitting or removing duplicate aids for the same session wouldn't push the metric. Normann answered [here](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363973#2029245)",
    "2086925": "Thank you for your nice information!\nI understand removing duplicate `aid` itself doesn't improve the metric.\nHowever, we can chose another item if we remove duplicated aid, which can improve the metrics.\nFor example:\n\nCase1: No duplicate\n- session1_click, aid1, aid2, aid3, ..., aid19, aid20\n-> We can select 20 aids\n\nCase2: Duplicate (aid1 is shown 2 times)\n- session1_click, aid1, aid2, aid3, ..., aid19, aid1\n-> We can select 19 aids\n\nTherefore, I think we should remove duplicated aid and we should chose distinct 20 aids for each session to get a good score.",
    "2086939": "I think of 2 possibilities, one is that the aids you added are not good enough to increase the metric, and the other is that the session < 20 aids/or is already covered by previous predicted items.",
    "2086948": "Yes.\nI totally agree with your ideas.\n\nIn addition, I'm also wondering we should be careful if we choose the candidate items by similar procedures which are used for the next ranking phase.\nI think most of the people just copied these notebook without understanding the fact that the results include the duplicated aids.\n\nAnyway, my correction have not improve the score so this issue does not cause the critical problem for now.",
    "2087226": "You are correct that removing duplicate aid entries and choosing distinct products for each session can improve the model's performance. This is because having duplicate aid entries in the same session can reduce the number of unique products that the model can choose from, which may limit its ability to accurately predict the correct products for that session.\n\nBy removing duplicate aid entries and ensuring that each session has a set of distinct products, you can give the model a better chance to make accurate predictions. This can potentially improve the overall score, depending on the specific metric being used to evaluate the model's performance.\n\nIt is always a good idea to ensure that the data used to train and evaluate a model is as clean and accurate as possible. Removing duplicate entries and addressing other issues with the data can help improve the model's performance and increase its reliability.",
    "2087409": "Thank you for your opinion.\nThat is just what I wanted to say!\nI think I have a deeper understanding of this idea now."
  },
  "source": "meta"
}