{
  "id": 384562,
  "title": "Another Coincidence: Abnormally High Similarity of Submissions",
  "url": "/competitions/otto-recommender-system/discussion/384562",
  "author_name": "sirius",
  "post_date": "2023-02-08T12:44:51.095000",
  "votes": 35,
  "comment_count": 8,
  "views": 0,
  "content": "<p>We, Team Bruce and Team Makotu, have tried our best to make a 100% re-production of their submissions using all the submission files we can access, but failed. Here is just another <code>coincidence</code> about them (more coincidences can be seen in <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/381321\" target=\"_blank\">thread1</a>, <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/381318\" target=\"_blank\">thread2</a>, <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/381360\" target=\"_blank\">thread3</a>): </p>\n<ul>\n<li><strong>Average similarity</strong> of submissions excepting RunningZ’s and Poteman’s: <strong>56.1%</strong><ul>\n<li><a href=\"https://www.kaggle.com/code/sirius81/otto-average-sim-of-irrelative-subs\" target=\"_blank\">Notebook</a></li>\n<li>I collected 6 submission files from <a href=\"https://www.kaggle.com/bruceqdu\" target=\"_blank\">@bruceqdu</a>, <a href=\"https://www.kaggle.com/mhyodo\" target=\"_blank\">@mhyodo</a>, <a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a>, <a href=\"https://www.kaggle.com/alvor\" target=\"_blank\">@alvor</a>, <a href=\"https://www.kaggle.com/h4211819\" target=\"_blank\">@h4211819</a> and myself, which all scored 0.598~0.600. Then I calculated the similarity of C(6,2) pairs and averaged them.</li>\n<li>In addition, the max value of them is 59.6%</li></ul></li>\n<li>The similarity of <strong>RunningZ-Makotu(Poteman’s teammate)</strong>: <strong>73.1%</strong><ul>\n<li><a href=\"https://www.kaggle.com/code/sirius81/otto-sub-verification\" target=\"_blank\">Notebook</a></li>\n<li>It is abnormally high, even higher than the similarity of RunningZ-Bruce(RunningZ’s teammate), which is 71.3%.</li></ul></li>\n<li>The similarity of <strong>Poteman-Bruce(RunningZ’s teammate)</strong>: <strong>74.6%</strong><ul>\n<li><a href=\"https://www.kaggle.com/code/sirius81/otto-sub-verification-p\" target=\"_blank\">Notebook</a></li>\n<li>It’s also abnormally high.</li></ul></li>\n</ul>\n<p>Here is the heatmap of all the similarities(<a href=\"https://www.kaggle.com/code/sirius81/otto-sub-verification-heatmap?scriptVersionId=118578717\" target=\"_blank\">Notebook</a>).<br>\n<img src=\"https://postimg.cc/cKtgdmq1\" alt=\"heamap.png\"><br>\n<img src=\"https://i.postimg.cc/xdtKyxtm/heatmap.png\" alt=\"heamap.png\"><br>\nP.S. As <a href=\"https://www.kaggle.com/shimacos\" target=\"_blank\">@shimacos</a> suggested, all the datasets will be published soon. If anyone is interested in the challenge of recreating submission, go ahead please! <br>\nP.P.S. Our final submissions contain nothing from these guys. We were working separately just since the next day of team forming. Call for a precise justice.</p>\n<p>---------------------------------------------update 2023-02-09--------------------------------<br>\n<strong>Datasets for anomaly detection</strong></p>\n<ol>\n<li>Bruce’s sub and RunningZ’s sub: <a href=\"https://www.kaggle.com/datasets/ilyalos/otto-submissions-runningz\" target=\"_blank\">https://www.kaggle.com/datasets/ilyalos/otto-submissions-runningz</a></li>\n<li>Sirius’ sub: <a href=\"https://www.kaggle.com/datasets/sirius81/otto-sirius-sub-before-runningz-scoring-601?select=sirius_sub_meta.csv\" target=\"_blank\">https://www.kaggle.com/datasets/sirius81/otto-sirius-sub-before-runningz-scoring-601?select=sirius_sub_meta.csv</a></li>\n<li>Makotu’s sub: <a href=\"https://www.kaggle.com/datasets/sirius81/otto-verification?select=makotu_doubtful_sub\" target=\"_blank\">https://www.kaggle.com/datasets/sirius81/otto-verification?select=makotu_doubtful_sub</a></li>\n<li>Alvor’s sub and Shimacos’ sub(only for runningZ’s 0.603): <a href=\"https://www.kaggle.com/datasets/sirius81/otto-sirius-sub-before-runningz-scoring-601?select=alvor_shimacos_sub\" target=\"_blank\">https://www.kaggle.com/datasets/sirius81/otto-sirius-sub-before-runningz-scoring-601?select=alvor_shimacos_sub</a></li>\n<li>Poteman's sub: <a href=\"https://www.kaggle.com/datasets/sirius81/otto-poteman-sub\" target=\"_blank\">https://www.kaggle.com/datasets/sirius81/otto-poteman-sub</a></li>\n<li>Some irrelative subs (for a baseline of similarity): <a href=\"https://www.kaggle.com/datasets/sirius81/otto-irrelative-subs\" target=\"_blank\">https://www.kaggle.com/datasets/sirius81/otto-irrelative-subs</a></li>\n</ol>",
  "messages": [
    {
      "id": 2135082,
      "postDate": "2023-02-08T12:44:51.097Z",
      "content": "<p>We, Team Bruce and Team Makotu, have tried our best to make a 100% re-production of their submissions using all the submission files we can access, but failed. Here is just another <code>coincidence</code> about them (more coincidences can be seen in <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/381321\" target=\"_blank\">thread1</a>, <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/381318\" target=\"_blank\">thread2</a>, <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/381360\" target=\"_blank\">thread3</a>): </p>\n<ul>\n<li><strong>Average similarity</strong> of submissions excepting RunningZ’s and Poteman’s: <strong>56.1%</strong><ul>\n<li><a href=\"https://www.kaggle.com/code/sirius81/otto-average-sim-of-irrelative-subs\" target=\"_blank\">Notebook</a></li>\n<li>I collected 6 submission files from <a href=\"https://www.kaggle.com/bruceqdu\" target=\"_blank\">@bruceqdu</a>, <a href=\"https://www.kaggle.com/mhyodo\" target=\"_blank\">@mhyodo</a>, <a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a>, <a href=\"https://www.kaggle.com/alvor\" target=\"_blank\">@alvor</a>, <a href=\"https://www.kaggle.com/h4211819\" target=\"_blank\">@h4211819</a> and myself, which all scored 0.598~0.600. Then I calculated the similarity of C(6,2) pairs and averaged them.</li>\n<li>In addition, the max value of them is 59.6%</li></ul></li>\n<li>The similarity of <strong>RunningZ-Makotu(Poteman’s teammate)</strong>: <strong>73.1%</strong><ul>\n<li><a href=\"https://www.kaggle.com/code/sirius81/otto-sub-verification\" target=\"_blank\">Notebook</a></li>\n<li>It is abnormally high, even higher than the similarity of RunningZ-Bruce(RunningZ’s teammate), which is 71.3%.</li></ul></li>\n<li>The similarity of <strong>Poteman-Bruce(RunningZ’s teammate)</strong>: <strong>74.6%</strong><ul>\n<li><a href=\"https://www.kaggle.com/code/sirius81/otto-sub-verification-p\" target=\"_blank\">Notebook</a></li>\n<li>It’s also abnormally high.</li></ul></li>\n</ul>\n<p>Here is the heatmap of all the similarities(<a href=\"https://www.kaggle.com/code/sirius81/otto-sub-verification-heatmap?scriptVersionId=118578717\" target=\"_blank\">Notebook</a>).<br>\n<img src=\"https://postimg.cc/cKtgdmq1\" alt=\"heamap.png\"><br>\n<img src=\"https://i.postimg.cc/xdtKyxtm/heatmap.png\" alt=\"heamap.png\"><br>\nP.S. As <a href=\"https://www.kaggle.com/shimacos\" target=\"_blank\">@shimacos</a> suggested, all the datasets will be published soon. If anyone is interested in the challenge of recreating submission, go ahead please! <br>\nP.P.S. Our final submissions contain nothing from these guys. We were working separately just since the next day of team forming. Call for a precise justice.</p>\n<p>---------------------------------------------update 2023-02-09--------------------------------<br>\n<strong>Datasets for anomaly detection</strong></p>\n<ol>\n<li>Bruce’s sub and RunningZ’s sub: <a href=\"https://www.kaggle.com/datasets/ilyalos/otto-submissions-runningz\" target=\"_blank\">https://www.kaggle.com/datasets/ilyalos/otto-submissions-runningz</a></li>\n<li>Sirius’ sub: <a href=\"https://www.kaggle.com/datasets/sirius81/otto-sirius-sub-before-runningz-scoring-601?select=sirius_sub_meta.csv\" target=\"_blank\">https://www.kaggle.com/datasets/sirius81/otto-sirius-sub-before-runningz-scoring-601?select=sirius_sub_meta.csv</a></li>\n<li>Makotu’s sub: <a href=\"https://www.kaggle.com/datasets/sirius81/otto-verification?select=makotu_doubtful_sub\" target=\"_blank\">https://www.kaggle.com/datasets/sirius81/otto-verification?select=makotu_doubtful_sub</a></li>\n<li>Alvor’s sub and Shimacos’ sub(only for runningZ’s 0.603): <a href=\"https://www.kaggle.com/datasets/sirius81/otto-sirius-sub-before-runningz-scoring-601?select=alvor_shimacos_sub\" target=\"_blank\">https://www.kaggle.com/datasets/sirius81/otto-sirius-sub-before-runningz-scoring-601?select=alvor_shimacos_sub</a></li>\n<li>Poteman's sub: <a href=\"https://www.kaggle.com/datasets/sirius81/otto-poteman-sub\" target=\"_blank\">https://www.kaggle.com/datasets/sirius81/otto-poteman-sub</a></li>\n<li>Some irrelative subs (for a baseline of similarity): <a href=\"https://www.kaggle.com/datasets/sirius81/otto-irrelative-subs\" target=\"_blank\">https://www.kaggle.com/datasets/sirius81/otto-irrelative-subs</a></li>\n</ol>",
      "rawMarkdown": "We, Team Bruce and Team Makotu, have tried our best to make a 100% re-production of their submissions using all the submission files we can access, but failed. Here is just another `coincidence` about them (more coincidences can be seen in [thread1](https://www.kaggle.com/competitions/otto-recommender-system/discussion/381321), [thread2](https://www.kaggle.com/competitions/otto-recommender-system/discussion/381318), [thread3](https://www.kaggle.com/competitions/otto-recommender-system/discussion/381360)): \n- **Average similarity** of submissions excepting RunningZ’s and Poteman’s: **56.1%**\n    - [Notebook](https://www.kaggle.com/code/sirius81/otto-average-sim-of-irrelative-subs)\n    - I collected 6 submission files from @bruceqdu, @mhyodo, @carnozhao, @alvor, @h4211819 and myself, which all scored 0.598~0.600. Then I calculated the similarity of C(6,2) pairs and averaged them.\n    - In addition, the max value of them is 59.6%\n- The similarity of **RunningZ-Makotu(Poteman’s teammate)**: **73.1%**\n    - [Notebook](https://www.kaggle.com/code/sirius81/otto-sub-verification)\n    - It is abnormally high, even higher than the similarity of RunningZ-Bruce(RunningZ’s teammate), which is 71.3%.\n- The similarity of **Poteman-Bruce(RunningZ’s teammate)**: **74.6%**\n    - [Notebook](https://www.kaggle.com/code/sirius81/otto-sub-verification-p)\n    - It’s also abnormally high.\n\nHere is the heatmap of all the similarities([Notebook](https://www.kaggle.com/code/sirius81/otto-sub-verification-heatmap?scriptVersionId=118578717)).\n![heamap.png](https://postimg.cc/cKtgdmq1)\n![heamap.png](https://i.postimg.cc/xdtKyxtm/heatmap.png)\nP.S. As @shimacos suggested, all the datasets will be published soon. If anyone is interested in the challenge of recreating submission, go ahead please! \nP.P.S. Our final submissions contain nothing from these guys. We were working separately just since the next day of team forming. Call for a precise justice.\n\n---------------------------------------------update 2023-02-09--------------------------------\n**Datasets for anomaly detection**\n1. Bruce’s sub and RunningZ’s sub: https://www.kaggle.com/datasets/ilyalos/otto-submissions-runningz\n2. Sirius’ sub: https://www.kaggle.com/datasets/sirius81/otto-sirius-sub-before-runningz-scoring-601?select=sirius_sub_meta.csv\n3. Makotu’s sub: https://www.kaggle.com/datasets/sirius81/otto-verification?select=makotu_doubtful_sub\n4. Alvor’s sub and Shimacos’ sub(only for runningZ’s 0.603): https://www.kaggle.com/datasets/sirius81/otto-sirius-sub-before-runningz-scoring-601?select=alvor_shimacos_sub\n5. Poteman's sub: https://www.kaggle.com/datasets/sirius81/otto-poteman-sub\n6. Some irrelative subs (for a baseline of similarity): https://www.kaggle.com/datasets/sirius81/otto-irrelative-subs\n",
      "votes": 35
    },
    {
      "id": 2135154,
      "postDate": "2023-02-08T13:37:34.033Z",
      "content": "<p>Some more data points using my team's submissions, obtained AFTER merging, i.e. we have shared ideas.<br>\nI compare the following submissions :</p>\n<ul>\n<li>One from Chris  - LB 0.599</li>\n<li>One from Benedikt - LB 0.599</li>\n<li>One from myself - LB 0.601</li>\n<li>A weighted blend of the 3 above, using weights 0.25, 0.25, 0.5 (theo) - LB 0.602</li>\n</ul>\n<p><strong>Similarities of our individual pipelines :</strong></p>\n<ul>\n<li>Chris / Theo : 0.625</li>\n<li>Chris / Benedikt : 0.584</li>\n<li>Theo / Benedikt : 0.625</li>\n</ul>\n<p><strong>Similarities with the ensemble :</strong></p>\n<ul>\n<li>Chris : 0.702</li>\n<li>Benedikt :  0.703</li>\n<li>Theo : 0.810</li>\n</ul>\n<p><strong>Observations :</strong></p>\n<ul>\n<li>Even with sharing features and candidates, our pipelines achieve ~0.6 similarity.</li>\n<li>0.7 similarities are reached when comparing submissions with a blend that contains it at a 0.25 weight.</li>\n<li>0.8  similarities are reached when comparing submissions with a blend that contains it at a 0.5 weight.</li>\n</ul>\n<p><strong>What we can infer regarind Poteman-Bruce and RunningZ-Makotu submissions:</strong></p>\n<ul>\n<li>Similarities that are higher than 0.7 indicate that the person used a very similar submission.</li>\n<li>We already know that Poteman and RunningZ love blending, so we can safely assume :<ul>\n<li>Poteman used one of Bruce's sub in his blend with a weight of approx. 1/3</li>\n<li>RunningZ used one of Makotu's sub  in his blend with a weight of approx. 1/3</li></ul></li>\n<li>Poteman most likely sent Makotu's sub to RunningZ and RunningZ gave Bruce's to Poteman.</li>\n</ul>\n<p>Poke <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> </p>\n<p>There is more than enough evidence that the two people in question have broken Kaggle rules.<br>\nDon't let this go unpunished.</p>",
      "rawMarkdown": "Some more data points using my team's submissions, obtained AFTER merging, i.e. we have shared ideas.\nI compare the following submissions :\n- One from Chris  - LB 0.599\n- One from Benedikt - LB 0.599\n- One from myself - LB 0.601\n- A weighted blend of the 3 above, using weights 0.25, 0.25, 0.5 (theo) - LB 0.602\n\n**Similarities of our individual pipelines :**\n- Chris / Theo : 0.625\n- Chris / Benedikt : 0.584\n- Theo / Benedikt : 0.625\n\n**Similarities with the ensemble :**\n- Chris : 0.702\n- Benedikt :  0.703\n- Theo : 0.810\n\n**Observations :**\n- Even with sharing features and candidates, our pipelines achieve ~0.6 similarity.\n- 0.7 similarities are reached when comparing submissions with a blend that contains it at a 0.25 weight.\n- 0.8  similarities are reached when comparing submissions with a blend that contains it at a 0.5 weight.\n\n**What we can infer regarind Poteman-Bruce and RunningZ-Makotu submissions:**\n- Similarities that are higher than 0.7 indicate that the person used a very similar submission.\n- We already know that Poteman and RunningZ love blending, so we can safely assume :\n  - Poteman used one of Bruce's sub in his blend with a weight of approx. 1/3\n  - RunningZ used one of Makotu's sub  in his blend with a weight of approx. 1/3\n- Poteman most likely sent Makotu's sub to RunningZ and RunningZ gave Bruce's to Poteman.\n\n\nPoke @inversion \n\nThere is more than enough evidence that the two people in question have broken Kaggle rules.\nDon't let this go unpunished.",
      "votes": 17,
      "replies": [
        {
          "id": 2135196,
          "postDate": "2023-02-08T14:05:32.717Z",
          "content": "<p>I am so surprised kaggle hasn't even clarified their decision </p>",
          "rawMarkdown": "I am so surprised kaggle hasn't even clarified their decision ",
          "votes": 6,
          "replies": [
            {
              "id": 2135439,
              "postDate": "2023-02-08T16:28:07.450Z",
              "content": "<p>Unfortunately they don't need to.<br>\nAs it is written in Terms of Use Kaggle team can eliminate or not eliminate anybody only based on their own judgement or feeling.<br>\nNo evidence is necessary from their side and no evidence can help an unlawfully eliminated participant to get the position back if Kaggle team decides.   </p>",
              "rawMarkdown": "Unfortunately they don't need to.\nAs it is written in Terms of Use Kaggle team can eliminate or not eliminate anybody only based on their own judgement or feeling.\nNo evidence is necessary from their side and no evidence can help an unlawfully eliminated participant to get the position back if Kaggle team decides.   "
            }
          ]
        }
      ]
    },
    {
      "id": 2135107,
      "postDate": "2023-02-08T13:01:50.690Z",
      "content": "<p>Maybe you can post a heatmap of all possible combinations of these submissions, which will be more easier to highlight the abnormality than just numbers.</p>",
      "rawMarkdown": "Maybe you can post a heatmap of all possible combinations of these submissions, which will be more easier to highlight the abnormality than just numbers.",
      "votes": 3,
      "replies": [
        {
          "id": 2135203,
          "postDate": "2023-02-08T14:07:56.907Z",
          "content": "<p>Thanks. Heatmap added</p>",
          "rawMarkdown": "Thanks. Heatmap added"
        }
      ]
    },
    {
      "id": 2135285,
      "postDate": "2023-02-08T14:45:12.623Z",
      "content": "<p>What's the public score of the runningz and poteman's sub ? </p>\n<p>If it's a bit lower, then it Would be interesting to see the cases where there is a mismatch. </p>\n<p>Have they been copied from a public high scoring notebook ? </p>\n<p>Has it been random replacement of aids, with a calculation to bring the score down by a fix amount. If some aids are added by random, they most probably would not increase the score at all, and x% of such aids would bring the score down by y%</p>",
      "rawMarkdown": "What's the public score of the runningz and poteman's sub ? \n\nIf it's a bit lower, then it Would be interesting to see the cases where there is a mismatch. \n\nHave they been copied from a public high scoring notebook ? \n\nHas it been random replacement of aids, with a calculation to bring the score down by a fix amount. If some aids are added by random, they most probably would not increase the score at all, and x% of such aids would bring the score down by y%",
      "replies": [
        {
          "id": 2136033,
          "postDate": "2023-02-09T03:58:54.970Z",
          "content": "<p>RunningZ's sub and Poteman's sub, used in the above analysis, scored 0.603 and 0.596 respectively. It's hard to tell whether they blended some public notebook with small weight.</p>",
          "rawMarkdown": "RunningZ's sub and Poteman's sub, used in the above analysis, scored 0.603 and 0.596 respectively. It's hard to tell whether they blended some public notebook with small weight."
        }
      ]
    },
    {
      "id": 2135097,
      "postDate": "2023-02-08T12:59:01.850Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2135154,
      "author_name": "Theo Viel",
      "author_url": "",
      "post_date": "2023-02-08T13:37:34.033000",
      "content": "<p>Some more data points using my team's submissions, obtained AFTER merging, i.e. we have shared ideas.<br>\nI compare the following submissions :</p>\n<ul>\n<li>One from Chris  - LB 0.599</li>\n<li>One from Benedikt - LB 0.599</li>\n<li>One from myself - LB 0.601</li>\n<li>A weighted blend of the 3 above, using weights 0.25, 0.25, 0.5 (theo) - LB 0.602</li>\n</ul>\n<p><strong>Similarities of our individual pipelines :</strong></p>\n<ul>\n<li>Chris / Theo : 0.625</li>\n<li>Chris / Benedikt : 0.584</li>\n<li>Theo / Benedikt : 0.625</li>\n</ul>\n<p><strong>Similarities with the ensemble :</strong></p>\n<ul>\n<li>Chris : 0.702</li>\n<li>Benedikt :  0.703</li>\n<li>Theo : 0.810</li>\n</ul>\n<p><strong>Observations :</strong></p>\n<ul>\n<li>Even with sharing features and candidates, our pipelines achieve ~0.6 similarity.</li>\n<li>0.7 similarities are reached when comparing submissions with a blend that contains it at a 0.25 weight.</li>\n<li>0.8  similarities are reached when comparing submissions with a blend that contains it at a 0.5 weight.</li>\n</ul>\n<p><strong>What we can infer regarind Poteman-Bruce and RunningZ-Makotu submissions:</strong></p>\n<ul>\n<li>Similarities that are higher than 0.7 indicate that the person used a very similar submission.</li>\n<li>We already know that Poteman and RunningZ love blending, so we can safely assume :<ul>\n<li>Poteman used one of Bruce's sub in his blend with a weight of approx. 1/3</li>\n<li>RunningZ used one of Makotu's sub  in his blend with a weight of approx. 1/3</li></ul></li>\n<li>Poteman most likely sent Makotu's sub to RunningZ and RunningZ gave Bruce's to Poteman.</li>\n</ul>\n<p>Poke <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> </p>\n<p>There is more than enough evidence that the two people in question have broken Kaggle rules.<br>\nDon't let this go unpunished.</p>",
      "votes": 17,
      "replies": [
        {
          "id": 2135196,
          "author_name": "NikhilMishra",
          "author_url": "",
          "post_date": "2023-02-08T14:05:32.717000",
          "content": "<p>I am so surprised kaggle hasn't even clarified their decision </p>",
          "votes": 6,
          "replies": [
            {
              "id": 2135439,
              "author_name": "Allie K.",
              "author_url": "",
              "post_date": "2023-02-08T16:28:07.450000",
              "content": "<p>Unfortunately they don't need to.<br>\nAs it is written in Terms of Use Kaggle team can eliminate or not eliminate anybody only based on their own judgement or feeling.<br>\nNo evidence is necessary from their side and no evidence can help an unlawfully eliminated participant to get the position back if Kaggle team decides.   </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2135107,
      "author_name": "Carno Zhao",
      "author_url": "",
      "post_date": "2023-02-08T13:01:50.690000",
      "content": "<p>Maybe you can post a heatmap of all possible combinations of these submissions, which will be more easier to highlight the abnormality than just numbers.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2135203,
          "author_name": "sirius",
          "author_url": "",
          "post_date": "2023-02-08T14:07:56.907000",
          "content": "<p>Thanks. Heatmap added</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2135285,
      "author_name": "NikhilMishra",
      "author_url": "",
      "post_date": "2023-02-08T14:45:12.623000",
      "content": "<p>What's the public score of the runningz and poteman's sub ? </p>\n<p>If it's a bit lower, then it Would be interesting to see the cases where there is a mismatch. </p>\n<p>Have they been copied from a public high scoring notebook ? </p>\n<p>Has it been random replacement of aids, with a calculation to bring the score down by a fix amount. If some aids are added by random, they most probably would not increase the score at all, and x% of such aids would bring the score down by y%</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2136033,
          "author_name": "sirius",
          "author_url": "",
          "post_date": "2023-02-09T03:58:54.970000",
          "content": "<p>RunningZ's sub and Poteman's sub, used in the above analysis, scored 0.603 and 0.596 respectively. It's hard to tell whether they blended some public notebook with small weight.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2135097,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-02-08T12:59:01.850000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2135082": "We, Team Bruce and Team Makotu, have tried our best to make a 100% re-production of their submissions using all the submission files we can access, but failed. Here is just another `coincidence` about them (more coincidences can be seen in [thread1](https://www.kaggle.com/competitions/otto-recommender-system/discussion/381321), [thread2](https://www.kaggle.com/competitions/otto-recommender-system/discussion/381318), [thread3](https://www.kaggle.com/competitions/otto-recommender-system/discussion/381360)): \n- **Average similarity** of submissions excepting RunningZ’s and Poteman’s: **56.1%**\n    - [Notebook](https://www.kaggle.com/code/sirius81/otto-average-sim-of-irrelative-subs)\n    - I collected 6 submission files from @bruceqdu, @mhyodo, @carnozhao, @alvor, @h4211819 and myself, which all scored 0.598~0.600. Then I calculated the similarity of C(6,2) pairs and averaged them.\n    - In addition, the max value of them is 59.6%\n- The similarity of **RunningZ-Makotu(Poteman’s teammate)**: **73.1%**\n    - [Notebook](https://www.kaggle.com/code/sirius81/otto-sub-verification)\n    - It is abnormally high, even higher than the similarity of RunningZ-Bruce(RunningZ’s teammate), which is 71.3%.\n- The similarity of **Poteman-Bruce(RunningZ’s teammate)**: **74.6%**\n    - [Notebook](https://www.kaggle.com/code/sirius81/otto-sub-verification-p)\n    - It’s also abnormally high.\n\nHere is the heatmap of all the similarities([Notebook](https://www.kaggle.com/code/sirius81/otto-sub-verification-heatmap?scriptVersionId=118578717)).\n![heamap.png](https://postimg.cc/cKtgdmq1)\n![heamap.png](https://i.postimg.cc/xdtKyxtm/heatmap.png)\nP.S. As @shimacos suggested, all the datasets will be published soon. If anyone is interested in the challenge of recreating submission, go ahead please! \nP.P.S. Our final submissions contain nothing from these guys. We were working separately just since the next day of team forming. Call for a precise justice.\n\n---------------------------------------------update 2023-02-09--------------------------------\n**Datasets for anomaly detection**\n1. Bruce’s sub and RunningZ’s sub: https://www.kaggle.com/datasets/ilyalos/otto-submissions-runningz\n2. Sirius’ sub: https://www.kaggle.com/datasets/sirius81/otto-sirius-sub-before-runningz-scoring-601?select=sirius_sub_meta.csv\n3. Makotu’s sub: https://www.kaggle.com/datasets/sirius81/otto-verification?select=makotu_doubtful_sub\n4. Alvor’s sub and Shimacos’ sub(only for runningZ’s 0.603): https://www.kaggle.com/datasets/sirius81/otto-sirius-sub-before-runningz-scoring-601?select=alvor_shimacos_sub\n5. Poteman's sub: https://www.kaggle.com/datasets/sirius81/otto-poteman-sub\n6. Some irrelative subs (for a baseline of similarity): https://www.kaggle.com/datasets/sirius81/otto-irrelative-subs\n",
    "2135154": "Some more data points using my team's submissions, obtained AFTER merging, i.e. we have shared ideas.\nI compare the following submissions :\n- One from Chris  - LB 0.599\n- One from Benedikt - LB 0.599\n- One from myself - LB 0.601\n- A weighted blend of the 3 above, using weights 0.25, 0.25, 0.5 (theo) - LB 0.602\n\n**Similarities of our individual pipelines :**\n- Chris / Theo : 0.625\n- Chris / Benedikt : 0.584\n- Theo / Benedikt : 0.625\n\n**Similarities with the ensemble :**\n- Chris : 0.702\n- Benedikt :  0.703\n- Theo : 0.810\n\n**Observations :**\n- Even with sharing features and candidates, our pipelines achieve ~0.6 similarity.\n- 0.7 similarities are reached when comparing submissions with a blend that contains it at a 0.25 weight.\n- 0.8  similarities are reached when comparing submissions with a blend that contains it at a 0.5 weight.\n\n**What we can infer regarind Poteman-Bruce and RunningZ-Makotu submissions:**\n- Similarities that are higher than 0.7 indicate that the person used a very similar submission.\n- We already know that Poteman and RunningZ love blending, so we can safely assume :\n  - Poteman used one of Bruce's sub in his blend with a weight of approx. 1/3\n  - RunningZ used one of Makotu's sub  in his blend with a weight of approx. 1/3\n- Poteman most likely sent Makotu's sub to RunningZ and RunningZ gave Bruce's to Poteman.\n\n\nPoke @inversion \n\nThere is more than enough evidence that the two people in question have broken Kaggle rules.\nDon't let this go unpunished.",
    "2135107": "Maybe you can post a heatmap of all possible combinations of these submissions, which will be more easier to highlight the abnormality than just numbers.",
    "2135285": "What's the public score of the runningz and poteman's sub ? \n\nIf it's a bit lower, then it Would be interesting to see the cases where there is a mismatch. \n\nHave they been copied from a public high scoring notebook ? \n\nHas it been random replacement of aids, with a calculation to bring the score down by a fix amount. If some aids are added by random, they most probably would not increase the score at all, and x% of such aids would bring the score down by y%",
    "2135097": ""
  }
}