{
  "id": 335698,
  "title": "An Evil Thought Experiment",
  "url": "/competitions/amex-default-prediction/discussion/335698",
  "author_name": "",
  "post_date": "2022-07-07T12:33:04.728975500Z",
  "votes": 22,
  "comment_count": 22,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>Imagine I am a bad actor in this competition. As a thought experiment, I will do the following:</p>\n<ol>\n<li>Figure out the datasets in test for public and private sections.<br>\nwe know the boundary of private and public datasets based on this <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327926\" target=\"_blank\">discussion</a>. </li>\n<li>Train a good model with high public score.</li>\n<li>Distort the predictions in the private dataset.</li>\n<li>Submit the code and predictions online in the code section of this competition.</li>\n</ol>\n<p>What will happen next is that some of the Kagglers will start noticing the high public score of my notebook and download the predictions to use it as an ensemble source. While their public scores get better and they see improvement in their LB ranking, their private predictions get worse. They will not realize it until the final ranking.</p>\n<p>This could shake-up the LB after the deadline and could waste a lot of efforts.<br>\nI was wondering if this happened before in the previous competitions and what could be done to detect these type of manipulations? I understand that this is very pessimistic, but anything that can go wrong will go wrong!</p>\n<p>Sincerely,<br>\nMo</p>",
  "messages": [
    {
      "id": "1846888",
      "postDate": "07/07/2022 12:33:04",
      "content": "<p>Hi everyone,</p>\n<p>Imagine I am a bad actor in this competition. As a thought experiment, I will do the following:</p>\n<ol>\n<li>Figure out the datasets in test for public and private sections.<br>\nwe know the boundary of private and public datasets based on this <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327926\" target=\"_blank\">discussion</a>. </li>\n<li>Train a good model with high public score.</li>\n<li>Distort the predictions in the private dataset.</li>\n<li>Submit the code and predictions online in the code section of this competition.</li>\n</ol>\n<p>What will happen next is that some of the Kagglers will start noticing the high public score of my notebook and download the predictions to use it as an ensemble source. While their public scores get better and they see improvement in their LB ranking, their private predictions get worse. They will not realize it until the final ranking.</p>\n<p>This could shake-up the LB after the deadline and could waste a lot of efforts.<br>\nI was wondering if this happened before in the previous competitions and what could be done to detect these type of manipulations? I understand that this is very pessimistic, but anything that can go wrong will go wrong!</p>\n<p>Sincerely,<br>\nMo</p>",
      "rawMarkdown": "Hi everyone,\n\nImagine I am a bad actor in this competition. As a thought experiment, I will do the following:\n1. Figure out the datasets in test for public and private sections.\nwe know the boundary of private and public datasets based on this [discussion](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327926). \n2. Train a good model with high public score.\n3. Distort the predictions in the private dataset.\n4. Submit the code and predictions online in the code section of this competition.\n\nWhat will happen next is that some of the Kagglers will start noticing the high public score of my notebook and download the predictions to use it as an ensemble source. While their public scores get better and they see improvement in their LB ranking, their private predictions get worse. They will not realize it until the final ranking.\n\nThis could shake-up the LB after the deadline and could waste a lot of efforts.\nI was wondering if this happened before in the previous competitions and what could be done to detect these type of manipulations? I understand that this is very pessimistic, but anything that can go wrong will go wrong!\n\nSincerely,\nMo",
      "votes": null
    },
    {
      "id": "1846913",
      "postDate": "07/07/2022 13:06:28",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/mohammadrahmati\" target=\"_blank\">@mohammadrahmati</a> </p>\n<p>Like I mentioned <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/332286#1827673\" target=\"_blank\">before</a>;  stand-alone public <code>submission.csv</code> files  are like sausages; people like eating them, but if you knew how they were made, maybe you would think twice.</p>\n<p>All the best,<br>\ncarl </p>",
      "rawMarkdown": "Dear @mohammadrahmati \n\nLike I mentioned [before](https://www.kaggle.com/competitions/amex-default-prediction/discussion/332286#1827673);  stand-alone public `submission.csv` files  are like sausages; people like eating them, but if you knew how they were made, maybe you would think twice.\n\nAll the best,\ncarl",
      "votes": null
    },
    {
      "id": "1846925",
      "postDate": "07/07/2022 13:20:29",
      "content": "<p>Excellent metaphor! 😄</p>",
      "rawMarkdown": "Excellent metaphor! 😄",
      "votes": null
    },
    {
      "id": "1847059",
      "postDate": "07/07/2022 15:37:56",
      "content": "<p>It was <a href=\"https://www.kaggle.com/competitions/ieee-fraud-detection/discussion/111108\" target=\"_blank\">done before</a></p>",
      "rawMarkdown": "It was [done before](https://www.kaggle.com/competitions/ieee-fraud-detection/discussion/111108)",
      "votes": null
    },
    {
      "id": "1847073",
      "postDate": "07/07/2022 15:55:02",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a> for sharing the link. </p>",
      "rawMarkdown": "Thanks @thedevastator for sharing the link.",
      "votes": null
    },
    {
      "id": "1847374",
      "postDate": "07/07/2022 21:48:55",
      "content": "<p>As a general rule of thumb you can check if the notebook was run (versus the dataset being added manually). <br>\nPs: this is not the case of the current top notebook. </p>",
      "rawMarkdown": "As a general rule of thumb you can check if the notebook was run (versus the dataset being added manually). \nPs: this is not the case of the current top notebook.",
      "votes": null
    },
    {
      "id": "1847741",
      "postDate": "07/08/2022 05:49:12",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/lucasmorin\" target=\"_blank\">@lucasmorin</a>!</p>",
      "rawMarkdown": "Thanks @lucasmorin!",
      "votes": null
    },
    {
      "id": "1847853",
      "postDate": "07/08/2022 07:51:34",
      "content": "<p>Public Kernels with high score can be quite tempting, but I always try to reproduce the result, implement it in my pipeline, then use it.(This way we can avoid any overfitting)</p>",
      "rawMarkdown": "Public Kernels with high score can be quite tempting, but I always try to reproduce the result, implement it in my pipeline, then use it.(This way we can avoid any overfitting)",
      "votes": null
    },
    {
      "id": "1848064",
      "postDate": "07/08/2022 10:19:59",
      "content": "<p>Bad actor, or penalizing people who put no effort in?</p>",
      "rawMarkdown": "Bad actor, or penalizing people who put no effort in?",
      "votes": null
    },
    {
      "id": "1848086",
      "postDate": "07/08/2022 10:36:56",
      "content": "<p>I agree with you but it is not fair this way. Some people may use the output for ensemble purposes which doesn't mean they are not putting enough effort.</p>",
      "rawMarkdown": "I agree with you but it is not fair this way. Some people may use the output for ensemble purposes which doesn't mean they are not putting enough effort.",
      "votes": null
    },
    {
      "id": "1848592",
      "postDate": "07/08/2022 18:32:14",
      "content": "<p>Kagglers can find a bad <code>submission.csv</code> file by checking the correlations between the <code>submission.csv</code> public and their own public. And then checking the correlations between <code>submission.csv</code> private and their own private. Bad <code>submission.csv</code> files will have significantly different correlations. This happened in a previous competition. Someone published a bad <code>submission.csv</code> file and Kagglers figured it out.</p>",
      "rawMarkdown": "Kagglers can find a bad `submission.csv` file by checking the correlations between the `submission.csv` public and their own public. And then checking the correlations between `submission.csv` private and their own private. Bad `submission.csv` files will have significantly different correlations. This happened in a previous competition. Someone published a bad `submission.csv` file and Kagglers figured it out.",
      "votes": null
    },
    {
      "id": "1848682",
      "postDate": "07/08/2022 20:25:24",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> !</p>",
      "rawMarkdown": "Thanks @cdeotte !",
      "votes": null
    },
    {
      "id": "1857078",
      "postDate": "07/15/2022 20:17:04",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/raj401\" target=\"_blank\">@raj401</a>!</p>",
      "rawMarkdown": "Thanks @raj401!",
      "votes": null
    },
    {
      "id": "1857098",
      "postDate": "07/15/2022 20:34:19",
      "content": "<p>While everything has been covered by others already, just my 2 cents, do look at the public kernel, try to reproduce the result , before using it as ensemble with your model. Also try to find (as suggested already), the correlation with your predictions/models, before using it.</p>\n<p>If major shakeup happens anyway, In hindsight, at least you will know what you used exactly :D</p>",
      "rawMarkdown": "While everything has been covered by others already, just my 2 cents, do look at the public kernel, try to reproduce the result , before using it as ensemble with your model. Also try to find (as suggested already), the correlation with your predictions/models, before using it.\n\nIf major shakeup happens anyway, In hindsight, at least you will know what you used exactly :D",
      "votes": null
    },
    {
      "id": "1857162",
      "postDate": "07/15/2022 22:59:37",
      "content": "<p>thanks <a href=\"https://www.kaggle.com/nitishraj\" target=\"_blank\">@nitishraj</a> !</p>",
      "rawMarkdown": "thanks @nitishraj !",
      "votes": null
    },
    {
      "id": "1899806",
      "postDate": "08/15/2022 14:11:06",
      "content": "<p>I think someone has tried to do this.</p>",
      "rawMarkdown": "I think someone has tried to do this.",
      "votes": null
    },
    {
      "id": "1899851",
      "postDate": "08/15/2022 14:38:24",
      "content": "<p>Yes, someone just did it today by sharing a single submission.csv with LB 0.8 ! As suggested by <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/335698#1848592\" target=\"_blank\">Chris Deotte</a>, I calculated the Spearman correlations between several predictions, just for fun. :-)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1133866%2Fd613d799cc988d57804c036eaa4e5744%2FUntitled.png?generation=1660573731639005&amp;alt=media\" alt=\"\"></p>\n<p>pub01, 02, 03 are good public submissions but pub04 is definitely evil ! In my case, even GRU gets ~0.97 correlation with LGBM.</p>",
      "rawMarkdown": "Yes, someone just did it today by sharing a single submission.csv with LB 0.8 ! As suggested by [Chris Deotte](https://www.kaggle.com/competitions/amex-default-prediction/discussion/335698#1848592), I calculated the Spearman correlations between several predictions, just for fun. :-)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1133866%2Fd613d799cc988d57804c036eaa4e5744%2FUntitled.png?generation=1660573731639005&alt=media)\n\npub01, 02, 03 are good public submissions but pub04 is definitely evil ! In my case, even GRU gets ~0.97 correlation with LGBM.",
      "votes": null
    },
    {
      "id": "1899972",
      "postDate": "08/15/2022 16:12:59",
      "content": "<p>A low correlation between LB 0.8 sub and your LGBM sub isn't necessarily bad. (It could be the result of using ranks or logits instead of probabilities for predictions). Below is how to search if someone has poisoned a submission's private test predictions:</p>\n<ul>\n<li>First identify which rows of test dataframe are public test and which are private test.</li>\n<li>Compute the correlation of LB 0.8 public test rows with your LGBM public test rows</li>\n<li>Compute the correlation of LB 0.8 private test rows with your LGBM private test rows.</li>\n<li>Compare the two values from steps 2 and 3 above. </li>\n<li>If these two numbers are different than we have a poisoned submission.</li>\n</ul>",
      "rawMarkdown": "A low correlation between LB 0.8 sub and your LGBM sub isn't necessarily bad. (It could be the result of using ranks or logits instead of probabilities for predictions). Below is how to search if someone has poisoned a submission's private test predictions:\n* First identify which rows of test dataframe are public test and which are private test.\n* Compute the correlation of LB 0.8 public test rows with your LGBM public test rows\n* Compute the correlation of LB 0.8 private test rows with your LGBM private test rows.\n* Compare the two values from steps 2 and 3 above. \n* If these two numbers are different than we have a poisoned submission.",
      "votes": null
    },
    {
      "id": "1899981",
      "postDate": "08/15/2022 16:26:16",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> you have very good points. BUT due to competition metric sensitivity to False Positive predictions it's enough to swipe 10 items from lowest rank to highest to completely ruin private lb score (and it will be completely unnoticed by any correlation study or analysis). I would suggest NOT to use any predictions that are not reproduced by yourself.</p>",
      "rawMarkdown": "cdeotte you have very good points. BUT due to competition metric sensitivity to False Positive predictions it's enough to swipe 10 items from lowest rank to highest to completely ruin private lb score (and it will be completely unnoticed by any correlation study or analysis). I would suggest NOT to use any predictions that are not reproduced by yourself.",
      "votes": null
    },
    {
      "id": "1900201",
      "postDate": "08/15/2022 19:14:39",
      "content": "<p>Anything that can go wrong will go wrong!</p>",
      "rawMarkdown": "Anything that can go wrong will go wrong!",
      "votes": null
    },
    {
      "id": "1900203",
      "postDate": "08/15/2022 19:20:27",
      "content": "<p>Delightful post title. </p>",
      "rawMarkdown": "Delightful post title.",
      "votes": null
    },
    {
      "id": "1900206",
      "postDate": "08/15/2022 19:26:40",
      "content": "<p>Unfortunately I am extra evil and thought of that way to detect poison, so used two highly correlated submissions but largely different LB scores. (or did I and is this just psyops and I actually released a full 0.800 submission?) </p>\n<p>Not the first time this has happened in competitions and will not be the last. 😄</p>",
      "rawMarkdown": "Unfortunately I am extra evil and thought of that way to detect poison, so used two highly correlated submissions but largely different LB scores. (or did I and is this just psyops and I actually released a full 0.800 submission?) \n\nNot the first time this has happened in competitions and will not be the last. 😄",
      "votes": null
    },
    {
      "id": "1900359",
      "postDate": "08/16/2022 00:14:07",
      "content": "<p>Thanks for the clarification <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> ! Yes, this way makes things clearer.</p>\n<p>Since the Amex metric is based on ranking (AUROC, TPR4% etc.), I intuitively think that the Spearman correlation is sensitive to the difference in the Amex metric. The pub01, pub02 and LGBM subs are 0.799 and their difference with the LB0.8 sub is only ~0.001, which seems unlikely to give a low Spearman correlation, so there's something wrong in the private set.</p>\n<p>Here is the search according to your suggestion:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1133866%2F975dfcc8705a8171e482dd244dfa5374%2Fc1.png?generation=1660608239102793&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1133866%2F9862df20f638a5d42a5027f76df7c87a%2Fc2.png?generation=1660608745632931&amp;alt=media\" alt=\"\"></p>\n<p>Well, the correlations are pretty high in both public and private sets, but ~0.01 difference in Spearman correlation may indicate a lot of difference in the Amex metric (e.g. high LB vs. low PB). As shown above, the distributions are quite different and thus give the overall low correlation. It is obvious that the private set was manipulated. Anyway, I've had my fun. 😃</p>",
      "rawMarkdown": "Thanks for the clarification @cdeotte ! Yes, this way makes things clearer.\n\nSince the Amex metric is based on ranking (AUROC, TPR4% etc.), I intuitively think that the Spearman correlation is sensitive to the difference in the Amex metric. The pub01, pub02 and LGBM subs are 0.799 and their difference with the LB0.8 sub is only ~0.001, which seems unlikely to give a low Spearman correlation, so there's something wrong in the private set.\n\nHere is the search according to your suggestion:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1133866%2F975dfcc8705a8171e482dd244dfa5374%2Fc1.png?generation=1660608239102793&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1133866%2F9862df20f638a5d42a5027f76df7c87a%2Fc2.png?generation=1660608745632931&alt=media)\n\nWell, the correlations are pretty high in both public and private sets, but ~0.01 difference in Spearman correlation may indicate a lot of difference in the Amex metric (e.g. high LB vs. low PB). As shown above, the distributions are quite different and thus give the overall low correlation. It is obvious that the private set was manipulated. Anyway, I've had my fun. 😃",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1846913,
      "author_name": "carlmcbrideellis",
      "author_url": "",
      "post_date": "07/07/2022 13:06:28",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/mohammadrahmati\" target=\"_blank\">@mohammadrahmati</a> </p>\n<p>Like I mentioned <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/332286#1827673\" target=\"_blank\">before</a>;  stand-alone public <code>submission.csv</code> files  are like sausages; people like eating them, but if you knew how they were made, maybe you would think twice.</p>\n<p>All the best,<br>\ncarl </p>",
      "votes": null,
      "replies": [
        {
          "id": 1846925,
          "author_name": "mohammadrahmati",
          "author_url": "",
          "post_date": "07/07/2022 13:20:29",
          "content": "<p>Excellent metaphor! 😄</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1847059,
      "author_name": "thedevastator",
      "author_url": "",
      "post_date": "07/07/2022 15:37:56",
      "content": "<p>It was <a href=\"https://www.kaggle.com/competitions/ieee-fraud-detection/discussion/111108\" target=\"_blank\">done before</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1847073,
          "author_name": "mohammadrahmati",
          "author_url": "",
          "post_date": "07/07/2022 15:55:02",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a> for sharing the link. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1847374,
      "author_name": "lucasmorin",
      "author_url": "",
      "post_date": "07/07/2022 21:48:55",
      "content": "<p>As a general rule of thumb you can check if the notebook was run (versus the dataset being added manually). <br>\nPs: this is not the case of the current top notebook. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1847741,
          "author_name": "mohammadrahmati",
          "author_url": "",
          "post_date": "07/08/2022 05:49:12",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/lucasmorin\" target=\"_blank\">@lucasmorin</a>!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1847853,
      "author_name": "raj401",
      "author_url": "",
      "post_date": "07/08/2022 07:51:34",
      "content": "<p>Public Kernels with high score can be quite tempting, but I always try to reproduce the result, implement it in my pipeline, then use it.(This way we can avoid any overfitting)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1857078,
          "author_name": "mohammadrahmati",
          "author_url": "",
          "post_date": "07/15/2022 20:17:04",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/raj401\" target=\"_blank\">@raj401</a>!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1848064,
      "author_name": "julianmukaj",
      "author_url": "",
      "post_date": "07/08/2022 10:19:59",
      "content": "<p>Bad actor, or penalizing people who put no effort in?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1848086,
          "author_name": "mohammadrahmati",
          "author_url": "",
          "post_date": "07/08/2022 10:36:56",
          "content": "<p>I agree with you but it is not fair this way. Some people may use the output for ensemble purposes which doesn't mean they are not putting enough effort.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1848592,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "07/08/2022 18:32:14",
      "content": "<p>Kagglers can find a bad <code>submission.csv</code> file by checking the correlations between the <code>submission.csv</code> public and their own public. And then checking the correlations between <code>submission.csv</code> private and their own private. Bad <code>submission.csv</code> files will have significantly different correlations. This happened in a previous competition. Someone published a bad <code>submission.csv</code> file and Kagglers figured it out.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1848682,
          "author_name": "mohammadrahmati",
          "author_url": "",
          "post_date": "07/08/2022 20:25:24",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> !</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1857098,
      "author_name": "nitishraj",
      "author_url": "",
      "post_date": "07/15/2022 20:34:19",
      "content": "<p>While everything has been covered by others already, just my 2 cents, do look at the public kernel, try to reproduce the result , before using it as ensemble with your model. Also try to find (as suggested already), the correlation with your predictions/models, before using it.</p>\n<p>If major shakeup happens anyway, In hindsight, at least you will know what you used exactly :D</p>",
      "votes": null,
      "replies": [
        {
          "id": 1857162,
          "author_name": "mohammadrahmati",
          "author_url": "",
          "post_date": "07/15/2022 22:59:37",
          "content": "<p>thanks <a href=\"https://www.kaggle.com/nitishraj\" target=\"_blank\">@nitishraj</a> !</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1899806,
      "author_name": "hongbinna",
      "author_url": "",
      "post_date": "08/15/2022 14:11:06",
      "content": "<p>I think someone has tried to do this.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1899851,
      "author_name": "franklamp",
      "author_url": "",
      "post_date": "08/15/2022 14:38:24",
      "content": "<p>Yes, someone just did it today by sharing a single submission.csv with LB 0.8 ! As suggested by <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/335698#1848592\" target=\"_blank\">Chris Deotte</a>, I calculated the Spearman correlations between several predictions, just for fun. :-)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1133866%2Fd613d799cc988d57804c036eaa4e5744%2FUntitled.png?generation=1660573731639005&amp;alt=media\" alt=\"\"></p>\n<p>pub01, 02, 03 are good public submissions but pub04 is definitely evil ! In my case, even GRU gets ~0.97 correlation with LGBM.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1899972,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "08/15/2022 16:12:59",
          "content": "<p>A low correlation between LB 0.8 sub and your LGBM sub isn't necessarily bad. (It could be the result of using ranks or logits instead of probabilities for predictions). Below is how to search if someone has poisoned a submission's private test predictions:</p>\n<ul>\n<li>First identify which rows of test dataframe are public test and which are private test.</li>\n<li>Compute the correlation of LB 0.8 public test rows with your LGBM public test rows</li>\n<li>Compute the correlation of LB 0.8 private test rows with your LGBM private test rows.</li>\n<li>Compare the two values from steps 2 and 3 above. </li>\n<li>If these two numbers are different than we have a poisoned submission.</li>\n</ul>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1899981,
          "author_name": "kyakovlev",
          "author_url": "",
          "post_date": "08/15/2022 16:26:16",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> you have very good points. BUT due to competition metric sensitivity to False Positive predictions it's enough to swipe 10 items from lowest rank to highest to completely ruin private lb score (and it will be completely unnoticed by any correlation study or analysis). I would suggest NOT to use any predictions that are not reproduced by yourself.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1900201,
          "author_name": "mohammadrahmati",
          "author_url": "",
          "post_date": "08/15/2022 19:14:39",
          "content": "<p>Anything that can go wrong will go wrong!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1900206,
          "author_name": "julianmukaj",
          "author_url": "",
          "post_date": "08/15/2022 19:26:40",
          "content": "<p>Unfortunately I am extra evil and thought of that way to detect poison, so used two highly correlated submissions but largely different LB scores. (or did I and is this just psyops and I actually released a full 0.800 submission?) </p>\n<p>Not the first time this has happened in competitions and will not be the last. 😄</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1900359,
          "author_name": "franklamp",
          "author_url": "",
          "post_date": "08/16/2022 00:14:07",
          "content": "<p>Thanks for the clarification <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> ! Yes, this way makes things clearer.</p>\n<p>Since the Amex metric is based on ranking (AUROC, TPR4% etc.), I intuitively think that the Spearman correlation is sensitive to the difference in the Amex metric. The pub01, pub02 and LGBM subs are 0.799 and their difference with the LB0.8 sub is only ~0.001, which seems unlikely to give a low Spearman correlation, so there's something wrong in the private set.</p>\n<p>Here is the search according to your suggestion:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1133866%2F975dfcc8705a8171e482dd244dfa5374%2Fc1.png?generation=1660608239102793&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1133866%2F9862df20f638a5d42a5027f76df7c87a%2Fc2.png?generation=1660608745632931&amp;alt=media\" alt=\"\"></p>\n<p>Well, the correlations are pretty high in both public and private sets, but ~0.01 difference in Spearman correlation may indicate a lot of difference in the Amex metric (e.g. high LB vs. low PB). As shown above, the distributions are quite different and thus give the overall low correlation. It is obvious that the private set was manipulated. Anyway, I've had my fun. 😃</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1900203,
      "author_name": "bellapisani",
      "author_url": "",
      "post_date": "08/15/2022 19:20:27",
      "content": "<p>Delightful post title. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1846888": "Hi everyone,\n\nImagine I am a bad actor in this competition. As a thought experiment, I will do the following:\n1. Figure out the datasets in test for public and private sections.\nwe know the boundary of private and public datasets based on this [discussion](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327926). \n2. Train a good model with high public score.\n3. Distort the predictions in the private dataset.\n4. Submit the code and predictions online in the code section of this competition.\n\nWhat will happen next is that some of the Kagglers will start noticing the high public score of my notebook and download the predictions to use it as an ensemble source. While their public scores get better and they see improvement in their LB ranking, their private predictions get worse. They will not realize it until the final ranking.\n\nThis could shake-up the LB after the deadline and could waste a lot of efforts.\nI was wondering if this happened before in the previous competitions and what could be done to detect these type of manipulations? I understand that this is very pessimistic, but anything that can go wrong will go wrong!\n\nSincerely,\nMo",
    "1846913": "Dear @mohammadrahmati \n\nLike I mentioned [before](https://www.kaggle.com/competitions/amex-default-prediction/discussion/332286#1827673);  stand-alone public `submission.csv` files  are like sausages; people like eating them, but if you knew how they were made, maybe you would think twice.\n\nAll the best,\ncarl",
    "1846925": "Excellent metaphor! 😄",
    "1847059": "It was [done before](https://www.kaggle.com/competitions/ieee-fraud-detection/discussion/111108)",
    "1847073": "Thanks @thedevastator for sharing the link.",
    "1847374": "As a general rule of thumb you can check if the notebook was run (versus the dataset being added manually). \nPs: this is not the case of the current top notebook.",
    "1847741": "Thanks @lucasmorin!",
    "1847853": "Public Kernels with high score can be quite tempting, but I always try to reproduce the result, implement it in my pipeline, then use it.(This way we can avoid any overfitting)",
    "1848064": "Bad actor, or penalizing people who put no effort in?",
    "1848086": "I agree with you but it is not fair this way. Some people may use the output for ensemble purposes which doesn't mean they are not putting enough effort.",
    "1848592": "Kagglers can find a bad `submission.csv` file by checking the correlations between the `submission.csv` public and their own public. And then checking the correlations between `submission.csv` private and their own private. Bad `submission.csv` files will have significantly different correlations. This happened in a previous competition. Someone published a bad `submission.csv` file and Kagglers figured it out.",
    "1848682": "Thanks @cdeotte !",
    "1857078": "Thanks @raj401!",
    "1857098": "While everything has been covered by others already, just my 2 cents, do look at the public kernel, try to reproduce the result , before using it as ensemble with your model. Also try to find (as suggested already), the correlation with your predictions/models, before using it.\n\nIf major shakeup happens anyway, In hindsight, at least you will know what you used exactly :D",
    "1857162": "thanks @nitishraj !",
    "1899806": "I think someone has tried to do this.",
    "1899851": "Yes, someone just did it today by sharing a single submission.csv with LB 0.8 ! As suggested by [Chris Deotte](https://www.kaggle.com/competitions/amex-default-prediction/discussion/335698#1848592), I calculated the Spearman correlations between several predictions, just for fun. :-)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1133866%2Fd613d799cc988d57804c036eaa4e5744%2FUntitled.png?generation=1660573731639005&alt=media)\n\npub01, 02, 03 are good public submissions but pub04 is definitely evil ! In my case, even GRU gets ~0.97 correlation with LGBM.",
    "1899972": "A low correlation between LB 0.8 sub and your LGBM sub isn't necessarily bad. (It could be the result of using ranks or logits instead of probabilities for predictions). Below is how to search if someone has poisoned a submission's private test predictions:\n* First identify which rows of test dataframe are public test and which are private test.\n* Compute the correlation of LB 0.8 public test rows with your LGBM public test rows\n* Compute the correlation of LB 0.8 private test rows with your LGBM private test rows.\n* Compare the two values from steps 2 and 3 above. \n* If these two numbers are different than we have a poisoned submission.",
    "1899981": "cdeotte you have very good points. BUT due to competition metric sensitivity to False Positive predictions it's enough to swipe 10 items from lowest rank to highest to completely ruin private lb score (and it will be completely unnoticed by any correlation study or analysis). I would suggest NOT to use any predictions that are not reproduced by yourself.",
    "1900201": "Anything that can go wrong will go wrong!",
    "1900203": "Delightful post title.",
    "1900206": "Unfortunately I am extra evil and thought of that way to detect poison, so used two highly correlated submissions but largely different LB scores. (or did I and is this just psyops and I actually released a full 0.800 submission?) \n\nNot the first time this has happened in competitions and will not be the last. 😄",
    "1900359": "Thanks for the clarification @cdeotte ! Yes, this way makes things clearer.\n\nSince the Amex metric is based on ranking (AUROC, TPR4% etc.), I intuitively think that the Spearman correlation is sensitive to the difference in the Amex metric. The pub01, pub02 and LGBM subs are 0.799 and their difference with the LB0.8 sub is only ~0.001, which seems unlikely to give a low Spearman correlation, so there's something wrong in the private set.\n\nHere is the search according to your suggestion:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1133866%2F975dfcc8705a8171e482dd244dfa5374%2Fc1.png?generation=1660608239102793&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1133866%2F9862df20f638a5d42a5027f76df7c87a%2Fc2.png?generation=1660608745632931&alt=media)\n\nWell, the correlations are pretty high in both public and private sets, but ~0.01 difference in Spearman correlation may indicate a lot of difference in the Amex metric (e.g. high LB vs. low PB). As shown above, the distributions are quite different and thus give the overall low correlation. It is obvious that the private set was manipulated. Anyway, I've had my fun. 😃"
  },
  "source": "meta"
}