{
  "id": 80086,
  "title": "Open source all leaks (135 samples) I detected",
  "url": "/competitions/humpback-whale-identification/discussion/80086",
  "author_name": "",
  "post_date": "2019-02-10T13:30:25.165651700Z",
  "votes": 61,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Hi all, the secret comes from the playground dataset (There are some duplicate images between test set and playground dataset) and motivated by <a href=\"/martinpiotte\">@martinpiotte</a> 's work ( <a href=\"https://www.kaggle.com/martinpiotte/whale-recognition-model-with-score-0-78563\">https://www.kaggle.com/martinpiotte/whale-recognition-model-with-score-0-78563</a> ). </p>\n\n<p>Some teams may not know the previous competition and haven't downloaded the playground images. Therefore, I decided to publish all leaks I detected for you for equity.</p>\n\n<p><strong>Usages</strong></p>\n\n<ul>\n<li>Simply replace these 135 samples' top 1 prediction</li>\n<li>You can also consider it as an extra validation set, calculate the MAP@5 for them to estimate your model's performance. In my experiment, this validation set is a good reference</li>\n<li>You can estimate the public/private distributions according to them (welcome to tell us the result, in order to save the submission times, I haven't taken this kind of test)</li>\n</ul>\n\n<p>Finally, I’d like to mention that:</p>\n\n<ul>\n<li>I'm not 100% sure that all of these results are true</li>\n<li>Maybe there are leaks I don't find, but I'll not continue this work because in my opinion, many other tasks are more valuable, like bounding box regression, segmentation, and key points regression</li>\n<li>the organizer may (or already) remove these samples in private LB, thus please don't down vote me after the competition if \"leaks\" are useless</li>\n</ul>",
  "messages": [
    {
      "id": "469091",
      "postDate": "02/10/2019 13:30:25",
      "content": "<p>Hi all, the secret comes from the playground dataset (There are some duplicate images between test set and playground dataset) and motivated by <a href=\"/martinpiotte\">@martinpiotte</a> 's work ( <a href=\"https://www.kaggle.com/martinpiotte/whale-recognition-model-with-score-0-78563\">https://www.kaggle.com/martinpiotte/whale-recognition-model-with-score-0-78563</a> ). </p>\n\n<p>Some teams may not know the previous competition and haven't downloaded the playground images. Therefore, I decided to publish all leaks I detected for you for equity.</p>\n\n<p><strong>Usages</strong></p>\n\n<ul>\n<li>Simply replace these 135 samples' top 1 prediction</li>\n<li>You can also consider it as an extra validation set, calculate the MAP@5 for them to estimate your model's performance. In my experiment, this validation set is a good reference</li>\n<li>You can estimate the public/private distributions according to them (welcome to tell us the result, in order to save the submission times, I haven't taken this kind of test)</li>\n</ul>\n\n<p>Finally, I’d like to mention that:</p>\n\n<ul>\n<li>I'm not 100% sure that all of these results are true</li>\n<li>Maybe there are leaks I don't find, but I'll not continue this work because in my opinion, many other tasks are more valuable, like bounding box regression, segmentation, and key points regression</li>\n<li>the organizer may (or already) remove these samples in private LB, thus please don't down vote me after the competition if \"leaks\" are useless</li>\n</ul>",
      "rawMarkdown": "Hi all, the secret comes from the playground dataset (There are some duplicate images between test set and playground dataset) and motivated by @martinpiotte 's work ( https://www.kaggle.com/martinpiotte/whale-recognition-model-with-score-0-78563 ). \n\nSome teams may not know the previous competition and haven't downloaded the playground images. Therefore, I decided to publish all leaks I detected for you for equity.\n\n**Usages**\n\n- Simply replace these 135 samples' top 1 prediction\n- You can also consider it as an extra validation set, calculate the MAP@5 for them to estimate your model's performance. In my experiment, this validation set is a good reference\n- You can estimate the public/private distributions according to them (welcome to tell us the result, in order to save the submission times, I haven't taken this kind of test)\n\nFinally, I’d like to mention that:\n\n- I'm not 100% sure that all of these results are true\n- Maybe there are leaks I don't find, but I'll not continue this work because in my opinion, many other tasks are more valuable, like bounding box regression, segmentation, and key points regression\n- the organizer may (or already) remove these samples in private LB, thus please don't down vote me after the competition if \"leaks\" are useless",
      "votes": null
    },
    {
      "id": "469126",
      "postDate": "02/10/2019 14:48:03",
      "content": "<p>one free leak from the test data itself  ... maybe they forget to remove bottom image in the composite test image</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/469126/11241/leak.png\" alt=\"enter image description here\"></p>",
      "rawMarkdown": "one free leak from the test data itself  ... maybe they forget to remove bottom image in the composite test image\n\n\n  ![enter image description here][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/469126/11241/leak.png",
      "votes": null
    },
    {
      "id": "469188",
      "postDate": "02/10/2019 17:07:06",
      "content": "<p>Thanks for sharing. But is it legal to use this kind of leak? I mean, the leak is extracted using data from another competition.</p>\n\n<p>-Edit- Actually it is legal if playground dataset is disclosed in the external data thread and the duplicates were found automatically, right?</p>",
      "rawMarkdown": "Thanks for sharing. But is it legal to use this kind of leak? I mean, the leak is extracted using data from another competition.\n\n-Edit- Actually it is legal if playground dataset is disclosed in the external data thread and the duplicates were found automatically, right?",
      "votes": null
    },
    {
      "id": "469289",
      "postDate": "02/10/2019 22:05:08",
      "content": "<p>Hi Eduardo, don't worry. </p>\n\n<p>The organizer said \"A Playground version of this competition was hosted a few months ago. This competition contains even more images and individual whales than the last. While you can certainly use the images in the previous competition if you wish\" </p>",
      "rawMarkdown": "Hi Eduardo, don't worry. \n\nThe organizer said \"A Playground version of this competition was hosted a few months ago. This competition contains even more images and individual whales than the last. While you can certainly use the images in the previous competition if you wish\"",
      "votes": null
    },
    {
      "id": "469310",
      "postDate": "02/10/2019 23:23:59",
      "content": "<p>Thanks Venn, very admirable.</p>",
      "rawMarkdown": "Thanks Venn, very admirable.",
      "votes": null
    },
    {
      "id": "469385",
      "postDate": "02/11/2019 05:27:43",
      "content": "<p>This is so awesome</p>",
      "rawMarkdown": "This is so awesome",
      "votes": null
    },
    {
      "id": "469674",
      "postDate": "02/11/2019 16:37:06",
      "content": "<p>Thank you for posting. For the model I am currently working on: LB score 0.789. After correcting 13 samples that my model got wrong LB 0.792, and after correcting the order of 10 other predictions that my model had in a different order, LB 0.794. I just realized that the MAP@5 score depends on the order of the predictions. </p>",
      "rawMarkdown": "Thank you for posting. For the model I am currently working on: LB score 0.789. After correcting 13 samples that my model got wrong LB 0.792, and after correcting the order of 10 other predictions that my model had in a different order, LB 0.794. I just realized that the MAP@5 score depends on the order of the predictions.",
      "votes": null
    },
    {
      "id": "469904",
      "postDate": "02/12/2019 03:01:47",
      "content": "<p>Hi Venn,\nThe labeling between this competition and the playground seem different. To generate the leaks.csv, we have to know the id mappings, may I know how did you get the mappings?</p>",
      "rawMarkdown": "Hi Venn,\nThe labeling between this competition and the playground seem different. To generate the leaks.csv, we have to know the id mappings, may I know how did you get the mappings?",
      "votes": null
    },
    {
      "id": "470031",
      "postDate": "02/12/2019 09:15:06",
      "content": "<p>Hi Guanshuo, yes, it needs to find the duplicates between training set and the playground dataset first to build the mappings.</p>",
      "rawMarkdown": "Hi Guanshuo, yes, it needs to find the duplicates between training set and the playground dataset first to build the mappings.",
      "votes": null
    },
    {
      "id": "470141",
      "postDate": "02/12/2019 13:06:58",
      "content": "<p>I see</p>",
      "rawMarkdown": "I see",
      "votes": null
    },
    {
      "id": "470199",
      "postDate": "02/12/2019 15:15:57",
      "content": "<p>Thank you for your sharing. I applied this leaked file to my submission. <br>\nBut decreased my public score.   </p>\n\n<pre><code>import pandas as pd\n# my submission file\ndf = pd.read_csv(\"xxx.csv\")\nleak_df = pd.read_csv(\"../data/leaks.csv\")\n\nleak_map = {}\n\nfor idx, row in leak_df.iterrows():\n    leak_map[row[\"b_test_img\"]] = row[\"b_label\"]\n\nsubmission_list = []\nfor idx, row in df.iterrows():\n    if row[\"Image\"] in leak_map:\n        id_list = row[\"Id\"].split(\" \")\n        if id_list[0] != leak_map[row[\"Image\"]]:\n            print(id_list[0], leak_map[row[\"Image\"]])\n            print(id_list)\n        id_list[0] = leak_map[row[\"Image\"]]\n        id_string = \" \".join(id_list)\n    else:\n        id_string = row[\"Id\"]\n    submission_list.append(id_string)\ndf[\"Id\"] = submission_list\n# modified output\ndf.to_csv(\"zzz.csv\", index=False)\n</code></pre>\n\n<p>```</p>",
      "rawMarkdown": "Thank you for your sharing. I applied this leaked file to my submission.  \nBut decreased my public score.   \n  \n    import pandas as pd\n    # my submission file\n    df = pd.read_csv(\"xxx.csv\")\n    leak_df = pd.read_csv(\"../data/leaks.csv\")\n    \n    leak_map = {}\n    \n    for idx, row in leak_df.iterrows():\n        leak_map[row[\"b_test_img\"]] = row[\"b_label\"]\n    \n    submission_list = []\n    for idx, row in df.iterrows():\n        if row[\"Image\"] in leak_map:\n            id_list = row[\"Id\"].split(\" \")\n            if id_list[0] != leak_map[row[\"Image\"]]:\n                print(id_list[0], leak_map[row[\"Image\"]])\n                print(id_list)\n            id_list[0] = leak_map[row[\"Image\"]]\n            id_string = \" \".join(id_list)\n        else:\n            id_string = row[\"Id\"]\n        submission_list.append(id_string)\n    df[\"Id\"] = submission_list\n    # modified output\n    df.to_csv(\"zzz.csv\", index=False)\n```",
      "votes": null
    },
    {
      "id": "474328",
      "postDate": "02/19/2019 08:30:11",
      "content": "<p>Thanks a lot! My LB has improved +0.002.</p>",
      "rawMarkdown": "Thanks a lot! My LB has improved +0.002.",
      "votes": null
    },
    {
      "id": "474756",
      "postDate": "02/19/2019 19:04:49",
      "content": "<p>Thanks a lot for sharing!</p>",
      "rawMarkdown": "Thanks a lot for sharing!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 469126,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "02/10/2019 14:48:03",
      "content": "<p>one free leak from the test data itself  ... maybe they forget to remove bottom image in the composite test image</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/469126/11241/leak.png\" alt=\"enter image description here\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 469188,
      "author_name": "arc144",
      "author_url": "",
      "post_date": "02/10/2019 17:07:06",
      "content": "<p>Thanks for sharing. But is it legal to use this kind of leak? I mean, the leak is extracted using data from another competition.</p>\n\n<p>-Edit- Actually it is legal if playground dataset is disclosed in the external data thread and the duplicates were found automatically, right?</p>",
      "votes": null,
      "replies": [
        {
          "id": 469289,
          "author_name": "yiheng",
          "author_url": "",
          "post_date": "02/10/2019 22:05:08",
          "content": "<p>Hi Eduardo, don't worry. </p>\n\n<p>The organizer said \"A Playground version of this competition was hosted a few months ago. This competition contains even more images and individual whales than the last. While you can certainly use the images in the previous competition if you wish\" </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 469310,
      "author_name": "interneuron",
      "author_url": "",
      "post_date": "02/10/2019 23:23:59",
      "content": "<p>Thanks Venn, very admirable.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 469385,
      "author_name": "dashnabanita",
      "author_url": "",
      "post_date": "02/11/2019 05:27:43",
      "content": "<p>This is so awesome</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 469674,
      "author_name": "msmelguizo",
      "author_url": "",
      "post_date": "02/11/2019 16:37:06",
      "content": "<p>Thank you for posting. For the model I am currently working on: LB score 0.789. After correcting 13 samples that my model got wrong LB 0.792, and after correcting the order of 10 other predictions that my model had in a different order, LB 0.794. I just realized that the MAP@5 score depends on the order of the predictions. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 469904,
      "author_name": "wowfattie",
      "author_url": "",
      "post_date": "02/12/2019 03:01:47",
      "content": "<p>Hi Venn,\nThe labeling between this competition and the playground seem different. To generate the leaks.csv, we have to know the id mappings, may I know how did you get the mappings?</p>",
      "votes": null,
      "replies": [
        {
          "id": 470031,
          "author_name": "yiheng",
          "author_url": "",
          "post_date": "02/12/2019 09:15:06",
          "content": "<p>Hi Guanshuo, yes, it needs to find the duplicates between training set and the playground dataset first to build the mappings.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 470141,
          "author_name": "wowfattie",
          "author_url": "",
          "post_date": "02/12/2019 13:06:58",
          "content": "<p>I see</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 470199,
      "author_name": "tereka",
      "author_url": "",
      "post_date": "02/12/2019 15:15:57",
      "content": "<p>Thank you for your sharing. I applied this leaked file to my submission. <br>\nBut decreased my public score.   </p>\n\n<pre><code>import pandas as pd\n# my submission file\ndf = pd.read_csv(\"xxx.csv\")\nleak_df = pd.read_csv(\"../data/leaks.csv\")\n\nleak_map = {}\n\nfor idx, row in leak_df.iterrows():\n    leak_map[row[\"b_test_img\"]] = row[\"b_label\"]\n\nsubmission_list = []\nfor idx, row in df.iterrows():\n    if row[\"Image\"] in leak_map:\n        id_list = row[\"Id\"].split(\" \")\n        if id_list[0] != leak_map[row[\"Image\"]]:\n            print(id_list[0], leak_map[row[\"Image\"]])\n            print(id_list)\n        id_list[0] = leak_map[row[\"Image\"]]\n        id_string = \" \".join(id_list)\n    else:\n        id_string = row[\"Id\"]\n    submission_list.append(id_string)\ndf[\"Id\"] = submission_list\n# modified output\ndf.to_csv(\"zzz.csv\", index=False)\n</code></pre>\n\n<p>```</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 474328,
      "author_name": "sophie0308",
      "author_url": "",
      "post_date": "02/19/2019 08:30:11",
      "content": "<p>Thanks a lot! My LB has improved +0.002.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 474756,
      "author_name": "kaepyro",
      "author_url": "",
      "post_date": "02/19/2019 19:04:49",
      "content": "<p>Thanks a lot for sharing!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "469091": "Hi all, the secret comes from the playground dataset (There are some duplicate images between test set and playground dataset) and motivated by @martinpiotte 's work ( https://www.kaggle.com/martinpiotte/whale-recognition-model-with-score-0-78563 ). \n\nSome teams may not know the previous competition and haven't downloaded the playground images. Therefore, I decided to publish all leaks I detected for you for equity.\n\n**Usages**\n\n- Simply replace these 135 samples' top 1 prediction\n- You can also consider it as an extra validation set, calculate the MAP@5 for them to estimate your model's performance. In my experiment, this validation set is a good reference\n- You can estimate the public/private distributions according to them (welcome to tell us the result, in order to save the submission times, I haven't taken this kind of test)\n\nFinally, I’d like to mention that:\n\n- I'm not 100% sure that all of these results are true\n- Maybe there are leaks I don't find, but I'll not continue this work because in my opinion, many other tasks are more valuable, like bounding box regression, segmentation, and key points regression\n- the organizer may (or already) remove these samples in private LB, thus please don't down vote me after the competition if \"leaks\" are useless",
    "469126": "one free leak from the test data itself  ... maybe they forget to remove bottom image in the composite test image\n\n\n  ![enter image description here][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/469126/11241/leak.png",
    "469188": "Thanks for sharing. But is it legal to use this kind of leak? I mean, the leak is extracted using data from another competition.\n\n-Edit- Actually it is legal if playground dataset is disclosed in the external data thread and the duplicates were found automatically, right?",
    "469289": "Hi Eduardo, don't worry. \n\nThe organizer said \"A Playground version of this competition was hosted a few months ago. This competition contains even more images and individual whales than the last. While you can certainly use the images in the previous competition if you wish\"",
    "469310": "Thanks Venn, very admirable.",
    "469385": "This is so awesome",
    "469674": "Thank you for posting. For the model I am currently working on: LB score 0.789. After correcting 13 samples that my model got wrong LB 0.792, and after correcting the order of 10 other predictions that my model had in a different order, LB 0.794. I just realized that the MAP@5 score depends on the order of the predictions.",
    "469904": "Hi Venn,\nThe labeling between this competition and the playground seem different. To generate the leaks.csv, we have to know the id mappings, may I know how did you get the mappings?",
    "470031": "Hi Guanshuo, yes, it needs to find the duplicates between training set and the playground dataset first to build the mappings.",
    "470141": "I see",
    "470199": "Thank you for your sharing. I applied this leaked file to my submission.  \nBut decreased my public score.   \n  \n    import pandas as pd\n    # my submission file\n    df = pd.read_csv(\"xxx.csv\")\n    leak_df = pd.read_csv(\"../data/leaks.csv\")\n    \n    leak_map = {}\n    \n    for idx, row in leak_df.iterrows():\n        leak_map[row[\"b_test_img\"]] = row[\"b_label\"]\n    \n    submission_list = []\n    for idx, row in df.iterrows():\n        if row[\"Image\"] in leak_map:\n            id_list = row[\"Id\"].split(\" \")\n            if id_list[0] != leak_map[row[\"Image\"]]:\n                print(id_list[0], leak_map[row[\"Image\"]])\n                print(id_list)\n            id_list[0] = leak_map[row[\"Image\"]]\n            id_string = \" \".join(id_list)\n        else:\n            id_string = row[\"Id\"]\n        submission_list.append(id_string)\n    df[\"Id\"] = submission_list\n    # modified output\n    df.to_csv(\"zzz.csv\", index=False)\n```",
    "474328": "Thanks a lot! My LB has improved +0.002.",
    "474756": "Thanks a lot for sharing!"
  },
  "source": "meta"
}