{
  "id": 73257,
  "title": "Some findings and sharing from Playground Competition",
  "url": "/competitions/humpback-whale-identification/discussion/73257",
  "author_name": "",
  "post_date": "2018-12-01T07:01:55.287884700Z",
  "votes": 48,
  "comment_count": 21,
  "views": 0,
  "content": "<ul>\n<li><strong>Fix:</strong> although these two competitions' pages show the same size --- training set: 4.16GB, test set: 1.35G. The data in playground one is less. If you click the download button, you will find the truth. For example, the test set for playground is 425MB actually. <a href=\"https://www.kaggle.com/c/whale-categorization-playground\">https://www.kaggle.com/c/whale-categorization-playground</a> </li>\n<li><strong>Experiences</strong> Previous wrap-up solutions: <a href=\"https://www.kaggle.com/c/whale-categorization-playground/discussion/47419\">https://www.kaggle.com/c/whale-categorization-playground/discussion/47419</a></li>\n<li><strong>Solution</strong> Previous 1st place kernel: <a href=\"https://www.kaggle.com/martinpiotte/whale-recognition-model-with-score-0-78563\">https://www.kaggle.com/martinpiotte/whale-recognition-model-with-score-0-78563</a></li>\n</ul>\n\n<p>Welcome to add more useful information : )</p>",
  "messages": [
    {
      "id": "430901",
      "postDate": "12/01/2018 07:01:55",
      "content": "<ul>\n<li><strong>Fix:</strong> although these two competitions' pages show the same size --- training set: 4.16GB, test set: 1.35G. The data in playground one is less. If you click the download button, you will find the truth. For example, the test set for playground is 425MB actually. <a href=\"https://www.kaggle.com/c/whale-categorization-playground\">https://www.kaggle.com/c/whale-categorization-playground</a> </li>\n<li><strong>Experiences</strong> Previous wrap-up solutions: <a href=\"https://www.kaggle.com/c/whale-categorization-playground/discussion/47419\">https://www.kaggle.com/c/whale-categorization-playground/discussion/47419</a></li>\n<li><strong>Solution</strong> Previous 1st place kernel: <a href=\"https://www.kaggle.com/martinpiotte/whale-recognition-model-with-score-0-78563\">https://www.kaggle.com/martinpiotte/whale-recognition-model-with-score-0-78563</a></li>\n</ul>\n\n<p>Welcome to add more useful information : )</p>",
      "rawMarkdown": "**Fix:** although these two competitions' pages show the same size --- training set: 4.16GB, test set: 1.35G. The data in playground one is less. If you click the download button, you will find the truth. For example, the test set for playground is 425MB actually. https://www.kaggle.com/c/whale-categorization-playground \n- **Experiences** Previous wrap-up solutions: https://www.kaggle.com/c/whale-categorization-playground/discussion/47419\n- **Solution** Previous 1st place kernel: https://www.kaggle.com/martinpiotte/whale-recognition-model-with-score-0-78563\n\nWelcome to add more useful information : )",
      "votes": null
    },
    {
      "id": "431095",
      "postDate": "12/01/2018 16:04:05",
      "content": "<p><a href=\"/yiheng\">@yiheng</a>, the first link is broken. ')' at the and. Thanks for sharing.</p>",
      "rawMarkdown": "yiheng, the first link is broken. ')' at the and. Thanks for sharing.",
      "votes": null
    },
    {
      "id": "431125",
      "postDate": "12/01/2018 16:46:40",
      "content": "<p>Thanks Peter, I modified it : )</p>",
      "rawMarkdown": "Thanks Peter, I modified it : )",
      "votes": null
    },
    {
      "id": "431154",
      "postDate": "12/01/2018 17:15:59",
      "content": "<p>It seems that the top solution used bounding boxex. Quite interesting...</p>",
      "rawMarkdown": "It seems that the top solution used bounding boxex. Quite interesting...",
      "votes": null
    },
    {
      "id": "431229",
      "postDate": "12/01/2018 22:05:57",
      "content": "<p>The finding about the identity of the dataset seems contradictory to the note in the competition description \"Note, this competition is similar in nature to this competition with an expanded and updated dataset.\" </p>\n\n<p>As well as to <a href=\"/inversion\">@inversion</a> statement in the welcome post \"This competition contains even more images and individual whales than the last\". </p>\n\n<p>Maybe it would be good to get a clarification early about the coverage of common examples in both competitions.</p>",
      "rawMarkdown": "The finding about the identity of the dataset seems contradictory to the note in the competition description \"Note, this competition is similar in nature to this competition with an expanded and updated dataset.\" \n\nAs well as to @inversion statement in the welcome post \"This competition contains even more images and individual whales than the last\". \n\nMaybe it would be good to get a clarification early about the coverage of common examples in both competitions.",
      "votes": null
    },
    {
      "id": "431231",
      "postDate": "12/01/2018 22:32:36",
      "content": "<p>What I believe right now is that the dataset is different in this competition.</p>\n\n<p>When on look at some of the playground's kernels we can see that the count of the dataset was different:\nExample: <a href=\"https://www.kaggle.com/mmrosenb/whales-an-exploration\">playground kernel</a> shows 9850 train and 15610 test images. which sums up to 25460 which is almost 25361 train images in this competition. </p>\n\n<p>What might have happened is that the dataset for playground was altered to become this dataset. Still it would be interesting to know what actually changed so that we can know how much knowledge from playground is applicable.</p>",
      "rawMarkdown": "What I believe right now is that the dataset is different in this competition.\n\nWhen on look at some of the playground's kernels we can see that the count of the dataset was different:\nExample: [playground kernel](https://www.kaggle.com/mmrosenb/whales-an-exploration) shows 9850 train and 15610 test images. which sums up to 25460 which is almost 25361 train images in this competition. \n\nWhat might have happened is that the dataset for playground was altered to become this dataset. Still it would be interesting to know what actually changed so that we can know how much knowledge from playground is applicable.",
      "votes": null
    },
    {
      "id": "431461",
      "postDate": "12/02/2018 10:23:00",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing.",
      "votes": null
    },
    {
      "id": "431503",
      "postDate": "12/02/2018 11:56:17",
      "content": "<p>Other finding is that the quality of the dataset has improved in terms of duplicate images.</p>\n\n<p>Previously there were lots of duplicate images (by pHash): <a href=\"https://www.kaggle.com/stehai/duplicate-images\">playground duplicates</a>. Now there is a negligible amount of duplicates, which I confirmed in <a href=\"https://www.kaggle.com/kretes/eda-distributions-images-and-no-duplicates\">this competition duplicates</a></p>",
      "rawMarkdown": "Other finding is that the quality of the dataset has improved in terms of duplicate images.\n\nPreviously there were lots of duplicate images (by pHash): [playground duplicates](https://www.kaggle.com/stehai/duplicate-images). Now there is a negligible amount of duplicates, which I confirmed in [this competition duplicates](https://www.kaggle.com/kretes/eda-distributions-images-and-no-duplicates)",
      "votes": null
    },
    {
      "id": "431968",
      "postDate": "12/03/2018 07:34:16",
      "content": "<p>Thanks for the links</p>",
      "rawMarkdown": "Thanks for the links",
      "votes": null
    },
    {
      "id": "433240",
      "postDate": "12/04/2018 22:05:50",
      "content": "<p>Nice!</p>",
      "rawMarkdown": "Nice!",
      "votes": null
    },
    {
      "id": "433510",
      "postDate": "12/05/2018 06:15:53",
      "content": "<p>Thanks for sharing <a href=\"/yiheng\">@yiheng</a> </p>",
      "rawMarkdown": "Thanks for sharing @yiheng",
      "votes": null
    },
    {
      "id": "434005",
      "postDate": "12/05/2018 18:56:21",
      "content": "<p>I am wondering how @LuYang got 0.889 accuracy even when the dataset is supposed to have increased from playground competition</p>",
      "rawMarkdown": "I am wondering how @LuYang got 0.889 accuracy even when the dataset is supposed to have increased from playground competition",
      "votes": null
    },
    {
      "id": "434010",
      "postDate": "12/05/2018 19:05:38",
      "content": "<p>As far as I understand he is using either bounding boxes or metric learning.</p>",
      "rawMarkdown": "As far as I understand he is using either bounding boxes or metric learning.",
      "votes": null
    },
    {
      "id": "434201",
      "postDate": "12/06/2018 03:34:53",
      "content": "<p>Thanks @Andrew. That seems interesting. I will try that approach soon</p>",
      "rawMarkdown": "Thanks @Andrew. That seems interesting. I will try that approach soon",
      "votes": null
    },
    {
      "id": "434305",
      "postDate": "12/06/2018 07:16:39",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing",
      "votes": null
    },
    {
      "id": "434675",
      "postDate": "12/06/2018 19:03:29",
      "content": "<p>Someone had already reached .91\nSeems like they have the bbox coords !</p>",
      "rawMarkdown": "Someone had already reached .91\nSeems like they have the bbox coords !",
      "votes": null
    },
    {
      "id": "435803",
      "postDate": "12/08/2018 19:49:35",
      "content": "<p>First point in topic header is not right. In the playground competition training set contains only  9&nbsp;850 items occupying  283,5 MB on the hard drive</p>",
      "rawMarkdown": "First point in topic header is not right. In the playground competition training set contains only  9&nbsp;850 items occupying  283,5 MB on the hard drive",
      "votes": null
    },
    {
      "id": "436261",
      "postDate": "12/10/2018 02:37:10",
      "content": "<p>Thanks, pi-null-mezon. I've fixed the word.</p>",
      "rawMarkdown": "Thanks, pi-null-mezon. I've fixed the word.",
      "votes": null
    },
    {
      "id": "436512",
      "postDate": "12/10/2018 13:11:12",
      "content": "<p>There is strange behavior...🤔 If I click the \"Download All\" button, same dataset is downloaded. On the other hand, different dataset can be downloaded from side bar.</p>",
      "rawMarkdown": "There is strange behavior...🤔 If I click the \"Download All\" button, same dataset is downloaded. On the other hand, different dataset can be downloaded from side bar.",
      "votes": null
    },
    {
      "id": "436833",
      "postDate": "12/11/2018 02:14:42",
      "content": "<p>Seems Kaggle wanna restrict our download of previous dataset, but forget sth?</p>",
      "rawMarkdown": "Seems Kaggle wanna restrict our download of previous dataset, but forget sth?",
      "votes": null
    },
    {
      "id": "437145",
      "postDate": "12/11/2018 12:44:12",
      "content": "<p>In that case, Kaggle restrict download completely in the future. We should download it soon.</p>",
      "rawMarkdown": "In that case, Kaggle restrict download completely in the future. We should download it soon.",
      "votes": null
    },
    {
      "id": "437148",
      "postDate": "12/11/2018 12:56:06",
      "content": "<p>Actually current dataset is kinda simpler than the playground one. For the instance one of my old learning pipeline learned on playground train dataset achives LB 0.57 at aplayground test dataset. Same pipeline learned on current train dataset achives LB 0.78 at current test dataset.  </p>",
      "rawMarkdown": "Actually current dataset is kinda simpler than the playground one. For the instance one of my old learning pipeline learned on playground train dataset achives LB 0.57 at aplayground test dataset. Same pipeline learned on current train dataset achives LB 0.78 at current test dataset.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 431095,
      "author_name": "pestipeti",
      "author_url": "",
      "post_date": "12/01/2018 16:04:05",
      "content": "<p><a href=\"/yiheng\">@yiheng</a>, the first link is broken. ')' at the and. Thanks for sharing.</p>",
      "votes": null,
      "replies": [
        {
          "id": 431125,
          "author_name": "yiheng",
          "author_url": "",
          "post_date": "12/01/2018 16:46:40",
          "content": "<p>Thanks Peter, I modified it : )</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 431154,
      "author_name": "artgor",
      "author_url": "",
      "post_date": "12/01/2018 17:15:59",
      "content": "<p>It seems that the top solution used bounding boxex. Quite interesting...</p>",
      "votes": null,
      "replies": [
        {
          "id": 434675,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "12/06/2018 19:03:29",
          "content": "<p>Someone had already reached .91\nSeems like they have the bbox coords !</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 431229,
      "author_name": "kretes",
      "author_url": "",
      "post_date": "12/01/2018 22:05:57",
      "content": "<p>The finding about the identity of the dataset seems contradictory to the note in the competition description \"Note, this competition is similar in nature to this competition with an expanded and updated dataset.\" </p>\n\n<p>As well as to <a href=\"/inversion\">@inversion</a> statement in the welcome post \"This competition contains even more images and individual whales than the last\". </p>\n\n<p>Maybe it would be good to get a clarification early about the coverage of common examples in both competitions.</p>",
      "votes": null,
      "replies": [
        {
          "id": 431231,
          "author_name": "kretes",
          "author_url": "",
          "post_date": "12/01/2018 22:32:36",
          "content": "<p>What I believe right now is that the dataset is different in this competition.</p>\n\n<p>When on look at some of the playground's kernels we can see that the count of the dataset was different:\nExample: <a href=\"https://www.kaggle.com/mmrosenb/whales-an-exploration\">playground kernel</a> shows 9850 train and 15610 test images. which sums up to 25460 which is almost 25361 train images in this competition. </p>\n\n<p>What might have happened is that the dataset for playground was altered to become this dataset. Still it would be interesting to know what actually changed so that we can know how much knowledge from playground is applicable.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 431503,
          "author_name": "kretes",
          "author_url": "",
          "post_date": "12/02/2018 11:56:17",
          "content": "<p>Other finding is that the quality of the dataset has improved in terms of duplicate images.</p>\n\n<p>Previously there were lots of duplicate images (by pHash): <a href=\"https://www.kaggle.com/stehai/duplicate-images\">playground duplicates</a>. Now there is a negligible amount of duplicates, which I confirmed in <a href=\"https://www.kaggle.com/kretes/eda-distributions-images-and-no-duplicates\">this competition duplicates</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 431461,
      "author_name": "sanikamal",
      "author_url": "",
      "post_date": "12/02/2018 10:23:00",
      "content": "<p>Thanks for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 431968,
      "author_name": "richarde",
      "author_url": "",
      "post_date": "12/03/2018 07:34:16",
      "content": "<p>Thanks for the links</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 433240,
      "author_name": "yanggu",
      "author_url": "",
      "post_date": "12/04/2018 22:05:50",
      "content": "<p>Nice!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 433510,
      "author_name": "phhasian0710",
      "author_url": "",
      "post_date": "12/05/2018 06:15:53",
      "content": "<p>Thanks for sharing <a href=\"/yiheng\">@yiheng</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 434005,
      "author_name": "kurianbenoy",
      "author_url": "",
      "post_date": "12/05/2018 18:56:21",
      "content": "<p>I am wondering how @LuYang got 0.889 accuracy even when the dataset is supposed to have increased from playground competition</p>",
      "votes": null,
      "replies": [
        {
          "id": 434010,
          "author_name": "artgor",
          "author_url": "",
          "post_date": "12/05/2018 19:05:38",
          "content": "<p>As far as I understand he is using either bounding boxes or metric learning.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 434201,
          "author_name": "kurianbenoy",
          "author_url": "",
          "post_date": "12/06/2018 03:34:53",
          "content": "<p>Thanks @Andrew. That seems interesting. I will try that approach soon</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 437148,
          "author_name": "pinullmezon",
          "author_url": "",
          "post_date": "12/11/2018 12:56:06",
          "content": "<p>Actually current dataset is kinda simpler than the playground one. For the instance one of my old learning pipeline learned on playground train dataset achives LB 0.57 at aplayground test dataset. Same pipeline learned on current train dataset achives LB 0.78 at current test dataset.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 434305,
      "author_name": "lamhoangtung",
      "author_url": "",
      "post_date": "12/06/2018 07:16:39",
      "content": "<p>Thanks for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 435803,
      "author_name": "pinullmezon",
      "author_url": "",
      "post_date": "12/08/2018 19:49:35",
      "content": "<p>First point in topic header is not right. In the playground competition training set contains only  9&nbsp;850 items occupying  283,5 MB on the hard drive</p>",
      "votes": null,
      "replies": [
        {
          "id": 436261,
          "author_name": "yiheng",
          "author_url": "",
          "post_date": "12/10/2018 02:37:10",
          "content": "<p>Thanks, pi-null-mezon. I've fixed the word.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 436512,
          "author_name": "toshik",
          "author_url": "",
          "post_date": "12/10/2018 13:11:12",
          "content": "<p>There is strange behavior...🤔 If I click the \"Download All\" button, same dataset is downloaded. On the other hand, different dataset can be downloaded from side bar.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 436833,
          "author_name": "yiheng",
          "author_url": "",
          "post_date": "12/11/2018 02:14:42",
          "content": "<p>Seems Kaggle wanna restrict our download of previous dataset, but forget sth?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 437145,
          "author_name": "toshik",
          "author_url": "",
          "post_date": "12/11/2018 12:44:12",
          "content": "<p>In that case, Kaggle restrict download completely in the future. We should download it soon.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "430901": "**Fix:** although these two competitions' pages show the same size --- training set: 4.16GB, test set: 1.35G. The data in playground one is less. If you click the download button, you will find the truth. For example, the test set for playground is 425MB actually. https://www.kaggle.com/c/whale-categorization-playground \n- **Experiences** Previous wrap-up solutions: https://www.kaggle.com/c/whale-categorization-playground/discussion/47419\n- **Solution** Previous 1st place kernel: https://www.kaggle.com/martinpiotte/whale-recognition-model-with-score-0-78563\n\nWelcome to add more useful information : )",
    "431095": "yiheng, the first link is broken. ')' at the and. Thanks for sharing.",
    "431125": "Thanks Peter, I modified it : )",
    "431154": "It seems that the top solution used bounding boxex. Quite interesting...",
    "431229": "The finding about the identity of the dataset seems contradictory to the note in the competition description \"Note, this competition is similar in nature to this competition with an expanded and updated dataset.\" \n\nAs well as to @inversion statement in the welcome post \"This competition contains even more images and individual whales than the last\". \n\nMaybe it would be good to get a clarification early about the coverage of common examples in both competitions.",
    "431231": "What I believe right now is that the dataset is different in this competition.\n\nWhen on look at some of the playground's kernels we can see that the count of the dataset was different:\nExample: [playground kernel](https://www.kaggle.com/mmrosenb/whales-an-exploration) shows 9850 train and 15610 test images. which sums up to 25460 which is almost 25361 train images in this competition. \n\nWhat might have happened is that the dataset for playground was altered to become this dataset. Still it would be interesting to know what actually changed so that we can know how much knowledge from playground is applicable.",
    "431461": "Thanks for sharing.",
    "431503": "Other finding is that the quality of the dataset has improved in terms of duplicate images.\n\nPreviously there were lots of duplicate images (by pHash): [playground duplicates](https://www.kaggle.com/stehai/duplicate-images). Now there is a negligible amount of duplicates, which I confirmed in [this competition duplicates](https://www.kaggle.com/kretes/eda-distributions-images-and-no-duplicates)",
    "431968": "Thanks for the links",
    "433240": "Nice!",
    "433510": "Thanks for sharing @yiheng",
    "434005": "I am wondering how @LuYang got 0.889 accuracy even when the dataset is supposed to have increased from playground competition",
    "434010": "As far as I understand he is using either bounding boxes or metric learning.",
    "434201": "Thanks @Andrew. That seems interesting. I will try that approach soon",
    "434305": "Thanks for sharing",
    "434675": "Someone had already reached .91\nSeems like they have the bbox coords !",
    "435803": "First point in topic header is not right. In the playground competition training set contains only  9&nbsp;850 items occupying  283,5 MB on the hard drive",
    "436261": "Thanks, pi-null-mezon. I've fixed the word.",
    "436512": "There is strange behavior...🤔 If I click the \"Download All\" button, same dataset is downloaded. On the other hand, different dataset can be downloaded from side bar.",
    "436833": "Seems Kaggle wanna restrict our download of previous dataset, but forget sth?",
    "437145": "In that case, Kaggle restrict download completely in the future. We should download it soon.",
    "437148": "Actually current dataset is kinda simpler than the playground one. For the instance one of my old learning pipeline learned on playground train dataset achives LB 0.57 at aplayground test dataset. Same pipeline learned on current train dataset achives LB 0.78 at current test dataset."
  },
  "source": "meta"
}