{
  "id": 305341,
  "title": "The Dataset seems to have good \"Inter-Species\" Separation",
  "url": "/competitions/happy-whale-and-dolphin/discussion/305341",
  "author_name": "",
  "post_date": "2022-02-04T21:00:36.422661300Z",
  "votes": 38,
  "comment_count": 8,
  "views": 0,
  "content": "<h1>Introduction</h1>\n<p>This competition is about learning the best embedding that brings together images of same mammal (belong to same species as well). But we will take things one at a time. </p>\n<h1>Step 1: Visualizing Similar Images by <code>individual_id</code>.</h1>\n<p>After the competition was launched, I quickly logged the images along with their <code>individual_id</code>s as W&amp;B Tables. The section, \"1. Log the Images as W&amp;B Table [Optional]\" of this <a href=\"https://www.kaggle.com/ayuraj/easy-pca-tsne-and-umap#1.-Log-the-Images-as-W&amp;B-Table-%5BOptional%5D\" target=\"_blank\">kernel</a> shows how I achieved this. Here's the <a href=\"https://wandb.ai/ayut/happywhale/runs/34086otd\" target=\"_blank\">logged table</a>. </p>\n<h2>Observations</h2>\n<ul>\n<li>There are multiple <code>individual_id</code>s with just 1 image. This makes learning relevant feature from it harder using regular ways. </li>\n<li>The visual cues boil down to the exposed part of the mammal. Is there a cut in the fin? Is the mammal grey in color with visible patch (body mark)? </li>\n<li>There are 15587 unique ids. </li>\n</ul>\n<p><img src=\"https://i.imgur.com/0S0cWe2.mp4\" alt=\"img\"></p>\n<h1>Step 2: Naive Approach: Inter-Species Classifier</h1>\n<p>Learn embedding meaningful enough to predict the species of the mammal. I have fine-tuned EfficientNetB0 on this competition's dataset. It was trained for 30 epochs with 5 fold stratified split of the dataset. No fancy augmentation or training techniques was used. <code>ReduceLRonPlateua</code> was used as a learning rate scheduler while training. You can find the trained models <a href=\"https://www.kaggle.com/ayuraj/happywhale-supervised\" target=\"_blank\">here</a>. The 128x128 dataset by <a href=\"https://www.kaggle.com/rdizzl3\" target=\"_blank\">@rdizzl3</a> was used to train the models.</p>\n<p>⌛️ I will soon publish the training kernel.</p>\n<h2>Remarks</h2>\n<ul>\n<li>From the metrics, we can see that it's easy to fit a model on this dataset to predict the species of the mammals. The model achieved a mean top 1 validation accuracy of 0.9548 with standard deviation of 0.0061. </li>\n<li>By using augmentation techniques, we can maybe learn even better model. </li>\n</ul>\n<p><img src=\"https://i.imgur.com/RCqv3Zz.png\" alt=\"img\"></p>\n<h1>Step 3: Visualize Embedding using W&amp;B Embedding Projector</h1>\n<p>I was waiting to try out the newly launched Embedding Projector by Weights and Biases. <strong>Check our my kernel <a href=\"https://www.kaggle.com/ayuraj/easy-pca-tsne-and-umap#Visualize-Embedding-using-W&amp;B-Embedding-Projector\" target=\"_blank\">Easy PCA, TSNE, and UMAP</a></strong> to see how to easily visualize the embeddings in 2D space. </p>\n<p>So what do we have?</p>\n<h2>Observations</h2>\n<ul>\n<li>Turns our there's a good inter-species separation. The UMAP algorithm gives the best 2D projection of embedding space. </li>\n<li>It would be interesting to find ways to achieve maximum intra-species separation. But that's a topic for the next discussion post. </li>\n</ul>\n<pre><code>💡The UMAP projection that you see below was achieved `Neighbors` parameter set to 5. Check out the section, \"4.0 How to Easily Create Embeddings\" in this [kernel](https://www.kaggle.com/ayuraj/easy-pca-tsne-and-umap#4.0-How-to-Easily-Create-Embeddings) to see how projections can easily be created. \n</code></pre>\n<p><img src=\"https://i.imgur.com/Hl301zy.png\" alt=\"img\"></p>",
  "messages": [
    {
      "id": "1676294",
      "postDate": "02/04/2022 21:00:36",
      "content": "<h1>Introduction</h1>\n<p>This competition is about learning the best embedding that brings together images of same mammal (belong to same species as well). But we will take things one at a time. </p>\n<h1>Step 1: Visualizing Similar Images by <code>individual_id</code>.</h1>\n<p>After the competition was launched, I quickly logged the images along with their <code>individual_id</code>s as W&amp;B Tables. The section, \"1. Log the Images as W&amp;B Table [Optional]\" of this <a href=\"https://www.kaggle.com/ayuraj/easy-pca-tsne-and-umap#1.-Log-the-Images-as-W&amp;B-Table-%5BOptional%5D\" target=\"_blank\">kernel</a> shows how I achieved this. Here's the <a href=\"https://wandb.ai/ayut/happywhale/runs/34086otd\" target=\"_blank\">logged table</a>. </p>\n<h2>Observations</h2>\n<ul>\n<li>There are multiple <code>individual_id</code>s with just 1 image. This makes learning relevant feature from it harder using regular ways. </li>\n<li>The visual cues boil down to the exposed part of the mammal. Is there a cut in the fin? Is the mammal grey in color with visible patch (body mark)? </li>\n<li>There are 15587 unique ids. </li>\n</ul>\n<p><img src=\"https://i.imgur.com/0S0cWe2.mp4\" alt=\"img\"></p>\n<h1>Step 2: Naive Approach: Inter-Species Classifier</h1>\n<p>Learn embedding meaningful enough to predict the species of the mammal. I have fine-tuned EfficientNetB0 on this competition's dataset. It was trained for 30 epochs with 5 fold stratified split of the dataset. No fancy augmentation or training techniques was used. <code>ReduceLRonPlateua</code> was used as a learning rate scheduler while training. You can find the trained models <a href=\"https://www.kaggle.com/ayuraj/happywhale-supervised\" target=\"_blank\">here</a>. The 128x128 dataset by <a href=\"https://www.kaggle.com/rdizzl3\" target=\"_blank\">@rdizzl3</a> was used to train the models.</p>\n<p>⌛️ I will soon publish the training kernel.</p>\n<h2>Remarks</h2>\n<ul>\n<li>From the metrics, we can see that it's easy to fit a model on this dataset to predict the species of the mammals. The model achieved a mean top 1 validation accuracy of 0.9548 with standard deviation of 0.0061. </li>\n<li>By using augmentation techniques, we can maybe learn even better model. </li>\n</ul>\n<p><img src=\"https://i.imgur.com/RCqv3Zz.png\" alt=\"img\"></p>\n<h1>Step 3: Visualize Embedding using W&amp;B Embedding Projector</h1>\n<p>I was waiting to try out the newly launched Embedding Projector by Weights and Biases. <strong>Check our my kernel <a href=\"https://www.kaggle.com/ayuraj/easy-pca-tsne-and-umap#Visualize-Embedding-using-W&amp;B-Embedding-Projector\" target=\"_blank\">Easy PCA, TSNE, and UMAP</a></strong> to see how to easily visualize the embeddings in 2D space. </p>\n<p>So what do we have?</p>\n<h2>Observations</h2>\n<ul>\n<li>Turns our there's a good inter-species separation. The UMAP algorithm gives the best 2D projection of embedding space. </li>\n<li>It would be interesting to find ways to achieve maximum intra-species separation. But that's a topic for the next discussion post. </li>\n</ul>\n<pre><code>💡The UMAP projection that you see below was achieved `Neighbors` parameter set to 5. Check out the section, \"4.0 How to Easily Create Embeddings\" in this [kernel](https://www.kaggle.com/ayuraj/easy-pca-tsne-and-umap#4.0-How-to-Easily-Create-Embeddings) to see how projections can easily be created. \n</code></pre>\n<p><img src=\"https://i.imgur.com/Hl301zy.png\" alt=\"img\"></p>",
      "rawMarkdown": "# Introduction\n\nThis competition is about learning the best embedding that brings together images of same mammal (belong to same species as well). But we will take things one at a time. \n\n# Step 1: Visualizing Similar Images by `individual_id`. \n\nAfter the competition was launched, I quickly logged the images along with their `individual_id`s as W&B Tables. The section, \"1. Log the Images as W&B Table [Optional]\" of this [kernel](https://www.kaggle.com/ayuraj/easy-pca-tsne-and-umap#1.-Log-the-Images-as-W&B-Table-%5BOptional%5D) shows how I achieved this. Here's the [logged table](https://wandb.ai/ayut/happywhale/runs/34086otd). \n\n## Observations\n\n* There are multiple `individual_id`s with just 1 image. This makes learning relevant feature from it harder using regular ways. \n* The visual cues boil down to the exposed part of the mammal. Is there a cut in the fin? Is the mammal grey in color with visible patch (body mark)? \n* There are 15587 unique ids. \n\n![img](https://i.imgur.com/0S0cWe2.mp4)\n\n# Step 2: Naive Approach: Inter-Species Classifier\n\nLearn embedding meaningful enough to predict the species of the mammal. I have fine-tuned EfficientNetB0 on this competition's dataset. It was trained for 30 epochs with 5 fold stratified split of the dataset. No fancy augmentation or training techniques was used. `ReduceLRonPlateua` was used as a learning rate scheduler while training. You can find the trained models [here](https://www.kaggle.com/ayuraj/happywhale-supervised). The 128x128 dataset by @rdizzl3 was used to train the models.\n\n⌛️ I will soon publish the training kernel.\n\n## Remarks\n\n* From the metrics, we can see that it's easy to fit a model on this dataset to predict the species of the mammals. The model achieved a mean top 1 validation accuracy of 0.9548 with standard deviation of 0.0061. \n* By using augmentation techniques, we can maybe learn even better model. \n\n![img](https://i.imgur.com/RCqv3Zz.png)\n\n\n# Step 3: Visualize Embedding using W&B Embedding Projector\n\nI was waiting to try out the newly launched Embedding Projector by Weights and Biases. **Check our my kernel [Easy PCA, TSNE, and UMAP](https://www.kaggle.com/ayuraj/easy-pca-tsne-and-umap#Visualize-Embedding-using-W&B-Embedding-Projector)** to see how to easily visualize the embeddings in 2D space. \n\nSo what do we have?\n\n## Observations\n\n* Turns our there's a good inter-species separation. The UMAP algorithm gives the best 2D projection of embedding space. \n* It would be interesting to find ways to achieve maximum intra-species separation. But that's a topic for the next discussion post. \n\n```\n💡The UMAP projection that you see below was achieved `Neighbors` parameter set to 5. Check out the section, \"4.0 How to Easily Create Embeddings\" in this [kernel](https://www.kaggle.com/ayuraj/easy-pca-tsne-and-umap#4.0-How-to-Easily-Create-Embeddings) to see how projections can easily be created. \n```\n\n![img](https://i.imgur.com/Hl301zy.png)",
      "votes": null
    },
    {
      "id": "1677092",
      "postDate": "02/05/2022 14:22:11",
      "content": "<p>You can get some ideas from my <a href=\"https://www.kaggle.com/kwentar/what-about-species\" target=\"_blank\">notebook</a> about spices, my observation:</p>\n<ol>\n<li>Here is a few typo: bottlenose_dolphin and bottlenose_dolpin, kiler_whale and killer_whale. It means at least two species should me merged in one class</li>\n<li>Some species are very similar and it looks like not possible to split it up effectively: long_finned_pilot_whale and short_finned_pilot_whale, also, pilot_whale is not specie in fact, I guess it is \"we don't know what exactly plot whale it is\". I guess it should be merged too</li>\n<li>Globis looks like \"others\", so, should be removed from species classification training</li>\n</ol>\n<p>UPD: in discussion we can get that globis is \"Short-finned pilot whale ~ Globicephala macrorhynchus\", so, globis could be merged in 2nd</p>",
      "rawMarkdown": "You can get some ideas from my [notebook](https://www.kaggle.com/kwentar/what-about-species) about spices, my observation:\n1. Here is a few typo: bottlenose_dolphin and bottlenose_dolpin, kiler_whale and killer_whale. It means at least two species should me merged in one class\n2. Some species are very similar and it looks like not possible to split it up effectively: long_finned_pilot_whale and short_finned_pilot_whale, also, pilot_whale is not specie in fact, I guess it is \"we don't know what exactly plot whale it is\". I guess it should be merged too\n3. Globis looks like \"others\", so, should be removed from species classification training\n\nUPD: in discussion we can get that globis is \"Short-finned pilot whale ~ Globicephala macrorhynchus\", so, globis could be merged in 2nd",
      "votes": null
    },
    {
      "id": "1677105",
      "postDate": "02/05/2022 14:36:26",
      "content": "<p>apologies for the typos</p>\n<p>both pilot_whale and globis are short_finned_pilot_whale and can be merged</p>",
      "rawMarkdown": "apologies for the typos\n\nboth pilot_whale and globis are short_finned_pilot_whale and can be merged",
      "votes": null
    },
    {
      "id": "1677141",
      "postDate": "02/05/2022 14:54:16",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/ayuraj\" target=\"_blank\">@ayuraj</a> , I think there are some overlaps in colors. The same 10 colors are repeated. You can check out this zoom version. </p>\n<ul>\n<li>As you can see, <code>beluga, frasiers_dolphin, pygmy_killer_whale</code> have the same color(<strong>blue</strong>).</li>\n<li>Similarly, <code>blue_whale, globis, rough_toothed_dolphin</code> have the same color(<strong>orange</strong>).</li>\n</ul>\n<p>Is there anyway to put unique colors here ?<br>\n<img src=\"https://i.ibb.co/qBBtxx0/colors.png\" alt=\"colors\"></p>",
      "rawMarkdown": "Hi @ayuraj , I think there are some overlaps in colors. The same 10 colors are repeated. You can check out this zoom version. \n* As you can see, `beluga, frasiers_dolphin, pygmy_killer_whale` have the same color(**blue**).\n* Similarly, `blue_whale, globis, rough_toothed_dolphin` have the same color(**orange**).\n\nIs there anyway to put unique colors here ?\n<img src=\"https://i.ibb.co/qBBtxx0/colors.png\" alt=\"colors\" border=\"0\" width=400>",
      "votes": null
    },
    {
      "id": "1677291",
      "postDate": "02/05/2022 16:21:59",
      "content": "<p>Nice catch <a href=\"https://www.kaggle.com/awsaf49\" target=\"_blank\">@awsaf49</a>. a common problem when plotting embedding space. You can easily tweak the code to get unrepeating colors but they still might look very similar. Its much better to handpick some distinct colors, but finding 26 distinct colors might be hard. So probably the best way is to use 2-3 different markers +,o,x</p>",
      "rawMarkdown": "Nice catch @awsaf49. a common problem when plotting embedding space. You can easily tweak the code to get unrepeating colors but they still might look very similar. Its much better to handpick some distinct colors, but finding 26 distinct colors might be hard. So probably the best way is to use 2-3 different markers +,o,x",
      "votes": null
    },
    {
      "id": "1677311",
      "postDate": "02/05/2022 16:32:46",
      "content": "<p>Thanks for flagging this <a href=\"https://www.kaggle.com/awsaf49\" target=\"_blank\">@awsaf49</a>. I will get back to you. </p>",
      "rawMarkdown": "Thanks for flagging this @awsaf49. I will get back to you.",
      "votes": null
    },
    {
      "id": "1677555",
      "postDate": "02/05/2022 20:00:43",
      "content": "<p>So we have 26 species in total in this dataset. </p>\n<p><code>['melon_headed_whale',\n 'humpback_whale',\n 'false_killer_whale',\n 'bottlenose_dolphin',\n 'beluga',\n 'minke_whale',\n 'fin_whale',\n 'blue_whale',\n 'gray_whale',\n 'southern_right_whale',\n 'common_dolphin',\n 'killer_whale',\n 'short_finned_pilot_whale',\n 'dusky_dolphin',\n 'long_finned_pilot_whale',\n 'sei_whale',\n 'spinner_dolphin',\n 'cuviers_beaked_whale',\n 'spotted_dolphin',\n 'brydes_whale',\n 'commersons_dolphin',\n 'white_sided_dolphin',\n 'rough_toothed_dolphin',\n 'pantropic_spotted_dolphin',\n 'pygmy_killer_whale',\n 'frasiers_dolphin']</code></p>",
      "rawMarkdown": "So we have 26 species in total in this dataset. \n\n`['melon_headed_whale',\n 'humpback_whale',\n 'false_killer_whale',\n 'bottlenose_dolphin',\n 'beluga',\n 'minke_whale',\n 'fin_whale',\n 'blue_whale',\n 'gray_whale',\n 'southern_right_whale',\n 'common_dolphin',\n 'killer_whale',\n 'short_finned_pilot_whale',\n 'dusky_dolphin',\n 'long_finned_pilot_whale',\n 'sei_whale',\n 'spinner_dolphin',\n 'cuviers_beaked_whale',\n 'spotted_dolphin',\n 'brydes_whale',\n 'commersons_dolphin',\n 'white_sided_dolphin',\n 'rough_toothed_dolphin',\n 'pantropic_spotted_dolphin',\n 'pygmy_killer_whale',\n 'frasiers_dolphin']`",
      "votes": null
    },
    {
      "id": "1678736",
      "postDate": "02/06/2022 18:50:59",
      "content": "<p>Great analysis. I think the big challenge is to be able to predict the species given the image and then predict the correct ID (should devise a validation procedure robust enough to not leak information when trying the 2-stages model etc).</p>",
      "rawMarkdown": "Great analysis. I think the big challenge is to be able to predict the species given the image and then predict the correct ID (should devise a validation procedure robust enough to not leak information when trying the 2-stages model etc).",
      "votes": null
    },
    {
      "id": "1680407",
      "postDate": "02/07/2022 19:26:57",
      "content": "<p>Thanks Yassine. </p>",
      "rawMarkdown": "Thanks Yassine.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1677092,
      "author_name": "kwentar",
      "author_url": "",
      "post_date": "02/05/2022 14:22:11",
      "content": "<p>You can get some ideas from my <a href=\"https://www.kaggle.com/kwentar/what-about-species\" target=\"_blank\">notebook</a> about spices, my observation:</p>\n<ol>\n<li>Here is a few typo: bottlenose_dolphin and bottlenose_dolpin, kiler_whale and killer_whale. It means at least two species should me merged in one class</li>\n<li>Some species are very similar and it looks like not possible to split it up effectively: long_finned_pilot_whale and short_finned_pilot_whale, also, pilot_whale is not specie in fact, I guess it is \"we don't know what exactly plot whale it is\". I guess it should be merged too</li>\n<li>Globis looks like \"others\", so, should be removed from species classification training</li>\n</ol>\n<p>UPD: in discussion we can get that globis is \"Short-finned pilot whale ~ Globicephala macrorhynchus\", so, globis could be merged in 2nd</p>",
      "votes": null,
      "replies": [
        {
          "id": 1677105,
          "author_name": "tedcheese",
          "author_url": "",
          "post_date": "02/05/2022 14:36:26",
          "content": "<p>apologies for the typos</p>\n<p>both pilot_whale and globis are short_finned_pilot_whale and can be merged</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1677555,
          "author_name": "ayuraj",
          "author_url": "",
          "post_date": "02/05/2022 20:00:43",
          "content": "<p>So we have 26 species in total in this dataset. </p>\n<p><code>['melon_headed_whale',\n 'humpback_whale',\n 'false_killer_whale',\n 'bottlenose_dolphin',\n 'beluga',\n 'minke_whale',\n 'fin_whale',\n 'blue_whale',\n 'gray_whale',\n 'southern_right_whale',\n 'common_dolphin',\n 'killer_whale',\n 'short_finned_pilot_whale',\n 'dusky_dolphin',\n 'long_finned_pilot_whale',\n 'sei_whale',\n 'spinner_dolphin',\n 'cuviers_beaked_whale',\n 'spotted_dolphin',\n 'brydes_whale',\n 'commersons_dolphin',\n 'white_sided_dolphin',\n 'rough_toothed_dolphin',\n 'pantropic_spotted_dolphin',\n 'pygmy_killer_whale',\n 'frasiers_dolphin']</code></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1677141,
      "author_name": "awsaf49",
      "author_url": "",
      "post_date": "02/05/2022 14:54:16",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/ayuraj\" target=\"_blank\">@ayuraj</a> , I think there are some overlaps in colors. The same 10 colors are repeated. You can check out this zoom version. </p>\n<ul>\n<li>As you can see, <code>beluga, frasiers_dolphin, pygmy_killer_whale</code> have the same color(<strong>blue</strong>).</li>\n<li>Similarly, <code>blue_whale, globis, rough_toothed_dolphin</code> have the same color(<strong>orange</strong>).</li>\n</ul>\n<p>Is there anyway to put unique colors here ?<br>\n<img src=\"https://i.ibb.co/qBBtxx0/colors.png\" alt=\"colors\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 1677291,
          "author_name": "bsridatta",
          "author_url": "",
          "post_date": "02/05/2022 16:21:59",
          "content": "<p>Nice catch <a href=\"https://www.kaggle.com/awsaf49\" target=\"_blank\">@awsaf49</a>. a common problem when plotting embedding space. You can easily tweak the code to get unrepeating colors but they still might look very similar. Its much better to handpick some distinct colors, but finding 26 distinct colors might be hard. So probably the best way is to use 2-3 different markers +,o,x</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1677311,
          "author_name": "ayuraj",
          "author_url": "",
          "post_date": "02/05/2022 16:32:46",
          "content": "<p>Thanks for flagging this <a href=\"https://www.kaggle.com/awsaf49\" target=\"_blank\">@awsaf49</a>. I will get back to you. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1678736,
      "author_name": "yassinealouini",
      "author_url": "",
      "post_date": "02/06/2022 18:50:59",
      "content": "<p>Great analysis. I think the big challenge is to be able to predict the species given the image and then predict the correct ID (should devise a validation procedure robust enough to not leak information when trying the 2-stages model etc).</p>",
      "votes": null,
      "replies": [
        {
          "id": 1680407,
          "author_name": "ayuraj",
          "author_url": "",
          "post_date": "02/07/2022 19:26:57",
          "content": "<p>Thanks Yassine. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1676294": "# Introduction\n\nThis competition is about learning the best embedding that brings together images of same mammal (belong to same species as well). But we will take things one at a time. \n\n# Step 1: Visualizing Similar Images by `individual_id`. \n\nAfter the competition was launched, I quickly logged the images along with their `individual_id`s as W&B Tables. The section, \"1. Log the Images as W&B Table [Optional]\" of this [kernel](https://www.kaggle.com/ayuraj/easy-pca-tsne-and-umap#1.-Log-the-Images-as-W&B-Table-%5BOptional%5D) shows how I achieved this. Here's the [logged table](https://wandb.ai/ayut/happywhale/runs/34086otd). \n\n## Observations\n\n* There are multiple `individual_id`s with just 1 image. This makes learning relevant feature from it harder using regular ways. \n* The visual cues boil down to the exposed part of the mammal. Is there a cut in the fin? Is the mammal grey in color with visible patch (body mark)? \n* There are 15587 unique ids. \n\n![img](https://i.imgur.com/0S0cWe2.mp4)\n\n# Step 2: Naive Approach: Inter-Species Classifier\n\nLearn embedding meaningful enough to predict the species of the mammal. I have fine-tuned EfficientNetB0 on this competition's dataset. It was trained for 30 epochs with 5 fold stratified split of the dataset. No fancy augmentation or training techniques was used. `ReduceLRonPlateua` was used as a learning rate scheduler while training. You can find the trained models [here](https://www.kaggle.com/ayuraj/happywhale-supervised). The 128x128 dataset by @rdizzl3 was used to train the models.\n\n⌛️ I will soon publish the training kernel.\n\n## Remarks\n\n* From the metrics, we can see that it's easy to fit a model on this dataset to predict the species of the mammals. The model achieved a mean top 1 validation accuracy of 0.9548 with standard deviation of 0.0061. \n* By using augmentation techniques, we can maybe learn even better model. \n\n![img](https://i.imgur.com/RCqv3Zz.png)\n\n\n# Step 3: Visualize Embedding using W&B Embedding Projector\n\nI was waiting to try out the newly launched Embedding Projector by Weights and Biases. **Check our my kernel [Easy PCA, TSNE, and UMAP](https://www.kaggle.com/ayuraj/easy-pca-tsne-and-umap#Visualize-Embedding-using-W&B-Embedding-Projector)** to see how to easily visualize the embeddings in 2D space. \n\nSo what do we have?\n\n## Observations\n\n* Turns our there's a good inter-species separation. The UMAP algorithm gives the best 2D projection of embedding space. \n* It would be interesting to find ways to achieve maximum intra-species separation. But that's a topic for the next discussion post. \n\n```\n💡The UMAP projection that you see below was achieved `Neighbors` parameter set to 5. Check out the section, \"4.0 How to Easily Create Embeddings\" in this [kernel](https://www.kaggle.com/ayuraj/easy-pca-tsne-and-umap#4.0-How-to-Easily-Create-Embeddings) to see how projections can easily be created. \n```\n\n![img](https://i.imgur.com/Hl301zy.png)",
    "1677092": "You can get some ideas from my [notebook](https://www.kaggle.com/kwentar/what-about-species) about spices, my observation:\n1. Here is a few typo: bottlenose_dolphin and bottlenose_dolpin, kiler_whale and killer_whale. It means at least two species should me merged in one class\n2. Some species are very similar and it looks like not possible to split it up effectively: long_finned_pilot_whale and short_finned_pilot_whale, also, pilot_whale is not specie in fact, I guess it is \"we don't know what exactly plot whale it is\". I guess it should be merged too\n3. Globis looks like \"others\", so, should be removed from species classification training\n\nUPD: in discussion we can get that globis is \"Short-finned pilot whale ~ Globicephala macrorhynchus\", so, globis could be merged in 2nd",
    "1677105": "apologies for the typos\n\nboth pilot_whale and globis are short_finned_pilot_whale and can be merged",
    "1677141": "Hi @ayuraj , I think there are some overlaps in colors. The same 10 colors are repeated. You can check out this zoom version. \n* As you can see, `beluga, frasiers_dolphin, pygmy_killer_whale` have the same color(**blue**).\n* Similarly, `blue_whale, globis, rough_toothed_dolphin` have the same color(**orange**).\n\nIs there anyway to put unique colors here ?\n<img src=\"https://i.ibb.co/qBBtxx0/colors.png\" alt=\"colors\" border=\"0\" width=400>",
    "1677291": "Nice catch @awsaf49. a common problem when plotting embedding space. You can easily tweak the code to get unrepeating colors but they still might look very similar. Its much better to handpick some distinct colors, but finding 26 distinct colors might be hard. So probably the best way is to use 2-3 different markers +,o,x",
    "1677311": "Thanks for flagging this @awsaf49. I will get back to you.",
    "1677555": "So we have 26 species in total in this dataset. \n\n`['melon_headed_whale',\n 'humpback_whale',\n 'false_killer_whale',\n 'bottlenose_dolphin',\n 'beluga',\n 'minke_whale',\n 'fin_whale',\n 'blue_whale',\n 'gray_whale',\n 'southern_right_whale',\n 'common_dolphin',\n 'killer_whale',\n 'short_finned_pilot_whale',\n 'dusky_dolphin',\n 'long_finned_pilot_whale',\n 'sei_whale',\n 'spinner_dolphin',\n 'cuviers_beaked_whale',\n 'spotted_dolphin',\n 'brydes_whale',\n 'commersons_dolphin',\n 'white_sided_dolphin',\n 'rough_toothed_dolphin',\n 'pantropic_spotted_dolphin',\n 'pygmy_killer_whale',\n 'frasiers_dolphin']`",
    "1678736": "Great analysis. I think the big challenge is to be able to predict the species given the image and then predict the correct ID (should devise a validation procedure robust enough to not leak information when trying the 2-stages model etc).",
    "1680407": "Thanks Yassine."
  },
  "source": "meta"
}