{
  "id": 80379,
  "title": "Here is how you can explore model's predictions",
  "url": "/competitions/humpback-whale-identification/discussion/80379",
  "author_name": "Haider Alwasiti",
  "post_date": "2019-02-13T06:01:06.090000",
  "votes": 20,
  "comment_count": 14,
  "views": 0,
  "content": "<p>I've created this code to explore the performance of the metric learning NN. The metric deep learning model that I am using is distance based which is based on @Iafoss kernel. The scores are interpreted as the distance between the image that we are testing with other images which  belong to known classes. The distance is a measure of the similarity of features in the embedding space. </p>\n\n<p>The function <code>image_matrix_draw</code> is going to associate two groups of images to each prediciton of your classifier's validation set predicitons. When you train a NN image classifier, you want to see how it is performing. One way to check its performance, and see why it failed on those wrong images, is to checkup the validation images that have been wrongly classified and to compare it with the images of the wrong class + with the images of the ground truth class. This notebook is doing just that.</p>\n\n<p>Let's take an example:</p>\n\n<p><strong>top-1 correct:</strong> is plotting the 1st image as the top-1 predicted image (which is bounded by a rectangle). And then plots the most similar same class images in an ascending scores (from the most similar that the model think is, to the less  similar images).  The 1st column is the repeated images of the image that we are asking the model to predict. The columns 2-7 are the prediction images sorted. (The columns 5-7 are the ground truth class images sorted if the prediciton is incorrect which is not the case here)</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/470532/11271/000000.37231_w_67a9841_w_67a9841_42a505fa7.jpg\" alt=\"top-1 correct\"></p>\n\n<p><strong>top-1 incorrect:</strong> Here the incorrectly classified top-1 image. And just like before, it plots the most similar same class images in an ascending scores. This time there are ground truth images (maximum 12 images) plotted in the same way (from the most similar to the less similar)</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/470532/11275/000032.75000_new_whale_w_42e0e40_da92fb1c1.jpg\" alt=\"top-1 incorrect\"></p>\n\n<p><strong>top-2 incorrect:</strong></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/470532/11272/000051.75000_w_e99ed06_w_c3e88ae_4a5aa11c9.jpg\" alt=\"top-2 incorrect\"></p>\n\n<p><strong>top-k correct:</strong> Here the plotted images are for the top-1 predicted image then the top-2, top-3 ...etc. Ground truth class plotted too in top-k correct plots, since we cannot see enough examples of top-1 images in the columns 2-4, so I thought it would be interesting to see correct class examples.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/470532/11274/000004.87500_w_7c5b20d_w_7c5b20d_0bd65a44d.jpg\" alt=\"top-k correct\"></p>\n\n<p><strong>top-k incorrect:</strong> </p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/470532/11273/000001.58105_w_3d1f606_new_whale_8b39ae55c.jpg\" alt=\"top-k incorrect\"></p>\n\n<p>All the above plots can be generated for the test set too. However, since there is no ground truth labels, it does not classify them into two correct/incorrect folders, and there are no ground truth images plotted.  It took ~10-20 minuted to generate each folder for the above categories on i9-9900k with 16 logical cores. </p>\n\n<p>The function to explore the prediciton's of a classifier can be used for other metric or Siamese learning NN or even used for classification models with a few modifications in the arguments passed.  </p>\n\n<p>Code:\n<a href=\"https://github.com/hwasiti/prediction_explorer\">https://github.com/hwasiti/prediction_explorer</a></p>\n\n<p>I think the notebook is self explanatory. If you have any question please feel free to ask. I worked ~1 week on this code. I think for any serious datascience project, there should be a proper way to know where is the weakness of the model. And how to remedy the issue. And this way you can get the intuition how to do that.</p>\n\n<p><strong>Sorting from most correct to most confused to most wrong predictions</strong>\nThe file name generated with the score being the 1st term. Which allows you to sort the file names in the folder of the top-1 correct (for instance), in ascending and you will get the <em>most correct items</em> in the top and <em>most confused</em> (barely correct) images at the bottom. The same sorting is useful for top-1 incorrect folder in descending. The top files will be the most confused (just barely classified) wrong, and the bottom the <em>most wrong</em> predictions.</p>\n\n<p>Please note that I have not applied distance cut (dcut) that I use to assign <code>new_whale</code> label when the distance of features in the embedding space is more than dcut parameter. (dcut is estimated by the best validation score checkup by brute force: check from dcut 1.0 to 30.0 to find which dcut is giving the best score). And the reason is, I did not want to distort the performance of the model by doing so. Dcut is a post-processing step, and I am afraid it will invalidate our insight on why and where the model is doing good and where it is failing.</p>\n\n<p><strong>P.S.:</strong> in case you are using @Iafoss kernel, <a href=\"https://www.kaggle.com/iafoss/similarity-densenet169-0-794lb-kernel-time-limit#470545\">here</a> is how to generate the pandas dataframes needed to generate these images. Otherwise you can modify the column names from your own model output as shown in the notebook in the github page.</p>",
  "messages": [
    {
      "id": 470532,
      "postDate": "2019-02-13T06:01:06.090Z",
      "content": "<p>I've created this code to explore the performance of the metric learning NN. The metric deep learning model that I am using is distance based which is based on @Iafoss kernel. The scores are interpreted as the distance between the image that we are testing with other images which  belong to known classes. The distance is a measure of the similarity of features in the embedding space. </p>\n\n<p>The function <code>image_matrix_draw</code> is going to associate two groups of images to each prediciton of your classifier's validation set predicitons. When you train a NN image classifier, you want to see how it is performing. One way to check its performance, and see why it failed on those wrong images, is to checkup the validation images that have been wrongly classified and to compare it with the images of the wrong class + with the images of the ground truth class. This notebook is doing just that.</p>\n\n<p>Let's take an example:</p>\n\n<p><strong>top-1 correct:</strong> is plotting the 1st image as the top-1 predicted image (which is bounded by a rectangle). And then plots the most similar same class images in an ascending scores (from the most similar that the model think is, to the less  similar images).  The 1st column is the repeated images of the image that we are asking the model to predict. The columns 2-7 are the prediction images sorted. (The columns 5-7 are the ground truth class images sorted if the prediciton is incorrect which is not the case here)</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/470532/11271/000000.37231_w_67a9841_w_67a9841_42a505fa7.jpg\" alt=\"top-1 correct\"></p>\n\n<p><strong>top-1 incorrect:</strong> Here the incorrectly classified top-1 image. And just like before, it plots the most similar same class images in an ascending scores. This time there are ground truth images (maximum 12 images) plotted in the same way (from the most similar to the less similar)</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/470532/11275/000032.75000_new_whale_w_42e0e40_da92fb1c1.jpg\" alt=\"top-1 incorrect\"></p>\n\n<p><strong>top-2 incorrect:</strong></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/470532/11272/000051.75000_w_e99ed06_w_c3e88ae_4a5aa11c9.jpg\" alt=\"top-2 incorrect\"></p>\n\n<p><strong>top-k correct:</strong> Here the plotted images are for the top-1 predicted image then the top-2, top-3 ...etc. Ground truth class plotted too in top-k correct plots, since we cannot see enough examples of top-1 images in the columns 2-4, so I thought it would be interesting to see correct class examples.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/470532/11274/000004.87500_w_7c5b20d_w_7c5b20d_0bd65a44d.jpg\" alt=\"top-k correct\"></p>\n\n<p><strong>top-k incorrect:</strong> </p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/470532/11273/000001.58105_w_3d1f606_new_whale_8b39ae55c.jpg\" alt=\"top-k incorrect\"></p>\n\n<p>All the above plots can be generated for the test set too. However, since there is no ground truth labels, it does not classify them into two correct/incorrect folders, and there are no ground truth images plotted.  It took ~10-20 minuted to generate each folder for the above categories on i9-9900k with 16 logical cores. </p>\n\n<p>The function to explore the prediciton's of a classifier can be used for other metric or Siamese learning NN or even used for classification models with a few modifications in the arguments passed.  </p>\n\n<p>Code:\n<a href=\"https://github.com/hwasiti/prediction_explorer\">https://github.com/hwasiti/prediction_explorer</a></p>\n\n<p>I think the notebook is self explanatory. If you have any question please feel free to ask. I worked ~1 week on this code. I think for any serious datascience project, there should be a proper way to know where is the weakness of the model. And how to remedy the issue. And this way you can get the intuition how to do that.</p>\n\n<p><strong>Sorting from most correct to most confused to most wrong predictions</strong>\nThe file name generated with the score being the 1st term. Which allows you to sort the file names in the folder of the top-1 correct (for instance), in ascending and you will get the <em>most correct items</em> in the top and <em>most confused</em> (barely correct) images at the bottom. The same sorting is useful for top-1 incorrect folder in descending. The top files will be the most confused (just barely classified) wrong, and the bottom the <em>most wrong</em> predictions.</p>\n\n<p>Please note that I have not applied distance cut (dcut) that I use to assign <code>new_whale</code> label when the distance of features in the embedding space is more than dcut parameter. (dcut is estimated by the best validation score checkup by brute force: check from dcut 1.0 to 30.0 to find which dcut is giving the best score). And the reason is, I did not want to distort the performance of the model by doing so. Dcut is a post-processing step, and I am afraid it will invalidate our insight on why and where the model is doing good and where it is failing.</p>\n\n<p><strong>P.S.:</strong> in case you are using @Iafoss kernel, <a href=\"https://www.kaggle.com/iafoss/similarity-densenet169-0-794lb-kernel-time-limit#470545\">here</a> is how to generate the pandas dataframes needed to generate these images. Otherwise you can modify the column names from your own model output as shown in the notebook in the github page.</p>",
      "rawMarkdown": "I've created this code to explore the performance of the metric learning NN. The metric deep learning model that I am using is distance based which is based on @Iafoss kernel. The scores are interpreted as the distance between the image that we are testing with other images which  belong to known classes. The distance is a measure of the similarity of features in the embedding space. \n\n The function `image_matrix_draw` is going to associate two groups of images to each prediciton of your classifier's validation set predicitons. When you train a NN image classifier, you want to see how it is performing. One way to check its performance, and see why it failed on those wrong images, is to checkup the validation images that have been wrongly classified and to compare it with the images of the wrong class + with the images of the ground truth class. This notebook is doing just that.\n \n Let's take an example:\n \n**top-1 correct:** is plotting the 1st image as the top-1 predicted image (which is bounded by a rectangle). And then plots the most similar same class images in an ascending scores (from the most similar that the model think is, to the less  similar images).  The 1st column is the repeated images of the image that we are asking the model to predict. The columns 2-7 are the prediction images sorted. (The columns 5-7 are the ground truth class images sorted if the prediciton is incorrect which is not the case here)\n\n![top-1 correct](https://storage.googleapis.com/kaggle-forum-message-attachments/470532/11271/000000.37231_w_67a9841_w_67a9841_42a505fa7.jpg)\n\n\n\n**top-1 incorrect:** Here the incorrectly classified top-1 image. And just like before, it plots the most similar same class images in an ascending scores. This time there are ground truth images (maximum 12 images) plotted in the same way (from the most similar to the less similar)\n\n![top-1 incorrect](https://storage.googleapis.com/kaggle-forum-message-attachments/470532/11275/000032.75000_new_whale_w_42e0e40_da92fb1c1.jpg)\n\n\n**top-2 incorrect:**\n\n![top-2 incorrect](https://storage.googleapis.com/kaggle-forum-message-attachments/470532/11272/000051.75000_w_e99ed06_w_c3e88ae_4a5aa11c9.jpg)\n\n\n**top-k correct:** Here the plotted images are for the top-1 predicted image then the top-2, top-3 ...etc. Ground truth class plotted too in top-k correct plots, since we cannot see enough examples of top-1 images in the columns 2-4, so I thought it would be interesting to see correct class examples.\n\n![top-k correct](https://storage.googleapis.com/kaggle-forum-message-attachments/470532/11274/000004.87500_w_7c5b20d_w_7c5b20d_0bd65a44d.jpg)\n\n\n**top-k incorrect:** \n\n![top-k incorrect](https://storage.googleapis.com/kaggle-forum-message-attachments/470532/11273/000001.58105_w_3d1f606_new_whale_8b39ae55c.jpg)\n \n \n \nAll the above plots can be generated for the test set too. However, since there is no ground truth labels, it does not classify them into two correct/incorrect folders, and there are no ground truth images plotted.  It took ~10-20 minuted to generate each folder for the above categories on i9-9900k with 16 logical cores. \n \nThe function to explore the prediciton's of a classifier can be used for other metric or Siamese learning NN or even used for classification models with a few modifications in the arguments passed.  \n\nCode:\n[https://github.com/hwasiti/prediction_explorer][1]\n\nI think the notebook is self explanatory. If you have any question please feel free to ask. I worked ~1 week on this code. I think for any serious datascience project, there should be a proper way to know where is the weakness of the model. And how to remedy the issue. And this way you can get the intuition how to do that.\n\n**Sorting from most correct to most confused to most wrong predictions**\nThe file name generated with the score being the 1st term. Which allows you to sort the file names in the folder of the top-1 correct (for instance), in ascending and you will get the *most correct items* in the top and *most confused* (barely correct) images at the bottom. The same sorting is useful for top-1 incorrect folder in descending. The top files will be the most confused (just barely classified) wrong, and the bottom the *most wrong* predictions.\n\nPlease note that I have not applied distance cut (dcut) that I use to assign `new_whale` label when the distance of features in the embedding space is more than dcut parameter. (dcut is estimated by the best validation score checkup by brute force: check from dcut 1.0 to 30.0 to find which dcut is giving the best score). And the reason is, I did not want to distort the performance of the model by doing so. Dcut is a post-processing step, and I am afraid it will invalidate our insight on why and where the model is doing good and where it is failing.\n\n\n**P.S.:** in case you are using @Iafoss kernel, [here][2] is how to generate the pandas dataframes needed to generate these images. Otherwise you can modify the column names from your own model output as shown in the notebook in the github page.\n\n\n  [1]: https://github.com/hwasiti/prediction_explorer\n  [2]: https://www.kaggle.com/iafoss/similarity-densenet169-0-794lb-kernel-time-limit#470545",
      "votes": 20
    },
    {
      "id": 470749,
      "postDate": "2019-02-13T13:50:40.773Z",
      "content": "<p>Really good article !</p>",
      "rawMarkdown": "Really good article !",
      "votes": 1
    },
    {
      "id": 470687,
      "postDate": "2019-02-13T11:34:55.833Z",
      "content": "<p>Great work!! I'll definitively going to use this, thank you! :)</p>",
      "rawMarkdown": "Great work!! I'll definitively going to use this, thank you! :)",
      "votes": 1
    },
    {
      "id": 470624,
      "postDate": "2019-02-13T09:09:23.720Z",
      "content": "<p>This one seems the same whale, but because the picture compared are black and white vs colored, the details are not quite clear for the model. Perhaps Martin was right turning all images into black and white could solve issues like this.\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/470624/11295/000001.07910_w_57db059_w_43fc74c_ce0de3213.jpg\" alt=\"enter image description here\"></p>",
      "rawMarkdown": "This one seems the same whale, but because the picture compared are black and white vs colored, the details are not quite clear for the model. Perhaps Martin was right turning all images into black and white could solve issues like this.\n![enter image description here][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/470624/11295/000001.07910_w_57db059_w_43fc74c_ce0de3213.jpg",
      "votes": 1
    },
    {
      "id": 470622,
      "postDate": "2019-02-13T09:03:12.380Z",
      "content": "<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/470622/11284/000000.24695_w_f91389a_w_7d8c37c_30dd15cf2.jpg\" alt=\"enter image description here\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/470622/11286/000000.29419_w_9573686_new_whale_75f28d098.jpg\" alt=\"enter link description here\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/470622/11285/000000.28540_w_1c55cc3_new_whale_28d2542db.jpg\" alt=\"enter image description here\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/470622/11287/000000.42505_w_2f1488c_w_c9b4607_8eccdf3d4.jpg\" alt=\"enter image description here\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/470622/11289/000000.85205_new_whale_w_05bf34e_c6a6fe666.jpg\" alt=\"enter image description here\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/470622/11288/000000.51074_new_whale_w_07125ed_17bb74c11.jpg\" alt=\"enter image description here\">\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/470622/11290/000000.63672_new_whale_w_faad3f8_929ee57b1.jpg\" alt=\"enter image description here\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/470622/11291/000000.70312_w_f5d4627_w_2623921_71148b1fc.jpg\" alt=\"enter image description here\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/470622/11292/000000.63916_w_228c7ee_w_f94fc92_e55dfcd2f.jpg\" alt=\"enter image description here\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/470622/11293/000000.78174_w_e73d9a5_new_whale_83b15630f.jpg\" alt=\"enter image description here\">\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/470622/11294/000000.75439_w_8d4c9f7_w_4135cb8_51616faf5.jpg\" alt=\"enter image description here\"></p>",
      "rawMarkdown": "\n![enter image description here][1]\n\n\n![enter link description here][2]\n\n![enter image description here][3]\n\n![enter image description here][4]\n\n![enter image description here][5]\n\n![enter image description here][6]\n![enter image description here][7]\n\n![enter image description here][8]\n\n![enter image description here][9]\n\n![enter image description here][10]\n![enter image description here][11]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/470622/11284/000000.24695_w_f91389a_w_7d8c37c_30dd15cf2.jpg\n  [2]: https://storage.googleapis.com/kaggle-forum-message-attachments/470622/11286/000000.29419_w_9573686_new_whale_75f28d098.jpg\n  [3]: https://storage.googleapis.com/kaggle-forum-message-attachments/470622/11285/000000.28540_w_1c55cc3_new_whale_28d2542db.jpg\n  [4]: https://storage.googleapis.com/kaggle-forum-message-attachments/470622/11287/000000.42505_w_2f1488c_w_c9b4607_8eccdf3d4.jpg\n  [5]: https://storage.googleapis.com/kaggle-forum-message-attachments/470622/11289/000000.85205_new_whale_w_05bf34e_c6a6fe666.jpg\n  [6]: https://storage.googleapis.com/kaggle-forum-message-attachments/470622/11288/000000.51074_new_whale_w_07125ed_17bb74c11.jpg\n  [7]: https://storage.googleapis.com/kaggle-forum-message-attachments/470622/11290/000000.63672_new_whale_w_faad3f8_929ee57b1.jpg\n  [8]: https://storage.googleapis.com/kaggle-forum-message-attachments/470622/11291/000000.70312_w_f5d4627_w_2623921_71148b1fc.jpg\n  [9]: https://storage.googleapis.com/kaggle-forum-message-attachments/470622/11292/000000.63916_w_228c7ee_w_f94fc92_e55dfcd2f.jpg\n  [10]: https://storage.googleapis.com/kaggle-forum-message-attachments/470622/11293/000000.78174_w_e73d9a5_new_whale_83b15630f.jpg\n  [11]: https://storage.googleapis.com/kaggle-forum-message-attachments/470622/11294/000000.75439_w_8d4c9f7_w_4135cb8_51616faf5.jpg",
      "votes": 1
    },
    {
      "id": 470610,
      "postDate": "2019-02-13T08:42:49.457Z",
      "content": "<p>The top-1 seems exactly the same picture cropped. However the rest of the class seems different. So the mislabelling I think  <code>new_whale</code> (944f124ae.jpg)  should be labelled: <code>w_d4d46bf</code></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/470610/11283/000000.21899_new_whale_w_d4d46bf_2b7fd4faf.jpg\" alt=\"enter image description here\"></p>",
      "rawMarkdown": "The top-1 seems exactly the same picture cropped. However the rest of the class seems different. So the mislabelling I think  `new_whale` (944f124ae.jpg)  should be labelled: `w_d4d46bf `\n\n![enter image description here][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/470610/11283/000000.21899_new_whale_w_d4d46bf_2b7fd4faf.jpg",
      "votes": 1
    },
    {
      "id": 472850,
      "postDate": "2019-02-16T19:06:03.600Z",
      "content": "<p><a href=\"/hwasiti\">@hwasiti</a> I can't upvote you enough for this. Thank you so much for this. Just like <a href=\"/iafoss\">@iafoss</a> i see great future (Kaggle Grandmaster) ahead for you. Thanks again.</p>",
      "rawMarkdown": "@hwasiti I can't upvote you enough for this. Thank you so much for this. Just like @iafoss i see great future (Kaggle Grandmaster) ahead for you. Thanks again.",
      "votes": 2,
      "replies": [
        {
          "id": 472902,
          "postDate": "2019-02-16T21:26:26.483Z",
          "content": "<p><a href=\"/liger1776\">@liger1776</a>\nThanks for the nice words...</p>\n\n<p>Btw, there are a lot more mislabeled whales... If you run the code  above (sorted according to the most confident validation predictions, but they are actually wrong: <code>top-1_incorrect</code> file names in ascending), it seems there are a LOT of mislabeled images. I posted here only ~15, but I got tired posting images, because they are keep getting more and more of wrong labeled training data. </p>",
          "rawMarkdown": "@liger1776\nThanks for the nice words...\n\nBtw, there are a lot more mislabeled whales... If you run the code  above (sorted according to the most confident validation predictions, but they are actually wrong: `top-1_incorrect` file names in ascending), it seems there are a LOT of mislabeled images. I posted here only ~15, but I got tired posting images, because they are keep getting more and more of wrong labeled training data. \n",
          "votes": 1
        },
        {
          "id": 472912,
          "postDate": "2019-02-16T22:00:03.563Z",
          "content": "<p>Indeed there are quite a lot, I've found at least two dozen pairs of labels that are different but contain images of the same whale. There are many more but I stopped counting and I have not really looked if there are examples of 3 or more labels all with the same whale, though I assume so.</p>\n\n<p>I don't really know how to deal with this if at all. In all cases that I found so far, if my model predicts one of them as top 1, the other will be 2nd. One thing I'm pretty sure about though, at least for public LB, is that one of them is right and the other (or others) is wrong, as flipping their order lowered the score.</p>",
          "rawMarkdown": "Indeed there are quite a lot, I've found at least two dozen pairs of labels that are different but contain images of the same whale. There are many more but I stopped counting and I have not really looked if there are examples of 3 or more labels all with the same whale, though I assume so.\n\nI don't really know how to deal with this if at all. In all cases that I found so far, if my model predicts one of them as top 1, the other will be 2nd. One thing I'm pretty sure about though, at least for public LB, is that one of them is right and the other (or others) is wrong, as flipping their order lowered the score.",
          "votes": 2
        },
        {
          "id": 472960,
          "postDate": "2019-02-17T01:42:46.343Z",
          "content": "<p>And I've tried to explore the top-1 test set predictions. Of course I would not be able to isolate the wrong predictions into a separate folder like what I did in the validation set. But I tried to investigate the (distance 15 and more) images (which is somehow the most confused predictions because my dcut = 17 for regarding everything farther as new whale). And it does not seem to me that it has &gt; 20% error even for those &gt;15 distance, which are the most confused predictions. \nFor 0.800 MAP5, what I understand is that the model has one 5th of its top-1 prediction wrong, or maybe more because there is some contribution from the top-2,3,4,5</p>\n\n<p>What I speculate is that the model is confused between 2 classes because those 2 classes has been used for training and contains mixed up labels..</p>\n\n<p>Perhaps with a clean dataset, our top LB could be ~ 1.000 !</p>\n\n<blockquote>\n  <p>I don't really know how to deal with this if at all. In all cases that I found so far, if my model predicts one of them as top 1, the other will be 2nd. One thing I'm pretty sure about though, at least for public LB, is that one of them is right and the other (or others) is wrong, as flipping their order lowered the score.</p>\n</blockquote>\n\n<p>Exactly, I was thinking to drop one of the mislabeled images in the training, but then, if the training set is noisy, the test set is noisy too.. I haven't tried it, and do not have a better idea what to do...</p>\n\n<p>If this was my own research, I would run 2-3 times on the whole dataset including the test set, and manually inspect the top 100 of <code>top1-incorrects</code> and manually correct the labels. Then I would have a cleaner dataset that I can go ahead in my adventure... And actually, this was my intention to make this code, and made it flexible enough to be used for future research or competitions..</p>",
          "rawMarkdown": "And I've tried to explore the top-1 test set predictions. Of course I would not be able to isolate the wrong predictions into a separate folder like what I did in the validation set. But I tried to investigate the (distance 15 and more) images (which is somehow the most confused predictions because my dcut = 17 for regarding everything farther as new whale). And it does not seem to me that it has &gt; 20% error even for those &gt;15 distance, which are the most confused predictions. \nFor 0.800 MAP5, what I understand is that the model has one 5th of its top-1 prediction wrong, or maybe more because there is some contribution from the top-2,3,4,5\n\nWhat I speculate is that the model is confused between 2 classes because those 2 classes has been used for training and contains mixed up labels..\n\nPerhaps with a clean dataset, our top LB could be ~ 1.000 !\n\n&gt; I don't really know how to deal with this if at all. In all cases that I found so far, if my model predicts one of them as top 1, the other will be 2nd. One thing I'm pretty sure about though, at least for public LB, is that one of them is right and the other (or others) is wrong, as flipping their order lowered the score.\n\nExactly, I was thinking to drop one of the mislabeled images in the training, but then, if the training set is noisy, the test set is noisy too.. I haven't tried it, and do not have a better idea what to do...\n\nIf this was my own research, I would run 2-3 times on the whole dataset including the test set, and manually inspect the top 100 of `top1-incorrects` and manually correct the labels. Then I would have a cleaner dataset that I can go ahead in my adventure... And actually, this was my intention to make this code, and made it flexible enough to be used for future research or competitions..",
          "votes": 1
        },
        {
          "id": 472964,
          "postDate": "2019-02-17T02:08:44.923Z",
          "content": "<p><a href=\"/interneuron\">@interneuron</a>\nbtw, did you use my code without issues for that? \nAppreciate from you and everybody using my code to give me feedback/criticism on the coding style, approach ...etc. or if you have found any issue...</p>\n\n<p>Here is an example of 3 labelled whales that should be one label. And yes, they are close with each other: top-1, top-2, top-3. And their score are below dcut (dcut for this case was 17.0), so all 3 labels will be included in the test time if such an image would be presented in the test.</p>\n\n<p><img src=\"https://www.dropbox.com/s/bkhvbip2peyv2lt/000000.51074_new_whale_w_07125ed_17bb74c11.jpg?dl=1\" alt=\"enter image description here\"></p>",
          "rawMarkdown": "@interneuron\nbtw, did you use my code without issues for that? \nAppreciate from you and everybody using my code to give me feedback/criticism on the coding style, approach ...etc. or if you have found any issue...\n\nHere is an example of 3 labelled whales that should be one label. And yes, they are close with each other: top-1, top-2, top-3. And their score are below dcut (dcut for this case was 17.0), so all 3 labels will be included in the test time if such an image would be presented in the test.\n\n![enter image description here][1]\n\n\n  [1]: https://www.dropbox.com/s/bkhvbip2peyv2lt/000000.51074_new_whale_w_07125ed_17bb74c11.jpg?dl=1",
          "votes": 1
        },
        {
          "id": 472976,
          "postDate": "2019-02-17T02:48:04.973Z",
          "content": "<p>Hi Haider, I have not had a chance to use your code yet, though I do have it waiting in another tab. Most of the duplicates I've found came from comparing the top-1 predictions of different models, especially when the top-2 are the same except for order, and indeed the same thing does happen with top-3 but much less often.</p>\n\n<p>It seems to me that many (though by no means all) of the mislabeled images come from the same camera burst, they are not exact duplicates but it seems odd they could have acquire different labels. This is pretty much par for biology research though. Even with perfect labels I think getting a 1.00 from these data is achievable only through luck as some of the images are so occluded/ambiguous that neither man nor machine could possible make a positive ID, maybe whale could.</p>\n\n<p>This is all good news for happywhale, as the organizers mentioned:</p>\n\n<blockquote>\n  <p><strong>Ted Cheeseman wrote</strong></p>\n  \n  <blockquote>\n    <p>It is true that there are false negatives in the dataset, whales where we failed to find an existing match, thus assigned a new ID. So, the fact that you are finding these is good news to us as it shows quality in your developing algorithm :) . But as @Andrzej Kuro posted, this doesn't affect the relative competitive result... keep at it!</p>\n  </blockquote>\n</blockquote>\n\n<p>Indeed it seems like the models are starting to outperform manual identification, which is great. The only thing I can think of to deal with the label discrepancies is to prefer the label with more images, in cases where the number of images is the same (which is the most often) I have no idea.</p>\n\n<p>I agree with you about treating the dataset, if it was mine I would probably farm it out to kaggle as well, it seems to be paying off. I'll let you know when I've had a chance to try your code, thanks for posting it! Good luck! </p>",
          "rawMarkdown": "Hi Haider, I have not had a chance to use your code yet, though I do have it waiting in another tab. Most of the duplicates I've found came from comparing the top-1 predictions of different models, especially when the top-2 are the same except for order, and indeed the same thing does happen with top-3 but much less often.\n\nIt seems to me that many (though by no means all) of the mislabeled images come from the same camera burst, they are not exact duplicates but it seems odd they could have acquire different labels. This is pretty much par for biology research though. Even with perfect labels I think getting a 1.00 from these data is achievable only through luck as some of the images are so occluded/ambiguous that neither man nor machine could possible make a positive ID, maybe whale could.\n\nThis is all good news for happywhale, as the organizers mentioned:\n\n&gt; **Ted Cheeseman wrote**\n&gt; \n&gt; &gt; It is true that there are false negatives in the dataset, whales where we failed to find an existing match, thus assigned a new ID. So, the fact that you are finding these is good news to us as it shows quality in your developing algorithm :) . But as @Andrzej Kuro posted, this doesn't affect the relative competitive result... keep at it!\n\nIndeed it seems like the models are starting to outperform manual identification, which is great. The only thing I can think of to deal with the label discrepancies is to prefer the label with more images, in cases where the number of images is the same (which is the most often) I have no idea.\n\nI agree with you about treating the dataset, if it was mine I would probably farm it out to kaggle as well, it seems to be paying off. I'll let you know when I've had a chance to try your code, thanks for posting it! Good luck! ",
          "votes": 1
        }
      ]
    },
    {
      "id": 470606,
      "postDate": "2019-02-13T08:38:55.357Z",
      "content": "<p>Here are my picks of most likely wrong labelled images (more will post later):\nI have sorted top-1_incorrect file names in ascending, and the top images are the most wrong. This one has a distance of similarity of  0.002  which is quite similar, but has a different label:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/470606/11281/000000.00278_w_456646a_new_whale_ada29d516.jpg\" alt=\"enter image description here\"></p>\n\n<p>Also this one:\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/470606/11282/000000.03656_w_6dbfd7a_w_993320b_a541664bc.jpg\" alt=\"enter image description here\"></p>",
      "rawMarkdown": "Here are my picks of most likely wrong labelled images (more will post later):\nI have sorted top-1_incorrect file names in ascending, and the top images are the most wrong. This one has a distance of similarity of  0.002  which is quite similar, but has a different label:\n\n![enter image description here][1]\n\nAlso this one:\n![enter image description here][2]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/470606/11281/000000.00278_w_456646a_new_whale_ada29d516.jpg\n  [2]: https://storage.googleapis.com/kaggle-forum-message-attachments/470606/11282/000000.03656_w_6dbfd7a_w_993320b_a541664bc.jpg",
      "votes": 2
    },
    {
      "id": 472903,
      "postDate": "2019-02-16T21:27:23.103Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 470547,
      "postDate": "2019-02-13T06:48:46.617Z",
      "content": "<p>This is really excellent. Thank you!</p>",
      "rawMarkdown": "This is really excellent. Thank you!",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 470749,
      "author_name": "nitin pawar",
      "author_url": "",
      "post_date": "2019-02-13T13:50:40.773000",
      "content": "<p>Really good article !</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 470687,
      "author_name": "Eduardo Rocha de Andrade",
      "author_url": "",
      "post_date": "2019-02-13T11:34:55.833000",
      "content": "<p>Great work!! I'll definitively going to use this, thank you! :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 470624,
      "author_name": "Haider Alwasiti",
      "author_url": "",
      "post_date": "2019-02-13T09:09:23.720000",
      "content": "<p>This one seems the same whale, but because the picture compared are black and white vs colored, the details are not quite clear for the model. Perhaps Martin was right turning all images into black and white could solve issues like this.\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/470624/11295/000001.07910_w_57db059_w_43fc74c_ce0de3213.jpg\" alt=\"enter image description here\"></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 470622,
      "author_name": "Haider Alwasiti",
      "author_url": "",
      "post_date": "2019-02-13T09:03:12.380000",
      "content": "<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/470622/11284/000000.24695_w_f91389a_w_7d8c37c_30dd15cf2.jpg\" alt=\"enter image description here\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/470622/11286/000000.29419_w_9573686_new_whale_75f28d098.jpg\" alt=\"enter link description here\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/470622/11285/000000.28540_w_1c55cc3_new_whale_28d2542db.jpg\" alt=\"enter image description here\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/470622/11287/000000.42505_w_2f1488c_w_c9b4607_8eccdf3d4.jpg\" alt=\"enter image description here\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/470622/11289/000000.85205_new_whale_w_05bf34e_c6a6fe666.jpg\" alt=\"enter image description here\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/470622/11288/000000.51074_new_whale_w_07125ed_17bb74c11.jpg\" alt=\"enter image description here\">\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/470622/11290/000000.63672_new_whale_w_faad3f8_929ee57b1.jpg\" alt=\"enter image description here\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/470622/11291/000000.70312_w_f5d4627_w_2623921_71148b1fc.jpg\" alt=\"enter image description here\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/470622/11292/000000.63916_w_228c7ee_w_f94fc92_e55dfcd2f.jpg\" alt=\"enter image description here\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/470622/11293/000000.78174_w_e73d9a5_new_whale_83b15630f.jpg\" alt=\"enter image description here\">\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/470622/11294/000000.75439_w_8d4c9f7_w_4135cb8_51616faf5.jpg\" alt=\"enter image description here\"></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 470610,
      "author_name": "Haider Alwasiti",
      "author_url": "",
      "post_date": "2019-02-13T08:42:49.457000",
      "content": "<p>The top-1 seems exactly the same picture cropped. However the rest of the class seems different. So the mislabelling I think  <code>new_whale</code> (944f124ae.jpg)  should be labelled: <code>w_d4d46bf</code></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/470610/11283/000000.21899_new_whale_w_d4d46bf_2b7fd4faf.jpg\" alt=\"enter image description here\"></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 472850,
      "author_name": "student",
      "author_url": "",
      "post_date": "2019-02-16T19:06:03.600000",
      "content": "<p><a href=\"/hwasiti\">@hwasiti</a> I can't upvote you enough for this. Thank you so much for this. Just like <a href=\"/iafoss\">@iafoss</a> i see great future (Kaggle Grandmaster) ahead for you. Thanks again.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 472902,
          "author_name": "Haider Alwasiti",
          "author_url": "",
          "post_date": "2019-02-16T21:26:26.483000",
          "content": "<p><a href=\"/liger1776\">@liger1776</a>\nThanks for the nice words...</p>\n\n<p>Btw, there are a lot more mislabeled whales... If you run the code  above (sorted according to the most confident validation predictions, but they are actually wrong: <code>top-1_incorrect</code> file names in ascending), it seems there are a LOT of mislabeled images. I posted here only ~15, but I got tired posting images, because they are keep getting more and more of wrong labeled training data. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 472912,
          "author_name": "interneuron",
          "author_url": "",
          "post_date": "2019-02-16T22:00:03.563000",
          "content": "<p>Indeed there are quite a lot, I've found at least two dozen pairs of labels that are different but contain images of the same whale. There are many more but I stopped counting and I have not really looked if there are examples of 3 or more labels all with the same whale, though I assume so.</p>\n\n<p>I don't really know how to deal with this if at all. In all cases that I found so far, if my model predicts one of them as top 1, the other will be 2nd. One thing I'm pretty sure about though, at least for public LB, is that one of them is right and the other (or others) is wrong, as flipping their order lowered the score.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 472960,
          "author_name": "Haider Alwasiti",
          "author_url": "",
          "post_date": "2019-02-17T01:42:46.343000",
          "content": "<p>And I've tried to explore the top-1 test set predictions. Of course I would not be able to isolate the wrong predictions into a separate folder like what I did in the validation set. But I tried to investigate the (distance 15 and more) images (which is somehow the most confused predictions because my dcut = 17 for regarding everything farther as new whale). And it does not seem to me that it has &gt; 20% error even for those &gt;15 distance, which are the most confused predictions. \nFor 0.800 MAP5, what I understand is that the model has one 5th of its top-1 prediction wrong, or maybe more because there is some contribution from the top-2,3,4,5</p>\n\n<p>What I speculate is that the model is confused between 2 classes because those 2 classes has been used for training and contains mixed up labels..</p>\n\n<p>Perhaps with a clean dataset, our top LB could be ~ 1.000 !</p>\n\n<blockquote>\n  <p>I don't really know how to deal with this if at all. In all cases that I found so far, if my model predicts one of them as top 1, the other will be 2nd. One thing I'm pretty sure about though, at least for public LB, is that one of them is right and the other (or others) is wrong, as flipping their order lowered the score.</p>\n</blockquote>\n\n<p>Exactly, I was thinking to drop one of the mislabeled images in the training, but then, if the training set is noisy, the test set is noisy too.. I haven't tried it, and do not have a better idea what to do...</p>\n\n<p>If this was my own research, I would run 2-3 times on the whole dataset including the test set, and manually inspect the top 100 of <code>top1-incorrects</code> and manually correct the labels. Then I would have a cleaner dataset that I can go ahead in my adventure... And actually, this was my intention to make this code, and made it flexible enough to be used for future research or competitions..</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 472964,
          "author_name": "Haider Alwasiti",
          "author_url": "",
          "post_date": "2019-02-17T02:08:44.923000",
          "content": "<p><a href=\"/interneuron\">@interneuron</a>\nbtw, did you use my code without issues for that? \nAppreciate from you and everybody using my code to give me feedback/criticism on the coding style, approach ...etc. or if you have found any issue...</p>\n\n<p>Here is an example of 3 labelled whales that should be one label. And yes, they are close with each other: top-1, top-2, top-3. And their score are below dcut (dcut for this case was 17.0), so all 3 labels will be included in the test time if such an image would be presented in the test.</p>\n\n<p><img src=\"https://www.dropbox.com/s/bkhvbip2peyv2lt/000000.51074_new_whale_w_07125ed_17bb74c11.jpg?dl=1\" alt=\"enter image description here\"></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 472976,
          "author_name": "interneuron",
          "author_url": "",
          "post_date": "2019-02-17T02:48:04.973000",
          "content": "<p>Hi Haider, I have not had a chance to use your code yet, though I do have it waiting in another tab. Most of the duplicates I've found came from comparing the top-1 predictions of different models, especially when the top-2 are the same except for order, and indeed the same thing does happen with top-3 but much less often.</p>\n\n<p>It seems to me that many (though by no means all) of the mislabeled images come from the same camera burst, they are not exact duplicates but it seems odd they could have acquire different labels. This is pretty much par for biology research though. Even with perfect labels I think getting a 1.00 from these data is achievable only through luck as some of the images are so occluded/ambiguous that neither man nor machine could possible make a positive ID, maybe whale could.</p>\n\n<p>This is all good news for happywhale, as the organizers mentioned:</p>\n\n<blockquote>\n  <p><strong>Ted Cheeseman wrote</strong></p>\n  \n  <blockquote>\n    <p>It is true that there are false negatives in the dataset, whales where we failed to find an existing match, thus assigned a new ID. So, the fact that you are finding these is good news to us as it shows quality in your developing algorithm :) . But as @Andrzej Kuro posted, this doesn't affect the relative competitive result... keep at it!</p>\n  </blockquote>\n</blockquote>\n\n<p>Indeed it seems like the models are starting to outperform manual identification, which is great. The only thing I can think of to deal with the label discrepancies is to prefer the label with more images, in cases where the number of images is the same (which is the most often) I have no idea.</p>\n\n<p>I agree with you about treating the dataset, if it was mine I would probably farm it out to kaggle as well, it seems to be paying off. I'll let you know when I've had a chance to try your code, thanks for posting it! Good luck! </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 470606,
      "author_name": "Haider Alwasiti",
      "author_url": "",
      "post_date": "2019-02-13T08:38:55.357000",
      "content": "<p>Here are my picks of most likely wrong labelled images (more will post later):\nI have sorted top-1_incorrect file names in ascending, and the top images are the most wrong. This one has a distance of similarity of  0.002  which is quite similar, but has a different label:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/470606/11281/000000.00278_w_456646a_new_whale_ada29d516.jpg\" alt=\"enter image description here\"></p>\n\n<p>Also this one:\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/470606/11282/000000.03656_w_6dbfd7a_w_993320b_a541664bc.jpg\" alt=\"enter image description here\"></p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 472903,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-16T21:27:23.103000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 470547,
      "author_name": "interneuron",
      "author_url": "",
      "post_date": "2019-02-13T06:48:46.617000",
      "content": "<p>This is really excellent. Thank you!</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "470532": "I've created this code to explore the performance of the metric learning NN. The metric deep learning model that I am using is distance based which is based on @Iafoss kernel. The scores are interpreted as the distance between the image that we are testing with other images which  belong to known classes. The distance is a measure of the similarity of features in the embedding space. \n\n The function `image_matrix_draw` is going to associate two groups of images to each prediciton of your classifier's validation set predicitons. When you train a NN image classifier, you want to see how it is performing. One way to check its performance, and see why it failed on those wrong images, is to checkup the validation images that have been wrongly classified and to compare it with the images of the wrong class + with the images of the ground truth class. This notebook is doing just that.\n \n Let's take an example:\n \n**top-1 correct:** is plotting the 1st image as the top-1 predicted image (which is bounded by a rectangle). And then plots the most similar same class images in an ascending scores (from the most similar that the model think is, to the less  similar images).  The 1st column is the repeated images of the image that we are asking the model to predict. The columns 2-7 are the prediction images sorted. (The columns 5-7 are the ground truth class images sorted if the prediciton is incorrect which is not the case here)\n\n![top-1 correct](https://storage.googleapis.com/kaggle-forum-message-attachments/470532/11271/000000.37231_w_67a9841_w_67a9841_42a505fa7.jpg)\n\n\n\n**top-1 incorrect:** Here the incorrectly classified top-1 image. And just like before, it plots the most similar same class images in an ascending scores. This time there are ground truth images (maximum 12 images) plotted in the same way (from the most similar to the less similar)\n\n![top-1 incorrect](https://storage.googleapis.com/kaggle-forum-message-attachments/470532/11275/000032.75000_new_whale_w_42e0e40_da92fb1c1.jpg)\n\n\n**top-2 incorrect:**\n\n![top-2 incorrect](https://storage.googleapis.com/kaggle-forum-message-attachments/470532/11272/000051.75000_w_e99ed06_w_c3e88ae_4a5aa11c9.jpg)\n\n\n**top-k correct:** Here the plotted images are for the top-1 predicted image then the top-2, top-3 ...etc. Ground truth class plotted too in top-k correct plots, since we cannot see enough examples of top-1 images in the columns 2-4, so I thought it would be interesting to see correct class examples.\n\n![top-k correct](https://storage.googleapis.com/kaggle-forum-message-attachments/470532/11274/000004.87500_w_7c5b20d_w_7c5b20d_0bd65a44d.jpg)\n\n\n**top-k incorrect:** \n\n![top-k incorrect](https://storage.googleapis.com/kaggle-forum-message-attachments/470532/11273/000001.58105_w_3d1f606_new_whale_8b39ae55c.jpg)\n \n \n \nAll the above plots can be generated for the test set too. However, since there is no ground truth labels, it does not classify them into two correct/incorrect folders, and there are no ground truth images plotted.  It took ~10-20 minuted to generate each folder for the above categories on i9-9900k with 16 logical cores. \n \nThe function to explore the prediciton's of a classifier can be used for other metric or Siamese learning NN or even used for classification models with a few modifications in the arguments passed.  \n\nCode:\n[https://github.com/hwasiti/prediction_explorer][1]\n\nI think the notebook is self explanatory. If you have any question please feel free to ask. I worked ~1 week on this code. I think for any serious datascience project, there should be a proper way to know where is the weakness of the model. And how to remedy the issue. And this way you can get the intuition how to do that.\n\n**Sorting from most correct to most confused to most wrong predictions**\nThe file name generated with the score being the 1st term. Which allows you to sort the file names in the folder of the top-1 correct (for instance), in ascending and you will get the *most correct items* in the top and *most confused* (barely correct) images at the bottom. The same sorting is useful for top-1 incorrect folder in descending. The top files will be the most confused (just barely classified) wrong, and the bottom the *most wrong* predictions.\n\nPlease note that I have not applied distance cut (dcut) that I use to assign `new_whale` label when the distance of features in the embedding space is more than dcut parameter. (dcut is estimated by the best validation score checkup by brute force: check from dcut 1.0 to 30.0 to find which dcut is giving the best score). And the reason is, I did not want to distort the performance of the model by doing so. Dcut is a post-processing step, and I am afraid it will invalidate our insight on why and where the model is doing good and where it is failing.\n\n\n**P.S.:** in case you are using @Iafoss kernel, [here][2] is how to generate the pandas dataframes needed to generate these images. Otherwise you can modify the column names from your own model output as shown in the notebook in the github page.\n\n\n  [1]: https://github.com/hwasiti/prediction_explorer\n  [2]: https://www.kaggle.com/iafoss/similarity-densenet169-0-794lb-kernel-time-limit#470545",
    "470749": "Really good article !",
    "470687": "Great work!! I'll definitively going to use this, thank you! :)",
    "470624": "This one seems the same whale, but because the picture compared are black and white vs colored, the details are not quite clear for the model. Perhaps Martin was right turning all images into black and white could solve issues like this.\n![enter image description here][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/470624/11295/000001.07910_w_57db059_w_43fc74c_ce0de3213.jpg",
    "470622": "\n![enter image description here][1]\n\n\n![enter link description here][2]\n\n![enter image description here][3]\n\n![enter image description here][4]\n\n![enter image description here][5]\n\n![enter image description here][6]\n![enter image description here][7]\n\n![enter image description here][8]\n\n![enter image description here][9]\n\n![enter image description here][10]\n![enter image description here][11]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/470622/11284/000000.24695_w_f91389a_w_7d8c37c_30dd15cf2.jpg\n  [2]: https://storage.googleapis.com/kaggle-forum-message-attachments/470622/11286/000000.29419_w_9573686_new_whale_75f28d098.jpg\n  [3]: https://storage.googleapis.com/kaggle-forum-message-attachments/470622/11285/000000.28540_w_1c55cc3_new_whale_28d2542db.jpg\n  [4]: https://storage.googleapis.com/kaggle-forum-message-attachments/470622/11287/000000.42505_w_2f1488c_w_c9b4607_8eccdf3d4.jpg\n  [5]: https://storage.googleapis.com/kaggle-forum-message-attachments/470622/11289/000000.85205_new_whale_w_05bf34e_c6a6fe666.jpg\n  [6]: https://storage.googleapis.com/kaggle-forum-message-attachments/470622/11288/000000.51074_new_whale_w_07125ed_17bb74c11.jpg\n  [7]: https://storage.googleapis.com/kaggle-forum-message-attachments/470622/11290/000000.63672_new_whale_w_faad3f8_929ee57b1.jpg\n  [8]: https://storage.googleapis.com/kaggle-forum-message-attachments/470622/11291/000000.70312_w_f5d4627_w_2623921_71148b1fc.jpg\n  [9]: https://storage.googleapis.com/kaggle-forum-message-attachments/470622/11292/000000.63916_w_228c7ee_w_f94fc92_e55dfcd2f.jpg\n  [10]: https://storage.googleapis.com/kaggle-forum-message-attachments/470622/11293/000000.78174_w_e73d9a5_new_whale_83b15630f.jpg\n  [11]: https://storage.googleapis.com/kaggle-forum-message-attachments/470622/11294/000000.75439_w_8d4c9f7_w_4135cb8_51616faf5.jpg",
    "470610": "The top-1 seems exactly the same picture cropped. However the rest of the class seems different. So the mislabelling I think  `new_whale` (944f124ae.jpg)  should be labelled: `w_d4d46bf `\n\n![enter image description here][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/470610/11283/000000.21899_new_whale_w_d4d46bf_2b7fd4faf.jpg",
    "472850": "@hwasiti I can't upvote you enough for this. Thank you so much for this. Just like @iafoss i see great future (Kaggle Grandmaster) ahead for you. Thanks again.",
    "470606": "Here are my picks of most likely wrong labelled images (more will post later):\nI have sorted top-1_incorrect file names in ascending, and the top images are the most wrong. This one has a distance of similarity of  0.002  which is quite similar, but has a different label:\n\n![enter image description here][1]\n\nAlso this one:\n![enter image description here][2]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/470606/11281/000000.00278_w_456646a_new_whale_ada29d516.jpg\n  [2]: https://storage.googleapis.com/kaggle-forum-message-attachments/470606/11282/000000.03656_w_6dbfd7a_w_993320b_a541664bc.jpg",
    "472903": "",
    "470547": "This is really excellent. Thank you!"
  }
}