{
  "id": 79773,
  "title": "Approaches to metric learning, let's discuss!",
  "url": "/competitions/humpback-whale-identification/discussion/79773",
  "author_name": "",
  "post_date": "2019-02-07T11:58:28.744982500Z",
  "votes": 8,
  "comment_count": 16,
  "views": 0,
  "content": "<p>Hi guys, from what I've observed there are two ways of doing metric learning and I'd like to know your thoughts and experience on that!</p>\n\n<p>There first, which can be found in the majority of papers, is: </p>\n\n<ul>\n<li><strong>Train</strong>: You take the CNN output, apply global pooling (optional: add some fully connected) and use it as an embedding optimizing it using some kind of metric loss like contrastive or triplet loss. \n<strong>Query</strong>: You run the net on all images to create their embeddings, compute the distance (e.g. euclidian) between query and gallery embeddings and use a K-NN algorithm or something like that to match.</li>\n</ul>\n\n<p>The second is similar to Martin's solution:</p>\n\n<ul>\n<li><strong>Train</strong>: You take the CNN output and apply global pooling. Run this model for 2 images and then attach a head model, which is responsible to measure a similarity score between these images.  Optimize it similar to a classification problem using cross-entropy . <strong>Query</strong>: You run the head model for all pairs and find the best scores.</li>\n</ul>\n\n<p>So far, from what I've been observing is that most of the solutions (including mine) that use the first approach along with online hard example mining falls closely below 800 public LB. As for the second approach, we can observe from Martin's and Martin's based solutions that it is likely to go up to 900 public LB as far as I know. Would it be due to approach 1 vs. approach 2 or more due to the way examples are mined?</p>",
  "messages": [
    {
      "id": "467600",
      "postDate": "02/07/2019 11:58:28",
      "content": "<p>Hi guys, from what I've observed there are two ways of doing metric learning and I'd like to know your thoughts and experience on that!</p>\n\n<p>There first, which can be found in the majority of papers, is: </p>\n\n<ul>\n<li><strong>Train</strong>: You take the CNN output, apply global pooling (optional: add some fully connected) and use it as an embedding optimizing it using some kind of metric loss like contrastive or triplet loss. \n<strong>Query</strong>: You run the net on all images to create their embeddings, compute the distance (e.g. euclidian) between query and gallery embeddings and use a K-NN algorithm or something like that to match.</li>\n</ul>\n\n<p>The second is similar to Martin's solution:</p>\n\n<ul>\n<li><strong>Train</strong>: You take the CNN output and apply global pooling. Run this model for 2 images and then attach a head model, which is responsible to measure a similarity score between these images.  Optimize it similar to a classification problem using cross-entropy . <strong>Query</strong>: You run the head model for all pairs and find the best scores.</li>\n</ul>\n\n<p>So far, from what I've been observing is that most of the solutions (including mine) that use the first approach along with online hard example mining falls closely below 800 public LB. As for the second approach, we can observe from Martin's and Martin's based solutions that it is likely to go up to 900 public LB as far as I know. Would it be due to approach 1 vs. approach 2 or more due to the way examples are mined?</p>",
      "rawMarkdown": "Hi guys, from what I've observed there are two ways of doing metric learning and I'd like to know your thoughts and experience on that!\n\nThere first, which can be found in the majority of papers, is: \n\n*  **Train**: You take the CNN output, apply global pooling (optional: add some fully connected) and use it as an embedding optimizing it using some kind of metric loss like contrastive or triplet loss. \n**Query**: You run the net on all images to create their embeddings, compute the distance (e.g. euclidian) between query and gallery embeddings and use a K-NN algorithm or something like that to match.\n\nThe second is similar to Martin's solution:\n\n*   **Train**: You take the CNN output and apply global pooling. Run this model for 2 images and then attach a head model, which is responsible to measure a similarity score between these images.  Optimize it similar to a classification problem using cross-entropy . **Query**: You run the head model for all pairs and find the best scores.\n\nSo far, from what I've been observing is that most of the solutions (including mine) that use the first approach along with online hard example mining falls closely below 800 public LB. As for the second approach, we can observe from Martin's and Martin's based solutions that it is likely to go up to 900 public LB as far as I know. Would it be due to approach 1 vs. approach 2 or more due to the way examples are mined?",
      "votes": null
    },
    {
      "id": "468010",
      "postDate": "02/08/2019 06:05:54",
      "content": "<p>You can check <a href=\"https://www.kaggle.com/c/humpback-whale-identification/discussion/79086\">this discussion</a>. I started with the first approach and when I added metric learning (like Martin does) on top of it I got ~0.05 boost. However, training the model stucks at ~0.8 LB. Another thing, I tried to run just a model of the second type (the same metric as Martin's one but ResNeXt50 backbone) and got worse results. So, right now I think that LAP is the key that defines if your model can get 0.8 or 0.9.</p>",
      "rawMarkdown": "You can check [this discussion][1]. I started with the first approach and when I added metric learning (like Martin does) on top of it I got ~0.05 boost. However, training the model stucks at ~0.8 LB. Another thing, I tried to run just a model of the second type (the same metric as Martin's one but ResNeXt50 backbone) and got worse results. So, right now I think that LAP is the key that defines if your model can get 0.8 or 0.9.\n\n\n  [1]: https://www.kaggle.com/c/humpback-whale-identification/discussion/79086",
      "votes": null
    },
    {
      "id": "468749",
      "postDate": "02/09/2019 15:28:24",
      "content": "<p>In martin's kernel, he said \"Linear sum assignment algorithm is used to find the most difficult matching.\" So do you mean the way to select training samples is the key to get 0.9 ?\n<a href=\"https://www.kaggle.com/martinpiotte/whale-recognition-model-with-score-0-78563\">https://www.kaggle.com/martinpiotte/whale-recognition-model-with-score-0-78563</a></p>",
      "rawMarkdown": "In martin's kernel, he said \"Linear sum assignment algorithm is used to find the most difficult matching.\" So do you mean the way to select training samples is the key to get 0.9 ?\nhttps://www.kaggle.com/martinpiotte/whale-recognition-model-with-score-0-78563",
      "votes": null
    },
    {
      "id": "468769",
      "postDate": "02/09/2019 16:04:21",
      "content": "<p>Yep, and it looks that it sd be LAP, simple selection of one of the most difficult examples for each image doesn't work well, as far as I saw. LAP allows to select the most difficult setup for entire training.</p>",
      "rawMarkdown": "Yep, and it looks that it sd be LAP, simple selection of one of the most difficult examples for each image doesn't work well, as far as I saw. LAP allows to select the most difficult setup for entire training.",
      "votes": null
    },
    {
      "id": "468784",
      "postDate": "02/09/2019 16:53:26",
      "content": "<p>Thank you for your answer. I think LAP is related to Negative Mining which is described in metric-learning papers.</p>",
      "rawMarkdown": "Thank you for your answer. I think LAP is related to Negative Mining which is described in metric-learning papers.",
      "votes": null
    },
    {
      "id": "468789",
      "postDate": "02/09/2019 17:08:56",
      "content": "<p>Yes, but I didn't see that a model using just regular negative mining can get ~0.9, the result is ~0.8. Earlier I thought that batch all/hard loss + selection of the most difficult negative examples would work the same as LAP (I tried to exclude this slow step from training). But I ended up with ~0.8 LB. When I tried martins metric and no head layer, the model ended up even with lower score. So the only remaining thing is LAP, which I'm checking right now.</p>",
      "rawMarkdown": "Yes, but I didn't see that a model using just regular negative mining can get ~0.9, the result is ~0.8. Earlier I thought that batch all/hard loss + selection of the most difficult negative examples would work the same as LAP (I tried to exclude this slow step from training). But I ended up with ~0.8 LB. When I tried martins metric and no head layer, the model ended up even with lower score. So the only remaining thing is LAP, which I'm checking right now.",
      "votes": null
    },
    {
      "id": "468872",
      "postDate": "02/09/2019 22:39:28",
      "content": "<p>Same here. I've tried online hard example mining and it worked to some extent but not well enough as Martin's solution. Then, I've tried to mine examples offline in simpler ways like higher loss, examples with more mistakes, etc.. It also didn't appear to work either. However, I haven't explored it a lot and there is a guy in the discussion forum that said it worked out for him. So maybe there is a mistake in my implementation.</p>\n\n<p>Another thing to note is that I'm directly mining the hardest example whereas Martin's solution do it progressively!</p>\n\n<p>Finally, do not mine the hardest positive examples!! They are too dissimilar, which end up breaking up the whole net.</p>",
      "rawMarkdown": "Same here. I've tried online hard example mining and it worked to some extent but not well enough as Martin's solution. Then, I've tried to mine examples offline in simpler ways like higher loss, examples with more mistakes, etc.. It also didn't appear to work either. However, I haven't explored it a lot and there is a guy in the discussion forum that said it worked out for him. So maybe there is a mistake in my implementation.\n\nAnother thing to note is that I'm directly mining the hardest example whereas Martin's solution do it progressively!\n\nFinally, do not mine the hardest positive examples!! They are too dissimilar, which end up breaking up the whole net.",
      "votes": null
    },
    {
      "id": "468890",
      "postDate": "02/10/2019 00:34:00",
      "content": "<p>check this:\n\"Sampling Matters in Deep Embedding Learning\"\n<a href=\"https://arxiv.org/pdf/1706.07567.pdf\">https://arxiv.org/pdf/1706.07567.pdf</a></p>\n\n<p>in training triple loss, we use what we called \"semi hard mining\".\n<a href=\"https://omoindrot.github.io/triplet-loss\">https://omoindrot.github.io/triplet-loss</a>\n<a href=\"https://www.tensorflow.org/api_docs/python/tf/contrib/losses/metric_learning/triplet_semihard_loss\">https://www.tensorflow.org/api_docs/python/tf/contrib/losses/metric_learning/triplet_semihard_loss</a></p>\n\n<p>it is like training object detection in faster-rcnn and ssd, in which we use: pos = sample with box overlap &gt; 0.5 and neg = sample with box overlap &lt;0.3. </p>\n\n<p>those with overlap between 0.3 and 0.5 are considered as down care, because they affect results .</p>\n\n<p>in theory, one can devise a new \"focal loss\" for metric learning to replace hrad mining</p>",
      "rawMarkdown": "check this:\n\"Sampling Matters in Deep Embedding Learning\"\nhttps://arxiv.org/pdf/1706.07567.pdf\n\nin training triple loss, we use what we called \"semi hard mining\".\nhttps://omoindrot.github.io/triplet-loss\nhttps://www.tensorflow.org/api_docs/python/tf/contrib/losses/metric_learning/triplet_semihard_loss\n\nit is like training object detection in faster-rcnn and ssd, in which we use: pos = sample with box overlap &gt; 0.5 and neg = sample with box overlap &lt;0.3. \n\nthose with overlap between 0.3 and 0.5 are considered as down care, because they affect results .\n\nin theory, one can devise a new \"focal loss\" for metric learning to replace hrad mining",
      "votes": null
    },
    {
      "id": "470497",
      "postDate": "02/13/2019 04:37:08",
      "content": "<p>I've tried mining hard positives beside hard negatives, and it did give very slight better score. My method is modified from Iafoss kernel method in mining negatives. Randomly picking from the hardest n=3 in my case. (Iafoss used n=64 for the hardest negative mining.) But it does not worth it. The issue is that, this dataset does not have a lot of positives. So whether I assign n=3 or do not assign any positive mining procedure, it will be the same image picked up anyway in most whales. They simply have only one other positive image or just a few.</p>\n\n<p>It seems that all of us are missing the gradual difficulty increase with every few epochs. I wonder whether that plays a role. You said picking up the most difficult positive images broke up the net. So it seems logical that mining negatives also, maybe some of them is too difficult, especially if it is introduced abruptly after easy training. I hope I get enough time to checkup Martin's method of  gradual increase in difficulty in negative mining with every few epochs.</p>",
      "rawMarkdown": "I've tried mining hard positives beside hard negatives, and it did give very slight better score. My method is modified from Iafoss kernel method in mining negatives. Randomly picking from the hardest n=3 in my case. (Iafoss used n=64 for the hardest negative mining.) But it does not worth it. The issue is that, this dataset does not have a lot of positives. So whether I assign n=3 or do not assign any positive mining procedure, it will be the same image picked up anyway in most whales. They simply have only one other positive image or just a few.\n\nIt seems that all of us are missing the gradual difficulty increase with every few epochs. I wonder whether that plays a role. You said picking up the most difficult positive images broke up the net. So it seems logical that mining negatives also, maybe some of them is too difficult, especially if it is introduced abruptly after easy training. I hope I get enough time to checkup Martin's method of  gradual increase in difficulty in negative mining with every few epochs.",
      "votes": null
    },
    {
      "id": "470668",
      "postDate": "02/13/2019 10:52:05",
      "content": "<p>Please, tell me if I misunderstood you. You mined the positives for online batch mining, right? If that's the case, yes it helps because the batch is just a small \"sampling space\" from the whole dataset. What I meant in my previous comment is that I've tried to mine hard positives offline in all dataset, which than contains some really dissimilar positive examples. That lowered my score even for models that were already pre-trained.</p>\n\n<p>I agree with you, probably the tricky is to do it gradually.</p>",
      "rawMarkdown": "Please, tell me if I misunderstood you. You mined the positives for online batch mining, right? If that's the case, yes it helps because the batch is just a small \"sampling space\" from the whole dataset. What I meant in my previous comment is that I've tried to mine hard positives offline in all dataset, which than contains some really dissimilar positive examples. That lowered my score even for models that were already pre-trained.\n\nI agree with you, probably the tricky is to do it gradually.",
      "votes": null
    },
    {
      "id": "470901",
      "postDate": "02/13/2019 18:27:20",
      "content": "<p>I mined hard positives offline in all dataset, following the same way @Iafoss mined the hard negatives.</p>\n\n<p>But I would not choose the most difficult. For each image, I am picking the 3 most difficult positives in the whole dataset (except the split out validation set), and I am taking one of these 3 by a random pick and not the hardest. Some of the hardest are not good at all. They are either distorted or something that if I had the energy and time, I would exclude it manually from the dataset.</p>",
      "rawMarkdown": "I mined hard positives offline in all dataset, following the same way @Iafoss mined the hard negatives.\n\nBut I would not choose the most difficult. For each image, I am picking the 3 most difficult positives in the whole dataset (except the split out validation set), and I am taking one of these 3 by a random pick and not the hardest. Some of the hardest are not good at all. They are either distorted or something that if I had the energy and time, I would exclude it manually from the dataset.",
      "votes": null
    },
    {
      "id": "470936",
      "postDate": "02/13/2019 19:45:17",
      "content": "<p>Hey <a href=\"/arc144\">@arc144</a> and @Iafoss \nWith my initial <a href=\"https://www.kaggle.com/c/humpback-whale-identification/discussion/80379\">exploration of the most wrong images</a> (sorted according to the most confident validation predictions, but they are actually wrong), it seems there are a LOT of mislabeled images. I posted there ~15 images, but I got tired posting images, because they are keep getting more and more of wrong labeled training data. I wonder how can we get a model with LB 0.90 with such wrong training labels..</p>\n\n<p>So, perhaps there is nothing wrong with your models. They are stuck in LB 0.80 because the training data is not as clean as Martin's kernel.</p>\n\n<p>Perhaps the under-estimated part of Martin's kernel that we did not use yet:\n<strong>Duplicate image identification</strong>\nCompute phash for each image\nGroup together images with equivalent phash, and replace by string format of phash (faster and more readable)\n....etc.</p>\n\n<p>In the same way I have tried to explore the predictions of the top-1 test set predictions. Of course I would not be able to isolate the wrong predictions into a separate folder like what I did in the validation set. But I tried to investigate the (distance 15 and more) images (which is somehow the most confused predictions because my dcut = 17 for regarding everything farther as new whale). And it does not seem to me that it has &gt; 20% error  even for those &gt;15 distance, which is the most confused predictions. For 0.800 MAP5, what I understand is that the model has one 5th of its top-1 prediction wrong, or maybe more because there is some contribution from the top-2,3,4,5</p>\n\n<p>What I speculate is that the model is confused between 2 classes because those 2 classes has been used for training and contains mixed up labels..</p>\n\n<p>I intended to try to replicate the same cleaning and augmentations used in Martin's kernel soon, but super busy with other stuff.. If you are interested to try it out, please share how it goes with you..</p>",
      "rawMarkdown": "Hey @arc144 and @Iafoss \nWith my initial [exploration of the most wrong images][1] (sorted according to the most confident validation predictions, but they are actually wrong), it seems there are a LOT of mislabeled images. I posted there ~15 images, but I got tired posting images, because they are keep getting more and more of wrong labeled training data. I wonder how can we get a model with LB 0.90 with such wrong training labels..\n\nSo, perhaps there is nothing wrong with your models. They are stuck in LB 0.80 because the training data is not as clean as Martin's kernel.\n\nPerhaps the under-estimated part of Martin's kernel that we did not use yet:\n**Duplicate image identification**\nCompute phash for each image\nGroup together images with equivalent phash, and replace by string format of phash (faster and more readable)\n....etc.\n\nIn the same way I have tried to explore the predictions of the top-1 test set predictions. Of course I would not be able to isolate the wrong predictions into a separate folder like what I did in the validation set. But I tried to investigate the (distance 15 and more) images (which is somehow the most confused predictions because my dcut = 17 for regarding everything farther as new whale). And it does not seem to me that it has &gt; 20% error  even for those &gt;15 distance, which is the most confused predictions. For 0.800 MAP5, what I understand is that the model has one 5th of its top-1 prediction wrong, or maybe more because there is some contribution from the top-2,3,4,5\n\nWhat I speculate is that the model is confused between 2 classes because those 2 classes has been used for training and contains mixed up labels..\n\n\nI intended to try to replicate the same cleaning and augmentations used in Martin's kernel soon, but super busy with other stuff.. If you are interested to try it out, please share how it goes with you..\n\n\n  [1]: https://www.kaggle.com/c/humpback-whale-identification/discussion/80379",
      "votes": null
    },
    {
      "id": "471010",
      "postDate": "02/13/2019 22:37:40",
      "content": "<p>It really does makes sense.. But I've somewhere that the playground competition really had a problem with duplicates . However, here its seems there are only a few. But anyway I'll take a look into that when I can and I'll update here! :)</p>",
      "rawMarkdown": "It really does makes sense.. But I've somewhere that the playground competition really had a problem with duplicates . However, here its seems there are only a few. But anyway I'll take a look into that when I can and I'll update here! :)",
      "votes": null
    },
    {
      "id": "471417",
      "postDate": "02/14/2019 12:32:51",
      "content": "<p>Hi <a href=\"/hwasiti\">@hwasiti</a>, I've analyzed the data cleaning step and that not seems to be the case. There are only 6 problematic images removed by the phash step. And if we take the entire Martin's preprocess step we end up with 13623 image ids whereas if we simply get the train df and remove new_whales and whale ids with a single picture we get 13624. Furthermore, Martin is not directly correcting or removing wrong labels.</p>",
      "rawMarkdown": "Hi @hwasiti, I've analyzed the data cleaning step and that not seems to be the case. There are only 6 problematic images removed by the phash step. And if we take the entire Martin's preprocess step we end up with 13623 image ids whereas if we simply get the train df and remove new_whales and whale ids with a single picture we get 13624. Furthermore, Martin is not directly correcting or removing wrong labels.",
      "votes": null
    },
    {
      "id": "472440",
      "postDate": "02/15/2019 23:26:08",
      "content": "<p>hmm... Thanks for sharing..</p>\n\n<p>It seems LAP is the key, like what @Iafoss expected:\n<a href=\"https://www.kaggle.com/iafoss/similarity-densenet169-0-800lb-kernel-time-limit/comments#471836\">https://www.kaggle.com/iafoss/similarity-densenet169-0-800lb-kernel-time-limit/comments#471836</a></p>\n\n<p><a href=\"/atom1231\">@atom1231</a>  :</p>\n\n<blockquote>\n  <p>(5) augmentation : V11 with some change got 0.82~ 0.83 (4) hard\n  negative mining is important  I did some test with Martin's solution \n  1. Martin's solution with default LAP package (from lap import lapjv) =&gt; 0.88\n  2.Martin's solution with simplify LAP package (from lap import lapjv) =&gt; 0.86\n  3.Martin's solution with another LAP package (from lapjv import lapjv)=&gt; 0.85x\n  4.Martin's solution with my own LAP =&gt; 0.84\n  5.Martin's solution with my own tricky hard negative mining =&gt; 0.82</p>\n</blockquote>",
      "rawMarkdown": "hmm... Thanks for sharing..\n\nIt seems LAP is the key, like what @Iafoss expected:\nhttps://www.kaggle.com/iafoss/similarity-densenet169-0-800lb-kernel-time-limit/comments#471836\n\n@atom1231  :\n&gt; (5) augmentation : V11 with some change got 0.82~ 0.83 (4) hard\n&gt; negative mining is important  I did some test with Martin's solution \n&gt; 1. Martin's solution with default LAP package (from lap import lapjv) =&gt; 0.88\n&gt; 2.Martin's solution with simplify LAP package (from lap import lapjv) =&gt; 0.86\n&gt; 3.Martin's solution with another LAP package (from lapjv import lapjv)=&gt; 0.85x\n&gt; 4.Martin's solution with my own LAP =&gt; 0.84\n&gt; 5.Martin's solution with my own tricky hard negative mining =&gt; 0.82",
      "votes": null
    },
    {
      "id": "472458",
      "postDate": "02/16/2019 00:55:51",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing",
      "votes": null
    },
    {
      "id": "2571028",
      "postDate": "12/22/2023 19:43:29",
      "content": "<p>By the way, here is an open-source related to the topic of metric learning - OpenMetricLearning (<a href=\"https://github.com/OML-Team/open-metric-learning\" target=\"_blank\">https://github.com/OML-Team/open-metric-learning</a>) which may be useful for people who will find this thread later </p>",
      "rawMarkdown": "By the way, here is an open-source related to the topic of metric learning - OpenMetricLearning (https://github.com/OML-Team/open-metric-learning) which may be useful for people who will find this thread later",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2571028,
      "author_name": "aglasis",
      "author_url": "",
      "post_date": "12/22/2023 19:43:29",
      "content": "<p>By the way, here is an open-source related to the topic of metric learning - OpenMetricLearning (<a href=\"https://github.com/OML-Team/open-metric-learning\" target=\"_blank\">https://github.com/OML-Team/open-metric-learning</a>) which may be useful for people who will find this thread later </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 468010,
      "author_name": "iafoss",
      "author_url": "",
      "post_date": "02/08/2019 06:05:54",
      "content": "<p>You can check <a href=\"https://www.kaggle.com/c/humpback-whale-identification/discussion/79086\">this discussion</a>. I started with the first approach and when I added metric learning (like Martin does) on top of it I got ~0.05 boost. However, training the model stucks at ~0.8 LB. Another thing, I tried to run just a model of the second type (the same metric as Martin's one but ResNeXt50 backbone) and got worse results. So, right now I think that LAP is the key that defines if your model can get 0.8 or 0.9.</p>",
      "votes": null,
      "replies": [
        {
          "id": 468749,
          "author_name": "toshik",
          "author_url": "",
          "post_date": "02/09/2019 15:28:24",
          "content": "<p>In martin's kernel, he said \"Linear sum assignment algorithm is used to find the most difficult matching.\" So do you mean the way to select training samples is the key to get 0.9 ?\n<a href=\"https://www.kaggle.com/martinpiotte/whale-recognition-model-with-score-0-78563\">https://www.kaggle.com/martinpiotte/whale-recognition-model-with-score-0-78563</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 468769,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "02/09/2019 16:04:21",
          "content": "<p>Yep, and it looks that it sd be LAP, simple selection of one of the most difficult examples for each image doesn't work well, as far as I saw. LAP allows to select the most difficult setup for entire training.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 468784,
          "author_name": "toshik",
          "author_url": "",
          "post_date": "02/09/2019 16:53:26",
          "content": "<p>Thank you for your answer. I think LAP is related to Negative Mining which is described in metric-learning papers.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 468789,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "02/09/2019 17:08:56",
          "content": "<p>Yes, but I didn't see that a model using just regular negative mining can get ~0.9, the result is ~0.8. Earlier I thought that batch all/hard loss + selection of the most difficult negative examples would work the same as LAP (I tried to exclude this slow step from training). But I ended up with ~0.8 LB. When I tried martins metric and no head layer, the model ended up even with lower score. So the only remaining thing is LAP, which I'm checking right now.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 468872,
          "author_name": "arc144",
          "author_url": "",
          "post_date": "02/09/2019 22:39:28",
          "content": "<p>Same here. I've tried online hard example mining and it worked to some extent but not well enough as Martin's solution. Then, I've tried to mine examples offline in simpler ways like higher loss, examples with more mistakes, etc.. It also didn't appear to work either. However, I haven't explored it a lot and there is a guy in the discussion forum that said it worked out for him. So maybe there is a mistake in my implementation.</p>\n\n<p>Another thing to note is that I'm directly mining the hardest example whereas Martin's solution do it progressively!</p>\n\n<p>Finally, do not mine the hardest positive examples!! They are too dissimilar, which end up breaking up the whole net.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 470497,
          "author_name": "hwasiti",
          "author_url": "",
          "post_date": "02/13/2019 04:37:08",
          "content": "<p>I've tried mining hard positives beside hard negatives, and it did give very slight better score. My method is modified from Iafoss kernel method in mining negatives. Randomly picking from the hardest n=3 in my case. (Iafoss used n=64 for the hardest negative mining.) But it does not worth it. The issue is that, this dataset does not have a lot of positives. So whether I assign n=3 or do not assign any positive mining procedure, it will be the same image picked up anyway in most whales. They simply have only one other positive image or just a few.</p>\n\n<p>It seems that all of us are missing the gradual difficulty increase with every few epochs. I wonder whether that plays a role. You said picking up the most difficult positive images broke up the net. So it seems logical that mining negatives also, maybe some of them is too difficult, especially if it is introduced abruptly after easy training. I hope I get enough time to checkup Martin's method of  gradual increase in difficulty in negative mining with every few epochs.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 470668,
          "author_name": "arc144",
          "author_url": "",
          "post_date": "02/13/2019 10:52:05",
          "content": "<p>Please, tell me if I misunderstood you. You mined the positives for online batch mining, right? If that's the case, yes it helps because the batch is just a small \"sampling space\" from the whole dataset. What I meant in my previous comment is that I've tried to mine hard positives offline in all dataset, which than contains some really dissimilar positive examples. That lowered my score even for models that were already pre-trained.</p>\n\n<p>I agree with you, probably the tricky is to do it gradually.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 470901,
          "author_name": "hwasiti",
          "author_url": "",
          "post_date": "02/13/2019 18:27:20",
          "content": "<p>I mined hard positives offline in all dataset, following the same way @Iafoss mined the hard negatives.</p>\n\n<p>But I would not choose the most difficult. For each image, I am picking the 3 most difficult positives in the whole dataset (except the split out validation set), and I am taking one of these 3 by a random pick and not the hardest. Some of the hardest are not good at all. They are either distorted or something that if I had the energy and time, I would exclude it manually from the dataset.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 468890,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "02/10/2019 00:34:00",
      "content": "<p>check this:\n\"Sampling Matters in Deep Embedding Learning\"\n<a href=\"https://arxiv.org/pdf/1706.07567.pdf\">https://arxiv.org/pdf/1706.07567.pdf</a></p>\n\n<p>in training triple loss, we use what we called \"semi hard mining\".\n<a href=\"https://omoindrot.github.io/triplet-loss\">https://omoindrot.github.io/triplet-loss</a>\n<a href=\"https://www.tensorflow.org/api_docs/python/tf/contrib/losses/metric_learning/triplet_semihard_loss\">https://www.tensorflow.org/api_docs/python/tf/contrib/losses/metric_learning/triplet_semihard_loss</a></p>\n\n<p>it is like training object detection in faster-rcnn and ssd, in which we use: pos = sample with box overlap &gt; 0.5 and neg = sample with box overlap &lt;0.3. </p>\n\n<p>those with overlap between 0.3 and 0.5 are considered as down care, because they affect results .</p>\n\n<p>in theory, one can devise a new \"focal loss\" for metric learning to replace hrad mining</p>",
      "votes": null,
      "replies": [
        {
          "id": 472458,
          "author_name": "siddharth5mn",
          "author_url": "",
          "post_date": "02/16/2019 00:55:51",
          "content": "<p>Thanks for sharing</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 470936,
      "author_name": "hwasiti",
      "author_url": "",
      "post_date": "02/13/2019 19:45:17",
      "content": "<p>Hey <a href=\"/arc144\">@arc144</a> and @Iafoss \nWith my initial <a href=\"https://www.kaggle.com/c/humpback-whale-identification/discussion/80379\">exploration of the most wrong images</a> (sorted according to the most confident validation predictions, but they are actually wrong), it seems there are a LOT of mislabeled images. I posted there ~15 images, but I got tired posting images, because they are keep getting more and more of wrong labeled training data. I wonder how can we get a model with LB 0.90 with such wrong training labels..</p>\n\n<p>So, perhaps there is nothing wrong with your models. They are stuck in LB 0.80 because the training data is not as clean as Martin's kernel.</p>\n\n<p>Perhaps the under-estimated part of Martin's kernel that we did not use yet:\n<strong>Duplicate image identification</strong>\nCompute phash for each image\nGroup together images with equivalent phash, and replace by string format of phash (faster and more readable)\n....etc.</p>\n\n<p>In the same way I have tried to explore the predictions of the top-1 test set predictions. Of course I would not be able to isolate the wrong predictions into a separate folder like what I did in the validation set. But I tried to investigate the (distance 15 and more) images (which is somehow the most confused predictions because my dcut = 17 for regarding everything farther as new whale). And it does not seem to me that it has &gt; 20% error  even for those &gt;15 distance, which is the most confused predictions. For 0.800 MAP5, what I understand is that the model has one 5th of its top-1 prediction wrong, or maybe more because there is some contribution from the top-2,3,4,5</p>\n\n<p>What I speculate is that the model is confused between 2 classes because those 2 classes has been used for training and contains mixed up labels..</p>\n\n<p>I intended to try to replicate the same cleaning and augmentations used in Martin's kernel soon, but super busy with other stuff.. If you are interested to try it out, please share how it goes with you..</p>",
      "votes": null,
      "replies": [
        {
          "id": 471010,
          "author_name": "arc144",
          "author_url": "",
          "post_date": "02/13/2019 22:37:40",
          "content": "<p>It really does makes sense.. But I've somewhere that the playground competition really had a problem with duplicates . However, here its seems there are only a few. But anyway I'll take a look into that when I can and I'll update here! :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 471417,
          "author_name": "arc144",
          "author_url": "",
          "post_date": "02/14/2019 12:32:51",
          "content": "<p>Hi <a href=\"/hwasiti\">@hwasiti</a>, I've analyzed the data cleaning step and that not seems to be the case. There are only 6 problematic images removed by the phash step. And if we take the entire Martin's preprocess step we end up with 13623 image ids whereas if we simply get the train df and remove new_whales and whale ids with a single picture we get 13624. Furthermore, Martin is not directly correcting or removing wrong labels.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 472440,
          "author_name": "hwasiti",
          "author_url": "",
          "post_date": "02/15/2019 23:26:08",
          "content": "<p>hmm... Thanks for sharing..</p>\n\n<p>It seems LAP is the key, like what @Iafoss expected:\n<a href=\"https://www.kaggle.com/iafoss/similarity-densenet169-0-800lb-kernel-time-limit/comments#471836\">https://www.kaggle.com/iafoss/similarity-densenet169-0-800lb-kernel-time-limit/comments#471836</a></p>\n\n<p><a href=\"/atom1231\">@atom1231</a>  :</p>\n\n<blockquote>\n  <p>(5) augmentation : V11 with some change got 0.82~ 0.83 (4) hard\n  negative mining is important  I did some test with Martin's solution \n  1. Martin's solution with default LAP package (from lap import lapjv) =&gt; 0.88\n  2.Martin's solution with simplify LAP package (from lap import lapjv) =&gt; 0.86\n  3.Martin's solution with another LAP package (from lapjv import lapjv)=&gt; 0.85x\n  4.Martin's solution with my own LAP =&gt; 0.84\n  5.Martin's solution with my own tricky hard negative mining =&gt; 0.82</p>\n</blockquote>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "467600": "Hi guys, from what I've observed there are two ways of doing metric learning and I'd like to know your thoughts and experience on that!\n\nThere first, which can be found in the majority of papers, is: \n\n*  **Train**: You take the CNN output, apply global pooling (optional: add some fully connected) and use it as an embedding optimizing it using some kind of metric loss like contrastive or triplet loss. \n**Query**: You run the net on all images to create their embeddings, compute the distance (e.g. euclidian) between query and gallery embeddings and use a K-NN algorithm or something like that to match.\n\nThe second is similar to Martin's solution:\n\n*   **Train**: You take the CNN output and apply global pooling. Run this model for 2 images and then attach a head model, which is responsible to measure a similarity score between these images.  Optimize it similar to a classification problem using cross-entropy . **Query**: You run the head model for all pairs and find the best scores.\n\nSo far, from what I've been observing is that most of the solutions (including mine) that use the first approach along with online hard example mining falls closely below 800 public LB. As for the second approach, we can observe from Martin's and Martin's based solutions that it is likely to go up to 900 public LB as far as I know. Would it be due to approach 1 vs. approach 2 or more due to the way examples are mined?",
    "468010": "You can check [this discussion][1]. I started with the first approach and when I added metric learning (like Martin does) on top of it I got ~0.05 boost. However, training the model stucks at ~0.8 LB. Another thing, I tried to run just a model of the second type (the same metric as Martin's one but ResNeXt50 backbone) and got worse results. So, right now I think that LAP is the key that defines if your model can get 0.8 or 0.9.\n\n\n  [1]: https://www.kaggle.com/c/humpback-whale-identification/discussion/79086",
    "468749": "In martin's kernel, he said \"Linear sum assignment algorithm is used to find the most difficult matching.\" So do you mean the way to select training samples is the key to get 0.9 ?\nhttps://www.kaggle.com/martinpiotte/whale-recognition-model-with-score-0-78563",
    "468769": "Yep, and it looks that it sd be LAP, simple selection of one of the most difficult examples for each image doesn't work well, as far as I saw. LAP allows to select the most difficult setup for entire training.",
    "468784": "Thank you for your answer. I think LAP is related to Negative Mining which is described in metric-learning papers.",
    "468789": "Yes, but I didn't see that a model using just regular negative mining can get ~0.9, the result is ~0.8. Earlier I thought that batch all/hard loss + selection of the most difficult negative examples would work the same as LAP (I tried to exclude this slow step from training). But I ended up with ~0.8 LB. When I tried martins metric and no head layer, the model ended up even with lower score. So the only remaining thing is LAP, which I'm checking right now.",
    "468872": "Same here. I've tried online hard example mining and it worked to some extent but not well enough as Martin's solution. Then, I've tried to mine examples offline in simpler ways like higher loss, examples with more mistakes, etc.. It also didn't appear to work either. However, I haven't explored it a lot and there is a guy in the discussion forum that said it worked out for him. So maybe there is a mistake in my implementation.\n\nAnother thing to note is that I'm directly mining the hardest example whereas Martin's solution do it progressively!\n\nFinally, do not mine the hardest positive examples!! They are too dissimilar, which end up breaking up the whole net.",
    "468890": "check this:\n\"Sampling Matters in Deep Embedding Learning\"\nhttps://arxiv.org/pdf/1706.07567.pdf\n\nin training triple loss, we use what we called \"semi hard mining\".\nhttps://omoindrot.github.io/triplet-loss\nhttps://www.tensorflow.org/api_docs/python/tf/contrib/losses/metric_learning/triplet_semihard_loss\n\nit is like training object detection in faster-rcnn and ssd, in which we use: pos = sample with box overlap &gt; 0.5 and neg = sample with box overlap &lt;0.3. \n\nthose with overlap between 0.3 and 0.5 are considered as down care, because they affect results .\n\nin theory, one can devise a new \"focal loss\" for metric learning to replace hrad mining",
    "470497": "I've tried mining hard positives beside hard negatives, and it did give very slight better score. My method is modified from Iafoss kernel method in mining negatives. Randomly picking from the hardest n=3 in my case. (Iafoss used n=64 for the hardest negative mining.) But it does not worth it. The issue is that, this dataset does not have a lot of positives. So whether I assign n=3 or do not assign any positive mining procedure, it will be the same image picked up anyway in most whales. They simply have only one other positive image or just a few.\n\nIt seems that all of us are missing the gradual difficulty increase with every few epochs. I wonder whether that plays a role. You said picking up the most difficult positive images broke up the net. So it seems logical that mining negatives also, maybe some of them is too difficult, especially if it is introduced abruptly after easy training. I hope I get enough time to checkup Martin's method of  gradual increase in difficulty in negative mining with every few epochs.",
    "470668": "Please, tell me if I misunderstood you. You mined the positives for online batch mining, right? If that's the case, yes it helps because the batch is just a small \"sampling space\" from the whole dataset. What I meant in my previous comment is that I've tried to mine hard positives offline in all dataset, which than contains some really dissimilar positive examples. That lowered my score even for models that were already pre-trained.\n\nI agree with you, probably the tricky is to do it gradually.",
    "470901": "I mined hard positives offline in all dataset, following the same way @Iafoss mined the hard negatives.\n\nBut I would not choose the most difficult. For each image, I am picking the 3 most difficult positives in the whole dataset (except the split out validation set), and I am taking one of these 3 by a random pick and not the hardest. Some of the hardest are not good at all. They are either distorted or something that if I had the energy and time, I would exclude it manually from the dataset.",
    "470936": "Hey @arc144 and @Iafoss \nWith my initial [exploration of the most wrong images][1] (sorted according to the most confident validation predictions, but they are actually wrong), it seems there are a LOT of mislabeled images. I posted there ~15 images, but I got tired posting images, because they are keep getting more and more of wrong labeled training data. I wonder how can we get a model with LB 0.90 with such wrong training labels..\n\nSo, perhaps there is nothing wrong with your models. They are stuck in LB 0.80 because the training data is not as clean as Martin's kernel.\n\nPerhaps the under-estimated part of Martin's kernel that we did not use yet:\n**Duplicate image identification**\nCompute phash for each image\nGroup together images with equivalent phash, and replace by string format of phash (faster and more readable)\n....etc.\n\nIn the same way I have tried to explore the predictions of the top-1 test set predictions. Of course I would not be able to isolate the wrong predictions into a separate folder like what I did in the validation set. But I tried to investigate the (distance 15 and more) images (which is somehow the most confused predictions because my dcut = 17 for regarding everything farther as new whale). And it does not seem to me that it has &gt; 20% error  even for those &gt;15 distance, which is the most confused predictions. For 0.800 MAP5, what I understand is that the model has one 5th of its top-1 prediction wrong, or maybe more because there is some contribution from the top-2,3,4,5\n\nWhat I speculate is that the model is confused between 2 classes because those 2 classes has been used for training and contains mixed up labels..\n\n\nI intended to try to replicate the same cleaning and augmentations used in Martin's kernel soon, but super busy with other stuff.. If you are interested to try it out, please share how it goes with you..\n\n\n  [1]: https://www.kaggle.com/c/humpback-whale-identification/discussion/80379",
    "471010": "It really does makes sense.. But I've somewhere that the playground competition really had a problem with duplicates . However, here its seems there are only a few. But anyway I'll take a look into that when I can and I'll update here! :)",
    "471417": "Hi @hwasiti, I've analyzed the data cleaning step and that not seems to be the case. There are only 6 problematic images removed by the phash step. And if we take the entire Martin's preprocess step we end up with 13623 image ids whereas if we simply get the train df and remove new_whales and whale ids with a single picture we get 13624. Furthermore, Martin is not directly correcting or removing wrong labels.",
    "472440": "hmm... Thanks for sharing..\n\nIt seems LAP is the key, like what @Iafoss expected:\nhttps://www.kaggle.com/iafoss/similarity-densenet169-0-800lb-kernel-time-limit/comments#471836\n\n@atom1231  :\n&gt; (5) augmentation : V11 with some change got 0.82~ 0.83 (4) hard\n&gt; negative mining is important  I did some test with Martin's solution \n&gt; 1. Martin's solution with default LAP package (from lap import lapjv) =&gt; 0.88\n&gt; 2.Martin's solution with simplify LAP package (from lap import lapjv) =&gt; 0.86\n&gt; 3.Martin's solution with another LAP package (from lapjv import lapjv)=&gt; 0.85x\n&gt; 4.Martin's solution with my own LAP =&gt; 0.84\n&gt; 5.Martin's solution with my own tricky hard negative mining =&gt; 0.82",
    "472458": "Thanks for sharing",
    "2571028": "By the way, here is an open-source related to the topic of metric learning - OpenMetricLearning (https://github.com/OML-Team/open-metric-learning) which may be useful for people who will find this thread later"
  },
  "source": "meta"
}