{
  "id": 78280,
  "title": "Metrics Learning Understand - Help Needed",
  "url": "/competitions/humpback-whale-identification/discussion/78280",
  "author_name": "",
  "post_date": "2019-01-21T22:05:34.379349500Z",
  "votes": 4,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi kagglers!</p>\n\n<p>I am leraning a lot with this comeptitions and I read a lot the forums and some papers... But I have one BIG doubt and is... How Metrics Learning runs? How can I use it in my solution :/ <br>\nI saw that first solution use it and that bestfitting will destroy us with them :D  -&gt; <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/78109\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/78109</a></p>\n\n<p>My best model is the siamese (<a href=\"https://www.kaggle.com/suicaokhoailang/martin-piotte-s-siamese-baseline\">https://www.kaggle.com/suicaokhoailang/martin-piotte-s-siamese-baseline</a>) which use some convolutional on the head mode... Is this what you refer as metric learning? these convolutionals?</p>\n\n<p>Thanks</p>",
  "messages": [
    {
      "id": "459504",
      "postDate": "01/21/2019 22:05:34",
      "content": "<p>Hi kagglers!</p>\n\n<p>I am leraning a lot with this comeptitions and I read a lot the forums and some papers... But I have one BIG doubt and is... How Metrics Learning runs? How can I use it in my solution :/ <br>\nI saw that first solution use it and that bestfitting will destroy us with them :D  -&gt; <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/78109\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/78109</a></p>\n\n<p>My best model is the siamese (<a href=\"https://www.kaggle.com/suicaokhoailang/martin-piotte-s-siamese-baseline\">https://www.kaggle.com/suicaokhoailang/martin-piotte-s-siamese-baseline</a>) which use some convolutional on the head mode... Is this what you refer as metric learning? these convolutionals?</p>\n\n<p>Thanks</p>",
      "rawMarkdown": "Hi kagglers!\n\nI am leraning a lot with this comeptitions and I read a lot the forums and some papers... But I have one BIG doubt and is... How Metrics Learning runs? How can I use it in my solution :/  \nI saw that first solution use it and that bestfitting will destroy us with them :D  -&gt; https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/78109\n\nMy best model is the siamese (https://www.kaggle.com/suicaokhoailang/martin-piotte-s-siamese-baseline) which use some convolutional on the head mode... Is this what you refer as metric learning? these convolutionals?\n\nThanks",
      "votes": null
    },
    {
      "id": "459520",
      "postDate": "01/21/2019 22:43:13",
      "content": "<p>Technically speaking, anything, which gives you descriptors to compare with some metric, e.g. Euclidean, is metric learning. This is in contrast to classification, which gives you model of the <em>class</em>, not the model of how similar or dissimilar inputs are.  \"Metric\" in <a href=\"https://www.kaggle.com/suicaokhoailang/martin-piotte-s-siamese-baseline\">https://www.kaggle.com/suicaokhoailang/martin-piotte-s-siamese-baseline</a> is the decision network. </p>\n\n<p>Things, which matter in metric learning are the same as in classification: architecture, data pre/post-processing, loss function. One important difference is sampling: that is why hard mining in playground competition is so important, since you are comparing tuples or triples. </p>",
      "rawMarkdown": "Technically speaking, anything, which gives you descriptors to compare with some metric, e.g. Euclidean, is metric learning. This is in contrast to classification, which gives you model of the _class_, not the model of how similar or dissimilar inputs are.  \"Metric\" in https://www.kaggle.com/suicaokhoailang/martin-piotte-s-siamese-baseline is the decision network. \n\nThings, which matter in metric learning are the same as in classification: architecture, data pre/post-processing, loss function. One important difference is sampling: that is why hard mining in playground competition is so important, since you are comparing tuples or triples.",
      "votes": null
    },
    {
      "id": "459522",
      "postDate": "01/21/2019 22:58:14",
      "content": "<p>So... If I change the head model and take the embedding produced by the branch model and learn the branch model using a loss function based on euclidean distance for example, I am using metric learning? Mayeb its not so clear for me on Martins approach because he use binary crossentropy.</p>\n\n<p>When you say 'Things, which matter in metric learning are the same as in classification: architecture, data pre/post-processing, loss function', you refer that \"you\" (to say someone) use a separated model from the embedder and train these one with the embeddings reducing for example the binary crossentropy... Like in Martin solution with branch (embedder) - head (¿metric learning?) models?</p>\n\n<p>I'm sorry but I want to understand it and I've read some other paper and it's not clear to me... I have changed the head model by a simple CNN for example and dont works. If people refer to metric learning to set a euclidean distance (for example) to head model and learn then branch model, as I understand you are not learning the metric, you are learning the embedding...</p>\n\n<p>Thanks for your answer. I don´t want to be annoying :(</p>",
      "rawMarkdown": "So... If I change the head model and take the embedding produced by the branch model and learn the branch model using a loss function based on euclidean distance for example, I am using metric learning? Mayeb its not so clear for me on Martins approach because he use binary crossentropy.\n\nWhen you say 'Things, which matter in metric learning are the same as in classification: architecture, data pre/post-processing, loss function', you refer that \"you\" (to say someone) use a separated model from the embedder and train these one with the embeddings reducing for example the binary crossentropy... Like in Martin solution with branch (embedder) - head (¿metric learning?) models?\n\nI'm sorry but I want to understand it and I've read some other paper and it's not clear to me... I have changed the head model by a simple CNN for example and dont works. If people refer to metric learning to set a euclidean distance (for example) to head model and learn then branch model, as I understand you are not learning the metric, you are learning the embedding...\n\nThanks for your answer. I don´t want to be annoying :(",
      "votes": null
    },
    {
      "id": "459525",
      "postDate": "01/21/2019 23:16:04",
      "content": "<p>You are right, terms are confusing. </p>\n\n<blockquote>\n  <p>So… If I change the head model and take the embedding produced by the branch model and learn the branch model using a loss function based on euclidean distance for example, I am using metric learning? </p>\n</blockquote>\n\n<p>Yes.  Embedding learning is the same as metric learning. Cross-entropy is valid choice of loss function, same as contrastive or triplet on L2 distance. The reason, why people, at least in production, prefer Euclidean metric over specific network as in Martin solution, is because for inference you can:\n1) extract descriptors, O(n) complexity\n2) Use approximate NN (as in faiss library) for finding the nearest neighbour + checking distance to it. Complexity varies, but usually O(N * log (N)). \nAnd if you use network for comparison, your complexity is O(N^2), which is not scalable. </p>\n\n<p>However, in this competition, unlike Google Landmarks, n^2 is not prohibitive and might bring one a bit extra accuracy.  </p>",
      "rawMarkdown": "You are right, terms are confusing. \n\n&gt;So… If I change the head model and take the embedding produced by the branch model and learn the branch model using a loss function based on euclidean distance for example, I am using metric learning? \n\nYes.  Embedding learning is the same as metric learning. Cross-entropy is valid choice of loss function, same as contrastive or triplet on L2 distance. The reason, why people, at least in production, prefer Euclidean metric over specific network as in Martin solution, is because for inference you can:\n1) extract descriptors, O(n) complexity\n2) Use approximate NN (as in faiss library) for finding the nearest neighbour + checking distance to it. Complexity varies, but usually O(N * log (N)). \nAnd if you use network for comparison, your complexity is O(N^2), which is not scalable. \n\nHowever, in this competition, unlike Google Landmarks, n^2 is not prohibitive and might bring one a bit extra accuracy.",
      "votes": null
    },
    {
      "id": "465437",
      "postDate": "02/03/2019 05:05:56",
      "content": "<p><code>\nHowever, in this competition, unlike Google Landmarks, n^2 is not prohibitive and might bring one a bit extra accuracy. \n</code></p>\n\n<p>How does it improve your score ? According to lafoss, his metric leaning method improve LB score 0.606-&gt;0.655.\n<a href=\"https://www.kaggle.com/iafoss/similarity-resnext50-0-740-lb-kernel-time-limit\">https://www.kaggle.com/iafoss/similarity-resnext50-0-740-lb-kernel-time-limit</a></p>",
      "rawMarkdown": "```\nHowever, in this competition, unlike Google Landmarks, n^2 is not prohibitive and might bring one a bit extra accuracy. \n```\n\nHow does it improve your score ? According to lafoss, his metric leaning method improve LB score 0.606-&gt;0.655.\nhttps://www.kaggle.com/iafoss/similarity-resnext50-0-740-lb-kernel-time-limit",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 459520,
      "author_name": "oldufo",
      "author_url": "",
      "post_date": "01/21/2019 22:43:13",
      "content": "<p>Technically speaking, anything, which gives you descriptors to compare with some metric, e.g. Euclidean, is metric learning. This is in contrast to classification, which gives you model of the <em>class</em>, not the model of how similar or dissimilar inputs are.  \"Metric\" in <a href=\"https://www.kaggle.com/suicaokhoailang/martin-piotte-s-siamese-baseline\">https://www.kaggle.com/suicaokhoailang/martin-piotte-s-siamese-baseline</a> is the decision network. </p>\n\n<p>Things, which matter in metric learning are the same as in classification: architecture, data pre/post-processing, loss function. One important difference is sampling: that is why hard mining in playground competition is so important, since you are comparing tuples or triples. </p>",
      "votes": null,
      "replies": [
        {
          "id": 459522,
          "author_name": "maparla",
          "author_url": "",
          "post_date": "01/21/2019 22:58:14",
          "content": "<p>So... If I change the head model and take the embedding produced by the branch model and learn the branch model using a loss function based on euclidean distance for example, I am using metric learning? Mayeb its not so clear for me on Martins approach because he use binary crossentropy.</p>\n\n<p>When you say 'Things, which matter in metric learning are the same as in classification: architecture, data pre/post-processing, loss function', you refer that \"you\" (to say someone) use a separated model from the embedder and train these one with the embeddings reducing for example the binary crossentropy... Like in Martin solution with branch (embedder) - head (¿metric learning?) models?</p>\n\n<p>I'm sorry but I want to understand it and I've read some other paper and it's not clear to me... I have changed the head model by a simple CNN for example and dont works. If people refer to metric learning to set a euclidean distance (for example) to head model and learn then branch model, as I understand you are not learning the metric, you are learning the embedding...</p>\n\n<p>Thanks for your answer. I don´t want to be annoying :(</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 459525,
          "author_name": "oldufo",
          "author_url": "",
          "post_date": "01/21/2019 23:16:04",
          "content": "<p>You are right, terms are confusing. </p>\n\n<blockquote>\n  <p>So… If I change the head model and take the embedding produced by the branch model and learn the branch model using a loss function based on euclidean distance for example, I am using metric learning? </p>\n</blockquote>\n\n<p>Yes.  Embedding learning is the same as metric learning. Cross-entropy is valid choice of loss function, same as contrastive or triplet on L2 distance. The reason, why people, at least in production, prefer Euclidean metric over specific network as in Martin solution, is because for inference you can:\n1) extract descriptors, O(n) complexity\n2) Use approximate NN (as in faiss library) for finding the nearest neighbour + checking distance to it. Complexity varies, but usually O(N * log (N)). \nAnd if you use network for comparison, your complexity is O(N^2), which is not scalable. </p>\n\n<p>However, in this competition, unlike Google Landmarks, n^2 is not prohibitive and might bring one a bit extra accuracy.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 465437,
          "author_name": "toshik",
          "author_url": "",
          "post_date": "02/03/2019 05:05:56",
          "content": "<p><code>\nHowever, in this competition, unlike Google Landmarks, n^2 is not prohibitive and might bring one a bit extra accuracy. \n</code></p>\n\n<p>How does it improve your score ? According to lafoss, his metric leaning method improve LB score 0.606-&gt;0.655.\n<a href=\"https://www.kaggle.com/iafoss/similarity-resnext50-0-740-lb-kernel-time-limit\">https://www.kaggle.com/iafoss/similarity-resnext50-0-740-lb-kernel-time-limit</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "459504": "Hi kagglers!\n\nI am leraning a lot with this comeptitions and I read a lot the forums and some papers... But I have one BIG doubt and is... How Metrics Learning runs? How can I use it in my solution :/  \nI saw that first solution use it and that bestfitting will destroy us with them :D  -&gt; https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/78109\n\nMy best model is the siamese (https://www.kaggle.com/suicaokhoailang/martin-piotte-s-siamese-baseline) which use some convolutional on the head mode... Is this what you refer as metric learning? these convolutionals?\n\nThanks",
    "459520": "Technically speaking, anything, which gives you descriptors to compare with some metric, e.g. Euclidean, is metric learning. This is in contrast to classification, which gives you model of the _class_, not the model of how similar or dissimilar inputs are.  \"Metric\" in https://www.kaggle.com/suicaokhoailang/martin-piotte-s-siamese-baseline is the decision network. \n\nThings, which matter in metric learning are the same as in classification: architecture, data pre/post-processing, loss function. One important difference is sampling: that is why hard mining in playground competition is so important, since you are comparing tuples or triples.",
    "459522": "So... If I change the head model and take the embedding produced by the branch model and learn the branch model using a loss function based on euclidean distance for example, I am using metric learning? Mayeb its not so clear for me on Martins approach because he use binary crossentropy.\n\nWhen you say 'Things, which matter in metric learning are the same as in classification: architecture, data pre/post-processing, loss function', you refer that \"you\" (to say someone) use a separated model from the embedder and train these one with the embeddings reducing for example the binary crossentropy... Like in Martin solution with branch (embedder) - head (¿metric learning?) models?\n\nI'm sorry but I want to understand it and I've read some other paper and it's not clear to me... I have changed the head model by a simple CNN for example and dont works. If people refer to metric learning to set a euclidean distance (for example) to head model and learn then branch model, as I understand you are not learning the metric, you are learning the embedding...\n\nThanks for your answer. I don´t want to be annoying :(",
    "459525": "You are right, terms are confusing. \n\n&gt;So… If I change the head model and take the embedding produced by the branch model and learn the branch model using a loss function based on euclidean distance for example, I am using metric learning? \n\nYes.  Embedding learning is the same as metric learning. Cross-entropy is valid choice of loss function, same as contrastive or triplet on L2 distance. The reason, why people, at least in production, prefer Euclidean metric over specific network as in Martin solution, is because for inference you can:\n1) extract descriptors, O(n) complexity\n2) Use approximate NN (as in faiss library) for finding the nearest neighbour + checking distance to it. Complexity varies, but usually O(N * log (N)). \nAnd if you use network for comparison, your complexity is O(N^2), which is not scalable. \n\nHowever, in this competition, unlike Google Landmarks, n^2 is not prohibitive and might bring one a bit extra accuracy.",
    "465437": "```\nHowever, in this competition, unlike Google Landmarks, n^2 is not prohibitive and might bring one a bit extra accuracy. \n```\n\nHow does it improve your score ? According to lafoss, his metric leaning method improve LB score 0.606-&gt;0.655.\nhttps://www.kaggle.com/iafoss/similarity-resnext50-0-740-lb-kernel-time-limit"
  },
  "source": "meta"
}