{
  "id": 167515,
  "title": "Anyone Using Metric Learning ? ",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/167515",
  "author_name": "",
  "post_date": "2020-07-16T20:01:38.512324Z",
  "votes": 3,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I'm trying to use CosFaceLoss + Focal Loss using Densenet121 with a custom head? I was wondering if anyone has tried out metric learning for this task ?? </p>\n\n<p>I'm also experimenting with attention mechanisms, they seem to be interesting but I'm yet to complete a full feldge model.</p>\n\n<p>So far I'm failing to properly leverage the power of metric learning. But I will keep trying since I'm learning so much reading papers. So far I have read around 8 papers which includes papers on ProtoNets , Siamese Nets, Relation Networks,etc and an amazing paper on \"Hyperbolic Image Embeddings\". </p>\n\n<p>I'm linking the papers \nCosFace :<a href=\"https://arxiv.org/abs/1801.09414\">CosFace</a>\nProtoNets : <a href=\"https://arxiv.org/abs/1703.05175\">ProtoNets</a>\nRelation networks : <a href=\"https://arxiv.org/abs/1711.06025\">Relation Networks</a>\nHyperbolic Image Embeddings : <a href=\"https://arxiv.org/abs/1904.02239\">Hyperbolic Embeddings</a></p>",
  "messages": [
    {
      "id": "932200",
      "postDate": "07/16/2020 20:01:38",
      "content": "<p>I'm trying to use CosFaceLoss + Focal Loss using Densenet121 with a custom head? I was wondering if anyone has tried out metric learning for this task ?? </p>\n\n<p>I'm also experimenting with attention mechanisms, they seem to be interesting but I'm yet to complete a full feldge model.</p>\n\n<p>So far I'm failing to properly leverage the power of metric learning. But I will keep trying since I'm learning so much reading papers. So far I have read around 8 papers which includes papers on ProtoNets , Siamese Nets, Relation Networks,etc and an amazing paper on \"Hyperbolic Image Embeddings\". </p>\n\n<p>I'm linking the papers \nCosFace :<a href=\"https://arxiv.org/abs/1801.09414\">CosFace</a>\nProtoNets : <a href=\"https://arxiv.org/abs/1703.05175\">ProtoNets</a>\nRelation networks : <a href=\"https://arxiv.org/abs/1711.06025\">Relation Networks</a>\nHyperbolic Image Embeddings : <a href=\"https://arxiv.org/abs/1904.02239\">Hyperbolic Embeddings</a></p>",
      "rawMarkdown": "I'm trying to use CosFaceLoss + Focal Loss using Densenet121 with a custom head? I was wondering if anyone has tried out metric learning for this task ?? \n\nI'm also experimenting with attention mechanisms, they seem to be interesting but I'm yet to complete a full feldge model.\n\nSo far I'm failing to properly leverage the power of metric learning. But I will keep trying since I'm learning so much reading papers. So far I have read around 8 papers which includes papers on ProtoNets , Siamese Nets, Relation Networks,etc and an amazing paper on \"Hyperbolic Image Embeddings\". \n\nI'm linking the papers \nCosFace :[CosFace]( https://arxiv.org/abs/1801.09414)\nProtoNets : [ProtoNets](https://arxiv.org/abs/1703.05175)\nRelation networks : [Relation Networks](https://arxiv.org/abs/1711.06025)\nHyperbolic Image Embeddings : [Hyperbolic Embeddings](https://arxiv.org/abs/1904.02239)",
      "votes": null
    },
    {
      "id": "932252",
      "postDate": "07/16/2020 21:44:38",
      "content": "<p>I experimented with this idea a little bit. For Proto Nets, Relation Nets, Matching Nets, etc. a few-shot setup is sort of required, but some ideas from those papers can be used. For example you can create two randomly initialized vectors and use these as \"prototypes\" for each class. Classification can be performed by softmaxing the negative squared euclidean distance between the input embedding and the two prototype embeddings. This can then be trained via CE or Focal Loss. Optionally, the prototype embeddings can be learned, and extra loss functions can be added such as triplet loss, or some sort of cosine loss or whatever else you can think of. This setup should encourage the model to reduce intra-class variance, and increase inter-class variance. I also found these types of models train better with L2 normalized embeddings. I observed a slight increase in performance over traditional fully connected classification, but nothing substantial, and adding extra loss functions didn't seem to change much so I moved in a different direction.</p>\n\n<p>There are likely other ways to make a metric-learning setup work well. I have experience with few-shot learning but not traditional metric learning, and it seems like most previous approaches are for the task of retrieval and not classification. I think it's a pretty cool idea though.</p>",
      "rawMarkdown": "I experimented with this idea a little bit. For Proto Nets, Relation Nets, Matching Nets, etc. a few-shot setup is sort of required, but some ideas from those papers can be used. For example you can create two randomly initialized vectors and use these as \"prototypes\" for each class. Classification can be performed by softmaxing the negative squared euclidean distance between the input embedding and the two prototype embeddings. This can then be trained via CE or Focal Loss. Optionally, the prototype embeddings can be learned, and extra loss functions can be added such as triplet loss, or some sort of cosine loss or whatever else you can think of. This setup should encourage the model to reduce intra-class variance, and increase inter-class variance. I also found these types of models train better with L2 normalized embeddings. I observed a slight increase in performance over traditional fully connected classification, but nothing substantial, and adding extra loss functions didn't seem to change much so I moved in a different direction.\n\nThere are likely other ways to make a metric-learning setup work well. I have experience with few-shot learning but not traditional metric learning, and it seems like most previous approaches are for the task of retrieval and not classification. I think it's a pretty cool idea though.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 932252,
      "author_name": "chriscareaga",
      "author_url": "",
      "post_date": "07/16/2020 21:44:38",
      "content": "<p>I experimented with this idea a little bit. For Proto Nets, Relation Nets, Matching Nets, etc. a few-shot setup is sort of required, but some ideas from those papers can be used. For example you can create two randomly initialized vectors and use these as \"prototypes\" for each class. Classification can be performed by softmaxing the negative squared euclidean distance between the input embedding and the two prototype embeddings. This can then be trained via CE or Focal Loss. Optionally, the prototype embeddings can be learned, and extra loss functions can be added such as triplet loss, or some sort of cosine loss or whatever else you can think of. This setup should encourage the model to reduce intra-class variance, and increase inter-class variance. I also found these types of models train better with L2 normalized embeddings. I observed a slight increase in performance over traditional fully connected classification, but nothing substantial, and adding extra loss functions didn't seem to change much so I moved in a different direction.</p>\n\n<p>There are likely other ways to make a metric-learning setup work well. I have experience with few-shot learning but not traditional metric learning, and it seems like most previous approaches are for the task of retrieval and not classification. I think it's a pretty cool idea though.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "932200": "I'm trying to use CosFaceLoss + Focal Loss using Densenet121 with a custom head? I was wondering if anyone has tried out metric learning for this task ?? \n\nI'm also experimenting with attention mechanisms, they seem to be interesting but I'm yet to complete a full feldge model.\n\nSo far I'm failing to properly leverage the power of metric learning. But I will keep trying since I'm learning so much reading papers. So far I have read around 8 papers which includes papers on ProtoNets , Siamese Nets, Relation Networks,etc and an amazing paper on \"Hyperbolic Image Embeddings\". \n\nI'm linking the papers \nCosFace :[CosFace]( https://arxiv.org/abs/1801.09414)\nProtoNets : [ProtoNets](https://arxiv.org/abs/1703.05175)\nRelation networks : [Relation Networks](https://arxiv.org/abs/1711.06025)\nHyperbolic Image Embeddings : [Hyperbolic Embeddings](https://arxiv.org/abs/1904.02239)",
    "932252": "I experimented with this idea a little bit. For Proto Nets, Relation Nets, Matching Nets, etc. a few-shot setup is sort of required, but some ideas from those papers can be used. For example you can create two randomly initialized vectors and use these as \"prototypes\" for each class. Classification can be performed by softmaxing the negative squared euclidean distance between the input embedding and the two prototype embeddings. This can then be trained via CE or Focal Loss. Optionally, the prototype embeddings can be learned, and extra loss functions can be added such as triplet loss, or some sort of cosine loss or whatever else you can think of. This setup should encourage the model to reduce intra-class variance, and increase inter-class variance. I also found these types of models train better with L2 normalized embeddings. I observed a slight increase in performance over traditional fully connected classification, but nothing substantial, and adding extra loss functions didn't seem to change much so I moved in a different direction.\n\nThere are likely other ways to make a metric-learning setup work well. I have experience with few-shot learning but not traditional metric learning, and it seems like most previous approaches are for the task of retrieval and not classification. I think it's a pretty cool idea though."
  },
  "source": "meta"
}