{
  "id": 194543,
  "title": "ArcFace and Image Embedding",
  "url": "/competitions/landmark-recognition-2020/discussion/194543",
  "author_name": "Long Luu",
  "post_date": "2020-11-02T07:12:27.235000",
  "votes": 2,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I am learning the ArcFace paper since I saw most winning image classification/retrieval solutions used it. I used it on MNIST, CIFAR10, CIFAR100 and I noticed the accuracy was very low compared to plain CNN, so I have some thoughts about it:</p>\n<ol>\n<li>Its purpose is to make the Image Embedding using ArcFace loss (instead of Softmax), so the training accuracy is not important. The embedding is the global descriptors.</li>\n<li>After training, it gets input as image and outputs as its embedding vector. We then use the output vector to measure the cosine similarities of the embedding matrix, get top k results and that is our final probability prediction. (same thing as Word Embedding). </li>\n</ol>\n<p>If I understand correctly, I have more questions:</p>\n<ol>\n<li>If we use ArcFace loss, but the model has to use (Sparse) Categorical Crossentropy, how do we know when the model is at its best? We use these 2 losses or we use another metric (accuracy for example)? If we use Accuracy as metric, is it reliable?</li>\n<li>Following the above, we optimize ArcFace loss but we still use Crossentropy, so which of the two should we consider as \"model is learning\"?</li>\n<li>If the two losses and Accuracy metric are not reliable, what do we use to evaluate the model? Is it Cosine similarity?</li>\n</ol>",
  "messages": [
    {
      "id": 1066883,
      "postDate": "2020-11-02T07:12:27.237Z",
      "content": "<p>I am learning the ArcFace paper since I saw most winning image classification/retrieval solutions used it. I used it on MNIST, CIFAR10, CIFAR100 and I noticed the accuracy was very low compared to plain CNN, so I have some thoughts about it:</p>\n<ol>\n<li>Its purpose is to make the Image Embedding using ArcFace loss (instead of Softmax), so the training accuracy is not important. The embedding is the global descriptors.</li>\n<li>After training, it gets input as image and outputs as its embedding vector. We then use the output vector to measure the cosine similarities of the embedding matrix, get top k results and that is our final probability prediction. (same thing as Word Embedding). </li>\n</ol>\n<p>If I understand correctly, I have more questions:</p>\n<ol>\n<li>If we use ArcFace loss, but the model has to use (Sparse) Categorical Crossentropy, how do we know when the model is at its best? We use these 2 losses or we use another metric (accuracy for example)? If we use Accuracy as metric, is it reliable?</li>\n<li>Following the above, we optimize ArcFace loss but we still use Crossentropy, so which of the two should we consider as \"model is learning\"?</li>\n<li>If the two losses and Accuracy metric are not reliable, what do we use to evaluate the model? Is it Cosine similarity?</li>\n</ol>",
      "rawMarkdown": "I am learning the ArcFace paper since I saw most winning image classification/retrieval solutions used it. I used it on MNIST, CIFAR10, CIFAR100 and I noticed the accuracy was very low compared to plain CNN, so I have some thoughts about it:\n1. Its purpose is to make the Image Embedding using ArcFace loss (instead of Softmax), so the training accuracy is not important. The embedding is the global descriptors.\n2. After training, it gets input as image and outputs as its embedding vector. We then use the output vector to measure the cosine similarities of the embedding matrix, get top k results and that is our final probability prediction. (same thing as Word Embedding). \n\nIf I understand correctly, I have more questions:\n1. If we use ArcFace loss, but the model has to use (Sparse) Categorical Crossentropy, how do we know when the model is at its best? We use these 2 losses or we use another metric (accuracy for example)? If we use Accuracy as metric, is it reliable?\n2. Following the above, we optimize ArcFace loss but we still use Crossentropy, so which of the two should we consider as \"model is learning\"?\n3. If the two losses and Accuracy metric are not reliable, what do we use to evaluate the model? Is it Cosine similarity?\n\n",
      "votes": 2
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1066883": "I am learning the ArcFace paper since I saw most winning image classification/retrieval solutions used it. I used it on MNIST, CIFAR10, CIFAR100 and I noticed the accuracy was very low compared to plain CNN, so I have some thoughts about it:\n1. Its purpose is to make the Image Embedding using ArcFace loss (instead of Softmax), so the training accuracy is not important. The embedding is the global descriptors.\n2. After training, it gets input as image and outputs as its embedding vector. We then use the output vector to measure the cosine similarities of the embedding matrix, get top k results and that is our final probability prediction. (same thing as Word Embedding). \n\nIf I understand correctly, I have more questions:\n1. If we use ArcFace loss, but the model has to use (Sparse) Categorical Crossentropy, how do we know when the model is at its best? We use these 2 losses or we use another metric (accuracy for example)? If we use Accuracy as metric, is it reliable?\n2. Following the above, we optimize ArcFace loss but we still use Crossentropy, so which of the two should we consider as \"model is learning\"?\n3. If the two losses and Accuracy metric are not reliable, what do we use to evaluate the model? Is it Cosine similarity?\n\n"
  }
}