{
  "id": 268227,
  "title": "Question about a metric learning (ArcFace)",
  "url": "/competitions/landmark-retrieval-2021/discussion/268227",
  "author_name": "",
  "post_date": "2021-08-26T14:20:09.270870600Z",
  "votes": 4,
  "comment_count": 2,
  "views": 0,
  "content": "<p>This is my first really big size image and metric learning competition.</p>\n<p>I have been modifying my model from just softmax classification to arcface metric learning classification for building my baseline model.</p>\n<p>Since there are too many classes in this competition, no margined model's score is too low( Validation accuracy : 0.889 but a leaderboard score is just 0.18x)</p>\n<p>So, I changed my model from softmax to arcface by referring to a kaggle code in Shopee competition which was held three months ago.</p>\n<p>When I added a margin(0.3), the model's accuracy score is similar to the previous model of softmax but the validation score dropped a lot(0.88 -&gt; 0.22)</p>\n<p>Is it natural? If an arcface version's score is lower than the softmax version that I mentioned above, can get a better score in submission with an arcface version?<br>\n(Since training takes a lot of time, I can't afford to do many experiments with my devices.</p>\n<p>=========================================================<br>\nMy Model:<br>\nEffcientNet3, GeM Polling(p=3), ArcFace(margin=0.5, scaler = 50)</p>\n<p>Optimizer:<br>\nCosineRAdam or SGD</p>\n<p>Batch size: 16 * 8 = 128</p>\n<p>Image size: 448<br>\nAugmentation : HorizontalFlip(p=0.5), RandomResizedCrop(scale(0.8,1.0))</p>",
  "messages": [
    {
      "id": "1491646",
      "postDate": "08/26/2021 14:20:09",
      "content": "<p>This is my first really big size image and metric learning competition.</p>\n<p>I have been modifying my model from just softmax classification to arcface metric learning classification for building my baseline model.</p>\n<p>Since there are too many classes in this competition, no margined model's score is too low( Validation accuracy : 0.889 but a leaderboard score is just 0.18x)</p>\n<p>So, I changed my model from softmax to arcface by referring to a kaggle code in Shopee competition which was held three months ago.</p>\n<p>When I added a margin(0.3), the model's accuracy score is similar to the previous model of softmax but the validation score dropped a lot(0.88 -&gt; 0.22)</p>\n<p>Is it natural? If an arcface version's score is lower than the softmax version that I mentioned above, can get a better score in submission with an arcface version?<br>\n(Since training takes a lot of time, I can't afford to do many experiments with my devices.</p>\n<p>=========================================================<br>\nMy Model:<br>\nEffcientNet3, GeM Polling(p=3), ArcFace(margin=0.5, scaler = 50)</p>\n<p>Optimizer:<br>\nCosineRAdam or SGD</p>\n<p>Batch size: 16 * 8 = 128</p>\n<p>Image size: 448<br>\nAugmentation : HorizontalFlip(p=0.5), RandomResizedCrop(scale(0.8,1.0))</p>",
      "rawMarkdown": "This is my first really big size image and metric learning competition.\n\nI have been modifying my model from just softmax classification to arcface metric learning classification for building my baseline model.\n\nSince there are too many classes in this competition, no margined model's score is too low( Validation accuracy : 0.889 but a leaderboard score is just 0.18x)\n\nSo, I changed my model from softmax to arcface by referring to a kaggle code in Shopee competition which was held three months ago.\n\nWhen I added a margin(0.3), the model's accuracy score is similar to the previous model of softmax but the validation score dropped a lot(0.88 -> 0.22)\n\nIs it natural? If an arcface version's score is lower than the softmax version that I mentioned above, can get a better score in submission with an arcface version?\n(Since training takes a lot of time, I can't afford to do many experiments with my devices.\n\n\n=========================================================\nMy Model:\nEffcientNet3, GeM Polling(p=3), ArcFace(margin=0.5, scaler = 50)\n\nOptimizer:\nCosineRAdam or SGD\n\nBatch size: 16 * 8 = 128\n\nImage size: 448\nAugmentation : HorizontalFlip(p=0.5), RandomResizedCrop(scale(0.8,1.0))",
      "votes": null
    },
    {
      "id": "1493634",
      "postDate": "08/28/2021 05:10:40",
      "content": "<p>I think using softmax, your model did not achieve the same class compactness and generalisation and hence over fitted on the marginal samples.</p>\n<p>With arcface there is a  \\sqrt{2m} decision boundary in cosine space, so maybe most of the validation samples are falling into that region and getting misclassified. While a larger decision boundary is good for class compactness and intra-class genralisation, setting it very high may lead to your samples start behaving as if the model is being trained on an open-set scenario.</p>",
      "rawMarkdown": "I think using softmax, your model did not achieve the same class compactness and generalisation and hence over fitted on the marginal samples.\n\nWith arcface there is a  \\sqrt{2m} decision boundary in cosine space, so maybe most of the validation samples are falling into that region and getting misclassified. While a larger decision boundary is good for class compactness and intra-class genralisation, setting it very high may lead to your samples start behaving as if the model is being trained on an open-set scenario.",
      "votes": null
    },
    {
      "id": "1493778",
      "postDate": "08/28/2021 06:41:40",
      "content": "<p>I did many experiments, setting scale = 50 and margin =0.3 seems like a good choice. I am expecting a good result when I make a submission.</p>",
      "rawMarkdown": "I did many experiments, setting scale = 50 and margin =0.3 seems like a good choice. I am expecting a good result when I make a submission.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1493634,
      "author_name": "prajwalsood",
      "author_url": "",
      "post_date": "08/28/2021 05:10:40",
      "content": "<p>I think using softmax, your model did not achieve the same class compactness and generalisation and hence over fitted on the marginal samples.</p>\n<p>With arcface there is a  \\sqrt{2m} decision boundary in cosine space, so maybe most of the validation samples are falling into that region and getting misclassified. While a larger decision boundary is good for class compactness and intra-class genralisation, setting it very high may lead to your samples start behaving as if the model is being trained on an open-set scenario.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1493778,
          "author_name": "deepkim",
          "author_url": "",
          "post_date": "08/28/2021 06:41:40",
          "content": "<p>I did many experiments, setting scale = 50 and margin =0.3 seems like a good choice. I am expecting a good result when I make a submission.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1491646": "This is my first really big size image and metric learning competition.\n\nI have been modifying my model from just softmax classification to arcface metric learning classification for building my baseline model.\n\nSince there are too many classes in this competition, no margined model's score is too low( Validation accuracy : 0.889 but a leaderboard score is just 0.18x)\n\nSo, I changed my model from softmax to arcface by referring to a kaggle code in Shopee competition which was held three months ago.\n\nWhen I added a margin(0.3), the model's accuracy score is similar to the previous model of softmax but the validation score dropped a lot(0.88 -> 0.22)\n\nIs it natural? If an arcface version's score is lower than the softmax version that I mentioned above, can get a better score in submission with an arcface version?\n(Since training takes a lot of time, I can't afford to do many experiments with my devices.\n\n\n=========================================================\nMy Model:\nEffcientNet3, GeM Polling(p=3), ArcFace(margin=0.5, scaler = 50)\n\nOptimizer:\nCosineRAdam or SGD\n\nBatch size: 16 * 8 = 128\n\nImage size: 448\nAugmentation : HorizontalFlip(p=0.5), RandomResizedCrop(scale(0.8,1.0))",
    "1493634": "I think using softmax, your model did not achieve the same class compactness and generalisation and hence over fitted on the marginal samples.\n\nWith arcface there is a  \\sqrt{2m} decision boundary in cosine space, so maybe most of the validation samples are falling into that region and getting misclassified. While a larger decision boundary is good for class compactness and intra-class genralisation, setting it very high may lead to your samples start behaving as if the model is being trained on an open-set scenario.",
    "1493778": "I did many experiments, setting scale = 50 and margin =0.3 seems like a good choice. I am expecting a good result when I make a submission."
  },
  "source": "meta"
}