{
  "id": 101392,
  "title": "ArcFace - Help needed",
  "url": "/competitions/recursion-cellular-image-classification/discussion/101392",
  "author_name": "",
  "post_date": "2019-07-25T12:43:25.420757900Z",
  "votes": 17,
  "comment_count": 21,
  "views": 0,
  "content": "<p>I want to try metric learning with ArcFace with pudae's arcface code. I modified it to work with my code and basemodel. During training the loss goes down, but the Accuracy stays always at 0.0 . I am using CrossEntropy-loss and SGD-optimizer with the ArcNet. I wonder what could be my fault?</p>\n\n<p>Another question i have, is that the ArcNet only during training seems to return the logits, in eval mode the forward() method returns only the features if i dont pass the labels to it. I saw, that there is special code for evaluation, why is that? </p>\n\n<p>Sorry if this are dumb questions, but this is the first time i want to try ArcFace and it seems to not work at all. I would be very thankful if someone could help me on explain for dummies how to train with ArcFace :-)</p>",
  "messages": [
    {
      "id": "584105",
      "postDate": "07/25/2019 12:43:25",
      "content": "<p>I want to try metric learning with ArcFace with pudae's arcface code. I modified it to work with my code and basemodel. During training the loss goes down, but the Accuracy stays always at 0.0 . I am using CrossEntropy-loss and SGD-optimizer with the ArcNet. I wonder what could be my fault?</p>\n\n<p>Another question i have, is that the ArcNet only during training seems to return the logits, in eval mode the forward() method returns only the features if i dont pass the labels to it. I saw, that there is special code for evaluation, why is that? </p>\n\n<p>Sorry if this are dumb questions, but this is the first time i want to try ArcFace and it seems to not work at all. I would be very thankful if someone could help me on explain for dummies how to train with ArcFace :-)</p>",
      "rawMarkdown": "I want to try metric learning with ArcFace with pudae's arcface code. I modified it to work with my code and basemodel. During training the loss goes down, but the Accuracy stays always at 0.0 . I am using CrossEntropy-loss and SGD-optimizer with the ArcNet. I wonder what could be my fault?\n\nAnother question i have, is that the ArcNet only during training seems to return the logits, in eval mode the forward() method returns only the features if i dont pass the labels to it. I saw, that there is special code for evaluation, why is that? \n\nSorry if this are dumb questions, but this is the first time i want to try ArcFace and it seems to not work at all. I would be very thankful if someone could help me on explain for dummies how to train with ArcFace :-)",
      "votes": null
    },
    {
      "id": "584216",
      "postDate": "07/25/2019 15:13:54",
      "content": "<p>hi marek. it is surprising to see that you are doing so well without using ArcFace yet. Can you please share your current approach? i.e., architecture?</p>",
      "rawMarkdown": "hi marek. it is surprising to see that you are doing so well without using ArcFace yet. Can you please share your current approach? i.e., architecture?",
      "votes": null
    },
    {
      "id": "584491",
      "postDate": "07/26/2019 04:23:45",
      "content": "<p>I'm having trouble with arcface too. \nMaybe just me but I see the implementation of <a href=\"/bestfitting\">@bestfitting</a> is easier to understand <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/78109#latest-560973\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/78109#latest-560973</a></p>\n\n<p>My result: Resnet18\n- Top line: Classification\n- Middle: ArcFace + classification (bestfitting)\n- Bottom line: ArcFace (pudea)</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F952137%2F785ca389d19731b7af9b981decd36a19%2FScreenshot_20190726_112037.png?generation=1564114864283490&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I'm having trouble with arcface too. \nMaybe just me but I see the implementation of @bestfitting is easier to understand https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/78109#latest-560973\n\nMy result: Resnet18\n- Top line: Classification\n- Middle: ArcFace + classification (bestfitting)\n- Bottom line: ArcFace (pudea)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F952137%2F785ca389d19731b7af9b981decd36a19%2FScreenshot_20190726_112037.png?generation=1564114864283490&amp;alt=media)",
      "votes": null
    },
    {
      "id": "584604",
      "postDate": "07/26/2019 08:03:52",
      "content": "<p><a href=\"/lkhphuc\">@lkhphuc</a> thank you very much, i will have a look at it. \n<a href=\"/wjshenggggg\">@wjshenggggg</a> i won't tell you my exact architecture, but i get with every net i tested at least over 0.4, even with low-resolution images (224x224) and CrossEntropy loss. I think before testing different architectures, you can improve your score a lot, if you do augmentation and average the predictions for both sites.</p>",
      "rawMarkdown": "lkhphuc thank you very much, i will have a look at it. \n@wjshenggggg i won't tell you my exact architecture, but i get with every net i tested at least over 0.4, even with low-resolution images (224x224) and CrossEntropy loss. I think before testing different architectures, you can improve your score a lot, if you do augmentation and average the predictions for both sites.",
      "votes": null
    },
    {
      "id": "584715",
      "postDate": "07/26/2019 11:24:17",
      "content": "<p>Hello, thanks for your kind information! Do you have any advice on what type of augmentations might be suitable for such 6-channel image data? </p>",
      "rawMarkdown": "Hello, thanks for your kind information! Do you have any advice on what type of augmentations might be suitable for such 6-channel image data?",
      "votes": null
    },
    {
      "id": "585115",
      "postDate": "07/27/2019 01:44:28",
      "content": "<p>Wish I could help, I’ve got the exact same problem trying to use arcnet. I fixed the thing with task.loss but I still get zero acc and immobile loss. Thinking it’s something in the loss dict...</p>",
      "rawMarkdown": "Wish I could help, I’ve got the exact same problem trying to use arcnet. I fixed the thing with task.loss but I still get zero acc and immobile loss. Thinking it’s something in the loss dict...",
      "votes": null
    },
    {
      "id": "585369",
      "postDate": "07/27/2019 10:22:59",
      "content": "<p><a href=\"/interneuron\">@interneuron</a> thank you very much for your feedback. At least i know, that i am not the only one getting zero acc, though my loss was decreasing :-/</p>",
      "rawMarkdown": "interneuron thank you very much for your feedback. At least i know, that i am not the only one getting zero acc, though my loss was decreasing :-/",
      "votes": null
    },
    {
      "id": "585478",
      "postDate": "07/27/2019 14:24:36",
      "content": "<p>I can confirm your results. The loss goes down but the accuracy stays 0. In my case, the initialization is  bad. I.e. the loss starts at ~43. But for random guessing it should be around ~<code>np.log(1108) = 7</code>.</p>\n\n<p>The inference is described here: <a href=\"https://www.kaggle.com/c/humpback-whale-identification/discussion/82484#latest-502552\">https://www.kaggle.com/c/humpback-whale-identification/discussion/82484#latest-502552</a></p>",
      "rawMarkdown": "I can confirm your results. The loss goes down but the accuracy stays 0. In my case, the initialization is  bad. I.e. the loss starts at ~43. But for random guessing it should be around ~`np.log(1108) = 7`.\n\nThe inference is described here: https://www.kaggle.com/c/humpback-whale-identification/discussion/82484#latest-502552",
      "votes": null
    },
    {
      "id": "585501",
      "postDate": "07/27/2019 15:01:58",
      "content": "<p>Thanks see-- , were you able to deal with the accuracy problem by changing the initialization? </p>\n\n<p>I admit I've found it a bit challenging to work with that code, removing the landmark stuff took me longer than it should have. I think the issue is I'm not used to working with code written as well as pudeas.</p>",
      "rawMarkdown": "Thanks see-- , were you able to deal with the accuracy problem by changing the initialization? \n\nI admit I've found it a bit challenging to work with that code, removing the landmark stuff took me longer than it should have. I think the issue is I'm not used to working with code written as well as pudeas.",
      "votes": null
    },
    {
      "id": "586127",
      "postDate": "07/28/2019 16:48:07",
      "content": "<p>&gt; Thanks see-- , were you able to deal with the accuracy problem by changing the initialization? </p>\n\n<p>It's probably not an issue. I just had to train longer. It's still worse than simple classification.</p>",
      "rawMarkdown": "&gt; Thanks see-- , were you able to deal with the accuracy problem by changing the initialization? \n\nIt's probably not an issue. I just had to train longer. It's still worse than simple classification.",
      "votes": null
    },
    {
      "id": "586607",
      "postDate": "07/29/2019 11:52:22",
      "content": "<p>Me too I am using arcface and it is giving me 0 accuracy...</p>",
      "rawMarkdown": "Me too I am using arcface and it is giving me 0 accuracy...",
      "votes": null
    },
    {
      "id": "592118",
      "postDate": "08/04/2019 20:47:14",
      "content": "<p>The 'dev' top5 accuracy stayed at 0 for me too. Going through pudae's code, I think this because of how 'task.metrics' works in eval mode, correct me if I'm wrong.</p>\n\n<p>In Eval mode, 'task.metrics' computes the similarity of a feature vector of an image to all other feature vectors in a single batch. So the score depends very much on a batch having many images of the same label, and in low batch sizes, this is not happening. In worst-case scenario, if all the labels in a batch are different, accuracy will be 0 irrespective of the model's performance.</p>",
      "rawMarkdown": "The 'dev' top5 accuracy stayed at 0 for me too. Going through pudae's code, I think this because of how 'task.metrics' works in eval mode, correct me if I'm wrong.\n\nIn Eval mode, 'task.metrics' computes the similarity of a feature vector of an image to all other feature vectors in a single batch. So the score depends very much on a batch having many images of the same label, and in low batch sizes, this is not happening. In worst-case scenario, if all the labels in a batch are different, accuracy will be 0 irrespective of the model's performance.",
      "votes": null
    },
    {
      "id": "602352",
      "postDate": "08/19/2019 01:16:58",
      "content": "<p>Just wanted to bump this up because I am exactly facing this issue.. <a href=\"/marekwyborski\">@marekwyborski</a> if you fixed this problem would you mind giving a hint? :)</p>",
      "rawMarkdown": "Just wanted to bump this up because I am exactly facing this issue.. @marekwyborski if you fixed this problem would you mind giving a hint? :)",
      "votes": null
    },
    {
      "id": "602581",
      "postDate": "08/19/2019 08:21:36",
      "content": "<p><a href=\"/hojaelee9212\">@hojaelee9212</a> i haven't tried again, but i think to make this work you would have to provide several classes with multiple images in one batch, like described here for Triplet Loss with Batch Hard strategy:\n<a href=\"https://omoindrot.github.io/triplet-loss\">https://omoindrot.github.io/triplet-loss</a>\ne.g. batch_size = 20 with 5 classes (sirnas) with each 4 images. This is how i understood <a href=\"/gautham11\">@gautham11</a> comment. I think just passing the batches with random sirnas is the problem, but like i said i haven't tried this again. Anyone please correct me if i am talking nonsense.</p>",
      "rawMarkdown": "hojaelee9212 i haven't tried again, but i think to make this work you would have to provide several classes with multiple images in one batch, like described here for Triplet Loss with Batch Hard strategy:\nhttps://omoindrot.github.io/triplet-loss\ne.g. batch_size = 20 with 5 classes (sirnas) with each 4 images. This is how i understood @gautham11 comment. I think just passing the batches with random sirnas is the problem, but like i said i haven't tried this again. Anyone please correct me if i am talking nonsense.",
      "votes": null
    },
    {
      "id": "602758",
      "postDate": "08/19/2019 13:11:29",
      "content": "<p><a href=\"/marekwyborski\">@marekwyborski</a> So at the moment your LB score comes from <em>simple</em> classification? Like, CNN with <code>CrossEntropyLoss</code> or something like that? Thanks.</p>",
      "rawMarkdown": "marekwyborski So at the moment your LB score comes from *simple* classification? Like, CNN with `CrossEntropyLoss` or something like that? Thanks.",
      "votes": null
    },
    {
      "id": "603537",
      "postDate": "08/20/2019 12:15:18",
      "content": "<p>Yes</p>",
      "rawMarkdown": "Yes",
      "votes": null
    },
    {
      "id": "607950",
      "postDate": "08/26/2019 05:48:42",
      "content": "<p>Is the <em>at least over 0.4</em> a LB score? Or CV?</p>",
      "rawMarkdown": "Is the *at least over 0.4* a LB score? Or CV?",
      "votes": null
    },
    {
      "id": "608655",
      "postDate": "08/27/2019 03:48:12",
      "content": "<p><a href=\"/udaykamal\">@udaykamal</a> Have you found a way to properly normalize the 6-channel images?</p>",
      "rawMarkdown": "udaykamal Have you found a way to properly normalize the 6-channel images?",
      "votes": null
    },
    {
      "id": "611100",
      "postDate": "08/29/2019 05:26:03",
      "content": "<p><a href=\"/marekwyborski\">@marekwyborski</a>  Is 0.4 the LB score using 224 RGB images .Is this without using the plates leak?\nThanks.</p>",
      "rawMarkdown": "marekwyborski  Is 0.4 the LB score using 224 RGB images .Is this without using the plates leak?\nThanks.",
      "votes": null
    },
    {
      "id": "612633",
      "postDate": "08/29/2019 23:01:31",
      "content": "<p><a href=\"/lorenzofabbri92\">@lorenzofabbri92</a> No, my current leaderboard score is just raw classification result, without any normalization. Using the Plate leak, the score got boosted from 0.483 to  0.599\nI am currently busy with another competition, so will get back to this one as soon as that ends and will let you know if I find anything interesting about normalization!</p>",
      "rawMarkdown": "lorenzofabbri92 No, my current leaderboard score is just raw classification result, without any normalization. Using the Plate leak, the score got boosted from 0.483 to  0.599\nI am currently busy with another competition, so will get back to this one as soon as that ends and will let you know if I find anything interesting about normalization!",
      "votes": null
    },
    {
      "id": "612822",
      "postDate": "08/30/2019 03:00:44",
      "content": "<p>Thanks! It's amazing how you people can go up to 0.6 with just raw classification, though... Best of luck then!</p>",
      "rawMarkdown": "Thanks! It's amazing how you people can go up to 0.6 with just raw classification, though... Best of luck then!",
      "votes": null
    },
    {
      "id": "613237",
      "postDate": "08/30/2019 10:39:21",
      "content": "<p><a href=\"/lorenzofabbri92\">@lorenzofabbri92</a> Thanks! umm, I am no expert in this field, but, just some tips for training:\n1) When using pre-trained model, you might first want to train the last layer (warmup) a few epochs first while freezing the remainings, then finetune the whole architecture end-to-end. This might help to faster convergence. \n2) Start with relatively small LR (say 1e-4) and stick to that until 10-20 epoch then use stepLR or other LR scheduler (as you wish) to reduce the LR.\n3) have patience and let the model train, say for at least 40-50 epochs. You might want to Reset (i.e. change the LR value to the initial one) the LR sometimes when you see no further change in Val loss. This might not sound like any systemic approach, but it works for my case.\nGood Luck to you too!! </p>",
      "rawMarkdown": "lorenzofabbri92 Thanks! umm, I am no expert in this field, but, just some tips for training:\n1) When using pre-trained model, you might first want to train the last layer (warmup) a few epochs first while freezing the remainings, then finetune the whole architecture end-to-end. This might help to faster convergence. \n2) Start with relatively small LR (say 1e-4) and stick to that until 10-20 epoch then use stepLR or other LR scheduler (as you wish) to reduce the LR.\n3) have patience and let the model train, say for at least 40-50 epochs. You might want to Reset (i.e. change the LR value to the initial one) the LR sometimes when you see no further change in Val loss. This might not sound like any systemic approach, but it works for my case.\nGood Luck to you too!!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 584216,
      "author_name": "wjshenggggg",
      "author_url": "",
      "post_date": "07/25/2019 15:13:54",
      "content": "<p>hi marek. it is surprising to see that you are doing so well without using ArcFace yet. Can you please share your current approach? i.e., architecture?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 584491,
      "author_name": "lkhphuc",
      "author_url": "",
      "post_date": "07/26/2019 04:23:45",
      "content": "<p>I'm having trouble with arcface too. \nMaybe just me but I see the implementation of <a href=\"/bestfitting\">@bestfitting</a> is easier to understand <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/78109#latest-560973\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/78109#latest-560973</a></p>\n\n<p>My result: Resnet18\n- Top line: Classification\n- Middle: ArcFace + classification (bestfitting)\n- Bottom line: ArcFace (pudea)</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F952137%2F785ca389d19731b7af9b981decd36a19%2FScreenshot_20190726_112037.png?generation=1564114864283490&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 584604,
      "author_name": "marekwyborski",
      "author_url": "",
      "post_date": "07/26/2019 08:03:52",
      "content": "<p><a href=\"/lkhphuc\">@lkhphuc</a> thank you very much, i will have a look at it. \n<a href=\"/wjshenggggg\">@wjshenggggg</a> i won't tell you my exact architecture, but i get with every net i tested at least over 0.4, even with low-resolution images (224x224) and CrossEntropy loss. I think before testing different architectures, you can improve your score a lot, if you do augmentation and average the predictions for both sites.</p>",
      "votes": null,
      "replies": [
        {
          "id": 584715,
          "author_name": "udaykamal",
          "author_url": "",
          "post_date": "07/26/2019 11:24:17",
          "content": "<p>Hello, thanks for your kind information! Do you have any advice on what type of augmentations might be suitable for such 6-channel image data? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 607950,
          "author_name": "lorenzofabbri92",
          "author_url": "",
          "post_date": "08/26/2019 05:48:42",
          "content": "<p>Is the <em>at least over 0.4</em> a LB score? Or CV?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 608655,
          "author_name": "lorenzofabbri92",
          "author_url": "",
          "post_date": "08/27/2019 03:48:12",
          "content": "<p><a href=\"/udaykamal\">@udaykamal</a> Have you found a way to properly normalize the 6-channel images?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 611100,
          "author_name": "ekansh",
          "author_url": "",
          "post_date": "08/29/2019 05:26:03",
          "content": "<p><a href=\"/marekwyborski\">@marekwyborski</a>  Is 0.4 the LB score using 224 RGB images .Is this without using the plates leak?\nThanks.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 612633,
          "author_name": "udaykamal",
          "author_url": "",
          "post_date": "08/29/2019 23:01:31",
          "content": "<p><a href=\"/lorenzofabbri92\">@lorenzofabbri92</a> No, my current leaderboard score is just raw classification result, without any normalization. Using the Plate leak, the score got boosted from 0.483 to  0.599\nI am currently busy with another competition, so will get back to this one as soon as that ends and will let you know if I find anything interesting about normalization!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 612822,
          "author_name": "lorenzofabbri92",
          "author_url": "",
          "post_date": "08/30/2019 03:00:44",
          "content": "<p>Thanks! It's amazing how you people can go up to 0.6 with just raw classification, though... Best of luck then!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 613237,
          "author_name": "udaykamal",
          "author_url": "",
          "post_date": "08/30/2019 10:39:21",
          "content": "<p><a href=\"/lorenzofabbri92\">@lorenzofabbri92</a> Thanks! umm, I am no expert in this field, but, just some tips for training:\n1) When using pre-trained model, you might first want to train the last layer (warmup) a few epochs first while freezing the remainings, then finetune the whole architecture end-to-end. This might help to faster convergence. \n2) Start with relatively small LR (say 1e-4) and stick to that until 10-20 epoch then use stepLR or other LR scheduler (as you wish) to reduce the LR.\n3) have patience and let the model train, say for at least 40-50 epochs. You might want to Reset (i.e. change the LR value to the initial one) the LR sometimes when you see no further change in Val loss. This might not sound like any systemic approach, but it works for my case.\nGood Luck to you too!! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 585115,
      "author_name": "interneuron",
      "author_url": "",
      "post_date": "07/27/2019 01:44:28",
      "content": "<p>Wish I could help, I’ve got the exact same problem trying to use arcnet. I fixed the thing with task.loss but I still get zero acc and immobile loss. Thinking it’s something in the loss dict...</p>",
      "votes": null,
      "replies": [
        {
          "id": 585369,
          "author_name": "marekwyborski",
          "author_url": "",
          "post_date": "07/27/2019 10:22:59",
          "content": "<p><a href=\"/interneuron\">@interneuron</a> thank you very much for your feedback. At least i know, that i am not the only one getting zero acc, though my loss was decreasing :-/</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 585478,
          "author_name": "seesee",
          "author_url": "",
          "post_date": "07/27/2019 14:24:36",
          "content": "<p>I can confirm your results. The loss goes down but the accuracy stays 0. In my case, the initialization is  bad. I.e. the loss starts at ~43. But for random guessing it should be around ~<code>np.log(1108) = 7</code>.</p>\n\n<p>The inference is described here: <a href=\"https://www.kaggle.com/c/humpback-whale-identification/discussion/82484#latest-502552\">https://www.kaggle.com/c/humpback-whale-identification/discussion/82484#latest-502552</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 585501,
          "author_name": "interneuron",
          "author_url": "",
          "post_date": "07/27/2019 15:01:58",
          "content": "<p>Thanks see-- , were you able to deal with the accuracy problem by changing the initialization? </p>\n\n<p>I admit I've found it a bit challenging to work with that code, removing the landmark stuff took me longer than it should have. I think the issue is I'm not used to working with code written as well as pudeas.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 586127,
          "author_name": "seesee",
          "author_url": "",
          "post_date": "07/28/2019 16:48:07",
          "content": "<p>&gt; Thanks see-- , were you able to deal with the accuracy problem by changing the initialization? </p>\n\n<p>It's probably not an issue. I just had to train longer. It's still worse than simple classification.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 592118,
          "author_name": "gautham11",
          "author_url": "",
          "post_date": "08/04/2019 20:47:14",
          "content": "<p>The 'dev' top5 accuracy stayed at 0 for me too. Going through pudae's code, I think this because of how 'task.metrics' works in eval mode, correct me if I'm wrong.</p>\n\n<p>In Eval mode, 'task.metrics' computes the similarity of a feature vector of an image to all other feature vectors in a single batch. So the score depends very much on a batch having many images of the same label, and in low batch sizes, this is not happening. In worst-case scenario, if all the labels in a batch are different, accuracy will be 0 irrespective of the model's performance.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 586607,
      "author_name": "rinnqd",
      "author_url": "",
      "post_date": "07/29/2019 11:52:22",
      "content": "<p>Me too I am using arcface and it is giving me 0 accuracy...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 602352,
      "author_name": "hojaelee9212",
      "author_url": "",
      "post_date": "08/19/2019 01:16:58",
      "content": "<p>Just wanted to bump this up because I am exactly facing this issue.. <a href=\"/marekwyborski\">@marekwyborski</a> if you fixed this problem would you mind giving a hint? :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 602581,
          "author_name": "marekwyborski",
          "author_url": "",
          "post_date": "08/19/2019 08:21:36",
          "content": "<p><a href=\"/hojaelee9212\">@hojaelee9212</a> i haven't tried again, but i think to make this work you would have to provide several classes with multiple images in one batch, like described here for Triplet Loss with Batch Hard strategy:\n<a href=\"https://omoindrot.github.io/triplet-loss\">https://omoindrot.github.io/triplet-loss</a>\ne.g. batch_size = 20 with 5 classes (sirnas) with each 4 images. This is how i understood <a href=\"/gautham11\">@gautham11</a> comment. I think just passing the batches with random sirnas is the problem, but like i said i haven't tried this again. Anyone please correct me if i am talking nonsense.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 602758,
          "author_name": "lorenzofabbri92",
          "author_url": "",
          "post_date": "08/19/2019 13:11:29",
          "content": "<p><a href=\"/marekwyborski\">@marekwyborski</a> So at the moment your LB score comes from <em>simple</em> classification? Like, CNN with <code>CrossEntropyLoss</code> or something like that? Thanks.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 603537,
          "author_name": "marekwyborski",
          "author_url": "",
          "post_date": "08/20/2019 12:15:18",
          "content": "<p>Yes</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "584105": "I want to try metric learning with ArcFace with pudae's arcface code. I modified it to work with my code and basemodel. During training the loss goes down, but the Accuracy stays always at 0.0 . I am using CrossEntropy-loss and SGD-optimizer with the ArcNet. I wonder what could be my fault?\n\nAnother question i have, is that the ArcNet only during training seems to return the logits, in eval mode the forward() method returns only the features if i dont pass the labels to it. I saw, that there is special code for evaluation, why is that? \n\nSorry if this are dumb questions, but this is the first time i want to try ArcFace and it seems to not work at all. I would be very thankful if someone could help me on explain for dummies how to train with ArcFace :-)",
    "584216": "hi marek. it is surprising to see that you are doing so well without using ArcFace yet. Can you please share your current approach? i.e., architecture?",
    "584491": "I'm having trouble with arcface too. \nMaybe just me but I see the implementation of @bestfitting is easier to understand https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/78109#latest-560973\n\nMy result: Resnet18\n- Top line: Classification\n- Middle: ArcFace + classification (bestfitting)\n- Bottom line: ArcFace (pudea)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F952137%2F785ca389d19731b7af9b981decd36a19%2FScreenshot_20190726_112037.png?generation=1564114864283490&amp;alt=media)",
    "584604": "lkhphuc thank you very much, i will have a look at it. \n@wjshenggggg i won't tell you my exact architecture, but i get with every net i tested at least over 0.4, even with low-resolution images (224x224) and CrossEntropy loss. I think before testing different architectures, you can improve your score a lot, if you do augmentation and average the predictions for both sites.",
    "584715": "Hello, thanks for your kind information! Do you have any advice on what type of augmentations might be suitable for such 6-channel image data?",
    "585115": "Wish I could help, I’ve got the exact same problem trying to use arcnet. I fixed the thing with task.loss but I still get zero acc and immobile loss. Thinking it’s something in the loss dict...",
    "585369": "interneuron thank you very much for your feedback. At least i know, that i am not the only one getting zero acc, though my loss was decreasing :-/",
    "585478": "I can confirm your results. The loss goes down but the accuracy stays 0. In my case, the initialization is  bad. I.e. the loss starts at ~43. But for random guessing it should be around ~`np.log(1108) = 7`.\n\nThe inference is described here: https://www.kaggle.com/c/humpback-whale-identification/discussion/82484#latest-502552",
    "585501": "Thanks see-- , were you able to deal with the accuracy problem by changing the initialization? \n\nI admit I've found it a bit challenging to work with that code, removing the landmark stuff took me longer than it should have. I think the issue is I'm not used to working with code written as well as pudeas.",
    "586127": "&gt; Thanks see-- , were you able to deal with the accuracy problem by changing the initialization? \n\nIt's probably not an issue. I just had to train longer. It's still worse than simple classification.",
    "586607": "Me too I am using arcface and it is giving me 0 accuracy...",
    "592118": "The 'dev' top5 accuracy stayed at 0 for me too. Going through pudae's code, I think this because of how 'task.metrics' works in eval mode, correct me if I'm wrong.\n\nIn Eval mode, 'task.metrics' computes the similarity of a feature vector of an image to all other feature vectors in a single batch. So the score depends very much on a batch having many images of the same label, and in low batch sizes, this is not happening. In worst-case scenario, if all the labels in a batch are different, accuracy will be 0 irrespective of the model's performance.",
    "602352": "Just wanted to bump this up because I am exactly facing this issue.. @marekwyborski if you fixed this problem would you mind giving a hint? :)",
    "602581": "hojaelee9212 i haven't tried again, but i think to make this work you would have to provide several classes with multiple images in one batch, like described here for Triplet Loss with Batch Hard strategy:\nhttps://omoindrot.github.io/triplet-loss\ne.g. batch_size = 20 with 5 classes (sirnas) with each 4 images. This is how i understood @gautham11 comment. I think just passing the batches with random sirnas is the problem, but like i said i haven't tried this again. Anyone please correct me if i am talking nonsense.",
    "602758": "marekwyborski So at the moment your LB score comes from *simple* classification? Like, CNN with `CrossEntropyLoss` or something like that? Thanks.",
    "603537": "Yes",
    "607950": "Is the *at least over 0.4* a LB score? Or CV?",
    "608655": "udaykamal Have you found a way to properly normalize the 6-channel images?",
    "611100": "marekwyborski  Is 0.4 the LB score using 224 RGB images .Is this without using the plates leak?\nThanks.",
    "612633": "lorenzofabbri92 No, my current leaderboard score is just raw classification result, without any normalization. Using the Plate leak, the score got boosted from 0.483 to  0.599\nI am currently busy with another competition, so will get back to this one as soon as that ends and will let you know if I find anything interesting about normalization!",
    "612822": "Thanks! It's amazing how you people can go up to 0.6 with just raw classification, though... Best of luck then!",
    "613237": "lorenzofabbri92 Thanks! umm, I am no expert in this field, but, just some tips for training:\n1) When using pre-trained model, you might first want to train the last layer (warmup) a few epochs first while freezing the remainings, then finetune the whole architecture end-to-end. This might help to faster convergence. \n2) Start with relatively small LR (say 1e-4) and stick to that until 10-20 epoch then use stepLR or other LR scheduler (as you wish) to reduce the LR.\n3) have patience and let the model train, say for at least 40-50 epochs. You might want to Reset (i.e. change the LR value to the initial one) the LR sometimes when you see no further change in Val loss. This might not sound like any systemic approach, but it works for my case.\nGood Luck to you too!!"
  },
  "source": "meta"
}