{
  "id": 307402,
  "title": "tensorflow score > pytorch score ??",
  "url": "/competitions/happy-whale-and-dolphin/discussion/307402",
  "author_name": "",
  "post_date": "2022-02-14T04:29:36.909289800Z",
  "votes": 19,
  "comment_count": 32,
  "views": 0,
  "content": "<p>I am trying to replicate the high scoring TensorFlow notebooks in pytorch, but its nowhere near the score of the public notebooks.<br>\nIs there some trick behind TF training? and has somebody been able to replicate those high scores on pytorch?<br>\nNeed tips from you guys.<br>\nThanks</p>\n<p>Edit<br>\nwas trying to replicate this notebook <a href=\"https://www.kaggle.com/ks2019/happywhale-arcface-baseline-tpu\" target=\"_blank\">https://www.kaggle.com/ks2019/happywhale-arcface-baseline-tpu</a> by <a href=\"https://www.kaggle.com/ks2019\" target=\"_blank\">@ks2019</a> </p>",
  "messages": [
    {
      "id": "1689180",
      "postDate": "02/14/2022 04:29:36",
      "content": "<p>I am trying to replicate the high scoring TensorFlow notebooks in pytorch, but its nowhere near the score of the public notebooks.<br>\nIs there some trick behind TF training? and has somebody been able to replicate those high scores on pytorch?<br>\nNeed tips from you guys.<br>\nThanks</p>\n<p>Edit<br>\nwas trying to replicate this notebook <a href=\"https://www.kaggle.com/ks2019/happywhale-arcface-baseline-tpu\" target=\"_blank\">https://www.kaggle.com/ks2019/happywhale-arcface-baseline-tpu</a> by <a href=\"https://www.kaggle.com/ks2019\" target=\"_blank\">@ks2019</a> </p>",
      "rawMarkdown": "I am trying to replicate the high scoring TensorFlow notebooks in pytorch, but its nowhere near the score of the public notebooks.\nIs there some trick behind TF training? and has somebody been able to replicate those high scores on pytorch?\nNeed tips from you guys.\nThanks\n\nEdit\nwas trying to replicate this notebook https://www.kaggle.com/ks2019/happywhale-arcface-baseline-tpu by @ks2019",
      "votes": null
    },
    {
      "id": "1689205",
      "postDate": "02/14/2022 04:48:10",
      "content": "<p>Could you please link to the notebooks that you're trying to replicate? </p>\n<p>Thanks! :)</p>",
      "rawMarkdown": "Could you please link to the notebooks that you're trying to replicate? \n\nThanks! :)",
      "votes": null
    },
    {
      "id": "1689244",
      "postDate": "02/14/2022 05:17:53",
      "content": "<p>hey I've been also trying the same but so far no luck ( TPU with pytorch is bad ) </p>",
      "rawMarkdown": "hey I've been also trying the same but so far no luck ( TPU with pytorch is bad )",
      "votes": null
    },
    {
      "id": "1689251",
      "postDate": "02/14/2022 05:40:15",
      "content": "<p>I was trying to replicate this <a href=\"https://www.kaggle.com/ks2019/happywhale-arcface-baseline-tpu\" target=\"_blank\">https://www.kaggle.com/ks2019/happywhale-arcface-baseline-tpu</a></p>",
      "rawMarkdown": "I was trying to replicate this https://www.kaggle.com/ks2019/happywhale-arcface-baseline-tpu",
      "votes": null
    },
    {
      "id": "1689380",
      "postDate": "02/14/2022 07:29:58",
      "content": "<p>One thing is I am just simply resizing to size 512, in the recent interview <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> with <a href=\"https://www.kaggle.com/init27\" target=\"_blank\">@init27</a> said that doing this may harm the aspect ratio, I don't know how much this is important this time.</p>",
      "rawMarkdown": "One thing is I am just simply resizing to size 512, in the recent interview @cdeotte with @init27 said that doing this may harm the aspect ratio, I don't know how much this is important this time.",
      "votes": null
    },
    {
      "id": "1689393",
      "postDate": "02/14/2022 07:39:47",
      "content": "<p>also, I fear random cropping images is not a good idea, as, in many images, random cropping will only result in the water, not the whale/dolphin, so don't know if its a good idea</p>",
      "rawMarkdown": "also, I fear random cropping images is not a good idea, as, in many images, random cropping will only result in the water, not the whale/dolphin, so don't know if its a good idea",
      "votes": null
    },
    {
      "id": "1689530",
      "postDate": "02/14/2022 09:50:38",
      "content": "<p>Did you try to replicate the inferencing part as well?<br>\nI tried my inference pipeline with my model and a baseline model trained with the notebook you mentioned and the results are quite different. So I am planning to replicate the inferencing part and compare the differences.<br>\nWill update once I do the comparisons</p>",
      "rawMarkdown": "Did you try to replicate the inferencing part as well?\nI tried my inference pipeline with my model and a baseline model trained with the notebook you mentioned and the results are quite different. So I am planning to replicate the inferencing part and compare the differences.\nWill update once I do the comparisons",
      "votes": null
    },
    {
      "id": "1689579",
      "postDate": "02/14/2022 10:37:37",
      "content": "<p>Not the inferencing part though, was trying with the training part.</p>",
      "rawMarkdown": "Not the inferencing part though, was trying with the training part.",
      "votes": null
    },
    {
      "id": "1689653",
      "postDate": "02/14/2022 11:45:05",
      "content": "<p>are you trying training with TPU or with GPU? </p>\n<p>again I think he did trained on different platform with cross validation make sure you take that in your consideration while training notebook </p>",
      "rawMarkdown": "are you trying training with TPU or with GPU? \n\nagain I think he did trained on different platform with cross validation make sure you take that in your consideration while training notebook",
      "votes": null
    },
    {
      "id": "1689681",
      "postDate": "02/14/2022 11:58:04",
      "content": "<p>there is dataset made with help of yolo it crops the images with the whales portion <a href=\"https://www.kaggle.com/awsaf49\" target=\"_blank\">@awsaf49</a> thanks for dataset </p>\n<h3>this notebook might help you <a href=\"https://www.kaggle.com/mrinath\" target=\"_blank\">@mrinath</a></h3>\n<p><a href=\"https://www.kaggle.com/awsaf49/happywhale-cropped-dataset-yolov5\" target=\"_blank\">https://www.kaggle.com/awsaf49/happywhale-cropped-dataset-yolov5</a></p>",
      "rawMarkdown": "there is dataset made with help of yolo it crops the images with the whales portion @awsaf49 thanks for dataset \n\n### this notebook might help you @mrinath \n[https://www.kaggle.com/awsaf49/happywhale-cropped-dataset-yolov5](https://www.kaggle.com/awsaf49/happywhale-cropped-dataset-yolov5)",
      "votes": null
    },
    {
      "id": "1689798",
      "postDate": "02/14/2022 13:59:40",
      "content": "<p>Using Pytorch training model  can reach LB 0.63 single model.  Configuration reference tf LB 0.60</p>",
      "rawMarkdown": "Using Pytorch training model  can reach LB 0.63 single model.  Configuration reference tf LB 0.60",
      "votes": null
    },
    {
      "id": "1689803",
      "postDate": "02/14/2022 14:03:23",
      "content": "<p>By the way, the optimizer warmup is very important.</p>",
      "rawMarkdown": "By the way, the optimizer warmup is very important.",
      "votes": null
    },
    {
      "id": "1689836",
      "postDate": "02/14/2022 14:28:59",
      "content": "<p>Can you elaborate on your inference pipeline?</p>",
      "rawMarkdown": "Can you elaborate on your inference pipeline?",
      "votes": null
    },
    {
      "id": "1690151",
      "postDate": "02/14/2022 18:25:05",
      "content": "<p><a href=\"https://www.kaggle.com/biglafe\" target=\"_blank\">@biglafe</a> whats your CV score, and your final val loss(assuming you are using ArcFace)</p>",
      "rawMarkdown": "biglafe whats your CV score, and your final val loss(assuming you are using ArcFace)",
      "votes": null
    },
    {
      "id": "1690152",
      "postDate": "02/14/2022 18:25:33",
      "content": "<p>training on GPU </p>",
      "rawMarkdown": "training on GPU",
      "votes": null
    },
    {
      "id": "1690285",
      "postDate": "02/14/2022 20:43:56",
      "content": "<p>Thanks for this hint! Are you using a metric learning approach similar to the tf public notebooks?</p>",
      "rawMarkdown": "Thanks for this hint! Are you using a metric learning approach similar to the tf public notebooks?",
      "votes": null
    },
    {
      "id": "1690287",
      "postDate": "02/14/2022 20:45:00",
      "content": "<p>How different are we talking? Interested to understand if their inference process has something I'm missing. It seems to make a big difference ;)</p>",
      "rawMarkdown": "How different are we talking? Interested to understand if their inference process has something I'm missing. It seems to make a big difference ;)",
      "votes": null
    },
    {
      "id": "1690658",
      "postDate": "02/15/2022 03:43:59",
      "content": "<p>I get an LB of 0.069 with PyTorch and 0.403 with tf. I am checking for bugs in my inference pipeline.</p>\n<p>This is what I was trying to do</p>\n<ol>\n<li>Get the top 50 neighbors (distance, indexes of the samples)</li>\n<li>Now I get the classes of the samples (so now I have distances, classes)</li>\n<li>If the class is already added to the list, skip (continue)</li>\n<li>Else if distance &gt; threshold, add the class to the list (distance refers to cosine similarity)</li>\n<li>Else add <code>new_individual</code></li>\n<li>Once <code>new_invidual</code> is added, add the rest of the classes</li>\n</ol>\n<pre><code>def get_preds(distances, indices, threshold=0.5):\n    classes = idx2class(indices)\n\n    new = 0\n    preds = []\n\n    for dist, class_ in zip(distances, classes):\n        if len(preds) == 5:\n            break\n        elif class_ in preds:\n            continue\n        elif dist &gt; threshold:\n            preds.append(class_)\n        elif not new:\n            preds.append('new_individual')\n            new = 1\n        else:\n            preds.append(class_)\n\n    return preds\n</code></pre>",
      "rawMarkdown": "I get an LB of 0.069 with PyTorch and 0.403 with tf. I am checking for bugs in my inference pipeline.\n\nThis is what I was trying to do\n1. Get the top 50 neighbors (distance, indexes of the samples)\n2. Now I get the classes of the samples (so now I have distances, classes)\n3. If the class is already added to the list, skip (continue)\n4. Else if distance > threshold, add the class to the list (distance refers to cosine similarity)\n5. Else add `new_individual`\n6. Once `new_invidual` is added, add the rest of the classes\n\n```\ndef get_preds(distances, indices, threshold=0.5):\n    classes = idx2class(indices)\n    \n    new = 0\n    preds = []\n    \n    for dist, class_ in zip(distances, classes):\n        if len(preds) == 5:\n            break\n        elif class_ in preds:\n            continue\n        elif dist > threshold:\n            preds.append(class_)\n        elif not new:\n            preds.append('new_individual')\n            new = 1\n        else:\n            preds.append(class_)\n            \n    return preds\n```",
      "votes": null
    },
    {
      "id": "1690710",
      "postDate": "02/15/2022 04:20:17",
      "content": "<p><code>I get an LB of 0.069 with PyTorch and 0.403 with tf. I am checking for bugs in my inference pipeline.</code><br>\nthe models were trained similarly?</p>",
      "rawMarkdown": "`I get an LB of 0.069 with PyTorch and 0.403 with tf. I am checking for bugs in my inference pipeline.`\nthe models were trained similarly?",
      "votes": null
    },
    {
      "id": "1690743",
      "postDate": "02/15/2022 04:43:02",
      "content": "<p>Not exactly. Same backbone using arcface in both. Image sizes are different 448 for pytorch, 512 for tf</p>",
      "rawMarkdown": "Not exactly. Same backbone using arcface in both. Image sizes are different 448 for pytorch, 512 for tf",
      "votes": null
    },
    {
      "id": "1690927",
      "postDate": "02/15/2022 06:53:30",
      "content": "<p>any idea <a href=\"https://www.kaggle.com/debarshichanda\" target=\"_blank\">@debarshichanda</a> why it's happening? </p>",
      "rawMarkdown": "any idea @debarshichanda why it's happening?",
      "votes": null
    },
    {
      "id": "1691091",
      "postDate": "02/15/2022 08:27:42",
      "content": "<p><a href=\"https://www.kaggle.com/somesh88\" target=\"_blank\">@somesh88</a> I think the model didn't learn anything. I checked the distances and they are all close to 0.9. Will also check the distances for the tf model once I get TPU access</p>",
      "rawMarkdown": "somesh88 I think the model didn't learn anything. I checked the distances and they are all close to 0.9. Will also check the distances for the tf model once I get TPU access",
      "votes": null
    },
    {
      "id": "1691542",
      "postDate": "02/15/2022 13:17:17",
      "content": "<p>val loss 9.6, inference pipeline complete references \b TF 0.6 <a href=\"https://www.kaggle.com/ks2019/happywhale-arcface-baseline-tpu\" target=\"_blank\">https://www.kaggle.com/ks2019/happywhale-arcface-baseline-tpu</a> by <a href=\"https://www.kaggle.com/ks2019\" target=\"_blank\">@ks2019</a></p>",
      "rawMarkdown": "val loss 9.6, inference pipeline complete references \b TF 0.6 https://www.kaggle.com/ks2019/happywhale-arcface-baseline-tpu by @ks2019",
      "votes": null
    },
    {
      "id": "1691550",
      "postDate": "02/15/2022 13:21:19",
      "content": "<p>Most of the training and testing structures refer to tf public notebooks, but some learning rates and optimizer methods are adjusted. <br>\nThe measurement method also uses KNN to find similarity. Later, I will try some post-processing methods similar to GLR competition</p>",
      "rawMarkdown": "Most of the training and testing structures refer to tf public notebooks, but some learning rates and optimizer methods are adjusted. \nThe measurement method also uses KNN to find similarity. Later, I will try some post-processing methods similar to GLR competition",
      "votes": null
    },
    {
      "id": "1692660",
      "postDate": "02/16/2022 07:00:08",
      "content": "<p>One doubt on inference, like to get the final predictions we are comparing test embeddings with the trained ones right,<br>\nsuppose I have used 5 folds when training, to get the train embeddings of all the train samples, do I have to find embeddings using all the 5 models and average it? I may be wrong, whats the correct way?</p>",
      "rawMarkdown": "One doubt on inference, like to get the final predictions we are comparing test embeddings with the trained ones right,\nsuppose I have used 5 folds when training, to get the train embeddings of all the train samples, do I have to find embeddings using all the 5 models and average it? I may be wrong, whats the correct way?",
      "votes": null
    },
    {
      "id": "1692941",
      "postDate": "02/16/2022 11:04:31",
      "content": "<p>What batch size do you use for training? Batch size is very important with ArcFace.</p>",
      "rawMarkdown": "What batch size do you use for training? Batch size is very important with ArcFace.",
      "votes": null
    },
    {
      "id": "1692969",
      "postDate": "02/16/2022 11:24:10",
      "content": "<p><a href=\"https://www.kaggle.com/ks2019\" target=\"_blank\">@ks2019</a> I am just using a batch size of 32 so maybe that is a problem. I will try using a higher batch size and revert.</p>",
      "rawMarkdown": "ks2019 I am just using a batch size of 32 so maybe that is a problem. I will try using a higher batch size and revert.",
      "votes": null
    },
    {
      "id": "1692972",
      "postDate": "02/16/2022 11:26:26",
      "content": "<p><a href=\"https://www.kaggle.com/mrinath\" target=\"_blank\">@mrinath</a> I think there are 2 ways. One would be to average the embeddings and the other to concatenate. I think averaging would make more sense if they are same architecture trained on different folds and concatenating might be good for different architectures</p>",
      "rawMarkdown": "mrinath I think there are 2 ways. One would be to average the embeddings and the other to concatenate. I think averaging would make more sense if they are same architecture trained on different folds and concatenating might be good for different architectures",
      "votes": null
    },
    {
      "id": "1692984",
      "postDate": "02/16/2022 11:33:02",
      "content": "<p><a href=\"https://www.kaggle.com/ks2019\" target=\"_blank\">@ks2019</a> more the batch size the better?</p>",
      "rawMarkdown": "ks2019 more the batch size the better?",
      "votes": null
    },
    {
      "id": "1693047",
      "postDate": "02/16/2022 12:09:41",
      "content": "<p><a href=\"https://www.kaggle.com/debarshichanda\" target=\"_blank\">@debarshichanda</a> I think 32 should be good to go, but try higher batch size once.<br>\n<a href=\"https://www.kaggle.com/mrinath\" target=\"_blank\">@mrinath</a> yes. With lower batch size, feature drift is higher. The ArcFace parameters (class centres) needs to be in sync with backbone. If for a class, we don't have any sample in many contiguous batches, we start observing a drift b/w feature representation and class centres representation. <br>\nNote that ArcFace doesn't employ sample-sample comparisons, but sample-prototype comparisons. The prototype stored in the ArcFace layer (or, the trainable parameter associated with ArcFace) represents the centre of each class. Training samples are enforced to be close to positive prototypes and far apart from negative prototypes by a clear margin (m). </p>",
      "rawMarkdown": "debarshichanda I think 32 should be good to go, but try higher batch size once.\n@mrinath yes. With lower batch size, feature drift is higher. The ArcFace parameters (class centres) needs to be in sync with backbone. If for a class, we don't have any sample in many contiguous batches, we start observing a drift b/w feature representation and class centres representation. \nNote that ArcFace doesn't employ sample-sample comparisons, but sample-prototype comparisons. The prototype stored in the ArcFace layer (or, the trainable parameter associated with ArcFace) represents the centre of each class. Training samples are enforced to be close to positive prototypes and far apart from negative prototypes by a clear margin (m).",
      "votes": null
    },
    {
      "id": "1706018",
      "postDate": "02/27/2022 05:09:50",
      "content": "<p>by optimizer warmup do you mean using LR schedulers? <a href=\"https://www.kaggle.com/biglafe\" target=\"_blank\">@biglafe</a> </p>",
      "rawMarkdown": "by optimizer warmup do you mean using LR schedulers? @biglafe",
      "votes": null
    },
    {
      "id": "1706024",
      "postDate": "02/27/2022 05:15:27",
      "content": "<p>Yes. When I use pytorch to train effb5 , it greatly affects my score.</p>",
      "rawMarkdown": "Yes. When I use pytorch to train effb5 , it greatly affects my score.",
      "votes": null
    },
    {
      "id": "1728699",
      "postDate": "03/19/2022 06:58:00",
      "content": "<p>I think so too.<br>\nI try to code using pytorch as same this Notebook 'kaggle kernels pull code/aikhmelnytskyy/happywhale-arcface-baseline-eff-net-kfold5-0-652' <a href=\"https://www.kaggle.com/Andrij\" target=\"_blank\">@Andrij</a> which is based on Tensorflow, but I got much more lower score than it.</p>\n<p>In addition, when I try to use tf_efficientnet_b7(like timm_create('tf_efficientnet_b7')) on colab TPU, it does not work.<br>\n(but in case of tf_efficientnet_b5, it works. I don't know why. It is bug?)<br>\nAnyway, for now, I think we should use Tensorflow if I want to use efficientnet on TPU envirionment.</p>",
      "rawMarkdown": "I think so too.\nI try to code using pytorch as same this Notebook 'kaggle kernels pull code/aikhmelnytskyy/happywhale-arcface-baseline-eff-net-kfold5-0-652' @Andrij which is based on Tensorflow, but I got much more lower score than it.\n\nIn addition, when I try to use tf_efficientnet_b7(like timm_create('tf_efficientnet_b7')) on colab TPU, it does not work.\n(but in case of tf_efficientnet_b5, it works. I don't know why. It is bug?)\nAnyway, for now, I think we should use Tensorflow if I want to use efficientnet on TPU envirionment.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1689205,
      "author_name": "init27",
      "author_url": "",
      "post_date": "02/14/2022 04:48:10",
      "content": "<p>Could you please link to the notebooks that you're trying to replicate? </p>\n<p>Thanks! :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1689251,
          "author_name": "mrinath",
          "author_url": "",
          "post_date": "02/14/2022 05:40:15",
          "content": "<p>I was trying to replicate this <a href=\"https://www.kaggle.com/ks2019/happywhale-arcface-baseline-tpu\" target=\"_blank\">https://www.kaggle.com/ks2019/happywhale-arcface-baseline-tpu</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1689653,
          "author_name": "somesh88",
          "author_url": "",
          "post_date": "02/14/2022 11:45:05",
          "content": "<p>are you trying training with TPU or with GPU? </p>\n<p>again I think he did trained on different platform with cross validation make sure you take that in your consideration while training notebook </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1690152,
          "author_name": "mrinath",
          "author_url": "",
          "post_date": "02/14/2022 18:25:33",
          "content": "<p>training on GPU </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1689244,
      "author_name": "somesh88",
      "author_url": "",
      "post_date": "02/14/2022 05:17:53",
      "content": "<p>hey I've been also trying the same but so far no luck ( TPU with pytorch is bad ) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1689380,
      "author_name": "mrinath",
      "author_url": "",
      "post_date": "02/14/2022 07:29:58",
      "content": "<p>One thing is I am just simply resizing to size 512, in the recent interview <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> with <a href=\"https://www.kaggle.com/init27\" target=\"_blank\">@init27</a> said that doing this may harm the aspect ratio, I don't know how much this is important this time.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1689393,
          "author_name": "mrinath",
          "author_url": "",
          "post_date": "02/14/2022 07:39:47",
          "content": "<p>also, I fear random cropping images is not a good idea, as, in many images, random cropping will only result in the water, not the whale/dolphin, so don't know if its a good idea</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1689681,
          "author_name": "somesh88",
          "author_url": "",
          "post_date": "02/14/2022 11:58:04",
          "content": "<p>there is dataset made with help of yolo it crops the images with the whales portion <a href=\"https://www.kaggle.com/awsaf49\" target=\"_blank\">@awsaf49</a> thanks for dataset </p>\n<h3>this notebook might help you <a href=\"https://www.kaggle.com/mrinath\" target=\"_blank\">@mrinath</a></h3>\n<p><a href=\"https://www.kaggle.com/awsaf49/happywhale-cropped-dataset-yolov5\" target=\"_blank\">https://www.kaggle.com/awsaf49/happywhale-cropped-dataset-yolov5</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1689530,
      "author_name": "debarshichanda",
      "author_url": "",
      "post_date": "02/14/2022 09:50:38",
      "content": "<p>Did you try to replicate the inferencing part as well?<br>\nI tried my inference pipeline with my model and a baseline model trained with the notebook you mentioned and the results are quite different. So I am planning to replicate the inferencing part and compare the differences.<br>\nWill update once I do the comparisons</p>",
      "votes": null,
      "replies": [
        {
          "id": 1689579,
          "author_name": "mrinath",
          "author_url": "",
          "post_date": "02/14/2022 10:37:37",
          "content": "<p>Not the inferencing part though, was trying with the training part.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1690287,
          "author_name": "taindow",
          "author_url": "",
          "post_date": "02/14/2022 20:45:00",
          "content": "<p>How different are we talking? Interested to understand if their inference process has something I'm missing. It seems to make a big difference ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1690658,
          "author_name": "debarshichanda",
          "author_url": "",
          "post_date": "02/15/2022 03:43:59",
          "content": "<p>I get an LB of 0.069 with PyTorch and 0.403 with tf. I am checking for bugs in my inference pipeline.</p>\n<p>This is what I was trying to do</p>\n<ol>\n<li>Get the top 50 neighbors (distance, indexes of the samples)</li>\n<li>Now I get the classes of the samples (so now I have distances, classes)</li>\n<li>If the class is already added to the list, skip (continue)</li>\n<li>Else if distance &gt; threshold, add the class to the list (distance refers to cosine similarity)</li>\n<li>Else add <code>new_individual</code></li>\n<li>Once <code>new_invidual</code> is added, add the rest of the classes</li>\n</ol>\n<pre><code>def get_preds(distances, indices, threshold=0.5):\n    classes = idx2class(indices)\n\n    new = 0\n    preds = []\n\n    for dist, class_ in zip(distances, classes):\n        if len(preds) == 5:\n            break\n        elif class_ in preds:\n            continue\n        elif dist &gt; threshold:\n            preds.append(class_)\n        elif not new:\n            preds.append('new_individual')\n            new = 1\n        else:\n            preds.append(class_)\n\n    return preds\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1690710,
          "author_name": "mrinath",
          "author_url": "",
          "post_date": "02/15/2022 04:20:17",
          "content": "<p><code>I get an LB of 0.069 with PyTorch and 0.403 with tf. I am checking for bugs in my inference pipeline.</code><br>\nthe models were trained similarly?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1690743,
          "author_name": "debarshichanda",
          "author_url": "",
          "post_date": "02/15/2022 04:43:02",
          "content": "<p>Not exactly. Same backbone using arcface in both. Image sizes are different 448 for pytorch, 512 for tf</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1690927,
          "author_name": "somesh88",
          "author_url": "",
          "post_date": "02/15/2022 06:53:30",
          "content": "<p>any idea <a href=\"https://www.kaggle.com/debarshichanda\" target=\"_blank\">@debarshichanda</a> why it's happening? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1691091,
          "author_name": "debarshichanda",
          "author_url": "",
          "post_date": "02/15/2022 08:27:42",
          "content": "<p><a href=\"https://www.kaggle.com/somesh88\" target=\"_blank\">@somesh88</a> I think the model didn't learn anything. I checked the distances and they are all close to 0.9. Will also check the distances for the tf model once I get TPU access</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1692660,
          "author_name": "mrinath",
          "author_url": "",
          "post_date": "02/16/2022 07:00:08",
          "content": "<p>One doubt on inference, like to get the final predictions we are comparing test embeddings with the trained ones right,<br>\nsuppose I have used 5 folds when training, to get the train embeddings of all the train samples, do I have to find embeddings using all the 5 models and average it? I may be wrong, whats the correct way?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1692941,
          "author_name": "ks2019",
          "author_url": "",
          "post_date": "02/16/2022 11:04:31",
          "content": "<p>What batch size do you use for training? Batch size is very important with ArcFace.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1692969,
          "author_name": "debarshichanda",
          "author_url": "",
          "post_date": "02/16/2022 11:24:10",
          "content": "<p><a href=\"https://www.kaggle.com/ks2019\" target=\"_blank\">@ks2019</a> I am just using a batch size of 32 so maybe that is a problem. I will try using a higher batch size and revert.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1692972,
          "author_name": "debarshichanda",
          "author_url": "",
          "post_date": "02/16/2022 11:26:26",
          "content": "<p><a href=\"https://www.kaggle.com/mrinath\" target=\"_blank\">@mrinath</a> I think there are 2 ways. One would be to average the embeddings and the other to concatenate. I think averaging would make more sense if they are same architecture trained on different folds and concatenating might be good for different architectures</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1692984,
          "author_name": "mrinath",
          "author_url": "",
          "post_date": "02/16/2022 11:33:02",
          "content": "<p><a href=\"https://www.kaggle.com/ks2019\" target=\"_blank\">@ks2019</a> more the batch size the better?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1693047,
          "author_name": "ks2019",
          "author_url": "",
          "post_date": "02/16/2022 12:09:41",
          "content": "<p><a href=\"https://www.kaggle.com/debarshichanda\" target=\"_blank\">@debarshichanda</a> I think 32 should be good to go, but try higher batch size once.<br>\n<a href=\"https://www.kaggle.com/mrinath\" target=\"_blank\">@mrinath</a> yes. With lower batch size, feature drift is higher. The ArcFace parameters (class centres) needs to be in sync with backbone. If for a class, we don't have any sample in many contiguous batches, we start observing a drift b/w feature representation and class centres representation. <br>\nNote that ArcFace doesn't employ sample-sample comparisons, but sample-prototype comparisons. The prototype stored in the ArcFace layer (or, the trainable parameter associated with ArcFace) represents the centre of each class. Training samples are enforced to be close to positive prototypes and far apart from negative prototypes by a clear margin (m). </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1689798,
      "author_name": "biglafe",
      "author_url": "",
      "post_date": "02/14/2022 13:59:40",
      "content": "<p>Using Pytorch training model  can reach LB 0.63 single model.  Configuration reference tf LB 0.60</p>",
      "votes": null,
      "replies": [
        {
          "id": 1689836,
          "author_name": "debarshichanda",
          "author_url": "",
          "post_date": "02/14/2022 14:28:59",
          "content": "<p>Can you elaborate on your inference pipeline?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1690151,
          "author_name": "mrinath",
          "author_url": "",
          "post_date": "02/14/2022 18:25:05",
          "content": "<p><a href=\"https://www.kaggle.com/biglafe\" target=\"_blank\">@biglafe</a> whats your CV score, and your final val loss(assuming you are using ArcFace)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1691542,
          "author_name": "biglafe",
          "author_url": "",
          "post_date": "02/15/2022 13:17:17",
          "content": "<p>val loss 9.6, inference pipeline complete references \b TF 0.6 <a href=\"https://www.kaggle.com/ks2019/happywhale-arcface-baseline-tpu\" target=\"_blank\">https://www.kaggle.com/ks2019/happywhale-arcface-baseline-tpu</a> by <a href=\"https://www.kaggle.com/ks2019\" target=\"_blank\">@ks2019</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1689803,
      "author_name": "biglafe",
      "author_url": "",
      "post_date": "02/14/2022 14:03:23",
      "content": "<p>By the way, the optimizer warmup is very important.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1690285,
          "author_name": "taindow",
          "author_url": "",
          "post_date": "02/14/2022 20:43:56",
          "content": "<p>Thanks for this hint! Are you using a metric learning approach similar to the tf public notebooks?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1691550,
          "author_name": "biglafe",
          "author_url": "",
          "post_date": "02/15/2022 13:21:19",
          "content": "<p>Most of the training and testing structures refer to tf public notebooks, but some learning rates and optimizer methods are adjusted. <br>\nThe measurement method also uses KNN to find similarity. Later, I will try some post-processing methods similar to GLR competition</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1706018,
          "author_name": "mrinath",
          "author_url": "",
          "post_date": "02/27/2022 05:09:50",
          "content": "<p>by optimizer warmup do you mean using LR schedulers? <a href=\"https://www.kaggle.com/biglafe\" target=\"_blank\">@biglafe</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1706024,
          "author_name": "biglafe",
          "author_url": "",
          "post_date": "02/27/2022 05:15:27",
          "content": "<p>Yes. When I use pytorch to train effb5 , it greatly affects my score.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1728699,
      "author_name": "haruki741",
      "author_url": "",
      "post_date": "03/19/2022 06:58:00",
      "content": "<p>I think so too.<br>\nI try to code using pytorch as same this Notebook 'kaggle kernels pull code/aikhmelnytskyy/happywhale-arcface-baseline-eff-net-kfold5-0-652' <a href=\"https://www.kaggle.com/Andrij\" target=\"_blank\">@Andrij</a> which is based on Tensorflow, but I got much more lower score than it.</p>\n<p>In addition, when I try to use tf_efficientnet_b7(like timm_create('tf_efficientnet_b7')) on colab TPU, it does not work.<br>\n(but in case of tf_efficientnet_b5, it works. I don't know why. It is bug?)<br>\nAnyway, for now, I think we should use Tensorflow if I want to use efficientnet on TPU envirionment.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1689180": "I am trying to replicate the high scoring TensorFlow notebooks in pytorch, but its nowhere near the score of the public notebooks.\nIs there some trick behind TF training? and has somebody been able to replicate those high scores on pytorch?\nNeed tips from you guys.\nThanks\n\nEdit\nwas trying to replicate this notebook https://www.kaggle.com/ks2019/happywhale-arcface-baseline-tpu by @ks2019",
    "1689205": "Could you please link to the notebooks that you're trying to replicate? \n\nThanks! :)",
    "1689244": "hey I've been also trying the same but so far no luck ( TPU with pytorch is bad )",
    "1689251": "I was trying to replicate this https://www.kaggle.com/ks2019/happywhale-arcface-baseline-tpu",
    "1689380": "One thing is I am just simply resizing to size 512, in the recent interview @cdeotte with @init27 said that doing this may harm the aspect ratio, I don't know how much this is important this time.",
    "1689393": "also, I fear random cropping images is not a good idea, as, in many images, random cropping will only result in the water, not the whale/dolphin, so don't know if its a good idea",
    "1689530": "Did you try to replicate the inferencing part as well?\nI tried my inference pipeline with my model and a baseline model trained with the notebook you mentioned and the results are quite different. So I am planning to replicate the inferencing part and compare the differences.\nWill update once I do the comparisons",
    "1689579": "Not the inferencing part though, was trying with the training part.",
    "1689653": "are you trying training with TPU or with GPU? \n\nagain I think he did trained on different platform with cross validation make sure you take that in your consideration while training notebook",
    "1689681": "there is dataset made with help of yolo it crops the images with the whales portion @awsaf49 thanks for dataset \n\n### this notebook might help you @mrinath \n[https://www.kaggle.com/awsaf49/happywhale-cropped-dataset-yolov5](https://www.kaggle.com/awsaf49/happywhale-cropped-dataset-yolov5)",
    "1689798": "Using Pytorch training model  can reach LB 0.63 single model.  Configuration reference tf LB 0.60",
    "1689803": "By the way, the optimizer warmup is very important.",
    "1689836": "Can you elaborate on your inference pipeline?",
    "1690151": "biglafe whats your CV score, and your final val loss(assuming you are using ArcFace)",
    "1690152": "training on GPU",
    "1690285": "Thanks for this hint! Are you using a metric learning approach similar to the tf public notebooks?",
    "1690287": "How different are we talking? Interested to understand if their inference process has something I'm missing. It seems to make a big difference ;)",
    "1690658": "I get an LB of 0.069 with PyTorch and 0.403 with tf. I am checking for bugs in my inference pipeline.\n\nThis is what I was trying to do\n1. Get the top 50 neighbors (distance, indexes of the samples)\n2. Now I get the classes of the samples (so now I have distances, classes)\n3. If the class is already added to the list, skip (continue)\n4. Else if distance > threshold, add the class to the list (distance refers to cosine similarity)\n5. Else add `new_individual`\n6. Once `new_invidual` is added, add the rest of the classes\n\n```\ndef get_preds(distances, indices, threshold=0.5):\n    classes = idx2class(indices)\n    \n    new = 0\n    preds = []\n    \n    for dist, class_ in zip(distances, classes):\n        if len(preds) == 5:\n            break\n        elif class_ in preds:\n            continue\n        elif dist > threshold:\n            preds.append(class_)\n        elif not new:\n            preds.append('new_individual')\n            new = 1\n        else:\n            preds.append(class_)\n            \n    return preds\n```",
    "1690710": "`I get an LB of 0.069 with PyTorch and 0.403 with tf. I am checking for bugs in my inference pipeline.`\nthe models were trained similarly?",
    "1690743": "Not exactly. Same backbone using arcface in both. Image sizes are different 448 for pytorch, 512 for tf",
    "1690927": "any idea @debarshichanda why it's happening?",
    "1691091": "somesh88 I think the model didn't learn anything. I checked the distances and they are all close to 0.9. Will also check the distances for the tf model once I get TPU access",
    "1691542": "val loss 9.6, inference pipeline complete references \b TF 0.6 https://www.kaggle.com/ks2019/happywhale-arcface-baseline-tpu by @ks2019",
    "1691550": "Most of the training and testing structures refer to tf public notebooks, but some learning rates and optimizer methods are adjusted. \nThe measurement method also uses KNN to find similarity. Later, I will try some post-processing methods similar to GLR competition",
    "1692660": "One doubt on inference, like to get the final predictions we are comparing test embeddings with the trained ones right,\nsuppose I have used 5 folds when training, to get the train embeddings of all the train samples, do I have to find embeddings using all the 5 models and average it? I may be wrong, whats the correct way?",
    "1692941": "What batch size do you use for training? Batch size is very important with ArcFace.",
    "1692969": "ks2019 I am just using a batch size of 32 so maybe that is a problem. I will try using a higher batch size and revert.",
    "1692972": "mrinath I think there are 2 ways. One would be to average the embeddings and the other to concatenate. I think averaging would make more sense if they are same architecture trained on different folds and concatenating might be good for different architectures",
    "1692984": "ks2019 more the batch size the better?",
    "1693047": "debarshichanda I think 32 should be good to go, but try higher batch size once.\n@mrinath yes. With lower batch size, feature drift is higher. The ArcFace parameters (class centres) needs to be in sync with backbone. If for a class, we don't have any sample in many contiguous batches, we start observing a drift b/w feature representation and class centres representation. \nNote that ArcFace doesn't employ sample-sample comparisons, but sample-prototype comparisons. The prototype stored in the ArcFace layer (or, the trainable parameter associated with ArcFace) represents the centre of each class. Training samples are enforced to be close to positive prototypes and far apart from negative prototypes by a clear margin (m).",
    "1706018": "by optimizer warmup do you mean using LR schedulers? @biglafe",
    "1706024": "Yes. When I use pytorch to train effb5 , it greatly affects my score.",
    "1728699": "I think so too.\nI try to code using pytorch as same this Notebook 'kaggle kernels pull code/aikhmelnytskyy/happywhale-arcface-baseline-eff-net-kfold5-0-652' @Andrij which is based on Tensorflow, but I got much more lower score than it.\n\nIn addition, when I try to use tf_efficientnet_b7(like timm_create('tf_efficientnet_b7')) on colab TPU, it does not work.\n(but in case of tf_efficientnet_b5, it works. I don't know why. It is bug?)\nAnyway, for now, I think we should use Tensorflow if I want to use efficientnet on TPU envirionment."
  },
  "source": "meta"
}