{
  "id": 315079,
  "title": "Tips for improving the score in my case",
  "url": "/competitions/happy-whale-and-dolphin/discussion/315079",
  "author_name": "kaggler",
  "post_date": "2022-03-26T04:52:44.056000",
  "votes": 65,
  "comment_count": 37,
  "views": 0,
  "content": "<p>I use TensorFlow framework in this competition<br>\nI want to share which parts help improve your score.</p>\n<ol>\n<li><p>Tensorflow<br>\nColab Tensorflow TPU is really fast. By using Tensorflow, You can really experiment with many things. I think learning tensorflow is better than using convenient PyTorch if you don't have many multiple GPUS.<br>\nI also have RTX 3090 but this isn't enough for this competition.</p></li>\n<li><p>Gem pooling<br>\nIn my case, gem pooling was slightly better than average pooling</p></li>\n<li><p>Arcface<br>\nI had tested other algorithms related to metric learning like Adacos, the Afcface is the best.</p></li>\n<li><p>Data &amp; Image size<br>\nWe can see a lot of datasets like original dataset of this competition, detic crop, backfin crop, full annotations, etc,,<br>\nI tested a lot and found an optimal one for me.<br>\nbut I have a plan to try backfin crop since the backfin crop hasn't been tried yet.<br>\nIt seems that the optimal size and algorithm are different for each data.<br>\nI would not tell what settings are best to be left for your try, but anyway, this part is the most crucial</p></li>\n<li><p>Batch Size<br>\nSince I am not sure about which part of batch size helps in model performance in deeply, anyway large Batch size worked for me</p></li>\n<li><p>Ensemble different architectures<br>\nFor example, when I ensembled efficientnet B5,B6, it boosted a score.<br>\nYou can concatenate embeddings from each model and run KNN algorithm to find matches.</p></li>\n</ol>\n<p>Still the listed are all I have tried and nothing deviates from here. (2022-03-26)<br>\nNow I am interested in post processing but currently i don't know how to do it </p>",
  "messages": [
    {
      "id": 1735284,
      "postDate": "2022-03-26T04:52:44.057Z",
      "content": "<p>I use TensorFlow framework in this competition<br>\nI want to share which parts help improve your score.</p>\n<ol>\n<li><p>Tensorflow<br>\nColab Tensorflow TPU is really fast. By using Tensorflow, You can really experiment with many things. I think learning tensorflow is better than using convenient PyTorch if you don't have many multiple GPUS.<br>\nI also have RTX 3090 but this isn't enough for this competition.</p></li>\n<li><p>Gem pooling<br>\nIn my case, gem pooling was slightly better than average pooling</p></li>\n<li><p>Arcface<br>\nI had tested other algorithms related to metric learning like Adacos, the Afcface is the best.</p></li>\n<li><p>Data &amp; Image size<br>\nWe can see a lot of datasets like original dataset of this competition, detic crop, backfin crop, full annotations, etc,,<br>\nI tested a lot and found an optimal one for me.<br>\nbut I have a plan to try backfin crop since the backfin crop hasn't been tried yet.<br>\nIt seems that the optimal size and algorithm are different for each data.<br>\nI would not tell what settings are best to be left for your try, but anyway, this part is the most crucial</p></li>\n<li><p>Batch Size<br>\nSince I am not sure about which part of batch size helps in model performance in deeply, anyway large Batch size worked for me</p></li>\n<li><p>Ensemble different architectures<br>\nFor example, when I ensembled efficientnet B5,B6, it boosted a score.<br>\nYou can concatenate embeddings from each model and run KNN algorithm to find matches.</p></li>\n</ol>\n<p>Still the listed are all I have tried and nothing deviates from here. (2022-03-26)<br>\nNow I am interested in post processing but currently i don't know how to do it </p>",
      "rawMarkdown": "I use TensorFlow framework in this competition\nI want to share which parts help improve your score.\n\n0. Tensorflow\nColab Tensorflow TPU is really fast. By using Tensorflow, You can really experiment with many things. I think learning tensorflow is better than using convenient PyTorch if you don't have many multiple GPUS.\nI also have RTX 3090 but this isn't enough for this competition.\n\n1. Gem pooling\nIn my case, gem pooling was slightly better than average pooling\n2. Arcface\nI had tested other algorithms related to metric learning like Adacos, the Afcface is the best.\n3. Data & Image size\nWe can see a lot of datasets like original dataset of this competition, detic crop, backfin crop, full annotations, etc,,\nI tested a lot and found an optimal one for me.\nbut I have a plan to try backfin crop since the backfin crop hasn't been tried yet.\nIt seems that the optimal size and algorithm are different for each data.\nI would not tell what settings are best to be left for your try, but anyway, this part is the most crucial\n4. Batch Size\nSince I am not sure about which part of batch size helps in model performance in deeply, anyway large Batch size worked for me\n4. Ensemble different architectures\nFor example, when I ensembled efficientnet B5,B6, it boosted a score.\nYou can concatenate embeddings from each model and run KNN algorithm to find matches.\n\nStill the listed are all I have tried and nothing deviates from here. (2022-03-26)\nNow I am interested in post processing but currently i don't know how to do it \n",
      "votes": 65
    },
    {
      "id": 1735331,
      "postDate": "2022-03-26T05:57:19.917Z",
      "content": "<p>Thank you for sharing. Comments to your post.</p>\n<ol>\n<li><p>Pytorch … I am sure people from TOP uses Pytorch. I know Tensorflow in this competition play important part of the game but … I prefer Pytorch over TF for many reasons …</p></li>\n<li><p>GemPooling - in my case we have better results using AvgPooling … but … you have better result so must check this again.</p></li>\n<li><p>I completly agree - main boost/jump we almost everytime get when we changed something in data.</p></li>\n</ol>",
      "rawMarkdown": "Thank you for sharing. Comments to your post.\n\n1. Pytorch ... I am sure people from TOP uses Pytorch. I know Tensorflow in this competition play important part of the game but ... I prefer Pytorch over TF for many reasons ...\n2. GemPooling - in my case we have better results using AvgPooling ... but ... you have better result so must check this again.\n\n4. I completly agree - main boost/jump we almost everytime get when we changed something in data.",
      "votes": 4,
      "replies": [
        {
          "id": 1735336,
          "postDate": "2022-03-26T06:03:10.387Z",
          "content": "<p>Can i ask you why many people prefer Pytorch over TF? I'm learning TF as my first framework and when it came to reading code, almost everyone use Pytorch. Is it about the accuracy or something. </p>",
          "rawMarkdown": "Can i ask you why many people prefer Pytorch over TF? I'm learning TF as my first framework and when it came to reading code, almost everyone use Pytorch. Is it about the accuracy or something. ",
          "votes": 2
        },
        {
          "id": 1735343,
          "postDate": "2022-03-26T06:13:23.613Z",
          "content": "<p>Flexibility …. \"Easy to do common things …. and still easy to do uncommon things\". TF \"Easy to do common things … and hard to do uncommon things\".</p>\n<p>And I feel that have control over everything ….</p>\n<p>TF … error … messages can be frustrating …. </p>",
          "rawMarkdown": "Flexibility .... \"Easy to do common things .... and still easy to do uncommon things\". TF \"Easy to do common things ... and hard to do uncommon things\".\n\nAnd I feel that have control over everything ....\n\nTF ... error ... messages can be frustrating .... ",
          "votes": 9
        },
        {
          "id": 1735387,
          "postDate": "2022-03-26T07:23:48.750Z",
          "content": "<p>In my case I also have better results using AvgPooling than Gem pooling….</p>",
          "rawMarkdown": "In my case I also have better results using AvgPooling than Gem pooling....",
          "votes": 2
        },
        {
          "id": 1735411,
          "postDate": "2022-03-26T08:05:43.367Z",
          "content": "<p><a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> I totally agree with you about  your thinking about pytorch. but with using tensorflow, at least we can try many things to check what option leads to a boost</p>",
          "rawMarkdown": "@remekkinas I totally agree with you about  your thinking about pytorch. but with using tensorflow, at least we can try many things to check what option leads to a boost"
        },
        {
          "id": 1735420,
          "postDate": "2022-03-26T08:18:47.267Z",
          "content": "<p>I was thinking this way until I started wasting my time solving puzzles in style “Segmentation fault #23339 error number #0098 at function blahblahnlahblah_impossible_to_check_blahh” </p>\n<p>Moreover I think that Pytorch community doing great job to provide us way to fast prototyping - lighting, timm, Catalyst, yolo series for CV … etc. etc. Look on TF object detection framework it is outdated … and what it is??? Joke? </p>\n<p>My first choice is Pytorch and spend a lot of time learning it. I use TF/Keras as well. I think that the ideal situation is then you learn and use both but … first choice is Pytorch for me.</p>",
          "rawMarkdown": "I was thinking this way until I started wasting my time solving puzzles in style “Segmentation fault #23339 error number #0098 at function blahblahnlahblah_impossible_to_check_blahh” \n\nMoreover I think that Pytorch community doing great job to provide us way to fast prototyping - lighting, timm, Catalyst, yolo series for CV … etc. etc. Look on TF object detection framework it is outdated … and what it is??? Joke? \n\nMy first choice is Pytorch and spend a lot of time learning it. I use TF/Keras as well. I think that the ideal situation is then you learn and use both but … first choice is Pytorch for me.",
          "votes": 1
        },
        {
          "id": 1735425,
          "postDate": "2022-03-26T08:38:04.460Z",
          "content": "<p>But sadly Pytorch + TPU is a nightmare</p>",
          "rawMarkdown": "But sadly Pytorch + TPU is a nightmare",
          "votes": 3
        },
        {
          "id": 1735442,
          "postDate": "2022-03-26T09:04:07.013Z",
          "content": "<p>Another one news … from tomorrow - Pytorch - Avalanche: <a href=\"https://medium.com/pytorch/avalanche-and-end-to-end-library-for-continual-learning-based-on-pytorch-a99cf5661a0d\" target=\"_blank\">https://medium.com/pytorch/avalanche-and-end-to-end-library-for-continual-learning-based-on-pytorch-a99cf5661a0d</a></p>",
          "rawMarkdown": "Another one news ... from tomorrow - Pytorch - Avalanche: https://medium.com/pytorch/avalanche-and-end-to-end-library-for-continual-learning-based-on-pytorch-a99cf5661a0d",
          "votes": 1
        },
        {
          "id": 1735598,
          "postDate": "2022-03-26T12:36:39.810Z",
          "content": "<p><a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> i'm deeply interested in the catastrophic forgetting when finetuning models  <br>\nthis can be a good source for me ! thanks!</p>",
          "rawMarkdown": "@remekkinas i'm deeply interested in the catastrophic forgetting when finetuning models  \nthis can be a good source for me ! thanks!"
        },
        {
          "id": 1736034,
          "postDate": "2022-03-26T21:02:11.613Z",
          "content": "<p>Good post Pytorch vs TF …. <a href=\"https://www.assemblyai.com/blog/pytorch-vs-tensorflow-in-2022/\" target=\"_blank\">https://www.assemblyai.com/blog/pytorch-vs-tensorflow-in-2022/</a></p>",
          "rawMarkdown": "Good post Pytorch vs TF …. https://www.assemblyai.com/blog/pytorch-vs-tensorflow-in-2022/",
          "votes": 2
        },
        {
          "id": 1737386,
          "postDate": "2022-03-28T11:50:39.383Z",
          "content": "<p>77% of Deep Learning solutions used PyTorch (up from 72% last year)</p>\n<p><a href=\"https://mlcontests.com/\" target=\"_blank\">https://mlcontests.com/</a></p>",
          "rawMarkdown": "77% of Deep Learning solutions used PyTorch (up from 72% last year)\n\nhttps://mlcontests.com/\n",
          "votes": 1
        },
        {
          "id": 1738556,
          "postDate": "2022-03-29T11:18:03.977Z",
          "content": "<p>Hello, how to use pytorch in TPU?</p>",
          "rawMarkdown": "Hello, how to use pytorch in TPU?"
        },
        {
          "id": 1739705,
          "postDate": "2022-03-30T09:20:10.713Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1741764,
          "postDate": "2022-04-01T06:06:40.520Z",
          "content": "<p>When I tried it, I found it isn't faster than GPU….</p>",
          "rawMarkdown": "When I tried it, I found it isn't faster than GPU....",
          "votes": 1
        }
      ]
    },
    {
      "id": 1744425,
      "postDate": "2022-04-04T01:13:31.517Z",
      "content": "<p>Thank you for your sharing. Waiting for update in future!</p>",
      "rawMarkdown": "Thank you for your sharing. Waiting for update in future!",
      "votes": 1
    },
    {
      "id": 1737323,
      "postDate": "2022-03-28T10:54:54.897Z",
      "content": "<p>Thanks for the post! I just wanna ask how you approached the ensembling of different architectures. Wanted to try it out myself for a different competition but not really sure as to where to start from </p>",
      "rawMarkdown": "Thanks for the post! I just wanna ask how you approached the ensembling of different architectures. Wanted to try it out myself for a different competition but not really sure as to where to start from ",
      "votes": 1,
      "replies": [
        {
          "id": 1737614,
          "postDate": "2022-03-28T15:00:32.537Z",
          "content": "<p>If you use arcface,<br>\nfor example, a network can be like this<br>\ninput -&gt; Efficientnet -&gt; Global Average pooling -&gt; Dense(512) -&gt; Arcface<br>\nYou can take the output of dense(512), then you have 512-dimensional embeddings<br>\nyou can train another model then again take 512-dimensional embeddings<br>\nthen you will have two 512-dimensional embeddings<br>\nyou can concatenate two embeddings, then embedding becomes 1024-dimensional embeddings<br>\nwith this, you can use KNN to find matches</p>",
          "rawMarkdown": "If you use arcface,\nfor example, a network can be like this\ninput -> Efficientnet -> Global Average pooling -> Dense(512) -> Arcface\nYou can take the output of dense(512), then you have 512-dimensional embeddings\nyou can train another model then again take 512-dimensional embeddings\nthen you will have two 512-dimensional embeddings\nyou can concatenate two embeddings, then embedding becomes 1024-dimensional embeddings\nwith this, you can use KNN to find matches",
          "votes": 10
        },
        {
          "id": 1737801,
          "postDate": "2022-03-28T18:08:34Z",
          "content": "<p>I have been concatenating the embeddings on the WRONG axis for so long. Thanks for the pointer :). Now I can see the major boost from +0.001 (wrong concatenation) to +0.008 (correct concatenation).</p>",
          "rawMarkdown": "I have been concatenating the embeddings on the WRONG axis for so long. Thanks for the pointer :). Now I can see the major boost from +0.001 (wrong concatenation) to +0.008 (correct concatenation).",
          "votes": 3
        },
        {
          "id": 1738143,
          "postDate": "2022-03-29T04:14:34.387Z",
          "content": "<p>Oh whao, utilizing KNN in this scenario? Is that common? Sorry for sounding very lost cause I just got started in deep learning and don't have much experience with image data.</p>",
          "rawMarkdown": "Oh whao, utilizing KNN in this scenario? Is that common? Sorry for sounding very lost cause I just got started in deep learning and don't have much experience with image data.",
          "votes": 1
        },
        {
          "id": 1738156,
          "postDate": "2022-03-29T04:28:35.680Z",
          "content": "<p><a href=\"https://www.kaggle.com/kimmik123\" target=\"_blank\">@kimmik123</a> can use either the softmax outcome or KNN by using embeddings.  In my case, KNN is better <br>\n<a href=\"https://www.kaggle.com/thakurudit\" target=\"_blank\">@thakurudit</a> nice improvement! keep going!</p>",
          "rawMarkdown": "@kimmik123 can use either the softmax outcome or KNN by using embeddings.  In my case, KNN is better \n@thakurudit nice improvement! keep going!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1736132,
      "postDate": "2022-03-27T02:49:28.933Z",
      "content": "<p>Thanks for sharing,  in my case better results using AvgPooling too. I am also use the colab + TPU and it is very help in this competition.</p>",
      "rawMarkdown": "Thanks for sharing,  in my case better results using AvgPooling too. I am also use the colab + TPU and it is very help in this competition.",
      "votes": 1
    },
    {
      "id": 1735440,
      "postDate": "2022-03-26T09:01:25.230Z",
      "content": "<p>Thanks for sharing the information.<br>\nAccording to my experiment, Avg Pooling performed better than GeM Pooling.<br>\nI will give it another try!</p>",
      "rawMarkdown": "Thanks for sharing the information.\nAccording to my experiment, Avg Pooling performed better than GeM Pooling.\nI will give it another try!",
      "votes": 1,
      "replies": [
        {
          "id": 1735607,
          "postDate": "2022-03-26T12:42:06.833Z",
          "content": "<p>I also want to know why the avg pooling doesn't work in my case!<br>\nGood Luck!</p>",
          "rawMarkdown": "I also want to know why the avg pooling doesn't work in my case!\nGood Luck!",
          "votes": 1
        },
        {
          "id": 1750740,
          "postDate": "2022-04-10T03:16:48.953Z",
          "content": "<p>I experimented again and found that GeM is better in some conditions.<br>\nOnce again, thank you for sharing this great information!</p>",
          "rawMarkdown": "I experimented again and found that GeM is better in some conditions.\nOnce again, thank you for sharing this great information!"
        }
      ]
    },
    {
      "id": 1743676,
      "postDate": "2022-04-03T07:47:16.730Z",
      "content": "<p><a href=\"https://www.kaggle.com/deepkim\" target=\"_blank\">@deepkim</a> Is it possible to train a few models and concat their embeddings and then take those as our inputs for neural network with dropout and arcface to get better accuracy on validation data? I had this thought and want to give it a try.<br>\nIf it is possible can anyone confirm?</p>\n<p>The other thing I thought was that an we make different models specific to the data for eg. dolphins and whales and the we generate embeddings specific to them and at the time of prediction we can rely on the threshold value to select the model to use for each specific test image so to get good prediction.</p>",
      "rawMarkdown": "@deepkim Is it possible to train a few models and concat their embeddings and then take those as our inputs for neural network with dropout and arcface to get better accuracy on validation data? I had this thought and want to give it a try.\nIf it is possible can anyone confirm?\n\nThe other thing I thought was that an we make different models specific to the data for eg. dolphins and whales and the we generate embeddings specific to them and at the time of prediction we can rely on the threshold value to select the model to use for each specific test image so to get good prediction.",
      "replies": [
        {
          "id": 1745124,
          "postDate": "2022-04-04T16:11:43.387Z",
          "content": "<p><a href=\"https://www.kaggle.com/jainishsavalia\" target=\"_blank\">@jainishsavalia</a> the second thing seems smart. but for the first one, i have no idea about it</p>",
          "rawMarkdown": "@jainishsavalia the second thing seems smart. but for the first one, i have no idea about it"
        }
      ]
    },
    {
      "id": 1737594,
      "postDate": "2022-03-28T14:51:32.463Z",
      "content": "<p>I didn't see major improvements in ensembling embeddings from different models - only a low boost of around +0.001.</p>",
      "rawMarkdown": "I didn't see major improvements in ensembling embeddings from different models - only a low boost of around +0.001.",
      "replies": [
        {
          "id": 1737619,
          "postDate": "2022-03-28T15:04:53.680Z",
          "content": "<p>For me, it boosted my score a lot</p>",
          "rawMarkdown": "For me, it boosted my score a lot\n",
          "votes": 1
        },
        {
          "id": 1737622,
          "postDate": "2022-03-28T15:06:58.613Z",
          "content": "<p>I tried two things:-</p>\n<ol>\n<li>Taking mean of all the embeddings np.mean(model 1+model 2): +0.001</li>\n<li>Concatenating both the embeddings: +0.001 (NVM was doing it wrong lmao)</li>\n</ol>\n<p>Will try the correct way.</p>",
          "rawMarkdown": "I tried two things:-\n1. Taking mean of all the embeddings np.mean(model 1+model 2): +0.001\n2. Concatenating both the embeddings: +0.001 (NVM was doing it wrong lmao)\n\nWill try the correct way.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1737520,
      "postDate": "2022-03-28T13:56:58.213Z",
      "content": "<p>Any reason why \"gem pooling was slightly better than average pooling\" ?</p>",
      "rawMarkdown": "Any reason why \"gem pooling was slightly better than average pooling\" ?",
      "replies": [
        {
          "id": 1745122,
          "postDate": "2022-04-04T16:10:04.857Z",
          "content": "<p><a href=\"https://www.kaggle.com/mohandass\" target=\"_blank\">@mohandass</a> sorry for the late reply, <br>\nI think gem pooling works when local features matter<br>\nIn the landmark competition held last year, many used gem pooling. so i adopted it!</p>",
          "rawMarkdown": "@mohandass sorry for the late reply, \nI think gem pooling works when local features matter\nIn the landmark competition held last year, many used gem pooling. so i adopted it!"
        }
      ]
    },
    {
      "id": 1736221,
      "postDate": "2022-03-27T05:42:34.733Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 1737378,
      "postDate": "2022-03-28T11:45:23.167Z",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing.",
      "votes": 1
    },
    {
      "id": 1736215,
      "postDate": "2022-03-27T05:33:27.543Z",
      "content": "<p>Thanks for ideas</p>",
      "rawMarkdown": "Thanks for ideas",
      "votes": 1
    },
    {
      "id": 1735286,
      "postDate": "2022-03-26T04:54:06.460Z",
      "content": "<p>Thanks for ideas👏</p>",
      "rawMarkdown": "Thanks for ideas👏",
      "votes": 1
    },
    {
      "id": 1751216,
      "postDate": "2022-04-10T14:07:49.263Z",
      "content": "<p>Thanks a lot for sharing this!</p>",
      "rawMarkdown": "Thanks a lot for sharing this!"
    },
    {
      "id": 1743013,
      "postDate": "2022-04-02T14:00:10.477Z",
      "content": "<p>Thanks for your insight!</p>",
      "rawMarkdown": "Thanks for your insight!"
    }
  ],
  "comments": [
    {
      "id": 1735331,
      "author_name": "Remek Kinas",
      "author_url": "",
      "post_date": "2022-03-26T05:57:19.917000",
      "content": "<p>Thank you for sharing. Comments to your post.</p>\n<ol>\n<li><p>Pytorch … I am sure people from TOP uses Pytorch. I know Tensorflow in this competition play important part of the game but … I prefer Pytorch over TF for many reasons …</p></li>\n<li><p>GemPooling - in my case we have better results using AvgPooling … but … you have better result so must check this again.</p></li>\n<li><p>I completly agree - main boost/jump we almost everytime get when we changed something in data.</p></li>\n</ol>",
      "votes": 4,
      "replies": [
        {
          "id": 1735336,
          "author_name": "Bao Loc Pham",
          "author_url": "",
          "post_date": "2022-03-26T06:03:10.387000",
          "content": "<p>Can i ask you why many people prefer Pytorch over TF? I'm learning TF as my first framework and when it came to reading code, almost everyone use Pytorch. Is it about the accuracy or something. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1735343,
          "author_name": "Remek Kinas",
          "author_url": "",
          "post_date": "2022-03-26T06:13:23.613000",
          "content": "<p>Flexibility …. \"Easy to do common things …. and still easy to do uncommon things\". TF \"Easy to do common things … and hard to do uncommon things\".</p>\n<p>And I feel that have control over everything ….</p>\n<p>TF … error … messages can be frustrating …. </p>",
          "votes": 9,
          "replies": []
        },
        {
          "id": 1735387,
          "author_name": "老肥",
          "author_url": "",
          "post_date": "2022-03-26T07:23:48.750000",
          "content": "<p>In my case I also have better results using AvgPooling than Gem pooling….</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1735411,
          "author_name": "kaggler",
          "author_url": "",
          "post_date": "2022-03-26T08:05:43.367000",
          "content": "<p><a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> I totally agree with you about  your thinking about pytorch. but with using tensorflow, at least we can try many things to check what option leads to a boost</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1735420,
          "author_name": "Remek Kinas",
          "author_url": "",
          "post_date": "2022-03-26T08:18:47.267000",
          "content": "<p>I was thinking this way until I started wasting my time solving puzzles in style “Segmentation fault #23339 error number #0098 at function blahblahnlahblah_impossible_to_check_blahh” </p>\n<p>Moreover I think that Pytorch community doing great job to provide us way to fast prototyping - lighting, timm, Catalyst, yolo series for CV … etc. etc. Look on TF object detection framework it is outdated … and what it is??? Joke? </p>\n<p>My first choice is Pytorch and spend a lot of time learning it. I use TF/Keras as well. I think that the ideal situation is then you learn and use both but … first choice is Pytorch for me.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1735425,
          "author_name": "Phat Tran",
          "author_url": "",
          "post_date": "2022-03-26T08:38:04.460000",
          "content": "<p>But sadly Pytorch + TPU is a nightmare</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1735442,
          "author_name": "Remek Kinas",
          "author_url": "",
          "post_date": "2022-03-26T09:04:07.013000",
          "content": "<p>Another one news … from tomorrow - Pytorch - Avalanche: <a href=\"https://medium.com/pytorch/avalanche-and-end-to-end-library-for-continual-learning-based-on-pytorch-a99cf5661a0d\" target=\"_blank\">https://medium.com/pytorch/avalanche-and-end-to-end-library-for-continual-learning-based-on-pytorch-a99cf5661a0d</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1735598,
          "author_name": "kaggler",
          "author_url": "",
          "post_date": "2022-03-26T12:36:39.810000",
          "content": "<p><a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> i'm deeply interested in the catastrophic forgetting when finetuning models  <br>\nthis can be a good source for me ! thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1736034,
          "author_name": "Remek Kinas",
          "author_url": "",
          "post_date": "2022-03-26T21:02:11.613000",
          "content": "<p>Good post Pytorch vs TF …. <a href=\"https://www.assemblyai.com/blog/pytorch-vs-tensorflow-in-2022/\" target=\"_blank\">https://www.assemblyai.com/blog/pytorch-vs-tensorflow-in-2022/</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1737386,
          "author_name": "Yangranran",
          "author_url": "",
          "post_date": "2022-03-28T11:50:39.383000",
          "content": "<p>77% of Deep Learning solutions used PyTorch (up from 72% last year)</p>\n<p><a href=\"https://mlcontests.com/\" target=\"_blank\">https://mlcontests.com/</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1738556,
          "author_name": "Zekun",
          "author_url": "",
          "post_date": "2022-03-29T11:18:03.977000",
          "content": "<p>Hello, how to use pytorch in TPU?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1739705,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-03-30T09:20:10.713000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1741764,
          "author_name": "Zekun",
          "author_url": "",
          "post_date": "2022-04-01T06:06:40.520000",
          "content": "<p>When I tried it, I found it isn't faster than GPU….</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1744425,
      "author_name": "Tan Phan",
      "author_url": "",
      "post_date": "2022-04-04T01:13:31.517000",
      "content": "<p>Thank you for your sharing. Waiting for update in future!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1737323,
      "author_name": "Kim Hyun Bin",
      "author_url": "",
      "post_date": "2022-03-28T10:54:54.897000",
      "content": "<p>Thanks for the post! I just wanna ask how you approached the ensembling of different architectures. Wanted to try it out myself for a different competition but not really sure as to where to start from </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1737614,
          "author_name": "kaggler",
          "author_url": "",
          "post_date": "2022-03-28T15:00:32.537000",
          "content": "<p>If you use arcface,<br>\nfor example, a network can be like this<br>\ninput -&gt; Efficientnet -&gt; Global Average pooling -&gt; Dense(512) -&gt; Arcface<br>\nYou can take the output of dense(512), then you have 512-dimensional embeddings<br>\nyou can train another model then again take 512-dimensional embeddings<br>\nthen you will have two 512-dimensional embeddings<br>\nyou can concatenate two embeddings, then embedding becomes 1024-dimensional embeddings<br>\nwith this, you can use KNN to find matches</p>",
          "votes": 10,
          "replies": []
        },
        {
          "id": 1737801,
          "author_name": "ilovepotatoes",
          "author_url": "",
          "post_date": "2022-03-28T18:08:34",
          "content": "<p>I have been concatenating the embeddings on the WRONG axis for so long. Thanks for the pointer :). Now I can see the major boost from +0.001 (wrong concatenation) to +0.008 (correct concatenation).</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1738143,
          "author_name": "Kim Hyun Bin",
          "author_url": "",
          "post_date": "2022-03-29T04:14:34.387000",
          "content": "<p>Oh whao, utilizing KNN in this scenario? Is that common? Sorry for sounding very lost cause I just got started in deep learning and don't have much experience with image data.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1738156,
          "author_name": "kaggler",
          "author_url": "",
          "post_date": "2022-03-29T04:28:35.680000",
          "content": "<p><a href=\"https://www.kaggle.com/kimmik123\" target=\"_blank\">@kimmik123</a> can use either the softmax outcome or KNN by using embeddings.  In my case, KNN is better <br>\n<a href=\"https://www.kaggle.com/thakurudit\" target=\"_blank\">@thakurudit</a> nice improvement! keep going!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1736132,
      "author_name": "liuzhangzhen",
      "author_url": "",
      "post_date": "2022-03-27T02:49:28.933000",
      "content": "<p>Thanks for sharing,  in my case better results using AvgPooling too. I am also use the colab + TPU and it is very help in this competition.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1735440,
      "author_name": "k_s",
      "author_url": "",
      "post_date": "2022-03-26T09:01:25.230000",
      "content": "<p>Thanks for sharing the information.<br>\nAccording to my experiment, Avg Pooling performed better than GeM Pooling.<br>\nI will give it another try!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1735607,
          "author_name": "kaggler",
          "author_url": "",
          "post_date": "2022-03-26T12:42:06.833000",
          "content": "<p>I also want to know why the avg pooling doesn't work in my case!<br>\nGood Luck!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1750740,
          "author_name": "k_s",
          "author_url": "",
          "post_date": "2022-04-10T03:16:48.953000",
          "content": "<p>I experimented again and found that GeM is better in some conditions.<br>\nOnce again, thank you for sharing this great information!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1743676,
      "author_name": "Jainish Savalia",
      "author_url": "",
      "post_date": "2022-04-03T07:47:16.730000",
      "content": "<p><a href=\"https://www.kaggle.com/deepkim\" target=\"_blank\">@deepkim</a> Is it possible to train a few models and concat their embeddings and then take those as our inputs for neural network with dropout and arcface to get better accuracy on validation data? I had this thought and want to give it a try.<br>\nIf it is possible can anyone confirm?</p>\n<p>The other thing I thought was that an we make different models specific to the data for eg. dolphins and whales and the we generate embeddings specific to them and at the time of prediction we can rely on the threshold value to select the model to use for each specific test image so to get good prediction.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1745124,
          "author_name": "kaggler",
          "author_url": "",
          "post_date": "2022-04-04T16:11:43.387000",
          "content": "<p><a href=\"https://www.kaggle.com/jainishsavalia\" target=\"_blank\">@jainishsavalia</a> the second thing seems smart. but for the first one, i have no idea about it</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1737594,
      "author_name": "ilovepotatoes",
      "author_url": "",
      "post_date": "2022-03-28T14:51:32.463000",
      "content": "<p>I didn't see major improvements in ensembling embeddings from different models - only a low boost of around +0.001.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1737619,
          "author_name": "kaggler",
          "author_url": "",
          "post_date": "2022-03-28T15:04:53.680000",
          "content": "<p>For me, it boosted my score a lot</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1737622,
          "author_name": "ilovepotatoes",
          "author_url": "",
          "post_date": "2022-03-28T15:06:58.613000",
          "content": "<p>I tried two things:-</p>\n<ol>\n<li>Taking mean of all the embeddings np.mean(model 1+model 2): +0.001</li>\n<li>Concatenating both the embeddings: +0.001 (NVM was doing it wrong lmao)</li>\n</ol>\n<p>Will try the correct way.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1737520,
      "author_name": "mohandass",
      "author_url": "",
      "post_date": "2022-03-28T13:56:58.213000",
      "content": "<p>Any reason why \"gem pooling was slightly better than average pooling\" ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1745122,
          "author_name": "kaggler",
          "author_url": "",
          "post_date": "2022-04-04T16:10:04.857000",
          "content": "<p><a href=\"https://www.kaggle.com/mohandass\" target=\"_blank\">@mohandass</a> sorry for the late reply, <br>\nI think gem pooling works when local features matter<br>\nIn the landmark competition held last year, many used gem pooling. so i adopted it!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1736221,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-03-27T05:42:34.733000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1737378,
      "author_name": "Yangranran",
      "author_url": "",
      "post_date": "2022-03-28T11:45:23.167000",
      "content": "<p>Thanks for sharing.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1736215,
      "author_name": "Fatma Gaber",
      "author_url": "",
      "post_date": "2022-03-27T05:33:27.543000",
      "content": "<p>Thanks for ideas</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1735286,
      "author_name": "AlanHaBrony",
      "author_url": "",
      "post_date": "2022-03-26T04:54:06.460000",
      "content": "<p>Thanks for ideas👏</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1751216,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-04-10T14:07:49.263000",
      "content": "<p>Thanks a lot for sharing this!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1743013,
      "author_name": "ME2MLE",
      "author_url": "",
      "post_date": "2022-04-02T14:00:10.477000",
      "content": "<p>Thanks for your insight!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1735284": "I use TensorFlow framework in this competition\nI want to share which parts help improve your score.\n\n0. Tensorflow\nColab Tensorflow TPU is really fast. By using Tensorflow, You can really experiment with many things. I think learning tensorflow is better than using convenient PyTorch if you don't have many multiple GPUS.\nI also have RTX 3090 but this isn't enough for this competition.\n\n1. Gem pooling\nIn my case, gem pooling was slightly better than average pooling\n2. Arcface\nI had tested other algorithms related to metric learning like Adacos, the Afcface is the best.\n3. Data & Image size\nWe can see a lot of datasets like original dataset of this competition, detic crop, backfin crop, full annotations, etc,,\nI tested a lot and found an optimal one for me.\nbut I have a plan to try backfin crop since the backfin crop hasn't been tried yet.\nIt seems that the optimal size and algorithm are different for each data.\nI would not tell what settings are best to be left for your try, but anyway, this part is the most crucial\n4. Batch Size\nSince I am not sure about which part of batch size helps in model performance in deeply, anyway large Batch size worked for me\n4. Ensemble different architectures\nFor example, when I ensembled efficientnet B5,B6, it boosted a score.\nYou can concatenate embeddings from each model and run KNN algorithm to find matches.\n\nStill the listed are all I have tried and nothing deviates from here. (2022-03-26)\nNow I am interested in post processing but currently i don't know how to do it \n",
    "1735331": "Thank you for sharing. Comments to your post.\n\n1. Pytorch ... I am sure people from TOP uses Pytorch. I know Tensorflow in this competition play important part of the game but ... I prefer Pytorch over TF for many reasons ...\n2. GemPooling - in my case we have better results using AvgPooling ... but ... you have better result so must check this again.\n\n4. I completly agree - main boost/jump we almost everytime get when we changed something in data.",
    "1744425": "Thank you for your sharing. Waiting for update in future!",
    "1737323": "Thanks for the post! I just wanna ask how you approached the ensembling of different architectures. Wanted to try it out myself for a different competition but not really sure as to where to start from ",
    "1736132": "Thanks for sharing,  in my case better results using AvgPooling too. I am also use the colab + TPU and it is very help in this competition.",
    "1735440": "Thanks for sharing the information.\nAccording to my experiment, Avg Pooling performed better than GeM Pooling.\nI will give it another try!",
    "1743676": "@deepkim Is it possible to train a few models and concat their embeddings and then take those as our inputs for neural network with dropout and arcface to get better accuracy on validation data? I had this thought and want to give it a try.\nIf it is possible can anyone confirm?\n\nThe other thing I thought was that an we make different models specific to the data for eg. dolphins and whales and the we generate embeddings specific to them and at the time of prediction we can rely on the threshold value to select the model to use for each specific test image so to get good prediction.",
    "1737594": "I didn't see major improvements in ensembling embeddings from different models - only a low boost of around +0.001.",
    "1737520": "Any reason why \"gem pooling was slightly better than average pooling\" ?",
    "1736221": "",
    "1737378": "Thanks for sharing.",
    "1736215": "Thanks for ideas",
    "1735286": "Thanks for ideas👏",
    "1751216": "Thanks a lot for sharing this!",
    "1743013": "Thanks for your insight!"
  }
}