{
  "id": 316555,
  "title": " 7 things that did not worked so far",
  "url": "/competitions/happy-whale-and-dolphin/discussion/316555",
  "author_name": "",
  "post_date": "2022-04-02T15:46:20.936288400Z",
  "votes": 19,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Among other things that give me decent result, here are 7 things that for me did not work, they either give the same cv/public leaderboard result as without them or lower.</p>\n<p><strong>1. SWA (Stochastic Weight Averaging)</strong><br>\nAlthough in the past my experience with SWA was positive, in this context it seems not to improve the results (for a fold with SWA 0.714 vs same fold without SWA 0.728)</p>\n<p>Intuition for SWA comes from empirical observation that local minima at the end of each learning rate cycle tend to accumulate at the border of areas on loss surface where loss value is low . By taking the average of several such points, it is possible to achieve a wide, generalizable solution with even lower loss.</p>\n<p><strong>2. EffnetsV2</strong><br>\nI have tried efficientnetv2-s-21k-ft1k and efficientnetv2-l-21k-ft1k. Probably with some fine tune it will provide good results, but all my tunning so far was made on EffnetV1</p>\n<p><strong>3. Concat</strong><br>\nStrangely it may sound, the concat ensembling is not working for me, I am 90% sure that is something that I do wrong but still did not figure out exactly what. When I take 2 models, each with 512 embeddings per image and create 1024 embeddings for an image the results are worse (with and without data normalization). Mean ensembling and Max ensembling are working much better for me. </p>\n<p>Mean ensembling<br>\nFor each image I have selected the best 100 predictions confidence from each model.<br>\nThen average confidences across all models and choose the ones with the highest mean of confidences</p>\n<p>Max ensembling<br>\nFor each image I have selected the best 100 predictions confidence from each model.<br>\nThen choose the top 5 predictions by the max confidence across all models.</p>\n<p><strong>4. EffnetsV1 lower than B4</strong><br>\nContext of the competition advantages bigger architectures. My best results were using EffNets B5 to B7 and ConvNext's architectures</p>\n<p><strong>5. Gradient accumulation</strong><br>\nGradient accumulation means running a configured number of steps without updating the model variables while accumulating the gradients of those steps and then using the accumulated gradients to compute the variable updates.<br>\nSounds simple and effective but it never worked for me. It always messes up batch normalization more than it helps with increasing gradients robustness</p>\n<p><strong>6. Combining multiple datasets for a single model in order to increase robustness</strong><br>\nThe idea was to merge the fins dataset with a dataset that is having the whole whale in the image for learning to recognize in both scenarios. For now, the results are not positive but I still have some ideas to test</p>\n<p><strong>7. Original dataset</strong><br>\nEither using the fins dataset or another whale cropping dataset, the results will be much better than with the original one. Noisy data present in large scales in images (sea, humans, boats, etc.) are severely affecting the classifier ability to recognize individuals. One interesting aspect here is that Pytorch approaches seems to run much better using the fins dataset, when TF using TPU's are not so sensitive to datasets (fins vs whole whale cropped)</p>",
  "messages": [
    {
      "id": "1743087",
      "postDate": "04/02/2022 15:46:20",
      "content": "<p>Among other things that give me decent result, here are 7 things that for me did not work, they either give the same cv/public leaderboard result as without them or lower.</p>\n<p><strong>1. SWA (Stochastic Weight Averaging)</strong><br>\nAlthough in the past my experience with SWA was positive, in this context it seems not to improve the results (for a fold with SWA 0.714 vs same fold without SWA 0.728)</p>\n<p>Intuition for SWA comes from empirical observation that local minima at the end of each learning rate cycle tend to accumulate at the border of areas on loss surface where loss value is low . By taking the average of several such points, it is possible to achieve a wide, generalizable solution with even lower loss.</p>\n<p><strong>2. EffnetsV2</strong><br>\nI have tried efficientnetv2-s-21k-ft1k and efficientnetv2-l-21k-ft1k. Probably with some fine tune it will provide good results, but all my tunning so far was made on EffnetV1</p>\n<p><strong>3. Concat</strong><br>\nStrangely it may sound, the concat ensembling is not working for me, I am 90% sure that is something that I do wrong but still did not figure out exactly what. When I take 2 models, each with 512 embeddings per image and create 1024 embeddings for an image the results are worse (with and without data normalization). Mean ensembling and Max ensembling are working much better for me. </p>\n<p>Mean ensembling<br>\nFor each image I have selected the best 100 predictions confidence from each model.<br>\nThen average confidences across all models and choose the ones with the highest mean of confidences</p>\n<p>Max ensembling<br>\nFor each image I have selected the best 100 predictions confidence from each model.<br>\nThen choose the top 5 predictions by the max confidence across all models.</p>\n<p><strong>4. EffnetsV1 lower than B4</strong><br>\nContext of the competition advantages bigger architectures. My best results were using EffNets B5 to B7 and ConvNext's architectures</p>\n<p><strong>5. Gradient accumulation</strong><br>\nGradient accumulation means running a configured number of steps without updating the model variables while accumulating the gradients of those steps and then using the accumulated gradients to compute the variable updates.<br>\nSounds simple and effective but it never worked for me. It always messes up batch normalization more than it helps with increasing gradients robustness</p>\n<p><strong>6. Combining multiple datasets for a single model in order to increase robustness</strong><br>\nThe idea was to merge the fins dataset with a dataset that is having the whole whale in the image for learning to recognize in both scenarios. For now, the results are not positive but I still have some ideas to test</p>\n<p><strong>7. Original dataset</strong><br>\nEither using the fins dataset or another whale cropping dataset, the results will be much better than with the original one. Noisy data present in large scales in images (sea, humans, boats, etc.) are severely affecting the classifier ability to recognize individuals. One interesting aspect here is that Pytorch approaches seems to run much better using the fins dataset, when TF using TPU's are not so sensitive to datasets (fins vs whole whale cropped)</p>",
      "rawMarkdown": "Among other things that give me decent result, here are 7 things that for me did not work, they either give the same cv/public leaderboard result as without them or lower.\n\n**1. SWA (Stochastic Weight Averaging)**\nAlthough in the past my experience with SWA was positive, in this context it seems not to improve the results (for a fold with SWA 0.714 vs same fold without SWA 0.728)\n\nIntuition for SWA comes from empirical observation that local minima at the end of each learning rate cycle tend to accumulate at the border of areas on loss surface where loss value is low . By taking the average of several such points, it is possible to achieve a wide, generalizable solution with even lower loss.\n\n**2. EffnetsV2**\nI have tried efficientnetv2-s-21k-ft1k and efficientnetv2-l-21k-ft1k. Probably with some fine tune it will provide good results, but all my tunning so far was made on EffnetV1\n\n**3. Concat**\nStrangely it may sound, the concat ensembling is not working for me, I am 90% sure that is something that I do wrong but still did not figure out exactly what. When I take 2 models, each with 512 embeddings per image and create 1024 embeddings for an image the results are worse (with and without data normalization). Mean ensembling and Max ensembling are working much better for me. \n\nMean ensembling\nFor each image I have selected the best 100 predictions confidence from each model.\nThen average confidences across all models and choose the ones with the highest mean of confidences\n\n Max ensembling\nFor each image I have selected the best 100 predictions confidence from each model.\nThen choose the top 5 predictions by the max confidence across all models.\n\n\n**4. EffnetsV1 lower than B4**\nContext of the competition advantages bigger architectures. My best results were using EffNets B5 to B7 and ConvNext's architectures\n\n**5. Gradient accumulation**\nGradient accumulation means running a configured number of steps without updating the model variables while accumulating the gradients of those steps and then using the accumulated gradients to compute the variable updates.\nSounds simple and effective but it never worked for me. It always messes up batch normalization more than it helps with increasing gradients robustness\n\n**6. Combining multiple datasets for a single model in order to increase robustness**\nThe idea was to merge the fins dataset with a dataset that is having the whole whale in the image for learning to recognize in both scenarios. For now, the results are not positive but I still have some ideas to test\n\n**7. Original dataset**\nEither using the fins dataset or another whale cropping dataset, the results will be much better than with the original one. Noisy data present in large scales in images (sea, humans, boats, etc.) are severely affecting the classifier ability to recognize individuals. One interesting aspect here is that Pytorch approaches seems to run much better using the fins dataset, when TF using TPU's are not so sensitive to datasets (fins vs whole whale cropped)",
      "votes": null
    },
    {
      "id": "1743105",
      "postDate": "04/02/2022 16:17:01",
      "content": "<p>Thanks for sharing!</p>\n<p>I learnt a lot 🔥</p>",
      "rawMarkdown": "Thanks for sharing!\n\nI learnt a lot 🔥",
      "votes": null
    },
    {
      "id": "1743379",
      "postDate": "04/02/2022 23:20:10",
      "content": "<p>In my experience, SWA works best with SGD. When averaging just the last epochs, after the model has essentially converged, it is often better than quenching the learning rate. Here, I only used Adam, did not try SWA.</p>",
      "rawMarkdown": "In my experience, SWA works best with SGD. When averaging just the last epochs, after the model has essentially converged, it is often better than quenching the learning rate. Here, I only used Adam, did not try SWA.",
      "votes": null
    },
    {
      "id": "1744294",
      "postDate": "04/03/2022 20:09:08",
      "content": "<p>For each image I have selected the best 100 predictions confidence from each model. - How? Did you train 100 models?</p>",
      "rawMarkdown": "For each image I have selected the best 100 predictions confidence from each model. - How? Did you train 100 models?",
      "votes": null
    },
    {
      "id": "1744743",
      "postDate": "04/04/2022 09:12:35",
      "content": "<p>I guess, for example, you have model1 and model2, and you have 100 predictions for test image from model1 and for model2, not 100 models but 100 predictions from train from each model. Num_modelsX100 will be 1X100 by mean or max operator</p>",
      "rawMarkdown": "I guess, for example, you have model1 and model2, and you have 100 predictions for test image from model1 and for model2, not 100 models but 100 predictions from train from each model. Num_modelsX100 will be 1X100 by mean or max operator",
      "votes": null
    },
    {
      "id": "1744746",
      "postDate": "04/04/2022 09:13:40",
      "content": "<p>In my case SWA with SGD doesn't work too but different. It doesn't change val and LB score, but no decreasing, just the same</p>",
      "rawMarkdown": "In my case SWA with SGD doesn't work too but different. It doesn't change val and LB score, but no decreasing, just the same",
      "votes": null
    },
    {
      "id": "1744952",
      "postDate": "04/04/2022 12:52:11",
      "content": "<p>When I am making the prediction for an image, first I am getting the embedding for that image and then using Nearest Neighbors method I am inferring 100 closest matching labels with their confidence (1- distance). Using a second model, I do the same methodology and then average (or get the max) from the same prediction.<br>\nFor example (use first 3 prediction instead of 100) model 1 predicts:<br>\na with confidence 0.9<br>\nb with cofidence 0.88<br>\nc with confidence 0.73 </p>\n<p>and model 2 predicts <br>\nb with confidence 0.9<br>\na with confidence 0.83<br>\nd with confidence 0.4</p>\n<p>Mean represents for a: (0.9+0.83)/2, for b (0.88+0.9)/2, for c (0.73+0)/2 , for d (0.4+0)/2<br>\nMax represents for a 0.9, for b 0.9, for c 0.73 and for d 0.4</p>",
      "rawMarkdown": "When I am making the prediction for an image, first I am getting the embedding for that image and then using Nearest Neighbors method I am inferring 100 closest matching labels with their confidence (1- distance). Using a second model, I do the same methodology and then average (or get the max) from the same prediction.\nFor example (use first 3 prediction instead of 100) model 1 predicts:\na with confidence 0.9\nb with cofidence 0.88\nc with confidence 0.73 \n\nand model 2 predicts \nb with confidence 0.9\na with confidence 0.83\nd with confidence 0.4\n\nMean represents for a: (0.9+0.83)/2, for b (0.88+0.9)/2, for c (0.73+0)/2 , for d (0.4+0)/2\nMax represents for a 0.9, for b 0.9, for c 0.73 and for d 0.4",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1743105,
      "author_name": "dqhdqmcttdqx",
      "author_url": "",
      "post_date": "04/02/2022 16:17:01",
      "content": "<p>Thanks for sharing!</p>\n<p>I learnt a lot 🔥</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1743379,
      "author_name": "greendolphin",
      "author_url": "",
      "post_date": "04/02/2022 23:20:10",
      "content": "<p>In my experience, SWA works best with SGD. When averaging just the last epochs, after the model has essentially converged, it is often better than quenching the learning rate. Here, I only used Adam, did not try SWA.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1744294,
      "author_name": "dimka11",
      "author_url": "",
      "post_date": "04/03/2022 20:09:08",
      "content": "<p>For each image I have selected the best 100 predictions confidence from each model. - How? Did you train 100 models?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1744743,
          "author_name": "kwentar",
          "author_url": "",
          "post_date": "04/04/2022 09:12:35",
          "content": "<p>I guess, for example, you have model1 and model2, and you have 100 predictions for test image from model1 and for model2, not 100 models but 100 predictions from train from each model. Num_modelsX100 will be 1X100 by mean or max operator</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1744952,
          "author_name": "vladvdv",
          "author_url": "",
          "post_date": "04/04/2022 12:52:11",
          "content": "<p>When I am making the prediction for an image, first I am getting the embedding for that image and then using Nearest Neighbors method I am inferring 100 closest matching labels with their confidence (1- distance). Using a second model, I do the same methodology and then average (or get the max) from the same prediction.<br>\nFor example (use first 3 prediction instead of 100) model 1 predicts:<br>\na with confidence 0.9<br>\nb with cofidence 0.88<br>\nc with confidence 0.73 </p>\n<p>and model 2 predicts <br>\nb with confidence 0.9<br>\na with confidence 0.83<br>\nd with confidence 0.4</p>\n<p>Mean represents for a: (0.9+0.83)/2, for b (0.88+0.9)/2, for c (0.73+0)/2 , for d (0.4+0)/2<br>\nMax represents for a 0.9, for b 0.9, for c 0.73 and for d 0.4</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1744746,
      "author_name": "kwentar",
      "author_url": "",
      "post_date": "04/04/2022 09:13:40",
      "content": "<p>In my case SWA with SGD doesn't work too but different. It doesn't change val and LB score, but no decreasing, just the same</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1743087": "Among other things that give me decent result, here are 7 things that for me did not work, they either give the same cv/public leaderboard result as without them or lower.\n\n**1. SWA (Stochastic Weight Averaging)**\nAlthough in the past my experience with SWA was positive, in this context it seems not to improve the results (for a fold with SWA 0.714 vs same fold without SWA 0.728)\n\nIntuition for SWA comes from empirical observation that local minima at the end of each learning rate cycle tend to accumulate at the border of areas on loss surface where loss value is low . By taking the average of several such points, it is possible to achieve a wide, generalizable solution with even lower loss.\n\n**2. EffnetsV2**\nI have tried efficientnetv2-s-21k-ft1k and efficientnetv2-l-21k-ft1k. Probably with some fine tune it will provide good results, but all my tunning so far was made on EffnetV1\n\n**3. Concat**\nStrangely it may sound, the concat ensembling is not working for me, I am 90% sure that is something that I do wrong but still did not figure out exactly what. When I take 2 models, each with 512 embeddings per image and create 1024 embeddings for an image the results are worse (with and without data normalization). Mean ensembling and Max ensembling are working much better for me. \n\nMean ensembling\nFor each image I have selected the best 100 predictions confidence from each model.\nThen average confidences across all models and choose the ones with the highest mean of confidences\n\n Max ensembling\nFor each image I have selected the best 100 predictions confidence from each model.\nThen choose the top 5 predictions by the max confidence across all models.\n\n\n**4. EffnetsV1 lower than B4**\nContext of the competition advantages bigger architectures. My best results were using EffNets B5 to B7 and ConvNext's architectures\n\n**5. Gradient accumulation**\nGradient accumulation means running a configured number of steps without updating the model variables while accumulating the gradients of those steps and then using the accumulated gradients to compute the variable updates.\nSounds simple and effective but it never worked for me. It always messes up batch normalization more than it helps with increasing gradients robustness\n\n**6. Combining multiple datasets for a single model in order to increase robustness**\nThe idea was to merge the fins dataset with a dataset that is having the whole whale in the image for learning to recognize in both scenarios. For now, the results are not positive but I still have some ideas to test\n\n**7. Original dataset**\nEither using the fins dataset or another whale cropping dataset, the results will be much better than with the original one. Noisy data present in large scales in images (sea, humans, boats, etc.) are severely affecting the classifier ability to recognize individuals. One interesting aspect here is that Pytorch approaches seems to run much better using the fins dataset, when TF using TPU's are not so sensitive to datasets (fins vs whole whale cropped)",
    "1743105": "Thanks for sharing!\n\nI learnt a lot 🔥",
    "1743379": "In my experience, SWA works best with SGD. When averaging just the last epochs, after the model has essentially converged, it is often better than quenching the learning rate. Here, I only used Adam, did not try SWA.",
    "1744294": "For each image I have selected the best 100 predictions confidence from each model. - How? Did you train 100 models?",
    "1744743": "I guess, for example, you have model1 and model2, and you have 100 predictions for test image from model1 and for model2, not 100 models but 100 predictions from train from each model. Num_modelsX100 will be 1X100 by mean or max operator",
    "1744746": "In my case SWA with SGD doesn't work too but different. It doesn't change val and LB score, but no decreasing, just the same",
    "1744952": "When I am making the prediction for an image, first I am getting the embedding for that image and then using Nearest Neighbors method I am inferring 100 closest matching labels with their confidence (1- distance). Using a second model, I do the same methodology and then average (or get the max) from the same prediction.\nFor example (use first 3 prediction instead of 100) model 1 predicts:\na with confidence 0.9\nb with cofidence 0.88\nc with confidence 0.73 \n\nand model 2 predicts \nb with confidence 0.9\na with confidence 0.83\nd with confidence 0.4\n\nMean represents for a: (0.9+0.83)/2, for b (0.88+0.9)/2, for c (0.73+0)/2 , for d (0.4+0)/2\nMax represents for a 0.9, for b 0.9, for c 0.73 and for d 0.4"
  },
  "source": "meta"
}