{
  "id": 314469,
  "title": "Model predictions",
  "url": "/competitions/happy-whale-and-dolphin/discussion/314469",
  "author_name": "",
  "post_date": "2022-03-22T19:15:54.734861400Z",
  "votes": 2,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I came here late, and as I see it, it is quite tough competition.<br>\nSince I am limited with resources, i have NVIDIA GeForce RTX 2080 Ti, in the first try, by using 512x512 image size, one epoch lasted ~45mins. </p>\n<p>I have decided to preprocess the dataset into numpy arrays stored in hdf5 file, and then to make a dataloader based on that hdf5 file. I was able to improve the training time to 10mins (here I also used gradient scaling, and gradient accumulation, since i am using a batch size of 16).</p>\n<p>Then I stored the trained models in a dataset which later on was imported into the inference notebook.</p>\n<p>The train and validation curves through epochs look like this (blue=train, red=val, sorry for not  putting a legend):</p>\n<p><a href=\"https://postimg.cc/mcVxpDMh\" target=\"_blank\"><img src=\"https://i.postimg.cc/JnC1nHz3/history.png\" alt=\"history.png\"></a></p>\n<p>It is clear that i did not overfit in here.</p>\n<p>Reading through multiple amazing notebooks, I have decided to use the PyTorch based implementation (🐋🐬 PyTorch ⚡ BackFin ConvNeXt ArcFace).</p>\n<p>The problem here is that the models always produce the same score despite using different thresholds, which appeared very strange to me, I did all the things correctly, almost everything is the same as the notebook that I am using. This also happens when I do the things locally, and also in the inference notebook:</p>\n<p><a href=\"https://postimg.cc/XBCzW7Zf\" target=\"_blank\"><img src=\"https://i.postimg.cc/hvyW9Xc5/inference.png\" alt=\"inference.png\"></a></p>\n<p>(Btw, instead of faiss , I  am using NearestNeighbors method from sklearn, by using metric=\"cosine\" and later on used confidence  = 1-distance for obtaining the scores.)</p>\n<p>On the public lb it scored 0.056, which proves that something I am missing in here.</p>\n<p>Then I decided to see the embeddings if the model learned something meaningful. It showed the opposite, the model has learned something by looking at the species (26 unique ones)</p>\n<p>The t-SNE embeddings:<br>\n<a href=\"https://postimg.cc/dkdM9Bqm\" target=\"_blank\"><img src=\"https://i.postimg.cc/0QX8PWF3/tsne.png\" alt=\"tsne.png\"></a></p>\n<p>Sorry for the long post, I need some guidance here, any ideas why this is happening ? Maybe its because I did the training locally and imported it as a external dataset in the inference notebook, or something else is the problem, I am struggling to see where I am missing something.</p>\n<p>Any help would be really appreciated !</p>",
  "messages": [
    {
      "id": "1731857",
      "postDate": "03/22/2022 19:15:54",
      "content": "<p>I came here late, and as I see it, it is quite tough competition.<br>\nSince I am limited with resources, i have NVIDIA GeForce RTX 2080 Ti, in the first try, by using 512x512 image size, one epoch lasted ~45mins. </p>\n<p>I have decided to preprocess the dataset into numpy arrays stored in hdf5 file, and then to make a dataloader based on that hdf5 file. I was able to improve the training time to 10mins (here I also used gradient scaling, and gradient accumulation, since i am using a batch size of 16).</p>\n<p>Then I stored the trained models in a dataset which later on was imported into the inference notebook.</p>\n<p>The train and validation curves through epochs look like this (blue=train, red=val, sorry for not  putting a legend):</p>\n<p><a href=\"https://postimg.cc/mcVxpDMh\" target=\"_blank\"><img src=\"https://i.postimg.cc/JnC1nHz3/history.png\" alt=\"history.png\"></a></p>\n<p>It is clear that i did not overfit in here.</p>\n<p>Reading through multiple amazing notebooks, I have decided to use the PyTorch based implementation (🐋🐬 PyTorch ⚡ BackFin ConvNeXt ArcFace).</p>\n<p>The problem here is that the models always produce the same score despite using different thresholds, which appeared very strange to me, I did all the things correctly, almost everything is the same as the notebook that I am using. This also happens when I do the things locally, and also in the inference notebook:</p>\n<p><a href=\"https://postimg.cc/XBCzW7Zf\" target=\"_blank\"><img src=\"https://i.postimg.cc/hvyW9Xc5/inference.png\" alt=\"inference.png\"></a></p>\n<p>(Btw, instead of faiss , I  am using NearestNeighbors method from sklearn, by using metric=\"cosine\" and later on used confidence  = 1-distance for obtaining the scores.)</p>\n<p>On the public lb it scored 0.056, which proves that something I am missing in here.</p>\n<p>Then I decided to see the embeddings if the model learned something meaningful. It showed the opposite, the model has learned something by looking at the species (26 unique ones)</p>\n<p>The t-SNE embeddings:<br>\n<a href=\"https://postimg.cc/dkdM9Bqm\" target=\"_blank\"><img src=\"https://i.postimg.cc/0QX8PWF3/tsne.png\" alt=\"tsne.png\"></a></p>\n<p>Sorry for the long post, I need some guidance here, any ideas why this is happening ? Maybe its because I did the training locally and imported it as a external dataset in the inference notebook, or something else is the problem, I am struggling to see where I am missing something.</p>\n<p>Any help would be really appreciated !</p>",
      "rawMarkdown": "I came here late, and as I see it, it is quite tough competition.\nSince I am limited with resources, i have NVIDIA GeForce RTX 2080 Ti, in the first try, by using 512x512 image size, one epoch lasted ~45mins. \n\nI have decided to preprocess the dataset into numpy arrays stored in hdf5 file, and then to make a dataloader based on that hdf5 file. I was able to improve the training time to 10mins (here I also used gradient scaling, and gradient accumulation, since i am using a batch size of 16).\n\nThen I stored the trained models in a dataset which later on was imported into the inference notebook.\n\nThe train and validation curves through epochs look like this (blue=train, red=val, sorry for not  putting a legend):\n\n[![history.png](https://i.postimg.cc/JnC1nHz3/history.png)](https://postimg.cc/mcVxpDMh)\n\nIt is clear that i did not overfit in here.\n\nReading through multiple amazing notebooks, I have decided to use the PyTorch based implementation (🐋🐬 PyTorch ⚡ BackFin ConvNeXt ArcFace).\n\nThe problem here is that the models always produce the same score despite using different thresholds, which appeared very strange to me, I did all the things correctly, almost everything is the same as the notebook that I am using. This also happens when I do the things locally, and also in the inference notebook:\n\n[![inference.png](https://i.postimg.cc/hvyW9Xc5/inference.png)](https://postimg.cc/XBCzW7Zf)\n\n(Btw, instead of faiss , I  am using NearestNeighbors method from sklearn, by using metric=\"cosine\" and later on used confidence  = 1-distance for obtaining the scores.)\n\nOn the public lb it scored 0.056, which proves that something I am missing in here.\n\nThen I decided to see the embeddings if the model learned something meaningful. It showed the opposite, the model has learned something by looking at the species (26 unique ones)\n\nThe t-SNE embeddings:\n[![tsne.png](https://i.postimg.cc/0QX8PWF3/tsne.png)](https://postimg.cc/dkdM9Bqm)\n\nSorry for the long post, I need some guidance here, any ideas why this is happening ? Maybe its because I did the training locally and imported it as a external dataset in the inference notebook, or something else is the problem, I am struggling to see where I am missing something.\n\nAny help would be really appreciated !",
      "votes": null
    },
    {
      "id": "1732020",
      "postDate": "03/23/2022 00:04:04",
      "content": "<p>I see that the score improves with a 1.0 threshold. it seems that the similarity between features is too high. use arcface!</p>",
      "rawMarkdown": "I see that the score improves with a 1.0 threshold. it seems that the similarity between features is too high. use arcface!",
      "votes": null
    },
    {
      "id": "1732022",
      "postDate": "03/23/2022 00:07:30",
      "content": "<p>Yes, yes. I did use ArcFace. It was strange to me since always the threshold of 1 improved a little bit, but none of the above 1 improved. What is other thing i should take care of? Thanks anyways.</p>",
      "rawMarkdown": "Yes, yes. I did use ArcFace. It was strange to me since always the threshold of 1 improved a little bit, but none of the above 1 improved. What is other thing i should take care of? Thanks anyways.",
      "votes": null
    },
    {
      "id": "1732028",
      "postDate": "03/23/2022 00:18:23",
      "content": "<p>Hmmm I don't know. I think it's a good idea to make a few small changes to the codes that will score reliably.</p>",
      "rawMarkdown": "Hmmm I don't know. I think it's a good idea to make a few small changes to the codes that will score reliably.",
      "votes": null
    },
    {
      "id": "1750162",
      "postDate": "04/09/2022 11:39:22",
      "content": "<p>Hi mane.stoimchev,<br>\nI think I am in a similar situation as well. Have you resolved this problem? Could you share some valuable insight with me as well?</p>\n<p>Thank you so much!</p>",
      "rawMarkdown": "Hi mane.stoimchev,\nI think I am in a similar situation as well. Have you resolved this problem? Could you share some valuable insight with me as well?\n\nThank you so much!",
      "votes": null
    },
    {
      "id": "1750249",
      "postDate": "04/09/2022 12:53:47",
      "content": "<p>You got high loss(if you use default setting of \"BackFin ConvNeXt ArcFace\" ) and low base cv score (0.23 ??).<br>\nTry to \"correct\" your training to got cv 0.5+ and the new_whale insertion mechanism would work as expectation.</p>\n<p>It should be ok to got cv 0.5+ with resolution 384x384 in about 20~25 epochs.</p>",
      "rawMarkdown": "You got high loss(if you use default setting of \"BackFin ConvNeXt ArcFace\" ) and low base cv score (0.23 ??).\nTry to \"correct\" your training to got cv 0.5+ and the new_whale insertion mechanism would work as expectation.\n\nIt should be ok to got cv 0.5+ with resolution 384x384 in about 20~25 epochs.",
      "votes": null
    },
    {
      "id": "1750260",
      "postDate": "04/09/2022 13:08:35",
      "content": "<p>Actually, here I used detic crop dataset 512x512, focal loss + ArcFace. Ill try to change the dataset to dorsal fins. However, I did change the architecture from effnet_b0 to b3 and got cv of around 0.67 (best_th = 0.6), its working now, trying to improve. Are you using the backfins dataset :) ? </p>",
      "rawMarkdown": "Actually, here I used detic crop dataset 512x512, focal loss + ArcFace. Ill try to change the dataset to dorsal fins. However, I did change the architecture from effnet_b0 to b3 and got cv of around 0.67 (best_th = 0.6), its working now, trying to improve. Are you using the backfins dataset :) ?",
      "votes": null
    },
    {
      "id": "1750261",
      "postDate": "04/09/2022 13:10:20",
      "content": "<p>Hi, try to change the architecture that was the case for me, and try to change the dataset as well. But make sure you use proper augmentations and generalizations, not to overfit. My results are not great , but was able to move from the dead point.</p>",
      "rawMarkdown": "Hi, try to change the architecture that was the case for me, and try to change the dataset as well. But make sure you use proper augmentations and generalizations, not to overfit. My results are not great , but was able to move from the dead point.",
      "votes": null
    },
    {
      "id": "1750265",
      "postDate": "04/09/2022 13:19:23",
      "content": "<p>I tried detic dataset, fullbody dataset ,backfin dataset and customized one . </p>",
      "rawMarkdown": "I tried detic dataset, fullbody dataset ,backfin dataset and customized one .",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1732020,
      "author_name": "yujiariyasu",
      "author_url": "",
      "post_date": "03/23/2022 00:04:04",
      "content": "<p>I see that the score improves with a 1.0 threshold. it seems that the similarity between features is too high. use arcface!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1732022,
          "author_name": "marjan1111",
          "author_url": "",
          "post_date": "03/23/2022 00:07:30",
          "content": "<p>Yes, yes. I did use ArcFace. It was strange to me since always the threshold of 1 improved a little bit, but none of the above 1 improved. What is other thing i should take care of? Thanks anyways.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1732028,
          "author_name": "yujiariyasu",
          "author_url": "",
          "post_date": "03/23/2022 00:18:23",
          "content": "<p>Hmmm I don't know. I think it's a good idea to make a few small changes to the codes that will score reliably.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1750162,
          "author_name": "junzhining",
          "author_url": "",
          "post_date": "04/09/2022 11:39:22",
          "content": "<p>Hi mane.stoimchev,<br>\nI think I am in a similar situation as well. Have you resolved this problem? Could you share some valuable insight with me as well?</p>\n<p>Thank you so much!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1750261,
          "author_name": "marjan1111",
          "author_url": "",
          "post_date": "04/09/2022 13:10:20",
          "content": "<p>Hi, try to change the architecture that was the case for me, and try to change the dataset as well. But make sure you use proper augmentations and generalizations, not to overfit. My results are not great , but was able to move from the dead point.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1750249,
      "author_name": "atom1231",
      "author_url": "",
      "post_date": "04/09/2022 12:53:47",
      "content": "<p>You got high loss(if you use default setting of \"BackFin ConvNeXt ArcFace\" ) and low base cv score (0.23 ??).<br>\nTry to \"correct\" your training to got cv 0.5+ and the new_whale insertion mechanism would work as expectation.</p>\n<p>It should be ok to got cv 0.5+ with resolution 384x384 in about 20~25 epochs.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1750260,
          "author_name": "marjan1111",
          "author_url": "",
          "post_date": "04/09/2022 13:08:35",
          "content": "<p>Actually, here I used detic crop dataset 512x512, focal loss + ArcFace. Ill try to change the dataset to dorsal fins. However, I did change the architecture from effnet_b0 to b3 and got cv of around 0.67 (best_th = 0.6), its working now, trying to improve. Are you using the backfins dataset :) ? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1750265,
          "author_name": "atom1231",
          "author_url": "",
          "post_date": "04/09/2022 13:19:23",
          "content": "<p>I tried detic dataset, fullbody dataset ,backfin dataset and customized one . </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1731857": "I came here late, and as I see it, it is quite tough competition.\nSince I am limited with resources, i have NVIDIA GeForce RTX 2080 Ti, in the first try, by using 512x512 image size, one epoch lasted ~45mins. \n\nI have decided to preprocess the dataset into numpy arrays stored in hdf5 file, and then to make a dataloader based on that hdf5 file. I was able to improve the training time to 10mins (here I also used gradient scaling, and gradient accumulation, since i am using a batch size of 16).\n\nThen I stored the trained models in a dataset which later on was imported into the inference notebook.\n\nThe train and validation curves through epochs look like this (blue=train, red=val, sorry for not  putting a legend):\n\n[![history.png](https://i.postimg.cc/JnC1nHz3/history.png)](https://postimg.cc/mcVxpDMh)\n\nIt is clear that i did not overfit in here.\n\nReading through multiple amazing notebooks, I have decided to use the PyTorch based implementation (🐋🐬 PyTorch ⚡ BackFin ConvNeXt ArcFace).\n\nThe problem here is that the models always produce the same score despite using different thresholds, which appeared very strange to me, I did all the things correctly, almost everything is the same as the notebook that I am using. This also happens when I do the things locally, and also in the inference notebook:\n\n[![inference.png](https://i.postimg.cc/hvyW9Xc5/inference.png)](https://postimg.cc/XBCzW7Zf)\n\n(Btw, instead of faiss , I  am using NearestNeighbors method from sklearn, by using metric=\"cosine\" and later on used confidence  = 1-distance for obtaining the scores.)\n\nOn the public lb it scored 0.056, which proves that something I am missing in here.\n\nThen I decided to see the embeddings if the model learned something meaningful. It showed the opposite, the model has learned something by looking at the species (26 unique ones)\n\nThe t-SNE embeddings:\n[![tsne.png](https://i.postimg.cc/0QX8PWF3/tsne.png)](https://postimg.cc/dkdM9Bqm)\n\nSorry for the long post, I need some guidance here, any ideas why this is happening ? Maybe its because I did the training locally and imported it as a external dataset in the inference notebook, or something else is the problem, I am struggling to see where I am missing something.\n\nAny help would be really appreciated !",
    "1732020": "I see that the score improves with a 1.0 threshold. it seems that the similarity between features is too high. use arcface!",
    "1732022": "Yes, yes. I did use ArcFace. It was strange to me since always the threshold of 1 improved a little bit, but none of the above 1 improved. What is other thing i should take care of? Thanks anyways.",
    "1732028": "Hmmm I don't know. I think it's a good idea to make a few small changes to the codes that will score reliably.",
    "1750162": "Hi mane.stoimchev,\nI think I am in a similar situation as well. Have you resolved this problem? Could you share some valuable insight with me as well?\n\nThank you so much!",
    "1750249": "You got high loss(if you use default setting of \"BackFin ConvNeXt ArcFace\" ) and low base cv score (0.23 ??).\nTry to \"correct\" your training to got cv 0.5+ and the new_whale insertion mechanism would work as expectation.\n\nIt should be ok to got cv 0.5+ with resolution 384x384 in about 20~25 epochs.",
    "1750260": "Actually, here I used detic crop dataset 512x512, focal loss + ArcFace. Ill try to change the dataset to dorsal fins. However, I did change the architecture from effnet_b0 to b3 and got cv of around 0.67 (best_th = 0.6), its working now, trying to improve. Are you using the backfins dataset :) ?",
    "1750261": "Hi, try to change the architecture that was the case for me, and try to change the dataset as well. But make sure you use proper augmentations and generalizations, not to overfit. My results are not great , but was able to move from the dead point.",
    "1750265": "I tried detic dataset, fullbody dataset ,backfin dataset and customized one ."
  },
  "source": "meta"
}