{
  "id": 305951,
  "title": "Could transformer work for this competition?",
  "url": "/competitions/happy-whale-and-dolphin/discussion/305951",
  "author_name": "",
  "post_date": "2022-02-07T16:00:38.423762800Z",
  "votes": 7,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I tried this notebook<a href=\"url\" target=\"_blank\">https://www.kaggle.com/debarshichanda/pytorch-arcface-gem-pooling-starter/comments</a>, changing the architecture from CNN to swin_transformer.<br>\nThen the loss just descended so slowly, from 24 to 22 after 3 epochs, even worse than 1 epoch with CNN.<br>\nDid I miss something?</p>",
  "messages": [
    {
      "id": "1680155",
      "postDate": "02/07/2022 16:00:38",
      "content": "<p>I tried this notebook<a href=\"url\" target=\"_blank\">https://www.kaggle.com/debarshichanda/pytorch-arcface-gem-pooling-starter/comments</a>, changing the architecture from CNN to swin_transformer.<br>\nThen the loss just descended so slowly, from 24 to 22 after 3 epochs, even worse than 1 epoch with CNN.<br>\nDid I miss something?</p>",
      "rawMarkdown": "I tried this notebook[https://www.kaggle.com/debarshichanda/pytorch-arcface-gem-pooling-starter/comments](url), changing the architecture from CNN to swin_transformer.\nThen the loss just descended so slowly, from 24 to 22 after 3 epochs, even worse than 1 epoch with CNN.\nDid I miss something?",
      "votes": null
    },
    {
      "id": "1680165",
      "postDate": "02/07/2022 16:05:57",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/jsrdcht\" target=\"_blank\">@jsrdcht</a>. Transformers probably won't work as well as CNNs in this competition, though you're welcome to try! Transformers are mainly used for sequential data, such as audio or text. The best models for image data, especially image information retreival, tend to be CNNs.</p>",
      "rawMarkdown": "Hello @jsrdcht. Transformers probably won't work as well as CNNs in this competition, though you're welcome to try! Transformers are mainly used for sequential data, such as audio or text. The best models for image data, especially image information retreival, tend to be CNNs.",
      "votes": null
    },
    {
      "id": "1680226",
      "postDate": "02/07/2022 16:58:56",
      "content": "<p>Thanks! I found vision transformers work well for some competitions, so I thought maybe it would work for this competition too.<br>\nAnyway, you say <strong>\"especially image information retrieval\"</strong>, is there any insight?</p>",
      "rawMarkdown": "Thanks! I found vision transformers work well for some competitions, so I thought maybe it would work for this competition too.\nAnyway, you say **\"especially image information retrieval\"**, is there any insight?",
      "votes": null
    },
    {
      "id": "1680243",
      "postDate": "02/07/2022 17:16:16",
      "content": "<p>Kind of. I suggest you read the following blogpost: <a href=\"https://towardsdatascience.com/a-gold-winning-solution-review-of-kaggle-humpback-whale-identification-challenge-53b0e3ba1e84\" target=\"_blank\">https://towardsdatascience.com/a-gold-winning-solution-review-of-kaggle-humpback-whale-identification-challenge-53b0e3ba1e84</a></p>\n<p>It describes the approaches and techniques used by the winning team on the first Happywhale competition. It seems most people might be using ArcFace/CosFace type of models. They seem to work for this even though they have been introduced as SOTA for face recognition (hence the names)</p>\n<p>Hope I was helpful.</p>",
      "rawMarkdown": "Kind of. I suggest you read the following blogpost: https://towardsdatascience.com/a-gold-winning-solution-review-of-kaggle-humpback-whale-identification-challenge-53b0e3ba1e84\n\nIt describes the approaches and techniques used by the winning team on the first Happywhale competition. It seems most people might be using ArcFace/CosFace type of models. They seem to work for this even though they have been introduced as SOTA for face recognition (hence the names)\n\nHope I was helpful.",
      "votes": null
    },
    {
      "id": "1682520",
      "postDate": "02/09/2022 08:13:22",
      "content": "<p>I disagree about this! </p>\n<p>Transformer models work really well on image tasks too. Please see recent competitions, petfinder, etc, and also the timm framework. </p>\n<p>However, they are known to take longer to converge, which is what ChenTuo observed too. </p>\n<p>Your argument about the winning solution is a bit misleading, the older competition had occurred just pre-transformer era. I had taken part and I remember the good old days of ResNet 50s and training loops shorter than 12 hours :')</p>",
      "rawMarkdown": "I disagree about this! \n\nTransformer models work really well on image tasks too. Please see recent competitions, petfinder, etc, and also the timm framework. \n\nHowever, they are known to take longer to converge, which is what ChenTuo observed too. \n\nYour argument about the winning solution is a bit misleading, the older competition had occurred just pre-transformer era. I had taken part and I remember the good old days of ResNet 50s and training loops shorter than 12 hours :')",
      "votes": null
    },
    {
      "id": "1682731",
      "postDate": "02/09/2022 10:58:06",
      "content": "<p>I see! I'm not that well versed on transformers or the current SoTA, so thanks for your insight <a href=\"https://www.kaggle.com/init27\" target=\"_blank\">@init27</a> Did not know about the timm framework.</p>",
      "rawMarkdown": "I see! I'm not that well versed on transformers or the current SoTA, so thanks for your insight @init27 Did not know about the timm framework.",
      "votes": null
    },
    {
      "id": "1682943",
      "postDate": "02/09/2022 13:33:55",
      "content": "<p>I agree about you. It's just really strange that training loss of Swin I tried converged to 22 after 3-5 epochs, while CNN can converge to 15 at least. I'm really confused about the matter. Maybe I should do more research. </p>",
      "rawMarkdown": "I agree about you. It's just really strange that training loss of Swin I tried converged to 22 after 3-5 epochs, while CNN can converge to 15 at least. I'm really confused about the matter. Maybe I should do more research.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1680165,
      "author_name": "asarvazyan",
      "author_url": "",
      "post_date": "02/07/2022 16:05:57",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/jsrdcht\" target=\"_blank\">@jsrdcht</a>. Transformers probably won't work as well as CNNs in this competition, though you're welcome to try! Transformers are mainly used for sequential data, such as audio or text. The best models for image data, especially image information retreival, tend to be CNNs.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1680226,
          "author_name": "jsrdcht",
          "author_url": "",
          "post_date": "02/07/2022 16:58:56",
          "content": "<p>Thanks! I found vision transformers work well for some competitions, so I thought maybe it would work for this competition too.<br>\nAnyway, you say <strong>\"especially image information retrieval\"</strong>, is there any insight?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1680243,
          "author_name": "asarvazyan",
          "author_url": "",
          "post_date": "02/07/2022 17:16:16",
          "content": "<p>Kind of. I suggest you read the following blogpost: <a href=\"https://towardsdatascience.com/a-gold-winning-solution-review-of-kaggle-humpback-whale-identification-challenge-53b0e3ba1e84\" target=\"_blank\">https://towardsdatascience.com/a-gold-winning-solution-review-of-kaggle-humpback-whale-identification-challenge-53b0e3ba1e84</a></p>\n<p>It describes the approaches and techniques used by the winning team on the first Happywhale competition. It seems most people might be using ArcFace/CosFace type of models. They seem to work for this even though they have been introduced as SOTA for face recognition (hence the names)</p>\n<p>Hope I was helpful.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1682520,
          "author_name": "init27",
          "author_url": "",
          "post_date": "02/09/2022 08:13:22",
          "content": "<p>I disagree about this! </p>\n<p>Transformer models work really well on image tasks too. Please see recent competitions, petfinder, etc, and also the timm framework. </p>\n<p>However, they are known to take longer to converge, which is what ChenTuo observed too. </p>\n<p>Your argument about the winning solution is a bit misleading, the older competition had occurred just pre-transformer era. I had taken part and I remember the good old days of ResNet 50s and training loops shorter than 12 hours :')</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1682731,
          "author_name": "asarvazyan",
          "author_url": "",
          "post_date": "02/09/2022 10:58:06",
          "content": "<p>I see! I'm not that well versed on transformers or the current SoTA, so thanks for your insight <a href=\"https://www.kaggle.com/init27\" target=\"_blank\">@init27</a> Did not know about the timm framework.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1682943,
          "author_name": "jsrdcht",
          "author_url": "",
          "post_date": "02/09/2022 13:33:55",
          "content": "<p>I agree about you. It's just really strange that training loss of Swin I tried converged to 22 after 3-5 epochs, while CNN can converge to 15 at least. I'm really confused about the matter. Maybe I should do more research. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1680155": "I tried this notebook[https://www.kaggle.com/debarshichanda/pytorch-arcface-gem-pooling-starter/comments](url), changing the architecture from CNN to swin_transformer.\nThen the loss just descended so slowly, from 24 to 22 after 3 epochs, even worse than 1 epoch with CNN.\nDid I miss something?",
    "1680165": "Hello @jsrdcht. Transformers probably won't work as well as CNNs in this competition, though you're welcome to try! Transformers are mainly used for sequential data, such as audio or text. The best models for image data, especially image information retreival, tend to be CNNs.",
    "1680226": "Thanks! I found vision transformers work well for some competitions, so I thought maybe it would work for this competition too.\nAnyway, you say **\"especially image information retrieval\"**, is there any insight?",
    "1680243": "Kind of. I suggest you read the following blogpost: https://towardsdatascience.com/a-gold-winning-solution-review-of-kaggle-humpback-whale-identification-challenge-53b0e3ba1e84\n\nIt describes the approaches and techniques used by the winning team on the first Happywhale competition. It seems most people might be using ArcFace/CosFace type of models. They seem to work for this even though they have been introduced as SOTA for face recognition (hence the names)\n\nHope I was helpful.",
    "1682520": "I disagree about this! \n\nTransformer models work really well on image tasks too. Please see recent competitions, petfinder, etc, and also the timm framework. \n\nHowever, they are known to take longer to converge, which is what ChenTuo observed too. \n\nYour argument about the winning solution is a bit misleading, the older competition had occurred just pre-transformer era. I had taken part and I remember the good old days of ResNet 50s and training loops shorter than 12 hours :')",
    "1682731": "I see! I'm not that well versed on transformers or the current SoTA, so thanks for your insight @init27 Did not know about the timm framework.",
    "1682943": "I agree about you. It's just really strange that training loss of Swin I tried converged to 22 after 3-5 epochs, while CNN can converge to 15 at least. I'm really confused about the matter. Maybe I should do more research."
  },
  "source": "meta"
}