{
  "id": 135974,
  "title": "Is se_resnext weird?",
  "url": "/competitions/bengaliai-cv19/discussion/135974",
  "author_name": "",
  "post_date": "2020-03-17T01:03:03.126836900Z",
  "votes": 3,
  "comment_count": 16,
  "views": 0,
  "content": "<p>  I choose the 'se resnext50 32x4d' (single model, single fold) as final submission. When this competition finished, it's out my expect that I shook down so many positions.  </p>\n\n<hr>\n\n<p>  However, I found one of my submissions, it's a densenet(fork of iafoss's kernel and fine tune) which can get 0.9355(private lb), it's enough to get silver. <br>\n  What's more, I discovery that seems most of the people who use se resnext shook down. (The one I think what net he/she uses is based on his/her comments in discussion area) <br>\nSo is se_resnext hard to tune? Or its network architecture doesn't suit for this task?</p>",
  "messages": [
    {
      "id": "775793",
      "postDate": "03/17/2020 01:03:03",
      "content": "<p>  I choose the 'se resnext50 32x4d' (single model, single fold) as final submission. When this competition finished, it's out my expect that I shook down so many positions.  </p>\n\n<hr>\n\n<p>  However, I found one of my submissions, it's a densenet(fork of iafoss's kernel and fine tune) which can get 0.9355(private lb), it's enough to get silver. <br>\n  What's more, I discovery that seems most of the people who use se resnext shook down. (The one I think what net he/she uses is based on his/her comments in discussion area) <br>\nSo is se_resnext hard to tune? Or its network architecture doesn't suit for this task?</p>",
      "rawMarkdown": "I choose the 'se resnext50 32x4d' (single model, single fold) as final submission. When this competition finished, it's out my expect that I shook down so many positions.  \n***\n  However, I found one of my submissions, it's a densenet(fork of iafoss's kernel and fine tune) which can get 0.9355(private lb), it's enough to get silver.  \n  What's more, I discovery that seems most of the people who use se resnext shook down. (The one I think what net he/she uses is based on his/her comments in discussion area)  \nSo is se_resnext hard to tune? Or its network architecture doesn't suit for this task?",
      "votes": null
    },
    {
      "id": "775795",
      "postDate": "03/17/2020 01:05:39",
      "content": "<p>I suspect there is more to it than simply the model however; I also noticed my densenet model performed significantly better than se-resnext50 (not that i was very active in this competition)</p>",
      "rawMarkdown": "I suspect there is more to it than simply the model however; I also noticed my densenet model performed significantly better than se-resnext50 (not that i was very active in this competition)",
      "votes": null
    },
    {
      "id": "775797",
      "postDate": "03/17/2020 01:06:52",
      "content": "<p>might be just coincidence. I only used seresnext50. No shake down</p>",
      "rawMarkdown": "might be just coincidence. I only used seresnext50. No shake down",
      "votes": null
    },
    {
      "id": "775805",
      "postDate": "03/17/2020 01:13:42",
      "content": "<p>Thank you for your fast reply, I will try some other network in future work ;)</p>",
      "rawMarkdown": "Thank you for your fast reply, I will try some other network in future work ;)",
      "votes": null
    },
    {
      "id": "775818",
      "postDate": "03/17/2020 01:23:13",
      "content": "<p>  Thanks for your fast reply and congrats to you! <br>\n  I'm curious about your solution after I heard it. Can you hint me little or you will disclosure it in the discussion area? (if the latter I will wait =)\n  And if limited to the time, how can I do experiments effectively to choose the network or do some fine tune? In this competition, I just wait 10~20 epochs after changing the type of network or doing fine tune.. <br>\n  Thanks again.</p>",
      "rawMarkdown": "Thanks for your fast reply and congrats to you!  \n  I'm curious about your solution after I heard it. Can you hint me little or you will disclosure it in the discussion area? (if the latter I will wait =)\n  And if limited to the time, how can I do experiments effectively to choose the network or do some fine tune? In this competition, I just wait 10~20 epochs after changing the type of network or doing fine tune..  \n  Thanks again.",
      "votes": null
    },
    {
      "id": "775886",
      "postDate": "03/17/2020 02:18:30",
      "content": "<p>I used se-resnext 50 and I actually jumped up 150 positions and ended up in sliver. But i ensembled 6 models. </p>",
      "rawMarkdown": "I used se-resnext 50 and I actually jumped up 150 positions and ended up in sliver. But i ensembled 6 models.",
      "votes": null
    },
    {
      "id": "775893",
      "postDate": "03/17/2020 02:22:23",
      "content": "<p>congrats! I will wait some solution about se-resnext50 ;)</p>",
      "rawMarkdown": "congrats! I will wait some solution about se-resnext50 ;)",
      "votes": null
    },
    {
      "id": "775902",
      "postDate": "03/17/2020 02:29:33",
      "content": "<p>I used 2x seresnext50 and 1 efficientnet b4. Turns out the b4 gave 0.9324 on its own and ensemble of the max of output of all 3 model gave 0.9327. Too bad I chose simple addition of the 3 models which performed worse. Overall I climbed 683 places. seresnext on its own wasn't that bad, gave 0.925.</p>",
      "rawMarkdown": "I used 2x seresnext50 and 1 efficientnet b4. Turns out the b4 gave 0.9324 on its own and ensemble of the max of output of all 3 model gave 0.9327. Too bad I chose simple addition of the 3 models which performed worse. Overall I climbed 683 places. seresnext on its own wasn't that bad, gave 0.925.",
      "votes": null
    },
    {
      "id": "775919",
      "postDate": "03/17/2020 02:53:27",
      "content": "<p>Yes, my best result for se-resnext is 0.9244, so I wanna know if se-resnext is more difficult to tune compared with other networks?(Because I stuck in seresnext, so I didn't do experiments, I will try later =) <br>\nThank you for your reply.</p>",
      "rawMarkdown": "Yes, my best result for se-resnext is 0.9244, so I wanna know if se-resnext is more difficult to tune compared with other networks?(Because I stuck in seresnext, so I didn't do experiments, I will try later =)  \nThank you for your reply.",
      "votes": null
    },
    {
      "id": "775974",
      "postDate": "03/17/2020 03:56:08",
      "content": "<p>Some other things I noticed:\n- seresnext101 was not performing as well as seresnext50. need to check the results with private set.\n- seresnexts took less time to train than efficientnets with mixed precision. i guess this is because some of the layers in efficientnet doesnt get benefit of cudnn. They end up being less efficient than seresnexts !\n- Same hyperparameters doesn't work on both of them which is kind of obvious i guess...\n- On public lb, efficientnets were just as good as seresnext if not better. on private lb efficientnets were marginally better.\n- Pipelines matters. Many of the guys at top used a mixture of efficientnets and seresnext and was not affected by shakedown.\n- Like you noticed with that densenet kernel, it was as good as seresnexts. I regret on dismissing densenets early on and didn't try any of the bigger models :/</p>",
      "rawMarkdown": "Some other things I noticed:\n- seresnext101 was not performing as well as seresnext50. need to check the results with private set.\n- seresnexts took less time to train than efficientnets with mixed precision. i guess this is because some of the layers in efficientnet doesnt get benefit of cudnn. They end up being less efficient than seresnexts !\n- Same hyperparameters doesn't work on both of them which is kind of obvious i guess...\n- On public lb, efficientnets were just as good as seresnext if not better. on private lb efficientnets were marginally better.\n- Pipelines matters. Many of the guys at top used a mixture of efficientnets and seresnext and was not affected by shakedown.\n- Like you noticed with that densenet kernel, it was as good as seresnexts. I regret on dismissing densenets early on and didn't try any of the bigger models :/",
      "votes": null
    },
    {
      "id": "776108",
      "postDate": "03/17/2020 05:59:53",
      "content": "<p>  Thank you for sharing your results. And it's the longest comment I have ever received 😄 <br>\n  I couldn't agree more about the training time, the one most important reason why I choose se-resnext50 is it trains so fast(18 min/epoch in P100, I use colab). So I can do more experiments and epochs for it(I finally train 80 epochs), but I didn't deal with the unseen grapheme, so I shook down heavily ;) <br>\n  It seems the next cv competition, we can start with densenet ;)</p>",
      "rawMarkdown": "Thank you for sharing your results. And it's the longest comment I have ever received 😄   \n  I couldn't agree more about the training time, the one most important reason why I choose se-resnext50 is it trains so fast(18 min/epoch in P100, I use colab). So I can do more experiments and epochs for it(I finally train 80 epochs), but I didn't deal with the unseen grapheme, so I shook down heavily ;)  \n  It seems the next cv competition, we can start with densenet ;)",
      "votes": null
    },
    {
      "id": "776120",
      "postDate": "03/17/2020 06:11:16",
      "content": "<p>I doubt that architecture can have such huge influence. Maybe it's because densenet tends to overfit less compared with resnext50 due to the model capability?</p>",
      "rawMarkdown": "I doubt that architecture can have such huge influence. Maybe it's because densenet tends to overfit less compared with resnext50 due to the model capability?",
      "votes": null
    },
    {
      "id": "776147",
      "postDate": "03/17/2020 06:37:20",
      "content": "<p>The 'capability' you mean is the depth? If it's, the reason what you explain might be reasonable, but se-resnext101 also work not so well... And as a novice I can't judge if it's because of the capability.. <br>\nBelow are some points of me <br>\n- If deal with unseen grapheme properly, the se-resnext will work well as some winner also use it. <br>\n- If not, it hard to say because of the layers(conv, lin, etc) added by every player are different.(variable is not unique) <br>\nHowever, if just a plain implement, se-resnext50 didn't work well for me...</p>\n\n<hr>\n\n<p>Conclusion: try more😂 </p>",
      "rawMarkdown": "The 'capability' you mean is the depth? If it's, the reason what you explain might be reasonable, but se-resnext101 also work not so well... And as a novice I can't judge if it's because of the capability..  \nBelow are some points of me  \n- If deal with unseen grapheme properly, the se-resnext will work well as some winner also use it.  \n- If not, it hard to say because of the layers(conv, lin, etc) added by every player are different.(variable is not unique)  \nHowever, if just a plain implement, se-resnext50 didn't work well for me...\n***\nConclusion: try more😂",
      "votes": null
    },
    {
      "id": "776266",
      "postDate": "03/17/2020 08:51:15",
      "content": "<p>I would love to understand better why this huge shake up occured, but I guess it's hard to blame se-resnext as I jump 50 places and my solution is based only on se resnext50 32x4d.\nI think that if your model (like mine) did not take care of the seen/unseen grapheme paradigm then your private score is kind of random (even if I see a strong correlation between my public and private scores).</p>\n\n<p>The only reason I can see for the shake up: there was more unseen graphemes in the private set and macro recall is a quite sensitive metric. Then your final score is your public score + some very noisy score for unseen graphemes.\nSee <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/136021\">https://www.kaggle.com/c/bengaliai-cv19/discussion/136021</a> about macro recall sensitivity.</p>",
      "rawMarkdown": "I would love to understand better why this huge shake up occured, but I guess it's hard to blame se-resnext as I jump 50 places and my solution is based only on se resnext50 32x4d.\nI think that if your model (like mine) did not take care of the seen/unseen grapheme paradigm then your private score is kind of random (even if I see a strong correlation between my public and private scores).\n\nThe only reason I can see for the shake up: there was more unseen graphemes in the private set and macro recall is a quite sensitive metric. Then your final score is your public score + some very noisy score for unseen graphemes.\nSee https://www.kaggle.com/c/bengaliai-cv19/discussion/136021 about macro recall sensitivity.",
      "votes": null
    },
    {
      "id": "776302",
      "postDate": "03/17/2020 09:34:44",
      "content": "<p>Thanks for your reply and congrats! <br>\nYou are right, deal with the unseen grapheme is the core and macro recall is so sensitive and I found it during the training. <br>\nAlso, the post-process of the discussion is amazing! <br>\nMy real meaning is that if it possible to change type to other network may get a better(or easy tune) result?(this question may be useless because it seems emphasized the 'if', lol) <br>\nCause for the time limited I choose the fastest train model, will try more experiments later ;)</p>",
      "rawMarkdown": "Thanks for your reply and congrats!  \nYou are right, deal with the unseen grapheme is the core and macro recall is so sensitive and I found it during the training.  \nAlso, the post-process of the discussion is amazing!  \nMy real meaning is that if it possible to change type to other network may get a better(or easy tune) result?(this question may be useless because it seems emphasized the 'if', lol)  \nCause for the time limited I choose the fastest train model, will try more experiments later ;)",
      "votes": null
    },
    {
      "id": "776372",
      "postDate": "03/17/2020 10:45:34",
      "content": "<p>I think the reason is SEResNeXt is a more expressive more model and overfits to the seen combinations more easily than focusing on the individual components, which made it shakedown for unseen graphemes.</p>",
      "rawMarkdown": "I think the reason is SEResNeXt is a more expressive more model and overfits to the seen combinations more easily than focusing on the individual components, which made it shakedown for unseen graphemes.",
      "votes": null
    },
    {
      "id": "777804",
      "postDate": "03/18/2020 00:31:20",
      "content": "<p>Thanks for your reply! <br>\nI agree with you, after I train it during 30-60 epochs, the submission jumped so quickly in public lb and it puzzle me from juding the result rightly.</p>",
      "rawMarkdown": "Thanks for your reply!  \nI agree with you, after I train it during 30-60 epochs, the submission jumped so quickly in public lb and it puzzle me from juding the result rightly.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 775795,
      "author_name": "nathanpw",
      "author_url": "",
      "post_date": "03/17/2020 01:05:39",
      "content": "<p>I suspect there is more to it than simply the model however; I also noticed my densenet model performed significantly better than se-resnext50 (not that i was very active in this competition)</p>",
      "votes": null,
      "replies": [
        {
          "id": 775805,
          "author_name": "cnzengshiyuan",
          "author_url": "",
          "post_date": "03/17/2020 01:13:42",
          "content": "<p>Thank you for your fast reply, I will try some other network in future work ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 775797,
      "author_name": "christofhenkel",
      "author_url": "",
      "post_date": "03/17/2020 01:06:52",
      "content": "<p>might be just coincidence. I only used seresnext50. No shake down</p>",
      "votes": null,
      "replies": [
        {
          "id": 775818,
          "author_name": "cnzengshiyuan",
          "author_url": "",
          "post_date": "03/17/2020 01:23:13",
          "content": "<p>  Thanks for your fast reply and congrats to you! <br>\n  I'm curious about your solution after I heard it. Can you hint me little or you will disclosure it in the discussion area? (if the latter I will wait =)\n  And if limited to the time, how can I do experiments effectively to choose the network or do some fine tune? In this competition, I just wait 10~20 epochs after changing the type of network or doing fine tune.. <br>\n  Thanks again.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 775886,
      "author_name": "yovinyahathugoda",
      "author_url": "",
      "post_date": "03/17/2020 02:18:30",
      "content": "<p>I used se-resnext 50 and I actually jumped up 150 positions and ended up in sliver. But i ensembled 6 models. </p>",
      "votes": null,
      "replies": [
        {
          "id": 775893,
          "author_name": "cnzengshiyuan",
          "author_url": "",
          "post_date": "03/17/2020 02:22:23",
          "content": "<p>congrats! I will wait some solution about se-resnext50 ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 775902,
      "author_name": "utsavnandi",
      "author_url": "",
      "post_date": "03/17/2020 02:29:33",
      "content": "<p>I used 2x seresnext50 and 1 efficientnet b4. Turns out the b4 gave 0.9324 on its own and ensemble of the max of output of all 3 model gave 0.9327. Too bad I chose simple addition of the 3 models which performed worse. Overall I climbed 683 places. seresnext on its own wasn't that bad, gave 0.925.</p>",
      "votes": null,
      "replies": [
        {
          "id": 775919,
          "author_name": "cnzengshiyuan",
          "author_url": "",
          "post_date": "03/17/2020 02:53:27",
          "content": "<p>Yes, my best result for se-resnext is 0.9244, so I wanna know if se-resnext is more difficult to tune compared with other networks?(Because I stuck in seresnext, so I didn't do experiments, I will try later =) <br>\nThank you for your reply.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 775974,
          "author_name": "utsavnandi",
          "author_url": "",
          "post_date": "03/17/2020 03:56:08",
          "content": "<p>Some other things I noticed:\n- seresnext101 was not performing as well as seresnext50. need to check the results with private set.\n- seresnexts took less time to train than efficientnets with mixed precision. i guess this is because some of the layers in efficientnet doesnt get benefit of cudnn. They end up being less efficient than seresnexts !\n- Same hyperparameters doesn't work on both of them which is kind of obvious i guess...\n- On public lb, efficientnets were just as good as seresnext if not better. on private lb efficientnets were marginally better.\n- Pipelines matters. Many of the guys at top used a mixture of efficientnets and seresnext and was not affected by shakedown.\n- Like you noticed with that densenet kernel, it was as good as seresnexts. I regret on dismissing densenets early on and didn't try any of the bigger models :/</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 776108,
          "author_name": "cnzengshiyuan",
          "author_url": "",
          "post_date": "03/17/2020 05:59:53",
          "content": "<p>  Thank you for sharing your results. And it's the longest comment I have ever received 😄 <br>\n  I couldn't agree more about the training time, the one most important reason why I choose se-resnext50 is it trains so fast(18 min/epoch in P100, I use colab). So I can do more experiments and epochs for it(I finally train 80 epochs), but I didn't deal with the unseen grapheme, so I shook down heavily ;) <br>\n  It seems the next cv competition, we can start with densenet ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 776120,
      "author_name": "syoya1997",
      "author_url": "",
      "post_date": "03/17/2020 06:11:16",
      "content": "<p>I doubt that architecture can have such huge influence. Maybe it's because densenet tends to overfit less compared with resnext50 due to the model capability?</p>",
      "votes": null,
      "replies": [
        {
          "id": 776147,
          "author_name": "cnzengshiyuan",
          "author_url": "",
          "post_date": "03/17/2020 06:37:20",
          "content": "<p>The 'capability' you mean is the depth? If it's, the reason what you explain might be reasonable, but se-resnext101 also work not so well... And as a novice I can't judge if it's because of the capability.. <br>\nBelow are some points of me <br>\n- If deal with unseen grapheme properly, the se-resnext will work well as some winner also use it. <br>\n- If not, it hard to say because of the layers(conv, lin, etc) added by every player are different.(variable is not unique) <br>\nHowever, if just a plain implement, se-resnext50 didn't work well for me...</p>\n\n<hr>\n\n<p>Conclusion: try more😂 </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 776266,
      "author_name": "optimo",
      "author_url": "",
      "post_date": "03/17/2020 08:51:15",
      "content": "<p>I would love to understand better why this huge shake up occured, but I guess it's hard to blame se-resnext as I jump 50 places and my solution is based only on se resnext50 32x4d.\nI think that if your model (like mine) did not take care of the seen/unseen grapheme paradigm then your private score is kind of random (even if I see a strong correlation between my public and private scores).</p>\n\n<p>The only reason I can see for the shake up: there was more unseen graphemes in the private set and macro recall is a quite sensitive metric. Then your final score is your public score + some very noisy score for unseen graphemes.\nSee <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/136021\">https://www.kaggle.com/c/bengaliai-cv19/discussion/136021</a> about macro recall sensitivity.</p>",
      "votes": null,
      "replies": [
        {
          "id": 776302,
          "author_name": "cnzengshiyuan",
          "author_url": "",
          "post_date": "03/17/2020 09:34:44",
          "content": "<p>Thanks for your reply and congrats! <br>\nYou are right, deal with the unseen grapheme is the core and macro recall is so sensitive and I found it during the training. <br>\nAlso, the post-process of the discussion is amazing! <br>\nMy real meaning is that if it possible to change type to other network may get a better(or easy tune) result?(this question may be useless because it seems emphasized the 'if', lol) <br>\nCause for the time limited I choose the fastest train model, will try more experiments later ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 776372,
      "author_name": "dipamc77",
      "author_url": "",
      "post_date": "03/17/2020 10:45:34",
      "content": "<p>I think the reason is SEResNeXt is a more expressive more model and overfits to the seen combinations more easily than focusing on the individual components, which made it shakedown for unseen graphemes.</p>",
      "votes": null,
      "replies": [
        {
          "id": 777804,
          "author_name": "cnzengshiyuan",
          "author_url": "",
          "post_date": "03/18/2020 00:31:20",
          "content": "<p>Thanks for your reply! <br>\nI agree with you, after I train it during 30-60 epochs, the submission jumped so quickly in public lb and it puzzle me from juding the result rightly.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "775793": "I choose the 'se resnext50 32x4d' (single model, single fold) as final submission. When this competition finished, it's out my expect that I shook down so many positions.  \n***\n  However, I found one of my submissions, it's a densenet(fork of iafoss's kernel and fine tune) which can get 0.9355(private lb), it's enough to get silver.  \n  What's more, I discovery that seems most of the people who use se resnext shook down. (The one I think what net he/she uses is based on his/her comments in discussion area)  \nSo is se_resnext hard to tune? Or its network architecture doesn't suit for this task?",
    "775795": "I suspect there is more to it than simply the model however; I also noticed my densenet model performed significantly better than se-resnext50 (not that i was very active in this competition)",
    "775797": "might be just coincidence. I only used seresnext50. No shake down",
    "775805": "Thank you for your fast reply, I will try some other network in future work ;)",
    "775818": "Thanks for your fast reply and congrats to you!  \n  I'm curious about your solution after I heard it. Can you hint me little or you will disclosure it in the discussion area? (if the latter I will wait =)\n  And if limited to the time, how can I do experiments effectively to choose the network or do some fine tune? In this competition, I just wait 10~20 epochs after changing the type of network or doing fine tune..  \n  Thanks again.",
    "775886": "I used se-resnext 50 and I actually jumped up 150 positions and ended up in sliver. But i ensembled 6 models.",
    "775893": "congrats! I will wait some solution about se-resnext50 ;)",
    "775902": "I used 2x seresnext50 and 1 efficientnet b4. Turns out the b4 gave 0.9324 on its own and ensemble of the max of output of all 3 model gave 0.9327. Too bad I chose simple addition of the 3 models which performed worse. Overall I climbed 683 places. seresnext on its own wasn't that bad, gave 0.925.",
    "775919": "Yes, my best result for se-resnext is 0.9244, so I wanna know if se-resnext is more difficult to tune compared with other networks?(Because I stuck in seresnext, so I didn't do experiments, I will try later =)  \nThank you for your reply.",
    "775974": "Some other things I noticed:\n- seresnext101 was not performing as well as seresnext50. need to check the results with private set.\n- seresnexts took less time to train than efficientnets with mixed precision. i guess this is because some of the layers in efficientnet doesnt get benefit of cudnn. They end up being less efficient than seresnexts !\n- Same hyperparameters doesn't work on both of them which is kind of obvious i guess...\n- On public lb, efficientnets were just as good as seresnext if not better. on private lb efficientnets were marginally better.\n- Pipelines matters. Many of the guys at top used a mixture of efficientnets and seresnext and was not affected by shakedown.\n- Like you noticed with that densenet kernel, it was as good as seresnexts. I regret on dismissing densenets early on and didn't try any of the bigger models :/",
    "776108": "Thank you for sharing your results. And it's the longest comment I have ever received 😄   \n  I couldn't agree more about the training time, the one most important reason why I choose se-resnext50 is it trains so fast(18 min/epoch in P100, I use colab). So I can do more experiments and epochs for it(I finally train 80 epochs), but I didn't deal with the unseen grapheme, so I shook down heavily ;)  \n  It seems the next cv competition, we can start with densenet ;)",
    "776120": "I doubt that architecture can have such huge influence. Maybe it's because densenet tends to overfit less compared with resnext50 due to the model capability?",
    "776147": "The 'capability' you mean is the depth? If it's, the reason what you explain might be reasonable, but se-resnext101 also work not so well... And as a novice I can't judge if it's because of the capability..  \nBelow are some points of me  \n- If deal with unseen grapheme properly, the se-resnext will work well as some winner also use it.  \n- If not, it hard to say because of the layers(conv, lin, etc) added by every player are different.(variable is not unique)  \nHowever, if just a plain implement, se-resnext50 didn't work well for me...\n***\nConclusion: try more😂",
    "776266": "I would love to understand better why this huge shake up occured, but I guess it's hard to blame se-resnext as I jump 50 places and my solution is based only on se resnext50 32x4d.\nI think that if your model (like mine) did not take care of the seen/unseen grapheme paradigm then your private score is kind of random (even if I see a strong correlation between my public and private scores).\n\nThe only reason I can see for the shake up: there was more unseen graphemes in the private set and macro recall is a quite sensitive metric. Then your final score is your public score + some very noisy score for unseen graphemes.\nSee https://www.kaggle.com/c/bengaliai-cv19/discussion/136021 about macro recall sensitivity.",
    "776302": "Thanks for your reply and congrats!  \nYou are right, deal with the unseen grapheme is the core and macro recall is so sensitive and I found it during the training.  \nAlso, the post-process of the discussion is amazing!  \nMy real meaning is that if it possible to change type to other network may get a better(or easy tune) result?(this question may be useless because it seems emphasized the 'if', lol)  \nCause for the time limited I choose the fastest train model, will try more experiments later ;)",
    "776372": "I think the reason is SEResNeXt is a more expressive more model and overfits to the seen combinations more easily than focusing on the individual components, which made it shakedown for unseen graphemes.",
    "777804": "Thanks for your reply!  \nI agree with you, after I train it during 30-60 epochs, the submission jumped so quickly in public lb and it puzzle me from juding the result rightly."
  },
  "source": "meta"
}