{
  "id": 131156,
  "title": "One classifier for all 3 outputs vs 3 distinct classifiers",
  "url": "/competitions/bengaliai-cv19/discussion/131156",
  "author_name": "",
  "post_date": "2020-02-18T15:42:31.783528500Z",
  "votes": 4,
  "comment_count": 15,
  "views": 0,
  "content": "<p>So far, in my tests configurations I designed architectures based on a final layer of 186 neurons (168 grapheme root outputs + 11 grapheme vowels + 7 grapheme consonants).\nI had tuned arhitectures like:\n- resnext50\n- resnext101\n- efficientnet-b0</p>\n\n<p>Activation functions:\n- relu \n- leaky relu</p>\n\n<p>Preprocessing techniques:\n- affine transformation\n- cutout\n- cutmix\n- mixup\n- bluring techniques</p>\n\n<p>Optimizers:\n- Adam\n- AdamW\n- RAdam </p>\n\n<p>My best results so far on the public leaderboard are 0.9750 on a 5 fold configuration and for my current training I hope a boost up to 0.9770.\nDid you guys which had time to experiment more, tried designing 3 different arhitectures for each output category ? (grapheme root, grapheme vowels, grapheme consonants). What results did you have ? I am courious if is this the key of going 0.98 +\nMy intuition tells my that the architecture with the 3 outputs combined should be more performant and robust (the model should understand also the correlations between outputs)</p>",
  "messages": [
    {
      "id": "749335",
      "postDate": "02/18/2020 15:42:31",
      "content": "<p>So far, in my tests configurations I designed architectures based on a final layer of 186 neurons (168 grapheme root outputs + 11 grapheme vowels + 7 grapheme consonants).\nI had tuned arhitectures like:\n- resnext50\n- resnext101\n- efficientnet-b0</p>\n\n<p>Activation functions:\n- relu \n- leaky relu</p>\n\n<p>Preprocessing techniques:\n- affine transformation\n- cutout\n- cutmix\n- mixup\n- bluring techniques</p>\n\n<p>Optimizers:\n- Adam\n- AdamW\n- RAdam </p>\n\n<p>My best results so far on the public leaderboard are 0.9750 on a 5 fold configuration and for my current training I hope a boost up to 0.9770.\nDid you guys which had time to experiment more, tried designing 3 different arhitectures for each output category ? (grapheme root, grapheme vowels, grapheme consonants). What results did you have ? I am courious if is this the key of going 0.98 +\nMy intuition tells my that the architecture with the 3 outputs combined should be more performant and robust (the model should understand also the correlations between outputs)</p>",
      "rawMarkdown": "So far, in my tests configurations I designed architectures based on a final layer of 186 neurons (168 grapheme root outputs + 11 grapheme vowels + 7 grapheme consonants).\nI had tuned arhitectures like:\n- resnext50\n- resnext101\n- efficientnet-b0\n\nActivation functions:\n- relu \n- leaky relu\n\nPreprocessing techniques:\n- affine transformation\n- cutout\n- cutmix\n- mixup\n- bluring techniques\n\nOptimizers:\n- Adam\n- AdamW\n- RAdam \n\nMy best results so far on the public leaderboard are 0.9750 on a 5 fold configuration and for my current training I hope a boost up to 0.9770.\nDid you guys which had time to experiment more, tried designing 3 different arhitectures for each output category ? (grapheme root, grapheme vowels, grapheme consonants). What results did you have ? I am courious if is this the key of going 0.98 +\nMy intuition tells my that the architecture with the 3 outputs combined should be more performant and robust (the model should understand also the correlations between outputs)",
      "votes": null
    },
    {
      "id": "749343",
      "postDate": "02/18/2020 15:59:56",
      "content": "<p>so you are performing 3 softmax separately? or you don't output probabilities?</p>\n\n<p>I personally tried 3 simple heads for each tasks.</p>",
      "rawMarkdown": "so you are performing 3 softmax separately? or you don't output probabilities?\n\nI personally tried 3 simple heads for each tasks.",
      "votes": null
    },
    {
      "id": "749374",
      "postDate": "02/18/2020 16:31:31",
      "content": "<p>Yes. I am doing 3 softmax functions separately (first 168 output neurons one softmax for grapheme root,  for the next 11 output neurons one softmax for grapheme vocal and for the last 7 output neurons one softmax for grapheme consonant).\n<a href=\"/optimo\">@optimo</a> Have you tried 3 different classifiers for each category ? This is next on my todo list</p>",
      "rawMarkdown": "Yes. I am doing 3 softmax functions separately (first 168 output neurons one softmax for grapheme root,  for the next 11 output neurons one softmax for grapheme vocal and for the last 7 output neurons one softmax for grapheme consonant).\n@optimo Have you tried 3 different classifiers for each category ? This is next on my todo list",
      "votes": null
    },
    {
      "id": "749534",
      "postDate": "02/18/2020 18:57:39",
      "content": "<p>This was discussed a little at the beginning of this comp: <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123956\">https://www.kaggle.com/c/bengaliai-cv19/discussion/123956</a></p>",
      "rawMarkdown": "This was discussed a little at the beginning of this comp: https://www.kaggle.com/c/bengaliai-cv19/discussion/123956",
      "votes": null
    },
    {
      "id": "749546",
      "postDate": "02/18/2020 19:12:50",
      "content": "<p>Cool, good to know, there are some interesting points of view there</p>",
      "rawMarkdown": "Cool, good to know, there are some interesting points of view there",
      "votes": null
    },
    {
      "id": "750043",
      "postDate": "02/19/2020 04:51:45",
      "content": "<p>Hi <a href=\"/vladvdv\">@vladvdv</a> ,</p>\n\n<p>May I ask how many epochs do you train and do you use any schedulers? \nSo far, I try with different models for each class and there is a slight improvement:\n<code>\nBackbone: Resnext50\nAug: Cutout/CutMix, RandomRotate\nOptim: SGD\nScheduler: Consine\nCV: 0.969785\nLB: 0.9672\n</code>\nHowever, my model for classify Grapheme reaches its highest only <code>0.97</code>. </p>\n\n<p>Now, I'm finding a way to go to <code>0.98</code>.\nIt's really frustrating since we have 20 days left.  😭 😭 </p>",
      "rawMarkdown": "Hi @vladvdv ,\n\nMay I ask how many epochs do you train and do you use any schedulers? \nSo far, I try with different models for each class and there is a slight improvement:\n```\nBackbone: Resnext50\nAug: Cutout/CutMix, RandomRotate\nOptim: SGD\nScheduler: Consine\nCV: 0.969785\nLB: 0.9672\n```\nHowever, my model for classify Grapheme reaches its highest only `0.97`. \n\nNow, I'm finding a way to go to `0.98`.\nIt's really frustrating since we have 20 days left.  😭 😭",
      "votes": null
    },
    {
      "id": "750225",
      "postDate": "02/19/2020 07:55:54",
      "content": "<p>nope it's would be too time consuming... I did not try</p>",
      "rawMarkdown": "nope it's would be too time consuming... I did not try",
      "votes": null
    },
    {
      "id": "750335",
      "postDate": "02/19/2020 10:02:14",
      "content": "<p>Multilabel classification?</p>",
      "rawMarkdown": "Multilabel classification?",
      "votes": null
    },
    {
      "id": "750402",
      "postDate": "02/19/2020 10:44:38",
      "content": "<p>I’m curious about why not trying a higher EfficientNet?</p>\n\n<p>B4 seems to give some really promising results on my side, though overfitting seems to be something to deal with. Is this the reason you’re staying at a B0, to have a more robust model?</p>",
      "rawMarkdown": "I’m curious about why not trying a higher EfficientNet?\n\nB4 seems to give some really promising results on my side, though overfitting seems to be something to deal with. Is this the reason you’re staying at a B0, to have a more robust model?",
      "votes": null
    },
    {
      "id": "750403",
      "postDate": "02/19/2020 10:45:15",
      "content": "<p>Hi <a href=\"/moximo13\">@moximo13</a> ,\nI am using ReduceLROnPlateau by a factor of 0.9 with patience 5 epochs, also I am using Adam as a optimizer\nMy frustration is on the time duration on training. It takes very long and I do not have enough time to tune every thing I got on the list so I choose based on my experience the most important aspects</p>",
      "rawMarkdown": "Hi @moximo13 ,\nI am using ReduceLROnPlateau by a factor of 0.9 with patience 5 epochs, also I am using Adam as a optimizer\nMy frustration is on the time duration on training. It takes very long and I do not have enough time to tune every thing I got on the list so I choose based on my experience the most important aspects",
      "votes": null
    },
    {
      "id": "750409",
      "postDate": "02/19/2020 10:49:50",
      "content": "<p>Hi <a href=\"/maxlenormand\">@maxlenormand</a> \nEfficientNet are pretty big architectures, especially if you go to B4, from my experience they give good results when your image has a bigger resolution, and to scale our small images images at 300x300, to augment them enough so the model will have diversity so it won't overfit, to wait around 150 epochs for each fold so the model will converge, it will take forever on my GPU</p>",
      "rawMarkdown": "Hi @maxlenormand \nEfficientNet are pretty big architectures, especially if you go to B4, from my experience they give good results when your image has a bigger resolution, and to scale our small images images at 300x300, to augment them enough so the model will have diversity so it won't overfit, to wait around 150 epochs for each fold so the model will converge, it will take forever on my GPU",
      "votes": null
    },
    {
      "id": "750418",
      "postDate": "02/19/2020 10:58:39",
      "content": "<p>Interesting!</p>\n\n<p>Once again I’m not at the same CV as you, I’m struggling to pass 0.963 and above, so maybe you’re right and I should lower the architecture I use and see if that helps.</p>\n\n<p>Either way, thanks for sharing! If I do find some interesting results on the use of different EfficientNet models I’ll come back to the discussion to share that.</p>",
      "rawMarkdown": "Interesting!\n\nOnce again I’m not at the same CV as you, I’m struggling to pass 0.963 and above, so maybe you’re right and I should lower the architecture I use and see if that helps.\n\nEither way, thanks for sharing! If I do find some interesting results on the use of different EfficientNet models I’ll come back to the discussion to share that.",
      "votes": null
    },
    {
      "id": "751617",
      "postDate": "02/20/2020 10:41:31",
      "content": "<p>I tried it out:\nMy first attempt was a simple model taken from <a href=\"https://www.kaggle.com/kaushal2896/bengali-graphemes-starter-eda-multi-output-cnn\">this</a> trained on 64x64 images cropped by an algorithm I developed myself. I later tried out training a separate network (same architecture, single head) just for the Grapheme Root. I suspected that this would allow the network to focus on the Root, which might improve results.</p>\n\n<p>It turned out it didn't: Despite focusing on the Grapheme Root the network actually learned slower than it did when it also had to focus on Vowel and Consonant (Root improved less quickly in a single NN). It also started overfitting more quickly (without reaching the same accuracy as it had with a multihead classifier) and performed worse in a submission. \nSo it seems that whatever the network finds that is useful for Vowel and Consonant is so useful for Root as well that combining them into one classifier is the way to go. </p>",
      "rawMarkdown": "I tried it out:\nMy first attempt was a simple model taken from [this](https://www.kaggle.com/kaushal2896/bengali-graphemes-starter-eda-multi-output-cnn) trained on 64x64 images cropped by an algorithm I developed myself. I later tried out training a separate network (same architecture, single head) just for the Grapheme Root. I suspected that this would allow the network to focus on the Root, which might improve results.\n\nIt turned out it didn't: Despite focusing on the Grapheme Root the network actually learned slower than it did when it also had to focus on Vowel and Consonant (Root improved less quickly in a single NN). It also started overfitting more quickly (without reaching the same accuracy as it had with a multihead classifier) and performed worse in a submission. \nSo it seems that whatever the network finds that is useful for Vowel and Consonant is so useful for Root as well that combining them into one classifier is the way to go.",
      "votes": null
    },
    {
      "id": "751626",
      "postDate": "02/20/2020 10:52:04",
      "content": "<p>Good feedback Lars,\nProbably the correlations are pretty strong between root and the other outputs</p>",
      "rawMarkdown": "Good feedback Lars,\nProbably the correlations are pretty strong between root and the other outputs",
      "votes": null
    },
    {
      "id": "753437",
      "postDate": "02/22/2020 07:20:45",
      "content": "<p>Also done some tail changing, not to much improvement</p>\n\n<p>For me \nmodel with 186(split to 3 output) output better than three independent branch.</p>",
      "rawMarkdown": "Also done some tail changing, not to much improvement\n\nFor me \nmodel with 186(split to 3 output) output better than three independent branch.",
      "votes": null
    },
    {
      "id": "753441",
      "postDate": "02/22/2020 07:27:36",
      "content": "<p>Good input. Thanks <a href=\"/cswwp347724\">@cswwp347724</a> </p>",
      "rawMarkdown": "Good input. Thanks @cswwp347724",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 749343,
      "author_name": "optimo",
      "author_url": "",
      "post_date": "02/18/2020 15:59:56",
      "content": "<p>so you are performing 3 softmax separately? or you don't output probabilities?</p>\n\n<p>I personally tried 3 simple heads for each tasks.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 749374,
      "author_name": "vladvdv",
      "author_url": "",
      "post_date": "02/18/2020 16:31:31",
      "content": "<p>Yes. I am doing 3 softmax functions separately (first 168 output neurons one softmax for grapheme root,  for the next 11 output neurons one softmax for grapheme vocal and for the last 7 output neurons one softmax for grapheme consonant).\n<a href=\"/optimo\">@optimo</a> Have you tried 3 different classifiers for each category ? This is next on my todo list</p>",
      "votes": null,
      "replies": [
        {
          "id": 750225,
          "author_name": "optimo",
          "author_url": "",
          "post_date": "02/19/2020 07:55:54",
          "content": "<p>nope it's would be too time consuming... I did not try</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 749534,
      "author_name": "greatgamedota",
      "author_url": "",
      "post_date": "02/18/2020 18:57:39",
      "content": "<p>This was discussed a little at the beginning of this comp: <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123956\">https://www.kaggle.com/c/bengaliai-cv19/discussion/123956</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 749546,
          "author_name": "vladvdv",
          "author_url": "",
          "post_date": "02/18/2020 19:12:50",
          "content": "<p>Cool, good to know, there are some interesting points of view there</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 750043,
      "author_name": "moximo13",
      "author_url": "",
      "post_date": "02/19/2020 04:51:45",
      "content": "<p>Hi <a href=\"/vladvdv\">@vladvdv</a> ,</p>\n\n<p>May I ask how many epochs do you train and do you use any schedulers? \nSo far, I try with different models for each class and there is a slight improvement:\n<code>\nBackbone: Resnext50\nAug: Cutout/CutMix, RandomRotate\nOptim: SGD\nScheduler: Consine\nCV: 0.969785\nLB: 0.9672\n</code>\nHowever, my model for classify Grapheme reaches its highest only <code>0.97</code>. </p>\n\n<p>Now, I'm finding a way to go to <code>0.98</code>.\nIt's really frustrating since we have 20 days left.  😭 😭 </p>",
      "votes": null,
      "replies": [
        {
          "id": 750403,
          "author_name": "vladvdv",
          "author_url": "",
          "post_date": "02/19/2020 10:45:15",
          "content": "<p>Hi <a href=\"/moximo13\">@moximo13</a> ,\nI am using ReduceLROnPlateau by a factor of 0.9 with patience 5 epochs, also I am using Adam as a optimizer\nMy frustration is on the time duration on training. It takes very long and I do not have enough time to tune every thing I got on the list so I choose based on my experience the most important aspects</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 750335,
      "author_name": "kupchanski",
      "author_url": "",
      "post_date": "02/19/2020 10:02:14",
      "content": "<p>Multilabel classification?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 750402,
      "author_name": "maxlenormand",
      "author_url": "",
      "post_date": "02/19/2020 10:44:38",
      "content": "<p>I’m curious about why not trying a higher EfficientNet?</p>\n\n<p>B4 seems to give some really promising results on my side, though overfitting seems to be something to deal with. Is this the reason you’re staying at a B0, to have a more robust model?</p>",
      "votes": null,
      "replies": [
        {
          "id": 750409,
          "author_name": "vladvdv",
          "author_url": "",
          "post_date": "02/19/2020 10:49:50",
          "content": "<p>Hi <a href=\"/maxlenormand\">@maxlenormand</a> \nEfficientNet are pretty big architectures, especially if you go to B4, from my experience they give good results when your image has a bigger resolution, and to scale our small images images at 300x300, to augment them enough so the model will have diversity so it won't overfit, to wait around 150 epochs for each fold so the model will converge, it will take forever on my GPU</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 750418,
          "author_name": "maxlenormand",
          "author_url": "",
          "post_date": "02/19/2020 10:58:39",
          "content": "<p>Interesting!</p>\n\n<p>Once again I’m not at the same CV as you, I’m struggling to pass 0.963 and above, so maybe you’re right and I should lower the architecture I use and see if that helps.</p>\n\n<p>Either way, thanks for sharing! If I do find some interesting results on the use of different EfficientNet models I’ll come back to the discussion to share that.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 751617,
      "author_name": "larswigger",
      "author_url": "",
      "post_date": "02/20/2020 10:41:31",
      "content": "<p>I tried it out:\nMy first attempt was a simple model taken from <a href=\"https://www.kaggle.com/kaushal2896/bengali-graphemes-starter-eda-multi-output-cnn\">this</a> trained on 64x64 images cropped by an algorithm I developed myself. I later tried out training a separate network (same architecture, single head) just for the Grapheme Root. I suspected that this would allow the network to focus on the Root, which might improve results.</p>\n\n<p>It turned out it didn't: Despite focusing on the Grapheme Root the network actually learned slower than it did when it also had to focus on Vowel and Consonant (Root improved less quickly in a single NN). It also started overfitting more quickly (without reaching the same accuracy as it had with a multihead classifier) and performed worse in a submission. \nSo it seems that whatever the network finds that is useful for Vowel and Consonant is so useful for Root as well that combining them into one classifier is the way to go. </p>",
      "votes": null,
      "replies": [
        {
          "id": 751626,
          "author_name": "vladvdv",
          "author_url": "",
          "post_date": "02/20/2020 10:52:04",
          "content": "<p>Good feedback Lars,\nProbably the correlations are pretty strong between root and the other outputs</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 753437,
      "author_name": "cswwp347724",
      "author_url": "",
      "post_date": "02/22/2020 07:20:45",
      "content": "<p>Also done some tail changing, not to much improvement</p>\n\n<p>For me \nmodel with 186(split to 3 output) output better than three independent branch.</p>",
      "votes": null,
      "replies": [
        {
          "id": 753441,
          "author_name": "vladvdv",
          "author_url": "",
          "post_date": "02/22/2020 07:27:36",
          "content": "<p>Good input. Thanks <a href=\"/cswwp347724\">@cswwp347724</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "749335": "So far, in my tests configurations I designed architectures based on a final layer of 186 neurons (168 grapheme root outputs + 11 grapheme vowels + 7 grapheme consonants).\nI had tuned arhitectures like:\n- resnext50\n- resnext101\n- efficientnet-b0\n\nActivation functions:\n- relu \n- leaky relu\n\nPreprocessing techniques:\n- affine transformation\n- cutout\n- cutmix\n- mixup\n- bluring techniques\n\nOptimizers:\n- Adam\n- AdamW\n- RAdam \n\nMy best results so far on the public leaderboard are 0.9750 on a 5 fold configuration and for my current training I hope a boost up to 0.9770.\nDid you guys which had time to experiment more, tried designing 3 different arhitectures for each output category ? (grapheme root, grapheme vowels, grapheme consonants). What results did you have ? I am courious if is this the key of going 0.98 +\nMy intuition tells my that the architecture with the 3 outputs combined should be more performant and robust (the model should understand also the correlations between outputs)",
    "749343": "so you are performing 3 softmax separately? or you don't output probabilities?\n\nI personally tried 3 simple heads for each tasks.",
    "749374": "Yes. I am doing 3 softmax functions separately (first 168 output neurons one softmax for grapheme root,  for the next 11 output neurons one softmax for grapheme vocal and for the last 7 output neurons one softmax for grapheme consonant).\n@optimo Have you tried 3 different classifiers for each category ? This is next on my todo list",
    "749534": "This was discussed a little at the beginning of this comp: https://www.kaggle.com/c/bengaliai-cv19/discussion/123956",
    "749546": "Cool, good to know, there are some interesting points of view there",
    "750043": "Hi @vladvdv ,\n\nMay I ask how many epochs do you train and do you use any schedulers? \nSo far, I try with different models for each class and there is a slight improvement:\n```\nBackbone: Resnext50\nAug: Cutout/CutMix, RandomRotate\nOptim: SGD\nScheduler: Consine\nCV: 0.969785\nLB: 0.9672\n```\nHowever, my model for classify Grapheme reaches its highest only `0.97`. \n\nNow, I'm finding a way to go to `0.98`.\nIt's really frustrating since we have 20 days left.  😭 😭",
    "750225": "nope it's would be too time consuming... I did not try",
    "750335": "Multilabel classification?",
    "750402": "I’m curious about why not trying a higher EfficientNet?\n\nB4 seems to give some really promising results on my side, though overfitting seems to be something to deal with. Is this the reason you’re staying at a B0, to have a more robust model?",
    "750403": "Hi @moximo13 ,\nI am using ReduceLROnPlateau by a factor of 0.9 with patience 5 epochs, also I am using Adam as a optimizer\nMy frustration is on the time duration on training. It takes very long and I do not have enough time to tune every thing I got on the list so I choose based on my experience the most important aspects",
    "750409": "Hi @maxlenormand \nEfficientNet are pretty big architectures, especially if you go to B4, from my experience they give good results when your image has a bigger resolution, and to scale our small images images at 300x300, to augment them enough so the model will have diversity so it won't overfit, to wait around 150 epochs for each fold so the model will converge, it will take forever on my GPU",
    "750418": "Interesting!\n\nOnce again I’m not at the same CV as you, I’m struggling to pass 0.963 and above, so maybe you’re right and I should lower the architecture I use and see if that helps.\n\nEither way, thanks for sharing! If I do find some interesting results on the use of different EfficientNet models I’ll come back to the discussion to share that.",
    "751617": "I tried it out:\nMy first attempt was a simple model taken from [this](https://www.kaggle.com/kaushal2896/bengali-graphemes-starter-eda-multi-output-cnn) trained on 64x64 images cropped by an algorithm I developed myself. I later tried out training a separate network (same architecture, single head) just for the Grapheme Root. I suspected that this would allow the network to focus on the Root, which might improve results.\n\nIt turned out it didn't: Despite focusing on the Grapheme Root the network actually learned slower than it did when it also had to focus on Vowel and Consonant (Root improved less quickly in a single NN). It also started overfitting more quickly (without reaching the same accuracy as it had with a multihead classifier) and performed worse in a submission. \nSo it seems that whatever the network finds that is useful for Vowel and Consonant is so useful for Root as well that combining them into one classifier is the way to go.",
    "751626": "Good feedback Lars,\nProbably the correlations are pretty strong between root and the other outputs",
    "753437": "Also done some tail changing, not to much improvement\n\nFor me \nmodel with 186(split to 3 output) output better than three independent branch.",
    "753441": "Good input. Thanks @cswwp347724"
  },
  "source": "meta"
}