{
  "id": 123956,
  "title": "One model with 3 ouputs vs 3 different models",
  "url": "/competitions/bengaliai-cv19/discussion/123956",
  "author_name": "Luis Herrera",
  "post_date": "2019-12-31T21:52:03.367000",
  "votes": 10,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Checking some of the current kernels I can see that some of them are training 3 separate models one for each target. Others are training one model but with 3 outputs. What are the pros/cons of each approach?\nThanks in advance</p>",
  "messages": [
    {
      "id": 707459,
      "postDate": "2019-12-31T21:52:03.367Z",
      "content": "<p>Checking some of the current kernels I can see that some of them are training 3 separate models one for each target. Others are training one model but with 3 outputs. What are the pros/cons of each approach?\nThanks in advance</p>",
      "rawMarkdown": "Checking some of the current kernels I can see that some of them are training 3 separate models one for each target. Others are training one model but with 3 outputs. What are the pros/cons of each approach?\nThanks in advance",
      "votes": 10
    },
    {
      "id": 710391,
      "postDate": "2020-01-04T17:17:25.230Z",
      "content": "<p>I have tried a unique custom made CNN (a unique combo of Inception and ResNet)</p>\n\n<h1>My Models</h1>\n\n<h2>3models</h2>\n\n<p><strong>Model</strong>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1550889%2F318694ade8a6ac22bac3e7e49a7e4c91%2Fmodel1.png?generation=1578157473472546&amp;alt=media\" alt=\"\">\n<strong>Only output layer changes for different model accordingly</strong></p>\n\n<h3>Graphemes model</h3>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1550889%2F818cfe531895232d5bf614d7fb7ffed8%2F3m_rootloss.png?generation=1578156939668534&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1550889%2Faf794941e0f5e4d85ce07fc3fbabd14c%2F3m_root%20accuracy.png?generation=1578156938700819&amp;alt=media\" alt=\"\"></p>\n\n<h3>Vowel model</h3>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1550889%2Fbbe32716fad026cd2e89b21aa469f5ca%2F3m_vowelloss.png?generation=1578157039488941&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1550889%2F13b6706a66407b7c4b9a26b85e0a3300%2F3m_vowelaccu.png?generation=1578157043613339&amp;alt=media\" alt=\"\"></p>\n\n<h3>consonant model</h3>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1550889%2F885de8de891119e319d1dfacbd40c47e%2F3m_consoloss.png?generation=1578157109272965&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1550889%2Facea537ade7569cb40f9af275a6a3276%2F3m_consoaccu.png?generation=1578157111023179&amp;alt=media\" alt=\"\"></p>\n\n<h2>1Model 3Output</h2>\n\n<p><strong>Model</strong>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1550889%2F251dd3d7a2d9bd124d19538c3a55fd28%2F3hmodel1.png?generation=1578157570669634&amp;alt=media\" alt=\"\"></p>\n\n<p><strong>Plot</strong>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1550889%2Fc055de851479c8d3d433e65e35dab9eb%2F3hloss.png?generation=1578157183018791&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1550889%2Fb3d9e96243a1da1bef02dae3177e7fe2%2F3haccu.png?generation=1578157184905331&amp;alt=media\" alt=\"\"></p>\n\n<h1>My Findings</h1>\n\n<p>-&gt; Models both give similar accuracy score with slight better points for 1Model 3Outpput. The model with 3 output will be better memory-optimized!! <strong>No Repetition of Convolution layers</strong>, Convolution layers are to extract the features and FCN is to classify them or to learn them!!, One Disadvantage what I can see is <strong>We cant set different epochs for 3 models Imagine 2 of the outputs are giving good result and the 3rd doesn't and when you wait for 3rd other two overfits !!!!</strong>, what will you do !! probably try to manage them with dropouts !! dealy other two models fitting!!</p>\n\n<p>Sorry If there are some mistakes in the above findings always open for a discussion 🤗 </p>\n\n<p>My similar work -&gt; <a href=\"https://www.kaggle.com/chekoduadarsh/bengali-ai-design-resnet-layer-by-layer/output\">https://www.kaggle.com/chekoduadarsh/bengali-ai-design-resnet-layer-by-layer/output</a></p>",
      "rawMarkdown": "I have tried a unique custom made CNN (a unique combo of Inception and ResNet)\n\n# My Models\n## 3models\n**Model**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1550889%2F318694ade8a6ac22bac3e7e49a7e4c91%2Fmodel1.png?generation=1578157473472546&amp;alt=media)\n**Only output layer changes for different model accordingly**\n###Graphemes model \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1550889%2F818cfe531895232d5bf614d7fb7ffed8%2F3m_rootloss.png?generation=1578156939668534&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1550889%2Faf794941e0f5e4d85ce07fc3fbabd14c%2F3m_root%20accuracy.png?generation=1578156938700819&amp;alt=media)\n###Vowel model \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1550889%2Fbbe32716fad026cd2e89b21aa469f5ca%2F3m_vowelloss.png?generation=1578157039488941&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1550889%2F13b6706a66407b7c4b9a26b85e0a3300%2F3m_vowelaccu.png?generation=1578157043613339&amp;alt=media)\n###consonant model \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1550889%2F885de8de891119e319d1dfacbd40c47e%2F3m_consoloss.png?generation=1578157109272965&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1550889%2Facea537ade7569cb40f9af275a6a3276%2F3m_consoaccu.png?generation=1578157111023179&amp;alt=media)\n\n## 1Model 3Output\n**Model**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1550889%2F251dd3d7a2d9bd124d19538c3a55fd28%2F3hmodel1.png?generation=1578157570669634&amp;alt=media)\n\n**Plot**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1550889%2Fc055de851479c8d3d433e65e35dab9eb%2F3hloss.png?generation=1578157183018791&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1550889%2Fb3d9e96243a1da1bef02dae3177e7fe2%2F3haccu.png?generation=1578157184905331&amp;alt=media)\n\n\n\n# My Findings\n\n-&gt; Models both give similar accuracy score with slight better points for 1Model 3Outpput. The model with 3 output will be better memory-optimized!! **No Repetition of Convolution layers**, Convolution layers are to extract the features and FCN is to classify them or to learn them!!, One Disadvantage what I can see is **We cant set different epochs for 3 models Imagine 2 of the outputs are giving good result and the 3rd doesn't and when you wait for 3rd other two overfits !!!!**, what will you do !! probably try to manage them with dropouts !! dealy other two models fitting!!\n\n\nSorry If there are some mistakes in the above findings always open for a discussion 🤗 \n\nMy similar work -&gt; https://www.kaggle.com/chekoduadarsh/bengali-ai-design-resnet-layer-by-layer/output",
      "votes": 8,
      "replies": [
        {
          "id": 710424,
          "postDate": "2020-01-04T17:56:17.593Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        },
        {
          "id": 710454,
          "postDate": "2020-01-04T18:54:49.243Z",
          "content": "<p><a href=\"/robga\">@robga</a> true we can do that, I just googled it 😃 </p>",
          "rawMarkdown": "@robga true we can do that, I just googled it 😃 "
        }
      ]
    },
    {
      "id": 707992,
      "postDate": "2020-01-01T20:11:13.733Z",
      "content": "<p>I have tried a lot of different tree-based approach (RandomForest, XGBoost, LGBoost, ExtraTrees...) with 3 separate models and it was not as successful as a Deep Learning approach with 3 outputs.</p>\n\n<p>I will discuss my opinion of positive and negative sides of having a single model</p>\n\n<p><strong>Positive</strong> : \n- Obvious one: <strong>easier to train</strong>, to manipulate and to save/export (one single file). However, 3 different RandomForest are not that hard to calculate\n- Having 3 outputs in a model will take into account that there is a <strong>strong correlation between the grapheme-vowel-consonant triplet</strong>. In an ideal situation, I would even try to design this relationship with a Bayesian Network but I cannot even imagine trying this with our dataset.. If for example, with only the provided data, our model has a strong belief that a given vowel is present, then this can be used to strengthen its belief in his prediction of the grapheme whereas in a 3 models-scenario, these 3 models do not interact with each other, hence there is not this exchange of information.\n- On a Deep Learning point of view, linked grapheme, vowel and consonants will tend to <strong>share the same features</strong>. So encoding the image in the same way to discriminate these 3 labels seems to make sense</p>\n\n<p><strong>Negative</strong> : \n- I strongly agree with Ahmed Imtiaz when he talks about <strong>risks of overfitting.</strong> The benefit of having 3 outputs of a single model is that you can deduct some of them based on the others. But these relationships may be <strong>totally artificial, inducing a huge risk of overfitting</strong>. Whereas, if you don't communicate about these rules, you have no risk of overfitting in this sense (don't get me wrong, you could still overfit, espacially given the fact that the dataset is strongly imbalanced).\n- Having one single model \"<em>kind of</em>\" imply that the <strong>3 labels are distributed the same way</strong>. For example, the consonant could be discriminated with only a simple linear model while the vowels may need a deep architecture. Using one single model for the 3 outputs means that you will, in a way, <strong>use the same \"<em>type of model</em>\" for these 3</strong>, even if it is an overkill. And as this competition has the notion of efficiency taken into account, it is important to note: if you can make it simpler, do it. </p>",
      "rawMarkdown": "I have tried a lot of different tree-based approach (RandomForest, XGBoost, LGBoost, ExtraTrees...) with 3 separate models and it was not as successful as a Deep Learning approach with 3 outputs.\n\nI will discuss my opinion of positive and negative sides of having a single model\n\n**Positive** : \n- Obvious one: **easier to train**, to manipulate and to save/export (one single file). However, 3 different RandomForest are not that hard to calculate\n- Having 3 outputs in a model will take into account that there is a **strong correlation between the grapheme-vowel-consonant triplet**. In an ideal situation, I would even try to design this relationship with a Bayesian Network but I cannot even imagine trying this with our dataset.. If for example, with only the provided data, our model has a strong belief that a given vowel is present, then this can be used to strengthen its belief in his prediction of the grapheme whereas in a 3 models-scenario, these 3 models do not interact with each other, hence there is not this exchange of information.\n- On a Deep Learning point of view, linked grapheme, vowel and consonants will tend to **share the same features**. So encoding the image in the same way to discriminate these 3 labels seems to make sense\n\n**Negative** : \n- I strongly agree with Ahmed Imtiaz when he talks about **risks of overfitting.** The benefit of having 3 outputs of a single model is that you can deduct some of them based on the others. But these relationships may be **totally artificial, inducing a huge risk of overfitting**. Whereas, if you don't communicate about these rules, you have no risk of overfitting in this sense (don't get me wrong, you could still overfit, espacially given the fact that the dataset is strongly imbalanced).\n- Having one single model \"*kind of*\" imply that the **3 labels are distributed the same way**. For example, the consonant could be discriminated with only a simple linear model while the vowels may need a deep architecture. Using one single model for the 3 outputs means that you will, in a way, **use the same \"*type of model*\" for these 3**, even if it is an overkill. And as this competition has the notion of efficiency taken into account, it is important to note: if you can make it simpler, do it. ",
      "votes": 3,
      "replies": [
        {
          "id": 714302,
          "postDate": "2020-01-09T09:30:13.763Z",
          "content": "<p>Nice try and thank you for discussing your approach. But I think using pretrained models or our own custom model gives good result than traditional approaches. I tried it but got worst results.</p>",
          "rawMarkdown": "Nice try and thank you for discussing your approach. But I think using pretrained models or our own custom model gives good result than traditional approaches. I tried it but got worst results."
        },
        {
          "id": 726173,
          "postDate": "2020-01-22T22:30:22.330Z",
          "content": "<p>I would be really curious on seeing the results of your more traditional ML approaches <a href=\"/dimartinot\">@dimartinot</a> ! Especially where the model gets ut right / wrong. </p>\n\n<p>As for the 3 model 1 output vs 1 model 3 outputs, I agree on most points, but it might also be interesting to do ensembling with the 2 approaches. There are some heavy limitations (would require running 4 predictions per image) but if the predictions are good in both case, this might be an interesting régularisation technique especially against unseen combinations of symbols. The predictions would need to be quite un-correlated for that to make sense though.</p>\n\n<p>Same for the ML algos. They might serve as some basis for ensembling. It could also be done with giving a lower weight to the predictions from ML algos compared to DL?</p>",
          "rawMarkdown": "I would be really curious on seeing the results of your more traditional ML approaches @dimartinot ! Especially where the model gets ut right / wrong. \n\nAs for the 3 model 1 output vs 1 model 3 outputs, I agree on most points, but it might also be interesting to do ensembling with the 2 approaches. There are some heavy limitations (would require running 4 predictions per image) but if the predictions are good in both case, this might be an interesting régularisation technique especially against unseen combinations of symbols. The predictions would need to be quite un-correlated for that to make sense though.\n\nSame for the ML algos. They might serve as some basis for ensembling. It could also be done with giving a lower weight to the predictions from ML algos compared to DL?",
          "votes": 1
        }
      ]
    },
    {
      "id": 707701,
      "postDate": "2020-01-01T10:21:16.827Z",
      "content": "<p>This is a question open to research!\nFrom the top of my head for 1 model with 3 outputs:</p>\n\n<p><strong>Pro:</strong> If the visual features corresponding to one of the 3 outputs (suppose the vowel diacritic output) is not clear in the image (meaning the diacritic is hard to recognize visually), the correlation of that particular visual feature with the other output labels (like how the grapheme root changes subtly everytime that vowel diacritic is present) could help resolve the presence of the occluded/unclear output.\n<strong>Con:</strong> Since there are graphemes in the test set which are not in training, you could end up overfitting on the training set distribution of grapheme root-vowel diacritic-consonant diacritic combinations.</p>\n\n<p>A balance between the two is what should be perused. I think feature sharing between 3 separate models could help. The models can also end up learning the features themselves but you'll be overparameterizing the models in that case. Separate models help handle the imbalances individually but we can't say for sure whether that is good or bad.</p>",
      "rawMarkdown": "This is a question open to research!\nFrom the top of my head for 1 model with 3 outputs:\n\n**Pro:** If the visual features corresponding to one of the 3 outputs (suppose the vowel diacritic output) is not clear in the image (meaning the diacritic is hard to recognize visually), the correlation of that particular visual feature with the other output labels (like how the grapheme root changes subtly everytime that vowel diacritic is present) could help resolve the presence of the occluded/unclear output.\n**Con:** Since there are graphemes in the test set which are not in training, you could end up overfitting on the training set distribution of grapheme root-vowel diacritic-consonant diacritic combinations.\n\nA balance between the two is what should be perused. I think feature sharing between 3 separate models could help. The models can also end up learning the features themselves but you'll be overparameterizing the models in that case. Separate models help handle the imbalances individually but we can't say for sure whether that is good or bad.",
      "votes": 4,
      "replies": [
        {
          "id": 710135,
          "postDate": "2020-01-04T10:51:20.283Z",
          "content": "<p>+\nYou can try training one model with three outputs in separate iterations for each of the outputs.</p>\n\n<p>Like in a naive sense during training, the gradient w.r.t each model parameter is calculated separately for all three outputs and then the weights are updated in one step.\nYou can split the weight update into three different steps with different number of iteration for each output, like how in GAN training you control the number of discriminator/generator updates to control training dynamics. By separating the updates you can try to balance the training mini-batches for each of the outputs individually.</p>",
          "rawMarkdown": "+\nYou can try training one model with three outputs in separate iterations for each of the outputs.\n\nLike in a naive sense during training, the gradient w.r.t each model parameter is calculated separately for all three outputs and then the weights are updated in one step.\nYou can split the weight update into three different steps with different number of iteration for each output, like how in GAN training you control the number of discriminator/generator updates to control training dynamics. By separating the updates you can try to balance the training mini-batches for each of the outputs individually.",
          "votes": 3
        }
      ]
    },
    {
      "id": 707548,
      "postDate": "2020-01-01T03:31:02.730Z",
      "content": "<p>My initial tests suggest that three different models are performing worse than a single model with three heads. Can't say the pros/cons for sure right now.</p>",
      "rawMarkdown": "My initial tests suggest that three different models are performing worse than a single model with three heads. Can't say the pros/cons for sure right now.",
      "votes": 2
    },
    {
      "id": 723131,
      "postDate": "2020-01-19T14:40:46.137Z",
      "content": "<p>keep up the good work!  upvoted..</p>",
      "rawMarkdown": "\nkeep up the good work!  upvoted.."
    },
    {
      "id": 710449,
      "postDate": "2020-01-04T18:51:53.433Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 710391,
      "author_name": "A/C",
      "author_url": "",
      "post_date": "2020-01-04T17:17:25.230000",
      "content": "<p>I have tried a unique custom made CNN (a unique combo of Inception and ResNet)</p>\n\n<h1>My Models</h1>\n\n<h2>3models</h2>\n\n<p><strong>Model</strong>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1550889%2F318694ade8a6ac22bac3e7e49a7e4c91%2Fmodel1.png?generation=1578157473472546&amp;alt=media\" alt=\"\">\n<strong>Only output layer changes for different model accordingly</strong></p>\n\n<h3>Graphemes model</h3>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1550889%2F818cfe531895232d5bf614d7fb7ffed8%2F3m_rootloss.png?generation=1578156939668534&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1550889%2Faf794941e0f5e4d85ce07fc3fbabd14c%2F3m_root%20accuracy.png?generation=1578156938700819&amp;alt=media\" alt=\"\"></p>\n\n<h3>Vowel model</h3>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1550889%2Fbbe32716fad026cd2e89b21aa469f5ca%2F3m_vowelloss.png?generation=1578157039488941&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1550889%2F13b6706a66407b7c4b9a26b85e0a3300%2F3m_vowelaccu.png?generation=1578157043613339&amp;alt=media\" alt=\"\"></p>\n\n<h3>consonant model</h3>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1550889%2F885de8de891119e319d1dfacbd40c47e%2F3m_consoloss.png?generation=1578157109272965&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1550889%2Facea537ade7569cb40f9af275a6a3276%2F3m_consoaccu.png?generation=1578157111023179&amp;alt=media\" alt=\"\"></p>\n\n<h2>1Model 3Output</h2>\n\n<p><strong>Model</strong>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1550889%2F251dd3d7a2d9bd124d19538c3a55fd28%2F3hmodel1.png?generation=1578157570669634&amp;alt=media\" alt=\"\"></p>\n\n<p><strong>Plot</strong>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1550889%2Fc055de851479c8d3d433e65e35dab9eb%2F3hloss.png?generation=1578157183018791&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1550889%2Fb3d9e96243a1da1bef02dae3177e7fe2%2F3haccu.png?generation=1578157184905331&amp;alt=media\" alt=\"\"></p>\n\n<h1>My Findings</h1>\n\n<p>-&gt; Models both give similar accuracy score with slight better points for 1Model 3Outpput. The model with 3 output will be better memory-optimized!! <strong>No Repetition of Convolution layers</strong>, Convolution layers are to extract the features and FCN is to classify them or to learn them!!, One Disadvantage what I can see is <strong>We cant set different epochs for 3 models Imagine 2 of the outputs are giving good result and the 3rd doesn't and when you wait for 3rd other two overfits !!!!</strong>, what will you do !! probably try to manage them with dropouts !! dealy other two models fitting!!</p>\n\n<p>Sorry If there are some mistakes in the above findings always open for a discussion 🤗 </p>\n\n<p>My similar work -&gt; <a href=\"https://www.kaggle.com/chekoduadarsh/bengali-ai-design-resnet-layer-by-layer/output\">https://www.kaggle.com/chekoduadarsh/bengali-ai-design-resnet-layer-by-layer/output</a></p>",
      "votes": 8,
      "replies": [
        {
          "id": 710424,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-04T17:56:17.593000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 710454,
          "author_name": "A/C",
          "author_url": "",
          "post_date": "2020-01-04T18:54:49.243000",
          "content": "<p><a href=\"/robga\">@robga</a> true we can do that, I just googled it 😃 </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 707992,
      "author_name": "Thomas Di Martino",
      "author_url": "",
      "post_date": "2020-01-01T20:11:13.733000",
      "content": "<p>I have tried a lot of different tree-based approach (RandomForest, XGBoost, LGBoost, ExtraTrees...) with 3 separate models and it was not as successful as a Deep Learning approach with 3 outputs.</p>\n\n<p>I will discuss my opinion of positive and negative sides of having a single model</p>\n\n<p><strong>Positive</strong> : \n- Obvious one: <strong>easier to train</strong>, to manipulate and to save/export (one single file). However, 3 different RandomForest are not that hard to calculate\n- Having 3 outputs in a model will take into account that there is a <strong>strong correlation between the grapheme-vowel-consonant triplet</strong>. In an ideal situation, I would even try to design this relationship with a Bayesian Network but I cannot even imagine trying this with our dataset.. If for example, with only the provided data, our model has a strong belief that a given vowel is present, then this can be used to strengthen its belief in his prediction of the grapheme whereas in a 3 models-scenario, these 3 models do not interact with each other, hence there is not this exchange of information.\n- On a Deep Learning point of view, linked grapheme, vowel and consonants will tend to <strong>share the same features</strong>. So encoding the image in the same way to discriminate these 3 labels seems to make sense</p>\n\n<p><strong>Negative</strong> : \n- I strongly agree with Ahmed Imtiaz when he talks about <strong>risks of overfitting.</strong> The benefit of having 3 outputs of a single model is that you can deduct some of them based on the others. But these relationships may be <strong>totally artificial, inducing a huge risk of overfitting</strong>. Whereas, if you don't communicate about these rules, you have no risk of overfitting in this sense (don't get me wrong, you could still overfit, espacially given the fact that the dataset is strongly imbalanced).\n- Having one single model \"<em>kind of</em>\" imply that the <strong>3 labels are distributed the same way</strong>. For example, the consonant could be discriminated with only a simple linear model while the vowels may need a deep architecture. Using one single model for the 3 outputs means that you will, in a way, <strong>use the same \"<em>type of model</em>\" for these 3</strong>, even if it is an overkill. And as this competition has the notion of efficiency taken into account, it is important to note: if you can make it simpler, do it. </p>",
      "votes": 3,
      "replies": [
        {
          "id": 714302,
          "author_name": "karan",
          "author_url": "",
          "post_date": "2020-01-09T09:30:13.763000",
          "content": "<p>Nice try and thank you for discussing your approach. But I think using pretrained models or our own custom model gives good result than traditional approaches. I tried it but got worst results.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 726173,
          "author_name": "Maxime Lenormand",
          "author_url": "",
          "post_date": "2020-01-22T22:30:22.330000",
          "content": "<p>I would be really curious on seeing the results of your more traditional ML approaches <a href=\"/dimartinot\">@dimartinot</a> ! Especially where the model gets ut right / wrong. </p>\n\n<p>As for the 3 model 1 output vs 1 model 3 outputs, I agree on most points, but it might also be interesting to do ensembling with the 2 approaches. There are some heavy limitations (would require running 4 predictions per image) but if the predictions are good in both case, this might be an interesting régularisation technique especially against unseen combinations of symbols. The predictions would need to be quite un-correlated for that to make sense though.</p>\n\n<p>Same for the ML algos. They might serve as some basis for ensembling. It could also be done with giving a lower weight to the predictions from ML algos compared to DL?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 707701,
      "author_name": "Ahmed Imtiaz Humayun",
      "author_url": "",
      "post_date": "2020-01-01T10:21:16.827000",
      "content": "<p>This is a question open to research!\nFrom the top of my head for 1 model with 3 outputs:</p>\n\n<p><strong>Pro:</strong> If the visual features corresponding to one of the 3 outputs (suppose the vowel diacritic output) is not clear in the image (meaning the diacritic is hard to recognize visually), the correlation of that particular visual feature with the other output labels (like how the grapheme root changes subtly everytime that vowel diacritic is present) could help resolve the presence of the occluded/unclear output.\n<strong>Con:</strong> Since there are graphemes in the test set which are not in training, you could end up overfitting on the training set distribution of grapheme root-vowel diacritic-consonant diacritic combinations.</p>\n\n<p>A balance between the two is what should be perused. I think feature sharing between 3 separate models could help. The models can also end up learning the features themselves but you'll be overparameterizing the models in that case. Separate models help handle the imbalances individually but we can't say for sure whether that is good or bad.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 710135,
          "author_name": "Ahmed Imtiaz Humayun",
          "author_url": "",
          "post_date": "2020-01-04T10:51:20.283000",
          "content": "<p>+\nYou can try training one model with three outputs in separate iterations for each of the outputs.</p>\n\n<p>Like in a naive sense during training, the gradient w.r.t each model parameter is calculated separately for all three outputs and then the weights are updated in one step.\nYou can split the weight update into three different steps with different number of iteration for each output, like how in GAN training you control the number of discriminator/generator updates to control training dynamics. By separating the updates you can try to balance the training mini-batches for each of the outputs individually.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 707548,
      "author_name": "Rafid Abyaad",
      "author_url": "",
      "post_date": "2020-01-01T03:31:02.730000",
      "content": "<p>My initial tests suggest that three different models are performing worse than a single model with three heads. Can't say the pros/cons for sure right now.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 723131,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-19T14:40:46.137000",
      "content": "<p>keep up the good work!  upvoted..</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 710449,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-04T18:51:53.433000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "707459": "Checking some of the current kernels I can see that some of them are training 3 separate models one for each target. Others are training one model but with 3 outputs. What are the pros/cons of each approach?\nThanks in advance",
    "710391": "I have tried a unique custom made CNN (a unique combo of Inception and ResNet)\n\n# My Models\n## 3models\n**Model**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1550889%2F318694ade8a6ac22bac3e7e49a7e4c91%2Fmodel1.png?generation=1578157473472546&amp;alt=media)\n**Only output layer changes for different model accordingly**\n###Graphemes model \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1550889%2F818cfe531895232d5bf614d7fb7ffed8%2F3m_rootloss.png?generation=1578156939668534&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1550889%2Faf794941e0f5e4d85ce07fc3fbabd14c%2F3m_root%20accuracy.png?generation=1578156938700819&amp;alt=media)\n###Vowel model \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1550889%2Fbbe32716fad026cd2e89b21aa469f5ca%2F3m_vowelloss.png?generation=1578157039488941&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1550889%2F13b6706a66407b7c4b9a26b85e0a3300%2F3m_vowelaccu.png?generation=1578157043613339&amp;alt=media)\n###consonant model \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1550889%2F885de8de891119e319d1dfacbd40c47e%2F3m_consoloss.png?generation=1578157109272965&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1550889%2Facea537ade7569cb40f9af275a6a3276%2F3m_consoaccu.png?generation=1578157111023179&amp;alt=media)\n\n## 1Model 3Output\n**Model**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1550889%2F251dd3d7a2d9bd124d19538c3a55fd28%2F3hmodel1.png?generation=1578157570669634&amp;alt=media)\n\n**Plot**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1550889%2Fc055de851479c8d3d433e65e35dab9eb%2F3hloss.png?generation=1578157183018791&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1550889%2Fb3d9e96243a1da1bef02dae3177e7fe2%2F3haccu.png?generation=1578157184905331&amp;alt=media)\n\n\n\n# My Findings\n\n-&gt; Models both give similar accuracy score with slight better points for 1Model 3Outpput. The model with 3 output will be better memory-optimized!! **No Repetition of Convolution layers**, Convolution layers are to extract the features and FCN is to classify them or to learn them!!, One Disadvantage what I can see is **We cant set different epochs for 3 models Imagine 2 of the outputs are giving good result and the 3rd doesn't and when you wait for 3rd other two overfits !!!!**, what will you do !! probably try to manage them with dropouts !! dealy other two models fitting!!\n\n\nSorry If there are some mistakes in the above findings always open for a discussion 🤗 \n\nMy similar work -&gt; https://www.kaggle.com/chekoduadarsh/bengali-ai-design-resnet-layer-by-layer/output",
    "707992": "I have tried a lot of different tree-based approach (RandomForest, XGBoost, LGBoost, ExtraTrees...) with 3 separate models and it was not as successful as a Deep Learning approach with 3 outputs.\n\nI will discuss my opinion of positive and negative sides of having a single model\n\n**Positive** : \n- Obvious one: **easier to train**, to manipulate and to save/export (one single file). However, 3 different RandomForest are not that hard to calculate\n- Having 3 outputs in a model will take into account that there is a **strong correlation between the grapheme-vowel-consonant triplet**. In an ideal situation, I would even try to design this relationship with a Bayesian Network but I cannot even imagine trying this with our dataset.. If for example, with only the provided data, our model has a strong belief that a given vowel is present, then this can be used to strengthen its belief in his prediction of the grapheme whereas in a 3 models-scenario, these 3 models do not interact with each other, hence there is not this exchange of information.\n- On a Deep Learning point of view, linked grapheme, vowel and consonants will tend to **share the same features**. So encoding the image in the same way to discriminate these 3 labels seems to make sense\n\n**Negative** : \n- I strongly agree with Ahmed Imtiaz when he talks about **risks of overfitting.** The benefit of having 3 outputs of a single model is that you can deduct some of them based on the others. But these relationships may be **totally artificial, inducing a huge risk of overfitting**. Whereas, if you don't communicate about these rules, you have no risk of overfitting in this sense (don't get me wrong, you could still overfit, espacially given the fact that the dataset is strongly imbalanced).\n- Having one single model \"*kind of*\" imply that the **3 labels are distributed the same way**. For example, the consonant could be discriminated with only a simple linear model while the vowels may need a deep architecture. Using one single model for the 3 outputs means that you will, in a way, **use the same \"*type of model*\" for these 3**, even if it is an overkill. And as this competition has the notion of efficiency taken into account, it is important to note: if you can make it simpler, do it. ",
    "707701": "This is a question open to research!\nFrom the top of my head for 1 model with 3 outputs:\n\n**Pro:** If the visual features corresponding to one of the 3 outputs (suppose the vowel diacritic output) is not clear in the image (meaning the diacritic is hard to recognize visually), the correlation of that particular visual feature with the other output labels (like how the grapheme root changes subtly everytime that vowel diacritic is present) could help resolve the presence of the occluded/unclear output.\n**Con:** Since there are graphemes in the test set which are not in training, you could end up overfitting on the training set distribution of grapheme root-vowel diacritic-consonant diacritic combinations.\n\nA balance between the two is what should be perused. I think feature sharing between 3 separate models could help. The models can also end up learning the features themselves but you'll be overparameterizing the models in that case. Separate models help handle the imbalances individually but we can't say for sure whether that is good or bad.",
    "707548": "My initial tests suggest that three different models are performing worse than a single model with three heads. Can't say the pros/cons for sure right now.",
    "723131": "\nkeep up the good work!  upvoted..",
    "710449": ""
  }
}