{
  "id": 124243,
  "title": "Multi-Output or Multi-Models?",
  "url": "/competitions/bengaliai-cv19/discussion/124243",
  "author_name": "",
  "post_date": "2020-01-02T20:40:31.659130700Z",
  "votes": 1,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Looks like there are 2 solutions: multi-output one model or one output three models. Did anyone try both of them? Which one have better accuracy? I was working the 'one output three models' choice but the accuracy is not looking good.</p>",
  "messages": [
    {
      "id": "708921",
      "postDate": "01/02/2020 20:40:31",
      "content": "<p>Looks like there are 2 solutions: multi-output one model or one output three models. Did anyone try both of them? Which one have better accuracy? I was working the 'one output three models' choice but the accuracy is not looking good.</p>",
      "rawMarkdown": "Looks like there are 2 solutions: multi-output one model or one output three models. Did anyone try both of them? Which one have better accuracy? I was working the 'one output three models' choice but the accuracy is not looking good.",
      "votes": null
    },
    {
      "id": "708947",
      "postDate": "01/02/2020 21:17:39",
      "content": "<p>I have not started yet, but I am asking the same question. One advantatge lf using 1 model over three separated is that we need less processing power since there is a shared learning in the model. Another advantatge in using 1 model is that image recognition ( I am supossing the use of CNNs) constist of several convolutional filters, each filter detects a different feature in the image, the first filters are trained to train more general features such as edges, taking this into account I think that even though there are three characteristics to predict the more general features are the same and we can take advantatge of this quality by sharing the first layers of the network. </p>",
      "rawMarkdown": "I have not started yet, but I am asking the same question. One advantatge lf using 1 model over three separated is that we need less processing power since there is a shared learning in the model. Another advantatge in using 1 model is that image recognition ( I am supossing the use of CNNs) constist of several convolutional filters, each filter detects a different feature in the image, the first filters are trained to train more general features such as edges, taking this into account I think that even though there are three characteristics to predict the more general features are the same and we can take advantatge of this quality by sharing the first layers of the network.",
      "votes": null
    },
    {
      "id": "708949",
      "postDate": "01/02/2020 21:31:01",
      "content": "<p>My thinking was: the conv layers that is on top of the FCN will extract features and the FCN will handle the features and make prediction. If they share the conv layers, then the features will be shared. The FCN will have to learn which ones are for grapheme_root, vowel_diacritic, or consonant_diacritic. That will take longer training times and make lower accuracy.</p>\n\n<p>Not sure whether it is correct or not. But looks like multi-output is the most popular choice in the public kernels area.</p>",
      "rawMarkdown": "My thinking was: the conv layers that is on top of the FCN will extract features and the FCN will handle the features and make prediction. If they share the conv layers, then the features will be shared. The FCN will have to learn which ones are for grapheme_root, vowel_diacritic, or consonant_diacritic. That will take longer training times and make lower accuracy.\n\nNot sure whether it is correct or not. But looks like multi-output is the most popular choice in the public kernels area.",
      "votes": null
    },
    {
      "id": "708971",
      "postDate": "01/02/2020 22:06:28",
      "content": "<p>Oops. This is a duplicate thread. Just found out. <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123956\">click me</a></p>",
      "rawMarkdown": "Oops. This is a duplicate thread. Just found out. [click me](https://www.kaggle.com/c/bengaliai-cv19/discussion/123956)",
      "votes": null
    },
    {
      "id": "709852",
      "postDate": "01/04/2020 02:18:29",
      "content": "<p>Multi-output is better, I think. Final submit must run in &lt;= 2 hours with GPU in kernel, multi-model is time-consuming.</p>",
      "rawMarkdown": "Multi-output is better, I think. Final submit must run in &lt;= 2 hours with GPU in kernel, multi-model is time-consuming.",
      "votes": null
    },
    {
      "id": "709952",
      "postDate": "01/04/2020 05:43:51",
      "content": "<p>Multi-output is better.</p>",
      "rawMarkdown": "Multi-output is better.",
      "votes": null
    },
    {
      "id": "709958",
      "postDate": "01/04/2020 06:02:39",
      "content": "<p>I tried NASNet Large(the largest network I know) for the multi-model and it ran in inference time! So inference time doesn't make a difference.</p>",
      "rawMarkdown": "I tried NASNet Large(the largest network I know) for the multi-model and it ran in inference time! So inference time doesn't make a difference.",
      "votes": null
    },
    {
      "id": "729036",
      "postDate": "01/25/2020 16:10:55",
      "content": "<p>I don't think most CNN will ONLY do feature extraction in reality. (depends on your kernel size and connectivity)</p>",
      "rawMarkdown": "I don't think most CNN will ONLY do feature extraction in reality. (depends on your kernel size and connectivity)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 708947,
      "author_name": "polmonroig",
      "author_url": "",
      "post_date": "01/02/2020 21:17:39",
      "content": "<p>I have not started yet, but I am asking the same question. One advantatge lf using 1 model over three separated is that we need less processing power since there is a shared learning in the model. Another advantatge in using 1 model is that image recognition ( I am supossing the use of CNNs) constist of several convolutional filters, each filter detects a different feature in the image, the first filters are trained to train more general features such as edges, taking this into account I think that even though there are three characteristics to predict the more general features are the same and we can take advantatge of this quality by sharing the first layers of the network. </p>",
      "votes": null,
      "replies": [
        {
          "id": 708949,
          "author_name": "unkownhihi",
          "author_url": "",
          "post_date": "01/02/2020 21:31:01",
          "content": "<p>My thinking was: the conv layers that is on top of the FCN will extract features and the FCN will handle the features and make prediction. If they share the conv layers, then the features will be shared. The FCN will have to learn which ones are for grapheme_root, vowel_diacritic, or consonant_diacritic. That will take longer training times and make lower accuracy.</p>\n\n<p>Not sure whether it is correct or not. But looks like multi-output is the most popular choice in the public kernels area.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 729036,
          "author_name": "kokecacao",
          "author_url": "",
          "post_date": "01/25/2020 16:10:55",
          "content": "<p>I don't think most CNN will ONLY do feature extraction in reality. (depends on your kernel size and connectivity)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 708971,
      "author_name": "unkownhihi",
      "author_url": "",
      "post_date": "01/02/2020 22:06:28",
      "content": "<p>Oops. This is a duplicate thread. Just found out. <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123956\">click me</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 709852,
      "author_name": "",
      "author_url": "",
      "post_date": "01/04/2020 02:18:29",
      "content": "<p>Multi-output is better, I think. Final submit must run in &lt;= 2 hours with GPU in kernel, multi-model is time-consuming.</p>",
      "votes": null,
      "replies": [
        {
          "id": 709958,
          "author_name": "unkownhihi",
          "author_url": "",
          "post_date": "01/04/2020 06:02:39",
          "content": "<p>I tried NASNet Large(the largest network I know) for the multi-model and it ran in inference time! So inference time doesn't make a difference.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 709952,
      "author_name": "sangthieuminh",
      "author_url": "",
      "post_date": "01/04/2020 05:43:51",
      "content": "<p>Multi-output is better.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "708921": "Looks like there are 2 solutions: multi-output one model or one output three models. Did anyone try both of them? Which one have better accuracy? I was working the 'one output three models' choice but the accuracy is not looking good.",
    "708947": "I have not started yet, but I am asking the same question. One advantatge lf using 1 model over three separated is that we need less processing power since there is a shared learning in the model. Another advantatge in using 1 model is that image recognition ( I am supossing the use of CNNs) constist of several convolutional filters, each filter detects a different feature in the image, the first filters are trained to train more general features such as edges, taking this into account I think that even though there are three characteristics to predict the more general features are the same and we can take advantatge of this quality by sharing the first layers of the network.",
    "708949": "My thinking was: the conv layers that is on top of the FCN will extract features and the FCN will handle the features and make prediction. If they share the conv layers, then the features will be shared. The FCN will have to learn which ones are for grapheme_root, vowel_diacritic, or consonant_diacritic. That will take longer training times and make lower accuracy.\n\nNot sure whether it is correct or not. But looks like multi-output is the most popular choice in the public kernels area.",
    "708971": "Oops. This is a duplicate thread. Just found out. [click me](https://www.kaggle.com/c/bengaliai-cv19/discussion/123956)",
    "709852": "Multi-output is better, I think. Final submit must run in &lt;= 2 hours with GPU in kernel, multi-model is time-consuming.",
    "709952": "Multi-output is better.",
    "709958": "I tried NASNet Large(the largest network I know) for the multi-model and it ran in inference time! So inference time doesn't make a difference.",
    "729036": "I don't think most CNN will ONLY do feature extraction in reality. (depends on your kernel size and connectivity)"
  },
  "source": "meta"
}