{
  "id": 224623,
  "title": "What is the correct approach to choose a model architecture?",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/224623",
  "author_name": "Abhishek Gupta",
  "post_date": "2021-03-09T08:45:01.786000",
  "votes": 4,
  "comment_count": 8,
  "views": 0,
  "content": "<p>In case of medical imaging, one should not lose any important feature of an image while feature extraction, so it is important to choose a model architecture to use such that we could identify important features and make predictions out of it. <br>\nAs far as I know, most of us would be using a hit and trial approach where one chooses a particular model architecture like Resnet or Effnet and make validation out of it, check the accuracy, if it doesn't work out then move to other architectures. There could be many architectures possible starting from Resnet 50 to Resnet 300 and EffNet B0 to Effnet B7. Now NFNets are the other good architectures to choose from.</p>\n<p>Based on the deep learning task, my approach of choosing a model is to see the accuracies of various models on a standard dataset like Imagenet in case of classification. Check my machine configuration if it could run the respective model or not and based on these two factors, I choose the model and proceed with it. If it doesn't work out then move to another architecture. But this too is a hit and trial approach.</p>\n<p><strong>From this thread I just want to know some organised way to choose a model architecture and what are the differnet approaches to choose a model architecture for a deep learning task.</strong></p>\n<p>Hoping to get some good insights from this thread</p>",
  "messages": [
    {
      "id": 1231790,
      "postDate": "2021-03-09T08:45:01.787Z",
      "content": "<p>In case of medical imaging, one should not lose any important feature of an image while feature extraction, so it is important to choose a model architecture to use such that we could identify important features and make predictions out of it. <br>\nAs far as I know, most of us would be using a hit and trial approach where one chooses a particular model architecture like Resnet or Effnet and make validation out of it, check the accuracy, if it doesn't work out then move to other architectures. There could be many architectures possible starting from Resnet 50 to Resnet 300 and EffNet B0 to Effnet B7. Now NFNets are the other good architectures to choose from.</p>\n<p>Based on the deep learning task, my approach of choosing a model is to see the accuracies of various models on a standard dataset like Imagenet in case of classification. Check my machine configuration if it could run the respective model or not and based on these two factors, I choose the model and proceed with it. If it doesn't work out then move to another architecture. But this too is a hit and trial approach.</p>\n<p><strong>From this thread I just want to know some organised way to choose a model architecture and what are the differnet approaches to choose a model architecture for a deep learning task.</strong></p>\n<p>Hoping to get some good insights from this thread</p>",
      "rawMarkdown": "In case of medical imaging, one should not lose any important feature of an image while feature extraction, so it is important to choose a model architecture to use such that we could identify important features and make predictions out of it. \nAs far as I know, most of us would be using a hit and trial approach where one chooses a particular model architecture like Resnet or Effnet and make validation out of it, check the accuracy, if it doesn't work out then move to other architectures. There could be many architectures possible starting from Resnet 50 to Resnet 300 and EffNet B0 to Effnet B7. Now NFNets are the other good architectures to choose from.\n\n\nBased on the deep learning task, my approach of choosing a model is to see the accuracies of various models on a standard dataset like Imagenet in case of classification. Check my machine configuration if it could run the respective model or not and based on these two factors, I choose the model and proceed with it. If it doesn't work out then move to another architecture. But this too is a hit and trial approach.\n\n**From this thread I just want to know some organised way to choose a model architecture and what are the differnet approaches to choose a model architecture for a deep learning task.**\n\nHoping to get some good insights from this thread",
      "votes": 3
    },
    {
      "id": 1233516,
      "postDate": "2021-03-10T13:50:26.487Z",
      "content": "<p>Hi! I personally tested 5 models at once (TF2,Keras) in one notebook,then went on to choose the ones with the best performance , then went on to fine tune them .</p>\n<p>Ended up using EfficientNetB6 and B7 ,LB 96.4. Exception was also performing well .</p>\n<p>I even tried a few models from tensorflow hub BiT_m_r101x3_1 , performance was good but not worth the computational power , plus the model is like 1.7gb. So basically it's an experimental thing.</p>\n<p>On the other hand it's seems some just check which model has best score , then the go knew to use that one .</p>",
      "rawMarkdown": "Hi! I personally tested 5 models at once (TF2,Keras) in one notebook,then went on to choose the ones with the best performance , then went on to fine tune them .\n\nEnded up using EfficientNetB6 and B7 ,LB 96.4. Exception was also performing well .\n\nI even tried a few models from tensorflow hub BiT_m_r101x3_1 , performance was good but not worth the computational power , plus the model is like 1.7gb. So basically it's an experimental thing.\n\nOn the other hand it's seems some just check which model has best score , then the go knew to use that one ."
    },
    {
      "id": 1233985,
      "postDate": "2021-03-10T20:38:58.647Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 1233994,
          "postDate": "2021-03-10T20:52:52.810Z",
          "content": "<p>Hi, i did'nt really get your question, can you elaborate?</p>",
          "rawMarkdown": "Hi, i did'nt really get your question, can you elaborate?"
        },
        {
          "id": 1234000,
          "postDate": "2021-03-10T21:03:13.070Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1234005,
          "postDate": "2021-03-10T21:11:56.140Z",
          "content": "<p>Yes , u are better using a subset just to check , after using the subset and choosing a model ,train the model with the full dataset using stratified kfold (5 folds).</p>\n<p>Yes it takes time , a lot if time tweeking and finetuning.</p>",
          "rawMarkdown": "Yes , u are better using a subset just to check , after using the subset and choosing a model ,train the model with the full dataset using stratified kfold (5 folds).\n\nYes it takes time , a lot if time tweeking and finetuning."
        },
        {
          "id": 1234023,
          "postDate": "2021-03-10T21:37:52.033Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1234028,
          "postDate": "2021-03-10T21:46:41.763Z",
          "content": "<p>Great! It's actually very vital .It makes sure your training is more even .Also check TTA , Test Time Augmentation also very helpful.</p>",
          "rawMarkdown": "Great! It's actually very vital .It makes sure your training is more even .Also check TTA , Test Time Augmentation also very helpful."
        },
        {
          "id": 1234115,
          "postDate": "2021-03-11T00:34:28.493Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1233516,
      "author_name": "Kudzayi Matinyarare",
      "author_url": "",
      "post_date": "2021-03-10T13:50:26.487000",
      "content": "<p>Hi! I personally tested 5 models at once (TF2,Keras) in one notebook,then went on to choose the ones with the best performance , then went on to fine tune them .</p>\n<p>Ended up using EfficientNetB6 and B7 ,LB 96.4. Exception was also performing well .</p>\n<p>I even tried a few models from tensorflow hub BiT_m_r101x3_1 , performance was good but not worth the computational power , plus the model is like 1.7gb. So basically it's an experimental thing.</p>\n<p>On the other hand it's seems some just check which model has best score , then the go knew to use that one .</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1233985,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-03-10T20:38:58.647000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1233994,
          "author_name": "Kudzayi Matinyarare",
          "author_url": "",
          "post_date": "2021-03-10T20:52:52.810000",
          "content": "<p>Hi, i did'nt really get your question, can you elaborate?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1234000,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-03-10T21:03:13.070000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1234005,
          "author_name": "Kudzayi Matinyarare",
          "author_url": "",
          "post_date": "2021-03-10T21:11:56.140000",
          "content": "<p>Yes , u are better using a subset just to check , after using the subset and choosing a model ,train the model with the full dataset using stratified kfold (5 folds).</p>\n<p>Yes it takes time , a lot if time tweeking and finetuning.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1234023,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-03-10T21:37:52.033000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1234028,
          "author_name": "Kudzayi Matinyarare",
          "author_url": "",
          "post_date": "2021-03-10T21:46:41.763000",
          "content": "<p>Great! It's actually very vital .It makes sure your training is more even .Also check TTA , Test Time Augmentation also very helpful.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1234115,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-03-11T00:34:28.493000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1231790": "In case of medical imaging, one should not lose any important feature of an image while feature extraction, so it is important to choose a model architecture to use such that we could identify important features and make predictions out of it. \nAs far as I know, most of us would be using a hit and trial approach where one chooses a particular model architecture like Resnet or Effnet and make validation out of it, check the accuracy, if it doesn't work out then move to other architectures. There could be many architectures possible starting from Resnet 50 to Resnet 300 and EffNet B0 to Effnet B7. Now NFNets are the other good architectures to choose from.\n\n\nBased on the deep learning task, my approach of choosing a model is to see the accuracies of various models on a standard dataset like Imagenet in case of classification. Check my machine configuration if it could run the respective model or not and based on these two factors, I choose the model and proceed with it. If it doesn't work out then move to another architecture. But this too is a hit and trial approach.\n\n**From this thread I just want to know some organised way to choose a model architecture and what are the differnet approaches to choose a model architecture for a deep learning task.**\n\nHoping to get some good insights from this thread",
    "1233516": "Hi! I personally tested 5 models at once (TF2,Keras) in one notebook,then went on to choose the ones with the best performance , then went on to fine tune them .\n\nEnded up using EfficientNetB6 and B7 ,LB 96.4. Exception was also performing well .\n\nI even tried a few models from tensorflow hub BiT_m_r101x3_1 , performance was good but not worth the computational power , plus the model is like 1.7gb. So basically it's an experimental thing.\n\nOn the other hand it's seems some just check which model has best score , then the go knew to use that one .",
    "1233985": ""
  }
}