{
  "id": 233211,
  "title": "This contest is a multi-category question or a multi-label question?",
  "url": "/competitions/plant-pathology-2021-fgvc8/discussion/233211",
  "author_name": "",
  "post_date": "2021-04-18T02:34:39.311228500Z",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>The game says the main objective of the competition is to develop machine learning-based models to accurately classify a given leaf image from the test dataset to a particular disease category, and to identify an individual disease from multiple disease symptoms on a single leaf image.</p>\n<p>Some images have more than one type of disease.<br>\nWhen we make a prediction, do we take the first n maximum probabilities, or just the maximum probabilities. Can someone come and tell me</p>",
  "messages": [
    {
      "id": "1276808",
      "postDate": "04/18/2021 02:34:39",
      "content": "<p>The game says the main objective of the competition is to develop machine learning-based models to accurately classify a given leaf image from the test dataset to a particular disease category, and to identify an individual disease from multiple disease symptoms on a single leaf image.</p>\n<p>Some images have more than one type of disease.<br>\nWhen we make a prediction, do we take the first n maximum probabilities, or just the maximum probabilities. Can someone come and tell me</p>",
      "rawMarkdown": "The game says the main objective of the competition is to develop machine learning-based models to accurately classify a given leaf image from the test dataset to a particular disease category, and to identify an individual disease from multiple disease symptoms on a single leaf image.\n\nSome images have more than one type of disease.\nWhen we make a prediction, do we take the first n maximum probabilities, or just the maximum probabilities. Can someone come and tell me",
      "votes": null
    },
    {
      "id": "1278062",
      "postDate": "04/19/2021 14:32:42",
      "content": "<p>Hello Dlidli</p>\n<p>It is something that I ve been thinking for a while. I think the best solution is defining a threshold and then make the prediction any of the illness which surpases that limit. In case there is none I am gonna choose the highest one.</p>",
      "rawMarkdown": "Hello Dlidli\n\nIt is something that I ve been thinking for a while. I think the best solution is defining a threshold and then make the prediction any of the illness which surpases that limit. In case there is none I am gonna choose the highest one.",
      "votes": null
    },
    {
      "id": "1278677",
      "postDate": "04/20/2021 06:57:55",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/ignaciomorenomarn\" target=\"_blank\">@ignaciomorenomarn</a> thanks for your answer, this has helped me tremendously</p>",
      "rawMarkdown": "Hello @ignaciomorenomarn thanks for your answer, this has helped me tremendously",
      "votes": null
    },
    {
      "id": "1279025",
      "postDate": "04/20/2021 14:13:42",
      "content": "<p>hi, do you mean the code should  may like this during the training step ? : <br>\ntrain_out = model(image)<br>\nloss_in = sigmoid(train_out)<br>\ntrain_loss = loss_fn(loss_in, label)<br>\npred_label = train_loss &gt; thredshold</p>",
      "rawMarkdown": "hi, do you mean the code should  may like this during the training step ? : \ntrain_out = model(image)\nloss_in = sigmoid(train_out)\ntrain_loss = loss_fn(loss_in, label)\npred_label = train_loss > thredshold",
      "votes": null
    },
    {
      "id": "1279182",
      "postDate": "04/20/2021 17:08:44",
      "content": "<p>train_out = model(image)<br>\nloss_in = sigmoid(train_out)<br>\ntrain_loss = loss_fn(loss_in, label)<br>\nimagine this example<br>\npred_label = <br>\n[[0,2  0,6  0,2  0,1]<br>\n[0,85  0,1   0,05  0,9]]</p>\n<p>With thredshold=0,7<br>\n[[0  1  0  0] #Because it is the largest (At least it have to be one)<br>\n[1  0  0  1]] #Because they are the larger the threshold</p>\n<p>Another alternative is to do a model for healthy or not healthy and then a second model for the illness but I think it is much more complicated.</p>",
      "rawMarkdown": "train_out = model(image)\nloss_in = sigmoid(train_out)\ntrain_loss = loss_fn(loss_in, label)\nimagine this example\npred_label = \n[[0,2  0,6  0,2  0,1]\n[0,85  0,1   0,05  0,9]]\n\nWith thredshold=0,7\n[[0  1  0  0] #Because it is the largest (At least it have to be one)\n[1  0  0  1]] #Because they are the larger the threshold\n\nAnother alternative is to do a model for healthy or not healthy and then a second model for the illness but I think it is much more complicated.",
      "votes": null
    },
    {
      "id": "1279518",
      "postDate": "04/21/2021 02:11:23",
      "content": "<p>the code i find on github.</p>\n<p>acc = []<br>\naccuracies = []<br>\nbest_threshold = np.zeros(out.shape[1])<br>\nfor i in range(out.shape[1]):<br>\n    y_prob = np.array(out[:,i])<br>\n    for j in threshold:<br>\n        y_pred = [1 if prob&gt;=j else 0 for prob in y_prob]<br>\n        acc.append( matthews_corrcoef(y_test[:,i],y_pred))<br>\n    acc   = np.array(acc)<br>\n    index = np.where(acc==acc.max()) <br>\n    accuracies.append(acc.max()) <br>\n    best_threshold[i] = threshold[index[0][0]]<br>\n    acc = []</p>\n<p><a href=\"https://github.com/suraj-deshmukh/Keras-Multi-Label-Image-Classification/blob/master/model.py\" target=\"_blank\">https://github.com/suraj-deshmukh/Keras-Multi-Label-Image-Classification/blob/master/model.py</a></p>",
      "rawMarkdown": "the code i find on github.\n\nacc = []\naccuracies = []\nbest_threshold = np.zeros(out.shape[1])\nfor i in range(out.shape[1]):\n    y_prob = np.array(out[:,i])\n    for j in threshold:\n        y_pred = [1 if prob>=j else 0 for prob in y_prob]\n        acc.append( matthews_corrcoef(y_test[:,i],y_pred))\n    acc   = np.array(acc)\n    index = np.where(acc==acc.max()) \n    accuracies.append(acc.max()) \n    best_threshold[i] = threshold[index[0][0]]\n    acc = []\n\nhttps://github.com/suraj-deshmukh/Keras-Multi-Label-Image-Classification/blob/master/model.py",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1278062,
      "author_name": "ignaciomorenomarn",
      "author_url": "",
      "post_date": "04/19/2021 14:32:42",
      "content": "<p>Hello Dlidli</p>\n<p>It is something that I ve been thinking for a while. I think the best solution is defining a threshold and then make the prediction any of the illness which surpases that limit. In case there is none I am gonna choose the highest one.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1279025,
          "author_name": "gjzhongdf163com",
          "author_url": "",
          "post_date": "04/20/2021 14:13:42",
          "content": "<p>hi, do you mean the code should  may like this during the training step ? : <br>\ntrain_out = model(image)<br>\nloss_in = sigmoid(train_out)<br>\ntrain_loss = loss_fn(loss_in, label)<br>\npred_label = train_loss &gt; thredshold</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1279182,
          "author_name": "ignaciomorenomarn",
          "author_url": "",
          "post_date": "04/20/2021 17:08:44",
          "content": "<p>train_out = model(image)<br>\nloss_in = sigmoid(train_out)<br>\ntrain_loss = loss_fn(loss_in, label)<br>\nimagine this example<br>\npred_label = <br>\n[[0,2  0,6  0,2  0,1]<br>\n[0,85  0,1   0,05  0,9]]</p>\n<p>With thredshold=0,7<br>\n[[0  1  0  0] #Because it is the largest (At least it have to be one)<br>\n[1  0  0  1]] #Because they are the larger the threshold</p>\n<p>Another alternative is to do a model for healthy or not healthy and then a second model for the illness but I think it is much more complicated.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1279518,
          "author_name": "gjzhongdf163com",
          "author_url": "",
          "post_date": "04/21/2021 02:11:23",
          "content": "<p>the code i find on github.</p>\n<p>acc = []<br>\naccuracies = []<br>\nbest_threshold = np.zeros(out.shape[1])<br>\nfor i in range(out.shape[1]):<br>\n    y_prob = np.array(out[:,i])<br>\n    for j in threshold:<br>\n        y_pred = [1 if prob&gt;=j else 0 for prob in y_prob]<br>\n        acc.append( matthews_corrcoef(y_test[:,i],y_pred))<br>\n    acc   = np.array(acc)<br>\n    index = np.where(acc==acc.max()) <br>\n    accuracies.append(acc.max()) <br>\n    best_threshold[i] = threshold[index[0][0]]<br>\n    acc = []</p>\n<p><a href=\"https://github.com/suraj-deshmukh/Keras-Multi-Label-Image-Classification/blob/master/model.py\" target=\"_blank\">https://github.com/suraj-deshmukh/Keras-Multi-Label-Image-Classification/blob/master/model.py</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1278677,
      "author_name": "dlidli",
      "author_url": "",
      "post_date": "04/20/2021 06:57:55",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/ignaciomorenomarn\" target=\"_blank\">@ignaciomorenomarn</a> thanks for your answer, this has helped me tremendously</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1276808": "The game says the main objective of the competition is to develop machine learning-based models to accurately classify a given leaf image from the test dataset to a particular disease category, and to identify an individual disease from multiple disease symptoms on a single leaf image.\n\nSome images have more than one type of disease.\nWhen we make a prediction, do we take the first n maximum probabilities, or just the maximum probabilities. Can someone come and tell me",
    "1278062": "Hello Dlidli\n\nIt is something that I ve been thinking for a while. I think the best solution is defining a threshold and then make the prediction any of the illness which surpases that limit. In case there is none I am gonna choose the highest one.",
    "1278677": "Hello @ignaciomorenomarn thanks for your answer, this has helped me tremendously",
    "1279025": "hi, do you mean the code should  may like this during the training step ? : \ntrain_out = model(image)\nloss_in = sigmoid(train_out)\ntrain_loss = loss_fn(loss_in, label)\npred_label = train_loss > thredshold",
    "1279182": "train_out = model(image)\nloss_in = sigmoid(train_out)\ntrain_loss = loss_fn(loss_in, label)\nimagine this example\npred_label = \n[[0,2  0,6  0,2  0,1]\n[0,85  0,1   0,05  0,9]]\n\nWith thredshold=0,7\n[[0  1  0  0] #Because it is the largest (At least it have to be one)\n[1  0  0  1]] #Because they are the larger the threshold\n\nAnother alternative is to do a model for healthy or not healthy and then a second model for the illness but I think it is much more complicated.",
    "1279518": "the code i find on github.\n\nacc = []\naccuracies = []\nbest_threshold = np.zeros(out.shape[1])\nfor i in range(out.shape[1]):\n    y_prob = np.array(out[:,i])\n    for j in threshold:\n        y_pred = [1 if prob>=j else 0 for prob in y_prob]\n        acc.append( matthews_corrcoef(y_test[:,i],y_pred))\n    acc   = np.array(acc)\n    index = np.where(acc==acc.max()) \n    accuracies.append(acc.max()) \n    best_threshold[i] = threshold[index[0][0]]\n    acc = []\n\nhttps://github.com/suraj-deshmukh/Keras-Multi-Label-Image-Classification/blob/master/model.py"
  },
  "source": "meta"
}