{
  "id": 209110,
  "title": "i am getting submission scoring error",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/209110",
  "author_name": "",
  "post_date": "2021-01-06T10:50:18.950315200Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I have tried submitting my notebook multiple times but i am getting this <strong>submission scoring error</strong>. My submission format and headers are same but i have no idea why its coming.Please checkout my notebook</p>\n<p><a href=\"https://www.kaggle.com/anandsm7/starter-code-pytorch-efficientnetb4-0-87\" target=\"_blank\">https://www.kaggle.com/anandsm7/starter-code-pytorch-efficientnetb4-0-87</a></p>\n<p>Thanks</p>",
  "messages": [
    {
      "id": "1140892",
      "postDate": "01/06/2021 10:50:18",
      "content": "<p>I have tried submitting my notebook multiple times but i am getting this <strong>submission scoring error</strong>. My submission format and headers are same but i have no idea why its coming.Please checkout my notebook</p>\n<p><a href=\"https://www.kaggle.com/anandsm7/starter-code-pytorch-efficientnetb4-0-87\" target=\"_blank\">https://www.kaggle.com/anandsm7/starter-code-pytorch-efficientnetb4-0-87</a></p>\n<p>Thanks</p>",
      "rawMarkdown": "I have tried submitting my notebook multiple times but i am getting this **submission scoring error**. My submission format and headers are same but i have no idea why its coming.Please checkout my notebook\n\nhttps://www.kaggle.com/anandsm7/starter-code-pytorch-efficientnetb4-0-87\n\nThanks",
      "votes": null
    },
    {
      "id": "1141098",
      "postDate": "01/06/2021 13:53:11",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/anandsm7\" target=\"_blank\">@anandsm7</a> your submission file is looking ok to me , Better you check the dtype for the 'label' column it should be 'int'</p>",
      "rawMarkdown": "Hi @anandsm7 your submission file is looking ok to me , Better you check the dtype for the 'label' column it should be 'int'",
      "votes": null
    },
    {
      "id": "1141692",
      "postDate": "01/06/2021 20:49:34",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/anandsm7\" target=\"_blank\">@anandsm7</a>,<br>\nproblem in you program is that your submission needs <strong>to be able to make prediction to multiple images</strong> because after submitting your notebook it will go throught same procces as when you submit it but there is over 15 000 images that makes final LB score.</p>\n<hr>\n<p>I would reccomend something like this:</p>\n<pre><code>samp_sub = pd.read_csv(\"../input/cassava-leaf-disease-classification/sample_submission.csv\")\n\ndef final_submission(model, dataloader, device, samp_sub):\n   model.eval()\n   img_preds_all = []\n\n   for index, (img) in enumerate(dataloader):\n      img = img.to(device).float()\n      output = model(img)\n\n      img_preds_all += [torch.softmax(output, 1).detach().cpu().numpy()]\n\n   img_preds_all = np.concatenate(img_preds_all, axis=0)\n\n   samp_sub[\"label\"] = np.argmax(img_preds_all, axis=1)\n   samp_sub.to_csv(\"submission.csv\", index=False)\n</code></pre>",
      "rawMarkdown": "Hello @anandsm7,\nproblem in you program is that your submission needs **to be able to make prediction to multiple images** because after submitting your notebook it will go throught same procces as when you submit it but there is over 15 000 images that makes final LB score.\n___\nI would reccomend something like this:\n```\nsamp_sub = pd.read_csv(\"../input/cassava-leaf-disease-classification/sample_submission.csv\")\n\ndef final_submission(model, dataloader, device, samp_sub):\n   model.eval()\n   img_preds_all = []\n\n   for index, (img) in enumerate(dataloader):\n      img = img.to(device).float()\n      output = model(img)\n\n      img_preds_all += [torch.softmax(output, 1).detach().cpu().numpy()]\n\n   img_preds_all = np.concatenate(img_preds_all, axis=0)\n\n   samp_sub[\"label\"] = np.argmax(img_preds_all, axis=1)\n   samp_sub.to_csv(\"submission.csv\", index=False)\n\n```",
      "votes": null
    },
    {
      "id": "1144387",
      "postDate": "01/08/2021 12:38:44",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/filchy\" target=\"_blank\">@filchy</a> <a href=\"https://www.kaggle.com/ssarkar445\" target=\"_blank\">@ssarkar445</a> </p>\n<p>I have tried the following</p>\n<pre><code>def final_submission(model, dataloader, device, samp_sub):\n    model.eval()\n    img_preds_all = []\n    with torch.no_grad():\n        for data in dataloader:\n            inputs = data['image'].to(device,dtype=torch.float)\n            output = model(img)\n            img_preds_all += [torch.softmax(output, 1).detach().cpu().numpy()]\n\n    img_preds_all = np.concatenate(img_preds_all, axis=0)\n    samp_sub[\"label\"] = np.argmax(img_preds_all, axis=1)\n    samp_sub.to_csv(\"submission.csv\", index=False)\nsamp_sub = pd.read_csv(\"../input/cassava-leaf-disease-classification/sample_submission.csv\")\nsamp_sub_img_id = samp_sub.image_id.values.tolist()\nsamp_sub_imgs = [os.path.join(TEST_PATH,img) for img in samp_sub_img_id]\n\n\nsamp_sub_dataset = LeafDiseaseDataset(img_path=samp_sub_imgs,\n                                    targets=None,\n                                    transforms=val_aug,\n                                    resize=(224,224),\n                                    test = True\n                                      )\nsamp_sub_loader = DataLoader(samp_sub_dataset,\n                        batch_size=batch_size,\n                        num_workers=num_workers)\n\n\nfinal_submission(test_model,samp_sub_loader,device,samp_sub)\n</code></pre>\n<p>Still the same issue<br>\nMy results<br>\nimage_id     label<br>\n0     2216849948.jpg  2</p>\n<p>INFO<br>\n<br>\nRangeIndex: 1 entries, 0 to 0<br>\nData columns (total 2 columns):<br>\n #   Column    Non-Null Count  Dtype <br>\n---  ------    --------------  ----- <br>\n 0   image_id  1 non-null      object<br>\n 1   label     1 non-null      int64 <br>\ndtypes: int64(1), object(1)<br>\nmemory usage: 144.0+ bytes</p>",
      "rawMarkdown": "Hi @filchy @ssarkar445 \n\nI have tried the following\n```\ndef final_submission(model, dataloader, device, samp_sub):\n    model.eval()\n    img_preds_all = []\n    with torch.no_grad():\n        for data in dataloader:\n            inputs = data['image'].to(device,dtype=torch.float)\n            output = model(img)\n            img_preds_all += [torch.softmax(output, 1).detach().cpu().numpy()]\n\n    img_preds_all = np.concatenate(img_preds_all, axis=0)\n    samp_sub[\"label\"] = np.argmax(img_preds_all, axis=1)\n    samp_sub.to_csv(\"submission.csv\", index=False)\nsamp_sub = pd.read_csv(\"../input/cassava-leaf-disease-classification/sample_submission.csv\")\nsamp_sub_img_id = samp_sub.image_id.values.tolist()\nsamp_sub_imgs = [os.path.join(TEST_PATH,img) for img in samp_sub_img_id]\n\n\nsamp_sub_dataset = LeafDiseaseDataset(img_path=samp_sub_imgs,\n                                    targets=None,\n                                    transforms=val_aug,\n                                    resize=(224,224),\n                                    test = True\n                                      )\nsamp_sub_loader = DataLoader(samp_sub_dataset,\n                        batch_size=batch_size,\n                        num_workers=num_workers)\n\n\nfinal_submission(test_model,samp_sub_loader,device,samp_sub)\n```\nStill the same issue\nMy results\nimage_id \tlabel\n0 \t2216849948.jpg \t2\n\nINFO\n<class 'pandas.core.frame.DataFrame'>\nRangeIndex: 1 entries, 0 to 0\nData columns (total 2 columns):\n #   Column    Non-Null Count  Dtype \n---  ------    --------------  ----- \n 0   image_id  1 non-null      object\n 1   label     1 non-null      int64 \ndtypes: int64(1), object(1)\nmemory usage: 144.0+ bytes",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1141098,
      "author_name": "ssarkar445",
      "author_url": "",
      "post_date": "01/06/2021 13:53:11",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/anandsm7\" target=\"_blank\">@anandsm7</a> your submission file is looking ok to me , Better you check the dtype for the 'label' column it should be 'int'</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1141692,
      "author_name": "filchy",
      "author_url": "",
      "post_date": "01/06/2021 20:49:34",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/anandsm7\" target=\"_blank\">@anandsm7</a>,<br>\nproblem in you program is that your submission needs <strong>to be able to make prediction to multiple images</strong> because after submitting your notebook it will go throught same procces as when you submit it but there is over 15 000 images that makes final LB score.</p>\n<hr>\n<p>I would reccomend something like this:</p>\n<pre><code>samp_sub = pd.read_csv(\"../input/cassava-leaf-disease-classification/sample_submission.csv\")\n\ndef final_submission(model, dataloader, device, samp_sub):\n   model.eval()\n   img_preds_all = []\n\n   for index, (img) in enumerate(dataloader):\n      img = img.to(device).float()\n      output = model(img)\n\n      img_preds_all += [torch.softmax(output, 1).detach().cpu().numpy()]\n\n   img_preds_all = np.concatenate(img_preds_all, axis=0)\n\n   samp_sub[\"label\"] = np.argmax(img_preds_all, axis=1)\n   samp_sub.to_csv(\"submission.csv\", index=False)\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1144387,
      "author_name": "anandsm7",
      "author_url": "",
      "post_date": "01/08/2021 12:38:44",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/filchy\" target=\"_blank\">@filchy</a> <a href=\"https://www.kaggle.com/ssarkar445\" target=\"_blank\">@ssarkar445</a> </p>\n<p>I have tried the following</p>\n<pre><code>def final_submission(model, dataloader, device, samp_sub):\n    model.eval()\n    img_preds_all = []\n    with torch.no_grad():\n        for data in dataloader:\n            inputs = data['image'].to(device,dtype=torch.float)\n            output = model(img)\n            img_preds_all += [torch.softmax(output, 1).detach().cpu().numpy()]\n\n    img_preds_all = np.concatenate(img_preds_all, axis=0)\n    samp_sub[\"label\"] = np.argmax(img_preds_all, axis=1)\n    samp_sub.to_csv(\"submission.csv\", index=False)\nsamp_sub = pd.read_csv(\"../input/cassava-leaf-disease-classification/sample_submission.csv\")\nsamp_sub_img_id = samp_sub.image_id.values.tolist()\nsamp_sub_imgs = [os.path.join(TEST_PATH,img) for img in samp_sub_img_id]\n\n\nsamp_sub_dataset = LeafDiseaseDataset(img_path=samp_sub_imgs,\n                                    targets=None,\n                                    transforms=val_aug,\n                                    resize=(224,224),\n                                    test = True\n                                      )\nsamp_sub_loader = DataLoader(samp_sub_dataset,\n                        batch_size=batch_size,\n                        num_workers=num_workers)\n\n\nfinal_submission(test_model,samp_sub_loader,device,samp_sub)\n</code></pre>\n<p>Still the same issue<br>\nMy results<br>\nimage_id     label<br>\n0     2216849948.jpg  2</p>\n<p>INFO<br>\n<br>\nRangeIndex: 1 entries, 0 to 0<br>\nData columns (total 2 columns):<br>\n #   Column    Non-Null Count  Dtype <br>\n---  ------    --------------  ----- <br>\n 0   image_id  1 non-null      object<br>\n 1   label     1 non-null      int64 <br>\ndtypes: int64(1), object(1)<br>\nmemory usage: 144.0+ bytes</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1140892": "I have tried submitting my notebook multiple times but i am getting this **submission scoring error**. My submission format and headers are same but i have no idea why its coming.Please checkout my notebook\n\nhttps://www.kaggle.com/anandsm7/starter-code-pytorch-efficientnetb4-0-87\n\nThanks",
    "1141098": "Hi @anandsm7 your submission file is looking ok to me , Better you check the dtype for the 'label' column it should be 'int'",
    "1141692": "Hello @anandsm7,\nproblem in you program is that your submission needs **to be able to make prediction to multiple images** because after submitting your notebook it will go throught same procces as when you submit it but there is over 15 000 images that makes final LB score.\n___\nI would reccomend something like this:\n```\nsamp_sub = pd.read_csv(\"../input/cassava-leaf-disease-classification/sample_submission.csv\")\n\ndef final_submission(model, dataloader, device, samp_sub):\n   model.eval()\n   img_preds_all = []\n\n   for index, (img) in enumerate(dataloader):\n      img = img.to(device).float()\n      output = model(img)\n\n      img_preds_all += [torch.softmax(output, 1).detach().cpu().numpy()]\n\n   img_preds_all = np.concatenate(img_preds_all, axis=0)\n\n   samp_sub[\"label\"] = np.argmax(img_preds_all, axis=1)\n   samp_sub.to_csv(\"submission.csv\", index=False)\n\n```",
    "1144387": "Hi @filchy @ssarkar445 \n\nI have tried the following\n```\ndef final_submission(model, dataloader, device, samp_sub):\n    model.eval()\n    img_preds_all = []\n    with torch.no_grad():\n        for data in dataloader:\n            inputs = data['image'].to(device,dtype=torch.float)\n            output = model(img)\n            img_preds_all += [torch.softmax(output, 1).detach().cpu().numpy()]\n\n    img_preds_all = np.concatenate(img_preds_all, axis=0)\n    samp_sub[\"label\"] = np.argmax(img_preds_all, axis=1)\n    samp_sub.to_csv(\"submission.csv\", index=False)\nsamp_sub = pd.read_csv(\"../input/cassava-leaf-disease-classification/sample_submission.csv\")\nsamp_sub_img_id = samp_sub.image_id.values.tolist()\nsamp_sub_imgs = [os.path.join(TEST_PATH,img) for img in samp_sub_img_id]\n\n\nsamp_sub_dataset = LeafDiseaseDataset(img_path=samp_sub_imgs,\n                                    targets=None,\n                                    transforms=val_aug,\n                                    resize=(224,224),\n                                    test = True\n                                      )\nsamp_sub_loader = DataLoader(samp_sub_dataset,\n                        batch_size=batch_size,\n                        num_workers=num_workers)\n\n\nfinal_submission(test_model,samp_sub_loader,device,samp_sub)\n```\nStill the same issue\nMy results\nimage_id \tlabel\n0 \t2216849948.jpg \t2\n\nINFO\n<class 'pandas.core.frame.DataFrame'>\nRangeIndex: 1 entries, 0 to 0\nData columns (total 2 columns):\n #   Column    Non-Null Count  Dtype \n---  ------    --------------  ----- \n 0   image_id  1 non-null      object\n 1   label     1 non-null      int64 \ndtypes: int64(1), object(1)\nmemory usage: 144.0+ bytes"
  },
  "source": "meta"
}