{
  "id": 134055,
  "title": "Submission CSV not found / Submission scoring error",
  "url": "/competitions/bengaliai-cv19/discussion/134055",
  "author_name": "",
  "post_date": "2020-03-05T16:47:10.536010800Z",
  "votes": 2,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hi all,</p>\n\n<p>this is my first Kaggle competition and I am struggling with the submission (as the topic title suggests).\nI read other topics about the same kind of issue and I found helpful fixes, but I still do not manage to submit successfully.</p>\n\n<p>I made sure that:\n- I load one parquet at a time and I run predictions in batches of size 32;\n- I delete all the intermediate output as soon as they are not useful anymore;\n- I run successfully the notebook on the training set, with no memory issues;\n- The row ids are in the correct order, and I list the outputs in the correct order (consonant_diacritic, grapheme_root, vowel_diacritic).\n- The dtypes of the output data frame are 'object' and 'int64';\n- The predicted labels have the correct range: consonant_diacritic from 0 to 6, grapheme_root from 0 to 168, vowel_diacritic from 0 to 10.</p>\n\n<p>Oddly, the submission is successful if I save the 'test.csv' adding a dummy column 'target' with, for instance, all zeros. The rest of the notebook runs also, indicating no CPU/RAM issues.</p>\n\n<p>Does anyone have any idea?</p>\n\n<p>I attached also the test output on the small set available for testing purposes.</p>\n\n<p>Cheers,\nAlessandro</p>",
  "messages": [
    {
      "id": "764615",
      "postDate": "03/05/2020 16:47:10",
      "content": "<p>Hi all,</p>\n\n<p>this is my first Kaggle competition and I am struggling with the submission (as the topic title suggests).\nI read other topics about the same kind of issue and I found helpful fixes, but I still do not manage to submit successfully.</p>\n\n<p>I made sure that:\n- I load one parquet at a time and I run predictions in batches of size 32;\n- I delete all the intermediate output as soon as they are not useful anymore;\n- I run successfully the notebook on the training set, with no memory issues;\n- The row ids are in the correct order, and I list the outputs in the correct order (consonant_diacritic, grapheme_root, vowel_diacritic).\n- The dtypes of the output data frame are 'object' and 'int64';\n- The predicted labels have the correct range: consonant_diacritic from 0 to 6, grapheme_root from 0 to 168, vowel_diacritic from 0 to 10.</p>\n\n<p>Oddly, the submission is successful if I save the 'test.csv' adding a dummy column 'target' with, for instance, all zeros. The rest of the notebook runs also, indicating no CPU/RAM issues.</p>\n\n<p>Does anyone have any idea?</p>\n\n<p>I attached also the test output on the small set available for testing purposes.</p>\n\n<p>Cheers,\nAlessandro</p>",
      "rawMarkdown": "Hi all,\n\nthis is my first Kaggle competition and I am struggling with the submission (as the topic title suggests).\nI read other topics about the same kind of issue and I found helpful fixes, but I still do not manage to submit successfully.\n\nI made sure that:\n- I load one parquet at a time and I run predictions in batches of size 32;\n- I delete all the intermediate output as soon as they are not useful anymore;\n- I run successfully the notebook on the training set, with no memory issues;\n- The row ids are in the correct order, and I list the outputs in the correct order (consonant_diacritic, grapheme_root, vowel_diacritic).\n- The dtypes of the output data frame are 'object' and 'int64';\n- The predicted labels have the correct range: consonant_diacritic from 0 to 6, grapheme_root from 0 to 168, vowel_diacritic from 0 to 10.\n\nOddly, the submission is successful if I save the 'test.csv' adding a dummy column 'target' with, for instance, all zeros. The rest of the notebook runs also, indicating no CPU/RAM issues.\n\nDoes anyone have any idea?\n\nI attached also the test output on the small set available for testing purposes.\n\nCheers,\nAlessandro",
      "votes": null
    },
    {
      "id": "764854",
      "postDate": "03/06/2020 01:36:40",
      "content": "<p>I have just solved this, thougt it could be a gpu oom problem. maybe you can try some code to reduce gpu memory</p>",
      "rawMarkdown": "I have just solved this, thougt it could be a gpu oom problem. maybe you can try some code to reduce gpu memory",
      "votes": null
    },
    {
      "id": "765133",
      "postDate": "03/06/2020 09:33:24",
      "content": "<p>Hi, thank you for your answer. Unfortunately, I'm quite sure that this is not the problem. If I add the following code:</p>\n\n<p>```\ndef get_fake_preds():\n    fake= pd.read_csv('/kaggle/input/bengaliai-cv19/test.csv')\n    fake= fake.drop(columns=['image_id', 'component'])\n    fake['target'] = 0\n    return fake</p>\n\n<p>LONG_PREDS = get_fake_preds()\nLONG_PREDS.to_csv('submission.csv', index=False)\n```</p>\n\n<p>The submission is successful, where I run the same computations, except I replace the output I generate with a dummy one.</p>",
      "rawMarkdown": "Hi, thank you for your answer. Unfortunately, I'm quite sure that this is not the problem. If I add the following code:\n\n```\ndef get_fake_preds():\n    fake= pd.read_csv('/kaggle/input/bengaliai-cv19/test.csv')\n    fake= fake.drop(columns=['image_id', 'component'])\n    fake['target'] = 0\n    return fake\n\nLONG_PREDS = get_fake_preds()\nLONG_PREDS.to_csv('submission.csv', index=False)\n```\n\nThe submission is successful, where I run the same computations, except I replace the output I generate with a dummy one.",
      "votes": null
    },
    {
      "id": "765697",
      "postDate": "03/07/2020 01:16:09",
      "content": "<p>in my case, i just reduce the batch size and added <code>with torch.no_grad():</code> (didn't realized i missed this before) and it worked well</p>",
      "rawMarkdown": "in my case, i just reduce the batch size and added `with torch.no_grad():` (didn't realized i missed this before) and it worked well",
      "votes": null
    },
    {
      "id": "765802",
      "postDate": "03/07/2020 06:00:42",
      "content": "<p>Did you solve the problem? (solved)</p>",
      "rawMarkdown": "Did you solve the problem? (solved)",
      "votes": null
    },
    {
      "id": "765849",
      "postDate": "03/07/2020 07:51:38",
      "content": "<p>from this <a href=\"https://www.kaggle.com/reasat/evaluating-model-performance-on-the-test-set\">NoteBooks</a> we can find that there is more than 12 image in the test, which lead to scoring error.</p>",
      "rawMarkdown": "from this [NoteBooks](https://www.kaggle.com/reasat/evaluating-model-performance-on-the-test-set) we can find that there is more than 12 image in the test, which lead to scoring error.",
      "votes": null
    },
    {
      "id": "765921",
      "postDate": "03/07/2020 11:13:22",
      "content": "<p>I met the same question,have you handled this question??</p>",
      "rawMarkdown": "I met the same question,have you handled this question??",
      "votes": null
    },
    {
      "id": "766457",
      "postDate": "03/08/2020 06:58:56",
      "content": "<p>I have the same problem and read all 37 topics related to \"submission error\" 😅  still not solved yet.\nAlready spend much time to solve the error without any information, not on improving my models😩 </p>",
      "rawMarkdown": "I have the same problem and read all 37 topics related to \"submission error\" 😅  still not solved yet.\nAlready spend much time to solve the error without any information, not on improving my models😩",
      "votes": null
    },
    {
      "id": "767865",
      "postDate": "03/10/2020 07:21:32",
      "content": "<p>when you submit the code, it will replace the test parquet file in the background. So there is more than 12 images need to predict. You must keep the code of reading image from the parquet file.</p>",
      "rawMarkdown": "when you submit the code, it will replace the test parquet file in the background. So there is more than 12 images need to predict. You must keep the code of reading image from the parquet file.",
      "votes": null
    },
    {
      "id": "770687",
      "postDate": "03/13/2020 09:05:06",
      "content": "<p>I am indeed taking into account a variable size for the input, creating the dataframe from the 4 parquet files. If they change, the submission table is going, in theory, to be according to the input data. But still no luck :/</p>",
      "rawMarkdown": "I am indeed taking into account a variable size for the input, creating the dataframe from the 4 parquet files. If they change, the submission table is going, in theory, to be according to the input data. But still no luck :/",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 764854,
      "author_name": "nemo0111",
      "author_url": "",
      "post_date": "03/06/2020 01:36:40",
      "content": "<p>I have just solved this, thougt it could be a gpu oom problem. maybe you can try some code to reduce gpu memory</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 765133,
      "author_name": "alessandrogentile",
      "author_url": "",
      "post_date": "03/06/2020 09:33:24",
      "content": "<p>Hi, thank you for your answer. Unfortunately, I'm quite sure that this is not the problem. If I add the following code:</p>\n\n<p>```\ndef get_fake_preds():\n    fake= pd.read_csv('/kaggle/input/bengaliai-cv19/test.csv')\n    fake= fake.drop(columns=['image_id', 'component'])\n    fake['target'] = 0\n    return fake</p>\n\n<p>LONG_PREDS = get_fake_preds()\nLONG_PREDS.to_csv('submission.csv', index=False)\n```</p>\n\n<p>The submission is successful, where I run the same computations, except I replace the output I generate with a dummy one.</p>",
      "votes": null,
      "replies": [
        {
          "id": 765697,
          "author_name": "nemo0111",
          "author_url": "",
          "post_date": "03/07/2020 01:16:09",
          "content": "<p>in my case, i just reduce the batch size and added <code>with torch.no_grad():</code> (didn't realized i missed this before) and it worked well</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 765802,
      "author_name": "junedgar",
      "author_url": "",
      "post_date": "03/07/2020 06:00:42",
      "content": "<p>Did you solve the problem? (solved)</p>",
      "votes": null,
      "replies": [
        {
          "id": 765849,
          "author_name": "junedgar",
          "author_url": "",
          "post_date": "03/07/2020 07:51:38",
          "content": "<p>from this <a href=\"https://www.kaggle.com/reasat/evaluating-model-performance-on-the-test-set\">NoteBooks</a> we can find that there is more than 12 image in the test, which lead to scoring error.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 765921,
      "author_name": "thefatcat",
      "author_url": "",
      "post_date": "03/07/2020 11:13:22",
      "content": "<p>I met the same question,have you handled this question??</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 766457,
      "author_name": "shinsei66",
      "author_url": "",
      "post_date": "03/08/2020 06:58:56",
      "content": "<p>I have the same problem and read all 37 topics related to \"submission error\" 😅  still not solved yet.\nAlready spend much time to solve the error without any information, not on improving my models😩 </p>",
      "votes": null,
      "replies": [
        {
          "id": 767865,
          "author_name": "junedgar",
          "author_url": "",
          "post_date": "03/10/2020 07:21:32",
          "content": "<p>when you submit the code, it will replace the test parquet file in the background. So there is more than 12 images need to predict. You must keep the code of reading image from the parquet file.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 770687,
      "author_name": "alessandrogentile",
      "author_url": "",
      "post_date": "03/13/2020 09:05:06",
      "content": "<p>I am indeed taking into account a variable size for the input, creating the dataframe from the 4 parquet files. If they change, the submission table is going, in theory, to be according to the input data. But still no luck :/</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "764615": "Hi all,\n\nthis is my first Kaggle competition and I am struggling with the submission (as the topic title suggests).\nI read other topics about the same kind of issue and I found helpful fixes, but I still do not manage to submit successfully.\n\nI made sure that:\n- I load one parquet at a time and I run predictions in batches of size 32;\n- I delete all the intermediate output as soon as they are not useful anymore;\n- I run successfully the notebook on the training set, with no memory issues;\n- The row ids are in the correct order, and I list the outputs in the correct order (consonant_diacritic, grapheme_root, vowel_diacritic).\n- The dtypes of the output data frame are 'object' and 'int64';\n- The predicted labels have the correct range: consonant_diacritic from 0 to 6, grapheme_root from 0 to 168, vowel_diacritic from 0 to 10.\n\nOddly, the submission is successful if I save the 'test.csv' adding a dummy column 'target' with, for instance, all zeros. The rest of the notebook runs also, indicating no CPU/RAM issues.\n\nDoes anyone have any idea?\n\nI attached also the test output on the small set available for testing purposes.\n\nCheers,\nAlessandro",
    "764854": "I have just solved this, thougt it could be a gpu oom problem. maybe you can try some code to reduce gpu memory",
    "765133": "Hi, thank you for your answer. Unfortunately, I'm quite sure that this is not the problem. If I add the following code:\n\n```\ndef get_fake_preds():\n    fake= pd.read_csv('/kaggle/input/bengaliai-cv19/test.csv')\n    fake= fake.drop(columns=['image_id', 'component'])\n    fake['target'] = 0\n    return fake\n\nLONG_PREDS = get_fake_preds()\nLONG_PREDS.to_csv('submission.csv', index=False)\n```\n\nThe submission is successful, where I run the same computations, except I replace the output I generate with a dummy one.",
    "765697": "in my case, i just reduce the batch size and added `with torch.no_grad():` (didn't realized i missed this before) and it worked well",
    "765802": "Did you solve the problem? (solved)",
    "765849": "from this [NoteBooks](https://www.kaggle.com/reasat/evaluating-model-performance-on-the-test-set) we can find that there is more than 12 image in the test, which lead to scoring error.",
    "765921": "I met the same question,have you handled this question??",
    "766457": "I have the same problem and read all 37 topics related to \"submission error\" 😅  still not solved yet.\nAlready spend much time to solve the error without any information, not on improving my models😩",
    "767865": "when you submit the code, it will replace the test parquet file in the background. So there is more than 12 images need to predict. You must keep the code of reading image from the parquet file.",
    "770687": "I am indeed taking into account a variable size for the input, creating the dataframe from the 4 parquet files. If they change, the submission table is going, in theory, to be according to the input data. But still no luck :/"
  },
  "source": "meta"
}