{
  "id": 123093,
  "title": "Problem scoring my submission file ",
  "url": "/competitions/bengaliai-cv19/discussion/123093",
  "author_name": "",
  "post_date": "2019-12-24T18:17:49.092652Z",
  "votes": 2,
  "comment_count": 13,
  "views": 0,
  "content": "<p>This is my first kernel competition and I'm having difficulties scoring my submission file. I train the model locally and then I upload my model parameters to my kaggle notebook and I make my predictions there. I save the \"submission.csv\" file as instructed, I commit the kernel, I click the button to submit and then it takes too long saying \"scoring your submissions\" and in the end it throws an error. Does anyone have any idea why this happens? Thanks!</p>",
  "messages": [
    {
      "id": "702478",
      "postDate": "12/24/2019 18:17:49",
      "content": "<p>This is my first kernel competition and I'm having difficulties scoring my submission file. I train the model locally and then I upload my model parameters to my kaggle notebook and I make my predictions there. I save the \"submission.csv\" file as instructed, I commit the kernel, I click the button to submit and then it takes too long saying \"scoring your submissions\" and in the end it throws an error. Does anyone have any idea why this happens? Thanks!</p>",
      "rawMarkdown": "This is my first kernel competition and I'm having difficulties scoring my submission file. I train the model locally and then I upload my model parameters to my kaggle notebook and I make my predictions there. I save the \"submission.csv\" file as instructed, I commit the kernel, I click the button to submit and then it takes too long saying \"scoring your submissions\" and in the end it throws an error. Does anyone have any idea why this happens? Thanks!",
      "votes": null
    },
    {
      "id": "702654",
      "postDate": "12/25/2019 01:05:48",
      "content": "<p>Most probably you are running past the allotted GPU time. There is a two hour limit on GPU use during inference.</p>",
      "rawMarkdown": "Most probably you are running past the allotted GPU time. There is a two hour limit on GPU use during inference.",
      "votes": null
    },
    {
      "id": "702700",
      "postDate": "12/25/2019 03:05:18",
      "content": "<blockquote>\n  <p><strong>TahsinReasat wrote:</strong></p>\n  \n  <p>Most probably you are running past the allotted GPU time. There is a two hour limit on GPU use during inference.</p>\n</blockquote>\n\n<p>I have the same problem. I only use CPU, and my kernel only runs for 300-400s, and then it ends normally</p>",
      "rawMarkdown": "&gt; **TahsinReasat wrote:**\n&gt; \n&gt; Most probably you are running past the allotted GPU time. There is a two hour limit on GPU use during inference.\n\nI have the same problem. I only use CPU, and my kernel only runs for 300-400s, and then it ends normally",
      "votes": null
    },
    {
      "id": "702716",
      "postDate": "12/25/2019 03:53:52",
      "content": "<p>try to run prediction for 200000 images by replicating kaggle configuration 1 gpu and 2 cpu with 15gb ram. See how long it takes, it should not exceed 9h... </p>",
      "rawMarkdown": "try to run prediction for 200000 images by replicating kaggle configuration 1 gpu and 2 cpu with 15gb ram. See how long it takes, it should not exceed 9h...",
      "votes": null
    },
    {
      "id": "703063",
      "postDate": "12/25/2019 14:35:35",
      "content": "<p>Sorry but I don't get it, the 4 test files we have have 3 rows each. How does that work, when you commit the kernel it runs on other files that we don't have access to? I'm asking because I check other public kernels for inference and in all of them it seems like their output file is always the same file with only the 12 tests.</p>",
      "rawMarkdown": "Sorry but I don't get it, the 4 test files we have have 3 rows each. How does that work, when you commit the kernel it runs on other files that we don't have access to? I'm asking because I check other public kernels for inference and in all of them it seems like their output file is always the same file with only the 12 tests.",
      "votes": null
    },
    {
      "id": "703085",
      "postDate": "12/25/2019 15:14:27",
      "content": "<p>when you submit, the test file get replaced by bigger test data =) </p>",
      "rawMarkdown": "when you submit, the test file get replaced by bigger test data =)",
      "votes": null
    },
    {
      "id": "703187",
      "postDate": "12/25/2019 18:48:57",
      "content": "<p>Facing the same issue. I am using CPU for submission only, so GPU limit is not a problem in my case. Not able to figure out what the problem is.</p>",
      "rawMarkdown": "Facing the same issue. I am using CPU for submission only, so GPU limit is not a problem in my case. Not able to figure out what the problem is.",
      "votes": null
    },
    {
      "id": "703197",
      "postDate": "12/25/2019 19:14:08",
      "content": "<p>Did you try to make an estimate of the inference time, following the suggestion from DrHB?</p>\n\n<blockquote>\n  <p>try to run prediction for 200000 images by replicating kaggle configuration 1 gpu and 2 cpu with 15gb ram. See how long it takes, it should not exceed 9h…</p>\n</blockquote>",
      "rawMarkdown": "Did you try to make an estimate of the inference time, following the suggestion from DrHB?\n&gt; try to run prediction for 200000 images by replicating kaggle configuration 1 gpu and 2 cpu with 15gb ram. See how long it takes, it should not exceed 9h…",
      "votes": null
    },
    {
      "id": "703233",
      "postDate": "12/25/2019 20:46:07",
      "content": "<p>But still it doesn't make so much sense... I mean I use google colab and to train the model for 20 epochs it takes about 45 minutes. And just predicting in Kaggle could take more than 2 hours? Is the kaggle GPU that much worse?</p>",
      "rawMarkdown": "But still it doesn't make so much sense... I mean I use google colab and to train the model for 20 epochs it takes about 45 minutes. And just predicting in Kaggle could take more than 2 hours? Is the kaggle GPU that much worse?",
      "votes": null
    },
    {
      "id": "703309",
      "postDate": "12/26/2019 01:40:00",
      "content": "<p>Hmm, that's odd. Kaggle servers use P100s (according to their documentation) which shouldn't be worse then colab (I believe they use K80). Would it be possible for you to share the training or inference code?</p>",
      "rawMarkdown": "Hmm, that's odd. Kaggle servers use P100s (according to their documentation) which shouldn't be worse then colab (I believe they use K80). Would it be possible for you to share the training or inference code?",
      "votes": null
    },
    {
      "id": "703354",
      "postDate": "12/26/2019 03:24:08",
      "content": "<p>`components = ['consonant_diacritic', 'grapheme_root', 'vowel_diacritic']\ntarget=[] # model predictions placeholder\nrow_id=[] # row_id place holder\nfor i in range(4):\n    df_test_img = pd.read_parquet('/kaggle/input/bengaliai-cv19/test_image_data_{}.parquet'.format(i)) \n    df_test_img.set_index('image_id', inplace=True)</p>\n\n<pre><code>X_test = resize(df_test_img, need_progress_bar=False)/255\n#X_test = X_test.values.reshape(-1, reHEIGHT,reWIDTH)\n#X_test = np.repeat(X_test,3)\nX_test = np.repeat(X_test.values,3).reshape(-1, reHEIGHT,reWIDTH, N_CHANNELS)\n\nfor pred in preds_dict:\n    preds_dict[pred]=np.argmax(model_dict[pred].predict(X_test), axis=1)\n\nfor k,id in enumerate(df_test_img.index.values):  \n    for i,comp in enumerate(components):\n        id_sample=id+'_'+comp\n        row_id.append(id_sample)\n        target.append(preds_dict[comp][k])\ndel df_test_img\ndel X_test\ngc.collect()\n</code></pre>\n\n<p>df_sample = pd.DataFrame(\n    {\n        'row_id': row_id,\n        'target':target\n    },\n    columns = ['row_id','target'] \n)\ndf_sample.to_csv('submission.csv',index=False)`</p>\n\n<p>Because I use the pre training model, which requires three channels, np.repeat (data, 3), the code is as follows:\n@TahsinReasat:Thanks first!\nHowever, the following error occurred: 《Notebook Exceeded Allowed Compute》</p>",
      "rawMarkdown": "`components = ['consonant_diacritic', 'grapheme_root', 'vowel_diacritic']\ntarget=[] # model predictions placeholder\nrow_id=[] # row_id place holder\nfor i in range(4):\n    df_test_img = pd.read_parquet('/kaggle/input/bengaliai-cv19/test_image_data_{}.parquet'.format(i)) \n    df_test_img.set_index('image_id', inplace=True)\n\n    X_test = resize(df_test_img, need_progress_bar=False)/255\n    #X_test = X_test.values.reshape(-1, reHEIGHT,reWIDTH)\n    #X_test = np.repeat(X_test,3)\n    X_test = np.repeat(X_test.values,3).reshape(-1, reHEIGHT,reWIDTH, N_CHANNELS)\n\n    for pred in preds_dict:\n        preds_dict[pred]=np.argmax(model_dict[pred].predict(X_test), axis=1)\n\n    for k,id in enumerate(df_test_img.index.values):  \n        for i,comp in enumerate(components):\n            id_sample=id+'_'+comp\n            row_id.append(id_sample)\n            target.append(preds_dict[comp][k])\n    del df_test_img\n    del X_test\n    gc.collect()\n\ndf_sample = pd.DataFrame(\n    {\n        'row_id': row_id,\n        'target':target\n    },\n    columns = ['row_id','target'] \n)\ndf_sample.to_csv('submission.csv',index=False)`\n\n\nBecause I use the pre training model, which requires three channels, np.repeat (data, 3), the code is as follows:\n@TahsinReasat:Thanks first!\nHowever, the following error occurred: 《Notebook Exceeded Allowed Compute》",
      "votes": null
    },
    {
      "id": "704386",
      "postDate": "12/27/2019 12:12:11",
      "content": "<p>please check the batchsize and number of workers you are using while doing prediction on the test set .  In the dummy test set batch_size=1 and no num_worker gives processing time as 5 minutes . However , the submission runs for 2+ hours and fails . It took me a while to find that bug . With batch_size 128 and num_worker 2 also the prediction in dummy test takes 5 minutes . But the submission runs for 1 hour and completes successfully . So be aware of that .</p>",
      "rawMarkdown": "please check the batchsize and number of workers you are using while doing prediction on the test set .  In the dummy test set batch_size=1 and no num_worker gives processing time as 5 minutes . However , the submission runs for 2+ hours and fails . It took me a while to find that bug . With batch_size 128 and num_worker 2 also the prediction in dummy test takes 5 minutes . But the submission runs for 1 hour and completes successfully . So be aware of that .",
      "votes": null
    },
    {
      "id": "752657",
      "postDate": "02/21/2020 09:14:45",
      "content": "<p>facing same problem.</p>\n\n<p>use some tricks and batch size of 32, it took less than 25min for inference on train set. so it's not the time problem.\nand I also use test files to generate row_id instead of sample submission file, failed either.\neven using test.csv, failed.\nwhat's the problem....</p>",
      "rawMarkdown": "facing same problem.\n\nuse some tricks and batch size of 32, it took less than 25min for inference on train set. so it's not the time problem.\nand I also use test files to generate row_id instead of sample submission file, failed either.\neven using test.csv, failed.\nwhat's the problem....",
      "votes": null
    },
    {
      "id": "768753",
      "postDate": "03/11/2020 06:38:50",
      "content": "<p>I ran the prediction for 200000 images on google colab and it ends up using less than 10GB of RAM and when i run the same thing on kaggle the ram overflows causing a submission error.\nNote: I'm not training, only running inference</p>\n\n<p>Any help would be appreciated!</p>",
      "rawMarkdown": "I ran the prediction for 200000 images on google colab and it ends up using less than 10GB of RAM and when i run the same thing on kaggle the ram overflows causing a submission error.\nNote: I'm not training, only running inference\n\nAny help would be appreciated!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 702654,
      "author_name": "reasat",
      "author_url": "",
      "post_date": "12/25/2019 01:05:48",
      "content": "<p>Most probably you are running past the allotted GPU time. There is a two hour limit on GPU use during inference.</p>",
      "votes": null,
      "replies": [
        {
          "id": 702700,
          "author_name": "zhangeng",
          "author_url": "",
          "post_date": "12/25/2019 03:05:18",
          "content": "<blockquote>\n  <p><strong>TahsinReasat wrote:</strong></p>\n  \n  <p>Most probably you are running past the allotted GPU time. There is a two hour limit on GPU use during inference.</p>\n</blockquote>\n\n<p>I have the same problem. I only use CPU, and my kernel only runs for 300-400s, and then it ends normally</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 703233,
          "author_name": "tzoukritzou",
          "author_url": "",
          "post_date": "12/25/2019 20:46:07",
          "content": "<p>But still it doesn't make so much sense... I mean I use google colab and to train the model for 20 epochs it takes about 45 minutes. And just predicting in Kaggle could take more than 2 hours? Is the kaggle GPU that much worse?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 703309,
          "author_name": "reasat",
          "author_url": "",
          "post_date": "12/26/2019 01:40:00",
          "content": "<p>Hmm, that's odd. Kaggle servers use P100s (according to their documentation) which shouldn't be worse then colab (I believe they use K80). Would it be possible for you to share the training or inference code?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 703354,
          "author_name": "zhangeng",
          "author_url": "",
          "post_date": "12/26/2019 03:24:08",
          "content": "<p>`components = ['consonant_diacritic', 'grapheme_root', 'vowel_diacritic']\ntarget=[] # model predictions placeholder\nrow_id=[] # row_id place holder\nfor i in range(4):\n    df_test_img = pd.read_parquet('/kaggle/input/bengaliai-cv19/test_image_data_{}.parquet'.format(i)) \n    df_test_img.set_index('image_id', inplace=True)</p>\n\n<pre><code>X_test = resize(df_test_img, need_progress_bar=False)/255\n#X_test = X_test.values.reshape(-1, reHEIGHT,reWIDTH)\n#X_test = np.repeat(X_test,3)\nX_test = np.repeat(X_test.values,3).reshape(-1, reHEIGHT,reWIDTH, N_CHANNELS)\n\nfor pred in preds_dict:\n    preds_dict[pred]=np.argmax(model_dict[pred].predict(X_test), axis=1)\n\nfor k,id in enumerate(df_test_img.index.values):  \n    for i,comp in enumerate(components):\n        id_sample=id+'_'+comp\n        row_id.append(id_sample)\n        target.append(preds_dict[comp][k])\ndel df_test_img\ndel X_test\ngc.collect()\n</code></pre>\n\n<p>df_sample = pd.DataFrame(\n    {\n        'row_id': row_id,\n        'target':target\n    },\n    columns = ['row_id','target'] \n)\ndf_sample.to_csv('submission.csv',index=False)`</p>\n\n<p>Because I use the pre training model, which requires three channels, np.repeat (data, 3), the code is as follows:\n@TahsinReasat:Thanks first!\nHowever, the following error occurred: 《Notebook Exceeded Allowed Compute》</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 704386,
          "author_name": "phoenix9032",
          "author_url": "",
          "post_date": "12/27/2019 12:12:11",
          "content": "<p>please check the batchsize and number of workers you are using while doing prediction on the test set .  In the dummy test set batch_size=1 and no num_worker gives processing time as 5 minutes . However , the submission runs for 2+ hours and fails . It took me a while to find that bug . With batch_size 128 and num_worker 2 also the prediction in dummy test takes 5 minutes . But the submission runs for 1 hour and completes successfully . So be aware of that .</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 702716,
      "author_name": "drhabib",
      "author_url": "",
      "post_date": "12/25/2019 03:53:52",
      "content": "<p>try to run prediction for 200000 images by replicating kaggle configuration 1 gpu and 2 cpu with 15gb ram. See how long it takes, it should not exceed 9h... </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 703063,
      "author_name": "tzoukritzou",
      "author_url": "",
      "post_date": "12/25/2019 14:35:35",
      "content": "<p>Sorry but I don't get it, the 4 test files we have have 3 rows each. How does that work, when you commit the kernel it runs on other files that we don't have access to? I'm asking because I check other public kernels for inference and in all of them it seems like their output file is always the same file with only the 12 tests.</p>",
      "votes": null,
      "replies": [
        {
          "id": 703085,
          "author_name": "drhabib",
          "author_url": "",
          "post_date": "12/25/2019 15:14:27",
          "content": "<p>when you submit, the test file get replaced by bigger test data =) </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 703187,
      "author_name": "bitthal",
      "author_url": "",
      "post_date": "12/25/2019 18:48:57",
      "content": "<p>Facing the same issue. I am using CPU for submission only, so GPU limit is not a problem in my case. Not able to figure out what the problem is.</p>",
      "votes": null,
      "replies": [
        {
          "id": 703197,
          "author_name": "reasat",
          "author_url": "",
          "post_date": "12/25/2019 19:14:08",
          "content": "<p>Did you try to make an estimate of the inference time, following the suggestion from DrHB?</p>\n\n<blockquote>\n  <p>try to run prediction for 200000 images by replicating kaggle configuration 1 gpu and 2 cpu with 15gb ram. See how long it takes, it should not exceed 9h…</p>\n</blockquote>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 752657,
      "author_name": "tiandaye",
      "author_url": "",
      "post_date": "02/21/2020 09:14:45",
      "content": "<p>facing same problem.</p>\n\n<p>use some tricks and batch size of 32, it took less than 25min for inference on train set. so it's not the time problem.\nand I also use test files to generate row_id instead of sample submission file, failed either.\neven using test.csv, failed.\nwhat's the problem....</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 768753,
      "author_name": "fsociety28",
      "author_url": "",
      "post_date": "03/11/2020 06:38:50",
      "content": "<p>I ran the prediction for 200000 images on google colab and it ends up using less than 10GB of RAM and when i run the same thing on kaggle the ram overflows causing a submission error.\nNote: I'm not training, only running inference</p>\n\n<p>Any help would be appreciated!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "702478": "This is my first kernel competition and I'm having difficulties scoring my submission file. I train the model locally and then I upload my model parameters to my kaggle notebook and I make my predictions there. I save the \"submission.csv\" file as instructed, I commit the kernel, I click the button to submit and then it takes too long saying \"scoring your submissions\" and in the end it throws an error. Does anyone have any idea why this happens? Thanks!",
    "702654": "Most probably you are running past the allotted GPU time. There is a two hour limit on GPU use during inference.",
    "702700": "&gt; **TahsinReasat wrote:**\n&gt; \n&gt; Most probably you are running past the allotted GPU time. There is a two hour limit on GPU use during inference.\n\nI have the same problem. I only use CPU, and my kernel only runs for 300-400s, and then it ends normally",
    "702716": "try to run prediction for 200000 images by replicating kaggle configuration 1 gpu and 2 cpu with 15gb ram. See how long it takes, it should not exceed 9h...",
    "703063": "Sorry but I don't get it, the 4 test files we have have 3 rows each. How does that work, when you commit the kernel it runs on other files that we don't have access to? I'm asking because I check other public kernels for inference and in all of them it seems like their output file is always the same file with only the 12 tests.",
    "703085": "when you submit, the test file get replaced by bigger test data =)",
    "703187": "Facing the same issue. I am using CPU for submission only, so GPU limit is not a problem in my case. Not able to figure out what the problem is.",
    "703197": "Did you try to make an estimate of the inference time, following the suggestion from DrHB?\n&gt; try to run prediction for 200000 images by replicating kaggle configuration 1 gpu and 2 cpu with 15gb ram. See how long it takes, it should not exceed 9h…",
    "703233": "But still it doesn't make so much sense... I mean I use google colab and to train the model for 20 epochs it takes about 45 minutes. And just predicting in Kaggle could take more than 2 hours? Is the kaggle GPU that much worse?",
    "703309": "Hmm, that's odd. Kaggle servers use P100s (according to their documentation) which shouldn't be worse then colab (I believe they use K80). Would it be possible for you to share the training or inference code?",
    "703354": "`components = ['consonant_diacritic', 'grapheme_root', 'vowel_diacritic']\ntarget=[] # model predictions placeholder\nrow_id=[] # row_id place holder\nfor i in range(4):\n    df_test_img = pd.read_parquet('/kaggle/input/bengaliai-cv19/test_image_data_{}.parquet'.format(i)) \n    df_test_img.set_index('image_id', inplace=True)\n\n    X_test = resize(df_test_img, need_progress_bar=False)/255\n    #X_test = X_test.values.reshape(-1, reHEIGHT,reWIDTH)\n    #X_test = np.repeat(X_test,3)\n    X_test = np.repeat(X_test.values,3).reshape(-1, reHEIGHT,reWIDTH, N_CHANNELS)\n\n    for pred in preds_dict:\n        preds_dict[pred]=np.argmax(model_dict[pred].predict(X_test), axis=1)\n\n    for k,id in enumerate(df_test_img.index.values):  \n        for i,comp in enumerate(components):\n            id_sample=id+'_'+comp\n            row_id.append(id_sample)\n            target.append(preds_dict[comp][k])\n    del df_test_img\n    del X_test\n    gc.collect()\n\ndf_sample = pd.DataFrame(\n    {\n        'row_id': row_id,\n        'target':target\n    },\n    columns = ['row_id','target'] \n)\ndf_sample.to_csv('submission.csv',index=False)`\n\n\nBecause I use the pre training model, which requires three channels, np.repeat (data, 3), the code is as follows:\n@TahsinReasat:Thanks first!\nHowever, the following error occurred: 《Notebook Exceeded Allowed Compute》",
    "704386": "please check the batchsize and number of workers you are using while doing prediction on the test set .  In the dummy test set batch_size=1 and no num_worker gives processing time as 5 minutes . However , the submission runs for 2+ hours and fails . It took me a while to find that bug . With batch_size 128 and num_worker 2 also the prediction in dummy test takes 5 minutes . But the submission runs for 1 hour and completes successfully . So be aware of that .",
    "752657": "facing same problem.\n\nuse some tricks and batch size of 32, it took less than 25min for inference on train set. so it's not the time problem.\nand I also use test files to generate row_id instead of sample submission file, failed either.\neven using test.csv, failed.\nwhat's the problem....",
    "768753": "I ran the prediction for 200000 images on google colab and it ends up using less than 10GB of RAM and when i run the same thing on kaggle the ram overflows causing a submission error.\nNote: I'm not training, only running inference\n\nAny help would be appreciated!"
  },
  "source": "meta"
}