{
  "id": 129045,
  "title": "How long does it take you to load test_image_data_.parquet?",
  "url": "/competitions/bengaliai-cv19/discussion/129045",
  "author_name": "PY",
  "post_date": "2020-02-05T03:36:59.392000",
  "votes": 0,
  "comment_count": 16,
  "views": 0,
  "content": "<p>When I'm running the submission script, I found that it took around 53 second to load each <code>test_image_data_[0-3].parquet</code>. Since the size of each file is around 1.2 MB, I found it weird that each file took more than 50 second to load. </p>\n\n<p>Not sure if it's related, but when I use a single model that's SE-ResNeXt 101 or larger, I got time out error for submission, maybe it took too long to load the public test data?</p>\n\n<p>Did you guys encounter similar situation? Thanks.</p>",
  "messages": [
    {
      "id": 737651,
      "postDate": "2020-02-05T16:09:50.593Z",
      "content": "<p>I wonder why dataset is not offered by feather or image format in this competition.</p>",
      "rawMarkdown": "I wonder why dataset is not offered by feather or image format in this competition.",
      "votes": 1
    },
    {
      "id": 738045,
      "postDate": "2020-02-06T04:55:45.520Z",
      "content": "<p>Thank you all for the help! I'll investigate the scripts provided by <a href=\"/pestipeti\">@pestipeti</a> and all to learn about the optimization for inference.</p>",
      "rawMarkdown": "Thank you all for the help! I'll investigate the scripts provided by @pestipeti and all to learn about the optimization for inference."
    },
    {
      "id": 737791,
      "postDate": "2020-02-05T19:19:00.063Z",
      "content": "<p>You could use fast-paraquet (<a href=\"https://www.kaggle.com/vladislavleketush/fast-parquet-loading-example\">https://www.kaggle.com/vladislavleketush/fast-parquet-loading-example</a>) it takes awhile to pip install but changes load time from 50-60s to 26-32s.</p>",
      "rawMarkdown": "You could use fast-paraquet (https://www.kaggle.com/vladislavleketush/fast-parquet-loading-example) it takes awhile to pip install but changes load time from 50-60s to 26-32s.",
      "replies": [
        {
          "id": 739335,
          "postDate": "2020-02-07T17:53:28.707Z",
          "content": "<p>Thank you, very helpful!</p>",
          "rawMarkdown": "Thank you, very helpful!"
        }
      ]
    },
    {
      "id": 737301,
      "postDate": "2020-02-05T06:47:12.267Z",
      "content": "<p>in case it helps, here is the part in which I read from the test image parquet files in the submission script:</p>\n\n<p>```\nfor i in range(0,4):</p>\n\n<pre><code>data_full = pd.read_parquet(os.path.join(DATA_PATH, 'test_image_data_{}.parquet'.format(i)),engine = 'pyarrow')\n\ntrain_image = GraphemeDataset(data_full, None, 128, None, 0.0)\ntrain_loader = torch.utils.data.DataLoader(train_image,batch_size=1,shuffle=False)\n\nfor idx, inputs in tqdm(enumerate(train_loader),total=len(train_loader)):\n    inputs = inputs.to(device)\n\n    with torch.no_grad():\n        outputs1, outputs2, outputs3 = model_candidate(inputs.unsqueeze(1).float())\n\n        predictions.append(outputs3.argmax(1).cpu().detach().numpy())\n        predictions.append(outputs2.argmax(1).cpu().detach().numpy())\n        predictions.append(outputs1.argmax(1).cpu().detach().numpy())\n\ndel data_full, train_image, train_loader\ngc.collect()\n</code></pre>\n\n<p>```</p>",
      "rawMarkdown": "in case it helps, here is the part in which I read from the test image parquet files in the submission script:\n\n```\nfor i in range(0,4):\n\n    data_full = pd.read_parquet(os.path.join(DATA_PATH, 'test_image_data_{}.parquet'.format(i)),engine = 'pyarrow')\n\n    train_image = GraphemeDataset(data_full, None, 128, None, 0.0)\n    train_loader = torch.utils.data.DataLoader(train_image,batch_size=1,shuffle=False)\n    \n    for idx, inputs in tqdm(enumerate(train_loader),total=len(train_loader)):\n        inputs = inputs.to(device)\n\n        with torch.no_grad():\n            outputs1, outputs2, outputs3 = model_candidate(inputs.unsqueeze(1).float())\n\n            predictions.append(outputs3.argmax(1).cpu().detach().numpy())\n            predictions.append(outputs2.argmax(1).cpu().detach().numpy())\n            predictions.append(outputs1.argmax(1).cpu().detach().numpy())\n\n    del data_full, train_image, train_loader\n    gc.collect()\n```"
    },
    {
      "id": 737284,
      "postDate": "2020-02-05T06:15:55.833Z",
      "content": "<p>but for submission, the format of the test data are in <code>.parquet</code>, right? So while it's okay to use feather in training, we can only load the <code>.parquet</code> file in the submission script. Or am I missing something?</p>",
      "rawMarkdown": "but for submission, the format of the test data are in `.parquet`, right? So while it's okay to use feather in training, we can only load the `.parquet` file in the submission script. Or am I missing something?",
      "replies": [
        {
          "id": 737307,
          "postDate": "2020-02-05T06:57:19.127Z",
          "content": "<p>It's correct you can use it for train and not for test.</p>\n\n<p>But the sample test files are not heavy so train on feature and test on paraquet</p>",
          "rawMarkdown": "It's correct you can use it for train and not for test.\n\nBut the sample test files are not heavy so train on feature and test on paraquet"
        },
        {
          "id": 737623,
          "postDate": "2020-02-05T15:40:57.643Z",
          "content": "<p>That's my point, when I'm loading the sample test files which is only 1.2MB, it took more than 50 sec. So I can imagine when the real test files were being loaded during submission, it must have taken a huge amount of time, which reduces the inference time (since each submission has a total time limit).</p>\n\n<p>Not sure if I'm doing anything wrong?</p>",
          "rawMarkdown": "That's my point, when I'm loading the sample test files which is only 1.2MB, it took more than 50 sec. So I can imagine when the real test files were being loaded during submission, it must have taken a huge amount of time, which reduces the inference time (since each submission has a total time limit).\n\nNot sure if I'm doing anything wrong?"
        },
        {
          "id": 737627,
          "postDate": "2020-02-05T15:44:48.400Z",
          "content": "<p>I used pretrained weights during submission and it took around 21 minutes for whole processing including loading of parquet test file, predicting and generating output file.\nBut everytime I submits this time varies from 20 25 minutes to an hour sometime.</p>\n\n<p>Hope this clarifies that it doesn't take much time to load heavy parquet files also.</p>",
          "rawMarkdown": "I used pretrained weights during submission and it took around 21 minutes for whole processing including loading of parquet test file, predicting and generating output file.\nBut everytime I submits this time varies from 20 25 minutes to an hour sometime.\n\nHope this clarifies that it doesn't take much time to load heavy parquet files also.\n\n"
        },
        {
          "id": 737695,
          "postDate": "2020-02-05T16:53:49.543Z",
          "content": "<p>Thank you Amit! I'm using pretrained weights as well, and so far all my GPU submission failed with Notebook Timeout Error, while it took around 8 hour for my CPU submission to run (with single model, ResNet50). </p>\n\n<p>If possible, do you mind sharing the section of your submission code in which you read the parquet files? Not sure if there is something obvious that I did wrong. Thank you!</p>",
          "rawMarkdown": "Thank you Amit! I'm using pretrained weights as well, and so far all my GPU submission failed with Notebook Timeout Error, while it took around 8 hour for my CPU submission to run (with single model, ResNet50). \n\nIf possible, do you mind sharing the section of your submission code in which you read the parquet files? Not sure if there is something obvious that I did wrong. Thank you!"
        },
        {
          "id": 737730,
          "postDate": "2020-02-05T17:40:32.920Z",
          "content": "<p>It takes 50 seconds because of the number of columns, not the rows. The private test has ~the same number of samples in the parquet files as we have in the train. You can check the reading time with the train files</p>\n\n<p><a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/126054\">Here is my topic about optimization</a> (Implementation linked)</p>\n\n<p>If you optimize your submission code, it is possible to submit a 5-fold ensembe (resnet34) within 20 minutes.</p>",
          "rawMarkdown": "It takes 50 seconds because of the number of columns, not the rows. The private test has ~the same number of samples in the parquet files as we have in the train. You can check the reading time with the train files\n\n[Here is my topic about optimization](https://www.kaggle.com/c/bengaliai-cv19/discussion/126054) (Implementation linked)\n\nIf you optimize your submission code, it is possible to submit a 5-fold ensembe (resnet34) within 20 minutes.\n"
        },
        {
          "id": 738108,
          "postDate": "2020-02-06T06:22:11.040Z",
          "content": "<p><a href=\"/pestipeti\">@pestipeti</a> Did you use GPU kernel for 5-fold ensemble ?</p>",
          "rawMarkdown": "@pestipeti Did you use GPU kernel for 5-fold ensemble ?"
        },
        {
          "id": 738154,
          "postDate": "2020-02-06T08:04:25.497Z",
          "content": "<p><a href=\"/toshik\">@toshik</a> Yes, I used GPU</p>",
          "rawMarkdown": "@toshik Yes, I used GPU",
          "votes": 1
        },
        {
          "id": 738730,
          "postDate": "2020-02-06T23:30:29.670Z",
          "content": "<p><a href=\"/pestipeti\">@pestipeti</a> Thank you ! I'll try it.</p>",
          "rawMarkdown": "@pestipeti Thank you ! I'll try it."
        }
      ]
    },
    {
      "id": 737235,
      "postDate": "2020-02-05T04:35:30.670Z",
      "content": "<p>Use feature file to reduce loading time.</p>",
      "rawMarkdown": "Use feature file to reduce loading time."
    },
    {
      "id": 737202,
      "postDate": "2020-02-05T03:36:59.393Z",
      "content": "<p>When I'm running the submission script, I found that it took around 53 second to load each <code>test_image_data_[0-3].parquet</code>. Since the size of each file is around 1.2 MB, I found it weird that each file took more than 50 second to load. </p>\n\n<p>Not sure if it's related, but when I use a single model that's SE-ResNeXt 101 or larger, I got time out error for submission, maybe it took too long to load the public test data?</p>\n\n<p>Did you guys encounter similar situation? Thanks.</p>",
      "rawMarkdown": "When I'm running the submission script, I found that it took around 53 second to load each `test_image_data_[0-3].parquet`. Since the size of each file is around 1.2 MB, I found it weird that each file took more than 50 second to load. \n\nNot sure if it's related, but when I use a single model that's SE-ResNeXt 101 or larger, I got time out error for submission, maybe it took too long to load the public test data?\n\nDid you guys encounter similar situation? Thanks."
    },
    {
      "id": 738543,
      "postDate": "2020-02-06T17:07:35.347Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 737651,
      "author_name": "toshi_k",
      "author_url": "",
      "post_date": "2020-02-05T16:09:50.593000",
      "content": "<p>I wonder why dataset is not offered by feather or image format in this competition.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 738045,
      "author_name": "PY",
      "author_url": "",
      "post_date": "2020-02-06T04:55:45.520000",
      "content": "<p>Thank you all for the help! I'll investigate the scripts provided by <a href=\"/pestipeti\">@pestipeti</a> and all to learn about the optimization for inference.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 737791,
      "author_name": "GreatGameDota",
      "author_url": "",
      "post_date": "2020-02-05T19:19:00.063000",
      "content": "<p>You could use fast-paraquet (<a href=\"https://www.kaggle.com/vladislavleketush/fast-parquet-loading-example\">https://www.kaggle.com/vladislavleketush/fast-parquet-loading-example</a>) it takes awhile to pip install but changes load time from 50-60s to 26-32s.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 739335,
          "author_name": "PY",
          "author_url": "",
          "post_date": "2020-02-07T17:53:28.707000",
          "content": "<p>Thank you, very helpful!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 737301,
      "author_name": "PY",
      "author_url": "",
      "post_date": "2020-02-05T06:47:12.267000",
      "content": "<p>in case it helps, here is the part in which I read from the test image parquet files in the submission script:</p>\n\n<p>```\nfor i in range(0,4):</p>\n\n<pre><code>data_full = pd.read_parquet(os.path.join(DATA_PATH, 'test_image_data_{}.parquet'.format(i)),engine = 'pyarrow')\n\ntrain_image = GraphemeDataset(data_full, None, 128, None, 0.0)\ntrain_loader = torch.utils.data.DataLoader(train_image,batch_size=1,shuffle=False)\n\nfor idx, inputs in tqdm(enumerate(train_loader),total=len(train_loader)):\n    inputs = inputs.to(device)\n\n    with torch.no_grad():\n        outputs1, outputs2, outputs3 = model_candidate(inputs.unsqueeze(1).float())\n\n        predictions.append(outputs3.argmax(1).cpu().detach().numpy())\n        predictions.append(outputs2.argmax(1).cpu().detach().numpy())\n        predictions.append(outputs1.argmax(1).cpu().detach().numpy())\n\ndel data_full, train_image, train_loader\ngc.collect()\n</code></pre>\n\n<p>```</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 737284,
      "author_name": "PY",
      "author_url": "",
      "post_date": "2020-02-05T06:15:55.833000",
      "content": "<p>but for submission, the format of the test data are in <code>.parquet</code>, right? So while it's okay to use feather in training, we can only load the <code>.parquet</code> file in the submission script. Or am I missing something?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 737307,
          "author_name": "Amit",
          "author_url": "",
          "post_date": "2020-02-05T06:57:19.127000",
          "content": "<p>It's correct you can use it for train and not for test.</p>\n\n<p>But the sample test files are not heavy so train on feature and test on paraquet</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 737623,
          "author_name": "PY",
          "author_url": "",
          "post_date": "2020-02-05T15:40:57.643000",
          "content": "<p>That's my point, when I'm loading the sample test files which is only 1.2MB, it took more than 50 sec. So I can imagine when the real test files were being loaded during submission, it must have taken a huge amount of time, which reduces the inference time (since each submission has a total time limit).</p>\n\n<p>Not sure if I'm doing anything wrong?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 737627,
          "author_name": "Amit",
          "author_url": "",
          "post_date": "2020-02-05T15:44:48.400000",
          "content": "<p>I used pretrained weights during submission and it took around 21 minutes for whole processing including loading of parquet test file, predicting and generating output file.\nBut everytime I submits this time varies from 20 25 minutes to an hour sometime.</p>\n\n<p>Hope this clarifies that it doesn't take much time to load heavy parquet files also.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 737695,
          "author_name": "PY",
          "author_url": "",
          "post_date": "2020-02-05T16:53:49.543000",
          "content": "<p>Thank you Amit! I'm using pretrained weights as well, and so far all my GPU submission failed with Notebook Timeout Error, while it took around 8 hour for my CPU submission to run (with single model, ResNet50). </p>\n\n<p>If possible, do you mind sharing the section of your submission code in which you read the parquet files? Not sure if there is something obvious that I did wrong. Thank you!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 737730,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2020-02-05T17:40:32.920000",
          "content": "<p>It takes 50 seconds because of the number of columns, not the rows. The private test has ~the same number of samples in the parquet files as we have in the train. You can check the reading time with the train files</p>\n\n<p><a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/126054\">Here is my topic about optimization</a> (Implementation linked)</p>\n\n<p>If you optimize your submission code, it is possible to submit a 5-fold ensembe (resnet34) within 20 minutes.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 738108,
          "author_name": "toshi_k",
          "author_url": "",
          "post_date": "2020-02-06T06:22:11.040000",
          "content": "<p><a href=\"/pestipeti\">@pestipeti</a> Did you use GPU kernel for 5-fold ensemble ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 738154,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2020-02-06T08:04:25.497000",
          "content": "<p><a href=\"/toshik\">@toshik</a> Yes, I used GPU</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 738730,
          "author_name": "toshi_k",
          "author_url": "",
          "post_date": "2020-02-06T23:30:29.670000",
          "content": "<p><a href=\"/pestipeti\">@pestipeti</a> Thank you ! I'll try it.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 737235,
      "author_name": "Amit",
      "author_url": "",
      "post_date": "2020-02-05T04:35:30.670000",
      "content": "<p>Use feature file to reduce loading time.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 738543,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-06T17:07:35.347000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "737651": "I wonder why dataset is not offered by feather or image format in this competition.",
    "738045": "Thank you all for the help! I'll investigate the scripts provided by @pestipeti and all to learn about the optimization for inference.",
    "737791": "You could use fast-paraquet (https://www.kaggle.com/vladislavleketush/fast-parquet-loading-example) it takes awhile to pip install but changes load time from 50-60s to 26-32s.",
    "737301": "in case it helps, here is the part in which I read from the test image parquet files in the submission script:\n\n```\nfor i in range(0,4):\n\n    data_full = pd.read_parquet(os.path.join(DATA_PATH, 'test_image_data_{}.parquet'.format(i)),engine = 'pyarrow')\n\n    train_image = GraphemeDataset(data_full, None, 128, None, 0.0)\n    train_loader = torch.utils.data.DataLoader(train_image,batch_size=1,shuffle=False)\n    \n    for idx, inputs in tqdm(enumerate(train_loader),total=len(train_loader)):\n        inputs = inputs.to(device)\n\n        with torch.no_grad():\n            outputs1, outputs2, outputs3 = model_candidate(inputs.unsqueeze(1).float())\n\n            predictions.append(outputs3.argmax(1).cpu().detach().numpy())\n            predictions.append(outputs2.argmax(1).cpu().detach().numpy())\n            predictions.append(outputs1.argmax(1).cpu().detach().numpy())\n\n    del data_full, train_image, train_loader\n    gc.collect()\n```",
    "737284": "but for submission, the format of the test data are in `.parquet`, right? So while it's okay to use feather in training, we can only load the `.parquet` file in the submission script. Or am I missing something?",
    "737235": "Use feature file to reduce loading time.",
    "737202": "When I'm running the submission script, I found that it took around 53 second to load each `test_image_data_[0-3].parquet`. Since the size of each file is around 1.2 MB, I found it weird that each file took more than 50 second to load. \n\nNot sure if it's related, but when I use a single model that's SE-ResNeXt 101 or larger, I got time out error for submission, maybe it took too long to load the public test data?\n\nDid you guys encounter similar situation? Thanks.",
    "738543": ""
  }
}