{
  "id": 134618,
  "title": "Finally solved submission errors",
  "url": "/competitions/bengaliai-cv19/discussion/134618",
  "author_name": "Shinsei66",
  "post_date": "2020-03-09T09:34:11.066000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>🙌 🙌 \nI was encountering submission errors up to 7 times and spent almost 30 hours of GPU time for these errors during this weekend😪 \nAlthough I've check all the 37 topics searched by \"submission errors\" that can be found in the forum, but nothing worked for me. The submissions still became \"submission scoring error\". </p>\n\n<p>Fortunately, I found <a href=\"https://www.kaggle.com/reasat/evaluating-model-performance-on-the-test-set\">this notebook</a> and modify the code a bit and the error is fixed eventually.\nHere I'd like to share the part of my code that worked, and what I've checked but not worked to you have the same situation.\nI guess the row_id mismatch caused the submission errors in my case. It needs to be obtained directly from the parquet files to avoid unexpected mismatch between hand-made row_id and the really row_id in the whole test set. (I've also tried obtain row_id from test.csv, but maybe something caused the unsynchronization of test.csv and test parquet, it didnt work. )</p>\n\n<p><strong>Refereces:</strong>\n- <a href=\"https://www.kaggle.com/reasat/evaluating-model-performance-on-the-test-set\">Evaluating model performance on the test set</a>\n- <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123632\">How to resolve \"Submission Error\" in my case</a>\n- <a href=\"https://www.kaggle.com/corochann/bengali-seresnext-prediction-with-pytorch\">Bengali.AI SEResNeXt prediction with pytorch</a></p>\n\n<p><strong>What I've checked and not worked</strong>\n- Notebook ran within 2 hours by replacing inference data to training dataset wi Batchsize=1 ==&gt; not a memory error, a time out error\n- Submission data type is matched with the sample_submission.csv ==&gt; not a scoring error\n- Output file is named \"submission.csv\" , and directly saved in the correct directly by <code>sub_df.to_csv('submission.csv', index=True)</code> ==&gt; not a CSV not found error</p>\n\n<p><strong>What worked for me</strong></p>\n\n<p>```\ndef prepare_image(datadir, featherdir, data_type='train',\n                  submission=True, indices=[0, 1, 2, 3]):\n    assert data_type in ['train', 'test']\n    if submission:\n        image_df_list = [pd.read_parquet(datadir / f'{data_type}_image_data_{i}.parquet')\n                         for i in indices]\n    else:\n        image_df_list = [pd.read_feather(featherdir / f'{data_type}_image_data_{i}.feather')\n                         for i in indices]</p>\n\n<pre><code>print('image_df_list', len(image_df_list))\nHEIGHT = 137\nWIDTH = 236\nimages = [df.iloc[:, 1:].values.reshape(-1, HEIGHT, WIDTH) for df in image_df_list]\n#del image_df_list\ngc.collect()\nimages = np.concatenate(images, axis=0)\nreturn images, image_df_list\n</code></pre>\n\n<p>def predict(model, dataloaders, phase, device):\n    model.eval()\n    output_list = []\n    label_list = []\n    with torch.no_grad():\n        if phase == 'test':\n            for i, inputs in enumerate(tqdm(dataloaders)):</p>\n\n<pre><code>            inputs = inputs.to(device)\n            outputs = model(inputs)\n            _, pred0 = torch.max(outputs[0], 1)\n            _, pred1 = torch.max(outputs[1], 1)\n            _, pred2 = torch.max(outputs[2], 1)\n            preds = (pred0.cpu().numpy(), pred1.cpu().numpy(), pred2.cpu().numpy())\n            output_list.append(preds)\n        return output_list\n    elif phase == 'val':\n        for i, (inputs, labels) in enumerate(tqdm(dataloaders)):\n\n            inputs = inputs.to(device)\n            outputs = model(inputs)\n            _, pred0 = torch.max(outputs[0], 1)\n            _, pred1 = torch.max(outputs[1], 1)\n            _, pred2 = torch.max(outputs[2], 1)\n            preds = (pred0, pred1, pred2)\n            output_list.append(preds)\n            label_list.append(labels.transpose(1,0))\n        return output_list, label_list\n</code></pre>\n\n<h1>--- Prediction ---</h1>\n\n<p>sts = time.time()\ndata_type = 'test'\ncomponents = ['consonant_diacritic', 'grapheme_root', 'vowel_diacritic']\ntarget=[] # model predictions placeholder\nrow_id=[] # row_id place holder</p>\n\n<p>for i in tqdm(range(4)):\n    indices = [i]\n   ## obtain row_id and images at the same time ##\n    test_images, df_test_img = prepare_image(\n        DATADIR, FEATHERDIR, data_type = data_type, submission=SUBMISSION, indices=indices)\n    n_dataset = len(test_images)\n    print(f'i={i}, n_dataset={n_dataset}')\n    test_dataset = BengaliAIDataset(\n    test_images, None,\n    transform=data_transforms[data_type])\n    print('test_dataset', len(test_dataset))\n    test_loader = DataLoader(test_dataset, batch_size=BATCH_SIZE, shuffle=False, num_workers=WORKER)</p>\n\n<pre><code>test_preds = predict(model_ft, test_loader, data_type,DEVICE)\np0 = np.array([test_pred[0] for test_pred in test_preds]) #grapheme\np1 = np.array([test_pred[1] for test_pred in test_preds]) #vowel\np2 = np.array([test_pred[2] for test_pred in test_preds]) #consonant\ntgt = [p2, p0, p1]\nfor idx, id in enumerate(df_test_img[0].image_id.values): # df_test_img[0].image_id.values has the test image_ids\n    #print(idx)\n    for i,comp in enumerate(components):\n        id_sample=id+'_'+comp\n        row_id.append(id_sample)\n        target.append(tgt[i].flatten()[idx]) \ndel test_images, df_test_img\ngc.collect()\nif DEBUG:\n    break\n</code></pre>\n\n<p>ed = time.time()\nprint('Predicted in {:.3f} min'.format((ed-sts)/60))\ndel p0, p1, p2</p>\n\n<h1>create a dataframe with the solutions</h1>\n\n<p>sub_df = pd.DataFrame(\n    {'row_id': row_id,\n    'target':target\n    },\n    columns =['row_id','target'] \n)</p>\n\n<p>sub_df.to_csv('submission.csv', index=False)\n```</p>",
  "messages": [
    {
      "id": 767191,
      "postDate": "2020-03-09T09:34:11.067Z",
      "content": "<p>🙌 🙌 \nI was encountering submission errors up to 7 times and spent almost 30 hours of GPU time for these errors during this weekend😪 \nAlthough I've check all the 37 topics searched by \"submission errors\" that can be found in the forum, but nothing worked for me. The submissions still became \"submission scoring error\". </p>\n\n<p>Fortunately, I found <a href=\"https://www.kaggle.com/reasat/evaluating-model-performance-on-the-test-set\">this notebook</a> and modify the code a bit and the error is fixed eventually.\nHere I'd like to share the part of my code that worked, and what I've checked but not worked to you have the same situation.\nI guess the row_id mismatch caused the submission errors in my case. It needs to be obtained directly from the parquet files to avoid unexpected mismatch between hand-made row_id and the really row_id in the whole test set. (I've also tried obtain row_id from test.csv, but maybe something caused the unsynchronization of test.csv and test parquet, it didnt work. )</p>\n\n<p><strong>Refereces:</strong>\n- <a href=\"https://www.kaggle.com/reasat/evaluating-model-performance-on-the-test-set\">Evaluating model performance on the test set</a>\n- <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123632\">How to resolve \"Submission Error\" in my case</a>\n- <a href=\"https://www.kaggle.com/corochann/bengali-seresnext-prediction-with-pytorch\">Bengali.AI SEResNeXt prediction with pytorch</a></p>\n\n<p><strong>What I've checked and not worked</strong>\n- Notebook ran within 2 hours by replacing inference data to training dataset wi Batchsize=1 ==&gt; not a memory error, a time out error\n- Submission data type is matched with the sample_submission.csv ==&gt; not a scoring error\n- Output file is named \"submission.csv\" , and directly saved in the correct directly by <code>sub_df.to_csv('submission.csv', index=True)</code> ==&gt; not a CSV not found error</p>\n\n<p><strong>What worked for me</strong></p>\n\n<p>```\ndef prepare_image(datadir, featherdir, data_type='train',\n                  submission=True, indices=[0, 1, 2, 3]):\n    assert data_type in ['train', 'test']\n    if submission:\n        image_df_list = [pd.read_parquet(datadir / f'{data_type}_image_data_{i}.parquet')\n                         for i in indices]\n    else:\n        image_df_list = [pd.read_feather(featherdir / f'{data_type}_image_data_{i}.feather')\n                         for i in indices]</p>\n\n<pre><code>print('image_df_list', len(image_df_list))\nHEIGHT = 137\nWIDTH = 236\nimages = [df.iloc[:, 1:].values.reshape(-1, HEIGHT, WIDTH) for df in image_df_list]\n#del image_df_list\ngc.collect()\nimages = np.concatenate(images, axis=0)\nreturn images, image_df_list\n</code></pre>\n\n<p>def predict(model, dataloaders, phase, device):\n    model.eval()\n    output_list = []\n    label_list = []\n    with torch.no_grad():\n        if phase == 'test':\n            for i, inputs in enumerate(tqdm(dataloaders)):</p>\n\n<pre><code>            inputs = inputs.to(device)\n            outputs = model(inputs)\n            _, pred0 = torch.max(outputs[0], 1)\n            _, pred1 = torch.max(outputs[1], 1)\n            _, pred2 = torch.max(outputs[2], 1)\n            preds = (pred0.cpu().numpy(), pred1.cpu().numpy(), pred2.cpu().numpy())\n            output_list.append(preds)\n        return output_list\n    elif phase == 'val':\n        for i, (inputs, labels) in enumerate(tqdm(dataloaders)):\n\n            inputs = inputs.to(device)\n            outputs = model(inputs)\n            _, pred0 = torch.max(outputs[0], 1)\n            _, pred1 = torch.max(outputs[1], 1)\n            _, pred2 = torch.max(outputs[2], 1)\n            preds = (pred0, pred1, pred2)\n            output_list.append(preds)\n            label_list.append(labels.transpose(1,0))\n        return output_list, label_list\n</code></pre>\n\n<h1>--- Prediction ---</h1>\n\n<p>sts = time.time()\ndata_type = 'test'\ncomponents = ['consonant_diacritic', 'grapheme_root', 'vowel_diacritic']\ntarget=[] # model predictions placeholder\nrow_id=[] # row_id place holder</p>\n\n<p>for i in tqdm(range(4)):\n    indices = [i]\n   ## obtain row_id and images at the same time ##\n    test_images, df_test_img = prepare_image(\n        DATADIR, FEATHERDIR, data_type = data_type, submission=SUBMISSION, indices=indices)\n    n_dataset = len(test_images)\n    print(f'i={i}, n_dataset={n_dataset}')\n    test_dataset = BengaliAIDataset(\n    test_images, None,\n    transform=data_transforms[data_type])\n    print('test_dataset', len(test_dataset))\n    test_loader = DataLoader(test_dataset, batch_size=BATCH_SIZE, shuffle=False, num_workers=WORKER)</p>\n\n<pre><code>test_preds = predict(model_ft, test_loader, data_type,DEVICE)\np0 = np.array([test_pred[0] for test_pred in test_preds]) #grapheme\np1 = np.array([test_pred[1] for test_pred in test_preds]) #vowel\np2 = np.array([test_pred[2] for test_pred in test_preds]) #consonant\ntgt = [p2, p0, p1]\nfor idx, id in enumerate(df_test_img[0].image_id.values): # df_test_img[0].image_id.values has the test image_ids\n    #print(idx)\n    for i,comp in enumerate(components):\n        id_sample=id+'_'+comp\n        row_id.append(id_sample)\n        target.append(tgt[i].flatten()[idx]) \ndel test_images, df_test_img\ngc.collect()\nif DEBUG:\n    break\n</code></pre>\n\n<p>ed = time.time()\nprint('Predicted in {:.3f} min'.format((ed-sts)/60))\ndel p0, p1, p2</p>\n\n<h1>create a dataframe with the solutions</h1>\n\n<p>sub_df = pd.DataFrame(\n    {'row_id': row_id,\n    'target':target\n    },\n    columns =['row_id','target'] \n)</p>\n\n<p>sub_df.to_csv('submission.csv', index=False)\n```</p>",
      "rawMarkdown": "🙌 🙌 \nI was encountering submission errors up to 7 times and spent almost 30 hours of GPU time for these errors during this weekend😪 \nAlthough I've check all the 37 topics searched by \"submission errors\" that can be found in the forum, but nothing worked for me. The submissions still became \"submission scoring error\". \n\nFortunately, I found [this notebook](https://www.kaggle.com/reasat/evaluating-model-performance-on-the-test-set) and modify the code a bit and the error is fixed eventually.\nHere I'd like to share the part of my code that worked, and what I've checked but not worked to you have the same situation.\nI guess the row_id mismatch caused the submission errors in my case. It needs to be obtained directly from the parquet files to avoid unexpected mismatch between hand-made row_id and the really row_id in the whole test set. (I've also tried obtain row_id from test.csv, but maybe something caused the unsynchronization of test.csv and test parquet, it didnt work. )\n\n**Refereces:**\n- [Evaluating model performance on the test set](https://www.kaggle.com/reasat/evaluating-model-performance-on-the-test-set)\n- [How to resolve \"Submission Error\" in my case](https://www.kaggle.com/c/bengaliai-cv19/discussion/123632)\n- [Bengali.AI SEResNeXt prediction with pytorch](https://www.kaggle.com/corochann/bengali-seresnext-prediction-with-pytorch)\n\n**What I've checked and not worked**\n- Notebook ran within 2 hours by replacing inference data to training dataset wi Batchsize=1 ==&gt; not a memory error, a time out error\n- Submission data type is matched with the sample_submission.csv ==&gt; not a scoring error\n- Output file is named \"submission.csv\" , and directly saved in the correct directly by `sub_df.to_csv('submission.csv', index=True)` ==&gt; not a CSV not found error\n\n**What worked for me**\n\n```\ndef prepare_image(datadir, featherdir, data_type='train',\n                  submission=True, indices=[0, 1, 2, 3]):\n    assert data_type in ['train', 'test']\n    if submission:\n        image_df_list = [pd.read_parquet(datadir / f'{data_type}_image_data_{i}.parquet')\n                         for i in indices]\n    else:\n        image_df_list = [pd.read_feather(featherdir / f'{data_type}_image_data_{i}.feather')\n                         for i in indices]\n\n    print('image_df_list', len(image_df_list))\n    HEIGHT = 137\n    WIDTH = 236\n    images = [df.iloc[:, 1:].values.reshape(-1, HEIGHT, WIDTH) for df in image_df_list]\n    #del image_df_list\n    gc.collect()\n    images = np.concatenate(images, axis=0)\n    return images, image_df_list\n\ndef predict(model, dataloaders, phase, device):\n    model.eval()\n    output_list = []\n    label_list = []\n    with torch.no_grad():\n        if phase == 'test':\n            for i, inputs in enumerate(tqdm(dataloaders)):\n                \n                inputs = inputs.to(device)\n                outputs = model(inputs)\n                _, pred0 = torch.max(outputs[0], 1)\n                _, pred1 = torch.max(outputs[1], 1)\n                _, pred2 = torch.max(outputs[2], 1)\n                preds = (pred0.cpu().numpy(), pred1.cpu().numpy(), pred2.cpu().numpy())\n                output_list.append(preds)\n            return output_list\n        elif phase == 'val':\n            for i, (inputs, labels) in enumerate(tqdm(dataloaders)):\n                \n                inputs = inputs.to(device)\n                outputs = model(inputs)\n                _, pred0 = torch.max(outputs[0], 1)\n                _, pred1 = torch.max(outputs[1], 1)\n                _, pred2 = torch.max(outputs[2], 1)\n                preds = (pred0, pred1, pred2)\n                output_list.append(preds)\n                label_list.append(labels.transpose(1,0))\n            return output_list, label_list\n\n# --- Prediction ---\n\nsts = time.time()\ndata_type = 'test'\ncomponents = ['consonant_diacritic', 'grapheme_root', 'vowel_diacritic']\ntarget=[] # model predictions placeholder\nrow_id=[] # row_id place holder\n\nfor i in tqdm(range(4)):\n    indices = [i]\n   ## obtain row_id and images at the same time ##\n    test_images, df_test_img = prepare_image(\n        DATADIR, FEATHERDIR, data_type = data_type, submission=SUBMISSION, indices=indices)\n    n_dataset = len(test_images)\n    print(f'i={i}, n_dataset={n_dataset}')\n    test_dataset = BengaliAIDataset(\n    test_images, None,\n    transform=data_transforms[data_type])\n    print('test_dataset', len(test_dataset))\n    test_loader = DataLoader(test_dataset, batch_size=BATCH_SIZE, shuffle=False, num_workers=WORKER)\n    \n    test_preds = predict(model_ft, test_loader, data_type,DEVICE)\n    p0 = np.array([test_pred[0] for test_pred in test_preds]) #grapheme\n    p1 = np.array([test_pred[1] for test_pred in test_preds]) #vowel\n    p2 = np.array([test_pred[2] for test_pred in test_preds]) #consonant\n    tgt = [p2, p0, p1]\n    for idx, id in enumerate(df_test_img[0].image_id.values): # df_test_img[0].image_id.values has the test image_ids\n        #print(idx)\n        for i,comp in enumerate(components):\n            id_sample=id+'_'+comp\n            row_id.append(id_sample)\n            target.append(tgt[i].flatten()[idx]) \n    del test_images, df_test_img\n    gc.collect()\n    if DEBUG:\n        break\ned = time.time()\nprint('Predicted in {:.3f} min'.format((ed-sts)/60))\ndel p0, p1, p2\n\n# create a dataframe with the solutions \nsub_df = pd.DataFrame(\n    {'row_id': row_id,\n    'target':target\n    },\n    columns =['row_id','target'] \n)\n\nsub_df.to_csv('submission.csv', index=False)\n```\n",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "767191": "🙌 🙌 \nI was encountering submission errors up to 7 times and spent almost 30 hours of GPU time for these errors during this weekend😪 \nAlthough I've check all the 37 topics searched by \"submission errors\" that can be found in the forum, but nothing worked for me. The submissions still became \"submission scoring error\". \n\nFortunately, I found [this notebook](https://www.kaggle.com/reasat/evaluating-model-performance-on-the-test-set) and modify the code a bit and the error is fixed eventually.\nHere I'd like to share the part of my code that worked, and what I've checked but not worked to you have the same situation.\nI guess the row_id mismatch caused the submission errors in my case. It needs to be obtained directly from the parquet files to avoid unexpected mismatch between hand-made row_id and the really row_id in the whole test set. (I've also tried obtain row_id from test.csv, but maybe something caused the unsynchronization of test.csv and test parquet, it didnt work. )\n\n**Refereces:**\n- [Evaluating model performance on the test set](https://www.kaggle.com/reasat/evaluating-model-performance-on-the-test-set)\n- [How to resolve \"Submission Error\" in my case](https://www.kaggle.com/c/bengaliai-cv19/discussion/123632)\n- [Bengali.AI SEResNeXt prediction with pytorch](https://www.kaggle.com/corochann/bengali-seresnext-prediction-with-pytorch)\n\n**What I've checked and not worked**\n- Notebook ran within 2 hours by replacing inference data to training dataset wi Batchsize=1 ==&gt; not a memory error, a time out error\n- Submission data type is matched with the sample_submission.csv ==&gt; not a scoring error\n- Output file is named \"submission.csv\" , and directly saved in the correct directly by `sub_df.to_csv('submission.csv', index=True)` ==&gt; not a CSV not found error\n\n**What worked for me**\n\n```\ndef prepare_image(datadir, featherdir, data_type='train',\n                  submission=True, indices=[0, 1, 2, 3]):\n    assert data_type in ['train', 'test']\n    if submission:\n        image_df_list = [pd.read_parquet(datadir / f'{data_type}_image_data_{i}.parquet')\n                         for i in indices]\n    else:\n        image_df_list = [pd.read_feather(featherdir / f'{data_type}_image_data_{i}.feather')\n                         for i in indices]\n\n    print('image_df_list', len(image_df_list))\n    HEIGHT = 137\n    WIDTH = 236\n    images = [df.iloc[:, 1:].values.reshape(-1, HEIGHT, WIDTH) for df in image_df_list]\n    #del image_df_list\n    gc.collect()\n    images = np.concatenate(images, axis=0)\n    return images, image_df_list\n\ndef predict(model, dataloaders, phase, device):\n    model.eval()\n    output_list = []\n    label_list = []\n    with torch.no_grad():\n        if phase == 'test':\n            for i, inputs in enumerate(tqdm(dataloaders)):\n                \n                inputs = inputs.to(device)\n                outputs = model(inputs)\n                _, pred0 = torch.max(outputs[0], 1)\n                _, pred1 = torch.max(outputs[1], 1)\n                _, pred2 = torch.max(outputs[2], 1)\n                preds = (pred0.cpu().numpy(), pred1.cpu().numpy(), pred2.cpu().numpy())\n                output_list.append(preds)\n            return output_list\n        elif phase == 'val':\n            for i, (inputs, labels) in enumerate(tqdm(dataloaders)):\n                \n                inputs = inputs.to(device)\n                outputs = model(inputs)\n                _, pred0 = torch.max(outputs[0], 1)\n                _, pred1 = torch.max(outputs[1], 1)\n                _, pred2 = torch.max(outputs[2], 1)\n                preds = (pred0, pred1, pred2)\n                output_list.append(preds)\n                label_list.append(labels.transpose(1,0))\n            return output_list, label_list\n\n# --- Prediction ---\n\nsts = time.time()\ndata_type = 'test'\ncomponents = ['consonant_diacritic', 'grapheme_root', 'vowel_diacritic']\ntarget=[] # model predictions placeholder\nrow_id=[] # row_id place holder\n\nfor i in tqdm(range(4)):\n    indices = [i]\n   ## obtain row_id and images at the same time ##\n    test_images, df_test_img = prepare_image(\n        DATADIR, FEATHERDIR, data_type = data_type, submission=SUBMISSION, indices=indices)\n    n_dataset = len(test_images)\n    print(f'i={i}, n_dataset={n_dataset}')\n    test_dataset = BengaliAIDataset(\n    test_images, None,\n    transform=data_transforms[data_type])\n    print('test_dataset', len(test_dataset))\n    test_loader = DataLoader(test_dataset, batch_size=BATCH_SIZE, shuffle=False, num_workers=WORKER)\n    \n    test_preds = predict(model_ft, test_loader, data_type,DEVICE)\n    p0 = np.array([test_pred[0] for test_pred in test_preds]) #grapheme\n    p1 = np.array([test_pred[1] for test_pred in test_preds]) #vowel\n    p2 = np.array([test_pred[2] for test_pred in test_preds]) #consonant\n    tgt = [p2, p0, p1]\n    for idx, id in enumerate(df_test_img[0].image_id.values): # df_test_img[0].image_id.values has the test image_ids\n        #print(idx)\n        for i,comp in enumerate(components):\n            id_sample=id+'_'+comp\n            row_id.append(id_sample)\n            target.append(tgt[i].flatten()[idx]) \n    del test_images, df_test_img\n    gc.collect()\n    if DEBUG:\n        break\ned = time.time()\nprint('Predicted in {:.3f} min'.format((ed-sts)/60))\ndel p0, p1, p2\n\n# create a dataframe with the solutions \nsub_df = pd.DataFrame(\n    {'row_id': row_id,\n    'target':target\n    },\n    columns =['row_id','target'] \n)\n\nsub_df.to_csv('submission.csv', index=False)\n```\n"
  }
}