{
  "id": 125991,
  "title": "Yet another Submission Scoring Error",
  "url": "/competitions/bengaliai-cv19/discussion/125991",
  "author_name": "",
  "post_date": "2020-01-15T02:02:46.978957700Z",
  "votes": 2,
  "comment_count": 12,
  "views": 0,
  "content": "<p>models trained locally, weights uploaded into a dataset, commits run successfully on both test parquets and train parquets so I would not suspect OOM.</p>\n\n<p>and this is how I create the submission df as I found it on a kernel already...\nI have waisted 2 days debugging but there is nothing wrong with it, it even worked in the past with a single model, output for test parquets is just identical to everyone else.</p>\n\n<p>Any ideas?</p>\n\n<p><code>\nrow_id = []\ntarget = []\nfor i in range(len(grapheme_root_labels)):\n    row_id += [f'Test_{i}_grapheme_root',\n               f'Test_{i}_vowel_diacritic',\n               f'Test_{i}_consonant_diacritic']\n    target += [grapheme_root_labels[i], vowel_diacritic_labels[i], consonant_diacritic_labels[i]]\nsubmission_df = pd.DataFrame({'row_id': row_id, 'target': target})\nsubmission_df\n</code></p>",
  "messages": [
    {
      "id": "718972",
      "postDate": "01/15/2020 02:02:46",
      "content": "<p>models trained locally, weights uploaded into a dataset, commits run successfully on both test parquets and train parquets so I would not suspect OOM.</p>\n\n<p>and this is how I create the submission df as I found it on a kernel already...\nI have waisted 2 days debugging but there is nothing wrong with it, it even worked in the past with a single model, output for test parquets is just identical to everyone else.</p>\n\n<p>Any ideas?</p>\n\n<p><code>\nrow_id = []\ntarget = []\nfor i in range(len(grapheme_root_labels)):\n    row_id += [f'Test_{i}_grapheme_root',\n               f'Test_{i}_vowel_diacritic',\n               f'Test_{i}_consonant_diacritic']\n    target += [grapheme_root_labels[i], vowel_diacritic_labels[i], consonant_diacritic_labels[i]]\nsubmission_df = pd.DataFrame({'row_id': row_id, 'target': target})\nsubmission_df\n</code></p>",
      "rawMarkdown": "models trained locally, weights uploaded into a dataset, commits run successfully on both test parquets and train parquets so I would not suspect OOM.\n\n and this is how I create the submission df as I found it on a kernel already...\nI have waisted 2 days debugging but there is nothing wrong with it, it even worked in the past with a single model, output for test parquets is just identical to everyone else.\n\nAny ideas?\n\n```\nrow_id = []\ntarget = []\nfor i in range(len(grapheme_root_labels)):\n    row_id += [f'Test_{i}_grapheme_root',\n               f'Test_{i}_vowel_diacritic',\n               f'Test_{i}_consonant_diacritic']\n    target += [grapheme_root_labels[i], vowel_diacritic_labels[i], consonant_diacritic_labels[i]]\nsubmission_df = pd.DataFrame({'row_id': row_id, 'target': target})\nsubmission_df\n```",
      "votes": null
    },
    {
      "id": "718994",
      "postDate": "01/15/2020 02:24:12",
      "content": "<p>No issues in above code. I still feels like it is memory error.\nCan you share logic before prediction, I mean loading and processing of input file.</p>",
      "rawMarkdown": "No issues in above code. I still feels like it is memory error.\nCan you share logic before prediction, I mean loading and processing of input file.",
      "votes": null
    },
    {
      "id": "719000",
      "postDate": "01/15/2020 02:30:00",
      "content": "<p>Can you spot anything? training dataset works just fine though..\n```\nfor i in range(4):\n    test = construct_img_arr(i) # &lt; returns image arr\n    for i, model in enumerate(models):\n        g, v, c = model.predict(test, batch_size=512, verbose=1)\n        grapheme_root[i].append(g.argmax(axis=1).astype(np.uint8))\n        vowel_diacritic[i].append(v.argmax(axis=1).astype(np.uint8))\n        consonant_diacritic[i].append(c.argmax(axis=1).astype(np.uint8))</p>\n\n<h1>ravel</h1>\n\n<p>for i in range(len(models)):\n    grapheme_root[i] = np.ravel(grapheme_root[i])\n    vowel_diacritic[i] = np.ravel(vowel_diacritic[i])\n    consonant_diacritic[i] = np.ravel(consonant_diacritic[i])</p>\n\n<h1>take votes</h1>\n\n<p>grapheme_root_labels = stats.mode(np.vstack(grapheme_root))[0].ravel()\nvowel_diacritic_labels = stats.mode(np.vstack(vowel_diacritic))[0].ravel()\n consonant_diacritic_labels = stats.mode(np.vstack(consonant_diacritic))[0].ravel()\n```</p>",
      "rawMarkdown": "Can you spot anything? training dataset works just fine though..\n```\nfor i in range(4):\n    test = construct_img_arr(i) # &lt; returns image arr\n    for i, model in enumerate(models):\n        g, v, c = model.predict(test, batch_size=512, verbose=1)\n        grapheme_root[i].append(g.argmax(axis=1).astype(np.uint8))\n        vowel_diacritic[i].append(v.argmax(axis=1).astype(np.uint8))\n        consonant_diacritic[i].append(c.argmax(axis=1).astype(np.uint8))\n\n#ravel\nfor i in range(len(models)):\n    grapheme_root[i] = np.ravel(grapheme_root[i])\n    vowel_diacritic[i] = np.ravel(vowel_diacritic[i])\n    consonant_diacritic[i] = np.ravel(consonant_diacritic[i])\n    \n# take votes\ngrapheme_root_labels = stats.mode(np.vstack(grapheme_root))[0].ravel()\nvowel_diacritic_labels = stats.mode(np.vstack(vowel_diacritic))[0].ravel()\n consonant_diacritic_labels = stats.mode(np.vstack(consonant_diacritic))[0].ravel()\n```",
      "votes": null
    },
    {
      "id": "719017",
      "postDate": "01/15/2020 03:05:25",
      "content": "<p>There are two areas where issue might occur 1st is construct_img_arr function and next is ravel() function.\nOne more thing the actual test set even larger than train set so might be you are consuming just little above margin memory causing this error.</p>",
      "rawMarkdown": "There are two areas where issue might occur 1st is construct_img_arr function and next is ravel() function.\nOne more thing the actual test set even larger than train set so might be you are consuming just little above margin memory causing this error.",
      "votes": null
    },
    {
      "id": "719018",
      "postDate": "01/15/2020 03:09:33",
      "content": "<p>do you know how many images are in the test? by what factor is it larger?</p>",
      "rawMarkdown": "do you know how many images are in the test? by what factor is it larger?",
      "votes": null
    },
    {
      "id": "719021",
      "postDate": "01/15/2020 03:17:50",
      "content": "<p>No I don't know the multiplication factor.</p>\n\n<p>As you said it worked earlier with single model, may be you can try and check with 2 models instead of multiple models in first go.</p>",
      "rawMarkdown": "No I don't know the multiplication factor.\n\nAs you said it worked earlier with single model, may be you can try and check with 2 models instead of multiple models in first go.",
      "votes": null
    },
    {
      "id": "719023",
      "postDate": "01/15/2020 03:21:55",
      "content": "<p>will give it a shot</p>",
      "rawMarkdown": "will give it a shot",
      "votes": null
    },
    {
      "id": "719496",
      "postDate": "01/15/2020 14:55:10",
      "content": "<p>Test images are about as many as the train set, i.e. 200840 (50210 x 4). Try predicting all train images and check memory and time limits. </p>",
      "rawMarkdown": "Test images are about as many as the train set, i.e. 200840 (50210 x 4). Try predicting all train images and check memory and time limits.",
      "votes": null
    },
    {
      "id": "719629",
      "postDate": "01/15/2020 17:04:34",
      "content": "<p>training images work without any problems, i am running out of ideas and the errors are not informative</p>",
      "rawMarkdown": "training images work without any problems, i am running out of ideas and the errors are not informative",
      "votes": null
    },
    {
      "id": "720025",
      "postDate": "01/16/2020 04:48:46",
      "content": "<p>hey, i had a similar issue of submission scoring error. The issue was due to hard-coding the **row_id **column .</p>\n\n<p>Change I did was to use the <strong>imageid column directly from test_imgs parquets to generate rowid column</strong></p>\n\n<p>Maybe you can try above to get your notebook work on private tests set as well.</p>",
      "rawMarkdown": "hey, i had a similar issue of submission scoring error. The issue was due to hard-coding the **row_id **column .\n\nChange I did was to use the **imageid column directly from test_imgs parquets to generate rowid column**\n\nMaybe you can try above to get your notebook work on private tests set as well.",
      "votes": null
    },
    {
      "id": "720034",
      "postDate": "01/16/2020 05:00:04",
      "content": "<p>you mean that test parquets do not follow the <code>Test_{i}</code> format?</p>",
      "rawMarkdown": "you mean that test parquets do not follow the `Test_{i}` format?",
      "votes": null
    },
    {
      "id": "733812",
      "postDate": "01/31/2020 14:46:23",
      "content": "<p>Did you solve your problem? I'm having the same trouble.</p>",
      "rawMarkdown": "Did you solve your problem? I'm having the same trouble.",
      "votes": null
    },
    {
      "id": "754878",
      "postDate": "02/24/2020 06:55:37",
      "content": "<p>Your sequence in <code>submission_df</code> is wrong. Your code\n<code>python\nrow_id += [f'Test_{i}_grapheme_root',\n           f'Test_{i}_vowel_diacritic',\n           f'Test_{i}_consonant_diacritic']\ntarget += [grapheme_root_labels[i], vowel_diacritic_labels[i], consonant_diacritic_labels[i]]\n</code>\nprints <code>grapheme_root</code> first. Whereas, the submission file expects <code>consonant_diacritic</code> to be the first value for each image in the submission.csv, then <code>grapheme_root</code> and last <code>vowel_diacritic</code>. You can replace the above snippet with the following and that should fix your submission error issue.\n<code>python\nrow_id += [f'Test_{i}_consonant_diacritic',\n           f'Test_{i}_grapheme_root',\n           f'Test_{i}_vowel_diacritic']\ntarget += [consonant_diacritic_labels[i], grapheme_root_labels[i], vowel_diacritic_labels[i]]\n</code></p>",
      "rawMarkdown": "Your sequence in `submission_df` is wrong. Your code\n```python\nrow_id += [f'Test_{i}_grapheme_root',\n           f'Test_{i}_vowel_diacritic',\n           f'Test_{i}_consonant_diacritic']\ntarget += [grapheme_root_labels[i], vowel_diacritic_labels[i], consonant_diacritic_labels[i]]\n```\nprints `grapheme_root` first. Whereas, the submission file expects `consonant_diacritic` to be the first value for each image in the submission.csv, then `grapheme_root` and last `vowel_diacritic`. You can replace the above snippet with the following and that should fix your submission error issue.\n```python\nrow_id += [f'Test_{i}_consonant_diacritic',\n           f'Test_{i}_grapheme_root',\n           f'Test_{i}_vowel_diacritic']\ntarget += [consonant_diacritic_labels[i], grapheme_root_labels[i], vowel_diacritic_labels[i]]\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 718994,
      "author_name": "amit9484",
      "author_url": "",
      "post_date": "01/15/2020 02:24:12",
      "content": "<p>No issues in above code. I still feels like it is memory error.\nCan you share logic before prediction, I mean loading and processing of input file.</p>",
      "votes": null,
      "replies": [
        {
          "id": 719000,
          "author_name": "ma7555",
          "author_url": "",
          "post_date": "01/15/2020 02:30:00",
          "content": "<p>Can you spot anything? training dataset works just fine though..\n```\nfor i in range(4):\n    test = construct_img_arr(i) # &lt; returns image arr\n    for i, model in enumerate(models):\n        g, v, c = model.predict(test, batch_size=512, verbose=1)\n        grapheme_root[i].append(g.argmax(axis=1).astype(np.uint8))\n        vowel_diacritic[i].append(v.argmax(axis=1).astype(np.uint8))\n        consonant_diacritic[i].append(c.argmax(axis=1).astype(np.uint8))</p>\n\n<h1>ravel</h1>\n\n<p>for i in range(len(models)):\n    grapheme_root[i] = np.ravel(grapheme_root[i])\n    vowel_diacritic[i] = np.ravel(vowel_diacritic[i])\n    consonant_diacritic[i] = np.ravel(consonant_diacritic[i])</p>\n\n<h1>take votes</h1>\n\n<p>grapheme_root_labels = stats.mode(np.vstack(grapheme_root))[0].ravel()\nvowel_diacritic_labels = stats.mode(np.vstack(vowel_diacritic))[0].ravel()\n consonant_diacritic_labels = stats.mode(np.vstack(consonant_diacritic))[0].ravel()\n```</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 719017,
          "author_name": "amit9484",
          "author_url": "",
          "post_date": "01/15/2020 03:05:25",
          "content": "<p>There are two areas where issue might occur 1st is construct_img_arr function and next is ravel() function.\nOne more thing the actual test set even larger than train set so might be you are consuming just little above margin memory causing this error.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 719018,
          "author_name": "ma7555",
          "author_url": "",
          "post_date": "01/15/2020 03:09:33",
          "content": "<p>do you know how many images are in the test? by what factor is it larger?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 719021,
          "author_name": "amit9484",
          "author_url": "",
          "post_date": "01/15/2020 03:17:50",
          "content": "<p>No I don't know the multiplication factor.</p>\n\n<p>As you said it worked earlier with single model, may be you can try and check with 2 models instead of multiple models in first go.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 719023,
          "author_name": "ma7555",
          "author_url": "",
          "post_date": "01/15/2020 03:21:55",
          "content": "<p>will give it a shot</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 719496,
          "author_name": "imeintanis",
          "author_url": "",
          "post_date": "01/15/2020 14:55:10",
          "content": "<p>Test images are about as many as the train set, i.e. 200840 (50210 x 4). Try predicting all train images and check memory and time limits. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 719629,
          "author_name": "ma7555",
          "author_url": "",
          "post_date": "01/15/2020 17:04:34",
          "content": "<p>training images work without any problems, i am running out of ideas and the errors are not informative</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 720025,
      "author_name": "ankitsajwan",
      "author_url": "",
      "post_date": "01/16/2020 04:48:46",
      "content": "<p>hey, i had a similar issue of submission scoring error. The issue was due to hard-coding the **row_id **column .</p>\n\n<p>Change I did was to use the <strong>imageid column directly from test_imgs parquets to generate rowid column</strong></p>\n\n<p>Maybe you can try above to get your notebook work on private tests set as well.</p>",
      "votes": null,
      "replies": [
        {
          "id": 720034,
          "author_name": "ma7555",
          "author_url": "",
          "post_date": "01/16/2020 05:00:04",
          "content": "<p>you mean that test parquets do not follow the <code>Test_{i}</code> format?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 733812,
      "author_name": "syoya1997",
      "author_url": "",
      "post_date": "01/31/2020 14:46:23",
      "content": "<p>Did you solve your problem? I'm having the same trouble.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 754878,
      "author_name": "santosh16k",
      "author_url": "",
      "post_date": "02/24/2020 06:55:37",
      "content": "<p>Your sequence in <code>submission_df</code> is wrong. Your code\n<code>python\nrow_id += [f'Test_{i}_grapheme_root',\n           f'Test_{i}_vowel_diacritic',\n           f'Test_{i}_consonant_diacritic']\ntarget += [grapheme_root_labels[i], vowel_diacritic_labels[i], consonant_diacritic_labels[i]]\n</code>\nprints <code>grapheme_root</code> first. Whereas, the submission file expects <code>consonant_diacritic</code> to be the first value for each image in the submission.csv, then <code>grapheme_root</code> and last <code>vowel_diacritic</code>. You can replace the above snippet with the following and that should fix your submission error issue.\n<code>python\nrow_id += [f'Test_{i}_consonant_diacritic',\n           f'Test_{i}_grapheme_root',\n           f'Test_{i}_vowel_diacritic']\ntarget += [consonant_diacritic_labels[i], grapheme_root_labels[i], vowel_diacritic_labels[i]]\n</code></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "718972": "models trained locally, weights uploaded into a dataset, commits run successfully on both test parquets and train parquets so I would not suspect OOM.\n\n and this is how I create the submission df as I found it on a kernel already...\nI have waisted 2 days debugging but there is nothing wrong with it, it even worked in the past with a single model, output for test parquets is just identical to everyone else.\n\nAny ideas?\n\n```\nrow_id = []\ntarget = []\nfor i in range(len(grapheme_root_labels)):\n    row_id += [f'Test_{i}_grapheme_root',\n               f'Test_{i}_vowel_diacritic',\n               f'Test_{i}_consonant_diacritic']\n    target += [grapheme_root_labels[i], vowel_diacritic_labels[i], consonant_diacritic_labels[i]]\nsubmission_df = pd.DataFrame({'row_id': row_id, 'target': target})\nsubmission_df\n```",
    "718994": "No issues in above code. I still feels like it is memory error.\nCan you share logic before prediction, I mean loading and processing of input file.",
    "719000": "Can you spot anything? training dataset works just fine though..\n```\nfor i in range(4):\n    test = construct_img_arr(i) # &lt; returns image arr\n    for i, model in enumerate(models):\n        g, v, c = model.predict(test, batch_size=512, verbose=1)\n        grapheme_root[i].append(g.argmax(axis=1).astype(np.uint8))\n        vowel_diacritic[i].append(v.argmax(axis=1).astype(np.uint8))\n        consonant_diacritic[i].append(c.argmax(axis=1).astype(np.uint8))\n\n#ravel\nfor i in range(len(models)):\n    grapheme_root[i] = np.ravel(grapheme_root[i])\n    vowel_diacritic[i] = np.ravel(vowel_diacritic[i])\n    consonant_diacritic[i] = np.ravel(consonant_diacritic[i])\n    \n# take votes\ngrapheme_root_labels = stats.mode(np.vstack(grapheme_root))[0].ravel()\nvowel_diacritic_labels = stats.mode(np.vstack(vowel_diacritic))[0].ravel()\n consonant_diacritic_labels = stats.mode(np.vstack(consonant_diacritic))[0].ravel()\n```",
    "719017": "There are two areas where issue might occur 1st is construct_img_arr function and next is ravel() function.\nOne more thing the actual test set even larger than train set so might be you are consuming just little above margin memory causing this error.",
    "719018": "do you know how many images are in the test? by what factor is it larger?",
    "719021": "No I don't know the multiplication factor.\n\nAs you said it worked earlier with single model, may be you can try and check with 2 models instead of multiple models in first go.",
    "719023": "will give it a shot",
    "719496": "Test images are about as many as the train set, i.e. 200840 (50210 x 4). Try predicting all train images and check memory and time limits.",
    "719629": "training images work without any problems, i am running out of ideas and the errors are not informative",
    "720025": "hey, i had a similar issue of submission scoring error. The issue was due to hard-coding the **row_id **column .\n\nChange I did was to use the **imageid column directly from test_imgs parquets to generate rowid column**\n\nMaybe you can try above to get your notebook work on private tests set as well.",
    "720034": "you mean that test parquets do not follow the `Test_{i}` format?",
    "733812": "Did you solve your problem? I'm having the same trouble.",
    "754878": "Your sequence in `submission_df` is wrong. Your code\n```python\nrow_id += [f'Test_{i}_grapheme_root',\n           f'Test_{i}_vowel_diacritic',\n           f'Test_{i}_consonant_diacritic']\ntarget += [grapheme_root_labels[i], vowel_diacritic_labels[i], consonant_diacritic_labels[i]]\n```\nprints `grapheme_root` first. Whereas, the submission file expects `consonant_diacritic` to be the first value for each image in the submission.csv, then `grapheme_root` and last `vowel_diacritic`. You can replace the above snippet with the following and that should fix your submission error issue.\n```python\nrow_id += [f'Test_{i}_consonant_diacritic',\n           f'Test_{i}_grapheme_root',\n           f'Test_{i}_vowel_diacritic']\ntarget += [consonant_diacritic_labels[i], grapheme_root_labels[i], vowel_diacritic_labels[i]]\n```"
  },
  "source": "meta"
}