{
  "id": 102297,
  "title": "when I use TTA ,I meet \"submission CSV not found\", but kernel success run success run, indeed help sincerely!",
  "url": "/competitions/aptos2019-blindness-detection/discussion/102297",
  "author_name": "",
  "post_date": "2019-08-01T07:10:25.345838400Z",
  "votes": 1,
  "comment_count": 12,
  "views": 0,
  "content": "<ul>\n<li>Batchsize:16</li>\n<li>image size:256</li>\n</ul>",
  "messages": [
    {
      "id": "589627",
      "postDate": "08/01/2019 07:10:25",
      "content": "<ul>\n<li>Batchsize:16</li>\n<li>image size:256</li>\n</ul>",
      "rawMarkdown": "Batchsize:16\n- image size:256",
      "votes": null
    },
    {
      "id": "589664",
      "postDate": "08/01/2019 08:10:18",
      "content": "<p>Re-submit the kernel may help. It may not cause by your code.</p>",
      "rawMarkdown": "Re-submit the kernel may help. It may not cause by your code.",
      "votes": null
    },
    {
      "id": "589675",
      "postDate": "08/01/2019 08:21:58",
      "content": "<p>when I delete the TTA code, I can commit success,  Addition, I try three model with TTA, unfortunately，they all fail, it make me anxious that waste my 3 times commit ,Do you meet the same problem?? thank you for your reply</p>",
      "rawMarkdown": "when I delete the TTA code, I can commit success,  Addition, I try three model with TTA, unfortunately，they all fail, it make me anxious that waste my 3 times commit ,Do you meet the same problem?? thank you for your reply",
      "votes": null
    },
    {
      "id": "589686",
      "postDate": "08/01/2019 08:42:31",
      "content": "<p>No I haven't try TTA yet. Can you show your TTA part code here?</p>",
      "rawMarkdown": "No I haven't try TTA yet. Can you show your TTA part code here?",
      "votes": null
    },
    {
      "id": "589708",
      "postDate": "08/01/2019 09:16:15",
      "content": "<p>TTA_steps = 10</p>\n\n<p>test_datagen = ImageDataGenerator(\n        featurewise_std_normalization= True,\n        zoom_range=0.15,\n        fill_mode='constant',\n        cval=0., # value used for fill_mode=\"constant\"\n        horizontal_flip = True,\n        vertical_flip = True,\n        rotation_range=360\n    )</p>\n\n<p>predictions = []\ntest_generator = test_datagen.flow(X_test, batch_size=16, shuffle=False)</p>\n\n<p>for i in range(TTA_steps): \n    print('<strong><em>*</em></strong>*TTA_{}'.format(i))\n    test_generator.reset()\n    preds = model.predict_generator(test_generator, steps=len(X_test)/16)\n    predictions.append(preds)\ny_test = np.mean(predictions, axis=0)</p>\n\n<p>y_test = np.argmax(y_test, axis=1)\ntest_df['diagnosis'] = y_test\ntest_df.to_csv('submission.csv',index=False)\nprint(test_df['diagnosis'].value_counts())\ntest_df['diagnosis'].hist()</p>",
      "rawMarkdown": "TTA_steps = 10\n\ntest_datagen = ImageDataGenerator(\n        featurewise_std_normalization= True,\n        zoom_range=0.15,\n        fill_mode='constant',\n        cval=0., # value used for fill_mode=\"constant\"\n        horizontal_flip = True,\n        vertical_flip = True,\n        rotation_range=360\n    )\n\npredictions = []\ntest_generator = test_datagen.flow(X_test, batch_size=16, shuffle=False)\n\nfor i in range(TTA_steps): \n    print('********TTA_{}'.format(i))\n    test_generator.reset()\n    preds = model.predict_generator(test_generator, steps=len(X_test)/16)\n    predictions.append(preds)\ny_test = np.mean(predictions, axis=0)\n\ny_test = np.argmax(y_test, axis=1)\ntest_df['diagnosis'] = y_test\ntest_df.to_csv('submission.csv',index=False)\nprint(test_df['diagnosis'].value_counts())\ntest_df['diagnosis'].hist()",
      "votes": null
    },
    {
      "id": "589711",
      "postDate": "08/01/2019 09:17:17",
      "content": "<p>thank your help me sinerely</p>",
      "rawMarkdown": "thank your help me sinerely",
      "votes": null
    },
    {
      "id": "589742",
      "postDate": "08/01/2019 09:53:33",
      "content": "<p>By the way, what's your commit kernel time cost?\nIn my experience, I got some submit problem when commit kernel time cost more than 400 secs.</p>",
      "rawMarkdown": "By the way, what's your commit kernel time cost?\nIn my experience, I got some submit problem when commit kernel time cost more than 400 secs.",
      "votes": null
    },
    {
      "id": "589783",
      "postDate": "08/01/2019 11:18:13",
      "content": "<p>I run code in kernel since this is a Kernel-only. Approximately, I spend 1 hour when I commit the kernel every times</p>",
      "rawMarkdown": "I run code in kernel since this is a Kernel-only. Approximately, I spend 1 hour when I commit the kernel every times",
      "votes": null
    },
    {
      "id": "590068",
      "postDate": "08/01/2019 19:04:41",
      "content": "<p>I had the exact same problem. The issue is that the private test dataset (the one used for the LB score) is ~6 times larger than the public one, that's why you can run it on your kernel but not on the private dataset. In my case this caused a memory overflow because I was using TTA with <code>flow</code> (which loads all the data in memory) instead of using <code>flow_from_dataframe</code>.</p>\n\n<p>This was the code that was causing the submission error: \n```\ndef TTA(model, generator, images, n_examples=10):\n    val_predictions = []\n    val_generator = generator.flow(images, shuffle=False, batch_size=1)</p>\n\n<pre><code>for i in tqdm(range(n_examples)):\n    val_generator.reset()\n    val_preds = model.predict_generator(\n        val_generator, \n        steps=images.shape[0]\n        ) \n    val_predictions.append(val_preds)\n</code></pre>\n\n<p>```\nand this was the fix:</p>\n\n<p><code>\ndef TTA(model, data_generator, images_df, n_examples=10):\n    predictions = []\n    generator = data_generator.flow_from_dataframe(\n        dataframe=images_df,\n        directory='../input/aptos2019-blindness-detection/test_images/',\n        x_col = 'id_code',\n        y_col=None,\n        target_size = IMAGE_SIZE,\n        shuffle=False, \n        batch_size=1,\n        seed=2019,\n        class_mode=None)\n</code></p>\n\n<p>I hope it helps! :)</p>",
      "rawMarkdown": "I had the exact same problem. The issue is that the private test dataset (the one used for the LB score) is ~6 times larger than the public one, that's why you can run it on your kernel but not on the private dataset. In my case this caused a memory overflow because I was using TTA with `flow` (which loads all the data in memory) instead of using `flow_from_dataframe`.\n\nThis was the code that was causing the submission error: \n```\ndef TTA(model, generator, images, n_examples=10):\n    val_predictions = []\n    val_generator = generator.flow(images, shuffle=False, batch_size=1)\n\n    for i in tqdm(range(n_examples)):\n        val_generator.reset()\n        val_preds = model.predict_generator(\n            val_generator, \n            steps=images.shape[0]\n            ) \n        val_predictions.append(val_preds)\n```\nand this was the fix:\n\n```\ndef TTA(model, data_generator, images_df, n_examples=10):\n    predictions = []\n    generator = data_generator.flow_from_dataframe(\n        dataframe=images_df,\n        directory='../input/aptos2019-blindness-detection/test_images/',\n        x_col = 'id_code',\n        y_col=None,\n        target_size = IMAGE_SIZE,\n        shuffle=False, \n        batch_size=1,\n        seed=2019,\n        class_mode=None)\n```\n\nI hope it helps! :)",
      "votes": null
    },
    {
      "id": "590266",
      "postDate": "08/02/2019 01:15:56",
      "content": "<p>Thank you! I will try it, hope it work!!</p>",
      "rawMarkdown": "Thank you! I will try it, hope it work!!",
      "votes": null
    },
    {
      "id": "590301",
      "postDate": "08/02/2019 03:21:47",
      "content": "<p>Additionlly,  I have another question that the train data should be deal as same as the test data?</p>",
      "rawMarkdown": "Additionlly,  I have another question that the train data should be deal as same as the test data?",
      "votes": null
    },
    {
      "id": "590415",
      "postDate": "08/02/2019 07:06:18",
      "content": "<p>In my case It wasn't necessary because the train set is ~3600 images and when I loaded it into an array and trained with <code>flow</code> it could fit in memory (my image size was 224).</p>",
      "rawMarkdown": "In my case It wasn't necessary because the train set is ~3600 images and when I loaded it into an array and trained with `flow` it could fit in memory (my image size was 224).",
      "votes": null
    },
    {
      "id": "590438",
      "postDate": "08/02/2019 07:36:56",
      "content": "<p>as you means, I can use flow( ) function to deal train data,and use flow_from_dataframe() function deal test data ,  that guarantee success commit and avoid \"Submissions CSV not found\". (I am now  try the flow_from_dataframe() to deal test Images to implement TTA). hope it work, and thank you sinerely, I am a beginner </p>",
      "rawMarkdown": "as you means, I can use flow( ) function to deal train data,and use flow_from_dataframe() function deal test data ,  that guarantee success commit and avoid \"Submissions CSV not found\". (I am now  try the flow_from_dataframe() to deal test Images to implement TTA). hope it work, and thank you sinerely, I am a beginner",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 589664,
      "author_name": "haqishen",
      "author_url": "",
      "post_date": "08/01/2019 08:10:18",
      "content": "<p>Re-submit the kernel may help. It may not cause by your code.</p>",
      "votes": null,
      "replies": [
        {
          "id": 589675,
          "author_name": "jiangkun2",
          "author_url": "",
          "post_date": "08/01/2019 08:21:58",
          "content": "<p>when I delete the TTA code, I can commit success,  Addition, I try three model with TTA, unfortunately，they all fail, it make me anxious that waste my 3 times commit ,Do you meet the same problem?? thank you for your reply</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 589686,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "08/01/2019 08:42:31",
          "content": "<p>No I haven't try TTA yet. Can you show your TTA part code here?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 589708,
          "author_name": "jiangkun2",
          "author_url": "",
          "post_date": "08/01/2019 09:16:15",
          "content": "<p>TTA_steps = 10</p>\n\n<p>test_datagen = ImageDataGenerator(\n        featurewise_std_normalization= True,\n        zoom_range=0.15,\n        fill_mode='constant',\n        cval=0., # value used for fill_mode=\"constant\"\n        horizontal_flip = True,\n        vertical_flip = True,\n        rotation_range=360\n    )</p>\n\n<p>predictions = []\ntest_generator = test_datagen.flow(X_test, batch_size=16, shuffle=False)</p>\n\n<p>for i in range(TTA_steps): \n    print('<strong><em>*</em></strong>*TTA_{}'.format(i))\n    test_generator.reset()\n    preds = model.predict_generator(test_generator, steps=len(X_test)/16)\n    predictions.append(preds)\ny_test = np.mean(predictions, axis=0)</p>\n\n<p>y_test = np.argmax(y_test, axis=1)\ntest_df['diagnosis'] = y_test\ntest_df.to_csv('submission.csv',index=False)\nprint(test_df['diagnosis'].value_counts())\ntest_df['diagnosis'].hist()</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 589711,
          "author_name": "jiangkun2",
          "author_url": "",
          "post_date": "08/01/2019 09:17:17",
          "content": "<p>thank your help me sinerely</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 589742,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "08/01/2019 09:53:33",
          "content": "<p>By the way, what's your commit kernel time cost?\nIn my experience, I got some submit problem when commit kernel time cost more than 400 secs.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 589783,
          "author_name": "jiangkun2",
          "author_url": "",
          "post_date": "08/01/2019 11:18:13",
          "content": "<p>I run code in kernel since this is a Kernel-only. Approximately, I spend 1 hour when I commit the kernel every times</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 590068,
      "author_name": "sam1320",
      "author_url": "",
      "post_date": "08/01/2019 19:04:41",
      "content": "<p>I had the exact same problem. The issue is that the private test dataset (the one used for the LB score) is ~6 times larger than the public one, that's why you can run it on your kernel but not on the private dataset. In my case this caused a memory overflow because I was using TTA with <code>flow</code> (which loads all the data in memory) instead of using <code>flow_from_dataframe</code>.</p>\n\n<p>This was the code that was causing the submission error: \n```\ndef TTA(model, generator, images, n_examples=10):\n    val_predictions = []\n    val_generator = generator.flow(images, shuffle=False, batch_size=1)</p>\n\n<pre><code>for i in tqdm(range(n_examples)):\n    val_generator.reset()\n    val_preds = model.predict_generator(\n        val_generator, \n        steps=images.shape[0]\n        ) \n    val_predictions.append(val_preds)\n</code></pre>\n\n<p>```\nand this was the fix:</p>\n\n<p><code>\ndef TTA(model, data_generator, images_df, n_examples=10):\n    predictions = []\n    generator = data_generator.flow_from_dataframe(\n        dataframe=images_df,\n        directory='../input/aptos2019-blindness-detection/test_images/',\n        x_col = 'id_code',\n        y_col=None,\n        target_size = IMAGE_SIZE,\n        shuffle=False, \n        batch_size=1,\n        seed=2019,\n        class_mode=None)\n</code></p>\n\n<p>I hope it helps! :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 590266,
          "author_name": "jiangkun2",
          "author_url": "",
          "post_date": "08/02/2019 01:15:56",
          "content": "<p>Thank you! I will try it, hope it work!!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 590301,
          "author_name": "jiangkun2",
          "author_url": "",
          "post_date": "08/02/2019 03:21:47",
          "content": "<p>Additionlly,  I have another question that the train data should be deal as same as the test data?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 590415,
          "author_name": "sam1320",
          "author_url": "",
          "post_date": "08/02/2019 07:06:18",
          "content": "<p>In my case It wasn't necessary because the train set is ~3600 images and when I loaded it into an array and trained with <code>flow</code> it could fit in memory (my image size was 224).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 590438,
          "author_name": "jiangkun2",
          "author_url": "",
          "post_date": "08/02/2019 07:36:56",
          "content": "<p>as you means, I can use flow( ) function to deal train data,and use flow_from_dataframe() function deal test data ,  that guarantee success commit and avoid \"Submissions CSV not found\". (I am now  try the flow_from_dataframe() to deal test Images to implement TTA). hope it work, and thank you sinerely, I am a beginner </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "589627": "Batchsize:16\n- image size:256",
    "589664": "Re-submit the kernel may help. It may not cause by your code.",
    "589675": "when I delete the TTA code, I can commit success,  Addition, I try three model with TTA, unfortunately，they all fail, it make me anxious that waste my 3 times commit ,Do you meet the same problem?? thank you for your reply",
    "589686": "No I haven't try TTA yet. Can you show your TTA part code here?",
    "589708": "TTA_steps = 10\n\ntest_datagen = ImageDataGenerator(\n        featurewise_std_normalization= True,\n        zoom_range=0.15,\n        fill_mode='constant',\n        cval=0., # value used for fill_mode=\"constant\"\n        horizontal_flip = True,\n        vertical_flip = True,\n        rotation_range=360\n    )\n\npredictions = []\ntest_generator = test_datagen.flow(X_test, batch_size=16, shuffle=False)\n\nfor i in range(TTA_steps): \n    print('********TTA_{}'.format(i))\n    test_generator.reset()\n    preds = model.predict_generator(test_generator, steps=len(X_test)/16)\n    predictions.append(preds)\ny_test = np.mean(predictions, axis=0)\n\ny_test = np.argmax(y_test, axis=1)\ntest_df['diagnosis'] = y_test\ntest_df.to_csv('submission.csv',index=False)\nprint(test_df['diagnosis'].value_counts())\ntest_df['diagnosis'].hist()",
    "589711": "thank your help me sinerely",
    "589742": "By the way, what's your commit kernel time cost?\nIn my experience, I got some submit problem when commit kernel time cost more than 400 secs.",
    "589783": "I run code in kernel since this is a Kernel-only. Approximately, I spend 1 hour when I commit the kernel every times",
    "590068": "I had the exact same problem. The issue is that the private test dataset (the one used for the LB score) is ~6 times larger than the public one, that's why you can run it on your kernel but not on the private dataset. In my case this caused a memory overflow because I was using TTA with `flow` (which loads all the data in memory) instead of using `flow_from_dataframe`.\n\nThis was the code that was causing the submission error: \n```\ndef TTA(model, generator, images, n_examples=10):\n    val_predictions = []\n    val_generator = generator.flow(images, shuffle=False, batch_size=1)\n\n    for i in tqdm(range(n_examples)):\n        val_generator.reset()\n        val_preds = model.predict_generator(\n            val_generator, \n            steps=images.shape[0]\n            ) \n        val_predictions.append(val_preds)\n```\nand this was the fix:\n\n```\ndef TTA(model, data_generator, images_df, n_examples=10):\n    predictions = []\n    generator = data_generator.flow_from_dataframe(\n        dataframe=images_df,\n        directory='../input/aptos2019-blindness-detection/test_images/',\n        x_col = 'id_code',\n        y_col=None,\n        target_size = IMAGE_SIZE,\n        shuffle=False, \n        batch_size=1,\n        seed=2019,\n        class_mode=None)\n```\n\nI hope it helps! :)",
    "590266": "Thank you! I will try it, hope it work!!",
    "590301": "Additionlly,  I have another question that the train data should be deal as same as the test data?",
    "590415": "In my case It wasn't necessary because the train set is ~3600 images and when I loaded it into an array and trained with `flow` it could fit in memory (my image size was 224).",
    "590438": "as you means, I can use flow( ) function to deal train data,and use flow_from_dataframe() function deal test data ,  that guarantee success commit and avoid \"Submissions CSV not found\". (I am now  try the flow_from_dataframe() to deal test Images to implement TTA). hope it work, and thank you sinerely, I am a beginner"
  },
  "source": "meta"
}