{
  "id": 199468,
  "title": "Large difference between CV and LB (closed)",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/199468",
  "author_name": "",
  "post_date": "2020-11-25T20:34:32.185304100Z",
  "votes": 1,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hello! Can anybody helps with my problem?</p>\n<p>I trained my simple model in this training notebook <a href=\"https://www.kaggle.com/khlevnov/efficientnetb0-training\" target=\"_blank\">EfficientNetB0 training</a> and take 0.800+ accuracy with hold-out CV. Results of the last epoch below (and you can see it in notebook):</p>\n<pre><code>Epoch 10/10\n1204/1204 - 236s - loss: 0.4054 - accuracy: 0.8602 - val_loss: 0.4383 - val_accuracy: 0.8570\n</code></pre>\n<p>Then I saved model in .h5 file and created <a href=\"https://www.kaggle.com/khlevnov/cassava-preptrained-model-2\" target=\"_blank\">public dataset</a> with notebook output (and saved model also).</p>\n<p>For submission I created separate notebook <a href=\"https://www.kaggle.com/khlevnov/efficientnetb0-submission\" target=\"_blank\">EfficientNetB0 submission</a> only for prediction using saved model from my freshly created dataset. But after submission… Boom, I have only 0.400+ accuracy on public leaderboard.</p>\n<pre><code>Succeeded\n0.417\n</code></pre>\n<p>Seems my public score divided by two(<br>\nWhere should I search my dumb error?</p>\n<p>Will be happy to listen any suggestions, thanks.</p>\n<p>UPDATE:<br>\nAs <a href=\"https://www.kaggle.com/richardepstein\" target=\"_blank\">@richardepstein</a> said, there are different order between <code>image_id</code> and <code>label</code> columns in my submission. Thanks to all.</p>",
  "messages": [
    {
      "id": "1091177",
      "postDate": "11/25/2020 20:34:32",
      "content": "<p>Hello! Can anybody helps with my problem?</p>\n<p>I trained my simple model in this training notebook <a href=\"https://www.kaggle.com/khlevnov/efficientnetb0-training\" target=\"_blank\">EfficientNetB0 training</a> and take 0.800+ accuracy with hold-out CV. Results of the last epoch below (and you can see it in notebook):</p>\n<pre><code>Epoch 10/10\n1204/1204 - 236s - loss: 0.4054 - accuracy: 0.8602 - val_loss: 0.4383 - val_accuracy: 0.8570\n</code></pre>\n<p>Then I saved model in .h5 file and created <a href=\"https://www.kaggle.com/khlevnov/cassava-preptrained-model-2\" target=\"_blank\">public dataset</a> with notebook output (and saved model also).</p>\n<p>For submission I created separate notebook <a href=\"https://www.kaggle.com/khlevnov/efficientnetb0-submission\" target=\"_blank\">EfficientNetB0 submission</a> only for prediction using saved model from my freshly created dataset. But after submission… Boom, I have only 0.400+ accuracy on public leaderboard.</p>\n<pre><code>Succeeded\n0.417\n</code></pre>\n<p>Seems my public score divided by two(<br>\nWhere should I search my dumb error?</p>\n<p>Will be happy to listen any suggestions, thanks.</p>\n<p>UPDATE:<br>\nAs <a href=\"https://www.kaggle.com/richardepstein\" target=\"_blank\">@richardepstein</a> said, there are different order between <code>image_id</code> and <code>label</code> columns in my submission. Thanks to all.</p>",
      "rawMarkdown": "Hello! Can anybody helps with my problem?\n\nI trained my simple model in this training notebook [EfficientNetB0 training](https://www.kaggle.com/khlevnov/efficientnetb0-training) and take 0.800+ accuracy with hold-out CV. Results of the last epoch below (and you can see it in notebook):\n```\nEpoch 10/10\n1204/1204 - 236s - loss: 0.4054 - accuracy: 0.8602 - val_loss: 0.4383 - val_accuracy: 0.8570\n```\nThen I saved model in .h5 file and created [public dataset](https://www.kaggle.com/khlevnov/cassava-preptrained-model-2) with notebook output (and saved model also).\n\nFor submission I created separate notebook [EfficientNetB0 submission](https://www.kaggle.com/khlevnov/efficientnetb0-submission) only for prediction using saved model from my freshly created dataset. But after submission... Boom, I have only 0.400+ accuracy on public leaderboard.\n```\nSucceeded\n0.417\n```\nSeems my public score divided by two(\nWhere should I search my dumb error?\n\nWill be happy to listen any suggestions, thanks.\n\nUPDATE:\nAs @richardepstein said, there are different order between `image_id` and `label` columns in my submission. Thanks to all.",
      "votes": null
    },
    {
      "id": "1091182",
      "postDate": "11/25/2020 20:41:33",
      "content": "<p>The data in tfrec format is a label error</p>",
      "rawMarkdown": "The data in tfrec format is a label error",
      "votes": null
    },
    {
      "id": "1091201",
      "postDate": "11/25/2020 21:08:45",
      "content": "<p>But I use images from <code>train_images</code>😦</p>",
      "rawMarkdown": "But I use images from `train_images`😦",
      "votes": null
    },
    {
      "id": "1091205",
      "postDate": "11/25/2020 21:18:39",
      "content": "<p>Do you have a local implementation of inference with model loading from a file and making predictions on hold out data?</p>",
      "rawMarkdown": "Do you have a local implementation of inference with model loading from a file and making predictions on hold out data?",
      "votes": null
    },
    {
      "id": "1091208",
      "postDate": "11/25/2020 21:25:55",
      "content": "<p>Is the order of images in sample submission.csv and the order returned by the directory the same? </p>",
      "rawMarkdown": "Is the order of images in sample submission.csv and the order returned by the directory the same?",
      "votes": null
    },
    {
      "id": "1091224",
      "postDate": "11/25/2020 21:50:58",
      "content": "<p>Thank you! My bad, I forgot to add shuffle arg</p>\n<pre><code>.flow_from_dataframe(\n    shuffle=False,\n    ...\n)\n</code></pre>\n<p>Now LB score updated to 0.852 for this baseline 😊</p>",
      "rawMarkdown": "Thank you! My bad, I forgot to add shuffle arg\n```\n.flow_from_dataframe(\n    shuffle=False,\n    ...\n)\n```\nNow LB score updated to 0.852 for this baseline 😊",
      "votes": null
    },
    {
      "id": "1091234",
      "postDate": "11/25/2020 22:06:34",
      "content": "<p>I've also tested tfrec format data, about the same as you.</p>",
      "rawMarkdown": "I've also tested tfrec format data, about the same as you.",
      "votes": null
    },
    {
      "id": "1091280",
      "postDate": "11/25/2020 22:59:51",
      "content": "<p>Thanks for advice! I hope it can helps me to prevent same issues in future.<br>\nI implemented simple solution and saw a difference on my local machine.</p>\n<pre><code>score = accuracy_score(hold_out_df[['label']], submission_df[['label']])\nprint(f'Score {score}')\n</code></pre>\n<p>Problem was when I shuffled labels in <code>.flow_from_dataframe</code> for prediction.</p>",
      "rawMarkdown": "Thanks for advice! I hope it can helps me to prevent same issues in future.\nI implemented simple solution and saw a difference on my local machine.\n```\nscore = accuracy_score(hold_out_df[['label']], submission_df[['label']])\nprint(f'Score {score}')\n```\nProblem was when I shuffled labels in `.flow_from_dataframe` for prediction.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1091182,
      "author_name": "zhangeng",
      "author_url": "",
      "post_date": "11/25/2020 20:41:33",
      "content": "<p>The data in tfrec format is a label error</p>",
      "votes": null,
      "replies": [
        {
          "id": 1091201,
          "author_name": "khlevnov",
          "author_url": "",
          "post_date": "11/25/2020 21:08:45",
          "content": "<p>But I use images from <code>train_images</code>😦</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1091234,
          "author_name": "zhangeng",
          "author_url": "",
          "post_date": "11/25/2020 22:06:34",
          "content": "<p>I've also tested tfrec format data, about the same as you.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1091205,
      "author_name": "opanichev",
      "author_url": "",
      "post_date": "11/25/2020 21:18:39",
      "content": "<p>Do you have a local implementation of inference with model loading from a file and making predictions on hold out data?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1091280,
          "author_name": "khlevnov",
          "author_url": "",
          "post_date": "11/25/2020 22:59:51",
          "content": "<p>Thanks for advice! I hope it can helps me to prevent same issues in future.<br>\nI implemented simple solution and saw a difference on my local machine.</p>\n<pre><code>score = accuracy_score(hold_out_df[['label']], submission_df[['label']])\nprint(f'Score {score}')\n</code></pre>\n<p>Problem was when I shuffled labels in <code>.flow_from_dataframe</code> for prediction.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1091208,
      "author_name": "richardepstein",
      "author_url": "",
      "post_date": "11/25/2020 21:25:55",
      "content": "<p>Is the order of images in sample submission.csv and the order returned by the directory the same? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1091224,
          "author_name": "khlevnov",
          "author_url": "",
          "post_date": "11/25/2020 21:50:58",
          "content": "<p>Thank you! My bad, I forgot to add shuffle arg</p>\n<pre><code>.flow_from_dataframe(\n    shuffle=False,\n    ...\n)\n</code></pre>\n<p>Now LB score updated to 0.852 for this baseline 😊</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1091177": "Hello! Can anybody helps with my problem?\n\nI trained my simple model in this training notebook [EfficientNetB0 training](https://www.kaggle.com/khlevnov/efficientnetb0-training) and take 0.800+ accuracy with hold-out CV. Results of the last epoch below (and you can see it in notebook):\n```\nEpoch 10/10\n1204/1204 - 236s - loss: 0.4054 - accuracy: 0.8602 - val_loss: 0.4383 - val_accuracy: 0.8570\n```\nThen I saved model in .h5 file and created [public dataset](https://www.kaggle.com/khlevnov/cassava-preptrained-model-2) with notebook output (and saved model also).\n\nFor submission I created separate notebook [EfficientNetB0 submission](https://www.kaggle.com/khlevnov/efficientnetb0-submission) only for prediction using saved model from my freshly created dataset. But after submission... Boom, I have only 0.400+ accuracy on public leaderboard.\n```\nSucceeded\n0.417\n```\nSeems my public score divided by two(\nWhere should I search my dumb error?\n\nWill be happy to listen any suggestions, thanks.\n\nUPDATE:\nAs @richardepstein said, there are different order between `image_id` and `label` columns in my submission. Thanks to all.",
    "1091182": "The data in tfrec format is a label error",
    "1091201": "But I use images from `train_images`😦",
    "1091205": "Do you have a local implementation of inference with model loading from a file and making predictions on hold out data?",
    "1091208": "Is the order of images in sample submission.csv and the order returned by the directory the same?",
    "1091224": "Thank you! My bad, I forgot to add shuffle arg\n```\n.flow_from_dataframe(\n    shuffle=False,\n    ...\n)\n```\nNow LB score updated to 0.852 for this baseline 😊",
    "1091234": "I've also tested tfrec format data, about the same as you.",
    "1091280": "Thanks for advice! I hope it can helps me to prevent same issues in future.\nI implemented simple solution and saw a difference on my local machine.\n```\nscore = accuracy_score(hold_out_df[['label']], submission_df[['label']])\nprint(f'Score {score}')\n```\nProblem was when I shuffled labels in `.flow_from_dataframe` for prediction."
  },
  "source": "meta"
}