{
  "id": 198996,
  "title": "A question about TPU",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/198996",
  "author_name": "",
  "post_date": "2020-11-24T00:46:03.215863700Z",
  "votes": 1,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Hello!  When I train a model on TPU, the cv score is 87.But when I inference on GPU,the LB score is 40-50.What is wrong?</p>",
  "messages": [
    {
      "id": "1088811",
      "postDate": "11/24/2020 00:46:03",
      "content": "<p>Hello!  When I train a model on TPU, the cv score is 87.But when I inference on GPU,the LB score is 40-50.What is wrong?</p>",
      "rawMarkdown": "Hello!  When I train a model on TPU, the cv score is 87.But when I inference on GPU,the LB score is 40-50.What is wrong?",
      "votes": null
    },
    {
      "id": "1088825",
      "postDate": "11/24/2020 00:57:32",
      "content": "<p>Not sure exactly what you mean by \"LB score is 40-50\". Do you mean 40th to 50th place? I see your current LB score is 21st (0.900). Since the first 71 places are mostly 0.900, with a hidden fourth decimal place, a score of 0.900 could seem like anywhere from 16th place to 71st place. I think the scores are sorted by more decimal points than they are displayed.</p>\n<p>Is your question why your CV is 0.87 but your LB score is 0.90?</p>\n<p>I think that is just the difference between the validation data and the test data.</p>\n<p>If you look at the \"CV vs LB\" thread, you'll see a lot of example of differences between CV and LB.</p>",
      "rawMarkdown": "Not sure exactly what you mean by \"LB score is 40-50\". Do you mean 40th to 50th place? I see your current LB score is 21st (0.900). Since the first 71 places are mostly 0.900, with a hidden fourth decimal place, a score of 0.900 could seem like anywhere from 16th place to 71st place. I think the scores are sorted by more decimal points than they are displayed.\n\nIs your question why your CV is 0.87 but your LB score is 0.90?\n\nI think that is just the difference between the validation data and the test data.\n\nIf you look at the \"CV vs LB\" thread, you'll see a lot of example of differences between CV and LB.",
      "votes": null
    },
    {
      "id": "1088873",
      "postDate": "11/24/2020 02:31:57",
      "content": "<p>The CV score is high does not mean your LB score/ranking will also be high. Your model has a chance to be overfit, and if you see worse LB score than your CV it is probably so.</p>",
      "rawMarkdown": "The CV score is high does not mean your LB score/ranking will also be high. Your model has a chance to be overfit, and if you see worse LB score than your CV it is probably so.",
      "votes": null
    },
    {
      "id": "1088926",
      "postDate": "11/24/2020 04:17:02",
      "content": "<p>But is much different，88 cv --&gt; 56LB.I think it is not overfit.</p>",
      "rawMarkdown": "But is much different，88 cv --> 56LB.I think it is not overfit.",
      "votes": null
    },
    {
      "id": "1088927",
      "postDate": "11/24/2020 04:18:54",
      "content": "<p>Thanks, but will the device cause a great change? Such as trainging on TPU,but testing on GPU and TPU.</p>",
      "rawMarkdown": "Thanks, but will the device cause a great change? Such as trainging on TPU,but testing on GPU and TPU.",
      "votes": null
    },
    {
      "id": "1088928",
      "postDate": "11/24/2020 04:21:16",
      "content": "<p>Are you using the provided TFRecords for the test data? I think there is a problem with the train TFRecords and possibly the test records also. Try your algorithm with the jpegs.</p>",
      "rawMarkdown": "Are you using the provided TFRecords for the test data? I think there is a problem with the train TFRecords and possibly the test records also. Try your algorithm with the jpegs.",
      "votes": null
    },
    {
      "id": "1088936",
      "postDate": "11/24/2020 04:30:06",
      "content": "<p>The results for TPU and GPU for inference should be the same. </p>\n<p>For training you could have a slight difference as neither are deterministic unless you use certain tensorflow settings. </p>",
      "rawMarkdown": "The results for TPU and GPU for inference should be the same. \n\nFor training you could have a slight difference as neither are deterministic unless you use certain tensorflow settings.",
      "votes": null
    },
    {
      "id": "1088967",
      "postDate": "11/24/2020 05:09:26",
      "content": "<p>Oh,thanks,I dont change that.I found a discussion that there are some question about them.</p>",
      "rawMarkdown": "Oh,thanks,I dont change that.I found a discussion that there are some question about them.",
      "votes": null
    },
    {
      "id": "1088968",
      "postDate": "11/24/2020 05:11:04",
      "content": "<p>It does not mean it is totally not overfit, but generally when CV &lt; LB we assume it is overfit. Btw if you are using TFRecords there are discussions about it is wrongly labelled, and the staff are looking into it.</p>",
      "rawMarkdown": "It does not mean it is totally not overfit, but generally when CV < LB we assume it is overfit. Btw if you are using TFRecords there are discussions about it is wrongly labelled, and the staff are looking into it.",
      "votes": null
    },
    {
      "id": "1088991",
      "postDate": "11/24/2020 05:51:18",
      "content": "<p>Thanks I think so.</p>",
      "rawMarkdown": "Thanks I think so.",
      "votes": null
    },
    {
      "id": "1089064",
      "postDate": "11/24/2020 07:18:59",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/zekunn\" target=\"_blank\">@zekunn</a>  I faced the same problem, what I was doing is, was using open cv's imread function or reading test images, I just replaced that image.open functionality of PIL,</p>\n<p>In other words this might be a problem of image loading logic which explains the discrepancy in scores.</p>\n<p>Thanks..</p>",
      "rawMarkdown": "Hi @zekunn  I faced the same problem, what I was doing is, was using open cv's imread function or reading test images, I just replaced that image.open functionality of PIL,\n\nIn other words this might be a problem of image loading logic which explains the discrepancy in scores.\n\nThanks..",
      "votes": null
    },
    {
      "id": "1089670",
      "postDate": "11/24/2020 17:08:29",
      "content": "<p>You are probably running into the same issues <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/198272\" target=\"_blank\">identified here</a>. We are working on a fix to the records. Please reference that thread for updates.</p>",
      "rawMarkdown": "You are probably running into the same issues [identified here](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/198272). We are working on a fix to the records. Please reference that thread for updates.",
      "votes": null
    },
    {
      "id": "1093787",
      "postDate": "11/28/2020 02:56:43",
      "content": "<p>Thanks!…</p>",
      "rawMarkdown": "Thanks!...",
      "votes": null
    },
    {
      "id": "1093788",
      "postDate": "11/28/2020 02:57:17",
      "content": "<p>It may caused by the wrong tfrecords.</p>",
      "rawMarkdown": "It may caused by the wrong tfrecords.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1088825,
      "author_name": "richardepstein",
      "author_url": "",
      "post_date": "11/24/2020 00:57:32",
      "content": "<p>Not sure exactly what you mean by \"LB score is 40-50\". Do you mean 40th to 50th place? I see your current LB score is 21st (0.900). Since the first 71 places are mostly 0.900, with a hidden fourth decimal place, a score of 0.900 could seem like anywhere from 16th place to 71st place. I think the scores are sorted by more decimal points than they are displayed.</p>\n<p>Is your question why your CV is 0.87 but your LB score is 0.90?</p>\n<p>I think that is just the difference between the validation data and the test data.</p>\n<p>If you look at the \"CV vs LB\" thread, you'll see a lot of example of differences between CV and LB.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1088927,
          "author_name": "zekunn",
          "author_url": "",
          "post_date": "11/24/2020 04:18:54",
          "content": "<p>Thanks, but will the device cause a great change? Such as trainging on TPU,but testing on GPU and TPU.</p>",
          "votes": null,
          "replies": [
            {
              "id": 1088928,
              "author_name": "richardepstein",
              "author_url": "",
              "post_date": "11/24/2020 04:21:16",
              "content": "<p>Are you using the provided TFRecords for the test data? I think there is a problem with the train TFRecords and possibly the test records also. Try your algorithm with the jpegs.</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 1088936,
              "author_name": "richardepstein",
              "author_url": "",
              "post_date": "11/24/2020 04:30:06",
              "content": "<p>The results for TPU and GPU for inference should be the same. </p>\n<p>For training you could have a slight difference as neither are deterministic unless you use certain tensorflow settings. </p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 1088967,
              "author_name": "zekunn",
              "author_url": "",
              "post_date": "11/24/2020 05:09:26",
              "content": "<p>Oh,thanks,I dont change that.I found a discussion that there are some question about them.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 1088873,
      "author_name": "aeryss",
      "author_url": "",
      "post_date": "11/24/2020 02:31:57",
      "content": "<p>The CV score is high does not mean your LB score/ranking will also be high. Your model has a chance to be overfit, and if you see worse LB score than your CV it is probably so.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1088926,
          "author_name": "zekunn",
          "author_url": "",
          "post_date": "11/24/2020 04:17:02",
          "content": "<p>But is much different，88 cv --&gt; 56LB.I think it is not overfit.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1088968,
          "author_name": "aeryss",
          "author_url": "",
          "post_date": "11/24/2020 05:11:04",
          "content": "<p>It does not mean it is totally not overfit, but generally when CV &lt; LB we assume it is overfit. Btw if you are using TFRecords there are discussions about it is wrongly labelled, and the staff are looking into it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1088991,
          "author_name": "zekunn",
          "author_url": "",
          "post_date": "11/24/2020 05:51:18",
          "content": "<p>Thanks I think so.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1089064,
      "author_name": "vanvalkenberg",
      "author_url": "",
      "post_date": "11/24/2020 07:18:59",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/zekunn\" target=\"_blank\">@zekunn</a>  I faced the same problem, what I was doing is, was using open cv's imread function or reading test images, I just replaced that image.open functionality of PIL,</p>\n<p>In other words this might be a problem of image loading logic which explains the discrepancy in scores.</p>\n<p>Thanks..</p>",
      "votes": null,
      "replies": [
        {
          "id": 1093788,
          "author_name": "zekunn",
          "author_url": "",
          "post_date": "11/28/2020 02:57:17",
          "content": "<p>It may caused by the wrong tfrecords.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1089670,
      "author_name": "juliaelliott",
      "author_url": "",
      "post_date": "11/24/2020 17:08:29",
      "content": "<p>You are probably running into the same issues <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/198272\" target=\"_blank\">identified here</a>. We are working on a fix to the records. Please reference that thread for updates.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1093787,
          "author_name": "zekunn",
          "author_url": "",
          "post_date": "11/28/2020 02:56:43",
          "content": "<p>Thanks!…</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1088811": "Hello!  When I train a model on TPU, the cv score is 87.But when I inference on GPU,the LB score is 40-50.What is wrong?",
    "1088825": "Not sure exactly what you mean by \"LB score is 40-50\". Do you mean 40th to 50th place? I see your current LB score is 21st (0.900). Since the first 71 places are mostly 0.900, with a hidden fourth decimal place, a score of 0.900 could seem like anywhere from 16th place to 71st place. I think the scores are sorted by more decimal points than they are displayed.\n\nIs your question why your CV is 0.87 but your LB score is 0.90?\n\nI think that is just the difference between the validation data and the test data.\n\nIf you look at the \"CV vs LB\" thread, you'll see a lot of example of differences between CV and LB.",
    "1088873": "The CV score is high does not mean your LB score/ranking will also be high. Your model has a chance to be overfit, and if you see worse LB score than your CV it is probably so.",
    "1088926": "But is much different，88 cv --> 56LB.I think it is not overfit.",
    "1088927": "Thanks, but will the device cause a great change? Such as trainging on TPU,but testing on GPU and TPU.",
    "1088928": "Are you using the provided TFRecords for the test data? I think there is a problem with the train TFRecords and possibly the test records also. Try your algorithm with the jpegs.",
    "1088936": "The results for TPU and GPU for inference should be the same. \n\nFor training you could have a slight difference as neither are deterministic unless you use certain tensorflow settings.",
    "1088967": "Oh,thanks,I dont change that.I found a discussion that there are some question about them.",
    "1088968": "It does not mean it is totally not overfit, but generally when CV < LB we assume it is overfit. Btw if you are using TFRecords there are discussions about it is wrongly labelled, and the staff are looking into it.",
    "1088991": "Thanks I think so.",
    "1089064": "Hi @zekunn  I faced the same problem, what I was doing is, was using open cv's imread function or reading test images, I just replaced that image.open functionality of PIL,\n\nIn other words this might be a problem of image loading logic which explains the discrepancy in scores.\n\nThanks..",
    "1089670": "You are probably running into the same issues [identified here](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/198272). We are working on a fix to the records. Please reference that thread for updates.",
    "1093787": "Thanks!...",
    "1093788": "It may caused by the wrong tfrecords."
  },
  "source": "meta"
}