{
  "id": 336429,
  "title": "A big gap between online and offline training",
  "url": "/competitions/hubmap-organ-segmentation/discussion/336429",
  "author_name": "",
  "post_date": "2022-07-11T05:59:41.411607400Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>This is my first Kaggle competition, I am quite new with using Kaggle kernel.<br>\nI have a question that there is a big gap between online and offline training.<br>\nThe ways of my online and offline training are as follow.<br>\n(a)online: train the model on Kaggle kernel, then inference test images and output the submission.csv.<br>\n(b)offline: upload the model locally trained well on my own machines through AddData, on kaggle kernel I only load model and do inference.</p>\n<p>The scores of online training is better than the other one. <br>\nDoes the training dataset also change after I submit the notebook?</p>",
  "messages": [
    {
      "id": "1851273",
      "postDate": "07/11/2022 05:59:41",
      "content": "<p>This is my first Kaggle competition, I am quite new with using Kaggle kernel.<br>\nI have a question that there is a big gap between online and offline training.<br>\nThe ways of my online and offline training are as follow.<br>\n(a)online: train the model on Kaggle kernel, then inference test images and output the submission.csv.<br>\n(b)offline: upload the model locally trained well on my own machines through AddData, on kaggle kernel I only load model and do inference.</p>\n<p>The scores of online training is better than the other one. <br>\nDoes the training dataset also change after I submit the notebook?</p>",
      "rawMarkdown": "This is my first Kaggle competition, I am quite new with using Kaggle kernel.\nI have a question that there is a big gap between online and offline training.\nThe ways of my online and offline training are as follow.\n(a)online: train the model on Kaggle kernel, then inference test images and output the submission.csv.\n(b)offline: upload the model locally trained well on my own machines through AddData, on kaggle kernel I only load model and do inference.\n\nThe scores of online training is better than the other one. \nDoes the training dataset also change after I submit the notebook?",
      "votes": null
    },
    {
      "id": "1851341",
      "postDate": "07/11/2022 07:12:38",
      "content": "<p>Hi <br>\nThe training data doesn't change after the submission.<br>\nI do train online (but not on kaggle) I use colab and sometimes jarvislabs.<br>\nThe gap hmm , maybe due different environments and libs versions, or you are loading the model with different settings … </p>",
      "rawMarkdown": "Hi \nThe training data doesn't change after the submission.\nI do train online (but not on kaggle) I use colab and sometimes jarvislabs.\nThe gap hmm , maybe due different environments and libs versions, or you are loading the model with different settings ...",
      "votes": null
    },
    {
      "id": "1851437",
      "postDate": "07/11/2022 08:39:08",
      "content": "<p>thanks! it is helpful.</p>",
      "rawMarkdown": "thanks! it is helpful.",
      "votes": null
    },
    {
      "id": "1851739",
      "postDate": "07/11/2022 13:44:13",
      "content": "<p>There could be a few reasons why this could happen.</p>\n<p>1) The train test split logic uses a random seed native to the environment. Use a train test split logic which is independent of the environment.<br>\n2) Use fixed random seeds for training and sampling in both your environments. Different seeds will sample different data and the outcome will be different.<br>\n3) Update your local environment with the one you are using in kaggle. You can use <code>pip3 freeze &gt; requirements.txt</code> in kaggle and use this file to update your required libraries in your local environment.</p>",
      "rawMarkdown": "There could be a few reasons why this could happen.\n\n1) The train test split logic uses a random seed native to the environment. Use a train test split logic which is independent of the environment.\n2) Use fixed random seeds for training and sampling in both your environments. Different seeds will sample different data and the outcome will be different.\n3) Update your local environment with the one you are using in kaggle. You can use `pip3 freeze > requirements.txt` in kaggle and use this file to update your required libraries in your local environment.",
      "votes": null
    },
    {
      "id": "1852636",
      "postDate": "07/12/2022 07:50:17",
      "content": "<p>Thanks！ 😊 ☺️ 🤩 </p>",
      "rawMarkdown": "Thanks！ 😊 ☺️ 🤩",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1851341,
      "author_name": "asalhi",
      "author_url": "",
      "post_date": "07/11/2022 07:12:38",
      "content": "<p>Hi <br>\nThe training data doesn't change after the submission.<br>\nI do train online (but not on kaggle) I use colab and sometimes jarvislabs.<br>\nThe gap hmm , maybe due different environments and libs versions, or you are loading the model with different settings … </p>",
      "votes": null,
      "replies": [
        {
          "id": 1851437,
          "author_name": "paijon",
          "author_url": "",
          "post_date": "07/11/2022 08:39:08",
          "content": "<p>thanks! it is helpful.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1851739,
      "author_name": "ezzzio",
      "author_url": "",
      "post_date": "07/11/2022 13:44:13",
      "content": "<p>There could be a few reasons why this could happen.</p>\n<p>1) The train test split logic uses a random seed native to the environment. Use a train test split logic which is independent of the environment.<br>\n2) Use fixed random seeds for training and sampling in both your environments. Different seeds will sample different data and the outcome will be different.<br>\n3) Update your local environment with the one you are using in kaggle. You can use <code>pip3 freeze &gt; requirements.txt</code> in kaggle and use this file to update your required libraries in your local environment.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1852636,
          "author_name": "paijon",
          "author_url": "",
          "post_date": "07/12/2022 07:50:17",
          "content": "<p>Thanks！ 😊 ☺️ 🤩 </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1851273": "This is my first Kaggle competition, I am quite new with using Kaggle kernel.\nI have a question that there is a big gap between online and offline training.\nThe ways of my online and offline training are as follow.\n(a)online: train the model on Kaggle kernel, then inference test images and output the submission.csv.\n(b)offline: upload the model locally trained well on my own machines through AddData, on kaggle kernel I only load model and do inference.\n\nThe scores of online training is better than the other one. \nDoes the training dataset also change after I submit the notebook?",
    "1851341": "Hi \nThe training data doesn't change after the submission.\nI do train online (but not on kaggle) I use colab and sometimes jarvislabs.\nThe gap hmm , maybe due different environments and libs versions, or you are loading the model with different settings ...",
    "1851437": "thanks! it is helpful.",
    "1851739": "There could be a few reasons why this could happen.\n\n1) The train test split logic uses a random seed native to the environment. Use a train test split logic which is independent of the environment.\n2) Use fixed random seeds for training and sampling in both your environments. Different seeds will sample different data and the outcome will be different.\n3) Update your local environment with the one you are using in kaggle. You can use `pip3 freeze > requirements.txt` in kaggle and use this file to update your required libraries in your local environment.",
    "1852636": "Thanks！ 😊 ☺️ 🤩"
  },
  "source": "meta"
}