{
  "id": 303984,
  "title": "Train vs Inference notebooks",
  "url": "/competitions/tensorflow-great-barrier-reef/discussion/303984",
  "author_name": "",
  "post_date": "2022-01-30T13:01:05.438577500Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi,</p>\n<p>A newbie general question - I commonly see people here splitting their training and inference code into two separate notebooks, but I'm used to seeing the same notebook first training and then predicting on the test.<br>\nIs the reason for this split a matter of the competition rules/restrictive time budget? Is it about privacy? Is it just a matter of better software engineering by decoupling both concepts?</p>\n<p>Is it valid to have a submission consisting of a notebook that only loads a pretrained network uploaded by someone else, then doing inference on the test set?</p>",
  "messages": [
    {
      "id": "1669184",
      "postDate": "01/30/2022 13:01:05",
      "content": "<p>Hi,</p>\n<p>A newbie general question - I commonly see people here splitting their training and inference code into two separate notebooks, but I'm used to seeing the same notebook first training and then predicting on the test.<br>\nIs the reason for this split a matter of the competition rules/restrictive time budget? Is it about privacy? Is it just a matter of better software engineering by decoupling both concepts?</p>\n<p>Is it valid to have a submission consisting of a notebook that only loads a pretrained network uploaded by someone else, then doing inference on the test set?</p>",
      "rawMarkdown": "Hi,\n\nA newbie general question - I commonly see people here splitting their training and inference code into two separate notebooks, but I'm used to seeing the same notebook first training and then predicting on the test.\nIs the reason for this split a matter of the competition rules/restrictive time budget? Is it about privacy? Is it just a matter of better software engineering by decoupling both concepts?\n\nIs it valid to have a submission consisting of a notebook that only loads a pretrained network uploaded by someone else, then doing inference on the test set?",
      "votes": null
    },
    {
      "id": "1669199",
      "postDate": "01/30/2022 13:26:51",
      "content": "<p>It is to reduce the test time because if you will put train and inference in your notebook, you will waste many hours of free GPU. So, Kagglers usually split their training and inference notebooks.</p>",
      "rawMarkdown": "It is to reduce the test time because if you will put train and inference in your notebook, you will waste many hours of free GPU. So, Kagglers usually split their training and inference notebooks.",
      "votes": null
    },
    {
      "id": "1674345",
      "postDate": "02/03/2022 12:12:20",
      "content": "<p>Thanks!<br>\nSo you usually run training on GPU, and then inference on CPU? In this competition, I thought both train and test datasets were large, so GPU would be needed for both - is that not the case? Or does the test API being sequential remove the need for a GPU at inference?</p>",
      "rawMarkdown": "Thanks!\nSo you usually run training on GPU, and then inference on CPU? In this competition, I thought both train and test datasets were large, so GPU would be needed for both - is that not the case? Or does the test API being sequential remove the need for a GPU at inference?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1669199,
      "author_name": "kocha1",
      "author_url": "",
      "post_date": "01/30/2022 13:26:51",
      "content": "<p>It is to reduce the test time because if you will put train and inference in your notebook, you will waste many hours of free GPU. So, Kagglers usually split their training and inference notebooks.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1674345,
          "author_name": "yiftachbeer",
          "author_url": "",
          "post_date": "02/03/2022 12:12:20",
          "content": "<p>Thanks!<br>\nSo you usually run training on GPU, and then inference on CPU? In this competition, I thought both train and test datasets were large, so GPU would be needed for both - is that not the case? Or does the test API being sequential remove the need for a GPU at inference?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1669184": "Hi,\n\nA newbie general question - I commonly see people here splitting their training and inference code into two separate notebooks, but I'm used to seeing the same notebook first training and then predicting on the test.\nIs the reason for this split a matter of the competition rules/restrictive time budget? Is it about privacy? Is it just a matter of better software engineering by decoupling both concepts?\n\nIs it valid to have a submission consisting of a notebook that only loads a pretrained network uploaded by someone else, then doing inference on the test set?",
    "1669199": "It is to reduce the test time because if you will put train and inference in your notebook, you will waste many hours of free GPU. So, Kagglers usually split their training and inference notebooks.",
    "1674345": "Thanks!\nSo you usually run training on GPU, and then inference on CPU? In this competition, I thought both train and test datasets were large, so GPU would be needed for both - is that not the case? Or does the test API being sequential remove the need for a GPU at inference?"
  },
  "source": "meta"
}