{
  "id": 238104,
  "title": "Training a model outside of Kaggle and how to evaluate on test set then",
  "url": "/competitions/birdclef-2021/discussion/238104",
  "author_name": "",
  "post_date": "2021-05-11T08:30:09.568930300Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>According to the notebook <a href=\"https://www.kaggle.com/stefankahl/birdclef2021-sample-submission\" target=\"_blank\">BirdClef2021 Sample Submission</a>, we can only evaluate on the test set, if we make a notebook and submit the notebook. Then the notebook will have access to the hidden test set files. So, in order to make a submission that will land you somewhere on the Leaderboard, this is a necessity, correct? </p>\n<p>My question is then: I would suppose that most of the top-scoring teams on the leaderboard, train outside of Kaggle, using the computing resources available at their research facility and possibly across several days. Something that's not possible with a Kaggle notebook. </p>\n<p>So, if training has been done off-Kaggle, am I correct then in assuming that one would need to make a notebook here on Kaggle, for evaluating on the test set, <strong><em>uploading the model that was trained outside of Kaggle</em></strong>, so that it appears under \"Data\" and then load and use it to do predictions? </p>\n<p>Is this indeed the approach of all the people training outside of a Kaggle notebook? </p>\n<p>Thanks :-) </p>",
  "messages": [
    {
      "id": "1301767",
      "postDate": "05/11/2021 08:30:09",
      "content": "<p>According to the notebook <a href=\"https://www.kaggle.com/stefankahl/birdclef2021-sample-submission\" target=\"_blank\">BirdClef2021 Sample Submission</a>, we can only evaluate on the test set, if we make a notebook and submit the notebook. Then the notebook will have access to the hidden test set files. So, in order to make a submission that will land you somewhere on the Leaderboard, this is a necessity, correct? </p>\n<p>My question is then: I would suppose that most of the top-scoring teams on the leaderboard, train outside of Kaggle, using the computing resources available at their research facility and possibly across several days. Something that's not possible with a Kaggle notebook. </p>\n<p>So, if training has been done off-Kaggle, am I correct then in assuming that one would need to make a notebook here on Kaggle, for evaluating on the test set, <strong><em>uploading the model that was trained outside of Kaggle</em></strong>, so that it appears under \"Data\" and then load and use it to do predictions? </p>\n<p>Is this indeed the approach of all the people training outside of a Kaggle notebook? </p>\n<p>Thanks :-) </p>",
      "rawMarkdown": "According to the notebook [BirdClef2021 Sample Submission](https://www.kaggle.com/stefankahl/birdclef2021-sample-submission), we can only evaluate on the test set, if we make a notebook and submit the notebook. Then the notebook will have access to the hidden test set files. So, in order to make a submission that will land you somewhere on the Leaderboard, this is a necessity, correct? \n\nMy question is then: I would suppose that most of the top-scoring teams on the leaderboard, train outside of Kaggle, using the computing resources available at their research facility and possibly across several days. Something that's not possible with a Kaggle notebook. \n\nSo, if training has been done off-Kaggle, am I correct then in assuming that one would need to make a notebook here on Kaggle, for evaluating on the test set, ***uploading the model that was trained outside of Kaggle***, so that it appears under \"Data\" and then load and use it to do predictions? \n\nIs this indeed the approach of all the people training outside of a Kaggle notebook? \n\nThanks :-)",
      "votes": null
    },
    {
      "id": "1301792",
      "postDate": "05/11/2021 08:45:57",
      "content": "<p>You are correct.  People who train models outside Kaggle upload their model (code and weights) in some kaggle dataset, then add this dataset to the inferencing notebook. You can probably find examples in the public notebooks for this competition.</p>",
      "rawMarkdown": "You are correct.  People who train models outside Kaggle upload their model (code and weights) in some kaggle dataset, then add this dataset to the inferencing notebook. You can probably find examples in the public notebooks for this competition.",
      "votes": null
    },
    {
      "id": "1303573",
      "postDate": "05/12/2021 06:24:14",
      "content": "<p>Here's a snippet for a resnet-like model in pytorch that may help you.<br>\nYou will need a dataset <code>predefinedmodels</code> and <code>checkpoints</code>.</p>\n<pre><code>def init_model(repo, model, checkpoint):\n    model = torch.hub.load(repo, model, pretrained=False, source='local')\n    n_features = model.fc.in_features\n    model.fc = torch.nn.Linear(in_features=n_features, out_features=n_labels, bias=True)\n    cp = torch.load(f\"../input/checkpoints/{checkpoint}.pt\", map_location=torch.device(device))\n    model.load_state_dict(cp['model_state_dict'])\n    model = model.to(device).eval()\n    return model\n...\nmodel = init_model('../input/predefinedmodels/pytorch_vision_v0.9.0', 'resnet18', 'resnet18-20210510-204737-latest')\n</code></pre>",
      "rawMarkdown": "Here's a snippet for a resnet-like model in pytorch that may help you.\nYou will need a dataset `predefinedmodels` and `checkpoints`.\n\n```\ndef init_model(repo, model, checkpoint):\n    model = torch.hub.load(repo, model, pretrained=False, source='local')\n    n_features = model.fc.in_features\n    model.fc = torch.nn.Linear(in_features=n_features, out_features=n_labels, bias=True)\n    cp = torch.load(f\"../input/checkpoints/{checkpoint}.pt\", map_location=torch.device(device))\n    model.load_state_dict(cp['model_state_dict'])\n    model = model.to(device).eval()\n    return model\n...\nmodel = init_model('../input/predefinedmodels/pytorch_vision_v0.9.0', 'resnet18', 'resnet18-20210510-204737-latest')\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1301792,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "05/11/2021 08:45:57",
      "content": "<p>You are correct.  People who train models outside Kaggle upload their model (code and weights) in some kaggle dataset, then add this dataset to the inferencing notebook. You can probably find examples in the public notebooks for this competition.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1303573,
      "author_name": "botkop",
      "author_url": "",
      "post_date": "05/12/2021 06:24:14",
      "content": "<p>Here's a snippet for a resnet-like model in pytorch that may help you.<br>\nYou will need a dataset <code>predefinedmodels</code> and <code>checkpoints</code>.</p>\n<pre><code>def init_model(repo, model, checkpoint):\n    model = torch.hub.load(repo, model, pretrained=False, source='local')\n    n_features = model.fc.in_features\n    model.fc = torch.nn.Linear(in_features=n_features, out_features=n_labels, bias=True)\n    cp = torch.load(f\"../input/checkpoints/{checkpoint}.pt\", map_location=torch.device(device))\n    model.load_state_dict(cp['model_state_dict'])\n    model = model.to(device).eval()\n    return model\n...\nmodel = init_model('../input/predefinedmodels/pytorch_vision_v0.9.0', 'resnet18', 'resnet18-20210510-204737-latest')\n</code></pre>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1301767": "According to the notebook [BirdClef2021 Sample Submission](https://www.kaggle.com/stefankahl/birdclef2021-sample-submission), we can only evaluate on the test set, if we make a notebook and submit the notebook. Then the notebook will have access to the hidden test set files. So, in order to make a submission that will land you somewhere on the Leaderboard, this is a necessity, correct? \n\nMy question is then: I would suppose that most of the top-scoring teams on the leaderboard, train outside of Kaggle, using the computing resources available at their research facility and possibly across several days. Something that's not possible with a Kaggle notebook. \n\nSo, if training has been done off-Kaggle, am I correct then in assuming that one would need to make a notebook here on Kaggle, for evaluating on the test set, ***uploading the model that was trained outside of Kaggle***, so that it appears under \"Data\" and then load and use it to do predictions? \n\nIs this indeed the approach of all the people training outside of a Kaggle notebook? \n\nThanks :-)",
    "1301792": "You are correct.  People who train models outside Kaggle upload their model (code and weights) in some kaggle dataset, then add this dataset to the inferencing notebook. You can probably find examples in the public notebooks for this competition.",
    "1303573": "Here's a snippet for a resnet-like model in pytorch that may help you.\nYou will need a dataset `predefinedmodels` and `checkpoints`.\n\n```\ndef init_model(repo, model, checkpoint):\n    model = torch.hub.load(repo, model, pretrained=False, source='local')\n    n_features = model.fc.in_features\n    model.fc = torch.nn.Linear(in_features=n_features, out_features=n_labels, bias=True)\n    cp = torch.load(f\"../input/checkpoints/{checkpoint}.pt\", map_location=torch.device(device))\n    model.load_state_dict(cp['model_state_dict'])\n    model = model.to(device).eval()\n    return model\n...\nmodel = init_model('../input/predefinedmodels/pytorch_vision_v0.9.0', 'resnet18', 'resnet18-20210510-204737-latest')\n```"
  },
  "source": "meta"
}