{
  "id": 189901,
  "title": "Reproduce Dataloaders for comparing models",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/189901",
  "author_name": "",
  "post_date": "2020-10-09T09:03:28.437397Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi,</p>\n<p>I am new to data science and excited to participate in my first kaggle competition. As discussed already in the forum, the data for the competition is very large and I rely only on kaggle GPUs for training purpose. I want to run different models on fixed subset of data (~25% training and ~25% validation sets) and then only train the more accurate models further, based on validation score.</p>\n<p>Basically, I want to train different models on same subset of data (but shuffled). I went through <a href=\"https://github.com/pytorch/pytorch/issues/7068\" target=\"_blank\">github page</a> but still couldn't understand as many variables are mentioned (np.random.seed, torch.manual_seed(seed)<br>\ntorch.cuda.manual_seed(seed) etc). </p>\n<p>Also, how can we compare if two subsets of data are same? </p>",
  "messages": [
    {
      "id": "1043811",
      "postDate": "10/09/2020 09:03:28",
      "content": "<p>Hi,</p>\n<p>I am new to data science and excited to participate in my first kaggle competition. As discussed already in the forum, the data for the competition is very large and I rely only on kaggle GPUs for training purpose. I want to run different models on fixed subset of data (~25% training and ~25% validation sets) and then only train the more accurate models further, based on validation score.</p>\n<p>Basically, I want to train different models on same subset of data (but shuffled). I went through <a href=\"https://github.com/pytorch/pytorch/issues/7068\" target=\"_blank\">github page</a> but still couldn't understand as many variables are mentioned (np.random.seed, torch.manual_seed(seed)<br>\ntorch.cuda.manual_seed(seed) etc). </p>\n<p>Also, how can we compare if two subsets of data are same? </p>",
      "rawMarkdown": "Hi,\n\nI am new to data science and excited to participate in my first kaggle competition. As discussed already in the forum, the data for the competition is very large and I rely only on kaggle GPUs for training purpose. I want to run different models on fixed subset of data (~25% training and ~25% validation sets) and then only train the more accurate models further, based on validation score.\n\nBasically, I want to train different models on same subset of data (but shuffled). I went through [github page](https://github.com/pytorch/pytorch/issues/7068) but still couldn't understand as many variables are mentioned (np.random.seed, torch.manual_seed(seed)\ntorch.cuda.manual_seed(seed) etc). \n\nAlso, how can we compare if two subsets of data are same?",
      "votes": null
    },
    {
      "id": "1047298",
      "postDate": "10/12/2020 12:59:12",
      "content": "<blockquote>\n  <p>Basically, I want to train different models on same subset of data (but shuffled)</p>\n</blockquote>\n<p>if you use pytorch, wrap the agent dataset using Dataloader with <a href=\"https://pytorch.org/docs/stable/data.html#torch.utils.data.SubsetRandomSampler\" target=\"_blank\">SubsetRandomSampler</a></p>",
      "rawMarkdown": "> Basically, I want to train different models on same subset of data (but shuffled)\n\nif you use pytorch, wrap the agent dataset using Dataloader with [SubsetRandomSampler](https://pytorch.org/docs/stable/data.html#torch.utils.data.SubsetRandomSampler)",
      "votes": null
    },
    {
      "id": "1053645",
      "postDate": "10/19/2020 07:34:00",
      "content": "<p>Thanks for the help <a href=\"https://www.kaggle.com/etareduce\" target=\"_blank\">@etareduce</a></p>",
      "rawMarkdown": "Thanks for the help @etareduce",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1047298,
      "author_name": "etareduce",
      "author_url": "",
      "post_date": "10/12/2020 12:59:12",
      "content": "<blockquote>\n  <p>Basically, I want to train different models on same subset of data (but shuffled)</p>\n</blockquote>\n<p>if you use pytorch, wrap the agent dataset using Dataloader with <a href=\"https://pytorch.org/docs/stable/data.html#torch.utils.data.SubsetRandomSampler\" target=\"_blank\">SubsetRandomSampler</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1053645,
          "author_name": "suryajrrafl",
          "author_url": "",
          "post_date": "10/19/2020 07:34:00",
          "content": "<p>Thanks for the help <a href=\"https://www.kaggle.com/etareduce\" target=\"_blank\">@etareduce</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1043811": "Hi,\n\nI am new to data science and excited to participate in my first kaggle competition. As discussed already in the forum, the data for the competition is very large and I rely only on kaggle GPUs for training purpose. I want to run different models on fixed subset of data (~25% training and ~25% validation sets) and then only train the more accurate models further, based on validation score.\n\nBasically, I want to train different models on same subset of data (but shuffled). I went through [github page](https://github.com/pytorch/pytorch/issues/7068) but still couldn't understand as many variables are mentioned (np.random.seed, torch.manual_seed(seed)\ntorch.cuda.manual_seed(seed) etc). \n\nAlso, how can we compare if two subsets of data are same?",
    "1047298": "> Basically, I want to train different models on same subset of data (but shuffled)\n\nif you use pytorch, wrap the agent dataset using Dataloader with [SubsetRandomSampler](https://pytorch.org/docs/stable/data.html#torch.utils.data.SubsetRandomSampler)",
    "1053645": "Thanks for the help @etareduce"
  },
  "source": "meta"
}