{
  "id": 61152,
  "title": "Anyone has a PyTorch tutorial on how to use Dataloaders?",
  "url": "/competitions/google-ai-open-images-object-detection-track/discussion/61152",
  "author_name": "",
  "post_date": "2018-07-14T22:09:13.905014900Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "",
  "messages": [
    {
      "id": "356957",
      "postDate": "07/14/2018 22:09:13",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "357362",
      "postDate": "07/15/2018 20:00:45",
      "content": "<p>A good way to conceptualize what a dataloader vs a dataset does is that a dataset prescribes how to fetch an example with a given ID and a dataloader handles doing the actual reading (reading in parallel for instance) and combining examples into batches.</p>\n\n<p>PyTorch documentation is a really good place to start and the base Dataset class - which comprises only a couple of lines of code - is a really good read.</p>\n\n<p>Having said that, unfortunately none of this will be very helpful for this competition on its own. The challenge here is how to process the bounding box information and how to present it so that we can run a loss function against our model's predictions. With image recognition - telling what class the main object in an image belongs to - this is easy, we just put a softmax at the end of our model and we are done. Here this is the part that requires much more thought.</p>\n\n<p>The best material that speaks to this that I am aware of is the beginning of <a href=\"http://course.fast.ai/part2.html\">part 2 of the fastai course</a> which is available on the Internet for free. All this is not all that complicated but explaining it is quite involved and I am not aware of there existing any tutorial doing a good job at this (might be beyond the scope of a single article, a lecture might work better).</p>\n\n<p>From a practical standpoint, you might be better off just going with the object detection functionality that the google team shares <a href=\"https://github.com/tensorflow/models/tree/master/research/object_detection\">here</a> which is developed in Tensorflow. The problem is that this code is not easy to understand and is really hard to learn from if you don't already have quite a bit of understanding of what is going on. For the results in this competition you might be best off by figuring out what is going on in the repo I linked - for building up your understanding you might benefit most from doing the fastai course or studying some other library.</p>",
      "rawMarkdown": "A good way to conceptualize what a dataloader vs a dataset does is that a dataset prescribes how to fetch an example with a given ID and a dataloader handles doing the actual reading (reading in parallel for instance) and combining examples into batches.\n\nPyTorch documentation is a really good place to start and the base Dataset class - which comprises only a couple of lines of code - is a really good read.\n\nHaving said that, unfortunately none of this will be very helpful for this competition on its own. The challenge here is how to process the bounding box information and how to present it so that we can run a loss function against our model's predictions. With image recognition - telling what class the main object in an image belongs to - this is easy, we just put a softmax at the end of our model and we are done. Here this is the part that requires much more thought.\n\nThe best material that speaks to this that I am aware of is the beginning of [part 2 of the fastai course][1] which is available on the Internet for free. All this is not all that complicated but explaining it is quite involved and I am not aware of there existing any tutorial doing a good job at this (might be beyond the scope of a single article, a lecture might work better).\n\nFrom a practical standpoint, you might be better off just going with the object detection functionality that the google team shares [here][2] which is developed in Tensorflow. The problem is that this code is not easy to understand and is really hard to learn from if you don't already have quite a bit of understanding of what is going on. For the results in this competition you might be best off by figuring out what is going on in the repo I linked - for building up your understanding you might benefit most from doing the fastai course or studying some other library.\n\n\n  [1]: http://course.fast.ai/part2.html\n  [2]: https://github.com/tensorflow/models/tree/master/research/object_detection",
      "votes": null
    },
    {
      "id": "358195",
      "postDate": "07/17/2018 17:37:51",
      "content": "<p>I thought this <a href=\"https://www.youtube.com/watch?v=zN49HdDxHi8\">YouTube Tutorial</a> was very clear. \nAfter that I'd take a look at <a href=\"https://stanford.edu/~shervine/blog/pytorch-how-to-generate-data-parallel.html\">Stanford's tutorial</a>, <a href=\"https://pytorch.org/tutorials/beginner/data_loading_tutorial.html\">PyTorch's website</a> and lastly, the PyTorch Forums have lots of examples and common confusions with great answers from the Dev Team &amp; community. For example, <a href=\"https://discuss.pytorch.org/t/trying-to-build-a-dataloader/12618/4\">here</a>.\nHope this helps!</p>",
      "rawMarkdown": "I thought this [YouTube Tutorial][1] was very clear. \nAfter that I'd take a look at [Stanford's tutorial][2], [PyTorch's website][3] and lastly, the PyTorch Forums have lots of examples and common confusions with great answers from the Dev Team &amp; community. For example, [here][4].\nHope this helps!\n\n\n  [1]: https://www.youtube.com/watch?v=zN49HdDxHi8\n  [2]: https://stanford.edu/~shervine/blog/pytorch-how-to-generate-data-parallel.html\n  [3]: https://pytorch.org/tutorials/beginner/data_loading_tutorial.html\n  [4]: https://discuss.pytorch.org/t/trying-to-build-a-dataloader/12618/4",
      "votes": null
    },
    {
      "id": "358983",
      "postDate": "07/19/2018 09:30:45",
      "content": "<p>Thank you. This cleared a lot of things</p>",
      "rawMarkdown": "Thank you. This cleared a lot of things",
      "votes": null
    },
    {
      "id": "359549",
      "postDate": "07/20/2018 10:29:46",
      "content": "<p>@Mario Aspiris</p>\n\n<p>We have implemented dataset/dataloaders for this competition in this open solution <a href=\"https://github.com/neptune-ml/open-solution-googleai-object-detection\">repo</a></p>\n\n<p>The datasets/loaders are <a href=\"https://github.com/neptune-ml/open-solution-googleai-object-detection/blob/master/src/loaders.py\">here</a></p>\n\n<p>The retinanet model itself is training and we should have first reasonable results after the weekend.\nYou can monitor the progress <a href=\"https://app.neptune.ml/neptune-ml/Google-AI-Object-Detection-Challenge\">here</a> if you like.</p>",
      "rawMarkdown": "Mario Aspiris\n\nWe have implemented dataset/dataloaders for this competition in this open solution [repo](https://github.com/neptune-ml/open-solution-googleai-object-detection)\n\nThe datasets/loaders are [here](https://github.com/neptune-ml/open-solution-googleai-object-detection/blob/master/src/loaders.py)\n\nThe retinanet model itself is training and we should have first reasonable results after the weekend.\nYou can monitor the progress [here](https://app.neptune.ml/neptune-ml/Google-AI-Object-Detection-Challenge) if you like.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 357362,
      "author_name": "radek1",
      "author_url": "",
      "post_date": "07/15/2018 20:00:45",
      "content": "<p>A good way to conceptualize what a dataloader vs a dataset does is that a dataset prescribes how to fetch an example with a given ID and a dataloader handles doing the actual reading (reading in parallel for instance) and combining examples into batches.</p>\n\n<p>PyTorch documentation is a really good place to start and the base Dataset class - which comprises only a couple of lines of code - is a really good read.</p>\n\n<p>Having said that, unfortunately none of this will be very helpful for this competition on its own. The challenge here is how to process the bounding box information and how to present it so that we can run a loss function against our model's predictions. With image recognition - telling what class the main object in an image belongs to - this is easy, we just put a softmax at the end of our model and we are done. Here this is the part that requires much more thought.</p>\n\n<p>The best material that speaks to this that I am aware of is the beginning of <a href=\"http://course.fast.ai/part2.html\">part 2 of the fastai course</a> which is available on the Internet for free. All this is not all that complicated but explaining it is quite involved and I am not aware of there existing any tutorial doing a good job at this (might be beyond the scope of a single article, a lecture might work better).</p>\n\n<p>From a practical standpoint, you might be better off just going with the object detection functionality that the google team shares <a href=\"https://github.com/tensorflow/models/tree/master/research/object_detection\">here</a> which is developed in Tensorflow. The problem is that this code is not easy to understand and is really hard to learn from if you don't already have quite a bit of understanding of what is going on. For the results in this competition you might be best off by figuring out what is going on in the repo I linked - for building up your understanding you might benefit most from doing the fastai course or studying some other library.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 358195,
      "author_name": "differentialmind",
      "author_url": "",
      "post_date": "07/17/2018 17:37:51",
      "content": "<p>I thought this <a href=\"https://www.youtube.com/watch?v=zN49HdDxHi8\">YouTube Tutorial</a> was very clear. \nAfter that I'd take a look at <a href=\"https://stanford.edu/~shervine/blog/pytorch-how-to-generate-data-parallel.html\">Stanford's tutorial</a>, <a href=\"https://pytorch.org/tutorials/beginner/data_loading_tutorial.html\">PyTorch's website</a> and lastly, the PyTorch Forums have lots of examples and common confusions with great answers from the Dev Team &amp; community. For example, <a href=\"https://discuss.pytorch.org/t/trying-to-build-a-dataloader/12618/4\">here</a>.\nHope this helps!</p>",
      "votes": null,
      "replies": [
        {
          "id": 358983,
          "author_name": "pejibaye",
          "author_url": "",
          "post_date": "07/19/2018 09:30:45",
          "content": "<p>Thank you. This cleared a lot of things</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 359549,
      "author_name": "jakubczakon",
      "author_url": "",
      "post_date": "07/20/2018 10:29:46",
      "content": "<p>@Mario Aspiris</p>\n\n<p>We have implemented dataset/dataloaders for this competition in this open solution <a href=\"https://github.com/neptune-ml/open-solution-googleai-object-detection\">repo</a></p>\n\n<p>The datasets/loaders are <a href=\"https://github.com/neptune-ml/open-solution-googleai-object-detection/blob/master/src/loaders.py\">here</a></p>\n\n<p>The retinanet model itself is training and we should have first reasonable results after the weekend.\nYou can monitor the progress <a href=\"https://app.neptune.ml/neptune-ml/Google-AI-Object-Detection-Challenge\">here</a> if you like.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "356957": "",
    "357362": "A good way to conceptualize what a dataloader vs a dataset does is that a dataset prescribes how to fetch an example with a given ID and a dataloader handles doing the actual reading (reading in parallel for instance) and combining examples into batches.\n\nPyTorch documentation is a really good place to start and the base Dataset class - which comprises only a couple of lines of code - is a really good read.\n\nHaving said that, unfortunately none of this will be very helpful for this competition on its own. The challenge here is how to process the bounding box information and how to present it so that we can run a loss function against our model's predictions. With image recognition - telling what class the main object in an image belongs to - this is easy, we just put a softmax at the end of our model and we are done. Here this is the part that requires much more thought.\n\nThe best material that speaks to this that I am aware of is the beginning of [part 2 of the fastai course][1] which is available on the Internet for free. All this is not all that complicated but explaining it is quite involved and I am not aware of there existing any tutorial doing a good job at this (might be beyond the scope of a single article, a lecture might work better).\n\nFrom a practical standpoint, you might be better off just going with the object detection functionality that the google team shares [here][2] which is developed in Tensorflow. The problem is that this code is not easy to understand and is really hard to learn from if you don't already have quite a bit of understanding of what is going on. For the results in this competition you might be best off by figuring out what is going on in the repo I linked - for building up your understanding you might benefit most from doing the fastai course or studying some other library.\n\n\n  [1]: http://course.fast.ai/part2.html\n  [2]: https://github.com/tensorflow/models/tree/master/research/object_detection",
    "358195": "I thought this [YouTube Tutorial][1] was very clear. \nAfter that I'd take a look at [Stanford's tutorial][2], [PyTorch's website][3] and lastly, the PyTorch Forums have lots of examples and common confusions with great answers from the Dev Team &amp; community. For example, [here][4].\nHope this helps!\n\n\n  [1]: https://www.youtube.com/watch?v=zN49HdDxHi8\n  [2]: https://stanford.edu/~shervine/blog/pytorch-how-to-generate-data-parallel.html\n  [3]: https://pytorch.org/tutorials/beginner/data_loading_tutorial.html\n  [4]: https://discuss.pytorch.org/t/trying-to-build-a-dataloader/12618/4",
    "358983": "Thank you. This cleared a lot of things",
    "359549": "Mario Aspiris\n\nWe have implemented dataset/dataloaders for this competition in this open solution [repo](https://github.com/neptune-ml/open-solution-googleai-object-detection)\n\nThe datasets/loaders are [here](https://github.com/neptune-ml/open-solution-googleai-object-detection/blob/master/src/loaders.py)\n\nThe retinanet model itself is training and we should have first reasonable results after the weekend.\nYou can monitor the progress [here](https://app.neptune.ml/neptune-ml/Google-AI-Object-Detection-Challenge) if you like."
  },
  "source": "meta"
}