{
  "id": 89293,
  "title": "Does lwlrap behave differently on batches vs. entire dataset",
  "url": "/competitions/freesound-audio-tagging-2019/discussion/89293",
  "author_name": "",
  "post_date": "2019-04-12T18:57:08.893362400Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi guys,</p>\n\n<p>I've never used lwlrap as a metric before and I'm working through the <a href=\"https://colab.research.google.com/drive/1AgPdhSp7ttY18O3fEoHOQKlt_3HJDLi8#scrollTo=FJv0Rtqfsu3X\">provided code</a> while trying to incorporate it into my approach.</p>\n\n<p>When I run lwlrap on batches of ~64 I receive much higher results than when I run it on my entire dataset. I also notice there an \"accumulator\" version of lwlrap in the provided code that mentions batches. Am I understanding it correctly that lwlrap will change if I run it on batches of the dataset when compared to running it on the entire dataset? </p>",
  "messages": [
    {
      "id": "515550",
      "postDate": "04/12/2019 18:57:08",
      "content": "<p>Hi guys,</p>\n\n<p>I've never used lwlrap as a metric before and I'm working through the <a href=\"https://colab.research.google.com/drive/1AgPdhSp7ttY18O3fEoHOQKlt_3HJDLi8#scrollTo=FJv0Rtqfsu3X\">provided code</a> while trying to incorporate it into my approach.</p>\n\n<p>When I run lwlrap on batches of ~64 I receive much higher results than when I run it on my entire dataset. I also notice there an \"accumulator\" version of lwlrap in the provided code that mentions batches. Am I understanding it correctly that lwlrap will change if I run it on batches of the dataset when compared to running it on the entire dataset? </p>",
      "rawMarkdown": "Hi guys,\n\nI've never used lwlrap as a metric before and I'm working through the [provided code](https://colab.research.google.com/drive/1AgPdhSp7ttY18O3fEoHOQKlt_3HJDLi8#scrollTo=FJv0Rtqfsu3X) while trying to incorporate it into my approach.\n\nWhen I run lwlrap on batches of ~64 I receive much higher results than when I run it on my entire dataset. I also notice there an \"accumulator\" version of lwlrap in the provided code that mentions batches. Am I understanding it correctly that lwlrap will change if I run it on batches of the dataset when compared to running it on the entire dataset?",
      "votes": null
    },
    {
      "id": "515555",
      "postDate": "04/12/2019 19:07:47",
      "content": "<p>How are you combining the lwlrap of each batch to produce the lwlrap of the entire dataset? lwlrap is designed so that you can combine the per-class lwlraps into an overall lwlrap, as demonstrated by the accumulator version that you found in the Colab notebook. If you're doing batch by batch computation, you should make sure to maintain per-class lwlraps.</p>\n\n<p>Although I will note that in practice, you will be computing lwlrap over a relatively small held-out validation set, and I would expect both the predicted scores and the ground truth labels of that entire set to fit in memory, so you would not need to do batch-by-batch computation for performance reasons. </p>\n\n<p>On the other hand, you should probably do per-class lwlrap analysis over your entire validation set because you can use that to get more information about how your models are performing at the class level, which might give you further ideas for improvement.</p>",
      "rawMarkdown": "How are you combining the lwlrap of each batch to produce the lwlrap of the entire dataset? lwlrap is designed so that you can combine the per-class lwlraps into an overall lwlrap, as demonstrated by the accumulator version that you found in the Colab notebook. If you're doing batch by batch computation, you should make sure to maintain per-class lwlraps.\n\nAlthough I will note that in practice, you will be computing lwlrap over a relatively small held-out validation set, and I would expect both the predicted scores and the ground truth labels of that entire set to fit in memory, so you would not need to do batch-by-batch computation for performance reasons. \n\nOn the other hand, you should probably do per-class lwlrap analysis over your entire validation set because you can use that to get more information about how your models are performing at the class level, which might give you further ideas for improvement.",
      "votes": null
    },
    {
      "id": "515569",
      "postDate": "04/12/2019 19:28:21",
      "content": "<p><a href=\"/dpwellis\">@dpwellis</a> our lwlrap expert</p>",
      "rawMarkdown": "dpwellis our lwlrap expert",
      "votes": null
    },
    {
      "id": "515571",
      "postDate": "04/12/2019 19:37:51",
      "content": "<blockquote>\n  <p>How are you combining the lwlrap of each batch to produce the lwlrap of the entire dataset?</p>\n</blockquote>\n\n<p>Originally I was passing <code>calculate_overall_lwlrap_sklearn</code> as a metric for fastai. As I understand it, fastai was simplying running that metric against each batch, calculating the score and then averaging them across all batches.</p>\n\n<p>I was also independently calling <code>calculate_overall_lwlrap_sklearn</code> with the entire validation set which gave me a lower score. </p>\n\n<p>It sounds like running it against the entire validation set is the easiest way to move forward for me, so I'll probably stick with that.</p>",
      "rawMarkdown": "&gt;How are you combining the lwlrap of each batch to produce the lwlrap of the entire dataset?\n\nOriginally I was passing `calculate_overall_lwlrap_sklearn` as a metric for fastai. As I understand it, fastai was simplying running that metric against each batch, calculating the score and then averaging them across all batches.\n\nI was also independently calling `calculate_overall_lwlrap_sklearn` with the entire validation set which gave me a lower score. \n\nIt sounds like running it against the entire validation set is the easiest way to move forward for me, so I'll probably stick with that.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 515555,
      "author_name": "plakal",
      "author_url": "",
      "post_date": "04/12/2019 19:07:47",
      "content": "<p>How are you combining the lwlrap of each batch to produce the lwlrap of the entire dataset? lwlrap is designed so that you can combine the per-class lwlraps into an overall lwlrap, as demonstrated by the accumulator version that you found in the Colab notebook. If you're doing batch by batch computation, you should make sure to maintain per-class lwlraps.</p>\n\n<p>Although I will note that in practice, you will be computing lwlrap over a relatively small held-out validation set, and I would expect both the predicted scores and the ground truth labels of that entire set to fit in memory, so you would not need to do batch-by-batch computation for performance reasons. </p>\n\n<p>On the other hand, you should probably do per-class lwlrap analysis over your entire validation set because you can use that to get more information about how your models are performing at the class level, which might give you further ideas for improvement.</p>",
      "votes": null,
      "replies": [
        {
          "id": 515571,
          "author_name": "joshvarty",
          "author_url": "",
          "post_date": "04/12/2019 19:37:51",
          "content": "<blockquote>\n  <p>How are you combining the lwlrap of each batch to produce the lwlrap of the entire dataset?</p>\n</blockquote>\n\n<p>Originally I was passing <code>calculate_overall_lwlrap_sklearn</code> as a metric for fastai. As I understand it, fastai was simplying running that metric against each batch, calculating the score and then averaging them across all batches.</p>\n\n<p>I was also independently calling <code>calculate_overall_lwlrap_sklearn</code> with the entire validation set which gave me a lower score. </p>\n\n<p>It sounds like running it against the entire validation set is the easiest way to move forward for me, so I'll probably stick with that.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 515569,
      "author_name": "plakal",
      "author_url": "",
      "post_date": "04/12/2019 19:28:21",
      "content": "<p><a href=\"/dpwellis\">@dpwellis</a> our lwlrap expert</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "515550": "Hi guys,\n\nI've never used lwlrap as a metric before and I'm working through the [provided code](https://colab.research.google.com/drive/1AgPdhSp7ttY18O3fEoHOQKlt_3HJDLi8#scrollTo=FJv0Rtqfsu3X) while trying to incorporate it into my approach.\n\nWhen I run lwlrap on batches of ~64 I receive much higher results than when I run it on my entire dataset. I also notice there an \"accumulator\" version of lwlrap in the provided code that mentions batches. Am I understanding it correctly that lwlrap will change if I run it on batches of the dataset when compared to running it on the entire dataset?",
    "515555": "How are you combining the lwlrap of each batch to produce the lwlrap of the entire dataset? lwlrap is designed so that you can combine the per-class lwlraps into an overall lwlrap, as demonstrated by the accumulator version that you found in the Colab notebook. If you're doing batch by batch computation, you should make sure to maintain per-class lwlraps.\n\nAlthough I will note that in practice, you will be computing lwlrap over a relatively small held-out validation set, and I would expect both the predicted scores and the ground truth labels of that entire set to fit in memory, so you would not need to do batch-by-batch computation for performance reasons. \n\nOn the other hand, you should probably do per-class lwlrap analysis over your entire validation set because you can use that to get more information about how your models are performing at the class level, which might give you further ideas for improvement.",
    "515569": "dpwellis our lwlrap expert",
    "515571": "&gt;How are you combining the lwlrap of each batch to produce the lwlrap of the entire dataset?\n\nOriginally I was passing `calculate_overall_lwlrap_sklearn` as a metric for fastai. As I understand it, fastai was simplying running that metric against each batch, calculating the score and then averaging them across all batches.\n\nI was also independently calling `calculate_overall_lwlrap_sklearn` with the entire validation set which gave me a lower score. \n\nIt sounds like running it against the entire validation set is the easiest way to move forward for me, so I'll probably stick with that."
  },
  "source": "meta"
}