{
  "id": 213376,
  "title": "Inference for PANNs (LB vs CV)",
  "url": "/competitions/rfcx-species-audio-detection/discussion/213376",
  "author_name": "",
  "post_date": "2021-01-22T15:32:19.521344100Z",
  "votes": 3,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I have recently started out on using PANN architecture for this competition and have followed closely on the discussions and notebooks relating to PANNs. I would like to thank the community for sharing their knowledge and work! However, i have been trying to wrap my head around the issue of a huge deviation between my Validation &amp; LB score, and in particular, i am not very sure of the inference process. I have used 10s crop on the test set by doing the following:<br>\n<code>1. test_batch = [y[i:i+effective_length].astype(np.float32) for i in range(i, 60*sr, effective_length)]</code>  <br>\n  <code>2. test_batch = torch.FloatTensor(np.stack(test_batch))</code><br>\nwhich generates 6 consecutive 10s crop for each test example with shape <em>(6,480000)</em>. I set <em>batch_size = 1</em>, thus, resulting in model output of <em>(6xTx24)</em> for each test example. By using <em>max(output, dim=T)</em>, i get the max probabilities for each 'crop' of the test example =&gt; <em>(6,24)</em>. It would then make sense for me to sum them up along the 0th axis to get the final output for 1 single test example. </p>\n<ol>\n<li>Is this way of cropping the right way to get prediction for my test? </li>\n<li>I am using the <code>framewise_output</code> for my validation metric, as suggested in <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/209684\" target=\"_blank\">this post</a>, and trained with weak labels. I don't seem to be able to narrow the gap between my validation and LB (difference of ~0.12)<br>\nI would really appreciate any advice on this:) </li>\n</ol>",
  "messages": [
    {
      "id": "1164851",
      "postDate": "01/22/2021 15:32:19",
      "content": "<p>I have recently started out on using PANN architecture for this competition and have followed closely on the discussions and notebooks relating to PANNs. I would like to thank the community for sharing their knowledge and work! However, i have been trying to wrap my head around the issue of a huge deviation between my Validation &amp; LB score, and in particular, i am not very sure of the inference process. I have used 10s crop on the test set by doing the following:<br>\n<code>1. test_batch = [y[i:i+effective_length].astype(np.float32) for i in range(i, 60*sr, effective_length)]</code>  <br>\n  <code>2. test_batch = torch.FloatTensor(np.stack(test_batch))</code><br>\nwhich generates 6 consecutive 10s crop for each test example with shape <em>(6,480000)</em>. I set <em>batch_size = 1</em>, thus, resulting in model output of <em>(6xTx24)</em> for each test example. By using <em>max(output, dim=T)</em>, i get the max probabilities for each 'crop' of the test example =&gt; <em>(6,24)</em>. It would then make sense for me to sum them up along the 0th axis to get the final output for 1 single test example. </p>\n<ol>\n<li>Is this way of cropping the right way to get prediction for my test? </li>\n<li>I am using the <code>framewise_output</code> for my validation metric, as suggested in <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/209684\" target=\"_blank\">this post</a>, and trained with weak labels. I don't seem to be able to narrow the gap between my validation and LB (difference of ~0.12)<br>\nI would really appreciate any advice on this:) </li>\n</ol>",
      "rawMarkdown": "I have recently started out on using PANN architecture for this competition and have followed closely on the discussions and notebooks relating to PANNs. I would like to thank the community for sharing their knowledge and work! However, i have been trying to wrap my head around the issue of a huge deviation between my Validation & LB score, and in particular, i am not very sure of the inference process. I have used 10s crop on the test set by doing the following:\n`1. test_batch = [y[i:i+effective_length].astype(np.float32) for i in range(i, 60*sr, effective_length)] `  \n  `2. test_batch = torch.FloatTensor(np.stack(test_batch))`\nwhich generates 6 consecutive 10s crop for each test example with shape *(6,480000)*. I set *batch_size = 1*, thus, resulting in model output of *(6xTx24)* for each test example. By using *max(output, dim=T)*, i get the max probabilities for each 'crop' of the test example => *(6,24)*. It would then make sense for me to sum them up along the 0th axis to get the final output for 1 single test example. \n1. Is this way of cropping the right way to get prediction for my test? \n2. I am using the `framewise_output` for my validation metric, as suggested in [this post](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/209684), and trained with weak labels. I don't seem to be able to narrow the gap between my validation and LB (difference of ~0.12)\nI would really appreciate any advice on this:)",
      "votes": null
    },
    {
      "id": "1183467",
      "postDate": "02/03/2021 02:58:04",
      "content": "<p>IMHO ~0.12 diffierence between CV and LB is quite normal, as the train dataset is small. Your inference crop is fine as most of discussion are using the same way, and as many discussion indicated that the key for this competition is how to crop the training data. If you generate training data smarter, and use more folds or models to ensemble your prediction, you may see stable CV and LB correlation. </p>",
      "rawMarkdown": "IMHO ~0.12 diffierence between CV and LB is quite normal, as the train dataset is small. Your inference crop is fine as most of discussion are using the same way, and as many discussion indicated that the key for this competition is how to crop the training data. If you generate training data smarter, and use more folds or models to ensemble your prediction, you may see stable CV and LB correlation.",
      "votes": null
    },
    {
      "id": "1183762",
      "postDate": "02/03/2021 07:51:26",
      "content": "<blockquote>\n  <p>huge deviation between my Validation &amp; LB score</p>\n</blockquote>\n<p>This is not necessarily an issue as long as they are correlated.  If your LB increases when you CV increases then the gap between the two isn't relevant.</p>\n<p>In this competition it is hard to get a great CV setting because we don't have ground truth for training data.  Indeed, we don't have recoding level info on the presence of all species.  </p>\n<p>TL;DR Dont' worry about CV LB gap as long as you have some correlation between the two.</p>",
      "rawMarkdown": "> huge deviation between my Validation & LB score\n\nThis is not necessarily an issue as long as they are correlated.  If your LB increases when you CV increases then the gap between the two isn't relevant.\n\nIn this competition it is hard to get a great CV setting because we don't have ground truth for training data.  Indeed, we don't have recoding level info on the presence of all species.  \n\nTL;DR Dont' worry about CV LB gap as long as you have some correlation between the two.",
      "votes": null
    },
    {
      "id": "1183776",
      "postDate": "02/03/2021 07:58:56",
      "content": "<p>Thank you for the advice! I guess i will have to change my approach :)</p>",
      "rawMarkdown": "Thank you for the advice! I guess i will have to change my approach :)",
      "votes": null
    },
    {
      "id": "1183779",
      "postDate": "02/03/2021 08:00:35",
      "content": "<p>Thank you for the explanation! :) With this i will not have to spend unnecessary amount of time pondering about it and focus on my approach instead!</p>",
      "rawMarkdown": "Thank you for the explanation! :) With this i will not have to spend unnecessary amount of time pondering about it and focus on my approach instead!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1183467,
      "author_name": "superchenhao",
      "author_url": "",
      "post_date": "02/03/2021 02:58:04",
      "content": "<p>IMHO ~0.12 diffierence between CV and LB is quite normal, as the train dataset is small. Your inference crop is fine as most of discussion are using the same way, and as many discussion indicated that the key for this competition is how to crop the training data. If you generate training data smarter, and use more folds or models to ensemble your prediction, you may see stable CV and LB correlation. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1183776,
          "author_name": "angqx95",
          "author_url": "",
          "post_date": "02/03/2021 07:58:56",
          "content": "<p>Thank you for the advice! I guess i will have to change my approach :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1183762,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "02/03/2021 07:51:26",
      "content": "<blockquote>\n  <p>huge deviation between my Validation &amp; LB score</p>\n</blockquote>\n<p>This is not necessarily an issue as long as they are correlated.  If your LB increases when you CV increases then the gap between the two isn't relevant.</p>\n<p>In this competition it is hard to get a great CV setting because we don't have ground truth for training data.  Indeed, we don't have recoding level info on the presence of all species.  </p>\n<p>TL;DR Dont' worry about CV LB gap as long as you have some correlation between the two.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1183779,
          "author_name": "angqx95",
          "author_url": "",
          "post_date": "02/03/2021 08:00:35",
          "content": "<p>Thank you for the explanation! :) With this i will not have to spend unnecessary amount of time pondering about it and focus on my approach instead!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1164851": "I have recently started out on using PANN architecture for this competition and have followed closely on the discussions and notebooks relating to PANNs. I would like to thank the community for sharing their knowledge and work! However, i have been trying to wrap my head around the issue of a huge deviation between my Validation & LB score, and in particular, i am not very sure of the inference process. I have used 10s crop on the test set by doing the following:\n`1. test_batch = [y[i:i+effective_length].astype(np.float32) for i in range(i, 60*sr, effective_length)] `  \n  `2. test_batch = torch.FloatTensor(np.stack(test_batch))`\nwhich generates 6 consecutive 10s crop for each test example with shape *(6,480000)*. I set *batch_size = 1*, thus, resulting in model output of *(6xTx24)* for each test example. By using *max(output, dim=T)*, i get the max probabilities for each 'crop' of the test example => *(6,24)*. It would then make sense for me to sum them up along the 0th axis to get the final output for 1 single test example. \n1. Is this way of cropping the right way to get prediction for my test? \n2. I am using the `framewise_output` for my validation metric, as suggested in [this post](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/209684), and trained with weak labels. I don't seem to be able to narrow the gap between my validation and LB (difference of ~0.12)\nI would really appreciate any advice on this:)",
    "1183467": "IMHO ~0.12 diffierence between CV and LB is quite normal, as the train dataset is small. Your inference crop is fine as most of discussion are using the same way, and as many discussion indicated that the key for this competition is how to crop the training data. If you generate training data smarter, and use more folds or models to ensemble your prediction, you may see stable CV and LB correlation.",
    "1183762": "> huge deviation between my Validation & LB score\n\nThis is not necessarily an issue as long as they are correlated.  If your LB increases when you CV increases then the gap between the two isn't relevant.\n\nIn this competition it is hard to get a great CV setting because we don't have ground truth for training data.  Indeed, we don't have recoding level info on the presence of all species.  \n\nTL;DR Dont' worry about CV LB gap as long as you have some correlation between the two.",
    "1183776": "Thank you for the advice! I guess i will have to change my approach :)",
    "1183779": "Thank you for the explanation! :) With this i will not have to spend unnecessary amount of time pondering about it and focus on my approach instead!"
  },
  "source": "meta"
}