{
  "id": 494142,
  "title": "efficientnetv2-b2 complete workflow [LB .61]",
  "url": "/competitions/birdclef-2024/discussion/494142",
  "author_name": "",
  "post_date": "2024-04-16T05:58:20.375167Z",
  "votes": 7,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Sharing my updated Spectrograms -&gt; ImageNet toolchain.</p>\n<p>There are a bunch of changes from <a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/492135\" target=\"_blank\">my prior effort</a> - that collectively got me from .57 up to .61 on the LB.</p>\n<p><a href=\"https://www.kaggle.com/code/richolson/birdclef-2024-contiguous-mel-spectrogram-generator\" target=\"_blank\">https://www.kaggle.com/code/richolson/birdclef-2024-contiguous-mel-spectrogram-generator</a><br>\n<a href=\"https://www.kaggle.com/code/richolson/birdclef-2024-train-v2-efficientnetv2-b2\" target=\"_blank\">https://www.kaggle.com/code/richolson/birdclef-2024-train-v2-efficientnetv2-b2</a><br>\n<a href=\"https://www.kaggle.com/code/richolson/birdclef-2024-run-v2-efficientnetv2-b2\" target=\"_blank\">https://www.kaggle.com/code/richolson/birdclef-2024-run-v2-efficientnetv2-b2</a></p>\n<p>Feedback / questions are welcome here or on the notebooks!</p>\n<p><strong>Generating Spectrograms</strong><br>\n<a href=\"https://www.kaggle.com/richolson/birdclef-2024-contiguous-mel-spectrogram-generator\" target=\"_blank\">https://www.kaggle.com/richolson/birdclef-2024-contiguous-mel-spectrogram-generator</a></p>\n<p>Spectrograms are now generated \"contiguously\" (if that's the right word).</p>\n<p>All files are 512-pixels-wide-per-5-second-interval and 224 high.  For example - a 50-second OGG file produces a PNG that is 5120x224.  I wasn't sure what time resolution I wanted - so I went relatively high horizontal resolution with the plan of scaling-down.</p>\n<p>I am also now generating them monochrome.  All files are processed - producing about 4GB of data.</p>\n<p>(I got some weird errors trying to generate a Dataset - so am just leaving this as Notebook output.  If so inclined - feel free to generate a Dataset from this / publish.)</p>\n<p><strong>Training</strong><br>\n<a href=\"https://www.kaggle.com/code/richolson/birdclef-2024-train-v2-efficientnetv2-b2\" target=\"_blank\">https://www.kaggle.com/code/richolson/birdclef-2024-train-v2-efficientnetv2-b2</a></p>\n<p>efficientnetv2_b2 is trained on random 10-second segments of spectrograms.  This assures the model gets trained with different images each epoch.  A 1024x224 image is randomly cropped from the full spectrogram image - and then scaled down to 224x224 for feeding into the imagenet.</p>\n<p>The model is fully enabled for training / no frozen layers.</p>\n<p>I went with 10-second segments because it seemed like too-frequently 5-second segments didn't include an obvious bird call.</p>\n<p>The maximum training files per class is capped at 3x the median for all classes.  This is done to help reduce over-representation of classes where there is a lot of data.</p>\n<p>Similarly - if a class has under a 30% file count of the median - references to files for the class are duplicated to help assure good representation in training.</p>\n<p>Any files that have \"secondary labels\" (multiple identified bird calls) in train_metadata.csv are removed from training.  The code is there to remove low-quality ratings - but it's not enabled.</p>\n<p>There is also some noise / contrast / brightness modification done to further augment the data.  I'm not doing any rotating / zoom - as I don't think it would help with spectrograms.</p>\n<p><strong>Running</strong><br>\n<a href=\"https://www.kaggle.com/richolson/birdclef-2024-run-v2-efficientnetv2-b2\" target=\"_blank\">https://www.kaggle.com/richolson/birdclef-2024-run-v2-efficientnetv2-b2</a></p>\n<p>The run notebook generates 10-second 224x224 spectrograms from the test data.  The model then predicts on the images - and duplicates each answer over 2 x 5-second prediction windows.</p>\n<p>This process takes about 5.5 seconds per 4-minute file.  The maximum allowed for 2 hours would be about 6.5 seconds.</p>\n<p><strong>Thoughts on further development</strong><br>\nThe current \"run\" notebook doesn't implement batching - so the CPU isn't that well utilized.  I did a test-run with batching implemented - and prediction time is more like 4-seconds per file.  So - I think there is opportunity to upgrade to a model with more parameters.</p>\n<p>Alternately - it might be possible to use a slightly-smaller model - and predict on each 5-second interval.  Maybe still training the model on 10-second samples - but overlapping the predictions somehow.</p>\n<p>Some effort to objectively quantify the importance of prediction frequency in this competition might be interesting…</p>\n<p>The model is probably a little overfit (no validation improvement the last several epochs) - but I don't think that hurts performance much.  Implementing early-stopping would probably be a good idea.  It may also be adjustment of learning rate / data augmentation settings might help.</p>\n<p>I current don't do anything with the geotagging of samples provided in the metadata.  An interesting experiment might be training the model -only- only on samples from the targeted Western Ghats - and then doing another model trained excluding any samples from the Western Ghats. </p>\n<p>Hope this is useful to someone!</p>\n<p>-Rich</p>",
  "messages": [
    {
      "id": "2754615",
      "postDate": "04/16/2024 05:58:20",
      "content": "<p>Sharing my updated Spectrograms -&gt; ImageNet toolchain.</p>\n<p>There are a bunch of changes from <a href=\"https://www.kaggle.com/competitions/birdclef-2024/discussion/492135\" target=\"_blank\">my prior effort</a> - that collectively got me from .57 up to .61 on the LB.</p>\n<p><a href=\"https://www.kaggle.com/code/richolson/birdclef-2024-contiguous-mel-spectrogram-generator\" target=\"_blank\">https://www.kaggle.com/code/richolson/birdclef-2024-contiguous-mel-spectrogram-generator</a><br>\n<a href=\"https://www.kaggle.com/code/richolson/birdclef-2024-train-v2-efficientnetv2-b2\" target=\"_blank\">https://www.kaggle.com/code/richolson/birdclef-2024-train-v2-efficientnetv2-b2</a><br>\n<a href=\"https://www.kaggle.com/code/richolson/birdclef-2024-run-v2-efficientnetv2-b2\" target=\"_blank\">https://www.kaggle.com/code/richolson/birdclef-2024-run-v2-efficientnetv2-b2</a></p>\n<p>Feedback / questions are welcome here or on the notebooks!</p>\n<p><strong>Generating Spectrograms</strong><br>\n<a href=\"https://www.kaggle.com/richolson/birdclef-2024-contiguous-mel-spectrogram-generator\" target=\"_blank\">https://www.kaggle.com/richolson/birdclef-2024-contiguous-mel-spectrogram-generator</a></p>\n<p>Spectrograms are now generated \"contiguously\" (if that's the right word).</p>\n<p>All files are 512-pixels-wide-per-5-second-interval and 224 high.  For example - a 50-second OGG file produces a PNG that is 5120x224.  I wasn't sure what time resolution I wanted - so I went relatively high horizontal resolution with the plan of scaling-down.</p>\n<p>I am also now generating them monochrome.  All files are processed - producing about 4GB of data.</p>\n<p>(I got some weird errors trying to generate a Dataset - so am just leaving this as Notebook output.  If so inclined - feel free to generate a Dataset from this / publish.)</p>\n<p><strong>Training</strong><br>\n<a href=\"https://www.kaggle.com/code/richolson/birdclef-2024-train-v2-efficientnetv2-b2\" target=\"_blank\">https://www.kaggle.com/code/richolson/birdclef-2024-train-v2-efficientnetv2-b2</a></p>\n<p>efficientnetv2_b2 is trained on random 10-second segments of spectrograms.  This assures the model gets trained with different images each epoch.  A 1024x224 image is randomly cropped from the full spectrogram image - and then scaled down to 224x224 for feeding into the imagenet.</p>\n<p>The model is fully enabled for training / no frozen layers.</p>\n<p>I went with 10-second segments because it seemed like too-frequently 5-second segments didn't include an obvious bird call.</p>\n<p>The maximum training files per class is capped at 3x the median for all classes.  This is done to help reduce over-representation of classes where there is a lot of data.</p>\n<p>Similarly - if a class has under a 30% file count of the median - references to files for the class are duplicated to help assure good representation in training.</p>\n<p>Any files that have \"secondary labels\" (multiple identified bird calls) in train_metadata.csv are removed from training.  The code is there to remove low-quality ratings - but it's not enabled.</p>\n<p>There is also some noise / contrast / brightness modification done to further augment the data.  I'm not doing any rotating / zoom - as I don't think it would help with spectrograms.</p>\n<p><strong>Running</strong><br>\n<a href=\"https://www.kaggle.com/richolson/birdclef-2024-run-v2-efficientnetv2-b2\" target=\"_blank\">https://www.kaggle.com/richolson/birdclef-2024-run-v2-efficientnetv2-b2</a></p>\n<p>The run notebook generates 10-second 224x224 spectrograms from the test data.  The model then predicts on the images - and duplicates each answer over 2 x 5-second prediction windows.</p>\n<p>This process takes about 5.5 seconds per 4-minute file.  The maximum allowed for 2 hours would be about 6.5 seconds.</p>\n<p><strong>Thoughts on further development</strong><br>\nThe current \"run\" notebook doesn't implement batching - so the CPU isn't that well utilized.  I did a test-run with batching implemented - and prediction time is more like 4-seconds per file.  So - I think there is opportunity to upgrade to a model with more parameters.</p>\n<p>Alternately - it might be possible to use a slightly-smaller model - and predict on each 5-second interval.  Maybe still training the model on 10-second samples - but overlapping the predictions somehow.</p>\n<p>Some effort to objectively quantify the importance of prediction frequency in this competition might be interesting…</p>\n<p>The model is probably a little overfit (no validation improvement the last several epochs) - but I don't think that hurts performance much.  Implementing early-stopping would probably be a good idea.  It may also be adjustment of learning rate / data augmentation settings might help.</p>\n<p>I current don't do anything with the geotagging of samples provided in the metadata.  An interesting experiment might be training the model -only- only on samples from the targeted Western Ghats - and then doing another model trained excluding any samples from the Western Ghats. </p>\n<p>Hope this is useful to someone!</p>\n<p>-Rich</p>",
      "rawMarkdown": "Sharing my updated Spectrograms -> ImageNet toolchain.\n\nThere are a bunch of changes from [my prior effort](https://www.kaggle.com/competitions/birdclef-2024/discussion/492135) - that collectively got me from .57 up to .61 on the LB.\n\nhttps://www.kaggle.com/code/richolson/birdclef-2024-contiguous-mel-spectrogram-generator\nhttps://www.kaggle.com/code/richolson/birdclef-2024-train-v2-efficientnetv2-b2\nhttps://www.kaggle.com/code/richolson/birdclef-2024-run-v2-efficientnetv2-b2\n\nFeedback / questions are welcome here or on the notebooks!\n\n**Generating Spectrograms**\nhttps://www.kaggle.com/richolson/birdclef-2024-contiguous-mel-spectrogram-generator\n\nSpectrograms are now generated \"contiguously\" (if that's the right word).\n\nAll files are 512-pixels-wide-per-5-second-interval and 224 high.  For example - a 50-second OGG file produces a PNG that is 5120x224.  I wasn't sure what time resolution I wanted - so I went relatively high horizontal resolution with the plan of scaling-down.\n\nI am also now generating them monochrome.  All files are processed - producing about 4GB of data.\n\n(I got some weird errors trying to generate a Dataset - so am just leaving this as Notebook output.  If so inclined - feel free to generate a Dataset from this / publish.)\n\n**Training**\nhttps://www.kaggle.com/code/richolson/birdclef-2024-train-v2-efficientnetv2-b2\n\nefficientnetv2_b2 is trained on random 10-second segments of spectrograms.  This assures the model gets trained with different images each epoch.  A 1024x224 image is randomly cropped from the full spectrogram image - and then scaled down to 224x224 for feeding into the imagenet.\n\nThe model is fully enabled for training / no frozen layers.\n\nI went with 10-second segments because it seemed like too-frequently 5-second segments didn't include an obvious bird call.\n\nThe maximum training files per class is capped at 3x the median for all classes.  This is done to help reduce over-representation of classes where there is a lot of data.\n\nSimilarly - if a class has under a 30% file count of the median - references to files for the class are duplicated to help assure good representation in training.\n\nAny files that have \"secondary labels\" (multiple identified bird calls) in train_metadata.csv are removed from training.  The code is there to remove low-quality ratings - but it's not enabled.\n\nThere is also some noise / contrast / brightness modification done to further augment the data.  I'm not doing any rotating / zoom - as I don't think it would help with spectrograms.\n\n**Running**\nhttps://www.kaggle.com/richolson/birdclef-2024-run-v2-efficientnetv2-b2\n\nThe run notebook generates 10-second 224x224 spectrograms from the test data.  The model then predicts on the images - and duplicates each answer over 2 x 5-second prediction windows.\n\nThis process takes about 5.5 seconds per 4-minute file.  The maximum allowed for 2 hours would be about 6.5 seconds.\n\n**Thoughts on further development**\nThe current \"run\" notebook doesn't implement batching - so the CPU isn't that well utilized.  I did a test-run with batching implemented - and prediction time is more like 4-seconds per file.  So - I think there is opportunity to upgrade to a model with more parameters.\n\nAlternately - it might be possible to use a slightly-smaller model - and predict on each 5-second interval.  Maybe still training the model on 10-second samples - but overlapping the predictions somehow.\n\nSome effort to objectively quantify the importance of prediction frequency in this competition might be interesting...\n\nThe model is probably a little overfit (no validation improvement the last several epochs) - but I don't think that hurts performance much.  Implementing early-stopping would probably be a good idea.  It may also be adjustment of learning rate / data augmentation settings might help.\n\nI current don't do anything with the geotagging of samples provided in the metadata.  An interesting experiment might be training the model -only- only on samples from the targeted Western Ghats - and then doing another model trained excluding any samples from the Western Ghats. \n\nHope this is useful to someone!\n\n-Rich",
      "votes": null
    },
    {
      "id": "2758053",
      "postDate": "04/17/2024 22:44:12",
      "content": "<p>It's   help</p>",
      "rawMarkdown": "It's   help",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2758053,
      "author_name": "",
      "author_url": "",
      "post_date": "04/17/2024 22:44:12",
      "content": "<p>It's   help</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2754615": "Sharing my updated Spectrograms -> ImageNet toolchain.\n\nThere are a bunch of changes from [my prior effort](https://www.kaggle.com/competitions/birdclef-2024/discussion/492135) - that collectively got me from .57 up to .61 on the LB.\n\nhttps://www.kaggle.com/code/richolson/birdclef-2024-contiguous-mel-spectrogram-generator\nhttps://www.kaggle.com/code/richolson/birdclef-2024-train-v2-efficientnetv2-b2\nhttps://www.kaggle.com/code/richolson/birdclef-2024-run-v2-efficientnetv2-b2\n\nFeedback / questions are welcome here or on the notebooks!\n\n**Generating Spectrograms**\nhttps://www.kaggle.com/richolson/birdclef-2024-contiguous-mel-spectrogram-generator\n\nSpectrograms are now generated \"contiguously\" (if that's the right word).\n\nAll files are 512-pixels-wide-per-5-second-interval and 224 high.  For example - a 50-second OGG file produces a PNG that is 5120x224.  I wasn't sure what time resolution I wanted - so I went relatively high horizontal resolution with the plan of scaling-down.\n\nI am also now generating them monochrome.  All files are processed - producing about 4GB of data.\n\n(I got some weird errors trying to generate a Dataset - so am just leaving this as Notebook output.  If so inclined - feel free to generate a Dataset from this / publish.)\n\n**Training**\nhttps://www.kaggle.com/code/richolson/birdclef-2024-train-v2-efficientnetv2-b2\n\nefficientnetv2_b2 is trained on random 10-second segments of spectrograms.  This assures the model gets trained with different images each epoch.  A 1024x224 image is randomly cropped from the full spectrogram image - and then scaled down to 224x224 for feeding into the imagenet.\n\nThe model is fully enabled for training / no frozen layers.\n\nI went with 10-second segments because it seemed like too-frequently 5-second segments didn't include an obvious bird call.\n\nThe maximum training files per class is capped at 3x the median for all classes.  This is done to help reduce over-representation of classes where there is a lot of data.\n\nSimilarly - if a class has under a 30% file count of the median - references to files for the class are duplicated to help assure good representation in training.\n\nAny files that have \"secondary labels\" (multiple identified bird calls) in train_metadata.csv are removed from training.  The code is there to remove low-quality ratings - but it's not enabled.\n\nThere is also some noise / contrast / brightness modification done to further augment the data.  I'm not doing any rotating / zoom - as I don't think it would help with spectrograms.\n\n**Running**\nhttps://www.kaggle.com/richolson/birdclef-2024-run-v2-efficientnetv2-b2\n\nThe run notebook generates 10-second 224x224 spectrograms from the test data.  The model then predicts on the images - and duplicates each answer over 2 x 5-second prediction windows.\n\nThis process takes about 5.5 seconds per 4-minute file.  The maximum allowed for 2 hours would be about 6.5 seconds.\n\n**Thoughts on further development**\nThe current \"run\" notebook doesn't implement batching - so the CPU isn't that well utilized.  I did a test-run with batching implemented - and prediction time is more like 4-seconds per file.  So - I think there is opportunity to upgrade to a model with more parameters.\n\nAlternately - it might be possible to use a slightly-smaller model - and predict on each 5-second interval.  Maybe still training the model on 10-second samples - but overlapping the predictions somehow.\n\nSome effort to objectively quantify the importance of prediction frequency in this competition might be interesting...\n\nThe model is probably a little overfit (no validation improvement the last several epochs) - but I don't think that hurts performance much.  Implementing early-stopping would probably be a good idea.  It may also be adjustment of learning rate / data augmentation settings might help.\n\nI current don't do anything with the geotagging of samples provided in the metadata.  An interesting experiment might be training the model -only- only on samples from the targeted Western Ghats - and then doing another model trained excluding any samples from the Western Ghats. \n\nHope this is useful to someone!\n\n-Rich",
    "2758053": "It's   help"
  },
  "source": "meta"
}