{
  "id": 168667,
  "title": "tackling domain adaptation head on is tough - what else can we do?",
  "url": "/competitions/birdsong-recognition/discussion/168667",
  "author_name": "Radek Osmulski",
  "post_date": "2020-07-21T12:22:31.323000",
  "votes": 22,
  "comment_count": 9,
  "views": 0,
  "content": "<p>The core of this competition is that train data comes from different sensors (mostly directional microphones I believe) than the test data.</p>\n\n<p>I tried a couple of techniques to address the domain shift problem (recalculating bn stats on the test set, technique outlined in this <a href=\"https://arxiv.org/abs/1409.7495\">paper</a> that I implement <a href=\"https://github.com/earthspecies/birdcall/blob/master/02k_train_on_melspectrograms_pytorch_lme_pool_all_classes_simple_minmax_domain_adaptation.ipynb\">here</a>) but not having access to the test set at train time (even if it should not be labeled) limits the options of what one can do quite a bit.</p>\n\n<p>One could try to do something at inference time, in the kernel, even doing some training using information from the now available test set, but this is tricky from a software engineering perspective and hard to say how much mileage one would get out of this (whether techniques for address domain adaptation from other domains would translate to audio, to this specific combination of recording devices).</p>\n\n<p>So how else can we improve the performance of our model(s)? Well, the other aspect probably worth taking a look at is the state of the training data - the labels are very noisy. In general, creating your own dataset is a great exercise I feel, one can learn a lot through this experience. Here I am curious whether we could get better results with better labels for the train set?</p>\n\n<p>I tried to annotate the data using the goto tool used by bioacoustic researchers, <a href=\"https://ravensoundsoftware.com/software/raven-lite/\">Raven</a>, but the process is quite cumbersome. It is great for what it was designed for - that is pixel perfect labeling on spectrograms, but for DL models I don't need this level of accuracy! The workflow I imagined was this -&gt; let me play a recording and press space when I here sound I want to label. At the end of the process I get a list of annotations that I can read into Raven or that I can train a model on. When one only needs a time stamp of when the call occurs, to within some temporal tolerance, this workflow leads to way more annotations per unit of effort. </p>\n\n<p>Anyhow, since I couldn't find such a tool, I quickly hacked an <a href=\"https://github.com/earthspecies/dl-annotator\">electron app</a>. It only has very limited capabilities for now, but I want to test drive the process, whether it makes sense.</p>\n\n<p>My first bird of choice will be the Mourning Dove ('moudov'). If any one else would like to give this a go, I am more than happy to report on progress and share the labels publicly. So far, the difference between the best performing submission and an all zero submission is so small, I wouldn't be surprised if actually this approach could be scaled to enough classes (maybe 10?) and perform competitively. We know from the organizers that one needs ~100 labels per species and this approach can enable improved sampling of the negative class. Also, in the past, multiple gold medalists (and I think also competition winners) have gotten there through labeling the train set - I participated in such an effort as well, for the fluke detection competition (I provided the model / train set for drawing a bounding box around the fluke which was used by at least one of the gold medalist teams). </p>\n\n<p>Anyhow, really curious to explore this approach as for now we haven't made too big of a jump from the all zero submission. Also, I think this approach can help address other scenarios, not only the challenge we are facing in this competition.</p>\n\n<p>As a baseline - I took my best performing model (0.565 LB) and asked it to output only 75 predictions for <code>moudov</code> on the test set. It scored 0.541. So the question now is whether a <code>moudov</code> detector trained on relabeled data can perform better than that with the same number of predictions? 🤔 The hope is the train set has enough examples of <code>moudov</code> and that the LB has enough precision for me to pick up whether this approach is working or not... no other way to find that out than try!</p>",
  "messages": [
    {
      "id": 938250,
      "postDate": "2020-07-21T12:22:31.323Z",
      "content": "<p>The core of this competition is that train data comes from different sensors (mostly directional microphones I believe) than the test data.</p>\n\n<p>I tried a couple of techniques to address the domain shift problem (recalculating bn stats on the test set, technique outlined in this <a href=\"https://arxiv.org/abs/1409.7495\">paper</a> that I implement <a href=\"https://github.com/earthspecies/birdcall/blob/master/02k_train_on_melspectrograms_pytorch_lme_pool_all_classes_simple_minmax_domain_adaptation.ipynb\">here</a>) but not having access to the test set at train time (even if it should not be labeled) limits the options of what one can do quite a bit.</p>\n\n<p>One could try to do something at inference time, in the kernel, even doing some training using information from the now available test set, but this is tricky from a software engineering perspective and hard to say how much mileage one would get out of this (whether techniques for address domain adaptation from other domains would translate to audio, to this specific combination of recording devices).</p>\n\n<p>So how else can we improve the performance of our model(s)? Well, the other aspect probably worth taking a look at is the state of the training data - the labels are very noisy. In general, creating your own dataset is a great exercise I feel, one can learn a lot through this experience. Here I am curious whether we could get better results with better labels for the train set?</p>\n\n<p>I tried to annotate the data using the goto tool used by bioacoustic researchers, <a href=\"https://ravensoundsoftware.com/software/raven-lite/\">Raven</a>, but the process is quite cumbersome. It is great for what it was designed for - that is pixel perfect labeling on spectrograms, but for DL models I don't need this level of accuracy! The workflow I imagined was this -&gt; let me play a recording and press space when I here sound I want to label. At the end of the process I get a list of annotations that I can read into Raven or that I can train a model on. When one only needs a time stamp of when the call occurs, to within some temporal tolerance, this workflow leads to way more annotations per unit of effort. </p>\n\n<p>Anyhow, since I couldn't find such a tool, I quickly hacked an <a href=\"https://github.com/earthspecies/dl-annotator\">electron app</a>. It only has very limited capabilities for now, but I want to test drive the process, whether it makes sense.</p>\n\n<p>My first bird of choice will be the Mourning Dove ('moudov'). If any one else would like to give this a go, I am more than happy to report on progress and share the labels publicly. So far, the difference between the best performing submission and an all zero submission is so small, I wouldn't be surprised if actually this approach could be scaled to enough classes (maybe 10?) and perform competitively. We know from the organizers that one needs ~100 labels per species and this approach can enable improved sampling of the negative class. Also, in the past, multiple gold medalists (and I think also competition winners) have gotten there through labeling the train set - I participated in such an effort as well, for the fluke detection competition (I provided the model / train set for drawing a bounding box around the fluke which was used by at least one of the gold medalist teams). </p>\n\n<p>Anyhow, really curious to explore this approach as for now we haven't made too big of a jump from the all zero submission. Also, I think this approach can help address other scenarios, not only the challenge we are facing in this competition.</p>\n\n<p>As a baseline - I took my best performing model (0.565 LB) and asked it to output only 75 predictions for <code>moudov</code> on the test set. It scored 0.541. So the question now is whether a <code>moudov</code> detector trained on relabeled data can perform better than that with the same number of predictions? 🤔 The hope is the train set has enough examples of <code>moudov</code> and that the LB has enough precision for me to pick up whether this approach is working or not... no other way to find that out than try!</p>",
      "rawMarkdown": "The core of this competition is that train data comes from different sensors (mostly directional microphones I believe) than the test data.\n\nI tried a couple of techniques to address the domain shift problem (recalculating bn stats on the test set, technique outlined in this [paper](https://arxiv.org/abs/1409.7495) that I implement [here](https://github.com/earthspecies/birdcall/blob/master/02k_train_on_melspectrograms_pytorch_lme_pool_all_classes_simple_minmax_domain_adaptation.ipynb)) but not having access to the test set at train time (even if it should not be labeled) limits the options of what one can do quite a bit.\n\nOne could try to do something at inference time, in the kernel, even doing some training using information from the now available test set, but this is tricky from a software engineering perspective and hard to say how much mileage one would get out of this (whether techniques for address domain adaptation from other domains would translate to audio, to this specific combination of recording devices).\n\nSo how else can we improve the performance of our model(s)? Well, the other aspect probably worth taking a look at is the state of the training data - the labels are very noisy. In general, creating your own dataset is a great exercise I feel, one can learn a lot through this experience. Here I am curious whether we could get better results with better labels for the train set?\n\nI tried to annotate the data using the goto tool used by bioacoustic researchers, [Raven](https://ravensoundsoftware.com/software/raven-lite/), but the process is quite cumbersome. It is great for what it was designed for - that is pixel perfect labeling on spectrograms, but for DL models I don't need this level of accuracy! The workflow I imagined was this -&gt; let me play a recording and press space when I here sound I want to label. At the end of the process I get a list of annotations that I can read into Raven or that I can train a model on. When one only needs a time stamp of when the call occurs, to within some temporal tolerance, this workflow leads to way more annotations per unit of effort. \n\nAnyhow, since I couldn't find such a tool, I quickly hacked an [electron app](https://github.com/earthspecies/dl-annotator). It only has very limited capabilities for now, but I want to test drive the process, whether it makes sense.\n\nMy first bird of choice will be the Mourning Dove ('moudov'). If any one else would like to give this a go, I am more than happy to report on progress and share the labels publicly. So far, the difference between the best performing submission and an all zero submission is so small, I wouldn't be surprised if actually this approach could be scaled to enough classes (maybe 10?) and perform competitively. We know from the organizers that one needs ~100 labels per species and this approach can enable improved sampling of the negative class. Also, in the past, multiple gold medalists (and I think also competition winners) have gotten there through labeling the train set - I participated in such an effort as well, for the fluke detection competition (I provided the model / train set for drawing a bounding box around the fluke which was used by at least one of the gold medalist teams). \n\nAnyhow, really curious to explore this approach as for now we haven't made too big of a jump from the all zero submission. Also, I think this approach can help address other scenarios, not only the challenge we are facing in this competition.\n\nAs a baseline - I took my best performing model (0.565 LB) and asked it to output only 75 predictions for `moudov` on the test set. It scored 0.541. So the question now is whether a `moudov` detector trained on relabeled data can perform better than that with the same number of predictions? 🤔 The hope is the train set has enough examples of `moudov` and that the LB has enough precision for me to pick up whether this approach is working or not... no other way to find that out than try!",
      "votes": 22
    },
    {
      "id": 942326,
      "postDate": "2020-07-23T17:43:15.867Z",
      "content": "<p>The previous BirdClef competitions (and my own work) indicate that augmentation is very important. The sounds in the training data are going to be (usually) cleaner than what we hear in soundscapes; the directional mics and recordist preferences are tending to pick out cleaner examples. Luckily, it's easier to dirty things up than the reverse.</p>\n\n<p>Here's a few common strategies for augmentation, <a href=\"http://ceur-ws.org/Vol-2125/paper_140.pdf\">mostly discussed here</a>:\na) Add noise at various SNR.\nb) Reduce the gain of the target signal.\nc) Mix multiple training examples to simulate the multi-label task.\nd) Apply a low-pass filter to the training example, with random cutoff and slope... Low frequency sound travels further than high frequency sound, so simply lowering gain isn't quite a realistic transform to simulate distance.\ne) Frequency/pitch shifting (either with a heavy duty algo or just shifting the spectrogram as a zero-th order approximation).</p>\n\n<p>You might also check whether your models are identifying the background species in the training set as a source of validation. Similar to soundscapes, these will be less impacted by recordist attention/preference.</p>",
      "rawMarkdown": "The previous BirdClef competitions (and my own work) indicate that augmentation is very important. The sounds in the training data are going to be (usually) cleaner than what we hear in soundscapes; the directional mics and recordist preferences are tending to pick out cleaner examples. Luckily, it's easier to dirty things up than the reverse.\n\nHere's a few common strategies for augmentation, [mostly discussed here](http://ceur-ws.org/Vol-2125/paper_140.pdf):\na) Add noise at various SNR.\nb) Reduce the gain of the target signal.\nc) Mix multiple training examples to simulate the multi-label task.\nd) Apply a low-pass filter to the training example, with random cutoff and slope... Low frequency sound travels further than high frequency sound, so simply lowering gain isn't quite a realistic transform to simulate distance.\ne) Frequency/pitch shifting (either with a heavy duty algo or just shifting the spectrogram as a zero-th order approximation).\n\nYou might also check whether your models are identifying the background species in the training set as a source of validation. Similar to soundscapes, these will be less impacted by recordist attention/preference.",
      "votes": 14
    },
    {
      "id": 938669,
      "postDate": "2020-07-21T17:01:22.593Z",
      "content": "<p>Have you inspected the output of your models on the sample test audio? After peeking at a few different kernels and inspecting my own model it is pretty apparent that the domain shift is completely breaking them. I inspected the spectrograms from train and compared to the sample test and also saw the prevalence of noise was so much higher. Like you mentioned maybe there is some way to adapt to the test or maybe there is some way to adapt the train to be more similar to the test instead.  </p>",
      "rawMarkdown": "Have you inspected the output of your models on the sample test audio? After peeking at a few different kernels and inspecting my own model it is pretty apparent that the domain shift is completely breaking them. I inspected the spectrograms from train and compared to the sample test and also saw the prevalence of noise was so much higher. Like you mentioned maybe there is some way to adapt to the test or maybe there is some way to adapt the train to be more similar to the test instead.  ",
      "votes": 3,
      "replies": [
        {
          "id": 938715,
          "postDate": "2020-07-21T17:40:33.403Z",
          "content": "<p>I am not sure if the audio files provided are actually sample test audio - AFAIU they come from a different organization than the test set. I am not sure if they are representative of what's in the test set.</p>\n\n<p>There were sample soundscape recordings shared by <a href=\"/stefankahl\">@stefankahl</a> (one of the organizers) <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/158877#911336\">here</a>. I think these could be more representative of the test set, but again hard to say. There are two sets of soundscape recordings there and they vary a lot between themselves.</p>\n\n<p>I constructed a validation set from the above soundscape recordings, but my model performs extremely poorly on them (~0.02 f1 when I look for the best threshold). The recordings do sound different and they (probably) look different on the spectrograms, not sure if I can trust my judgment on this as there is quite a bit of variability in the train set as well.</p>\n\n<p>One thing interesting that the best performing NBs from <a href=\"/hidehisaarai1213\">@hidehisaarai1213</a> and <a href=\"/ttahara\">@ttahara</a> are doing is the transformation of the spectrogram below:</p>\n\n<p>```\ndef mono_to_color(\n    X: np.ndarray, mean=None, std=None,\n    norm_max=None, norm_min=None, eps=1e-6\n):\n    # Stack X as [X,X,X]\n    X = np.stack([X, X, X], axis=-1)</p>\n\n<pre><code># Standardize\nmean = mean or X.mean()\nX = X - mean\nstd = std or X.std()\nXstd = X / (std + eps)\n_min, _max = Xstd.min(), Xstd.max()\nnorm_max = norm_max or _max\nnorm_min = norm_min or _min\nif (_max - _min) &gt; eps:\n    # Normalize to [0, 255]\n    V = Xstd\n    V[V &lt; norm_min] = norm_min\n    V[V &gt; norm_max] = norm_max\n    V = 255 * (V - norm_min) / (norm_max - norm_min)\n    V = V.astype(np.uint8)\nelse:\n    # Just zero\n    V = np.zeros_like(Xstd, dtype=np.uint8)\nreturn V\n</code></pre>\n\n<p>```</p>\n\n<p>Going to integers from 0 to 255 is in a sense completely pointless, but then I started thinking - maybe instead of having continuous floats, through this quanitization of sorts, this removes some variability from the data, simplifies it in some way? Could this explain why the public kernels perform so well?</p>\n\n<p>But so far the results I am seeing seem to suggest that this is not the case, that this additional step does not add much value. Currently testing on upscaled specs... hoping this is not what can explain the difference in performance for reasons I mention in my post above.</p>\n\n<p>To return to your point - I think you are on the right track, making the data (or representation of data say before the classifier in a cnn) seem more alike between train and test seems to me as well as the way to go. But I haven't explored any techniques to this end apart from <a href=\"https://arxiv.org/abs/1409.7495\">https://arxiv.org/abs/1409.7495</a> and recalculating BN stats before inference on test.</p>",
          "rawMarkdown": "I am not sure if the audio files provided are actually sample test audio - AFAIU they come from a different organization than the test set. I am not sure if they are representative of what's in the test set.\n\nThere were sample soundscape recordings shared by @stefankahl (one of the organizers) [here](https://www.kaggle.com/c/birdsong-recognition/discussion/158877#911336). I think these could be more representative of the test set, but again hard to say. There are two sets of soundscape recordings there and they vary a lot between themselves.\n\nI constructed a validation set from the above soundscape recordings, but my model performs extremely poorly on them (~0.02 f1 when I look for the best threshold). The recordings do sound different and they (probably) look different on the spectrograms, not sure if I can trust my judgment on this as there is quite a bit of variability in the train set as well.\n\nOne thing interesting that the best performing NBs from @hidehisaarai1213 and @ttahara are doing is the transformation of the spectrogram below:\n\n```\ndef mono_to_color(\n    X: np.ndarray, mean=None, std=None,\n    norm_max=None, norm_min=None, eps=1e-6\n):\n    # Stack X as [X,X,X]\n    X = np.stack([X, X, X], axis=-1)\n\n    # Standardize\n    mean = mean or X.mean()\n    X = X - mean\n    std = std or X.std()\n    Xstd = X / (std + eps)\n    _min, _max = Xstd.min(), Xstd.max()\n    norm_max = norm_max or _max\n    norm_min = norm_min or _min\n    if (_max - _min) &gt; eps:\n        # Normalize to [0, 255]\n        V = Xstd\n        V[V &lt; norm_min] = norm_min\n        V[V &gt; norm_max] = norm_max\n        V = 255 * (V - norm_min) / (norm_max - norm_min)\n        V = V.astype(np.uint8)\n    else:\n        # Just zero\n        V = np.zeros_like(Xstd, dtype=np.uint8)\n    return V\n```\n\nGoing to integers from 0 to 255 is in a sense completely pointless, but then I started thinking - maybe instead of having continuous floats, through this quanitization of sorts, this removes some variability from the data, simplifies it in some way? Could this explain why the public kernels perform so well?\n\nBut so far the results I am seeing seem to suggest that this is not the case, that this additional step does not add much value. Currently testing on upscaled specs... hoping this is not what can explain the difference in performance for reasons I mention in my post above.\n\nTo return to your point - I think you are on the right track, making the data (or representation of data say before the classifier in a cnn) seem more alike between train and test seems to me as well as the way to go. But I haven't explored any techniques to this end apart from https://arxiv.org/abs/1409.7495 and recalculating BN stats before inference on test.",
          "votes": 2
        },
        {
          "id": 941357,
          "postDate": "2020-07-23T07:23:39.920Z",
          "content": "<p>I've had the same experience on the sample test data. Even when removing the classes we don't have in our training data like squirrels the performance to me looks completely random. An f1 of .02 after threshold tuning to me means our model are doing random guessing. </p>\n\n<p>If someone solves that problem they will have a huge lead. </p>",
          "rawMarkdown": "I've had the same experience on the sample test data. Even when removing the classes we don't have in our training data like squirrels the performance to me looks completely random. An f1 of .02 after threshold tuning to me means our model are doing random guessing. \n\nIf someone solves that problem they will have a huge lead. ",
          "votes": 3
        },
        {
          "id": 962836,
          "postDate": "2020-08-08T13:19:47.693Z",
          "content": "<p><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> I feel that I'm responsible for explaining the mono-to-color function which I wrote for Freesound competition last year.</p>\n<p>Basic design of the function: By simply making a sound available as an image, we could simply make the best use of various image data augmentation (like Albumentations for example) without change.</p>\n<p>This would be obvious to you, but I'd like to clarify for some people here who didn't notice my intension.</p>\n<p>One more intention for the function is, it is trying to make every sound have the same basic stat, making all samples have the mean 0.0 and the std 1.0 (it's unlike what is done for usual images that are normalized according to ImageNet stats). This might help tackling domain shift if this is part of the problem.</p>\n<p>( Now I've joined this competition with few budget of time.)</p>",
          "rawMarkdown": "@radek1 I feel that I'm responsible for explaining the mono-to-color function which I wrote for Freesound competition last year.\n\nBasic design of the function: By simply making a sound available as an image, we could simply make the best use of various image data augmentation (like Albumentations for example) without change.\n\nThis would be obvious to you, but I'd like to clarify for some people here who didn't notice my intension.\n\nOne more intention for the function is, it is trying to make every sound have the same basic stat, making all samples have the mean 0.0 and the std 1.0 (it's unlike what is done for usual images that are normalized according to ImageNet stats). This might help tackling domain shift if this is part of the problem.\n\n(~~Anyway I have to say that I don't have time to listen samples in this competition so far... excuse me about that.~~ Now I've joined this competition with few budget of time.)",
          "votes": 7
        }
      ]
    },
    {
      "id": 941556,
      "postDate": "2020-07-23T09:45:10.170Z",
      "content": "<p>this is what is meant by directional microphone</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F4f731797dfc7b93b43bb24022a90450f%2Fd3c0ed08-7daa-4e95-8cb0-2b117515d52a-1020x612.jpeg?generation=1595514708157834&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fff4541bb47cb92284c958c53b8a96fe6%2Fde870a51-8bca-4090-9227-6741337be36c-2060x1236.jpeg?generation=1595514705817819&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "this is what is meant by directional microphone\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F4f731797dfc7b93b43bb24022a90450f%2Fd3c0ed08-7daa-4e95-8cb0-2b117515d52a-1020x612.jpeg?generation=1595514708157834&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fff4541bb47cb92284c958c53b8a96fe6%2Fde870a51-8bca-4090-9227-6741337be36c-2060x1236.jpeg?generation=1595514705817819&amp;alt=media)\n ",
      "votes": 4
    },
    {
      "id": 942118,
      "postDate": "2020-07-23T15:36:26.497Z",
      "content": "<p>it is real problem that real-life data exhibits domain shift from train data. This challenge attempts to solve this by:</p>\n\n<p>1) given some train data, make a robust model to perform well in some test data that is domain shifted and not seen in train.</p>\n\n<p>I think this is difficult. if we have absolutely no information about the test data, it is very difficult to ensure robustness. But i think the real problem can be solved in another way, domain transfer/few shot/etc:</p>\n\n<p>2) given some train data,  make a model. Given a little data from the test domain,  adapt model to the test easily.</p>\n\n<p>In that case, kaggle should provide a large pool of train data and some test data. This still helps to make better tools for bird monitoring as we need not annotate so much new data when there is domain shift.</p>\n\n<hr>\n\n<p>It may be interesting for kaggle to hold a competition that rank model by how little of data you need for domain adaption. e.g.\na. data for domain adaption is divided into 1% 5% 10% 15% .... etc.\nb. some baseline results is given\nc. score = (100-'% of data you use' )+ ('your score'- 'baseline score')</p>",
      "rawMarkdown": "it is real problem that real-life data exhibits domain shift from train data. This challenge attempts to solve this by:\n\n1) given some train data, make a robust model to perform well in some test data that is domain shifted and not seen in train.\n\nI think this is difficult. if we have absolutely no information about the test data, it is very difficult to ensure robustness. But i think the real problem can be solved in another way, domain transfer/few shot/etc:\n\n2) given some train data,  make a model. Given a little data from the test domain,  adapt model to the test easily.\n\nIn that case, kaggle should provide a large pool of train data and some test data. This still helps to make better tools for bird monitoring as we need not annotate so much new data when there is domain shift.\n\n---\n\n It may be interesting for kaggle to hold a competition that rank model by how little of data you need for domain adaption. e.g.\na. data for domain adaption is divided into 1% 5% 10% 15% .... etc.\nb. some baseline results is given\nc. score = (100-'% of data you use' )+ ('your score'- 'baseline score')",
      "votes": 1
    },
    {
      "id": 976090,
      "postDate": "2020-08-18T16:20:05.147Z",
      "content": "<blockquote>\n  <p>recalculating bn stats on the test set,</p>\n</blockquote>\n<p>Sorry, what's bn?</p>",
      "rawMarkdown": "> recalculating bn stats on the test set,\n\nSorry, what's bn?"
    },
    {
      "id": 939431,
      "postDate": "2020-07-22T08:36:36.070Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 942326,
      "author_name": "Tom Denton",
      "author_url": "",
      "post_date": "2020-07-23T17:43:15.867000",
      "content": "<p>The previous BirdClef competitions (and my own work) indicate that augmentation is very important. The sounds in the training data are going to be (usually) cleaner than what we hear in soundscapes; the directional mics and recordist preferences are tending to pick out cleaner examples. Luckily, it's easier to dirty things up than the reverse.</p>\n\n<p>Here's a few common strategies for augmentation, <a href=\"http://ceur-ws.org/Vol-2125/paper_140.pdf\">mostly discussed here</a>:\na) Add noise at various SNR.\nb) Reduce the gain of the target signal.\nc) Mix multiple training examples to simulate the multi-label task.\nd) Apply a low-pass filter to the training example, with random cutoff and slope... Low frequency sound travels further than high frequency sound, so simply lowering gain isn't quite a realistic transform to simulate distance.\ne) Frequency/pitch shifting (either with a heavy duty algo or just shifting the spectrogram as a zero-th order approximation).</p>\n\n<p>You might also check whether your models are identifying the background species in the training set as a source of validation. Similar to soundscapes, these will be less impacted by recordist attention/preference.</p>",
      "votes": 14,
      "replies": []
    },
    {
      "id": 938669,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "2020-07-21T17:01:22.593000",
      "content": "<p>Have you inspected the output of your models on the sample test audio? After peeking at a few different kernels and inspecting my own model it is pretty apparent that the domain shift is completely breaking them. I inspected the spectrograms from train and compared to the sample test and also saw the prevalence of noise was so much higher. Like you mentioned maybe there is some way to adapt to the test or maybe there is some way to adapt the train to be more similar to the test instead.  </p>",
      "votes": 3,
      "replies": [
        {
          "id": 938715,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2020-07-21T17:40:33.403000",
          "content": "<p>I am not sure if the audio files provided are actually sample test audio - AFAIU they come from a different organization than the test set. I am not sure if they are representative of what's in the test set.</p>\n\n<p>There were sample soundscape recordings shared by <a href=\"/stefankahl\">@stefankahl</a> (one of the organizers) <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/158877#911336\">here</a>. I think these could be more representative of the test set, but again hard to say. There are two sets of soundscape recordings there and they vary a lot between themselves.</p>\n\n<p>I constructed a validation set from the above soundscape recordings, but my model performs extremely poorly on them (~0.02 f1 when I look for the best threshold). The recordings do sound different and they (probably) look different on the spectrograms, not sure if I can trust my judgment on this as there is quite a bit of variability in the train set as well.</p>\n\n<p>One thing interesting that the best performing NBs from <a href=\"/hidehisaarai1213\">@hidehisaarai1213</a> and <a href=\"/ttahara\">@ttahara</a> are doing is the transformation of the spectrogram below:</p>\n\n<p>```\ndef mono_to_color(\n    X: np.ndarray, mean=None, std=None,\n    norm_max=None, norm_min=None, eps=1e-6\n):\n    # Stack X as [X,X,X]\n    X = np.stack([X, X, X], axis=-1)</p>\n\n<pre><code># Standardize\nmean = mean or X.mean()\nX = X - mean\nstd = std or X.std()\nXstd = X / (std + eps)\n_min, _max = Xstd.min(), Xstd.max()\nnorm_max = norm_max or _max\nnorm_min = norm_min or _min\nif (_max - _min) &gt; eps:\n    # Normalize to [0, 255]\n    V = Xstd\n    V[V &lt; norm_min] = norm_min\n    V[V &gt; norm_max] = norm_max\n    V = 255 * (V - norm_min) / (norm_max - norm_min)\n    V = V.astype(np.uint8)\nelse:\n    # Just zero\n    V = np.zeros_like(Xstd, dtype=np.uint8)\nreturn V\n</code></pre>\n\n<p>```</p>\n\n<p>Going to integers from 0 to 255 is in a sense completely pointless, but then I started thinking - maybe instead of having continuous floats, through this quanitization of sorts, this removes some variability from the data, simplifies it in some way? Could this explain why the public kernels perform so well?</p>\n\n<p>But so far the results I am seeing seem to suggest that this is not the case, that this additional step does not add much value. Currently testing on upscaled specs... hoping this is not what can explain the difference in performance for reasons I mention in my post above.</p>\n\n<p>To return to your point - I think you are on the right track, making the data (or representation of data say before the classifier in a cnn) seem more alike between train and test seems to me as well as the way to go. But I haven't explored any techniques to this end apart from <a href=\"https://arxiv.org/abs/1409.7495\">https://arxiv.org/abs/1409.7495</a> and recalculating BN stats before inference on test.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 941357,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2020-07-23T07:23:39.920000",
          "content": "<p>I've had the same experience on the sample test data. Even when removing the classes we don't have in our training data like squirrels the performance to me looks completely random. An f1 of .02 after threshold tuning to me means our model are doing random guessing. </p>\n\n<p>If someone solves that problem they will have a huge lead. </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 962836,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2020-08-08T13:19:47.693000",
          "content": "<p><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> I feel that I'm responsible for explaining the mono-to-color function which I wrote for Freesound competition last year.</p>\n<p>Basic design of the function: By simply making a sound available as an image, we could simply make the best use of various image data augmentation (like Albumentations for example) without change.</p>\n<p>This would be obvious to you, but I'd like to clarify for some people here who didn't notice my intension.</p>\n<p>One more intention for the function is, it is trying to make every sound have the same basic stat, making all samples have the mean 0.0 and the std 1.0 (it's unlike what is done for usual images that are normalized according to ImageNet stats). This might help tackling domain shift if this is part of the problem.</p>\n<p>( Now I've joined this competition with few budget of time.)</p>",
          "votes": 7,
          "replies": []
        }
      ]
    },
    {
      "id": 941556,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-07-23T09:45:10.170000",
      "content": "<p>this is what is meant by directional microphone</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F4f731797dfc7b93b43bb24022a90450f%2Fd3c0ed08-7daa-4e95-8cb0-2b117515d52a-1020x612.jpeg?generation=1595514708157834&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fff4541bb47cb92284c958c53b8a96fe6%2Fde870a51-8bca-4090-9227-6741337be36c-2060x1236.jpeg?generation=1595514705817819&amp;alt=media\" alt=\"\"></p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 942118,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-07-23T15:36:26.497000",
      "content": "<p>it is real problem that real-life data exhibits domain shift from train data. This challenge attempts to solve this by:</p>\n\n<p>1) given some train data, make a robust model to perform well in some test data that is domain shifted and not seen in train.</p>\n\n<p>I think this is difficult. if we have absolutely no information about the test data, it is very difficult to ensure robustness. But i think the real problem can be solved in another way, domain transfer/few shot/etc:</p>\n\n<p>2) given some train data,  make a model. Given a little data from the test domain,  adapt model to the test easily.</p>\n\n<p>In that case, kaggle should provide a large pool of train data and some test data. This still helps to make better tools for bird monitoring as we need not annotate so much new data when there is domain shift.</p>\n\n<hr>\n\n<p>It may be interesting for kaggle to hold a competition that rank model by how little of data you need for domain adaption. e.g.\na. data for domain adaption is divided into 1% 5% 10% 15% .... etc.\nb. some baseline results is given\nc. score = (100-'% of data you use' )+ ('your score'- 'baseline score')</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 976090,
      "author_name": "Marco Gorelli",
      "author_url": "",
      "post_date": "2020-08-18T16:20:05.147000",
      "content": "<blockquote>\n  <p>recalculating bn stats on the test set,</p>\n</blockquote>\n<p>Sorry, what's bn?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 939431,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-22T08:36:36.070000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "938250": "The core of this competition is that train data comes from different sensors (mostly directional microphones I believe) than the test data.\n\nI tried a couple of techniques to address the domain shift problem (recalculating bn stats on the test set, technique outlined in this [paper](https://arxiv.org/abs/1409.7495) that I implement [here](https://github.com/earthspecies/birdcall/blob/master/02k_train_on_melspectrograms_pytorch_lme_pool_all_classes_simple_minmax_domain_adaptation.ipynb)) but not having access to the test set at train time (even if it should not be labeled) limits the options of what one can do quite a bit.\n\nOne could try to do something at inference time, in the kernel, even doing some training using information from the now available test set, but this is tricky from a software engineering perspective and hard to say how much mileage one would get out of this (whether techniques for address domain adaptation from other domains would translate to audio, to this specific combination of recording devices).\n\nSo how else can we improve the performance of our model(s)? Well, the other aspect probably worth taking a look at is the state of the training data - the labels are very noisy. In general, creating your own dataset is a great exercise I feel, one can learn a lot through this experience. Here I am curious whether we could get better results with better labels for the train set?\n\nI tried to annotate the data using the goto tool used by bioacoustic researchers, [Raven](https://ravensoundsoftware.com/software/raven-lite/), but the process is quite cumbersome. It is great for what it was designed for - that is pixel perfect labeling on spectrograms, but for DL models I don't need this level of accuracy! The workflow I imagined was this -&gt; let me play a recording and press space when I here sound I want to label. At the end of the process I get a list of annotations that I can read into Raven or that I can train a model on. When one only needs a time stamp of when the call occurs, to within some temporal tolerance, this workflow leads to way more annotations per unit of effort. \n\nAnyhow, since I couldn't find such a tool, I quickly hacked an [electron app](https://github.com/earthspecies/dl-annotator). It only has very limited capabilities for now, but I want to test drive the process, whether it makes sense.\n\nMy first bird of choice will be the Mourning Dove ('moudov'). If any one else would like to give this a go, I am more than happy to report on progress and share the labels publicly. So far, the difference between the best performing submission and an all zero submission is so small, I wouldn't be surprised if actually this approach could be scaled to enough classes (maybe 10?) and perform competitively. We know from the organizers that one needs ~100 labels per species and this approach can enable improved sampling of the negative class. Also, in the past, multiple gold medalists (and I think also competition winners) have gotten there through labeling the train set - I participated in such an effort as well, for the fluke detection competition (I provided the model / train set for drawing a bounding box around the fluke which was used by at least one of the gold medalist teams). \n\nAnyhow, really curious to explore this approach as for now we haven't made too big of a jump from the all zero submission. Also, I think this approach can help address other scenarios, not only the challenge we are facing in this competition.\n\nAs a baseline - I took my best performing model (0.565 LB) and asked it to output only 75 predictions for `moudov` on the test set. It scored 0.541. So the question now is whether a `moudov` detector trained on relabeled data can perform better than that with the same number of predictions? 🤔 The hope is the train set has enough examples of `moudov` and that the LB has enough precision for me to pick up whether this approach is working or not... no other way to find that out than try!",
    "942326": "The previous BirdClef competitions (and my own work) indicate that augmentation is very important. The sounds in the training data are going to be (usually) cleaner than what we hear in soundscapes; the directional mics and recordist preferences are tending to pick out cleaner examples. Luckily, it's easier to dirty things up than the reverse.\n\nHere's a few common strategies for augmentation, [mostly discussed here](http://ceur-ws.org/Vol-2125/paper_140.pdf):\na) Add noise at various SNR.\nb) Reduce the gain of the target signal.\nc) Mix multiple training examples to simulate the multi-label task.\nd) Apply a low-pass filter to the training example, with random cutoff and slope... Low frequency sound travels further than high frequency sound, so simply lowering gain isn't quite a realistic transform to simulate distance.\ne) Frequency/pitch shifting (either with a heavy duty algo or just shifting the spectrogram as a zero-th order approximation).\n\nYou might also check whether your models are identifying the background species in the training set as a source of validation. Similar to soundscapes, these will be less impacted by recordist attention/preference.",
    "938669": "Have you inspected the output of your models on the sample test audio? After peeking at a few different kernels and inspecting my own model it is pretty apparent that the domain shift is completely breaking them. I inspected the spectrograms from train and compared to the sample test and also saw the prevalence of noise was so much higher. Like you mentioned maybe there is some way to adapt to the test or maybe there is some way to adapt the train to be more similar to the test instead.  ",
    "941556": "this is what is meant by directional microphone\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F4f731797dfc7b93b43bb24022a90450f%2Fd3c0ed08-7daa-4e95-8cb0-2b117515d52a-1020x612.jpeg?generation=1595514708157834&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fff4541bb47cb92284c958c53b8a96fe6%2Fde870a51-8bca-4090-9227-6741337be36c-2060x1236.jpeg?generation=1595514705817819&amp;alt=media)\n ",
    "942118": "it is real problem that real-life data exhibits domain shift from train data. This challenge attempts to solve this by:\n\n1) given some train data, make a robust model to perform well in some test data that is domain shifted and not seen in train.\n\nI think this is difficult. if we have absolutely no information about the test data, it is very difficult to ensure robustness. But i think the real problem can be solved in another way, domain transfer/few shot/etc:\n\n2) given some train data,  make a model. Given a little data from the test domain,  adapt model to the test easily.\n\nIn that case, kaggle should provide a large pool of train data and some test data. This still helps to make better tools for bird monitoring as we need not annotate so much new data when there is domain shift.\n\n---\n\n It may be interesting for kaggle to hold a competition that rank model by how little of data you need for domain adaption. e.g.\na. data for domain adaption is divided into 1% 5% 10% 15% .... etc.\nb. some baseline results is given\nc. score = (100-'% of data you use' )+ ('your score'- 'baseline score')",
    "976090": "> recalculating bn stats on the test set,\n\nSorry, what's bn?",
    "939431": ""
  }
}