{
  "id": 178144,
  "title": "Domain Shifting",
  "url": "/competitions/birdsong-recognition/discussion/178144",
  "author_name": "",
  "post_date": "2020-08-28T18:59:06.121666800Z",
  "votes": 4,
  "comment_count": 7,
  "views": 0,
  "content": "<p>All you need to know about this competition:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2F972c6101917c46a82d6c177fc5e6350c%2FTopic.png?generation=1598641035034500&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li><strong>test set</strong>: only top5 birds from <em>example_test_audio</em> recordings (20 slices with 5s length)</li>\n<li><strong>train/val set</strong>:  only top5 from training set (400 audio clips) splitted with StratifiedKFold</li>\n<li><strong>model</strong>: tiny MobileNetV1</li>\n</ul>",
  "messages": [
    {
      "id": "989398",
      "postDate": "08/28/2020 18:59:06",
      "content": "<p>All you need to know about this competition:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2F972c6101917c46a82d6c177fc5e6350c%2FTopic.png?generation=1598641035034500&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li><strong>test set</strong>: only top5 birds from <em>example_test_audio</em> recordings (20 slices with 5s length)</li>\n<li><strong>train/val set</strong>:  only top5 from training set (400 audio clips) splitted with StratifiedKFold</li>\n<li><strong>model</strong>: tiny MobileNetV1</li>\n</ul>",
      "rawMarkdown": "All you need to know about this competition:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2F972c6101917c46a82d6c177fc5e6350c%2FTopic.png?generation=1598641035034500&alt=media)\n\n- **test set**: only top5 birds from *example_test_audio* recordings (20 slices with 5s length)\n- **train/val set**:  only top5 from training set (400 audio clips) splitted with StratifiedKFold\n- **model**: tiny MobileNetV1",
      "votes": null
    },
    {
      "id": "989413",
      "postDate": "08/28/2020 19:10:17",
      "content": "<p>That's a good insight. Thank you </p>",
      "rawMarkdown": "That's a good insight. Thank you",
      "votes": null
    },
    {
      "id": "989417",
      "postDate": "08/28/2020 19:11:53",
      "content": "<p>You're welcome!</p>",
      "rawMarkdown": "You're welcome!",
      "votes": null
    },
    {
      "id": "989426",
      "postDate": "08/28/2020 19:15:48",
      "content": "<p>Honestly speaking it ends up that way (but domain shifting still remains):</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2F35886ca90a7f265b2748f99b1bff318c%2FTopic2.png?generation=1598642136879456&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Honestly speaking it ends up that way (but domain shifting still remains):\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2F35886ca90a7f265b2748f99b1bff318c%2FTopic2.png?generation=1598642136879456&alt=media)",
      "votes": null
    },
    {
      "id": "989497",
      "postDate": "08/28/2020 20:57:25",
      "content": "<p>That looks very interesting, but I don't quite follow. This tells us that models can perform well on the training set but poorly on the example test set. I suppose the implication is that the domain of the validation set is very different from the domain of the test set? Is that all?</p>",
      "rawMarkdown": "That looks very interesting, but I don't quite follow. This tells us that models can perform well on the training set but poorly on the example test set. I suppose the implication is that the domain of the validation set is very different from the domain of the test set? Is that all?",
      "votes": null
    },
    {
      "id": "989549",
      "postDate": "08/28/2020 22:38:32",
      "content": "<p><a href=\"https://www.kaggle.com/koza4ukdmitrij\" target=\"_blank\">@koza4ukdmitrij</a>  Can you please tell me how you trained this? Is it on xenocanto, validated on the example test set? And you are seeing that F1 score only reaches its best values for validation after training for like 100 epochs??</p>",
      "rawMarkdown": "koza4ukdmitrij  Can you please tell me how you trained this? Is it on xenocanto, validated on the example test set? And you are seeing that F1 score only reaches its best values for validation after training for like 100 epochs??",
      "votes": null
    },
    {
      "id": "990179",
      "postDate": "08/29/2020 12:08:16",
      "content": "<p>Yes, you're right. </p>",
      "rawMarkdown": "Yes, you're right.",
      "votes": null
    },
    {
      "id": "990182",
      "postDate": "08/29/2020 12:10:37",
      "content": "<p>I took available audio for top5 labels and splitted it into train and validation part. In the inference time I plotted metrics not only for validation set but for test set too (2 available test audio splitted on slides with slides only from top5 label). I trained model in absolutely common way from one of the public notebooks. That was surprisingly for me too that I've got 0.4 result for this tiny task only in the end of the training. </p>",
      "rawMarkdown": "I took available audio for top5 labels and splitted it into train and validation part. In the inference time I plotted metrics not only for validation set but for test set too (2 available test audio splitted on slides with slides only from top5 label). I trained model in absolutely common way from one of the public notebooks. That was surprisingly for me too that I've got 0.4 result for this tiny task only in the end of the training.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 989413,
      "author_name": "sukanthen",
      "author_url": "",
      "post_date": "08/28/2020 19:10:17",
      "content": "<p>That's a good insight. Thank you </p>",
      "votes": null,
      "replies": [
        {
          "id": 989417,
          "author_name": "koza4ukdmitrij",
          "author_url": "",
          "post_date": "08/28/2020 19:11:53",
          "content": "<p>You're welcome!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 989426,
      "author_name": "koza4ukdmitrij",
      "author_url": "",
      "post_date": "08/28/2020 19:15:48",
      "content": "<p>Honestly speaking it ends up that way (but domain shifting still remains):</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2F35886ca90a7f265b2748f99b1bff318c%2FTopic2.png?generation=1598642136879456&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 989549,
          "author_name": "returnofsputnik",
          "author_url": "",
          "post_date": "08/28/2020 22:38:32",
          "content": "<p><a href=\"https://www.kaggle.com/koza4ukdmitrij\" target=\"_blank\">@koza4ukdmitrij</a>  Can you please tell me how you trained this? Is it on xenocanto, validated on the example test set? And you are seeing that F1 score only reaches its best values for validation after training for like 100 epochs??</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 990182,
          "author_name": "koza4ukdmitrij",
          "author_url": "",
          "post_date": "08/29/2020 12:10:37",
          "content": "<p>I took available audio for top5 labels and splitted it into train and validation part. In the inference time I plotted metrics not only for validation set but for test set too (2 available test audio splitted on slides with slides only from top5 label). I trained model in absolutely common way from one of the public notebooks. That was surprisingly for me too that I've got 0.4 result for this tiny task only in the end of the training. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 989497,
      "author_name": "lewington",
      "author_url": "",
      "post_date": "08/28/2020 20:57:25",
      "content": "<p>That looks very interesting, but I don't quite follow. This tells us that models can perform well on the training set but poorly on the example test set. I suppose the implication is that the domain of the validation set is very different from the domain of the test set? Is that all?</p>",
      "votes": null,
      "replies": [
        {
          "id": 990179,
          "author_name": "koza4ukdmitrij",
          "author_url": "",
          "post_date": "08/29/2020 12:08:16",
          "content": "<p>Yes, you're right. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "989398": "All you need to know about this competition:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2F972c6101917c46a82d6c177fc5e6350c%2FTopic.png?generation=1598641035034500&alt=media)\n\n- **test set**: only top5 birds from *example_test_audio* recordings (20 slices with 5s length)\n- **train/val set**:  only top5 from training set (400 audio clips) splitted with StratifiedKFold\n- **model**: tiny MobileNetV1",
    "989413": "That's a good insight. Thank you",
    "989417": "You're welcome!",
    "989426": "Honestly speaking it ends up that way (but domain shifting still remains):\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2F35886ca90a7f265b2748f99b1bff318c%2FTopic2.png?generation=1598642136879456&alt=media)",
    "989497": "That looks very interesting, but I don't quite follow. This tells us that models can perform well on the training set but poorly on the example test set. I suppose the implication is that the domain of the validation set is very different from the domain of the test set? Is that all?",
    "989549": "koza4ukdmitrij  Can you please tell me how you trained this? Is it on xenocanto, validated on the example test set? And you are seeing that F1 score only reaches its best values for validation after training for like 100 epochs??",
    "990179": "Yes, you're right.",
    "990182": "I took available audio for top5 labels and splitted it into train and validation part. In the inference time I plotted metrics not only for validation set but for test set too (2 available test audio splitted on slides with slides only from top5 label). I trained model in absolutely common way from one of the public notebooks. That was surprisingly for me too that I've got 0.4 result for this tiny task only in the end of the training."
  },
  "source": "meta"
}