{
  "id": 176490,
  "title": "SED? That just sounds like Audio Tagging with extra steps",
  "url": "/competitions/birdsong-recognition/discussion/176490",
  "author_name": "",
  "post_date": "2020-08-22T00:10:34.253763400Z",
  "votes": 11,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I just finished reading Hidehisa's <a href=\"https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection/data?\" target=\"_blank\">most recent excellent kernel</a>, and I'm just wondering what the advantage of Sound Event Detection might be over simple clip-wise predictions in this competition.</p>\n<p>Besides being more complex and parameter-hungry, a SED model has to do a bunch of additional learning to home-in on when events begin and end, which, if I've read Hidehisa's code correctly, more or less just gets sampled back down to the clip-wise level in the end anyway. </p>\n<p>What's the supposed advantage exactly? </p>\n<p>The only possibility in my mind is that it is somehow easier to train SED models, but since we only have access to clip-wise labels the training benefits can't be straightforward. I personally can't think of anything that has a good chance of improving over training a clipwise network. </p>\n<p>Is the <em>whole</em> idea behind SED <em>just</em> to get access to <a href=\"https://github.com/qiuqiangkong/audioset_tagging_cnn/\" target=\"_blank\">juicy juicy pretrained SED</a> networks?  Because that's what it looks like from where I'm standing.</p>",
  "messages": [
    {
      "id": "980869",
      "postDate": "08/22/2020 00:10:34",
      "content": "<p>I just finished reading Hidehisa's <a href=\"https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection/data?\" target=\"_blank\">most recent excellent kernel</a>, and I'm just wondering what the advantage of Sound Event Detection might be over simple clip-wise predictions in this competition.</p>\n<p>Besides being more complex and parameter-hungry, a SED model has to do a bunch of additional learning to home-in on when events begin and end, which, if I've read Hidehisa's code correctly, more or less just gets sampled back down to the clip-wise level in the end anyway. </p>\n<p>What's the supposed advantage exactly? </p>\n<p>The only possibility in my mind is that it is somehow easier to train SED models, but since we only have access to clip-wise labels the training benefits can't be straightforward. I personally can't think of anything that has a good chance of improving over training a clipwise network. </p>\n<p>Is the <em>whole</em> idea behind SED <em>just</em> to get access to <a href=\"https://github.com/qiuqiangkong/audioset_tagging_cnn/\" target=\"_blank\">juicy juicy pretrained SED</a> networks?  Because that's what it looks like from where I'm standing.</p>",
      "rawMarkdown": "I just finished reading Hidehisa's [most recent excellent kernel](https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection/data?), and I'm just wondering what the advantage of Sound Event Detection might be over simple clip-wise predictions in this competition.\n\nBesides being more complex and parameter-hungry, a SED model has to do a bunch of additional learning to home-in on when events begin and end, which, if I've read Hidehisa's code correctly, more or less just gets sampled back down to the clip-wise level in the end anyway. \n\nWhat's the supposed advantage exactly? \n\nThe only possibility in my mind is that it is somehow easier to train SED models, but since we only have access to clip-wise labels the training benefits can't be straightforward. I personally can't think of anything that has a good chance of improving over training a clipwise network. \n\nIs the *whole* idea behind SED *just* to get access to [juicy juicy pretrained SED](https://github.com/qiuqiangkong/audioset_tagging_cnn/) networks?  Because that's what it looks like from where I'm standing.",
      "votes": null
    },
    {
      "id": "982061",
      "postDate": "08/23/2020 02:46:30",
      "content": "<p>To get clipwise prediction, I don't think SED model does better than CNN encoder + GlobalMaxPool like aggregation + Classification head architecture (which is usually used for audio tagging).</p>\n<p>What I'd like to tell you in that notebook is there's also a chance to get strong label (time interval label) from only weak label (clipwise label). Obtaining strong label is beneficial in some ways that:</p>\n<ol>\n<li>we can use the metric used for test dataset (it requires strong label)</li>\n<li>we can smartly sample from audio clip</li>\n</ol>\n<blockquote>\n  <p>Is the whole idea behind SED just to get access to juicy juicy pretrained SED networks? Because that's what it looks like from where I'm standing.</p>\n</blockquote>\n<p>After some experiments, I conclude that pretrained PANNs models do not work better than ImageNet pretrained models like ResNet, EfficientNet, etc. Their expressiveness is almost the same for this dataset. Using pretrained model is just a start line, but what matters is good training methods.</p>",
      "rawMarkdown": "To get clipwise prediction, I don't think SED model does better than CNN encoder + GlobalMaxPool like aggregation + Classification head architecture (which is usually used for audio tagging).\n\nWhat I'd like to tell you in that notebook is there's also a chance to get strong label (time interval label) from only weak label (clipwise label). Obtaining strong label is beneficial in some ways that:\n\n1. we can use the metric used for test dataset (it requires strong label)\n2. we can smartly sample from audio clip\n\n> Is the whole idea behind SED just to get access to juicy juicy pretrained SED networks? Because that's what it looks like from where I'm standing.\n\nAfter some experiments, I conclude that pretrained PANNs models do not work better than ImageNet pretrained models like ResNet, EfficientNet, etc. Their expressiveness is almost the same for this dataset. Using pretrained model is just a start line, but what matters is good training methods.",
      "votes": null
    },
    {
      "id": "983371",
      "postDate": "08/24/2020 08:39:38",
      "content": "<p>I think you are right. I've tried ResNeSt and PANNs model and they worked exactly the same on public LB. But codes of learning strong labels from weak label are really impressive. </p>\n<p>However, I feel a bit blurry when it comes to \"good training methods\". Training methods stands for model arch or training set ups like optimizer&amp;scheduler or callbacks or model ensemble or the mix of them? In fact I've tried many methods but mostly I'm just altering the params&amp;functions&amp;set ups in the notebook. I'll be really thankful if you can help me clarify this conception a bit.</p>",
      "rawMarkdown": "I think you are right. I've tried ResNeSt and PANNs model and they worked exactly the same on public LB. But codes of learning strong labels from weak label are really impressive. \n\nHowever, I feel a bit blurry when it comes to \"good training methods\". Training methods stands for model arch or training set ups like optimizer&scheduler or callbacks or model ensemble or the mix of them? In fact I've tried many methods but mostly I'm just altering the params&functions&set ups in the notebook. I'll be really thankful if you can help me clarify this conception a bit.",
      "votes": null
    },
    {
      "id": "984065",
      "postDate": "08/24/2020 19:56:20",
      "content": "<p>Thanks for taking the time to answer Hidehisha, and to answer in English as well! It must be very hard explaining all this in a second language. </p>\n<blockquote>\n  <p>we can smartly sample from audio clip</p>\n</blockquote>\n<p>This makes sense, you're saying that if we have a good SED model, we can use it to do clever things with the data,  a good clip-wise model isn't as useful right? But</p>\n<blockquote>\n  <p>we can use the metric used for test dataset (it requires strong label)</p>\n</blockquote>\n<p>I don't understand this though, site 3 has been labelled VERY weakly, and we don't know how sites 1 and 2 were actually labelled. They might have been labelled strongly (SED-style) and then split up, or maybe they were split up first and THEN labelled weakly. We don't know which one right?</p>",
      "rawMarkdown": "Thanks for taking the time to answer Hidehisha, and to answer in English as well! It must be very hard explaining all this in a second language. \n\n> we can smartly sample from audio clip\n\nThis makes sense, you're saying that if we have a good SED model, we can use it to do clever things with the data,  a good clip-wise model isn't as useful right? But\n\n> we can use the metric used for test dataset (it requires strong label)\n\nI don't understand this though, site 3 has been labelled VERY weakly, and we don't know how sites 1 and 2 were actually labelled. They might have been labelled strongly (SED-style) and then split up, or maybe they were split up first and THEN labelled weakly. We don't know which one right?",
      "votes": null
    },
    {
      "id": "984173",
      "postDate": "08/24/2020 22:56:24",
      "content": "<blockquote>\n  <p>However, I feel a bit blurry when it comes to \"good training methods\". Training methods stands for model arch or training set ups like optimizer&amp;scheduler or callbacks or model ensemble or the mix of them? In fact I've tried many methods but mostly I'm just altering the params&amp;functions&amp;set ups in the notebook. I'll be really thankful if you can help me clarify this conception a bit.</p>\n</blockquote>\n<p>It's basically about the label - how we treat the given label. I cannot tell you in more detail now since that's the key I found to get my current position.</p>",
      "rawMarkdown": "> However, I feel a bit blurry when it comes to \"good training methods\". Training methods stands for model arch or training set ups like optimizer&scheduler or callbacks or model ensemble or the mix of them? In fact I've tried many methods but mostly I'm just altering the params&functions&set ups in the notebook. I'll be really thankful if you can help me clarify this conception a bit.\n\nIt's basically about the label - how we treat the given label. I cannot tell you in more detail now since that's the key I found to get my current position.",
      "votes": null
    },
    {
      "id": "984177",
      "postDate": "08/24/2020 23:06:06",
      "content": "<blockquote>\n  <p>and to answer in English as well! It must be very hard explaining all this in a second language.</p>\n</blockquote>\n<p>Do you have any idea how someone feel when you tell them a thing like this? It sounds a bit offensive to me.</p>\n<blockquote>\n  <p>and we don't know how sites 1 and 2 were actually labelled. They might have been labelled strongly (SED-style) and then split up, or maybe they were split up first and THEN labelled weakly. We don't know which one right?</p>\n</blockquote>\n<p>No, we don't, but at least we need 5 sec chunk level annotation. We don't have it either for train dataset. </p>",
      "rawMarkdown": "> and to answer in English as well! It must be very hard explaining all this in a second language.\n\nDo you have any idea how someone feel when you tell them a thing like this? It sounds a bit offensive to me.\n\n> and we don't know how sites 1 and 2 were actually labelled. They might have been labelled strongly (SED-style) and then split up, or maybe they were split up first and THEN labelled weakly. We don't know which one right?\n\nNo, we don't, but at least we need 5 sec chunk level annotation. We don't have it either for train dataset.",
      "votes": null
    },
    {
      "id": "984235",
      "postDate": "08/25/2020 01:22:40",
      "content": "<p>I understand. Thank you for clarify my doubt.</p>",
      "rawMarkdown": "I understand. Thank you for clarify my doubt.",
      "votes": null
    },
    {
      "id": "985665",
      "postDate": "08/25/2020 23:17:34",
      "content": "<p><a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> I was trying to compliment you on being so fluent in English, sorry if I phrased it badly! I think it's very impressive, I would have no chance talking about machine learning in Japanese, let alone guiding a whole competition.</p>\n<blockquote>\n  <p>No, we don't, but at least we need 5 sec chunk level annotation. We don't have it either for train dataset. </p>\n</blockquote>\n<p>Ok that makes perfect sense to me, you could easily be right. I still think that it's possible to train on longer clips and then transfer to shorter clips, but I guess we'll find out depending on my lb score.</p>",
      "rawMarkdown": "hidehisaarai1213 I was trying to compliment you on being so fluent in English, sorry if I phrased it badly! I think it's very impressive, I would have no chance talking about machine learning in Japanese, let alone guiding a whole competition.\n\n> No, we don't, but at least we need 5 sec chunk level annotation. We don't have it either for train dataset. \n\nOk that makes perfect sense to me, you could easily be right. I still think that it's possible to train on longer clips and then transfer to shorter clips, but I guess we'll find out depending on my lb score.",
      "votes": null
    },
    {
      "id": "986858",
      "postDate": "08/26/2020 20:55:55",
      "content": "<p>People still seem to be downvoting… if I've said something mean please let me know what it is. The last thing I want to do is go around offending people inadvertently. </p>",
      "rawMarkdown": "People still seem to be downvoting... if I've said something mean please let me know what it is. The last thing I want to do is go around offending people inadvertently.",
      "votes": null
    },
    {
      "id": "987008",
      "postDate": "08/26/2020 22:56:09",
      "content": "<p>Writing about English specifically in a comment made me feel a little bit weird and uneasy that I might have made some stupid mistakes which native English speakers would not make. 'cause you won't tell such a thing to native English speaker, right?</p>\n<p>Well it's ok, I don't care any more.</p>\n<blockquote>\n  <p>I still think that it's possible to train on longer clips and then transfer to shorter clips, </p>\n</blockquote>\n<p>Yes, I also think this is also possible. What I've introduced in the notebook was a SED model which uses Attention pooling but there's also SED model variants which use Max pooling or Average pooling. These are essentially the same as usual audio tagging model, but they works. So I think it's worth trying.</p>",
      "rawMarkdown": "Writing about English specifically in a comment made me feel a little bit weird and uneasy that I might have made some stupid mistakes which native English speakers would not make. 'cause you won't tell such a thing to native English speaker, right?\n\nWell it's ok, I don't care any more.\n\n>  I still think that it's possible to train on longer clips and then transfer to shorter clips, \n\nYes, I also think this is also possible. What I've introduced in the notebook was a SED model which uses Attention pooling but there's also SED model variants which use Max pooling or Average pooling. These are essentially the same as usual audio tagging model, but they works. So I think it's worth trying.",
      "votes": null
    },
    {
      "id": "987221",
      "postDate": "08/27/2020 05:12:35",
      "content": "<p>Yes, I fully intend to use SED once I get ANY kind of model working properly haha. Thanks for clarifying.</p>",
      "rawMarkdown": "Yes, I fully intend to use SED once I get ANY kind of model working properly haha. Thanks for clarifying.",
      "votes": null
    },
    {
      "id": "992035",
      "postDate": "08/30/2020 20:44:39",
      "content": "<p>Arai San, you share a lot, this is impressive.  I hope you'll still win this despite sharing so much.</p>",
      "rawMarkdown": "Arai San, you share a lot, this is impressive.  I hope you'll still win this despite sharing so much.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 982061,
      "author_name": "hidehisaarai1213",
      "author_url": "",
      "post_date": "08/23/2020 02:46:30",
      "content": "<p>To get clipwise prediction, I don't think SED model does better than CNN encoder + GlobalMaxPool like aggregation + Classification head architecture (which is usually used for audio tagging).</p>\n<p>What I'd like to tell you in that notebook is there's also a chance to get strong label (time interval label) from only weak label (clipwise label). Obtaining strong label is beneficial in some ways that:</p>\n<ol>\n<li>we can use the metric used for test dataset (it requires strong label)</li>\n<li>we can smartly sample from audio clip</li>\n</ol>\n<blockquote>\n  <p>Is the whole idea behind SED just to get access to juicy juicy pretrained SED networks? Because that's what it looks like from where I'm standing.</p>\n</blockquote>\n<p>After some experiments, I conclude that pretrained PANNs models do not work better than ImageNet pretrained models like ResNet, EfficientNet, etc. Their expressiveness is almost the same for this dataset. Using pretrained model is just a start line, but what matters is good training methods.</p>",
      "votes": null,
      "replies": [
        {
          "id": 983371,
          "author_name": "hzhaobang",
          "author_url": "",
          "post_date": "08/24/2020 08:39:38",
          "content": "<p>I think you are right. I've tried ResNeSt and PANNs model and they worked exactly the same on public LB. But codes of learning strong labels from weak label are really impressive. </p>\n<p>However, I feel a bit blurry when it comes to \"good training methods\". Training methods stands for model arch or training set ups like optimizer&amp;scheduler or callbacks or model ensemble or the mix of them? In fact I've tried many methods but mostly I'm just altering the params&amp;functions&amp;set ups in the notebook. I'll be really thankful if you can help me clarify this conception a bit.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 984065,
          "author_name": "lewington",
          "author_url": "",
          "post_date": "08/24/2020 19:56:20",
          "content": "<p>Thanks for taking the time to answer Hidehisha, and to answer in English as well! It must be very hard explaining all this in a second language. </p>\n<blockquote>\n  <p>we can smartly sample from audio clip</p>\n</blockquote>\n<p>This makes sense, you're saying that if we have a good SED model, we can use it to do clever things with the data,  a good clip-wise model isn't as useful right? But</p>\n<blockquote>\n  <p>we can use the metric used for test dataset (it requires strong label)</p>\n</blockquote>\n<p>I don't understand this though, site 3 has been labelled VERY weakly, and we don't know how sites 1 and 2 were actually labelled. They might have been labelled strongly (SED-style) and then split up, or maybe they were split up first and THEN labelled weakly. We don't know which one right?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 984173,
          "author_name": "hidehisaarai1213",
          "author_url": "",
          "post_date": "08/24/2020 22:56:24",
          "content": "<blockquote>\n  <p>However, I feel a bit blurry when it comes to \"good training methods\". Training methods stands for model arch or training set ups like optimizer&amp;scheduler or callbacks or model ensemble or the mix of them? In fact I've tried many methods but mostly I'm just altering the params&amp;functions&amp;set ups in the notebook. I'll be really thankful if you can help me clarify this conception a bit.</p>\n</blockquote>\n<p>It's basically about the label - how we treat the given label. I cannot tell you in more detail now since that's the key I found to get my current position.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 984177,
          "author_name": "hidehisaarai1213",
          "author_url": "",
          "post_date": "08/24/2020 23:06:06",
          "content": "<blockquote>\n  <p>and to answer in English as well! It must be very hard explaining all this in a second language.</p>\n</blockquote>\n<p>Do you have any idea how someone feel when you tell them a thing like this? It sounds a bit offensive to me.</p>\n<blockquote>\n  <p>and we don't know how sites 1 and 2 were actually labelled. They might have been labelled strongly (SED-style) and then split up, or maybe they were split up first and THEN labelled weakly. We don't know which one right?</p>\n</blockquote>\n<p>No, we don't, but at least we need 5 sec chunk level annotation. We don't have it either for train dataset. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 984235,
          "author_name": "hzhaobang",
          "author_url": "",
          "post_date": "08/25/2020 01:22:40",
          "content": "<p>I understand. Thank you for clarify my doubt.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 985665,
          "author_name": "lewington",
          "author_url": "",
          "post_date": "08/25/2020 23:17:34",
          "content": "<p><a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> I was trying to compliment you on being so fluent in English, sorry if I phrased it badly! I think it's very impressive, I would have no chance talking about machine learning in Japanese, let alone guiding a whole competition.</p>\n<blockquote>\n  <p>No, we don't, but at least we need 5 sec chunk level annotation. We don't have it either for train dataset. </p>\n</blockquote>\n<p>Ok that makes perfect sense to me, you could easily be right. I still think that it's possible to train on longer clips and then transfer to shorter clips, but I guess we'll find out depending on my lb score.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 986858,
          "author_name": "lewington",
          "author_url": "",
          "post_date": "08/26/2020 20:55:55",
          "content": "<p>People still seem to be downvoting… if I've said something mean please let me know what it is. The last thing I want to do is go around offending people inadvertently. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 987008,
          "author_name": "hidehisaarai1213",
          "author_url": "",
          "post_date": "08/26/2020 22:56:09",
          "content": "<p>Writing about English specifically in a comment made me feel a little bit weird and uneasy that I might have made some stupid mistakes which native English speakers would not make. 'cause you won't tell such a thing to native English speaker, right?</p>\n<p>Well it's ok, I don't care any more.</p>\n<blockquote>\n  <p>I still think that it's possible to train on longer clips and then transfer to shorter clips, </p>\n</blockquote>\n<p>Yes, I also think this is also possible. What I've introduced in the notebook was a SED model which uses Attention pooling but there's also SED model variants which use Max pooling or Average pooling. These are essentially the same as usual audio tagging model, but they works. So I think it's worth trying.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 987221,
          "author_name": "lewington",
          "author_url": "",
          "post_date": "08/27/2020 05:12:35",
          "content": "<p>Yes, I fully intend to use SED once I get ANY kind of model working properly haha. Thanks for clarifying.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 992035,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/30/2020 20:44:39",
          "content": "<p>Arai San, you share a lot, this is impressive.  I hope you'll still win this despite sharing so much.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "980869": "I just finished reading Hidehisa's [most recent excellent kernel](https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection/data?), and I'm just wondering what the advantage of Sound Event Detection might be over simple clip-wise predictions in this competition.\n\nBesides being more complex and parameter-hungry, a SED model has to do a bunch of additional learning to home-in on when events begin and end, which, if I've read Hidehisa's code correctly, more or less just gets sampled back down to the clip-wise level in the end anyway. \n\nWhat's the supposed advantage exactly? \n\nThe only possibility in my mind is that it is somehow easier to train SED models, but since we only have access to clip-wise labels the training benefits can't be straightforward. I personally can't think of anything that has a good chance of improving over training a clipwise network. \n\nIs the *whole* idea behind SED *just* to get access to [juicy juicy pretrained SED](https://github.com/qiuqiangkong/audioset_tagging_cnn/) networks?  Because that's what it looks like from where I'm standing.",
    "982061": "To get clipwise prediction, I don't think SED model does better than CNN encoder + GlobalMaxPool like aggregation + Classification head architecture (which is usually used for audio tagging).\n\nWhat I'd like to tell you in that notebook is there's also a chance to get strong label (time interval label) from only weak label (clipwise label). Obtaining strong label is beneficial in some ways that:\n\n1. we can use the metric used for test dataset (it requires strong label)\n2. we can smartly sample from audio clip\n\n> Is the whole idea behind SED just to get access to juicy juicy pretrained SED networks? Because that's what it looks like from where I'm standing.\n\nAfter some experiments, I conclude that pretrained PANNs models do not work better than ImageNet pretrained models like ResNet, EfficientNet, etc. Their expressiveness is almost the same for this dataset. Using pretrained model is just a start line, but what matters is good training methods.",
    "983371": "I think you are right. I've tried ResNeSt and PANNs model and they worked exactly the same on public LB. But codes of learning strong labels from weak label are really impressive. \n\nHowever, I feel a bit blurry when it comes to \"good training methods\". Training methods stands for model arch or training set ups like optimizer&scheduler or callbacks or model ensemble or the mix of them? In fact I've tried many methods but mostly I'm just altering the params&functions&set ups in the notebook. I'll be really thankful if you can help me clarify this conception a bit.",
    "984065": "Thanks for taking the time to answer Hidehisha, and to answer in English as well! It must be very hard explaining all this in a second language. \n\n> we can smartly sample from audio clip\n\nThis makes sense, you're saying that if we have a good SED model, we can use it to do clever things with the data,  a good clip-wise model isn't as useful right? But\n\n> we can use the metric used for test dataset (it requires strong label)\n\nI don't understand this though, site 3 has been labelled VERY weakly, and we don't know how sites 1 and 2 were actually labelled. They might have been labelled strongly (SED-style) and then split up, or maybe they were split up first and THEN labelled weakly. We don't know which one right?",
    "984173": "> However, I feel a bit blurry when it comes to \"good training methods\". Training methods stands for model arch or training set ups like optimizer&scheduler or callbacks or model ensemble or the mix of them? In fact I've tried many methods but mostly I'm just altering the params&functions&set ups in the notebook. I'll be really thankful if you can help me clarify this conception a bit.\n\nIt's basically about the label - how we treat the given label. I cannot tell you in more detail now since that's the key I found to get my current position.",
    "984177": "> and to answer in English as well! It must be very hard explaining all this in a second language.\n\nDo you have any idea how someone feel when you tell them a thing like this? It sounds a bit offensive to me.\n\n> and we don't know how sites 1 and 2 were actually labelled. They might have been labelled strongly (SED-style) and then split up, or maybe they were split up first and THEN labelled weakly. We don't know which one right?\n\nNo, we don't, but at least we need 5 sec chunk level annotation. We don't have it either for train dataset.",
    "984235": "I understand. Thank you for clarify my doubt.",
    "985665": "hidehisaarai1213 I was trying to compliment you on being so fluent in English, sorry if I phrased it badly! I think it's very impressive, I would have no chance talking about machine learning in Japanese, let alone guiding a whole competition.\n\n> No, we don't, but at least we need 5 sec chunk level annotation. We don't have it either for train dataset. \n\nOk that makes perfect sense to me, you could easily be right. I still think that it's possible to train on longer clips and then transfer to shorter clips, but I guess we'll find out depending on my lb score.",
    "986858": "People still seem to be downvoting... if I've said something mean please let me know what it is. The last thing I want to do is go around offending people inadvertently.",
    "987008": "Writing about English specifically in a comment made me feel a little bit weird and uneasy that I might have made some stupid mistakes which native English speakers would not make. 'cause you won't tell such a thing to native English speaker, right?\n\nWell it's ok, I don't care any more.\n\n>  I still think that it's possible to train on longer clips and then transfer to shorter clips, \n\nYes, I also think this is also possible. What I've introduced in the notebook was a SED model which uses Attention pooling but there's also SED model variants which use Max pooling or Average pooling. These are essentially the same as usual audio tagging model, but they works. So I think it's worth trying.",
    "987221": "Yes, I fully intend to use SED once I get ANY kind of model working properly haha. Thanks for clarifying.",
    "992035": "Arai San, you share a lot, this is impressive.  I hope you'll still win this despite sharing so much."
  },
  "source": "meta"
}