{
  "id": 97815,
  "title": "2nd place Solution",
  "url": "/competitions/freesound-audio-tagging-2019/writeups/the-art-of-ensemble-2nd-place-solution",
  "author_name": "",
  "post_date": "2019-07-09T01:36:31.537Z",
  "votes": 59,
  "comment_count": 19,
  "views": 0,
  "content": "<p>Thanks for the competition and congratulations of all participants! Here is a summary of our solution.\ncode: <a href=\"https://github.com/qrfaction/1st-Freesound-Audio-Tagging-2019\">https://github.com/qrfaction/1st-Freesound-Audio-Tagging-2019</a></p>\n\n<h3>Solution</h3>\n\n<p>single model CV:  0.89763\nensemble CV: 0.9108</p>\n\nfeature engineering\n\n<ul>\n<li>log mel  (441,64)  (time,mels)</li>\n<li>global feature (128,12) (Split the clip evenly, and create 12 features for each frame. local cv +0.005)</li>\n<li>length</li>\n</ul>\n\n<p><code>\ndef get_global_feat(x,num_steps):\n    stride = len(x)/num_steps\n    ts = []\n    for s in range(num_steps):\n        i = s * stride\n        wl = max(0,int(i - stride/2))\n        wr = int(i + 1.5*stride)\n        local_x = x[wl:wr]\n        percent_feat = np.percentile(local_x, [0, 1, 25, 30, 50, 60, 75, 99, 100]).tolist()\n        range_feat = local_x.max()-local_x.min()\n        ts.append([np.mean(local_x),np.std(local_x),range_feat]+percent_feat)\n    ts = np.array(ts)\n    assert ts.shape == (128,12),(len(x),ts.shape)\n    return ts\n</code></p>\n\nprepocess\n\n<ul>\n<li>audio clips are first trimmed of leading and trailing silence</li>\n<li>random select a 5s clip from audio clip</li>\n</ul>\n\nmodel\n\n<p>For details, please refer to code/models.py\n* Melspectrogram Layer(code from kapre,We use it to search the hyperparameter of log mel end2end)\n* Our main model is a 9-layer CNN. \nIn this competition, we consider that the two axes of the log mel feature have different physical meanings, \nso the max pooling and average pooling in the model are replaced \nby one axis using max pooling and the other axis using average pooling.\n(Our local cv gain a lot from it, but the exact number is forgotten).\n* global pooling: pixelshuffle + max pooling in time axes + ave pooling in mel axes.\n* se block（several of our models use se block)\n* highway + 1*1 conv（several of our models use se block)\n* label smoothing</p>\n\n<p>```</p>\n\n<h1>log mel layer</h1>\n\n<p>x_mel = Melspectrogram(n_dft=1024, n_hop=cfg.stride, input_shape=(1, K.int_shape(x_in)[1]),\n                           # n_hop -&gt; stride   n_dft kernel_size\n                           padding='same', sr=44100, n_mels=64,\n                           power_melgram=2, return_decibel_melgram=True,\n                           trainable_fb=False, trainable_kernel=False,\n                           image_data_format='channels_last', trainable=False)(x)\n<code>\n</code></p>\n\n<h1>pooling mode</h1>\n\n<p>x = AveragePooling2D(pool_size=(pool_size1,1), padding='same', strides=(stride,1))(x)\nx = MaxPool2D(pool_size=(1,pool_size2), padding='same', strides=(1,stride))(x)\n<code>\n</code></p>\n\n<h1>model head</h1>\n\n<p>def pixelShuffle(x):\n    _,h,w,c = K.int_shape(x)\n    bs = K.shape(x)[0]\n    assert w%2==0\n    x = K.reshape(x,(bs,h,w//2,c*2))</p>\n\n<pre><code># assert h % 2 == 0\n# x = K.permute_dimensions(x,(0,2,1,3))\n# x = K.reshape(x,(bs,w//2,h//2,c*4))\n# x = K.permute_dimensions(x,(0,2,1,3))\nreturn x\n</code></pre>\n\n<p>x = Lambda(pixelShuffle)(x)\nx = Lambda(lambda x: K.max(x, axis=1))(x)\nx = Lambda(lambda x: K.mean(x, axis=1))(x)\n```</p>\n\ndata augmentation\n\n<ul>\n<li>mixup (local cv +0.002, lb +0.008)</li>\n<li>random select 5s clip + random padding</li>\n<li>3TTA</li>\n</ul>\n\npretrain\n\n<ul>\n<li>train a model only on train_noisy as pretrained model</li>\n</ul>\n\nensemble\n\n<p>For details, please refer to code/ensemble.py\n* We use nn for stacking, \nwhich uses localconnect1D to learn the ensemble weights of each class, \nthen use fully connect to learn about label correlation, \nusing some initialization and weight constraint tricks.\n```\ndef stacker(cfg,n):\n    def kinit(shape, name=None):\n        value = np.zeros(shape)\n        value[:, -1] = 1\n        return K.variable(value, name=name)</p>\n\n<pre><code>x_in = Input((80,n))\nx = x_in\n# x = Lambda(lambda x: 1.5*x)(x)\nx = LocallyConnected1D(1,1,kernel_initializer=kinit,kernel_constraint=normNorm(1),use_bias=False)(x)\nx = Flatten()(x)\nx = Dense(80, use_bias=False, kernel_initializer=Identity(1))(x)\nx = Lambda(lambda x: (x - 1.6))(x)\nx = Activation('tanh')(x)\nx = Lambda(lambda x:(x+1)*0.5)(x)\n\nmodel = Model(inputs=x_in, outputs=x)\nmodel.compile(\n    loss='binary_crossentropy',\n    optimizer=Nadam(lr=cfg.lr),\n)\nreturn model\n</code></pre>\n\n<p>```</p>",
  "messages": [
    {
      "id": "564113",
      "postDate": "06/29/2019 02:00:47",
      "content": "<p>Thanks for the competition and congratulations of all participants! Here is a summary of our solution.\ncode: <a href=\"https://github.com/qrfaction/1st-Freesound-Audio-Tagging-2019\">https://github.com/qrfaction/1st-Freesound-Audio-Tagging-2019</a></p>\n\n<h3>Solution</h3>\n\n<p>single model CV:  0.89763\nensemble CV: 0.9108</p>\n\nfeature engineering\n\n<ul>\n<li>log mel  (441,64)  (time,mels)</li>\n<li>global feature (128,12) (Split the clip evenly, and create 12 features for each frame. local cv +0.005)</li>\n<li>length</li>\n</ul>\n\n<p><code>\ndef get_global_feat(x,num_steps):\n    stride = len(x)/num_steps\n    ts = []\n    for s in range(num_steps):\n        i = s * stride\n        wl = max(0,int(i - stride/2))\n        wr = int(i + 1.5*stride)\n        local_x = x[wl:wr]\n        percent_feat = np.percentile(local_x, [0, 1, 25, 30, 50, 60, 75, 99, 100]).tolist()\n        range_feat = local_x.max()-local_x.min()\n        ts.append([np.mean(local_x),np.std(local_x),range_feat]+percent_feat)\n    ts = np.array(ts)\n    assert ts.shape == (128,12),(len(x),ts.shape)\n    return ts\n</code></p>\n\nprepocess\n\n<ul>\n<li>audio clips are first trimmed of leading and trailing silence</li>\n<li>random select a 5s clip from audio clip</li>\n</ul>\n\nmodel\n\n<p>For details, please refer to code/models.py\n* Melspectrogram Layer(code from kapre,We use it to search the hyperparameter of log mel end2end)\n* Our main model is a 9-layer CNN. \nIn this competition, we consider that the two axes of the log mel feature have different physical meanings, \nso the max pooling and average pooling in the model are replaced \nby one axis using max pooling and the other axis using average pooling.\n(Our local cv gain a lot from it, but the exact number is forgotten).\n* global pooling: pixelshuffle + max pooling in time axes + ave pooling in mel axes.\n* se block（several of our models use se block)\n* highway + 1*1 conv（several of our models use se block)\n* label smoothing</p>\n\n<p>```</p>\n\n<h1>log mel layer</h1>\n\n<p>x_mel = Melspectrogram(n_dft=1024, n_hop=cfg.stride, input_shape=(1, K.int_shape(x_in)[1]),\n                           # n_hop -&gt; stride   n_dft kernel_size\n                           padding='same', sr=44100, n_mels=64,\n                           power_melgram=2, return_decibel_melgram=True,\n                           trainable_fb=False, trainable_kernel=False,\n                           image_data_format='channels_last', trainable=False)(x)\n<code>\n</code></p>\n\n<h1>pooling mode</h1>\n\n<p>x = AveragePooling2D(pool_size=(pool_size1,1), padding='same', strides=(stride,1))(x)\nx = MaxPool2D(pool_size=(1,pool_size2), padding='same', strides=(1,stride))(x)\n<code>\n</code></p>\n\n<h1>model head</h1>\n\n<p>def pixelShuffle(x):\n    _,h,w,c = K.int_shape(x)\n    bs = K.shape(x)[0]\n    assert w%2==0\n    x = K.reshape(x,(bs,h,w//2,c*2))</p>\n\n<pre><code># assert h % 2 == 0\n# x = K.permute_dimensions(x,(0,2,1,3))\n# x = K.reshape(x,(bs,w//2,h//2,c*4))\n# x = K.permute_dimensions(x,(0,2,1,3))\nreturn x\n</code></pre>\n\n<p>x = Lambda(pixelShuffle)(x)\nx = Lambda(lambda x: K.max(x, axis=1))(x)\nx = Lambda(lambda x: K.mean(x, axis=1))(x)\n```</p>\n\ndata augmentation\n\n<ul>\n<li>mixup (local cv +0.002, lb +0.008)</li>\n<li>random select 5s clip + random padding</li>\n<li>3TTA</li>\n</ul>\n\npretrain\n\n<ul>\n<li>train a model only on train_noisy as pretrained model</li>\n</ul>\n\nensemble\n\n<p>For details, please refer to code/ensemble.py\n* We use nn for stacking, \nwhich uses localconnect1D to learn the ensemble weights of each class, \nthen use fully connect to learn about label correlation, \nusing some initialization and weight constraint tricks.\n```\ndef stacker(cfg,n):\n    def kinit(shape, name=None):\n        value = np.zeros(shape)\n        value[:, -1] = 1\n        return K.variable(value, name=name)</p>\n\n<pre><code>x_in = Input((80,n))\nx = x_in\n# x = Lambda(lambda x: 1.5*x)(x)\nx = LocallyConnected1D(1,1,kernel_initializer=kinit,kernel_constraint=normNorm(1),use_bias=False)(x)\nx = Flatten()(x)\nx = Dense(80, use_bias=False, kernel_initializer=Identity(1))(x)\nx = Lambda(lambda x: (x - 1.6))(x)\nx = Activation('tanh')(x)\nx = Lambda(lambda x:(x+1)*0.5)(x)\n\nmodel = Model(inputs=x_in, outputs=x)\nmodel.compile(\n    loss='binary_crossentropy',\n    optimizer=Nadam(lr=cfg.lr),\n)\nreturn model\n</code></pre>\n\n<p>```</p>",
      "rawMarkdown": "Thanks for the competition and congratulations of all participants! Here is a summary of our solution.\ncode: https://github.com/qrfaction/1st-Freesound-Audio-Tagging-2019\n\n### Solution\nsingle model CV:  0.89763\nensemble CV: 0.9108\n\n#### feature engineering\n* log mel  (441,64)  (time,mels)\n* global feature (128,12) (Split the clip evenly, and create 12 features for each frame. local cv +0.005)\n* length\n\n```\ndef get_global_feat(x,num_steps):\n    stride = len(x)/num_steps\n    ts = []\n    for s in range(num_steps):\n        i = s * stride\n        wl = max(0,int(i - stride/2))\n        wr = int(i + 1.5*stride)\n        local_x = x[wl:wr]\n        percent_feat = np.percentile(local_x, [0, 1, 25, 30, 50, 60, 75, 99, 100]).tolist()\n        range_feat = local_x.max()-local_x.min()\n        ts.append([np.mean(local_x),np.std(local_x),range_feat]+percent_feat)\n    ts = np.array(ts)\n    assert ts.shape == (128,12),(len(x),ts.shape)\n    return ts\n```\n\n\n#### prepocess\n* audio clips are first trimmed of leading and trailing silence\n* random select a 5s clip from audio clip\n\n#### model\nFor details, please refer to code/models.py\n* Melspectrogram Layer(code from kapre,We use it to search the hyperparameter of log mel end2end)\n* Our main model is a 9-layer CNN. \nIn this competition, we consider that the two axes of the log mel feature have different physical meanings, \nso the max pooling and average pooling in the model are replaced \nby one axis using max pooling and the other axis using average pooling.\n(Our local cv gain a lot from it, but the exact number is forgotten).\n* global pooling: pixelshuffle + max pooling in time axes + ave pooling in mel axes.\n* se block（several of our models use se block)\n* highway + 1*1 conv（several of our models use se block)\n* label smoothing\n\n```\n# log mel layer\nx_mel = Melspectrogram(n_dft=1024, n_hop=cfg.stride, input_shape=(1, K.int_shape(x_in)[1]),\n                           # n_hop -&gt; stride   n_dft kernel_size\n                           padding='same', sr=44100, n_mels=64,\n                           power_melgram=2, return_decibel_melgram=True,\n                           trainable_fb=False, trainable_kernel=False,\n                           image_data_format='channels_last', trainable=False)(x)\n```\n```\n# pooling mode\nx = AveragePooling2D(pool_size=(pool_size1,1), padding='same', strides=(stride,1))(x)\nx = MaxPool2D(pool_size=(1,pool_size2), padding='same', strides=(1,stride))(x)\n```\n```\n# model head\ndef pixelShuffle(x):\n    _,h,w,c = K.int_shape(x)\n    bs = K.shape(x)[0]\n    assert w%2==0\n    x = K.reshape(x,(bs,h,w//2,c*2))\n\n    # assert h % 2 == 0\n    # x = K.permute_dimensions(x,(0,2,1,3))\n    # x = K.reshape(x,(bs,w//2,h//2,c*4))\n    # x = K.permute_dimensions(x,(0,2,1,3))\n    return x\nx = Lambda(pixelShuffle)(x)\nx = Lambda(lambda x: K.max(x, axis=1))(x)\nx = Lambda(lambda x: K.mean(x, axis=1))(x)\n```\n\n#### data augmentation\n* mixup (local cv +0.002, lb +0.008)\n* random select 5s clip + random padding\n* 3TTA\n\n#### pretrain\n* train a model only on train_noisy as pretrained model\n\n#### ensemble\nFor details, please refer to code/ensemble.py\n* We use nn for stacking, \nwhich uses localconnect1D to learn the ensemble weights of each class, \nthen use fully connect to learn about label correlation, \nusing some initialization and weight constraint tricks.\n```\ndef stacker(cfg,n):\n    def kinit(shape, name=None):\n        value = np.zeros(shape)\n        value[:, -1] = 1\n        return K.variable(value, name=name)\n\n\n    x_in = Input((80,n))\n    x = x_in\n    # x = Lambda(lambda x: 1.5*x)(x)\n    x = LocallyConnected1D(1,1,kernel_initializer=kinit,kernel_constraint=normNorm(1),use_bias=False)(x)\n    x = Flatten()(x)\n    x = Dense(80, use_bias=False, kernel_initializer=Identity(1))(x)\n    x = Lambda(lambda x: (x - 1.6))(x)\n    x = Activation('tanh')(x)\n    x = Lambda(lambda x:(x+1)*0.5)(x)\n\n    model = Model(inputs=x_in, outputs=x)\n    model.compile(\n        loss='binary_crossentropy',\n        optimizer=Nadam(lr=cfg.lr),\n    )\n    return model\n\n```",
      "votes": null
    },
    {
      "id": "564141",
      "postDate": "06/29/2019 03:06:18",
      "content": "<p>Congrats becoming a GGGGGGGGGGGGM!!!!</p>",
      "rawMarkdown": "Congrats becoming a GGGGGGGGGGGGM!!!!",
      "votes": null
    },
    {
      "id": "564173",
      "postDate": "06/29/2019 04:44:44",
      "content": "<p>Congrats on GM </p>",
      "rawMarkdown": "Congrats on GM",
      "votes": null
    },
    {
      "id": "564188",
      "postDate": "06/29/2019 05:21:38",
      "content": "<p>tks!</p>",
      "rawMarkdown": "tks!",
      "votes": null
    },
    {
      "id": "564189",
      "postDate": "06/29/2019 05:22:06",
      "content": "<p>notks</p>",
      "rawMarkdown": "notks",
      "votes": null
    },
    {
      "id": "564227",
      "postDate": "06/29/2019 06:49:47",
      "content": "<p>Congratulations! I'm sure its one of the best feelings to become a GM with a #1 winning gold :-)\nAnd thanks a lot for sharing your solution.</p>",
      "rawMarkdown": "Congratulations! I'm sure its one of the best feelings to become a GM with a #1 winning gold :-)\nAnd thanks a lot for sharing your solution.",
      "votes": null
    },
    {
      "id": "564230",
      "postDate": "06/29/2019 06:51:49",
      "content": "<p>Yes, very happy. tks</p>",
      "rawMarkdown": "Yes, very happy. tks",
      "votes": null
    },
    {
      "id": "564539",
      "postDate": "06/29/2019 15:15:38",
      "content": "<p>Great work!!!\nCongrats becoming  GM!!!!  (为钱老师打call)</p>",
      "rawMarkdown": "Great work!!!\nCongrats becoming  GM!!!!  (为钱老师打call)",
      "votes": null
    },
    {
      "id": "564569",
      "postDate": "06/29/2019 16:15:45",
      "content": "<p>hah</p>",
      "rawMarkdown": "hah",
      "votes": null
    },
    {
      "id": "564669",
      "postDate": "06/29/2019 19:25:43",
      "content": "<p>Heartily congrates and thanks a lot for sharing the solution </p>",
      "rawMarkdown": "Heartily congrates and thanks a lot for sharing the solution",
      "votes": null
    },
    {
      "id": "565680",
      "postDate": "07/01/2019 09:17:11",
      "content": "<p>congratulations</p>",
      "rawMarkdown": "congratulations",
      "votes": null
    },
    {
      "id": "566368",
      "postDate": "07/02/2019 05:20:43",
      "content": "<p>Congrats Dude</p>",
      "rawMarkdown": "Congrats Dude",
      "votes": null
    },
    {
      "id": "566655",
      "postDate": "07/02/2019 12:27:07",
      "content": "<p>Coooooool solution.</p>",
      "rawMarkdown": "Coooooool solution.",
      "votes": null
    },
    {
      "id": "566673",
      "postDate": "07/02/2019 12:44:57",
      "content": "<p>😀 </p>",
      "rawMarkdown": "😀",
      "votes": null
    },
    {
      "id": "568391",
      "postDate": "07/04/2019 20:44:29",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!",
      "votes": null
    },
    {
      "id": "569705",
      "postDate": "07/07/2019 07:28:45",
      "content": "<p>Why MaxPool time and AvgPool freq specifically? Is there an intuitive reason for doing this? Congratulations!</p>",
      "rawMarkdown": "Why MaxPool time and AvgPool freq specifically? Is there an intuitive reason for doing this? Congratulations!",
      "votes": null
    },
    {
      "id": "570199",
      "postDate": "07/08/2019 01:46:11",
      "content": "<p>I only considered the physical meaning, ave and max pooling are the results of my experiment.</p>",
      "rawMarkdown": "I only considered the physical meaning, ave and max pooling are the results of my experiment.",
      "votes": null
    },
    {
      "id": "582300",
      "postDate": "07/23/2019 03:24:06",
      "content": "<p>Big congrats!!! Thanks for shareing!! By the way, I am trying to reproduce your result, but when I encounter an error in  \"x_mel = Melspectrogram(n_dft=1024, n_hop=cfg.stride, input_shape=(1, K.int_shape(x_in)[1]),\" there is no stride in config.py, cfg.stride does not exists, is this a typo?</p>",
      "rawMarkdown": "Big congrats!!! Thanks for shareing!! By the way, I am trying to reproduce your result, but when I encounter an error in  \"x_mel = Melspectrogram(n_dft=1024, n_hop=cfg.stride, input_shape=(1, K.int_shape(x_in)[1]),\" there is no stride in config.py, cfg.stride does not exists, is this a typo?",
      "votes": null
    },
    {
      "id": "756802",
      "postDate": "02/26/2020 05:03:40",
      "content": "<p>hi, my name is Wubin Bai. and I have been doing your 2nd place solution to the 2019 fsd problem for a while.  I accidently noticed that you are also from China. May we chat more on this? Do you have a wechat account so maybe we could talk more about this competition, or maybeI could learn something from you. Thanks!</p>",
      "rawMarkdown": "hi, my name is Wubin Bai. and I have been doing your 2nd place solution to the 2019 fsd problem for a while.  I accidently noticed that you are also from China. May we chat more on this? Do you have a wechat account so maybe we could talk more about this competition, or maybeI could learn something from you. Thanks!",
      "votes": null
    },
    {
      "id": "761407",
      "postDate": "03/02/2020 13:34:03",
      "content": "<p>I have sent you a WeChat account via email</p>",
      "rawMarkdown": "I have sent you a WeChat account via email",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 564141,
      "author_name": "hadxu123",
      "author_url": "",
      "post_date": "06/29/2019 03:06:18",
      "content": "<p>Congrats becoming a GGGGGGGGGGGGM!!!!</p>",
      "votes": null,
      "replies": [
        {
          "id": 564189,
          "author_name": "action",
          "author_url": "",
          "post_date": "06/29/2019 05:22:06",
          "content": "<p>notks</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 564173,
      "author_name": "zjucor",
      "author_url": "",
      "post_date": "06/29/2019 04:44:44",
      "content": "<p>Congrats on GM </p>",
      "votes": null,
      "replies": [
        {
          "id": 564188,
          "author_name": "action",
          "author_url": "",
          "post_date": "06/29/2019 05:21:38",
          "content": "<p>tks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 564227,
      "author_name": "rohanrao",
      "author_url": "",
      "post_date": "06/29/2019 06:49:47",
      "content": "<p>Congratulations! I'm sure its one of the best feelings to become a GM with a #1 winning gold :-)\nAnd thanks a lot for sharing your solution.</p>",
      "votes": null,
      "replies": [
        {
          "id": 564230,
          "author_name": "action",
          "author_url": "",
          "post_date": "06/29/2019 06:51:49",
          "content": "<p>Yes, very happy. tks</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 564539,
      "author_name": "chizhu2018",
      "author_url": "",
      "post_date": "06/29/2019 15:15:38",
      "content": "<p>Great work!!!\nCongrats becoming  GM!!!!  (为钱老师打call)</p>",
      "votes": null,
      "replies": [
        {
          "id": 564569,
          "author_name": "action",
          "author_url": "",
          "post_date": "06/29/2019 16:15:45",
          "content": "<p>hah</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 756802,
          "author_name": "",
          "author_url": "",
          "post_date": "02/26/2020 05:03:40",
          "content": "<p>hi, my name is Wubin Bai. and I have been doing your 2nd place solution to the 2019 fsd problem for a while.  I accidently noticed that you are also from China. May we chat more on this? Do you have a wechat account so maybe we could talk more about this competition, or maybeI could learn something from you. Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 761407,
          "author_name": "action",
          "author_url": "",
          "post_date": "03/02/2020 13:34:03",
          "content": "<p>I have sent you a WeChat account via email</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 564669,
      "author_name": "dataraj",
      "author_url": "",
      "post_date": "06/29/2019 19:25:43",
      "content": "<p>Heartily congrates and thanks a lot for sharing the solution </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 565680,
      "author_name": "krishnakatyal",
      "author_url": "",
      "post_date": "07/01/2019 09:17:11",
      "content": "<p>congratulations</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 566368,
      "author_name": "maliabbas",
      "author_url": "",
      "post_date": "07/02/2019 05:20:43",
      "content": "<p>Congrats Dude</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 566655,
      "author_name": "baomengjiao",
      "author_url": "",
      "post_date": "07/02/2019 12:27:07",
      "content": "<p>Coooooool solution.</p>",
      "votes": null,
      "replies": [
        {
          "id": 566673,
          "author_name": "action",
          "author_url": "",
          "post_date": "07/02/2019 12:44:57",
          "content": "<p>😀 </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 568391,
      "author_name": "v100p100",
      "author_url": "",
      "post_date": "07/04/2019 20:44:29",
      "content": "<p>Congratulations!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 569705,
      "author_name": "erosis",
      "author_url": "",
      "post_date": "07/07/2019 07:28:45",
      "content": "<p>Why MaxPool time and AvgPool freq specifically? Is there an intuitive reason for doing this? Congratulations!</p>",
      "votes": null,
      "replies": [
        {
          "id": 570199,
          "author_name": "action",
          "author_url": "",
          "post_date": "07/08/2019 01:46:11",
          "content": "<p>I only considered the physical meaning, ave and max pooling are the results of my experiment.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 582300,
      "author_name": "riemann",
      "author_url": "",
      "post_date": "07/23/2019 03:24:06",
      "content": "<p>Big congrats!!! Thanks for shareing!! By the way, I am trying to reproduce your result, but when I encounter an error in  \"x_mel = Melspectrogram(n_dft=1024, n_hop=cfg.stride, input_shape=(1, K.int_shape(x_in)[1]),\" there is no stride in config.py, cfg.stride does not exists, is this a typo?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "564113": "Thanks for the competition and congratulations of all participants! Here is a summary of our solution.\ncode: https://github.com/qrfaction/1st-Freesound-Audio-Tagging-2019\n\n### Solution\nsingle model CV:  0.89763\nensemble CV: 0.9108\n\n#### feature engineering\n* log mel  (441,64)  (time,mels)\n* global feature (128,12) (Split the clip evenly, and create 12 features for each frame. local cv +0.005)\n* length\n\n```\ndef get_global_feat(x,num_steps):\n    stride = len(x)/num_steps\n    ts = []\n    for s in range(num_steps):\n        i = s * stride\n        wl = max(0,int(i - stride/2))\n        wr = int(i + 1.5*stride)\n        local_x = x[wl:wr]\n        percent_feat = np.percentile(local_x, [0, 1, 25, 30, 50, 60, 75, 99, 100]).tolist()\n        range_feat = local_x.max()-local_x.min()\n        ts.append([np.mean(local_x),np.std(local_x),range_feat]+percent_feat)\n    ts = np.array(ts)\n    assert ts.shape == (128,12),(len(x),ts.shape)\n    return ts\n```\n\n\n#### prepocess\n* audio clips are first trimmed of leading and trailing silence\n* random select a 5s clip from audio clip\n\n#### model\nFor details, please refer to code/models.py\n* Melspectrogram Layer(code from kapre,We use it to search the hyperparameter of log mel end2end)\n* Our main model is a 9-layer CNN. \nIn this competition, we consider that the two axes of the log mel feature have different physical meanings, \nso the max pooling and average pooling in the model are replaced \nby one axis using max pooling and the other axis using average pooling.\n(Our local cv gain a lot from it, but the exact number is forgotten).\n* global pooling: pixelshuffle + max pooling in time axes + ave pooling in mel axes.\n* se block（several of our models use se block)\n* highway + 1*1 conv（several of our models use se block)\n* label smoothing\n\n```\n# log mel layer\nx_mel = Melspectrogram(n_dft=1024, n_hop=cfg.stride, input_shape=(1, K.int_shape(x_in)[1]),\n                           # n_hop -&gt; stride   n_dft kernel_size\n                           padding='same', sr=44100, n_mels=64,\n                           power_melgram=2, return_decibel_melgram=True,\n                           trainable_fb=False, trainable_kernel=False,\n                           image_data_format='channels_last', trainable=False)(x)\n```\n```\n# pooling mode\nx = AveragePooling2D(pool_size=(pool_size1,1), padding='same', strides=(stride,1))(x)\nx = MaxPool2D(pool_size=(1,pool_size2), padding='same', strides=(1,stride))(x)\n```\n```\n# model head\ndef pixelShuffle(x):\n    _,h,w,c = K.int_shape(x)\n    bs = K.shape(x)[0]\n    assert w%2==0\n    x = K.reshape(x,(bs,h,w//2,c*2))\n\n    # assert h % 2 == 0\n    # x = K.permute_dimensions(x,(0,2,1,3))\n    # x = K.reshape(x,(bs,w//2,h//2,c*4))\n    # x = K.permute_dimensions(x,(0,2,1,3))\n    return x\nx = Lambda(pixelShuffle)(x)\nx = Lambda(lambda x: K.max(x, axis=1))(x)\nx = Lambda(lambda x: K.mean(x, axis=1))(x)\n```\n\n#### data augmentation\n* mixup (local cv +0.002, lb +0.008)\n* random select 5s clip + random padding\n* 3TTA\n\n#### pretrain\n* train a model only on train_noisy as pretrained model\n\n#### ensemble\nFor details, please refer to code/ensemble.py\n* We use nn for stacking, \nwhich uses localconnect1D to learn the ensemble weights of each class, \nthen use fully connect to learn about label correlation, \nusing some initialization and weight constraint tricks.\n```\ndef stacker(cfg,n):\n    def kinit(shape, name=None):\n        value = np.zeros(shape)\n        value[:, -1] = 1\n        return K.variable(value, name=name)\n\n\n    x_in = Input((80,n))\n    x = x_in\n    # x = Lambda(lambda x: 1.5*x)(x)\n    x = LocallyConnected1D(1,1,kernel_initializer=kinit,kernel_constraint=normNorm(1),use_bias=False)(x)\n    x = Flatten()(x)\n    x = Dense(80, use_bias=False, kernel_initializer=Identity(1))(x)\n    x = Lambda(lambda x: (x - 1.6))(x)\n    x = Activation('tanh')(x)\n    x = Lambda(lambda x:(x+1)*0.5)(x)\n\n    model = Model(inputs=x_in, outputs=x)\n    model.compile(\n        loss='binary_crossentropy',\n        optimizer=Nadam(lr=cfg.lr),\n    )\n    return model\n\n```",
    "564141": "Congrats becoming a GGGGGGGGGGGGM!!!!",
    "564173": "Congrats on GM",
    "564188": "tks!",
    "564189": "notks",
    "564227": "Congratulations! I'm sure its one of the best feelings to become a GM with a #1 winning gold :-)\nAnd thanks a lot for sharing your solution.",
    "564230": "Yes, very happy. tks",
    "564539": "Great work!!!\nCongrats becoming  GM!!!!  (为钱老师打call)",
    "564569": "hah",
    "564669": "Heartily congrates and thanks a lot for sharing the solution",
    "565680": "congratulations",
    "566368": "Congrats Dude",
    "566655": "Coooooool solution.",
    "566673": "😀",
    "568391": "Congratulations!",
    "569705": "Why MaxPool time and AvgPool freq specifically? Is there an intuitive reason for doing this? Congratulations!",
    "570199": "I only considered the physical meaning, ave and max pooling are the results of my experiment.",
    "582300": "Big congrats!!! Thanks for shareing!! By the way, I am trying to reproduce your result, but when I encounter an error in  \"x_mel = Melspectrogram(n_dft=1024, n_hop=cfg.stride, input_shape=(1, K.int_shape(x_in)[1]),\" there is no stride in config.py, cfg.stride does not exists, is this a typo?",
    "756802": "hi, my name is Wubin Bai. and I have been doing your 2nd place solution to the 2019 fsd problem for a while.  I accidently noticed that you are also from China. May we chat more on this? Do you have a wechat account so maybe we could talk more about this competition, or maybeI could learn something from you. Thanks!",
    "761407": "I have sent you a WeChat account via email"
  },
  "source": "meta"
}