{
  "id": 97804,
  "title": "A Naive Solution (28th in Private LB)",
  "url": "/competitions/freesound-audio-tagging-2019/writeups/maxwell-is-observing-a-naive-solution-28th-in-priv",
  "author_name": "",
  "post_date": "2019-06-29T00:01:50.615200400Z",
  "votes": 10,
  "comment_count": 6,
  "views": 0,
  "content": "<p>First of all, I'd like to thank organizers and all participants, and congratulate winners on excellent results !\nFor me, this is the first image(sound) competition in Kaggle and I have learned a lot.</p>\n\n<p>Many top teams unveiled their nice solutions(also scores). So I'm ashamed to publish my solution here, but I will do for a memento.  </p>\n\n<hr>\n\n<h1>Local Validation</h1>\n\n<p>I created 80class-balanced 5 folds validation and have evaluated the model performance with BCE. At the early stage in this competition, I have checked the correlation of BCE and LWLRAP. Below figure shows an almost good correlation. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F01d7ab43832e5eb5eec8a139fdb69c19%2Ffig1.png?generation=1561761823509063&amp;alt=media\" alt=\"\"></p>\n\n<h1>Features</h1>\n\n<ol>\n<li>Remove silent audios</li>\n<li>Trim silent parts</li>\n<li>Log Mel Spectrogram ( SR 44.1kHz, FFT window size 80ms, Hop 10ms, Mel Bands 64 )</li>\n<li>Clustered frequency-wise statistical features</li>\n</ol>\n\n<p>I did not do nothing special. But I will explain about <strong>4</strong>. <br>\nI thought CNN can catch an audio property in frequency space like Fourier Analysis, but less in  statistics values ( max, min, ... ). And labels of train-noisy are unreliable, so I created frequency-wise statistical features from spectrograms and compute cluster distances without using label information(Unsupervised).  The procedure is following,  </p>\n\n<p>a. Compute 25 statistical values (Max, Min, Primary difference, Secondary difference, etc... ) and flatten 64 x 25 (=1600) features \nb. Compute 200 cluster distances with MiniBatchKMeans ( dimensional reduction from 1600 to 200 )</p>\n\n<p>This feature pushed my score about 0.5 ~ 1% in each model.  </p>\n\n<h1>Models, Train and Prediction</h1>\n\n<p>In this competition, we do not have a lot of time to make inferences ( less than 1 GPU hour ). So I selected 3 relatively light-weight models.  </p>\n\n<ul>\n<li>Mobile Net V2   with/without clustered features</li>\n<li>ResNet50   with/without clustered features</li>\n<li>DenseNet121   with/without clustered features</li>\n</ul>\n\n<p>My final submission is the ensemble of above 6 models ( 3 models x 2 type of features ). Here is a model pipeline. The setting of train and prediction with TTA is written in this figure.  </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2Fd21b6e837c1be18cc4a66bbbc82c3174%2Ffs2019_model-pipeline_01.png?generation=1561760818137707&amp;alt=media\" alt=\"\"></p>\n\n<h1>Performance</h1>\n\n<p>I used the weighted geometric averaging to blend 6 model predictions. Below table shows each performance and the blending coefficients. Those coefficients are computed with optimization based on Out Of Fold predictions.  LWLRAP values are calculated with 5 fold OOF predictions on train-curated data.</p>\n\n<p>| Model | LWLRAP on train-curated | Blending Coefficient |\n| --- | --- | --- |\n| MobileNetV2 with clustered features( cf ) | 0.84519 | 0.246 |\n| MobileNetV2 without cf | 0.82940 | 0.217 |\n| ResNet50 with cf | 0.84490 | 0.149 |\n| ResNet50 without cf | 0.83006 | 0.072 |\n| DenseNet121 with cf | 0.84353 | 0.115 |\n| DenseNet121 without cf | 0.83501 | 0.201 |\n| <strong>Blended</strong> | 0.87611 ( Private 0.72820 ) | --- |</p>\n\n<p>Thank you very much for reading to the end. <br>\nSee you at the next competition !</p>",
  "messages": [
    {
      "id": "564067",
      "postDate": "06/29/2019 00:01:50",
      "content": "<p>First of all, I'd like to thank organizers and all participants, and congratulate winners on excellent results !\nFor me, this is the first image(sound) competition in Kaggle and I have learned a lot.</p>\n\n<p>Many top teams unveiled their nice solutions(also scores). So I'm ashamed to publish my solution here, but I will do for a memento.  </p>\n\n<hr>\n\n<h1>Local Validation</h1>\n\n<p>I created 80class-balanced 5 folds validation and have evaluated the model performance with BCE. At the early stage in this competition, I have checked the correlation of BCE and LWLRAP. Below figure shows an almost good correlation. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F01d7ab43832e5eb5eec8a139fdb69c19%2Ffig1.png?generation=1561761823509063&amp;alt=media\" alt=\"\"></p>\n\n<h1>Features</h1>\n\n<ol>\n<li>Remove silent audios</li>\n<li>Trim silent parts</li>\n<li>Log Mel Spectrogram ( SR 44.1kHz, FFT window size 80ms, Hop 10ms, Mel Bands 64 )</li>\n<li>Clustered frequency-wise statistical features</li>\n</ol>\n\n<p>I did not do nothing special. But I will explain about <strong>4</strong>. <br>\nI thought CNN can catch an audio property in frequency space like Fourier Analysis, but less in  statistics values ( max, min, ... ). And labels of train-noisy are unreliable, so I created frequency-wise statistical features from spectrograms and compute cluster distances without using label information(Unsupervised).  The procedure is following,  </p>\n\n<p>a. Compute 25 statistical values (Max, Min, Primary difference, Secondary difference, etc... ) and flatten 64 x 25 (=1600) features \nb. Compute 200 cluster distances with MiniBatchKMeans ( dimensional reduction from 1600 to 200 )</p>\n\n<p>This feature pushed my score about 0.5 ~ 1% in each model.  </p>\n\n<h1>Models, Train and Prediction</h1>\n\n<p>In this competition, we do not have a lot of time to make inferences ( less than 1 GPU hour ). So I selected 3 relatively light-weight models.  </p>\n\n<ul>\n<li>Mobile Net V2   with/without clustered features</li>\n<li>ResNet50   with/without clustered features</li>\n<li>DenseNet121   with/without clustered features</li>\n</ul>\n\n<p>My final submission is the ensemble of above 6 models ( 3 models x 2 type of features ). Here is a model pipeline. The setting of train and prediction with TTA is written in this figure.  </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2Fd21b6e837c1be18cc4a66bbbc82c3174%2Ffs2019_model-pipeline_01.png?generation=1561760818137707&amp;alt=media\" alt=\"\"></p>\n\n<h1>Performance</h1>\n\n<p>I used the weighted geometric averaging to blend 6 model predictions. Below table shows each performance and the blending coefficients. Those coefficients are computed with optimization based on Out Of Fold predictions.  LWLRAP values are calculated with 5 fold OOF predictions on train-curated data.</p>\n\n<p>| Model | LWLRAP on train-curated | Blending Coefficient |\n| --- | --- | --- |\n| MobileNetV2 with clustered features( cf ) | 0.84519 | 0.246 |\n| MobileNetV2 without cf | 0.82940 | 0.217 |\n| ResNet50 with cf | 0.84490 | 0.149 |\n| ResNet50 without cf | 0.83006 | 0.072 |\n| DenseNet121 with cf | 0.84353 | 0.115 |\n| DenseNet121 without cf | 0.83501 | 0.201 |\n| <strong>Blended</strong> | 0.87611 ( Private 0.72820 ) | --- |</p>\n\n<p>Thank you very much for reading to the end. <br>\nSee you at the next competition !</p>",
      "rawMarkdown": "First of all, I'd like to thank organizers and all participants, and congratulate winners on excellent results !\nFor me, this is the first image(sound) competition in Kaggle and I have learned a lot.\n\nMany top teams unveiled their nice solutions(also scores). So I'm ashamed to publish my solution here, but I will do for a memento.  \n\n---\n\n# Local Validation\n\nI created 80class-balanced 5 folds validation and have evaluated the model performance with BCE. At the early stage in this competition, I have checked the correlation of BCE and LWLRAP. Below figure shows an almost good correlation.  \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F01d7ab43832e5eb5eec8a139fdb69c19%2Ffig1.png?generation=1561761823509063&amp;alt=media)\n\n\n# Features\n\n1. Remove silent audios\n2. Trim silent parts\n3. Log Mel Spectrogram ( SR 44.1kHz, FFT window size 80ms, Hop 10ms, Mel Bands 64 )\n4. Clustered frequency-wise statistical features\n\nI did not do nothing special. But I will explain about **4**.  \nI thought CNN can catch an audio property in frequency space like Fourier Analysis, but less in  statistics values ( max, min, ... ). And labels of train-noisy are unreliable, so I created frequency-wise statistical features from spectrograms and compute cluster distances without using label information(Unsupervised).  The procedure is following,  \n\na. Compute 25 statistical values (Max, Min, Primary difference, Secondary difference, etc... ) and flatten 64 x 25 (=1600) features \nb. Compute 200 cluster distances with MiniBatchKMeans ( dimensional reduction from 1600 to 200 )\n\nThis feature pushed my score about 0.5 ~ 1% in each model.  \n\n\n# Models, Train and Prediction\n\nIn this competition, we do not have a lot of time to make inferences ( less than 1 GPU hour ). So I selected 3 relatively light-weight models.  \n\n-  Mobile Net V2   with/without clustered features\n- ResNet50   with/without clustered features\n- DenseNet121   with/without clustered features\n\nMy final submission is the ensemble of above 6 models ( 3 models x 2 type of features ). Here is a model pipeline. The setting of train and prediction with TTA is written in this figure.  \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2Fd21b6e837c1be18cc4a66bbbc82c3174%2Ffs2019_model-pipeline_01.png?generation=1561760818137707&amp;alt=media)\n\n\n# Performance  \n\nI used the weighted geometric averaging to blend 6 model predictions. Below table shows each performance and the blending coefficients. Those coefficients are computed with optimization based on Out Of Fold predictions.  LWLRAP values are calculated with 5 fold OOF predictions on train-curated data.\n\n| Model | LWLRAP on train-curated | Blending Coefficient |\n| --- | --- | --- |\n| MobileNetV2 with clustered features( cf ) | 0.84519 | 0.246 |\n| MobileNetV2 without cf | 0.82940 | 0.217 |\n| ResNet50 with cf | 0.84490 | 0.149 |\n| ResNet50 without cf | 0.83006 | 0.072 |\n| DenseNet121 with cf | 0.84353 | 0.115 |\n| DenseNet121 without cf | 0.83501 | 0.201 |\n| **Blended** | 0.87611 ( Private 0.72820 ) | --- |\n\n\nThank you very much for reading to the end.  \nSee you at the next competition !",
      "votes": null
    },
    {
      "id": "564076",
      "postDate": "06/29/2019 00:24:05",
      "content": "<p>Congrats. A pretty clear explanation of your approach. I have a question about that figure, what software did you use to create it.</p>",
      "rawMarkdown": "Congrats. A pretty clear explanation of your approach. I have a question about that figure, what software did you use to create it.",
      "votes": null
    },
    {
      "id": "564079",
      "postDate": "06/29/2019 00:29:24",
      "content": "<p><a href=\"/jmourad100\">@jmourad100</a> \nThanks!\nIt may be unexpected for you, I used a just power point to create this figure. <br>\n( Someone called me <code>Power Point Master</code> not kaggle master ;) )</p>",
      "rawMarkdown": "jmourad100 \nThanks!\nIt may be unexpected for you, I used a just power point to create this figure.  \n( Someone called me `Power Point Master` not kaggle master ;) )",
      "votes": null
    },
    {
      "id": "564088",
      "postDate": "06/29/2019 00:55:07",
      "content": "<p>Congrats.\nI am interested in what kind of features can be made.\nIf possible, could you tell me more about the content of \"25 statistical values\"?</p>",
      "rawMarkdown": "Congrats.\nI am interested in what kind of features can be made.\nIf possible, could you tell me more about the content of \"25 statistical values\"?",
      "votes": null
    },
    {
      "id": "564098",
      "postDate": "06/29/2019 01:24:24",
      "content": "<p><a href=\"/bossimuimu\">@bossimuimu</a> \nThanks for your question. <br>\nI made a helper function, modified version in <a href=\"https://www.kaggle.com/titericz/lightgbm-simple-solution-lb-0-203\">this kernel</a></p>\n\n<p>Hope this is what you asked.  </p>\n\n<p>```\ndef create_features(logmel_fn):\n    ## load log-mel array\n    lm_ar = mel_0_1(np.load(logmel_fn))\n    lm_df = pd.DataFrame(lm_ar)</p>\n\n<pre><code>stat_l = []\nstat_l.append(lm_df.mean(axis=1).values)\nstat_l.append(lm_df.median(axis=1).values)\nstat_l.append(lm_df.std(axis=1).values)\nstat_l.append(lm_df.max(axis=1).values)\nstat_l.append(lm_df.min(axis=1).values)\nstat_l.append(lm_df.skew(axis=1).values)\nstat_l.append(lm_df.mad(axis=1).values)\nstat_l.append(lm_df.kurtosis(axis=1).values)\n\nstat_l.append(np.abs(lm_ar).max(axis=1))\nstat_l.append(np.abs(lm_ar).min(axis=1))\nstat_l.append(np.abs(lm_ar).mean(axis=1))\nstat_l.append(np.abs(lm_ar).std(axis=1))\n\nstat_l.append(lm_ar.max(axis=1) / (np.abs(lm_ar.min(axis=1)) + 1e-32))\nstat_l.append(lm_ar.max(axis=1) - np.abs(lm_ar.min(axis=1)))\nstat_l.append(lm_ar.sum(axis=1))\n\nstat_l.append(lm_df.quantile(0.99, axis=1).values)\nstat_l.append(lm_df.quantile(0.95, axis=1).values)\nstat_l.append(lm_df.quantile(0.1, axis=1).values)\nstat_l.append(lm_df.quantile(0.05, axis=1).values)\n\nstat_l.append(np.mean(np.diff(lm_ar), axis=1))\nstat_l.append(np.max(np.diff(lm_ar), axis=1))\nstat_l.append(np.min(np.diff(lm_ar), axis=1))\nstat_l.append(np.mean(np.diff(np.diff(lm_ar)), axis=1))\nstat_l.append(np.max(np.diff(np.diff(lm_ar)), axis=1))\nstat_l.append(np.min(np.diff(np.diff(lm_ar)), axis=1))\n\nreturn np.array(stat_l).T\n</code></pre>\n\n<p>```</p>",
      "rawMarkdown": "bossimuimu \nThanks for your question.  \nI made a helper function, modified version in [this kernel](https://www.kaggle.com/titericz/lightgbm-simple-solution-lb-0-203)\n\nHope this is what you asked.  \n\n```\ndef create_features(logmel_fn):\n    ## load log-mel array\n    lm_ar = mel_0_1(np.load(logmel_fn))\n    lm_df = pd.DataFrame(lm_ar)\n\n    stat_l = []\n    stat_l.append(lm_df.mean(axis=1).values)\n    stat_l.append(lm_df.median(axis=1).values)\n    stat_l.append(lm_df.std(axis=1).values)\n    stat_l.append(lm_df.max(axis=1).values)\n    stat_l.append(lm_df.min(axis=1).values)\n    stat_l.append(lm_df.skew(axis=1).values)\n    stat_l.append(lm_df.mad(axis=1).values)\n    stat_l.append(lm_df.kurtosis(axis=1).values)\n\n    stat_l.append(np.abs(lm_ar).max(axis=1))\n    stat_l.append(np.abs(lm_ar).min(axis=1))\n    stat_l.append(np.abs(lm_ar).mean(axis=1))\n    stat_l.append(np.abs(lm_ar).std(axis=1))\n\n    stat_l.append(lm_ar.max(axis=1) / (np.abs(lm_ar.min(axis=1)) + 1e-32))\n    stat_l.append(lm_ar.max(axis=1) - np.abs(lm_ar.min(axis=1)))\n    stat_l.append(lm_ar.sum(axis=1))\n\n    stat_l.append(lm_df.quantile(0.99, axis=1).values)\n    stat_l.append(lm_df.quantile(0.95, axis=1).values)\n    stat_l.append(lm_df.quantile(0.1, axis=1).values)\n    stat_l.append(lm_df.quantile(0.05, axis=1).values)\n\n    stat_l.append(np.mean(np.diff(lm_ar), axis=1))\n    stat_l.append(np.max(np.diff(lm_ar), axis=1))\n    stat_l.append(np.min(np.diff(lm_ar), axis=1))\n    stat_l.append(np.mean(np.diff(np.diff(lm_ar)), axis=1))\n    stat_l.append(np.max(np.diff(np.diff(lm_ar)), axis=1))\n    stat_l.append(np.min(np.diff(np.diff(lm_ar)), axis=1))\n\n    return np.array(stat_l).T\n```",
      "votes": null
    },
    {
      "id": "564818",
      "postDate": "06/30/2019 03:36:17",
      "content": "<p>Thank you! It was helpful!</p>",
      "rawMarkdown": "Thank you! It was helpful!",
      "votes": null
    },
    {
      "id": "569320",
      "postDate": "07/06/2019 12:55:46",
      "content": "<p>Congrats.Power Point Master &amp; kaggle master\nyour models seem very complex\nhow to combine so many models into one?</p>",
      "rawMarkdown": "Congrats.Power Point Master &amp; kaggle master\nyour models seem very complex\nhow to combine so many models into one?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 564076,
      "author_name": "jmourad100",
      "author_url": "",
      "post_date": "06/29/2019 00:24:05",
      "content": "<p>Congrats. A pretty clear explanation of your approach. I have a question about that figure, what software did you use to create it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 564079,
          "author_name": "maxwell110",
          "author_url": "",
          "post_date": "06/29/2019 00:29:24",
          "content": "<p><a href=\"/jmourad100\">@jmourad100</a> \nThanks!\nIt may be unexpected for you, I used a just power point to create this figure. <br>\n( Someone called me <code>Power Point Master</code> not kaggle master ;) )</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 564088,
      "author_name": "bossimuimu",
      "author_url": "",
      "post_date": "06/29/2019 00:55:07",
      "content": "<p>Congrats.\nI am interested in what kind of features can be made.\nIf possible, could you tell me more about the content of \"25 statistical values\"?</p>",
      "votes": null,
      "replies": [
        {
          "id": 564098,
          "author_name": "maxwell110",
          "author_url": "",
          "post_date": "06/29/2019 01:24:24",
          "content": "<p><a href=\"/bossimuimu\">@bossimuimu</a> \nThanks for your question. <br>\nI made a helper function, modified version in <a href=\"https://www.kaggle.com/titericz/lightgbm-simple-solution-lb-0-203\">this kernel</a></p>\n\n<p>Hope this is what you asked.  </p>\n\n<p>```\ndef create_features(logmel_fn):\n    ## load log-mel array\n    lm_ar = mel_0_1(np.load(logmel_fn))\n    lm_df = pd.DataFrame(lm_ar)</p>\n\n<pre><code>stat_l = []\nstat_l.append(lm_df.mean(axis=1).values)\nstat_l.append(lm_df.median(axis=1).values)\nstat_l.append(lm_df.std(axis=1).values)\nstat_l.append(lm_df.max(axis=1).values)\nstat_l.append(lm_df.min(axis=1).values)\nstat_l.append(lm_df.skew(axis=1).values)\nstat_l.append(lm_df.mad(axis=1).values)\nstat_l.append(lm_df.kurtosis(axis=1).values)\n\nstat_l.append(np.abs(lm_ar).max(axis=1))\nstat_l.append(np.abs(lm_ar).min(axis=1))\nstat_l.append(np.abs(lm_ar).mean(axis=1))\nstat_l.append(np.abs(lm_ar).std(axis=1))\n\nstat_l.append(lm_ar.max(axis=1) / (np.abs(lm_ar.min(axis=1)) + 1e-32))\nstat_l.append(lm_ar.max(axis=1) - np.abs(lm_ar.min(axis=1)))\nstat_l.append(lm_ar.sum(axis=1))\n\nstat_l.append(lm_df.quantile(0.99, axis=1).values)\nstat_l.append(lm_df.quantile(0.95, axis=1).values)\nstat_l.append(lm_df.quantile(0.1, axis=1).values)\nstat_l.append(lm_df.quantile(0.05, axis=1).values)\n\nstat_l.append(np.mean(np.diff(lm_ar), axis=1))\nstat_l.append(np.max(np.diff(lm_ar), axis=1))\nstat_l.append(np.min(np.diff(lm_ar), axis=1))\nstat_l.append(np.mean(np.diff(np.diff(lm_ar)), axis=1))\nstat_l.append(np.max(np.diff(np.diff(lm_ar)), axis=1))\nstat_l.append(np.min(np.diff(np.diff(lm_ar)), axis=1))\n\nreturn np.array(stat_l).T\n</code></pre>\n\n<p>```</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 564818,
          "author_name": "bossimuimu",
          "author_url": "",
          "post_date": "06/30/2019 03:36:17",
          "content": "<p>Thank you! It was helpful!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 569320,
      "author_name": "abnerzhang",
      "author_url": "",
      "post_date": "07/06/2019 12:55:46",
      "content": "<p>Congrats.Power Point Master &amp; kaggle master\nyour models seem very complex\nhow to combine so many models into one?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "564067": "First of all, I'd like to thank organizers and all participants, and congratulate winners on excellent results !\nFor me, this is the first image(sound) competition in Kaggle and I have learned a lot.\n\nMany top teams unveiled their nice solutions(also scores). So I'm ashamed to publish my solution here, but I will do for a memento.  \n\n---\n\n# Local Validation\n\nI created 80class-balanced 5 folds validation and have evaluated the model performance with BCE. At the early stage in this competition, I have checked the correlation of BCE and LWLRAP. Below figure shows an almost good correlation.  \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F01d7ab43832e5eb5eec8a139fdb69c19%2Ffig1.png?generation=1561761823509063&amp;alt=media)\n\n\n# Features\n\n1. Remove silent audios\n2. Trim silent parts\n3. Log Mel Spectrogram ( SR 44.1kHz, FFT window size 80ms, Hop 10ms, Mel Bands 64 )\n4. Clustered frequency-wise statistical features\n\nI did not do nothing special. But I will explain about **4**.  \nI thought CNN can catch an audio property in frequency space like Fourier Analysis, but less in  statistics values ( max, min, ... ). And labels of train-noisy are unreliable, so I created frequency-wise statistical features from spectrograms and compute cluster distances without using label information(Unsupervised).  The procedure is following,  \n\na. Compute 25 statistical values (Max, Min, Primary difference, Secondary difference, etc... ) and flatten 64 x 25 (=1600) features \nb. Compute 200 cluster distances with MiniBatchKMeans ( dimensional reduction from 1600 to 200 )\n\nThis feature pushed my score about 0.5 ~ 1% in each model.  \n\n\n# Models, Train and Prediction\n\nIn this competition, we do not have a lot of time to make inferences ( less than 1 GPU hour ). So I selected 3 relatively light-weight models.  \n\n-  Mobile Net V2   with/without clustered features\n- ResNet50   with/without clustered features\n- DenseNet121   with/without clustered features\n\nMy final submission is the ensemble of above 6 models ( 3 models x 2 type of features ). Here is a model pipeline. The setting of train and prediction with TTA is written in this figure.  \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2Fd21b6e837c1be18cc4a66bbbc82c3174%2Ffs2019_model-pipeline_01.png?generation=1561760818137707&amp;alt=media)\n\n\n# Performance  \n\nI used the weighted geometric averaging to blend 6 model predictions. Below table shows each performance and the blending coefficients. Those coefficients are computed with optimization based on Out Of Fold predictions.  LWLRAP values are calculated with 5 fold OOF predictions on train-curated data.\n\n| Model | LWLRAP on train-curated | Blending Coefficient |\n| --- | --- | --- |\n| MobileNetV2 with clustered features( cf ) | 0.84519 | 0.246 |\n| MobileNetV2 without cf | 0.82940 | 0.217 |\n| ResNet50 with cf | 0.84490 | 0.149 |\n| ResNet50 without cf | 0.83006 | 0.072 |\n| DenseNet121 with cf | 0.84353 | 0.115 |\n| DenseNet121 without cf | 0.83501 | 0.201 |\n| **Blended** | 0.87611 ( Private 0.72820 ) | --- |\n\n\nThank you very much for reading to the end.  \nSee you at the next competition !",
    "564076": "Congrats. A pretty clear explanation of your approach. I have a question about that figure, what software did you use to create it.",
    "564079": "jmourad100 \nThanks!\nIt may be unexpected for you, I used a just power point to create this figure.  \n( Someone called me `Power Point Master` not kaggle master ;) )",
    "564088": "Congrats.\nI am interested in what kind of features can be made.\nIf possible, could you tell me more about the content of \"25 statistical values\"?",
    "564098": "bossimuimu \nThanks for your question.  \nI made a helper function, modified version in [this kernel](https://www.kaggle.com/titericz/lightgbm-simple-solution-lb-0-203)\n\nHope this is what you asked.  \n\n```\ndef create_features(logmel_fn):\n    ## load log-mel array\n    lm_ar = mel_0_1(np.load(logmel_fn))\n    lm_df = pd.DataFrame(lm_ar)\n\n    stat_l = []\n    stat_l.append(lm_df.mean(axis=1).values)\n    stat_l.append(lm_df.median(axis=1).values)\n    stat_l.append(lm_df.std(axis=1).values)\n    stat_l.append(lm_df.max(axis=1).values)\n    stat_l.append(lm_df.min(axis=1).values)\n    stat_l.append(lm_df.skew(axis=1).values)\n    stat_l.append(lm_df.mad(axis=1).values)\n    stat_l.append(lm_df.kurtosis(axis=1).values)\n\n    stat_l.append(np.abs(lm_ar).max(axis=1))\n    stat_l.append(np.abs(lm_ar).min(axis=1))\n    stat_l.append(np.abs(lm_ar).mean(axis=1))\n    stat_l.append(np.abs(lm_ar).std(axis=1))\n\n    stat_l.append(lm_ar.max(axis=1) / (np.abs(lm_ar.min(axis=1)) + 1e-32))\n    stat_l.append(lm_ar.max(axis=1) - np.abs(lm_ar.min(axis=1)))\n    stat_l.append(lm_ar.sum(axis=1))\n\n    stat_l.append(lm_df.quantile(0.99, axis=1).values)\n    stat_l.append(lm_df.quantile(0.95, axis=1).values)\n    stat_l.append(lm_df.quantile(0.1, axis=1).values)\n    stat_l.append(lm_df.quantile(0.05, axis=1).values)\n\n    stat_l.append(np.mean(np.diff(lm_ar), axis=1))\n    stat_l.append(np.max(np.diff(lm_ar), axis=1))\n    stat_l.append(np.min(np.diff(lm_ar), axis=1))\n    stat_l.append(np.mean(np.diff(np.diff(lm_ar)), axis=1))\n    stat_l.append(np.max(np.diff(np.diff(lm_ar)), axis=1))\n    stat_l.append(np.min(np.diff(np.diff(lm_ar)), axis=1))\n\n    return np.array(stat_l).T\n```",
    "564818": "Thank you! It was helpful!",
    "569320": "Congrats.Power Point Master &amp; kaggle master\nyour models seem very complex\nhow to combine so many models into one?"
  },
  "source": "meta"
}