{
  "id": 93295,
  "title": "Train with melspectrogram feature vectors or images?",
  "url": "/competitions/freesound-audio-tagging-2019/discussion/93295",
  "author_name": "",
  "post_date": "2019-05-25T11:14:14.457164200Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p>The baseline system use only feature vectors extracted from the wav files, which has only 1 channel. Some kernels convert the feature vectors into spectrogram images with 3 channels. </p>\n\n<p>Personally I find using feature vectors achieve very limited performance, only slightly above baseline system.</p>\n\n<p>What kind of input do you guys adopt? Is there any performance difference?</p>",
  "messages": [
    {
      "id": "536819",
      "postDate": "05/25/2019 11:14:14",
      "content": "<p>The baseline system use only feature vectors extracted from the wav files, which has only 1 channel. Some kernels convert the feature vectors into spectrogram images with 3 channels. </p>\n\n<p>Personally I find using feature vectors achieve very limited performance, only slightly above baseline system.</p>\n\n<p>What kind of input do you guys adopt? Is there any performance difference?</p>",
      "rawMarkdown": "The baseline system use only feature vectors extracted from the wav files, which has only 1 channel. Some kernels convert the feature vectors into spectrogram images with 3 channels. \n\nPersonally I find using feature vectors achieve very limited performance, only slightly above baseline system.\n \nWhat kind of input do you guys adopt? Is there any performance difference?",
      "votes": null
    },
    {
      "id": "537171",
      "postDate": "05/26/2019 10:40:58",
      "content": "<p>Converting to 3-channels is necessary if you use one of the well known model architectures for image classification without modifications. For me, custom model with 1-channel input works better.</p>",
      "rawMarkdown": "Converting to 3-channels is necessary if you use one of the well known model architectures for image classification without modifications. For me, custom model with 1-channel input works better.",
      "votes": null
    },
    {
      "id": "537187",
      "postDate": "05/26/2019 11:46:46",
      "content": "<p>I have never tried the  3-ch conversion because it's not plausible. At least, the 1-ch log-mel feature is enough to get 1st place in public LB.</p>",
      "rawMarkdown": "I have never tried the  3-ch conversion because it's not plausible. At least, the 1-ch log-mel feature is enough to get 1st place in public LB.",
      "votes": null
    },
    {
      "id": "537219",
      "postDate": "05/26/2019 13:31:49",
      "content": "<p>Thanks for your reply! Now I'm sure that my problem is not resulted from my data. </p>",
      "rawMarkdown": "Thanks for your reply! Now I'm sure that my problem is not resulted from my data.",
      "votes": null
    },
    {
      "id": "537220",
      "postDate": "05/26/2019 13:32:16",
      "content": "<p>Thanks a lot. I'll continue to work on my model.</p>",
      "rawMarkdown": "Thanks a lot. I'll continue to work on my model.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 537171,
      "author_name": "vzaguskin",
      "author_url": "",
      "post_date": "05/26/2019 10:40:58",
      "content": "<p>Converting to 3-channels is necessary if you use one of the well known model architectures for image classification without modifications. For me, custom model with 1-channel input works better.</p>",
      "votes": null,
      "replies": [
        {
          "id": 537220,
          "author_name": "alfredwei",
          "author_url": "",
          "post_date": "05/26/2019 13:32:16",
          "content": "<p>Thanks a lot. I'll continue to work on my model.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 537187,
      "author_name": "osciiart",
      "author_url": "",
      "post_date": "05/26/2019 11:46:46",
      "content": "<p>I have never tried the  3-ch conversion because it's not plausible. At least, the 1-ch log-mel feature is enough to get 1st place in public LB.</p>",
      "votes": null,
      "replies": [
        {
          "id": 537219,
          "author_name": "alfredwei",
          "author_url": "",
          "post_date": "05/26/2019 13:31:49",
          "content": "<p>Thanks for your reply! Now I'm sure that my problem is not resulted from my data. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "536819": "The baseline system use only feature vectors extracted from the wav files, which has only 1 channel. Some kernels convert the feature vectors into spectrogram images with 3 channels. \n\nPersonally I find using feature vectors achieve very limited performance, only slightly above baseline system.\n \nWhat kind of input do you guys adopt? Is there any performance difference?",
    "537171": "Converting to 3-channels is necessary if you use one of the well known model architectures for image classification without modifications. For me, custom model with 1-channel input works better.",
    "537187": "I have never tried the  3-ch conversion because it's not plausible. At least, the 1-ch log-mel feature is enough to get 1st place in public LB.",
    "537219": "Thanks for your reply! Now I'm sure that my problem is not resulted from my data.",
    "537220": "Thanks a lot. I'll continue to work on my model."
  },
  "source": "meta"
}