{
  "id": 93337,
  "title": "Xie, Junyuan, et al. “Bag of Tricks for Image Classification with Convolutional Neural Networks.”",
  "url": "/competitions/freesound-audio-tagging-2019/discussion/93337",
  "author_name": "",
  "post_date": "2019-05-26T01:07:31.441416900Z",
  "votes": 3,
  "comment_count": 3,
  "views": 0,
  "content": "<p>This competition is basically an image classification problem if we convert audio to a spectrogram. This paper covers a lot of tricks including mixup: <a href=\"https://arxiv.org/pdf/1812.01187.pdf\">https://arxiv.org/pdf/1812.01187.pdf</a></p>",
  "messages": [
    {
      "id": "537012",
      "postDate": "05/26/2019 01:07:31",
      "content": "<p>This competition is basically an image classification problem if we convert audio to a spectrogram. This paper covers a lot of tricks including mixup: <a href=\"https://arxiv.org/pdf/1812.01187.pdf\">https://arxiv.org/pdf/1812.01187.pdf</a></p>",
      "rawMarkdown": "This competition is basically an image classification problem if we convert audio to a spectrogram. This paper covers a lot of tricks including mixup: https://arxiv.org/pdf/1812.01187.pdf",
      "votes": null
    },
    {
      "id": "537297",
      "postDate": "05/26/2019 16:35:46",
      "content": "<p>&gt; This competition is basically an image classification problem if we convert audio to a spectrogram.</p>\n\n<p>I wouldn't be too sure about that :)</p>\n\n<p>It doesn't make these methods useless, though.</p>",
      "rawMarkdown": "&gt; This competition is basically an image classification problem if we convert audio to a spectrogram.\n\nI wouldn't be too sure about that :)\n\nIt doesn't make these methods useless, though.",
      "votes": null
    },
    {
      "id": "537348",
      "postDate": "05/26/2019 19:43:42",
      "content": "<p>For extra features,  I was wondering all the time if it's possible to add them into additional dimension (4th dimension for RGB spectrogram). Problem is that you would need to crop/pad tensors to the same size.</p>\n\n<p>Probably, the only option is to stack mel-trained network (freezing its weights) with another classifier and feed additional features there...</p>\n\n<p>...and there is already good kernel on feature extraction <a href=\"https://www.kaggle.com/varanr/audio-feature-extraction\">https://www.kaggle.com/varanr/audio-feature-extraction</a></p>\n\n<p>P.S. however I'm mostly guessing since I know nothing about audio processing :)</p>",
      "rawMarkdown": "For extra features,  I was wondering all the time if it's possible to add them into additional dimension (4th dimension for RGB spectrogram). Problem is that you would need to crop/pad tensors to the same size.\n\nProbably, the only option is to stack mel-trained network (freezing its weights) with another classifier and feed additional features there...\n\n...and there is already good kernel on feature extraction https://www.kaggle.com/varanr/audio-feature-extraction\n\nP.S. however I'm mostly guessing since I know nothing about audio processing :)",
      "votes": null
    },
    {
      "id": "537350",
      "postDate": "05/26/2019 19:47:23",
      "content": "<p>My own question about this challenge (and MEL in particular) is what size of network is sufficient?\nI believe last year winners used variations of MobileNet. Current public kernel either use small custom CNN or Inception3, which is relatively huge. </p>",
      "rawMarkdown": "My own question about this challenge (and MEL in particular) is what size of network is sufficient?\nI believe last year winners used variations of MobileNet. Current public kernel either use small custom CNN or Inception3, which is relatively huge.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 537297,
      "author_name": "ddanevskyi",
      "author_url": "",
      "post_date": "05/26/2019 16:35:46",
      "content": "<p>&gt; This competition is basically an image classification problem if we convert audio to a spectrogram.</p>\n\n<p>I wouldn't be too sure about that :)</p>\n\n<p>It doesn't make these methods useless, though.</p>",
      "votes": null,
      "replies": [
        {
          "id": 537348,
          "author_name": "vandalko",
          "author_url": "",
          "post_date": "05/26/2019 19:43:42",
          "content": "<p>For extra features,  I was wondering all the time if it's possible to add them into additional dimension (4th dimension for RGB spectrogram). Problem is that you would need to crop/pad tensors to the same size.</p>\n\n<p>Probably, the only option is to stack mel-trained network (freezing its weights) with another classifier and feed additional features there...</p>\n\n<p>...and there is already good kernel on feature extraction <a href=\"https://www.kaggle.com/varanr/audio-feature-extraction\">https://www.kaggle.com/varanr/audio-feature-extraction</a></p>\n\n<p>P.S. however I'm mostly guessing since I know nothing about audio processing :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 537350,
      "author_name": "vandalko",
      "author_url": "",
      "post_date": "05/26/2019 19:47:23",
      "content": "<p>My own question about this challenge (and MEL in particular) is what size of network is sufficient?\nI believe last year winners used variations of MobileNet. Current public kernel either use small custom CNN or Inception3, which is relatively huge. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "537012": "This competition is basically an image classification problem if we convert audio to a spectrogram. This paper covers a lot of tricks including mixup: https://arxiv.org/pdf/1812.01187.pdf",
    "537297": "&gt; This competition is basically an image classification problem if we convert audio to a spectrogram.\n\nI wouldn't be too sure about that :)\n\nIt doesn't make these methods useless, though.",
    "537348": "For extra features,  I was wondering all the time if it's possible to add them into additional dimension (4th dimension for RGB spectrogram). Problem is that you would need to crop/pad tensors to the same size.\n\nProbably, the only option is to stack mel-trained network (freezing its weights) with another classifier and feed additional features there...\n\n...and there is already good kernel on feature extraction https://www.kaggle.com/varanr/audio-feature-extraction\n\nP.S. however I'm mostly guessing since I know nothing about audio processing :)",
    "537350": "My own question about this challenge (and MEL in particular) is what size of network is sufficient?\nI believe last year winners used variations of MobileNet. Current public kernel either use small custom CNN or Inception3, which is relatively huge."
  },
  "source": "meta"
}