{
  "id": 89927,
  "title": "Challenge Baseline system released",
  "url": "/competitions/freesound-audio-tagging-2019/discussion/89927",
  "author_name": "",
  "post_date": "2019-04-18T19:22:13.646871400Z",
  "votes": 27,
  "comment_count": 2,
  "views": 0,
  "content": "<p>The challenge baseline system is now available at\n<a href=\"https://github.com/DCASE-REPO/dcase2019_task2_baseline\">https://github.com/DCASE-REPO/dcase2019_task2_baseline</a></p>\n\n<p>Some highlights</p>\n\n<ul>\n<li><p>The baseline model is a variant of the MobileNet v1 architecture. This is ~8x smaller than a ResNet-50 with ~4x less compute. This architecture can be run in real-time on a mobile or other resource-constrained device without a large loss in accuracy.</p></li>\n<li><p>We use log mel spectrogram input features which are computed on the fly in TensorFlow from WAV files. This allows you to play with feature generation hyperparameters in your grid searches, and also could be sped up if your TensorFlow installation includes GPU-accelerated kernels for FFT and other signal processing ops.</p></li>\n<li><p>We use label smoothing and dropout to handle label noise, and warm-start training from another training run to allow transfer learning from noisy to curated datasets which combats the domain mismatch. We have included some ideas for further improvement in the documentation.</p></li>\n<li><p>The baseline achieves lwlraps of ~0.546 on the full test set and ~0.537 on the public leaderboard. Inference time in a kernel on the full test set (4x the size of the public test set) is ~2 min on GPU and ~45 min on CPU. </p></li>\n</ul>\n\n<p>Good luck!</p>\n\n<p>Manoj\n(on behalf of all the challenge organizers)</p>",
  "messages": [
    {
      "id": "519342",
      "postDate": "04/18/2019 19:22:13",
      "content": "<p>The challenge baseline system is now available at\n<a href=\"https://github.com/DCASE-REPO/dcase2019_task2_baseline\">https://github.com/DCASE-REPO/dcase2019_task2_baseline</a></p>\n\n<p>Some highlights</p>\n\n<ul>\n<li><p>The baseline model is a variant of the MobileNet v1 architecture. This is ~8x smaller than a ResNet-50 with ~4x less compute. This architecture can be run in real-time on a mobile or other resource-constrained device without a large loss in accuracy.</p></li>\n<li><p>We use log mel spectrogram input features which are computed on the fly in TensorFlow from WAV files. This allows you to play with feature generation hyperparameters in your grid searches, and also could be sped up if your TensorFlow installation includes GPU-accelerated kernels for FFT and other signal processing ops.</p></li>\n<li><p>We use label smoothing and dropout to handle label noise, and warm-start training from another training run to allow transfer learning from noisy to curated datasets which combats the domain mismatch. We have included some ideas for further improvement in the documentation.</p></li>\n<li><p>The baseline achieves lwlraps of ~0.546 on the full test set and ~0.537 on the public leaderboard. Inference time in a kernel on the full test set (4x the size of the public test set) is ~2 min on GPU and ~45 min on CPU. </p></li>\n</ul>\n\n<p>Good luck!</p>\n\n<p>Manoj\n(on behalf of all the challenge organizers)</p>",
      "rawMarkdown": "The challenge baseline system is now available at\nhttps://github.com/DCASE-REPO/dcase2019_task2_baseline\n\nSome highlights\n\n- The baseline model is a variant of the MobileNet v1 architecture. This is ~8x smaller than a ResNet-50 with ~4x less compute. This architecture can be run in real-time on a mobile or other resource-constrained device without a large loss in accuracy.\n\n- We use log mel spectrogram input features which are computed on the fly in TensorFlow from WAV files. This allows you to play with feature generation hyperparameters in your grid searches, and also could be sped up if your TensorFlow installation includes GPU-accelerated kernels for FFT and other signal processing ops.\n\n- We use label smoothing and dropout to handle label noise, and warm-start training from another training run to allow transfer learning from noisy to curated datasets which combats the domain mismatch. We have included some ideas for further improvement in the documentation.\n\n- The baseline achieves lwlraps of ~0.546 on the full test set and ~0.537 on the public leaderboard. Inference time in a kernel on the full test set (4x the size of the public test set) is ~2 min on GPU and ~45 min on CPU. \n\nGood luck!\n\nManoj\n(on behalf of all the challenge organizers)",
      "votes": null
    },
    {
      "id": "519353",
      "postDate": "04/18/2019 19:47:24",
      "content": "<p>Awesome thanks</p>",
      "rawMarkdown": "Awesome thanks",
      "votes": null
    },
    {
      "id": "539381",
      "postDate": "05/30/2019 02:20:26",
      "content": "<p>Good sharing, thanks</p>",
      "rawMarkdown": "Good sharing, thanks",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 519353,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "04/18/2019 19:47:24",
      "content": "<p>Awesome thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 539381,
      "author_name": "guoyang0601",
      "author_url": "",
      "post_date": "05/30/2019 02:20:26",
      "content": "<p>Good sharing, thanks</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "519342": "The challenge baseline system is now available at\nhttps://github.com/DCASE-REPO/dcase2019_task2_baseline\n\nSome highlights\n\n- The baseline model is a variant of the MobileNet v1 architecture. This is ~8x smaller than a ResNet-50 with ~4x less compute. This architecture can be run in real-time on a mobile or other resource-constrained device without a large loss in accuracy.\n\n- We use log mel spectrogram input features which are computed on the fly in TensorFlow from WAV files. This allows you to play with feature generation hyperparameters in your grid searches, and also could be sped up if your TensorFlow installation includes GPU-accelerated kernels for FFT and other signal processing ops.\n\n- We use label smoothing and dropout to handle label noise, and warm-start training from another training run to allow transfer learning from noisy to curated datasets which combats the domain mismatch. We have included some ideas for further improvement in the documentation.\n\n- The baseline achieves lwlraps of ~0.546 on the full test set and ~0.537 on the public leaderboard. Inference time in a kernel on the full test set (4x the size of the public test set) is ~2 min on GPU and ~45 min on CPU. \n\nGood luck!\n\nManoj\n(on behalf of all the challenge organizers)",
    "519353": "Awesome thanks",
    "539381": "Good sharing, thanks"
  },
  "source": "meta"
}