{
  "id": 216564,
  "title": "Audio Cropping is critical",
  "url": "/competitions/rfcx-species-audio-detection/discussion/216564",
  "author_name": "",
  "post_date": "2021-02-03T08:56:24.945199900Z",
  "votes": 16,
  "comment_count": 4,
  "views": 0,
  "content": "<p>In this competition many people crop the test data by a few seconds and use it for prediction. However, the sound to be predicted can be on the borders of clips. This makes prediction difficult.</p>\n<p>So during the prediction time you can slide the clip with some margin and predict many times.</p>\n<p>I slide the clip two seconds at a time, make three predictions, and then sum the predictions. This improves scores slightly. In my case from <code>0.847</code> to <code>0.853</code>(LB).</p>\n<p>And you can also increase the performance of the training process by adding more data with a little bit of sound cut off.</p>\n<p>Sample Code</p>\n<pre><code>SR = 32000\nHOP_LENGTH = 512\nTIME = 6\n\n\nclass TestDataFactory:\n    def __init__(self):\n        self.data_type = \"test\"\n        base_path = Path(\n            \"/this/is/My/Melspec/path\"\n        )\n        self.test_files = load_test_files(base_path)\n        self.recording_ids = list(self.test_files.keys())\n\n    def create_dataloader(self, batch_size, margin, per_file):\n        slide = slide\n        data = split_test_files(file_dict=self.test_files, per_file=per_file,\n                                margin=margin)\n        dataset = TestDataset(data)\n        return DataLoader(dataset, batch_size=batch_size * per_file,\n                          shuffle=False, drop_last=False)\n\n    def predict(self, model, device, batch_size, slide):\n        per_file = int((60.0 - slide) // TIME)\n        test_dataloader = self.create_dataloader(batch_size, slide, per_file)\n        recording_ids = []\n        probs = []\n        for row in test_dataloader:\n            mel = row[\"mel\"].to(device).float()\n            with torch.no_grad():\n                out = model.predict(mel)\n            # aggregate results by file\n            out = out.reshape((-1, per_file, CLASS_N))\n            out = torch.max(out, dim=1)[0].cpu().detach()\n            recording_id_batch = np.array(row[\"recording_id\"])\n            for rec in recording_id_batch:\n                assert len(np.unique(rec)) == 1\n            recording_id_batch = recording_id_batch.reshape((-1, per_file))\n            # recording_id_batch: (BATCH, PER_FILE)\n            recording_ids.append(recording_id_batch[:, 0])\n            probs.append(out.numpy())\n        recording_ids = np.concatenate(recording_ids)\n        probs = np.concatenate(probs)\n        df = pd.DataFrame({\n            'recording_id': recording_ids.tolist(),\n            **{f's{i}': probs[:, i] for i in range(CLASS_N)}\n        })\n        df.sort_values(\"recording_id\")\n        df = df.set_index(\"recording_id\")\n        return df\n\n\nclass TestDataset(Dataset):\n    def __init__(self, data):\n        self.data = data\n\n    def __len__(self):\n        return len(self.data)\n\n    def __getitem__(self, index):\n        return self.data[index]\n\n\ndef load_test_files(base_path: Path):\n    files = base_path.glob(\"*.npy\")\n    file_dicts = dict()\n    for file_path in files:\n        recording_id = file_path.stem\n        mel = np.load(str(file_path))\n        file_dicts[recording_id] = mel\n    return file_dicts\n\n\ndef split_test_files(file_dict, per_file, margin):\n    unit = SR / HOP_LENGTH\n    mel_length = int(TIME * unit)\n    return_values = []\n    for recording_id, mel in file_dict.items():\n        for i in range(per_file):\n            start_index = int(np.round(i * unit + margin * unit))\n            end_index = start_index + mel_length\n            image = spec_to_image(mel[:, start_index:end_index], HEIGHT, WIDTH)\n            assert image.shape[1] == mel_length\n            image = add_channel(image).astype(np.float32)\n            # mel = self.transform(torch.from_numpy(mel / 255.0))\n            image = image / 255.0\n            return_values.append(\n                dict(mel=image,\n                     recording_id=recording_id))\n    assert len(return_values) == per_file * len(file_dict)\n    return return_values\n</code></pre>",
  "messages": [
    {
      "id": "1183843",
      "postDate": "02/03/2021 08:56:24",
      "content": "<p>In this competition many people crop the test data by a few seconds and use it for prediction. However, the sound to be predicted can be on the borders of clips. This makes prediction difficult.</p>\n<p>So during the prediction time you can slide the clip with some margin and predict many times.</p>\n<p>I slide the clip two seconds at a time, make three predictions, and then sum the predictions. This improves scores slightly. In my case from <code>0.847</code> to <code>0.853</code>(LB).</p>\n<p>And you can also increase the performance of the training process by adding more data with a little bit of sound cut off.</p>\n<p>Sample Code</p>\n<pre><code>SR = 32000\nHOP_LENGTH = 512\nTIME = 6\n\n\nclass TestDataFactory:\n    def __init__(self):\n        self.data_type = \"test\"\n        base_path = Path(\n            \"/this/is/My/Melspec/path\"\n        )\n        self.test_files = load_test_files(base_path)\n        self.recording_ids = list(self.test_files.keys())\n\n    def create_dataloader(self, batch_size, margin, per_file):\n        slide = slide\n        data = split_test_files(file_dict=self.test_files, per_file=per_file,\n                                margin=margin)\n        dataset = TestDataset(data)\n        return DataLoader(dataset, batch_size=batch_size * per_file,\n                          shuffle=False, drop_last=False)\n\n    def predict(self, model, device, batch_size, slide):\n        per_file = int((60.0 - slide) // TIME)\n        test_dataloader = self.create_dataloader(batch_size, slide, per_file)\n        recording_ids = []\n        probs = []\n        for row in test_dataloader:\n            mel = row[\"mel\"].to(device).float()\n            with torch.no_grad():\n                out = model.predict(mel)\n            # aggregate results by file\n            out = out.reshape((-1, per_file, CLASS_N))\n            out = torch.max(out, dim=1)[0].cpu().detach()\n            recording_id_batch = np.array(row[\"recording_id\"])\n            for rec in recording_id_batch:\n                assert len(np.unique(rec)) == 1\n            recording_id_batch = recording_id_batch.reshape((-1, per_file))\n            # recording_id_batch: (BATCH, PER_FILE)\n            recording_ids.append(recording_id_batch[:, 0])\n            probs.append(out.numpy())\n        recording_ids = np.concatenate(recording_ids)\n        probs = np.concatenate(probs)\n        df = pd.DataFrame({\n            'recording_id': recording_ids.tolist(),\n            **{f's{i}': probs[:, i] for i in range(CLASS_N)}\n        })\n        df.sort_values(\"recording_id\")\n        df = df.set_index(\"recording_id\")\n        return df\n\n\nclass TestDataset(Dataset):\n    def __init__(self, data):\n        self.data = data\n\n    def __len__(self):\n        return len(self.data)\n\n    def __getitem__(self, index):\n        return self.data[index]\n\n\ndef load_test_files(base_path: Path):\n    files = base_path.glob(\"*.npy\")\n    file_dicts = dict()\n    for file_path in files:\n        recording_id = file_path.stem\n        mel = np.load(str(file_path))\n        file_dicts[recording_id] = mel\n    return file_dicts\n\n\ndef split_test_files(file_dict, per_file, margin):\n    unit = SR / HOP_LENGTH\n    mel_length = int(TIME * unit)\n    return_values = []\n    for recording_id, mel in file_dict.items():\n        for i in range(per_file):\n            start_index = int(np.round(i * unit + margin * unit))\n            end_index = start_index + mel_length\n            image = spec_to_image(mel[:, start_index:end_index], HEIGHT, WIDTH)\n            assert image.shape[1] == mel_length\n            image = add_channel(image).astype(np.float32)\n            # mel = self.transform(torch.from_numpy(mel / 255.0))\n            image = image / 255.0\n            return_values.append(\n                dict(mel=image,\n                     recording_id=recording_id))\n    assert len(return_values) == per_file * len(file_dict)\n    return return_values\n</code></pre>",
      "rawMarkdown": "In this competition many people crop the test data by a few seconds and use it for prediction. However, the sound to be predicted can be on the borders of clips. This makes prediction difficult.\n\nSo during the prediction time you can slide the clip with some margin and predict many times.\n\nI slide the clip two seconds at a time, make three predictions, and then sum the predictions. This improves scores slightly. In my case from `0.847 ` to `0.853`(LB).\n\nAnd you can also increase the performance of the training process by adding more data with a little bit of sound cut off.\n\nSample Code\n\n```python\nSR = 32000\nHOP_LENGTH = 512\nTIME = 6\n\n\nclass TestDataFactory:\n    def __init__(self):\n        self.data_type = \"test\"\n        base_path = Path(\n            \"/this/is/My/Melspec/path\"\n        )\n        self.test_files = load_test_files(base_path)\n        self.recording_ids = list(self.test_files.keys())\n\n    def create_dataloader(self, batch_size, margin, per_file):\n        slide = slide\n        data = split_test_files(file_dict=self.test_files, per_file=per_file,\n                                margin=margin)\n        dataset = TestDataset(data)\n        return DataLoader(dataset, batch_size=batch_size * per_file,\n                          shuffle=False, drop_last=False)\n\n    def predict(self, model, device, batch_size, slide):\n        per_file = int((60.0 - slide) // TIME)\n        test_dataloader = self.create_dataloader(batch_size, slide, per_file)\n        recording_ids = []\n        probs = []\n        for row in test_dataloader:\n            mel = row[\"mel\"].to(device).float()\n            with torch.no_grad():\n                out = model.predict(mel)\n            # aggregate results by file\n            out = out.reshape((-1, per_file, CLASS_N))\n            out = torch.max(out, dim=1)[0].cpu().detach()\n            recording_id_batch = np.array(row[\"recording_id\"])\n            for rec in recording_id_batch:\n                assert len(np.unique(rec)) == 1\n            recording_id_batch = recording_id_batch.reshape((-1, per_file))\n            # recording_id_batch: (BATCH, PER_FILE)\n            recording_ids.append(recording_id_batch[:, 0])\n            probs.append(out.numpy())\n        recording_ids = np.concatenate(recording_ids)\n        probs = np.concatenate(probs)\n        df = pd.DataFrame({\n            'recording_id': recording_ids.tolist(),\n            **{f's{i}': probs[:, i] for i in range(CLASS_N)}\n        })\n        df.sort_values(\"recording_id\")\n        df = df.set_index(\"recording_id\")\n        return df\n\n\nclass TestDataset(Dataset):\n    def __init__(self, data):\n        self.data = data\n\n    def __len__(self):\n        return len(self.data)\n\n    def __getitem__(self, index):\n        return self.data[index]\n\n\ndef load_test_files(base_path: Path):\n    files = base_path.glob(\"*.npy\")\n    file_dicts = dict()\n    for file_path in files:\n        recording_id = file_path.stem\n        mel = np.load(str(file_path))\n        file_dicts[recording_id] = mel\n    return file_dicts\n\n\ndef split_test_files(file_dict, per_file, margin):\n    unit = SR / HOP_LENGTH\n    mel_length = int(TIME * unit)\n    return_values = []\n    for recording_id, mel in file_dict.items():\n        for i in range(per_file):\n            start_index = int(np.round(i * unit + margin * unit))\n            end_index = start_index + mel_length\n            image = spec_to_image(mel[:, start_index:end_index], HEIGHT, WIDTH)\n            assert image.shape[1] == mel_length\n            image = add_channel(image).astype(np.float32)\n            # mel = self.transform(torch.from_numpy(mel / 255.0))\n            image = image / 255.0\n            return_values.append(\n                dict(mel=image,\n                     recording_id=recording_id))\n    assert len(return_values) == per_file * len(file_dict)\n    return return_values\n```",
      "votes": null
    },
    {
      "id": "1184661",
      "postDate": "02/03/2021 16:55:11",
      "content": "<p>Thank, I will try it, </p>\n<p>How do you crop the audio in the training?</p>",
      "rawMarkdown": "Thank, I will try it, \n\nHow do you crop the audio in the training?",
      "votes": null
    },
    {
      "id": "1185139",
      "postDate": "02/04/2021 01:23:59",
      "content": "<p>It depends on the length of the target sound. If the target sound is long, I determine the window size to include at least 2 or 3 seconds. If the target sound is short, I include everything.</p>",
      "rawMarkdown": "It depends on the length of the target sound. If the target sound is long, I determine the window size to include at least 2 or 3 seconds. If the target sound is short, I include everything.",
      "votes": null
    },
    {
      "id": "1186281",
      "postDate": "02/04/2021 18:08:01",
      "content": "<p>I applied this to one of my already trained model, and the LB increased 0.5 in my case as well.</p>\n<p>Thanks! 👌</p>",
      "rawMarkdown": "I applied this to one of my already trained model, and the LB increased 0.5 in my case as well.\n\nThanks! 👌",
      "votes": null
    },
    {
      "id": "1190376",
      "postDate": "02/07/2021 17:13:06",
      "content": "<p>Curious to see how top players crop. </p>",
      "rawMarkdown": "Curious to see how top players crop.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1184661,
      "author_name": "truonghoang",
      "author_url": "",
      "post_date": "02/03/2021 16:55:11",
      "content": "<p>Thank, I will try it, </p>\n<p>How do you crop the audio in the training?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1185139,
          "author_name": "takadaat",
          "author_url": "",
          "post_date": "02/04/2021 01:23:59",
          "content": "<p>It depends on the length of the target sound. If the target sound is long, I determine the window size to include at least 2 or 3 seconds. If the target sound is short, I include everything.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1186281,
      "author_name": "snnclsr",
      "author_url": "",
      "post_date": "02/04/2021 18:08:01",
      "content": "<p>I applied this to one of my already trained model, and the LB increased 0.5 in my case as well.</p>\n<p>Thanks! 👌</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1190376,
      "author_name": "dannywu375",
      "author_url": "",
      "post_date": "02/07/2021 17:13:06",
      "content": "<p>Curious to see how top players crop. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1183843": "In this competition many people crop the test data by a few seconds and use it for prediction. However, the sound to be predicted can be on the borders of clips. This makes prediction difficult.\n\nSo during the prediction time you can slide the clip with some margin and predict many times.\n\nI slide the clip two seconds at a time, make three predictions, and then sum the predictions. This improves scores slightly. In my case from `0.847 ` to `0.853`(LB).\n\nAnd you can also increase the performance of the training process by adding more data with a little bit of sound cut off.\n\nSample Code\n\n```python\nSR = 32000\nHOP_LENGTH = 512\nTIME = 6\n\n\nclass TestDataFactory:\n    def __init__(self):\n        self.data_type = \"test\"\n        base_path = Path(\n            \"/this/is/My/Melspec/path\"\n        )\n        self.test_files = load_test_files(base_path)\n        self.recording_ids = list(self.test_files.keys())\n\n    def create_dataloader(self, batch_size, margin, per_file):\n        slide = slide\n        data = split_test_files(file_dict=self.test_files, per_file=per_file,\n                                margin=margin)\n        dataset = TestDataset(data)\n        return DataLoader(dataset, batch_size=batch_size * per_file,\n                          shuffle=False, drop_last=False)\n\n    def predict(self, model, device, batch_size, slide):\n        per_file = int((60.0 - slide) // TIME)\n        test_dataloader = self.create_dataloader(batch_size, slide, per_file)\n        recording_ids = []\n        probs = []\n        for row in test_dataloader:\n            mel = row[\"mel\"].to(device).float()\n            with torch.no_grad():\n                out = model.predict(mel)\n            # aggregate results by file\n            out = out.reshape((-1, per_file, CLASS_N))\n            out = torch.max(out, dim=1)[0].cpu().detach()\n            recording_id_batch = np.array(row[\"recording_id\"])\n            for rec in recording_id_batch:\n                assert len(np.unique(rec)) == 1\n            recording_id_batch = recording_id_batch.reshape((-1, per_file))\n            # recording_id_batch: (BATCH, PER_FILE)\n            recording_ids.append(recording_id_batch[:, 0])\n            probs.append(out.numpy())\n        recording_ids = np.concatenate(recording_ids)\n        probs = np.concatenate(probs)\n        df = pd.DataFrame({\n            'recording_id': recording_ids.tolist(),\n            **{f's{i}': probs[:, i] for i in range(CLASS_N)}\n        })\n        df.sort_values(\"recording_id\")\n        df = df.set_index(\"recording_id\")\n        return df\n\n\nclass TestDataset(Dataset):\n    def __init__(self, data):\n        self.data = data\n\n    def __len__(self):\n        return len(self.data)\n\n    def __getitem__(self, index):\n        return self.data[index]\n\n\ndef load_test_files(base_path: Path):\n    files = base_path.glob(\"*.npy\")\n    file_dicts = dict()\n    for file_path in files:\n        recording_id = file_path.stem\n        mel = np.load(str(file_path))\n        file_dicts[recording_id] = mel\n    return file_dicts\n\n\ndef split_test_files(file_dict, per_file, margin):\n    unit = SR / HOP_LENGTH\n    mel_length = int(TIME * unit)\n    return_values = []\n    for recording_id, mel in file_dict.items():\n        for i in range(per_file):\n            start_index = int(np.round(i * unit + margin * unit))\n            end_index = start_index + mel_length\n            image = spec_to_image(mel[:, start_index:end_index], HEIGHT, WIDTH)\n            assert image.shape[1] == mel_length\n            image = add_channel(image).astype(np.float32)\n            # mel = self.transform(torch.from_numpy(mel / 255.0))\n            image = image / 255.0\n            return_values.append(\n                dict(mel=image,\n                     recording_id=recording_id))\n    assert len(return_values) == per_file * len(file_dict)\n    return return_values\n```",
    "1184661": "Thank, I will try it, \n\nHow do you crop the audio in the training?",
    "1185139": "It depends on the length of the target sound. If the target sound is long, I determine the window size to include at least 2 or 3 seconds. If the target sound is short, I include everything.",
    "1186281": "I applied this to one of my already trained model, and the LB increased 0.5 in my case as well.\n\nThanks! 👌",
    "1190376": "Curious to see how top players crop."
  },
  "source": "meta"
}