{
  "id": 243378,
  "title": "Postprocessing",
  "url": "/competitions/birdclef-2021/discussion/243378",
  "author_name": "",
  "post_date": "2021-06-02T09:37:02.885201200Z",
  "votes": 23,
  "comment_count": 4,
  "views": 0,
  "content": "<p>After reading some top solutions writeup I see that there are two classes.  <a href=\"https://www.kaggle.com/tezdhar\" target=\"_blank\">@tezdhar</a> me, and probably others trained on 5 seconds clips while others used longer clips with SED model variants.  The former misses both long term interactions and species interactions.  It must be completed by some post processing to capture them.  Here is the one I used to post process soundscape predictions.  It is based on two ideas:</p>\n<ol>\n<li><p>If a bird is heard in one of the 5 second clips of a soundscape then it is more likely to be hear in others.  This was used by top solutions in Cornell competition.  We compute the maximum prediction for each species across all 5 second clips to determine presence of bird species in the soundscape.  If a bird species is detected then the logits for this species are increased for all 5 seconds clips. </p></li>\n<li><p>If two bird species are known to co exist then hearing one makes the other one more likely.  For this I computed a co occurrence matrix using primary and secondary labels.  If a species is detected then logits for all co occurring species are increased.</p></li>\n</ol>\n<p>This logic is controlled by three thresholds tuned on train soundscape:</p>\n<ul>\n<li><code>thr</code>.   Logits greater than it correspond to a presence</li>\n<li>'incr'.  Logits above it for a present species are rounded to 1 in prediction.</li>\n<li>ocur_thr.  Amount by which logits of co occurring species are augmented.</li>\n</ul>\n<p>Example values are:</p>\n<pre><code>     'thr':2.1, \n\n      'incr':-0.2,\n\n      'occur_thr':0.9,\n</code></pre>\n<p>The code of my prediction function is this.  It takes as input the logits for all 5 second clips for a soundscape.  This postprocessing worked extremely well (maybe too well).</p>\n<pre><code>occur = np.zeros((len(BIRD_CODE), len(BIRD_CODE)), dtype='int')\n\ndf = train[train.primary_label != 'nocall']\n\nfor primary, secondary in zip(df.primary_label, df.secondary_labels):\n    for label in secondary:\n        occur[BIRD_CODE[primary], BIRD_CODE[label]] += 1\n        occur[BIRD_CODE[label], BIRD_CODE[primary]] += 1\n    #occur[BIRD_CODE[primary], BIRD_CODE[primary]] = 1\n\noccur = occur.clip(0,1)\n\ndef compute_pred(LOGITS, thr, incr, occur_thr):\n    all_preds = []\n    for file_path, logits in zip(test_audio, LOGITS):\n        fileinfo = file_path.split(os.sep)[-1].rsplit('.', 1)[0].split('_')\n        audio_id = int(fileinfo[0])\n        site = fileinfo[1]\n        logits_max = logits.max(0)\n        logits = logits.copy()\n        if occur_thr is not None and occur_thr != 0:\n            present = ((logits_max &gt; thr) * 1).reshape((1, -1))\n            secondary = np.matmul(present, occur)\n            for j in range(logits.shape[1]):\n                if secondary[:, j] &gt; 0:\n                    logits[:, j] += occur_thr\n            logits_max = logits.max(0)\n        for j in range(logits.shape[1]):\n            if logits_max[j] &gt; thr:\n                logits[:, j] += thr - incr\n        for seconds, logit in zip(range(5, 605, 5), logits):\n            row_id = str(audio_id )+ '_' + site + '_' + str(seconds)\n            birds = list(INV_BIRD_CODE[logit &gt;= thr])\n            if 'nocall' in birds and len(birds) &gt; 1:\n                birds = [p for p in birds if p!= 'nocall']\n            elif len(birds) == 0:\n                birds = ['nocall']\n            birds = ' '.join(birds)\n\n            all_preds.append((row_id, site, audio_id, seconds, birds))\n    df = pd.DataFrame().from_records(all_preds)\n    df.columns = ['row_id', 'site', 'audio_id', 'seconds', 'birds']\n    df = df.sort_values(['site', 'audio_id', 'seconds']).reset_index(drop=True)\n    return df\n</code></pre>",
  "messages": [
    {
      "id": "1332773",
      "postDate": "06/02/2021 09:37:02",
      "content": "<p>After reading some top solutions writeup I see that there are two classes.  <a href=\"https://www.kaggle.com/tezdhar\" target=\"_blank\">@tezdhar</a> me, and probably others trained on 5 seconds clips while others used longer clips with SED model variants.  The former misses both long term interactions and species interactions.  It must be completed by some post processing to capture them.  Here is the one I used to post process soundscape predictions.  It is based on two ideas:</p>\n<ol>\n<li><p>If a bird is heard in one of the 5 second clips of a soundscape then it is more likely to be hear in others.  This was used by top solutions in Cornell competition.  We compute the maximum prediction for each species across all 5 second clips to determine presence of bird species in the soundscape.  If a bird species is detected then the logits for this species are increased for all 5 seconds clips. </p></li>\n<li><p>If two bird species are known to co exist then hearing one makes the other one more likely.  For this I computed a co occurrence matrix using primary and secondary labels.  If a species is detected then logits for all co occurring species are increased.</p></li>\n</ol>\n<p>This logic is controlled by three thresholds tuned on train soundscape:</p>\n<ul>\n<li><code>thr</code>.   Logits greater than it correspond to a presence</li>\n<li>'incr'.  Logits above it for a present species are rounded to 1 in prediction.</li>\n<li>ocur_thr.  Amount by which logits of co occurring species are augmented.</li>\n</ul>\n<p>Example values are:</p>\n<pre><code>     'thr':2.1, \n\n      'incr':-0.2,\n\n      'occur_thr':0.9,\n</code></pre>\n<p>The code of my prediction function is this.  It takes as input the logits for all 5 second clips for a soundscape.  This postprocessing worked extremely well (maybe too well).</p>\n<pre><code>occur = np.zeros((len(BIRD_CODE), len(BIRD_CODE)), dtype='int')\n\ndf = train[train.primary_label != 'nocall']\n\nfor primary, secondary in zip(df.primary_label, df.secondary_labels):\n    for label in secondary:\n        occur[BIRD_CODE[primary], BIRD_CODE[label]] += 1\n        occur[BIRD_CODE[label], BIRD_CODE[primary]] += 1\n    #occur[BIRD_CODE[primary], BIRD_CODE[primary]] = 1\n\noccur = occur.clip(0,1)\n\ndef compute_pred(LOGITS, thr, incr, occur_thr):\n    all_preds = []\n    for file_path, logits in zip(test_audio, LOGITS):\n        fileinfo = file_path.split(os.sep)[-1].rsplit('.', 1)[0].split('_')\n        audio_id = int(fileinfo[0])\n        site = fileinfo[1]\n        logits_max = logits.max(0)\n        logits = logits.copy()\n        if occur_thr is not None and occur_thr != 0:\n            present = ((logits_max &gt; thr) * 1).reshape((1, -1))\n            secondary = np.matmul(present, occur)\n            for j in range(logits.shape[1]):\n                if secondary[:, j] &gt; 0:\n                    logits[:, j] += occur_thr\n            logits_max = logits.max(0)\n        for j in range(logits.shape[1]):\n            if logits_max[j] &gt; thr:\n                logits[:, j] += thr - incr\n        for seconds, logit in zip(range(5, 605, 5), logits):\n            row_id = str(audio_id )+ '_' + site + '_' + str(seconds)\n            birds = list(INV_BIRD_CODE[logit &gt;= thr])\n            if 'nocall' in birds and len(birds) &gt; 1:\n                birds = [p for p in birds if p!= 'nocall']\n            elif len(birds) == 0:\n                birds = ['nocall']\n            birds = ' '.join(birds)\n\n            all_preds.append((row_id, site, audio_id, seconds, birds))\n    df = pd.DataFrame().from_records(all_preds)\n    df.columns = ['row_id', 'site', 'audio_id', 'seconds', 'birds']\n    df = df.sort_values(['site', 'audio_id', 'seconds']).reset_index(drop=True)\n    return df\n</code></pre>",
      "rawMarkdown": "After reading some top solutions writeup I see that there are two classes.  @tezdhar me, and probably others trained on 5 seconds clips while others used longer clips with SED model variants.  The former misses both long term interactions and species interactions.  It must be completed by some post processing to capture them.  Here is the one I used to post process soundscape predictions.  It is based on two ideas:\n\n1. If a bird is heard in one of the 5 second clips of a soundscape then it is more likely to be hear in others.  This was used by top solutions in Cornell competition.  We compute the maximum prediction for each species across all 5 second clips to determine presence of bird species in the soundscape.  If a bird species is detected then the logits for this species are increased for all 5 seconds clips. \n\n2. If two bird species are known to co exist then hearing one makes the other one more likely.  For this I computed a co occurrence matrix using primary and secondary labels.  If a species is detected then logits for all co occurring species are increased.\n\nThis logic is controlled by three thresholds tuned on train soundscape:\n\n- `thr`.   Logits greater than it correspond to a presence\n- 'incr'.  Logits above it for a present species are rounded to 1 in prediction.\n- ocur_thr.  Amount by which logits of co occurring species are augmented.\n\nExample values are:\n\n         'thr':2.1, \n\n          'incr':-0.2,\n\n          'occur_thr':0.9,\n\nThe code of my prediction function is this.  It takes as input the logits for all 5 second clips for a soundscape.  This postprocessing worked extremely well (maybe too well).\n\n```\n\noccur = np.zeros((len(BIRD_CODE), len(BIRD_CODE)), dtype='int')\n\ndf = train[train.primary_label != 'nocall']\n\nfor primary, secondary in zip(df.primary_label, df.secondary_labels):\n    for label in secondary:\n        occur[BIRD_CODE[primary], BIRD_CODE[label]] += 1\n        occur[BIRD_CODE[label], BIRD_CODE[primary]] += 1\n    #occur[BIRD_CODE[primary], BIRD_CODE[primary]] = 1\n        \noccur = occur.clip(0,1)\n\ndef compute_pred(LOGITS, thr, incr, occur_thr):\n    all_preds = []\n    for file_path, logits in zip(test_audio, LOGITS):\n        fileinfo = file_path.split(os.sep)[-1].rsplit('.', 1)[0].split('_')\n        audio_id = int(fileinfo[0])\n        site = fileinfo[1]\n        logits_max = logits.max(0)\n        logits = logits.copy()\n        if occur_thr is not None and occur_thr != 0:\n            present = ((logits_max > thr) * 1).reshape((1, -1))\n            secondary = np.matmul(present, occur)\n            for j in range(logits.shape[1]):\n                if secondary[:, j] > 0:\n                    logits[:, j] += occur_thr\n            logits_max = logits.max(0)\n        for j in range(logits.shape[1]):\n            if logits_max[j] > thr:\n                logits[:, j] += thr - incr\n        for seconds, logit in zip(range(5, 605, 5), logits):\n            row_id = str(audio_id )+ '_' + site + '_' + str(seconds)\n            birds = list(INV_BIRD_CODE[logit >= thr])\n            if 'nocall' in birds and len(birds) > 1:\n                birds = [p for p in birds if p!= 'nocall']\n            elif len(birds) == 0:\n                birds = ['nocall']\n            birds = ' '.join(birds)\n\n            all_preds.append((row_id, site, audio_id, seconds, birds))\n    df = pd.DataFrame().from_records(all_preds)\n    df.columns = ['row_id', 'site', 'audio_id', 'seconds', 'birds']\n    df = df.sort_values(['site', 'audio_id', 'seconds']).reset_index(drop=True)\n    return df\n\n```",
      "votes": null
    },
    {
      "id": "1332996",
      "postDate": "06/02/2021 12:21:41",
      "content": "<p>we tried both, short and long clips both with/ without PP and used what gave the best cv (which was 8sec for a binary bird/ no-bird classifier and 30sec for our \"normal\" model. I also tried different variants in switching length by reshapeing within the model. Will probably talk about that a bit more in our summary</p>",
      "rawMarkdown": "we tried both, short and long clips both with/ without PP and used what gave the best cv (which was 8sec for a binary bird/ no-bird classifier and 30sec for our \"normal\" model. I also tried different variants in switching length by reshapeing within the model. Will probably talk about that a bit more in our summary",
      "votes": null
    },
    {
      "id": "1333085",
      "postDate": "06/02/2021 13:31:22",
      "content": "<p>I am now convinced that a long range model is needed. ;)</p>",
      "rawMarkdown": "I am now convinced that a long range model is needed. ;)",
      "votes": null
    },
    {
      "id": "1336173",
      "postDate": "06/04/2021 17:23:26",
      "content": "<p>Thank you for posting this, my post-processing was quite similar too. Keeping a higher threshold to get the most confident birds in a soundscape and increase their confidence with a large factor(5), and then again apply a moderate threshold. It gave us around (0.05) boost on both validation/Lb. </p>\n<p>How much boost did you get by this? </p>",
      "rawMarkdown": "Thank you for posting this, my post-processing was quite similar too. Keeping a higher threshold to get the most confident birds in a soundscape and increase their confidence with a large factor(5), and then again apply a moderate threshold. It gave us around (0.05) boost on both validation/Lb. \n\nHow much boost did you get by this?",
      "votes": null
    },
    {
      "id": "1336215",
      "postDate": "06/04/2021 18:10:33",
      "content": "<p>I got a boost similar to yours.</p>",
      "rawMarkdown": "I got a boost similar to yours.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1332996,
      "author_name": "christofhenkel",
      "author_url": "",
      "post_date": "06/02/2021 12:21:41",
      "content": "<p>we tried both, short and long clips both with/ without PP and used what gave the best cv (which was 8sec for a binary bird/ no-bird classifier and 30sec for our \"normal\" model. I also tried different variants in switching length by reshapeing within the model. Will probably talk about that a bit more in our summary</p>",
      "votes": null,
      "replies": [
        {
          "id": 1333085,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/02/2021 13:31:22",
          "content": "<p>I am now convinced that a long range model is needed. ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1336173,
      "author_name": "nischaydnk",
      "author_url": "",
      "post_date": "06/04/2021 17:23:26",
      "content": "<p>Thank you for posting this, my post-processing was quite similar too. Keeping a higher threshold to get the most confident birds in a soundscape and increase their confidence with a large factor(5), and then again apply a moderate threshold. It gave us around (0.05) boost on both validation/Lb. </p>\n<p>How much boost did you get by this? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1336215,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/04/2021 18:10:33",
          "content": "<p>I got a boost similar to yours.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1332773": "After reading some top solutions writeup I see that there are two classes.  @tezdhar me, and probably others trained on 5 seconds clips while others used longer clips with SED model variants.  The former misses both long term interactions and species interactions.  It must be completed by some post processing to capture them.  Here is the one I used to post process soundscape predictions.  It is based on two ideas:\n\n1. If a bird is heard in one of the 5 second clips of a soundscape then it is more likely to be hear in others.  This was used by top solutions in Cornell competition.  We compute the maximum prediction for each species across all 5 second clips to determine presence of bird species in the soundscape.  If a bird species is detected then the logits for this species are increased for all 5 seconds clips. \n\n2. If two bird species are known to co exist then hearing one makes the other one more likely.  For this I computed a co occurrence matrix using primary and secondary labels.  If a species is detected then logits for all co occurring species are increased.\n\nThis logic is controlled by three thresholds tuned on train soundscape:\n\n- `thr`.   Logits greater than it correspond to a presence\n- 'incr'.  Logits above it for a present species are rounded to 1 in prediction.\n- ocur_thr.  Amount by which logits of co occurring species are augmented.\n\nExample values are:\n\n         'thr':2.1, \n\n          'incr':-0.2,\n\n          'occur_thr':0.9,\n\nThe code of my prediction function is this.  It takes as input the logits for all 5 second clips for a soundscape.  This postprocessing worked extremely well (maybe too well).\n\n```\n\noccur = np.zeros((len(BIRD_CODE), len(BIRD_CODE)), dtype='int')\n\ndf = train[train.primary_label != 'nocall']\n\nfor primary, secondary in zip(df.primary_label, df.secondary_labels):\n    for label in secondary:\n        occur[BIRD_CODE[primary], BIRD_CODE[label]] += 1\n        occur[BIRD_CODE[label], BIRD_CODE[primary]] += 1\n    #occur[BIRD_CODE[primary], BIRD_CODE[primary]] = 1\n        \noccur = occur.clip(0,1)\n\ndef compute_pred(LOGITS, thr, incr, occur_thr):\n    all_preds = []\n    for file_path, logits in zip(test_audio, LOGITS):\n        fileinfo = file_path.split(os.sep)[-1].rsplit('.', 1)[0].split('_')\n        audio_id = int(fileinfo[0])\n        site = fileinfo[1]\n        logits_max = logits.max(0)\n        logits = logits.copy()\n        if occur_thr is not None and occur_thr != 0:\n            present = ((logits_max > thr) * 1).reshape((1, -1))\n            secondary = np.matmul(present, occur)\n            for j in range(logits.shape[1]):\n                if secondary[:, j] > 0:\n                    logits[:, j] += occur_thr\n            logits_max = logits.max(0)\n        for j in range(logits.shape[1]):\n            if logits_max[j] > thr:\n                logits[:, j] += thr - incr\n        for seconds, logit in zip(range(5, 605, 5), logits):\n            row_id = str(audio_id )+ '_' + site + '_' + str(seconds)\n            birds = list(INV_BIRD_CODE[logit >= thr])\n            if 'nocall' in birds and len(birds) > 1:\n                birds = [p for p in birds if p!= 'nocall']\n            elif len(birds) == 0:\n                birds = ['nocall']\n            birds = ' '.join(birds)\n\n            all_preds.append((row_id, site, audio_id, seconds, birds))\n    df = pd.DataFrame().from_records(all_preds)\n    df.columns = ['row_id', 'site', 'audio_id', 'seconds', 'birds']\n    df = df.sort_values(['site', 'audio_id', 'seconds']).reset_index(drop=True)\n    return df\n\n```",
    "1332996": "we tried both, short and long clips both with/ without PP and used what gave the best cv (which was 8sec for a binary bird/ no-bird classifier and 30sec for our \"normal\" model. I also tried different variants in switching length by reshapeing within the model. Will probably talk about that a bit more in our summary",
    "1333085": "I am now convinced that a long range model is needed. ;)",
    "1336173": "Thank you for posting this, my post-processing was quite similar too. Keeping a higher threshold to get the most confident birds in a soundscape and increase their confidence with a large factor(5), and then again apply a moderate threshold. It gave us around (0.05) boost on both validation/Lb. \n\nHow much boost did you get by this?",
    "1336215": "I got a boost similar to yours."
  },
  "source": "meta"
}