{
  "id": 179592,
  "title": "Accelerating librosa read",
  "url": "/competitions/birdsong-recognition/discussion/179592",
  "author_name": "CPMP",
  "post_date": "2020-09-02T14:20:09.145000",
  "votes": 45,
  "comment_count": 33,
  "views": 0,
  "content": "<p>Librosa issues a warning when reading mp3 files: <br>\n<code>warnings.warn(\"PySoundFile failed. Trying audioread instead.\")</code></p>\n<p>We can skip the failed try by calling directly audioread. Here is a sample code, inspired by librosa.load source code at <a href=\"https://github.com/librosa/librosa/blob/0.8.0/librosa/core/audio.py\" target=\"_blank\">https://github.com/librosa/librosa/blob/0.8.0/librosa/core/audio.py</a></p>\n<pre><code>def get_clip_sr(audio_id, data_path):\n    path = data_path / (audio_id + '.mp3')\n    clip, sr_native = librosa.core.audio.__audioread_load(path, offset=0.0, duration=None, dtype=np.float32)\n    clip = librosa.to_mono(clip)\n    sr = 32000\n    if sr_native &gt; 0:\n        clip = librosa.resample(clip, sr_native, sr, res_type='kaiser_fast')\n    return clip, sr\n</code></pre>",
  "messages": [
    {
      "id": 995527,
      "postDate": "2020-09-02T14:20:09.147Z",
      "content": "<p>Librosa issues a warning when reading mp3 files: <br>\n<code>warnings.warn(\"PySoundFile failed. Trying audioread instead.\")</code></p>\n<p>We can skip the failed try by calling directly audioread. Here is a sample code, inspired by librosa.load source code at <a href=\"https://github.com/librosa/librosa/blob/0.8.0/librosa/core/audio.py\" target=\"_blank\">https://github.com/librosa/librosa/blob/0.8.0/librosa/core/audio.py</a></p>\n<pre><code>def get_clip_sr(audio_id, data_path):\n    path = data_path / (audio_id + '.mp3')\n    clip, sr_native = librosa.core.audio.__audioread_load(path, offset=0.0, duration=None, dtype=np.float32)\n    clip = librosa.to_mono(clip)\n    sr = 32000\n    if sr_native &gt; 0:\n        clip = librosa.resample(clip, sr_native, sr, res_type='kaiser_fast')\n    return clip, sr\n</code></pre>",
      "rawMarkdown": "Librosa issues a warning when reading mp3 files: \n`warnings.warn(\"PySoundFile failed. Trying audioread instead.\")`\n\nWe can skip the failed try by calling directly audioread. Here is a sample code, inspired by librosa.load source code at https://github.com/librosa/librosa/blob/0.8.0/librosa/core/audio.py\n\n```\ndef get_clip_sr(audio_id, data_path):\n    path = data_path / (audio_id + '.mp3')\n    clip, sr_native = librosa.core.audio.__audioread_load(path, offset=0.0, duration=None, dtype=np.float32)\n    clip = librosa.to_mono(clip)\n    sr = 32000\n    if sr_native > 0:\n        clip = librosa.resample(clip, sr_native, sr, res_type='kaiser_fast')\n    return clip, sr\n```\n\n",
      "votes": 44
    },
    {
      "id": 995840,
      "postDate": "2020-09-02T20:43:05.907Z",
      "content": "<p>Just to confirm that using the above I could get valid submissions.</p>",
      "rawMarkdown": "Just to confirm that using the above I could get valid submissions.",
      "votes": 3
    },
    {
      "id": 1005403,
      "postDate": "2020-09-10T13:25:41.223Z",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> thanks, this is much faster than the librosa.load<br>\nCan you point me please where the host said the test data is sampled at 32kHz? 32kHz is quite an unusual sampling frequency. </p>",
      "rawMarkdown": "@cpmpml thanks, this is much faster than the librosa.load\nCan you point me please where the host said the test data is sampled at 32kHz? 32kHz is quite an unusual sampling frequency. ",
      "votes": 1,
      "replies": [
        {
          "id": 1005409,
          "postDate": "2020-09-10T13:28:07.943Z",
          "content": "<p>You can find it here:<br>\n<a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/179253\" target=\"_blank\">https://www.kaggle.com/c/birdsong-recognition/discussion/179253</a></p>",
          "rawMarkdown": "You can find it here:\nhttps://www.kaggle.com/c/birdsong-recognition/discussion/179253",
          "votes": 2
        },
        {
          "id": 1005417,
          "postDate": "2020-09-10T13:34:14.093Z",
          "content": "<p>Thanks for the quick reply!</p>",
          "rawMarkdown": "Thanks for the quick reply!"
        }
      ]
    },
    {
      "id": 1003786,
      "postDate": "2020-09-09T08:45:05.123Z",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a></p>\n<p>Thanks for the tip. This is definitely faster :)</p>\n<blockquote>\n  <p>if sr_native &gt; 0:</p>\n</blockquote>\n<p>Why do you check for the above condition ?<br>\nAlso, please confirm if this works for you in the submission process.</p>",
      "rawMarkdown": "@cpmpml\n\nThanks for the tip. This is definitely faster :)\n>if sr_native > 0:\n\nWhy do you check for the above condition ?\nAlso, please confirm if this works for you in the submission process.\n\n",
      "votes": 1,
      "replies": [
        {
          "id": 1003857,
          "postDate": "2020-09-09T09:59:37.230Z",
          "content": "<ol>\n<li><p>It seems that some test files have a null sampling rate in the file, see another post where I ask why people resample test files.</p></li>\n<li><p>Yes, I use this code in my submissions.</p></li>\n</ol>",
          "rawMarkdown": "1. It seems that some test files have a null sampling rate in the file, see another post where I ask why people resample test files.\n\n2. Yes, I use this code in my submissions.",
          "votes": 1
        },
        {
          "id": 1003868,
          "postDate": "2020-09-09T10:13:05.007Z",
          "content": "<p>Thank you. <br>\nSo how is resampling done when sr_native is None ? Since my understanding is that resampling to 32000 is recommended.</p>",
          "rawMarkdown": "Thank you. \nSo how is resampling done when sr_native is None ? Since my understanding is that resampling to 32000 is recommended."
        },
        {
          "id": 1004565,
          "postDate": "2020-09-09T19:45:11.747Z",
          "content": "<p>host says data is already sampled at 32 kHz. </p>\n<p>I am not convinced we need to resample to be honest.  I asked the hosts but they don't answer.</p>",
          "rawMarkdown": "host says data is already sampled at 32 kHz. \n\nI am not convinced we need to resample to be honest.  I asked the hosts but they don't answer.",
          "votes": 1
        }
      ]
    },
    {
      "id": 997266,
      "postDate": "2020-09-03T21:49:34.107Z",
      "content": "<p><a href=\"https://www.kaggle.com/CPMP\" target=\"_blank\">@CPMP</a>: thanks a lot for sharing your workaround! Really helpful.</p>\n<p>Just for the record, I got this type of issues when trying to do librosa-facilitated parallelized feature extraction from audio files, using Ray.</p>\n<p>When I did the same processing with different parallelized computation frameworks (either using Dask or classical multiprocessing), I have not got such warnings at all.</p>\n<p>That's what I observed both on Kaggle kernels and my local Windows 10-based notebook (with Python 3.7 and Anaconda as a runtime environment).</p>",
      "rawMarkdown": "@CPMP: thanks a lot for sharing your workaround! Really helpful.\n\nJust for the record, I got this type of issues when trying to do librosa-facilitated parallelized feature extraction from audio files, using Ray.\n\nWhen I did the same processing with different parallelized computation frameworks (either using Dask or classical multiprocessing), I have not got such warnings at all.\n\nThat's what I observed both on Kaggle kernels and my local Windows 10-based notebook (with Python 3.7 and Anaconda as a runtime environment).",
      "votes": 1,
      "replies": [
        {
          "id": 997269,
          "postDate": "2020-09-03T21:57:03.407Z",
          "content": "<p>Thanks.  I think warnings are printed on each worked stdout.  There is no reason you see them in the master process.  Master only see the error messages that are printed on stderr.</p>",
          "rawMarkdown": "Thanks.  I think warnings are printed on each worked stdout.  There is no reason you see them in the master process.  Master only see the error messages that are printed on stderr.",
          "votes": 1
        },
        {
          "id": 997275,
          "postDate": "2020-09-03T22:14:16.437Z",
          "content": "<p>Thanks, <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> </p>\n<p>I think the runtime environment/platform may also contribute to such effects. When I tried doing librosa-based feature extraction just in a single-process fashion, using my local Windows 10-based notebook (with Python 3.7 and Anaconda as a runtime environment), it did not display any warnings of this sort.</p>\n<p>Anyways, it is great to see your deep technical insights - they have been helping the community on a number of competions, both now and years back.</p>",
          "rawMarkdown": "Thanks, @cpmpml \n\nI think the runtime environment/platform may also contribute to such effects. When I tried doing librosa-based feature extraction just in a single-process fashion, using my local Windows 10-based notebook (with Python 3.7 and Anaconda as a runtime environment), it did not display any warnings of this sort.\n\nAnyways, it is great to see your deep technical insights - they have been helping the community on a number of competions, both now and years back.",
          "votes": 1
        },
        {
          "id": 997605,
          "postDate": "2020-09-04T06:07:41.827Z",
          "content": "<p>The warning disappears with more recent version as they rewrote the load code.  In Kaggle we use 0.8.0.  Maybe you used a more recent version of librosa on Windows?</p>",
          "rawMarkdown": "The warning disappears with more recent version as they rewrote the load code.  In Kaggle we use 0.8.0.  Maybe you used a more recent version of librosa on Windows?",
          "votes": 1
        },
        {
          "id": 997774,
          "postDate": "2020-09-04T08:24:19.787Z",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> : yeah, the librosa version discrepancy could be a determinant indeed.</p>\n<p>In my local Windows 10 environment , I use librosa 0.7.2 as well as the downgraded version of numba (see <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/167792\" target=\"_blank\">https://www.kaggle.com/c/birdsong-recognition/discussion/167792</a> for more context on it).</p>\n<p>thank you.</p>\n<p>P. S. Just to clarify on my earlier comments, I ran the Ray-based process on Kaggle kernels only as Ray is not fully tuned to run under Windows 10 or equally Linux subsystem for Windows 10 (see the latest comments in <a href=\"https://github.com/ray-project/ray/issues/631)\" target=\"_blank\">https://github.com/ray-project/ray/issues/631)</a>. My processing with Dask and multiprocessing was done in both environments, in turn.</p>",
          "rawMarkdown": "@cpmpml : yeah, the librosa version discrepancy could be a determinant indeed.\n\nIn my local Windows 10 environment , I use librosa 0.7.2 as well as the downgraded version of numba (see https://www.kaggle.com/c/birdsong-recognition/discussion/167792 for more context on it).\n\nthank you.\n\nP. S. Just to clarify on my earlier comments, I ran the Ray-based process on Kaggle kernels only as Ray is not fully tuned to run under Windows 10 or equally Linux subsystem for Windows 10 (see the latest comments in https://github.com/ray-project/ray/issues/631). My processing with Dask and multiprocessing was done in both environments, in turn.",
          "votes": 1
        }
      ]
    },
    {
      "id": 995933,
      "postDate": "2020-09-03T00:55:37.087Z",
      "content": "<p>I'm doing this too!</p>",
      "rawMarkdown": "I'm doing this too!",
      "votes": 1
    },
    {
      "id": 995751,
      "postDate": "2020-09-02T18:33:59.983Z",
      "content": "<p>Just a side question: Does anyone use AudioSegment?</p>",
      "rawMarkdown": "Just a side question: Does anyone use AudioSegment?",
      "votes": 1,
      "replies": [
        {
          "id": 995761,
          "postDate": "2020-09-02T18:42:45.077Z",
          "content": "<p>You should create a topic for your question given it is unrelated to my topic.</p>",
          "rawMarkdown": "You should create a topic for your question given it is unrelated to my topic.",
          "votes": 1
        },
        {
          "id": 996401,
          "postDate": "2020-09-03T09:07:54.813Z",
          "content": "<p>What is wrong with the advice above.  Asking a question unrelated to a topic in a comment of the topic has very little chances to get an answer.  </p>",
          "rawMarkdown": "What is wrong with the advice above.  Asking a question unrelated to a topic in a comment of the topic has very little chances to get an answer.  ",
          "votes": 1
        },
        {
          "id": 997283,
          "postDate": "2020-09-03T22:38:12.537Z",
          "content": "<p>Imagine the day when one has decrypted the understanding to certain downvotes, probably remains a dream.  Keep up the good work!</p>",
          "rawMarkdown": "Imagine the day when one has decrypted the understanding to certain downvotes, probably remains a dream.  Keep up the good work!",
          "votes": 1
        },
        {
          "id": 997540,
          "postDate": "2020-09-04T05:27:54.697Z",
          "content": "<p>I even get a downvote for thanking you!  Someone is really sick.</p>",
          "rawMarkdown": "I even get a downvote for thanking you!  Someone is really sick."
        },
        {
          "id": 997645,
          "postDate": "2020-09-04T06:35:46.983Z",
          "content": "<p>Downvote on a positive segment, then we can rule out that's an algorithm, even a bad trained model would handle it better.</p>",
          "rawMarkdown": "Downvote on a positive segment, then we can rule out that's an algorithm, even a bad trained model would handle it better."
        },
        {
          "id": 997650,
          "postDate": "2020-09-04T06:38:51.073Z",
          "content": "<p>The algorithm is simple: if cpmp posts, then downvote.  </p>\n<p>But it must be a human as this only happens during a specific time zone.</p>",
          "rawMarkdown": "The algorithm is simple: if cpmp posts, then downvote.  \n\nBut it must be a human as this only happens during a specific time zone.\n"
        },
        {
          "id": 997686,
          "postDate": "2020-09-04T07:11:53.590Z",
          "content": "<p>Try to report it, misuse of downvote, actively repeatedly downvotes on someone's posts. </p>",
          "rawMarkdown": "Try to report it, misuse of downvote, actively repeatedly downvotes on someone's posts. ",
          "votes": 1
        },
        {
          "id": 998232,
          "postDate": "2020-09-04T15:33:09.060Z",
          "content": "<p>I reported it.  I reported it sometimes when it was systematic, but I never saw an effect.  Thanks for trying to help though.  Most people are annoyed if I mention downvotes…</p>",
          "rawMarkdown": "I reported it.  I reported it sometimes when it was systematic, but I never saw an effect.  Thanks for trying to help though.  Most people are annoyed if I mention downvotes...",
          "votes": 1
        },
        {
          "id": 999018,
          "postDate": "2020-09-05T09:14:23.637Z",
          "content": "<p>No problem. \"How\" is as important as \"What\". </p>",
          "rawMarkdown": "No problem. \"How\" is as important as \"What\". "
        },
        {
          "id": 1006939,
          "postDate": "2020-09-11T16:55:57.017Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 995608,
      "postDate": "2020-09-02T15:52:58.980Z",
      "content": "<p>This is what i really want!</p>",
      "rawMarkdown": "This is what i really want!",
      "votes": 1
    },
    {
      "id": 995579,
      "postDate": "2020-09-02T15:30:23.777Z",
      "content": "<p>You are probably where you usually are soon, at the top ☝️</p>",
      "rawMarkdown": "You are probably where you usually are soon, at the top ☝️",
      "votes": 2,
      "replies": [
        {
          "id": 995582,
          "postDate": "2020-09-02T15:33:09.593Z",
          "content": "<p>Thank you for the nice words.</p>",
          "rawMarkdown": "Thank you for the nice words.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1005864,
      "postDate": "2020-09-10T20:27:47.137Z",
      "content": "<p>This worked better than the standard resampling:</p>\n<p>librosa.load(path, offset=start_time, <br>\n    duration=duration, <br>\n    sr=nSAMPLERATE,<br>\n    mono=True, <br>\n    res_type='kaiser_fast') </p>\n<p>Kaiser_fast resampling is faster than the normal version.</p>\n<p>I used 6 threads and kept the whole file in memory and parsed out the 5 seconds intervals separately to speed up.</p>",
      "rawMarkdown": "This worked better than the standard resampling:\n\n librosa.load(path, offset=start_time, \n    duration=duration, \n    sr=nSAMPLERATE,\n    mono=True, \n    res_type='kaiser_fast') \n\n Kaiser_fast resampling is faster than the normal version.\n\nI used 6 threads and kept the whole file in memory and parsed out the 5 seconds intervals separately to speed up.\n\n\n\n",
      "replies": [
        {
          "id": 1005870,
          "postDate": "2020-09-10T20:33:10.820Z",
          "content": "<p>??</p>\n<p>Your code is the slow code I started from.</p>\n<p>Also, your code formatting makes it almost impossible to read.</p>",
          "rawMarkdown": "??\n\nYour code is the slow code I started from.\n\nAlso, your code formatting makes it almost impossible to read.",
          "votes": -1
        },
        {
          "id": 1005880,
          "postDate": "2020-09-10T20:46:35.933Z",
          "content": "<p>Ok, I rushed putting it out, it greatly sped up things for me, but I was using a lower resampling already.</p>",
          "rawMarkdown": "Ok, I rushed putting it out, it greatly sped up things for me, but I was using a lower resampling already."
        },
        {
          "id": 1005884,
          "postDate": "2020-09-10T20:57:12.777Z",
          "content": "<p>Sorry, misunderstood you last comment. Thanks for confirming it speeded up thing.</p>",
          "rawMarkdown": "Sorry, misunderstood you last comment. Thanks for confirming it speeded up thing.",
          "votes": -1
        }
      ]
    },
    {
      "id": 999548,
      "postDate": "2020-09-05T18:30:51.317Z",
      "content": "<p>very good sir</p>",
      "rawMarkdown": "very good sir"
    }
  ],
  "comments": [
    {
      "id": 995840,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2020-09-02T20:43:05.907000",
      "content": "<p>Just to confirm that using the above I could get valid submissions.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1005403,
      "author_name": "Krisztián Fekete",
      "author_url": "",
      "post_date": "2020-09-10T13:25:41.223000",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> thanks, this is much faster than the librosa.load<br>\nCan you point me please where the host said the test data is sampled at 32kHz? 32kHz is quite an unusual sampling frequency. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1005409,
          "author_name": "Ali Abdin",
          "author_url": "",
          "post_date": "2020-09-10T13:28:07.943000",
          "content": "<p>You can find it here:<br>\n<a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/179253\" target=\"_blank\">https://www.kaggle.com/c/birdsong-recognition/discussion/179253</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1005417,
          "author_name": "Krisztián Fekete",
          "author_url": "",
          "post_date": "2020-09-10T13:34:14.093000",
          "content": "<p>Thanks for the quick reply!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1003786,
      "author_name": "Vee",
      "author_url": "",
      "post_date": "2020-09-09T08:45:05.123000",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a></p>\n<p>Thanks for the tip. This is definitely faster :)</p>\n<blockquote>\n  <p>if sr_native &gt; 0:</p>\n</blockquote>\n<p>Why do you check for the above condition ?<br>\nAlso, please confirm if this works for you in the submission process.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1003857,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-09-09T09:59:37.230000",
          "content": "<ol>\n<li><p>It seems that some test files have a null sampling rate in the file, see another post where I ask why people resample test files.</p></li>\n<li><p>Yes, I use this code in my submissions.</p></li>\n</ol>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1003868,
          "author_name": "Vee",
          "author_url": "",
          "post_date": "2020-09-09T10:13:05.007000",
          "content": "<p>Thank you. <br>\nSo how is resampling done when sr_native is None ? Since my understanding is that resampling to 32000 is recommended.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1004565,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-09-09T19:45:11.747000",
          "content": "<p>host says data is already sampled at 32 kHz. </p>\n<p>I am not convinced we need to resample to be honest.  I asked the hosts but they don't answer.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 997266,
      "author_name": "Georgii Vyshnia",
      "author_url": "",
      "post_date": "2020-09-03T21:49:34.107000",
      "content": "<p><a href=\"https://www.kaggle.com/CPMP\" target=\"_blank\">@CPMP</a>: thanks a lot for sharing your workaround! Really helpful.</p>\n<p>Just for the record, I got this type of issues when trying to do librosa-facilitated parallelized feature extraction from audio files, using Ray.</p>\n<p>When I did the same processing with different parallelized computation frameworks (either using Dask or classical multiprocessing), I have not got such warnings at all.</p>\n<p>That's what I observed both on Kaggle kernels and my local Windows 10-based notebook (with Python 3.7 and Anaconda as a runtime environment).</p>",
      "votes": 1,
      "replies": [
        {
          "id": 997269,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-09-03T21:57:03.407000",
          "content": "<p>Thanks.  I think warnings are printed on each worked stdout.  There is no reason you see them in the master process.  Master only see the error messages that are printed on stderr.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 997275,
          "author_name": "Georgii Vyshnia",
          "author_url": "",
          "post_date": "2020-09-03T22:14:16.437000",
          "content": "<p>Thanks, <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> </p>\n<p>I think the runtime environment/platform may also contribute to such effects. When I tried doing librosa-based feature extraction just in a single-process fashion, using my local Windows 10-based notebook (with Python 3.7 and Anaconda as a runtime environment), it did not display any warnings of this sort.</p>\n<p>Anyways, it is great to see your deep technical insights - they have been helping the community on a number of competions, both now and years back.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 997605,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-09-04T06:07:41.827000",
          "content": "<p>The warning disappears with more recent version as they rewrote the load code.  In Kaggle we use 0.8.0.  Maybe you used a more recent version of librosa on Windows?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 997774,
          "author_name": "Georgii Vyshnia",
          "author_url": "",
          "post_date": "2020-09-04T08:24:19.787000",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> : yeah, the librosa version discrepancy could be a determinant indeed.</p>\n<p>In my local Windows 10 environment , I use librosa 0.7.2 as well as the downgraded version of numba (see <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/167792\" target=\"_blank\">https://www.kaggle.com/c/birdsong-recognition/discussion/167792</a> for more context on it).</p>\n<p>thank you.</p>\n<p>P. S. Just to clarify on my earlier comments, I ran the Ray-based process on Kaggle kernels only as Ray is not fully tuned to run under Windows 10 or equally Linux subsystem for Windows 10 (see the latest comments in <a href=\"https://github.com/ray-project/ray/issues/631)\" target=\"_blank\">https://github.com/ray-project/ray/issues/631)</a>. My processing with Dask and multiprocessing was done in both environments, in turn.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 995933,
      "author_name": "Quan",
      "author_url": "",
      "post_date": "2020-09-03T00:55:37.087000",
      "content": "<p>I'm doing this too!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 995751,
      "author_name": "Ali Abdin",
      "author_url": "",
      "post_date": "2020-09-02T18:33:59.983000",
      "content": "<p>Just a side question: Does anyone use AudioSegment?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 995761,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-09-02T18:42:45.077000",
          "content": "<p>You should create a topic for your question given it is unrelated to my topic.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 996401,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-09-03T09:07:54.813000",
          "content": "<p>What is wrong with the advice above.  Asking a question unrelated to a topic in a comment of the topic has very little chances to get an answer.  </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 997283,
          "author_name": "Kirderf",
          "author_url": "",
          "post_date": "2020-09-03T22:38:12.537000",
          "content": "<p>Imagine the day when one has decrypted the understanding to certain downvotes, probably remains a dream.  Keep up the good work!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 997540,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-09-04T05:27:54.697000",
          "content": "<p>I even get a downvote for thanking you!  Someone is really sick.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 997645,
          "author_name": "Kirderf",
          "author_url": "",
          "post_date": "2020-09-04T06:35:46.983000",
          "content": "<p>Downvote on a positive segment, then we can rule out that's an algorithm, even a bad trained model would handle it better.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 997650,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-09-04T06:38:51.073000",
          "content": "<p>The algorithm is simple: if cpmp posts, then downvote.  </p>\n<p>But it must be a human as this only happens during a specific time zone.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 997686,
          "author_name": "Kirderf",
          "author_url": "",
          "post_date": "2020-09-04T07:11:53.590000",
          "content": "<p>Try to report it, misuse of downvote, actively repeatedly downvotes on someone's posts. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 998232,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-09-04T15:33:09.060000",
          "content": "<p>I reported it.  I reported it sometimes when it was systematic, but I never saw an effect.  Thanks for trying to help though.  Most people are annoyed if I mention downvotes…</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 999018,
          "author_name": "Kirderf",
          "author_url": "",
          "post_date": "2020-09-05T09:14:23.637000",
          "content": "<p>No problem. \"How\" is as important as \"What\". </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1006939,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-09-11T16:55:57.017000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 995608,
      "author_name": "kaggler",
      "author_url": "",
      "post_date": "2020-09-02T15:52:58.980000",
      "content": "<p>This is what i really want!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 995579,
      "author_name": "Kirderf",
      "author_url": "",
      "post_date": "2020-09-02T15:30:23.777000",
      "content": "<p>You are probably where you usually are soon, at the top ☝️</p>",
      "votes": 2,
      "replies": [
        {
          "id": 995582,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-09-02T15:33:09.593000",
          "content": "<p>Thank you for the nice words.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1005864,
      "author_name": "Mark Eckdahl",
      "author_url": "",
      "post_date": "2020-09-10T20:27:47.137000",
      "content": "<p>This worked better than the standard resampling:</p>\n<p>librosa.load(path, offset=start_time, <br>\n    duration=duration, <br>\n    sr=nSAMPLERATE,<br>\n    mono=True, <br>\n    res_type='kaiser_fast') </p>\n<p>Kaiser_fast resampling is faster than the normal version.</p>\n<p>I used 6 threads and kept the whole file in memory and parsed out the 5 seconds intervals separately to speed up.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1005870,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-09-10T20:33:10.820000",
          "content": "<p>??</p>\n<p>Your code is the slow code I started from.</p>\n<p>Also, your code formatting makes it almost impossible to read.</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 1005880,
          "author_name": "Mark Eckdahl",
          "author_url": "",
          "post_date": "2020-09-10T20:46:35.933000",
          "content": "<p>Ok, I rushed putting it out, it greatly sped up things for me, but I was using a lower resampling already.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1005884,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-09-10T20:57:12.777000",
          "content": "<p>Sorry, misunderstood you last comment. Thanks for confirming it speeded up thing.</p>",
          "votes": -1,
          "replies": []
        }
      ]
    },
    {
      "id": 999548,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-05T18:30:51.317000",
      "content": "<p>very good sir</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "995527": "Librosa issues a warning when reading mp3 files: \n`warnings.warn(\"PySoundFile failed. Trying audioread instead.\")`\n\nWe can skip the failed try by calling directly audioread. Here is a sample code, inspired by librosa.load source code at https://github.com/librosa/librosa/blob/0.8.0/librosa/core/audio.py\n\n```\ndef get_clip_sr(audio_id, data_path):\n    path = data_path / (audio_id + '.mp3')\n    clip, sr_native = librosa.core.audio.__audioread_load(path, offset=0.0, duration=None, dtype=np.float32)\n    clip = librosa.to_mono(clip)\n    sr = 32000\n    if sr_native > 0:\n        clip = librosa.resample(clip, sr_native, sr, res_type='kaiser_fast')\n    return clip, sr\n```\n\n",
    "995840": "Just to confirm that using the above I could get valid submissions.",
    "1005403": "@cpmpml thanks, this is much faster than the librosa.load\nCan you point me please where the host said the test data is sampled at 32kHz? 32kHz is quite an unusual sampling frequency. ",
    "1003786": "@cpmpml\n\nThanks for the tip. This is definitely faster :)\n>if sr_native > 0:\n\nWhy do you check for the above condition ?\nAlso, please confirm if this works for you in the submission process.\n\n",
    "997266": "@CPMP: thanks a lot for sharing your workaround! Really helpful.\n\nJust for the record, I got this type of issues when trying to do librosa-facilitated parallelized feature extraction from audio files, using Ray.\n\nWhen I did the same processing with different parallelized computation frameworks (either using Dask or classical multiprocessing), I have not got such warnings at all.\n\nThat's what I observed both on Kaggle kernels and my local Windows 10-based notebook (with Python 3.7 and Anaconda as a runtime environment).",
    "995933": "I'm doing this too!",
    "995751": "Just a side question: Does anyone use AudioSegment?",
    "995608": "This is what i really want!",
    "995579": "You are probably where you usually are soon, at the top ☝️",
    "1005864": "This worked better than the standard resampling:\n\n librosa.load(path, offset=start_time, \n    duration=duration, \n    sr=nSAMPLERATE,\n    mono=True, \n    res_type='kaiser_fast') \n\n Kaiser_fast resampling is faster than the normal version.\n\nI used 6 threads and kept the whole file in memory and parsed out the 5 seconds intervals separately to speed up.\n\n\n\n",
    "999548": "very good sir"
  }
}