{
  "id": 398827,
  "title": "Packages won't work on R notebook",
  "url": "/competitions/birdclef-2023/discussion/398827",
  "author_name": "",
  "post_date": "2023-04-01T04:02:30.264543700Z",
  "votes": null,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Bummer, seems like another competition where I've built out a full R model and one of the CRAN package won't work in the notebook.  The package I need is called 'av' it seems to be the only package to convert an ogg to wav.  Any tips?  Is it a known thing that Kaggle notebook competitions are bad for R?</p>",
  "messages": [
    {
      "id": "2204869",
      "postDate": "04/01/2023 04:02:30",
      "content": "<p>Bummer, seems like another competition where I've built out a full R model and one of the CRAN package won't work in the notebook.  The package I need is called 'av' it seems to be the only package to convert an ogg to wav.  Any tips?  Is it a known thing that Kaggle notebook competitions are bad for R?</p>",
      "rawMarkdown": "Bummer, seems like another competition where I've built out a full R model and one of the CRAN package won't work in the notebook.  The package I need is called 'av' it seems to be the only package to convert an ogg to wav.  Any tips?  Is it a known thing that Kaggle notebook competitions are bad for R?",
      "votes": null
    },
    {
      "id": "2204881",
      "postDate": "04/01/2023 04:13:30",
      "content": "<p>torchaudio is part of my pipeline and doesn't work either :*(</p>",
      "rawMarkdown": "torchaudio is part of my pipeline and doesn't work either :*(",
      "votes": null
    },
    {
      "id": "2205068",
      "postDate": "04/01/2023 07:59:55",
      "content": "<p>Hey Andy,</p>\n<p>would you mind sharing the exact Error message? There might be a solution to your problem if you elaborate a bit on what is happening!</p>\n<p>Best,<br>\nJan</p>",
      "rawMarkdown": "Hey Andy,\n\nwould you mind sharing the exact Error message? There might be a solution to your problem if you elaborate a bit on what is happening!\n\nBest,\nJan",
      "votes": null
    },
    {
      "id": "2205069",
      "postDate": "04/01/2023 08:03:42",
      "content": "<p>Hey Andy, <br>\nwhat exactly are you trying to do? </p>\n<p>Best,<br>\nJan</p>",
      "rawMarkdown": "Hey Andy, \nwhat exactly are you trying to do? \n\nBest,\nJan",
      "votes": null
    },
    {
      "id": "2205403",
      "postDate": "04/01/2023 14:02:37",
      "content": "<p>Hi Jan, thanks for your reply!  Here's what I've tried.</p>\n<p>library(av)<br>\n<em>Error in library(av): there is no package called ‘av’</em></p>\n<p>install.packages(\"av\",lib = \"/kaggle/working\")<br>\n<em>Warning message in install.packages(\"av\", lib = \"/kaggle/working\"):\n“installation of package ‘av’ had non-zero exit status”</em></p>\n<p>library(av)<br>\n<em>Error in library(av): there is no package called ‘av’</em></p>\n<p>library(torchaudio)<br>\n<em>Error in library(torchaudio): there is no package called ‘torchaudio’</em></p>\n<p>install.packages('torchaudio', lib = \"/kaggle/working\")<br>\n<em>also installing the dependency ‘av’\nWarning message in install.packages(\"torchaudio\", lib = \"/kaggle/working\"):\n“installation of package ‘av’ had non-zero exit status”\nWarning message in install.packages(\"torchaudio\", lib = \"/kaggle/working\"):\n“installation of package ‘torchaudio’ had non-zero exit status”</em></p>",
      "rawMarkdown": "Hi Jan, thanks for your reply!  Here's what I've tried.\n\nlibrary(av)\n*Error in library(av): there is no package called ‘av’*\n\ninstall.packages(\"av\",lib = \"/kaggle/working\")\n*Warning message in install.packages(\"av\", lib = \"/kaggle/working\"):\n“installation of package ‘av’ had non-zero exit status”*\n\nlibrary(av)\n*Error in library(av): there is no package called ‘av’*\n\nlibrary(torchaudio)\n*Error in library(torchaudio): there is no package called ‘torchaudio’*\n\ninstall.packages('torchaudio', lib = \"/kaggle/working\")\n*also installing the dependency ‘av’\nWarning message in install.packages(\"torchaudio\", lib = \"/kaggle/working\"):\n“installation of package ‘av’ had non-zero exit status”\nWarning message in install.packages(\"torchaudio\", lib = \"/kaggle/working\"):\n“installation of package ‘torchaudio’ had non-zero exit status”*",
      "votes": null
    },
    {
      "id": "2205416",
      "postDate": "04/01/2023 14:17:34",
      "content": "<p>Load ogg files to a vector.  I found a few ways to do this but by far the fastest way was:</p>\n<p>av::av_audio_convert('filename.ogg','filename.wav'))<br>\nmy_vector&lt;-audio::load.wave('filename.wav')</p>\n<p>The second command above executes nearly instantly on my computer, and the first command is needed to get the wav.</p>\n<p>The av and torchaudio packages each have a way to load directly from ogg to RAM but not quickly enough to load 200 10 min files in 2 hours.  </p>\n<p>So it seems I need the 'av' package to convert ogg to wav for a workable solution.</p>\n<p>I also use torchaudio to get spectrograms within my training routine.  Understood there are other ways, so I could work around if I could get the av package or another quick ogg load method. Then I could rebuild my pipeline around a different spectrogram function.</p>",
      "rawMarkdown": "Load ogg files to a vector.  I found a few ways to do this but by far the fastest way was:\n\nav::av_audio_convert('filename.ogg','filename.wav'))\nmy_vector<-audio::load.wave('filename.wav')\n\nThe second command above executes nearly instantly on my computer, and the first command is needed to get the wav.\n\nThe av and torchaudio packages each have a way to load directly from ogg to RAM but not quickly enough to load 200 10 min files in 2 hours.  \n\nSo it seems I need the 'av' package to convert ogg to wav for a workable solution.\n\nI also use torchaudio to get spectrograms within my training routine.  Understood there are other ways, so I could work around if I could get the av package or another quick ogg load method. Then I could rebuild my pipeline around a different spectrogram function.",
      "votes": null
    },
    {
      "id": "2206842",
      "postDate": "04/02/2023 23:56:32",
      "content": "<p>`One option would be to train in RStudio and upload model(s)  as dataset<br>\nin python such as Awsaf's .78 infer<br>\n<a href=\"https://www.kaggle.com/code/awsaf49/birdclef23-effnet-fsr-cutmixup-infer\" target=\"_blank\">https://www.kaggle.com/code/awsaf49/birdclef23-effnet-fsr-cutmixup-infer</a></p>\n<p>I used Python Script like in notebook to get spectrogram in Spyder<br>\nspec = Audio2Spec(audio)<br>\nbb = spec.numpy()</p>\n<h1>convert -50 to 50 output from audio2spec to 0 to 255</h1>\n<h1>so can save as image for use as flow from directory</h1>\n<p>bb = (bb+50)*2.55<br>\noo = specdf['fout'][n]<br>\ncv2.imwrite(oo,bb)</p>\n<p>divide image input by 255 for model input 0-1</p>\n<p>In python infer:<br>\nspecs = Audio2Spec(chunks)<br>\nspecs = Spec2Img(specs)</p>\n<h1>convert -50 to 50 output from audio2spec to 0 to 1 for model</h1>\n<p>b= (specs.numpy() + 50)/100<br>\nchunk_preds = np.zeros(shape=(len(specs), 264), dtype=np.float32)<br>\nrec_preds = model(b, training=False).numpy()`</p>\n<p>If using efficientnet no need to divide by 255 and infer would be b= (specs.numpy() + 50) * 2.55</p>\n<p>Also if using efficientnet in tensorflow 2.10 in RStudio or Spyder (because tensorflow&gt;= 2.11 has no gpu support) may get error when save model.<br>\nThe fix is: <a href=\"https://github.com/keras-team/keras/pull/17498/files\" target=\"_blank\">https://github.com/keras-team/keras/pull/17498/files</a><br>\n It involves changing line in efficientnet.py</p>",
      "rawMarkdown": "`One option would be to train in RStudio and upload model(s)  as dataset\nin python such as Awsaf's .78 infer\nhttps://www.kaggle.com/code/awsaf49/birdclef23-effnet-fsr-cutmixup-infer\n\nI used Python Script like in notebook to get spectrogram in Spyder\nspec = Audio2Spec(audio)\nbb = spec.numpy()\n# convert -50 to 50 output from audio2spec to 0 to 255\n# so can save as image for use as flow from directory \nbb = (bb+50)*2.55\noo = specdf['fout'][n]\ncv2.imwrite(oo,bb)\n\ndivide image input by 255 for model input 0-1\n\nIn python infer:\nspecs = Audio2Spec(chunks)\nspecs = Spec2Img(specs)\n# convert -50 to 50 output from audio2spec to 0 to 1 for model\nb= (specs.numpy() + 50)/100\nchunk_preds = np.zeros(shape=(len(specs), 264), dtype=np.float32)\nrec_preds = model(b, training=False).numpy()`\n\nIf using efficientnet no need to divide by 255 and infer would be b= (specs.numpy() + 50) * 2.55\n\nAlso if using efficientnet in tensorflow 2.10 in RStudio or Spyder (because tensorflow>= 2.11 has no gpu support) may get error when save model.\nThe fix is: https://github.com/keras-team/keras/pull/17498/files\n It involves changing line in efficientnet.py",
      "votes": null
    },
    {
      "id": "2211395",
      "postDate": "04/06/2023 02:30:49",
      "content": "<p>Thanks Thomas, I will give this mixed approach a try.  I'm a little worried about using a different spectrogram function for training and inference.  I think I'll see if I can use the reticulate package to embed python into my R script for training and inference.</p>",
      "rawMarkdown": "Thanks Thomas, I will give this mixed approach a try.  I'm a little worried about using a different spectrogram function for training and inference.  I think I'll see if I can use the reticulate package to embed python into my R script for training and inference.",
      "votes": null
    },
    {
      "id": "2212516",
      "postDate": "04/06/2023 20:21:22",
      "content": "<p>May still need to infer with python as tensorflow_io and librosa<br>\nare not in kaggle reticulate.</p>\n<p>in kaggle notebook:<br>\nlibrary(\"reticulate\")<br>\nlist.files('/root/.local/share/r-miniconda/envs/r-reticulate/lib/python3.8/site-packages')<br>\npy_config()</p>\n<p>I used this at first, now i use Spyder.<br>\nlibrary(keras)<br>\nlibrary(reticulate)<br>\nsource_python(\"C:\\r\\py2rbird.py\")<br>\nn&lt;- 1L<br>\nre &lt;- array( dim = c(n,128,192))<br>\nre[n,,]&lt;- getspec(\"C:\\bi\\train_audio\\abethr1\\XC128013.ogg\",0L)<br>\nsummary(re)<br>\n  Min. 1st Qu.  Median    Mean 3rd Qu.    Max. </p>\n<h2>  -32.786  -7.894  -1.001  -2.112   4.834  47.214 </h2>\n<pre><code>py2rbird.py\n</code></pre>\n<p>import sys, os<br>\n  import tensorflow as tf<br>\n  #tf.get_logger().setLevel('ERROR')<br>\n  #tf.autograph.set_verbosity(0)<br>\n  import pandas as pd<br>\n  import numpy as np<br>\n  import random<br>\n  from glob import glob<br>\n  from tqdm import tqdm<br>\n  #tqdm.pandas()<br>\n  import gc<br>\n  import librosa<br>\n  import sklearn<br>\n  import time<br>\n  #import matplotlib as mpl<br>\n  #import matplotlib.pyplot as plt<br>\n  #import librosa.display as lid<br>\n  #<br>\n  #import tensorflow as tf<br>\n  #tf.config.optimizer.set_jit(True) # enable xla for speed up<br>\n  import tensorflow_io as tfio<br>\n  #import tensorflow.keras.backend as K<br>\n  #import efficientnet.tfkeras as efn<br>\n  #import matplotlib.image<br>\n  import cv2</p>\n<p>class CFG:<br>\n    # Input image size and batch size<br>\n    img_size = [128, 384]<br>\n  batch_size = 16<br>\n  infer_bs = 2<br>\n  tta = 1<br>\n  drop_remainder = True</p>\n<p># STFT parameters<br>\n  duration = 5 # duration for test<br>\n  train_duration = 10<br>\n  sample_rate = 32000<br>\n  downsample = 1<br>\n  trim = True<br>\n  audio_len = duration<em>sample_rate\n  nfft = 2028\n  window = 2048\n  hop_length = train_duration</em>32000 // (img_size[1] - 1)<br>\n  fmin = 20<br>\n  fmax = 16000<br>\n  normalize = True</p>\n<p>def load_audio(filepath, sr=32000, normalize=True):<br>\n    audio, orig_sr = librosa.load(filepath, sr=None)<br>\n  if sr!=orig_sr:<br>\n    audio = librosa.resample( orig_sr, sr)<br>\n  audio = audio.astype('float32').ravel()<br>\n  audio = tf.convert_to_tensor(audio)<br>\n  if normalize:<br>\n    audio = Normalize(audio)<br>\n  return audio</p>\n<p><a href=\"https://www.kaggle.com/tf.function\" target=\"_blank\">@tf.function</a>(jit_compile=True)<br>\n  def Normalize(data, min_max=True):<br>\n    # Compute the mean and standard deviation of the data<br>\n    MEAN = tf.math.reduce_mean(data)<br>\n    STD = tf.math.reduce_std(data)<br>\n    # Standardize the data<br>\n    data = tf.math.divide_no_nan(data - MEAN, STD)<br>\n    # Normalize to [0, 1]<br>\n    if min_max:<br>\n     MIN = tf.math.reduce_min(data)<br>\n    MAX = tf.math.reduce_max(data)<br>\n    data = tf.math.divide_no_nan(data - MIN, MAX - MIN)<br>\n    return data</p>\n<p><a href=\"https://www.kaggle.com/tf.function\" target=\"_blank\">@tf.function</a>(jit_compile=True)<br>\n  def Audio2Spec(audio, spec_shape = CFG.img_size, sr=CFG.sample_rate, <br>\n                 nfft=CFG.nfft, window=CFG.window, fmin=CFG.fmin, fmax=CFG.fmax, return_img=True):<br>\n    spec_height = spec_shape[0]<br>\n    spec_width = spec_shape[1]<br>\n    hop_length = tf.cast(CFG.hop_length, tf.int32) # sample rate * duration / spec width - 1 == 627<br>\n    spec = tfio.audio.spectrogram(audio, nfft=nfft, window=window, stride=hop_length)<br>\n    mel_spec = tfio.audio.melscale(spec, rate=sr, mels=spec_height, fmin=fmin, fmax=fmax)<br>\n    db_mel_spec = tfio.audio.dbscale(mel_spec, top_db=80)<br>\n    db_mel_spec = tf.linalg.matrix_transpose(db_mel_spec) # to keep it (batch, mel, time)<br>\n    return db_mel_spec</p>\n<p>def MakeFrame(audio, duration=5, sr=32000):<br>\n    frame_length = int(duration * sr)<br>\n    frame_step = int(duration * sr)<br>\n    chunks = tf.signal.frame(audio, frame_length, frame_step, pad_end=True)<br>\n    return chunks</p>\n<p>def getspec(ff,st):         <br>\n    audio = load_audio(ff) <br>\n    audio = audio[(st * 32000):((st+5) * 32000)]<br>\n    spec = Audio2Spec(audio)<br>\n    bb = spec.numpy()<br>\n    return bb</p>",
      "rawMarkdown": "May still need to infer with python as tensorflow_io and librosa\nare not in kaggle reticulate.\n\nin kaggle notebook:\nlibrary(\"reticulate\")\nlist.files('/root/.local/share/r-miniconda/envs/r-reticulate/lib/python3.8/site-packages')\npy_config()\n\nI used this at first, now i use Spyder.\nlibrary(keras)\nlibrary(reticulate)\nsource_python(\"C:\\\\r\\\\py2rbird.py\")\nn<- 1L\nre <- array( dim = c(n,128,192))\nre[n,,]<- getspec(\"C:\\\\bi\\\\train_audio\\\\abethr1\\\\XC128013.ogg\",0L)\nsummary(re)\n  Min. 1st Qu.  Median    Mean 3rd Qu.    Max. \n  -32.786  -7.894  -1.001  -2.112   4.834  47.214 \n-------------------------------------------------------------\n    py2rbird.py\n  import sys, os\n  import tensorflow as tf\n  #tf.get_logger().setLevel('ERROR')\n  #tf.autograph.set_verbosity(0)\n  import pandas as pd\n  import numpy as np\n  import random\n  from glob import glob\n  from tqdm import tqdm\n  #tqdm.pandas()\n  import gc\n  import librosa\n  import sklearn\n  import time\n  #import matplotlib as mpl\n  #import matplotlib.pyplot as plt\n  #import librosa.display as lid\n  #\n  #import tensorflow as tf\n  #tf.config.optimizer.set_jit(True) # enable xla for speed up\n  import tensorflow_io as tfio\n  #import tensorflow.keras.backend as K\n  #import efficientnet.tfkeras as efn\n  #import matplotlib.image\n  import cv2\n  \n  class CFG:\n    # Input image size and batch size\n    img_size = [128, 384]\n  batch_size = 16\n  infer_bs = 2\n  tta = 1\n  drop_remainder = True\n  \n  # STFT parameters\n  duration = 5 # duration for test\n  train_duration = 10\n  sample_rate = 32000\n  downsample = 1\n  trim = True\n  audio_len = duration*sample_rate\n  nfft = 2028\n  window = 2048\n  hop_length = train_duration*32000 // (img_size[1] - 1)\n  fmin = 20\n  fmax = 16000\n  normalize = True\n  \n  def load_audio(filepath, sr=32000, normalize=True):\n    audio, orig_sr = librosa.load(filepath, sr=None)\n  if sr!=orig_sr:\n    audio = librosa.resample( orig_sr, sr)\n  audio = audio.astype('float32').ravel()\n  audio = tf.convert_to_tensor(audio)\n  if normalize:\n    audio = Normalize(audio)\n  return audio\n  \n  @tf.function(jit_compile=True)\n  def Normalize(data, min_max=True):\n    # Compute the mean and standard deviation of the data\n    MEAN = tf.math.reduce_mean(data)\n    STD = tf.math.reduce_std(data)\n    # Standardize the data\n    data = tf.math.divide_no_nan(data - MEAN, STD)\n    # Normalize to [0, 1]\n    if min_max:\n     MIN = tf.math.reduce_min(data)\n    MAX = tf.math.reduce_max(data)\n    data = tf.math.divide_no_nan(data - MIN, MAX - MIN)\n    return data\n  \n  @tf.function(jit_compile=True)\n  def Audio2Spec(audio, spec_shape = CFG.img_size, sr=CFG.sample_rate, \n                 nfft=CFG.nfft, window=CFG.window, fmin=CFG.fmin, fmax=CFG.fmax, return_img=True):\n    spec_height = spec_shape[0]\n    spec_width = spec_shape[1]\n    hop_length = tf.cast(CFG.hop_length, tf.int32) # sample rate * duration / spec width - 1 == 627\n    spec = tfio.audio.spectrogram(audio, nfft=nfft, window=window, stride=hop_length)\n    mel_spec = tfio.audio.melscale(spec, rate=sr, mels=spec_height, fmin=fmin, fmax=fmax)\n    db_mel_spec = tfio.audio.dbscale(mel_spec, top_db=80)\n    db_mel_spec = tf.linalg.matrix_transpose(db_mel_spec) # to keep it (batch, mel, time)\n    return db_mel_spec\n  \n  def MakeFrame(audio, duration=5, sr=32000):\n    frame_length = int(duration * sr)\n    frame_step = int(duration * sr)\n    chunks = tf.signal.frame(audio, frame_length, frame_step, pad_end=True)\n    return chunks\n\n  \n  def getspec(ff,st):         \n    audio = load_audio(ff) \n    audio = audio[(st * 32000):((st+5) * 32000)]\n    spec = Audio2Spec(audio)\n    bb = spec.numpy()\n    return bb",
      "votes": null
    },
    {
      "id": "2223273",
      "postDate": "04/16/2023 03:53:43",
      "content": "<p>Big thanks Tom!  You got me back on track!  I'm going the route of doing inference with python now like you suggested.  I'm nearly there I just need to reproduce my R Torch model architecture in python exactly, hoping this doesn't give me issues.  You saved me a lot of time by avoiding trying reticulate in the kaggle notebook environment.  Also, you really gave me hope that I could use both R and python together.  I'm going to use reticulate when training on my computer so that I can use the identical python audio load and spectrogram function for training and inference.</p>\n<p>I think I've built up my python skills in this inference exercise so that if I'm unable to get my R torch model working with python for some reason, I'm pretty sure i can just go back and build it in python with a little effort.</p>",
      "rawMarkdown": "Big thanks Tom!  You got me back on track!  I'm going the route of doing inference with python now like you suggested.  I'm nearly there I just need to reproduce my R Torch model architecture in python exactly, hoping this doesn't give me issues.  You saved me a lot of time by avoiding trying reticulate in the kaggle notebook environment.  Also, you really gave me hope that I could use both R and python together.  I'm going to use reticulate when training on my computer so that I can use the identical python audio load and spectrogram function for training and inference.\n\nI think I've built up my python skills in this inference exercise so that if I'm unable to get my R torch model working with python for some reason, I'm pretty sure i can just go back and build it in python with a little effort.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2204881,
      "author_name": "andyatkinson",
      "author_url": "",
      "post_date": "04/01/2023 04:13:30",
      "content": "<p>torchaudio is part of my pipeline and doesn't work either :*(</p>",
      "votes": null,
      "replies": [
        {
          "id": 2205069,
          "author_name": "janbrederecke",
          "author_url": "",
          "post_date": "04/01/2023 08:03:42",
          "content": "<p>Hey Andy, <br>\nwhat exactly are you trying to do? </p>\n<p>Best,<br>\nJan</p>",
          "votes": null,
          "replies": [
            {
              "id": 2205416,
              "author_name": "andyatkinson",
              "author_url": "",
              "post_date": "04/01/2023 14:17:34",
              "content": "<p>Load ogg files to a vector.  I found a few ways to do this but by far the fastest way was:</p>\n<p>av::av_audio_convert('filename.ogg','filename.wav'))<br>\nmy_vector&lt;-audio::load.wave('filename.wav')</p>\n<p>The second command above executes nearly instantly on my computer, and the first command is needed to get the wav.</p>\n<p>The av and torchaudio packages each have a way to load directly from ogg to RAM but not quickly enough to load 200 10 min files in 2 hours.  </p>\n<p>So it seems I need the 'av' package to convert ogg to wav for a workable solution.</p>\n<p>I also use torchaudio to get spectrograms within my training routine.  Understood there are other ways, so I could work around if I could get the av package or another quick ogg load method. Then I could rebuild my pipeline around a different spectrogram function.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2205068,
      "author_name": "janbrederecke",
      "author_url": "",
      "post_date": "04/01/2023 07:59:55",
      "content": "<p>Hey Andy,</p>\n<p>would you mind sharing the exact Error message? There might be a solution to your problem if you elaborate a bit on what is happening!</p>\n<p>Best,<br>\nJan</p>",
      "votes": null,
      "replies": [
        {
          "id": 2205403,
          "author_name": "andyatkinson",
          "author_url": "",
          "post_date": "04/01/2023 14:02:37",
          "content": "<p>Hi Jan, thanks for your reply!  Here's what I've tried.</p>\n<p>library(av)<br>\n<em>Error in library(av): there is no package called ‘av’</em></p>\n<p>install.packages(\"av\",lib = \"/kaggle/working\")<br>\n<em>Warning message in install.packages(\"av\", lib = \"/kaggle/working\"):\n“installation of package ‘av’ had non-zero exit status”</em></p>\n<p>library(av)<br>\n<em>Error in library(av): there is no package called ‘av’</em></p>\n<p>library(torchaudio)<br>\n<em>Error in library(torchaudio): there is no package called ‘torchaudio’</em></p>\n<p>install.packages('torchaudio', lib = \"/kaggle/working\")<br>\n<em>also installing the dependency ‘av’\nWarning message in install.packages(\"torchaudio\", lib = \"/kaggle/working\"):\n“installation of package ‘av’ had non-zero exit status”\nWarning message in install.packages(\"torchaudio\", lib = \"/kaggle/working\"):\n“installation of package ‘torchaudio’ had non-zero exit status”</em></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2206842,
      "author_name": "tomkkk",
      "author_url": "",
      "post_date": "04/02/2023 23:56:32",
      "content": "<p>`One option would be to train in RStudio and upload model(s)  as dataset<br>\nin python such as Awsaf's .78 infer<br>\n<a href=\"https://www.kaggle.com/code/awsaf49/birdclef23-effnet-fsr-cutmixup-infer\" target=\"_blank\">https://www.kaggle.com/code/awsaf49/birdclef23-effnet-fsr-cutmixup-infer</a></p>\n<p>I used Python Script like in notebook to get spectrogram in Spyder<br>\nspec = Audio2Spec(audio)<br>\nbb = spec.numpy()</p>\n<h1>convert -50 to 50 output from audio2spec to 0 to 255</h1>\n<h1>so can save as image for use as flow from directory</h1>\n<p>bb = (bb+50)*2.55<br>\noo = specdf['fout'][n]<br>\ncv2.imwrite(oo,bb)</p>\n<p>divide image input by 255 for model input 0-1</p>\n<p>In python infer:<br>\nspecs = Audio2Spec(chunks)<br>\nspecs = Spec2Img(specs)</p>\n<h1>convert -50 to 50 output from audio2spec to 0 to 1 for model</h1>\n<p>b= (specs.numpy() + 50)/100<br>\nchunk_preds = np.zeros(shape=(len(specs), 264), dtype=np.float32)<br>\nrec_preds = model(b, training=False).numpy()`</p>\n<p>If using efficientnet no need to divide by 255 and infer would be b= (specs.numpy() + 50) * 2.55</p>\n<p>Also if using efficientnet in tensorflow 2.10 in RStudio or Spyder (because tensorflow&gt;= 2.11 has no gpu support) may get error when save model.<br>\nThe fix is: <a href=\"https://github.com/keras-team/keras/pull/17498/files\" target=\"_blank\">https://github.com/keras-team/keras/pull/17498/files</a><br>\n It involves changing line in efficientnet.py</p>",
      "votes": null,
      "replies": [
        {
          "id": 2211395,
          "author_name": "andyatkinson",
          "author_url": "",
          "post_date": "04/06/2023 02:30:49",
          "content": "<p>Thanks Thomas, I will give this mixed approach a try.  I'm a little worried about using a different spectrogram function for training and inference.  I think I'll see if I can use the reticulate package to embed python into my R script for training and inference.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2212516,
              "author_name": "tomkkk",
              "author_url": "",
              "post_date": "04/06/2023 20:21:22",
              "content": "<p>May still need to infer with python as tensorflow_io and librosa<br>\nare not in kaggle reticulate.</p>\n<p>in kaggle notebook:<br>\nlibrary(\"reticulate\")<br>\nlist.files('/root/.local/share/r-miniconda/envs/r-reticulate/lib/python3.8/site-packages')<br>\npy_config()</p>\n<p>I used this at first, now i use Spyder.<br>\nlibrary(keras)<br>\nlibrary(reticulate)<br>\nsource_python(\"C:\\r\\py2rbird.py\")<br>\nn&lt;- 1L<br>\nre &lt;- array( dim = c(n,128,192))<br>\nre[n,,]&lt;- getspec(\"C:\\bi\\train_audio\\abethr1\\XC128013.ogg\",0L)<br>\nsummary(re)<br>\n  Min. 1st Qu.  Median    Mean 3rd Qu.    Max. </p>\n<h2>  -32.786  -7.894  -1.001  -2.112   4.834  47.214 </h2>\n<pre><code>py2rbird.py\n</code></pre>\n<p>import sys, os<br>\n  import tensorflow as tf<br>\n  #tf.get_logger().setLevel('ERROR')<br>\n  #tf.autograph.set_verbosity(0)<br>\n  import pandas as pd<br>\n  import numpy as np<br>\n  import random<br>\n  from glob import glob<br>\n  from tqdm import tqdm<br>\n  #tqdm.pandas()<br>\n  import gc<br>\n  import librosa<br>\n  import sklearn<br>\n  import time<br>\n  #import matplotlib as mpl<br>\n  #import matplotlib.pyplot as plt<br>\n  #import librosa.display as lid<br>\n  #<br>\n  #import tensorflow as tf<br>\n  #tf.config.optimizer.set_jit(True) # enable xla for speed up<br>\n  import tensorflow_io as tfio<br>\n  #import tensorflow.keras.backend as K<br>\n  #import efficientnet.tfkeras as efn<br>\n  #import matplotlib.image<br>\n  import cv2</p>\n<p>class CFG:<br>\n    # Input image size and batch size<br>\n    img_size = [128, 384]<br>\n  batch_size = 16<br>\n  infer_bs = 2<br>\n  tta = 1<br>\n  drop_remainder = True</p>\n<p># STFT parameters<br>\n  duration = 5 # duration for test<br>\n  train_duration = 10<br>\n  sample_rate = 32000<br>\n  downsample = 1<br>\n  trim = True<br>\n  audio_len = duration<em>sample_rate\n  nfft = 2028\n  window = 2048\n  hop_length = train_duration</em>32000 // (img_size[1] - 1)<br>\n  fmin = 20<br>\n  fmax = 16000<br>\n  normalize = True</p>\n<p>def load_audio(filepath, sr=32000, normalize=True):<br>\n    audio, orig_sr = librosa.load(filepath, sr=None)<br>\n  if sr!=orig_sr:<br>\n    audio = librosa.resample( orig_sr, sr)<br>\n  audio = audio.astype('float32').ravel()<br>\n  audio = tf.convert_to_tensor(audio)<br>\n  if normalize:<br>\n    audio = Normalize(audio)<br>\n  return audio</p>\n<p><a href=\"https://www.kaggle.com/tf.function\" target=\"_blank\">@tf.function</a>(jit_compile=True)<br>\n  def Normalize(data, min_max=True):<br>\n    # Compute the mean and standard deviation of the data<br>\n    MEAN = tf.math.reduce_mean(data)<br>\n    STD = tf.math.reduce_std(data)<br>\n    # Standardize the data<br>\n    data = tf.math.divide_no_nan(data - MEAN, STD)<br>\n    # Normalize to [0, 1]<br>\n    if min_max:<br>\n     MIN = tf.math.reduce_min(data)<br>\n    MAX = tf.math.reduce_max(data)<br>\n    data = tf.math.divide_no_nan(data - MIN, MAX - MIN)<br>\n    return data</p>\n<p><a href=\"https://www.kaggle.com/tf.function\" target=\"_blank\">@tf.function</a>(jit_compile=True)<br>\n  def Audio2Spec(audio, spec_shape = CFG.img_size, sr=CFG.sample_rate, <br>\n                 nfft=CFG.nfft, window=CFG.window, fmin=CFG.fmin, fmax=CFG.fmax, return_img=True):<br>\n    spec_height = spec_shape[0]<br>\n    spec_width = spec_shape[1]<br>\n    hop_length = tf.cast(CFG.hop_length, tf.int32) # sample rate * duration / spec width - 1 == 627<br>\n    spec = tfio.audio.spectrogram(audio, nfft=nfft, window=window, stride=hop_length)<br>\n    mel_spec = tfio.audio.melscale(spec, rate=sr, mels=spec_height, fmin=fmin, fmax=fmax)<br>\n    db_mel_spec = tfio.audio.dbscale(mel_spec, top_db=80)<br>\n    db_mel_spec = tf.linalg.matrix_transpose(db_mel_spec) # to keep it (batch, mel, time)<br>\n    return db_mel_spec</p>\n<p>def MakeFrame(audio, duration=5, sr=32000):<br>\n    frame_length = int(duration * sr)<br>\n    frame_step = int(duration * sr)<br>\n    chunks = tf.signal.frame(audio, frame_length, frame_step, pad_end=True)<br>\n    return chunks</p>\n<p>def getspec(ff,st):         <br>\n    audio = load_audio(ff) <br>\n    audio = audio[(st * 32000):((st+5) * 32000)]<br>\n    spec = Audio2Spec(audio)<br>\n    bb = spec.numpy()<br>\n    return bb</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2223273,
                  "author_name": "andyatkinson",
                  "author_url": "",
                  "post_date": "04/16/2023 03:53:43",
                  "content": "<p>Big thanks Tom!  You got me back on track!  I'm going the route of doing inference with python now like you suggested.  I'm nearly there I just need to reproduce my R Torch model architecture in python exactly, hoping this doesn't give me issues.  You saved me a lot of time by avoiding trying reticulate in the kaggle notebook environment.  Also, you really gave me hope that I could use both R and python together.  I'm going to use reticulate when training on my computer so that I can use the identical python audio load and spectrogram function for training and inference.</p>\n<p>I think I've built up my python skills in this inference exercise so that if I'm unable to get my R torch model working with python for some reason, I'm pretty sure i can just go back and build it in python with a little effort.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2204869": "Bummer, seems like another competition where I've built out a full R model and one of the CRAN package won't work in the notebook.  The package I need is called 'av' it seems to be the only package to convert an ogg to wav.  Any tips?  Is it a known thing that Kaggle notebook competitions are bad for R?",
    "2204881": "torchaudio is part of my pipeline and doesn't work either :*(",
    "2205068": "Hey Andy,\n\nwould you mind sharing the exact Error message? There might be a solution to your problem if you elaborate a bit on what is happening!\n\nBest,\nJan",
    "2205069": "Hey Andy, \nwhat exactly are you trying to do? \n\nBest,\nJan",
    "2205403": "Hi Jan, thanks for your reply!  Here's what I've tried.\n\nlibrary(av)\n*Error in library(av): there is no package called ‘av’*\n\ninstall.packages(\"av\",lib = \"/kaggle/working\")\n*Warning message in install.packages(\"av\", lib = \"/kaggle/working\"):\n“installation of package ‘av’ had non-zero exit status”*\n\nlibrary(av)\n*Error in library(av): there is no package called ‘av’*\n\nlibrary(torchaudio)\n*Error in library(torchaudio): there is no package called ‘torchaudio’*\n\ninstall.packages('torchaudio', lib = \"/kaggle/working\")\n*also installing the dependency ‘av’\nWarning message in install.packages(\"torchaudio\", lib = \"/kaggle/working\"):\n“installation of package ‘av’ had non-zero exit status”\nWarning message in install.packages(\"torchaudio\", lib = \"/kaggle/working\"):\n“installation of package ‘torchaudio’ had non-zero exit status”*",
    "2205416": "Load ogg files to a vector.  I found a few ways to do this but by far the fastest way was:\n\nav::av_audio_convert('filename.ogg','filename.wav'))\nmy_vector<-audio::load.wave('filename.wav')\n\nThe second command above executes nearly instantly on my computer, and the first command is needed to get the wav.\n\nThe av and torchaudio packages each have a way to load directly from ogg to RAM but not quickly enough to load 200 10 min files in 2 hours.  \n\nSo it seems I need the 'av' package to convert ogg to wav for a workable solution.\n\nI also use torchaudio to get spectrograms within my training routine.  Understood there are other ways, so I could work around if I could get the av package or another quick ogg load method. Then I could rebuild my pipeline around a different spectrogram function.",
    "2206842": "`One option would be to train in RStudio and upload model(s)  as dataset\nin python such as Awsaf's .78 infer\nhttps://www.kaggle.com/code/awsaf49/birdclef23-effnet-fsr-cutmixup-infer\n\nI used Python Script like in notebook to get spectrogram in Spyder\nspec = Audio2Spec(audio)\nbb = spec.numpy()\n# convert -50 to 50 output from audio2spec to 0 to 255\n# so can save as image for use as flow from directory \nbb = (bb+50)*2.55\noo = specdf['fout'][n]\ncv2.imwrite(oo,bb)\n\ndivide image input by 255 for model input 0-1\n\nIn python infer:\nspecs = Audio2Spec(chunks)\nspecs = Spec2Img(specs)\n# convert -50 to 50 output from audio2spec to 0 to 1 for model\nb= (specs.numpy() + 50)/100\nchunk_preds = np.zeros(shape=(len(specs), 264), dtype=np.float32)\nrec_preds = model(b, training=False).numpy()`\n\nIf using efficientnet no need to divide by 255 and infer would be b= (specs.numpy() + 50) * 2.55\n\nAlso if using efficientnet in tensorflow 2.10 in RStudio or Spyder (because tensorflow>= 2.11 has no gpu support) may get error when save model.\nThe fix is: https://github.com/keras-team/keras/pull/17498/files\n It involves changing line in efficientnet.py",
    "2211395": "Thanks Thomas, I will give this mixed approach a try.  I'm a little worried about using a different spectrogram function for training and inference.  I think I'll see if I can use the reticulate package to embed python into my R script for training and inference.",
    "2212516": "May still need to infer with python as tensorflow_io and librosa\nare not in kaggle reticulate.\n\nin kaggle notebook:\nlibrary(\"reticulate\")\nlist.files('/root/.local/share/r-miniconda/envs/r-reticulate/lib/python3.8/site-packages')\npy_config()\n\nI used this at first, now i use Spyder.\nlibrary(keras)\nlibrary(reticulate)\nsource_python(\"C:\\\\r\\\\py2rbird.py\")\nn<- 1L\nre <- array( dim = c(n,128,192))\nre[n,,]<- getspec(\"C:\\\\bi\\\\train_audio\\\\abethr1\\\\XC128013.ogg\",0L)\nsummary(re)\n  Min. 1st Qu.  Median    Mean 3rd Qu.    Max. \n  -32.786  -7.894  -1.001  -2.112   4.834  47.214 \n-------------------------------------------------------------\n    py2rbird.py\n  import sys, os\n  import tensorflow as tf\n  #tf.get_logger().setLevel('ERROR')\n  #tf.autograph.set_verbosity(0)\n  import pandas as pd\n  import numpy as np\n  import random\n  from glob import glob\n  from tqdm import tqdm\n  #tqdm.pandas()\n  import gc\n  import librosa\n  import sklearn\n  import time\n  #import matplotlib as mpl\n  #import matplotlib.pyplot as plt\n  #import librosa.display as lid\n  #\n  #import tensorflow as tf\n  #tf.config.optimizer.set_jit(True) # enable xla for speed up\n  import tensorflow_io as tfio\n  #import tensorflow.keras.backend as K\n  #import efficientnet.tfkeras as efn\n  #import matplotlib.image\n  import cv2\n  \n  class CFG:\n    # Input image size and batch size\n    img_size = [128, 384]\n  batch_size = 16\n  infer_bs = 2\n  tta = 1\n  drop_remainder = True\n  \n  # STFT parameters\n  duration = 5 # duration for test\n  train_duration = 10\n  sample_rate = 32000\n  downsample = 1\n  trim = True\n  audio_len = duration*sample_rate\n  nfft = 2028\n  window = 2048\n  hop_length = train_duration*32000 // (img_size[1] - 1)\n  fmin = 20\n  fmax = 16000\n  normalize = True\n  \n  def load_audio(filepath, sr=32000, normalize=True):\n    audio, orig_sr = librosa.load(filepath, sr=None)\n  if sr!=orig_sr:\n    audio = librosa.resample( orig_sr, sr)\n  audio = audio.astype('float32').ravel()\n  audio = tf.convert_to_tensor(audio)\n  if normalize:\n    audio = Normalize(audio)\n  return audio\n  \n  @tf.function(jit_compile=True)\n  def Normalize(data, min_max=True):\n    # Compute the mean and standard deviation of the data\n    MEAN = tf.math.reduce_mean(data)\n    STD = tf.math.reduce_std(data)\n    # Standardize the data\n    data = tf.math.divide_no_nan(data - MEAN, STD)\n    # Normalize to [0, 1]\n    if min_max:\n     MIN = tf.math.reduce_min(data)\n    MAX = tf.math.reduce_max(data)\n    data = tf.math.divide_no_nan(data - MIN, MAX - MIN)\n    return data\n  \n  @tf.function(jit_compile=True)\n  def Audio2Spec(audio, spec_shape = CFG.img_size, sr=CFG.sample_rate, \n                 nfft=CFG.nfft, window=CFG.window, fmin=CFG.fmin, fmax=CFG.fmax, return_img=True):\n    spec_height = spec_shape[0]\n    spec_width = spec_shape[1]\n    hop_length = tf.cast(CFG.hop_length, tf.int32) # sample rate * duration / spec width - 1 == 627\n    spec = tfio.audio.spectrogram(audio, nfft=nfft, window=window, stride=hop_length)\n    mel_spec = tfio.audio.melscale(spec, rate=sr, mels=spec_height, fmin=fmin, fmax=fmax)\n    db_mel_spec = tfio.audio.dbscale(mel_spec, top_db=80)\n    db_mel_spec = tf.linalg.matrix_transpose(db_mel_spec) # to keep it (batch, mel, time)\n    return db_mel_spec\n  \n  def MakeFrame(audio, duration=5, sr=32000):\n    frame_length = int(duration * sr)\n    frame_step = int(duration * sr)\n    chunks = tf.signal.frame(audio, frame_length, frame_step, pad_end=True)\n    return chunks\n\n  \n  def getspec(ff,st):         \n    audio = load_audio(ff) \n    audio = audio[(st * 32000):((st+5) * 32000)]\n    spec = Audio2Spec(audio)\n    bb = spec.numpy()\n    return bb",
    "2223273": "Big thanks Tom!  You got me back on track!  I'm going the route of doing inference with python now like you suggested.  I'm nearly there I just need to reproduce my R Torch model architecture in python exactly, hoping this doesn't give me issues.  You saved me a lot of time by avoiding trying reticulate in the kaggle notebook environment.  Also, you really gave me hope that I could use both R and python together.  I'm going to use reticulate when training on my computer so that I can use the identical python audio load and spectrogram function for training and inference.\n\nI think I've built up my python skills in this inference exercise so that if I'm unable to get my R torch model working with python for some reason, I'm pretty sure i can just go back and build it in python with a little effort."
  },
  "source": "meta"
}