{
  "id": 443046,
  "title": "Data reading and preprocessing trouble",
  "url": "/competitions/bengaliai-speech/discussion/443046",
  "author_name": "Prabhakar_Nimmagadda",
  "post_date": "2023-09-25T11:52:07.027000",
  "votes": 0,
  "comment_count": 0,
  "views": 0,
  "content": "<p>At the stage of preprocessing of the training data, we have to load the data feed the each file to the preprocessed function for extracting MFCCs and spectrogram, here I stuck at the feeding each file to the function because the files are huge more than 9 Lakhs reading each file with loop is consuming more time. Like to process the 15 % of the files from the train directory it is taking more than a whole day using librosa library for data loading. <br>\nCan any body help me in this to speedup the data loading and pre processing step. <br>\nhere is my initial approach approach:</p>\n<p>train_file = pd.read_csv('/kaggle/input/bengaliai-speech/train.csv')<br>\nfile_id = train_file['id'].unique()</p>\n<h1>Select the first 15% of the file IDs</h1>\n<p>file_ids_15pct = file_id[:int(0.01 * len(file_id))]</p>\n<p>for A_id in tqdm(file_ids_15pct):<br>\n    processed_Ad_Wo = prepro_WoAUG(A_id)</p>\n<h1>Define the number of CPU cores to use for parallel processing</h1>\n<p>num_cores = multiprocessing.cpu_count()</p>\n<h1>Process the audio files in parallel</h1>\n<p>pool = Pool(num_cores)<br>\npool.map(prepro_WoAUG, file_ids_15pct)</p>\n<h1>Close the pool</h1>\n<p>pool.close()<br>\npool.join()</p>",
  "messages": [
    {
      "id": 2455237,
      "postDate": "2023-09-25T11:52:07.027Z",
      "content": "<p>At the stage of preprocessing of the training data, we have to load the data feed the each file to the preprocessed function for extracting MFCCs and spectrogram, here I stuck at the feeding each file to the function because the files are huge more than 9 Lakhs reading each file with loop is consuming more time. Like to process the 15 % of the files from the train directory it is taking more than a whole day using librosa library for data loading. <br>\nCan any body help me in this to speedup the data loading and pre processing step. <br>\nhere is my initial approach approach:</p>\n<p>train_file = pd.read_csv('/kaggle/input/bengaliai-speech/train.csv')<br>\nfile_id = train_file['id'].unique()</p>\n<h1>Select the first 15% of the file IDs</h1>\n<p>file_ids_15pct = file_id[:int(0.01 * len(file_id))]</p>\n<p>for A_id in tqdm(file_ids_15pct):<br>\n    processed_Ad_Wo = prepro_WoAUG(A_id)</p>\n<h1>Define the number of CPU cores to use for parallel processing</h1>\n<p>num_cores = multiprocessing.cpu_count()</p>\n<h1>Process the audio files in parallel</h1>\n<p>pool = Pool(num_cores)<br>\npool.map(prepro_WoAUG, file_ids_15pct)</p>\n<h1>Close the pool</h1>\n<p>pool.close()<br>\npool.join()</p>",
      "rawMarkdown": "At the stage of preprocessing of the training data, we have to load the data feed the each file to the preprocessed function for extracting MFCCs and spectrogram, here I stuck at the feeding each file to the function because the files are huge more than 9 Lakhs reading each file with loop is consuming more time. Like to process the 15 % of the files from the train directory it is taking more than a whole day using librosa library for data loading. \nCan any body help me in this to speedup the data loading and pre processing step. \nhere is my initial approach approach:\n\ntrain_file = pd.read_csv('/kaggle/input/bengaliai-speech/train.csv')\nfile_id = train_file['id'].unique()\n\n# Select the first 15% of the file IDs\nfile_ids_15pct = file_id[:int(0.01 * len(file_id))]\n\n\nfor A_id in tqdm(file_ids_15pct):\n    processed_Ad_Wo = prepro_WoAUG(A_id)\n\n\n# Define the number of CPU cores to use for parallel processing\nnum_cores = multiprocessing.cpu_count()\n\n# Process the audio files in parallel\npool = Pool(num_cores)\npool.map(prepro_WoAUG, file_ids_15pct)\n\n# Close the pool\npool.close()\npool.join()\n"
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2455237": "At the stage of preprocessing of the training data, we have to load the data feed the each file to the preprocessed function for extracting MFCCs and spectrogram, here I stuck at the feeding each file to the function because the files are huge more than 9 Lakhs reading each file with loop is consuming more time. Like to process the 15 % of the files from the train directory it is taking more than a whole day using librosa library for data loading. \nCan any body help me in this to speedup the data loading and pre processing step. \nhere is my initial approach approach:\n\ntrain_file = pd.read_csv('/kaggle/input/bengaliai-speech/train.csv')\nfile_id = train_file['id'].unique()\n\n# Select the first 15% of the file IDs\nfile_ids_15pct = file_id[:int(0.01 * len(file_id))]\n\n\nfor A_id in tqdm(file_ids_15pct):\n    processed_Ad_Wo = prepro_WoAUG(A_id)\n\n\n# Define the number of CPU cores to use for parallel processing\nnum_cores = multiprocessing.cpu_count()\n\n# Process the audio files in parallel\npool = Pool(num_cores)\npool.map(prepro_WoAUG, file_ids_15pct)\n\n# Close the pool\npool.close()\npool.join()\n"
  }
}