{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.11.11"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":91844,"databundleVersionId":11361821,"sourceType":"competition"},{"sourceId":12051777,"sourceType":"datasetVersion","datasetId":7505901}],"dockerImageVersionId":31040,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false},"papermill":{"default_parameters":{},"duration":66.275304,"end_time":"2025-05-20T09:23:45.886660","environment_variables":{},"exception":null,"input_path":"__notebook__.ipynb","output_path":"__notebook__.ipynb","parameters":{},"start_time":"2025-05-20T09:22:39.611356","version":"2.6.0"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## 📌 Introduction\n\nThis notebook performs **inference** on the BirdCLEF 2025 test soundscapes using a trained Bird Sound Event Detection (SED) model. The pipeline includes:\n\n- ✅ Voice Activity Detection (VAD) to remove silent and human-voice segments  \n- 🔁 Chunking test audio files into 5-second windows  \n- 🎛️ Mel spectrogram generation for each chunk  \n- 🧠 Inference using a pre-trained model  \n- 📉 Prediction smoothing for temporal consistency  \n- 📄 Generating the final `submission.csv` file in the required format\n\nThis setup ensures efficient, high-quality inference and is fully aligned with BirdCLEF 2025 submission standards.\n\n---","metadata":{}},{"cell_type":"markdown","source":"## 🔗 BirdCLEF 2025 - Project Notebook Links\n\nHere are the different stages of my BirdCLEF 2025 pipeline, organized by functionality:\n\n### 📊 Data Preparation\n- [BirdCLEF 2025 - Data Preparation](https://www.kaggle.com/code/sheemamasood/birdclef-2025-data-prepartion)\n\n### 🎛️ Mel Spectrogram Generation\n- [BirdCLEF 2025 - Mel Generation](https://www.kaggle.com/code/sheemamasood/birdclef2025-mel-generation)\n\n### 🏷️ Pseudo Labelling for SSL\n- [BirdCLEF 2025 - Pseudo Labelling for SSL](https://www.kaggle.com/code/sheemamasood/birdclef2025-psedolabelling-for-ssl)\n\n### 🧠 Model Training\n- [BirdCLEF 2025 - Model Training (Phase 1)](https://www.kaggle.com/code/sheemamasood/birdclef2025-model-training-phase1)\n\n### 📦 Inference & Submissions\n- [BirdCLEF 2025 - Submissions](https://www.kaggle.com/code/sheemamasood/birdclef2025-submissions)\n","metadata":{}},{"cell_type":"code","source":"# Standard Libraries\nimport os\nimport gc\nimport time\nimport math\nimport random\nimport warnings\nimport logging\nfrom pathlib import Path\nfrom glob import glob\nfrom typing import Union\nimport copy\n\n# Data Handling\nimport numpy as np\nimport pandas as pd\nimport joblib\nimport pickle\nimport collections\n\n# Audio Processing\nimport librosa\nimport librosa.display\nimport torchaudio\n\n# Machine Learning & PyTorch\nimport torch\nimport torch.nn as nn\nimport torchaudio.transforms as T\nimport torchaudio.functional as F\nimport torch.optim as optim\nfrom torch.utils.data import Dataset, DataLoader, WeightedRandomSampler\nimport timm\n\n# Visualization\nimport cv2\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# Progress Bars\nfrom tqdm import tqdm\nfrom tqdm.notebook import tqdm as notebook_tqdm\n\n# Logging and Warnings\nwarnings.filterwarnings(\"ignore\")\nlogging.basicConfig(level=logging.ERROR)\n\n# Check versions\nprint(f\"librosa version : {librosa.__version__}\")\nprint(f\"librosa files : {librosa.__file__}\")\nprint(\"✅ All libraries imported in the environment.\")\n\nimport sys\nsys.path.append('/kaggle/input/birdcleft-clean-and-vad-filtered-data')\n\nfrom utils_vad import get_speech_timestamps","metadata":{"execution":{"iopub.status.busy":"2025-06-17T00:21:17.278309Z","iopub.execute_input":"2025-06-17T00:21:17.278624Z","iopub.status.idle":"2025-06-17T00:21:17.286865Z","shell.execute_reply.started":"2025-06-17T00:21:17.278601Z","shell.execute_reply":"2025-06-17T00:21:17.285855Z"},"papermill":{"duration":23.046365,"end_time":"2025-05-20T09:23:07.831968","exception":false,"start_time":"2025-05-20T09:22:44.785603","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ===================== CONFIG =====================\nclass Config:\n    # General\n    seed = 42\n    print_freq = 100\n    num_workers = 4\n    device = 'cuda' if torch.cuda.is_available() else 'cpu'\n\n    # Mel spectrogram parameters (for converting audio to image)\n    N_FFT = 1024       # FFT window size\n    HOP_LENGTH = 512   # Step size for each frame\n    FS = 32000\n    FMIN = 50          # Minimum Mel frequency\n    FMAX = 14000       # Maximum Mel frequency\n    N_MELS = 128\n     \n    SR = 32000\n    TARGET_DURATION= 5\n    train_duration = 10\n    # Mel/Image\n    MEL_SHAPE = (3, 256, 256)\n    TARGET_SHAPE = (3, 256, 256)\n\n    DEVICE = \"cuda\" if torch.cuda.is_available() else \"cpu\"\n    \n    # Paths (change these if needed)\n    test_soundscapes = \"/kaggle/input/birdclef-2025/test_soundscapes\"\n    submission_csv = \"/kaggle/input/birdclef-2025/sample_submission.csv\"\n    model_path = \"/kaggle/input/birdcleft-clean-and-vad-filtered-data/best_model_187.pth\"\n    backbone_weights = \"/kaggle/input/birdcleft-clean-and-vad-filtered-data/seresnext_backbone_weights.pth\"\n    master_labels = \"/kaggle/input/birdcleft-clean-and-vad-filtered-data/valid_labels.pkl\"\n    vad_model_path = \"/kaggle/input/birdcleft-clean-and-vad-filtered-data/silero_vad.jit\"\n    \nconfig = Config()\n\ndevice = torch.device(config.device)\n\nspecies_ids = pd.read_csv(config.submission_csv).columns[1:].tolist()\n\nprint(\"test_soundscapes path:\", config.test_soundscapes)","metadata":{"papermill":{"duration":0.012329,"end_time":"2025-05-20T09:23:07.848146","exception":false,"start_time":"2025-05-20T09:23:07.835817","status":"completed"},"tags":[],"trusted":true,"execution":{"iopub.status.busy":"2025-06-17T00:21:17.576293Z","iopub.execute_input":"2025-06-17T00:21:17.576612Z","iopub.status.idle":"2025-06-17T00:21:17.590566Z","shell.execute_reply.started":"2025-06-17T00:21:17.576563Z","shell.execute_reply":"2025-06-17T00:21:17.589643Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Data Utils","metadata":{}},{"cell_type":"code","source":"# ===================== DATA UTILS =====================\ndef load_species_ids(submission_csv):\n    \"\"\"Load 206 submission class names from sample_submission.csv (skip 'row_id').\"\"\"\n    return pd.read_csv(submission_csv).columns[1:].tolist()\n\ndef load_master_labels(master_labels):\n    \"\"\"Load 187 model class labels (output order) from a pickle file.\"\"\"\n    with open(master_labels, 'rb') as f:\n        return pickle.load(f)\n\ndef load_vad_model(jit_path, device='cpu'):\n    \"\"\"\n    Load Silero VAD JIT model from file.\n    Args:\n        jit_path (str): Path to the .jit model file.\n        device (str or torch.device): 'cpu' or 'cuda'\n    Returns:\n        torch.jit.ScriptModule: Loaded model in eval mode.\n    \"\"\"\n    model = torch.jit.load(jit_path, map_location=device)\n    model.eval()\n    return model\n\n# ====== MEL SPECTROGRAM TRANSFORM ======\ndef get_mel_transform(cfg):\n    return T.MelSpectrogram(\n        sample_rate=cfg.FS,\n        n_fft=cfg.N_FFT,\n        hop_length=cfg.HOP_LENGTH,\n        n_mels=cfg.N_MELS,\n        f_min=cfg.FMIN,\n        f_max=cfg.FMAX,\n        power=2.0,\n    ).to(cfg.DEVICE)\n\nmel_transform = get_mel_transform(config)\n\ndef audio2melspec_gpu(audio_data, cfg, mel_transform):\n    if np.isnan(audio_data).any():\n        mean_val = np.nanmean(audio_data)\n        audio_data = np.nan_to_num(audio_data, nan=mean_val)\n    waveform = torch.tensor(audio_data, dtype=torch.float32).unsqueeze(0).to(cfg.DEVICE)\n    mel = mel_transform(waveform)\n    #mel_db = F.amplitude_to_DB(mel, multiplier=10.0, amin=1e-10, db_multiplier=0.0)\n    mel_db = F.amplitude_to_DB(mel, multiplier=10.0, amin=1e-10, db_multiplier=0.0)\n    mel_db = (mel_db - mel_db.min()) / (mel_db.max() - mel_db.min() + 1e-8)\n    return mel_db.squeeze(0).cpu().numpy()\n\ndef prepare_audio(audio, target_len):\n    current_len = len(audio)\n    if current_len < target_len:\n        pad_left = (target_len - current_len) // 2\n        pad_right = target_len - current_len - pad_left\n        audio = np.pad(audio, (pad_left, pad_right), mode='constant')\n    elif current_len > target_len:\n        start = (current_len - target_len) // 2\n        audio = audio[start: start + target_len]\n    return audio\n\n# ======  GET TEST FILES ======\ndef get_test_files(test_dir, fallback_dir=None, fallback_n=10):\n    test_files = list(Path(test_dir).glob('*.ogg'))\n    if len(test_files) == 0 and fallback_dir:\n        test_files = sorted(list(Path(fallback_dir).glob('*.ogg')))[:fallback_n]\n    return test_files\n\n# ======  CHUNKING ======\ndef get_chunks_for_files(soundscape_files, sample_rate=32000, chunk_len_sec=5):\n    chunk_records = []\n    for file_path in tqdm(soundscape_files, desc=\"Processing test soundscapes\"):\n        file = Path(file_path).name\n        try:\n            y, sr = librosa.load(file_path, sr=sample_rate)\n            duration_sec = librosa.get_duration(y=y, sr=sr)\n            samples_per_chunk = int(chunk_len_sec * sr)\n            num_chunks = int(duration_sec // chunk_len_sec)\n            samplename = file.replace(\".ogg\", \"\")\n            for i in range(num_chunks):\n                start_sample = i * samples_per_chunk\n                end_sample = (i + 1) * samples_per_chunk\n                chunk_records.append({\n                    'chunk_id': f\"{samplename}_chunk{i}\",\n                    'filename': file,\n                    'filepath': str(file_path),\n                    'samplename': samplename,\n                    'start_sec': i * chunk_len_sec,\n                    'end_sec': (i + 1) * chunk_len_sec,\n                    'start_sample': start_sample,\n                    'end_sample': end_sample,\n                    'duration': duration_sec,\n                    'row_id': f\"{samplename}_{(i + 1) * chunk_len_sec}\",  \n                })\n        except Exception as e:\n            print(f\"❌ Failed to process {file}: {e}\")\n    chunked_df = pd.DataFrame(chunk_records)\n    return chunked_df\n\n\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-17T00:21:17.991333Z","iopub.execute_input":"2025-06-17T00:21:17.991671Z","iopub.status.idle":"2025-06-17T00:21:18.042981Z","shell.execute_reply.started":"2025-06-17T00:21:17.991644Z","shell.execute_reply":"2025-06-17T00:21:18.042122Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ====== 3. DATASET CLASS ======\nclass BirdTestMelDataset(Dataset):\n    def __init__(self, chunked_df, cfg, mel_transform):\n        self.df = chunked_df.reset_index(drop=True)\n        self.cfg = cfg\n        self.mel_transform = mel_transform\n\n    def __len__(self):\n        return len(self.df)\n\n    def __getitem__(self, idx):\n        row = self.df.iloc[idx]\n        audio, sr = torchaudio.load(row.filepath)\n        audio = audio.mean(dim=0).numpy()\n        if sr != self.cfg.FS:\n            audio = torchaudio.functional.resample(\n                torch.tensor(audio), orig_freq=sr, new_freq=self.cfg.FS\n            ).numpy()\n        chunk_audio = audio[row.start_sample:row.end_sample]\n        chunk_audio = prepare_audio(chunk_audio, int(self.cfg.TARGET_DURATION * self.cfg.FS))\n        mel = audio2melspec_gpu(chunk_audio, self.cfg, self.mel_transform)\n        mel = cv2.resize(mel, self.cfg.MEL_SHAPE[1:])\n        mel = cv2.cvtColor((mel * 255).astype(np.uint8), cv2.COLOR_GRAY2RGB)\n        mel = mel.transpose(2, 0, 1)\n        mel = mel.astype(np.float32) \n        mel_tensor = torch.tensor(mel, dtype=torch.float32)\n\n        return mel_tensor, row['row_id']\n\ndef vad_and_silence_filter(\n    chunked_df,\n    vad_model,\n    get_speech_timestamps,\n    device,\n    silence_threshold=0.01,\n    sampling_rate_col='sample_rate',\n    save_csv_path=None,\n    description=\"VAD and silence filter\"\n):\n    \"\"\"\n    Remove silent and human-voice chunks from the DataFrame using VAD and amplitude threshold.\n\n    Args:\n        chunked_df (pd.DataFrame): DataFrame with columns ['filepath', 'start_sample', 'end_sample', ...].\n        vad_model: Trained VAD PyTorch model.\n        get_speech_timestamps: Silero function for speech detection.\n        device: torch.device\n        silence_threshold (float): Mean amplitude below which chunk is considered silent.\n        sampling_rate_col (str): Column name for sample rate (if per-row), else use torchaudio.load default.\n        save_csv_path (str, optional): If provided, saves cleaned DataFrame to this path.\n        description (str): For logging.\n\n    Returns:\n        pd.DataFrame: Filtered DataFrame with only clean, non-silent chunks.\n    \"\"\"\n    vad_model.eval()\n    clean_chunks = []\n\n    print(f\"\\n== Cleaning chunks: Removing silent & human-voice segments ({description}) ==\")\n\n    for idx, row in tqdm(chunked_df.iterrows(), total=len(chunked_df)):\n        # Load audio\n        waveform, sr = torchaudio.load(row['filepath'])\n        # Use per-row sample_rate if available\n        if sampling_rate_col in row:\n            sr = int(row[sampling_rate_col])\n\n        start_sample = int(row['start_sample'])\n        end_sample = int(row['end_sample'])\n        chunk_audio = waveform[0, start_sample:end_sample].to(device)\n\n        # Silence check\n        if chunk_audio.abs().mean().item() < silence_threshold:\n            continue\n\n        # VAD (human voice) check\n        speech_timestamps = get_speech_timestamps(chunk_audio, vad_model, sampling_rate=sr)\n        if len(speech_timestamps) == 0:\n            clean_chunks.append(row)\n\n    clean_df = pd.DataFrame(clean_chunks)\n\n    if save_csv_path is not None:\n        clean_df.to_csv(save_csv_path, index=False)\n        print(f\"📁 Saved to {save_csv_path}\")\n\n    print(f\"✅ Clean and non-silent chunks count: {len(clean_df)}\")\n\n    return clean_df\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-17T00:21:18.211354Z","iopub.execute_input":"2025-06-17T00:21:18.211701Z","iopub.status.idle":"2025-06-17T00:21:18.223790Z","shell.execute_reply.started":"2025-06-17T00:21:18.211674Z","shell.execute_reply":"2025-06-17T00:21:18.222899Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def run_test_pipeline(\n    test_dir,\n    fallback_dir=None,\n    chunk_len_sec=5,\n    batch_size=4,\n    fallback_n=10,\n    n_print=2,\n    do_cleaning=True,\n    vad_model=None,\n    get_speech_timestamps=None,\n    silence_threshold=0.01,\n    cleaning_desc=\"VAD filtering\"\n):\n    print(\"== Getting test files ==\")\n    test_files = get_test_files(test_dir, fallback_dir, fallback_n)\n    print(f\"Found {len(test_files)} test files.\")\n    if len(test_files) == 0:\n        raise RuntimeError(\"No test files found.\")\n\n    print(\"\\n== Chunking test files ==\")\n    chunked_df = get_chunks_for_files(test_files, sample_rate=config.FS, chunk_len_sec=chunk_len_sec)\n    display(chunked_df.head())\n\n    # === CLEANING STEP ===\n    clean_df = chunked_df\n    if do_cleaning:\n        if vad_model is None or get_speech_timestamps is None:\n            raise ValueError(\"vad_model and get_speech_timestamps must be provided for cleaning.\")\n        print(\"\\n== Cleaning chunks: Removing silent & human-voice segments ==\")\n        clean_df = vad_and_silence_filter(\n            chunked_df,\n            vad_model,\n            get_speech_timestamps,\n            device=device,\n            silence_threshold=silence_threshold,\n            description=cleaning_desc\n        )\n\n    print(\"\\n== Initializing Dataset & Loader ==\")\n    dataset = BirdTestMelDataset(clean_df, config, mel_transform)\n    loader = DataLoader(dataset, batch_size=batch_size, shuffle=False)\n\n    print(f\"\\n== Dataset shape check ==\")\n    for i in range(min(n_print, len(dataset))):\n        mel, row_id = dataset[i]\n        print(f\"RowID: {row_id}\")\n        print(f\"  MEL shape: {mel.shape}\")\n        print(f\"  MEL dtype: {mel.dtype}\")\n        print(f\"  MEL min/max: {mel.min()} / {mel.max()}\")\n        print(f\"  MEL first 5 vals: {mel.flatten()[:5]}\")\n        print(\"-\" * 40)\n\n    print(\"\\n== Loader batch check ==\")\n    for batch in loader:\n        mels, row_ids = batch\n        print(f\"Batch MELS shape: {mels.shape}\")  # (B, 3, 256, 256)\n        print(f\"Batch MELS dtype: {mels.dtype}\")\n        print(f\"Batch first row_id: {row_ids[0]}\")\n        print(f\"Batch MELS min/max: {mels.min().item()} / {mels.max().item()}\")\n        break  # just first batch\n\n    print(\"All checks done.\")\n    # RETURN BOTH dfs!\n    return chunked_df, clean_df, dataset, loader\n\n# 3. Model load karo (bahar, function ke andar nahi)\nvad_model= load_vad_model('/kaggle/input/birdcleft-clean-and-vad-filtered-data/silero_vad.jit')  \n\n##============TEST RUN=====================\n# Make sure vad_model and get_speech_timestamps are defined, then pass to your pipeline:\n\n# 4. Pipeline run karo\nchunked_df ,clean_df, test_dataset, test_loader = run_test_pipeline(\n    test_dir=\"/kaggle/input/birdclef-2025/test_soundscapes\",\n    fallback_dir=\"/kaggle/input/birdclef-2025/train_soundscapes\",\n    chunk_len_sec=5,\n    batch_size=4,\n    fallback_n=10,\n    n_print=2,\n    do_cleaning=True,\n    vad_model=vad_model,  # Yahan pass karo\n    get_speech_timestamps=get_speech_timestamps,\n    silence_threshold=0.01,\n    cleaning_desc=\"Filtering test soundscapes\"\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-17T00:21:18.394458Z","iopub.execute_input":"2025-06-17T00:21:18.394775Z","iopub.status.idle":"2025-06-17T00:21:47.524142Z","shell.execute_reply.started":"2025-06-17T00:21:18.394751Z","shell.execute_reply":"2025-06-17T00:21:47.523225Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def plot_multiple_mels(dataset, indices, title_prefix='Sample Mel Spectrogram'):\n    n = len(indices)\n    fig, axes = plt.subplots(1, n, figsize=(4*n, 4))  # 1 row, n columns\n\n    for i, idx in enumerate(indices):\n        mel_tensor, _ = dataset[idx]\n        mel = mel_tensor[0].numpy()\n\n        ax = axes[i] if n > 1 else axes\n        im = ax.imshow(mel, aspect='auto', origin='lower')\n        ax.set_title(f\"{title_prefix} #{idx}\")\n        ax.set_xlabel(\"Time Frames\")\n        ax.set_ylabel(\"Mel Bands\")\n        ax.label_outer()  # Only show outer labels for clean look\n\n    fig.colorbar(im, ax=axes, orientation='vertical', fraction=0.02, pad=0.04)\n    plt.tight_layout()\n    plt.show()\n\n    # Example usage:\nprint(\"Test Samples:\")\nplot_multiple_mels(test_dataset, indices=list(range(5)), title_prefix='test Sample Mel')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-17T00:21:47.525459Z","iopub.execute_input":"2025-06-17T00:21:47.525953Z","iopub.status.idle":"2025-06-17T00:21:49.149739Z","shell.execute_reply.started":"2025-06-17T00:21:47.525929Z","shell.execute_reply":"2025-06-17T00:21:49.148647Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#=========INFERENCE CODE=============\ndef infer_model(model, \n                dataloader, \n                device, \n                master_labels, \n                save_preds_csv=None,\n                num_batches=None):\n    \"\"\"\n    Run inference on test dataloader, return DataFrame with predictions.\n    Assumes dataloader yields (mel_tensor, row_id).\n    \"\"\"\n    model.eval()\n    all_row_ids = []\n    all_probs = []\n\n    with torch.no_grad():\n        for batch_idx, batch in enumerate(tqdm(dataloader, desc=\"Inference\")):\n            # Unpack batch: (inputs, row_ids)\n            inputs, row_ids = batch\n            inputs = inputs.to(device)\n\n            outputs = model(inputs)\n            probs = torch.sigmoid(outputs).cpu().numpy()  # shape [B, num_classes]\n\n            all_row_ids.extend(list(row_ids))\n            all_probs.append(probs)\n\n            if num_batches and batch_idx + 1 >= num_batches:\n                break\n\n    all_probs = np.vstack(all_probs)  # shape: [num_samples, num_classes]\n    # Make sure DataFrame columns use master_labels:\n    df = pd.DataFrame(all_probs, columns=master_labels)\n    df.insert(0, \"row_id\", all_row_ids)\n    \n\n    if save_preds_csv:\n        df.to_csv(save_preds_csv, index=False)\n        print(f\"Saved predictions to: {save_preds_csv}\")\n\n    # Optional: Print top-5 for first few samples\n    print(\"\\nSample predictions:\")\n    for i in range(min(3, len(df))):\n        top5_idx = df.iloc[i, 1:].values.argsort()[-5:][::-1]\n        top5_probs = df.iloc[i, 1:].values[top5_idx]\n        print(f\"\\nRow: {df.loc[i, 'row_id']}\")\n        for rank, (cls_idx, prob) in enumerate(zip(top5_idx, top5_probs), 1):\n            print(f\"  {rank}. {master_labels[cls_idx]} ({prob:.3f})\")\n\n    return df","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-17T00:21:49.150929Z","iopub.execute_input":"2025-06-17T00:21:49.151207Z","iopub.status.idle":"2025-06-17T00:21:49.160855Z","shell.execute_reply.started":"2025-06-17T00:21:49.151186Z","shell.execute_reply":"2025-06-17T00:21:49.159943Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Model","metadata":{}},{"cell_type":"code","source":"# ===================== MODEL =====================\nclass ImprovedBirdCLEFModel(nn.Module):\n    def __init__(self, num_classes=187, backbone_weights=None, device=device):\n        super().__init__()\n        self.backbone = timm.create_model(\n            \"seresnext26t_32x4d\",\n            pretrained=False,\n            in_chans=3,\n            num_classes=0\n        )\n        if backbone_weights:\n            state_dict = torch.load(backbone_weights, map_location=device)\n            self.backbone.load_state_dict(state_dict, strict=False)\n        self.classifier = nn.Sequential(\n            nn.Linear(self.backbone.num_features, 512),\n            nn.BatchNorm1d(512),\n            nn.ReLU(inplace=True),\n            nn.Dropout(0.3),\n            nn.Linear(512, num_classes)\n        )\n    def forward(self, x):\n        x = self.backbone(x)\n        x = self.classifier(x)\n        return x\n\ndef load_model(model_path, device, num_classes, backbone_weights):\n    model = ImprovedBirdCLEFModel(num_classes=num_classes, backbone_weights=backbone_weights, device=device)\n    state_dict = torch.load(model_path, map_location=device)\n    model.load_state_dict(state_dict)\n    model = model.to(device)\n    model.eval()\n    return model\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-17T00:21:49.162548Z","iopub.execute_input":"2025-06-17T00:21:49.162843Z","iopub.status.idle":"2025-06-17T00:21:49.176720Z","shell.execute_reply.started":"2025-06-17T00:21:49.162822Z","shell.execute_reply":"2025-06-17T00:21:49.175754Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#=========================SUBMISSIONS CODE================\ndef create_submission(row_ids, predictions, master_labels, species_ids, submission_csv):\n    \"\"\"\n    row_ids: list of row_id strings\n    predictions: numpy array shape (N, 187) [model output]\n    master_labels: list of 187 model class labels (your cleaned data)\n    species_ids: list of 206 submission class labels (from sample_submission.csv)\n    sample_sub_csv: path to sample_submission.csv\n    \"\"\"\n    print(\"Creating submission dataframe...\")\n\n    # Mapping: for each submission class, find its index in model output, or -1 if missing\n    model2sub_idx = [master_labels.index(lbl) if lbl in master_labels else -1 for lbl in species_ids]\n\n    sub_dict = {'row_id': row_ids}\n    for i, lbl in enumerate(species_ids):\n        idx = model2sub_idx[i]\n        if idx != -1:\n            sub_dict[lbl] = predictions[:, idx]\n        else:\n            sub_dict[lbl] = np.zeros(len(row_ids), dtype=np.float32)\n\n    sub_df = pd.DataFrame(sub_dict)\n\n    # Ensure all columns (same order) as sample_submission\n    sample_sub = pd.read_csv(config.submission_csv)\n    for col in sample_sub.columns:\n        if col not in sub_df.columns:\n            sub_df[col] = 0.0\n    sub_df = sub_df[sample_sub.columns]\n    return sub_df\n\ndef smooth_submission(submission_path):\n    \"\"\"\n    submission_path: path to CSV file (will overwrite with smoothed)\n    \"\"\"\n    print(\"Smoothing submission predictions...\")\n    sub = pd.read_csv(submission_path)\n    cols = sub.columns[1:]  # all label columns\n    groups = sub['row_id'].str.rsplit('_', n=1).str[0]\n\n    for group, idx in sub.groupby(groups).groups.items():\n        arr = sub.loc[idx, cols].values\n        arr_sm = arr.copy()\n        if len(arr) > 1:\n            arr_sm[0] = 0.8 * arr[0] + 0.2 * arr[1]\n            arr_sm[-1] = 0.8 * arr[-1] + 0.2 * arr[-2]\n            for i in range(1, len(arr) - 1):\n                arr_sm[i] = 0.2 * arr[i-1] + 0.6 * arr[i] + 0.2 * arr[i+1]\n        sub.loc[idx, cols] = arr_sm\n\n    sub.to_csv(submission_path, index=False)\n    print(f\"Smoothed submission saved to {submission_path}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-17T00:21:49.177657Z","iopub.execute_input":"2025-06-17T00:21:49.177892Z","iopub.status.idle":"2025-06-17T00:21:49.193517Z","shell.execute_reply.started":"2025-06-17T00:21:49.177874Z","shell.execute_reply":"2025-06-17T00:21:49.192540Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"device = 'cuda' if torch.cuda.is_available() else 'cpu'","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-17T00:21:49.194450Z","iopub.execute_input":"2025-06-17T00:21:49.194741Z","iopub.status.idle":"2025-06-17T00:21:49.215188Z","shell.execute_reply.started":"2025-06-17T00:21:49.194711Z","shell.execute_reply":"2025-06-17T00:21:49.214204Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if __name__ == \"__main__\":\n    # 1. Load VAD model & utility\n    vad_model = load_vad_model('/kaggle/input/birdcleft-clean-and-vad-filtered-data/silero_vad.jit')  \n    print(\"VAD loaded\")\n\n    # 2. Prepare test chunks & loader (return BOTH dfs!)\n    chunked_df, clean_df, test_dataset, test_loader = run_test_pipeline(\n        test_dir=\"/kaggle/input/birdclef-2025/test_soundscapes\",\n        fallback_dir=\"/kaggle/input/birdclef-2025/train_soundscapes\",\n        chunk_len_sec=5,\n        batch_size=32,\n        fallback_n=10,\n        n_print=2,\n        do_cleaning=True,\n        vad_model=vad_model,\n        get_speech_timestamps=get_speech_timestamps,\n        silence_threshold=0.01,\n        cleaning_desc=\"Filtering test soundscapes\"\n    )\n\n    # 3. Load model\n    model = load_model(\n        model_path=config.model_path,\n        device=device,\n        num_classes=187,\n        backbone_weights=config.backbone_weights\n    )\n    print(\"✅ Model loaded!\")\n\n    # 4. Load label lists\n    master_labels = load_master_labels(config.master_labels)        # 187 model classes\n    species_ids = load_species_ids(config.submission_csv)           # 206 submission columns\n\n    # 5. Run inference on only clean chunks\n    test_preds_df = infer_model(\n        model=model,\n        dataloader=test_loader,\n        device=config.DEVICE,\n        master_labels=master_labels,\n        save_preds_csv=\"test_preds.csv\",\n        num_batches=None\n    )\n    # test_preds_df columns: [\"row_id\"] + master_labels (187)\n\n    # 6. Zero-fill for all chunks (incl. filtered ones)\n    all_row_ids = chunked_df[\"row_id\"].tolist()  # <-- use original chunked_df for all 120 row_ids qk submission me subki prediction chhaye chahy zero hi ho\n    full_preds_df = pd.DataFrame({'row_id': all_row_ids}).merge(\n        test_preds_df, on='row_id', how='left'\n    )\n    full_preds_df = full_preds_df.fillna(0.0)\n\n    # 7. Create & save submission\n    predictions = full_preds_df[master_labels].values  # shape: (N, 187)\n    row_ids = full_preds_df[\"row_id\"].tolist()\n    submission_df = create_submission(\n        row_ids=row_ids,\n        predictions=predictions,\n        master_labels=master_labels,\n        species_ids=species_ids,\n        submission_csv=\"sample_submission.csv\"\n    )\n    submission_df.to_csv('submission.csv', index=False)\n    print(\"Submission file saved as submission.csv\")\n\n    # 8. (Optional) BirdCLEF smoothing\n    smooth_submission('submission.csv')\n    print(\"Smoothed submission saved.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-17T00:21:49.216279Z","iopub.execute_input":"2025-06-17T00:21:49.216699Z","iopub.status.idle":"2025-06-17T00:22:21.759597Z","shell.execute_reply.started":"2025-06-17T00:21:49.216675Z","shell.execute_reply":"2025-06-17T00:22:21.758643Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"pd.read_csv(\"submission.csv\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-17T00:22:44.434291Z","iopub.execute_input":"2025-06-17T00:22:44.435093Z","iopub.status.idle":"2025-06-17T00:22:44.468981Z","shell.execute_reply.started":"2025-06-17T00:22:44.435062Z","shell.execute_reply":"2025-06-17T00:22:44.468137Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## ✅ Conclusion\n\n- 🎯 Successfully ran inference on all test chunks after cleaning and preprocessing  \n- 🔢 Processed **120 chunks**, each with mel spectrograms and predicted species probabilities  \n- 📦 Created `submission.csv` with smoothed outputs for leaderboard evaluation  \n- 🐦 Model confidently predicted bird species from environmental soundscapes\n\nThis notebook completes the final phase of the BirdCLEF 2025 pipeline following:\n- Data preparation  \n- Mel spectrogram generation  \n","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}