{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.12.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceType":"competition","sourceId":59093,"databundleVersionId":7469972},{"sourceType":"datasetVersion","sourceId":7447509,"datasetId":4334995,"databundleVersionId":7539394},{"sourceType":"datasetVersion","sourceId":7450712,"datasetId":4336944,"databundleVersionId":7542629},{"sourceType":"datasetVersion","sourceId":7403069,"datasetId":4304949,"databundleVersionId":7494186},{"sourceType":"datasetVersion","sourceId":7392775,"datasetId":4297782,"databundleVersionId":7483780},{"sourceType":"datasetVersion","sourceId":7392733,"datasetId":4297749,"databundleVersionId":7483738},{"sourceType":"datasetVersion","sourceId":7402356,"datasetId":4304475,"databundleVersionId":7493457},{"sourceType":"kernelVersion","sourceId":158958765}],"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Klasifikasi Aktivitas Otak Berbahaya dengan EfficientNetV2 (v5)\n\n## Perubahan dari v4\n\n- ✅ **Save snapshot per-fold** — setiap fold selesai langsung disimpan\n- ✅ **Print verbose hasil per fold** — bisa direkonstruksi dari Versions log\n- ✅ **Auto-warning di akhir** untuk reminder save version\n\n## ⚠️ KRITIS: Cara Pakai yang AMAN di Kaggle\n\n**Sebelum Klik Save Version:**\n1. Pastikan Notebook Settings → **Save Output** = ON\n2. Pastikan **Internet** = ON\n3. Pilih **Accelerator** = GPU T4 x2\n\n**Setelah Run Selesai:**\n1. JANGAN tutup browser/tab Kaggle dulu\n2. Cek tab **Output** di sidebar kanan, pastikan ada file `.pkl`, `.json`, `.h5`, `.png`\n3. **DOWNLOAD SEMUA FILE** ke lokal SEBELUM tutup notebook\n4. Sebagai backup: **Save Version** lagi (commit) supaya output tersimpan ke Versions\n\n## Cara Pakai\n\n1. Ubah `MODEL_VARIANT` dan `USE_AUGMENTATION` di Cell 1\n2. Klik **Save Version** → **Save & Run All (Commit)**\n3. Tunggu selesai, **JANGAN TUTUP BROWSER**\n4. Download semua file di tab Output","metadata":{}},{"cell_type":"markdown","source":"## Cell 1: Konfigurasi Eksperimen\n\n⚠️ **UBAH BAGIAN INI UNTUK TIAP SKENARIO**","metadata":{}},{"cell_type":"code","source":"# ===================================================================\n# KONFIGURASI EKSPERIMEN\n# ===================================================================\nMODEL_VARIANT = 'B0'        # Pilih: 'B0', 'B1', atau 'B2'\nUSE_AUGMENTATION = True    # True = Augmented, False = Baseline\n# ===================================================================\n\n# Quick test mode — set True untuk sanity check (hanya 1 fold, 3 epoch)\nQUICK_TEST = False\n\n# Otomatis — jangan diubah\nSCENARIO_NAME = f\"{'Augmented' if USE_AUGMENTATION else 'Baseline'}_{MODEL_VARIANT}\"\nprint(f'SKENARIO: {SCENARIO_NAME}')\nprint(f'MODEL_VARIANT: {MODEL_VARIANT}')\nprint(f'USE_AUGMENTATION: {USE_AUGMENTATION}')\nprint(f'QUICK_TEST: {QUICK_TEST}')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-22T11:31:10.402230Z","iopub.execute_input":"2026-04-22T11:31:10.402855Z","iopub.status.idle":"2026-04-22T11:31:10.416617Z","shell.execute_reply.started":"2026-04-22T11:31:10.402824Z","shell.execute_reply":"2026-04-22T11:31:10.415894Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Cell 2: Import & Setup","metadata":{}},{"cell_type":"code","source":"import os, gc, math, pickle, json\nfrom datetime import datetime\n\nos.environ['CUDA_VISIBLE_DEVICES'] = '0,1'\n\nimport tensorflow as tf\nimport pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport albumentations as albu\n\nprint(f'TensorFlow version: {tf.__version__}')\nprint(f'Albumentations version: {albu.__version__}')\n\n# Setup GPU strategy\ngpus = tf.config.list_physical_devices('GPU')\nif len(gpus) <= 1:\n    strategy = tf.distribute.OneDeviceStrategy(device='/gpu:0')\n    print(f'Using {len(gpus)} GPU')\nelse:\n    strategy = tf.distribute.MirroredStrategy()\n    print(f'Using {len(gpus)} GPUs')\n\n# Mixed precision\ntf.config.optimizer.set_experimental_options({'auto_mixed_precision': True})\nprint('Mixed precision enabled')\n\n# Global flags\nUSE_KAGGLE_SPECTROGRAMS = True\nUSE_EEG_SPECTROGRAMS = True\nLOAD_MODELS_FROM = None","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-22T11:31:10.418224Z","iopub.execute_input":"2026-04-22T11:31:10.418582Z","iopub.status.idle":"2026-04-22T11:31:46.133937Z","shell.execute_reply.started":"2026-04-22T11:31:10.418540Z","shell.execute_reply":"2026-04-22T11:31:46.133126Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Cell 3: Load Data Train CSV","metadata":{}},{"cell_type":"code","source":"df = pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/train.csv')\nTARGETS = df.columns[-6:]\nprint(f'Train shape: {df.shape}')\nprint(f'Targets: {list(TARGETS)}')\ndf.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-22T11:31:46.135030Z","iopub.execute_input":"2026-04-22T11:31:46.135631Z","iopub.status.idle":"2026-04-22T11:31:46.405512Z","shell.execute_reply.started":"2026-04-22T11:31:46.135604Z","shell.execute_reply":"2026-04-22T11:31:46.404650Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Cell 4: Aggregate per EEG ID\n\nKonversi ke format 1 baris per EEG ID dengan label berupa distribusi probabilitas hasil voting expert.","metadata":{}},{"cell_type":"code","source":"train = df.groupby('eeg_id')[['spectrogram_id', 'spectrogram_label_offset_seconds']].agg(\n    {'spectrogram_id': 'first', 'spectrogram_label_offset_seconds': 'min'})\ntrain.columns = ['spec_id', 'min']\n\ntmp = df.groupby('eeg_id')[['spectrogram_id', 'spectrogram_label_offset_seconds']].agg(\n    {'spectrogram_label_offset_seconds': 'max'})\ntrain['max'] = tmp\n\ntmp = df.groupby('eeg_id')[['patient_id']].agg('first')\ntrain['patient_id'] = tmp\n\ntmp = df.groupby('eeg_id')[TARGETS].agg('sum')\nfor t in TARGETS:\n    train[t] = tmp[t].values\n\n# Normalisasi label ke distribusi probabilitas (soft labels)\ny_data = train[TARGETS].values\ny_data = y_data / y_data.sum(axis=1, keepdims=True)\ntrain[TARGETS] = y_data\n\ntmp = df.groupby('eeg_id')[['expert_consensus']].agg('first')\ntrain['target'] = tmp\n\ntrain = train.reset_index()\nprint(f'Train non-overlap eeg_id shape: {train.shape}')\nprint(f'Jumlah unique patient: {train.patient_id.nunique()}')\ntrain.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-22T11:31:46.406638Z","iopub.execute_input":"2026-04-22T11:31:46.407156Z","iopub.status.idle":"2026-04-22T11:31:46.482356Z","shell.execute_reply.started":"2026-04-22T11:31:46.407125Z","shell.execute_reply":"2026-04-22T11:31:46.481733Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Cell 5: Load Spectrograms","metadata":{}},{"cell_type":"code","source":"%%time\nspectrograms = np.load(\n    '/kaggle/input/brain-spectrograms/specs.npy',\n    allow_pickle=True\n).item()\nprint(f'Jumlah spectrograms: {len(spectrograms)}')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-22T11:31:46.484458Z","iopub.execute_input":"2026-04-22T11:31:46.485038Z","iopub.status.idle":"2026-04-22T11:32:37.581703Z","shell.execute_reply.started":"2026-04-22T11:31:46.485008Z","shell.execute_reply":"2026-04-22T11:32:37.580977Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"%%time\nall_eegs = np.load(\n    '/kaggle/input/brain-eeg-spectrograms/eeg_specs.npy',\n    allow_pickle=True\n).item()\nprint(f'Jumlah EEG spectrograms: {len(all_eegs)}')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-22T11:32:37.582640Z","iopub.execute_input":"2026-04-22T11:32:37.582892Z","iopub.status.idle":"2026-04-22T11:33:27.817038Z","shell.execute_reply.started":"2026-04-22T11:32:37.582867Z","shell.execute_reply":"2026-04-22T11:33:27.816262Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Cell 6: Data Generator dengan Augmentasi Lengkap\n\nAugmentasi sesuai Tabel 3.3 (revisi):\n- **HorizontalFlip** (p=0.5)\n- **XYMasking** dengan proporsi disesuaikan SpecAugment best practice:\n  - num_masks_x, num_masks_y: (1, 2)\n  - mask_x_length (time): (10, 40) → ~15% dari 256 pixel\n  - mask_y_length (freq): (10, 30) → ~23% dari 128 pixel\n- **GaussNoise**: var_limit (5, 25), disesuaikan dengan skala signal ternormalisasi","metadata":{}},{"cell_type":"code","source":"TARS = {'Seizure': 0, 'LPD': 1, 'GPD': 2, 'LRDA': 3, 'GRDA': 4, 'Other': 5}\nTARS2 = {x: y for y, x in TARS.items()}\n\nclass DataGenerator(tf.keras.utils.Sequence):\n    '''Generator untuk Keras dengan augmentasi opsional'''\n    \n    def __init__(self, data, batch_size=32, shuffle=False, augment=False, mode='train',\n                 specs=None, eeg_specs=None):\n        self.data = data\n        self.batch_size = batch_size\n        self.shuffle = shuffle\n        self.augment = augment\n        self.mode = mode\n        self.specs = specs\n        self.eeg_specs = eeg_specs\n        self.on_epoch_end()\n    \n    def __len__(self):\n        return int(np.ceil(len(self.data) / self.batch_size))\n    \n    def __getitem__(self, index):\n        indexes = self.indexes[index*self.batch_size:(index+1)*self.batch_size]\n        X, y = self.__data_generation(indexes)\n        if self.augment:\n            X = self.__augment_batch(X)\n        return X, y\n    \n    def on_epoch_end(self):\n        self.indexes = np.arange(len(self.data))\n        if self.shuffle:\n            np.random.shuffle(self.indexes)\n    \n    def __data_generation(self, indexes):\n        X = np.zeros((len(indexes), 128, 256, 8), dtype='float32')\n        y = np.zeros((len(indexes), 6), dtype='float32')\n        img = np.ones((128, 256), dtype='float32')\n        \n        for j, i in enumerate(indexes):\n            row = self.data.iloc[i]\n            if self.mode == 'test':\n                r = 0\n            else:\n                r = int((row['min'] + row['max']) // 4)\n            \n            for k in range(4):\n                img = self.specs[row.spec_id][r:r+300, k*100:(k+1)*100].T\n                \n                # LOG TRANSFORM\n                img = np.clip(img, np.exp(-4), np.exp(8))\n                img = np.log(img)\n                \n                # NORMALIZATION\n                ep = 1e-6\n                m = np.nanmean(img.flatten())\n                s = np.nanstd(img.flatten())\n                img = (img - m) / (s + ep)\n                img = np.nan_to_num(img, nan=0.0)\n                \n                # CROP\n                X[j, 14:-14, :, k] = img[:, 22:-22] / 2.0\n            \n            img = self.eeg_specs[row.eeg_id]\n            X[j, :, :, 4:] = img\n            \n            if self.mode != 'test':\n                y[j,] = row[TARGETS]\n        \n        return X, y\n    \n    def __random_transform(self, img):\n        '''\n        Augmentasi sesuai Tabel 3.3 (revisi):\n        - HorizontalFlip (p=0.5)\n        - XYMasking: num_masks (1-2), \n          mask_x_length (10-40) untuk waktu (dim=256),\n          mask_y_length (10-30) untuk frekuensi (dim=128)\n        - GaussNoise: var_limit (5, 25), mean=0, p=0.5\n        \n        Proporsi masking disesuaikan dengan best practice SpecAugment:\n        ~15-25% dari total dimensi agar tidak merusak informasi signal.\n        '''\n        composition = albu.Compose([\n            albu.HorizontalFlip(p=0.5),\n            albu.XYMasking(\n                num_masks_x=(1, 2),\n                num_masks_y=(1, 2),\n                mask_x_length=(10, 40),\n                mask_y_length=(10, 30),\n                fill_value=0,\n                p=0.5\n            ),\n            albu.GaussNoise(\n                var_limit=(5.0, 25.0),\n                mean=0,\n                p=0.5\n            ),\n        ])\n        return composition(image=img)['image']\n    \n    def __augment_batch(self, img_batch):\n        for i in range(img_batch.shape[0]):\n            img_batch[i] = self.__random_transform(img_batch[i])\n        return img_batch\n\nprint('DataGenerator siap')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-22T11:33:27.818143Z","iopub.execute_input":"2026-04-22T11:33:27.818566Z","iopub.status.idle":"2026-04-22T11:33:27.835641Z","shell.execute_reply.started":"2026-04-22T11:33:27.818525Z","shell.execute_reply":"2026-04-22T11:33:27.834791Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Cell 7: Learning Rate Scheduler — Single-Stage Cosine Decay\n\nStrategi single-stage cosine decay selama 10 epoch:\n- LR Max: 1×10⁻³\n- LR Min: 1×10⁻⁶\n- Decay: smooth cosine sepanjang 10 epoch\n\nTidak ada warm-up atau step decay. LR langsung dimulai tinggi lalu turun mulus mengikuti kurva cosine.","metadata":{}},{"cell_type":"code","source":"LR_MAX = 1e-3\nLR_MIN = 1e-6\nEPOCHS = 10\n\ndef lrfn(epoch):\n    '''Cosine decay dari LR_MAX ke LR_MIN selama EPOCHS'''\n    decay_total_epochs = EPOCHS - 1\n    phase = math.pi * epoch / decay_total_epochs\n    cosine_decay = 0.5 * (1 + math.cos(phase))\n    lr = (LR_MAX - LR_MIN) * cosine_decay + LR_MIN\n    return lr\n\nLR_scheduler = tf.keras.callbacks.LearningRateScheduler(lrfn, verbose=1)\n\n# Visualisasi\nrng = [i for i in range(EPOCHS)]\nlr_values = [lrfn(x) for x in rng]\n\nplt.figure(figsize=(10, 4))\nplt.plot(rng, lr_values, 'o-', color='#1D9E75', linewidth=2, markersize=6)\nplt.xlabel('Epoch', fontsize=12)\nplt.ylabel('Learning Rate', fontsize=12)\nplt.title(f'Cosine Decay LR Schedule ({EPOCHS} epochs)', fontsize=14)\nplt.yscale('log')\nplt.grid(alpha=0.3)\n\n# Annotasi nilai di titik penting\nfor idx in [0, EPOCHS//2, EPOCHS-1]:\n    plt.annotate(f'{lr_values[idx]:.1e}',\n                 xy=(idx, lr_values[idx]),\n                 xytext=(5, 5), textcoords='offset points',\n                 fontsize=9)\n\nplt.tight_layout()\nplt.show()\n\nprint(f'Total epochs: {EPOCHS}')\nprint(f'LR awal: {lr_values[0]:.1e}, LR akhir: {lr_values[-1]:.1e}')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-22T11:33:27.836728Z","iopub.execute_input":"2026-04-22T11:33:27.837488Z","iopub.status.idle":"2026-04-22T11:33:28.549927Z","shell.execute_reply.started":"2026-04-22T11:33:27.837462Z","shell.execute_reply":"2026-04-22T11:33:28.549157Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Cell 8: Build Model EfficientNetV2\n\nMenggunakan `tf.keras.applications.EfficientNetV2B0/B1/B2` (built-in TensorFlow).","metadata":{}},{"cell_type":"code","source":"def build_model():\n    inp = tf.keras.Input(shape=(128, 256, 8))\n    \n    # Pilih varian EfficientNetV2\n    if MODEL_VARIANT == 'B0':\n        base_model = tf.keras.applications.EfficientNetV2B0(\n            include_top=False, weights='imagenet', input_shape=(512, 512, 3)\n        )\n    elif MODEL_VARIANT == 'B1':\n        base_model = tf.keras.applications.EfficientNetV2B1(\n            include_top=False, weights='imagenet', input_shape=(512, 512, 3)\n        )\n    elif MODEL_VARIANT == 'B2':\n        base_model = tf.keras.applications.EfficientNetV2B2(\n            include_top=False, weights='imagenet', input_shape=(512, 512, 3)\n        )\n    else:\n        raise ValueError(f'MODEL_VARIANT harus B0, B1, atau B2. Got: {MODEL_VARIANT}')\n    \n    # Reshape 128x256x8 → 512x512x3\n    x1 = [inp[:, :, :, i:i+1] for i in range(4)]\n    x1 = tf.keras.layers.Concatenate(axis=1)(x1)\n    \n    x2 = [inp[:, :, :, i+4:i+5] for i in range(4)]\n    x2 = tf.keras.layers.Concatenate(axis=1)(x2)\n    \n    if USE_KAGGLE_SPECTROGRAMS & USE_EEG_SPECTROGRAMS:\n        x = tf.keras.layers.Concatenate(axis=2)([x1, x2])\n    elif USE_EEG_SPECTROGRAMS:\n        x = x2\n    else:\n        x = x1\n    \n    x = tf.keras.layers.Concatenate(axis=3)([x, x, x])\n    \n    x = base_model(x)\n    x = tf.keras.layers.GlobalAveragePooling2D()(x)\n    x = tf.keras.layers.Dense(6, activation='softmax', dtype='float32')(x)\n    \n    model = tf.keras.Model(inputs=inp, outputs=x)\n    opt = tf.keras.optimizers.Adam(learning_rate=LR_MAX)\n    loss = tf.keras.losses.KLDivergence()\n    model.compile(loss=loss, optimizer=opt)\n    \n    return model\n\n# Sanity check\nwith strategy.scope():\n    _test_model = build_model()\n    total_params = _test_model.count_params()\n    print(f'Model: EfficientNetV2-{MODEL_VARIANT}')\n    print(f'Total parameters: {total_params:,}')\n    del _test_model\n    gc.collect()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-22T11:33:28.551004Z","iopub.execute_input":"2026-04-22T11:33:28.551379Z","iopub.status.idle":"2026-04-22T11:33:34.134816Z","shell.execute_reply.started":"2026-04-22T11:33:28.551351Z","shell.execute_reply":"2026-04-22T11:33:34.133957Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Cell 9: Training Loop dengan ModelCheckpoint\n\n**Perbaikan penting**:\n- Single-stage cosine decay (10 epoch)\n- **ModelCheckpoint**: simpan weights dengan val loss terendah (sesuai Tabel 3.5)\n- Setelah training, load kembali best weights sebelum OOF predictions","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import GroupKFold\nimport tensorflow.keras.backend as K\nimport time\nimport pickle as pkl_lib\n\nall_oof = []\nall_true = []\nall_histories = []\nall_fold_times = []\n\ngkf = GroupKFold(n_splits=5)\ngkf_splits = list(gkf.split(train, train.target, train.patient_id))\n\nif QUICK_TEST:\n    print('⚠️  QUICK TEST MODE: hanya 1 fold, 3 epoch')\n    gkf_splits = gkf_splits[:1]\n    EPOCHS_ACTUAL = 3\nelse:\n    EPOCHS_ACTUAL = EPOCHS\n\nclass TimingCallback(tf.keras.callbacks.Callback):\n    def __init__(self):\n        super().__init__()\n        self.epoch_times = []\n    def on_epoch_begin(self, epoch, logs=None):\n        self.epoch_start = time.time()\n    def on_epoch_end(self, epoch, logs=None):\n        self.epoch_times.append(time.time() - self.epoch_start)\n\ntotal_start_time = time.time()\n\nfor i, (train_index, valid_index) in enumerate(gkf_splits):\n    print('\\n' + '#' * 70)\n    print(f'### FOLD {i+1}/{len(gkf_splits)} | Skenario: {SCENARIO_NAME}')\n    print(f'### train size: {len(train_index)} | valid size: {len(valid_index)}')\n    print('#' * 70)\n    \n    train_gen = DataGenerator(\n        train.iloc[train_index],\n        shuffle=True,\n        batch_size=32,\n        augment=USE_AUGMENTATION,\n        specs=spectrograms,\n        eeg_specs=all_eegs\n    )\n    \n    valid_gen = DataGenerator(\n        train.iloc[valid_index],\n        shuffle=False,\n        batch_size=64,\n        mode='valid',\n        augment=False,\n        specs=spectrograms,\n        eeg_specs=all_eegs\n    )\n    \n    K.clear_session()\n    with strategy.scope():\n        model = build_model()\n    \n    if LOAD_MODELS_FROM is None:\n        weight_name = f'EffNetV2{MODEL_VARIANT}_aug{int(USE_AUGMENTATION)}_f{i}.weights.h5'\n        checkpoint = tf.keras.callbacks.ModelCheckpoint(\n            filepath=weight_name,\n            monitor='val_loss',\n            mode='min',\n            save_best_only=True,\n            save_weights_only=True,\n            verbose=1\n        )\n        timing_cb = TimingCallback()\n        \n        fold_start = time.time()\n        \n        print(f'\\n--- Training {EPOCHS_ACTUAL} epochs dengan Cosine Decay ---')\n        history = model.fit(\n            train_gen,\n            verbose=1,\n            validation_data=valid_gen,\n            epochs=EPOCHS_ACTUAL,\n            callbacks=[LR_scheduler, checkpoint, timing_cb]\n        )\n        \n        fold_duration = time.time() - fold_start\n        avg_epoch_time = sum(timing_cb.epoch_times) / len(timing_cb.epoch_times)\n        \n        all_fold_times.append({\n            'fold_duration_sec': fold_duration,\n            'avg_epoch_sec': avg_epoch_time,\n            'epoch_times': timing_cb.epoch_times\n        })\n        \n        print(f'\\n⏱  Fold {i+1} duration: {fold_duration/60:.2f} menit')\n        print(f'⏱  Avg epoch time: {avg_epoch_time:.1f} detik')\n        \n        all_histories.append({\n            'loss': history.history['loss'],\n            'val_loss': history.history['val_loss'],\n            'lr': history.history.get('lr', [])\n        })\n        \n        print(f'\\n✓ Loading best weights: {weight_name}')\n        model.load_weights(weight_name)\n    else:\n        weight_name = f'{LOAD_MODELS_FROM}EffNetV2{MODEL_VARIANT}_aug{int(USE_AUGMENTATION)}_f{i}.weights.h5'\n        model.load_weights(weight_name)\n    \n    oof = model.predict(valid_gen, verbose=1)\n    all_oof.append(oof)\n    all_true.append(train.iloc[valid_index][TARGETS].values)\n    \n    # ===== SNAPSHOT PARTIAL HASIL setelah setiap fold =====\n    snapshot = {\n        'scenario_name': SCENARIO_NAME,\n        'completed_folds': i + 1,\n        'all_oof': np.concatenate(all_oof),\n        'all_true': np.concatenate(all_true),\n        'all_histories': all_histories,\n        'all_fold_times': all_fold_times,\n    }\n    snapshot_name = f'snapshot_{SCENARIO_NAME}_after_fold{i+1}.pkl'\n    with open(snapshot_name, 'wb') as f:\n        pkl_lib.dump(snapshot, f)\n    print(f'💾 Snapshot disimpan: {snapshot_name}')\n    \n    # ===== PRINT VERBOSE INFO supaya bisa direkonstruksi dari log =====\n    best_val = min(history.history['val_loss'])\n    final_train = history.history['loss'][-1]\n    final_val = history.history['val_loss'][-1]\n    print(f'\\n📊 RINGKASAN FOLD {i+1}:')\n    print(f'   Best val loss: {best_val:.6f}')\n    print(f'   Final train loss: {final_train:.6f}')\n    print(f'   Final val loss: {final_val:.6f}')\n    print(f'   Train losses per epoch: {[round(x, 6) for x in history.history[\"loss\"]]}')\n    print(f'   Val losses per epoch: {[round(x, 6) for x in history.history[\"val_loss\"]]}')\n    print(f'   Epoch times (sec): {[round(x, 1) for x in timing_cb.epoch_times]}')\n    \n    del model, oof\n    gc.collect()\n    K.clear_session()\n\ntotal_duration = time.time() - total_start_time\n\nall_oof = np.concatenate(all_oof)\nall_true = np.concatenate(all_true)\n\nprint(f'\\n\\n=== Training Selesai ===')\nprint(f'Total OOF predictions: {all_oof.shape}')\nprint(f'Total true labels: {all_true.shape}')\nprint(f'⏱  TOTAL TRAINING TIME: {total_duration/60:.2f} menit ({total_duration/3600:.2f} jam)')\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-04-22T11:33:34.135905Z","iopub.execute_input":"2026-04-22T11:33:34.136467Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Cell 10: Hitung CV KL Divergence","metadata":{}},{"cell_type":"code","source":"import sys\nsys.path.append('/kaggle/input/kaggle-kl-div')\nfrom kaggle_kl_div import score\n\noof_df = pd.DataFrame(all_oof.copy())\noof_df['id'] = np.arange(len(oof_df))\n\ntrue_df = pd.DataFrame(all_true.copy())\ntrue_df['id'] = np.arange(len(true_df))\n\ncv_score = score(solution=true_df, submission=oof_df, row_id_column_name='id')\nprint(f'\\n{\"=\"*60}')\nprint(f'SKENARIO: {SCENARIO_NAME}')\nprint(f'CV KL-Divergence: {cv_score:.4f}')\nprint(f'{\"=\"*60}')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Cell 11: Visualisasi Training Curves","metadata":{}},{"cell_type":"code","source":"n_folds = len(all_histories)\nfig, axes = plt.subplots(1, n_folds, figsize=(4.5*n_folds, 4))\n\nif n_folds == 1:\n    axes = [axes]\n\nfor i, hist in enumerate(all_histories):\n    ax = axes[i]\n    epochs_total = len(hist['loss'])\n    best_val_epoch = int(np.argmin(hist['val_loss']))\n    best_val_loss = min(hist['val_loss'])\n    \n    ax.plot(range(epochs_total), hist['loss'], 'b-o', label='Train Loss', markersize=5)\n    ax.plot(range(epochs_total), hist['val_loss'], 'r-o', label='Val Loss', markersize=5)\n    ax.axvline(x=best_val_epoch, color='green', linestyle=':', alpha=0.6, label=f'Best val (epoch {best_val_epoch})')\n    \n    ax.set_title(f'Fold {i+1} | Best Val: {best_val_loss:.4f}')\n    ax.set_xlabel('Epoch')\n    ax.set_ylabel('KL Divergence Loss')\n    ax.legend(fontsize=8)\n    ax.grid(alpha=0.3)\n\nplt.suptitle(f'Training Curves — {SCENARIO_NAME}', fontsize=14, y=1.02)\nplt.tight_layout()\nplt.savefig(f'training_curves_{SCENARIO_NAME}.png', dpi=150, bbox_inches='tight')\nplt.show()\n\n# Ringkasan per fold\nprint('\\nRingkasan per fold:')\nprint(f'{\"Fold\":<8}{\"Best Val\":<15}{\"Best Epoch\":<14}{\"Final Train\":<16}{\"Final Val\":<14}{\"Gap\":<10}')\nprint('-' * 77)\nfor i, hist in enumerate(all_histories):\n    best_val = min(hist['val_loss'])\n    best_epoch = int(np.argmin(hist['val_loss']))\n    final_train = hist['loss'][-1]\n    final_val = hist['val_loss'][-1]\n    gap = final_val - final_train\n    print(f'{i+1:<8}{best_val:<15.4f}{best_epoch:<14}{final_train:<16.4f}{final_val:<14.4f}{gap:<10.4f}')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Cell 12: Simpan Hasil Eksperimen untuk Bab 4","metadata":{}},{"cell_type":"code","source":"timestamp = datetime.now().strftime('%Y%m%d_%H%M')\nresult_filename = f'results_{SCENARIO_NAME}_{timestamp}.pkl'\n\n# Hitung statistik timing\nimport numpy as np\nfold_durations = [t['fold_duration_sec'] for t in all_fold_times]\nepoch_times_all = [t['avg_epoch_sec'] for t in all_fold_times]\n\nresults = {\n    'scenario_name': SCENARIO_NAME,\n    'model_variant': MODEL_VARIANT,\n    'use_augmentation': USE_AUGMENTATION,\n    'cv_score': float(cv_score),\n    'all_oof': all_oof,\n    'all_true': all_true,\n    'all_histories': all_histories,\n    'all_fold_times': all_fold_times,  # NEW\n    'total_duration_sec': total_duration,  # NEW\n    'config': {\n        'batch_size_train': 32,\n        'batch_size_valid': 64,\n        'epochs': EPOCHS_ACTUAL,\n        'lr_schedule': 'Single-stage Cosine Decay',\n        'lr_max': LR_MAX,\n        'lr_min': LR_MIN,\n        'optimizer': 'Adam',\n        'loss': 'KLDivergence',\n        'n_folds': 5,\n        'cv_type': 'GroupKFold by patient_id',\n        'model_checkpoint': 'monitor=val_loss, mode=min, save_best_only=True',\n    },\n    'timestamp': timestamp,\n}\n\nwith open(result_filename, 'wb') as f:\n    pickle.dump(results, f)\n\nprint(f'✓ Hasil disimpan: {result_filename}')\n\nsummary = {\n    'scenario_name': SCENARIO_NAME,\n    'model_variant': MODEL_VARIANT,\n    'use_augmentation': USE_AUGMENTATION,\n    'cv_score': float(cv_score),\n    'per_fold_best_val_loss': [float(min(h['val_loss'])) for h in all_histories],\n    'per_fold_best_epoch': [int(np.argmin(h['val_loss'])) for h in all_histories],\n    'per_fold_final_val_loss': [float(h['val_loss'][-1]) for h in all_histories],\n    'per_fold_final_train_loss': [float(h['loss'][-1]) for h in all_histories],\n    # NEW: timing info\n    'per_fold_duration_sec': [float(d) for d in fold_durations],\n    'per_fold_avg_epoch_sec': [float(t) for t in epoch_times_all],\n    'total_duration_sec': float(total_duration),\n    'total_duration_min': float(total_duration / 60),\n    'mean_epoch_time_sec': float(np.mean(epoch_times_all)),\n    'timestamp': timestamp,\n}\n\nsummary_filename = f'summary_{SCENARIO_NAME}_{timestamp}.json'\nwith open(summary_filename, 'w') as f:\n    json.dump(summary, f, indent=2)\n\nprint(f'✓ Ringkasan disimpan: {summary_filename}')\nprint(f'\\n=== EKSPERIMEN SELESAI: {SCENARIO_NAME} | CV KL = {cv_score:.4f} ===')\nprint(f'⏱  Total: {total_duration/60:.2f} menit | Avg epoch: {np.mean(epoch_times_all):.1f} detik')\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## ⚠️ PENTING: Sebelum Tutup Notebook\n\n**JANGAN TUTUP TAB BROWSER SEBELUM:**\n\n1. ✅ Cek tab **Output** di sidebar kanan\n2. ✅ Pastikan ada file:\n   - `summary_*.json`\n   - `results_*.pkl`\n   - `training_curves_*.png`\n   - `EffNetV2*_f0.weights.h5` sampai `f4.weights.h5` (5 file)\n   - `snapshot_*.pkl` (5 file backup per fold)\n3. ✅ **Download semua file** ke laptop lokal\n4. ✅ **Save Version** sekali lagi sebagai backup di Versions\n\nKalau session terhenti SEBELUM Anda download:\n- File-file yang Anda lihat di Output session **AKAN HILANG**\n- Tapi file dari **Save & Run All (Commit)** sebelumnya **MASIH AMAN** di Versions tab\n- Bisa di-download dari sana sebagai backup\n","metadata":{}},{"cell_type":"code","source":"import os\n\nprint('=' * 70)\nprint('📁 FILES YANG HARUS DI-DOWNLOAD:')\nprint('=' * 70)\n\nfiles_to_check = [f for f in os.listdir('.') if f.startswith(('summary_', 'results_', 'snapshot_', 'training_curves_', 'EffNetV2'))]\nfiles_to_check.sort()\n\nfor f in files_to_check:\n    size_mb = os.path.getsize(f) / (1024 * 1024)\n    print(f'  {f} ({size_mb:.2f} MB)')\n\nprint('\\n' + '=' * 70)\nprint('⚠️  SEKARANG:')\nprint('   1. Buka tab Output di sidebar kanan')\nprint('   2. Klik kanan pada masing-masing file → Download')\nprint('   3. Atau: Klik tombol Download All (kalau ada)')\nprint('   4. Save Version sekali lagi sebagai backup')\nprint('=' * 70)","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}