{"metadata":{"kernelspec":{"name":"python3","display_name":"Python 3","language":"python"},"language_info":{"name":"python","version":"3.11.11","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":19991,"databundleVersionId":1117522,"isSourceIdPinned":false,"sourceType":"competition"}],"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# BÁO CÁO ĐỒ ÁN MÔN HỌC CSC14101 - Ẩn dữ liệu & Chia sẻ thông tin\n\n**Lưu ý: Notebook chỉ chạy được ngay trên Kaggle**\n**Nếu chạy local hoặc Google Colab thì phải tải dataset về**","metadata":{}},{"cell_type":"markdown","source":"## Giới thiệu về đồ án\n\nSteganalysis là một lĩnh vực quan trọng trong an ninh máy tính, là quá trình phát hiện thông tin ẩn trong hình ảnh, video,..., và các digital media nói chung (tạm dịch là dữ liệu số được truyền qua internet).\n\nTrong bài toán này, ta cần phải xây dựng giải pháp, cụ thể là dùng machine learning, để phát hiện thông tin ẩn trong các hình ảnh kỹ thuật số trong bộ dữ liệu ALASKA2 một cách hiệu quả, ít sai sót nhất có thể.\n\n### Bộ dữ liệu\n\nBộ dữ liệu ALASKA2 gồm 5 thư mục:\n- Thư mục Cover gồm 75000 ảnh gốc (không qua chỉnh sửa)\n- 3 thư mục JMiPOD, JUNIWARD, UERD mỗi cái gồm 75000 ảnh đã được giấu tin tương ứng với 3 thuật toán steganographic\n- Thư mục Test gồm 5000 ảnh dùng để kiểm tra, nhiệm vụ là xác định ảnh nào có thông tin ẩn\nvà sample_submission.csv: File mẫu cho định dạng kết quả cần nộp\n\nMỗi thuật toán giấu tin được sử dụng với xác suất như nhau.\n\nĐộ dài tin nhắn (payload) được điều chỉnh để độ khó xấp xỉ như nhau, bất kể nội dung ảnh. Ảnh có nội dung mịn sẽ giấu tin nhắn ngắn hơn, trong khi ảnh có nhiều chi tiết (textured) sẽ giấu được nhiều bit hơn.\n\nĐộ dài tin nhắn trung bình là 0.4 bit trên mỗi hệ số AC DCT khác 0.\n\nTất cả ảnh được nén với một trong ba chất lượng JPEG: 95, 90 hoặc 75.\n\n### Cách tính điểm\n\nCuộc thi sẽ sử dụng diện tích dưới đường cong ROC có trọng số (weighted AUC), nhằm tập trung vào khả năng phát hiện đáng tin cậy với tỷ lệ báo động sai thấp.\n\nViệc tính toán weighted AUC được thực hiện như sau:\n- Đường cong ROC được chia thành các vùng dựa trên ngưỡng True Positive Rate (TPR). Cụ thể, các ngưỡng TPR là 0.0, 0.4 và 1.0.\n- Các vùng này được gán trọng số khác nhau:\n  - Vùng có TPR từ 0 đến 0.4 được nhân trọng số 2 lần.\n  - Vùng có TPR từ 0.4 đến 1.0 được nhân trọng số 1 lần.\n- Tổng diện tích có trọng số sau đó được chuẩn hóa bằng tổng các trọng số, đảm bảo giá trị weighted AUC cuối cùng nằm trong khoảng từ 0 đến 1.\n","metadata":{}},{"cell_type":"markdown","source":"## Thành viên & phân công\n\n**Công việc chung**\n\n- Phân tích đồ án\n- Tìm hiểu những giải pháp (phương pháp, code) có sẵn và từ đó tìm hướng đi phù hợp cho nhóm\n- Học cách tận dụng tài nguyên của Kaggle, các tác vụ machine learning\n\n**Mai Vinh Hiển - 20120076**\n\n- Xử lý dữ liệu\n- Học PyTorch\n- Xây dựng giải pháp tận dụng được GPU của Kaggle\n\n**Dương Đình Bảo Hoàng - 20120087**\n\n- Học Tensorflow\n- Xây dựng giải pháp tận dụng được TPU của Kaggle\n\n**Những gì học được qua đồ án này**\n\n- Hiểu rõ hơn về steganalysis, cách ứng dụng machine learning vào steganalysis, những thách thức của steganalysis\n- Nhận thức về tầm quan trọng của dữ liệu đối với bài toán này\n- Kinh nghiệm xây dựng và huấn luyện mô hình deep learning với PyTorch, Tensorflow, và các thư viện liên quan như albumentations\n- Kinh nghiệm sử dụng Kaggle, hiểu được quy trình huấn luyện model liên tục trong thời gian dài","metadata":{}},{"cell_type":"markdown","source":"## Setup & Imports","metadata":{}},{"cell_type":"code","source":"!pip install --upgrade pip\n!pip install -qU timm albumentations","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-08T12:40:40.069075Z","iopub.execute_input":"2025-06-08T12:40:40.069349Z","iopub.status.idle":"2025-06-08T12:40:44.196092Z","shell.execute_reply.started":"2025-06-08T12:40:40.069329Z","shell.execute_reply":"2025-06-08T12:40:44.195195Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import os\nimport random\nimport numpy as np\nimport pandas as pd\nimport torch\nimport torch.nn as nn\nfrom torch.utils.data import Dataset, DataLoader\nimport cv2\nimport albumentations as A\nfrom albumentations.pytorch import ToTensorV2\nfrom sklearn.model_selection import StratifiedKFold\nfrom sklearn.metrics import roc_curve\nfrom tqdm.notebook import tqdm\nimport timm\nimport glob\n\n# for reproducibility\nimport warnings\nwarnings.filterwarnings('ignore')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-08T12:40:44.197922Z","iopub.execute_input":"2025-06-08T12:40:44.198155Z","iopub.status.idle":"2025-06-08T12:40:44.203922Z","shell.execute_reply.started":"2025-06-08T12:40:44.198133Z","shell.execute_reply":"2025-06-08T12:40:44.203110Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Config","metadata":{}},{"cell_type":"code","source":"class CFG:\n    # General\n    seed = 42\n    device = torch.device('cuda' if torch.cuda.is_available() else 'cpu')\n    \n    # Data\n    data_path = '/kaggle/input/alaska2-image-steganalysis/'\n\n    # Take a subset of the original dataset\n    n_stego_samples_per_type = 20000\n    \n    # To create a 1:1 balance, we need an equal number of cover images\n    n_cover_samples = 3 * n_stego_samples_per_type\n    \n    image_size = 512\n    \n    # Model\n    model_name = 'tf_efficientnet_b1_ns' \n    num_classes = 1\n    \n    # Training\n    n_folds = 5\n    fold_to_train = 0\n    epochs = 1\n    train_batch_size = 16\n    valid_batch_size = 32\n    \n    # Optimizer & Scheduler\n    lr = 1e-4\n    weight_decay = 1e-6\n    T_0 = 5 \n    eta_min = 1e-6\n    \n    # Checkpointing\n    checkpoint_save_path = '/kaggle/working/latest_checkpoint.pth'\n    best_model_save_path = f'/kaggle/working/best_model_fold_{fold_to_train}.pth'\n    checkpoint_load_path = glob.glob('/kaggle/input/*/latest_checkpoint.pth')\n    if len(checkpoint_load_path) > 0:\n        checkpoint_load_path = checkpoint_load_path[0]\n    else:\n        checkpoint_load_path = None\n\n\ndef set_seed(seed):\n    random.seed(seed)\n    os.environ['PYTHONHASHSEED'] = str(seed)\n    np.random.seed(seed)\n    torch.manual_seed(seed)\n    torch.cuda.manual_seed(seed)\n    torch.backends.cudnn.deterministic = True\n    torch.backends.cudnn.benchmark = True\n\nset_seed(CFG.seed)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-08T12:40:44.204747Z","iopub.execute_input":"2025-06-08T12:40:44.204998Z","iopub.status.idle":"2025-06-08T12:40:44.222982Z","shell.execute_reply.started":"2025-06-08T12:40:44.204977Z","shell.execute_reply":"2025-06-08T12:40:44.222283Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Weighted AUC Implementation","metadata":{}},{"cell_type":"code","source":"def alaska_weighted_auc(y_true, y_pred):\n    \"\"\"\n    Calculates the weighted AUC score for the ALASKA2 competition.\n    \"\"\"\n    tpr_thresholds = [0.0, 0.4, 1.0]\n    weights = [2, 1]\n    fpr, tpr, thresholds = roc_curve(y_true, y_pred, pos_label=1)\n    if len(fpr) < 2: return 0.5\n    \n    areas = np.array([0.0] * len(weights))\n    for i, lower in enumerate(tpr_thresholds[:-1]):\n        upper = tpr_thresholds[i+1]\n        mask = (tpr >= lower) & (tpr < upper)\n        if np.any(mask):\n            mask_indices = np.where(mask)[0]\n            start_idx, end_idx = mask_indices[0], mask_indices[-1]\n            tpr_slice = np.concatenate([[lower], tpr[start_idx:end_idx+1], [upper]])\n            fpr_slice = np.concatenate([[np.interp(lower, tpr, fpr)], fpr[start_idx:end_idx+1], [np.interp(upper, tpr, fpr)]])\n            tpr_slice, unique_indices = np.unique(tpr_slice, return_index=True)\n            fpr_slice = fpr_slice[unique_indices]\n            areas[i] = np.trapz(fpr_slice, tpr_slice)\n            \n    return np.sum(areas * weights) / np.sum(weights)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-08T12:40:44.224722Z","iopub.execute_input":"2025-06-08T12:40:44.224934Z","iopub.status.idle":"2025-06-08T12:40:44.240057Z","shell.execute_reply.started":"2025-06-08T12:40:44.224919Z","shell.execute_reply":"2025-06-08T12:40:44.239463Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Data Preprocessing","metadata":{}},{"cell_type":"code","source":"all_files = []\n\n# --- Sample Cover Images ---\ncover_folder = os.path.join(CFG.data_path, 'Cover')\ncover_files = [os.path.join(cover_folder, f) for f in os.listdir(cover_folder)]\nrandom.shuffle(cover_files)\nall_files.extend(cover_files[:CFG.n_cover_samples])\nprint(f\"Sampled {len(cover_files[:CFG.n_cover_samples])} images from Cover.\")\n\n# --- Sample Stego Images ---\nstego_folders = ['JMiPOD', 'JUNIWARD', 'UERD']\nprint(f\"Sampling {CFG.n_stego_samples_per_type} images from each stego type...\")\nfor folder in stego_folders:\n    folder_path = os.path.join(CFG.data_path, folder)\n    files_in_folder = [os.path.join(folder_path, f) for f in os.listdir(folder_path)]\n    random.shuffle(files_in_folder)\n    all_files.extend(files_in_folder[:CFG.n_stego_samples_per_type])\n    print(f\"  - Took {len(files_in_folder[:CFG.n_stego_samples_per_type])} images from {folder}\")\n\n# --- Create DataFrame and Folds ---\ndf = pd.DataFrame({'image_path': all_files})\ndf['label'] = df['image_path'].apply(lambda x: 0 if 'Cover' in x else 1)\ndf['image_id'] = df['image_path'].apply(os.path.basename)\n\n# Now the dataset is balanced 1:1.\n# StratifiedKFold will ensure this ratio is maintained in train/valid splits.\nskf = StratifiedKFold(n_splits=CFG.n_folds, shuffle=True, random_state=CFG.seed)\ndf['fold'] = -1\nfor fold, (train_idx, val_idx) in enumerate(skf.split(df, df['label'])):\n    df.loc[val_idx, 'fold'] = fold\n\nprint(\"\\nBalanced subset dataset distribution:\")\nprint(df['label'].value_counts())\nprint(\"\\nFold distribution:\")\nprint(df.groupby('fold')['label'].value_counts())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-08T12:40:44.240761Z","iopub.execute_input":"2025-06-08T12:40:44.240971Z","iopub.status.idle":"2025-06-08T12:40:47.526252Z","shell.execute_reply.started":"2025-06-08T12:40:44.240949Z","shell.execute_reply":"2025-06-08T12:40:47.525585Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Augmentations and Dataset\ndef get_transforms(data_type='train'):\n    if data_type == 'train':\n        return A.Compose([\n            A.HorizontalFlip(p=0.5),\n            A.VerticalFlip(p=0.5),\n            A.Resize(height=CFG.image_size, width=CFG.image_size, always_apply=True),\n            A.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]),\n            ToTensorV2(),\n        ])\n    else:\n        return A.Compose([\n            A.Resize(height=CFG.image_size, width=CFG.image_size, always_apply=True),\n            A.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]),\n            ToTensorV2(),\n        ])\n\nclass AlaskaDataset(Dataset):\n    def __init__(self, df, transforms=None):\n        self.df = df\n        self.image_paths = df['image_path'].values\n        self.labels = df['label'].values\n        self.transforms = transforms\n    def __len__(self): return len(self.df)\n    def __getitem__(self, idx):\n        image_path = self.image_paths[idx]\n        label = torch.tensor(self.labels[idx], dtype=torch.float)\n        image = cv2.imread(image_path)\n        image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\n        if self.transforms: image = self.transforms(image=image)['image']\n        return image, label","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-08T12:40:47.527016Z","iopub.execute_input":"2025-06-08T12:40:47.527205Z","iopub.status.idle":"2025-06-08T12:40:47.534424Z","shell.execute_reply.started":"2025-06-08T12:40:47.527190Z","shell.execute_reply":"2025-06-08T12:40:47.533589Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Training","metadata":{}},{"cell_type":"code","source":"def train_fn(loader, model, criterion, optimizer, scheduler, device):\n    model.train()\n    running_loss = 0.0\n    pbar = tqdm(loader, desc=\"Training\")\n    for images, labels in pbar:\n        images, labels = images.to(device), labels.to(device).unsqueeze(1)\n        optimizer.zero_grad()\n        outputs = model(images)\n        loss = criterion(outputs, labels)\n        loss.backward()\n        optimizer.step()\n        if scheduler: scheduler.step()\n        running_loss += loss.item()\n        pbar.set_postfix(loss=loss.item(), lr=optimizer.param_groups[0]['lr'])\n    return running_loss / len(loader)\n\ndef eval_fn(loader, model, criterion, device):\n    model.eval()\n    running_loss, all_preds, all_labels = 0.0, [], []\n    with torch.no_grad():\n        pbar = tqdm(loader, desc=\"Evaluating\")\n        for images, labels in pbar:\n            images, labels = images.to(device), labels.to(device).unsqueeze(1)\n            outputs = model(images)\n            loss = criterion(outputs, labels)\n            running_loss += loss.item()\n            all_preds.append(torch.sigmoid(outputs).cpu().numpy())\n            all_labels.append(labels.cpu().numpy())\n    all_preds = np.concatenate(all_preds).flatten()\n    all_labels = np.concatenate(all_labels).flatten()\n    val_loss = running_loss / len(loader)\n    score = alaska_weighted_auc(all_labels, all_preds)\n    return val_loss, score","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-08T12:40:47.535144Z","iopub.execute_input":"2025-06-08T12:40:47.535448Z","iopub.status.idle":"2025-06-08T12:40:47.555984Z","shell.execute_reply.started":"2025-06-08T12:40:47.535425Z","shell.execute_reply":"2025-06-08T12:40:47.555298Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Training Loop with Checkpointing\ndef run_training(fold):\n    print(f\"========== Starting Training for Fold {fold} ==========\")\n    \n    # --- Data Setup ---\n    train_df = df[df['fold'] != fold].reset_index(drop=True)\n    valid_df = df[df['fold'] == fold].reset_index(drop=True)\n    train_dataset = AlaskaDataset(train_df, transforms=get_transforms('train'))\n    valid_dataset = AlaskaDataset(valid_df, transforms=get_transforms('valid'))\n    train_loader = DataLoader(train_dataset, batch_size=CFG.train_batch_size, shuffle=True, num_workers=2, pin_memory=True)\n    valid_loader = DataLoader(valid_dataset, batch_size=CFG.valid_batch_size, shuffle=False, num_workers=2, pin_memory=True)\n    \n    # --- Model, Optimizer, Scheduler Setup ---\n    model = timm.create_model(CFG.model_name, pretrained=True, num_classes=CFG.num_classes).to(CFG.device)\n    optimizer = torch.optim.AdamW(model.parameters(), lr=CFG.lr, weight_decay=CFG.weight_decay)\n    scheduler = torch.optim.lr_scheduler.CosineAnnealingWarmRestarts(optimizer, T_0=CFG.T_0 * len(train_loader), eta_min=CFG.eta_min)\n    criterion = nn.BCEWithLogitsLoss()\n\n    # --- Checkpoint Loading ---\n    start_epoch = 0\n    best_score = 0.0\n    if CFG.checkpoint_load_path and os.path.exists(CFG.checkpoint_load_path):\n        print(f\"Resuming training from checkpoint: {CFG.checkpoint_load_path}\")\n        checkpoint = torch.load(CFG.checkpoint_load_path, map_location=CFG.device, weights_only=False)\n        model.load_state_dict(checkpoint['model_state'])\n        optimizer.load_state_dict(checkpoint['optimizer_state'])\n        scheduler.load_state_dict(checkpoint['scheduler_state'])\n        start_epoch = 10 # checkpoint['epoch'] + 1 # Start from the next epoch\n        best_score = checkpoint['best_score']\n        print(f\"Loaded model from epoch {start_epoch-1} with best score: {best_score:.4f}\")\n    else:\n        print(\"No checkpoint found, starting training from scratch.\")\n\n    # --- Main Loop ---\n    for epoch in range(start_epoch, CFG.epochs):\n        print(f\"\\n--- Epoch {epoch+1}/{CFG.epochs} ---\")\n        train_loss = train_fn(train_loader, model, criterion, optimizer, scheduler, CFG.device)\n        val_loss, val_score = eval_fn(valid_loader, model, criterion, CFG.device)\n        \n        print(f\"Epoch {epoch+1} -> Train Loss: {train_loss:.4f}, Valid Loss: {val_loss:.4f}, Valid Weighted AUC: {val_score:.4f}\")\n\n        # Save best model based on validation score\n        if val_score > best_score:\n            print(f\"Validation score improved! ({best_score:.4f} -> {val_score:.4f}). Saving best model...\")\n            best_score = val_score\n            torch.save(model.state_dict(), CFG.best_model_save_path)\n        \n        # Save current state for resuming\n        checkpoint = {\n            'epoch': epoch,\n            'model_state': model.state_dict(),\n            'optimizer_state': optimizer.state_dict(),\n            'scheduler_state': scheduler.state_dict(),\n            'best_score': best_score,\n        }\n        torch.save(checkpoint, CFG.checkpoint_save_path)\n        print(f\"Epoch {epoch+1} state saved to checkpoint: {CFG.checkpoint_save_path}\")\n\n    print(f\"\\n========== Finished Training for Fold {fold}. Best Score: {best_score:.4f} ==========\")\n\n# Start the training process\nrun_training(CFG.fold_to_train)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-08T12:40:48.238014Z","iopub.execute_input":"2025-06-08T12:40:48.238230Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Testing & submission","metadata":{}},{"cell_type":"code","source":"def create_submission():\n    print(\"\\nStarting inference on the test set...\")\n    \n    test_folder = os.path.join(CFG.data_path, 'Test')\n    test_image_ids = [f for f in os.listdir(test_folder) if f.endswith('.jpg')]\n    \n    test_df = pd.DataFrame({'image_id': test_image_ids})\n    test_df['image_path'] = test_df['image_id'].apply(lambda x: os.path.join(test_folder, x))\n    test_df['label'] = 0 # Dummy label\n    \n    test_dataset = AlaskaDataset(test_df, transforms=get_transforms('test'))\n    test_loader = DataLoader(test_dataset, batch_size=CFG.valid_batch_size, shuffle=False, num_workers=2)\n    \n    # Load the best performing model for inference\n    model = timm.create_model(CFG.model_name, pretrained=False, num_classes=CFG.num_classes)\n    model.load_state_dict(torch.load(CFG.best_model_save_path))\n    model.to(CFG.device)\n    model.eval()\n    \n    predictions = []\n    with torch.no_grad():\n        pbar = tqdm(test_loader, desc=\"Predicting\")\n        for images, _ in pbar:\n            images = images.to(CFG.device)\n            outputs = model(images)\n            predictions.extend(outputs.cpu().numpy().flatten())\n            \n    submission_df = pd.DataFrame({'Id': test_image_ids, 'Label': predictions}).sort_values('Id')\n    submission_df.to_csv('submission.csv', index=False)\n    \n    print(\"\\nSubmission file created successfully!\")\n    print(submission_df)\n\ncreate_submission()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Result\n\nWe got 0.689 as our best score (running on 3 epochs and full dataset)","metadata":{}},{"cell_type":"markdown","source":"## Future improvements\n\n1. Sử dụng mô hình khác có độ chính xác cao hơn (và khi ta có GPU xịn)\n2. Tìm hiểu các cách trích xuất đặc trưng khác không làm ảnh hưởng đến thông tin ẩn trong ảnh\n3. Tìm cách cải tiến solution này để tận dụng được TPU của Kaggle","metadata":{}}]}