{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.10.0"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceType":"competition","sourceId":4117,"databundleVersionId":46665},{"sourceType":"datasetVersion","sourceId":16164731,"datasetId":10364823,"databundleVersionId":17140759},{"sourceType":"datasetVersion","sourceId":16164482,"datasetId":10364690,"databundleVersionId":17140493},{"sourceType":"datasetVersion","sourceId":15522655,"datasetId":9931127,"databundleVersionId":16449701},{"sourceType":"datasetVersion","sourceId":15743447,"datasetId":10088103,"databundleVersionId":16686091}],"dockerImageVersionId":31329,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"id":"a1b2c3d4","cell_type":"markdown","source":"# B4.2 — Channel Occlusion Attribution (XAI chính)\n**PIXEL Pipeline v3 — Inference-only, ~1 giờ T4**\n\nVới mỗi sample trong tập **test**:\n- `p₀  = p(y | R, G, B)` — full image\n- `p^R = p(y | 0, G, B)` — che kênh R (set raw = 0)\n- `p^G = p(y | R, 0, B)` — che kênh G\n- `p^B = p(y | R, G, 0)` — che kênh B\n- `drop_R = p₀ − p^R`,  `drop_G = p₀ − p^G`,  `drop_B = p₀ − p^B`\n\n**Output 1** — Bảng mean attribution (2 dataset × 3 kênh)  \n**Output 2** — Heatmap family × channel (figure XAI chính của paper)  \n**Output 3** — Diễn giải: family nào phụ thuộc entropy / texture","metadata":{}},{"id":"b2c3d4e5","cell_type":"code","source":"# ============================================================\n# CELL 1: Imports\n# ============================================================\nimport os, json, gc, warnings\nfrom contextlib import nullcontext\n\nimport numpy as np\nimport pandas as pd\nimport torch\nimport torch.nn as nn\nfrom torch.amp import autocast\nfrom torch.utils.data import DataLoader, random_split, Subset, Dataset\nfrom torchvision import datasets, transforms, models\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom PIL import Image\nfrom tqdm.auto import tqdm\n\nwarnings.filterwarnings(\"ignore\")\n\nDEVICE      = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\nAMP_ENABLED = torch.cuda.is_available()\nprint(f\"Device : {DEVICE}\")\nprint(f\"AMP    : {AMP_ENABLED}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-08T21:56:53.761371Z","iopub.execute_input":"2026-05-08T21:56:53.762141Z","iopub.status.idle":"2026-05-08T21:56:53.769359Z","shell.execute_reply.started":"2026-05-08T21:56:53.762108Z","shell.execute_reply":"2026-05-08T21:56:53.768425Z"}},"outputs":[],"execution_count":null},{"id":"c3d4e5f6","cell_type":"code","source":"# ============================================================\n# CELL 2: Config\n# ============================================================\n# ── Đường dẫn dataset (giống training notebook) ──────────────\nDATASET_CONFIGS = {\n    \"microsoft_rgb\": {\n        \"path\"      : \"/kaggle/input/datasets/vnhtbo/microsoft/train_rgb/train_rgb\",\n        \"csv_path\"  : \"/kaggle/input/competitions/malware-classification/trainLabels.csv\",\n        \"is_rgb\"    : True,\n        \"num_classes\": 9,\n        \"description\": \"Microsoft BIG-2015\",\n        # TODO: đổi thành đường dẫn checkpoint EB3 tương ứng nếu cần\n        \"ckpt_path\" : \"/kaggle/input/datasets/vnhtbo/eb3-train/best_eb3_teacher_run1.pt\",\n    },\n    \"malimg_rgb\": {\n        \"path\"      : \"/kaggle/input/datasets/dongquan/malimg-rgb/kaggle/working/malimg_rgb\",\n        \"csv_path\"  : None,\n        \"is_rgb\"    : True,\n        \"num_classes\": 25,\n        \"description\": \"Malimg\",\n        # TODO: thay bằng đường dẫn checkpoint EB3 train trên malimg\n        \"ckpt_path\" : \"/kaggle/input/datasets/vnhtbo/eb3-train-malimg/best_eb3_teacher_run1 (1).pt\",\n    },\n}\n\n# ── Hyperparams (phải khớp với training notebook) ─────────────\nIMG_SIZE   = 300\nBATCH_SIZE = 32\nVAL_RATIO  = 0.15\nTEST_RATIO = 0.15\nSEED       = 42\n\n# Đường dẫn lưu split từ training notebook (nếu có)\nCKPT_DIR   = \"/kaggle/working/checkpoints\"\nOUTPUT_DIR = \"/kaggle/working\"\nos.makedirs(OUTPUT_DIR, exist_ok=True)\n\n# ── ImageNet normalization (khớp với training notebook) ───────\nNORM_MEAN = [0.485, 0.456, 0.406]\nNORM_STD  = [0.229, 0.224, 0.225]\n\n# Giá trị normalized tương ứng raw pixel = 0\n# (dùng để occlude: set kênh về 0 trước normalize)\nOCCLUDE_VALS = [\n    (0.0 - NORM_MEAN[0]) / NORM_STD[0],   # channel R  ≈ -2.1179\n    (0.0 - NORM_MEAN[1]) / NORM_STD[1],   # channel G  ≈ -2.0357\n    (0.0 - NORM_MEAN[2]) / NORM_STD[2],   # channel B  ≈ -1.8044\n]\nprint(f\"Occlude vals — R: {OCCLUDE_VALS[0]:.4f} | G: {OCCLUDE_VALS[1]:.4f} | B: {OCCLUDE_VALS[2]:.4f}\")\n\n# ── Mapping số → tên family (Microsoft BIG-2015) ──────────────\nMS_FAMILY_NAMES = {\n    '1': 'Ramnit',\n    '2': 'Lollipop',\n    '3': 'Kelihos_v3',\n    '4': 'Vundo',\n    '5': 'Simda',\n    '6': 'Tracur',\n    '7': 'Kelihos_v1',\n    '8': 'Obfuscator.ACY',\n    '9': 'Gatak',\n}","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-08T21:56:53.771847Z","iopub.execute_input":"2026-05-08T21:56:53.772306Z","iopub.status.idle":"2026-05-08T21:56:53.789913Z","shell.execute_reply.started":"2026-05-08T21:56:53.772273Z","shell.execute_reply":"2026-05-08T21:56:53.789091Z"}},"outputs":[],"execution_count":null},{"id":"d4e5f6a7","cell_type":"code","source":"# ============================================================\n# CELL 3: Helper functions (khớp với training notebook)\n# ============================================================\n\nclass CSVImageDataset(Dataset):\n    \"\"\"Đọc ảnh từ thư mục phẳng + CSV label (copy từ training notebook)\"\"\"\n    def __init__(self, root, csv_path, transform=None):\n        df = pd.read_csv(csv_path)\n        unique_classes     = sorted(df[\"Class\"].unique())\n        self.class_to_idx  = {c: i for i, c in enumerate(unique_classes)}\n        self.classes       = [str(c) for c in unique_classes]\n        existing           = set(os.listdir(root))\n        self.samples, self.targets = [], []\n        for _, row in df.iterrows():\n            fname = f\"{row['Id']}.png\"\n            if fname in existing:\n                label = self.class_to_idx[row[\"Class\"]]\n                self.samples.append((os.path.join(root, fname), label))\n                self.targets.append(label)\n        self.transform = transform\n        print(f\"  Loaded {len(self.samples)}/{len(df)} images, {len(self.classes)} classes\")\n\n    def __len__(self):  return len(self.samples)\n\n    def __getitem__(self, idx):\n        path, label = self.samples[idx]\n        img = Image.open(path).convert(\"RGB\")\n        if self.transform:\n            img = self.transform(img)\n        return img, label\n\n\ndef get_val_transform(is_rgb: bool):\n    \"\"\"Val/test transform — không augmentation (khớp với training notebook)\"\"\"\n    norm_mean = NORM_MEAN if is_rgb else [0.5, 0.5, 0.5]\n    norm_std  = NORM_STD  if is_rgb else [0.5, 0.5, 0.5]\n    base = [transforms.Resize((IMG_SIZE, IMG_SIZE))]\n    if not is_rgb:\n        base.append(transforms.Grayscale(num_output_channels=3))\n    return transforms.Compose(base + [\n        transforms.ToTensor(),\n        transforms.Normalize(norm_mean, norm_std),\n    ])\n\n\ndef build_efficientnet_b3(num_classes: int) -> nn.Module:\n    \"\"\"Xây dựng EB3 với kiến trúc giống training notebook\"\"\"\n    model = models.efficientnet_b3(weights=models.EfficientNet_B3_Weights.IMAGENET1K_V1)\n    for param in model.parameters():\n        param.requires_grad = False\n    in_features = model.classifier[1].in_features\n    model.classifier = nn.Sequential(\n        nn.Dropout(p=0.3, inplace=True),\n        nn.Linear(in_features, num_classes),\n    )\n    return model.to(DEVICE)\n\n\ndef get_test_indices(full_ds, dataset_name):\n    \"\"\"\n    Lấy test split indices.\n    Ưu tiên load từ split file đã lưu (cùng session với training).\n    Nếu không có → tái tạo với seed=42 (kết quả giống hệt training notebook).\n    \"\"\"\n    split_file = os.path.join(CKPT_DIR, f\"split_{dataset_name}.json\")\n    n       = len(full_ds)\n    n_test  = int(n * TEST_RATIO)\n    n_val   = int(n * VAL_RATIO)\n    n_train = n - n_val - n_test\n\n    if os.path.exists(split_file):\n        saved = json.load(open(split_file))\n        print(f\"  📂 Loaded split: {split_file}  (seed={saved['seed']}, total={saved['total']})\")\n        indices = saved[\"test\"]\n    else:\n        gen = torch.Generator().manual_seed(SEED)\n        _, _, test_ds = random_split(full_ds, [n_train, n_val, n_test], generator=gen)\n        indices = test_ds.indices\n        print(f\"  ⚠️  Split file not found — regenerated deterministically (seed={SEED})\")\n\n    print(f\"  Test set: {len(indices)} samples\")\n    return indices\n\n\nprint(\"✅ Helper functions ready\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-08T21:56:53.791322Z","iopub.execute_input":"2026-05-08T21:56:53.791686Z","iopub.status.idle":"2026-05-08T21:56:53.809077Z","shell.execute_reply.started":"2026-05-08T21:56:53.791661Z","shell.execute_reply":"2026-05-08T21:56:53.808201Z"}},"outputs":[],"execution_count":null},{"id":"e5f6a7b8","cell_type":"code","source":"# ============================================================\n# CELL 4: Channel Occlusion Inference\n# ============================================================\n\n@torch.no_grad()\ndef run_channel_occlusion(model, loader, class_names):\n    \"\"\"\n    Chạy 4 forward pass cho mỗi batch:\n      full  : ảnh gốc (R, G, B)\n      no_R  : kênh R set về OCCLUDE_VALS[0] (tương đương raw=0)\n      no_G  : kênh G set về OCCLUDE_VALS[1]\n      no_B  : kênh B set về OCCLUDE_VALS[2]\n\n    Trả về DataFrame với per-sample attribution scores.\n    \"\"\"\n    model.eval()\n    amp_ctx = autocast('cuda') if AMP_ENABLED else nullcontext()\n\n    occ_R = float(OCCLUDE_VALS[0])\n    occ_G = float(OCCLUDE_VALS[1])\n    occ_B = float(OCCLUDE_VALS[2])\n\n    records = []\n\n    for imgs, labels in tqdm(loader, desc=\"  Occlusion\", leave=True):\n        imgs   = imgs.to(DEVICE)\n        labels = labels.to(DEVICE)\n        bsz    = imgs.size(0)\n        idx    = torch.arange(bsz, device=DEVICE)\n\n        # ── Full image ────────────────────────────────────────\n        with amp_ctx:\n            logits_full = model(imgs)\n        probs_full = torch.softmax(logits_full, dim=1)\n        p0         = probs_full[idx, labels]          # shape [B]\n        preds      = logits_full.argmax(dim=1)\n\n        # ── Occlude R (channel 0) ─────────────────────────────\n        imgs_occ = imgs.clone()\n        imgs_occ[:, 0, :, :] = occ_R\n        with amp_ctx:\n            p_R = torch.softmax(model(imgs_occ), dim=1)[idx, labels]\n\n        # ── Occlude G (channel 1) ─────────────────────────────\n        imgs_occ = imgs.clone()\n        imgs_occ[:, 1, :, :] = occ_G\n        with amp_ctx:\n            p_G = torch.softmax(model(imgs_occ), dim=1)[idx, labels]\n\n        # ── Occlude B (channel 2) ─────────────────────────────\n        imgs_occ = imgs.clone()\n        imgs_occ[:, 2, :, :] = occ_B\n        with amp_ctx:\n            p_B = torch.softmax(model(imgs_occ), dim=1)[idx, labels]\n\n        # ── Attribution drops ─────────────────────────────────\n        drop_R = (p0 - p_R).cpu().numpy()\n        drop_G = (p0 - p_G).cpu().numpy()\n        drop_B = (p0 - p_B).cpu().numpy()\n        p0_np  = p0.cpu().numpy()\n        lbl_np = labels.cpu().numpy()\n        prd_np = preds.cpu().numpy()\n\n        for i in range(bsz):\n            records.append({\n                \"true_class_idx\"  : int(lbl_np[i]),\n                \"true_class_name\" : class_names[lbl_np[i]],\n                \"pred_class_idx\"  : int(prd_np[i]),\n                \"correct\"         : bool(lbl_np[i] == prd_np[i]),\n                \"p0\"              : float(p0_np[i]),\n                \"drop_R\"          : float(drop_R[i]),\n                \"drop_G\"          : float(drop_G[i]),\n                \"drop_B\"          : float(drop_B[i]),\n            })\n\n        del imgs_occ\n        torch.cuda.empty_cache()\n\n    return pd.DataFrame(records)\n\n\nprint(\"✅ Channel occlusion function ready\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-08T21:56:53.963209Z","iopub.execute_input":"2026-05-08T21:56:53.963814Z","iopub.status.idle":"2026-05-08T21:56:53.975963Z","shell.execute_reply.started":"2026-05-08T21:56:53.963780Z","shell.execute_reply":"2026-05-08T21:56:53.975019Z"}},"outputs":[],"execution_count":null},{"id":"f6a7b8c9","cell_type":"code","source":"# ============================================================\n# CELL 5: Main loop — chạy trên cả 2 dataset\n# ============================================================\n\nall_results = {}\n\nfor ds_name, cfg in DATASET_CONFIGS.items():\n\n    print(f\"\\n{'='*60}\")\n    print(f\"  {cfg['description']} ({ds_name})\")\n    print(f\"{'='*60}\")\n\n    # ── Kiểm tra paths ────────────────────────────────────────\n    if not os.path.exists(cfg[\"path\"]):\n        print(f\"  ❌ SKIP — dataset path không tìm thấy: {cfg['path']}\")\n        continue\n    if not os.path.exists(cfg[\"ckpt_path\"]):\n        print(f\"  ❌ SKIP — checkpoint không tìm thấy: {cfg['ckpt_path']}\")\n        continue\n\n    # ── Load dataset (val/test transform — không augmentation) ──\n    val_tf = get_val_transform(cfg[\"is_rgb\"])\n\n    if cfg.get(\"csv_path\"):\n        full_ds = CSVImageDataset(cfg[\"path\"], cfg[\"csv_path\"], transform=val_tf)\n    else:\n        full_ds = datasets.ImageFolder(cfg[\"path\"], transform=val_tf)\n\n    test_indices = get_test_indices(full_ds, ds_name)\n    test_ds      = Subset(full_ds, test_indices)\n    test_loader  = DataLoader(\n        test_ds, batch_size=BATCH_SIZE, shuffle=False,\n        num_workers=2, pin_memory=True,\n    )\n\n    # ── Load model ────────────────────────────────────────────\n    model = build_efficientnet_b3(cfg[\"num_classes\"])\n    state = torch.load(cfg[\"ckpt_path\"], map_location=DEVICE)\n    model.load_state_dict(state)\n    model.eval()\n    total_params = sum(p.numel() for p in model.parameters()) / 1e6\n    print(f\"  ✅ Checkpoint loaded — {total_params:.1f}M params\")\n\n    # ── Resolve class names ───────────────────────────────────\n    raw_classes = full_ds.classes\n    if ds_name == \"microsoft_rgb\":\n        # Map số '1'–'9' → tên family thực sự\n        display_classes = [MS_FAMILY_NAMES.get(c, c) for c in raw_classes]\n    else:\n        display_classes = list(raw_classes)\n\n    # ── Run occlusion ─────────────────────────────────────────\n    df = run_channel_occlusion(model, test_loader, display_classes)\n\n    # ── Save raw results ──────────────────────────────────────\n    csv_out = os.path.join(OUTPUT_DIR, f\"channel_occlusion_{ds_name}.csv\")\n    df.to_csv(csv_out, index=False)\n    acc = df[\"correct\"].mean()\n    print(f\"  💾 Saved: {csv_out}  ({len(df)} samples, acc={acc:.4f})\")\n\n    all_results[ds_name] = {\n        \"df\"         : df,\n        \"cfg\"        : cfg,\n        \"classes\"    : display_classes,\n    }\n\n    del model, full_ds, test_ds, test_loader\n    gc.collect()\n    torch.cuda.empty_cache()\n\nprint(f\"\\n✅ Done — {len(all_results)} dataset(s) processed\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-08T21:56:53.977686Z","iopub.execute_input":"2026-05-08T21:56:53.978187Z","iopub.status.idle":"2026-05-08T21:58:22.623625Z","shell.execute_reply.started":"2026-05-08T21:56:53.978162Z","shell.execute_reply":"2026-05-08T21:58:22.622500Z"}},"outputs":[],"execution_count":null},{"id":"a7b8c9d0","cell_type":"markdown","source":"---\n## Output 1 — Bảng Attribution trung bình per Dataset","metadata":{}},{"id":"b8c9d0e1","cell_type":"code","source":"# ============================================================\n# CELL 6: Output 1 — Attribution table (2 rows × 3 cols)\n# ============================================================\n\nrows = []\nfor ds_name, res in all_results.items():\n    df = res[\"df\"]\n    rows.append({\n        \"Dataset\"       : res[\"cfg\"][\"description\"],\n        \"n_samples\"     : len(df),\n        \"Accuracy\"      : round(df[\"correct\"].mean(), 4),\n        \"drop_R mean\"   : round(df[\"drop_R\"].mean(), 4),\n        \"drop_R std\"    : round(df[\"drop_R\"].std(),  4),\n        \"drop_G mean\"   : round(df[\"drop_G\"].mean(), 4),\n        \"drop_G std\"    : round(df[\"drop_G\"].std(),  4),\n        \"drop_B mean\"   : round(df[\"drop_B\"].mean(), 4),\n        \"drop_B std\"    : round(df[\"drop_B\"].std(),  4),\n    })\n\nsummary_df = pd.DataFrame(rows)\nprint(\"=\" * 80)\nprint(\"OUTPUT 1: Mean Channel Attribution Score\")\nprint(\"=\" * 80)\nprint(summary_df.to_string(index=False))\n\nsummary_path = os.path.join(OUTPUT_DIR, \"attribution_summary.csv\")\nsummary_df.to_csv(summary_path, index=False)\nprint(f\"\\n💾 Saved: {summary_path}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-08T21:58:22.625352Z","iopub.execute_input":"2026-05-08T21:58:22.625782Z","iopub.status.idle":"2026-05-08T21:58:22.650337Z","shell.execute_reply.started":"2026-05-08T21:58:22.625724Z","shell.execute_reply":"2026-05-08T21:58:22.649450Z"}},"outputs":[],"execution_count":null},{"id":"c9d0e1f2","cell_type":"markdown","source":"---\n## Output 2 — Heatmap Family × Channel","metadata":{}},{"id":"d0e1f2a3","cell_type":"code","source":"# ============================================================\n# CELL 7: Output 2 — Heatmap family × channel (Figure XAI)\n# ============================================================\n\nn_ds = len(all_results)\nif n_ds == 0:\n    print(\"Không có kết quả để vẽ.\")\nelse:\n    fig, axes = plt.subplots(1, n_ds, figsize=(6 * n_ds, max(6, 0.4 * 25)))\n    if n_ds == 1:\n        axes = [axes]\n\n    for ax, (ds_name, res) in zip(axes, all_results.items()):\n        df = res[\"df\"]\n\n        # Mean attribution per (family, channel)\n        family_attr = (\n            df.groupby(\"true_class_name\")[[\"drop_R\", \"drop_G\", \"drop_B\"]]\n            .mean()\n            .rename(columns={\"drop_R\": \"R\\n(Byte)\",\n                              \"drop_G\": \"G\\n(Entropy)\",\n                              \"drop_B\": \"B\\n(LBP)\"})\n        )\n\n        # Sắp xếp family theo drop_G giảm dần\n        # (family obfuscated sẽ nổi lên trên)\n        family_attr = family_attr.sort_values(\"G\\n(Entropy)\", ascending=False)\n\n        g = sns.heatmap(\n            family_attr,\n            annot=True, fmt=\".3f\",\n            cmap=\"YlOrRd\",\n            ax=ax,\n            linewidths=0.5,\n            cbar_kws={\"label\": \"Attribution drop (p₀ − p^c)\", \"shrink\": 0.7},\n            vmin=0,\n        )\n        ax.set_title(\n            f\"{res['cfg']['description']}\\nChannel Occlusion Attribution\",\n            fontsize=12, fontweight=\"bold\", pad=10,\n        )\n        ax.set_xlabel(\"Channel\", fontsize=10)\n        ax.set_ylabel(\"Malware Family\", fontsize=10)\n        ax.tick_params(axis=\"y\", labelsize=8)\n        ax.tick_params(axis=\"x\", labelsize=9)\n\n    plt.tight_layout()\n    fig_path = os.path.join(OUTPUT_DIR, \"channel_attribution_heatmap.png\")\n    plt.savefig(fig_path, dpi=300, bbox_inches=\"tight\")\n    plt.show()\n    print(f\"💾 Saved: {fig_path}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-08T21:58:22.652665Z","iopub.execute_input":"2026-05-08T21:58:22.653128Z","iopub.status.idle":"2026-05-08T21:58:24.783813Z","shell.execute_reply.started":"2026-05-08T21:58:22.653090Z","shell.execute_reply":"2026-05-08T21:58:24.783005Z"}},"outputs":[],"execution_count":null},{"id":"e1f2a3b4","cell_type":"markdown","source":"---\n## Output 3 — Family Analysis: Entropy vs Texture","metadata":{}},{"id":"f2a3b4c5","cell_type":"code","source":"# ============================================================\n# CELL 8: Output 3 — Diễn giải family nào phụ thuộc kênh nào\n# ============================================================\n\nCHANNEL_LABELS = {\n    \"drop_R\": \"R — Byte amplitude\",\n    \"drop_G\": \"G — Shannon entropy (obfuscated/packed)\",\n    \"drop_B\": \"B — LBP texture (structured code)\",\n}\n\nall_analysis = []\n\nfor ds_name, res in all_results.items():\n    df  = res[\"df\"]\n    print(f\"\\n{'='*60}\")\n    print(f\"  {res['cfg']['description']}\")\n    print(f\"{'='*60}\")\n\n    family_attr = (\n        df.groupby(\"true_class_name\")[[\"drop_R\", \"drop_G\", \"drop_B\"]].mean()\n    )\n\n    # Kênh dominant (attribution cao nhất) của mỗi family\n    dominant_col = family_attr.idxmax(axis=1)\n\n    groups = {col: [] for col in [\"drop_R\", \"drop_G\", \"drop_B\"]}\n    for fam, col in dominant_col.items():\n        groups[col].append(fam)\n\n    for col, fams in groups.items():\n        if not fams:\n            continue\n        print(f\"\\n  [{CHANNEL_LABELS[col]}] — {len(fams)} famil{'y' if len(fams)==1 else 'ies'}:\")\n        for f in sorted(fams):\n            r = family_attr.loc[f, \"drop_R\"]\n            g = family_attr.loc[f, \"drop_G\"]\n            b = family_attr.loc[f, \"drop_B\"]\n            print(f\"    • {f:<22}  drop_R={r:.4f}  drop_G={g:.4f}  drop_B={b:.4f}\")\n\n    # --- Insight tổng hợp ---\n    g_fams = groups[\"drop_G\"]\n    b_fams = groups[\"drop_B\"]\n    r_fams = groups[\"drop_R\"]\n\n    print(f\"\\n  Insight:\")\n    if g_fams:\n        print(f\"    → {', '.join(g_fams)}\")\n        print(f\"      High entropy attribution → likely obfuscated/packed sections.\")\n        print(f\"      Validates kênh G hoạt động đúng mục đích thiết kế.\")\n    if b_fams:\n        print(f\"    → {', '.join(b_fams)}\")\n        print(f\"      High LBP attribution → distinct spatial texture patterns.\")\n        print(f\"      Validates kênh B bổ sung thông tin cấu trúc spatial.\")\n    if r_fams:\n        print(f\"    → {', '.join(r_fams)}\")\n        print(f\"      Primarily classified by raw byte amplitude (kênh R).\")\n\n    # Per-family table cho paper\n    family_detail = family_attr.copy()\n    family_detail.columns = [\"drop_R\", \"drop_G\", \"drop_B\"]\n    family_detail[\"dominant\"] = dominant_col.map({\n        \"drop_R\": \"R\", \"drop_G\": \"G\", \"drop_B\": \"B\",\n    })\n    family_detail[\"dataset\"] = res[\"cfg\"][\"description\"]\n    all_analysis.append(family_detail.reset_index())\n\n# Gộp và lưu\nif all_analysis:\n    analysis_df = pd.concat(all_analysis, ignore_index=True)\n    analysis_path = os.path.join(OUTPUT_DIR, \"family_attribution_analysis.csv\")\n    analysis_df.to_csv(analysis_path, index=False)\n    print(f\"\\n💾 Saved full family analysis: {analysis_path}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-08T21:58:24.784857Z","iopub.execute_input":"2026-05-08T21:58:24.785219Z","iopub.status.idle":"2026-05-08T21:58:24.809833Z","shell.execute_reply.started":"2026-05-08T21:58:24.785194Z","shell.execute_reply":"2026-05-08T21:58:24.808992Z"}},"outputs":[],"execution_count":null},{"id":"a3b4c5d6","cell_type":"code","source":"# ============================================================\n# CELL 9: (Bonus) Per-dataset bar chart — Top/bottom families\n# ============================================================\n\nfor ds_name, res in all_results.items():\n    df = res[\"df\"]\n    family_attr = (\n        df.groupby(\"true_class_name\")[[\"drop_R\", \"drop_G\", \"drop_B\"]].mean()\n        .rename(columns={\"drop_R\": \"R (Byte)\",\n                          \"drop_G\": \"G (Entropy)\",\n                          \"drop_B\": \"B (LBP)\"})\n        .sort_values(\"G (Entropy)\", ascending=True)\n    )\n\n    ax = family_attr.plot(\n        kind=\"barh\",\n        figsize=(9, max(4, len(family_attr) * 0.35)),\n        color=[\"#e74c3c\", \"#2ecc71\", \"#3498db\"],\n        edgecolor=\"white\",\n        linewidth=0.5,\n    )\n    ax.set_title(\n        f\"{res['cfg']['description']} — Attribution drop per channel\",\n        fontsize=11, fontweight=\"bold\",\n    )\n    ax.set_xlabel(\"Mean attribution drop (p₀ − p^c)\")\n    ax.set_ylabel(\"Malware Family\")\n    ax.axvline(0, color=\"black\", linewidth=0.8, linestyle=\"--\")\n    plt.tight_layout()\n\n    bar_path = os.path.join(OUTPUT_DIR, f\"attribution_bar_{ds_name}.png\")\n    plt.savefig(bar_path, dpi=200, bbox_inches=\"tight\")\n    plt.show()\n    print(f\"💾 {bar_path}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-08T21:58:24.811047Z","iopub.execute_input":"2026-05-08T21:58:24.811422Z","iopub.status.idle":"2026-05-08T21:58:26.177934Z","shell.execute_reply.started":"2026-05-08T21:58:24.811396Z","shell.execute_reply":"2026-05-08T21:58:26.177218Z"}},"outputs":[],"execution_count":null},{"id":"b4c5d6e7","cell_type":"markdown","source":"---\n## Files đầu ra\n\n| File | Nội dung |\n|---|---|\n| `channel_occlusion_microsoft_rgb.csv` | Per-sample: p0, drop_R/G/B, correct |\n| `channel_occlusion_malimg_rgb.csv` | Per-sample: p0, drop_R/G/B, correct |\n| `attribution_summary.csv` | **Output 1**: mean±std per dataset×channel |\n| `channel_attribution_heatmap.png` | **Output 2**: heatmap family×channel (300dpi) |\n| `family_attribution_analysis.csv` | **Output 3**: dominant channel per family |\n| `attribution_bar_*.png` | Bar chart bổ sung |\n\n**Claim an toàn cho paper:**  \n*\"channel occlusion attribution reveals that [entropy-dominant families] rely primarily on the G channel, while [texture-dominant families] rely on B — validating that the complementary channel design is actively exploited during inference.\"*","metadata":{}}]}