{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.10"},"sambit":{"notebook_id":"P00","schema_version":"2.0","status":"MAIN - required"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# P00 - Build Manifest, Stratified Folds and Shards\n\n> **Trạng thái:** MAIN - required  \n> **Mục tiêu:** Tạo manifest bất biến cho toàn bộ pipeline, gán 5 fold phân tầng và 64 shard ổn định.\n\n## Kaggle setup\n\n- **Accelerator:** CPU; cần dung lượng đĩa lớn khi lập chỉ mục archive\n- **Internet:** Off sau khi đã attach đủ dataset\n- **Output dataset gợi ý:** `sambit-manifest-v1`\n\n## Input bắt buộc\n\n- BIG 2015 raw data: train.7z, test.7z, trainLabels.csv, sampleSubmission.csv\n- Hoặc thư mục đã giải nén chứa các file .bytes/.asm\n\n### Dataset cần attach\n\n- `Microsoft Malware Classification Challenge dataset (raw hoặc pre-extracted)`\n\n## Output chính\n\n- `manifest.parquet / manifest.csv`\n- `manifest_summary.json`\n- `shard_summary.csv`\n- `manifest_config.json`\n\n## Bước kế tiếp\n\n`P01_extract_section_aware_tiles_sharded.ipynb`\n\n## Quy tắc chạy và tái lập\n\n- Không thay đổi SEED, số fold hoặc manifest sau khi đã chạy các bước tiếp theo.\n- Publish toàn bộ thư mục /kaggle/working/sambit_manifest thành Kaggle Dataset.\n\n","metadata":{}},{"cell_type":"code","source":"from pathlib import Path\nimport json, os, subprocess, sys\n\n# ========================= CẤU HÌNH =========================\nCFG = {\n    # Thư mục gốc chứa dữ liệu đã Attach vào notebook.\n    \"INPUT_ROOT\": \"/kaggle/input\",\n\n    # auto | preextracted | archive\n    \"SOURCE_MODE\": \"auto\",\n\n    # Với pre-extracted: đặt đúng thư mục chứa train/test .bytes/.asm.\n    # Có thể để None để quét toàn bộ /kaggle/input, nhưng sẽ chậm hơn.\n    \"PREEXTRACTED_ROOT\": None,\n\n    # Với archive: đặt đường dẫn cụ thể nếu auto-detect không tìm đúng.\n    \"TRAIN_ARCHIVE\": None,  # ví dụ /kaggle/input/malware-classification/train.7z\n    \"TEST_ARCHIVE\": None,   # ví dụ /kaggle/input/malware-classification/test.7z\n\n    \"TRAIN_LABELS\": None,\n    \"SAMPLE_SUBMISSION\": None,\n    \"OUTPUT_ROOT\": \"/kaggle/working/sambit_manifest\",\n    \"NUM_SHARDS\": 64,\n    \"N_FOLDS\": 5,\n    \"SEED\": 2026,\n}\n\ncfg_path = Path('/kaggle/working/manifest_config.json')\ncfg_path.write_text(json.dumps(CFG, indent=2), encoding='utf-8')\nprint(cfg_path.read_text())","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"%%writefile /kaggle/working/sambit_manifest.py\nfrom __future__ import annotations\n\nimport argparse\nimport hashlib\nimport json\nimport os\nimport re\nimport shutil\nimport subprocess\nfrom pathlib import Path\nfrom typing import Dict, Iterable, Optional\n\nimport numpy as np\nimport pandas as pd\nfrom sklearn.model_selection import StratifiedKFold\n\n\ndef atomic_json(path: Path, obj: dict) -> None:\n    path.parent.mkdir(parents=True, exist_ok=True)\n    tmp = path.with_suffix(path.suffix + '.tmp')\n    tmp.write_text(json.dumps(obj, indent=2, ensure_ascii=False), encoding='utf-8')\n    os.replace(tmp, path)\n\n\ndef stable_shard(sample_id: str, num_shards: int, seed: int) -> int:\n    payload = f'{seed}:{sample_id}'.encode('utf-8')\n    value = int.from_bytes(hashlib.blake2b(payload, digest_size=8).digest(), 'little')\n    return value % num_shards\n\n\ndef find_unique(root: Path, names: Iterable[str]) -> Path:\n    lowered = {n.lower() for n in names}\n    matches = [p for p in root.rglob('*') if p.is_file() and p.name.lower() in lowered]\n    if not matches:\n        raise FileNotFoundError(f'Không tìm thấy một trong {sorted(names)} dưới {root}')\n    matches.sort(key=lambda p: (len(p.parts), str(p)))\n    return matches[0]\n\n\ndef build_file_index(root: Path, cache_path: Path) -> Dict[str, dict]:\n    if cache_path.exists():\n        print(f'[resume] Load file index: {cache_path}')\n        return json.loads(cache_path.read_text(encoding='utf-8'))\n\n    index: Dict[str, dict] = {}\n    candidates = [p for p in root.rglob('*') if p.is_file() and p.suffix.lower() in {'.bytes', '.asm'}]\n    print(f'Đang lập chỉ mục {len(candidates):,} file dưới {root}')\n    for i, path in enumerate(candidates, 1):\n        sid = path.stem\n        rec = index.setdefault(sid, {})\n        rec[path.suffix.lower().lstrip('.')] = str(path)\n        try:\n            rec[path.suffix.lower().lstrip('.') + '_size'] = path.stat().st_size\n        except OSError:\n            pass\n        if i % 2000 == 0:\n            atomic_json(cache_path, index)\n            print(f'  indexed {i:,}/{len(candidates):,}')\n    atomic_json(cache_path, index)\n    return index\n\n\ndef list_7z_members(archive: Path, cache_path: Path) -> Dict[str, dict]:\n    if cache_path.exists():\n        print(f'[resume] Load archive index: {cache_path}')\n        return json.loads(cache_path.read_text(encoding='utf-8'))\n    sevenzip = shutil.which('7z') or shutil.which('7zz') or shutil.which('7za')\n    if sevenzip is None:\n        raise RuntimeError('Không tìm thấy 7z/7zz/7za. Hãy dùng dataset pre-extracted hoặc cài p7zip.')\n\n    print(f'Đang đọc danh sách thành viên archive: {archive}')\n    proc = subprocess.run([sevenzip, 'l', '-slt', str(archive)], check=True, text=True,\n                          stdout=subprocess.PIPE, stderr=subprocess.STDOUT, errors='replace')\n    index: Dict[str, dict] = {}\n    current_path: Optional[str] = None\n    current_size: Optional[int] = None\n    for line in proc.stdout.splitlines():\n        if line.startswith('Path = '):\n            if current_path:\n                p = Path(current_path)\n                if p.suffix.lower() in {'.bytes', '.asm'}:\n                    rec = index.setdefault(p.stem, {})\n                    key = p.suffix.lower().lstrip('.')\n                    rec[key] = current_path\n                    if current_size is not None:\n                        rec[key + '_size'] = current_size\n            current_path = line[7:].strip()\n            current_size = None\n        elif line.startswith('Size = '):\n            try:\n                current_size = int(line[7:].strip())\n            except ValueError:\n                current_size = None\n    if current_path:\n        p = Path(current_path)\n        if p.suffix.lower() in {'.bytes', '.asm'}:\n            rec = index.setdefault(p.stem, {})\n            key = p.suffix.lower().lstrip('.')\n            rec[key] = current_path\n            if current_size is not None:\n                rec[key + '_size'] = current_size\n    atomic_json(cache_path, index)\n    print(f'Archive index có {len(index):,} sample ID')\n    return index\n\n\ndef save_table(df: pd.DataFrame, path_stem: Path) -> None:\n    csv_path = path_stem.with_suffix('.csv')\n    df.to_csv(csv_path, index=False)\n    try:\n        df.to_parquet(path_stem.with_suffix('.parquet'), index=False)\n    except Exception as exc:\n        print(f'[note] Không ghi được parquet ({exc}); CSV vẫn đầy đủ.')\n\n\ndef main() -> None:\n    parser = argparse.ArgumentParser()\n    parser.add_argument('--config', required=True)\n    args = parser.parse_args()\n    cfg = json.loads(Path(args.config).read_text(encoding='utf-8'))\n\n    input_root = Path(cfg.get('INPUT_ROOT', '/kaggle/input'))\n    output_root = Path(cfg.get('OUTPUT_ROOT', '/kaggle/working/sambit_manifest'))\n    output_root.mkdir(parents=True, exist_ok=True)\n    seed = int(cfg.get('SEED', 2026))\n    num_shards = int(cfg.get('NUM_SHARDS', 64))\n    n_folds = int(cfg.get('N_FOLDS', 5))\n    source_mode = cfg.get('SOURCE_MODE', 'auto').lower()\n\n    labels_path = Path(cfg['TRAIN_LABELS']) if cfg.get('TRAIN_LABELS') else find_unique(input_root, ['trainLabels.csv'])\n    sample_path = Path(cfg['SAMPLE_SUBMISSION']) if cfg.get('SAMPLE_SUBMISSION') else find_unique(input_root, ['sampleSubmission.csv'])\n    labels = pd.read_csv(labels_path)\n    sample = pd.read_csv(sample_path)\n    labels.columns = [c.strip('\"') for c in labels.columns]\n    sample.columns = [c.strip('\"') for c in sample.columns]\n    labels['Id'] = labels['Id'].astype(str)\n    sample['Id'] = sample['Id'].astype(str)\n\n    pre_root = Path(cfg['PREEXTRACTED_ROOT']) if cfg.get('PREEXTRACTED_ROOT') else None\n    train_archive = Path(cfg['TRAIN_ARCHIVE']) if cfg.get('TRAIN_ARCHIVE') else None\n    test_archive = Path(cfg['TEST_ARCHIVE']) if cfg.get('TEST_ARCHIVE') else None\n\n    if source_mode == 'auto':\n        if pre_root and pre_root.exists() and any(pre_root.rglob('*.bytes')):\n            source_mode = 'preextracted'\n        else:\n            if train_archive is None:\n                candidates = sorted(input_root.rglob('train.7z'))\n                train_archive = candidates[0] if candidates else None\n            if test_archive is None:\n                candidates = sorted(input_root.rglob('test.7z'))\n                test_archive = candidates[0] if candidates else None\n            source_mode = 'archive' if train_archive and test_archive else 'preextracted'\n            if source_mode == 'preextracted' and pre_root is None:\n                pre_root = input_root\n\n    if source_mode == 'preextracted':\n        pre_root = pre_root or input_root\n        source_index = build_file_index(pre_root, output_root / 'file_index_checkpoint.json')\n        train_index = source_index\n        test_index = source_index\n    elif source_mode == 'archive':\n        if not train_archive or not test_archive:\n            raise FileNotFoundError('Archive mode cần TRAIN_ARCHIVE và TEST_ARCHIVE.')\n        train_index = list_7z_members(train_archive, output_root / 'train_archive_index_checkpoint.json')\n        test_index = list_7z_members(test_archive, output_root / 'test_archive_index_checkpoint.json')\n    else:\n        raise ValueError(f'SOURCE_MODE không hợp lệ: {source_mode}')\n\n    train_df = labels.rename(columns={'Class': 'label'}).copy()\n    train_df['split'] = 'train'\n    test_df = sample[['Id']].copy()\n    test_df['label'] = -1\n    test_df['split'] = 'test'\n    manifest = pd.concat([train_df[['Id', 'label', 'split']], test_df], ignore_index=True)\n\n    def enrich(row: pd.Series) -> pd.Series:\n        idx = train_index if row['split'] == 'train' else test_index\n        rec = idx.get(row['Id'], {})\n        return pd.Series({\n            'bytes_path': rec.get('bytes', '') if source_mode == 'preextracted' else '',\n            'asm_path': rec.get('asm', '') if source_mode == 'preextracted' else '',\n            'bytes_member': rec.get('bytes', '') if source_mode == 'archive' else '',\n            'asm_member': rec.get('asm', '') if source_mode == 'archive' else '',\n            'bytes_size': int(rec.get('bytes_size', -1)),\n            'asm_size': int(rec.get('asm_size', -1)),\n            'has_bytes': bool(rec.get('bytes')),\n            'has_asm': bool(rec.get('asm')),\n        })\n\n    extra = manifest.apply(enrich, axis=1)\n    manifest = pd.concat([manifest, extra], axis=1)\n    manifest['shard'] = manifest['Id'].map(lambda x: stable_shard(x, num_shards, seed)).astype(np.int16)\n    manifest['fold'] = -1\n    train_mask = manifest['split'].eq('train')\n    train_positions = np.flatnonzero(train_mask.to_numpy())\n    y = manifest.loc[train_mask, 'label'].to_numpy()\n    min_class_count = int(pd.Series(y).value_counts().min()) if len(y) else 0\n    effective_folds = min(n_folds, min_class_count)\n    if effective_folds >= 2:\n        skf = StratifiedKFold(n_splits=effective_folds, shuffle=True, random_state=seed)\n        for fold, (_, val_idx) in enumerate(skf.split(np.zeros(len(y)), y)):\n            manifest.loc[train_positions[val_idx], 'fold'] = fold\n    else:\n        # Chỉ dành cho smoke test cực nhỏ; dữ liệu BIG 2015 thật đủ mẫu cho 5 folds.\n        manifest.loc[train_positions, 'fold'] = 0\n        effective_folds = 1\n    manifest['fold'] = manifest['fold'].astype(np.int8)\n\n    save_table(manifest, output_root / 'manifest')\n    shards_root = output_root / 'shards'\n    shards_root.mkdir(exist_ok=True)\n    for shard in range(num_shards):\n        part = manifest[manifest['shard'].eq(shard)].copy()\n        part.to_csv(shards_root / f'shard_{shard:03d}.csv', index=False)\n\n    summary = {\n        'source_mode': source_mode,\n        'labels_path': str(labels_path),\n        'sample_submission': str(sample_path),\n        'train_archive': str(train_archive) if train_archive else None,\n        'test_archive': str(test_archive) if test_archive else None,\n        'preextracted_root': str(pre_root) if pre_root else None,\n        'num_rows': int(len(manifest)),\n        'num_train': int(train_mask.sum()),\n        'num_test': int((~train_mask).sum()),\n        'num_shards': num_shards,\n        'n_folds_requested': n_folds,\n        'n_folds_effective': effective_folds,\n        'missing_bytes': int((~manifest['has_bytes']).sum()),\n        'missing_asm': int((~manifest['has_asm']).sum()),\n        'class_counts': manifest.loc[train_mask, 'label'].value_counts().sort_index().to_dict(),\n        'shard_counts': manifest.groupby(['split', 'shard']).size().unstack(0, fill_value=0).to_dict(),\n    }\n    atomic_json(output_root / 'manifest_summary.json', summary)\n    print(json.dumps(summary, indent=2, ensure_ascii=False))\n    if summary['missing_bytes']:\n        print('[warning] Có sample thiếu .bytes. Kiểm tra đúng dataset/đường dẫn trước bước 01.')\n\n\nif __name__ == '__main__':\n    main()","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"subprocess.run([\n    sys.executable, '/kaggle/working/sambit_manifest.py',\n    '--config', '/kaggle/working/manifest_config.json'\n], check=True)","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import pandas as pd\nout = Path(CFG['OUTPUT_ROOT'])\nmanifest_path = out / 'manifest.parquet'\nif not manifest_path.exists():\n    manifest_path = out / 'manifest.csv'\nmanifest = pd.read_parquet(manifest_path) if manifest_path.suffix == '.parquet' else pd.read_csv(manifest_path)\nprint(manifest_path)\ndisplay(manifest.head())\ndisplay(manifest.groupby(['split', 'shard']).size().unstack(0, fill_value=0).describe())\ndisplay(manifest[manifest.split.eq('train')].groupby(['fold', 'label']).size().unstack(fill_value=0))\nprint((out / 'manifest_summary.json').read_text())","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Sau khi chạy xong\n\n1. Bấm **Save Version → Save & Run All**.\n2. Tạo một Kaggle Dataset từ output `sambit_manifest` hoặc dùng trực tiếp Notebook Output.\n3. Gắn output này vào tất cả notebook 01–05.\n4. Không sửa `NUM_SHARDS`, `N_FOLDS`, `SEED` giữa các máy.\n\nNếu dùng nhiều thành viên trong team, tất cả phải dùng đúng bản manifest này. Không dùng nhiều tài khoản cá nhân để né quota hoặc quy định cuộc thi.","metadata":{}},{"cell_type":"markdown","source":"## Checklist sau khi chạy P00\n\n1. Kiểm tra toàn bộ assert/preflight đều pass.\n2. Kiểm tra số dòng, số fold, probability sum và file summary theo phần **Output chính** ở đầu notebook.\n3. Save Version bằng **Save & Run All**.\n4. Publish output thành Kaggle Dataset `sambit-manifest-v1` nếu bước sau cần attach.\n5. Ghi public notebook URL và public artifact URL vào `source/notebooks.txt`.\n6. Không đổi manifest/fold hoặc trộn OOF và test artifact từ hai run khác nhau.\n","metadata":{}}]}