{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":106680,"databundleVersionId":13374319,"sourceType":"competition"}],"dockerImageVersionId":31192,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# AIRR-ML🧬25: EDA & Convert dataset to parquet\n\n### 🔍 What it does\n- Performs **Exploratory Data Analysis (EDA)** on immunology-related datasets.\n- Checks dataset structure, size, missing values, and feature distributions.\n- Uses simple visualizations to spot anomalies and patterns.\n\n### 📦 Why Parquet?\n- Converts dataset into **Parquet format** for:\n  - ⚡ Faster read/write\n  - 📉 Smaller storage size (compression)\n  - 🔄 Better compatibility with big-data tools (Spark, Dask)\n\n### 🧪 Key Insight\n- It’s **important to test even the simplest assumptions** —  \n  ✅ Sometimes naive checks (like counting rows, checking nulls) reveal big issues.  \n  ✅ These basic steps prevent wasted effort later in ML pipelines.\n\n### 🚀 Workflow Benefits\n- Ensures data is **understood** before modeling.\n- Guarantees it’s stored in an **efficient format** for scalable experiments.\n- Builds a reproducible pipeline for biomedical ML tasks.\n\n---\n\n✨ **Takeaway:** Don’t skip the “boring” steps — even naive checks can save your project!\n","metadata":{}},{"cell_type":"code","source":"import os\n\nPATH_DATASET = \"/kaggle/input/adaptive-immune-profiling-challenge-2025\"\nPATH_TRAIN_DATASETS = os.path.join(PATH_DATASET, 'train_datasets', 'train_datasets')\ntrain_datasets = sorted(os.listdir(PATH_TRAIN_DATASETS))\nprint(train_datasets)\nPATH_TEST_DATASETS = os.path.join(PATH_DATASET, 'test_datasets', 'test_datasets')\ntest_datasets = sorted(os.listdir(PATH_TEST_DATASETS))\nprint(test_datasets)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2025-11-13T09:16:50.042973Z","iopub.execute_input":"2025-11-13T09:16:50.043909Z","iopub.status.idle":"2025-11-13T09:16:50.067979Z","shell.execute_reply.started":"2025-11-13T09:16:50.043872Z","shell.execute_reply":"2025-11-13T09:16:50.066595Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Why merge and convert to Parquet?\n\nMerging multiple small TSV files into a single DataFrame and then saving it to Parquet format offers several advantages:\n\n1.  **Efficiency in Reading/Writing**: Parquet is a columnar storage format, which means data is stored column by column. This is highly efficient for analytical queries as it allows reading only the necessary columns, significantly reducing I/O operations compared to row-oriented formats like CSV or TSV. It also typically provides better compression.\n2.  **Performance**: Reading from a single, large Parquet file is generally much faster than reading and concatenating many small TSV files repeatedly. This is especially true when dealing with big data, where file system overhead for numerous small files can be substantial.\n3.  **Data Type Preservation**: Parquet stores schema information along with the data, ensuring that data types are preserved when the data is read back, avoiding issues with type inference that can occur with text-based formats like TSV.\n4.  **Simplified Data Handling**: Instead of managing hundreds or thousands of individual TSV files, you now have one consolidated and optimized data file, simplifying data management and subsequent processing steps.\n5.  **Disk Space Savings**: Due to its columnar nature and efficient compression algorithms, Parquet files often take up significantly less disk space than their TSV counterparts.","metadata":{}},{"cell_type":"code","source":"import os\nimport shutil\nimport seaborn as sns\nimport pandas as pd\nfrom tqdm.auto import tqdm\nimport matplotlib.pyplot as plt\nfrom typing import List, Optional\n\n\ndef load_tsv_files_export_parquet(folder_path: str, output_path: str, show_hist: Optional[List[str]] = None):\n    folder = os.path.basename(folder_path) # Derive folder name from folder_path\n    # List all files in the directory\n    files = os.listdir(folder_path)\n\n    # Filter for .tsv files\n    tsv_files = [f for f in files if f.endswith('.tsv')]\n    other_files = [f.name for f in os.scandir(folder_path) if not f.name.endswith('.tsv')]\n    print(f'Loading {len(tsv_files)} .tsv files from {folder} (remaining: {other_files}).')\n\n    # Iterate through each TSV file, load it into a DataFrame, and print column names\n    dfs = []\n    for tsv_file in tqdm(tsv_files, desc=\"Loading TSV files\"):\n        file_path = os.path.join(folder_path, tsv_file)\n        file_name, _ = os.path.splitext(tsv_file)\n        try:\n            df = pd.read_csv(file_path, sep='\\t')\n            df['repertoire_id'] = file_name\n            dfs.append(df)\n        except Exception as e:\n            print(f\"Error loading {tsv_file}: {e}\")\n\n    merged_df = pd.concat(dfs, ignore_index=True)\n    del dfs # Free up memory\n\n    print(f\"Merged DataFrame shape: {merged_df.shape}\")\n    for col in merged_df.columns:\n        print(f\"Unique values in column '{col}': {len(merged_df[col].unique())}\")\n    print(\"Merged DataFrame head:\")\n    display(merged_df.head())\n\n    os.makedirs(output_path, exist_ok=True)\n    merged_df.to_parquet(f'{output_path}/{folder}.parquet')\n\n    # Plot histograms for specified columns if show_hist is provided\n    if not isinstance(show_hist, list) and not show_hist:\n        return\n    print(f\"Plotting histograms for columns: {', '.join(show_hist)}\")\n    for col in show_hist:\n        if col not in merged_df.columns:\n            print(f\"Warning: Column '{col}' not found in the DataFrame for {folder}.\")\n            continue\n        # Get all value counts\n        all_counts = merged_df[col].value_counts()\n        \n        plt.figure(figsize=(min(12, len(all_counts) * 0.3), 4)) # Adjust figure size dynamically\n        sns.barplot(x=all_counts.index, y=all_counts.values, palette='viridis')\n        plt.title(f'Value Counts for {col} in {folder}')\n        plt.xlabel(col)\n        plt.ylabel('Count')\n        plt.xticks(rotation=90, ha='center') # Rotate labels more for many categories\n        plt.grid(True)\n        plt.tight_layout()\n    plt.show()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Training dataset","metadata":{}},{"cell_type":"code","source":"# Iterate over all sub-datasets\nfor folder in tqdm(train_datasets):\n    path_dataset_ = os.path.join(PATH_TRAIN_DATASETS, folder)\n    load_tsv_files_export_parquet(\n        path_dataset_, output_path='train_dataset', show_hist=['v_call', 'j_call', 'd_call'])\n    new_meta_csv = os.path.join(\"train_dataset\", f\"{folder}-metadata.csv\")\n    shutil.copy(os.path.join(path_dataset_, \"metadata.csv\"), new_meta_csv)\n    df_meta = pd.read_csv(new_meta_csv)\n    display(df_meta)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-13T09:25:48.335074Z","iopub.execute_input":"2025-11-13T09:25:48.335401Z","execution_failed":"2025-11-13T09:26:35.662Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Test dataset","metadata":{}},{"cell_type":"code","source":"# Iterate over all sub-datasets\nfor folder in tqdm(test_datasets):\n    path_dataset_ = os.path.join(PATH_TEST_DATASETS, folder)\n    load_tsv_files_export_parquet(\n        path_dataset_, output_path='test_dataset', show_hist=['v_call', 'j_call', 'd_call'])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-11-13T09:17:05.893483Z","iopub.status.idle":"2025-11-13T09:17:05.893772Z","shell.execute_reply.started":"2025-11-13T09:17:05.893636Z","shell.execute_reply":"2025-11-13T09:17:05.893649Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Construct the path to the sample submission file\nsample_submission_path = os.path.join(PATH_DATASET, 'sample_submissions.csv')\n\n# Load the sample submission file\nsample_submission_df = pd.read_csv(sample_submission_path)\n\n# You can save this to a CSV file if needed\nsample_submission_df.to_csv('submission.csv', index=False)","metadata":{"trusted":true,"_kg_hide-input":true},"outputs":[],"execution_count":null}]}