{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":108394,"databundleVersionId":13172641,"sourceType":"competition"}],"dockerImageVersionId":31089,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# **Introduction**\n\n\nWelcome to the **Great Daxinzhuang Pottery Puzzle AI Challenge**!  \n\nIn this groundbreaking Kaggle competition, we explore how **machine learning** can revolutionize **archaeology**, especially in solving one of its most time-consuming tasks — the matching and reconstruction of ancient pottery sherds.\n\nParticipants will work with a truly exceptional archaeological dataset from **Daxinzhuang**, a major Shang Dynasty site in eastern China. The data includes nearly **20,000 pottery fragments** from **pit H690**, an ancient granary turned refuse pit rich in cultural artifacts — from pottery to animal bones and even gold foil. This offers a unique glimpse into the Shang Dynasty’s expansion and cultural interactions.\n\nHowever, traditional manual classification and reconstruction of these pottery sherds is **extremely labor-intensive** and limited. This is where **AI** steps in — to assist, accelerate, and even enhance the work of archaeologists.\n\n## **Challenge Goals**\n\nThis competition focuses on two core tasks:\n\n1. **Fracture Edge Matching and Assembly**  \n   Develop models to automatically match pottery sherds based on fracture edges and metadata, reconstructing their adjacency relationships.\n\n2. **3D Vessel Reconstruction**  \n   Use the matched sherds to digitally restore the original vessels as accurately as possible.","metadata":{"_kg_hide-input":false}},{"cell_type":"markdown","source":"# **Part One: EDA**\n## **1.1 Dataset Overview**\n\n","metadata":{}},{"cell_type":"code","source":"#1.1 Dataset Overview\n\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom IPython.display import display, HTML\n\n# --- Plotting Configuration ---\n# Set the visual style for the plots\nsns.set(style=\"whitegrid\")\n\n# Set font to display Chinese characters correctly in plots. \n# 'SimHei' is a commonly available font in Kaggle environments.\nplt.rcParams['font.sans-serif'] = ['SimHei', 'DejaVu Sans', 'sans-serif']\n\n# Ensure that minus signs are displayed correctly\nplt.rcParams['axes.unicode_minus'] = False  \n\n# --- Load the Data ---\n# Define the correct path to the CSV file within the Kaggle environment\nfile_path = '/kaggle/input/h690/h690/jd_sherds_info.csv'\n\ntry:\n    # Load the dataset into a pandas DataFrame\n    df_info = pd.read_csv(file_path)\n    \n    \n    # Display beautiful data preview\n    print(\"📊 DATASET PREVIEW - First 5 Rows\")\n    print(\"=\" * 60)\n    \n    # Create a styled table for the first 5 rows\n    styled_head = df_info.head().style.set_properties(**{\n        'background-color': '#f8f9fa',\n        'color': '#212529',\n        'border': '1px solid #dee2e6',\n        'text-align': 'center'\n    }).set_table_styles([\n        {'selector': 'th', 'props': [\n            ('background-color', '#007bff'),\n            ('color', 'white'),\n            ('font-weight', 'bold'),\n            ('text-align', 'center'),\n            ('border', '1px solid #0056b3')\n        ]},\n        {'selector': 'td', 'props': [\n            ('border', '1px solid #dee2e6'),\n            ('padding', '8px')\n        ]},\n        {'selector': 'table', 'props': [\n            ('border-collapse', 'collapse'),\n            ('margin', '10px auto'),\n            ('width', '100%')\n        ]}\n    ])\n    \n    display(styled_head)\n    \n    print(\"\\n\" + \"=\" * 60)\n    print(\"📋 DATASET INFORMATION SUMMARY\")\n    print(\"=\" * 60)\n    \n    # Create a beautiful info summary\n    info_data = {\n        'Metric': ['Total Rows', 'Total Columns', 'Memory Usage', 'Data Types'],\n        'Value': [\n            f\"{len(df_info):,}\",\n            f\"{len(df_info.columns)}\",\n            f\"{df_info.memory_usage(deep=True).sum() / 1024:.2f} KB\",\n            f\"{len(df_info.dtypes.unique())} unique types\"\n        ]\n    }\n    \n    info_df = pd.DataFrame(info_data)\n    \n    styled_info = info_df.style.set_properties(**{\n        'background-color': '#f8f9fa',\n        'color': '#495057',\n        'border': '1px solid #dee2e6',\n        'text-align': 'center',\n        'font-size': '14px'\n    }).set_table_styles([\n        {'selector': 'th', 'props': [\n            ('background-color', '#28a745'),\n            ('color', 'white'),\n            ('font-weight', 'bold'),\n            ('text-align', 'center'),\n            ('border', '1px solid #1e7e34')\n        ]},\n        {'selector': 'td', 'props': [\n            ('border', '1px solid #dee2e6'),\n            ('padding', '10px'),\n            ('font-weight', '500')\n        ]},\n        {'selector': 'table', 'props': [\n            ('border-collapse', 'collapse'),\n            ('margin', '10px auto'),\n            ('width', '60%')\n        ]}\n    ])\n    \n    display(styled_info)\n    \n    # Column information table\n    print(\"\\n\" + \"=\" * 60)\n    print(\"🔍 COLUMN DETAILS\")\n    print(\"=\" * 60)\n    \n    column_info = pd.DataFrame({\n        'Data Type': df_info.dtypes.astype(str),\n        'Non-Null Count': df_info.count(),\n        'Null Count': df_info.isnull().sum(),\n        'Null Percentage': (df_info.isnull().sum() / len(df_info) * 100).round(2)\n    })\n    \n    styled_columns = column_info.style.set_properties(**{\n        'background-color': '#fff3cd',\n        'color': '#856404',\n        'border': '1px solid #ffeaa7',\n        'text-align': 'center'\n    }).set_table_styles([\n        {'selector': 'th', 'props': [\n            ('background-color', '#ffc107'),\n            ('color', '#212529'),\n            ('font-weight', 'bold'),\n            ('text-align', 'center'),\n            ('border', '1px solid #e0a800')\n        ]},\n        {'selector': 'td', 'props': [\n            ('border', '1px solid #ffeaa7'),\n            ('padding', '8px')\n        ]},\n        {'selector': 'table', 'props': [\n            ('border-collapse', 'collapse'),\n            ('margin', '10px auto'),\n            ('width', '100%')\n        ]}\n    ]).highlight_null(color='#f8d7da')\n    \n    display(styled_columns)\n    \n    \nexcept FileNotFoundError:\n    print(\"❌ Error: Could not find the file at '{}'\".format(file_path))\n    print(\"💡 Please double-check the input data path in your Kaggle Notebook.\")\n    # In case of an error, create an empty DataFrame to prevent subsequent code from crashing.\n    df_info = pd.DataFrame()\n    \n    # Display error message in a styled format\n    error_html = \"\"\"\n    <div style=\"background-color: #f8d7da; color: #721c24; padding: 15px; border: 1px solid #f5c6cb; border-radius: 5px; margin: 10px 0;\">\n        <h3 style=\"margin-top: 0;\">⚠️ File Not Found Error</h3>\n        <p><strong>Path:</strong> {}</p>\n        <p><strong>Solution:</strong> Please verify the file path and ensure the dataset is uploaded correctly.</p>\n    </div>\n    \"\"\".format(file_path)\n    \n    display(HTML(error_html))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-30T10:00:47.313811Z","iopub.execute_input":"2025-07-30T10:00:47.314053Z","iopub.status.idle":"2025-07-30T10:00:48.879909Z","shell.execute_reply.started":"2025-07-30T10:00:47.314033Z","shell.execute_reply":"2025-07-30T10:00:48.878807Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"- The table contains a total of 35,375 records and 10 fields, covering key information such as image ID, pottery shard ID, and part information\n- There are a large number of missing values (approximately 96.79%) in the fields type and type_C.","metadata":{}},{"cell_type":"markdown","source":"## **1.2 Data Integrity Check**","metadata":{}},{"cell_type":"code","source":"#Data Integrity Check\n\nimport pandas as pd\nfrom IPython.display import display, HTML\nimport warnings\nwarnings.filterwarnings('ignore')\n\ndef data_integrity_check(df):\n    \"\"\"\n    Performs data integrity checks with simple, consistent formatting\n    \"\"\"\n    if df.empty:\n        print(\"❌ No data available for integrity check.\")\n        return\n    \n    \n    # Check 1: Missing Values Analysis\n    print(f\"📊 Missing Values Analysis:\")\n    missing_values = df.isnull().sum()\n    total_missing = missing_values.sum()\n    \n    if total_missing == 0:\n        print(\"   • No missing values found in the dataset\")\n    else:\n        print(f\"   • Found missing values in {(missing_values > 0).sum()} columns:\")\n        for col, count in missing_values[missing_values > 0].items():\n            percentage = (count / len(df)) * 100\n            print(f\"     - {col}: {count:,} ({percentage:.1f}%)\")\n    \n    # Check 2: Sherd Side Completeness\n    if 'sherd_id' in df.columns and 'image_side' in df.columns:\n        sherd_sides = df.groupby('sherd_id')['image_side'].nunique()\n        incomplete_sherds = sherd_sides[sherd_sides < 2]\n        \n        print(f\"\\n📊 Sherd Side Completeness:\")\n        if len(incomplete_sherds) == 0:\n            print(\"   • All sherd_ids have records for both interior and exterior sides\")\n        else:\n            print(f\"   • Warning: Found {len(incomplete_sherds):,} sherd_ids missing one side\")\n            print(\"   • Examples of incomplete sherds:\")\n            for sherd_id in incomplete_sherds.head(5).index:\n                available_sides = df[df['sherd_id'] == sherd_id]['image_side'].unique()\n                print(f\"     - {sherd_id}: only has {', '.join(available_sides)}\")\n    \n    # Check 3: Duplicate Records Check\n    if 'sherd_id' in df.columns and 'image_side' in df.columns:\n        side_counts = df.groupby('sherd_id')['image_side'].count()\n        multiple_entries = side_counts[side_counts > 2]\n        \n        print(f\"\\n📊 Duplicate Records Check:\")\n        if len(multiple_entries) == 0:\n            print(\"   • Each sherd_id has at most two records (one interior, one exterior)\")\n        else:\n            print(f\"   • Warning: Found {len(multiple_entries):,} sherd_ids with more than two records\")\n            print(\"   • Examples of sherds with multiple records:\")\n            for sherd_id, count in multiple_entries.head(5).items():\n                print(f\"     - {sherd_id}: {count} records\")\n    \n\n# Execute the simple data integrity check\nif not df_info.empty:\n    data_integrity_check(df_info)\nelse:\n    print(\"❌ No data available for integrity check.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-30T10:00:48.880984Z","iopub.execute_input":"2025-07-30T10:00:48.881417Z","iopub.status.idle":"2025-07-30T10:00:48.972826Z","shell.execute_reply.started":"2025-07-30T10:00:48.881393Z","shell.execute_reply":"2025-07-30T10:00:48.971465Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **1.3 Univariate Distribution Analysis**\n### **Vessel Types Distribution**","metadata":{}},{"cell_type":"code","source":"# Vessel Types Distribution Analysis\n\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport pandas as pd\nimport matplotlib.patches as mpatches\nfrom matplotlib.patches import Rectangle\nimport warnings\nfrom IPython.display import display, HTML\nwarnings.filterwarnings('ignore')\n\ndef analyze_vessel_types(df):\n    \"\"\"\n    Analyze vessel type distribution\n    \"\"\"\n    if df.empty or 'type' not in df.columns:\n        print(\"❌ DataFrame is empty or 'type' column not found.\")\n        return\n    \n    \n    # Calculate frequency and percentage\n    counts = df['type'].value_counts()\n    percentages = df['type'].value_counts(normalize=True) * 100\n    \n    # Create a beautiful summary table\n    dist_summary = pd.DataFrame({\n        'Category': counts.index,\n        'Frequency': counts.values,\n        'Percentage (%)': percentages.round(2).values,\n        'Cumulative %': percentages.cumsum().round(2).values\n    })\n    \n    # Style the summary table\n    styled_summary = dist_summary.style.set_properties(**{\n        'background-color': '#f8f9fa',\n        'color': '#495057',\n        'border': '1px solid #dee2e6',\n        'text-align': 'center',\n        'font-size': '12px'\n    }).set_table_styles([\n        {'selector': 'th', 'props': [\n            ('background-color', '#17a2b8'),\n            ('color', 'white'),\n            ('font-weight', 'bold'),\n            ('text-align', 'center'),\n            ('border', '1px solid #138496'),\n            ('padding', '8px')\n        ]},\n        {'selector': 'td', 'props': [\n            ('border', '1px solid #dee2e6'),\n            ('padding', '6px'),\n            ('font-weight', '500')\n        ]},\n        {'selector': 'table', 'props': [\n            ('border-collapse', 'collapse'),\n            ('margin', '10px auto'),\n            ('width', '90%'),\n            ('box-shadow', '0 2px 4px rgba(0,0,0,0.1)')\n        ]}\n    ]).format({'Percentage (%)': '{:.2f}%', 'Cumulative %': '{:.2f}%'})\n    \n    display(styled_summary)\n    \n    # Create the plot\n    fig, ax = plt.subplots(1, 1, figsize=(12, 8))\n    \n    # Custom color palette\n    colors = ['#FF6B6B', '#4ECDC4', '#45B7D1', '#96CEB4', '#FFEAA7', '#DDA0DD', '#98D8C8', '#F7DC6F', '#FF9999', '#87CEEB', '#FFB6C1', '#98FB98', '#F0E68C', '#DDA0DD', '#B0E0E6', '#FFA07A', '#20B2AA', '#87CEFA', '#9370DB', '#3CB371']\n    \n    # Create gradient effect for bars\n    bars = ax.barh(range(len(counts)), counts.values, \n                   color=colors[:len(counts)], \n                   height=0.7,\n                   edgecolor='white', \n                   linewidth=1.5)\n    \n    # Set labels and ticks\n    ax.set_yticks(range(len(counts)))\n    ax.set_yticklabels(counts.index, fontsize=10, fontweight='600')\n    ax.set_xlabel('Frequency Count', fontsize=12, fontweight='bold')\n    ax.set_title('Vessel Type Distribution', fontsize=14, fontweight='bold', pad=20)\n    \n    # Add value annotations on bars with improved styling\n    for i, (bar, count, pct) in enumerate(zip(bars, counts.values, percentages.values)):\n        width = bar.get_width()\n        # Position text inside bar if bar is long enough, otherwise outside\n        if width > max(counts) * 0.1:\n            ax.text(width * 0.95, bar.get_y() + bar.get_height()/2, \n                    f'{count:,} ({pct:.1f}%)', \n                    ha='right', va='center', fontweight='bold', fontsize=9,\n                    color='white', \n                    bbox=dict(boxstyle='round,pad=0.2', facecolor='black', alpha=0.7))\n        else:\n            ax.text(width + max(counts) * 0.01, bar.get_y() + bar.get_height()/2, \n                    f'{count:,} ({pct:.1f}%)', \n                    ha='left', va='center', fontweight='bold', fontsize=9,\n                    bbox=dict(boxstyle='round,pad=0.2', facecolor='white', alpha=0.9, edgecolor='gray'))\n    \n    # Enhanced grid and styling\n    ax.grid(True, alpha=0.3, linestyle='--', axis='x')\n    ax.set_axisbelow(True)\n    ax.spines['top'].set_visible(False)\n    ax.spines['right'].set_visible(False)\n    ax.spines['left'].set_linewidth(2)\n    ax.spines['bottom'].set_linewidth(2)\n    ax.spines['left'].set_color('#333333')\n    ax.spines['bottom'].set_color('#333333')\n    \n    # Add subtle background color\n    ax.set_facecolor('#fafafa')\n    \n    plt.tight_layout()\n    plt.show()\n    \n    # Print key insights\n    total_count = len(df)\n    unique_categories = len(counts)\n    most_common = counts.index[0]\n    most_common_pct = percentages.iloc[0]\n    \n    print(f\"🔍 KEY INSIGHTS:\")\n    print(f\"   • Most frequent category: '{most_common}' ({most_common_pct:.1f}%)\")\n    print(f\"   • Total unique categories: {unique_categories}\")\n    print(f\"   • Data coverage: {total_count:,} records analyzed\")\n\n# Execute analysis\nif not df_info.empty:\n    analyze_vessel_types(df_info)\nelse:\n    print(\"❌ No data available for analysis. Please check your data loading process.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-30T10:15:23.695169Z","iopub.execute_input":"2025-07-30T10:15:23.695663Z","iopub.status.idle":"2025-07-30T10:15:24.434084Z","shell.execute_reply.started":"2025-07-30T10:15:23.695624Z","shell.execute_reply":"2025-07-30T10:15:24.432944Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### **Parts Distribution**","metadata":{}},{"cell_type":"code","source":"# Parts Distribution Analysis\n\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport pandas as pd\nimport matplotlib.patches as mpatches\nfrom matplotlib.patches import Rectangle\nimport warnings\nfrom IPython.display import display, HTML\nwarnings.filterwarnings('ignore')\n\ndef analyze_parts(df):\n    \"\"\"\n    Analyze part distribution\n    \"\"\"\n    if df.empty or 'part' not in df.columns:\n        print(\"❌ DataFrame is empty or 'part' column not found.\")\n        return\n\n    \n    # Calculate frequency and percentage\n    counts = df['part'].value_counts()\n    percentages = df['part'].value_counts(normalize=True) * 100\n    \n    # Create a beautiful summary table\n    dist_summary = pd.DataFrame({\n        'Category': counts.index,\n        'Frequency': counts.values,\n        'Percentage (%)': percentages.round(2).values,\n        'Cumulative %': percentages.cumsum().round(2).values\n    })\n    \n    # Style the summary table\n    styled_summary = dist_summary.style.set_properties(**{\n        'background-color': '#f8f9fa',\n        'color': '#495057',\n        'border': '1px solid #dee2e6',\n        'text-align': 'center',\n        'font-size': '12px'\n    }).set_table_styles([\n        {'selector': 'th', 'props': [\n            ('background-color', '#28a745'),  # Green theme for parts\n            ('color', 'white'),\n            ('font-weight', 'bold'),\n            ('text-align', 'center'),\n            ('border', '1px solid #1e7e34'),\n            ('padding', '8px')\n        ]},\n        {'selector': 'td', 'props': [\n            ('border', '1px solid #dee2e6'),\n            ('padding', '6px'),\n            ('font-weight', '500')\n        ]},\n        {'selector': 'table', 'props': [\n            ('border-collapse', 'collapse'),\n            ('margin', '10px auto'),\n            ('width', '90%'),\n            ('box-shadow', '0 2px 4px rgba(0,0,0,0.1)')\n        ]}\n    ]).format({'Percentage (%)': '{:.2f}%', 'Cumulative %': '{:.2f}%'})\n    \n    display(styled_summary)\n    \n    # Create the plot\n    fig, ax = plt.subplots(1, 1, figsize=(12, 8))\n    \n    # Custom color palette - different from vessel types\n    colors = ['#28a745', '#20c997', '#17a2b8', '#6f42c1', '#e83e8c', '#fd7e14', '#ffc107', '#6c757d', '#dc3545', '#007bff', '#198754', '#0dcaf0', '#6610f2', '#d63384', '#fd7e14', '#ffca2c', '#6c757d', '#dc3545', '#0d6efd', '#198754']\n    \n    # Create gradient effect for bars\n    bars = ax.barh(range(len(counts)), counts.values, \n                   color=colors[:len(counts)], \n                   height=0.7,\n                   edgecolor='white', \n                   linewidth=1.5)\n    \n    # Set labels and ticks\n    ax.set_yticks(range(len(counts)))\n    ax.set_yticklabels(counts.index, fontsize=10, fontweight='600')\n    ax.set_xlabel('Frequency Count', fontsize=12, fontweight='bold')\n    ax.set_title('Part Distribution', fontsize=14, fontweight='bold', pad=20)\n    \n    # Add value annotations on bars with improved styling\n    for i, (bar, count, pct) in enumerate(zip(bars, counts.values, percentages.values)):\n        width = bar.get_width()\n        # Position text inside bar if bar is long enough, otherwise outside\n        if width > max(counts) * 0.1:\n            ax.text(width * 0.95, bar.get_y() + bar.get_height()/2, \n                    f'{count:,} ({pct:.1f}%)', \n                    ha='right', va='center', fontweight='bold', fontsize=9,\n                    color='white', \n                    bbox=dict(boxstyle='round,pad=0.2', facecolor='black', alpha=0.7))\n        else:\n            ax.text(width + max(counts) * 0.01, bar.get_y() + bar.get_height()/2, \n                    f'{count:,} ({pct:.1f}%)', \n                    ha='left', va='center', fontweight='bold', fontsize=9,\n                    bbox=dict(boxstyle='round,pad=0.2', facecolor='white', alpha=0.9, edgecolor='gray'))\n    \n    # Enhanced grid and styling\n    ax.grid(True, alpha=0.3, linestyle='--', axis='x')\n    ax.set_axisbelow(True)\n    ax.spines['top'].set_visible(False)\n    ax.spines['right'].set_visible(False)\n    ax.spines['left'].set_linewidth(2)\n    ax.spines['bottom'].set_linewidth(2)\n    ax.spines['left'].set_color('#333333')\n    ax.spines['bottom'].set_color('#333333')\n    \n    # Add subtle background color\n    ax.set_facecolor('#fafafa')\n    \n    plt.tight_layout()\n    plt.show()\n    \n    # Print key insights\n    total_count = len(df)\n    unique_categories = len(counts)\n    most_common = counts.index[0]\n    most_common_pct = percentages.iloc[0]\n    \n    print(f\"🔍 KEY INSIGHTS:\")\n    print(f\"   • Most frequent category: '{most_common}' ({most_common_pct:.1f}%)\")\n    print(f\"   • Total unique categories: {unique_categories}\")\n    print(f\"   • Data coverage: {total_count:,} records analyzed\")\n\n# Execute analysis\nif not df_info.empty:\n    analyze_parts(df_info)\nelse:\n    print(\"❌ No data available for analysis. Please check your data loading process.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-30T10:15:40.802374Z","iopub.execute_input":"2025-07-30T10:15:40.802753Z","iopub.status.idle":"2025-07-30T10:15:41.302843Z","shell.execute_reply.started":"2025-07-30T10:15:40.802730Z","shell.execute_reply":"2025-07-30T10:15:41.301545Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### **Units Distribution**","metadata":{}},{"cell_type":"code","source":"# Units Distribution Analysis\n\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport pandas as pd\nimport matplotlib.patches as mpatches\nfrom matplotlib.patches import Rectangle\nimport warnings\nfrom IPython.display import display, HTML\nwarnings.filterwarnings('ignore')\n\ndef analyze_units(df):\n    \"\"\"\n    Analyze unit distribution\n    \"\"\"\n    if df.empty or 'unit' not in df.columns:\n        print(\"❌ DataFrame is empty or 'unit' column not found.\")\n        return\n    \n    \n    # Calculate frequency and percentage\n    counts = df['unit'].value_counts()\n    percentages = df['unit'].value_counts(normalize=True) * 100\n    \n    # Create a beautiful summary table\n    dist_summary = pd.DataFrame({\n        'Category': counts.index,\n        'Frequency': counts.values,\n        'Percentage (%)': percentages.round(2).values,\n        'Cumulative %': percentages.cumsum().round(2).values\n    })\n    \n    # Style the summary table\n    styled_summary = dist_summary.style.set_properties(**{\n        'background-color': '#f8f9fa',\n        'color': '#495057',\n        'border': '1px solid #dee2e6',\n        'text-align': 'center',\n        'font-size': '12px'\n    }).set_table_styles([\n        {'selector': 'th', 'props': [\n            ('background-color', '#6f42c1'),  # Purple theme for units\n            ('color', 'white'),\n            ('font-weight', 'bold'),\n            ('text-align', 'center'),\n            ('border', '1px solid #59359a'),\n            ('padding', '8px')\n        ]},\n        {'selector': 'td', 'props': [\n            ('border', '1px solid #dee2e6'),\n            ('padding', '6px'),\n            ('font-weight', '500')\n        ]},\n        {'selector': 'table', 'props': [\n            ('border-collapse', 'collapse'),\n            ('margin', '10px auto'),\n            ('width', '90%'),\n            ('box-shadow', '0 2px 4px rgba(0,0,0,0.1)')\n        ]}\n    ]).format({'Percentage (%)': '{:.2f}%', 'Cumulative %': '{:.2f}%'})\n    \n    display(styled_summary)\n    \n    # Create the plot\n    fig, ax = plt.subplots(1, 1, figsize=(12, 8))\n    \n    # Custom color palette - different from previous analyses\n    colors = ['#6f42c1', '#e83e8c', '#fd7e14', '#20c997', '#17a2b8', '#ffc107', '#dc3545', '#198754', '#0dcaf0', '#6610f2', '#d63384', '#fd7e14', '#ffca2c', '#6c757d', '#0d6efd', '#198754', '#20c997', '#6f42c1', '#e83e8c', '#fd7e14']\n    \n    # Create gradient effect for bars\n    bars = ax.barh(range(len(counts)), counts.values, \n                   color=colors[:len(counts)], \n                   height=0.7,\n                   edgecolor='white', \n                   linewidth=1.5)\n    \n    # Set labels and ticks\n    ax.set_yticks(range(len(counts)))\n    ax.set_yticklabels(counts.index, fontsize=10, fontweight='600')\n    ax.set_xlabel('Frequency Count', fontsize=12, fontweight='bold')\n    ax.set_title('Unit Distribution', fontsize=14, fontweight='bold', pad=20)\n    \n    # Add value annotations on bars with improved styling\n    for i, (bar, count, pct) in enumerate(zip(bars, counts.values, percentages.values)):\n        width = bar.get_width()\n        # Position text inside bar if bar is long enough, otherwise outside\n        if width > max(counts) * 0.1:\n            ax.text(width * 0.95, bar.get_y() + bar.get_height()/2, \n                    f'{count:,} ({pct:.1f}%)', \n                    ha='right', va='center', fontweight='bold', fontsize=9,\n                    color='white', \n                    bbox=dict(boxstyle='round,pad=0.2', facecolor='black', alpha=0.7))\n        else:\n            ax.text(width + max(counts) * 0.01, bar.get_y() + bar.get_height()/2, \n                    f'{count:,} ({pct:.1f}%)', \n                    ha='left', va='center', fontweight='bold', fontsize=9,\n                    bbox=dict(boxstyle='round,pad=0.2', facecolor='white', alpha=0.9, edgecolor='gray'))\n    \n    # Enhanced grid and styling\n    ax.grid(True, alpha=0.3, linestyle='--', axis='x')\n    ax.set_axisbelow(True)\n    ax.spines['top'].set_visible(False)\n    ax.spines['right'].set_visible(False)\n    ax.spines['left'].set_linewidth(2)\n    ax.spines['bottom'].set_linewidth(2)\n    ax.spines['left'].set_color('#333333')\n    ax.spines['bottom'].set_color('#333333')\n    \n    # Add subtle background color\n    ax.set_facecolor('#fafafa')\n    \n    plt.tight_layout()\n    plt.show()\n    \n    # Print key insights\n    total_count = len(df)\n    unique_categories = len(counts)\n    most_common = counts.index[0]\n    most_common_pct = percentages.iloc[0]\n    \n    print(f\"🔍 KEY INSIGHTS:\")\n    print(f\"   • Most frequent category: '{most_common}' ({most_common_pct:.1f}%)\")\n    print(f\"   • Total unique categories: {unique_categories}\")\n    print(f\"   • Data coverage: {total_count:,} records analyzed\")\n\n# Execute analysis\nif not df_info.empty:\n    analyze_units(df_info)\nelse:\n    print(\"❌ No data available for analysis. Please check your data loading process.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-30T10:16:20.552458Z","iopub.execute_input":"2025-07-30T10:16:20.552839Z","iopub.status.idle":"2025-07-30T10:16:21.218352Z","shell.execute_reply.started":"2025-07-30T10:16:20.552817Z","shell.execute_reply":"2025-07-30T10:16:21.217331Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"- The data is clearly concentrated on certain specific types (such as \"li\"), parts (such as \"body\"), and units (such as \"L08\"). This imbalance needs to be addressed during the analysis, modeling, or training phases, for example, through resampling, weighting, or specifically focusing on the main class","metadata":{}},{"cell_type":"markdown","source":"## **1.4 Bivariate and Multivariate Relationship Exploration**\nThis part focus on analyzing the relationship between variables with a high missing rate (**type**) and other features to help us better understand the data structure and identify potential patterns or anomalies. \n### **Vessel Type × Part**","metadata":{}},{"cell_type":"code","source":"# Vessel Type × Part Composition Analysis\n\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport numpy as np\nfrom IPython.display import display, HTML\nimport warnings\nwarnings.filterwarnings('ignore')\n\ndef analyze_vessel_part_composition(df, min_frequency=1):\n    \"\"\"\n    Analyze vessel type × part composition relationships\n    \"\"\"\n    if df.empty:\n        print(\"❌ DataFrame is empty.\")\n        return\n    \n    print(\"VESSEL TYPE × PART COMPOSITION ANALYSIS\")\n    \n    # Filter low-frequency categories\n    type_counts = df['type'].value_counts()\n    frequent_types = type_counts[type_counts >= min_frequency].index\n    df_filtered = df[df['type'].isin(frequent_types)]\n    \n    print(f\"📊 Analysis Info:\")\n    print(f\"   • Vessel types analyzed: {len(frequent_types)} (≥{min_frequency} occurrences)\")\n    print(f\"   • Records included: {len(df_filtered):,} out of {len(df):,} ({len(df_filtered)/len(df)*100:.1f}%)\")\n    \n    # Create cross-tabulation\n    crosstab = pd.crosstab(df_filtered['type'], df_filtered['part'])\n    \n    # Sort by total frequency (most common vessels first)\n    vessel_totals = crosstab.sum(axis=1).sort_values(ascending=True)\n    crosstab_sorted = crosstab.loc[vessel_totals.index]\n    \n    # Create the visualization\n    fig, ax = plt.subplots(1, 1, figsize=(14, 8))\n    \n    # Plot: Absolute counts (stacked horizontal bar)\n    colors = ['#FF6B6B', '#4ECDC4', '#45B7D1', '#96CEB4', '#FFEAA7', '#DDA0DD', \n             '#98D8C8', '#F7DC6F', '#FF9999', '#87CEEB', '#FFB6C1', '#98FB98']\n    \n    crosstab_sorted.plot(kind='barh', stacked=True, ax=ax, \n                        color=colors[:len(crosstab_sorted.columns)], \n                        width=0.8)\n    \n    ax.set_title('Vessel Type × Part Composition', fontsize=16, fontweight='bold', pad=20)\n    ax.set_xlabel('Number of Fragments', fontsize=12, fontweight='bold')\n    ax.set_ylabel('Vessel Type', fontsize=12, fontweight='bold')\n    ax.legend(title='Part', bbox_to_anchor=(1.05, 1), loc='upper left', fontsize=9)\n    ax.grid(True, alpha=0.3, axis='x')\n    ax.set_facecolor('#fafafa')\n    \n    # Enhanced styling\n    ax.spines['top'].set_visible(False)\n    ax.spines['right'].set_visible(False)\n    ax.spines['left'].set_linewidth(2)\n    ax.spines['bottom'].set_linewidth(2)\n    ax.spines['left'].set_color('#333333')\n    ax.spines['bottom'].set_color('#333333')\n    \n    # Add total labels\n    for i, (idx, row) in enumerate(crosstab_sorted.iterrows()):\n        total = row.sum()\n        ax.text(total + max(crosstab_sorted.sum(axis=1)) * 0.01, i, \n                f'{total:,}', va='center', ha='left', fontweight='bold', fontsize=9)\n    \n    plt.tight_layout()\n    plt.show()\n    \n    # Summary statistics table\n    print(f\"\\n📋 COMPOSITION SUMMARY:\")\n    \n    # Create summary with key metrics\n    summary_data = []\n    for vessel_type in crosstab_sorted.index:\n        row_data = crosstab_sorted.loc[vessel_type]\n        total_fragments = row_data.sum()\n        most_common_part = row_data.idxmax()\n        most_common_pct = (row_data.max() / total_fragments) * 100\n        part_diversity = (row_data > 0).sum()\n        \n        summary_data.append({\n            'Vessel Type': vessel_type,\n            'Total Fragments': total_fragments,\n            'Most Common Part': most_common_part,\n            'Dominance (%)': f\"{most_common_pct:.1f}%\",\n            'Part Diversity': part_diversity\n        })\n    \n    summary_df = pd.DataFrame(summary_data)\n    \n    # Style the summary table\n    styled_summary = summary_df.style.set_properties(**{\n        'background-color': '#f8f9fa',\n        'color': '#495057',\n        'border': '1px solid #dee2e6',\n        'text-align': 'center',\n        'font-size': '11px'\n    }).set_table_styles([\n        {'selector': 'th', 'props': [\n            ('background-color', '#dc3545'),  # Red theme for vessel-part analysis\n            ('color', 'white'),\n            ('font-weight', 'bold'),\n            ('text-align', 'center'),\n            ('border', '1px solid #c82333'),\n            ('padding', '8px')\n        ]},\n        {'selector': 'td', 'props': [\n            ('border', '1px solid #dee2e6'),\n            ('padding', '6px'),\n            ('font-weight', '500')\n        ]},\n        {'selector': 'table', 'props': [\n            ('border-collapse', 'collapse'),\n            ('margin', '10px auto'),\n            ('width', '90%'),\n            ('box-shadow', '0 2px 4px rgba(0,0,0,0.1)')\n        ]}\n    ])\n    \n    display(styled_summary)\n    \n    # Key insights\n    print(f\"\\n🔍 KEY INSIGHTS:\")\n    if len(summary_df) > 0:\n        most_fragments = summary_df.loc[summary_df['Total Fragments'].idxmax()]\n        most_diverse = summary_df.loc[summary_df['Part Diversity'].idxmax()]\n        \n        print(f\"   • Most fragmented vessel: '{most_fragments['Vessel Type']}' ({most_fragments['Total Fragments']:,} fragments)\")\n        print(f\"   • Most diverse parts: '{most_diverse['Vessel Type']}' ({most_diverse['Part Diversity']} different parts)\")\n        print(f\"   • Total vessel-part combinations: {len(summary_df)} vessel types analyzed\")\n        \n        # Additional insights\n        avg_diversity = summary_df['Part Diversity'].mean()\n        total_fragments = summary_df['Total Fragments'].sum()\n        print(f\"   • Average part diversity per vessel: {avg_diversity:.1f} parts\")\n        print(f\"   • Total fragments in analysis: {total_fragments:,}\")\n\n# Execute analysis\nif not df_info.empty:\n    analyze_vessel_part_composition(df_info, min_frequency=1)\nelse:\n    print(\"❌ No data available for analysis. Please check your data loading process.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-30T10:27:09.322152Z","iopub.execute_input":"2025-07-30T10:27:09.322589Z","iopub.status.idle":"2025-07-30T10:27:10.590043Z","shell.execute_reply.started":"2025-07-30T10:27:09.322553Z","shell.execute_reply":"2025-07-30T10:27:10.588844Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### **Vessel Type × Unit**","metadata":{}},{"cell_type":"code","source":"# Excavation Unit × Vessel Type Distribution Analysis\n\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport numpy as np\nfrom IPython.display import display, HTML\nimport warnings\nwarnings.filterwarnings('ignore')\n\ndef analyze_unit_vessel_distribution(df, top_n_units=15, min_frequency=1):\n    \"\"\"\n    Analyze excavation unit × vessel type distribution relationships\n    \"\"\"\n    if df.empty:\n        print(\"❌ DataFrame is empty.\")\n        return\n    \n    print(\"EXCAVATION UNIT × VESSEL TYPE DISTRIBUTION\")\n    \n    # Filter to most active units and frequent vessel types\n    unit_counts = df['unit'].value_counts()\n    type_counts = df['type'].value_counts()\n    \n    top_units = unit_counts.head(top_n_units).index\n    frequent_types = type_counts[type_counts >= min_frequency].index\n    \n    df_filtered = df[df['unit'].isin(top_units) & df['type'].isin(frequent_types)]\n    \n    print(f\"📊 Analysis Info:\")\n    print(f\"   • Top excavation units: {len(top_units)} out of {df['unit'].nunique()}\")\n    print(f\"   • Vessel types included: {len(frequent_types)} (≥{min_frequency} occurrences)\")\n    print(f\"   • Records analyzed: {len(df_filtered):,} out of {len(df):,} ({len(df_filtered)/len(df)*100:.1f}%)\")\n    \n    # Create cross-tabulation\n    crosstab = pd.crosstab(df_filtered['unit'], df_filtered['type'])\n    \n    # Sort units by total activity and types by frequency\n    unit_totals = crosstab.sum(axis=1).sort_values(ascending=False)\n    type_totals = crosstab.sum(axis=0).sort_values(ascending=False)\n    \n    crosstab_sorted = crosstab.loc[unit_totals.index, type_totals.index]\n    \n    # Create the visualization\n    fig, ax = plt.subplots(1, 1, figsize=(16, 10))\n    \n    # Plot: Absolute counts heatmap\n    sns.heatmap(crosstab_sorted, annot=True, fmt='d', cmap='Blues', \n                linewidths=0.5, linecolor='white', ax=ax,\n                cbar_kws={'label': 'Number of Fragments', 'shrink': 0.8})\n    \n    ax.set_title('Unit × Vessel Type Distribution', fontsize=16, fontweight='bold', pad=20)\n    ax.set_xlabel('Vessel Type', fontsize=12, fontweight='bold')\n    ax.set_ylabel('Excavation Unit', fontsize=12, fontweight='bold')\n    ax.tick_params(axis='x', rotation=45, labelsize=9)\n    ax.tick_params(axis='y', rotation=0, labelsize=9)\n    \n    plt.tight_layout()\n    plt.show()\n    \n    # Unit specialization analysis\n    print(f\"\\n📈 UNIT SPECIALIZATION ANALYSIS:\")\n    \n    # Calculate percentage by unit (row percentages)\n    crosstab_pct = crosstab_sorted.div(crosstab_sorted.sum(axis=1), axis=0) * 100\n    \n    # Find units with highest concentration in specific vessel types\n    specialization_data = []\n    for unit in crosstab_sorted.index:\n        unit_row = crosstab_pct.loc[unit]\n        if unit_row.sum() > 0:  # Avoid empty units\n            max_type = unit_row.idxmax()\n            max_pct = unit_row.max()\n            total_fragments = crosstab_sorted.loc[unit].sum()\n            diversity = (crosstab_sorted.loc[unit] > 0).sum()\n            \n            specialization_data.append({\n                'Unit': unit,\n                'Specialization': max_type,\n                'Concentration (%)': f\"{max_pct:.1f}%\",\n                'Total Fragments': total_fragments,\n                'Type Diversity': diversity\n            })\n    \n    spec_df = pd.DataFrame(specialization_data)\n    spec_df = spec_df.sort_values('Total Fragments', ascending=False)\n    \n    # Style the specialization table\n    styled_spec = spec_df.style.set_properties(**{\n        'background-color': '#f8f9fa',\n        'color': '#495057',\n        'border': '1px solid #dee2e6',\n        'text-align': 'center',\n        'font-size': '11px'\n    }).set_table_styles([\n        {'selector': 'th', 'props': [\n            ('background-color', '#007bff'),  # Blue theme for unit-vessel analysis\n            ('color', 'white'),\n            ('font-weight', 'bold'),\n            ('text-align', 'center'),\n            ('border', '1px solid #0056b3'),\n            ('padding', '8px')\n        ]},\n        {'selector': 'td', 'props': [\n            ('border', '1px solid #dee2e6'),\n            ('padding', '6px'),\n            ('font-weight', '500')\n        ]},\n        {'selector': 'table', 'props': [\n            ('border-collapse', 'collapse'),\n            ('margin', '10px auto'),\n            ('width', '90%'),\n            ('box-shadow', '0 2px 4px rgba(0,0,0,0.1)')\n        ]}\n    ])\n    \n    display(styled_spec)\n    \n    # Additional analysis: Create a summary of unit patterns\n    print(f\"\\n📊 EXCAVATION PATTERNS:\")\n    \n    # Overall statistics\n    total_units_analyzed = len(spec_df)\n    avg_fragments_per_unit = spec_df['Total Fragments'].mean()\n    avg_diversity_per_unit = spec_df['Type Diversity'].mean()\n    \n    # Concentration analysis\n    high_concentration_units = spec_df[\n        spec_df['Concentration (%)'].str.rstrip('%').astype(float) >= 50\n    ]\n    \n    print(f\"   • Units with high specialization (≥50%): {len(high_concentration_units)}/{total_units_analyzed}\")\n    print(f\"   • Average fragments per unit: {avg_fragments_per_unit:.1f}\")\n    print(f\"   • Average vessel type diversity: {avg_diversity_per_unit:.1f} types per unit\")\n    \n    # Key insights\n    print(f\"\\n🔍 KEY INSIGHTS:\")\n    if len(spec_df) > 0:\n        most_active = spec_df.iloc[0]\n        most_specialized = spec_df.loc[spec_df['Concentration (%)'].str.rstrip('%').astype(float).idxmax()]\n        most_diverse = spec_df.loc[spec_df['Type Diversity'].idxmax()]\n        \n        print(f\"   • Most active unit: '{most_active['Unit']}' ({most_active['Total Fragments']} fragments)\")\n        print(f\"   • Most specialized: '{most_specialized['Unit']}' ({most_specialized['Concentration (%)']} {most_specialized['Specialization']})\")\n        print(f\"   • Most diverse unit: '{most_diverse['Unit']}' ({most_diverse['Type Diversity']} different types)\")\n        \n        # Pattern insights\n        if len(high_concentration_units) > 0:\n            print(f\"   • Specialization pattern: {len(high_concentration_units)} units show strong preference for specific vessel types\")\n        else:\n            print(f\"   • Distribution pattern: Units show relatively balanced vessel type distributions\")\n\n# Execute analysis\nif not df_info.empty:\n    analyze_unit_vessel_distribution(df_info, top_n_units=15, min_frequency=1)\nelse:\n    print(\"❌ No data available for analysis. Please check your data loading process.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-30T10:20:12.499408Z","iopub.execute_input":"2025-07-30T10:20:12.500396Z","iopub.status.idle":"2025-07-30T10:20:13.945083Z","shell.execute_reply.started":"2025-07-30T10:20:12.500344Z","shell.execute_reply":"2025-07-30T10:20:13.944031Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# **Part Two: Basic image processing**\n\n## **2.1 Automated detection of targets**\n\nA critical prerequisite for any detailed analysis in this competition is the ability to **accurately isolate the pottery sherd** from its surrounding environment in each image. This process, known as **segmentation**, allows us to focus exclusively on the artifact, enabling precise measurements of its **physical dimensions, color, and texture**, which are essential for both classification and reconstruction.\n\nMy initial exploration involved several **fully automated computer vision techniques**:\n\n- **Canny edge detection**\n- **Global thresholding** (Otsu's method)\n\nHowever, as my trials demonstrated, these approaches proved **insufficient** for this specific dataset. The complex surface texture of the sherds and the presence of high-contrast calibration tools (the scale bar and color checker) often resulted in **fragmented or inaccurate object outlines**, making reliable identification impossible.\n\nEven more advanced methods like **adaptive thresholding**, while an improvement, still **struggled to consistently differentiate the sherd** from the other elements in the image.\n\n\n### **A Hybrid Strategy**\n\nThis iterative process revealed that a more robust, **hybrid strategy** was necessary. The final, successful approach combines a simple manual step with a powerful automated algorithm.\n\n#### **1. Manual Masking of Fixed Interferences**\n\nI first acknowledge that the **calibration tools are consistently placed** in the lower portion of each image. A fixed rectangular area encompassing these tools is defined and programmatically **\"masked out\" or ignored**. This simple step removes the **primary sources of algorithmic confusion**. And the marking of these calibration tools are **manually** done after the shred marking.\n \n#### **2. Automated Segmentation of the Target**\n\nWith the interferences removed, a combination of **adaptive thresholding** and **morphological operations** is applied to the image. This cleanly separates the remaining **foreground object (the sherd)** from the uniform background. The **largest resulting contour** is then confidently identified and marked as the **pottery sherd**.\n\n","metadata":{}},{"cell_type":"code","source":"#Automated detection of targets\n\nimport cv2\nimport numpy as np\nimport matplotlib.pyplot as plt\n# --- Config ---\n# List of image paths to process\nimage_paths = [\n   '/kaggle/input/h690/h690/sherd_images/JD00001_exterior.jpg',\n   '/kaggle/input/h690/h690/sherd_images/JD00002_interior.jpg'\n]\n# Define the rectangular region covering the scale bar and color checker\n# Format: (top-left x, top-left y, width, height)\ncalibration_tools_mask_area = (260, 730, 480, 200)\n# --- Loop over each image ---\nfor image_path in image_paths:\n   print(f\"\\n📂 Processing image: {image_path}\")\n   \n   try:\n       image = cv2.imread(image_path)\n       image_rgb = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\n       print(\"Image loaded successfully.\")\n   except Exception as e:\n       print(f\"Failed to load image: {e}\")\n       continue\n   # Convert to grayscale\n   gray = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY)\n   # Apply Gaussian blur\n   blurred = cv2.GaussianBlur(gray, (5, 5), 0)\n   # Adaptive thresholding to segment the image\n   adaptive_thresh = cv2.adaptiveThreshold(\n       blurred, 255, cv2.ADAPTIVE_THRESH_GAUSSIAN_C, \n       cv2.THRESH_BINARY_INV, 51, 9\n   )\n   # Morphological closing to fill small holes\n   kernel = np.ones((5, 5), np.uint8)\n   closing = cv2.morphologyEx(adaptive_thresh, cv2.MORPH_CLOSE, kernel, iterations=2)\n   # Mask out the region that contains calibration tools (scale bar + color checker)\n   (x, y, w, h) = calibration_tools_mask_area\n   cv2.rectangle(closing, (x, y), (x + w, y + h), (0, 0, 0), -1)  # -1 means filled rectangle\n   # Find contours in the cleaned-up mask\n   contours, hierarchy = cv2.findContours(closing.copy(), cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE)\n   print(f\"🔎 Found {len(contours)} contours after masking calibration tools.\")\n   output_image = image_rgb.copy()\n   \n   if contours:\n       # Choose the largest contour, assuming it is the sherd\n       sherd_contour = max(contours, key=cv2.contourArea)\n       \n       # Get bounding box of the sherd\n       (sx, sy, sw, sh) = cv2.boundingRect(sherd_contour)\n       \n       # Draw rectangle and label on the original image\n       cv2.rectangle(output_image, (sx, sy), (sx + sw, sy + sh), (255, 0, 255), 4)  # Purple box\n       cv2.putText(output_image, \"Automated Sherd Mark\", (sx, sy - 10), \n                   cv2.FONT_HERSHEY_SIMPLEX, 0.8, (255, 0, 255), 2)\n       print(\"Successfully marked sherd region.\")\n       print(f\"📍 Sherd coordinates: (x={sx}, y={sy}, width={sw}, height={sh})\")\n   else:\n       print(\"No sherd contour detected.\")\n   \n   # Add scale bar and color checker marks\n   # Scale Bar\n   scale_coords = (275, 780, 465, 88)\n   cv2.rectangle(output_image, (scale_coords[0], scale_coords[1]), \n                 (scale_coords[0] + scale_coords[2], scale_coords[1] + scale_coords[3]), (0, 255, 0), 4)\n   cv2.putText(output_image, \"Scale Bar\", (scale_coords[0], scale_coords[1] - 5), \n               cv2.FONT_HERSHEY_SIMPLEX, 0.8, (0, 255, 0), 2)\n   print(f\"📍 Scale Bar coordinates: (x={scale_coords[0]}, y={scale_coords[1]}, width={scale_coords[2]}, height={scale_coords[3]})\")\n   \n   # Color Checker\n   color_coords = (275, 865, 465, 40)\n   cv2.rectangle(output_image, (color_coords[0], color_coords[1]), \n                 (color_coords[0] + color_coords[2], color_coords[1] + color_coords[3]), (255, 0, 0), 4)\n   cv2.putText(output_image, \"Color Checker\", (color_coords[0], color_coords[1] - 5), \n               cv2.FONT_HERSHEY_SIMPLEX, 0.8, (255, 0, 0), 2)\n   print(f\"📍 Color Checker coordinates: (x={color_coords[0]}, y={color_coords[1]}, width={color_coords[2]}, height={color_coords[3]})\")\n   \n   # --- Visualization ---\n   plt.figure(figsize=(18, 9))\n   plt.subplot(1, 3, 1)\n   plt.imshow(image_rgb)\n   plt.title('1. Original Image')\n   plt.axis('off')\n   plt.subplot(1, 3, 2)\n   plt.imshow(closing, cmap='gray')\n   plt.title(\"2. Segmentation with Calibration Mask\")\n   plt.axis('off')\n   plt.subplot(1, 3, 3)\n   plt.imshow(output_image)\n   plt.title('3. Detected Sherd')\n   plt.axis('off')\n   plt.tight_layout()\n   plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-30T10:00:53.389460Z","iopub.execute_input":"2025-07-30T10:00:53.389810Z","iopub.status.idle":"2025-07-30T10:00:55.850099Z","shell.execute_reply.started":"2025-07-30T10:00:53.389780Z","shell.execute_reply":"2025-07-30T10:00:55.848379Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **2.2 Real-world physical measurement**\nThe original image pixel data itself is not scientifically comparable. By using the corresponding scale, we can perform physical calibration and convert the pixel information into standardized and meaningful measurements.\n### **Methodology:**\n\n#### 1. Identify the Scale Bar Location\n- Use the predefined scale bar region in the image to extract its pixel length  and associate it with its known real-world physical length (e.g., `6 cm`).\n\n#### 2. Calculate Pixel Density\n\n#### 3. Extract Sherd Contour and Measure Real Dimensions\n\n- Use contour detection (`cv2.findContours`) and morphological operations to accurately extract the sherd boundary.\n\n- Obtain sherd width and height using the minimum bounding rectangle (`cv2.boundingRect`).\n\n- Use contour area (`cv2.contourArea`) to calculate the actual surface area (in square centimeters).","metadata":{}},{"cell_type":"code","source":"# Real-world physical measurement\n\nimport cv2\nimport numpy as np\nimport matplotlib.pyplot as plt\n\n\nimage_path = '/kaggle/input/h690/h690/sherd_images/JD00002_interior.jpg'\n\n# Define the fixed mask area covering the calibration tools\ncalibration_tools_mask_area = (260, 730, 480, 200)\n\nprint(f\"\\n📂 Processing image: {image_path}\")\n\ntry:\n   image = cv2.imread(image_path)\n   image_rgb = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\n   print(\"Image loaded successfully.\")\nexcept Exception as e:\n   print(f\"Failed to load image: {e}\")\n   exit()\n\n# --- Image segmentation and sherd detection ---\ngray = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY)\nblurred = cv2.GaussianBlur(gray, (5, 5), 0)\nadaptive_thresh = cv2.adaptiveThreshold(\n   blurred, 255, cv2.ADAPTIVE_THRESH_GAUSSIAN_C, \n   cv2.THRESH_BINARY_INV, 51, 9\n)\nkernel = np.ones((5, 5), np.uint8)\nclosing = cv2.morphologyEx(adaptive_thresh, cv2.MORPH_CLOSE, kernel, iterations=2)\n\n# Apply masking to exclude calibration tools\n(x, y, w, h) = calibration_tools_mask_area\ncv2.rectangle(closing, (x, y), (x + w, y + h), (0, 0, 0), -1)\n\ncontours, hierarchy = cv2.findContours(closing.copy(), cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE)\nprint(f\"🔎 Found {len(contours)} contours after masking calibration tools.\")\n\noutput_image = image_rgb.copy()\n\n# --- Real-world physical measurements ---\npixels_per_metric = None\n\n# 1. Compute pixel-to-cm ratio using known scale bar\nscale_coords = (275, 780, 465, 88)\nscale_bar_pixel_width = scale_coords[2]  \nknown_width_cm = 6.0  # Known physical length of scale bar in cm\npixels_per_metric = scale_bar_pixel_width / known_width_cm\nprint(f\"📏 Computed scale: {pixels_per_metric:.2f} pixels/cm\")\n\nif contours:\n   # Identify the largest contour, assuming it's the sherd\n   sherd_contour = max(contours, key=cv2.contourArea)\n   \n   # 2. Calculate physical dimensions using bounding box\n   (sx, sy, sw, sh) = cv2.boundingRect(sherd_contour)\n   sherd_width_cm = sw / pixels_per_metric\n   sherd_height_cm = sh / pixels_per_metric\n   \n   # 3. Calculate actual sherd area using contour (not bounding box)\n   sherd_area_pixels = cv2.contourArea(sherd_contour)  \n   sherd_area_cm2 = sherd_area_pixels / (pixels_per_metric ** 2)\n   \n   # 4. Calculate bounding box area for comparison\n   bbox_area_pixels = sw * sh\n   bbox_area_cm2 = bbox_area_pixels / (pixels_per_metric ** 2)\n   \n   # 5. Calculate fill ratio (actual area vs bounding box area)\n   fill_ratio = (sherd_area_pixels / bbox_area_pixels) * 100\n   \n   print(\"✅ Sherd contour successfully identified.\")\n   print(f\"📍 Sherd pixel coordinates: (x={sx}, y={sy}, width={sw}, height={sh})\")\n   print(\"--- Real-World Physical Measurements ---\")\n   print(f\"Bounding Box Width: {sherd_width_cm:.2f} cm\")\n   print(f\"Bounding Box Height: {sherd_height_cm:.2f} cm\")\n   print(f\"Bounding Box Area: {bbox_area_cm2:.2f} cm²\")\n   print(f\"🎯 Actual Sherd Area: {sherd_area_cm2:.2f} cm²\")\n   print(f\"📊 Fill Ratio: {fill_ratio:.1f}% (sherd area / bounding box area)\")\n   \n   # Draw bounding box\n   cv2.rectangle(output_image, (sx, sy), (sx + sw, sy + sh), (255, 0, 255), 2)\n   \n   # Draw actual sherd contour\n   cv2.drawContours(output_image, [sherd_contour], -1, (0, 255, 0), 3)\n   \n   # Add measurement text \n   dimension_text = f\"W: {sherd_width_cm:.1f}cm  H: {sherd_height_cm:.1f}cm\"\n   cv2.putText(output_image, dimension_text, (sx, sy - 10), \n               cv2.FONT_HERSHEY_SIMPLEX, 0.7, (255, 0, 255), 2)\n   \n   # Add legend\n   cv2.putText(output_image, \"Green: Actual Sherd Contour\", (10, 30), \n               cv2.FONT_HERSHEY_SIMPLEX, 0.7, (0, 255, 0), 2)\n   cv2.putText(output_image, \"Purple: Bounding Box\", (10, 60), \n               cv2.FONT_HERSHEY_SIMPLEX, 0.7, (255, 0, 255), 2)\n   \nelse:\n   print(\"⚠️ No sherd contour detected.\")\n\n# Add scale bar marking for reference\nscale_coords = (275, 780, 465, 88)\ncv2.rectangle(output_image, (scale_coords[0], scale_coords[1]), \n             (scale_coords[0] + scale_coords[2], scale_coords[1] + scale_coords[3]), (0, 255, 0), 2)\ncv2.putText(output_image, \"Scale Bar (6cm)\", (scale_coords[0], scale_coords[1] - 5), \n           cv2.FONT_HERSHEY_SIMPLEX, 0.6, (0, 255, 0), 2)\n\n# --- Visualization ---\nplt.figure(figsize=(15, 10))\nplt.imshow(output_image)\nplt.title(f'Physical Measurement Results for JD00002_interior.jpg')\nplt.axis('off')\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-30T10:00:55.851413Z","iopub.execute_input":"2025-07-30T10:00:55.851697Z","iopub.status.idle":"2025-07-30T10:00:56.944955Z","shell.execute_reply.started":"2025-07-30T10:00:55.851674Z","shell.execute_reply":"2025-07-30T10:00:56.943582Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **2.3 Standardized color analysis**\n**olor correction** is crucial for eliminating the influence of environmental lighting and camera characteristics on the RGB values of an image. When the same artifact is photographed under different lighting conditions, its RGB values may vary significantly. By using the **Color Checker** in the image, color information is automatically extracted and the color of the entire image is corrected, making the image closer to the real physical color\n### **Methodalogy:**\n#### 1. Automatic Patch Extraction\n\n- Automatically divide the color checker region into 12 vertical color patches.  \n- Compute the **average RGB value** for each patch.  \n\n#### 2. Color Correction Matrix (CCM) Calculation\n\n- Use the last few grayscale patches (assumed to be white, gray, and black) to compute white balance factors.  \n- Construct a simplified **diagonal color correction matrix (CCM)**.  \n\n\n#### 3. Image Color Correction\n\n- Apply the CCM to the entire image via matrix multiplication.  \n- Ensure that the output RGB values stay within the valid range `[0, 255]`.\n","metadata":{}},{"cell_type":"code","source":"# Standardized color analysis\n\nimport cv2\nimport numpy as np\nimport matplotlib.pyplot as plt\nfrom sklearn.linear_model import LinearRegression\n\ndef extract_color_patches_auto(color_checker_image):\n    \"\"\"Automatically detect and extract color patches from a color checker\"\"\"\n    h, w = color_checker_image.shape[:2]\n    possible_cols = [12, 13, 14, 15, 16]  # Try different patch counts\n\n    colors = []\n    best_cols = 12  # Default\n\n    for cols in possible_cols:\n        patch_w = w // cols\n        if patch_w > 10:\n            best_cols = cols\n            break\n\n    print(f\"🎨 Detected {best_cols} color patches\")\n    patch_w = w // best_cols\n\n    for j in range(best_cols):\n        margin = 2\n        x1 = max(0, j * patch_w + margin)\n        x2 = min(w, (j + 1) * patch_w - margin)\n        y1 = margin\n        y2 = h - margin\n\n        if y2 > y1 and x2 > x1:\n            patch = color_checker_image[y1:y2, x1:x2]\n            avg_color = np.mean(patch, axis=(0, 1))\n            colors.append(avg_color)\n            print(f\"Patch {j+1}: RGB({avg_color[0]:.0f}, {avg_color[1]:.0f}, {avg_color[2]:.0f})\")\n\n    return np.array(colors, dtype=\"float32\")\n\ndef compute_color_correction_matrix_simple(observed_colors):\n    \"\"\"\n    Simplified color correction using white balance assumption\n    \"\"\"\n    num_colors = len(observed_colors)\n    \n    if num_colors >= 3:\n        white_patch = observed_colors[-4] if num_colors >= 4 else observed_colors[-1]\n        gray_patch = observed_colors[-2] if num_colors >= 2 else observed_colors[0]\n\n        target_white = np.array([240, 240, 240], dtype=np.float32)\n        target_gray = np.array([128, 128, 128], dtype=np.float32)\n\n        white_correction = target_white / np.maximum(white_patch, 1.0)\n        ccm = np.diag(white_correction)\n\n        print(f\"🎯 Detected white patch: RGB({white_patch[0]:.0f}, {white_patch[1]:.0f}, {white_patch[2]:.0f})\")\n        print(f\"📊 White balance correction: R={white_correction[0]:.3f}, G={white_correction[1]:.3f}, B={white_correction[2]:.3f}\")\n    else:\n        ccm = np.eye(3, dtype=np.float32)\n        print(\"⚠️ Insufficient color patches for correction. Using identity matrix.\")\n\n    return ccm\n\ndef apply_color_correction_matrix(image, ccm):\n    \"\"\"\n    Apply the color correction matrix to the entire image\n    \"\"\"\n    original_shape = image.shape\n    image_flat = image.reshape(-1, 3).astype(np.float32)\n    corrected_flat = np.dot(image_flat, ccm.T)\n    corrected_flat = np.clip(corrected_flat, 0, 255)\n    corrected_image = corrected_flat.reshape(original_shape).astype(np.uint8)\n    return corrected_image\n\n# --- Setup and Load Image ---\nimage_path = '/kaggle/input/h690/h690/sherd_images/JD00002_interior.jpg'\n\n# Define manual ROIs\nROIs = {\n    \"JD00002_interior\": {\n        \"Pottery Sherd\": {\"coords\": (222, 253, 527, 386)},\n        \"Color Checker\": {\"coords\": (275, 865, 465, 40)},\n    }\n}\n\nprint(f\"\\n📂 Processing image: {image_path}\")\nimage_name = image_path.split('/')[-1].replace('.jpg', '')\n\ntry:\n    image = cv2.imread(image_path)\n    image_rgb = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\n    print(f\"✅ Image loaded successfully: {image_path}\")\nexcept Exception as e:\n    print(f\"❌ Error loading image: {e}\")\n\n# Extract ROIs\ncurrent_roi = ROIs[\"JD00002_interior\"]\n(x, y, w, h) = current_roi[\"Color Checker\"][\"coords\"]\nobserved_checker = image_rgb[y:y+h, x:x+w]\nprint(f\"Color checker area shape: {observed_checker.shape}\")\n\nif observed_checker.size == 0:\n    print(\"❌ Color checker region is empty. Skipping correction.\")\n    corrected_image_uint8 = image_rgb.copy()\nelse:\n    observed_colors = extract_color_patches_auto(observed_checker)\n    print(f\"✅ Successfully extracted {len(observed_colors)} patches\")\n\n    print(\"🎨 Computing color correction matrix...\")\n    try:\n        ccm = compute_color_correction_matrix_simple(observed_colors)\n        print(f\"CCM matrix:\\n{ccm}\")\n\n        corrected_image_uint8 = apply_color_correction_matrix(image_rgb, ccm)\n    except Exception as e:\n        print(f\"❌ Color correction failed: {e}\")\n        corrected_image_uint8 = image_rgb.copy()\n\n# Visualize Results\n(sx, sy, sw, sh) = current_roi[\"Pottery Sherd\"][\"coords\"]\noriginal_sherd = image_rgb[sy:sy+sh, sx:sx+sw]\ncorrected_sherd = corrected_image_uint8[sy:sy+sh, sx:sx+sw]\nobserved_display = cv2.resize(observed_checker, (600, 100))\n\nplt.figure(figsize=(20, 12))\n\nplt.subplot(2, 3, 1)\nplt.imshow(image_rgb)\nplt.title(f\"1. Original Full Image - {image_name}\", fontsize=14)\nplt.axis('off')\n\nplt.subplot(2, 3, 2)\nplt.imshow(corrected_image_uint8)\nplt.title(f\"2. Color Corrected Full Image - {image_name}\", fontsize=14)\nplt.axis('off')\n\nplt.subplot(2, 3, 3)\nplt.imshow(observed_checker)\nplt.title(\"3. Observed Color Checker\", fontsize=14)\nplt.axis('off')\n\nplt.subplot(2, 3, 4)\nplt.imshow(original_sherd)\nplt.title(\"4. Sherd - Before Correction\", fontsize=14)\nplt.axis('off')\n\nplt.subplot(2, 3, 5)\nplt.imshow(corrected_sherd)\nplt.title(\"5. Sherd - After Correction\", fontsize=14)\nplt.axis('off')\n\nplt.subplot(2, 3, 6)\nplt.imshow(observed_display)\nplt.title(\"6. Color Checker (Zoomed)\", fontsize=14)\nplt.axis('off')\n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-30T10:00:56.946138Z","iopub.execute_input":"2025-07-30T10:00:56.946455Z","iopub.status.idle":"2025-07-30T10:00:59.564257Z","shell.execute_reply.started":"2025-07-30T10:00:56.946430Z","shell.execute_reply":"2025-07-30T10:00:59.561817Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<p align=\"center\">\n  <img src=\"https://media.giphy.com/media/APHFMUIaTnLIA/giphy.gif\" width=\"300\" />\n</p>\n\n---\n\n### 🙌 If you found this notebook helpful...\n**Please consider leaving an _upvote_ 👍**\n\n---\n\n","metadata":{}}]}