{"cells":[{"metadata":{},"cell_type":"markdown","source":"### Motivation\n\nNow, we [know it officially](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161943) that there are duplicate images in the data set. The hosts have officially provided a list of duplicates which I put in a [public data set](https://www.kaggle.com/graf10a/siim-list-of-duplicates). The purpose of this notebook is to check the consistency of the tabular data between the orignal images and their duplicates. ","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"### Loadng libraries and data","execution_count":null},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import os\nimport numpy as np\nimport pandas as pd\nfrom pathlib import Path\n\npd.set_option('max_columns', None)\npd.set_option('max_rows', None)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"show_files=0\n\nif show_files:\n    for dirname, _, filenames in os.walk('/kaggle/input'):\n        for filename in filenames:\n            print(os.path.join(dirname, filename))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"train=pd.read_csv('/kaggle/input/siim-isic-melanoma-classification/train.csv')\ntest=pd.read_csv('/kaggle/input/siim-isic-melanoma-classification/test.csv')\ndups=pd.read_csv('/kaggle/input/siim-list-of-duplicates/2020_Challenge_duplicates.csv')\n\n# TRAIN_IMG=Path('/kaggle/input/siim-isic-melanoma-classification/jpeg/train/')\n# TEST_IMG=Path('/kaggle/input/siim-isic-melanoma-classification/jpeg/test/')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### A quick peek at the list of duplicates","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"dups.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"len(dups)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"unq=np.unique(dups['ISIC_id'].values)\nlen(unq)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### The consistency check","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"Here is a function that we will use to make data frames holding tabular data for original and paired images present in a given partition (`train` or `test`). ","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"def extract_data(partition, verbose=1):\n    \n    print(f\"Partition: {partition}\")\n    \n    if partition=='train':\n        df=train.copy()\n    else:\n        df=test.copy()\n        \n    mask_df=dups['partition']==partition\n\n    \n    original=dups[mask_df].merge(df, left_on='ISIC_id', right_on='image_name')\n\n        \n    paired=dups[mask_df].merge(df, left_on='ISIC_id_paired', right_on='image_name')\n    \n    if verbose:\n        print(f\"The total number of entries: {mask_df.sum()}\")    \n        print(f\"The length of 'original': {len(original)}\")    \n        print(f\"The length of 'paired': {len(paired)}\")\n    \n    return original, paired, df","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"original_train, paired_train, _ = extract_data('train')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"original_train.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Now, let's make a function for consistency checking -- we want to make sure that the original and the paired tabular data are the same (or almost the same).","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"def check_consistency(partition, verbose=0):\n    \n    original, paired, df = extract_data(partition, verbose=verbose)\n    \n    cols=[c for c in df.columns if c not in ['image_name']]\n\n    for c in cols:\n        print(\"=\"*100)\n        print(f\"{c}:\")\n        mask_c=np.equal(original[c].fillna('na').values, paired[c].fillna('na').values)\n        if mask_c.all():\n            print(f\"The values of '{c}' are in a perfect agreement between the original and paired images.\")\n        else:\n            print(f\"The values of {c} differ between the original and paired images.\\n\")\n            df_cols=['ISIC_id', 'ISIC_id_paired', c]\n            df=original.loc[~mask_c, df_cols]\n            df[c+'_o']=df.pop(c)\n            df[c+'_p']=paired.loc[~mask_c, c]\n            print(df)\n    print(\"=\"*100)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"check_consistency('train')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"check_consistency('test')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Conclusion\n\nWe see that there are some tabular data that differ between the original and paired images. Fortunately, the number of such data points is small, so they should not affect our final results in any significant way. ","execution_count":null}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}