{"cells":[{"metadata":{},"cell_type":"markdown","source":"In an attempt to understand this competition better, I was reading the paper describing Alaska-I competition winner solution from [this](https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/147039) discussion post. I noticed that out of many things that winner tried, one of them was splitting the dataset into three sets, namely training set (TRN), validation set (VAL), and test set (TST).","execution_count":null},{"metadata":{},"cell_type":"markdown","source":">The training set (TRN), validation set (VAL), and test set (TST)\ncontained respectively 42,500, 3,500, and 3,500 cover images (around\n500 cover images were not used because they were corrupted or\nfailed the processing pipeline). The TRN, VAL, and TST sets were\ncreated for each quality factor and each stego scheme in TILEdouble,\nTILEbase, and ARBITRARYbase","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"They split the dataset into TRN, VAL and TST for **each quality factor and each stego scheme**. In our case this would mean splitting the dataset separately based on each quality factor of `75, 90, 95` and also based on each stego scheme of `JMiPOD, UERD, JUNIWARD`. The pseudo-code for this split would look something like this:","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"```python\n# split based on quality factors\nfor quality_factor in [75, 90, 95]:\n    split_data_into(TRN, VAL, TST)\n    \n# split based on stego scheme\nfor stego in ['JMiPOD', 'UERD', 'JUNIWARD']:\n    split_data_into(TRN, VAL, TST)\n\n```\n \n    ","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"This kernel is my attempt to split the dataset based on each quality factor and stego scheme. I will call TST, VAL and TST as `train`, `valid` and `test_val` respectively and will use `sklean`'s `train_test_split` to split the dataset with a split percentage of `70`, `20` and `10` respectively. The `train-valid` split will be used for training different classifiers while the `test_val` split can be used for ensembling of those classifiers. ","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"I will use the dataset from [this amazing kernel](https://www.kaggle.com/meaninglesslives/alaska2-cnn-multiclass-classifier) by @meaninglesslives as it contains information about quality factors.","execution_count":null},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import os\nimport pandas as pd\nimport numpy as np\nfrom sklearn.model_selection import train_test_split","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"!ls ../input/alaska2-image-steganalysis","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"!ls ../input/alaska2trainvalsplit","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"split_dir = '../input/alaska2trainvalsplit'","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_split = pd.read_csv(f'{split_dir}/alaska2_train_df.csv')\ntrain_split.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def transform(x):\n    split = x.split('/')\n    path = split[-2] + '/' + split[-1]\n    return path","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"I will modify the `ImageFileName` column as it contains the path of image files which is kaggle-style.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"train_split['ImageFileName'] = train_split['ImageFileName'].transform(transform)\ntrain_split.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_split['Label'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"valid_split = pd.read_csv(f'{split_dir}/alaska2_val_df.csv')\nvalid_split.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"valid_split['ImageFileName'] = valid_split['ImageFileName'].transform(transform)\nvalid_split.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# combine train and valid split\ndf_all = pd.concat([train_split, valid_split])\ndf_all.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# sanity check that train_split + valid_split = combined \ntrain_split.shape[0] + valid_split.shape[0] == df_all.shape[0]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def transform(x):\n    split = x.split('/')[-2]\n    path = split[-2] + '/' + split[-1]\n    return path","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# make a column representing the stego-scheme\ndf_all['Stego'] = df_all['ImageFileName'].transform(lambda x: x.split('/')[-2])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df_all.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df_all['Stego'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def qf_transform(x):\n    if x in [1,4,7]:\n        return 75\n    elif x in [2,5,8]:\n        return 90\n    elif x in [3,6,9]:\n        return 95\n    else:\n        return x","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# make a column representing the quality factors; for Cover images, I've set quality factor = 0\ndf_all['quality_factor'] = df_all['Label'].transform(qf_transform)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df_all.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df_all['quality_factor'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# save the combined df \ndf_all.to_csv('df_all.csv', index=False)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# split the data based on quality factor of 75, 90 and 95\nfor qf in [75, 90, 95]:\n    df_qf = df_all[df_all['quality_factor']==qf]\n    df_qf_tr, df_qf_val_test = train_test_split(df_qf, test_size=0.3, random_state=1234, stratify=df_qf['Label'].values)\n    df_qf_val, df_qf_test  = train_test_split(df_qf_val_test, test_size=0.2, random_state=1234, stratify=df_qf_val_test['Label'].values)\n    print(f'Split for quality factor of {qf}...')\n    #print(df_qf_tr['Label'].value_counts())\n    #print(df_qf_val['Label'].value_counts())\n    #print(df_qf_test['Label'].value_counts())\n    print('Shape of train split: ', df_qf_tr.shape)\n    print('Shape of valid split: ', df_qf_val.shape)\n    print('Shape of val_test split: ', df_qf_test.shape)\n    print('*'*35)\n    \n    #save the splits\n    df_qf_tr.to_csv(f'train_split_qf_{qf}.csv', index=False)\n    df_qf_val.to_csv(f'valid_split_qf_{qf}.csv', index=False)\n    df_qf_test.to_csv(f'test_val_split_qf_{qf}.csv', index=False)\n    ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# split the data based on stego scheme of JMiPOD, UERD, JUNIWARD\nfor stego in ['JMiPOD', 'UERD', 'JUNIWARD']:\n    df_stego = df_all[df_all['Stego']==stego]\n    df_stego_tr, df_stego_val_test = train_test_split(df_stego, test_size=0.3, random_state=1234, stratify=df_stego['Label'].values)\n    df_stego_val, df_stego_test  = train_test_split(df_stego_val_test, test_size=0.2, random_state=1234, stratify=df_stego_val_test['Label'].values)\n    print(f'Split for Stego type {stego}...')\n    #print(df_stego_tr['Label'].value_counts())\n    #print(df_stego_val['Label'].value_counts())\n    #print(df_stego_test['Label'].value_counts())\n    print('Shape of train split: ', df_stego_tr.shape)\n    print('Shape of valid split: ', df_stego_val.shape)\n    print('Shape of val_test split: ', df_stego_test.shape)\n    print('*'*35)\n    \n    #save the splits\n    df_stego_tr.to_csv(f'train_split_stego_{stego}.csv', index=False)\n    df_stego_val.to_csv(f'valid_split_stego_{stego}.csv', index=False)\n    df_stego_test.to_csv(f'test_val_split_stego_{stego}.csv', index=False)\n    ","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"These splits can be used to train different classifiers based on quality factor and stego-scheme using `train-valid` sets and finally combine them using `test_val` set. \n>This kernel just shows one of the many ways in which the dataset can be split. If you other ways, please let me know me in the comments","execution_count":null}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}