{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"Note book này được sao chép và chỉnh sửa lại từ project của [Nick Kuzmenkov ](https://www.kaggle.com/nickuzmenkov)\n\nNotebook này tập trung vào việc xử lý dữ liệu, bao gồm các công việc\n1. Loại bỏ các ảnh bị trùng lặp.\n2. Định dạng lại nhãn.\n3. Tạo các fold để áp dụng phương pháp Cross Validation.\n4. Tạo pre-augmentation cho dữ liệu ảnh.\n5. Đưa kết quả vào định dạng file Tfrecord.\n### Các notebook khác của nhóm:\n\n[Revealing Duplicates notebook](https://www.kaggle.com/nvlinhh/int3414-22-n11-revealing-duplicate)\n \n[Tranning notebook](https://www.kaggle.com/congnguyen8201/int3414-22-n11-training)\n \n[Submission notebook](https://www.kaggle.com/congnguyen8201/int3414-22-n11-submission)","metadata":{}},{"cell_type":"markdown","source":"### Imports","metadata":{}},{"cell_type":"code","source":"from sklearn.preprocessing import MultiLabelBinarizer\nfrom sklearn.model_selection import StratifiedKFold\nfrom tqdm.notebook import tqdm\nimport matplotlib.pyplot as plt\nimport tensorflow as tf\nimport albumentations\nimport pandas as pd\nimport numpy as np\nimport shutil\nimport os","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Configuration\n","metadata":{}},{"cell_type":"code","source":"plt.style.use('fivethirtyeight')\nprint(f'Using TensorFlow {tf.__version__}')\n\nclass CFG:\n    \n    '''\n    keep these\n    '''\n    root = '../input/plant-pathology-2021-fgvc8/train_images'\n    classes = [\n        'complex', \n        'frog_eye_leaf_spot', \n        'powdery_mildew', \n        'rust', \n        'scab',\n        'healthy']\n    strategy = tf.distribute.get_strategy()#phân phối việc huấn luyện của Tensorflow\n    batch_size = 16\n    \n    '''\n    tune these\n    '''\n    img_size = 600 # kích thước ảnh\n    folds = 5 # số lượng KFold n_splits\n    seed = 42 # random seed (chỉ dành cho KFold)\n    subfolds = 16 # số file tfrec trong mỗi fold\n    transform = True # có áp dụng pre-augmentation\n    epochs = 5 # (>=5) số lần áp dụng pre-augmetaion khi transform = True\n","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 1. Removing duplicates\nViệc xác định sự trùng lặp của các ảnh được thực hiên ở **[other notebook](https://www.kaggle.com/nickuzmenkov/pp2021-duplicates-revealing)** và nhận được 50 ảnh trừn lặp. Dưới đây là cách nhóm xử lỷ các ảnh bị trùng:\n1. Chỉ giữ lại một nếu các ảnh giống nhau có chung nhãn.\n2. Xóa tất cả các ảnh nếu các ảnh giống nhau khác nhãn.","metadata":{}},{"cell_type":"code","source":"df = pd.read_csv('../input/plant-pathology-2021-fgvc8/train.csv', index_col='image')\ninit_len = len(df)\n\n#đọc file\nwith open('../input/int3414-22-n11-revealing-duplicate/duplicates.csv', 'r') as file:\n    duplicates = [x.strip().split(',') for x in file.readlines()]\nprint(len(duplicates))\n\nfor row in duplicates:\n    unique_labels = df.loc[row].drop_duplicates().values\n    if len(unique_labels) == 1:\n        df = df.drop(row[1:], axis=0)\n    else:\n        df = df.drop(row, axis=0)\n        \nprint(f'Dropping {init_len - len(df)} duplicate samples.')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df['labels'].value_counts()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2. Label formatting\nThe initial format of space-separated string labels is inapplicable for model training. Here we change the format via `MultiLabelBinarizer` instance","metadata":{}},{"cell_type":"code","source":"original_labels = df['labels'].values.copy()\noriginal_labels","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Chuyển label từ dạng string sang list \ndf['labels'] = [x.split(' ') for x in df['labels']]\nprint(df['labels'])\n\n#tạo onehot vector\nlabels = MultiLabelBinarizer(classes=CFG.classes).fit_transform(df['labels'].values)\ndf = pd.DataFrame(columns=CFG.classes, data=labels, index=df.index)\n\ndf.to_csv('train.csv')\ndisplay(df.head())","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n## 3. Making stratified folds\nĐể tránh imbalanced class distribution ở các fold nên nhóm sử dụng StratifiedKfolds của thư viện SKlearn. \n","metadata":{}},{"cell_type":"code","source":"kfold = StratifiedKFold(n_splits=CFG.folds, shuffle=True, random_state=CFG.seed)\nfold = np.zeros((len(df),))\n\n#Kfold sau khi trộn tập data, fold được tạo ra để đánh dấu từng phần tử của df thuộc tập test trong fold nào.\nfor i, (train_index, val_index) in enumerate(kfold.split(df.index, original_labels)):#đếm các phần tử trong tập val thuộc fold nào.\n    fold[val_index] = i\n    \n#đêm số lần các labels suất hiện ở từng fold(giá trị 1)\nvalue_counts = lambda x: pd.Series.value_counts(x, normalize=True)\ndf_occurence = pd.DataFrame({\n    'origin': df.apply(value_counts).loc[1],\n    'fold_0': df[fold == 0].apply(value_counts).loc[1],\n    'fold_1': df[fold == 1].apply(value_counts).loc[1],\n    'fold_2': df[fold == 2].apply(value_counts).loc[1],\n    'fold_3': df[fold == 3].apply(value_counts).loc[1],\n    'fold_4': df[fold == 4].apply(value_counts).loc[1]})\nprint(df_occurence)\n\nbar = df_occurence.plot.barh(figsize=[15, 5], colormap='plasma')\n\nfolds = pd.DataFrame({\n    'image': df.index,\n    'fold': fold})\n\nfolds.to_csv('folds.csv', index=False)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Đối với mảng fold, giá trị j ở vị trí i thể hiện cho việc phần tử ở vị trí i thuộc fold j.\nSau khi thực hiện việc chia thành các fold, một nhãn có tỉ lệ gần bằng nhau ở các fold. \n## 4. Data pre-augmentation\nĐối với việc tạo pre-augmentation cho tập dữ liệu, nhóm chọn thư viện albumentations.","metadata":{}},{"cell_type":"code","source":"if CFG.transform:\n    transform = albumentations.Compose([\n        #ngẫu nhiên áp dụng các phương pháp dưới đây với tham số xác xuất là p.\n       albumentations.RandomResizedCrop(CFG.img_size, CFG.img_size, scale=(0.9, 1), p=1), #cắt ảnh và resize\n       albumentations.HorizontalFlip(p=0.5),#lật ngang\n       albumentations.VerticalFlip(p=0.5),#lật dọc\n       albumentations.ShiftScaleRotate(p=0.5),#ngẫu nhiên áp dụng affine transform\n       albumentations.HueSaturationValue(hue_shift_limit=10, sat_shift_limit=10, val_shift_limit=10, p=0.7),#ngẫu nhiên thay đổi các giá trị trong không gian màu HSB(hue, saturation, brightness)\n       albumentations.RandomBrightnessContrast(brightness_limit=(-0.2,0.2), contrast_limit=(-0.2, 0.2), p=0.7),#Thay đổi độ tương phản\n       albumentations.CLAHE(clip_limit=(1,4), p=0.5),\n       albumentations.OneOf([\n           albumentations.OpticalDistortion(distort_limit=1.0),\n           albumentations.GridDistortion(num_steps=5, distort_limit=1.),\n           albumentations.ElasticTransform(alpha=3),\n       ], p=0.2),\n        #thêm các noise vào ảnh\n       albumentations.OneOf([\n           albumentations.GaussNoise(var_limit=[10, 50]),\n           albumentations.GaussianBlur(),\n           albumentations.MotionBlur(),\n           albumentations.MedianBlur(),\n       ], p=0.2),\n      albumentations.Resize(CFG.img_size, CFG.img_size),\n      albumentations.OneOf([\n          albumentations.JpegCompression(),\n          albumentations.Downscale(scale_min=0.1, scale_max=0.15),\n      ], p=0.2),\n      albumentations.IAAPiecewiseAffine(p=0.2),\n      albumentations.IAASharpen(p=0.2),\n        #tạo ra những ô vuông nhỏ trong ảnh để ngăn overfitting  \n      albumentations.Cutout(max_h_size=int(CFG.img_size * 0.1), max_w_size=int(CFG.img_size * 0.1), num_holes=5, p=0.5),\n    ])\nelse:\n    transform = None","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Ví dụ về `albumentations`","metadata":{}},{"cell_type":"code","source":"figure, axes = plt.subplots(5, 5, figsize=[15, 15])\naxes = axes.reshape(-1,)\n\nif transform is None:\n    for i in range(len(axes)):\n        image = tf.io.read_file(os.path.join(CFG.root, df.index[i]))\n        image = tf.image.decode_jpeg(image, channels=3)\n        image = tf.image.resize(image, [CFG.img_size, CFG.img_size])\n        image = tf.cast(image, tf.uint8)\n        \n        axes[i].imshow(image.numpy())\n        axes[i].axis('off')\n\nelse:\n    image = tf.io.read_file(os.path.join(CFG.root, df.index[CFG.seed]))\n    image = tf.image.decode_jpeg(image, channels=3)\n    image = tf.image.resize(image, [CFG.img_size, CFG.img_size])\n    image = tf.cast(image, tf.uint8)\n\n    for i in range(len(axes)):\n        axes[i].imshow(transform(image=image.numpy())['image'])\n        axes[i].axis('off')\n\nplt.show()","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 5. Making TFRecords\n### Helper functions (serialization)","metadata":{}},{"cell_type":"code","source":"#decode ảnh sang ma trận, chuyển size và cast sang dạng unit8, tạo augmentation\ndef _serialize_image(path, transform=None):\n    image = tf.io.read_file(path)\n    image = tf.image.decode_jpeg(image, channels=3)\n    image = tf.image.resize(image, [CFG.img_size, CFG.img_size])\n    image = tf.cast(image, tf.uint8)\n    \n    if transform is not None:\n        #tạo Augmentation cho ảnh\n        image = transform(image=image.numpy())['image']\n        \n    return tf.image.encode_jpeg(image).numpy()\n\n\ndef _serialize_sample(image, image_name, label):\n    feature = {\n        'image': tf.train.Feature(bytes_list=tf.train.BytesList(value=[image])),\n        'image_name': tf.train.Feature(bytes_list=tf.train.BytesList(value=[image_name])),\n        'complex': tf.train.Feature(int64_list=tf.train.Int64List(value=[label[0]])),\n        'frog_eye_leaf_spot': tf.train.Feature(int64_list=tf.train.Int64List(value=[label[1]])),\n        'powdery_mildew': tf.train.Feature(int64_list=tf.train.Int64List(value=[label[2]])),\n        'rust': tf.train.Feature(int64_list=tf.train.Int64List(value=[label[3]])),\n        'scab': tf.train.Feature(int64_list=tf.train.Int64List(value=[label[4]])),\n        'healthy': tf.train.Feature(int64_list=tf.train.Int64List(value=[label[5]]))}\n    sample = tf.train.Example(features=tf.train.Features(feature=feature))#định dạng lưu trữ tfrec, trong đó có các feature\n    return sample.SerializeToString()\n\n\ndef serialize_fold(fold, name, transform=None, bar=None):\n    samples = []\n    #Duyệt các ảnh trong fold\n    for image_name, labels in fold.iterrows():\n        path = os.path.join(CFG.root, image_name)#lấy đường dẫn ảnh\n        image = _serialize_image(path, transform=transform)\n        samples.append(_serialize_sample(image, image_name.encode(), labels))\n    \n    #viết vào file\n    with tf.io.TFRecordWriter(name + '.tfrec') as writer:\n        [writer.write(x) for x in samples]\n        \n    if bar is not None:\n        bar.update(1)","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"total = CFG.folds * CFG.subfolds if transform is None else CFG.folds * CFG.subfolds * CFG.epochs\n\nwith tqdm(total=total) as bar:\n\n    for i in range(CFG.folds):\n\n        df_fold = df[fold == i]\n        #print(df_fold)\n        \n        folder = f'fold_{i}'\n        \n        #tạo folder cho các fold\n        try:\n            os.mkdir(folder)\n        except FileExistsError:\n            shutil.rmtree(folder)\n            os.mkdir(folder)\n        \n        #chạy 80 lần khi ko tạo augmentation\n        if transform is None:\n            for k, subfold in enumerate(np.array_split(df_fold, CFG.subfolds)):\n                name=os.path.join(folder, '%.2i-%.3i' % (k, len(subfold)))\n                serialize_fold(subfold, name=name, bar=bar)\n                \n        #chạy total (400) lần do x5 lần tạo augmentation \n        else:\n            for j in range(CFG.epochs):\n                for k, subfold in enumerate(np.array_split(df_fold, CFG.subfolds)):\n                    name=os.path.join(folder, '%.2i-%.3i' % (j * CFG.subfolds + k, len(subfold)))\n                    serialize_fold(subfold, name=name, transform=transform, bar=bar)\n#                     print(k,len(subfolds))\n#                     print(name)","metadata":{"trusted":true},"execution_count":null,"outputs":[]}]}