{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Kuzushiji Page Generator\n### This notebook introduces a new way to automatically generate pages filled with Kuzushiji symbols using the great [KMNIST dataset](https://github.com/rois-codh/kmnist). The aim is to help pretraining models learning both character detection and classification at the same time before moving to the pages from the original competition dataset. Of course, the resulting pages would not make any sense!\n\n### I have generated 20,000 pages using this script and made them available in the [following dataset](https://www.kaggle.com/frlemarchand/synthetic-kmnist-pages).\n\n### If you find this notebook useful, please feel free to give it an upvote!"},{"metadata":{},"cell_type":"markdown","source":"version 9 notes:\n* add variation in the symbol size within the same page\n* fix bug in the generation of the groundtruth .csv file\n* the groundtruth .csv file now follow the exact same format as the [Kuzushiji competition dataset](https://www.kaggle.com/c/kuzushiji-recognition/data)"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport cv2\nimport os\nimport matplotlib.pyplot as plt\nimport random\nfrom PIL import Image\nfrom tqdm import tqdm","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Extract datasets from the compressed files and load them"},{"metadata":{"trusted":true},"cell_type":"code","source":"!unzip ../input/kuzushiji/k49-train-imgs.npz && mv ../working/arr_0.npy k49-train-imgs.npy\n!unzip ../input/kuzushiji/k49-train-labels.npz && mv ../working/arr_0.npy k49-train-labels.npy\n!unzip ../input/kuzushiji/k49-train-imgs.npz && mv ../working/arr_0.npy k49-test-imgs.npy\n!unzip ../input/kuzushiji/k49-train-labels.npz && mv ../working/arr_0.npy k49-test-labels.npy","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"!unzip ../input/kuzushiji/kmnist-train-imgs.npz && mv ../working/arr_0.npy kmnist-train-imgs.npy\n!unzip ../input/kuzushiji/kmnist-train-labels.npz && mv ../working/arr_0.npy kmnist-train-labels.npy\n!unzip ../input/kuzushiji/kmnist-train-imgs.npz && mv ../working/arr_0.npy kmnist-test-imgs.npy\n!unzip ../input/kuzushiji/kmnist-train-labels.npz && mv ../working/arr_0.npy kmnist-test-labels.npy","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"os.listdir(\"../input/kuzushiji\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"k49 = np.load('../working/k49-train-imgs.npy')\nk49_labels = np.load('../working/k49-train-labels.npy')\nk49_mapping = pd.read_csv(\"../input/kuzushiji/k49_classmap.csv\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"kmnist = np.load('../working/kmnist-train-imgs.npy')\nkmnist_labels = np.load('../working/kmnist-train-labels.npy')\nkmnist_mapping = pd.read_csv(\"../input/kuzushiji/kmnist_classmap.csv\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Functions to generate a symbol from one of the three subsets"},{"metadata":{"trusted":true},"cell_type":"code","source":"def get_k49(show=False):\n    idx = random.randint(0,len(k49)-1)\n    img = k49[idx]\n    if show:\n        plt.imshow(img)\n        plt.show()\n    return img, k49_mapping.iloc[k49_labels[idx]].codepoint","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sample = get_k49(show=True)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def get_kmnist(show=False):\n    idx = random.randint(0,len(kmnist)-1)\n    img = kmnist[idx]\n    if show:\n        plt.imshow(img)\n        plt.show()\n    return img, kmnist_mapping.iloc[kmnist_labels[idx]].codepoint","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sample = get_kmnist(show=True)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def get_kuzushiji_kanji(show=False):\n    kanji_list = os.listdir(\"../input/kuzushiji/kkanji/kkanji2/\")\n    selected_kanji = random.choice(kanji_list)\n    image_list = os.listdir(\"../input/kuzushiji/kkanji/kkanji2/\"+selected_kanji)\n    selected_image = random.choice(image_list)\n    image=cv2.imread(\"../input/kuzushiji/kkanji/kkanji2/{}/{}\".format(selected_kanji,selected_image))\n    if show:\n        plt.imshow(image)\n        plt.show()\n    return image, selected_kanji","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sample = get_kuzushiji_kanji(show=True)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Create a synthetic Kuzushiji-filled page"},{"metadata":{"trusted":true},"cell_type":"code","source":"def get_new_page(page_dimensions = (3900,2400), binary_mask=True):\n    \n    page = np.zeros(page_dimensions)\n    labels = \"\"\n    \n    number_of_columns = random.randint(3,8)\n    symbols_per_columns = random.randint(10,20)\n    margin = random.randint(30,200)\n    symbol_size = random.randint(100,250)\n    \n    for row in range(1,symbols_per_columns):\n        for col in range(1,number_of_columns):\n            x_location = int((page.shape[1]-margin*2)*col/number_of_columns)\n            y_location = int((page.shape[0]-margin*2)*row/symbols_per_columns)\n            symbol_size_variation = random.randint(0,10)\n            #randomly pick a subtype from the KMNIST dataset.\n            condition = random.randint(1,3)\n            if condition==1:\n                symbol, label = get_kmnist()\n                symbol = cv2.resize(symbol, (symbol_size-symbol_size_variation, symbol_size-symbol_size_variation)) \n                if binary_mask:\n                    ret,symbol = cv2.threshold(symbol.astype('uint8'), 0, 255, cv2.THRESH_BINARY | cv2.THRESH_OTSU)\n                page[y_location:y_location+symbol_size-symbol_size_variation,x_location:x_location+symbol_size-symbol_size_variation] = symbol\n            elif condition==2:\n                symbol, label = get_k49()\n                symbol = cv2.resize(symbol, (symbol_size-symbol_size_variation, symbol_size-symbol_size_variation)) \n                if binary_mask:\n                    ret,symbol = cv2.threshold(symbol.astype('uint8'), 0, 255, cv2.THRESH_BINARY | cv2.THRESH_OTSU)\n                page[y_location:y_location+symbol_size-symbol_size_variation,x_location:x_location+symbol_size-symbol_size_variation] = symbol\n            elif condition==3:\n                symbol, label = get_kuzushiji_kanji()\n                symbol = cv2.resize(symbol, (symbol_size-symbol_size_variation, symbol_size-symbol_size_variation)) \n                symbol = symbol[:,:,0]\n                if binary_mask:\n                    ret,symbol = cv2.threshold(symbol.astype('uint8'), 0, 255, cv2.THRESH_BINARY | cv2.THRESH_OTSU)\n                page[y_location:y_location+symbol_size-symbol_size_variation,x_location:x_location+symbol_size-symbol_size_variation] = symbol\n            #Bug fixed in version 9. \n            labels += \"{} {} {} {} {} \".format(label,str(x_location),str(y_location),str(symbol_size-symbol_size_variation),str(symbol_size-symbol_size_variation))\n\n    return page, labels","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def print_random_pages():\n    sample_number = 10\n    fig = plt.figure(figsize = (20,sample_number))\n    for i in range(0,sample_number):\n        ax = fig.add_subplot(2, 5, i+1)\n        ax.imshow(get_new_page()[0])\n    plt.tight_layout()\n    plt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Examples of binary images generated"},{"metadata":{"trusted":true},"cell_type":"code","source":"print_random_pages()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"dest_dir = \"../working/synthetic-kmnist-pages\"\nos.mkdir(dest_dir)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The dataset is saved in the working directory, as well as the dataframe into a .csv file. The number of generated pages is set to only 100 due to hard drive space limitation."},{"metadata":{"trusted":true},"cell_type":"code","source":"kuzushiji_df = pd.DataFrame(columns=[\"image_id\",\"labels\"])\n#The number of output files is limited to 500\nnumber_of_pages = 400\nwith tqdm(total=number_of_pages) as pbar:\n    for idx in range(0,number_of_pages):\n        pbar.update(1)\n        filename = \"{}.png\".format(idx)\n        binary_image, labels = get_new_page()\n        cv2.imwrite(\"{}/{}\".format(dest_dir, filename), binary_image)\n        kuzushiji_df = kuzushiji_df.append({\"image_id\":filename,\"labels\":labels}, ignore_index=True)\nkuzushiji_df.to_csv(\"{}/synthetic_kmnist_pages.csv\".format(dest_dir),index=False)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"kuzushiji_df.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Please check out the [generated dataset](https://www.kaggle.com/frlemarchand/synthetic-kmnist-pages) with 20,000 synthetic pages. If you use the dataset to improve a solution, I would love to hear about it. :)\n\n### If this contribution was of any help, please give an upvote to help me know whether this type of kernel is useful to the community!"}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":1}