{"cells":[{"metadata":{},"cell_type":"markdown","source":"# order charactors by clustering columns "},{"metadata":{},"cell_type":"markdown","source":"Hello, kagglers!\n\nMy idea for dealing with the task is ordering charactors at first. Because sequences of charactors are seamless which means orders can help us to predict charactors from its relations. \n\nJapanese sentences can be wrote longitudinal (from up-right to down left). Columns could possibly  clustered and ordered then charactors are ordered one-dimensional.\n\nFirst image '100241706_00004_2.jpg' is successful as below. \n\n'自序 若い時の気強に己やれと思ふ た細工も老武者のかなしさは 息子に及ず浮世を裏の三畳に 避て正風の俳諧を楽しめ ども根が職人の文盲だけこそけれ '\n\nHowever, generalization of this pipeline has something wrong. In this trial, 3 documents in 10 was not successful. In addition, execution time for this pipeline is long.\n\nI keep going to get more sophisticated clustering.\n\n(20190820) I found better solution. With the solution, 19 in 20 seem to be successful. Now, we can get orders of charactors and use them for learning."},{"metadata":{},"cell_type":"markdown","source":"## index\n- load data\n- split unicode and coordinate\n- convert box coordinate to center coordinate\n- k-mean clustering for each column\n- generate full sentence string\n- generalize above"},{"metadata":{"trusted":true},"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport os\nfrom tqdm import tqdm\nimport cv2\nimport matplotlib.pyplot as plt\nimport time","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"FOLDER = '/kaggle/input/'\nIMAGES = FOLDER + 'train_images/'\nprint(os.listdir(FOLDER))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"###  load data"},{"metadata":{"trusted":true},"cell_type":"code","source":"df_train = pd.read_csv(FOLDER + 'train.csv')\ndf_train_idx = df_train.set_index(\"image_id\")\nidx_train = df_train['image_id']\nunicode_map = {codepoint: char for codepoint, char in pd.read_csv(FOLDER + 'unicode_translation.csv').values}","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### split unicode and coordinate\nlabels are aligned as 'unicode, x, y, w, t, unicode, x, y, w, h,,,' which can be reshaped matrix (num_char, (code, pos(4-dim))) or (-1, 5)"},{"metadata":{"trusted":true},"cell_type":"code","source":"def label_reader(label):\n    try:\n        code_arr = np.array(label['labels'].split(' ')).reshape(-1, 5)\n    except:\n        return\n    return code_arr","execution_count":null,"outputs":[]},{"metadata":{"scrolled":true,"trusted":true},"cell_type":"code","source":"idx = idx_train[0]\ndf_code = pd.DataFrame(label_reader(df_train_idx.loc[idx]))\ndf_code['image_id'] = idx\ndf_code.columns = ['char', 'x', 'y', 'w', 'h', 'image_id']\ndf_code[['x', 'y', 'w', 'h']] = df_code[['x', 'y', 'w', 'h']].astype('int')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### convert box coordinate to center coordinate\n\nCharactors order may be important to predict charactors, not only from those shape. Becuase our task is not a simple object detection, charactors are wrote seamlessly.\n\nmy idea is to order charactors at first, focusing those center. These can be obtained by converting coordinate as below."},{"metadata":{"trusted":true},"cell_type":"code","source":"def get_center(coord):\n    return np.vstack([coord[:, 0] + coord[:, 2] //2, coord[:, 1] + coord[:, 3] //2]).T","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"coord = df_code.query('image_id == \"{}\"'.format(idx))[['x', 'y','w','h']].values\ncenters =get_center(coord)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"let's show coordinate the centers. As you know, Japanese sentences can be wrote 'columns-wise' which is red from up-right to down-left, especially in historical documents."},{"metadata":{"trusted":true},"cell_type":"code","source":"image_path = IMAGES + idx + '.jpg'\nimg = cv2.imread(image_path)\nimg = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)\nplt.scatter(centers[:,0], centers[:,1])\nplt.imshow(img)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"###  k-mean clustering for each column\nit looks easy to cluster the columns if we find num of clusters. However, num of columns is different for each document. I'd like to casually use k-mean method, but num of cluster muse be specified before applying the method. \n\nMy simple idea is below\n- Transverse variance in the same columns might be small and the variance between columns might be large. \n- Longitudinal distance is small compered to transverse distance. \n- It's bclustering for each column \netter to get longitudinal scale small which means shrinking columns to split cluster transverse-wise easily.\n\nAccodring to above, several k-mean clusterings are carried out for searching for better num of clusters. Idea for the function is that difference of SD shall be large when column clustering get to be successfull while searching n_clusters from small number."},{"metadata":{"trusted":true},"cell_type":"code","source":"from sklearn.cluster import KMeans","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def get_cluster_n(centers, min_n=3, max_n=10):\n    stds_list = []\n    for n in range(min_n, max_n):\n        X = centers.copy()\n        X[:, 1] = X[:, 1]/100\n\n        df_center = pd.DataFrame(centers)\n        df_center['col_n'] = KMeans(n_clusters=n).fit(X).labels_\n        stds_list.append(df_center.groupby('col_n').std().mean().values)\n\n    stds = np.array(stds_list)\n    xsm = np.log(stds[:,0])\n    n_xsm = np.argmin(xsm[1:] - xsm[:-1]) + 1\n    \n    return n_xsm + min_n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"get_cluster_n(centers)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"n = get_cluster_n(centers)\nX = centers.copy().astype('float')\nX[:, 1] = X[:, 1]/100\ndf_center = pd.DataFrame(centers)\ndf_center['col_n'] = KMeans(n_clusters=n).fit(X).labels_\ncols = df_center['col_n'].unique()\nfor col in cols:\n    temp = df_center.query('col_n == {}'.format(col))\n    plt.scatter(temp[0], temp[1])\n    plt.imshow(img)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### generate full sentence string\n\nLooks good. Total 6 columns are there in this document, including title column. The we can get ordered string as a sentence."},{"metadata":{"trusted":true},"cell_type":"code","source":"df_center['char'] = df_code.query('image_id == \"{}\"'.format(idx))['char'] # add unicode\ncols = df_center.sort_values(0, ascending=False)['col_n'].unique() # sort by center_x because clustering labels are random.\nchars = []\nfor col in cols:\n    chars.extend(df_center.query('col_n == {}'.format(col)).sort_values(1)['char'].replace(unicode_map))\n    chars.append(' ')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"string = ''\nfor c in chars:\n    string += c","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"string","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"It's difficult to understand this poem, even though I'm a Japanese.\n\nThe meaning may be below,????\n\n\"Looking back. how sad old warrior feels, even though he'd considered he was strong in his young, is strong as he is craftsman-like and　illiteracy even though he has fun from ideal poem behind his son.\"\n\nIt's diffcult..."},{"metadata":{"trusted":true},"cell_type":"markdown","source":"## generalize above\nGeneralizing can be done easily, with iteration of 'image_id' below. However, the splitting columns is not successfull, such as '100241706_00007_2'\n\nI need more sophisticated clustering method.."},{"metadata":{"trusted":true},"cell_type":"code","source":"def gen_df_code(df_idx, idx):\n    df_code = pd.DataFrame(label_reader(df_idx.loc[idx]), columns = ['char', 'x', 'y', 'w', 'h'])\n    df_code['image_id'] = idx\n    df_code = df_code.reset_index()\n    df_code[['x','y','w','h']] = df_code[['x','y','w','h']].astype('int')\n\n    centers = get_center(df_code[['x','y','w','h']].values)\n    df_code[['center_x', 'center_y']] = pd.DataFrame(centers)\n\n    X = centers.copy().astype('float')\n    X[:, 1] = X[:, 1]/100\n    df_code['col_n'] =  KMeans(n_clusters=get_cluster_n(centers)).fit(X).labels_\n    \n    new_col_n = np.zeros(0)\n    new_index = np.zeros(0)\n    cols = df_code.sort_values('center_x', ascending=False)['col_n'].unique()\n    for i, col in enumerate(cols):\n        temp = df_code.query('col_n == {}'.format(col))\n        new_index = np.hstack([new_index, temp['index'].values])\n        new_col_n = np.hstack([new_col_n, np.ones(len(temp)) * i])\n\n    del df_code['col_n']\n    df_new_idx = pd.DataFrame([new_index, new_col_n]).T\n    df_new_idx.columns = ['index', 'col_n']\n    df_code = pd.merge(df_code, df_new_idx, on='index').sort_values('col_n').reset_index(drop=True)\n    del df_code['index']\n    df_code['col_n'] = df_code['col_n'].astype('int')\n\n    return df_code","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def gen_string(df_code):\n    cols = df_code['col_n'].unique()\n    chars = []\n    for col in cols:\n        chars.extend(df_code.query('col_n == {}'.format(col)).sort_values('center_y')['char'].replace(unicode_map))\n        chars.append(' ')\n\n    string = ''\n    for c in chars:\n        string += c\n\n    print(string)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"for idx in tqdm(idx_train[:40]):\n    df_code = gen_df_code(df_train_idx, idx)\n    gen_string(df_code)\n\n    image_path = IMAGES + idx + '.jpg'\n    img = cv2.imread(image_path)\n    img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)\n    cols = df_code['col_n'].unique()\n    for col in cols:\n        centers = df_code.query('col_n == {}'.format(col))[['center_x','center_y']].values\n        plt.scatter(centers[:,0], centers[:,1])\n    plt.imshow(img)\n    plt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.7.3"}},"nbformat":4,"nbformat_minor":1}