{"cells":[{"metadata":{"_uuid":"7bf57fcfa23f5332c15412cbe71a3954539be259"},"cell_type":"markdown","source":"In the pursuit of a nice juicy \"leakage\" I give you a simple and straightforward text detection facility using plain OpenCV (EAST text detection caffe model).It manages to automatically locate text in both train and test images. The majority of the code was taken from the great Pyimagesearh [OpenCV Text Detection (EAST text detector)](https://www.pyimagesearch.com/2018/08/20/opencv-text-detection-east-text-detector/). No character/word recognition yet. I believe there must be a correlation between text appearance and `new_whale` tag.\n\nTo reduce false positives we restrict our search to the the bottom area of each image (see `offsetY/H < 0.80` below)"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport numpy as np\nimport pandas as pd\nimport cv2\nimport os\n\nfrom matplotlib import pyplot as plt\nfrom tqdm import tqdm_notebook\nfrom glob import glob\nimport multiprocessing\n\nimport os\n \n# Any results you write to the current directory are saved as output.","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"whale = pd.read_csv('../input/humpback-whale-identification/train.csv')\nwhale.head()\nprint(len(whale))\nlen(np.unique(whale.Id))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"edcb6bea226697b27d35a978ecf9e825a8c07692"},"cell_type":"code","source":"unknown_whale = whale[whale.Id=='new_whale']\nunknown_whale.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"7ede04b40c1d50c3958b4a663e7c749452c847de"},"cell_type":"code","source":"train_path = '../input/humpback-whale-identification/train/'\ntrain_images = unknown_whale.Image.values#os.listdir(train_path)\ntest_path = '../input/humpback-whale-identification/test/'\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"89ffce775272e2ffcb9e563a927df43512ae3e3c"},"cell_type":"code","source":"whale_dict = dict(zip(whale.Image, whale.Id))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"6b40f306be21c70d5052908d845c9c429656cc60"},"cell_type":"code","source":"layerNames = [\n\t\"feature_fusion/Conv_7/Sigmoid\",\n\t\"feature_fusion/concat_3\"]\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"24f4d8549a15d457b07e46c8fece16bd023b1bb0"},"cell_type":"code","source":"!pip install imutils","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ba1e7108821aa546b31c2c971e2f3dcd1d75d7bb"},"cell_type":"markdown","source":"The architecture of EAST text detector is depicted in the image below. Among the outputs is a set of text  boxes. We loop through this test for increasing `y` values and we can get an estimate of the number of lines.  \n![image](https://www.pyimagesearch.com/wp-content/uploads/2018/08/opencv_text_detection_east.jpg)\n"},{"metadata":{"trusted":true,"_uuid":"a44b223d79d14f649ad6056d962b587efd26f7f8"},"cell_type":"code","source":"import time\nfrom imutils.object_detection import non_max_suppression\nnet = cv2.dnn.readNet('../input/frozen-east-text-detection/frozen_east_text_detection.pb')\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c04cfd7b89388069a0ce442ed3d7178b76f18c0d"},"cell_type":"code","source":"WW = 320\nHH = 160\ndef get_images_with_text(path, with_class=True, WW=320, HH=160):\n\n    image_files = os.listdir(path)\n    FOUND  = []\n    new_whale_count = 0\n    for image_file in tqdm_notebook(image_files):\n\n        # load the input image and grab the image dimensions\n        image = cv2.imread(path + image_file)\n        orig = image.copy()\n        (H, W) = image.shape[:2]\n\n        # set the new width and height and then determine the ratio in change\n        # for both the width and height\n        (newW, newH) = (WW, HH)\n        rW = W / float(newW)\n        rH = H / float(newH)\n\n        # resize the image and grab the new image dimensions\n        image = cv2.resize(image, (newW, newH))\n        (H, W) = image.shape[:2]\n\n\n        # construct a blob from the image and then perform a forward pass of\n        # the model to obtain the two output layer sets\n        blob = cv2.dnn.blobFromImage(image, 1.0, (W, H),\n            (123.68, 116.78, 103.94), swapRB=False, crop=False)\n        start = time.time()\n        net.setInput(blob)\n        (scores, geometry) = net.forward(layerNames)\n        end = time.time()\n\n        # show timing information on text prediction\n\n\n        (numRows, numCols) = scores.shape[2:4]\n        rects = []\n        confidences = []\n\n        text_found = 0\n        text_lines = 0\n        # loop over the number of rows\n        for y in range(0, numRows):\n\n            # extract the scores (probabilities), followed by the geometrical\n            # data used to derive potential bounding box coordinates that\n            # surround text\n            scoresData = scores[0, 0, y]\n            xData0 = geometry[0, 0, y]\n            xData1 = geometry[0, 1, y]\n            xData2 = geometry[0, 2, y]\n            xData3 = geometry[0, 3, y]\n            anglesData = geometry[0, 4, y]\n\n            # loop over the number of columns\n            found = False\n            for x in range(0, numCols):\n                # if our score does not have sufficient probability, ignore it\n                if scoresData[x] < 0.8:\n                    continue\n\n                # compute the offset factor as our resulting feature maps will\n                # be 4x smaller than the input image\n                (offsetX, offsetY) = (x * 4.0, y * 4.0)\n\n                if offsetY/H < 0.80:\n                    continue\n\n\n                # extract the rotation angle for the prediction and then\n                # compute the sin and cosine\n                angle = anglesData[x]\n                cos = np.cos(angle)\n                sin = np.sin(angle)\n\n                # use the geometry volume to derive the width and height of\n                # the bounding box\n                h = xData0[x] + xData2[x]\n                w = xData1[x] + xData3[x]\n\n                # compute both the starting and ending (x, y)-coordinates for\n                # the text prediction bounding box\n                endX = int(offsetX + (cos * xData1[x]) + (sin * xData2[x]))\n                endY = int(offsetY - (sin * xData1[x]) + (cos * xData2[x]))\n                startX = int(endX - w)\n                startY = int(endY - h)\n\n                # add the bounding box coordinates and probability score to\n                # our respective lists\n                rects.append((startX, startY, endX, endY))\n                confidences.append(scoresData[x])\n                found = True\n\n            if found == True:\n                text_lines += 1\n\n\n        boxes = non_max_suppression(np.array(rects), probs=confidences)\n        #if len(boxes)>0:\n        #    print (image_file, text_lines)\n        # loop over the bounding boxes\n        for (startX, startY, endX, endY) in boxes:\n            # scale the bounding box coordinates based on the respective\n            # ratios\n            startX = int(startX * rW)\n            startY = int(startY * rH)\n            endX = int(endX * rW)\n            endY = int(endY * rH)\n\n        if len(boxes) > 0:\n            if with_class==True:\n                FOUND.append([image_file, whale_dict[image_file], text_lines])\n                if whale_dict[image_file] == 'new_whale':\n                    new_whale_count = new_whale_count + 1\n            else:\n                FOUND.append([image_file, text_lines])\n    return FOUND","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"520821cd10abfc93770dabfa2899c5d4a6383723"},"cell_type":"code","source":"df_train = get_images_with_text(train_path)\ndf_test = get_images_with_text(test_path, with_class=False)\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"909aaa496eb3d23ce3acc463d41c1043ef8542d4"},"cell_type":"code","source":"df_train= pd.DataFrame(df_train)\ndf_train.columns = ['image', 'class', 'line_count']\ndf_train.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c0757b3891bcca2063e5d9a0f7ef0c432172efec"},"cell_type":"code","source":"df_test= pd.DataFrame(df_test)\ndf_test.columns = ['image', 'line_count']\ndf_test.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"adddf8d030e52bf2601a96d97aae4f64798bc386"},"cell_type":"code","source":"df_train.to_csv('train_text.csv')\ndf_test.to_csv('train_text.csv')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5f8be0c215374fd0ebda232706fdc0be74e52747"},"cell_type":"markdown","source":"Let us plot some training images along with class annotations and number of lines."},{"metadata":{"trusted":true,"scrolled":false,"_uuid":"3e619b85a9a0ceea8b4e18bf2b9966dbdc8540c1"},"cell_type":"code","source":"fig, axes = plt.subplots(5, 5)\n \nfig.set_figwidth(20)\nfig.set_figheight(20)\n\nfor i, row in df_train.iterrows():\n    if i >= 25:\n        break\n    img = cv2.imread(train_path + row['image'])\n    axes[int(i/5), i%5].imshow(img)\n    axes[int(i/5), i%5].set_title(row['image']  + '-' + str(whale_dict[row['image']]) + ' (' + str(row['line_count']) + ')')\n    axes[int(i/5), i%5].axis('off')\n\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"59987069f8f12ed6ecd0fed496948bf3c5e6d988"},"cell_type":"markdown","source":"Let us plot some tet images along with the  number of lines. It seems that text detector is producing more false positives."},{"metadata":{"trusted":true,"_uuid":"142751b8acf4e0719d88f9e34dfedfdb29319acc"},"cell_type":"code","source":"fig, axes = plt.subplots(5, 5)\n \nfig.set_figwidth(20)\nfig.set_figheight(20)\n\nfor i, row in df_test.iterrows():\n    if i >= 25:\n        break\n    img = cv2.imread(test_path + row['image'])\n    axes[int(i/5), i%5].imshow(img)\n    axes[int(i/5), i%5].set_title( row['image']  + '-' +  ' (' + str(row['line_count']) + ')')\n    axes[int(i/5), i%5].axis('off')\n\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"61e6d9ee392a8182ee43cc39b397447fa780d8c6"},"cell_type":"markdown","source":"## Some observations:\n### 1. By looking at the training images there seems to be great correlation between images with many lines of text  and identified whales"},{"metadata":{"trusted":true,"_uuid":"19fe73a04b02911dbe4df450bb0ed2335ef97579"},"cell_type":"code","source":"df_train[df_train.line_count>3]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"4626c68bdad821739dbd4953fa556908d673ffcd"},"cell_type":"markdown","source":"### 2. Text that starts with '#' followed by a four digit number is most probably new_whale (see  #3332, #0518 in training images above) "},{"metadata":{"trusted":true,"_uuid":"beb68dadf15f5e0b650e698a3ca983b54d6bd060"},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}