{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Wikipedia - Image/Caption Matching EDA\n\nThis looks like a very interesting competition (too bad it doesn't award points!!). Let's try to take a look at the data and try to submit a super simple baseline to check if our understanding is fine.\n\n**Note this is a new version, previously we've helped discover some issues with test data :)**","metadata":{}},{"cell_type":"code","source":"!pip install rapidfuzz -qq","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2021-10-26T20:05:02.437225Z","iopub.execute_input":"2021-10-26T20:05:02.437807Z","iopub.status.idle":"2021-10-26T20:05:13.659255Z","shell.execute_reply.started":"2021-10-26T20:05:02.437703Z","shell.execute_reply":"2021-10-26T20:05:13.658059Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Test data\n\nLet's start with **sample submission**. We have an *id* column and *caption_title_and_reference_description* column, and for each id we predict 5 captions. We need to select them from a predefined set of captions we'll see next.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nsub = pd.read_csv('../input/wikipedia-image-caption/sample_submission.csv')\nsub.head(10)","metadata":{"execution":{"iopub.status.busy":"2021-10-26T20:05:13.661476Z","iopub.execute_input":"2021-10-26T20:05:13.661751Z","iopub.status.idle":"2021-10-26T20:05:14.176291Z","shell.execute_reply.started":"2021-10-26T20:05:13.661720Z","shell.execute_reply":"2021-10-26T20:05:14.175386Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"captions = pd.read_csv('../input/wikipedia-image-caption/test_caption_list.csv')\nprint(len(captions))\ncaptions.head()","metadata":{"execution":{"iopub.status.busy":"2021-10-26T20:05:14.177300Z","iopub.execute_input":"2021-10-26T20:05:14.177510Z","iopub.status.idle":"2021-10-26T20:05:14.563409Z","shell.execute_reply.started":"2021-10-26T20:05:14.177487Z","shell.execute_reply":"2021-10-26T20:05:14.562851Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The test file contains a list of id's and image urls. Let's print a few of those urls. ","metadata":{}},{"cell_type":"code","source":"test = pd.read_csv('../input/wikipedia-image-caption/test.tsv', sep='\\t')\ntest.head()","metadata":{"execution":{"iopub.status.busy":"2021-10-26T20:05:14.564815Z","iopub.execute_input":"2021-10-26T20:05:14.565007Z","iopub.status.idle":"2021-10-26T20:05:14.780997Z","shell.execute_reply.started":"2021-10-26T20:05:14.564984Z","shell.execute_reply":"2021-10-26T20:05:14.780158Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i in range(5):\n    print(test.image_url.loc[i])","metadata":{"execution":{"iopub.status.busy":"2021-10-26T20:05:14.782192Z","iopub.execute_input":"2021-10-26T20:05:14.782426Z","iopub.status.idle":"2021-10-26T20:05:14.792047Z","shell.execute_reply.started":"2021-10-26T20:05:14.782401Z","shell.execute_reply":"2021-10-26T20:05:14.791127Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's take a look at some of the **test images**. It looks like we have the test image data in 5 csv files, and these include links to the images (upload and commons varieties) as well as the base64 string encoded version of the images. ","metadata":{}},{"cell_type":"code","source":"ls ../input/wikipedia-image-caption/image_data_test/image_pixels","metadata":{"execution":{"iopub.status.busy":"2021-10-26T20:05:14.793196Z","iopub.execute_input":"2021-10-26T20:05:14.793439Z","iopub.status.idle":"2021-10-26T20:05:15.598299Z","shell.execute_reply.started":"2021-10-26T20:05:14.793412Z","shell.execute_reply":"2021-10-26T20:05:15.597569Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tst0 = pd.read_csv('../input/wikipedia-image-caption/image_data_test/image_pixels/test_image_pixels_part-00000.csv', sep='\\t', header=None)\ntst0.head()","metadata":{"execution":{"iopub.status.busy":"2021-10-26T20:05:15.599549Z","iopub.execute_input":"2021-10-26T20:05:15.600033Z","iopub.status.idle":"2021-10-26T20:05:22.472562Z","shell.execute_reply.started":"2021-10-26T20:05:15.599995Z","shell.execute_reply":"2021-10-26T20:05:22.471671Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import base64 \nfrom PIL import Image\nimport io\n\nimage_64_decode = base64.b64decode(tst0[1].loc[0])\nimg = Image.open(io.BytesIO(image_64_decode))\nimg","metadata":{"execution":{"iopub.status.busy":"2021-10-26T20:05:22.473827Z","iopub.execute_input":"2021-10-26T20:05:22.474147Z","iopub.status.idle":"2021-10-26T20:05:22.532115Z","shell.execute_reply.started":"2021-10-26T20:05:22.474107Z","shell.execute_reply":"2021-10-26T20:05:22.531175Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Train data\n\nThanks to [this kernel](https://www.kaggle.com/udbhavpangotra/reading-the-data-datatable-works-like-a-charm) for showing how to accelerate data reading with datatable, please give it an upvote! Here we have a bunch of columns including page and image urls, as well as what looks as our target: caption_title_and_reference_description. There is a [SEP] string at the end of each entry here, or in the middle if multiple entries are provided. Not fully sure where this is coming from. \n","metadata":{}},{"cell_type":"code","source":"import datatable as dt\ntrain0 = dt.fread('../input/wikipedia-image-caption/train-00000-of-00005.tsv')\ntrain0.head()","metadata":{"execution":{"iopub.status.busy":"2021-10-26T20:05:22.533066Z","iopub.execute_input":"2021-10-26T20:05:22.533297Z","iopub.status.idle":"2021-10-26T20:06:32.784132Z","shell.execute_reply.started":"2021-10-26T20:05:22.533273Z","shell.execute_reply":"2021-10-26T20:06:32.783188Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Baseline\n\nLooking at the train data, it seems that there is a connect between page url and our target. We also have page url in our test data, so let's try to exploit that and make a caption prediction only based on page url, without looking at the image. We'll verify that heuristic on train data, then we'll try to fuzzy match that caption prediction with the list of test captions.","metadata":{}},{"cell_type":"code","source":"from urllib.parse import unquote\n\nt = test.image_url.loc[2]\n\ndef convert(t):\n    t = t.rsplit('/',1)[1]\n    t = unquote(t)\n    t = t.replace('_', ' ')\n    t = t + ' [SEP]'\n    return(t)","metadata":{"execution":{"iopub.status.busy":"2021-10-26T20:06:32.786727Z","iopub.execute_input":"2021-10-26T20:06:32.787004Z","iopub.status.idle":"2021-10-26T20:06:32.793318Z","shell.execute_reply.started":"2021-10-26T20:06:32.786979Z","shell.execute_reply":"2021-10-26T20:06:32.791742Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i in range(5):\n    print(f'target: {train0[i,-1]}')\n    print(f'prediction: {convert(train0[i,1])}')\n    print()","metadata":{"execution":{"iopub.status.busy":"2021-10-26T20:06:32.794727Z","iopub.execute_input":"2021-10-26T20:06:32.794955Z","iopub.status.idle":"2021-10-26T20:06:32.809829Z","shell.execute_reply.started":"2021-10-26T20:06:32.794931Z","shell.execute_reply":"2021-10-26T20:06:32.808674Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test['prediction'] = test['image_url'].apply(convert)","metadata":{"execution":{"iopub.status.busy":"2021-10-26T20:06:32.811883Z","iopub.execute_input":"2021-10-26T20:06:32.812339Z","iopub.status.idle":"2021-10-26T20:06:33.112654Z","shell.execute_reply.started":"2021-10-26T20:06:32.812308Z","shell.execute_reply":"2021-10-26T20:06:33.111747Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test.head()","metadata":{"execution":{"iopub.status.busy":"2021-10-26T20:06:33.114281Z","iopub.execute_input":"2021-10-26T20:06:33.114506Z","iopub.status.idle":"2021-10-26T20:06:33.127101Z","shell.execute_reply.started":"2021-10-26T20:06:33.114479Z","shell.execute_reply":"2021-10-26T20:06:33.126051Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"CAPTIONS = captions.caption_title_and_reference_description.values.tolist()\nlen(CAPTIONS)","metadata":{"execution":{"iopub.status.busy":"2021-10-26T20:06:33.128732Z","iopub.execute_input":"2021-10-26T20:06:33.129020Z","iopub.status.idle":"2021-10-26T20:06:33.145234Z","shell.execute_reply.started":"2021-10-26T20:06:33.128994Z","shell.execute_reply":"2021-10-26T20:06:33.144282Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from rapidfuzz import process, fuzz","metadata":{"execution":{"iopub.status.busy":"2021-10-26T20:06:33.146408Z","iopub.execute_input":"2021-10-26T20:06:33.146651Z","iopub.status.idle":"2021-10-26T20:06:33.253862Z","shell.execute_reply.started":"2021-10-26T20:06:33.146615Z","shell.execute_reply":"2021-10-26T20:06:33.253190Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\nfor i in range(5):\n    s = test.prediction.loc[i]\n    print(f'image_url: {s}')\n    res = process.extract(s, CAPTIONS, scorer=fuzz.ratio, processor=None, limit=5)\n    print(f'closest captions:')\n    for c in res:\n        print(c[0])\n    print('*'*60)\n    print()  ","metadata":{"execution":{"iopub.status.busy":"2021-10-26T20:06:33.254875Z","iopub.execute_input":"2021-10-26T20:06:33.255656Z","iopub.status.idle":"2021-10-26T20:06:33.354955Z","shell.execute_reply.started":"2021-10-26T20:06:33.255626Z","shell.execute_reply":"2021-10-26T20:06:33.353966Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def find_closest_match(s):\n    res = process.extract(s, CAPTIONS, scorer=fuzz.ratio, processor=None, limit=5)\n    res = [x[0] for x in res]\n    return res","metadata":{"execution":{"iopub.status.busy":"2021-10-26T20:06:33.356147Z","iopub.execute_input":"2021-10-26T20:06:33.356374Z","iopub.status.idle":"2021-10-26T20:06:33.360630Z","shell.execute_reply.started":"2021-10-26T20:06:33.356348Z","shell.execute_reply":"2021-10-26T20:06:33.359936Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from tqdm.auto import tqdm\ntqdm.pandas()","metadata":{"execution":{"iopub.status.busy":"2021-10-26T20:06:33.361355Z","iopub.execute_input":"2021-10-26T20:06:33.361649Z","iopub.status.idle":"2021-10-26T20:06:33.446758Z","shell.execute_reply.started":"2021-10-26T20:06:33.361621Z","shell.execute_reply":"2021-10-26T20:06:33.445588Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test['caption_title_and_reference_description'] = test['prediction'].progress_apply(find_closest_match)","metadata":{"execution":{"iopub.status.busy":"2021-10-26T20:06:33.447947Z","iopub.execute_input":"2021-10-26T20:06:33.448261Z","iopub.status.idle":"2021-10-26T20:35:07.804464Z","shell.execute_reply.started":"2021-10-26T20:06:33.448221Z","shell.execute_reply":"2021-10-26T20:35:07.802256Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub = test[['id', 'caption_title_and_reference_description']]\nsub = sub.explode('caption_title_and_reference_description')\nsub.head()","metadata":{"execution":{"iopub.status.busy":"2021-10-26T20:35:07.806019Z","iopub.execute_input":"2021-10-26T20:35:07.806624Z","iopub.status.idle":"2021-10-26T20:35:07.984895Z","shell.execute_reply.started":"2021-10-26T20:35:07.806591Z","shell.execute_reply":"2021-10-26T20:35:07.984119Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2021-10-26T20:35:07.986238Z","iopub.execute_input":"2021-10-26T20:35:07.986610Z","iopub.status.idle":"2021-10-26T20:35:08.987525Z","shell.execute_reply.started":"2021-10-26T20:35:07.986569Z","shell.execute_reply":"2021-10-26T20:35:08.986599Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We've now been able to score above 0.0000 on the leaderboard. To be honest, I'm not sure if using page_url is expected by the host so I asked that question on the forum. For sure, the more challenging and interesting aspect is matching the captions directly with images, and we'll try to tackle that next :) ","metadata":{}},{"cell_type":"markdown","source":"## to be continued ...","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}