{"cells":[{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"markdown","source":"As stated in the data description we cannot acces the full test set in this format of the competiton\n> This is a synchronous rerun code competition, you can assume that the complete test set will contain essentially the same size and number of images as the training set. Consider performing inference on just one batch at a time to avoid memory errors. Only the first few rows/images in the test set and sample submission files can be downloaded. These samples provided so you can review the basic structure of the files and to ensure consistency between the publicly available set of file names and those your code will have access to while it is being rerun for scoring."},{"metadata":{},"cell_type":"markdown","source":"You should structure your code so that it returns predictions for the partial test set images (12 images available in .parquet format) in the format specified by the public sample_submission.csv (which has id for the 12 images), but does not hard code aspects like the id or number of rows. When Kaggle runs your Kernel privately, it substitutes the partial test set with the full version and evaluates the model on that. This is done to prevent test set manipulation."},{"metadata":{"trusted":true},"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"for i in range(4):\n    df_test_img = pd.read_parquet('/kaggle/input/bengaliai-cv19/test_image_data_{}.parquet'.format(i))\n    print(df_test_img.shape)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"As you can see the publicly available test set has 3 images (the image pixels are in the rows) in each of the 4 parquet files. Total 3x4=12"},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"df_test = pd.read_csv('/kaggle/input/bengaliai-cv19/test.csv')\ndf_test.shape","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"And the test.csv have 3x12=36 labels"},{"metadata":{},"cell_type":"markdown","source":"Let's also look at the first few entries in the image dataframe and label dataframe to get acquainted with the data structure"},{"metadata":{"trusted":true},"cell_type":"code","source":"print('Data')\ndisplay(df_test_img.head())\nprint('label')\ndisplay(df_test.head())","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"So, to write the inference function, you have to make sure the .csv structure is similar to the given test.csv but it has to account for the fact that there are more then 12 images."},{"metadata":{},"cell_type":"markdown","source":"Here's one way to do that"},{"metadata":{"trusted":true},"cell_type":"code","source":"components = ['consonant_diacritic', 'grapheme_root', 'vowel_diacritic']\ntarget=[] # model predictions placeholder\nrow_id=[] # row_id place holder\nn_cls = [7,168,11] # number of classes in each of the 3 targets\nfor i in range(4):\n    df_test_img = pd.read_parquet('/kaggle/input/bengaliai-cv19/test_image_data_{}.parquet'.format(i)) # read image data\n    df_test_img.set_index('image_id', inplace=True) # set image_id as index value\n    for id in df_test_img.index.values: # df_test_img.index.values has the test image_ids \n        for i,comp in enumerate(components):\n            id_sample=id+'_'+comp\n            row_id.append(id_sample)\n            target.append(np.random.randint(0,n_cls[i])) # our model is a random integer generator between 0 and n_cls","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# create a dataframe with the solutions \ndf_sample = pd.DataFrame(\n    {'row_id': row_id,\n    'target':target\n    },\n    columns =['row_id','target'] \n)\ndf_sample.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# create submission file\ndf_sample.to_csv('submission.csv',index=False)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"After creating the submit file, click the `commit` button on the upper right corner and let it complete. It will run the whole code and save stable version of it.\n<img src=\"https://i.imgur.com/aACt0QZ.png\" alt=\"Smiley face\" align=\"center\" width=\"400\" height=\"500\">\n"},{"metadata":{},"cell_type":"markdown","source":"Click the `Open Version` button.\n<img src=\"https://i.imgur.com/nvEOQVR.png\" alt=\"Smiley face\" align=\"center\" width=\"700\" height=\"900\">"},{"metadata":{},"cell_type":"markdown","source":" and navigate back to your commited notebook. \n<img src=\"https://i.imgur.com/P71iGHj.png\" alt=\"Smiley face\" align=\"center\" width=\"700\" height=\"900\">"},{"metadata":{},"cell_type":"markdown","source":"Scroll down to the bottom and you will see the option to `submit to competition` button.\n<img src=\"https://i.imgur.com/j56xJLW.png\" alt=\"Smiley face\" align=\"center\" width=\"700\" height=\"500\">"},{"metadata":{},"cell_type":"markdown","source":" Clicking that button will do multiple things. It will replace the partial test set dataset with the full version, rerun the kernel and evaluate the metric.\n <img src=\"https://i.imgur.com/5HzYeiC.png\" alt=\"Smiley face\" align=\"center\" width=\"400\" height=\"700\">"},{"metadata":{},"cell_type":"markdown","source":"This is might take time as there are around 200k test images. After the computation is finished the result will be shown. It takes around 10-15 minutes for me. Sometimes the computation is finished but the running status is still there which is a bit annoying. Refresh the page regularly."}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":1}