{"cells":[{"metadata":{},"cell_type":"markdown","source":"## CSV Direct Submit\n\nIn this competition I am training remotely using the TPU quota provided by GCP. Inference is done on GPU in a Kaggle kernel. As we reach the last 10 days of the comp, even the generous allocation of 30 hours of free GPU time from Kaggle looks like it might be a bit tight.\n\nConsequently, I have started doing inference offline as well, and submitting csv predictions directly through this kernel. This has the added advantage of provideng fast feedback on the leaderboard with no kaggle GPU time consumed. Obviously this kernel cannot be used for a final submission, as it will score zero on the private dataset. In the `first50.csv` sample provided, there are only 50 predictions to keep the LB score low. Replace this with your own dataset and csv predictions.\n\n**PLEASE NOTE: This kernel will score ZERO on the final leaderboard. It is shared only as a time-saving utility to get quick LB feedback from remote prediction**"},{"metadata":{"_cell_guid":"","_uuid":"","trusted":true},"cell_type":"code","source":"import pandas as pd\npred_path = \"../input/nq-sample-csv/first50.csv\" #replace this with your own dataset.","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"This creates a dummy submission using the `examle_ids` from `simplified-nq-test.jsonl`. This will work for private and public test sets, so that commit and submit both run correctly."},{"metadata":{"trusted":true},"cell_type":"code","source":"df = pd.read_json('../input/tensorflow2-question-answering/simplified-nq-test.jsonl', lines = True, dtype={'example_id':'Object'})\nsubmission = pd.DataFrame(index=pd.concat([df.example_id + '_short', df.example_id + '_long']), columns=['PredictionString']).sort_index()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"This updates the submission from your csv, based on an intersection of `example_ids`. For the public LB, this will submit your csv. For the private test set, there will be no intersection, so all predictions are left blank and the score will be zero."},{"metadata":{"trusted":true,"collapsed":true},"cell_type":"code","source":"updates = pd.read_csv(pred_path, na_filter=False).set_index('example_id').sort_index()\nsubmission.loc[updates.index.intersection(submission.index),'PredictionString'] = updates['PredictionString']","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"submission.head(50)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"submission.to_csv('submission.csv')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":1}