{"cells":[{"metadata":{},"cell_type":"markdown","source":"Update: June.18.2020  \n\nAs in [this discussion](https://www.kaggle.com/c/birdsong-recognition/discussion/159123#890189) or [host comments](https://www.kaggle.com/c/birdsong-recognition/discussion/159123#890675), `background` column or `secondary_labels` column in train.csv have multi-label information.  \nI think both columns are almost same, after processing like below I did here. But using `secondary_labels` will be preferable based on host's comment. (Here I stick to using `background`)\n\n[host](https://www.kaggle.com/stefankahl) comments:\n> Overlapping vocalizations are a major issue and Xeno-canto recordings may or may not contain background species and they may or may not have an appropriate label (typically primary and secondary labels in the metadata). \n\nI'm not sure using secondary labels makes our score better or not, but for curiosity I made an dataframe for multi-label task. Please let me know, if I'm wrong.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"import warnings, re\n\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\n\npd.set_option('display.max_columns', 100)\npd.set_option('display.max_rows', 100)\npd.options.mode.chained_assignment = None\n# dir(pd.options.display)\nwarnings.simplefilter(action='ignore', category=FutureWarning)\n\nplt.style.use('ggplot')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"train = pd.read_csv('../input/birdsong-recognition/train.csv')\nprint(train.shape)\ntrain.head(3)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"train['background'].value_counts(dropna=True, sort=True).to_frame().head(20).plot.bar(\n    color='deeppink', figsize=(15, 5)\n);","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Make a neet function to extract multi label information from train.background column.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"def get_multi_label(s, bird_l):\n    if type(s) != str: s = str(s)\n    return [b in re.sub(r' \\([^()]*\\)', '', s).split('; ') for b in bird_l]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"bird_l = train.species.unique().tolist()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# default label\nlabel_df = pd.get_dummies(train.species).set_index(train.xc_id)\nlabel_df.head(3)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"%%time\nbackground_arr = np.array([get_multi_label(r, bird_l) for r in train.background])\nbackground_arr.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"label_df.iloc[:, :] = label_df.values + background_arr\nlabel_df['label_n'] = label_df.sum(axis=1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"label_df.label_n.value_counts().plot.bar(figsize=(10, 5), color='deeppink');","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"label_df.drop('label_n', axis=1).to_csv('multi-label.csv')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"As a result, almost of all instances are still mono-labeled. But one third of all train data is multi-labeled.\n\nOne strategy may be, first we train our models with mono-labeled data, then fine-tune with multi-labeled data.\nI'm not sure this may be good or not. But hope this information will help you.  \n\nHappy Kaggling!!!\n\n<img src=\"https://storage.googleapis.com/kaggle-avatars/images/2080166-kg.png\" width=100 align='left'>","execution_count":null}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}