{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Species Audio Detection -Aim\nThe dataset consists of data for a lot of species in a forest. We are to predict the species present in a particular audio file. In this kernel, we will be analysing the data and trying to create an initial model.\n\nSources:\n* [Kernel by Bojan Tunguz](https://www.kaggle.com/tunguz/rainforest-rapids-baseline)"},{"metadata":{"trusted":true},"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport librosa","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"# input training csv file audio labels(True Positives)\ntruePositiveData = pd.read_csv(\"../input/rfcx-species-audio-detection/train_tp.csv\")\ntruePositiveData.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# EDA"},{"metadata":{"trusted":true},"cell_type":"code","source":"#Check for null values\ntruePositiveData.isnull().any()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Are some records repeated or not?\ntruePositiveData['recording_id'].nunique() == len(truePositiveData)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Possible causes for lengths not matching:\n* There are duplicates\n* There are multiple species for some audio. Each species for one audio has been put separately."},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"# Unique species id Count Plot\nplt.figure(figsize = (10, 5), dpi = 300)\nsns.countplot(truePositiveData['species_id'], palette='autumn')\nplt.title('species counts');","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Inference**\n\nThe Distribution of species if mostly uniform. However, species 23 has a high number of occurences, so we need to probe further."},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"# Song type analysis\nplt.figure(figsize = (10, 5), dpi = 300)\nsns.countplot(truePositiveData['songtype_id'], palette='gist_earth')\nplt.title('songtype_id Counts');","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**Inference**\n\nThere is a huge difference between the songtype id's. Class 1 has about 1000+ counts while class 4 has about 150 or so.\nThis might create a problem for models."},{"metadata":{"trusted":true},"cell_type":"code","source":"# Sense of the numerical data\ntruePositiveData.describe()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Distribution of f_min and f_max\nplt.figure(figsize = (10, 5), dpi = 300)\nplt.style.use('ggplot')\nsns.distplot(truePositiveData['f_min'], color='black')\nsns.distplot(truePositiveData['f_max'], color='red')\nplt.title('Min and Max frequencies')\nplt.legend(['Min_Frequency', 'Max_Frequency']);","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# log of f_min\nplt.figure(figsize = (10, 5), dpi = 300)\nplt.style.use('seaborn-paper')\nmax_of_min_freq = truePositiveData['f_min'].max()\nsns.distplot(np.exp(truePositiveData['f_min'] / max_of_min_freq), color='olive')\nplt.title('exp(truePositiveData[\"f_min\"] / max_of_min_freq)');","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Tried plotting some plots to see the frequency distributions."},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}