{"cells": [{"outputs": [], "metadata": {"_uuid": "871ad81ad74fcce73e8ab88c3f543a811535d2d0", "_cell_guid": "09ee6395-d4a8-4d8c-a1cd-77d577768e8d"}, "cell_type": "markdown", "source": "## Organizing Labels with pandas\n\nThis shows how the data labels for the Stage 1 training scans can be organized in a pandas dataframe.  By separating the scan IDs and body zones, we can easily get information such as how many threats are in each zone or how many threats are in the scans.\n\nFirst, let's import some dependencies.\n", "execution_count": null}, {"outputs": [], "metadata": {"_uuid": "44ff6f0adc63da62e4f637ba7898a7b2b65b4009", "_execution_state": "idle", "_cell_guid": "4980f94b-bb19-480b-bb6e-f94cecd7f1ba", "trusted": false}, "cell_type": "code", "source": "import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport os\n%matplotlib inline\nplt.style.use('ggplot')\npd.options.display.max_rows = 20", "execution_count": 1}, {"outputs": [], "metadata": {"_uuid": "9d4821e228ce1e73bf83c64bd983a89045115699", "_cell_guid": "5510fb35-1ab2-446c-bfae-27afab0561ef"}, "cell_type": "markdown", "source": "\nRead in the data and create a dataframe that has the body zones for column names and the scan IDs for indices.  Note that the IDs are in the csv file are in alphanumeric order, so the zones are in order of Zone1, Zone10, Zone11, etc.\n", "execution_count": null}, {"outputs": [], "metadata": {"_uuid": "33de701f28fa685bc4e2314b23f9ea6fff460c93", "_execution_state": "idle", "_cell_guid": "b973410f-f516-45da-8166-fea64ef6e5aa", "trusted": false}, "cell_type": "code", "source": "# Read in labels\nbase_dir = os.path.join('..', 'input')\nunsorted_df = pd.read_csv(os.path.join(base_dir, 'stage1_labels.csv'))\n\n# Get IDs for rows\ns = list(range(0,len(unsorted_df),17))\nobs = unsorted_df.loc[s,'Id'].str.split('_')\nscanID = [x[0] for x in obs]\n\n# Put zones in columns\ncolumns = sorted(['Zone'+str(i) for i in range(1,18)])\n\ndf = pd.DataFrame(index=scanID, columns=columns)\n\n# Sort labels by zone\nfor i in range(17):\n    s = list(range(i,len(unsorted_df),17))\n    df.iloc[:,i] = unsorted_df.iloc[s,1].values\n\nprint('Number of labeled scans:', len(df))\ndf.head()", "execution_count": 3}, {"outputs": [], "metadata": {"_uuid": "d4523da282030967ff5815461c971359d72aa9ac", "_cell_guid": "3f9b679a-e0f1-4f2b-a0b9-4b5fa155d728"}, "cell_type": "markdown", "source": "\nIt's now easier to look at what data is available to us for training.  For example, how many threats are in each zone:\n", "execution_count": null}, {"outputs": [], "metadata": {"_uuid": "95a87732c797f1ea57e21decec69c292fa441f23", "_execution_state": "idle", "_cell_guid": "f4d4d1ee-7258-4d90-8dc2-7e08f026a751", "trusted": false}, "cell_type": "code", "source": "nobj_zone = df.sum()\nprint(nobj_zone)\n\nnobj_zone.plot(kind='bar', width=.75, title='Threat Count in Each Zone')\nplt.ylabel('Number of Threats')\nplt.xlabel('Zone')\nplt.show()", "execution_count": null}, {"outputs": [], "metadata": {"_uuid": "1b5585e0f91fae1959314425027ed4dd448491d3", "_cell_guid": "b7971fce-5494-4a46-a1e8-1c6dc6179dde"}, "cell_type": "markdown", "source": "\nWe can also see how many threats are being used in the scans.  \n", "execution_count": null}, {"outputs": [], "metadata": {"_uuid": "027a8ea800c6d11f761e5e943a51bce809bae65d", "_execution_state": "idle", "_cell_guid": "4b353890-5735-43b2-b6f7-7ea8c54305c4", "trusted": false}, "cell_type": "code", "source": "nobj_scan = df.sum(1).value_counts().sort_values()\nprint(nobj_scan)\n\nnobj_scan.plot(kind='bar', width=.75, title='Frequency of Threat Counts')\nplt.ylabel('Number of Scans')\nplt.xlabel('Threat Count')\nplt.show()", "execution_count": null}, {"outputs": [], "metadata": {"_uuid": "ebcb1038f7501bb34f187933ee4dc337ce8d7c25", "_cell_guid": "934aa2d8-becb-4d37-9da2-da37a915d4c7", "collapsed": true}, "cell_type": "markdown", "source": "\nUsing this dataframe would also make it easier to get scans with threats in desired zones...\n", "execution_count": null}, {"outputs": [], "metadata": {"_uuid": "94fb8288254cbcf38350de4c2201ae7edf8b73d3", "_execution_state": "idle", "_cell_guid": "308748ac-41e3-497e-9905-e661bde7b648", "trusted": false}, "cell_type": "code", "source": "df[df['Zone1']==1]", "execution_count": null}, {"outputs": [], "metadata": {"_uuid": "83c262c8feaa7e5878c21739f4ae1e188251b238", "_cell_guid": "b3772e67-72cb-43fa-bf3a-4601b668fd32"}, "cell_type": "markdown", "source": "\n...or to search for any correlations...\n", "execution_count": null}, {"outputs": [], "metadata": {"_uuid": "0d683f68eb432f008ea4d458d7067a1f8674dcda", "_execution_state": "idle", "_cell_guid": "7a661ebe-d3e6-43ce-b02d-2cff51f2d233", "trusted": false}, "cell_type": "code", "source": "df.corr()", "execution_count": null}, {"outputs": [], "metadata": {"_uuid": "96207e46913b5beea661384d87b31fff5787ccb9", "_cell_guid": "f6ccc40f-87e8-4fc4-ace0-f14bb1a73b47"}, "cell_type": "markdown", "source": "\n\n... although correlations probably shouldn't be taken too seriously unless there's a reason to think it could be real, but it's just to demonstrate how organizing the labels in a dataframe like this is more useful.\n", "execution_count": null}], "nbformat_minor": 1, "metadata": {"anaconda-cloud": {}, "language_info": {"mimetype": "text/x-python", "pygments_lexer": "ipython3", "name": "python", "codemirror_mode": {"name": "ipython", "version": 3}, "nbconvert_exporter": "python", "file_extension": ".py", "version": "3.6.1"}, "kernelspec": {"name": "python3", "language": "python", "display_name": "Python 3"}}, "nbformat": 4}