{"cells":[{"metadata":{"_uuid":"410bd90db0dfa35c046d26f6c97dd7c08b27cd8b"},"cell_type":"markdown","source":"### This fork extends the explanations and shows the magic behind the initial code.\n\nForked from: https://www.kaggle.com/kyleboone/naive-benchmark-galactic-vs-extragalactic\n\nInitial kernel author: Kyle Boone (https://www.kaggle.com/kyleboone)\n\nCurrent kernel author: Ilya Khristoforov (https://www.kaggle.com/darbin)\n\n# Galactic vs Extragalactic Objects\n\nThe astronomical transients that appear in this challenge can be separated into two distinct groups: ones that are in our Milky Way galaxy (galactic) and ones that are outside of our galaxy (extragalactic). As described in the data note, all of the galactic objects have been assigned a host galaxy photometric redshift of 0. We can use this information to immediately classify every object as either galactic or extragalactic and remove a lot of potential options from the classification. Doing so results in matching the naive benchmark.\n\nWe find that all of the classes are either uniquely galactic or extragalactic except for class 99 which represents the unknown objects that aren't in the training set."},{"metadata":{"_uuid":"7d93b0f69b90cfe528f974b5f5009de8b7c49719"},"cell_type":"markdown","source":"## Load the data\n\nFor this notebook, we'll only need the metadata (**training_set_metadata.csv** and **test_set_metadata.csv**)."},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"07771416a764cb921cf9fca1ae5d38053714a6f3"},"cell_type":"code","source":"# matplotlib setting: create static png graphs instead of interactive ones\n%matplotlib inline","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"meta_data = pd.read_csv('../input/training_set_metadata.csv')\ntest_meta_data = pd.read_csv('../input/test_set_metadata.csv')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d49d04cc170cba612a9e096dabdf9acc8cc6cc3b"},"cell_type":"markdown","source":"Map the classes to the range 0-14. We manually add in the 99 class that doesn't show up in the training data."},{"metadata":{"trusted":true,"_uuid":"892d3243b91748486be716e4bf867ea639b65546"},"cell_type":"code","source":"# create list of all unique classes and add class 99 (undefiend class) to the list\nclasses = np.unique(meta_data['target'])\nclasses_all = np.hstack([classes, [99]])\nclasses_all","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"scrolled":true,"_uuid":"bdcf1301b1c318fd40a789cd6e3671c4b6b5dcfd"},"cell_type":"code","source":"# create a dictionary {class : index} to map class number with the index \n# (index will be used for submission columns like 0, 1, 2 ... 14)\ntarget_map = {j:i for i, j in enumerate(classes_all)}\ntarget_map","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"scrolled":true,"_uuid":"d2859cea2d4e1d1389ca9bcbfae208ce48b0d35a"},"cell_type":"code","source":"# check the type of the target_map\ntype(target_map)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"0370a63f801e40b6abc0b81e24bcea00ee431a7a","scrolled":false},"cell_type":"code","source":"# create 'target_id' column to map with 'target' classes\n# target_id is the index defined in previous step: see dictionary target_map\n# this column will be used later as index for the columns in the final submission\ntarget_ids = [target_map[i] for i in meta_data['target']]\nmeta_data['target_id'] = target_ids\nmeta_data.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"348a2c78d850ff1d5021f483876ae2422c8e0d3c"},"cell_type":"markdown","source":"Let's look at which classes show up in galactic vs extragalactic hosts. We can use the hostgal_specz key which is 0 for galactic objects."},{"metadata":{"trusted":true,"_uuid":"a112b08530ad43013c08783198981461398d0e2b"},"cell_type":"code","source":"# hostgal_specz == 0 means that the object is situated in our Milky Way Galaxy, these objects are labeled as 'Galactic'\n# other objects exist outside of our Galaxy and thus are labeld as 'Extragalactic'\n# hostgal_specz == 0 is the mask which will be used to separate Galactic objects from Extragalactic\n# hostgal_specz - is 'host galaxy spectral redshift'. The bigger the redshift - the farrer the object from us.\ngalactic_cut = meta_data['hostgal_specz'] == 0\n\n# create figure object to hold an image 10x8 (width x height)\nplt.figure(figsize=(10, 8))\n\n# create 2 histograms on one image showing the number of objects in each class\n# label - is a legend label\n# alpha - is a transparency (50%) to see the overlapping between the 2 histograms\n# 15 - is the number of bins for hist\n# (0, 15) - is the range\n\n# histogram 1 for hostgal_specz == 0 (Galactic objects)\nplt.hist(meta_data[galactic_cut]['target_id'], 15, (0, 15), alpha=0.5, label='Galactic')\n# histogram 2 for hostgal_specz <> 0 (Extragalactic objects)\nplt.hist(meta_data[~galactic_cut]['target_id'], 15, (0, 15), alpha=0.5, label='Extragalactic')\n\n# labels for x axis\nplt.xticks(np.arange(15)+0.5, classes_all)\n\n# show y axis in a log format to narrow the distances between the counts for each bar: \n# this simplifies the presentation and helps to determine the overlapings between the two hists\nplt.gca().set_yscale(\"log\")\n\n# print axes labels\nplt.xlabel('Class')\nplt.ylabel('Counts')\n\n# range for x axis (0 - 15)\nplt.xlim(0, 15)\n\n# print a legend\nplt.legend();\n\n# as we can see classes 6, 16, 53,65 and 92 are all Galactic\n# and classes 15, 42, 52, 62,64, 67, 88, 90 and 95 are all Extragalactic\n# there is no overlapping between the classes","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3714915a09d1aab61d0bc1a15a7ecd654a09886e"},"cell_type":"markdown","source":"There is no overlap at all between the galactic and extragalactic objects in the training set. Class 99 isn't represented in the training set at all. Let's make a classifier that checks if an object is galactic or extragalactic and then assigns a flat probability to each class in that group. We'll include class 99 in both the galactic and extragalactic groups."},{"metadata":{"trusted":true,"_uuid":"c8eeaf0e1cecbcf11fcbcbdaa7c8b6a5c7e9fb3e"},"cell_type":"code","source":"# Build the flat probability arrays for both the galactic and extragalactic groups\n\n# Extract galactic and Extragalactic classes\ngalactic_cut = meta_data['hostgal_specz'] == 0\ngalactic_data = meta_data[galactic_cut]\nextragalactic_data = meta_data[~galactic_cut]\n\ngalactic_classes = np.unique(galactic_data['target_id'])\nextragalactic_classes = np.unique(extragalactic_data['target_id'])\n\nprint('Galactic classes:', galactic_classes)\nprint('Extragalactic classes:', extragalactic_classes)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"scrolled":true,"_uuid":"4c0318771db7e783b23ff2bfcab4727dd2fdf2fa"},"cell_type":"code","source":"# Add class 99 (id=14) to both groups (Galactic and Extragalactic classes).\ngalactic_classes = np.append(galactic_classes, 14)\nextragalactic_classes = np.append(extragalactic_classes, 14)\n\nprint('Galactic classes:', galactic_classes)\nprint('Extragalactic classes:', extragalactic_classes)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"scrolled":true,"_uuid":"6c7aec24e7a5940a01fcd2511bd8f90a90c60563"},"cell_type":"code","source":"# create a 15 zeros array 'galactic_probabilities'\ngalactic_probabilities = np.zeros(15)\nprint('Zeros for Galactic probabilities:', galactic_probabilities)\n\n# suppose that probability for the Galactic object to have certain Galactic class is evenly distributed \n# create an array of probabilities for the galactic object to belong to a certain Galactic class\ngalactic_probabilities[galactic_classes] = 1 / len(galactic_classes)\nprint('Galactic flat probabilities: ',galactic_probabilities)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"eff5e8fe2122841c00fa001df70d305d4d4d2bec","scrolled":true},"cell_type":"code","source":"# create a 15 zeros array 'extragalactic_probabilities'\nextragalactic_probabilities = np.zeros(15)\nprint('Zeros for Extragalactic probabilities:', extragalactic_probabilities)\n\n# suppose that probability for the Extragalactic object to have certain Extragalactic class is evenly distributed \n# create an array of probabilities for the galactic object to belong to a certain Extragalactic class\nextragalactic_probabilities[extragalactic_classes] = 1 / len(extragalactic_classes)\nprint('Extragalactic flat probabilities: ', extragalactic_probabilities)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d5b1920af33fbfd7ba008e6f159718f0f279760b"},"cell_type":"markdown","source":"Apply this prediction to the data. We simply choose which of the two probability arrays to use based off of whether the object is galactic or extragalactic."},{"metadata":{"trusted":true,"_uuid":"0f60b092518adac8bfdaa12d374525c84d428188"},"cell_type":"code","source":"# Apply this prediction to a table\n\n# import progress bar package\nimport tqdm\n\ndef do_prediction(table):\n    probs = []\n    for index, row in tqdm.tqdm(table.iterrows(), total=len(table)):\n        \n        # we use 'hostgal_photoz' (photometric redshift) here instead of 'hostgal_specz' (spectral redshift)\n        # it is the same redshift measure but made faster on a larger area and thus less accurate, but we have this measure for all of the objects \n        \n        # if object is in the Milky Way Galaxy\n        if row['hostgal_photoz'] == 0:\n            prob = galactic_probabilities\n            \n        # if object is out of the Milky Way Galaxy\n        else:\n            prob = extragalactic_probabilities\n        probs.append(prob)\n    \n    return np.array(probs)\n\npred = do_prediction(meta_data)\ntest_pred = do_prediction(test_meta_data)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"743383451a349de2ca262a584a4fdf082fda7298"},"cell_type":"markdown","source":"Now write the prediction out and submit it. This notebook gets a score of 2.158 which matches the naive benchmark."},{"metadata":{"trusted":true,"_uuid":"b932871549617570eb343776b938ede025464583"},"cell_type":"code","source":"test_df = pd.DataFrame(index=test_meta_data['object_id'], data=test_pred, columns=['class_%d' % i for i in classes_all])\ntest_df.to_csv('./naive_benchmark.csv')","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}