{"cells":[{"metadata":{"toc":true,"_uuid":"5d64801e514fd0241dd66f947c0572c86d8e9832"},"cell_type":"markdown","source":"<h1>Table of Contents<span class=\"tocSkip\"></span></h1>\n<div class=\"toc\"><ul class=\"toc-item\"><li><span><a href=\"#The-PLAsTiCC-Astronomy-&quot;Starter-Kit&quot;-Classification-Demo\" data-toc-modified-id=\"The-PLAsTiCC-Astronomy-&quot;Starter-Kit&quot;-Classification-Demo-1\"><span class=\"toc-item-num\">1&nbsp;&nbsp;</span><strong>The PLAsTiCC Astronomy \"Starter Kit\" Classification Demo</strong></a></span><ul class=\"toc-item\"><li><ul class=\"toc-item\"><li><span><a href=\"#-Gautham-Narayan,-Renée-Hložek,-Emilie-Ishida-20180831\" data-toc-modified-id=\"-Gautham-Narayan,-Renée-Hložek,-Emilie-Ishida-20180831-1.0.1\"><span class=\"toc-item-num\">1.0.1&nbsp;&nbsp;</span>-Gautham Narayan, Renée Hložek, Emilie Ishida 20180831</a></span></li></ul></li></ul></li><li><span><a href=\"#Resources\" data-toc-modified-id=\"Resources-2\"><span class=\"toc-item-num\">2&nbsp;&nbsp;</span>Resources</a></span><ul class=\"toc-item\"><li><ul class=\"toc-item\"><li><ul class=\"toc-item\"><li><span><a href=\"#Useful-Packages-for-Astrophysics---particularly-feature-extraction\" data-toc-modified-id=\"Useful-Packages-for-Astrophysics---particularly-feature-extraction-2.0.0.1\"><span class=\"toc-item-num\">2.0.0.1&nbsp;&nbsp;</span>Useful Packages for Astrophysics - particularly feature extraction</a></span></li><li><span><a href=\"#External-Data-Sources\" data-toc-modified-id=\"External-Data-Sources-2.0.0.2\"><span class=\"toc-item-num\">2.0.0.2&nbsp;&nbsp;</span>External Data Sources</a></span></li></ul></li></ul></li></ul></li></ul></div>"},{"metadata":{"_uuid":"948d9e4ae0db32d7be52627089a14cd7cef9bba4"},"cell_type":"markdown","source":"# **The PLAsTiCC Astronomy \"Starter Kit\" Classification Demo**\n\n### -Gautham Narayan, Renée Hložek, Emilie Ishida 20180831\n\nThis kernel was developed for PLAsTiCC on Kaggle. The original version of this kernel is available as a Jupyter Notebook on <a href=\"https://github.com/LSSTDESC/plasticc-kit\">LSST DESC GitHub</a>. \n\n***"},{"metadata":{"trusted":true,"_uuid":"ed8e3bf0f7105eea63496e5a92879bb9e45335c9"},"cell_type":"code","source":"# You can edit the font size here to make rendered text more comfortable to read\n# It was built on a 13\" retina screen with 18px\nfrom IPython.core.display import display, HTML\ndisplay(HTML(\"<style>.rendered_html { font-size: 18px; }</style>\"))","execution_count":null,"outputs":[]},{"metadata":{"slideshow":{"slide_type":"slide"},"_uuid":"9c07291cbf42b1b34ddbe655a9c4288d6c53dbcf"},"cell_type":"markdown","source":"In this Kernel, we'll look at how astronomers have approached light curve classification, and provide you some references and useful packages if you want to get a sense for what has been done in the past.\n\nIt's almost certain that some astronomers will approach PLAsTiCC with these kinds of methods, so if you want to win, you'll likely have to do something better!\n\nWe'll begin by importing some packages that might be useful. "},{"metadata":{"trusted":true,"_uuid":"3373cff9dcf88614c09ebce9af571b256ec84d8a"},"cell_type":"code","source":"%matplotlib inline\n\nimport os\nfrom collections import Counter, OrderedDict\nimport numpy as np\nfrom operator import itemgetter\nimport matplotlib.pyplot as plt\nfrom astropy.table import Table\nimport multiprocessing\nfrom cesium.time_series import TimeSeries\nimport cesium.featurize as featurize\nfrom tqdm import tnrange, tqdm_notebook\nimport sklearn \nfrom sklearn.model_selection import StratifiedShuffleSplit\nfrom sklearn.decomposition import PCA\nfrom sklearn.ensemble import RandomForestClassifier\nfrom sklearn.metrics import confusion_matrix\nimport seaborn as sns","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"82d28648e155449a91493f510b5c02caec1f9185"},"cell_type":"markdown","source":"We described how the the data is split into a light curve table and a metadata table in the astronomy starter kit. We've described the format of these files there, so if you need a refresher, take a look there again.\n\nWe described a mapping from string passband name to integer passband IDs. Here, we'll create an `OrderedDict` to invert that transformation. "},{"metadata":{"trusted":true,"_uuid":"ecc57b1b025584f785e0225c01c8e12603941e99"},"cell_type":"code","source":"pbmap = OrderedDict([(0,'u'), (1,'g'), (2,'r'), (3,'i'), (4, 'z'), (5, 'Y')])\n\n# it also helps to have passbands associated with a color\npbcols = OrderedDict([(0,'blueviolet'), (1,'green'), (2,'red'),\\\n                      (3,'orange'), (4, 'black'), (5, 'brown')])\n\npbnames = list(pbmap.values())","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"bec689906a4a8e548f08445304fdc15b8eb49338"},"cell_type":"markdown","source":"Next, we'll read the metadata table."},{"metadata":{"trusted":true,"_uuid":"725fe76763a92f5b179b5b4ba592678751f746c5"},"cell_type":"code","source":"datadir = '../input/plasticc-astronomy-starter-kit-media'\nmetafilename = '../input/PLAsTiCC-2018/training_set_metadata.csv'\nmetadata = Table.read(metafilename, format='csv')\nnobjects = len(metadata)\nmetadata","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"44ebe6e7a083bf67a381ed0e0c4ec780535248e7"},"cell_type":"markdown","source":"Since we can split between extragalactic and galactic sources on the basis of redshift, you may choose to build separate classifiers for the two sets. \n\nOr you might split the deep-drilling fields (DDF i.e. `ddf_bool` = 1) up from the wide-fast-deep fields (WFD i.e. `ddf_bool` = 0). \n\nYou could even try making classifiers for different redshift bins, if you believe some classes will not be present at some redshifts. \n\nOr you might use that some of the training _and test_ data includes spectroscopic and photometric redshift, and see if you can use that information to derive bias corrections and apply non-linear transformations to the data before processing it.\n\nBecause there are definitely biases... \n\n(Also, if you made this kind of redshift-redshift plot from the test data, you'd see how non-representative the training set was.)"},{"metadata":{"trusted":true,"_uuid":"a3ad48f188162aab1d10f103d21d914542f37263"},"cell_type":"code","source":"extragal = metadata['hostgal_specz'] != 0.\ng = sns.jointplot(metadata['hostgal_specz'][extragal],\\\n              metadata['hostgal_photoz'][extragal], kind='hex',\\\n                  xlim=(-0.01, 3.01), ylim=(-0.01,3.01), height=8)\n\noutliers = np.abs(metadata['hostgal_specz'] - metadata['hostgal_photoz']) > 0.1\nfig = g.fig\nfig.axes[0].scatter(metadata['hostgal_specz'][outliers],\\\n                    metadata['hostgal_photoz'][outliers], color='C3', alpha=0.05)\nfig.tight_layout()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"52ef6062a12acbec51451b2a42a0e9aadca03761"},"cell_type":"markdown","source":"There's many potential choices. \n\nFor now we'll leave everything together."},{"metadata":{"trusted":true,"_uuid":"10cf5dad681a83162c9ad24c53cb92376f3af60b"},"cell_type":"code","source":"counts = Counter(metadata['target'])\nlabels, values = zip(*sorted(counts.items(), key=itemgetter(1)))\nfig, ax = plt.subplots()\n\ncmap = plt.cm.tab20\nnlines = len(labels)\nclasscolor =  list(cmap(np.linspace(0,1,nlines)))[::-1]\n\n# we'll create a mapping between class and color\nclasscolmap = dict(zip(labels, classcolor))\n\nindexes = np.arange(nlines)\nwidth = 1\nax.bar(indexes, values, width, edgecolor='k',\\\n       linewidth=1.5, tick_label=labels, log=True, color=classcolor)\nax.set_xlabel('Target')\nax.set_ylabel('Number of Objects')\nfig.tight_layout()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8308bb5a8ff255c9c62b241e4eebbf4cdc8b7b24"},"cell_type":"markdown","source":"You can see the class distribution in the training set is imbalanced. This reflects reality. The Universe doesn't produce all kinds of events at equal rates, and even if it did, some events are fainter than others, so we'd naturally find fewer of them than bright events.\n\nNext, we'll read the light curve data. All the objects in the training set are in a single file:"},{"metadata":{"trusted":true,"_uuid":"fb917b102cbf2b3a3a4277f72efd15d00f0c2774"},"cell_type":"code","source":"lcfilename = '../input/PLAsTiCC-2018/training_set.csv'\nlcdata = Table.read(lcfilename, format='csv')\nlcdata","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5dad142886ec11c29fe0c8bb27a5aefb6eff2c57"},"cell_type":"markdown","source":"Next, we'll make a `Timeseries` object using the `cesium` python package for each lightcurve."},{"metadata":{"trusted":true,"_uuid":"30de80cfd41bf159a704a6e51083c51e9451f635"},"cell_type":"code","source":"tsdict = OrderedDict()\nfor i in tnrange(nobjects, desc='Building Timeseries'):\n    row = metadata[i]\n    thisid = row['object_id']\n    target = row['target']\n    \n    meta = {'z':row['hostgal_photoz'],\\\n            'zerr':row['hostgal_photoz_err'],\\\n            'mwebv':row['mwebv']}\n    \n    ind = (lcdata['object_id'] == thisid)\n    thislc = lcdata[ind]\n\n    pbind = [(thislc['passband'] == pb) for pb in pbmap]\n    t = [thislc['mjd'][mask].data for mask in pbind ]\n    m = [thislc['flux'][mask].data for mask in pbind ]\n    e = [thislc['flux_err'][mask].data for mask in pbind ]\n\n    tsdict[thisid] = TimeSeries(t=t, m=m, e=e,\\\n                        label=target, name=thisid, meta_features=meta,\\\n                        channel_names=pbnames )\n    \ndel lcdata","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7680458549eaf1531d844243eedd34200098ca5d"},"cell_type":"markdown","source":"The list of features available with packages like `cesium` is <a href=\"http://cesium-ml.org/docs/feature_table.html\">huge</a>. We'll only compute a subset now."},{"metadata":{"trusted":true,"_uuid":"8e477543cacdec1789b8371681429d4bf071285e"},"cell_type":"code","source":"features_to_use = [\"amplitude\",\n                   \"percent_beyond_1_std\",\n                   \"maximum\",\n                   \"max_slope\",\n                   \"median\",\n                   \"median_absolute_deviation\",\n                   \"percent_close_to_median\",\n                   \"minimum\",\n                   \"skew\",\n                   \"std\",\n                   \"weighted_average\"]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"1ee6399f5feca7da10dd3e4eae4c6cc6d20c7149"},"cell_type":"code","source":"# we'll turn off warnings for a bit, because numpy can be whiny. \nimport warnings\nwarnings.simplefilter('ignore')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"1033eac9f933b2aa07a7acb68838dcaefdb1fb53"},"cell_type":"markdown","source":"We'll start computing the features. This takes a while, so it's good to do this only the once, and save state. \n\nIn principle you can avoid loading the data even if you run this again, but it's always a good idea to have access to your raw data while doing exploratory data analysis.\n\nIf you have access to a multiprocessing capabilities, it is usually a good idea to divide the feature computation task up. Particularly when you are dealing with the test set with about 3.5 million objects... \n\nWe'll use whatever cores you have on your machine."},{"metadata":{"trusted":true,"_uuid":"fe7f5d5efa17ce44a1293cca6a2e3159b3224c6e"},"cell_type":"code","source":"def worker(tsobj):\n    global features_to_use\n    thisfeats = featurize.featurize_single_ts(tsobj,\\\n    features_to_use=features_to_use,\n    raise_exceptions=False)\n    return thisfeats","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"39f1c91ec5fdc485694b24ec7e4c283928d33ff3"},"cell_type":"code","source":"featurefile = f'{datadir}/plasticc_featuretable.npz'\nif os.path.exists(featurefile):\n    featuretable, _ = featurize.load_featureset(featurefile)\nelse:\n    features_list = []\n    with tqdm_notebook(total=nobjects, desc=\"Computing Features\") as pbar:\n        with multiprocessing.Pool() as pool:  \n            results = pool.imap(worker, list(tsdict.values()))\n            for res in results:\n                features_list.append(res)\n                pbar.update()\n            \n    featuretable = featurize.assemble_featureset(features_list=features_list,\\\n                              time_series=tsdict.values())\n    featurize.save_featureset(fset=featuretable, path=featurefile)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"50385555ae9603370cccd207f53a750c2dc57831"},"cell_type":"markdown","source":"The computed feature table has descriptive statistics for each object - a low-dimensional encoding of the information in the light curves. `cesium` assembles it as a `MultiIndex` which does not make `sklearn` happy, so we'll make a simpler table out of it."},{"metadata":{"trusted":true,"_uuid":"c5ab43e0675ccbc953737a69797003508cbe2047"},"cell_type":"code","source":"old_names = featuretable.columns.values\nnew_names = ['{}_{}'.format(x, pbmap.get(y,'meta')) for x,y in old_names]\ncols = [featuretable[col] for col in old_names]\nallfeats = Table(cols, names=new_names)\ndel featuretable","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e283559574df23a14df6698107b7bb0c6426ba34"},"cell_type":"markdown","source":"We'll split the training set into two - one for training in this demo and the other for testing.\n\nWe'll do this preserving class balance in the full training set, but remember the training set might not be representative. The class balance in the training set reflects whatever we've followed up spectroscopically prior to LSST science operations starting."},{"metadata":{"trusted":true,"_uuid":"b8479004a43027ac4f59227ca0cf3f72bf0291dc"},"cell_type":"code","source":"splitter = StratifiedShuffleSplit(n_splits=1, test_size=0.3, random_state=42)\nsplits = list(splitter.split(allfeats, metadata['target']))[0]\ntrain_ind, test_ind = splits","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8d100082e4368f34450c356f4e4dc7971544267d"},"cell_type":"markdown","source":"Since there's multiple passbands in the data, and the passbands are all sampling the same underlying astrophysical event, it's reasonable to expect that the features we've computed to be correlated. \n\nTo prevent overfitting, it's a probably a good idea to reduce the dimensionality of the dataset. We'll look at this correlation structure."},{"metadata":{"trusted":true,"_uuid":"4e0f6b19c00a65698d07d46dd56f1eea3d9278ec"},"cell_type":"code","source":"corr = allfeats.to_pandas().corr()\n\n# Generate a mask for the upper triangle\nmask = np.zeros_like(corr, dtype=np.bool)\nmask[np.triu_indices_from(mask)] = True\n\n# Set up the matplotlib figure\nfig, ax = plt.subplots(figsize=(10, 8))\n\n# Draw the heatmap with the mask and correct aspect ratio\ncorr_plot = sns.heatmap(corr, mask=mask, cmap='RdBu', center=0,\n                square=True, linewidths=.2, cbar_kws={\"shrink\": .5})","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9b19b601db5d2ecfdf912241e028e4ed3cb4fadd"},"cell_type":"markdown","source":"You can see some of the bands are strongly correlated with each other, so lets try to reduce the dimensionality of the dataset looking for a two effective passbands."},{"metadata":{"trusted":true,"_uuid":"1c8c985243c3b5551590b9cba4f70100c4b7d342"},"cell_type":"code","source":"Xtrain = np.array(allfeats[train_ind].as_array().tolist())\nYtrain = np.array(metadata['target'][train_ind].tolist())\n\nXtest  = np.array(allfeats[test_ind].as_array().tolist())\nYtest  = np.array(metadata['target'][test_ind].tolist())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"53977495abd3e83de01d8326a6f86f8b6f59e016"},"cell_type":"code","source":"ncols = len(new_names)\nnpca  = (ncols  - 3)//len(pbnames)  + 3","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b47e964a62c02e98f1581b952eef90027135e413"},"cell_type":"code","source":"pca = PCA(n_components=npca, whiten=True, svd_solver=\"full\", random_state=42)\nXtrain_pca = pca.fit_transform(Xtrain)\nXtest_pca = pca.transform(Xtest)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"0867df59177e308f965fe33e81e23891393990f6"},"cell_type":"code","source":"fig, ax = plt.subplots()\nax.plot(np.arange(npca), pca.explained_variance_ratio_, color='C0')\nax2 = ax.twinx()\nax2.plot(np.arange(npca), np.cumsum(pca.explained_variance_ratio_), color='C1')\nax.set_yscale('log')\nax.set_xlabel('PCA Component')\nax.set_ylabel('Explained Variance Ratio')\nax2.set_ylabel('Cumulative Explained Ratio')\nfig.tight_layout()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"78ce061ca7e799d56060a08fa94d99d854c3eecb"},"cell_type":"markdown","source":"Looks like a few features seem to hold all the weight! You might think to yourself that this is trivial! Lets build a simple classifier - we'll use a random forest here, but you can switch this out for whatever you like."},{"metadata":{"trusted":true,"_uuid":"ca92a01be3c842edb897dda330f8158a4edf5d6c"},"cell_type":"code","source":"clf = RandomForestClassifier(n_estimators=200, criterion='gini',\\\n                       oob_score=True, n_jobs=-1, random_state=42,\\\n                      verbose=1, class_weight='balanced', max_features='sqrt')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"872cf2de7eef06d9e0185dc7dffa8db8029ce645"},"cell_type":"markdown","source":"Now let's apply the classifier to the data:"},{"metadata":{"trusted":true,"_uuid":"a7a1bbe1e51916ce940531216e32232a34501e8a"},"cell_type":"code","source":"clf.fit(Xtrain_pca, Ytrain)\nYpred = clf.predict(Xtest_pca)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"64dbde739db476bd680a1a616a849593b8d29e7d"},"cell_type":"markdown","source":"And see just how well it worked by constructing a confusion matrix."},{"metadata":{"trusted":true,"_uuid":"40d46b4c55908949eb8a8fd3e0862f21816d276e"},"cell_type":"code","source":"cm = confusion_matrix(Ytest, Ypred, labels=labels)\ncm = cm.astype('float') / cm.sum(axis=1)[:, np.newaxis]\nannot = np.around(cm, 2)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"957fe37d56dd06134d1b1357d3f12dc57d3e1071"},"cell_type":"markdown","source":"Which we can plot up..."},{"metadata":{"trusted":true,"_uuid":"ec748a6dbc857ee84fdca9991de9b6b4a7bfe512"},"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(9,7))\nsns.heatmap(cm, xticklabels=labels, yticklabels=labels, cmap='Blues', annot=annot, lw=0.5)\nax.set_xlabel('Predicted Label')\nax.set_ylabel('True Label')\nax.set_aspect('equal')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"849e2e8914e3f39e1224d512842e0abfdd663056"},"cell_type":"markdown","source":"Ooer... Some classes seem to be easy to pick out. Others are being confused very badly. This reflects what happens in reality to astronomers. Some classes are easier to distinguish than others. \n\nYou can also see the performance (or lack thereof) on the test data. \n\nNow this classifier is not the best we can do (by far), and it should be trivial for you to beat, but the fact is we can definitely use your expertise on this problem. There's many methods you can use to tackle this problem, and how you approach different aspects is going to be interesting to us regardless of how you do on the Kaggle leader board.\n\nFrom data augmentation to feature engineering to re-balancing classes to outlier identification and hierarchical multi-class probabilistic classification, there's much we can learn from this.\n\nWe'll leave you with a few resources and references that might be useful to tackle this challenge:"},{"metadata":{"_uuid":"46f4ad9c62fdaae546103d3460de79fc5b34c307"},"cell_type":"markdown","source":"# Resources\n\n|Reference   | Light Curve fit  | Dimensionality <br> Reduction  | Classification <br> algorithm  | Use redshift  | Source Code | \n|---|---|---|---|---|---|\n|[Poznanski *et al*, 2006](https://arxiv.org/pdf/astro-ph/0610129.pdf)| --- | --- | Template fit | Yes | [pSNiD II](https://www.sas.upenn.edu/~gladney/html-physics/psnid/psnidII/)|\n|[Newling *et al*, 2010](https://arxiv.org/pdf/1010.1005.pdf)   | parametric | parameters from fit  | Kernel Density Estimation <br> Boosting  | Yes  | No |\n|[Richards *et al*, 2011](https://arxiv.org/pdf/1103.6034.pdf)   | spline  | diffusion maps  | Random Forest  | Yes  | No |\n|[Karpenka *et al*, 2012](https://arxiv.org/pdf/1208.1264.pdf)   |parametric | parameters from fit  | Neural Network  | No  | No |\n|[Ishida & de Souza, 2013](https://arxiv.org/pdf/1201.6676.pdf)| spline | kernel PCA | Nearest Neighbor | No |[github](https://github.com/emilleishida/snclass) |\n|[Mislis *et al*, 2015](https://arxiv.org/abs/1511.03456) | --- | descriptive statistics | Random Forest | No | No | \n|[Varughese *et al*, 2015](https://arxiv.org/pdf/1504.00015.pdf)| spline | Wavelets | Nearest Neighbor <br> Support Vector Machine | No| No |\n|[Hernitschek *et al*, 2016](http://iopscience.iop.org/article/10.3847/0004-637X/817/1/73/meta)| $\\chi^{2}$ | --- | Random Forest | No | No | \n|[Lochner *et al*, 2016](https://arxiv.org/pdf/1603.00882.pdf)|parametric <br> Gaussian Process | Wavelets<br> PCA <br> Model Fit | Naive Bayes<br> Nearest Neighbor <br> Support Vector Machine <br> Boosted Decision Trees |No| No |\n|[Moller *et al*, 2016](https://arxiv.org/pdf/1608.05423.pdf) | parametric | parameters from fit| Boosted Decision Trees  <br> Random Forest | Yes | No |\n|[Charnok and Moss, 2017](https://arxiv.org/pdf/1606.07442.pdf)| --- | --- | Recurrent Neural Network | No|[github](https://github.com/adammoss/supernovae) |\n|[Mahabal *et al*, 2017](https://arxiv.org/abs/1709.06257)| rate of change | ---| Neural Network | No | No |\n|[Narayan *et al*, 2018](https://arxiv.org/abs/1801.07323) | parametric <br> Gaussian Process <br> | Wavelets<br> PCA <br> | Random Forest | No | No |\n|[Revsbech *et al*, 2018](https://arxiv.org/pdf/1706.03811.pdf)|Gaussian Process|Diffusion Maps| Random Forest |Yes| [github](https://github.com/rtrotta/STACCATO)|\n| [Dai *et al*, 2018](https://arxiv.org/pdf/1701.05689.pdf)|parametric|parameters from fit| Random Forest| No| No|\n\nWe're not suggesting reading all of these papers in detail and the references therein, but they will give you an overview of the sort of techniques astronomers have tried in the past, and will likely employ in this challenge.\n\n\n#### Useful Packages for Astrophysics - particularly feature extraction\n\n- [astropy](http://www.astropy.org/)\n- [astroML](http://www.astroml.org/)\n- [cesium](https://github.com/cesium-ml/cesium)\n- [celerite](https://celerite.readthedocs.io/en/stable/)\n- [FATS](https://github.com/isadoranun/FATS)\n- [feets](https://github.com/carpyncho/feets)\n- [gatspy](https://www.astroml.org/gatspy/)\n- [sncosmo](http://sncosmo.github.io/)\n- [tsfresh](https://tsfresh.readthedocs.io/en/latest/)\n- [vartools](https://www.astro.princeton.edu/~jhartman/vartools.html)\n\nAll of these packages were developed for astrophysics and in particular light curve processing. Celerite is a general 1-D Gaussian Process package but the particular class of kernels it offers is suitable for astrophysics and allows it to be much faster than many other packages.\n\n#### External Data Sources\n\nIf you are determined to try and augment the training set, or to learn more about the astrophysical sources in the data before building your data, then these resources might be useful to you. Note that you build in what biases are present in the training set into your classifier.\n\n- [AAVSO](https://www.aavso.org/research-portal)\n- [Open Astronomy Catalogs](https://astrocats.space/)\n- [OGLE Collection of Variable Stars](http://ogledb.astrouw.edu.pl/~ogle/OCVS/)\n- [SDSS DR12 Variables and Transients](https://www.sdss.org/dr12/algorithms/ancillary/boss/transient82/)\n\nRemember that LSST may discover objects that are predicted but have never been seen yet. We can't give you examples of things that haven't been seen of course.\n\n***"},{"metadata":{"_uuid":"76c5df069174226e5b4125fc73e3867610802a38"},"cell_type":"markdown","source":"Again, find us on the Kaggle forums if you have questions we can answer.\n\nGood luck!"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"toc":{"base_numbering":1,"nav_menu":{},"number_sections":true,"sideBar":true,"skip_h1_title":false,"title_cell":"Table of Contents","title_sidebar":"Contents","toc_cell":true,"toc_position":{},"toc_section_display":true,"toc_window_display":true}},"nbformat":4,"nbformat_minor":1}