{"cells":[{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"markdown","source":""},{"metadata":{"_uuid":"8521a8d26036a6e9b66942f0279491ef5c0c168f"},"cell_type":"markdown","source":"# data_note.pdfの日本語\n- google翻訳しただけ"},{"metadata":{"_uuid":"31900686252b2f515d7a4e37f0c87e7c14b7cdc3"},"cell_type":"markdown","source":"- 1.はじめに \nPLAsTiCCは大規模なデータ課題であり、参加者は天文学時系列データこれらのシミュレートされた時系列、すなわち「ライトカーブ」は、物体の明るさを時間の関数として6つの異なる光子束を通過帯域と呼ばれる天文学フィルタで測定している。   \nこれらの通過帯域には、光スペクトルの紫外、赤外領域を含む。  \n物理的プロセスで駆動されていて多くの異なる天体タイプや天文学のクラスがある。 \n課題はライトカーブの各セットを分析することを決定し、それぞれがオブジェクトはこれらのどのクラスに属しているか。  \n時系列データは、これからのLarge Synoptic Survey望遠鏡から提供されている。   \nLSSTはCerro Pachonと呼ばれる山にチリ北部の砂漠で建設中。  \n完了すると、LSSTは8メートルの望遠鏡を使用し、30億画素のカメラでおおよそ数夜おきに、そして10年間の期間にわたって、南の空全体をイメージします。  \nデータの猛攻撃に備えるために数百人の科学者が過渡的および可変的な星のワーキンググループを形成していて、このチャレンジの時系列データはライトカーブと呼ばれます。  \nこのチャレンジのために捨てられた移動オブジェクト（小惑星）は同じ位置に留まりますが、明るさは異なります。  \nオブジェクトの最初の状態に応じて爆発または明るくなった場合は時間と共に増加または減少することがある。  \n特定のフラックスが変化する様子（明るくなる時間の長さ、オブジェクトが異なるパスバンドで明るくなる、フェードする時間など）は、オブジェクトの基本型。  \n我々は、これらのライトカーブを使用して、LSSTから一連の訓練データ（ライトカーブ）を用いて分類され、ラベルが与えられる。  \nPLAsTiCCの参加者は、そのデータを15のクラスに分け、そのうち14個はトレーニングサンプルで表されます。  \nこれらのクラスを生成するために使用されるモデルについては、次のチャレンジが完了すると解放されます。  \n最後のクラスは面白い存在すると仮定されているが観察されたことがないオブジェクトではないトレーニングセット。   \n分類アルゴリズムがトレーニングセットで調整されると、参加者は以前に「test」ライトカーブデータが見えなかった同じアルゴリズムで分類することです。  \n図1では、トレーニングセットからの3つのライトカーブ例を示します。上の2つのパネル短時間で明るくフェードインするソースを表示します。下のパネルには、一時的にフェードすることができますが、常に再び明るくなる可変オブジェクトです。それは注目に値します。  \nすべての変数が明るくなったり消えたりするわけではなく、再び明るくてフェードしたりします。  \nLSSTには2つの異なる種類の調査地域があります：いわゆる深掘削畑は、大きな深度を達成するために頻繁にサンプリングされる空の小さなパッチです。  \nこれらのDDFパッチ内のオブジェクトは非常によく決定され、したがって小さいフラックスの誤差。WFD調査は、大部分の空をカバーしていて、それほど頻繁に観測されることはありません。  \nポイントはより大きな不確実性を持つ点です。 "},{"metadata":{"_uuid":"7429dc861b8bdc88011d561036524f33c4b21a17"},"cell_type":"markdown","source":"1.  Introduction  \nPLAsTiCC is a large data challenge for which participants are asked to classify astronomical\ntime series data. These simulated time series, or ‘light curves’, are measurements of\nan object’s brightness as a function of time - by measuring the photon flux in six different\nastronomical filters (commonly referred to as passbands). These passbands include\nultra-violet, optical and infrared regions of the light spectrum. There are many different\ntypes of astronomical objects (that are driven by different physical processes) that we separate\ninto astronomical classes. The challenge is to analyze each set of light curves (1\nlight curve per passband, 6 passbands per object) and determine a probability that each\nobject belongs to each of these classes. The time-series data provided are simulations of\nwhat we expect from the upcoming Large Synoptic Survey Telescope (LSST). LSST is\nunder construction high in the deserts of northern Chile on a mountain called Cerro Pachon.\nWhen complete, LSST will use an 8 meter telescope, with a 3 billion pixel camera\nto image the entire Southern sky roughly every few nights and over a ten-year duration.\nIn order to prepare for the data onslaught, hundreds of scientists are joining forces to\nform collaborations and working groups including the Transient and Variable Stars Collaboration\n(TVS) and the Dark Energy Science Collaboration (DESC). As part of the\ncollaborative effort, these two collaborations developed the PLAsTiCC team.\nThe time-series data in this challenge are called light curves. These light curves are the\nresult of difference imaging: two images are taken of the same region on different nights,\nand the images are subtracted from each other. This differencing procedure catches\nboth moving objects (asteroids) that are discarded for this challenge, and also objects\nthat stay at the same position but vary in brightness. Depending on when the object first\nexploded or brightened, the flux may be increasing or decreasing with time. The specific\nmanner in which the flux changes (the length of time over which it brightens, the way the\nobject brightens in different passbands, the time to fade, etc.) is a good indicator of the\nunderlying type of the object. We use these light curves to classify the variable sources\nfrom LSST. The data are classified using a set of training data (light curves) for which\nthe true classifications (i.e. labels) are given. The participants in PLAsTiCC are asked to\nseparate the data into 15 classes, 14 of which are represented in the training sample. The\nmodels used to generate these classes will be described in an upcoming paper that will be\nreleased once the challenge is complete. The final class is meant to capture interesting\nobjects that are hypothesized to exist but have never been observed and are thus not in\nthe training set. This class would encompass any objects not seen before (i.e. there may\nbe more than one type of object in this class).\nOnce their classification algorithms are tuned on the training set, the participants will apply\nthose same algorithms to previously unseen ‘test’ light curve data. The goal is to classify\nthose objects into classes.\nIn Figure 1, we show three example light curves from the training set. The top two panels\nshow sources that brighten and fade over a short time period. The lower panel shows a\nvariable object that can fade temporarily, but always brightens again. It is worth noting\nthat not all variables brighten and fade, some could just brighten and fade, never to rise\nagain. Also note the gaps between observations: small gaps (from minutes to days or\nweeks) from the time between telescope observing a given patch of sky, and large gaps\n(> 6 months) where the object is not visible at night from the LSST site.\nLSST will have two different kinds of survey regions: the so-called Deep Drilling Fields\n(DDF) are small patches of the sky that will be sampled often to achieve great depth\n(i.e. to be able to measure the flux from fainter objects). Objects in these DDF patches\nwill have light-curve points that are extremely well determined and therefore have small\nerrors in flux. The Wide-Fast-Deep (WFD) survey covers a larger part of the sky (almost\n400 times the area of the deep fields) that will be observed less frequently (and so lightcurve\npoints will have larger uncertainties), but will discover many more objects over the\nlarger area.\nFigure 1. Example light curves in the PLAsTiCC data set. The three example objects display\ndifferent changes in flux with time that are typical of real objects. The top row shows the same\ntype of object that brightens from some ‘nominal’ brightness and then fades away. The difference\nbetween the panels in the top row show the effects of observation gaps (linked to the telescope\nsurvey strategy): the top-right panel illustrates that the brightening of the flux can occur near\nobservation gaps, and therefore may not include the full time period of brightening (or fading) for\nthe object. The bottom figure illustrates an object that brightens and fades periodically rather than\nas a ‘once-off’ event. In addition, all three panels show that seasonal gaps over a roughly two-year\nperiod, and the cadence of observations as set by the LSST survey strategy can introduce gaps\nin the light curve.\n."},{"metadata":{"_uuid":"eacc71670c6d0ba2d404e27e6764c6dc1b1a215f"},"cell_type":"markdown","source":"- 2.天文学の背景  \n私たちは夜の空を静かなものとして遠くの星を考えていますが、数秒から数分の時間や月年スケールで明るさが変化する光源で空は満ちています。  \nこれらのイベントの一部はトランジェントと呼ばれ、天文学現象の多種多様となる。  \n例えば、星が爆発すると明るい「超新星」が生成され、時間と共に消え繰り返さない。  \n他のイベントは輝度が繰り返し変化するため周期的なやり方で、または活動銀河核を含むエピソード的に変化する物体（AGN）銀河の心臓で、セフェイドと呼ばれる脈動する星は互いの光を遮るため交互に見える。  \nこれらの明るい情報源におけるこれらの変化は、彼ら自身についての重要な手がかりを提供することができる。  \nそしてその環境 や全体としての宇宙の進化。  \n例えば、タイプ1aの超新星光カーブの測定は、加速された宇宙の拡大。  \n各タイプの一時変数と変数は、星がどのように進化するか、恒星の爆発の物理学、化学宇宙の豊かさ、そして宇宙の加速的な拡大が含まれます。  \nしたがって、観測源の適切な分類は、観測天文学における重要な課題です。  \n次世代の天文観測に期待される膨大なデータ量の光(LSSTを含む)。  \nこの課題で我々が取り組む問題は、天文学的な分類データを模倣するように設計されたシミュレートされたLSSTライトカーブデータセットからの変数から決定すること。  \n分類は大規模なテストセットで行われますが、トレーニングデータは完全なデータの小さなサブセットになりますが、私たちが観測に直面する課題を模倣するためのテストセットです。    \n   \n- 2.1. 天体観測のためのさまざまな方法  \nここでは、LSSTの詳細と手元の課題について説明します。  \n2つの重要なモード  \n天体からの光を特徴付けることを「分光法」および「測光」と呼ぶ。  \n分光法は、フラックスを波長の関数として測定し、現代の同等物であるプリズムを使用して光のビームを高精度に色の虹に分離すること。  \n特定の物質を示す排出および吸収の特徴を特定することを可能にする測定オブジェクトに存在する化学元素。  \n分光法もまた最も正確で天体の過渡現象や変数の分類を可能にする信頼性の高いツールです。   \nしかしながら分級作業のために最も重要なことであるが、分光法は非常に時間がかかるプロセス でオブジェクトを検出するのに必要な時間よりもはるかに長い露光時間で同等の大きさの望遠鏡上のフィルタを必要とする。  \n今後の大規模なスカイサーベイから期待されるデータ量を考えると、あらゆる物体の分光観測は実現不可能である。  \n代わりの方法は、さまざまなフィルタ（パスバンドとも呼ばれます）を通してオブジェクトの画像を撮影します。  \n各通過帯域は、特定の（広い）波長範囲内の光を選択する。  \nこのアプローチで得られたデータを測光データと呼ぶ。  \nLSSTには、u、g、r、i、z、yと表される6つの通過帯域があり、異なる波長範囲：  \nuバンドに対して300〜400ナノメートルの波長、 \ngパスバンドについて400〜600nm、rバンドについて500〜700nm、 \niバンドの場合は650〜850、zバンドの場合は800〜950nm、 \ny帯域については950nmおよび1050nmである。  \nフィルター効率対波長が示されている。  \n参考までに、人間の目はgバンドとrバンドの光に敏感です。  \nフラックス時間の関数として測定される各通過帯域内の光の光度は、光の曲線である。  \nこれらのライトカーブ上で実行される。  \n分光測定は素晴らしいことがありますが、通過帯域データは、1回の露光につき1バンドあたり1つの輝度測定値です。  \nチャレンジ低（スペクトル）解像度の光カーブデータでオブジェクトを分類することです。  \n比較した分光法では、通過帯域観測による光カーブ測定の利点私たちははるかに遠く、より暗い物体を観察することができ、その物体を同時に（測定するのではなく）より大きな視野の中の多くの物体を観察することができる。  \n観測は、月光、夕暮れ、雲、および大気の影響（私たちは「見る」と呼ぶ）。これらの劣化は、より大きな磁束不確実性を生じ、この情報はライトカーブデータに含まれます（下記第3章参照）。   \nチャレンジデータセットは、ライトカーブに加えて、他の2つの情報が各オブジェクトごとに用意されています。   \n訓練データには対象物の正確な赤方偏移が含まれるが、試験データの赤方偏移は銀河からの醜い通過帯域測定値に基づく近似測定で、オブジェクト（空の同じ位置にある銀河、オブジェクトと同じ赤方偏移）。   \n光源が明るくなったり消えたりする間に、ホスト銀河のフラックス変化しないので、含まれる変数オブジェクトが明るくなる前に測定することができます。   \nテストデータの赤方偏移の数パーセントが壊滅的であることに注意してください。   \n赤方偏移の不確実性は、測定値と真値との差を大きく過小評価する赤方偏移。   \n銀河が地球に到達したときの光に対する赤方偏移の影響を図2に示します。黒曲線は、0.01の赤方偏移における近くのタイプ1aの超新星スペクトルを示し、対応する140億光年の距離に。 「近くに」という言葉が奇妙に見えるかもしれませんが、この場合、この距離は実際には宇宙の全範囲と比較して近くにあります。   \n超新星スペクトルとフィルター効率の目視検査最大フラックス（フィルタ上で合計されたスペクトル）がグリーン（g）フィルタ内にあることを示します。   \n破線曲線はより遠い超新星からのスペクトルを示し、赤色シフトに対応する0.5、または51億光年離れています.   \n最大光束は赤色（r）フィルターに移動します。   \n赤方偏移および距離が増加するにつれて、最大フラックスはより赤色のフィルタに現れる。   \nしたがって「赤方偏移」という言葉。   \n2番目の追加情報は、知られている私たちの銀河からの絶滅に関連しています。   \n天の川として私たちの光カーブの測定値は大気と望遠鏡の透過率（望遠鏡が標準星から較正されていると仮定して、これに関する知識はPLAsTiCCに参加する必要はありません）。さらに、それぞれを修正する。   \n地球に向かう途中で天の川「塵」を通って進む光の吸収のための光の曲線。   \nこの吸収は、紫外線フィルタで最も強く、赤外線フィルタでは最も弱い。   \n各オブジェクトの空の座標には、天の川の消滅の値が含まれています。   \nデータリリースでは、ラベル「MWEBV」で、それはどのように天文学的な尺度です。   \n塵のない天の川に比べてはるかに赤い物体が現れます。より大きなMWEBV値はオブジェクトへの視線に沿ったより多くのミルキーウェイダストに対応し、オブジェクトは赤く表示されます。 PLAsTiCCデータ内のすべてのオブジェクトは、MWEBVを持つように選択されます。"},{"metadata":{"_uuid":"93a3edd20a2912c4982126ecfd431c01479ed4fe"},"cell_type":"markdown","source":"2. Astronomy background\nWhile we think of the night sky and the distant stars it contains as static, the sky is filled\nwith sources of light that vary in brightness on timescales from seconds and minutes to\nmonths and years.\nSome of these events are called transients, and are the observational consequences of a\nlarge variety of astronomical phenomena. For example, the cataclysmic event that occurs\nwhen a star explodes generates a bright ‘supernova’ signal that fades with time, but does\nnot repeat. Other events are called variables, since they vary repeatedly in brightness\neither in a periodic way, or episodically variable objects including active galactic nuclei\n(AGN) at the hearts of galaxies, pulsating stars known as Cepheids, and eclipsing binary\nstars that alternate blocking out each other’s light from view.\nThese variation in these bright sources can provide important clues about themselves\nand their environment - as well as the evolution of the universe as a whole. For example,\nmeasurements of type Ia supernovae light curves provided the first evidence of accelerated\nexpansion of the Universe. Each type of transient and variable provides a different\nclue that helps us study how stars evolve, the physics of stellar explosions, the chemical\nenrichment of the cosmos, and the accelerating expansion of the universe. Therefore, the\nproper classification of sources is a crucial task in observational astronomy - especially in\nlight of the large data volumes expected for the next generation of astronomical surveys -\nthat includes LSST.\nThe question we address in this challenge is: how well can we classify astronomical\ntransients and variables from a simulated light curve data set designed to mimic the data\nfrom LSST? Crucially, the classifications will occur on a large test set, but the training\ndata will be a small subset of the full data, and will also be a poor representation of the\ntest set, to mimic the challenges we face observationally.\n2.1. Different methods for observing astronomical objects\nHere we give more detail on LSST, and the challenge at hand. Two important modes for\ncharacterizing light from astronomical objects are called ‘spectroscopy’ and ‘photometry.’\nSpectroscopy measures the flux as a function of wavelength and is the modern equivalent\nof using a prism to separate a beam of light into a rainbow of colours. It is a high-precision\nmeasurement that allows us to identify emission & absorption features indicative of specific\nchemical elements present in an object. Spectroscopy is also the most accurate and\nreliable tool that enables classification of astronomical transients and variables. However\nbeing paramount for the classification task, spectroscopy is an extremely time-consuming\nprocess - with exposure times that are much longer than needed to discover objects with\nfilters on an equivalently sized telescope.\nGiven the volume of data expected from the upcoming large-scale sky surveys, obtaining\nspectroscopic observations for every object is not feasible. An alternative approach is to\ntake an image of the object through different filters (also known as passbands), where\neach passband selects light within a specific (broad) wavelength range. This approach is\ncalled photometry, and data obtained through this method are called photometric data.\nFor LSST there are six passbands denoted u, g, r, i, z, y, that select light within different\nwavelength ranges: wavelengths between 300 and 400 nanometers for the u band,\nbetween 400 and 600 nm for the g passband, between 500 and 700 nm for the r band,\nbetween 650 and 850 for the i band, between 800 and 950 nm for the z band, and between\n950 and 1050 nm for the y band. The filter efficiencies vs. wavelength are shown\nin Fig. 2. For reference, the human eye is sensitive to light in the g and r bands. The flux\nof light in each passband, measured as a function of time, is a light curve. Classification\nis performed on these light curves. While a spectroscopic measurement can have great\ndetail, passband data are one brightness measurement per band per exposure. The challenge\nis to classify objects with the low (spectral) resolution light-curve data. Compared\nwith spectroscopy, the advantage of measuring light curves with passband observations\nis that we can observe objects that are much further away and much fainter, and that one\ncan observe many objects in a larger field of view at the same time (rather than measuring\na spectrum of one or a few objects at a time).\nBeware that observations are sometimes degraded by moonlight, twilight, clouds, and\natmospheric effects (that we call ‘seeing’). These degradations result in larger flux uncertainties,\nand this information is included in the light-curve data (see Sec. 3 below). In\nthe challenge data sets, in addition to light curves, two other pieces of information are\nprovided for each object.\nThe training data include accurate redshifts for the objects, but the test-data redshifts are\napproximate measurements based on ugrizy passband measurements from the galaxy\nthat contains the object (a galaxy that is located at the same position on the sky, and at\nthe same redshift as the object). While sources brighten and fade, the host galaxy fluxes\ndon’t change and can thus be measured before a variable object it contains gets bright\nenough to be detected.\nBeware that a few percent of the test data redshifts are catastrophic, meaning that some\nredshift uncertainties greatly underestimate the difference between measured and true\nredshift.\nThe effect of redshift on light from a galaxy reaching earth is illustrated in Fig. 2. The black\ncurve shows a nearby Type Ia supernova spectrum at a redshift of 0.01, corresponding\nto a distance of 140 million light years. While the term ‘nearby’ may seem strange in\nthis case, this distance is indeed nearby when compared with the whole range of cosmic\ndistances. Visual inspection of the supernova spectrum and the filter efficiencies shows\nthat the maximum flux (spectrum summed over filter) is in the green (g) filter. The dashed\ncurve shows a spectrum from a more distant supernova, corresponding to a redshift of\n0.5, or 5.1 billion light years away. The maximum flux is now shifted to the red (r) filter.\nAs the redshift and distance increase, the maximum flux appears in a redder filter: hence\nthe term ‘redshift.’\nThe second piece of additional information is related to extinction from our Galaxy, known\nas the Milky Way. Our light curve measurements are corrected for the atmosphere and\ntelescope transmission (assuming that the telescope is calibrated off standard stars,\nknowledge of this isn’t needed to participate in PLAsTiCC). In addition, we correct each\nlight curve for the absorption of light traveling through Milky Way ‘dust’ on its way to Earth.\nThis absorption is strongest in the ultra-violet u-filter, and weakest in the infrared filters\n(izy). We include the value of the Milky Way extinction at the sky coordinates of each object\nin the data release, with the label ‘MWEBV’, that is an astronomical measure of how\nmuch redder an object appears compared to a Milky Way without dust. Larger MWEBV\nvalues correspond to more Milky Way dust along the line of sight to the objects, making\nthe objects appearing redder. All objects in the PLAsTiCC data are selected to have\nMWEBV < 3, to ensure that we are not looking too close to the disc of the Milky Way, or\nsimilarly to ensure we are not looking through a large amount of dust"},{"metadata":{"_uuid":"1384a5a62386b59da32047da4da13a6b07664108"},"cell_type":"markdown","source":"- 3.データ    \nPLAsTiCCデータは、訓練データセットとテストセットとに分離される。   \n後者は分類されていないデータを分類する必要があります。   \nデータは、複数CSVファイルはKaggleのウェブサイトからアクセスできます。   \nファイルには2種類あります：   \n1.オブジェクトに関する要約（天文学的）情報を含むヘッダーファイル   \n2.6つのフィルタにおける時系列のフラックスからなる各物体のライトカーブデータ（ugrizy）、フラックスの不確実性を含む   \nヘッダーファイルには、一意の識別子 'objid'で索引付けされたデータの各ソースがリストされます。   \n整数。表の各行には、次のようにソースのプロパティがリストされています。   \n•オブジェクトID：オブジェクトID、一意の識別子（int32番号で指定）。   \n•ra：右上がり、天空座標：経度、単位は度です（float32）。   \n•decl：偏角、スカイ座標：緯度、単位は度です（float32の数値）。   \n•gal l：銀河系の経度。単位は度です（浮動小数点数として与えられます）。   \n•gal b：銀河系の緯度。単位は度です（float 32の数値）。   \n•ddf：オブジェクトがDDF測量エリアから来たものであることを識別するブールフラグ。   \nDDFの値ddf = 1）。 DDFフィールドは、完全なWFD調査区域では、DDF欄は著しく小さい不確かさを有している。   \n特定の夜にすべての観測値の追加としてデータが提供されます。   \n•hostgal specz：光源の分光赤方偏移。      \nレッドシフトの尺度、訓練セットに提供され、テストセット（浮動小数点数として与えられます）。   \n•hostgal photoz：天文学のホスト銀河の測光赤方偏移ソース。これはhostgal speczのプロキシであることを意図していますが、両者とホステル写真の違いははるかに小さいとみなされるべきです   \nhostgal speczの正確なバージョン。 hostgal photozはfloat32の数字で与えられます。   \n•hostgal photoz err：LSST調査に基づくhostgal photozの不確実性float32の数値で与えられる投影。   \n•distmod：これ以降のhostgal photozから計算された距離（モジュラス）  \nredshiftはすべてのオブジェクトに対して与えられます（float32の数値として与えられます）。距離の計算モジュラスは一般相対性理論の知識を必要とし、暗宇宙のエネルギーと暗黒物質の内容セクション。   \n•MWEBV = MW E（B-V）：光のこの「消滅」は、天の川（MW）の性質であり、天体の視線に沿ったダストであり、したがって、ソースraの空の座標は、宣言します。これは、通過帯域を決定するために使用されます   \n記載されているように、天文学的光源からの光の依存する調光および赤色化サブセクション2.1で浮動小数点数として与えられています。   \n•target：天文情報源のクラス。これはトレーニングデータで提供されます。目標を正しく決定する（分類確率を正しく割り当てるオブジェクト）は、テストデータの分類課題の目標です。ターゲットint8の数値として与えられます。   \n第2の時系列データの表には、情報源とその情報時間の関数として異なる通過帯域内の輝度、すなわちそれはライトカーブデータである。各行は、特定の時刻におけるソースの観測に対応し、通過帯域。   \nトレーニングデータには1つのライトカーブファイルがありますが、11つのライトカーブファイルがあります。   \n大きいテストセット。テストテーブルには、DDFオブジェクトとWFDオブジェクトが混在しています（1×\nDDFおよび10×WFD表）。   \nDDFの相対的な面積はWFDと比較してLSSTは1/400であり、シミュレートされたPLAsTiCCデータの約1％はDDFサブセットからのものです   \nこのライトカーブテーブルには、次の情報が含まれています。   \n•オブジェクトID：上記のメタデータテーブルと同じキーです（int32の数値として指定）。   \n•mjd：観測の変更ユリウス日（MJD）の時刻。 MJDはユニットです   \n1957年にスミソニアン天体望遠鏡によって導入された時間の記録スプートニクの軌道。 MJDは深夜0時から始まり1858年11月17日。2018年9月25日に58386のMJDを有する.MJDは、式unix time =（MJD40587）86400でUnixエポック時間に変換されます。   \n単位は日であり、数値は浮動小数点数として与えられます。   \n•通過帯域：u、g、r、i、z、y = 0,1,2,3,4,5のような特定のLSST通過帯域幅の整数それが見られた。これらはint8番号で与えられます。   \n•flux：観測された通過帯域の測定されたフラックス（輝度）。   \nパスバンド列。フラックスはMWEBVに対して補正されるが、大きな塵の消滅については訂正にもかかわらず、不確実性ははるかに大きくなります。   \nほこりはfloat32 number（fluxとfluxの両方の単位は任意です）。   \n•フラックスエラー：上記のフラックス測定の不確実性float32番号。   \n•detected：検出された場合= 1、オブジェクトの明るさは3σ基準テンプレートに対して相対的に高レベルである。これはブール値フラグとして与えられます。   \nライトカーブに関するいくつかの注意点は次のとおりです。   \n•データギャップ：さまざまな時間帯でさまざまな通過帯域が使用されます。   \n•銀河系と銀河系との関係：私たち自身の銀河系における物体の赤方偏移銀河はゼロとして与えられる。    \n•負のフラックス：統計的な変動（例えば、空の明るさ）および明るさが評価されるように、光束は暗い光源に対して負であり、真のフラックスはゼロに近い。 さらに、事前調査画像に実際にその真の「ゼロ」よりも光束が明るい場合、これは計算される。   \n3.1.データの取得と分類の採点   \n参加者は、確率的分類のマトリックスを提出する必要があります。   \n≈350万行と15列（トレーニングデータにすでに表示されている14個と、他のクラス）、行（オブジェクト）あたりのすべてのクラスの確率の合計は1です。   \nセクション4では、分類確率を評価するために使用されるメトリックについて説明します。   \nPLAsTiCCのメトリックの選択肢については、次の論文で説明し、1つのソースに集中するのではなく、クラス間でバランスのとれた指標を確保する   \n興味挑戦の一環として、Jupyterノートブックの「スターターキット」の例を紹介します。   \nより多くの入門資料と別のノートブックを使用して、チャレンジ。   \n3.2.レーニングデータとテストデータトレーニングデータは上記の説明に従い、8000個の天文情報源の集合であり、高価な分光法が可能となる。テストデータは、すべてのデータ分光器を持たないもので、≈350万個の非常に大きなセットです。   \nこのため、テストデータは、hostgal specz列の「NULL」エントリを数％以外の数パーセントで持っています   \nテストデータ内のオブジェクトのすべてのテスト・データのターゲット列はNULLです。   \nさらに、訓練データ特性は、試験データセットの分布の非代表的なものである。   \nトレーニングデータは、ほとんどが近くの低赤方偏移でより明るいオブジェクトで構成され、テストデータには、より遠い（より高い赤方偏移）およびより暗いオブジェクトが含まれています。   \nしたがって、トレーニングデータに対応するものを持たないテストデータ内のオブジェクト。   "},{"metadata":{"_uuid":"52a82616e8e959fbd4f29cdc24440fce1f647c43"},"cell_type":"markdown","source":"3. The data\nThe PLAsTiCC data are separated into a training data set and a test set; the latter is the\ndata without classifications that needs to be classified. The data are provided in multiple\nCSV files, that are accessed from the Kaggle website.7 There are two types of files:\n1. header files that contain summary (astronomical) information about the objects\n2. light-curve data for each object consisting of a time series of fluxes in six filters\n(ugrizy), including flux uncertainties\nThe header file lists each source in the data indexed by a unique identifier ’objid’, that is\nan integer. Each row of the table lists the properties of the source as follows:\n• object id: the Object ID, unique identifier (given as int32 numbers).\n• ra: right ascension, sky coordinate: longitude, units are degrees (given as float32\nnumbers).\n• decl: declination, sky coordinate: latitude, units are degrees (given as float32 numbers).\n• gal l: Galactic longitude, units are degrees (given as float32 numbers).\n• gal b: Galactic lattitude, units are degrees (given as float32 numbers).\n• ddf: A Boolean flag to identify the object as coming from the DDF survey area (with\nvalue ddf = 1 for the DDF). Note that while the DDF fields are contained within the\nfull WFD survey area, the DDF fields have significantly smaller uncertainties, given\nthat the data are provided as additions of all observations in a given night.\n• hostgal specz: the spectroscopic redshift of the source8\n. This is an extremely accurate\nmeasure of redshift, provided for the training set and a small fraction of the\ntest set (given as float32 numbers).\n• hostgal photoz: The photometric redshift of the host galaxy of the astronomical\nsource. While this is meant to be a proxy for hostgal specz, there can be large\ndifferences between the two and hostgal photoz should be regarded as a far less\naccurate version of hostgal specz. The hostgal photoz is given as float32 numbers.\n• hostgal photoz err: The uncertainty on the hostgal photoz based on LSST survey\nprojections, given as float32 numbers.\n• distmod: The distance (modulus) calculated from the hostgal photoz since this\nredshift is given for all objects (given as float32 numbers). Computing the distance\nmodulus requires knowledge of General Relativity, and assumed values of the dark\nenergy and dark matter content of the Universe, as mentioned in the introduction\nsection.\n• MWEBV = MW E(B-V): this ‘extinction’ of light is a property of the Milky Way (MW)\ndust along the line of sight to the astronomical source, and is thus a function of\nthe sky coordinates of the source ra, decl. This is used to determine a passband\ndependent dimming and reddening of light from astronomical sources as described\nin subsection 2.1, and is given as float32 numbers.\n• target: The class of the astronomical source. This is provided in the training data.\nCorrectly determining the target (correctly assigning classification probabilities to\nthe objects) is the goal of the classification challenge for the test data. The target\nis given as int8 numbers.\nThe second table of time-series data contains information about the sources and their\nbrightness in different passbands as a function of time i.e. it is the light-curve data. Each\nrow of this table corresponds to an observation of the source at a particular time and\npassband. There is one light-curve file for the training data, but 11 light-curve files for\nthe (much larger) test set. The test tables contain a mix of DDF and WFD objects (1×\nDDF and 10× WFD tables). While the relative areas of the DDF compared to the WFD for\nLSST will be 1/400, roughly 1% of the simulated PLAsTiCC data are from the DDF subset.\nThis light-curve tables include the following information:\n• object id: Same key as in the metadata table above, given as int32 numbers.\n• mjd: the time in Modified Julian Date (MJD) of the observation. The MJD is a unit\nof time introduced by the Smithsonian Astrophysical Observatory in 1957 to record\nthe orbit of Sputnik. The MJD is defined to have a starting point of midnight on\nNovember 17, 1858. The 25th of September 2018 has an MJD of 58386. The MJD can\nbe converted to Unix epoch time with the formula unix time = (MJD40587)86400.\nThe units are days, and the numbers are given as float64 numbers.\n• passband: The specific LSST passband integer, such that u, g, r, i, z, y = 0, 1, 2, 3, 4, 5\nin which it was viewed. These are given as int8 numbers.\n• flux: the measured flux (brightness) in the passband of observation as listed in the\npassband column. The flux is corrected for MWEBV, but for large dust extinctions the\nuncertainty will be much larger in spite of the correction. The dust is given as a\nfloat32 number (note that the units for both flux and flux err are arbitrary).\n• flux err: the uncertainty on the measurement of the flux listed above, given as\nfloat32 number.\n• detected: If detected= 1, the object’s brightness is significantly different at the 3σ\nlevel relative to the reference template. This is given as a Boolean flag.\nA few caveats about the light-curve data are as follows:\n• Data gaps: Different passbands are taken at different times, sometimes many days\napart.\n• Galactic vs extragalactic: The given redshift for objects in our own Milky Way\ngalaxy is given as zero.\n• Negative Flux: Due to statistical fluctuations (of e.g. the sky brightness) and the\nway the brightness is estimated, the flux may be negative for dim sources, where\nthe true flux is close to zero. Additionally, if the pre-survey image actually contains a\nflux brighter than its true ‘zero’, this can lead to a negative flux when the difference\nis computed.\n3.1. Obtaining the data and scoring a classification\nParticipants will be required to submit a matrix of probabilistic classifications: a table with\n≈ 3.5 million rows and 15 columns (the 14 already seen in the training data and one\nothers class), where the sum of probabilities across all classes per row (object) is unity.\nSection 4 describes the metric that will be used to evaluate the classification probabilities.\nThe metric choices for PLAsTiCC are described in an upcoming paper and focused on\nensuring a balanced metric across classes, rather than just focusing on one source of\ninterest. As part of the challenge, we provide an example Jupyter notebook ‘starter kit’\nwith more introductory material and another notebook to compute the metrics for the\nchallenge9\n.\n3.2. Training and test data\nThe training data follow the description above and have the properties and light curves of\na set of 8000 astronomical sources and are meant to represent the brighter objects for\nwhich obtaining expensive spectroscopy is possible. The test data represent all the data\nthat have no spectroscopy, and is a much larger set of ≈ 3.5 million objects. Therefore,\nthe test data have ‘NULL’ entries for the hostgal specz column for all but a few percent\nof object in the test data. The target column is ‘NULL’ for all test data. Moreover, the\ntraining data properties are non-representative of distributions of the the test data set.\nThe training data are mostly composed of nearby, low-redshift, brighter objects while the\ntest data contain more distant (higher redshift) and fainter objects. Therefore, there are\nobjects in the test data that do not have counterparts in the training data."},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","collapsed":true,"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":false},"cell_type":"markdown","source":"- 4.チャレンジへの参加   \nPLASTiCCチャレンジエントリは、ヘッダー内の各ソースを分類するために必要です   \nファイルの特性と光のカーブに基づいてテストデータのセットを作成します。分類は確率Pij、P（class | datai、トレーニングデータ、知識）、テストデータ内のi番目のソースがクラスjのメンバーである確率\nオブジェクト、トレーニングデータセット、および参加者が他の場所で取得した可能性がある外部知識のプロパティとライトカーブデータの組み合わせで天文学の知識はPLAsTiCCに参加する必要はありません。   \nエントリを指定するには単一のソースについて、参加者は、そのソースの確率をトレーニングセット内の互いに排他的な（重複しない）14のクラスのそれぞれと、トレーニングセットのいずれかのクラスに属するので、他のクラスによって示されるクラス。   \n特定のクラスjおよびオブジェクトiに対するPijの高い値は、参加者i番目のソースはj番目のクラスのメンバーである可能性が高いと考えています。真の確率として、量Pijは以下の基準を満たさなければならない。  PLAsTiCCエントリが有効であるためには、各クラスの確率を含む必要があります。   \n天文学的情報源、すなわち、エントリーは、任意の情報源またはクラスの確率を除外することができない。   \n課題に勝つためには、すべてのエントリーが有効で、PLAsTiCCメトリックスコアを最小限に抑える必要があります。   \nチャレンジに使用されるメトリックは、重み付けされたログ損失メトリックです   \nここで、i番目のオブジェクトがj番目のクラスに由来する場合はτi、j = 1、それ以外の場合は0、Nj\nそれは任意のクラスj内のオブジェクトの数、およびwjは、クラスごとの個別の重みです。   \n全体的なメトリックへの相対的な貢献を反映する（例えば、それがいかに望ましいかに依存する\nクラスj内のオブジェクトを正しく分類する）。これらのwjはチャレンジ。  \nNjを平均化することは、PLAsTiCCが1つの特定のクラスに集中するのではなく、すべてのオブジェクトを平均してよく分類します。   \nこの添付の論文（Malz、Hlozek et alこのメモと一緒にリリースされました）。   \nある参加者が、あるオブジェクトが特定のクラスであると判断する分類子を使用する場合確率を提供するものよりも、参加者は自分自身の処方箋を使用して確定的な分類から確率を定義する。たとえば、分類された情報源第1のクラスは1.0の確率を与えられ、他の全てのクラス確率は0.0であってもよいし、第1のクラスに0.9を割り当ててもよく、他のすべてのクラスには一様確率が1になるように確率   \nたとえば、3つのオブジェクトのセットを2つのクラスの「スター」に分類することが課題であった場合、\nまたは 'galaxy'クラス（およびその他のクラス）の場合、返される分類テーブルは3×3マトリックス   \n行の確率は1に正規化されるべきである。   \nPLAsTiCCのデータ、競技会、エントリーのルールについては、Kaggle web PLAsTiCCのページ。    \nこのコンテストには、3つの指標ベース（「一般」）の受賞者と、追加の科学賞を受賞した彼らの経験レベル：天文学の分野での以前の経験は必要ない参加。"},{"metadata":{"trusted":true,"_uuid":"2bdd6064bfb2303695f16b393c771b2643175496"},"cell_type":"markdown","source":"4. Challenge participation\nPLAsTiCC challenge entries are required to classify each of the sources in the header\nfile of the test data set based on their properties and light curves. The classification is\ndone though the assignment of probabilities Pij , P(class|datai\n,training data, knowledge),\nthe probability that the ith source in the test data is a member of the class j based on\nthe combination of properties and light-curve data for the object, the training data set,\nand any outside knowledge the participant may have acquired elsewhere, although prior\nknowledge of astronomy is not required to participate in PLAsTiCC. To specify the entry\nfor a single source, the participant provides the probabilities of that source belonging to\neach of the mutually exclusive (non-overlapping) 14 classes in the training set, and of not\nbelonging to any of the classes in the training set and therefore denoted by the others\nclass. High values of Pij for a particular class j and object i indicate that the participant\nbelieves that the i-th source is likely to be a member of the jth class. As true probabilities,\nthe quantities Pij must satisfy the following criteria:\nFor a PLAsTiCC entry to be valid, it will have to include probabilities for each class and\nastronomical source, i.e. an entry cannot leave out probabilities on any source, or class.\nTo win the challenge, all entries should be valid and minimize the PLAsTiCC metric score.\nThe metric used for the challenge is a weighted log-loss metric\nwhere τi,j = 1 if the ith object comes from the jth class and 0 otherwise, and Nj\nis the\nnumber of objects in any given class j, and wj are individual weights per class which\nreflect relative contribution to the overall metric (depending on e.g. how desireable it is\nto have objects in class j classified correctly). These wj are hidden to the participants of\nthe challenge. The averaging over Nj reflects that PLAsTiCC is designed to reward those\nwho classify all objects well on average, rather than focusing on one particular class. This\nis discussed more in an accompanying paper (see Malz, Hlozek ˇ et al. 2018 which is\nreleased with this note).\nIf a participant uses a classifier that decides that an object is of a particular class rather\nthan one that provides probabilities, that participant has to use their own prescription to\ndefine probabilities from deterministic classifications. For example, a source classified\nas the first class may be given a probability of 1.0 and all other classes probabilities of\n0.0, or the first class may be assigned 0.9 and all other classes may be given uniform\nprobabilities so that the probabilities sum to unity.\nThe PLAsTiCC data, competition and rules for entry are described on the Kaggle web\npage for PLAsTiCC. The competition will feature three metric-based (‘general’) winners and\nadditional science-focused prizes, and we encourage those interested to enter regardless\nof their experience level: previous experience in the field of astronomy is not required for\nparticipation."},{"metadata":{"trusted":true,"_uuid":"ac937e5b64557a563d454c4680d85b18ee140f04"},"cell_type":"code","source":"%matplotlib inline\n#モジュールの読み込み\nimport os\nimport pandas as pd\n\nimport numpy as np\nimport matplotlib.pyplot as plt\n\nprint(os.listdir(\"../input\"))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"caf159163dea5fb50de6e2a3f13ab1660830739c"},"cell_type":"code","source":"#train_set = pd.read_csv('../input/training_set.csv', nrows=100)\ntrain_meta = pd.read_csv('../input/training_set_metadata.csv')\ntrain_meta.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"18ee42a723a636798e1704817ffeb6da217cc5a2"},"cell_type":"code","source":"plt.scatter(train_meta.hostgal_specz,train_meta.target)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f06897e4a40bd37195d1158bcb702d6a7282ce86"},"cell_type":"code","source":"plt.scatter(train_meta.hostgal_photoz,train_meta.target)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"74170f8bf93f3f6b505005fa454b7c9eacfee1a5"},"cell_type":"code","source":"plt.scatter(train_meta.target, (train_meta.hostgal_photoz*train_meta.hostgal_specz))\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b1581c533cc8656cc32b41dc76f0485e0da20a82"},"cell_type":"code","source":"from mpl_toolkits.mplot3d import Axes3D\nfig = plt.figure()\nax = fig.add_subplot(111, projection='3d')\nax.scatter(train_meta.hostgal_photoz, train_meta.target, train_meta.hostgal_specz)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"7ce0d7d313df19718897c1b9eb586df08a56122d"},"cell_type":"code","source":"plt.scatter(train_meta.hostgal_photoz,train_meta.hostgal_specz)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"64ab0271f0d53bf748849cdd8fb5ea76a4ceb7dc"},"cell_type":"code","source":"plt.scatter((train_meta.ra**2+train_meta.decl**2),train_meta.target)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"bea1608c019eac0d5517d9b92de6479759e45a5e"},"cell_type":"code","source":"fig = plt.figure()\nax = fig.add_subplot(111, projection='3d')\nax.scatter(train_meta.target, (train_meta.hostgal_photoz*train_meta.hostgal_specz),(train_meta.ra**2+train_meta.decl**2))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"5da77eccffc4ae12bf1e12582bb2817c289bd447"},"cell_type":"code","source":"train_meta2 = train_meta[train_meta['hostgal_specz'] >0.0]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f0d840632e029217833498fecc09ec19cf94e5b0"},"cell_type":"code","source":"plt.scatter(train_meta2.hostgal_specz,train_meta2.target)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ab46551caaa0e715260d3072c71f35b24a93b5eb"},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}