{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<div style=\"padding:20px;color:#6A79BA;margin:0;font-size:180%;text-align:center;display:fill;border-radius:5px;background-color:white;overflow:hidden;font-weight:600\">JPX Tokyo Stock Exchange Prediction</div>\n\n<img src=\"https://storage.googleapis.com/kaggle-competitions/kaggle/34349/logos/header.png?t=2022-03-09-00-33-57\" style=\"border-radius:5px\">\n\n# <div style=\"padding:20px;color:white;margin:0;font-size:100%;text-align:left;display:fill;border-radius:5px;background-color:#6A79BA;overflow:hidden\">1 | Competition Overview</div>\n\nSuccess in any financial market requires one to identify solid investments. When a stock or derivative is undervalued, it makes sense to buy. If it's overvalued, perhaps it's time to sell. While these finance decisions were historically made manually by professionals, technology has ushered in new opportunities for retail investors. Data scientists, specifically, may be interested to explore quantitative trading, where decisions are executed programmatically based on predictions from trained models. In this competition, financial data for the Japanese market will be provided, allowing retail investors to analyze the market to the fullest extent. \n\nJapan Exchange Group, Inc. (JPX) is a holding company operating one of the largest stock exchanges in the world, Tokyo Stock Exchange (TSE), and derivatives exchanges Osaka Exchange (OSE) and Tokyo Commodity Exchange (TOCOM). JPX is [hosting this competition](https://www.kaggle.com/competitions/jpx-tokyo-stock-exchange-prediction) and is supported by AI technology company AlpacaJapan Co., Ltd.\n\nThis competition will involve building portfolios from the stocks eligible for predictions. Specifically, each participant ranks the stocks from highest to lowest expected returns and is evaluated on the difference in returns between the top and bottom 200 stocks. You'll have access to financial data from the Japanese market, such as stock information and historical stock prices to train and test your model. After the training phase is complete, the competition will compare your models against real future returns.\n\nThe training data begins in January 2017 with about 1,860 stocks with new stocks added through December 2020 for a total of 2,000 stocks. Below is a brief overview and descriptive statistics of the variables in the training data. ","metadata":{}},{"cell_type":"code","source":"import warnings, gc\nimport numpy as np \nimport pandas as pd\nimport matplotlib.colors\nimport seaborn as sns\nimport plotly.express as px\nimport plotly.graph_objects as go\nfrom plotly.subplots import make_subplots\nfrom plotly.offline import init_notebook_mode\nfrom datetime import datetime, timedelta\nfrom sklearn.model_selection import TimeSeriesSplit\nfrom sklearn.metrics import mean_squared_error,mean_absolute_error\nfrom lightgbm import LGBMRegressor\nfrom decimal import ROUND_HALF_UP, Decimal\nwarnings.filterwarnings(\"ignore\")\nimport plotly.figure_factory as ff\n\ninit_notebook_mode(connected=True)\ntemp = dict(layout=go.Layout(font=dict(family=\"Franklin Gothic\", size=12), width=800))\ncolors=px.colors.qualitative.Plotly\n\ntrain=pd.read_csv(\"../input/jpx-tokyo-stock-exchange-prediction/train_files/stock_prices.csv\", parse_dates=['Date'])\nstock_list=pd.read_csv(\"../input/jpx-tokyo-stock-exchange-prediction/stock_list.csv\")\n\nprint(\"The training data begins on {} and ends on {}.\\n\".format(train.Date.min(),train.Date.max()))\ndisplay(train.describe().style.format('{:,.2f}'))","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-06-27T09:23:17.273882Z","iopub.execute_input":"2022-06-27T09:23:17.274202Z","iopub.status.idle":"2022-06-27T09:23:28.394127Z","shell.execute_reply.started":"2022-06-27T09:23:17.274124Z","shell.execute_reply":"2022-06-27T09:23:28.393267Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"padding:20px;color:white;margin:0;font-size:100%;text-align:left;display:fill;border-radius:5px;background-color:#6A79BA;overflow:hidden\">2 | Exploratory Data Analysis</div>","metadata":{}},{"cell_type":"code","source":"train_date=train.Date.unique()\nreturns=train.groupby('Date')['Target'].mean().mul(100).rename('Average Return')\nclose_avg=train.groupby('Date')['Close'].mean().rename('Closing Price')\nvol_avg=train.groupby('Date')['Volume'].mean().rename('Volume')\n\nfig = make_subplots(rows=3, cols=1, \n                    shared_xaxes=True)\nfor i, j in enumerate([returns, close_avg, vol_avg]):\n    fig.add_trace(go.Scatter(x=train_date, y=j, mode='lines',\n                             name=j.name, marker_color=colors[i]), row=i+1, col=1)\nfig.update_xaxes(rangeslider_visible=False,\n                 rangeselector=dict(\n                     buttons=list([\n                         dict(count=6, label=\"6m\", step=\"month\", stepmode=\"backward\"),\n                         dict(count=1, label=\"1y\", step=\"year\", stepmode=\"backward\"),\n                         dict(count=2, label=\"2y\", step=\"year\", stepmode=\"backward\"),\n                         dict(step=\"all\")])),\n                 row=1,col=1)\nfig.update_layout(template=temp,title='JPX Market Average Stock Return, Closing Price, and Shares Traded', \n                  hovermode='x unified', height=700, \n                  yaxis1=dict(title='Stock Return', ticksuffix='%'), \n                  yaxis2_title='Closing Price', yaxis3_title='Shares Traded',\n                  showlegend=False)\nfig.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-06-27T09:23:28.396128Z","iopub.execute_input":"2022-06-27T09:23:28.3964Z","iopub.status.idle":"2022-06-27T09:23:28.89161Z","shell.execute_reply.started":"2022-06-27T09:23:28.396366Z","shell.execute_reply":"2022-06-27T09:23:28.888657Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The graphs above show the market's average stock return, closing price, and shares traded since January 2017. While there has been much fluctuation over the past four years, the number of shares traded has slightly decreased from the volume in early 2017.","metadata":{}},{"cell_type":"code","source":"stock_list['SectorName']=[i.rstrip().lower().capitalize() for i in stock_list['17SectorName']]\nstock_list['Name']=[i.rstrip().lower().capitalize() for i in stock_list['Name']]\ntrain_df = train.merge(stock_list[['SecuritiesCode','Name','SectorName']], on='SecuritiesCode', how='left')\ntrain_df['Year'] = train_df['Date'].dt.year\nyears = {year: pd.DataFrame() for year in train_df.Year.unique()[::-1]}\nfor key in years.keys():\n    df=train_df[train_df.Year == key]\n    years[key] = df.groupby('SectorName')['Target'].mean().mul(100).rename(\"Avg_return_{}\".format(key))\ndf=pd.concat((years[i].to_frame() for i in years.keys()), axis=1)\ndf=df.sort_values(by=\"Avg_return_2021\")\n\nfig = make_subplots(rows=1, cols=5, shared_yaxes=True)\nfor i, col in enumerate(df.columns):\n    x = df[col]\n    mask = x<=0\n    fig.add_trace(go.Bar(x=x[mask], y=df.index[mask],orientation='h', \n                         text=x[mask], texttemplate='%{text:.2f}%',textposition='auto',\n                         hovertemplate='Average Return in %{y} Stocks = %{x:.4f}%',\n                         marker=dict(color='red', opacity=0.7),name=col[-4:]), \n                  row=1, col=i+1)\n    fig.add_trace(go.Bar(x=x[~mask], y=df.index[~mask],orientation='h', \n                         text=x[~mask], texttemplate='%{text:.2f}%', textposition='auto', \n                         hovertemplate='Average Return in %{y} Stocks = %{x:.4f}%',\n                         marker=dict(color='green', opacity=0.7),name=col[-4:]), \n                  row=1, col=i+1)\n    fig.update_xaxes(range=(x.min()-.15,x.max()+.15), title='{} Returns'.format(col[-4:]), \n                     showticklabels=False, row=1, col=i+1)\nfig.update_layout(template=temp,title='Yearly Average Stock Returns by Sector', \n                  hovermode='closest',margin=dict(l=250,r=50),\n                  height=600, width=1000, showlegend=False)\nfig.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-06-27T09:23:28.892559Z","iopub.execute_input":"2022-06-27T09:23:28.892773Z","iopub.status.idle":"2022-06-27T09:23:30.604014Z","shell.execute_reply.started":"2022-06-27T09:23:28.892742Z","shell.execute_reply":"2022-06-27T09:23:30.603245Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In 2021, nearly all industries saw a positive return on average, with the highest in Energy Resources at about 0.13% overall, while in 2018, all sectors saw a negative return except for Electric Power & Gas. \n\nSince some of the stocks were added in December 2020, I will use the data filtered after this date so that the data will consist of 231 days of stock prices for all 2,000 stocks.","metadata":{}},{"cell_type":"code","source":"train_df=train_df[train_df.Date>'2020-12-23']\nprint(\"New Train Shape {}.\\nMissing values in Target = {}\".format(train_df.shape,train_df['Target'].isna().sum()))","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-06-27T09:23:30.606128Z","iopub.execute_input":"2022-06-27T09:23:30.606382Z","iopub.status.idle":"2022-06-27T09:23:30.673704Z","shell.execute_reply.started":"2022-06-27T09:23:30.60635Z","shell.execute_reply":"2022-06-27T09:23:30.672722Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = go.Figure()\nx_hist=train_df['Target']\nfig.add_trace(go.Histogram(x=x_hist*100,\n                           marker=dict(color=colors[0], opacity=0.7, \n                                       line=dict(width=1, color=colors[0])),\n                           xbins=dict(start=-40,end=40,size=1)))\nfig.update_layout(template=temp,title='Target Distribution', \n                  xaxis=dict(title='Stock Return',ticksuffix='%'), height=450)\nfig.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-06-27T09:23:30.675337Z","iopub.execute_input":"2022-06-27T09:23:30.675599Z","iopub.status.idle":"2022-06-27T09:23:31.342236Z","shell.execute_reply.started":"2022-06-27T09:23:30.675563Z","shell.execute_reply":"2022-06-27T09:23:31.34119Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pal = ['hsl('+str(h)+',50%'+',50%)' for h in np.linspace(0, 360, 18)]\nfig = go.Figure()\nfor i, sector in enumerate(df.index[::-1]):\n    y_data=train_df[train_df['SectorName']==sector]['Target']\n    fig.add_trace(go.Box(y=y_data*100, name=sector,\n                         marker_color=pal[i], showlegend=False))\nfig.update_layout(template=temp, title='Target Distribution by Sector',\n                  yaxis=dict(title='Stock Return',ticksuffix='%'),\n                  margin=dict(b=150), height=750, width=900)\nfig.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-06-27T09:23:31.343482Z","iopub.execute_input":"2022-06-27T09:23:31.344235Z","iopub.status.idle":"2022-06-27T09:23:33.490516Z","shell.execute_reply.started":"2022-06-27T09:23:31.344177Z","shell.execute_reply":"2022-06-27T09:23:33.488857Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"While most sectors have returns between 1% and -1%, there are quite a few outliers across all industries, with some returns as high as 62% in Commercial & Wholesale Trade and others as low as -31% in IT & Services sector. The graph below shows the stock price movements within each sector.","metadata":{}},{"cell_type":"code","source":"train_date=train_df.Date.unique()\nsectors=train_df.SectorName.unique().tolist()\nsectors.insert(0, 'All')\nopen_avg=train_df.groupby('Date')['Open'].mean()\nhigh_avg=train_df.groupby('Date')['High'].mean()\nlow_avg=train_df.groupby('Date')['Low'].mean()\nclose_avg=train_df.groupby('Date')['Close'].mean() \nbuttons=[]\n\nfig = go.Figure()\nfor i in range(18):\n    if i != 0:\n        open_avg=train_df[train_df.SectorName==sectors[i]].groupby('Date')['Open'].mean()\n        high_avg=train_df[train_df.SectorName==sectors[i]].groupby('Date')['High'].mean()\n        low_avg=train_df[train_df.SectorName==sectors[i]].groupby('Date')['Low'].mean()\n        close_avg=train_df[train_df.SectorName==sectors[i]].groupby('Date')['Close'].mean()        \n    \n    fig.add_trace(go.Candlestick(x=train_date, open=open_avg, high=high_avg,\n                                 low=low_avg, close=close_avg, name=sectors[i],\n                                 visible=(True if i==0 else False)))\n    \n    visibility=[False]*len(sectors)\n    visibility[i]=True\n    button = dict(label = sectors[i],\n                  method = \"update\",\n                  args=[{\"visible\": visibility}])\n    buttons.append(button)\n    \nfig.update_xaxes(rangeslider_visible=True,\n                 rangeselector=dict(\n                     buttons=list([\n                         dict(count=3, label=\"3m\", step=\"month\", stepmode=\"backward\"),\n                         dict(count=6, label=\"6m\", step=\"month\", stepmode=\"backward\"),\n                         dict(step=\"all\")]), xanchor='left',yanchor='bottom', y=1.16, x=.01))\nfig.update_layout(template=temp,title='Stock Price Movements by Sector', \n                  hovermode='x unified', showlegend=False, width=1000,\n                  updatemenus=[dict(active=0, type=\"dropdown\",\n                                    buttons=buttons, xanchor='left',\n                                    yanchor='bottom', y=1.01, x=.01)],\n                  yaxis=dict(title='Stock Price'))\nfig.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-06-27T09:23:33.49289Z","iopub.execute_input":"2022-06-27T09:23:33.493367Z","iopub.status.idle":"2022-06-27T09:23:39.021999Z","shell.execute_reply.started":"2022-06-27T09:23:33.493304Z","shell.execute_reply":"2022-06-27T09:23:39.021367Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In the candlestick charts above, the boxes represent the daily spread between the open and close prices and the lines represent the spread between the low and high prices. The color of the boxes indicates whether the close price was greater or lower than the open price, with green indicating a higher closing price on that day and red indicating a lower closing price. In late August, the market saw a consecutive 14-day period where the close price was greater than the open price.","metadata":{}},{"cell_type":"code","source":"stock=train_df.groupby('Name')['Target'].mean().mul(100)\nstock_low=stock.nsmallest(7)[::-1].rename(\"Return\")\nstock_high=stock.nlargest(7).rename(\"Return\")\nstock=pd.concat([stock_high, stock_low], axis=0).reset_index()\nstock['Sector']='All'\nfor i in train_df.SectorName.unique():\n    sector=train_df[train_df.SectorName==i].groupby('Name')['Target'].mean().mul(100)\n    stock_low=sector.nsmallest(7)[::-1].rename(\"Return\")\n    stock_high=sector.nlargest(7).rename(\"Return\")\n    sector_stock=pd.concat([stock_high, stock_low], axis=0).reset_index()\n    sector_stock['Sector']=i\n    stock=stock.append(sector_stock,ignore_index=True)\n    \nfig=go.Figure()\nbuttons = []\nfor i, sector in enumerate(stock.Sector.unique()):\n    \n    x=stock[stock.Sector==sector]['Name']\n    y=stock[stock.Sector==sector]['Return']\n    mask=y>0\n    fig.add_trace(go.Bar(x=x[mask], y=y[mask], text=y[mask], \n                         texttemplate='%{text:.2f}%',\n                         textposition='auto',\n                         name=sector, visible=(False if i != 0 else True),\n                         hovertemplate='%{x} average return: %{y:.3f}%',\n                         marker=dict(color='green', opacity=0.7)))\n    fig.add_trace(go.Bar(x=x[~mask], y=y[~mask], text=y[~mask], \n                         texttemplate='%{text:.2f}%',\n                         textposition='auto',\n                         name=sector, visible=(False if i != 0 else True),\n                         hovertemplate='%{x} average return: %{y:.3f}%',\n                         marker=dict(color='red', opacity=0.7)))\n    \n    visibility=[False]*2*len(stock.Sector.unique())\n    visibility[i*2],visibility[i*2+1]=True,True\n    button = dict(label = sector,\n                  method = \"update\",\n                  args=[{\"visible\": visibility}])\n    buttons.append(button)\n\nfig.update_layout(title='Stocks with Highest and Lowest Returns by Sector',\n                  template=temp, yaxis=dict(title='Average Return', ticksuffix='%'),\n                  updatemenus=[dict(active=0, type=\"dropdown\",\n                                    buttons=buttons, xanchor='left',\n                                    yanchor='bottom', y=1.01, x=.01)], \n                  margin=dict(b=150),showlegend=False,height=700, width=900)\nfig.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-06-27T09:23:39.023252Z","iopub.execute_input":"2022-06-27T09:23:39.02363Z","iopub.status.idle":"2022-06-27T09:23:40.358027Z","shell.execute_reply.started":"2022-06-27T09:23:39.023595Z","shell.execute_reply":"2022-06-27T09:23:40.357384Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Among stocks with the highest return on average since December 2020, 6 of the 7 were in the IT & Services sector and one was in the Pharmaceutical sector. The IT & Services and Pharmaceutical sectors also make up 6 of the stocks with the lowest returns on average. The graph below shows the relationships between the top 5 stocks with the highest average returns, Enechange Ltd., For Startups, Inc., Symbio Pharmaceuticals, Fronteo, Inc., and Emnet Japan Co. Ltd.","metadata":{}},{"cell_type":"code","source":"stocks=train_df[train_df.SecuritiesCode.isin([4169,7089,4582,2158,7036])]\ndf_pivot=stocks.pivot_table(index='Date', columns='Name', values='Close').reset_index()\npal=['rgb'+str(i) for i in sns.color_palette(\"coolwarm\", len(df_pivot))]\n\nfig = ff.create_scatterplotmatrix(df_pivot.iloc[:,1:], diag='histogram', name='')\nfig.update_traces(marker=dict(color=pal, opacity=0.9, line_color='white', line_width=.5))\nfig.update_layout(template=temp, title='Scatterplots of Highest Performing Stocks', \n                  height=1000, width=1000, showlegend=False)\nfig.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-06-27T09:23:40.359261Z","iopub.execute_input":"2022-06-27T09:23:40.359662Z","iopub.status.idle":"2022-06-27T09:23:41.022555Z","shell.execute_reply.started":"2022-06-27T09:23:40.359626Z","shell.execute_reply":"2022-06-27T09:23:41.021786Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"corr=train_df.groupby('SecuritiesCode')[['Target','Close']].corr().unstack().iloc[:,1]\nstocks=corr.nlargest(10).rename(\"Return\").reset_index()\nstocks=stocks.merge(train_df[['Name','SecuritiesCode']], on='SecuritiesCode').drop_duplicates()\npal=sns.color_palette(\"magma_r\", 14).as_hex()\nrgb=['rgba'+str(matplotlib.colors.to_rgba(i,0.7)) for i in pal]\n\nfig = go.Figure()\nfig.add_trace(go.Bar(x=stocks.Name, y=stocks.Return, text=stocks.Return, \n                     texttemplate='%{text:.2f}', name='', width=0.8,\n                     textposition='outside',marker=dict(color=rgb, line=dict(color=pal,width=1)),\n                     hovertemplate='Correlation of %{x} with target = %{y:.3f}'))\nfig.update_layout(template=temp, title='Most Correlated Stocks with Target Variable',\n                  yaxis=dict(title='Correlation',showticklabels=False), \n                  xaxis=dict(title='Stock',tickangle=45), margin=dict(b=100),\n                  width=800,height=500)\nfig.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-06-27T09:23:41.024947Z","iopub.execute_input":"2022-06-27T09:23:41.025638Z","iopub.status.idle":"2022-06-27T09:23:41.769995Z","shell.execute_reply.started":"2022-06-27T09:23:41.0256Z","shell.execute_reply":"2022-06-27T09:23:41.769336Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_pivot=train_df.pivot_table(index='Date', columns='SectorName', values='Close').reset_index()\ncorr=df_pivot.corr().round(2)\nmask=np.triu(np.ones_like(corr, dtype=bool))\nc_mask = np.where(~mask, corr, 100)\nc=[]\nfor i in c_mask.tolist()[1:]:\n    c.append([x for x in i if x != 100])\n    \ncor=c[::-1]\nx=corr.index.tolist()[:-1]\ny=corr.columns.tolist()[1:][::-1]\nfig=ff.create_annotated_heatmap(z=cor, x=x, y=y, \n                                hovertemplate='Correlation between %{x} and %{y} stocks = %{z}',\n                                colorscale='viridis', name='')\nfig.update_layout(template=temp, title='Stock Correlation between Sectors',\n                  margin=dict(l=250,t=270),height=800,width=900,\n                  yaxis=dict(showgrid=False, autorange='reversed'),\n                  xaxis=dict(showgrid=False))\nfig.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-06-27T09:23:41.771025Z","iopub.execute_input":"2022-06-27T09:23:41.77142Z","iopub.status.idle":"2022-06-27T09:23:41.9477Z","shell.execute_reply.started":"2022-06-27T09:23:41.771384Z","shell.execute_reply":"2022-06-27T09:23:41.947051Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"padding:20px;color:white;margin:0;font-size:100%;text-align:left;display:fill;border-radius:5px;background-color:#6A79BA;overflow:hidden\">3 | Feature Engineering</div>\n\nBefore creating additional features, I will first generate the adjusted close price using the function provided by the competition hosts in the [Train Demo](https://www.kaggle.com/code/smeitoma/train-demo#Generating-AdjustedClose-price) Notebook, which will adjust the close price to account for any stock splits or reverse splits. Using the stock's adjusted close price, I will create a set of new features, including the stock price moving average, exponential moving average, return, and volatility, each over a period of 5, 10, 20, 30, and 50 days. These features are shown in the graphs below across each sector.","metadata":{}},{"cell_type":"code","source":"def adjust_price(price):\n    \"\"\"\n    Args:\n        price (pd.DataFrame)  : pd.DataFrame include stock_price\n    Returns:\n        price DataFrame (pd.DataFrame): stock_price with generated AdjustedClose\n    \"\"\"\n    # transform Date column into datetime\n    price.loc[: ,\"Date\"] = pd.to_datetime(price.loc[: ,\"Date\"], format=\"%Y-%m-%d\")\n\n    def generate_adjusted_close(df):\n        \"\"\"\n        Args:\n            df (pd.DataFrame)  : stock_price for a single SecuritiesCode\n        Returns:\n            df (pd.DataFrame): stock_price with AdjustedClose for a single SecuritiesCode\n        \"\"\"\n        # sort data to generate CumulativeAdjustmentFactor\n        df = df.sort_values(\"Date\", ascending=False)\n        # generate CumulativeAdjustmentFactor\n        df.loc[:, \"CumulativeAdjustmentFactor\"] = df[\"AdjustmentFactor\"].cumprod()\n        # generate AdjustedClose\n        df.loc[:, \"AdjustedClose\"] = (\n            df[\"CumulativeAdjustmentFactor\"] * df[\"Close\"]\n        ).map(lambda x: float(\n            Decimal(str(x)).quantize(Decimal('0.1'), rounding=ROUND_HALF_UP)\n        ))\n        # reverse order\n        df = df.sort_values(\"Date\")\n        # to fill AdjustedClose, replace 0 into np.nan\n        df.loc[df[\"AdjustedClose\"] == 0, \"AdjustedClose\"] = np.nan\n        # forward fill AdjustedClose\n        df.loc[:, \"AdjustedClose\"] = df.loc[:, \"AdjustedClose\"].ffill()\n        return df\n    \n    # generate AdjustedClose\n    price = price.sort_values([\"SecuritiesCode\", \"Date\"])\n    price = price.groupby(\"SecuritiesCode\").apply(generate_adjusted_close).reset_index(drop=True)\n    return price\n\ntrain=train.drop('ExpectedDividend',axis=1).fillna(0)\nprices=adjust_price(train)","metadata":{"execution":{"iopub.status.busy":"2022-06-27T09:23:41.948843Z","iopub.execute_input":"2022-06-27T09:23:41.949413Z","iopub.status.idle":"2022-06-27T09:24:15.229302Z","shell.execute_reply.started":"2022-06-27T09:23:41.949374Z","shell.execute_reply":"2022-06-27T09:24:15.228449Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def create_features(df):\n    df=df.copy()\n    col='AdjustedClose'\n    periods=[5,10,20,30,50]\n    for period in periods:\n        df.loc[:,\"Return_{}Day\".format(period)] = df.groupby(\"SecuritiesCode\")[col].pct_change(period)\n        df.loc[:,\"MovingAvg_{}Day\".format(period)] = df.groupby(\"SecuritiesCode\")[col].rolling(window=period).mean().values\n        df.loc[:,\"ExpMovingAvg_{}Day\".format(period)] = df.groupby(\"SecuritiesCode\")[col].ewm(span=period,adjust=False).mean().values\n        df.loc[:,\"Volatility_{}Day\".format(period)] = np.log(df[col]).groupby(df[\"SecuritiesCode\"]).diff().rolling(period).std()\n    return df\n\nprice_features=create_features(df=prices)\nprice_features.drop(['RowId','SupervisionFlag','AdjustmentFactor','CumulativeAdjustmentFactor','Close'],axis=1,inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-06-27T09:24:15.231482Z","iopub.execute_input":"2022-06-27T09:24:15.231677Z","iopub.status.idle":"2022-06-27T09:24:26.534528Z","shell.execute_reply.started":"2022-06-27T09:24:15.231651Z","shell.execute_reply":"2022-06-27T09:24:26.533777Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"price_names=price_features.merge(stock_list[['SecuritiesCode','Name','SectorName']], on='SecuritiesCode').set_index('Date')\nprice_names=price_names[price_names.index>='2020-12-29']\nprice_names.fillna(0, inplace=True)\n\nfeatures=['MovingAvg','ExpMovingAvg','Return', 'Volatility']\nnames=['Average', 'Exp. Moving Average', 'Period', 'Volatility']\nbuttons=[]\n\nfig = make_subplots(rows=2, cols=2, \n                    shared_xaxes=True, \n                    vertical_spacing=0.1,\n                    subplot_titles=('Adjusted Close Moving Average',\n                                    'Exponential Moving Average',\n                                    'Stock Return', 'Stock Volatility'))\n\nfor i, sector in enumerate(price_names.SectorName.unique()):\n    \n    sector_df=price_names[price_names.SectorName==sector]\n    periods=[0,10,30,50]\n    colors=px.colors.qualitative.Vivid\n    dash=['solid','dash', 'longdash', 'dashdot', 'longdashdot']\n    row,col=1,1\n    \n    for j, (feature, name) in enumerate(zip(features, names)):\n        if j>=2:\n            row,periods=2,[10,30,50]\n            colors=px.colors.qualitative.Bold[1:]\n        if j%2==0:\n            col=1\n        else:\n            col=2\n        \n        for k, period in enumerate(periods):\n            if (k==0)&(j<2):\n                plot_data=sector_df.groupby(sector_df.index)['AdjustedClose'].mean().rename('Adjusted Close')\n            elif j>=2:\n                plot_data=sector_df.groupby(sector_df.index)['{}_{}Day'.format(feature,period)].mean().mul(100).rename('{}-day {}'.format(period,name))\n            else:\n                plot_data=sector_df.groupby(sector_df.index)['{}_{}Day'.format(feature,period)].mean().rename('{}-day {}'.format(period,name))\n            fig.add_trace(go.Scatter(x=plot_data.index, y=plot_data, mode='lines',\n                                     name=plot_data.name, marker_color=colors[k+1],\n                                     line=dict(width=2,dash=(dash[k] if j<2 else 'solid')), \n                                     showlegend=(True if (j==0) or (j==2) else False), legendgroup=row,\n                                     visible=(False if i != 0 else True)), row=row, col=col)\n            \n    visibility=[False]*14*len(price_names.SectorName.unique())\n    for l in range(i*14, i*14+14):\n        visibility[l]=True\n    button = dict(label = sector,\n                  method = \"update\",\n                  args=[{\"visible\": visibility}])\n    buttons.append(button)\n\nfig.update_layout(title='Stock Price Moving Average, Return,<br>and Volatility by Sector',\n                  template=temp, yaxis3_ticksuffix='%', yaxis4_ticksuffix='%',\n                  legend_title_text='Period', legend_tracegroupgap=250,\n                  updatemenus=[dict(active=0, type=\"dropdown\",\n                                    buttons=buttons, xanchor='left',\n                                    yanchor='bottom', y=1.105, x=.01)], \n                  hovermode='x unified', height=800,width=1200, margin=dict(t=150))\nfig.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-06-27T09:24:26.536003Z","iopub.execute_input":"2022-06-27T09:24:26.536369Z","iopub.status.idle":"2022-06-27T09:24:35.001115Z","shell.execute_reply.started":"2022-06-27T09:24:26.536303Z","shell.execute_reply":"2022-06-27T09:24:35.000279Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In the graphs of the stock price Moving Averages (MA) and Exponential Moving Averages (EMA), when the shorter period, the 10-Day average, crosses above the longer period, the 50-day average, the closing price tends to decrease, which is typically indicative of a sell signal. Conversely, when the 10-Day MA/EMA crosses the 50-Day MA/EMA from below, the closing price increases, which typically indicates a buy signal. In addition, the Exponential Moving Averages tend to respond to stock price changes more quickly as it puts greater weight on more recent observations. In the graphs of the Stock Return, there is greater fluctuation between longer periods, while Stock Volatility tends to be more stable over time, with more fluctuation in shorter periods.\n\n# <div style=\"padding:20px;color:white;margin:0;font-size:100%;text-align:left;display:fill;border-radius:5px;background-color:#6A79BA;overflow:hidden\">4 | Stock Price Prediction</div>\n\nSubmissions for this competition are evaluated based on the [Sharpe Ratio](https://en.wikipedia.org/wiki/Sharpe_ratio) of the daily spread of returns. For each forecast day, all active stocks will be ranked in order of their predicted return. The returns for a single day treat the 200 highest (e.g. 0 to 199) ranked stocks as purchased and the lowest (e.g. 1800 to 1999) ranked 200 stocks as shorted. The stocks are then weighted based on their ranks and the total returns for the portfolio are calculated assuming the stocks were purchased the next day and sold the day after that. \n\nSince risk control is also an important element of investment, the competing score is the mean/standard deviation of the time series of daily spread returns, rather than the simple mean or sum of daily spread returns. This makes it necessary to build a model that can respond to changes in the distribution of data and produce stable and high performance, rather than a model that only wins big on certain days. You can find a python implementation of this metric [here](https://www.kaggle.com/code/smeitoma/jpx-competition-metric-definition). ","metadata":{}},{"cell_type":"code","source":"def calc_spread_return_sharpe(df: pd.DataFrame, portfolio_size: int = 200, toprank_weight_ratio: float = 2) -> float:\n    \"\"\"\n    Args:\n        df (pd.DataFrame): predicted results\n        portfolio_size (int): # of equities to buy/sell\n        toprank_weight_ratio (float): the relative weight of the most highly ranked stock compared to the least.\n    Returns:\n        (float): sharpe ratio\n    \"\"\"\n    def _calc_spread_return_per_day(df, portfolio_size, toprank_weight_ratio):\n        \"\"\"\n        Args:\n            df (pd.DataFrame): predicted results\n            portfolio_size (int): # of equities to buy/sell\n            toprank_weight_ratio (float): the relative weight of the most highly ranked stock compared to the least.\n        Returns:\n            (float): spread return\n        \"\"\"\n        assert df['Rank'].min() == 0\n        assert df['Rank'].max() == len(df['Rank']) - 1\n        weights = np.linspace(start=toprank_weight_ratio, stop=1, num=portfolio_size)\n        purchase = (df.sort_values(by='Rank')['Target'][:portfolio_size] * weights).sum() / weights.mean()\n        short = (df.sort_values(by='Rank', ascending=False)['Target'][:portfolio_size] * weights).sum() / weights.mean()\n        return purchase - short\n\n    buf = df.groupby('Date').apply(_calc_spread_return_per_day, portfolio_size, toprank_weight_ratio)\n    sharpe_ratio = buf.mean() / buf.std()\n    return sharpe_ratio","metadata":{"execution":{"iopub.status.busy":"2022-06-27T09:24:35.0025Z","iopub.execute_input":"2022-06-27T09:24:35.002901Z","iopub.status.idle":"2022-06-27T09:24:35.019184Z","shell.execute_reply.started":"2022-06-27T09:24:35.002864Z","shell.execute_reply":"2022-06-27T09:24:35.01832Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ts_fold = TimeSeriesSplit(n_splits=10, gap=10000)\nprices=price_features.dropna().sort_values(['Date','SecuritiesCode'])\ny=prices['Target'].to_numpy()\nX=prices.drop(['Target'],axis=1)\n\nfeat_importance=pd.DataFrame()\nsharpe_ratio=[]\n    \nfor fold, (train_idx, val_idx) in enumerate(ts_fold.split(X, y)):\n    \n    print(\"\\n========================== Fold {} ==========================\".format(fold+1))\n    X_train, y_train = X.iloc[train_idx,:], y[train_idx]\n    X_valid, y_val = X.iloc[val_idx,:], y[val_idx]\n    \n    print(\"Train Date range: {} to {}\".format(X_train.Date.min(),X_train.Date.max()))\n    print(\"Valid Date range: {} to {}\".format(X_valid.Date.min(),X_valid.Date.max()))\n    \n    X_train.drop(['Date','SecuritiesCode'], axis=1, inplace=True)\n    X_val=X_valid[X_valid.columns[~X_valid.columns.isin(['Date','SecuritiesCode'])]]\n    val_dates=X_valid.Date.unique()[1:-1]\n    print(\"\\nTrain Shape: {} {}, Valid Shape: {} {}\".format(X_train.shape, y_train.shape, X_val.shape, y_val.shape))\n    \n    params = {'n_estimators': 1500,\n               'num_leaves' : 100,\n              'learning_rate': 0.1,\n              'colsample_bytree': 0.9,\n              'subsample': 0.8,\n              'reg_alpha': 0.4,\n              'metric': 'mae',\n              'random_state': 21}\n    \n    gbm = LGBMRegressor(**params).fit(X_train, y_train, \n                                      eval_set=[(X_train, y_train), (X_val, y_val)],\n                                      verbose=300, \n                                      eval_metric=['mae','mse'])\n    y_pred = gbm.predict(X_val)\n    rmse = np.sqrt(mean_squared_error(y_val, y_pred))\n    mae = mean_absolute_error(y_val, y_pred)\n    feat_importance[\"Importance_Fold\"+str(fold)]=gbm.feature_importances_\n    feat_importance.set_index(X_train.columns, inplace=True)\n    \n    rank=[]\n    X_val_df=X_valid[X_valid.Date.isin(val_dates)]\n    for i in X_val_df.Date.unique():\n        temp_df = X_val_df[X_val_df.Date == i].drop(['Date','SecuritiesCode'],axis=1)\n        temp_df[\"pred\"] = gbm.predict(temp_df)\n        temp_df[\"Rank\"] = (temp_df[\"pred\"].rank(method=\"first\", ascending=False)-1).astype(int)\n        rank.append(temp_df[\"Rank\"].values)\n\n    stock_rank=pd.Series([x for y in rank for x in y], name=\"Rank\")\n    df=pd.concat([X_val_df.reset_index(drop=True),stock_rank,\n                  prices[prices.Date.isin(val_dates)]['Target'].reset_index(drop=True)], axis=1)\n    sharpe=calc_spread_return_sharpe(df)\n    sharpe_ratio.append(sharpe)\n    print(\"Valid Sharpe: {}, RMSE: {}, MAE: {}\".format(sharpe,rmse,mae))\n    \n    del X_train, y_train,  X_val, y_val\n    gc.collect()\n    \nprint(\"\\nAverage cross-validation Sharpe Ratio: {:.4f}, standard deviation = {:.2f}.\".format(np.mean(sharpe_ratio),np.std(sharpe_ratio)))","metadata":{"execution":{"iopub.status.busy":"2022-06-27T09:24:35.020975Z","iopub.execute_input":"2022-06-27T09:24:35.021354Z","iopub.status.idle":"2022-06-27T09:26:34.122158Z","shell.execute_reply.started":"2022-06-27T09:24:35.021316Z","shell.execute_reply":"2022-06-27T09:26:34.119642Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Overall, the models have a cross-validation Sharpe Ratio of about 0.129 with a standard deviation of 0.13. Below are the most important features in the models, averaged across each fold.","metadata":{}},{"cell_type":"code","source":"feat_importance['avg'] = feat_importance.mean(axis=1)\nfeat_importance = feat_importance.sort_values(by='avg',ascending=True)\npal=sns.color_palette(\"plasma_r\", 29).as_hex()[2:]\n\nfig=go.Figure()\nfor i in range(len(feat_importance.index)):\n    fig.add_shape(dict(type=\"line\", y0=i, y1=i, x0=0, x1=feat_importance['avg'][i], \n                       line_color=pal[::-1][i],opacity=0.7,line_width=4))\nfig.add_trace(go.Scatter(x=feat_importance['avg'], y=feat_importance.index, mode='markers', \n                         marker_color=pal[::-1], marker_size=8,\n                         hovertemplate='%{y} Importance = %{x:.0f}<extra></extra>'))\nfig.update_layout(template=temp,title='Overall Feature Importance', \n                  xaxis=dict(title='Average Importance',zeroline=False),\n                  yaxis_showgrid=False, margin=dict(l=120,t=80),\n                  height=700, width=800)\nfig.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-06-27T09:26:34.124829Z","iopub.status.idle":"2022-06-27T09:26:34.126871Z","shell.execute_reply.started":"2022-06-27T09:26:34.126633Z","shell.execute_reply":"2022-06-27T09:26:34.126661Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cols_fin=feat_importance.avg.nlargest(3).index.tolist()\ncols_fin.extend(('Open','High','Low'))\nX_train=prices[cols_fin]\ny_train=prices['Target']\ngbm = LGBMRegressor(**params).fit(X_train, y_train)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-06-27T09:26:34.130355Z","iopub.status.idle":"2022-06-27T09:26:34.132389Z","shell.execute_reply.started":"2022-06-27T09:26:34.132131Z","shell.execute_reply":"2022-06-27T09:26:34.132158Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <div style=\"padding:20px;color:white;margin:0;font-size:100%;text-align:left;display:fill;border-radius:5px;background-color:#6A79BA;overflow:hidden\">5 | API Submission</div>","metadata":{}},{"cell_type":"code","source":"import jpx_tokyo_market_prediction\nenv = jpx_tokyo_market_prediction.make_env()\niter_test = env.iter_test()\n\ncols=['Date','SecuritiesCode','Open','High','Low','Close','Volume','AdjustmentFactor']\ntrain=train[train.Date>='2021-08-01'][cols]\n\ncounter = 0\nfor (prices, options, financials, trades, secondary_prices, sample_prediction) in iter_test:\n\n    current_date = prices[\"Date\"].iloc[0]\n    if counter == 0:\n        df_price_raw = train.loc[train[\"Date\"] < current_date]\n    df_price_raw = pd.concat([df_price_raw, prices[cols]]).reset_index(drop=True)\n    df_price = adjust_price(df_price_raw)\n    features = create_features(df=df_price)\n    feat = features[features.Date == current_date][cols_fin]\n    feat[\"pred\"] = gbm.predict(feat)\n    feat[\"Rank\"] = (feat[\"pred\"].rank(method=\"first\", ascending=False)-1).astype(int)\n    sample_prediction[\"Rank\"] = feat[\"Rank\"].values\n    display(sample_prediction.head())\n    \n    assert sample_prediction[\"Rank\"].notna().all()\n    assert sample_prediction[\"Rank\"].min() == 0\n    assert sample_prediction[\"Rank\"].max() == len(sample_prediction[\"Rank\"]) - 1\n    \n    env.predict(sample_prediction)\n    counter += 1","metadata":{"execution":{"iopub.status.busy":"2022-06-27T09:26:34.134994Z","iopub.status.idle":"2022-06-27T09:26:34.135604Z","shell.execute_reply.started":"2022-06-27T09:26:34.135381Z","shell.execute_reply":"2022-06-27T09:26:34.135406Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <div style=\"padding:20px;color:#6A79BA;margin:0;font-size:100%;text-align:center;display:fill;border-radius:5px;background-color:white;overflow:hidden\">Thank you for reading!<br>Please let me know if you have any feedback. 🙂</div>\n### <span style=\"padding:20px;color:#6A79BA;margin:0;font-size:100%;text-align:center;display:fill;border-radius:5px;background-color:white;overflow:hidden\">Notebook References</span>\n- https://www.kaggle.com/code/smeitoma/train-demo\n- https://www.kaggle.com/code/smeitoma/submission-demo\n- https://www.kaggle.com/code/smeitoma/jpx-competition-metric-definition\n- https://www.kaggle.com/code/chumajin/easy-to-understand-the-competition\n- https://www.kaggle.com/code/riteshsinha/useful-features-in-predicting-stock-prices","metadata":{}}]}