{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.11.9"}},"nbformat_minor":4,"nbformat":4,"cells":[{"id":"7ce67172","cell_type":"markdown","source":"<h1 style=\"color:#000000;font-weight:700;font-family:Arial,Helvetica,sans-serif;font-size:30px\">H&M Recommendation: From Multi-Recall &amp; LambdaRank to Generative Semantic Signals</h1>\n\n| Evaluation | Scope | Result |\n| --- | --- | --- |\n| Kaggle 官方全量评分 | 全品类、1,371,980位客户；Multi-Recall + LambdaRank | Public MAP@12 **0.02466**；Private **0.02443** |\n| 生成式机制受控实验 | 20k 裤装用户，开发窗口中1,150位购买用户 | Joint MAP@12 **0.087891**；相对行为基线 **+4.19%** |\n\n全量 Public 分数处于历史榜分数的 **Top 10%** 区间，不是正式比赛排名。全量提交未使用生成模块；裤装子集用于机制分析，其 MAP@12 与官方全量评分不可直接比较。\n\n<h3 style=\"color:#000000;font-weight:700;font-family:Arial,Helvetica,sans-serif\">Contributions</h3>\n\n1. 独立构建 **Multi-Recall + Feature Engineering + LightGBM LambdaRank** 两阶段系统，完成时间隔离、动态 catalog 与全客户推理。\n2. 构建 **Semantic ID + lightweight T5**，拆分为 **Candidate Generation / Semantic Scoring** 两个可独立检验的接口。\n3. 设计 **Score-only / Candidate-only / Joint** 消融，结合 shuffle control、candidate coverage 和 paired bootstrap 分析增益来源。\n\n**运行说明**：Run All 使用所附 Report Evidence Input 重算子集表格与 Matplotlib 图，不启动训练或提交。","metadata":{}},{"id":"5b78dc6d","cell_type":"markdown","source":"<h2 id=\"section1\" style=\"color:#000000;font-weight:700;font-family:Arial,Helvetica,sans-serif\">1. Task & Data</h2>\n\n任务是在 cutoff 时刻预测用户下一周购买的商品集合，输出 Top12。受控实验限定 H&M 裤装：根据预先设定的历史区间选取20,000名用户，在开发窗口评估其中下一周有购买的1,150人；每人保留最近20次购买。\n\nMAP@12 的分母按完整真实购买集合计算，候选未覆盖的目标仍保留。candidate recall 衡量排序前的目标覆盖，与最终排序指标分开报告。","metadata":{}},{"id":"e53c180a","cell_type":"code","source":"import base64, io, json\nimport numpy as np\nimport pandas as pd\nfrom pathlib import Path\n\n# Kaggle Input contains the frozen report assets; local runs use report/data.\nasset_roots = [Path.cwd() / 'report' / 'data']\nif Path('/kaggle/input').exists():\n    asset_roots += [p.parent for p in Path('/kaggle/input').rglob('notebook_examples.json')]\nasset_roots = [p for p in asset_roots if all((p / name).is_file() for name in\n    ['evidence.npz', 'notebook_examples.json', 'product_images.json', 'figure_data.json'])]\nif len(asset_roots) != 1:\n    raise RuntimeError('Attach the report evidence Input, or run locally with report/data.')\nasset_dir = asset_roots[0]\nevidence = np.load(asset_dir / 'evidence.npz')\nexamples = json.loads((asset_dir / 'notebook_examples.json').read_text(encoding='utf-8'))\nproduct_images = json.loads((asset_dir / 'product_images.json').read_text(encoding='utf-8'))\nfigure_data = json.loads((asset_dir / 'figure_data.json').read_text(encoding='utf-8'))\nassert evidence['baseline'].shape == (1150, 2)\nprint('1,150 users; two illustrative cases; 23 product images loaded.')","metadata":{"collapsed":true,"jupyter":{"source_hidden":true,"outputs_hidden":true},"tags":["hide_input"]},"outputs":[],"execution_count":null},{"id":"ab8d6185","cell_type":"code","source":"from matplotlib import pyplot as plt\nfrom matplotlib import image as mpimg\nimport textwrap\n\nBLUE, ORANGE, TEAL = '#023e8a', '#ffa500', '#5f9ea0'\nplt.rcParams.update({\n    'font.family': 'DejaVu Sans', 'font.size': 11,\n    'axes.facecolor': '#EAEAF2', 'axes.edgecolor': 'white',\n    'axes.grid': True, 'grid.color': 'white', 'grid.linewidth': 1,\n    'axes.titlecolor': 'black', 'axes.titleweight': 'bold', 'axes.spines.top': False,\n    'axes.spines.right': False, 'svg.fonttype': 'none', 'pdf.fonttype': 42,\n})\nfigure_dir = Path.cwd() / 'report' / 'figures'\nfigure_dir.mkdir(parents=True, exist_ok=True)\n\ndef show(fig, name):\n    # Save exactly the same figure that is displayed below the cell.\n    for extension in ['png', 'svg', 'pdf']:\n        fig.savefig(figure_dir / f'{name}.{extension}', dpi=180, bbox_inches='tight')\n    plt.show()","metadata":{},"outputs":[],"execution_count":null},{"id":"5f2eb349","cell_type":"code","source":"lengths = examples['population']['history_lengths']\nfig, ax = plt.subplots(figsize=(10, 3.5))\nax.hist(lengths, bins=np.arange(0.5, 21.5), color=ORANGE, edgecolor='white')\nax.set(xlabel='Retained history length (capped at 20)', ylabel='Users',\n       title='How much history do we keep?', xticks=[1, 5, 10, 15, 20])\nshow(fig, '10_history')","metadata":{},"outputs":[],"execution_count":null},{"id":"f9b449a4","cell_type":"code","source":"targets = np.asarray(examples['population']['target_counts'])\ncounts = pd.Series(np.minimum(targets, 8)).value_counts().sort_index()\nfig, ax = plt.subplots(figsize=(10, 3.5))\nax.bar(counts.index, counts.values, color=ORANGE, edgecolor='white')\nax.set(xticks=range(1, 9), xticklabels=['1','2','3','4','5','6','7','8+'],\n       xlabel='Distinct items purchased next week', ylabel='Users',\n       title='One user can have several correct recommendations')\nshow(fig, '11_targets')","metadata":{},"outputs":[],"execution_count":null},{"id":"46697cbf","cell_type":"markdown","source":"<h2 id=\"section2\" style=\"color:#000000;font-weight:700;font-family:Arial,Helvetica,sans-serif\">2. Two-stage Baseline</h2>\n\n我们实现多路 Recall，将复购、同购与热门商品合并为 Top100；Ranking 使用用户—商品特征，从候选内部选出 Top12。LambdaRank 学习同一用户候选的相对次序，适合 Top-K 排序目标，不将输出分数解释为购买概率。\n\n| Feature group | Features in the controlled baseline |\n| --- | --- |\n| Popularity | 7天、14天及历史累计购买计数（log1p） |\n| User-Item | 用户购买该商品的次数（log1p） |\n| Recency | 用户最近购买该商品、商品最近交易距 cutoff 的天数（截断至365天） |\n| Co-occurrence | 近期同日共购计数（log1p） |\n| Recall Prior | 去重后候选位置 / 100 |\n| User Context | 保留的历史长度，上限20 |\n\n动态 catalog、时间隔离与 candidate recall 共同构成评估协议：目录随 cutoff 更新，候选和特征仅使用此前交易，未来目标不补入候选。","metadata":{}},{"id":"a98e569d","cell_type":"code","source":"pd.DataFrame({\n    'Stage': ['Recall', 'Rank', 'Evaluate'],\n    'Input / rule': ['Repeat + co-purchase + popularity',\n                     '9 behavioral features; LightGBM LambdaRank',\n                     'Top12 against the full next-week purchase set'],\n    'Size': ['100 candidates', '150 trees; 15 leaves', '1,150 purchasing users'],\n})","metadata":{},"outputs":[],"execution_count":null},{"id":"ceb71adc","cell_type":"markdown","source":"<h2 id=\"section3\" style=\"color:#000000;font-weight:700;font-family:Arial,Helvetica,sans-serif\">3. Generative Semantic Extension</h2>\n\n我们构建低成本、可审计的 Semantic-ID 模块，评估 generative signal 对既有 cascade recommender 的增量。商品名、描述和颜色经 TF-IDF/SVD、两级残差 KMeans 及唯一序号编码；lightweight T5 根据近期商品码预测下一购买日的商品码。\n\n**同一个生成模型拆为两个接口：**\n\n| Interface | Integration |\n| --- | --- |\n| Candidate Generation | 生成商品码并映射回商品，保留原80件候选，最多替换20个位置 |\n| Semantic Scoring | 计算候选完整商品码的 log probability，以用户内中心化分数与相对排名加入 LambdaRank |\n\nT5 随机初始化，隐藏维64，2层编码器、1层解码器；15,887组历史样本、4轮、两个种子。生成器训练和选模完成后再训练排序器。","metadata":{}},{"id":"45722d92","cell_type":"markdown","source":"<h2 id=\"section4\" style=\"color:#000000;font-weight:700;font-family:Arial,Helvetica,sans-serif\">4. Controlled Experiments</h2>\n\n<h3 style=\"color:#000000;font-weight:700;font-family:Arial,Helvetica,sans-serif\">Controlled Performance</h3>\n\n我们固定树参数，分别接入两个接口及其组合。Semantic Scoring 的 log probability 未经购买概率校准；Candidate-only 和 Joint 还会改变部分排序训练样本。\n\n每位用户的指标先在两个种子间取平均，再对用户取平均，不做模型集成。95%区间采用按用户配对的 paired bootstrap，抽样5,000次。","metadata":{}},{"id":"14db36c1","cell_type":"code","source":"pd.DataFrame({\n    'Recipe': ['Baseline', 'Score-only', 'Candidate-only', 'Joint'],\n    'Candidates': ['Original 100', 'Original 100', 'Keep 80 + up to 20 generated', 'Same mixed 100'],\n    'Ranker features': ['Behavior', 'Behavior + generation score', 'Behavior', 'Behavior + generation score'],\n})","metadata":{},"outputs":[],"execution_count":null},{"id":"f6cd5f62","cell_type":"code","source":"baseline = evidence['baseline'][:, 0]\nrecipes = ['score_only', 'candidate_only', 'candidate_score']\nnames = ['Baseline', 'Score-only', 'Candidate-only', 'Joint']\nmeans = {r: np.mean([evidence[f'{r}_{s}'] for s in [42, 2026]], axis=0) for r in recipes}\nvalues = np.vstack([evidence['baseline'].mean(axis=0)] + [means[r].mean(axis=0) for r in recipes])\npd.DataFrame(values, index=names, columns=['MAP@12', 'Candidate Recall@100']).round(6)","metadata":{},"outputs":[],"execution_count":null},{"id":"931f65c3","cell_type":"code","source":"fig, ax = plt.subplots(figsize=(10, 4))\nax.barh(names, values[:, 0], color=['#9DA3AE', ORANGE, ORANGE, ORANGE])\nfor i, value in enumerate(values[:, 0]):\n    ax.text(value + .001, i, f'{value:.5f}', va='center')\nax.set(xlim=(0, .105), xlabel='Development MAP@12', title='Development MAP@12 under controlled ablations')\nax.invert_yaxis()\nshow(fig, '12_results')","metadata":{},"outputs":[],"execution_count":null},{"id":"329eb6d1","cell_type":"code","source":"def paired_interval(delta):\n    rng = np.random.default_rng(123)\n    draws = [rng.choice(delta, len(delta), replace=True).mean() for _ in range(5000)]\n    return np.quantile(draws, [.025, .975])\n\nfig, ax = plt.subplots(figsize=(10, 3.5))\nfor i, recipe in enumerate(recipes):\n    delta = means[recipe][:, 0] - baseline\n    low, high = paired_interval(delta)\n    ax.plot([low, high], [i, i], color=ORANGE, linewidth=3)\n    ax.scatter(delta.mean(), i, color=BLUE, zorder=3)\nax.axvline(0, color=TEAL, linestyle='--')\nax.set(yticks=range(3), yticklabels=names[1:], xlabel='Paired change in MAP@12',\n       title='95% paired user-bootstrap intervals')\nax.invert_yaxis()\nshow(fig, '13_intervals')","metadata":{},"outputs":[],"execution_count":null},{"id":"19658d17","cell_type":"markdown","source":"Joint 开发集 MAP@12 为 **0.087891**，行为基线为 **0.084353**，相对提升 **4.19%**。Joint 对基线的区间下界略高于零；相对传统共现替换和 Top200 控制，差值区间均跨零。该窗口用于开发，区间未做多重比较修正。","metadata":{}},{"id":"0274a318","cell_type":"markdown","source":"<h2 id=\"section5\" style=\"color:#000000;font-weight:700;font-family:Arial,Helvetica,sans-serif\">5. Mechanism Analysis</h2>\n\n","metadata":{}},{"id":"b4a06699","cell_type":"code","source":"fig, ax = plt.subplots(figsize=(10, 3.5))\nax.bar(names, values[:, 1], color=[TEAL, TEAL, ORANGE, ORANGE], width=.6)\nfor i, value in enumerate(values[:, 1]):\n    ax.text(i, value + .008, f'{value:.3f}', ha='center')\nax.set(ylim=(0, .5), ylabel='Candidate Recall@100', title='Joint improvement is not explained by higher candidate coverage')\nshow(fig, '14_coverage')","metadata":{},"outputs":[],"execution_count":null},{"id":"d160b211","cell_type":"code","source":"fig, ax = plt.subplots(figsize=(10, 4))\nlabels = ['Targets gained', 'Targets lost', 'New correct Top12 hits']\nfor offset, row, color in zip([-.18, .18], figure_data['candidate_changes'], [ORANGE, TEAL]):\n    counts = [row['gained_target_pairs'], row['lost_target_pairs'],\n              row['correct_final_recommendations_outside_original100']]\n    label = 'Seed ' + row['model'].split('_')[1]\n    x = np.arange(3) + offset\n    ax.bar(x, counts, width=.34, color=color, label=label)\n    for xx, count in zip(x, counts):\n        ax.text(xx, count + 1.5, str(count), ha='center')\nax.set(xticks=range(3), xticklabels=labels, ylabel='User-item pairs', ylim=(0, 76),\n       title='Generated candidates contribute no additional correct Top-12 hits')\nax.legend()\nshow(fig, '15_newhits')","metadata":{},"outputs":[],"execution_count":null},{"id":"877269cb","cell_type":"markdown","source":"Joint 的 candidate recall 下降，原 Top100 外新增生成候选没有贡献新的正确 Top12 命中。当前结果更支持**已有候选空间中的排序侧增量**，而非成功的 generative retrieval；尚未完全分离 semantic score 本身与候选变化引起的训练分布变化。\n\n**Shuffle controls.** 混合候选内打乱生成分数后，MAP@12 为0.084187；真实分数与打乱分数的差值区间仍跨零。固定原 Top100 的补充验证中，真实历史分数方案为0.08180，基线为0.08072，差值区间跨零；真实历史顺序相对按购买日块打乱也未显示优势。","metadata":{}},{"id":"683efe0d","cell_type":"markdown","source":"<h3 style=\"color:#000000;font-weight:700;font-family:Arial,Helvetica,sans-serif\">Recommendation example</h3>\n\n固定种子42，在 Joint 有提升的用户中按匿名标识取第一位。展示最近4件不同商品、两组推荐的前4件及真实购买；AP 按完整 Top12 计算。案例用于展示重排行为，不作为总体性能证据。商品名称、颜色与照片来自竞赛数据。","metadata":{}},{"id":"ee33dbdd","cell_type":"code","source":"def show_case(case, filename):\n    fig, axes = plt.subplots(4, 4, figsize=(12, 15))\n    for row_index, (group, item_ids) in enumerate(case['groups'].items()):\n        for column in range(4):\n            ax = axes[row_index, column]\n            ax.axis('off')\n            if column == 0:\n                ax.text(-.08, .5, group, transform=ax.transAxes, rotation=90,\n                        ha='right', va='center', color='black', fontsize=12, weight='bold')\n            if column >= len(item_ids):\n                continue\n            article = item_ids[column]\n            image = mpimg.imread(io.BytesIO(base64.b64decode(product_images[article])), format='jpg')\n            ax.imshow(image)\n            item = examples['items'][article]\n            title = textwrap.fill(item['name'], 23) + '\\n' + item['color']\n            ax.set_title(title, fontsize=10, color='black', fontweight='bold', pad=6)\n            note = article\n            if group.startswith(('Baseline', 'Joint')) and article in case['targets']:\n                note += '  |  PURCHASED'\n            if group == 'Next-week purchases':\n                b = case['baseline_top12']; j = case['joint_top12']\n                note += f\"\\nBaseline rank: {b.index(article)+1 if article in b else '—'}; Joint rank: {j.index(article)+1 if article in j else '—'}\"\n            ax.text(.5, -.025, note, transform=ax.transAxes, ha='center', va='top', fontsize=8)\n    fig.suptitle(f\"{case['label']}  |  AP@12: baseline {case['baseline_ap']:.3f} → Joint {case['joint_ap']:.3f}\",\n                 color='black', fontweight='bold', fontsize=16, y=.995)\n    fig.subplots_adjust(hspace=.65, wspace=.3, top=.94, bottom=.035)\n    show(fig, filename)\n\nshow_case(examples['cases'][0], '16_case_improved')","metadata":{},"outputs":[],"execution_count":null},{"id":"9beb5d24","cell_type":"markdown","source":"命中的 **0783346001** 原本就在 Top100 中，经 Joint 重排至第2位，AP@12=0.125；基线 Top12 未命中。","metadata":{}},{"id":"645d6080","cell_type":"markdown","source":"<h2 id=\"section6\" style=\"color:#000000;font-weight:700;font-family:Arial,Helvetica,sans-serif\">6. Key Findings</h2>\n\n1. **行为召回 + LambdaRank 是系统基础。** 全量行为系统获得 Public 0.02466 / Private 0.02443；历史分数处于 Top 10% 区间。\n2. **生成语义信号在开发实验中表现出排序侧增量。** Joint 改善最终排序，现有对照尚未唯一确定增益来源。\n3. **新生成候选和历史顺序尚无稳定增量证据。** 新候选没有新增正确 Top12，顺序打乱对照未显示真实顺序优势。","metadata":{}},{"id":"3b5a3609","cell_type":"markdown","source":"<h2 id=\"section7\" style=\"color:#000000;font-weight:700;font-family:Arial,Helvetica,sans-serif\">7. Limitations & Reproducibility</h2>\n\n<h3 style=\"color:#000000;font-weight:700;font-family:Arial,Helvetica,sans-serif\">Limitations</h3>\n\n受控实验与补充验证均使用参与过开发的窗口。补充验证覆盖 Score-only，Joint 的跨周稳定性不在本次验证范围内；现有对照尚未完全分离生成分数效应与候选、训练组变化。\n\n<h3 style=\"color:#000000;font-weight:700;font-family:Arial,Helvetica,sans-serif\">Reproducibility</h3>\n\n子集实验的七个排序模型已在 CPU 上精确复现；清理后的完整生成模型 GPU 训练流程未重跑。全量行为基线已单独完成训练与 Kaggle 评分。实例照片用于本地报告。\n\n<h3 style=\"color:#000000;font-weight:700;font-family:Arial,Helvetica,sans-serif\">Related Work & References</h3>\n\n- 本项目独立实现上述系统。[H-M-Fashion-RecSys](https://github.com/Wp-Zhang/H-M-Fashion-RecSys) 用于核对 H&M 中常见的 multi-recall + learning-to-rank 范式。\n- [TIGER](https://arxiv.org/abs/2305.05065)、[LETTER](https://arxiv.org/abs/2405.07314) 将商品编码与生成式推荐联系起来。本项目使用显式内容编码与轻量 T5，重点检验生成候选和评分两个接口在两阶段系统中的增量，不是对上述模型的完整复现。\n- [OneRec（2025）](https://arxiv.org/abs/2502.18965) 探索用生成模型统一 retrieve + rank；[ActionPiece（ICML 2025）](https://proceedings.mlr.press/v267/hou25f.html) 研究上下文相关的 action tokenization。本项目选择在传统 cascade 中隔离研究 semantic signal 的边际价值。","metadata":{}},{"id":"8be0569d","cell_type":"markdown","source":"<h2 style=\"color:#000000;font-weight:700;font-family:Arial,Helvetica,sans-serif\">Appendix · Additional recommendation example</h2>\n\n固定种子42，在 Joint 无提升的用户组中按匿名标识取第一位。该用户下一周购买2件商品，两套 Top12 均未命中；选择规则与正文一致。","metadata":{}},{"id":"644d9e82","cell_type":"code","source":"show_case(examples['cases'][1], '17_case_unchanged')","metadata":{},"outputs":[],"execution_count":null}]}