⏱ ~69 min

知识库分析报告(analytics)

自动生成:python3 scripts/analyze_kb.py。基于 papers.jsonl(975 篇)+ papers_pdf/(746 篇本地全文)+ papers/terse/P-*(43 篇深笔记)。
本报告是知识库「分析层」的体检,用于找:关键影响力工作、元数据缺口、引用图断链、交叉学习富矿、待补遗漏。

1. 概览

学习仪表盘

面向学习者的入口:每条线索的浅入口密度、推荐起点、领域图谱摘要、进阶学习路径。数据来自 site/content/*/.md(浅入口)+ clusters.jsonl / bridges.jsonl / influence.jsonl(图谱摘要)。

浅入口覆盖(按线索)

已发布入门教学页合计 294 篇,分布在 14 条线索下。入门页=一句话直觉 + 五分钟读懂,是最易切入的浅入口;该数高 = 这条线索的学习素材更密。

浅入口覆盖条形图

线索浅入口入口
T0118T01
T0222T02
T036T03
T0434T04
T0523T05
T0665T06
T0716T07
T0820T08
T0916T09
T1010T10
T1136T11
T1215T12
T136T13
T147T14

推荐起点(浅入口覆盖最高的 3 条线索)

知识图谱摘要

5 个研究群落(label-propagation + 模块度合并;覆盖库内 754 篇):

群落规模主导方向主题
C0265Foundation视觉-语言-动作模型(VLA)
C1202Motion Prior人形机器人 RL / 扩散模型 / 动作跟踪
C2134Foundation世界模型 / 极限地形/parkour / locomotion
C387Action Gen动作重定向/映射 / 扩散模型
C466Dexterous灵巧手操作

桥接工作 Top-5(邻居跨 ≥2 群落;最值得深挖的交叉节点):

工作桥接度span方向
Proximal Policy Optimization Algorithms (PPO)4.94Locomotion
RMA: Rapid Motor Adaptation for Legged Robots4.54Locomotion
Isaac Gym: High Performance GPU-Based Physics Simulation For4.34Locomotion
AMASS: Archive of Motion Capture as Surface Shapes4.03Motion Prior
GR00T N1: An Open Foundation Model for Generalist Humanoid R3.93Foundation

影响力论文 Top-5(按 PageRank):

工作方向in-degPageRank
AMP: Adversarial Motion Priors for Stylized Physics-Based ChAction Gen170.0099
Extreme Parkour with Legged RobotsLocomotion130.0065
BeyondMimic: From Motion Tracking to Versatile Humanoid ContAction Gen140.0064
ASE: Large-Scale Reusable Adversarial Skill Embeddings for PAction Gen150.0059
ThriftyDAgger: Budget-aware novelty and risk gating for inteImitation30.0055

领域分布速览

论文按方向 / 按年份的分布(与 §6 覆盖度同源数据,这里以视觉形式给出学习者一个直觉)。

论文按方向分布环形图

论文按年份分布条形图

学习路径入口

2. 元数据完整度

字段key填充占比
标题title975/975100%
作者authors689/97571%
机构org655/97567%
年份year975/975100%
venuevenue967/97599%
arxiv idarxiv787/97581%
代码code225/97523%
项目页project204/97521%
开源open_source975/975100%
方向category975/975100%
子方向subfield670/97569%
中文贡献contribution_zh975/975100%
⭐为何重要why_important_zh967/97599%
builds_onbuilds_on948/97597%
leads_toleads_to633/97565%
输入模态input_modality254/97526%
模型规模model_size42/9754%
上真机on_robot282/97529%
物理感知physics_aware282/97529%
实时real_time198/97520%
机器人robot184/97519%
数据集dataset181/97519%

薄弱字段(<50%,按优先级)

3. 影响力工作榜

3.1 被「builds_on」最多(影响力 Top 20)

次数 = 多少篇后续工作显式声明建立在该工作之上。

被引次数工作方向深笔记本地PDF
60OpenVLA: An Open-Source Vision-Language-Action ModelFoundation
39Diffusion Policy: Visuomotor Policy Learning via Action DiffusionFoundation
37Open X-Embodiment: Robotic Learning Datasets and RT-X ModelsFoundation
36π0: A Vision-Language-Action Flow Model for General Robot ControlFoundation
35RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic ControlFoundation
30AMP: Adversarial Motion Priors for Stylized Physics-Based Character ControlAction Gen
30OmniH2O: Universal and Dexterous Human-to-Humanoid Whole-Body Teleoperation and LearningAction Gen
27DeepMimic: Example-Guided Deep Reinforcement Learning of Physics-Based Character SkillsLocomotion
27HOVER: Versatile Neural Whole-Body Controller for Humanoid RobotsLocomotion
24Learning Transferable Visual Models From Natural Language Supervision (CLIP)Foundation
24Expressive Whole-Body Control for Humanoid Robots (ExBody)Action Gen
23Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (ACT / ALOHA)Foundation
21Perpetual Humanoid Control for Real-time Simulated AvatarsAction Gen
20Attention Is All You Need (Transformer)Foundation
19Human Motion Diffusion Model (MDM)Action Gen
17Auto-Encoding Variational Bayes (VAE)Foundation
16World ModelsFoundation
15Proximal Policy Optimization Algorithms (PPO)Locomotion
15AMASS: Archive of Motion Capture as Surface ShapesMotion Prior
14Denoising Diffusion Probabilistic Models (DDPM)Foundation

3.2 「leads_to」最多(催生后续工作最多 Top 12)

催生数工作方向
7AMP: Adversarial Motion Priors for Stylized Physics-Based Character ControlAction Gen
6Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (ACT / ALOHA)Foundation
5Generative Adversarial Nets (GAN)Foundation
5SAPIEN: A SimulAted Part-based Interactive ENvironmentFoundation
5Denoising Diffusion Probabilistic Models (DDPM)Foundation
5Learning to Walk in Minutes Using Massively Parallel Deep Reinforcement LearningLocomotion
5Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online VideosImitation
4Animation of Dynamic Legged LocomotionMotion Prior
4Auto-Encoding Variational Bayes (VAE)Foundation
4SMPL: A Skinned Multi-person Linear ModelMotion Prior
4One-Shot Imitation LearningImitation
4Constrained Policy OptimizationLocomotion

3.3 高影响力但尚无深笔记(待补 terse/P-*)

4. 引用图健康

未解析标签按 kb_common.classify 二分类:

4.1 真实漏收录(suspected_missing,按出现次数 Top 20)

出现次数标签分类
3PULSEsuspected_missing
2Allegro Handsuspected_missing
2CPOsuspected_missing
2DALL-E 2suspected_missing
2HM3Dsuspected_missing
2LIBERO-PROsuspected_missing
2Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrationssuspected_missing
2VPPsuspected_missing
2YOLOv5suspected_missing
13D Diffuser Actorsuspected_missing
16-DOF AUV controlsuspected_missing
1Action chunking (ACT)suspected_missing
1Action chunking (ACT) for reactive robot controlsuspected_missing
1Agent57suspected_missing
1AlphaStarsuspected_missing
1Ambient Diffusion (theory)suspected_missing
1AnyRotatesuspected_missing
1AnySkinsuspected_missing
1Apple Vision Pro hand trackingsuspected_missing
1Asymmetric actor-critic PPO for dexterous manipulationsuspected_missing

4.2 概念标签(concept,按出现次数 Top 20;非漏收录)

出现次数标签分类
26Behavior Cloning (BC)concept
14VLAconcept
14deep RLconcept
13reinforcement learningconcept
12deep learningconcept
7Diffusion modelsconcept
7domain randomizationconcept
7language modelsconcept
7sim2realconcept
7tactile sensingconcept
6BCconcept
6ExBody2: Advanced Expressive Humanoid Whole-Body Controlconcept
6VLA modelsconcept
6video diffusionconcept
5ANYmal-blind-locomotionconcept
5Contrastive learningconcept
5Hierarchical ILconcept
5MPCconcept
5actor-criticconcept
4Policy Gradientconcept

5. 跨方向引用(交叉学习机会)

行方向 → 列方向:行方向的论文有多少 builds_on 指向列方向的论文。高数值 = 两个方向深度耦合 = 交叉学习富矿

行 \ 列Action GenDexterousFoundationImitationLocomotionMotion PriorTeleop
Action Gen··4721382
Dexterous4·2846·7
Foundation174·3313·2
Imitation1288·313
Locomotion28·93·82
Motion Prior62·16533··
Teleop20·18103··

最耦合的方向对(双向合计):

6. 覆盖度

6.1 按方向

方向数量占比
Foundation33534%
Motion Prior13214%
Imitation13214%
Locomotion11011%
Action Gen10811%
Dexterous10511%
Teleop535%

6.2 按年份(近 8 年)

年份数量
2026316
202586
202499
2023114
202278
202142
202041
201937

7. 查漏补缺(reference_expansion/ 汇总)

8. 关键指标

关键指标的操作化仪表盘。前 4 项是元数据达标核心,达标状态实时计算:3/4 达标(✓=达标 / ✗=未达)。
浅入口覆盖已接入实测(扫 site/content 已发布教学页);
后 2 项(高引收录率 / 事实 grounding 率)当前未接入实测
标「未接入」,是否达标列记「—」——不冒充绿

指标操作定义当前值目标是否达标
枢纽 why_important 完整率枢纽(被引≥3)有 why_important_zh 占比133/133 = 100.0%100%
枢纽结构化字段完整率枢纽有 contribution_type+paper_type+reproducibility+primary_thread 占比111/133 = 83.5%≥80%
引用解析率resolved/(resolved+真实漏收录);概念标签不计分母(与 validate_citations 门同源)2238/2489 = 89.9%(原始 2238/3742 = 59.8%;概念 1253,真实漏收录 251)≥90%
孤立节点率无 builds_on 且无 leads_to 论文占比(与 §4 同定义)8/975 = 0.8%<10%
浅入口覆盖枢纽(被引≥3)有已发布入门教学页占比65/133 = 48.9%100%
高引工作收录率近十年领域高引收录占比未接入(需外部高引榜)≥95%
事实 grounding 率断言可回链 KB/PDF/权威源 或标 unconfirmed 占比未接入(需 cite 提取管线)≥95%(余显式 unconfirmed)

读法:✓/✗ 由当前值对照「目标」列实时计算(不写死);「—」= 该指标尚未接入实测,不计入达标计数。

9. 研究群落与桥接工作

在引用图(citations_normalized.jsonl,1803 条 builds_on/leads_to 边)上跑 label-propagation 社区检测 + 模块度合并,把密集互引的论文聚成研究群落;再识别桥接工作——邻居跨越 ≥2 个群落的论文,即「交叉学习」的关键节点。由 scripts/cluster_analysis.py 产出 clusters.jsonl + bridges.jsonl,本节读其结果。

群落主表:5 个群落,覆盖库内 754 篇论文 (其余为孤立/小群落,<5 篇不计入主表)。群落编号 = 按规模降序。

#规模主导方向主导子方向主题(启发式)Top-3 代表论文(PageRank)
C0265FoundationVLA视觉-语言-动作模型(VLA)· RT-2: Vision-Language-Action Models Transfer Web Knowledge t(pr=0.0051)
· Open X-Embodiment: Robotic Learning Datasets and RT-X Models(pr=0.0045)
· OpenVLA: An Open-Source Vision-Language-Action Model(pr=0.0044)
C1202Motion Priorhumanoid_rl人形机器人 RL / 扩散模型 / 动作跟踪· AMP: Adversarial Motion Priors for Stylized Physics-Based Ch(pr=0.0107)
· BeyondMimic: From Motion Tracking to Versatile Humanoid Cont(pr=0.0075)
· ASE: Large-Scale Reusable Adversarial Skill Embeddings for P(pr=0.0061)
C2134Foundationworld_model世界模型 / 极限地形/parkour / locomotion· Extreme Parkour with Legged Robots(pr=0.0070)
· Parkour in the Wild: Learning a General and Extensible Agile(pr=0.0039)
· World Action Models: The Next Frontier in Embodied AI(pr=0.0032)
C387Action Genretarget动作重定向/映射 / 扩散模型· MaskedMimic: Unified Physics-Based Character Control Through(pr=0.0054)
· Human Motion Diffusion Model (MDM)(pr=0.0050)
· MotionGPT: Human Motion as a Foreign Language(pr=0.0037)
C466Dexterousin_hand_rl灵巧手内操作· Learning Dexterous In-Hand Manipulation(pr=0.0036)
· Learning Robust Dexterous In-Hand Manipulation from Vision a(pr=0.0030)
· DexImit: Learning Bimanual Dexterous Manipulation from Monoc(pr=0.0026)

桥接工作 Top-15(按桥接度 = 邻居群落数 span + 0.1×跨群落邻居数 cross 排序):

桥接度spancrossPageRank工作方向所属群落
4.9490.0009Proximal Policy Optimization Algorithms (PPO)LocomotionC1
4.5450.0015RMA: Rapid Motor Adaptation for Legged RobotsLocomotionC2
4.3430.0006Isaac Gym: High Performance GPU-Based Physics Simulation ForLocomotionC2
4.03100.0014AMASS: Archive of Motion Capture as Surface ShapesMotion PriorC3
3.9390.0009GR00T N1: An Open Foundation Model for Generalist Humanoid RFoundationC1
3.6360.0044OpenVLA: An Open-Source Vision-Language-Action ModelFoundationC0
3.6360.0006World ModelsFoundationC2
3.5350.0026DeepMimic: Example-Guided Deep Reinforcement Learning of PhyLocomotionC1
3.4340.0107AMP: Adversarial Motion Priors for Stylized Physics-Based ChLocomotionC1
3.4340.0028Cosmos World Foundation Model Platform for Physical AIFoundationC2
3.4340.0008Language Models are Few-Shot Learners (GPT-3)FoundationC2
3.4340.0006Auto-Encoding Variational Bayes (VAE)FoundationC1
3.3330.0039Learning Fine-Grained Bimanual Manipulation with Low-Cost HaTeleopC0
3.3330.0032GR00T N1.x (NVIDIA humanoid VLA foundation model, 2026 iteraImitationC1
3.3330.0019π0: A Vision-Language-Action Flow Model for General Robot CoFoundationC0

读法:群落编号 C0–C4 与上表一致;span = 该论文邻居所属的不同群落数(含自身群落,≥2 即桥接);cross = 跨群落的邻居数;桥接度 = span + 0.1×cross,越高越关键。桥接工作通常是 PPO / Transformer / AMASS / MuJoCo / OpenVLA / Isaac Gym 等被多个方向共同引用的基础工作——它们正是「交叉学习富矿」最值得深挖的节点。